fix(build-vllm): let a dispatch raise the QA selection floors - #270
Closed
wbrennan899 wants to merge 1 commit into
Closed
wbrennan899 wants to merge 1 commit into
wbrennan899 wants to merge 1 commit into
Conversation
The vLLM QA template declares compute_cap.gte=800 — the generic vLLM
floor, sm_80. Offers are searched cheapest-first, so for a model whose
kernels are built for a newer arch the gate rents an Ampere card and the
engine dies on startup:
torch.AcceleratorError: CUDA error: no kernel image is available
for execution on the device
It then redraws, excluding the failed machine, and draws another cheap
card from the same pool — so it fails identically every time. That is the
signature the redraw rule uses to separate a real defect from a bad host
(ADR 0029), and here it produces exactly the wrong reading: the image is
reported broken having never once run on a GPU it has kernels for.
Measured on hy4-preview (run 33509520098): three draws, A10 / RTX 3080 /
RTX 4000 Ada — cc 860, 860, 890 — same error each time, both CUDA
variants, while every non-vLLM test in the suite passed.
qa-gate already takes set_filters for this, raise-only and already used
by the base, pytorch and llama-cpp gates for their per-config driver
floors. It was simply never plumbed through here, so a vLLM dispatch had
no way to say "this build needs Blackwell". Both cells get it, and they
are deliberately not forked: a floor the standard cell honours and the
serverless cell ignores would leave one red cell per run that means
nothing about the image.
max_price comes along because it is not independent of the change.
Raising compute_cap can empty the offer pool, and the gate answers an
empty pool by waiting and redrawing — so without a matching ceiling the
override converts a wrong-GPU red into a run that spends its retries on
"no offers matched the floors" and still reports no verdict.
Defaults reproduce current behaviour exactly: scheduled runs and
dispatches that set neither input pass an empty set_filters and the same
2.00 ceiling the gate already defaulted to.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Collaborator
|
Thanks, Will. The need here is real, so we've taken this over in #297 rather than bounce review rounds. Your commit is the first of two there, authorship kept. The second scopes the override to custom-tag builds and |
robballantyne
added a commit
that referenced
this pull request
Sep 24, 2026
… bounded price - on every custom-tag QA workflow (ADR 0047, L099, L100) (#297) * fix(build-vllm): let a dispatch raise the QA selection floors The vLLM QA template declares compute_cap.gte=800 — the generic vLLM floor, sm_80. Offers are searched cheapest-first, so for a model whose kernels are built for a newer arch the gate rents an Ampere card and the engine dies on startup: torch.AcceleratorError: CUDA error: no kernel image is available for execution on the device It then redraws, excluding the failed machine, and draws another cheap card from the same pool — so it fails identically every time. That is the signature the redraw rule uses to separate a real defect from a bad host (ADR 0029), and here it produces exactly the wrong reading: the image is reported broken having never once run on a GPU it has kernels for. Measured on hy4-preview (run 33509520098): three draws, A10 / RTX 3080 / RTX 4000 Ada — cc 860, 860, 890 — same error each time, both CUDA variants, while every non-vLLM test in the suite passed. qa-gate already takes set_filters for this, raise-only and already used by the base, pytorch and llama-cpp gates for their per-config driver floors. It was simply never plumbed through here, so a vLLM dispatch had no way to say "this build needs Blackwell". Both cells get it, and they are deliberately not forked: a floor the standard cell honours and the serverless cell ignores would leave one red cell per run that means nothing about the image. max_price comes along because it is not independent of the change. Raising compute_cap can empty the offer pool, and the gate answers an empty pool by waiting and redrawing — so without a matching ceiling the override converts a wrong-GPU red into a run that spends its retries on "no offers matched the floors" and still reports no verdict. Defaults reproduce current behaviour exactly: scheduled runs and dispatches that set neither input pass an empty set_filters and the same 2.00 ceiling the gate already defaulted to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): a dispatch narrows QA only for a custom tag, visibly, with a bounded price (ADR 0047, L099) The previous commit (from PR #270) plumbed QA_SET_FILTERS / QA_MAX_PRICE straight from dispatch inputs into both vLLM QA cells. A design review found the need real and the default path inert, but four problems with the wiring. ADR 0047 records the decision; this makes it hold. - Route around QA: nothing tied the override to a custom tag, so a mainline version or nightly could be certified on Blackwell only and published to tags rented under the sm_80 floor. create.py's raise-only check stops QA widening, not narrowing. Preflight now refuses an override without CUSTOM_IMAGE_TAG, and any key but compute_cap.gte/lte (a new key such as machine_id is merged unchecked). - Invisible: a narrowed pass read "vLLM promoted - live-GPU QA passed". The step summary and a prefix on the Slack headline now name the floors and ceiling; an ordinary run renders byte-identically. - max_price: free text spliced into a run: script holding VAST_API_KEY; 0 removed the cap (`if args.max_price:`); a malformed value exited 2, read as "no offers". Preflight bounds it to (0, QA_MAX_PRICE_HARD_MAX = 15.00]; qa-gate refuses a non-positive price for every caller and passes it via env:; the client's --max-price refuses 0, negatives, nan and inf. - Guidance: offers are ranked smallest-VRAM then lowest compute_cap, not cheapest-first, so compute_cap.gte=1000 alone draws a consumer sm_120 card before any sm_100 one. Comments and input text now say so and give the range form. The QA cells read preflight's validated outputs. L099 fails any qa-gate caller that feeds set_filters or max_price from a dispatch input directly; mutated on the real workflow (restoring the original wiring fires on both cells). test_qa_floor_override.py executes the preflight step for every refusal; each guard was mutated and its test goes red. * feat(ci): make the validated QA override the pattern for every custom-tag QA workflow (ADR 0047, L100) The QA override was vLLM-only. The need is not: any image built under a custom tag for a newer architecture fails every draw on its template floor, as hy4-preview did. A hatch that exists in one workflow gets copied by hand into the next and drifts, so this makes it one shared thing and requires it. - One composite action, .github/actions/validate-qa-override, now holds the rules and the hard maximum (15.00). build-vllm's inline step is replaced by it. - Rolled out to every build workflow with a QA cell and a CUSTOM_IMAGE_TAG input: sglang, llama-cpp, vllm-omni, comfyui, unsloth-studio, aio-studio, aio-studio-base (plus vllm). The action runs in the job the QA cells already need (preflight; resolve-refs on aio-studio; build on aio-studio-base). llama-cpp adds the range on top of its committed cuda_max_good floor. - notify-slack.yml gains a qa-override-note input and prefixes it to the header, the caller's headline or the default, so unsloth-studio (no headline) is covered. build-vllm's format() wrapper is reverted. - L100 requires the full wiring in every in-scope workflow; mutated arm by arm on six real workflows. The promotion gates are out of scope: no custom tag, mainline only, where condition 1 refuses every override. - imagegen new scaffolds the wiring; test_generate asserts a fresh scaffold passes L099 and L100 untouched. - ADR 0047 amended; invariants.md section; lint-rules.md regenerated. --------- Co-authored-by: Will <wbrennan899@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
external/vllm/templates/vllm-qa/template.ymldeclares the generic vLLM selection floor:That is correct for almost every vLLM build. It is wrong for one whose kernels are compiled for a newer architecture, because offers are searched cheapest-first: the gate rents an Ampere card, and the engine dies before serving.
The redraw then makes it worse rather than better. It excludes the failed machine and draws again from the same pool, so it lands on another card the image has no kernels for and fails identically. Per ADR 0029 the discriminator between an image defect and a bad host is reproducibility — "an image defect fails every draw" — and this failure is reproducible for a reason that has nothing to do with the image. The gate reports the build broken without ever having run it on hardware it supports.
Measured
hy4-preview, run 33509520098 — a Blackwell-targeted model:Same error on every draw, on both CUDA variants. The rest of the suite was green —
24 passed, 2 failed, 3 skipped, and the only failures werevllm.d/10-vllm-servingandvllm.d/12-vllm-contract.base/60-gpu-cuda,base/61-cuda-computeandbase/62-gpu-librariesall passed, because the driver and CUDA userland were fine; it was only ever the kernels.Change
qa-gate.ymlalready has the input for this —set_filters, raise-only, rejected bycreate.pyif an override would widen selection.promote-base-image.yml,promote-pytorch.ymlandbuild-llama-cpp.ymlall use it for their per-config driver floors. It was just never plumbed throughbuild-vllm.yml, so a vLLM dispatch had no way to say "this build needs Blackwell".Two new
workflow_dispatchinputs, passed to both QA cells:Both cells, not one. A floor the standard cell honours and the serverless cell ignores would leave every run of such an image carrying one red cell that says nothing about the image — the same misreading this PR exists to remove, just narrower.
max_priceis not an independent knob. Raisingcompute_capcan empty the offer pool, and the gate's response to an empty pool is to wait and redraw. Without a matching ceiling the override trades a wrong-GPU red for a run that spends its retries onno offers matched the floorsand reports no verdict at all.Blast radius
None by default. Scheduled runs, and dispatches that set neither input, pass an empty
set_filtersand the same2.00ceilingqa-gate.ymlalready defaulted to.set_filtersis raise-only, so a dispatch cannot loosen the linted floors even deliberately.imagegen lint --all:27 images | 0 errors, baseline clean.Using it
Two things this PR deliberately does not decide:
1000is sm_100; if its kernels are sm_120-only then an RTX 5090 is required and1200is the floor. Worth confirming against the upstream build before the re-run.Neither is knowable from the failed run, since it never reached a card in that class.
Not addressed here
vastai/vllm:hy4-preview-cuda-13.0and-cuda-12.9are published and currently unverified — the manifests merged from a run whose QA never tested them meaningfully. They are not known-bad; they are untested. Re-running the gate with a real floor is what settles that, and that is a dispatch rather than a code change.