Skip to content

fix(build-vllm): let a dispatch raise the QA selection floors - #270

Closed
wbrennan899 wants to merge 1 commit into
vast-ai:mainfrom
wbrennan899:qa-compute-cap-override
Closed

wbrennan899 wants to merge 1 commit into
vast-ai:mainfrom
wbrennan899:qa-compute-cap-override

Conversation

@wbrennan899

Copy link
Copy Markdown
Contributor

Problem

external/vllm/templates/vllm-qa/template.yml declares the generic vLLM selection floor:

extra_filters:
  compute_cap:
    gte: 800        # sm_80 — matches the production vLLM floor (vLLM kernels)

That is correct for almost every vLLM build. It is wrong for one whose kernels are compiled for a newer architecture, because offers are searched cheapest-first: the gate rents an Ampere card, and the engine dies before serving.

torch.AcceleratorError: CUDA error: no kernel image is available for execution on the device
RuntimeError: Engine core initialization failed.

The redraw then makes it worse rather than better. It excludes the failed machine and draws again from the same pool, so it lands on another card the image has no kernels for and fails identically. Per ADR 0029 the discriminator between an image defect and a bad host is reproducibility — "an image defect fails every draw" — and this failure is reproducible for a reason that has nothing to do with the image. The gate reports the build broken without ever having run it on hardware it supports.

Measured

hy4-preview, run 33509520098 — a Blackwell-targeted model:

Draw Offer GPU Compute cap Price
1 47574078 A10 860 $0.243/hr
2 49411012 RTX 3080 860 $0.123/hr
3 31921484 RTX 4000 Ada 890 $0.164/hr

Same error on every draw, on both CUDA variants. The rest of the suite was green — 24 passed, 2 failed, 3 skipped, and the only failures were vllm.d/10-vllm-serving and vllm.d/12-vllm-contract. base/60-gpu-cuda, base/61-cuda-compute and base/62-gpu-libraries all passed, because the driver and CUDA userland were fine; it was only ever the kernels.

Change

qa-gate.yml already has the input for this — set_filters, raise-only, rejected by create.py if an override would widen selection. promote-base-image.yml, promote-pytorch.yml and build-llama-cpp.yml all use it for their per-config driver floors. It was just never plumbed through build-vllm.yml, so a vLLM dispatch had no way to say "this build needs Blackwell".

Two new workflow_dispatch inputs, passed to both QA cells:

QA_SET_FILTERS:   # e.g. compute_cap.gte=1000
QA_MAX_PRICE:     # e.g. 12.00

Both cells, not one. A floor the standard cell honours and the serverless cell ignores would leave every run of such an image carrying one red cell that says nothing about the image — the same misreading this PR exists to remove, just narrower.

max_price is not an independent knob. Raising compute_cap can empty the offer pool, and the gate's response to an empty pool is to wait and redraw. Without a matching ceiling the override trades a wrong-GPU red for a run that spends its retries on no offers matched the floors and reports no verdict at all.

Blast radius

None by default. Scheduled runs, and dispatches that set neither input, pass an empty set_filters and the same 2.00 ceiling qa-gate.yml already defaulted to. set_filters is raise-only, so a dispatch cannot loosen the linted floors even deliberately.

imagegen lint --all: 27 images | 0 errors, baseline clean.

Using it

build-vllm.yml
  VLLM_VERSION=nightly
  CUSTOM_IMAGE_TAG=hy4-preview
  QA_SET_FILTERS=compute_cap.gte=1000
  QA_MAX_PRICE=<ceiling that reaches a Blackwell card>

Two things this PR deliberately does not decide:

  • The right floor for hy4. 1000 is sm_100; if its kernels are sm_120-only then an RTX 5090 is required and 1200 is the floor. Worth confirming against the upstream build before the re-run.
  • The right ceiling. A 5090 fits comfortably under $2/hr; a B200 does not. If hy4 turns out to need sm_100 specifically, the ceiling has to clear datacentre Blackwell pricing or the pool is empty.

Neither is knowable from the failed run, since it never reached a card in that class.

Not addressed here

vastai/vllm:hy4-preview-cuda-13.0 and -cuda-12.9 are published and currently unverified — the manifests merged from a run whose QA never tested them meaningfully. They are not known-bad; they are untested. Re-running the gate with a real floor is what settles that, and that is a dispatch rather than a code change.

The vLLM QA template declares compute_cap.gte=800 — the generic vLLM
floor, sm_80. Offers are searched cheapest-first, so for a model whose
kernels are built for a newer arch the gate rents an Ampere card and the
engine dies on startup:

    torch.AcceleratorError: CUDA error: no kernel image is available
    for execution on the device

It then redraws, excluding the failed machine, and draws another cheap
card from the same pool — so it fails identically every time. That is the
signature the redraw rule uses to separate a real defect from a bad host
(ADR 0029), and here it produces exactly the wrong reading: the image is
reported broken having never once run on a GPU it has kernels for.

Measured on hy4-preview (run 33509520098): three draws, A10 / RTX 3080 /
RTX 4000 Ada — cc 860, 860, 890 — same error each time, both CUDA
variants, while every non-vLLM test in the suite passed.

qa-gate already takes set_filters for this, raise-only and already used
by the base, pytorch and llama-cpp gates for their per-config driver
floors. It was simply never plumbed through here, so a vLLM dispatch had
no way to say "this build needs Blackwell". Both cells get it, and they
are deliberately not forked: a floor the standard cell honours and the
serverless cell ignores would leave one red cell per run that means
nothing about the image.

max_price comes along because it is not independent of the change.
Raising compute_cap can empty the offer pool, and the gate answers an
empty pool by waiting and redrawing — so without a matching ceiling the
override converts a wrong-GPU red into a run that spends its retries on
"no offers matched the floors" and still reports no verdict.

Defaults reproduce current behaviour exactly: scheduled runs and
dispatches that set neither input pass an empty set_filters and the same
2.00 ceiling the gate already defaulted to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@robballantyne

Copy link
Copy Markdown
Collaborator

Thanks, Will. The need here is real, so we've taken this over in #297 rather than bounce review rounds. Your commit is the first of two there, authorship kept. The second scopes the override to custom-tag builds and compute_cap only, bounds and validates max_price (it was spliced into a run: script, and 0 removed the cap), and makes a narrowed QA pass visible in the summary and Slack headline. The reasoning is in ADR 0047, and a new lint rule (L099) enforces the wiring. Closing in favour of #297.

robballantyne added a commit that referenced this pull request Sep 24, 2026
… bounded price - on every custom-tag QA workflow (ADR 0047, L099, L100) (#297)

* fix(build-vllm): let a dispatch raise the QA selection floors

The vLLM QA template declares compute_cap.gte=800 — the generic vLLM
floor, sm_80. Offers are searched cheapest-first, so for a model whose
kernels are built for a newer arch the gate rents an Ampere card and the
engine dies on startup:

    torch.AcceleratorError: CUDA error: no kernel image is available
    for execution on the device

It then redraws, excluding the failed machine, and draws another cheap
card from the same pool — so it fails identically every time. That is the
signature the redraw rule uses to separate a real defect from a bad host
(ADR 0029), and here it produces exactly the wrong reading: the image is
reported broken having never once run on a GPU it has kernels for.

Measured on hy4-preview (run 33509520098): three draws, A10 / RTX 3080 /
RTX 4000 Ada — cc 860, 860, 890 — same error each time, both CUDA
variants, while every non-vLLM test in the suite passed.

qa-gate already takes set_filters for this, raise-only and already used
by the base, pytorch and llama-cpp gates for their per-config driver
floors. It was simply never plumbed through here, so a vLLM dispatch had
no way to say "this build needs Blackwell". Both cells get it, and they
are deliberately not forked: a floor the standard cell honours and the
serverless cell ignores would leave one red cell per run that means
nothing about the image.

max_price comes along because it is not independent of the change.
Raising compute_cap can empty the offer pool, and the gate answers an
empty pool by waiting and redrawing — so without a matching ceiling the
override converts a wrong-GPU red into a run that spends its retries on
"no offers matched the floors" and still reports no verdict.

Defaults reproduce current behaviour exactly: scheduled runs and
dispatches that set neither input pass an empty set_filters and the same
2.00 ceiling the gate already defaulted to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(ci): a dispatch narrows QA only for a custom tag, visibly, with a bounded price (ADR 0047, L099)

The previous commit (from PR #270) plumbed QA_SET_FILTERS / QA_MAX_PRICE
straight from dispatch inputs into both vLLM QA cells. A design review found
the need real and the default path inert, but four problems with the wiring.
ADR 0047 records the decision; this makes it hold.

- Route around QA: nothing tied the override to a custom tag, so a mainline
  version or nightly could be certified on Blackwell only and published to
  tags rented under the sm_80 floor. create.py's raise-only check stops QA
  widening, not narrowing. Preflight now refuses an override without
  CUSTOM_IMAGE_TAG, and any key but compute_cap.gte/lte (a new key such as
  machine_id is merged unchecked).
- Invisible: a narrowed pass read "vLLM promoted - live-GPU QA passed". The
  step summary and a prefix on the Slack headline now name the floors and
  ceiling; an ordinary run renders byte-identically.
- max_price: free text spliced into a run: script holding VAST_API_KEY; 0
  removed the cap (`if args.max_price:`); a malformed value exited 2, read as
  "no offers". Preflight bounds it to (0, QA_MAX_PRICE_HARD_MAX = 15.00];
  qa-gate refuses a non-positive price for every caller and passes it via
  env:; the client's --max-price refuses 0, negatives, nan and inf.
- Guidance: offers are ranked smallest-VRAM then lowest compute_cap, not
  cheapest-first, so compute_cap.gte=1000 alone draws a consumer sm_120 card
  before any sm_100 one. Comments and input text now say so and give the
  range form.

The QA cells read preflight's validated outputs. L099 fails any qa-gate
caller that feeds set_filters or max_price from a dispatch input directly;
mutated on the real workflow (restoring the original wiring fires on both
cells). test_qa_floor_override.py executes the preflight step for every
refusal; each guard was mutated and its test goes red.

* feat(ci): make the validated QA override the pattern for every custom-tag QA workflow (ADR 0047, L100)

The QA override was vLLM-only. The need is not: any image built under a custom
tag for a newer architecture fails every draw on its template floor, as
hy4-preview did. A hatch that exists in one workflow gets copied by hand into
the next and drifts, so this makes it one shared thing and requires it.

- One composite action, .github/actions/validate-qa-override, now holds the
  rules and the hard maximum (15.00). build-vllm's inline step is replaced by
  it.
- Rolled out to every build workflow with a QA cell and a CUSTOM_IMAGE_TAG
  input: sglang, llama-cpp, vllm-omni, comfyui, unsloth-studio, aio-studio,
  aio-studio-base (plus vllm). The action runs in the job the QA cells already
  need (preflight; resolve-refs on aio-studio; build on aio-studio-base).
  llama-cpp adds the range on top of its committed cuda_max_good floor.
- notify-slack.yml gains a qa-override-note input and prefixes it to the
  header, the caller's headline or the default, so unsloth-studio (no
  headline) is covered. build-vllm's format() wrapper is reverted.
- L100 requires the full wiring in every in-scope workflow; mutated arm by
  arm on six real workflows. The promotion gates are out of scope: no custom
  tag, mainline only, where condition 1 refuses every override.
- imagegen new scaffolds the wiring; test_generate asserts a fresh scaffold
  passes L099 and L100 untouched.
- ADR 0047 amended; invariants.md section; lint-rules.md regenerated.

---------

Co-authored-by: Will <wbrennan899@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants