Skip to content

Add regulated profile for deterministic behavior - #98

Merged
csvance merged 2 commits into
mainfrom
regulated-profile
Oct 8, 2026
Merged

csvance merged 2 commits into
mainfrom
regulated-profile

Conversation

@csvance

@csvance csvance commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Sources of non-determinism:

  1. autotuning
  2. different batch size programs - handled by always padding to the largest batch size
  3. --xla_gpu_exclude_nondeterministic_ops

Add a new "regulated" profile which handles 2 and 3 automatically, and also asserts you specify the expected precision (ie the server fails to start if you declare tf32 but its not available or f32 and vice versa)

autotuning can be made deterministic via a persistent autotuning cache

csvance and others added 2 commits October 8, 2026 03:48
…flags

Validated deployments need a row's result to be independent of how many
other requests were coalesced with it. With several compiled batch sizes the
scheduler picks a different program depending on queue depth, so the same
input can run through different kernels. runtime.batch_sizes: largest loads
only the largest declared size per input-shape variant (the others are not
read and need not exist); every dispatch then runs that one program, padded
with zero rows, at the cost of full-batch compute on every dispatch.

runtime.xla_flags passes XLA DebugOptions overrides to every compile. Names
and value types are checked against the linked XLA's proto at startup so a
typo fails loudly instead of being silently ignored, and the flags join the
executable cache key (only when set, so existing cache entries keep their
keys).

runtime.profile: regulated bundles these as defaults (largest batch only,
xla_gpu_exclude_nondeterministic_ops) while anything set explicitly wins. It
deliberately does not force numerics: f32; instead it logs a bannered warning
naming the matmul/conv precision actually in effect, using the TF32 probe,
since only f32 is invariant across GPU generations.

The executable cache used to sweep the programs of any module a worker did
not read. A largest-only worker sharing a bundle directory with a normal one
would have deleted the smaller sizes' programs on every start, so declared
but unread modules now keep their cache records.

Also corrects bundles.md, which claimed a request's batch size must equal a
compiled size; coalescable models pad up to the next one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eport

numerics=tf32 only checked that the GPU could run TF32 (compute capability
>= 8.0). The startup probe that shows whether TF32 is actually used was
informational, so a capable worker with NVIDIA_TF32_OVERRIDE=0 in its
environment, or a stack whose kernel choice stopped using TF32, would start
and serve full-f32 results to a deployment validated on TF32.

Under tf32 the probe now compiles through the configured pool and fails
startup unless the TF32 signature is observed (or if the probe cannot run),
and NVIDIA_TF32_OVERRIDE=0 is refused before the client is created so the
operator sees the actual cause.

The regulated-profile startup report assumed f32 was the goal and told a
tf32 deployment to switch. f32 and tf32 are now both attested modes and get
a bannered info line naming the precision; only auto, which guarantees
neither, keeps the bannered warning.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@csvance
csvance merged commit 842bb76 into main Oct 8, 2026
11 checks passed
@csvance
csvance deleted the regulated-profile branch October 8, 2026 05:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant