Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .cursor/rules/python.mdc
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,10 @@ Python 3.11+ async AI benchmarking tool for measuring LLM inference server perfo

Numeric metric values crossing a serialization boundary or feeding a numerical algorithm must be finite or explicitly `None`. Use `aiperf.common.finite` (`FiniteFloat`, `scrub_non_finite`, `nan_safe_mean`/`std`, `is_finite_value`). Mechanical CI invariants in `tests/unit/property/test_finite_invariants.py` reject new violations and ratchet existing debt to zero via baseline files. See [`docs/dev/patterns.md`](docs/dev/patterns.md) § "NaN/Inf Discipline Pattern" and [`docs/dev/global-invariants.md`](docs/dev/global-invariants.md) for the full contract.

## CLI Flag Routing Under `--config`

A CLI flag passed alongside `--config` MUST either change the resolved config or raise an error naming the flag — never be silently ignored, which for a benchmarking tool means publishing numbers that do not match what was asked for. `CLIConfig` is the source of truth, structurally: `resolve_config` receives a `CLIConfig` and a path and nothing else, so a flag reaches `AIPerfConfig` only by being a field on it. Every field must be classified in `src/aiperf/config/flags/_config_flag_routing.py` as `ROUTED_UNDER_CONFIG`, `UNROUTED_UNDER_CONFIG`, `EXEMPT_FROM_CONFIG_ROUTING`, or one of the conditional sets (`COMPANION_ROUTED`, `MAGIC_LIST_ONLY_UNDER_CONFIG`). Adding a flag without classifying it fails `test_every_cli_config_field_is_classified`; claiming a flag is routed when nothing routes it fails `test_routed_field_never_silently_no_ops`. Prefer routing the flag (check whether a builder already exists) over listing it as unrouted. See [`docs/dev/global-invariants.md`](docs/dev/global-invariants.md) § "CLI flag routing under `--config`".

## Build and Test Commands

```bash
Expand Down
4 changes: 4 additions & 0 deletions .github/copilot-instructions.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,10 @@ Python 3.11+ async AI benchmarking tool for measuring LLM inference server perfo

Numeric metric values crossing a serialization boundary or feeding a numerical algorithm must be finite or explicitly `None`. Use `aiperf.common.finite` (`FiniteFloat`, `scrub_non_finite`, `nan_safe_mean`/`std`, `is_finite_value`). Mechanical CI invariants in `tests/unit/property/test_finite_invariants.py` reject new violations and ratchet existing debt to zero via baseline files. See [`docs/dev/patterns.md`](docs/dev/patterns.md) § "NaN/Inf Discipline Pattern" and [`docs/dev/global-invariants.md`](docs/dev/global-invariants.md) for the full contract.

## CLI Flag Routing Under `--config`

A CLI flag passed alongside `--config` MUST either change the resolved config or raise an error naming the flag — never be silently ignored, which for a benchmarking tool means publishing numbers that do not match what was asked for. `CLIConfig` is the source of truth, structurally: `resolve_config` receives a `CLIConfig` and a path and nothing else, so a flag reaches `AIPerfConfig` only by being a field on it. Every field must be classified in `src/aiperf/config/flags/_config_flag_routing.py` as `ROUTED_UNDER_CONFIG`, `UNROUTED_UNDER_CONFIG`, `EXEMPT_FROM_CONFIG_ROUTING`, or one of the conditional sets (`COMPANION_ROUTED`, `MAGIC_LIST_ONLY_UNDER_CONFIG`). Adding a flag without classifying it fails `test_every_cli_config_field_is_classified`; claiming a flag is routed when nothing routes it fails `test_routed_field_never_silently_no_ops`. Prefer routing the flag (check whether a builder already exists) over listing it as unrouted. See [`docs/dev/global-invariants.md`](docs/dev/global-invariants.md) § "CLI flag routing under `--config`".

## Build and Test Commands

```bash
Expand Down
4 changes: 4 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,10 @@ Python 3.11+ async AI benchmarking tool for measuring LLM inference server perfo

Numeric metric values crossing a serialization boundary or feeding a numerical algorithm must be finite or explicitly `None`. Use `aiperf.common.finite` (`FiniteFloat`, `scrub_non_finite`, `nan_safe_mean`/`std`, `is_finite_value`). Mechanical CI invariants in `tests/unit/property/test_finite_invariants.py` reject new violations and ratchet existing debt to zero via baseline files. See [`docs/dev/patterns.md`](docs/dev/patterns.md) § "NaN/Inf Discipline Pattern" and [`docs/dev/global-invariants.md`](docs/dev/global-invariants.md) for the full contract.

## CLI Flag Routing Under `--config`

A CLI flag passed alongside `--config` MUST either change the resolved config or raise an error naming the flag — never be silently ignored, which for a benchmarking tool means publishing numbers that do not match what was asked for. `CLIConfig` is the source of truth, structurally: `resolve_config` receives a `CLIConfig` and a path and nothing else, so a flag reaches `AIPerfConfig` only by being a field on it. Every field must be classified in `src/aiperf/config/flags/_config_flag_routing.py` as `ROUTED_UNDER_CONFIG`, `UNROUTED_UNDER_CONFIG`, `EXEMPT_FROM_CONFIG_ROUTING`, or one of the conditional sets (`COMPANION_ROUTED`, `MAGIC_LIST_ONLY_UNDER_CONFIG`). Adding a flag without classifying it fails `test_every_cli_config_field_is_classified`; claiming a flag is routed when nothing routes it fails `test_routed_field_never_silently_no_ops`. Prefer routing the flag (check whether a builder already exists) over listing it as unrouted. See [`docs/dev/global-invariants.md`](docs/dev/global-invariants.md) § "CLI flag routing under `--config`".

## Build and Test Commands

```bash
Expand Down
4 changes: 4 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,10 @@ Python 3.11+ async AI benchmarking tool for measuring LLM inference server perfo

Numeric metric values crossing a serialization boundary or feeding a numerical algorithm must be finite or explicitly `None`. Use `aiperf.common.finite` (`FiniteFloat`, `scrub_non_finite`, `nan_safe_mean`/`std`, `is_finite_value`). Mechanical CI invariants in `tests/unit/property/test_finite_invariants.py` reject new violations and ratchet existing debt to zero via baseline files. See [`docs/dev/patterns.md`](docs/dev/patterns.md) § "NaN/Inf Discipline Pattern" and [`docs/dev/global-invariants.md`](docs/dev/global-invariants.md) for the full contract.

## CLI Flag Routing Under `--config`

A CLI flag passed alongside `--config` MUST either change the resolved config or raise an error naming the flag — never be silently ignored, which for a benchmarking tool means publishing numbers that do not match what was asked for. `CLIConfig` is the source of truth, structurally: `resolve_config` receives a `CLIConfig` and a path and nothing else, so a flag reaches `AIPerfConfig` only by being a field on it. Every field must be classified in `src/aiperf/config/flags/_config_flag_routing.py` as `ROUTED_UNDER_CONFIG`, `UNROUTED_UNDER_CONFIG`, `EXEMPT_FROM_CONFIG_ROUTING`, or one of the conditional sets (`COMPANION_ROUTED`, `MAGIC_LIST_ONLY_UNDER_CONFIG`). Adding a flag without classifying it fails `test_every_cli_config_field_is_classified`; claiming a flag is routed when nothing routes it fails `test_routed_field_never_silently_no_ops`. Prefer routing the flag (check whether a builder already exists) over listing it as unrouted. See [`docs/dev/global-invariants.md`](docs/dev/global-invariants.md) § "CLI flag routing under `--config`".

## Build and Test Commands

```bash
Expand Down
4 changes: 2 additions & 2 deletions docs/cli-options.md
Original file line number Diff line number Diff line change
Expand Up @@ -577,7 +577,7 @@ Fraction of trace sessions to keep for replay, sampled whole-session to preserve

#### `-f`, `--config` `<str>`

Path to a YAML configuration file. CLI flags override values from the config file.
Path to a YAML configuration file. CLI flags override values from the config file. A flag that cannot be applied on top of a config file is rejected by name rather than ignored, so a run never silently differs from what was asked for.

### Fixed Schedule

Expand Down Expand Up @@ -2168,7 +2168,7 @@ Fraction of trace sessions to keep for replay, sampled whole-session to preserve

#### `-f`, `--config` `<str>`

Path to a YAML configuration file. CLI flags override values from the config file.
Path to a YAML configuration file. CLI flags override values from the config file. A flag that cannot be applied on top of a config file is rejected by name rather than ignored, so a run never silently differs from what was asked for.

### Fixed Schedule

Expand Down
125 changes: 125 additions & 0 deletions docs/dev/global-invariants.md
Original file line number Diff line number Diff line change
Expand Up @@ -152,6 +152,131 @@ When you fix a field, just delete its line from the baseline; the
test starts enforcing the constraint on that field on the next CI
run.

## CLI flag routing under `--config`

`resolve_config` merges a YAML config file with explicitly-set CLI
flags. Historically it only knew how to route a subset of `CLIConfig`
into the merged dict, and **every other flag was discarded without a
word** — `aiperf profile -f base.yaml --random-seed 42` ran with a
different seed than the user asked for and said nothing. For a
benchmarking tool that silently corrupts published numbers.

The invariant now in force:

> A CLI flag passed alongside `--config` either changes the resolved
> config, or raises an error naming the flag. It is never silently
> ignored.

### `CLIConfig` is the source of truth

The guarantee is scoped to what `resolve_config` can see, and its
signature is the boundary:

```python
def resolve_config(cli_config: CLIConfig, config_file: Path | None = None) -> AIPerfConfig
```

Both entry points — `aiperf profile` and `aiperf service` — call it with
nothing but their `CLIConfig` and a path. So a CLI option can influence
the resolved `AIPerfConfig` **only** by being a field on `CLIConfig`,
and that is a structural property of the call, not a convention someone
has to remember. It is what lets the classification test enumerate
`CLIConfig.model_fields` and claim completeness: there is no second
source of flags for it to miss.

Two consequences worth stating plainly:

- **Command-level parameters are out of scope, by construction.**
`aiperf service` declares `--service-id`, `--health-host`,
`--health-port` and its `service_type` argument directly on the
command function rather than on `CLIConfig`. They are service-instance
plumbing, never passed to `resolve_config`, and so cannot be dropped
by this mechanism — they never enter it. Anything that *is* benchmark
configuration belongs on `CLIConfig`, precisely so this guarantee
covers it.
- **The assumption is falsifiable, and this is how it would break.**
Giving `resolve_config` another parameter that carries user intent, or
having it read flags from module state or the environment, puts values
into the resolved config that no test enumerates. If you need to do
that, extend the classification to cover the new source in the same
breath — otherwise the suite will keep reporting a completeness it no
longer has.

Every field on `CLIConfig` must be classified in
[`_config_flag_routing.py`](https://github.com/ai-dynamo/aiperf/tree/main/src/aiperf/config/flags/_config_flag_routing.py):

| Set | Meaning |
| --- | --- |
| `ROUTED_UNDER_CONFIG` | Reaches `AIPerfConfig`. Derived from the resolver's own routing tables where possible, so it cannot drift from them. |
| `UNROUTED_UNDER_CONFIG` | Known not to route. Raises `ConfigurationError` naming the flag. |
| `EXEMPT_FROM_CONFIG_ROUTING` | Not benchmark config at all (`--config` itself). Each entry needs a stated reason. |
| `COMPANION_ROUTED` | Routes only alongside another flag (`--model-selection-strategy` needs `--model-names`). Rejected when the companion is absent. |
| `MAGIC_LIST_ONLY_UNDER_CONFIG` | Routes in list form only (`--isl 128 256` becomes a sweep parameter; scalar `--isl 128` goes to the dataset). Decided per value. |

### Why the guarantee holds

Four layers, each covering the previous one's blind spot:

1. **The universe is derived, not maintained.** Every set is
subtracted from `frozenset(CLIConfig.model_fields)`. You cannot add
a CLI option without adding a field, so nothing is checked against
a list someone has to remember to update.
2. **Runtime is default-deny.** `reject_unrouted_cli_flags` computes
`model_fields_set - ROUTED - EXEMPT` and raises on the remainder —
"is this known-good?", not "is this known-bad?". It gates on
`model_fields_set`, never truthiness, so `--prompt-batch-size 0`
counts as set.
3. **Classification is mandatory.**
`test_every_cli_config_field_is_classified` fails when a field
belongs to none of the sets, so a new flag fails CI at authoring
time.
4. **The classification is verified, not trusted.** Layer 3 only
proves someone applied a label.
`test_routed_field_never_silently_no_ops` drives every routed field
against both a synthetic and a file dataset and asserts it changes
the config or raises. This is what catches whole-section
classifications (`OUTPUT`/`TOKENIZER`/`ACCURACY`/`SWEEPING` are
marked routed wholesale, so a new member is auto-classified and
layer 3 passes vacuously).

Two further invariants guard the merge itself:
`test_override_emits_no_key_the_user_did_not_set` runs `build_dataset`
twice with different values for one field — any emitted key identical
across both runs is a materialized default that would overwrite a YAML
value the user never mentioned. `test_routed_dataset_field_actually_changes_the_resolved_config`
covers the dataset block specifically.

### Adding a new CLI flag

1. Add the field to `CLIConfig` as usual. CI now fails with
`CLI fields are unclassified for the --config path: [...]`.
2. Decide what the flag should do under `--config`:
- **Route it** (preferred). Dataset-shaping fields flow through
`build_dataset` automatically. Otherwise wire it into
`build_cli_overrides` — check first whether a builder already
exists, as `build_mlflow`/`build_otel`/`build_network_latency`
did while going uncalled.
- **Or list it in `UNROUTED_UNDER_CONFIG`**, so users get an error
naming the flag rather than a wrong benchmark.
3. If the invariant suite cannot generate a value from the
annotation (`typing.Any`, free-form strings), add one to
`FIELD_PROBE_VALUES` or list the field in `UNDRIVABLE_FIELDS` with
a reason. Do not let it report as skipped — a skip is invisible in
CI, which is exactly how a coverage hole hides.

### Known gaps

`--sweep-type` and `--disable-auto-fixed-schedule` are unrouted and
loud: the first needs a sweep block to attach to, the second is
consumed during phase construction, which this path does not rebuild.

The flags in `SWEEP_FIELDS_NOT_ROUTED` (`--concurrency-min/max/steps`,
`--isl-*`/`--osl-*`, the `*-sla-ms` filters, `--parameter-sweep-*`)
resolve cleanly and change nothing, so they are rejected. Verified
individually — notably `--ttft-sla-ms` does **not** take effect even
alongside `--search-recipe` and `--streaming`. Routing them is
follow-up work; until then the failure is loud.

## Extending the suite

### Adding a new mechanical invariant
Expand Down
Loading
Loading