Skip to content

feat(eval): add model-quality benchmark fixtures and metrics - #267

Closed
zachshallbetter wants to merge 1 commit into
nikopueringer:mainfrom
zachshallbetter:feat/eval-benchmark-harness
Closed

zachshallbetter wants to merge 1 commit into
nikopueringer:mainfrom
zachshallbetter:feat/eval-benchmark-harness

Conversation

@zachshallbetter

@zachshallbetter zachshallbetter commented Sep 21, 2026 •

Copy link
Copy Markdown

What does this change?

The repository benchmarks runtime speed in benchmarks/bench_inference.py (PR #221), measuring FPS, latency, and VRAM scaling. It doesn't measure whether an inference result is correct.

Latency isn't key quality. It's a separate measurement.

Without a quality harness, evaluating candidate changes relies on eyeballing frames. Eyeballing doesn't scale across subtle edge cases. We're already seeing the limits of visual inspection in open discussions:

This PR adds a standalone, reproducible model-quality evaluation harness in corridorkey.eval. It measures alpha error, straight foreground recovery, edge boundary error, and physical recomposition. It doesn't touch the training pipeline.

Area Change
corridorkey/eval/cases.py MattingEvalSample dataclass contract and dataset loaders. Validates dimensions, channel counts, and spatial alignment.
corridorkey/eval/metrics.py Four core quality metrics: alpha_mae, foreground_mae (where A > 0.01), boundary_alpha_mae (0.01 < A < 0.99 plus morphological band), and recomposition_mae (C_hat = alpha_hat * F_hat + (1 - alpha_hat) * B).
corridorkey/eval/fixtures.py Generates 7 canonical diagnostic cases on demand: opaque_subject, soft_hair, motion_blur, green_foreground, blue_foreground, bad_hint_eroded, and bad_hint_dilated. Zero binary files tracked in git.
corridorkey/eval/runner.py Engine adapter supporting CorridorKeyEngine, PyTorch modules, and callables/mocks. Aggregates global means and tag breakdowns.
corridorkey/eval/report.py Formats structured JSON benchmark receipts and terminal summary tables.
corridorkey/eval/__main__.py CLI entry point: python -m corridorkey.eval and console script corridorkey-eval.
pyproject.toml Packages corridorkey and backend into wheel targets; adds corridorkey-eval script.
.gitignore Ignores generated benchmark datasets, eval receipts (eval*.json), and cache artifacts.
tests/eval/ 24 unit and integration tests covering metrics, loaders, reporting, and CLI.

How was it tested?

  1. Unit and integration suite:

    uv run pytest tests/eval/ -v

    24 passed in 3.46s across metric math, boundary masking, sample contracts, and CLI execution.

  2. CLI execution:

    uv run corridorkey-eval --checkpoint mock --dataset /tmp/ck_eval --generate-fixtures --output /tmp/ck_eval/receipt.json

    Generates fixtures on the fly, evaluates all 7 cases, prints the terminal summary table, and writes a valid JSON receipt.

  3. Wheel build and packaging:

    uv build --wheel
    unzip -l dist/corridorkey-1.0.0-py3-none-any.whl

    Confirmed corridorkey, corridorkey/eval, backend, and CorridorKeyModule are properly packaged at the root.

  4. Backwards compatibility:
    Existing inference pipelines, wizard prompts, and CLI options are completely untouched.

Checklist

  • uv run pytest passes
  • uv run ruff check passes
  • uv run ruff format --check passes

The repository benchmarks runtime speed. It does not measure whether
a model change makes the key better.

This change adds a reproducible evaluation harness:
- Define MattingEvalSample data contract in corridorkey.eval.cases
- Implement 4 foundational metrics in corridorkey.eval.metrics:
  * alpha_mae: Mean Absolute Error across all pixels
  * foreground_mae: Mean Absolute Error for recovered straight foreground
  * boundary_alpha_mae: Alpha MAE evaluated strictly on the boundary band
    (fractional alpha plus morphological edge transition)
  * recomposition_mae: Error reconstructing C_hat = alpha*F + (1-alpha)*B
- Add deterministic generator for 7 canonical diagnostic benchmark cases:
  opaque_subject, soft_hair, motion_blur, green_foreground,
  blue_foreground, bad_hint_eroded, and bad_hint_dilated
- Provide CLI runner (python -m corridorkey.eval / corridorkey-eval)
  with structured JSON receipt export and terminal summary tables
- Update pyproject.toml packaging to include corridorkey in wheel targets
  and expose console script entrypoint
- Update .gitignore to ignore generated benchmark datasets, eval receipts,
  and cache artifacts
- Add unit test suite covering metrics, loaders, and runner (24 passing)
@zachshallbetter zachshallbetter closed this by deleting the head repository Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant