Repository navigation
feat(eval): add model-quality benchmark fixtures and metrics - #267
Closed
zachshallbetter wants to merge 1 commit into
Closed
zachshallbetter wants to merge 1 commit into
zachshallbetter wants to merge 1 commit into
Conversation
The repository benchmarks runtime speed. It does not measure whether
a model change makes the key better.
This change adds a reproducible evaluation harness:
- Define MattingEvalSample data contract in corridorkey.eval.cases
- Implement 4 foundational metrics in corridorkey.eval.metrics:
* alpha_mae: Mean Absolute Error across all pixels
* foreground_mae: Mean Absolute Error for recovered straight foreground
* boundary_alpha_mae: Alpha MAE evaluated strictly on the boundary band
(fractional alpha plus morphological edge transition)
* recomposition_mae: Error reconstructing C_hat = alpha*F + (1-alpha)*B
- Add deterministic generator for 7 canonical diagnostic benchmark cases:
opaque_subject, soft_hair, motion_blur, green_foreground,
blue_foreground, bad_hint_eroded, and bad_hint_dilated
- Provide CLI runner (python -m corridorkey.eval / corridorkey-eval)
with structured JSON receipt export and terminal summary tables
- Update pyproject.toml packaging to include corridorkey in wheel targets
and expose console script entrypoint
- Update .gitignore to ignore generated benchmark datasets, eval receipts,
and cache artifacts
- Add unit test suite covering metrics, loaders, and runner (24 passing)
3 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this change?
The repository benchmarks runtime speed in
benchmarks/bench_inference.py(PR #221), measuring FPS, latency, and VRAM scaling. It doesn't measure whether an inference result is correct.Latency isn't key quality. It's a separate measurement.
Without a quality harness, evaluating candidate changes relies on eyeballing frames. Eyeballing doesn't scale across subtle edge cases. We're already seeing the limits of visual inspection in open discussions:
This PR adds a standalone, reproducible model-quality evaluation harness in
corridorkey.eval. It measures alpha error, straight foreground recovery, edge boundary error, and physical recomposition. It doesn't touch the training pipeline.corridorkey/eval/cases.pyMattingEvalSampledataclass contract and dataset loaders. Validates dimensions, channel counts, and spatial alignment.corridorkey/eval/metrics.pyalpha_mae,foreground_mae(whereA > 0.01),boundary_alpha_mae(0.01 < A < 0.99plus morphological band), andrecomposition_mae(C_hat = alpha_hat * F_hat + (1 - alpha_hat) * B).corridorkey/eval/fixtures.pyopaque_subject,soft_hair,motion_blur,green_foreground,blue_foreground,bad_hint_eroded, andbad_hint_dilated. Zero binary files tracked in git.corridorkey/eval/runner.pyCorridorKeyEngine, PyTorch modules, and callables/mocks. Aggregates global means and tag breakdowns.corridorkey/eval/report.pycorridorkey/eval/__main__.pypython -m corridorkey.evaland console scriptcorridorkey-eval.pyproject.tomlcorridorkeyandbackendinto wheel targets; addscorridorkey-evalscript..gitignoreeval*.json), and cache artifacts.tests/eval/How was it tested?
Unit and integration suite:
24 passed in 3.46s across metric math, boundary masking, sample contracts, and CLI execution.
CLI execution:
Generates fixtures on the fly, evaluates all 7 cases, prints the terminal summary table, and writes a valid JSON receipt.
Wheel build and packaging:
Confirmed
corridorkey,corridorkey/eval,backend, andCorridorKeyModuleare properly packaged at the root.Backwards compatibility:
Existing inference pipelines, wizard prompts, and CLI options are completely untouched.
Checklist
uv run pytestpassesuv run ruff checkpassesuv run ruff format --checkpasses