Repository navigation
feat(eval): parameterized alpha-hint corruption benchmark - #268
Closed
zachshallbetter wants to merge 2 commits into
Closed
zachshallbetter wants to merge 2 commits into
zachshallbetter wants to merge 2 commits into
Conversation
The repository benchmarks runtime speed. It does not measure whether
a model change makes the key better.
This change adds a reproducible evaluation harness:
- Define MattingEvalSample data contract in corridorkey.eval.cases
- Implement 4 foundational metrics in corridorkey.eval.metrics:
* alpha_mae: Mean Absolute Error across all pixels
* foreground_mae: Mean Absolute Error for recovered straight foreground
* boundary_alpha_mae: Alpha MAE evaluated strictly on the boundary band
(fractional alpha plus morphological edge transition)
* recomposition_mae: Error reconstructing C_hat = alpha*F + (1-alpha)*B
- Add deterministic generator for 7 canonical diagnostic benchmark cases:
opaque_subject, soft_hair, motion_blur, green_foreground,
blue_foreground, bad_hint_eroded, and bad_hint_dilated
- Provide CLI runner (python -m corridorkey.eval / corridorkey-eval)
with structured JSON receipt export and terminal summary tables
- Update pyproject.toml packaging to include corridorkey in wheel targets
and expose console script entrypoint
- Update .gitignore to ignore generated benchmark datasets, eval receipts,
and cache artifacts
- Add unit test suite covering metrics, loaders, and runner (24 passing)
GreenFormer hints during training were coarse, eroded masks. The model learned to grow missing foreground well. It didn't learn to cut away false-positive background islands. This adds a corruption harness to quantify that asymmetry: - HintCorruptionSpec with erosion, dilation, holes, and false islands - compute_asymmetric_error to break out false positives vs false negatives - Standard test profiles from 4px to 16px radius - CLI --hint-sweep flag with formatted tables and JSON receipts - Full unit test coverage for corruption routines and metrics
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this change?
Relates to #184 and #195. Builds on top of #267 (
feat/eval-benchmark-harness).In #184, Niko noted that GreenFormer behaves inconsistently when alpha guide hints drift. During training, hints were coarse, eroded masks. The network learned to expand foreground and recover fine details well. It didn't learn to subtract false-positive background islands.
Our baseline benchmark harness (#267) measures global MAE, but global MAE hides this directional bias. A 4px dilation error produces roughly the same global error as a 4px erosion error. We need a way to measure foreground recovery separately from background rejection.
This PR adds a parameterized alpha-hint corruption harness and asymmetry sweep in
corridorkey.eval:corridorkey/eval/hint_corruption.pyHintCorruptionSpecdataclass and corruption engine supporting morphological erosion, dilation, interior holes, false islands, Gaussian blur, and edge jitter.corridorkey/eval/hint_corruption.pycompute_hint_iouandcompute_asymmetric_error, separatingfp_alpha(background contamination) fromfn_alpha(foreground loss).corridorkey/eval/hint_corruption.pyCANONICAL_HINT_PROFILESdictionary with 12 standard corruption specs (4px to 16px radius across undersegmentation, oversegmentation, and topological noise).corridorkey/eval/hint_corruption.pyrun_hint_sweeprunner aggregating metrics across models and corruption profiles.corridorkey/eval/report.pyFP AlphaandFN Alphacolumns.corridorkey/eval/__main__.py--hint-sweepto run matrix sweeps and save structured JSON receipts.tests/eval/test_hint_corruption.pyHow was it tested?
Unit and integration suite:
37 passed in 64s across all baseline metrics, corruption operations, asymmetric error separation, and CLI execution.
CLI execution and sweep output:
Confirmed all 12 profiles run against the mock engine, producing formatted summary tables and saving structured JSON receipts.
Asymmetry validation:
Verified that erosion profiles isolate
fn_alpha(e.g.erosion_16pxyieldsfn_alpha=0.4174withfp_alpha=0.0041), while dilation profiles isolatefp_alpha(dilation_16pxyieldsfp_alpha=0.2055withfn_alpha=0.0100).Backwards compatibility:
Standard evaluation (
python -m corridorkey.eval) and existing pipelines remain completely unaffected.Checklist
uv run pytestpassesuv run ruff checkpassesuv run ruff format --checkpasses