Skip to content

feat(eval): parameterized alpha-hint corruption benchmark - #268

Closed
zachshallbetter wants to merge 2 commits into
nikopueringer:mainfrom
zachshallbetter:feat/eval-hint-corruption
Closed

zachshallbetter wants to merge 2 commits into
nikopueringer:mainfrom
zachshallbetter:feat/eval-hint-corruption

Conversation

@zachshallbetter

Copy link
Copy Markdown

What does this change?

Relates to #184 and #195. Builds on top of #267 (feat/eval-benchmark-harness).

In #184, Niko noted that GreenFormer behaves inconsistently when alpha guide hints drift. During training, hints were coarse, eroded masks. The network learned to expand foreground and recover fine details well. It didn't learn to subtract false-positive background islands.

Our baseline benchmark harness (#267) measures global MAE, but global MAE hides this directional bias. A 4px dilation error produces roughly the same global error as a 4px erosion error. We need a way to measure foreground recovery separately from background rejection.

This PR adds a parameterized alpha-hint corruption harness and asymmetry sweep in corridorkey.eval:

Area Change
corridorkey/eval/hint_corruption.py HintCorruptionSpec dataclass and corruption engine supporting morphological erosion, dilation, interior holes, false islands, Gaussian blur, and edge jitter.
corridorkey/eval/hint_corruption.py Asymmetry metrics compute_hint_iou and compute_asymmetric_error, separating fp_alpha (background contamination) from fn_alpha (foreground loss).
corridorkey/eval/hint_corruption.py CANONICAL_HINT_PROFILES dictionary with 12 standard corruption specs (4px to 16px radius across undersegmentation, oversegmentation, and topological noise).
corridorkey/eval/hint_corruption.py run_hint_sweep runner aggregating metrics across models and corruption profiles.
corridorkey/eval/report.py Formats ASCII tables and Rich tables with dedicated FP Alpha and FN Alpha columns.
corridorkey/eval/__main__.py CLI flag --hint-sweep to run matrix sweeps and save structured JSON receipts.
tests/eval/test_hint_corruption.py 13 unit tests covering each corruption kernel, IoU, asymmetric error separation, dimension squeezing, empty sample validation, and CLI execution.

How was it tested?

  1. Unit and integration suite:

    uv run pytest tests/eval/ -v

    37 passed in 64s across all baseline metrics, corruption operations, asymmetric error separation, and CLI execution.

  2. CLI execution and sweep output:

    uv run python -m corridorkey.eval --checkpoint mock --dataset /tmp/hint_test --generate-fixtures --hint-sweep --output /tmp/hint_test/sweep.json

    Confirmed all 12 profiles run against the mock engine, producing formatted summary tables and saving structured JSON receipts.

  3. Asymmetry validation:
    Verified that erosion profiles isolate fn_alpha (e.g. erosion_16px yields fn_alpha=0.4174 with fp_alpha=0.0041), while dilation profiles isolate fp_alpha (dilation_16px yields fp_alpha=0.2055 with fn_alpha=0.0100).

  4. Backwards compatibility:
    Standard evaluation (python -m corridorkey.eval) and existing pipelines remain completely unaffected.

Checklist

  • uv run pytest passes
  • uv run ruff check passes
  • uv run ruff format --check passes

The repository benchmarks runtime speed. It does not measure whether
a model change makes the key better.

This change adds a reproducible evaluation harness:
- Define MattingEvalSample data contract in corridorkey.eval.cases
- Implement 4 foundational metrics in corridorkey.eval.metrics:
  * alpha_mae: Mean Absolute Error across all pixels
  * foreground_mae: Mean Absolute Error for recovered straight foreground
  * boundary_alpha_mae: Alpha MAE evaluated strictly on the boundary band
    (fractional alpha plus morphological edge transition)
  * recomposition_mae: Error reconstructing C_hat = alpha*F + (1-alpha)*B
- Add deterministic generator for 7 canonical diagnostic benchmark cases:
  opaque_subject, soft_hair, motion_blur, green_foreground,
  blue_foreground, bad_hint_eroded, and bad_hint_dilated
- Provide CLI runner (python -m corridorkey.eval / corridorkey-eval)
  with structured JSON receipt export and terminal summary tables
- Update pyproject.toml packaging to include corridorkey in wheel targets
  and expose console script entrypoint
- Update .gitignore to ignore generated benchmark datasets, eval receipts,
  and cache artifacts
- Add unit test suite covering metrics, loaders, and runner (24 passing)
GreenFormer hints during training were coarse, eroded masks. The model
learned to grow missing foreground well. It didn't learn to cut away
false-positive background islands.

This adds a corruption harness to quantify that asymmetry:
- HintCorruptionSpec with erosion, dilation, holes, and false islands
- compute_asymmetric_error to break out false positives vs false negatives
- Standard test profiles from 4px to 16px radius
- CLI --hint-sweep flag with formatted tables and JSON receipts
- Full unit test coverage for corruption routines and metrics
@zachshallbetter zachshallbetter closed this by deleting the head repository Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant