Skip to content

Semantic validator is never constructed: wire it or remove it #35

Description

@teslamint

Problem

run_screening runs the semantic validator only when one is passed (screening.py:641-642), and nothing in src/ ever constructs one. _build_services builds the screening stage without the argument (cli.py:440-444), so the field is always None.

grep -rn "CalibratedSemanticValidator(" src/ returns no call site; the class at screening_semantics.py:70 is reached only from tests.

So every screening published through the CLI skips semantic validation. The gates that do run are the quality gate, the structure check, the match-table parse, and the evidence-row check. The semantic layer is written, tested, and unreachable.

require_semantic_validation exists and raises when it is set without a validator, but nothing in the CLI sets it either, so it cannot surface the gap.

Decision needed

Either wire it or remove it. Both are defensible and the choice is not obvious from the code:

Wire it. The CLI builds a judge and passes a validator. Cost: the judge is an LLM, so every screening pays extra latency, and a non-deterministic component becomes a publication gate. CalibratedSemanticValidator.validate also runs evaluate_semantic_judge on first use and raises semantic-eval-threshold-not-met when calibration fails, which turns a judge regression into a screening outage. If it ships, it needs a documented opt-in flag and a decision about whether a judge failure blocks publication or only annotates it.

Remove it. Delete the class, the semantic_validator/require_semantic_validation parameters, and their tests. The remaining gates keep enforcing citation-to-source verbatim matching, which is the check that actually rejects unsupported claims today.

Leaving it as-is is the worst option: the code reads as if screening is semantically validated when no run ever is.

Acceptance

Whichever direction is chosen:

  • No parameter remains whose only effect is to be permanently None.
  • If wired: the flag that enables it is documented, a judge failure has a defined outcome, and the added per-screening cost is measured and recorded.
  • If removed: no test asserts behavior of a code path the application cannot reach.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions