Skip to content

Quality-gate rejections are not retried with feedback #36

Description

@teslamint

Problem

run_screening retries the provider only on AssessmentContractError (parse-level contract failure): the _assessment_retry_prefix retry wraps only that exception. A ScreeningQualityError raised by validate_assessment_quality propagates out of the screening stage, and the completed provider response is discarded. The model never sees the issue codes, so the next call — if any — starts from zero feedback.

Observed behavior

Measured on a real workspace run this week (counts cited, examples withheld):

  • A batch of fourteen strong-provider screenings passed eight and produced six quality-gate rejections.
  • The same record was rejected with different issue codes across attempts — a summary-count violation in one call and a citation-level violation in another — showing the gate input is stable while the response varies.
  • A later screening-only run under a fixed gate re-screened six records from that batch (four previously rejected, two previously passed): five passed with unchanged pipeline inputs, one was rejected again with a summary-count violation.
  • One record that had been rejected twice passed on its third attempt without any change other than the re-run itself.
  • Every rejected attempt still consumed a full provider call (one to six minutes).

Three violation classes remain as genuine output variance: extra prose after citations, a summary count sentence disagreeing with the model's own matches, and citations that fail the corpus check (then demotion makes the count sentence stale). A fourth class observed during the same runs — hold-condition lines that name several requirements in one marker — turned out to be a deterministic gate bug in marker parsing and is fixed separately; it is not evidence of variance.

Why it matters

  • A single bad-format response costs the whole record, even though the next response would likely pass.
  • A later run with an unwrapped provider environment can then write a fallback placeholder over the empty state.

Fix candidate

On ScreeningQualityError, retry once with the stable issue codes prefixed to the prompt, mirroring the existing _assessment_retry_prefix for parse errors. The feedback contains issue codes only — no source content — so the gate stays fail-closed and the allowlist semantics do not change.

Two open questions before implementing:

  1. Should the retry cap be one attempt or measured (the measured-rejection corpus from a corpus run is needed before choosing)?
  2. Should provider telemetry count a quality-feedback retry as a second attempt of the same provider?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions