Problem
run_screening retries the provider only on AssessmentContractError (parse-level contract failure): the _assessment_retry_prefix retry wraps only that exception. A ScreeningQualityError raised by validate_assessment_quality propagates out of the screening stage, and the completed provider response is discarded. The model never sees the issue codes, so the next call — if any — starts from zero feedback.
Observed behavior
Measured on a real workspace run this week (counts cited, examples withheld):
- A batch of fourteen strong-provider screenings passed eight and produced six quality-gate rejections.
- The same record was rejected with different issue codes across attempts — a summary-count violation in one call and a citation-level violation in another — showing the gate input is stable while the response varies.
- A later screening-only run under a fixed gate re-screened six records from that batch (four previously rejected, two previously passed): five passed with unchanged pipeline inputs, one was rejected again with a summary-count violation.
- One record that had been rejected twice passed on its third attempt without any change other than the re-run itself.
- Every rejected attempt still consumed a full provider call (one to six minutes).
Three violation classes remain as genuine output variance: extra prose after citations, a summary count sentence disagreeing with the model's own matches, and citations that fail the corpus check (then demotion makes the count sentence stale). A fourth class observed during the same runs — hold-condition lines that name several requirements in one marker — turned out to be a deterministic gate bug in marker parsing and is fixed separately; it is not evidence of variance.
Why it matters
- A single bad-format response costs the whole record, even though the next response would likely pass.
- A later run with an unwrapped provider environment can then write a fallback placeholder over the empty state.
Fix candidate
On ScreeningQualityError, retry once with the stable issue codes prefixed to the prompt, mirroring the existing _assessment_retry_prefix for parse errors. The feedback contains issue codes only — no source content — so the gate stays fail-closed and the allowlist semantics do not change.
Two open questions before implementing:
- Should the retry cap be one attempt or measured (the measured-rejection corpus from a corpus run is needed before choosing)?
- Should provider telemetry count a quality-feedback retry as a second attempt of the same provider?
Problem
run_screeningretries the provider only onAssessmentContractError(parse-level contract failure): the_assessment_retry_prefixretry wraps only that exception. AScreeningQualityErrorraised byvalidate_assessment_qualitypropagates out of the screening stage, and the completed provider response is discarded. The model never sees the issue codes, so the next call — if any — starts from zero feedback.Observed behavior
Measured on a real workspace run this week (counts cited, examples withheld):
Three violation classes remain as genuine output variance: extra prose after citations, a summary count sentence disagreeing with the model's own matches, and citations that fail the corpus check (then demotion makes the count sentence stale). A fourth class observed during the same runs — hold-condition lines that name several requirements in one marker — turned out to be a deterministic gate bug in marker parsing and is fixed separately; it is not evidence of variance.
Why it matters
Fix candidate
On
ScreeningQualityError, retry once with the stable issue codes prefixed to the prompt, mirroring the existing_assessment_retry_prefixfor parse errors. The feedback contains issue codes only — no source content — so the gate stays fail-closed and the allowlist semantics do not change.Two open questions before implementing: