Follow-up to #168 / #169. Those close the load path: the S-series is reachable from every host, and npm run validate fails if that stops being true. This issue closes the reason nobody noticed the load path was broken for two releases.
The gap
No S-code has an example in this repo. examples/python/semantic_cases.py, examples/javascript/semantic_cases.js, and examples/typescript/semantic_cases.ts all declare "Cases: 10, 11, 12, 15, 18" in their header and cover only those five. examples/robot/ has no semantic file at all. grep -rn "\bS[0-9][0-9]\?\b" examples test returns nothing.
So 19 of the 24 advertised semantic codes ship with no example that proves they fire, and none with a look-alike that proves they do not. The contributing checklist asks for both on any new pattern:
For a new pattern: an example proves it fires on the bad test AND an example proves it does NOT fire on the legitimate look-alike.
The S-series predates that rule and never got them.
Why this is the root cause of #168
The load-path defect was structurally invisible. Nothing exercised an S-code, so nothing could detect that no S-code was loading. #169 adds a static gate that catches the specific regression, but a gate over instruction text cannot tell you whether a model, given the codes, applies them correctly. Without examples the next budget optimization can silently degrade S-code detection the same way, and the only signal will be a Codex reviewer noticing prose.
Scope
One positive plus one look-alike negative per S-code, per language where the code applies. The taxonomy split from #168 holds: S8, S12 and S18 need a test-double facility, so they apply to Python, TypeScript, JavaScript and weakly to Robot, but not to Gherkin or Tavern. The other 16 apply everywhere.
Priority order, highest value first:
- S17, the only HIGH-severity code in the series, and the one whose look-alike boundary is narrowest:
pytest.raises(SpecificError, match=...) bound to the SUT line is exempt, a bare pytest.raises(Exception) on a documented error contract is not.
- The contract-dependent codes S8, S11, S12, S13, S18. Their exemptions carry the most conditions, so they have the most room for a false positive. S11 is the live example: the compact rule asks for a paired positive assertion, the exemption permits a negative-only check on a filter whose contract is to drop the input entirely. That pair belongs in
examples/ as two files, not only in prose.
- The remaining 13.
Definition of done
- Each S-code has a positive example that a run flags and a look-alike that a run leaves alone, in at least one language.
- The look-alike files are named so the intent is obvious from the path, following the existing
family_* convention.
examples/robot/ gains a semantic file.
- The
semantic_cases.* headers stop claiming "Cases: 10, 11, 12, 15, 18" once they cover more.
npm run validate stays green; if a new check is warranted, extend check-catalog-consistency.mjs rather than adding a script.
Refs #168, #169.
Follow-up to #168 / #169. Those close the load path: the S-series is reachable from every host, and
npm run validatefails if that stops being true. This issue closes the reason nobody noticed the load path was broken for two releases.The gap
No S-code has an example in this repo.
examples/python/semantic_cases.py,examples/javascript/semantic_cases.js, andexamples/typescript/semantic_cases.tsall declare "Cases: 10, 11, 12, 15, 18" in their header and cover only those five.examples/robot/has no semantic file at all.grep -rn "\bS[0-9][0-9]\?\b" examples testreturns nothing.So 19 of the 24 advertised semantic codes ship with no example that proves they fire, and none with a look-alike that proves they do not. The contributing checklist asks for both on any new pattern:
The S-series predates that rule and never got them.
Why this is the root cause of #168
The load-path defect was structurally invisible. Nothing exercised an S-code, so nothing could detect that no S-code was loading. #169 adds a static gate that catches the specific regression, but a gate over instruction text cannot tell you whether a model, given the codes, applies them correctly. Without examples the next budget optimization can silently degrade S-code detection the same way, and the only signal will be a Codex reviewer noticing prose.
Scope
One positive plus one look-alike negative per S-code, per language where the code applies. The taxonomy split from #168 holds: S8, S12 and S18 need a test-double facility, so they apply to Python, TypeScript, JavaScript and weakly to Robot, but not to Gherkin or Tavern. The other 16 apply everywhere.
Priority order, highest value first:
pytest.raises(SpecificError, match=...)bound to the SUT line is exempt, a barepytest.raises(Exception)on a documented error contract is not.examples/as two files, not only in prose.Definition of done
family_*convention.examples/robot/gains a semantic file.semantic_cases.*headers stop claiming "Cases: 10, 11, 12, 15, 18" once they cover more.npm run validatestays green; if a new check is warranted, extendcheck-catalog-consistency.mjsrather than adding a script.Refs #168, #169.