Skip to content

The S-series ships with no examples: 19 codes have neither a firing case nor a look-alike #170

Description

@vinicq

Follow-up to #168 / #169. Those close the load path: the S-series is reachable from every host, and npm run validate fails if that stops being true. This issue closes the reason nobody noticed the load path was broken for two releases.

The gap

No S-code has an example in this repo. examples/python/semantic_cases.py, examples/javascript/semantic_cases.js, and examples/typescript/semantic_cases.ts all declare "Cases: 10, 11, 12, 15, 18" in their header and cover only those five. examples/robot/ has no semantic file at all. grep -rn "\bS[0-9][0-9]\?\b" examples test returns nothing.

So 19 of the 24 advertised semantic codes ship with no example that proves they fire, and none with a look-alike that proves they do not. The contributing checklist asks for both on any new pattern:

For a new pattern: an example proves it fires on the bad test AND an example proves it does NOT fire on the legitimate look-alike.

The S-series predates that rule and never got them.

Why this is the root cause of #168

The load-path defect was structurally invisible. Nothing exercised an S-code, so nothing could detect that no S-code was loading. #169 adds a static gate that catches the specific regression, but a gate over instruction text cannot tell you whether a model, given the codes, applies them correctly. Without examples the next budget optimization can silently degrade S-code detection the same way, and the only signal will be a Codex reviewer noticing prose.

Scope

One positive plus one look-alike negative per S-code, per language where the code applies. The taxonomy split from #168 holds: S8, S12 and S18 need a test-double facility, so they apply to Python, TypeScript, JavaScript and weakly to Robot, but not to Gherkin or Tavern. The other 16 apply everywhere.

Priority order, highest value first:

  1. S17, the only HIGH-severity code in the series, and the one whose look-alike boundary is narrowest: pytest.raises(SpecificError, match=...) bound to the SUT line is exempt, a bare pytest.raises(Exception) on a documented error contract is not.
  2. The contract-dependent codes S8, S11, S12, S13, S18. Their exemptions carry the most conditions, so they have the most room for a false positive. S11 is the live example: the compact rule asks for a paired positive assertion, the exemption permits a negative-only check on a filter whose contract is to drop the input entirely. That pair belongs in examples/ as two files, not only in prose.
  3. The remaining 13.

Definition of done

  • Each S-code has a positive example that a run flags and a look-alike that a run leaves alone, in at least one language.
  • The look-alike files are named so the intent is obvious from the path, following the existing family_* convention.
  • examples/robot/ gains a semantic file.
  • The semantic_cases.* headers stop claiming "Cases: 10, 11, 12, 15, 18" once they cover more.
  • npm run validate stays green; if a new check is warranted, extend check-catalog-consistency.mjs rather than adding a script.

Refs #168, #169.

Activity

  1. vinicq commented on Jul 29, 2026

    @vinicq
    OwnerAuthor

    Scope amendment from the #169 review.

    The corpus has to cover the structural codes that were undiscoverable too, not only the S-series. #169 found the same reachability defect in the structural catalog: SKILL.md Step 2b carried 10 of the 24 JS-codes and Robot had no table at all, so 23 structural codes were reachable only by loading a whole reference.md language section. That is fixed by a generated index, but the codes still have no labeled example proving they fire or a look-alike proving they do not.

    If this issue covers only the S-series, its ground truth bakes in the identical blind spot and we rediscover it in the review of whatever PR closes it.

    Revised scope: one positive plus one look-alike negative per code, for the 19 S-codes and for the JS, R and PL codes absent from the compact tables before #169. Priority order unchanged for the semantic side; on the structural side start with the codes that had no summary-table entry at all, since those are the ones no reader could have found.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions