Skip to content

add a first-class data-access domain for cht-datasource memories #151

Description

@Hareet

Describe the issue

cht-datasource (shared-libs/cht-datasource) work has no clean home among the nine
CHTDomain values, so memories about the data-access layer are scattered across whichever
product domain their tickets happened to touch. This is the context-selection gap raised in
#86: the distiller injects the wrong patterns for these tickets.

The scatter is wider than previously scoped. Counting drafts whose body references
cht-datasource:

Location Count
already merged on main 4 (9835, 10344, 10074 in contacts; 10266 in data-sync)
#122 forms-and-reports 6
#123 tasks-and-targets 2
#132 contacts 15

Of these, six have anchor PRs that touch only shared-libs/cht-datasource — 10071
(4/4 files), 10180 (2/2), 10043 (2/2), 10057 (2/2), 9266 (17/17), 9281 (12/12).
Another ten are library work that also wires the api controller and route backing the remote
path — 10141 (14/18), 9368 (16/22), 10390 (13/21), 10064 (11/18), 10173 (9/14),
9295 (10/15), 9625 (49/74), 9090 (34/82), 9177 (11/61), 9390 (2/5).

The classification rule proposed in #122 — primary domain is data-access when the anchor
PR touches only shared-libs/cht-datasource
— is not decidable against that population. It
misclassifies its own worked example: 10064 is named as an extender but touches
api/src/controllers/report.js, two webapp files, and shared-libs/lineage/src/index.d.ts.

Describe the improvement you'd like

A tenth domain, data-access, plus an optional secondaryDomains: CHTDomain[] so a memory
can declare the product areas it also serves without needing two co-equal primaries.

Rule, restated so it decides the real population: primary domain is data-access when the
PR adds or changes the library's public API surface (src/<entity>.ts, src/local/,
src/remote/, src/qualifier.ts), whether or not it also wires the matching api controller
and route
— that triad is the library's own idiom, documented as such in the 10390
draft's Code Patterns section. A draft that merely consumes the library keeps its product
domain with subDomain: cht-datasource (e.g. 9755, which touches zero cht-datasource
files).

Sequenced so that a schema change, a content change, and a behaviour change are never in
flight together, and so that nothing on the critical path waits on #135:

Why L comes first. #135 is not cht-conf work — it is the loader/pipeline contract bug,
and it gates memory retrieval for every domain. Only its current location is
cht-conf-specific: the fix is written and pushed on origin/134-cht-conf with no PR of its
own, cleanly separable at 3 files / +234-52 against 17 files / +1241 for that branch as a
whole. Landing it independently is worth doing on its own merits — until it lands, nothing
reads the memory corpus at all, so no amount of taxonomy work is observable. It also makes C
testable: with the loader live, a lossy re-key shows up as a failing test rather than as a
silent regression discovered months later.

C carries a hard constraint. Retrieval selects candidates by primary-domain directory,
before any scoring:

const draftsPath = path.join(AGENT_MEMORY_PATH, 'domains', domain, 'issues');
return scanDraftsForIssues(draftsPath, domain);

Moving ~20 drafts into domains/data-access/issues/ therefore removes them from every
contacts / forms-and-reports / tasks-and-targets query, and a secondary-domain weight inside
calculateSimilarityScore cannot recover a candidate the selector already excluded. C must
ship the union-selection change together with the moves
— primary-domain directory ∪ drafts
whose secondaryDomains include the query domain — or it is a net retrieval regression.

open-review-pr.ts already stages into path.join(domainsDir, domain, 'issues'), so the
promotion tooling needs no change for a new domain.

Two content problems to resolve in C, not silently:

  • 10390 documents 21 files named target-interval.* that exist nowhere in cht-core —
    git ls-files | grep -c target-interval returns 0. PR 10390 merged into the epic branch
    10140_previous-month-targets, which renamed them all to target.* before landing as
    #10423 / 622c625427. Its source_sha (60ca9634…) does not resolve. chore(memory): promote strong-fit tasks-and-targets drafts from memory-pipeline for review #123's 10423
    draft already documents the same work with correct filenames, so 10390 should be
    dropped, not rewritten.
  • 9835 carries source_prs: [10022, 10081, 10083, 10222, 10246], and 10344 carries
    [10432]. Both predate the Domain Rationale convention and have no such section, as does
    10074.

Describe alternatives you've considered

  • subDomain: cht-datasource alone. Free-text, scalar, and read by zero code
    (grep -rn subDomain src/ test/ → no hits). It can mark a consumer but cannot express
    "primarily data-access, also relevant to forms".
  • related_workflows. Wrong axis — a CHTWorkflow, not a CHTDomain. Also written by
    the distiller and never read in retrieval.
  • Two co-equal primary domains. Breaks the directory-per-domain layout and the
    domain singular contract the whole pipeline is built on.
  • One coordinated PR (schema + moves + rewrites + retrieval weighting, as originally
    proposed). Mixes a schema change, ~20 file moves rendered as delete+add, a content rewrite,
    and a behaviour change into one diff; if any part stalls, all of it reverts. The A/B/C split
    gets the same end state with each step independently reviewable and revertible.
  • Landing the schema after the content PRs, to avoid rebases. Measured: there are no
    rebases to avoid. Applying the enum + field to main and then merging all three PR heads
    produces zero conflicts — the content PRs touch schema.json only at the source_prs
    block (~line 138), the enum is at line 10 and the new field at ~line 76. The three
    source_prs additions are byte-identical across the branches, so they don't conflict with
    each other either.

Acceptance criteria

  • data-access in CHTDomain, CHT_DOMAINS, and every exhaustive Record<CHTDomain, …> (PR A)
  • secondaryDomains in the schema, optional and backward-compatible (PR A)
  • Classification rule documented, with the mixed extend+wire case decided (PR A/C)
  • Context analysis agent doesn't load memory-pipeline drafts (wrong path + schema) #135 loader fix landed independently of the cht-conf work (PR L)
  • Extenders re-keyed and moved; consumers keep their product domain with subDomain (PR C)
  • Candidate selection reads secondaryDomains — primary directory ∪ secondary matches — shipped with the moves (PR C)
  • 10390 dropped as a duplicate of 10423 (PR C)
  • Domain Rationale added for 9835, 10344, 10074 (PR C)

Known adjacent gaps (not blocking, worth tracking)

  • code-context-agent.mock-data.ts types its map Record<CHTDomain, MockCodeContextData>
    but builds it via a JSON.parse(...) as cast, so exhaustiveness is not enforced — it
    silently lacks both infrastructure and data-access. Reads are guarded by
    || EMPTY_MOCK_CODE_CONTEXT_DATA.
  • agent-memory/indices/component-to-domains.json is an empty stub.
  • The test suite appends to the tracked file agent-memory/_skipped.ndjson on every run.

Activity

  1. added 4 commits that reference this issue on Aug 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions