Skip to content

Latest commit

 

History

History
331 lines (228 loc) · 23.8 KB

File metadata and controls

331 lines (228 loc) · 23.8 KB

Roadmap

Trusted compiler configuration discovery

Compilation databases record the compiler command selected by the build system, but compiler wrappers may add arguments only when they execute. Nix compiler wrappers commonly inject glibc, C++ standard-library, GCC internal-header, target, sysroot, and toolchain paths that therefore do not appear in compile_commands.json.

Scalps currently passes the recorded arguments directly to an in-process Clang frontend. It does not execute the named compiler, so it cannot observe those wrapper additions. The robust solution is an opt-in, GCC-compatible compiler query facility similar to clangd's --query-driver.

The implementation should remain generic rather than parse Nix-specific wrapper files:

compile_commands.json
        ↓ load and validate
RecordedCompilationCommand
        ↓ resolve and optionally query an authorized driver
EffectiveCompilationCommand
        ↓ hash, freshness planning, and extraction
ClangTool

Compilation-database loading should remain free of process execution. A separate preparation stage should decorate commands before extraction, command hashing, or freshness planning.

User interface and trust boundary

index and status should accept a repeatable query-driver allowlist:

scalps index --query-driver '/nix/store/**-gcc-wrapper-*/bin/g++'
scalps status --query-driver '/nix/store/**-gcc-wrapper-*/bin/g++'

Compiler querying must be disabled by default. Allowlist patterns must be absolute and should match the resolved, normalized compiler path. Bare compiler names in compilation commands should be resolved through PATH before matching. * should match within a path component, while ** may cross path separators.

No executable named by a compilation database may run unless its resolved path matches the explicit allowlist. An allowlist supplied by the user that matches no compilation commands should be an error so the command does not falsely appear to have enabled discovery.

The driver must be invoked directly without a shell, with empty standard input, a bounded runtime, and a captured-output limit. It must run in the compilation command's working directory and inherit the current environment because compiler wrappers can depend on that environment. Exit failures, timeouts, and malformed output need distinct diagnostics.

Query contract

Each distinct compiler-query key should run the GCC-compatible discovery commands once per scalps invocation:

<driver> -E -v -x <language> -
<driver> -print-file-name=include

The key should contain the resolved driver, working directory, effective language, standard-include suppression flags, explicit target and sysroot options, standard-library selection, and a reviewed set of multilib-selection options.

Only an audited set of discovery-related arguments may be forwarded. Arbitrary project arguments such as compiler plugins, -Xclang -load, specifications files, or project-controlled program search paths could turn execution of an authorized compiler into execution of untrusted code.

The parser should extract and validate:

  • the effective target reported by the driver;
  • the ordered directories between the GCC-compatible system-include markers;
  • the compiler builtin-header directory reported by -print-file-name=include.

Include directories must be absolute and normalized. Their reported order is semantically significant and must be preserved. The queried compiler's builtin-header directory should be removed from the discovered include list because scalps must continue using the builtin headers coupled to its embedded Clang frontend.

This contract recovers the target and system-header configuration needed by the frontend. It does not claim to reproduce arbitrary flags injected by every possible wrapper. Drivers that do not provide GCC-compatible discovery output remain unsupported unless a future query backend defines another explicit, testable protocol.

Effective Clang commands

One component should construct the deterministic arguments passed to Clang:

  1. Recorded arguments after response-file expansion.
  2. The queried target when the recorded command has no explicit target.
  3. Queried include directories as ordered -isystem pairs.
  4. Scalps' configured Clang resource directory.

Derived arguments must be inserted before a -- terminator. Explicit targets retain precedence. The existing resource-directory behavior remains independent from compiler querying.

Hashing, persistence, and freshness

Derived arguments must participate in compilation-command identity rather than being appended only during extraction. The input-manifest and storage schema should be extended to record, for every translation unit:

  • the recorded-command hash;
  • the effective-command hash;
  • the resolved queried-driver path, when present;
  • the compiler-query protocol version;
  • a hash of the derived target and ordered include list.

The effective-command hash should identify reusable translation-unit artifacts. A change in wrapper environment, resolved compiler, target, or discovered include list must therefore invalidate the affected translation units.

Compiler queries should be cached only within one invocation. Later index and status runs must query again so current wrapper behavior is compared with the stored configuration.

Freshness behavior should follow these rules:

  • An unchanged recorded command and unchanged derived configuration are fresh.
  • A changed query result makes only the affected translation units stale.
  • If an index used compiler querying but status lacks current authorization to repeat it, freshness is unknown and the driver is not executed implicitly.
  • A failed authorized query makes status unknown and makes index fail before staging or replacing an index.
  • A prior queried index should not silently be replaced by an unqueried generation. Disabling discovery should require an explicit clean-rebuild decision.

Sessions and commits

Session 1: contract and secure query core

  1. Stage trusted compiler configuration discovery
    • Add the roadmap phase, supported GCC-compatible contract, trust boundary, and explicit non-goals.
  2. Add query-driver authorization and resolution
    • Implement absolute glob validation, compiler resolution, canonical matching, query-key construction, and diagnostics.
  3. Query GCC-compatible compiler configuration
    • Add bounded shell-free execution, output parsing, target validation, builtin-directory removal, per-run caching, and fake-driver tests proving that untrusted drivers are never executed.

Session 2: effective commands and extraction

  1. Prepare effective compilation commands
    • Introduce an explicit prepared-command representation, centralize response-file expansion, and construct the final argument vector in one place.
  2. Apply queried toolchains during extraction
    • Pass prepared arguments into the one-command Clang database while preserving resource-directory handling and explicit-target precedence.
  3. Extract through a hidden-include compiler wrapper
    • Add an integration fixture whose header is available only through queried wrapper configuration, covering include order, target propagation, standard-include suppression, malformed output, process failure, timeout, and caching.

Session 3: persistence and incremental correctness

  1. Persist effective command provenance
    • Bump the storage schema and input-manifest version, store recorded and effective hashes plus query fingerprints, and extend storage validation and corruption tests.
  2. Include queried configuration in freshness
    • Add query-specific invalidation reasons and ensure a changed query result selectively re-extracts the commands that share it.
  3. Cover queried incremental index lifecycle
    • Verify no-op reuse, changed query output, missing authorization, failed probes, multiple commands for one source, and preservation of the prior complete index after failures.

Session 4: command-line behavior and documentation

  1. Expose trusted query-driver options
    • Add the repeatable option to index, status, help, diagnostics, and verbose summaries of resolved drivers and derived configuration.
  2. Cover query-driver CLI security and status
    • Test unmatched patterns, bare compiler resolution, quoting, shell metacharacters, mixed queried and unqueried commands, freshness diagnostics, and exit statuses.
  3. Document compiler configuration discovery
    • Update the README, indexing, installation, support, and security guidance, including the risks of broad allowlist globs and the need to authorize queries again for status.

Session 5: Nix verification

  1. Run tests through compiler-wrapper discovery
    • Supply the build's exact wrapper path to the relevant extraction and CLI tests, re-enable the test suite in the flake package, and remove the obsolete statement that Nix package tests cannot run.
  2. Verify wrapped and explicit commands agree
    • Compare indexes produced from wrapper-dependent commands with equivalent commands containing explicit toolchain arguments and require the complete existing and new test suites to pass in both supported and Nix builds.

Acceptance criteria

  • Existing behavior is unchanged without --query-driver.
  • No executable from compile_commands.json runs without an explicit match against its resolved path.
  • Wrapper-dependent extraction and CLI tests pass through the real Nix compiler wrapper.
  • Changes to derived targets or include paths make status stale and selectively invalidate affected artifacts.
  • Query failures never replace the prior complete index.
  • Scalps' configured Clang resource directory remains authoritative.
  • Documentation states that the facility supports GCC-compatible target and include discovery, not arbitrary wrapper emulation.

Coordinated index access

Index publication is already atomic: a full or incremental update builds a complete generation in a temporary database beside the destination, then installs it with a same-directory rename. Readers therefore never open a half-written SQLite file. That guarantee does not coordinate concurrent scalps processes.

Without coordination, two index commands that target the same resolved index path can race: both may plan against the prior generation, both may extract and stage, and whichever rename runs last wins. A reader that opens the prior complete generation during a refresh cannot tell that a replacement is actively being produced, so it may report freshness or search results against data that is already known to be outdated for the current sources.

Goals

  • Allow only one refresh at a time for a given resolved index path.
  • Let status, search, and context detect an active refresh and report it instead of treating the prior generation as the current truth.
  • Preserve atomic publication and the rule that failed or interrupted updates leave the prior complete index in place.
  • Release the coordination lock when the holding process exits or crashes, so a leftover sidecar file cannot permanently block the index.
  • Keep distinct resolved index paths independent; locking one path must not affect another.

Lock contract

Coordination uses a persistent sidecar lock derived from the resolved index path: if the index is .scalps/index.sqlite, the lock file is .scalps/index.sqlite.lock. The sidecar remains on disk after use. The operating system releases the advisory exclusive lock when the process exits or crashes; an unlocked leftover file is harmless and is not treated as an active refresh.

Locks are whole-file advisory locks acquired through the supported system API (not shell commands or PID-file cleanup). Acquisition must distinguish contention from path, permission, and other system failures.

Command Behavior when indexing is active
index Exit immediately with an “indexing already in progress” error
status Print Indexing is in progress. and return status 1
search Leave stdout empty, report the active refresh on stderr, return status 2
context Report the active refresh on stderr and return status 2

index resolves and validates the index path, creates its parent directory, and acquires the exclusive lock before loading compilation commands or inspecting the prior generation. It holds the lock through planning, extraction, no-op completion, publication, and every failure path. Readers probe the sidecar before the existing-index check so a concurrent initial build reports active indexing rather than “index does not exist.” Machine-readable search keeps the rule that failures leave stdout empty.

Failing fast avoids introducing another untracked long-running command session. Atomic replacement still protects a reader that opened the prior generation immediately before an indexer acquired the lock.

Sessions and commits

Session 1: lock contract and core

  1. Stage coordinated index access
    • Add this roadmap phase, including atomic publication, the concurrent-writer race, lock location, command behavior, exit statuses, and crash handling.
  2. Add an index refresh lock
    • Add a small RAII lock component backed by the system advisory-lock API, with tests for contention, independent paths, normal and process-exit release, and invalid lock paths.

Session 2: CLI coordination

  1. Guard concurrent index refreshes
    • Acquire the exclusive lock for the full index command body after path validation and parent creation.
  2. Report indexing in progress to readers
    • Probe the sidecar from status, search, and context before opening the index.
  3. Cover concurrent command behavior
    • Hold a real index process after lock acquisition and exercise concurrent commands, including explicit --index paths and JSON search.

Session 3: user and agent workflow

  1. Document coordinated index access
    • Update the README, indexing lifecycle documentation, help text, and exit-status descriptions.
  2. Make the scalps skill await indexing
    • Require direct scalps index invocation, poll until completion, and forbid treating a yielded process as finished.

Acceptance criteria

  • Only one scalps index process can operate on a given index path.
  • Readers report an active refresh instead of using the previous generation unknowingly.
  • Existing atomic publication and failure-preservation tests continue to pass.
  • A crash cannot leave the index permanently locked.
  • Different index paths do not block one another.
  • The skill cannot mistake a yielded indexing process for a completed one.
  • Full build, targeted concurrency tests, and the complete ctest suite pass.

Parallel extraction

Indexing currently parses translation units one at a time. On this repository, a clean full rebuild spends about 52 seconds extracting 43 independent translation units, making frontend parsing the main indexing bottleneck.

Each translation-unit extraction already owns its compilation database, Clang frontend, matchers, diagnostics, and dependency capture. Parallel extraction should preserve that isolation boundary and add a bounded scheduler around the existing unit operation. Results must still be composed in compilation-command order so worker completion order cannot change counters, appearances, origins, dependencies, command indexes, exclusions, or diagnostics.

The public library remains sequential by default. A job count of zero selects at least one available hardware thread, while a positive count is capped by the number of commands. The command line will use automatic selection by default and offer an explicit one-job escape hatch for lower memory use. Empty, single-command, and one-job inputs should take a direct sequential path without starting worker threads.

The scheduler claims command indexes and reports progress under one mutex, then runs each isolated extraction outside the mutex and stores its result at the claimed index. Workers must keep claiming work after an unexpected unit or progress-callback exception. After every worker joins, the scheduler reports the exception for the lowest command index. Failure while creating threads must stop further claims and join every thread that was already created.

CompilerQueryService must be safe for concurrent callers before the worker pool reaches production. Its successful and failed caches and process-start counter require synchronization, and equal query keys must not launch duplicate probes. Because a cache miss forks a query process while extraction workers exist, every C++ value and raw argument pointer needed by the child must be prepared before fork(). The child may perform only the audited descriptor, directory, exec, and _exit operations.

clang::tooling::AllTUsToolExecutor is not suitable for this work. It assumes a shared compilation database and does not preserve scalps' prepared driver-query arguments, per-unit dependency capture, diagnostics, progress contract, or incremental subset database. A fallback that serializes every tool.run(...) call also does not complete this phase because AST parsing must actually overlap.

This phase does not add preamble or module reuse, narrow the matcher set, change incremental FTS updates, alter the schema, extractor, or ranking versions, modify the refresh lock, or add an automatic RAM-based job limit. Existing completeness checks, incremental reuse, query-driver behavior, atomic publication, and the exclusive refresh lock remain unchanged.

Sessions and commits

Session 1: concurrency foundations

  1. Stage parallel extraction
    • Record the bottleneck, isolated unit boundary, deterministic ordering, worker-pool contract, query-process prerequisite, rejected alternatives, and non-goals.
  2. Make compiler queries safe after fork
    • Prepare all child-used argument storage before fork() and test process launch while unrelated threads are alive.
  3. Make the compiler query cache thread-safe
    • Synchronize successful and failed cache access plus process-start accounting, and cover concurrent successful, failed, and unauthorized queries.
  4. Add an ordered extraction scheduler
    • Add a private injectable scheduler with bounded workers, ordered progress, deterministic exception selection, and guaranteed joins, plus focused scheduler tests.
  5. Extract translation units on the worker pool
    • Add the library job-count parameter, connect the scheduler to unit extraction, and link thread support from the core library in test and non-test builds.

Session 2: extraction identity and incremental indexing

  1. Cover parallel extraction identity
    • Compare complete extraction results across sequential, fixed parallel, and automatic job counts, including malformed and response-file failures.
  2. Use the extraction pool for incremental updates
    • Replace the duplicate incremental loop with pooled extraction over the ordered subset database while preserving all subset metadata.
  3. Cover parallel incremental publication
    • Exercise every replacement class in parallel and preserve the prior complete generation after failed extraction.

Session 3: command line and user contract

  1. Expose --jobs on index
    • Add index-only --jobs N and -j N parsing, automatic CLI selection, validation, and explicit rejection by other commands.
  2. Cover --jobs command behavior
    • Test valid, repeated, missing, malformed, zero, negative, overflow, and out-of-scope uses, then compare indexes from sequential, parallel, and automatic runs.
  3. Document parallel extraction
    • Document automatic jobs, the sequential escape hatch, ordered progress, memory scaling, and unchanged no-op, completeness, lock, and publication behavior.

Acceptance criteria

  • Two requested workers demonstrably overlap unit operations, while empty, single-command, and one-worker inputs start no worker threads.
  • Job counts zero and greater than one produce the same complete extraction result as sequential extraction.
  • Progress is reported once per command in compilation-command order.
  • Concurrent compiler-query callers execute each successful key's process pair once and replay one cached failure without another process.
  • No C++ allocation or locking remains in the query child between fork() and execv().
  • Unexpected worker failures join all threads and select the lowest failing command index.
  • scalps index defaults to automatic jobs, and --jobs 1 preserves sequential extraction.
  • status, search, and context reject both jobs option spellings.
  • Incremental completeness, refresh locking, failure preservation, and atomic publication remain intact.
  • A non-test build configures and links with thread support owned by scalps_core.
  • Targeted tests and the complete test suite pass; supported ThreadSanitizer checks have no race in scalps code.

Potential future enhancements

Current scalps roughly works like this:

query words → literal/identifier/FTS candidates → structural ranker → results

Future retrieval could add another candidate source while keeping the existing ranker’s strong preference for exact names, matching operations, and structural evidence.

Fuzzy retrieval

Fuzzy retrieval tolerates small differences between the query and indexed terms. Techniques could include edit distance, prefix matching, token similarity, or broader synonym normalization.

Examples:

  • serialise response could find SerializeResponse
  • initialze parser could survive the typo and find InitializeParser
  • completion popup might find code using autocomplete list
  • disconnect client might find CloseConnection

This would offer better typo tolerance and help when the user remembers an approximate name. It is comparatively simple, local, and explainable. Its main risk is producing many superficially similar candidates, especially in a large codebase.

Learned retrieval

Learned retrieval uses examples of successful searches to recognize relationships not encoded in fixed rules. The roadmap’s phrase “from observed workflows” suggests collecting data such as:

  • the query someone entered;
  • which result they opened or accepted;
  • unsuccessful queries followed by a revised query;
  • recurring query-to-definition associations.

That information could support several implementations:

  • Learn project-specific aliases, such as dialog → Window or buffer → Document.
  • Train or tune a candidate-scoring model.
  • Use embeddings to retrieve definitions whose descriptions or implementations express a similar idea with different words.
  • Learn which fields are useful for particular kinds of query.

For example, after observing that developers searching for refresh syntax colors consistently choose StyleNeeded, the system could retrieve that definition even though few query words appear literally in its name.

This would offer broader natural-language matching and adaptation to a project’s vocabulary. The costs are greater complexity, training data requirements, reduced explainability, model storage and runtime expense, and possible privacy concerns around recording searches.

Why they should only add candidates

The roadmap explicitly says these methods should not override strong literal and structural evidence. A likely design is:

literal candidates ─┐
fuzzy candidates   ─┼─ union/deduplicate → existing evidence-based ranking
learned candidates ─┘

Thus, an exact SetWrapMode query should still strongly favor SetWrapMode. Fuzzy or learned retrieval would mainly rescue definitions that never entered today’s candidate set; it would not promote a vague model guess above an exact structural match.

In short: fuzzy retrieval handles approximate spelling and wording, while learned retrieval handles conceptual and project-specific vocabulary gaps. Both would primarily improve recall—the chance that the correct definition is considered at all—while the current ranker protects result precision. The roadmap wisely leaves their exact form open until real search failures and performance costs can be measured.