Knowledge-graph memory for agents. A local model turns documents into subject relation object facts in one non-autoregressive pass (no LLM tokens), and the cogito-mcp server
answers questions from that graph and the sentences behind it, so an agent reads a few
hundred tokens instead of a whole document.
pip install "cogito-estella[mcp]" # add [pdf] for PDF inputClient configuration (Claude Code .mcp.json, Cursor, any MCP client):
{"mcpServers": {"cogito": {"command": "cogito-mcp", "args": ["--dir", ".cogito"]}}}On first run the server downloads the spaCy model and the default weights from
Hugging Face: Meta's frozen M2M-100 418M
encoder (MIT, ~1.9 GB) under a learned pooling, and three decoder heads (Apache-2.0).
Python 3.12; a GPU is optional (--device cpu). The graph lives in .cogito/graph.json;
add that directory to .gitignore.
ingest(path_or_text)— a file, a directory (txt/md/html/pdf) or raw text. Documents are deduplicated by content hash and replaced when they change.ask(question, budget=600, use_dense=False)— start here. The graph facts for the entities in the question, a--line, then the source sentences that answer it, withinbudgettokens (use_sonaris the older name of the flag). Sentences are shortlisted by IDF and ordered by a learned relevance scorer;use_dense=Trueranks them by encoder cosine instead.query(entity, hops, limit),provenance(edge_ids),search(term),entities(prefix)andstats()go deeper when oneaskis not enough. Facts listed under~ class-only, verify with provenance:carry only a coarse relation label.
entities: blt, layer · scorer=learned
blt encode byte #12
--
2412.09871.txt s112: "We use SwiGLU activation in the feed-forward layers, as in Llama 3."
Flags: --encoder sonar for the SONAR-space heads (research only: pip install "cogito-estella[sonar]", CC-BY-NC 4.0 at runtime); --checkpoint, --vocab and --pool
for local weights; --device; --no-download. COGITO_ASK_SCORER=lexical keeps the
plain IDF order.
50 questions over five arXiv papers, answered by Claude Code headless (Haiku) under the same rules in every branch, judged blind to the branch, with the Claude Code prefix calibrated the same day:
| branch | accuracy | raw input tokens / question | net tokens / question |
|---|---|---|---|
| whole paper in the prompt | 1.00 | 38,231 | 12,810 |
native Read/Grep on the file |
0.98 | 89,493 | 4,753 |
one ask reply (600-token budget) |
0.78 | 26,126 | 613 |
Raw tokens are what the API bills before caching, measured end to end; the ~25k-token
Claude Code prefix dominates every branch. Net tokens subtract that calibrated prefix and
are an estimate. One ask reply answers 78 % of the questions (abstaining on 16 %) at
1.46× fewer raw tokens and about 21× fewer net tokens than pasting the paper; the
Read/Grep agent is accurate but costs 2.3× the raw tokens of pasting the paper.
Extraction quality: the default three-head ensemble reaches Triple F1 0.852 on held-out sentences whose entity/relation combinations never appeared in training; on 5,000 sentences held out under both protocols it scores 0.819 against 0.730 for the SONAR-space ensemble.
A single ask is the cheapest route but abstains on some questions. The header's p=
says how sure the scorer is, so a caller can escalate before answering; both escalation
policies were measured on the same 50 questions, reusing the rows above:
| policy | accuracy | context tokens / question | per correct answer |
|---|---|---|---|
one ask at 600 |
0.78 | 588 | 754 (15.0×) |
ask → read the document when it abstains |
0.94 | 2,353 | 2,503 (4.5×) |
ask → ask at 1200 when p < 0.3 → document |
0.94 | 1,380 | 1,468 (7.7×) |
ask → document when p < 0.3 or it abstains |
0.96 | 2,486 | 2,590 (4.4×) |
A second ask spends no LLM tokens, so escalating the budget is far cheaper than reading
the document: it leaves 5 documents read out of 50 instead of 8. Sending low-confidence
questions straight to the document instead is the accurate end of the trade: it answers
0.96 and leaves two answers wrong-without-warning rather than three.
from cogito_estella.integrations.llamaindex_connector import CogitoGraphExtractor
from cogito_estella.mcp.weights import resolve
w = resolve(None, None) # the default assets, downloaded once
ex = CogitoGraphExtractor([str(c) for c in w.checkpoints], str(w.vocab), pool_path=w.pool,
threshold=w.operating_point[0], adj_threshold=w.operating_point[1])
ex.extract("The committee approved the new budget.") # [(subject, relation, object), ...]extract_with_literals keeps exact literals (IDs, hashes, amounts) verbatim and
literals_to_neo4j stores them with provenance; to_neo4j writes triples into Neo4j. Call
ensure_schema(driver) once per database before concurrent writers: its uniqueness
constraints make parallel MERGEs deterministic, and they cannot be created while the
database still holds duplicate nodes. cogito_estella.graph_summary.GraphSummarizer
writes community summaries over the triples and flags (accepted=False) the ones that
mention entities absent from their cluster.
Each head is a compact non-autoregressive decoder (38M parameters) on a frozen sentence encoder: the candidate entities found in the sentence act as queries over its vector, and existence plus a 76-relation adjacency come out in one forward pass (about 0.01 ms per sentence on an RTX 5070). Relations are named from the sentence's dependency syntax when a pattern is recognized; otherwise the fact keeps the class label only. Every edge points back to its source sentence and character spans.
| Encoder | Assets on the Hub | Licence |
|---|---|---|
m2m100-pool (default) |
m2m100-pool/cogito-prose-ontology-m2mpool{,-s2,-s3}.safetensors, pool.safetensors, vocab-onto-m2mpool.json; manifest encoders.json |
heads and pooling Apache-2.0; M2M-100 encoder MIT |
sonar |
cogito-prose-ontology{,-s2,-s3}.pt, vocab-onto.json |
heads Apache-2.0; the SONAR encoder they need is CC-BY-NC 4.0 |
Every checkpoint carries its encoder, revision, width and normalization; a checkpoint decoded in the wrong space stops the server, and a startup canary re-checks the encoder against shipped reference cosines.
Three standalone research heads (not part of the default ask/ingest path) are also
ported to the permissive encoder, each measured above its SONAR original: tool-call
extraction (Triple F1 1.000), entity-conditioned prose on the 60-verb vocabulary
(0.878 vs 0.827, 5-model ensemble, CogitoGraphExtractor loads it unchanged), and code
call/import extraction (0.979 vs 0.777) — the last via a LoRA-adapted M2M-100 encoder,
so it too carries no non-commercial term. Details and a loader in the Hub model card.
uv sync --all-extras
uv run python -m spacy download en_core_web_sm
uv run pytest tests/573 tests with every extra installed; modules that need the [sonar] extra and the tests
that load a real encoder on CUDA skip themselves without them. Package layout:
model/ (decoder heads), encoders/ (encoder contract, adapters, pooling, canary),
mcp/ (store, readers, scorers, server), integrations/ (the extractor). Training data,
experiments and logs are untracked; see CHANGELOG.md for the version history.
Apache License 2.0 (see LICENSE). Weight licensing is described above.
