Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

obsidian-rag

Local-first hybrid retrieval over an Obsidian vault, exposed as an MCP server so Claude Code (or any MCP client) can search your notes.

Read this first: your vault is probably too small to need it.

If your vault is under roughly 50k tokens, the honest answer is that retrieval will make your results worse, not better — it can only remove information a model would otherwise have seen in full. This tool knows that and says so: vault_stats reports a recommended strategy, and below the threshold it tells you to outline the vault and read the two notes that matter instead of searching.

The retrieval pipeline is here for when you grow into it. vault_stats will tell you when that is.

What it does

  • Hybrid retrieval — BGE-M3 dense + BGE-M3 lexical (sparse) from a single forward pass, fused with Reciprocal Rank Fusion (c=60). RRF is scale-free, so a cosine in [-1,1] and an unbounded lexical dot product need no cross-channel normalisation.
  • Multilingual by construction — built for a mixed Chinese/English vault. Measured cross-lingual similarity on this vault's own terms: 碳盤查成本 ↔ "carbon inventory cost" = 0.71, versus 0.32 against an unrelated Chinese phrase. BM25 is deliberately not the sparse channel: it needs word segmentation, and a whitespace tokenizer treats an entire Chinese paragraph as a single token.
  • Cross-encoder reranking with bge-reranker-v2-m3, driven directly through transformers rather than FlagEmbedding's reranker (which calls a tokenizer method current transformers no longer exposes).
  • Structure-aware, CJK-safe chunking — heading-aware with breadcrumb prefixes; files that are really N independent records are detected and each record kept whole; separator ladder carries full-width punctuation; token counts come from a tokenizer, never from character counts.
  • Prompt-injection containment — trust assigned at ingest, filtered in SQL, and every returned string passes a single framing chokepoint. Tested as a contract, not asserted as a claim.
  • Incremental indexing — content-hash based; renaming a note re-embeds nothing.
  • Honest evaluation — metrics stratified by query/document lexical overlap, because generated questions echo their source; unverified items labelled as such; the report refuses to imply that a heuristically generated dataset proves anything.
  • No Docker, no server, no cloud — SQLite plus a memory-mapped numpy matrix.

Install

git clone https://github.com/CYC2002tommy/obsidian-rag
cd obsidian-rag
uv venv --python 3.11
uv pip install -e .            # model-free: ingest + keyword search work now
uv pip install -e ".[local]"   # BGE-M3 + reranker (~3.4 GB of weights)
uv pip install -e ".[mcp]"     # the MCP server

Copy obsidian-rag.example.toml to obsidian-rag.toml and set your vault path.

uv run obsidian-rag doctor --trust   # environment + trust review
uv run obsidian-rag index            # walk, chunk, store
uv run obsidian-rag embed            # compute vectors
uv run obsidian-rag search "碳盤查成本" "carbon inventory cost" -k 5

GPU note

On an RTX 5070 (Blackwell, sm_120) the default PyPI torch wheel may ship without matching kernels and fall back to CPU silently — 10–50× slower, easy to miss. pyproject.toml pins the CUDA 12.8 index, and doctor compares your device capability against the wheel's arch list and warns loudly on a mismatch.

MCP

"obsidian-rag": {
  "type": "stdio",
  "command": "/abs/path/obsidian-rag/.venv/Scripts/python.exe",
  "args": ["-m", "obsidian_rag.mcp_server"],
  "env": {
    "OBSIDIAN_RAG_CONFIG": "/abs/path/obsidian-rag/obsidian-rag.toml",
    "HF_HUB_DISABLE_PROGRESS_BARS": "1",
    "TRANSFORMERS_VERBOSITY": "error",
    "TOKENIZERS_PARALLELISM": "false"
  }
}

Those env vars are not cosmetic: inside an MCP server stdout is the JSON-RPC channel, and one stray progress bar corrupts the protocol. Model loading is lazy, so vault_stats, vault_outline, get_note and vault_dump never touch a model and answer instantly on a cold start.

Tools

tool purpose
vault_stats corpus size, trust breakdown, recommended strategy — call first
vault_outline every note's path/title/tags/headings, no bodies
search_vault hybrid search; takes several query variants, fused by RRF
get_note full note after search located it
vault_dump everything, or a hard refusal above the limit — never silent truncation
index_vault incremental refresh

search_vault takes queries: string[] deliberately. The calling model already knows the conversation, the user's terminology, and both of a bilingual vault's languages — asking a small local model to rewrite the query would be strictly worse. It also dissolves the "intersection of zero" failure, where four entities in one query match nothing while four atomic queries fused by RRF return the union.

Security

A vault that mirrors AI prompts, agent configs or scraped instructions contains text that reads as a system message when injected into a model's context. The vault this was developed against holds nineteen such files.

Three layers, all tested:

  1. Exclusion. Trust is assigned at ingest and filtered in SQL, so a defect in the retrieval pipeline cannot leak quarantined text. Ships quarantining **/system_prompts/**, **/prompts/**, **/CLAUDE.md, **/AGENTS.md.
  2. Escaping. Anything not fully trusted has its markup escaped — <system-reminder> arrives inert.
  3. Containment. Framing delimiters are nonce-suffixed and stripped from content, so a document cannot close its own block and address the model.

The classifier is heuristic, needs no network, and can only escalate. It is overridden by an explicit trusted_globs entry — which exists because a user's own imperative prose scores high, and losing your most important notes to a false positive is its own kind of failure.

Status

Working: config, ingest, trust classification, chunking, SQLite store, CJK segmentation, incremental indexing, BGE-M3 embeddings, hybrid retrieval with RRF, cross-encoder reranking, MCP server, injection containment, evaluation harness.

Not done: LLM-backed evaluation query generation (the heuristic generator is a stand-in and the report says so), C-RAG style retrieval correction, watch mode.

License

MIT

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages