Markdown chunking with single-call LLM enrichment or ultra-fast spaCy processing for RAG pipelines. Paper: MDKeyChunker: Single-Call LLM Enrichment with Rolling Keys and Key-Based Restructuring for High-Accuracy RAG (arXiv:2603.23533)
- Chunks Markdown into structural units (headers, code blocks, tables, lists)
- Enriches each chunk with ONE call → title, summary, keywords, entities, questions, semantic key
- LLM Mode: High-fidelity metadata using GPT/Claude/Ollama
- spaCy Mode: Fast, local, and free semantic key extraction
- Passes rolling keys forward so the enricher has context about prior topics
- Restructures by merging chunks that share the same specific-subtopic key
git clone https://github.com/bhavik-mangla/MDKeyChunker.git
cd MDKeyChunker
pip install -e . # add [anthropic] or [spacy] extras as needed
cp .env.sample .env # then edit .env with your API key
mdkeychunker demo.mdspaCy mode needs the extra and a model: pip install -e ".[spacy]" && python -m spacy download en_core_web_md.
Or programmatically:
from mdkeychunker import Pipeline, Config
# High-fidelity mode (GPT-4o-mini)
config = Config.from_env()
pipeline = Pipeline(config, enricher_mode="llm")
chunks = pipeline.process_file("document.md")
# Ultra-fast / Free mode (spaCy)
pipeline_fast = Pipeline(config, enricher_mode="spacy")
chunks_fast = pipeline_fast.process_file("document.md")# Basic usage
mdkeychunker document.md
# Save to specific output
mdkeychunker document.md -o chunks.jsonl
# With summary file and stats
mdkeychunker document.md --summary summary.txt --stats
# Disable merging
mdkeychunker document.md --no-merge
# Override LLM provider (e.g., local Ollama)
mdkeychunker document.md --provider openai_compatible --base-url http://localhost:11434/v1 --model llama3Set via .env file or environment variables:
| Variable | Default | Description |
|---|---|---|
LLM_PROVIDER |
openai |
openai | anthropic | openai_compatible |
LLM_API_KEY |
— | API key (not needed for local Ollama) |
LLM_BASE_URL |
— | Base URL for Ollama/vLLM/LM Studio |
LLM_MODEL |
gpt-4o-mini |
Model name |
MIN_CHUNK_SIZE |
100 |
Minimum characters per chunk |
MAX_CHUNK_SIZE |
1500 |
Soft max characters per chunk |
MERGE_BY_KEYS |
true |
Merge chunks sharing the same key |
MAX_MERGED_SIZE |
3000 |
Max combined size after merging |
MIN_ORPHAN_SIZE |
200 |
Below this, orphan chunks get context enrichment |
LOG_LEVEL |
INFO |
Logging verbosity |
# OpenAI
LLM_PROVIDER=openai
LLM_API_KEY=sk-...
LLM_MODEL=gpt-4o-mini
# Anthropic
LLM_PROVIDER=anthropic
LLM_API_KEY=sk-ant-...
LLM_MODEL=claude-3-haiku-20240307
# Ollama (local)
LLM_PROVIDER=openai_compatible
LLM_BASE_URL=http://localhost:11434/v1
LLM_MODEL=llama3Markdown → Chunker → Enricher (1 LLM call/chunk) → Restructurer → Enriched Chunks
↑
Rolling Keys
(context from prior chunks)
Key design: The key field means the specific subtopic that distinguishes a chunk — e.g. "admissions process", "oauth token flow", "gradient descent optimization" — never the broad document topic. Chunks with matching keys are merged globally, regardless of position.
{
"chunk_id": "a3f2b1c4d5e6f7a8",
"text": "The admissions process begins in March...",
"section_title": "Admissions",
"title": "Spring Admissions Timeline",
"summary": "Describes the March start of the admissions process...",
"keywords": ["admissions", "application deadline", "March intake"],
"entities": [{"name": "March", "type": "EVENT"}],
"questions": ["When does the admissions process begin?"],
"key": "admissions process",
"related_keys": ["curriculum framework"],
"content_types": ["paragraph"],
"position_index": 2,
"previous_chunk_id": "9b8c7d6e5f4a3b2c",
"next_chunk_id": "1a2b3c4d5e6f7a8b",
"token_count": 187,
"start_line": 12,
"end_line": 28
}The paper's evaluation (30 queries over an 18-document Markdown corpus, Configs A–D) is described in arXiv:2603.23533. The code version used in the paper is commit e3e1b86.
benchmarks/scifact.py runs a SciFact sanity check in spaCy mode (pip install -e ".[benchmark,spacy]", then python benchmarks/scifact.py). Embeddings are BAAI/bge-small-en-v1.5 with exact cosine search.
| Strategy | Library | Recall@5 | nDCG@5 | Chunks/Doc |
|---|---|---|---|---|
| MDKeyChunker (spaCy mode) | — | 0.762 | 0.681 | 1.00 |
| Recursive Character | LangChain | 0.760 | 0.678 | 3.43 |
| Semantic Chunker | LangChain | 0.775 | 0.681 | 2.03 |
| Fixed Token (512) | Standard | 0.762 | 0.681 | 1.05 |
Read this table with care. SciFact abstracts are short single-paragraph plain text, so MDKeyChunker emits one chunk per abstract and this run effectively measures whole-abstract retrieval. It does not exercise Markdown structure, LLM enrichment, or key-based restructuring, and it is not evidence for those stages. The baseline rows were produced with scripts that are not yet in this repository.
from mdkeychunker import Pipeline, Config, Chunk
# Config
config = Config(llm_provider="openai", llm_model="gpt-4o-mini", merge_by_keys=True)
# Pipeline
pipeline = Pipeline(config)
chunks: list[Chunk] = pipeline.process_text(markdown_text)
chunks: list[Chunk] = pipeline.process_file("path/to/file.md")
# Save output
pipeline.save_jsonl(chunks, "output.jsonl")
pipeline.save_summary(chunks, "summary.txt")pip install -e ".[dev]"
pytest tests/ -vIf you use MDKeyChunker, please cite:
@misc{mangla2026mdkeychunker,
title = {{MDKeyChunker}: Single-Call {LLM} Enrichment with Rolling Keys and Key-Based Restructuring for High-Accuracy {RAG}},
author = {Mangla, Bhavik},
year = {2026},
eprint = {2603.23533},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2603.23533}
}GitHub's "Cite this repository" button uses CITATION.cff.
MIT
