diff --git a/assets/media/Sirchmunk_Architecture.png b/assets/media/Sirchmunk_Architecture.png index ac86d1e..7346622 100644 Binary files a/assets/media/Sirchmunk_Architecture.png and b/assets/media/Sirchmunk_Architecture.png differ diff --git a/assets/media/Sirchmunk_Knowledge_Graph.png b/assets/media/Sirchmunk_Knowledge_Graph.png new file mode 100644 index 0000000..f53916d Binary files /dev/null and b/assets/media/Sirchmunk_Knowledge_Graph.png differ diff --git a/assets/media/Sirchmunk_LENS_Framework.png b/assets/media/Sirchmunk_LENS_Framework.png new file mode 100644 index 0000000..7346622 Binary files /dev/null and b/assets/media/Sirchmunk_LENS_Framework.png differ diff --git a/assets/media/Sirchmunk_MonteCarloSamplingAlgo.png b/assets/media/Sirchmunk_MonteCarloSamplingAlgo.png index aca393f..7346622 100644 Binary files a/assets/media/Sirchmunk_MonteCarloSamplingAlgo.png and b/assets/media/Sirchmunk_MonteCarloSamplingAlgo.png differ diff --git a/config/_default/menus.yaml b/config/_default/menus.yaml index 940b1da..f17dfed 100644 --- a/config/_default/menus.yaml +++ b/config/_default/menus.yaml @@ -17,6 +17,11 @@ main: - name: Community url: community/ weight: 50 + - name: Paper + url: "https://arxiv.org/abs/2608.16185" + weight: 55 + params: + external: true sidebar: - identifier: more diff --git a/config/_default/menus.zh.yaml b/config/_default/menus.zh.yaml index b6af600..b4e6544 100644 --- a/config/_default/menus.zh.yaml +++ b/config/_default/menus.zh.yaml @@ -16,6 +16,11 @@ main: - name: 社区 url: community/ weight: 50 + - name: 论文 + url: "https://arxiv.org/abs/2608.16185" + weight: 55 + params: + external: true sidebar: - identifier: more diff --git a/content/_index.md b/content/_index.md index 1ea2bcf..85b8a82 100644 --- a/content/_index.md +++ b/content/_index.md @@ -17,10 +17,10 @@ sections: url: docs/getting-started/ icon: rocket-launch secondary_action: - text: Read the Technical Report - url: blog/technical-deep-dive/ + text: Read the Paper + url: https://arxiv.org/abs/2608.16185 announcement: - text: "Sirchmunk v0.0.8 — Knowledge Compile (Beta), DEEP Mode Generalization & I/O Optimization" + text: "Sirchmunk v0.2.0 — LENS Paper on arXiv · Multi-Path DEEP Retrieval · Large Corpus Robustness" link: text: "View all releases" url: "https://github.com/modelscope/sirchmunk/releases" @@ -62,22 +62,22 @@ sections: items: - name: Embedding-Free Retrieval icon: magnifying-glass - description: "Work directly with raw data. No vector database, no pre-indexing, no ETL pipeline. Just drop your files and search immediately." + description: "Work directly with raw data — no vector database, no pre-indexing, no ETL pipeline. Drop your files and search immediately with full source fidelity." - name: Self-Evolving Knowledge icon: arrow-path - description: "Knowledge clusters compound with every search. The system learns and improves over time, delivering faster and richer results." - - name: Monte Carlo Evidence Sampling + description: "Every search produces a reusable KnowledgeCluster. Clusters merge, broaden, and form meta-communities over time — the system literally gets smarter as you use it." + - name: "LENS: Latent Evidence Exploration" icon: chart-bar - description: "Strategically sample documents using exploration-exploitation methods. Extract precise evidence without reading entire files." - - name: ReAct Agent Fallback + description: "Budgeted evidence localization over a query-conditioned latent evidence space. The LENS framework locates source-grounded evidence from raw dynamic documents under explicit cost constraints." + - name: Multi-Path DEEP Retrieval icon: cpu-chip - description: "When standard retrieval falls short, an autonomous ReAct agent iteratively explores alternative strategies until answers are found." + description: "Parallel lexical, entity, directory, structural, and topic-graph retrieval routes fused by confidence-weighted RRF — with soft route-collapse for high-confidence single-file lookups." + - name: Large Corpus Robustness + icon: shield-check + description: "Bounded per-file and per-query retrieval cost: tiered rg-first scan, adapter whitelist, file-size cap, per-file match limits, and hard token budgets keep huge corpora fast." - name: Multi-Surface Integration icon: globe-alt - description: "MCP protocol, OpenClaw skill, REST API, WebSocket real-time chat, CLI, and a modern Web UI — all built in." - - name: Token-Efficient Design - icon: bolt - description: "LLM inference triggered only when necessary. Monte Carlo sampling and knowledge reuse minimize costs while maximizing intelligence." + description: "MCP protocol, OpenClaw skill, REST API, WebSocket real-time chat, CLI, and a modern Web UI with knowledge graph visualization — all built in." - block: cta-card content: title: "Start Searching with Sirchmunk" diff --git a/content/_index.zh.md b/content/_index.zh.md index 79064d5..b1f2cb1 100644 --- a/content/_index.zh.md +++ b/content/_index.zh.md @@ -17,10 +17,10 @@ sections: url: docs/getting-started/ icon: rocket-launch secondary_action: - text: 阅读技术报告 - url: blog/technical-deep-dive/ + text: 阅读论文 + url: https://arxiv.org/abs/2608.16185 announcement: - text: "Sirchmunk v0.0.8 — 知识编译(Beta)、DEEP 模式泛化增强与 I/O 优化" + text: "Sirchmunk v0.2.0 — LENS 论文已发布 · 多路 DEEP 检索 · 大语料鲁棒性" link: text: "查看所有版本" url: "https://github.com/modelscope/sirchmunk/releases" @@ -62,22 +62,22 @@ sections: items: - name: 无嵌入检索 icon: magnifying-glass - description: "直接处理原始数据。无需向量数据库、无需预索引、无需 ETL 管线。只需放入文件即可立即搜索。" + description: "直接处理原始数据 — 无需向量数据库、无需预索引、无需 ETL 管线。放入文件即可搜索,完整保留源数据保真度。" - name: 自进化知识 icon: arrow-path - description: "知识簇随每次搜索不断积累。系统持续学习与改进,搜索结果越来越快、越来越丰富。" - - name: 蒙特卡洛证据采样 + description: "每次搜索都生成可复用的 KnowledgeCluster。聚类随使用不断合并、拓展并形成元社区 — 系统在使用中持续变得更聪明。" + - name: "LENS:隐式证据探索" icon: chart-bar - description: "使用探索-利用策略对文档进行智能采样,无需阅读整个文件即可提取精确证据。" - - name: ReAct 智能体自适应检索 + description: "在查询条件下的隐式证据空间中进行预算约束证据定位。LENS 框架在显式成本约束下从原始动态文档中定位源可追溯证据。" + - name: 多路 DEEP 检索 icon: cpu-chip - description: "当常规检索不足时,ReAct 智能体自主迭代推理并探索替代检索策略,直到找到答案。" + description: "词法、实体、目录、结构与主题图多条并行检索路径,通过置信度加权 RRF 融合 — 高置信单文件查询触发 soft 路由收缩快通道。" + - name: 大语料鲁棒性 + icon: shield-check + description: "单文件与单次查询的检索成本均设界:rg 优先分层扫描、适配器白名单、文件大小上限、单文件匹配上限与硬 token 预算,确保大规模语料高效稳定。" - name: 多接口集成 icon: globe-alt - description: "MCP 协议、OpenClaw 技能,REST API、WebSocket 实时聊天、CLI 与现代 Web UI — 全部内置。" - - name: Token 高效设计 - icon: bolt - description: "仅在必要时触发 LLM 推理。蒙特卡洛采样和知识复用最大程度降低成本,同时最大化智能。" + description: "MCP 协议、OpenClaw 技能、REST API、WebSocket 实时聊天、CLI 与知识图谱可视化的现代 Web UI — 全部内置。" - block: cta-card content: title: "开始使用 Sirchmunk 搜索" diff --git a/content/blog/in-context-search/index.md b/content/blog/in-context-search/index.md index bfa12a8..bdc1f64 100644 --- a/content/blog/in-context-search/index.md +++ b/content/blog/in-context-search/index.md @@ -18,6 +18,8 @@ image: With the evolution of RAG (Retrieval-Augmented Generation) technology, a new paradigm called **In-Context Search (ICS)** is redefining how LLMs interact with external knowledge. This post compares traditional Graph-based RAG with next-generation ICS approaches represented by **[PageIndex](https://github.com/VectifyAI/PageIndex)** and **[Sirchmunk](https://github.com/modelscope/sirchmunk)**. +**Update (September 2026):** The theoretical foundations described in this article have been formalized in our research paper: [LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents](https://arxiv.org/abs/2608.16185) (arXiv:2608.16185). + ### Abstract @@ -47,8 +49,8 @@ The industry is now pivoting toward **In-Context Search (ICS)**. In this paradig | Dimension | **LightRAG** (Advanced Graph-RAG) | **PageIndex** (Reasoning-based ICS) | **Sirchmunk** (Indexless / Self-Evolving) | | --- | --- | --- |-------------------------------------------| | **Architectural Philosophy** | Index-Centric (Graph Topologies) | Reasoning-Centric (Hierarchical ICS) | Agile-Centric (Raw Search & Evolution) | -| **Primary Mechanism** | Dual-level Graph Traversal | Agentic Tree Navigation | Greedy Cascade (FAST) / Monte Carlo Sampling (DEEP) | -| **Retrieval Depth** | Global & Local via Graph Edges | Structural Pathfinding | Statistical Importance Extraction | +| **Primary Mechanism** | Dual-level Graph Traversal | Agentic Tree Navigation | Multi-path DEEP Retrieval (5 paths) / Greedy Cascade (FAST) | +| **Retrieval Depth** | Global & Local via Graph Edges | Structural Pathfinding | Confidence-weighted RRF Fusion + Statistical Importance Extraction | | **Indexing Overhead** | High (Graph Construction) | Moderate (Tree Metadata) | **Minimal to Zero** | | **Context Fidelity** | High (Entity-Relationship) | **Maximum** (Structural Integrity) | **Full Fidelity** (Raw Data Access) | | **Data Freshness** | Low (Re-indexing required) | Moderate (Incremental updates) | **Real-time** (Direct File Access) | @@ -101,9 +103,9 @@ PageIndex implements an **Agentic Loop** that mimics human research patterns—n Sirchmunk represents the **"Agile Hunter"** philosophy in the In-Context Search (ICS) landscape. It prioritizes **data freshness** and **operational speed**, completely bypassing the static tree-building phase. Instead, it treats the file system as a live, queryable environment, leveraging statistical mechanics and agentic reflection. -Sirchmunk operates in two distinct search modes. **FAST mode** (default) employs a greedy strategy with 2-level keyword cascade and context-window sampling, achieving retrieval in 2–5 seconds with only 2 LLM calls — a **~10x speedup** over the comprehensive mode. **DEEP mode** activates the full Monte Carlo evidence sampling pipeline with multi-round ReAct refinement for maximum recall on complex queries (10–30 seconds). +Sirchmunk operates in two distinct search modes. **FAST mode** employs a greedy strategy with 2-level keyword cascade and context-window sampling, achieving retrieval in 2–5 seconds with only 2 LLM calls — a **~10x speedup** over the comprehensive mode. **DEEP mode** (default since v0.1.0) runs five complementary retrieval paths (lexical, entity, directory, structural, topic-graph) fused via confidence-weighted Reciprocal Rank Fusion, followed by Monte Carlo evidence sampling and multi-round ReAct refinement for maximum recall on complex queries (10–30 seconds). -As of v0.0.6post1, Sirchmunk also ships as an OpenClaw skill — enabling any OpenClaw-compatible agent to invoke its search capability via natural language. From v0.0.6 onward, the stack further includes **multi-turn conversation** with context management, **document summarization**, and **cross-lingual retrieval** alongside the FAST/DEEP search modes above. +As of v0.2.0, Sirchmunk also ships as an OpenClaw skill — enabling any OpenClaw-compatible agent to invoke its search capability via natural language. The stack includes **multi-turn conversation** with context management, **document summarization**, and **cross-lingual retrieval** alongside the FAST/DEEP search modes above. --- @@ -155,7 +157,20 @@ Sirchmunk utilizes a "Post-hoc Indexing" strategy. It doesn't index before you a | **Just-in-Time Indexing** | Builds a dynamic map of data based on actual usage patterns. | Python-Native | | **Knowledge Reuse** | Hits the cache for similar future queries, evolving from brute-force to high-speed retrieval. | DuckDB SQL | -Through this mechanism, Sirchmunk evolves from a "brute-force hunter" into a "sophisticated librarian" organically, without the maintenance overhead of traditional pre-indexed databases. +Through this mechanism, Sirchmunk evolves from a “brute-force hunter” into a “sophisticated librarian” organically, without the maintenance overhead of traditional pre-indexed databases. + +--- + +### 3.5 Experimental Validation + +The LENS framework’s effectiveness has been validated through rigorous controlled evaluation: + +| Setting | LENS (Sirchmunk DEEP) | ReAct Baseline | +| --- | --- | --- | +| **500-question controlled eval** | 62.4% EM, 84.8% evidence recall | 65.2% EM, 50.4% evidence recall | +| **150-question fullwiki** (raw Wikipedia, zero indexing) | 43.3% EM, 84.0% evidence recall | 42.7% EM, 70.7% evidence recall | + +These results reveal a key insight: LENS trades a small margin on headline accuracy for dramatically better evidence grounding. In the fullwiki setting — where no indexing or preprocessing is performed — LENS achieves comparable EM while providing 13+ percentage points more evidence recall, validating the source-fidelity claims of the in-context search paradigm. --- @@ -231,7 +246,8 @@ The era of treating RAG as a static database lookup is ending. By embracing **In 1. Lewis, P., Perez, E., Piktus, A., et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.* NeurIPS 2020. [arXiv:2005.11401](https://arxiv.org/abs/2005.11401) 2. Guo, Z., Qian, C., et al. (2024). *LightRAG: Simple and Fast Retrieval-Augmented Generation.* [arXiv:2410.05779](https://arxiv.org/abs/2410.05779) | [GitHub](https://github.com/HKUDS/LightRAG) 3. VectifyAI. (2025). *PageIndex: Extracting and Understanding Financial Reports with LLM.* [GitHub](https://github.com/VectifyAI/PageIndex) -4. ModelScope. (2025). *Sirchmunk: An Embedding-Free, Agentic Search Engine for Raw Data.* [GitHub](https://github.com/modelscope/sirchmunk) -5. Yao, S., Zhao, J., Yu, D., et al. (2023). *ReAct: Synergizing Reasoning and Acting in Language Models.* ICLR 2023. [arXiv:2210.03629](https://arxiv.org/abs/2210.03629) -6. Anthropic. (2024). *Model Context Protocol (MCP) Specification.* [Documentation](https://modelcontextprotocol.io) -7. Kaddour, J., Harris, J., Mozes, M., et al. (2023). *Challenges and Applications of Large Language Models.* [arXiv:2307.10169](https://arxiv.org/abs/2307.10169) +4. ModelScope. (2026). *Sirchmunk: An Embedding-Free, Agentic Search Engine for Raw Data.* [GitHub](https://github.com/modelscope/sirchmunk) +5. Wang, X., et al. (2026). *LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents.* [arXiv:2608.16185](https://arxiv.org/abs/2608.16185) +6. Yao, S., Zhao, J., Yu, D., et al. (2023). *ReAct: Synergizing Reasoning and Acting in Language Models.* ICLR 2023. [arXiv:2210.03629](https://arxiv.org/abs/2210.03629) +7. Anthropic. (2024). *Model Context Protocol (MCP) Specification.* [Documentation](https://modelcontextprotocol.io) +8. Kaddour, J., Harris, J., Mozes, M., et al. (2023). *Challenges and Applications of Large Language Models.* [arXiv:2307.10169](https://arxiv.org/abs/2307.10169) diff --git a/content/blog/in-context-search/index.zh.md b/content/blog/in-context-search/index.zh.md index 963e3f9..513f915 100644 --- a/content/blog/in-context-search/index.zh.md +++ b/content/blog/in-context-search/index.zh.md @@ -18,6 +18,8 @@ image: 随着 RAG (Retrieval-Augmented Generation) 技术的演进,一种名为 **上下文搜索(In-Context Search, ICS)** 的新范式正在重新定义 LLM 与外部知识的交互方式。本文对比了传统 Graph-based RAG 与以 **PageIndex** 和 **Sirchmunk** 为代表的下一代 ICS 方案。 +**更新(2026年9月):** 本文所述理论基础已在研究论文中形式化:[LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents](https://arxiv.org/abs/2608.16185)(arXiv:2608.16185)。 + @@ -48,8 +50,8 @@ image: | 维度 | **LightRAG**(高级 Graph-RAG) | **PageIndex**(基于推理的 ICS) | **Sirchmunk**(无索引 / 自进化) | | --- | --- | --- |--------------------------| | **架构理念** | 索引中心(图拓扑) | 推理中心(层级化 ICS) | 敏捷中心(原始搜索与进化) | -| **核心机制** | 双层图遍历 | 智能体树导航 | 贪心级联(FAST)/ 蒙特卡洛采样(DEEP) | -| **检索深度** | 通过图边实现全局与局部检索 | 结构化路径搜索 | 统计重要性提取 | +| **核心机制** | 双层图遍历 | 智能体树导航 | 多路 DEEP 检索(5 条路径)/ 贪心级联(FAST) | +| **检索深度** | 通过图边实现全局与局部检索 | 结构化路径搜索 | 置信度加权 RRF 融合 + 统计重要性提取 | | **索引开销** | 高(图构建) | 中等(树元数据) | **极低至零** | | **上下文保真度** | 高(实体-关系) | **最高**(结构完整性) | **完全保真**(原始数据直接访问) | | **数据新鲜度** | 低(需重新索引) | 中等(增量更新) | **实时**(直接文件访问) | @@ -101,9 +103,9 @@ PageIndex 实现了一个 **智能体循环**,模拟人类的研究模式 — Sirchmunk 代表了上下文搜索(ICS)领域的 **"敏捷猎手"** 理念。它优先保障 **数据新鲜度** 与 **操作速度**,完全跳过静态树构建阶段。它将文件系统视为一个实时可查询的环境,借助统计力学和智能体反思机制进行搜索。 -Sirchmunk 提供两种搜索模式。**FAST 模式**(默认)采用贪心策略,结合两级关键词级联与上下文窗口采样,仅需 2 次 LLM 调用即可在 2–5 秒内完成检索,速度约为全面模式的 **10 倍**。**DEEP 模式** 则激活完整的蒙特卡洛证据采样管线,配合多轮 ReAct 自适应推理,适用于复杂查询场景下的最大召回(10–30 秒)。 +Sirchmunk 提供两种搜索模式。**FAST 模式** 采用贪心策略,结合两级关键词级联与上下文窗口采样,仅需 2 次 LLM 调用即可在 2–5 秒内完成检索,速度约为全面模式的 **10 倍**。**DEEP 模式**(自 v0.1.0 起为默认模式)并行运行五条互补检索路径(词法、实体、目录、结构、主题图),通过置信度加权倒数排名融合,继而进行蒙特卡洛证据采样和多轮 ReAct 自适应推理,适用于复杂查询场景下的最大召回(10–30 秒)。 -自 v0.0.6post1 起,Sirchmunk 已发布为 OpenClaw 技能 — 任何兼容 OpenClaw 的 Agent 均可通过自然语言调用其搜索能力。自 v0.0.6 起,在 FAST/DEEP 搜索模式之上还增加了**多轮对话**与上下文管理、**文档摘要**与**跨语言检索**。 +自 v0.2.0 起,Sirchmunk 已发布为 OpenClaw 技能 — 任何兼容 OpenClaw 的 Agent 均可通过自然语言调用其搜索能力。在 FAST/DEEP 搜索模式之上还增加了**多轮对话**与上下文管理、**文档摘要**与**跨语言检索**。 --- @@ -155,7 +157,20 @@ Sirchmunk 采用"事后索引"策略。它不在你提问之前建索引,而 | **即时索引** | 根据实际使用模式动态构建数据地图。 | Python 原生 | | **知识复用** | 对相似的后续查询命中缓存,从暴力搜索进化为高速检索。 | DuckDB SQL | -通过这一机制,Sirchmunk 从"暴力猎手"有机地进化为"精明的图书管理员",无需传统预索引数据库的维护开销。 +通过这一机制,Sirchmunk 从“暴力猎手”有机地进化为“精明的图书管理员”,无需传统预索引数据库的维护开销。 + +--- + +### 3.5 实验验证 + +LENS 框架的有效性已通过严格的受控评估验证: + +| 设置 | LENS(Sirchmunk DEEP) | ReAct 基线 | +| --- | --- | --- | +| **500 题受控评估** | 62.4% EM,84.8% 证据召回 | 65.2% EM,50.4% 证据召回 | +| **150 题 fullwiki**(原始 Wikipedia,零索引) | 43.3% EM,84.0% 证据召回 | 42.7% EM,70.7% 证据召回 | + +这些结果揭示了一个关键洞察:LENS 以微小的准确率差距换取了显著更好的证据接地能力。在 fullwiki 场景中 — 不执行任何索引或预处理 — LENS 在实现可比的 EM 的同时,提供了多 13 个百分点以上的证据召回率,验证了上下文搜索范式的源保真度主张。 --- @@ -229,7 +244,8 @@ ICS 范式以预处理时间换取查询时的智能,这使得性能瓶颈发 1. Lewis, P., Perez, E., Piktus, A., 等. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.* NeurIPS 2020. [arXiv:2005.11401](https://arxiv.org/abs/2005.11401) 2. Guo, Z., Qian, C., 等. (2024). *LightRAG: Simple and Fast Retrieval-Augmented Generation.* [arXiv:2410.05779](https://arxiv.org/abs/2410.05779) | [GitHub](https://github.com/HKUDS/LightRAG) 3. VectifyAI. (2025). *PageIndex: Extracting and Understanding Financial Reports with LLM.* [GitHub](https://github.com/VectifyAI/PageIndex) -4. ModelScope. (2025). *Sirchmunk:一个无嵌入的、智能体驱动的原始数据搜索引擎。* [GitHub](https://github.com/modelscope/sirchmunk) -5. Yao, S., Zhao, J., Yu, D., 等. (2023). *ReAct: Synergizing Reasoning and Acting in Language Models.* ICLR 2023. [arXiv:2210.03629](https://arxiv.org/abs/2210.03629) -6. Anthropic. (2024). *模型上下文协议(MCP)规范。* [官方文档](https://modelcontextprotocol.io) -7. Kaddour, J., Harris, J., Mozes, M., 等. (2023). *Challenges and Applications of Large Language Models.* [arXiv:2307.10169](https://arxiv.org/abs/2307.10169) +4. ModelScope. (2026). *Sirchmunk:一个无嵌入的、智能体驱动的原始数据搜索引擎。* [GitHub](https://github.com/modelscope/sirchmunk) +5. Wang, X., et al. (2026). *LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents.* [arXiv:2608.16185](https://arxiv.org/abs/2608.16185) +6. Yao, S., Zhao, J., Yu, D., 等. (2023). *ReAct: Synergizing Reasoning and Acting in Language Models.* ICLR 2023. [arXiv:2210.03629](https://arxiv.org/abs/2210.03629) +7. Anthropic. (2024). *模型上下文协议(MCP)规范。* [官方文档](https://modelcontextprotocol.io) +8. Kaddour, J., Harris, J., Mozes, M., 等. (2023). *Challenges and Applications of Large Language Models.* [arXiv:2307.10169](https://arxiv.org/abs/2307.10169) diff --git a/content/blog/technical-deep-dive/Sirchmunk_Architecture.png b/content/blog/technical-deep-dive/Sirchmunk_Architecture.png index ac86d1e..7346622 100644 Binary files a/content/blog/technical-deep-dive/Sirchmunk_Architecture.png and b/content/blog/technical-deep-dive/Sirchmunk_Architecture.png differ diff --git a/content/blog/technical-deep-dive/Sirchmunk_LENS_Framework.png b/content/blog/technical-deep-dive/Sirchmunk_LENS_Framework.png new file mode 100644 index 0000000..7346622 Binary files /dev/null and b/content/blog/technical-deep-dive/Sirchmunk_LENS_Framework.png differ diff --git a/content/blog/technical-deep-dive/Sirchmunk_MonteCarloSamplingAlgo.png b/content/blog/technical-deep-dive/Sirchmunk_MonteCarloSamplingAlgo.png index aca393f..7346622 100644 Binary files a/content/blog/technical-deep-dive/Sirchmunk_MonteCarloSamplingAlgo.png and b/content/blog/technical-deep-dive/Sirchmunk_MonteCarloSamplingAlgo.png differ diff --git a/content/blog/technical-deep-dive/featured.png b/content/blog/technical-deep-dive/featured.png index ac86d1e..7346622 100644 Binary files a/content/blog/technical-deep-dive/featured.png and b/content/blog/technical-deep-dive/featured.png differ diff --git a/content/blog/technical-deep-dive/index.md b/content/blog/technical-deep-dive/index.md index f0c9a48..3a55f15 100644 --- a/content/blog/technical-deep-dive/index.md +++ b/content/blog/technical-deep-dive/index.md @@ -61,11 +61,25 @@ Sirchmunk's design philosophy rests on three pillars: ![Sirchmunk Architecture](Sirchmunk_Architecture.png "Fig. 1 — Sirchmunk high-level architecture. The system is organized into cleanly separated layers: an Integration layer for external surfaces, an Orchestration layer for the search pipeline, an Intelligence layer for evidence extraction and knowledge synthesis, and a Storage layer for persistence.") +## 3.5 LENS Framework + +![LENS Framework](Sirchmunk_LENS_Framework.png "Fig. 1b — LENS reframes in-context search as budgeted evidence exploration over a latent evidence space induced by dynamic raw documents.") + +The theoretical foundations of Sirchmunk's retrieval engine are formalized in the LENS framework — **Latent Evidence Navigation and Search** — published as ["LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents"](https://arxiv.org/abs/2608.16185) (arXiv:2608.16185). + +LENS formulates in-context search as **Budgeted Evidence Localization**: given a query and a raw-document corpus, the system maintains a query-conditioned belief over a latent evidence space and iteratively refines it through: + +- **Proposal policies**: Candidate evidence regions are nominated by multiple retrieval strategies (lexical, entity, directory, structural, topic-graph). +- **LLM relevance oracle**: Each candidate is evaluated by the LLM, updating the belief state. +- **Budget-aware stopping**: The system halts when the evidence is deemed sufficient or the token budget is exhausted. + +Key experimental results: 62.4% Exact Match with 84.8% evidence recall on a 500-question controlled evaluation; 43.3% EM on a raw Wikipedia dump (fullwiki) with zero indexing — demonstrating competitive quality under strict source-fidelity constraints. + ## 4. Core Search Pipeline At the heart of Sirchmunk is a multi-phase search pipeline. The core design principle is **maximum parallelism within each phase** combined with **strict phase dependencies** between them. This achieves both speed and correctness: independent tasks race concurrently, while each phase builds on the converged output of the previous one. -Sirchmunk supports two search modes: **FAST mode** (default) uses a greedy strategy with 2-level keyword cascade, context-window sampling, and early stopping — completing retrieval in 2–5 seconds with only 2 LLM calls (~10x faster than DEEP mode). **DEEP mode** activates the full pipeline described below, including Monte Carlo evidence sampling and multi-round ReAct refinement, for maximum recall on complex queries (10–30 seconds). +Sirchmunk supports two search modes: **FAST mode** uses a greedy strategy with 2-level keyword cascade, context-window sampling, and early stopping — completing retrieval in 2–5 seconds with only 2 LLM calls (~10x faster than DEEP mode). **DEEP mode** (default since v0.1.0) activates the full pipeline described below, including multi-path retrieval (lexical, entity, directory, structural, topic-graph) fused via confidence-weighted Reciprocal Rank Fusion (RRF), Monte Carlo evidence sampling, and multi-round ReAct refinement, for maximum recall on complex queries (10–30 seconds). A soft route-collapse mechanism dynamically disables low-yield paths when a single path produces a high-confidence match. ```text Query @@ -263,6 +277,8 @@ Clusters do not exist in isolation. Two types of semantic edges weave them into - **Weak semantic edges**: Lightweight, weighted connections that indicate topical relatedness. - **Rich cognitive edges**: Typed relationships (Pathway, Barrier, Analogy, Shortcut, Resolution) that model how insights relate to each other at a deeper level — enabling future capabilities like cognitive graph navigation and chain-of-thought retrieval. +The Web UI now includes an **interactive Knowledge Graph visualization** powered by Cytoscape.js. Users can explore the cluster topology, filter by lifecycle states, and trace semantic relationships — making the system's evolving intelligence tangible and inspectable. + > **Design insight:** By modeling knowledge with lifecycle states and abstraction levels, Sirchmunk treats its knowledge base as a living organism rather than a dead archive. Knowledge can be born, grow, and eventually retire — mirroring how human expertise evolves. ## 9. Storage & Persistence Philosophy @@ -283,7 +299,7 @@ Each knowledge cluster's embedding vector (384 dimensions) is stored alongside t ## 10. Integration Layer: MCP, OpenClaw, API & Beyond -Sirchmunk is designed to be **consumed, not deployed**. Rather than requiring users to build applications around it, it exposes its intelligence through multiple standard interfaces — meeting users where they already work. **As of v0.0.6post1**, Sirchmunk is also published as an **OpenClaw skill**, so any OpenClaw-compatible agent can invoke its search workflow via natural language alongside MCP-native and HTTP clients. +Sirchmunk is designed to be **consumed, not deployed**. Rather than requiring users to build applications around it, it exposes its intelligence through multiple standard interfaces — meeting users where they already work. **As of v0.2.0**, Sirchmunk is also published as an **OpenClaw skill**, so any OpenClaw-compatible agent can invoke its search workflow via natural language alongside MCP-native and HTTP clients. ### Model Context Protocol (MCP) @@ -334,7 +350,7 @@ Sirchmunk represents a paradigm shift in how we think about retrieval-augmented --- -*This technical report was generated by analyzing the Sirchmunk source code (v0.0.6post1).* +*This technical report was generated by analyzing the Sirchmunk source code (v0.2.0).* *[github.com/modelscope/sirchmunk](https://github.com/modelscope/sirchmunk) · [ModelScope](https://github.com/modelscope)* *Sirchmunk: Raw data to self-evolving intelligence, real-time.* diff --git a/content/blog/technical-deep-dive/index.zh.md b/content/blog/technical-deep-dive/index.zh.md index 1c2a30d..f867e49 100644 --- a/content/blog/technical-deep-dive/index.zh.md +++ b/content/blog/technical-deep-dive/index.zh.md @@ -61,11 +61,25 @@ Sirchmunk 的设计哲学建立在三个支柱之上: ![Sirchmunk 架构](Sirchmunk_Architecture.png "图 1 — Sirchmunk 高层架构。系统采用清晰分离的层次结构:用于外部接口的集成层、用于搜索管线的编排层、用于证据提取和知识合成的智能层,以及用于持久化的存储层。") +## 3.5 LENS 框架 + +![LENS 框架](Sirchmunk_LENS_Framework.png "图 1b — LENS 将 in-context search 表述为动态原始文档诱导的隐式证据空间上的预算约束证据探索。") + +Sirchmunk 检索引擎的理论基础在 LENS 框架中得以形式化 — **Latent Evidence Navigation and Search** — 论文发表为 ["LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents"](https://arxiv.org/abs/2608.16185)(arXiv:2608.16185)。 + +LENS 将 in-context search 表述为**预算约束证据定位**:给定查询和原始文档语料,系统在隐式证据空间上维护查询条件信念,并通过以下机制迭代精炼: + +- **提议策略**:多种检索策略(词法、实体、目录、结构、主题图)提名候选证据区域。 +- **LLM 相关性预言机**:LLM 评估每个候选,更新信念状态。 +- **预算感知停止**:当证据被判定为充分或 token 预算耗尽时,系统停止。 + +关键实验结果:500 题受控评估中 62.4% 精确匹配,84.8% 证据召回率;原始 Wikipedia 转储(fullwiki)零索引场景下 43.3% EM — 在严格的源保真度约束下展现出竞争力的质量。 + ## 4. 核心搜索管线 Sirchmunk 的核心是一个多阶段搜索管线。核心设计原则是**每个阶段内最大并行化**与阶段之间**严格的依赖关系**。这同时实现了速度和正确性:独立任务并发竞争,而每个阶段都建立在前一阶段汇聚的输出之上。 -Sirchmunk 支持两种搜索模式:**FAST 模式**(默认)采用贪心策略,结合两级关键词级联、上下文窗口采样和 early stopping,仅需 2 次 LLM 调用即可在 2–5 秒内完成检索(速度约为 DEEP 模式的 10 倍)。**DEEP 模式** 则激活下文所述的完整管线,包括蒙特卡洛证据采样和多轮 ReAct 自适应推理,适用于复杂查询场景下的最大召回(10–30 秒)。 +Sirchmunk 支持两种搜索模式:**FAST 模式** 采用贪心策略,结合两级关键词级联、上下文窗口采样和 early stopping,仅需 2 次 LLM 调用即可在 2–5 秒内完成检索(速度约为 DEEP 模式的 10 倍)。**DEEP 模式**(自 v0.1.0 起为默认模式)则激活下文所述的完整管线,包括多路检索(词法、实体、目录、结构、主题图)通过置信度加权倒数排名融合(RRF)融合、蒙特卡洛证据采样和多轮 ReAct 自适应推理,适用于复杂查询场景下的最大召回(10–30 秒)。soft 路由收缩机制在单条路径产生高置信度匹配时动态禁用低产出路径。 ```text Query @@ -263,6 +277,8 @@ Sirchmunk 中的知识不是静态索引 — 它是一个随每次搜索不断 - **弱语义边**:轻量级的加权连接,指示主题相关性。 - **丰富认知边**:类型化关系(路径、障碍、类比、捷径、解决方案),在更深层次建模洞察之间的关系 — 为未来的认知图谱导航和思维链检索等能力赋能。 +Web UI 现已包含基于 Cytoscape.js 的**交互式知识图谱可视化**。用户可以探索聚类拓扑、按生命周期状态过滤、追踪语义关系 — 使系统进化中的智能变得可见和可检查。 + > **设计洞察:** 通过使用生命周期状态和抽象级别建模知识,Sirchmunk 将其知识库视为一个活的有机体而非死的档案。知识可以诞生、成长并最终退休 — 映射人类专业知识的演进方式。 ## 9. 存储与持久化哲学 @@ -283,7 +299,7 @@ Sirchmunk 的存储设计遵循我们称之为**"默认快速,重要时持久" ## 10. 集成层:MCP、OpenClaw、API 及更多 -Sirchmunk 的设计理念是**被使用,而非被部署**。它不要求用户围绕它构建应用,而是通过多个标准接口暴露其智能 — 在用户已经工作的地方与他们会面。**自 v0.0.6post1 起**,Sirchmunk 亦以 **OpenClaw 技能** 形式发布,任何兼容 OpenClaw 的 Agent 均可通过自然语言调用其搜索流程,与 MCP 原生及 HTTP 客户端并存。 +Sirchmunk 的设计理念是**被使用,而非被部署**。它不要求用户围绕它构建应用,而是通过多个标准接口暴露其智能 — 在用户已经工作的地方与他们会面。**自 v0.2.0 起**,Sirchmunk 亦以 **OpenClaw 技能** 形式发布,任何兼容 OpenClaw 的 Agent 均可通过自然语言调用其搜索流程,与 MCP 原生及 HTTP 客户端并存。 ### 模型上下文协议 (MCP) @@ -334,7 +350,7 @@ Sirchmunk 代表了我们思考检索增强生成方式的范式转变。通过 --- -*本技术报告通过分析 Sirchmunk 源代码(v0.0.6post1)生成。* +*本技术报告通过分析 Sirchmunk 源代码(v0.2.0)生成。* *[github.com/modelscope/sirchmunk](https://github.com/modelscope/sirchmunk) · [ModelScope](https://github.com/modelscope)* *Sirchmunk:从原始数据到自进化智能,实时。* diff --git a/content/blog/v0.2.0/Sirchmunk_Architecture.png b/content/blog/v0.2.0/Sirchmunk_Architecture.png new file mode 100644 index 0000000..7346622 Binary files /dev/null and b/content/blog/v0.2.0/Sirchmunk_Architecture.png differ diff --git a/content/blog/v0.2.0/featured.png b/content/blog/v0.2.0/featured.png new file mode 100644 index 0000000..7346622 Binary files /dev/null and b/content/blog/v0.2.0/featured.png differ diff --git a/content/blog/v0.2.0/index.md b/content/blog/v0.2.0/index.md new file mode 100644 index 0000000..6f626b4 --- /dev/null +++ b/content/blog/v0.2.0/index.md @@ -0,0 +1,96 @@ +--- +title: "Sirchmunk v0.2.0: LENS Paper, Multi-Path DEEP Retrieval & Large Corpus Robustness" +summary: "Sirchmunk v0.2.0 is a major milestone — the LENS research paper is published on arXiv, the retrieval engine gains multi-path DEEP fusion and bounded retrieval cost invariants, and the system adopts a generalization-first design philosophy." +date: 2026-09-20 +authors: + - admin +tags: + - Release + - LENS + - DEEP + - Retrieval +image: + caption: 'LENS Framework' +--- + +Sirchmunk v0.2.0 marks a turning point for the project. The core algorithm behind Sirchmunk's retrieval engine has been formalized in a research paper and published on arXiv, the DEEP search pipeline has been rebuilt around multi-path fusion, and the system now enforces strict retrieval cost invariants that keep performance predictable on corpora of any size. + + + +## LENS: The Research Paper + +The theoretical foundations of Sirchmunk's in-context search have been formalized in **"LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents"** ([arXiv:2608.16185](https://arxiv.org/abs/2608.16185)). + +LENS reframes retrieval as *Budgeted Evidence Localization*: given a query and a raw-document corpus, the system maintains a query-conditioned belief over a latent evidence space and iteratively refines it through proposal policies and an LLM relevance oracle — all under an explicit token budget. + +Key results from the controlled evaluation: + +- **500-question evaluation**: 62.4% Exact Match with 84.8% evidence recall (vs. ReAct baseline at 65.2% EM but only 50.4% evidence recall). +- **150-question fullwiki subset** (raw Wikipedia dump, zero indexing): LENS achieves 43.3% EM vs. ReAct's 42.7% EM, with substantially stronger evidence grounding (84.0% vs. 70.7%). + +These numbers demonstrate that LENS trades a modest amount of accuracy for dramatically better evidence traceability — the answer is not only correct, but *provably grounded* in source material. + +## Multi-Path DEEP Retrieval + +DEEP mode has been fundamentally restructured. Instead of a single retrieval strategy, it now runs **five complementary retrieval paths** in parallel: + +1. **Lexical** — keyword-driven content matching with IDF-weighted scoring. +2. **Entity** — exact-match probes for named entities, identifiers, and structured values. +3. **Directory** — file-system structure analysis using naming conventions, path hierarchy, and modification timestamps. +4. **Structural** — heuristic document tree navigation (v2) with structure anchors for DOCX, RST, and other formatted sources. +5. **Topic-Graph** — cross-document topic-map routing that leverages the self-evolving knowledge graph to identify relevant document clusters. + +Results from all five paths are fused via **confidence-weighted Reciprocal Rank Fusion (RRF)**. A **soft route-collapse** mechanism dynamically disables low-yield paths when a single path produces a high-confidence match, cutting latency and token consumption without sacrificing answer quality. + +![Sirchmunk Architecture](Sirchmunk_Architecture.png "Sirchmunk v0.2.0 system architecture with multi-path DEEP retrieval and LENS framework integration.") + +## Large Corpus Robustness + +Previous versions could stall or time out on very large, archive-heavy corpora. v0.2.0 introduces **retrieval cost invariants** — hard bounds that ensure per-file and per-query cost never grows unbounded: + +- **`GREP_RGA_ADAPTERS`**: Capability-based adapter whitelist. Only bounded document extractors (poppler, pandoc) are enabled on the query hot path; unbounded recursive adapters (decompress, zip, tar, sqlite, ffmpeg) are restricted to offline extraction. +- **`GREP_MAX_FILESIZE_MB`**: Per-file size cap. Files exceeding the threshold are skipped during live queries. +- **`GREP_TIERED_SCAN`**: A fast native-`rg` pass over all files is unioned with an `rga` pass restricted to rich-format extensions, so adapter dispatch never walks the entire tree. +- **Fail-fast timeout budgets**: `GREP_TEXT_TIMEOUT` for the text pass and `GREP_TIMEOUT` for the rich pass; on timeout, search degrades gracefully to native `rg` rather than hanging. + +Directory scanning is now enabled by default for stronger filename routing, and a hard token budget keeps each query within a configurable limit. + +## Generalization-First Design + +v0.2.0 adopts a principled stance against benchmark-specific hard rules. Benchmark-tied logic such as hardcoded entity patterns, fixed column positions, and English-only stop words has been replaced with generalizable alternatives: + +- **Grounded numeric verification**: Computation answers are re-checked deterministically from model-disclosed, evidence-grounded operands — entirely corpus-agnostic. +- **Injectable tokenizers**: Tokenizers and lexical policies are dependency-injected with a general default implementation, rather than embedded inside modules. +- **Corpus-adaptive statistics**: Document-frequency-based adaptive stop-word pruning replaces fixed word lists, working natively across languages and domains. + +Every replacement behavior is gated behind an environment switch, defaults to the new implementation, and allows single-item rollback on failure. + +## Knowledge Graph Visualization + +The Web UI now includes an **interactive knowledge graph** powered by Cytoscape.js. The graph visualizes self-evolving knowledge clusters and their semantic relationships, including lifecycle states (Emerging, Stable, Meta) and edge weights. Users can explore, filter, and drill down into the cluster topology to understand how the system's knowledge evolves with use. + +## Other Improvements + +- **Broader format coverage**: Native exact-match fallback for LOG, PPTX, and XLSX files, plus heuristic document tree v2 (including DOCX/RST) with structure anchors guiding evidence extraction. +- **DeepSeek V4 compatibility**: Full support for DeepSeek V4's thinking mode (`thinking_content`) in the OpenAI-compatible client. +- **Stabilized large-corpus retrieval**: Combined improvements in tiered scanning, per-file match caps, and adapter whitelisting eliminate the timeout and stall issues observed on large corpora in earlier versions. + +## Get Started + +```bash +pip install --upgrade sirchmunk +``` + +Or install with all extras: + +```bash +pip install "sirchmunk[all]" +``` + +- **Documentation**: [modelscope.github.io/sirchmunk-web](https://modelscope.github.io/sirchmunk-web/) +- **GitHub**: [github.com/modelscope/sirchmunk](https://github.com/modelscope/sirchmunk) +- **Paper**: [arXiv:2608.16185](https://arxiv.org/abs/2608.16185) + +--- + +*[GitHub Repository](https://github.com/modelscope/sirchmunk) · [ModelScope](https://github.com/modelscope)* diff --git a/content/blog/v0.2.0/index.zh.md b/content/blog/v0.2.0/index.zh.md new file mode 100644 index 0000000..ed78237 --- /dev/null +++ b/content/blog/v0.2.0/index.zh.md @@ -0,0 +1,96 @@ +--- +title: "Sirchmunk v0.2.0:LENS 论文发布、多路 DEEP 检索与大语料鲁棒性" +summary: "Sirchmunk v0.2.0 是一个重要里程碑 — LENS 研究论文已发布至 arXiv,检索引擎引入多路 DEEP 融合与有界检索成本不变量,系统全面采用泛化优先的设计理念。" +date: 2026-09-20 +authors: + - admin +tags: + - Release + - LENS + - DEEP + - Retrieval +image: + caption: 'LENS 框架' +--- + +Sirchmunk v0.2.0 是项目的一个转折点。Sirchmunk 检索引擎背后的核心算法已在研究论文中形式化并发布至 arXiv,DEEP 搜索管线围绕多路融合进行了重构,系统现在强制执行严格的检索成本不变量,确保在任意规模的语料上都能保持可预测的性能。 + + + +## LENS:研究论文 + +Sirchmunk 上下文搜索的理论基础已在 **"LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents"**([arXiv:2608.16185](https://arxiv.org/abs/2608.16185))中形式化。 + +LENS 将检索重新表述为*预算约束证据定位*:给定查询和原始文档语料,系统在隐式证据空间上维护一个查询条件信念,通过提议策略和 LLM 相关性预言机迭代精炼 — 全程受显式 token 预算约束。 + +受控评估的关键结果: + +- **500 题评估**:62.4% 精确匹配(EM),84.8% 证据召回率(对比 ReAct 基线 65.2% EM,但证据召回仅 50.4%)。 +- **150 题 fullwiki 子集**(原始 Wikipedia 转储,零索引):LENS 达到 43.3% EM,对比 ReAct 的 42.7% EM,且证据接地能力显著更强(84.0% vs. 70.7%)。 + +这些数据表明 LENS 以微小的准确率代价换取了显著更好的证据可追溯性 — 答案不仅正确,而且*可证地接地*于源材料。 + +## 多路 DEEP 检索 + +DEEP 模式经过了根本性重构。取代单一检索策略,现在并行运行 **五条互补的检索路径**: + +1. **词法路径** — 基于关键词的内容匹配,采用 IDF 加权评分。 +2. **实体路径** — 对命名实体、标识符和结构化值的精确匹配探测。 +3. **目录路径** — 利用命名约定、路径层次和修改时间戳进行文件系统结构分析。 +4. **结构路径** — 启发式文档树导航(v2),带有针对 DOCX、RST 等格式源的结构锚点。 +5. **主题图路径** — 跨文档主题图路由,利用自进化知识图谱识别相关文档集群。 + +五条路径的结果通过**置信度加权倒数排名融合(RRF)**进行融合。**soft 路由收缩**机制在单条路径产生高置信度匹配时动态禁用低产出路径,在不损失答案质量的前提下降低延迟和 token 消耗。 + +![Sirchmunk 架构](Sirchmunk_Architecture.png "Sirchmunk v0.2.0 系统架构:多路 DEEP 检索与 LENS 框架集成。") + +## 大语料鲁棒性 + +早期版本在超大、含归档文件的语料上可能卡顿或超时。v0.2.0 引入了**检索成本不变量** — 确保单文件和单次查询成本永远不会无界增长的硬约束: + +- **`GREP_RGA_ADAPTERS`**:基于能力的适配器白名单。查询热路径上仅启用有界文档提取器(poppler、pandoc);无界递归适配器(decompress、zip、tar、sqlite、ffmpeg)限制为离线提取。 +- **`GREP_MAX_FILESIZE_MB`**:单文件大小上限。超过阈值的文件在实时查询中被跳过。 +- **`GREP_TIERED_SCAN`**:对所有文件执行快速原生 `rg` 扫描,与限定富格式扩展名的 `rga` 扫描取并集,避免适配器调度遍历整棵文件树。 +- **快速失败超时预算**:`GREP_TEXT_TIMEOUT` 用于文本扫描,`GREP_TIMEOUT` 用于富文档扫描;超时后搜索优雅降级为原生 `rg`,而非挂起。 + +目录扫描现已默认开启以增强文件名路由,硬 token 预算将每次查询限制在可配置额度内。 + +## 泛化优先设计 + +v0.2.0 对基准特定硬规则采取了原则性立场。硬编码的实体模式、固定列位置、仅限英语的停用词等基准绑定逻辑已被可泛化的替代方案取代: + +- **接地数值校验**:计算类答案基于模型披露且证据接地的操作数进行确定性复算 — 完全语料无关。 +- **可注入分词器**:分词器和词法策略通过依赖注入提供通用默认实现,而非嵌入模块内部。 +- **语料自适应统计**:基于文档频率的自适应停用词剪枝替代固定词表,原生适配所有语言和领域。 + +每个替换行为都通过环境开关控制,默认启用新实现,并允许单项故障回滚。 + +## 知识图谱可视化 + +Web UI 新增基于 Cytoscape.js 的**交互式知识图谱**。图谱可视化自进化知识聚类及其语义关系,包括生命周期状态(Emerging、Stable、Meta)和边权重。用户可以探索、过滤和深入聚类拓扑,直观了解系统知识随使用的演化过程。 + +## 其他改进 + +- **更广格式覆盖**:LOG、PPTX、XLSX 文件原生精确匹配回退,启发式文档树 v2(含 DOCX/RST)并以结构锚点引导证据抽取。 +- **DeepSeek V4 兼容**:OpenAI 兼容客户端完整支持 DeepSeek V4 思考模式(`thinking_content`)。 +- **大语料检索稳定性**:分层扫描、单文件匹配上限和适配器白名单的综合改进,消除了早期版本在大语料上的超时和卡顿问题。 + +## 开始使用 + +```bash +pip install --upgrade sirchmunk +``` + +或安装全部附加组件: + +```bash +pip install "sirchmunk[all]" +``` + +- **文档**:[modelscope.github.io/sirchmunk-web/zh](https://modelscope.github.io/sirchmunk-web/zh/) +- **GitHub**:[github.com/modelscope/sirchmunk](https://github.com/modelscope/sirchmunk) +- **论文**:[arXiv:2608.16185](https://arxiv.org/abs/2608.16185) + +--- + +*[GitHub 仓库](https://github.com/modelscope/sirchmunk) · [ModelScope](https://github.com/modelscope)* diff --git a/content/community/index.md b/content/community/index.md index 0faf7b4..f774a87 100644 --- a/content/community/index.md +++ b/content/community/index.md @@ -14,6 +14,10 @@ Get help and connect with the Sirchmunk community. We welcome contributions, que - View the [Sirchmunk Documentation](/docs/) - Read the [Technical Deep Dive](/blog/technical-deep-dive/) +## Research Paper {#paper} + +- [LENS Paper (arXiv:2608.16185)](https://arxiv.org/abs/2608.16185) — *LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents* — the research paper behind Sirchmunk's core algorithm. + ## Source Code {#source} - [Sirchmunk on GitHub](https://github.com/modelscope/sirchmunk) — Star the repo and follow development diff --git a/content/community/index.zh.md b/content/community/index.zh.md index 2f91e0b..f5e0318 100644 --- a/content/community/index.zh.md +++ b/content/community/index.zh.md @@ -14,6 +14,10 @@ pager: false - 查看 [Sirchmunk 文档](/zh/docs/) - 阅读 [技术深度报告](/zh/blog/technical-deep-dive/) +## 研究论文 {#paper} + +- [LENS 论文 (arXiv:2608.16185)](https://arxiv.org/abs/2608.16185) — *LENS:基于动态原始文档的隐式证据探索上下文搜索* — Sirchmunk 核心算法的研究论文。 + ## 源代码 {#source} - [Sirchmunk GitHub](https://github.com/modelscope/sirchmunk) — 给仓库点个 Star,关注开发进展 diff --git a/content/docs/guide/Sirchmunk_Architecture.png b/content/docs/guide/Sirchmunk_Architecture.png index ac86d1e..7346622 100644 Binary files a/content/docs/guide/Sirchmunk_Architecture.png and b/content/docs/guide/Sirchmunk_Architecture.png differ diff --git a/content/docs/guide/Sirchmunk_Knowledge_Graph.png b/content/docs/guide/Sirchmunk_Knowledge_Graph.png new file mode 100644 index 0000000..f53916d Binary files /dev/null and b/content/docs/guide/Sirchmunk_Knowledge_Graph.png differ diff --git a/content/docs/guide/Sirchmunk_LENS_Framework.png b/content/docs/guide/Sirchmunk_LENS_Framework.png new file mode 100644 index 0000000..7346622 Binary files /dev/null and b/content/docs/guide/Sirchmunk_LENS_Framework.png differ diff --git a/content/docs/guide/Sirchmunk_MonteCarloSamplingAlgo.png b/content/docs/guide/Sirchmunk_MonteCarloSamplingAlgo.png index aca393f..7346622 100644 Binary files a/content/docs/guide/Sirchmunk_MonteCarloSamplingAlgo.png and b/content/docs/guide/Sirchmunk_MonteCarloSamplingAlgo.png differ diff --git a/content/docs/guide/architecture.md b/content/docs/guide/architecture.md index aa912f1..7e69435 100644 --- a/content/docs/guide/architecture.md +++ b/content/docs/guide/architecture.md @@ -9,13 +9,39 @@ Sirchmunk's architecture is organized into cleanly separated layers, following t ![Sirchmunk Architecture](Sirchmunk_Architecture.png "Sirchmunk high-level architecture diagram") +## LENS Framework + +![LENS Framework](Sirchmunk_LENS_Framework.png "LENS: budgeted evidence exploration over latent evidence space") + +**LENS** (Latent Evidence Exploration and Search) is an index-free retrieval framework that formulates in-context search as **Budgeted Evidence Localization** over a latent evidence space induced by dynamic raw documents. Instead of pre-materializing evidence via embedding indexes or chunk stores, LENS maintains a query-conditioned belief over candidate evidence units and iteratively: + +1. **Proposes** candidates via complementary lexical, local, and exploratory proposal policies +2. **Updates** the belief via an LLM relevance oracle +3. **Narrows** toward high-posterior regions under a controllable token budget + +This formulation makes the search process adaptive, budget-aware, and fully grounded in source documents — without any pre-built index infrastructure. + +> **Paper:** [LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents](https://arxiv.org/abs/2608.16185) (arXiv 2026) + +## Knowledge Graph + +![Knowledge Graph](Sirchmunk_Knowledge_Graph.png "Interactive knowledge cluster visualization") + +The **Knowledge Graph** provides an interactive visualization of self-evolving knowledge clusters built incrementally from search interactions. Powered by Cytoscape.js, it displays: + +- **Cluster relationships** — semantic edges linking related knowledge units +- **Lifecycle states** — visual encoding of Emerging → Stable → Deprecated transitions +- **Leiden meta-clustering** — higher-level community structure discovered via the Leiden algorithm + +Users can explore, filter, and drill into clusters directly from the Web UI. + ## Core Components | Component | Description | |:----------------------|:-------------------------------------------------------------------------| -| **AgenticSearch** | Search orchestrator with LLM-enhanced retrieval capabilities | -| **KnowledgeBase** | Transforms raw results into structured knowledge clusters with evidences | -| **EvidenceProcessor** | Evidence processing based on the Monte Carlo Importance Sampling | +| **AgenticSearch** | Search orchestrator with FAST / DEEP / FILENAME_ONLY modes and budget-aware evidence localization | +| **KnowledgeBase** | Persists source-grounded evidence clusters as reusable warm priors for later queries | +| **EvidenceProcessor** | Consolidates candidate regions into compact, traceable evidence units | | **GrepRetriever** | High-performance _indexless_ file search with parallel processing | | **OpenAIChat** | Unified LLM interface supporting streaming and usage tracking | | **KnowledgeCompiler** | Offline document compilation into tree indices and knowledge clusters (Beta) | @@ -34,28 +60,35 @@ This is not merely a cache — it is the beginning of knowledge *compounding*. E ### Phase 1 — Parallel Probing -Four independent probes launch concurrently to gather diverse signals: +Multiple independent probes launch concurrently to gather diverse signals: -1. **LLM Keyword Extraction** — The LLM decomposes the query into multi-level keywords, from coarse (high recall) to fine (high precision), each annotated with an estimated rarity score. This multi-granularity approach ensures that both broad topics and specific terms are captured. -2. **Directory Structure Scan** — The file system is traversed to collect path metadata: file names, sizes, modification times, and content previews. This is the foundation for *intelligent inference* — using structural cues (naming conventions, directory hierarchy, file types) to narrow down the most promising candidates before ever reading their content. +1. **LLM Keyword Extraction** — The LLM decomposes the query into multi-level keywords, from coarse (high recall) to fine (high precision), each annotated with an estimated rarity score. +2. **Directory Structure Scan** — The file system is traversed to collect path metadata: file names, sizes, modification times, and content previews — the foundation for *intelligent inference* using structural cues. 3. **Knowledge Cache Lookup** — Partial match search across existing clusters for potential reuse of previously acquired knowledge. -4. **Spec-Path Context Load** — Previously computed context for known paths is loaded from cache. +4. **Compile Artifact Load** — Previously computed tree indices, topic maps, and summary indexes are loaded when available. + +### Phase 2 — Multi-Path DEEP Retrieval -### Phase 2 — Retrieval & Ranking +In v0.2.0, DEEP mode runs **5 complementary retrieval paths** in parallel: -Two complementary strategies run in parallel: +| Path | Signal | +|:-----|:-------| +| **Lexical** | IDF-weighted keyword search through raw file contents | +| **Entity** | Exact-match entity lookup for named entities and identifiers | +| **Directory** | Structure-based ranking via LLM-guided metadata evaluation | +| **Structural** | Heuristic document tree navigation using compile artifacts | +| **Topic-Graph** | Cross-document topic-map traversal for multi-hop discovery | -- **Content-based retrieval** — IDF-weighted keyword search through raw file contents -- **Structure-based ranking** — LLM-guided evaluation of candidate files by metadata +Results from all paths are fused via **confidence-weighted Reciprocal Rank Fusion (RRF)**. A **soft route collapse** mechanism dynamically disables low-yield paths — when a single path achieves high confidence with strong margin, the remaining paths are collapsed to cut latency and tokens while preserving answer quality. -### Phase 3 — Knowledge Cluster Construction +### Phase 3 — Evidence Localization & Cluster Construction -Results are merged, deduplicated, and processed through Monte Carlo evidence sampling. The LLM synthesizes evidence fragments into structured Knowledge Clusters. +Results are merged, deduplicated, and processed through budgeted evidence localization. The LLM synthesizes evidence fragments into structured Knowledge Clusters. ### Phase 4 — Summarization or ReAct Refinement -- **Evidence found** → LLM generates a structured briefing -- **No evidence** → ReAct agent activates for iterative exploration +- **Evidence found** → LLM generates a structured briefing with source-linked evidence +- **No evidence** → ReAct agent activates for iterative exploration with hop-aware strategies ### Phase 5 — Persist @@ -63,25 +96,25 @@ Valuable clusters are saved with their embeddings for future reuse. ## Key Algorithms -### Monte Carlo Evidence Sampling +### Budgeted Evidence Exploration -Traditional retrieval systems read entire documents or rely on fixed-size chunks, leading to either wasted tokens or lost context. Sirchmunk takes a fundamentally different approach inspired by **Monte Carlo methods** — treating evidence extraction as a **sampling problem** rather than a parsing problem. +Traditional retrieval systems read entire documents or rely on fixed-size chunks, leading to either wasted tokens or lost context. LENS instead treats the relevant evidence as **latent and query-conditioned**: the system first forms a low-cost prior over likely evidence regions, then spends LLM calls only where observations are most useful. -![Monte Carlo Evidence Sampling Algorithm](Sirchmunk_MonteCarloSamplingAlgo.png "Monte Carlo Evidence Sampling: Three-act exploration–exploitation strategy") +![Budgeted Evidence Exploration](Sirchmunk_MonteCarloSamplingAlgo.png "Budgeted evidence exploration: three-layer workflow") -The algorithm operates in three phases: +The workflow has three layers: -1. **Phase 1 — Cast the Net (Exploration):** Fuzzy anchor matching combined with stratified random sampling. The system identifies seed regions of potential relevance while maintaining broad coverage through randomized probing — ensuring no high-value region is missed. +1. **Low-cost prior:** Lexical anchors, document-path structure, compiled summaries, historical source-grounded evidence, and lightweight corpus scans narrow the candidate subspace before expensive oracle calls. -2. **Phase 2 — Focus (Exploitation):** Gaussian importance sampling centered around high-scoring seeds from Phase 1. The sampling density concentrates on the most promising regions, extracting surrounding context and scoring each snippet for relevance. +2. **Budget-constrained sequential inference:** Candidate regions are proposed, observed by an LLM relevance oracle, and used to update the belief state until the budget-aware stopping rule says the evidence is sufficient. -3. **Phase 3 — Synthesize:** The top-K scored snippets are passed to the LLM, which synthesizes them into a coherent Region of Interest (ROI) summary with a confidence flag — enabling the pipeline to decide whether evidence is sufficient or a ReAct agent should be invoked for deeper exploration. +3. **Consolidation and synthesis:** Selected regions are merged into a compact source-grounded evidence set, synthesized into an answer, and optionally persisted as reusable knowledge for follow-up queries. **Key properties:** -- **Document-agnostic:** The same algorithm works equally well on a 2-page memo and a 500-page technical manual — no document-specific chunking heuristics needed. -- **Token-efficient:** Only the most relevant regions are sent to the LLM, dramatically reducing token consumption compared to full-document approaches. -- **Exploration-exploitation balance:** Random exploration prevents tunnel vision, while importance sampling ensures depth where it matters most. +- **Index-free over raw documents:** Search can run directly over dynamic files without pre-materializing a persistent embedding or chunk index. +- **Source-grounded:** The final answer is paired with traceable evidence regions instead of opaque vector hits. +- **Budget-aware:** LLM calls are spent adaptively on uncertain or high-value evidence regions, with explicit telemetry for cost and latency. ### ReAct Agent @@ -101,7 +134,7 @@ A KnowledgeCluster is a richly annotated object that captures the full cognitive | Field | Purpose | |:------|:--------| -| **Evidences** | Source-linked snippets extracted via Monte Carlo sampling, each with file path, summary, and raw text | +| **Evidences** | Source-linked evidence regions localized by LENS, each with file path, summary, and raw text | | **Content** | LLM-synthesized markdown with structured analysis and references | | **Patterns** | 3–5 distilled design principles or mechanisms identified from the evidence | | **Confidence** | A consensus score \[0, 1\] indicating the reliability of the cluster | @@ -123,7 +156,7 @@ A KnowledgeCluster is a richly annotated object that captures the full cognitive │ ┌──────────────────────────────┐ │ │ Phase 1–3: Full Search │ │ │ (keywords → retrieval → │ - │ │ Monte Carlo → LLM synth) │ + │ │ evidence localization → │ │ └──────────┬───────────────────┘ │ ▼ │ ┌──────────────────────────────┐ @@ -140,7 +173,7 @@ A KnowledgeCluster is a richly annotated object that captures the full cognitive 1. **Reuse Check (Phase 0):** Before any retrieval, the query is embedded and compared against all stored clusters via cosine similarity. If a high-confidence match is found, the existing cluster is returned instantly — saving LLM tokens and search time entirely. -2. **Creation (Phase 1–3):** When no reuse match is found, the full pipeline runs: keyword extraction, file retrieval, Monte Carlo evidence sampling, and LLM synthesis produce a new `KnowledgeCluster`. +2. **Creation (Phase 1–3):** When no reuse match is found, the full pipeline runs: keyword extraction, file retrieval, budgeted evidence localization, and LLM synthesis produce a new `KnowledgeCluster`. 3. **Persistence (Phase 5):** The cluster is stored in an in-memory DuckDB table and periodically flushed to Parquet files. Atomic writes and mtime-based reload ensure multi-process safety. @@ -190,3 +223,7 @@ Sirchmunk adheres to **SOLID principles**: - **Dependency Inversion** — High-level logic depends on abstractions For a comprehensive technical analysis, read the [Technical Deep Dive](/blog/technical-deep-dive/). + +--- + +> **Paper:** [LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents](https://arxiv.org/abs/2608.16185) (arXiv 2026) diff --git a/content/docs/guide/architecture.zh.md b/content/docs/guide/architecture.zh.md index 6975a7b..a11ba41 100644 --- a/content/docs/guide/architecture.zh.md +++ b/content/docs/guide/architecture.zh.md @@ -9,13 +9,39 @@ Sirchmunk 采用清晰分离的层次化架构,遵循**关注点分离**原则 ![Sirchmunk 架构](Sirchmunk_Architecture.png "Sirchmunk 高层架构图") +## LENS 框架 + +![LENS 框架](Sirchmunk_LENS_Framework.png "LENS:基于隐式证据空间的预算约束证据探索") + +**LENS**(Latent Evidence Exploration and Search,隐式证据探索与搜索)是一个无索引检索框架,将上下文内搜索形式化为在动态原始文档诱导的隐式证据空间上进行**预算约束的证据定位**。LENS 无需通过嵌入索引或分块存储预物化证据,而是维护一个查询条件化的候选证据置信度,并迭代地: + +1. 通过互补的词法、局部和探索性提案策略**提议**候选项 +2. 经由 LLM 相关性预言机**更新**置信度 +3. 在可控 token 预算下向高后验区域**收敛** + +这一形式化使搜索过程具备自适应性、预算感知能力,并完全基于源文档——无需任何预构建的索引基础设施。 + +> **论文:** [LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents](https://arxiv.org/abs/2608.16185)(arXiv 2026) + +## 知识图谱 + +![知识图谱](Sirchmunk_Knowledge_Graph.png "交互式知识簇可视化") + +**知识图谱**提供了自进化知识簇的交互式可视化,这些知识簇从搜索交互中增量构建。基于 Cytoscape.js 实现,它展示: + +- **知识簇关系** — 连接相关知识单元的语义边 +- **生命周期状态** — 萌芽 → 稳定 → 弃用 转换的可视化编码 +- **Leiden 元聚类** — 通过 Leiden 算法发现的更高层级社区结构 + +用户可以直接在 Web UI 中探索、筛选和深入查看知识簇。 + ## 核心组件 | 组件 | 说明 | |:------------------------|:-----------------------------------------------------------------------| -| **AgenticSearch** | 搜索编排器,具备 LLM 增强检索能力 | -| **KnowledgeBase** | 将原始结果转化为结构化知识簇并附带证据 | -| **EvidenceProcessor** | 基于蒙特卡洛重要性采样的证据处理 | +| **AgenticSearch** | 搜索编排器,支持 FAST / DEEP / FILENAME_ONLY 模式与预算感知的证据定位 | +| **KnowledgeBase** | 将源头证据簇持久化为可复用的热先验,供后续查询使用 | +| **EvidenceProcessor** | 将候选区域整合为紧凑、可追溯的证据单元 | | **GrepRetriever** | 高性能 _无索引_ 文件检索,支持并行处理 | | **OpenAIChat** | 统一 LLM 接口,支持流式与用量统计 | | **KnowledgeCompiler** | 离线文档编译为树索引和知识簇(Beta) | @@ -34,28 +60,35 @@ Sirchmunk 的核心是一个多阶段搜索管线,其设计原则是**每个 ### 阶段 1 — 并行探测 -四个独立的探测并发启动,收集多样化信号: +多个独立的探测并发启动,收集多样化信号: -1. **LLM 关键词提取** — LLM 将查询分解为多层级关键词,从粗粒度(高召回)到细粒度(高精度),每个关键词附带稀有度评分。这种多粒度方法确保同时捕获广泛主题和特定术语。 -2. **目录结构扫描** — 遍历文件系统收集路径元数据:文件名、大小、修改时间和内容预览。这是*智能推断*的基础——利用结构线索(命名规范、目录层级、文件类型)在读取内容之前就缩小最有可能的候选范围。 +1. **LLM 关键词提取** — LLM 将查询分解为多层级关键词,从粗粒度(高召回)到细粒度(高精度),每个关键词附带稀有度评分。 +2. **目录结构扫描** — 遍历文件系统收集路径元数据:文件名、大小、修改时间和内容预览——这是利用结构线索进行*智能推断*的基础。 3. **知识缓存查找** — 在现有知识簇中进行部分匹配搜索,复用先前获取的知识。 -4. **指定路径上下文加载** — 从缓存中加载已知路径的预计算上下文。 +4. **编译产物加载** — 加载可用的预计算树索引、主题图和摘要索引。 + +### 阶段 2 — 多路 DEEP 检索 -### 阶段 2 — 检索与排序 +v0.2.0 中,DEEP 模式并行运行 **5 条互补检索路径**: -两种互补策略并行运行: +| 路径 | 信号 | +|:-----|:-----| +| **词法路径** | IDF 加权关键词搜索原始文件内容 | +| **实体路径** | 命名实体和标识符的精确匹配查找 | +| **目录路径** | 基于 LLM 引导的元数据评估进行结构排序 | +| **结构路径** | 利用编译产物的启发式文档树导航 | +| **主题图路径** | 跨文档主题图遍历,用于多跳发现 | -- **基于内容的检索** — IDF 加权关键词搜索原始文件内容 -- **基于结构的排序** — LLM 引导的候选文件元数据评估 +所有路径的结果通过**置信度加权的倒数排名融合(RRF)**进行融合。**软路由折叠**机制会动态禁用低收益路径——当单一路径以高置信度和强边际达成共识时,其余路径被折叠以减少延迟和 token 消耗,同时保持回答质量。 -### 阶段 3 — 知识簇构建 +### 阶段 3 — 证据定位与知识簇构建 -结果被合并、去重,并通过蒙特卡洛证据采样处理。LLM 将证据片段合成为结构化的知识簇。 +结果被合并、去重,并通过预算约束的证据定位处理。LLM 将证据片段合成为结构化的知识簇。 ### 阶段 4 — 摘要或 ReAct 精炼 -- **找到证据** → LLM 生成结构化简报 -- **未找到证据** → ReAct 智能体启动进行迭代探索 +- **找到证据** → LLM 生成带有源头链接证据的结构化简报 +- **未找到证据** → ReAct 智能体启动,使用跳数感知策略进行迭代探索 ### 阶段 5 — 持久化 @@ -63,25 +96,25 @@ Sirchmunk 的核心是一个多阶段搜索管线,其设计原则是**每个 ## 核心算法 -### 蒙特卡洛证据采样 +### 预算约束的证据探索 -传统检索系统要么读取完整文档,要么依赖固定大小的分块,导致 Token 浪费或上下文丢失。Sirchmunk 借鉴**蒙特卡洛方法**,采用了截然不同的策略——将证据提取视为一个**采样问题**而非解析问题。 +传统检索系统要么读取完整文档,要么依赖固定大小的分块,导致 Token 浪费或上下文丢失。LENS 将相关证据视为**隐式且查询条件化的**:系统首先以低成本形成关于可能证据区域的先验,然后仅在最有价值的观测位置花费 LLM 调用。 -![蒙特卡洛证据采样算法](Sirchmunk_MonteCarloSamplingAlgo.png "蒙特卡洛证据采样:三阶段启发式探索-利用策略") +![预算约束的证据探索](Sirchmunk_MonteCarloSamplingAlgo.png "预算约束的证据探索:三层工作流") -该算法分为三个阶段: +该工作流分为三个层次: -1. **第一阶段 — 撒网(探索):** 模糊锚定匹配结合分层随机采样。系统在识别潜在相关种子区域的同时,通过随机探测保持广泛覆盖,确保不会遗漏高价值区域。 +1. **低成本先验:** 词法锚点、文档路径结构、编译摘要、历史源头证据和轻量级语料库扫描,在昂贵的预言机调用之前缩小候选子空间。 -2. **第二阶段 — 聚焦(利用):** 以第一阶段高分种子为中心进行高斯重要性采样。采样密度集中在最有前景的区域,提取上下文并对每个片段评分。 +2. **预算约束的序贯推断:** 提议候选区域,由 LLM 相关性预言机进行观测,并用于更新置信状态,直到预算感知的停止规则判定证据充分。 -3. **第三阶段 — 合成:** 将 Top-K 评分片段传递给 LLM,合成为连贯的兴趣区域(ROI)摘要,并附带置信度标志——使管线能够判断证据是否充分,或是否需要启用 ReAct 智能体进行更深层的自适应检索。 +3. **整合与合成:** 选定区域被合并为紧凑的源头证据集,合成为答案,并可选地持久化为可复用知识供后续查询使用。 **核心特性:** -- **文档无关性:** 同一算法在 2 页备忘录和 500 页技术手册上同样有效,无需针对特定文档的分块启发式规则。 -- **Token 高效:** 仅将最相关的区域发送给 LLM,相比全文档方案大幅降低 Token 消耗。 -- **探索-利用平衡:** 随机探索防止视野盲区,重要性采样确保在关键区域深入挖掘。 +- **基于原始文档的无索引搜索:** 可直接在动态文件上运行搜索,无需预物化持久化嵌入或分块索引。 +- **源头可追溯:** 最终答案附带可追溯的证据区域,而非不透明的向量命中。 +- **预算感知:** LLM 调用自适应地花费在不确定或高价值的证据区域上,并提供显式的成本和延迟遥测。 ### ReAct 智能体 @@ -101,7 +134,7 @@ KnowledgeCluster 是一个丰富标注的对象,完整记录了单次搜索周 | 字段 | 用途 | |:-----|:-----| -| **Evidences(证据)** | 通过蒙特卡洛采样提取的源文件片段,包含文件路径、摘要和原始文本 | +| **Evidences(证据)** | 通过 LENS 定位的源文件证据区域,包含文件路径、摘要和原始文本 | | **Content(内容)** | LLM 合成的结构化 Markdown 分析,附带引用 | | **Patterns(模式)** | 从证据中提炼的 3–5 条设计原则或核心机制 | | **Confidence(置信度)** | 共识评分 \[0, 1\],指示知识簇的可靠性 | @@ -123,7 +156,7 @@ KnowledgeCluster 是一个丰富标注的对象,完整记录了单次搜索周 │ ┌───────────────────────────────┐ │ │ 阶段 1–3:完整搜索 │ │ │ (关键词 → 检索 → │ - │ │ 蒙特卡洛采样 → LLM 合成) │ + │ │ 证据定位 → LLM 合成) │ │ └──────────┬────────────────────┘ │ ▼ │ ┌───────────────────────────────┐ @@ -140,7 +173,7 @@ KnowledgeCluster 是一个丰富标注的对象,完整记录了单次搜索周 1. **复用检查(阶段 0):** 在任何检索开始之前,查询会被嵌入并通过余弦相似度与所有已存储知识簇进行比对。若发现高置信度匹配,系统直接返回已有知识簇——完全省去 LLM 推理和搜索开销。 -2. **创建(阶段 1–3):** 当无复用匹配时,完整管线运行:关键词提取、文件检索、蒙特卡洛证据采样、LLM 合成,最终生成新的 `KnowledgeCluster`。 +2. **创建(阶段 1–3):** 当无复用匹配时,完整管线运行:关键词提取、文件检索、预算约束证据定位、LLM 合成,最终生成新的 `KnowledgeCluster`。 3. **持久化(阶段 5):** 知识簇存储在内存中的 DuckDB 表中,并定期刷写为 Parquet 文件。原子写入和基于文件修改时间的重载机制确保多进程安全。 @@ -190,3 +223,7 @@ Sirchmunk 遵循 **SOLID 原则**: - **依赖倒置** — 高层逻辑依赖抽象 欲了解全面的技术分析,请阅读 [技术深度报告](/zh/blog/technical-deep-dive/)。 + +--- + +> **论文:** [LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents](https://arxiv.org/abs/2608.16185)(arXiv 2026) diff --git a/content/docs/guide/cli.md b/content/docs/guide/cli.md index 79b9ff4..3f515e9 100644 --- a/content/docs/guide/cli.md +++ b/content/docs/guide/cli.md @@ -52,14 +52,14 @@ sirchmunk serve --host 0.0.0.0 --port 8000 Perform search queries directly from the terminal. ```bash -# Search in current directory (FAST mode by default) +# Search in current directory (DEEP mode with rich output by default) sirchmunk search "How does authentication work?" # Search in specific paths sirchmunk search "find all API endpoints" ./src ./docs -# DEEP mode: comprehensive agentic retrieval analysis -sirchmunk search "database architecture" --mode DEEP +# FAST mode: greedy search with 2 LLM calls +sirchmunk search "database architecture" --mode FAST # Quick filename search (no LLM required) sirchmunk search "config" --mode FILENAME_ONLY diff --git a/content/docs/guide/cli.zh.md b/content/docs/guide/cli.zh.md index 07aa3a3..7110aef 100644 --- a/content/docs/guide/cli.zh.md +++ b/content/docs/guide/cli.zh.md @@ -52,14 +52,14 @@ sirchmunk serve --host 0.0.0.0 --port 8000 直接从终端执行搜索查询。 ```bash -# 在当前目录中搜索(默认 FAST 模式) +# 在当前目录中搜索(默认 DEEP 模式,富文本输出) sirchmunk search "How does authentication work?" # 在指定路径中搜索 sirchmunk search "find all API endpoints" ./src ./docs -# DEEP 模式:智能体检索全面分析 -sirchmunk search "数据库架构" --mode DEEP +# FAST 模式:贪心搜索,2 次 LLM 调用 +sirchmunk search "数据库架构" --mode FAST # 快速文件名搜索(无需 LLM) sirchmunk search "config" --mode FILENAME_ONLY diff --git a/content/docs/guide/configuration.md b/content/docs/guide/configuration.md index b095f3c..4b729f1 100644 --- a/content/docs/guide/configuration.md +++ b/content/docs/guide/configuration.md @@ -15,7 +15,7 @@ Sirchmunk is configured through environment variables stored in a `.env` file. A |----------|-------------|---------| | `LLM_API_KEY` | Your LLM API key (required for FAST and DEEP modes) | — | | `LLM_BASE_URL` | OpenAI-compatible API base URL | `https://api.openai.com/v1` | -| `LLM_MODEL` | Model name to use | `gpt-4o` | +| `LLM_MODEL_NAME` | Model name to use | `gpt-5.2` | ### Search Configuration @@ -26,7 +26,18 @@ Sirchmunk is configured through environment variables stored in a `.env` file. A | `SIRCHMUNK_MAX_DEPTH` | Maximum directory traversal depth | `10` | | `SIRCHMUNK_TOP_K_FILES` | Number of top files to analyze | `20` | | `SIRCHMUNK_MAX_CONCURRENT_SEARCHES` | Max concurrent search tasks | `3` | -| `SIRCHMUNK_ENABLE_CLUSTER_REUSE` | Enable knowledge cluster reuse | `false` | +| `SIRCHMUNK_ENABLE_CLUSTER_REUSE` | Enable knowledge cluster reuse | `true` | + +### Retrieval Cost Configuration + +| Variable | Description | Default | +|----------|-------------|--------| +| `GREP_MAX_FILESIZE_MB` | Per-file size cap (MB); files over this limit are skipped on the query hot path | `64` | +| `GREP_RGA_ADAPTERS` | Allowed rga adapters (bounded only; archives disabled on hot path) | `poppler,pandoc,postprocpagebreaks` | +| `GREP_TIERED_SCAN` | Enable tiered scan: fast native-rg pass + bounded rga pass for rich formats | `true` | +| `GREP_TEXT_TIMEOUT` | Timeout (seconds) for the native rg text pass | `15.0` | +| `GREP_TIMEOUT` | Timeout (seconds) for the rga rich pass | `60.0` | +| `GREP_RICH_EXTENSIONS` | File extensions routed to the rga rich pass | `pdf,docx,epub,odt` | ### Chat Configuration @@ -74,7 +85,7 @@ When invoking search (via SDK, CLI, or API), the following parameters are availa |-----------|------|---------|-------------| | `query` | `string` | *required* | Search query or question | | `paths` | `string \| string[]` | *optional* | Directories or files to search; falls back to `SIRCHMUNK_SEARCH_PATHS`, then cwd | -| `mode` | `string` | `FAST` | `FAST` (greedy, 2-5s), `DEEP` (agentic retrieval, 10-30s), or `FILENAME_ONLY` | +| `mode` | `string` | `DEEP` | `DEEP` (agentic retrieval with budgeted evidence exploration), `FAST` (greedy, 2-5s), or `FILENAME_ONLY` | | `max_depth` | `int` | `null` | Maximum directory depth | | `top_k_files` | `int` | `null` | Number of top files to return | | `enable_dir_scan` | `bool` | `true` | Enable directory scanning | @@ -82,7 +93,7 @@ When invoking search (via SDK, CLI, or API), the following parameters are availa | `max_token_budget` | `int` | `null` | DEEP mode token budget (default 128K when unset) | | `include_patterns` | `string[]` | `null` | File glob patterns to include | | `exclude_patterns` | `string[]` | `null` | File glob patterns to exclude | -| `return_context` | `bool` | `false` | Return SearchContext with cluster and telemetry | +| `response_format` | `string` | `rich` | `"rich"` Markdown report, `"minimal"` short answer, `"context"` SearchContext, or `"json"` serialized context | > [!NOTE] -> `FILENAME_ONLY` mode does not require an LLM API key. `FAST` and `DEEP` modes require a configured LLM. `FAST` mode uses a greedy strategy with 2-level keyword cascade and early stopping — approximately **10x faster** than `DEEP` mode. +> `FILENAME_ONLY` mode does not require an LLM API key. `FAST` and `DEEP` modes require a configured LLM. Default mode is `DEEP`, which performs budgeted evidence exploration with multi-path retrieval. diff --git a/content/docs/guide/configuration.zh.md b/content/docs/guide/configuration.zh.md index fdd1762..7d80d53 100644 --- a/content/docs/guide/configuration.zh.md +++ b/content/docs/guide/configuration.zh.md @@ -15,7 +15,7 @@ Sirchmunk 通过存储在 `.env` 文件中的环境变量进行配置。运行 ` |------|------|--------| | `LLM_API_KEY` | LLM API 密钥(FAST 和 DEEP 模式必需) | — | | `LLM_BASE_URL` | OpenAI 兼容 API 基础 URL | `https://api.openai.com/v1` | -| `LLM_MODEL` | 使用的模型名称 | `gpt-4o` | +| `LLM_MODEL_NAME` | 使用的模型名称 | `gpt-5.2` | ### 搜索配置 @@ -26,7 +26,18 @@ Sirchmunk 通过存储在 `.env` 文件中的环境变量进行配置。运行 ` | `SIRCHMUNK_MAX_DEPTH` | 最大目录遍历深度 | `10` | | `SIRCHMUNK_TOP_K_FILES` | 分析的最大文件数 | `20` | | `SIRCHMUNK_MAX_CONCURRENT_SEARCHES` | 最大并发搜索任务数 | `3` | -| `SIRCHMUNK_ENABLE_CLUSTER_REUSE` | 启用知识簇复用 | `false` | +| `SIRCHMUNK_ENABLE_CLUSTER_REUSE` | 启用知识簇复用 | `true` | + +### 检索成本配置 + +| 变量 | 描述 | 默认值 | +|------|------|--------| +| `GREP_MAX_FILESIZE_MB` | 单文件大小上限(MB);超过此限制的文件在查询热路径上被跳过 | `64` | +| `GREP_RGA_ADAPTERS` | 允许的 rga 适配器(仅有界的;查询热路径禁用归档解压) | `poppler,pandoc,postprocpagebreaks` | +| `GREP_TIERED_SCAN` | 启用分层扫描:快速原生 rg 扫描 + 有界 rga 富格式扫描 | `true` | +| `GREP_TEXT_TIMEOUT` | 原生 rg 文本扫描超时(秒) | `15.0` | +| `GREP_TIMEOUT` | rga 富格式扫描超时(秒) | `60.0` | +| `GREP_RICH_EXTENSIONS` | 路由到 rga 富格式扫描的文件扩展名 | `pdf,docx,epub,odt` | ### 对话配置 @@ -74,7 +85,7 @@ Sirchmunk 通过存储在 `.env` 文件中的环境变量进行配置。运行 ` |------|------|--------|------| | `query` | `string` | *必填* | 搜索查询或问题 | | `paths` | `string \| string[]` | *可选* | 要搜索的目录或文件;未设置时依次回退到 `SIRCHMUNK_SEARCH_PATHS`、当前工作目录 | -| `mode` | `string` | `FAST` | `FAST`(贪心搜索,2-5s)、`DEEP`(智能体检索,10-30s)或 `FILENAME_ONLY` | +| `mode` | `string` | `DEEP` | `DEEP`(预算约束的证据探索智能体检索)、`FAST`(贪心搜索,2-5s)或 `FILENAME_ONLY` | | `max_depth` | `int` | `null` | 最大目录深度 | | `top_k_files` | `int` | `null` | 返回的文件数量 | | `enable_dir_scan` | `bool` | `true` | 是否启用目录扫描 | @@ -82,7 +93,7 @@ Sirchmunk 通过存储在 `.env` 文件中的环境变量进行配置。运行 ` | `max_token_budget` | `int` | `null` | DEEP 模式 token 预算(未设置时默认 128K) | | `include_patterns` | `string[]` | `null` | 要包含的文件 glob 模式 | | `exclude_patterns` | `string[]` | `null` | 要排除的文件 glob 模式 | -| `return_context` | `bool` | `false` | 返回包含知识簇与遥测的 SearchContext | +| `response_format` | `string` | `rich` | `"rich"` Markdown 报告、`"minimal"` 简短回答、`"context"` SearchContext 对象、`"json"` 序列化上下文 | > [!NOTE] -> `FILENAME_ONLY` 模式不需要 LLM API 密钥。`FAST` 和 `DEEP` 模式需要配置 LLM。`FAST` 模式采用贪心策略,结合两级关键词级联与 early stopping,速度约为 `DEEP` 模式的 **10 倍**。 +> `FILENAME_ONLY` 模式不需要 LLM API 密钥。`FAST` 和 `DEEP` 模式需要配置 LLM。默认模式为 `DEEP`,执行多路检索的预算约束证据探索。 diff --git a/content/docs/guide/docker.md b/content/docs/guide/docker.md index 03c598e..9b97d97 100644 --- a/content/docs/guide/docker.md +++ b/content/docs/guide/docker.md @@ -75,5 +75,8 @@ print(response.json()) | `SIRCHMUNK_SEARCH_PATHS` | Default search paths (comma-separated) | — | | `SIRCHMUNK_MAX_CONCURRENT_SEARCHES` | Max concurrent search tasks | `3` | +> [!TIP] +> The image tag `ubuntu22.04-py312-0.0.7` is the latest pre-built image at the time of writing. Check the [Alibaba Cloud Container Registry](https://cr.console.aliyun.com/) or the [Sirchmunk README](https://github.com/modelscope/sirchmunk#-docker-deployment) for the latest available tag. + > [!TIP] > For full Docker parameters and advanced usage, see the [docker/README.md](https://github.com/modelscope/sirchmunk/blob/main/docker/README.md) in the Sirchmunk repository. diff --git a/content/docs/guide/docker.zh.md b/content/docs/guide/docker.zh.md index 44bc025..1f0e354 100644 --- a/content/docs/guide/docker.zh.md +++ b/content/docs/guide/docker.zh.md @@ -75,5 +75,8 @@ print(response.json()) | `SIRCHMUNK_SEARCH_PATHS` | 默认搜索路径(逗号分隔) | — | | `SIRCHMUNK_MAX_CONCURRENT_SEARCHES` | 最大并发搜索任务数 | `3` | +> [!TIP] +> 镜像标签 `ubuntu22.04-py312-0.0.7` 是截至撰写时最新的预构建镜像。请查看 [阿里云容器镜像服务](https://cr.console.aliyun.com/) 或 [Sirchmunk README](https://github.com/modelscope/sirchmunk#-docker-deployment) 获取最新可用标签。 + > [!TIP] > 完整 Docker 参数和高级用法,请参阅 Sirchmunk 仓库中的 [docker/README.md](https://github.com/modelscope/sirchmunk/blob/main/docker/README.md)。 diff --git a/content/docs/guide/python-sdk.md b/content/docs/guide/python-sdk.md index 1566c66..0b51dc9 100644 --- a/content/docs/guide/python-sdk.md +++ b/content/docs/guide/python-sdk.md @@ -22,23 +22,23 @@ from sirchmunk.llm import OpenAIChat llm = OpenAIChat( api_key="your-api-key", base_url="your-base-url", # e.g., https://api.openai.com/v1 - model="your-model-name" # e.g., gpt-4o + model="your-model-name" # e.g., gpt-5.2 ) async def main(): searcher = AgenticSearch(llm=llm) - # FAST mode (default): greedy search, 2 LLM calls, 2-5s + # DEEP mode (default): rich Markdown report with budgeted evidence exploration result: str = await searcher.search( query="How does transformer attention work?", paths=["/path/to/documents"], ) - # DEEP mode: comprehensive agentic retrieval analysis, 10-30s - result_deep: str = await searcher.search( + # FAST mode: greedy search, 2 LLM calls, 2-5s + result_fast: str = await searcher.search( query="How does transformer attention work?", paths=["/path/to/documents"], - mode="DEEP", + mode="FAST", ) print(result) @@ -60,13 +60,13 @@ asyncio.run(main()) result = await searcher.search( query="database connection pooling", # Required: search question paths=["/path/to/project/src"], # Optional: directories (env, then cwd) - mode="FAST", # FAST (default), DEEP, or FILENAME_ONLY + mode="DEEP", # DEEP (default), FAST, or FILENAME_ONLY max_depth=10, # Max directory depth top_k_files=20, # Number of top files max_loops=10, # Max search loops include_patterns=["*.py", "*.java"], # File patterns to include exclude_patterns=["*test*", "*__pycache__*"], # Patterns to exclude - return_context=True, # Return full SearchContext + response_format="context", # Return format: rich, minimal, context, json ) ``` @@ -87,7 +87,7 @@ print(result) result = await searcher.search( query="...", paths=["..."], - return_context=True, + response_format="context", ) # Access context metadata @@ -111,9 +111,12 @@ for usage in searcher.llm_usages: Sirchmunk works with any OpenAI-compatible API endpoint: -- **OpenAI** — GPT-4, GPT-4o, GPT-5.2 -- **MiniMax** — MiniMax-M2.7, MiniMax-M2.7-highspeed +- **OpenAI** — GPT-4o, GPT-5.2 +- **MiniMax** — MiniMax-M3, MiniMax-M2.7, MiniMax-M2.7-highspeed - **DeepSeek** — DeepSeek-V3, DeepSeek-R1 and other DeepSeek chat models +- **Google Gemini**, **Zhipu (GLM)**, **Baichuan**, **Yi**, **SiliconFlow**, **Volcengine** +- **Moonshot**, **Mistral**, **Groq**, **Together AI**, **Cohere** +- **Azure OpenAI** - **Local models** — Ollama, llama.cpp, vLLM, SGLang - **Claude** — Via API proxy - **Any OpenAI-compatible endpoint** diff --git a/content/docs/guide/python-sdk.zh.md b/content/docs/guide/python-sdk.zh.md index 975d57b..5ab3551 100644 --- a/content/docs/guide/python-sdk.zh.md +++ b/content/docs/guide/python-sdk.zh.md @@ -22,23 +22,23 @@ from sirchmunk.llm import OpenAIChat llm = OpenAIChat( api_key="your-api-key", base_url="your-base-url", # 如 https://api.openai.com/v1 - model="your-model-name" # 如 gpt-4o + model="your-model-name" # 如 gpt-5.2 ) async def main(): searcher = AgenticSearch(llm=llm) - # FAST 模式(默认):贪心搜索,2 次 LLM 调用,2-5s + # DEEP 模式(默认):富文本 Markdown 报告,预算约束证据探索 result: str = await searcher.search( query="How does transformer attention work?", paths=["/path/to/documents"], ) - # DEEP 模式:智能体检索全面分析,10-30s - result_deep: str = await searcher.search( + # FAST 模式:贪心搜索,2 次 LLM 调用,2-5s + result_fast: str = await searcher.search( query="How does transformer attention work?", paths=["/path/to/documents"], - mode="DEEP", + mode="FAST", ) print(result) @@ -57,13 +57,13 @@ asyncio.run(main()) result = await searcher.search( query="database connection pooling", # 必填:搜索问题 paths=["/path/to/project/src"], # 可选:省略时回退到 SIRCHMUNK_SEARCH_PATHS → cwd - mode="FAST", # FAST(默认)、DEEP 或 FILENAME_ONLY + mode="DEEP", # DEEP(默认)、FAST 或 FILENAME_ONLY max_depth=10, # 最大目录深度 top_k_files=20, # 最大文件数 max_loops=10, # ReAct 最大迭代次数(DEEP 模式) include_patterns=["*.py", "*.java"], # 要包含的文件模式 exclude_patterns=["*test*", "*__pycache__*"], # 要排除的文件模式 - return_context=True, # 返回 SearchContext(含 KnowledgeCluster 和遥测数据) + response_format="context", # 返回格式:rich、minimal、context、json ) ``` @@ -84,7 +84,7 @@ print(result) result = await searcher.search( query="...", paths=["..."], - return_context=True, + response_format="context", ) # 访问上下文元数据 @@ -108,9 +108,12 @@ for usage in searcher.llm_usages: Sirchmunk 适用于任何 OpenAI 兼容的 API 端点: -- **OpenAI** — GPT-4、GPT-4o、GPT-5.2 -- **MiniMax** — MiniMax-M2.7、MiniMax-M2.7-highspeed +- **OpenAI** — GPT-4o、GPT-5.2 +- **MiniMax** — MiniMax-M3、MiniMax-M2.7、MiniMax-M2.7-highspeed - **DeepSeek** — DeepSeek-V3、DeepSeek-R1 及其他 DeepSeek 对话模型 +- **Google Gemini**、**智谱(GLM)**、**百川**、**零一万物**、**硅基流动**、**火山引擎** +- **Moonshot**、**Mistral**、**Groq**、**Together AI**、**Cohere** +- **Azure OpenAI** - **本地模型** — Ollama、llama.cpp、vLLM、SGLang - **Claude** — 通过 API 代理 - **其他供应商** — 提供 OpenAI 兼容 HTTP API 的服务商 diff --git a/content/docs/guide/web-ui.md b/content/docs/guide/web-ui.md index 9c3b0cb..f485f68 100644 --- a/content/docs/guide/web-ui.md +++ b/content/docs/guide/web-ui.md @@ -66,6 +66,15 @@ Browse and manage knowledge clusters: - Track lifecycle states (Emerging → Stable → Deprecated) - Monitor hotness scores and query histories +### Knowledge Graph + +Interactive visualization of knowledge clusters, their relationships and lifecycle states. Accessible from the sidebar navigation. + +- **Cluster network** — Cytoscape.js-powered graph showing semantic edges between clusters +- **Lifecycle encoding** — Visual distinction of Emerging, Stable, and Deprecated states +- **Leiden meta-clustering** — Higher-level community structure overlay +- **Interactive exploration** — Click, zoom, and filter clusters in real-time + ### Monitor — System Dashboard Real-time system health and usage metrics: diff --git a/content/docs/guide/web-ui.zh.md b/content/docs/guide/web-ui.zh.md index fc5f62d..73a1ec7 100644 --- a/content/docs/guide/web-ui.zh.md +++ b/content/docs/guide/web-ui.zh.md @@ -66,6 +66,15 @@ python scripts/stop_web.py - 追踪生命周期状态(萌芽 → 稳定 → 弃用) - 监控热度分数和查询历史 +### 知识图谱 + +知识簇的交互式可视化,展示知识簇之间的关系和生命周期状态。可从侧边栏导航访问。 + +- **簇网络** — 基于 Cytoscape.js 的图形,展示知识簇之间的语义边 +- **生命周期编码** — 萌芽、稳定和弃用状态的可视化区分 +- **Leiden 元聚类** — 更高层级的社区结构叠加层 +- **交互式探索** — 实时点击、缩放和筛选知识簇 + ### 监控 — 系统仪表盘 实时系统健康和使用指标: diff --git a/content/docs/reference/i18n.md b/content/docs/reference/i18n.md index 620905e..d1d1a64 100644 --- a/content/docs/reference/i18n.md +++ b/content/docs/reference/i18n.md @@ -14,7 +14,7 @@ Sirchmunk takes an **indexless approach**: 1. **No pre-indexing**: Direct file search without vector database setup 2. **Self-evolving**: Knowledge clusters evolve based on search patterns 3. **Multi-level retrieval**: Adaptive keyword granularity for better recall -4. **Evidence-based**: Monte Carlo sampling for precise content extraction +4. **Evidence-based**: Budgeted evidence localization for precise, source-grounded extraction ## What LLM providers are supported? @@ -23,7 +23,9 @@ Any OpenAI-compatible API endpoint, including: - **OpenAI** (GPT-4, GPT-4o, GPT-5.2) - **Local models** served via Ollama, llama.cpp, vLLM, SGLang - **Claude** via API proxy -- **MiniMax**, **DeepSeek**, **Moonshot**, **Mistral**, **Groq**, **Together AI**, **Cohere**, **Google Gemini**, **Zhipu (GLM)**, **Baichuan**, **Yi**, **SiliconFlow**, **Volcengine**, **Azure OpenAI** +- **MiniMax** (MiniMax-M3, MiniMax-M2.7, MiniMax-M2.7-highspeed) +- **DeepSeek**, **Moonshot**, **Mistral**, **Groq**, **Together AI**, **Cohere** +- **Google Gemini**, **Zhipu (GLM)**, **Baichuan**, **Yi**, **SiliconFlow**, **Volcengine**, **Azure OpenAI** - Any other OpenAI-compatible provider ## How do I add documents to search? @@ -57,7 +59,7 @@ Three ways: ## Does FILENAME_ONLY mode require an LLM? -No. `FILENAME_ONLY` mode performs fast filename-based search without any LLM calls. Both `FAST` and `DEEP` modes require a configured LLM API key. `FAST` mode (default) uses a greedy strategy with early stopping for ~10x faster retrieval than `DEEP` mode. +No. `FILENAME_ONLY` mode performs fast filename-based search without any LLM calls. Both `FAST` and `DEEP` modes require a configured LLM API key. Default mode is now `DEEP`, which performs budgeted evidence exploration with multi-path retrieval. ## What file formats are supported? @@ -78,3 +80,7 @@ Knowledge clusters follow a natural lifecycle: 2. **Reuse** — Similar queries match and enhance existing clusters 3. **Maturation** — Repeated validation transitions clusters from Emerging to Stable 4. **Deprecation** — Unsupported clusters transition to Contested or Deprecated + +## What is the LENS paper and how does it relate to Sirchmunk? + +LENS (Latent Evidence Exploration and Search) is the research paper that formalizes Sirchmunk's core retrieval algorithm. Published as arXiv:2608.16185, it describes how Sirchmunk performs budgeted evidence localization over dynamic raw documents without pre-built indexes. Read the full paper at https://arxiv.org/abs/2608.16185. diff --git a/content/docs/reference/i18n.zh.md b/content/docs/reference/i18n.zh.md index 33a61c1..175dd33 100644 --- a/content/docs/reference/i18n.zh.md +++ b/content/docs/reference/i18n.zh.md @@ -14,7 +14,7 @@ Sirchmunk 采用**无索引方法**: 1. **无需预索引**:无需设置向量数据库,直接文件搜索 2. **自进化**:知识簇基于搜索模式不断演进 3. **多级检索**:自适应关键词粒度以获得更好的召回率 -4. **基于证据**:蒙特卡洛采样实现精确内容提取 +4. **基于证据**:预算约束的证据定位实现精确、源头可追溯的内容提取 ## 支持哪些 LLM 供应商? @@ -23,7 +23,9 @@ Sirchmunk 采用**无索引方法**: - **OpenAI**(GPT-4、GPT-4o、GPT-5.2) - 通过 Ollama、llama.cpp、vLLM、SGLang 提供的**本地模型** - **Claude** 通过 API 代理 -- **MiniMax**、**DeepSeek**、**Moonshot**、**Mistral**、**Groq**、**Together AI**、**Cohere**、**Google Gemini**、**智谱(GLM)**、**百川**、**零一万物**、**硅基流动**、**火山引擎**、**Azure OpenAI** +- **MiniMax**(MiniMax-M3、MiniMax-M2.7、MiniMax-M2.7-highspeed) +- **DeepSeek**、**Moonshot**、**Mistral**、**Groq**、**Together AI**、**Cohere** +- **Google Gemini**、**智谱(GLM)**、**百川**、**零一万物**、**硅基流动**、**火山引擎**、**Azure OpenAI** - 其他任何 OpenAI 兼容供应商 ## 如何添加搜索文档? @@ -57,7 +59,7 @@ result = await searcher.search( ## FILENAME_ONLY 模式需要 LLM 吗? -不需要。`FILENAME_ONLY` 模式执行快速文件名搜索,不进行任何 LLM 调用。`FAST` 和 `DEEP` 模式均需要配置 LLM API 密钥。`FAST` 模式(默认)采用贪心策略和 early stopping,速度约为 `DEEP` 模式的 10 倍。 +不需要。`FILENAME_ONLY` 模式执行快速文件名搜索,不进行任何 LLM 调用。`FAST` 和 `DEEP` 模式均需要配置 LLM API 密钥。默认模式现在是 `DEEP`,执行多路检索的预算约束证据探索。 ## 支持哪些文件格式? @@ -77,4 +79,8 @@ Sirchmunk 利用 ripgrep-all 搜索 **100 多种文件格式**,包括: 1. **创建** — 新证据生成新的知识簇 2. **复用** — 相似查询匹配并增强现有知识簇 3. **成熟** — 经过多次查询验证,知识簇从"萌芽"过渡到"稳定" -4. **弃用** — 当底层数据变化且证据不再支持时,知识簇可能过渡到"有争议"或"已弃用" +4. **弃用** — 当底层数据变化且证据不再支持时,知识簇可能过渡到“有争议”或“已弃用” + +## LENS 论文是什么?它与 Sirchmunk 有什么关系? + +LENS(隐式证据探索与搜索)是将 Sirchmunk 核心检索算法形式化的研究论文。论文编号 arXiv:2608.16185,描述了 Sirchmunk 如何在无需预建索引的条件下,对动态原始文档进行预算约束的证据定位。阅读完整论文:https://arxiv.org/abs/2608.16185。 diff --git a/content/showcase/knowledge-graph/Sirchmunk_Knowledge_Graph.png b/content/showcase/knowledge-graph/Sirchmunk_Knowledge_Graph.png new file mode 100644 index 0000000..f53916d Binary files /dev/null and b/content/showcase/knowledge-graph/Sirchmunk_Knowledge_Graph.png differ diff --git a/content/showcase/knowledge-graph/index.md b/content/showcase/knowledge-graph/index.md new file mode 100644 index 0000000..6724fd2 --- /dev/null +++ b/content/showcase/knowledge-graph/index.md @@ -0,0 +1,14 @@ +--- +title: "Knowledge Graph Visualization" +summary: "Interactive knowledge graph powered by Cytoscape.js — visualize self-evolving knowledge clusters, their relationships, and lifecycle states." +date: 2026-07-21 +image: + filename: Sirchmunk_Knowledge_Graph.png + caption: "Sirchmunk Knowledge Graph" +--- + +Sirchmunk's Web UI includes an interactive knowledge graph that visualizes the self-evolving knowledge clusters built incrementally from search interactions. Powered by [Cytoscape.js](https://js.cytoscape.org/), the graph presents semantic relationships between clusters and their lifecycle states — from **Emerging** to **Stable** — giving you a live view of how the system's intelligence grows over time. + +Knowledge clusters are persisted in **DuckDB + Parquet** and evolve through a four-phase runtime cycle: connect & merge, edge refresh, meta-cluster discovery (via the Leiden community detection algorithm), and global recalibration. The graph UI lets you explore these communities interactively, filter by lifecycle stage, and trace how queries contribute to cluster formation. + +![Sirchmunk Knowledge Graph](Sirchmunk_Knowledge_Graph.png "Knowledge Graph — Interactive visualization of knowledge clusters with lifecycle stages") diff --git a/content/showcase/knowledge-graph/index.zh.md b/content/showcase/knowledge-graph/index.zh.md new file mode 100644 index 0000000..9f209c8 --- /dev/null +++ b/content/showcase/knowledge-graph/index.zh.md @@ -0,0 +1,14 @@ +--- +title: "知识图谱可视化" +summary: "基于 Cytoscape.js 的交互式知识图谱 — 可视化自进化知识聚类、语义关联与生命周期状态。" +date: 2026-07-21 +image: + filename: Sirchmunk_Knowledge_Graph.png + caption: "Sirchmunk 知识图谱" +--- + +Sirchmunk 的 Web UI 内置了一套交互式知识图谱,可视化展示随搜索交互增量构建的自进化知识聚类。图谱由 [Cytoscape.js](https://js.cytoscape.org/) 驱动,直观呈现聚类之间的语义关联及其生命周期状态 — 从 **Emerging(新兴)** 到 **Stable(稳定)** — 让你实时观察系统智能如何随使用不断增长。 + +知识聚类持久化于 **DuckDB + Parquet**,通过四阶段运行时周期持续进化:连接合并、边刷新、元聚类发现(基于 Leiden 社区发现算法)与全局重校准。图谱界面支持交互式探索社区结构、按生命周期阶段过滤,并追踪查询如何促成聚类的形成与演化。 + +![Sirchmunk 知识图谱](Sirchmunk_Knowledge_Graph.png "知识图谱 — 知识聚类交互式可视化与生命周期阶段")