Skip to content

[Docs] Update guide - #2

Merged
wangxingjun778 merged 1 commit into
mainfrom
dev
Feb 12, 2026
Merged

wangxingjun778 merged 1 commit into
mainfrom
dev

Conversation

@wangxingjun778

Copy link
Copy Markdown
Member

No description provided.

@wangxingjun778
wangxingjun778 merged commit 9f70a14 into main Feb 12, 2026
1 check passed
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @wangxingjun778, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly expands and clarifies the Sirchmunk architecture documentation. It introduces a detailed breakdown of core components, elaborates on the multi-phase search pipeline, and provides an in-depth explanation of the Monte Carlo evidence sampling method. A new section on 'Self-Evolving Knowledge Clusters' describes how the system continuously learns and improves, complemented by a clear overview of data storage. These updates aim to provide a more comprehensive and accessible understanding of Sirchmunk's design principles and operational mechanisms.

Highlights

  • Core Components Section Added: A new 'Core Components' section was introduced, providing a table that describes key system components such as AgenticSearch, KnowledgeBase, and GrepRetriever, clarifying their roles within the Sirchmunk architecture.
  • Enhanced Knowledge Cluster Reuse Description: The description for 'Phase 0 — Knowledge Cluster Reuse' was expanded to emphasize the concept of knowledge compounding and highlight the achievement of sub-second response times for repeated queries.
  • Detailed Parallel Probing Explanations: More detailed explanations were provided for each step within 'Phase 1 — Parallel Probing,' specifically clarifying the functions of LLM Keyword Extraction and Directory Structure Scan.
  • Revised Monte Carlo Evidence Sampling: The 'Monte Carlo Evidence Sampling' section was completely revised, now detailing its three distinct phases (Cast the Net, Focus, Synthesize) and outlining key properties like document-agnosticism and token efficiency.
  • New Self-Evolving Knowledge Clusters Section: A comprehensive new section on 'Self-Evolving Knowledge Clusters' was added, defining KnowledgeClusters, illustrating their lifecycle with a diagram, and highlighting their properties for query-driven embeddings and semantic broadening.
  • Data Storage Overview: A new 'Data Storage' section was included, specifying the default work path and detailing the directory structure for cache, knowledge clusters, and settings.
  • Chinese Guide Title Update: The title of the Chinese architecture guide (architecture.zh.md) was updated from '架构' to '架构设计' for improved clarity.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Changelog
  • content/docs/guide/architecture.md
    • Added a "Core Components" section with a table.
    • Expanded the description of "Phase 0 — Knowledge Cluster Reuse" to include knowledge compounding.
    • Detailed the four steps of "Phase 1 — Parallel Probing."
    • Rewrote the "Monte Carlo Evidence Sampling" section, adding a new explanation, three phases, and key properties.
    • Introduced a new "Self-Evolving Knowledge Clusters" section, including its definition, lifecycle, and properties.
    • Added a "Data Storage" section with the directory structure.
  • content/docs/guide/architecture.zh.md
    • Updated the document title from "架构" to "架构设计".
    • Added a "核心组件" (Core Components) section with a table.
    • Expanded the description of "阶段 0 — 知识簇复用" (Phase 0 — Knowledge Cluster Reuse).
    • Detailed the four steps of "阶段 1 — 并行探测" (Phase 1 — Parallel Probing).
    • Rewrote the "蒙特卡洛证据采样" (Monte Carlo Evidence Sampling) section, adding a new explanation, three phases, and core characteristics.
    • Introduced a new "自进化知识簇" (Self-Evolving Knowledge Clusters) section, including its definition, lifecycle, and core characteristics.
    • Added a "数据存储" (Data Storage) section with the directory structure.
Activity
  • No human activity has been recorded for this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request significantly enhances the architecture documentation by adding a core components overview and detailing the self-evolving knowledge clusters and data storage. The new content is well-structured and informative.

My review focuses on two main points to improve clarity and correctness:

  1. Consistency in Terminology: In the 'Monte Carlo Evidence Sampling' section of both English and Chinese documents, the steps are named 'Phases' (or '阶段'), which conflicts with the main pipeline's phase naming. I've suggested using 'Acts' ('幕') for consistency with other documentation.
  2. Architectural Accuracy: The new 'Self-Evolving Knowledge Clusters' lifecycle diagram and description in both documents omit the crucial 'Phase 4' (Summarization/ReAct). I've highlighted this as a high-priority issue to ensure the documentation accurately reflects the system's workflow.

Addressing these points will make the documentation even more clear and precise.

Comment on lines +110 to +143
### Lifecycle: From Creation to Evolution

```text
┌─────── New Query ───────┐
│ ▼
│ ┌──────────────────────────────┐
│ │ Phase 0: Semantic Reuse │──── Match found ──→ Return cached cluster
│ │ (cosine similarity ≥ 0.85) │ + update hotness/queries/embedding
│ └──────────┬───────────────────┘
│ No match
│ ▼
│ ┌──────────────────────────────┐
│ │ Phase 1–3: Full Search │
│ │ (keywords → retrieval → │
│ │ Monte Carlo → LLM synth) │
│ └──────────┬───────────────────┘
│ ▼
│ ┌──────────────────────────────┐
│ │ Build New Cluster │
│ │ Deterministic ID: C{sha256} │
│ └──────────┬───────────────────┘
│ ▼
│ ┌──────────────────────────────┐
│ │ Phase 5: Persist │
│ │ Embed queries → DuckDB → │
│ │ Parquet (atomic sync) │
└─────└──────────────────────────────┘
```

1. **Reuse Check (Phase 0):** Before any retrieval, the query is embedded and compared against all stored clusters via cosine similarity. If a high-confidence match is found, the existing cluster is returned instantly — saving LLM tokens and search time entirely.

2. **Creation (Phase 1–3):** When no reuse match is found, the full pipeline runs: keyword extraction, file retrieval, Monte Carlo evidence sampling, and LLM synthesis produce a new `KnowledgeCluster`.

3. **Persistence (Phase 5):** The cluster is stored in an in-memory DuckDB table and periodically flushed to Parquet files. Atomic writes and mtime-based reload ensure multi-process safety.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The lifecycle diagram and the numbered list description for 'Self-Evolving Knowledge Clusters' are missing a crucial step: Phase 4 — Summarization or ReAct Refinement. The current flow incorrectly shows a direct path from cluster creation (Phases 1-3) to persistence (Phase 5). This omits the decision-making logic for summarization or invoking the ReAct agent, which is a key part of the architecture. Please update both the diagram and the description to accurately reflect the complete pipeline, including Phase 4.

Comment on lines +110 to +143
### 生命周期:从创建到进化

```text
┌─────── 新查询 ───────┐
│ ▼
│ ┌───────────────────────────────┐
│ │ 阶段 0:语义复用 │──── 匹配命中 ──→ 返回缓存知识簇
│ │ (余弦相似度 ≥ 0.85) │ + 更新热度/查询/嵌入
│ └──────────┬────────────────────┘
│ 未匹配
│ ▼
│ ┌───────────────────────────────┐
│ │ 阶段 1–3:完整搜索 │
│ │ (关键词 → 检索 → │
│ │ 蒙特卡洛采样 → LLM 合成) │
│ └──────────┬────────────────────┘
│ ▼
│ ┌───────────────────────────────┐
│ │ 构建新知识簇 │
│ │ 确定性 ID: C{sha256} │
│ └──────────┬────────────────────┘
│ ▼
│ ┌───────────────────────────────┐
│ │ 阶段 5:持久化 │
│ │ 嵌入查询 → DuckDB → │
│ │ Parquet(原子写入同步) │
└─────└───────────────────────────────┘
```

1. **复用检查(阶段 0):** 在任何检索开始之前,查询会被嵌入并通过余弦相似度与所有已存储知识簇进行比对。若发现高置信度匹配,系统直接返回已有知识簇——完全省去 LLM 推理和搜索开销。

2. **创建(阶段 1–3):** 当无复用匹配时,完整管线运行:关键词提取、文件检索、蒙特卡洛证据采样、LLM 合成,最终生成新的 `KnowledgeCluster`。

3. **持久化(阶段 5):** 知识簇存储在内存中的 DuckDB 表中,并定期刷写为 Parquet 文件。原子写入和基于文件修改时间的重载机制确保多进程安全。

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The lifecycle diagram and description for '自进化知识簇' (Self-Evolving Knowledge Clusters) are missing 阶段 4 — 摘要或 ReAct 精炼 (Phase 4 — Summarization or ReAct Refinement). The current documentation incorrectly shows a direct flow from creation (Phases 1-3) to persistence (Phase 5), which is a misleading representation of the architecture. Please update both the diagram and the description to include Phase 4 and show the complete process.

Comment on lines +70 to +76
The algorithm operates in three phases:

1. **Phase 1 — Cast the Net (Exploration):** Fuzzy anchor matching combined with stratified random sampling. The system identifies seed regions of potential relevance while maintaining broad coverage through randomized probing — ensuring no high-value region is missed.

2. **Phase 2 — Focus (Exploitation):** Gaussian importance sampling centered around high-scoring seeds from Phase 1. The sampling density concentrates on the most promising regions, extracting surrounding context and scoring each snippet for relevance.

3. **Phase 3 — Synthesize:** The top-K scored snippets are passed to the LLM, which synthesizes them into a coherent Region of Interest (ROI) summary with a confidence flag — enabling the pipeline to decide whether evidence is sufficient or a ReAct agent should be invoked for deeper exploration.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using 'Phase 1', 'Phase 2', and 'Phase 3' for the steps of the Monte Carlo algorithm can be confusing, as the main search pipeline is also described using 'Phases'. For clarity and to avoid ambiguity, consider using a different term for these steps. The technical-deep-dive/index.md document refers to these as 'Acts' ('Act 1', 'Act 2', 'Act 3'), which seems like a good alternative to maintain consistency across the documentation.

Suggested change
The algorithm operates in three phases:
1. **Phase 1 — Cast the Net (Exploration):** Fuzzy anchor matching combined with stratified random sampling. The system identifies seed regions of potential relevance while maintaining broad coverage through randomized probing — ensuring no high-value region is missed.
2. **Phase 2 — Focus (Exploitation):** Gaussian importance sampling centered around high-scoring seeds from Phase 1. The sampling density concentrates on the most promising regions, extracting surrounding context and scoring each snippet for relevance.
3. **Phase 3 — Synthesize:** The top-K scored snippets are passed to the LLM, which synthesizes them into a coherent Region of Interest (ROI) summary with a confidence flag — enabling the pipeline to decide whether evidence is sufficient or a ReAct agent should be invoked for deeper exploration.
The algorithm operates in three acts:
1. **Act 1 — Cast the Net (Exploration):** Fuzzy anchor matching combined with stratified random sampling. The system identifies seed regions of potential relevance while maintaining broad coverage through randomized probing — ensuring no high-value region is missed.
2. **Act 2 — Focus (Exploitation):** Gaussian importance sampling centered around high-scoring seeds from Act 1. The sampling density concentrates on the most promising regions, extracting surrounding context and scoring each snippet for relevance.
3. **Act 3 — Synthesize:** The top-K scored snippets are passed to the LLM, which synthesizes them into a coherent Region of Interest (ROI) summary with a confidence flag — enabling the pipeline to decide whether evidence is sufficient or a ReAct agent should be invoked for deeper exploration.

Comment on lines +70 to +76
该算法分为三个阶段:

1. **第一阶段 — 撒网(探索):** 模糊锚定匹配结合分层随机采样。系统在识别潜在相关种子区域的同时,通过随机探测保持广泛覆盖,确保不会遗漏高价值区域。

2. **第二阶段 — 聚焦(利用):** 以第一阶段高分种子为中心进行高斯重要性采样。采样密度集中在最有前景的区域,提取上下文并对每个片段评分。

3. **第三阶段 — 合成:** 将 Top-K 评分片段传递给 LLM,合成为连贯的兴趣区域(ROI)摘要,并附带置信度标志——使管线能够判断证据是否充分,或是否需要启用 ReAct 智能体进行更深层的自适应检索。

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Similar to the English version, using '阶段' for the steps of the Monte Carlo algorithm can be confusing, as the main search pipeline also uses '阶段' for its stages. To improve clarity, consider using a different term. The technical-deep-dive documentation suggests the concept of 'Acts'. Translating this as '幕' (e.g., '第一幕', '第二幕') would be more distinct and consistent with the overall documentation.

Suggested change
该算法分为三个阶段:
1. **第一阶段 — 撒网(探索):** 模糊锚定匹配结合分层随机采样。系统在识别潜在相关种子区域的同时,通过随机探测保持广泛覆盖,确保不会遗漏高价值区域。
2. **第二阶段 — 聚焦(利用):** 以第一阶段高分种子为中心进行高斯重要性采样。采样密度集中在最有前景的区域,提取上下文并对每个片段评分。
3. **第三阶段 — 合成:** 将 Top-K 评分片段传递给 LLM,合成为连贯的兴趣区域(ROI)摘要,并附带置信度标志——使管线能够判断证据是否充分,或是否需要启用 ReAct 智能体进行更深层的自适应检索。
该算法分为三个幕:
1. **第一幕 — 撒网(探索):** 模糊锚定匹配结合分层随机采样。系统在识别潜在相关种子区域的同时,通过随机探测保持广泛覆盖,确保不会遗漏高价值区域。
2. **第二幕 — 聚焦(利用):** 以第一幕高分种子为中心进行高斯重要性采样。采样密度集中在最有前景的区域,提取上下文并对每个片段评分。
3. **第三幕 — 合成:** 将 Top-K 评分片段传递给 LLM,合成为连贯的兴趣区域(ROI)摘要,并附带置信度标志——使管线能够判断证据是否充分,或是否需要启用 ReAct 智能体进行更深层的自适应检索。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant