Repository navigation
[Docs] Update guide - #2
Conversation
Summary of ChangesHello @wangxingjun778, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly expands and clarifies the Sirchmunk architecture documentation. It introduces a detailed breakdown of core components, elaborates on the multi-phase search pipeline, and provides an in-depth explanation of the Monte Carlo evidence sampling method. A new section on 'Self-Evolving Knowledge Clusters' describes how the system continuously learns and improves, complemented by a clear overview of data storage. These updates aim to provide a more comprehensive and accessible understanding of Sirchmunk's design principles and operational mechanisms. Highlights
🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request significantly enhances the architecture documentation by adding a core components overview and detailing the self-evolving knowledge clusters and data storage. The new content is well-structured and informative.
My review focuses on two main points to improve clarity and correctness:
- Consistency in Terminology: In the 'Monte Carlo Evidence Sampling' section of both English and Chinese documents, the steps are named 'Phases' (or '阶段'), which conflicts with the main pipeline's phase naming. I've suggested using 'Acts' ('幕') for consistency with other documentation.
- Architectural Accuracy: The new 'Self-Evolving Knowledge Clusters' lifecycle diagram and description in both documents omit the crucial 'Phase 4' (Summarization/ReAct). I've highlighted this as a high-priority issue to ensure the documentation accurately reflects the system's workflow.
Addressing these points will make the documentation even more clear and precise.
| ### Lifecycle: From Creation to Evolution | ||
|
|
||
| ```text | ||
| ┌─────── New Query ───────┐ | ||
| │ ▼ | ||
| │ ┌──────────────────────────────┐ | ||
| │ │ Phase 0: Semantic Reuse │──── Match found ──→ Return cached cluster | ||
| │ │ (cosine similarity ≥ 0.85) │ + update hotness/queries/embedding | ||
| │ └──────────┬───────────────────┘ | ||
| │ No match | ||
| │ ▼ | ||
| │ ┌──────────────────────────────┐ | ||
| │ │ Phase 1–3: Full Search │ | ||
| │ │ (keywords → retrieval → │ | ||
| │ │ Monte Carlo → LLM synth) │ | ||
| │ └──────────┬───────────────────┘ | ||
| │ ▼ | ||
| │ ┌──────────────────────────────┐ | ||
| │ │ Build New Cluster │ | ||
| │ │ Deterministic ID: C{sha256} │ | ||
| │ └──────────┬───────────────────┘ | ||
| │ ▼ | ||
| │ ┌──────────────────────────────┐ | ||
| │ │ Phase 5: Persist │ | ||
| │ │ Embed queries → DuckDB → │ | ||
| │ │ Parquet (atomic sync) │ | ||
| └─────└──────────────────────────────┘ | ||
| ``` | ||
|
|
||
| 1. **Reuse Check (Phase 0):** Before any retrieval, the query is embedded and compared against all stored clusters via cosine similarity. If a high-confidence match is found, the existing cluster is returned instantly — saving LLM tokens and search time entirely. | ||
|
|
||
| 2. **Creation (Phase 1–3):** When no reuse match is found, the full pipeline runs: keyword extraction, file retrieval, Monte Carlo evidence sampling, and LLM synthesis produce a new `KnowledgeCluster`. | ||
|
|
||
| 3. **Persistence (Phase 5):** The cluster is stored in an in-memory DuckDB table and periodically flushed to Parquet files. Atomic writes and mtime-based reload ensure multi-process safety. |
There was a problem hiding this comment.
The lifecycle diagram and the numbered list description for 'Self-Evolving Knowledge Clusters' are missing a crucial step: Phase 4 — Summarization or ReAct Refinement. The current flow incorrectly shows a direct path from cluster creation (Phases 1-3) to persistence (Phase 5). This omits the decision-making logic for summarization or invoking the ReAct agent, which is a key part of the architecture. Please update both the diagram and the description to accurately reflect the complete pipeline, including Phase 4.
| ### 生命周期:从创建到进化 | ||
|
|
||
| ```text | ||
| ┌─────── 新查询 ───────┐ | ||
| │ ▼ | ||
| │ ┌───────────────────────────────┐ | ||
| │ │ 阶段 0:语义复用 │──── 匹配命中 ──→ 返回缓存知识簇 | ||
| │ │ (余弦相似度 ≥ 0.85) │ + 更新热度/查询/嵌入 | ||
| │ └──────────┬────────────────────┘ | ||
| │ 未匹配 | ||
| │ ▼ | ||
| │ ┌───────────────────────────────┐ | ||
| │ │ 阶段 1–3:完整搜索 │ | ||
| │ │ (关键词 → 检索 → │ | ||
| │ │ 蒙特卡洛采样 → LLM 合成) │ | ||
| │ └──────────┬────────────────────┘ | ||
| │ ▼ | ||
| │ ┌───────────────────────────────┐ | ||
| │ │ 构建新知识簇 │ | ||
| │ │ 确定性 ID: C{sha256} │ | ||
| │ └──────────┬────────────────────┘ | ||
| │ ▼ | ||
| │ ┌───────────────────────────────┐ | ||
| │ │ 阶段 5:持久化 │ | ||
| │ │ 嵌入查询 → DuckDB → │ | ||
| │ │ Parquet(原子写入同步) │ | ||
| └─────└───────────────────────────────┘ | ||
| ``` | ||
|
|
||
| 1. **复用检查(阶段 0):** 在任何检索开始之前,查询会被嵌入并通过余弦相似度与所有已存储知识簇进行比对。若发现高置信度匹配,系统直接返回已有知识簇——完全省去 LLM 推理和搜索开销。 | ||
|
|
||
| 2. **创建(阶段 1–3):** 当无复用匹配时,完整管线运行:关键词提取、文件检索、蒙特卡洛证据采样、LLM 合成,最终生成新的 `KnowledgeCluster`。 | ||
|
|
||
| 3. **持久化(阶段 5):** 知识簇存储在内存中的 DuckDB 表中,并定期刷写为 Parquet 文件。原子写入和基于文件修改时间的重载机制确保多进程安全。 |
There was a problem hiding this comment.
The lifecycle diagram and description for '自进化知识簇' (Self-Evolving Knowledge Clusters) are missing 阶段 4 — 摘要或 ReAct 精炼 (Phase 4 — Summarization or ReAct Refinement). The current documentation incorrectly shows a direct flow from creation (Phases 1-3) to persistence (Phase 5), which is a misleading representation of the architecture. Please update both the diagram and the description to include Phase 4 and show the complete process.
| The algorithm operates in three phases: | ||
|
|
||
| 1. **Phase 1 — Cast the Net (Exploration):** Fuzzy anchor matching combined with stratified random sampling. The system identifies seed regions of potential relevance while maintaining broad coverage through randomized probing — ensuring no high-value region is missed. | ||
|
|
||
| 2. **Phase 2 — Focus (Exploitation):** Gaussian importance sampling centered around high-scoring seeds from Phase 1. The sampling density concentrates on the most promising regions, extracting surrounding context and scoring each snippet for relevance. | ||
|
|
||
| 3. **Phase 3 — Synthesize:** The top-K scored snippets are passed to the LLM, which synthesizes them into a coherent Region of Interest (ROI) summary with a confidence flag — enabling the pipeline to decide whether evidence is sufficient or a ReAct agent should be invoked for deeper exploration. |
There was a problem hiding this comment.
Using 'Phase 1', 'Phase 2', and 'Phase 3' for the steps of the Monte Carlo algorithm can be confusing, as the main search pipeline is also described using 'Phases'. For clarity and to avoid ambiguity, consider using a different term for these steps. The technical-deep-dive/index.md document refers to these as 'Acts' ('Act 1', 'Act 2', 'Act 3'), which seems like a good alternative to maintain consistency across the documentation.
| The algorithm operates in three phases: | |
| 1. **Phase 1 — Cast the Net (Exploration):** Fuzzy anchor matching combined with stratified random sampling. The system identifies seed regions of potential relevance while maintaining broad coverage through randomized probing — ensuring no high-value region is missed. | |
| 2. **Phase 2 — Focus (Exploitation):** Gaussian importance sampling centered around high-scoring seeds from Phase 1. The sampling density concentrates on the most promising regions, extracting surrounding context and scoring each snippet for relevance. | |
| 3. **Phase 3 — Synthesize:** The top-K scored snippets are passed to the LLM, which synthesizes them into a coherent Region of Interest (ROI) summary with a confidence flag — enabling the pipeline to decide whether evidence is sufficient or a ReAct agent should be invoked for deeper exploration. | |
| The algorithm operates in three acts: | |
| 1. **Act 1 — Cast the Net (Exploration):** Fuzzy anchor matching combined with stratified random sampling. The system identifies seed regions of potential relevance while maintaining broad coverage through randomized probing — ensuring no high-value region is missed. | |
| 2. **Act 2 — Focus (Exploitation):** Gaussian importance sampling centered around high-scoring seeds from Act 1. The sampling density concentrates on the most promising regions, extracting surrounding context and scoring each snippet for relevance. | |
| 3. **Act 3 — Synthesize:** The top-K scored snippets are passed to the LLM, which synthesizes them into a coherent Region of Interest (ROI) summary with a confidence flag — enabling the pipeline to decide whether evidence is sufficient or a ReAct agent should be invoked for deeper exploration. |
| 该算法分为三个阶段: | ||
|
|
||
| 1. **第一阶段 — 撒网(探索):** 模糊锚定匹配结合分层随机采样。系统在识别潜在相关种子区域的同时,通过随机探测保持广泛覆盖,确保不会遗漏高价值区域。 | ||
|
|
||
| 2. **第二阶段 — 聚焦(利用):** 以第一阶段高分种子为中心进行高斯重要性采样。采样密度集中在最有前景的区域,提取上下文并对每个片段评分。 | ||
|
|
||
| 3. **第三阶段 — 合成:** 将 Top-K 评分片段传递给 LLM,合成为连贯的兴趣区域(ROI)摘要,并附带置信度标志——使管线能够判断证据是否充分,或是否需要启用 ReAct 智能体进行更深层的自适应检索。 |
There was a problem hiding this comment.
Similar to the English version, using '阶段' for the steps of the Monte Carlo algorithm can be confusing, as the main search pipeline also uses '阶段' for its stages. To improve clarity, consider using a different term. The technical-deep-dive documentation suggests the concept of 'Acts'. Translating this as '幕' (e.g., '第一幕', '第二幕') would be more distinct and consistent with the overall documentation.
| 该算法分为三个阶段: | |
| 1. **第一阶段 — 撒网(探索):** 模糊锚定匹配结合分层随机采样。系统在识别潜在相关种子区域的同时,通过随机探测保持广泛覆盖,确保不会遗漏高价值区域。 | |
| 2. **第二阶段 — 聚焦(利用):** 以第一阶段高分种子为中心进行高斯重要性采样。采样密度集中在最有前景的区域,提取上下文并对每个片段评分。 | |
| 3. **第三阶段 — 合成:** 将 Top-K 评分片段传递给 LLM,合成为连贯的兴趣区域(ROI)摘要,并附带置信度标志——使管线能够判断证据是否充分,或是否需要启用 ReAct 智能体进行更深层的自适应检索。 | |
| 该算法分为三个幕: | |
| 1. **第一幕 — 撒网(探索):** 模糊锚定匹配结合分层随机采样。系统在识别潜在相关种子区域的同时,通过随机探测保持广泛覆盖,确保不会遗漏高价值区域。 | |
| 2. **第二幕 — 聚焦(利用):** 以第一幕高分种子为中心进行高斯重要性采样。采样密度集中在最有前景的区域,提取上下文并对每个片段评分。 | |
| 3. **第三幕 — 合成:** 将 Top-K 评分片段传递给 LLM,合成为连贯的兴趣区域(ROI)摘要,并附带置信度标志——使管线能够判断证据是否充分,或是否需要启用 ReAct 智能体进行更深层的自适应检索。 |
No description provided.