diff --git a/content/docs/guide/architecture.md b/content/docs/guide/architecture.md index 1bea21f..958dc37 100644 --- a/content/docs/guide/architecture.md +++ b/content/docs/guide/architecture.md @@ -9,22 +9,35 @@ Sirchmunk's architecture is organized into cleanly separated layers, following t ![Sirchmunk Architecture](Sirchmunk_Architecture.png "Sirchmunk high-level architecture diagram") +## Core Components + +| Component | Description | +|:----------------------|:-------------------------------------------------------------------------| +| **AgenticSearch** | Search orchestrator with LLM-enhanced retrieval capabilities | +| **KnowledgeBase** | Transforms raw results into structured knowledge clusters with evidences | +| **EvidenceProcessor** | Evidence processing based on the Monte Carlo Importance Sampling | +| **GrepRetriever** | High-performance _indexless_ file search with parallel processing | +| **OpenAIChat** | Unified LLM interface supporting streaming and usage tracking | +| **MonitorTracker** | Real-time system and application metrics collection | + ## Multi-Phase Search Pipeline At the heart of Sirchmunk is a multi-phase search pipeline designed around **maximum parallelism within each phase** and **strict phase dependencies** between them. ### Phase 0 — Knowledge Cluster Reuse -Before any computation begins, the system checks whether a semantically similar query has been answered before. A lightweight embedding of the query is compared via cosine similarity against stored knowledge clusters. If a close match is found, the cached cluster is returned in sub-second time. +Before any computation begins, the system checks whether a semantically similar query has been answered before. A lightweight embedding of the query is compared via cosine similarity against stored knowledge clusters. If a close match is found (above a configurable threshold), the cached cluster is returned immediately — providing **sub-second response times** for repeated or paraphrased queries. + +This is not merely a cache — it is the beginning of knowledge *compounding*. Each reuse appends the new query to the cluster's history, so the system remembers which questions led to which insights. ### Phase 1 — Parallel Probing Four independent probes launch concurrently to gather diverse signals: -1. **LLM Keyword Extraction** — Multi-granularity keyword decomposition -2. **Directory Structure Scan** — File metadata collection and structural analysis -3. **Knowledge Cache Lookup** — Partial match search across existing clusters -4. **Spec-Path Context Load** — Previously computed context for known paths +1. **LLM Keyword Extraction** — The LLM decomposes the query into multi-level keywords, from coarse (high recall) to fine (high precision), each annotated with an estimated rarity score. This multi-granularity approach ensures that both broad topics and specific terms are captured. +2. **Directory Structure Scan** — The file system is traversed to collect path metadata: file names, sizes, modification times, and content previews. This is the foundation for *intelligent inference* — using structural cues (naming conventions, directory hierarchy, file types) to narrow down the most promising candidates before ever reading their content. +3. **Knowledge Cache Lookup** — Partial match search across existing clusters for potential reuse of previously acquired knowledge. +4. **Spec-Path Context Load** — Previously computed context for known paths is loaded from cache. ### Phase 2 — Retrieval & Ranking @@ -50,13 +63,23 @@ Valuable clusters are saved with their embeddings for future reuse. ### Monte Carlo Evidence Sampling -An exploration–exploitation approach for extracting evidence from large documents: +Traditional retrieval systems read entire documents or rely on fixed-size chunks, leading to either wasted tokens or lost context. Sirchmunk takes a fundamentally different approach inspired by **Monte Carlo methods** — treating evidence extraction as a **sampling problem** rather than a parsing problem. ![Monte Carlo Evidence Sampling Algorithm](Sirchmunk_MonteCarloSamplingAlgo.png "Monte Carlo Evidence Sampling: Three-act exploration–exploitation strategy") -1. **Act 1 — Casting the Net**: Fuzzy anchoring + stratified random sampling -2. **Act 2 — Zooming In**: Gaussian importance sampling around high-scoring seeds -3. **Act 3 — Synthesis**: Top-K snippets → LLM summary +The algorithm operates in three phases: + +1. **Phase 1 — Cast the Net (Exploration):** Fuzzy anchor matching combined with stratified random sampling. The system identifies seed regions of potential relevance while maintaining broad coverage through randomized probing — ensuring no high-value region is missed. + +2. **Phase 2 — Focus (Exploitation):** Gaussian importance sampling centered around high-scoring seeds from Phase 1. The sampling density concentrates on the most promising regions, extracting surrounding context and scoring each snippet for relevance. + +3. **Phase 3 — Synthesize:** The top-K scored snippets are passed to the LLM, which synthesizes them into a coherent Region of Interest (ROI) summary with a confidence flag — enabling the pipeline to decide whether evidence is sufficient or a ReAct agent should be invoked for deeper exploration. + +**Key properties:** + +- **Document-agnostic:** The same algorithm works equally well on a 2-page memo and a 500-page technical manual — no document-specific chunking heuristics needed. +- **Token-efficient:** Only the most relevant regions are sent to the LLM, dramatically reducing token consumption compared to full-document approaches. +- **Exploration-exploitation balance:** Random exploration prevents tunnel vision, while importance sampling ensures depth where it matters most. ### ReAct Agent @@ -66,6 +89,87 @@ An autonomous Think → Act → Observe loop with: - Dual-budget mechanism (token budget + loop count) - Memory of previously explored avenues +## Self-Evolving Knowledge Clusters + +Sirchmunk does not discard search results after answering a query. Instead, every search produces a **KnowledgeCluster** — a structured, reusable knowledge unit that grows smarter over time. This is what makes the system _self-evolving_. + +### What is a KnowledgeCluster? + +A KnowledgeCluster is a richly annotated object that captures the full cognitive output of a single search cycle: + +| Field | Purpose | +|:------|:--------| +| **Evidences** | Source-linked snippets extracted via Monte Carlo sampling, each with file path, summary, and raw text | +| **Content** | LLM-synthesized markdown with structured analysis and references | +| **Patterns** | 3–5 distilled design principles or mechanisms identified from the evidence | +| **Confidence** | A consensus score \[0, 1\] indicating the reliability of the cluster | +| **Queries** | Historical queries that contributed to or reused this cluster (FIFO, max 5) | +| **Hotness** | Activity score reflecting query frequency and recency | +| **Embedding** | 384-dim vector derived from accumulated queries, enabling semantic retrieval | + +### Lifecycle: From Creation to Evolution + +```text + ┌─────── New Query ───────┐ + │ ▼ + │ ┌──────────────────────────────┐ + │ │ Phase 0: Semantic Reuse │──── Match found ──→ Return cached cluster + │ │ (cosine similarity ≥ 0.85) │ + update hotness/queries/embedding + │ └──────────┬───────────────────┘ + │ No match + │ ▼ + │ ┌──────────────────────────────┐ + │ │ Phase 1–3: Full Search │ + │ │ (keywords → retrieval → │ + │ │ Monte Carlo → LLM synth) │ + │ └──────────┬───────────────────┘ + │ ▼ + │ ┌──────────────────────────────┐ + │ │ Build New Cluster │ + │ │ Deterministic ID: C{sha256} │ + │ └──────────┬───────────────────┘ + │ ▼ + │ ┌──────────────────────────────┐ + │ │ Phase 5: Persist │ + │ │ Embed queries → DuckDB → │ + │ │ Parquet (atomic sync) │ + └─────└──────────────────────────────┘ +``` + +1. **Reuse Check (Phase 0):** Before any retrieval, the query is embedded and compared against all stored clusters via cosine similarity. If a high-confidence match is found, the existing cluster is returned instantly — saving LLM tokens and search time entirely. + +2. **Creation (Phase 1–3):** When no reuse match is found, the full pipeline runs: keyword extraction, file retrieval, Monte Carlo evidence sampling, and LLM synthesis produce a new `KnowledgeCluster`. + +3. **Persistence (Phase 5):** The cluster is stored in an in-memory DuckDB table and periodically flushed to Parquet files. Atomic writes and mtime-based reload ensure multi-process safety. + +4. **Evolution on Reuse:** Each time a cluster is reused, the system: + - Appends the new query to the cluster's query history (FIFO, max 5) + - Increases hotness (`+0.1`, capped at 1.0) + - Recomputes the embedding from the updated query set — broadening the cluster's semantic catchment area + - Updates version and timestamp + +### Key Properties + +- **Zero-cost acceleration:** Repeated or semantically similar queries are answered from cached clusters without any LLM inference, making subsequent searches near-instantaneous. +- **Query-driven embeddings:** Cluster embeddings are derived from _queries_ rather than content, ensuring that retrieval aligns with how users actually ask questions — not how documents are written. +- **Semantic broadening:** As diverse queries reuse the same cluster, its embedding drifts to cover a wider semantic neighborhood, naturally improving recall for related future queries. +- **Lightweight persistence:** DuckDB in-memory + Parquet on disk — no external database infrastructure required. Background daemon sync with configurable flush intervals keeps overhead minimal. + +## Data Storage + +All persistent data is stored in the configured `SIRCHMUNK_WORK_PATH` (default: `~/.sirchmunk/`): + +```text +{SIRCHMUNK_WORK_PATH}/ + ├── .cache/ + ├── history/ # Chat session history (DuckDB) + │ └── chat_history.db + ├── knowledge/ # Knowledge clusters (Parquet) + │ └── knowledge_clusters.parquet + └── settings/ # User settings (DuckDB) + └── settings.db +``` + ## Design Principles Sirchmunk adheres to **SOLID principles**: diff --git a/content/docs/guide/architecture.zh.md b/content/docs/guide/architecture.zh.md index 8d17b55..0098d52 100644 --- a/content/docs/guide/architecture.zh.md +++ b/content/docs/guide/architecture.zh.md @@ -1,5 +1,5 @@ --- -title: 架构 +title: 架构设计 weight: 0 --- @@ -9,22 +9,35 @@ Sirchmunk 采用清晰分离的层次化架构,遵循**关注点分离**原则 ![Sirchmunk 架构](Sirchmunk_Architecture.png "Sirchmunk 高层架构图") +## 核心组件 + +| 组件 | 说明 | +|:------------------------|:-----------------------------------------------------------------------| +| **AgenticSearch** | 搜索编排器,具备 LLM 增强检索能力 | +| **KnowledgeBase** | 将原始结果转化为结构化知识簇并附带证据 | +| **EvidenceProcessor** | 基于蒙特卡洛重要性采样的证据处理 | +| **GrepRetriever** | 高性能 _无索引_ 文件检索,支持并行处理 | +| **OpenAIChat** | 统一 LLM 接口,支持流式与用量统计 | +| **MonitorTracker** | 实时系统与应用指标采集 | + ## 多阶段搜索管线 Sirchmunk 的核心是一个多阶段搜索管线,其设计原则是**每个阶段内最大并行化**与阶段之间**严格的依赖关系**。 ### 阶段 0 — 知识簇复用 -在任何计算开始之前,系统会检查是否有语义相似的查询已被回答过。将查询的轻量级嵌入通过余弦相似度与存储的知识簇进行比较。如果找到近似匹配,缓存的知识簇会在亚秒级时间内返回。 +在任何计算开始之前,系统会检查是否有语义相似的查询已被回答过。将查询的轻量级嵌入通过余弦相似度与存储的知识簇进行比较。如果找到近似匹配(超过可配置的相似度阈值),缓存的知识簇会立即返回——为重复或改述的查询提供**亚秒级响应时间**。 + +这不仅仅是缓存——它是知识*复利积累*的起点。每次复用都会将新查询追加到知识簇的历史中,让系统记住哪些问题导向了哪些洞见。 ### 阶段 1 — 并行探测 四个独立的探测并发启动,收集多样化信号: -1. **LLM 关键词提取** — 多粒度关键词分解 -2. **目录结构扫描** — 文件元数据收集和结构分析 -3. **知识缓存查找** — 在现有知识簇中进行部分匹配搜索 -4. **指定路径上下文加载** — 已知路径的预计算上下文 +1. **LLM 关键词提取** — LLM 将查询分解为多层级关键词,从粗粒度(高召回)到细粒度(高精度),每个关键词附带稀有度评分。这种多粒度方法确保同时捕获广泛主题和特定术语。 +2. **目录结构扫描** — 遍历文件系统收集路径元数据:文件名、大小、修改时间和内容预览。这是*智能推断*的基础——利用结构线索(命名规范、目录层级、文件类型)在读取内容之前就缩小最有可能的候选范围。 +3. **知识缓存查找** — 在现有知识簇中进行部分匹配搜索,复用先前获取的知识。 +4. **指定路径上下文加载** — 从缓存中加载已知路径的预计算上下文。 ### 阶段 2 — 检索与排序 @@ -50,13 +63,23 @@ Sirchmunk 的核心是一个多阶段搜索管线,其设计原则是**每个 ### 蒙特卡洛证据采样 -一种用于从大文档中提取证据的探索-利用方法: +传统检索系统要么读取完整文档,要么依赖固定大小的分块,导致 Token 浪费或上下文丢失。Sirchmunk 借鉴**蒙特卡洛方法**,采用了截然不同的策略——将证据提取视为一个**采样问题**而非解析问题。 ![蒙特卡洛证据采样算法](Sirchmunk_MonteCarloSamplingAlgo.png "蒙特卡洛证据采样:三阶段启发式探索-利用策略") -1. **第一阶段 — 撒网**:模糊锚定 + 分层随机采样 -2. **第二阶段 — 聚焦**:围绕高分种子的高斯重要性采样 -3. **第三阶段 — 合成**:Top-K 片段 → LLM 摘要 +该算法分为三个阶段: + +1. **第一阶段 — 撒网(探索):** 模糊锚定匹配结合分层随机采样。系统在识别潜在相关种子区域的同时,通过随机探测保持广泛覆盖,确保不会遗漏高价值区域。 + +2. **第二阶段 — 聚焦(利用):** 以第一阶段高分种子为中心进行高斯重要性采样。采样密度集中在最有前景的区域,提取上下文并对每个片段评分。 + +3. **第三阶段 — 合成:** 将 Top-K 评分片段传递给 LLM,合成为连贯的兴趣区域(ROI)摘要,并附带置信度标志——使管线能够判断证据是否充分,或是否需要启用 ReAct 智能体进行更深层的自适应检索。 + +**核心特性:** + +- **文档无关性:** 同一算法在 2 页备忘录和 500 页技术手册上同样有效,无需针对特定文档的分块启发式规则。 +- **Token 高效:** 仅将最相关的区域发送给 LLM,相比全文档方案大幅降低 Token 消耗。 +- **探索-利用平衡:** 随机探索防止视野盲区,重要性采样确保在关键区域深入挖掘。 ### ReAct 智能体 @@ -66,6 +89,87 @@ Sirchmunk 的核心是一个多阶段搜索管线,其设计原则是**每个 - 双预算机制(Token 预算 + 循环次数) - 记忆已探索的途径 +## 自进化知识簇(Knowledge Cluster) + +Sirchmunk 不会在回答完查询后丢弃搜索结果。相反,每次搜索都会产生一个 **KnowledgeCluster(知识簇)**——一个结构化、可复用的知识单元,随着使用不断变得更加智能。这正是系统具备_自进化_能力的核心机制。 + +### 什么是 KnowledgeCluster? + +KnowledgeCluster 是一个丰富标注的对象,完整记录了单次搜索周期的认知产出: + +| 字段 | 用途 | +|:-----|:-----| +| **Evidences(证据)** | 通过蒙特卡洛采样提取的源文件片段,包含文件路径、摘要和原始文本 | +| **Content(内容)** | LLM 合成的结构化 Markdown 分析,附带引用 | +| **Patterns(模式)** | 从证据中提炼的 3–5 条设计原则或核心机制 | +| **Confidence(置信度)** | 共识评分 \[0, 1\],指示知识簇的可靠性 | +| **Queries(查询历史)** | 贡献或复用该知识簇的历史查询(FIFO,最多 5 条) | +| **Hotness(热度)** | 反映查询频率和时效性的活跃度评分 | +| **Embedding(嵌入向量)** | 由累积查询生成的 384 维向量,用于语义检索 | + +### 生命周期:从创建到进化 + +```text + ┌─────── 新查询 ───────┐ + │ ▼ + │ ┌───────────────────────────────┐ + │ │ 阶段 0:语义复用 │──── 匹配命中 ──→ 返回缓存知识簇 + │ │ (余弦相似度 ≥ 0.85) │ + 更新热度/查询/嵌入 + │ └──────────┬────────────────────┘ + │ 未匹配 + │ ▼ + │ ┌───────────────────────────────┐ + │ │ 阶段 1–3:完整搜索 │ + │ │ (关键词 → 检索 → │ + │ │ 蒙特卡洛采样 → LLM 合成) │ + │ └──────────┬────────────────────┘ + │ ▼ + │ ┌───────────────────────────────┐ + │ │ 构建新知识簇 │ + │ │ 确定性 ID: C{sha256} │ + │ └──────────┬────────────────────┘ + │ ▼ + │ ┌───────────────────────────────┐ + │ │ 阶段 5:持久化 │ + │ │ 嵌入查询 → DuckDB → │ + │ │ Parquet(原子写入同步) │ + └─────└───────────────────────────────┘ +``` + +1. **复用检查(阶段 0):** 在任何检索开始之前,查询会被嵌入并通过余弦相似度与所有已存储知识簇进行比对。若发现高置信度匹配,系统直接返回已有知识簇——完全省去 LLM 推理和搜索开销。 + +2. **创建(阶段 1–3):** 当无复用匹配时,完整管线运行:关键词提取、文件检索、蒙特卡洛证据采样、LLM 合成,最终生成新的 `KnowledgeCluster`。 + +3. **持久化(阶段 5):** 知识簇存储在内存中的 DuckDB 表中,并定期刷写为 Parquet 文件。原子写入和基于文件修改时间的重载机制确保多进程安全。 + +4. **复用时进化:** 每当知识簇被复用时,系统会: + - 将新查询追加到知识簇的查询历史中(FIFO,最多 5 条) + - 提升热度(+0.1,上限 1.0) + - 基于更新后的查询集重新计算嵌入——扩展知识簇的语义覆盖范围 + - 更新版本号和时间戳 + +### 核心特性 + +- **零成本加速:** 重复或语义相似的查询直接从缓存知识簇获取答案,无需任何 LLM 推理,后续搜索几乎瞬时完成。 +- **查询驱动的嵌入:** 知识簇嵌入基于_查询_而非内容生成,确保检索与用户的实际提问方式对齐——而非文档的书写方式。 +- **语义拓展:** 随着多样化查询复用同一知识簇,其嵌入会漂移以覆盖更广的语义邻域,自然提升相关未来查询的召回率。 +- **轻量级持久化:** DuckDB 内存存储 + Parquet 磁盘持久化——无需外部数据库基础设施。后台守护线程同步,可配置刷写间隔,开销极小。 + +## 数据存储 + +所有持久化数据存储在配置的 `SIRCHMUNK_WORK_PATH`(默认:`~/.sirchmunk/`): + +```text +{SIRCHMUNK_WORK_PATH}/ + ├── .cache/ + ├── history/ # 聊天会话历史(DuckDB) + │ └── chat_history.db + ├── knowledge/ # 知识簇(Parquet) + │ └── knowledge_clusters.parquet + └── settings/ # 用户设置(DuckDB) + └── settings.db +``` + ## 设计原则 Sirchmunk 遵循 **SOLID 原则**: