feat(knowledge): breadcrumb分块上下文 & LookupKnowledgeTool日志优化
- DocumentChunk新增breadcrumb字段,分块时构建完整标题层级路径 - DocumentChunkService splitByHeadings维护标题层级栈算法 - VectorIndexService 将breadcrumb写入Milvus metadata - LookupKnowledgeTool日志替换为结构化摘要,替代原始MD预览 - L0返回策略:唯一匹配用正文摘要,多匹配+L1有结果仅元数据(不读文件) - 新增buildCompactSummary / buildMetadataOnlySummary方法 - 安装frontend-design skill - 创建mvp/文档目录(架构设计+会话存储方案) - 更新测试适配新逻辑
This commit is contained in:
@@ -0,0 +1,57 @@
|
||||
# Frontend Design — Complete Guidance
|
||||
|
||||
This document provides a comprehensive framework for creating visually distinctive, non-templated UI designs. Here's the full breakdown:
|
||||
|
||||
## Foundational Approach
|
||||
|
||||
Act as the design lead for a studio known for unique client identities — the client has already turned down template-like proposals. Every choice about palette, typography, and layout must be specific to the brief, including "one real aesthetic risk you can justify."
|
||||
|
||||
## Grounding in Subject Matter
|
||||
|
||||
If the brief is vague about the product or subject, pin it down yourself: name the subject, its audience, and the page's single job. Draw inspiration from "the subject's own world, its materials, instruments, artifacts, and vernacular." Use any known context about the human's preferences or past designs as hints.
|
||||
|
||||
## Design Principles
|
||||
|
||||
- **Hero as thesis**: Open with "the most characteristic thing in the subject's world" — avoid default choices like a big number with a small label and gradient accent unless truly optimal.
|
||||
- **Typography**: Pair display and body faces deliberately, not from your usual repertoire. Set a clear type scale with intentional weights, widths, and spacing. "Make the type treatment itself a memorable part of the design."
|
||||
- **Structure as information**: Numbering, eyebrows, dividers must encode something true about the content. Question whether numbered markers (01/02/03) actually make sense before using them — only appropriate for real sequences.
|
||||
- **Motion**: Consider where animation serves the subject. "An orchestrated moment usually lands harder than scattered effects." Sometimes less is better to avoid an AI-generated feel.
|
||||
- **Complexity**: Match execution to the vision — maximalist needs elaborate execution, minimal needs precision.
|
||||
- **Content**: Come up with copy if the brief lacks it. Poor copy makes a design feel as templated as poor layout.
|
||||
|
||||
## AI-Generated Design Traps
|
||||
|
||||
Three common AI-default looks to watch for: (1) warm cream background (~#F4F1EA) with serif display and terracotta accent; (2) near-black with bright acid-green or vermilion; (3) broadsheet layout with hairline rules, zero border-radius, and dense columns. "All three are legitimate for some briefs, but they are defaults rather than choices." Where the brief leaves an axis free, don't spend that freedom on a default.
|
||||
|
||||
## Two-Pass Process
|
||||
|
||||
**Pass 1 — Plan**: Create a compact token system:
|
||||
|
||||
1. **Color**: 4–6 named hex values
|
||||
2. **Type**: Characterful display face (used with restraint), complementary body face, utility face for captions/data
|
||||
3. **Layout**: One-sentence prose descriptions + ASCII wireframes
|
||||
4. **Signature**: The single unique element the page will be remembered by
|
||||
|
||||
Review the plan against the brief. If any part reads like what you'd produce for any similar page, revise it. Only then write code.
|
||||
|
||||
**Pass 2 — Build**: Follow the revised plan exactly. Watch for CSS selector specificity conflicts (e.g., `.section` and `.cta` fighting over padding/margins). Do most planning internally, only sharing ideas when confident.
|
||||
|
||||
## Restraint & Self-Critique
|
||||
|
||||
"Spend your boldness in one place" — let the signature element be the one memorable thing; keep everything else quiet. "Not taking a risk can be a risk itself!" Build responsively down to mobile, with visible keyboard focus and reduced motion respected. Critique as you build. Follow Chanel's advice: before finishing, remove one accessory. Jot notes about what you've tried to avoid repeating yourself.
|
||||
|
||||
## Writing in Design
|
||||
|
||||
Words exist to make the design understandable and usable — they're "design material, not decoration." Write from the end user's perspective, naming things by what people control and recognize, never by how the system is built.
|
||||
|
||||
- Use active voice as default
|
||||
- A control should say exactly what happens: "Save changes," not "Submit"
|
||||
- Maintain consistent vocabulary throughout flows (button says "Publish," toast says "Published")
|
||||
- Treat errors as guidance, not mood — explain what went wrong and how to fix it
|
||||
- Empty screens are invitations to act
|
||||
- Keep the register conversational: "plain verbs, sentence case, no filler"
|
||||
- Let each element do exactly one job — "a label labels, an example demonstrates"
|
||||
|
||||
## License
|
||||
|
||||
Apache License 2.0 — see LICENSE.txt
|
||||
@@ -229,5 +229,5 @@ category: api
|
||||
|
||||
## 八、参考文档
|
||||
|
||||
- [知识库检索架构说明](../mvp/architecture/knowledge-retrieval-architecture.md)
|
||||
- [知识库检索架构说明](./mvp/architecture/knowledge-retrieval-architecture.md)
|
||||
- [AI Ops 核心设计 Essence 报告](../docs/learning/01-AI-Ops-核心设计-Essence报告.md)
|
||||
|
||||
@@ -0,0 +1,421 @@
|
||||
# 知识库检索架构(L0 + L1)
|
||||
|
||||
**更新日期**: 2026-06-25
|
||||
|
||||
---
|
||||
|
||||
## 一、概述
|
||||
|
||||
`LookupKnowledgeTool` 实现两阶段混合检索:
|
||||
|
||||
- **L0 精确匹配**:基于内存索引的关键词匹配(< 10ms),索引从数据库加载
|
||||
- **L1 语义检索**:基于 Milvus 向量数据库的相似度搜索(200-500ms)
|
||||
|
||||
---
|
||||
|
||||
## 二、完整流程
|
||||
|
||||
```
|
||||
用户查询
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────┐
|
||||
│ L0: 关键词精确匹配 │ (< 10ms)
|
||||
│ • 从内存索引做关键词匹配 │
|
||||
│ • 索引来源: ApiDocument DB│
|
||||
└──────────┬──────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────┴──────┐
|
||||
│ matches=1 │ ← 唯一匹配(高置信度)
|
||||
└──────┬──────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────┐ ┌──────────────────┐
|
||||
│ 跳过 L1 │ │ L0 返回正文摘要 │
|
||||
│ 置信度: high │ │ buildCompactSummary│
|
||||
└──────────────┘ └──────────────────┘
|
||||
|
||||
|
||||
┌──────┴──────┐
|
||||
│ matches=0 │ ← 无匹配
|
||||
└──────┬──────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────┐ ┌──────────────────┐
|
||||
│ 触发 L1 │ │ L0 无结果 │
|
||||
│ L1 语义检索 │ │ 仅有 L1 补充结果 │
|
||||
└──────────────┘ └──────────────────┘
|
||||
|
||||
|
||||
┌──────┴──────┐
|
||||
│ matches>=2 │ ← 多匹配
|
||||
└──────┬──────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────┐
|
||||
│ 触发 L1 │
|
||||
│ L1 语义检索 │
|
||||
└──────┬──────┘
|
||||
│
|
||||
┌─────┴─────┐
|
||||
│ ║ │
|
||||
▼ ▼
|
||||
L1 有结果 L1 无结果
|
||||
│ │
|
||||
▼ ▼
|
||||
元数据摘要 正文摘要
|
||||
(不读文件) (读文件)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 三、L0 返回内容策略
|
||||
|
||||
根据匹配场景决定 L0 返回给 LLM 的上下文内容量。
|
||||
|
||||
### 3.1 唯一匹配(高置信度,matches=1)
|
||||
|
||||
**策略**: `buildCompactSummary()`
|
||||
|
||||
L1 被跳过,LLM 只有 L0 信息来源,需要提供足够的正文内容。
|
||||
|
||||
```
|
||||
文档: 支付网关错误码定义
|
||||
摘要: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
章节:
|
||||
- 超时类错误
|
||||
- 业务类错误
|
||||
- 签名类错误
|
||||
---
|
||||
**含义**:支付网关请求超时
|
||||
**常见原因**:网络延迟、第三方服务响应慢
|
||||
...
|
||||
```
|
||||
|
||||
| 组成部分 | 说明 | 大小 |
|
||||
|---------|------|------|
|
||||
| title + summary | 从内存索引获取 | ~50-100 字符 |
|
||||
| 章节标题列表 | 从文件解析 `##` 标题 | ~50-200 字符 |
|
||||
| 正文片段 | 去 frontmatter/标题行/空行,短文档 800/长文档 500 字符截断 | ~300-800 字符 |
|
||||
| **总计** | | **~400-1000 字符** |
|
||||
|
||||
### 3.2 多匹配 + L1 有结果
|
||||
|
||||
**策略**: `buildMetadataOnlySummary()`
|
||||
|
||||
L1 已有语义内容片段,L0 仅需告知 LLM 命中了哪些文档。**不读文件**,仅用内存索引。
|
||||
|
||||
```
|
||||
文档: 支付网关错误码定义
|
||||
摘要: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
关键词: ERR_TIMEOUT, 超时, 支付网关
|
||||
来源: api/payment-errors.md
|
||||
```
|
||||
|
||||
| 组成部分 | 说明 | 大小 |
|
||||
|---------|------|------|
|
||||
| title + summary + keywords | 全部从内存索引获取 | ~100-200 字符 |
|
||||
| **总计** | | **~100-200 字符** |
|
||||
|
||||
### 3.3 多匹配 + L1 无结果
|
||||
|
||||
**策略**: `buildCompactSummary()`(同 3.1)
|
||||
|
||||
L1 未返回结果, L0 作为兜底提供正文内容。
|
||||
|
||||
---
|
||||
|
||||
## 四、决策矩阵
|
||||
|
||||
```
|
||||
needFullContent = highConfidence || !hasL1
|
||||
```
|
||||
|
||||
| 场景 | matches | L1 结果 | needFullContent | L0 策略 | 是否读文件 | 上下文大小 |
|
||||
|------|:-------:|:--------:|:---------------:|---------|:---------:|:--------:|
|
||||
| 唯一匹配 | 1 | 未执行 | true | `buildCompactSummary` | 是 | ~600 字符 |
|
||||
| 多匹配 + L1 有结果 | 2+ | 有 | false | `buildMetadataOnlySummary` | **否** | ~150 字符 |
|
||||
| 多匹配 + L1 无结果 | 2+ | 无 | true | `buildCompactSummary` | 是 | ~600 字符 |
|
||||
| 无匹配 | 0 | 有 | — | 无 L0,仅 L1 | 否 | 0 |
|
||||
|
||||
---
|
||||
|
||||
## 五、代码结构
|
||||
|
||||
```
|
||||
LookupKnowledgeTool
|
||||
├── lookupKnowledge(query) # 入口:编排 L0 + L1
|
||||
├── buildResult(l0, l1, confidence) # 组装结果,选择摘要策略
|
||||
├── buildCompactSummary(entry) # 元数据 + 章节 + 正文片段(读文件)
|
||||
├── buildMetadataOnlySummary(entry) # 仅元数据(不读文件)
|
||||
├── countMdHeadings(content) # 统计章节数(日志用)
|
||||
└── extractFirstMeaningfulLine(...) # 提取首个有意义文本行(日志用)
|
||||
```
|
||||
|
||||
### 关键逻辑(buildResult)
|
||||
|
||||
```java
|
||||
boolean needFullContent = highConfidence || !hasL1;
|
||||
String content = needFullContent
|
||||
? buildCompactSummary(first)
|
||||
: buildMetadataOnlySummary(first);
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 六、日志输出示例
|
||||
|
||||
### 多匹配场景(matches=2, L1 有结果)
|
||||
|
||||
```
|
||||
[L0 精确匹配] 完成: matches=2, time=3ms
|
||||
[置信度判断] highConfidence=false, reason=多个或零个匹配
|
||||
[L1 语义检索] L0非唯一匹配,触发L1语义检索...
|
||||
[L1 语义检索] 完成: matches=1, time=245ms
|
||||
----------------------------------------
|
||||
<<< [工具返回] lookup_knowledge
|
||||
<<< [L0 主结果] 标题: 支付网关错误码定义
|
||||
<<< [L0 主结果] 摘要: 记录了支付网关所有核心错误码的含义及排查方向 ← 仅元数据
|
||||
<<< [L0 主结果] 内容: 126 字符, 0 个章节 ← 约150字符
|
||||
<<< [L1 补充] 相似度: 0.8234
|
||||
<<< [L1 补充] 内容片段: 支付网关请求超时... ← L1 提供具体内容
|
||||
```
|
||||
|
||||
### 唯一匹配场景(matches=1, 跳过 L1)
|
||||
|
||||
```
|
||||
[L0 精确匹配] 完成: matches=1, time=2ms
|
||||
[置信度判断] highConfidence=true, reason=唯一匹配
|
||||
[L1 语义检索] L0唯一匹配,跳过L1检索
|
||||
----------------------------------------
|
||||
<<< [工具返回] lookup_knowledge
|
||||
<<< [L0 主结果] 标题: 支付网关错误码定义
|
||||
<<< [L0 主结果] 摘要: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
<<< [L0 主结果] 内容: 725 字符, 3 个章节 ← 约700字符
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 七、MVP 效率评估 & 改进方向
|
||||
|
||||
### 7.1 当前效率评估
|
||||
|
||||
| 维度 | 评分 | 说明 |
|
||||
|------|:----:|------|
|
||||
| L0 匹配速度 | ★★★★★ | 内存索引,< 10ms,几乎没有优化空间 |
|
||||
| L1 检索速度 | ★★★★☆ | Milvus 向量检索,200-500ms,取决于数据量 |
|
||||
| L0 匹配准确率 | ★★☆☆☆ | 子串匹配,无排序无评分,匹配即返回 |
|
||||
| L1 检索准确率 | ★★★☆☆ | 语义相似度,但分块缺少上下文信息 |
|
||||
| 召回率(查全) | ★★★☆☆ | L0+L1 两阶段覆盖大多数场景,但缺乏融合重排 |
|
||||
| 上下文利用率 | ★★★★☆ | 根据场景动态控制 L0 内容量,已优化 |
|
||||
| **综合** | **★★★☆☆** | **MVP 可用,但检索质量有提升空间** |
|
||||
|
||||
### 7.2 关键瓶颈
|
||||
|
||||
#### 瓶颈 1:分块丢失上下文(✅ 已修复—见下方 7.5)
|
||||
|
||||
当前每个 Chunk 只记录最近的 `##` 标题:
|
||||
|
||||
```json
|
||||
{
|
||||
"content": "**含义**:支付网关请求超时\n**常见原因**:网络延迟",
|
||||
"title": "超时类错误",
|
||||
"chunkIndex": 2
|
||||
}
|
||||
```
|
||||
|
||||
LLM 收到这个片段时**不知道**它属于"支付网关错误码定义"这个文档,也不知道具体错误码名称是 ERR_TIMEOUT。如果同时检索了多个文档的片段,LLM 容易混淆。
|
||||
|
||||
#### 瓶颈 2:L0 关键词匹配过于简单
|
||||
|
||||
当前 `KnowledgeIndexService.matchesKeywords()` 只做子串包含匹配,没有:
|
||||
- 排序/评分(多个匹配时按什么顺序?)
|
||||
- 权重(标题匹配 > 正文匹配)
|
||||
- 部分匹配("timeout" 匹配 "ERR_TIMEOUT")
|
||||
|
||||
#### 瓶颈 3:L0 和 L1 无交叉融合
|
||||
|
||||
两阶段检索结果只是简单的"1位L0 + 1位L1"拼接,没有:
|
||||
- RRF 或加权融合重排
|
||||
- 重复内容去重
|
||||
- 根据相关性选择 top-K
|
||||
|
||||
---
|
||||
|
||||
### 7.3 改进方向分析
|
||||
|
||||
#### 方向 A:面包屑导航(Chunk 携带层级上下文)
|
||||
|
||||
**做法**:分块时记录完整的标题层级路径作为 `breadcrumb`。
|
||||
|
||||
当前分块 metadata:
|
||||
```json
|
||||
{ "title": "超时类错误" }
|
||||
```
|
||||
|
||||
改进后:
|
||||
```json
|
||||
{
|
||||
"title": "超时类错误",
|
||||
"breadcrumb": "支付网关错误码定义 > 超时类错误 > ERR_TIMEOUT",
|
||||
"heading_h1": "支付网关错误码定义",
|
||||
"heading_h2": "超时类错误",
|
||||
"heading_h3": "ERR_TIMEOUT"
|
||||
}
|
||||
```
|
||||
|
||||
**收益评估**:
|
||||
|
||||
| 场景 | 无面包屑的问题 | 有面包屑的改善 | 提升幅度 |
|
||||
|------|---------------|---------------|:--------:|
|
||||
| 单文档多分块 | LLM 知道标题但不知道层级关系 | 清楚"文档>章节>条目"归属 | 中等 |
|
||||
| 跨文档混合结果 | 分块看不出源文档 | breadcrumb 第一段就是文档标题 | 大 |
|
||||
| 深层嵌套文档(3+ 级) | 分块内容难以定位 | 完整路径一目了然 | 显著 |
|
||||
| 向量检索相关性 | 只对 chunk content 做 embedding | breadcrumb 可拼入 content 做 embedding 或单独索引 | 中等 |
|
||||
|
||||
**MVP 阶段价值**:当前文档结构较浅(2-3级),breadcrumb 对 LLM 理解帮助中等。但如果后续文档层级加深(像你提到的"排障指南 > 支付网关 > 502错误处理"),价值会显著提升。
|
||||
|
||||
**实现成本**:低。修改 `DocumentChunkService` 的分块逻辑,积累当前标题栈,写入 `DocumentChunk` 和 Milvus metadata。
|
||||
|
||||
#### 方向 B:混合检索 + RRF 重排
|
||||
|
||||
**做法**:L0 关键词和 L1 向量检索并行执行 → 结果用 Reciprocal Rank Fusion 统一排序 → 取 top-K。
|
||||
|
||||
```
|
||||
用户查询 → 并行的:
|
||||
├── L0 关键词匹配 → 得分向量 S₀
|
||||
└── L1 向量检索 → 得分向量 S₁
|
||||
↓
|
||||
RRF 融合重排
|
||||
↓
|
||||
top-K 统一结果
|
||||
```
|
||||
|
||||
RRF 公式:对每个文档 d,`score(d) = Σ 1/(k + rank_r(d))`,其中 k=60(常数)。
|
||||
|
||||
**收益评估**:
|
||||
|
||||
| 场景 | 当前的问题 | 混合 + RRF | 提升幅度 |
|
||||
|------|-----------|-----------|:--------:|
|
||||
| 精确关键词("ERR_TIMEOUT") | L0 匹配但不排序,L1 可能不匹配 | L0 高排名 → RRF 拉到顶部 | 大 |
|
||||
| 语义查询("支付超时如何处理") | L0 可能不匹配,全靠 L1 | L1 兜底不受影响 | 无变化 |
|
||||
| 混合查询("ERR_TIMEOUT 支付网关超时") | L0 匹配一个、L1 匹配一个,无融合 | RRF 统一排序,更合理 | 中等 |
|
||||
| 多文档匹配 | L0 返回无序列表 + L1 独立结果 | 统一排序、去重 | 大 |
|
||||
|
||||
**MVP 阶段价值**:RRF 的实现成本和维护成本较高,而当前 MVP 数据量小(6 个文档),人工检查即可确定哪些匹配是好的。**建议数据量 > 50 个文档时引入**。
|
||||
|
||||
#### 方向 C:Breadcrumb + Embedding 增强
|
||||
|
||||
**做法**:将 breadcrumb 拼入 chunk content 后再做 embedding,让向量包含层级语义。
|
||||
|
||||
```java
|
||||
// 当前
|
||||
embeddingService.generateEmbedding(chunk.getContent())
|
||||
|
||||
// 改进
|
||||
String augmentedContent = chunk.getBreadcrumb() + "\n" + chunk.getContent();
|
||||
embeddingService.generateEmbedding(augmentedContent);
|
||||
```
|
||||
|
||||
这样搜索"ERR_TIMEOUT"时,"支付网关错误码定义 > 超时类错误 > ERR_TIMEOUT" 也会匹配到,而不只是 chunk 正文。
|
||||
|
||||
| 场景 | 当前 | Breadcrumb + Embedding | 提升 |
|
||||
|------|------|------------------------|:----:|
|
||||
| 搜索"支付网关超时" | 匹配到正文含"超时"和"支付网关"的 chunk | breadcrumb 直接含"支付网关",匹配更准 | 中等 |
|
||||
| 搜索"错误码定义" | 可能匹配不到具体错误内容的 chunk | breadcrumb 含"错误码定义",相关性更高 | 大 |
|
||||
|
||||
---
|
||||
|
||||
### 7.4 实施优先级建议
|
||||
|
||||
| 优先级 | 改进项 | 复杂度 | 收益 | 状态 |
|
||||
|:------:|--------|:------:|:----:|:----:|
|
||||
| P0 | **Breadcrumb 上下文**(方向 A) | 低 | 中 | **✅ 已实现 (2026-06-26)** |
|
||||
| P1 | 下个版本 | 低 | 中-大 | 待定 |
|
||||
| P1 | L0 排序(匹配评分 + 排序) | 低 | 中 | 待定 |
|
||||
| P2 | 混合检索 + RRF 重排 | 高 | 大 | 数据量 > 50 文档时引入 |
|
||||
|
||||
### 7.5 Breadcrumb 实现说明
|
||||
|
||||
已于 2026-06-26 实现。改动范围:
|
||||
|
||||
| 文件 | 改动 |
|
||||
|------|------|
|
||||
| `DocumentChunk.java` | 新增 `breadcrumb` 字段 |
|
||||
| `DocumentChunkService.java` | `splitByHeadings()` 维护标题层级栈,`Section` 新增 `level`/`breadcrumb`,`chunkSection()` 和 `saveChunkAndGetNextStart()` 透传 Breadcrumb |
|
||||
| `VectorIndexService.java` | `buildMetadata()` 和 `buildDocumentMetadata()` 将 breadcrumb 写入 Milvus metadata |
|
||||
|
||||
#### 层级栈算法
|
||||
|
||||
```java
|
||||
// 在 splitByHeadings() 中,每次匹配到标题时:
|
||||
while (!headingStack.isEmpty() && headingStack.size() >= level) {
|
||||
headingStack.remove(headingStack.size() - 1); // 弹出同级或更高级
|
||||
}
|
||||
headingStack.add(title); // 追加当前标题
|
||||
currentBreadcrumb = String.join(" > ", headingStack);
|
||||
```
|
||||
|
||||
示例:处理 `fault-diagnosis-process.md` 的完整面包屑路径──
|
||||
|
||||
```json
|
||||
// 分块 "应急响应流程 > 1. 初步评估"
|
||||
{ "breadcrumb": "故障诊断流程规范 > 应急响应流程 > 1. 初步评估" }
|
||||
|
||||
// 分块 "根因分析方法 > 5-Why 分析法"
|
||||
{ "breadcrumb": "故障诊断流程规范 > 根因分析方法 > 5-Why 分析法" }
|
||||
```
|
||||
|
||||
#### 当前 metadata 结构(Milvus)
|
||||
|
||||
```json
|
||||
{
|
||||
"_source": "knowledge_base/api/payment-errors.md",
|
||||
"_file_name": "payment-errors.md",
|
||||
"category": "api",
|
||||
"chunkIndex": 2,
|
||||
"totalChunks": 5,
|
||||
"title": "超时类错误",
|
||||
"breadcrumb": "支付网关错误码定义 > 超时类错误 > ERR_TIMEOUT"
|
||||
}
|
||||
```
|
||||
|
||||
以 `fault-diagnosis-process.md` 为例:
|
||||
|
||||
```markdown
|
||||
# 故障诊断流程规范 ← heading_h1
|
||||
|
||||
## 应急响应流程 ← heading_h2
|
||||
|
||||
### 1. 初步评估 ← heading_h3(分块1)
|
||||
内容...
|
||||
|
||||
### 2. 快速止血 ← heading_h3(分块2)
|
||||
内容...
|
||||
|
||||
## 根因分析方法 ← heading_h2
|
||||
|
||||
### 5-Why 分析法 ← heading_h3(分块3)
|
||||
内容...
|
||||
```
|
||||
|
||||
改造后每个分块的 metadata:
|
||||
|
||||
```
|
||||
分块1: breadcrumb = "故障诊断流程规范 > 应急响应流程 > 1. 初步评估"
|
||||
分块2: breadcrumb = "故障诊断流程规范 > 应急响应流程 > 2. 快速止血"
|
||||
分块3: breadcrumb = "故障诊断流程规范 > 根因分析方法 > 5-Why 分析法"
|
||||
```
|
||||
|
||||
LLM 视角受益:当检索到 "2. 快速止血" 时,LLM 立刻知道它属于"故障诊断流程规范 > 应急响应流程"体系,不需要额外读取其他分块来推断上下文。
|
||||
|
||||
| 文件 | 说明 |
|
||||
|------|------|
|
||||
| `LookupKnowledgeTool.java` | 检索工具入口 |
|
||||
| `KnowledgeIndexService.java` | L0 内存索引管理 |
|
||||
| `VectorSearchService.java` | L1 向量检索(Milvus) |
|
||||
| `KnowledgeEntry.java` | 索引条目 DTO(含 title, summary, keywords) |
|
||||
| `LookupResult.java` | 查询结果 DTO |
|
||||
| `PrimaryResult.java` | L0 结果 DTO |
|
||||
| `SupplementResult.java` | L1 结果 DTO |
|
||||
@@ -0,0 +1,266 @@
|
||||
# 会话存储方案设计
|
||||
|
||||
**日期**: 2026-06-26
|
||||
**类型**: 架构设计
|
||||
**状态**: 待评审
|
||||
|
||||
---
|
||||
|
||||
## 一、背景与目标
|
||||
|
||||
### 1.1 现状问题
|
||||
|
||||
当前仅有一张 `diagnosis_record` 表,存在以下问题:
|
||||
|
||||
| 问题 | 说明 |
|
||||
|------|------|
|
||||
| **语义耦合** | `fault_category`、`error_code`、`root_cause`、`solution` 等字段耦合在"告警分析"领域语义,ChatService 通用问答场景用不上 |
|
||||
| **Agent 维度缺失** | 只有一个 `tool_calls` JSON 字段,存不下两个 Agent 的多轮决策链 |
|
||||
| **检索质量不可追溯** | 没有记录 L0/L1 命中层、截断信息、召回内容长度 |
|
||||
| **指标不完整** | 有 `duration` 和 `confidence`,但缺 token 用量、自评信号、采纳率 |
|
||||
|
||||
### 1.2 存储范围
|
||||
|
||||
需要覆盖四个层面的数据:
|
||||
|
||||
```
|
||||
诊断级元数据
|
||||
├── 单次诊断的唯一 ID、查询问题、状态
|
||||
├── 会话级决策链
|
||||
│ ├── agent_step:每个 Agent 的每一步(输入、输出、延迟、Token)
|
||||
│ └── tool_invocation:每次工具调用(参数、结果、耗时)
|
||||
├── 检索质量明细
|
||||
│ └── 每次 lookup_knowledge 的命中层(L0/L1)、内容长度、是否截断
|
||||
└── 自评估信号
|
||||
└── LLM 对结论的置信度自评
|
||||
```
|
||||
|
||||
### 1.3 设计目标
|
||||
|
||||
- **可观测**:Debug 时能回溯完整决策链
|
||||
- **可评估**:能统计 L0/L1 命中率、平均 Token 消耗、工具采纳率等指标
|
||||
- **可演进**:覆盖当前两个 Agent(ChatService / AiOpsService),未来新增 Agent 也能接入
|
||||
|
||||
---
|
||||
|
||||
## 二、存储选型分析
|
||||
|
||||
### 2.1 方案对比
|
||||
|
||||
| 维度 | SQL + JSON 列 | NoSQL 文档库 |
|
||||
|------|:------------:|:-----------:|
|
||||
| 基础设施 | 已有的 MySQL,零新增 | 需新部署 MongoDB 等 |
|
||||
| 层级查询 | `WHERE session_id=? AND agent_name=?` 高效 | 需二级索引 |
|
||||
| 指标聚合 | `AVG(token_count) GROUP BY agent_name` 原生支持 | 聚合管道,学习成本 |
|
||||
| 非结构化内容 | JSON 列(MySQL 8+ 支持良好) | 天然支持 |
|
||||
| MVP 迭代速度 | JPA Entity + Flyway 快速迭代 | 新 ORM 学习成本 |
|
||||
|
||||
### 2.2 结论
|
||||
|
||||
**采用 MySQL + JSON 列**。结构化字段做查询和聚合,JSON 列存非结构化载荷。MVP 阶段数据量可控,等后续 > 百万级或需要更灵活 schema 时再评估 NoSQL。
|
||||
|
||||
---
|
||||
|
||||
## 三、存储模型
|
||||
|
||||
### 3.1 整体关系
|
||||
|
||||
```
|
||||
diagnosis_session (1)
|
||||
│
|
||||
└── agent_step (0:N) —— 单次诊断的每一步 Agent 决策
|
||||
│
|
||||
└── tool_invocation (0:N) —— 每步中的工具调用
|
||||
```
|
||||
|
||||
### 3.2 表设计
|
||||
|
||||
#### 表 1:diagnosis_session(诊断会话)
|
||||
|
||||
```sql
|
||||
CREATE TABLE diagnosis_session (
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
session_id VARCHAR(64) UNIQUE NOT NULL COMMENT '会话唯一 ID',
|
||||
|
||||
-- 请求
|
||||
query TEXT NOT NULL COMMENT '用户原始问题',
|
||||
status VARCHAR(16) DEFAULT 'PENDING' COMMENT 'PENDING / RUNNING / SUCCESS / FAILED',
|
||||
agent_flow VARCHAR(32) COMMENT 'CHAT / AI_OPS',
|
||||
|
||||
-- 汇总指标
|
||||
total_duration_ms INT COMMENT '总耗时(毫秒)',
|
||||
total_token_count INT COMMENT '总 Token 消耗',
|
||||
step_count INT COMMENT 'Agent 步数',
|
||||
tool_call_count INT COMMENT '工具调用次数',
|
||||
|
||||
-- 自评估信号(模型对结论的置信度自评)
|
||||
self_evaluation JSON COMMENT '{"confidence": 0-100, "reasoning": "...", "evidence_count": 3}',
|
||||
|
||||
-- 用户反馈
|
||||
feedback VARCHAR(16) COMMENT 'useful / not_useful / null',
|
||||
|
||||
-- 元数据
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
|
||||
|
||||
INDEX idx_created_at (created_at),
|
||||
INDEX idx_status (status),
|
||||
INDEX idx_agent_flow (agent_flow)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='诊断会话表';
|
||||
```
|
||||
|
||||
#### 表 2:agent_step(Agent 决策步骤)
|
||||
|
||||
```sql
|
||||
CREATE TABLE agent_step (
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
session_id VARCHAR(64) NOT NULL COMMENT '关联 diagnosis_session',
|
||||
|
||||
step_index INT NOT NULL COMMENT '当前 Agent 的第几步(从0开始)',
|
||||
agent_name VARCHAR(32) NOT NULL COMMENT 'intelligent_assistant / planner / executor / supervisor',
|
||||
|
||||
-- 模型调用(输入输出摘要,非完整消息体)
|
||||
model_input JSON COMMENT '模型输入摘要 [{role, content_truncated}, ...]',
|
||||
model_output JSON COMMENT '模型输出摘要 {text, tool_calls, ...}',
|
||||
thought TEXT COMMENT 'Agent 思考过程文本',
|
||||
has_tool_call BOOLEAN DEFAULT FALSE COMMENT '本轮是否调用了工具',
|
||||
|
||||
-- 性能指标
|
||||
duration_ms INT COMMENT '本轮耗时',
|
||||
token_count INT COMMENT '本轮 Token 消耗',
|
||||
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
|
||||
INDEX idx_session_step (session_id, step_index),
|
||||
INDEX idx_agent_name (agent_name)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Agent 决策步骤表';
|
||||
```
|
||||
|
||||
#### 表 3:tool_invocation(工具调用明细)
|
||||
|
||||
```sql
|
||||
CREATE TABLE tool_invocation (
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
session_id VARCHAR(64) NOT NULL COMMENT '关联 diagnosis_session',
|
||||
step_id BIGINT COMMENT '关联 agent_step.id(可为空,不强制外键)',
|
||||
|
||||
tool_name VARCHAR(64) NOT NULL COMMENT 'lookup_knowledge / queryPrometheusAlerts / 等',
|
||||
|
||||
-- 调用信息
|
||||
input_params JSON NOT NULL COMMENT '工具入参',
|
||||
output_preview TEXT COMMENT '输出前500字符(可观测用,不存完整输出)',
|
||||
output_length INT COMMENT '输出总字符数',
|
||||
|
||||
-- 检索质量(仅 lookup_knowledge 时有意义)
|
||||
retrieval_layer VARCHAR(8) COMMENT 'L0 / L1 / L0+L1',
|
||||
l0_match_count INT COMMENT 'L0 匹配数',
|
||||
l1_match_count INT COMMENT 'L1 匹配数',
|
||||
is_truncated BOOLEAN DEFAULT FALSE COMMENT '返回内容是否被截断',
|
||||
retrieval_details JSON COMMENT '{"l0_titles":[], "l1_scores":[], "has_supplement": true}',
|
||||
|
||||
-- 性能 & 状态
|
||||
duration_ms INT COMMENT '工具执行耗时',
|
||||
success BOOLEAN DEFAULT TRUE COMMENT '是否成功',
|
||||
error_message TEXT COMMENT '失败原因',
|
||||
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
|
||||
INDEX idx_session_id (session_id),
|
||||
INDEX idx_tool_name (tool_name),
|
||||
INDEX idx_retrieval_layer (retrieval_layer)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='工具调用明细表';
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 四、数据流设计
|
||||
|
||||
### 4.1 完整链路
|
||||
|
||||
```
|
||||
用户请求
|
||||
│
|
||||
▼
|
||||
1. 创建 diagnosis_session(status=RUNNING)
|
||||
│
|
||||
▼
|
||||
2. Agent Loop(可能多轮)
|
||||
│
|
||||
├── beforeModel()
|
||||
│ └── AgentLoggingHook 记录 model_input + 开始时间 → 写入 agent_step(先创建,duration 待填)
|
||||
│
|
||||
├── afterModel()
|
||||
│ └── AgentLoggingHook 记录 model_output + token_count + 工具调用决策 → 更新 agent_step
|
||||
│
|
||||
├── 工具执行(如 lookup_knowledge)
|
||||
│ └── LookupKnowledgeTool 记录 tool_invocation(L0/L1 明细、耗时、是否截断)
|
||||
│
|
||||
└── 循环直到模型不再调用工具
|
||||
│
|
||||
▼
|
||||
3. 诊断完成 → 更新 diagnosis_session
|
||||
├── status = SUCCESS / FAILED
|
||||
├── 汇总指标:total_duration_ms / total_token_count / step_count / tool_call_count
|
||||
└── self_evaluation(可选,由 LLM 自评)
|
||||
```
|
||||
|
||||
### 4.2 变更点
|
||||
|
||||
| 模块 | 当前行为 | 改造后 |
|
||||
|------|---------|--------|
|
||||
| `AgentLoggingHook` | 只打日志到 stdout | 同时写入 `agent_step` 表 |
|
||||
| `LookupKnowledgeTool` | 只打日志到 stdout | 同时写入 `tool_invocation` 表 |
|
||||
| `ChatService` / `AiOpsService` | 执行前后无持久化 | 创建 + 更新 `diagnosis_session` |
|
||||
|
||||
---
|
||||
|
||||
## 五、可观测能力
|
||||
|
||||
### 5.1 查询场景
|
||||
|
||||
| 需求 | SQL | 说明 |
|
||||
|------|-----|------|
|
||||
| 某次诊断用了哪些工具 | `SELECT * FROM tool_invocation WHERE session_id=?` | 按 session 关联 |
|
||||
| lookup_knowledge 的 L0/L1 命中率 | `SELECT retrieval_layer, COUNT(*) FROM tool_invocation WHERE tool_name='lookup_knowledge' GROUP BY retrieval_layer` | 聚合检索层分布 |
|
||||
| 某个 Agent 的平均思考耗时 | `SELECT AVG(duration_ms) FROM agent_step WHERE agent_name=?` | 按 Agent 分组 |
|
||||
| 某次诊断的完整决策链 | `SELECT * FROM agent_step WHERE session_id=? ORDER BY step_index` | 按步骤号排序 |
|
||||
| 被截断的检索占比 | `SELECT COUNT(*) FROM tool_invocation WHERE is_truncated=true AND tool_name='lookup_knowledge'` | 条件计数 |
|
||||
| 高置信度但用户反馈 negative | `SELECT * FROM diagnosis_session WHERE JSON_EXTRACT(self_evaluation, '$.confidence') > 80 AND feedback='not_useful'` | JSON 条件查询 |
|
||||
|
||||
### 5.2 评估指标
|
||||
|
||||
| 指标 | 计算方式 | 数据来源 |
|
||||
|------|---------|---------|
|
||||
| 平均诊断耗时 | `AVG(total_duration_ms)` | diagnosis_session |
|
||||
| 平均 Token 消耗 | `AVG(total_token_count)` | diagnosis_session |
|
||||
| 工具采纳率 | `tools_accepted / tools_proposed` | self_evaluation |
|
||||
| L0 命中率 | `l0_match_count > 0 的比例` | tool_invocation |
|
||||
| 截断率 | `is_truncated=true 的比例` | tool_invocation |
|
||||
| 用户满意度 | `feedback='useful' 的比例` | diagnosis_session |
|
||||
|
||||
---
|
||||
|
||||
## 六、与现有表的关系
|
||||
|
||||
### 6.1 diagnosis_session vs 现有 diagnosis_record
|
||||
|
||||
- **`diagnosis_record`** 保持不动,继续用于"告警分析"场景的领域字段(root_cause、solution 等)
|
||||
- **`diagnosis_session`** 是通用会话存储,覆盖 ChatService 和 AiOpsService
|
||||
- 两者通过 `session_id` 可关联
|
||||
|
||||
### 6.2 迁移策略
|
||||
|
||||
| 阶段 | 动作 |
|
||||
|:----:|------|
|
||||
| MVP | 新建三张表,新代码写入新表 |
|
||||
| V1.1 | 评估是否将 diagnosis_record 合并回 diagnosis_session(加 fault 相关字段到 JSON) |
|
||||
| V1.2 | 数据量 > 10 万时评估是否需要归档或迁移 |
|
||||
|
||||
---
|
||||
|
||||
## 七、未完成事项
|
||||
|
||||
- [ ] AI Ops Supervisor 的 Agent 执行步骤如何对应 agent_step 表(Supervisor 内嵌的子 Agent 步骤归到同一个 session 还是独立)
|
||||
- [ ] self_evaluation 的 confidence 自评通过什么方式获取(单独的 LLM 调用还是在 prompt 中要求输出)
|
||||
- [ ] feedback 字段和前端的交互方式
|
||||
- [ ] Tool_invocation 的 output_preview 截断策略(当前建议 500 字符)
|
||||
@@ -1,409 +1,421 @@
|
||||
# 知识库检索架构说明
|
||||
# 知识库检索架构(L0 + L1)
|
||||
|
||||
## 一、架构位置
|
||||
|
||||
知识库检索是 Agent 工具层的一部分,为所有 Agent 提供知识查询能力。
|
||||
|
||||
```
|
||||
Agent 层
|
||||
├── Supervisor Agent
|
||||
├── Planner Agent
|
||||
├── SubAgents (ExternalApi, InternalError, Database...)
|
||||
└── Verifier Agent
|
||||
↓ 调用
|
||||
工具层 (Tools)
|
||||
├── searchDoc (文档检索 - L1 向量检索)
|
||||
├── lookup_knowledge (混合检索 - L0+L1) ← 新增
|
||||
├── queryLogs (日志查询)
|
||||
├── queryTrace (链路追踪)
|
||||
└── queryOrder (订单查询)
|
||||
↓ 依赖
|
||||
服务层 (Services)
|
||||
├── VectorSearchService (L1 语义检索 - Milvus)
|
||||
├── KnowledgeIndexService (L0 精确匹配 - 内存) ← 新增
|
||||
├── FrontmatterParser (元数据解析) ← 新增
|
||||
└── DocumentManagementService (文档管理)
|
||||
↓ 持久化
|
||||
数据层
|
||||
├── MySQL (api_document + metadata 字段) ← 增强
|
||||
├── Milvus (向量索引)
|
||||
└── Local Files (knowledge_base/) ← 新增
|
||||
```
|
||||
**更新日期**: 2026-06-25
|
||||
|
||||
---
|
||||
|
||||
## 二、L0+L1 混合检索架构
|
||||
## 一、概述
|
||||
|
||||
### 2.1 检索流程
|
||||
`LookupKnowledgeTool` 实现两阶段混合检索:
|
||||
|
||||
- **L0 精确匹配**:基于内存索引的关键词匹配(< 10ms),索引从数据库加载
|
||||
- **L1 语义检索**:基于 Milvus 向量数据库的相似度搜索(200-500ms)
|
||||
|
||||
---
|
||||
|
||||
## 二、完整流程
|
||||
|
||||
```
|
||||
Agent 调用 lookup_knowledge(query)
|
||||
↓
|
||||
┌─────────────────────────────────────────┐
|
||||
│ LookupKnowledgeTool │
|
||||
│ (工具入口) │
|
||||
└────────────┬────────────────────────────┘
|
||||
用户查询
|
||||
│
|
||||
↓
|
||||
┌────────────────┐
|
||||
│ Step 1: L0 精确匹配 │ < 10ms
|
||||
│ (内存索引) │
|
||||
└────────┬───────────┘
|
||||
▼
|
||||
┌─────────────────────────────┐
|
||||
│ L0: 关键词精确匹配 │ (< 10ms)
|
||||
│ • 从内存索引做关键词匹配 │
|
||||
│ • 索引来源: ApiDocument DB│
|
||||
└──────────┬──────────────────┘
|
||||
│
|
||||
┌───────┴────────┐
|
||||
│ │
|
||||
唯一匹配 多个/零个匹配
|
||||
│ │
|
||||
↓ ↓
|
||||
高置信度 低置信度
|
||||
(不调用L1) (调用L1补充)
|
||||
│ │
|
||||
│ ┌──────────────────┐
|
||||
│ │ Step 2: L1 语义检索 │ 200-500ms
|
||||
│ │ (Milvus) │
|
||||
│ └──────────┬─────────┘
|
||||
│ │
|
||||
└────────┬───────────┘
|
||||
↓
|
||||
┌─────────────────────┐
|
||||
│ Step 3: 组装结果 │
|
||||
│ primary + supplement │
|
||||
└─────────────────────┘
|
||||
↓
|
||||
返回给 Agent
|
||||
```
|
||||
▼
|
||||
┌──────┴──────┐
|
||||
│ matches=1 │ ← 唯一匹配(高置信度)
|
||||
└──────┬──────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────┐ ┌──────────────────┐
|
||||
│ 跳过 L1 │ │ L0 返回正文摘要 │
|
||||
│ 置信度: high │ │ buildCompactSummary│
|
||||
└──────────────┘ └──────────────────┘
|
||||
|
||||
### 2.2 数据流
|
||||
|
||||
```
|
||||
文档上传流程:
|
||||
POST /api/documents/upload
|
||||
↓
|
||||
DocumentManagementService.uploadDocument()
|
||||
↓
|
||||
1. 文本提取
|
||||
2. 保存原始文件 → knowledge_base/{category}/{filename}
|
||||
3. 解析 frontmatter (FrontmatterParser)
|
||||
4. 分块 → 向量化 → Milvus 索引 (L1)
|
||||
5. 元数据存 MySQL (metadata 字段 JSON)
|
||||
6. 更新 L0 内存索引 (KnowledgeIndexService)
|
||||
↓
|
||||
完成
|
||||
┌──────┴──────┐
|
||||
│ matches=0 │ ← 无匹配
|
||||
└──────┬──────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────┐ ┌──────────────────┐
|
||||
│ 触发 L1 │ │ L0 无结果 │
|
||||
│ L1 语义检索 │ │ 仅有 L1 补充结果 │
|
||||
└──────────────┘ └──────────────────┘
|
||||
|
||||
文档查询流程:
|
||||
Agent 调用 lookup_knowledge("ERR_TIMEOUT")
|
||||
↓
|
||||
KnowledgeIndexService.exactMatch()
|
||||
↓
|
||||
遍历内存索引 (keywords 精确匹配)
|
||||
↓
|
||||
找到唯一匹配 → 读取本地文件 (前 2000 字符)
|
||||
↓
|
||||
返回 primary (高置信度)
|
||||
|
||||
┌──────┴──────┐
|
||||
│ matches>=2 │ ← 多匹配
|
||||
└──────┬──────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────┐
|
||||
│ 触发 L1 │
|
||||
│ L1 语义检索 │
|
||||
└──────┬──────┘
|
||||
│
|
||||
┌─────┴─────┐
|
||||
│ ║ │
|
||||
▼ ▼
|
||||
L1 有结果 L1 无结果
|
||||
│ │
|
||||
▼ ▼
|
||||
元数据摘要 正文摘要
|
||||
(不读文件) (读文件)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 三、核心组件说明
|
||||
## 三、L0 返回内容策略
|
||||
|
||||
### 3.1 FrontmatterParser
|
||||
根据匹配场景决定 L0 返回给 LLM 的上下文内容量。
|
||||
|
||||
**职责**:解析 Markdown 文件头的 YAML frontmatter
|
||||
### 3.1 唯一匹配(高置信度,matches=1)
|
||||
|
||||
**输入**:
|
||||
```markdown
|
||||
**策略**: `buildCompactSummary()`
|
||||
|
||||
L1 被跳过,LLM 只有 L0 信息来源,需要提供足够的正文内容。
|
||||
|
||||
```
|
||||
文档: 支付网关错误码定义
|
||||
摘要: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
章节:
|
||||
- 超时类错误
|
||||
- 业务类错误
|
||||
- 签名类错误
|
||||
---
|
||||
title: 支付网关错误码定义
|
||||
keywords: [ERR_TIMEOUT, 超时, 支付网关]
|
||||
summary: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
category: api
|
||||
**含义**:支付网关请求超时
|
||||
**常见原因**:网络延迟、第三方服务响应慢
|
||||
...
|
||||
```
|
||||
|
||||
| 组成部分 | 说明 | 大小 |
|
||||
|---------|------|------|
|
||||
| title + summary | 从内存索引获取 | ~50-100 字符 |
|
||||
| 章节标题列表 | 从文件解析 `##` 标题 | ~50-200 字符 |
|
||||
| 正文片段 | 去 frontmatter/标题行/空行,短文档 800/长文档 500 字符截断 | ~300-800 字符 |
|
||||
| **总计** | | **~400-1000 字符** |
|
||||
|
||||
### 3.2 多匹配 + L1 有结果
|
||||
|
||||
**策略**: `buildMetadataOnlySummary()`
|
||||
|
||||
L1 已有语义内容片段,L0 仅需告知 LLM 命中了哪些文档。**不读文件**,仅用内存索引。
|
||||
|
||||
```
|
||||
文档: 支付网关错误码定义
|
||||
摘要: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
关键词: ERR_TIMEOUT, 超时, 支付网关
|
||||
来源: api/payment-errors.md
|
||||
```
|
||||
|
||||
| 组成部分 | 说明 | 大小 |
|
||||
|---------|------|------|
|
||||
| title + summary + keywords | 全部从内存索引获取 | ~100-200 字符 |
|
||||
| **总计** | | **~100-200 字符** |
|
||||
|
||||
### 3.3 多匹配 + L1 无结果
|
||||
|
||||
**策略**: `buildCompactSummary()`(同 3.1)
|
||||
|
||||
L1 未返回结果, L0 作为兜底提供正文内容。
|
||||
|
||||
---
|
||||
|
||||
# 正文内容
|
||||
## 四、决策矩阵
|
||||
|
||||
```
|
||||
needFullContent = highConfidence || !hasL1
|
||||
```
|
||||
|
||||
**输出**:
|
||||
| 场景 | matches | L1 结果 | needFullContent | L0 策略 | 是否读文件 | 上下文大小 |
|
||||
|------|:-------:|:--------:|:---------------:|---------|:---------:|:--------:|
|
||||
| 唯一匹配 | 1 | 未执行 | true | `buildCompactSummary` | 是 | ~600 字符 |
|
||||
| 多匹配 + L1 有结果 | 2+ | 有 | false | `buildMetadataOnlySummary` | **否** | ~150 字符 |
|
||||
| 多匹配 + L1 无结果 | 2+ | 无 | true | `buildCompactSummary` | 是 | ~600 字符 |
|
||||
| 无匹配 | 0 | 有 | — | 无 L0,仅 L1 | 否 | 0 |
|
||||
|
||||
---
|
||||
|
||||
## 五、代码结构
|
||||
|
||||
```
|
||||
LookupKnowledgeTool
|
||||
├── lookupKnowledge(query) # 入口:编排 L0 + L1
|
||||
├── buildResult(l0, l1, confidence) # 组装结果,选择摘要策略
|
||||
├── buildCompactSummary(entry) # 元数据 + 章节 + 正文片段(读文件)
|
||||
├── buildMetadataOnlySummary(entry) # 仅元数据(不读文件)
|
||||
├── countMdHeadings(content) # 统计章节数(日志用)
|
||||
└── extractFirstMeaningfulLine(...) # 提取首个有意义文本行(日志用)
|
||||
```
|
||||
|
||||
### 关键逻辑(buildResult)
|
||||
|
||||
```java
|
||||
Frontmatter {
|
||||
title: "支付网关错误码定义",
|
||||
keywords: ["ERR_TIMEOUT", "超时", "支付网关"],
|
||||
summary: "...",
|
||||
category: "api"
|
||||
}
|
||||
boolean needFullContent = highConfidence || !hasL1;
|
||||
String content = needFullContent
|
||||
? buildCompactSummary(first)
|
||||
: buildMetadataOnlySummary(first);
|
||||
```
|
||||
|
||||
### 3.2 KnowledgeIndexService
|
||||
---
|
||||
|
||||
**职责**:维护 L0 内存索引,提供精确关键词匹配
|
||||
## 六、日志输出示例
|
||||
|
||||
**核心方法**:
|
||||
- `@PostConstruct loadIndex()` - 启动时扫描 knowledge_base/
|
||||
- `exactMatch(String query)` - 精确匹配(不区分大小写)
|
||||
- `readDocument(String filePath, int maxChars)` - 读取文档内容
|
||||
- `addToIndex(KnowledgeEntry entry)` - 添加到索引
|
||||
- `removeFromIndex(String filePath)` - 从索引移除
|
||||
### 多匹配场景(matches=2, L1 有结果)
|
||||
|
||||
**数据结构**:
|
||||
```java
|
||||
List<KnowledgeEntry> knowledgeIndex = new CopyOnWriteArrayList<>();
|
||||
|
||||
KnowledgeEntry {
|
||||
filePath: "knowledge_base/api/payment-errors.md",
|
||||
title: "支付网关错误码定义",
|
||||
keywords: ["ERR_TIMEOUT", "超时", "支付网关"],
|
||||
summary: "...",
|
||||
category: "api"
|
||||
}
|
||||
```
|
||||
[L0 精确匹配] 完成: matches=2, time=3ms
|
||||
[置信度判断] highConfidence=false, reason=多个或零个匹配
|
||||
[L1 语义检索] L0非唯一匹配,触发L1语义检索...
|
||||
[L1 语义检索] 完成: matches=1, time=245ms
|
||||
----------------------------------------
|
||||
<<< [工具返回] lookup_knowledge
|
||||
<<< [L0 主结果] 标题: 支付网关错误码定义
|
||||
<<< [L0 主结果] 摘要: 记录了支付网关所有核心错误码的含义及排查方向 ← 仅元数据
|
||||
<<< [L0 主结果] 内容: 126 字符, 0 个章节 ← 约150字符
|
||||
<<< [L1 补充] 相似度: 0.8234
|
||||
<<< [L1 补充] 内容片段: 支付网关请求超时... ← L1 提供具体内容
|
||||
```
|
||||
|
||||
### 3.3 LookupKnowledgeTool
|
||||
### 唯一匹配场景(matches=1, 跳过 L1)
|
||||
|
||||
**职责**:L0+L1 混合检索工具,Agent 可调用
|
||||
|
||||
**工具定义**:
|
||||
```java
|
||||
@Tool(description = "查询知识库文档。优先精确匹配关键词,未命中或多个匹配时自动补充语义相关片段。" +
|
||||
"参数 query: 查询关键词,例如 'ERR_TIMEOUT'、'支付网关超时'")
|
||||
public LookupResult lookupKnowledge(String query)
|
||||
```
|
||||
[L0 精确匹配] 完成: matches=1, time=2ms
|
||||
[置信度判断] highConfidence=true, reason=唯一匹配
|
||||
[L1 语义检索] L0唯一匹配,跳过L1检索
|
||||
----------------------------------------
|
||||
<<< [工具返回] lookup_knowledge
|
||||
<<< [L0 主结果] 标题: 支付网关错误码定义
|
||||
<<< [L0 主结果] 摘要: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
<<< [L0 主结果] 内容: 725 字符, 3 个章节 ← 约700字符
|
||||
```
|
||||
|
||||
**返回格式**:
|
||||
---
|
||||
|
||||
## 七、MVP 效率评估 & 改进方向
|
||||
|
||||
### 7.1 当前效率评估
|
||||
|
||||
| 维度 | 评分 | 说明 |
|
||||
|------|:----:|------|
|
||||
| L0 匹配速度 | ★★★★★ | 内存索引,< 10ms,几乎没有优化空间 |
|
||||
| L1 检索速度 | ★★★★☆ | Milvus 向量检索,200-500ms,取决于数据量 |
|
||||
| L0 匹配准确率 | ★★☆☆☆ | 子串匹配,无排序无评分,匹配即返回 |
|
||||
| L1 检索准确率 | ★★★☆☆ | 语义相似度,但分块缺少上下文信息 |
|
||||
| 召回率(查全) | ★★★☆☆ | L0+L1 两阶段覆盖大多数场景,但缺乏融合重排 |
|
||||
| 上下文利用率 | ★★★★☆ | 根据场景动态控制 L0 内容量,已优化 |
|
||||
| **综合** | **★★★☆☆** | **MVP 可用,但检索质量有提升空间** |
|
||||
|
||||
### 7.2 关键瓶颈
|
||||
|
||||
#### 瓶颈 1:分块丢失上下文(✅ 已修复—见下方 7.5)
|
||||
|
||||
当前每个 Chunk 只记录最近的 `##` 标题:
|
||||
|
||||
```json
|
||||
{
|
||||
"found": true,
|
||||
"primary": {
|
||||
"content": "文档内容(前 2000 字符)",
|
||||
"source": "knowledge_base/api/payment-errors.md",
|
||||
"matchType": "exact_L0",
|
||||
"confidence": "high"
|
||||
},
|
||||
"supplement": {
|
||||
"content": "语义相关片段(L1)",
|
||||
"source": "metadata",
|
||||
"matchType": "semantic_L1"
|
||||
}
|
||||
"content": "**含义**:支付网关请求超时\n**常见原因**:网络延迟",
|
||||
"title": "超时类错误",
|
||||
"chunkIndex": 2
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
LLM 收到这个片段时**不知道**它属于"支付网关错误码定义"这个文档,也不知道具体错误码名称是 ERR_TIMEOUT。如果同时检索了多个文档的片段,LLM 容易混淆。
|
||||
|
||||
## 四、与现有架构的集成
|
||||
#### 瓶颈 2:L0 关键词匹配过于简单
|
||||
|
||||
### 4.1 Agent 使用场景
|
||||
当前 `KnowledgeIndexService.matchesKeywords()` 只做子串包含匹配,没有:
|
||||
- 排序/评分(多个匹配时按什么顺序?)
|
||||
- 权重(标题匹配 > 正文匹配)
|
||||
- 部分匹配("timeout" 匹配 "ERR_TIMEOUT")
|
||||
|
||||
**ExternalApiSubAgent** (接口专家):
|
||||
```
|
||||
诊断步骤:
|
||||
1. 提取错误码(如 "ERR_TIMEOUT")
|
||||
2. 调用 lookup_knowledge("ERR_TIMEOUT")
|
||||
3. 获得完整错误码定义和排查方向
|
||||
4. 结合日志/链路追踪进行分析
|
||||
```
|
||||
#### 瓶颈 3:L0 和 L1 无交叉融合
|
||||
|
||||
**DatabaseSubAgent** (数据库专家):
|
||||
```
|
||||
诊断步骤:
|
||||
1. 识别数据库问题(如 "连接池满")
|
||||
2. 调用 lookup_knowledge("HikariCP")
|
||||
3. 获得连接池配置最佳实践
|
||||
4. 提供优化建议
|
||||
```
|
||||
|
||||
**Planner Agent** (规划者):
|
||||
```
|
||||
规划阶段:
|
||||
1. 分析问题类型
|
||||
2. 调用 lookup_knowledge("故障诊断")
|
||||
3. 获得标准诊断流程
|
||||
4. 制定排查策略
|
||||
```
|
||||
|
||||
### 4.2 与现有工具对比
|
||||
|
||||
| 工具 | 检索方式 | 响应时间 | 适用场景 | 置信度 |
|
||||
|------|---------|---------|---------|--------|
|
||||
| searchDoc | L1 语义检索 | 200-500ms | 模糊查询、语义理解 | 依赖相似度 |
|
||||
| lookup_knowledge | L0+L1 混合 | < 10ms (高置信) | 精确关键词 + 语义补充 | high/low |
|
||||
|
||||
**推荐使用策略**:
|
||||
- 已知精确关键词(错误码、配置项)→ `lookup_knowledge`
|
||||
- 模糊描述、需要语义理解 → `searchDoc`
|
||||
两阶段检索结果只是简单的"1位L0 + 1位L1"拼接,没有:
|
||||
- RRF 或加权融合重排
|
||||
- 重复内容去重
|
||||
- 根据相关性选择 top-K
|
||||
|
||||
---
|
||||
|
||||
## 五、数据库变更
|
||||
### 7.3 改进方向分析
|
||||
|
||||
### 5.1 api_document 表增强
|
||||
#### 方向 A:面包屑导航(Chunk 携带层级上下文)
|
||||
|
||||
**新增字段**:
|
||||
```sql
|
||||
ALTER TABLE api_document
|
||||
ADD COLUMN metadata TEXT COMMENT 'Frontmatter 元数据 (JSON)';
|
||||
**做法**:分块时记录完整的标题层级路径作为 `breadcrumb`。
|
||||
|
||||
当前分块 metadata:
|
||||
```json
|
||||
{ "title": "超时类错误" }
|
||||
```
|
||||
|
||||
**字段说明**:
|
||||
- 类型:TEXT(最大 64KB)
|
||||
- 格式:JSON 字符串
|
||||
- 内容:frontmatter 解析结果
|
||||
|
||||
**示例数据**:
|
||||
改进后:
|
||||
```json
|
||||
{
|
||||
"title": "支付网关错误码定义",
|
||||
"keywords": ["ERR_TIMEOUT", "超时", "支付网关"],
|
||||
"summary": "记录了支付网关所有核心错误码的含义及排查方向",
|
||||
"title": "超时类错误",
|
||||
"breadcrumb": "支付网关错误码定义 > 超时类错误 > ERR_TIMEOUT",
|
||||
"heading_h1": "支付网关错误码定义",
|
||||
"heading_h2": "超时类错误",
|
||||
"heading_h3": "ERR_TIMEOUT"
|
||||
}
|
||||
```
|
||||
|
||||
**收益评估**:
|
||||
|
||||
| 场景 | 无面包屑的问题 | 有面包屑的改善 | 提升幅度 |
|
||||
|------|---------------|---------------|:--------:|
|
||||
| 单文档多分块 | LLM 知道标题但不知道层级关系 | 清楚"文档>章节>条目"归属 | 中等 |
|
||||
| 跨文档混合结果 | 分块看不出源文档 | breadcrumb 第一段就是文档标题 | 大 |
|
||||
| 深层嵌套文档(3+ 级) | 分块内容难以定位 | 完整路径一目了然 | 显著 |
|
||||
| 向量检索相关性 | 只对 chunk content 做 embedding | breadcrumb 可拼入 content 做 embedding 或单独索引 | 中等 |
|
||||
|
||||
**MVP 阶段价值**:当前文档结构较浅(2-3级),breadcrumb 对 LLM 理解帮助中等。但如果后续文档层级加深(像你提到的"排障指南 > 支付网关 > 502错误处理"),价值会显著提升。
|
||||
|
||||
**实现成本**:低。修改 `DocumentChunkService` 的分块逻辑,积累当前标题栈,写入 `DocumentChunk` 和 Milvus metadata。
|
||||
|
||||
#### 方向 B:混合检索 + RRF 重排
|
||||
|
||||
**做法**:L0 关键词和 L1 向量检索并行执行 → 结果用 Reciprocal Rank Fusion 统一排序 → 取 top-K。
|
||||
|
||||
```
|
||||
用户查询 → 并行的:
|
||||
├── L0 关键词匹配 → 得分向量 S₀
|
||||
└── L1 向量检索 → 得分向量 S₁
|
||||
↓
|
||||
RRF 融合重排
|
||||
↓
|
||||
top-K 统一结果
|
||||
```
|
||||
|
||||
RRF 公式:对每个文档 d,`score(d) = Σ 1/(k + rank_r(d))`,其中 k=60(常数)。
|
||||
|
||||
**收益评估**:
|
||||
|
||||
| 场景 | 当前的问题 | 混合 + RRF | 提升幅度 |
|
||||
|------|-----------|-----------|:--------:|
|
||||
| 精确关键词("ERR_TIMEOUT") | L0 匹配但不排序,L1 可能不匹配 | L0 高排名 → RRF 拉到顶部 | 大 |
|
||||
| 语义查询("支付超时如何处理") | L0 可能不匹配,全靠 L1 | L1 兜底不受影响 | 无变化 |
|
||||
| 混合查询("ERR_TIMEOUT 支付网关超时") | L0 匹配一个、L1 匹配一个,无融合 | RRF 统一排序,更合理 | 中等 |
|
||||
| 多文档匹配 | L0 返回无序列表 + L1 独立结果 | 统一排序、去重 | 大 |
|
||||
|
||||
**MVP 阶段价值**:RRF 的实现成本和维护成本较高,而当前 MVP 数据量小(6 个文档),人工检查即可确定哪些匹配是好的。**建议数据量 > 50 个文档时引入**。
|
||||
|
||||
#### 方向 C:Breadcrumb + Embedding 增强
|
||||
|
||||
**做法**:将 breadcrumb 拼入 chunk content 后再做 embedding,让向量包含层级语义。
|
||||
|
||||
```java
|
||||
// 当前
|
||||
embeddingService.generateEmbedding(chunk.getContent())
|
||||
|
||||
// 改进
|
||||
String augmentedContent = chunk.getBreadcrumb() + "\n" + chunk.getContent();
|
||||
embeddingService.generateEmbedding(augmentedContent);
|
||||
```
|
||||
|
||||
这样搜索"ERR_TIMEOUT"时,"支付网关错误码定义 > 超时类错误 > ERR_TIMEOUT" 也会匹配到,而不只是 chunk 正文。
|
||||
|
||||
| 场景 | 当前 | Breadcrumb + Embedding | 提升 |
|
||||
|------|------|------------------------|:----:|
|
||||
| 搜索"支付网关超时" | 匹配到正文含"超时"和"支付网关"的 chunk | breadcrumb 直接含"支付网关",匹配更准 | 中等 |
|
||||
| 搜索"错误码定义" | 可能匹配不到具体错误内容的 chunk | breadcrumb 含"错误码定义",相关性更高 | 大 |
|
||||
|
||||
---
|
||||
|
||||
### 7.4 实施优先级建议
|
||||
|
||||
| 优先级 | 改进项 | 复杂度 | 收益 | 状态 |
|
||||
|:------:|--------|:------:|:----:|:----:|
|
||||
| P0 | **Breadcrumb 上下文**(方向 A) | 低 | 中 | **✅ 已实现 (2026-06-26)** |
|
||||
| P1 | 下个版本 | 低 | 中-大 | 待定 |
|
||||
| P1 | L0 排序(匹配评分 + 排序) | 低 | 中 | 待定 |
|
||||
| P2 | 混合检索 + RRF 重排 | 高 | 大 | 数据量 > 50 文档时引入 |
|
||||
|
||||
### 7.5 Breadcrumb 实现说明
|
||||
|
||||
已于 2026-06-26 实现。改动范围:
|
||||
|
||||
| 文件 | 改动 |
|
||||
|------|------|
|
||||
| `DocumentChunk.java` | 新增 `breadcrumb` 字段 |
|
||||
| `DocumentChunkService.java` | `splitByHeadings()` 维护标题层级栈,`Section` 新增 `level`/`breadcrumb`,`chunkSection()` 和 `saveChunkAndGetNextStart()` 透传 Breadcrumb |
|
||||
| `VectorIndexService.java` | `buildMetadata()` 和 `buildDocumentMetadata()` 将 breadcrumb 写入 Milvus metadata |
|
||||
|
||||
#### 层级栈算法
|
||||
|
||||
```java
|
||||
// 在 splitByHeadings() 中,每次匹配到标题时:
|
||||
while (!headingStack.isEmpty() && headingStack.size() >= level) {
|
||||
headingStack.remove(headingStack.size() - 1); // 弹出同级或更高级
|
||||
}
|
||||
headingStack.add(title); // 追加当前标题
|
||||
currentBreadcrumb = String.join(" > ", headingStack);
|
||||
```
|
||||
|
||||
示例:处理 `fault-diagnosis-process.md` 的完整面包屑路径──
|
||||
|
||||
```json
|
||||
// 分块 "应急响应流程 > 1. 初步评估"
|
||||
{ "breadcrumb": "故障诊断流程规范 > 应急响应流程 > 1. 初步评估" }
|
||||
|
||||
// 分块 "根因分析方法 > 5-Why 分析法"
|
||||
{ "breadcrumb": "故障诊断流程规范 > 根因分析方法 > 5-Why 分析法" }
|
||||
```
|
||||
|
||||
#### 当前 metadata 结构(Milvus)
|
||||
|
||||
```json
|
||||
{
|
||||
"_source": "knowledge_base/api/payment-errors.md",
|
||||
"_file_name": "payment-errors.md",
|
||||
"category": "api",
|
||||
"version": "1.0",
|
||||
"author": "zhangsan"
|
||||
"chunkIndex": 2,
|
||||
"totalChunks": 5,
|
||||
"title": "超时类错误",
|
||||
"breadcrumb": "支付网关错误码定义 > 超时类错误 > ERR_TIMEOUT"
|
||||
}
|
||||
```
|
||||
|
||||
### 5.2 filePath 字段用途变更
|
||||
以 `fault-diagnosis-process.md` 为例:
|
||||
|
||||
**原用途**:存储相对路径或 URL
|
||||
```markdown
|
||||
# 故障诊断流程规范 ← heading_h1
|
||||
|
||||
**新用途**:存储本地文件绝对路径
|
||||
```
|
||||
knowledge_base/api/payment-errors.md
|
||||
knowledge_base/infrastructure/redis-config.md
|
||||
## 应急响应流程 ← heading_h2
|
||||
|
||||
### 1. 初步评估 ← heading_h3(分块1)
|
||||
内容...
|
||||
|
||||
### 2. 快速止血 ← heading_h3(分块2)
|
||||
内容...
|
||||
|
||||
## 根因分析方法 ← heading_h2
|
||||
|
||||
### 5-Why 分析法 ← heading_h3(分块3)
|
||||
内容...
|
||||
```
|
||||
|
||||
**用途**:
|
||||
1. L0 索引读取完整文档
|
||||
2. 支持未来的章节锚点功能
|
||||
|
||||
---
|
||||
|
||||
## 六、配置说明
|
||||
|
||||
### 6.1 application.yml 新增配置
|
||||
|
||||
```yaml
|
||||
knowledge:
|
||||
base-path: knowledge_base/
|
||||
```
|
||||
|
||||
**说明**:
|
||||
- 相对于项目根目录
|
||||
- 启动时递归扫描此目录
|
||||
- 建议按 category 组织子目录
|
||||
|
||||
### 6.2 目录结构规范
|
||||
改造后每个分块的 metadata:
|
||||
|
||||
```
|
||||
knowledge_base/
|
||||
├── api/ # API 相关文档
|
||||
│ └── payment-errors.md
|
||||
├── infrastructure/ # 基础设施配置
|
||||
│ ├── redis-config.md
|
||||
│ ├── mysql-connection-pool.md
|
||||
│ └── flyway-best-practices.md
|
||||
├── domain/ # 领域知识
|
||||
│ └── spring-ai-tool-best-practices.md
|
||||
└── troubleshooting/ # 故障排查
|
||||
└── fault-diagnosis-process.md
|
||||
分块1: breadcrumb = "故障诊断流程规范 > 应急响应流程 > 1. 初步评估"
|
||||
分块2: breadcrumb = "故障诊断流程规范 > 应急响应流程 > 2. 快速止血"
|
||||
分块3: breadcrumb = "故障诊断流程规范 > 根因分析方法 > 5-Why 分析法"
|
||||
```
|
||||
|
||||
---
|
||||
LLM 视角受益:当检索到 "2. 快速止血" 时,LLM 立刻知道它属于"故障诊断流程规范 > 应急响应流程"体系,不需要额外读取其他分块来推断上下文。
|
||||
|
||||
## 七、性能指标
|
||||
|
||||
### 7.1 查询性能
|
||||
|
||||
| 场景 | L0 耗时 | L1 耗时 | 总耗时 |
|
||||
|------|---------|---------|--------|
|
||||
| 唯一匹配(高置信) | < 5ms | 0 (不调用) | < 10ms |
|
||||
| 多个匹配(低置信) | < 5ms | 200-500ms | < 500ms |
|
||||
| 未匹配(仅L1) | < 5ms | 200-500ms | < 500ms |
|
||||
|
||||
### 7.2 索引性能
|
||||
|
||||
| 指标 | 实测值 | 目标值 |
|
||||
|------|--------|--------|
|
||||
| 启动扫描时间 | < 20ms (6 个文档) | < 1s (500 个文档) |
|
||||
| 内存占用 | < 1MB (6 个文档) | < 5MB (500 个文档) |
|
||||
| L0 匹配时间 | < 5ms | < 10ms |
|
||||
|
||||
---
|
||||
|
||||
## 八、可观测性
|
||||
|
||||
### 8.1 日志追踪
|
||||
|
||||
所有查询都带 requestId(8 位 UUID),可追踪完整流程:
|
||||
|
||||
```
|
||||
[a1b2c3d4] 收到知识库查询请求: query=ERR_TIMEOUT
|
||||
[a1b2c3d4] L0精确匹配完成: matches=1, time=2ms
|
||||
[a1b2c3d4] 置信度判断: highConfidence=true, reason=唯一匹配
|
||||
[a1b2c3d4] L0唯一匹配,跳过L1检索
|
||||
[a1b2c3d4] 查询完成: found=true, confidence=high, totalTime=5ms
|
||||
```
|
||||
|
||||
### 8.2 关键指标
|
||||
|
||||
**监控指标**:
|
||||
- L0 查询耗时(P50/P95/P99)
|
||||
- L1 调用频率(低置信度比例)
|
||||
- 查询总耗时(端到端)
|
||||
- 高置信度命中率
|
||||
|
||||
**告警阈值**:
|
||||
- 查询总耗时 > 2s
|
||||
- L0 索引加载失败
|
||||
- 高置信度命中率 < 20%
|
||||
|
||||
---
|
||||
|
||||
## 九、限制与注意事项
|
||||
|
||||
### 9.1 MVP 阶段限制
|
||||
|
||||
1. **L0 索引无持久化**
|
||||
- 应用重启需要重新扫描
|
||||
- 缓解:启动扫描通常 < 1s
|
||||
|
||||
2. **章节锚点未实现**
|
||||
- sectionTitle 参数预留
|
||||
- availableSections 返回 null
|
||||
|
||||
3. **批量导入不支持**
|
||||
- 当前仅支持单文件上传
|
||||
|
||||
### 9.2 最佳实践
|
||||
|
||||
1. **编写高质量 frontmatter**
|
||||
- keywords 精准且全面
|
||||
- 避免关键词重复(导致多匹配)
|
||||
|
||||
2. **知识库目录组织**
|
||||
- 按 category 分类
|
||||
- 文件命名语义化
|
||||
|
||||
3. **监控告警配置**
|
||||
- 慢查询告警
|
||||
- L0 索引加载失败告警
|
||||
|
||||
---
|
||||
|
||||
## 十、后续增强方向(Phase 2)
|
||||
|
||||
1. **章节锚点**
|
||||
- 支持 sectionTitle 参数
|
||||
- 直接定位到文档特定章节
|
||||
|
||||
2. **L0 索引持久化**
|
||||
- 序列化到文件
|
||||
- 避免重启扫描
|
||||
|
||||
3. **批量导入工具**
|
||||
- 支持目录批量导入
|
||||
- 进度监控
|
||||
|
||||
4. **知识库管理 API**
|
||||
- CRUD 接口
|
||||
- 在线编辑
|
||||
|
||||
5. **向量化元数据**
|
||||
- title/summary 也参与 L1 检索
|
||||
- 提升语义检索准确度
|
||||
| 文件 | 说明 |
|
||||
|------|------|
|
||||
| `LookupKnowledgeTool.java` | 检索工具入口 |
|
||||
| `KnowledgeIndexService.java` | L0 内存索引管理 |
|
||||
| `VectorSearchService.java` | L1 向量检索(Milvus) |
|
||||
| `KnowledgeEntry.java` | 索引条目 DTO(含 title, summary, keywords) |
|
||||
| `LookupResult.java` | 查询结果 DTO |
|
||||
| `PrimaryResult.java` | L0 结果 DTO |
|
||||
| `SupplementResult.java` | L1 结果 DTO |
|
||||
|
||||
@@ -0,0 +1,266 @@
|
||||
# 会话存储方案设计
|
||||
|
||||
**日期**: 2026-06-26
|
||||
**类型**: 架构设计
|
||||
**状态**: 待评审
|
||||
|
||||
---
|
||||
|
||||
## 一、背景与目标
|
||||
|
||||
### 1.1 现状问题
|
||||
|
||||
当前仅有一张 `diagnosis_record` 表,存在以下问题:
|
||||
|
||||
| 问题 | 说明 |
|
||||
|------|------|
|
||||
| **语义耦合** | `fault_category`、`error_code`、`root_cause`、`solution` 等字段耦合在"告警分析"领域语义,ChatService 通用问答场景用不上 |
|
||||
| **Agent 维度缺失** | 只有一个 `tool_calls` JSON 字段,存不下两个 Agent 的多轮决策链 |
|
||||
| **检索质量不可追溯** | 没有记录 L0/L1 命中层、截断信息、召回内容长度 |
|
||||
| **指标不完整** | 有 `duration` 和 `confidence`,但缺 token 用量、自评信号、采纳率 |
|
||||
|
||||
### 1.2 存储范围
|
||||
|
||||
需要覆盖四个层面的数据:
|
||||
|
||||
```
|
||||
诊断级元数据
|
||||
├── 单次诊断的唯一 ID、查询问题、状态
|
||||
├── 会话级决策链
|
||||
│ ├── agent_step:每个 Agent 的每一步(输入、输出、延迟、Token)
|
||||
│ └── tool_invocation:每次工具调用(参数、结果、耗时)
|
||||
├── 检索质量明细
|
||||
│ └── 每次 lookup_knowledge 的命中层(L0/L1)、内容长度、是否截断
|
||||
└── 自评估信号
|
||||
└── LLM 对结论的置信度自评
|
||||
```
|
||||
|
||||
### 1.3 设计目标
|
||||
|
||||
- **可观测**:Debug 时能回溯完整决策链
|
||||
- **可评估**:能统计 L0/L1 命中率、平均 Token 消耗、工具采纳率等指标
|
||||
- **可演进**:覆盖当前两个 Agent(ChatService / AiOpsService),未来新增 Agent 也能接入
|
||||
|
||||
---
|
||||
|
||||
## 二、存储选型分析
|
||||
|
||||
### 2.1 方案对比
|
||||
|
||||
| 维度 | SQL + JSON 列 | NoSQL 文档库 |
|
||||
|------|:------------:|:-----------:|
|
||||
| 基础设施 | 已有的 MySQL,零新增 | 需新部署 MongoDB 等 |
|
||||
| 层级查询 | `WHERE session_id=? AND agent_name=?` 高效 | 需二级索引 |
|
||||
| 指标聚合 | `AVG(token_count) GROUP BY agent_name` 原生支持 | 聚合管道,学习成本 |
|
||||
| 非结构化内容 | JSON 列(MySQL 8+ 支持良好) | 天然支持 |
|
||||
| MVP 迭代速度 | JPA Entity + Flyway 快速迭代 | 新 ORM 学习成本 |
|
||||
|
||||
### 2.2 结论
|
||||
|
||||
**采用 MySQL + JSON 列**。结构化字段做查询和聚合,JSON 列存非结构化载荷。MVP 阶段数据量可控,等后续 > 百万级或需要更灵活 schema 时再评估 NoSQL。
|
||||
|
||||
---
|
||||
|
||||
## 三、存储模型
|
||||
|
||||
### 3.1 整体关系
|
||||
|
||||
```
|
||||
diagnosis_session (1)
|
||||
│
|
||||
└── agent_step (0:N) —— 单次诊断的每一步 Agent 决策
|
||||
│
|
||||
└── tool_invocation (0:N) —— 每步中的工具调用
|
||||
```
|
||||
|
||||
### 3.2 表设计
|
||||
|
||||
#### 表 1:diagnosis_session(诊断会话)
|
||||
|
||||
```sql
|
||||
CREATE TABLE diagnosis_session (
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
session_id VARCHAR(64) UNIQUE NOT NULL COMMENT '会话唯一 ID',
|
||||
|
||||
-- 请求
|
||||
query TEXT NOT NULL COMMENT '用户原始问题',
|
||||
status VARCHAR(16) DEFAULT 'PENDING' COMMENT 'PENDING / RUNNING / SUCCESS / FAILED',
|
||||
agent_flow VARCHAR(32) COMMENT 'CHAT / AI_OPS',
|
||||
|
||||
-- 汇总指标
|
||||
total_duration_ms INT COMMENT '总耗时(毫秒)',
|
||||
total_token_count INT COMMENT '总 Token 消耗',
|
||||
step_count INT COMMENT 'Agent 步数',
|
||||
tool_call_count INT COMMENT '工具调用次数',
|
||||
|
||||
-- 自评估信号(模型对结论的置信度自评)
|
||||
self_evaluation JSON COMMENT '{"confidence": 0-100, "reasoning": "...", "evidence_count": 3}',
|
||||
|
||||
-- 用户反馈
|
||||
feedback VARCHAR(16) COMMENT 'useful / not_useful / null',
|
||||
|
||||
-- 元数据
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
|
||||
|
||||
INDEX idx_created_at (created_at),
|
||||
INDEX idx_status (status),
|
||||
INDEX idx_agent_flow (agent_flow)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='诊断会话表';
|
||||
```
|
||||
|
||||
#### 表 2:agent_step(Agent 决策步骤)
|
||||
|
||||
```sql
|
||||
CREATE TABLE agent_step (
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
session_id VARCHAR(64) NOT NULL COMMENT '关联 diagnosis_session',
|
||||
|
||||
step_index INT NOT NULL COMMENT '当前 Agent 的第几步(从0开始)',
|
||||
agent_name VARCHAR(32) NOT NULL COMMENT 'intelligent_assistant / planner / executor / supervisor',
|
||||
|
||||
-- 模型调用(输入输出摘要,非完整消息体)
|
||||
model_input JSON COMMENT '模型输入摘要 [{role, content_truncated}, ...]',
|
||||
model_output JSON COMMENT '模型输出摘要 {text, tool_calls, ...}',
|
||||
thought TEXT COMMENT 'Agent 思考过程文本',
|
||||
has_tool_call BOOLEAN DEFAULT FALSE COMMENT '本轮是否调用了工具',
|
||||
|
||||
-- 性能指标
|
||||
duration_ms INT COMMENT '本轮耗时',
|
||||
token_count INT COMMENT '本轮 Token 消耗',
|
||||
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
|
||||
INDEX idx_session_step (session_id, step_index),
|
||||
INDEX idx_agent_name (agent_name)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Agent 决策步骤表';
|
||||
```
|
||||
|
||||
#### 表 3:tool_invocation(工具调用明细)
|
||||
|
||||
```sql
|
||||
CREATE TABLE tool_invocation (
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
session_id VARCHAR(64) NOT NULL COMMENT '关联 diagnosis_session',
|
||||
step_id BIGINT COMMENT '关联 agent_step.id(可为空,不强制外键)',
|
||||
|
||||
tool_name VARCHAR(64) NOT NULL COMMENT 'lookup_knowledge / queryPrometheusAlerts / 等',
|
||||
|
||||
-- 调用信息
|
||||
input_params JSON NOT NULL COMMENT '工具入参',
|
||||
output_preview TEXT COMMENT '输出前500字符(可观测用,不存完整输出)',
|
||||
output_length INT COMMENT '输出总字符数',
|
||||
|
||||
-- 检索质量(仅 lookup_knowledge 时有意义)
|
||||
retrieval_layer VARCHAR(8) COMMENT 'L0 / L1 / L0+L1',
|
||||
l0_match_count INT COMMENT 'L0 匹配数',
|
||||
l1_match_count INT COMMENT 'L1 匹配数',
|
||||
is_truncated BOOLEAN DEFAULT FALSE COMMENT '返回内容是否被截断',
|
||||
retrieval_details JSON COMMENT '{"l0_titles":[], "l1_scores":[], "has_supplement": true}',
|
||||
|
||||
-- 性能 & 状态
|
||||
duration_ms INT COMMENT '工具执行耗时',
|
||||
success BOOLEAN DEFAULT TRUE COMMENT '是否成功',
|
||||
error_message TEXT COMMENT '失败原因',
|
||||
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
|
||||
INDEX idx_session_id (session_id),
|
||||
INDEX idx_tool_name (tool_name),
|
||||
INDEX idx_retrieval_layer (retrieval_layer)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='工具调用明细表';
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 四、数据流设计
|
||||
|
||||
### 4.1 完整链路
|
||||
|
||||
```
|
||||
用户请求
|
||||
│
|
||||
▼
|
||||
1. 创建 diagnosis_session(status=RUNNING)
|
||||
│
|
||||
▼
|
||||
2. Agent Loop(可能多轮)
|
||||
│
|
||||
├── beforeModel()
|
||||
│ └── AgentLoggingHook 记录 model_input + 开始时间 → 写入 agent_step(先创建,duration 待填)
|
||||
│
|
||||
├── afterModel()
|
||||
│ └── AgentLoggingHook 记录 model_output + token_count + 工具调用决策 → 更新 agent_step
|
||||
│
|
||||
├── 工具执行(如 lookup_knowledge)
|
||||
│ └── LookupKnowledgeTool 记录 tool_invocation(L0/L1 明细、耗时、是否截断)
|
||||
│
|
||||
└── 循环直到模型不再调用工具
|
||||
│
|
||||
▼
|
||||
3. 诊断完成 → 更新 diagnosis_session
|
||||
├── status = SUCCESS / FAILED
|
||||
├── 汇总指标:total_duration_ms / total_token_count / step_count / tool_call_count
|
||||
└── self_evaluation(可选,由 LLM 自评)
|
||||
```
|
||||
|
||||
### 4.2 变更点
|
||||
|
||||
| 模块 | 当前行为 | 改造后 |
|
||||
|------|---------|--------|
|
||||
| `AgentLoggingHook` | 只打日志到 stdout | 同时写入 `agent_step` 表 |
|
||||
| `LookupKnowledgeTool` | 只打日志到 stdout | 同时写入 `tool_invocation` 表 |
|
||||
| `ChatService` / `AiOpsService` | 执行前后无持久化 | 创建 + 更新 `diagnosis_session` |
|
||||
|
||||
---
|
||||
|
||||
## 五、可观测能力
|
||||
|
||||
### 5.1 查询场景
|
||||
|
||||
| 需求 | SQL | 说明 |
|
||||
|------|-----|------|
|
||||
| 某次诊断用了哪些工具 | `SELECT * FROM tool_invocation WHERE session_id=?` | 按 session 关联 |
|
||||
| lookup_knowledge 的 L0/L1 命中率 | `SELECT retrieval_layer, COUNT(*) FROM tool_invocation WHERE tool_name='lookup_knowledge' GROUP BY retrieval_layer` | 聚合检索层分布 |
|
||||
| 某个 Agent 的平均思考耗时 | `SELECT AVG(duration_ms) FROM agent_step WHERE agent_name=?` | 按 Agent 分组 |
|
||||
| 某次诊断的完整决策链 | `SELECT * FROM agent_step WHERE session_id=? ORDER BY step_index` | 按步骤号排序 |
|
||||
| 被截断的检索占比 | `SELECT COUNT(*) FROM tool_invocation WHERE is_truncated=true AND tool_name='lookup_knowledge'` | 条件计数 |
|
||||
| 高置信度但用户反馈 negative | `SELECT * FROM diagnosis_session WHERE JSON_EXTRACT(self_evaluation, '$.confidence') > 80 AND feedback='not_useful'` | JSON 条件查询 |
|
||||
|
||||
### 5.2 评估指标
|
||||
|
||||
| 指标 | 计算方式 | 数据来源 |
|
||||
|------|---------|---------|
|
||||
| 平均诊断耗时 | `AVG(total_duration_ms)` | diagnosis_session |
|
||||
| 平均 Token 消耗 | `AVG(total_token_count)` | diagnosis_session |
|
||||
| 工具采纳率 | `tools_accepted / tools_proposed` | self_evaluation |
|
||||
| L0 命中率 | `l0_match_count > 0 的比例` | tool_invocation |
|
||||
| 截断率 | `is_truncated=true 的比例` | tool_invocation |
|
||||
| 用户满意度 | `feedback='useful' 的比例` | diagnosis_session |
|
||||
|
||||
---
|
||||
|
||||
## 六、与现有表的关系
|
||||
|
||||
### 6.1 diagnosis_session vs 现有 diagnosis_record
|
||||
|
||||
- **`diagnosis_record`** 保持不动,继续用于"告警分析"场景的领域字段(root_cause、solution 等)
|
||||
- **`diagnosis_session`** 是通用会话存储,覆盖 ChatService 和 AiOpsService
|
||||
- 两者通过 `session_id` 可关联
|
||||
|
||||
### 6.2 迁移策略
|
||||
|
||||
| 阶段 | 动作 |
|
||||
|:----:|------|
|
||||
| MVP | 新建三张表,新代码写入新表 |
|
||||
| V1.1 | 评估是否将 diagnosis_record 合并回 diagnosis_session(加 fault 相关字段到 JSON) |
|
||||
| V1.2 | 数据量 > 10 万时评估是否需要归档或迁移 |
|
||||
|
||||
---
|
||||
|
||||
## 七、未完成事项
|
||||
|
||||
- [ ] AI Ops Supervisor 的 Agent 执行步骤如何对应 agent_step 表(Supervisor 内嵌的子 Agent 步骤归到同一个 session 还是独立)
|
||||
- [ ] self_evaluation 的 confidence 自评通过什么方式获取(单独的 LLM 调用还是在 prompt 中要求输出)
|
||||
- [ ] feedback 字段和前端的交互方式
|
||||
- [ ] Tool_invocation 的 output_preview 截断策略(当前建议 500 字符)
|
||||
@@ -38,4 +38,10 @@ public class DocumentChunk {
|
||||
* 分片标题或上下文信息
|
||||
*/
|
||||
private String title;
|
||||
|
||||
/**
|
||||
* 面包屑导航(完整标题层级路径)
|
||||
* 例如: "故障诊断流程规范 > 应急响应流程 > 1. 初步评估"
|
||||
*/
|
||||
private String breadcrumb;
|
||||
}
|
||||
|
||||
@@ -56,7 +56,7 @@ public class DocumentChunkService {
|
||||
}
|
||||
|
||||
/**
|
||||
* 按照 Markdown 标题分割文档
|
||||
* 按照 Markdown 标题分割文档,同时构建面包屑层级路径
|
||||
*/
|
||||
private List<Section> splitByHeadings(String content) {
|
||||
List<Section> sections = new ArrayList<>();
|
||||
@@ -65,20 +65,34 @@ public class DocumentChunkService {
|
||||
Pattern headingPattern = Pattern.compile("^(#{1,6})\\s+(.+)$", Pattern.MULTILINE);
|
||||
Matcher matcher = headingPattern.matcher(content);
|
||||
|
||||
// 标题层级栈:维护当前标题的完整路径
|
||||
List<String> headingStack = new ArrayList<>();
|
||||
int lastEnd = 0;
|
||||
String currentTitle = null;
|
||||
String currentBreadcrumb = null;
|
||||
|
||||
while (matcher.find()) {
|
||||
int level = matcher.group(1).length(); // #→1, ##→2, ###→3 ...
|
||||
String title = matcher.group(2).trim();
|
||||
|
||||
// 保存上一个章节
|
||||
if (lastEnd < matcher.start()) {
|
||||
String sectionContent = content.substring(lastEnd, matcher.start()).trim();
|
||||
if (!sectionContent.isEmpty()) {
|
||||
sections.add(new Section(currentTitle, sectionContent, lastEnd));
|
||||
sections.add(new Section(
|
||||
headingStack.isEmpty() ? null : headingStack.get(headingStack.size() - 1),
|
||||
level,
|
||||
currentBreadcrumb,
|
||||
sectionContent,
|
||||
lastEnd));
|
||||
}
|
||||
}
|
||||
|
||||
// 更新当前标题
|
||||
currentTitle = matcher.group(2).trim();
|
||||
// 维护层级栈:同级别或更高级别 → 弹出,低级 → 追加
|
||||
while (!headingStack.isEmpty() && headingStack.size() >= level) {
|
||||
headingStack.remove(headingStack.size() - 1);
|
||||
}
|
||||
headingStack.add(title);
|
||||
currentBreadcrumb = String.join(" > ", headingStack);
|
||||
lastEnd = matcher.start();
|
||||
}
|
||||
|
||||
@@ -86,13 +100,18 @@ public class DocumentChunkService {
|
||||
if (lastEnd < content.length()) {
|
||||
String sectionContent = content.substring(lastEnd).trim();
|
||||
if (!sectionContent.isEmpty()) {
|
||||
sections.add(new Section(currentTitle, sectionContent, lastEnd));
|
||||
sections.add(new Section(
|
||||
headingStack.isEmpty() ? null : headingStack.get(headingStack.size() - 1),
|
||||
headingStack.size(),
|
||||
currentBreadcrumb,
|
||||
sectionContent,
|
||||
lastEnd));
|
||||
}
|
||||
}
|
||||
|
||||
// 如果没有找到任何标题,将整个文档作为一个章节
|
||||
if (sections.isEmpty()) {
|
||||
sections.add(new Section(null, content, 0));
|
||||
sections.add(new Section(null, 0, null, content, 0));
|
||||
}
|
||||
|
||||
return sections;
|
||||
@@ -111,6 +130,7 @@ public class DocumentChunkService {
|
||||
List<DocumentChunk> chunks = new ArrayList<>();
|
||||
String content = section.content;
|
||||
String title = section.title;
|
||||
String breadcrumb = section.breadcrumb;
|
||||
|
||||
// 短章节直接作为一个分片(用 token 估算替代字符数做短路判断)
|
||||
if (content.length() <= chunkConfig.getMaxSize()
|
||||
@@ -121,6 +141,7 @@ public class DocumentChunkService {
|
||||
.endOffset(section.startIndex + content.length())
|
||||
.chunkIndex(startChunkIndex)
|
||||
.title(title)
|
||||
.breadcrumb(breadcrumb)
|
||||
.build();
|
||||
chunks.add(chunk);
|
||||
return chunks;
|
||||
@@ -155,7 +176,7 @@ public class DocumentChunkService {
|
||||
logger.debug(" 触及硬上限 ({} tokens),强制切分", tokenCount + paraTokens);
|
||||
chunkParaStart = saveChunkAndGetNextStart(
|
||||
chunks, section, paraPositions,
|
||||
chunkParaStart, i, title, chunkIndex);
|
||||
chunkParaStart, i, title, breadcrumb, chunkIndex);
|
||||
chunkIndex++;
|
||||
|
||||
String prevChunkContent = chunks.get(chunks.size() - 1).getContent();
|
||||
@@ -168,7 +189,7 @@ public class DocumentChunkService {
|
||||
// 安全切点:段落边界
|
||||
chunkParaStart = saveChunkAndGetNextStart(
|
||||
chunks, section, paraPositions,
|
||||
chunkParaStart, i, title, chunkIndex);
|
||||
chunkParaStart, i, title, breadcrumb, chunkIndex);
|
||||
chunkIndex++;
|
||||
|
||||
// 新分片以重叠文本开头
|
||||
@@ -194,6 +215,7 @@ public class DocumentChunkService {
|
||||
.endOffset(section.startIndex + actualEnd)
|
||||
.chunkIndex(chunkIndex)
|
||||
.title(title)
|
||||
.breadcrumb(breadcrumb)
|
||||
.build();
|
||||
chunks.add(chunk);
|
||||
}
|
||||
@@ -213,6 +235,7 @@ public class DocumentChunkService {
|
||||
int fromPara,
|
||||
int toPara,
|
||||
String title,
|
||||
String breadcrumb,
|
||||
int chunkIndex) {
|
||||
|
||||
int actualStart = paraPositions.get(fromPara).start;
|
||||
@@ -225,6 +248,7 @@ public class DocumentChunkService {
|
||||
.endOffset(section.startIndex + actualEnd)
|
||||
.chunkIndex(chunkIndex)
|
||||
.title(title)
|
||||
.breadcrumb(breadcrumb)
|
||||
.build();
|
||||
chunks.add(chunk);
|
||||
|
||||
@@ -392,12 +416,16 @@ public class DocumentChunkService {
|
||||
* 章节数据类
|
||||
*/
|
||||
private static class Section {
|
||||
String title;
|
||||
String content;
|
||||
int startIndex;
|
||||
String title; // 最近一级标题名称
|
||||
int level; // 标题级别(1-6),0=无标题
|
||||
String breadcrumb; // 完整面包屑路径
|
||||
String content; // 章节内容
|
||||
int startIndex; // 在原文中的起始偏移
|
||||
|
||||
Section(String title, String content, int startIndex) {
|
||||
Section(String title, int level, String breadcrumb, String content, int startIndex) {
|
||||
this.title = title;
|
||||
this.level = level;
|
||||
this.breadcrumb = breadcrumb;
|
||||
this.content = content;
|
||||
this.startIndex = startIndex;
|
||||
}
|
||||
|
||||
@@ -269,6 +269,11 @@ public class VectorIndexService {
|
||||
metadata.put("title", chunk.getTitle());
|
||||
}
|
||||
|
||||
// 面包屑导航(完整标题层级路径)
|
||||
if (chunk.getBreadcrumb() != null && !chunk.getBreadcrumb().isEmpty()) {
|
||||
metadata.put("breadcrumb", chunk.getBreadcrumb());
|
||||
}
|
||||
|
||||
// 文档类别
|
||||
metadata.put("category", category != null && !category.isBlank() ? category : "upload");
|
||||
|
||||
@@ -360,6 +365,11 @@ public class VectorIndexService {
|
||||
metadata.put("title", chunk.getTitle());
|
||||
}
|
||||
|
||||
// 面包屑导航(完整标题层级路径)
|
||||
if (chunk.getBreadcrumb() != null && !chunk.getBreadcrumb().isEmpty()) {
|
||||
metadata.put("breadcrumb", chunk.getBreadcrumb());
|
||||
}
|
||||
|
||||
return metadata;
|
||||
}
|
||||
|
||||
|
||||
@@ -9,6 +9,7 @@ import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.stereotype.Component;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* 知识库查询工具
|
||||
@@ -91,23 +92,46 @@ public class LookupKnowledgeTool {
|
||||
// Step 4: 组装结果
|
||||
LookupResult result = buildResult(l0Matches, l1Results, highConfidence);
|
||||
|
||||
// 记录完整结果
|
||||
// 记录结构化结果摘要(替代原始 MD 内容预览)
|
||||
long totalTime = System.currentTimeMillis() - startTime;
|
||||
log.info("----------------------------------------");
|
||||
log.info("<<< [工具返回] lookup_knowledge");
|
||||
log.info("<<< 结果: found={}, matchType={}, confidence={}",
|
||||
result.isFound(),
|
||||
result.getPrimary() != null ? result.getPrimary().getMatchType() : "N/A",
|
||||
result.getPrimary() != null ? result.getPrimary().getConfidence() : "N/A");
|
||||
log.info("<<< 总耗时: {}ms (L0={}ms, L1={}ms)",
|
||||
totalTime, l0Time, l1Results != null ? (totalTime - l0Time) : 0);
|
||||
if (result.isFound() && result.getPrimary() != null) {
|
||||
String content = result.getPrimary().getContent();
|
||||
log.info("<<< 返回内容长度: {} 字符", content != null ? content.length() : 0);
|
||||
if (content != null && content.length() > 200) {
|
||||
log.info("<<< 内容预览: {}", content.substring(0, 200) + "...");
|
||||
log.info("<<< 结果: found={}, 耗时: {}ms (L0={}ms, L1={}ms)",
|
||||
result.isFound(), totalTime, l0Time,
|
||||
l1Results != null ? System.currentTimeMillis() - startTime - l0Time : 0);
|
||||
|
||||
// L0 精确匹配摘要
|
||||
if (!l0Matches.isEmpty()) {
|
||||
KnowledgeEntry top = l0Matches.get(0);
|
||||
log.info("<<< [L0 主结果] 标题: {}", top.getTitle());
|
||||
log.info("<<< [L0 主结果] 来源: {}", top.getFilePath());
|
||||
if (top.getSummary() != null) {
|
||||
log.info("<<< [L0 主结果] 摘要: {}", top.getSummary());
|
||||
}
|
||||
if (top.getKeywords() != null && !top.getKeywords().isEmpty()) {
|
||||
log.info("<<< [L0 主结果] 关键词: {}", String.join(", ", top.getKeywords()));
|
||||
}
|
||||
// 内容概况:长度 + 章节数
|
||||
String content = result.getPrimary() != null ? result.getPrimary().getContent() : null;
|
||||
if (content != null) {
|
||||
int headingCount = countMdHeadings(content);
|
||||
log.info("<<< [L0 主结果] 内容: {} 字符, {} 个章节",
|
||||
content.length(), headingCount);
|
||||
}
|
||||
}
|
||||
|
||||
// L1 语义检索摘要
|
||||
if (l1Results != null && !l1Results.isEmpty()) {
|
||||
VectorSearchService.SearchResult topL1 = l1Results.get(0);
|
||||
log.info("<<< [L1 补充] 来源: {}", topL1.getMetadata() != null ? topL1.getMetadata() : topL1.getId());
|
||||
log.info("<<< [L1 补充] 相似度: {}", String.format("%.4f", topL1.getScore()));
|
||||
if (topL1.getContent() != null) {
|
||||
String snippet = extractFirstMeaningfulLine(topL1.getContent(), 120);
|
||||
log.info("<<< [L1 补充] 内容片段: {}", snippet);
|
||||
log.info("<<< [L1 补充] 片段长度: {} 字符", topL1.getContent().length());
|
||||
}
|
||||
}
|
||||
|
||||
log.info("========================================");
|
||||
|
||||
return result;
|
||||
@@ -132,7 +156,13 @@ public class LookupKnowledgeTool {
|
||||
PrimaryResult primary = null;
|
||||
if (l0Matches != null && !l0Matches.isEmpty()) {
|
||||
KnowledgeEntry first = l0Matches.get(0);
|
||||
String content = knowledgeIndexService.readDocument(first.getFilePath(), 2000);
|
||||
boolean hasL1 = l1Results != null && !l1Results.isEmpty();
|
||||
|
||||
// 场景决策:唯一匹配或 L1 无结果 → LLM 需要正文内容;多匹配且有 L1 → 只需元数据
|
||||
boolean needFullContent = highConfidence || !hasL1;
|
||||
String content = needFullContent
|
||||
? buildCompactSummary(first)
|
||||
: buildMetadataOnlySummary(first);
|
||||
|
||||
if (content != null) {
|
||||
primary = PrimaryResult.builder()
|
||||
@@ -169,4 +199,118 @@ public class LookupKnowledgeTool {
|
||||
|
||||
return builder.build();
|
||||
}
|
||||
|
||||
/**
|
||||
* 统计 MD 文档中的章节数(二级标题 ## 数量)
|
||||
*/
|
||||
private int countMdHeadings(String content) {
|
||||
if (content == null) return 0;
|
||||
return (int) content.lines()
|
||||
.filter(l -> l.trim().startsWith("##"))
|
||||
.count();
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建紧凑文档摘要(替代原始 MD 全文,节省上下文窗口)
|
||||
* 组合:title/summary + 章节结构 + 正文片段(~500 字符)
|
||||
*/
|
||||
private String buildCompactSummary(KnowledgeEntry entry) {
|
||||
String rawContent = knowledgeIndexService.readDocument(entry.getFilePath(), 2000);
|
||||
if (rawContent == null) return null;
|
||||
|
||||
// 跳过 YAML frontmatter 得到正文
|
||||
String body = rawContent;
|
||||
if (body.startsWith("---")) {
|
||||
int end = body.indexOf("---", 3);
|
||||
if (end != -1) {
|
||||
body = body.substring(end + 3).trim();
|
||||
}
|
||||
}
|
||||
|
||||
StringBuilder sb = new StringBuilder();
|
||||
|
||||
// 1. 元数据头(始终包含)
|
||||
sb.append("文档: ").append(entry.getTitle()).append("\n");
|
||||
if (entry.getSummary() != null) {
|
||||
sb.append("摘要: ").append(entry.getSummary()).append("\n");
|
||||
}
|
||||
|
||||
// 2. 章节结构(## 标题列表)
|
||||
String headings = body.lines()
|
||||
.filter(l -> l.trim().startsWith("##"))
|
||||
.map(l -> " - " + l.trim().replaceAll("^#+\\s*", ""))
|
||||
.collect(Collectors.joining("\n"));
|
||||
if (!headings.isEmpty()) {
|
||||
sb.append("章节:\n").append(headings).append("\n");
|
||||
}
|
||||
sb.append("---\n");
|
||||
|
||||
// 3. 正文片段(去标题行、去空行,智能截断)
|
||||
String textContent = body.lines()
|
||||
.filter(l -> !l.trim().startsWith("#") && !l.trim().isEmpty())
|
||||
.collect(Collectors.joining("\n"))
|
||||
.trim();
|
||||
|
||||
// 短文档保留更多内容,长文档节省上下文
|
||||
int maxBodyChars = body.length() < 500 ? 800 : 500;
|
||||
if (textContent.length() > maxBodyChars) {
|
||||
sb.append(textContent, 0, maxBodyChars).append("...");
|
||||
} else {
|
||||
sb.append(textContent);
|
||||
}
|
||||
|
||||
return sb.toString();
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建纯元数据摘要(不读文件,仅用内存索引信息)
|
||||
* 多匹配且有 L1 补充时使用,L0 只需告知 LLM 命中了哪些文档
|
||||
*/
|
||||
private String buildMetadataOnlySummary(KnowledgeEntry entry) {
|
||||
StringBuilder sb = new StringBuilder();
|
||||
sb.append("文档: ").append(entry.getTitle()).append("\n");
|
||||
if (entry.getSummary() != null) {
|
||||
sb.append("摘要: ").append(entry.getSummary()).append("\n");
|
||||
}
|
||||
if (entry.getKeywords() != null && !entry.getKeywords().isEmpty()) {
|
||||
sb.append("关键词: ").append(String.join(", ", entry.getKeywords())).append("\n");
|
||||
}
|
||||
sb.append("来源: ").append(entry.getFilePath()).append("\n");
|
||||
return sb.toString();
|
||||
}
|
||||
|
||||
/**
|
||||
* 提取 MD 内容中第一个有意义的文本行(跳过 frontmatter 和标题行)
|
||||
*/
|
||||
private String extractFirstMeaningfulLine(String content, int maxLen) {
|
||||
if (content == null || content.isBlank()) return "(空)";
|
||||
|
||||
String text = content.trim();
|
||||
// 跳过 YAML frontmatter (--- ... ---)
|
||||
if (text.startsWith("---")) {
|
||||
int end = text.indexOf("---", 3);
|
||||
if (end != -1) {
|
||||
text = text.substring(end + 3);
|
||||
}
|
||||
}
|
||||
|
||||
// 查找第一个非空、非标题行
|
||||
String[] lines = text.split("\n");
|
||||
for (String line : lines) {
|
||||
String tl = line.trim();
|
||||
if (!tl.isEmpty() && !tl.startsWith("#")) {
|
||||
return tl.length() <= maxLen ? tl : tl.substring(0, maxLen) + "...";
|
||||
}
|
||||
}
|
||||
|
||||
// 兜底:第一行非空行
|
||||
for (String line : lines) {
|
||||
if (!line.trim().isEmpty()) {
|
||||
String tl = line.trim();
|
||||
return tl.length() <= maxLen ? tl : tl.substring(0, maxLen) + "...";
|
||||
}
|
||||
}
|
||||
|
||||
return "(无有效内容)";
|
||||
}
|
||||
}
|
||||
|
||||
@@ -49,7 +49,7 @@ class LookupKnowledgeToolTest {
|
||||
when(knowledgeIndexService.exactMatch("ERR_TIMEOUT"))
|
||||
.thenReturn(List.of(entry));
|
||||
when(knowledgeIndexService.readDocument("test.md", 2000))
|
||||
.thenReturn("Test content");
|
||||
.thenReturn("Test content * 用于构建紧凑摘要 * keyword2");
|
||||
|
||||
// 执行查询
|
||||
LookupResult result = tool.lookupKnowledge("ERR_TIMEOUT");
|
||||
@@ -59,7 +59,11 @@ class LookupKnowledgeToolTest {
|
||||
assertNotNull(result.getPrimary());
|
||||
assertEquals("high", result.getPrimary().getConfidence());
|
||||
assertEquals("exact_L0", result.getPrimary().getMatchType());
|
||||
assertEquals("Test content", result.getPrimary().getContent());
|
||||
// 唯一匹配 → buildCompactSummary(),内容为结构化摘要
|
||||
String content = result.getPrimary().getContent();
|
||||
assertTrue(content.contains("文档: Test Doc"));
|
||||
assertTrue(content.contains("摘要: Test summary"));
|
||||
assertTrue(content.contains("Test content"));
|
||||
assertNull(result.getSupplement()); // 高置信度不调用 L1
|
||||
|
||||
// 验证 L1 未被调用
|
||||
@@ -71,7 +75,9 @@ class LookupKnowledgeToolTest {
|
||||
// 准备 L0 多个匹配
|
||||
KnowledgeEntry entry1 = KnowledgeEntry.builder()
|
||||
.filePath("doc1.md")
|
||||
.title("测试文档1")
|
||||
.keywords(List.of("超时"))
|
||||
.summary("这是一个测试文档")
|
||||
.build();
|
||||
|
||||
KnowledgeEntry entry2 = KnowledgeEntry.builder()
|
||||
@@ -81,8 +87,7 @@ class LookupKnowledgeToolTest {
|
||||
|
||||
when(knowledgeIndexService.exactMatch("超时"))
|
||||
.thenReturn(List.of(entry1, entry2));
|
||||
when(knowledgeIndexService.readDocument("doc1.md", 2000))
|
||||
.thenReturn("Content 1");
|
||||
// 多匹配 + L1 有结果 → buildMetadataOnlySummary(),不读文件,不调用 readDocument
|
||||
|
||||
// 准备 L1 结果
|
||||
VectorSearchService.SearchResult l1Result = new VectorSearchService.SearchResult();
|
||||
@@ -99,14 +104,20 @@ class LookupKnowledgeToolTest {
|
||||
assertTrue(result.isFound());
|
||||
assertNotNull(result.getPrimary());
|
||||
assertEquals("low", result.getPrimary().getConfidence()); // 多个匹配 = 低置信度
|
||||
assertEquals("Content 1", result.getPrimary().getContent());
|
||||
// 多匹配 + L1 有结果 → 仅元数据摘要
|
||||
String content = result.getPrimary().getContent();
|
||||
assertTrue(content.contains("文档: 测试文档1"));
|
||||
assertTrue(content.contains("摘要: 这是一个测试文档"));
|
||||
assertTrue(content.contains("关键词: 超时"));
|
||||
assertTrue(content.contains("来源: doc1.md"));
|
||||
|
||||
assertNotNull(result.getSupplement()); // 低置信度调用 L1
|
||||
assertEquals("L1 content", result.getSupplement().getContent());
|
||||
assertEquals("semantic_L1", result.getSupplement().getMatchType());
|
||||
|
||||
// 验证 L1 被调用
|
||||
// 验证 L1 被调用,readDocument 未被调用(多匹配不走 buildCompactSummary)
|
||||
verify(vectorSearchService).searchSimilarDocuments("超时", 3, null);
|
||||
verify(knowledgeIndexService, never()).readDocument(anyString(), anyInt());
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
Reference in New Issue
Block a user