Compare commits

..
Author SHA1 Message Date
aruo 88e0a6c944 docs: reorganize MVP interview documentation 2026-07-05 15:29:28 +08:00
aruo b22f2d22c8 docs: archive historical openspec changes 2026-07-05 14:04:32 +08:00
aruo 63b62b28a2 docs: archive aiops lightweight verifier change 2026-07-05 14:01:33 +08:00
aruo d5902a0499 docs: update mvp architecture snapshot 2026-07-05 13:56:51 +08:00
aruo ed267d753d feat: add aiops lightweight verifier 2026-07-05 13:44:30 +08:00
aruo 2658742119 docs: archive rag and aiops query changes 2026-07-05 13:12:47 +08:00
aruo 72a3dbf8c5 feat: add aiops payload query augmentation 2026-07-05 12:56:20 +08:00
aruo 674dd27a48 feat: add rag post-reindex acceptance 2026-07-05 12:29:24 +08:00
aruo c7e2fc2ee2 feat: include breadcrumb in embedding text 2026-07-05 12:15:27 +08:00
aruo 1bfe1a17b4 docs: add rag refactor story 2026-07-05 12:09:32 +08:00
aruo 9dd6823fe7 docs: add rag retrieval quality report 2026-07-05 11:29:16 +08:00
aruo 9376448804 docs: add rag vectorstore interview notes 2026-07-05 11:13:50 +08:00
aruo f2bae0382c fix: align vectorstore live retrieval 2026-07-05 10:53:45 +08:00
aruo 5c71f5fc79 feat: integrate spring ai vectorstore fallback 2026-07-05 10:20:29 +08:00
aruo b9ec07de57 feat: add spring ai retrieval sidecar 2026-07-05 03:22:03 +08:00
aruo 5197712719 feat: add rag evidence postprocess blocks 2026-07-05 03:03:18 +08:00
aruo 4a94c14feb feat: treat l0 retrieval as domain hint 2026-07-05 02:18:40 +08:00
aruo 9a2a44d1b5 test: add rag retrieval baseline 2026-07-05 02:02:27 +08:00
aruo 79feed3314 Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0
# Conflicts:
#	mvp/issues/README.md
2026-07-05 01:42:42 +08:00
aruo 98155ae1d8 Merge branch 'aiops-trace-scope' into refactor/mvp1.0 2026-07-05 01:42:23 +08:00
aruo 2609c5a5ab docs: consolidate rag refactor issues 2026-07-05 01:40:08 +08:00
aruo bf5286c8f4 docs: add interview project materials 2026-07-05 01:39:37 +08:00
aruo 26e12a8d6b Archive MVP demo interview runbook 2026-07-05 01:34:04 +08:00
aruo cbef3ddd3c Add MVP demo interview runbook 2026-07-05 01:25:20 +08:00
aruo 69deb15330 Add diagnosis eval baseline diff 2026-07-05 00:59:53 +08:00
aruo 4c7c53b024 Expand diagnosis eval fixtures 2026-07-05 00:27:57 +08:00
aruo ca5c61fabf Add diagnosis eval harness 2026-07-04 23:51:43 +08:00
aruo 23ee05c7c3 feat: add traceable scoped AIOps diagnosis 2026-07-04 22:57:28 +08:00
aruo dc6cd32a67 Harden evidence trace semantics 2026-07-04 22:36:30 +08:00
zhuyongxin 246c99b954 add query sql script 2026-07-04 20:14:57 +08:00
zhuyongxin f01866c1a2 refactor: use sequential agent for chat workflow 2026-07-03 18:03:06 +08:00
zhuyongxin 6919092b83 feat: archive mvp demo trace acceptance 2026-07-03 16:25:00 +08:00
zhuyongxin 5b827fe90e fix: avoid low-confidence supervisor retry by default 2026-07-03 15:32:09 +08:00
zhuyongxin b0f288ae36 use supervisor agent for complex chat 2026-07-03 14:10:55 +08:00
zhuyongxin 1ff7f09d25 fix chat session traces and document paths 2026-07-03 13:53:16 +08:00
zhuyongxin fd89d84fc0 docs: add MVP review issue 2026-07-03 11:21:24 +08:00
zhuyongxin 9050487307 feat: add chat verifier agent 2026-07-03 10:54:33 +08:00
zhuyongxin 4f5316d473 chore(cleanup): 清理临时文件和已归档标记
- .gitignore 添加 *.stackdump 和 NUL 规则
- 移除 bash.exe.stackdump 跟踪
- 移除已归档的 .archive-ready 旧标记
2026-07-01 18:32:12 +08:00
zhuyongxin 2a7164288f chore(docs): 补充 ISS-001 架构设计文档到 mvp
- 新增 mvp/architecture/session-dedup-knowledge-map.md
- 更新 mvp/README.md 文档导航
- ISS-001 issue 关联架构文档
2026-07-01 18:28:19 +08:00
zhuyongxin a1c896ebda chore(docs): 归档 ISS-001 session-dedup-knowledge-map + ISS-002 mvp 文档
- 移动 session-dedup-knowledge-map OpenSpec 到 archive 目录
- 提交 ISS-001 遗留的 devflow 档案文件
- 更新 ISS-002 状态为已修复
- 新增 mvp/architecture/action-memory-relevance.md 设计文档
- 更新 mvp/README.md 文档导航
- 更新 devflow/index.md OpenSpec 链接指向 archive
2026-07-01 18:27:04 +08:00
zhuyongxin e438df4355 feat(knowledge): Executor 行动记忆 + 归一化质量等级解决 ISS-002 重复检索
- RetrievedDocTracker 升级为域级+文档级双层记录(Map<sessionId, Map<domain, Set<filePath>>>)
- LookupKnowledgeTool 新增 Min-Max 归一化层(BGE-M3 L2 距离→[0,1] similarity)
- 三等级 relevanceLevel:PRECISE / HIGHLY_RELEVANT / REFERENCE + completenessHint 兜底信号
- LookupResult 新增 relevanceLevel、completenessHint、retrievedDomainsThisSession
- Executor prompt 重写:4 条检索约束 + 合法出口不查全不追责,重复检索才惩罚
- 入库可观测性:V010 迁移 + retrieval_details JSON 扩展
- 归档 executor-action-memory-relevance change
2026-07-01 18:24:41 +08:00
zhuyongxin e4f37cb9e6 fix(knowledge): 修复循环依赖 + 归档 session-dedup-knowledge-map + 记录 ISS-002
- KnowledgeIndexService: 域级生成从 @PostConstruct 移到 @EventListener(ApplicationReadyEvent),解决 KnowledgeIndexService ↔ KnowledgeDomainService 循环依赖
- devflow 归档: evidence.md + acceptance.md(含运行验证结果)
- devflow/index.md: session-dedup-knowledge-map 状态改为 archived
- openspec .archive-ready 标记
- mvp/issues/ISS-002: Executor 无约束重复调用 lookup_knowledge
2026-07-01 14:26:59 +08:00
zhuyongxin 354ffc1947 feat(feedback): 补提交 feedback 相关源码(漏提交的新建文件) 2026-07-01 10:57:49 +08:00
zhuyongxin bb44140901 feat(knowledge): 会话级去重 + 知识域地图注入 Planner 解决 ISS-001 重复检索
- RetrievedDocTracker: sessionId → Set<filePath> 会话级去重,LookupKnowledgeTool Step 5 过滤已检索文档
- KnowledgeDomainService: 域级聚合,LLM 生成 when_to_retrieve,构建 knowledge map YAML
- DocumentFieldEnricher: 上传时 LLM 补全 covers + whenToRetrieve(含同域文档排除上下文)
- KnowledgeDomain entity + V009 迁移: 域级元数据持久化,避免重启重复 LLM 调用
- ChatService: 注入 knowledge map 到 Planner prompt,会话结束时清理去重状态
- KnowledgeIndexService: 手写 JSON 解析替换为 Jackson ObjectMapper,启动时补建缺失域记录
- chat-planner-prompt: 新增知识库检索规则(按域 when_to_retrieve 判断,每域最多一次检索)
- doc-field-enricher-prompt / domain-summary-prompt: 外部化 LLM 提示词
2026-07-01 10:47:46 +08:00
zhuyongxin 2a796da490 feat(feedback): 置信度评分与用户反馈机制 & 归档 confidence-feedback change 2026-06-30 18:01:06 +08:00
zhuyongxin 3ffa5cc366 Merge branch 'emdash/afraid-geese-carry-h5718' into refactor/mvp1.0
# Conflicts:
#	devflow/index.md
2026-06-26 17:36:59 +08:00
zhuyongxin e3f20b1f06 chore: 归档 session-storage change
- 创建 devflow 项目档案(brief/evidence/decisions/acceptance)
- 更新 devflow/index.md 索引
- 移动 OpenSpec 到 archive
2026-06-26 17:33:56 +08:00
zhuyongxin 9b52afce07 docs: 合并 .docs/mvp 到根 mvp 目录并更新文档
- 删除 .docs/mvp 目录,内容合并到根目录 mvp/
- 更新 session-storage-design.md 实现变更记录
- 修复引用路径
2026-06-26 17:29:27 +08:00
zhuyongxin a3abe3f7a2 refactor(session): 清理代码 & RunnableConfig 传 sessionId
- AgentLoggingHook 改为从 config.metadata 读取 sessionId(线程安全)
- 移除 AgentLoggingHook 调试用的 metadata 日志
- TokenTrackingChatModel 日志降为 debug
- SessionContextHolder 移除未使用的 setAgentName/getAgentName
- ChatService 清理无用 import
- 修复 stream 路径下 ThreadLocal NPE
2026-06-26 17:28:30 +08:00
zhuyongxin 0d9cce75f9 feat(session): 会话存储体系实现 & Chat多Agent路由
- 新增诊断会话(diagnosis_session/agent_step/tool_invocation)三表
- AgentLoggingHook 持久化 agent_step,记录决策链和耗时
- LookupKnowledgeTool 写入 tool_invocation,记录L0/L1检索质量
- TokenTrackingChatModel 捕获真实token用量
- Chat接口支持意图路由:简单问题单Agent,复杂问题多Agent(Planner+Executor)
- Prompt外置到 src/main/resources/prompts/
- 删除旧 diagnosis_record 表及相关文件
- 新增SessionContextHolder(ThreadLocal传递sessionId)
- QuestionComplexity 复杂度判断工具
- 测试覆盖三张新表的Repository
2026-06-26 16:22:05 +08:00
zhuyongxin a74ccea5be feat(knowledge): breadcrumb分块上下文 & LookupKnowledgeTool日志优化
- DocumentChunk新增breadcrumb字段,分块时构建完整标题层级路径
- DocumentChunkService splitByHeadings维护标题层级栈算法
- VectorIndexService 将breadcrumb写入Milvus metadata
- LookupKnowledgeTool日志替换为结构化摘要,替代原始MD预览
- L0返回策略:唯一匹配用正文摘要,多匹配+L1有结果仅元数据(不读文件)
- 新增buildCompactSummary / buildMetadataOnlySummary方法
- 安装frontend-design skill
- 创建mvp/文档目录(架构设计+会话存储方案)
- 更新测试适配新逻辑
2026-06-26 13:56:10 +08:00
zhuyongxin a1876286fd fix(observability): 增强模型文本提取,支持多种方式并输出调试信息
## 改动内容

### 增强 extractTextContent() 方法

支持 6 种提取方式,依次尝试:

```java
// 方法 1: 反射获取 text 字段
Field textField = message.getClass().getDeclaredField("text");

// 方法 2: 反射获取 content 字段
Field contentField = message.getClass().getDeclaredField("content");

// 方法 3: 调用 getText() 方法
Method getTextMethod = message.getClass().getMethod("getText");

// 方法 4: 调用 getContent() 方法
Method getContentMethod = message.getClass().getMethod("getContent");

// 方法 5: 打印类结构信息(帮助调试)
log.warn("字段列表: ...");
log.warn("方法列表: ...");

// 方法 6: toString() 兜底
return message.toString();
```

---

## 调试信息输出

### 当提取失败时

```
[WARN] 无法提取 AssistantMessage 文本内容,打印类信息:
[WARN] 类名: org.springframework.ai.chat.messages.AssistantMessage
[WARN] 字段列表:
[WARN]   - text: String
[WARN]   - toolCalls: List
[WARN]   - metadata: Map
[WARN] 方法列表:
[WARN]   - getText(): String
[WARN]   - getToolCalls(): List
[WARN]   - getMetadata(): Map
```

**用途**:
- 帮助快速定位正确的字段/方法名
- 不同 Spring AI 版本可能有不同实现
- 一次调试,永久修复

---

### 当提取成功时

```
[DEBUG] 通过 text 字段提取成功
[INFO] *** [Agent 思考] 模型返回文本: 我需要查询知识库...
```

---

## 适配不同 Spring AI 版本

| 版本 | 字段/方法 | 提取方式 |
|------|----------|---------|
| **Spring AI 0.x** | `text` 字段 | 方法 1 ✅ |
| **Spring AI 1.x** | `content` 字段 | 方法 2 ✅ |
| **阿里云版本** | `getText()` 方法 | 方法 3 ✅ |
| **自定义实现** | `getContent()` 方法 | 方法 4 ✅ |
| **未知版本** | 打印类信息 | 方法 5 → 手动适配 |

---

## 错误处理

### 提取失败但不中断

```java
catch (Exception e) {
    log.error("提取 AssistantMessage 文本内容时出错", e);
    return null;
}

// 调用处
String textContent = extractTextContent(lastAssistant);
if (textContent != null && !textContent.isEmpty()) {
    log.info("*** [Agent 思考] 模型返回文本: {}", textContent);
} else {
    // 跳过,不打印
}
```

**不会中断流程**:
- 提取失败 → 返回 null
- null 检查 → 跳过日志输出
- 继续执行后续逻辑

---

## 使用场景

### 场景 1:首次运行,不确定字段名

```bash
# 启动应用
mvn spring-boot:run

# 发起请求
curl -X POST http://localhost:9900/api/chat \
  -d '{"id":"test","question":"测试"}'

# 查看日志
tail -f logs/application.log | grep "模型返回文本\|字段列表\|方法列表"
```

**如果看到**:
```
[WARN] 无法提取 AssistantMessage 文本内容,打印类信息:
[WARN] 方法列表:
[WARN]   - getTextContent(): String  ← 找到了!
```

**修复**:在 `extractTextContent()` 中添加方法 7:
```java
// 方法 7: 尝试 getTextContent()
Method method = message.getClass().getMethod("getTextContent");
Object value = method.invoke(message);
```

---

### 场景 2:提取成功

```
[DEBUG] 通过 text 字段提取成功
[INFO] *** [Agent 思考] 模型返回文本: 我需要查询知识库来了解支付失败的具体原因
```

正常使用,无需调整。

---

## 性能考虑

### 反射开销

- 反射调用比直接调用慢 ~10-100 倍
- 但只在日志输出时使用,不在热路径
- Agent 调用频率低(秒级),性能影响可忽略

### 优化建议(可选)

缓存反射结果:

```java
private static Field cachedTextField = null;

private String extractTextContent(AssistantMessage message) {
    if (cachedTextField == null) {
        cachedTextField = message.getClass().getDeclaredField("text");
        cachedTextField.setAccessible(true);
    }
    return (String) cachedTextField.get(message);
}
```

**当前未实现**,因为:
- 日志场景无性能瓶颈
- 简单实现更易维护
- 如需优化再添加

---

## 提交历史

```
当前 fix(observability): 增强模型文本提取,支持多种方式并输出调试信息
934d8ee feat(observability): 在 Hook 中输出模型返回的文本内容
7c8758d refactor(observability): 简化 ChatService 日志,避免与 Hook 重复
```
2026-06-25 17:41:37 +08:00
zhuyongxin 934d8eee29 feat(observability): 在 Hook 中输出模型返回的文本内容
## 改动内容

### 增强 AgentLoggingHook.afterModel()

在模型调用完成后,提取并输出模型返回的文本内容:

```java
@Override
public AgentCommand afterModel(List<Message> messages, RunnableConfig config) {
    // 查找最后一条 AssistantMessage
    AssistantMessage lastAssistant = ...;

    // 提取文本内容
    String textContent = extractTextContent(lastAssistant);
    log.info("*** [Agent 思考] 模型返回文本: {}", textContent);

    // 检查工具调用
    if (hasToolCalls) {
        log.info("*** [Agent 思考] 模型决定调用 N 个工具");
    } else {
        log.info("*** [Agent 思考] 这是最终答案");
    }
}
```

---

### extractTextContent() 实现

通过反射提取 AssistantMessage 的文本内容:

```java
private String extractTextContent(AssistantMessage message) {
    try {
        // 尝试获取 text 或 content 字段
        Field textField = message.getClass().getDeclaredField("text");
        textField.setAccessible(true);
        Object value = textField.get(message);
        return value != null ? value.toString() : null;
    } catch (NoSuchFieldException e) {
        // 尝试 content 字段
        try {
            Field contentField = message.getClass().getDeclaredField("content");
            // ...
        } catch (NoSuchFieldException ex) {
            // 字段不存在,返回 null
        }
    }
}
```

**为什么用反射?**
- Spring AI 的 `AssistantMessage` 没有公开的 `getText()` 或 `getContent()` 方法
- 不同版本可能使用 `text` 或 `content` 字段
- 反射可以兼容不同实现

---

## 日志输出示例

### 第 1 轮:模型决定调用工具

```
========================================
*** [Agent 思考] 第 1 轮思考完成
*** [Agent 思考] 模型返回文本: 我需要查询知识库来了解支付失败的原因
*** [Agent 思考] 模型决定调用 1 个工具:
  - 工具: lookup_knowledge, 参数: {"query":"支付失败原因"}
*** [Agent 思考] 等待工具执行结果...
========================================
```

---

### 第 2 轮:模型返回最终答案

```
========================================
*** [Agent 思考] 第 2 轮思考完成
*** [Agent 思考] 模型返回文本: 根据知识库的记录,支付失败的主要原因包括:
1. ERR_TIMEOUT - 支付网关响应超时,通常是网络问题或第三方服务不稳定
2. ERR_INVALID_SIGNATURE - 签名验证失败,检查密钥配置
3. ERR_INSUFFICIENT_BALANCE - 账户余额不足
... (已截断,总长度: 1234)
*** [Agent 思考] 模型决定不调用工具
*** [Agent 思考] 这是最终答案,准备返回给用户
========================================
```

---

## 关键观测点

| 轮次 | 模型输出内容 | 决策 |
|------|-------------|------|
| **第 1 轮** | 模型的推理过程(通常很短) | 决定调用工具 |
| **第 2 轮** | 模型的最终答案(完整回复) | 不调用工具 |

---

## 内容截断策略

- 长度 ≤ 500:完整输出
- 长度 > 500:截断前 500 字符,显示总长度

```
模型返回文本: 根据知识库的记录,支付失败的主要原因包括...
(前 500 字符)
... (已截断,总长度: 1234)
```

---

## 异常处理

如果反射失败(字段不存在或访问被拒绝):

```java
catch (Exception e) {
    log.debug("无法提取 AssistantMessage 文本内容: {}", e.getMessage());
    return null;
}
```

日志输出:
```
*** [Agent 思考] 模型返回文本: (无法提取)
```

不会中断程序,只是跳过文本输出。

---

## 完整的思考流程日志

```
📝 用户问题: 支付为什么会失败?

*** [Agent 思考] 第 1 轮思考开始
*** [Agent 思考] 准备调用模型...
*** [Agent 思考] 第 1 轮思考完成
*** [Agent 思考] 模型返回文本: 我需要查询知识库
*** [Agent 思考] 模型决定调用 1 个工具:
  - 工具: lookup_knowledge, 参数: {"query":"支付失败"}

>>> [工具调用] lookup_knowledge
<<< [工具返回] lookup_knowledge
<<< 结果: found=true

*** [Agent 思考] 第 2 轮思考开始
*** [Agent 思考] 准备调用模型...
*** [Agent 思考] 第 2 轮思考完成
*** [Agent 思考] 模型返回文本: 根据知识库的记录,支付失败...
*** [Agent 思考] 模型决定不调用工具
*** [Agent 思考] 这是最终答案,准备返回给用户

⏱️  总耗时: 1523 ms
📏 输出长度: 456 字符
```

---

## 提交历史

```
当前 feat(observability): 在 Hook 中输出模型返回的文本内容
7c8758d refactor(observability): 简化 ChatService 日志,避免与 Hook 重复
b3ea6e2 feat(observability): 添加 Agent 思考过程日志 Hook
```
2026-06-25 17:27:48 +08:00
zhuyongxin 7c8758d7fa refactor(observability): 简化 ChatService 日志,避免与 Hook 重复
## 改动内容

### 修改前:重复的日志

```java
// ChatService.executeChat()
logger.info("========== Agent 执行开始 ==========");
logger.info("📝 用户问题: {}", question);
logger.info("🚀 执行 ReactAgent.call() - 自动处理工具调用");

// ... Agent 执行 ...

logger.info("========== Agent 执行完成 ==========");
logger.info("⏱️  执行耗时: {} ms", duration);
logger.info("📤 最终输出内容:");
// 打印完整答案
```

**问题**:
- 与 `AgentLoggingHook` 的日志重复
- 日志过于冗长
- Hook 已经覆盖了 Agent 的思考过程

---

### 修改后:精简的日志

```java
// ChatService.executeChat() - 只保留最外层框架
logger.info("========================================");
logger.info("📝 用户问题: {}", question);

// ... Agent 执行(Hook 负责内部日志)...

logger.info("⏱️  总耗时: {} ms", duration);
logger.info("📏 输出长度: {} 字符", answer.length());
logger.info("========================================");
```

**优势**:
- 职责清晰:ChatService 只记录最外层信息
- 避免重复:思考过程由 Hook 负责
- 更简洁:减少噪音日志

---

## 日志分层

| 层级 | 负责类 | 职责 |
|------|--------|------|
| **外层框架** | `ChatService` | 用户问题、总耗时、输出长度 |
| **思考过程** | `AgentLoggingHook` | 每轮思考、模型决策、消息流转 |
| **工具执行** | `LookupKnowledgeTool` | L0/L1 检索、工具参数/返回 |

---

## 日志输出对比

### 修改前(重复冗长)

```
========================================
========== Agent 执行开始 ==========      ← ChatService
========================================
📝 用户问题: 支付为什么会失败?          ← ChatService
----------------------------------------
🚀 执行 ReactAgent.call()               ← ChatService

*** [Agent 思考] 第 1 轮思考开始         ← Hook
...
*** [Agent 思考] 第 1 轮思考完成         ← Hook

>>> [工具调用] lookup_knowledge         ← Tool
...
<<< [工具返回] lookup_knowledge         ← Tool

*** [Agent 思考] 第 2 轮思考开始         ← Hook
...
*** [Agent 思考] 第 2 轮思考完成         ← Hook

========================================
========== Agent 执行完成 ==========      ← ChatService (重复)
========================================
⏱️  执行耗时: 1523 ms                   ← ChatService
📤 最终输出内容:                         ← ChatService
根据知识库的记录...(完整答案)          ← ChatService (太长)
========================================
```

---

### 修改后(清晰简洁)

```
========================================
📝 用户问题: 支付为什么会失败?          ← ChatService (简洁)

*** [Agent 思考] 第 1 轮思考开始         ← Hook
...
*** [Agent 思考] 第 1 轮思考完成         ← Hook

>>> [工具调用] lookup_knowledge         ← Tool
...
<<< [工具返回] lookup_knowledge         ← Tool

*** [Agent 思考] 第 2 轮思考开始         ← Hook
...
*** [Agent 思考] 第 2 轮思考完成         ← Hook
*** [Agent 思考] 这是最终答案            ← Hook (已说明)

⏱️  总耗时: 1523 ms                     ← ChatService (简洁)
📏 输出长度: 456 字符                    ← ChatService (摘要)
========================================
```

**改进**:
- ✅ 去掉重复的"开始/完成"标记
- ✅ 不再打印完整答案(通过 HTTP 响应已返回)
- ✅ Hook 已说明"这是最终答案"
- ✅ 日志更紧凑,信噪比更高

---

## 设计原则

### 1. 单一职责

- **ChatService**:顶层编排,只记录执行框架
- **Hook**:Agent 内部状态,记录思考过程
- **Tool**:工具执行细节,记录检索过程

### 2. 避免重复

- 不在多处打印相同信息
- Hook 已说明"最终答案",ChatService 不再重复

### 3. 信息密度

- 关键信息:保留(用户问题、耗时、长度)
- 冗余信息:删除(重复标题、完整答案)

---

## 查看日志

```bash
# 完整日志
tail -f logs/application.log

# 只看框架
tail -f logs/application.log | grep "📝\|⏱️\|📏"

# 只看思考过程
tail -f logs/application.log | grep "Agent 思考"

# 只看工具调用
tail -f logs/application.log | grep "工具调用\|工具返回"
```

---

## 提交历史

```
当前 refactor(observability): 简化 ChatService 日志,避免与 Hook 重复
b3ea6e2 feat(observability): 添加 Agent 思考过程日志 Hook
8890cd2 feat(observability): 增强 Agent 和工具调用的可观测日志
```
2026-06-25 17:24:59 +08:00
zhuyongxin b3ea6e202d feat(observability): 添加 Agent 思考过程日志 Hook
## 改动内容

### 1. 创建 AgentLoggingHook

基于 Spring AI Alibaba 的 `MessagesModelHook` 实现:

```java
@HookPositions({HookPosition.BEFORE_MODEL, HookPosition.AFTER_MODEL})
public class AgentLoggingHook extends MessagesModelHook {

    // 在模型调用前
    public AgentCommand beforeModel(List<Message> messages, RunnableConfig config)

    // 在模型调用后
    public AgentCommand afterModel(List<Message> messages, RunnableConfig config)
}
```

---

### 2. 集成到 ReactAgent

在 `ChatService.createReactAgent()` 中添加 Hook:

```java
ReactAgent.builder()
    .name("intelligent_assistant")
    .model(chatModel)
    .hooks(new AgentLoggingHook())  // ✅ 添加日志 Hook
    .build();
```

---

## 日志输出示例

### 完整的 Agent 思考流程

```
========================================
========== Agent 执行开始 ==========
========================================
📝 用户问题: 支付为什么会失败?
----------------------------------------
🚀 执行 ReactAgent.call() - 自动处理工具调用

========================================
*** [Agent 思考] 第 1 轮思考开始
*** [Agent 思考] 当前消息数量: 2
*** [Agent 思考] 最近 2 条消息:
  [1] 角色: User(用户), 类型: UserMessage
  [2] 角色: User(用户), 类型: UserMessage
*** [Agent 思考] 准备调用模型...
========================================

========================================
*** [Agent 思考] 第 1 轮思考完成
*** [Agent 思考] 模型输出: <AssistantMessage>
*** [Agent 思考] 模型决定调用 1 个工具:
  - 工具: lookup_knowledge, 参数: {"query":"支付失败原因"}
*** [Agent 思考] 等待工具执行结果...
========================================

========================================
>>> [工具调用] lookup_knowledge
>>> 参数: query = "支付失败原因"
>>> RequestId: a3b4c5d6
----------------------------------------
[L0 精确匹配] 完成: matches=0, time=2ms
[L1 语义检索] L0非唯一匹配,触发L1语义检索...
[L1 语义检索] 完成: matches=1, time=245ms
<<< [工具返回] lookup_knowledge
<<< 结果: found=true, matchType=semantic_L1, confidence=medium
========================================

========================================
*** [Agent 思考] 第 2 轮思考开始
*** [Agent 思考] 当前消息数量: 4
*** [Agent 思考] 最近 3 条消息:
  [1] 角色: User(用户), 类型: UserMessage
  [2] 角色: Assistant(模型), 类型: AssistantMessage
  [3] 角色: Tool(工具返回), 类型: ToolResponseMessage
*** [Agent 思考] 准备调用模型...
========================================

========================================
*** [Agent 思考] 第 2 轮思考完成
*** [Agent 思考] 模型输出: <AssistantMessage>
*** [Agent 思考] 模型决定不调用工具
*** [Agent 思考] 这是最终答案,准备返回给用户
========================================

========================================
========== Agent 执行完成 ==========
========================================
⏱️  执行耗时: 1523 ms
📏 最终输出长度: 456 字符
📤 最终输出内容:
根据知识库的记录,支付失败的主要原因包括...
========================================
```

---

## 核心观测点

| 阶段 | 日志标识 | 信息 |
|------|---------|------|
| **Agent 开始** | `Agent 执行开始` | 用户问题 |
| **思考开始** | `第 N 轮思考开始` | 消息数量、最近消息 |
| **思考完成** | `第 N 轮思考完成` | 模型决策(调用工具 or 返回答案) |
| **工具调用** | `工具调用 lookup_knowledge` | 工具名称、参数 |
| **工具返回** | `工具返回 lookup_knowledge` | 结果摘要、耗时 |
| **Agent 完成** | `Agent 执行完成` | 总耗时、最终输出 |

---

## Hook 机制说明

### MessagesModelHook

- **触发时机**:
  - `BEFORE_MODEL`:模型调用前
  - `AFTER_MODEL`:模型调用后

- **消息流转**:
  ```
  用户问题
      ↓
  [第1轮] beforeModel → 模型决定调用工具 → afterModel
      ↓
  工具执行(lookup_knowledge)
      ↓
  [第2轮] beforeModel → 模型生成最终答案 → afterModel
      ↓
  返回给用户
  ```

- **轮次统计**:
  - 每次调用模型计为一轮
  - 通常需要 2 轮:第 1 轮调用工具,第 2 轮生成答案

---

## 技术细节

### 1. 为什么不用 ModelHook?

`ModelHook` 需要处理 `OverAllState`,更复杂。`MessagesModelHook` 直接操作消息列表,更简单。

### 2. 为什么跳过消息内容?

Spring AI 的 `Message` 接口没有统一的 `getContent()` 方法,不同实现类有不同的访问方式。工具调用的详细内容已在工具层日志体现。

### 3. 消息类型识别

```java
UserMessage          → "User(用户)"
AssistantMessage     → "Assistant(模型)"
ToolResponseMessage  → "Tool(工具返回)"
```

---

## 验证方法

```bash
# 1. 启动应用
mvn spring-boot:run

# 2. 提问
curl -X POST http://localhost:9900/api/chat \
  -H "Content-Type: application/json" \
  -d '{"id":"test","question":"支付为什么会失败?"}'

# 3. 查看完整日志
tail -f logs/application.log

# 4. 过滤关键日志
tail -f logs/application.log | grep -E "Agent|思考|工具|输出"
```

---

## 提交历史

```
当前 feat(observability): 添加 Agent 思考过程日志 Hook
8890cd2 feat(observability): 增强 Agent 和工具调用的可观测日志
f4f0c63 fix(knowledge): 修复 readDocument 文件路径拼接问题
```
2026-06-25 17:02:13 +08:00
zhuyongxin 8890cd2806 feat(observability): 增强 Agent 和工具调用的可观测日志
## 改动内容

### 1. ChatService - Agent 执行日志

在 `executeChat` 方法中添加:

```
========================================
========== Agent 执行开始 ==========
========================================
📝 用户问题: 支付为什么会失败?
----------------------------------------
🚀 执行 ReactAgent.call() - 自动处理工具调用
========================================
========== Agent 执行完成 ==========
========================================
⏱️  执行耗时: 1523 ms
📏 最终输出长度: 456 字符
----------------------------------------
📤 最终输出内容:
根据知识库的记录,支付失败的主要原因是...
========================================
```

**关键信息**:
- 用户问题
- 执行耗时
- 最终输出长度和内容

---

### 2. LookupKnowledgeTool - 工具调用详细日志

```
========================================
>>> [工具调用] lookup_knowledge
>>> 参数: query = "支付为什么会失败?"
>>> RequestId: a3b4c5d6
----------------------------------------
[L0 精确匹配] 完成: matches=0, time=3ms
[置信度判断] highConfidence=false, reason=多个或零个匹配
[L1 语义检索] L0非唯一匹配,触发L1语义检索...
[L1 语义检索] 完成: matches=1, time=245ms
[L1 语义检索] 找到文档:
  - [1] 文档ID: doc-123, 相似度得分: 0.82
----------------------------------------
<<< [工具返回] lookup_knowledge
<<< 结果: found=true, matchType=semantic_L1, confidence=medium
<<< 总耗时: 248ms (L0=3ms, L1=245ms)
<<< 返回内容长度: 1234 字符
<<< 内容预览: ## 支付网关错误码定义...
========================================
```

**关键信息**:
- 工具名称和参数
- L0/L1 执行时间和结果
- 匹配文档列表
- 返回结果摘要

---

## 日志格式说明

### 符号约定

- `>>>` - 工具调用(入参)
- `<<<` - 工具返回(出参)
- `***` - Agent 思考过程(暂未实现)
- `📝` - 用户输入
- `📤` - Agent 输出
- `⏱️` - 性能指标

### 日志级别

- `INFO` - 关键节点和结果
- `DEBUG` - 详细的中间状态(已设置但默认不显示)

---

## 使用场景

### 1. 调试工具调用

```bash
# 查看工具调用详情
grep "工具调用\|工具返回" logs/application.log

# 输出示例
>>> [工具调用] lookup_knowledge
>>> 参数: query = "ERR_TIMEOUT"
<<< [工具返回] lookup_knowledge
<<< 结果: found=true, matchType=exact_L0, confidence=high
```

### 2. 性能分析

```bash
# 查看执行耗时
grep "执行耗时\|总耗时" logs/application.log

# 输出示例
⏱️  执行耗时: 1523 ms
<<< 总耗时: 248ms (L0=3ms, L1=245ms)
```

### 3. L0/L1 验证

```bash
# 查看检索路径
grep "L0精确匹配\|L1语义检索" logs/application.log

# 示例 - L0 命中
[L0 精确匹配] 完成: matches=1, time=3ms
[L0 精确匹配] 找到文档:
  - [1] 标题: 支付网关错误码定义, 路径: api/payment-errors.md
[L1 语义检索] L0唯一匹配,跳过L1检索

# 示例 - L1 命中
[L0 精确匹配] 完成: matches=0, time=2ms
[L1 语义检索] L0非唯一匹配,触发L1语义检索...
[L1 语义检索] 完成: matches=1, time=245ms
```

---

## 后续优化

### 可能的增强(未实现)

由于阿里云 ReactAgent 不支持内置监听器,以下功能暂时无法实现:

- ❌ Agent 思考过程实时监听(`onStateUpdate`)
- ❌ 工具调用前拦截(`onToolCall`)
- ❌ 工具返回后拦截(`onToolResponse`)

如需这些功能,需要:
1. 包装每个工具,统一添加日志
2. 或使用支持监听器的 Agent 框架

当前实现已满足基本可观测需求。

---

## 验证

```bash
# 1. 启动应用
mvn spring-boot:run

# 2. 发起对话
curl -X POST http://localhost:9900/api/chat \
  -H "Content-Type: application/json" \
  -d '{"id":"test","question":"支付为什么会失败?"}'

# 3. 查看日志
tail -f logs/application.log | grep -E "Agent|工具|输出"
```
2026-06-25 16:16:46 +08:00
zhuyongxin f4f0c63325 fix(knowledge): 修复 readDocument 文件路径拼接问题
## 问题

L1 语义检索返回文档后,尝试读取文档内容时报错:

```
读取文档失败: api/payment-errors.md
java.nio.file.NoSuchFileException: api\payment-errors.md
```

**原因**:
- `KnowledgeEntry.filePath` 存储的是相对路径(如 `api/payment-errors.md`)
- `readDocument()` 方法直接使用相对路径读取,未拼接 `knowledge_base` 前缀
- 导致找不到文件

## 修复内容

### 1. 添加 knowledge.base-path 配置

```java
@Value("${knowledge.base-path:knowledge_base}")
private String knowledgeBasePath;
```

**默认值**:`knowledge_base`(当前工作目录下)

### 2. 修复 readDocument 方法

```java
public String readDocument(String filePath, int maxChars) {
    // 拼接完整路径:knowledge_base + 相对路径
    Path fullPath = Paths.get(knowledgeBasePath, filePath);
    String content = Files.readString(fullPath);
    // ...
}
```

**修复前**:
```
读取: api/payment-errors.md
实际路径: <当前目录>/api/payment-errors.md ❌
```

**修复后**:
```
读取: api/payment-errors.md
实际路径: knowledge_base/api/payment-errors.md ✅
```

## 数据流

```
L1 语义检索
    ↓
返回 KnowledgeEntry
filePath = "api/payment-errors.md"
    ↓
readDocument("api/payment-errors.md", 2000)
    ↓
拼接路径: Paths.get("knowledge_base", "api/payment-errors.md")
    ↓
完整路径: "knowledge_base/api/payment-errors.md"
    ↓
Files.readString(fullPath)
    ↓
返回文档内容 ✅
```

## 配置

在 `application.yml` 中可以自定义路径:

```yaml
knowledge:
  base-path: ./knowledge_base  # 默认值
```

或者绝对路径:

```yaml
knowledge:
  base-path: /data/knowledge_base
```

## 验证

```bash
# 1. 启动应用
mvn spring-boot:run

# 2. 提问触发 L1
你:支付为什么会失败?

# 3. Agent 应该:
# - lookup_knowledge("支付为什么会失败")
# - L0 失败 → L1 语义检索
# - 找到 api/payment-errors.md
# - 读取文件内容成功 ✅
# - 返回文档内容给用户
```

## 相关代码路径

- `KnowledgeIndexService.readDocument()` - 文件读取
- `KnowledgeEntry.filePath` - 存储相对路径
- `LookupKnowledgeTool` - 调用 readDocument
2026-06-25 15:56:49 +08:00
zhuyongxin f02a1389c8 refactor(knowledge): KnowledgeIndexService 从数据库加载索引
## 改动内容

### 修改前:从文件系统扫描

```java
@Value("${knowledge.base-path}")
private String knowledgeBasePath;

@Autowired
private FrontmatterParser frontmatterParser;

@PostConstruct
public void loadIndex() {
    // 1. 扫描 knowledge_base 目录
    // 2. 读取每个 .md 文件
    // 3. 解析 frontmatter
    // 4. 构建内存索引
}
```

**问题**:
- 依赖文件系统,无法利用数据库已有数据
- 启动时需要重新扫描和解析所有文件
- 文件和数据库可能不一致

---

### 修改后:从数据库加载

```java
@Autowired
private ApiDocumentRepository apiDocumentRepository;

@PostConstruct
public void loadIndex() {
    // 1. 从数据库读取所有文档
    List<ApiDocument> documents = apiDocumentRepository.findAll();

    // 2. 解析 metadata JSON
    // 3. 构建内存索引
}
```

**优势**:
- ✅ 数据源统一:数据库是唯一真实数据源
- ✅ 启动更快:无需重新扫描文件和解析 frontmatter
- ✅ 数据一致:L0 索引与数据库完全同步
- ✅ 支持动态更新:初始化接口更新数据库后,L0 索引也会更新

---

## 核心方法

### 1. parseDocumentToEntry

从 `ApiDocument` 转换为 `KnowledgeEntry`:

```java
private KnowledgeEntry parseDocumentToEntry(ApiDocument doc) {
    String metadata = doc.getMetadata();

    String title = extractJsonValue(metadata, "title");
    String summary = extractJsonValue(metadata, "summary");
    String category = extractJsonValue(metadata, "category");
    List<String> keywords = extractJsonArray(metadata, "keywords");

    return KnowledgeEntry.builder()
        .filePath(doc.getFilePath())
        .title(title)
        .keywords(keywords)
        .summary(summary)
        .category(category)
        .build();
}
```

### 2. 简单的 JSON 解析

```java
private String extractJsonValue(String json, String key) {
    // 提取 "key":"value" 格式
}

private List<String> extractJsonArray(String json, String key) {
    // 提取 "key":["v1","v2"] 格式
}
```

**注意**:使用简单的字符串解析,避免引入 JSON 库依赖

---

## 数据流

```
应用启动
    ↓
KnowledgeIndexService.loadIndex()
    ↓
apiDocumentRepository.findAll()
    ↓
读取所有 api_document 记录
    ↓
解析每条记录的 metadata JSON
    ↓
构建 KnowledgeEntry
    ↓
加入内存索引(CopyOnWriteArrayList)
    ↓
L0 索引就绪
```

---

## 启动日志

```
[INFO] 开始从数据库加载知识库索引
[INFO] 知识库索引加载完成,共 6 个文档
```

---

## 与初始化接口的配合

### 流程 1:首次启动(数据库为空)

```
1. 应用启动
2. KnowledgeIndexService.loadIndex() → 0 个文档
3. 调用 POST /api/knowledge/init
4. 写入数据库 + 调用 knowledgeIndexService.addToIndex()
5. L0 索引更新为 6 个文档
```

### 流程 2:重启应用(数据库有数据)

```
1. 应用启动
2. KnowledgeIndexService.loadIndex() → 从数据库加载 6 个文档
3. L0 索引已就绪,无需再调用初始化接口
```

---

## 兼容性

### metadata JSON 示例

```json
{
  "title": "支付网关错误码定义",
  "summary": "记录了支付网关所有核心错误码的含义及排查方向",
  "category": "api",
  "keywords": ["ERR_TIMEOUT","超时","支付网关"]
}
```

### 字段映射

| metadata | KnowledgeEntry | 说明 |
|----------|---------------|------|
| title | title | 文档标题 |
| summary | summary | 文档摘要 |
| category | category | 文档分类 |
| keywords | keywords | 关键词列表 |

---

## 删除的代码

- ❌ `@Value("${knowledge.base-path}")`:不再需要文件路径配置
- ❌ `FrontmatterParser` 依赖:不再扫描文件
- ❌ `indexFile()` 方法:不再读取文件
- ❌ `extractCategoryFromPath()` 方法:从 metadata 获取

---

## 验证

```bash
# 1. 清空数据库
# DELETE FROM api_document;

# 2. 启动应用
mvn spring-boot:run

# 3. 查看日志
# [INFO] 知识库索引加载完成,共 0 个文档

# 4. 初始化
curl -X POST http://localhost:9900/api/knowledge/init

# 5. 重启应用
# [INFO] 知识库索引加载完成,共 6 个文档
```
2026-06-25 15:05:29 +08:00
zhuyongxin 91931363d4 refactor(knowledge): 重构 FaultCategory 枚举为文档分类
## 改动内容

### 1. 重构 FaultCategory 枚举

**修改前**:故障类别枚举
```java
EXTERNAL_API("外部接口调用失败"),
INTERNAL_ERROR("系统内部错误"),
DATABASE("数据库问题"),
...
```

**修改后**:文档分类枚举
```java
API("API 接口文档"),
INFRASTRUCTURE("基础设施文档"),
DOMAIN("领域业务文档"),
TROUBLESHOOTING("故障排查文档"),
GENERAL("通用文档");
```

### 2. 新增 fromString 映射方法

```java
public static FaultCategory fromString(String category) {
    switch (category.toLowerCase()) {
        case "api": return API;
        case "infrastructure": return INFRASTRUCTURE;
        case "domain": return DOMAIN;
        case "troubleshooting": return TROUBLESHOOTING;
        default: return GENERAL;
    }
}
```

### 3. 更新所有引用

- `ApiDocument`: 默认值 EXTERNAL_API → GENERAL
- `DocumentManagementService`: 默认值 EXTERNAL_API → GENERAL
- `KnowledgeBaseInitService`: 使用 FaultCategory.fromString() 映射

### 4. 字段映射关系

| Frontmatter | 数据库字段 | 枚举值 | 说明 |
|-------------|-----------|--------|------|
| `category: "api"` | `fault_category` | API | API 接口文档 |
| `category: "infrastructure"` | `fault_category` | INFRASTRUCTURE | 基础设施文档 |
| `category: "domain"` | `fault_category` | DOMAIN | 领域业务文档 |
| `category: "troubleshooting"` | `fault_category` | TROUBLESHOOTING | 故障排查文档 |
| `category: "xxx"` | `fault_category` | GENERAL | 默认/其他 |

## 数据库影响

**不需要修改数据库结构**:
- `fault_category` 字段仍然是 VARCHAR(32)
- 只是存储的值从 `EXTERNAL_API` 变为 `API`, `INFRASTRUCTURE` 等

**已存在的数据**:
- 旧数据中的 `EXTERNAL_API` 仍可以正常读取(枚举向后兼容)
- 新导入的文档会使用新的枚举值

## 验证

```bash
# 1. 重新初始化
curl -X POST http://localhost:9900/api/knowledge/init?force=true

# 2. 查询统计
curl http://localhost:9900/api/knowledge/stats

# 3. 响应
{
  "categories": {
    "API": 1,
    "INFRASTRUCTURE": 3,
    "DOMAIN": 1,
    "TROUBLESHOOTING": 1
  }
}
```

## 数据库查询

```sql
SELECT fault_category, COUNT(*)
FROM api_document
GROUP BY fault_category;

-- 结果
API             | 1
INFRASTRUCTURE  | 3
DOMAIN          | 1
TROUBLESHOOTING | 1
```
2026-06-25 14:23:57 +08:00
zhuyongxin 4e3502a51b fix(knowledge): 将 category 存储到 fault_source 字段
## 问题

数据库表 api_document 没有独立的 category 字段,导致 frontmatter 的 category 信息无法正确存储。

### 表结构分析

```sql
CREATE TABLE api_document (
    fault_category VARCHAR(32) DEFAULT 'EXTERNAL_API',  -- 固定枚举,不合适存储自定义分类
    fault_source VARCHAR(128),                           -- 可以存储自定义分类
    ...
)
```

## 解决方案

使用 `fault_source` 字段存储 frontmatter 的 category:

```java
// 保存时
document.setFaultSource(category);  // api, infrastructure, domain, troubleshooting

// 统计时
Map<String, Long> categoryCount = apiDocumentRepository.findAll().stream()
    .collect(Collectors.groupingBy(
        doc -> doc.getFaultSource() != null ? doc.getFaultSource() : "general",
        Collectors.counting()
    ));
```

## 字段映射关系

| Frontmatter | 数据库字段 | 示例值 |
|-------------|-----------|--------|
| `title` | `api_name` | "支付网关错误码定义" |
| `category` | `fault_source` | "api" / "infrastructure" |
| `keywords` | `metadata` (JSON) | ["ERR_TIMEOUT","超时"] |
| `summary` | `metadata` (JSON) | "记录了..." |

## 优势

1. **充分利用现有字段**:fault_source (VARCHAR 128) 足够存储分类
2. **避免枚举限制**:不受 FaultCategory 枚举约束
3. **查询方便**:直接通过 fault_source 字段查询和统计
4. **向后兼容**:metadata 中仍保留完整的 frontmatter 信息

## 验证

```bash
# 初始化
curl -X POST http://localhost:9900/api/knowledge/init

# 查询统计
curl http://localhost:9900/api/knowledge/stats

# 响应
{
  "categories": {
    "api": 1,
    "infrastructure": 3,
    "domain": 1,
    "troubleshooting": 1
  }
}
```

## 数据库查询

```sql
-- 按分类统计
SELECT fault_source, COUNT(*)
FROM api_document
GROUP BY fault_source;

-- 结果
api             | 1
infrastructure  | 3
domain          | 1
troubleshooting | 1
```
2026-06-25 14:08:50 +08:00
zhuyongxin b01f133efb fix(knowledge): 修复状态字段和 indexed_at 时间戳设置
## 问题

1. **状态字段不正确**:
   - 保存到数据库时直接设置 status="INDEXED"
   - 实际上此时还未索引到 Milvus
   - 应该先设置为 "PENDING",索引成功后更新为 "INDEXED"

2. **indexed_at 时间戳过早**:
   - 在保存数据库时就设置了 indexed_at
   - 应该在 Milvus 索引成功后才设置

3. **fault_category 字段说明**:
   - fault_category 是枚举类型(EXTERNAL_API, DATABASE, CACHE 等)
   - frontmatter 的 category 是自定义分类(api, infrastructure, domain 等)
   - 两者不匹配,保持 fault_category 默认值
   - 真实的分类信息保存在 metadata JSON 中

## 修复内容

### 1. 状态流转正确

```java
// 保存到数据库时
document.setStatus("PENDING");  // 初始状态

// Milvus 索引成功后
document.setStatus("INDEXED");
document.setChunkCount(chunks.size());
document.setIndexedAt(LocalDateTime.now());  // 此时才设置时间戳

// Milvus 索引失败后
document.setStatus("FAILED");
document.setErrorMessage(e.getMessage());
```

### 2. metadata 结构说明

```json
{
  "title": "支付网关错误码定义",
  "summary": "记录了支付网关所有核心错误码的含义及排查方向",
  "category": "api",  // 自定义分类,不是 fault_category
  "keywords": ["ERR_TIMEOUT","超时","支付网关"]
}
```

### 3. 数据库字段含义

- `fault_category`:固定枚举(EXTERNAL_API, DATABASE 等),保持默认值
- `metadata.category`:frontmatter 自定义分类(api, infrastructure, domain 等)
- `status`:索引状态(PENDING → INDEXED / FAILED)
- `indexed_at`:索引完成时间(索引成功后设置)

## 验证

```bash
# 1. 启动应用(Milvus 可以不启动)
mvn spring-boot:run

# 2. 初始化
curl -X POST http://localhost:9900/api/knowledge/init

# 3. 检查数据库
# - Milvus 未启动:status = "FAILED", indexed_at = NULL
# - Milvus 已启动:status = "INDEXED", indexed_at = 实际时间
# - fault_category:始终为 "EXTERNAL_API"(默认值)
# - metadata:包含真实的 category 信息
```
2026-06-25 14:04:10 +08:00
zhuyongxin 3ed48e38cd feat(knowledge): 完整实现知识库初始化 - 包含 Milvus 向量索引
## 核心改动

在上一版本基础上,补充完整的 Milvus (L1) 向量索引功能。

### 新增依赖注入

```java
@Autowired
private DocumentChunkService documentChunkService;

@Autowired
private VectorIndexService vectorIndexService;

@Autowired
private VectorEmbeddingService vectorEmbeddingService;
```

### 完整的数据流

```
knowledge_base/*.md
    ↓ 1. 扫描 & 解析 frontmatter
    ↓ 2. 保存到 MySQL (api_document)
    ↓ 3. 提取正文 & 文档分块
    ↓ 4. 生成向量并索引到 Milvus
    ↓ 5. 加入 L0 内存索引
完成 (L0 + L1 双层索引)
```

### 关键代码

```java
// 1. 提取正文(去除 frontmatter)
String body = extractBody(content);

// 2. 文档分块
List<DocumentChunk> chunks = documentChunkService.chunkDocument(body, relativePath);

// 3. 上传到 Milvus
vectorIndexService.indexDocumentChunks(document.getDocId(), chunks, category);

// 4. 更新状态
document.setStatus("INDEXED");
document.setChunkCount(chunks.size());
```

### 错误处理

- Milvus 索引失败时:
  - 更新文档状态为 FAILED
  - 记录错误信息到 error_message 字段
  - 继续处理下一个文档(不中断整个流程)

### 响应示例

```json
{
  "success": true,
  "scanned": 6,
  "inserted": 6,
  "failed": 0,
  "details": {
    "api/payment-errors.md": "导入成功(L0+L1)"
  }
}
```

### 数据库字段

新增:
- `chunk_count`:分块数量
- `error_message`:错误信息(失败时)

## 验证步骤

```bash
# 1. 启动应用(确保 Milvus 已运行)
mvn spring-boot:run

# 2. 初始化知识库
curl -X POST http://localhost:9900/api/knowledge/init

# 3. 验证结果
# - MySQL: 检查 api_document 表
# - Milvus: 检查 knowledge_base_collection
# - L0: 日志显示"知识库索引加载完成,共 6 个文档"

# 4. 测试 L1 语义检索
# lookup_knowledge("支付为什么会失败")
# 应该返回 semantic_L1 结果
```

## 文档更新

- 更新使用文档,删除"暂未实现 L1"的说明
- 添加 Milvus 数据结构说明
- 添加 Milvus 相关错误处理
2026-06-25 11:00:10 +08:00
zhuyongxin dec587959c feat(knowledge): 添加知识库批量初始化接口
## 新增功能

1. **KnowledgeBaseController**
   - POST /api/knowledge/init - 批量初始化知识库
   - GET /api/knowledge/stats - 查询统计信息

2. **KnowledgeBaseInitService**
   - 递归扫描 knowledge_base 目录所有 .md 文件
   - 解析 frontmatter 提取元数据
   - 自动去重(基于文件路径)
   - 数据入库到 api_document 表
   - 自动加入 L0 内存索引

## 核心特性

### 去重机制
- 基于文件相对路径去重
- 支持 force=true 强制重新导入
- 跳过已存在文档,避免重复插入

### 数据存储
- 数据库:保存文档元数据(title、keywords、summary 等)
- L0 索引:加入 KnowledgeIndexService 内存索引
- L1 索引:暂未实现(TODO)

### 错误处理
- 格式无效:frontmatter 解析失败
- 缺少标题:必填字段验证
- 详细的错误信息反馈

## API 示例

```bash
# 首次导入
curl -X POST http://localhost:9900/api/knowledge/init

# 强制重新导入
curl -X POST http://localhost:9900/api/knowledge/init?force=true

# 查询统计
curl http://localhost:9900/api/knowledge/stats
```

## 响应示例

```json
{
  "success": true,
  "scanned": 6,
  "skipped": 0,
  "inserted": 6,
  "failed": 0,
  "details": {
    "api/payment-errors.md": "导入成功(L0)"
  }
}
```

## 后续扩展

- [ ] L1 向量索引(Milvus)集成
- [ ] 文档更新检测(基于文件哈希)
- [ ] 批量删除接口
- [ ] 进度回调支持

## 文档

- 使用文档:.docs/2026-06-25-knowledge-base-init-api.md
2026-06-25 10:49:30 +08:00
zhuyongxin f002571629 refactor(ai-ops): 增强 lookup_knowledge 工具描述并调整工具优先级
## 主要改动

1. 增强 lookup_knowledge 工具描述
   - 参考 queryLogs 的详细描述格式
   - 添加 IMPORTANT 关键词强调优先使用场景
   - 详细列举 4 种查询场景及示例:
     * 错误码定义(ERR_TIMEOUT)
     * 接口文档(payment-gateway)
     * 排障步骤(支付超时排查)
     * 配置说明(HikariCP)
   - 保留性能优势说明(L0 < 10ms, L1 200-500ms)

2. 调整工具数组顺序
   - 将 lookupKnowledgeTool 提前到第 2 位(仅次于 dateTimeTools)
   - Mock 模式顺序:dateTimeTools → lookupKnowledgeTool → queryMetricsTools → queryLogsTools
   - 真实模式顺序:dateTimeTools → lookupKnowledgeTool → queryMetricsTools → internalDocsTools
   - 弃用工具(internalDocsTools)放在最后

## 设计目标

解决 Executor 优先选择 queryLogsTools 的问题:
- 工具描述对等:lookup_knowledge 与 queryLogs 同等详细
- 位置优先:知识库查询排在日志查询之前
- 明确引导:IMPORTANT 关键词强调使用时机

## 预期效果

Executor 在遇到错误码、配置项、排障问题时,应优先调用 lookup_knowledge,
而不是直接查询日志。
2026-06-25 09:55:02 +08:00
zhuyongxin 553d1d1faf refactor(ai-ops): 重构 Executor Prompt - 强化行为准则和证据驱动
## 主要改动

1. 重构为更清晰的三段式结构
   - 角色定位:明确"诊断流程的执行者"
   - 核心行为准则:严格按步执行、必须调用工具、证据综合分析
   - 任务执行规范:每步产出要求、最终报告格式

2. 强化关键约束
   - 永远不要凭记忆回答错误码含义、接口定义、排障步骤
   - 结论必须基于至少两个独立证据源
   - 提供证据链格式示例

3. 简化 lookup_knowledge 说明
   - 修正参数名:query_text → query
   - 简化返回字段说明(保留核心信息)
   - 保留三级使用规则(必须/必须/建议)

## 设计理念

- 从"技术细节"转向"行为准则"
- 从"字段说明"转向"证据驱动"
- 提供具体的输出格式示例,减少 Agent 的不确定性

## 参考

基于 mvp/discuss/Executor_Prompt.md 微调
2026-06-24 18:53:42 +08:00
zhuyongxin c88b287f83 refactor(ai-ops): 优化 lookup_knowledge 工具描述和 Executor Prompt
## 主要改动

1. 优化工具描述(中等版)
   - 保留 L0/L1 两阶段检索机制说明
   - 增加适用场景列举(错误码、接口文档、排障步骤等)
   - 简化为核心信息,减少 token 消耗

2. Executor Prompt 新增详细使用说明
   - 添加 lookup_knowledge 返回结果字段说明
   - 提供 confidence 和 match_type 的使用建议
   - 明确 found=false 的处理方式
   - 结构化组织:工具选择 → 结果处理 → 执行反馈

## 设计思路

- 工具描述:简洁,快速理解核心用途
- Executor Prompt:详细,指导正确使用
- 分层设计:减少重复信息,降低 token 消耗
2026-06-24 18:48:04 +08:00
zhuyongxin 463d8b817b fix: 恢复 queryInternalDocs 的原始工具描述
保留原始 @Tool description,仅在 Java 层面标记 @Deprecated。
这样可以保留完整的工具提示词用于后续对比分析。
2026-06-24 18:41:02 +08:00
zhuyongxin 363767d3e7 refactor(ai-ops): 标记 queryInternalDocs 为弃用,统一使用 lookup_knowledge
## 改动说明

1. 标记 InternalDocsTools 为 @Deprecated
   - 添加弃用注解和说明文档
   - 工具描述中明确提示使用 lookup_knowledge 替代

2. 简化 Executor Prompt
   - 移除 queryInternalDocs 相关的工具选择逻辑
   - 统一使用 lookup_knowledge 处理所有知识库查询
   - 精确关键词、模糊概念、故障流程都使用同一个工具

## 理由

lookup_knowledge 已经支持:
- L0 精确匹配(< 10ms,高置信度)
- L1 语义检索(自动兜底)

功能完全覆盖 queryInternalDocs(纯 L1 检索),且性能更优。
保留 queryInternalDocs 会导致:
- 工具功能重叠,Agent 决策困难
- 维护两套相似的代码逻辑

## 迁移路径

- 当前:标记为弃用,但保持可用
- 验证:观察 lookup_knowledge 是否能完全替代
- 未来:确认无问题后,在下个版本中移除
2026-06-24 18:38:39 +08:00
zhuyongxin c4d23c3bd8 feat(ai-ops): Prompt 配置化 & 集成 LookupKnowledgeTool
## 主要改动

1. Prompt 配置化
   - 从硬编码改为独立 Markdown 文件管理
   - 新增 AiOpsPromptProperties 配置类,使用 @PostConstruct 加载
   - 创建 prompts/{planner,executor,supervisor}-prompt.md

2. 集成 LookupKnowledgeTool
   - 在 AiOpsService 中注入 LookupKnowledgeTool
   - 添加到工具数组,只给 Executor Agent 使用
   - 符合 3-Agent 协同分析模式

3. Executor Prompt 增强
   - 添加工具选择指南(精确关键词 vs 模糊概念)
   - 明确降级策略(lookup_knowledge 未找到时降级到 queryInternalDocs)

## 优势

- 易于维护:Prompt 修改不需重新编译
- 格式友好:Markdown 原生支持代码块和表格
- 性能优化:精确关键词查询 < 10ms(L0 匹配)

## 文件变更

- 新增:AiOpsPromptProperties.java
- 新增:prompts/planner-prompt.md
- 新增:prompts/executor-prompt.md
- 新增:prompts/supervisor-prompt.md
- 修改:AiOpsService.java(-136 行硬编码,+5 行配置引用)
2026-06-24 18:29:24 +08:00
425 changed files with 30462 additions and 1657 deletions
+57
View File
@@ -0,0 +1,57 @@
# Frontend Design — Complete Guidance
This document provides a comprehensive framework for creating visually distinctive, non-templated UI designs. Here's the full breakdown:
## Foundational Approach
Act as the design lead for a studio known for unique client identities — the client has already turned down template-like proposals. Every choice about palette, typography, and layout must be specific to the brief, including "one real aesthetic risk you can justify."
## Grounding in Subject Matter
If the brief is vague about the product or subject, pin it down yourself: name the subject, its audience, and the page's single job. Draw inspiration from "the subject's own world, its materials, instruments, artifacts, and vernacular." Use any known context about the human's preferences or past designs as hints.
## Design Principles
- **Hero as thesis**: Open with "the most characteristic thing in the subject's world" — avoid default choices like a big number with a small label and gradient accent unless truly optimal.
- **Typography**: Pair display and body faces deliberately, not from your usual repertoire. Set a clear type scale with intentional weights, widths, and spacing. "Make the type treatment itself a memorable part of the design."
- **Structure as information**: Numbering, eyebrows, dividers must encode something true about the content. Question whether numbered markers (01/02/03) actually make sense before using them — only appropriate for real sequences.
- **Motion**: Consider where animation serves the subject. "An orchestrated moment usually lands harder than scattered effects." Sometimes less is better to avoid an AI-generated feel.
- **Complexity**: Match execution to the vision — maximalist needs elaborate execution, minimal needs precision.
- **Content**: Come up with copy if the brief lacks it. Poor copy makes a design feel as templated as poor layout.
## AI-Generated Design Traps
Three common AI-default looks to watch for: (1) warm cream background (~#F4F1EA) with serif display and terracotta accent; (2) near-black with bright acid-green or vermilion; (3) broadsheet layout with hairline rules, zero border-radius, and dense columns. "All three are legitimate for some briefs, but they are defaults rather than choices." Where the brief leaves an axis free, don't spend that freedom on a default.
## Two-Pass Process
**Pass 1 — Plan**: Create a compact token system:
1. **Color**: 4–6 named hex values
2. **Type**: Characterful display face (used with restraint), complementary body face, utility face for captions/data
3. **Layout**: One-sentence prose descriptions + ASCII wireframes
4. **Signature**: The single unique element the page will be remembered by
Review the plan against the brief. If any part reads like what you'd produce for any similar page, revise it. Only then write code.
**Pass 2 — Build**: Follow the revised plan exactly. Watch for CSS selector specificity conflicts (e.g., `.section` and `.cta` fighting over padding/margins). Do most planning internally, only sharing ideas when confident.
## Restraint & Self-Critique
"Spend your boldness in one place" — let the signature element be the one memorable thing; keep everything else quiet. "Not taking a risk can be a risk itself!" Build responsively down to mobile, with visible keyboard focus and reduced motion respected. Critique as you build. Follow Chanel's advice: before finishing, remove one accessory. Jot notes about what you've tried to avoid repeating yourself.
## Writing in Design
Words exist to make the design understandable and usable — they're "design material, not decoration." Write from the end user's perspective, naming things by what people control and recognize, never by how the system is built.
- Use active voice as default
- A control should say exactly what happens: "Save changes," not "Submit"
- Maintain consistent vocabulary throughout flows (button says "Publish," toast says "Published")
- Treat errors as guidance, not mood — explain what went wrong and how to fix it
- Empty screens are invitations to act
- Keep the register conversational: "plain verbs, sentence case, no filler"
- Let each element do exactly one job — "a label labels, an example demonstrates"
## License
Apache License 2.0 — see LICENSE.txt
@@ -0,0 +1,156 @@
---
name: openspec-apply-change
description: Implement tasks from an OpenSpec change. Use when the user wants to start implementing, continue implementation, or work through tasks.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Implement tasks from an OpenSpec change.
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
**Steps**
1. **Select the change**
If a name is provided, use it. Otherwise:
- Infer from conversation context if the user mentioned a change
- Auto-select if only one active change exists
- If ambiguous, run `openspec list --json` to get available changes and use the **AskUserQuestion tool** to let the user select
Always announce: "Using change: <name>" and how to override (e.g., `/opsx:apply <other>`).
2. **Check status to understand the schema**
```bash
openspec status --change "<name>" --json
```
Parse the JSON to understand:
- `schemaName`: The workflow being used (e.g., "spec-driven")
- Which artifact contains the tasks (typically "tasks" for spec-driven, check status for others)
3. **Get apply instructions**
```bash
openspec instructions apply --change "<name>" --json
```
This returns:
- `contextFiles`: artifact ID -> array of concrete file paths (varies by schema - could be proposal/specs/design/tasks or spec/tests/implementation/docs)
- Progress (total, complete, remaining)
- Task list with status
- Dynamic instruction based on current state
**Handle states:**
- If `state: "blocked"` (missing artifacts): show message, suggest using openspec-continue-change
- If `state: "all_done"`: congratulate, suggest archive
- Otherwise: proceed to implementation
4. **Read context files**
Read every file path listed under `contextFiles` from the apply instructions output.
The files depend on the schema being used:
- **spec-driven**: proposal, specs, design, tasks
- Other schemas: follow the contextFiles from CLI output
5. **Show current progress**
Display:
- Schema being used
- Progress: "N/M tasks complete"
- Remaining tasks overview
- Dynamic instruction from CLI
6. **Implement tasks (loop until done or blocked)**
For each pending task:
- Show which task is being worked on
- Make the code changes required
- Keep changes minimal and focused
- Mark task complete in the tasks file: `- [ ]` → `- [x]`
- Continue to next task
**Pause if:**
- Task is unclear → ask for clarification
- Implementation reveals a design issue → suggest updating artifacts
- Error or blocker encountered → report and wait for guidance
- User interrupts
7. **On completion or pause, show status**
Display:
- Tasks completed this session
- Overall progress: "N/M tasks complete"
- If all done: suggest archive
- If paused: explain why and wait for guidance
**Output During Implementation**
```
## Implementing: <change-name> (schema: <schema-name>)
Working on task 3/7: <task description>
[...implementation happening...]
✓ Task complete
Working on task 4/7: <task description>
[...implementation happening...]
✓ Task complete
```
**Output On Completion**
```
## Implementation Complete
**Change:** <change-name>
**Schema:** <schema-name>
**Progress:** 7/7 tasks complete ✓
### Completed This Session
- [x] Task 1
- [x] Task 2
...
All tasks complete! Ready to archive this change.
```
**Output On Pause (Issue Encountered)**
```
## Implementation Paused
**Change:** <change-name>
**Schema:** <schema-name>
**Progress:** 4/7 tasks complete
### Issue Encountered
<description of the issue>
**Options:**
1. <option 1>
2. <option 2>
3. Other approach
What would you like to do?
```
**Guardrails**
- Keep going through tasks until done or blocked
- Always read context files before starting (from the apply instructions output)
- If task is ambiguous, pause and ask before implementing
- If implementation reveals issues, pause and suggest artifact updates
- Keep code changes minimal and scoped to each task
- Update task checkbox immediately after completing each task
- Pause on errors, blockers, or unclear requirements - don't guess
- Use contextFiles from CLI output, don't assume specific file names
**Fluid Workflow Integration**
This skill supports the "actions on a change" model:
- **Can be invoked anytime**: Before all artifacts are done (if tasks exist), after partial implementation, interleaved with other actions
- **Allows artifact updates**: If implementation reveals design issues, suggest updating artifacts - not phase-locked, work fluidly
@@ -0,0 +1,114 @@
---
name: openspec-archive-change
description: Archive a completed change in the experimental workflow. Use when the user wants to finalize and archive a change after implementation is complete.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Archive a completed change in the experimental workflow.
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
**Steps**
1. **If no change name provided, prompt for selection**
Run `openspec list --json` to get available changes. Use the **AskUserQuestion tool** to let the user select.
Show only active changes (not already archived).
Include the schema used for each change if available.
**IMPORTANT**: Do NOT guess or auto-select a change. Always let the user choose.
2. **Check artifact completion status**
Run `openspec status --change "<name>" --json` to check artifact completion.
Parse the JSON to understand:
- `schemaName`: The workflow being used
- `artifacts`: List of artifacts with their status (`done` or other)
**If any artifacts are not `done`:**
- Display warning listing incomplete artifacts
- Use **AskUserQuestion tool** to confirm user wants to proceed
- Proceed if user confirms
3. **Check task completion status**
Read the tasks file (typically `tasks.md`) to check for incomplete tasks.
Count tasks marked with `- [ ]` (incomplete) vs `- [x]` (complete).
**If incomplete tasks found:**
- Display warning showing count of incomplete tasks
- Use **AskUserQuestion tool** to confirm user wants to proceed
- Proceed if user confirms
**If no tasks file exists:** Proceed without task-related warning.
4. **Assess delta spec sync state**
Check for delta specs at `openspec/changes/<name>/specs/`. If none exist, proceed without sync prompt.
**If delta specs exist:**
- Compare each delta spec with its corresponding main spec at `openspec/specs/<capability>/spec.md`
- Determine what changes would be applied (adds, modifications, removals, renames)
- Show a combined summary before prompting
**Prompt options:**
- If changes needed: "Sync now (recommended)", "Archive without syncing"
- If already synced: "Archive now", "Sync anyway", "Cancel"
If user chooses sync, use Task tool (subagent_type: "general-purpose", prompt: "Use Skill tool to invoke openspec-sync-specs for change '<name>'. Delta spec analysis: <include the analyzed delta spec summary>"). Proceed to archive regardless of choice.
5. **Perform the archive**
Create the archive directory if it doesn't exist:
```bash
mkdir -p openspec/changes/archive
```
Generate target name using current date: `YYYY-MM-DD-<change-name>`
**Check if target already exists:**
- If yes: Fail with error, suggest renaming existing archive or using different date
- If no: Move the change directory to archive
```bash
mv openspec/changes/<name> openspec/changes/archive/YYYY-MM-DD-<name>
```
6. **Display summary**
Show archive completion summary including:
- Change name
- Schema that was used
- Archive location
- Whether specs were synced (if applicable)
- Note about any warnings (incomplete artifacts/tasks)
**Output On Success**
```
## Archive Complete
**Change:** <change-name>
**Schema:** <schema-name>
**Archived to:** openspec/changes/archive/YYYY-MM-DD-<name>/
**Specs:** ✓ Synced to main specs (or "No delta specs" or "Sync skipped")
All artifacts complete. All tasks complete.
```
**Guardrails**
- Always prompt for change selection if not provided
- Use artifact graph (openspec status --json) for completion checking
- Don't block archive on warnings - just inform and confirm
- Preserve .openspec.yaml when moving to archive (it moves with the directory)
- Show clear summary of what happened
- If sync is requested, use openspec-sync-specs approach (agent-driven)
- If delta specs exist, always run the sync assessment and show the combined summary before prompting
+288
View File
@@ -0,0 +1,288 @@
---
name: openspec-explore
description: Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements. Use when the user wants to think through something before or during a change.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Enter explore mode. Think deeply. Visualize freely. Follow the conversation wherever it goes.
**IMPORTANT: Explore mode is for thinking, not implementing.** You may read files, search code, and investigate the codebase, but you must NEVER write code or implement features. If the user asks you to implement something, remind them to exit explore mode first and create a change proposal. You MAY create OpenSpec artifacts (proposals, designs, specs) if the user asks—that's capturing thinking, not implementing.
**This is a stance, not a workflow.** There are no fixed steps, no required sequence, no mandatory outputs. You're a thinking partner helping the user explore.
---
## The Stance
- **Curious, not prescriptive** - Ask questions that emerge naturally, don't follow a script
- **Open threads, not interrogations** - Surface multiple interesting directions and let the user follow what resonates. Don't funnel them through a single path of questions.
- **Visual** - Use ASCII diagrams liberally when they'd help clarify thinking
- **Adaptive** - Follow interesting threads, pivot when new information emerges
- **Patient** - Don't rush to conclusions, let the shape of the problem emerge
- **Grounded** - Explore the actual codebase when relevant, don't just theorize
---
## What You Might Do
Depending on what the user brings, you might:
**Explore the problem space**
- Ask clarifying questions that emerge from what they said
- Challenge assumptions
- Reframe the problem
- Find analogies
**Investigate the codebase**
- Map existing architecture relevant to the discussion
- Find integration points
- Identify patterns already in use
- Surface hidden complexity
**Compare options**
- Brainstorm multiple approaches
- Build comparison tables
- Sketch tradeoffs
- Recommend a path (if asked)
**Visualize**
```
┌─────────────────────────────────────────┐
│ Use ASCII diagrams liberally │
├─────────────────────────────────────────┤
│ │
│ ┌────────┐ ┌────────┐ │
│ │ State │────────▶│ State │ │
│ │ A │ │ B │ │
│ └────────┘ └────────┘ │
│ │
│ System diagrams, state machines, │
│ data flows, architecture sketches, │
│ dependency graphs, comparison tables │
│ │
└─────────────────────────────────────────┘
```
**Surface risks and unknowns**
- Identify what could go wrong
- Find gaps in understanding
- Suggest spikes or investigations
---
## OpenSpec Awareness
You have full context of the OpenSpec system. Use it naturally, don't force it.
### Check for context
At the start, quickly check what exists:
```bash
openspec list --json
```
This tells you:
- If there are active changes
- Their names, schemas, and status
- What the user might be working on
### When no change exists
Think freely. When insights crystallize, you might offer:
- "This feels solid enough to start a change. Want me to create a proposal?"
- Or keep exploring - no pressure to formalize
### When a change exists
If the user mentions a change or you detect one is relevant:
1. **Read existing artifacts for context**
- `openspec/changes/<name>/proposal.md`
- `openspec/changes/<name>/design.md`
- `openspec/changes/<name>/tasks.md`
- etc.
2. **Reference them naturally in conversation**
- "Your design mentions using Redis, but we just realized SQLite fits better..."
- "The proposal scopes this to premium users, but we're now thinking everyone..."
3. **Offer to capture when decisions are made**
| Insight Type | Where to Capture |
|----------------------------|--------------------------------|
| New requirement discovered | `specs/<capability>/spec.md` |
| Requirement changed | `specs/<capability>/spec.md` |
| Design decision made | `design.md` |
| Scope changed | `proposal.md` |
| New work identified | `tasks.md` |
| Assumption invalidated | Relevant artifact |
Example offers:
- "That's a design decision. Capture it in design.md?"
- "This is a new requirement. Add it to specs?"
- "This changes scope. Update the proposal?"
4. **The user decides** - Offer and move on. Don't pressure. Don't auto-capture.
---
## What You Don't Have To Do
- Follow a script
- Ask the same questions every time
- Produce a specific artifact
- Reach a conclusion
- Stay on topic if a tangent is valuable
- Be brief (this is thinking time)
---
## Handling Different Entry Points
**User brings a vague idea:**
```
User: I'm thinking about adding real-time collaboration
You: Real-time collab is a big space. Let me think about this...
COLLABORATION SPECTRUM
════════════════════════════════════════════
Awareness Coordination Sync
│ │ │
▼ ▼ ▼
┌────────┐ ┌────────┐ ┌────────┐
│Presence│ │Cursors │ │ CRDT │
│ "3 │ │ Multi │ │Conflict│
│online" │ │ select │ │ free │
└────────┘ └────────┘ └────────┘
│ │ │
trivial moderate complex
Where's your head at?
```
**User brings a specific problem:**
```
User: The auth system is a mess
You: [reads codebase]
┌─────────────────────────────────────────────┐
│ CURRENT AUTH FLOW │
└─────────────────────────────────────────────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
┌─────────┐ ┌─────────┐ ┌─────────┐
│ Google │ │ GitHub │ │ Email │
│ OAuth │ │ OAuth │ │ Magic │
└────┬────┘ └────┬────┘ └────┬────┘
│ │ │
└─────────────┼─────────────┘
▼
┌───────────┐
│ Session │
└─────┬─────┘
│
▼
┌───────────┐
│ Perms │
└───────────┘
I see three tangles. Which one's burning?
```
**User is stuck mid-implementation:**
```
User: /opsx:explore add-auth-system
The OAuth integration is more complex than expected
You: [reads change artifacts]
You're on task 4: "Implement OAuth flow"
Let me trace what's involved...
[draws diagram, explores options, suggests paths]
Want to update the design to reflect this?
Or add a spike task to investigate?
```
**User wants to compare options:**
```
User: Should we use Postgres or SQLite?
You: Generic answer is boring. What's the context?
User: A CLI tool that tracks local dev environments
You: That changes everything.
┌─────────────────────────────────────────────────┐
│ CLI TOOL DATA STORAGE │
└─────────────────────────────────────────────────┘
Key constraints:
• No daemon running
• Must work offline
• Single user
SQLite Postgres
Deployment embedded ✓ needs server ✗
Offline yes ✓ no ✗
Single file yes ✓ no ✗
SQLite. Not even close.
Unless... is there a sync component?
```
---
## Ending Discovery
There's no required ending. Discovery might:
- **Flow into a proposal**: "Ready to start? I can create a change proposal."
- **Result in artifact updates**: "Updated design.md with these decisions"
- **Just provide clarity**: User has what they need, moves on
- **Continue later**: "We can pick this up anytime"
When it feels like things are crystallizing, you might summarize:
```
## What We Figured Out
**The problem**: [crystallized understanding]
**The approach**: [if one emerged]
**Open questions**: [if any remain]
**Next steps** (if ready):
- Create a change proposal
- Keep exploring: just keep talking
```
But this summary is optional. Sometimes the thinking IS the value.
---
## Guardrails
- **Don't implement** - Never write code or implement features. Creating OpenSpec artifacts is fine, writing application code is not.
- **Don't fake understanding** - If something is unclear, dig deeper
- **Don't rush** - Discovery is thinking time, not task time
- **Don't force structure** - Let patterns emerge naturally
- **Don't auto-capture** - Offer to save insights, don't just do it
- **Do visualize** - A good diagram is worth many paragraphs
- **Do explore the codebase** - Ground discussions in reality
- **Do question assumptions** - Including the user's and your own
+110
View File
@@ -0,0 +1,110 @@
---
name: openspec-propose
description: Propose a new change with all artifacts generated in one step. Use when the user wants to quickly describe what they want to build and get a complete proposal with design, specs, and tasks ready for implementation.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Propose a new change - create the change and generate all artifacts in one step.
I'll create a change with artifacts:
- proposal.md (what & why)
- design.md (how)
- tasks.md (implementation steps)
When ready to implement, run /opsx:apply
---
**Input**: The user's request should include a change name (kebab-case) OR a description of what they want to build.
**Steps**
1. **If no clear input provided, ask what they want to build**
Use the **AskUserQuestion tool** (open-ended, no preset options) to ask:
> "What change do you want to work on? Describe what you want to build or fix."
From their description, derive a kebab-case name (e.g., "add user authentication" → `add-user-auth`).
**IMPORTANT**: Do NOT proceed without understanding what the user wants to build.
2. **Create the change directory**
```bash
openspec new change "<name>"
```
This creates a scaffolded change at `openspec/changes/<name>/` with `.openspec.yaml`.
3. **Get the artifact build order**
```bash
openspec status --change "<name>" --json
```
Parse the JSON to get:
- `applyRequires`: array of artifact IDs needed before implementation (e.g., `["tasks"]`)
- `artifacts`: list of all artifacts with their status and dependencies
4. **Create artifacts in sequence until apply-ready**
Use the **TodoWrite tool** to track progress through the artifacts.
Loop through artifacts in dependency order (artifacts with no pending dependencies first):
a. **For each artifact that is `ready` (dependencies satisfied)**:
- Get instructions:
```bash
openspec instructions <artifact-id> --change "<name>" --json
```
- The instructions JSON includes:
- `context`: Project background (constraints for you - do NOT include in output)
- `rules`: Artifact-specific rules (constraints for you - do NOT include in output)
- `template`: The structure to use for your output file
- `instruction`: Schema-specific guidance for this artifact type
- `outputPath`: Where to write the artifact
- `dependencies`: Completed artifacts to read for context
- Read any completed dependency files for context
- Create the artifact file using `template` as the structure
- Apply `context` and `rules` as constraints - but do NOT copy them into the file
- Show brief progress: "Created <artifact-id>"
b. **Continue until all `applyRequires` artifacts are complete**
- After creating each artifact, re-run `openspec status --change "<name>" --json`
- Check if every artifact ID in `applyRequires` has `status: "done"` in the artifacts array
- Stop when all `applyRequires` artifacts are done
c. **If an artifact requires user input** (unclear context):
- Use **AskUserQuestion tool** to clarify
- Then continue with creation
5. **Show final status**
```bash
openspec status --change "<name>"
```
**Output**
After completing all artifacts, summarize:
- Change name and location
- List of artifacts created with brief descriptions
- What's ready: "All artifacts created! Ready for implementation."
- Prompt: "Run `/opsx:apply` or ask me to implement to start working on the tasks."
**Artifact Creation Guidelines**
- Follow the `instruction` field from `openspec instructions` for each artifact type
- The schema defines what each artifact should contain - follow it
- Read dependency artifacts for context before creating new ones
- Use `template` as the structure for your output file - fill in its sections
- **IMPORTANT**: `context` and `rules` are constraints for YOU, not content for the file
- Do NOT copy `<context>`, `<rules>`, `<project_context>` blocks into the artifact
- These guide what you write, but should never appear in the output
**Guardrails**
- Create ALL artifacts needed for implementation (as defined by schema's `apply.requires`)
- Always read dependency artifacts before creating a new one
- If context is critically unclear, ask the user - but prefer making reasonable decisions to keep momentum
- If a change with that name already exists, ask if user wants to continue it or create a new one
- Verify each artifact file exists after writing before proceeding to next
@@ -0,0 +1,233 @@
# AI Ops Prompt 配置化 & LookupKnowledgeTool 集成
**日期**: 2026-06-24
**类型**: 功能增强 + 架构优化
**影响范围**: AI Ops 服务
---
## 一、变更背景
### 1.1 问题
- **硬编码 Prompt**:Planner、Executor、Supervisor 的系统提示词硬编码在 `AiOpsService.java` 中,难以维护和版本控制
- **缺少知识库精确检索**:现有 `InternalDocsTools` 只支持 L1 语义检索(200-500ms),对于错误码、配置项等精确关键词查询效率较低
### 1.2 解决方案
1. **Prompt 配置化**:将所有 Agent 的 Prompt 抽取到 `prompts/ai-ops-prompts.yml` 配置文件
2. **集成 L0+L1 混合检索**:引入 `LookupKnowledgeTool`,支持精确关键词匹配(< 10ms)+ 语义检索补充
---
## 二、架构变更
### 2.1 Prompt 配置化架构
```
AiOpsService
↓ 注入
AiOpsPromptProperties (配置类)
↓ @PostConstruct 加载
ClassPathResource 读取 Markdown 文件
↓ 读取
prompts/
├── planner-prompt.md
├── executor-prompt.md
└── supervisor-prompt.md
```
**优点**:
- 易于维护:Prompt 修改不需要重新编译
- 格式友好:Markdown 格式支持代码块、表格,无 YAML 转义问题
- 版本控制:配置文件独立管理
- 易于扩展:后续可按环境区分(dev/prod)
### 2.2 工具层增强
```
原有工具:
- queryInternalDocs (纯 L1 语义检索,200-500ms)
新增工具:
- lookup_knowledge (L0 精确匹配 + L1 补充,< 10ms 高置信度)
```
**使用策略**:
- 精确关键词(错误码、配置项)→ `lookup_knowledge`,未找到时降级到 `queryInternalDocs`
- 模糊概念、故障流程 → 直接使用 `queryInternalDocs`
---
## 三、核心改动
### 3.1 新增文件
#### `AiOpsPromptProperties.java`
```java
@Configuration
public class AiOpsPromptProperties {
private String planner;
private String executor;
private String supervisor;
@PostConstruct
public void loadPrompts() {
planner = loadPromptFromFile("prompts/planner-prompt.md");
executor = loadPromptFromFile("prompts/executor-prompt.md");
supervisor = loadPromptFromFile("prompts/supervisor-prompt.md");
}
private String loadPromptFromFile(String path) throws IOException {
ClassPathResource resource = new ClassPathResource(path);
return new String(resource.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
}
}
```
#### `prompts/*.md`
三个独立的 Markdown 文件,包含 Agent 的完整系统提示词:
- `planner-prompt.md` - Planner Agent 系统提示词
- `executor-prompt.md` - Executor Agent 系统提示词(含工具选择指南)
- `supervisor-prompt.md` - Supervisor Agent 系统提示词
### 3.2 修改文件
#### `AiOpsService.java`
**注入新组件**:
```java
@Autowired
private LookupKnowledgeTool lookupKnowledgeTool;
@Autowired
private AiOpsPromptProperties promptProperties;
```
**使用配置化 Prompt**:
```java
// 原来
.systemPrompt(buildPlannerPrompt())
// 改为
.systemPrompt(promptProperties.getPlanner())
```
**添加工具到工具数组**:
```java
return new Object[]{
dateTimeTools,
internalDocsTools,
queryMetricsTools,
lookupKnowledgeTool // 新增
};
```
**删除方法**:
- `buildPlannerPrompt()`
- `buildExecutorPrompt()`
- `buildSupervisorSystemPrompt()`
---
## 四、Executor Prompt 变更详情
### 4.1 新增工具选择指南
```yaml
- 根据查询内容选择合适的工具:
* 精确关键词(错误码、配置项名称)→ 优先使用 lookup_knowledge,未找到时降级到 queryInternalDocs
* 模糊概念、故障流程 → 直接使用 queryInternalDocs
* 告警数据 → queryPrometheusAlerts
* 日志数据 → queryLogs
```
### 4.2 降级策略
关键改进:明确了 `lookup_knowledge` 未找到时的降级策略。
**流程**:
```
1. Planner: "查询 ERR_TIMEOUT 定义"
2. Executor: 调用 lookup_knowledge("ERR_TIMEOUT")
3a. 如果 found=true, confidence=high → 使用 primary.content
3b. 如果 found=false → 自动降级到 queryInternalDocs("ERR_TIMEOUT 超时错误")
4. 返回 feedback 给 Planner
```
---
## 五、兼容性说明
### 5.1 向后兼容
✅ **完全兼容**:
- 现有工具调用逻辑不变
- 3-Agent 协同模式不变
- Planner/Executor/Supervisor 的职责边界不变
### 5.2 新增依赖
- `LookupKnowledgeTool` 依赖 `KnowledgeIndexService` 和 `VectorSearchService`
- 需要 `knowledge_base/` 目录存在(已在 `application.yml` 中配置)
---
## 六、验证清单
### 6.1 编译验证
```bash
mvn clean compile -DskipTests
```
✅ **结果**: BUILD SUCCESS
### 6.2 运行时验证(待完成)
- [ ] 启动应用,验证 Prompt 配置加载成功
- [ ] 触发 AI Ops 流程,验证 `lookup_knowledge` 工具可调用
- [ ] 测试精确关键词查询(如 "ERR_TIMEOUT")
- [ ] 测试降级策略(查询不存在的关键词)
---
## 七、后续工作
### 7.1 知识库内容准备
当前 `knowledge_base/` 目录需要补充文档:
- 错误码定义(支付网关、订单系统等)
- 配置最佳实践(Redis、HikariCP、Flyway 等)
- 故障排查流程
**文档格式示例**:
```markdown
---
title: 支付网关错误码定义
keywords: [ERR_TIMEOUT, 超时, 支付网关]
summary: 记录了支付网关所有核心错误码的含义及排查方向
category: api
---
# 支付网关错误码定义
## ERR_TIMEOUT
...
```
### 7.2 Prompt 优化
基于实际运行反馈,持续优化 `prompts/ai-ops-prompts.yml` 中的提示词。
### 7.3 可观测性增强
- 监控 `lookup_knowledge` 的调用频率和命中率
- 记录降级场景(L0 未找到 → L1 补充)
---
## 八、参考文档
- [知识库检索架构说明](../mvp/architecture/knowledge-retrieval-architecture.md)
- [AI Ops 核心设计 Essence 报告](../docs/learning/01-AI-Ops-核心设计-Essence报告.md)
+100
View File
@@ -0,0 +1,100 @@
# Prompt 配置化改进总结
**日期**: 2026-06-24
**改进**: 从 YAML 配置改为 Markdown 文件
---
## 改进原因
YAML 格式存在以下问题:
1. **多行字符串缩进敏感**:容易出现格式错误
2. **转义字符复杂**:代码块、表格需要转义处理
3. **可读性差**:长文本在 YAML 中难以阅读和维护
Markdown 格式优势:
- ✅ 原生支持代码块、表格、列表
- ✅ 无需转义,所见即所得
- ✅ 版本控制 diff 更清晰
- ✅ 编辑器语法高亮支持好
---
## 最终方案
### 文件结构
```
src/main/resources/prompts/
├── planner-prompt.md # Planner Agent 系统提示词
├── executor-prompt.md # Executor Agent 系统提示词
└── supervisor-prompt.md # Supervisor Agent 系统提示词
```
### 加载方式
```java
@Configuration
public class AiOpsPromptProperties {
@PostConstruct
public void loadPrompts() {
planner = loadPromptFromFile("prompts/planner-prompt.md");
executor = loadPromptFromFile("prompts/executor-prompt.md");
supervisor = loadPromptFromFile("prompts/supervisor-prompt.md");
}
private String loadPromptFromFile(String path) throws IOException {
ClassPathResource resource = new ClassPathResource(path);
return new String(resource.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
}
}
```
### 使用方式
```java
@Autowired
private AiOpsPromptProperties promptProperties;
// 直接使用
.systemPrompt(promptProperties.getPlanner())
```
---
## 编译验证
```bash
mvn clean compile -DskipTests
```
✅ **结果**: BUILD SUCCESS
---
## 完整改动清单
| 文件 | 改动 |
|------|------|
| `AiOpsService.java` | 注入 `LookupKnowledgeTool` + `AiOpsPromptProperties` |
| `AiOpsPromptProperties.java` | 从 Markdown 文件加载 Prompt(使用 `@PostConstruct`)|
| `prompts/planner-prompt.md` | 新增:Planner 系统提示词 |
| `prompts/executor-prompt.md` | 新增:Executor 系统提示词(含工具选择指南)|
| `prompts/supervisor-prompt.md` | 新增:Supervisor 系统提示词 |
| ~~`YamlPropertySourceFactory.java`~~ | 已删除(不再需要)|
| ~~`prompts/ai-ops-prompts.yml`~~ | 已删除(改用 Markdown)|
---
## Executor Prompt 关键改进
新增工具选择指南:
```markdown
- 根据查询内容选择合适的工具:
* 精确关键词(错误码、配置项名称)→ 优先使用 lookup_knowledge,未找到时降级到 queryInternalDocs
* 模糊概念、故障流程 → 直接使用 queryInternalDocs
* 告警数据 → queryPrometheusAlerts
* 日志数据 → queryLogs
```
降级策略:
- `lookup_knowledge` 未找到 → 自动降级到 `queryInternalDocs`
- 确保查询不会因为知识库缺少内容而失败
+469
View File
@@ -0,0 +1,469 @@
# 知识库初始化 API 使用文档
## 概述
提供了知识库批量初始化接口,用于将 `knowledge_base` 目录下的所有 Markdown 文档导入到数据库和向量索引(L0 + L1)。
**功能特点**:
1. ✅ **批量扫描**:递归扫描 knowledge_base 目录下所有 .md 文件
2. ✅ **自动去重**:基于文件路径检查,避免重复导入
3. ✅ **数据入库**:保存文档元数据到 MySQL
4. ✅ **L0 索引**:自动加入内存精确匹配索引
5. ✅ **L1 索引**:文档分块并上传到 Milvus 向量数据库
---
## API 接口
### 1. 初始化知识库
**端点**:
```
POST /api/knowledge/init?force=false
```
**参数**:
- `force`(可选):是否强制重新导入,跳过去重检查
- `false`(默认):跳过已存在的文档
- `true`:强制重新导入所有文档
**请求示例**:
```bash
# 首次导入(去重模式)
curl -X POST http://localhost:9900/api/knowledge/init
# 强制重新导入
curl -X POST http://localhost:9900/api/knowledge/init?force=true
```
**响应示例**:
```json
{
"success": true,
"message": "知识库初始化完成",
"scanned": 6,
"skipped": 0,
"inserted": 6,
"failed": 0,
"details": {
"api/payment-errors.md": "导入成功(L0+L1)",
"domain/spring-ai-tool-best-practices.md": "导入成功(L0+L1)",
"infrastructure/flyway-best-practices.md": "导入成功(L0+L1)",
"infrastructure/mysql-connection-pool.md": "导入成功(L0+L1)",
"infrastructure/redis-config.md": "导入成功(L0+L1)",
"troubleshooting/fault-diagnosis-process.md": "导入成功(L0+L1)"
}
}
```
**字段说明**:
- `scanned`:扫描到的文件总数
- `skipped`:跳过的文件数量(已存在)
- `inserted`:成功导入的文件数量
- `failed`:失败的文件数量
- `details`:每个文件的处理结果详情
---
### 2. 查询知识库统计
**端点**:
```
GET /api/knowledge/stats
```
**请求示例**:
```bash
curl http://localhost:9900/api/knowledge/stats
```
**响应示例**:
```json
{
"success": true,
"totalDocuments": 6,
"totalVectors": 48,
"categories": {
"api": 1,
"domain": 1,
"infrastructure": 3,
"troubleshooting": 1
}
}
```
**字段说明**:
- `totalDocuments`:数据库中的文档总数
- `totalVectors`:Milvus 中的向量总数(chunk 数量)
- `categories`:按分类统计的文档数量
---
## 使用场景
### 场景 1:项目启动时初始化
```bash
# 1. 启动应用
mvn spring-boot:run
# 2. 等待应用启动完成(约 10 秒)
# 3. 调用初始化接口
curl -X POST http://localhost:9900/api/knowledge/init
# 4. 查看结果
# 日志输出:知识库初始化完成: 扫描=6, 跳过=0, 新增=6, 失败=0
```
---
### 场景 2:添加新文档后重新初始化
```bash
# 1. 添加新文档到 knowledge_base 目录
echo "---
title: 新文档
keywords: [测试, test]
summary: 这是一个测试文档
category: test
---
# 新文档内容
" > knowledge_base/test/new-doc.md
# 2. 调用初始化接口(去重模式)
curl -X POST http://localhost:9900/api/knowledge/init
# 3. 查看结果
# 只会导入新文档,跳过已存在的 6 个文档
# 响应: scanned=7, skipped=6, inserted=1, failed=0
```
---
### 场景 3:强制重新导入所有文档
```bash
# 适用场景:
# - 数据库被清空,需要重新导入
# - 文档内容有更新,需要刷新
# - 索引损坏,需要重建
curl -X POST http://localhost:9900/api/knowledge/init?force=true
# 响应: scanned=6, skipped=0, inserted=6, failed=0
```
---
## 去重机制
### 去重依据
- **文件路径**:相对于 `knowledge_base` 目录的相对路径
- 示例:`api/payment-errors.md`
### 去重逻辑
```
if (!force && existingFilePaths.contains(relativePath)) {
跳过该文档
} else {
导入该文档
}
```
### 注意事项
1. **文件移动会被视为新文档**:
```bash
# 移动前:api/payment-errors.md
# 移动后:errors/payment-errors.md
# 结果:会被当作两个不同的文档
```
2. **文件重命名会被视为新文档**:
```bash
# 重命名前:payment-errors.md
# 重命名后:payment-error-codes.md
# 结果:会被当作两个不同的文档
```
3. **内容更新不触发重新导入**(非 force 模式):
```bash
# 修改文件内容后调用 init(非 force)
# 结果:跳过该文档,数据库中仍是旧内容
# 解决:使用 force=true 强制重新导入
```
---
## 数据存储
### 完整的数据流
```
knowledge_base/*.md
↓ 1. 扫描
KnowledgeBaseInitService
↓ 2. 解析 frontmatter
Frontmatter (title, keywords, summary)
↓ 3. 保存到数据库
MySQL (api_document)
↓ 4. 提取正文 & 分块
DocumentChunkService
↓ 5. 生成向量
VectorEmbeddingService
↓ 6. 索引到 Milvus
Milvus (L1 向量索引)
↓ 7. 加入内存索引
KnowledgeIndexService (L0)
```
---
### 数据库表结构(api_document)
| 字段 | 类型 | 说明 | 示例 |
|------|------|------|------|
| `id` | BIGINT | 主键 | 1 |
| `doc_id` | VARCHAR(64) | 文档唯一标识 | uuid |
| `file_name` | VARCHAR(256) | 文件名 | payment-errors.md |
| `file_path` | VARCHAR(512) | 相对路径 | api/payment-errors.md |
| `api_name` | VARCHAR(128) | 文档标题 | 支付网关错误码定义 |
| `status` | VARCHAR(16) | 状态 | INDEXED / FAILED |
| `chunk_count` | INT | 分块数量 | 8 |
| `error_message` | TEXT | 错误信息 | null |
| `metadata` | TEXT | Frontmatter JSON | {"title":"...","keywords":[...]} |
| `file_size` | BIGINT | 文件大小(字节) | 2048 |
| `indexed_at` | DATETIME | 索引时间 | 2026-06-25 10:00:00 |
### metadata JSON 结构
```json
{
"title": "支付网关错误码定义",
"summary": "记录了支付网关所有核心错误码的含义及排查方向",
"category": "api",
"keywords": ["ERR_TIMEOUT","超时","支付网关"]
}
```
---
### Milvus 向量索引
每个文档会被分块(chunk)并生成向量,存储到 Milvus 集合中:
**Collection**: `knowledge_base_collection`
**字段**:
- `doc_id`:文档 ID
- `chunk_id`:分块 ID
- `chunk_text`:分块文本内容
- `embedding`:768 维向量
- `category`:文档分类
- `file_path`:文件路径
**分块策略**:
- Chunk Size:根据 `DocumentChunkConfig` 配置(默认 500 token)
- Overlap:重叠区域(默认 50 token)
---
## L0 内存索引
导入过程会自动将文档加入 `KnowledgeIndexService` 的内存索引:
```java
KnowledgeEntry entry = KnowledgeEntry.builder()
.filePath(relativePath)
.title(title)
.keywords(keywords)
.summary(summary)
.category(category)
.build();
knowledgeIndexService.addToIndex(entry);
```
**验证 L0 索引**:
```bash
# 应用启动后查看日志
grep "知识库索引加载完成" logs/application.log
# 输出示例:
# [INFO] 知识库索引加载完成,共 6 个文档
```
---
## 错误处理
### 常见错误
#### 1. 目录不存在
```json
{
"success": false,
"message": "初始化失败: 知识库目录不存在: knowledge_base"
}
```
**解决**:
```bash
mkdir -p knowledge_base/api
mkdir -p knowledge_base/infrastructure
mkdir -p knowledge_base/domain
mkdir -p knowledge_base/troubleshooting
```
---
#### 2. 文档格式无效
```json
{
"success": true,
"scanned": 6,
"inserted": 5,
"failed": 1,
"details": {
"test/invalid.md": "格式无效: frontmatter 解析失败"
}
}
```
**原因**:
- 缺少 frontmatter
- YAML 格式错误
- 缺少必填字段(title, keywords, summary)
**解决**:
```markdown
---
title: 文档标题
keywords: [关键词1, 关键词2]
summary: 文档摘要
category: api
---
# 正文内容
```
---
### 问题 4: Milvus 连接失败
**症状**:
```json
{
"success": true,
"scanned": 6,
"inserted": 0,
"failed": 6,
"details": {
"api/payment-errors.md": "Milvus 索引失败: Connection refused"
}
}
```
**原因**:
- Milvus 服务未启动
- 网络连接问题
- 配置错误
**解决**:
```bash
# 检查 Milvus 是否运行
docker ps | grep milvus
# 检查配置
grep milvus application.yml
# 启动 Milvus
docker-compose up -d milvus-standalone
```
---
### 问题 5: 文档分块失败
**症状**:
```json
{
"details": {
"test/large-doc.md": "Milvus 索引失败: Document too large"
}
}
```
**原因**:
- 文档内容过大
- 分块配置不当
**解决**:
- 检查 `DocumentChunkConfig` 配置
- 调整 chunk size 和 overlap
---
#### 3. 文档缺少标题
```json
{
"details": {
"test/no-title.md": "缺少标题"
}
}
```
**解决**:在 frontmatter 中添加 `title` 字段。
---
## 最佳实践
### ✅ 推荐做法
1. **首次启动后立即初始化**:
```bash
mvn spring-boot:run
sleep 15 # 等待启动完成
curl -X POST http://localhost:9900/api/knowledge/init
```
2. **新增文档后增量导入**:
```bash
# 不使用 force,只导入新文档
curl -X POST http://localhost:9900/api/knowledge/init
```
3. **定期检查统计信息**:
```bash
curl http://localhost:9900/api/knowledge/stats
```
4. **更新文档内容后强制刷新**:
```bash
curl -X POST http://localhost:9900/api/knowledge/init?force=true
```
---
### ❌ 避免做法
1. **不检查响应就认为成功**:
- 始终检查 `failed` 字段
- 查看 `details` 了解具体失败原因
2. **频繁使用 force=true**:
- 会重复插入数据(违反唯一约束)
- 建议先清理数据库,再使用 force
3. **不检查文档格式就导入**:
- 先手动验证 frontmatter 格式
- 确保必填字段完整
---
## 相关文档
- **知识库使用指南**:`mvp/architecture/knowledge-retrieval-usage.md`
- **知识库架构**:`mvp/architecture/knowledge-retrieval-architecture.md`
- **Executor Prompt**:`src/main/resources/prompts/executor-prompt.md`
+9
View File
@@ -55,3 +55,12 @@ uploads/
/volumes
/server.pid
.claude/settings.local.json
.opencode/plugins/emdash-notifications.js
### Windows / Runtime Artifacts
*.stackdump
NUL
### MVP Demo Generated Outputs
mvp/demo/output/*.json
!mvp/demo/output/README.md
+1 -1
View File
@@ -1,7 +1,7 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **SuperBizAgent-java** (1528 symbols, 2828 relationships, 87 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+1 -1
View File
@@ -115,7 +115,7 @@ trailing off into the following information in 99% of cases:
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **SuperBizAgent-java** (1001 symbols, 2043 relationships, 78 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
View File
-29
View File
@@ -1,29 +0,0 @@
Stack trace:
Frame Function Args
0007FFFFB920 00021005FE8E (000210285F68, 00021026AB6E, 000000000000, 0007FFFFA820) msys-2.0.dll+0x1FE8E
0007FFFFB920 0002100467F9 (000000000000, 000000000000, 000000000000, 0007FFFFBBF8) msys-2.0.dll+0x67F9
0007FFFFB920 000210046832 (000210286019, 0007FFFFB7D8, 000000000000, 000000000000) msys-2.0.dll+0x6832
0007FFFFB920 000210068CF6 (000000000000, 000000000000, 000000000000, 000000000000) msys-2.0.dll+0x28CF6
0007FFFFB920 000210068E24 (0007FFFFB930, 000000000000, 000000000000, 000000000000) msys-2.0.dll+0x28E24
0007FFFFBC00 00021006A225 (0007FFFFB930, 000000000000, 000000000000, 000000000000) msys-2.0.dll+0x2A225
End of stack trace
Loaded modules:
000100400000 bash.exe
7FF9B93D0000 ntdll.dll
7FF9B79A0000 KERNEL32.DLL
7FF9B6860000 KERNELBASE.dll
7FF9B8740000 USER32.dll
7FF9B6830000 win32u.dll
7FF9B84F0000 GDI32.dll
7FF9B6CD0000 gdi32full.dll
7FF9B6790000 msvcp_win.dll
7FF9B7000000 ucrtbase.dll
000210040000 msys-2.0.dll
7FF9B7370000 advapi32.dll
7FF9B8E40000 msvcrt.dll
7FF9B85B0000 sechost.dll
7FF9B6FD0000 bcrypt.dll
7FF9B90F0000 RPCRT4.dll
7FF9B5F20000 CRYPTBASE.DLL
7FF9B6710000 bcryptPrimitives.dll
7FF9B86E0000 IMM32.DLL
+13
View File
@@ -4,7 +4,20 @@
| 日期 | slug | 领域 | 关键词 | 状态 |
|---|---|---|---|---|
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
@@ -0,0 +1,28 @@
# 验收记录
## 验证情况
### 静态验证
- [x] 编译通过(`mvn compile`)
- [x] 42 个测试全部通过(DocumentChunkService / LookupKnowledgeTool / Repository)
- [x] 三张新表通过 Flyway 成功创建
### 脚本验证
- [x] `/api/chat` — 单 Agent 正常响应,agent_step 记录正确
- [x] `/api/chat` — 复杂问题路由到多 Agent(Planner + Executor)
- [x] `/api/ai_ops` — 多 Agent 流程正常,planner 步骤写入 agent_step
- [x] Tool_invocation L0/L1 检索质量明细正确
- [x] diagnosis_session 汇总指标(total_token_count / step_count / tool_call_count)正确
- [x] TokenTrackingChatModel 捕获实际 token 数(已验证 total=827)
- [x] 旧 diagnosis_record 表删除成功
### 未验证
- `/api/chat_stream`(SSE 流式)— 未接入 session 存储,不在本次范围,后续覆盖
- `self_evaluation` / `feedback` — 无前端交互入口
## 剩余风险
| 风险 | 说明 |
|------|------|
| Token 累加 | 当前每步独立记录,汇总在 `backfillSessionMetrics`,未在 Hook 层累加 |
| Async 优化 | 同步写 DB 在低并发下无问题,后续可引入 @Async |
@@ -0,0 +1,21 @@
# 会话存储体系
## 背景
当前 `diagnosis_record` 单表字段耦合在"告警分析"领域,无法支撑通用会话存储。缺少 Agent 决策链维度、检索质量明细、Token 消耗等可观测指标。
## 目标
将单表拆分为三表体系,覆盖 ChatService 和 AiOpsService 两个 Agent 的完整决策链记录,支撑可观测和评估。
## 范围
- 新建 3 张表(diagnosis_session / agent_step / tool_invocation)
- Flyway 迁移 + JPA Entity + Repository
- 改造 AgentLoggingHook 持久化 agent_step
- 改造 LookupKnowledgeTool 写入 tool_invocation
- ChatService / AiOpsService 支持 diagnosis_session 生命周期
- Token 用量追踪(TokenTrackingChatModel)
- 意图识别路由(单 Agent / 多 Agent)
- 删除旧 diagnosis_record 表
## 非目标
- 不涉及 UI 层面的会话展示
- 不涉及历史数据迁移
@@ -0,0 +1,22 @@
# 会话存储 — 决策记录
## 关键决策
| 决策 | 选择 | 理由 |
|------|------|------|
| AgentLoggingHook 创建方式 | POJO(构造注入),非 @Component | 需为 ChatService/AiOpsService 创建多个实例(不同 agentName) |
| AiOpsService 记录粒度 | 只记子 Agent(Planner/Executor),不记 Supervisor | Supervisor 编排日志已有体现,单独记录增加噪音 |
| sessionId 传递 | RunnableConfig.metadata(优先)+ ThreadLocal(兜底) | RunnableConfig 线程安全,异步兼容 |
| Tool 获取 sessionId | SessionContextHolder(ThreadLocal) | Tool 不在调用链中,无法通过 RunnableConfig 获取 |
| Token 追踪 | TokenTrackingChatModel 包装器拦截 ChatModel.call() | 框架 _TOKEN_USAGE_ 仅 stream 路径可用 |
| Chat 复杂度路由 | 关键词 + 长度判断 | MVP 简化实现 |
| 多 Agent Planner 无工具 | 不注入 methodTools/tools | 防止 Planner 自己执行,强制通过 Executor 执行 |
| 旧表处理 | V007 Flyway 迁移删除 diagnosis_record | 被三表替代,不再使用 |
## 风险
| 风险 | 等级 | 说明 |
|------|:----:|------|
| Hook 同步写 DB | 低 | MVP 阶段数据量小,后续可异步化 |
| token_count 依赖 ChatResponse.usage | 低 | DeepSeek 已确认返回实际用量 |
| stream 路径 session 记录 | 低 | 当前 call 路径正常,stream 需确认 RunnableConfig 传播 |
@@ -0,0 +1,23 @@
# 证据记录
## Evidence-Driven 查证
### E1: AgentLoggingHook 创建方式
- **发现**: ChatService 通过 `new AgentLoggingHook()` 创建,非 Spring 管理,无法注入 Repository
- **结论**: 需要改造为可注入的 POJO(构造注入)
- **影响**: Hook 重构为构造注入 Repository + agentName
### E2: AiOpsService 未使用 Hook
- **发现**: AiOpsService 的 Planner / Executor / Supervisor 均未配置 AgentLoggingHook
- **结论**: 需要补齐,每个子 Agent 加 Hook
- **影响**: Planner 和 Executor 各加 Hook,Supervisor 不加
### E3: 项目无异步基础设施
- **发现**: 全局搜索 `@Async` / `@EnableAsync` 均无匹配
- **结论**: MVP 阶段同步写 DB,后续优化
- **影响**: 标记为技术债
### E4: RunnableConfig 支持 metadata
- **发现**: `RunnableConfig` 的 `metadata` 为 `ConcurrentMap`,可在构建时设置
- **结论**: sessionId 通过 `config.addMetadata("sessionId", id)` 传递,线程安全
- **影响**: 取代 ThreadLocal 方案
@@ -0,0 +1,64 @@
# acceptance.md — confidence-feedback
## 实现清单
| 任务 | 文件 | 状态 |
|---|---|---|
| T0:Flyway V008 + answer 字段 | `V008__add_answer_to_diagnosis_session.sql`、`DiagnosisSession.java` | 完成 |
| T1:EvaluationService(规则引擎) | `EvaluationService.java` | 完成 |
| T2:ChatService 后置调用 | `ChatService.java` | 完成 |
| T3:FeedbackController + FeedbackService | `FeedbackController.java`、`FeedbackService.java`、`FeedbackRequest.java`、`FeedbackResponse.java` | 完成 |
| T4:CaseLibraryService | `CaseLibraryService.java` | 完成 |
| T5:AsyncConfig | `AsyncConfig.java` | 完成 |
## 验证记录
### 静态验证(已通过)
- `mvn compile` BUILD SUCCESS(2026-06-30)
- 无新增 ERROR,存量 WARNING 与本次改动无关
- import 完整性人工检查通过
### 脚本验证(已通过,2026-06-30)
验证工具:`scripts/query_mysql.py`(本次新建)
| 步骤 | 操作 | 结果 |
|---|---|---|
| 1 | POST /api/chat 发送问题 | 200,answer 有值 |
| 2 | 等 5 秒查 diagnosis_session | self_evaluation 写入规则引擎结果,answer 写入完整回答 |
| 3 | POST /api/feedback useful | 200,返回 caseId;case_library 新增一行,feedback=useful,status=SUCCESS |
| 4 | POST /api/feedback not_useful | 200,feedback=not_useful,status 仍为 SUCCESS(未被改写) |
| 5(边界)| 重复提交 useful | 返回同一 caseId,case_library 无重复插入 |
| 6(边界)| 非法 feedback 值 | HTTP 400 |
### Flyway V008 迁移
- 服务启动后 diagnosis_session 表存在 answer 列,验证通过(步骤 2 能写入 answer)
### 浏览器/人工验证(已通过,2026-06-30)
| 步骤 | 操作 | 结果 |
|---|---|---|
| 1 | 发送"今天天气怎么样" | AI 回复下方出现"有用/无用"按钮 |
| 2 | 点击"有用" | 按钮区域替换为"已标记为有用" |
| 3 | 网络请求确认 | POST /api/feedback 返回 HTTP 200,`success: true` |
### 前端反馈按钮(追加,2026-06-30)
**改动文件**:`app.js`、`styles.css`
关键设计:
- `ChatResult` record 新增(`ChatService`),`ChatResponse` 增加 `sessionId` 字段(`ChatController`)
- `sendQuickMessage` 读取 `chatResponse.sessionId` 存为 `this.lastSessionId`
- `createFeedbackBar(sessionId)` 闭包绑定 sessionId,避免多轮对话时 sessionId 错位
- `submitFeedback(feedback, barElement, sessionId)` 直接用传入参数,不依赖全局状态
- 流式模式(`/api/chat_stream`)反馈按钮会渲染,但 sessionId 为空,点击不生效(已知限制)
## 已知限制
- 非检索工具(DateTimeTools 等)不写 tool_invocation,evidence_score = 0(已接受,符合"证据充分度"定义)
- `@Async` 失败时 selfEvaluation 为 null,前端需处理 null(已接受)
- CaseLibrary 的 faultCategory 固定为 GENERAL,需人工补充(已接受,Phase 2 优化)
- LLM 观点层未实现,selfEvaluation JSON 预留 llm_opinion 扩展位(Phase 2)
- 流式模式反馈按钮 sessionId 缺失,暂不处理(已知,后续处理流式接口时一并解决)
@@ -0,0 +1,37 @@
# brief.md — confidence-feedback
## 背景
DiagnosisSession 已预留 `selfEvaluation`(JSON)和 `feedback`(VARCHAR 16)两个字段,但完全为空。Agent 完成对话后不计算证据评分,也没有接收用户反馈的 API,无法支撑报告质量评估和 BadCase 追踪。
## 目标
1. 给每次对话结果自动打一个基于事实的证据充分度评分(evidence_score)
2. 提供用户反馈 API(useful/not_useful),useful 触发案例自动沉淀,not_useful 标记 BadCase
## 范围
- `DiagnosisSession` 加 `answer` 字段(Flyway V008)
- `EvaluationService`:基于 tool_invocation 的规则引擎,@Async 写 selfEvaluation
- `FeedbackController` + `FeedbackService`:POST /api/feedback
- `CaseLibraryService.createFromSession`:幂等案例沉淀
- `AsyncConfig`:@EnableAsync
- `ChatService`:SUCCESS 分支写 answer + 触发 evaluate;新增 `ChatResult` record 回传 sessionId
- `ChatController.ChatResponse` 增加 `sessionId` 字段
- 前端 `app.js`:AI 回复下方反馈按钮,点击调用 `/api/feedback`,闭包绑定 sessionId
- 前端 `styles.css`:反馈栏样式
## 非目标
- 不实现 Verifier Agent 完整链路
- 不实现 LLM 自评(预留扩展位,Phase 2 再做)
- 不实现案例结构化字段自动填充(faultCategory 等暂时填 GENERAL)
- 不实现 BadCase 自动分析或 Prompt 优化
## 分档
standard
## 关联 OpenSpec
`openspec/changes/confidence-feedback/`
@@ -0,0 +1,115 @@
# decisions.md — confidence-feedback
## Question Pool(grill 阶段)
| # | 问题 | 模式 | 状态 |
|---|---|---|---|
| Q1 | 置信度由谁计算 | user-interview | 已确认 |
| Q2 | 反馈触发哪些后端操作 | user-interview | 已确认 |
| Q3 | CaseLibrary 结构化字段从哪里填 | evidence-driven | 已确认(方案变更) |
| Q4 | 验收口径 | user-interview | 已确认 |
---
## Evidence-Driven 结论
### Q3:CaseLibrary 内容来源
**初始结论**:从 `agent_step.thought` 提取(grill 阶段)
**修正(apply 阶段讨论后)**:
- 代码证据:`agent_step.thought` 截断为 2000 字符,`modelOutput` 截断为 500 字符,均不是完整答案
- `ChatService.executeChat` 第 269 行已有完整答案 `answer = response.getText()`,但未持久化
- 决策:给 `DiagnosisSession` 加 `answer TEXT` 字段,Flyway V008 迁移,案例内容直接从 `session.answer` 取
---
## User-Interview 确认记录
### Q1 — 置信度由谁评估
- 用户原话(grill):"两者都要:规则兜底 + Verifier 主打分"
- **apply 后修正**:讨论后决定去掉 LLM 自评,仅用规则引擎(见"apply 阶段决策")
- 最终实现:`EvaluationService` 纯规则,预留 `llm_opinion` 扩展位
### Q2 — 反馈触发操作
- 用户原话:"写入 DiagnosisSession.feedback 字段, not_useful → 打 BAD_CASE 标记"
- **apply 后修正**:BAD_CASE 不改 status,feedback 字段本身即为标记(见"apply 阶段决策")
- 最终实现:`FeedbackService` 只写 feedback + 可选写 case_library,不改 status
### Q4 — 验收口径
- 用户原话:"端到端可验证:发一次 chat → 查 DB 看 selfEvaluation 有值 → 提交 feedback → 查 DB 看 feedback + case_library"
- 确认状态:已确认,未变化
---
## Apply 阶段决策(post-grill 重要变更)
### 决策 A:DiagnosisSession 加 answer 字段
- **问题**:案例沉淀需要完整答案,agent_step.thought 被截断,不可用
- **决策**:新增 `answer LONGTEXT` 字段,ChatService SUCCESS 分支写入
- **影响**:V008 Flyway 迁移,CaseLibraryService 直接读 session.answer
### 决策 B:去掉 LLM 自评,只用规则引擎
- **问题**:LLM 评估自己的答案系统性偏高分;多一次调用消耗 token;Verifier Agent 当前未实现
- **决策**:MVP 阶段仅用基于 tool_invocation 的规则引擎
- **理由**:规则可解释、可复现、不撒谎;Verifier 留待诊断全链路实现时再做
- **预留**:`selfEvaluation` JSON 结构保留 `llm_opinion` 扩展位,代码底部注释说明接入点
### 决策 C:BAD_CASE 不改 status 字段
- **问题**:status 是执行状态语义(RUNNING/SUCCESS/FAILED),BAD_CASE 是质量标签,两个维度不同;覆盖 status 会破坏统计
- **决策**:`not_useful` 通过 `feedback` 字段本身标识,查 BadCase 用 `WHERE feedback = 'not_useful'`
### 决策 D:评分字段重命名为 evidence_score
- **问题**:原名 confidence 容易误解为"答案准确性",实际衡量的是"证据收集充分度"
- **决策**:重命名为 `evidence_score`,明确语义边界
- **边界说明**:工具调用能证明 Agent 有尝试收集证据,但无法证明答案无幻觉;这个分数过滤最差情况(无工具调用就给答案),不能识别"调用了工具但结论仍错误"
### 决策 E:规则输入来源仅限 tool_invocation 事实
- **问题**:DateTimeTools、QueryMetricsTools 等非检索工具调用未写入 tool_invocation
- **接受**:evidence_score 定义本来就是检索证据充分度,非检索工具排除在外是合理的,不是 bug
- **已知限制**:调用了时间工具但 evidence_score = 0 的 session 存在
---
## 架构审计记录
- 接口影响:`POST /api/feedback` 是新接口(L2);ChatService 主流程返回值不变(L1)
- 时序验证:tool_invocation 在工具执行时同步写入,evaluate @Async 在 Agent 完成后触发,无竞态问题
- 已接受风险:
- `@Async` 失败时 selfEvaluation 保持 null,前端需处理 null
- 案例结构化字段(faultCategory 等)暂时填 GENERAL,后续可人工补充
- LLM 自评预留但未实现,Phase 2 再迭代
### 决策 F:ChatResult record + ChatResponse.sessionId 回传
- **问题**:`ChatService` 内部生成 8 位 sessionId,但从不返回给前端;前端用自己的 sessionId 调 feedback 接口,后端查不到 session(400)
- **决策**:新增 `ChatResult(answer, sessionId)` record,`executeChatWithStrategy` 链路全部返回 `ChatResult`;`ChatResponse` 增加 `sessionId` 字段;前端读取并闭包绑定至对应消息的反馈按钮
- **影响**:`ChatService` 三个方法签名变更(内部链路),`ChatController` 调用方更新,前端 `app.js` 读取新字段
### 决策 G:反馈 sessionId 闭包绑定而非全局变量
- **问题**:最初实现用 `this.lastSessionId` 全局变量,多轮对话时点击早期消息的反馈按钮会提交最新 sessionId
- **决策**:`createFeedbackBar(sessionId)` 接收 sessionId 参数,`submitFeedback(feedback, bar, sessionId)` 直接用传入值,不读全局状态
- **效果**:每条 AI 回复绑定自己那轮的 sessionId,多轮对话下行为正确
### 项目技术栈清单
- ChatModel 注入:`@Autowired ChatModel chatModel`,通过 `ModelRoutingConfig` 路由
- Repository:Spring Data JPA,`Optional<T>` 返回,方法命名约定
- DTO:独立文件放 `dto/` 包
- 异步:新建 `AsyncConfig.java` 加 `@EnableAsync`(项目原无此配置)
- 无 MQ,无加密,工具类直接用 UUID.randomUUID()
- 日志:SLF4J Logger,`LoggerFactory.getLogger()`
- `ToolInvocationRepository.findBySessionId` 已有,可直接用
### 参考实现文件
- `ChatService.java`:executeChat/executeChatComplex 流程
- `CaseLibraryRepository.findByDiagnosisId`:幂等检查用
- `DiagnosisSessionRepository.findBySessionId`
- `ToolInvocationRepository.findBySessionId`
@@ -0,0 +1,52 @@
# evidence.md — confidence-feedback
## 代码证据
### agent_step.thought 不可作为案例内容
- 文件:`AgentLoggingHook.java:135`
- 证据:`thought` 在写入前截断为 2000 字符,`modelOutput` 截断为 500 字符
- 结论:两者均不是返回给用户的完整答案,案例质量低
### ChatService 已有完整答案未持久化
- 文件:`ChatService.java:269`(executeChat)、`ChatService.java:353`(executeChatComplex)
- 证据:`String answer = response.getText()` 只用于返回前端,未写入任何持久化存储
- 结论:加 `DiagnosisSession.answer` 字段是最干净的方案
### ToolInvocationRepository 已有 findBySessionId
- 文件:`ToolInvocationRepository.java`
- 证据:`findBySessionId(String sessionId)` 已实现,返回 `List<ToolInvocation>`
- 结论:规则引擎可直接读取 tool_invocation 事实,无需新增查询方法
### tool_invocation 写入时序安全
- 文件:`LookupKnowledgeTool.java:144`
- 证据:`saveToolInvocation` 在工具执行时同步调用,早于 ChatService 的 SUCCESS 分支
- 结论:@Async evaluate 触发时 tool_invocation 数据已在库,无竞态
### 项目原无 @EnableAsync
- 证据:`grep -rn "EnableAsync"` 无任何命中(apply 前)
- 结论:需要新建 `AsyncConfig.java`
### CaseLibraryRepository.findByDiagnosisId 已有幂等检查支持
- 文件:`CaseLibraryRepository.java`
- 证据:`findByDiagnosisId(String diagnosisId)` 已实现
- 结论:useful 重复提交时可用此方法检查,不重复插入
## 设计推导
### evidence_score vs confidence 命名
- 基于工具调用的分数衡量的是证据收集充分度,不是答案准确性
- "confidence" 容易误解,改为 "evidence_score" 更准确
- LLM 自评才适合叫 confidence,但当前未实现
### BAD_CASE 不应混入 status
- status 有明确执行状态语义(RUNNING/SUCCESS/FAILED)
- 一个 SUCCESS 的 session 被标为 BAD_CASE 后,按 status 做的统计会失真
- feedback 字段本身就够,`WHERE feedback = 'not_useful'` 即可查 BadCase
@@ -0,0 +1,58 @@
# Acceptance: session-dedup-knowledge-map
## 静态验证
| 项目 | 结果 | 说明 |
|------|------|------|
| 编译检查 | PASS | `mvn compile -q` exit code 0,所有 17 个变更文件无编译错误 |
| 代码结构检查 | PASS | 6 个新文件(RetrievedDocTracker, DocumentFieldEnricher, KnowledgeDomainService, KnowledgeDomain, KnowledgeDomainRepository, V009 迁移)均存在且路径正确 |
| Prompt 外部化 | PASS | `doc-field-enricher-prompt.md` 和 `domain-summary-prompt.md` 位于 `src/main/resources/prompts/`,Java 代码通过 `@PostConstruct` + `ClassPathResource` 加载 |
| Flyway 迁移脚本 | PASS | `V009__add_knowledge_domain.sql` 存在,表结构完整 |
| DTO 字段 | PASS | Frontmatter / KnowledgeEntry / LookupResult 新增字段均已添加 |
| 解析器扩展 | PASS | FrontmatterParser 解析 `covers` 和 `when_to_retrieve` |
| Jackson 替换 | PASS | KnowledgeIndexService 不再包含 extractJsonValue/extractJsonArray,改用 objectMapper.readValue |
| Prompt 检索规则 | PASS | chat-planner-prompt.md 新增"知识库检索规则"区块(4 条规则) |
## 脚本验证
| 项目 | 结果 | 说明 |
|------|------|------|
| 单元测试 | 未运行 | 项目当前无针对本 change 的单元测试 |
| 集成测试 | 未运行 | 需启动应用 + Milvus + MySQL 验证完整链路 |
## 浏览器/人工验证
| 项目 | 结果 | 说明 |
|------|------|------|
| V009 迁移 | PASS | Flyway 日志:`Successfully applied 1 migration to schema superbiz_agent, now at version v009` |
| knowledge_domain 表数据 | PASS | 4 个域全部 LLM 生成 when_to_retrieve 成功(api/domain/infrastructure/troubleshooting),内容包含跨域边界引用 |
| knowledge map 注入 Planner | PASS | 多 Agent 路径正常触发 `Supervisor → chat_planner → chat_executor`,Planner 能按域做检索规划 |
| session 级去重 | PASS | 两个 session 均验证去重生效:session `7c517329` 去 4 次重拦截,session `9693b9fb` 6 次去重拦截 |
| LLM 字段生成 | 未验证 | 需上传新文档后检查 metadata JSON 中是否包含 covers 和 whenToRetrieve |
## 未验证项
| 项目 | 风险 | 建议补验步骤 |
|------|------|-------------|
| LLM 字段生成 | 中 — 依赖外部 LLM 服务 | 上传新文档,检查 metadata JSON 中是否包含 covers 和 whenToRetrieve |
## 启动问题修复
| 问题 | 修复 | 状态 |
|------|------|------|
| `@PostConstruct` 中调用 `knowledgeDomainService.onDocumentChange()` 导致循环依赖 | 将域级生成从 `@PostConstruct` 移到 `@EventListener(ApplicationReadyEvent.class)` | 已修复,编译通过 |
## 任务完成状态
14/14 任务全部完成 (T1-1 ~ T6-2)。
## 遗留问题
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/ISS-002-executor-unconstrained-lookup.md`。
## 已知限制
1. **RetrievedDocTracker 为 JVM 内存存储**:应用重启后去重状态丢失,同一会话内重启无法继续去重(可接受,会话通常短于重启间隔)
2. **Planner 只看域级 when_to_retrieve**:文档级细粒度筛选留 Phase 2
3. **文档级 prompt 依赖同域其他文档**:首个上传到某域的文档无法获得同域参照(此时 prompt 输出"无同域其他文档")
4. **域级 prompt 依赖其他域已入库**:首次启动且 DB 为空时,其他域信息从 L0 索引 category 列表兜底
@@ -0,0 +1,33 @@
# Brief: session-dedup-knowledge-map
## 背景
ISS-001:Executor 在单次对话中重复调用 `lookup_knowledge` 多达 20 次,同一文档被召回 13 次。原因是工具层无状态、Planner 无知识边界感知。
## 目标
1. 彻底消除 session 内重复文档召回(Part A)
2. 给 Planner 注入知识图谱,让其在规划阶段就能判断需要检索哪个域、只检索一次(Part B)
## 范围
- `LookupKnowledgeTool`:session 级去重
- `Frontmatter` / `KnowledgeEntry`:新增 covers + whenToRetrieve
- `DocumentManagementService`:上传时 LLM 生成文档级字段
- `KnowledgeDomainService`(新):域级聚合与 DB 存储
- `knowledge_domain` 表(新)
- `ChatService` + `chat-planner-prompt.md`:注入 knowledge map
## 非目标(Phase 2)
- Executor 文档级 when_to_retrieve 细粒度筛选
- RRF 混合重排
- 文档 frontmatter 自动生成(手动覆盖 LLM 优先已支持)
## 分档
standard
## 关联 OpenSpec
openspec/changes/session-dedup-knowledge-map/
@@ -0,0 +1,60 @@
# decisions.md — session-dedup-knowledge-map
## Question Pool
| # | 问题 | 类型 | 状态 |
|---|---|---|---|
| Q1 | domain.when_to_retrieve 来源(手动/自动聚合/LLM上传时生成) | user-interview | 已确认 |
| Q2 | LLM 生成时机(同步上传 vs 异步补全) | user-interview | 已确认 |
| Q3 | knowledge map 结构(域级平铺 vs 两层) | user-interview | 已确认 |
| Q4 | domain.when_to_retrieve 存储(内存 vs DB) | user-interview | 已确认 |
| Q5 | Executor 文档级细粒度筛选是否进 MVP | user-interview | 已确认 |
| E1 | ThreadLocal 在多 Agent 路径是否安全 | evidence-driven | 已汇报 |
| E2 | 6 个文档是否全部有 category 字段 | evidence-driven | 已汇报 |
| E3 | 去重 key 设计 | evidence-driven | 已汇报 |
| E4 | Planner prompt token 增量是否可接受 | evidence-driven | 已汇报 |
| E5 | EvaluationService.tool_call_count 影响 | evidence-driven | 已汇报 |
## Evidence-Driven 结论
- **E1**:`AsyncConfig` 只启用 `@EnableAsync`,无 TaskDecorator。`SupervisorAgent.invoke()` 是同步阻塞调用,工具调用与主线程同线程,ThreadLocal 当前路径安全。异步扩展时需补 TaskDecorator。
- **E2**:全部 6 个文档均有 `category` 字段:api(1)、domain(1)、infrastructure(3)、troubleshooting(1)。
- **E3**:`KnowledgeEntry.filePath` 在 L0 内唯一,L1 `_source` 字段也是 filePath,统一用 filePath 作去重 key。
- **E4**:当前 planner prompt 21 行,注入 knowledge map 约增加 200-400 字符,可接受。
- **E5**:去重后 `agent_step.has_tool_call` 减少,`tool_call_count` 降低,这是修复效果,`EvaluationService` 评分规则无需改动。
## User-Interview 确认记录
**Q1** — doc.when_to_retrieve 来源
用户原话:选 C(上传时 LLM 自动生成)
确认状态:已确认
**Q2** — LLM 生成时机
用户原话:选 X(同步,上传时当场生成)
确认状态:已确认
**Q3** — knowledge map 结构
用户原话:认可两层结构(domain → documents[])
确认状态:已确认
补充:Planner 只注入域级 when_to_retrieve,文档级 when_to_retrieve 留 Executor 筛选(Phase 2)
**Q4** — domain.when_to_retrieve 存储
用户原话:存 DB,这样每次启动都不用让 LLM 再总结一次
确认状态:已确认 → 新建 knowledge_domain 表,Flyway 迁移脚本
**Q5** — Executor 文档级细粒度筛选
用户原话:留 Phase 2
确认状态:已确认,MVP 不做
## Pre-apply 补充决策
- **P1:KnowledgeIndexService.parseDocumentToEntry 替换为 Jackson**:`extractJsonValue` / `extractJsonArray` 手写解析器遇到含逗号、引号的自然语言字段(whenToRetrieve)会截断。全量替换为 `objectMapper.readValue(metadata, Frontmatter.class)`,影响范围仅 `KnowledgeIndexService`,行为更健壮。(用户确认)
- **P2:LookupResult 新增 message 字段**:去重命中时 `found=false` + `message="文档已在本会话中检索过:xxx"`,不复用 `primary.content`。语义清晰,LLM 能理解原因不会重试。(用户确认)
## 关键设计决策
1. **两级 when_to_retrieve**:文档级(upload 时 LLM 生成,存 metadata)+ 域级(文档变更时 LLM 聚合,存 knowledge_domain 表)
2. **域级重算触发**:文档上传后、文档删除后,只重算受影响的域(不是全量);`loadIndex()` 时如果某域在 DB 没有记录,则触发生成
3. **注入 Planner 只给域级**:knowledge map 只包含域级 when_to_retrieve + documents[](title + covers),不暴露文档级 when_to_retrieve
4. **去重 key**:filePath(L0+L1 统一)
5. **去重状态存储**:JVM 内 `ConcurrentHashMap<sessionId, Set<filePath>>`,`SessionContextHolder.clear()` 时同步清理
@@ -0,0 +1,87 @@
# Evidence: session-dedup-knowledge-map
## E1: ThreadLocal 在多 Agent 路径是否安全
**问题**:`SessionContextHolder` 基于 ThreadLocal,多 Agent 异步路径可能导致 sessionId 丢失。
**证据**:
- `AsyncConfig` 只启用 `@EnableAsync`,无 `TaskDecorator`
- `SupervisorAgent.invoke()` 是同步阻塞调用,工具调用与主线程同线程
- 当前路径下 ThreadLocal 安全
**结论**:当前同步路径安全。未来引入异步扩展时需补 `TaskDecorator` 传递 ThreadLocal。
---
## E2: 6 个文档是否全部有 category 字段
**问题**:域聚合依赖 `category` 字段分组,需确认现有文档是否都有值。
**证据**:
- 全部 6 个文档均有 `category` 字段:api(1)、domain(1)、infrastructure(3)、troubleshooting(1)
**结论**:现有文档无需修补,category 覆盖率 100%。
---
## E3: 去重 key 设计
**问题**:用什么字段唯一标识一个文档用于去重。
**证据**:
- `KnowledgeEntry.filePath` 在 L0 索引内唯一
- L1 向量索引的 `_source` 字段也是 filePath
- 上传时 `saveToLocal()` 生成 `knowledge_base/{category}/{fileName}` 路径
**结论**:统一用 `filePath` 作去重 key,L0 和 L1 一致。
---
## E4: Planner prompt token 增量是否可接受
**问题**:knowledge map YAML 注入 Planner prompt 会增加固定 token 开销。
**证据**:
- 当前 planner prompt 21 行
- 注入 knowledge map 约增加 200-400 字符(6 个文档场景)
- 相比 Planner 整体 prompt + 历史消息,增量占比 < 5%
**结论**:可接受,不构成性能瓶颈。
---
## E5: EvaluationService.tool_call_count 影响
**问题**:去重后 `tool_call_count` 降低,是否影响 `EvaluationService` 评分逻辑。
**证据**:
- `EvaluationService` 使用 `tool_call_count` 作为评分因子
- 去重导致重复调用被过滤,`tool_call_count` 下降
- 这是修复效果(消除了无意义的重复调用),不是回归
**结论**:`EvaluationService` 评分规则无需改动。下降的 `tool_call_count` 反映了真实效率提升。
---
## P1: 手写 JSON 解析器脆弱性
**问题**:`KnowledgeIndexService.extractJsonValue` / `extractJsonArray` 在遇到含逗号、引号的自然语言字段时会截断。
**证据**:
- `whenToRetrieve` 字段由 LLM 生成,内容为自然语言(含逗号、分号等标点)
- 手写解析器以 `"` 和 `,` 作分隔符,自然语言中的标点会导致提前截断
- Jackson `ObjectMapper.readValue(metadata, Frontmatter.class)` 是项目已有依赖
**结论**:全量替换为 Jackson,影响范围仅 `KnowledgeIndexService.parseDocumentToEntry()`,行为更健壮。
---
## P2: LookupResult 去重提示字段
**问题**:去重命中时如何向 LLM 返回"不要重试"的信号。
**证据**:
- 复用 `primary.content` 语义不清,LLM 可能理解为正常检索结果
- 独立 `message` 字段 + `found=false` 语义明确,LLM 能理解"已检索过"不再重试
**结论**:`LookupResult` 新增 `String message` 字段,去重时填入提示文本。
@@ -0,0 +1,70 @@
# Acceptance: executor-action-memory-relevance
## 分档
standard
## 任务完成状态
| 任务 | 状态 | 说明 |
|------|------|------|
| T1: RetrievedDocTracker 域级升级 | ✅ 完成 | 双层 Map 结构,域级+文档级记录 |
| T2: LookupResult 新增字段 | ✅ 完成 | relevanceLevel / completenessHint / retrievedDomainsThisSession |
| T3: 归一化计算逻辑 | ✅ 完成 | Min-Max 归一化 + 三等级判定 |
| T4: LookupKnowledgeTool 集成 | ✅ 完成 | 归一化层 + 行动记忆注入 + 域拦截 |
| T5: Executor Prompt 重写 | ✅ 完成 | 4 条检索约束,无 knowledge map |
| T6: 入库可观测性 | ✅ 完成 | V010 + Entity + JSON 扩展 |
| T7: BGE-M3 归一化验证测试 | ✅ 完成 | 范数=1.00000002,测试通过 |
## 静态验证
- [x] **语法/编译检查**: 所有 Java 文件编译通过
- [x] **Impact Analysis**: LookupKnowledgeTool、RetrievedDocTracker 变更范围经 `gitnexus_impact` 检查,均为 L2 内部接口影响
- [x] **Cross-artifact 对齐检查**: brief → proposal → design → specs → tasks 闭环,无 gap
- [x] **Prompt 约束检查**: chat-executor-prompt.md 不包含 knowledge map,包含 4 条检索约束
## 脚本验证
- [x] **V010 Flyway 迁移**: 迁移成功,`relevance_level` 和 `dedup_reason` 列已添加
```sql
ALTER TABLE tool_invocation
ADD COLUMN relevance_level VARCHAR(20),
ADD COLUMN dedup_reason VARCHAR(32);
```
- [x] **FullPipelineSmokeTest**: BGE-M3 归一化测试通过(范数=1.00000002)
- [x] **数据库数据校验**:
- `relevance_level` 列已写入 HIGHLY_RELEVANT / REFERENCE
- `dedup_reason` 列已写入 doc_retrieved / null
- `retrieval_details` JSON 包含 l1_top_similarity、completeness_hint、retrieved_domains、dedup_reason
## 浏览器/人工验证
- [x] **应用启动验证**: Spring Boot 应用正常启动,端口 9900
- [x] **Chat API 调用验证**: 通过 curl 测试 chat 接口,lookup_knowledge 调用链完整
```
curl -X POST "http://localhost:9900/api/chat/send" \
-H "Content-Type: application/json" \
-d '{"sessionId": "b66d799e", "question": "..."}'
```
- [x] **日志验证**: 应用日志可观察到 relevanceLevel、retrievedDomainsThisSession 输出
- [x] **归一化数学验证**: l1_top_score=0.383 → l1_top_similarity=0.8085(`1 - 0.383/2.0 = 0.8085`)✅
- [x] **域追踪验证**: `[infrastructure]` → `[infrastructure, api]` 域列表正常扩展
## 未验证
| 场景 | 原因 | 风险 | 补验建议 |
|------|------|------|---------|
| PRECISE 等级(L0 唯一精确匹配) | 测试会话无精确匹配场景 | 低 — L0 matchCount=1 的判断逻辑与 HIGHLY_RELEVANT 共用,实现确定性强 | 构造一条 L0 精确匹配的知识库文档后测试 |
| domain_retrieved 域级去重 | 需要同一域全部文档已检索再查该域才触发 | 低 — isDomainRetrieved 逻辑简单,与 isDocRetrieved 等价 | Phase 2 启用域级硬限流时测试 |
| DEDUPED 等级 | 当前 code path 去重时仍写 REFERENCE,DEDUPED 未被使用 | 低 — 设计预留,当前未启用 | Phase 2 若启用 DEDUPED 等级时验证 |
| Phase 2 域级硬限流 | 非本次范围 | 中 — 当前仅有软约束(prompt),LLM 仍可能在 REFERENCE 下继续检索 | 实测观察,如果 lookup 调用仍偏高,启动 Phase 2 |
## 剩余风险
1. **Prompt 软约束局限性**:实测 10 次调用中 9 次为 REFERENCE,说明 LLM 仍倾向于继续检索。如果 prompt 约束效果不足,需启用 Phase 2 域级硬限流。
2. **L1 Metadata 解析兼容性**:L1 domain 兜底路径解析 metadata JSON,如果知识库文档 frontmatter 格式不一致可能解析失败,已有 try-catch 兜底。
## 归档状态
- [ ] OpenSpec change 尚未归档
- [ ] devflow/index.md 状态为 `implemented`,待改为 `archived`
@@ -0,0 +1,35 @@
# Brief: executor-action-memory-relevance
## 背景
ISS-002:Executor 在单次会话中调用 `lookup_knowledge` 20+ 次,大部分是同域换变体的冗余调用。前序 change `session-dedup-knowledge-map` 解决了文档级重复召回(ISS-001),但未解决 Executor 重复调用问题。
## 目标
- Executor 获得行动记忆(知道自己本次会话已检索了哪些域)
- 检索结果提供归一化质量等级(PRECISE/HIGHLY_RELEVANT/REFERENCE)+ 兜底信号
- Executor prompt 提供明确的检索约束和"放弃检索"的合法出口
- 原始分数入库保留可观测性,但不暴露给 LLM
## 范围
- `RetrievedDocTracker`:域级 + 文档级双层记录
- `LookupKnowledgeTool`:归一化层 + 行动记忆注入
- `LookupResult`:新增 relevanceLevel / completenessHint / retrievedDomainsThisSession
- `chat-executor-prompt.md`:检索约束重写
- `ToolInvocation` + V010:入库可观测性
## 非目标
- 不给 Executor 注入 knowledge map(保持 Agent 边界)
- 不修改 Planner prompt 或 Planner 逻辑
- 不修改 PrimaryResult / SupplementResult 的字段(不暴露原始分数)
- Phase 2 域级硬限制暂不实施
## 分档
standard
## 关联 OpenSpec change
openspec/changes/executor-action-memory-relevance
@@ -0,0 +1,83 @@
# Decisions: executor-action-memory-relevance
## 过程日志
### Clarify 阶段
**入口摘要**:ISS-002 Executor 无约束重复调用 lookup_knowledge(单会话 20+ 次),需要行动记忆 + 归一化质量等级 + prompt 约束来解决。
**slug**: `executor-action-memory-relevance`
**规模分档**: `standard`(涉及 7 个文件,跨 DTO/工具层/持久化/Prompt,有设计决策需澄清)
### Context 阶段
**devflow/index.md 使用状态**: 已命中。前序 change `session-dedup-knowledge-map`(archived)提供了 RetrievedDocTracker、KnowledgeDomainService、ISS-002 文档。
**相关 ADR**: 无直接 ADR,但 `session-dedup-knowledge-map` 的 decisions.md 和 evidence.md 记录了文档级去重和 knowledge map 注入的决策。
**不能违反的历史决策**:
1. RetrievedDocTracker 的文档级去重必须保留
2. knowledge map 只注入 Planner,不注入 Executor(本次讨论确认)
3. L0/L1 原始分数不暴露给 LLM,只在归一化层内部使用(本次讨论确认)
**需进入 OpenSpec 的上下文点**:
1. L1 score 是 L2 距离(值域 [0,+∞)),不是归一化分数——阈值设计需基于实际分布
2. L0 的 category 可从 KnowledgeEntry.getCategory() 直接获取;L1 需解析 metadata JSON
3. ReactAgent 是自主决策工具调用的 Agent,Prompt 约束是软约束
### Grill 阶段 — Question Pool
**维度:术语**
1. [evidence-driven] `relevanceLevel` 三个等级(PRECISE/HIGHLY_RELEVANT/REFERENCE)的边界是否清晰,是否存在 LLM 误解的可能? → **已查证**:三个等级语义明确,PRECISE=唯一匹配、HIGHLY_RELEVANT=高分命中、REFERENCE=低置信度参考。LLM 理解风险低。
**维度:边界**
2. [evidence-driven] L1 score 是 L2 距离(值域 [0,+∞)),当前代码无阈值判断。归一化阈值如何设计? → **已查证**:L2 距离典型范围取决于 BGE-M3 1024 维 embedding 的尺度,需从 `tool_invocation.retrieval_details` 中查询实际 `l1_scores` 分布才能定阈值。当前先以常量定义,标记为"需实测校准"。
3. [evidence-driven] L1 结果的 category 提取需要解析 metadata JSON 字符串,当前 `SearchResult.metadata` 是 `toString()` 的结果。归一化层是否需要 L1 的 domain? → **已查证**:L1 的 domain 主要用于 RetrievedDocTracker 的域级记录。如果 L0 已命中且包含 category,可直接用 L0 的 category;如果仅 L1 命中,需解析 metadata 提取 category。当前知识库中 L0 大概率先命中,L1 domain 提取作为兜底路径。
4. [user-interview] 归一化阈值(L1 score 分界线)在实测数据不足时,是否接受先用保守初始值 + 后续调优的策略? → **用户待确认**
**维度:验收**
5. [evidence-driven] 现有 `tool_invocation` 表 `retrieval_details` JSON 中 `l1_scores` 存的是 L2 距离原始值,新增的 `relevance_level` 和 `completeness_hint` 入库后是否需要回填历史数据? → **已查证**:不需要回填历史数据,新列 nullable 即可,历史记录 relevance_level=null。
### Grill 结论
**evidence-driven 汇报**:
- E1: relevanceLevel 三等级语义清晰,LLM 误解风险低
- E2: L1 score 是 L2 距离,值域不固定,阈值需实测校准
- E3: L0 category 直接可用,L1 category 需解析 metadata(兜底路径)
- E4: 历史数据不回填,新列 nullable
**user-interview 已确认**:
- Q4: 归一化阈值先用保守初始值 + 后续调优 → **用户已确认**,并建议用 Min-Max 归一化到 [0,1]
### Specify 阶段补充
**BGE-M3 L2 归一化实测验证**:
- FullPipelineSmokeTest.embeddingBgeM3Works() 新增 L2 范数断言
- 结果:范数=1.00000002,误差 < 0.01,测试通过
- 结论:BGE-M3 输出为 L2 归一化单位向量,L2 距离数学硬上界 = 2.0
- Min-Max 归一化公式:`similarity = 1 - min(l2Score, 2.0) / 2.0`
**Cross-artifact 对齐检查**:
| 对齐项 | 状态 |
|--------|------|
| brief 目标/范围/非目标 → proposal 覆盖 | 已对齐 |
| proposal 范围/约束 → design 覆盖 | 已对齐 |
| design 归一化/行动记忆/接口影响 → specs 覆盖 | 已对齐 |
| specs 可观察行为 → tasks 覆盖 | 已对齐 |
**接口影响分级**:
- RetrievedDocTracker 数据结构升级 → L2(内部接口,消费者只有 LookupKnowledgeTool)
- LookupResult 新增 3 字段 → L2(工具返回值,无跨模块调用方)
- tool_invocation 新增 2 列 → L2(Flyway nullable,不影响现有查询)
- chat-executor-prompt.md 更新 → L1(Prompt 文本变更)
### Audit 阶段
**架构风险评估**(5 句以内):
1. 归一化层嵌入 LookupKnowledgeTool 内部(静态方法),无跨模块耦合风险。
2. RetrievedDocTracker 升级为双层结构,数据量级不变(文档数 × session 数),内存无风险。
3. L1 metadata 解析 category 是兜底路径,如果 JSON 格式不一致可能解析失败——已有 try-catch 兜底。
4. 归一化阈值 yml 配置化,运行时调优不需要改代码和重启——运维友好。
5. Prompt 约束仍依赖 LLM 遵守——如果 Phase 1 效果不足,Phase 2 域级硬限制的 isDomainRetrieved 已就绪,无需额外改造。
@@ -0,0 +1,65 @@
# Evidence: executor-action-memory-relevance
## Evidence-driven 结论
### E1: relevanceLevel 三等级语义清晰度
- **来源**: Grill 阶段 Question Pool #1
- **查证结果**: 三个等级语义明确,边界清晰:
- PRECISE:L0 唯一精确匹配,LLM 应直接使用
- HIGHLY_RELEVANT:归一化 similarity ≥ 0.75,高度相关
- REFERENCE:归一化 similarity ≥ 0.5,相关参考
- **结论**: LLM 误解风险低,语义边界足够清晰
### E2: L1 Score 值域与归一化阈值
- **来源**: Grill 阶段 Question Pool #2
- **查证结果**:
- L1 score 是 L2 距离,值域 [0, +∞)
- BGE-M3 输出为 L2 归一化单位向量(实测范数=1.00000002),L2 距离数学硬上界 = 2.0
- Min-Max 归一化公式:`similarity = 1 - min(l2Score, 2.0) / 2.0`
- **结论**: 使用 `maxL2Distance=2.0` 作为归一化上界,阈值 yml 可配置
### E3: L1 Domain 提取兜底路径
- **来源**: Grill 阶段 Question Pool #3
- **查证结果**:
- L0 的 domain 可从 `KnowledgeEntry.getCategory()` 直接获取
- L1 结果的 domain 需解析 `SearchResult.metadata` JSON 字符串
- 当前知识库设计下 L0 大概率先命中,L1 domain 提取作为兜底
- **结论**: 先尝试 L0 category,失败时解析 L1 metadata JSON(try-catch 兜底)
### E4: 历史数据不回填
- **来源**: Grill 阶段 Question Pool #5
- **查证结果**: 新列 `relevance_level` 和 `dedup_reason` 均为 nullable,不影响现有查询
- **结论**: 历史记录保持 null,不需要回填迁移
### E5: BGE-M3 L2 归一化实测验证
- **来源**: Specify 阶段 + FullPipelineSmokeTest
- **查证结果**:
- embeddingBgeM3Works() 测试新增 L2 范数断言
- 实测范数 = 1.00000002,误差 < 0.01
- 测试通过,BGE-M3 输出确认为 L2 归一化单位向量
- **结论**: L2 距离上界 = 2.0 的数学依据成立
### E6: V010 迁移验证
- **来源**: Apply 阶段运行时验证
- **查证结果**:
- Flyway V010 迁移成功执行
- `relevance_level` VARCHAR(20) 列可空,已正确写入
- `dedup_reason` VARCHAR(32) 列可空,已正确写入
- `retrieval_details` JSON 扩展字段(l1_top_similarity、relevance_level、completeness_hint、retrieved_domains、dedup_reason)全部写入
- **结论**: 入库可观测性符合设计
### E7: 数据库数据校验
- **来源**: Apply 阶段运行时验证
- **查证结果**:
- session `b66d799e` 共 10 条 lookup_knowledge 调用
- id=138: L2=0.383 → similarity=0.8085 → HIGHLY_RELEVANT(符合预期)
- id=139-147: 主要为 REFERENCE,doc_retrieved 去重正常触发
- retrieved_domains 域追踪:`[infrastructure]` → `[infrastructure, api]` 正常扩展
- **结论**: 归一化、行动记忆、去重机制数据层面全部验证通过
@@ -0,0 +1,58 @@
# Acceptance: chat-verifier-agent
## Classification
standard
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Verifier prompt | Done | Strict JSON schema, verdict matrix, fact classifications, and `evidence_refs` are defined. |
| VerifierInputHook | Done | Explicit verifier payload replaces raw conversation history. |
| ChatService integration | Done | Planner, executor, and verifier are called explicitly with max two rounds. |
| Verdict routing | Done | PASS, LOW_CONFID, and REJECT paths are handled in code. |
| Trace summary | Done | Evidence summaries include `trace_ref` and `source_invocation_ids`. |
| self_evaluation merge | Done | `rule_evaluation` and `verifier_evaluation` are preserved independently. |
| Verifier observability | Done | `verifier_evaluation` persists facts, evidence refs, trace summary, rationale, score, and round. |
## Static Verification
- [x] OpenSpec artifacts exist: `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, `tasks.md`, `.committed`.
- [x] `change.json` exists and has `metadata.status = committed`.
- [x] `.archive-ready` exists.
- [x] devflow archive-prep files exist: `brief.md`, `evidence.md`, `decisions.md`, `acceptance.md`.
- [x] `devflow/index.md` contains `chat-verifier-agent` with status `archived`.
## Script Verification
- [x] `mvn -q -DskipTests compile` passed.
## Runtime Verification
- [x] POST `/api/chat` with a complex question returned successfully.
- [x] Runtime session `9138f064` showed planner, executor, and verifier execution in logs.
- [x] Runtime session `9138f064` wrote `verifier_evaluation.verdict = LOW_CONFID`.
- [x] Runtime session `9138f064` wrote `facts_checked[*].evidence_refs`.
- [x] Runtime session `9138f064` wrote `tool_trace_summary[*].source_invocation_ids`.
- [x] LOW_CONFID final answer included disclaimer and verifier-derived evidence gaps.
## Unverified
| Scenario | Reason | Risk | Follow-up |
| --- | --- | --- | --- |
| PASS runtime path | The exercised complex runtime case produced LOW_CONFID. | Low; PASS routing is simple pass-through after parsed verifier decision. | Add a fixture or deterministic verifier test if this becomes product-critical. |
| REJECT runtime path | No forced contradiction case was run after traceability changes. | Medium; REJECT is the safety-critical degraded path. | Add a targeted test with a fabricated claim and evidence contradiction. |
| Document-path-level evidence mapping | Current implementation records invocation ids and source document labels, not guaranteed canonical document paths for every retrieval mode. | Low for current audit need; medium for future UI drill-down. | Extend retrieval details with canonical document paths in a later change. |
## Remaining Risks
1. Verifier output still depends on model compliance with JSON schema; code falls back to LOW_CONFID on missing or invalid output.
2. `AgentLoggingHook` is shared by several agent paths; current changes preserve compile and runtime behavior but should be watched in AiOps flows.
3. `SupervisorAgent` construction remains as legacy residue in `ChatService`; runtime orchestration is explicit, but a later cleanup should remove unused supervisor construction.
## Archive State
- [x] OpenSpec change is archive-ready.
- [x] OpenSpec change has been moved to `openspec/changes/archive/2026-07-03-chat-verifier-agent/`.
- [x] Main spec exists at `openspec/specs/chat-verifier-agent/spec.md`.
@@ -0,0 +1,36 @@
# Brief: chat-verifier-agent
## Background
The complex Chat path previously returned Executor answers without a synchronous quality gate. Existing rule scoring was asynchronous and post-hoc, so it could not prevent unsupported answers from reaching users.
## Goals
1. Add a Verifier Agent after Executor in the complex chat path.
2. Require structured verifier output with `PASS`, `LOW_CONFID`, or `REJECT`.
3. Route final user output in code based on verifier verdict.
4. Persist verifier results under `diagnosis_session.self_evaluation.verifier_evaluation`.
5. Preserve rule scoring under `rule_evaluation`.
6. Make verifier decisions traceable to real tool invocations through `evidence_refs` and `source_invocation_ids`.
## Scope
- `ChatService`: explicit `planner -> executor -> verifier` orchestration, max two rounds, verdict routing, retry context, verifier persistence.
- `VerifierInputHook`: explicit verifier input payload.
- `ToolTraceSummaryService`: evidence summary from persisted tool calls.
- `VerifierContextHolder`: round-local verifier context.
- `SelfEvaluationMergeService`: safe JSON merge for evaluation channels.
- `AgentLoggingHook`: concise verifier thought and fuller structured output retention.
- `chat-verifier-prompt.md`: verifier contract, verdict matrix, and traceability schema.
## Non-Goals
- Verifier does not call tools.
- Verifier does not rewrite Executor output.
- Single-agent chat path remains outside this change.
- No database schema migration is included.
- Document-path-level evidence attribution is deferred; current traceability is invocation-level with source document labels.
## Related OpenSpec
`openspec/changes/archive/2026-07-03-chat-verifier-agent/`
@@ -0,0 +1,123 @@
# Decisions: chat-verifier-agent
## 过程日志
### Clarify 阶段
**入口摘要**: 在 Chat 多 Agent 链路中新增 Verifier Agent,作为 Executor 输出后的质量门禁,做事实核查。
**slug**: `chat-verifier-agent`
**规模分档**: standard
### Context 阶段
**devflow/index.md 使用状态**: 已命中。前序 change `executor-action-memory-relevance`(archived)提供了 Chat 多 Agent 当前链路(Supervisor → Planner → Executor)。
**不能违反的历史决策**:
1. Executor 已有完整的行动记忆和归一化质量等级,Verifier 不需要重复验证检索质量
2. Chat Supervisor 的职责是调度,Verifier 作为子 Agent 加入后不改变 Supervisor 的定位
3. 已有 evidence_score 做事后评分,Verifier 是事前门禁,两者不冲突
**需进入 OpenSpec 的上下文点**:
1. Verifier 不需要工具调用,只是一个质量核查 Agent
2. Verifier 需要访问 Executor 的输出 + 工具调用记录
3. Supervisor prompt 需要重写以包含 Verifier 调度规则
4. groundedness_score 的阈值需要在代码中定义
### Grill 阶段 — Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|------|------|------|------|
| Q1 | 术语 | evidence_score(事后评分)与 Verifier(事前门禁)职责是否冲突? | evidence-driven | 已解决 |
| Q2 | 边界 | Verifier 需要的"工具调用记录"在 SupervisorAgent 中是否自动传递? | evidence-driven | 已解决 |
| Q3 | 边界 | LOW_CONFID < 0.5 回调 Planner 后的新输出是否再次走 Verifier?循环上限多少? | user-interview | 已解决 |
| Q4 | 验收 | Verifier 判决结果如何可观测?是否写入 agent_step 或 tool_invocation? | user-interview | 已解决 |
| Q5 | 验收 | 当前 Supervisor 硬编码 prompt 是否支持多 Agent 路由变更? | evidence-driven | 已解决 |
| Q6 | 技术 | Verifier 如何隔离 Executor 的中间推理过程,只看到干净的 query + tool 记录 + 最终答案? | user-interview | 已解决 |
| Q7 | 验收 | groundedness_score 阈值(0.5)是否需要配置化? | user-interview | 已解决 |
### Evidence-driven 结论
| 结论 | 证据来源 | 是否已汇报用户 |
|------|---------|-------------|
| evidence_score(异步事后)与 Verifier(同步事前门禁)不冲突 | EvaluationService.java: @Async 注解 | 已汇报 |
| SupervisorAgent 自动传递完整对话状态,Verifier 无需额外传递工具记录 | Spring AI Alibaba SupervisorAgent 实现 | 已汇报 |
| Supervisor prompt 为字符串字面量,直接修改即可 | ChatService.java:353 .systemPrompt("...") | 已汇报 |
### User-interview 记录
| 问题 | 用户原话 | 确认状态 | OpenSpec 回写 |
|------|---------|---------|-------------|
| Q3: LOW_CONFID < 0.5 回调 Planner 循环上限? | "可以,回调一次" | 已确认 | 已回写 proposal |
| Q4: Verifier 判决写入哪里做可观测? | "可以"(写入 diagnosis_session.self_evaluation JSON) | 已确认 | 已回写 proposal |
| Q6: Verifier 如何隔离 Executor 中间推理? | "用 MessagesModelHook 过滤 messages" | 已确认 | 已回写 design |
| Q7: groundedness_score 阈值是否需要配置化? | "需要配置化" | 已确认 | 已回写 design |
### Specify 阶段 — Cross-Artifact 对齐检查
| 上游 → 下游 | 检查内容 | 状态 |
|---|---|---|
| proposal → design | 范围、约束、关键承诺是否进入 design | 已对齐 |
| design → specs | 关键决策、模块地图是否进入 specs | 已对齐 |
| specs → tasks | 可观察行为是否被 tasks 覆盖为可执行切片 | 已对齐 |
**接口影响分级**:
- buildChatVerifierAgent() 新增方法 → L1(内部方法,无外部消费者)
- VerifierInputHook 类 → L1(内部 Hook,无外部消费者)
- Supervisor prompt 重写 → L1(仅影响 Chat 多 Agent 内部调度)
- subAgents 列表变更 → L1(Supervisor 内部配置)
- verifier.low-confidence-threshold 配置 → L1(新增配置项,不改已有配置)
### Audit 阶段
**模块链路**:
```
用户 → Supervisor → Planner(步骤) → Executor(答案+工具记录)
│
Supervisor 调用 Verifier
│
[VerifierInputHook BEFORE_MODEL]
├─ 保留:system prompt + user query
├─ 保留:tool call 记录(输入+返回)
├─ 保留:Executor 最终答案
└─ 去除:Executor 中间推理、Planner 规划过程
│
Verifier 判决
│
┌─── PASS ───→ 直接输出
├─── LOW_CONFID≥0.5 → 带声明输出
├─── LOW_CONFID<0.5 → 回调 Planner(一次)
└─── REJECT → 降级输出
│
写入 self_evaluation JSON
```
**架构风险评估**(5 句以内):
1. Verifier 是轻量 Agent(无工具、无外部依赖),架构风险低。
2. MessagesModelHook 纯过滤逻辑,不引入新数据源。
3. LOW_CONFID 分级处理 + 回调仅一次的设计,避免无限循环风险。
4. REJECT 降级确保编造内容不到达用户。
5. 审计结论不影响现有 design/tasks,无需回写。
### 关键取舍
- 决策:LOW_CONFID < 0.5 回调 Planner 一次
- 原因:给系统一次修正机会,但避免无限循环
- 影响:Supervisor prompt 需维护"已回调"状态
- 风险接受:用户已确认
- 决策:Verifier 判决写入 diagnosis_session.self_evaluation JSON
- 原因:不改表结构,与 evidence_score 统一可观测体系
- 影响:ChatService 后处理需追加 JSON
- 风险接受:用户已确认
### Archive-Ready Update
- 实现调整:最终运行链路由 `ChatService` 显式调用 `planner -> executor -> verifier`,不再依赖 Supervisor prompt 保证 verifier 被调用。
- 可追溯性补充:`tool_trace_summary` 增加 `trace_ref`、`source_invocation_ids`、查询样本、检索层级、相关性等级和来源文档标签。
- 可追溯性补充:`facts_checked[*].evidence_refs` 被 prompt 要求、代码解析并持久化。
- 验证记录:`mvn -q -DskipTests compile` 通过。
- 验证记录:运行会话 `9138f064` 走通 planner、executor、verifier,并持久化 `verifier_evaluation.facts_checked[*].evidence_refs` 与 `tool_trace_summary[*].source_invocation_ids`。
- 当前状态:OpenSpec change 已归档到 `openspec/changes/archive/2026-07-03-chat-verifier-agent/`,主规格已同步到 `openspec/specs/chat-verifier-agent/spec.md`。
@@ -0,0 +1,52 @@
# Evidence: chat-verifier-agent
## Code Evidence
### Complex chat path now invokes verifier deterministically
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
- Evidence: `executeChatComplex` calls planner, executor, then verifier directly through `callAgent(...)`.
- Conclusion: runtime no longer depends on prompt-only Supervisor behavior to call verifier.
### Verifier receives explicit inputs
- File: `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
- Evidence: the hook builds a JSON payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- Conclusion: verifier input is stable and does not depend on guessing the last assistant message from raw history.
### Tool evidence is traceable to persisted invocations
- File: `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
- Evidence: summaries include `trace_ref`, `source_invocation_ids`, `query_samples`, `retrieval_layers`, `relevance_levels`, and `source_documents`.
- Conclusion: verifier facts can be correlated with actual `tool_invocation` rows.
### Verifier facts preserve evidence references
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
- Evidence: verifier parsing preserves `facts_checked[*].evidence_refs` and persists `tool_trace_summary` under `verifier_evaluation`.
- Conclusion: `self_evaluation` now contains both verifier judgments and the evidence index used to form them.
### Evaluation channels no longer overwrite each other
- File: `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
- Evidence: rule and verifier evaluations are merged into separate keys.
- Conclusion: asynchronous rule scoring preserves verifier output.
### Verifier logging is less noisy
- File: `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
- Evidence: verifier `thought` stores a concise verdict summary, while fuller model output remains available in structured storage.
- Conclusion: `agent_step.thought` is no longer a misleading place for full verifier JSON.
## Runtime Evidence
- Compile verification passed: `mvn -q -DskipTests compile`.
- Runtime session `9138f064` executed `planner -> executor -> verifier`.
- Runtime session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
- Runtime session `9138f064` persisted `verifier_evaluation.tool_trace_summary[*].source_invocation_ids`.
## Design Evidence
- `LOW_CONFID` returns a fixed disclaimer and verifier-derived gaps.
- `REJECT` returns degraded output and does not pass through the raw Executor answer.
- `retry_context` is derived from verifier-identified missing evidence facts.
@@ -0,0 +1,65 @@
# MVP Demo Trace Acceptance
## Result
Accepted for implementation scope.
## Verification
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
- Notes: New trace controller, service, DTO, profile, verifier fallback, and test sources compile with the project.
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisTraceServiceTest,ChatServiceSupervisorAgentTest" test`
- Result: passed
- Notes: Covers successful trace aggregation, missing-session 404 path via `SessionNotFoundException`, low-confidence no-retry behavior, method-tool injection, and verifier fallback when Supervisor skips `chat_verifier`.
### OpenSpec Verification
- Command: `openspec validate mvp-demo-trace-acceptance --strict`
- Result: passed
### GitNexus Verification
- Result: skipped by user decision
- Notes: User requested subsequent project flow to bypass GitNexus.
### Manual / Runtime Verification
- Steps: Follow `mvp/demo/README.md` with `--spring.profiles.active=mvp-demo`.
- Result: passed
- Notes:
- Session `mvp-demo-payment-timeout-20260703-rerun2` completed as `SUCCESS`.
- Chat request returned `code=200`, `success=true`, and the same `sessionId`.
- Chat duration was `96316 ms`; persisted session duration was `95028 ms`.
- Trace API returned `code=200`, `returnedSteps=13`, `returnedTools=12`, `hasVerifier=true`, and `verifierVerdict=LOW_CONFID`.
- Trace agents included `planner,executor,verifier`.
- Trace tools included `lookup_knowledge,query_logs,query_metrics`.
- Feedback submission returned success, and a follow-up trace query showed `feedback=useful`.
- MySQL verification confirmed `agent_step` count `13` with agents `executor,planner,verifier`.
- MySQL verification confirmed `tool_invocation` count `12` with tools `lookup_knowledge,query_logs,query_metrics`.
## Completed Scope
- Added `GET /api/diagnosis/{sessionId}/trace`.
- Added read-only trace aggregation from persisted diagnosis tables.
- Added `mvp-demo` profile overlay.
- Added payment-timeout demo acceptance documentation.
- Added MVP note for interview storytelling.
- Added verifier fallback so runtime trace remains complete when Supervisor returns without `verifier_output`.
## Known Limits
- `mvp-demo` is not a fully offline mock runtime.
- Runtime still depends on available MySQL, Redis, Milvus/Zilliz, model, and embedding configuration.
- Sensitive configuration cleanup remains intentionally deferred.
- Supervisor can still make inefficient routing choices inside a single round; `ChatService` now invokes `chat_verifier` as a fallback when Supervisor returns without `verifier_output`, so trace completeness is preserved for the MVP demo.
## Handoff
- Runtime demo passed with current infrastructure.
- OpenSpec archive confirmation: requested by user after successful rerun.
@@ -0,0 +1,35 @@
# MVP Demo Trace Acceptance Brief
## Background
- User goal: make the MVP runnable, observable, and explainable for an Agent Engineer interview.
- Current problem: the system can execute diagnosis, but reviewers need a simple way to replay one session from final answer back to agent steps and tool evidence.
- Associated OpenSpec: `openspec/changes/mvp-demo-trace-acceptance/`
- Devflow scale: standard-light.
## Scope
- In scope:
- `mvp-demo` Spring profile overlay.
- `GET /api/diagnosis/{sessionId}/trace` read-only API.
- Trace aggregation DTO/service/controller.
- Focused service tests.
- Demo and acceptance documentation.
- Out of scope:
- Sensitive configuration cleanup.
- Full offline LLM/vector/database mock runtime.
- Database schema migration.
- Changes to chat execution, verifier routing, upload, or feedback behavior.
- Impact area:
- `src/main/java/com/superbiz/agent/controller`
- `src/main/java/com/superbiz/agent/service`
- `src/main/java/com/superbiz/agent/dto`
- `src/main/resources/application-mvp-demo.yml`
- `mvp/demo`
- `mvp/notes`
## OpenSpec Alignment
- proposal coverage: covered
- specs coverage: covered
- tasks coverage: covered
@@ -0,0 +1,87 @@
# MVP Demo Trace Acceptance Decisions
## Clarify
- Entry summary: continue the MVP toward a runnable and explainable demo by adding an `mvp-demo` profile, an end-to-end acceptance case, and a trace query API.
- Slug: `mvp-demo-trace-acceptance`
- Devflow scale: standard-light. The change adds a public read-only API and documentation, but does not alter core chat execution or persistence schemas.
## Context
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
## Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Terminology | Should "trace" mean persisted diagnosis execution evidence instead of transient frontend chat history? | evidence-driven | Resolved |
| Q2 | Boundary | Should this change modify chat execution or only expose existing persisted evidence? | evidence-driven | Resolved |
| Q3 | Acceptance | What proves the MVP flow is end-to-end enough for demo/interview use? | evidence-driven | Resolved |
| Q4 | Interface | What is the API impact level for `GET /api/diagnosis/{sessionId}/trace`? | evidence-driven | Resolved |
## Evidence-driven
| Conclusion | Evidence Source | Reported To User |
|---|---|---|
| Trace should aggregate persisted diagnosis evidence, not Redis-only chat history. | `DiagnosisSession`, `AgentStep`, `ToolInvocation` entities and repositories | Reported in progress update |
| Core chat execution does not need to change for this slice. | Existing unified chat path and SupervisorAgent commits; requested scope is demo/profile/trace/acceptance | Reported in progress update |
| End-to-end acceptance should cover start -> chat -> trace -> feedback. | `ChatController`, `FeedbackController`, traceable session id decision in MVP notes | Reported in progress update |
| Trace API is additive L3 because it is a new HTTP API for frontend/demo consumers. | sm-flow interface impact rules | Recorded in OpenSpec design |
## User-interview
| Question | User Words | Confirmation | OpenSpec Writeback |
|---|---|---|---|
| Should security/sensitive config cleanup be included? | "安全问题先不考虑"; "敏感配置先不做" | Confirmed | Non-goal |
| Should this be implemented under sm-flow? | "按照 sm-flow 的流程来实现吧" | Confirmed | This change follows sm-flow artifacts |
## Key Decisions
- Decision: Add a new trace API instead of embedding trace details in `/api/chat`.
- Reason: Chat execution and observability should stay decoupled.
- Impact: Demo can query trace after any successful chat request using the same session id.
- Risk accepted: Response shape is new and should be treated as demo-facing contract.
- Decision: Keep `mvp-demo` profile as configuration overlay, not a fully mocked standalone runtime.
- Reason: The current MVP still depends on real DB/Redis/Milvus/LLM for full chat execution; this change avoids inventing a fake runtime that hides integration behavior.
- Impact: Demo profile improves repeatability for logs/metrics, while docs remain explicit about required external services.
- Risk accepted: End-to-end acceptance may still require valid infrastructure and keys.
## Cross-Artifact Alignment
| Upstream -> Downstream | Check | Status |
|---|---|---|
| brief/prd -> proposal | Goal, scope, non-goals, and acceptance expectation are in proposal | Aligned |
| proposal -> design | Scope, constraints, and API impact are in design | Aligned |
| design -> specs/tasks | Trace DTO, controller/service, demo profile, and docs are represented | Aligned |
| specs -> tasks | Observable behavior is covered by executable tasks | Aligned |
## Architecture Audit
- Data path: HTTP trace request -> controller -> trace service -> repositories -> aggregate DTO -> `Result.success`.
- The service is read-only and does not mutate diagnosis, step, tool, or feedback state.
- No schema change is needed because all required fields already exist in `diagnosis_session`, `agent_step`, and `tool_invocation`.
- Main risk is response size for large sessions; MVP mitigates by returning previews already persisted by tools rather than raw external logs.
- The additive API is acceptable for MVP because old callers remain unaffected.
## Pre-apply Research
- Reference implementations read:
- `ChatController` for `/api` controller conventions.
- `FeedbackController` for simple API controller shape.
- `GlobalExceptionHandler` and `SessionNotFoundException` for 404 handling.
- `DiagnosisSessionRepository`, `AgentStepRepository`, `ToolInvocationRepository` for available queries.
- `DiagnosisSession`, `AgentStep`, `ToolInvocation` for fields.
- Impact analysis:
- `DiagnosisSessionRepository`: LOW, direct imports in service/controller paths.
- `AgentStepRepository`: HIGH because it participates in chat/AiOps flows. This change only consumes existing query methods and does not modify the repository.
- `ToolInvocationRepository`: LOW.
## Commit Gate
- OpenSpec proposal/design/specs/tasks exist.
- API impact: L3 additive collaboration API, documented in design and spec.
- User-confirmed non-goal: sensitive configuration cleanup remains out of scope.
- No unresolved user-interview questions remain for this slice.
@@ -0,0 +1,25 @@
# MVP Demo Trace Acceptance Evidence
## Evidence
| Source | Evidence | Conclusion | Reported |
|---|---|---|---|
| `DiagnosisSessionRepository` | Existing `findBySessionId(String)` query | Trace can locate the session without new repository methods | Yes |
| `AgentStepRepository` | Existing `findBySessionIdOrderByStepIndex(String)` query | Agent steps can be returned in execution order | Yes |
| `ToolInvocationRepository` | Existing `findBySessionIdOrderByIdAsc(String)` query | Tool evidence can be returned in persisted order | Yes |
| `GlobalExceptionHandler` | Handles `SessionNotFoundException` as HTTP 404 with `Result.error(404, ...)` | Missing trace can reuse existing error contract | Yes |
| `mvn -q "-Dtest=DiagnosisTraceServiceTest" test` | Command passed | Trace aggregation behavior is covered offline | Yes |
| `mvn -q -DskipTests compile` | Command passed | New code compiles with the full project | Yes |
| `gitnexus detect-changes --repo SuperBizAgent-java` | Command completed with `No changes detected` and line-ending warnings | Required GitNexus check ran; output likely does not capture newly added files | Yes |
## Evidence-driven Conclusions
- Conclusion: No database migration is required.
- Evidence: All trace fields are available from existing `diagnosis_session`, `agent_step`, and `tool_invocation` entities.
- Risk: Response shape becomes a new API contract.
- User confirmation: Not required; additive L3 API recorded in OpenSpec.
- Conclusion: Trace aggregation can be tested without external infrastructure.
- Evidence: `DiagnosisTraceServiceTest` uses mocked repositories and an `ObjectMapper`.
- Risk: Runtime integration still depends on configured infrastructure.
- User confirmation: Not required; limitation recorded in acceptance docs.
@@ -0,0 +1,14 @@
# Acceptance: aiops-alert-scope-control
## Verification
- [x] Payload-mode prompt focuses the final report on the supplied alert.
- [x] No-payload prompt requires active-alert discovery first.
- [x] Targeted tests pass.
- [x] Compile passes.
- [x] OpenSpec validates.
## Known Limits
- Prompt-only scope control may still require runtime observation.
- AIOps Verifier remains deferred.
@@ -0,0 +1,28 @@
# Brief: aiops-alert-scope-control
## Background
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
## Goal
Make AIOps scope explicit:
- Payload present -> targeted diagnosis for the supplied alert.
- Payload absent -> automatic active-alert discovery and diagnosis.
## Scope
- In scope:
- `AiOpsService.buildTaskPrompt(...)` scope rules.
- Focused tests.
- Demo acceptance wording.
- Out of scope:
- Verifier integration.
- Java-side filtering of tool results.
- API shape changes.
- Database changes.
## Related OpenSpec
`openspec/changes/aiops-alert-scope-control/`
@@ -0,0 +1,42 @@
# Decisions: aiops-alert-scope-control
## Clarify
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
- Slug: `aiops-alert-scope-control`
- Scale: standard-light.
## Context
- AIOps traceability is implemented and verified.
- Mock Prometheus returns multiple active alerts.
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
## GitNexus
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
## Key Decisions
- Payload mode is detected when any alert field is present.
- Payload mode final report must focus on the supplied alert.
- No-payload mode must first call `queryPrometheusAlerts`.
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
@@ -0,0 +1,44 @@
# Evidence: aiops-alert-scope-control
## Local Impact Analysis
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
- No controller, DTO, repository, or database changes are required.
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
## Verification Results
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
- `mvn -q -DskipTests compile` passed.
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
## Runtime Verification
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
- `diagnosis_session` persisted:
- `agent_flow = AI_OPS`
- `status = SUCCESS`
- `total_duration_ms = 69875`
- `step_count = 5`
- `tool_call_count = 8`
- Tool invocation counts:
- `query_metrics = 1`
- `lookup_knowledge = 1`
- `query_logs = 6`
- Report scope check:
- `告警根因分析 - HighCPUUsage` exists.
- `告警根因分析 - HighMemoryUsage` does not exist.
- `告警根因分析 - SlowResponse` does not exist.
- `相关风险告警` exists.
## Runtime Fix
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
- `maximum-pool-size: 5`
- `minimum-idle: 1`
- `connection-timeout: 10000`
- `validation-timeout: 5000`
- `idle-timeout: 60000`
- `max-lifetime: 120000`
- `keepalive-time: 30000`
@@ -0,0 +1,18 @@
# Acceptance: aiops-traceable-diagnosis-entry
## Verification
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
- [x] Targeted AIOps service tests pass.
- [x] Compile verification passes.
- [x] Demo docs describe AIOps request -> session id -> trace query.
## Result
Accepted for implementation scope.
## Known Limits
- AIOps Verifier integration is deferred.
- Runtime still depends on configured model and infrastructure.
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
@@ -0,0 +1,28 @@
# Brief: aiops-traceable-diagnosis-entry
## Background
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
## Goal
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
## Scope
- In scope:
- Optional AIOps alert request body.
- Stable request/session id propagation.
- Persisted AIOps query summary and final answer.
- SSE session id event.
- Demo documentation and focused tests.
- Out of scope:
- Full AIOps and ChatService unification.
- AIOps Verifier integration.
- Database schema changes.
- Sensitive configuration cleanup.
- Fully offline runtime.
## Related OpenSpec
`openspec/changes/aiops-traceable-diagnosis-entry/`
@@ -0,0 +1,68 @@
# Decisions: aiops-traceable-diagnosis-entry
## Clarify
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
- Slug: `aiops-traceable-diagnosis-entry`
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
## Context
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
## User-Interview Confirmations
| Topic | User Words | Decision |
|---|---|---|
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
## Key Decisions
- Keep `/api/ai_ops` and make its body optional.
- Emit `SseMessage.type=session` before long-running analysis starts.
- Store AIOps request summary in `diagnosis_session.query`.
- Store final report in `diagnosis_session.answer`.
- Defer AIOps Verifier to a later change so this slice stays focused.
## Architecture Audit
```text
POST /api/ai_ops
-> optional AIOpsRequest
-> resolve sessionId
-> create diagnosis_session(agentFlow=AI_OPS)
-> run ai_ops_supervisor(planner, executor)
-> AgentLoggingHook persists steps
-> tools persist invocations under SessionContextHolder
-> extract final report
-> persist answer
-> GET /api/diagnosis/{sessionId}/trace replays the run
```
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
@@ -0,0 +1,37 @@
# Evidence: aiops-traceable-diagnosis-entry
## Local Impact Analysis
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
## GitNexus
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
## Expected Verification
- Focused unit tests for AIOps request/session/report helper behavior.
- Compile verification.
- OpenSpec validation if CLI is available.
## Verification Results
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
- `mvn -q -DskipTests compile`: passed.
## Demo Alignment
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
## Metric Alignment Follow-up
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
- Targeted verification:
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
- `mvn -q -DskipTests compile`: passed.
@@ -0,0 +1,36 @@
# Acceptance: diagnosis-eval-harness
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | `ISS-006` and initial OpenSpec artifacts were created. |
| Implementation | Done | Added fixed cases, fixture-mode trace evaluation, aggregate metrics, and JSON / Markdown report writer. |
| Verification | Done | Targeted evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- First implementation uses fixture-mode evaluation.
- Live trace API polling remains a follow-up option.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed
- Notes: Covers fixed case loading, fixture evaluation, missing fixture reporting, reject degraded-output validation, and report writing.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate diagnosis-eval-harness --strict`
- Result: passed
@@ -0,0 +1,31 @@
# Brief: diagnosis-eval-harness
## Background
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
## Goals
1. Define fixed diagnosis cases for the MVP demo domain.
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
3. Produce JSON and Markdown reports for interview and regression use.
4. Keep the first version offline by supporting trace fixtures.
## Scope
- Evaluation case definitions
- Trace fixture shape
- Rule-based evaluator
- JSON / Markdown report output
- Focused offline tests and docs
## Non-Goals
- No LLM-as-judge
- No live end-to-end runtime requirement
- No production API
- No chat or verifier runtime change
## Related OpenSpec
`openspec/changes/diagnosis-eval-harness/`
@@ -0,0 +1,28 @@
# Diagnosis Eval Harness Decisions
## Clarify
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
- Slug: `diagnosis-eval-harness`
- Devflow scale: standard-light
## Context
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
## Key Decisions
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
- Reason: The first regression signal should be deterministic and tied to trace contracts.
- Decision: Support offline fixture traces first.
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
- Decision: Output both JSON and Markdown.
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
## Open Questions
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
@@ -0,0 +1,10 @@
# Diagnosis Eval Harness Evidence
## Evidence
| Source | Evidence | Conclusion | Reported |
|---|---|---|---|
| `openspec/specs/evidence-trace-hardening/spec.md` | Defines stable evidence states and summary behavior | Evaluation can rely on trace semantics rather than ad hoc log parsing | Yes |
| `mvp/demo/README.md` | Documents an end-to-end demo flow with chat, trace, and feedback | Existing demo flow provides the runtime story, but not a reusable evaluation baseline | Yes |
| `DiagnosisTraceService` | Aggregates session, steps, tools, and self-evaluation | Trace response shape can be reused as evaluation input | Yes |
| `ToolTraceSummaryService` | Builds verifier-facing evidence summaries from persisted tool rows | Evaluator can check evidence coverage through persisted trace artifacts | Yes |
@@ -0,0 +1,32 @@
# Acceptance: evidence-trace-hardening
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | `ISS-005` and the initial OpenSpec artifacts were created. |
| Implementation | Done | Recorder contract, lookup persistence path, evidence summary semantics, and degraded-path tests were implemented. |
| Verification | Done | Targeted offline tests and compile verification passed. |
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,LookupKnowledgeToolTest" test`
- Result: passed
- Notes: Covers recorder contract, summary semantics for success/failure/no-evidence, and `ChatService` fallback / degraded paths.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
## Open Questions
| Question | Current position |
| --- | --- |
| Should deduped retrievals be counted separately from generic no-hit events in future evaluation metrics? | Deferred to P1-B; this change preserves enough structure to decide later. |
@@ -0,0 +1,32 @@
# Brief: evidence-trace-hardening
## Background
The MVP already has persisted tool traces and a verifier, but the evidence contract is still only partially standardized. For interview-focused hardening, the project now needs a tighter contract for evidence persistence, no-evidence / failure semantics, and degraded-output behavior.
## Goals
1. Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
2. Make verifier-facing summaries distinguish failed calls, no-hit calls, deduped retrievals, and actual supporting evidence.
3. Add offline tests for verifier fallback and degraded-output paths.
## Scope
- `ToolInvocationRecorder`
- `LookupKnowledgeTool`
- `QueryLogsTools`
- `QueryMetricsTools`
- `ToolTraceSummaryService`
- `ChatService`
- Focused offline tests
## Non-Goals
- No new API or schema
- No evaluation harness yet
- No trace UI
- No security/config cleanup
## Related OpenSpec
`openspec/changes/evidence-trace-hardening/`
@@ -0,0 +1,24 @@
# Evidence Trace Hardening Decisions
## Clarify
- Entry summary: harden the MVP evidence contract before building the P1-B evaluation harness.
- Slug: `evidence-trace-hardening`
- Devflow scale: standard-light
## Context
- `ISS-003` raised verifier traceability and failure-path concerns.
- Current code inspection shows `QueryLogsTools` and `QueryMetricsTools` already use `ToolInvocationRecorder`, while `LookupKnowledgeTool` still persists rows through a local helper.
- `ChatService` already contains fallback behavior for missing/invalid `verifier_output`, but coverage is narrow.
## Key Decisions
- Decision: Treat this as a contract-hardening change, not a new feature change.
- Reason: The project already has the necessary runtime pieces; the gap is semantic consistency and testability.
- Decision: Keep the scope before P1-B.
- Reason: The evaluation harness will rely on stable evidence semantics, so this contract slice should land first.
- Decision: Preserve schema and API stability.
- Reason: The interview value here is engineering rigor, not more surface area.
@@ -0,0 +1,11 @@
# Evidence Trace Hardening Evidence
## Evidence
| Source | Evidence | Conclusion | Reported |
|---|---|---|---|
| `ToolInvocationRecorder` | Provides a common persistence seam for evidence tools | Contract hardening should build on the existing recorder instead of introducing a new store path | Yes |
| `LookupKnowledgeTool` | Still constructs `ToolInvocation` rows through a local helper | Retrieval-aware evidence persistence is not yet unified with the recorder contract | Yes |
| `QueryLogsTools` / `QueryMetricsTools` | Already record evidence invocations through `recordEvidenceTool(...)` | Current gap is semantic alignment, not missing persistence | Yes |
| `ToolTraceSummaryService` | Merges rows by tool and topic domain and infers evidence level heuristically | Summary rules need explicit handling for failure, no-hit, and dedup cases | Yes |
| `ChatService` | Falls back to `LOW_CONFID` when verifier output is missing or invalid | These degraded paths exist and should now be covered by focused offline tests | Yes |
@@ -0,0 +1,37 @@
# Acceptance: expand-diagnosis-eval-fixtures
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
| Implementation | Done | Added remaining fixtures, full baseline reports, and documentation updates. |
| Verification | Done | Evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- Fixture coverage is complete for the five fixed diagnosis cases.
- Baseline reports are saved under `mvp/eval/reports`.
- No production runtime behavior has been changed.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed
- Notes: Covers full fixture coverage, baseline report matching, reject degraded-output validation, and report writing.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate expand-diagnosis-eval-fixtures --strict`
- Result: passed
@@ -0,0 +1,31 @@
# Brief: expand-diagnosis-eval-fixtures
## Background
The diagnosis eval harness is implemented and archived, but the fixed baseline is incomplete because three of the five diagnosis cases still reference missing fixtures.
## Goals
1. Add representative trace fixtures for all remaining fixed diagnosis cases.
2. Save a reproducible baseline report in JSON and Markdown.
3. Document how to regenerate and interpret the baseline.
4. Keep evaluation offline and deterministic.
## Scope
- Redis timeout fixture
- Slow response fixture
- JVM memory risk fixture
- Baseline reports under `mvp/eval/reports`
- Focused tests for full fixture coverage and report generation
## Non-Goals
- No new diagnosis cases
- No production Agent runtime changes
- No LLM-as-judge
- No live infrastructure requirement
## Related OpenSpec
`openspec/changes/expand-diagnosis-eval-fixtures/`
@@ -0,0 +1,28 @@
# Expand Diagnosis Eval Fixtures Decisions
## Clarify
- Entry summary: complete the fixed diagnosis eval baseline after the harness is in place.
- Slug: `expand-diagnosis-eval-fixtures`
- Devflow scale: standard-light
## Context
- `diagnosis-eval-harness` created the evaluator, case file, fixture mode, and report writer.
- The first baseline still has missing fixtures by design.
- This follow-up turns that partial baseline into a full fixed-case baseline.
## Key Decisions
- Decision: Keep this change data-focused.
- Reason: the evaluator rules already landed; this change should not blur fixture expansion with harness behavior changes.
- Decision: Save baseline reports in the repository.
- Reason: interview review and future diffs are easier when the expected baseline is visible.
- Decision: Use deterministic fixture traces instead of live trace generation.
- Reason: this baseline should run without infrastructure or external model calls.
## Open Questions
- Whether a future change should add a CLI or Maven goal for report regeneration.
@@ -0,0 +1,11 @@
# Evidence: expand-diagnosis-eval-fixtures
## Evidence Log
- 2026-07-04: Created slug-based issue `expand-diagnosis-eval-fixtures.md`.
- 2026-07-04: Created OpenSpec change `expand-diagnosis-eval-fixtures`.
- 2026-07-04: Added Redis timeout, slow response, and JVM memory risk fixtures.
- 2026-07-04: Added baseline JSON and Markdown reports under `mvp/eval/reports`.
- 2026-07-04: Verification passed with `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`.
- 2026-07-04: Verification passed with `mvn -q -DskipTests compile`.
- 2026-07-04: Verification passed with `openspec validate expand-diagnosis-eval-fixtures --strict`.
@@ -0,0 +1,37 @@
# Acceptance: diagnosis-eval-baseline-diff
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
- JSON and Markdown diff output are available.
- No production runtime behavior has been changed.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
- Result: passed
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
- Result: passed
@@ -0,0 +1,30 @@
# Brief: diagnosis-eval-baseline-diff
## Background
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
## Goals
1. Compare baseline and current `DiagnosisEvalReport` objects.
2. Detect aggregate and per-case regressions.
3. Output JSON and Markdown diff reports.
4. Document how to read the diff in interview and engineering terms.
## Scope
- Diff data structures
- Deterministic report comparison
- JSON / Markdown diff output
- Focused tests and eval docs
## Non-Goals
- No live Agent execution
- No LLM-as-judge
- No evaluator scoring rule changes
- No production API changes
## Related OpenSpec
`openspec/changes/diagnosis-eval-baseline-diff/`
@@ -0,0 +1,28 @@
# Diagnosis Eval Baseline Diff Decisions
## Clarify
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
- Slug: `diagnosis-eval-baseline-diff`
- Devflow scale: standard-light
## Context
- `diagnosis-eval-harness` created deterministic fixture evaluation.
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
- This change compares new reports against that baseline.
## Key Decisions
- Decision: Diff report DTOs instead of raw traces.
- Reason: the report is the stable contract for regression review.
- Decision: Use deterministic code rules instead of LLM-as-judge.
- Reason: baseline regression checks should be repeatable and explainable.
- Decision: Output both JSON and Markdown.
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
## Open Questions
- Whether a future change should expose this through a CLI or Maven goal.
@@ -0,0 +1,11 @@
# Evidence: diagnosis-eval-baseline-diff
## Evidence Log
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
@@ -0,0 +1,25 @@
# Acceptance: mvp-demo-interview-runbook
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
| Verification | Done | OpenSpec validation passed. |
## Current State
- No backend runtime behavior has been changed.
- Demo is packaged under `mvp/demo` for interview use.
## Verification
### OpenSpec Verification
- Command: `openspec validate mvp-demo-interview-runbook --strict`
- Result: passed
@@ -0,0 +1,29 @@
# Brief: mvp-demo-interview-runbook
## Background
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
## Goals
1. Provide a fixed payment-timeout request payload.
2. Provide a PowerShell script that runs chat, trace, and feedback.
3. Save demo responses under `mvp/demo/output`.
4. Add interview walkthrough and trace checklist.
## Scope
- Demo docs and scripts only
- Existing local APIs only
- Existing `mvp-demo` profile only
## Non-Goals
- No backend code changes
- No eval extension
- No secret cleanup
- No full offline runtime
## Related OpenSpec
`openspec/changes/mvp-demo-interview-runbook/`
@@ -0,0 +1,27 @@
# MVP Demo Interview Runbook Decisions
## Clarify
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
- Slug: `mvp-demo-interview-runbook`
- Devflow scale: standard-light
## Context
- Evidence trace and eval baseline work are already done.
- The next useful step is not more eval tooling, but a runnable demo path.
## Key Decisions
- Decision: Keep this change documentation/script-only.
- Reason: Plan C is about demo packaging, not new runtime capability.
- Decision: Use a stable session id.
- Reason: it makes trace lookup and saved output predictable.
- Decision: Save outputs to `mvp/demo/output`.
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
## Open Questions
- Whether a later change should add a truly offline stubbed demo mode.
@@ -0,0 +1,9 @@
# Evidence: mvp-demo-interview-runbook
## Evidence Log
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
- 2026-07-05: Added fixed payment-timeout request payload.
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
+86
View File
@@ -0,0 +1,86 @@
# RAG Retrieval Baseline
This directory contains the offline retrieval baseline for the RAG refactor.
The baseline is intentionally narrower than full diagnosis evaluation. It checks
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
evidence keywords before changing L0 behavior, query augmentation, evidence
post-processing, or Spring AI VectorStore integration.
## Layout
```text
eval/rag-retrieval/
cases/golden-cases.json Fixed retrieval golden cases
fixtures/*.json Saved retrieval candidates for each case
reports/baseline.json Machine-readable baseline report
reports/baseline.md Human-readable baseline report
reports/live-post-reindex.* Optional live acceptance reports
```
## Run
From the repository root:
```bash
python scripts/eval_rag_retrieval.py
```
Custom paths are also supported:
```bash
python scripts/eval_rag_retrieval.py \
--cases eval/rag-retrieval/cases/golden-cases.json \
--fixtures eval/rag-retrieval/fixtures \
--json-report eval/rag-retrieval/reports/baseline.json \
--markdown-report eval/rag-retrieval/reports/baseline.md
```
## Hit Levels
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
- `weak`: expected evidence keyword is found, but expected document is missing.
- `miss`: expected document and expected evidence are not found.
`Recall@K` counts `strong` and `medium` as retrieved.
## Scope
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
or the Spring Boot application. It is a regression harness for retrieval behavior,
not a claim that live production retrieval accuracy is complete.
## Live Post-Reindex Acceptance
When embedding input changes, existing vectors do not update by themselves. For
example, after adding `title` and `breadcrumb` to the embedding text, the live
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
semantic signal.
Use this optional live acceptance flow after the application is running and the
knowledge base has been reindexed:
```bash
python scripts/eval_rag_live_acceptance.py
```
Custom service URL and output paths are supported:
```bash
python scripts/eval_rag_live_acceptance.py \
--base-url http://127.0.0.1:9900 \
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
```
The script calls:
```text
GET /api/search/similar
```
It writes JSON and Markdown reports with query, topK, result count, top
candidates, breadcrumb, score labels, and raw response fields. This is a live
smoke check for environment readiness and post-reindex behavior; it does not
replace the deterministic offline baseline above.
@@ -0,0 +1,61 @@
{
"version": 1,
"description": "Offline golden retrieval cases for RAG refactor baseline.",
"topK": 5,
"cases": [
{
"caseId": "chat-mysql-connection-pool",
"scenario": "chat",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"expectedDocIds": ["mysql-connection-pool"],
"expectedBreadcrumbs": ["Database > MySQL > Connection Pool"],
"expectedKeywords": ["connection pool", "max_connections", "HikariCP"],
"notes": "Covers precise database troubleshooting retrieval."
},
{
"caseId": "chat-diagnosis-flow",
"scenario": "chat",
"query": "What is the standard troubleshooting flow for an application incident?",
"expectedDocIds": ["incident-diagnosis-flow"],
"expectedBreadcrumbs": ["AIOps > Diagnosis Flow"],
"expectedKeywords": ["collect evidence", "verify", "remediation"],
"notes": "Covers process-style knowledge where breadcrumb matters."
},
{
"caseId": "aiops-payment-latency-alert",
"scenario": "aiops",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"expectedDocIds": ["payment-service-latency"],
"expectedBreadcrumbs": ["AIOps > Service Alerts > Payment Latency"],
"expectedKeywords": ["p95 latency", "payment-service", "downstream dependency"],
"notes": "Covers alert payload terms that should become retrieval hints."
},
{
"caseId": "aiops-prometheus-alert-scope",
"scenario": "aiops",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"expectedDocIds": ["aiops-alert-scope-control"],
"expectedBreadcrumbs": ["AIOps > Alert Scope Control"],
"expectedKeywords": ["payload", "unrelated active alerts", "scope"],
"notes": "Covers scoped alert diagnosis behavior."
},
{
"caseId": "chat-rag-chunk-context",
"scenario": "chat",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"expectedDocIds": ["rag-chunk-context-reconstruction"],
"expectedBreadcrumbs": ["RAG > Chunking > Context Reconstruction"],
"expectedKeywords": ["neighbor chunk", "same section", "breadcrumb"],
"notes": "Covers the known RAG refactor issue around context reconstruction."
},
{
"caseId": "chat-l0-domain-hint",
"scenario": "chat",
"query": "Should L0 keyword matching decide the final retrieval result?",
"expectedDocIds": ["rag-l0-domain-entity-hint"],
"expectedBreadcrumbs": ["RAG > L0 > Domain Entity Hint"],
"expectedKeywords": ["domain detector", "entity extractor", "metadata filter"],
"notes": "Covers the target L0 role after refactor."
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "aiops-payment-latency-alert",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "payment-service-latency",
"title": "Payment Service Latency Alert Playbook",
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
"score": 0.84,
"retrievalLayer": "L1"
},
{
"rank": 2,
"docId": "mysql-connection-pool",
"title": "MySQL Connection Pool Troubleshooting",
"breadcrumb": "Database > MySQL > Connection Pool",
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
"score": 0.68,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,16 @@
{
"caseId": "aiops-prometheus-alert-scope",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "aiops-alert-scope-control",
"title": "AIOps Alert Scope Control",
"breadcrumb": "AIOps > Alert Scope Control",
"content": "When payload mode is active, queryPrometheusAlerts can verify the supplied alert, but unrelated active alerts must remain scoped context and should not become full diagnoses.",
"score": 0.9,
"retrievalLayer": "L0+L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-diagnosis-flow",
"query": "What is the standard troubleshooting flow for an application incident?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "incident-diagnosis-flow",
"title": "Incident Diagnosis Flow",
"breadcrumb": "AIOps > Diagnosis Flow",
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
"score": 0.82,
"retrievalLayer": "L1"
},
{
"rank": 2,
"docId": "rag-chunk-context-reconstruction",
"title": "RAG Chunk Context Reconstruction",
"breadcrumb": "RAG > Chunking > Context Reconstruction",
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
"score": 0.55,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-l0-domain-hint",
"query": "Should L0 keyword matching decide the final retrieval result?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "rag-l0-domain-entity-hint",
"title": "RAG L0 Domain Entity Hint",
"breadcrumb": "RAG > L0 > Domain Entity Hint",
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
"score": 0.88,
"retrievalLayer": "L0"
},
{
"rank": 2,
"docId": "rag-l0-l1-fusion-ranking",
"title": "RAG L0 L1 Fusion Ranking",
"breadcrumb": "RAG > Ranking > Fusion",
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
"score": 0.75,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-mysql-connection-pool",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "mysql-connection-pool",
"title": "MySQL Connection Pool Troubleshooting",
"breadcrumb": "Database > MySQL > Connection Pool",
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
"score": 0.86,
"retrievalLayer": "L0+L1"
},
{
"rank": 2,
"docId": "incident-diagnosis-flow",
"title": "Incident Diagnosis Flow",
"breadcrumb": "AIOps > Diagnosis Flow",
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
"score": 0.61,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-rag-chunk-context",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "rag-chunk-context-reconstruction",
"title": "RAG Chunk Context Reconstruction",
"breadcrumb": "RAG > Chunking > Context Reconstruction",
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
"score": 0.79,
"retrievalLayer": "L1"
},
{
"rank": 2,
"docId": "rag-breadcrumb-embedding-gap",
"title": "RAG Breadcrumb Embedding Gap",
"breadcrumb": "RAG > Embedding > Breadcrumb",
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
"score": 0.72,
"retrievalLayer": "L1"
}
]
}
+1
View File
@@ -0,0 +1 @@
+131
View File
@@ -0,0 +1,131 @@
{
"generatedAt": "2026-07-04T17:59:52.172759+00:00",
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
"fixtureDir": "eval/rag-retrieval/fixtures",
"aggregate": {
"caseCount": 6,
"topK": 5,
"strongHitCount": 6,
"mediumHitCount": 0,
"weakHitCount": 0,
"missCount": 0,
"recallAtK": 1.0,
"strongHitRate": 1.0,
"averageFirstHitRank": 1.0
},
"results": [
{
"caseId": "chat-mysql-connection-pool",
"scenario": "chat",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:mysql-connection-pool",
"2:incident-diagnosis-flow"
],
"matchedKeywords": [
"connection pool",
"max_connections",
"hikaricp"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "chat-diagnosis-flow",
"scenario": "chat",
"query": "What is the standard troubleshooting flow for an application incident?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:incident-diagnosis-flow",
"2:rag-chunk-context-reconstruction"
],
"matchedKeywords": [
"collect evidence",
"verify",
"remediation"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "aiops-payment-latency-alert",
"scenario": "aiops",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:payment-service-latency",
"2:mysql-connection-pool"
],
"matchedKeywords": [
"p95 latency",
"payment-service",
"downstream dependency"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "aiops-prometheus-alert-scope",
"scenario": "aiops",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:aiops-alert-scope-control"
],
"matchedKeywords": [
"payload",
"unrelated active alerts",
"scope"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "chat-rag-chunk-context",
"scenario": "chat",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:rag-chunk-context-reconstruction",
"2:rag-breadcrumb-embedding-gap"
],
"matchedKeywords": [
"neighbor chunk",
"same section",
"breadcrumb"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "chat-l0-domain-hint",
"scenario": "chat",
"query": "Should L0 keyword matching decide the final retrieval result?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:rag-l0-domain-entity-hint",
"2:rag-l0-l1-fusion-ranking"
],
"matchedKeywords": [
"domain detector",
"entity extractor",
"metadata filter"
],
"breadcrumbMatched": true,
"failedChecks": []
}
]
}
+28
View File
@@ -0,0 +1,28 @@
# RAG Retrieval Baseline
Generated at: `2026-07-04T17:59:52.172759+00:00`
## Aggregate
| Metric | Value |
|---|---:|
| Cases | 6 |
| Top K | 5 |
| Recall@K | 1.0 |
| Strong hit rate | 1.0 |
| Strong hits | 6 |
| Medium hits | 0 |
| Weak hits | 0 |
| Misses | 0 |
| Average first hit rank | 1.0 |
## Cases
| Case | Scenario | Hit | First Expected Rank | Top Candidates | Failed Checks |
|---|---|---|---:|---|---|
| chat-mysql-connection-pool | chat | strong | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
| chat-diagnosis-flow | chat | strong | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
| aiops-payment-latency-alert | aiops | strong | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
| aiops-prometheus-alert-scope | aiops | strong | 1 | 1:aiops-alert-scope-control | |
| chat-rag-chunk-context | chat | strong | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
| chat-l0-domain-hint | chat | strong | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
+66
View File
@@ -0,0 +1,66 @@
# SuperBizAgent 面试资料包
## 一句话定位
SuperBizAgent 是一个面向企业故障诊断场景的 Agent 工程项目。它把用户问题或 AIOps 告警转换成可追踪的 Agent 执行链路,并把工具证据、模型步骤、最终答案、自评估和用户反馈统一沉淀到诊断 Trace 中。
## 面试重点
- **Agent 编排**:Chat 复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 `Supervisor -> Planner / Executor`。
- **工具证据链**:知识库、日志、指标、Prometheus 告警都通过显式工具调用进入链路,并记录到 `tool_invocation`。
- **可追踪诊断**:一次诊断对应一个 `sessionId`,可通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
- **质量门禁**:Chat Verifier 校验 groundedness;AIOps 规则评估检查报告完整性、payload 聚焦和证据工具覆盖。
- **RAG 工程化**:`lookup_knowledge` 是显式 Agent Tool,底层通过 Spring AI VectorStore 主路径 + Milvus SDK fallback。
- **反馈闭环**:用户反馈 `useful` 会沉淀 `case_library`,`not_useful` 保留 bad case 信号。
## 推荐阅读顺序
1. `mvp/architecture/interview-one-pager.md`:一页式架构图和 2-5 分钟讲解。
2. `mvp/demo/ten-minute-interview-demo.md`:10 分钟现场演示脚本。
3. `interview/story-cases.md`:可复用的面试故事案例。
4. `interview/architecture.md`:面试版系统架构。
5. `interview/design-tradeoffs.md`:关键设计取舍。
6. `interview/demo-script.md`:更细的命令式演示脚本。
7. `interview/acceptance-checklist.md`:面试前验收清单。
8. RAG 专题文档:`rag-refactor-story.md`、`rag-vectorstore-interview-notes.md`、`rag-retrieval-quality-report.md`。
## 核心演示链路
### Chat 诊断
```text
POST /api/chat
-> ChatService
-> Planner -> Executor -> Verifier
-> lookup_knowledge / query_logs / query_metrics
-> diagnosis_session + agent_step + tool_invocation
-> GET /api/diagnosis/{sessionId}/trace
-> POST /api/feedback
```
### AIOps 告警诊断
```text
POST /api/ai_ops
-> AiOpsService
-> PAYLOAD_TARGETED / AUTO_DISCOVERY
-> ai_ops_supervisor
-> planner_agent / executor_agent
-> queryPrometheusAlerts + logs + metrics + lookup_knowledge
-> alert report
-> aiops_rule_evaluation
-> GET /api/diagnosis/{sessionId}/trace
```
## 当前完成度
- Chat 诊断链路:可运行、可追踪、有 Verifier。
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
- RAG 检索链路:Spring AI VectorStore 主路径、Milvus SDK fallback、L0 hint、检索评测 baseline。
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
- Demo 材料:`mvp/demo/README.md`、`mvp/demo/ten-minute-interview-demo.md`。
## 主叙事
这个项目不是简单调用大模型,而是在做一个可审计、可验证、可回归的 Agent 诊断系统。模型可以规划和推理,但每一步工具证据、最终结论、Verifier 结果和用户反馈都能被 Trace API 回放。面试时重点展示“从问题到证据到答案到验证再到反馈”的闭环。
+166
View File
@@ -0,0 +1,166 @@
# 面试前验收清单
## 1. 环境检查
- 当前分支包含最新架构文档和面试材料。
- MySQL 可连接。
- Redis 可连接。
- Milvus/Zilliz 可连接。
- 模型 API key 可用。
- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。
启动服务:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
编译检查:
```powershell
mvn -q -DskipTests compile
```
目标测试:
```powershell
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest,VectorSearchServiceTest,LookupKnowledgeToolTest" test
```
## 2. Chat Demo 验收
请求:
```powershell
$sessionId = "interview-chat-payment-timeout-001"
$body = @{
Id = $sessionId
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
验收:
- 返回 `data.success = true`。
- 返回 `data.sessionId = interview-chat-payment-timeout-001`。
- `diagnosis_session.agent_flow = CHAT`。
- Trace API 返回 session、steps、toolInvocations。
- 复杂问题下 trace 中能看到 verifier 相关数据。
SQL:
```powershell
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'"
```
## 3. AIOps Demo 验收
请求:
```powershell
$aiopsSessionId = "interview-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
验收:
- SSE 首条包含 `type=session`。
- SSE 最后包含 `type=done`。
- `diagnosis_session.agent_flow = AI_OPS`。
- `diagnosis_session.status = SUCCESS`。
- `diagnosis_session.answer` 有最终报告。
- Trace API 返回 AIOps steps 和 tool invocations。
- 报告主章节聚焦 `HighCPUUsage/payment-service`。
- 其他 active alerts 不应展开成独立主根因章节。
- `self_evaluation.aiops_rule_evaluation` 存在。
SQL:
```powershell
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, total_duration_ms, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
```
```powershell
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name ORDER BY tool_name"
```
Scope 检查:
```powershell
python scripts/query_mysql.py "SELECT (answer LIKE '%HighCPUUsage%') AS has_main_alert, (answer LIKE '%payment-service%') AS has_service FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
```
## 4. Trace API 验收
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
```
如果 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
```powershell
curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
```
## 5. RAG 验收
```powershell
Invoke-RestMethod `
-Uri "http://127.0.0.1:9900/api/search/similar?query=ERR_TIMEOUT&topK=3" `
-Method Get
```
验收:
- 返回 `code = 200`。
- top candidates 中包含 `ERR_TIMEOUT` 相关文档。
- `scoreLabel` 能体现当前检索路径语义。
- 如果走 VectorStore,日志应出现 Spring AI VectorStore search。
## 6. 常见问题
### MySQL stale connection
现象:
```text
HikariPool - Connection is not available
No operations allowed after connection closed
```
处理:
- 重启服务。
- 确认 HikariPool 使用当前配置启动成功。
- 再跑 trace 或 AIOps 请求。
### SSE 客户端显示异常
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 Trace API 验证结果。
### OpenSpec 全量校验失败
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试演示主要依赖已归档 spec、MVP trace 和 RAG 验收材料,可以单独验证相关 spec。
+74
View File
@@ -0,0 +1,74 @@
# AIOps 轻量规则验证器
## 1. 改动是什么
AIOps 现在有一个确定性的后置质量门禁。最终告警报告持久化后,`AiOpsRuleEvaluationService` 会检查:
- 最终报告是否存在,且不是明显过短。
- payload 模式下,报告是否提到输入的告警和服务。
- 是否有证据工具调用,例如 `lookup_knowledge`、`query_metrics`、`query_logs`。
结果写入:
```text
diagnosis_session.self_evaluation.aiops_rule_evaluation
```
Trace API 会通过 session self-evaluation 展示这个结果。
## 2. 为什么先做规则型
这还不是完整 LLM Verifier。
AIOps 第一阶段质量风险比较具体,适合先用规则:
- 报告有没有生成。
- 报告有没有聚焦 payload。
- 有没有使用证据工具。
- 有没有把无关告警展开成主诊断对象。
规则验证稳定、便宜、容易解释,也不会在当前链路里额外引入一次隐藏模型调用。
## 3. 判定结果
当前评估器输出:
```text
PASS
WARN
FAIL
```
含义:
- `PASS`:核心检查通过。
- `WARN`:报告存在,但可能缺少 payload 关键词或证据工具。
- `FAIL`:缺少最终报告、报告过短等关键问题。
缺少 payload 关键词或证据工具先给 `WARN`,因为 demo/mock 环境下证据可能不可用,且报告措辞可能与 payload 字段不完全一致。
## 4. 面试回答
如果被问:为什么 AIOps 也需要验证器?
```text
Chat 已经有 LLM Verifier,因为用户问题开放度高。
AIOps 的第一阶段质量风险更明确:报告是否聚焦输入告警、是否使用证据工具、报告是否完整。
所以我先做了轻量规则验证器,把结果写入 self_evaluation,让 Trace 不只展示 Agent 做了什么,也展示输出是否通过基础质量门。
```
如果被问:为什么不直接复用 Chat Verifier?
```text
AIOps 验证语义和 Chat 不一样。
它要检查 alert scope、payload focus、证据工具覆盖,以及是否过度展开无关 active alerts。
直接复用 Chat Verifier 会混淆这些语义。
规则评估先提供稳定质量门,后续 AIOps LLM Verifier 可以基于同一套 trace contract 扩展。
```
## 5. 后续增强
- 引入 AIOps LLM Verifier,逐条校验根因和建议是否有 evidence refs。
- 把 rule evaluation 的 checks 在 Trace API 中结构化展示。
- 将 payload scope violation 沉淀为 bad case。
+75
View File
@@ -0,0 +1,75 @@
# AIOps 查询增强说明
## 1. 改动是什么
AIOps 在 `PAYLOAD_TARGETED` 模式下,会从告警 payload 中稳定生成一条推荐知识库检索 query。
参与拼接的非空字段:
```text
alertName service severity description timeRange userRequest
```
示例:
```text
HighCPUUsage payment-service P1 CPU 使用率超过 80% last_15m
```
最终会进入 Prompt:
```text
Recommended lookup_knowledge query: ...
```
## 2. 为什么重要
AIOps payload 里包含高价值检索词:
- 告警名称。
- 服务名。
- 严重等级。
- 症状描述。
- 时间范围。
- 用户补充请求。
如果完全让 Agent 从长 Prompt 里自己组织检索 query,可能遗漏服务名或告警名。推荐 query 让检索种子更稳定。
## 3. 设计取舍
这是 Prompt 层 query augmentation,不是隐藏检索。
我没有在 Agent 运行前自动调用 `lookup_knowledge`,原因是项目强调可追踪性:工具调用应该由 Agent 显式发起,并记录到 `tool_invocation`。
当前设计:
```text
AIOps payload
-> deterministic recommended retrieval query
-> Agent prompt
-> Agent 显式调用 lookup_knowledge
-> tool_invocation 记录真实检索行为
```
## 4. 面试回答
如果被问:AIOps payload 怎么提升 RAG 检索?
```text
我没有把告警 payload 粗暴替换成一个宽泛领域,而是提取 alertName、service、severity、description、timeRange 等高信号字段,拼成推荐的 lookup_knowledge query。
Agent 仍然显式调用工具,所以 trace 仍然能看到真实检索行为,但 query 不再完全依赖模型临场发挥。
```
如果被问:为什么不自动检索?
```text
自动检索会在 Agent 真正决策前制造一份隐藏证据。
这个项目的重点是可观测 Agent 执行,所以我选择 Prompt 层增强:给 Agent 一个更好的 query seed,但不改变工具调用必须显式可追踪的契约。
```
## 5. 后续增强
- 将 recommended query 写入 trace 的结构化字段,便于对比 Agent 实际 query。
- 对 payload 字段加权,例如 alertName/service 权重大于 timeRange。
- 后续接入 Query Transformer 时,保留原始 query、推荐 query、改写 query 三者的可追踪关系。
+97
View File
@@ -0,0 +1,97 @@
# 面试版系统架构
## 1. 系统分层
```mermaid
flowchart TB
API["API 层\nChatController / DiagnosisTraceController / SearchController"] --> Service["应用服务层\nChatService / AiOpsService / DiagnosisTraceService"]
Service --> Agent["Agent 编排层\nPlanner / Executor / Verifier / Supervisor"]
Agent --> Tools["工具层\nlookup_knowledge / query_logs / query_metrics / Prometheus"]
Tools --> RAG["RAG 检索\nL0 hint + VectorSearchService"]
RAG --> VectorStore["Spring AI VectorStore"]
RAG --> SDK["Milvus SDK fallback"]
Agent --> Trace["Trace 持久化"]
Tools --> Trace
Trace --> Session["diagnosis_session"]
Trace --> Step["agent_step"]
Trace --> Invocation["tool_invocation"]
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
Step --> TraceAPI
Invocation --> TraceAPI
```
## 2. Chat 链路
```mermaid
flowchart TD
User["用户问题"] --> ChatAPI["POST /api/chat"]
ChatAPI --> Strategy["ChatService.executeChatWithStrategy"]
Strategy --> Complexity{"复杂问题?"}
Complexity -->|否| Single["单 ReactAgent 快速回答"]
Complexity -->|是| Planner["chat_planner"]
Planner --> Executor["chat_executor"]
Executor --> Tools["证据工具"]
Tools --> Executor
Executor --> Verifier["chat_verifier"]
Verifier --> Decision{"PASS / LOW_CONFID / REJECT"}
Decision --> Answer["最终答复"]
Planner --> Step["agent_step"]
Executor --> Step
Verifier --> Step
Tools --> Invocation["tool_invocation"]
Answer --> Session["diagnosis_session"]
```
讲解重点:
- Planner 拆解问题和排查方向。
- Executor 必须通过工具收集证据。
- Verifier 只基于 `tool_trace_summary` 校验答案,不做新检索。
- Trace API 能回放模型步骤和工具证据。
## 3. AIOps 链路
```mermaid
flowchart TD
Alert["告警 payload 或空请求"] --> API["POST /api/ai_ops"]
API --> AiOps["AiOpsService"]
AiOps --> Mode{"是否有 payload?"}
Mode -->|有| Targeted["PAYLOAD_TARGETED\n聚焦输入告警"]
Mode -->|无| Discovery["AUTO_DISCOVERY\n先发现活跃告警"]
Targeted --> Supervisor["ai_ops_supervisor"]
Discovery --> Supervisor
Supervisor --> Planner["planner_agent"]
Supervisor --> Executor["executor_agent"]
Planner --> Tools["Prometheus / 日志 / 知识库"]
Executor --> Tools
Tools --> Report["告警分析报告"]
Report --> Eval["AiOpsRuleEvaluationService"]
Eval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
```
讲解重点:
- AIOps 有明确产品边界:有 payload 时必须聚焦该告警。
- payload 字段会生成 recommended `lookup_knowledge` query。
- 当前 AIOps 先用规则评估做质量门,后续再扩展 LLM Verifier。
## 4. Trace 数据模型
| 表 | 作用 |
|---|---|
| `diagnosis_session` | 一次诊断的主记录:问题、状态、答案、自评估、反馈 |
| `agent_step` | Agent 模型调用记录:输入、输出、耗时、token、是否有工具调用 |
| `tool_invocation` | 工具调用事实:工具名、入参、输出预览、检索层、相关性、成功状态 |
## 5. 为什么 Trace 是核心
故障诊断系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
- 模型为什么这么答。
- 调了哪些工具。
- 工具返回了什么证据。
- Verifier 如何判断答案可信度。
- 用户反馈如何回写到同一个 session。
这就是它区别于普通 Chatbot 的地方。
+145
View File
@@ -0,0 +1,145 @@
# 面试演示脚本
## 1. 30 秒开场
```text
这是一个 Agent 工程项目,场景是企业故障诊断。
它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。
项目重点不是单次模型回答,而是把 Agent 编排、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 Trace。
```
## 2. 启动服务
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
服务地址:
```text
http://localhost:9900
```
`mvp-demo` profile 下:
- Prometheus 告警使用 mock 数据。
- CLS 日志使用 mock 数据。
- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。
## 3. Demo 1:Chat 诊断
目标:展示用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 Trace。
```powershell
$sessionId = "interview-chat-payment-timeout-001"
$body = @{
Id = $sessionId
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
讲解点:
- `ChatService` 会根据问题复杂度选择轻量回答或复杂 Agent 流程。
- 复杂问题走 `Planner -> Executor -> Verifier`。
- Executor 调用知识库、日志、指标等证据工具。
- Verifier 基于 `tool_trace_summary` 生成 groundedness 评估。
- 最终写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
查询 Trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
```
展示点:
- `data.session.agentFlow = CHAT`
- `data.steps` 中能看到 planner/executor/verifier
- `data.toolInvocations` 中能看到证据工具
- `data.session.selfEvaluation` 中有 verifier 结果
## 4. Demo 2:AIOps 告警诊断
目标:展示告警 payload 如何触发 AIOps,并且报告聚焦目标告警。
```powershell
$aiopsSessionId = "interview-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
讲解点:
- `/api/ai_ops` 接受可选 `AIOpsRequest`。
- 首条 SSE 消息会返回 `type=session`。
- `AiOpsService` 根据 payload 判断模式:
- `PAYLOAD_TARGETED`:聚焦传入告警。
- `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。
- AIOps 当前用 rule evaluation 检查报告完整性、payload 聚焦和证据工具覆盖。
查询 Trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
```
展示点:
- `data.session.agentFlow = AI_OPS`
- `data.session.answer` 有最终告警报告
- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge`
- 报告主线聚焦 `HighCPUUsage/payment-service`
## 5. Demo 3:反馈闭环
```powershell
$feedback = @{
sessionId = $sessionId
feedback = "useful"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/feedback" `
-ContentType "application/json" `
-Body $feedback
```
讲解点:
- feedback 写回同一个 `diagnosis_session`。
- `useful` 会沉淀 `case_library`。
- `not_useful` 不改变 `status`,只作为质量信号。
## 6. 收尾总结
```text
这个 Demo 展示的是完整 Agent 闭环:
用户问题或告警 -> Agent 编排 -> 工具证据 -> 自评估 -> Trace 回放 -> 用户反馈 -> 案例沉淀。
我关注的不是一次回答,而是这个回答能否被审计、验证和持续改进。
```
+107
View File
@@ -0,0 +1,107 @@
# 关键设计取舍
## 1. 为什么先做 Trace,而不是只返回答案
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答:
- 结论是什么。
- 证据来自哪里。
- 哪些步骤由哪个 Agent 完成。
- 如果答案不可靠,系统怎么降级。
因此项目把一次会话拆成:
- `diagnosis_session`:会话摘要、最终答案、自评估、用户反馈。
- `agent_step`:模型输入输出、耗时、token 和工具调用标记。
- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。
代价是实现复杂度上升,收益是可回放、可调试、可演示。
## 2. 为什么 RAG 不直接隐藏在 Advisor 里
Spring AI Advisor 可以让 RAG 更隐式,但本项目的核心是 Agent 证据链。`lookup_knowledge` 必须作为显式工具调用出现,这样 Trace 里才能看到:
- Agent 什么时候决定检索。
- 用了什么 query。
- 命中了哪些文档。
- 相关性等级是什么。
- 证据如何支撑最终答案。
所以当前设计是:
```text
Executor -> lookup_knowledge -> VectorSearchService -> VectorStore / SDK fallback
```
这牺牲了一点框架自动化,但保留了可审计性。
## 3. 为什么 L0 只做 hint,不直接返回
旧版 L0 关键词唯一命中时可能直接跳过 L1。这个策略速度快,但风险是:关键词子串命中不等于最终语义相关。
当前改成:
```text
L0 = domain/entity hint
L1 = semantic retrieval
postprocess = evidence shaping + trace
```
L0 仍然有价值:错误码、服务名、告警名、指标名都很适合做精确 hint。但最终证据仍需要 L1 和后处理支撑。
## 4. 为什么保留 Milvus SDK fallback
Spring AI VectorStore 是当前读路径主方向,但 SDK fallback 没有删除,原因有三点:
- 迁移安全:旧 SDK 路径已经被验证过。
- 运行韧性:VectorStore 配置、schema、collection 出问题时可以回退。
- 面试稳定:检索抽象迁移不应该破坏主 demo。
这不是“没有迁完”,而是分阶段迁移:先稳定读路径,再决定是否迁移写入和索引。
## 5. 为什么 Chat 有 Verifier,AIOps 先用规则评估
Chat 问题更开放,容易出现跨领域推理,所以需要 LLM Verifier 做 groundedness 校验。
AIOps 当前优先解决更具体的问题:
- 最终报告是否存在。
- payload 模式是否聚焦输入告警。
- 是否使用了证据工具。
- 是否把无关活跃告警展开成主根因。
这些用规则就能稳定检查。后续可以在同一个 `self_evaluation` 容器下增加 AIOps LLM Verifier。
## 6. 为什么 AIOps payload scope 先用 Prompt + Rule
真实告警环境里可能同时有多个 active alerts。用户传入 `HighCPUUsage/payment-service` 时,Agent 如果把所有告警都展开分析,报告会跑偏。
当前选择:
- Prompt 中加入 `PAYLOAD_TARGETED`。
- 从 payload 生成 recommended `lookup_knowledge` query。
- 用 `AiOpsRuleEvaluationService` 检查报告是否聚焦输入告警。
没有先做硬过滤,是因为有些相关告警可以作为风险背景。目标不是屏蔽上下文,而是控制主诊断对象。
## 7. 为什么反馈不改 status
`status` 表示执行状态,`feedback` 表示用户评价。一个执行成功但用户觉得没用的诊断,应该是:
```text
status = SUCCESS
feedback = not_useful
```
这样才能区分系统异常和质量问题。`useful` 反馈会沉淀 `case_library`,`not_useful` 作为 bad case 信号保留。
## 8. 可以主动承认的限制
- AIOps 还没有完整 LLM Verifier。
- RAG 还没有 hybrid search、rerank、邻居 chunk 扩展。
- `case_library` 的 rootCause/solution 仍需要结构化抽取。
- `tool_invocation.step_id` 关联还可以更严格。
- `mvp-demo` profile 使用 mock 日志和指标,主要服务稳定面试演示。
主动讲清这些限制,能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
@@ -0,0 +1,91 @@
# RAG Breadcrumb Embedding 验收说明
## 1. 改动是什么
索引路径现在构造 embedding 文本时,不只使用 chunk 内容,还会把结构上下文拼进去:
```text
Title: {title}
Path: {breadcrumb}
Content:
{content}
```
Milvus 中存储的 `content` 字段仍然保留原始 chunk 内容。这样展示和证据输出保持干净,而向量本身携带章节语义。
## 2. 为什么必须重新索引
Embedding 是索引时物化的。已有向量是用旧的 content-only 文本生成的,所以只有代码变化并不会改变线上检索结果。
验收关键点:
```text
只改代码 != live retrieval 已变化
代码改动 + 重新索引 + live query report = 行为验收完成
```
## 3. 如何验证
1. 启动 Spring Boot 应用。
2. 通过现有索引路径重新索引知识库。
3. 运行:
```bash
python scripts/eval_rag_live_acceptance.py
```
脚本输出:
```text
eval/rag-retrieval/reports/live-post-reindex.json
eval/rag-retrieval/reports/live-post-reindex.md
```
默认覆盖:
- breadcrumb 敏感的 RAG chunk context query。
- 需要章节路径的诊断流程问题。
- `ERR_TIMEOUT` 精确错误码检索。
- MySQL 连接池排障。
- AIOps payment-service 延迟告警检索。
## 4. 看什么结果
对 breadcrumb 敏感 case:
- top candidates 是否暴露预期 `title`。
- top candidates 是否暴露预期 `breadcrumb`。
- 命中内容是否能看出所属章节。
对核心排障 case:
- 结果数量是否稳定。
- top candidates 是否仍然命中核心文档。
- 没有因为拼接 title/breadcrumb 导致核心检索退化。
## 5. 面试回答
如果被问:你怎么验证 breadcrumb 参与 embedding 后真的生效?
```text
我把 deterministic regression 和 live acceptance 分开。
离线 fixture baseline 不依赖服务,可以做稳定回归。
但 embedding 改动只会影响新生成的向量,所以我另外加了 live post-reindex acceptance 脚本。
脚本会调用真实 /api/search/similar,对 breadcrumb 敏感、排障和 AIOps query 生成 JSON/Markdown 报告。
这样能证明代码改了,也能证明 live vector collection 已经刷新。
```
如果被问:为什么脚本不自动 reindex?
```text
reindex 会修改向量库,而且依赖环境中的知识库数据。
我把 reindex 保持为显式动作,验收脚本只做读取验证。
这样如果检索没有改善,我能区分是代码问题、索引未刷新,还是运行时检索行为问题。
```
## 6. 后续增强
- 将 live acceptance 结果加入面试 Demo 输出。
- 增加 breadcrumb hit rate 统计。
- 对同章节 chunk 做邻居扩展,进一步利用 breadcrumb。
+127
View File
@@ -0,0 +1,127 @@
# RAG 重构故事
## 1. 起点
原始 RAG 实现已经能支撑 MVP:
- 文档可以上传、切片、向量化,并写入 Milvus/Zilliz。
- Agent 可以显式调用 `lookup_knowledge`。
- AIOps 诊断能在告警流程里检索排障知识。
- 工具调用会落到 `tool_invocation`,检索步骤可见。
但它有几个工程问题:
- 检索实现过于依赖 Milvus SDK,业务代码承担了太多底层搜索细节。
- L0 和 L1 职责不清,L0 关键词命中容易被当作最终召回决策。
- chunk 级检索容易丢失章节上下文。
- `breadcrumb` 存在 metadata 中,但没有充分参与 embedding、filter 和上下文重建。
- 检索质量主要靠手工接口和日志判断,缺少可重复的 golden cases。
所以重构目标不是“全盘替换成框架”,而是:
```text
通用 RAG 基础设施交给 Spring AI,
业务可观测链路保留在项目里。
```
## 2. 我如何拆解问题
我把迁移拆成几个阶段,因为 RAG 同时影响 Agent 工具层、AIOps、向量检索、证据打包和 Trace。
第一步是建立 baseline。`eval/rag-retrieval/` 中的 golden cases 用来对比后续改动,而不是只靠直觉判断检索有没有变好。
第二步是明确职责:
```text
L0 = domain/entity hint
L1 = semantic retrieval
postprocess = evidence shaping + trace-friendly output
```
L0 仍然有价值,但不再默认绕过语义检索。它更适合提取服务名、告警名、错误码、领域和 metadata filter。
第三步是增强 evidence 输出。Agent 不应该只拿到 raw chunk,而应该拿到带 source、title、breadcrumb、score、hit reason 的证据块。
最后,我把 Spring AI `VectorStore` 接入为读取主路径,同时保留原 Milvus SDK 作为 fallback。
## 3. 当前架构
```text
Agent / API
-> lookup_knowledge or /api/search/similar
-> L0 domain/entity hint
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
-> relevance normalization
-> tool_invocation trace
```
`VectorSearchService` 仍然是公共检索门面。Agent 工具层不需要知道底层是 SDK 还是 Spring AI。
支持三种模式:
```text
auto -> 优先 Spring AI VectorStore,失败后 fallback 到 SDK
spring-ai -> 强制 Spring AI VectorStore
sdk -> 强制 Milvus SDK
```
## 4. 关键取舍
### 保留显式工具
我没有把检索藏进 Spring AI Advisor。原因是这个项目强调 Agent 执行可见性:`lookup_knowledge` 的 query、命中文档、相关性和证据预览都要进入 Trace。
### 保留 SDK fallback
SDK fallback 不是废代码,而是迁移安全网。实际验证时,第一次 VectorStore 指向了错误 collection,`auto` 模式 fallback 到 SDK 后仍能返回结果。修正 collection 后,Spring AI 路径成为主路径。
### L0 降权
生产事故中经常有精确标识:错误码、告警名、服务名、指标名。L0 适合做 hint,但不应该做最终裁判。
### 分数语义拆开
SDK 使用 L2 distance,Spring AI 暴露 similarity。混在一个字段里会让 relevance normalization 出错。
当前拆成:
```text
score -> 兼容旧逻辑的距离型分数
rawScore -> 底层原始分数
scoreLabel -> rawScore 的语义
```
### 暂不迁移写入
写入和索引仍走 SDK。这是有意分阶段:先验证读路径,再评估 `VectorStore.add(...)` 是否适合现有 metadata 和 chunk 模型。
## 5. 验证方式
我用了三层验证:
- 单元测试:SDK mode、Spring AI mode、auto fallback、category filter、distance metadata mapping。
- Live API:`GET /api/search/similar?query=ERR_TIMEOUT&topK=3`。
- 代表性 query 对比:错误码、支付超时、MySQL 连接池、AIOps 告警式 query、抽象 RAG 设计问题。
核心排障和 AIOps query 在 SDK 与 VectorStore 下 top3 一致。差异主要集中在抽象设计类问题和 metadata taxonomy,这些被记录为后续质量工作。
## 6. 面试短版
```text
这个 RAG 系统最初是基于 Milvus SDK 的自研 MVP。它能跑,但底层检索细节过多地散落在业务代码里,L0/L1 职责也不够清晰。
我按阶段重构:先加 retrieval baseline,再把 L0 降级为 domain/entity hint,再增强 evidence postprocess,最后把读取主路径切到 Spring AI VectorStore,并保留 SDK fallback。
我没有把 lookup_knowledge 替换成隐式 Advisor,因为这个项目的核心是可追踪 Agent:面试官可以看到什么时候检索、检索了什么、证据如何支撑诊断。
```
## 7. 可主动承认的不足
- metadata taxonomy 还需要清理,例如 `database` 与 `infrastructure`。
- 抽象设计问题可能需要 query rewrite 或更好的文档索引。
- 邻居 chunk / 同章节上下文扩展还不完整。
- rerank、RRF、BM25、hybrid retrieval 还没有接入。
- 写入路径仍使用 SDK。
这些不是当前迁移阻塞项,而是后续检索质量优化方向。
+148
View File
@@ -0,0 +1,148 @@
# RAG 检索质量报告
## 1. 目的
这份报告回答一个面试关键问题:
```text
迁移到 Spring AI VectorStore 后,怎么证明检索质量没有退化?
```
这不是完整 benchmark,而是针对当前 Milvus/Zilliz collection 的代表性 live smoke comparison。
## 2. 验证设置
服务端点:
```text
GET http://127.0.0.1:9900/api/search/similar
```
collection:
```text
biz
```
对比模式:
```text
retrieval.vector-store.mode=sdk
retrieval.vector-store.mode=spring-ai
```
每个 case:
```text
topK=3
```
## 3. 测试案例
| Case | Query | 目的 |
|---|---|---|
| `err-timeout` | `ERR_TIMEOUT` | 精确错误码检索 |
| `payment-service-timeout` | `payment-service timeout` | 服务超时排障 |
| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | 数据库排障 |
| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps 告警式检索 |
| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | 抽象 RAG 设计问题 |
| `database-filter` | `mysql timeout`, category=`database` | metadata filter 行为 |
## 4. 对比摘要
| Case | SDK 数量 | VectorStore 数量 | Top1 一致 | TopK 重叠 | 结论 |
|---|---:|---:|---|---:|---|
| `err-timeout` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
| `payment-service-timeout` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
| `mysql-connection-pool` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
| `high-cpu-payment` | 3 | 3 | 是 | 3/3 | AIOps 核心 query 一致 |
| `rag-l0-l1` | 3 | 1 | 是 | 1/3 | VectorStore 尾部结果更少 |
| `database-filter` | 0 | 0 | 不适用 | 不适用 | filter 行为一致,taxonomy 有问题 |
## 5. 代表性结果
### ERR_TIMEOUT
SDK:
```text
1. ERR_TIMEOUT score=0.5659486 label=l2_distance
2. ERR_GATEWAY_TIMEOUT score=0.6048740 label=l2_distance
3. Error handling score=0.7735061 label=l2_distance
```
VectorStore:
```text
1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity
2. ERR_GATEWAY_TIMEOUT score=0.6048740 rawScore=0.3951259 label=similarity
3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity
```
解释:
- 排序一致。
- 兼容 `score` 与 SDK L2 distance 一致。
- `rawScore` 暴露 Spring AI similarity。
### MySQL connection pool
两条路径都返回:
```text
1. MySQL connection pool config
2. wait_timeout timeout
3. idle-timeout
```
说明迁移保留了核心基础设施排障检索能力。
### HighCPUUsage payment-service
两条路径都返回 payment-service 高 CPU 相关排障文档,说明 AIOps 告警式 query 没有退化。
### rag-l0-l1
VectorStore 只返回一个候选,但 Top1 与 SDK 一致。这说明抽象设计类 query 需要后续 query rewrite、补充索引或 threshold 调整。
### database-filter
两条路径都返回 0,因为相关 MySQL 文档当前分类是 `infrastructure`,不是 `database`。这是 metadata taxonomy 问题,不是 VectorStore 回归。
## 6. 分数兼容结论
对比验证了当前分数设计:
```text
SDK:
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
VectorStore:
score = Milvus metadata.distance
rawScore = Spring AI similarity
scoreLabel = similarity
```
这样既保持 `lookup_knowledge` 原有归一化逻辑,又能暴露 VectorStore 语义。
## 7. 验收结论
Spring AI VectorStore 读路径可以接受用于当前 MVP/面试:
- 核心排障和 AIOps case 与 SDK top3 一致。
- 分数兼容性保留。
- VectorStore 语义通过 `rawScore` 和 `scoreLabel` 可观察。
- SDK fallback 仍保留运行安全。
后续检索质量工作不阻塞这次迁移,应作为独立优化继续推进。
## 8. 下一步
- 增加自动 live comparison 脚本。
- 在 offline evaluator 中加入 topK overlap、top1 hit、MRR。
- 规范 metadata category,例如 `database` 与 `infrastructure`。
- 为抽象设计类 query 增加 query rewriting。
- 后续再评估是否迁移写入路径到 `VectorStore.add(...)`。
@@ -0,0 +1,119 @@
# RAG VectorStore 面试要点
## 1. 60 秒讲法
```text
我把 RAG 检索从 Milvus SDK-only 重构为 Spring AI VectorStore 主路径,同时保留 SDK fallback。
关键不是换了一个依赖,而是保留 VectorSearchService 作为边界,所以 lookup_knowledge 和 Agent workflow 不需要改。
现在支持 auto、spring-ai、sdk 三种模式。auto 会优先尝试 VectorStore,失败后 fallback 到 SDK。
```
现场验证时,第一次发现 VectorStore 指向了错误 collection:`business_knowledge`,而实际 Zilliz collection 是 `biz`。fallback 生效,所以系统仍能通过 SDK 返回结果。修正 collection 后,同一个 query 成功走 Spring AI VectorStore。
## 2. 架构回答
```text
Agent / API
-> lookup_knowledge or /api/search/similar
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
-> Milvus/Zilliz collection: biz
```
关键设计:`VectorSearchService` 是检索门面,避免 Spring AI 或 SDK 细节扩散到 Agent 工具层。
## 3. 为什么保留 SDK
- 迁移安全:原 SDK 路径已验证可用。
- 运行韧性:VectorStore schema、filter 或配置失败时,检索仍可用。
- Demo 稳定:检索抽象变化不应该破坏主诊断演示。
这在实际验证中发挥了作用:VectorStore 配置错时,`auto` 模式 fallback 到 SDK,API 没有失败。
## 4. 为什么引入 Spring AI VectorStore
使用 `VectorStore` 可以让项目更接近标准 RAG 抽象:
- 业务代码不再持有全部 Milvus search 细节。
- 后续 QueryTransformer、DocumentPostProcessor、Retriever 等能力更容易接入。
- 面试中也更容易解释和 Spring AI 生态的关系。
但我没有一次性迁移写入,因为读写同时迁移会让问题难定位。当前先稳定读路径。
## 5. 为什么保留 L0
L0 现在不是最终答案来源,而是确定性 hint 层:
- 提取 domain/entity。
- 在可能时生成 category filter。
- 给 trace 提供解释信号。
当前职责:
```text
L0 = domain/entity hint
L1 = semantic retrieval through VectorStore/SDK
postprocess = evidence trace + relevance normalization
```
真实故障诊断里有很多精确标识,完全只靠向量检索并不稳。
## 6. 为什么不用隐藏 Advisor
`lookup_knowledge` 保持显式工具,因为:
- Trace 要展示什么时候检索。
- `tool_invocation` 要记录输入、输出预览、相关性和 metadata。
- 面试故事是可审计 Agent 执行,而不只是答案质量。
Advisor 后续可以接入,但需要先解决可观测性。
## 7. 分数设计
当前结果故意拆成:
```text
score -> 兼容旧 relevance normalization 的分数
rawScore -> 当前检索实现原始分数
scoreLabel -> rawScore 的语义
```
SDK:
```text
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
```
VectorStore:
```text
score = metadata.distance if present
rawScore = Spring AI document score
scoreLabel = similarity
```
这样避免把 similarity 当成 L2 distance 的隐蔽 bug。
## 8. 如何证明 VectorStore 被使用
- 日志出现 `Starting Spring AI VectorStore search` 和 `Spring AI VectorStore search complete`。
- API 响应中 `scoreLabel=similarity`。
- `rawScore` 是 Spring AI similarity,`score` 仍是兼容 distance。
## 9. 常见追问
### 为什么不删 SDK?
这是迁移,不是重写。fallback 提供回滚安全,并且已经证明配置错误时仍能保证主链路可用。
### `lookup_knowledge` 变了吗?
外部契约没变。它仍然调用 `VectorSearchService.searchSimilarDocuments(...)`,变化在门面背后的实现。
### 这是完整 Spring AI RAG 了吗?
还不是。当前是 Spring AI VectorStore 读路径 + 显式工具 + 自定义 evidence trace + SDK 写入。这样做是为了保留审计能力和分阶段迁移安全。
@@ -0,0 +1,198 @@
# RAG VectorStore Live 验收说明
## 1. 目的
本文记录 RAG 检索重构的 live 验收结论。
这次重构的目标不只是接入 Spring AI 抽象,而是证明线上读路径能够:
- 优先使用 Spring AI `VectorStore` 做 Milvus 检索。
- 保留原 Milvus SDK 作为 fallback。
- 保持 `lookup_knowledge` 工具契约稳定。
- 保持基于 L2 distance 的相关性归一化兼容。
## 2. 当前检索形态
```text
lookup_knowledge / /api/search/similar
-> VectorSearchService.searchSimilarDocuments(...)
-> retrieval.vector-store.mode
-> auto
-> Spring AI VectorStore
-> VectorStore 失败时 fallback 到 Milvus SDK
-> spring-ai
-> 只走 Spring AI VectorStore
-> sdk
-> 只走 Milvus SDK
```
## 3. 已验证配置
live Milvus/Zilliz 数据库中存在 collection:
```text
biz
```
Spring AI VectorStore 配置与 SDK 使用的 collection 对齐:
```yaml
spring:
ai:
vectorstore:
type: milvus
milvus:
initialize-schema: false
database-name: ${milvus.database}
collection-name: biz
embedding-dimension: ${milvus.vector-dim}
metric-type: L2
id-field-name: id
content-field-name: content
metadata-field-name: metadata
embedding-field-name: vector
```
为什么重要:早期配置使用 `business_knowledge`,而真实 collection 是 `biz`。这个错配证明了 fallback 生效,但也说明修正前 VectorStore 不是成功主路径。
## 4. 验收命令
健康检查:
```powershell
Invoke-RestMethod `
-Uri "http://127.0.0.1:9900/milvus/health" `
-Method Get
```
期望:
```json
{
"collections": ["biz"],
"message": "ok"
}
```
直接检索:
```powershell
Invoke-RestMethod `
-Uri "http://127.0.0.1:9900/api/search/similar?query=ERR_TIMEOUT&topK=3" `
-Method Get
```
期望结果形态:
```json
{
"code": 200,
"message": "success",
"data": [
{
"content": "### ERR_TIMEOUT ...",
"score": 0.5662,
"rawScore": 0.4337,
"scoreLabel": "similarity",
"metadata": {
"distance": 0.5662,
"title": "ERR_TIMEOUT",
"category": "api"
}
}
]
}
```
## 5. 日志证明了什么
collection 修正前:
```text
Starting Spring AI VectorStore search
Spring AI VectorStore retrieval failed, falling back to Milvus SDK
Starting Milvus SDK search
```
collection 修正后:
```text
Starting Spring AI VectorStore search: query=ERR_TIMEOUT
Spring AI VectorStore search complete, candidates=3
```
这证明:
- `auto` 模式确实先尝试 VectorStore。
- VectorStore 失败时 fallback 可用。
- 配置对齐后,主路径是 Spring AI VectorStore,而不是 SDK fallback。
## 6. 分数语义
项目保留三个分数字段:
```text
rawScore -> 当前检索实现的原始分数
scoreLabel -> rawScore 的语义
score -> lookup relevance normalization 使用的兼容分数
```
SDK:
```text
rawScore = L2 distance
scoreLabel = l2_distance
score = L2 distance
```
Spring AI VectorStore:
```text
rawScore = Spring AI similarity score
scoreLabel = similarity
score = Milvus distance metadata when available
```
使用 `metadata.distance` 的原因:`LookupKnowledgeTool` 已经基于 L2 distance 做相关性归一化。Spring AI Milvus 主分数是 similarity,但 metadata 中仍有 Milvus distance。用 distance 保持旧逻辑稳定,同时通过 `rawScore` 暴露新语义。
## 7. 回归检查
目标测试:
```powershell
mvn -q "-Dtest=VectorSearchServiceTest,LookupKnowledgeToolTest" test
```
相关 spec:
```powershell
openspec.cmd validate rag-knowledge-retrieval --specs
openspec.cmd validate rag-retrieval-evaluation --specs
```
diff 检查:
```powershell
git diff --check
```
验收结论:
```text
目标测试通过。
相关 spec 通过。
diff-check 无错误。
```
## 8. 验收结论
VectorStore 读路径可以接受:
- Spring AI VectorStore 已集成,并在 `auto` 模式中优先使用。
- SDK fallback 保留且已被实际验证。
- live collection 配置与现有 Milvus collection 对齐。
- `lookup_knowledge` 对外契约保持稳定。
- 旧的 L2 relevance normalization 仍兼容。
写入和索引路径仍使用 Milvus SDK。这是有意的分阶段迁移,不是验收失败项。
+225
View File
@@ -0,0 +1,225 @@
# 面试故事案例
**用途**:把项目能力讲成可被面试官理解的工程故事
**使用方式**:按问题选择一个故事,不需要从头到尾背诵
## 故事 1:从黑盒 Chatbot 到可追踪 Agent
### 面试官问题
```text
这个项目和普通调用大模型有什么区别?
```
### 30 秒回答
```text
普通 Chatbot 只给最终答案,出了问题很难解释答案怎么来的。
我这个项目把诊断过程拆成 Planner、Executor、Verifier,并把每个 Agent 步骤和每次工具调用落库。
最后通过 Trace API 可以回放:模型怎么规划、调用了哪些工具、工具返回了什么证据、Verifier 怎么判断答案可信。
```
### 展开讲法
一开始最容易做的是:用户问题进来,直接让模型回答。但故障诊断场景不能只看答案,因为答案可能看起来合理却没有证据支撑。
所以我把系统拆成三层:
- `diagnosis_session` 记录一次诊断的主状态和最终答案。
- `agent_step` 记录 Planner、Executor、Verifier 的模型调用。
- `tool_invocation` 记录知识库、日志、指标等真实工具证据。
这样就能做到:答案不是孤立文本,而是一条可审计的执行链。
### 可展示文件
- `mvp/architecture/interview-one-pager.md`
- `mvp/architecture/session-trace-lifecycle.md`
- `mvp/demo/output/trace-response.json`
### 主动说不足
```text
当前 tool_invocation.step_id 还不是每次都强绑定具体 agent_step,后续可以加 runId 和更严格的 step 关联,让多轮同 session 诊断更清晰。
```
## 故事 2:RAG 从自研 SDK 检索迁移到 Spring AI VectorStore
### 面试官问题
```text
你的 RAG 是怎么设计的?为什么不用框架全包?
```
### 30 秒回答
```text
我把 RAG 分成两部分:通用检索基础设施尽量交给 Spring AI VectorStore,业务可观测链路留在项目里。
所以 Agent 仍然显式调用 lookup_knowledge,底层通过 VectorSearchService 走 Spring AI VectorStore,失败时 fallback 到原 Milvus SDK。
这样既能减少自研检索代码,又不会丢失工具调用 trace。
```
### 展开讲法
旧实现里,Milvus SDK 查询、topK、filter、score 映射都在业务代码里。它能跑,但后续扩展成本高。
我没有直接把 RAG 隐藏进 Advisor,因为这个项目的核心是 Agent 工程,需要知道 Agent 何时检索、检索了什么、证据怎么支撑诊断。
于是我保留了边界:
```text
Executor -> lookup_knowledge -> VectorSearchService -> VectorStore / SDK fallback
```
同时把 L0 从“直接返回结果”降级为 domain/entity hint,降低关键词误召回的风险。
### 可展示文件
- `mvp/architecture/rag-architecture.md`
- `mvp/architecture/retrieval-observability.md`
- `interview/rag-refactor-story.md`
### 主动说不足
```text
当前还没有完整 hybrid search 和 rerank。
我先做 golden cases、VectorStore 主路径和 SDK fallback,是为了让每一步迁移都能被验证。
```
## 故事 3:Verifier 如何降低幻觉风险
### 面试官问题
```text
Agent 怎么保证不胡说?
```
### 30 秒回答
```text
我没有假设模型天然可靠,而是加了 Verifier。
Executor 给出答案后,Verifier 只拿 executor_final_answer 和 tool_trace_summary,不允许做新检索。
它把答案里的关键事实逐条校验,输出 PASS、LOW_CONFID 或 REJECT。
这个结果会写回 self_evaluation,Trace API 可以看到。
```
### 展开讲法
Verifier 的关键不是再问一次模型“你觉得对吗”,而是让它基于真实工具调用做 groundedness 检查。
`ToolTraceSummaryService` 会从 `tool_invocation` 里整理证据索引,包含:
- 工具名。
- 输入摘要。
- 输出摘要。
- evidence level。
- source invocation ids。
Verifier 输出结构化 JSON,ChatService 根据 verdict 决定是否输出、补证据或降级。
### 可展示文件
- `mvp/architecture/harness-quality-gates.md`
- `mvp/architecture/feedback-architecture.md`
- `src/main/resources/prompts/chat-verifier-prompt.md`
### 主动说不足
```text
AIOps 当前还是轻量 rule evaluation,不是完整 LLM Verifier。
这是有意收敛:先用规则保证 payload 聚焦和工具证据使用,后续再加 AIOps LLM Verifier。
```
## 故事 4:AIOps 告警为什么要做 payload scope control
### 面试官问题
```text
AIOps 场景和普通 Chat 有什么区别?
```
### 30 秒回答
```text
AIOps 告警有一个很关键的问题:环境里可能同时有很多活跃告警,Agent 容易跑偏。
所以我把 AIOps 分成 PAYLOAD_TARGETED 和 AUTO_DISCOVERY。
如果请求带 alert payload,最终报告必须聚焦输入告警,并且会把 alertName、service、severity、description 拼成 recommended lookup_knowledge query。
```
### 展开讲法
没有 payload 时,Agent 可以先查询活跃告警,再选择目标排查。
但有 payload 时,用户已经告诉系统“我要查这个告警”。这时如果 Agent 把其他活跃告警写成主根因,产品体验会很差。
所以我做了两件事:
- Prompt 中明确 `PAYLOAD_TARGETED` 范围。
- `AiOpsRuleEvaluationService` 检查最终报告是否聚焦输入告警,以及是否使用证据工具。
### 可展示文件
- `mvp/architecture/agent-orchestration.md`
- `mvp/architecture/current-mvp-architecture.md`
- `interview/aiops-query-augmentation.md`
- `interview/aiops-lightweight-verifier.md`
### 主动说不足
```text
当前 scope control 主要靠 prompt 和规则评估。
后续可以把 AIOps 也接入类似 Chat Verifier 的事实校验,让告警报告的每个根因和建议都有 evidence refs。
```
## 故事 5:反馈不是点赞按钮,而是案例沉淀入口
### 面试官问题
```text
用户反馈在系统里有什么用?
```
### 30 秒回答
```text
反馈不只是前端按钮。
用户提交 useful 后,系统会把同一个 diagnosis_session 沉淀为 case_library。
not_useful 不会改执行状态,而是作为 bad case 信号保留。
这样 status、self_evaluation、feedback 三个维度是分开的。
```
### 展开讲法
我刻意没有把 `not_useful` 写成 `FAILED`。因为失败表示系统执行异常,而用户觉得不好用是质量标签。
当前设计里:
```text
status -> 执行是否成功
self_evaluation -> 系统自己判断证据和事实支撑度
feedback -> 用户是否认可
```
`useful` 会进入 `CaseLibraryService.createFromSession`,生成可复用案例。后续可以做相似案例推荐或高质量样本积累。
### 可展示文件
- `mvp/architecture/feedback-architecture.md`
- `mvp/architecture/data-model.md`
- `mvp/demo/output/feedback-response.json`
### 主动说不足
```text
当前 case_library 的 rootCause 和 solution 还直接使用完整 answer。
后续应该从报告中结构化抽取 rootCause、solution、errorCode 和 service,提高案例复用质量。
```
## 结尾万能总结
```text
这个项目我最想展示的不是某一个模型效果,而是 Agent 工程化能力:
一个诊断答案从哪里来、用了什么证据、是否被验证、用户是否认可、后续怎么沉淀和回归。
这些链路都被结构化记录下来,所以它可以继续演进,而不是一次性 demo。
```
@@ -0,0 +1,81 @@
---
title: AIOps 告警排障 Runbook
keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs]
summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。
category: troubleshooting
---
# AIOps 告警排障 Runbook
## 1. 告警输入处理原则
AIOps 诊断入口有两种触发方式:
- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。
- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。
最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。
## 2. Mock 告警与日志主题映射
| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 |
|---|---|---|---|
| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` |
| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` |
| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` |
| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` |
## 3. HighCPUUsage / payment-service 排障步骤
### 3.1 现象确认
先确认 Prometheus 活动告警中是否存在:
- `alert_name = HighCPUUsage`
- `service = payment-service`
- CPU 使用率超过 80%
- 状态为 firing
如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。
### 3.2 指标与日志取证
推荐工具调用顺序:
1. `queryPrometheusAlerts`:确认当前活动告警。
2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。
3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。
### 3.3 根因判断
可接受的根因结论必须至少满足一项:
- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。
- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。
- 告警持续时间与日志时间线一致。
如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。
## 4. 处理建议
### 临时止血
- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。
- 对高耗时接口开启限流或降级非核心功能。
- 如果近期有发布,检查变更窗口并准备回滚。
### 根因修复
- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。
- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。
- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。
## 5. 报告要求
告警分析报告至少包含:
- 活跃告警清单。
- 告警根因分析。
- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。
- 已执行或建议执行的处理方案。
- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。
+96 -137
View File
@@ -1,157 +1,116 @@
# 数据库设计文档
# SuperBizAgent MVP 文档
## 📚 文档导航
**更新日期**:2026-07-05
### 核心表设计
- [diagnosis_record](tables/diagnosis_record.md) - 诊断记录表(核心)
- [case_library](tables/case_library.md) - 案例库表
- [api_document](tables/api_document.md) - 文档元数据表
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前架构入口已经整理到 `mvp/architecture/`,旧版架构材料已归档,避免继续把历史方案当成当前实现。
### 架构设计
- [Agent 架构设计](architecture/agent-architecture.md) - Agent 协作 + Skill + Harness
- [知识库检索架构](architecture/knowledge-retrieval-architecture.md) - L0+L1 混合检索架构 ⭐新增
- [知识库检索使用指南](architecture/knowledge-retrieval-usage.md) - 文档编写和使用说明 ⭐新增
- [会话管理](architecture/session-management.md) - Redis + MySQL 会话管理
- [实施规划](architecture/implementation-plan.md) - 分阶段实施计划
## 当前入口
---
| 目录/文档 | 用途 |
|---|---|
| [architecture/README.md](architecture/README.md) | 当前 MVP 架构入口 |
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
| [architecture/feedback-architecture.md](architecture/feedback-architecture.md) | 反馈与自评估架构 |
| [architecture/session-trace-lifecycle.md](architecture/session-trace-lifecycle.md) | 会话与 Trace 生命周期 |
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
| [issues/rag-refactor-plan.md](issues/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
| [eval/README.md](eval/README.md) | 诊断评测材料 |
| [issues/README.md](issues/README.md) | MVP issue 索引 |
## 一、设计原则
## 当前系统一句话
### 1.1 核心原则
- ✅ **简单优先**:满足诊断流程需要,避免过度设计
- ✅ **渐进增强**:先实现核心功能,再逐步扩展
- ✅ **数据分离**:诊断结果持久化(MySQL),会话上下文临时化(Redis)
- ✅ **适度冗余**:避免过度范式化,适当冗余提升查询性能
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据,诊断过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
### 1.2 系统定位
**自动化诊断系统**
- 核心:一键诊断 → 返回完整报告
- 辅助:支持追问,但不是主要场景
- 特点:大部分用户单次诊断即结束,少数用户会追问细节
## 文档结构
---
## 二、表结构总览
### 2.1 核心表关系
```
┌─────────────────────┐
│ diagnosis_record │ 诊断记录(核心)
│ - 每次诊断一条 │
└──────────┬──────────┘
│ 1:1
↓
┌─────────────────────┐
│ case_library │ 案例库(知识沉淀)
│ - 诊断成功→案例 │
└─────────────────────┘
┌─────────────────────┐
│ api_document │ 文档元数据(管理层)
│ - 状态追踪/去重 │
└──────────┬──────────┘
│ doc_id
↓
┌─────────────────────┐
│ Milvus │ 文档内容(检索层)
│ - 向量检索 │
└─────────────────────┘
┌─────────────────────┐
│ Redis Session │ 会话管理(临时)
│ - 30分钟过期 │
│ - 支持追问 │
└─────────────────────┘
```text
mvp/
architecture/
README.md
current-mvp-architecture.md
interview-one-pager.md
agent-orchestration.md
harness-quality-gates.md
rag-architecture.md
retrieval-observability.md
feedback-architecture.md
session-trace-lifecycle.md
knowledge-base-authoring.md
data-model.md
evolution-roadmap.md
archive/2026-07-05-legacy/
issues/
README.md
rag-refactor-plan.md
ISS-*.md
rag-*.md
demo/
README.md
ten-minute-interview-demo.md
requests/
scripts/
output/
eval/
README.md
schema.md
cases/
fixtures/
reports/
notes/
plan/
tables/
```
### 2.2 表统计
## 当前核心设计
| 表名 | 类型 | 预估数据量 | 用途 |
|------|------|-----------|------|
| diagnosis_record | 核心 | 3.6万/年 | 诊断记录 |
| case_library | 核心 | 500-1000 | 案例库 |
| api_document | 核心 | 100-200 | 文档管理 |
- `lookup_knowledge` 保持显式 Agent Tool,不隐藏到 Chat Advisor。
- L0 降级为 domain/entity hint,不再默认承担最终召回决策。
- `VectorSearchService` 是检索稳定门面。
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
- AIOps payload 会生成推荐知识库 query,保留业务语义。
- Trace API 聚合 session、step、tool invocation 和 self evaluation。
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
---
## 关键运行链路
## 三、技术栈
```text
Chat
-> ChatService
-> Planner / Executor / Verifier
-> evidence tools
-> diagnosis_session / agent_step / tool_invocation
-> DiagnosisTraceService
### 3.1 数据存储
```
MySQL 8.0+
├─ 元数据管理
├─ 事务支持
└─ JSON 字段支持
AIOps
-> AiOpsService
-> PAYLOAD_TARGETED or AUTO_DISCOVERY
-> Planner / Executor
-> Prometheus / logs / lookup_knowledge
-> AiOpsRuleEvaluationService
-> DiagnosisTraceService
Redis 6.0+
├─ 会话存储
├─ 缓存
└─ TTL 自动过期
Milvus 2.6+
├─ 向量存储
├─ 语义检索
└─ 混合检索
RAG
-> lookup_knowledge
-> L0 domain/entity hint
-> VectorSearchService
-> Spring AI VectorStore / Milvus SDK fallback
-> relevance normalization
-> tool_invocation
```
### 3.2 开发框架
```
Spring Boot 3.2
Spring AI Alibaba 1.1.0
Milvus SDK Java 2.6.10
DashScope SDK
```
## 旧文档说明
---
旧版架构文档已移动到:
## 四、快速开始
- [architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
### 4.1 创建数据库
```sql
-- 1. 创建数据库
CREATE DATABASE diagnosis_system CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;
-- 2. 执行建表脚本(按顺序)
SOURCE tables/diagnosis_record.sql;
SOURCE tables/case_library.sql;
SOURCE tables/api_document.sql;
```
### 4.2 初始化 Milvus
```java
// 创建 Collection
MilvusClientFactory.createCollection();
```
### 4.3 配置 Redis
```yaml
spring:
redis:
host: localhost
port: 6379
database: 0
```
---
## 五、版本历史
| 版本 | 日期 | 变更内容 |
|------|------|---------|
| v1.0 | 2024-06-15 | 初版,定义核心表结构 |
| v2.0 | 2024-06-15 | diagnosis_record 字段泛化,支持多种故障类型 |
| v2.1 | 2024-06-22 | 文档拆分,增加 api_document 表 |
---
## 六、维护说明
- 每个表的详细设计在 `tables/` 目录下
- 架构设计文档在 `architecture/` 目录下
- 修改表结构时,同步更新对应的 Markdown 文档
- 重大变更需记录在版本历史中
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/current-mvp-architecture.md` 与 `architecture/rag-architecture.md` 为准。
+42
View File
@@ -0,0 +1,42 @@
# MVP 架构文档
**更新日期**:2026-07-05
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
- `mvp/architecture/archive/2026-07-05-legacy/`
归档材料只作为设计历史阅读,不再作为当前实现依据。
## 当前文档
| 文档 | 用途 |
|---|---|
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库文档编写与维护规范,覆盖 frontmatter、category、chunk、reindex |
| [data-model.md](data-model.md) | 数据模型总览,覆盖 Trace、知识库、反馈沉淀和 Milvus metadata |
| [evolution-roadmap.md](evolution-roadmap.md) | 从旧版 Agent 蓝图继承的后续演进路线,不代表当前已实现 |
## 当前架构一句话
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
## 阅读顺序
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
6. 继续读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
7. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
8. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
9. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
10. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
+187
View File
@@ -0,0 +1,187 @@
# Agent 编排架构
**更新日期**:2026-07-05
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 设计定位
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Verifier`。
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
## 2. 当前 Agent 全景
```mermaid
flowchart TB
subgraph Chat["Chat diagnosis"]
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
ChatService --> ChatPlanner["chat_planner"]
ChatPlanner --> ChatExecutor["chat_executor"]
ChatExecutor --> ChatTools["evidence tools"]
ChatTools --> ChatExecutor
ChatExecutor --> ChatVerifier["chat_verifier"]
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
ChatDecision --> ChatAnswer["final answer"]
end
subgraph AiOps["AIOps diagnosis"]
AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
AiOpsService --> Supervisor["ai_ops_supervisor"]
Supervisor --> AiOpsPlanner["planner_agent"]
Supervisor --> AiOpsExecutor["executor_agent"]
AiOpsPlanner --> AiOpsExecutor
AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
AiOpsTools --> AiOpsReport["alert report"]
AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
end
subgraph Trace["Trace persistence"]
Session["diagnosis_session"]
Step["agent_step"]
Invocation["tool_invocation"]
SelfEval["self_evaluation"]
end
ChatService --> Session
ChatPlanner --> Step
ChatExecutor --> Step
ChatVerifier --> Step
ChatTools --> Invocation
ChatDecision --> SelfEval
AiOpsService --> Session
AiOpsPlanner --> Step
AiOpsExecutor --> Step
AiOpsTools --> Invocation
AiOpsRule --> SelfEval
```
## 3. Chat 编排
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
```text
chat_planner
-> chat_executor
-> lookup_knowledge / query_logs / query_metrics / date_time
-> chat_verifier
-> reads tool_trace_summary
-> outputs verifier JSON
```
关键行为:
| 角色 | 当前职责 | 输出 |
|---|---|---|
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
| `chat_executor` | 按计划调用证据工具,组合工具返回形成诊断答复 | `executor_feedback` |
| `chat_verifier` | 只基于已有证据校验 Executor 答案,不做新检索 | `verifier_output` |
Chat 链路最多支持两轮验证:
```mermaid
sequenceDiagram
autonumber
participant C as ChatService
participant P as chat_planner
participant E as chat_executor
participant T as tools
participant V as chat_verifier
participant S as diagnosis_session
C->>P: 原始问题 + history + retry_context
P-->>C: planner_plan
C->>E: planner_plan + 上下文
E->>T: 调用证据工具
T-->>E: 证据结果
E-->>C: executor_feedback
C->>V: executor_final_answer + tool_trace_summary
V-->>C: PASS / LOW_CONFID / REJECT
C->>S: 写入 verifier_evaluation
alt LOW_CONFID 且允许补证据
C->>P: retry_context: 仅补缺失证据
else PASS 或 REJECT
C-->>S: 保存最终 answer
end
```
决策语义:
| Verdict | 行为 |
|---|---|
| `PASS` | 输出 Executor 答案 |
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
## 4. AIOps 编排
AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
```text
ai_ops_supervisor
-> planner_agent
-> executor_agent
-> final report
-> AiOpsRuleEvaluationService
```
与 Chat 的差异:
- AIOps 的输入可能是结构化告警 payload。
- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
## 5. 工具边界
当前 Executor 可用工具来自两类:
```text
methodTools
-> dateTimeTools
-> lookupKnowledgeTool
-> queryMetricsTools
-> queryLogsTools when mock enabled
ToolCallbackProvider
-> framework-discovered tools
```
工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
- L0/L1 命中数量。
- 检索层。
- relevance level。
- retrieved domains。
- dedup reason。
## 6. 与旧版设计的差异
| 旧版设想 | 当前实现 |
|---|---|
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Verifier;AIOps: Supervisor + Planner + Executor |
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
| Skill 驱动不同诊断流程 | 当前以 Prompt、知识域地图、工具调用和评测 baseline 控制 |
## 7. 后续演进
当诊断场景和工具复杂度继续上升时,再考虑拆分:
- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
拆分前提:
- 当前 Executor prompt 已难以维护。
- 不同故障类型的工具权限明显不同。
- Trace 能证明某类问题需要独立的推理策略。
- 评测集能覆盖拆分前后的行为差异。
@@ -0,0 +1,28 @@
# 旧版架构文档归档
**归档日期**:2026-07-05
本目录保存 `mvp/architecture` 下的旧版架构文档。它们包含早期 MVP 设计、旧 RAG 方案、会话存储设计、行动记忆和实施计划等历史材料。
这些文档不再作为当前实现依据。当前架构请阅读:
- `mvp/architecture/README.md`
- `mvp/architecture/current-mvp-architecture.md`
- `mvp/architecture/rag-architecture.md`
## 归档文件
| 文件 | 说明 |
|---|---|
| `agent-architecture.md` | 早期完整 Agent 设想,包含较多超出当前 MVP 的 SubAgent 设计 |
| `agent-architecture-mvp.md` | 早期 MVP Agent 设计 |
| `knowledge-retrieval-architecture.md` | 旧版 L0 + L1 检索架构,包含 L0 唯一命中跳过 L1 的旧逻辑 |
| `knowledge-retrieval-usage.md` | 旧版知识库检索使用说明 |
| `current-mvp-architecture.md` | 归档前的当前架构快照 |
| `implementation-plan.md` | 早期实施计划 |
| `implementation-detail.md` | 早期完整实施计划 |
| `session-management.md` | 会话管理旧设计 |
| `session-dedup-knowledge-map.md` | 会话去重和知识域地图设计 |
| `confidence-feedback.md` | 证据评分和用户反馈旧设计 |
| `action-memory-relevance.md` | 行动记忆和检索质量归一化旧设计 |
@@ -0,0 +1,275 @@
# 行动记忆与检索质量归一化
Executor 行动记忆 + 归一化质量等级设计,解决 ISS-002 Executor 无约束重复检索问题。
---
## 一、问题背景
ISS-001 修复文档级去重后,Executor 在单次会话中仍调用 `lookup_knowledge` 20+ 次。根因:
1. **行动记忆缺失**:Executor 不知道自己已检索过哪些域
2. **质量信号缺失**:检索结果没有给 LLM 判断"结果够不够"的信号
3. **Prompt 缺少合法出口**:原 prompt 要求"所有外部信息都必须调用工具",LLM 不敢停止检索
---
## 二、整体架构
```
lookup_knowledge(query)
│
├─ Step 1: L0 精确匹配(keywords 索引)
├─ Step 2: L1 语义检索(Milvus 向量)
├─ Step 3: computeRelevance()
│ ├─ 归一化:L2 → similarity [0,1]
│ └─ 判定:PRECISE / HIGHLY_RELEVANT / REFERENCE
├─ Step 4: RetrievedDocTracker 检查
│ ├─ 文档级去重 → isDocRetrieved(sessionId, docKey)
│ ├─ 域级检查 → isDomainRetrieved(sessionId, domain)
│ └─ 记录 → markRetrieved(sessionId, domain, docKey)
└─ Step 5: 返回 LookupResult
├─ primary / supplement(原始内容,不含分数)
├─ relevanceLevel(PRECISE / HIGHLY_RELEVANT / REFERENCE)
├─ completenessHint(兜底信号)
└─ retrievedDomainsThisSession(行动记忆)
```
### 设计原则
| 原则 | 说明 |
|------|------|
| **Agent 边界清晰** | 不给 Executor 注入 knowledge map,Executor 只知道做了什么,不用知道有什么 |
| **分数封装** | L0/L1 原始分数不在 LookupResult 中返回 LLM,只在归一化层内部使用 |
| **原始分数只入库** | 原始 L2 距离写进 `tool_invocation.retrieval_details` JSON 用于可观测 |
| **软约束 + 硬拦截** | Prompt 约束(软)+ 工具层域级去重(硬)两层防御 |
---
## 三、归一化质量等级
### L2 距离归一化
BGE-M3 输出为 L2 归一化单位向量(实测范数=1.00000002),L2 距离数学硬上界 = 2.0。
```
similarity = 1 - min(l2Score, maxL2Distance) / maxL2Distance
```
| L2 距离 | similarity | 等级 |
|---------|-----------|------|
| 0.0 | 1.0 | PRECISE |
| 0.383 | 0.8085 | HIGHLY_RELEVANT |
| 0.5 | 0.75 | HIGHLY_RELEVANT |
| 0.6031 | 0.6984 | REFERENCE |
| 1.0 | 0.5 | REFERENCE 边界 |
| 2.0+ | 0.0 | 不视为有效结果 |
### 三等级判定
| 等级 | 条件 | completenessHint | LLM 行为 |
|------|------|-----------------|---------|
| PRECISE | L0 matchCount == 1 | "知识库中不存在比上述结果更精准的文档" | 直接使用,禁止再检索 |
| HIGHLY_RELEVANT | L0 命中 + similarity ≥ 0.75,或仅 L1 similarity ≥ 0.75 | "当前结果已高度相关,继续检索不太可能找到更精准的文档" | 可综合推理,大概率不需要继续查 |
| REFERENCE | 其余命中(similarity ≥ 0.5) | "当前结果为相关参考,如需更精准信息请明确缺少的具体维度" | 可参考,如需更精准请指出缺少的维度后定向补充 |
### 阈值配置
```yaml
retrieval:
normalization:
max-l2-distance: 2.0 # L2 距离上界
highly-relevant-threshold: 0.75 # similarity ≥ 0.75 → HIGHLY_RELEVANT
reference-threshold: 0.5 # similarity ≥ 0.5 → REFERENCE
```
---
## 四、行动记忆
### RetrievedDocTracker 数据结构
```java
// 从单层升级为双层:session → domain → filePath 集合
ConcurrentHashMap<String, Map<String, Set<String>>> sessionRetrievals;
```
### API
| 方法 | 作用 |
|------|------|
| `markRetrieved(sessionId, domain, filePath)` | 记录一次检索 |
| `isDocRetrieved(sessionId, filePath)` | 文档级去重 |
| `isDomainRetrieved(sessionId, domain)` | 域级检查 |
| `getRetrievedDomains(sessionId)` | 获取已检索域列表 |
| `clearSession(sessionId)` | 清理会话记录 |
### LookupResult 返回
```java
LookupResult.builder()
.found(true)
.primary(primaryResult)
.supplement(supplementResult)
.relevanceLevel("HIGHLY_RELEVANT") // PRECISE / HIGHLY_RELEVANT / REFERENCE
.completenessHint("当前结果已高度相关...") // 兜底信号
.retrievedDomainsThisSession(["infrastructure", "api"]) // 行动记忆
.message("...")
.build();
```
---
## 五、Executor Prompt 约束
### 4 条检索约束
1. **判断重复**:基于 `retrievedDomainsThisSession` 判断语义重叠
2. **重复了怎么办**:禁止换关键词重查;先指缺少的维度,再定向补充
3. **合法出口**:"不查全不会被追责,重复检索才会被惩罚"
4. **利用质量信号**:PRECISE → 停止;HIGHLY_RELEVANT + 域已检索 → 禁止;REFERENCE → 指出缺少维度
### 关键变化
原有 prompt:"所有需要外部信息的地方,都必须调用对应的工具"
→ 改为:"需要外部信息时调用工具,但须遵守下方的检索约束"
---
## 六、数据库变更
### V010
```sql
ALTER TABLE tool_invocation
ADD COLUMN relevance_level VARCHAR(20) COMMENT 'PRECISE/HIGHLY_RELEVANT/REFERENCE/DEDUPED',
ADD COLUMN dedup_reason VARCHAR(32) COMMENT 'doc_retrieved/domain_retrieved/null';
```
### retrieval_details JSON 扩展
```json
{
"l0_titles": ["MySQL 数据库连接池配置", "Redis 缓存配置指南"],
"l1_scores": [0.383, 0.4502, 0.7011],
"l1_top_score": 0.383,
"l1_top_similarity": 0.8085,
"relevance_level": "HIGHLY_RELEVANT",
"completeness_hint": "当前结果已高度相关,继续检索不太可能找到更精准的文档",
"retrieved_domains": ["infrastructure"]
}
```
扩展字段使用方式:
| 字段 | 用途 |
|------|------|
| `l1_top_score` | 原始 L2 距离最小值(可观测性) |
| `l1_top_similarity` | 归一化后的相似度 [0,1] |
| `relevance_level` | 归一化质量等级 |
| `completeness_hint` | 兜底信号 |
| `retrieved_domains` | 已检索域列表 |
| `dedup_reason` | 去重原因(如有) |
---
## 七、使用场景
### 场景 1:正常检索
```
用户:数据库连接池怎么配置?
Executor 内部:
1. lookup_knowledge("数据库连接池配置")
→ relevanceLevel=HIGHLY_RELEVANT (similarity=0.8085)
→ completenessHint="当前结果已高度相关..."
→ retrievedDomainsThisSession=["infrastructure"]
2. 基于已有信息直接回答,不再检索
```
### 场景 2:行动记忆阻止重复
```
Executor 步骤列表:
- 查数据库连接池配置
- 查 HikariCP 参数
- 查连接池耗尽排查
实际行为:
1. lookup("数据库连接池") → relevance=HIGHLY_RELEVANT, domains=["infrastructure"]
2. lookup("HikariCP 参数") → retrievedDomainsThisSession=["infrastructure"]
LLM 判断:infrastructure 域已检索过,禁止换关键词重查
→ 基于已有信息回答,指出缺少的具体维度
3. lookup("连接池耗尽") → 同域,被 prompt 约束拦截或工具层去重拦截
```
### 场景 3:PRECISE 精确匹配
```
用户:ERR_TIMEOUT 是什么?
Executor 内部:
1. lookup_knowledge("ERR_TIMEOUT")
→ L0 matchCount=1(唯一精确匹配)
→ relevanceLevel=PRECISE
→ completenessHint="知识库中不存在比上述结果更精准的文档"
2. 直接使用,不再检索
```
### 场景 4:REFERENCE + 定向补充
```
用户:如何排查生产故障?
Executor 内部:
1. lookup_knowledge("故障排查")
→ relevanceLevel=REFERENCE (similarity=0.6)
→ retrievedDomainsThisSession=["troubleshooting"]
2. LLM 判断:信息不足,缺少"日志分析"维度的具体步骤
3. lookup_knowledge("日志分析步骤")
→ 定向补充,不盲目换关键词
```
---
## 八、可观测性
### 查询质量分布
```sql
SELECT relevance_level, COUNT(*) AS cnt
FROM tool_invocation
WHERE tool_name = 'lookup_knowledge'
GROUP BY relevance_level;
```
### 去重原因分布
```sql
SELECT dedup_reason, COUNT(*) AS cnt
FROM tool_invocation
WHERE tool_name = 'lookup_knowledge'
GROUP BY dedup_reason;
```
### 归一化分数分布
```sql
SELECT
JSON_EXTRACT(retrieval_details, '$.l1_top_similarity') AS similarity,
COUNT(*) AS cnt
FROM tool_invocation
WHERE tool_name = 'lookup_knowledge'
AND retrieval_details IS NOT NULL
GROUP BY similarity
ORDER BY similarity;
```
---
## 九、扩展方向(Phase 2)
- **域级硬限流**:`isDomainRetrieved` 已就绪,在 LookupKnowledgeTool 入口直接拦截同域调用,不依赖 LLM 遵守 prompt
- **DEDUPED 等级**:去重时单独标记为 DEDUPED 等级,与 REFERENCE 区分
- **分数反馈调优**:基于 feedback 数据优化归一化阈值
@@ -0,0 +1,157 @@
# 证据评分与用户反馈架构
## 一、整体架构
```
用户对话
↓
ChatService.executeChat / executeChatComplex
↓ SUCCESS 后写入 answer,异步触发
EvaluationService.evaluate(sessionId, answer)
└─ 读取 tool_invocation 事实 → 规则引擎 → 写 selfEvaluation
用户提交反馈
↓
POST /api/feedback { sessionId, feedback: "useful" | "not_useful" }
↓
FeedbackService.submitFeedback
├─ 写 DiagnosisSession.feedback
├─ useful → CaseLibraryService.createFromSession → 写 case_library
└─ not_useful → 仅写 feedback,status 不变
```
---
## 二、评分规则(evidence_score)
### 定位
`evidence_score` 衡量的是**证据收集充分度**,不是答案准确性。
- 能证明的:Agent 是否有尝试收集证据、检索是否命中
- 不能证明的:答案是否有幻觉、推理是否正确
### 数据来源
规则引擎只消费 `tool_invocation` 表的事实记录,不依赖 LLM 判断。
### 规则定义
| 规则名 | 条件 | delta |
|---|---|---|
| `no_tool_call` | 无任何工具调用 | 直接 0 分,不参与加权 |
| `execution_failed` | status = FAILED | 直接 0 分,不参与加权 |
| `has_successful_tool_call` | 至少 1 次成功调用 | +30 |
| `l0_exact_match` | 任意调用有 L0 精确匹配命中 | +35 |
| `l1_semantic_match` | 无 L0 命中但有 L1 语义匹配 | +20 |
| `retrieval_no_hit` | 有检索调用但无任何命中 | -10 |
| `all_tool_calls_failed` | 全部调用失败 | -20 |
> L0 和 L1 互斥取高优先级(L0 命中时跳过 L1 分支)。
### selfEvaluation 字段格式
```json
{
"evidence_score": 65,
"source": "rule",
"factors": [
{"name": "has_successful_tool_call", "delta": 30, "description": "有成功的工具调用(20次)"},
{"name": "l0_exact_match", "delta": 35, "description": "L0 精确匹配命中"}
]
}
```
| 字段 | 说明 |
|---|---|
| `evidence_score` | 0-100 整数 |
| `source` | 当前固定为 `"rule"`;预留 `"llm"` 供后续扩展 |
| `factors` | 命中的规则列表,含 name / delta / description |
| `llm_opinion` | 预留字段(未实现),LLM 观点叠加时在此扩展 |
### 已知边界
- 非检索工具(DateTimeTools、QueryMetricsTools 等)不写 `tool_invocation`,这类 session 的 evidence_score = 0,属于设计边界
- 评分为异步写入(`@Async`),失败时 `selfEvaluation` 保持 null,前端需处理 null
---
## 三、反馈机制
### API
```
POST /api/feedback
Content-Type: application/json
{
"sessionId": "xxx",
"feedback": "useful" | "not_useful"
}
```
**响应**
```json
{
"success": true,
"message": "反馈已记录",
"caseId": "uuid 或 null"
}
```
### 后端行为
| feedback 值 | 操作 |
|---|---|
| `useful` | 写 `DiagnosisSession.feedback = "useful"`,生成 `CaseLibrary` 记录,返回 caseId |
| `not_useful` | 写 `DiagnosisSession.feedback = "not_useful"`,status 不变 |
| 其他值 | 返回 HTTP 400 |
### 重要设计决策
**BAD_CASE 不改 status 字段**
`status` 表示执行状态(RUNNING/SUCCESS/FAILED),是独立维度,不能被质量标签覆盖。
查询 BadCase 使用:`WHERE feedback = 'not_useful'`
**useful 触发案例沉淀规则**
| CaseLibrary 字段 | 来源 |
|---|---|
| caseId | UUID |
| diagnosisId | DiagnosisSession.sessionId |
| sourceType | AUTO |
| faultCategory | GENERAL(暂时,后续人工补充) |
| title | query 前 100 字符 |
| rootCause / solution | DiagnosisSession.answer(完整答案) |
| createdBy | "system" |
**幂等性**:同一 sessionId 重复提交 useful,返回已有 caseId,不重复插入 case_library。
---
## 四、数据库变更
### V008(新增)
```sql
ALTER TABLE diagnosis_session ADD COLUMN answer LONGTEXT COMMENT 'Agent 返回给用户的完整答案';
```
### diagnosis_session 关键字段
| 字段 | 类型 | 说明 |
|---|---|---|
| `answer` | LONGTEXT | Agent 完整回答,useful 案例沉淀的内容来源 |
| `self_evaluation` | JSON | 证据评分结果,格式见上 |
| `feedback` | VARCHAR(16) | useful / not_useful / null |
| `status` | VARCHAR(16) | 执行状态,不受 feedback 影响 |
---
## 五、扩展方向(Phase 2)
- **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑
- **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../issues/ISS-001-duplicate-retrieval.md))
@@ -0,0 +1,220 @@
# Current MVP Architecture Snapshot
**Updated**: 2026-07-05
This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.
## 1. Positioning
The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.
Core goals:
- Support normal chat-based diagnosis.
- Support AIOps alert-triggered diagnosis.
- Keep tool calls explicit and traceable.
- Keep RAG retrieval observable through `lookup_knowledge`.
- Persist enough execution evidence for replay, evaluation, and interview explanation.
## 2. Runtime Architecture
```text
HTTP API
-> ChatService / AiOpsService
-> Agent orchestration
-> Supervisor / Planner / Executor / Verifier
-> Tools
-> lookup_knowledge
-> query_logs
-> query_metrics
-> other diagnosis tools
-> Persistence
-> diagnosis_session
-> agent_step
-> tool_invocation
-> Trace API
-> DiagnosisTraceService
```
Current entry points:
- `ChatService`: user-driven troubleshooting and follow-up diagnosis.
- `AiOpsService`: alert-driven diagnosis, including payload mode and auto-discovery mode.
- `DiagnosisTraceService`: trace view of session, steps, tool calls, and self-evaluation.
## 3. Chat Diagnosis Flow
```text
User question
-> ChatService
-> simple response or diagnosis flow
-> Planner creates investigation direction
-> Executor calls tools for evidence
-> lookup_knowledge
-> query_logs
-> query_metrics
-> Verifier checks final diagnosis quality
-> self_evaluation.verifier_evaluation
-> diagnosis trace
```
The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under `diagnosis_session.self_evaluation.verifier_evaluation`.
## 4. AIOps Diagnosis Flow
```text
AIOps request
-> AiOpsService
-> payload mode or auto-discovery mode
-> build alert-focused diagnosis prompt
-> append recommended lookup_knowledge query when payload exists
-> Agent diagnosis flow
-> Supervisor / Planner / Executor
-> evidence tools
-> final report
-> AiOpsRuleEvaluationService
-> self_evaluation.aiops_rule_evaluation
-> diagnosis trace
```
AIOps keeps two modes:
- Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
- Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.
The AIOps verifier is currently lightweight and rule-based. It checks:
- Whether the final report exists.
- Whether the result stays focused on the alert payload when payload exists.
- Whether evidence tools were used, especially `lookup_knowledge`, `query_logs`, and `query_metrics`.
## 5. RAG Architecture
```text
lookup_knowledge
-> L0 domain/entity hint
-> matched domain
-> matched keywords/entities
-> metadata filter signal
-> VectorSearchService
-> Spring AI VectorStore path
-> Milvus SDK fallback path
-> evidence post-processing
-> score / rawScore / scoreLabel
-> source metadata
-> title / breadcrumb / content evidence block
-> tool_invocation record
```
Important decisions:
- `lookup_knowledge` remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
- L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
- L1 retrieval now goes through `VectorSearchService`.
- Spring AI `VectorStore` is the preferred retrieval path.
- The original Milvus SDK path is retained as fallback and compatibility path.
- `title`, `breadcrumb`, and `content` participate in embedding text so chunk context is less likely to be lost.
- Retrieval output keeps compatibility fields: `score`, `rawScore`, and `scoreLabel`.
Vector retrieval modes:
```text
retrieval.vector-store.mode=auto # Prefer Spring AI VectorStore, fallback to SDK
retrieval.vector-store.mode=spring-ai # Use Spring AI VectorStore only
retrieval.vector-store.mode=sdk # Use original Milvus SDK path
```
## 6. Persistence And Trace
Current trace-related persistence:
```text
diagnosis_session
-> final_report
-> self_evaluation
-> verifier_evaluation
-> aiops_rule_evaluation
agent_step
-> role
-> step input/output
-> execution order
tool_invocation
-> tool_name
-> query
-> retrieval_layer
-> retrieval_details
-> evidence blocks
-> duration
```
Trace API aggregates these records into a session-level view:
- Agent step sequence.
- Tool calls and retrieval details.
- Final diagnosis report.
- Chat verifier status.
- AIOps rule verifier status.
## 7. Quality Gates
Current quality gates:
- Chat verifier: LLM-based final answer verification for normal diagnosis.
- AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
- Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
- RAG retrieval baseline: golden query set with offline baseline report.
- Live RAG acceptance: post-reindex script for validating retrieval against the running stack.
These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.
## 8. Current Completion State
Completed for the current MVP stage:
- Explicit `lookup_knowledge` Agent tool.
- L0 + L1 retrieval shape retained.
- L0 downgraded to domain/entity hint.
- Spring AI VectorStore retrieval path integrated.
- Milvus SDK fallback retained.
- RAG evidence post-processing added.
- Breadcrumb/title/content embedding text improved.
- RAG offline baseline and live acceptance script added.
- AIOps payload query augmentation added.
- AIOps lightweight verifier added.
- Trace summary includes both chat verifier and AIOps verifier signals.
Deferred future enhancements:
- LLM QueryTransformer / MultiQuery.
- BM25, RRF, and reranker.
- Neighbor chunk or section-level context expansion.
- VectorStore write path migration.
- Full LLM-based AIOps verifier.
- More complete golden set for recall, MRR, and nDCG metrics.
## 9. Key Code References
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
## 10. Supporting Materials
- `mvp/issues/rag-refactor-plan.md`
- `eval/rag-retrieval/README.md`
- `scripts/eval_rag_live_acceptance.py`
- `interview/rag-refactor-story.md`
- `interview/rag-vectorstore-interview-notes.md`
- `interview/rag-retrieval-quality-report.md`
- `interview/rag-breadcrumb-embedding-acceptance.md`
- `interview/aiops-query-augmentation.md`
- `interview/aiops-lightweight-verifier.md`

Some files were not shown because too many files have changed in this diff Show More