docs: reorganize MVP interview documentation
This commit is contained in:
@@ -1,7 +1,7 @@
|
|||||||
<!-- gitnexus:start -->
|
<!-- gitnexus:start -->
|
||||||
# GitNexus — Code Intelligence
|
# GitNexus — Code Intelligence
|
||||||
|
|
||||||
This project is indexed by GitNexus as **SuperBizAgent-java** (1528 symbols, 2828 relationships, 87 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||||
|
|
||||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||||
|
|
||||||
|
|||||||
@@ -115,7 +115,7 @@ trailing off into the following information in 99% of cases:
|
|||||||
<!-- gitnexus:start -->
|
<!-- gitnexus:start -->
|
||||||
# GitNexus — Code Intelligence
|
# GitNexus — Code Intelligence
|
||||||
|
|
||||||
This project is indexed by GitNexus as **SuperBizAgent-java** (1001 symbols, 2043 relationships, 78 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||||
|
|
||||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||||
|
|
||||||
|
|||||||
+32
-25
@@ -1,48 +1,54 @@
|
|||||||
# SuperBizAgent Interview Guide
|
# SuperBizAgent 面试资料包
|
||||||
|
|
||||||
## 一句话定位
|
## 一句话定位
|
||||||
|
|
||||||
SuperBizAgent 是一个面向企业故障诊断场景的 Agent Engineering 项目:它把用户问题或告警事件转成可追踪的多 Agent 执行链路,并把工具证据、模型步骤、最终答案和反馈统一落到诊断 trace 中。
|
SuperBizAgent 是一个面向企业故障诊断场景的 Agent 工程项目。它把用户问题或 AIOps 告警转换成可追踪的 Agent 执行链路,并把工具证据、模型步骤、最终答案、自评估和用户反馈统一沉淀到诊断 Trace 中。
|
||||||
|
|
||||||
## 面试重点
|
## 面试重点
|
||||||
|
|
||||||
- **多 Agent 编排**:普通 Chat 的复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 Supervisor 调度 Planner/Executor。
|
- **Agent 编排**:Chat 复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 `Supervisor -> Planner / Executor`。
|
||||||
- **工具证据链**:知识库、日志、指标和 Prometheus 告警都通过工具调用进入链路,并记录到 `tool_invocation`。
|
- **工具证据链**:知识库、日志、指标、Prometheus 告警都通过显式工具调用进入链路,并记录到 `tool_invocation`。
|
||||||
- **可追踪诊断**:一次会话对应一个 `sessionId`,最终可以通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
|
- **可追踪诊断**:一次诊断对应一个 `sessionId`,可通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
|
||||||
- **质量门**:Chat 链路包含 Verifier,把 groundedness、facts checked 和 evidence refs 写回 `diagnosis_session.self_evaluation`。
|
- **质量门禁**:Chat Verifier 校验 groundedness;AIOps 规则评估检查报告完整性、payload 聚焦和证据工具覆盖。
|
||||||
- **AIOps 产品边界**:有告警 payload 时聚焦该告警;没有 payload 时先自动发现 active alerts。
|
- **RAG 工程化**:`lookup_knowledge` 是显式 Agent Tool,底层通过 Spring AI VectorStore 主路径 + Milvus SDK fallback。
|
||||||
- **可复现 Demo**:`mvp-demo` profile 使用 mock Prometheus 和 mock CLS,让面试演示不依赖真实线上故障。
|
- **反馈闭环**:用户反馈 `useful` 会沉淀 `case_library`,`not_useful` 保留 bad case 信号。
|
||||||
|
|
||||||
## 推荐阅读顺序
|
## 推荐阅读顺序
|
||||||
|
|
||||||
1. `interview/demo-script.md`:面试现场怎么讲、怎么演示。
|
1. `mvp/architecture/interview-one-pager.md`:一页式架构图和 2-5 分钟讲解。
|
||||||
2. `interview/architecture.md`:系统架构和两条主链路。
|
2. `mvp/demo/ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||||
3. `interview/design-tradeoffs.md`:关键设计取舍和可被追问的问题。
|
3. `interview/story-cases.md`:可复用的面试故事案例。
|
||||||
4. `interview/acceptance-checklist.md`:面试前验证清单。
|
4. `interview/architecture.md`:面试版系统架构。
|
||||||
5. `mvp/demo/README.md`:更细的 MVP 可执行 runbook。
|
5. `interview/design-tradeoffs.md`:关键设计取舍。
|
||||||
|
6. `interview/demo-script.md`:更细的命令式演示脚本。
|
||||||
|
7. `interview/acceptance-checklist.md`:面试前验收清单。
|
||||||
|
8. RAG 专题文档:`rag-refactor-story.md`、`rag-vectorstore-interview-notes.md`、`rag-retrieval-quality-report.md`。
|
||||||
|
|
||||||
## 核心 Demo
|
## 核心演示链路
|
||||||
|
|
||||||
### Chat Diagnosis
|
### Chat 诊断
|
||||||
|
|
||||||
```text
|
```text
|
||||||
POST /api/chat
|
POST /api/chat
|
||||||
-> ChatService.executeChatWithStrategy(...)
|
-> ChatService
|
||||||
-> simple ReactAgent or Planner -> Executor -> Verifier
|
-> Planner -> Executor -> Verifier
|
||||||
-> lookup_knowledge / query_logs / query_metrics
|
-> lookup_knowledge / query_logs / query_metrics
|
||||||
-> diagnosis_session + agent_step + tool_invocation
|
-> diagnosis_session + agent_step + tool_invocation
|
||||||
-> GET /api/diagnosis/{sessionId}/trace
|
-> GET /api/diagnosis/{sessionId}/trace
|
||||||
|
-> POST /api/feedback
|
||||||
```
|
```
|
||||||
|
|
||||||
### AIOps Alert Diagnosis
|
### AIOps 告警诊断
|
||||||
|
|
||||||
```text
|
```text
|
||||||
POST /api/ai_ops
|
POST /api/ai_ops
|
||||||
-> AiOpsService.executeAiOpsAnalysis(...)
|
-> AiOpsService
|
||||||
|
-> PAYLOAD_TARGETED / AUTO_DISCOVERY
|
||||||
-> ai_ops_supervisor
|
-> ai_ops_supervisor
|
||||||
-> planner_agent / executor_agent
|
-> planner_agent / executor_agent
|
||||||
-> queryPrometheusAlerts + logs + knowledge
|
-> queryPrometheusAlerts + logs + metrics + lookup_knowledge
|
||||||
-> scoped alert report
|
-> alert report
|
||||||
|
-> aiops_rule_evaluation
|
||||||
-> GET /api/diagnosis/{sessionId}/trace
|
-> GET /api/diagnosis/{sessionId}/trace
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -50,10 +56,11 @@ POST /api/ai_ops
|
|||||||
|
|
||||||
- Chat 诊断链路:可运行、可追踪、有 Verifier。
|
- Chat 诊断链路:可运行、可追踪、有 Verifier。
|
||||||
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
|
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
|
||||||
|
- RAG 检索链路:Spring AI VectorStore 主路径、Milvus SDK fallback、L0 hint、检索评测 baseline。
|
||||||
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
|
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
|
||||||
- Demo 文档:`mvp/demo/README.md` 和 `mvp/demo/aiops-alert-acceptance.md`。
|
- Demo 材料:`mvp/demo/README.md`、`mvp/demo/ten-minute-interview-demo.md`。
|
||||||
- Devflow 沉淀:`devflow/index.md` 记录了 MVP、Verifier、AIOps trace 和 AIOps scope-control 的演进。
|
|
||||||
|
|
||||||
## 面试时的主叙事
|
## 主叙事
|
||||||
|
|
||||||
|
这个项目不是简单调用大模型,而是在做一个可审计、可验证、可回归的 Agent 诊断系统。模型可以规划和推理,但每一步工具证据、最终结论、Verifier 结果和用户反馈都能被 Trace API 回放。面试时重点展示“从问题到证据到答案到验证再到反馈”的闭环。
|
||||||
|
|
||||||
这个项目不是简单调用大模型,而是在做一个可审计的 Agent 诊断系统。核心价值是:模型可以规划和推理,但每一步工具证据、最终结论和质量评估都能被 trace API 回放。面试时重点展示“从问题到证据到答案到验证”的完整闭环。
|
|
||||||
|
|||||||
@@ -1,15 +1,15 @@
|
|||||||
# Acceptance Checklist
|
# 面试前验收清单
|
||||||
|
|
||||||
## 面试前环境检查
|
## 1. 环境检查
|
||||||
|
|
||||||
- 当前分支包含最新 AIOps trace/scope 变更。
|
- 当前分支包含最新架构文档和面试材料。
|
||||||
- MySQL 可连接。
|
- MySQL 可连接。
|
||||||
- Redis 可连接。
|
- Redis 可连接。
|
||||||
- Milvus/Zilliz 可连接。
|
- Milvus/Zilliz 可连接。
|
||||||
- 模型 API key 可用。
|
- 模型 API key 可用。
|
||||||
- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。
|
- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。
|
||||||
|
|
||||||
启动:
|
启动服务:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||||
@@ -24,10 +24,10 @@ mvn -q -DskipTests compile
|
|||||||
目标测试:
|
目标测试:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test
|
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest,VectorSearchServiceTest,LookupKnowledgeToolTest" test
|
||||||
```
|
```
|
||||||
|
|
||||||
## Chat Demo 验收
|
## 2. Chat Demo 验收
|
||||||
|
|
||||||
请求:
|
请求:
|
||||||
|
|
||||||
@@ -50,7 +50,7 @@ Invoke-RestMethod `
|
|||||||
- 返回 `data.success = true`。
|
- 返回 `data.success = true`。
|
||||||
- 返回 `data.sessionId = interview-chat-payment-timeout-001`。
|
- 返回 `data.sessionId = interview-chat-payment-timeout-001`。
|
||||||
- `diagnosis_session.agent_flow = CHAT`。
|
- `diagnosis_session.agent_flow = CHAT`。
|
||||||
- trace API 返回 session、steps、toolInvocations。
|
- Trace API 返回 session、steps、toolInvocations。
|
||||||
- 复杂问题下 trace 中能看到 verifier 相关数据。
|
- 复杂问题下 trace 中能看到 verifier 相关数据。
|
||||||
|
|
||||||
SQL:
|
SQL:
|
||||||
@@ -59,7 +59,7 @@ SQL:
|
|||||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'"
|
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'"
|
||||||
```
|
```
|
||||||
|
|
||||||
## AIOps Demo 验收
|
## 3. AIOps Demo 验收
|
||||||
|
|
||||||
请求:
|
请求:
|
||||||
|
|
||||||
@@ -89,11 +89,10 @@ Invoke-WebRequest `
|
|||||||
- `diagnosis_session.agent_flow = AI_OPS`。
|
- `diagnosis_session.agent_flow = AI_OPS`。
|
||||||
- `diagnosis_session.status = SUCCESS`。
|
- `diagnosis_session.status = SUCCESS`。
|
||||||
- `diagnosis_session.answer` 有最终报告。
|
- `diagnosis_session.answer` 有最终报告。
|
||||||
- trace API 返回 AIOps steps 和 tool invocations。
|
- Trace API 返回 AIOps steps 和 tool invocations。
|
||||||
- 报告主章节聚焦 `HighCPUUsage/payment-service`。
|
- 报告主章节聚焦 `HighCPUUsage/payment-service`。
|
||||||
- 无 `告警根因分析 - HighMemoryUsage` 独立章节。
|
- 其他 active alerts 不应展开成独立主根因章节。
|
||||||
- 无 `告警根因分析 - SlowResponse` 独立章节。
|
- `self_evaluation.aiops_rule_evaluation` 存在。
|
||||||
- 有“相关风险告警”或类似上下文说明。
|
|
||||||
|
|
||||||
SQL:
|
SQL:
|
||||||
|
|
||||||
@@ -108,19 +107,10 @@ python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invoc
|
|||||||
Scope 检查:
|
Scope 检查:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
python scripts/query_mysql.py "SELECT (answer LIKE '%告警根因分析 - HighCPUUsage%') AS has_main_root_cause, (answer LIKE '%告警根因分析 - HighMemoryUsage%') AS has_memory_root_cause, (answer LIKE '%告警根因分析 - SlowResponse%') AS has_slow_root_cause, (answer LIKE '%相关风险告警%') AS has_related_risk FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
python scripts/query_mysql.py "SELECT (answer LIKE '%HighCPUUsage%') AS has_main_alert, (answer LIKE '%payment-service%') AS has_service FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
||||||
```
|
```
|
||||||
|
|
||||||
期望:
|
## 4. Trace API 验收
|
||||||
|
|
||||||
```text
|
|
||||||
has_main_root_cause = 1
|
|
||||||
has_memory_root_cause = 0
|
|
||||||
has_slow_root_cause = 0
|
|
||||||
has_related_risk = 1
|
|
||||||
```
|
|
||||||
|
|
||||||
## Trace API 验收
|
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
@@ -128,13 +118,28 @@ Invoke-RestMethod `
|
|||||||
-Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
-Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||||
```
|
```
|
||||||
|
|
||||||
若 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
|
如果 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||||
```
|
```
|
||||||
|
|
||||||
## 常见问题
|
## 5. RAG 验收
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Uri "http://127.0.0.1:9900/api/search/similar?query=ERR_TIMEOUT&topK=3" `
|
||||||
|
-Method Get
|
||||||
|
```
|
||||||
|
|
||||||
|
验收:
|
||||||
|
|
||||||
|
- 返回 `code = 200`。
|
||||||
|
- top candidates 中包含 `ERR_TIMEOUT` 相关文档。
|
||||||
|
- `scoreLabel` 能体现当前检索路径语义。
|
||||||
|
- 如果走 VectorStore,日志应出现 Spring AI VectorStore search。
|
||||||
|
|
||||||
|
## 6. 常见问题
|
||||||
|
|
||||||
### MySQL stale connection
|
### MySQL stale connection
|
||||||
|
|
||||||
@@ -145,26 +150,17 @@ HikariPool - Connection is not available
|
|||||||
No operations allowed after connection closed
|
No operations allowed after connection closed
|
||||||
```
|
```
|
||||||
|
|
||||||
当前已在 `application.yml` 配置:
|
|
||||||
|
|
||||||
- `maximum-pool-size: 5`
|
|
||||||
- `minimum-idle: 1`
|
|
||||||
- `connection-timeout: 10000`
|
|
||||||
- `validation-timeout: 5000`
|
|
||||||
- `idle-timeout: 60000`
|
|
||||||
- `max-lifetime: 120000`
|
|
||||||
- `keepalive-time: 30000`
|
|
||||||
|
|
||||||
处理:
|
处理:
|
||||||
|
|
||||||
- 重新编译或重启服务。
|
- 重启服务。
|
||||||
- 确认日志中新的 HikariPool 启动成功。
|
- 确认 HikariPool 使用当前配置启动成功。
|
||||||
- 再跑 trace 或 AIOps 请求。
|
- 再跑 trace 或 AIOps 请求。
|
||||||
|
|
||||||
### SSE 客户端显示异常
|
### SSE 客户端显示异常
|
||||||
|
|
||||||
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 trace API 验证结果。
|
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 Trace API 验证结果。
|
||||||
|
|
||||||
### OpenSpec 全量校验失败
|
### OpenSpec 全量校验失败
|
||||||
|
|
||||||
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试材料主要依赖已归档的 AIOps spec 和 MVP trace spec,可以单独验证相关 spec。
|
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试演示主要依赖已归档 spec、MVP trace 和 RAG 验收材料,可以单独验证相关 spec。
|
||||||
|
|
||||||
|
|||||||
@@ -1,38 +1,37 @@
|
|||||||
# AIOps Lightweight Verifier
|
# AIOps 轻量规则验证器
|
||||||
|
|
||||||
## What Changed
|
## 1. 改动是什么
|
||||||
|
|
||||||
AIOps now has a deterministic post-run quality gate.
|
AIOps 现在有一个确定性的后置质量门禁。最终告警报告持久化后,`AiOpsRuleEvaluationService` 会检查:
|
||||||
|
|
||||||
After the final AIOps report is persisted, the service evaluates:
|
- 最终报告是否存在,且不是明显过短。
|
||||||
|
- payload 模式下,报告是否提到输入的告警和服务。
|
||||||
|
- 是否有证据工具调用,例如 `lookup_knowledge`、`query_metrics`、`query_logs`。
|
||||||
|
|
||||||
- whether the final report exists and is not trivially short
|
结果写入:
|
||||||
- whether a payload-targeted report mentions the supplied alert and service
|
|
||||||
- whether evidence tools such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted
|
|
||||||
|
|
||||||
The result is stored under:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||||
```
|
```
|
||||||
|
|
||||||
The trace API returns this payload through the existing session self-evaluation field.
|
Trace API 会通过 session self-evaluation 展示这个结果。
|
||||||
|
|
||||||
## Why Rule-Based First
|
## 2. 为什么先做规则型
|
||||||
|
|
||||||
This is not a full LLM verifier yet.
|
这还不是完整 LLM Verifier。
|
||||||
|
|
||||||
The first AIOps quality risks are concrete and easy to check with rules:
|
AIOps 第一阶段质量风险比较具体,适合先用规则:
|
||||||
|
|
||||||
- Did the report stay focused on the payload?
|
- 报告有没有生成。
|
||||||
- Did the run use evidence tools?
|
- 报告有没有聚焦 payload。
|
||||||
- Did the system produce a usable final report?
|
- 有没有使用证据工具。
|
||||||
|
- 有没有把无关告警展开成主诊断对象。
|
||||||
|
|
||||||
Rule evaluation is stable, cheap, and easy to explain. It also avoids adding another hidden model call to the AIOps flow before the current trace contract is mature.
|
规则验证稳定、便宜、容易解释,也不会在当前链路里额外引入一次隐藏模型调用。
|
||||||
|
|
||||||
## Verdicts
|
## 3. 判定结果
|
||||||
|
|
||||||
The evaluator emits:
|
当前评估器输出:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
PASS
|
PASS
|
||||||
@@ -40,14 +39,36 @@ WARN
|
|||||||
FAIL
|
FAIL
|
||||||
```
|
```
|
||||||
|
|
||||||
`FAIL` is reserved for critical issues such as a missing or too-short report. Missing payload focus terms or missing evidence tools currently produce `WARN`, because valid reports may use slightly different wording or evidence may be unavailable in a mock/demo environment.
|
含义:
|
||||||
|
|
||||||
## Interview Answer
|
- `PASS`:核心检查通过。
|
||||||
|
- `WARN`:报告存在,但可能缺少 payload 关键词或证据工具。
|
||||||
|
- `FAIL`:缺少最终报告、报告过短等关键问题。
|
||||||
|
|
||||||
If asked why AIOps has a verifier now:
|
缺少 payload 关键词或证据工具先给 `WARN`,因为 demo/mock 环境下证据可能不可用,且报告措辞可能与 payload 字段不完全一致。
|
||||||
|
|
||||||
> Chat already has an LLM verifier because the user questions are open-ended. For AIOps, I started with a lighter rule-based verifier because the first quality checks are very concrete: payload focus, evidence coverage, and report completeness. The evaluation is persisted into `self_evaluation`, so the trace can show not only what the Agent did, but also whether the output passed basic quality gates.
|
## 4. 面试回答
|
||||||
|
|
||||||
If asked why not use the Chat verifier directly:
|
如果被问:为什么 AIOps 也需要验证器?
|
||||||
|
|
||||||
|
```text
|
||||||
|
Chat 已经有 LLM Verifier,因为用户问题开放度高。
|
||||||
|
AIOps 的第一阶段质量风险更明确:报告是否聚焦输入告警、是否使用证据工具、报告是否完整。
|
||||||
|
所以我先做了轻量规则验证器,把结果写入 self_evaluation,让 Trace 不只展示 Agent 做了什么,也展示输出是否通过基础质量门。
|
||||||
|
```
|
||||||
|
|
||||||
|
如果被问:为什么不直接复用 Chat Verifier?
|
||||||
|
|
||||||
|
```text
|
||||||
|
AIOps 验证语义和 Chat 不一样。
|
||||||
|
它要检查 alert scope、payload focus、证据工具覆盖,以及是否过度展开无关 active alerts。
|
||||||
|
直接复用 Chat Verifier 会混淆这些语义。
|
||||||
|
规则评估先提供稳定质量门,后续 AIOps LLM Verifier 可以基于同一套 trace contract 扩展。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 5. 后续增强
|
||||||
|
|
||||||
|
- 引入 AIOps LLM Verifier,逐条校验根因和建议是否有 evidence refs。
|
||||||
|
- 把 rule evaluation 的 checks 在 Trace API 中结构化展示。
|
||||||
|
- 将 payload scope violation 沉淀为 bad case。
|
||||||
|
|
||||||
> AIOps verification is different from Chat verification. It needs to check alert scope, evidence tool coverage, and whether unrelated active alerts were over-expanded. Reusing the Chat verifier directly would blur those semantics. The rule-based evaluator gives us a stable first quality gate; a later AIOps LLM verifier can build on the same trace contract.
|
|
||||||
|
|||||||
@@ -1,62 +1,75 @@
|
|||||||
# AIOps Query Augmentation
|
# AIOps 查询增强说明
|
||||||
|
|
||||||
## What Changed
|
## 1. 改动是什么
|
||||||
|
|
||||||
Payload-targeted AIOps prompts now include a deterministic recommended knowledge query.
|
AIOps 在 `PAYLOAD_TARGETED` 模式下,会从告警 payload 中稳定生成一条推荐知识库检索 query。
|
||||||
|
|
||||||
The query is built from the non-blank payload fields:
|
参与拼接的非空字段:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
alertName service severity description timeRange userRequest
|
alertName service severity description timeRange userRequest
|
||||||
```
|
```
|
||||||
|
|
||||||
Example:
|
示例:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
HighCPUUsage payment-service P1 CPU usage is above 80% last_15m
|
HighCPUUsage payment-service P1 CPU 使用率超过 80% last_15m
|
||||||
```
|
```
|
||||||
|
|
||||||
## Why This Matters
|
最终会进入 Prompt:
|
||||||
|
|
||||||
AIOps payload fields contain high-value retrieval terms:
|
|
||||||
|
|
||||||
- alert name
|
|
||||||
- service name
|
|
||||||
- severity
|
|
||||||
- symptom description
|
|
||||||
- time range
|
|
||||||
- operator request
|
|
||||||
|
|
||||||
Before this change, the Agent still had to invent its own `lookup_knowledge` query from the full prompt. That can work, but it may omit important terms such as the service name or alert name.
|
|
||||||
|
|
||||||
The new prompt makes the retrieval seed explicit:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Recommended lookup_knowledge query: ...
|
Recommended lookup_knowledge query: ...
|
||||||
```
|
```
|
||||||
|
|
||||||
## Design Choice
|
## 2. 为什么重要
|
||||||
|
|
||||||
This is prompt-level query augmentation, not hidden retrieval.
|
AIOps payload 里包含高价值检索词:
|
||||||
|
|
||||||
I intentionally did not call `lookup_knowledge` automatically before the Agent runs. The project values traceability: tool calls should appear as Agent actions, with their inputs and outputs recorded in `tool_invocation`.
|
- 告警名称。
|
||||||
|
- 服务名。
|
||||||
|
- 严重等级。
|
||||||
|
- 症状描述。
|
||||||
|
- 时间范围。
|
||||||
|
- 用户补充请求。
|
||||||
|
|
||||||
So the design is:
|
如果完全让 Agent 从长 Prompt 里自己组织检索 query,可能遗漏服务名或告警名。推荐 query 让检索种子更稳定。
|
||||||
|
|
||||||
|
## 3. 设计取舍
|
||||||
|
|
||||||
|
这是 Prompt 层 query augmentation,不是隐藏检索。
|
||||||
|
|
||||||
|
我没有在 Agent 运行前自动调用 `lookup_knowledge`,原因是项目强调可追踪性:工具调用应该由 Agent 显式发起,并记录到 `tool_invocation`。
|
||||||
|
|
||||||
|
当前设计:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
AIOps payload
|
AIOps payload
|
||||||
-> deterministic recommended retrieval query
|
-> deterministic recommended retrieval query
|
||||||
-> Agent prompt
|
-> Agent prompt
|
||||||
-> Agent may call lookup_knowledge explicitly
|
-> Agent 显式调用 lookup_knowledge
|
||||||
-> tool_invocation records the real retrieval action
|
-> tool_invocation 记录真实检索行为
|
||||||
```
|
```
|
||||||
|
|
||||||
## Interview Answer
|
## 4. 面试回答
|
||||||
|
|
||||||
If asked how AIOps payload improves RAG retrieval:
|
如果被问:AIOps payload 怎么提升 RAG 检索?
|
||||||
|
|
||||||
> I do not replace the user query with a broad domain. I extract the high-signal alert terms from the payload, such as alertName, service, severity, symptom, and time range, and put them into a compact recommended lookup query. The Agent still calls `lookup_knowledge` explicitly, so the trace remains auditable, but the retrieval query is less dependent on model improvisation.
|
```text
|
||||||
|
我没有把告警 payload 粗暴替换成一个宽泛领域,而是提取 alertName、service、severity、description、timeRange 等高信号字段,拼成推荐的 lookup_knowledge query。
|
||||||
|
Agent 仍然显式调用工具,所以 trace 仍然能看到真实检索行为,但 query 不再完全依赖模型临场发挥。
|
||||||
|
```
|
||||||
|
|
||||||
If asked why not auto-call retrieval:
|
如果被问:为什么不自动检索?
|
||||||
|
|
||||||
|
```text
|
||||||
|
自动检索会在 Agent 真正决策前制造一份隐藏证据。
|
||||||
|
这个项目的重点是可观测 Agent 执行,所以我选择 Prompt 层增强:给 Agent 一个更好的 query seed,但不改变工具调用必须显式可追踪的契约。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 5. 后续增强
|
||||||
|
|
||||||
|
- 将 recommended query 写入 trace 的结构化字段,便于对比 Agent 实际 query。
|
||||||
|
- 对 payload 字段加权,例如 alertName/service 权重大于 timeRange。
|
||||||
|
- 后续接入 Query Transformer 时,保留原始 query、推荐 query、改写 query 三者的可追踪关系。
|
||||||
|
|
||||||
> Auto-calling retrieval would create hidden evidence before the Agent actually decides to use a tool. For this project, explicit tool invocation is more important because the interview story is about observable Agent execution. Prompt-level augmentation gives the Agent a better query seed without changing the trace contract.
|
|
||||||
|
|||||||
+70
-120
@@ -1,147 +1,97 @@
|
|||||||
# Architecture
|
# 面试版系统架构
|
||||||
|
|
||||||
## 系统分层
|
## 1. 系统分层
|
||||||
|
|
||||||
```text
|
```mermaid
|
||||||
API Layer
|
flowchart TB
|
||||||
-> ChatController / DiagnosisTraceController
|
API["API 层\nChatController / DiagnosisTraceController / SearchController"] --> Service["应用服务层\nChatService / AiOpsService / DiagnosisTraceService"]
|
||||||
|
Service --> Agent["Agent 编排层\nPlanner / Executor / Verifier / Supervisor"]
|
||||||
Agent Orchestration
|
Agent --> Tools["工具层\nlookup_knowledge / query_logs / query_metrics / Prometheus"]
|
||||||
-> ChatService / AiOpsService
|
Tools --> RAG["RAG 检索\nL0 hint + VectorSearchService"]
|
||||||
|
RAG --> VectorStore["Spring AI VectorStore"]
|
||||||
Tools
|
RAG --> SDK["Milvus SDK fallback"]
|
||||||
-> lookupKnowledgeTool / queryLogs / queryMetrics / queryPrometheusAlerts
|
Agent --> Trace["Trace 持久化"]
|
||||||
|
Tools --> Trace
|
||||||
Persistence
|
Trace --> Session["diagnosis_session"]
|
||||||
-> diagnosis_session / agent_step / tool_invocation
|
Trace --> Step["agent_step"]
|
||||||
|
Trace --> Invocation["tool_invocation"]
|
||||||
Trace
|
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
|
||||||
-> GET /api/diagnosis/{sessionId}/trace
|
Step --> TraceAPI
|
||||||
|
Invocation --> TraceAPI
|
||||||
```
|
```
|
||||||
|
|
||||||
## Chat 链路
|
## 2. Chat 链路
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart TD
|
flowchart TD
|
||||||
User[User Question] --> ChatAPI[POST /api/chat]
|
User["用户问题"] --> ChatAPI["POST /api/chat"]
|
||||||
ChatAPI --> Strategy[ChatService.executeChatWithStrategy]
|
ChatAPI --> Strategy["ChatService.executeChatWithStrategy"]
|
||||||
Strategy --> Complexity{QuestionComplexity}
|
Strategy --> Complexity{"复杂问题?"}
|
||||||
Complexity -->|simple| Single[ReactAgent]
|
Complexity -->|否| Single["单 ReactAgent 快速回答"]
|
||||||
Complexity -->|complex| Planner[Planner Agent]
|
Complexity -->|是| Planner["chat_planner"]
|
||||||
Planner --> Executor[Executor Agent]
|
Planner --> Executor["chat_executor"]
|
||||||
Executor --> Tools[Evidence Tools]
|
Executor --> Tools["证据工具"]
|
||||||
Tools --> Executor
|
Tools --> Executor
|
||||||
Executor --> Verifier[Verifier Agent]
|
Executor --> Verifier["chat_verifier"]
|
||||||
Verifier --> Answer[Final Answer]
|
Verifier --> Decision{"PASS / LOW_CONFID / REJECT"}
|
||||||
Answer --> Session[diagnosis_session]
|
Decision --> Answer["最终答复"]
|
||||||
Planner --> Steps[agent_step]
|
Planner --> Step["agent_step"]
|
||||||
Executor --> Steps
|
Executor --> Step
|
||||||
Verifier --> Steps
|
Verifier --> Step
|
||||||
Tools --> Invocations[tool_invocation]
|
Tools --> Invocation["tool_invocation"]
|
||||||
Session --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
Answer --> Session["diagnosis_session"]
|
||||||
Steps --> Trace
|
|
||||||
Invocations --> Trace
|
|
||||||
```
|
```
|
||||||
|
|
||||||
关键代码:
|
讲解重点:
|
||||||
|
|
||||||
- `ChatController.chat(...)`
|
- Planner 拆解问题和排查方向。
|
||||||
- `ChatService.executeChatWithStrategy(...)`
|
- Executor 必须通过工具收集证据。
|
||||||
- `ChatService.executeChatComplex(...)`
|
- Verifier 只基于 `tool_trace_summary` 校验答案,不做新检索。
|
||||||
- `AgentLoggingHook`
|
- Trace API 能回放模型步骤和工具证据。
|
||||||
- `ToolInvocationRecorder`
|
|
||||||
- `DiagnosisTraceService.getTrace(...)`
|
|
||||||
|
|
||||||
## AIOps 链路
|
## 3. AIOps 链路
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart TD
|
flowchart TD
|
||||||
Alert[Alert Payload or Empty Request] --> AiOpsAPI[POST /api/ai_ops]
|
Alert["告警 payload 或空请求"] --> API["POST /api/ai_ops"]
|
||||||
AiOpsAPI --> SessionEvent[SSE session event]
|
API --> AiOps["AiOpsService"]
|
||||||
AiOpsAPI --> AiOpsService[AiOpsService.executeAiOpsAnalysis]
|
AiOps --> Mode{"是否有 payload?"}
|
||||||
AiOpsService --> PromptMode{Payload?}
|
Mode -->|有| Targeted["PAYLOAD_TARGETED\n聚焦输入告警"]
|
||||||
PromptMode -->|yes| Targeted[PAYLOAD_TARGETED]
|
Mode -->|无| Discovery["AUTO_DISCOVERY\n先发现活跃告警"]
|
||||||
PromptMode -->|no| Discovery[AUTO_DISCOVERY]
|
Targeted --> Supervisor["ai_ops_supervisor"]
|
||||||
Targeted --> Supervisor[ai_ops_supervisor]
|
|
||||||
Discovery --> Supervisor
|
Discovery --> Supervisor
|
||||||
Supervisor --> Planner[planner_agent]
|
Supervisor --> Planner["planner_agent"]
|
||||||
Supervisor --> Executor[executor_agent]
|
Supervisor --> Executor["executor_agent"]
|
||||||
Planner --> Tools[Prometheus / Logs / Knowledge]
|
Planner --> Tools["Prometheus / 日志 / 知识库"]
|
||||||
Executor --> Tools
|
Executor --> Tools
|
||||||
Tools --> Report[Alert Report]
|
Tools --> Report["告警分析报告"]
|
||||||
Report --> Persist[diagnosis_session.answer]
|
Report --> Eval["AiOpsRuleEvaluationService"]
|
||||||
Planner --> Steps[agent_step]
|
Eval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
|
||||||
Executor --> Steps
|
|
||||||
Tools --> Invocations[tool_invocation]
|
|
||||||
Persist --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
|
||||||
Steps --> Trace
|
|
||||||
Invocations --> Trace
|
|
||||||
```
|
```
|
||||||
|
|
||||||
关键代码:
|
讲解重点:
|
||||||
|
|
||||||
- `ChatController.aiOps(...)`
|
- AIOps 有明确产品边界:有 payload 时必须聚焦该告警。
|
||||||
- `AIOpsRequest`
|
- payload 字段会生成 recommended `lookup_knowledge` query。
|
||||||
- `AiOpsService.resolveSessionId(...)`
|
- 当前 AIOps 先用规则评估做质量门,后续再扩展 LLM Verifier。
|
||||||
- `AiOpsService.buildTaskPrompt(...)`
|
|
||||||
- `AiOpsService.hasAlertPayload(...)`
|
|
||||||
- `AiOpsService.persistFinalReport(...)`
|
|
||||||
|
|
||||||
## Trace 数据模型
|
## 4. Trace 数据模型
|
||||||
|
|
||||||
### `diagnosis_session`
|
| 表 | 作用 |
|
||||||
|
|---|---|
|
||||||
|
| `diagnosis_session` | 一次诊断的主记录:问题、状态、答案、自评估、反馈 |
|
||||||
|
| `agent_step` | Agent 模型调用记录:输入、输出、耗时、token、是否有工具调用 |
|
||||||
|
| `tool_invocation` | 工具调用事实:工具名、入参、输出预览、检索层、相关性、成功状态 |
|
||||||
|
|
||||||
记录一次诊断会话的主信息:
|
## 5. 为什么 Trace 是核心
|
||||||
|
|
||||||
- `session_id`
|
故障诊断系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
|
||||||
- `query`
|
|
||||||
- `status`
|
|
||||||
- `agent_flow`
|
|
||||||
- `total_duration_ms`
|
|
||||||
- `total_token_count`
|
|
||||||
- `step_count`
|
|
||||||
- `tool_call_count`
|
|
||||||
- `answer`
|
|
||||||
- `self_evaluation`
|
|
||||||
- `feedback`
|
|
||||||
|
|
||||||
### `agent_step`
|
- 模型为什么这么答。
|
||||||
|
- 调了哪些工具。
|
||||||
|
- 工具返回了什么证据。
|
||||||
|
- Verifier 如何判断答案可信度。
|
||||||
|
- 用户反馈如何回写到同一个 session。
|
||||||
|
|
||||||
记录 Agent 模型调用过程:
|
这就是它区别于普通 Chatbot 的地方。
|
||||||
|
|
||||||
- `session_id`
|
|
||||||
- `step_index`
|
|
||||||
- `agent_name`
|
|
||||||
- `model_input`
|
|
||||||
- `model_output`
|
|
||||||
- `thought`
|
|
||||||
- `has_tool_call`
|
|
||||||
- `duration_ms`
|
|
||||||
- `token_count`
|
|
||||||
|
|
||||||
### `tool_invocation`
|
|
||||||
|
|
||||||
记录真实工具调用:
|
|
||||||
|
|
||||||
- `session_id`
|
|
||||||
- `tool_name`
|
|
||||||
- `input_params`
|
|
||||||
- `output_preview`
|
|
||||||
- `output_length`
|
|
||||||
- `retrieval_layer`
|
|
||||||
- `relevance_level`
|
|
||||||
- `duration_ms`
|
|
||||||
- `success`
|
|
||||||
- `error_message`
|
|
||||||
|
|
||||||
## 为什么 trace 是核心
|
|
||||||
|
|
||||||
Agent 系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
|
|
||||||
|
|
||||||
- 模型为什么这么答
|
|
||||||
- 调了哪些工具
|
|
||||||
- 工具返回了什么证据
|
|
||||||
- Verifier 如何判断答案可信度
|
|
||||||
- 用户反馈如何回写到同一个 session
|
|
||||||
|
|
||||||
这就是项目区别于普通 Chatbot 的地方。
|
|
||||||
|
|||||||
+45
-32
@@ -1,18 +1,20 @@
|
|||||||
# Interview Demo Script
|
# 面试演示脚本
|
||||||
|
|
||||||
## 30 秒开场
|
## 1. 30 秒开场
|
||||||
|
|
||||||
这是一个 Agent Engineering 项目,场景是企业故障诊断。它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。项目重点不是单次回答,而是把多 Agent 执行、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 trace。
|
```text
|
||||||
|
这是一个 Agent 工程项目,场景是企业故障诊断。
|
||||||
|
它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。
|
||||||
|
项目重点不是单次模型回答,而是把 Agent 编排、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 Trace。
|
||||||
|
```
|
||||||
|
|
||||||
## Demo 准备
|
## 2. 启动服务
|
||||||
|
|
||||||
启动服务:
|
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||||
```
|
```
|
||||||
|
|
||||||
确认服务地址:
|
服务地址:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
http://localhost:9900
|
http://localhost:9900
|
||||||
@@ -24,11 +26,9 @@ http://localhost:9900
|
|||||||
- CLS 日志使用 mock 数据。
|
- CLS 日志使用 mock 数据。
|
||||||
- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。
|
- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。
|
||||||
|
|
||||||
## Demo 1: Chat 诊断
|
## 3. Demo 1:Chat 诊断
|
||||||
|
|
||||||
目标:展示普通用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 trace。
|
目标:展示用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 Trace。
|
||||||
|
|
||||||
请求:
|
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
$sessionId = "interview-chat-payment-timeout-001"
|
$sessionId = "interview-chat-payment-timeout-001"
|
||||||
@@ -46,13 +46,13 @@ Invoke-RestMethod `
|
|||||||
|
|
||||||
讲解点:
|
讲解点:
|
||||||
|
|
||||||
- `ChatController` 把请求交给 `ChatService.executeChatWithStrategy(...)`。
|
- `ChatService` 会根据问题复杂度选择轻量回答或复杂 Agent 流程。
|
||||||
- 简单问题走单 ReactAgent,复杂问题走 `Planner -> Executor -> Verifier`。
|
- 复杂问题走 `Planner -> Executor -> Verifier`。
|
||||||
- Executor 可以调用知识库、日志、指标等工具。
|
- Executor 调用知识库、日志、指标等证据工具。
|
||||||
- Verifier 会基于工具证据生成 groundedness 评估。
|
- Verifier 基于 `tool_trace_summary` 生成 groundedness 评估。
|
||||||
- 最终会写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
|
- 最终写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
|
||||||
|
|
||||||
查询 trace:
|
查询 Trace:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
@@ -67,11 +67,9 @@ Invoke-RestMethod `
|
|||||||
- `data.toolInvocations` 中能看到证据工具
|
- `data.toolInvocations` 中能看到证据工具
|
||||||
- `data.session.selfEvaluation` 中有 verifier 结果
|
- `data.session.selfEvaluation` 中有 verifier 结果
|
||||||
|
|
||||||
## Demo 2: AIOps 告警诊断
|
## 4. Demo 2:AIOps 告警诊断
|
||||||
|
|
||||||
目标:展示告警 payload 如何触发 AIOps 入口,并且报告只聚焦目标告警。
|
目标:展示告警 payload 如何触发 AIOps,并且报告聚焦目标告警。
|
||||||
|
|
||||||
请求:
|
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
$aiopsSessionId = "interview-aiops-payment-cpu-001"
|
$aiopsSessionId = "interview-aiops-payment-cpu-001"
|
||||||
@@ -99,9 +97,9 @@ Invoke-WebRequest `
|
|||||||
- `AiOpsService` 根据 payload 判断模式:
|
- `AiOpsService` 根据 payload 判断模式:
|
||||||
- `PAYLOAD_TARGETED`:聚焦传入告警。
|
- `PAYLOAD_TARGETED`:聚焦传入告警。
|
||||||
- `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。
|
- `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。
|
||||||
- AIOps 暂时不加 Verifier,先保证告警入口、证据工具和 trace 可用。
|
- AIOps 当前用 rule evaluation 检查报告完整性、payload 聚焦和证据工具覆盖。
|
||||||
|
|
||||||
查询 trace:
|
查询 Trace:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
@@ -114,19 +112,34 @@ Invoke-RestMethod `
|
|||||||
- `data.session.agentFlow = AI_OPS`
|
- `data.session.agentFlow = AI_OPS`
|
||||||
- `data.session.answer` 有最终告警报告
|
- `data.session.answer` 有最终告警报告
|
||||||
- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge`
|
- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge`
|
||||||
- 报告有 `HighCPUUsage/payment-service` 的完整根因分析
|
- 报告主线聚焦 `HighCPUUsage/payment-service`
|
||||||
- 其他 active alerts 只作为相关风险出现,不展开成独立根因章节
|
|
||||||
|
|
||||||
## MySQL 验证
|
## 5. Demo 3:反馈闭环
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session ORDER BY id DESC LIMIT 5"
|
$feedback = @{
|
||||||
|
sessionId = $sessionId
|
||||||
|
feedback = "useful"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "http://localhost:9900/api/feedback" `
|
||||||
|
-ContentType "application/json" `
|
||||||
|
-Body $feedback
|
||||||
```
|
```
|
||||||
|
|
||||||
```powershell
|
讲解点:
|
||||||
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name"
|
|
||||||
|
- feedback 写回同一个 `diagnosis_session`。
|
||||||
|
- `useful` 会沉淀 `case_library`。
|
||||||
|
- `not_useful` 不改变 `status`,只作为质量信号。
|
||||||
|
|
||||||
|
## 6. 收尾总结
|
||||||
|
|
||||||
|
```text
|
||||||
|
这个 Demo 展示的是完整 Agent 闭环:
|
||||||
|
用户问题或告警 -> Agent 编排 -> 工具证据 -> 自评估 -> Trace 回放 -> 用户反馈 -> 案例沉淀。
|
||||||
|
我关注的不是一次回答,而是这个回答能否被审计、验证和持续改进。
|
||||||
```
|
```
|
||||||
|
|
||||||
## 收尾总结
|
|
||||||
|
|
||||||
这套 Demo 展示的是一个完整 Agent 系统,而不是一次模型问答:入口有明确场景边界,Agent 负责规划和执行,工具提供证据,Verifier 提供质量门,trace API 提供审计和复盘能力。AIOps 入口进一步证明它可以从用户问答扩展到事件驱动诊断。
|
|
||||||
|
|||||||
@@ -1,99 +1,107 @@
|
|||||||
# Design Tradeoffs
|
# 关键设计取舍
|
||||||
|
|
||||||
## 1. 为什么要做 trace,而不是只返回答案
|
## 1. 为什么先做 Trace,而不是只返回答案
|
||||||
|
|
||||||
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答三件事:
|
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答:
|
||||||
|
|
||||||
- 结论是什么
|
- 结论是什么。
|
||||||
- 证据来自哪里
|
- 证据来自哪里。
|
||||||
- 哪些步骤由哪个 Agent 完成
|
- 哪些步骤由哪个 Agent 完成。
|
||||||
|
- 如果答案不可靠,系统怎么降级。
|
||||||
|
|
||||||
因此项目把一次会话拆成:
|
因此项目把一次会话拆成:
|
||||||
|
|
||||||
- `diagnosis_session`:会话级摘要、最终答案、质量评估、反馈。
|
- `diagnosis_session`:会话摘要、最终答案、自评估、用户反馈。
|
||||||
- `agent_step`:Agent 模型输入输出、耗时、token 和工具调用标记。
|
- `agent_step`:模型输入输出、耗时、token 和工具调用标记。
|
||||||
- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。
|
- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。
|
||||||
|
|
||||||
这个设计牺牲了一些实现复杂度,但换来了可回放、可调试、可演示。
|
代价是实现复杂度上升,收益是可回放、可调试、可演示。
|
||||||
|
|
||||||
## 2. 为什么 Chat 有 Verifier,AIOps 暂时没有
|
## 2. 为什么 RAG 不直接隐藏在 Advisor 里
|
||||||
|
|
||||||
Chat 入口的问题更开放,用户可能要求复杂推理或跨领域结论,所以 Verifier 是必要的质量门。当前 Chat 链路通过 `Planner -> Executor -> Verifier` 固定流程,把 groundedness 和 facts checked 写入 `self_evaluation`。
|
Spring AI Advisor 可以让 RAG 更隐式,但本项目的核心是 Agent 证据链。`lookup_knowledge` 必须作为显式工具调用出现,这样 Trace 里才能看到:
|
||||||
|
|
||||||
AIOps 当前阶段先不加 Verifier,原因是:
|
- Agent 什么时候决定检索。
|
||||||
|
- 用了什么 query。
|
||||||
|
- 命中了哪些文档。
|
||||||
|
- 相关性等级是什么。
|
||||||
|
- 证据如何支撑最终答案。
|
||||||
|
|
||||||
- AIOps 刚完成从“自动跑告警”到“可追踪告警入口”的改造。
|
所以当前设计是:
|
||||||
- 先要确认告警 payload、工具证据、最终报告和 trace 能闭环。
|
|
||||||
- AIOps Verifier 的规则不同于 Chat Verifier,需要检查告警 scope、证据覆盖和处置建议,不宜直接复用。
|
|
||||||
|
|
||||||
后续可以做 lightweight AIOps Verifier,检查报告是否聚焦 payload、是否引用工具证据、是否误展开无关告警。
|
|
||||||
|
|
||||||
## 3. 为什么 AIOps payload scope 先用 prompt 控制
|
|
||||||
|
|
||||||
运行验证发现:传入 `HighCPUUsage/payment-service` 后,Agent 仍可能把 mock Prometheus 返回的所有 active alerts 都展开分析。这个问题的本质是任务边界不清晰。
|
|
||||||
|
|
||||||
当前选择 prompt-level scope control:
|
|
||||||
|
|
||||||
- 有 payload:`PAYLOAD_TARGETED`,最终报告围绕传入告警。
|
|
||||||
- 无 payload:`AUTO_DISCOVERY`,先调用 `queryPrometheusAlerts` 自动发现告警。
|
|
||||||
|
|
||||||
没有先做 Java 侧过滤,是因为:
|
|
||||||
|
|
||||||
- 过滤工具结果会降低 Agent 发现关联风险的能力。
|
|
||||||
- 目前需要的是报告主线聚焦,而不是完全屏蔽上下文。
|
|
||||||
- Prompt 改动小,风险低,能保留 Agent 灵活性。
|
|
||||||
|
|
||||||
已验证结果:主报告有 `HighCPUUsage/payment-service` 的完整根因分析,`HighMemoryUsage` 和 `SlowResponse` 只作为相关风险出现。
|
|
||||||
|
|
||||||
## 4. 为什么用 `tool_invocation` 统计真实工具调用次数
|
|
||||||
|
|
||||||
早期可以通过 `agent_step.hasToolCall` 粗略判断是否调用工具,但它统计的是“哪些模型步骤包含工具调用”,不是“真实调用了几次工具”。
|
|
||||||
|
|
||||||
现在 `tool_call_count` 来自:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
ToolInvocationRepository.countBySessionId(sessionId)
|
Executor -> lookup_knowledge -> VectorSearchService -> VectorStore / SDK fallback
|
||||||
```
|
```
|
||||||
|
|
||||||
这样更符合 trace 语义:
|
这牺牲了一点框架自动化,但保留了可审计性。
|
||||||
|
|
||||||
- 一个 step 可能调用多个工具。
|
## 3. 为什么 L0 只做 hint,不直接返回
|
||||||
- 工具可能来自不同来源:知识库、日志、指标、Prometheus。
|
|
||||||
- 面试时可以把 `tool_call_count` 和 trace 中返回的工具明细对上。
|
|
||||||
|
|
||||||
## 5. 为什么保留 mock Prometheus 和 mock CLS
|
旧版 L0 关键词唯一命中时可能直接跳过 L1。这个策略速度快,但风险是:关键词子串命中不等于最终语义相关。
|
||||||
|
|
||||||
面试 Demo 最怕不稳定。真实 Prometheus、日志平台和线上故障都有不可控因素,所以 MVP profile 保留 mock 工具:
|
当前改成:
|
||||||
|
|
||||||
- `prometheus.mock-enabled=true`
|
```text
|
||||||
- `cls.mock-enabled=true`
|
L0 = domain/entity hint
|
||||||
|
L1 = semantic retrieval
|
||||||
|
postprocess = evidence shaping + trace
|
||||||
|
```
|
||||||
|
|
||||||
这样可以稳定复现:
|
L0 仍然有价值:错误码、服务名、告警名、指标名都很适合做精确 hint。但最终证据仍需要 L1 和后处理支撑。
|
||||||
|
|
||||||
- `HighCPUUsage/payment-service`
|
## 4. 为什么保留 Milvus SDK fallback
|
||||||
- `HighMemoryUsage/order-service`
|
|
||||||
- `SlowResponse/user-service`
|
|
||||||
- system-metrics、application-logs、database-slow-query 等日志证据
|
|
||||||
|
|
||||||
这不是逃避真实集成,而是把“Agent 编排和证据追踪”作为面试演示的主目标。
|
Spring AI VectorStore 是当前读路径主方向,但 SDK fallback 没有删除,原因有三点:
|
||||||
|
|
||||||
## 6. 为什么把面试材料单独放 `interview/`
|
- 迁移安全:旧 SDK 路径已经被验证过。
|
||||||
|
- 运行韧性:VectorStore 配置、schema、collection 出问题时可以回退。
|
||||||
|
- 面试稳定:检索抽象迁移不应该破坏主 demo。
|
||||||
|
|
||||||
`mvp/` 是持续迭代现场,包含过程文档、验收记录和 runbook。面试材料的目标不同,它应该是可讲、可演示、可评估的展示层。
|
这不是“没有迁完”,而是分阶段迁移:先稳定读路径,再决定是否迁移写入和索引。
|
||||||
|
|
||||||
因此:
|
## 5. 为什么 Chat 有 Verifier,AIOps 先用规则评估
|
||||||
|
|
||||||
- `mvp/` 保留真实演进材料。
|
Chat 问题更开放,容易出现跨领域推理,所以需要 LLM Verifier 做 groundedness 校验。
|
||||||
- `devflow/` 保留决策沉淀。
|
|
||||||
- `interview/` 只组织面试叙事和演示脚本。
|
|
||||||
|
|
||||||
这样后续继续做 AIOps Verifier、UI、更多工具集成时,不会污染面试讲稿。
|
AIOps 当前优先解决更具体的问题:
|
||||||
|
|
||||||
## 7. 可以主动承认的限制
|
- 最终报告是否存在。
|
||||||
|
- payload 模式是否聚焦输入告警。
|
||||||
|
- 是否使用了证据工具。
|
||||||
|
- 是否把无关活跃告警展开成主根因。
|
||||||
|
|
||||||
- AIOps 还没有 Verifier。
|
这些用规则就能稳定检查。后续可以在同一个 `self_evaluation` 容器下增加 AIOps LLM Verifier。
|
||||||
- Prompt-level scope control 不能做到强约束,只能通过 trace 和测试观察遵循情况。
|
|
||||||
- 当前 mock 数据适合 demo,不代表生产接入已经完成。
|
## 6. 为什么 AIOps payload scope 先用 Prompt + Rule
|
||||||
- Hikari 连接池已经加了短生命周期和 keepalive,但真实生产还需要按数据库 wait_timeout 和连接数预算调优。
|
|
||||||
|
真实告警环境里可能同时有多个 active alerts。用户传入 `HighCPUUsage/payment-service` 时,Agent 如果把所有告警都展开分析,报告会跑偏。
|
||||||
|
|
||||||
|
当前选择:
|
||||||
|
|
||||||
|
- Prompt 中加入 `PAYLOAD_TARGETED`。
|
||||||
|
- 从 payload 生成 recommended `lookup_knowledge` query。
|
||||||
|
- 用 `AiOpsRuleEvaluationService` 检查报告是否聚焦输入告警。
|
||||||
|
|
||||||
|
没有先做硬过滤,是因为有些相关告警可以作为风险背景。目标不是屏蔽上下文,而是控制主诊断对象。
|
||||||
|
|
||||||
|
## 7. 为什么反馈不改 status
|
||||||
|
|
||||||
|
`status` 表示执行状态,`feedback` 表示用户评价。一个执行成功但用户觉得没用的诊断,应该是:
|
||||||
|
|
||||||
|
```text
|
||||||
|
status = SUCCESS
|
||||||
|
feedback = not_useful
|
||||||
|
```
|
||||||
|
|
||||||
|
这样才能区分系统异常和质量问题。`useful` 反馈会沉淀 `case_library`,`not_useful` 作为 bad case 信号保留。
|
||||||
|
|
||||||
|
## 8. 可以主动承认的限制
|
||||||
|
|
||||||
|
- AIOps 还没有完整 LLM Verifier。
|
||||||
|
- RAG 还没有 hybrid search、rerank、邻居 chunk 扩展。
|
||||||
|
- `case_library` 的 rootCause/solution 仍需要结构化抽取。
|
||||||
|
- `tool_invocation.step_id` 关联还可以更严格。
|
||||||
|
- `mvp-demo` profile 使用 mock 日志和指标,主要服务稳定面试演示。
|
||||||
|
|
||||||
|
主动讲清这些限制,能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
|
||||||
|
|
||||||
主动讲清这些限制,反而能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
|
|
||||||
|
|||||||
@@ -1,8 +1,8 @@
|
|||||||
# RAG Breadcrumb Embedding Acceptance
|
# RAG Breadcrumb Embedding 验收说明
|
||||||
|
|
||||||
## What Changed
|
## 1. 改动是什么
|
||||||
|
|
||||||
The indexing path now builds embedding text from chunk structure plus content:
|
索引路径现在构造 embedding 文本时,不只使用 chunk 内容,还会把结构上下文拼进去:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Title: {title}
|
Title: {title}
|
||||||
@@ -11,56 +11,81 @@ Content:
|
|||||||
{content}
|
{content}
|
||||||
```
|
```
|
||||||
|
|
||||||
The stored Milvus `content` field remains the original chunk content. This keeps display and evidence output clean while allowing the vector to carry section-level semantics.
|
Milvus 中存储的 `content` 字段仍然保留原始 chunk 内容。这样展示和证据输出保持干净,而向量本身携带章节语义。
|
||||||
|
|
||||||
## Why Reindex Is Required
|
## 2. 为什么必须重新索引
|
||||||
|
|
||||||
Embeddings are materialized at index time. Existing vectors were generated from the previous content-only text, so they cannot benefit from `title` and `breadcrumb` until the knowledge base is reindexed.
|
Embedding 是索引时物化的。已有向量是用旧的 content-only 文本生成的,所以只有代码变化并不会改变线上检索结果。
|
||||||
|
|
||||||
This is the key acceptance point:
|
验收关键点:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
code change alone != live retrieval changed
|
只改代码 != live retrieval 已变化
|
||||||
code change + reindex + live query report = accepted behavior
|
代码改动 + 重新索引 + live query report = 行为验收完成
|
||||||
```
|
```
|
||||||
|
|
||||||
## How To Validate
|
## 3. 如何验证
|
||||||
|
|
||||||
1. Start the Spring Boot application.
|
1. 启动 Spring Boot 应用。
|
||||||
2. Reindex the knowledge base through the existing indexing path.
|
2. 通过现有索引路径重新索引知识库。
|
||||||
3. Run:
|
3. 运行:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scripts/eval_rag_live_acceptance.py
|
python scripts/eval_rag_live_acceptance.py
|
||||||
```
|
```
|
||||||
|
|
||||||
The script writes:
|
脚本输出:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
eval/rag-retrieval/reports/live-post-reindex.json
|
eval/rag-retrieval/reports/live-post-reindex.json
|
||||||
eval/rag-retrieval/reports/live-post-reindex.md
|
eval/rag-retrieval/reports/live-post-reindex.md
|
||||||
```
|
```
|
||||||
|
|
||||||
The default cases cover:
|
默认覆盖:
|
||||||
|
|
||||||
- RAG chunk context questions where breadcrumb matters.
|
- breadcrumb 敏感的 RAG chunk context query。
|
||||||
- Diagnosis flow questions where section path matters.
|
- 需要章节路径的诊断流程问题。
|
||||||
- `ERR_TIMEOUT` exact error-code retrieval.
|
- `ERR_TIMEOUT` 精确错误码检索。
|
||||||
- MySQL connection pool troubleshooting.
|
- MySQL 连接池排障。
|
||||||
- AIOps payment-service latency alert retrieval.
|
- AIOps payment-service 延迟告警检索。
|
||||||
|
|
||||||
## What To Look For
|
## 4. 看什么结果
|
||||||
|
|
||||||
For breadcrumb-sensitive cases, inspect whether top candidates expose expected `title` and `breadcrumb` values in the report.
|
对 breadcrumb 敏感 case:
|
||||||
|
|
||||||
For core troubleshooting cases, check that result counts and top candidates remain stable. The goal is not to prove a full benchmark; it is to prove that reindexing did not obviously break important demo retrieval paths.
|
- top candidates 是否暴露预期 `title`。
|
||||||
|
- top candidates 是否暴露预期 `breadcrumb`。
|
||||||
|
- 命中内容是否能看出所属章节。
|
||||||
|
|
||||||
## Interview Answer
|
对核心排障 case:
|
||||||
|
|
||||||
If asked how I verified the breadcrumb embedding change:
|
- 结果数量是否稳定。
|
||||||
|
- top candidates 是否仍然命中核心文档。
|
||||||
|
- 没有因为拼接 title/breadcrumb 导致核心检索退化。
|
||||||
|
|
||||||
> I separated deterministic regression from live acceptance. The offline fixture baseline still runs without services. But because embedding changes only affect newly indexed vectors, I added a live post-reindex acceptance script. It calls the real `/api/search/similar` endpoint against representative breadcrumb-sensitive, troubleshooting, and AIOps queries, then writes JSON and Markdown reports. This lets me prove both that the code changed and that the live vector collection was refreshed.
|
## 5. 面试回答
|
||||||
|
|
||||||
If asked why the script does not reindex automatically:
|
如果被问:你怎么验证 breadcrumb 参与 embedding 后真的生效?
|
||||||
|
|
||||||
|
```text
|
||||||
|
我把 deterministic regression 和 live acceptance 分开。
|
||||||
|
离线 fixture baseline 不依赖服务,可以做稳定回归。
|
||||||
|
但 embedding 改动只会影响新生成的向量,所以我另外加了 live post-reindex acceptance 脚本。
|
||||||
|
脚本会调用真实 /api/search/similar,对 breadcrumb 敏感、排障和 AIOps query 生成 JSON/Markdown 报告。
|
||||||
|
这样能证明代码改了,也能证明 live vector collection 已经刷新。
|
||||||
|
```
|
||||||
|
|
||||||
|
如果被问:为什么脚本不自动 reindex?
|
||||||
|
|
||||||
|
```text
|
||||||
|
reindex 会修改向量库,而且依赖环境中的知识库数据。
|
||||||
|
我把 reindex 保持为显式动作,验收脚本只做读取验证。
|
||||||
|
这样如果检索没有改善,我能区分是代码问题、索引未刷新,还是运行时检索行为问题。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 6. 后续增强
|
||||||
|
|
||||||
|
- 将 live acceptance 结果加入面试 Demo 输出。
|
||||||
|
- 增加 breadcrumb hit rate 统计。
|
||||||
|
- 对同章节 chunk 做邻居扩展,进一步利用 breadcrumb。
|
||||||
|
|
||||||
> Reindexing mutates the vector store and depends on environment-specific data. I kept mutation explicit and made the script validation-only. That makes failures easier to diagnose: if retrieval does not improve, I can distinguish code changes, reindex state, and runtime retrieval behavior.
|
|
||||||
|
|||||||
+72
-154
@@ -1,47 +1,50 @@
|
|||||||
# RAG Refactor Story
|
# RAG 重构故事
|
||||||
|
|
||||||
## The Starting Point
|
## 1. 起点
|
||||||
|
|
||||||
The original RAG implementation was already usable for the MVP:
|
原始 RAG 实现已经能支撑 MVP:
|
||||||
|
|
||||||
- Documents could be uploaded, chunked, embedded, and written to Milvus/Zilliz.
|
- 文档可以上传、切片、向量化,并写入 Milvus/Zilliz。
|
||||||
- The Agent could call `lookup_knowledge` as an explicit tool.
|
- Agent 可以显式调用 `lookup_knowledge`。
|
||||||
- AIOps diagnosis could retrieve troubleshooting knowledge during an alert workflow.
|
- AIOps 诊断能在告警流程里检索排障知识。
|
||||||
- Tool invocations were persisted, so the retrieval step was visible in the execution trace.
|
- 工具调用会落到 `tool_invocation`,检索步骤可见。
|
||||||
|
|
||||||
But the design had several engineering problems:
|
但它有几个工程问题:
|
||||||
|
|
||||||
- Retrieval was too SDK-specific. The business code directly owned many Milvus search details.
|
- 检索实现过于依赖 Milvus SDK,业务代码承担了太多底层搜索细节。
|
||||||
- L0 and L1 responsibilities were blurry. L0 keyword matching could look like a final retrieval decision instead of a hint.
|
- L0 和 L1 职责不清,L0 关键词命中容易被当作最终召回决策。
|
||||||
- Chunk-level retrieval could lose section context when one section was split into multiple chunks.
|
- chunk 级检索容易丢失章节上下文。
|
||||||
- Metadata such as `breadcrumb` existed, but it was not fully used in retrieval, filtering, or context reconstruction.
|
- `breadcrumb` 存在 metadata 中,但没有充分参与 embedding、filter 和上下文重建。
|
||||||
- Retrieval quality was mostly checked by manual API calls and logs, not by repeatable cases.
|
- 检索质量主要靠手工接口和日志判断,缺少可重复的 golden cases。
|
||||||
|
|
||||||
So the refactor goal was not "replace everything with a framework." The goal was to move generic RAG infrastructure toward Spring AI while keeping the project-specific Agent evidence chain.
|
所以重构目标不是“全盘替换成框架”,而是:
|
||||||
|
|
||||||
## How I Broke The Problem Down
|
```text
|
||||||
|
通用 RAG 基础设施交给 Spring AI,
|
||||||
|
业务可观测链路保留在项目里。
|
||||||
|
```
|
||||||
|
|
||||||
I treated this as a staged migration, because RAG touches the Agent tool layer, AIOps diagnosis, vector retrieval, evidence packing, and database traces.
|
## 2. 我如何拆解问题
|
||||||
|
|
||||||
The first step was to establish a baseline. I added retrieval evaluation cases under `eval/rag-retrieval/` so future changes could be compared against known queries instead of judged only by intuition.
|
我把迁移拆成几个阶段,因为 RAG 同时影响 Agent 工具层、AIOps、向量检索、证据打包和 Trace。
|
||||||
|
|
||||||
Then I clarified the retrieval roles:
|
第一步是建立 baseline。`eval/rag-retrieval/` 中的 golden cases 用来对比后续改动,而不是只靠直觉判断检索有没有变好。
|
||||||
|
|
||||||
|
第二步是明确职责:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
L0 = domain/entity hint
|
L0 = domain/entity hint
|
||||||
L1 = semantic retrieval
|
L1 = semantic retrieval
|
||||||
postprocess = evidence shaping and trace-friendly output
|
postprocess = evidence shaping + trace-friendly output
|
||||||
```
|
```
|
||||||
|
|
||||||
That means L0 is still valuable, but it should not bypass semantic retrieval as the default path. It is better used to extract service names, alert names, error codes, domains, and metadata hints.
|
L0 仍然有价值,但不再默认绕过语义检索。它更适合提取服务名、告警名、错误码、领域和 metadata filter。
|
||||||
|
|
||||||
After that, I added evidence postprocessing. The Agent should not just receive raw chunks; it should receive structured evidence with source, title, breadcrumb, score, hit reason, and content. This makes the result easier to inspect and easier to explain in an interview.
|
第三步是增强 evidence 输出。Agent 不应该只拿到 raw chunk,而应该拿到带 source、title、breadcrumb、score、hit reason 的证据块。
|
||||||
|
|
||||||
Finally, I integrated Spring AI `VectorStore` as the main read path while preserving the original Milvus SDK implementation as fallback.
|
最后,我把 Spring AI `VectorStore` 接入为读取主路径,同时保留原 Milvus SDK 作为 fallback。
|
||||||
|
|
||||||
## Current Architecture
|
## 3. 当前架构
|
||||||
|
|
||||||
The current retrieval path is:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Agent / API
|
Agent / API
|
||||||
@@ -50,160 +53,75 @@ Agent / API
|
|||||||
-> VectorSearchService
|
-> VectorSearchService
|
||||||
-> Spring AI VectorStore
|
-> Spring AI VectorStore
|
||||||
-> Milvus SDK fallback
|
-> Milvus SDK fallback
|
||||||
-> evidence postprocess
|
-> relevance normalization
|
||||||
-> tool_invocation trace
|
-> tool_invocation trace
|
||||||
```
|
```
|
||||||
|
|
||||||
`VectorSearchService` is still the public retrieval facade. This is deliberate: the Agent tool layer does not need to know whether the underlying retrieval engine is SDK-based or Spring AI-based.
|
`VectorSearchService` 仍然是公共检索门面。Agent 工具层不需要知道底层是 SDK 还是 Spring AI。
|
||||||
|
|
||||||
The supported retrieval modes are:
|
支持三种模式:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
auto -> try Spring AI VectorStore, fallback to SDK
|
auto -> 优先 Spring AI VectorStore,失败后 fallback 到 SDK
|
||||||
spring-ai -> force Spring AI VectorStore
|
spring-ai -> 强制 Spring AI VectorStore
|
||||||
sdk -> force Milvus SDK
|
sdk -> 强制 Milvus SDK
|
||||||
```
|
```
|
||||||
|
|
||||||
This keeps the migration reversible and testable.
|
## 4. 关键取舍
|
||||||
|
|
||||||
## Key Tradeoffs
|
### 保留显式工具
|
||||||
|
|
||||||
### Keep The Explicit Tool
|
我没有把检索藏进 Spring AI Advisor。原因是这个项目强调 Agent 执行可见性:`lookup_knowledge` 的 query、命中文档、相关性和证据预览都要进入 Trace。
|
||||||
|
|
||||||
I did not hide retrieval inside a Spring AI Advisor.
|
### 保留 SDK fallback
|
||||||
|
|
||||||
For this project, `lookup_knowledge` is part of the Agent execution story. It records what query was used, which evidence was retrieved, how relevant it looked, and how it supported diagnosis. If retrieval is hidden inside an advisor, the answer may still work, but the audit trail becomes harder to show.
|
SDK fallback 不是废代码,而是迁移安全网。实际验证时,第一次 VectorStore 指向了错误 collection,`auto` 模式 fallback 到 SDK 后仍能返回结果。修正 collection 后,Spring AI 路径成为主路径。
|
||||||
|
|
||||||
### Keep SDK Fallback
|
### L0 降权
|
||||||
|
|
||||||
The SDK path is not dead code. It is a safety net during migration.
|
生产事故中经常有精确标识:错误码、告警名、服务名、指标名。L0 适合做 hint,但不应该做最终裁判。
|
||||||
|
|
||||||
This proved useful during live validation. The first VectorStore run pointed at the wrong collection name, but `auto` mode fell back to SDK and still returned results. After the collection was corrected to `biz`, the Spring AI path worked as the main path.
|
### 分数语义拆开
|
||||||
|
|
||||||
### Keep L0, But Reduce Its Authority
|
SDK 使用 L2 distance,Spring AI 暴露 similarity。混在一个字段里会让 relevance normalization 出错。
|
||||||
|
|
||||||
L0 is worth keeping because production incidents often contain exact identifiers:
|
当前拆成:
|
||||||
|
|
||||||
- error code
|
|
||||||
- alert name
|
|
||||||
- service name
|
|
||||||
- metric name
|
|
||||||
- domain tag
|
|
||||||
|
|
||||||
But L0 should not be the final judge of retrieval quality. Its role is now closer to domain hint, entity extraction, metadata filtering, and explainability signal.
|
|
||||||
|
|
||||||
### Split Score Semantics
|
|
||||||
|
|
||||||
The old SDK path used L2 distance. Spring AI exposes similarity. Treating those as the same number would quietly break relevance normalization.
|
|
||||||
|
|
||||||
So the result separates:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
score -> compatibility score used by existing logic
|
score -> 兼容旧逻辑的距离型分数
|
||||||
rawScore -> raw score from the retrieval implementation
|
rawScore -> 底层原始分数
|
||||||
scoreLabel -> semantic label for rawScore
|
scoreLabel -> rawScore 的语义
|
||||||
```
|
```
|
||||||
|
|
||||||
For SDK:
|
### 暂不迁移写入
|
||||||
|
|
||||||
|
写入和索引仍走 SDK。这是有意分阶段:先验证读路径,再评估 `VectorStore.add(...)` 是否适合现有 metadata 和 chunk 模型。
|
||||||
|
|
||||||
|
## 5. 验证方式
|
||||||
|
|
||||||
|
我用了三层验证:
|
||||||
|
|
||||||
|
- 单元测试:SDK mode、Spring AI mode、auto fallback、category filter、distance metadata mapping。
|
||||||
|
- Live API:`GET /api/search/similar?query=ERR_TIMEOUT&topK=3`。
|
||||||
|
- 代表性 query 对比:错误码、支付超时、MySQL 连接池、AIOps 告警式 query、抽象 RAG 设计问题。
|
||||||
|
|
||||||
|
核心排障和 AIOps query 在 SDK 与 VectorStore 下 top3 一致。差异主要集中在抽象设计类问题和 metadata taxonomy,这些被记录为后续质量工作。
|
||||||
|
|
||||||
|
## 6. 面试短版
|
||||||
|
|
||||||
```text
|
```text
|
||||||
score = L2 distance
|
这个 RAG 系统最初是基于 Milvus SDK 的自研 MVP。它能跑,但底层检索细节过多地散落在业务代码里,L0/L1 职责也不够清晰。
|
||||||
rawScore = L2 distance
|
我按阶段重构:先加 retrieval baseline,再把 L0 降级为 domain/entity hint,再增强 evidence postprocess,最后把读取主路径切到 Spring AI VectorStore,并保留 SDK fallback。
|
||||||
scoreLabel = l2_distance
|
我没有把 lookup_knowledge 替换成隐式 Advisor,因为这个项目的核心是可追踪 Agent:面试官可以看到什么时候检索、检索了什么、证据如何支撑诊断。
|
||||||
```
|
```
|
||||||
|
|
||||||
For VectorStore:
|
## 7. 可主动承认的不足
|
||||||
|
|
||||||
```text
|
- metadata taxonomy 还需要清理,例如 `database` 与 `infrastructure`。
|
||||||
score = Milvus metadata.distance when available
|
- 抽象设计问题可能需要 query rewrite 或更好的文档索引。
|
||||||
rawScore = Spring AI similarity
|
- 邻居 chunk / 同章节上下文扩展还不完整。
|
||||||
scoreLabel = similarity
|
- rerank、RRF、BM25、hybrid retrieval 还没有接入。
|
||||||
```
|
- 写入路径仍使用 SDK。
|
||||||
|
|
||||||
This makes the migration inspectable instead of hiding score changes behind one overloaded field.
|
这些不是当前迁移阻塞项,而是后续检索质量优化方向。
|
||||||
|
|
||||||
### Do Not Migrate Writes Yet
|
|
||||||
|
|
||||||
Writes and indexing still use the SDK path.
|
|
||||||
|
|
||||||
That is intentional. Migrating reads and writes at the same time would make debugging harder. The read path can be validated first; write-path migration can happen later if Spring AI `VectorStore.add(...)` fits the existing metadata and chunk model.
|
|
||||||
|
|
||||||
## Validation Story
|
|
||||||
|
|
||||||
I validated the refactor at multiple levels.
|
|
||||||
|
|
||||||
Unit tests cover:
|
|
||||||
|
|
||||||
- SDK mode.
|
|
||||||
- Spring AI mode.
|
|
||||||
- `auto` fallback.
|
|
||||||
- category filter behavior.
|
|
||||||
- distance metadata mapping.
|
|
||||||
|
|
||||||
Live API verification used:
|
|
||||||
|
|
||||||
```text
|
|
||||||
GET /api/search/similar?query=ERR_TIMEOUT&topK=3
|
|
||||||
```
|
|
||||||
|
|
||||||
Logs confirmed when the Spring AI VectorStore path was used and when fallback happened.
|
|
||||||
|
|
||||||
Then I compared SDK and VectorStore retrieval quality on representative queries:
|
|
||||||
|
|
||||||
| Query Type | Result |
|
|
||||||
| --- | --- |
|
|
||||||
| exact error code | same top3 |
|
|
||||||
| payment-service timeout | same top3 |
|
|
||||||
| MySQL connection pool | same top3 |
|
|
||||||
| AIOps alert-style query | same top3 |
|
|
||||||
| abstract RAG design query | same top1, VectorStore returned fewer tail results |
|
|
||||||
| category filter | both returned zero because metadata taxonomy did not match |
|
|
||||||
|
|
||||||
The acceptance decision was that Spring AI VectorStore is good enough for the current MVP read path, with SDK fallback preserved.
|
|
||||||
|
|
||||||
## Known Gaps
|
|
||||||
|
|
||||||
The refactor improved the architecture, but it did not solve every retrieval-quality problem.
|
|
||||||
|
|
||||||
Known gaps:
|
|
||||||
|
|
||||||
- Metadata taxonomy still needs cleanup, for example `database` vs `infrastructure`.
|
|
||||||
- Abstract design questions may need query rewriting or better indexed interview/devflow documents.
|
|
||||||
- Chunk context reconstruction is still limited when one logical section spans multiple chunks.
|
|
||||||
- `breadcrumb` now participates in embedding text, but it can still be used more strongly in context expansion, rerank, and evidence packing.
|
|
||||||
- Rerank, RRF, BM25, and hybrid retrieval are not implemented yet.
|
|
||||||
- Indexing writes still use SDK.
|
|
||||||
|
|
||||||
These are good follow-up issues because they are retrieval-quality improvements, not blockers for the VectorStore migration.
|
|
||||||
|
|
||||||
## How I Present This In An Interview
|
|
||||||
|
|
||||||
My short version would be:
|
|
||||||
|
|
||||||
> This RAG system started as a self-built MVP around Milvus SDK retrieval. It worked, but too much infrastructure logic lived in business code, and L0/L1 responsibilities were unclear. I refactored it in stages: first I added baseline retrieval cases, then made L0 a domain/entity hint instead of a final decision layer, then added evidence postprocessing, and finally moved the main read path to Spring AI VectorStore with SDK fallback. I kept `lookup_knowledge` as an explicit Agent tool because the project values traceability: the interviewer can see when retrieval happened, what evidence was found, and how it supported the diagnosis. The result is closer to standard Spring AI RAG while still preserving business-specific observability.
|
|
||||||
|
|
||||||
If asked why this is not a full framework migration:
|
|
||||||
|
|
||||||
> I intentionally did not migrate everything at once. Reads moved first because they are easier to compare using golden queries. Writes/indexing stayed on SDK to avoid mixing schema and retrieval behavior changes in one step. Advisors were not used as the main interface because hidden retrieval would weaken the Agent trace.
|
|
||||||
|
|
||||||
If asked what I would improve next:
|
|
||||||
|
|
||||||
> I would add query transformation for AIOps payloads, improve metadata taxonomy, use breadcrumb and section metadata for context expansion, and then evaluate whether hybrid retrieval or rerank is necessary based on measured recall and topK overlap.
|
|
||||||
|
|
||||||
## Interview Follow-Up Questions
|
|
||||||
|
|
||||||
### Why introduce Spring AI VectorStore if the SDK path already worked?
|
|
||||||
|
|
||||||
Because SDK-only retrieval made the project own too much low-level RAG infrastructure. `VectorStore` gives a standard abstraction for retrieval and makes future Spring AI features easier to adopt, while the facade keeps the Agent layer stable.
|
|
||||||
|
|
||||||
### Why keep custom code at all?
|
|
||||||
|
|
||||||
The custom code is where the Agent engineering value lives: AIOps payload mapping, L0 hints, evidence packing, score compatibility, and tool invocation tracing. Those are domain-specific and should remain visible.
|
|
||||||
|
|
||||||
### How do you know quality did not regress?
|
|
||||||
|
|
||||||
I compared SDK and VectorStore modes on representative live queries. Core troubleshooting and AIOps cases returned the same top3 documents in the same order. The differences were isolated to abstract design queries and metadata taxonomy, which are documented follow-up work.
|
|
||||||
|
|
||||||
### What is the most important design decision?
|
|
||||||
|
|
||||||
Keeping a stable boundary: `lookup_knowledge` calls `VectorSearchService`, and `VectorSearchService` decides whether to use Spring AI or SDK. That boundary made the migration small enough to validate and explain.
|
|
||||||
|
|||||||
@@ -1,71 +1,69 @@
|
|||||||
# RAG Retrieval Quality Report
|
# RAG 检索质量报告
|
||||||
|
|
||||||
## Purpose
|
## 1. 目的
|
||||||
|
|
||||||
This report compares the live retrieval behavior of the original Milvus SDK path and the new Spring AI VectorStore path.
|
这份报告回答一个面试关键问题:
|
||||||
|
|
||||||
The goal is to answer an interview-critical question:
|
```text
|
||||||
|
迁移到 Spring AI VectorStore 后,怎么证明检索质量没有退化?
|
||||||
|
```
|
||||||
|
|
||||||
> After moving retrieval to Spring AI VectorStore, how do we know retrieval quality did not regress?
|
这不是完整 benchmark,而是针对当前 Milvus/Zilliz collection 的代表性 live smoke comparison。
|
||||||
|
|
||||||
This is not a full benchmark yet. It is a focused live smoke comparison using representative RAG queries against the current Milvus/Zilliz collection.
|
## 2. 验证设置
|
||||||
|
|
||||||
## Setup
|
服务端点:
|
||||||
|
|
||||||
Service endpoint:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
GET http://127.0.0.1:9900/api/search/similar
|
GET http://127.0.0.1:9900/api/search/similar
|
||||||
```
|
```
|
||||||
|
|
||||||
Collection:
|
collection:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
biz
|
biz
|
||||||
```
|
```
|
||||||
|
|
||||||
Compared modes:
|
对比模式:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
retrieval.vector-store.mode=sdk
|
retrieval.vector-store.mode=sdk
|
||||||
retrieval.vector-store.mode=spring-ai
|
retrieval.vector-store.mode=spring-ai
|
||||||
```
|
```
|
||||||
|
|
||||||
Each case used:
|
每个 case:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
topK=3
|
topK=3
|
||||||
```
|
```
|
||||||
|
|
||||||
The application was restarted once per mode using command-line configuration so no repository config file had to be changed.
|
## 3. 测试案例
|
||||||
|
|
||||||
## Cases
|
| Case | Query | 目的 |
|
||||||
|
|---|---|---|
|
||||||
|
| `err-timeout` | `ERR_TIMEOUT` | 精确错误码检索 |
|
||||||
|
| `payment-service-timeout` | `payment-service timeout` | 服务超时排障 |
|
||||||
|
| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | 数据库排障 |
|
||||||
|
| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps 告警式检索 |
|
||||||
|
| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | 抽象 RAG 设计问题 |
|
||||||
|
| `database-filter` | `mysql timeout`, category=`database` | metadata filter 行为 |
|
||||||
|
|
||||||
| Case | Query | Purpose |
|
## 4. 对比摘要
|
||||||
| --- | --- | --- |
|
|
||||||
| `err-timeout` | `ERR_TIMEOUT` | Exact error-code retrieval |
|
|
||||||
| `payment-service-timeout` | `payment-service timeout` | Service timeout troubleshooting |
|
|
||||||
| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | Database troubleshooting |
|
|
||||||
| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps alert-style retrieval |
|
|
||||||
| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | Abstract RAG design query |
|
|
||||||
| `database-filter` | `mysql timeout`, category=`database` | Metadata filter behavior |
|
|
||||||
|
|
||||||
## Summary
|
| Case | SDK 数量 | VectorStore 数量 | Top1 一致 | TopK 重叠 | 结论 |
|
||||||
|
|---|---:|---:|---|---:|---|
|
||||||
|
| `err-timeout` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
|
||||||
|
| `payment-service-timeout` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
|
||||||
|
| `mysql-connection-pool` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
|
||||||
|
| `high-cpu-payment` | 3 | 3 | 是 | 3/3 | AIOps 核心 query 一致 |
|
||||||
|
| `rag-l0-l1` | 3 | 1 | 是 | 1/3 | VectorStore 尾部结果更少 |
|
||||||
|
| `database-filter` | 0 | 0 | 不适用 | 不适用 | filter 行为一致,taxonomy 有问题 |
|
||||||
|
|
||||||
| Case | SDK Count | VectorStore Count | Top1 Same | TopK Overlap | Notes |
|
## 5. 代表性结果
|
||||||
| --- | ---: | ---: | --- | ---: | --- |
|
|
||||||
| `err-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
|
||||||
| `payment-service-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
|
||||||
| `mysql-connection-pool` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
|
||||||
| `high-cpu-payment` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
|
||||||
| `rag-l0-l1` | 3 | 1 | Yes | 1/3 | VectorStore returned only the strongest candidate |
|
|
||||||
| `database-filter` | 0 | 0 | N/A | N/A | Both paths applied the filter consistently; no live docs matched `category=database` |
|
|
||||||
|
|
||||||
## Representative Results
|
### ERR_TIMEOUT
|
||||||
|
|
||||||
### `ERR_TIMEOUT`
|
SDK:
|
||||||
|
|
||||||
SDK:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
1. ERR_TIMEOUT score=0.5659486 label=l2_distance
|
1. ERR_TIMEOUT score=0.5659486 label=l2_distance
|
||||||
@@ -73,7 +71,7 @@ SDK:
|
|||||||
3. Error handling score=0.7735061 label=l2_distance
|
3. Error handling score=0.7735061 label=l2_distance
|
||||||
```
|
```
|
||||||
|
|
||||||
VectorStore:
|
VectorStore:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity
|
1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity
|
||||||
@@ -81,15 +79,15 @@ VectorStore:
|
|||||||
3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity
|
3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity
|
||||||
```
|
```
|
||||||
|
|
||||||
Interpretation:
|
解释:
|
||||||
|
|
||||||
- Document ordering is identical.
|
- 排序一致。
|
||||||
- Compatibility `score` is identical to SDK L2 distance.
|
- 兼容 `score` 与 SDK L2 distance 一致。
|
||||||
- VectorStore `rawScore` exposes Spring AI similarity separately.
|
- `rawScore` 暴露 Spring AI similarity。
|
||||||
|
|
||||||
### `MySQL connection pool`
|
### MySQL connection pool
|
||||||
|
|
||||||
Both paths returned:
|
两条路径都返回:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
1. MySQL connection pool config
|
1. MySQL connection pool config
|
||||||
@@ -97,58 +95,23 @@ Both paths returned:
|
|||||||
3. idle-timeout
|
3. idle-timeout
|
||||||
```
|
```
|
||||||
|
|
||||||
Interpretation:
|
说明迁移保留了核心基础设施排障检索能力。
|
||||||
|
|
||||||
- The migration preserves a precise infrastructure troubleshooting retrieval case.
|
### HighCPUUsage payment-service
|
||||||
- Metadata fields such as title, category, and source remain available.
|
|
||||||
|
|
||||||
### `HighCPUUsage payment-service`
|
两条路径都返回 payment-service 高 CPU 相关排障文档,说明 AIOps 告警式 query 没有退化。
|
||||||
|
|
||||||
Both paths returned:
|
### rag-l0-l1
|
||||||
|
|
||||||
```text
|
VectorStore 只返回一个候选,但 Top1 与 SDK 一致。这说明抽象设计类 query 需要后续 query rewrite、补充索引或 threshold 调整。
|
||||||
1. 3. HighCPUUsage / payment-service troubleshooting steps
|
|
||||||
2. evidence mapping table row for HighCPUUsage/payment-service
|
|
||||||
3. 3.1 Symptom confirmation
|
|
||||||
```
|
|
||||||
|
|
||||||
Interpretation:
|
### database-filter
|
||||||
|
|
||||||
- AIOps-style alert terms still retrieve the expected troubleshooting document.
|
两条路径都返回 0,因为相关 MySQL 文档当前分类是 `infrastructure`,不是 `database`。这是 metadata taxonomy 问题,不是 VectorStore 回归。
|
||||||
- This is important because AIOps diagnosis depends on knowledge retrieval plus metrics/log evidence.
|
|
||||||
|
|
||||||
### `rag-l0-l1`
|
## 6. 分数兼容结论
|
||||||
|
|
||||||
SDK returned three results, while VectorStore returned one:
|
对比验证了当前分数设计:
|
||||||
|
|
||||||
```text
|
|
||||||
Top1: Return error information
|
|
||||||
```
|
|
||||||
|
|
||||||
Interpretation:
|
|
||||||
|
|
||||||
- Top1 did not regress.
|
|
||||||
- VectorStore appears stricter for low-similarity tail results because the Spring AI path uses `similarityThresholdAll()`.
|
|
||||||
- This is acceptable for current read-path migration, but it is worth tracking because abstract design questions may need query rewriting, better indexed docs, or adjusted threshold behavior.
|
|
||||||
|
|
||||||
### `database-filter`
|
|
||||||
|
|
||||||
Both paths returned zero results for:
|
|
||||||
|
|
||||||
```text
|
|
||||||
query=mysql timeout
|
|
||||||
category=database
|
|
||||||
```
|
|
||||||
|
|
||||||
Interpretation:
|
|
||||||
|
|
||||||
- The filter path is consistent.
|
|
||||||
- The live indexed MySQL docs are categorized as `infrastructure`, not `database`.
|
|
||||||
- This highlights a metadata taxonomy issue rather than a VectorStore migration regression.
|
|
||||||
|
|
||||||
## Score Compatibility
|
|
||||||
|
|
||||||
The comparison validates the score design:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
SDK:
|
SDK:
|
||||||
@@ -162,50 +125,24 @@ VectorStore:
|
|||||||
scoreLabel = similarity
|
scoreLabel = similarity
|
||||||
```
|
```
|
||||||
|
|
||||||
This keeps `lookup_knowledge` relevance normalization stable while still exposing the VectorStore score semantics for trace/debugging.
|
这样既保持 `lookup_knowledge` 原有归一化逻辑,又能暴露 VectorStore 语义。
|
||||||
|
|
||||||
## Findings
|
## 7. 验收结论
|
||||||
|
|
||||||
### Finding 1: Main live cases are equivalent
|
Spring AI VectorStore 读路径可以接受用于当前 MVP/面试:
|
||||||
|
|
||||||
For exact error code, service timeout, MySQL troubleshooting, and AIOps alert-style retrieval, SDK and VectorStore returned identical top3 documents in identical order.
|
- 核心排障和 AIOps case 与 SDK top3 一致。
|
||||||
|
- 分数兼容性保留。
|
||||||
|
- VectorStore 语义通过 `rawScore` 和 `scoreLabel` 可观察。
|
||||||
|
- SDK fallback 仍保留运行安全。
|
||||||
|
|
||||||
This is strong evidence that the read-path migration did not regress the most important demo and troubleshooting cases.
|
后续检索质量工作不阻塞这次迁移,应作为独立优化继续推进。
|
||||||
|
|
||||||
### Finding 2: Abstract RAG design queries need better retrieval support
|
## 8. 下一步
|
||||||
|
|
||||||
The `rag-l0-l1` query only returned one VectorStore candidate. The top result matched SDK top1, but the tail differed.
|
- 增加自动 live comparison 脚本。
|
||||||
|
- 在 offline evaluator 中加入 topK overlap、top1 hit、MRR。
|
||||||
|
- 规范 metadata category,例如 `database` 与 `infrastructure`。
|
||||||
|
- 为抽象设计类 query 增加 query rewriting。
|
||||||
|
- 后续再评估是否迁移写入路径到 `VectorStore.add(...)`。
|
||||||
|
|
||||||
This suggests the next quality work should focus on:
|
|
||||||
|
|
||||||
- Query transformation for abstract design questions.
|
|
||||||
- Better indexing of interview/devflow RAG design docs.
|
|
||||||
- Context expansion around same-section chunks.
|
|
||||||
- Possibly tuning VectorStore threshold behavior.
|
|
||||||
|
|
||||||
### Finding 3: Metadata taxonomy matters
|
|
||||||
|
|
||||||
The category filter case returned zero results in both modes because the relevant MySQL docs are categorized as `infrastructure`, not `database`.
|
|
||||||
|
|
||||||
This supports a previous RAG issue: category/domain metadata should be normalized before it is used as a hard filter.
|
|
||||||
|
|
||||||
## Acceptance Decision
|
|
||||||
|
|
||||||
The Spring AI VectorStore read path is accepted for current MVP/interview use:
|
|
||||||
|
|
||||||
- Core troubleshooting cases match SDK behavior.
|
|
||||||
- Score compatibility is preserved.
|
|
||||||
- The VectorStore path exposes better score semantics without changing the `lookup_knowledge` API.
|
|
||||||
- SDK fallback remains available for runtime safety.
|
|
||||||
|
|
||||||
The next retrieval-quality improvements should not block this migration. They should be handled as separate RAG quality work.
|
|
||||||
|
|
||||||
## Next Work
|
|
||||||
|
|
||||||
Recommended next steps:
|
|
||||||
|
|
||||||
- Add a small automated live comparison script if repeated validation becomes common.
|
|
||||||
- Add topK overlap and top1 hit metrics to the offline evaluator.
|
|
||||||
- Normalize metadata categories such as `database` vs `infrastructure`.
|
|
||||||
- Add query rewriting for abstract RAG questions.
|
|
||||||
- Decide later whether to migrate indexing writes to Spring AI `VectorStore.add(...)`.
|
|
||||||
|
|||||||
@@ -1,22 +1,16 @@
|
|||||||
# RAG VectorStore Interview Notes
|
# RAG VectorStore 面试要点
|
||||||
|
|
||||||
## 60-Second Explanation
|
## 1. 60 秒讲法
|
||||||
|
|
||||||
I refactored the RAG retrieval path from a direct Milvus SDK-only implementation to a Spring AI `VectorStore` main path, while keeping the SDK path as a fallback.
|
|
||||||
|
|
||||||
The important part is not just the dependency change. I kept `VectorSearchService` as the boundary, so `lookup_knowledge` and the Agent workflow did not need to change. The system now supports three modes:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
auto -> try Spring AI VectorStore, fallback to SDK
|
我把 RAG 检索从 Milvus SDK-only 重构为 Spring AI VectorStore 主路径,同时保留 SDK fallback。
|
||||||
spring-ai -> force VectorStore
|
关键不是换了一个依赖,而是保留 VectorSearchService 作为边界,所以 lookup_knowledge 和 Agent workflow 不需要改。
|
||||||
sdk -> force SDK
|
现在支持 auto、spring-ai、sdk 三种模式。auto 会优先尝试 VectorStore,失败后 fallback 到 SDK。
|
||||||
```
|
```
|
||||||
|
|
||||||
During live verification, the first run found a real config mismatch: VectorStore was pointed at `business_knowledge`, but the real Zilliz collection was `biz`. The fallback worked, so the system still returned results through SDK. After aligning the collection name, the same query went through Spring AI VectorStore successfully.
|
现场验证时,第一次发现 VectorStore 指向了错误 collection:`business_knowledge`,而实际 Zilliz collection 是 `biz`。fallback 生效,所以系统仍能通过 SDK 返回结果。修正 collection 后,同一个 query 成功走 Spring AI VectorStore。
|
||||||
|
|
||||||
I also fixed score compatibility. Spring AI Milvus exposes similarity as the document score, but the old `lookup_knowledge` logic expects L2 distance. So I preserve `rawScore` and `scoreLabel`, and use Milvus `metadata.distance` as the compatibility `score` when available.
|
## 2. 架构回答
|
||||||
|
|
||||||
## Architecture Answer
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Agent / API
|
Agent / API
|
||||||
@@ -27,69 +21,65 @@ Agent / API
|
|||||||
-> Milvus/Zilliz collection: biz
|
-> Milvus/Zilliz collection: biz
|
||||||
```
|
```
|
||||||
|
|
||||||
The key design choice is that `VectorSearchService` remains the retrieval facade. This avoids spreading framework-specific code into the Agent tool layer.
|
关键设计:`VectorSearchService` 是检索门面,避免 Spring AI 或 SDK 细节扩散到 Agent 工具层。
|
||||||
|
|
||||||
## Why Keep The SDK Path?
|
## 3. 为什么保留 SDK
|
||||||
|
|
||||||
I kept SDK fallback for three reasons:
|
- 迁移安全:原 SDK 路径已验证可用。
|
||||||
|
- 运行韧性:VectorStore schema、filter 或配置失败时,检索仍可用。
|
||||||
|
- Demo 稳定:检索抽象变化不应该破坏主诊断演示。
|
||||||
|
|
||||||
- Migration safety: the existing SDK path was already proven against the live collection.
|
这在实际验证中发挥了作用:VectorStore 配置错时,`auto` 模式 fallback 到 SDK,API 没有失败。
|
||||||
- Runtime resilience: if VectorStore schema mapping or filtering fails, retrieval still works.
|
|
||||||
- Interview/demo stability: a retrieval abstraction change should not break the main Agent diagnosis demo.
|
|
||||||
|
|
||||||
This was validated in practice. When VectorStore pointed at the wrong collection, `auto` mode fell back to SDK and still returned results.
|
## 4. 为什么引入 Spring AI VectorStore
|
||||||
|
|
||||||
## Why Use Spring AI VectorStore At All?
|
使用 `VectorStore` 可以让项目更接近标准 RAG 抽象:
|
||||||
|
|
||||||
Using Spring AI `VectorStore` moves the project closer to a standard RAG abstraction:
|
- 业务代码不再持有全部 Milvus search 细节。
|
||||||
|
- 后续 QueryTransformer、DocumentPostProcessor、Retriever 等能力更容易接入。
|
||||||
|
- 面试中也更容易解释和 Spring AI 生态的关系。
|
||||||
|
|
||||||
- Retrieval code no longer needs to own all Milvus-specific search details.
|
但我没有一次性迁移写入,因为读写同时迁移会让问题难定位。当前先稳定读路径。
|
||||||
- Later features such as query transformers, document postprocessors, advisors, or retrievers can be introduced more naturally.
|
|
||||||
- The code becomes easier to compare with common Spring AI RAG patterns in an interview.
|
|
||||||
|
|
||||||
But I did not blindly replace everything. Writes/indexing still use SDK because changing read and write paths at the same time would make failures harder to isolate.
|
## 5. 为什么保留 L0
|
||||||
|
|
||||||
## Why Keep L0?
|
L0 现在不是最终答案来源,而是确定性 hint 层:
|
||||||
|
|
||||||
L0 is no longer treated as the final source of truth. It is a deterministic hint layer:
|
- 提取 domain/entity。
|
||||||
|
- 在可能时生成 category filter。
|
||||||
|
- 给 trace 提供解释信号。
|
||||||
|
|
||||||
- It extracts domain/entity hints from indexed metadata.
|
当前职责:
|
||||||
- It helps constrain L1 retrieval by category when possible.
|
|
||||||
- It gives the Agent a stable clue even when semantic retrieval is weak.
|
|
||||||
|
|
||||||
The current design is:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
L0 = domain/entity hint
|
L0 = domain/entity hint
|
||||||
L1 = semantic retrieval through VectorStore/SDK
|
L1 = semantic retrieval through VectorStore/SDK
|
||||||
postprocess = evidence trace and relevance normalization
|
postprocess = evidence trace + relevance normalization
|
||||||
```
|
```
|
||||||
|
|
||||||
This is easier to defend than saying "we only use vector search." Real incident diagnosis often has exact identifiers, error codes, service names, and alert names. L0 is useful for those.
|
真实故障诊断里有很多精确标识,完全只靠向量检索并不稳。
|
||||||
|
|
||||||
## Why Not Use Hidden Spring AI Advisors Directly?
|
## 6. 为什么不用隐藏 Advisor
|
||||||
|
|
||||||
For this project, `lookup_knowledge` remains an explicit tool.
|
`lookup_knowledge` 保持显式工具,因为:
|
||||||
|
|
||||||
Reason:
|
- Trace 要展示什么时候检索。
|
||||||
|
- `tool_invocation` 要记录输入、输出预览、相关性和 metadata。
|
||||||
|
- 面试故事是可审计 Agent 执行,而不只是答案质量。
|
||||||
|
|
||||||
- The Agent trace needs to show when knowledge was retrieved.
|
Advisor 后续可以接入,但需要先解决可观测性。
|
||||||
- `tool_invocation` records input, output preview, relevance level, and evidence metadata.
|
|
||||||
- The interview story is about auditable Agent execution, not only answer quality.
|
|
||||||
|
|
||||||
Spring AI Advisors may be useful later, but hiding retrieval inside an advisor would make the evidence chain less visible unless we rebuild trace hooks around it.
|
## 7. 分数设计
|
||||||
|
|
||||||
## Score Design
|
当前结果故意拆成:
|
||||||
|
|
||||||
The result object intentionally separates these fields:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
score -> compatibility score used by old relevance normalization
|
score -> 兼容旧 relevance normalization 的分数
|
||||||
rawScore -> raw score from the retrieval implementation
|
rawScore -> 当前检索实现原始分数
|
||||||
scoreLabel -> semantic label for rawScore
|
scoreLabel -> rawScore 的语义
|
||||||
```
|
```
|
||||||
|
|
||||||
For SDK:
|
SDK:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
score = L2 distance
|
score = L2 distance
|
||||||
@@ -97,7 +87,7 @@ rawScore = L2 distance
|
|||||||
scoreLabel = l2_distance
|
scoreLabel = l2_distance
|
||||||
```
|
```
|
||||||
|
|
||||||
For VectorStore:
|
VectorStore:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
score = metadata.distance if present
|
score = metadata.distance if present
|
||||||
@@ -105,63 +95,25 @@ rawScore = Spring AI document score
|
|||||||
scoreLabel = similarity
|
scoreLabel = similarity
|
||||||
```
|
```
|
||||||
|
|
||||||
This prevents a subtle bug: if we treat Spring AI similarity as L2 distance, relevance becomes wrong. If we only expose distance, we lose the ability to compare Spring AI behavior. Keeping both makes the migration inspectable.
|
这样避免把 similarity 当成 L2 distance 的隐蔽 bug。
|
||||||
|
|
||||||
## How I Verified It
|
## 8. 如何证明 VectorStore 被使用
|
||||||
|
|
||||||
I verified at three levels:
|
- 日志出现 `Starting Spring AI VectorStore search` 和 `Spring AI VectorStore search complete`。
|
||||||
|
- API 响应中 `scoreLabel=similarity`。
|
||||||
|
- `rawScore` 是 Spring AI similarity,`score` 仍是兼容 distance。
|
||||||
|
|
||||||
- Unit tests: SDK mode, auto VectorStore mode, fallback mode, category filter, distance metadata mapping.
|
## 9. 常见追问
|
||||||
- Live API: `/api/search/similar?query=ERR_TIMEOUT&topK=3`.
|
|
||||||
- Logs: confirmed whether the path was VectorStore success or SDK fallback.
|
|
||||||
|
|
||||||
The live API returned:
|
### 为什么不删 SDK?
|
||||||
|
|
||||||
```text
|
这是迁移,不是重写。fallback 提供回滚安全,并且已经证明配置错误时仍能保证主链路可用。
|
||||||
scoreLabel = similarity
|
|
||||||
rawScore = Spring AI similarity
|
|
||||||
score = Milvus distance metadata
|
|
||||||
```
|
|
||||||
|
|
||||||
That means the main path was Spring AI VectorStore and compatibility scoring remained stable.
|
### `lookup_knowledge` 变了吗?
|
||||||
|
|
||||||
## What I Would Do Next
|
外部契约没变。它仍然调用 `VectorSearchService.searchSimilarDocuments(...)`,变化在门面背后的实现。
|
||||||
|
|
||||||
I would not immediately migrate indexing writes. The next responsible steps are:
|
### 这是完整 Spring AI RAG 了吗?
|
||||||
|
|
||||||
- Add a small live acceptance report for several golden queries.
|
还不是。当前是 Spring AI VectorStore 读路径 + 显式工具 + 自定义 evidence trace + SDK 写入。这样做是为了保留审计能力和分阶段迁移安全。
|
||||||
- Compare `sdk` and `spring-ai` mode side by side for topK overlap.
|
|
||||||
- Decide whether `VectorIndexService` should move to `VectorStore.add(...)`.
|
|
||||||
- Add query transformation or hybrid retrieval only after we have baseline metrics.
|
|
||||||
|
|
||||||
This staged approach is intentional: first stabilize the read path, then evaluate retrieval quality, then migrate writes if the abstraction proves reliable.
|
|
||||||
|
|
||||||
## Interview Questions And Short Answers
|
|
||||||
|
|
||||||
### Why did you not remove the SDK?
|
|
||||||
|
|
||||||
Because this is a migration, not a rewrite. SDK fallback gives rollback safety and proved useful when VectorStore config was initially wrong.
|
|
||||||
|
|
||||||
### What changed for `lookup_knowledge`?
|
|
||||||
|
|
||||||
The public contract did not change. It still calls `VectorSearchService.searchSimilarDocuments(...)`. The implementation behind that facade changed.
|
|
||||||
|
|
||||||
### How do you know VectorStore is actually used?
|
|
||||||
|
|
||||||
The logs show `Starting Spring AI VectorStore search` followed by `Spring AI VectorStore search complete`. The API response also has `scoreLabel=similarity`, which only comes from the VectorStore path.
|
|
||||||
|
|
||||||
### What was the main bug found during live validation?
|
|
||||||
|
|
||||||
The configured collection name was wrong. Spring AI looked for `business_knowledge`, but the actual Milvus collection was `biz`.
|
|
||||||
|
|
||||||
### What did fallback prove?
|
|
||||||
|
|
||||||
It proved that `auto` mode is resilient: VectorStore failed, SDK search still returned valid results, and the API did not fail.
|
|
||||||
|
|
||||||
### Why is `metadata.distance` important?
|
|
||||||
|
|
||||||
Because `lookup_knowledge` uses L2 distance normalization. Spring AI returns similarity as the main document score, but the Milvus distance is available in metadata. Using it preserves old relevance behavior.
|
|
||||||
|
|
||||||
### Is this full Spring AI RAG now?
|
|
||||||
|
|
||||||
Not yet. It uses Spring AI VectorStore for the main read path, but keeps explicit tools, custom evidence trace, L0 hints, and SDK indexing. That is deliberate because the project values auditability and staged migration.
|
|
||||||
|
|||||||
@@ -1,17 +1,17 @@
|
|||||||
# RAG VectorStore Live Acceptance
|
# RAG VectorStore Live 验收说明
|
||||||
|
|
||||||
## Purpose
|
## 1. 目的
|
||||||
|
|
||||||
This note records the live acceptance result for the RAG retrieval refactor.
|
本文记录 RAG 检索重构的 live 验收结论。
|
||||||
|
|
||||||
The goal of this refactor was not only to add a Spring AI abstraction, but to prove that the production retrieval path can:
|
这次重构的目标不只是接入 Spring AI 抽象,而是证明线上读路径能够:
|
||||||
|
|
||||||
- Prefer Spring AI `VectorStore` for Milvus retrieval.
|
- 优先使用 Spring AI `VectorStore` 做 Milvus 检索。
|
||||||
- Preserve the existing Milvus SDK path as fallback.
|
- 保留原 Milvus SDK 作为 fallback。
|
||||||
- Keep the `lookup_knowledge` tool contract stable.
|
- 保持 `lookup_knowledge` 工具契约稳定。
|
||||||
- Keep L2-distance based relevance normalization compatible.
|
- 保持基于 L2 distance 的相关性归一化兼容。
|
||||||
|
|
||||||
## Current Retrieval Shape
|
## 2. 当前检索形态
|
||||||
|
|
||||||
```text
|
```text
|
||||||
lookup_knowledge / /api/search/similar
|
lookup_knowledge / /api/search/similar
|
||||||
@@ -19,22 +19,22 @@ lookup_knowledge / /api/search/similar
|
|||||||
-> retrieval.vector-store.mode
|
-> retrieval.vector-store.mode
|
||||||
-> auto
|
-> auto
|
||||||
-> Spring AI VectorStore
|
-> Spring AI VectorStore
|
||||||
-> fallback to Milvus SDK if VectorStore fails
|
-> VectorStore 失败时 fallback 到 Milvus SDK
|
||||||
-> spring-ai
|
-> spring-ai
|
||||||
-> Spring AI VectorStore only
|
-> 只走 Spring AI VectorStore
|
||||||
-> sdk
|
-> sdk
|
||||||
-> Milvus SDK only
|
-> 只走 Milvus SDK
|
||||||
```
|
```
|
||||||
|
|
||||||
## Configuration Verified
|
## 3. 已验证配置
|
||||||
|
|
||||||
The live Milvus/Zilliz database contains the collection:
|
live Milvus/Zilliz 数据库中存在 collection:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
biz
|
biz
|
||||||
```
|
```
|
||||||
|
|
||||||
The Spring AI VectorStore configuration was aligned with the existing SDK collection:
|
Spring AI VectorStore 配置与 SDK 使用的 collection 对齐:
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
spring:
|
spring:
|
||||||
@@ -53,11 +53,11 @@ spring:
|
|||||||
embedding-field-name: vector
|
embedding-field-name: vector
|
||||||
```
|
```
|
||||||
|
|
||||||
Why this matters: the earlier config used `business_knowledge`, but the SDK path and real collection use `biz`. That mismatch proved the fallback worked, but it also meant VectorStore was not the successful main path until the config was corrected.
|
为什么重要:早期配置使用 `business_knowledge`,而真实 collection 是 `biz`。这个错配证明了 fallback 生效,但也说明修正前 VectorStore 不是成功主路径。
|
||||||
|
|
||||||
## Commands Used
|
## 4. 验收命令
|
||||||
|
|
||||||
Health check:
|
健康检查:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
@@ -65,7 +65,7 @@ Invoke-RestMethod `
|
|||||||
-Method Get
|
-Method Get
|
||||||
```
|
```
|
||||||
|
|
||||||
Observed result:
|
期望:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@@ -74,7 +74,7 @@ Observed result:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
Direct retrieval check:
|
直接检索:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
@@ -82,7 +82,7 @@ Invoke-RestMethod `
|
|||||||
-Method Get
|
-Method Get
|
||||||
```
|
```
|
||||||
|
|
||||||
Observed result shape:
|
期望结果形态:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@@ -90,7 +90,6 @@ Observed result shape:
|
|||||||
"message": "success",
|
"message": "success",
|
||||||
"data": [
|
"data": [
|
||||||
{
|
{
|
||||||
"id": "f7dff7c8-5665-3145-9f75-ef741528b914",
|
|
||||||
"content": "### ERR_TIMEOUT ...",
|
"content": "### ERR_TIMEOUT ...",
|
||||||
"score": 0.5662,
|
"score": 0.5662,
|
||||||
"rawScore": 0.4337,
|
"rawScore": 0.4337,
|
||||||
@@ -105,41 +104,40 @@ Observed result shape:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
## What The Logs Proved
|
## 5. 日志证明了什么
|
||||||
|
|
||||||
Before collection alignment:
|
collection 修正前:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Starting Spring AI VectorStore search
|
Starting Spring AI VectorStore search
|
||||||
SearchRequest collectionName:business_knowledge failed
|
|
||||||
Spring AI VectorStore retrieval failed, falling back to Milvus SDK
|
Spring AI VectorStore retrieval failed, falling back to Milvus SDK
|
||||||
Starting Milvus SDK search
|
Starting Milvus SDK search
|
||||||
```
|
```
|
||||||
|
|
||||||
After collection alignment:
|
collection 修正后:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Starting Spring AI VectorStore search: query=ERR_TIMEOUT
|
Starting Spring AI VectorStore search: query=ERR_TIMEOUT
|
||||||
Spring AI VectorStore search complete, candidates=3
|
Spring AI VectorStore search complete, candidates=3
|
||||||
```
|
```
|
||||||
|
|
||||||
This proves:
|
这证明:
|
||||||
|
|
||||||
- `auto` mode really attempts VectorStore first.
|
- `auto` 模式确实先尝试 VectorStore。
|
||||||
- The fallback is functional when VectorStore fails.
|
- VectorStore 失败时 fallback 可用。
|
||||||
- After config alignment, the main path is Spring AI VectorStore rather than SDK fallback.
|
- 配置对齐后,主路径是 Spring AI VectorStore,而不是 SDK fallback。
|
||||||
|
|
||||||
## Score Semantics
|
## 6. 分数语义
|
||||||
|
|
||||||
The project keeps three score fields intentionally:
|
项目保留三个分数字段:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
rawScore -> the raw score from the active retrieval implementation
|
rawScore -> 当前检索实现的原始分数
|
||||||
scoreLabel -> the semantic meaning of rawScore
|
scoreLabel -> rawScore 的语义
|
||||||
score -> compatibility score used by existing lookup relevance normalization
|
score -> lookup relevance normalization 使用的兼容分数
|
||||||
```
|
```
|
||||||
|
|
||||||
For SDK retrieval:
|
SDK:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
rawScore = L2 distance
|
rawScore = L2 distance
|
||||||
@@ -147,7 +145,7 @@ scoreLabel = l2_distance
|
|||||||
score = L2 distance
|
score = L2 distance
|
||||||
```
|
```
|
||||||
|
|
||||||
For Spring AI VectorStore retrieval:
|
Spring AI VectorStore:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
rawScore = Spring AI similarity score
|
rawScore = Spring AI similarity score
|
||||||
@@ -155,45 +153,46 @@ scoreLabel = similarity
|
|||||||
score = Milvus distance metadata when available
|
score = Milvus distance metadata when available
|
||||||
```
|
```
|
||||||
|
|
||||||
Why use `metadata.distance` for `score`: `LookupKnowledgeTool` already normalizes relevance from L2 distance. Spring AI Milvus returns similarity as the document score, but also includes the Milvus distance in metadata. Using distance preserves the old relevance behavior while still exposing the new VectorStore score semantics through `rawScore` and `scoreLabel`.
|
使用 `metadata.distance` 的原因:`LookupKnowledgeTool` 已经基于 L2 distance 做相关性归一化。Spring AI Milvus 主分数是 similarity,但 metadata 中仍有 Milvus distance。用 distance 保持旧逻辑稳定,同时通过 `rawScore` 暴露新语义。
|
||||||
|
|
||||||
## Regression Checks
|
## 7. 回归检查
|
||||||
|
|
||||||
Targeted tests:
|
目标测试:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
mvn -q "-Dtest=VectorSearchServiceTest,LookupKnowledgeToolTest" test
|
mvn -q "-Dtest=VectorSearchServiceTest,LookupKnowledgeToolTest" test
|
||||||
```
|
```
|
||||||
|
|
||||||
Spec validation:
|
相关 spec:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
openspec.cmd validate rag-knowledge-retrieval --specs
|
openspec.cmd validate rag-knowledge-retrieval --specs
|
||||||
openspec.cmd validate rag-retrieval-evaluation --specs
|
openspec.cmd validate rag-retrieval-evaluation --specs
|
||||||
```
|
```
|
||||||
|
|
||||||
Whitespace check:
|
diff 检查:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
git diff --check
|
git diff --check
|
||||||
```
|
```
|
||||||
|
|
||||||
Observed result:
|
验收结论:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
All targeted tests passed.
|
目标测试通过。
|
||||||
All related specs passed.
|
相关 spec 通过。
|
||||||
No diff-check errors.
|
diff-check 无错误。
|
||||||
```
|
```
|
||||||
|
|
||||||
## Acceptance Conclusion
|
## 8. 验收结论
|
||||||
|
|
||||||
The VectorStore refactor is accepted for the read path:
|
VectorStore 读路径可以接受:
|
||||||
|
|
||||||
- Spring AI VectorStore is integrated and selected in `auto` mode.
|
- Spring AI VectorStore 已集成,并在 `auto` 模式中优先使用。
|
||||||
- The SDK path remains available and was proven by fallback behavior.
|
- SDK fallback 保留且已被实际验证。
|
||||||
- The live collection configuration is aligned with the existing Milvus collection.
|
- live collection 配置与现有 Milvus collection 对齐。
|
||||||
- The `lookup_knowledge` public contract remains stable.
|
- `lookup_knowledge` 对外契约保持稳定。
|
||||||
- Existing L2-based relevance normalization remains compatible.
|
- 旧的 L2 relevance normalization 仍兼容。
|
||||||
|
|
||||||
|
写入和索引路径仍使用 Milvus SDK。这是有意的分阶段迁移,不是验收失败项。
|
||||||
|
|
||||||
The write/indexing path still uses the Milvus SDK. That is an intentional staged migration decision, not a failed acceptance item.
|
|
||||||
|
|||||||
@@ -0,0 +1,225 @@
|
|||||||
|
# 面试故事案例
|
||||||
|
|
||||||
|
**用途**:把项目能力讲成可被面试官理解的工程故事
|
||||||
|
**使用方式**:按问题选择一个故事,不需要从头到尾背诵
|
||||||
|
|
||||||
|
## 故事 1:从黑盒 Chatbot 到可追踪 Agent
|
||||||
|
|
||||||
|
### 面试官问题
|
||||||
|
|
||||||
|
```text
|
||||||
|
这个项目和普通调用大模型有什么区别?
|
||||||
|
```
|
||||||
|
|
||||||
|
### 30 秒回答
|
||||||
|
|
||||||
|
```text
|
||||||
|
普通 Chatbot 只给最终答案,出了问题很难解释答案怎么来的。
|
||||||
|
我这个项目把诊断过程拆成 Planner、Executor、Verifier,并把每个 Agent 步骤和每次工具调用落库。
|
||||||
|
最后通过 Trace API 可以回放:模型怎么规划、调用了哪些工具、工具返回了什么证据、Verifier 怎么判断答案可信。
|
||||||
|
```
|
||||||
|
|
||||||
|
### 展开讲法
|
||||||
|
|
||||||
|
一开始最容易做的是:用户问题进来,直接让模型回答。但故障诊断场景不能只看答案,因为答案可能看起来合理却没有证据支撑。
|
||||||
|
|
||||||
|
所以我把系统拆成三层:
|
||||||
|
|
||||||
|
- `diagnosis_session` 记录一次诊断的主状态和最终答案。
|
||||||
|
- `agent_step` 记录 Planner、Executor、Verifier 的模型调用。
|
||||||
|
- `tool_invocation` 记录知识库、日志、指标等真实工具证据。
|
||||||
|
|
||||||
|
这样就能做到:答案不是孤立文本,而是一条可审计的执行链。
|
||||||
|
|
||||||
|
### 可展示文件
|
||||||
|
|
||||||
|
- `mvp/architecture/interview-one-pager.md`
|
||||||
|
- `mvp/architecture/session-trace-lifecycle.md`
|
||||||
|
- `mvp/demo/output/trace-response.json`
|
||||||
|
|
||||||
|
### 主动说不足
|
||||||
|
|
||||||
|
```text
|
||||||
|
当前 tool_invocation.step_id 还不是每次都强绑定具体 agent_step,后续可以加 runId 和更严格的 step 关联,让多轮同 session 诊断更清晰。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 故事 2:RAG 从自研 SDK 检索迁移到 Spring AI VectorStore
|
||||||
|
|
||||||
|
### 面试官问题
|
||||||
|
|
||||||
|
```text
|
||||||
|
你的 RAG 是怎么设计的?为什么不用框架全包?
|
||||||
|
```
|
||||||
|
|
||||||
|
### 30 秒回答
|
||||||
|
|
||||||
|
```text
|
||||||
|
我把 RAG 分成两部分:通用检索基础设施尽量交给 Spring AI VectorStore,业务可观测链路留在项目里。
|
||||||
|
所以 Agent 仍然显式调用 lookup_knowledge,底层通过 VectorSearchService 走 Spring AI VectorStore,失败时 fallback 到原 Milvus SDK。
|
||||||
|
这样既能减少自研检索代码,又不会丢失工具调用 trace。
|
||||||
|
```
|
||||||
|
|
||||||
|
### 展开讲法
|
||||||
|
|
||||||
|
旧实现里,Milvus SDK 查询、topK、filter、score 映射都在业务代码里。它能跑,但后续扩展成本高。
|
||||||
|
|
||||||
|
我没有直接把 RAG 隐藏进 Advisor,因为这个项目的核心是 Agent 工程,需要知道 Agent 何时检索、检索了什么、证据怎么支撑诊断。
|
||||||
|
|
||||||
|
于是我保留了边界:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Executor -> lookup_knowledge -> VectorSearchService -> VectorStore / SDK fallback
|
||||||
|
```
|
||||||
|
|
||||||
|
同时把 L0 从“直接返回结果”降级为 domain/entity hint,降低关键词误召回的风险。
|
||||||
|
|
||||||
|
### 可展示文件
|
||||||
|
|
||||||
|
- `mvp/architecture/rag-architecture.md`
|
||||||
|
- `mvp/architecture/retrieval-observability.md`
|
||||||
|
- `interview/rag-refactor-story.md`
|
||||||
|
|
||||||
|
### 主动说不足
|
||||||
|
|
||||||
|
```text
|
||||||
|
当前还没有完整 hybrid search 和 rerank。
|
||||||
|
我先做 golden cases、VectorStore 主路径和 SDK fallback,是为了让每一步迁移都能被验证。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 故事 3:Verifier 如何降低幻觉风险
|
||||||
|
|
||||||
|
### 面试官问题
|
||||||
|
|
||||||
|
```text
|
||||||
|
Agent 怎么保证不胡说?
|
||||||
|
```
|
||||||
|
|
||||||
|
### 30 秒回答
|
||||||
|
|
||||||
|
```text
|
||||||
|
我没有假设模型天然可靠,而是加了 Verifier。
|
||||||
|
Executor 给出答案后,Verifier 只拿 executor_final_answer 和 tool_trace_summary,不允许做新检索。
|
||||||
|
它把答案里的关键事实逐条校验,输出 PASS、LOW_CONFID 或 REJECT。
|
||||||
|
这个结果会写回 self_evaluation,Trace API 可以看到。
|
||||||
|
```
|
||||||
|
|
||||||
|
### 展开讲法
|
||||||
|
|
||||||
|
Verifier 的关键不是再问一次模型“你觉得对吗”,而是让它基于真实工具调用做 groundedness 检查。
|
||||||
|
|
||||||
|
`ToolTraceSummaryService` 会从 `tool_invocation` 里整理证据索引,包含:
|
||||||
|
|
||||||
|
- 工具名。
|
||||||
|
- 输入摘要。
|
||||||
|
- 输出摘要。
|
||||||
|
- evidence level。
|
||||||
|
- source invocation ids。
|
||||||
|
|
||||||
|
Verifier 输出结构化 JSON,ChatService 根据 verdict 决定是否输出、补证据或降级。
|
||||||
|
|
||||||
|
### 可展示文件
|
||||||
|
|
||||||
|
- `mvp/architecture/harness-quality-gates.md`
|
||||||
|
- `mvp/architecture/feedback-architecture.md`
|
||||||
|
- `src/main/resources/prompts/chat-verifier-prompt.md`
|
||||||
|
|
||||||
|
### 主动说不足
|
||||||
|
|
||||||
|
```text
|
||||||
|
AIOps 当前还是轻量 rule evaluation,不是完整 LLM Verifier。
|
||||||
|
这是有意收敛:先用规则保证 payload 聚焦和工具证据使用,后续再加 AIOps LLM Verifier。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 故事 4:AIOps 告警为什么要做 payload scope control
|
||||||
|
|
||||||
|
### 面试官问题
|
||||||
|
|
||||||
|
```text
|
||||||
|
AIOps 场景和普通 Chat 有什么区别?
|
||||||
|
```
|
||||||
|
|
||||||
|
### 30 秒回答
|
||||||
|
|
||||||
|
```text
|
||||||
|
AIOps 告警有一个很关键的问题:环境里可能同时有很多活跃告警,Agent 容易跑偏。
|
||||||
|
所以我把 AIOps 分成 PAYLOAD_TARGETED 和 AUTO_DISCOVERY。
|
||||||
|
如果请求带 alert payload,最终报告必须聚焦输入告警,并且会把 alertName、service、severity、description 拼成 recommended lookup_knowledge query。
|
||||||
|
```
|
||||||
|
|
||||||
|
### 展开讲法
|
||||||
|
|
||||||
|
没有 payload 时,Agent 可以先查询活跃告警,再选择目标排查。
|
||||||
|
|
||||||
|
但有 payload 时,用户已经告诉系统“我要查这个告警”。这时如果 Agent 把其他活跃告警写成主根因,产品体验会很差。
|
||||||
|
|
||||||
|
所以我做了两件事:
|
||||||
|
|
||||||
|
- Prompt 中明确 `PAYLOAD_TARGETED` 范围。
|
||||||
|
- `AiOpsRuleEvaluationService` 检查最终报告是否聚焦输入告警,以及是否使用证据工具。
|
||||||
|
|
||||||
|
### 可展示文件
|
||||||
|
|
||||||
|
- `mvp/architecture/agent-orchestration.md`
|
||||||
|
- `mvp/architecture/current-mvp-architecture.md`
|
||||||
|
- `interview/aiops-query-augmentation.md`
|
||||||
|
- `interview/aiops-lightweight-verifier.md`
|
||||||
|
|
||||||
|
### 主动说不足
|
||||||
|
|
||||||
|
```text
|
||||||
|
当前 scope control 主要靠 prompt 和规则评估。
|
||||||
|
后续可以把 AIOps 也接入类似 Chat Verifier 的事实校验,让告警报告的每个根因和建议都有 evidence refs。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 故事 5:反馈不是点赞按钮,而是案例沉淀入口
|
||||||
|
|
||||||
|
### 面试官问题
|
||||||
|
|
||||||
|
```text
|
||||||
|
用户反馈在系统里有什么用?
|
||||||
|
```
|
||||||
|
|
||||||
|
### 30 秒回答
|
||||||
|
|
||||||
|
```text
|
||||||
|
反馈不只是前端按钮。
|
||||||
|
用户提交 useful 后,系统会把同一个 diagnosis_session 沉淀为 case_library。
|
||||||
|
not_useful 不会改执行状态,而是作为 bad case 信号保留。
|
||||||
|
这样 status、self_evaluation、feedback 三个维度是分开的。
|
||||||
|
```
|
||||||
|
|
||||||
|
### 展开讲法
|
||||||
|
|
||||||
|
我刻意没有把 `not_useful` 写成 `FAILED`。因为失败表示系统执行异常,而用户觉得不好用是质量标签。
|
||||||
|
|
||||||
|
当前设计里:
|
||||||
|
|
||||||
|
```text
|
||||||
|
status -> 执行是否成功
|
||||||
|
self_evaluation -> 系统自己判断证据和事实支撑度
|
||||||
|
feedback -> 用户是否认可
|
||||||
|
```
|
||||||
|
|
||||||
|
`useful` 会进入 `CaseLibraryService.createFromSession`,生成可复用案例。后续可以做相似案例推荐或高质量样本积累。
|
||||||
|
|
||||||
|
### 可展示文件
|
||||||
|
|
||||||
|
- `mvp/architecture/feedback-architecture.md`
|
||||||
|
- `mvp/architecture/data-model.md`
|
||||||
|
- `mvp/demo/output/feedback-response.json`
|
||||||
|
|
||||||
|
### 主动说不足
|
||||||
|
|
||||||
|
```text
|
||||||
|
当前 case_library 的 rootCause 和 solution 还直接使用完整 answer。
|
||||||
|
后续应该从报告中结构化抽取 rootCause、solution、errorCode 和 service,提高案例复用质量。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 结尾万能总结
|
||||||
|
|
||||||
|
```text
|
||||||
|
这个项目我最想展示的不是某一个模型效果,而是 Agent 工程化能力:
|
||||||
|
一个诊断答案从哪里来、用了什么证据、是否被验证、用户是否认可、后续怎么沉淀和回归。
|
||||||
|
这些链路都被结构化记录下来,所以它可以继续演进,而不是一次性 demo。
|
||||||
|
```
|
||||||
|
|
||||||
+96
-142
@@ -1,162 +1,116 @@
|
|||||||
# 数据库设计文档
|
# SuperBizAgent MVP 文档
|
||||||
|
|
||||||
> 当前架构快照:[mvp/architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md)
|
**更新日期**:2026-07-05
|
||||||
|
|
||||||
## 📚 文档导航
|
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前架构入口已经整理到 `mvp/architecture/`,旧版架构材料已归档,避免继续把历史方案当成当前实现。
|
||||||
|
|
||||||
### 核心表设计
|
## 当前入口
|
||||||
- [diagnosis_record](tables/diagnosis_record.md) - 诊断记录表(核心)
|
|
||||||
- [case_library](tables/case_library.md) - 案例库表
|
|
||||||
- [api_document](tables/api_document.md) - 文档元数据表
|
|
||||||
|
|
||||||
### 架构设计
|
| 目录/文档 | 用途 |
|
||||||
- [Agent 架构设计](architecture/agent-architecture.md) - Agent 协作 + Skill + Harness
|
|---|---|
|
||||||
- [知识库检索架构](architecture/knowledge-retrieval-architecture.md) - L0+L1 混合检索架构 ⭐新增
|
| [architecture/README.md](architecture/README.md) | 当前 MVP 架构入口 |
|
||||||
- [知识库检索使用指南](architecture/knowledge-retrieval-usage.md) - 文档编写和使用说明 ⭐新增
|
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
||||||
- [会话管理](architecture/session-management.md) - Redis + MySQL 会话管理
|
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
|
||||||
- [实施规划](architecture/implementation-plan.md) - 分阶段实施计划
|
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
||||||
- [会话级去重与知识域地图](architecture/session-dedup-knowledge-map.md) - 文档级去重 + Planner 知识域地图注入解决 ISS-001 ⭐新增
|
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
|
||||||
- [证据评分与用户反馈](architecture/confidence-feedback.md) - evidence_score 规则引擎 + feedback API ⭐新增
|
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
|
||||||
- [行动记忆与检索归一化](architecture/action-memory-relevance.md) - Executor 行动记忆 + 归一化质量等级解决 ISS-002 ⭐新增
|
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
|
||||||
|
| [architecture/feedback-architecture.md](architecture/feedback-architecture.md) | 反馈与自评估架构 |
|
||||||
|
| [architecture/session-trace-lifecycle.md](architecture/session-trace-lifecycle.md) | 会话与 Trace 生命周期 |
|
||||||
|
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
|
||||||
|
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
|
||||||
|
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
|
||||||
|
| [issues/rag-refactor-plan.md](issues/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
|
||||||
|
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
|
||||||
|
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
|
||||||
|
| [eval/README.md](eval/README.md) | 诊断评测材料 |
|
||||||
|
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
||||||
|
|
||||||
---
|
## 当前系统一句话
|
||||||
|
|
||||||
## 一、设计原则
|
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据,诊断过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
||||||
|
|
||||||
### 1.1 核心原则
|
## 文档结构
|
||||||
- ✅ **简单优先**:满足诊断流程需要,避免过度设计
|
|
||||||
- ✅ **渐进增强**:先实现核心功能,再逐步扩展
|
|
||||||
- ✅ **数据分离**:诊断结果持久化(MySQL),会话上下文临时化(Redis)
|
|
||||||
- ✅ **适度冗余**:避免过度范式化,适当冗余提升查询性能
|
|
||||||
|
|
||||||
### 1.2 系统定位
|
```text
|
||||||
**自动化诊断系统**
|
mvp/
|
||||||
- 核心:一键诊断 → 返回完整报告
|
architecture/
|
||||||
- 辅助:支持追问,但不是主要场景
|
README.md
|
||||||
- 特点:大部分用户单次诊断即结束,少数用户会追问细节
|
current-mvp-architecture.md
|
||||||
|
interview-one-pager.md
|
||||||
---
|
agent-orchestration.md
|
||||||
|
harness-quality-gates.md
|
||||||
## 二、表结构总览
|
rag-architecture.md
|
||||||
|
retrieval-observability.md
|
||||||
### 2.1 核心表关系
|
feedback-architecture.md
|
||||||
|
session-trace-lifecycle.md
|
||||||
```
|
knowledge-base-authoring.md
|
||||||
┌─────────────────────┐
|
data-model.md
|
||||||
│ diagnosis_record │ 诊断记录(核心)
|
evolution-roadmap.md
|
||||||
│ - 每次诊断一条 │
|
archive/2026-07-05-legacy/
|
||||||
└──────────┬──────────┘
|
issues/
|
||||||
│ 1:1
|
README.md
|
||||||
↓
|
rag-refactor-plan.md
|
||||||
┌─────────────────────┐
|
ISS-*.md
|
||||||
│ case_library │ 案例库(知识沉淀)
|
rag-*.md
|
||||||
│ - 诊断成功→案例 │
|
demo/
|
||||||
└─────────────────────┘
|
README.md
|
||||||
|
ten-minute-interview-demo.md
|
||||||
┌─────────────────────┐
|
requests/
|
||||||
│ api_document │ 文档元数据(管理层)
|
scripts/
|
||||||
│ - 状态追踪/去重 │
|
output/
|
||||||
└──────────┬──────────┘
|
eval/
|
||||||
│ doc_id
|
README.md
|
||||||
↓
|
schema.md
|
||||||
┌─────────────────────┐
|
cases/
|
||||||
│ Milvus │ 文档内容(检索层)
|
fixtures/
|
||||||
│ - 向量检索 │
|
reports/
|
||||||
└─────────────────────┘
|
notes/
|
||||||
|
plan/
|
||||||
┌─────────────────────┐
|
tables/
|
||||||
│ Redis Session │ 会话管理(临时)
|
|
||||||
│ - 30分钟过期 │
|
|
||||||
│ - 支持追问 │
|
|
||||||
└─────────────────────┘
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### 2.2 表统计
|
## 当前核心设计
|
||||||
|
|
||||||
| 表名 | 类型 | 预估数据量 | 用途 |
|
- `lookup_knowledge` 保持显式 Agent Tool,不隐藏到 Chat Advisor。
|
||||||
|------|------|-----------|------|
|
- L0 降级为 domain/entity hint,不再默认承担最终召回决策。
|
||||||
| diagnosis_record | 核心 | 3.6万/年 | 诊断记录 |
|
- `VectorSearchService` 是检索稳定门面。
|
||||||
| case_library | 核心 | 500-1000 | 案例库 |
|
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
|
||||||
| api_document | 核心 | 100-200 | 文档管理 |
|
- AIOps payload 会生成推荐知识库 query,保留业务语义。
|
||||||
|
- Trace API 聚合 session、step、tool invocation 和 self evaluation。
|
||||||
|
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
|
||||||
|
|
||||||
---
|
## 关键运行链路
|
||||||
|
|
||||||
## 三、技术栈
|
```text
|
||||||
|
Chat
|
||||||
|
-> ChatService
|
||||||
|
-> Planner / Executor / Verifier
|
||||||
|
-> evidence tools
|
||||||
|
-> diagnosis_session / agent_step / tool_invocation
|
||||||
|
-> DiagnosisTraceService
|
||||||
|
|
||||||
### 3.1 数据存储
|
AIOps
|
||||||
```
|
-> AiOpsService
|
||||||
MySQL 8.0+
|
-> PAYLOAD_TARGETED or AUTO_DISCOVERY
|
||||||
├─ 元数据管理
|
-> Planner / Executor
|
||||||
├─ 事务支持
|
-> Prometheus / logs / lookup_knowledge
|
||||||
└─ JSON 字段支持
|
-> AiOpsRuleEvaluationService
|
||||||
|
-> DiagnosisTraceService
|
||||||
|
|
||||||
Redis 6.0+
|
RAG
|
||||||
├─ 会话存储
|
-> lookup_knowledge
|
||||||
├─ 缓存
|
-> L0 domain/entity hint
|
||||||
└─ TTL 自动过期
|
-> VectorSearchService
|
||||||
|
-> Spring AI VectorStore / Milvus SDK fallback
|
||||||
Milvus 2.6+
|
-> relevance normalization
|
||||||
├─ 向量存储
|
-> tool_invocation
|
||||||
├─ 语义检索
|
|
||||||
└─ 混合检索
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### 3.2 开发框架
|
## 旧文档说明
|
||||||
```
|
|
||||||
Spring Boot 3.2
|
|
||||||
Spring AI Alibaba 1.1.0
|
|
||||||
Milvus SDK Java 2.6.10
|
|
||||||
DashScope SDK
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
旧版架构文档已移动到:
|
||||||
|
|
||||||
## 四、快速开始
|
- [architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
||||||
|
|
||||||
### 4.1 创建数据库
|
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/current-mvp-architecture.md` 与 `architecture/rag-architecture.md` 为准。
|
||||||
|
|
||||||
```sql
|
|
||||||
-- 1. 创建数据库
|
|
||||||
CREATE DATABASE diagnosis_system CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;
|
|
||||||
|
|
||||||
-- 2. 执行建表脚本(按顺序)
|
|
||||||
SOURCE tables/diagnosis_record.sql;
|
|
||||||
SOURCE tables/case_library.sql;
|
|
||||||
SOURCE tables/api_document.sql;
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4.2 初始化 Milvus
|
|
||||||
|
|
||||||
```java
|
|
||||||
// 创建 Collection
|
|
||||||
MilvusClientFactory.createCollection();
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4.3 配置 Redis
|
|
||||||
|
|
||||||
```yaml
|
|
||||||
spring:
|
|
||||||
redis:
|
|
||||||
host: localhost
|
|
||||||
port: 6379
|
|
||||||
database: 0
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 五、版本历史
|
|
||||||
|
|
||||||
| 版本 | 日期 | 变更内容 |
|
|
||||||
|------|------|---------|
|
|
||||||
| v1.0 | 2024-06-15 | 初版,定义核心表结构 |
|
|
||||||
| v2.0 | 2024-06-15 | diagnosis_record 字段泛化,支持多种故障类型 |
|
|
||||||
| v2.1 | 2024-06-22 | 文档拆分,增加 api_document 表 |
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 六、维护说明
|
|
||||||
|
|
||||||
- 每个表的详细设计在 `tables/` 目录下
|
|
||||||
- 架构设计文档在 `architecture/` 目录下
|
|
||||||
- 修改表结构时,同步更新对应的 Markdown 文档
|
|
||||||
- 重大变更需记录在版本历史中
|
|
||||||
|
|||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# MVP 架构文档
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
|
||||||
|
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||||
|
|
||||||
|
- `mvp/architecture/archive/2026-07-05-legacy/`
|
||||||
|
|
||||||
|
归档材料只作为设计历史阅读,不再作为当前实现依据。
|
||||||
|
|
||||||
|
## 当前文档
|
||||||
|
|
||||||
|
| 文档 | 用途 |
|
||||||
|
|---|---|
|
||||||
|
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
|
||||||
|
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
|
||||||
|
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
||||||
|
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
|
||||||
|
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||||
|
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
|
||||||
|
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
|
||||||
|
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
|
||||||
|
| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库文档编写与维护规范,覆盖 frontmatter、category、chunk、reindex |
|
||||||
|
| [data-model.md](data-model.md) | 数据模型总览,覆盖 Trace、知识库、反馈沉淀和 Milvus metadata |
|
||||||
|
| [evolution-roadmap.md](evolution-roadmap.md) | 从旧版 Agent 蓝图继承的后续演进路线,不代表当前已实现 |
|
||||||
|
|
||||||
|
## 当前架构一句话
|
||||||
|
|
||||||
|
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||||
|
|
||||||
|
## 阅读顺序
|
||||||
|
|
||||||
|
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
|
||||||
|
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
|
||||||
|
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
||||||
|
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||||
|
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||||
|
6. 继续读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||||
|
7. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||||
|
8. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||||
|
9. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||||
|
10. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||||
@@ -0,0 +1,187 @@
|
|||||||
|
# Agent 编排架构
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:当前可运行架构
|
||||||
|
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||||
|
|
||||||
|
## 1. 设计定位
|
||||||
|
|
||||||
|
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
||||||
|
|
||||||
|
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Verifier`。
|
||||||
|
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
||||||
|
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
||||||
|
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
||||||
|
|
||||||
|
## 2. 当前 Agent 全景
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
subgraph Chat["Chat diagnosis"]
|
||||||
|
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
|
||||||
|
ChatService --> ChatPlanner["chat_planner"]
|
||||||
|
ChatPlanner --> ChatExecutor["chat_executor"]
|
||||||
|
ChatExecutor --> ChatTools["evidence tools"]
|
||||||
|
ChatTools --> ChatExecutor
|
||||||
|
ChatExecutor --> ChatVerifier["chat_verifier"]
|
||||||
|
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
||||||
|
ChatDecision --> ChatAnswer["final answer"]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph AiOps["AIOps diagnosis"]
|
||||||
|
AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
|
||||||
|
AiOpsService --> Supervisor["ai_ops_supervisor"]
|
||||||
|
Supervisor --> AiOpsPlanner["planner_agent"]
|
||||||
|
Supervisor --> AiOpsExecutor["executor_agent"]
|
||||||
|
AiOpsPlanner --> AiOpsExecutor
|
||||||
|
AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
|
||||||
|
AiOpsTools --> AiOpsReport["alert report"]
|
||||||
|
AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph Trace["Trace persistence"]
|
||||||
|
Session["diagnosis_session"]
|
||||||
|
Step["agent_step"]
|
||||||
|
Invocation["tool_invocation"]
|
||||||
|
SelfEval["self_evaluation"]
|
||||||
|
end
|
||||||
|
|
||||||
|
ChatService --> Session
|
||||||
|
ChatPlanner --> Step
|
||||||
|
ChatExecutor --> Step
|
||||||
|
ChatVerifier --> Step
|
||||||
|
ChatTools --> Invocation
|
||||||
|
ChatDecision --> SelfEval
|
||||||
|
|
||||||
|
AiOpsService --> Session
|
||||||
|
AiOpsPlanner --> Step
|
||||||
|
AiOpsExecutor --> Step
|
||||||
|
AiOpsTools --> Invocation
|
||||||
|
AiOpsRule --> SelfEval
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. Chat 编排
|
||||||
|
|
||||||
|
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
|
||||||
|
|
||||||
|
```text
|
||||||
|
chat_planner
|
||||||
|
-> chat_executor
|
||||||
|
-> lookup_knowledge / query_logs / query_metrics / date_time
|
||||||
|
-> chat_verifier
|
||||||
|
-> reads tool_trace_summary
|
||||||
|
-> outputs verifier JSON
|
||||||
|
```
|
||||||
|
|
||||||
|
关键行为:
|
||||||
|
|
||||||
|
| 角色 | 当前职责 | 输出 |
|
||||||
|
|---|---|---|
|
||||||
|
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
||||||
|
| `chat_executor` | 按计划调用证据工具,组合工具返回形成诊断答复 | `executor_feedback` |
|
||||||
|
| `chat_verifier` | 只基于已有证据校验 Executor 答案,不做新检索 | `verifier_output` |
|
||||||
|
|
||||||
|
Chat 链路最多支持两轮验证:
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
sequenceDiagram
|
||||||
|
autonumber
|
||||||
|
participant C as ChatService
|
||||||
|
participant P as chat_planner
|
||||||
|
participant E as chat_executor
|
||||||
|
participant T as tools
|
||||||
|
participant V as chat_verifier
|
||||||
|
participant S as diagnosis_session
|
||||||
|
|
||||||
|
C->>P: 原始问题 + history + retry_context
|
||||||
|
P-->>C: planner_plan
|
||||||
|
C->>E: planner_plan + 上下文
|
||||||
|
E->>T: 调用证据工具
|
||||||
|
T-->>E: 证据结果
|
||||||
|
E-->>C: executor_feedback
|
||||||
|
C->>V: executor_final_answer + tool_trace_summary
|
||||||
|
V-->>C: PASS / LOW_CONFID / REJECT
|
||||||
|
C->>S: 写入 verifier_evaluation
|
||||||
|
alt LOW_CONFID 且允许补证据
|
||||||
|
C->>P: retry_context: 仅补缺失证据
|
||||||
|
else PASS 或 REJECT
|
||||||
|
C-->>S: 保存最终 answer
|
||||||
|
end
|
||||||
|
```
|
||||||
|
|
||||||
|
决策语义:
|
||||||
|
|
||||||
|
| Verdict | 行为 |
|
||||||
|
|---|---|
|
||||||
|
| `PASS` | 输出 Executor 答案 |
|
||||||
|
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
|
||||||
|
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
|
||||||
|
|
||||||
|
## 4. AIOps 编排
|
||||||
|
|
||||||
|
AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
|
||||||
|
|
||||||
|
```text
|
||||||
|
ai_ops_supervisor
|
||||||
|
-> planner_agent
|
||||||
|
-> executor_agent
|
||||||
|
-> final report
|
||||||
|
-> AiOpsRuleEvaluationService
|
||||||
|
```
|
||||||
|
|
||||||
|
与 Chat 的差异:
|
||||||
|
|
||||||
|
- AIOps 的输入可能是结构化告警 payload。
|
||||||
|
- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
|
||||||
|
- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
|
||||||
|
- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
|
||||||
|
|
||||||
|
## 5. 工具边界
|
||||||
|
|
||||||
|
当前 Executor 可用工具来自两类:
|
||||||
|
|
||||||
|
```text
|
||||||
|
methodTools
|
||||||
|
-> dateTimeTools
|
||||||
|
-> lookupKnowledgeTool
|
||||||
|
-> queryMetricsTools
|
||||||
|
-> queryLogsTools when mock enabled
|
||||||
|
|
||||||
|
ToolCallbackProvider
|
||||||
|
-> framework-discovered tools
|
||||||
|
```
|
||||||
|
|
||||||
|
工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
|
||||||
|
|
||||||
|
- L0/L1 命中数量。
|
||||||
|
- 检索层。
|
||||||
|
- relevance level。
|
||||||
|
- retrieved domains。
|
||||||
|
- dedup reason。
|
||||||
|
|
||||||
|
## 6. 与旧版设计的差异
|
||||||
|
|
||||||
|
| 旧版设想 | 当前实现 |
|
||||||
|
|---|---|
|
||||||
|
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Verifier;AIOps: Supervisor + Planner + Executor |
|
||||||
|
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
||||||
|
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
||||||
|
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
||||||
|
| Skill 驱动不同诊断流程 | 当前以 Prompt、知识域地图、工具调用和评测 baseline 控制 |
|
||||||
|
|
||||||
|
## 7. 后续演进
|
||||||
|
|
||||||
|
当诊断场景和工具复杂度继续上升时,再考虑拆分:
|
||||||
|
|
||||||
|
- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
|
||||||
|
- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
|
||||||
|
- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
|
||||||
|
- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
|
||||||
|
|
||||||
|
拆分前提:
|
||||||
|
|
||||||
|
- 当前 Executor prompt 已难以维护。
|
||||||
|
- 不同故障类型的工具权限明显不同。
|
||||||
|
- Trace 能证明某类问题需要独立的推理策略。
|
||||||
|
- 评测集能覆盖拆分前后的行为差异。
|
||||||
|
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
# 旧版架构文档归档
|
||||||
|
|
||||||
|
**归档日期**:2026-07-05
|
||||||
|
|
||||||
|
本目录保存 `mvp/architecture` 下的旧版架构文档。它们包含早期 MVP 设计、旧 RAG 方案、会话存储设计、行动记忆和实施计划等历史材料。
|
||||||
|
|
||||||
|
这些文档不再作为当前实现依据。当前架构请阅读:
|
||||||
|
|
||||||
|
- `mvp/architecture/README.md`
|
||||||
|
- `mvp/architecture/current-mvp-architecture.md`
|
||||||
|
- `mvp/architecture/rag-architecture.md`
|
||||||
|
|
||||||
|
## 归档文件
|
||||||
|
|
||||||
|
| 文件 | 说明 |
|
||||||
|
|---|---|
|
||||||
|
| `agent-architecture.md` | 早期完整 Agent 设想,包含较多超出当前 MVP 的 SubAgent 设计 |
|
||||||
|
| `agent-architecture-mvp.md` | 早期 MVP Agent 设计 |
|
||||||
|
| `knowledge-retrieval-architecture.md` | 旧版 L0 + L1 检索架构,包含 L0 唯一命中跳过 L1 的旧逻辑 |
|
||||||
|
| `knowledge-retrieval-usage.md` | 旧版知识库检索使用说明 |
|
||||||
|
| `current-mvp-architecture.md` | 归档前的当前架构快照 |
|
||||||
|
| `implementation-plan.md` | 早期实施计划 |
|
||||||
|
| `implementation-detail.md` | 早期完整实施计划 |
|
||||||
|
| `session-management.md` | 会话管理旧设计 |
|
||||||
|
| `session-dedup-knowledge-map.md` | 会话去重和知识域地图设计 |
|
||||||
|
| `confidence-feedback.md` | 证据评分和用户反馈旧设计 |
|
||||||
|
| `action-memory-relevance.md` | 行动记忆和检索质量归一化旧设计 |
|
||||||
|
|
||||||
@@ -0,0 +1,220 @@
|
|||||||
|
# Current MVP Architecture Snapshot
|
||||||
|
|
||||||
|
**Updated**: 2026-07-05
|
||||||
|
|
||||||
|
This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.
|
||||||
|
|
||||||
|
## 1. Positioning
|
||||||
|
|
||||||
|
The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.
|
||||||
|
|
||||||
|
Core goals:
|
||||||
|
|
||||||
|
- Support normal chat-based diagnosis.
|
||||||
|
- Support AIOps alert-triggered diagnosis.
|
||||||
|
- Keep tool calls explicit and traceable.
|
||||||
|
- Keep RAG retrieval observable through `lookup_knowledge`.
|
||||||
|
- Persist enough execution evidence for replay, evaluation, and interview explanation.
|
||||||
|
|
||||||
|
## 2. Runtime Architecture
|
||||||
|
|
||||||
|
```text
|
||||||
|
HTTP API
|
||||||
|
-> ChatService / AiOpsService
|
||||||
|
-> Agent orchestration
|
||||||
|
-> Supervisor / Planner / Executor / Verifier
|
||||||
|
-> Tools
|
||||||
|
-> lookup_knowledge
|
||||||
|
-> query_logs
|
||||||
|
-> query_metrics
|
||||||
|
-> other diagnosis tools
|
||||||
|
-> Persistence
|
||||||
|
-> diagnosis_session
|
||||||
|
-> agent_step
|
||||||
|
-> tool_invocation
|
||||||
|
-> Trace API
|
||||||
|
-> DiagnosisTraceService
|
||||||
|
```
|
||||||
|
|
||||||
|
Current entry points:
|
||||||
|
|
||||||
|
- `ChatService`: user-driven troubleshooting and follow-up diagnosis.
|
||||||
|
- `AiOpsService`: alert-driven diagnosis, including payload mode and auto-discovery mode.
|
||||||
|
- `DiagnosisTraceService`: trace view of session, steps, tool calls, and self-evaluation.
|
||||||
|
|
||||||
|
## 3. Chat Diagnosis Flow
|
||||||
|
|
||||||
|
```text
|
||||||
|
User question
|
||||||
|
-> ChatService
|
||||||
|
-> simple response or diagnosis flow
|
||||||
|
-> Planner creates investigation direction
|
||||||
|
-> Executor calls tools for evidence
|
||||||
|
-> lookup_knowledge
|
||||||
|
-> query_logs
|
||||||
|
-> query_metrics
|
||||||
|
-> Verifier checks final diagnosis quality
|
||||||
|
-> self_evaluation.verifier_evaluation
|
||||||
|
-> diagnosis trace
|
||||||
|
```
|
||||||
|
|
||||||
|
The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||||
|
|
||||||
|
## 4. AIOps Diagnosis Flow
|
||||||
|
|
||||||
|
```text
|
||||||
|
AIOps request
|
||||||
|
-> AiOpsService
|
||||||
|
-> payload mode or auto-discovery mode
|
||||||
|
-> build alert-focused diagnosis prompt
|
||||||
|
-> append recommended lookup_knowledge query when payload exists
|
||||||
|
-> Agent diagnosis flow
|
||||||
|
-> Supervisor / Planner / Executor
|
||||||
|
-> evidence tools
|
||||||
|
-> final report
|
||||||
|
-> AiOpsRuleEvaluationService
|
||||||
|
-> self_evaluation.aiops_rule_evaluation
|
||||||
|
-> diagnosis trace
|
||||||
|
```
|
||||||
|
|
||||||
|
AIOps keeps two modes:
|
||||||
|
|
||||||
|
- Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
|
||||||
|
- Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.
|
||||||
|
|
||||||
|
The AIOps verifier is currently lightweight and rule-based. It checks:
|
||||||
|
|
||||||
|
- Whether the final report exists.
|
||||||
|
- Whether the result stays focused on the alert payload when payload exists.
|
||||||
|
- Whether evidence tools were used, especially `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||||
|
|
||||||
|
## 5. RAG Architecture
|
||||||
|
|
||||||
|
```text
|
||||||
|
lookup_knowledge
|
||||||
|
-> L0 domain/entity hint
|
||||||
|
-> matched domain
|
||||||
|
-> matched keywords/entities
|
||||||
|
-> metadata filter signal
|
||||||
|
-> VectorSearchService
|
||||||
|
-> Spring AI VectorStore path
|
||||||
|
-> Milvus SDK fallback path
|
||||||
|
-> evidence post-processing
|
||||||
|
-> score / rawScore / scoreLabel
|
||||||
|
-> source metadata
|
||||||
|
-> title / breadcrumb / content evidence block
|
||||||
|
-> tool_invocation record
|
||||||
|
```
|
||||||
|
|
||||||
|
Important decisions:
|
||||||
|
|
||||||
|
- `lookup_knowledge` remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
|
||||||
|
- L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
|
||||||
|
- L1 retrieval now goes through `VectorSearchService`.
|
||||||
|
- Spring AI `VectorStore` is the preferred retrieval path.
|
||||||
|
- The original Milvus SDK path is retained as fallback and compatibility path.
|
||||||
|
- `title`, `breadcrumb`, and `content` participate in embedding text so chunk context is less likely to be lost.
|
||||||
|
- Retrieval output keeps compatibility fields: `score`, `rawScore`, and `scoreLabel`.
|
||||||
|
|
||||||
|
Vector retrieval modes:
|
||||||
|
|
||||||
|
```text
|
||||||
|
retrieval.vector-store.mode=auto # Prefer Spring AI VectorStore, fallback to SDK
|
||||||
|
retrieval.vector-store.mode=spring-ai # Use Spring AI VectorStore only
|
||||||
|
retrieval.vector-store.mode=sdk # Use original Milvus SDK path
|
||||||
|
```
|
||||||
|
|
||||||
|
## 6. Persistence And Trace
|
||||||
|
|
||||||
|
Current trace-related persistence:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session
|
||||||
|
-> final_report
|
||||||
|
-> self_evaluation
|
||||||
|
-> verifier_evaluation
|
||||||
|
-> aiops_rule_evaluation
|
||||||
|
|
||||||
|
agent_step
|
||||||
|
-> role
|
||||||
|
-> step input/output
|
||||||
|
-> execution order
|
||||||
|
|
||||||
|
tool_invocation
|
||||||
|
-> tool_name
|
||||||
|
-> query
|
||||||
|
-> retrieval_layer
|
||||||
|
-> retrieval_details
|
||||||
|
-> evidence blocks
|
||||||
|
-> duration
|
||||||
|
```
|
||||||
|
|
||||||
|
Trace API aggregates these records into a session-level view:
|
||||||
|
|
||||||
|
- Agent step sequence.
|
||||||
|
- Tool calls and retrieval details.
|
||||||
|
- Final diagnosis report.
|
||||||
|
- Chat verifier status.
|
||||||
|
- AIOps rule verifier status.
|
||||||
|
|
||||||
|
## 7. Quality Gates
|
||||||
|
|
||||||
|
Current quality gates:
|
||||||
|
|
||||||
|
- Chat verifier: LLM-based final answer verification for normal diagnosis.
|
||||||
|
- AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
|
||||||
|
- Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
|
||||||
|
- RAG retrieval baseline: golden query set with offline baseline report.
|
||||||
|
- Live RAG acceptance: post-reindex script for validating retrieval against the running stack.
|
||||||
|
|
||||||
|
These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.
|
||||||
|
|
||||||
|
## 8. Current Completion State
|
||||||
|
|
||||||
|
Completed for the current MVP stage:
|
||||||
|
|
||||||
|
- Explicit `lookup_knowledge` Agent tool.
|
||||||
|
- L0 + L1 retrieval shape retained.
|
||||||
|
- L0 downgraded to domain/entity hint.
|
||||||
|
- Spring AI VectorStore retrieval path integrated.
|
||||||
|
- Milvus SDK fallback retained.
|
||||||
|
- RAG evidence post-processing added.
|
||||||
|
- Breadcrumb/title/content embedding text improved.
|
||||||
|
- RAG offline baseline and live acceptance script added.
|
||||||
|
- AIOps payload query augmentation added.
|
||||||
|
- AIOps lightweight verifier added.
|
||||||
|
- Trace summary includes both chat verifier and AIOps verifier signals.
|
||||||
|
|
||||||
|
Deferred future enhancements:
|
||||||
|
|
||||||
|
- LLM QueryTransformer / MultiQuery.
|
||||||
|
- BM25, RRF, and reranker.
|
||||||
|
- Neighbor chunk or section-level context expansion.
|
||||||
|
- VectorStore write path migration.
|
||||||
|
- Full LLM-based AIOps verifier.
|
||||||
|
- More complete golden set for recall, MRR, and nDCG metrics.
|
||||||
|
|
||||||
|
## 9. Key Code References
|
||||||
|
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||||
|
|
||||||
|
## 10. Supporting Materials
|
||||||
|
|
||||||
|
- `mvp/issues/rag-refactor-plan.md`
|
||||||
|
- `eval/rag-retrieval/README.md`
|
||||||
|
- `scripts/eval_rag_live_acceptance.py`
|
||||||
|
- `interview/rag-refactor-story.md`
|
||||||
|
- `interview/rag-vectorstore-interview-notes.md`
|
||||||
|
- `interview/rag-retrieval-quality-report.md`
|
||||||
|
- `interview/rag-breadcrumb-embedding-acceptance.md`
|
||||||
|
- `interview/aiops-query-augmentation.md`
|
||||||
|
- `interview/aiops-lightweight-verifier.md`
|
||||||
@@ -1,220 +1,391 @@
|
|||||||
# Current MVP Architecture Snapshot
|
# 当前 MVP 架构
|
||||||
|
|
||||||
**Updated**: 2026-07-05
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:当前可运行架构
|
||||||
|
**适用范围**:Demo、面试讲解、后续迭代规划
|
||||||
|
|
||||||
This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.
|
## 1. 系统定位
|
||||||
|
|
||||||
## 1. Positioning
|
SuperBizAgent MVP 不是通用 Chatbot,而是面向故障诊断的 Agent 工程项目。
|
||||||
|
|
||||||
The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.
|
核心目标:
|
||||||
|
|
||||||
Core goals:
|
- 支持用户主动发起的 Chat 诊断。
|
||||||
|
- 支持 AIOps 告警触发的自动诊断。
|
||||||
|
- 保留 Agent 的规划、执行、验证过程。
|
||||||
|
- 工具调用必须显式、可追踪、可回放。
|
||||||
|
- RAG 检索必须通过 `lookup_knowledge` 暴露证据链。
|
||||||
|
- 每次诊断都沉淀 session、step、tool invocation 和 self evaluation。
|
||||||
|
|
||||||
- Support normal chat-based diagnosis.
|
## 2. 总体分层
|
||||||
- Support AIOps alert-triggered diagnosis.
|
|
||||||
- Keep tool calls explicit and traceable.
|
|
||||||
- Keep RAG retrieval observable through `lookup_knowledge`.
|
|
||||||
- Persist enough execution evidence for replay, evaluation, and interview explanation.
|
|
||||||
|
|
||||||
## 2. Runtime Architecture
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
subgraph API["API Layer"]
|
||||||
|
ChatController["ChatController"]
|
||||||
|
TraceController["DiagnosisTraceController"]
|
||||||
|
SearchController["SearchController"]
|
||||||
|
DocumentController["DocumentController"]
|
||||||
|
end
|
||||||
|
|
||||||
```text
|
subgraph App["Application Service"]
|
||||||
HTTP API
|
ChatService["ChatService"]
|
||||||
-> ChatService / AiOpsService
|
AiOpsService["AiOpsService"]
|
||||||
-> Agent orchestration
|
TraceService["DiagnosisTraceService"]
|
||||||
-> Supervisor / Planner / Executor / Verifier
|
end
|
||||||
-> Tools
|
|
||||||
-> lookup_knowledge
|
subgraph Agent["Agent Orchestration"]
|
||||||
-> query_logs
|
Supervisor["Supervisor"]
|
||||||
-> query_metrics
|
Planner["Planner"]
|
||||||
-> other diagnosis tools
|
Executor["Executor"]
|
||||||
-> Persistence
|
Verifier["Verifier"]
|
||||||
-> diagnosis_session
|
end
|
||||||
-> agent_step
|
|
||||||
-> tool_invocation
|
subgraph Tools["Evidence Tools"]
|
||||||
-> Trace API
|
KnowledgeTool["lookup_knowledge"]
|
||||||
-> DiagnosisTraceService
|
LogsTool["query_logs"]
|
||||||
|
MetricsTool["query_metrics"]
|
||||||
|
AlertsTool["queryPrometheusAlerts"]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph RAG["RAG Retrieval"]
|
||||||
|
L0["KnowledgeIndexService"]
|
||||||
|
VectorSearch["VectorSearchService"]
|
||||||
|
VectorStore["Spring AI VectorStore"]
|
||||||
|
SdkFallback["Milvus SDK fallback"]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph Store["Persistence and Trace"]
|
||||||
|
Session["diagnosis_session"]
|
||||||
|
Step["agent_step"]
|
||||||
|
Invocation["tool_invocation"]
|
||||||
|
ApiDoc["api_document"]
|
||||||
|
Milvus["Milvus/Zilliz"]
|
||||||
|
end
|
||||||
|
|
||||||
|
API --> App
|
||||||
|
ChatService --> Agent
|
||||||
|
AiOpsService --> Agent
|
||||||
|
Agent --> Tools
|
||||||
|
KnowledgeTool --> RAG
|
||||||
|
RAG --> Store
|
||||||
|
Tools --> Invocation
|
||||||
|
Agent --> Step
|
||||||
|
App --> Session
|
||||||
|
TraceService --> Session
|
||||||
|
TraceService --> Step
|
||||||
|
TraceService --> Invocation
|
||||||
```
|
```
|
||||||
|
|
||||||
Current entry points:
|
|
||||||
|
|
||||||
- `ChatService`: user-driven troubleshooting and follow-up diagnosis.
|
|
||||||
- `AiOpsService`: alert-driven diagnosis, including payload mode and auto-discovery mode.
|
|
||||||
- `DiagnosisTraceService`: trace view of session, steps, tool calls, and self-evaluation.
|
|
||||||
|
|
||||||
## 3. Chat Diagnosis Flow
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
User question
|
API Layer
|
||||||
|
-> ChatController
|
||||||
|
-> DiagnosisTraceController
|
||||||
|
-> SearchController
|
||||||
|
-> DocumentController
|
||||||
|
|
||||||
|
Application Service
|
||||||
-> ChatService
|
-> ChatService
|
||||||
-> simple response or diagnosis flow
|
|
||||||
-> Planner creates investigation direction
|
|
||||||
-> Executor calls tools for evidence
|
|
||||||
-> lookup_knowledge
|
|
||||||
-> query_logs
|
|
||||||
-> query_metrics
|
|
||||||
-> Verifier checks final diagnosis quality
|
|
||||||
-> self_evaluation.verifier_evaluation
|
|
||||||
-> diagnosis trace
|
|
||||||
```
|
|
||||||
|
|
||||||
The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
|
||||||
|
|
||||||
## 4. AIOps Diagnosis Flow
|
|
||||||
|
|
||||||
```text
|
|
||||||
AIOps request
|
|
||||||
-> AiOpsService
|
-> AiOpsService
|
||||||
-> payload mode or auto-discovery mode
|
-> DiagnosisTraceService
|
||||||
-> build alert-focused diagnosis prompt
|
|
||||||
-> append recommended lookup_knowledge query when payload exists
|
|
||||||
-> Agent diagnosis flow
|
|
||||||
-> Supervisor / Planner / Executor
|
|
||||||
-> evidence tools
|
|
||||||
-> final report
|
|
||||||
-> AiOpsRuleEvaluationService
|
|
||||||
-> self_evaluation.aiops_rule_evaluation
|
|
||||||
-> diagnosis trace
|
|
||||||
```
|
|
||||||
|
|
||||||
AIOps keeps two modes:
|
Agent Orchestration
|
||||||
|
-> Supervisor
|
||||||
|
-> Planner
|
||||||
|
-> Executor
|
||||||
|
-> Verifier
|
||||||
|
|
||||||
- Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
|
Evidence Tools
|
||||||
- Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.
|
-> lookup_knowledge
|
||||||
|
-> query_logs
|
||||||
|
-> query_metrics
|
||||||
|
-> queryPrometheusAlerts
|
||||||
|
|
||||||
The AIOps verifier is currently lightweight and rule-based. It checks:
|
RAG Retrieval
|
||||||
|
-> KnowledgeIndexService
|
||||||
- Whether the final report exists.
|
|
||||||
- Whether the result stays focused on the alert payload when payload exists.
|
|
||||||
- Whether evidence tools were used, especially `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
|
||||||
|
|
||||||
## 5. RAG Architecture
|
|
||||||
|
|
||||||
```text
|
|
||||||
lookup_knowledge
|
|
||||||
-> L0 domain/entity hint
|
|
||||||
-> matched domain
|
|
||||||
-> matched keywords/entities
|
|
||||||
-> metadata filter signal
|
|
||||||
-> VectorSearchService
|
-> VectorSearchService
|
||||||
-> Spring AI VectorStore path
|
-> Spring AI VectorStore
|
||||||
-> Milvus SDK fallback path
|
-> Milvus SDK fallback
|
||||||
-> evidence post-processing
|
|
||||||
-> score / rawScore / scoreLabel
|
Persistence
|
||||||
-> source metadata
|
-> diagnosis_session
|
||||||
-> title / breadcrumb / content evidence block
|
-> agent_step
|
||||||
-> tool_invocation record
|
-> tool_invocation
|
||||||
|
-> api_document
|
||||||
|
-> Milvus/Zilliz collection
|
||||||
|
|
||||||
|
Quality Gates
|
||||||
|
-> chat verifier
|
||||||
|
-> AIOps rule evaluation
|
||||||
|
-> diagnosis eval baseline
|
||||||
|
-> RAG retrieval baseline
|
||||||
```
|
```
|
||||||
|
|
||||||
Important decisions:
|
## 3. Chat 诊断链路
|
||||||
|
|
||||||
- `lookup_knowledge` remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
|
```mermaid
|
||||||
- L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
|
sequenceDiagram
|
||||||
- L1 retrieval now goes through `VectorSearchService`.
|
autonumber
|
||||||
- Spring AI `VectorStore` is the preferred retrieval path.
|
actor User as 用户
|
||||||
- The original Milvus SDK path is retained as fallback and compatibility path.
|
participant API as POST /api/chat
|
||||||
- `title`, `breadcrumb`, and `content` participate in embedding text so chunk context is less likely to be lost.
|
participant Chat as ChatService
|
||||||
- Retrieval output keeps compatibility fields: `score`, `rawScore`, and `scoreLabel`.
|
participant Planner as Planner Agent
|
||||||
|
participant Executor as Executor Agent
|
||||||
|
participant Tool as Evidence Tools
|
||||||
|
participant Verifier as Verifier Agent
|
||||||
|
participant DB as Trace Tables
|
||||||
|
participant Trace as Trace API
|
||||||
|
|
||||||
Vector retrieval modes:
|
User->>API: 提交诊断问题
|
||||||
|
API->>Chat: execute chat strategy
|
||||||
|
Chat->>Planner: 复杂问题进入规划
|
||||||
|
Planner->>DB: 写入 agent_step
|
||||||
|
Planner->>Executor: 下发排查方向
|
||||||
|
Executor->>Tool: lookup_knowledge / logs / metrics
|
||||||
|
Tool->>DB: 写入 tool_invocation
|
||||||
|
Tool-->>Executor: 返回证据
|
||||||
|
Executor->>Verifier: 生成候选诊断并校验
|
||||||
|
Verifier->>DB: 合并 self_evaluation.verifier_evaluation
|
||||||
|
Chat->>DB: 保存 diagnosis_session.answer
|
||||||
|
User->>Trace: GET /api/diagnosis/{sessionId}/trace
|
||||||
|
Trace->>DB: 聚合 session / step / tool
|
||||||
|
Trace-->>User: 返回可回放诊断链路
|
||||||
|
```
|
||||||
|
|
||||||
```text
|
```text
|
||||||
retrieval.vector-store.mode=auto # Prefer Spring AI VectorStore, fallback to SDK
|
POST /api/chat
|
||||||
retrieval.vector-store.mode=spring-ai # Use Spring AI VectorStore only
|
-> ChatService
|
||||||
retrieval.vector-store.mode=sdk # Use original Milvus SDK path
|
-> 简单问题:轻量回答
|
||||||
|
-> 复杂诊断:Agent 编排
|
||||||
|
-> Planner 制定排查方向
|
||||||
|
-> Executor 调用证据工具
|
||||||
|
-> lookup_knowledge
|
||||||
|
-> query_logs
|
||||||
|
-> query_metrics
|
||||||
|
-> Verifier 校验最终诊断
|
||||||
|
-> 保存 diagnosis_session
|
||||||
|
-> 保存 agent_step
|
||||||
|
-> 保存 tool_invocation
|
||||||
|
-> 合并 self_evaluation.verifier_evaluation
|
||||||
```
|
```
|
||||||
|
|
||||||
## 6. Persistence And Trace
|
Chat 链路的质量门禁是 LLM Verifier。Verifier 输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||||
|
|
||||||
Current trace-related persistence:
|
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
||||||
|
|
||||||
|
关键代码:
|
||||||
|
|
||||||
|
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||||
|
|
||||||
|
## 4. AIOps 诊断链路
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Request["POST /api/ai_ops"] --> Payload{"包含告警 payload?"}
|
||||||
|
Payload -->|是| Targeted["PAYLOAD_TARGETED"]
|
||||||
|
Payload -->|否| Discovery["AUTO_DISCOVERY"]
|
||||||
|
|
||||||
|
Targeted --> BuildPrompt["构造聚焦 payload 的诊断 prompt"]
|
||||||
|
Targeted --> QueryAug["生成 recommended lookup_knowledge query"]
|
||||||
|
Discovery --> DiscoverAlert["通过 queryPrometheusAlerts 发现活跃告警"]
|
||||||
|
|
||||||
|
BuildPrompt --> Plan["Planner 规划排查"]
|
||||||
|
QueryAug --> Plan
|
||||||
|
DiscoverAlert --> Plan
|
||||||
|
|
||||||
|
Plan --> Execute["Executor 收集证据"]
|
||||||
|
Execute --> Knowledge["lookup_knowledge"]
|
||||||
|
Execute --> Metrics["query_metrics / Prometheus"]
|
||||||
|
Execute --> Logs["query_logs"]
|
||||||
|
|
||||||
|
Knowledge --> Report["告警分析报告"]
|
||||||
|
Metrics --> Report
|
||||||
|
Logs --> Report
|
||||||
|
|
||||||
|
Report --> RuleEval["AiOpsRuleEvaluationService"]
|
||||||
|
RuleEval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
|
||||||
|
Report --> Trace["DiagnosisTraceService"]
|
||||||
|
SelfEval --> Trace
|
||||||
|
```
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/ai_ops
|
||||||
|
-> AiOpsService
|
||||||
|
-> 判断是否有告警 payload
|
||||||
|
-> PAYLOAD_TARGETED
|
||||||
|
-> AUTO_DISCOVERY
|
||||||
|
-> 构造 AIOps 诊断 prompt
|
||||||
|
-> payload 模式补充 recommended lookup_knowledge query
|
||||||
|
-> Agent 编排
|
||||||
|
-> Planner / Executor
|
||||||
|
-> Prometheus / logs / knowledge tools
|
||||||
|
-> 生成告警分析报告
|
||||||
|
-> AiOpsRuleEvaluationService
|
||||||
|
-> 合并 self_evaluation.aiops_rule_evaluation
|
||||||
|
-> Trace API 可查看全链路
|
||||||
|
```
|
||||||
|
|
||||||
|
AIOps 保留两种模式:
|
||||||
|
|
||||||
|
| 模式 | 触发条件 | 行为 |
|
||||||
|
|---|---|---|
|
||||||
|
| `PAYLOAD_TARGETED` | 请求包含 alertName、service、severity、description、timeRange 等字段 | 以 payload 为唯一主诊断对象,并生成推荐知识库 query |
|
||||||
|
| `AUTO_DISCOVERY` | 请求没有明确告警 payload | 先查询当前活跃告警,再选择目标排查 |
|
||||||
|
|
||||||
|
AIOps 当前使用轻量规则验证器,重点检查:
|
||||||
|
|
||||||
|
- 最终报告是否存在。
|
||||||
|
- payload 模式是否聚焦输入告警。
|
||||||
|
- 是否使用关键证据工具,例如 `lookup_knowledge`、日志、指标。
|
||||||
|
|
||||||
|
关键代码:
|
||||||
|
|
||||||
|
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
|
||||||
|
|
||||||
|
## 5. RAG 位置
|
||||||
|
|
||||||
|
RAG 不是隐藏在 Chat Advisor 里的隐式能力,而是 Executor 可以显式调用的工具:
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
Executor["Executor Agent"] --> Tool["lookup_knowledge Tool"]
|
||||||
|
Tool --> L0["L0 domain/entity hint"]
|
||||||
|
Tool --> Search["VectorSearchService"]
|
||||||
|
L0 --> Search
|
||||||
|
Search --> VectorStore["Spring AI VectorStore"]
|
||||||
|
Search --> Fallback["Milvus SDK fallback"]
|
||||||
|
VectorStore --> Normalize["score/rawScore/scoreLabel"]
|
||||||
|
Fallback --> Normalize
|
||||||
|
Normalize --> Evidence["evidence output"]
|
||||||
|
Evidence --> Invocation["tool_invocation"]
|
||||||
|
Evidence --> Executor
|
||||||
|
```
|
||||||
|
|
||||||
|
```text
|
||||||
|
Executor
|
||||||
|
-> lookup_knowledge(query)
|
||||||
|
-> L0 domain/entity hint
|
||||||
|
-> VectorSearchService
|
||||||
|
-> Spring AI VectorStore
|
||||||
|
-> Milvus SDK fallback
|
||||||
|
-> evidence shaping
|
||||||
|
-> tool_invocation
|
||||||
|
```
|
||||||
|
|
||||||
|
保留显式工具的原因:
|
||||||
|
|
||||||
|
- Agent 何时检索、检索什么、证据是什么,必须能在 trace 中解释。
|
||||||
|
- AIOps payload 到 query 的业务映射需要项目内控制。
|
||||||
|
- `tool_invocation` 是后续评测、回放和面试讲解的核心材料。
|
||||||
|
|
||||||
|
RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细节见 [retrieval-observability.md](retrieval-observability.md)。
|
||||||
|
|
||||||
|
## 6. 持久化模型
|
||||||
|
|
||||||
|
当前诊断持久化以三张表为核心:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
diagnosis_session
|
diagnosis_session
|
||||||
-> final_report
|
-> 一次诊断会话的主记录
|
||||||
|
-> query / status / agent_flow / answer
|
||||||
-> self_evaluation
|
-> self_evaluation
|
||||||
-> verifier_evaluation
|
-> step_count / tool_call_count / duration
|
||||||
-> aiops_rule_evaluation
|
|
||||||
|
|
||||||
agent_step
|
agent_step
|
||||||
-> role
|
-> Agent 模型调用步骤
|
||||||
-> step input/output
|
-> step_index / agent_name
|
||||||
-> execution order
|
-> model_input / model_output / thought
|
||||||
|
-> duration / token_count
|
||||||
|
|
||||||
tool_invocation
|
tool_invocation
|
||||||
-> tool_name
|
-> 工具调用事实
|
||||||
-> query
|
-> tool_name / input_params / output_preview
|
||||||
-> retrieval_layer
|
-> retrieval_layer / retrieval_details
|
||||||
-> retrieval_details
|
-> relevance_level / dedup_reason
|
||||||
-> evidence blocks
|
-> duration / success
|
||||||
-> duration
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Trace API aggregates these records into a session-level view:
|
说明:
|
||||||
|
|
||||||
- Agent step sequence.
|
- 旧的 `diagnosis_record` 已不是当前主模型,迁移脚本中已经由 `diagnosis_session + agent_step + tool_invocation` 取代。
|
||||||
- Tool calls and retrieval details.
|
- `api_document` 仍用于文档元数据管理。
|
||||||
- Final diagnosis report.
|
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
||||||
- Chat verifier status.
|
|
||||||
- AIOps rule verifier status.
|
|
||||||
|
|
||||||
## 7. Quality Gates
|
会话和 Trace 生命周期见 [session-trace-lifecycle.md](session-trace-lifecycle.md),完整数据关系见 [data-model.md](data-model.md)。
|
||||||
|
|
||||||
Current quality gates:
|
## 7. Trace API
|
||||||
|
|
||||||
- Chat verifier: LLM-based final answer verification for normal diagnosis.
|
```text
|
||||||
- AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
|
GET /api/diagnosis/{sessionId}/trace
|
||||||
- Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
|
```
|
||||||
- RAG retrieval baseline: golden query set with offline baseline report.
|
|
||||||
- Live RAG acceptance: post-reindex script for validating retrieval against the running stack.
|
|
||||||
|
|
||||||
These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.
|
Trace API 聚合:
|
||||||
|
|
||||||
## 8. Current Completion State
|
- 会话状态和最终报告。
|
||||||
|
- Agent step 序列。
|
||||||
|
- 工具调用和检索细节。
|
||||||
|
- Chat verifier 结果。
|
||||||
|
- AIOps rule evaluation 结果。
|
||||||
|
|
||||||
Completed for the current MVP stage:
|
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
|
||||||
|
|
||||||
- Explicit `lookup_knowledge` Agent tool.
|
Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||||
- L0 + L1 retrieval shape retained.
|
|
||||||
- L0 downgraded to domain/entity hint.
|
|
||||||
- Spring AI VectorStore retrieval path integrated.
|
|
||||||
- Milvus SDK fallback retained.
|
|
||||||
- RAG evidence post-processing added.
|
|
||||||
- Breadcrumb/title/content embedding text improved.
|
|
||||||
- RAG offline baseline and live acceptance script added.
|
|
||||||
- AIOps payload query augmentation added.
|
|
||||||
- AIOps lightweight verifier added.
|
|
||||||
- Trace summary includes both chat verifier and AIOps verifier signals.
|
|
||||||
|
|
||||||
Deferred future enhancements:
|
## 8. 质量门禁
|
||||||
|
|
||||||
- LLM QueryTransformer / MultiQuery.
|
当前质量门禁分层如下:
|
||||||
- BM25, RRF, and reranker.
|
|
||||||
- Neighbor chunk or section-level context expansion.
|
|
||||||
- VectorStore write path migration.
|
|
||||||
- Full LLM-based AIOps verifier.
|
|
||||||
- More complete golden set for recall, MRR, and nDCG metrics.
|
|
||||||
|
|
||||||
## 9. Key Code References
|
| 门禁 | 位置 | 作用 |
|
||||||
|
|---|---|---|
|
||||||
|
| Chat Verifier | `ChatService` | 校验普通诊断回答质量 |
|
||||||
|
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
|
||||||
|
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
|
||||||
|
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
|
||||||
|
| Live RAG Acceptance | `scripts/eval_rag_live_acceptance.py` | 在运行环境中验证重建索引后的真实检索 |
|
||||||
|
|
||||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
## 9. 当前完成状态
|
||||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
|
||||||
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
|
|
||||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
|
||||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
|
||||||
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
|
||||||
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
|
|
||||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
|
||||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
|
||||||
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
|
||||||
|
|
||||||
## 10. Supporting Materials
|
已经完成:
|
||||||
|
|
||||||
- `mvp/issues/rag-refactor-plan.md`
|
- Chat 和 AIOps 两条入口链路。
|
||||||
- `eval/rag-retrieval/README.md`
|
- 显式 `lookup_knowledge` Agent Tool。
|
||||||
- `scripts/eval_rag_live_acceptance.py`
|
- L0 从最终决策降级为 domain/entity hint。
|
||||||
- `interview/rag-refactor-story.md`
|
- `VectorSearchService` 作为稳定检索门面。
|
||||||
- `interview/rag-vectorstore-interview-notes.md`
|
- Spring AI VectorStore 读取路径。
|
||||||
- `interview/rag-retrieval-quality-report.md`
|
- Milvus SDK fallback。
|
||||||
- `interview/rag-breadcrumb-embedding-acceptance.md`
|
- `score` / `rawScore` / `scoreLabel` 分数语义拆分。
|
||||||
- `interview/aiops-query-augmentation.md`
|
- `title`、`breadcrumb`、`content` 参与 embedding 文本。
|
||||||
- `interview/aiops-lightweight-verifier.md`
|
- `tool_invocation` 记录检索层、relevance level、dedup reason。
|
||||||
|
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
|
||||||
|
- RAG offline baseline 和 live acceptance 脚本。
|
||||||
|
|
||||||
|
暂不作为当前已完成能力声明:
|
||||||
|
|
||||||
|
- 完整 QueryTransformer / MultiQuery。
|
||||||
|
- BM25、RRF、cross-encoder rerank。
|
||||||
|
- 完整邻居 chunk / section context expansion。
|
||||||
|
- VectorStore 写入路径全面迁移。
|
||||||
|
- 完整 LLM-based AIOps verifier。
|
||||||
|
|
||||||
|
后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
|
||||||
|
|
||||||
|
## 10. 关键代码索引
|
||||||
|
|
||||||
|
| 能力 | 代码 |
|
||||||
|
|---|---|
|
||||||
|
| Chat 入口与编排 | `ChatController`, `ChatService` |
|
||||||
|
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
|
||||||
|
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
|
||||||
|
| 知识库工具 | `LookupKnowledgeTool` |
|
||||||
|
| L0 hint | `KnowledgeIndexService` |
|
||||||
|
| 向量检索门面 | `VectorSearchService` |
|
||||||
|
| 文档切片 | `DocumentChunkService` |
|
||||||
|
| 向量写入 | `VectorIndexService` |
|
||||||
|
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
|
||||||
|
| Trace 聚合 | `DiagnosisTraceService` |
|
||||||
|
| 工具调用记录 | `ToolInvocationRecorder` |
|
||||||
|
| self_evaluation 合并 | `SelfEvaluationMergeService` |
|
||||||
|
|||||||
@@ -0,0 +1,258 @@
|
|||||||
|
# 数据模型总览
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:当前可运行架构
|
||||||
|
|
||||||
|
## 1. 定位
|
||||||
|
|
||||||
|
本文从架构角度说明当前 MVP 的核心数据模型。详细字段仍以 Flyway migration 和 `mvp/tables/` 为准。
|
||||||
|
|
||||||
|
核心数据分三组:
|
||||||
|
|
||||||
|
- 诊断 Trace:`diagnosis_session`、`agent_step`、`tool_invocation`
|
||||||
|
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
|
||||||
|
- 反馈沉淀:`case_library`
|
||||||
|
|
||||||
|
## 2. 总体关系
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
erDiagram
|
||||||
|
diagnosis_session ||--o{ agent_step : has
|
||||||
|
diagnosis_session ||--o{ tool_invocation : has
|
||||||
|
diagnosis_session ||--o| case_library : creates_when_useful
|
||||||
|
api_document ||--o{ milvus_chunk : indexed_as
|
||||||
|
knowledge_domain ||--o{ api_document : groups
|
||||||
|
|
||||||
|
diagnosis_session {
|
||||||
|
bigint id
|
||||||
|
varchar session_id
|
||||||
|
text query
|
||||||
|
varchar status
|
||||||
|
varchar agent_flow
|
||||||
|
longtext answer
|
||||||
|
json self_evaluation
|
||||||
|
varchar feedback
|
||||||
|
}
|
||||||
|
|
||||||
|
agent_step {
|
||||||
|
bigint id
|
||||||
|
varchar session_id
|
||||||
|
int step_index
|
||||||
|
varchar agent_name
|
||||||
|
text model_input
|
||||||
|
text model_output
|
||||||
|
text thought
|
||||||
|
boolean has_tool_call
|
||||||
|
}
|
||||||
|
|
||||||
|
tool_invocation {
|
||||||
|
bigint id
|
||||||
|
varchar session_id
|
||||||
|
varchar tool_name
|
||||||
|
json input_params
|
||||||
|
text output_preview
|
||||||
|
varchar retrieval_layer
|
||||||
|
json retrieval_details
|
||||||
|
varchar relevance_level
|
||||||
|
varchar dedup_reason
|
||||||
|
}
|
||||||
|
|
||||||
|
api_document {
|
||||||
|
bigint id
|
||||||
|
varchar doc_id
|
||||||
|
varchar file_name
|
||||||
|
varchar file_path
|
||||||
|
varchar status
|
||||||
|
int chunk_count
|
||||||
|
text metadata
|
||||||
|
}
|
||||||
|
|
||||||
|
knowledge_domain {
|
||||||
|
bigint id
|
||||||
|
varchar domain_id
|
||||||
|
varchar description
|
||||||
|
text when_to_retrieve
|
||||||
|
int document_count
|
||||||
|
}
|
||||||
|
|
||||||
|
case_library {
|
||||||
|
bigint id
|
||||||
|
varchar case_id
|
||||||
|
varchar diagnosis_id
|
||||||
|
varchar source_type
|
||||||
|
varchar fault_category
|
||||||
|
text root_cause
|
||||||
|
text solution
|
||||||
|
}
|
||||||
|
|
||||||
|
milvus_chunk {
|
||||||
|
varchar id
|
||||||
|
text content
|
||||||
|
json metadata
|
||||||
|
vector vector
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
|
||||||
|
|
||||||
|
## 3. 诊断 Trace 模型
|
||||||
|
|
||||||
|
### diagnosis_session
|
||||||
|
|
||||||
|
会话级主记录。
|
||||||
|
|
||||||
|
关键字段:
|
||||||
|
|
||||||
|
| 字段 | 说明 |
|
||||||
|
|---|---|
|
||||||
|
| `session_id` | 外部关联键,Trace 和 Feedback 都使用它 |
|
||||||
|
| `query` | 用户原始问题或 AIOps 输入摘要 |
|
||||||
|
| `status` | 执行状态 |
|
||||||
|
| `agent_flow` | `CHAT` / `AI_OPS` |
|
||||||
|
| `answer` | 最终答复或告警报告 |
|
||||||
|
| `self_evaluation` | rule/verifier/aiops 自评估容器 |
|
||||||
|
| `feedback` | 用户反馈 |
|
||||||
|
|
||||||
|
### agent_step
|
||||||
|
|
||||||
|
记录模型调用步骤。
|
||||||
|
|
||||||
|
用途:
|
||||||
|
|
||||||
|
- 回放 Agent 推理过程。
|
||||||
|
- 查看 Planner / Executor / Verifier 的输入输出摘要。
|
||||||
|
- 统计 step count、duration、token count。
|
||||||
|
|
||||||
|
### tool_invocation
|
||||||
|
|
||||||
|
记录工具调用事实。
|
||||||
|
|
||||||
|
用途:
|
||||||
|
|
||||||
|
- 给 Trace API 展示证据。
|
||||||
|
- 给 Verifier 构造 `tool_trace_summary`。
|
||||||
|
- 给 `EvaluationService` 计算 evidence score。
|
||||||
|
- 给 RAG eval 和人工排查提供检索细节。
|
||||||
|
|
||||||
|
## 4. 知识库模型
|
||||||
|
|
||||||
|
### api_document
|
||||||
|
|
||||||
|
MySQL 中的文档元数据表。
|
||||||
|
|
||||||
|
职责:
|
||||||
|
|
||||||
|
- 管理上传文件。
|
||||||
|
- 保存 file hash,用于去重。
|
||||||
|
- 记录索引状态和 chunk 数量。
|
||||||
|
- 保存 frontmatter JSON。
|
||||||
|
|
||||||
|
### knowledge_domain
|
||||||
|
|
||||||
|
领域级元数据。
|
||||||
|
|
||||||
|
职责:
|
||||||
|
|
||||||
|
- 按 category 聚合文档。
|
||||||
|
- 存储领域描述。
|
||||||
|
- 存储 `when_to_retrieve`,辅助 Planner/Executor 判断什么时候检索该领域。
|
||||||
|
|
||||||
|
### Milvus/Zilliz metadata
|
||||||
|
|
||||||
|
向量 collection 中每个 chunk 的 metadata 主要包括:
|
||||||
|
|
||||||
|
```text
|
||||||
|
docId
|
||||||
|
_source
|
||||||
|
chunkIndex
|
||||||
|
totalChunks
|
||||||
|
title
|
||||||
|
breadcrumb
|
||||||
|
category
|
||||||
|
```
|
||||||
|
|
||||||
|
这些字段支撑:
|
||||||
|
|
||||||
|
- category filter。
|
||||||
|
- source 展示。
|
||||||
|
- breadcrumb 上下文。
|
||||||
|
- docId 删除和重建索引。
|
||||||
|
- evidence block 构造。
|
||||||
|
|
||||||
|
## 5. 反馈沉淀模型
|
||||||
|
|
||||||
|
### case_library
|
||||||
|
|
||||||
|
`useful` 反馈会触发 `CaseLibraryService.createFromSession`。
|
||||||
|
|
||||||
|
当前自动映射:
|
||||||
|
|
||||||
|
| 字段 | 来源 |
|
||||||
|
|---|---|
|
||||||
|
| `case_id` | UUID |
|
||||||
|
| `diagnosis_id` | `diagnosis_session.session_id` |
|
||||||
|
| `source_type` | `AUTO` |
|
||||||
|
| `fault_category` | 当前默认 `GENERAL` |
|
||||||
|
| `title` | session query 前 100 字符 |
|
||||||
|
| `root_cause` | session answer |
|
||||||
|
| `solution` | session answer |
|
||||||
|
| `created_by` | `system` |
|
||||||
|
|
||||||
|
## 6. self_evaluation 结构
|
||||||
|
|
||||||
|
`diagnosis_session.self_evaluation` 是 JSON 容器:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"rule_evaluation": {},
|
||||||
|
"verifier_evaluation": {},
|
||||||
|
"aiops_rule_evaluation": {}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
边界:
|
||||||
|
|
||||||
|
- `rule_evaluation` 评估证据收集充分度。
|
||||||
|
- `verifier_evaluation` 评估 Chat 答案关键事实是否有证据支撑。
|
||||||
|
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
|
||||||
|
|
||||||
|
## 7. 数据写入时序
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
sequenceDiagram
|
||||||
|
autonumber
|
||||||
|
participant API as API
|
||||||
|
participant Svc as ChatService/AiOpsService
|
||||||
|
participant Session as diagnosis_session
|
||||||
|
participant Agent as Agent
|
||||||
|
participant Step as agent_step
|
||||||
|
participant Tool as tool_invocation
|
||||||
|
participant Eval as self_evaluation
|
||||||
|
participant Feedback as case_library
|
||||||
|
|
||||||
|
API->>Svc: request
|
||||||
|
Svc->>Session: create/update RUNNING
|
||||||
|
Agent->>Step: before/after model
|
||||||
|
Agent->>Tool: tool call record
|
||||||
|
Svc->>Session: SUCCESS/FAILED + answer
|
||||||
|
Svc->>Eval: merge evaluation
|
||||||
|
API->>Svc: feedback useful
|
||||||
|
Svc->>Feedback: create case
|
||||||
|
```
|
||||||
|
|
||||||
|
## 8. 当前边界和后续
|
||||||
|
|
||||||
|
当前边界:
|
||||||
|
|
||||||
|
- `agent_step.session_id` 和 `tool_invocation.session_id` 通过 sessionId 关联,不强制外键。
|
||||||
|
- `tool_invocation.step_id` 可为空。
|
||||||
|
- Milvus chunk 与 `api_document` 通过 metadata.docId 逻辑关联。
|
||||||
|
- `case_library` 与 session 通过 `diagnosis_id=session_id` 关联。
|
||||||
|
|
||||||
|
后续可增强:
|
||||||
|
|
||||||
|
1. 增加 run id,支持同 session 多次独立诊断。
|
||||||
|
2. 强化 `tool_invocation.step_id` 关联。
|
||||||
|
3. 将 evidence block 结构化保存。
|
||||||
|
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
||||||
|
|
||||||
@@ -0,0 +1,179 @@
|
|||||||
|
# Agent 架构演进路线
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:后续演进设计,不代表当前已实现
|
||||||
|
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||||
|
|
||||||
|
## 1. 为什么需要演进路线
|
||||||
|
|
||||||
|
旧版 `agent-architecture.md` 包含很多生产级设想:专科 SubAgent、Skill 体系、进程隔离、回退路由、MCP 工具协议化、进化引擎。它们不应作为当前 MVP 事实写入主架构,但可以作为后续扩展路线。
|
||||||
|
|
||||||
|
当前原则:
|
||||||
|
|
||||||
|
- 当前文档只声明已经可运行或明确落地的能力。
|
||||||
|
- 演进路线记录未来方向和触发条件。
|
||||||
|
- 每个演进项必须有可验证收益,不能只因为“架构更炫”就拆。
|
||||||
|
|
||||||
|
## 2. 演进总图
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
MVP["Current MVP: Planner + Executor + Verifier"] --> Split{"Executor 是否过载?"}
|
||||||
|
Split -->|是| SubAgents["专科 SubAgent"]
|
||||||
|
Split -->|否| Keep["继续强化通用 Executor"]
|
||||||
|
|
||||||
|
SubAgents --> Skills["Skill / Playbook 体系"]
|
||||||
|
Skills --> Fallback["回退路由"]
|
||||||
|
Fallback --> Isolation["进程或 Pod 隔离"]
|
||||||
|
|
||||||
|
MVP --> ToolGrowth{"工具数量和来源是否增长?"}
|
||||||
|
ToolGrowth -->|是| MCP["MCP / Tool Server 协议化"]
|
||||||
|
ToolGrowth -->|否| ToolCallbacks["继续使用 @Tool / ToolCallback"]
|
||||||
|
|
||||||
|
MVP --> EvalGrowth{"评测数据是否足够?"}
|
||||||
|
EvalGrowth -->|是| Evolution["Prompt / Skill 进化引擎"]
|
||||||
|
EvalGrowth -->|否| Baseline["先扩大 baseline"]
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. 专科 SubAgent
|
||||||
|
|
||||||
|
### 触发条件
|
||||||
|
|
||||||
|
- Executor prompt 变得臃肿,难以同时覆盖接口、数据库、缓存、网络等场景。
|
||||||
|
- 不同故障类型需要明显不同的工具权限。
|
||||||
|
- Trace 显示某些场景经常走错排查路径。
|
||||||
|
- 评测集已经能衡量拆分前后的收益。
|
||||||
|
|
||||||
|
### 候选 SubAgent
|
||||||
|
|
||||||
|
| SubAgent | 场景 | 工具倾向 |
|
||||||
|
|---|---|---|
|
||||||
|
| `ExternalApiSubAgent` | 错误码、接口参数、第三方调用失败 | `lookup_knowledge`, logs, trace |
|
||||||
|
| `DatabaseSubAgent` | 连接池、慢 SQL、死锁、数据库不可用 | metrics, logs, knowledge |
|
||||||
|
| `CacheSubAgent` | Redis 超时、热点 key、内存风险 | metrics, logs, knowledge |
|
||||||
|
| `GenericDiagnosisSubAgent` | 兜底诊断 | 全量只读证据工具 |
|
||||||
|
|
||||||
|
### 不立即拆分的原因
|
||||||
|
|
||||||
|
- 当前 MVP 的工具规模还可由通用 Executor 管理。
|
||||||
|
- 过早拆分会增加 Prompt、评测和 trace 分析成本。
|
||||||
|
- 没有足够分类评测前,拆分可能只是移动复杂度。
|
||||||
|
|
||||||
|
## 4. Skill / Playbook 体系
|
||||||
|
|
||||||
|
旧版设计中的 Skill 可以在当前项目中演进为可版本化的诊断 Playbook。
|
||||||
|
|
||||||
|
```text
|
||||||
|
fault_category
|
||||||
|
-> playbook
|
||||||
|
-> required evidence
|
||||||
|
-> tool sequence
|
||||||
|
-> stop condition
|
||||||
|
-> report template
|
||||||
|
-> evaluation checks
|
||||||
|
```
|
||||||
|
|
||||||
|
优先落地方向:
|
||||||
|
|
||||||
|
- AIOps 告警处理 Playbook。
|
||||||
|
- 支付超时 Playbook。
|
||||||
|
- MySQL 连接池风险 Playbook。
|
||||||
|
- Redis timeout Playbook。
|
||||||
|
|
||||||
|
落地前提:
|
||||||
|
|
||||||
|
- 每个 Playbook 至少有 3-5 个 eval case。
|
||||||
|
- Playbook 失败时可以回退到通用 Executor。
|
||||||
|
- Trace 中能标记使用了哪个 Playbook 和哪个版本。
|
||||||
|
|
||||||
|
## 5. 回退路由
|
||||||
|
|
||||||
|
当前 Chat 已有低置信补证据和 REJECT 降级输出。后续如果引入 SubAgent,可扩展为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Specialized SubAgent
|
||||||
|
-> failed / low confidence
|
||||||
|
-> another specialized SubAgent
|
||||||
|
-> GenericDiagnosisSubAgent
|
||||||
|
-> degraded answer with confirmed facts only
|
||||||
|
```
|
||||||
|
|
||||||
|
回退依据:
|
||||||
|
|
||||||
|
- 工具连续失败。
|
||||||
|
- Verifier `REJECT`。
|
||||||
|
- Verifier `LOW_CONFID` 且补证据失败。
|
||||||
|
- Agent 输出缺失关键报告字段。
|
||||||
|
|
||||||
|
## 6. 进程隔离
|
||||||
|
|
||||||
|
当前所有 Agent 在同一 JVM 内运行。生产级隔离可以考虑:
|
||||||
|
|
||||||
|
```text
|
||||||
|
API service
|
||||||
|
-> Supervisor service
|
||||||
|
-> Planner service
|
||||||
|
-> SubAgent services
|
||||||
|
-> Verifier service
|
||||||
|
```
|
||||||
|
|
||||||
|
触发条件:
|
||||||
|
|
||||||
|
- 某类 Agent 需要独立扩缩容。
|
||||||
|
- 某类工具依赖不稳定,可能拖垮主应用。
|
||||||
|
- 不同 Agent 需要不同权限和网络访问策略。
|
||||||
|
- 单 JVM 内资源隔离不足。
|
||||||
|
|
||||||
|
MVP 阶段暂不拆分进程,优先保证 trace、评测和工具边界清晰。
|
||||||
|
|
||||||
|
## 7. MCP / Tool Server 协议化
|
||||||
|
|
||||||
|
当前工具主要通过 `@Tool`、`methodTools` 和 `ToolCallbackProvider` 暴露。工具数量增加后,可演进为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Agent
|
||||||
|
-> Tool registry
|
||||||
|
-> MCP / tool server
|
||||||
|
-> log server
|
||||||
|
-> metrics server
|
||||||
|
-> knowledge server
|
||||||
|
-> ticket/change server
|
||||||
|
```
|
||||||
|
|
||||||
|
收益:
|
||||||
|
|
||||||
|
- 工具独立部署。
|
||||||
|
- 新工具上线不必重发主应用。
|
||||||
|
- 不同 Agent 可获得不同工具子集。
|
||||||
|
- 工具调用协议统一,更利于审计。
|
||||||
|
|
||||||
|
风险:
|
||||||
|
|
||||||
|
- 调用链更长。
|
||||||
|
- 权限和超时治理更复杂。
|
||||||
|
- 本地开发和 Demo 成本上升。
|
||||||
|
|
||||||
|
## 8. 进化引擎
|
||||||
|
|
||||||
|
旧版文档提到从诊断中学习。当前可以拆成更务实的步骤:
|
||||||
|
|
||||||
|
1. 先扩大 diagnosis eval 和 RAG eval。
|
||||||
|
2. 从失败 trace 中标注 bad case。
|
||||||
|
3. 将高频失败沉淀为 Playbook 或 Prompt 规则。
|
||||||
|
4. 对 Prompt 版本做离线对比。
|
||||||
|
5. 足够稳定后再考虑线上 A/B。
|
||||||
|
|
||||||
|
不建议 MVP 直接做自动 Prompt 自优化。没有可靠评测和回滚机制时,自动优化更容易引入不可解释变化。
|
||||||
|
|
||||||
|
## 9. 演进优先级
|
||||||
|
|
||||||
|
| 优先级 | 项目 | 原因 |
|
||||||
|
|---|---|---|
|
||||||
|
| P0 | 扩大 eval baseline | 没有评测,拆任何架构都难以证明收益 |
|
||||||
|
| P1 | Playbook 化高频故障 | 可控、可解释、比拆 SubAgent 更轻 |
|
||||||
|
| P1 | 完整 evidence block | 提升 Verifier 和 Trace 质量 |
|
||||||
|
| P2 | 专科 SubAgent | 等问题类型和工具权限差异足够明显 |
|
||||||
|
| P2 | AIOps LLM Verifier | 规则门禁不足时再引入 |
|
||||||
|
| P3 | MCP 工具协议化 | 工具来源复杂后再做 |
|
||||||
|
| P3 | 进程隔离 | 生产负载和权限隔离需要明确后再做 |
|
||||||
|
|
||||||
@@ -0,0 +1,251 @@
|
|||||||
|
# 反馈与自评估架构
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:当前可运行架构
|
||||||
|
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
||||||
|
|
||||||
|
## 1. 定位
|
||||||
|
|
||||||
|
反馈架构包含两条闭环:
|
||||||
|
|
||||||
|
1. 系统自评估:基于工具调用、Verifier、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。
|
||||||
|
2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。
|
||||||
|
|
||||||
|
当前重要边界:
|
||||||
|
|
||||||
|
- `status` 表示执行状态,不表示答案质量。
|
||||||
|
- `feedback` 表示用户反馈,不覆盖 `status`。
|
||||||
|
- `self_evaluation` 是 JSON 容器,内部按来源分层,不再把所有评分字段平铺在根节点。
|
||||||
|
|
||||||
|
## 2. 总体闭环
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Answer["Chat / AIOps final answer"] --> Session["diagnosis_session.answer"]
|
||||||
|
|
||||||
|
subgraph SelfEval["Self evaluation"]
|
||||||
|
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
||||||
|
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||||
|
TraceSummary --> Verifier["chat_verifier"]
|
||||||
|
Verifier --> VerifierEval["verifier_evaluation"]
|
||||||
|
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||||
|
AiOpsRule --> AiOpsEval["aiops_rule_evaluation"]
|
||||||
|
end
|
||||||
|
|
||||||
|
RuleEval --> Merge["SelfEvaluationMergeService"]
|
||||||
|
VerifierEval --> Merge
|
||||||
|
AiOpsEval --> Merge
|
||||||
|
Merge --> SelfJson["diagnosis_session.self_evaluation"]
|
||||||
|
|
||||||
|
subgraph UserFeedback["User feedback"]
|
||||||
|
UI["Feedback bar"] --> API["POST /api/feedback"]
|
||||||
|
API --> FeedbackService["FeedbackService"]
|
||||||
|
FeedbackService --> FeedbackField["diagnosis_session.feedback"]
|
||||||
|
FeedbackService --> Useful{"feedback == useful?"}
|
||||||
|
Useful -->|yes| CaseService["CaseLibraryService.createFromSession"]
|
||||||
|
CaseService --> Case["case_library"]
|
||||||
|
Useful -->|no| BadCase["Bad case by feedback=not_useful"]
|
||||||
|
end
|
||||||
|
|
||||||
|
Session --> UI
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. self_evaluation JSON
|
||||||
|
|
||||||
|
`SelfEvaluationMergeService` 统一维护 `diagnosis_session.self_evaluation`。
|
||||||
|
|
||||||
|
当前结构:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"rule_evaluation": {
|
||||||
|
"evidence_score": 65,
|
||||||
|
"source": "rule",
|
||||||
|
"factors": []
|
||||||
|
},
|
||||||
|
"verifier_evaluation": {
|
||||||
|
"verdict": "PASS",
|
||||||
|
"groundedness_score": 0.8,
|
||||||
|
"critical_fact_count": 2,
|
||||||
|
"facts_checked": [],
|
||||||
|
"rationale": "...",
|
||||||
|
"round": 1,
|
||||||
|
"traceability_version": "v1",
|
||||||
|
"tool_trace_summary": []
|
||||||
|
},
|
||||||
|
"aiops_rule_evaluation": {
|
||||||
|
"verdict": "...",
|
||||||
|
"checks": []
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
兼容逻辑:
|
||||||
|
|
||||||
|
- 如果旧 JSON 根节点包含 `evidence_score`,会被包进 `rule_evaluation`。
|
||||||
|
- 如果旧 JSON 根节点包含 `verdict` / `groundedness_score`,会被包进 `verifier_evaluation`。
|
||||||
|
|
||||||
|
## 4. 规则评分
|
||||||
|
|
||||||
|
`EvaluationService` 只消费 `tool_invocation` 和 session 状态,输出 `rule_evaluation`。
|
||||||
|
|
||||||
|
定位:
|
||||||
|
|
||||||
|
- 衡量证据收集充分度。
|
||||||
|
- 不直接证明答案是否推理正确。
|
||||||
|
- 不依赖 LLM。
|
||||||
|
|
||||||
|
规则:
|
||||||
|
|
||||||
|
| 规则名 | 条件 | 分数变化 |
|
||||||
|
|---|---|---|
|
||||||
|
| `execution_failed` | session status = `FAILED` | 直接 0 |
|
||||||
|
| `no_tool_call` | 没有工具调用 | 直接 0 |
|
||||||
|
| `has_successful_tool_call` | 至少一次工具成功 | +30 |
|
||||||
|
| `l0_exact_match` | 任意工具调用有 L0 命中 | +35 |
|
||||||
|
| `l1_semantic_match` | 无 L0 命中但有 L1 命中 | +20 |
|
||||||
|
| `retrieval_no_hit` | 有检索调用但无命中 | -10 |
|
||||||
|
| `all_tool_calls_failed` | 工具全部失败 | -20 |
|
||||||
|
|
||||||
|
最终分数裁剪到 `[0, 100]`。
|
||||||
|
|
||||||
|
说明:
|
||||||
|
|
||||||
|
- 当前 `rule_evaluation` 是异步写入,失败时 `self_evaluation` 可能暂时为空或缺少该节点。
|
||||||
|
- L0/L1 分支互斥:有 L0 命中时优先记 L0。
|
||||||
|
- 更强的答案真实性校验由 Chat Verifier 承担。
|
||||||
|
|
||||||
|
## 5. Chat Verifier 自评估
|
||||||
|
|
||||||
|
Chat Verifier 校验 Executor 的最终答案是否被证据支撑。
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
Answer["executor_final_answer"] --> Verifier["chat_verifier"]
|
||||||
|
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
|
||||||
|
Summary --> Evidence["tool_trace_summary"]
|
||||||
|
Evidence --> Verifier
|
||||||
|
Verifier --> Output["verifier_output JSON"]
|
||||||
|
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
|
||||||
|
Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"]
|
||||||
|
```
|
||||||
|
|
||||||
|
Verifier 输出:
|
||||||
|
|
||||||
|
| 字段 | 说明 |
|
||||||
|
|---|---|
|
||||||
|
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||||
|
| `groundedness_score` | 关键事实证据支撑度 |
|
||||||
|
| `critical_fact_count` | 关键事实数量 |
|
||||||
|
| `facts_checked` | 逐条事实校验 |
|
||||||
|
| `rationale` | 判定原因 |
|
||||||
|
| `tool_trace_summary` | 本次校验使用的证据索引 |
|
||||||
|
|
||||||
|
ChatService 根据 verdict 决定:
|
||||||
|
|
||||||
|
- `PASS`:输出 Executor 答案。
|
||||||
|
- `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。
|
||||||
|
- `REJECT`:降级输出,只保留已确认信息。
|
||||||
|
|
||||||
|
## 6. AIOps 规则自评估
|
||||||
|
|
||||||
|
AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。
|
||||||
|
|
||||||
|
检查重点:
|
||||||
|
|
||||||
|
- 是否有最终报告。
|
||||||
|
- payload 模式是否聚焦输入告警。
|
||||||
|
- 是否调用 `lookup_knowledge`、日志、指标等证据工具。
|
||||||
|
- 是否把无关活跃告警扩展成主诊断对象。
|
||||||
|
|
||||||
|
这是轻量规则检查,不等价于完整 LLM Verifier。完整 AIOps Verifier 是后续增强项。
|
||||||
|
|
||||||
|
## 7. 用户反馈 API
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/feedback
|
||||||
|
Content-Type: application/json
|
||||||
|
|
||||||
|
{
|
||||||
|
"sessionId": "xxx",
|
||||||
|
"feedback": "useful" | "not_useful"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
响应:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"success": true,
|
||||||
|
"message": "反馈已记录",
|
||||||
|
"caseId": "uuid 或 null"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
后端行为:
|
||||||
|
|
||||||
|
| feedback | 行为 |
|
||||||
|
|---|---|
|
||||||
|
| `useful` | 写入 `DiagnosisSession.feedback`,调用 `CaseLibraryService.createFromSession` |
|
||||||
|
| `not_useful` | 写入 `DiagnosisSession.feedback`,不改变 session status |
|
||||||
|
| 其他值 | 返回 HTTP 400 |
|
||||||
|
|
||||||
|
## 8. 案例沉淀
|
||||||
|
|
||||||
|
`useful` 反馈会生成或复用 `case_library` 记录。
|
||||||
|
|
||||||
|
字段映射:
|
||||||
|
|
||||||
|
| CaseLibrary 字段 | 来源 |
|
||||||
|
|---|---|
|
||||||
|
| `caseId` | UUID |
|
||||||
|
| `diagnosisId` | `DiagnosisSession.sessionId` |
|
||||||
|
| `sourceType` | `AUTO` |
|
||||||
|
| `faultCategory` | 当前固定为 `GENERAL` |
|
||||||
|
| `title` | `query` 前 100 字符 |
|
||||||
|
| `rootCause` | `answer` |
|
||||||
|
| `solution` | `answer` |
|
||||||
|
| `createdBy` | `system` |
|
||||||
|
|
||||||
|
幂等性:
|
||||||
|
|
||||||
|
```text
|
||||||
|
case_library.diagnosisId == sessionId
|
||||||
|
-> existing case: return existing
|
||||||
|
-> missing case: create new
|
||||||
|
```
|
||||||
|
|
||||||
|
## 9. Trace 呈现
|
||||||
|
|
||||||
|
Trace API 会展示:
|
||||||
|
|
||||||
|
- `feedback`
|
||||||
|
- `hasFeedback`
|
||||||
|
- `hasVerifierEvaluation`
|
||||||
|
- `hasAiOpsRuleEvaluation`
|
||||||
|
- session、step、tool invocation 明细
|
||||||
|
|
||||||
|
这让一次诊断可以被分成三种视角查看:
|
||||||
|
|
||||||
|
| 视角 | 数据来源 |
|
||||||
|
|---|---|
|
||||||
|
| 执行是否成功 | `diagnosis_session.status` |
|
||||||
|
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
|
||||||
|
| 用户是否认可 | `diagnosis_session.feedback` |
|
||||||
|
|
||||||
|
## 10. 后续增强
|
||||||
|
|
||||||
|
近期优先:
|
||||||
|
|
||||||
|
1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。
|
||||||
|
2. `not_useful` 反馈沉淀 bad case,而不是只写字段。
|
||||||
|
3. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
|
||||||
|
4. AIOps 引入 LLM Verifier。
|
||||||
|
5. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
|
||||||
|
|
||||||
|
暂不优先:
|
||||||
|
|
||||||
|
- 用用户反馈直接修改 session status。
|
||||||
|
- 仅凭 `evidence_score` 判断答案正确。
|
||||||
|
- 在没有人工审核时自动把 bad case 反向写入 Prompt。
|
||||||
|
|
||||||
@@ -0,0 +1,205 @@
|
|||||||
|
# Harness 与质量门禁架构
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:当前可运行架构 + 后续门禁规划
|
||||||
|
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||||
|
|
||||||
|
## 1. 设计目标
|
||||||
|
|
||||||
|
Agent 系统的核心风险不是“没有答案”,而是:
|
||||||
|
|
||||||
|
- 答案引用了不存在的证据。
|
||||||
|
- 工具调用失败后仍然编造结论。
|
||||||
|
- 检索结果相关性不足但被当作强证据。
|
||||||
|
- 多轮诊断重复检索同一文档,浪费上下文。
|
||||||
|
- 最终报告无法回放执行过程。
|
||||||
|
|
||||||
|
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Prompt contract
|
||||||
|
+ Tool boundary
|
||||||
|
+ Agent hooks
|
||||||
|
+ Trace persistence
|
||||||
|
+ Verifier / rule evaluation
|
||||||
|
+ Eval baseline
|
||||||
|
```
|
||||||
|
|
||||||
|
## 2. Harness 总图
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||||||
|
Prompt --> Agent["Planner / Executor / Verifier"]
|
||||||
|
Agent --> Tools["Evidence tools"]
|
||||||
|
Tools --> Invocation["tool_invocation"]
|
||||||
|
Agent --> StepHook["AgentLoggingHook"]
|
||||||
|
StepHook --> Step["agent_step"]
|
||||||
|
Agent --> Session["diagnosis_session"]
|
||||||
|
|
||||||
|
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||||
|
TraceSummary --> Verifier["chat_verifier"]
|
||||||
|
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||||||
|
|
||||||
|
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||||
|
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||||
|
|
||||||
|
Session --> TraceAPI["DiagnosisTraceService"]
|
||||||
|
Step --> TraceAPI
|
||||||
|
Invocation --> TraceAPI
|
||||||
|
SelfEval --> TraceAPI
|
||||||
|
AiOpsEval --> TraceAPI
|
||||||
|
|
||||||
|
TraceAPI --> Eval["diagnosis eval / RAG eval"]
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. Prompt Contract
|
||||||
|
|
||||||
|
当前 Prompt 按角色拆分:
|
||||||
|
|
||||||
|
| Prompt | 用途 |
|
||||||
|
|---|---|
|
||||||
|
| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
|
||||||
|
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
|
||||||
|
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
|
||||||
|
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
|
||||||
|
| `chat-executor-prompt.md` | Chat 执行工具并形成诊断答复 |
|
||||||
|
| `chat-verifier-prompt.md` | 校验 Executor 答案是否被工具证据支撑 |
|
||||||
|
|
||||||
|
Prompt 层当前承担的门禁:
|
||||||
|
|
||||||
|
- 禁止凭记忆回答错误码、接口定义、排障步骤。
|
||||||
|
- 需要外部信息时必须调用工具。
|
||||||
|
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
|
||||||
|
- Chat Verifier 不允许做新检索,只能校验已有证据。
|
||||||
|
- AIOps payload 模式必须聚焦输入告警。
|
||||||
|
|
||||||
|
## 4. Trace Hooks
|
||||||
|
|
||||||
|
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
sequenceDiagram
|
||||||
|
autonumber
|
||||||
|
participant A as Agent
|
||||||
|
participant H as AgentLoggingHook
|
||||||
|
participant DB as agent_step
|
||||||
|
|
||||||
|
A->>H: before_model(messages, sessionId)
|
||||||
|
H->>DB: 写入 model_input / step_index / agent_name
|
||||||
|
A-->>A: LLM 推理
|
||||||
|
A->>H: after_model(messages, sessionId)
|
||||||
|
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||||||
|
```
|
||||||
|
|
||||||
|
记录内容:
|
||||||
|
|
||||||
|
- 最近输入消息摘要。
|
||||||
|
- Agent 输出摘要。
|
||||||
|
- 是否包含 tool call。
|
||||||
|
- duration。
|
||||||
|
- token count。
|
||||||
|
- Verifier 的 JSON 输出摘要。
|
||||||
|
|
||||||
|
## 5. Tool Invocation 门禁
|
||||||
|
|
||||||
|
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||||||
|
|
||||||
|
核心记录:
|
||||||
|
|
||||||
|
```text
|
||||||
|
tool_name
|
||||||
|
input_params
|
||||||
|
output_preview
|
||||||
|
retrieval_layer
|
||||||
|
l0_match_count
|
||||||
|
l1_match_count
|
||||||
|
retrieval_details
|
||||||
|
relevance_level
|
||||||
|
dedup_reason
|
||||||
|
duration_ms
|
||||||
|
success
|
||||||
|
error_message
|
||||||
|
```
|
||||||
|
|
||||||
|
对 `lookup_knowledge` 的质量约束:
|
||||||
|
|
||||||
|
- L0 只作为 hint,不绕过 L1。
|
||||||
|
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
|
||||||
|
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
|
||||||
|
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
|
||||||
|
|
||||||
|
## 6. Verifier 门禁
|
||||||
|
|
||||||
|
Chat Verifier 的输入不是原始工具日志,而是 `ToolTraceSummaryService` 构造的证据索引。
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
|
||||||
|
Summary --> EvidenceIndex["tool_trace_summary"]
|
||||||
|
EvidenceIndex --> Verifier["chat_verifier"]
|
||||||
|
ExecutorAnswer["executor_final_answer"] --> Verifier
|
||||||
|
Verifier --> Verdict{"verdict"}
|
||||||
|
Verdict -->|PASS| Pass["输出原答案"]
|
||||||
|
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
|
||||||
|
Verdict -->|REJECT| Reject["降级输出"]
|
||||||
|
```
|
||||||
|
|
||||||
|
Verifier 输出:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"verdict": "PASS|LOW_CONFID|REJECT",
|
||||||
|
"groundedness_score": 0.8,
|
||||||
|
"critical_fact_count": 2,
|
||||||
|
"facts_checked": [],
|
||||||
|
"rationale": "..."
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
结果写入:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session.self_evaluation.verifier_evaluation
|
||||||
|
```
|
||||||
|
|
||||||
|
## 7. AIOps 规则门禁
|
||||||
|
|
||||||
|
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
|
||||||
|
|
||||||
|
检查重点:
|
||||||
|
|
||||||
|
- 最终报告是否存在。
|
||||||
|
- payload 模式是否围绕输入告警展开。
|
||||||
|
- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
|
||||||
|
- 是否把无关活跃告警扩展成主诊断对象。
|
||||||
|
|
||||||
|
结果写入:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||||
|
```
|
||||||
|
|
||||||
|
## 8. Eval Baseline
|
||||||
|
|
||||||
|
当前质量门禁还包括离线评测资产:
|
||||||
|
|
||||||
|
| 评测 | 位置 | 作用 |
|
||||||
|
|---|---|---|
|
||||||
|
| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
|
||||||
|
| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
|
||||||
|
| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
|
||||||
|
|
||||||
|
## 9. 后续门禁规划
|
||||||
|
|
||||||
|
从旧版设计继承但尚未完整实现的门禁:
|
||||||
|
|
||||||
|
- 工具参数 schema 校验。
|
||||||
|
- 同一工具调用次数上限。
|
||||||
|
- 工具超时的统一熔断。
|
||||||
|
- 报告中的数值与工具返回值自动对齐校验。
|
||||||
|
- Prompt 版本记录和回滚。
|
||||||
|
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||||||
|
|
||||||
|
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|
||||||
|
|
||||||
@@ -0,0 +1,113 @@
|
|||||||
|
# 面试一页式架构讲解
|
||||||
|
|
||||||
|
**用途**:面试现场 2-5 分钟讲清项目
|
||||||
|
**适合场景**:开场介绍、架构追问、Demo 前铺垫
|
||||||
|
|
||||||
|
## 1. 一句话
|
||||||
|
|
||||||
|
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Verifier 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。
|
||||||
|
|
||||||
|
## 2. 一张图
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
User["用户问题 / AIOps 告警"] --> API["API Layer"]
|
||||||
|
|
||||||
|
API --> Chat["ChatService"]
|
||||||
|
API --> AiOps["AiOpsService"]
|
||||||
|
|
||||||
|
Chat --> ChatFlow["Chat: Planner -> Executor -> Verifier"]
|
||||||
|
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
|
||||||
|
|
||||||
|
ChatFlow --> Tools["Evidence Tools"]
|
||||||
|
AiOpsFlow --> Tools
|
||||||
|
|
||||||
|
Tools --> Knowledge["lookup_knowledge"]
|
||||||
|
Tools --> Logs["query_logs"]
|
||||||
|
Tools --> Metrics["query_metrics / Prometheus"]
|
||||||
|
|
||||||
|
Knowledge --> RAG["RAG: L0 hint + VectorSearchService"]
|
||||||
|
RAG --> VectorStore["Spring AI VectorStore"]
|
||||||
|
RAG --> SDK["Milvus SDK fallback"]
|
||||||
|
|
||||||
|
ChatFlow --> Trace["Trace Persistence"]
|
||||||
|
AiOpsFlow --> Trace
|
||||||
|
Tools --> Trace
|
||||||
|
|
||||||
|
Trace --> Session["diagnosis_session"]
|
||||||
|
Trace --> Step["agent_step"]
|
||||||
|
Trace --> Invocation["tool_invocation"]
|
||||||
|
|
||||||
|
Invocation --> Verifier["Verifier / Rule Evaluation"]
|
||||||
|
Verifier --> SelfEval["self_evaluation"]
|
||||||
|
|
||||||
|
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
|
||||||
|
Step --> TraceAPI
|
||||||
|
Invocation --> TraceAPI
|
||||||
|
SelfEval --> TraceAPI
|
||||||
|
|
||||||
|
TraceAPI --> Feedback["POST /api/feedback"]
|
||||||
|
Feedback --> Case["useful -> case_library"]
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. 面试讲法
|
||||||
|
|
||||||
|
```text
|
||||||
|
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
|
||||||
|
|
||||||
|
Chat 复杂问题走 Planner -> Executor -> Verifier:
|
||||||
|
Planner 负责拆解,Executor 负责调用知识库、日志和指标工具,Verifier 只基于已有工具证据校验最终答案。
|
||||||
|
|
||||||
|
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
|
||||||
|
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
|
||||||
|
|
||||||
|
所有过程都会落到 diagnosis_session、agent_step、tool_invocation。
|
||||||
|
所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Verifier 怎么判定、用户最后是否反馈有用。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 4. 五个亮点
|
||||||
|
|
||||||
|
| 亮点 | 怎么讲 |
|
||||||
|
|---|---|
|
||||||
|
| 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool |
|
||||||
|
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
|
||||||
|
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
|
||||||
|
| 质量门禁 | Chat Verifier 校验 groundedness,AIOps rule evaluation 控制告警聚焦 |
|
||||||
|
| 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 |
|
||||||
|
|
||||||
|
## 5. 三个关键取舍
|
||||||
|
|
||||||
|
### 取舍 1:为什么不用隐式 Advisor 做 RAG?
|
||||||
|
|
||||||
|
因为这个项目强调 Agent 决策可见性。`lookup_knowledge` 必须作为显式工具调用被记录,这样才能解释“什么时候检索、检索了什么、证据如何支撑结论”。
|
||||||
|
|
||||||
|
### 取舍 2:为什么保留 Milvus SDK fallback?
|
||||||
|
|
||||||
|
因为迁移到 Spring AI VectorStore 期间,schema、collection、score 语义都可能变化。`auto` 模式先走 VectorStore,失败时 fallback 到 SDK,保证 MVP 主链路可运行,也方便对比新旧检索质量。
|
||||||
|
|
||||||
|
### 取舍 3:为什么 self_evaluation 分三层?
|
||||||
|
|
||||||
|
因为三类评估回答的问题不同:
|
||||||
|
|
||||||
|
```text
|
||||||
|
rule_evaluation -> 工具证据是否充分
|
||||||
|
verifier_evaluation -> Chat 答案关键事实是否有证据支撑
|
||||||
|
aiops_rule_evaluation -> AIOps 报告是否聚焦告警并使用证据
|
||||||
|
```
|
||||||
|
|
||||||
|
## 6. 面试官可能追问
|
||||||
|
|
||||||
|
| 追问 | 回答方向 |
|
||||||
|
|---|---|
|
||||||
|
| 怎么防止幻觉? | Executor 必须用工具;Verifier 只基于 `tool_trace_summary` 校验;LOW_CONFID/REJECT 会降级输出 |
|
||||||
|
| RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 |
|
||||||
|
| 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter |
|
||||||
|
| AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 |
|
||||||
|
| 下一步怎么演进? | evidence block、邻居 chunk、Playbook、AIOps LLM Verifier、MCP 工具协议化 |
|
||||||
|
|
||||||
|
## 7. 现场演示入口
|
||||||
|
|
||||||
|
- Demo 脚本:`mvp/demo/ten-minute-interview-demo.md`
|
||||||
|
- 故事案例:`interview/story-cases.md`
|
||||||
|
- 架构细节:`mvp/architecture/README.md`
|
||||||
|
|
||||||
@@ -0,0 +1,240 @@
|
|||||||
|
# 知识库文档编写与维护
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:当前建议规范
|
||||||
|
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-usage.md`
|
||||||
|
|
||||||
|
## 1. 定位
|
||||||
|
|
||||||
|
知识库文档不是普通 Markdown 资料堆叠,而是 RAG 检索的输入资产。写得好的文档会提升:
|
||||||
|
|
||||||
|
- L0 hint 的关键词和领域识别。
|
||||||
|
- L1 向量召回质量。
|
||||||
|
- `breadcrumb` 上下文恢复能力。
|
||||||
|
- Verifier 可引用的证据质量。
|
||||||
|
|
||||||
|
当前推荐写法:结构化 Markdown + frontmatter + 明确分类 + 可检索关键词。
|
||||||
|
|
||||||
|
## 2. 文档进入系统的链路
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Markdown["Markdown file"] --> Upload["POST /api/documents/upload"]
|
||||||
|
Upload --> Parse["FrontmatterParser"]
|
||||||
|
Parse --> Enrich["DocumentFieldEnricher"]
|
||||||
|
Enrich --> Metadata["api_document.metadata"]
|
||||||
|
Upload --> Chunk["DocumentChunkService"]
|
||||||
|
Chunk --> Breadcrumb["title / breadcrumb / chunkIndex"]
|
||||||
|
Breadcrumb --> Embedding["VectorIndexService embedding text"]
|
||||||
|
Embedding --> Milvus["Milvus/Zilliz"]
|
||||||
|
Metadata --> L0["KnowledgeIndexService L0 index"]
|
||||||
|
Milvus --> L1["VectorSearchService L1 retrieval"]
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. Frontmatter
|
||||||
|
|
||||||
|
推荐模板:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
---
|
||||||
|
title: 支付网关错误码定义
|
||||||
|
keywords: [ERR_TIMEOUT, 支付超时, payment timeout, 支付网关]
|
||||||
|
summary: 记录支付网关核心错误码的含义、常见原因和排查步骤
|
||||||
|
category: api
|
||||||
|
version: 1.0
|
||||||
|
author: sre-team
|
||||||
|
---
|
||||||
|
|
||||||
|
# 支付网关错误码定义
|
||||||
|
|
||||||
|
...
|
||||||
|
```
|
||||||
|
|
||||||
|
字段说明:
|
||||||
|
|
||||||
|
| 字段 | 必填 | 用途 |
|
||||||
|
|---|---:|---|
|
||||||
|
| `title` | 是 | 文档标题,进入 L0 索引和 embedding 上下文 |
|
||||||
|
| `keywords` | 是 | L0 hint 的主要来源 |
|
||||||
|
| `summary` | 是 | 文档摘要,进入知识域描述和 Agent 上下文 |
|
||||||
|
| `category` | 建议 | 知识域、metadata filter、上传目录 |
|
||||||
|
| `version` | 可选 | 文档版本 |
|
||||||
|
| `author` | 可选 | 维护人 |
|
||||||
|
|
||||||
|
当前解析器会提示缺少 `title`、`keywords`、`summary` 的情况;缺失不一定阻断上传,但会降低检索质量。
|
||||||
|
|
||||||
|
## 4. category 建议
|
||||||
|
|
||||||
|
`category` 会影响:
|
||||||
|
|
||||||
|
- 上传文件本地目录。
|
||||||
|
- Milvus metadata。
|
||||||
|
- L0 domain hint。
|
||||||
|
- `knowledge_domain` 聚合。
|
||||||
|
- VectorStore / SDK category filter。
|
||||||
|
|
||||||
|
推荐保持稳定,不要频繁换名。
|
||||||
|
|
||||||
|
| category | 用途 |
|
||||||
|
|---|---|
|
||||||
|
| `api` | 接口、错误码、请求/响应协议 |
|
||||||
|
| `infrastructure` | MySQL、Redis、JVM、网络、中间件 |
|
||||||
|
| `troubleshooting` | 通用排障流程、Runbook |
|
||||||
|
| `domain` | 业务领域规则 |
|
||||||
|
| `spring-ai` | Spring AI / Agent / 工具最佳实践 |
|
||||||
|
|
||||||
|
注意:分类过细会导致 filter 召回不足;分类过粗会降低 L0 hint 解释力。
|
||||||
|
|
||||||
|
## 5. 关键词写法
|
||||||
|
|
||||||
|
好的关键词应该覆盖:
|
||||||
|
|
||||||
|
- 精确实体:错误码、服务名、指标名。
|
||||||
|
- 常用中文说法。
|
||||||
|
- 英文别名。
|
||||||
|
- 组合词。
|
||||||
|
|
||||||
|
示例:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
keywords: [ERR_TIMEOUT, timeout, 支付超时, 支付网关超时, payment-service, gateway timeout]
|
||||||
|
```
|
||||||
|
|
||||||
|
避免:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
keywords: [错误, 问题, 系统]
|
||||||
|
```
|
||||||
|
|
||||||
|
原因:过宽关键词会让 L0 hint 变脏,多个文档同时命中,影响解释性和 category filter。
|
||||||
|
|
||||||
|
## 6. Markdown 结构
|
||||||
|
|
||||||
|
推荐结构:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# 文档总标题
|
||||||
|
|
||||||
|
## 场景或错误码
|
||||||
|
|
||||||
|
### 含义
|
||||||
|
|
||||||
|
### 常见原因
|
||||||
|
|
||||||
|
### 排查步骤
|
||||||
|
|
||||||
|
### 处理方案
|
||||||
|
|
||||||
|
### 日志示例
|
||||||
|
```
|
||||||
|
|
||||||
|
为什么这样写:
|
||||||
|
|
||||||
|
- `DocumentChunkService` 会按 Markdown 标题切分。
|
||||||
|
- 标题层级会生成 `breadcrumb`。
|
||||||
|
- `title + breadcrumb + content` 会一起进入 embedding 文本。
|
||||||
|
- 命中 chunk 时,Agent 更容易知道证据属于哪个章节。
|
||||||
|
|
||||||
|
## 7. 内容建议
|
||||||
|
|
||||||
|
每个可诊断条目尽量包含:
|
||||||
|
|
||||||
|
- 现象。
|
||||||
|
- 判断条件。
|
||||||
|
- 可能原因。
|
||||||
|
- 证据来源。
|
||||||
|
- 排查步骤。
|
||||||
|
- 处理建议。
|
||||||
|
- 日志或配置示例。
|
||||||
|
|
||||||
|
示例:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
## ERR_TIMEOUT
|
||||||
|
|
||||||
|
### 含义
|
||||||
|
|
||||||
|
支付网关请求超过本地或上游超时时间。
|
||||||
|
|
||||||
|
### 常见原因
|
||||||
|
|
||||||
|
1. 第三方支付服务响应慢。
|
||||||
|
2. 本地 timeout 配置过短。
|
||||||
|
3. 网络链路抖动。
|
||||||
|
|
||||||
|
### 排查步骤
|
||||||
|
|
||||||
|
1. 查询 payment-service 日志中的请求耗时。
|
||||||
|
2. 查看网关 5xx 和 timeout 指标。
|
||||||
|
3. 对比当前 timeout 配置。
|
||||||
|
|
||||||
|
### 处理建议
|
||||||
|
|
||||||
|
- 短期:重试受影响订单。
|
||||||
|
- 长期:调整 timeout 和重试策略,并监控上游延迟。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 8. 上传与索引
|
||||||
|
|
||||||
|
上传接口:
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/documents/upload
|
||||||
|
Content-Type: multipart/form-data
|
||||||
|
|
||||||
|
file=<markdown file>
|
||||||
|
category=<category>
|
||||||
|
```
|
||||||
|
|
||||||
|
系统处理:
|
||||||
|
|
||||||
|
1. 计算文件 hash,避免重复上传。
|
||||||
|
2. 保存原始文件。
|
||||||
|
3. 解析 frontmatter。
|
||||||
|
4. 补全文档字段。
|
||||||
|
5. 写入 `api_document`。
|
||||||
|
6. Markdown-aware chunking。
|
||||||
|
7. 写入 Milvus/Zilliz。
|
||||||
|
8. 更新 L0 索引和 `knowledge_domain`。
|
||||||
|
|
||||||
|
## 9. 重建索引注意事项
|
||||||
|
|
||||||
|
当以下内容变化时,需要重新索引:
|
||||||
|
|
||||||
|
- 正文内容。
|
||||||
|
- 标题层级。
|
||||||
|
- `category`。
|
||||||
|
- `title`、`summary`、`keywords`。
|
||||||
|
- embedding 输入策略,例如加入 `breadcrumb`。
|
||||||
|
|
||||||
|
特别注意:
|
||||||
|
|
||||||
|
```text
|
||||||
|
修改 Markdown 或 embedding 输入策略,不会自动改变已有向量。
|
||||||
|
必须重新上传或重建索引后,live retrieval 才能体现变化。
|
||||||
|
```
|
||||||
|
|
||||||
|
可用 live 验收:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/eval_rag_live_acceptance.py
|
||||||
|
```
|
||||||
|
|
||||||
|
## 10. 维护 checklist
|
||||||
|
|
||||||
|
新增文档前检查:
|
||||||
|
|
||||||
|
- frontmatter 是否包含 `title`、`keywords`、`summary`。
|
||||||
|
- `category` 是否属于现有稳定分类。
|
||||||
|
- 关键词是否既有精确词也有常用表达。
|
||||||
|
- Markdown 标题层级是否清晰。
|
||||||
|
- 每个故障条目是否包含可执行排查步骤。
|
||||||
|
- 日志/配置示例是否脱敏。
|
||||||
|
|
||||||
|
更新文档后检查:
|
||||||
|
|
||||||
|
- `api_document.status` 是否为 `INDEXED`。
|
||||||
|
- `/api/search/similar` 是否能搜到目标文档。
|
||||||
|
- `eval/rag-retrieval` 是否需要新增 golden case。
|
||||||
|
- Trace 中 `tool_invocation` 是否记录到正确 source 和 breadcrumb。
|
||||||
|
|
||||||
@@ -0,0 +1,414 @@
|
|||||||
|
# RAG 新架构
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:当前主架构 + 后续演进边界
|
||||||
|
**关联计划**:`mvp/issues/rag-refactor-plan.md`
|
||||||
|
|
||||||
|
## 1. 架构目标
|
||||||
|
|
||||||
|
RAG 重构的目标不是把所有能力交给框架,也不是继续维护一套完全自研检索框架,而是形成:
|
||||||
|
|
||||||
|
```text
|
||||||
|
成熟框架能力 + 业务可观测编排
|
||||||
|
```
|
||||||
|
|
||||||
|
具体原则:
|
||||||
|
|
||||||
|
- 通用向量检索能力交给 Spring AI `VectorStore`。
|
||||||
|
- 项目保留 Agent Tool 入口、AIOps 业务 query 映射、证据打包、trace 记录。
|
||||||
|
- `lookup_knowledge` 继续是显式工具,不替换成隐式 Advisor。
|
||||||
|
- Spring AI 读取路径作为主路径,Milvus SDK 作为 fallback。
|
||||||
|
- 所有检索行为必须可评测、可回放、可解释。
|
||||||
|
|
||||||
|
## 2. 当前主链路
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Agent["Agent Executor"] --> Tool["lookup_knowledge(query)"]
|
||||||
|
|
||||||
|
Tool --> L0["KnowledgeIndexService.analyzeQuery"]
|
||||||
|
L0 --> Hint["L0 hint: domain / entities / matchedKeywords"]
|
||||||
|
Hint --> Filter["category filter candidate"]
|
||||||
|
|
||||||
|
Tool --> Search["VectorSearchService.searchSimilarDocuments"]
|
||||||
|
Filter --> Search
|
||||||
|
|
||||||
|
Search --> Mode{"retrieval.vector-store.mode"}
|
||||||
|
Mode -->|auto| SpringTry["try Spring AI VectorStore"]
|
||||||
|
SpringTry -->|success| Results["SearchResult list"]
|
||||||
|
SpringTry -->|failure| SdkFallback["Milvus SDK fallback"]
|
||||||
|
Mode -->|spring-ai| SpringOnly["Spring AI VectorStore only"]
|
||||||
|
Mode -->|sdk| SdkOnly["Milvus SDK only"]
|
||||||
|
|
||||||
|
SpringOnly --> Results
|
||||||
|
SdkFallback --> Results
|
||||||
|
SdkOnly --> Results
|
||||||
|
|
||||||
|
Results --> Normalize["relevance normalization"]
|
||||||
|
Normalize --> Dedup["session dedup: RetrievedDocTracker"]
|
||||||
|
Dedup --> Output["LookupResult"]
|
||||||
|
Output --> Record["tool_invocation record"]
|
||||||
|
Output --> Agent
|
||||||
|
```
|
||||||
|
|
||||||
|
```text
|
||||||
|
Agent Executor
|
||||||
|
-> lookup_knowledge(query)
|
||||||
|
-> KnowledgeIndexService.analyzeQuery
|
||||||
|
-> L0 domain/entity hint
|
||||||
|
-> matchedKeywords
|
||||||
|
-> category filter candidate
|
||||||
|
-> VectorSearchService.searchSimilarDocuments
|
||||||
|
-> mode=auto
|
||||||
|
-> Spring AI VectorStore
|
||||||
|
-> fallback: Milvus SDK
|
||||||
|
-> mode=spring-ai
|
||||||
|
-> Spring AI VectorStore only
|
||||||
|
-> mode=sdk
|
||||||
|
-> Milvus SDK only
|
||||||
|
-> result normalization
|
||||||
|
-> relevanceLevel
|
||||||
|
-> completenessHint
|
||||||
|
-> score/rawScore/scoreLabel
|
||||||
|
-> session dedup
|
||||||
|
-> RetrievedDocTracker
|
||||||
|
-> tool_invocation record
|
||||||
|
```
|
||||||
|
|
||||||
|
运行配置:
|
||||||
|
|
||||||
|
```properties
|
||||||
|
retrieval.vector-store.mode=auto
|
||||||
|
retrieval.normalization.max-l2-distance=2.0
|
||||||
|
retrieval.normalization.highly-relevant-threshold=0.75
|
||||||
|
retrieval.normalization.reference-threshold=0.5
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. 稳定边界
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
subgraph AgentBoundary["Agent boundary"]
|
||||||
|
Executor["Executor Agent"]
|
||||||
|
Tool["LookupKnowledgeTool"]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph RetrievalBoundary["Retrieval boundary"]
|
||||||
|
Search["VectorSearchService"]
|
||||||
|
Spring["Spring AI VectorStore"]
|
||||||
|
SDK["Milvus SDK"]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph ObservabilityBoundary["Observability boundary"]
|
||||||
|
Invocation["tool_invocation"]
|
||||||
|
Eval["RAG baseline / trace inspection"]
|
||||||
|
end
|
||||||
|
|
||||||
|
Executor --> Tool
|
||||||
|
Tool --> Search
|
||||||
|
Search --> Spring
|
||||||
|
Search --> SDK
|
||||||
|
Tool --> Invocation
|
||||||
|
Invocation --> Eval
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.1 Agent 边界
|
||||||
|
|
||||||
|
Agent 只知道自己可以调用 `lookup_knowledge`,不直接关心底层是 Spring AI VectorStore 还是 Milvus SDK。
|
||||||
|
|
||||||
|
```text
|
||||||
|
Executor -> LookupKnowledgeTool -> VectorSearchService
|
||||||
|
```
|
||||||
|
|
||||||
|
这个边界让 RAG 底层迁移不影响 Agent prompt、工具声明和 trace 数据结构。
|
||||||
|
|
||||||
|
### 3.2 检索边界
|
||||||
|
|
||||||
|
`VectorSearchService` 是当前检索门面:
|
||||||
|
|
||||||
|
- `auto`:优先 Spring AI VectorStore,失败后 fallback 到 SDK。
|
||||||
|
- `spring-ai`:只走 Spring AI VectorStore。
|
||||||
|
- `sdk`:只走原 Milvus SDK。
|
||||||
|
|
||||||
|
这样可以在不改 Agent 工具的情况下切换检索实现,并支持线上验证和回退。
|
||||||
|
|
||||||
|
### 3.3 可观测边界
|
||||||
|
|
||||||
|
无论底层检索路径如何变化,都必须写入 `tool_invocation`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
sessionId
|
||||||
|
toolName
|
||||||
|
inputParams
|
||||||
|
outputPreview
|
||||||
|
retrievalLayer
|
||||||
|
l0MatchCount
|
||||||
|
l1MatchCount
|
||||||
|
retrievalDetails
|
||||||
|
relevanceLevel
|
||||||
|
dedupReason
|
||||||
|
duration
|
||||||
|
success
|
||||||
|
```
|
||||||
|
|
||||||
|
## 4. L0 的新职责
|
||||||
|
|
||||||
|
旧版 L0 容易承担过重职责,例如唯一匹配后直接跳过 L1。当前架构中 L0 被降级为 hint 层。
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Input["query / AIOps payload"] --> L0["L0 hint analysis"]
|
||||||
|
L0 --> Domain["domain detector"]
|
||||||
|
L0 --> Entity["entity extractor"]
|
||||||
|
L0 --> Keyword["matched keyword explanation"]
|
||||||
|
L0 --> Filter["metadata/category filter candidate"]
|
||||||
|
|
||||||
|
Domain --> Retrieval["L1 semantic retrieval"]
|
||||||
|
Entity --> Retrieval
|
||||||
|
Keyword --> Trace["hit reason in tool_invocation"]
|
||||||
|
Filter --> Retrieval
|
||||||
|
|
||||||
|
Retrieval --> Normalize["relevance normalization"]
|
||||||
|
Normalize --> Evidence["evidence returned to Agent"]
|
||||||
|
```
|
||||||
|
|
||||||
|
L0 负责:
|
||||||
|
|
||||||
|
- domain detector
|
||||||
|
- entity extractor
|
||||||
|
- matched keyword explanation
|
||||||
|
- metadata/category filter candidate
|
||||||
|
- trace 中的 hit reason
|
||||||
|
|
||||||
|
L0 不再默认负责:
|
||||||
|
|
||||||
|
```text
|
||||||
|
L0 unique hit -> 直接作为最终检索结果
|
||||||
|
```
|
||||||
|
|
||||||
|
当前职责是:
|
||||||
|
|
||||||
|
```text
|
||||||
|
query / AIOps payload
|
||||||
|
-> L0 matched keywords / domains / entities
|
||||||
|
-> category filter candidate
|
||||||
|
-> L1 semantic retrieval
|
||||||
|
-> relevance normalization
|
||||||
|
```
|
||||||
|
|
||||||
|
这样既保留精确关键词和领域 hint 的价值,也避免 L0 误召回直接污染最终证据。
|
||||||
|
|
||||||
|
## 5. L1 向量检索
|
||||||
|
|
||||||
|
L1 语义检索通过 `VectorSearchService` 调度。
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Search["VectorSearchService"] --> Request["SearchRequest: query / topK / threshold / filter"]
|
||||||
|
Request --> VectorStore["Spring AI VectorStore"]
|
||||||
|
VectorStore --> Docs["Document results"]
|
||||||
|
Docs --> Map["map to SearchResult"]
|
||||||
|
Map --> Score["score compatibility mapping"]
|
||||||
|
|
||||||
|
Search --> SDK["Milvus SDK fallback"]
|
||||||
|
SDK --> SdkRows["id / content / metadata / L2 distance"]
|
||||||
|
SdkRows --> Map
|
||||||
|
|
||||||
|
Score --> Output["id / content / metadata / score / rawScore / scoreLabel"]
|
||||||
|
```
|
||||||
|
|
||||||
|
### Spring AI VectorStore 路径
|
||||||
|
|
||||||
|
```text
|
||||||
|
SearchRequest
|
||||||
|
-> query
|
||||||
|
-> topK
|
||||||
|
-> similarityThresholdAll
|
||||||
|
-> optional filterExpression: category == '...'
|
||||||
|
-> VectorStore.similaritySearch
|
||||||
|
```
|
||||||
|
|
||||||
|
返回结果会映射为项目兼容结构:
|
||||||
|
|
||||||
|
```text
|
||||||
|
id
|
||||||
|
content
|
||||||
|
metadata
|
||||||
|
score
|
||||||
|
rawScore
|
||||||
|
scoreLabel
|
||||||
|
```
|
||||||
|
|
||||||
|
### Milvus SDK fallback
|
||||||
|
|
||||||
|
SDK 路径仍保留:
|
||||||
|
|
||||||
|
- 用于 `auto` 模式兜底。
|
||||||
|
- 用于与旧链路对比。
|
||||||
|
- 用于 VectorStore 配置或 collection schema 异常时保证 MVP 可运行。
|
||||||
|
|
||||||
|
## 6. 分数语义
|
||||||
|
|
||||||
|
旧 SDK 使用 L2 distance,Spring AI 返回 similarity。两者不能混用为同一个含义。
|
||||||
|
|
||||||
|
当前统一输出:
|
||||||
|
|
||||||
|
| 字段 | 含义 |
|
||||||
|
|---|---|
|
||||||
|
| `score` | 兼容旧逻辑的距离型分数,越小越近 |
|
||||||
|
| `rawScore` | 底层实现的原始分数 |
|
||||||
|
| `scoreLabel` | `l2_distance` 或 `similarity` |
|
||||||
|
|
||||||
|
SDK 路径:
|
||||||
|
|
||||||
|
```text
|
||||||
|
score = L2 distance
|
||||||
|
rawScore = L2 distance
|
||||||
|
scoreLabel = l2_distance
|
||||||
|
```
|
||||||
|
|
||||||
|
VectorStore 路径:
|
||||||
|
|
||||||
|
```text
|
||||||
|
rawScore = Spring AI similarity
|
||||||
|
scoreLabel = similarity
|
||||||
|
score = metadata.distance if available else compatible distance
|
||||||
|
```
|
||||||
|
|
||||||
|
## 7. 文档切片和 embedding 输入
|
||||||
|
|
||||||
|
当前保留 Markdown-aware chunking:
|
||||||
|
|
||||||
|
- 识别 Markdown 标题层级。
|
||||||
|
- 生成 `title`。
|
||||||
|
- 生成 `breadcrumb`。
|
||||||
|
- 保留 `chunkIndex`。
|
||||||
|
- 使用 token 估算和软/硬上限控制 chunk 大小。
|
||||||
|
- 尽量不打断列表和代码块。
|
||||||
|
|
||||||
|
embedding 输入中已经加强:
|
||||||
|
|
||||||
|
```text
|
||||||
|
title + breadcrumb + content
|
||||||
|
```
|
||||||
|
|
||||||
|
这样可以降低单个 chunk 脱离章节上下文后的召回损失。
|
||||||
|
|
||||||
|
## 8. AIOps query 增强
|
||||||
|
|
||||||
|
AIOps payload 中的业务字段不能完全交给通用检索框架隐式理解。
|
||||||
|
|
||||||
|
payload 模式会把以下字段拼成推荐知识库 query:
|
||||||
|
|
||||||
|
- `alertName`
|
||||||
|
- `service`
|
||||||
|
- `severity`
|
||||||
|
- `description`
|
||||||
|
- `timeRange`
|
||||||
|
- `userRequest`
|
||||||
|
|
||||||
|
Prompt 会明确要求 Agent 在需要知识库证据时,优先使用推荐 query 或保留 alertName/service 的更窄 query。
|
||||||
|
|
||||||
|
```text
|
||||||
|
AIOps payload
|
||||||
|
-> buildKnowledgeRetrievalQuery
|
||||||
|
-> Recommended lookup_knowledge query
|
||||||
|
-> lookup_knowledge
|
||||||
|
-> tool_invocation
|
||||||
|
```
|
||||||
|
|
||||||
|
## 9. Evidence 与去重
|
||||||
|
|
||||||
|
当前 evidence 输出仍以 `LookupResult` 和工具返回文本为主,已经具备:
|
||||||
|
|
||||||
|
- L0/L1 命中数量。
|
||||||
|
- 检索层记录。
|
||||||
|
- relevance level。
|
||||||
|
- completeness hint。
|
||||||
|
- session 级文档去重。
|
||||||
|
- domain 行动记忆。
|
||||||
|
- `tool_invocation` 明细记录。
|
||||||
|
|
||||||
|
后续更完整的 evidence block 目标:
|
||||||
|
|
||||||
|
```text
|
||||||
|
source
|
||||||
|
docId
|
||||||
|
chunkIndex
|
||||||
|
title
|
||||||
|
breadcrumb
|
||||||
|
score
|
||||||
|
rawScore
|
||||||
|
scoreLabel
|
||||||
|
hitReason
|
||||||
|
content
|
||||||
|
expandedFrom
|
||||||
|
```
|
||||||
|
|
||||||
|
这部分应作为下一阶段增强,而不是当前已完全完成能力。
|
||||||
|
|
||||||
|
## 10. 评测与验收
|
||||||
|
|
||||||
|
RAG 架构变更必须先过评测,再认为可合入主链路。
|
||||||
|
|
||||||
|
当前评测资产:
|
||||||
|
|
||||||
|
- `eval/rag-retrieval/cases/golden-cases.json`
|
||||||
|
- `eval/rag-retrieval/fixtures/`
|
||||||
|
- `eval/rag-retrieval/reports/baseline.json`
|
||||||
|
- `eval/rag-retrieval/reports/baseline.md`
|
||||||
|
- `scripts/eval_rag_retrieval.py`
|
||||||
|
- `scripts/eval_rag_live_acceptance.py`
|
||||||
|
|
||||||
|
评测层次:
|
||||||
|
|
||||||
|
| 层次 | 作用 |
|
||||||
|
|---|---|
|
||||||
|
| Offline baseline | 不依赖 MySQL、Redis、Milvus、LLM,用固定 fixtures 检查召回行为 |
|
||||||
|
| Live acceptance | 应用运行并重建索引后,调用 `/api/search/similar` 验证真实检索 |
|
||||||
|
| Trace inspection | 通过 `tool_invocation` 检查 Agent 是否真的使用了证据 |
|
||||||
|
|
||||||
|
## 11. 当前已完成
|
||||||
|
|
||||||
|
- `lookup_knowledge` 保持显式 Agent Tool。
|
||||||
|
- L0 降级为 domain/entity hint。
|
||||||
|
- L1 默认执行语义检索。
|
||||||
|
- `VectorSearchService` 支持 `auto`、`spring-ai`、`sdk` 三种模式。
|
||||||
|
- Spring AI VectorStore 成为读取主路径。
|
||||||
|
- Milvus SDK fallback 保留。
|
||||||
|
- 分数语义拆成 `score`、`rawScore`、`scoreLabel`。
|
||||||
|
- Markdown chunk 保留 `title` 和 `breadcrumb`。
|
||||||
|
- embedding 输入包含 `title`、`breadcrumb` 和 `content`。
|
||||||
|
- AIOps payload 生成推荐知识库 query。
|
||||||
|
- `tool_invocation` 记录 relevance level 和 dedup reason。
|
||||||
|
- RAG offline baseline 和 live acceptance 脚本已补齐。
|
||||||
|
|
||||||
|
## 12. 后续演进
|
||||||
|
|
||||||
|
近期优先:
|
||||||
|
|
||||||
|
1. 完整 evidence block 结构化输出。
|
||||||
|
2. 命中 chunk 的相邻 chunk / 同章节上下文扩展。
|
||||||
|
3. metadata taxonomy 清理,例如 `database` 与 `infrastructure` 的分类边界。
|
||||||
|
4. Query Transformer / MultiQuery 的可回退接入。
|
||||||
|
5. VectorStore 写入路径评估。
|
||||||
|
|
||||||
|
暂不优先:
|
||||||
|
|
||||||
|
- 把 `lookup_knowledge` 替换成隐式 Advisor。
|
||||||
|
- 完整自研 RRF 框架。
|
||||||
|
- 立即引入 Elasticsearch / OpenSearch。
|
||||||
|
- 立即引入 cross-encoder 或 LLM rerank。
|
||||||
|
|
||||||
|
## 13. 关键代码索引
|
||||||
|
|
||||||
|
| 能力 | 代码 |
|
||||||
|
|---|---|
|
||||||
|
| Agent 工具入口 | `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java` |
|
||||||
|
| L0 hint | `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java` |
|
||||||
|
| 向量检索门面 | `src/main/java/com/superbiz/agent/service/VectorSearchService.java` |
|
||||||
|
| 文档切片 | `src/main/java/com/superbiz/agent/service/DocumentChunkService.java` |
|
||||||
|
| 文档管理 | `src/main/java/com/superbiz/agent/service/DocumentManagementService.java` |
|
||||||
|
| 向量写入 | `src/main/java/com/superbiz/agent/service/VectorIndexService.java` |
|
||||||
|
| AIOps query 增强 | `src/main/java/com/superbiz/agent/service/AiOpsService.java` |
|
||||||
|
| 工具调用记录 | `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java` |
|
||||||
@@ -0,0 +1,266 @@
|
|||||||
|
# 检索与可观测性架构
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:当前可运行架构
|
||||||
|
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-architecture.md`
|
||||||
|
|
||||||
|
## 1. 定位
|
||||||
|
|
||||||
|
本文补充 [rag-architecture.md](rag-architecture.md) 中的检索细节,重点回答:
|
||||||
|
|
||||||
|
- 查询如何进入 `lookup_knowledge`。
|
||||||
|
- L0 和 L1 当前分别承担什么职责。
|
||||||
|
- 检索结果如何归一化、去重、记录。
|
||||||
|
- 如何通过 trace 和 eval 判断检索质量。
|
||||||
|
|
||||||
|
当前架构与旧版最大的差异是:L0 不再因为唯一命中而默认跳过 L1。L0 是 hint 和解释信号,L1 语义检索是默认召回路径。
|
||||||
|
|
||||||
|
## 2. 检索总图
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Query["Agent query / AIOps recommended query"] --> Tool["LookupKnowledgeTool"]
|
||||||
|
|
||||||
|
Tool --> L0["KnowledgeIndexService.analyzeQuery"]
|
||||||
|
L0 --> L0Result["L0 hint: matches / domains / keywords"]
|
||||||
|
L0Result --> Filter["singleDomainOrNull -> category filter"]
|
||||||
|
|
||||||
|
Tool --> L1["VectorSearchService.searchSimilarDocuments"]
|
||||||
|
Filter --> L1
|
||||||
|
L1 --> Mode{"retrieval.vector-store.mode"}
|
||||||
|
Mode -->|auto| Spring["Spring AI VectorStore"]
|
||||||
|
Spring -->|failure| SDK["Milvus SDK fallback"]
|
||||||
|
Mode -->|spring-ai| Spring
|
||||||
|
Mode -->|sdk| SDK
|
||||||
|
|
||||||
|
Spring --> Candidates["L1 candidates"]
|
||||||
|
SDK --> Candidates
|
||||||
|
Candidates --> Normalize["relevance normalization"]
|
||||||
|
L0Result --> Normalize
|
||||||
|
Normalize --> Result["LookupResult"]
|
||||||
|
|
||||||
|
Result --> Dedup["RetrievedDocTracker session dedup"]
|
||||||
|
Dedup --> Final["final tool output"]
|
||||||
|
Final --> Invocation["tool_invocation"]
|
||||||
|
Final --> Agent["Agent Executor"]
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. L0 Hint 层
|
||||||
|
|
||||||
|
L0 的输入是原始 query,输出是解释性结构:
|
||||||
|
|
||||||
|
```text
|
||||||
|
matches
|
||||||
|
matchedKeywords
|
||||||
|
domains
|
||||||
|
singleDomainOrNull
|
||||||
|
```
|
||||||
|
|
||||||
|
当前职责:
|
||||||
|
|
||||||
|
| 职责 | 说明 |
|
||||||
|
|---|---|
|
||||||
|
| domain hint | 判断 query 可能属于哪个知识域 |
|
||||||
|
| entity / keyword hint | 记录命中的关键词、错误码、服务名等 |
|
||||||
|
| category filter candidate | 当只有单一领域时,给 L1 一个 metadata filter 候选 |
|
||||||
|
| trace explanation | 写入 `tool_invocation.retrieval_details`,用于解释检索为什么这么走 |
|
||||||
|
|
||||||
|
不再承担:
|
||||||
|
|
||||||
|
```text
|
||||||
|
matches=1 -> skip L1 -> 直接返回 L0 文档正文
|
||||||
|
```
|
||||||
|
|
||||||
|
原因:
|
||||||
|
|
||||||
|
- 子串命中不等价于最终相关性。
|
||||||
|
- L0 没有稳定排序和语义相似度。
|
||||||
|
- AIOps query 往往包含多个字段,单点关键词命中容易误导。
|
||||||
|
|
||||||
|
## 4. L1 语义检索层
|
||||||
|
|
||||||
|
L1 通过 `VectorSearchService` 调度,支持三种模式:
|
||||||
|
|
||||||
|
| 模式 | 行为 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| `auto` | 优先 Spring AI VectorStore,失败 fallback 到 SDK | 默认运行模式 |
|
||||||
|
| `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
|
||||||
|
| `sdk` | 只走 Milvus SDK | 对比旧链路或临时回退 |
|
||||||
|
|
||||||
|
### Spring AI VectorStore 路径
|
||||||
|
|
||||||
|
```text
|
||||||
|
SearchRequest
|
||||||
|
-> query
|
||||||
|
-> topK
|
||||||
|
-> similarityThresholdAll
|
||||||
|
-> optional filterExpression
|
||||||
|
-> VectorStore.similaritySearch
|
||||||
|
```
|
||||||
|
|
||||||
|
### Milvus SDK fallback
|
||||||
|
|
||||||
|
```text
|
||||||
|
query
|
||||||
|
-> VectorEmbeddingService.generateQueryVector
|
||||||
|
-> Milvus search(vector, topK, L2)
|
||||||
|
-> id / content / metadata
|
||||||
|
```
|
||||||
|
|
||||||
|
SDK fallback 保留的价值:
|
||||||
|
|
||||||
|
- VectorStore bean 缺失时不让 MVP 主链路中断。
|
||||||
|
- Spring AI collection/schema 配置异常时可回退。
|
||||||
|
- 便于 SDK 与 VectorStore 的结果对比。
|
||||||
|
|
||||||
|
## 5. 分数与相关性归一化
|
||||||
|
|
||||||
|
检索结果输出三类分数字段:
|
||||||
|
|
||||||
|
| 字段 | 说明 |
|
||||||
|
|---|---|
|
||||||
|
| `score` | 兼容旧逻辑的距离型分数 |
|
||||||
|
| `rawScore` | 底层检索实现原始分数 |
|
||||||
|
| `scoreLabel` | 原始分数语义,例如 `similarity` 或 `l2_distance` |
|
||||||
|
|
||||||
|
工具层再把 L0/L1 情况归一为:
|
||||||
|
|
||||||
|
| relevanceLevel | 含义 |
|
||||||
|
|---|---|
|
||||||
|
| `PRECISE` | L0 单命中且 L1 相似度高 |
|
||||||
|
| `HIGHLY_RELEVANT` | L1 相似度高,或 L0 多命中且 L1 支撑强 |
|
||||||
|
| `REFERENCE` | 可作为参考,但不足以声明强证据 |
|
||||||
|
| `DEDUPED` | 同 session 中已检索过,不重复注入上下文 |
|
||||||
|
|
||||||
|
归一化结果用于:
|
||||||
|
|
||||||
|
- 给 Agent 输出 completeness hint。
|
||||||
|
- 写入 `tool_invocation.relevance_level`。
|
||||||
|
- 给 Verifier 构造 `tool_trace_summary`。
|
||||||
|
- 供 EvaluationService 计算 evidence score。
|
||||||
|
|
||||||
|
## 6. 文档切片和 metadata
|
||||||
|
|
||||||
|
当前保留 Markdown-aware chunking。
|
||||||
|
|
||||||
|
关键 metadata:
|
||||||
|
|
||||||
|
```text
|
||||||
|
docId
|
||||||
|
chunkIndex
|
||||||
|
totalChunks
|
||||||
|
title
|
||||||
|
breadcrumb
|
||||||
|
category
|
||||||
|
source
|
||||||
|
```
|
||||||
|
|
||||||
|
embedding 输入已经增强为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
title + breadcrumb + content
|
||||||
|
```
|
||||||
|
|
||||||
|
这解决旧版检索中的一个主要问题:单个 chunk 被召回后,LLM 不知道它属于哪个文档、哪个章节。
|
||||||
|
|
||||||
|
## 7. 输出和记录
|
||||||
|
|
||||||
|
`lookup_knowledge` 的输出会进入两条路径:
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
LookupResult["LookupResult"] --> Agent["Agent context"]
|
||||||
|
LookupResult --> Recorder["ToolInvocationRecorder"]
|
||||||
|
Recorder --> Invocation["tool_invocation"]
|
||||||
|
Invocation --> Trace["DiagnosisTraceService"]
|
||||||
|
Invocation --> Summary["ToolTraceSummaryService"]
|
||||||
|
Summary --> Verifier["chat_verifier"]
|
||||||
|
Invocation --> Eval["EvaluationService / RAG eval"]
|
||||||
|
```
|
||||||
|
|
||||||
|
`tool_invocation` 中与检索相关的字段:
|
||||||
|
|
||||||
|
```text
|
||||||
|
retrieval_layer
|
||||||
|
l0_match_count
|
||||||
|
l1_match_count
|
||||||
|
retrieval_details
|
||||||
|
relevance_level
|
||||||
|
dedup_reason
|
||||||
|
output_preview
|
||||||
|
duration_ms
|
||||||
|
success
|
||||||
|
```
|
||||||
|
|
||||||
|
`retrieval_details` 承载更细信息,例如:
|
||||||
|
|
||||||
|
- L0 命中文档标题和路径。
|
||||||
|
- L1 分数。
|
||||||
|
- retrieved domains。
|
||||||
|
- evidence status。
|
||||||
|
- dedup reason。
|
||||||
|
|
||||||
|
## 8. 去重与行动记忆
|
||||||
|
|
||||||
|
当前 session 级去重由 `RetrievedDocTracker` 负责。
|
||||||
|
|
||||||
|
```text
|
||||||
|
sessionId + docKey
|
||||||
|
-> already retrieved?
|
||||||
|
-> yes: return dedup message and record dedup_reason
|
||||||
|
-> no: mark retrieved and return evidence
|
||||||
|
```
|
||||||
|
|
||||||
|
去重目的:
|
||||||
|
|
||||||
|
- 避免同一文档反复进入上下文。
|
||||||
|
- 降低 token 浪费。
|
||||||
|
- 给 Executor 一个“这个方向已经查过”的行动记忆。
|
||||||
|
|
||||||
|
注意:去重不是全局缓存,只在当前诊断 session 内生效。
|
||||||
|
|
||||||
|
## 9. 检索质量评测
|
||||||
|
|
||||||
|
检索质量不能只看一次接口返回,需要用固定 query 回归。
|
||||||
|
|
||||||
|
当前评测资产:
|
||||||
|
|
||||||
|
| 资产 | 用途 |
|
||||||
|
|---|---|
|
||||||
|
| `eval/rag-retrieval/cases/golden-cases.json` | 固定 query 和期望证据 |
|
||||||
|
| `eval/rag-retrieval/fixtures/` | 离线候选结果 |
|
||||||
|
| `eval/rag-retrieval/reports/baseline.md` | 人类可读基线 |
|
||||||
|
| `scripts/eval_rag_retrieval.py` | 离线回归 |
|
||||||
|
| `scripts/eval_rag_live_acceptance.py` | 运行环境验收 |
|
||||||
|
|
||||||
|
评测层次:
|
||||||
|
|
||||||
|
```text
|
||||||
|
offline baseline
|
||||||
|
-> 不依赖服务和外部组件
|
||||||
|
|
||||||
|
live acceptance
|
||||||
|
-> 调用 /api/search/similar
|
||||||
|
-> 验证重建索引后的真实检索
|
||||||
|
|
||||||
|
trace inspection
|
||||||
|
-> 检查 Agent 是否真的调用 lookup_knowledge
|
||||||
|
-> 检查 tool_invocation 证据是否完整
|
||||||
|
```
|
||||||
|
|
||||||
|
## 10. 后续增强
|
||||||
|
|
||||||
|
近期优先:
|
||||||
|
|
||||||
|
1. 完整 evidence block 输出。
|
||||||
|
2. 邻居 chunk / 同章节上下文扩展。
|
||||||
|
3. metadata taxonomy 清理。
|
||||||
|
4. Query Transformer / MultiQuery 可回退接入。
|
||||||
|
5. 更完整的 Recall@K、MRR、nDCG 报告。
|
||||||
|
|
||||||
|
暂不优先:
|
||||||
|
|
||||||
|
- 重新引入 L0 直接返回。
|
||||||
|
- 一次性迁移所有写入路径。
|
||||||
|
- 在没有评测收益前引入 rerank / RRF / BM25。
|
||||||
|
|
||||||
@@ -0,0 +1,196 @@
|
|||||||
|
# 会话与 Trace 生命周期
|
||||||
|
|
||||||
|
**更新日期**:2026-07-05
|
||||||
|
**状态**:当前可运行架构
|
||||||
|
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||||
|
|
||||||
|
## 1. 定位
|
||||||
|
|
||||||
|
旧版会话设计以 Redis 会话为主,MySQL 作为可选长期沉淀。当前 MVP 的可追踪诊断已经转为 MySQL Trace 三表为主:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session
|
||||||
|
-> agent_step
|
||||||
|
-> tool_invocation
|
||||||
|
```
|
||||||
|
|
||||||
|
因此本文描述的是当前可运行链路:
|
||||||
|
|
||||||
|
- `sessionId` 是一次诊断和后续 trace/feedback 的关联键。
|
||||||
|
- `diagnosis_session` 保存会话级状态、问题、答案、自评估和反馈。
|
||||||
|
- `agent_step` 保存每个 Agent 模型调用。
|
||||||
|
- `tool_invocation` 保存工具调用事实。
|
||||||
|
- `DiagnosisTraceService` 聚合三类记录,形成可回放 trace。
|
||||||
|
|
||||||
|
## 2. 生命周期总图
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
|
||||||
|
Resolve --> Create["create or reset diagnosis_session"]
|
||||||
|
Create --> Running["status = RUNNING"]
|
||||||
|
|
||||||
|
Running --> Agent["Agent workflow"]
|
||||||
|
Agent --> StepHook["AgentLoggingHook"]
|
||||||
|
StepHook --> Step["agent_step"]
|
||||||
|
Agent --> Tool["Evidence tools"]
|
||||||
|
Tool --> Invocation["tool_invocation"]
|
||||||
|
|
||||||
|
Agent --> Final{"workflow result"}
|
||||||
|
Final -->|success| Success["status = SUCCESS, answer saved"]
|
||||||
|
Final -->|failed| Failed["status = FAILED"]
|
||||||
|
|
||||||
|
Success --> Evaluation["self_evaluation merge"]
|
||||||
|
Failed --> Evaluation
|
||||||
|
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace"]
|
||||||
|
Success --> Feedback["POST /api/feedback"]
|
||||||
|
Feedback --> Case["useful -> case_library"]
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. sessionId 规则
|
||||||
|
|
||||||
|
| 链路 | sessionId 来源 |
|
||||||
|
|---|---|
|
||||||
|
| Chat | 如果请求带 sessionId,则复用;否则生成短 UUID |
|
||||||
|
| AIOps | 如果 payload 带 sessionId,则复用;否则生成 UUID |
|
||||||
|
| Trace | URL path 中的 `{sessionId}` |
|
||||||
|
| Feedback | request body 中的 `sessionId` |
|
||||||
|
|
||||||
|
设计含义:
|
||||||
|
|
||||||
|
- 同一个 `sessionId` 可以贯穿诊断、trace 查询和用户反馈。
|
||||||
|
- 当前诊断开始时会重置当前 session 的运行态字段,例如 answer、duration、step/tool count。
|
||||||
|
- `sessionId` 是业务关联键,不依赖数据库自增 ID 暴露给外部。
|
||||||
|
|
||||||
|
## 4. 状态流转
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
stateDiagram-v2
|
||||||
|
[*] --> PENDING
|
||||||
|
PENDING --> RUNNING: start diagnosis
|
||||||
|
RUNNING --> SUCCESS: workflow completed
|
||||||
|
RUNNING --> FAILED: exception / empty state
|
||||||
|
SUCCESS --> SUCCESS: feedback submitted
|
||||||
|
FAILED --> FAILED: feedback submitted
|
||||||
|
```
|
||||||
|
|
||||||
|
字段边界:
|
||||||
|
|
||||||
|
| 字段 | 含义 |
|
||||||
|
|---|---|
|
||||||
|
| `status` | 执行状态:`PENDING` / `RUNNING` / `SUCCESS` / `FAILED` |
|
||||||
|
| `answer` | Agent 最终返回给用户的报告或答复 |
|
||||||
|
| `self_evaluation` | 系统自评估 JSON |
|
||||||
|
| `feedback` | 用户反馈:`useful` / `not_useful` / null |
|
||||||
|
|
||||||
|
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 session,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||||
|
|
||||||
|
## 5. agent_step 写入
|
||||||
|
|
||||||
|
`AgentLoggingHook` 在模型调用前后写入和回填 `agent_step`。
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
sequenceDiagram
|
||||||
|
autonumber
|
||||||
|
participant Agent as ReactAgent
|
||||||
|
participant Hook as AgentLoggingHook
|
||||||
|
participant DB as agent_step
|
||||||
|
|
||||||
|
Agent->>Hook: before_model(messages, sessionId)
|
||||||
|
Hook->>DB: insert step_index / agent_name / model_input
|
||||||
|
Agent-->>Agent: model call
|
||||||
|
Agent->>Hook: after_model(messages, sessionId)
|
||||||
|
Hook->>DB: update model_output / thought / has_tool_call / duration / token_count
|
||||||
|
```
|
||||||
|
|
||||||
|
当前记录:
|
||||||
|
|
||||||
|
- `session_id`
|
||||||
|
- `step_index`
|
||||||
|
- `agent_name`
|
||||||
|
- `model_input`
|
||||||
|
- `model_output`
|
||||||
|
- `thought`
|
||||||
|
- `has_tool_call`
|
||||||
|
- `duration_ms`
|
||||||
|
- `token_count`
|
||||||
|
|
||||||
|
## 6. tool_invocation 写入
|
||||||
|
|
||||||
|
工具调用记录真实工具事实,不记录模型猜测。
|
||||||
|
|
||||||
|
关键字段:
|
||||||
|
|
||||||
|
```text
|
||||||
|
session_id
|
||||||
|
step_id
|
||||||
|
tool_name
|
||||||
|
input_params
|
||||||
|
output_preview
|
||||||
|
output_length
|
||||||
|
retrieval_layer
|
||||||
|
l0_match_count
|
||||||
|
l1_match_count
|
||||||
|
retrieval_details
|
||||||
|
relevance_level
|
||||||
|
dedup_reason
|
||||||
|
duration_ms
|
||||||
|
success
|
||||||
|
error_message
|
||||||
|
```
|
||||||
|
|
||||||
|
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对非检索工具,检索字段可以为空。
|
||||||
|
|
||||||
|
## 7. Trace API 聚合
|
||||||
|
|
||||||
|
```text
|
||||||
|
GET /api/diagnosis/{sessionId}/trace
|
||||||
|
```
|
||||||
|
|
||||||
|
聚合逻辑:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session by sessionId
|
||||||
|
+ agent_step ordered by step_index
|
||||||
|
+ tool_invocation ordered by id
|
||||||
|
-> DiagnosisTraceResponse
|
||||||
|
```
|
||||||
|
|
||||||
|
Trace 视图回答的问题:
|
||||||
|
|
||||||
|
- 这次诊断是否成功?
|
||||||
|
- 哪些 Agent 参与了?
|
||||||
|
- 每一步模型输入输出是什么摘要?
|
||||||
|
- 调用了哪些工具?
|
||||||
|
- 工具返回了什么证据?
|
||||||
|
- Verifier / AIOps rule 是否通过?
|
||||||
|
- 用户是否反馈有用?
|
||||||
|
|
||||||
|
## 8. Chat 与 AIOps 差异
|
||||||
|
|
||||||
|
| 维度 | Chat | AIOps |
|
||||||
|
|---|---|---|
|
||||||
|
| `agent_flow` | `CHAT` | `AI_OPS` |
|
||||||
|
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Verifier | `SupervisorAgent`: Planner + Executor |
|
||||||
|
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||||
|
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||||
|
| payload | 用户自然语言 + history | alert payload 或 auto-discovery |
|
||||||
|
|
||||||
|
## 9. 清理与边界
|
||||||
|
|
||||||
|
当前会话持久化边界:
|
||||||
|
|
||||||
|
- MySQL Trace 记录是主要可回放来源。
|
||||||
|
- Chat 历史仍可作为请求上下文传入 Agent,但不是本文档的主持久化模型。
|
||||||
|
- Redis 主会话存储是历史设计,不作为当前架构事实。
|
||||||
|
- `RetrievedDocTracker` 是 session 级运行时去重状态,诊断结束后清理。
|
||||||
|
|
||||||
|
## 10. 后续增强
|
||||||
|
|
||||||
|
可考虑:
|
||||||
|
|
||||||
|
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
||||||
|
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
||||||
|
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。
|
||||||
|
4. 为 Trace 增加导出能力,服务面试演示和回归分析。
|
||||||
|
|
||||||
+61
-57
@@ -1,41 +1,42 @@
|
|||||||
# MVP Demo Runbook
|
# MVP 演示手册
|
||||||
|
|
||||||
This demo proves the MVP flow from user question to persisted diagnosis trace.
|
本目录用于演示 MVP 从用户问题到诊断 Trace 的完整闭环。
|
||||||
|
|
||||||
For interview use, start with:
|
面试时建议先读:
|
||||||
|
|
||||||
- `interview-walkthrough.md` for the talk track
|
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||||
- `trace-inspection-checklist.md` for fields to inspect
|
- `interview-walkthrough.md`:面试讲解话术。
|
||||||
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
|
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||||
- `requests/payment-timeout-chat.json` for the fixed request payload
|
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||||
|
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||||
|
|
||||||
## Prerequisites
|
## 1. 前置条件
|
||||||
|
|
||||||
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
|
- MySQL、Redis、Milvus/Zilliz、LLM 和 embedding 配置可用。
|
||||||
- Security and secret cleanup are intentionally out of scope for this MVP slice.
|
- 安全和密钥清理不属于当前 MVP 演示范围。
|
||||||
- The `mvp-demo` profile enables mock Prometheus and CLS providers so log and metric tools can return repeatable evidence.
|
- `mvp-demo` profile 会启用 mock Prometheus 和 mock CLS,让日志和指标工具返回可复现证据。
|
||||||
|
|
||||||
## Start
|
## 2. 启动服务
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||||
```
|
```
|
||||||
|
|
||||||
The service listens on:
|
服务地址:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
http://localhost:9900
|
http://localhost:9900
|
||||||
```
|
```
|
||||||
|
|
||||||
## 1. Run Chat Diagnosis
|
## 3. Chat 诊断 Demo
|
||||||
|
|
||||||
Fast path:
|
最快方式:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||||
```
|
```
|
||||||
|
|
||||||
This writes:
|
脚本会生成:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
mvp/demo/output/chat-response.json
|
mvp/demo/output/chat-response.json
|
||||||
@@ -43,7 +44,7 @@ mvp/demo/output/trace-response.json
|
|||||||
mvp/demo/output/feedback-response.json
|
mvp/demo/output/feedback-response.json
|
||||||
```
|
```
|
||||||
|
|
||||||
Manual path:
|
手动请求:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
$sessionId = "mvp-demo-payment-timeout-001"
|
$sessionId = "mvp-demo-payment-timeout-001"
|
||||||
@@ -59,13 +60,13 @@ Invoke-RestMethod `
|
|||||||
-Body $body
|
-Body $body
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected result:
|
期望结果:
|
||||||
|
|
||||||
- `data.success` is `true`.
|
- `data.success = true`
|
||||||
- `data.sessionId` equals `mvp-demo-payment-timeout-001`.
|
- `data.sessionId = mvp-demo-payment-timeout-001`
|
||||||
- `data.answer` contains a diagnosis answer.
|
- `data.answer` 包含诊断答复
|
||||||
|
|
||||||
## 2. Query Trace
|
## 4. 查询 Trace
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
@@ -73,15 +74,15 @@ Invoke-RestMethod `
|
|||||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected result:
|
期望结果:
|
||||||
|
|
||||||
- `code` is `200`.
|
- `code = 200`
|
||||||
- `data.session.sessionId` equals the chat session id.
|
- `data.session.sessionId` 等于 Chat session id
|
||||||
- `data.steps` contains planner/executor/verifier records for complex questions.
|
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||||
- `data.toolInvocations` contains evidence tool calls such as `lookup_knowledge`, `query_logs`, or `query_metrics`.
|
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||||
- `data.session.selfEvaluation` contains verifier or rule evaluation when available.
|
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||||
|
|
||||||
## 3. Submit Feedback
|
## 5. 提交反馈
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
$feedback = @{
|
$feedback = @{
|
||||||
@@ -96,12 +97,13 @@ Invoke-RestMethod `
|
|||||||
-Body $feedback
|
-Body $feedback
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected result:
|
期望结果:
|
||||||
|
|
||||||
- `success` is `true`.
|
- `success = true`
|
||||||
- A later trace query shows `data.session.feedback` as `useful`.
|
- 后续 Trace 中 `data.session.feedback = useful`
|
||||||
|
- useful 反馈会尝试沉淀 `case_library`
|
||||||
|
|
||||||
## 4. Run AIOps Alert Diagnosis
|
## 6. AIOps 告警诊断 Demo
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
|
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
|
||||||
@@ -122,15 +124,16 @@ Invoke-WebRequest `
|
|||||||
-Body $aiopsBody
|
-Body $aiopsBody
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected result:
|
期望结果:
|
||||||
|
|
||||||
- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`.
|
- SSE 首条包含 `session` 消息,sessionId 为 `mvp-demo-aiops-payment-cpu-001`
|
||||||
- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload.
|
- 后续流式输出包含 AIOps 告警分析报告
|
||||||
- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`.
|
- 报告聚焦输入的 `HighCPUUsage/payment-service`
|
||||||
- `data.session.answer` contains the final alert analysis report when a report is generated.
|
- 同一 session 的 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||||
- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them.
|
- `data.session.answer` 包含最终告警报告
|
||||||
|
- `data.toolInvocations` 包含证据工具调用
|
||||||
|
|
||||||
Query the AIOps trace:
|
查询 AIOps Trace:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
@@ -138,28 +141,29 @@ Invoke-RestMethod `
|
|||||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||||
```
|
```
|
||||||
|
|
||||||
## Demo Story
|
## 7. Demo 主线
|
||||||
|
|
||||||
The important interview story is:
|
Chat 主线:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
one session id
|
一个 session id
|
||||||
-> user question
|
-> 用户问题
|
||||||
-> multi-agent execution
|
-> 多 Agent 执行
|
||||||
-> evidence tools
|
-> 证据工具
|
||||||
-> verifier/self-evaluation
|
-> Verifier / self_evaluation
|
||||||
-> final answer
|
-> 最终答案
|
||||||
-> feedback
|
-> 用户反馈
|
||||||
-> trace API for replay and audit
|
-> Trace API 回放
|
||||||
```
|
```
|
||||||
|
|
||||||
The AIOps story uses the same audit spine:
|
AIOps 主线:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
one session id
|
一个 session id
|
||||||
-> alert payload
|
-> 告警 payload
|
||||||
-> AIOps planner/executor execution
|
-> AIOps Planner / Executor
|
||||||
-> evidence tools
|
-> 证据工具
|
||||||
-> alert analysis report
|
-> 告警分析报告
|
||||||
-> trace API for replay and audit
|
-> AIOps rule evaluation
|
||||||
|
-> Trace API 回放
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -1,15 +1,15 @@
|
|||||||
# AIOps Alert Acceptance Case
|
# AIOps 告警验收用例
|
||||||
|
|
||||||
## Goal
|
## 1. 目标
|
||||||
|
|
||||||
Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry.
|
验证旧版 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
||||||
|
|
||||||
## Input
|
## 2. 输入
|
||||||
|
|
||||||
- Session id: `mvp-demo-aiops-payment-cpu-001`
|
- Session id:`mvp-demo-aiops-payment-cpu-001`
|
||||||
- Endpoint: `POST /api/ai_ops`
|
- Endpoint:`POST /api/ai_ops`
|
||||||
- Profile: `mvp-demo`
|
- Profile:`mvp-demo`
|
||||||
- Alert:
|
- 告警 payload:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@@ -23,16 +23,20 @@ Validate that the legacy AIOps endpoint can act as a traceable alert-triggered d
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
## Acceptance Criteria
|
## 3. 验收标准
|
||||||
|
|
||||||
1. The SSE stream emits a `session` message containing the requested session id.
|
1. SSE 流输出 `session` 消息,且包含请求中的 session id。
|
||||||
2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`.
|
2. AIOps 执行创建或更新 `diagnosis_session`,并写入 `agent_flow = AI_OPS`。
|
||||||
3. The persisted session query contains the alert name, service, severity, time range, and description.
|
3. 持久化的 session query 包含告警名、服务名、等级、时间范围和描述。
|
||||||
4. If a final report is generated, `diagnosis_session.answer` contains that report.
|
4. 如果生成最终报告,`diagnosis_session.answer` 包含该报告。
|
||||||
5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations.
|
5. `GET /api/diagnosis/{sessionId}/trace` 返回 AIOps session、按顺序排列的 agent steps 和 tool invocations。
|
||||||
6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections.
|
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
|
||||||
|
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
|
||||||
|
8. `self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
||||||
|
|
||||||
## Known Limits
|
## 4. 已知边界
|
||||||
|
|
||||||
|
- 当前 AIOps 使用轻量规则评估器,不是完整 LLM Verifier。
|
||||||
|
- 完整运行仍依赖有效的 DB、Redis、Milvus/Zilliz、模型和 embedding 配置。
|
||||||
|
- `mvp-demo` profile 使用 mock Prometheus 和 mock CLS,主要用于稳定演示。
|
||||||
|
|
||||||
- This slice does not add a Verifier Agent to AIOps.
|
|
||||||
- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration.
|
|
||||||
|
|||||||
@@ -1,146 +1,144 @@
|
|||||||
# Interview Walkthrough: MVP Diagnosis Agent
|
# 面试演示讲解稿
|
||||||
|
|
||||||
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
|
这是一份短时间 Agent 工程面试用讲解稿,不是完整系统文档。
|
||||||
|
|
||||||
## 30-Second Summary
|
## 1. 30 秒摘要
|
||||||
|
|
||||||
```text
|
```text
|
||||||
This is an enterprise diagnosis Agent MVP.
|
这是一个企业故障诊断 Agent MVP。
|
||||||
It takes a payment-timeout question, plans the investigation, calls evidence tools,
|
它接收支付超时问题,规划排查步骤,调用证据工具,
|
||||||
checks the answer through a verifier, persists the full trace, and accepts feedback.
|
用 Verifier 检查答案,把完整 Trace 持久化,并支持用户反馈。
|
||||||
```
|
```
|
||||||
|
|
||||||
The important claim is not "the model answered once." The claim is:
|
关键主张不是“模型回答了一次”,而是:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
The system can show what evidence was used, how the answer was checked, and how to replay the session.
|
系统能展示用了什么证据、答案如何被检查、如何用 sessionId 回放整次诊断。
|
||||||
```
|
```
|
||||||
|
|
||||||
## Demo Flow
|
## 2. Demo 流程
|
||||||
|
|
||||||
1. Start the service with the `mvp-demo` profile.
|
1. 用 `mvp-demo` profile 启动服务。
|
||||||
2. Run the fixed payment-timeout request.
|
2. 运行固定的支付超时请求。
|
||||||
3. Open `mvp/demo/output/chat-response.json`.
|
3. 打开 `mvp/demo/output/chat-response.json`。
|
||||||
4. Open `mvp/demo/output/trace-response.json`.
|
4. 打开 `mvp/demo/output/trace-response.json`。
|
||||||
5. Point to evidence tools and verifier evaluation.
|
5. 指出证据工具和 verifier evaluation。
|
||||||
6. Submit feedback and show it is attached to the same session.
|
6. 提交 feedback,并展示它挂在同一个 session 上。
|
||||||
|
|
||||||
## Commands
|
## 3. 命令
|
||||||
|
|
||||||
Start service:
|
启动服务:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||||
```
|
```
|
||||||
|
|
||||||
Run the demo from another terminal:
|
另开终端运行 Demo:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||||
```
|
```
|
||||||
|
|
||||||
Optional custom session:
|
可选自定义 session:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
||||||
```
|
```
|
||||||
|
|
||||||
## What To Show
|
## 4. 展示什么
|
||||||
|
|
||||||
### 1. User-Facing Answer
|
### 4.1 用户侧答案
|
||||||
|
|
||||||
File:
|
文件:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
mvp/demo/output/chat-response.json
|
mvp/demo/output/chat-response.json
|
||||||
```
|
```
|
||||||
|
|
||||||
Say:
|
话术:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
|
这是用户看到的答案。这里的 sessionId 是稳定的,所以我后面可以追踪这一次回答是怎么来的。
|
||||||
```
|
```
|
||||||
|
|
||||||
### 2. Evidence Trace
|
### 4.2 证据 Trace
|
||||||
|
|
||||||
File:
|
文件:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
mvp/demo/output/trace-response.json
|
mvp/demo/output/trace-response.json
|
||||||
```
|
```
|
||||||
|
|
||||||
Say:
|
话术:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
This is the important Agent engineering part.
|
这才是 Agent 工程最重要的部分。
|
||||||
I can inspect which tools were called, what inputs they received,
|
我可以检查 Agent 调用了哪些工具、每个工具拿到什么入参、是否成功、返回了什么证据预览。
|
||||||
whether they succeeded, and what evidence preview was persisted.
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Point to:
|
重点字段:
|
||||||
|
|
||||||
- `data.toolInvocations[*].toolName`
|
- `data.toolInvocations[*].toolName`
|
||||||
- `data.toolInvocations[*].inputParams`
|
- `data.toolInvocations[*].inputParams`
|
||||||
- `data.toolInvocations[*].outputPreview`
|
- `data.toolInvocations[*].outputPreview`
|
||||||
- `data.toolInvocations[*].success`
|
- `data.toolInvocations[*].success`
|
||||||
|
|
||||||
### 3. Verifier / Self-Evaluation
|
### 4.3 Verifier / 自评估
|
||||||
|
|
||||||
Point to:
|
重点字段:
|
||||||
|
|
||||||
- `data.session.selfEvaluation`
|
- `data.session.selfEvaluation`
|
||||||
- `data.summary.hasVerifierEvaluation`
|
- `data.summary.hasVerifierEvaluation`
|
||||||
|
|
||||||
Say:
|
话术:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
The final answer is not just raw Executor output.
|
最终答案不是 Executor 原始输出直接返回。
|
||||||
It is checked by a verifier or self-evaluation layer using the persisted trace.
|
系统会基于持久化的工具 trace 做 Verifier 或规则自评估。
|
||||||
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
|
这样系统可以区分 PASS、LOW_CONFID、REJECT,而不是假装每个答案都同样可信。
|
||||||
```
|
```
|
||||||
|
|
||||||
### 4. Feedback Loop
|
### 4.4 反馈闭环
|
||||||
|
|
||||||
File:
|
文件:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
mvp/demo/output/feedback-response.json
|
mvp/demo/output/feedback-response.json
|
||||||
```
|
```
|
||||||
|
|
||||||
Then re-query trace if needed.
|
必要时重新查询 Trace。
|
||||||
|
|
||||||
Say:
|
话术:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Feedback is attached to the same diagnosis session.
|
feedback 会挂在同一个 diagnosis session 上。
|
||||||
That makes it possible to mine useful / not useful cases later.
|
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
|
||||||
```
|
```
|
||||||
|
|
||||||
### 5. Regression Story
|
### 4.5 回归故事
|
||||||
|
|
||||||
Mention, do not deep dive unless asked:
|
如果被问到稳定性,可以补充:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
For repeatability, I also built an offline eval baseline.
|
我把运行时 Demo 和离线 eval 分开。
|
||||||
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
|
Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可以回归。
|
||||||
The two are separate on purpose: demo for human review, eval for automated signal.
|
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
|
||||||
```
|
```
|
||||||
|
|
||||||
## Strong Interview Framing
|
## 5. 强面试表达
|
||||||
|
|
||||||
Use this phrasing:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
I focused on the Agent engineering surface:
|
我关注的是 Agent 工程表面:
|
||||||
traceability, evidence persistence, verifier gating, feedback, and regression checks.
|
traceability、evidence persistence、verifier gating、feedback 和 regression checks。
|
||||||
The model answer is only one part of the system.
|
模型答案只是系统的一部分。
|
||||||
The more important part is whether we can audit and improve the answer after it is produced.
|
更重要的是答案产出后,能否被审计、验证和持续改进。
|
||||||
```
|
```
|
||||||
|
|
||||||
## Known Limits To Say Proactively
|
## 6. 主动说明限制
|
||||||
|
|
||||||
```text
|
```text
|
||||||
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
|
这个 MVP 仍依赖 MySQL、Redis、Milvus 和模型凭证。
|
||||||
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
|
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
|
||||||
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
|
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -1,11 +1,12 @@
|
|||||||
# Demo Output
|
# Demo 输出目录
|
||||||
|
|
||||||
This directory is the default output location for local demo responses.
|
本目录是本地 Demo 响应的默认输出位置。
|
||||||
|
|
||||||
Generated files are intentionally ignored by Git:
|
生成文件会被 Git 忽略:
|
||||||
|
|
||||||
- `chat-response.json`
|
- `chat-response.json`
|
||||||
- `trace-response.json`
|
- `trace-response.json`
|
||||||
- `feedback-response.json`
|
- `feedback-response.json`
|
||||||
|
|
||||||
Keep this README so the directory exists in the repository.
|
保留此 README 是为了让目录存在于仓库中。
|
||||||
|
|
||||||
|
|||||||
@@ -1,24 +1,24 @@
|
|||||||
# Payment Timeout Acceptance Case
|
# 支付超时诊断验收用例
|
||||||
|
|
||||||
## Goal
|
## 1. 目标
|
||||||
|
|
||||||
Validate that the MVP can diagnose a payment timeout incident and expose the complete trace for replay.
|
验证 MVP 能诊断支付超时问题,并暴露完整 Trace 供回放。
|
||||||
|
|
||||||
## Input
|
## 2. 输入
|
||||||
|
|
||||||
- Session id: `mvp-demo-payment-timeout-001`
|
- Session id:`mvp-demo-payment-timeout-001`
|
||||||
- Question: `支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
|
- 问题:`支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
|
||||||
- Profile: `mvp-demo`
|
- Profile:`mvp-demo`
|
||||||
|
|
||||||
## Acceptance Criteria
|
## 3. 验收标准
|
||||||
|
|
||||||
1. Chat returns a successful answer with the same session id.
|
1. Chat 返回成功答复,且 session id 与请求一致。
|
||||||
2. Trace API returns session metadata, final answer, ordered agent steps, and ordered tool invocations.
|
2. Trace API 返回 session 元数据、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
||||||
3. Trace contains enough evidence to explain which tools were used and whether verifier/self-evaluation was persisted.
|
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
|
||||||
4. Feedback can be submitted for the same session id.
|
4. 可以使用同一个 session id 提交反馈。
|
||||||
5. A follow-up trace query shows the persisted feedback value.
|
5. 后续 Trace 查询能看到已持久化的 feedback 值。
|
||||||
|
|
||||||
## Trace Fields To Inspect
|
## 4. 需要检查的 Trace 字段
|
||||||
|
|
||||||
- `data.session.query`
|
- `data.session.query`
|
||||||
- `data.session.answer`
|
- `data.session.answer`
|
||||||
@@ -32,8 +32,9 @@ Validate that the MVP can diagnose a payment timeout incident and expose the com
|
|||||||
- `data.toolInvocations[*].retrievalDetails`
|
- `data.toolInvocations[*].retrievalDetails`
|
||||||
- `data.summary`
|
- `data.summary`
|
||||||
|
|
||||||
## Known Limits
|
## 5. 已知边界
|
||||||
|
|
||||||
|
- 这不是完整离线测试,仍需要有效的 chat、持久化、向量检索和模型调用环境。
|
||||||
|
- `mvp-demo` profile 启用 mock 日志和指标,让证据工具返回更稳定。
|
||||||
|
- 敏感配置清理不属于当前 MVP 优先级。
|
||||||
|
|
||||||
- This case is not a full offline test. It still requires valid infrastructure for chat, persistence, vector search, and model calls.
|
|
||||||
- Mock logs and metrics are enabled by the `mvp-demo` profile to make those evidence tools repeatable.
|
|
||||||
- Sensitive configuration cleanup is deferred by current MVP priority.
|
|
||||||
|
|||||||
@@ -13,7 +13,7 @@ $request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
|||||||
$request.Id = $SessionId
|
$request.Id = $SessionId
|
||||||
$body = $request | ConvertTo-Json -Depth 8
|
$body = $request | ConvertTo-Json -Depth 8
|
||||||
|
|
||||||
Write-Host "Running payment-timeout chat demo..."
|
Write-Host "正在运行支付超时 Chat 诊断 Demo..."
|
||||||
Write-Host "BaseUrl: $BaseUrl"
|
Write-Host "BaseUrl: $BaseUrl"
|
||||||
Write-Host "SessionId: $SessionId"
|
Write-Host "SessionId: $SessionId"
|
||||||
|
|
||||||
@@ -24,14 +24,14 @@ $chat = Invoke-RestMethod `
|
|||||||
-Body $body
|
-Body $body
|
||||||
|
|
||||||
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||||
Write-Host "Saved chat response: $OutputDir/chat-response.json"
|
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
|
||||||
|
|
||||||
$trace = Invoke-RestMethod `
|
$trace = Invoke-RestMethod `
|
||||||
-Method Get `
|
-Method Get `
|
||||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||||
|
|
||||||
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||||
Write-Host "Saved trace response: $OutputDir/trace-response.json"
|
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
|
||||||
|
|
||||||
$feedbackBody = @{
|
$feedbackBody = @{
|
||||||
sessionId = $SessionId
|
sessionId = $SessionId
|
||||||
@@ -45,10 +45,10 @@ $feedback = Invoke-RestMethod `
|
|||||||
-Body $feedbackBody
|
-Body $feedbackBody
|
||||||
|
|
||||||
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
|
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
|
||||||
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
|
Write-Host "已保存反馈响应: $OutputDir/feedback-response.json"
|
||||||
|
|
||||||
Write-Host ""
|
Write-Host ""
|
||||||
Write-Host "Demo completed. Review:"
|
Write-Host "Demo 已完成,请检查:"
|
||||||
Write-Host "- mvp/demo/output/chat-response.json"
|
Write-Host "- mvp/demo/output/chat-response.json"
|
||||||
Write-Host "- mvp/demo/output/trace-response.json"
|
Write-Host "- mvp/demo/output/trace-response.json"
|
||||||
Write-Host "- mvp/demo/output/feedback-response.json"
|
Write-Host "- mvp/demo/output/feedback-response.json"
|
||||||
|
|||||||
@@ -0,0 +1,237 @@
|
|||||||
|
# 10 分钟面试演示脚本
|
||||||
|
|
||||||
|
**用途**:面试现场按步骤演示
|
||||||
|
**目标**:展示从问题到证据、验证、Trace、反馈的闭环
|
||||||
|
**前置条件**:服务以 `mvp-demo` profile 启动
|
||||||
|
|
||||||
|
更完整的 runbook 见 [README.md](README.md),字段检查见 [trace-inspection-checklist.md](trace-inspection-checklist.md)。
|
||||||
|
|
||||||
|
## 0. 开场话术
|
||||||
|
|
||||||
|
```text
|
||||||
|
我会演示一个支付超时诊断。
|
||||||
|
重点不是看模型给出一段答案,而是看这个答案背后的 Agent 执行链路:
|
||||||
|
Planner 怎么拆解,Executor 调了哪些工具,Verifier 如何判断证据是否支撑答案,以及最终如何通过 sessionId 回放。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 1. 启动服务
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||||
|
```
|
||||||
|
|
||||||
|
服务地址:
|
||||||
|
|
||||||
|
```text
|
||||||
|
http://localhost:9900
|
||||||
|
```
|
||||||
|
|
||||||
|
说明:
|
||||||
|
|
||||||
|
- `mvp-demo` profile 使用 mock Prometheus 和 mock CLS。
|
||||||
|
- 演示不依赖真实线上故障。
|
||||||
|
- MySQL、Redis、Milvus/Zilliz 和模型配置仍需要可用。
|
||||||
|
|
||||||
|
## 2. 演示 Chat 诊断
|
||||||
|
|
||||||
|
推荐使用固定脚本:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||||
|
```
|
||||||
|
|
||||||
|
脚本会写出:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/output/chat-response.json
|
||||||
|
mvp/demo/output/trace-response.json
|
||||||
|
mvp/demo/output/feedback-response.json
|
||||||
|
```
|
||||||
|
|
||||||
|
现场话术:
|
||||||
|
|
||||||
|
```text
|
||||||
|
这里我用固定 sessionId 跑一个支付接口超时问题。
|
||||||
|
固定 sessionId 的好处是,后面 trace 和 feedback 都能关联到同一次诊断。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. 展示用户答案
|
||||||
|
|
||||||
|
打开:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/output/chat-response.json
|
||||||
|
```
|
||||||
|
|
||||||
|
重点看:
|
||||||
|
|
||||||
|
```text
|
||||||
|
data.sessionId
|
||||||
|
data.answer
|
||||||
|
```
|
||||||
|
|
||||||
|
现场话术:
|
||||||
|
|
||||||
|
```text
|
||||||
|
这是用户看到的答案。
|
||||||
|
但这个项目的重点不是这段文字,而是这段文字是否有证据链。
|
||||||
|
接下来我用同一个 sessionId 查 trace。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 4. 展示 Trace
|
||||||
|
|
||||||
|
打开:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/output/trace-response.json
|
||||||
|
```
|
||||||
|
|
||||||
|
重点看:
|
||||||
|
|
||||||
|
```text
|
||||||
|
data.session.sessionId
|
||||||
|
data.session.agentFlow
|
||||||
|
data.steps[*].agentName
|
||||||
|
data.toolInvocations[*].toolName
|
||||||
|
data.toolInvocations[*].inputParams
|
||||||
|
data.toolInvocations[*].outputPreview
|
||||||
|
data.toolInvocations[*].retrievalLayer
|
||||||
|
data.toolInvocations[*].relevanceLevel
|
||||||
|
data.summary.hasVerifierEvaluation
|
||||||
|
```
|
||||||
|
|
||||||
|
现场话术:
|
||||||
|
|
||||||
|
```text
|
||||||
|
这里能看到三个层次:
|
||||||
|
第一,session 记录了这次诊断的问题、答案、耗时和自评估。
|
||||||
|
第二,agent_step 记录 Planner、Executor、Verifier 的模型步骤。
|
||||||
|
第三,tool_invocation 记录真实工具调用,包括 lookup_knowledge、日志和指标。
|
||||||
|
|
||||||
|
所以这不是一个黑盒 Chatbot,而是一条可以回放的诊断链路。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 5. 展示知识库检索
|
||||||
|
|
||||||
|
在 trace 中找到 `lookup_knowledge`。
|
||||||
|
|
||||||
|
重点看:
|
||||||
|
|
||||||
|
```text
|
||||||
|
toolName = lookup_knowledge
|
||||||
|
inputParams.query
|
||||||
|
retrievalLayer
|
||||||
|
l0MatchCount
|
||||||
|
l1MatchCount
|
||||||
|
relevanceLevel
|
||||||
|
retrievalDetails
|
||||||
|
outputPreview
|
||||||
|
```
|
||||||
|
|
||||||
|
现场话术:
|
||||||
|
|
||||||
|
```text
|
||||||
|
知识库检索保留为显式工具,而不是藏在 Advisor 里。
|
||||||
|
这样面试官或线上排查人员能看到:Agent 查了什么 query,命中了哪个知识域,检索层是 L0/L1 还是混合,相关性等级是什么。
|
||||||
|
|
||||||
|
底层检索现在走 VectorSearchService,优先 Spring AI VectorStore,失败时 fallback 到 Milvus SDK。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 6. 展示 Verifier
|
||||||
|
|
||||||
|
在 trace 中查看:
|
||||||
|
|
||||||
|
```text
|
||||||
|
data.session.selfEvaluation
|
||||||
|
data.summary.hasVerifierEvaluation
|
||||||
|
```
|
||||||
|
|
||||||
|
现场话术:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Verifier 不做新检索,只看工具 trace 汇总。
|
||||||
|
它会把 Executor 答案里的关键事实拆出来,判断每条事实是 direct_evidence、indirect_support、no_evidence 还是 contradicted。
|
||||||
|
|
||||||
|
如果 PASS,就输出原答案。
|
||||||
|
如果 LOW_CONFID,可以补证据或加低置信提示。
|
||||||
|
如果 REJECT,就降级输出,只保留已确认信息。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 7. 展示反馈闭环
|
||||||
|
|
||||||
|
打开:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/output/feedback-response.json
|
||||||
|
```
|
||||||
|
|
||||||
|
重点看:
|
||||||
|
|
||||||
|
```text
|
||||||
|
success
|
||||||
|
caseId
|
||||||
|
```
|
||||||
|
|
||||||
|
现场话术:
|
||||||
|
|
||||||
|
```text
|
||||||
|
用户反馈 useful 会写回同一个 diagnosis_session。
|
||||||
|
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
|
||||||
|
|
||||||
|
这里 status 和 feedback 是分开的:
|
||||||
|
status 表示执行是否成功,feedback 表示用户是否认可。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 8. 可选演示 AIOps
|
||||||
|
|
||||||
|
如果时间允许,再演示 AIOps payload。
|
||||||
|
|
||||||
|
请求示例见:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/README.md
|
||||||
|
```
|
||||||
|
|
||||||
|
现场话术:
|
||||||
|
|
||||||
|
```text
|
||||||
|
AIOps 有两个模式。
|
||||||
|
有 payload 时进入 PAYLOAD_TARGETED,报告必须聚焦这个告警。
|
||||||
|
没有 payload 时进入 AUTO_DISCOVERY,先发现活跃告警再排查。
|
||||||
|
|
||||||
|
我专门加了 recommended lookup_knowledge query,把 alertName、service、severity、description 等字段稳定送入知识库检索,避免 Agent 随意扩展问题范围。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 9. 结束总结
|
||||||
|
|
||||||
|
```text
|
||||||
|
这个 Demo 展示的是一个完整闭环:
|
||||||
|
|
||||||
|
用户问题
|
||||||
|
-> Agent 规划和执行
|
||||||
|
-> 显式工具证据
|
||||||
|
-> Verifier / self_evaluation
|
||||||
|
-> Trace 回放
|
||||||
|
-> 用户反馈
|
||||||
|
-> 案例沉淀
|
||||||
|
|
||||||
|
我把重点放在 Agent 工程能力:可追踪、可验证、可回归、可演进。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 10. 如果现场失败
|
||||||
|
|
||||||
|
如果模型或外部组件不可用,不要硬跑。可以直接打开上一次输出:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/output/chat-response.json
|
||||||
|
mvp/demo/output/trace-response.json
|
||||||
|
mvp/demo/output/feedback-response.json
|
||||||
|
```
|
||||||
|
|
||||||
|
降级话术:
|
||||||
|
|
||||||
|
```text
|
||||||
|
现场环境依赖 MySQL、Redis、Milvus 和模型服务。
|
||||||
|
如果外部服务不可用,我会用固定输出讲 trace 结构。
|
||||||
|
因为这个项目的核心不是一次在线请求,而是诊断链路如何被记录、检查和回放。
|
||||||
|
```
|
||||||
@@ -1,52 +1,53 @@
|
|||||||
# Trace Inspection Checklist
|
# Trace 检查清单
|
||||||
|
|
||||||
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
|
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。
|
||||||
|
|
||||||
## Session
|
## 1. Session
|
||||||
|
|
||||||
| JSON path | What to check | Interview point |
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
| --- | --- | --- |
|
|---|---|---|
|
||||||
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
|
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | 一个 session id 串起 chat、工具、verifier、feedback 和 trace |
|
||||||
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
|
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||||
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
|
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||||
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
|
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||||
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
|
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
|
||||||
|
|
||||||
## Agent Steps
|
## 2. Agent 步骤
|
||||||
|
|
||||||
| JSON path | What to check | Interview point |
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
| --- | --- | --- |
|
|---|---|---|
|
||||||
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
|
| `data.steps[*].agentName` | 是否有 Planner / Executor / Verifier 或等价步骤 | 流程被拆成可检查的 Agent 步骤 |
|
||||||
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
|
| `data.steps[*].thought` | 是否有高层步骤摘要 | 内部过程可审计,不只看最终文本 |
|
||||||
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
|
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
||||||
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
|
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
||||||
|
|
||||||
## Tool Evidence
|
## 3. 工具证据
|
||||||
|
|
||||||
| JSON path | What to check | Interview point |
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
| --- | --- | --- |
|
|---|---|---|
|
||||||
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
|
| `data.toolInvocations[*].toolName` | 是否包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 | Agent 通过工具收集证据,而不是无依据猜测 |
|
||||||
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
|
| `data.toolInvocations[*].inputParams` | 是否能看到每个工具的入参 | 工具输入可审计、可调试 |
|
||||||
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
|
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
|
||||||
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
|
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
|
||||||
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
|
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
|
||||||
|
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||||
|
|
||||||
## Summary
|
## 4. Summary
|
||||||
|
|
||||||
| JSON path | What to check | Interview point |
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
| --- | --- | --- |
|
|---|---|---|
|
||||||
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
|
| `data.summary.persistedStepCount` | step 行是否持久化 | Trace 来自存储,不是响应内存 |
|
||||||
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
|
| `data.summary.persistedToolCallCount` | tool 行是否持久化 | 工具证据在请求结束后仍可回放 |
|
||||||
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
|
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
||||||
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
|
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
||||||
|
|
||||||
## What Good Looks Like
|
## 5. 好的结果长什么样
|
||||||
|
|
||||||
```text
|
```text
|
||||||
same session id
|
同一个 session id
|
||||||
-> final answer
|
-> 最终答案
|
||||||
-> persisted agent steps
|
-> 持久化 agent steps
|
||||||
-> persisted evidence tool calls
|
-> 持久化 evidence tool calls
|
||||||
-> verifier/self-evaluation
|
-> verifier / self-evaluation
|
||||||
-> feedback attached to the same session
|
-> feedback attached to the same session
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -4,7 +4,7 @@
|
|||||||
**严重程度**:中(影响 token 消耗和上下文质量,不影响功能正确性)
|
**严重程度**:中(影响 token 消耗和上下文质量,不影响功能正确性)
|
||||||
**发现时间**:2026-06-30
|
**发现时间**:2026-06-30
|
||||||
**修复版本**:session-dedup-knowledge-map
|
**修复版本**:session-dedup-knowledge-map
|
||||||
**架构文档**:[会话级去重与知识域地图](../architecture/session-dedup-knowledge-map.md)
|
**历史架构文档**:[会话级去重与知识域地图](../architecture/archive/2026-07-05-legacy/session-dedup-knowledge-map.md)
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -58,7 +58,7 @@
|
|||||||
|
|
||||||
### P1:会话管理设计与实现不一致
|
### P1:会话管理设计与实现不一致
|
||||||
|
|
||||||
`mvp/architecture/session-management.md` 设计 Redis 作为主会话存储,带 `session:{session_id}` 和 TTL。
|
`mvp/architecture/archive/2026-07-05-legacy/session-management.md` 设计 Redis 作为主会话存储,带 `session:{session_id}` 和 TTL。
|
||||||
|
|
||||||
实际 `/api/chat` 在 `ChatController` 中使用 JVM 内存 `ConcurrentHashMap` 管理历史消息,`RedisSessionManager` 虽然存在但没有接入 controller。
|
实际 `/api/chat` 在 `ChatController` 中使用 JVM 内存 `ConcurrentHashMap` 管理历史消息,`RedisSessionManager` 虽然存在但没有接入 controller。
|
||||||
|
|
||||||
|
|||||||
@@ -56,4 +56,4 @@ Prompt 软约束依赖 LLM 自觉遵守。在 ReactAgent 自主决策模式下
|
|||||||
- `src/main/java/com/superbiz/agent/tool/RetrievedDocTracker.java`
|
- `src/main/java/com/superbiz/agent/tool/RetrievedDocTracker.java`
|
||||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||||
- `src/main/resources/prompts/chat-executor-prompt.md`
|
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||||
- `mvp/architecture/action-memory-relevance.md`
|
- `mvp/architecture/archive/2026-07-05-legacy/action-memory-relevance.md`
|
||||||
|
|||||||
Reference in New Issue
Block a user