docs: reorganize MVP interview documentation
This commit is contained in:
+32
-25
@@ -1,48 +1,54 @@
|
||||
# SuperBizAgent Interview Guide
|
||||
# SuperBizAgent 面试资料包
|
||||
|
||||
## 一句话定位
|
||||
|
||||
SuperBizAgent 是一个面向企业故障诊断场景的 Agent Engineering 项目:它把用户问题或告警事件转成可追踪的多 Agent 执行链路,并把工具证据、模型步骤、最终答案和反馈统一落到诊断 trace 中。
|
||||
SuperBizAgent 是一个面向企业故障诊断场景的 Agent 工程项目。它把用户问题或 AIOps 告警转换成可追踪的 Agent 执行链路,并把工具证据、模型步骤、最终答案、自评估和用户反馈统一沉淀到诊断 Trace 中。
|
||||
|
||||
## 面试重点
|
||||
|
||||
- **多 Agent 编排**:普通 Chat 的复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 Supervisor 调度 Planner/Executor。
|
||||
- **工具证据链**:知识库、日志、指标和 Prometheus 告警都通过工具调用进入链路,并记录到 `tool_invocation`。
|
||||
- **可追踪诊断**:一次会话对应一个 `sessionId`,最终可以通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
|
||||
- **质量门**:Chat 链路包含 Verifier,把 groundedness、facts checked 和 evidence refs 写回 `diagnosis_session.self_evaluation`。
|
||||
- **AIOps 产品边界**:有告警 payload 时聚焦该告警;没有 payload 时先自动发现 active alerts。
|
||||
- **可复现 Demo**:`mvp-demo` profile 使用 mock Prometheus 和 mock CLS,让面试演示不依赖真实线上故障。
|
||||
- **Agent 编排**:Chat 复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 `Supervisor -> Planner / Executor`。
|
||||
- **工具证据链**:知识库、日志、指标、Prometheus 告警都通过显式工具调用进入链路,并记录到 `tool_invocation`。
|
||||
- **可追踪诊断**:一次诊断对应一个 `sessionId`,可通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
|
||||
- **质量门禁**:Chat Verifier 校验 groundedness;AIOps 规则评估检查报告完整性、payload 聚焦和证据工具覆盖。
|
||||
- **RAG 工程化**:`lookup_knowledge` 是显式 Agent Tool,底层通过 Spring AI VectorStore 主路径 + Milvus SDK fallback。
|
||||
- **反馈闭环**:用户反馈 `useful` 会沉淀 `case_library`,`not_useful` 保留 bad case 信号。
|
||||
|
||||
## 推荐阅读顺序
|
||||
|
||||
1. `interview/demo-script.md`:面试现场怎么讲、怎么演示。
|
||||
2. `interview/architecture.md`:系统架构和两条主链路。
|
||||
3. `interview/design-tradeoffs.md`:关键设计取舍和可被追问的问题。
|
||||
4. `interview/acceptance-checklist.md`:面试前验证清单。
|
||||
5. `mvp/demo/README.md`:更细的 MVP 可执行 runbook。
|
||||
1. `mvp/architecture/interview-one-pager.md`:一页式架构图和 2-5 分钟讲解。
|
||||
2. `mvp/demo/ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||
3. `interview/story-cases.md`:可复用的面试故事案例。
|
||||
4. `interview/architecture.md`:面试版系统架构。
|
||||
5. `interview/design-tradeoffs.md`:关键设计取舍。
|
||||
6. `interview/demo-script.md`:更细的命令式演示脚本。
|
||||
7. `interview/acceptance-checklist.md`:面试前验收清单。
|
||||
8. RAG 专题文档:`rag-refactor-story.md`、`rag-vectorstore-interview-notes.md`、`rag-retrieval-quality-report.md`。
|
||||
|
||||
## 核心 Demo
|
||||
## 核心演示链路
|
||||
|
||||
### Chat Diagnosis
|
||||
### Chat 诊断
|
||||
|
||||
```text
|
||||
POST /api/chat
|
||||
-> ChatService.executeChatWithStrategy(...)
|
||||
-> simple ReactAgent or Planner -> Executor -> Verifier
|
||||
-> ChatService
|
||||
-> Planner -> Executor -> Verifier
|
||||
-> lookup_knowledge / query_logs / query_metrics
|
||||
-> diagnosis_session + agent_step + tool_invocation
|
||||
-> GET /api/diagnosis/{sessionId}/trace
|
||||
-> POST /api/feedback
|
||||
```
|
||||
|
||||
### AIOps Alert Diagnosis
|
||||
### AIOps 告警诊断
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> AiOpsService.executeAiOpsAnalysis(...)
|
||||
-> AiOpsService
|
||||
-> PAYLOAD_TARGETED / AUTO_DISCOVERY
|
||||
-> ai_ops_supervisor
|
||||
-> planner_agent / executor_agent
|
||||
-> queryPrometheusAlerts + logs + knowledge
|
||||
-> scoped alert report
|
||||
-> queryPrometheusAlerts + logs + metrics + lookup_knowledge
|
||||
-> alert report
|
||||
-> aiops_rule_evaluation
|
||||
-> GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
@@ -50,10 +56,11 @@ POST /api/ai_ops
|
||||
|
||||
- Chat 诊断链路:可运行、可追踪、有 Verifier。
|
||||
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
|
||||
- RAG 检索链路:Spring AI VectorStore 主路径、Milvus SDK fallback、L0 hint、检索评测 baseline。
|
||||
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
|
||||
- Demo 文档:`mvp/demo/README.md` 和 `mvp/demo/aiops-alert-acceptance.md`。
|
||||
- Devflow 沉淀:`devflow/index.md` 记录了 MVP、Verifier、AIOps trace 和 AIOps scope-control 的演进。
|
||||
- Demo 材料:`mvp/demo/README.md`、`mvp/demo/ten-minute-interview-demo.md`。
|
||||
|
||||
## 面试时的主叙事
|
||||
## 主叙事
|
||||
|
||||
这个项目不是简单调用大模型,而是在做一个可审计、可验证、可回归的 Agent 诊断系统。模型可以规划和推理,但每一步工具证据、最终结论、Verifier 结果和用户反馈都能被 Trace API 回放。面试时重点展示“从问题到证据到答案到验证再到反馈”的闭环。
|
||||
|
||||
这个项目不是简单调用大模型,而是在做一个可审计的 Agent 诊断系统。核心价值是:模型可以规划和推理,但每一步工具证据、最终结论和质量评估都能被 trace API 回放。面试时重点展示“从问题到证据到答案到验证”的完整闭环。
|
||||
|
||||
@@ -1,15 +1,15 @@
|
||||
# Acceptance Checklist
|
||||
# 面试前验收清单
|
||||
|
||||
## 面试前环境检查
|
||||
## 1. 环境检查
|
||||
|
||||
- 当前分支包含最新 AIOps trace/scope 变更。
|
||||
- 当前分支包含最新架构文档和面试材料。
|
||||
- MySQL 可连接。
|
||||
- Redis 可连接。
|
||||
- Milvus/Zilliz 可连接。
|
||||
- 模型 API key 可用。
|
||||
- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。
|
||||
|
||||
启动:
|
||||
启动服务:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
@@ -24,10 +24,10 @@ mvn -q -DskipTests compile
|
||||
目标测试:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test
|
||||
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest,VectorSearchServiceTest,LookupKnowledgeToolTest" test
|
||||
```
|
||||
|
||||
## Chat Demo 验收
|
||||
## 2. Chat Demo 验收
|
||||
|
||||
请求:
|
||||
|
||||
@@ -50,7 +50,7 @@ Invoke-RestMethod `
|
||||
- 返回 `data.success = true`。
|
||||
- 返回 `data.sessionId = interview-chat-payment-timeout-001`。
|
||||
- `diagnosis_session.agent_flow = CHAT`。
|
||||
- trace API 返回 session、steps、toolInvocations。
|
||||
- Trace API 返回 session、steps、toolInvocations。
|
||||
- 复杂问题下 trace 中能看到 verifier 相关数据。
|
||||
|
||||
SQL:
|
||||
@@ -59,7 +59,7 @@ SQL:
|
||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'"
|
||||
```
|
||||
|
||||
## AIOps Demo 验收
|
||||
## 3. AIOps Demo 验收
|
||||
|
||||
请求:
|
||||
|
||||
@@ -89,11 +89,10 @@ Invoke-WebRequest `
|
||||
- `diagnosis_session.agent_flow = AI_OPS`。
|
||||
- `diagnosis_session.status = SUCCESS`。
|
||||
- `diagnosis_session.answer` 有最终报告。
|
||||
- trace API 返回 AIOps steps 和 tool invocations。
|
||||
- Trace API 返回 AIOps steps 和 tool invocations。
|
||||
- 报告主章节聚焦 `HighCPUUsage/payment-service`。
|
||||
- 无 `告警根因分析 - HighMemoryUsage` 独立章节。
|
||||
- 无 `告警根因分析 - SlowResponse` 独立章节。
|
||||
- 有“相关风险告警”或类似上下文说明。
|
||||
- 其他 active alerts 不应展开成独立主根因章节。
|
||||
- `self_evaluation.aiops_rule_evaluation` 存在。
|
||||
|
||||
SQL:
|
||||
|
||||
@@ -108,19 +107,10 @@ python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invoc
|
||||
Scope 检查:
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT (answer LIKE '%告警根因分析 - HighCPUUsage%') AS has_main_root_cause, (answer LIKE '%告警根因分析 - HighMemoryUsage%') AS has_memory_root_cause, (answer LIKE '%告警根因分析 - SlowResponse%') AS has_slow_root_cause, (answer LIKE '%相关风险告警%') AS has_related_risk FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
||||
python scripts/query_mysql.py "SELECT (answer LIKE '%HighCPUUsage%') AS has_main_alert, (answer LIKE '%payment-service%') AS has_service FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
||||
```
|
||||
|
||||
期望:
|
||||
|
||||
```text
|
||||
has_main_root_cause = 1
|
||||
has_memory_root_cause = 0
|
||||
has_slow_root_cause = 0
|
||||
has_related_risk = 1
|
||||
```
|
||||
|
||||
## Trace API 验收
|
||||
## 4. Trace API 验收
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
@@ -128,13 +118,28 @@ Invoke-RestMethod `
|
||||
-Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||
```
|
||||
|
||||
若 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
|
||||
如果 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
|
||||
|
||||
```powershell
|
||||
curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||
```
|
||||
|
||||
## 常见问题
|
||||
## 5. RAG 验收
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Uri "http://127.0.0.1:9900/api/search/similar?query=ERR_TIMEOUT&topK=3" `
|
||||
-Method Get
|
||||
```
|
||||
|
||||
验收:
|
||||
|
||||
- 返回 `code = 200`。
|
||||
- top candidates 中包含 `ERR_TIMEOUT` 相关文档。
|
||||
- `scoreLabel` 能体现当前检索路径语义。
|
||||
- 如果走 VectorStore,日志应出现 Spring AI VectorStore search。
|
||||
|
||||
## 6. 常见问题
|
||||
|
||||
### MySQL stale connection
|
||||
|
||||
@@ -145,26 +150,17 @@ HikariPool - Connection is not available
|
||||
No operations allowed after connection closed
|
||||
```
|
||||
|
||||
当前已在 `application.yml` 配置:
|
||||
|
||||
- `maximum-pool-size: 5`
|
||||
- `minimum-idle: 1`
|
||||
- `connection-timeout: 10000`
|
||||
- `validation-timeout: 5000`
|
||||
- `idle-timeout: 60000`
|
||||
- `max-lifetime: 120000`
|
||||
- `keepalive-time: 30000`
|
||||
|
||||
处理:
|
||||
|
||||
- 重新编译或重启服务。
|
||||
- 确认日志中新的 HikariPool 启动成功。
|
||||
- 重启服务。
|
||||
- 确认 HikariPool 使用当前配置启动成功。
|
||||
- 再跑 trace 或 AIOps 请求。
|
||||
|
||||
### SSE 客户端显示异常
|
||||
|
||||
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 trace API 验证结果。
|
||||
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 Trace API 验证结果。
|
||||
|
||||
### OpenSpec 全量校验失败
|
||||
|
||||
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试材料主要依赖已归档的 AIOps spec 和 MVP trace spec,可以单独验证相关 spec。
|
||||
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试演示主要依赖已归档 spec、MVP trace 和 RAG 验收材料,可以单独验证相关 spec。
|
||||
|
||||
|
||||
@@ -1,38 +1,37 @@
|
||||
# AIOps Lightweight Verifier
|
||||
# AIOps 轻量规则验证器
|
||||
|
||||
## What Changed
|
||||
## 1. 改动是什么
|
||||
|
||||
AIOps now has a deterministic post-run quality gate.
|
||||
AIOps 现在有一个确定性的后置质量门禁。最终告警报告持久化后,`AiOpsRuleEvaluationService` 会检查:
|
||||
|
||||
After the final AIOps report is persisted, the service evaluates:
|
||||
- 最终报告是否存在,且不是明显过短。
|
||||
- payload 模式下,报告是否提到输入的告警和服务。
|
||||
- 是否有证据工具调用,例如 `lookup_knowledge`、`query_metrics`、`query_logs`。
|
||||
|
||||
- whether the final report exists and is not trivially short
|
||||
- whether a payload-targeted report mentions the supplied alert and service
|
||||
- whether evidence tools such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted
|
||||
|
||||
The result is stored under:
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
The trace API returns this payload through the existing session self-evaluation field.
|
||||
Trace API 会通过 session self-evaluation 展示这个结果。
|
||||
|
||||
## Why Rule-Based First
|
||||
## 2. 为什么先做规则型
|
||||
|
||||
This is not a full LLM verifier yet.
|
||||
这还不是完整 LLM Verifier。
|
||||
|
||||
The first AIOps quality risks are concrete and easy to check with rules:
|
||||
AIOps 第一阶段质量风险比较具体,适合先用规则:
|
||||
|
||||
- Did the report stay focused on the payload?
|
||||
- Did the run use evidence tools?
|
||||
- Did the system produce a usable final report?
|
||||
- 报告有没有生成。
|
||||
- 报告有没有聚焦 payload。
|
||||
- 有没有使用证据工具。
|
||||
- 有没有把无关告警展开成主诊断对象。
|
||||
|
||||
Rule evaluation is stable, cheap, and easy to explain. It also avoids adding another hidden model call to the AIOps flow before the current trace contract is mature.
|
||||
规则验证稳定、便宜、容易解释,也不会在当前链路里额外引入一次隐藏模型调用。
|
||||
|
||||
## Verdicts
|
||||
## 3. 判定结果
|
||||
|
||||
The evaluator emits:
|
||||
当前评估器输出:
|
||||
|
||||
```text
|
||||
PASS
|
||||
@@ -40,14 +39,36 @@ WARN
|
||||
FAIL
|
||||
```
|
||||
|
||||
`FAIL` is reserved for critical issues such as a missing or too-short report. Missing payload focus terms or missing evidence tools currently produce `WARN`, because valid reports may use slightly different wording or evidence may be unavailable in a mock/demo environment.
|
||||
含义:
|
||||
|
||||
## Interview Answer
|
||||
- `PASS`:核心检查通过。
|
||||
- `WARN`:报告存在,但可能缺少 payload 关键词或证据工具。
|
||||
- `FAIL`:缺少最终报告、报告过短等关键问题。
|
||||
|
||||
If asked why AIOps has a verifier now:
|
||||
缺少 payload 关键词或证据工具先给 `WARN`,因为 demo/mock 环境下证据可能不可用,且报告措辞可能与 payload 字段不完全一致。
|
||||
|
||||
> Chat already has an LLM verifier because the user questions are open-ended. For AIOps, I started with a lighter rule-based verifier because the first quality checks are very concrete: payload focus, evidence coverage, and report completeness. The evaluation is persisted into `self_evaluation`, so the trace can show not only what the Agent did, but also whether the output passed basic quality gates.
|
||||
## 4. 面试回答
|
||||
|
||||
If asked why not use the Chat verifier directly:
|
||||
如果被问:为什么 AIOps 也需要验证器?
|
||||
|
||||
```text
|
||||
Chat 已经有 LLM Verifier,因为用户问题开放度高。
|
||||
AIOps 的第一阶段质量风险更明确:报告是否聚焦输入告警、是否使用证据工具、报告是否完整。
|
||||
所以我先做了轻量规则验证器,把结果写入 self_evaluation,让 Trace 不只展示 Agent 做了什么,也展示输出是否通过基础质量门。
|
||||
```
|
||||
|
||||
如果被问:为什么不直接复用 Chat Verifier?
|
||||
|
||||
```text
|
||||
AIOps 验证语义和 Chat 不一样。
|
||||
它要检查 alert scope、payload focus、证据工具覆盖,以及是否过度展开无关 active alerts。
|
||||
直接复用 Chat Verifier 会混淆这些语义。
|
||||
规则评估先提供稳定质量门,后续 AIOps LLM Verifier 可以基于同一套 trace contract 扩展。
|
||||
```
|
||||
|
||||
## 5. 后续增强
|
||||
|
||||
- 引入 AIOps LLM Verifier,逐条校验根因和建议是否有 evidence refs。
|
||||
- 把 rule evaluation 的 checks 在 Trace API 中结构化展示。
|
||||
- 将 payload scope violation 沉淀为 bad case。
|
||||
|
||||
> AIOps verification is different from Chat verification. It needs to check alert scope, evidence tool coverage, and whether unrelated active alerts were over-expanded. Reusing the Chat verifier directly would blur those semantics. The rule-based evaluator gives us a stable first quality gate; a later AIOps LLM verifier can build on the same trace contract.
|
||||
|
||||
@@ -1,62 +1,75 @@
|
||||
# AIOps Query Augmentation
|
||||
# AIOps 查询增强说明
|
||||
|
||||
## What Changed
|
||||
## 1. 改动是什么
|
||||
|
||||
Payload-targeted AIOps prompts now include a deterministic recommended knowledge query.
|
||||
AIOps 在 `PAYLOAD_TARGETED` 模式下,会从告警 payload 中稳定生成一条推荐知识库检索 query。
|
||||
|
||||
The query is built from the non-blank payload fields:
|
||||
参与拼接的非空字段:
|
||||
|
||||
```text
|
||||
alertName service severity description timeRange userRequest
|
||||
```
|
||||
|
||||
Example:
|
||||
示例:
|
||||
|
||||
```text
|
||||
HighCPUUsage payment-service P1 CPU usage is above 80% last_15m
|
||||
HighCPUUsage payment-service P1 CPU 使用率超过 80% last_15m
|
||||
```
|
||||
|
||||
## Why This Matters
|
||||
|
||||
AIOps payload fields contain high-value retrieval terms:
|
||||
|
||||
- alert name
|
||||
- service name
|
||||
- severity
|
||||
- symptom description
|
||||
- time range
|
||||
- operator request
|
||||
|
||||
Before this change, the Agent still had to invent its own `lookup_knowledge` query from the full prompt. That can work, but it may omit important terms such as the service name or alert name.
|
||||
|
||||
The new prompt makes the retrieval seed explicit:
|
||||
最终会进入 Prompt:
|
||||
|
||||
```text
|
||||
Recommended lookup_knowledge query: ...
|
||||
```
|
||||
|
||||
## Design Choice
|
||||
## 2. 为什么重要
|
||||
|
||||
This is prompt-level query augmentation, not hidden retrieval.
|
||||
AIOps payload 里包含高价值检索词:
|
||||
|
||||
I intentionally did not call `lookup_knowledge` automatically before the Agent runs. The project values traceability: tool calls should appear as Agent actions, with their inputs and outputs recorded in `tool_invocation`.
|
||||
- 告警名称。
|
||||
- 服务名。
|
||||
- 严重等级。
|
||||
- 症状描述。
|
||||
- 时间范围。
|
||||
- 用户补充请求。
|
||||
|
||||
So the design is:
|
||||
如果完全让 Agent 从长 Prompt 里自己组织检索 query,可能遗漏服务名或告警名。推荐 query 让检索种子更稳定。
|
||||
|
||||
## 3. 设计取舍
|
||||
|
||||
这是 Prompt 层 query augmentation,不是隐藏检索。
|
||||
|
||||
我没有在 Agent 运行前自动调用 `lookup_knowledge`,原因是项目强调可追踪性:工具调用应该由 Agent 显式发起,并记录到 `tool_invocation`。
|
||||
|
||||
当前设计:
|
||||
|
||||
```text
|
||||
AIOps payload
|
||||
-> deterministic recommended retrieval query
|
||||
-> Agent prompt
|
||||
-> Agent may call lookup_knowledge explicitly
|
||||
-> tool_invocation records the real retrieval action
|
||||
-> Agent 显式调用 lookup_knowledge
|
||||
-> tool_invocation 记录真实检索行为
|
||||
```
|
||||
|
||||
## Interview Answer
|
||||
## 4. 面试回答
|
||||
|
||||
If asked how AIOps payload improves RAG retrieval:
|
||||
如果被问:AIOps payload 怎么提升 RAG 检索?
|
||||
|
||||
> I do not replace the user query with a broad domain. I extract the high-signal alert terms from the payload, such as alertName, service, severity, symptom, and time range, and put them into a compact recommended lookup query. The Agent still calls `lookup_knowledge` explicitly, so the trace remains auditable, but the retrieval query is less dependent on model improvisation.
|
||||
```text
|
||||
我没有把告警 payload 粗暴替换成一个宽泛领域,而是提取 alertName、service、severity、description、timeRange 等高信号字段,拼成推荐的 lookup_knowledge query。
|
||||
Agent 仍然显式调用工具,所以 trace 仍然能看到真实检索行为,但 query 不再完全依赖模型临场发挥。
|
||||
```
|
||||
|
||||
If asked why not auto-call retrieval:
|
||||
如果被问:为什么不自动检索?
|
||||
|
||||
```text
|
||||
自动检索会在 Agent 真正决策前制造一份隐藏证据。
|
||||
这个项目的重点是可观测 Agent 执行,所以我选择 Prompt 层增强:给 Agent 一个更好的 query seed,但不改变工具调用必须显式可追踪的契约。
|
||||
```
|
||||
|
||||
## 5. 后续增强
|
||||
|
||||
- 将 recommended query 写入 trace 的结构化字段,便于对比 Agent 实际 query。
|
||||
- 对 payload 字段加权,例如 alertName/service 权重大于 timeRange。
|
||||
- 后续接入 Query Transformer 时,保留原始 query、推荐 query、改写 query 三者的可追踪关系。
|
||||
|
||||
> Auto-calling retrieval would create hidden evidence before the Agent actually decides to use a tool. For this project, explicit tool invocation is more important because the interview story is about observable Agent execution. Prompt-level augmentation gives the Agent a better query seed without changing the trace contract.
|
||||
|
||||
+70
-120
@@ -1,147 +1,97 @@
|
||||
# Architecture
|
||||
# 面试版系统架构
|
||||
|
||||
## 系统分层
|
||||
## 1. 系统分层
|
||||
|
||||
```text
|
||||
API Layer
|
||||
-> ChatController / DiagnosisTraceController
|
||||
|
||||
Agent Orchestration
|
||||
-> ChatService / AiOpsService
|
||||
|
||||
Tools
|
||||
-> lookupKnowledgeTool / queryLogs / queryMetrics / queryPrometheusAlerts
|
||||
|
||||
Persistence
|
||||
-> diagnosis_session / agent_step / tool_invocation
|
||||
|
||||
Trace
|
||||
-> GET /api/diagnosis/{sessionId}/trace
|
||||
```mermaid
|
||||
flowchart TB
|
||||
API["API 层\nChatController / DiagnosisTraceController / SearchController"] --> Service["应用服务层\nChatService / AiOpsService / DiagnosisTraceService"]
|
||||
Service --> Agent["Agent 编排层\nPlanner / Executor / Verifier / Supervisor"]
|
||||
Agent --> Tools["工具层\nlookup_knowledge / query_logs / query_metrics / Prometheus"]
|
||||
Tools --> RAG["RAG 检索\nL0 hint + VectorSearchService"]
|
||||
RAG --> VectorStore["Spring AI VectorStore"]
|
||||
RAG --> SDK["Milvus SDK fallback"]
|
||||
Agent --> Trace["Trace 持久化"]
|
||||
Tools --> Trace
|
||||
Trace --> Session["diagnosis_session"]
|
||||
Trace --> Step["agent_step"]
|
||||
Trace --> Invocation["tool_invocation"]
|
||||
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
```
|
||||
|
||||
## Chat 链路
|
||||
## 2. Chat 链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
User[User Question] --> ChatAPI[POST /api/chat]
|
||||
ChatAPI --> Strategy[ChatService.executeChatWithStrategy]
|
||||
Strategy --> Complexity{QuestionComplexity}
|
||||
Complexity -->|simple| Single[ReactAgent]
|
||||
Complexity -->|complex| Planner[Planner Agent]
|
||||
Planner --> Executor[Executor Agent]
|
||||
Executor --> Tools[Evidence Tools]
|
||||
User["用户问题"] --> ChatAPI["POST /api/chat"]
|
||||
ChatAPI --> Strategy["ChatService.executeChatWithStrategy"]
|
||||
Strategy --> Complexity{"复杂问题?"}
|
||||
Complexity -->|否| Single["单 ReactAgent 快速回答"]
|
||||
Complexity -->|是| Planner["chat_planner"]
|
||||
Planner --> Executor["chat_executor"]
|
||||
Executor --> Tools["证据工具"]
|
||||
Tools --> Executor
|
||||
Executor --> Verifier[Verifier Agent]
|
||||
Verifier --> Answer[Final Answer]
|
||||
Answer --> Session[diagnosis_session]
|
||||
Planner --> Steps[agent_step]
|
||||
Executor --> Steps
|
||||
Verifier --> Steps
|
||||
Tools --> Invocations[tool_invocation]
|
||||
Session --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
||||
Steps --> Trace
|
||||
Invocations --> Trace
|
||||
Executor --> Verifier["chat_verifier"]
|
||||
Verifier --> Decision{"PASS / LOW_CONFID / REJECT"}
|
||||
Decision --> Answer["最终答复"]
|
||||
Planner --> Step["agent_step"]
|
||||
Executor --> Step
|
||||
Verifier --> Step
|
||||
Tools --> Invocation["tool_invocation"]
|
||||
Answer --> Session["diagnosis_session"]
|
||||
```
|
||||
|
||||
关键代码:
|
||||
讲解重点:
|
||||
|
||||
- `ChatController.chat(...)`
|
||||
- `ChatService.executeChatWithStrategy(...)`
|
||||
- `ChatService.executeChatComplex(...)`
|
||||
- `AgentLoggingHook`
|
||||
- `ToolInvocationRecorder`
|
||||
- `DiagnosisTraceService.getTrace(...)`
|
||||
- Planner 拆解问题和排查方向。
|
||||
- Executor 必须通过工具收集证据。
|
||||
- Verifier 只基于 `tool_trace_summary` 校验答案,不做新检索。
|
||||
- Trace API 能回放模型步骤和工具证据。
|
||||
|
||||
## AIOps 链路
|
||||
## 3. AIOps 链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Alert[Alert Payload or Empty Request] --> AiOpsAPI[POST /api/ai_ops]
|
||||
AiOpsAPI --> SessionEvent[SSE session event]
|
||||
AiOpsAPI --> AiOpsService[AiOpsService.executeAiOpsAnalysis]
|
||||
AiOpsService --> PromptMode{Payload?}
|
||||
PromptMode -->|yes| Targeted[PAYLOAD_TARGETED]
|
||||
PromptMode -->|no| Discovery[AUTO_DISCOVERY]
|
||||
Targeted --> Supervisor[ai_ops_supervisor]
|
||||
Alert["告警 payload 或空请求"] --> API["POST /api/ai_ops"]
|
||||
API --> AiOps["AiOpsService"]
|
||||
AiOps --> Mode{"是否有 payload?"}
|
||||
Mode -->|有| Targeted["PAYLOAD_TARGETED\n聚焦输入告警"]
|
||||
Mode -->|无| Discovery["AUTO_DISCOVERY\n先发现活跃告警"]
|
||||
Targeted --> Supervisor["ai_ops_supervisor"]
|
||||
Discovery --> Supervisor
|
||||
Supervisor --> Planner[planner_agent]
|
||||
Supervisor --> Executor[executor_agent]
|
||||
Planner --> Tools[Prometheus / Logs / Knowledge]
|
||||
Supervisor --> Planner["planner_agent"]
|
||||
Supervisor --> Executor["executor_agent"]
|
||||
Planner --> Tools["Prometheus / 日志 / 知识库"]
|
||||
Executor --> Tools
|
||||
Tools --> Report[Alert Report]
|
||||
Report --> Persist[diagnosis_session.answer]
|
||||
Planner --> Steps[agent_step]
|
||||
Executor --> Steps
|
||||
Tools --> Invocations[tool_invocation]
|
||||
Persist --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
||||
Steps --> Trace
|
||||
Invocations --> Trace
|
||||
Tools --> Report["告警分析报告"]
|
||||
Report --> Eval["AiOpsRuleEvaluationService"]
|
||||
Eval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
|
||||
```
|
||||
|
||||
关键代码:
|
||||
讲解重点:
|
||||
|
||||
- `ChatController.aiOps(...)`
|
||||
- `AIOpsRequest`
|
||||
- `AiOpsService.resolveSessionId(...)`
|
||||
- `AiOpsService.buildTaskPrompt(...)`
|
||||
- `AiOpsService.hasAlertPayload(...)`
|
||||
- `AiOpsService.persistFinalReport(...)`
|
||||
- AIOps 有明确产品边界:有 payload 时必须聚焦该告警。
|
||||
- payload 字段会生成 recommended `lookup_knowledge` query。
|
||||
- 当前 AIOps 先用规则评估做质量门,后续再扩展 LLM Verifier。
|
||||
|
||||
## Trace 数据模型
|
||||
## 4. Trace 数据模型
|
||||
|
||||
### `diagnosis_session`
|
||||
| 表 | 作用 |
|
||||
|---|---|
|
||||
| `diagnosis_session` | 一次诊断的主记录:问题、状态、答案、自评估、反馈 |
|
||||
| `agent_step` | Agent 模型调用记录:输入、输出、耗时、token、是否有工具调用 |
|
||||
| `tool_invocation` | 工具调用事实:工具名、入参、输出预览、检索层、相关性、成功状态 |
|
||||
|
||||
记录一次诊断会话的主信息:
|
||||
## 5. 为什么 Trace 是核心
|
||||
|
||||
- `session_id`
|
||||
- `query`
|
||||
- `status`
|
||||
- `agent_flow`
|
||||
- `total_duration_ms`
|
||||
- `total_token_count`
|
||||
- `step_count`
|
||||
- `tool_call_count`
|
||||
- `answer`
|
||||
- `self_evaluation`
|
||||
- `feedback`
|
||||
故障诊断系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
|
||||
|
||||
### `agent_step`
|
||||
- 模型为什么这么答。
|
||||
- 调了哪些工具。
|
||||
- 工具返回了什么证据。
|
||||
- Verifier 如何判断答案可信度。
|
||||
- 用户反馈如何回写到同一个 session。
|
||||
|
||||
记录 Agent 模型调用过程:
|
||||
这就是它区别于普通 Chatbot 的地方。
|
||||
|
||||
- `session_id`
|
||||
- `step_index`
|
||||
- `agent_name`
|
||||
- `model_input`
|
||||
- `model_output`
|
||||
- `thought`
|
||||
- `has_tool_call`
|
||||
- `duration_ms`
|
||||
- `token_count`
|
||||
|
||||
### `tool_invocation`
|
||||
|
||||
记录真实工具调用:
|
||||
|
||||
- `session_id`
|
||||
- `tool_name`
|
||||
- `input_params`
|
||||
- `output_preview`
|
||||
- `output_length`
|
||||
- `retrieval_layer`
|
||||
- `relevance_level`
|
||||
- `duration_ms`
|
||||
- `success`
|
||||
- `error_message`
|
||||
|
||||
## 为什么 trace 是核心
|
||||
|
||||
Agent 系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
|
||||
|
||||
- 模型为什么这么答
|
||||
- 调了哪些工具
|
||||
- 工具返回了什么证据
|
||||
- Verifier 如何判断答案可信度
|
||||
- 用户反馈如何回写到同一个 session
|
||||
|
||||
这就是项目区别于普通 Chatbot 的地方。
|
||||
|
||||
+45
-32
@@ -1,18 +1,20 @@
|
||||
# Interview Demo Script
|
||||
# 面试演示脚本
|
||||
|
||||
## 30 秒开场
|
||||
## 1. 30 秒开场
|
||||
|
||||
这是一个 Agent Engineering 项目,场景是企业故障诊断。它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。项目重点不是单次回答,而是把多 Agent 执行、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 trace。
|
||||
```text
|
||||
这是一个 Agent 工程项目,场景是企业故障诊断。
|
||||
它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。
|
||||
项目重点不是单次模型回答,而是把 Agent 编排、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 Trace。
|
||||
```
|
||||
|
||||
## Demo 准备
|
||||
|
||||
启动服务:
|
||||
## 2. 启动服务
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
确认服务地址:
|
||||
服务地址:
|
||||
|
||||
```text
|
||||
http://localhost:9900
|
||||
@@ -24,11 +26,9 @@ http://localhost:9900
|
||||
- CLS 日志使用 mock 数据。
|
||||
- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。
|
||||
|
||||
## Demo 1: Chat 诊断
|
||||
## 3. Demo 1:Chat 诊断
|
||||
|
||||
目标:展示普通用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 trace。
|
||||
|
||||
请求:
|
||||
目标:展示用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 Trace。
|
||||
|
||||
```powershell
|
||||
$sessionId = "interview-chat-payment-timeout-001"
|
||||
@@ -46,13 +46,13 @@ Invoke-RestMethod `
|
||||
|
||||
讲解点:
|
||||
|
||||
- `ChatController` 把请求交给 `ChatService.executeChatWithStrategy(...)`。
|
||||
- 简单问题走单 ReactAgent,复杂问题走 `Planner -> Executor -> Verifier`。
|
||||
- Executor 可以调用知识库、日志、指标等工具。
|
||||
- Verifier 会基于工具证据生成 groundedness 评估。
|
||||
- 最终会写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
|
||||
- `ChatService` 会根据问题复杂度选择轻量回答或复杂 Agent 流程。
|
||||
- 复杂问题走 `Planner -> Executor -> Verifier`。
|
||||
- Executor 调用知识库、日志、指标等证据工具。
|
||||
- Verifier 基于 `tool_trace_summary` 生成 groundedness 评估。
|
||||
- 最终写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
|
||||
|
||||
查询 trace:
|
||||
查询 Trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
@@ -67,11 +67,9 @@ Invoke-RestMethod `
|
||||
- `data.toolInvocations` 中能看到证据工具
|
||||
- `data.session.selfEvaluation` 中有 verifier 结果
|
||||
|
||||
## Demo 2: AIOps 告警诊断
|
||||
## 4. Demo 2:AIOps 告警诊断
|
||||
|
||||
目标:展示告警 payload 如何触发 AIOps 入口,并且报告只聚焦目标告警。
|
||||
|
||||
请求:
|
||||
目标:展示告警 payload 如何触发 AIOps,并且报告聚焦目标告警。
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "interview-aiops-payment-cpu-001"
|
||||
@@ -99,9 +97,9 @@ Invoke-WebRequest `
|
||||
- `AiOpsService` 根据 payload 判断模式:
|
||||
- `PAYLOAD_TARGETED`:聚焦传入告警。
|
||||
- `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。
|
||||
- AIOps 暂时不加 Verifier,先保证告警入口、证据工具和 trace 可用。
|
||||
- AIOps 当前用 rule evaluation 检查报告完整性、payload 聚焦和证据工具覆盖。
|
||||
|
||||
查询 trace:
|
||||
查询 Trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
@@ -114,19 +112,34 @@ Invoke-RestMethod `
|
||||
- `data.session.agentFlow = AI_OPS`
|
||||
- `data.session.answer` 有最终告警报告
|
||||
- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge`
|
||||
- 报告有 `HighCPUUsage/payment-service` 的完整根因分析
|
||||
- 其他 active alerts 只作为相关风险出现,不展开成独立根因章节
|
||||
- 报告主线聚焦 `HighCPUUsage/payment-service`
|
||||
|
||||
## MySQL 验证
|
||||
## 5. Demo 3:反馈闭环
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session ORDER BY id DESC LIMIT 5"
|
||||
$feedback = @{
|
||||
sessionId = $sessionId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/feedback" `
|
||||
-ContentType "application/json" `
|
||||
-Body $feedback
|
||||
```
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name"
|
||||
讲解点:
|
||||
|
||||
- feedback 写回同一个 `diagnosis_session`。
|
||||
- `useful` 会沉淀 `case_library`。
|
||||
- `not_useful` 不改变 `status`,只作为质量信号。
|
||||
|
||||
## 6. 收尾总结
|
||||
|
||||
```text
|
||||
这个 Demo 展示的是完整 Agent 闭环:
|
||||
用户问题或告警 -> Agent 编排 -> 工具证据 -> 自评估 -> Trace 回放 -> 用户反馈 -> 案例沉淀。
|
||||
我关注的不是一次回答,而是这个回答能否被审计、验证和持续改进。
|
||||
```
|
||||
|
||||
## 收尾总结
|
||||
|
||||
这套 Demo 展示的是一个完整 Agent 系统,而不是一次模型问答:入口有明确场景边界,Agent 负责规划和执行,工具提供证据,Verifier 提供质量门,trace API 提供审计和复盘能力。AIOps 入口进一步证明它可以从用户问答扩展到事件驱动诊断。
|
||||
|
||||
@@ -1,99 +1,107 @@
|
||||
# Design Tradeoffs
|
||||
# 关键设计取舍
|
||||
|
||||
## 1. 为什么要做 trace,而不是只返回答案
|
||||
## 1. 为什么先做 Trace,而不是只返回答案
|
||||
|
||||
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答三件事:
|
||||
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答:
|
||||
|
||||
- 结论是什么
|
||||
- 证据来自哪里
|
||||
- 哪些步骤由哪个 Agent 完成
|
||||
- 结论是什么。
|
||||
- 证据来自哪里。
|
||||
- 哪些步骤由哪个 Agent 完成。
|
||||
- 如果答案不可靠,系统怎么降级。
|
||||
|
||||
因此项目把一次会话拆成:
|
||||
|
||||
- `diagnosis_session`:会话级摘要、最终答案、质量评估、反馈。
|
||||
- `agent_step`:Agent 模型输入输出、耗时、token 和工具调用标记。
|
||||
- `diagnosis_session`:会话摘要、最终答案、自评估、用户反馈。
|
||||
- `agent_step`:模型输入输出、耗时、token 和工具调用标记。
|
||||
- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。
|
||||
|
||||
这个设计牺牲了一些实现复杂度,但换来了可回放、可调试、可演示。
|
||||
代价是实现复杂度上升,收益是可回放、可调试、可演示。
|
||||
|
||||
## 2. 为什么 Chat 有 Verifier,AIOps 暂时没有
|
||||
## 2. 为什么 RAG 不直接隐藏在 Advisor 里
|
||||
|
||||
Chat 入口的问题更开放,用户可能要求复杂推理或跨领域结论,所以 Verifier 是必要的质量门。当前 Chat 链路通过 `Planner -> Executor -> Verifier` 固定流程,把 groundedness 和 facts checked 写入 `self_evaluation`。
|
||||
Spring AI Advisor 可以让 RAG 更隐式,但本项目的核心是 Agent 证据链。`lookup_knowledge` 必须作为显式工具调用出现,这样 Trace 里才能看到:
|
||||
|
||||
AIOps 当前阶段先不加 Verifier,原因是:
|
||||
- Agent 什么时候决定检索。
|
||||
- 用了什么 query。
|
||||
- 命中了哪些文档。
|
||||
- 相关性等级是什么。
|
||||
- 证据如何支撑最终答案。
|
||||
|
||||
- AIOps 刚完成从“自动跑告警”到“可追踪告警入口”的改造。
|
||||
- 先要确认告警 payload、工具证据、最终报告和 trace 能闭环。
|
||||
- AIOps Verifier 的规则不同于 Chat Verifier,需要检查告警 scope、证据覆盖和处置建议,不宜直接复用。
|
||||
|
||||
后续可以做 lightweight AIOps Verifier,检查报告是否聚焦 payload、是否引用工具证据、是否误展开无关告警。
|
||||
|
||||
## 3. 为什么 AIOps payload scope 先用 prompt 控制
|
||||
|
||||
运行验证发现:传入 `HighCPUUsage/payment-service` 后,Agent 仍可能把 mock Prometheus 返回的所有 active alerts 都展开分析。这个问题的本质是任务边界不清晰。
|
||||
|
||||
当前选择 prompt-level scope control:
|
||||
|
||||
- 有 payload:`PAYLOAD_TARGETED`,最终报告围绕传入告警。
|
||||
- 无 payload:`AUTO_DISCOVERY`,先调用 `queryPrometheusAlerts` 自动发现告警。
|
||||
|
||||
没有先做 Java 侧过滤,是因为:
|
||||
|
||||
- 过滤工具结果会降低 Agent 发现关联风险的能力。
|
||||
- 目前需要的是报告主线聚焦,而不是完全屏蔽上下文。
|
||||
- Prompt 改动小,风险低,能保留 Agent 灵活性。
|
||||
|
||||
已验证结果:主报告有 `HighCPUUsage/payment-service` 的完整根因分析,`HighMemoryUsage` 和 `SlowResponse` 只作为相关风险出现。
|
||||
|
||||
## 4. 为什么用 `tool_invocation` 统计真实工具调用次数
|
||||
|
||||
早期可以通过 `agent_step.hasToolCall` 粗略判断是否调用工具,但它统计的是“哪些模型步骤包含工具调用”,不是“真实调用了几次工具”。
|
||||
|
||||
现在 `tool_call_count` 来自:
|
||||
所以当前设计是:
|
||||
|
||||
```text
|
||||
ToolInvocationRepository.countBySessionId(sessionId)
|
||||
Executor -> lookup_knowledge -> VectorSearchService -> VectorStore / SDK fallback
|
||||
```
|
||||
|
||||
这样更符合 trace 语义:
|
||||
这牺牲了一点框架自动化,但保留了可审计性。
|
||||
|
||||
- 一个 step 可能调用多个工具。
|
||||
- 工具可能来自不同来源:知识库、日志、指标、Prometheus。
|
||||
- 面试时可以把 `tool_call_count` 和 trace 中返回的工具明细对上。
|
||||
## 3. 为什么 L0 只做 hint,不直接返回
|
||||
|
||||
## 5. 为什么保留 mock Prometheus 和 mock CLS
|
||||
旧版 L0 关键词唯一命中时可能直接跳过 L1。这个策略速度快,但风险是:关键词子串命中不等于最终语义相关。
|
||||
|
||||
面试 Demo 最怕不稳定。真实 Prometheus、日志平台和线上故障都有不可控因素,所以 MVP profile 保留 mock 工具:
|
||||
当前改成:
|
||||
|
||||
- `prometheus.mock-enabled=true`
|
||||
- `cls.mock-enabled=true`
|
||||
```text
|
||||
L0 = domain/entity hint
|
||||
L1 = semantic retrieval
|
||||
postprocess = evidence shaping + trace
|
||||
```
|
||||
|
||||
这样可以稳定复现:
|
||||
L0 仍然有价值:错误码、服务名、告警名、指标名都很适合做精确 hint。但最终证据仍需要 L1 和后处理支撑。
|
||||
|
||||
- `HighCPUUsage/payment-service`
|
||||
- `HighMemoryUsage/order-service`
|
||||
- `SlowResponse/user-service`
|
||||
- system-metrics、application-logs、database-slow-query 等日志证据
|
||||
## 4. 为什么保留 Milvus SDK fallback
|
||||
|
||||
这不是逃避真实集成,而是把“Agent 编排和证据追踪”作为面试演示的主目标。
|
||||
Spring AI VectorStore 是当前读路径主方向,但 SDK fallback 没有删除,原因有三点:
|
||||
|
||||
## 6. 为什么把面试材料单独放 `interview/`
|
||||
- 迁移安全:旧 SDK 路径已经被验证过。
|
||||
- 运行韧性:VectorStore 配置、schema、collection 出问题时可以回退。
|
||||
- 面试稳定:检索抽象迁移不应该破坏主 demo。
|
||||
|
||||
`mvp/` 是持续迭代现场,包含过程文档、验收记录和 runbook。面试材料的目标不同,它应该是可讲、可演示、可评估的展示层。
|
||||
这不是“没有迁完”,而是分阶段迁移:先稳定读路径,再决定是否迁移写入和索引。
|
||||
|
||||
因此:
|
||||
## 5. 为什么 Chat 有 Verifier,AIOps 先用规则评估
|
||||
|
||||
- `mvp/` 保留真实演进材料。
|
||||
- `devflow/` 保留决策沉淀。
|
||||
- `interview/` 只组织面试叙事和演示脚本。
|
||||
Chat 问题更开放,容易出现跨领域推理,所以需要 LLM Verifier 做 groundedness 校验。
|
||||
|
||||
这样后续继续做 AIOps Verifier、UI、更多工具集成时,不会污染面试讲稿。
|
||||
AIOps 当前优先解决更具体的问题:
|
||||
|
||||
## 7. 可以主动承认的限制
|
||||
- 最终报告是否存在。
|
||||
- payload 模式是否聚焦输入告警。
|
||||
- 是否使用了证据工具。
|
||||
- 是否把无关活跃告警展开成主根因。
|
||||
|
||||
- AIOps 还没有 Verifier。
|
||||
- Prompt-level scope control 不能做到强约束,只能通过 trace 和测试观察遵循情况。
|
||||
- 当前 mock 数据适合 demo,不代表生产接入已经完成。
|
||||
- Hikari 连接池已经加了短生命周期和 keepalive,但真实生产还需要按数据库 wait_timeout 和连接数预算调优。
|
||||
这些用规则就能稳定检查。后续可以在同一个 `self_evaluation` 容器下增加 AIOps LLM Verifier。
|
||||
|
||||
## 6. 为什么 AIOps payload scope 先用 Prompt + Rule
|
||||
|
||||
真实告警环境里可能同时有多个 active alerts。用户传入 `HighCPUUsage/payment-service` 时,Agent 如果把所有告警都展开分析,报告会跑偏。
|
||||
|
||||
当前选择:
|
||||
|
||||
- Prompt 中加入 `PAYLOAD_TARGETED`。
|
||||
- 从 payload 生成 recommended `lookup_knowledge` query。
|
||||
- 用 `AiOpsRuleEvaluationService` 检查报告是否聚焦输入告警。
|
||||
|
||||
没有先做硬过滤,是因为有些相关告警可以作为风险背景。目标不是屏蔽上下文,而是控制主诊断对象。
|
||||
|
||||
## 7. 为什么反馈不改 status
|
||||
|
||||
`status` 表示执行状态,`feedback` 表示用户评价。一个执行成功但用户觉得没用的诊断,应该是:
|
||||
|
||||
```text
|
||||
status = SUCCESS
|
||||
feedback = not_useful
|
||||
```
|
||||
|
||||
这样才能区分系统异常和质量问题。`useful` 反馈会沉淀 `case_library`,`not_useful` 作为 bad case 信号保留。
|
||||
|
||||
## 8. 可以主动承认的限制
|
||||
|
||||
- AIOps 还没有完整 LLM Verifier。
|
||||
- RAG 还没有 hybrid search、rerank、邻居 chunk 扩展。
|
||||
- `case_library` 的 rootCause/solution 仍需要结构化抽取。
|
||||
- `tool_invocation.step_id` 关联还可以更严格。
|
||||
- `mvp-demo` profile 使用 mock 日志和指标,主要服务稳定面试演示。
|
||||
|
||||
主动讲清这些限制,能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
|
||||
|
||||
主动讲清这些限制,反而能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
# RAG Breadcrumb Embedding Acceptance
|
||||
# RAG Breadcrumb Embedding 验收说明
|
||||
|
||||
## What Changed
|
||||
## 1. 改动是什么
|
||||
|
||||
The indexing path now builds embedding text from chunk structure plus content:
|
||||
索引路径现在构造 embedding 文本时,不只使用 chunk 内容,还会把结构上下文拼进去:
|
||||
|
||||
```text
|
||||
Title: {title}
|
||||
@@ -11,56 +11,81 @@ Content:
|
||||
{content}
|
||||
```
|
||||
|
||||
The stored Milvus `content` field remains the original chunk content. This keeps display and evidence output clean while allowing the vector to carry section-level semantics.
|
||||
Milvus 中存储的 `content` 字段仍然保留原始 chunk 内容。这样展示和证据输出保持干净,而向量本身携带章节语义。
|
||||
|
||||
## Why Reindex Is Required
|
||||
## 2. 为什么必须重新索引
|
||||
|
||||
Embeddings are materialized at index time. Existing vectors were generated from the previous content-only text, so they cannot benefit from `title` and `breadcrumb` until the knowledge base is reindexed.
|
||||
Embedding 是索引时物化的。已有向量是用旧的 content-only 文本生成的,所以只有代码变化并不会改变线上检索结果。
|
||||
|
||||
This is the key acceptance point:
|
||||
验收关键点:
|
||||
|
||||
```text
|
||||
code change alone != live retrieval changed
|
||||
code change + reindex + live query report = accepted behavior
|
||||
只改代码 != live retrieval 已变化
|
||||
代码改动 + 重新索引 + live query report = 行为验收完成
|
||||
```
|
||||
|
||||
## How To Validate
|
||||
## 3. 如何验证
|
||||
|
||||
1. Start the Spring Boot application.
|
||||
2. Reindex the knowledge base through the existing indexing path.
|
||||
3. Run:
|
||||
1. 启动 Spring Boot 应用。
|
||||
2. 通过现有索引路径重新索引知识库。
|
||||
3. 运行:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py
|
||||
```
|
||||
|
||||
The script writes:
|
||||
脚本输出:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/reports/live-post-reindex.json
|
||||
eval/rag-retrieval/reports/live-post-reindex.md
|
||||
```
|
||||
|
||||
The default cases cover:
|
||||
默认覆盖:
|
||||
|
||||
- RAG chunk context questions where breadcrumb matters.
|
||||
- Diagnosis flow questions where section path matters.
|
||||
- `ERR_TIMEOUT` exact error-code retrieval.
|
||||
- MySQL connection pool troubleshooting.
|
||||
- AIOps payment-service latency alert retrieval.
|
||||
- breadcrumb 敏感的 RAG chunk context query。
|
||||
- 需要章节路径的诊断流程问题。
|
||||
- `ERR_TIMEOUT` 精确错误码检索。
|
||||
- MySQL 连接池排障。
|
||||
- AIOps payment-service 延迟告警检索。
|
||||
|
||||
## What To Look For
|
||||
## 4. 看什么结果
|
||||
|
||||
For breadcrumb-sensitive cases, inspect whether top candidates expose expected `title` and `breadcrumb` values in the report.
|
||||
对 breadcrumb 敏感 case:
|
||||
|
||||
For core troubleshooting cases, check that result counts and top candidates remain stable. The goal is not to prove a full benchmark; it is to prove that reindexing did not obviously break important demo retrieval paths.
|
||||
- top candidates 是否暴露预期 `title`。
|
||||
- top candidates 是否暴露预期 `breadcrumb`。
|
||||
- 命中内容是否能看出所属章节。
|
||||
|
||||
## Interview Answer
|
||||
对核心排障 case:
|
||||
|
||||
If asked how I verified the breadcrumb embedding change:
|
||||
- 结果数量是否稳定。
|
||||
- top candidates 是否仍然命中核心文档。
|
||||
- 没有因为拼接 title/breadcrumb 导致核心检索退化。
|
||||
|
||||
> I separated deterministic regression from live acceptance. The offline fixture baseline still runs without services. But because embedding changes only affect newly indexed vectors, I added a live post-reindex acceptance script. It calls the real `/api/search/similar` endpoint against representative breadcrumb-sensitive, troubleshooting, and AIOps queries, then writes JSON and Markdown reports. This lets me prove both that the code changed and that the live vector collection was refreshed.
|
||||
## 5. 面试回答
|
||||
|
||||
If asked why the script does not reindex automatically:
|
||||
如果被问:你怎么验证 breadcrumb 参与 embedding 后真的生效?
|
||||
|
||||
```text
|
||||
我把 deterministic regression 和 live acceptance 分开。
|
||||
离线 fixture baseline 不依赖服务,可以做稳定回归。
|
||||
但 embedding 改动只会影响新生成的向量,所以我另外加了 live post-reindex acceptance 脚本。
|
||||
脚本会调用真实 /api/search/similar,对 breadcrumb 敏感、排障和 AIOps query 生成 JSON/Markdown 报告。
|
||||
这样能证明代码改了,也能证明 live vector collection 已经刷新。
|
||||
```
|
||||
|
||||
如果被问:为什么脚本不自动 reindex?
|
||||
|
||||
```text
|
||||
reindex 会修改向量库,而且依赖环境中的知识库数据。
|
||||
我把 reindex 保持为显式动作,验收脚本只做读取验证。
|
||||
这样如果检索没有改善,我能区分是代码问题、索引未刷新,还是运行时检索行为问题。
|
||||
```
|
||||
|
||||
## 6. 后续增强
|
||||
|
||||
- 将 live acceptance 结果加入面试 Demo 输出。
|
||||
- 增加 breadcrumb hit rate 统计。
|
||||
- 对同章节 chunk 做邻居扩展,进一步利用 breadcrumb。
|
||||
|
||||
> Reindexing mutates the vector store and depends on environment-specific data. I kept mutation explicit and made the script validation-only. That makes failures easier to diagnose: if retrieval does not improve, I can distinguish code changes, reindex state, and runtime retrieval behavior.
|
||||
|
||||
+72
-154
@@ -1,47 +1,50 @@
|
||||
# RAG Refactor Story
|
||||
# RAG 重构故事
|
||||
|
||||
## The Starting Point
|
||||
## 1. 起点
|
||||
|
||||
The original RAG implementation was already usable for the MVP:
|
||||
原始 RAG 实现已经能支撑 MVP:
|
||||
|
||||
- Documents could be uploaded, chunked, embedded, and written to Milvus/Zilliz.
|
||||
- The Agent could call `lookup_knowledge` as an explicit tool.
|
||||
- AIOps diagnosis could retrieve troubleshooting knowledge during an alert workflow.
|
||||
- Tool invocations were persisted, so the retrieval step was visible in the execution trace.
|
||||
- 文档可以上传、切片、向量化,并写入 Milvus/Zilliz。
|
||||
- Agent 可以显式调用 `lookup_knowledge`。
|
||||
- AIOps 诊断能在告警流程里检索排障知识。
|
||||
- 工具调用会落到 `tool_invocation`,检索步骤可见。
|
||||
|
||||
But the design had several engineering problems:
|
||||
但它有几个工程问题:
|
||||
|
||||
- Retrieval was too SDK-specific. The business code directly owned many Milvus search details.
|
||||
- L0 and L1 responsibilities were blurry. L0 keyword matching could look like a final retrieval decision instead of a hint.
|
||||
- Chunk-level retrieval could lose section context when one section was split into multiple chunks.
|
||||
- Metadata such as `breadcrumb` existed, but it was not fully used in retrieval, filtering, or context reconstruction.
|
||||
- Retrieval quality was mostly checked by manual API calls and logs, not by repeatable cases.
|
||||
- 检索实现过于依赖 Milvus SDK,业务代码承担了太多底层搜索细节。
|
||||
- L0 和 L1 职责不清,L0 关键词命中容易被当作最终召回决策。
|
||||
- chunk 级检索容易丢失章节上下文。
|
||||
- `breadcrumb` 存在 metadata 中,但没有充分参与 embedding、filter 和上下文重建。
|
||||
- 检索质量主要靠手工接口和日志判断,缺少可重复的 golden cases。
|
||||
|
||||
So the refactor goal was not "replace everything with a framework." The goal was to move generic RAG infrastructure toward Spring AI while keeping the project-specific Agent evidence chain.
|
||||
所以重构目标不是“全盘替换成框架”,而是:
|
||||
|
||||
## How I Broke The Problem Down
|
||||
```text
|
||||
通用 RAG 基础设施交给 Spring AI,
|
||||
业务可观测链路保留在项目里。
|
||||
```
|
||||
|
||||
I treated this as a staged migration, because RAG touches the Agent tool layer, AIOps diagnosis, vector retrieval, evidence packing, and database traces.
|
||||
## 2. 我如何拆解问题
|
||||
|
||||
The first step was to establish a baseline. I added retrieval evaluation cases under `eval/rag-retrieval/` so future changes could be compared against known queries instead of judged only by intuition.
|
||||
我把迁移拆成几个阶段,因为 RAG 同时影响 Agent 工具层、AIOps、向量检索、证据打包和 Trace。
|
||||
|
||||
Then I clarified the retrieval roles:
|
||||
第一步是建立 baseline。`eval/rag-retrieval/` 中的 golden cases 用来对比后续改动,而不是只靠直觉判断检索有没有变好。
|
||||
|
||||
第二步是明确职责:
|
||||
|
||||
```text
|
||||
L0 = domain/entity hint
|
||||
L1 = semantic retrieval
|
||||
postprocess = evidence shaping and trace-friendly output
|
||||
postprocess = evidence shaping + trace-friendly output
|
||||
```
|
||||
|
||||
That means L0 is still valuable, but it should not bypass semantic retrieval as the default path. It is better used to extract service names, alert names, error codes, domains, and metadata hints.
|
||||
L0 仍然有价值,但不再默认绕过语义检索。它更适合提取服务名、告警名、错误码、领域和 metadata filter。
|
||||
|
||||
After that, I added evidence postprocessing. The Agent should not just receive raw chunks; it should receive structured evidence with source, title, breadcrumb, score, hit reason, and content. This makes the result easier to inspect and easier to explain in an interview.
|
||||
第三步是增强 evidence 输出。Agent 不应该只拿到 raw chunk,而应该拿到带 source、title、breadcrumb、score、hit reason 的证据块。
|
||||
|
||||
Finally, I integrated Spring AI `VectorStore` as the main read path while preserving the original Milvus SDK implementation as fallback.
|
||||
最后,我把 Spring AI `VectorStore` 接入为读取主路径,同时保留原 Milvus SDK 作为 fallback。
|
||||
|
||||
## Current Architecture
|
||||
|
||||
The current retrieval path is:
|
||||
## 3. 当前架构
|
||||
|
||||
```text
|
||||
Agent / API
|
||||
@@ -50,160 +53,75 @@ Agent / API
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
-> evidence postprocess
|
||||
-> relevance normalization
|
||||
-> tool_invocation trace
|
||||
```
|
||||
|
||||
`VectorSearchService` is still the public retrieval facade. This is deliberate: the Agent tool layer does not need to know whether the underlying retrieval engine is SDK-based or Spring AI-based.
|
||||
`VectorSearchService` 仍然是公共检索门面。Agent 工具层不需要知道底层是 SDK 还是 Spring AI。
|
||||
|
||||
The supported retrieval modes are:
|
||||
支持三种模式:
|
||||
|
||||
```text
|
||||
auto -> try Spring AI VectorStore, fallback to SDK
|
||||
spring-ai -> force Spring AI VectorStore
|
||||
sdk -> force Milvus SDK
|
||||
auto -> 优先 Spring AI VectorStore,失败后 fallback 到 SDK
|
||||
spring-ai -> 强制 Spring AI VectorStore
|
||||
sdk -> 强制 Milvus SDK
|
||||
```
|
||||
|
||||
This keeps the migration reversible and testable.
|
||||
## 4. 关键取舍
|
||||
|
||||
## Key Tradeoffs
|
||||
### 保留显式工具
|
||||
|
||||
### Keep The Explicit Tool
|
||||
我没有把检索藏进 Spring AI Advisor。原因是这个项目强调 Agent 执行可见性:`lookup_knowledge` 的 query、命中文档、相关性和证据预览都要进入 Trace。
|
||||
|
||||
I did not hide retrieval inside a Spring AI Advisor.
|
||||
### 保留 SDK fallback
|
||||
|
||||
For this project, `lookup_knowledge` is part of the Agent execution story. It records what query was used, which evidence was retrieved, how relevant it looked, and how it supported diagnosis. If retrieval is hidden inside an advisor, the answer may still work, but the audit trail becomes harder to show.
|
||||
SDK fallback 不是废代码,而是迁移安全网。实际验证时,第一次 VectorStore 指向了错误 collection,`auto` 模式 fallback 到 SDK 后仍能返回结果。修正 collection 后,Spring AI 路径成为主路径。
|
||||
|
||||
### Keep SDK Fallback
|
||||
### L0 降权
|
||||
|
||||
The SDK path is not dead code. It is a safety net during migration.
|
||||
生产事故中经常有精确标识:错误码、告警名、服务名、指标名。L0 适合做 hint,但不应该做最终裁判。
|
||||
|
||||
This proved useful during live validation. The first VectorStore run pointed at the wrong collection name, but `auto` mode fell back to SDK and still returned results. After the collection was corrected to `biz`, the Spring AI path worked as the main path.
|
||||
### 分数语义拆开
|
||||
|
||||
### Keep L0, But Reduce Its Authority
|
||||
SDK 使用 L2 distance,Spring AI 暴露 similarity。混在一个字段里会让 relevance normalization 出错。
|
||||
|
||||
L0 is worth keeping because production incidents often contain exact identifiers:
|
||||
|
||||
- error code
|
||||
- alert name
|
||||
- service name
|
||||
- metric name
|
||||
- domain tag
|
||||
|
||||
But L0 should not be the final judge of retrieval quality. Its role is now closer to domain hint, entity extraction, metadata filtering, and explainability signal.
|
||||
|
||||
### Split Score Semantics
|
||||
|
||||
The old SDK path used L2 distance. Spring AI exposes similarity. Treating those as the same number would quietly break relevance normalization.
|
||||
|
||||
So the result separates:
|
||||
当前拆成:
|
||||
|
||||
```text
|
||||
score -> compatibility score used by existing logic
|
||||
rawScore -> raw score from the retrieval implementation
|
||||
scoreLabel -> semantic label for rawScore
|
||||
score -> 兼容旧逻辑的距离型分数
|
||||
rawScore -> 底层原始分数
|
||||
scoreLabel -> rawScore 的语义
|
||||
```
|
||||
|
||||
For SDK:
|
||||
### 暂不迁移写入
|
||||
|
||||
写入和索引仍走 SDK。这是有意分阶段:先验证读路径,再评估 `VectorStore.add(...)` 是否适合现有 metadata 和 chunk 模型。
|
||||
|
||||
## 5. 验证方式
|
||||
|
||||
我用了三层验证:
|
||||
|
||||
- 单元测试:SDK mode、Spring AI mode、auto fallback、category filter、distance metadata mapping。
|
||||
- Live API:`GET /api/search/similar?query=ERR_TIMEOUT&topK=3`。
|
||||
- 代表性 query 对比:错误码、支付超时、MySQL 连接池、AIOps 告警式 query、抽象 RAG 设计问题。
|
||||
|
||||
核心排障和 AIOps query 在 SDK 与 VectorStore 下 top3 一致。差异主要集中在抽象设计类问题和 metadata taxonomy,这些被记录为后续质量工作。
|
||||
|
||||
## 6. 面试短版
|
||||
|
||||
```text
|
||||
score = L2 distance
|
||||
rawScore = L2 distance
|
||||
scoreLabel = l2_distance
|
||||
这个 RAG 系统最初是基于 Milvus SDK 的自研 MVP。它能跑,但底层检索细节过多地散落在业务代码里,L0/L1 职责也不够清晰。
|
||||
我按阶段重构:先加 retrieval baseline,再把 L0 降级为 domain/entity hint,再增强 evidence postprocess,最后把读取主路径切到 Spring AI VectorStore,并保留 SDK fallback。
|
||||
我没有把 lookup_knowledge 替换成隐式 Advisor,因为这个项目的核心是可追踪 Agent:面试官可以看到什么时候检索、检索了什么、证据如何支撑诊断。
|
||||
```
|
||||
|
||||
For VectorStore:
|
||||
## 7. 可主动承认的不足
|
||||
|
||||
```text
|
||||
score = Milvus metadata.distance when available
|
||||
rawScore = Spring AI similarity
|
||||
scoreLabel = similarity
|
||||
```
|
||||
- metadata taxonomy 还需要清理,例如 `database` 与 `infrastructure`。
|
||||
- 抽象设计问题可能需要 query rewrite 或更好的文档索引。
|
||||
- 邻居 chunk / 同章节上下文扩展还不完整。
|
||||
- rerank、RRF、BM25、hybrid retrieval 还没有接入。
|
||||
- 写入路径仍使用 SDK。
|
||||
|
||||
This makes the migration inspectable instead of hiding score changes behind one overloaded field.
|
||||
这些不是当前迁移阻塞项,而是后续检索质量优化方向。
|
||||
|
||||
### Do Not Migrate Writes Yet
|
||||
|
||||
Writes and indexing still use the SDK path.
|
||||
|
||||
That is intentional. Migrating reads and writes at the same time would make debugging harder. The read path can be validated first; write-path migration can happen later if Spring AI `VectorStore.add(...)` fits the existing metadata and chunk model.
|
||||
|
||||
## Validation Story
|
||||
|
||||
I validated the refactor at multiple levels.
|
||||
|
||||
Unit tests cover:
|
||||
|
||||
- SDK mode.
|
||||
- Spring AI mode.
|
||||
- `auto` fallback.
|
||||
- category filter behavior.
|
||||
- distance metadata mapping.
|
||||
|
||||
Live API verification used:
|
||||
|
||||
```text
|
||||
GET /api/search/similar?query=ERR_TIMEOUT&topK=3
|
||||
```
|
||||
|
||||
Logs confirmed when the Spring AI VectorStore path was used and when fallback happened.
|
||||
|
||||
Then I compared SDK and VectorStore retrieval quality on representative queries:
|
||||
|
||||
| Query Type | Result |
|
||||
| --- | --- |
|
||||
| exact error code | same top3 |
|
||||
| payment-service timeout | same top3 |
|
||||
| MySQL connection pool | same top3 |
|
||||
| AIOps alert-style query | same top3 |
|
||||
| abstract RAG design query | same top1, VectorStore returned fewer tail results |
|
||||
| category filter | both returned zero because metadata taxonomy did not match |
|
||||
|
||||
The acceptance decision was that Spring AI VectorStore is good enough for the current MVP read path, with SDK fallback preserved.
|
||||
|
||||
## Known Gaps
|
||||
|
||||
The refactor improved the architecture, but it did not solve every retrieval-quality problem.
|
||||
|
||||
Known gaps:
|
||||
|
||||
- Metadata taxonomy still needs cleanup, for example `database` vs `infrastructure`.
|
||||
- Abstract design questions may need query rewriting or better indexed interview/devflow documents.
|
||||
- Chunk context reconstruction is still limited when one logical section spans multiple chunks.
|
||||
- `breadcrumb` now participates in embedding text, but it can still be used more strongly in context expansion, rerank, and evidence packing.
|
||||
- Rerank, RRF, BM25, and hybrid retrieval are not implemented yet.
|
||||
- Indexing writes still use SDK.
|
||||
|
||||
These are good follow-up issues because they are retrieval-quality improvements, not blockers for the VectorStore migration.
|
||||
|
||||
## How I Present This In An Interview
|
||||
|
||||
My short version would be:
|
||||
|
||||
> This RAG system started as a self-built MVP around Milvus SDK retrieval. It worked, but too much infrastructure logic lived in business code, and L0/L1 responsibilities were unclear. I refactored it in stages: first I added baseline retrieval cases, then made L0 a domain/entity hint instead of a final decision layer, then added evidence postprocessing, and finally moved the main read path to Spring AI VectorStore with SDK fallback. I kept `lookup_knowledge` as an explicit Agent tool because the project values traceability: the interviewer can see when retrieval happened, what evidence was found, and how it supported the diagnosis. The result is closer to standard Spring AI RAG while still preserving business-specific observability.
|
||||
|
||||
If asked why this is not a full framework migration:
|
||||
|
||||
> I intentionally did not migrate everything at once. Reads moved first because they are easier to compare using golden queries. Writes/indexing stayed on SDK to avoid mixing schema and retrieval behavior changes in one step. Advisors were not used as the main interface because hidden retrieval would weaken the Agent trace.
|
||||
|
||||
If asked what I would improve next:
|
||||
|
||||
> I would add query transformation for AIOps payloads, improve metadata taxonomy, use breadcrumb and section metadata for context expansion, and then evaluate whether hybrid retrieval or rerank is necessary based on measured recall and topK overlap.
|
||||
|
||||
## Interview Follow-Up Questions
|
||||
|
||||
### Why introduce Spring AI VectorStore if the SDK path already worked?
|
||||
|
||||
Because SDK-only retrieval made the project own too much low-level RAG infrastructure. `VectorStore` gives a standard abstraction for retrieval and makes future Spring AI features easier to adopt, while the facade keeps the Agent layer stable.
|
||||
|
||||
### Why keep custom code at all?
|
||||
|
||||
The custom code is where the Agent engineering value lives: AIOps payload mapping, L0 hints, evidence packing, score compatibility, and tool invocation tracing. Those are domain-specific and should remain visible.
|
||||
|
||||
### How do you know quality did not regress?
|
||||
|
||||
I compared SDK and VectorStore modes on representative live queries. Core troubleshooting and AIOps cases returned the same top3 documents in the same order. The differences were isolated to abstract design queries and metadata taxonomy, which are documented follow-up work.
|
||||
|
||||
### What is the most important design decision?
|
||||
|
||||
Keeping a stable boundary: `lookup_knowledge` calls `VectorSearchService`, and `VectorSearchService` decides whether to use Spring AI or SDK. That boundary made the migration small enough to validate and explain.
|
||||
|
||||
@@ -1,71 +1,69 @@
|
||||
# RAG Retrieval Quality Report
|
||||
# RAG 检索质量报告
|
||||
|
||||
## Purpose
|
||||
## 1. 目的
|
||||
|
||||
This report compares the live retrieval behavior of the original Milvus SDK path and the new Spring AI VectorStore path.
|
||||
这份报告回答一个面试关键问题:
|
||||
|
||||
The goal is to answer an interview-critical question:
|
||||
```text
|
||||
迁移到 Spring AI VectorStore 后,怎么证明检索质量没有退化?
|
||||
```
|
||||
|
||||
> After moving retrieval to Spring AI VectorStore, how do we know retrieval quality did not regress?
|
||||
这不是完整 benchmark,而是针对当前 Milvus/Zilliz collection 的代表性 live smoke comparison。
|
||||
|
||||
This is not a full benchmark yet. It is a focused live smoke comparison using representative RAG queries against the current Milvus/Zilliz collection.
|
||||
## 2. 验证设置
|
||||
|
||||
## Setup
|
||||
|
||||
Service endpoint:
|
||||
服务端点:
|
||||
|
||||
```text
|
||||
GET http://127.0.0.1:9900/api/search/similar
|
||||
```
|
||||
|
||||
Collection:
|
||||
collection:
|
||||
|
||||
```text
|
||||
biz
|
||||
```
|
||||
|
||||
Compared modes:
|
||||
对比模式:
|
||||
|
||||
```text
|
||||
retrieval.vector-store.mode=sdk
|
||||
retrieval.vector-store.mode=spring-ai
|
||||
```
|
||||
|
||||
Each case used:
|
||||
每个 case:
|
||||
|
||||
```text
|
||||
topK=3
|
||||
```
|
||||
|
||||
The application was restarted once per mode using command-line configuration so no repository config file had to be changed.
|
||||
## 3. 测试案例
|
||||
|
||||
## Cases
|
||||
| Case | Query | 目的 |
|
||||
|---|---|---|
|
||||
| `err-timeout` | `ERR_TIMEOUT` | 精确错误码检索 |
|
||||
| `payment-service-timeout` | `payment-service timeout` | 服务超时排障 |
|
||||
| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | 数据库排障 |
|
||||
| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps 告警式检索 |
|
||||
| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | 抽象 RAG 设计问题 |
|
||||
| `database-filter` | `mysql timeout`, category=`database` | metadata filter 行为 |
|
||||
|
||||
| Case | Query | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `err-timeout` | `ERR_TIMEOUT` | Exact error-code retrieval |
|
||||
| `payment-service-timeout` | `payment-service timeout` | Service timeout troubleshooting |
|
||||
| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | Database troubleshooting |
|
||||
| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps alert-style retrieval |
|
||||
| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | Abstract RAG design query |
|
||||
| `database-filter` | `mysql timeout`, category=`database` | Metadata filter behavior |
|
||||
## 4. 对比摘要
|
||||
|
||||
## Summary
|
||||
| Case | SDK 数量 | VectorStore 数量 | Top1 一致 | TopK 重叠 | 结论 |
|
||||
|---|---:|---:|---|---:|---|
|
||||
| `err-timeout` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
|
||||
| `payment-service-timeout` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
|
||||
| `mysql-connection-pool` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
|
||||
| `high-cpu-payment` | 3 | 3 | 是 | 3/3 | AIOps 核心 query 一致 |
|
||||
| `rag-l0-l1` | 3 | 1 | 是 | 1/3 | VectorStore 尾部结果更少 |
|
||||
| `database-filter` | 0 | 0 | 不适用 | 不适用 | filter 行为一致,taxonomy 有问题 |
|
||||
|
||||
| Case | SDK Count | VectorStore Count | Top1 Same | TopK Overlap | Notes |
|
||||
| --- | ---: | ---: | --- | ---: | --- |
|
||||
| `err-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
||||
| `payment-service-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
||||
| `mysql-connection-pool` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
||||
| `high-cpu-payment` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
||||
| `rag-l0-l1` | 3 | 1 | Yes | 1/3 | VectorStore returned only the strongest candidate |
|
||||
| `database-filter` | 0 | 0 | N/A | N/A | Both paths applied the filter consistently; no live docs matched `category=database` |
|
||||
## 5. 代表性结果
|
||||
|
||||
## Representative Results
|
||||
### ERR_TIMEOUT
|
||||
|
||||
### `ERR_TIMEOUT`
|
||||
|
||||
SDK:
|
||||
SDK:
|
||||
|
||||
```text
|
||||
1. ERR_TIMEOUT score=0.5659486 label=l2_distance
|
||||
@@ -73,7 +71,7 @@ SDK:
|
||||
3. Error handling score=0.7735061 label=l2_distance
|
||||
```
|
||||
|
||||
VectorStore:
|
||||
VectorStore:
|
||||
|
||||
```text
|
||||
1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity
|
||||
@@ -81,15 +79,15 @@ VectorStore:
|
||||
3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
解释:
|
||||
|
||||
- Document ordering is identical.
|
||||
- Compatibility `score` is identical to SDK L2 distance.
|
||||
- VectorStore `rawScore` exposes Spring AI similarity separately.
|
||||
- 排序一致。
|
||||
- 兼容 `score` 与 SDK L2 distance 一致。
|
||||
- `rawScore` 暴露 Spring AI similarity。
|
||||
|
||||
### `MySQL connection pool`
|
||||
### MySQL connection pool
|
||||
|
||||
Both paths returned:
|
||||
两条路径都返回:
|
||||
|
||||
```text
|
||||
1. MySQL connection pool config
|
||||
@@ -97,58 +95,23 @@ Both paths returned:
|
||||
3. idle-timeout
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
说明迁移保留了核心基础设施排障检索能力。
|
||||
|
||||
- The migration preserves a precise infrastructure troubleshooting retrieval case.
|
||||
- Metadata fields such as title, category, and source remain available.
|
||||
### HighCPUUsage payment-service
|
||||
|
||||
### `HighCPUUsage payment-service`
|
||||
两条路径都返回 payment-service 高 CPU 相关排障文档,说明 AIOps 告警式 query 没有退化。
|
||||
|
||||
Both paths returned:
|
||||
### rag-l0-l1
|
||||
|
||||
```text
|
||||
1. 3. HighCPUUsage / payment-service troubleshooting steps
|
||||
2. evidence mapping table row for HighCPUUsage/payment-service
|
||||
3. 3.1 Symptom confirmation
|
||||
```
|
||||
VectorStore 只返回一个候选,但 Top1 与 SDK 一致。这说明抽象设计类 query 需要后续 query rewrite、补充索引或 threshold 调整。
|
||||
|
||||
Interpretation:
|
||||
### database-filter
|
||||
|
||||
- AIOps-style alert terms still retrieve the expected troubleshooting document.
|
||||
- This is important because AIOps diagnosis depends on knowledge retrieval plus metrics/log evidence.
|
||||
两条路径都返回 0,因为相关 MySQL 文档当前分类是 `infrastructure`,不是 `database`。这是 metadata taxonomy 问题,不是 VectorStore 回归。
|
||||
|
||||
### `rag-l0-l1`
|
||||
## 6. 分数兼容结论
|
||||
|
||||
SDK returned three results, while VectorStore returned one:
|
||||
|
||||
```text
|
||||
Top1: Return error information
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
|
||||
- Top1 did not regress.
|
||||
- VectorStore appears stricter for low-similarity tail results because the Spring AI path uses `similarityThresholdAll()`.
|
||||
- This is acceptable for current read-path migration, but it is worth tracking because abstract design questions may need query rewriting, better indexed docs, or adjusted threshold behavior.
|
||||
|
||||
### `database-filter`
|
||||
|
||||
Both paths returned zero results for:
|
||||
|
||||
```text
|
||||
query=mysql timeout
|
||||
category=database
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
|
||||
- The filter path is consistent.
|
||||
- The live indexed MySQL docs are categorized as `infrastructure`, not `database`.
|
||||
- This highlights a metadata taxonomy issue rather than a VectorStore migration regression.
|
||||
|
||||
## Score Compatibility
|
||||
|
||||
The comparison validates the score design:
|
||||
对比验证了当前分数设计:
|
||||
|
||||
```text
|
||||
SDK:
|
||||
@@ -162,50 +125,24 @@ VectorStore:
|
||||
scoreLabel = similarity
|
||||
```
|
||||
|
||||
This keeps `lookup_knowledge` relevance normalization stable while still exposing the VectorStore score semantics for trace/debugging.
|
||||
这样既保持 `lookup_knowledge` 原有归一化逻辑,又能暴露 VectorStore 语义。
|
||||
|
||||
## Findings
|
||||
## 7. 验收结论
|
||||
|
||||
### Finding 1: Main live cases are equivalent
|
||||
Spring AI VectorStore 读路径可以接受用于当前 MVP/面试:
|
||||
|
||||
For exact error code, service timeout, MySQL troubleshooting, and AIOps alert-style retrieval, SDK and VectorStore returned identical top3 documents in identical order.
|
||||
- 核心排障和 AIOps case 与 SDK top3 一致。
|
||||
- 分数兼容性保留。
|
||||
- VectorStore 语义通过 `rawScore` 和 `scoreLabel` 可观察。
|
||||
- SDK fallback 仍保留运行安全。
|
||||
|
||||
This is strong evidence that the read-path migration did not regress the most important demo and troubleshooting cases.
|
||||
后续检索质量工作不阻塞这次迁移,应作为独立优化继续推进。
|
||||
|
||||
### Finding 2: Abstract RAG design queries need better retrieval support
|
||||
## 8. 下一步
|
||||
|
||||
The `rag-l0-l1` query only returned one VectorStore candidate. The top result matched SDK top1, but the tail differed.
|
||||
- 增加自动 live comparison 脚本。
|
||||
- 在 offline evaluator 中加入 topK overlap、top1 hit、MRR。
|
||||
- 规范 metadata category,例如 `database` 与 `infrastructure`。
|
||||
- 为抽象设计类 query 增加 query rewriting。
|
||||
- 后续再评估是否迁移写入路径到 `VectorStore.add(...)`。
|
||||
|
||||
This suggests the next quality work should focus on:
|
||||
|
||||
- Query transformation for abstract design questions.
|
||||
- Better indexing of interview/devflow RAG design docs.
|
||||
- Context expansion around same-section chunks.
|
||||
- Possibly tuning VectorStore threshold behavior.
|
||||
|
||||
### Finding 3: Metadata taxonomy matters
|
||||
|
||||
The category filter case returned zero results in both modes because the relevant MySQL docs are categorized as `infrastructure`, not `database`.
|
||||
|
||||
This supports a previous RAG issue: category/domain metadata should be normalized before it is used as a hard filter.
|
||||
|
||||
## Acceptance Decision
|
||||
|
||||
The Spring AI VectorStore read path is accepted for current MVP/interview use:
|
||||
|
||||
- Core troubleshooting cases match SDK behavior.
|
||||
- Score compatibility is preserved.
|
||||
- The VectorStore path exposes better score semantics without changing the `lookup_knowledge` API.
|
||||
- SDK fallback remains available for runtime safety.
|
||||
|
||||
The next retrieval-quality improvements should not block this migration. They should be handled as separate RAG quality work.
|
||||
|
||||
## Next Work
|
||||
|
||||
Recommended next steps:
|
||||
|
||||
- Add a small automated live comparison script if repeated validation becomes common.
|
||||
- Add topK overlap and top1 hit metrics to the offline evaluator.
|
||||
- Normalize metadata categories such as `database` vs `infrastructure`.
|
||||
- Add query rewriting for abstract RAG questions.
|
||||
- Decide later whether to migrate indexing writes to Spring AI `VectorStore.add(...)`.
|
||||
|
||||
@@ -1,22 +1,16 @@
|
||||
# RAG VectorStore Interview Notes
|
||||
# RAG VectorStore 面试要点
|
||||
|
||||
## 60-Second Explanation
|
||||
|
||||
I refactored the RAG retrieval path from a direct Milvus SDK-only implementation to a Spring AI `VectorStore` main path, while keeping the SDK path as a fallback.
|
||||
|
||||
The important part is not just the dependency change. I kept `VectorSearchService` as the boundary, so `lookup_knowledge` and the Agent workflow did not need to change. The system now supports three modes:
|
||||
## 1. 60 秒讲法
|
||||
|
||||
```text
|
||||
auto -> try Spring AI VectorStore, fallback to SDK
|
||||
spring-ai -> force VectorStore
|
||||
sdk -> force SDK
|
||||
我把 RAG 检索从 Milvus SDK-only 重构为 Spring AI VectorStore 主路径,同时保留 SDK fallback。
|
||||
关键不是换了一个依赖,而是保留 VectorSearchService 作为边界,所以 lookup_knowledge 和 Agent workflow 不需要改。
|
||||
现在支持 auto、spring-ai、sdk 三种模式。auto 会优先尝试 VectorStore,失败后 fallback 到 SDK。
|
||||
```
|
||||
|
||||
During live verification, the first run found a real config mismatch: VectorStore was pointed at `business_knowledge`, but the real Zilliz collection was `biz`. The fallback worked, so the system still returned results through SDK. After aligning the collection name, the same query went through Spring AI VectorStore successfully.
|
||||
现场验证时,第一次发现 VectorStore 指向了错误 collection:`business_knowledge`,而实际 Zilliz collection 是 `biz`。fallback 生效,所以系统仍能通过 SDK 返回结果。修正 collection 后,同一个 query 成功走 Spring AI VectorStore。
|
||||
|
||||
I also fixed score compatibility. Spring AI Milvus exposes similarity as the document score, but the old `lookup_knowledge` logic expects L2 distance. So I preserve `rawScore` and `scoreLabel`, and use Milvus `metadata.distance` as the compatibility `score` when available.
|
||||
|
||||
## Architecture Answer
|
||||
## 2. 架构回答
|
||||
|
||||
```text
|
||||
Agent / API
|
||||
@@ -27,69 +21,65 @@ Agent / API
|
||||
-> Milvus/Zilliz collection: biz
|
||||
```
|
||||
|
||||
The key design choice is that `VectorSearchService` remains the retrieval facade. This avoids spreading framework-specific code into the Agent tool layer.
|
||||
关键设计:`VectorSearchService` 是检索门面,避免 Spring AI 或 SDK 细节扩散到 Agent 工具层。
|
||||
|
||||
## Why Keep The SDK Path?
|
||||
## 3. 为什么保留 SDK
|
||||
|
||||
I kept SDK fallback for three reasons:
|
||||
- 迁移安全:原 SDK 路径已验证可用。
|
||||
- 运行韧性:VectorStore schema、filter 或配置失败时,检索仍可用。
|
||||
- Demo 稳定:检索抽象变化不应该破坏主诊断演示。
|
||||
|
||||
- Migration safety: the existing SDK path was already proven against the live collection.
|
||||
- Runtime resilience: if VectorStore schema mapping or filtering fails, retrieval still works.
|
||||
- Interview/demo stability: a retrieval abstraction change should not break the main Agent diagnosis demo.
|
||||
这在实际验证中发挥了作用:VectorStore 配置错时,`auto` 模式 fallback 到 SDK,API 没有失败。
|
||||
|
||||
This was validated in practice. When VectorStore pointed at the wrong collection, `auto` mode fell back to SDK and still returned results.
|
||||
## 4. 为什么引入 Spring AI VectorStore
|
||||
|
||||
## Why Use Spring AI VectorStore At All?
|
||||
使用 `VectorStore` 可以让项目更接近标准 RAG 抽象:
|
||||
|
||||
Using Spring AI `VectorStore` moves the project closer to a standard RAG abstraction:
|
||||
- 业务代码不再持有全部 Milvus search 细节。
|
||||
- 后续 QueryTransformer、DocumentPostProcessor、Retriever 等能力更容易接入。
|
||||
- 面试中也更容易解释和 Spring AI 生态的关系。
|
||||
|
||||
- Retrieval code no longer needs to own all Milvus-specific search details.
|
||||
- Later features such as query transformers, document postprocessors, advisors, or retrievers can be introduced more naturally.
|
||||
- The code becomes easier to compare with common Spring AI RAG patterns in an interview.
|
||||
但我没有一次性迁移写入,因为读写同时迁移会让问题难定位。当前先稳定读路径。
|
||||
|
||||
But I did not blindly replace everything. Writes/indexing still use SDK because changing read and write paths at the same time would make failures harder to isolate.
|
||||
## 5. 为什么保留 L0
|
||||
|
||||
## Why Keep L0?
|
||||
L0 现在不是最终答案来源,而是确定性 hint 层:
|
||||
|
||||
L0 is no longer treated as the final source of truth. It is a deterministic hint layer:
|
||||
- 提取 domain/entity。
|
||||
- 在可能时生成 category filter。
|
||||
- 给 trace 提供解释信号。
|
||||
|
||||
- It extracts domain/entity hints from indexed metadata.
|
||||
- It helps constrain L1 retrieval by category when possible.
|
||||
- It gives the Agent a stable clue even when semantic retrieval is weak.
|
||||
|
||||
The current design is:
|
||||
当前职责:
|
||||
|
||||
```text
|
||||
L0 = domain/entity hint
|
||||
L1 = semantic retrieval through VectorStore/SDK
|
||||
postprocess = evidence trace and relevance normalization
|
||||
postprocess = evidence trace + relevance normalization
|
||||
```
|
||||
|
||||
This is easier to defend than saying "we only use vector search." Real incident diagnosis often has exact identifiers, error codes, service names, and alert names. L0 is useful for those.
|
||||
真实故障诊断里有很多精确标识,完全只靠向量检索并不稳。
|
||||
|
||||
## Why Not Use Hidden Spring AI Advisors Directly?
|
||||
## 6. 为什么不用隐藏 Advisor
|
||||
|
||||
For this project, `lookup_knowledge` remains an explicit tool.
|
||||
`lookup_knowledge` 保持显式工具,因为:
|
||||
|
||||
Reason:
|
||||
- Trace 要展示什么时候检索。
|
||||
- `tool_invocation` 要记录输入、输出预览、相关性和 metadata。
|
||||
- 面试故事是可审计 Agent 执行,而不只是答案质量。
|
||||
|
||||
- The Agent trace needs to show when knowledge was retrieved.
|
||||
- `tool_invocation` records input, output preview, relevance level, and evidence metadata.
|
||||
- The interview story is about auditable Agent execution, not only answer quality.
|
||||
Advisor 后续可以接入,但需要先解决可观测性。
|
||||
|
||||
Spring AI Advisors may be useful later, but hiding retrieval inside an advisor would make the evidence chain less visible unless we rebuild trace hooks around it.
|
||||
## 7. 分数设计
|
||||
|
||||
## Score Design
|
||||
|
||||
The result object intentionally separates these fields:
|
||||
当前结果故意拆成:
|
||||
|
||||
```text
|
||||
score -> compatibility score used by old relevance normalization
|
||||
rawScore -> raw score from the retrieval implementation
|
||||
scoreLabel -> semantic label for rawScore
|
||||
score -> 兼容旧 relevance normalization 的分数
|
||||
rawScore -> 当前检索实现原始分数
|
||||
scoreLabel -> rawScore 的语义
|
||||
```
|
||||
|
||||
For SDK:
|
||||
SDK:
|
||||
|
||||
```text
|
||||
score = L2 distance
|
||||
@@ -97,7 +87,7 @@ rawScore = L2 distance
|
||||
scoreLabel = l2_distance
|
||||
```
|
||||
|
||||
For VectorStore:
|
||||
VectorStore:
|
||||
|
||||
```text
|
||||
score = metadata.distance if present
|
||||
@@ -105,63 +95,25 @@ rawScore = Spring AI document score
|
||||
scoreLabel = similarity
|
||||
```
|
||||
|
||||
This prevents a subtle bug: if we treat Spring AI similarity as L2 distance, relevance becomes wrong. If we only expose distance, we lose the ability to compare Spring AI behavior. Keeping both makes the migration inspectable.
|
||||
这样避免把 similarity 当成 L2 distance 的隐蔽 bug。
|
||||
|
||||
## How I Verified It
|
||||
## 8. 如何证明 VectorStore 被使用
|
||||
|
||||
I verified at three levels:
|
||||
- 日志出现 `Starting Spring AI VectorStore search` 和 `Spring AI VectorStore search complete`。
|
||||
- API 响应中 `scoreLabel=similarity`。
|
||||
- `rawScore` 是 Spring AI similarity,`score` 仍是兼容 distance。
|
||||
|
||||
- Unit tests: SDK mode, auto VectorStore mode, fallback mode, category filter, distance metadata mapping.
|
||||
- Live API: `/api/search/similar?query=ERR_TIMEOUT&topK=3`.
|
||||
- Logs: confirmed whether the path was VectorStore success or SDK fallback.
|
||||
## 9. 常见追问
|
||||
|
||||
The live API returned:
|
||||
### 为什么不删 SDK?
|
||||
|
||||
```text
|
||||
scoreLabel = similarity
|
||||
rawScore = Spring AI similarity
|
||||
score = Milvus distance metadata
|
||||
```
|
||||
这是迁移,不是重写。fallback 提供回滚安全,并且已经证明配置错误时仍能保证主链路可用。
|
||||
|
||||
That means the main path was Spring AI VectorStore and compatibility scoring remained stable.
|
||||
### `lookup_knowledge` 变了吗?
|
||||
|
||||
## What I Would Do Next
|
||||
外部契约没变。它仍然调用 `VectorSearchService.searchSimilarDocuments(...)`,变化在门面背后的实现。
|
||||
|
||||
I would not immediately migrate indexing writes. The next responsible steps are:
|
||||
### 这是完整 Spring AI RAG 了吗?
|
||||
|
||||
- Add a small live acceptance report for several golden queries.
|
||||
- Compare `sdk` and `spring-ai` mode side by side for topK overlap.
|
||||
- Decide whether `VectorIndexService` should move to `VectorStore.add(...)`.
|
||||
- Add query transformation or hybrid retrieval only after we have baseline metrics.
|
||||
还不是。当前是 Spring AI VectorStore 读路径 + 显式工具 + 自定义 evidence trace + SDK 写入。这样做是为了保留审计能力和分阶段迁移安全。
|
||||
|
||||
This staged approach is intentional: first stabilize the read path, then evaluate retrieval quality, then migrate writes if the abstraction proves reliable.
|
||||
|
||||
## Interview Questions And Short Answers
|
||||
|
||||
### Why did you not remove the SDK?
|
||||
|
||||
Because this is a migration, not a rewrite. SDK fallback gives rollback safety and proved useful when VectorStore config was initially wrong.
|
||||
|
||||
### What changed for `lookup_knowledge`?
|
||||
|
||||
The public contract did not change. It still calls `VectorSearchService.searchSimilarDocuments(...)`. The implementation behind that facade changed.
|
||||
|
||||
### How do you know VectorStore is actually used?
|
||||
|
||||
The logs show `Starting Spring AI VectorStore search` followed by `Spring AI VectorStore search complete`. The API response also has `scoreLabel=similarity`, which only comes from the VectorStore path.
|
||||
|
||||
### What was the main bug found during live validation?
|
||||
|
||||
The configured collection name was wrong. Spring AI looked for `business_knowledge`, but the actual Milvus collection was `biz`.
|
||||
|
||||
### What did fallback prove?
|
||||
|
||||
It proved that `auto` mode is resilient: VectorStore failed, SDK search still returned valid results, and the API did not fail.
|
||||
|
||||
### Why is `metadata.distance` important?
|
||||
|
||||
Because `lookup_knowledge` uses L2 distance normalization. Spring AI returns similarity as the main document score, but the Milvus distance is available in metadata. Using it preserves old relevance behavior.
|
||||
|
||||
### Is this full Spring AI RAG now?
|
||||
|
||||
Not yet. It uses Spring AI VectorStore for the main read path, but keeps explicit tools, custom evidence trace, L0 hints, and SDK indexing. That is deliberate because the project values auditability and staged migration.
|
||||
|
||||
@@ -1,17 +1,17 @@
|
||||
# RAG VectorStore Live Acceptance
|
||||
# RAG VectorStore Live 验收说明
|
||||
|
||||
## Purpose
|
||||
## 1. 目的
|
||||
|
||||
This note records the live acceptance result for the RAG retrieval refactor.
|
||||
本文记录 RAG 检索重构的 live 验收结论。
|
||||
|
||||
The goal of this refactor was not only to add a Spring AI abstraction, but to prove that the production retrieval path can:
|
||||
这次重构的目标不只是接入 Spring AI 抽象,而是证明线上读路径能够:
|
||||
|
||||
- Prefer Spring AI `VectorStore` for Milvus retrieval.
|
||||
- Preserve the existing Milvus SDK path as fallback.
|
||||
- Keep the `lookup_knowledge` tool contract stable.
|
||||
- Keep L2-distance based relevance normalization compatible.
|
||||
- 优先使用 Spring AI `VectorStore` 做 Milvus 检索。
|
||||
- 保留原 Milvus SDK 作为 fallback。
|
||||
- 保持 `lookup_knowledge` 工具契约稳定。
|
||||
- 保持基于 L2 distance 的相关性归一化兼容。
|
||||
|
||||
## Current Retrieval Shape
|
||||
## 2. 当前检索形态
|
||||
|
||||
```text
|
||||
lookup_knowledge / /api/search/similar
|
||||
@@ -19,22 +19,22 @@ lookup_knowledge / /api/search/similar
|
||||
-> retrieval.vector-store.mode
|
||||
-> auto
|
||||
-> Spring AI VectorStore
|
||||
-> fallback to Milvus SDK if VectorStore fails
|
||||
-> VectorStore 失败时 fallback 到 Milvus SDK
|
||||
-> spring-ai
|
||||
-> Spring AI VectorStore only
|
||||
-> 只走 Spring AI VectorStore
|
||||
-> sdk
|
||||
-> Milvus SDK only
|
||||
-> 只走 Milvus SDK
|
||||
```
|
||||
|
||||
## Configuration Verified
|
||||
## 3. 已验证配置
|
||||
|
||||
The live Milvus/Zilliz database contains the collection:
|
||||
live Milvus/Zilliz 数据库中存在 collection:
|
||||
|
||||
```text
|
||||
biz
|
||||
```
|
||||
|
||||
The Spring AI VectorStore configuration was aligned with the existing SDK collection:
|
||||
Spring AI VectorStore 配置与 SDK 使用的 collection 对齐:
|
||||
|
||||
```yaml
|
||||
spring:
|
||||
@@ -53,11 +53,11 @@ spring:
|
||||
embedding-field-name: vector
|
||||
```
|
||||
|
||||
Why this matters: the earlier config used `business_knowledge`, but the SDK path and real collection use `biz`. That mismatch proved the fallback worked, but it also meant VectorStore was not the successful main path until the config was corrected.
|
||||
为什么重要:早期配置使用 `business_knowledge`,而真实 collection 是 `biz`。这个错配证明了 fallback 生效,但也说明修正前 VectorStore 不是成功主路径。
|
||||
|
||||
## Commands Used
|
||||
## 4. 验收命令
|
||||
|
||||
Health check:
|
||||
健康检查:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
@@ -65,7 +65,7 @@ Invoke-RestMethod `
|
||||
-Method Get
|
||||
```
|
||||
|
||||
Observed result:
|
||||
期望:
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -74,7 +74,7 @@ Observed result:
|
||||
}
|
||||
```
|
||||
|
||||
Direct retrieval check:
|
||||
直接检索:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
@@ -82,7 +82,7 @@ Invoke-RestMethod `
|
||||
-Method Get
|
||||
```
|
||||
|
||||
Observed result shape:
|
||||
期望结果形态:
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -90,7 +90,6 @@ Observed result shape:
|
||||
"message": "success",
|
||||
"data": [
|
||||
{
|
||||
"id": "f7dff7c8-5665-3145-9f75-ef741528b914",
|
||||
"content": "### ERR_TIMEOUT ...",
|
||||
"score": 0.5662,
|
||||
"rawScore": 0.4337,
|
||||
@@ -105,41 +104,40 @@ Observed result shape:
|
||||
}
|
||||
```
|
||||
|
||||
## What The Logs Proved
|
||||
## 5. 日志证明了什么
|
||||
|
||||
Before collection alignment:
|
||||
collection 修正前:
|
||||
|
||||
```text
|
||||
Starting Spring AI VectorStore search
|
||||
SearchRequest collectionName:business_knowledge failed
|
||||
Spring AI VectorStore retrieval failed, falling back to Milvus SDK
|
||||
Starting Milvus SDK search
|
||||
```
|
||||
|
||||
After collection alignment:
|
||||
collection 修正后:
|
||||
|
||||
```text
|
||||
Starting Spring AI VectorStore search: query=ERR_TIMEOUT
|
||||
Spring AI VectorStore search complete, candidates=3
|
||||
```
|
||||
|
||||
This proves:
|
||||
这证明:
|
||||
|
||||
- `auto` mode really attempts VectorStore first.
|
||||
- The fallback is functional when VectorStore fails.
|
||||
- After config alignment, the main path is Spring AI VectorStore rather than SDK fallback.
|
||||
- `auto` 模式确实先尝试 VectorStore。
|
||||
- VectorStore 失败时 fallback 可用。
|
||||
- 配置对齐后,主路径是 Spring AI VectorStore,而不是 SDK fallback。
|
||||
|
||||
## Score Semantics
|
||||
## 6. 分数语义
|
||||
|
||||
The project keeps three score fields intentionally:
|
||||
项目保留三个分数字段:
|
||||
|
||||
```text
|
||||
rawScore -> the raw score from the active retrieval implementation
|
||||
scoreLabel -> the semantic meaning of rawScore
|
||||
score -> compatibility score used by existing lookup relevance normalization
|
||||
rawScore -> 当前检索实现的原始分数
|
||||
scoreLabel -> rawScore 的语义
|
||||
score -> lookup relevance normalization 使用的兼容分数
|
||||
```
|
||||
|
||||
For SDK retrieval:
|
||||
SDK:
|
||||
|
||||
```text
|
||||
rawScore = L2 distance
|
||||
@@ -147,7 +145,7 @@ scoreLabel = l2_distance
|
||||
score = L2 distance
|
||||
```
|
||||
|
||||
For Spring AI VectorStore retrieval:
|
||||
Spring AI VectorStore:
|
||||
|
||||
```text
|
||||
rawScore = Spring AI similarity score
|
||||
@@ -155,45 +153,46 @@ scoreLabel = similarity
|
||||
score = Milvus distance metadata when available
|
||||
```
|
||||
|
||||
Why use `metadata.distance` for `score`: `LookupKnowledgeTool` already normalizes relevance from L2 distance. Spring AI Milvus returns similarity as the document score, but also includes the Milvus distance in metadata. Using distance preserves the old relevance behavior while still exposing the new VectorStore score semantics through `rawScore` and `scoreLabel`.
|
||||
使用 `metadata.distance` 的原因:`LookupKnowledgeTool` 已经基于 L2 distance 做相关性归一化。Spring AI Milvus 主分数是 similarity,但 metadata 中仍有 Milvus distance。用 distance 保持旧逻辑稳定,同时通过 `rawScore` 暴露新语义。
|
||||
|
||||
## Regression Checks
|
||||
## 7. 回归检查
|
||||
|
||||
Targeted tests:
|
||||
目标测试:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=VectorSearchServiceTest,LookupKnowledgeToolTest" test
|
||||
```
|
||||
|
||||
Spec validation:
|
||||
相关 spec:
|
||||
|
||||
```powershell
|
||||
openspec.cmd validate rag-knowledge-retrieval --specs
|
||||
openspec.cmd validate rag-retrieval-evaluation --specs
|
||||
```
|
||||
|
||||
Whitespace check:
|
||||
diff 检查:
|
||||
|
||||
```powershell
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Observed result:
|
||||
验收结论:
|
||||
|
||||
```text
|
||||
All targeted tests passed.
|
||||
All related specs passed.
|
||||
No diff-check errors.
|
||||
目标测试通过。
|
||||
相关 spec 通过。
|
||||
diff-check 无错误。
|
||||
```
|
||||
|
||||
## Acceptance Conclusion
|
||||
## 8. 验收结论
|
||||
|
||||
The VectorStore refactor is accepted for the read path:
|
||||
VectorStore 读路径可以接受:
|
||||
|
||||
- Spring AI VectorStore is integrated and selected in `auto` mode.
|
||||
- The SDK path remains available and was proven by fallback behavior.
|
||||
- The live collection configuration is aligned with the existing Milvus collection.
|
||||
- The `lookup_knowledge` public contract remains stable.
|
||||
- Existing L2-based relevance normalization remains compatible.
|
||||
- Spring AI VectorStore 已集成,并在 `auto` 模式中优先使用。
|
||||
- SDK fallback 保留且已被实际验证。
|
||||
- live collection 配置与现有 Milvus collection 对齐。
|
||||
- `lookup_knowledge` 对外契约保持稳定。
|
||||
- 旧的 L2 relevance normalization 仍兼容。
|
||||
|
||||
写入和索引路径仍使用 Milvus SDK。这是有意的分阶段迁移,不是验收失败项。
|
||||
|
||||
The write/indexing path still uses the Milvus SDK. That is an intentional staged migration decision, not a failed acceptance item.
|
||||
|
||||
@@ -0,0 +1,225 @@
|
||||
# 面试故事案例
|
||||
|
||||
**用途**:把项目能力讲成可被面试官理解的工程故事
|
||||
**使用方式**:按问题选择一个故事,不需要从头到尾背诵
|
||||
|
||||
## 故事 1:从黑盒 Chatbot 到可追踪 Agent
|
||||
|
||||
### 面试官问题
|
||||
|
||||
```text
|
||||
这个项目和普通调用大模型有什么区别?
|
||||
```
|
||||
|
||||
### 30 秒回答
|
||||
|
||||
```text
|
||||
普通 Chatbot 只给最终答案,出了问题很难解释答案怎么来的。
|
||||
我这个项目把诊断过程拆成 Planner、Executor、Verifier,并把每个 Agent 步骤和每次工具调用落库。
|
||||
最后通过 Trace API 可以回放:模型怎么规划、调用了哪些工具、工具返回了什么证据、Verifier 怎么判断答案可信。
|
||||
```
|
||||
|
||||
### 展开讲法
|
||||
|
||||
一开始最容易做的是:用户问题进来,直接让模型回答。但故障诊断场景不能只看答案,因为答案可能看起来合理却没有证据支撑。
|
||||
|
||||
所以我把系统拆成三层:
|
||||
|
||||
- `diagnosis_session` 记录一次诊断的主状态和最终答案。
|
||||
- `agent_step` 记录 Planner、Executor、Verifier 的模型调用。
|
||||
- `tool_invocation` 记录知识库、日志、指标等真实工具证据。
|
||||
|
||||
这样就能做到:答案不是孤立文本,而是一条可审计的执行链。
|
||||
|
||||
### 可展示文件
|
||||
|
||||
- `mvp/architecture/interview-one-pager.md`
|
||||
- `mvp/architecture/session-trace-lifecycle.md`
|
||||
- `mvp/demo/output/trace-response.json`
|
||||
|
||||
### 主动说不足
|
||||
|
||||
```text
|
||||
当前 tool_invocation.step_id 还不是每次都强绑定具体 agent_step,后续可以加 runId 和更严格的 step 关联,让多轮同 session 诊断更清晰。
|
||||
```
|
||||
|
||||
## 故事 2:RAG 从自研 SDK 检索迁移到 Spring AI VectorStore
|
||||
|
||||
### 面试官问题
|
||||
|
||||
```text
|
||||
你的 RAG 是怎么设计的?为什么不用框架全包?
|
||||
```
|
||||
|
||||
### 30 秒回答
|
||||
|
||||
```text
|
||||
我把 RAG 分成两部分:通用检索基础设施尽量交给 Spring AI VectorStore,业务可观测链路留在项目里。
|
||||
所以 Agent 仍然显式调用 lookup_knowledge,底层通过 VectorSearchService 走 Spring AI VectorStore,失败时 fallback 到原 Milvus SDK。
|
||||
这样既能减少自研检索代码,又不会丢失工具调用 trace。
|
||||
```
|
||||
|
||||
### 展开讲法
|
||||
|
||||
旧实现里,Milvus SDK 查询、topK、filter、score 映射都在业务代码里。它能跑,但后续扩展成本高。
|
||||
|
||||
我没有直接把 RAG 隐藏进 Advisor,因为这个项目的核心是 Agent 工程,需要知道 Agent 何时检索、检索了什么、证据怎么支撑诊断。
|
||||
|
||||
于是我保留了边界:
|
||||
|
||||
```text
|
||||
Executor -> lookup_knowledge -> VectorSearchService -> VectorStore / SDK fallback
|
||||
```
|
||||
|
||||
同时把 L0 从“直接返回结果”降级为 domain/entity hint,降低关键词误召回的风险。
|
||||
|
||||
### 可展示文件
|
||||
|
||||
- `mvp/architecture/rag-architecture.md`
|
||||
- `mvp/architecture/retrieval-observability.md`
|
||||
- `interview/rag-refactor-story.md`
|
||||
|
||||
### 主动说不足
|
||||
|
||||
```text
|
||||
当前还没有完整 hybrid search 和 rerank。
|
||||
我先做 golden cases、VectorStore 主路径和 SDK fallback,是为了让每一步迁移都能被验证。
|
||||
```
|
||||
|
||||
## 故事 3:Verifier 如何降低幻觉风险
|
||||
|
||||
### 面试官问题
|
||||
|
||||
```text
|
||||
Agent 怎么保证不胡说?
|
||||
```
|
||||
|
||||
### 30 秒回答
|
||||
|
||||
```text
|
||||
我没有假设模型天然可靠,而是加了 Verifier。
|
||||
Executor 给出答案后,Verifier 只拿 executor_final_answer 和 tool_trace_summary,不允许做新检索。
|
||||
它把答案里的关键事实逐条校验,输出 PASS、LOW_CONFID 或 REJECT。
|
||||
这个结果会写回 self_evaluation,Trace API 可以看到。
|
||||
```
|
||||
|
||||
### 展开讲法
|
||||
|
||||
Verifier 的关键不是再问一次模型“你觉得对吗”,而是让它基于真实工具调用做 groundedness 检查。
|
||||
|
||||
`ToolTraceSummaryService` 会从 `tool_invocation` 里整理证据索引,包含:
|
||||
|
||||
- 工具名。
|
||||
- 输入摘要。
|
||||
- 输出摘要。
|
||||
- evidence level。
|
||||
- source invocation ids。
|
||||
|
||||
Verifier 输出结构化 JSON,ChatService 根据 verdict 决定是否输出、补证据或降级。
|
||||
|
||||
### 可展示文件
|
||||
|
||||
- `mvp/architecture/harness-quality-gates.md`
|
||||
- `mvp/architecture/feedback-architecture.md`
|
||||
- `src/main/resources/prompts/chat-verifier-prompt.md`
|
||||
|
||||
### 主动说不足
|
||||
|
||||
```text
|
||||
AIOps 当前还是轻量 rule evaluation,不是完整 LLM Verifier。
|
||||
这是有意收敛:先用规则保证 payload 聚焦和工具证据使用,后续再加 AIOps LLM Verifier。
|
||||
```
|
||||
|
||||
## 故事 4:AIOps 告警为什么要做 payload scope control
|
||||
|
||||
### 面试官问题
|
||||
|
||||
```text
|
||||
AIOps 场景和普通 Chat 有什么区别?
|
||||
```
|
||||
|
||||
### 30 秒回答
|
||||
|
||||
```text
|
||||
AIOps 告警有一个很关键的问题:环境里可能同时有很多活跃告警,Agent 容易跑偏。
|
||||
所以我把 AIOps 分成 PAYLOAD_TARGETED 和 AUTO_DISCOVERY。
|
||||
如果请求带 alert payload,最终报告必须聚焦输入告警,并且会把 alertName、service、severity、description 拼成 recommended lookup_knowledge query。
|
||||
```
|
||||
|
||||
### 展开讲法
|
||||
|
||||
没有 payload 时,Agent 可以先查询活跃告警,再选择目标排查。
|
||||
|
||||
但有 payload 时,用户已经告诉系统“我要查这个告警”。这时如果 Agent 把其他活跃告警写成主根因,产品体验会很差。
|
||||
|
||||
所以我做了两件事:
|
||||
|
||||
- Prompt 中明确 `PAYLOAD_TARGETED` 范围。
|
||||
- `AiOpsRuleEvaluationService` 检查最终报告是否聚焦输入告警,以及是否使用证据工具。
|
||||
|
||||
### 可展示文件
|
||||
|
||||
- `mvp/architecture/agent-orchestration.md`
|
||||
- `mvp/architecture/current-mvp-architecture.md`
|
||||
- `interview/aiops-query-augmentation.md`
|
||||
- `interview/aiops-lightweight-verifier.md`
|
||||
|
||||
### 主动说不足
|
||||
|
||||
```text
|
||||
当前 scope control 主要靠 prompt 和规则评估。
|
||||
后续可以把 AIOps 也接入类似 Chat Verifier 的事实校验,让告警报告的每个根因和建议都有 evidence refs。
|
||||
```
|
||||
|
||||
## 故事 5:反馈不是点赞按钮,而是案例沉淀入口
|
||||
|
||||
### 面试官问题
|
||||
|
||||
```text
|
||||
用户反馈在系统里有什么用?
|
||||
```
|
||||
|
||||
### 30 秒回答
|
||||
|
||||
```text
|
||||
反馈不只是前端按钮。
|
||||
用户提交 useful 后,系统会把同一个 diagnosis_session 沉淀为 case_library。
|
||||
not_useful 不会改执行状态,而是作为 bad case 信号保留。
|
||||
这样 status、self_evaluation、feedback 三个维度是分开的。
|
||||
```
|
||||
|
||||
### 展开讲法
|
||||
|
||||
我刻意没有把 `not_useful` 写成 `FAILED`。因为失败表示系统执行异常,而用户觉得不好用是质量标签。
|
||||
|
||||
当前设计里:
|
||||
|
||||
```text
|
||||
status -> 执行是否成功
|
||||
self_evaluation -> 系统自己判断证据和事实支撑度
|
||||
feedback -> 用户是否认可
|
||||
```
|
||||
|
||||
`useful` 会进入 `CaseLibraryService.createFromSession`,生成可复用案例。后续可以做相似案例推荐或高质量样本积累。
|
||||
|
||||
### 可展示文件
|
||||
|
||||
- `mvp/architecture/feedback-architecture.md`
|
||||
- `mvp/architecture/data-model.md`
|
||||
- `mvp/demo/output/feedback-response.json`
|
||||
|
||||
### 主动说不足
|
||||
|
||||
```text
|
||||
当前 case_library 的 rootCause 和 solution 还直接使用完整 answer。
|
||||
后续应该从报告中结构化抽取 rootCause、solution、errorCode 和 service,提高案例复用质量。
|
||||
```
|
||||
|
||||
## 结尾万能总结
|
||||
|
||||
```text
|
||||
这个项目我最想展示的不是某一个模型效果,而是 Agent 工程化能力:
|
||||
一个诊断答案从哪里来、用了什么证据、是否被验证、用户是否认可、后续怎么沉淀和回归。
|
||||
这些链路都被结构化记录下来,所以它可以继续演进,而不是一次性 demo。
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user