docs: reorganize MVP interview documentation

This commit is contained in:
aruo
2026-07-05 15:29:28 +08:00
parent b22f2d22c8
commit 88e0a6c944
51 changed files with 4352 additions and 1318 deletions
+61 -57
View File
@@ -1,41 +1,42 @@
# MVP Demo Runbook
# MVP 演示手册
This demo proves the MVP flow from user question to persisted diagnosis trace.
本目录用于演示 MVP 从用户问题到诊断 Trace 的完整闭环。
For interview use, start with:
面试时建议先读:
- `interview-walkthrough.md` for the talk track
- `trace-inspection-checklist.md` for fields to inspect
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
- `requests/payment-timeout-chat.json` for the fixed request payload
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
- `interview-walkthrough.md`:面试讲解话术。
- `trace-inspection-checklist.md`:Trace 字段检查清单。
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
## Prerequisites
## 1. 前置条件
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
- Security and secret cleanup are intentionally out of scope for this MVP slice.
- The `mvp-demo` profile enables mock Prometheus and CLS providers so log and metric tools can return repeatable evidence.
- MySQL、Redis、Milvus/Zilliz、LLM 和 embedding 配置可用。
- 安全和密钥清理不属于当前 MVP 演示范围。
- `mvp-demo` profile 会启用 mock Prometheus 和 mock CLS,让日志和指标工具返回可复现证据。
## Start
## 2. 启动服务
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
The service listens on:
服务地址:
```text
http://localhost:9900
```
## 1. Run Chat Diagnosis
## 3. Chat 诊断 Demo
Fast path:
最快方式:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
This writes:
脚本会生成:
```text
mvp/demo/output/chat-response.json
@@ -43,7 +44,7 @@ mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
Manual path:
手动请求:
```powershell
$sessionId = "mvp-demo-payment-timeout-001"
@@ -59,13 +60,13 @@ Invoke-RestMethod `
-Body $body
```
Expected result:
期望结果:
- `data.success` is `true`.
- `data.sessionId` equals `mvp-demo-payment-timeout-001`.
- `data.answer` contains a diagnosis answer.
- `data.success = true`
- `data.sessionId = mvp-demo-payment-timeout-001`
- `data.answer` 包含诊断答复
## 2. Query Trace
## 4. 查询 Trace
```powershell
Invoke-RestMethod `
@@ -73,15 +74,15 @@ Invoke-RestMethod `
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
```
Expected result:
期望结果:
- `code` is `200`.
- `data.session.sessionId` equals the chat session id.
- `data.steps` contains planner/executor/verifier records for complex questions.
- `data.toolInvocations` contains evidence tool calls such as `lookup_knowledge`, `query_logs`, or `query_metrics`.
- `data.session.selfEvaluation` contains verifier or rule evaluation when available.
- `code = 200`
- `data.session.sessionId` 等于 Chat session id
- `data.steps` 包含 planner / executor / verifier 等步骤
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
## 3. Submit Feedback
## 5. 提交反馈
```powershell
$feedback = @{
@@ -96,12 +97,13 @@ Invoke-RestMethod `
-Body $feedback
```
Expected result:
期望结果:
- `success` is `true`.
- A later trace query shows `data.session.feedback` as `useful`.
- `success = true`
- 后续 Trace 中 `data.session.feedback = useful`
- useful 反馈会尝试沉淀 `case_library`
## 4. Run AIOps Alert Diagnosis
## 6. AIOps 告警诊断 Demo
```powershell
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
@@ -122,15 +124,16 @@ Invoke-WebRequest `
-Body $aiopsBody
```
Expected result:
期望结果:
- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`.
- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload.
- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`.
- `data.session.answer` contains the final alert analysis report when a report is generated.
- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them.
- SSE 首条包含 `session` 消息,sessionId 为 `mvp-demo-aiops-payment-cpu-001`
- 后续流式输出包含 AIOps 告警分析报告
- 报告聚焦输入的 `HighCPUUsage/payment-service`
- 同一 session 的 Trace 中 `data.session.agentFlow = AI_OPS`
- `data.session.answer` 包含最终告警报告
- `data.toolInvocations` 包含证据工具调用
Query the AIOps trace:
查询 AIOps Trace:
```powershell
Invoke-RestMethod `
@@ -138,28 +141,29 @@ Invoke-RestMethod `
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
```
## Demo Story
## 7. Demo 主线
The important interview story is:
Chat 主线:
```text
one session id
-> user question
-> multi-agent execution
-> evidence tools
-> verifier/self-evaluation
-> final answer
-> feedback
-> trace API for replay and audit
一个 session id
-> 用户问题
-> 多 Agent 执行
-> 证据工具
-> Verifier / self_evaluation
-> 最终答案
-> 用户反馈
-> Trace API 回放
```
The AIOps story uses the same audit spine:
AIOps 主线:
```text
one session id
-> alert payload
-> AIOps planner/executor execution
-> evidence tools
-> alert analysis report
-> trace API for replay and audit
一个 session id
-> 告警 payload
-> AIOps Planner / Executor
-> 证据工具
-> 告警分析报告
-> AIOps rule evaluation
-> Trace API 回放
```
+22 -18
View File
@@ -1,15 +1,15 @@
# AIOps Alert Acceptance Case
# AIOps 告警验收用例
## Goal
## 1. 目标
Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry.
验证旧版 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
## Input
## 2. 输入
- Session id: `mvp-demo-aiops-payment-cpu-001`
- Endpoint: `POST /api/ai_ops`
- Profile: `mvp-demo`
- Alert:
- Session id:`mvp-demo-aiops-payment-cpu-001`
- Endpoint:`POST /api/ai_ops`
- Profile:`mvp-demo`
- 告警 payload:
```json
{
@@ -23,16 +23,20 @@ Validate that the legacy AIOps endpoint can act as a traceable alert-triggered d
}
```
## Acceptance Criteria
## 3. 验收标准
1. The SSE stream emits a `session` message containing the requested session id.
2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`.
3. The persisted session query contains the alert name, service, severity, time range, and description.
4. If a final report is generated, `diagnosis_session.answer` contains that report.
5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations.
6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections.
1. SSE 流输出 `session` 消息,且包含请求中的 session id。
2. AIOps 执行创建或更新 `diagnosis_session`,并写入 `agent_flow = AI_OPS`。
3. 持久化的 session query 包含告警名、服务名、等级、时间范围和描述。
4. 如果生成最终报告,`diagnosis_session.answer` 包含该报告。
5. `GET /api/diagnosis/{sessionId}/trace` 返回 AIOps session、按顺序排列的 agent steps 和 tool invocations。
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
8. `self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
## Known Limits
## 4. 已知边界
- 当前 AIOps 使用轻量规则评估器,不是完整 LLM Verifier。
- 完整运行仍依赖有效的 DB、Redis、Milvus/Zilliz、模型和 embedding 配置。
- `mvp-demo` profile 使用 mock Prometheus 和 mock CLS,主要用于稳定演示。
- This slice does not add a Verifier Agent to AIOps.
- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration.
+57 -59
View File
@@ -1,146 +1,144 @@
# Interview Walkthrough: MVP Diagnosis Agent
# 面试演示讲解稿
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
这是一份短时间 Agent 工程面试用讲解稿,不是完整系统文档。
## 30-Second Summary
## 1. 30 秒摘要
```text
This is an enterprise diagnosis Agent MVP.
It takes a payment-timeout question, plans the investigation, calls evidence tools,
checks the answer through a verifier, persists the full trace, and accepts feedback.
这是一个企业故障诊断 Agent MVP。
它接收支付超时问题,规划排查步骤,调用证据工具,
用 Verifier 检查答案,把完整 Trace 持久化,并支持用户反馈。
```
The important claim is not "the model answered once." The claim is:
关键主张不是“模型回答了一次”,而是:
```text
The system can show what evidence was used, how the answer was checked, and how to replay the session.
系统能展示用了什么证据、答案如何被检查、如何用 sessionId 回放整次诊断。
```
## Demo Flow
## 2. Demo 流程
1. Start the service with the `mvp-demo` profile.
2. Run the fixed payment-timeout request.
3. Open `mvp/demo/output/chat-response.json`.
4. Open `mvp/demo/output/trace-response.json`.
5. Point to evidence tools and verifier evaluation.
6. Submit feedback and show it is attached to the same session.
1. 用 `mvp-demo` profile 启动服务。
2. 运行固定的支付超时请求。
3. 打开 `mvp/demo/output/chat-response.json`。
4. 打开 `mvp/demo/output/trace-response.json`。
5. 指出证据工具和 verifier evaluation。
6. 提交 feedback,并展示它挂在同一个 session 上。
## Commands
## 3. 命令
Start service:
启动服务:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
Run the demo from another terminal:
另开终端运行 Demo:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
Optional custom session:
可选自定义 session:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
```
## What To Show
## 4. 展示什么
### 1. User-Facing Answer
### 4.1 用户侧答案
File:
文件:
```text
mvp/demo/output/chat-response.json
```
Say:
话术:
```text
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
这是用户看到的答案。这里的 sessionId 是稳定的,所以我后面可以追踪这一次回答是怎么来的。
```
### 2. Evidence Trace
### 4.2 证据 Trace
File:
文件:
```text
mvp/demo/output/trace-response.json
```
Say:
话术:
```text
This is the important Agent engineering part.
I can inspect which tools were called, what inputs they received,
whether they succeeded, and what evidence preview was persisted.
这才是 Agent 工程最重要的部分。
我可以检查 Agent 调用了哪些工具、每个工具拿到什么入参、是否成功、返回了什么证据预览。
```
Point to:
重点字段:
- `data.toolInvocations[*].toolName`
- `data.toolInvocations[*].inputParams`
- `data.toolInvocations[*].outputPreview`
- `data.toolInvocations[*].success`
### 3. Verifier / Self-Evaluation
### 4.3 Verifier / 自评估
Point to:
重点字段:
- `data.session.selfEvaluation`
- `data.summary.hasVerifierEvaluation`
Say:
话术:
```text
The final answer is not just raw Executor output.
It is checked by a verifier or self-evaluation layer using the persisted trace.
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
最终答案不是 Executor 原始输出直接返回。
系统会基于持久化的工具 trace 做 Verifier 或规则自评估。
这样系统可以区分 PASS、LOW_CONFID、REJECT,而不是假装每个答案都同样可信。
```
### 4. Feedback Loop
### 4.4 反馈闭环
File:
文件:
```text
mvp/demo/output/feedback-response.json
```
Then re-query trace if needed.
必要时重新查询 Trace。
Say:
话术:
```text
Feedback is attached to the same diagnosis session.
That makes it possible to mine useful / not useful cases later.
feedback 会挂在同一个 diagnosis session 上。
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
```
### 5. Regression Story
### 4.5 回归故事
Mention, do not deep dive unless asked:
如果被问到稳定性,可以补充:
```text
For repeatability, I also built an offline eval baseline.
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
The two are separate on purpose: demo for human review, eval for automated signal.
我把运行时 Demo 和离线 eval 分开。
Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可以回归。
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
```
## Strong Interview Framing
Use this phrasing:
## 5. 强面试表达
```text
I focused on the Agent engineering surface:
traceability, evidence persistence, verifier gating, feedback, and regression checks.
The model answer is only one part of the system.
The more important part is whether we can audit and improve the answer after it is produced.
我关注的是 Agent 工程表面:
traceability、evidence persistence、verifier gating、feedback 和 regression checks。
模型答案只是系统的一部分。
更重要的是答案产出后,能否被审计、验证和持续改进。
```
## Known Limits To Say Proactively
## 6. 主动说明限制
```text
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
这个 MVP 仍依赖 MySQL、Redis、Milvus 和模型凭证。
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
```
+5 -4
View File
@@ -1,11 +1,12 @@
# Demo Output
# Demo 输出目录
This directory is the default output location for local demo responses.
本目录是本地 Demo 响应的默认输出位置。
Generated files are intentionally ignored by Git:
生成文件会被 Git 忽略:
- `chat-response.json`
- `trace-response.json`
- `feedback-response.json`
Keep this README so the directory exists in the repository.
保留此 README 是为了让目录存在于仓库中。
+19 -18
View File
@@ -1,24 +1,24 @@
# Payment Timeout Acceptance Case
# 支付超时诊断验收用例
## Goal
## 1. 目标
Validate that the MVP can diagnose a payment timeout incident and expose the complete trace for replay.
验证 MVP 能诊断支付超时问题,并暴露完整 Trace 供回放。
## Input
## 2. 输入
- Session id: `mvp-demo-payment-timeout-001`
- Question: `支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
- Profile: `mvp-demo`
- Session id:`mvp-demo-payment-timeout-001`
- 问题:`支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
- Profile:`mvp-demo`
## Acceptance Criteria
## 3. 验收标准
1. Chat returns a successful answer with the same session id.
2. Trace API returns session metadata, final answer, ordered agent steps, and ordered tool invocations.
3. Trace contains enough evidence to explain which tools were used and whether verifier/self-evaluation was persisted.
4. Feedback can be submitted for the same session id.
5. A follow-up trace query shows the persisted feedback value.
1. Chat 返回成功答复,且 session id 与请求一致。
2. Trace API 返回 session 元数据、最终答案、按顺序排列的 agent steps 和 tool invocations。
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
4. 可以使用同一个 session id 提交反馈。
5. 后续 Trace 查询能看到已持久化的 feedback 值。
## Trace Fields To Inspect
## 4. 需要检查的 Trace 字段
- `data.session.query`
- `data.session.answer`
@@ -32,8 +32,9 @@ Validate that the MVP can diagnose a payment timeout incident and expose the com
- `data.toolInvocations[*].retrievalDetails`
- `data.summary`
## Known Limits
## 5. 已知边界
- 这不是完整离线测试,仍需要有效的 chat、持久化、向量检索和模型调用环境。
- `mvp-demo` profile 启用 mock 日志和指标,让证据工具返回更稳定。
- 敏感配置清理不属于当前 MVP 优先级。
- This case is not a full offline test. It still requires valid infrastructure for chat, persistence, vector search, and model calls.
- Mock logs and metrics are enabled by the `mvp-demo` profile to make those evidence tools repeatable.
- Sensitive configuration cleanup is deferred by current MVP priority.
@@ -13,7 +13,7 @@ $request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
$request.Id = $SessionId
$body = $request | ConvertTo-Json -Depth 8
Write-Host "Running payment-timeout chat demo..."
Write-Host "正在运行支付超时 Chat 诊断 Demo..."
Write-Host "BaseUrl: $BaseUrl"
Write-Host "SessionId: $SessionId"
@@ -24,14 +24,14 @@ $chat = Invoke-RestMethod `
-Body $body
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
Write-Host "Saved chat response: $OutputDir/chat-response.json"
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
$trace = Invoke-RestMethod `
-Method Get `
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
Write-Host "Saved trace response: $OutputDir/trace-response.json"
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
$feedbackBody = @{
sessionId = $SessionId
@@ -45,10 +45,10 @@ $feedback = Invoke-RestMethod `
-Body $feedbackBody
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
Write-Host "已保存反馈响应: $OutputDir/feedback-response.json"
Write-Host ""
Write-Host "Demo completed. Review:"
Write-Host "Demo 已完成,请检查:"
Write-Host "- mvp/demo/output/chat-response.json"
Write-Host "- mvp/demo/output/trace-response.json"
Write-Host "- mvp/demo/output/feedback-response.json"
+237
View File
@@ -0,0 +1,237 @@
# 10 分钟面试演示脚本
**用途**:面试现场按步骤演示
**目标**:展示从问题到证据、验证、Trace、反馈的闭环
**前置条件**:服务以 `mvp-demo` profile 启动
更完整的 runbook 见 [README.md](README.md),字段检查见 [trace-inspection-checklist.md](trace-inspection-checklist.md)。
## 0. 开场话术
```text
我会演示一个支付超时诊断。
重点不是看模型给出一段答案,而是看这个答案背后的 Agent 执行链路:
Planner 怎么拆解,Executor 调了哪些工具,Verifier 如何判断证据是否支撑答案,以及最终如何通过 sessionId 回放。
```
## 1. 启动服务
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
服务地址:
```text
http://localhost:9900
```
说明:
- `mvp-demo` profile 使用 mock Prometheus 和 mock CLS。
- 演示不依赖真实线上故障。
- MySQL、Redis、Milvus/Zilliz 和模型配置仍需要可用。
## 2. 演示 Chat 诊断
推荐使用固定脚本:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
脚本会写出:
```text
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
现场话术:
```text
这里我用固定 sessionId 跑一个支付接口超时问题。
固定 sessionId 的好处是,后面 trace 和 feedback 都能关联到同一次诊断。
```
## 3. 展示用户答案
打开:
```text
mvp/demo/output/chat-response.json
```
重点看:
```text
data.sessionId
data.answer
```
现场话术:
```text
这是用户看到的答案。
但这个项目的重点不是这段文字,而是这段文字是否有证据链。
接下来我用同一个 sessionId 查 trace。
```
## 4. 展示 Trace
打开:
```text
mvp/demo/output/trace-response.json
```
重点看:
```text
data.session.sessionId
data.session.agentFlow
data.steps[*].agentName
data.toolInvocations[*].toolName
data.toolInvocations[*].inputParams
data.toolInvocations[*].outputPreview
data.toolInvocations[*].retrievalLayer
data.toolInvocations[*].relevanceLevel
data.summary.hasVerifierEvaluation
```
现场话术:
```text
这里能看到三个层次:
第一,session 记录了这次诊断的问题、答案、耗时和自评估。
第二,agent_step 记录 Planner、Executor、Verifier 的模型步骤。
第三,tool_invocation 记录真实工具调用,包括 lookup_knowledge、日志和指标。
所以这不是一个黑盒 Chatbot,而是一条可以回放的诊断链路。
```
## 5. 展示知识库检索
在 trace 中找到 `lookup_knowledge`。
重点看:
```text
toolName = lookup_knowledge
inputParams.query
retrievalLayer
l0MatchCount
l1MatchCount
relevanceLevel
retrievalDetails
outputPreview
```
现场话术:
```text
知识库检索保留为显式工具,而不是藏在 Advisor 里。
这样面试官或线上排查人员能看到:Agent 查了什么 query,命中了哪个知识域,检索层是 L0/L1 还是混合,相关性等级是什么。
底层检索现在走 VectorSearchService,优先 Spring AI VectorStore,失败时 fallback 到 Milvus SDK。
```
## 6. 展示 Verifier
在 trace 中查看:
```text
data.session.selfEvaluation
data.summary.hasVerifierEvaluation
```
现场话术:
```text
Verifier 不做新检索,只看工具 trace 汇总。
它会把 Executor 答案里的关键事实拆出来,判断每条事实是 direct_evidence、indirect_support、no_evidence 还是 contradicted。
如果 PASS,就输出原答案。
如果 LOW_CONFID,可以补证据或加低置信提示。
如果 REJECT,就降级输出,只保留已确认信息。
```
## 7. 展示反馈闭环
打开:
```text
mvp/demo/output/feedback-response.json
```
重点看:
```text
success
caseId
```
现场话术:
```text
用户反馈 useful 会写回同一个 diagnosis_session。
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
这里 status 和 feedback 是分开的:
status 表示执行是否成功,feedback 表示用户是否认可。
```
## 8. 可选演示 AIOps
如果时间允许,再演示 AIOps payload。
请求示例见:
```text
mvp/demo/README.md
```
现场话术:
```text
AIOps 有两个模式。
有 payload 时进入 PAYLOAD_TARGETED,报告必须聚焦这个告警。
没有 payload 时进入 AUTO_DISCOVERY,先发现活跃告警再排查。
我专门加了 recommended lookup_knowledge query,把 alertName、service、severity、description 等字段稳定送入知识库检索,避免 Agent 随意扩展问题范围。
```
## 9. 结束总结
```text
这个 Demo 展示的是一个完整闭环:
用户问题
-> Agent 规划和执行
-> 显式工具证据
-> Verifier / self_evaluation
-> Trace 回放
-> 用户反馈
-> 案例沉淀
我把重点放在 Agent 工程能力:可追踪、可验证、可回归、可演进。
```
## 10. 如果现场失败
如果模型或外部组件不可用,不要硬跑。可以直接打开上一次输出:
```text
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
降级话术:
```text
现场环境依赖 MySQL、Redis、Milvus 和模型服务。
如果外部服务不可用,我会用固定输出讲 trace 结构。
因为这个项目的核心不是一次在线请求,而是诊断链路如何被记录、检查和回放。
```
+39 -38
View File
@@ -1,52 +1,53 @@
# Trace Inspection Checklist
# Trace 检查清单
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。
## Session
## 1. Session
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | 一个 session id 串起 chat、工具、verifier、feedback 和 trace |
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
## Agent Steps
## 2. Agent 步骤
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.steps[*].agentName` | 是否有 Planner / Executor / Verifier 或等价步骤 | 流程被拆成可检查的 Agent 步骤 |
| `data.steps[*].thought` | 是否有高层步骤摘要 | 内部过程可审计,不只看最终文本 |
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
## Tool Evidence
## 3. 工具证据
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.toolInvocations[*].toolName` | 是否包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 | Agent 通过工具收集证据,而不是无依据猜测 |
| `data.toolInvocations[*].inputParams` | 是否能看到每个工具的入参 | 工具输入可审计、可调试 |
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
## Summary
## 4. Summary
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.summary.persistedStepCount` | step 行是否持久化 | Trace 来自存储,不是响应内存 |
| `data.summary.persistedToolCallCount` | tool 行是否持久化 | 工具证据在请求结束后仍可回放 |
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
## What Good Looks Like
## 5. 好的结果长什么样
```text
same session id
-> final answer
-> persisted agent steps
-> persisted evidence tool calls
-> verifier/self-evaluation
同一个 session id
-> 最终答案
-> 持久化 agent steps
-> 持久化 evidence tool calls
-> verifier / self-evaluation
-> feedback attached to the same session
```