docs: reorganize MVP interview documentation

This commit is contained in:
aruo
2026-07-05 15:29:28 +08:00
parent b22f2d22c8
commit 88e0a6c944
51 changed files with 4352 additions and 1318 deletions
+57 -59
View File
@@ -1,146 +1,144 @@
# Interview Walkthrough: MVP Diagnosis Agent
# 面试演示讲解稿
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
这是一份短时间 Agent 工程面试用讲解稿,不是完整系统文档。
## 30-Second Summary
## 1. 30 秒摘要
```text
This is an enterprise diagnosis Agent MVP.
It takes a payment-timeout question, plans the investigation, calls evidence tools,
checks the answer through a verifier, persists the full trace, and accepts feedback.
这是一个企业故障诊断 Agent MVP。
它接收支付超时问题,规划排查步骤,调用证据工具,
用 Verifier 检查答案,把完整 Trace 持久化,并支持用户反馈。
```
The important claim is not "the model answered once." The claim is:
关键主张不是“模型回答了一次”,而是:
```text
The system can show what evidence was used, how the answer was checked, and how to replay the session.
系统能展示用了什么证据、答案如何被检查、如何用 sessionId 回放整次诊断。
```
## Demo Flow
## 2. Demo 流程
1. Start the service with the `mvp-demo` profile.
2. Run the fixed payment-timeout request.
3. Open `mvp/demo/output/chat-response.json`.
4. Open `mvp/demo/output/trace-response.json`.
5. Point to evidence tools and verifier evaluation.
6. Submit feedback and show it is attached to the same session.
1. 用 `mvp-demo` profile 启动服务。
2. 运行固定的支付超时请求。
3. 打开 `mvp/demo/output/chat-response.json`。
4. 打开 `mvp/demo/output/trace-response.json`。
5. 指出证据工具和 verifier evaluation。
6. 提交 feedback,并展示它挂在同一个 session 上。
## Commands
## 3. 命令
Start service:
启动服务:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
Run the demo from another terminal:
另开终端运行 Demo:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
Optional custom session:
可选自定义 session:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
```
## What To Show
## 4. 展示什么
### 1. User-Facing Answer
### 4.1 用户侧答案
File:
文件:
```text
mvp/demo/output/chat-response.json
```
Say:
话术:
```text
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
这是用户看到的答案。这里的 sessionId 是稳定的,所以我后面可以追踪这一次回答是怎么来的。
```
### 2. Evidence Trace
### 4.2 证据 Trace
File:
文件:
```text
mvp/demo/output/trace-response.json
```
Say:
话术:
```text
This is the important Agent engineering part.
I can inspect which tools were called, what inputs they received,
whether they succeeded, and what evidence preview was persisted.
这才是 Agent 工程最重要的部分。
我可以检查 Agent 调用了哪些工具、每个工具拿到什么入参、是否成功、返回了什么证据预览。
```
Point to:
重点字段:
- `data.toolInvocations[*].toolName`
- `data.toolInvocations[*].inputParams`
- `data.toolInvocations[*].outputPreview`
- `data.toolInvocations[*].success`
### 3. Verifier / Self-Evaluation
### 4.3 Verifier / 自评估
Point to:
重点字段:
- `data.session.selfEvaluation`
- `data.summary.hasVerifierEvaluation`
Say:
话术:
```text
The final answer is not just raw Executor output.
It is checked by a verifier or self-evaluation layer using the persisted trace.
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
最终答案不是 Executor 原始输出直接返回。
系统会基于持久化的工具 trace 做 Verifier 或规则自评估。
这样系统可以区分 PASS、LOW_CONFID、REJECT,而不是假装每个答案都同样可信。
```
### 4. Feedback Loop
### 4.4 反馈闭环
File:
文件:
```text
mvp/demo/output/feedback-response.json
```
Then re-query trace if needed.
必要时重新查询 Trace。
Say:
话术:
```text
Feedback is attached to the same diagnosis session.
That makes it possible to mine useful / not useful cases later.
feedback 会挂在同一个 diagnosis session 上。
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
```
### 5. Regression Story
### 4.5 回归故事
Mention, do not deep dive unless asked:
如果被问到稳定性,可以补充:
```text
For repeatability, I also built an offline eval baseline.
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
The two are separate on purpose: demo for human review, eval for automated signal.
我把运行时 Demo 和离线 eval 分开。
Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可以回归。
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
```
## Strong Interview Framing
Use this phrasing:
## 5. 强面试表达
```text
I focused on the Agent engineering surface:
traceability, evidence persistence, verifier gating, feedback, and regression checks.
The model answer is only one part of the system.
The more important part is whether we can audit and improve the answer after it is produced.
我关注的是 Agent 工程表面:
traceability、evidence persistence、verifier gating、feedback 和 regression checks。
模型答案只是系统的一部分。
更重要的是答案产出后,能否被审计、验证和持续改进。
```
## Known Limits To Say Proactively
## 6. 主动说明限制
```text
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
这个 MVP 仍依赖 MySQL、Redis、Milvus 和模型凭证。
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
```