refactor(harness): remove legacy agent architecture
This commit is contained in:
@@ -0,0 +1,205 @@
|
||||
# MVP 演示手册
|
||||
|
||||
本目录用于演示 MVP 从用户问题到诊断 Trace 的完整闭环。
|
||||
|
||||
面试时建议先读:
|
||||
|
||||
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||
- `interview-walkthrough.md`:面试讲解话术。
|
||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
|
||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
|
||||
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
|
||||
- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
|
||||
|
||||
## 1. 前置条件
|
||||
|
||||
- MySQL、Redis、Milvus/Zilliz、LLM 和 embedding 配置可用。
|
||||
- 安全和密钥清理不属于当前 MVP 演示范围。
|
||||
- `mvp-demo` profile 会启用 mock Prometheus 和 mock CLS,让日志和指标工具返回可复现证据。
|
||||
|
||||
## 2. 启动服务
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
服务地址:
|
||||
|
||||
```text
|
||||
http://localhost:9900
|
||||
```
|
||||
|
||||
## 3. Chat 诊断 Demo
|
||||
|
||||
最快方式:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会生成:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
手动请求:
|
||||
|
||||
```powershell
|
||||
$sessionId = "mvp-demo-payment-timeout-001"
|
||||
$body = @{
|
||||
Id = $sessionId
|
||||
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
|
||||
|
||||
```powershell
|
||||
$chat = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
|
||||
$runId = $chat.data.runId
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `data.success = true`
|
||||
- `data.sessionId = mvp-demo-payment-timeout-001`
|
||||
- `data.runId` 为本次诊断运行的唯一 ID
|
||||
- `data.answer` 包含诊断答复
|
||||
|
||||
## 4. 查询 Trace
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `code = 200`
|
||||
- `data.runId` 等于 `$runId`
|
||||
- `data.session.sessionId` 等于 Chat session id
|
||||
- `data.run.runId` 等于 `$runId`
|
||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
|
||||
|
||||
## 5. 提交反馈
|
||||
|
||||
```powershell
|
||||
$feedback = @{
|
||||
sessionId = $sessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/feedback" `
|
||||
-ContentType "application/json" `
|
||||
-Body $feedback
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `success = true`
|
||||
- `runId = $runId`
|
||||
- 后续精确 Trace 中 `data.session.feedback = useful`
|
||||
- useful 反馈会尝试沉淀 `case_library`
|
||||
|
||||
## 6. AIOps 告警诊断 Demo
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
|
||||
- 后续流式输出包含 AIOps 告警分析报告
|
||||
- 报告聚焦输入的 `HighCPUUsage/payment-service`
|
||||
- 精确 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||
- `data.session.answer` 包含最终告警报告
|
||||
- `data.toolInvocations` 包含证据工具调用
|
||||
|
||||
查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
|
||||
```
|
||||
|
||||
## 7. Demo 主线
|
||||
|
||||
Chat 主线:
|
||||
|
||||
```text
|
||||
一个 session id + 一个 run id
|
||||
-> 用户问题
|
||||
-> 多 Agent 执行
|
||||
-> 证据工具
|
||||
-> Verifier / self_evaluation
|
||||
-> 最终答案
|
||||
-> 用户反馈
|
||||
-> Trace API 回放
|
||||
```
|
||||
|
||||
AIOps 主线:
|
||||
|
||||
```text
|
||||
一个 session id + 一个 run id
|
||||
-> 告警 payload
|
||||
-> AIOps Planner / Executor
|
||||
-> 证据工具
|
||||
-> 告警分析报告
|
||||
-> AIOps rule evaluation
|
||||
-> Trace API 回放
|
||||
```
|
||||
|
||||
## 8. Evidence Pipeline 场景矩阵
|
||||
|
||||
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
|
||||
|
||||
- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
|
||||
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
|
||||
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
|
||||
|
||||
这样可以同时展示真实链路和确定性回归能力。
|
||||
@@ -0,0 +1,3 @@
|
||||
# Archive Note
|
||||
|
||||
本目录保存旧多角色与旧诊断入口 demo。当前 demo 以 `mvp/demo/README.md` 和唯一 `/api/chat` named SSE 为准。
|
||||
@@ -0,0 +1,42 @@
|
||||
# AIOps 告警验收用例
|
||||
|
||||
## 1. 目标
|
||||
|
||||
验证 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
||||
|
||||
## 2. 输入
|
||||
|
||||
- Session id:`mvp-demo-aiops-payment-cpu-001`
|
||||
- Endpoint:`POST /api/ai_ops`
|
||||
- Profile:`mvp-demo`
|
||||
- 告警 payload:
|
||||
|
||||
```json
|
||||
{
|
||||
"sessionId": "mvp-demo-aiops-payment-cpu-001",
|
||||
"alertName": "HighCPUUsage",
|
||||
"service": "payment-service",
|
||||
"severity": "P1",
|
||||
"description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。",
|
||||
"timeRange": "last_15m",
|
||||
"userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
}
|
||||
```
|
||||
|
||||
## 3. 验收标准
|
||||
|
||||
1. SSE 流首条输出 `type=metadata` 的 `message` 事件,且包含请求中的 session id 和本次 AIOps run id。
|
||||
2. AIOps 执行创建 `diagnosis_run`,并写入 `agent_flow = AI_OPS`。
|
||||
3. 持久化的 run query 包含告警名、服务名、等级、时间范围和描述。
|
||||
4. 如果生成最终报告,`diagnosis_run.answer` 包含该报告。
|
||||
5. `GET /api/diagnosis/{sessionId}/trace?runId=...` 返回 AIOps run、按顺序排列的 agent steps 和 tool invocations。
|
||||
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
|
||||
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
|
||||
8. `diagnosis_run.self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
||||
|
||||
## 4. 已知边界
|
||||
|
||||
- 当前 AIOps 使用轻量规则评估器,不是完整 LLM Verifier。
|
||||
- 完整运行仍依赖有效的 DB、Redis、Milvus/Zilliz、模型和 embedding 配置。
|
||||
- `mvp-demo` profile 使用 mock Prometheus 和 mock CLS,主要用于稳定演示。
|
||||
|
||||
@@ -0,0 +1,63 @@
|
||||
# Evidence Pipeline Demo Scenarios
|
||||
|
||||
这份清单用于面试时说明 Chat 证据链路如何覆盖 `PASS`、`LOW_CONFID`、`REJECT` 和 no-evidence 场景。
|
||||
|
||||
重点区别:
|
||||
|
||||
- Live demo 证明本地服务、工具、Trace、Feedback 主链路能跑通。
|
||||
- Fixture-backed eval 证明固定安全场景可以确定性回归,不依赖 LLM 当场随机输出。
|
||||
|
||||
## Scenario Matrix
|
||||
|
||||
| 场景 | 类型 | 输入/证据 | 期望讲点 |
|
||||
|---|---|---|---|
|
||||
| Payment timeout | Live 主路径 | `requests/payment-timeout-chat.json` | 完整 Chat -> Trace -> Feedback 闭环 |
|
||||
| Narrow HighCPU observation | Live 可尝试 + fixture-backed | `requests/narrow-highcpu-chat.json` / `mvp/eval/fixtures/narrow-highcpu-observation-pass.json` | Executor 只输出观察类 claim,Gatekeeper 验引用,Verifier PASS |
|
||||
| Hikari no-evidence | Live 可尝试 + fixture-backed | `requests/hikari-no-evidence-chat.json` / `mvp/eval/fixtures/hikari-no-evidence-negative-observation-pass.json` | `$.no_evidence` 只表示本次查询无匹配证据,Composer 不说“已排除” |
|
||||
| Unsupported claim filtering | Fixture-backed | `requests/safety-unsupported-claim-chat.json` / `mvp/eval/fixtures/unsupported-claim-filtering-low-confid.json` | Verifier 将 unsupported claim 降为 LOW_CONFID,最终答案不确认“主库故障” |
|
||||
| Fabricated invocation reject | Fixture-backed | `mvp/eval/fixtures/gatekeeper-fabricated-invocation-reject.json` | Gatekeeper 拦截伪造 invocation,最终 REJECT/降级 |
|
||||
| Composer fallback | Fixture-backed | `mvp/eval/fixtures/composer-fallback-no-raw-json-low-confid.json` | 即使 Composer 输出异常,也不能把 Executor JSON 泄漏给用户 |
|
||||
|
||||
## Trace Fields To Inspect
|
||||
|
||||
| 能力 | JSON path |
|
||||
|---|---|
|
||||
| Executor V2 输出 | `data.session.selfEvaluation.verifier_evaluation.executor_structured_output` |
|
||||
| Gatekeeper 结果 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` |
|
||||
| Gatekeeper 规则版本 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` |
|
||||
| 证据绑定校验 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.checked_bindings` |
|
||||
| Verifier claim checks | `data.session.selfEvaluation.verifier_evaluation.claim_checks` |
|
||||
| Composer 输出 | `data.session.selfEvaluation.verifier_evaluation.composer_output` |
|
||||
| 工具证据引用 | `data.toolInvocations[*].retrievalDetails.evidence_refs` |
|
||||
|
||||
## How To Present It
|
||||
|
||||
```text
|
||||
我把现场 demo 和固定 eval 分开。
|
||||
现场 demo 证明系统能跑通真实链路;
|
||||
fixture-backed eval 证明反幻觉安全场景可以稳定回归。
|
||||
Gatekeeper 的规则版本也进入 trace,所以后续调整阈值或规则时可以审计。
|
||||
```
|
||||
|
||||
## Optional Live Requests
|
||||
|
||||
手动发送某个请求样例:
|
||||
|
||||
```powershell
|
||||
$body = Get-Content -Raw -Encoding UTF8 "mvp/demo/requests/narrow-highcpu-chat.json"
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
然后保留响应里的 `runId`,查询同一 run 的 Trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace?runId=$runId"
|
||||
```
|
||||
|
||||
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
|
||||
@@ -0,0 +1,33 @@
|
||||
# 面试追问 Q&A
|
||||
|
||||
## 为什么不用普通 Chatbot?
|
||||
|
||||
这个项目的重点不是生成一段诊断文本,而是把诊断拆成可审计链路:Planner 拆解问题,Executor 调工具拿证据,Gatekeeper 用代码核验证据引用,Verifier 判断可推导性,Composer 生成最终表达。`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行,Trace 和 Feedback 都可以按 `sessionId + runId` 回放和定位。
|
||||
|
||||
## 为什么 RAG 要做成显式工具?
|
||||
|
||||
`lookup_knowledge` 保持显式工具调用,才能在 `tool_invocation` 里看到 Agent 查了什么、命中了什么、相关性等级是什么,以及最终答案是否真的使用了这些证据。隐式 Advisor 更方便,但不利于审计 Agent 决策。
|
||||
|
||||
## 怎么防止 Executor 幻觉?
|
||||
|
||||
Executor 不直接负责最终用户答案,而是输出 `executor_evidence_v2` 的微观事实和证据引用。Gatekeeper 会校验 `source_invocation_id`、`raw_path`、`evidence_excerpt` 是否真实存在;Verifier 再判断 claim 是否能由已验真的证据推出;Composer 只表达 Verifier 允许输出的内容。
|
||||
|
||||
## LOW_CONFID 是失败吗?
|
||||
|
||||
不是。`LOW_CONFID` 表示当前证据不足以支撑强结论,但系统仍然可以安全表达已确认事实和缺失信息。面试时可以把它作为“没有证据就不强答”的质量门禁,而不是模型能力失败。
|
||||
|
||||
## Prompt 改了怎么审计?
|
||||
|
||||
Chat verifier evaluation 里会记录 `prompt_audit.version`,并列出 planner、executor、verifier、composer 的 Prompt 版本和资源路径。它不保存完整 Prompt 文本,只保留用于回放和回归解释的紧凑元数据。
|
||||
|
||||
## Gatekeeper 改了怎么审计?
|
||||
|
||||
Gatekeeper 结果里记录 `gatekeeper_result.rule_set_version` 和已启用规则元数据摘要。规则执行仍是确定性 Java 代码,版本和规则元数据用于解释“这次引用验真用的是哪套规则”。
|
||||
|
||||
## 为什么现在不拆 SubAgent?
|
||||
|
||||
当前 MVP 的主要风险不是 Agent 数量不够,而是证据、验证和回归是否稳定。文档里的演进路线把 SubAgent 放在 P2:等故障类型、工具权限和评测集足够明确后再拆,避免只是移动复杂度。
|
||||
|
||||
## 为什么 baseline 比 live demo 更重要?
|
||||
|
||||
live demo 证明链路在当前环境能跑通,但 LLM 和外部依赖会波动。`mvp/eval` 的固定 fixture baseline 是确定性回归来源,用来判断 Prompt、工具、Gatekeeper、Verifier 或 Composer 的改动有没有让系统退化。
|
||||
@@ -0,0 +1,152 @@
|
||||
# 面试演示讲解稿
|
||||
|
||||
这是一份短时间 Agent 工程面试用讲解稿,不是完整系统文档。
|
||||
|
||||
## 1. 30 秒摘要
|
||||
|
||||
```text
|
||||
这是一个企业故障诊断 Agent MVP。
|
||||
它接收支付超时问题,规划排查步骤,调用证据工具,
|
||||
用 Verifier 检查答案,把完整 Trace 持久化,并支持用户反馈。
|
||||
```
|
||||
|
||||
关键主张不是“模型回答了一次”,而是:
|
||||
|
||||
```text
|
||||
系统能展示用了什么证据、答案如何被检查、如何用 sessionId + runId 精确回放这次诊断。
|
||||
```
|
||||
|
||||
## 2. Demo 流程
|
||||
|
||||
1. 用 `mvp-demo` profile 启动服务。
|
||||
2. 运行固定的支付超时请求。
|
||||
3. 打开 `mvp/demo/output/chat-response.json`。
|
||||
4. 打开 `mvp/demo/output/trace-response.json`。
|
||||
5. 指出证据工具和 verifier evaluation。
|
||||
6. 提交 feedback,并展示它挂在当前 run 上。
|
||||
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
|
||||
|
||||
## 3. 命令
|
||||
|
||||
启动服务:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
另开终端运行 Demo:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
可选自定义 session:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
||||
```
|
||||
|
||||
## 4. 展示什么
|
||||
|
||||
### 4.1 用户侧答案
|
||||
|
||||
文件:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
```
|
||||
|
||||
话术:
|
||||
|
||||
```text
|
||||
这是用户看到的答案。这里的 sessionId 是稳定的,同时响应里会返回 runId,所以我后面可以精确追踪这一次回答是怎么来的。
|
||||
```
|
||||
|
||||
### 4.2 证据 Trace
|
||||
|
||||
文件:
|
||||
|
||||
```text
|
||||
mvp/demo/output/trace-response.json
|
||||
```
|
||||
|
||||
话术:
|
||||
|
||||
```text
|
||||
这才是 Agent 工程最重要的部分。
|
||||
我可以检查 Agent 调用了哪些工具、每个工具拿到什么入参、是否成功、返回了什么证据预览。
|
||||
```
|
||||
|
||||
重点字段:
|
||||
|
||||
- `data.toolInvocations[*].toolName`
|
||||
- `data.toolInvocations[*].inputParams`
|
||||
- `data.toolInvocations[*].outputPreview`
|
||||
- `data.toolInvocations[*].success`
|
||||
|
||||
### 4.3 Verifier / 自评估
|
||||
|
||||
重点字段:
|
||||
|
||||
- `data.session.selfEvaluation`
|
||||
- `data.summary.hasVerifierEvaluation`
|
||||
|
||||
话术:
|
||||
|
||||
```text
|
||||
最终答案不是 Executor 原始输出直接返回。
|
||||
系统会基于持久化的工具 trace 做 Verifier 或规则自评估。
|
||||
这样系统可以区分 PASS、LOW_CONFID、REJECT,而不是假装每个答案都同样可信。
|
||||
```
|
||||
|
||||
### 4.4 反馈闭环
|
||||
|
||||
文件:
|
||||
|
||||
```text
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
必要时重新查询 Trace。
|
||||
|
||||
话术:
|
||||
|
||||
```text
|
||||
feedback 会挂在当前 diagnosis run 上。
|
||||
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
|
||||
```
|
||||
|
||||
### 4.5 回归故事
|
||||
|
||||
如果被问到稳定性,可以补充:
|
||||
|
||||
```text
|
||||
我把运行时 Demo 和离线 eval 分开。
|
||||
Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可以回归。
|
||||
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
|
||||
```
|
||||
|
||||
如果被问到怎么防止证据归因幻觉,可以补充:
|
||||
|
||||
```text
|
||||
Executor 的 claim 必须绑定 source_invocation_id、raw_path 和 evidence_excerpt。
|
||||
Gatekeeper 用代码核验这些引用,并把 rule_set_version 写进 trace。
|
||||
Verifier 只判断已核验证据能否推出 claim,Composer 只表达允许输出的内容。
|
||||
```
|
||||
|
||||
## 5. 强面试表达
|
||||
|
||||
```text
|
||||
我关注的是 Agent 工程表面:
|
||||
traceability、evidence persistence、verifier gating、feedback 和 regression checks。
|
||||
模型答案只是系统的一部分。
|
||||
更重要的是答案产出后,能否被审计、验证和持续改进。
|
||||
```
|
||||
|
||||
## 6. 主动说明限制
|
||||
|
||||
```text
|
||||
这个 MVP 仍依赖 MySQL、Redis、Milvus 和模型凭证。
|
||||
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
|
||||
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
|
||||
```
|
||||
@@ -0,0 +1,168 @@
|
||||
param(
|
||||
[string]$BaseUrl = "http://localhost:9900",
|
||||
[string]$SessionId = "mvp-demo-interview-payment-timeout-001",
|
||||
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
|
||||
[string]$OutputDir = "$PSScriptRoot/../output"
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
function Test-ServiceReachable {
|
||||
param([string]$Url)
|
||||
|
||||
try {
|
||||
$request = [System.Net.WebRequest]::Create($Url)
|
||||
$request.Method = "GET"
|
||||
$request.Timeout = 5000
|
||||
$response = $request.GetResponse()
|
||||
$response.Close()
|
||||
return $true
|
||||
} catch [System.Net.WebException] {
|
||||
if ($_.Exception.Response -ne $null) {
|
||||
$_.Exception.Response.Close()
|
||||
return $true
|
||||
}
|
||||
return $false
|
||||
}
|
||||
}
|
||||
|
||||
function Get-TraceData {
|
||||
param($TraceResponse)
|
||||
|
||||
if ($TraceResponse.PSObject.Properties.Name -contains "data") {
|
||||
return $TraceResponse.data
|
||||
}
|
||||
return $TraceResponse
|
||||
}
|
||||
|
||||
function Get-SelfEvaluation {
|
||||
param($TraceData)
|
||||
|
||||
if ($null -eq $TraceData -or $null -eq $TraceData.session) {
|
||||
return $null
|
||||
}
|
||||
return $TraceData.session.selfEvaluation
|
||||
}
|
||||
|
||||
function Get-ToolNames {
|
||||
param($TraceData)
|
||||
|
||||
if ($null -eq $TraceData -or $null -eq $TraceData.toolInvocations) {
|
||||
return @()
|
||||
}
|
||||
return @($TraceData.toolInvocations | ForEach-Object { $_.toolName } | Where-Object { $_ } | Sort-Object -Unique)
|
||||
}
|
||||
|
||||
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||
|
||||
Write-Host "Running interview demo preflight..."
|
||||
Write-Host "BaseUrl: $BaseUrl"
|
||||
Write-Host "SessionId: $SessionId"
|
||||
|
||||
if (-not (Test-ServiceReachable -Url $BaseUrl)) {
|
||||
throw "Service is not reachable: $BaseUrl. Start the app with mvp-demo profile first: mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo"
|
||||
}
|
||||
|
||||
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||
$request.Id = $SessionId
|
||||
$body = $request | ConvertTo-Json -Depth 8
|
||||
|
||||
$chatRequest = @{
|
||||
Method = "Post"
|
||||
Uri = "$BaseUrl/api/chat"
|
||||
ContentType = "application/json; charset=utf-8"
|
||||
Body = $body
|
||||
}
|
||||
$chat = Invoke-RestMethod @chatRequest
|
||||
|
||||
$chatPath = Join-Path $OutputDir "chat-response.json"
|
||||
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if (-not $runId) {
|
||||
throw "Chat response did not include runId; exact trace verification cannot continue."
|
||||
}
|
||||
|
||||
$traceRequest = @{
|
||||
Method = "Get"
|
||||
Uri = "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||
}
|
||||
$trace = Invoke-RestMethod @traceRequest
|
||||
|
||||
$tracePath = Join-Path $OutputDir "trace-response.json"
|
||||
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
$feedbackRequest = @{
|
||||
Method = "Post"
|
||||
Uri = "$BaseUrl/api/feedback"
|
||||
ContentType = "application/json; charset=utf-8"
|
||||
Body = $feedbackBody
|
||||
}
|
||||
$feedback = Invoke-RestMethod @feedbackRequest
|
||||
|
||||
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
|
||||
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
|
||||
|
||||
$traceData = Get-TraceData -TraceResponse $trace
|
||||
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
|
||||
$verifierEvaluation = $null
|
||||
if ($null -ne $selfEvaluation) {
|
||||
$verifierEvaluation = $selfEvaluation.verifier_evaluation
|
||||
}
|
||||
|
||||
$gatekeeperResult = $null
|
||||
$promptAudit = $null
|
||||
if ($null -ne $verifierEvaluation) {
|
||||
$gatekeeperResult = $verifierEvaluation.gatekeeper_result
|
||||
$promptAudit = $verifierEvaluation.prompt_audit
|
||||
}
|
||||
|
||||
$verdict = $null
|
||||
$gatekeeperStatus = $null
|
||||
$gatekeeperRuleSetVersion = $null
|
||||
$promptAuditVersion = $null
|
||||
if ($null -ne $verifierEvaluation) {
|
||||
$verdict = $verifierEvaluation.verdict
|
||||
}
|
||||
if ($null -ne $gatekeeperResult) {
|
||||
$gatekeeperStatus = $gatekeeperResult.status
|
||||
$gatekeeperRuleSetVersion = $gatekeeperResult.rule_set_version
|
||||
}
|
||||
if ($null -ne $promptAudit) {
|
||||
$promptAuditVersion = $promptAudit.version
|
||||
}
|
||||
$toolNames = Get-ToolNames -TraceData $traceData
|
||||
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
||||
|
||||
$summary = [ordered]@{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
baseUrl = $BaseUrl
|
||||
chatSuccess = $chat.data.success
|
||||
verdict = $verdict
|
||||
gatekeeperStatus = $gatekeeperStatus
|
||||
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
|
||||
promptAuditVersion = $promptAuditVersion
|
||||
toolNames = $toolNames
|
||||
paths = [ordered]@{
|
||||
chat = $chatPath
|
||||
trace = $tracePath
|
||||
feedback = $feedbackPath
|
||||
summary = $summaryPath
|
||||
}
|
||||
}
|
||||
|
||||
$summary | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path $summaryPath
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "Interview demo preflight completed."
|
||||
Write-Host "Verdict: $($summary.verdict)"
|
||||
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
|
||||
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
|
||||
Write-Host "Summary: $summaryPath"
|
||||
@@ -0,0 +1,246 @@
|
||||
# 10 分钟面试演示脚本
|
||||
|
||||
**用途**:面试现场按步骤演示
|
||||
**目标**:展示从问题到证据、验证、Trace、反馈的闭环
|
||||
**前置条件**:服务以 `mvp-demo` profile 启动
|
||||
|
||||
更完整的 runbook 见 [README.md](README.md),字段检查见 [trace-inspection-checklist.md](trace-inspection-checklist.md)。
|
||||
|
||||
## 0. 开场话术
|
||||
|
||||
```text
|
||||
我会演示一个支付超时诊断。
|
||||
重点不是看模型给出一段答案,而是看这个答案背后的 Agent 执行链路:
|
||||
Planner 怎么拆解,Executor 调了哪些工具,Verifier 如何判断证据是否支撑答案,以及最终如何通过 sessionId 回放。
|
||||
```
|
||||
|
||||
## 1. 启动服务
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
服务地址:
|
||||
|
||||
```text
|
||||
http://localhost:9900
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- `mvp-demo` profile 使用 mock Prometheus 和 mock CLS。
|
||||
- 演示不依赖真实线上故障。
|
||||
- MySQL、Redis、Milvus/Zilliz 和模型配置仍需要可用。
|
||||
|
||||
## 2. 演示 Chat 诊断
|
||||
|
||||
推荐使用固定脚本:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会写出:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
这里我用固定 sessionId 跑一个支付接口超时问题。
|
||||
固定 sessionId 的好处是保留多轮上下文;每次诊断还会返回 runId,后面 trace 和 feedback 都用这个 runId 精确关联到同一次运行。
|
||||
```
|
||||
|
||||
## 3. 展示用户答案
|
||||
|
||||
打开:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
```
|
||||
|
||||
重点看:
|
||||
|
||||
```text
|
||||
data.sessionId
|
||||
data.runId
|
||||
data.answer
|
||||
```
|
||||
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
这是用户看到的答案。
|
||||
但这个项目的重点不是这段文字,而是这段文字是否有证据链。
|
||||
接下来我用同一个 sessionId 加 runId 查 trace。
|
||||
```
|
||||
|
||||
## 4. 展示 Trace
|
||||
|
||||
打开:
|
||||
|
||||
```text
|
||||
mvp/demo/output/trace-response.json
|
||||
```
|
||||
|
||||
重点看:
|
||||
|
||||
```text
|
||||
data.session.sessionId
|
||||
data.session.agentFlow
|
||||
data.steps[*].agentName
|
||||
data.toolInvocations[*].toolName
|
||||
data.toolInvocations[*].inputParams
|
||||
data.toolInvocations[*].outputPreview
|
||||
data.toolInvocations[*].retrievalLayer
|
||||
data.toolInvocations[*].relevanceLevel
|
||||
data.summary.hasVerifierEvaluation
|
||||
data.session.selfEvaluation.verifier_evaluation.prompt_audit.version
|
||||
data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version
|
||||
```
|
||||
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
这里能看到三个层次:
|
||||
第一,session 记录了这次诊断的问题、答案、耗时和自评估。
|
||||
第二,agent_step 记录 Planner、Executor、Verifier 的模型步骤。
|
||||
第三,tool_invocation 记录真实工具调用,包括 lookup_knowledge、日志和指标。
|
||||
|
||||
所以这不是一个黑盒 Chatbot,而是一条可以回放的诊断链路。
|
||||
```
|
||||
|
||||
## 5. 展示知识库检索
|
||||
|
||||
在 trace 中找到 `lookup_knowledge`。
|
||||
|
||||
重点看:
|
||||
|
||||
```text
|
||||
toolName = lookup_knowledge
|
||||
inputParams.query
|
||||
retrievalLayer
|
||||
l0MatchCount
|
||||
l1MatchCount
|
||||
relevanceLevel
|
||||
retrievalDetails
|
||||
outputPreview
|
||||
```
|
||||
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
知识库检索保留为显式工具,而不是藏在 Advisor 里。
|
||||
这样面试官或线上排查人员能看到:Agent 查了什么 query,命中了哪个知识域,检索层是 L0/L1 还是混合,相关性等级是什么。
|
||||
|
||||
底层检索现在走 VectorSearchService,优先 Spring AI VectorStore,失败时 fallback 到 Milvus SDK。
|
||||
```
|
||||
|
||||
## 6. 展示 Verifier
|
||||
|
||||
在 trace 中查看:
|
||||
|
||||
```text
|
||||
data.session.selfEvaluation
|
||||
data.summary.hasVerifierEvaluation
|
||||
```
|
||||
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
Verifier 不做新检索,只看工具 trace 汇总。
|
||||
它会把 Executor 答案里的关键事实拆出来,判断每条事实是 direct_evidence、indirect_support、no_evidence 还是 contradicted。
|
||||
|
||||
如果 PASS,就输出原答案。
|
||||
如果 LOW_CONFID,可以补证据或加低置信提示。
|
||||
如果 REJECT,就降级输出,只保留已确认信息。
|
||||
|
||||
Prompt 和 Gatekeeper 的版本也会进入 trace。
|
||||
`prompt_audit.version` 用于说明本次 Chat 使用哪套 Prompt 契约,`gatekeeper_result.rule_set_version` 用于说明引用验真的规则版本。
|
||||
固定 fixture baseline 是回归判断来源,live demo 主要证明当前环境链路可跑通。
|
||||
```
|
||||
|
||||
## 7. 展示反馈闭环
|
||||
|
||||
打开:
|
||||
|
||||
```text
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
重点看:
|
||||
|
||||
```text
|
||||
success
|
||||
caseId
|
||||
```
|
||||
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
用户反馈 useful 会写回当前 diagnosis_run。
|
||||
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
|
||||
|
||||
这里 status 和 feedback 是分开的:
|
||||
status 表示执行是否成功,feedback 表示用户是否认可。
|
||||
```
|
||||
|
||||
## 8. 可选演示 AIOps
|
||||
|
||||
如果时间允许,再演示 AIOps payload。
|
||||
|
||||
请求示例见:
|
||||
|
||||
```text
|
||||
mvp/demo/README.md
|
||||
```
|
||||
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
AIOps 有两个模式。
|
||||
有 payload 时进入 PAYLOAD_TARGETED,报告必须聚焦这个告警。
|
||||
没有 payload 时进入 AUTO_DISCOVERY,先发现活跃告警再排查。
|
||||
|
||||
我专门加了 recommended lookup_knowledge query,把 alertName、service、severity、description 等字段稳定送入知识库检索,避免 Agent 随意扩展问题范围。
|
||||
```
|
||||
|
||||
## 9. 结束总结
|
||||
|
||||
```text
|
||||
这个 Demo 展示的是一个完整闭环:
|
||||
|
||||
用户问题
|
||||
-> Agent 规划和执行
|
||||
-> 显式工具证据
|
||||
-> Verifier / self_evaluation
|
||||
-> Trace 回放
|
||||
-> 用户反馈
|
||||
-> 案例沉淀
|
||||
|
||||
我把重点放在 Agent 工程能力:可追踪、可验证、可回归、可演进。
|
||||
```
|
||||
|
||||
## 10. 如果现场失败
|
||||
|
||||
如果模型或外部组件不可用,不要硬跑。可以直接打开上一次输出:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
降级话术:
|
||||
|
||||
```text
|
||||
现场环境依赖 MySQL、Redis、Milvus 和模型服务。
|
||||
如果外部服务不可用,我会用固定输出讲 trace 结构。
|
||||
因为这个项目的核心不是一次在线请求,而是诊断链路如何被记录、检查和回放。
|
||||
```
|
||||
@@ -0,0 +1,58 @@
|
||||
# Trace 检查清单
|
||||
|
||||
运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
|
||||
|
||||
## 1. Session
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.steps[*].agentName` | 是否有 Planner / Executor / Verifier 或等价步骤 | 流程被拆成可检查的 Agent 步骤 |
|
||||
| `data.steps[*].thought` | 是否有高层步骤摘要 | 内部过程可审计,不只看最终文本 |
|
||||
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
||||
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
||||
|
||||
## 3. 工具证据
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.toolInvocations[*].toolName` | 是否包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 | Agent 通过工具收集证据,而不是无依据猜测 |
|
||||
| `data.toolInvocations[*].inputParams` | 是否能看到每个工具的入参 | 工具输入可审计、可调试 |
|
||||
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
|
||||
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
|
||||
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
|
||||
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||
|
||||
## 4. Summary
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.summary.persistedStepCount` | step 行是否持久化 | Trace 来自存储,不是响应内存 |
|
||||
| `data.summary.persistedToolCallCount` | tool 行是否持久化 | 工具证据在请求结束后仍可回放 |
|
||||
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
||||
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
||||
|
||||
## 5. 好的结果长什么样
|
||||
|
||||
```text
|
||||
同一个 session id + run id
|
||||
-> 最终答案
|
||||
-> 持久化 agent steps
|
||||
-> 持久化 evidence tool calls
|
||||
-> verifier / self-evaluation
|
||||
-> feedback attached to the same run
|
||||
```
|
||||
Reference in New Issue
Block a user