feat: archive mvp demo trace acceptance

This commit is contained in:
zhuyongxin
2026-07-03 16:25:00 +08:00
parent 5b827fe90e
commit 6919092b83
24 changed files with 1343 additions and 2 deletions
+94
View File
@@ -0,0 +1,94 @@
# MVP Demo Runbook
This demo proves the MVP flow from user question to persisted diagnosis trace.
## Prerequisites
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
- Security and secret cleanup are intentionally out of scope for this MVP slice.
- The `mvp-demo` profile enables mock Prometheus and CLS providers so log and metric tools can return repeatable evidence.
## Start
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
The service listens on:
```text
http://localhost:9900
```
## 1. Run Chat Diagnosis
```powershell
$sessionId = "mvp-demo-payment-timeout-001"
$body = @{
Id = $sessionId
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
Expected result:
- `data.success` is `true`.
- `data.sessionId` equals `mvp-demo-payment-timeout-001`.
- `data.answer` contains a diagnosis answer.
## 2. Query Trace
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
```
Expected result:
- `code` is `200`.
- `data.session.sessionId` equals the chat session id.
- `data.steps` contains planner/executor/verifier records for complex questions.
- `data.toolInvocations` contains evidence tool calls such as `lookup_knowledge`, `query_logs`, or `query_metrics`.
- `data.session.selfEvaluation` contains verifier or rule evaluation when available.
## 3. Submit Feedback
```powershell
$feedback = @{
sessionId = $sessionId
feedback = "useful"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/feedback" `
-ContentType "application/json" `
-Body $feedback
```
Expected result:
- `success` is `true`.
- A later trace query shows `data.session.feedback` as `useful`.
## Demo Story
The important interview story is:
```text
one session id
-> user question
-> multi-agent execution
-> evidence tools
-> verifier/self-evaluation
-> final answer
-> feedback
-> trace API for replay and audit
```
+39
View File
@@ -0,0 +1,39 @@
# Payment Timeout Acceptance Case
## Goal
Validate that the MVP can diagnose a payment timeout incident and expose the complete trace for replay.
## Input
- Session id: `mvp-demo-payment-timeout-001`
- Question: `支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
- Profile: `mvp-demo`
## Acceptance Criteria
1. Chat returns a successful answer with the same session id.
2. Trace API returns session metadata, final answer, ordered agent steps, and ordered tool invocations.
3. Trace contains enough evidence to explain which tools were used and whether verifier/self-evaluation was persisted.
4. Feedback can be submitted for the same session id.
5. A follow-up trace query shows the persisted feedback value.
## Trace Fields To Inspect
- `data.session.query`
- `data.session.answer`
- `data.session.selfEvaluation`
- `data.session.feedback`
- `data.steps[*].agentName`
- `data.steps[*].thought`
- `data.toolInvocations[*].toolName`
- `data.toolInvocations[*].inputParams`
- `data.toolInvocations[*].outputPreview`
- `data.toolInvocations[*].retrievalDetails`
- `data.summary`
## Known Limits
- This case is not a full offline test. It still requires valid infrastructure for chat, persistence, vector search, and model calls.
- Mock logs and metrics are enabled by the `mvp-demo` profile to make those evidence tools repeatable.
- Sensitive configuration cleanup is deferred by current MVP priority.
+268
View File
@@ -0,0 +1,268 @@
# MVP Agent 工程决策记录
本文记录 MVP 实现过程中已经落地的一些关键修复、取舍和工程判断。目标不是写流水账,而是沉淀面试时可以讲清楚的 Agent 工程思路。
---
## 1. 统一流式与非流式 Chat 主链路
### 背景
早期 `/api/chat` 和 `/api/chat_stream` 是两条不同实现:
- 非流式接口会走复杂度判断,并可能进入 Planner / Executor / Verifier 多 Agent 流程。
- 流式接口直接创建单个 ReactAgent,然后 `agent.stream()` 输出 token。
这导致两个接口表面都是 chat,实际能力不一致:流式接口不会进入 verifier、不会沉淀完整诊断链路,也不容易和 `diagnosis_session`、`tool_invocation` 对齐。
### 决策
将两个接口统一到同一条核心链路:
```text
getOrCreateSession
-> 读取会话历史
-> ChatService.executeChatWithStrategy(...)
-> 写回会话历史
```
接口差异只保留在传输层:
- `/api/chat` 返回完整 JSON。
- `/api/chat_stream` 通过 SSE 分块发送最终答案。
### 取舍
这样会牺牲原来的 token 级实时流式体验,但换来业务行为一致、诊断链路一致、Verifier 和 evidence trace 一致。
对 MVP 来说,优先保证“同一个问题不因接口不同而进入不同智能链路”,比 token 级流式更重要。
---
## 2. 会话 ID 与诊断链路统一
### 背景
原实现中:
- `ChatController` 用前端传入的 `Id` 在 JVM 内存里维护历史消息。
- `ChatService` 每次执行又生成新的 8 位 sessionId,作为 `diagnosis_session` 和工具调用追踪 ID。
这会造成前端会话、后端诊断会话、工具证据链三者分裂。
### 决策
将前端 chat session id 作为后端诊断链路的主 session id:
- Redis `SessionContext` 保存聊天历史。
- `diagnosis_session.session_id` 复用同一个 id。
- `RunnableConfig.metadata.sessionId` 和 `SessionContextHolder` 也使用同一个 id。
- `tool_invocation`、`agent_step`、verifier evaluation 都可按同一 session id 串起来。
### 企业级意义
Agent 系统最怕“答得出来但查不清”。统一 session id 后,一次用户请求可以完整追踪:
```text
用户问题 -> Agent 步骤 -> 工具调用 -> Verifier 判断 -> 最终答案 -> 用户反馈
```
这是可观测、可审计、可复盘的基础。
---
## 3. 引入统一 ToolInvocationRecorder
### 背景
Verifier 需要结构化证据链,但原实现只有 `lookup_knowledge` 主动写入 `tool_invocation`。
`query_logs`、`query_metrics` 虽然返回 JSON,但没有统一落库,导致 verifier 看不到日志、指标等 evidence tool 的稳定记录。
### 决策
新增 `ToolInvocationRecorder`,作为所有 evidence tool 的统一落库入口。
当前接入:
- `lookup_knowledge`
- `query_logs`
- `query_metrics`
记录字段包括:
- tool name
- input params
- output preview
- output length
- success
- error message
- duration
- trace id / domain details
### 企业级意义
这一步把 Agent 从“模型说它查过”推进到“系统能证明它查过”。
后续 verifier 不应该依赖模型自由文本回忆工具调用,而应该消费结构化 trace summary。
---
## 4. Verifier 作为事实约束层
### 背景
普通 Agent 很容易在工具调用后直接生成答案,但企业场景更关心:
- 关键结论有没有证据
- 证据是直接证据还是间接支持
- 哪些事实缺口需要人工介入
- 工具失败时是否诚实降级
### 决策
保留 Planner / Executor / Verifier 三角色:
- Planner 负责拆解问题。
- Executor 负责执行查询与形成初稿。
- Verifier 负责基于 `tool_trace_summary` 做事实核查。
Verifier 输出结构化 JSON,包括:
- verdict
- groundedness_score
- critical_fact_count
- facts_checked
- rationale
### 取舍
Verifier 会增加一次模型调用成本,但换来可解释性和质量约束。对企业级 Agent 来说,这是值得的。
---
## 5. 从手写编排切换到 SupervisorAgent
### 背景
之前 `ChatService.executeChatComplex()` 中构建了 `SupervisorAgent`,但实际仍然手写调用:
```text
planner -> executor -> verifier
```
这会造成代码与设计不一致,维护者容易误以为当前已经由 Supervisor 调度。
### 决策
复杂问题真正切换到 `SupervisorAgent.invoke(...)`。
Supervisor 负责路由:
```text
chat_supervisor -> chat_planner
chat_supervisor -> chat_executor
chat_supervisor -> chat_verifier
chat_supervisor -> FINISH
```
外层仍保留:
- verifier 输出解析
- PASS / LOW_CONFID / REJECT 判定
- retry context
- fallback
- evaluation 入库
### 验证
新增离线专项测试 `ChatServiceSupervisorAgentTest`,使用 scripted `ChatModel` 验证真实 SupervisorAgent 路由顺序,不依赖真实 LLM、MySQL、Redis。
### 企业级意义
这让项目不只是“自己写 if/else 多 Agent”,而是使用框架原生 multi-agent orchestration,同时保留业务层的质量门控。
---
## 6. 文档上传路径语义统一
### 背景
上传文档时,`DocumentManagementService.saveToLocal()` 返回带 `knowledge_base` 前缀的路径。
而 `KnowledgeIndexService.readDocument()` 又执行:
```java
Paths.get(knowledgeBasePath, filePath)
```
这可能拼出:
```text
knowledge_base/knowledge_base/...
```
最终表现为 L0 命中文档,但读取原文失败。
### 决策
统一路径语义:
- 新上传文档存相对 `knowledge.base-path` 的路径,例如 `payment/runbook.md`。
- `readDocument()` 兼容新旧路径:
- 相对路径
- 已带 base path 的旧相对路径
- 绝对路径
### 企业级意义
知识库检索不能只看“命中”,还要保证命中后的内容可读、可引用、可追踪。
这是 RAG / Agent 系统里很典型的工程细节:检索质量问题不一定来自模型,也可能来自路径、元数据、索引和原文之间的语义不一致。
---
## 7. MVP 阶段的优先级取舍
当前主动暂缓的问题:
- 敏感配置外置与密钥轮换
- CORS / Redis 反序列化安全边界
- 默认 `mvn test` 离线化
原因不是这些不重要,而是当前目标是先跑通并讲清楚 MVP Agent 工程闭环。
短期优先目标:
```text
可演示 -> 可观测 -> 可验证 -> 可复盘
```
安全和完整测试体系属于企业落地必须项,但可以在 MVP 主链路稳定后作为下一阶段补齐。
---
## 8. 后续建议
下一阶段建议聚焦“可复现 MVP Demo”:
1. 增加 `local-demo` 或 `mvp-demo` profile。
2. 准备固定诊断 case,例如“支付接口超时”。
3. 提供一键初始化知识库样例。
4. 提供一键触发复杂诊断请求的脚本。
5. 增加 trace 查询接口:
```text
GET /api/diagnosis/{sessionId}/trace
```
该接口聚合:
- diagnosis_session
- agent_step
- tool_invocation
- verifier evaluation
- final answer
- feedback
这样 MVP 就能从“功能实现”升级为“企业级 Agent 工程作品”。
+39
View File
@@ -0,0 +1,39 @@
# MVP Demo Profile 与 Trace 查询接口
## 背景
MVP 已经能跑多 Agent 诊断、工具调用、Verifier 和反馈,但对外展示时仍然缺少一个稳定的复盘入口。面试官或评审如果想确认一次 Agent 回答是否可信,不能只看最终答案,还需要看到用户原始问题、Agent 步骤顺序、工具调用证据、Verifier / self-evaluation、最终答案和用户反馈。
## 决策
新增 `mvp-demo` profile 和 trace 查询接口:
```text
GET /api/diagnosis/{sessionId}/trace
```
接口聚合:
- `diagnosis_session`
- `agent_step`
- `tool_invocation`
- `self_evaluation`
- `feedback`
同时在 `mvp/demo` 下沉淀端到端验收 case,把启动、提问、查 trace、提交 feedback 串成一条可演示路径。
## 取舍
`mvp-demo` profile 不是完整离线 mock 环境,仍然复用当前真实 DB / Redis / Milvus / LLM 配置,只显式打开日志和指标 mock。原因是当前阶段目标是展示企业级 Agent 工程闭环,不是隐藏真实集成复杂度。
这让 MVP 的讲述从“我实现了一个聊天接口”升级为:
```text
我实现了一条可执行、可观测、可验收、可复盘的 Agent 诊断链路。
```
## 面试表达
- 我没有把 trace 塞进 chat 返回值,而是做成独立只读观测接口,保持执行链路和观测链路解耦。
- Trace API 复用已经沉淀的 `diagnosis_session`、`agent_step`、`tool_invocation` 三张表,没有引入新的 schema 风险。
- Demo profile 只做最小 overlay,让日志和指标工具可重复,保留真实基础设施集成,方便说明 MVP 与生产化之间的差距。