diff --git a/devflow/glossary/CONTEXT.md b/devflow/glossary/CONTEXT.md
index 756a03c..45b8951 100644
--- a/devflow/glossary/CONTEXT.md
+++ b/devflow/glossary/CONTEXT.md
@@ -175,6 +175,10 @@
- 定义:Tool 调用及结果投影的生命周期状态,固定为 `PROJECTING`、`READY`、`ERROR`。
- 边界:它只说明调用记录是否完成,不说明结果是否包含证据。
+### Durable Audit
+- 定义:为 Diagnosis Trace 长期保存的 Run、Agent 模型步骤和 Tool 调用元数据,用于 exact sessionId/runId 回放、评测和运维核对。
+- 边界:只保存有界、脱敏、可长期保留的身份、状态、耗时、预算和结果摘要;不保存 Prompt、Thought、完整 Tool 参数、raw response 或 Redis canonical invocation。
+
### Evidence Status
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
- 边界:`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
diff --git a/devflow/index.md b/devflow/index.md
index 8923237..aecb3d7 100644
--- a/devflow/index.md
+++ b/devflow/index.md
@@ -43,3 +43,4 @@
| 2026-07-21 | single-react-evidence-semantic-guards | Deterministic evidence validation, isolated semantic review and fail-closed diagnosis release | Harness/EvidenceGuard/SemanticGuard/Release | ISS-014, EvidenceGuard, verified snapshot, SemanticGuard, repair, fallback, release policy | openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards | archived |
| 2026-07-21 | single-react-chat-application-usecase | Internal Chat application use case with isolated routing, fixed executors and safe PreviousTurn | Harness/Chat application/Run persistence | ISS-014, Intent Router, PreviousTurn, PublishedResult, V012, observer, cancellation | openspec/changes/archive/2026-07-21-single-react-chat-application-usecase | archived |
| 2026-07-21 | single-react-chat-sse-cutover | Unique named-event Chat SSE endpoint, bounded production Harness wiring and strict frontend consumer | Chat/SSE/Harness production wiring | ISS-014, /api/chat, SSE, metadata, status, content, failure, done, disconnect, bounded executor | openspec/changes/archive/2026-07-22-single-react-chat-sse-cutover | archived |
+| 2026-07-22 | single-react-cleanup-e2e | Remove legacy Agent paths, add bounded Harness audit, and complete exact-run live acceptance | Chat/Harness/cleanup/E2E | ISS-014, single ReAct Agent, durable audit, named SSE, exact run, Flyway V013 | openspec/changes/archive/2026-07-22-single-react-cleanup-e2e | archived |
diff --git a/devflow/projects/2026-07-22-single-react-cleanup-e2e/acceptance.md b/devflow/projects/2026-07-22-single-react-cleanup-e2e/acceptance.md
new file mode 100644
index 0000000..c629fab
--- /dev/null
+++ b/devflow/projects/2026-07-22-single-react-cleanup-e2e/acceptance.md
@@ -0,0 +1,55 @@
+# Acceptance: single-react-cleanup-e2e
+
+## Result
+
+Accepted. 阶段 7 的清理、Harness-native audit、文档收口和 live E2E 已完成;最终发布为安全 `FALLBACK`,没有释放未验证 Draft。
+
+## Static Verification
+
+- `openspec validate --all --strict`: 22/22 passed。
+- `node --check src/main/resources/static/app.js`: passed。
+- `git diff --check`: passed。
+- production legacy path、旧 Agent-facing Tool contract 和持久化敏感 payload 扫描:0 matches。
+- 最终 Tool 时间窗日志扫描:原始 query、request/response、Tool Call ID、Prompt、stack、debug instrumentation 均为 0;只保留有界 metadata。
+
+## Script Verification
+
+- Deterministic Maven regression: 60 suites / 221 tests,0 failures,0 errors,3 skipped。
+- `mvn -q -DskipTests compile`: passed。
+- `mvn -q -DskipTests package`: passed。
+- `mvn -q -Dtest=QueryLogsToolsTest,QueryLogsResultProjectorTest,CanonicalInvocationStoreTest test`: passed。
+- Flyway 9.22.3 repair + migrate: V012 checksum repaired,V013 applied,schema current version 013。
+- Maven `spring-boot:run` with `mvp-demo`: Flyway 13 migrations validated,JPA schema validation passed,Tomcat 9900 started,Harness/Audit wiring active。
+- `run-payment-timeout-demo.ps1`: strict named SSE validation passed for exact final session/run。
+- `scripts/query_mysql.py`: exact diagnosis_run、agent_step、tool_invocation queries passed。
+
+## Final Live Evidence
+
+- sessionId: `mvp-demo-payment-timeout-stage7-20260722-1741`
+- runId: `363f481c-33b8-42e7-8699-428a6ec61806`
+- SSE: `metadata -> status -> status -> status -> content -> done`
+- diagnosis_run: `status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=2`、`total_token_count=11764`。
+- AgentStep: 2 exact-run rows,唯一 agent `diagnosis_agent`,`thought IS NULL`,只含角色/数量/Tool name metadata。
+- ToolInvocation: `lookup_knowledge` 和 Mock `query_logs` 各 1 行,均 `READY/EVIDENCE_FOUND/success=1`;input/details 只有 framework Tool Call ID、状态和字节数。
+
+## Browser / Manual Verification
+
+- 未执行浏览器点击验收;本阶段公开协议由真实 HTTP SSE 脚本和前端 JavaScript syntax/contract tests 覆盖。
+
+## Remaining Risks
+
+- live 结果为 EvidenceGuard/SemanticGuard 约束下的安全 `FALLBACK`,不是业务根因成功发布;这是允许的最终释放状态。
+- `diagnosis_run.step_count` 仍为空,但 exact AgentStep 查询返回 2 行;该历史汇总字段不作为本阶段 release gate。
+- `query_logs` 使用 Mock;真实 CLS 与生产业务 `query_mysql` datasource 仍属于后续接入范围。
+- `application-local.yml` 含本地内部配置且被 Git 忽略;未 stage、未提交、未输出凭据。
+
+## Migration And Rollback
+
+- L4 endpoint/frontend cleanup 必须整体回滚阶段 7 commit,不恢复双轨 endpoint 或旧 Tool annotations。
+- V013 仅幂等增加缺失列/索引;应用回滚时保留新增 nullable 列,避免破坏已写数据,不执行 destructive down migration。
+- Flyway repair 已将远端 V012 checksum 对齐当前迁移;V013 保证旧库与 fresh database 最终 schema 一致。
+
+## Archive
+
+- OpenSpec archive: completed at `openspec/changes/archive/2026-07-22-single-react-cleanup-e2e`。
+- Delta specs synced: created main `single-react-cleanup-e2e` spec and removed the temporary AiOps-preservation requirement from `single-react-chat-sse-cutover`。
diff --git a/devflow/projects/2026-07-22-single-react-cleanup-e2e/brief.md b/devflow/projects/2026-07-22-single-react-cleanup-e2e/brief.md
new file mode 100644
index 0000000..ec23c4d
--- /dev/null
+++ b/devflow/projects/2026-07-22-single-react-cleanup-e2e/brief.md
@@ -0,0 +1,30 @@
+# Brief: single-react-cleanup-e2e
+
+## Background
+
+ISS-014 阶段 0-6B 已建立单一 Diagnosis ReAct Agent、Harness、ACI Tool、Guard 和 named SSE,但仓库仍存在 legacy AiOps/Sequential/Redis Session 链、旧 Tool contract、副作用式审计和过时文档。阶段 7 负责物理清理、durable metadata audit 和最终 live E2E。
+
+## Goals
+
+- 公开诊断只保留 `POST /api/chat`,业务代码只保留一个拥有 Tool loop 的 `diagnosis_agent`。
+- RAG/log 仅作为 Harness backend;Agent-facing Tool 只来自 `HarnessEvidenceTools`。
+- AgentStep 与 ToolInvocation 使用 exact sessionId/runId,长期持久化仅包含有界 metadata。
+- 通过 Maven 启动、named SSE、日志与 MySQL exact-run 查询完成最终验收。
+
+## Scope
+
+- 删除 legacy Controller、Service、Hook、ThreadLocal、prompt、前端入口、测试和死文档。
+- 新增 Harness-native Agent/Tool durable audit,修正文档、issue 与 OpenSpec strict 缺陷。
+- 为旧 V012 数据库增加幂等 V013 兼容迁移,并校准可重复 payment-timeout demo。
+
+## Non-goals
+
+- 不实现真实 CLS 或生产业务 MySQL Tool datasource。
+- 不修改三类 Intent、DiagnosisDraft、EvidenceGuard、SemanticGuard 或 Release Policy 语义。
+- 不保留 legacy endpoint、兼容分支、Graph 或第二套 Tool Call ID。
+
+## Classification
+
+- Scale: `complex`
+- Interface impact: L4 breaking HTTP/frontend cleanup
+- OpenSpec: `single-react-cleanup-e2e`
diff --git a/devflow/projects/2026-07-22-single-react-cleanup-e2e/decisions.md b/devflow/projects/2026-07-22-single-react-cleanup-e2e/decisions.md
new file mode 100644
index 0000000..f12875b
--- /dev/null
+++ b/devflow/projects/2026-07-22-single-react-cleanup-e2e/decisions.md
@@ -0,0 +1,111 @@
+# Decisions: single-react-cleanup-e2e
+
+## Discover Status
+
+- Checkpoint: Discover
+- Capability source: `sm-flow` + `grill-with-docs` + `gitnexus-refactoring`。
+- GitNexus index 停在 `2362665`,当前为 `bc36248`;刷新会改写用户已修改的 `AGENTS.md`,因此不安全。使用 OpenSpec/devflow、`rg` 全引用扫描、源码阅读、编译和测试作为 fallback。
+- AGENTS.md 指定的 `codebase-retrieval` 与 LSP 工具在当前环境不可用,已显式记录限制。
+- Scale: complex。涉及 L4 endpoint 删除、跨模块物理清理、Tool/Trace ownership、安全持久化和 live E2E。
+
+## Question Pool
+
+| # | 维度 | 问题 | 模式 | 状态 |
+|---|---|---|---|---|
+| Q1 | 范围 | 哪些旧 Agent/Service/Hook/Session 只有自引用测试,哪些仍在生产可达? | evidence-driven | 已解决 |
+| Q2 | 协议 | 阶段 7 是否必须删除 `/api/ai_ops` 与旧 Session endpoints? | evidence-driven | 已解决 |
+| Q3 | Tool | RAG/log 旧实现哪些可复用,哪些 Agent-facing contract/副作用必须删除? | evidence-driven | 已解决 |
+| Q4 | Trace | 新 Harness 如何在不泄漏 raw/prompt/thought 的前提下满足 AgentStep/ToolInvocation E2E? | evidence-driven | 已解决 |
+| Q5 | 文档 | ISS-012/ISS-013 和现有架构/Demo 文档如何收口? | evidence-driven | 已解决 |
+| Q6 | 验收 | 最终 live E2E 必须证明哪些 exact-run 事实,哪些外部系统明确不声称 live? | evidence-driven | 已解决 |
+| Q7 | 回滚 | L4 endpoint 删除如何迁移与回滚? | evidence-driven | 已解决 |
+
+## Evidence-driven
+
+| 结论 | 证据来源 | 是否已汇报用户 |
+|---|---|---|
+| `ChatService` 没有生产调用方,只剩自身单元/Smoke tests。 | `rg ChatService` | 已汇报 |
+| `/api/ai_ops` 仍由 bundled frontend 按钮调用,并运行 Supervisor + Planner + Executor 多 Agent。 | `AiOpsController`、`AiOpsService`、`app.js`、`index.html` | 已汇报 |
+| ISS-014 总体验收要求旧多 Agent/Graph 不存在,阶段 6B 只暂时保持 AiOps,阶段 7 负责清理/处置。 | ISS-014、阶段 6B design/acceptance | 已汇报 |
+| `/api/chat/clear` 与 session info/runs 没有 frontend caller;Redis SessionManager 只由该 Controller 和测试使用。 | controller/frontend/session 引用扫描 | 已汇报 |
+| `LookupKnowledgeTool` 和 `QueryLogsTools` 被新 Harness adapter 复用,但仍携带旧 `@Tool`、ThreadLocal/recorder 副作用。 | `HarnessChatConfiguration`、Tool source | 已汇报 |
+| 新 ToolBoundary 写 Redis canonical invocation,但没有 `tool_invocation` durable audit;最终 DB E2E 会缺 Tool rows。 | Harness boundary/config 引用扫描 | 已汇报 |
+| 旧 `AgentLoggingHook` 仍回退 ThreadLocal,并持久化 thought/model正文;不满足新安全边界。 | `AgentLoggingHook.java` | 已汇报 |
+| 当前架构、Agent、Harness 和 lifecycle 文档仍描述 Planner/Executor/Verifier/Composer 与 AIOps 双入口。 | `mvp/architecture/*.md` | 已汇报 |
+
+## User-interview
+
+- 无新增 user-interview。唯一公开 Chat、旧多 Agent/Graph 物理删除、无兼容分支、最终 E2E 和外部 Mock 边界均已由 ISS-014 与用户的逐阶段自动执行授权冻结。
+
+## Key Decisions
+
+- 删除 legacy AiOps endpoint 而不是迁移到第二个 use case;所有诊断统一进入 `/api/chat` 的 Intent Router。
+- 删除旧 Session endpoints/Redis conversation context;安全 PreviousTurn 只来自 `diagnosis_run.published_result`。
+- RAG/log 查询实现保留为 Harness backend,移除 `@Tool` 和旧 recorder/session dedup;Agent 只看 ACI callbacks。
+- 新 Agent audit hook 只写角色/数量/Tool 名称/耗时等 metadata,不写模型输入正文、输出正文、Thought 或 Tool arguments。
+- Tool durable audit 通过 Harness port + JPA adapter fail-open 写入;Redis canonical store failure 仍 fail-closed,DB audit failure 只记录日志,不改变 Tool observation。
+- 删除 endpoint 的迁移无兼容层;bundled frontend 同 commit 删除按钮/consumer,回滚整体回滚 commit。
+- 不创建 ADR:方向已由 ISS-014 冻结,本 change 只完成最终落地与验收。
+
+## OpenSpec Backfill
+
+- 需进入 design/spec/tasks:删除清单、唯一 endpoint、Harness-native trace、Tool durable audit schema/safety、Tool backend 解耦、文档/issue 收口、strict validation、live E2E/log/DB acceptance 与回滚。
+
+## Cross-artifact Alignment
+
+| 上游 -> 下游 | 检查内容 | 状态 |
+|---|---|---|
+| ISS-014/brief -> proposal | 阶段 7 物理清理、文档、最终 Maven/log/DB E2E、Mock 外部 Tool 边界 | 已对齐 |
+| proposal -> design | 删除闭包、Tool backend 复用、安全 Agent/Tool audit、L4 migration/rollback | 已对齐 |
+| design -> specs/tasks | ownership、禁止泄漏、exact identity、strict validation、live E2E 均有 requirement 与切片 | 已对齐 |
+| specs -> tasks | 8 组可观察 requirements 覆盖删除、audit、docs、verification 和最终 E2E | 已对齐 |
+
+## Architecture Audit
+
+- Capability source: `zoom-out`,使用 Diagnosis Harness、Diagnosis Agent、Canonical Invocation、Durable Audit、Diagnosis Trace 和 Chat SSE Contract 术语。
+- 最终链路为 `browser -> ChatController -> ChatApplicationUseCase -> Harness -> Diagnosis Agent -> ACI Tools -> Guards -> SSE`,没有第二业务入口或业务 Graph。
+- Redis canonical invocation 是短期完整 Tool 真理源;MySQL ToolInvocation 是长期有界 metadata audit,两者禁止双写 raw payload。
+- `chat_session` JPA entity 属于当前 Run/PreviousTurn 目录,Redis SessionContext 属于旧 conversation memory;删除时必须区分。
+- 最大风险是 backend 的旧 Tool annotation/recorder 隐式暴露和 audit 内容泄漏;design/tasks 已加入独占 discovery、negative serialization、context startup 与 live DB inspection。
+
+## Interface Impact
+
+- Level: L4 breaking HTTP/frontend contract。
+- 删除 `/api/ai_ops`、`/api/chat/clear`、`/api/chat/session/{sessionId}` 和 `/runs`;保留唯一 `/api/chat` SSE 与 Trace API。
+- bundled frontend 同 commit 删除 AiOps 按钮/consumer;外部 caller 迁移到 `/api/chat`,无兼容 branch。
+- 回滚必须整体回滚阶段 7 commit,不能单独恢复旧 endpoint/Tool annotations。
+
+## Apply Verification Status
+
+- Follow-up current-document audit found and corrected stale runtime semantics in the tracked payment-timeout PowerShell demo and `mvp/tables`: old JSON Chat parsing, Redis SessionContext history, Planner/Verifier identities, full Tool payload persistence, and Chat/AiOps Run descriptions are no longer presented as current behavior.
+- `mvp/demo/scripts/run-payment-timeout-demo.ps1` now strictly validates `metadata -> status* -> content|failure -> done`, captures exact session/run identity, and fetches exact Trace without writing feedback. PowerShell parser validation and two in-memory SSE contract samples passed.
+- Current architecture/demo/table scan excluding explicit `archive/` and ignored local `output/` artifacts: 0 legacy runtime matches.
+- Follow-up `openspec validate --all --strict`: 22 passed, 0 failed; `git diff --check`: passed.
+- Initial environment checks reported missing injected variables. Per user direction, credentials were restored to ignored `application-local.yml`; no credential was added to Git or emitted in the archive.
+- `mvn -q -DskipTests compile`: passed.
+- `node --check src/main/resources/static/app.js`: passed.
+- `openspec validate --all --strict`: 22 passed, 0 failed.
+- Deterministic regression suite excluding credential-dependent MySQL/Redis/Milvus tests: 60 suites, 221 tests, 0 failures, 0 errors, 3 skipped.
+- `SemanticGuardTest` interruption assertion failed once under the credential-dependent full-suite run, then passed three isolated repetitions and the deterministic regression suite; classified as load-sensitive test timing, not a reproduced product regression.
+- `mvn -q -DskipTests package`: passed.
+- Production legacy path scan, audit sensitive-payload scan, and legacy backend Tool contract scan: 0 matches.
+- The approved external-network run connected to MySQL/Redis/Milvus/model services. Flyway repair aligned the old V012 checksum and V013 reconciled the missing release-contract columns/index.
+- Maven startup validated all 13 migrations, passed JPA schema validation and started Tomcat 9900 with Harness/Audit beans.
+- Final exact live run: session `mvp-demo-payment-timeout-stage7-20260722-1741`, run `363f481c-33b8-42e7-8699-428a6ec61806`, SSE `metadata -> status -> status -> status -> content -> done`, outcome `FALLBACK`.
+- Exact MySQL evidence: one `DIAGNOSIS/SUCCESS/FALLBACK` run, two metadata-only `diagnosis_agent` steps with null Thought, and two READY/EVIDENCE_FOUND Tool audits for `lookup_knowledge` and Mock `query_logs`.
+- Tasks 4.2-4.5 are complete. Real CLS and production business MySQL Tool datasource remain explicit non-goals.
+
+## Apply Conflict Classification
+
+- **Code deviation**: `DiagnosisRun` enum mapping expected native ENUM while V012/V013 define VARCHAR. Fixed ORM column definitions; OpenSpec unchanged.
+- **Code deviation**: production `ObjectMapper` lacked Java Time modules, causing canonical Redis `STORE_ERROR`. Fixed mapper registration and bound the store test to the production mapper.
+- **Code deviation**: Mock empty log results used `success=false`, conflicting with the specified `NO_EVIDENCE` projection. Fixed backend success semantics and the stale test assertion.
+- **Code deviation**: RAG backend logs exposed raw query/rewritten query content. Replaced with bounded counts/category metadata and verified the final Tool execution window contains no prohibited payload.
+- **Acceptance fixture drift**: the payment-timeout demo requested the removed metrics Tool and encouraged an unbounded investigation. Updated the fixture to the current two-Tool Mock acceptance scope and explicit stopping boundary.
+
+## Commit Gate Preflight
+
+- proposal、design、两份 specs 和 tasks 完整;新 capability 与 modified capability 的范围无 gap。
+- Question pool 全部 evidence-driven 并已汇报,无 user-interview、未判级接口或未接受架构风险。
+- 删除清单区分 current JPA metadata 与 legacy Redis Session,Tool backend 与 Agent-facing contract,canonical truth 与 durable audit。
+- live E2E 明确要求真实应用/模型链和 exact ID;真实 CLS/生产业务 MySQL 明确不在验收声称范围。
diff --git a/devflow/projects/2026-07-22-single-react-cleanup-e2e/evidence.md b/devflow/projects/2026-07-22-single-react-cleanup-e2e/evidence.md
new file mode 100644
index 0000000..54be48d
--- /dev/null
+++ b/devflow/projects/2026-07-22-single-react-cleanup-e2e/evidence.md
@@ -0,0 +1,26 @@
+# Evidence: single-react-cleanup-e2e
+
+## Repository Evidence
+
+- 生产引用扫描确认旧 `ChatService` 仅剩自身测试,`/api/ai_ops` 与旧 Session endpoints 属于第二条公开/状态链。
+- Harness adapters 复用 `LookupKnowledgeTool`/`QueryLogsTools` backend;旧 `@Tool`、ThreadLocal、recorder 和 topic discovery contract 已移除。
+- `HarnessAgentAuditHook` 只写 message count/roles、Tool names、text presence 和 duration;`thought` 保持空。
+- ToolBoundary durable audit 只写 exact identity、状态、稳定错误码、耗时和字节数;Redis canonical invocation 仍是短期完整 Tool 真理源。
+
+## Live Findings
+
+- 远端库已登记旧 V012 checksum,但缺少 `diagnosis_run.intent/release_outcome/published_result`。Flyway repair 后,幂等 V013 补齐三列与索引,fresh database 上为 no-op。
+- Hibernate 6 将无 `columnDefinition` 的字符串枚举校验为原生 ENUM;`DiagnosisRun` 已显式映射到 V012/V013 的 `VARCHAR(32/16)`。
+- 生产 `WebConfig` 的裸 `ObjectMapper` 无法序列化 canonical record 的 `Instant`,导致所有 Tool 在 begin 阶段返回 `STORE_ERROR`;改为自动注册模块,并让 store 测试使用生产 mapper。
+- Mock `query_logs` 把 0 命中错误表达为 `success=false`,与 `NO_EVIDENCE` contract 冲突;现以成功查询 + 空数组表达无证据。
+- live 日志发现 RAG backend 打印原始 query/rewrittenQuery/keywords;已改为字符数、命中数、类别数和耗时 metadata。
+
+## Final Exact-run Evidence
+
+- sessionId: `mvp-demo-payment-timeout-stage7-20260722-1741`
+- runId: `363f481c-33b8-42e7-8699-428a6ec61806`
+- SSE: `metadata -> status -> status -> status -> content -> done`
+- outcome: `FALLBACK`; diagnosis_run: `DIAGNOSIS / SUCCESS / FALLBACK`
+- AgentStep: 2 rows,均为 `diagnosis_agent`,`thought IS NULL`,model input/output 仅 metadata。
+- ToolInvocation: 2 rows,`lookup_knowledge` 与 `query_logs` 均为 `READY/EVIDENCE_FOUND`,同一 exact identity,无错误。
+- `query_logs` 明确为 Mock;Agent-facing `query_mysql` 未配置生产业务 datasource,也未声称 live。
diff --git a/mvp/architecture/README.md b/mvp/architecture/README.md
index 568e2a8..8ac9e68 100644
--- a/mvp/architecture/README.md
+++ b/mvp/architecture/README.md
@@ -1,48 +1,15 @@
# MVP 架构文档
-**更新日期**:2026-07-10
+**更新日期**:2026-07-22
+**状态**:当前单 Diagnosis Agent + Harness 架构
-这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
+当前文档入口:
-- `mvp/architecture/archive/2026-07-05-legacy/`
-
-归档材料只作为设计历史阅读,不再作为当前实现依据。
-
-## 当前文档
-
-| 文档 | 用途 |
+| 文档 | 内容 |
|---|---|
-| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
-| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
-| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
-| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
-| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
-| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
-| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
-| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
-| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
-| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
-| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
-| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库文档编写与维护规范,覆盖 frontmatter、category、chunk、reindex |
-| [data-model.md](data-model.md) | 数据模型总览,覆盖 Trace、知识库、反馈沉淀和 Milvus metadata |
-| [evolution-roadmap.md](evolution-roadmap.md) | 从旧版 Agent 蓝图继承的后续演进路线,不代表当前已实现 |
+| [current-mvp-architecture.md](current-mvp-architecture.md) | 系统分层、请求主链与 API surface |
+| [agent-orchestration.md](agent-orchestration.md) | 单 Diagnosis ReAct Agent 的职责和执行方式 |
+| [harness-quality-gates.md](harness-quality-gates.md) | Run、Tool、Evidence、Semantic 与 Release 门禁 |
+| [session-trace-lifecycle.md](session-trace-lifecycle.md) | sessionId/runId、SSE 和 durable audit 生命周期 |
-## 当前架构一句话
-
-SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
-
-## 阅读顺序
-
-1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
-2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
-3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
-4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
-5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
-6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
-7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
-8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
-9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
-10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
-11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
-12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
-13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
+2026-07-22 前的多角色编排、双入口和旧证据链文档已移动到 `archive/2026-07-22-legacy/`,仅用于历史决策追溯,不代表当前运行时。
diff --git a/mvp/architecture/agent-orchestration.md b/mvp/architecture/agent-orchestration.md
index 602e05a..fb9eed1 100644
--- a/mvp/architecture/agent-orchestration.md
+++ b/mvp/architecture/agent-orchestration.md
@@ -1,237 +1,54 @@
-# Agent 编排架构
+# Diagnosis Agent 执行架构
-**更新日期**:2026-07-08
+**更新日期**:2026-07-22
**状态**:当前可运行架构
-**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
-## 1. 设计定位
+## 1. 单 Agent 原则
-旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
+当前业务诊断只有一个 `Diagnosis Agent`。它使用框架 `ReactAgent` 完成规划、行动、观察和最终 Draft,但项目不在外层复制 ReAct 状态机,也不使用业务 Graph 或多角色协作链。
-- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
-- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
-- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
-- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
+## 2. 职责
-## 2. 当前 Agent 全景
+Diagnosis Agent:
-```mermaid
-flowchart TB
- subgraph Chat["Chat diagnosis"]
- ChatIn["POST /api/chat"] --> ChatService["ChatService"]
- ChatService --> ChatPlanner["chat_planner"]
- ChatPlanner --> ChatExecutor["chat_executor"]
- ChatExecutor --> ChatTools["evidence tools"]
- ChatTools --> ChatExecutor
- ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
- ChatGatekeeper --> ChatVerifier["chat_verifier"]
- ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
- ChatDecision --> ChatComposer["chat_composer"]
- ChatComposer --> ChatAnswer["final answer"]
- end
+- 接收当前 query 与可选、受限的安全 PreviousTurn。
+- 自主选择只读 evidence Tool。
+- 根据 Agent projection 判断是否需要继续查询。
+- 输出结构化 `DiagnosisDraft`,每条 analysis 绑定 framework `tool_call_id`。
+- 证据不足时明确限制,不补造事实。
- subgraph AiOps["AIOps diagnosis"]
- AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
- AiOpsService --> Supervisor["ai_ops_supervisor"]
- Supervisor --> AiOpsPlanner["planner_agent"]
- Supervisor --> AiOpsExecutor["executor_agent"]
- AiOpsPlanner --> AiOpsExecutor
- AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
- AiOpsTools --> AiOpsReport["alert report"]
- AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
- end
+Diagnosis Agent 不负责:
- subgraph Trace["Trace persistence"]
- ChatSession["chat_session"]
- Run["diagnosis_run"]
- Step["agent_step"]
- Invocation["tool_invocation"]
- SelfEval["self_evaluation"]
- end
+- HTTP/SSE、Session/Run 生命周期和持久化。
+- 模型/Tool/Token/timeout/cancel 预算。
+- Tool 参数授权、raw response 投影或证据物理验真。
+- SemanticGuard 与最终发布决定。
- ChatService --> ChatSession
- ChatService --> Run
- ChatPlanner --> Step
- ChatExecutor --> Step
- ChatGatekeeper --> SelfEval
- ChatVerifier --> Step
- ChatTools --> Invocation
- ChatDecision --> SelfEval
- ChatComposer --> Step
-
- AiOpsService --> ChatSession
- AiOpsService --> Run
- AiOpsPlanner --> Step
- AiOpsExecutor --> Step
- AiOpsTools --> Invocation
- AiOpsRule --> SelfEval
-```
-
-## 3. Chat 编排
-
-Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
-
-```text
-chat_planner
- -> chat_executor
- -> lookup_knowledge / query_logs / query_metrics / date_time
- -> outputs executor_evidence_v2
- -> VerifierInputHook / ExecutorGatekeeperService
- -> validates source_invocation_id / raw_path / evidence_excerpt
- -> chat_verifier
- -> judges whether verified evidence can derive claims
- -> chat_composer
- -> writes final user-facing answer
-```
-
-关键行为:
-
-| 角色 | 当前职责 | 输出 |
-|---|---|---|
-| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
-| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
-| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
-| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
-| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
-
-Chat 链路最多支持两轮验证:
+## 3. 执行序列
```mermaid
sequenceDiagram
- autonumber
- participant C as ChatService
- participant P as chat_planner
- participant E as chat_executor
- participant T as tools
- participant G as gatekeeper
- participant V as chat_verifier
- participant M as chat_composer
- participant R as diagnosis_run
+ participant App as Chat Application
+ participant Core as Harness Core
+ participant Agent as Diagnosis Agent
+ participant Tool as ACI Tool Boundary
+ participant EG as EvidenceGuard
+ participant SG as SemanticGuard
+ participant Release as Release Policy
- C->>P: 原始问题 + history + retry_context
- P-->>C: planner_plan
- C->>E: planner_plan + 上下文
- E->>T: 调用证据工具
- T-->>E: 证据结果
- E-->>C: executor_evidence_v2
- C->>G: executor_structured_output + tool_invocation.evidence_refs
- G-->>C: gatekeeper_result
- C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
- V-->>C: PASS / LOW_CONFID / REJECT
- C->>R: 写入 verifier_evaluation
- alt LOW_CONFID 且允许补证据
- C->>P: retry_context: 仅补缺失证据
- else PASS 或 REJECT
- C->>M: allowed_claims + missing_info + recommended_actions
- M-->>C: composer_output
- C->>R: 保存 Composer 最终 answer
- end
+ App->>Core: start RunContext
+ App->>Agent: query + safe previous_turn
+ Agent->>Tool: tool name + framework tool_call_id + typed args
+ Tool-->>Agent: bounded agent_result
+ Agent-->>App: DiagnosisDraft
+ App->>EG: Draft + current Run canonical invocations
+ EG-->>App: verified snapshot or deterministic failure
+ App->>SG: query + full Draft + verified snapshot
+ SG-->>App: SUPPORTED / UNSUPPORTED
+ App->>Release: decide public content
+ Release-->>App: report or fixed fallback
```
-决策语义:
+## 4. PreviousTurn
-| Verdict | 行为 |
-|---|---|
-| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
-| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
-| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
-
-## 4. AIOps 编排
-
-AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
-
-```text
-ai_ops_supervisor
- -> planner_agent
- -> executor_agent
- -> final report
- -> AiOpsRuleEvaluationService
-```
-
-与 Chat 的差异:
-
-- AIOps 的输入可能是结构化告警 payload。
-- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
-- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
-- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
-
-## 5. 工具边界
-
-当前 Executor 可用工具来自两类:
-
-```text
-methodTools
- -> dateTimeTools
- -> lookupKnowledgeTool
- -> queryMetricsTools
- -> queryLogsTools when mock enabled
-
-ToolCallbackProvider
- -> framework-discovered tools
-```
-
-工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
-
-- L0/L1 命中数量。
-- 检索层。
-- relevance level。
-- retrieved domains。
-- dedup reason。
-
-## 6. Skill / Playbook 流程
-
-当前 Skill 是诊断流程编排提示,不是事实证据来源。Planner 只能看到 `SkillRegistry.listAll()` 暴露的 name/description 元数据;Executor 才能通过 Spring AI Alibaba 官方 `SkillsAgentHook` 使用 `read_skill` 读取完整 `SKILL.md`。
-
-```mermaid
-flowchart LR
- Registry["SkillRegistry
active skill metadata"] --> PlannerHook["PlannerSkillMetadataHook"]
- PlannerHook --> Planner["Planner
metadata only"]
- Planner --> Plan["planner_plan
selected_skill + steps"]
-
- Registry --> ExecutorHook["SkillsAgentHook"]
- ExecutorHook --> ReadSkill["read_skill"]
- Plan --> Executor["Executor"]
- Executor --> ReadSkill
- ReadSkill --> SkillBody["SKILL.md workflow"]
- SkillBody --> Executor
- Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
- EvidenceTools --> ToolTrace["tool_invocation evidence"]
- Executor --> Gatekeeper["Gatekeeper"]
- Gatekeeper --> Verifier["Verifier"]
- ToolTrace --> Verifier
- Verifier --> Composer["Composer"]
-```
-
-| 角色 | Skill 可见性 | 工具权限 |
-|---|---|---|
-| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
-| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
-| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
-| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
-| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
-
-## 7. 与旧版设计的差异
-
-| 旧版设想 | 当前实现 |
-|---|---|
-| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
-| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
-| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
-| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
-| Skill 驱动不同诊断流程 | 当前以 Planner 元数据选择 + Executor 读取 playbook 的方式接入 |
-
-## 8. 后续演进
-
-当诊断场景和工具复杂度继续上升时,再考虑拆分:
-
-- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
-- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
-- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
-- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
-
-拆分前提:
-
-- 当前 Executor prompt 已难以维护。
-- 不同故障类型的工具权限明显不同。
-- Trace 能证明某类问题需要独立的推理策略。
-- 评测集能覆盖拆分前后的行为差异。
+PreviousTurn 只来自同一 Session 最近一个 `DIAGNOSIS + SUCCESS + published_result`。Fallback、失败、取消、raw evidence 和完整历史都不能进入下一轮;字段与字节上限由 Harness 配置控制。
diff --git a/mvp/architecture/archive/2026-07-22-legacy/README.md b/mvp/architecture/archive/2026-07-22-legacy/README.md
new file mode 100644
index 0000000..568e2a8
--- /dev/null
+++ b/mvp/architecture/archive/2026-07-22-legacy/README.md
@@ -0,0 +1,48 @@
+# MVP 架构文档
+
+**更新日期**:2026-07-10
+
+这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
+
+- `mvp/architecture/archive/2026-07-05-legacy/`
+
+归档材料只作为设计历史阅读,不再作为当前实现依据。
+
+## 当前文档
+
+| 文档 | 用途 |
+|---|---|
+| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
+| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
+| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
+| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
+| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
+| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
+| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
+| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
+| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
+| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
+| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
+| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库文档编写与维护规范,覆盖 frontmatter、category、chunk、reindex |
+| [data-model.md](data-model.md) | 数据模型总览,覆盖 Trace、知识库、反馈沉淀和 Milvus metadata |
+| [evolution-roadmap.md](evolution-roadmap.md) | 从旧版 Agent 蓝图继承的后续演进路线,不代表当前已实现 |
+
+## 当前架构一句话
+
+SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
+
+## 阅读顺序
+
+1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
+2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
+3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
+4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
+5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
+6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
+7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
+8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
+9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
+10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
+11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
+12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
+13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
diff --git a/mvp/architecture/archive/2026-07-22-legacy/_ARCHIVE_NOTE.md b/mvp/architecture/archive/2026-07-22-legacy/_ARCHIVE_NOTE.md
new file mode 100644
index 0000000..cc09e4d
--- /dev/null
+++ b/mvp/architecture/archive/2026-07-22-legacy/_ARCHIVE_NOTE.md
@@ -0,0 +1,3 @@
+# Archive Note
+
+本目录保存 2026-07-22 单 Diagnosis Agent + Harness 切换前的当前架构文档。内容用于历史决策追溯,不代表现行 runtime、API 或验收口径。
diff --git a/mvp/architecture/archive/2026-07-22-legacy/agent-orchestration.md b/mvp/architecture/archive/2026-07-22-legacy/agent-orchestration.md
new file mode 100644
index 0000000..602e05a
--- /dev/null
+++ b/mvp/architecture/archive/2026-07-22-legacy/agent-orchestration.md
@@ -0,0 +1,237 @@
+# Agent 编排架构
+
+**更新日期**:2026-07-08
+**状态**:当前可运行架构
+**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
+
+## 1. 设计定位
+
+旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
+
+- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
+- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
+- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
+- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
+
+## 2. 当前 Agent 全景
+
+```mermaid
+flowchart TB
+ subgraph Chat["Chat diagnosis"]
+ ChatIn["POST /api/chat"] --> ChatService["ChatService"]
+ ChatService --> ChatPlanner["chat_planner"]
+ ChatPlanner --> ChatExecutor["chat_executor"]
+ ChatExecutor --> ChatTools["evidence tools"]
+ ChatTools --> ChatExecutor
+ ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
+ ChatGatekeeper --> ChatVerifier["chat_verifier"]
+ ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
+ ChatDecision --> ChatComposer["chat_composer"]
+ ChatComposer --> ChatAnswer["final answer"]
+ end
+
+ subgraph AiOps["AIOps diagnosis"]
+ AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
+ AiOpsService --> Supervisor["ai_ops_supervisor"]
+ Supervisor --> AiOpsPlanner["planner_agent"]
+ Supervisor --> AiOpsExecutor["executor_agent"]
+ AiOpsPlanner --> AiOpsExecutor
+ AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
+ AiOpsTools --> AiOpsReport["alert report"]
+ AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
+ end
+
+ subgraph Trace["Trace persistence"]
+ ChatSession["chat_session"]
+ Run["diagnosis_run"]
+ Step["agent_step"]
+ Invocation["tool_invocation"]
+ SelfEval["self_evaluation"]
+ end
+
+ ChatService --> ChatSession
+ ChatService --> Run
+ ChatPlanner --> Step
+ ChatExecutor --> Step
+ ChatGatekeeper --> SelfEval
+ ChatVerifier --> Step
+ ChatTools --> Invocation
+ ChatDecision --> SelfEval
+ ChatComposer --> Step
+
+ AiOpsService --> ChatSession
+ AiOpsService --> Run
+ AiOpsPlanner --> Step
+ AiOpsExecutor --> Step
+ AiOpsTools --> Invocation
+ AiOpsRule --> SelfEval
+```
+
+## 3. Chat 编排
+
+Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
+
+```text
+chat_planner
+ -> chat_executor
+ -> lookup_knowledge / query_logs / query_metrics / date_time
+ -> outputs executor_evidence_v2
+ -> VerifierInputHook / ExecutorGatekeeperService
+ -> validates source_invocation_id / raw_path / evidence_excerpt
+ -> chat_verifier
+ -> judges whether verified evidence can derive claims
+ -> chat_composer
+ -> writes final user-facing answer
+```
+
+关键行为:
+
+| 角色 | 当前职责 | 输出 |
+|---|---|---|
+| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
+| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
+| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
+| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
+| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
+
+Chat 链路最多支持两轮验证:
+
+```mermaid
+sequenceDiagram
+ autonumber
+ participant C as ChatService
+ participant P as chat_planner
+ participant E as chat_executor
+ participant T as tools
+ participant G as gatekeeper
+ participant V as chat_verifier
+ participant M as chat_composer
+ participant R as diagnosis_run
+
+ C->>P: 原始问题 + history + retry_context
+ P-->>C: planner_plan
+ C->>E: planner_plan + 上下文
+ E->>T: 调用证据工具
+ T-->>E: 证据结果
+ E-->>C: executor_evidence_v2
+ C->>G: executor_structured_output + tool_invocation.evidence_refs
+ G-->>C: gatekeeper_result
+ C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
+ V-->>C: PASS / LOW_CONFID / REJECT
+ C->>R: 写入 verifier_evaluation
+ alt LOW_CONFID 且允许补证据
+ C->>P: retry_context: 仅补缺失证据
+ else PASS 或 REJECT
+ C->>M: allowed_claims + missing_info + recommended_actions
+ M-->>C: composer_output
+ C->>R: 保存 Composer 最终 answer
+ end
+```
+
+决策语义:
+
+| Verdict | 行为 |
+|---|---|
+| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
+| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
+| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
+
+## 4. AIOps 编排
+
+AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
+
+```text
+ai_ops_supervisor
+ -> planner_agent
+ -> executor_agent
+ -> final report
+ -> AiOpsRuleEvaluationService
+```
+
+与 Chat 的差异:
+
+- AIOps 的输入可能是结构化告警 payload。
+- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
+- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
+- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
+
+## 5. 工具边界
+
+当前 Executor 可用工具来自两类:
+
+```text
+methodTools
+ -> dateTimeTools
+ -> lookupKnowledgeTool
+ -> queryMetricsTools
+ -> queryLogsTools when mock enabled
+
+ToolCallbackProvider
+ -> framework-discovered tools
+```
+
+工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
+
+- L0/L1 命中数量。
+- 检索层。
+- relevance level。
+- retrieved domains。
+- dedup reason。
+
+## 6. Skill / Playbook 流程
+
+当前 Skill 是诊断流程编排提示,不是事实证据来源。Planner 只能看到 `SkillRegistry.listAll()` 暴露的 name/description 元数据;Executor 才能通过 Spring AI Alibaba 官方 `SkillsAgentHook` 使用 `read_skill` 读取完整 `SKILL.md`。
+
+```mermaid
+flowchart LR
+ Registry["SkillRegistry
active skill metadata"] --> PlannerHook["PlannerSkillMetadataHook"]
+ PlannerHook --> Planner["Planner
metadata only"]
+ Planner --> Plan["planner_plan
selected_skill + steps"]
+
+ Registry --> ExecutorHook["SkillsAgentHook"]
+ ExecutorHook --> ReadSkill["read_skill"]
+ Plan --> Executor["Executor"]
+ Executor --> ReadSkill
+ ReadSkill --> SkillBody["SKILL.md workflow"]
+ SkillBody --> Executor
+ Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
+ EvidenceTools --> ToolTrace["tool_invocation evidence"]
+ Executor --> Gatekeeper["Gatekeeper"]
+ Gatekeeper --> Verifier["Verifier"]
+ ToolTrace --> Verifier
+ Verifier --> Composer["Composer"]
+```
+
+| 角色 | Skill 可见性 | 工具权限 |
+|---|---|---|
+| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
+| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
+| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
+| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
+| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
+
+## 7. 与旧版设计的差异
+
+| 旧版设想 | 当前实现 |
+|---|---|
+| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
+| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
+| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
+| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
+| Skill 驱动不同诊断流程 | 当前以 Planner 元数据选择 + Executor 读取 playbook 的方式接入 |
+
+## 8. 后续演进
+
+当诊断场景和工具复杂度继续上升时,再考虑拆分:
+
+- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
+- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
+- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
+- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
+
+拆分前提:
+
+- 当前 Executor prompt 已难以维护。
+- 不同故障类型的工具权限明显不同。
+- Trace 能证明某类问题需要独立的推理策略。
+- 评测集能覆盖拆分前后的行为差异。
diff --git a/mvp/architecture/archive/2026-07-22-legacy/current-mvp-architecture.md b/mvp/architecture/archive/2026-07-22-legacy/current-mvp-architecture.md
new file mode 100644
index 0000000..a4bebb8
--- /dev/null
+++ b/mvp/architecture/archive/2026-07-22-legacy/current-mvp-architecture.md
@@ -0,0 +1,444 @@
+# 当前 MVP 架构
+
+**更新日期**:2026-07-08
+**状态**:当前可运行架构
+**适用范围**:Demo、面试讲解、后续迭代规划
+
+## 1. 系统定位
+
+SuperBizAgent MVP 不是通用 Chatbot,而是面向故障诊断的 Agent 工程项目。
+
+核心目标:
+
+- 支持用户主动发起的 Chat 诊断。
+- 支持 AIOps 告警触发的自动诊断。
+- 保留 Agent 的规划、执行、验证过程。
+- 工具调用必须显式、可追踪、可回放。
+- RAG 检索必须通过 `lookup_knowledge` 暴露证据链。
+- 每次诊断都沉淀 session、step、tool invocation 和 self evaluation。
+
+## 2. 总体分层
+
+```mermaid
+flowchart TB
+ subgraph API["API Layer"]
+ ChatController["ChatController"]
+ TraceController["DiagnosisTraceController"]
+ SearchController["SearchController"]
+ DocumentController["DocumentController"]
+ end
+
+ subgraph App["Application Service"]
+ ChatService["ChatService"]
+ AiOpsService["AiOpsService"]
+ TraceService["DiagnosisTraceService"]
+ end
+
+ subgraph Agent["Agent Orchestration"]
+ Supervisor["Supervisor"]
+ Planner["Planner"]
+ Executor["Executor"]
+ Gatekeeper["Gatekeeper"]
+ Verifier["Verifier"]
+ Composer["Composer"]
+ end
+
+ subgraph Tools["Evidence Tools"]
+ KnowledgeTool["lookup_knowledge"]
+ LogsTool["query_logs"]
+ MetricsTool["query_metrics"]
+ AlertsTool["queryPrometheusAlerts"]
+ end
+
+ subgraph Skills["Skill / Playbook"]
+ SkillRegistry["SkillRegistry"]
+ PlannerSkillHook["PlannerSkillMetadataHook"]
+ SkillsHook["SkillsAgentHook"]
+ ReadSkill["read_skill"]
+ end
+
+ subgraph RAG["RAG Retrieval"]
+ L0["KnowledgeIndexService"]
+ VectorSearch["VectorSearchService"]
+ VectorStore["Spring AI VectorStore"]
+ SdkFallback["Milvus SDK fallback"]
+ end
+
+ subgraph Store["Persistence and Trace"]
+ ChatSession["chat_session"]
+ Run["diagnosis_run"]
+ Step["agent_step"]
+ Invocation["tool_invocation"]
+ ApiDoc["api_document"]
+ Milvus["Milvus/Zilliz"]
+ end
+
+ API --> App
+ ChatService --> Agent
+ AiOpsService --> Agent
+ SkillRegistry --> PlannerSkillHook
+ PlannerSkillHook --> Planner
+ SkillRegistry --> SkillsHook
+ SkillsHook --> Executor
+ Executor --> ReadSkill
+ Agent --> Tools
+ KnowledgeTool --> RAG
+ RAG --> Store
+ Tools --> Invocation
+ Agent --> Step
+ App --> Session
+ TraceService --> Session
+ TraceService --> Step
+ TraceService --> Invocation
+```
+
+```text
+API Layer
+ -> ChatController
+ -> DiagnosisTraceController
+ -> SearchController
+ -> DocumentController
+
+Application Service
+ -> ChatService
+ -> AiOpsService
+ -> DiagnosisTraceService
+
+Agent Orchestration
+ -> Supervisor
+ -> Planner
+ -> Executor
+ -> Gatekeeper
+ -> Verifier
+ -> Composer
+
+Evidence Tools
+ -> lookup_knowledge
+ -> query_logs
+ -> query_metrics
+ -> queryPrometheusAlerts
+
+Skill / Playbook
+ -> SkillRegistry
+ -> PlannerSkillMetadataHook gives Planner name/description only
+ -> SkillsAgentHook gives Executor read_skill
+ -> Verifier is isolated from skills
+
+RAG Retrieval
+ -> KnowledgeIndexService
+ -> VectorSearchService
+ -> Spring AI VectorStore
+ -> Milvus SDK fallback
+
+Persistence
+ -> chat_session
+ -> diagnosis_run
+ -> agent_step.run_id
+ -> tool_invocation.run_id
+ -> api_document
+ -> Milvus/Zilliz collection
+
+Quality Gates
+ -> executor gatekeeper
+ -> chat verifier
+ -> AIOps rule evaluation
+ -> diagnosis eval baseline
+ -> RAG retrieval baseline
+```
+
+## 3. Chat 诊断链路
+
+```mermaid
+sequenceDiagram
+ autonumber
+ actor User as 用户
+ participant API as POST /api/chat
+ participant Chat as ChatService
+ participant Planner as Planner Agent
+ participant Executor as Executor Agent
+ participant Tool as Evidence Tools
+ participant Gatekeeper as Gatekeeper Hook
+ participant Verifier as Verifier Agent
+ participant Composer as Composer Agent
+ participant DB as Trace Tables
+ participant Trace as Trace API
+
+ User->>API: 提交诊断问题
+ API->>Chat: execute chat strategy
+ Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
+ Chat->>Planner: 复杂问题进入规划
+ Planner->>DB: 写入 agent_step.run_id
+ Planner->>Executor: 下发排查方向
+ Executor->>Tool: lookup_knowledge / logs / metrics
+ Tool->>DB: 写入 tool_invocation.run_id
+ Tool-->>Executor: 返回证据
+ Executor->>Gatekeeper: 输出 executor_evidence_v2
+ Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
+ Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
+ Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
+ Verifier->>Composer: 传入 allowed_claims / missing_info / actions
+ Composer->>Chat: 生成最终用户答复
+ Chat->>DB: 保存 diagnosis_run.answer
+ User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
+ Trace->>DB: 聚合 run / step / tool
+ Trace-->>User: 返回可回放诊断链路
+```
+
+```text
+POST /api/chat
+ -> ChatService
+ -> 简单问题:轻量回答
+ -> 复杂诊断:Agent 编排
+ -> Planner 制定排查方向
+ -> Executor 调用证据工具
+ -> lookup_knowledge
+ -> query_logs
+ -> query_metrics
+ -> Gatekeeper 校验 Executor 证据引用真实性
+ -> Verifier 判断 claim 是否能由已核验证据推出
+ -> Composer 生成最终用户答复
+ -> 保存 chat_session metadata
+ -> 保存 diagnosis_run
+ -> 保存 agent_step.run_id
+ -> 保存 tool_invocation.run_id
+ -> 合并 diagnosis_run.self_evaluation.verifier_evaluation
+```
+
+Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
+
+Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
+
+关键代码:
+
+- `src/main/java/com/superbiz/agent/controller/ChatController.java`
+- `src/main/java/com/superbiz/agent/service/ChatService.java`
+- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
+- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
+
+## 4. AIOps 诊断链路
+
+```mermaid
+flowchart TD
+ Request["POST /api/ai_ops"] --> Payload{"包含告警 payload?"}
+ Payload -->|是| Targeted["PAYLOAD_TARGETED"]
+ Payload -->|否| Discovery["AUTO_DISCOVERY"]
+
+ Targeted --> BuildPrompt["构造聚焦 payload 的诊断 prompt"]
+ Targeted --> QueryAug["生成 recommended lookup_knowledge query"]
+ Discovery --> DiscoverAlert["通过 queryPrometheusAlerts 发现活跃告警"]
+
+ BuildPrompt --> Plan["Planner 规划排查"]
+ QueryAug --> Plan
+ DiscoverAlert --> Plan
+
+ Plan --> Execute["Executor 收集证据"]
+ Execute --> Knowledge["lookup_knowledge"]
+ Execute --> Metrics["query_metrics / Prometheus"]
+ Execute --> Logs["query_logs"]
+
+ Knowledge --> Report["告警分析报告"]
+ Metrics --> Report
+ Logs --> Report
+
+ Report --> RuleEval["AiOpsRuleEvaluationService"]
+ RuleEval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
+ Report --> Trace["DiagnosisTraceService"]
+ SelfEval --> Trace
+```
+
+```text
+POST /api/ai_ops
+ -> AiOpsService
+ -> 判断是否有告警 payload
+ -> PAYLOAD_TARGETED
+ -> AUTO_DISCOVERY
+ -> 构造 AIOps 诊断 prompt
+ -> payload 模式补充 recommended lookup_knowledge query
+ -> Agent 编排
+ -> Planner / Executor
+ -> Prometheus / logs / knowledge tools
+ -> 生成告警分析报告
+ -> AiOpsRuleEvaluationService
+ -> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
+ -> Trace API 可查看全链路
+```
+
+AIOps 保留两种模式:
+
+| 模式 | 触发条件 | 行为 |
+|---|---|---|
+| `PAYLOAD_TARGETED` | 请求包含 alertName、service、severity、description、timeRange 等字段 | 以 payload 为唯一主诊断对象,并生成推荐知识库 query |
+| `AUTO_DISCOVERY` | 请求没有明确告警 payload | 先查询当前活跃告警,再选择目标排查 |
+
+AIOps 当前使用轻量规则验证器,重点检查:
+
+- 最终报告是否存在。
+- payload 模式是否聚焦输入告警。
+- 是否使用关键证据工具,例如 `lookup_knowledge`、日志、指标。
+
+关键代码:
+
+- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
+- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
+
+## 5. RAG 位置
+
+RAG 不是隐藏在 Chat Advisor 里的隐式能力,而是 Executor 可以显式调用的工具:
+
+```mermaid
+flowchart LR
+ Executor["Executor Agent"] --> Tool["lookup_knowledge Tool"]
+ Tool --> L0["L0 domain/entity hint"]
+ Tool --> Search["VectorSearchService"]
+ L0 --> Search
+ Search --> VectorStore["Spring AI VectorStore"]
+ Search --> Fallback["Milvus SDK fallback"]
+ VectorStore --> Normalize["score/rawScore/scoreLabel"]
+ Fallback --> Normalize
+ Normalize --> Evidence["evidence output"]
+ Evidence --> Invocation["tool_invocation"]
+ Evidence --> Executor
+```
+
+```text
+Executor
+ -> lookup_knowledge(query)
+ -> L0 domain/entity hint
+ -> VectorSearchService
+ -> Spring AI VectorStore
+ -> Milvus SDK fallback
+ -> evidence shaping
+ -> tool_invocation
+```
+
+保留显式工具的原因:
+
+- Agent 何时检索、检索什么、证据是什么,必须能在 trace 中解释。
+- AIOps payload 到 query 的业务映射需要项目内控制。
+- `tool_invocation` 是后续评测、回放和面试讲解的核心材料。
+
+RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细节见 [retrieval-observability.md](retrieval-observability.md)。
+
+## 6. 持久化模型
+
+当前诊断持久化以 session/run/trace 明细为核心:
+
+```text
+chat_session
+ -> 多轮会话目录和元数据
+ -> session_id / status / message_pair_count
+
+diagnosis_run
+ -> 一次诊断运行的主记录
+ -> run_id / session_id
+ -> query / status / agent_flow / answer
+ -> self_evaluation
+ -> step_count / tool_call_count / duration
+
+agent_step
+ -> Agent 模型调用步骤
+ -> session_id / run_id
+ -> step_index / agent_name
+ -> model_input / model_output / thought
+ -> duration / token_count
+
+tool_invocation
+ -> 工具调用事实
+ -> session_id / run_id
+ -> tool_name / input_params / output_preview
+ -> retrieval_layer / retrieval_details
+ -> retrieval_details.evidence_refs
+ -> relevance_level / dedup_reason
+ -> duration / success
+```
+
+说明:
+
+- 旧的 `diagnosis_record` 已不是当前主模型。
+- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
+- `api_document` 仍用于文档元数据管理。
+- 文档向量内容存放在 Milvus/Zilliz collection 中。
+
+会话和 Trace 生命周期见 [session-trace-lifecycle.md](session-trace-lifecycle.md),完整数据关系见 [data-model.md](data-model.md)。
+
+## 7. Trace API
+
+```text
+GET /api/diagnosis/{sessionId}/trace
+GET /api/diagnosis/{sessionId}/trace?runId=run-...
+```
+
+Trace API 聚合:
+
+- 会话元数据、运行状态和最终报告。
+- Agent step 序列。
+- 工具调用和检索细节。
+- Chat Gatekeeper / Verifier / Composer 结果。
+- AIOps rule evaluation 结果。
+
+Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
+
+Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
+
+## 8. 质量门禁
+
+当前质量门禁分层如下:
+
+| 门禁 | 位置 | 作用 |
+|---|---|---|
+| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
+| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
+| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
+| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
+| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
+| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
+| Live RAG Acceptance | `scripts/eval_rag_live_acceptance.py` | 在运行环境中验证重建索引后的真实检索 |
+
+## 9. 当前完成状态
+
+已经完成:
+
+- Chat 和 AIOps 两条入口链路。
+- 显式 `lookup_knowledge` Agent Tool。
+- L0 从最终决策降级为 domain/entity hint。
+- `VectorSearchService` 作为稳定检索门面。
+- Spring AI VectorStore 读取路径。
+- Milvus SDK fallback。
+- `score` / `rawScore` / `scoreLabel` 分数语义拆分。
+- `title`、`breadcrumb`、`content` 参与 embedding 文本。
+- `tool_invocation` 记录检索层、relevance level、dedup reason。
+- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
+- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
+- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
+- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
+- Verifier 只判断可推导性。
+- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
+- RAG offline baseline 和 live acceptance 脚本。
+
+暂不作为当前已完成能力声明:
+
+- 完整 QueryTransformer / MultiQuery。
+- BM25、RRF、cross-encoder rerank。
+- 完整邻居 chunk / section context expansion。
+- VectorStore 写入路径全面迁移。
+- 完整 LLM-based AIOps verifier。
+
+后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
+
+## 10. 关键代码索引
+
+| 能力 | 代码 |
+|---|---|
+| Chat 入口与编排 | `ChatController`, `ChatService` |
+| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
+| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
+| 知识库工具 | `LookupKnowledgeTool` |
+| L0 hint | `KnowledgeIndexService` |
+| 向量检索门面 | `VectorSearchService` |
+| 文档切片 | `DocumentChunkService` |
+| 向量写入 | `VectorIndexService` |
+| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
+| Trace 聚合 | `DiagnosisTraceService` |
+| 工具调用记录 | `ToolInvocationRecorder` |
+| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
+| self_evaluation 合并 | `SelfEvaluationMergeService` |
diff --git a/mvp/architecture/data-model.md b/mvp/architecture/archive/2026-07-22-legacy/data-model.md
similarity index 100%
rename from mvp/architecture/data-model.md
rename to mvp/architecture/archive/2026-07-22-legacy/data-model.md
diff --git a/mvp/architecture/evolution-roadmap.md b/mvp/architecture/archive/2026-07-22-legacy/evolution-roadmap.md
similarity index 100%
rename from mvp/architecture/evolution-roadmap.md
rename to mvp/architecture/archive/2026-07-22-legacy/evolution-roadmap.md
diff --git a/mvp/architecture/executor-evidence-pipeline-refactor.md b/mvp/architecture/archive/2026-07-22-legacy/executor-evidence-pipeline-refactor.md
similarity index 100%
rename from mvp/architecture/executor-evidence-pipeline-refactor.md
rename to mvp/architecture/archive/2026-07-22-legacy/executor-evidence-pipeline-refactor.md
diff --git a/mvp/architecture/feedback-architecture.md b/mvp/architecture/archive/2026-07-22-legacy/feedback-architecture.md
similarity index 100%
rename from mvp/architecture/feedback-architecture.md
rename to mvp/architecture/archive/2026-07-22-legacy/feedback-architecture.md
diff --git a/mvp/architecture/archive/2026-07-22-legacy/harness-quality-gates.md b/mvp/architecture/archive/2026-07-22-legacy/harness-quality-gates.md
new file mode 100644
index 0000000..b91d4e3
--- /dev/null
+++ b/mvp/architecture/archive/2026-07-22-legacy/harness-quality-gates.md
@@ -0,0 +1,258 @@
+# Harness 与质量门禁架构
+
+**更新日期**:2026-07-08
+**状态**:当前可运行架构 + 后续门禁规划
+**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
+
+## 1. 设计目标
+
+Agent 系统的核心风险不是“没有答案”,而是:
+
+- 答案引用了不存在的证据。
+- 工具调用失败后仍然编造结论。
+- 检索结果相关性不足但被当作强证据。
+- 多轮诊断重复检索同一文档,浪费上下文。
+- 最终报告无法回放执行过程。
+
+因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
+
+```text
+Prompt contract
+ + Tool boundary
+ + Agent hooks
+ + Trace persistence
+ + Gatekeeper deterministic validation
+ + Verifier / rule evaluation
+ + Eval baseline
+```
+
+## 2. Harness 总图
+
+```mermaid
+flowchart TB
+ Input["User / AIOps input"] --> Prompt["Prompt contract"]
+ Prompt --> Agent["Planner / Executor / Verifier / Composer"]
+ Agent --> Tools["Evidence tools"]
+ Tools --> Invocation["tool_invocation"]
+ Agent --> StepHook["AgentLoggingHook"]
+ StepHook --> Step["agent_step"]
+ Agent --> Run["diagnosis_run"]
+
+ Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
+ EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
+ Agent --> Gatekeeper
+ Invocation --> TraceSummary["ToolTraceSummaryService"]
+ Gatekeeper --> Verifier["chat_verifier"]
+ TraceSummary --> Verifier
+ Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
+
+ Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
+ AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
+
+ Run --> TraceAPI["DiagnosisTraceService"]
+ Step --> TraceAPI
+ Invocation --> TraceAPI
+ SelfEval --> TraceAPI
+ AiOpsEval --> TraceAPI
+
+ TraceAPI --> Eval["diagnosis eval / RAG eval"]
+```
+
+## 3. Prompt Contract
+
+当前 Prompt 按角色拆分:
+
+| Prompt | 用途 |
+|---|---|
+| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
+| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
+| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
+| `chat-planner-prompt.md` | Chat 复杂问题规划 |
+| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
+| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
+| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
+
+Prompt 层当前承担的门禁:
+
+- 禁止凭记忆回答错误码、接口定义、排障步骤。
+- 需要外部信息时必须调用工具。
+- 工具连续失败或返回空结果时,最终报告必须诚实说明。
+- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
+- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
+- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
+- AIOps payload 模式必须聚焦输入告警。
+
+Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
+
+```json
+{
+ "version": "chat-prompts-v1",
+ "prompts": [
+ {
+ "name": "chat_executor",
+ "version": "chat-executor-v2",
+ "resource": "prompts/chat-executor-prompt.md"
+ }
+ ]
+}
+```
+
+该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
+
+## 4. Trace Hooks
+
+`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
+
+```mermaid
+sequenceDiagram
+ autonumber
+ participant A as Agent
+ participant H as AgentLoggingHook
+ participant DB as agent_step
+
+ A->>H: before_model(messages, sessionId, runId)
+ H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
+ A-->>A: LLM 推理
+ A->>H: after_model(messages, sessionId, runId)
+ H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
+```
+
+记录内容:
+
+- 最近输入消息摘要。
+- Agent 输出摘要。
+- 是否包含 tool call。
+- duration。
+- token count。
+- Verifier 的 JSON 输出摘要。
+
+新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
+
+## 5. Tool Invocation 门禁
+
+工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
+
+核心记录:
+
+```text
+tool_name
+input_params
+output_preview
+retrieval_layer
+l0_match_count
+l1_match_count
+retrieval_details
+ -> evidence_refs
+relevance_level
+dedup_reason
+duration_ms
+success
+error_message
+```
+
+对 `lookup_knowledge` 的质量约束:
+
+- L0 只作为 hint,不绕过 L1。
+- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
+- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
+- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
+- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
+- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
+
+## 6. Gatekeeper 与 Verifier 门禁
+
+Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
+
+```mermaid
+flowchart LR
+ Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
+ ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
+ EvidenceRefs --> Gatekeeper
+ Gatekeeper --> GateResult["gatekeeper_result"]
+ Invocation --> Summary["ToolTraceSummaryService"]
+ Summary --> EvidenceIndex["tool_trace_summary"]
+ GateResult --> Verifier["chat_verifier"]
+ ExecutorOutput --> Verifier
+ EvidenceIndex --> Verifier
+ Verifier --> Verdict{"verdict"}
+ Verdict -->|PASS| Composer["chat_composer"]
+ Composer --> Pass["输出最终答复"]
+ Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
+ Verdict -->|REJECT| Reject["降级输出"]
+```
+
+Gatekeeper 检查:
+
+| 检查 | 失败语义 |
+|---|---|
+| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
+| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
+| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
+| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
+| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
+| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
+
+Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
+
+Verifier 输出:
+
+```json
+{
+ "verdict": "PASS|LOW_CONFID|REJECT",
+ "groundedness_score": 0.8,
+ "critical_fact_count": 2,
+ "claim_checks": [],
+ "facts_checked": [],
+ "rationale": "..."
+}
+```
+
+Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
+
+结果写入:
+
+```text
+diagnosis_run.self_evaluation.verifier_evaluation
+```
+
+其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
+
+## 7. AIOps 规则门禁
+
+AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
+
+检查重点:
+
+- 最终报告是否存在。
+- payload 模式是否围绕输入告警展开。
+- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
+- 是否把无关活跃告警扩展成主诊断对象。
+
+结果写入:
+
+```text
+diagnosis_run.self_evaluation.aiops_rule_evaluation
+```
+
+## 8. Eval Baseline
+
+当前质量门禁还包括离线评测资产:
+
+| 评测 | 位置 | 作用 |
+|---|---|---|
+| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
+| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
+| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
+
+## 9. 后续门禁规划
+
+从旧版设计继承但尚未完整实现的门禁:
+
+- 工具参数 schema 校验。
+- 同一工具调用次数上限。
+- 工具超时的统一熔断。
+- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
+- Prompt 版本回滚和更细粒度变更审计。
+- Verifier 对 AIOps 报告的 LLM 级事实校验。
+
+这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
diff --git a/mvp/architecture/interview-one-pager.md b/mvp/architecture/archive/2026-07-22-legacy/interview-one-pager.md
similarity index 100%
rename from mvp/architecture/interview-one-pager.md
rename to mvp/architecture/archive/2026-07-22-legacy/interview-one-pager.md
diff --git a/mvp/architecture/knowledge-base-authoring.md b/mvp/architecture/archive/2026-07-22-legacy/knowledge-base-authoring.md
similarity index 100%
rename from mvp/architecture/knowledge-base-authoring.md
rename to mvp/architecture/archive/2026-07-22-legacy/knowledge-base-authoring.md
diff --git a/mvp/architecture/modular-rag-pipeline.md b/mvp/architecture/archive/2026-07-22-legacy/modular-rag-pipeline.md
similarity index 100%
rename from mvp/architecture/modular-rag-pipeline.md
rename to mvp/architecture/archive/2026-07-22-legacy/modular-rag-pipeline.md
diff --git a/mvp/architecture/rag-architecture.md b/mvp/architecture/archive/2026-07-22-legacy/rag-architecture.md
similarity index 100%
rename from mvp/architecture/rag-architecture.md
rename to mvp/architecture/archive/2026-07-22-legacy/rag-architecture.md
diff --git a/mvp/architecture/rag-eval-closure.md b/mvp/architecture/archive/2026-07-22-legacy/rag-eval-closure.md
similarity index 100%
rename from mvp/architecture/rag-eval-closure.md
rename to mvp/architecture/archive/2026-07-22-legacy/rag-eval-closure.md
diff --git a/mvp/architecture/retrieval-observability.md b/mvp/architecture/archive/2026-07-22-legacy/retrieval-observability.md
similarity index 100%
rename from mvp/architecture/retrieval-observability.md
rename to mvp/architecture/archive/2026-07-22-legacy/retrieval-observability.md
diff --git a/mvp/architecture/archive/2026-07-22-legacy/session-trace-lifecycle.md b/mvp/architecture/archive/2026-07-22-legacy/session-trace-lifecycle.md
new file mode 100644
index 0000000..cbb61c4
--- /dev/null
+++ b/mvp/architecture/archive/2026-07-22-legacy/session-trace-lifecycle.md
@@ -0,0 +1,160 @@
+# 会话与 Trace 生命周期
+
+**更新日期**:2026-07-10
+**状态**:当前可运行架构
+**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
+
+## 1. 定位
+
+当前 MVP 把“会话态”和“运行态”拆开:
+
+```text
+chat_session(sessionId)
+ -> diagnosis_run(runId)
+ -> agent_step(runId)
+ -> tool_invocation(runId)
+```
+
+- `sessionId` 表示多轮会话目录和 Redis 上下文。
+- `runId` 表示一次可回放诊断执行。
+- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
+- `diagnosis_session` 只保留为历史兼容和回滚表。
+
+## 2. 生命周期总图
+
+```mermaid
+flowchart TD
+ Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
+ Resolve --> Session["ensure chat_session metadata"]
+ Session --> Run["create diagnosis_run(runId)"]
+ Run --> Running["run.status = RUNNING"]
+
+ Running --> Agent["Agent workflow"]
+ Agent --> Context["execution context(sessionId, runId)"]
+ Context --> StepHook["AgentLoggingHook"]
+ StepHook --> Step["agent_step(session_id, run_id)"]
+ Context --> Tool["Evidence tools"]
+ Tool --> Invocation["tool_invocation(session_id, run_id)"]
+ Invocation --> Gatekeeper["Gatekeeper evidence validation"]
+
+ Agent --> Final{"workflow result"}
+ Final -->|success| Success["run.status = SUCCESS, answer saved"]
+ Final -->|failed| Failed["run.status = FAILED"]
+
+ Success --> Evaluation["diagnosis_run.self_evaluation merge"]
+ Failed --> Evaluation
+ Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
+ Success --> Feedback["POST /api/feedback(sessionId, runId)"]
+ Feedback --> Case["useful -> case_library(run_id)"]
+```
+
+## 3. ID 规则
+
+| ID | 来源 | 含义 |
+|---|---|---|
+| `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
+| `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
+
+设计含义:
+
+- 同一个 `sessionId` 可以贯穿多轮 Chat。
+- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
+- Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
+- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
+
+## 4. 运行状态流转
+
+```mermaid
+stateDiagram-v2
+ [*] --> PENDING
+ PENDING --> RUNNING: start diagnosis
+ RUNNING --> SUCCESS: workflow completed
+ RUNNING --> FAILED: exception / empty state
+ SUCCESS --> SUCCESS: feedback submitted
+ FAILED --> FAILED: feedback submitted
+```
+
+字段边界:
+
+| 字段 | 所属表 | 含义 |
+|---|---|---|
+| `status` | `diagnosis_run` | 单次运行执行状态 |
+| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
+| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
+| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
+
+`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
+
+## 5. agent_step 写入
+
+`AgentLoggingHook` 在模型调用前后写入和回填 `agent_step`。
+
+```mermaid
+sequenceDiagram
+ autonumber
+ participant Agent as Agent
+ participant Hook as AgentLoggingHook
+ participant DB as agent_step
+
+ Agent->>Hook: before_model(messages, sessionId, runId)
+ Hook->>DB: insert step(session_id, run_id, model_input, step_index)
+ Agent->>Hook: after_model(output, sessionId, runId)
+ Hook->>DB: update model_output, duration, token_count, has_tool_call
+```
+
+新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
+
+## 6. tool_invocation 写入
+
+工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
+
+```text
+ToolInvocationRecorder
+ -> tool_invocation.session_id
+ -> tool_invocation.run_id
+ -> retrieval_details / evidence_refs
+```
+
+Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
+
+## 7. Trace API 聚合
+
+```text
+GET /api/diagnosis/{sessionId}/trace
+GET /api/diagnosis/{sessionId}/trace?runId=run-...
+```
+
+聚合逻辑:
+
+```text
+diagnosis_run by sessionId + runId
+ + chat_session metadata when available
+ + agent_step where run_id = runId, ordered by the Trace API
+ + tool_invocation where run_id = runId order by id
+ -> DiagnosisTraceResponse
+```
+
+当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
+
+## 8. Chat 与 AIOps 差异
+
+| 维度 | Chat | AIOps |
+|---|---|---|
+| `agent_flow` | `CHAT` | `AI_OPS` |
+| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
+| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
+| 答案字段 | Chat 最终答复 | 告警分析报告 |
+| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
+
+## 9. 清理与边界
+
+- Redis 会话历史用于多轮上下文,不是长期审计记录。
+- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
+- `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
+- `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
+
+## 10. 后续增强
+
+1. Trace API 增加更结构化的 `self_evaluation` 展示。
+2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
+3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
diff --git a/mvp/architecture/current-mvp-architecture.md b/mvp/architecture/current-mvp-architecture.md
index a4bebb8..51f9bfe 100644
--- a/mvp/architecture/current-mvp-architecture.md
+++ b/mvp/architecture/current-mvp-architecture.md
@@ -1,444 +1,86 @@
# 当前 MVP 架构
-**更新日期**:2026-07-08
+**更新日期**:2026-07-22
**状态**:当前可运行架构
-**适用范围**:Demo、面试讲解、后续迭代规划
## 1. 系统定位
-SuperBizAgent MVP 不是通用 Chatbot,而是面向故障诊断的 Agent 工程项目。
+SuperBizAgent 是面向故障诊断的可追踪 Agent 应用。当前系统只保留一个拥有 Tool loop 的 `Diagnosis Agent`;Harness 负责确定性的预算、取消、工具边界、证据验真、语义审查和安全发布。
-核心目标:
-
-- 支持用户主动发起的 Chat 诊断。
-- 支持 AIOps 告警触发的自动诊断。
-- 保留 Agent 的规划、执行、验证过程。
-- 工具调用必须显式、可追踪、可回放。
-- RAG 检索必须通过 `lookup_knowledge` 暴露证据链。
-- 每次诊断都沉淀 session、step、tool invocation 和 self evaluation。
-
-## 2. 总体分层
+## 2. 分层
```mermaid
flowchart TB
- subgraph API["API Layer"]
- ChatController["ChatController"]
- TraceController["DiagnosisTraceController"]
- SearchController["SearchController"]
- DocumentController["DocumentController"]
- end
-
- subgraph App["Application Service"]
- ChatService["ChatService"]
- AiOpsService["AiOpsService"]
- TraceService["DiagnosisTraceService"]
- end
-
- subgraph Agent["Agent Orchestration"]
- Supervisor["Supervisor"]
- Planner["Planner"]
- Executor["Executor"]
- Gatekeeper["Gatekeeper"]
- Verifier["Verifier"]
- Composer["Composer"]
- end
-
- subgraph Tools["Evidence Tools"]
- KnowledgeTool["lookup_knowledge"]
- LogsTool["query_logs"]
- MetricsTool["query_metrics"]
- AlertsTool["queryPrometheusAlerts"]
- end
-
- subgraph Skills["Skill / Playbook"]
- SkillRegistry["SkillRegistry"]
- PlannerSkillHook["PlannerSkillMetadataHook"]
- SkillsHook["SkillsAgentHook"]
- ReadSkill["read_skill"]
- end
-
- subgraph RAG["RAG Retrieval"]
- L0["KnowledgeIndexService"]
- VectorSearch["VectorSearchService"]
- VectorStore["Spring AI VectorStore"]
- SdkFallback["Milvus SDK fallback"]
- end
-
- subgraph Store["Persistence and Trace"]
- ChatSession["chat_session"]
- Run["diagnosis_run"]
- Step["agent_step"]
- Invocation["tool_invocation"]
- ApiDoc["api_document"]
- Milvus["Milvus/Zilliz"]
- end
-
- API --> App
- ChatService --> Agent
- AiOpsService --> Agent
- SkillRegistry --> PlannerSkillHook
- PlannerSkillHook --> Planner
- SkillRegistry --> SkillsHook
- SkillsHook --> Executor
- Executor --> ReadSkill
- Agent --> Tools
- KnowledgeTool --> RAG
- RAG --> Store
- Tools --> Invocation
- Agent --> Step
- App --> Session
- TraceService --> Session
- TraceService --> Step
- TraceService --> Invocation
+ Browser["Browser / API client"] --> Chat["POST /api/chat named SSE"]
+ Chat --> App["ChatApplicationUseCase"]
+ App --> Router["Intent Router"]
+ Router --> System["System Chat"]
+ Router --> Knowledge["Knowledge Query"]
+ Router --> Diagnosis["Diagnosis Agent"]
+ Diagnosis --> Tools["Harness ACI Tools"]
+ Tools --> Canonical["Redis canonical invocation"]
+ Diagnosis --> Evidence["EvidenceGuard"]
+ Evidence --> Semantic["SemanticGuard"]
+ Semantic --> Release["Release Policy"]
+ Release --> Chat
+ App --> Run["diagnosis_run"]
+ Diagnosis --> Step["agent_step metadata audit"]
+ Tools --> Invocation["tool_invocation metadata audit"]
+ Run --> Trace["Diagnosis Trace API"]
+ Step --> Trace
+ Invocation --> Trace
```
-```text
-API Layer
- -> ChatController
- -> DiagnosisTraceController
- -> SearchController
- -> DocumentController
-
-Application Service
- -> ChatService
- -> AiOpsService
- -> DiagnosisTraceService
-
-Agent Orchestration
- -> Supervisor
- -> Planner
- -> Executor
- -> Gatekeeper
- -> Verifier
- -> Composer
-
-Evidence Tools
- -> lookup_knowledge
- -> query_logs
- -> query_metrics
- -> queryPrometheusAlerts
-
-Skill / Playbook
- -> SkillRegistry
- -> PlannerSkillMetadataHook gives Planner name/description only
- -> SkillsAgentHook gives Executor read_skill
- -> Verifier is isolated from skills
-
-RAG Retrieval
- -> KnowledgeIndexService
- -> VectorSearchService
- -> Spring AI VectorStore
- -> Milvus SDK fallback
-
-Persistence
- -> chat_session
- -> diagnosis_run
- -> agent_step.run_id
- -> tool_invocation.run_id
- -> api_document
- -> Milvus/Zilliz collection
-
-Quality Gates
- -> executor gatekeeper
- -> chat verifier
- -> AIOps rule evaluation
- -> diagnosis eval baseline
- -> RAG retrieval baseline
-```
-
-## 3. Chat 诊断链路
-
-```mermaid
-sequenceDiagram
- autonumber
- actor User as 用户
- participant API as POST /api/chat
- participant Chat as ChatService
- participant Planner as Planner Agent
- participant Executor as Executor Agent
- participant Tool as Evidence Tools
- participant Gatekeeper as Gatekeeper Hook
- participant Verifier as Verifier Agent
- participant Composer as Composer Agent
- participant DB as Trace Tables
- participant Trace as Trace API
-
- User->>API: 提交诊断问题
- API->>Chat: execute chat strategy
- Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
- Chat->>Planner: 复杂问题进入规划
- Planner->>DB: 写入 agent_step.run_id
- Planner->>Executor: 下发排查方向
- Executor->>Tool: lookup_knowledge / logs / metrics
- Tool->>DB: 写入 tool_invocation.run_id
- Tool-->>Executor: 返回证据
- Executor->>Gatekeeper: 输出 executor_evidence_v2
- Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
- Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
- Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
- Verifier->>Composer: 传入 allowed_claims / missing_info / actions
- Composer->>Chat: 生成最终用户答复
- Chat->>DB: 保存 diagnosis_run.answer
- User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
- Trace->>DB: 聚合 run / step / tool
- Trace-->>User: 返回可回放诊断链路
-```
+## 3. 唯一 Chat 主链
```text
POST /api/chat
- -> ChatService
- -> 简单问题:轻量回答
- -> 复杂诊断:Agent 编排
- -> Planner 制定排查方向
- -> Executor 调用证据工具
- -> lookup_knowledge
- -> query_logs
- -> query_metrics
- -> Gatekeeper 校验 Executor 证据引用真实性
- -> Verifier 判断 claim 是否能由已核验证据推出
- -> Composer 生成最终用户答复
- -> 保存 chat_session metadata
- -> 保存 diagnosis_run
- -> 保存 agent_step.run_id
- -> 保存 tool_invocation.run_id
- -> 合并 diagnosis_run.self_evaluation.verifier_evaluation
+ -> metadata(session_id, run_id)
+ -> status*
+ -> ChatApplicationUseCase
+ -> SYSTEM_CHAT | KNOWLEDGE_QUERY | DIAGNOSIS
+ -> content | failure
+ -> done(SUCCESS | FALLBACK | FAILED)
```
-Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
+- Controller 只处理请求校验、bounded worker、SSE 和 disconnect。
+- Application Use Case 拥有 Session/Run、路由、PreviousTurn 和终态持久化。
+- Diagnosis Agent 是唯一报告作者和唯一拥有 evidence Tool loop 的业务 Agent。
+- EvidenceGuard 只做确定性结构/引用验真;SemanticGuard 在隔离上下文做整份报告语义审查。
+- 未通过 Release Policy 的 Draft 永不进入公开 SSE。
-Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
+## 4. Tool 与数据边界
-关键代码:
+Agent 只看到三个固定 Tool:
-- `src/main/java/com/superbiz/agent/controller/ChatController.java`
-- `src/main/java/com/superbiz/agent/service/ChatService.java`
-- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
-- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
+- `lookup_knowledge`
+- `query_logs`
+- `query_mysql`
-## 4. AIOps 诊断链路
+每次调用由框架提供 `tool_call_id`,Harness 校验 exact run、只读、Schema、预算和容量。Redis 保存 TTL 内完整 canonical invocation;MySQL `tool_invocation` 只保存长期有界 metadata,不保存完整参数、SQL/日志正文、raw response 或 Agent projection。
-```mermaid
-flowchart TD
- Request["POST /api/ai_ops"] --> Payload{"包含告警 payload?"}
- Payload -->|是| Targeted["PAYLOAD_TARGETED"]
- Payload -->|否| Discovery["AUTO_DISCOVERY"]
-
- Targeted --> BuildPrompt["构造聚焦 payload 的诊断 prompt"]
- Targeted --> QueryAug["生成 recommended lookup_knowledge query"]
- Discovery --> DiscoverAlert["通过 queryPrometheusAlerts 发现活跃告警"]
-
- BuildPrompt --> Plan["Planner 规划排查"]
- QueryAug --> Plan
- DiscoverAlert --> Plan
-
- Plan --> Execute["Executor 收集证据"]
- Execute --> Knowledge["lookup_knowledge"]
- Execute --> Metrics["query_metrics / Prometheus"]
- Execute --> Logs["query_logs"]
-
- Knowledge --> Report["告警分析报告"]
- Metrics --> Report
- Logs --> Report
-
- Report --> RuleEval["AiOpsRuleEvaluationService"]
- RuleEval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
- Report --> Trace["DiagnosisTraceService"]
- SelfEval --> Trace
-```
+## 5. Trace 与持久化
```text
-POST /api/ai_ops
- -> AiOpsService
- -> 判断是否有告警 payload
- -> PAYLOAD_TARGETED
- -> AUTO_DISCOVERY
- -> 构造 AIOps 诊断 prompt
- -> payload 模式补充 recommended lookup_knowledge query
- -> Agent 编排
- -> Planner / Executor
- -> Prometheus / logs / knowledge tools
- -> 生成告警分析报告
- -> AiOpsRuleEvaluationService
- -> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
- -> Trace API 可查看全链路
+chat_session(sessionId)
+ -> diagnosis_run(runId)
+ -> agent_step(runId)
+ -> tool_invocation(runId)
```
-AIOps 保留两种模式:
+- `chat_session` 是 JPA Run 目录与多轮 metadata,不保存完整对话历史。
+- `diagnosis_run` 是 Run 状态、intent、release outcome、安全发布结果和预算汇总真理源。
+- `agent_step` 只保存模型步骤 metadata,不保存 Prompt、消息正文、模型正文或 Thought。
+- `tool_invocation` 只保存 Tool durable audit metadata;完整调用由 Redis canonical store 短期保存。
-| 模式 | 触发条件 | 行为 |
-|---|---|---|
-| `PAYLOAD_TARGETED` | 请求包含 alertName、service、severity、description、timeRange 等字段 | 以 payload 为唯一主诊断对象,并生成推荐知识库 query |
-| `AUTO_DISCOVERY` | 请求没有明确告警 payload | 先查询当前活跃告警,再选择目标排查 |
+## 6. 公开 API
-AIOps 当前使用轻量规则验证器,重点检查:
+当前诊断执行入口只有 `POST /api/chat`。Trace、feedback、文档与检索 API 保持独立;已删除的旧诊断和 Redis conversation Session endpoint 不提供兼容分支。
-- 最终报告是否存在。
-- payload 模式是否聚焦输入告警。
-- 是否使用关键证据工具,例如 `lookup_knowledge`、日志、指标。
+## 7. 安全边界
-关键代码:
-
-- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
-- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
-
-## 5. RAG 位置
-
-RAG 不是隐藏在 Chat Advisor 里的隐式能力,而是 Executor 可以显式调用的工具:
-
-```mermaid
-flowchart LR
- Executor["Executor Agent"] --> Tool["lookup_knowledge Tool"]
- Tool --> L0["L0 domain/entity hint"]
- Tool --> Search["VectorSearchService"]
- L0 --> Search
- Search --> VectorStore["Spring AI VectorStore"]
- Search --> Fallback["Milvus SDK fallback"]
- VectorStore --> Normalize["score/rawScore/scoreLabel"]
- Fallback --> Normalize
- Normalize --> Evidence["evidence output"]
- Evidence --> Invocation["tool_invocation"]
- Evidence --> Executor
-```
-
-```text
-Executor
- -> lookup_knowledge(query)
- -> L0 domain/entity hint
- -> VectorSearchService
- -> Spring AI VectorStore
- -> Milvus SDK fallback
- -> evidence shaping
- -> tool_invocation
-```
-
-保留显式工具的原因:
-
-- Agent 何时检索、检索什么、证据是什么,必须能在 trace 中解释。
-- AIOps payload 到 query 的业务映射需要项目内控制。
-- `tool_invocation` 是后续评测、回放和面试讲解的核心材料。
-
-RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细节见 [retrieval-observability.md](retrieval-observability.md)。
-
-## 6. 持久化模型
-
-当前诊断持久化以 session/run/trace 明细为核心:
-
-```text
-chat_session
- -> 多轮会话目录和元数据
- -> session_id / status / message_pair_count
-
-diagnosis_run
- -> 一次诊断运行的主记录
- -> run_id / session_id
- -> query / status / agent_flow / answer
- -> self_evaluation
- -> step_count / tool_call_count / duration
-
-agent_step
- -> Agent 模型调用步骤
- -> session_id / run_id
- -> step_index / agent_name
- -> model_input / model_output / thought
- -> duration / token_count
-
-tool_invocation
- -> 工具调用事实
- -> session_id / run_id
- -> tool_name / input_params / output_preview
- -> retrieval_layer / retrieval_details
- -> retrieval_details.evidence_refs
- -> relevance_level / dedup_reason
- -> duration / success
-```
-
-说明:
-
-- 旧的 `diagnosis_record` 已不是当前主模型。
-- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
-- `api_document` 仍用于文档元数据管理。
-- 文档向量内容存放在 Milvus/Zilliz collection 中。
-
-会话和 Trace 生命周期见 [session-trace-lifecycle.md](session-trace-lifecycle.md),完整数据关系见 [data-model.md](data-model.md)。
-
-## 7. Trace API
-
-```text
-GET /api/diagnosis/{sessionId}/trace
-GET /api/diagnosis/{sessionId}/trace?runId=run-...
-```
-
-Trace API 聚合:
-
-- 会话元数据、运行状态和最终报告。
-- Agent step 序列。
-- 工具调用和检索细节。
-- Chat Gatekeeper / Verifier / Composer 结果。
-- AIOps rule evaluation 结果。
-
-Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
-
-Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
-
-## 8. 质量门禁
-
-当前质量门禁分层如下:
-
-| 门禁 | 位置 | 作用 |
-|---|---|---|
-| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
-| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
-| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
-| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
-| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
-| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
-| Live RAG Acceptance | `scripts/eval_rag_live_acceptance.py` | 在运行环境中验证重建索引后的真实检索 |
-
-## 9. 当前完成状态
-
-已经完成:
-
-- Chat 和 AIOps 两条入口链路。
-- 显式 `lookup_knowledge` Agent Tool。
-- L0 从最终决策降级为 domain/entity hint。
-- `VectorSearchService` 作为稳定检索门面。
-- Spring AI VectorStore 读取路径。
-- Milvus SDK fallback。
-- `score` / `rawScore` / `scoreLabel` 分数语义拆分。
-- `title`、`breadcrumb`、`content` 参与 embedding 文本。
-- `tool_invocation` 记录检索层、relevance level、dedup reason。
-- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
-- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
-- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
-- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
-- Verifier 只判断可推导性。
-- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
-- RAG offline baseline 和 live acceptance 脚本。
-
-暂不作为当前已完成能力声明:
-
-- 完整 QueryTransformer / MultiQuery。
-- BM25、RRF、cross-encoder rerank。
-- 完整邻居 chunk / section context expansion。
-- VectorStore 写入路径全面迁移。
-- 完整 LLM-based AIOps verifier。
-
-后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
-
-## 10. 关键代码索引
-
-| 能力 | 代码 |
-|---|---|
-| Chat 入口与编排 | `ChatController`, `ChatService` |
-| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
-| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
-| 知识库工具 | `LookupKnowledgeTool` |
-| L0 hint | `KnowledgeIndexService` |
-| 向量检索门面 | `VectorSearchService` |
-| 文档切片 | `DocumentChunkService` |
-| 向量写入 | `VectorIndexService` |
-| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
-| Trace 聚合 | `DiagnosisTraceService` |
-| 工具调用记录 | `ToolInvocationRecorder` |
-| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
-| self_evaluation 合并 | `SelfEvaluationMergeService` |
+- 不输出或长期持久化 Chain of Thought。
+- 不向 Agent 暴露 Redis、canonical key、完整 Tool 请求/响应或数据库凭据。
+- EvidenceGuard 只接受当前 Run 的 READY canonical invocation。
+- SemanticGuard 无 Tool、无记忆、无回调主 Agent 能力。
+- technical failure 与 guard rejection 只能产生 stable failure 或固定 safe fallback。
diff --git a/mvp/architecture/harness-quality-gates.md b/mvp/architecture/harness-quality-gates.md
index b91d4e3..64eb09a 100644
--- a/mvp/architecture/harness-quality-gates.md
+++ b/mvp/architecture/harness-quality-gates.md
@@ -1,258 +1,48 @@
-# Harness 与质量门禁架构
+# Harness 与质量门禁
-**更新日期**:2026-07-08
-**状态**:当前可运行架构 + 后续门禁规划
-**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
+**更新日期**:2026-07-22
+**状态**:当前可运行架构
-## 1. 设计目标
+## 1. Harness 定位
-Agent 系统的核心风险不是“没有答案”,而是:
+Harness 是确定性执行边界,不承担业务推理。它统一管理:
-- 答案引用了不存在的证据。
-- 工具调用失败后仍然编造结论。
-- 检索结果相关性不足但被当作强证据。
-- 多轮诊断重复检索同一文档,浪费上下文。
-- 最终报告无法回放执行过程。
+- RunContext、deadline、first-terminal-wins lifecycle 与客户端取消。
+- 模型调用、Tool 调用、Token、字节数和单 Tool 次数预算。
+- 类型化 retry policy;Diagnosis Agent 和 Tool 调用不自动重试。
+- ToolBoundary、canonical invocation 与 Agent projection。
+- EvidenceGuard、Evidence repair、SemanticGuard 与 Release Policy。
+- metadata-only durable audit。
-因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
+## 2. ToolBoundary
```text
-Prompt contract
- + Tool boundary
- + Agent hooks
- + Trace persistence
- + Gatekeeper deterministic validation
- + Verifier / rule evaluation
- + Eval baseline
+framework tool_call_id
+ -> exact Run / schema / authorization / read-only / budget
+ -> backend execution
+ -> raw response -> Redis canonical invocation
+ -> projector -> bounded agent_result
+ -> ToolInvocation durable metadata audit
+ -> Agent observation
```
-## 2. Harness 总图
+Redis canonical invocation 可在 TTL 内保存完整 request/raw_response/agent_result,受独立前缀、容量和 Harness-only 访问保护。Durable audit 只保存 identity、Tool 名、状态、耗时和字节数;audit 写入失败可观测但不改变 canonical Tool 结果。
-```mermaid
-flowchart TB
- Input["User / AIOps input"] --> Prompt["Prompt contract"]
- Prompt --> Agent["Planner / Executor / Verifier / Composer"]
- Agent --> Tools["Evidence tools"]
- Tools --> Invocation["tool_invocation"]
- Agent --> StepHook["AgentLoggingHook"]
- StepHook --> Step["agent_step"]
- Agent --> Run["diagnosis_run"]
+## 3. EvidenceGuard
- Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
- EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
- Agent --> Gatekeeper
- Invocation --> TraceSummary["ToolTraceSummaryService"]
- Gatekeeper --> Verifier["chat_verifier"]
- TraceSummary --> Verifier
- Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
+EvidenceGuard 不调用模型。它校验 Draft schema、analysis ID、当前 Run Tool ownership、READY 状态、evidence status 和每条结论的引用闭包,并生成只包含 Agent projection 的 verified snapshot。
- Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
- AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
+## 4. SemanticGuard
- Run --> TraceAPI["DiagnosisTraceService"]
- Step --> TraceAPI
- Invocation --> TraceAPI
- SelfEval --> TraceAPI
- AiOpsEval --> TraceAPI
+SemanticGuard 使用隔离的单轮模型调用,只接收原始 query、完整 Draft 和 verified snapshot。它无 Tool、无记忆、不访问 Redis、不改写报告;技术失败最多按相同输入重试一次,仍失败则安全降级。
- TraceAPI --> Eval["diagnosis eval / RAG eval"]
-```
+## 5. Release Policy
-## 3. Prompt Contract
+- `SUPPORTED`:发布 Diagnosis Agent 原始安全 Draft 的 typed report。
+- `UNSUPPORTED` 或 evidence failure:发布固定 SAFE_FALLBACK。
+- technical failure:发布 stable failure,不泄漏内部异常。
+- cancel/timeout:结束 exact Run,禁止 late content。
-当前 Prompt 按角色拆分:
+## 6. Audit 安全
-| Prompt | 用途 |
-|---|---|
-| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
-| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
-| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
-| `chat-planner-prompt.md` | Chat 复杂问题规划 |
-| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
-| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
-| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
-
-Prompt 层当前承担的门禁:
-
-- 禁止凭记忆回答错误码、接口定义、排障步骤。
-- 需要外部信息时必须调用工具。
-- 工具连续失败或返回空结果时,最终报告必须诚实说明。
-- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
-- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
-- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
-- AIOps payload 模式必须聚焦输入告警。
-
-Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
-
-```json
-{
- "version": "chat-prompts-v1",
- "prompts": [
- {
- "name": "chat_executor",
- "version": "chat-executor-v2",
- "resource": "prompts/chat-executor-prompt.md"
- }
- ]
-}
-```
-
-该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
-
-## 4. Trace Hooks
-
-`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
-
-```mermaid
-sequenceDiagram
- autonumber
- participant A as Agent
- participant H as AgentLoggingHook
- participant DB as agent_step
-
- A->>H: before_model(messages, sessionId, runId)
- H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
- A-->>A: LLM 推理
- A->>H: after_model(messages, sessionId, runId)
- H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
-```
-
-记录内容:
-
-- 最近输入消息摘要。
-- Agent 输出摘要。
-- 是否包含 tool call。
-- duration。
-- token count。
-- Verifier 的 JSON 输出摘要。
-
-新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
-
-## 5. Tool Invocation 门禁
-
-工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
-
-核心记录:
-
-```text
-tool_name
-input_params
-output_preview
-retrieval_layer
-l0_match_count
-l1_match_count
-retrieval_details
- -> evidence_refs
-relevance_level
-dedup_reason
-duration_ms
-success
-error_message
-```
-
-对 `lookup_knowledge` 的质量约束:
-
-- L0 只作为 hint,不绕过 L1。
-- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
-- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
-- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
-- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
-- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
-
-## 6. Gatekeeper 与 Verifier 门禁
-
-Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
-
-```mermaid
-flowchart LR
- Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
- ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
- EvidenceRefs --> Gatekeeper
- Gatekeeper --> GateResult["gatekeeper_result"]
- Invocation --> Summary["ToolTraceSummaryService"]
- Summary --> EvidenceIndex["tool_trace_summary"]
- GateResult --> Verifier["chat_verifier"]
- ExecutorOutput --> Verifier
- EvidenceIndex --> Verifier
- Verifier --> Verdict{"verdict"}
- Verdict -->|PASS| Composer["chat_composer"]
- Composer --> Pass["输出最终答复"]
- Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
- Verdict -->|REJECT| Reject["降级输出"]
-```
-
-Gatekeeper 检查:
-
-| 检查 | 失败语义 |
-|---|---|
-| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
-| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
-| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
-| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
-| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
-| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
-
-Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
-
-Verifier 输出:
-
-```json
-{
- "verdict": "PASS|LOW_CONFID|REJECT",
- "groundedness_score": 0.8,
- "critical_fact_count": 2,
- "claim_checks": [],
- "facts_checked": [],
- "rationale": "..."
-}
-```
-
-Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
-
-结果写入:
-
-```text
-diagnosis_run.self_evaluation.verifier_evaluation
-```
-
-其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
-
-## 7. AIOps 规则门禁
-
-AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
-
-检查重点:
-
-- 最终报告是否存在。
-- payload 模式是否围绕输入告警展开。
-- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
-- 是否把无关活跃告警扩展成主诊断对象。
-
-结果写入:
-
-```text
-diagnosis_run.self_evaluation.aiops_rule_evaluation
-```
-
-## 8. Eval Baseline
-
-当前质量门禁还包括离线评测资产:
-
-| 评测 | 位置 | 作用 |
-|---|---|---|
-| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
-| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
-| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
-
-## 9. 后续门禁规划
-
-从旧版设计继承但尚未完整实现的门禁:
-
-- 工具参数 schema 校验。
-- 同一工具调用次数上限。
-- 工具超时的统一熔断。
-- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
-- Prompt 版本回滚和更细粒度变更审计。
-- Verifier 对 AIOps 报告的 LLM 级事实校验。
-
-这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
+AgentStep 不保存 Prompt、消息正文、模型正文、Tool arguments 或 Thought。ToolInvocation 不保存完整 request、SQL/日志 query、raw response 或 Agent projection。应用日志不得打印这些字段。
diff --git a/mvp/architecture/session-trace-lifecycle.md b/mvp/architecture/session-trace-lifecycle.md
index cbb61c4..526c627 100644
--- a/mvp/architecture/session-trace-lifecycle.md
+++ b/mvp/architecture/session-trace-lifecycle.md
@@ -1,160 +1,39 @@
-# 会话与 Trace 生命周期
+# Session、Run 与 Trace 生命周期
-**更新日期**:2026-07-10
+**更新日期**:2026-07-22
**状态**:当前可运行架构
-**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
-## 1. 定位
+## 1. Identity
-当前 MVP 把“会话态”和“运行态”拆开:
+- `sessionId`:多轮对话目录,由客户端传入或应用生成。
+- `runId`:一次 Chat 执行,由 Harness 生成并在 SSE metadata 首事件返回。
+- 所有 Run、AgentStep、ToolInvocation 和 Trace 查询必须使用同一个 exact ID;禁止用“最新一条”替代。
+
+## 2. 生命周期
```text
-chat_session(sessionId)
- -> diagnosis_run(runId)
- -> agent_step(runId)
- -> tool_invocation(runId)
+request accepted
+ -> start RunContext
+ -> persist diagnosis_run RUNNING
+ -> metadata(session_id, run_id)
+ -> route / execute / guard / release
+ -> SUCCESS | FALLBACK | FAILED | CANCELLED
+ -> persist terminal state and budget usage
```
-- `sessionId` 表示多轮会话目录和 Redis 上下文。
-- `runId` 表示一次可回放诊断执行。
-- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
-- `diagnosis_session` 只保留为历史兼容和回滚表。
+disconnect、timeout 与 send failure 通过同一个 `ChatRunControl` 请求取消。正常 SSE complete 在 callback 前标记 terminal,避免误取消;late content 被 state machine 拒绝。
-## 2. 生命周期总图
+## 3. Trace 聚合
-```mermaid
-flowchart TD
- Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
- Resolve --> Session["ensure chat_session metadata"]
- Session --> Run["create diagnosis_run(runId)"]
- Run --> Running["run.status = RUNNING"]
+`GET /api/diagnosis/{sessionId}/trace?runId={runId}` 聚合:
- Running --> Agent["Agent workflow"]
- Agent --> Context["execution context(sessionId, runId)"]
- Context --> StepHook["AgentLoggingHook"]
- StepHook --> Step["agent_step(session_id, run_id)"]
- Context --> Tool["Evidence tools"]
- Tool --> Invocation["tool_invocation(session_id, run_id)"]
- Invocation --> Gatekeeper["Gatekeeper evidence validation"]
+- `chat_session` metadata。
+- exact `diagnosis_run` 状态、intent、release outcome、安全 answer 与预算。
+- `agent_step` metadata-only 模型步骤。
+- `tool_invocation` metadata-only Tool durable audit。
- Agent --> Final{"workflow result"}
- Final -->|success| Success["run.status = SUCCESS, answer saved"]
- Final -->|failed| Failed["run.status = FAILED"]
+Redis canonical invocation 不是 Trace API 的长期响应内容;它只供当前 Run EvidenceGuard 验真。
- Success --> Evaluation["diagnosis_run.self_evaluation merge"]
- Failed --> Evaluation
- Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
- Success --> Feedback["POST /api/feedback(sessionId, runId)"]
- Feedback --> Case["useful -> case_library(run_id)"]
-```
+## 4. PreviousTurn
-## 3. ID 规则
-
-| ID | 来源 | 含义 |
-|---|---|---|
-| `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
-| `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
-
-设计含义:
-
-- 同一个 `sessionId` 可以贯穿多轮 Chat。
-- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
-- Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
-- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
-
-## 4. 运行状态流转
-
-```mermaid
-stateDiagram-v2
- [*] --> PENDING
- PENDING --> RUNNING: start diagnosis
- RUNNING --> SUCCESS: workflow completed
- RUNNING --> FAILED: exception / empty state
- SUCCESS --> SUCCESS: feedback submitted
- FAILED --> FAILED: feedback submitted
-```
-
-字段边界:
-
-| 字段 | 所属表 | 含义 |
-|---|---|---|
-| `status` | `diagnosis_run` | 单次运行执行状态 |
-| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
-| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
-| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
-
-`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
-
-## 5. agent_step 写入
-
-`AgentLoggingHook` 在模型调用前后写入和回填 `agent_step`。
-
-```mermaid
-sequenceDiagram
- autonumber
- participant Agent as Agent
- participant Hook as AgentLoggingHook
- participant DB as agent_step
-
- Agent->>Hook: before_model(messages, sessionId, runId)
- Hook->>DB: insert step(session_id, run_id, model_input, step_index)
- Agent->>Hook: after_model(output, sessionId, runId)
- Hook->>DB: update model_output, duration, token_count, has_tool_call
-```
-
-新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
-
-## 6. tool_invocation 写入
-
-工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
-
-```text
-ToolInvocationRecorder
- -> tool_invocation.session_id
- -> tool_invocation.run_id
- -> retrieval_details / evidence_refs
-```
-
-Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
-
-## 7. Trace API 聚合
-
-```text
-GET /api/diagnosis/{sessionId}/trace
-GET /api/diagnosis/{sessionId}/trace?runId=run-...
-```
-
-聚合逻辑:
-
-```text
-diagnosis_run by sessionId + runId
- + chat_session metadata when available
- + agent_step where run_id = runId, ordered by the Trace API
- + tool_invocation where run_id = runId order by id
- -> DiagnosisTraceResponse
-```
-
-当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
-
-## 8. Chat 与 AIOps 差异
-
-| 维度 | Chat | AIOps |
-|---|---|---|
-| `agent_flow` | `CHAT` | `AI_OPS` |
-| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
-| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
-| 答案字段 | Chat 最终答复 | 告警分析报告 |
-| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
-
-## 9. 清理与边界
-
-- Redis 会话历史用于多轮上下文,不是长期审计记录。
-- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
-- `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
-- `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
-
-## 10. 后续增强
-
-1. Trace API 增加更结构化的 `self_evaluation` 展示。
-2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
-3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
+应用在创建当前 Run 前读取同 Session 最近安全发布结果。只允许结构化 PublishedResult 的固定字段进入 PreviousTurn,且执行字节上限;完整历史、失败、Fallback、Tool raw data 和 guard reason 均排除。
diff --git a/mvp/demo/README.md b/mvp/demo/README.md
index f4757a8..9518848 100644
--- a/mvp/demo/README.md
+++ b/mvp/demo/README.md
@@ -1,205 +1,29 @@
-# MVP 演示手册
+# 单 Diagnosis Agent Demo
-本目录用于演示 MVP 从用户问题到诊断 Trace 的完整闭环。
+**更新日期**:2026-07-22
-面试时建议先读:
+## 运行前提
-- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
-- `interview-walkthrough.md`:面试讲解话术。
-- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
-- `trace-inspection-checklist.md`:Trace 字段检查清单。
-- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
-- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
-- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
-- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
-- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
-- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
-- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
+- 应用、MySQL、Redis、Milvus 与模型配置可用。
+- `cls.mock-enabled=true` 用于 query_logs Mock 证据。
+- query_mysql 只使用 `mysql-tool.datasources` 配置的隔离只读数据源;不查询应用数据库。
-## 1. 前置条件
+## 主流程
-- MySQL、Redis、Milvus/Zilliz、LLM 和 embedding 配置可用。
-- 安全和密钥清理不属于当前 MVP 演示范围。
-- `mvp-demo` profile 会启用 mock Prometheus 和 mock CLS,让日志和指标工具返回可复现证据。
+1. 启动应用:`mvn spring-boot:run`。
+2. 向 `POST /api/chat` 提交 `requests/payment-timeout-chat.json`。
+3. 验证 SSE:`metadata -> status* -> content|failure -> done`。
+4. 保存 metadata 的 exact `session_id` 与 `run_id`。
+5. 检查 `logs/application.log` 的 Run/Tool/Guard/Release 状态,确认无 Prompt、Thought 或 raw Tool payload。
+6. 使用 `scripts/query_mysql.py` 按 exact runId 查询 `diagnosis_run`、`agent_step`、`tool_invocation`。
+7. 打开 `/trace.html?sessionId=...&runId=...` 检查聚合 Trace。
-## 2. 启动服务
+## 验收重点
-```powershell
-mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
-```
+- 只有 `diagnosis_agent` 具有 Tool loop。
+- Tool 只包含 `lookup_knowledge`、`query_logs`、`query_mysql`。
+- EvidenceGuard/SemanticGuard 完成前没有 content。
+- SUCCESS 发布 typed diagnosis report;证据或语义不支持时发布固定 safe fallback。
+- query_logs 为 Mock;不把它表述为真实 CLS live 结果。
-服务地址:
-
-```text
-http://localhost:9900
-```
-
-## 3. Chat 诊断 Demo
-
-最快方式:
-
-```powershell
-powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
-```
-
-脚本会生成:
-
-```text
-mvp/demo/output/chat-response.json
-mvp/demo/output/trace-response.json
-mvp/demo/output/feedback-response.json
-mvp/demo/output/interview-demo-summary.json
-```
-
-手动请求:
-
-```powershell
-$sessionId = "mvp-demo-payment-timeout-001"
-$body = @{
- Id = $sessionId
- Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
-} | ConvertTo-Json
-
-Invoke-RestMethod `
- -Method Post `
- -Uri "http://localhost:9900/api/chat" `
- -ContentType "application/json" `
- -Body $body
-```
-
-如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
-
-```powershell
-$chat = Invoke-RestMethod `
- -Method Post `
- -Uri "http://localhost:9900/api/chat" `
- -ContentType "application/json" `
- -Body $body
-
-$runId = $chat.data.runId
-```
-
-期望结果:
-
-- `data.success = true`
-- `data.sessionId = mvp-demo-payment-timeout-001`
-- `data.runId` 为本次诊断运行的唯一 ID
-- `data.answer` 包含诊断答复
-
-## 4. 查询 Trace
-
-```powershell
-Invoke-RestMethod `
- -Method Get `
- -Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
-```
-
-期望结果:
-
-- `code = 200`
-- `data.runId` 等于 `$runId`
-- `data.session.sessionId` 等于 Chat session id
-- `data.run.runId` 等于 `$runId`
-- `data.steps` 包含 planner / executor / verifier 等步骤
-- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
-- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
-- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
-- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
-
-## 5. 提交反馈
-
-```powershell
-$feedback = @{
- sessionId = $sessionId
- runId = $runId
- feedback = "useful"
-} | ConvertTo-Json
-
-Invoke-RestMethod `
- -Method Post `
- -Uri "http://localhost:9900/api/feedback" `
- -ContentType "application/json" `
- -Body $feedback
-```
-
-期望结果:
-
-- `success = true`
-- `runId = $runId`
-- 后续精确 Trace 中 `data.session.feedback = useful`
-- useful 反馈会尝试沉淀 `case_library`
-
-## 6. AIOps 告警诊断 Demo
-
-```powershell
-$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
-$aiopsBody = @{
- sessionId = $aiopsSessionId
- alertName = "HighCPUUsage"
- service = "payment-service"
- severity = "P1"
- description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
- timeRange = "last_15m"
- userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
-} | ConvertTo-Json
-
-Invoke-WebRequest `
- -Method Post `
- -Uri "http://localhost:9900/api/ai_ops" `
- -ContentType "application/json" `
- -Body $aiopsBody
-```
-
-期望结果:
-
-- SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
-- 后续流式输出包含 AIOps 告警分析报告
-- 报告聚焦输入的 `HighCPUUsage/payment-service`
-- 精确 Trace 中 `data.session.agentFlow = AI_OPS`
-- `data.session.answer` 包含最终告警报告
-- `data.toolInvocations` 包含证据工具调用
-
-查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
-
-```powershell
-Invoke-RestMethod `
- -Method Get `
- -Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
-```
-
-## 7. Demo 主线
-
-Chat 主线:
-
-```text
-一个 session id + 一个 run id
--> 用户问题
--> 多 Agent 执行
--> 证据工具
--> Verifier / self_evaluation
--> 最终答案
--> 用户反馈
--> Trace API 回放
-```
-
-AIOps 主线:
-
-```text
-一个 session id + 一个 run id
--> 告警 payload
--> AIOps Planner / Executor
--> 证据工具
--> 告警分析报告
--> AIOps rule evaluation
--> Trace API 回放
-```
-
-## 8. Evidence Pipeline 场景矩阵
-
-面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
-
-- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
-- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
-- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
-
-这样可以同时展示真实链路和确定性回归能力。
+旧多角色与第二诊断入口的 demo 已保存在 `archive/2026-07-22-legacy/`,不代表当前运行时。
diff --git a/mvp/demo/archive/2026-07-22-legacy/README.md b/mvp/demo/archive/2026-07-22-legacy/README.md
new file mode 100644
index 0000000..f4757a8
--- /dev/null
+++ b/mvp/demo/archive/2026-07-22-legacy/README.md
@@ -0,0 +1,205 @@
+# MVP 演示手册
+
+本目录用于演示 MVP 从用户问题到诊断 Trace 的完整闭环。
+
+面试时建议先读:
+
+- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
+- `interview-walkthrough.md`:面试讲解话术。
+- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
+- `trace-inspection-checklist.md`:Trace 字段检查清单。
+- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
+- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
+- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
+- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
+- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
+- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
+- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
+
+## 1. 前置条件
+
+- MySQL、Redis、Milvus/Zilliz、LLM 和 embedding 配置可用。
+- 安全和密钥清理不属于当前 MVP 演示范围。
+- `mvp-demo` profile 会启用 mock Prometheus 和 mock CLS,让日志和指标工具返回可复现证据。
+
+## 2. 启动服务
+
+```powershell
+mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
+```
+
+服务地址:
+
+```text
+http://localhost:9900
+```
+
+## 3. Chat 诊断 Demo
+
+最快方式:
+
+```powershell
+powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
+```
+
+脚本会生成:
+
+```text
+mvp/demo/output/chat-response.json
+mvp/demo/output/trace-response.json
+mvp/demo/output/feedback-response.json
+mvp/demo/output/interview-demo-summary.json
+```
+
+手动请求:
+
+```powershell
+$sessionId = "mvp-demo-payment-timeout-001"
+$body = @{
+ Id = $sessionId
+ Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
+} | ConvertTo-Json
+
+Invoke-RestMethod `
+ -Method Post `
+ -Uri "http://localhost:9900/api/chat" `
+ -ContentType "application/json" `
+ -Body $body
+```
+
+如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
+
+```powershell
+$chat = Invoke-RestMethod `
+ -Method Post `
+ -Uri "http://localhost:9900/api/chat" `
+ -ContentType "application/json" `
+ -Body $body
+
+$runId = $chat.data.runId
+```
+
+期望结果:
+
+- `data.success = true`
+- `data.sessionId = mvp-demo-payment-timeout-001`
+- `data.runId` 为本次诊断运行的唯一 ID
+- `data.answer` 包含诊断答复
+
+## 4. 查询 Trace
+
+```powershell
+Invoke-RestMethod `
+ -Method Get `
+ -Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
+```
+
+期望结果:
+
+- `code = 200`
+- `data.runId` 等于 `$runId`
+- `data.session.sessionId` 等于 Chat session id
+- `data.run.runId` 等于 `$runId`
+- `data.steps` 包含 planner / executor / verifier 等步骤
+- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
+- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
+- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
+- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
+
+## 5. 提交反馈
+
+```powershell
+$feedback = @{
+ sessionId = $sessionId
+ runId = $runId
+ feedback = "useful"
+} | ConvertTo-Json
+
+Invoke-RestMethod `
+ -Method Post `
+ -Uri "http://localhost:9900/api/feedback" `
+ -ContentType "application/json" `
+ -Body $feedback
+```
+
+期望结果:
+
+- `success = true`
+- `runId = $runId`
+- 后续精确 Trace 中 `data.session.feedback = useful`
+- useful 反馈会尝试沉淀 `case_library`
+
+## 6. AIOps 告警诊断 Demo
+
+```powershell
+$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
+$aiopsBody = @{
+ sessionId = $aiopsSessionId
+ alertName = "HighCPUUsage"
+ service = "payment-service"
+ severity = "P1"
+ description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
+ timeRange = "last_15m"
+ userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
+} | ConvertTo-Json
+
+Invoke-WebRequest `
+ -Method Post `
+ -Uri "http://localhost:9900/api/ai_ops" `
+ -ContentType "application/json" `
+ -Body $aiopsBody
+```
+
+期望结果:
+
+- SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
+- 后续流式输出包含 AIOps 告警分析报告
+- 报告聚焦输入的 `HighCPUUsage/payment-service`
+- 精确 Trace 中 `data.session.agentFlow = AI_OPS`
+- `data.session.answer` 包含最终告警报告
+- `data.toolInvocations` 包含证据工具调用
+
+查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
+
+```powershell
+Invoke-RestMethod `
+ -Method Get `
+ -Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
+```
+
+## 7. Demo 主线
+
+Chat 主线:
+
+```text
+一个 session id + 一个 run id
+-> 用户问题
+-> 多 Agent 执行
+-> 证据工具
+-> Verifier / self_evaluation
+-> 最终答案
+-> 用户反馈
+-> Trace API 回放
+```
+
+AIOps 主线:
+
+```text
+一个 session id + 一个 run id
+-> 告警 payload
+-> AIOps Planner / Executor
+-> 证据工具
+-> 告警分析报告
+-> AIOps rule evaluation
+-> Trace API 回放
+```
+
+## 8. Evidence Pipeline 场景矩阵
+
+面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
+
+- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
+- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
+- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
+
+这样可以同时展示真实链路和确定性回归能力。
diff --git a/mvp/demo/archive/2026-07-22-legacy/_ARCHIVE_NOTE.md b/mvp/demo/archive/2026-07-22-legacy/_ARCHIVE_NOTE.md
new file mode 100644
index 0000000..cf3bd15
--- /dev/null
+++ b/mvp/demo/archive/2026-07-22-legacy/_ARCHIVE_NOTE.md
@@ -0,0 +1,3 @@
+# Archive Note
+
+本目录保存旧多角色与旧诊断入口 demo。当前 demo 以 `mvp/demo/README.md` 和唯一 `/api/chat` named SSE 为准。
diff --git a/mvp/demo/aiops-alert-acceptance.md b/mvp/demo/archive/2026-07-22-legacy/aiops-alert-acceptance.md
similarity index 100%
rename from mvp/demo/aiops-alert-acceptance.md
rename to mvp/demo/archive/2026-07-22-legacy/aiops-alert-acceptance.md
diff --git a/mvp/demo/evidence-pipeline-scenarios.md b/mvp/demo/archive/2026-07-22-legacy/evidence-pipeline-scenarios.md
similarity index 100%
rename from mvp/demo/evidence-pipeline-scenarios.md
rename to mvp/demo/archive/2026-07-22-legacy/evidence-pipeline-scenarios.md
diff --git a/mvp/demo/interview-q-and-a.md b/mvp/demo/archive/2026-07-22-legacy/interview-q-and-a.md
similarity index 100%
rename from mvp/demo/interview-q-and-a.md
rename to mvp/demo/archive/2026-07-22-legacy/interview-q-and-a.md
diff --git a/mvp/demo/interview-walkthrough.md b/mvp/demo/archive/2026-07-22-legacy/interview-walkthrough.md
similarity index 100%
rename from mvp/demo/interview-walkthrough.md
rename to mvp/demo/archive/2026-07-22-legacy/interview-walkthrough.md
diff --git a/mvp/demo/scripts/run-interview-demo-check.ps1 b/mvp/demo/archive/2026-07-22-legacy/run-interview-demo-check.ps1
similarity index 100%
rename from mvp/demo/scripts/run-interview-demo-check.ps1
rename to mvp/demo/archive/2026-07-22-legacy/run-interview-demo-check.ps1
diff --git a/mvp/demo/ten-minute-interview-demo.md b/mvp/demo/archive/2026-07-22-legacy/ten-minute-interview-demo.md
similarity index 100%
rename from mvp/demo/ten-minute-interview-demo.md
rename to mvp/demo/archive/2026-07-22-legacy/ten-minute-interview-demo.md
diff --git a/mvp/demo/archive/2026-07-22-legacy/trace-inspection-checklist.md b/mvp/demo/archive/2026-07-22-legacy/trace-inspection-checklist.md
new file mode 100644
index 0000000..5ebc134
--- /dev/null
+++ b/mvp/demo/archive/2026-07-22-legacy/trace-inspection-checklist.md
@@ -0,0 +1,58 @@
+# Trace 检查清单
+
+运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
+
+## 1. Session
+
+| JSON path | 检查点 | 面试讲点 |
+|---|---|---|
+| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
+| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
+| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
+| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
+| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
+| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
+| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
+| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
+| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
+
+## 2. Agent 步骤
+
+| JSON path | 检查点 | 面试讲点 |
+|---|---|---|
+| `data.steps[*].agentName` | 是否有 Planner / Executor / Verifier 或等价步骤 | 流程被拆成可检查的 Agent 步骤 |
+| `data.steps[*].thought` | 是否有高层步骤摘要 | 内部过程可审计,不只看最终文本 |
+| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
+| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
+
+## 3. 工具证据
+
+| JSON path | 检查点 | 面试讲点 |
+|---|---|---|
+| `data.toolInvocations[*].toolName` | 是否包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 | Agent 通过工具收集证据,而不是无依据猜测 |
+| `data.toolInvocations[*].inputParams` | 是否能看到每个工具的入参 | 工具输入可审计、可调试 |
+| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
+| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
+| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
+| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
+| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
+
+## 4. Summary
+
+| JSON path | 检查点 | 面试讲点 |
+|---|---|---|
+| `data.summary.persistedStepCount` | step 行是否持久化 | Trace 来自存储,不是响应内存 |
+| `data.summary.persistedToolCallCount` | tool 行是否持久化 | 工具证据在请求结束后仍可回放 |
+| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
+| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
+
+## 5. 好的结果长什么样
+
+```text
+同一个 session id + run id
+-> 最终答案
+-> 持久化 agent steps
+-> 持久化 evidence tool calls
+-> verifier / self-evaluation
+-> feedback attached to the same run
+```
diff --git a/mvp/demo/output/README.md b/mvp/demo/output/README.md
index adf49a2..26958c0 100644
--- a/mvp/demo/output/README.md
+++ b/mvp/demo/output/README.md
@@ -2,11 +2,13 @@
本目录是本地 Demo 响应的默认输出位置。
-生成文件会被 Git 忽略:
+当前 named SSE Demo 生成以下文件,均被 Git 忽略:
-- `chat-response.json`
+- `chat-sse.txt`
+- `chat-events.json`
- `trace-response.json`
-- `feedback-response.json`
+
+目录中可能存在旧版 Demo 生成的 `chat-response.json`、`feedback-response.json` 或历史 Trace;它们不是当前架构的验收证据。阶段验收必须使用本次 SSE metadata 返回的 exact `session_id + run_id` 重新生成结果。
保留此 README 是为了让目录存在于仓库中。
diff --git a/mvp/demo/payment-timeout-acceptance.md b/mvp/demo/payment-timeout-acceptance.md
index 0befb75..077ceed 100644
--- a/mvp/demo/payment-timeout-acceptance.md
+++ b/mvp/demo/payment-timeout-acceptance.md
@@ -2,41 +2,34 @@
## 1. 目标
-验证 MVP 能诊断支付超时问题,并暴露完整 Trace 供回放。
+验证当前单 Diagnosis Agent 能诊断支付超时问题,并为一次精确 Run 暴露安全、可核对的 Trace。
## 2. 输入
-- Session id:`mvp-demo-payment-timeout-001`
-- 问题:`支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
+- `session_id`:`mvp-demo-payment-timeout-001`
+- 问题:结合知识库和日志证据诊断支付超时,并给出有边界的修复建议
- Profile:`mvp-demo`
## 3. 验收标准
-1. Chat 返回成功答复,且 session id 与请求一致,并返回本次诊断的 run id。
-2. Trace API 使用 `sessionId + runId` 返回会话元数据、运行摘要、最终答案、按顺序排列的 agent steps 和 tool invocations。
-3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
-4. 可以使用同一个 session id 和本次 run id 提交反馈。
-5. 后续精确 Trace 查询能看到已持久化的 feedback 值。
+1. `POST /api/chat` 按 `metadata -> status* -> content|failure -> done` 顺序发送 named SSE;`content` 与 `failure` 必须互斥且只出现一次。
+2. `metadata.session_id` 和 `metadata.run_id` 非空;该精确 ID 对能唯一定位持久化 Run 与 Trace。
+3. Run 的 `intent=DIAGNOSIS`,终态、`release_outcome` 与 SSE `done` outcome 一致。
+4. `agent_step.agent_name` 只出现 `diagnosis_agent`。AgentStep 只保存消息数量/角色、输出是否存在、Tool 名称等 metadata,不保存 Prompt、消息正文、模型正文或 Thought。
+5. 每条 `tool_invocation` 都属于精确 `run_id`,Tool 名称属于 ACI allowlist,只保存有界的身份、状态、错误码、耗时和字节数 metadata;不得包含 SQL、日志查询正文、raw response、凭据或 evidence body。
+6. `content` 只在 Harness guards 与 Release Policy 完成后发布;guard 或技术失败只能发布固定安全 fallback,不能泄漏 Agent 或 Tool 原始 JSON。
+7. `query_logs` 明确标记为 Mock。`query_mysql` 只针对配置的隔离只读数据源验证;本验收不声称接入真实 CLS 或生产业务 MySQL。
## 4. 需要检查的 Trace 字段
-- `data.runId`
-- `data.run.runId`
-- `data.session.query`
-- `data.session.answer`
-- `data.session.selfEvaluation`
-- `data.session.feedback`
-- `data.steps[*].agentName`
-- `data.steps[*].thought`
-- `data.toolInvocations[*].toolName`
-- `data.toolInvocations[*].inputParams`
-- `data.toolInvocations[*].outputPreview`
-- `data.toolInvocations[*].retrievalDetails`
+- `data.runId` 与 `data.run.runId`
+- `data.session.sessionId`、`data.session.intent`、`data.session.releaseOutcome`
+- `data.steps[*].agentName`、`data.steps[*].thought` 与 metadata 字段
+- `data.toolInvocations[*].toolName`、identity、status、error code、duration 与 size 字段
- `data.summary`
## 5. 已知边界
-- 这不是完整离线测试,仍需要有效的 chat、持久化、向量检索和模型调用环境。
-- `mvp-demo` profile 启用 mock 日志和指标,让证据工具返回更稳定。
-- 敏感配置清理不属于当前 MVP 优先级。
-
+- 这是一次真实应用 E2E 验收,不能替代确定性的单元与契约测试。
+- 外部模型、Redis、Milvus 和隔离数据源的可用性可能影响 live Run;失败时必须记录精确 session/run identity。
+- 敏感配置治理和生产 CLS/业务 MySQL 接入不属于本次 MVP 验收范围。
diff --git a/mvp/demo/requests/payment-timeout-chat.json b/mvp/demo/requests/payment-timeout-chat.json
index 1b22113..8c80fc4 100644
--- a/mvp/demo/requests/payment-timeout-chat.json
+++ b/mvp/demo/requests/payment-timeout-chat.json
@@ -1,4 +1,4 @@
{
"Id": "mvp-demo-payment-timeout-001",
- "Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
+ "Question": "支付接口最近出现超时。在给出结论前必须实际调用 lookup_knowledge 和 query_logs 各一次,日志范围使用 APPLICATION 并查询 payment-service error slow database;随后结束诊断,证据不足时明确说明缺口。"
}
diff --git a/mvp/demo/scripts/run-payment-timeout-demo.ps1 b/mvp/demo/scripts/run-payment-timeout-demo.ps1
index c963fd5..bcdc0bf 100644
--- a/mvp/demo/scripts/run-payment-timeout-demo.ps1
+++ b/mvp/demo/scripts/run-payment-timeout-demo.ps1
@@ -7,55 +7,122 @@ param(
$ErrorActionPreference = "Stop"
+function ConvertFrom-NamedSse {
+ param([Parameter(Mandatory = $true)][string]$Content)
+
+ $events = @()
+ foreach ($frame in ($Content -split "(?:\r?\n){2,}")) {
+ if ([string]::IsNullOrWhiteSpace($frame)) {
+ continue
+ }
+
+ $name = $null
+ $dataLines = @()
+ foreach ($line in ($frame -split "\r?\n")) {
+ if ($line.StartsWith("event:")) {
+ $name = $line.Substring(6).Trim()
+ } elseif ($line.StartsWith("data:")) {
+ $dataLines += $line.Substring(5).TrimStart()
+ }
+ }
+
+ if ([string]::IsNullOrWhiteSpace($name) -or $dataLines.Count -eq 0) {
+ throw "Invalid named SSE frame: $frame"
+ }
+ $rawData = $dataLines -join "`n"
+ $events += [pscustomobject]@{
+ name = $name
+ payload = $rawData | ConvertFrom-Json
+ }
+ }
+ return @($events)
+}
+
+function Assert-ChatSseContract {
+ param(
+ [Parameter(Mandatory = $true)][array]$Events,
+ [Parameter(Mandatory = $true)][string]$ExpectedSessionId
+ )
+
+ if ($Events.Count -lt 3) {
+ throw "Chat SSE must contain metadata, a terminal event, and done"
+ }
+ if ($Events[0].name -ne "metadata") {
+ throw "First Chat SSE event must be metadata"
+ }
+ if ($Events[-1].name -ne "done") {
+ throw "Last Chat SSE event must be done"
+ }
+
+ $terminalEvents = @($Events | Where-Object { $_.name -in @("content", "failure") })
+ if ($terminalEvents.Count -ne 1 -or $Events[-2].name -ne $terminalEvents[0].name) {
+ throw "Chat SSE must contain exactly one content or failure immediately before done"
+ }
+
+ $allowed = @("metadata", "status", "content", "failure", "done")
+ $unknown = @($Events | Where-Object { $_.name -notin $allowed })
+ if ($unknown.Count -gt 0) {
+ throw "Chat SSE contains unknown events: $($unknown.name -join ', ')"
+ }
+ $invalidMiddle = @()
+ if ($Events.Count -gt 3) {
+ $invalidMiddle = @($Events[1..($Events.Count - 3)] |
+ Where-Object { $_.name -ne "status" })
+ }
+ if ($invalidMiddle.Count -gt 0) {
+ throw "Only status events are allowed between metadata and the terminal event"
+ }
+
+ $metadata = $Events[0].payload
+ if ($metadata.session_id -ne $ExpectedSessionId) {
+ throw "SSE session_id does not match the requested SessionId"
+ }
+ if ([string]::IsNullOrWhiteSpace([string]$metadata.run_id)) {
+ throw "SSE metadata is missing run_id"
+ }
+
+ $outcome = [string]$Events[-1].payload.outcome
+ if ($terminalEvents[0].name -eq "failure" -and $outcome -ne "FAILED") {
+ throw "A failure event must end with outcome FAILED"
+ }
+ if ($terminalEvents[0].name -eq "content" -and $outcome -notin @("SUCCESS", "FALLBACK")) {
+ throw "A content event must end with outcome SUCCESS or FALLBACK"
+ }
+}
+
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
$request.Id = $SessionId
$body = $request | ConvertTo-Json -Depth 8
-Write-Host "正在运行支付超时 Chat 诊断 Demo..."
+Write-Host "Running payment-timeout Chat E2E"
Write-Host "BaseUrl: $BaseUrl"
Write-Host "SessionId: $SessionId"
-$chat = Invoke-RestMethod `
+$response = Invoke-WebRequest `
+ -UseBasicParsing `
-Method Post `
-Uri "$BaseUrl/api/chat" `
+ -Headers @{ Accept = "text/event-stream" } `
-ContentType "application/json; charset=utf-8" `
-Body $body
-$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
-Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
+$response.Content | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-sse.txt"
+$events = @(ConvertFrom-NamedSse -Content $response.Content)
+Assert-ChatSseContract -Events $events -ExpectedSessionId $SessionId
+$events | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-events.json"
-$runId = $chat.data.runId
-if (-not $runId) {
- throw "Chat 响应缺少 runId,无法查询精确 Trace。"
-}
+$runId = [string]$events[0].payload.run_id
Write-Host "RunId: $runId"
+Write-Host "SSE sequence: $($events.name -join ' -> ')"
$trace = Invoke-RestMethod `
-Method Get `
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
-
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
-Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
-$feedbackBody = @{
- sessionId = $SessionId
- runId = $runId
- feedback = "useful"
-} | ConvertTo-Json
-
-$feedback = Invoke-RestMethod `
- -Method Post `
- -Uri "$BaseUrl/api/feedback" `
- -ContentType "application/json; charset=utf-8" `
- -Body $feedbackBody
-
-$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
-Write-Host "已保存反馈响应: $OutputDir/feedback-response.json"
-
-Write-Host ""
-Write-Host "Demo 已完成,请检查:"
-Write-Host "- mvp/demo/output/chat-response.json"
-Write-Host "- mvp/demo/output/trace-response.json"
-Write-Host "- mvp/demo/output/feedback-response.json"
+Write-Host "E2E artifacts:"
+Write-Host "- $OutputDir/chat-sse.txt"
+Write-Host "- $OutputDir/chat-events.json"
+Write-Host "- $OutputDir/trace-response.json"
diff --git a/mvp/demo/trace-inspection-checklist.md b/mvp/demo/trace-inspection-checklist.md
index 5ebc134..cf61398 100644
--- a/mvp/demo/trace-inspection-checklist.md
+++ b/mvp/demo/trace-inspection-checklist.md
@@ -1,58 +1,12 @@
# Trace 检查清单
-运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
-
-## 1. Session
-
-| JSON path | 检查点 | 面试讲点 |
-|---|---|---|
-| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
-| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
-| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
-| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
-| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
-| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
-| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
-| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
-| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
-
-## 2. Agent 步骤
-
-| JSON path | 检查点 | 面试讲点 |
-|---|---|---|
-| `data.steps[*].agentName` | 是否有 Planner / Executor / Verifier 或等价步骤 | 流程被拆成可检查的 Agent 步骤 |
-| `data.steps[*].thought` | 是否有高层步骤摘要 | 内部过程可审计,不只看最终文本 |
-| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
-| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
-
-## 3. 工具证据
-
-| JSON path | 检查点 | 面试讲点 |
-|---|---|---|
-| `data.toolInvocations[*].toolName` | 是否包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 | Agent 通过工具收集证据,而不是无依据猜测 |
-| `data.toolInvocations[*].inputParams` | 是否能看到每个工具的入参 | 工具输入可审计、可调试 |
-| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
-| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
-| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
-| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
-| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
-
-## 4. Summary
-
-| JSON path | 检查点 | 面试讲点 |
-|---|---|---|
-| `data.summary.persistedStepCount` | step 行是否持久化 | Trace 来自存储,不是响应内存 |
-| `data.summary.persistedToolCallCount` | tool 行是否持久化 | 工具证据在请求结束后仍可回放 |
-| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
-| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
-
-## 5. 好的结果长什么样
-
-```text
-同一个 session id + run id
--> 最终答案
--> 持久化 agent steps
--> 持久化 evidence tool calls
--> verifier / self-evaluation
--> feedback attached to the same run
-```
+- [ ] SSE metadata 的 `session_id`、`run_id` 非空且与数据库完全一致。
+- [ ] `diagnosis_run.intent=DIAGNOSIS`,status/release_outcome 与 done outcome 一致。
+- [ ] `agent_step.agent_name` 只出现 `diagnosis_agent`。
+- [ ] AgentStep model_input/model_output 只含 metadata,thought 为空。
+- [ ] ToolInvocation 全部属于 exact runId,Tool 名在 ACI allowlist 内。
+- [ ] ToolInvocation input/output/retrieval details 不含 SQL、日志 query、raw response 或 evidence body。
+- [ ] Run 的模型/Tool/Token/字节预算均未超过集中配置。
+- [ ] content 只出现一次并来自 Release Policy;failure 与 content 互斥。
+- [ ] `logs/application.log` 不含 Prompt、Thought、完整 Tool 参数、raw response、vendor exception 或 stack 泄漏。
+- [ ] query_logs 标记为 Mock;query_mysql 只使用隔离只读 datasource contract。
diff --git a/mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md b/mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md
index 064cac5..7c9a22c 100644
--- a/mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md
+++ b/mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md
@@ -1,6 +1,6 @@
# ISS-014 单体 ReAct Agent、Harness 与 ACI 工具瘦身
-**状态**:实施中(阶段 0-6B 已归档,下一阶段 7)
+**状态**:已完成(阶段 0-7 已验收并归档)
**严重程度**:高
**发现时间**:2026-07-20
**目标分支**:`refactor/chat-single-react-harness`
@@ -1307,27 +1307,27 @@ ISS-014 是总设计 Issue,不创建跨阶段共享的 OpenSpec change。以
## 14. 总体验收标准
-- [ ] 复杂诊断只存在一个拥有工具循环的 Diagnosis ReAct Agent。
-- [ ] 不存在业务 StateGraph 或 Planner/Executor/Composer 多 Agent 主链路。
-- [ ] Harness 不承担业务推理,不演变为工作流引擎。
-- [ ] EvidenceGuard 是 Harness 内的确定性能力,不是独立编排节点。
-- [ ] SemanticGuard 使用完全隔离上下文,无工具、无记忆、无回调循环。
-- [ ] SemanticGuard 不使用未校准数值置信度控制在线释放。
-- [ ] 所有 Agent-facing evidence Tool 符合 ACI 状态和 Tool Call ID 契约。
-- [ ] RAG 不再向 Agent 返回 ContextPack/Trace/Rerank 等审计数据。
-- [ ] query_logs 不要求 Agent 先调用 Topic discovery,且返回聚合、抽样、脱敏结果。
-- [ ] MySQL Tool 只读、安全解析、参数绑定、allowlist、超时和结果上限全部生效。
-- [ ] Tool 原始结果不会未经有界投影进入 Agent 上下文;Redis canonical evidence 当前可暂不脱敏,但不得被 Agent 直接读取。
-- [ ] Redis 每次 Tool Call 单 Key 保存,状态、TTL、容量、ACL 和日志禁泄漏规则均有测试。
-- [ ] 每个 Tool Call ID 可以按 exact runId 和调用记录状态验真,并取得对应 `agent_result`。
-- [ ] 只保留一个 `/api/chat` SSE 接口。
-- [ ] SSE 不输出 Thought、Prompt、原始 Tool 载荷和未验证结论。
-- [ ] 客户端断开、模型/Tool/SemanticGuard 超时均有明确取消和 Run 终态。
-- [ ] 所有重试由 Harness 按类型化策略装配和记录,无 SDK/HTTP/数据库隐藏重试或整个 Diagnosis Agent 重跑。
-- [ ] 最终 E2E 能按 sessionId/runId 对齐 SSE、日志、AgentStep、ToolInvocation 和最终答案。
-- [ ] Token、工具调用、Tool 投影、Redis TTL 和总延迟预算均来自集中配置,并有可验证的强制上限和耗尽原因。
-- [ ] 11 个 OpenSpec changes 均已独立 Archive,并分别对应一个范围清晰的 Git commit。
-- [ ] 不保留旧兼容分支、注释代码、本地 refs 卸载和硬编码凭据。
+- [x] 复杂诊断只存在一个拥有工具循环的 Diagnosis ReAct Agent。
+- [x] 不存在业务 StateGraph 或 Planner/Executor/Composer 多 Agent 主链路。
+- [x] Harness 不承担业务推理,不演变为工作流引擎。
+- [x] EvidenceGuard 是 Harness 内的确定性能力,不是独立编排节点。
+- [x] SemanticGuard 使用完全隔离上下文,无工具、无记忆、无回调循环。
+- [x] SemanticGuard 不使用未校准数值置信度控制在线释放。
+- [x] 所有 Agent-facing evidence Tool 符合 ACI 状态和 Tool Call ID 契约。
+- [x] RAG 不再向 Agent 返回 ContextPack/Trace/Rerank 等审计数据。
+- [x] query_logs 不要求 Agent 先调用 Topic discovery,且返回聚合、抽样、脱敏结果。
+- [x] MySQL Tool 只读、安全解析、参数绑定、allowlist、超时和结果上限全部生效。
+- [x] Tool 原始结果不会未经有界投影进入 Agent 上下文;Redis canonical evidence 当前可暂不脱敏,但不得被 Agent 直接读取。
+- [x] Redis 每次 Tool Call 单 Key 保存,状态、TTL、容量、ACL 和日志禁泄漏规则均有测试。
+- [x] 每个 Tool Call ID 可以按 exact runId 和调用记录状态验真,并取得对应 `agent_result`。
+- [x] 只保留一个 `/api/chat` SSE 接口。
+- [x] SSE 不输出 Thought、Prompt、原始 Tool 载荷和未验证结论。
+- [x] 客户端断开、模型/Tool/SemanticGuard 超时均有明确取消和 Run 终态。
+- [x] 所有重试由 Harness 按类型化策略装配和记录,无 SDK/HTTP/数据库隐藏重试或整个 Diagnosis Agent 重跑。
+- [x] 最终 E2E 能按 sessionId/runId 对齐 SSE、日志、AgentStep、ToolInvocation 和最终答案。
+- [x] Token、工具调用、Tool 投影、Redis TTL 和总延迟预算均来自集中配置,并有可验证的强制上限和耗尽原因。
+- [x] 11 个 OpenSpec changes 均已独立 Archive,并分别对应一个范围清晰的 Git commit。
+- [x] 不保留旧兼容分支、注释代码、本地 refs 卸载和硬编码凭据。
## 15. 非目标
diff --git a/mvp/issues/active/ISS-012-executor-token-budget-and-context-growth.md b/mvp/issues/archived/ISS-012-executor-token-budget-and-context-growth.md
similarity index 95%
rename from mvp/issues/active/ISS-012-executor-token-budget-and-context-growth.md
rename to mvp/issues/archived/ISS-012-executor-token-budget-and-context-growth.md
index d73d53f..0d48a70 100644
--- a/mvp/issues/active/ISS-012-executor-token-budget-and-context-growth.md
+++ b/mvp/issues/archived/ISS-012-executor-token-budget-and-context-growth.md
@@ -1,6 +1,7 @@
# ISS-012 Executor Token 预算与上下文膨胀
-**状态**:待规划
+**状态**:已被 ISS-014 吸收并归档(2026-07-22)
+**吸收结果**:单 Diagnosis Agent、Harness 集中预算、ACI Tool projection、exact Run audit 与安全 Fallback 已替代本 Issue 的旧 Executor 方案;最终 live 数据见 ISS-014 阶段 7 验收。
**严重程度**:高
**发现时间**:2026-07-20
**关联**:ISS-002、ISS-004、ISS-011
diff --git a/mvp/issues/active/ISS-013-chat-entry-decoupling-and-sse.md b/mvp/issues/archived/ISS-013-chat-entry-decoupling-and-sse.md
similarity index 94%
rename from mvp/issues/active/ISS-013-chat-entry-decoupling-and-sse.md
rename to mvp/issues/archived/ISS-013-chat-entry-decoupling-and-sse.md
index f32f5d3..205cff0 100644
--- a/mvp/issues/active/ISS-013-chat-entry-decoupling-and-sse.md
+++ b/mvp/issues/archived/ISS-013-chat-entry-decoupling-and-sse.md
@@ -1,6 +1,7 @@
# ISS-013 Chat 入口解耦与真正 SSE 收敛
-**状态**:待规划
+**状态**:已被 ISS-014 吸收并归档(2026-07-22)
+**吸收结果**:唯一 `/api/chat` named SSE、Chat Application Use Case、bounded executor、exact Run cancel 和前端单 consumer 已由 ISS-014 阶段 6A/6B 完成,最终 E2E 归入阶段 7。
**严重程度**:高
**发现时间**:2026-07-20
**关联**:ISS-011、ISS-012
diff --git a/mvp/tables/Agent步骤表-agent_step.md b/mvp/tables/Agent步骤表-agent_step.md
index 51cde2c..d7fbacf 100644
--- a/mvp/tables/Agent步骤表-agent_step.md
+++ b/mvp/tables/Agent步骤表-agent_step.md
@@ -5,7 +5,7 @@
## 定位
-`agent_step` 记录一次诊断运行中每个 Agent 步骤的模型输入、输出、耗时和 Token 消耗。`run_id` 是执行隔离边界;Trace 页面展示顺序以 Trace API 返回顺序为准。
+`agent_step` 记录 Diagnosis Agent 模型步骤的有界审计 metadata。`run_id` 是执行隔离边界;当前写入不得保存 Prompt、消息正文、模型正文、Tool arguments 或 Thought。
## 字段
@@ -15,10 +15,10 @@
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
| `step_index` | INT | 是 | 步骤序号,从 0 开始 |
-| `agent_name` | VARCHAR(32) | 是 | Agent 名称,例如 planner、executor、verifier、composer |
-| `model_input` | TEXT | 否 | 模型输入摘要;`V006` 已从 JSON 改为 TEXT |
-| `model_output` | TEXT | 否 | 模型输出摘要;`V006` 已从 JSON 改为 TEXT |
-| `thought` | TEXT | 否 | Agent 思考过程或调试摘要 |
+| `agent_name` | VARCHAR(32) | 是 | 当前 Harness 写入固定为 `diagnosis_agent` |
+| `model_input` | TEXT | 否 | JSON metadata,仅包含 message count 与 roles |
+| `model_output` | TEXT | 否 | JSON metadata,仅包含 text presence 与 Tool names |
+| `thought` | TEXT | 否 | 当前 Harness 必须写空;字段仅保留历史兼容 |
| `has_tool_call` | BOOLEAN | 否 | 本步骤是否触发工具调用 |
| `duration_ms` | INT | 否 | 本步骤耗时 |
| `token_count` | INT | 否 | 本步骤 Token 消耗 |
@@ -41,5 +41,5 @@
## 注意点
- 前端展示步骤时应使用 Trace API 返回顺序;服务端会在同一 `run_id` 范围内整理步骤顺序。
-- 新 Trace、Verifier 和评测读路径应按 `run_id` 取数,避免同一 `sessionId` 多轮诊断混入。
-- Verifier 应在 Executor 循环完成后出现;如果 `step_index` 中 Verifier 提前,通常意味着编排或记录顺序有问题。
+- 新 Trace 和验收读路径必须按 exact `run_id` 取数,避免同一 `sessionId` 多次运行混入。
+- 当前 Run 若出现 `diagnosis_agent` 之外的新写入,或 `thought` 非空,视为审计边界违规。
diff --git a/mvp/tables/README.md b/mvp/tables/README.md
index ac5f147..8f9782a 100644
--- a/mvp/tables/README.md
+++ b/mvp/tables/README.md
@@ -9,10 +9,10 @@
| 表 | 用途 | 文档 |
|---|---|---|
-| `chat_session` | 会话目录元数据,保存同一个 `sessionId` 的多轮会话状态快照 | [聊天会话表-chat_session.md](聊天会话表-chat_session.md) |
-| `diagnosis_run` | 运行级主记录,保存一次 Chat/AIOps 诊断的 query、状态、答案、自评估和反馈 | [诊断运行表-diagnosis_run.md](诊断运行表-diagnosis_run.md) |
-| `agent_step` | Agent 步骤记录,按 `run_id` 隔离回放执行链路 | [Agent步骤表-agent_step.md](Agent步骤表-agent_step.md) |
-| `tool_invocation` | 工具调用记录,按 `run_id` 支撑 Trace、Verifier 和评测 | [工具调用表-tool_invocation.md](工具调用表-tool_invocation.md) |
+| `chat_session` | Chat 会话目录 metadata;不保存完整消息历史 | [聊天会话表-chat_session.md](聊天会话表-chat_session.md) |
+| `diagnosis_run` | `/api/chat` 运行主记录,保存 intent、终态与安全发布结果 | [诊断运行表-diagnosis_run.md](诊断运行表-diagnosis_run.md) |
+| `agent_step` | Diagnosis Agent metadata-only 模型步骤审计 | [Agent步骤表-agent_step.md](Agent步骤表-agent_step.md) |
+| `tool_invocation` | Harness ToolBoundary metadata-only 长期审计 | [工具调用表-tool_invocation.md](工具调用表-tool_invocation.md) |
| `api_document` | 知识库文档元数据,和向量库 chunk 通过 `doc_id` 关联 | [文档元数据表-api_document.md](文档元数据表-api_document.md) |
| `knowledge_domain` | 知识域元数据,支撑 RAG domain hint 和检索策略 | [知识域表-knowledge_domain.md](知识域表-knowledge_domain.md) |
| `case_library` | 用户反馈沉淀出的高质量诊断案例 | [案例库表-case_library.md](案例库表-case_library.md) |
diff --git a/mvp/tables/工具调用表-tool_invocation.md b/mvp/tables/工具调用表-tool_invocation.md
index adb1499..8eeb34a 100644
--- a/mvp/tables/工具调用表-tool_invocation.md
+++ b/mvp/tables/工具调用表-tool_invocation.md
@@ -5,7 +5,7 @@
## 定位
-`tool_invocation` 记录 Agent 在一次诊断运行中显式调用工具的事实,包括工具名、入参、输出摘要、检索层级、证据引用和失败信息。它是 Trace、Verifier、评测和人工排查的共同数据源。
+`tool_invocation` 是 Harness ToolBoundary 的长期 metadata-only 审计表。它记录 exact Run/Tool identity、状态、稳定错误码、耗时与字节数;完整请求、raw response 和 Agent projection 只短期存在于 Redis canonical invocation,不写入本表。
## 字段
@@ -15,20 +15,20 @@
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
| `step_id` | BIGINT | 否 | 可关联 `agent_step.id` |
-| `tool_name` | VARCHAR(64) | 是 | 工具名称,例如 `lookup_knowledge`、日志查询、指标查询 |
-| `input_params` | JSON | 是 | 工具入参 |
-| `output_preview` | TEXT | 否 | 工具输出摘要或前缀 |
-| `output_length` | INT | 否 | 工具输出字符数 |
-| `retrieval_layer` | VARCHAR(8) | 否 | 检索层级,例如 `L0`、`L1`、`L0+L1` |
+| `tool_name` | VARCHAR(64) | 是 | ACI Tool 名:`lookup_knowledge`、`query_logs` 或 `query_mysql` |
+| `input_params` | JSON | 是 | 仅 `tool_call_id` 与 `request_bytes` metadata,不含 Tool 参数正文 |
+| `output_preview` | TEXT | 否 | 仅 invocation/evidence status metadata |
+| `output_length` | INT | 否 | Agent projection UTF-8 字节数 |
+| `retrieval_layer` | VARCHAR(8) | 否 | 当前 Harness 审计固定为 `HARNESS` |
| `l0_match_count` | INT | 否 | L0 命中数量 |
| `l1_match_count` | INT | 否 | L1 命中数量 |
| `is_truncated` | BOOLEAN | 否 | 输出是否被截断 |
-| `relevance_level` | VARCHAR(20) | 否 | 归一化质量等级:`PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`、`DEDUPED` |
-| `dedup_reason` | VARCHAR(32) | 否 | 去重原因,例如 `doc_retrieved`、`domain_retrieved` |
-| `retrieval_details` | JSON | 否 | 检索明细、证据引用、Gatekeeper 可用导航信息 |
+| `relevance_level` | VARCHAR(20) | 否 | 当前 Harness 复用该字段保存 evidence status |
+| `dedup_reason` | VARCHAR(32) | 否 | 历史字段;当前 Harness 不写入 |
+| `retrieval_details` | JSON | 否 | `tool_call_id`、status、evidence status、result bytes 与可选稳定错误码 |
| `duration_ms` | INT | 否 | 工具耗时 |
| `success` | BOOLEAN | 否 | 工具是否成功 |
-| `error_message` | TEXT | 否 | 失败原因 |
+| `error_message` | TEXT | 否 | 仅稳定错误码,不保存内部异常或 vendor message |
| `created_at` | DATETIME | 是 | 创建时间 |
## 索引
@@ -36,7 +36,7 @@
| 索引 | 字段 | 用途 |
|---|---|---|
| `idx_session_id` | `session_id` | 历史兼容和粗粒度排查 |
-| `idx_tool_invocation_run_id` | `run_id, id` | Trace、Verifier、评测按运行查询工具调用 |
+| `idx_tool_invocation_run_id` | `run_id, id` | Trace 与验收按 exact Run 查询 Tool 审计 |
| `idx_tool_name` | `tool_name` | 按工具类型排查 |
| `idx_retrieval_layer` | `retrieval_layer` | 观察 RAG L0/L1 行为 |
@@ -48,22 +48,19 @@
## 关键 JSON
-`retrieval_details` 是扩展字段。当前重要结构包括:
+当前 Harness 写入的 `retrieval_details` 结构为:
```json
{
- "evidence_status": "supported",
- "evidence_refs": [
- {
- "raw_path": "$.logs[0]",
- "text": "工具返回中可核对的最小证据文本"
- }
- ]
+ "tool_call_id": "framework-call-id",
+ "status": "READY",
+ "evidence_status": "EVIDENCE_FOUND",
+ "agent_result_bytes": 512
}
```
## 注意点
-- Verifier 不应只信任 RAG 证据;所有工具只要能提供 `evidence_refs`,都应该进入可校验证据链。
-- `output_preview` 只适合展示和排查,不应被当成完整原始输出。
-- `$.no_evidence` 只代表“本次工具未命中证据”,不能推导为“故障不存在”。
+- 本表不是完整证据真理源,EvidenceGuard 只读取当前 Run 的 Redis canonical invocation。
+- `input_params`、`output_preview` 和 `retrieval_details` 均不得出现 SQL、日志 query、evidence body、凭据或 raw response。
+- audit 写入失败应记录安全 warning,但不能改变已确定的 canonical Tool 结果。
diff --git a/mvp/tables/知识域表-knowledge_domain.md b/mvp/tables/知识域表-knowledge_domain.md
index 978e74d..5242371 100644
--- a/mvp/tables/知识域表-knowledge_domain.md
+++ b/mvp/tables/知识域表-knowledge_domain.md
@@ -5,7 +5,7 @@
## 定位
-`knowledge_domain` 保存知识库领域级元数据,用来帮助 Planner/Executor 判断什么时候检索某一类知识,并为 RAG 的 domain hint、去重和可观测性提供基础信息。
+`knowledge_domain` 保存知识库领域级元数据,为 RAG backend 的 domain hint、检索选择和可观测性提供基础信息;它不是 Agent-facing Tool contract。
## 字段
@@ -33,4 +33,4 @@
## 注意点
- `when_to_retrieve` 是检索策略提示,不是事实证据。
-- Executor / Verifier 不能把领域描述当作诊断结论依据;事实仍应来自工具返回的证据块或证据引用。
+- Diagnosis Agent 与 SemanticGuard 不能把领域描述当作诊断结论依据;事实仍应来自当前 Run 经 EvidenceGuard 验真的 Tool evidence。
diff --git a/mvp/tables/聊天会话表-chat_session.md b/mvp/tables/聊天会话表-chat_session.md
index 38c36d3..8cc8557 100644
--- a/mvp/tables/聊天会话表-chat_session.md
+++ b/mvp/tables/聊天会话表-chat_session.md
@@ -5,7 +5,7 @@
## 定位
-`chat_session` 保存多轮 Chat 会话的元数据,用于把同一个 `sessionId` 下的多次诊断运行组织在一起。它不保存完整对话历史;正文消息仍由 Redis `SessionContext.messageHistory` 管理。
+`chat_session` 保存 Chat 会话目录元数据,用于把同一个 `sessionId` 下的多次运行组织在一起。它不保存完整对话历史;当前多轮只从最近一次安全发布的 `diagnosis_run.published_result` 构造有界 `PreviousTurn`,不再使用 Redis `SessionContext`。
## 字段
@@ -14,10 +14,10 @@
| `id` | BIGINT | 是 | 自增主键 |
| `session_id` | VARCHAR(64) | 是 | 会话目录 ID,外部 API 仍通过它定位会话 |
| `status` | VARCHAR(16) | 否 | `ACTIVE`、`EXPIRED`、`CLOSED` |
-| `message_pair_count` | INT | 否 | Redis 会话中问答轮次数的快照 |
+| `message_pair_count` | INT | 否 | 历史兼容计数;当前运行不依赖它恢复消息正文 |
| `created_at` | DATETIME | 是 | 创建时间 |
| `last_active_at` | DATETIME | 否 | 最近活跃时间 |
-| `expires_at` | DATETIME | 否 | 目录元数据,可为空;Redis 消息历史可独立过期 |
+| `expires_at` | DATETIME | 否 | 会话目录过期元数据,可为空 |
## 索引
@@ -38,3 +38,4 @@
- `chat_session` 是会话元数据,不是诊断执行记录。
- 不要把 query、answer、self_evaluation、feedback 写入该表;这些属于 `diagnosis_run`。
+- 不要从该表或 Redis 恢复完整对话正文;安全追问上下文只来自成功发布的结构化结果。
diff --git a/mvp/tables/诊断会话表-diagnosis_session.md b/mvp/tables/诊断会话表-diagnosis_session.md
index 94d371f..a224943 100644
--- a/mvp/tables/诊断会话表-diagnosis_session.md
+++ b/mvp/tables/诊断会话表-diagnosis_session.md
@@ -5,7 +5,7 @@
## 定位
-`diagnosis_session` 是旧版 session 级诊断主记录。`V011` 之后,新 Chat/AIOps 执行的运行态写入已经切到 `chat_session + diagnosis_run`;本表保留用于历史兼容、迁移回填和回滚比较。
+`diagnosis_session` 是旧版 session 级诊断主记录。`V011` 之后,当前 `/api/chat` 执行的运行态写入已经切到 `chat_session + diagnosis_run`;本表保留用于历史兼容、迁移回填和回滚比较。
## 字段
diff --git a/mvp/tables/诊断运行表-diagnosis_run.md b/mvp/tables/诊断运行表-diagnosis_run.md
index fb462b5..ee59103 100644
--- a/mvp/tables/诊断运行表-diagnosis_run.md
+++ b/mvp/tables/诊断运行表-diagnosis_run.md
@@ -5,7 +5,7 @@
## 定位
-`diagnosis_run` 表示一次可回放的 Chat 或 AIOps 诊断执行。`run_id` 是运行级边界,Trace、反馈、自评估、案例沉淀和统计都应优先按 `run_id` 绑定。
+`diagnosis_run` 表示一次 `/api/chat` 应用执行。`run_id` 是运行级边界,SSE、Trace、AgentStep 与 ToolInvocation 必须按 metadata 返回的 exact `run_id` 绑定。
## 字段
@@ -14,11 +14,14 @@
| `id` | BIGINT | 是 | 自增主键 |
| `run_id` | VARCHAR(64) | 是 | 运行唯一 ID,格式为 `run-` + UUID |
| `session_id` | VARCHAR(64) | 是 | 所属 `chat_session.session_id` |
-| `query` | TEXT | 是 | 本次 Chat 问题或 AIOps 告警摘要 |
-| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FAILED` |
-| `agent_flow` | VARCHAR(32) | 否 | `CHAT` 或 `AI_OPS` |
-| `answer` | LONGTEXT | 否 | 本次运行的最终答复或告警报告 |
-| `self_evaluation` | JSON | 否 | 本次运行的 rule、verifier、aiops 自评估容器 |
+| `query` | TEXT | 是 | 本次 Chat 用户问题 |
+| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` |
+| `agent_flow` | VARCHAR(32) | 否 | 历史兼容字段;当前公开执行统一来自 Chat Harness |
+| `answer` | LONGTEXT | 否 | 安全发布的最终文本兼容字段 |
+| `intent` | VARCHAR(32) | 否 | `SYSTEM_CHAT`、`KNOWLEDGE_QUERY` 或 `DIAGNOSIS` |
+| `release_outcome` | VARCHAR(16) | 否 | `SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` |
+| `published_result` | JSON | 否 | Release Policy 允许发布的结构化安全结果 |
+| `self_evaluation` | JSON | 否 | 历史兼容字段;当前 Harness 不写入旧 verifier/AiOps 结构 |
| `feedback` | VARCHAR(16) | 否 | 本次运行的用户反馈 |
| `total_duration_ms` | INT | 否 | 本次运行总耗时 |
| `total_token_count` | INT | 否 | 本次运行 Token 消耗 |
@@ -35,7 +38,7 @@
| `idx_diagnosis_run_session_created` | `session_id, created_at, id` | session 下最新运行解析和运行列表 |
| `idx_diagnosis_run_session_run` | `session_id, run_id` | exact trace / feedback ownership 校验 |
| `idx_diagnosis_run_status` | `status` | 状态筛选 |
-| `idx_diagnosis_run_agent_flow` | `agent_flow` | 区分 Chat / AIOps |
+| `idx_diagnosis_run_agent_flow` | `agent_flow` | 历史兼容筛选 |
## 关系
@@ -49,3 +52,4 @@
- `GET /api/diagnosis/{sessionId}/trace` 未带 `runId` 时只为兼容解析 latest run;新 demo 和新客户端应传 `runId`。
- latest run 排序使用 `created_at DESC, id DESC`,避免 feedback 或自评估更新 `updated_at` 后改变回放目标。
- 历史 `diagnosis_session` 会被迁移成兼容 run,但旧混合数据不能被还原成真实多轮边界。
+- 当前诊断发布结果以 `release_outcome + published_result` 为准,不得从旧 self-evaluation 推断 Release Policy 结果。
diff --git a/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/.archive-ready b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/.archive-ready
new file mode 100644
index 0000000..395527d
--- /dev/null
+++ b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/.archive-ready
@@ -0,0 +1 @@
+ready
diff --git a/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/.committed b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/.committed
new file mode 100644
index 0000000..d0fe822
--- /dev/null
+++ b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/.committed
@@ -0,0 +1 @@
+committed
diff --git a/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/.openspec.yaml b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/.openspec.yaml
new file mode 100644
index 0000000..7250f8f
--- /dev/null
+++ b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/.openspec.yaml
@@ -0,0 +1,2 @@
+schema: spec-driven
+created: 2026-07-22
diff --git a/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/design.md b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/design.md
new file mode 100644
index 0000000..20500fd
--- /dev/null
+++ b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/design.md
@@ -0,0 +1,124 @@
+## Context
+
+阶段 6B 后,`POST /api/chat` 已经只调用新 `ChatApplicationUseCase`,但仓库仍有四类遗留:一是 bundled frontend 可达的 `/api/ai_ops` Supervisor/Planner/Executor 链;二是只有测试引用的 `ChatService`、Redis Session API、Verifier/Gatekeeper/ThreadLocal;三是被新 Harness adapter 复用但仍带旧 `@Tool` 和 DB recorder 副作用的 RAG/log backend;四是仍描述旧架构的文档和 demo。
+
+新 Harness 已以 Redis canonical invocation 保存当前 Run 完整 Tool 调用并用 EvidenceGuard 验真,但 `ToolBoundary` 没有写长期 `tool_invocation` audit。当前 `AgentLoggingHook` 又会持久化/日志输出模型正文、Tool arguments 和 `thought`,且回退到 Session ThreadLocal。阶段 7 必须同时删除旧链和收紧审计,否则代码结构与最终 E2E 都不能满足 ISS-014。
+
+这是 L4 协议/前端删除和跨模块物理重构。历史数据库表与历史数据保持不变;现有 `chat_session` JPA 元数据仍由 `JpaChatRunStore` 使用,不属于 Redis conversation Session 遗留。
+
+## Goals / Non-Goals
+
+**Goals:**
+
+- 业务代码只剩一个拥有 Tool loop 的 Diagnosis ReAct Agent,公开诊断只剩 `/api/chat`。
+- 删除旧 Chat/AiOps/Session 编排、Hook、ThreadLocal、prompts、Tool compatibility annotations 和过时测试。
+- RAG/log backend 只通过 Harness adapters 调用,不自行写旧 ToolInvocation。
+- AgentStep 与 ToolInvocation durable audit 按 exact Run 持久化有界脱敏 metadata。
+- 文档、issues、OpenSpec strict 状态与最终代码一致。
+- 真实启动应用并完成 SSE、日志、MySQL exact-run E2E。
+
+**Non-Goals:**
+
+- 不改变 Intent Router、Diagnosis Draft、EvidenceGuard、SemanticGuard、Fallback 或 SSE 五事件语义。
+- 不删除历史表/数据,不做 schema destructive migration。
+- 不实现真实 CLS 或生产业务 MySQL datasource live 连接。
+- 不改 RAG 算法、Trace API response schema 或 UI 视觉设计。
+
+## Decisions
+
+### 1. 删除 legacy AiOps 和 Redis Session surface,不迁移第二条诊断链
+
+删除 `AiOpsController`、`AiOpsService`、AIOps DTO/config/rule evaluation、三个 prompts、前端按钮/consumer 和相应 tests。删除 `ChatSessionController`、Redis `SessionManager`/`SessionContext` 和 tests。告警诊断通过 `/api/chat` 提交自然语言或结构化文本,由同一个 Intent Router 进入 Diagnosis。
+
+替代方案是把 `/api/ai_ops` 适配到新 use case;拒绝,因为它保留第二个公开诊断协议和前端双轨,违反 ISS-014 唯一入口与无兼容分支目标。`chat_session` entity/repository 不删除,因为它是当前 Run 目录与 PreviousTurn 生命周期的一部分。
+
+### 2. 旧 Chat orchestration 按闭包删除
+
+删除 `ChatService` 及其 Sequential Planner/Executor/Verifier/Composer tests、旧 prompts、Verifier/Gatekeeper helpers、Skill metadata hook、Token ThreadLocal wrapper、旧 Tool trace summary/recorder。删除顺序为公开 caller -> service/orchestration -> helper/hook/context -> resources/tests,并在每层后编译/引用扫描。
+
+`SelfEvaluationMergeService`、`DiagnosisTraceService`、repositories 和 eval 基础设施若仍有非旧链调用则保留。替代仅取消 Spring annotation 而保留死类;拒绝,因为阶段 7 明确要求物理清理。
+
+### 3. RAG/log 是 backend,不是 Agent-facing Tool
+
+保留 `LookupKnowledgeTool.lookupKnowledge` 的检索管线和 `QueryLogsTools.queryLogs` 的 Mock/边界实现,供 `RagToolAdapter`/`QueryLogsToolAdapter` 调用;移除所有 `@Tool/@ToolParam`、旧 topic discovery contract、Session ThreadLocal、session dedup 和 `ToolInvocationRecorder`。Agent-facing schema/description 只来自 `HarnessEvidenceTools` 与 `AgentToolContracts`。
+
+替代完全重写 RAG/log;拒绝,因为阶段 3B 已验证 adapter/projector contract,阶段 7 只需清除双 contract 与副作用。
+
+### 4. Agent audit 使用 Harness-native metadata-only Hook
+
+新增 Harness audit hook,强制从 `RunnableConfig` 读取 exact sessionId/runId。before-model 只保存 message count/roles;after-model 只保存是否有文本、是否有 Tool call、Tool names 和 duration,不保存 message content、model output text、Tool arguments、Prompt 或 Thought。缺少 metadata 时跳过持久化并记录安全 warning,不回退 ThreadLocal。
+
+替代继续修补 `AgentLoggingHook`;拒绝,因为旧类包含 verifier 特例、Token ThreadLocal、正文日志和 thought persistence,保留会让旧边界继续存在。
+
+### 5. ToolBoundary 通过安全 port 写 durable audit
+
+新增 `ToolInvocationAuditSink` 和不可变 audit event。`ToolBoundary.execute` 在 canonical 结果确定后调用 sink;event 只包含 sessionId、runId、toolCallId、toolName、InvocationStatus、EvidenceStatus、stable error code、duration、request/agent-result byte count。audit 调用 fail-open:JPA 审计失败写 warning,但不改变已经确定的 Tool observation;canonical Redis store 仍按原设计 fail-closed。
+
+JPA adapter 复用现有 `tool_invocation` 表:`input_params` 只保存 tool_call_id/request_bytes,`output_preview` 只保存 status/evidence_status,`output_length` 保存 agent result bytes,`error_message` 只保存 stable error code。禁止 raw response、完整 request、SQL、日志正文和模型内容。
+
+替代让旧 backend recorder 继续双写;拒绝,因为它依赖 ThreadLocal、不同 Tool 各自实现且可能保存大 payload。替代审计失败阻断业务;拒绝,因为 durable audit 不是 canonical evidence truth source。
+
+### 6. 文档与 issue 以当前单链架构为准
+
+重写当前 MVP、Agent、Harness、Trace/Demo 文档中的旧双入口与多 Agent 描述;ISS-012/ISS-013 标记由 ISS-014 吸收并归档。历史架构文档保留在 archive,不篡改历史。
+
+### 7. 最终 E2E 使用 exact metadata identity
+
+启动 Spring Boot 后向 `/api/chat` 发送固定 payment-timeout 诊断,严格解析 named SSE,取得 metadata sessionId/runId 并要求 `content|failure -> done`。检查 `logs/` 不含 Prompt/Thought/raw Tool payload/stack leakage;使用 `scripts/query_mysql.py` 按 exact runId 查询 `diagnosis_run`、`agent_step`、`tool_invocation`,核对 intent/outcome/status、唯一 Diagnosis Agent、Tool 名称/次数、同一 identity 和最终 safe content。
+
+query_logs 使用 Mock,query_mysql 使用配置的隔离 datasource contract;不声称真实 CLS/生产业务库 live。若真实模型选择非 Diagnosis intent,则使用明确诊断 query 重试一次,不伪造数据库结果。
+
+## Module and Ownership Audit
+
+```text
+Browser POST /api/chat
+ -> ChatController / ChatSseSession
+ -> ChatApplicationUseCase
+ -> Intent Router
+ -> System / Knowledge / Diagnosis executor
+ -> Diagnosis Harness
+ -> Diagnosis Agent (only Tool loop)
+ -> HarnessEvidenceTools
+ -> ToolBoundary
+ -> Redis canonical invocation (full, TTL)
+ -> Durable audit sink (bounded metadata)
+ -> EvidenceGuard / SemanticGuard / Release
+ -> named SSE safe content
+ -> diagnosis_run + agent_step + tool_invocation exact-run Trace
+```
+
+- Application owns Session/Run/PreviousTurn;Controller 不拥有业务状态。
+- Core owns budget/cancel/lifecycle;Diagnosis Agent owns diagnosis authorship;Guards own verification/release。
+- Redis canonical invocation owns short-lived complete Tool truth;MySQL ToolInvocation owns long-lived metadata audit。
+- `chat_session` is current JPA run directory;Redis SessionContext is legacy conversation memory and is removed。
+- 最大耦合风险是复用 backend 时旧 annotations/recorder 被 Spring 自动发现;static scan + context startup 覆盖。
+
+## Interface Impact
+
+- Level: L4 breaking HTTP/frontend contract。
+- Removed: `POST /api/ai_ops`、`POST /api/chat/clear`、`GET /api/chat/session/{sessionId}`、`GET /api/chat/session/{sessionId}/runs`、legacy AiOps message-wrapper SSE。
+- Retained: `POST /api/chat` named SSE、Trace/feedback/document/search APIs。
+- Consumer migration: bundled frontend 同 commit 删除 AiOps 按钮和 consumer;外部调用方迁移到 `/api/chat`。
+- Rollback: 整体回滚阶段 7 commit;不单独恢复旧 endpoint 或 old Tool discovery。
+
+## Risks / Trade-offs
+
+- [外部 AiOps caller 失败] -> L4 文档明确迁移到 `/api/chat`,不提供双轨。
+- [删除 bean 导致 context 启动失败] -> compile、focused context test、全量 tests、真实 Spring startup 分层验证。
+- [backend 仍被 ToolCallbackProvider 发现] -> 删除 annotations/imports 并静态扫描旧 Tool 名/description。
+- [audit 泄漏敏感内容] -> typed metadata-only event + serialization tests + DB query preview inspection。
+- [audit 写失败丢 Trace 明细] -> warning + Run 主记录仍完成;验收环境要求 audit row 存在,生产运行可观测 failure。
+- [live 环境依赖不可用] -> 先检查服务与启动日志,分类环境/实现失败;不降低为假 E2E。
+
+## Migration Plan
+
+1. 删除 frontend/Controller/Service 可达旧入口和旧 orchestration 闭包。
+2. 解耦 RAG/log backend,替换 Agent/Tool durable audit。
+3. 删除剩余 dead resources/tests/config,跑 compile/focused/full regression 和 static scans。
+4. 更新 docs/issues/OpenSpec,并修复已知 strict validation。
+5. 启动应用,执行 final SSE/log/DB E2E,记录 exact identity evidence。
+6. 同一 commit 部署所有删除与文档;回滚整体回滚该 commit。
+
+## Open Questions
+
+- None。live 默认值是否需要校准由 E2E 数据决定,但不会改变 architecture direction。
diff --git a/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/proposal.md b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/proposal.md
new file mode 100644
index 0000000..10a783b
--- /dev/null
+++ b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/proposal.md
@@ -0,0 +1,55 @@
+## Why
+
+阶段 0-6B 已把公开 Chat 切换到单一 Diagnosis ReAct Agent + Harness,但仓库仍保留可达的 `/api/ai_ops` Supervisor/Planner/Executor 多 Agent 链、无生产调用方的旧 `ChatService`、Redis Session endpoint、ThreadLocal、Verifier/Gatekeeper Hook、旧 Agent-facing Tool annotations 和过时文档。新 Harness 的 ToolBoundary 也只写 Redis canonical invocation,尚缺最终 E2E 所需的脱敏 `tool_invocation` durable audit。
+
+## What Changes
+
+- 删除 `/api/ai_ops`、bundled frontend AI Ops 按钮/consumer,以及对应 AiOps service、DTO、prompt、rule evaluation 和测试;公开诊断只保留 `/api/chat`。
+- 删除无生产调用方的旧 `ChatService`、Planner/Executor/Verifier/Composer prompts、Gatekeeper/Verifier helpers、旧 session controller/manager/context 和过时测试。
+- 保留 RAG 与 Mock logs 的底层查询能力,但移除旧 `@Tool` contract、Session ThreadLocal、session dedup 和 `ToolInvocationRecorder` 副作用;Diagnosis Agent 只能看到 `HarnessEvidenceTools` 的三项 ACI contract。
+- 用 Harness-native、metadata-only Agent trace hook 替代旧 `AgentLoggingHook`,禁止持久化 Thought、Prompt、Tool arguments 或模型正文。
+- 为 ToolBoundary 增加 fail-open durable audit port 与 JPA adapter,只保存 exact session/run/tool_call identity、工具名、状态、耗时和有界脱敏摘要,不保存 raw response 或完整请求。
+- 删除不再使用的配置、logback logger、资源和测试,修正阶段 3B/3C 已知 OpenSpec strict 格式问题。
+- 更新当前架构、Agent/Harness、Trace、Demo 和 API 文档;将已被 ISS-014 吸收的 ISS-012/ISS-013 归档。
+- 启动真实 Spring Boot 应用,执行最终 Chat SSE E2E,并用 `logs/` 与 `scripts/query_mysql.py` 核对 exact sessionId/runId、Run、AgentStep、ToolInvocation 和安全发布结果。
+
+## Capabilities
+
+### New Capabilities
+
+- `single-react-cleanup-e2e`: 定义旧链物理删除、唯一公开 Chat surface、Harness durable audit、安全 Trace 和最终 live E2E 验收。
+
+### Modified Capabilities
+
+- `single-react-chat-sse-cutover`: 删除阶段 6B 暂时保留的 legacy AiOps 与 Session endpoints,使唯一 Chat surface 成为最终状态。
+
+## Scope
+
+- Controller/service/hook/util/tool/resource/test 的旧架构删除与 Harness-native audit replacement。
+- bundled frontend AI Ops 路径清理。
+- MVP architecture/demo/API/issue 文档收口。
+- OpenSpec strict validation cleanup。
+- Spring Boot live Chat SSE、日志和 MySQL exact-run 验收。
+
+## Non-goals
+
+- 不修改三类 Intent、Diagnosis Agent Draft、EvidenceGuard、SemanticGuard 或 Release Policy 业务语义。
+- 不新增 Graph、工作流 DSL、兼容 endpoint、双轨开关或 Token streaming。
+- 不删除历史数据库表或历史数据,不迁移生产数据。
+- 不声称完成真实 CLS 或生产业务 MySQL datasource live E2E;query_logs 和 query_mysql 继续按 ISS-014 使用 Mock/隔离契约验收。
+- 不重新设计 RAG 检索算法、Trace API 或前端视觉体验。
+
+## Context Constraints
+
+- Diagnosis Agent 是唯一拥有 Tool loop 的业务 Agent;项目业务代码不得保留 Supervisor/Sequential/Planner/Executor/Verifier/Composer 编排。
+- Agent-facing Tools 只能来自 `HarnessEvidenceTools`,底层 RAG/log 实现不是 Agent contract。
+- Durable audit 只能保存有界脱敏元数据;Redis canonical invocation 仍是当前 Run 完整调用记录真理源。
+- live E2E 必须对齐 SSE metadata 的 exact sessionId/runId,不能用“最新一条”替代。
+- L4 endpoint 删除不提供兼容分支;回滚只能整体回滚阶段 7 commit。
+
+## Risks
+
+- 旧类存在隐式 Spring bean 或动态 Tool discovery,删除不完整会继续把旧 Tools 暴露给模型。
+- 旧 Tool recorder 与新 canonical store 并存会产生双写或缺少 exact run 的数据库审计。
+- Agent trace hook 若保存模型正文或 Tool arguments,会违反不持久化 Thought/Prompt/raw payload 的边界。
+- live 环境依赖模型、Redis、Milvus 和 MySQL;必须区分环境失败、实现失败与验收限制。
diff --git a/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/specs/single-react-chat-sse-cutover/spec.md b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/specs/single-react-chat-sse-cutover/spec.md
new file mode 100644
index 0000000..2ca80a5
--- /dev/null
+++ b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/specs/single-react-chat-sse-cutover/spec.md
@@ -0,0 +1,6 @@
+## REMOVED Requirements
+
+### Requirement: AiOps public behavior SHALL remain isolated
+**Reason**: Stage 6B temporarily preserved the legacy `/api/ai_ops` path only to make the Chat SSE cutover atomic. Stage 7 removes the remaining Supervisor/Planner/Executor orchestration so `/api/chat` is the single diagnosis entry required by ISS-014.
+
+**Migration**: Bundled and external consumers MUST submit diagnosis requests to the named-event SSE `POST /api/chat`. There is no compatibility endpoint or legacy message-wrapper parser.
diff --git a/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/specs/single-react-cleanup-e2e/spec.md b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/specs/single-react-cleanup-e2e/spec.md
new file mode 100644
index 0000000..0a223a5
--- /dev/null
+++ b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/specs/single-react-cleanup-e2e/spec.md
@@ -0,0 +1,85 @@
+## ADDED Requirements
+
+### Requirement: Legacy diagnosis orchestration SHALL be physically absent
+Production and test source SHALL contain no legacy ChatService, AiOps Supervisor/Planner/Executor, Sequential Planner/Executor/Verifier/Composer, legacy Gatekeeper/Verifier helpers, ThreadLocal execution context, or their dedicated prompts and obsolete tests. The only business Agent with a Tool loop SHALL be the Diagnosis Agent constructed inside the Harness.
+
+#### Scenario: Source inventory after cleanup
+- **WHEN** production source, resources and tests are scanned
+- **THEN** legacy orchestration classes, prompts, ThreadLocals and self-only tests are absent rather than disabled or commented out
+
+#### Scenario: Agent construction inventory
+- **WHEN** business Agent builders and framework orchestration types are inspected
+- **THEN** only the Diagnosis Agent has evidence Tool callbacks and no Supervisor/Sequential business workflow remains
+
+### Requirement: Public diagnosis SHALL use one endpoint
+The application SHALL expose `POST /api/chat` as the only public diagnosis execution endpoint. `/api/ai_ops`, legacy Chat Session management endpoints and bundled frontend callers for those endpoints MUST be removed without a compatibility branch.
+
+#### Scenario: Bundled frontend diagnosis
+- **WHEN** a user submits a diagnosis from the bundled frontend
+- **THEN** it uses the named-event SSE `/api/chat` consumer and exposes no AiOps mode or button
+
+#### Scenario: Removed endpoint scan
+- **WHEN** Controller mappings and frontend request targets are inspected
+- **THEN** no `/api/ai_ops`, `/api/chat/clear` or `/api/chat/session` execution/management mapping remains
+
+### Requirement: Evidence backends SHALL not be Agent contracts
+RAG and log query implementations MAY be reused behind Harness adapters, but MUST NOT expose `@Tool`, `ToolCallbackProvider`, topic-discovery-first behavior, legacy Tool descriptions, Session ThreadLocal, session dedup or per-backend durable recorder side effects. Agent-facing Tool names, schemas and descriptions SHALL come only from `HarnessEvidenceTools` and `AgentToolContracts`.
+
+#### Scenario: Tool discovery inspection
+- **WHEN** Spring Tool annotations and Agent callback registration are inspected
+- **THEN** only `lookup_knowledge`, `query_logs` and `query_mysql` Harness ACI callbacks are available to the Diagnosis Agent
+
+#### Scenario: Backend invocation
+- **WHEN** a Harness adapter calls RAG or Mock logs
+- **THEN** the backend returns raw adapter input without reading ThreadLocal or writing a second ToolInvocation record
+
+### Requirement: Agent durable audit SHALL be metadata-only
+Every persisted Diagnosis AgentStep SHALL use the exact RunnableConfig sessionId/runId and MAY contain only bounded metadata such as message count/roles, Tool names, text presence, duration and budget counters. It MUST NOT persist or log Prompt text, message content, model output text, Tool arguments, raw evidence or Thought, and MUST NOT fall back to ThreadLocal identity.
+
+#### Scenario: Model step persistence
+- **WHEN** the Diagnosis Agent performs model calls
+- **THEN** AgentStep rows use the SSE metadata identity and contain no user query, evidence body, Tool argument or chain-of-thought text
+
+#### Scenario: Missing metadata
+- **WHEN** an audit hook is invoked without exact sessionId/runId metadata
+- **THEN** it skips persistence with a safe warning rather than inventing or reading implicit identity
+
+### Requirement: ToolBoundary SHALL write safe durable audit
+For each accepted Harness evidence Tool call, ToolBoundary SHALL attempt to write one durable audit row with exact sessionId/runId/toolCallId/toolName, invocation/evidence status, stable error code, duration and byte counts. Durable audit MUST NOT contain the complete request, SQL/log query body, raw response or Agent projection content. Audit persistence failure SHALL be observable but MUST NOT change the canonical Tool result.
+
+#### Scenario: Successful Tool call
+- **WHEN** ToolBoundary completes a READY invocation
+- **THEN** Redis retains the canonical record and MySQL receives one metadata-only ToolInvocation row for the same Run and Tool call
+
+#### Scenario: Failed Tool call
+- **WHEN** ToolBoundary returns a stable error after an accepted request
+- **THEN** the durable row records only stable status/error metadata and no internal exception or raw payload
+
+#### Scenario: Audit database failure
+- **WHEN** durable audit persistence throws after canonical result determination
+- **THEN** ToolBoundary logs a safe warning and returns the unchanged canonical Tool result
+
+### Requirement: Current documentation SHALL describe the single Harness architecture
+Current MVP architecture, Agent/Harness, Trace, API and demo documents SHALL describe the single Diagnosis Agent, explicit Harness ownership, ACI Tools, named SSE and safe durable audit. ISS-012 and ISS-013 SHALL be recorded as absorbed by ISS-014; historical archived documents MAY retain historical descriptions.
+
+#### Scenario: Current documentation scan
+- **WHEN** non-archived current architecture and demo documents are inspected
+- **THEN** they do not present Planner/Executor/Verifier/Composer or `/api/ai_ops` as current runtime behavior
+
+### Requirement: Repository verification SHALL be clean
+The final implementation SHALL compile, pass focused and relevant full regressions, pass JavaScript syntax and static legacy scans, and pass strict OpenSpec validation for all current specs. No dead imports, temporary files, hardcoded credentials, compatibility flags or unexplained legacy references may be introduced.
+
+#### Scenario: Automated verification
+- **WHEN** the stage 7 verification suite runs
+- **THEN** all selected tests/build/static/OpenSpec checks pass with no legacy runtime token in production source
+
+### Requirement: Final live E2E SHALL correlate exact Run evidence
+The final acceptance SHALL start the Spring Boot application, send a real diagnosis request through `/api/chat`, strictly parse the named SSE sequence, inspect `logs/`, and query MySQL with `scripts/query_mysql.py` using the exact metadata sessionId/runId. It SHALL verify Run intent/outcome/status, one Diagnosis Agent identity, bounded model/Tool counts, ToolInvocation identity, safe final content/fallback and no prohibited content leakage.
+
+#### Scenario: Live successful or safe fallback diagnosis
+- **WHEN** the configured model and infrastructure process the fixed diagnosis request
+- **THEN** SSE, logs, diagnosis_run, agent_step and tool_invocation evidence agree on the exact identity and final release is SUCCESS or documented safe FALLBACK
+
+#### Scenario: External Tool boundary statement
+- **WHEN** final acceptance is archived
+- **THEN** it distinguishes Mock query_logs and isolated query_mysql contract evidence from unverified real CLS or production business MySQL integration
diff --git a/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/tasks.md b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/tasks.md
new file mode 100644
index 0000000..91652e3
--- /dev/null
+++ b/openspec/changes/archive/2026-07-22-single-react-cleanup-e2e/tasks.md
@@ -0,0 +1,28 @@
+## 1. Remove legacy public and orchestration paths
+
+- [x] 1.1 Remove `/api/ai_ops`, bundled frontend AiOps controls/consumer, AiOps DTO/config/service/rule evaluation/prompts/tests and obsolete logger/config references.
+- [x] 1.2 Remove legacy `ChatService`, Sequential Planner/Executor/Verifier/Composer prompts, Gatekeeper/Verifier/Skill/Token helpers, ThreadLocals and self-only tests/resources.
+- [x] 1.3 Remove legacy Chat Session controller, Redis SessionManager/SessionContext and tests while retaining current JPA `chat_session` Run metadata.
+- [x] 1.4 Compile and run reference/static scans proving no production Supervisor/Sequential legacy orchestration or removed endpoint remains.
+
+## 2. Make Tool and audit boundaries Harness-native
+
+- [x] 2.1 Refactor `LookupKnowledgeTool` into a Harness-only backend by removing old Tool annotation, Session ThreadLocal, session dedup and legacy ToolInvocation recording while preserving retrieval behavior tests.
+- [x] 2.2 Refactor `QueryLogsTools` into a Harness-only Mock/log backend by removing old Tool annotations, topic-discovery-first contract and legacy recorder side effects while preserving adapter contract tests.
+- [x] 2.3 Replace `AgentLoggingHook` with a metadata-only Harness audit hook that requires exact RunnableConfig IDs and persists no Prompt, message/model content, Tool arguments or Thought.
+- [x] 2.4 Add ToolBoundary audit event/sink and a fail-open JPA adapter that persists one bounded metadata-only ToolInvocation row per accepted Harness call.
+- [x] 2.5 Add focused negative/safety tests for exact audit identity, no sensitive payload persistence, audit failure isolation and exclusive Harness Tool discovery.
+
+## 3. Close documentation and repository contracts
+
+- [x] 3.1 Update current MVP architecture, Agent/Harness, Trace/API and demo documents to the single Diagnosis Agent, ACI Tool and named SSE architecture.
+- [x] 3.2 Mark ISS-012 and ISS-013 absorbed by ISS-014 and archive them; update ISS-014 stage/completion state only after final acceptance.
+- [x] 3.3 Fix known `mysql-readonly-tool` and `rag-log-projections` main spec strict-validation defects without changing their behavior.
+
+## 4. Final verification and live acceptance
+
+- [x] 4.1 Run focused stage 7 tests, relevant stage 1-6B regressions, Maven compile/package, JavaScript syntax, static legacy/sensitive scans and strict OpenSpec validation.
+- [x] 4.2 Start the Spring Boot application with Maven and verify clean context startup plus bounded Harness/Audit bean wiring in `logs/`.
+- [x] 4.3 Execute a real `/api/chat` diagnosis, strictly validate named SSE order/outcome and capture exact metadata sessionId/runId.
+- [x] 4.4 Use `logs/` and `scripts/query_mysql.py` to verify exact diagnosis_run, Diagnosis AgentStep, ToolInvocation identity/count/status/safe fields and released content/fallback; record Mock/isolated external Tool limits.
+- [x] 4.5 Record final cleanup inventory, validation outputs, live E2E evidence, remaining risks and L4 migration/rollback in devflow acceptance before archive.
diff --git a/openspec/specs/mysql-readonly-tool/spec.md b/openspec/specs/mysql-readonly-tool/spec.md
index 5a6b381..379fcf3 100644
--- a/openspec/specs/mysql-readonly-tool/spec.md
+++ b/openspec/specs/mysql-readonly-tool/spec.md
@@ -10,22 +10,52 @@ Define a fail-closed, parameterized, read-only MySQL evidence Tool that reuses T
The Tool SHALL parse exactly one SQL statement with JSqlParser and SHALL accept only a single `SELECT` with explicit projection columns, supported predicates/grouping/ordering, `INNER JOIN` or `LEFT JOIN`, parameter placeholders and allowlisted functions. It SHALL reject writes, CTEs, subqueries, set operations, wildcard projections except `COUNT(*)`, metadata discovery, unsupported joins/functions, `FOR UPDATE`, multiple statements and unknown/ambiguous AST structures.
+#### Scenario: Reject unsupported SQL before execution
+
+- **WHEN** an agent submits a write statement, a second statement, a subquery, or an unallowlisted function
+- **THEN** validation fails closed and no JDBC execution is attempted
+
### Requirement: Data-source and identifier authorization SHALL use exact independent allowlists
The Tool SHALL accept only a logical `data_source` ID and SHALL authorize every schema, table and column used in projection, join, predicate, grouping and ordering against the configured exact allowlist. It SHALL reject unknown data sources, schemas, tables, columns, aliases and ambiguous unqualified columns. Agent input SHALL NOT provide JDBC coordinates or authorization controls.
+#### Scenario: Reject an identifier outside the configured allowlist
+
+- **WHEN** a valid-looking query references an unknown data source, table, column, alias, or ambiguous unqualified column
+- **THEN** authorization fails before connection creation and agent-provided JDBC coordinates are ignored
+
### Requirement: Parameter binding and JDBC execution SHALL be read-only and bounded
The executor SHALL use a configured logical datasource, a read-only JDBC connection, `PreparedStatement` parameter binding, query timeout, max rows and Run cancellation/deadline checks. Placeholder count SHALL exactly match `params`. The executor SHALL not expose connection details or raw JDBC failures to the Agent.
+#### Scenario: Execute an authorized query within read-only bounds
+
+- **WHEN** an authorized query has exactly matching parameters and an active Run
+- **THEN** it executes through a read-only prepared statement with timeout and row limits, while cancellation or a deadline stops execution and raw JDBC details remain hidden
+
### Requirement: MySQL projection SHALL be bounded and evidence-aware
The projector SHALL expose only ordered columns, bounded JSON-safe rows, returned count and truncation. It SHALL enforce row, cell and total UTF-8 limits, redact sensitive column values, return `NO_EVIDENCE` for a successful empty result, and never expose raw JDBC metadata or credentials.
+#### Scenario: Project bounded rows and empty evidence safely
+
+- **WHEN** JDBC returns rows containing oversized or sensitive values, or returns a successful empty result
+- **THEN** the projection redacts and bounds values, reports truncation when data is removed, and returns `NO_EVIDENCE` for the empty result without exposing metadata or credentials
+
### Requirement: MySQL Tool SHALL reuse canonical Harness ownership
The adapter SHALL pass the exact framework `tool_call_id` and RunContext through the existing ToolBoundary and canonical invocation store. It SHALL not create a second ID, use a parallel store, return raw SQL results, or modify legacy audit/public runtime paths.
+#### Scenario: Preserve framework invocation identity
+
+- **WHEN** the adapter invokes an authorized MySQL query
+- **THEN** ToolBoundary and the canonical invocation store receive the exact framework `tool_call_id` and RunContext, with no second identifier or parallel raw-result path
+
### Requirement: The query helper script SHALL be read-only and secret-free by default
The repository query helper SHALL require connection values from environment variables, reject non-SELECT and metadata discovery SQL before connection, and SHALL NOT commit writes or expose hardcoded external connection defaults.
+
+#### Scenario: Reject unsafe helper SQL without connecting
+
+- **WHEN** the helper receives a non-`SELECT` or metadata-discovery statement, or missing environment connection values
+- **THEN** it exits before opening a connection and does not reveal or invent credentials or external connection defaults
diff --git a/openspec/specs/rag-log-projections/spec.md b/openspec/specs/rag-log-projections/spec.md
index a168c70..740b431 100644
--- a/openspec/specs/rag-log-projections/spec.md
+++ b/openspec/specs/rag-log-projections/spec.md
@@ -10,14 +10,34 @@ Define bounded RAG and Mock query-log projections that execute through the stage
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result.
+#### Scenario: Project only bounded RAG evidence
+
+- **WHEN** the adapter receives a logical query and the knowledge backend returns a result
+- **THEN** it invokes through ToolBoundary and exposes only the bounded `RagToolResult` evidence fields, excluding retrieval internals and full document bodies
+
### Requirement: Query-log projection SHALL preserve logical scope and Mock provenance
The query-log adapter SHALL accept only logical topic, query and optional lookback minutes, execute the existing Mock source through ToolBoundary, and project `source_kind=MOCK`, complete scope, match count, returned count, bounded patterns, bounded timeline events and truncation.
+#### Scenario: Return a bounded Mock query-log projection
+
+- **WHEN** the adapter receives a logical topic, query, and optional lookback
+- **THEN** it invokes the Mock source through ToolBoundary and returns `source_kind=MOCK`, the complete logical scope, bounded matches and timeline events, counts, and truncation
+
### Requirement: Projection SHALL redact and bound sensitive log content
The log projector SHALL exclude instance and metrics fields and redact credentials, token-like values, host/pod identifiers, PIDs, IP addresses, SQL literals and stack-like suffixes from Agent-facing messages. It SHALL enforce per-item, collection and total UTF-8 bounds and set `truncated=true` when any bound removes data.
+#### Scenario: Redact sensitive log fields and mark truncation
+
+- **WHEN** a log result contains credentials, identifiers, SQL literals, stack suffixes, or content beyond configured UTF-8 bounds
+- **THEN** those values are omitted or redacted and `truncated=true` indicates removed content
+
### Requirement: Adapters SHALL reuse canonical boundary ownership
Both adapters SHALL pass the framework `tool_call_id` and RunContext to the existing ToolBoundary and SHALL NOT create a second ID, write a parallel store, return raw responses, or modify legacy audit paths.
+
+#### Scenario: Preserve canonical identity for both adapters
+
+- **WHEN** either the RAG or query-log adapter executes
+- **THEN** it passes the exact framework `tool_call_id` and RunContext through ToolBoundary without creating a second identifier, parallel store, raw response path, or legacy audit side effect
diff --git a/openspec/specs/single-react-chat-sse-cutover/spec.md b/openspec/specs/single-react-chat-sse-cutover/spec.md
index 2cb7ac0..7009b6c 100644
--- a/openspec/specs/single-react-chat-sse-cutover/spec.md
+++ b/openspec/specs/single-react-chat-sse-cutover/spec.md
@@ -117,10 +117,3 @@ The bundled frontend SHALL send every Chat message to `/api/chat`, parse complet
#### Scenario: Frontend request target
- **WHEN** static Chat consumer source is inspected
- **THEN** it contains one `/chat` streaming request and no `/chat_stream` or synchronous Chat consumer
-
-### Requirement: AiOps public behavior SHALL remain isolated
-The `/api/ai_ops` URL, request schema and event behavior SHALL remain unchanged in stage 6B. Any internal model/Tool dependency movement needed to keep Chat Controller protocol-only MUST preserve existing AiOps observable behavior.
-
-#### Scenario: AiOps regression
-- **WHEN** existing AiOps Controller and service tests run after Chat cutover
-- **THEN** existing metadata and analysis behavior remains compatible
diff --git a/openspec/specs/single-react-cleanup-e2e/spec.md b/openspec/specs/single-react-cleanup-e2e/spec.md
new file mode 100644
index 0000000..03f1a4b
--- /dev/null
+++ b/openspec/specs/single-react-cleanup-e2e/spec.md
@@ -0,0 +1,88 @@
+# single-react-cleanup-e2e Specification
+
+## Purpose
+TBD - created by archiving change single-react-cleanup-e2e. Update Purpose after archive.
+## Requirements
+### Requirement: Legacy diagnosis orchestration SHALL be physically absent
+Production and test source SHALL contain no legacy ChatService, AiOps Supervisor/Planner/Executor, Sequential Planner/Executor/Verifier/Composer, legacy Gatekeeper/Verifier helpers, ThreadLocal execution context, or their dedicated prompts and obsolete tests. The only business Agent with a Tool loop SHALL be the Diagnosis Agent constructed inside the Harness.
+
+#### Scenario: Source inventory after cleanup
+- **WHEN** production source, resources and tests are scanned
+- **THEN** legacy orchestration classes, prompts, ThreadLocals and self-only tests are absent rather than disabled or commented out
+
+#### Scenario: Agent construction inventory
+- **WHEN** business Agent builders and framework orchestration types are inspected
+- **THEN** only the Diagnosis Agent has evidence Tool callbacks and no Supervisor/Sequential business workflow remains
+
+### Requirement: Public diagnosis SHALL use one endpoint
+The application SHALL expose `POST /api/chat` as the only public diagnosis execution endpoint. `/api/ai_ops`, legacy Chat Session management endpoints and bundled frontend callers for those endpoints MUST be removed without a compatibility branch.
+
+#### Scenario: Bundled frontend diagnosis
+- **WHEN** a user submits a diagnosis from the bundled frontend
+- **THEN** it uses the named-event SSE `/api/chat` consumer and exposes no AiOps mode or button
+
+#### Scenario: Removed endpoint scan
+- **WHEN** Controller mappings and frontend request targets are inspected
+- **THEN** no `/api/ai_ops`, `/api/chat/clear` or `/api/chat/session` execution/management mapping remains
+
+### Requirement: Evidence backends SHALL not be Agent contracts
+RAG and log query implementations MAY be reused behind Harness adapters, but MUST NOT expose `@Tool`, `ToolCallbackProvider`, topic-discovery-first behavior, legacy Tool descriptions, Session ThreadLocal, session dedup or per-backend durable recorder side effects. Agent-facing Tool names, schemas and descriptions SHALL come only from `HarnessEvidenceTools` and `AgentToolContracts`.
+
+#### Scenario: Tool discovery inspection
+- **WHEN** Spring Tool annotations and Agent callback registration are inspected
+- **THEN** only `lookup_knowledge`, `query_logs` and `query_mysql` Harness ACI callbacks are available to the Diagnosis Agent
+
+#### Scenario: Backend invocation
+- **WHEN** a Harness adapter calls RAG or Mock logs
+- **THEN** the backend returns raw adapter input without reading ThreadLocal or writing a second ToolInvocation record
+
+### Requirement: Agent durable audit SHALL be metadata-only
+Every persisted Diagnosis AgentStep SHALL use the exact RunnableConfig sessionId/runId and MAY contain only bounded metadata such as message count/roles, Tool names, text presence, duration and budget counters. It MUST NOT persist or log Prompt text, message content, model output text, Tool arguments, raw evidence or Thought, and MUST NOT fall back to ThreadLocal identity.
+
+#### Scenario: Model step persistence
+- **WHEN** the Diagnosis Agent performs model calls
+- **THEN** AgentStep rows use the SSE metadata identity and contain no user query, evidence body, Tool argument or chain-of-thought text
+
+#### Scenario: Missing metadata
+- **WHEN** an audit hook is invoked without exact sessionId/runId metadata
+- **THEN** it skips persistence with a safe warning rather than inventing or reading implicit identity
+
+### Requirement: ToolBoundary SHALL write safe durable audit
+For each accepted Harness evidence Tool call, ToolBoundary SHALL attempt to write one durable audit row with exact sessionId/runId/toolCallId/toolName, invocation/evidence status, stable error code, duration and byte counts. Durable audit MUST NOT contain the complete request, SQL/log query body, raw response or Agent projection content. Audit persistence failure SHALL be observable but MUST NOT change the canonical Tool result.
+
+#### Scenario: Successful Tool call
+- **WHEN** ToolBoundary completes a READY invocation
+- **THEN** Redis retains the canonical record and MySQL receives one metadata-only ToolInvocation row for the same Run and Tool call
+
+#### Scenario: Failed Tool call
+- **WHEN** ToolBoundary returns a stable error after an accepted request
+- **THEN** the durable row records only stable status/error metadata and no internal exception or raw payload
+
+#### Scenario: Audit database failure
+- **WHEN** durable audit persistence throws after canonical result determination
+- **THEN** ToolBoundary logs a safe warning and returns the unchanged canonical Tool result
+
+### Requirement: Current documentation SHALL describe the single Harness architecture
+Current MVP architecture, Agent/Harness, Trace, API and demo documents SHALL describe the single Diagnosis Agent, explicit Harness ownership, ACI Tools, named SSE and safe durable audit. ISS-012 and ISS-013 SHALL be recorded as absorbed by ISS-014; historical archived documents MAY retain historical descriptions.
+
+#### Scenario: Current documentation scan
+- **WHEN** non-archived current architecture and demo documents are inspected
+- **THEN** they do not present Planner/Executor/Verifier/Composer or `/api/ai_ops` as current runtime behavior
+
+### Requirement: Repository verification SHALL be clean
+The final implementation SHALL compile, pass focused and relevant full regressions, pass JavaScript syntax and static legacy scans, and pass strict OpenSpec validation for all current specs. No dead imports, temporary files, hardcoded credentials, compatibility flags or unexplained legacy references may be introduced.
+
+#### Scenario: Automated verification
+- **WHEN** the stage 7 verification suite runs
+- **THEN** all selected tests/build/static/OpenSpec checks pass with no legacy runtime token in production source
+
+### Requirement: Final live E2E SHALL correlate exact Run evidence
+The final acceptance SHALL start the Spring Boot application, send a real diagnosis request through `/api/chat`, strictly parse the named SSE sequence, inspect `logs/`, and query MySQL with `scripts/query_mysql.py` using the exact metadata sessionId/runId. It SHALL verify Run intent/outcome/status, one Diagnosis Agent identity, bounded model/Tool counts, ToolInvocation identity, safe final content/fallback and no prohibited content leakage.
+
+#### Scenario: Live successful or safe fallback diagnosis
+- **WHEN** the configured model and infrastructure process the fixed diagnosis request
+- **THEN** SSE, logs, diagnosis_run, agent_step and tool_invocation evidence agree on the exact identity and final release is SUCCESS or documented safe FALLBACK
+
+#### Scenario: External Tool boundary statement
+- **WHEN** final acceptance is archived
+- **THEN** it distinguishes Mock query_logs and isolated query_mysql contract evidence from unverified real CLS or production business MySQL integration
diff --git a/src/main/java/com/superbiz/agent/agent/tool/DateTimeTools.java b/src/main/java/com/superbiz/agent/agent/tool/DateTimeTools.java
deleted file mode 100644
index 1087820..0000000
--- a/src/main/java/com/superbiz/agent/agent/tool/DateTimeTools.java
+++ /dev/null
@@ -1,27 +0,0 @@
-package com.superbiz.agent.agent.tool;
-
-import org.slf4j.Logger;
-import org.slf4j.LoggerFactory;
-import org.springframework.ai.tool.annotation.Tool;
-import org.springframework.context.i18n.LocaleContextHolder;
-import org.springframework.stereotype.Component;
-
-import java.time.LocalDateTime;
-
-@Component
-public class DateTimeTools {
-
- private static final Logger logger = LoggerFactory.getLogger(DateTimeTools.class);
-
- /** 工具名常量,用于动态构建提示词 */
- public static final String TOOL_GET_CURRENT_DATETIME = "getCurrentDateTime";
-
- @Tool(description = "Get the current date and time in the user's timezone. " +
- "IMPORTANT: Time changes constantly. Always call this tool when user asks about time, " +
- "even if there's a recent time query in the conversation history.")
- public String getCurrentDateTime() {
- String currentTime = LocalDateTime.now().atZone(LocaleContextHolder.getTimeZone().toZoneId()).toString();
- logger.debug("🕐 getCurrentDateTime 调用 - 返回时间: {}", currentTime);
- return currentTime;
- }
-}
diff --git a/src/main/java/com/superbiz/agent/agent/tool/InternalDocsTools.java b/src/main/java/com/superbiz/agent/agent/tool/InternalDocsTools.java
deleted file mode 100644
index a520374..0000000
--- a/src/main/java/com/superbiz/agent/agent/tool/InternalDocsTools.java
+++ /dev/null
@@ -1,86 +0,0 @@
-package com.superbiz.agent.agent.tool;
-
-import com.fasterxml.jackson.databind.ObjectMapper;
-import com.superbiz.agent.service.VectorSearchService;
-import org.slf4j.Logger;
-import org.slf4j.LoggerFactory;
-import org.springframework.ai.tool.annotation.Tool;
-import org.springframework.ai.tool.annotation.ToolParam;
-import org.springframework.beans.factory.annotation.Autowired;
-import org.springframework.beans.factory.annotation.Value;
-import org.springframework.stereotype.Component;
-
-import java.util.List;
-
-/**
- * 内部文档查询工具
- * 使用 RAG (Retrieval-Augmented Generation) 从内部知识库检索相关文档
- *
- * @deprecated 请使用 {@link com.superbiz.agent.tool.LookupKnowledgeTool} 替代。
- * lookup_knowledge 支持 L0 精确匹配 + L1 语义检索,性能更优且功能更全面。
- * 计划在下一个版本中移除此工具。
- */
-@Deprecated
-@Component
-public class InternalDocsTools {
-
- private static final Logger logger = LoggerFactory.getLogger(InternalDocsTools.class);
-
- /** 工具名常量,用于动态构建提示词 */
- public static final String TOOL_QUERY_INTERNAL_DOCS = "queryInternalDocs";
-
- private final VectorSearchService vectorSearchService;
-
- @Value("${rag.top-k:3}")
- private int topK = 3; // 默认值
-
- private final ObjectMapper objectMapper = new ObjectMapper();
-
- /**
- * 构造函数注入依赖
- * Spring 会自动注入 VectorSearchService
- */
- @Autowired
- public InternalDocsTools(VectorSearchService vectorSearchService) {
- this.vectorSearchService = vectorSearchService;
- }
-
- /**
- * 查询内部文档工具
- *
- * @param query 搜索查询,描述您要查找的信息
- * @return JSON 格式的搜索结果,包含相关文档内容、相似度分数和元数据
- * @deprecated 请使用 {@link com.superbiz.agent.tool.LookupKnowledgeTool#lookupKnowledge(String)} 替代
- */
- @Deprecated
- @Tool(description = "Use this tool to search internal documentation and knowledge base for relevant information. " +
- "It performs RAG (Retrieval-Augmented Generation) to find similar documents and extract processing steps. " +
- "This is useful when you need to understand internal procedures, best practices, or step-by-step guides " +
- "stored in the company's documentation.")
- public String queryInternalDocs(
- @ToolParam(description = "Search query describing what information you are looking for")
- String query) {
-
-
- try {
- // 使用向量搜索服务检索相关文档
- List searchResults =
- vectorSearchService.searchSimilarDocuments(query, topK);
-
- if (searchResults.isEmpty()) {
- return "{\"status\": \"no_results\", \"message\": \"No relevant documents found in the knowledge base.\"}";
- }
-
- // 将搜索结果转换为 JSON 格式
- String resultJson = objectMapper.writeValueAsString(searchResults);
-
-
- return resultJson;
-
- } catch (Exception e) {
- logger.error("[工具错误] queryInternalDocs 执行失败", e);
- return String.format("{\"status\": \"error\", \"message\": \"Failed to query internal docs: %s\"}",
- e.getMessage());
- }
- }
-}
diff --git a/src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java b/src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java
index 6b00c78..e3c5cf0 100644
--- a/src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java
+++ b/src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java
@@ -2,12 +2,9 @@ package com.superbiz.agent.agent.tool;
import com.fasterxml.jackson.annotation.JsonProperty;
import com.fasterxml.jackson.databind.ObjectMapper;
-import com.superbiz.agent.service.ToolInvocationRecorder;
import lombok.Data;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
-import org.springframework.ai.tool.annotation.Tool;
-import org.springframework.ai.tool.annotation.ToolParam;
import org.springframework.beans.factory.annotation.Value;
import org.springframework.stereotype.Component;
@@ -30,16 +27,7 @@ public class QueryLogsTools {
private static final Logger logger = LoggerFactory.getLogger(QueryLogsTools.class);
- /** 工具名常量,用于动态构建提示词 */
- public static final String TOOL_QUERY_LOGS = "queryLogs";
- public static final String TOOL_GET_AVAILABLE_LOG_TOPICS = "getAvailableLogTopics";
-
private final ObjectMapper objectMapper = new ObjectMapper();
- private final ToolInvocationRecorder toolInvocationRecorder;
-
- public QueryLogsTools(ToolInvocationRecorder toolInvocationRecorder) {
- this.toolInvocationRecorder = toolInvocationRecorder;
- }
@Value("${cls.mock-enabled:false}")
private boolean mockEnabled;
@@ -53,97 +41,6 @@ public class QueryLogsTools {
logger.info("✅ QueryLogsTools 初始化成功, Mock模式: {}", mockEnabled);
}
- /**
- * 获取可用的日志主题列表
- * 用于查询前先了解有哪些日志主题可供查询
- */
- @Tool(description = "Get all available log topics and their descriptions. " +
- "Call this tool first before querying logs to understand what log topics are available. " +
- "Returns a list of log topics with their names, descriptions, and example queries.")
- public String getAvailableLogTopics() {
- long startTime = System.currentTimeMillis();
- logger.info("获取可用的日志主题列表");
-
- try {
- List topics = new ArrayList<>();
-
- // 系统指标日志
- LogTopicInfo systemMetrics = new LogTopicInfo();
- systemMetrics.setTopicName("system-metrics");
- systemMetrics.setDescription("系统指标日志,包含 CPU、内存、磁盘使用率等系统资源监控数据");
- systemMetrics.setExampleQueries(List.of(
- "cpu_usage:>80",
- "memory_usage:>85",
- "disk_usage:>90",
- "level:WARN AND service:payment-service"
- ));
- systemMetrics.setRelatedAlerts(List.of("HighCPUUsage", "HighMemoryUsage", "HighDiskUsage"));
- topics.add(systemMetrics);
-
- // 应用日志
- LogTopicInfo applicationLogs = new LogTopicInfo();
- applicationLogs.setTopicName("application-logs");
- applicationLogs.setDescription("应用日志,包含应用程序的错误日志、警告日志、慢请求日志、下游依赖调用日志等");
- applicationLogs.setExampleQueries(List.of(
- "level:ERROR",
- "level:FATAL",
- "http_status:500",
- "response_time:>3000",
- "slow",
- "downstream OR redis OR database OR mq"
- ));
- applicationLogs.setRelatedAlerts(List.of("ServiceUnavailable", "SlowResponse", "HighMemoryUsage"));
- topics.add(applicationLogs);
-
- // 数据库慢查询日志
- LogTopicInfo dbSlowQuery = new LogTopicInfo();
- dbSlowQuery.setTopicName("database-slow-query");
- dbSlowQuery.setDescription("数据库慢查询日志,包含执行时间较长的 SQL 查询,可用于分析数据库性能问题");
- dbSlowQuery.setExampleQueries(List.of(
- "query_time:>2",
- "table:orders",
- "query_type:SELECT",
- "*" // 查询所有慢查询
- ));
- dbSlowQuery.setRelatedAlerts(List.of("SlowResponse", "ServiceUnavailable"));
- topics.add(dbSlowQuery);
-
- // 系统事件日志
- LogTopicInfo systemEvents = new LogTopicInfo();
- systemEvents.setTopicName("system-events");
- systemEvents.setDescription("系统事件日志,包含 Kubernetes Pod 重启、OOM Kill、容器崩溃等系统级事件");
- systemEvents.setExampleQueries(List.of(
- "restart OR crash",
- "oom_kill",
- "event_type:PodRestart",
- "reason:OOMKilled"
- ));
- systemEvents.setRelatedAlerts(List.of("ServiceUnavailable", "HighMemoryUsage"));
- topics.add(systemEvents);
-
- // 构建输出
- LogTopicsOutput output = new LogTopicsOutput();
- output.setSuccess(true);
- output.setTopics(topics);
- output.setAvailableRegions(List.of("ap-guangzhou", "ap-shanghai", "ap-beijing", "ap-chengdu"));
- output.setDefaultRegion("ap-guangzhou");
-
- output.setMessage(String.format("共有 %d 个可用的日志主题。建议使用默认地域 'ap-guangzhou' 或省略 region 参数", topics.size()));
-
- String response = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
- recordInvocation("get_available_log_topics", startTime, "get_available_log_topics", null, null, null,
- response, true, null, "logs", ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED);
- return response;
-
- } catch (Exception e) {
- logger.error("获取日志主题列表失败", e);
- String response = "{\"success\":false,\"message\":\"获取日志主题列表失败: " + e.getMessage() + "\"}";
- recordInvocation("get_available_log_topics", startTime, "get_available_log_topics", null, null, null,
- response, false, e.getMessage(), "logs", ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
- return response;
- }
- }
-
/**
* 查询日志
* 从云日志服务查询指定条件的日志
@@ -160,24 +57,9 @@ public class QueryLogsTools {
private static final String DEFAULT_REGION = "ap-guangzhou";
- @Tool(description = "Query logs from Cloud Log Service (CLS). " +
- "Use this tool to search application logs, system metrics, and other log data. " +
- "IMPORTANT: Before calling this tool, you should call getAvailableLogTopics to understand what log topics are available. " +
- "Available log topics: " +
- "1) 'system-metrics' - System metrics logs (CPU, memory, disk usage, etc. Related to HighCPUUsage, HighMemoryUsage, HighDiskUsage alerts); " +
- "2) 'application-logs' - Application logs (error logs, slow request logs, downstream dependency logs. Related to ServiceUnavailable, SlowResponse alerts); " +
- "3) 'database-slow-query' - Database slow query logs (SQL queries with long execution time. Related to SlowResponse alerts); " +
- "4) 'system-events' - System event logs (Pod restart, OOM Kill, container crash. Related to ServiceUnavailable, HighMemoryUsage alerts). " +
- "logTopic (required, one of the above topics or their CLS topicId), " +
- "query (optional, defaults to a curated search if empty), " +
- "limit (optional, default 20, max 100).")
public String queryLogs(
- @ToolParam(description = "地域,可选值: ap-guangzhou, ap-shanghai, ap-beijing, ap-chengdu。默认 ap-guangzhou") String region,
- @ToolParam(description = "日志主题,如 system-metrics, application-logs, database-slow-query, system-events,也支持 CLS TopicId") String logTopic,
- @ToolParam(description = "查询条件,支持 Lucene 语法,如 level:ERROR OR cpu_usage:>80;为空时返回该主题近 5 条核心日志") String query,
- @ToolParam(description = "返回日志条数,默认20,最大100") Integer limit) {
+ String region, String logTopic, String query, Integer limit) {
- long startTime = System.currentTimeMillis();
int actualLimit = (limit == null || limit <= 0) ? 20 : Math.min(limit, 100);
String safeQuery = query == null ? "" : query;
@@ -193,15 +75,12 @@ public class QueryLogsTools {
} else {
// 真实模式:调用 CLS API(这里预留接口,后续实现)
String response = buildErrorResponse("CLS 真实查询尚未实现,请启用 mock 模式进行测试");
- recordInvocation("query_logs", startTime, safeQuery, region, logTopic, actualLimit, response, false,
- "CLS 真实查询尚未实现,请启用 mock 模式进行测试", normalizeTopicDomain(logTopic),
- ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
return response;
}
// 构建成功响应
QueryLogsOutput output = new QueryLogsOutput();
- output.setSuccess(!logEntries.isEmpty());
+ output.setSuccess(true);
output.setRegion(region);
output.setLogTopic(logTopic);
output.setQuery(safeQuery.isBlank() ? "DEFAULT_QUERY" : safeQuery);
@@ -211,56 +90,15 @@ public class QueryLogsTools {
String jsonResult = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
logger.info("日志查询完成: 找到 {} 条日志", logEntries.size());
- recordInvocation("query_logs", startTime, safeQuery, region, logTopic, actualLimit, jsonResult,
- true, null, normalizeTopicDomain(logTopic),
- logEntries.isEmpty()
- ? ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE
- : ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED);
-
return jsonResult;
} catch (Exception e) {
logger.error("查询日志失败", e);
String response = buildErrorResponse("查询失败: " + e.getMessage());
- recordInvocation("query_logs", startTime, safeQuery, region, logTopic, actualLimit, response, false,
- e.getMessage(), normalizeTopicDomain(logTopic), ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
return response;
}
}
- private void recordInvocation(String toolName, long startTime, String query, String region, String logTopic, Integer limit,
- String output, boolean success, String errorMessage, String topicDomain,
- String evidenceStatus) {
- Map input = new HashMap<>();
- input.put("query", query == null || query.isBlank() ? "DEFAULT_QUERY" : query);
- if (region != null) {
- input.put("region", region);
- }
- if (logTopic != null) {
- input.put("log_topic", logTopic);
- }
- if (limit != null) {
- input.put("limit", limit);
- }
- input.put("mock_enabled", mockEnabled);
-
- toolInvocationRecorder.recordEvidenceTool(
- toolName,
- input,
- output,
- success,
- startTime,
- errorMessage,
- topicDomain,
- evidenceStatus,
- Map.of("log_topic", logTopic == null ? "" : logTopic)
- );
- }
-
- private String normalizeTopicDomain(String logTopic) {
- return logTopic == null || logTopic.isBlank() ? "logs" : logTopic;
- }
-
/**
* 构建 Mock 日志数据
@@ -764,42 +602,4 @@ public class QueryLogsTools {
private String message;
}
- /**
- * 日志主题信息
- */
- @Data
- public static class LogTopicInfo {
- @JsonProperty("topic_name")
- private String topicName;
-
- @JsonProperty("description")
- private String description;
-
- @JsonProperty("example_queries")
- private List exampleQueries;
-
- @JsonProperty("related_alerts")
- private List relatedAlerts;
- }
-
- /**
- * 日志主题列表输出
- */
- @Data
- public static class LogTopicsOutput {
- @JsonProperty("success")
- private boolean success;
-
- @JsonProperty("topics")
- private List topics;
-
- @JsonProperty("available_regions")
- private List availableRegions;
-
- @JsonProperty("default_region")
- private String defaultRegion;
-
- @JsonProperty("message")
- private String message;
- }
}
diff --git a/src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java b/src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java
deleted file mode 100644
index 7d2f9af..0000000
--- a/src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java
+++ /dev/null
@@ -1,332 +0,0 @@
-package com.superbiz.agent.agent.tool;
-
-import com.fasterxml.jackson.annotation.JsonProperty;
-import com.fasterxml.jackson.databind.ObjectMapper;
-import com.superbiz.agent.service.ToolInvocationRecorder;
-import lombok.Data;
-import okhttp3.OkHttpClient;
-import okhttp3.Request;
-import okhttp3.Response;
-import org.slf4j.Logger;
-import org.slf4j.LoggerFactory;
-import org.springframework.ai.tool.annotation.Tool;
-import org.springframework.beans.factory.annotation.Value;
-import org.springframework.stereotype.Component;
-
-import java.time.Duration;
-import java.time.Instant;
-import java.time.temporal.ChronoUnit;
-import java.util.*;
-
-/**
- * Prometheus 告警查询工具
- * 用于查询 Prometheus 的活动告警信息
- */
-@Component
-public class QueryMetricsTools {
-
- private static final Logger logger = LoggerFactory.getLogger(QueryMetricsTools.class);
-
- /** 工具名常量,用于动态构建提示词 */
- public static final String TOOL_QUERY_PROMETHEUS_ALERTS = "queryPrometheusAlerts";
-
- private final ObjectMapper objectMapper = new ObjectMapper();
- private final ToolInvocationRecorder toolInvocationRecorder;
-
- public QueryMetricsTools(ToolInvocationRecorder toolInvocationRecorder) {
- this.toolInvocationRecorder = toolInvocationRecorder;
- }
-
- @Value("${prometheus.base-url}")
- private String prometheusBaseUrl;
-
- @Value("${prometheus.timeout:10}")
- private int timeout;
-
- @Value("${prometheus.mock-enabled:false}")
- private boolean mockEnabled;
-
- private OkHttpClient httpClient;
-
- @jakarta.annotation.PostConstruct
- public void init() {
- this.httpClient = new OkHttpClient.Builder()
- .connectTimeout(Duration.ofSeconds(timeout))
- .readTimeout(Duration.ofSeconds(timeout))
- .build();
- logger.info("✅ QueryMetricsTools 初始化成功, Prometheus URL: {}, Mock模式: {}", prometheusBaseUrl, mockEnabled);
- }
-
- /**
- * 查询 Prometheus 活动告警
- * 该工具从 Prometheus 告警系统检索所有当前活动/触发的告警,包括标签、注释、状态和值
- */
- @Tool(description = "Query active alerts from Prometheus alerting system. " +
- "This tool retrieves all currently active/firing alerts including their labels, annotations, state, and values. " +
- "Use this tool when you need to check what alerts are currently firing, investigate alert conditions, or monitor alert status.")
- public String queryPrometheusAlerts() {
- long startTime = System.currentTimeMillis();
- logger.info("开始查询 Prometheus 活动告警, Mock模式: {}", mockEnabled);
-
- try {
- List simplifiedAlerts;
-
- if (mockEnabled) {
- // Mock 模式:返回与文档关联的模拟告警数据
- simplifiedAlerts = buildMockAlerts();
- logger.info("使用 Mock 数据,返回 {} 个模拟告警", simplifiedAlerts.size());
- } else {
- // 真实模式:调用 Prometheus Alerts API
- PrometheusAlertsResult result = fetchPrometheusAlerts();
-
- if (!"success".equals(result.getStatus())) {
- String response = buildErrorResponse("Prometheus API 返回非成功状态: " + result.getStatus(), result.getError());
- recordInvocation(startTime, response, false, result.getError(), ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
- return response;
- }
-
- // 转换为简化格式,对于相同的 alertname,只保留第一个
- Set seenAlertNames = new HashSet<>();
- simplifiedAlerts = new ArrayList<>();
-
- for (PrometheusAlert alert : result.getData().getAlerts()) {
- String alertName = alert.getLabels().get("alertname");
-
- // 如果这个 alertname 已经存在,跳过
- if (seenAlertNames.contains(alertName)) {
- continue;
- }
-
- // 标记为已见过
- seenAlertNames.add(alertName);
-
- SimplifiedAlert simplified = new SimplifiedAlert();
- simplified.setAlertName(alertName);
- simplified.setDescription(alert.getAnnotations().getOrDefault("description", ""));
- simplified.setState(alert.getState());
- simplified.setActiveAt(alert.getActiveAt());
- simplified.setDuration(calculateDuration(alert.getActiveAt()));
-
- simplifiedAlerts.add(simplified);
- }
- }
-
- // 构建成功响应
- PrometheusAlertsOutput output = new PrometheusAlertsOutput();
- output.setSuccess(true);
- output.setAlerts(simplifiedAlerts);
- output.setMessage(String.format("成功检索到 %d 个活动告警", simplifiedAlerts.size()));
-
- String jsonResult = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
- logger.info("Prometheus 告警查询完成: 找到 {} 个告警", simplifiedAlerts.size());
- recordInvocation(startTime, jsonResult, true, null,
- simplifiedAlerts.isEmpty()
- ? ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE
- : ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED);
-
- return jsonResult;
-
- } catch (Exception e) {
- logger.error("查询 Prometheus 告警失败", e);
- String response = buildErrorResponse("查询失败", e.getMessage());
- recordInvocation(startTime, response, false, e.getMessage(), ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
- return response;
- }
- }
-
- private void recordInvocation(long startTime, String output, boolean success, String errorMessage, String evidenceStatus) {
- toolInvocationRecorder.recordEvidenceTool(
- "query_metrics",
- Map.of("query", "active_prometheus_alerts", "mock_enabled", mockEnabled),
- output,
- success,
- startTime,
- errorMessage,
- "prometheus_alerts",
- evidenceStatus,
- Map.of("metric_family", "prometheus_alerts")
- );
- }
-
- /**
- * 构建 Mock 告警数据
- * 与 aiops-docs 文档中的告警类型对应:
- * - HighCPUUsage: CPU使用率过高
- * - HighMemoryUsage: 内存使用率过高
- * - HighDiskUsage: 磁盘使用率过高
- * - ServiceUnavailable: 服务不可用
- * - SlowResponse: 响应时间过长
- */
- private List buildMockAlerts() {
- List alerts = new ArrayList<>();
- Instant now = Instant.now();
-
- // 告警1: CPU使用率过高 - 持续约25分钟
- SimplifiedAlert cpuAlert = new SimplifiedAlert();
- cpuAlert.setAlertName("HighCPUUsage");
- cpuAlert.setDescription("服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。" +
- "实例: pod-payment-service-7d8f9c6b5-x2k4m,命名空间: production");
- cpuAlert.setState("firing");
- Instant cpuActiveAt = now.minus(25, ChronoUnit.MINUTES);
- cpuAlert.setActiveAt(cpuActiveAt.toString());
- cpuAlert.setDuration(calculateDuration(cpuActiveAt.toString()));
- alerts.add(cpuAlert);
-
- // 告警2: 内存使用率过高 - 持续约15分钟
- SimplifiedAlert memoryAlert = new SimplifiedAlert();
- memoryAlert.setAlertName("HighMemoryUsage");
- memoryAlert.setDescription("服务 order-service 的内存使用率持续超过 85%,当前值为 91%。" +
- "JVM堆内存使用: 3.8GB/4GB,可能存在内存泄漏风险。" +
- "实例: pod-order-service-5c7d8e9f1-m3n2p,命名空间: production");
- memoryAlert.setState("firing");
- Instant memoryActiveAt = now.minus(15, ChronoUnit.MINUTES);
- memoryAlert.setActiveAt(memoryActiveAt.toString());
- memoryAlert.setDuration(calculateDuration(memoryActiveAt.toString()));
- alerts.add(memoryAlert);
-
- // 告警3: 响应时间过长 - 持续约10分钟
- SimplifiedAlert slowAlert = new SimplifiedAlert();
- slowAlert.setAlertName("SlowResponse");
- slowAlert.setDescription("服务 user-service 的 P99 响应时间持续超过 3 秒,当前值为 4.2 秒。" +
- "受影响接口: /api/v1/users/profile, /api/v1/users/orders。" +
- "可能原因:数据库慢查询或下游服务延迟");
- slowAlert.setState("firing");
- Instant slowActiveAt = now.minus(10, ChronoUnit.MINUTES);
- slowAlert.setActiveAt(slowActiveAt.toString());
- slowAlert.setDuration(calculateDuration(slowActiveAt.toString()));
- alerts.add(slowAlert);
-
- return alerts;
- }
-
- /**
- * 从 Prometheus API 获取告警数据
- */
- private PrometheusAlertsResult fetchPrometheusAlerts() throws Exception {
- String apiUrl = prometheusBaseUrl + "/api/v1/alerts";
- logger.debug("请求 Prometheus API: {}", apiUrl);
-
- Request request = new Request.Builder()
- .url(apiUrl)
- .get()
- .build();
-
- try (Response response = httpClient.newCall(request).execute()) {
- if (!response.isSuccessful()) {
- throw new RuntimeException("HTTP 请求失败: " + response.code());
- }
-
- String responseBody = response.body().string();
- return objectMapper.readValue(responseBody, PrometheusAlertsResult.class);
- }
- }
-
- /**
- * 计算从 activeAt 到现在的持续时间
- */
- private String calculateDuration(String activeAtStr) {
- try {
- Instant activeAt = Instant.parse(activeAtStr);
- Duration duration = Duration.between(activeAt, Instant.now());
-
- long hours = duration.toHours();
- long minutes = duration.toMinutes() % 60;
- long seconds = duration.getSeconds() % 60;
-
- if (hours > 0) {
- return String.format("%dh%dm%ds", hours, minutes, seconds);
- } else if (minutes > 0) {
- return String.format("%dm%ds", minutes, seconds);
- } else {
- return String.format("%ds", seconds);
- }
- } catch (Exception e) {
- logger.warn("解析时间失败: {}", activeAtStr, e);
- return "unknown";
- }
- }
-
- /**
- * 构建错误响应
- */
- private String buildErrorResponse(String message, String error) {
- try {
- PrometheusAlertsOutput output = new PrometheusAlertsOutput();
- output.setSuccess(false);
- output.setMessage(message);
- output.setError(error);
- return objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
- } catch (Exception e) {
- return String.format("{\"success\":false,\"message\":\"%s\",\"error\":\"%s\"}", message, error);
- }
- }
-
- // ==================== 数据模型 ====================
-
- /**
- * Prometheus 告警信息结构
- */
- @Data
- public static class PrometheusAlert {
- private Map labels;
- private Map annotations;
- private String state;
- private String activeAt;
- private String value;
- }
-
- /**
- * Prometheus 告警查询结果
- */
- @Data
- public static class PrometheusAlertsResult {
- private String status;
- private AlertsData data;
- private String error;
- private String errorType;
- }
-
- @Data
- public static class AlertsData {
- private List alerts = new ArrayList<>();
- }
-
- /**
- * 简化的告警信息
- */
- @Data
- public static class SimplifiedAlert {
- @JsonProperty("alert_name")
- private String alertName;
-
- @JsonProperty("description")
- private String description;
-
- @JsonProperty("state")
- private String state;
-
- @JsonProperty("active_at")
- private String activeAt;
-
- @JsonProperty("duration")
- private String duration;
- }
-
- /**
- * 告警查询输出
- */
- @Data
- public static class PrometheusAlertsOutput {
- @JsonProperty("success")
- private boolean success;
-
- @JsonProperty("alerts")
- private List alerts;
-
- @JsonProperty("message")
- private String message;
-
- @JsonProperty("error")
- private String error;
- }
-}
diff --git a/src/main/java/com/superbiz/agent/config/AiOpsPromptProperties.java b/src/main/java/com/superbiz/agent/config/AiOpsPromptProperties.java
deleted file mode 100644
index d3a800a..0000000
--- a/src/main/java/com/superbiz/agent/config/AiOpsPromptProperties.java
+++ /dev/null
@@ -1,56 +0,0 @@
-package com.superbiz.agent.config;
-
-import lombok.extern.slf4j.Slf4j;
-import org.springframework.context.annotation.Configuration;
-import org.springframework.core.io.ClassPathResource;
-
-import jakarta.annotation.PostConstruct;
-import java.io.IOException;
-import java.nio.charset.StandardCharsets;
-
-/**
- * AI Ops Agent Prompt 配置
- * 从独立的 Markdown 文件加载 Prompt 模板
- */
-@Slf4j
-@Configuration
-public class AiOpsPromptProperties {
-
- private String planner;
- private String executor;
- private String supervisor;
-
- @PostConstruct
- public void loadPrompts() {
- try {
- planner = loadPromptFromFile("prompts/planner-prompt.md");
- executor = loadPromptFromFile("prompts/executor-prompt.md");
- supervisor = loadPromptFromFile("prompts/supervisor-prompt.md");
-
- log.info("AI Ops Prompts 加载成功");
- log.debug("Planner Prompt 长度: {} 字符", planner.length());
- log.debug("Executor Prompt 长度: {} 字符", executor.length());
- log.debug("Supervisor Prompt 长度: {} 字符", supervisor.length());
- } catch (IOException e) {
- log.error("加载 Prompt 文件失败", e);
- throw new RuntimeException("Failed to load AI Ops prompts", e);
- }
- }
-
- private String loadPromptFromFile(String path) throws IOException {
- ClassPathResource resource = new ClassPathResource(path);
- return new String(resource.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
- }
-
- public String getPlanner() {
- return planner;
- }
-
- public String getExecutor() {
- return executor;
- }
-
- public String getSupervisor() {
- return supervisor;
- }
-}
diff --git a/src/main/java/com/superbiz/agent/config/HarnessChatConfiguration.java b/src/main/java/com/superbiz/agent/config/HarnessChatConfiguration.java
index b7dbe13..0e1972f 100644
--- a/src/main/java/com/superbiz/agent/config/HarnessChatConfiguration.java
+++ b/src/main/java/com/superbiz/agent/config/HarnessChatConfiguration.java
@@ -49,7 +49,8 @@ import com.superbiz.agent.harness.application.DiagnosisOperation;
import com.superbiz.agent.harness.application.KnowledgeQueryOperation;
import com.superbiz.agent.harness.application.SystemChatOperation;
import com.superbiz.agent.harness.application.IntentRouting;
-import com.superbiz.agent.hook.AgentLoggingHook;
+import com.superbiz.agent.harness.audit.HarnessAgentAuditHook;
+import com.superbiz.agent.harness.audit.ToolInvocationAuditSink;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.tool.LookupKnowledgeTool;
import com.superbiz.agent.agent.tool.QueryLogsTools;
@@ -137,8 +138,9 @@ public class HarnessChatConfiguration {
@Bean
public ToolBoundary toolBoundary(DiagnosisHarnessCore core, ToolCallKeyFactory keyFactory,
- CanonicalInvocationStore store, ObjectMapper objectMapper, Clock clock) {
- return new ToolBoundary(core, keyFactory, store, objectMapper, clock);
+ CanonicalInvocationStore store, ObjectMapper objectMapper, Clock clock,
+ ToolInvocationAuditSink auditSink) {
+ return new ToolBoundary(core, keyFactory, store, objectMapper, clock, auditSink);
}
@Bean
@@ -158,17 +160,17 @@ public class HarnessChatConfiguration {
@Bean
public RagToolAdapter ragToolAdapter(ToolBoundary boundary, ObjectMapper mapper,
- RagResultProjector projector, LookupKnowledgeTool legacy) {
- return new RagToolAdapter(boundary, mapper, projector, legacy::lookupKnowledge);
+ RagResultProjector projector, LookupKnowledgeTool backend) {
+ return new RagToolAdapter(boundary, mapper, projector, backend::lookupKnowledge);
}
@Bean
public QueryLogsToolAdapter queryLogsToolAdapter(ToolBoundary boundary, ObjectMapper mapper,
QueryLogsResultProjector projector,
- ObjectProvider legacy, Clock clock) {
+ ObjectProvider backend, Clock clock) {
return new QueryLogsToolAdapter(boundary, mapper, projector,
(region, topic, query, limit) -> {
- QueryLogsTools tools = legacy.getIfAvailable();
+ QueryLogsTools tools = backend.getIfAvailable();
if (tools == null) {
return "{\"success\":false,\"logs\":[],\"total\":0}";
}
@@ -230,10 +232,11 @@ public class HarnessChatConfiguration {
@Bean
public DiagnosisAgentFactory diagnosisAgentFactory(ChatModel chatModel, DiagnosisHarnessCore core,
- HarnessEvidenceTools tools, ObjectMapper mapper,
- AgentStepRepository steps) {
+ HarnessEvidenceTools tools, ObjectMapper mapper,
+ AgentStepRepository steps) {
return new DiagnosisAgentFactory(chatModel, core, tools, mapper,
- List.of(new AgentLoggingHook(steps, DiagnosisAgentFactory.AGENT_NAME)));
+ List.of(new HarnessAgentAuditHook(
+ steps, mapper, DiagnosisAgentFactory.AGENT_NAME)));
}
@Bean
diff --git a/src/main/java/com/superbiz/agent/config/WebConfig.java b/src/main/java/com/superbiz/agent/config/WebConfig.java
index 5b845c0..ec7f050 100644
--- a/src/main/java/com/superbiz/agent/config/WebConfig.java
+++ b/src/main/java/com/superbiz/agent/config/WebConfig.java
@@ -33,6 +33,6 @@ public class WebConfig implements WebMvcConfigurer {
@Bean
public ObjectMapper objectMapper() {
- return new ObjectMapper();
+ return new ObjectMapper().findAndRegisterModules();
}
}
diff --git a/src/main/java/com/superbiz/agent/controller/AiOpsController.java b/src/main/java/com/superbiz/agent/controller/AiOpsController.java
deleted file mode 100644
index 3406ed0..0000000
--- a/src/main/java/com/superbiz/agent/controller/AiOpsController.java
+++ /dev/null
@@ -1,124 +0,0 @@
-package com.superbiz.agent.controller;
-
-import com.alibaba.cloud.ai.graph.OverAllState;
-import com.superbiz.agent.dto.AIOpsRequest;
-import com.superbiz.agent.service.AiOpsService;
-import lombok.Getter;
-import lombok.Setter;
-import org.springframework.beans.factory.annotation.Qualifier;
-import org.springframework.http.MediaType;
-import org.springframework.web.bind.annotation.PostMapping;
-import org.springframework.web.bind.annotation.RequestBody;
-import org.springframework.web.bind.annotation.RequestMapping;
-import org.springframework.web.bind.annotation.RestController;
-import org.springframework.web.servlet.mvc.method.annotation.SseEmitter;
-
-import java.io.IOException;
-import java.util.Optional;
-import java.util.concurrent.RejectedExecutionException;
-import java.util.concurrent.ThreadPoolExecutor;
-
-/** Preserves the legacy AiOps SSE protocol independently from Chat. */
-@RestController
-@RequestMapping("/api")
-public class AiOpsController {
-
- private final AiOpsService aiOpsService;
- private final ThreadPoolExecutor workerExecutor;
-
- public AiOpsController(AiOpsService aiOpsService,
- @Qualifier("chatWorkerExecutor") ThreadPoolExecutor workerExecutor) {
- this.aiOpsService = aiOpsService;
- this.workerExecutor = workerExecutor;
- }
-
- @PostMapping(value = "/ai_ops", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
- public SseEmitter aiOps(@RequestBody(required = false) AIOpsRequest request) {
- SseEmitter emitter = new SseEmitter(600_000L);
- String sessionId = aiOpsService.resolveSessionId(request);
- String runId = aiOpsService.newRunId();
-
- emitter.onTimeout(emitter::complete);
- try {
- workerExecutor.execute(() -> executeAiOps(request, sessionId, runId, emitter));
- } catch (RejectedExecutionException rejected) {
- emitter.completeWithError(rejected);
- }
- return emitter;
- }
-
- private void executeAiOps(AIOpsRequest request, String sessionId, String runId, SseEmitter emitter) {
- try {
- emitter.send(SseEmitter.event().name("message")
- .data(SseMessage.metadata(sessionId, runId), MediaType.APPLICATION_JSON));
- emitter.send(SseEmitter.event().name("message")
- .data(SseMessage.content("正在读取告警并拆解任务...\n"), MediaType.APPLICATION_JSON));
-
- Optional state = aiOpsService.executeAiOpsAnalysis(request, sessionId, runId);
- if (state.isEmpty()) {
- emitter.send(SseEmitter.event().name("message")
- .data(SseMessage.error("多 Agent 编排未获取到有效结果"), MediaType.APPLICATION_JSON));
- emitter.complete();
- return;
- }
-
- Optional report = aiOpsService.extractFinalReport(state.get());
- if (report.isPresent()) {
- aiOpsService.persistFinalReport(sessionId, runId, report.get(), request);
- emitter.send(SseEmitter.event().name("message")
- .data(SseMessage.content(report.get()), MediaType.APPLICATION_JSON));
- } else {
- emitter.send(SseEmitter.event().name("message")
- .data(SseMessage.content("多 Agent 流程已完成,但未能生成最终报告。"), MediaType.APPLICATION_JSON));
- }
- emitter.send(SseEmitter.event().name("message")
- .data(SseMessage.done(), MediaType.APPLICATION_JSON));
- emitter.complete();
- } catch (Exception exception) {
- try {
- emitter.send(SseEmitter.event().name("message")
- .data(SseMessage.error("AI Ops 流程失败"), MediaType.APPLICATION_JSON));
- } catch (IOException ignored) {
- // Client has already gone away.
- }
- emitter.completeWithError(exception);
- }
- }
-
- @Getter
- @Setter
- public static class SseMessage {
- private String type;
- private String data;
- private String sessionId;
- private String runId;
-
- public static SseMessage content(String data) {
- SseMessage message = new SseMessage();
- message.type = "content";
- message.data = data;
- return message;
- }
-
- public static SseMessage metadata(String sessionId, String runId) {
- SseMessage message = new SseMessage();
- message.type = "metadata";
- message.sessionId = sessionId;
- message.runId = runId;
- return message;
- }
-
- public static SseMessage error(String data) {
- SseMessage message = new SseMessage();
- message.type = "error";
- message.data = data;
- return message;
- }
-
- public static SseMessage done() {
- SseMessage message = new SseMessage();
- message.type = "done";
- return message;
- }
- }
-}
diff --git a/src/main/java/com/superbiz/agent/controller/ChatSessionController.java b/src/main/java/com/superbiz/agent/controller/ChatSessionController.java
deleted file mode 100644
index 15506e9..0000000
--- a/src/main/java/com/superbiz/agent/controller/ChatSessionController.java
+++ /dev/null
@@ -1,122 +0,0 @@
-package com.superbiz.agent.controller;
-
-import com.fasterxml.jackson.annotation.JsonAlias;
-import com.fasterxml.jackson.annotation.JsonProperty;
-import com.superbiz.agent.dto.DiagnosisTraceResponse;
-import com.superbiz.agent.service.DiagnosisTraceService;
-import com.superbiz.agent.service.session.SessionManager;
-import lombok.Getter;
-import lombok.Setter;
-import org.springframework.http.ResponseEntity;
-import org.springframework.web.bind.annotation.GetMapping;
-import org.springframework.web.bind.annotation.PathVariable;
-import org.springframework.web.bind.annotation.PostMapping;
-import org.springframework.web.bind.annotation.RequestBody;
-import org.springframework.web.bind.annotation.RequestMapping;
-import org.springframework.web.bind.annotation.RestController;
-
-import java.time.LocalDateTime;
-import java.time.ZoneId;
-import java.util.List;
-import java.util.Optional;
-
-/** Legacy Chat session management endpoints, isolated from the Chat execution adapter. */
-@RestController
-@RequestMapping("/api/chat")
-public class ChatSessionController {
-
- private final DiagnosisTraceService diagnosisTraceService;
- private final SessionManager sessionManager;
-
- public ChatSessionController(DiagnosisTraceService diagnosisTraceService, SessionManager sessionManager) {
- this.diagnosisTraceService = diagnosisTraceService;
- this.sessionManager = sessionManager;
- }
-
- @PostMapping("/clear")
- public ResponseEntity> clearChatHistory(@RequestBody ClearRequest request) {
- if (request == null || request.getId() == null || request.getId().isBlank()) {
- return ResponseEntity.ok(ApiResponse.error("会话ID不能为空"));
- }
- try {
- Optional session =
- sessionManager.getSession(request.getId());
- if (session.isEmpty()) {
- return ResponseEntity.ok(ApiResponse.error("会话不存在"));
- }
- var context = session.get();
- context.clearMessageHistory();
- sessionManager.updateSession(context);
- return ResponseEntity.ok(ApiResponse.success("会话历史已清空"));
- } catch (RuntimeException exception) {
- return ResponseEntity.ok(ApiResponse.error("会话历史清理失败"));
- }
- }
-
- @GetMapping("/session/{sessionId}")
- public ResponseEntity> getSessionInfo(@PathVariable String sessionId) {
- try {
- Optional session = sessionManager.getSession(sessionId);
- if (session.isEmpty()) {
- return ResponseEntity.ok(ApiResponse.error("会话不存在"));
- }
- var context = session.get();
- SessionInfoResponse response = new SessionInfoResponse();
- response.setSessionId(sessionId);
- response.setMessagePairCount(context.getMessagePairCount());
- response.setCreateTime(toEpochMillis(context.getCreatedAt()));
- return ResponseEntity.ok(ApiResponse.success(response));
- } catch (RuntimeException exception) {
- return ResponseEntity.ok(ApiResponse.error("会话读取失败"));
- }
- }
-
- @GetMapping("/session/{sessionId}/runs")
- public ResponseEntity>> listSessionRuns(
- @PathVariable String sessionId) {
- return ResponseEntity.ok(ApiResponse.success(diagnosisTraceService.listRunSummaries(sessionId)));
- }
-
- private long toEpochMillis(LocalDateTime time) {
- return time == null ? 0L : time.atZone(ZoneId.systemDefault()).toInstant().toEpochMilli();
- }
-
- @Getter
- @Setter
- public static class ClearRequest {
- @JsonProperty("Id")
- @JsonAlias({"id", "ID"})
- private String id;
- }
-
- @Getter
- @Setter
- public static class SessionInfoResponse {
- private String sessionId;
- private int messagePairCount;
- private long createTime;
- }
-
- @Getter
- @Setter
- public static class ApiResponse {
- private int code;
- private String message;
- private T data;
-
- public static ApiResponse success(T data) {
- ApiResponse response = new ApiResponse<>();
- response.code = 200;
- response.message = "success";
- response.data = data;
- return response;
- }
-
- public static ApiResponse error(String message) {
- ApiResponse response = new ApiResponse<>();
- response.code = 500;
- response.message = message;
- return response;
- }
- }
-}
diff --git a/src/main/java/com/superbiz/agent/domain/entity/ChatSession.java b/src/main/java/com/superbiz/agent/domain/entity/ChatSession.java
index f1117f0..11b8109 100644
--- a/src/main/java/com/superbiz/agent/domain/entity/ChatSession.java
+++ b/src/main/java/com/superbiz/agent/domain/entity/ChatSession.java
@@ -10,7 +10,7 @@ import java.time.LocalDateTime;
/**
* Chat session metadata entity.
- * Full message history remains in Redis SessionContext.
+ * Stores current Chat Run session metadata; message history is not persisted here.
*/
@Entity
@Table(name = "chat_session", indexes = {
@@ -61,4 +61,3 @@ public class ChatSession {
}
}
}
-
diff --git a/src/main/java/com/superbiz/agent/domain/entity/DiagnosisRun.java b/src/main/java/com/superbiz/agent/domain/entity/DiagnosisRun.java
index 35b8a36..211e5f6 100644
--- a/src/main/java/com/superbiz/agent/domain/entity/DiagnosisRun.java
+++ b/src/main/java/com/superbiz/agent/domain/entity/DiagnosisRun.java
@@ -51,11 +51,11 @@ public class DiagnosisRun {
private String answer;
@Enumerated(EnumType.STRING)
- @Column(name = "intent", length = 32)
+ @Column(name = "intent", length = 32, columnDefinition = "VARCHAR(32)")
private IntentType intent;
@Enumerated(EnumType.STRING)
- @Column(name = "release_outcome", length = 16)
+ @Column(name = "release_outcome", length = 16, columnDefinition = "VARCHAR(16)")
private ReleaseOutcome releaseOutcome;
@JdbcTypeCode(SqlTypes.JSON)
diff --git a/src/main/java/com/superbiz/agent/domain/model/SessionContext.java b/src/main/java/com/superbiz/agent/domain/model/SessionContext.java
deleted file mode 100644
index 00f02bd..0000000
--- a/src/main/java/com/superbiz/agent/domain/model/SessionContext.java
+++ /dev/null
@@ -1,159 +0,0 @@
-package com.superbiz.agent.domain.model;
-
-import com.fasterxml.jackson.annotation.JsonIgnore;
-import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
-import lombok.AllArgsConstructor;
-import lombok.Builder;
-import lombok.Data;
-import lombok.NoArgsConstructor;
-
-import java.io.Serializable;
-import java.time.LocalDateTime;
-import java.util.ArrayList;
-import java.util.HashMap;
-import java.util.List;
-import java.util.Map;
-
-/**
- * 会话上下文数据类
- * 存储在 Redis 中的会话数据
- */
-@Data
-@Builder
-@NoArgsConstructor
-@AllArgsConstructor
-@JsonIgnoreProperties(ignoreUnknown = true)
-public class SessionContext implements Serializable {
-
- private static final long serialVersionUID = 1L;
-
- /**
- * 会话ID
- */
- private String sessionId;
-
- /**
- * 用户ID
- */
- private String userId;
-
- /**
- * 业务ID(订单号/请求ID等)
- */
- private String businessId;
-
- /**
- * 链路追踪ID
- */
- private String traceId;
-
- /**
- * 会话状态(ACTIVE/COMPLETED/EXPIRED)
- */
- private String status;
-
- /**
- * 工具调用历史
- */
- @Builder.Default
- private List toolCalls = new ArrayList<>();
-
- /**
- * 聊天消息历史:[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]
- */
- @Builder.Default
- private List