Compare commits

..
Author SHA1 Message Date
zhuyongxin 30d3296043 docs(mvp): archive session run trace issue 2026-07-11 16:46:20 +08:00
zhuyongxin 3578709896 docs(openspec): archive session run isolation 2026-07-10 22:52:39 +08:00
zhuyongxin f9df94377b feat(trace): finish run-aware demo verification 2026-07-10 21:37:43 +08:00
zhuyongxin 78c1477198 feat(trace): isolate aiops runs 2026-07-10 20:57:46 +08:00
zhuyongxin d928a1968a feat(trace): bind feedback to runs 2026-07-10 20:29:56 +08:00
zhuyongxin 027aed1eeb feat(trace): add run-scoped trace reads 2026-07-10 20:07:23 +08:00
zhuyongxin 26d5529280 feat(trace): isolate chat runs 2026-07-10 19:02:04 +08:00
zhuyongxin 6fdbd34bab docs(openspec): tighten run isolation contract 2026-07-10 17:56:51 +08:00
zhuyongxin 52bf0302c6 feat(trace): add session run isolation schema 2026-07-10 17:47:56 +08:00
zhuyongxin 841437fa06 docs(mvp): organize mvp documentation 2026-07-09 13:40:08 +08:00
zhuyongxin 9c9a0024d4 feat(demo): add interview quality audit 2026-07-09 11:18:49 +08:00
zhuyongxin a6c2d4459c docs(openspec): propose interview demo quality audit 2026-07-09 10:34:33 +08:00
zhuyongxin da45fa3fb0 docs(devflow): sort index by date 2026-07-09 10:24:35 +08:00
aruo db0f229285 feat(eval): add evidence pipeline acceptance closure 2026-07-09 00:47:48 +08:00
aruo a77c947cd4 docs(architecture): align evidence pipeline design 2026-07-08 23:53:21 +08:00
aruo 9a84b3de34 feat(agent): support no-evidence references 2026-07-08 23:30:00 +08:00
aruo 7b8c75e571 feat(agent): harden verifier evidence references 2026-07-08 16:12:56 +08:00
aruo a08672b31e feat(eval): add executor audit closure checks 2026-07-08 10:21:39 +08:00
aruo 6015bcbf6f feat(agent): add composer final answer 2026-07-08 09:51:07 +08:00
aruo a5b4502c72 docs(openspec): propose executor composer final answer 2026-07-08 02:49:46 +08:00
aruo 39c0c5f8be docs(mvp): clarify executor v2 implementation issue 2026-07-08 02:43:12 +08:00
aruo 1b31e78be5 feat(agent): add verifier claim checks 2026-07-08 02:33:02 +08:00
aruo c5e496e715 feat(agent): add executor gatekeeper hook 2026-07-08 02:01:49 +08:00
aruo 050cbc8fee feat(agent): add executor evidence v2 contract 2026-07-08 01:37:15 +08:00
zhuyongxin a6afbfaa9d chore: add editorconfig 2026-07-07 21:08:51 +08:00
zhuyongxin 0ee27eb523 feat(trace): improve session workbench review 2026-07-07 19:06:18 +08:00
zhuyongxin 3b62a8940c chore(openspec): archive executor evidence output contract 2026-07-07 19:04:24 +08:00
zhuyongxin 04eb50e2b4 feat(agent): add executor evidence output contract 2026-07-07 19:02:02 +08:00
274 changed files with 22440 additions and 1918 deletions
+13
View File
@@ -0,0 +1,13 @@
root = true
[*]
charset = utf-8
end_of_line = crlf
insert_final_newline = true
trim_trailing_whitespace = true
[*.md]
trim_trailing_whitespace = false
[*.{java,xml,yml,yaml,properties,json,sql,txt,ps1}]
charset = utf-8
+17 -1
View File
@@ -72,9 +72,10 @@
### SessionContext ### SessionContext
- 定义:会话上下文数据类,存储在 Redis 中的会话数据 - 定义:会话上下文数据类,存储在 Redis 中的会话数据
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、TTL - 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、messageHistory、TTL
- 序列化方式:JSON(GenericJackson2JsonRedisSerializer) - 序列化方式:JSON(GenericJackson2JsonRedisSerializer)
- 使用场景:多轮对话上下文管理、工具调用历史追踪 - 使用场景:多轮对话上下文管理、工具调用历史追踪
- 边界:messageHistory 是热路径对话历史缓存,用于下一轮 prompt 上下文;长期审计的问题和答案应落到 Diagnosis Run,而不是依赖 Redis TTL 内的上下文正文。
### ToolCall ### ToolCall
- 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果 - 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果
@@ -87,6 +88,21 @@
- 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall - 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall
- 使用场景:分布式会话管理、Agent 状态维护 - 使用场景:分布式会话管理、Agent 状态维护
### Chat Session
- 定义:一次多轮对话上下文,由 `sessionId` 唯一标识。
- 使用场景:保存用户连续对话的上下文窗口、会话状态和最近活跃时间。
- 边界:Chat Session 不代表一次诊断执行;同一个 Chat Session 可以包含多次 Diagnosis Run。
### Diagnosis Run
- 定义:一次独立诊断执行,由 `runId` 唯一标识,属于一个 Chat Session。
- 使用场景:保存某一轮诊断的 query、answer、status、耗时、token、反馈和自评估结果。
- 边界:Diagnosis Run 是 Trace、Feedback 和 Evidence score 的绑定对象;多轮对话中的每次 `/api/chat` 或 `/api/ai_ops` 执行都应创建新的 Diagnosis Run。
### Diagnosis Trace
- 定义:一次 Diagnosis Run 的可回放执行轨迹,由 run 主记录、AgentStep 和 ToolInvocation 聚合形成。
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
### Flyway ### Flyway
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行 - 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true - 配置:spring.flyway.enabled=true, baseline-on-migrate=true
+31 -22
View File
@@ -2,25 +2,34 @@
## 项目 ## 项目
| 日期 | slug | 领域 | 关键词 | 状态 | | 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|---|---|---|---|---| |---|---|---|---|---|---|---|
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented | | 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived | | 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived | | 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived | | 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | 收敛诊断评测、稳定 demo 场景和 Gatekeeper 审计元数据。 | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived | | 2026-07-08 | verifier-evidence-reference-fidelity | 强化 Verifier 对 evidence_refs、raw_path 和 no_evidence 的保真校验。 | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived | | 2026-07-07 | executor-evidence-output-contract | 设计 Executor 结构化证据输出,解决证据归因幻觉和 LOW_CONFID 问题。 | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived | | 2026-07-07 | executor-v2-output-contract | 将 Executor 输出升级为 V2 契约,移除面向用户的最终回答字段。 | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived | | 2026-07-07 | executor-gatekeeper-hook | 在 Executor 与 Verifier 之间接入 Gatekeeper,校验证据绑定来源。 | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived | | 2026-07-07 | executor-verifier-claim-checks | 增加 Verifier claim_checks 和事实校验兼容逻辑。 | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived | | 2026-07-06 | rag-eval-pipeline-closure | 建立 RAG 评测闭环,加入 fixture、快照和 baseline diff。 | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived | | 2026-07-06 | modular-rag-pipeline | 将 lookup_knowledge 改造成模块化 RAG 管线,补齐证据块和检索追踪。 | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived | | 2026-07-05 | diagnosis-playbook-skills | 增加诊断 Playbook Skill,沉淀支付超时、MySQL 池、Redis 超时等套路。 | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived | | 2026-07-05 | mvp-demo-interview-runbook | 准备可复现的 MVP 面试演示包、运行手册和 Trace 检查清单。 | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived | | 2026-07-05 | diagnosis-eval-baseline-diff | 增加诊断评测 baseline diff,用于判断回归和证据覆盖变化。 | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived | | 2026-07-04 | expand-diagnosis-eval-fixtures | 扩充诊断评测 fixture,覆盖 Redis、慢响应和 JVM 内存风险。 | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived | | 2026-07-04 | diagnosis-eval-harness | 建立固定诊断评测 Harness,输出 trace、证据覆盖和 verdict 分布。 | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived | | 2026-07-04 | evidence-trace-hardening | 强化工具调用证据链、降级契约和离线验证能力。 | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived | | 2026-07-04 | aiops-traceable-diagnosis-entry | 增加可追踪的 AIOps 告警诊断入口,打通 sessionId 和 Trace API。 | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived | | 2026-07-04 | aiops-alert-scope-control | 收敛 AIOps 告警诊断范围,区分 payload 定向和自动发现模式。 | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived | | 2026-07-03 | mvp-demo-trace-acceptance | 增加 MVP demo 的 Trace 验收,覆盖会话、步骤、工具和反馈链路。 | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-02 | chat-verifier-agent | 增加 Chat Verifier Agent,用 groundedness 和 evidence_refs 校验回答。 | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
| 2026-07-01 | executor-action-memory-relevance | 增加行动记忆和相关性信号,约束 Executor 重复检索。 | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
| 2026-06-30 | session-dedup-knowledge-map | 引入会话级去重和知识域地图,减少重复召回。 | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
| 2026-06-29 | confidence-feedback | 建立质量评估和用户反馈机制,并把有用反馈沉淀为案例。 | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
| 2026-06-26 | session-storage | 建立通用会话存储,记录 session、agent step 和 tool invocation。 | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
| 2026-06-25 | doc-management-ui | 实现文档管理页面,支持文档 CRUD、状态监控和 API 集成。 | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | - | archived |
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
@@ -48,7 +48,7 @@
## 遗留问题 ## 遗留问题
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/ISS-002-executor-unconstrained-lookup.md`。 ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/archived/ISS-002-executor-unconstrained-lookup.md`。
## 已知限制 ## 已知限制
@@ -9,8 +9,8 @@
## Context ## Context
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`. - `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`. - `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision. - `mvp/issues/active/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
## Question Pool ## Question Pool
@@ -0,0 +1,16 @@
# Acceptance: executor-evidence-output-contract
## Draft Acceptance
- [x] Issue exists: `mvp/issues/active/executor-evidence-attribution-hallucination.md`.
- [x] OpenSpec change artifacts exist.
- [x] devflow tracking files exist.
- [x] OpenSpec validation passes.
- [x] Implementation updates Chat Executor prompt.
- [x] Implementation passes structured Executor output to Verifier.
- [x] Verifier prefers structured claims and still falls back safely.
- [x] Focused tests cover parsing, payload assembly, and unsupported confirmed claims.
## Notes
This project is currently in proposal/design stage. Runtime code is intentionally not changed yet.
@@ -0,0 +1,23 @@
# Brief: executor-evidence-output-contract
## Summary
Executor currently returns natural-language diagnosis answers that may mix confirmed evidence, runbook guidance, historical patterns, and unsupported inference. Verifier catches many unsupported facts, but only after extracting claims from prose.
This project defines a structured Executor evidence-attribution contract and updates the Verifier input/verification path to consume it.
## Goal
Make Chat Executor output machine-checkable so confirmed claims are explicitly bound to current-session evidence, while hypotheses and evidence gaps remain visibly separate.
## Scope
- Chat Executor prompt contract.
- Executor structured output parsing.
- Verifier payload extension.
- Chat Verifier prompt behavior.
- Focused tests/eval fixtures.
## Related OpenSpec
`openspec/changes/executor-evidence-output-contract/`
@@ -0,0 +1,71 @@
# Decisions: executor-evidence-output-contract
## sm-flow Progress
### Clarify
Entry summary: recent Chat diagnosis sessions are `LOW_CONFID` because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls.
Slug: `executor-evidence-output-contract`
Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change.
### Context
Relevant history:
- `executor-action-memory-relevance`: Executor already has retrieval quality constraints and should avoid repeated `lookup_knowledge`.
- `chat-verifier-agent`: Verifier should not see intermediate reasoning; it receives explicit `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- `evidence-trace-hardening`: evidence-bearing tools persist stable traces and no-evidence semantics.
- `modular-rag-pipeline`: `lookup_knowledge` exposes evidence blocks and context packs; L0 hints are not fact evidence.
Current code shape:
- `src/main/resources/prompts/chat-executor-prompt.md` is the Chat Executor prompt.
- `src/main/resources/prompts/executor-prompt.md` is for the AiOps flow and is not the target of this Chat change.
- `VerifierInputHook` currently builds a payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- Verifier prompt currently extracts facts from `executor_final_answer`.
### Grill
Question: Should Executor output only JSON or JSON plus readable answer?
Decision: use one JSON object containing both machine fields and `user_facing_answer`. This avoids losing a readable Chinese answer while giving Verifier structured claims.
Question: Should evidence binding use `chunk_id`?
Decision: no. Use generic binding fields because `query_logs` and `query_metrics` do not naturally expose RAG chunks.
Question: Should Verifier trust Executor-provided claims completely?
Decision: no. Verifier should verify structured claims first, then scan `user_facing_answer` for extra confirmed-sounding facts omitted from `claims`.
Question: What happens when Executor JSON is malformed?
Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope
- `design.md`: contract, verifier behavior, risks
- `specs/chat-verifier-agent/spec.md`: modified and added requirements
- `tasks.md`: implementation checklist
### Audit
Cross-artifact alignment:
- Issue describes evidence attribution hallucination.
- Proposal scopes the fix to Executor output and Verifier consumption.
- Design preserves existing verifier isolation.
- Spec adds observable behavior without changing database schema.
- Tasks remain implementation-oriented and unchecked.
Interface impact:
- Prompt/output contract: L2 internal Agent contract change.
- Verifier payload: L2 internal structured input extension.
- Database schema: no change.
- External HTTP API: no intended change.
@@ -0,0 +1,17 @@
# Evidence: executor-evidence-output-contract
## Repository Evidence
- `chat-executor-prompt.md` currently requires using real tool data but does not require a structured evidence-attribution output.
- `chat-verifier-prompt.md` currently extracts facts from `executor_final_answer` prose and compares them with `tool_trace_summary`.
- `VerifierInputHook` currently provides `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- `openspec/specs/chat-verifier-agent/spec.md` already requires explicit verifier inputs, auditable evidence refs, fixed verdicts, and low-confidence handling.
- `openspec/specs/evidence-trace-hardening/spec.md` already distinguishes failed, no-hit, deduped, and successful evidence-tool traces.
## Runtime Evidence From Recent Sessions
Recent MySQL inspection showed repeated `LOW_CONFID` verifier results with many `no_evidence` facts. Typical unsupported claims included OOM, Full GC frequency, specific slow SQL timings, lock waits, and service-specific timeout details that were not supported by current-session tool traces.
## Design Evidence
This change preserves the previous design that Verifier should not inspect intermediate reasoning. The new structured output is still final Executor output, not hidden chain-of-thought.
@@ -0,0 +1,48 @@
# Acceptance: executor-gatekeeper-hook
## Implementation Result
Completed stage two of Executor Structured Output V2.
- Added `ExecutorGatekeeperService`.
- Added initial `schema.executor_v2` and `evidence.invocation_ref` rules.
- Added `gatekeeper_result` to Verifier payload.
- Stored `gatekeeper_result` in `VerifierContextHolder`.
- Persisted `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated Verifier prompt so Gatekeeper fail must not produce PASS.
- Added focused tests for schema failure, valid pass, fabricated invocation ids, tool name mismatch, hook payload, and persistence.
## Static Verification
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
- Coverage: OpenSpec change validity.
## Script Verification
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests for Gatekeeper service, hook integration, and ChatService persistence.
## Browser / Manual Verification
Not run. This stage changes backend validation and audit behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
- Excerpt similarity or phrase/utilization rules.
Reason: This phase intentionally covers deterministic schema and invocation-reference validation. Full live verification is better after Verifier V2 and Composer are implemented.
## Remaining Work
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer final-answer generation.
- Phase five: eval fixtures and full audit closure.
## Archive Status
Devflow archive files created for stage two. OpenSpec archive is expected before moving to stage three.
@@ -0,0 +1,44 @@
# Brief: executor-gatekeeper-hook
## Background
Stage one of Executor Structured Output V2 changed Chat Executor output to `executor_evidence_v2`, removing final-expression fields from Executor. That made the output structured, but it did not yet prevent deterministic evidence attribution failures such as fabricated invocation ids, removed fields, empty evidence bindings, or mismatched tool names.
## Goal
Add a deterministic Gatekeeper between Executor output parsing and Verifier model execution.
The Gatekeeper should:
- Validate the initial Executor V2 schema.
- Validate `claims[].evidence_bindings[].source_invocation_ids` against current-session `tool_invocation` rows.
- Validate evidence binding `tool_name` against the persisted invocation tool name.
- Expose a small `gatekeeper_result` to Verifier and audit persistence.
## Scope
Included:
- New Gatekeeper validation service.
- `schema.executor_v2` initial rule.
- `evidence.invocation_ref` initial rule.
- `VerifierInputHook` payload integration.
- `VerifierContextHolder` storage.
- `ChatService` verifier evaluation persistence.
- Minimal verifier prompt update.
- Focused tests for Gatekeeper, hook payload, fabricated invocation ids, tool name mismatch, and persistence.
Excluded:
- No Executor retry on Gatekeeper failure.
- No excerpt similarity rule in this phase.
- No hallucination phrase or evidence utilization rule in this phase.
- No Verifier V2 `claim_checks`.
- No Composer.
- No database schema changes.
## OpenSpec
- Change: `openspec/changes/executor-gatekeeper-hook`
- Parent stage: `openspec/changes/archive/2026-07-07-executor-v2-output-contract`
@@ -0,0 +1,51 @@
# Decisions: executor-gatekeeper-hook
## Key Decisions
### Gatekeeper stays in VerifierInputHook
Decision: Gatekeeper is integrated inside `VerifierInputHook`, after Executor output parsing and before Verifier model execution.
Reason: The user explicitly chose to keep this version in the Verifier hook and not move validation into Executor hook. This preserves the current workflow orchestration.
### No retry in this phase
Decision: Gatekeeper failure does not trigger automatic Executor retry.
Reason: Retry behavior is intentionally deferred. This phase only validates, exposes, and audits deterministic failures.
### Initial rule set is intentionally small
Decision: Stage two implements only `schema.executor_v2` and `evidence.invocation_ref` as hard checks.
Reason: These rules catch the highest-confidence physical failures with low implementation risk. Excerpt similarity, hallucination phrases, and evidence utilization remain later enhancements.
### No new database schema
Decision: Persist `gatekeeper_result` in existing `diagnosis_session.self_evaluation.verifier_evaluation`.
Reason: The user asked to keep database fields minimal. Existing JSON audit storage is enough for this phase.
### Internal interface impact
Decision: This is an L2 internal interface extension.
Impact:
- Verifier payload gains `gatekeeper_result`.
- `VerifierContextHolder` gains Gatekeeper result storage.
- `self_evaluation.verifier_evaluation` gains `gatekeeper_result`.
- No external API, DTO, database table, or schema migration changes.
## Deferred Decisions
- Whether Gatekeeper should later trigger Executor retry.
- Whether `evidence.excerpt_similarity` should be hard fail or warn-only.
- Whether hallucination phrase and evidence utilization rules should be config-driven from metadata files.
- How Verifier V2 `claim_checks` should enforce Gatekeeper failures in code, beyond prompt instruction.
## Remaining Risks
- Verifier prompt compliance is not a deterministic guarantee; stage three should make Gatekeeper fail incompatible with PASS in Verifier V2 behavior.
- Excerpt authenticity is not checked in this phase, so real invocation ids can still be paired with misleading excerpt text until a later rule is implemented.
@@ -0,0 +1,54 @@
# Evidence: executor-gatekeeper-hook
## Context Evidence
- `executor-v2-output-contract` established `executor_evidence_v2` and removed Executor final-expression fields.
- `VerifierInputHook` is the existing integration point for explicit Verifier payload construction.
- `ChatService.persistVerifierEvaluation(...)` is the existing persistence path for verifier audit snapshots.
- `ToolInvocationRepository.findBySessionIdOrderByIdAsc(...)` provides the current-session invocation pool used by Gatekeeper.
## Implementation Evidence
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
- Implements `schema.executor_v2`.
- Implements `evidence.invocation_ref`.
- Returns `status`, `failed_rules`, `warnings`, and `errors`.
- `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
- Runs Gatekeeper after parsing Executor output and building trace summary.
- Adds `gatekeeper_result` to Verifier payload.
- Stores `gatekeeper_result` in `VerifierContextHolder`.
- `src/main/java/com/superbiz/agent/util/VerifierContextHolder.java`
- Stores per-request Gatekeeper result for later persistence.
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- Wires `ExecutorGatekeeperService` into verifier hook construction.
- Persists `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- `src/main/resources/prompts/chat-verifier-prompt.md`
- Documents `gatekeeper_result` as an input.
- States Gatekeeper fail must not produce PASS.
## Test Evidence
- `src/test/java/com/superbiz/agent/service/ExecutorGatekeeperServiceTest.java`
- Covers schema failure and valid pass behavior.
- Covers fabricated invocation ids and tool name mismatch.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`
- Covers Verifier payload containing `gatekeeper_result`.
- Covers hook behavior for fabricated invocation ids.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
- Covers persistence of `gatekeeper_result` into verifier evaluation.
## Validation Evidence
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests.
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
@@ -0,0 +1,41 @@
# Acceptance: executor-v2-output-contract
## Implementation Result
Completed stage one of Executor Structured Output V2.
- Executor prompt now emits `executor_evidence_v2`.
- Executor output no longer includes `diagnosis_summary` or `user_facing_answer`.
- ChatService PASS path renders V2 structured output into readable Chinese.
- VerifierInputHook remains parse-only and accepts V2 output without final-expression fields.
## Static Verification
- `cmd /c openspec validate executor-v2-output-contract`
- Result: passed.
- Coverage: OpenSpec syntax and change validity.
## Script Verification
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: VerifierInputHook V2 parsing; ChatService PASS rendering for V2; existing sequential workflow tests.
## Browser / Manual Verification
Not run. This stage changes backend prompt/runtime contract and unit-level behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
Reason: Stage one is covered by focused unit tests; live verification is more useful after Gatekeeper and Composer phases.
## Remaining Work
- Phase two: Gatekeeper in `VerifierInputHook`.
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer.
- Phase five: eval fixtures and full audit closure.
@@ -0,0 +1,32 @@
# Brief: executor-v2-output-contract
## Background
`executor_evidence_v1` still made Chat Executor produce both evidence attribution and final user-facing prose through `diagnosis_summary` and `user_facing_answer`.
This kept Executor in a "diagnose and narrate" role and left room for unsupported conclusions to appear before later verification and composition stages.
## Goal
Narrow Chat Executor output to `executor_evidence_v2`: structured diagnostic material only, with final expression removed from Executor.
## Scope
- Update Chat Executor prompt to emit `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
- Keep `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and evidence bindings.
- Add temporary ChatService rendering for PASS + V2 output so normal users do not see raw JSON.
- Preserve V1 `user_facing_answer` extraction for compatibility.
## Non-Goals
- No Gatekeeper implementation.
- No Verifier V2 `claim_checks`.
- No Composer.
- No Planner changes.
- No database schema changes.
- No evidence tool signature changes.
## Related OpenSpec
`openspec/changes/archive/2026-07-07-executor-v2-output-contract/`
@@ -0,0 +1,39 @@
# Decisions: executor-v2-output-contract
## Key Decisions
### Executor V2 removes final-expression fields
Decision: Chat Executor final output now uses `executor_evidence_v2` and must not include `diagnosis_summary` or `user_facing_answer`.
Reason: Executor should collect evidence and produce structured diagnostic material, not write final user-facing conclusions.
### Temporary renderer bridges the gap before Composer
Decision: `ChatService` renders V2 structured fields into readable Chinese only when Verifier returns `PASS`.
Reason: Composer is a later phase, but external users must not receive raw JSON during this intermediate stage.
### V1 compatibility remains
Decision: Existing V1 `user_facing_answer` extraction remains.
Reason: It keeps old tests and any lingering V1 output compatible while the staged migration continues.
### Gatekeeper and Verifier V2 are deferred
Decision: This phase does not add Gatekeeper or `claim_checks`.
Reason: The user requested phase-by-phase implementation with archive and commit after each phase. Gatekeeper is phase two.
## Interface Impact
- Internal Agent output contract: L4, because fields are removed.
- Verifier payload: L2, because raw `executor_final_answer` and parsed `executor_structured_output` remain.
- External Chat answer: compatible intent; users still get readable Chinese.
## Risks
- The temporary renderer is not a full Composer and should be replaced in the Composer phase.
- Verifier prompt still uses V1 `facts_checked`; Verifier V2 is a later phase.
@@ -0,0 +1,21 @@
# Evidence: executor-v2-output-contract
## Context Used
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
- `mvp/issues/design-notes/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
## Code Evidence
- `src/main/resources/prompts/chat-executor-prompt.md`: V2 contract now uses `answer_version="executor_evidence_v2"` and removes final-expression fields.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: PASS path now tries V1 `user_facing_answer`, then renders V2 structured output to readable Chinese.
- `src/main/resources/prompts/chat-verifier-prompt.md`: `user_facing_answer` is now described as compatibility-only.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`: V2 structured output without `user_facing_answer` parses successfully.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: PASS + V2 output renders Chinese and does not expose raw JSON.
## Key Finding
The previous V1 contract intentionally kept `user_facing_answer`, but the V2 staged design intentionally removes it. This is an internal Agent contract break, mitigated by a temporary renderer until Composer is implemented.
@@ -0,0 +1,58 @@
# Acceptance
## Implementation Result
Implemented stage three of Executor Structured Output V2:
- Verifier prompt now validates claim derivability rather than scanning final natural-language output.
- Verifier output supports `claim_checks`.
- `ChatService` derives compatibility `facts_checked` from `claim_checks`.
- `ChatService` persists both `claim_checks` and `facts_checked`.
- Effective verdict guardrails prevent Gatekeeper failures and malformed structured output from remaining `PASS`.
- Verifier logging summarizes `claim_checks`.
## Static Verification
- Reviewed `git diff --stat` and changed files are scoped to stage three implementation, tests, OpenSpec/devflow, and the issue handoff document.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
- After OpenSpec archive, `cmd /c openspec validate --specs` passed.
## Script Verification
Passed:
```powershell
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Coverage:
- Verifier payload and Gatekeeper hook behavior.
- Gatekeeper schema/invocation validation.
- `claim_checks` parsing and persistence.
- `claim_checks` to `facts_checked` compatibility mapping.
- all claim verification mapping classes.
- Gatekeeper failure downgrade from model `PASS`.
- malformed Executor output downgrade from model `PASS`.
The targeted Maven test command was re-run after OpenSpec archive and passed.
## Browser / Manual Verification
Not run. This stage changes backend prompt, parser, audit, and tests only.
## OpenSpec Archive Status
Archived:
```text
openspec/changes/archive/2026-07-07-executor-verifier-claim-checks
```
Archive follow-up: the generated canonical `chat-verifier-agent` spec was reviewed and amended to preserve pre-existing verifier input and Gatekeeper schema scenarios while adding the new claim-check scenarios.
## Remaining Risks
- Composer is not implemented in this stage; final PASS rendering still uses the temporary V2 renderer until stage four.
- Full eval fixture expansion is deferred to stage five.
- Existing Maven warnings about duplicate test dependency and Lombok builder defaults remain outside this stage.
@@ -0,0 +1,41 @@
# Executor Verifier Claim Checks
## Background
Stage one moved Chat Executor to `executor_evidence_v2`, and stage two added deterministic Gatekeeper checks before Verifier. After those stages, Verifier still primarily used the legacy `facts_checked` contract and could still treat `executor_final_answer` as a fact source.
That left two risks:
- Verifier could still extract extra confirmed facts from natural-language Executor output.
- Downstream audit and retry consumers could not distinguish V2 claim-level verification from legacy fact checks.
## Goal
Make Verifier V2 claim-oriented:
- verify `executor_structured_output.claims` as the primary target;
- emit `claim_checks` as the authoritative V2 result;
- keep `facts_checked` only as a compatibility projection;
- enforce code-side guardrails so Gatekeeper failures or malformed structured output cannot remain effective `PASS`.
## Scope
- Updated `chat-verifier-prompt.md` to frame verification as claim derivability.
- Extended `ChatService` to parse, normalize, map, and persist `claim_checks`.
- Added effective verdict guardrails for Gatekeeper failure and malformed/missing Executor structured output.
- Updated verifier logging summaries to count `claim_checks`.
- Updated sequential workflow tests to cover claim mapping and downgrade behavior.
## Non-Goals
- No Composer integration in this phase.
- No final-answer material filtering beyond existing templates and temporary V2 renderer.
- No Executor retry behavior change.
- No Gatekeeper rule expansion.
- No database schema migration.
## OpenSpec
- Active change before archive: `openspec/changes/executor-verifier-claim-checks`
- Capability: `chat-verifier-agent`
- Scale: standard
@@ -0,0 +1,57 @@
# Decisions
## Scope Decision
Stage three is limited to Verifier V2 claim checks. Composer is explicitly deferred to stage four.
Reason: Composer requires stable verifier output and allowed-material filtering; mixing it into this stage would make rollback and acceptance unclear.
## Contract Decision
`claim_checks` is the authoritative V2 verifier output.
`facts_checked` remains as a compatibility projection generated from `claim_checks` when present.
Reason: existing low-confidence rendering, retry context, trace output, and evaluation code still depend on `facts_checked`.
## Mapping Decision
Claim verification maps to legacy facts as follows:
| claim verification | legacy facts_checked verification |
|---|---|
| `direct_observation` | `direct_evidence` |
| `reasonable_inference` | `indirect_support` |
| `overstated` | `indirect_support` |
| `unsupported` | `no_evidence` |
| `external_unknown` | `no_evidence` |
| `contradicted` | `contradicted` |
## Guardrail Decision
Effective verdict is enforced in code:
- `gatekeeper_result.status=fail` cannot remain `PASS`.
- `evidence.invocation_ref` failure downgrades to `REJECT`.
- Other Gatekeeper failures downgrade at least to `LOW_CONFID`.
- missing/malformed Executor structured output cannot remain `PASS`.
Reason: prompt compliance is not deterministic enough for safety-critical evidence attribution.
## Apply Fix Record
Initial targeted Maven verification failed because older tests expected PASS to return Executor natural-language output or V1 `user_facing_answer`.
Classification: test drift from the committed OpenSpec, not a design blocker.
Resolution: update tests to use valid Executor V2 output for PASS paths and assert downgrade behavior for malformed or Gatekeeper-failed outputs.
## Interface Impact
L2 internal contract extension:
- Verifier output gains `claim_checks`.
- Existing `facts_checked` remains available.
- Persistence JSON gains `claim_checks` under existing `self_evaluation`.
No external API or database schema changes.
@@ -0,0 +1,35 @@
# Evidence
## Relevant History
- `executor-v2-output-contract`: Executor emits `executor_evidence_v2` and no longer emits final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates schema and invocation references before Verifier and persists `gatekeeper_result`.
- `chat-verifier-agent`: Existing Verifier used `facts_checked`, low-confidence rendering, retry context, and verifier audit.
## Code Evidence
- `src/main/resources/prompts/chat-verifier-prompt.md`: Verifier prompt now makes `executor_structured_output.claims` primary and treats `executor_final_answer` as debug/fallback only.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: parses `claim_checks`, maps them to compatibility `facts_checked`, persists both, and applies effective verdict guardrails.
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`: verifier thought summaries now include `claim_checks`.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: covers claim-check mapping, Gatekeeper downgrade, malformed output downgrade, and V2 PASS paths.
## Evidence-Driven Conclusions
- `facts_checked` cannot be removed yet because existing low-confidence templates, retry context, trace tooling, and eval paths still consume it.
- Prompt-only prevention is insufficient for Gatekeeper failures; `ChatService` must enforce effective verdict downgrades in code.
- No database schema migration is needed because `claim_checks` is persisted inside existing `diagnosis_session.self_evaluation`.
- Composer remains stage four and must not be mixed into this stage.
## Verification Evidence
Script verification passed:
```powershell
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-verifier-claim-checks
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -0,0 +1,48 @@
# diagnosis-eval-demo-gatekeeper-closure Acceptance
## Static / Structure Verification
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`
- Result: passed.
- `cmd /c openspec validate --specs`
- Result: passed, 10 specs passed.
## Script Verification
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
- Result: 36 tests, 0 failures, 0 errors.
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
- Result: 61 tests, 0 failures, 0 errors.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result after E2E startup fix: 23 tests, 0 failures, 0 errors.
## Live E2E Verification
- Start command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`.
- Demo command: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`.
- Result: chat, trace, and feedback requests completed successfully.
- Output files:
- `mvp/demo/output/chat-response.json`
- `mvp/demo/output/trace-response.json`
- `mvp/demo/output/feedback-response.json`
- Trace observations:
- `hasVerifierEvaluation=true`
- `gatekeeper_result.rule_set_version=gatekeeper-rules-v1`
## Fixed During Verification
- E2E startup initially failed because Spring could not instantiate `ExecutorGatekeeperService`.
- Root cause: two public constructors and no explicit `@Autowired` constructor.
- Fix: annotate the production constructor with `@Autowired`.
## Residual Risk
- The live payment-timeout path can still produce `LOW_CONFID` because model-generated evidence bindings may omit some explicit `source_invocation_id` values.
- This is not a blocker for this change because deterministic matrix behavior is covered by saved fixtures and baseline evaluation.
- Existing Maven warnings remain: duplicate `spring-boot-starter-test` declaration and Lombok `@Builder` default warnings.
## Archive Status
- Devflow archive artifacts created.
- OpenSpec change archived to `openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure`.
- Main specs synced by `cmd /c openspec archive diagnosis-eval-demo-gatekeeper-closure --yes`.
@@ -0,0 +1,33 @@
# diagnosis-eval-demo-gatekeeper-closure Brief
## Background
The Chat evidence pipeline already had Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The missing piece was an interview-ready acceptance story that made the anti-hallucination behavior easy to demonstrate and regress.
## Goal
Close the next three interview-readiness gaps together:
- diagnosis eval fixture matrix
- stable demo data set
- Gatekeeper rule configuration and audit version
## Scope
- Expand `mvp/eval` with matrix-oriented cases, fixtures, and baseline reports.
- Add stable demo request payloads and scenario documentation.
- Add a lightweight local Gatekeeper rule catalog with `rule_set_version` and rule metadata in `gatekeeper_result`.
- Update architecture, demo, and eval docs to describe the current implementation.
## Non-goals
- No new public HTTP endpoint.
- No new database table.
- No Planner `scope_contract`.
- No Gatekeeper retry loop.
- No remote or dynamic rule execution engine.
## OpenSpec
- Change: `openspec/changes/diagnosis-eval-demo-gatekeeper-closure`
- Interface impact: L2 internal contract change.
@@ -0,0 +1,144 @@
# diagnosis-eval-demo-gatekeeper-closure Decisions
## Clarify
- Entry summary: implement the next three interview-readiness items together: diagnosis eval fixture matrix, stable demo data set, and Gatekeeper rule configuration/audit version.
- Slug: `diagnosis-eval-demo-gatekeeper-closure`
- Devflow scale: `standard`
- Interface impact: expected L2 internal contract change because `gatekeeper_result` audit JSON will gain rule metadata/version fields.
## Context
- `devflow/index.md` used: related entries found for diagnosis eval harness, fixture expansion, MVP demo runbook, Gatekeeper hook, and verifier evidence reference fidelity.
- Relevant glossary:
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Verifier should not use skills/runbooks as incident evidence.
- `tool_invocation.retrieval_details` is the structured evidence/audit home for tool-specific details.
- Historical constraints that must enter OpenSpec:
- Diagnosis eval is offline and deterministic; no LLM-as-judge.
- Demo assets should be runnable, but fixed regression should use saved fixtures.
- Gatekeeper remains in the Verifier hook path.
- No new database table for Gatekeeper audit; use `self_evaluation.verifier_evaluation.gatekeeper_result`.
- `$.no_evidence` is a query no-hit signal, not proof that a problem is impossible.
## Question Pool
| ID | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | Terminology | evidence-driven | What names should this change use for the matrix, demo set, and Gatekeeper rule metadata? | Resolved |
| Q2 | Boundary | evidence-driven | Should this change alter public APIs, database schema, Planner output, or retry behavior? | Resolved |
| Q3 | Acceptance | evidence-driven | Which existing tests and baseline assets define the current acceptance style? | Resolved |
| Q4 | Technical | evidence-driven | Where should Gatekeeper rule metadata live with minimal implementation risk? | Pending code research |
| Q5 | Scope | user-interview | Should the stable demo set be documentation/payloads only, or should it include live E2E scripts for all scenarios? | Confirmed |
## Evidence-driven Conclusions
- Q1 conclusion: use `diagnosis eval matrix`, `stable demo scenarios`, and `Gatekeeper rule set version` as terms.
- Q2 conclusion: keep this as an internal contract change. Do not add public endpoints, tables, Planner `scope_contract`, or Gatekeeper retry.
- Q3 conclusion: existing `DiagnosisTraceEvaluatorTest`, `ExecutorGatekeeperServiceTest`, `VerifierInputHookTest`, `ToolInvocationRecorderTest`, and `mvp/eval/reports` define the current acceptance style.
- Q4 conclusion: Gatekeeper metadata should live behind a small rule catalog loaded by `ExecutorGatekeeperService`; the audit output should include a rule set version and enabled rule metadata summary, without adding tables or remote registry.
## User-interview Confirmations
- Q5 confirmed by resumed objective: complete items 1/2/3 with sm-flow, archive, submit, and run end-to-end if necessary.
- Implementation interpretation: stable demo scenarios will be fixed request payloads and runbook docs plus deterministic fixture-backed eval. Live E2E remains necessary only for at least one main path or where unit/fixture evidence is insufficient.
## OpenSpec Backfill
- Created Draft proposal at `openspec/changes/diagnosis-eval-demo-gatekeeper-closure/proposal.md`.
- Context constraints from historical devflow entries were written into the proposal.
- Scope confirmation and Gatekeeper catalog placement were written into the proposal/design.
## Current Checkpoint
- Discover completed.
- No implementation files changed yet.
## Specify / Alignment
### Cross-artifact Alignment
| Check | Status | Notes |
|---|---|---|
| brief/proposal goals -> proposal | Aligned | Proposal covers eval matrix, stable demo scenarios, and Gatekeeper rule catalog/audit version. |
| proposal scope/constraints -> design | Aligned | Design records offline deterministic eval, fixture-backed demo distinction, local rule catalog, and no new table/API. |
| design decisions -> specs/tasks | Aligned | Specs cover eval matrix, rule set version validation, demo scenarios, and Gatekeeper rule metadata; tasks cover matching implementation slices. |
| specs observable behavior -> tasks | Aligned | Each requirement has an executable task and acceptance check. |
### Interface Impact
- Level: L2 internal contract change.
- Reason: `gatekeeper_result` internal audit JSON gains `rule_set_version` and rule metadata summary. Eval case/result fields may gain optional rule set checks. No public HTTP API, database schema, or external DTO contract changes.
## Audit
Input -> processing -> output chain:
```text
mvp/demo request docs + mvp/eval fixtures
-> DiagnosisTraceEvaluator
-> baseline reports
-> interview/demo evidence
Gatekeeper rule catalog
-> ExecutorGatekeeperService
-> VerifierInputHook / ChatService persisted self_evaluation
-> Trace and eval audit
```
Architecture risk assessment:
1. The change is intentionally internal and should not add new public consumers.
2. Gatekeeper catalog must stay metadata-only; dynamic rule execution would be a different, riskier architecture.
3. Fixture-backed demo scenarios should be documented as deterministic regression artifacts, not live LLM guarantees.
4. Baseline report churn is expected and must be committed with case/fixture changes.
5. No devflow/OpenSpec conflict found.
## Commit Gate
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`: passed.
- `cmd /c openspec validate --specs`: passed, 10 specs passed.
- File completeness:
- proposal.md: present.
- design.md: present.
- specs: present for `diagnosis-eval-harness`, `mvp-demo-trace-acceptance`, `chat-verifier-agent`.
- tasks.md: present.
- Consistency:
- Proposal concepts have corresponding design sections.
- Design decisions are reflected in specs/tasks.
- Task acceptance checks are verifiable.
## Current Checkpoint
- Commit completed.
- `.committed` marker created.
## Apply Verification
- Focused verification passed:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
- Result: 36 tests, 0 failures, 0 errors.
- Broader relevant regression passed:
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
- Result: 61 tests, 0 failures, 0 errors.
- E2E startup repro found a Spring bean construction issue:
- Command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
- Failure: `ExecutorGatekeeperService` had two public constructors and no annotated constructor, so Spring attempted a no-arg constructor and failed with `No default constructor found`.
- Classification: code deviation from OpenSpec implementation intent, not a spec gap.
- Fix: annotate the production constructor with `@Autowired`.
- Post-fix focused regression passed:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result: 23 tests, 0 failures, 0 errors.
- Live E2E passed for demo compatibility:
- Start: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
- Run: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`
- Result: `/api/chat`, `/api/diagnosis/{sessionId}/trace`, and `/api/feedback` completed successfully.
- Trace summary included `hasVerifierEvaluation=true`.
- Persisted Gatekeeper audit included `rule_set_version=gatekeeper-rules-v1`.
- Residual quality note: the live payment-timeout response remained `LOW_CONFID` because some model-produced evidence bindings still lacked explicit `source_invocation_id`; deterministic PASS/LOW_CONFID/REJECT claims are covered by fixture-backed eval.
## Archive Readiness
- OpenSpec tasks 1-4 completed.
- Verification is recorded in devflow acceptance artifacts.
- Remaining known risk: live LLM output is not deterministic and may still produce LOW_CONFID on the payment-timeout path; this is intentionally documented as demo compatibility, not a fixed PASS guarantee.
@@ -0,0 +1,22 @@
# diagnosis-eval-demo-gatekeeper-closure Evidence
## Code And Artifact Evidence
- Gatekeeper rule metadata lives in `src/main/resources/gatekeeper/gatekeeper-rules.json`.
- `ExecutorGatekeeperService` loads the local catalog, uses configured threshold parameters, and emits `rule_set_version` plus enabled rule metadata.
- `VerifierInputHook` and `ChatService` preserve Gatekeeper audit metadata in fallback/default paths.
- `DiagnosisTraceEvaluator` can optionally validate expected Gatekeeper rule set version.
- `mvp/eval/cases/diagnosis-cases.json` now includes narrow-scope and no-evidence matrix cases.
- `mvp/eval/reports/baseline-report.json` and `.md` were regenerated for the expanded fixed matrix.
- `mvp/demo/evidence-pipeline-scenarios.md` documents live vs fixture-backed demo scenarios.
## Decisions
- Keep this phase internal: no public API, no DB schema, no Planner output change.
- Keep Gatekeeper deterministic Java validation; the catalog is metadata/config only.
- Treat live demo as compatibility evidence and fixture-backed eval as deterministic regression evidence.
- Persist audit under the existing `self_evaluation.verifier_evaluation.gatekeeper_result` structure.
## Runtime Finding
The first Maven E2E startup found a real integration issue: `ExecutorGatekeeperService` had multiple public constructors without an annotated constructor, so Spring could not instantiate the service. The fix was to annotate the production constructor with `@Autowired`.
@@ -0,0 +1,57 @@
# Acceptance
## Implementation Result
Implemented stage four of Executor Structured Output V2:
- Added `chat_composer` prompt and Agent.
- Final answers for PASS, LOW_CONFID, and REJECT now use Composer when Verifier decision is valid.
- Composer input is filtered from Verifier decision and Executor structured output.
- Unsupported, external-unknown, and contradicted claims are excluded from confirmed final-answer material.
- REJECT Composer input has `allowed_hypotheses=[]`.
- Malformed Composer output uses deterministic safe fallback.
- Fallback does not expose raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
- `composer_output` is persisted in verifier audit.
## Static Verification
- Reviewed implementation diff for stage-four scope.
- `cmd /c openspec validate executor-composer-final-answer` passed.
- `cmd /c openspec validate --specs` passed before archive.
## Script Verification
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Coverage:
- Composer prompt loading and invocation.
- valid Composer output as final answer source.
- malformed Composer output fallback.
- PASS no raw Executor JSON leakage.
- LOW_CONFID separation of confirmed material, possible directions, and gaps.
- REJECT safe output without raw Executor answer.
- Gatekeeper and Verifier stage compatibility.
## Browser / Manual Verification
Not run. This stage changes backend prompt, routing, parser, audit, and tests only.
## OpenSpec Archive Status
Archived:
```text
openspec/changes/archive/2026-07-08-executor-composer-final-answer
```
## Remaining Risks
- Stage five still needs broader eval fixture coverage for full evidence-attribution regressions.
- Composer prompt quality can be improved after real run traces are collected.
- Existing Maven warnings about duplicate test dependency and Lombok builder defaults remain outside this stage.
@@ -0,0 +1,54 @@
# Executor Composer Final Answer
## Background
Stages one to three moved the Chat diagnosis chain to structured Executor output, deterministic Gatekeeper validation, and Verifier `claim_checks`.
Before this stage, `ChatService` still owned final answer rendering. PASS paths could use a temporary V2 renderer, while LOW_CONFID and REJECT paths used templates. That left final user-facing expression too close to Executor material and made it harder to prove that only Verifier-allowed claims reached the user.
## Goal
Add a Composer expression layer after Verifier:
```text
chat_planner
-> chat_executor
-> VerifierInputHook + Gatekeeper
-> chat_verifier
-> chat_composer
-> final answer
```
Composer produces user-facing answers from filtered material only:
- `allowed_claims`
- `allowed_hypotheses`
- `missing_info`
- `recommended_actions`
- `rationale`
## Scope
- Added `chat-composer-prompt.md`.
- Added `chat_composer` Agent construction in `ChatService`.
- Added Composer input filtering from Verifier decision and Executor structured output.
- Replaced PASS temporary V2 renderer usage with Composer-or-safe-fallback rendering.
- Routed LOW_CONFID and REJECT final answers through Composer when Verifier output is valid.
- Added deterministic fallback for malformed Composer output.
- Persisted `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated sequential workflow tests.
## Non-Goals
- No Planner changes.
- No Executor retry changes.
- No Gatekeeper rule expansion.
- No Verifier classification expansion.
- No database schema migration.
- No stage-five eval fixture expansion.
## OpenSpec
- Active change before archive: `openspec/changes/executor-composer-final-answer`
- Capabilities: `chat-composer-agent`, `chat-verifier-agent`
- Scale: standard
@@ -0,0 +1,70 @@
# Decisions
## Scope Decision
Stage four is limited to Composer final-answer generation and routing.
Reason: Gatekeeper and Verifier contracts were stabilized in earlier stages; this phase should only close the final-expression path.
## Composer Responsibility
Composer is an expression layer, not a diagnosis layer.
It may rephrase and organize only filtered material. It must not call tools, introduce new facts, rejudge root cause, or read raw Executor/tool output.
## Filtering Decision
`ChatService` owns Composer input filtering:
| Verifier classification | Composer handling |
|---|---|
| `direct_observation` | `allowed_claims` |
| `reasonable_inference` | `allowed_claims`, with bounded wording |
| `overstated` | `allowed_hypotheses` or `missing_info` |
| `unsupported` | `missing_info` |
| `external_unknown` | `missing_info` |
| `contradicted` | `missing_info` / REJECT-safe output |
For REJECT, `allowed_hypotheses` is always empty.
## Fallback Decision
Malformed Composer output falls back to deterministic rendering from filtered Composer input.
Fallback must never expose:
- raw Composer JSON;
- raw Executor JSON;
- Executor `user_facing_answer`;
- full unscreened tool output.
## Audit Decision
No new table is added. Composer output is persisted under:
```text
diagnosis_session.self_evaluation.verifier_evaluation.composer_output
```
The audit snapshot is intentionally compact and stores status plus parsed user-facing fields.
## Apply Fix Record
Initial targeted verification exposed test drift:
- test file had a UTF-8 BOM and failed Java compilation;
- scripted chat model did not recognize `COMPOSER_TEST_PROMPT`;
- older tests expected three-Agent execution and temporary V2 renderer behavior;
- LOW_CONFID assertions required indirect support to disappear instead of appearing as a possible direction.
Resolution: remove BOM, add Composer script branch, and update assertions to match the committed Composer contract.
## Interface Impact
L2 internal behavior change:
- external Chat API still returns a final answer string;
- internal final-answer source changes from Executor/temporary renderer to Composer or safe fallback;
- audit JSON gains `composer_output` under existing `self_evaluation`.
No database schema change.
@@ -0,0 +1,36 @@
# Evidence
## Relevant History
- `executor-v2-output-contract`: Executor emits structured diagnostic material and no final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates deterministic evidence failures before Verifier.
- `executor-verifier-claim-checks`: Verifier emits `claim_checks` and effective verdict guardrails.
## Code Evidence
- `src/main/resources/prompts/chat-composer-prompt.md`: defines Composer as an expression layer with strict JSON output.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: loads Composer prompt, invokes `chat_composer`, filters Composer input, parses Composer output, falls back safely, and persists Composer audit.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: covers Composer invocation, fallback, REJECT/LOW_CONFID behavior, and no raw JSON leakage.
## Evidence-Driven Conclusions
- Composer must be after Verifier because Verifier `claim_checks` are the authority for allowed final-answer material.
- Composer must not receive raw tool output or full unscreened Executor output because that would re-open the evidence attribution problem.
- Verifier malformed/missing output should not invoke Composer because there is no trustworthy decision to filter with.
- Fixed fallback remains necessary because Composer is an LLM call with a strict JSON contract and can produce malformed output.
## Verification Evidence
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-composer-final-answer
cmd /c openspec validate --specs
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -0,0 +1,52 @@
# Acceptance
## Static Verification
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
- `openspec validate --specs`: passed.
## Script Verification
- `mvn "-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,QueryLogsToolsTest,ChatServiceSequentialAgentTest" test`
- Passed: 38 tests.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest" test`
- Passed: 9 tests.
- `mvn test`
- Failed on unrelated environment-gated `MilvusConnectionTest.connect`: `MILVUS_TOKEN` environment variable was not set.
- Other executed tests in the run progressed until that single failure; focused tests for this change passed.
## End-to-End Verification
The Java service was restarted with `mvn spring-boot:run`; logs were written under `logs/`.
| Case | Session | Result | Gatekeeper Audit |
|---|---|---|---|
| HikariCP positive | `iss007-hikari-positive-20260708-1553` | PASS; confirmed order-service HikariCP timeout and pool saturation logs | `pass / none`, checked_bindings=2 |
| HikariCP negative | `iss007-hikari-negative-20260708-1555` | LOW_CONFID; no `generic-service`; no false positive for inventory-service | `fail / low_confid` |
| HighMemoryUsage positive | `iss007-memory-positive-20260708-1558` | PASS; confirmed HighMemoryUsage 91%, did not confirm memory leak | `pass / none`, checked_bindings=1 |
| SlowResponse positive | `iss007-slow-positive-20260708-1600` | PASS; confirmed SlowResponse and slow request logs, no DB pool root cause | `pass / none`, checked_bindings=7 |
| Narrow HighCPUUsage | `iss007-narrow-highcpu-20260708-1602` | PASS; only covered payment-service HighCPUUsage | `pass / none`, checked_bindings=1 |
## Database Audit
`scripts/query_mysql.py` was used to verify:
- `diagnosis_session.self_evaluation.verifier_evaluation.verdict`
- `gatekeeper_result.status`
- `gatekeeper_result.severity`
- `gatekeeper_result.checked_bindings`
- no-hit HikariCP query rows persist `evidence_status=no_evidence`
## Remaining Risk
- Negative no-hit claims still have incomplete precise references when Executor uses `$.logs` for empty arrays. Gatekeeper correctly downgrades to `LOW_CONFID`.
- Prompt-only scope control is improved but not a hard contract. A future `scope_contract` may still be needed.
- Full test suite requires `MILVUS_TOKEN` to pass `MilvusConnectionTest`.
## OpenSpec Archive
- `openspec archive verifier-evidence-reference-fidelity --yes`: succeeded.
- Main specs updated:
- `openspec/specs/chat-verifier-agent/spec.md`
- `openspec/specs/evidence-trace-hardening/spec.md`
- Non-blocking warning: proposal did not use OpenSpec's preferred `## Why` / `## What Changes` headers, but archive completed.
@@ -0,0 +1,44 @@
# Verifier Evidence Reference Fidelity
## Background
ISS-007 came from end-to-end diagnosis cases where raw tool output and Executor `evidence_excerpt` contained enough facts, but Verifier still returned `LOW_CONFID` because the verifier-facing summary compressed away key details.
The affected flow is:
```text
chat_planner
-> chat_executor
-> VerifierInputHook / Gatekeeper
-> chat_verifier
-> chat_composer
```
The change hardens the evidence handoff between Executor, Gatekeeper, and Verifier.
## Goal
Make Executor cite concrete tool evidence, make Gatekeeper validate that citation with code, and make Verifier judge whether verified evidence can derive the claim.
## Scope
- Persist `tool_invocation.retrieval_details.evidence_refs`.
- Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence binding.
- Add Gatekeeper `severity` and checked binding audit.
- Keep Gatekeeper in the Verifier hook path.
- Keep `tool_trace_summary` as navigation/audit context, not the only evidence source.
- Fix HikariCP mock positive/no-hit behavior.
- Tighten Executor/Verifier prompts for narrow-scope evidence handling.
## Non-goals
- No Planner `scope_contract`.
- No new database table.
- No full JSONPath engine.
- No change to external HTTP API.
- No retry rollback from Gatekeeper to Executor in this phase.
## OpenSpec
- Change: `openspec/changes/verifier-evidence-reference-fidelity`
- Source issue: `mvp/issues/archived/ISS-007-verifier-evidence-summary-fidelity.md`
@@ -0,0 +1,54 @@
# Decisions
## Evidence Reference
Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence reference for Executor claim bindings.
Reason:
- Invocation ID alone only identifies a tool call, not the evidence inside it.
- `raw_path` is enough for the first version when paired with `retrieval_details.evidence_refs`.
- `evidence_excerpt` remains the text Verifier reads, but only after Gatekeeper validates it.
## Raw Path
Only support stable locators in the first version:
- `$.alerts[i]`
- `$.logs[i]`
- `$.evidence_blocks[i]`
No full JSONPath engine is introduced.
## Gatekeeper Severity
Gatekeeper output includes:
- `status`
- `severity`
- `checked_bindings`
- `failed_rules`
- `warnings`
- `errors`
Severity meaning:
- `none`: precise references passed.
- `low_confid`: evidence is missing or incomplete, but not fabricated.
- `reject`: fabricated ID, wrong tool, unknown raw path, or mismatched excerpt.
## Verifier Boundary
Verifier uses verified claim-local excerpts as primary derivability evidence. `tool_trace_summary` remains available for navigation and audit, but no longer needs to carry every concrete fact.
## Hook Placement
Gatekeeper remains in the Verifier input hook path. This version does not retry Executor on Gatekeeper failure.
## Planner
Planner is not changed. `scope_contract` remains a later-stage idea. This phase uses prompt constraints to reduce narrow-scope over-expansion.
## Database
No new tables. Evidence refs are stored in `tool_invocation.retrieval_details.evidence_refs`; audit is stored in `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
@@ -0,0 +1,33 @@
# Evidence
## Existing Context
- Existing `chat-verifier-agent` spec still used `source_invocation_ids` and `tool_trace_summary` as the main verifier evidence context.
- Existing `evidence-trace-hardening` spec already established `tool_invocation.retrieval_details` as the right place for structured tool-specific facts.
- Prior devflow projects established that runbook/skill content is guidance, not incident evidence.
## Code Findings
- `VerifierInputHook` previously backfilled plural `source_invocation_ids` from `tool_trace_summary` by tool name.
- `ExecutorGatekeeperService` previously validated invocation existence and tool name, but not `raw_path` or excerpt authenticity.
- `ToolInvocationRecorder` persisted retrieval details but did not generate claim-addressable `evidence_refs`.
- `QueryLogsTools` could fall back to `generic-service` placeholder logs on no-hit.
## Implementation Evidence
- `ToolInvocationRecorder` now extracts:
- `$.alerts[i]` for `query_metrics`
- `$.logs[i]` for `query_logs`
- `$.evidence_blocks[i]` for `lookup_knowledge`
- `ExecutorGatekeeperService` now validates:
- invocation existence
- tool name
- raw path presence
- `retrieval_details.evidence_refs`
- excerpt similarity/support
- `VerifierInputHook` only auto-fills a singular `source_invocation_id` when exactly one candidate exists and never invents `raw_path`.
- `QueryLogsTools` returns HikariCP mock logs for `order-service` and returns empty no-hit results for unrelated services.
## Residual Finding
The HikariCP negative E2E no longer has generic-service pollution, but the model still issued an extra broad HikariCP query without the service filter and used order-service as context. This is a remaining narrow-scope behavior issue, not a mock evidence pollution issue.
@@ -0,0 +1,46 @@
# Acceptance
## Static Verification
- `openspec validate interview-demo-quality-audit --strict`
- Result: passed.
- Coverage: OpenSpec proposal/design/spec/tasks consistency.
- PowerShell parser/runtime readiness check:
- Command: `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://127.0.0.1:1 -OutputDir target/demo-check-syntax`
- Result: expected failure with actionable readiness message.
- Coverage: script parses under Windows PowerShell and fails before issuing diagnosis requests when service is unreachable.
## Script Verification
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed.
- Coverage: 12/12 fixed eval fixtures, Prompt audit evaluator checks, Gatekeeper rule metadata checks, regenerated baseline reports.
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: Chat verifier evaluation persists `prompt_audit`.
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result: passed.
- Coverage: broader eval, baseline diff, Chat sequential flow, Gatekeeper, and Verifier input hook regression set.
- `mvn -q -DskipTests compile`
- Result: passed.
- Coverage: main source compilation.
## E2E Verification
- Started service with:
- `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`
- Ran:
- `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://localhost:9900 -SessionId mvp-demo-interview-quality-audit-001`
- Result: passed.
- Summary:
- `chatSuccess=true`
- `verdict=LOW_CONFID`
- `gatekeeperStatus=fail`
- `gatekeeperRuleSetVersion=gatekeeper-rules-v1`
- `promptAuditVersion=chat-prompts-v1`
- tools included `lookup_knowledge`, `query_logs`, `query_metrics`, and `get_available_log_topics`
- Note: live E2E remains a compatibility check, not the deterministic PASS oracle. The fixed fixture baseline is the regression source of truth.
## Not Verified
- Browser UI inspection was not required for this change because the scope is backend trace/eval/demo script documentation, not frontend behavior.
@@ -0,0 +1,23 @@
# Interview Demo Quality Audit Brief
## Background
The MVP already demonstrates traceable Agent diagnosis with Planner, Executor, Gatekeeper, Verifier, Composer, evidence tools, trace persistence, and deterministic eval fixtures. The remaining interview-readiness gap is not a new Agent architecture; it is making the demo easier to run and making prompt/rule changes easier to audit.
## Goal
Stabilize the interview demo path, expand fixture-backed evaluation, and persist prompt/Gatekeeper audit metadata so the project can explain and verify Agent behavior during interviews.
## Scope
- Add prompt audit metadata to Chat verifier evaluation.
- Extend deterministic eval cases and baseline reports.
- Add an interview demo preflight/check script.
- Update MVP demo and architecture documentation.
## Non-goals
- No public API or database schema changes.
- No new SubAgent split, MCP migration, process isolation, or AIOps LLM Verifier.
- No guarantee that every live LLM run returns PASS.
@@ -0,0 +1,111 @@
# interview-demo-quality-audit Decisions
## Clarify
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
- Slug: `interview-demo-quality-audit`.
- Devflow scale: `standard`.
- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change.
## Context
- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`.
- Relevant glossary:
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Verifier should not use skills/runbooks as incident evidence.
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
- Historical constraints that must enter OpenSpec:
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
- Composer is the final expression layer and must not leak raw Executor JSON.
## Question Pool
| ID | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
## Evidence-driven Conclusions
- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`.
- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation.
- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
## Specify / Alignment
| Check | Status | Notes |
|---|---|---|
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
## Audit
Input -> processing -> output chain:
```text
prompt resource metadata
-> ChatService / PromptAudit snapshot
-> verifier_evaluation.prompt_audit
-> Trace API / eval fixtures
-> DiagnosisTraceEvaluator baseline
run-interview-demo-check.ps1
-> service readiness
-> chat / trace / feedback
-> mvp/demo/output summary
```
Architecture risk assessment:
1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
4. No devflow/OpenSpec conflict found.
## Commit Gate
- `openspec validate interview-demo-quality-audit --strict`: passed.
- File completeness:
- `proposal.md`: present.
- `design.md`: present.
- `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`.
- `tasks.md`: present.
- Consistency:
- Proposal goals map to design sections.
- Design decisions map to spec requirements and executable tasks.
- Task acceptance checks are verifiable.
- `.committed` marker created.
## Current Checkpoint
- Commit completed.
- Apply is authorized by the original objective: "完成后归档提交".
## Pre-apply Research
- Capability source: sm-flow built-in apply protocol. `openspec-apply-change` was not invoked directly in this session.
- Repository semantic search/LSP note: the requested `codebase-retrieval` and LSP tools were not available in the exposed toolset, so impact analysis used `rg`, direct file reads, OpenSpec/devflow artifacts, and targeted tests.
- Reference implementation and reuse:
- `ChatService.persistVerifierEvaluation(...)` is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.
- `DiagnosisTraceEvaluator` and `DiagnosisEvalReportWriter` are the deterministic eval extension points; no LLM judge was introduced.
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` provided the request/trace/feedback flow reused by the new interview preflight script.
- Interface impact remains L2 internal trace contract: `verifier_evaluation.prompt_audit` and eval report fields are added; no public endpoint, table, or request DTO changed.
## Apply Notes
- Added compact Chat prompt audit metadata: `chat-prompts-v1`, with planner/executor/verifier/composer prompt versions and resource paths.
- Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
- Added two fixture-backed audit cases:
- `prompt-gatekeeper-audit-closure`
- `audit-metadata-low-confid`
- Added `mvp/demo/scripts/run-interview-demo-check.ps1` to run service readiness, Chat, Trace, feedback, and summary output.
- Updated MVP demo/eval/architecture docs to explain `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and deterministic fixture baseline.
@@ -0,0 +1,58 @@
# Evidence
## Context Files Read
- `devflow/index.md`
- `devflow/glossary/CONTEXT.md`
- `devflow/projects/2026-07-08-diagnosis-eval-demo-gatekeeper-closure/decisions.md`
- `devflow/projects/2026-07-08-executor-composer-final-answer/decisions.md`
- `mvp/architecture/current-mvp-architecture.md`
- `mvp/architecture/agent-orchestration.md`
- `mvp/architecture/executor-evidence-pipeline-refactor.md`
- `mvp/architecture/harness-quality-gates.md`
- `mvp/demo/README.md`
- `mvp/demo/ten-minute-interview-demo.md`
- `mvp/eval/README.md`
- `mvp/eval/cases/diagnosis-cases.json`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- `src/main/resources/gatekeeper/gatekeeper-rules.json`
## Tooling Note
The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead.
## Implementation Evidence
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- Adds `prompt_audit` under `verifier_evaluation` through the shared `persistVerifierEvaluation(...)` path.
- Uses compact metadata only: audit version, prompt names, prompt versions, and resource paths.
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- Adds deterministic checks for `requirePromptAudit`, `expectedPromptAuditVersion`, `expectedPromptVersions`, and `requireGatekeeperRules`.
- `src/main/java/com/superbiz/agent/eval/DiagnosisEvalReportWriter.java`
- Adds Prompt Audit and Gatekeeper rule count columns to Markdown reports.
- `mvp/eval/cases/diagnosis-cases.json`
- Expands fixed baseline to 12 fixture-backed cases.
- `mvp/eval/fixtures/prompt-gatekeeper-audit-closure-pass.json`
- Positive PASS fixture proving Prompt audit and Gatekeeper rule metadata closure.
- `mvp/eval/fixtures/audit-metadata-low-confid.json`
- LOW_CONFID fixture proving safe answer behavior while audit metadata remains present.
- `mvp/demo/scripts/run-interview-demo-check.ps1`
- Adds service readiness, Chat, Trace, feedback, and summary output for interview preflight.
## Verification Evidence
- OpenSpec:
- `openspec validate interview-demo-quality-audit --strict`: passed before archive.
- `openspec validate --specs --strict`: 10 specs passed after merging deltas into main specs.
- Unit/eval:
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`: passed.
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`: passed.
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`: passed.
- Compile:
- `mvn -q -DskipTests compile`: passed.
- E2E:
- Started `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`.
- Ran `mvp/demo/scripts/run-interview-demo-check.ps1` against `http://localhost:9900`.
- Summary recorded `chatSuccess=true`, `verdict=LOW_CONFID`, `gatekeeperRuleSetVersion=gatekeeper-rules-v1`, and `promptAuditVersion=chat-prompts-v1`.
@@ -0,0 +1,103 @@
# Acceptance
## 实现结果
- OpenSpec tasks: `42/42` complete。
- Phase commits:
- `52bf030 feat(trace): add session run isolation schema`
- `6fdbd34 docs(openspec): tighten run isolation contract`
- `26d5529 feat(trace): isolate chat runs`
- `027aed1 feat(trace): add run-scoped trace reads`
- `d928a19 feat(trace): bind feedback to runs`
- `78c1477 feat(trace): isolate aiops runs`
- `f9df943 feat(trace): finish run-aware demo verification`
- OpenSpec archive: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
## 静态验证
```powershell
node --check src\main\resources\static\app.js
node --check src\main\resources\static\trace.js
openspec validate session-run-trace-isolation --strict
git diff --check -- . ':!devflow/index.md'
```
结果:通过。
## 脚本验证
PowerShell demo 脚本解析:
```powershell
$scripts = @(
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
'mvp\demo\scripts\run-interview-demo-check.ps1'
)
foreach ($script in $scripts) {
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
}
```
结果:通过。
Focused tests:
```powershell
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
```
结果:通过。
Baseline / regression:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
结果:通过,无 baseline drift。
## E2E 验证
使用 Maven 启动:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
E2E 使用同一 `sessionId` 连续两轮 Chat:
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
验证结果:
- Chat1 / Chat2 均成功。
- run1 exact trace 返回 run1。
- run2 exact trace 返回 run2。
- session-only latest trace 返回 run2。
- DB 中同一 session 有两条 `diagnosis_run`。
- step/tool rows 按 `run_id` 隔离,mixed row check 为 0。
- `chat_session.message_pair_count = 2`。
## 日志验证
检查:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `logs/application.log`
- `logs/chat.log`
结果:能找到 E2E `sessionId`、两个 `runId`、Chat execution、run persistence 和 trace lookup 相关日志。
## 浏览器/人工验证
未单独进行浏览器点击验证。Trace UI 的本次验收通过静态语法检查、URL/runId 参数代码审查和后端 exact trace E2E 共同覆盖。建议后续手动打开 `trace.html?sessionId=...&runId=...` 做展示层冒烟。
## 剩余风险 / 后续事项
- 缺少 `runId` 的 Feedback fallback 是短期兼容路径,客户端全部迁移后可收紧。
- `diagnosis_session` 仍保留为历史兼容和回滚表,后续需要观察窗口后再评估约束收紧或归档策略。
- `case_library.diagnosis_id` 仍是过渡字段,旧值可能为 `session_id`,新自动值为 `run_id`。
- 历史 mixed trace 不能恢复真实多轮边界,只能按 compatibility run 查询。
@@ -0,0 +1,41 @@
# Session / Run / Trace Isolation
## 背景
同一个 `sessionId` 以前同时代表多轮 Chat 上下文和一次持久化诊断 Trace。端到端验证发现,同一 `sessionId` 连续两轮 Chat 时,Redis 多轮上下文是正确的,但 MySQL 中 `diagnosis_session` 会被后一轮覆盖,`agent_step` 和 `tool_invocation` 会按同一个 `session_id` 混在一起。
这会导致 Trace 回放、Verifier/Evaluation 读数、Feedback 绑定和 `case_library` 来源都可能跨轮污染。
## 目标
- 将会话态和运行态拆开:`chat_session` 保存会话元数据,`diagnosis_run` 保存一次诊断运行。
- 引入正式 API 字段 `runId`,作为一次可回放诊断执行的边界。
- `agent_step` 和 `tool_invocation` 保留原 Trace 明细角色,新增 `run_id` 并按 run 隔离读写。
- Trace、Feedback、CaseLibrary、AIOps、demo 脚本和 Trace UI 都支持 run-aware 流程。
- 保留旧 `diagnosis_session` 作为历史兼容和回滚表。
- 完成 Maven E2E、DB 检查、日志检查和 baseline drift 验证。
## 范围
- Flyway/JPA 增加 `chat_session`、`diagnosis_run`,并给 `agent_step`、`tool_invocation` 增加 `run_id`。
- Chat 每次有效执行创建一个新的 `diagnosis_run`,响应返回 `sessionId + runId`。
- Trace API 支持 latest-run fallback 和 exact-run 查询:`GET /api/diagnosis/{sessionId}/trace?runId=...`。
- 新增 run list API:`GET /api/chat/session/{sessionId}/runs`。
- Feedback 优先绑定 `runId`,缺省时短期 fallback 到 latest run 并返回 `fallbackToLatestRun=true`。
- AIOps 每次有效执行创建并透出 `runId`,SSE 保持 `message` event name 并发送 `type=metadata`。
- MVP demo、Trace UI、表文档和架构文档统一为 `chat_session -> diagnosis_run -> trace detail(run_id)`。
## 非目标
- 不新增 `diagnosis_trace` 或 `trace_event` 主表。
- 不实现完整 run-list UI。
- 不删除旧 `diagnosis_session`。
- 不改变 Redis 对话历史窗口策略。
- 不把完整多轮正文历史持久化到 MySQL。
- 不尝试把历史混合 trace 还原成真实多轮边界。
## 关联
- OpenSpec: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
- Change slug: `session-run-trace-isolation`
- 分档: complex
@@ -0,0 +1,48 @@
# Decisions
## 核心决策
| 决策 | 选择 | 理由 |
|---|---|---|
| 领域拆分 | 新增 `chat_session` 和 `diagnosis_run` | 会话元数据和一次诊断执行的生命周期不同,继续塞在一张表会导致上下文膨胀和边界混淆 |
| Trace 明细 | 复用 `agent_step` / `tool_invocation`,增加 `run_id` | 现有明细表已经能表达 Trace,隔离需要 run key,不需要新事件模型 |
| API 身份 | `runId = "run-" + UUID` | 外部 ID 不依赖数据库自增 ID,碰撞风险低 |
| Trace 兼容 | 缺少 `runId` 时按 `created_at DESC, id DESC` 解析 latest run | 保留旧客户端兼容性,避免 feedback/eval 更新 `updated_at` 后改变 latest 判定 |
| 历史迁移 | 每条旧 `diagnosis_session` 生成一条 compatibility run | 旧混合数据没有真实轮次边界,不能伪造多 run 历史 |
| Feedback fallback | 缺少 `runId` 时短期绑定 latest run 并返回 `fallbackToLatestRun=true` | 老客户端可继续工作,同时让歧义可观测 |
| Case provenance | 新自动案例写 `case_library.diagnosis_id = run_id` | 保留旧列,文档声明过渡语义 |
| AIOps 范围 | 同一个 change 内完成 AIOps run isolation | AIOps 是一等 Trace 入口,不能留下同类混合 trace bug |
| 所有权校验 | 服务层校验 run/session ownership,暂不加 DB 外键 | 兼容历史 orphan rows 和回滚窗口 |
## 用户确认
- 选择拆 `chat_session` 和 `diagnosis_run`,不只是在旧表加字段。
- `chat_session` 第一阶段只保存元数据,不保存完整对话正文。
- 完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`。
- `runId` 是正式 API 字段。
- Trace 缺少 `runId` 时短期默认查 latest run。
- Feedback 缺少 `runId` 时短期 fallback,长期可再收紧。
- 每次有效 Chat/AIOps 都创建 run。
- 不新增 `diagnosis_trace` / `trace_event` 主表。
- 旧 `diagnosis_session` 保留用于历史和回滚,新代码不再写新执行态。
- demo 脚本和 Trace UI 做最小 `runId` 支持。
## 接口影响
级别:L4。
- 新 API 响应字段:`runId`。
- Trace API 新 query 参数:`runId`。
- 新 API:`GET /api/chat/session/{sessionId}/runs`。
- Feedback request 新增 optional/preferred `runId`。
- Feedback response 新增 bound `runId` 和 `fallbackToLatestRun`。
- `/api/ai_ops` SSE 保持 event name `message`,新增 `type=metadata` 消息。
- DB contract 新增两张表和两个 `run_id` 列。
- 旧 `sessionId` only 调用仍兼容,但 fallback 必须可观测。
## 风险接受
- 历史混合 trace 无法真实拆分,只能作为 compatibility run。
- 上下文传播同时依赖 `RunnableConfig.metadata` 和 `SessionContextHolder`,后续改动必须注意 `sessionId/runId` 同步。
- `case_library.diagnosis_id` 在过渡期存在 `session_id` 和 `run_id` 两种语义。
- 缺少 `runId` 的 Feedback 仍有歧义,后续客户端迁移完成后可收紧为参数错误。
@@ -0,0 +1,76 @@
# Evidence
## 上下文证据
- `SessionContext.messageHistory` 和 `getMessagePairCount()` 证明 Redis 承载热对话历史;MySQL 只需要长期审计的会话目录和运行记录。
- `CaseLibraryService.createFromSession` 原先按 `DiagnosisSession.sessionId` 去重并映射 query/answer,因此 run 隔离后需要新增 `createFromRun`。
- 旧 `mvp/architecture/data-model.md` 把 `case_library.diagnosis_id` 解释为 `diagnosis_session.session_id`,本次改为过渡语义:旧数据可能是 `session_id`,新自动案例是 `run_id`。
- 既有 Trace OpenSpec 要求 `GET /api/diagnosis/{sessionId}/trace` 是只读端点;latest-run 和 exact-run 查询都必须保持只读。
- ISS-010 的 E2E 事实显示同一 `sessionId` 两轮 Chat 会产生 MySQL Trace 混合,是本 change 的直接触发证据。
## 实现证据
- Phase 1 增加 `V011__add_session_run_isolation.sql`,创建 `chat_session`、`diagnosis_run`,并为 `agent_step` / `tool_invocation` 增加 nullable `run_id`。
- Phase 2 将 Chat 写路径切到 `chat_session + diagnosis_run`,并让 Hook/Tool/Evaluation/Gatekeeper 使用 run-scoped 数据。
- Phase 3 将 Trace API 改为 latest-run / exact-run 双模式,并加入 lightweight run summaries。
- Phase 4 将 Feedback 和 CaseLibrary 绑定到 run,保留没有 run-backed 数据时的 legacy fallback。
- Phase 5 将 AIOps 接入 run isolation,SSE metadata 暴露 `sessionId + runId`。
- Phase 6 更新 demo 脚本、Trace UI、MVP 架构文档和表文档,并修正 review 后发现的 session-only 文档残留。
## E2E 证据
Maven 启动命令:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
日志:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `target/e2e/phase6-mvn-20260710-211831.err.log`
- `logs/application.log`
- `logs/chat.log`
E2E session:
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
结果:
- 两轮 Chat 都成功,并复用同一个 `sessionId`。
- 两轮返回不同 `runId`。
- run1 exact trace 只返回 run1。
- run2 exact trace 只返回 run2。
- session-only Trace latest fallback 返回 run2。
- `chat_session.message_pair_count = 2`,证明多轮上下文连续。
## DB 证据
通过 `scripts/query_mysql.py` 检查:
- `diagnosis_run` 中该 E2E session 有 2 条 `SUCCESS / CHAT` 运行。
- `agent_step` 按 run 分组:run1 `10` 行,run2 `9` 行。
- `tool_invocation` 按 run 分组:run1 `14` 行,run2 `8` 行。
- mixed row check 为 `0`,没有 NULL 或 unexpected `run_id` 混入该 E2E session。
## Baseline 证据
运行:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
结果:
- 两组 baseline / regression 命令通过。
- baseline harness 使用离线 fixture,不依赖 live DB/session tables。
- 未观察到 baseline drift。
## 工具限制
AGENTS 要求的 `codebase-retrieval` 和 LSP 工具在本会话不可用。替代验证使用 OpenSpec、`rg`、定向阅读、 focused tests、E2E、DB 查询和日志检查。
+28 -18
View File
@@ -1,8 +1,8 @@
# SuperBizAgent MVP 文档 # SuperBizAgent MVP 文档
**更新日期**:2026-07-05 **更新日期**:2026-07-10
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前架构入口已经整理到 `mvp/architecture/`,旧版架构材料已归档,避免继续把历史方案当成当前实现。 本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
## 当前入口 ## 当前入口
@@ -12,6 +12,7 @@
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 | | [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 | | [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 | | [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
| [architecture/executor-evidence-pipeline-refactor.md](architecture/executor-evidence-pipeline-refactor.md) | Executor 证据链路改造记录 |
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 | | [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 | | [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 | | [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
@@ -20,15 +21,16 @@
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 | | [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 | | [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 | | [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
| [issues/rag-refactor-plan.md](issues/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 | | [issues/README.md](issues/README.md) | MVP issue 索引 |
| [issues/active/rag-refactor-plan.md](issues/active/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
| [tables/README.md](tables/README.md) | 当前 MySQL 表说明 |
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 | | [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 | | [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
| [eval/README.md](eval/README.md) | 诊断评测材料 | | [eval/README.md](eval/README.md) | 诊断评测材料 |
| [issues/README.md](issues/README.md) | MVP issue 索引 |
## 当前系统一句话 ## 当前系统一句话
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据,诊断过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。 SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
## 文档结构 ## 文档结构
@@ -39,6 +41,7 @@ mvp/
current-mvp-architecture.md current-mvp-architecture.md
interview-one-pager.md interview-one-pager.md
agent-orchestration.md agent-orchestration.md
executor-evidence-pipeline-refactor.md
harness-quality-gates.md harness-quality-gates.md
rag-architecture.md rag-architecture.md
retrieval-observability.md retrieval-observability.md
@@ -47,12 +50,17 @@ mvp/
knowledge-base-authoring.md knowledge-base-authoring.md
data-model.md data-model.md
evolution-roadmap.md evolution-roadmap.md
archive/2026-07-05-legacy/ archive/
issues/ issues/
README.md README.md
rag-refactor-plan.md active/
ISS-*.md archived/
rag-*.md design-notes/
rag/
tables/
README.md
*表-*.md
archive/
demo/ demo/
README.md README.md
ten-minute-interview-demo.md ten-minute-interview-demo.md
@@ -65,9 +73,7 @@ mvp/
cases/ cases/
fixtures/ fixtures/
reports/ reports/
notes/ archive/
plan/
tables/
``` ```
## 当前核心设计 ## 当前核心设计
@@ -77,7 +83,8 @@ mvp/
- `VectorSearchService` 是检索稳定门面。 - `VectorSearchService` 是检索稳定门面。
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。 - Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
- AIOps payload 会生成推荐知识库 query,保留业务语义。 - AIOps payload 会生成推荐知识库 query,保留业务语义。
- Trace API 聚合 session、step、tool invocation 和 self evaluation。 - `sessionId` 表示多轮会话上下文,`runId` 表示一次可回放诊断运行。
- Trace API 聚合 `diagnosis_run`、`agent_step.run_id`、`tool_invocation.run_id` 和 self evaluation。
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。 - RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
## 关键运行链路 ## 关键运行链路
@@ -87,7 +94,8 @@ Chat
-> ChatService -> ChatService
-> Planner / Executor / Verifier -> Planner / Executor / Verifier
-> evidence tools -> evidence tools
-> diagnosis_session / agent_step / tool_invocation -> chat_session / diagnosis_run
-> agent_step.run_id / tool_invocation.run_id
-> DiagnosisTraceService -> DiagnosisTraceService
AIOps AIOps
@@ -96,6 +104,7 @@ AIOps
-> Planner / Executor -> Planner / Executor
-> Prometheus / logs / lookup_knowledge -> Prometheus / logs / lookup_knowledge
-> AiOpsRuleEvaluationService -> AiOpsRuleEvaluationService
-> diagnosis_run(agent_flow=AI_OPS)
-> DiagnosisTraceService -> DiagnosisTraceService
RAG RAG
@@ -107,10 +116,11 @@ RAG
-> tool_invocation -> tool_invocation
``` ```
## 旧文档说明 ## 归档说明
旧版架构文档已移动到: 历史材料分两类:
- [architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/) - 旧架构文档:[architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
- 本次文档清理归档:[archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/)
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/current-mvp-architecture.md` 与 `architecture/rag-architecture.md` 为准。 归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/`、`issues/README.md`、`tables/README.md` 和 OpenSpec/devflow 的最新记录为准。
+14 -12
View File
@@ -1,6 +1,6 @@
# MVP 架构文档 # MVP 架构文档
**更新日期**:2026-07-06 **更新日期**:2026-07-10
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到: 这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
@@ -15,7 +15,8 @@
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 | | [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 | | [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 | | [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 | | [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 | | [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace | | [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance | | [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
@@ -28,19 +29,20 @@
## 当前架构一句话 ## 当前架构一句话
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。 SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
## 阅读顺序 ## 阅读顺序
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。 1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。 2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。 3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。 4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。 5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
6. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。 6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
7. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。 7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
8. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。 8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
9. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。 9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
10. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。 10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
11. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。 11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
12. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。 12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
+44 -21
View File
@@ -1,14 +1,14 @@
# Agent 编排架构 # Agent 编排架构
**更新日期**:2026-07-05 **更新日期**:2026-07-08
**状态**:当前可运行架构 **状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md` **参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 设计定位 ## 1. 设计定位
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛: 旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Verifier`。 - Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。 - AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。 - 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。 - 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
@@ -23,9 +23,11 @@ flowchart TB
ChatPlanner --> ChatExecutor["chat_executor"] ChatPlanner --> ChatExecutor["chat_executor"]
ChatExecutor --> ChatTools["evidence tools"] ChatExecutor --> ChatTools["evidence tools"]
ChatTools --> ChatExecutor ChatTools --> ChatExecutor
ChatExecutor --> ChatVerifier["chat_verifier"] ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
ChatGatekeeper --> ChatVerifier["chat_verifier"]
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"} ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
ChatDecision --> ChatAnswer["final answer"] ChatDecision --> ChatComposer["chat_composer"]
ChatComposer --> ChatAnswer["final answer"]
end end
subgraph AiOps["AIOps diagnosis"] subgraph AiOps["AIOps diagnosis"]
@@ -40,20 +42,25 @@ flowchart TB
end end
subgraph Trace["Trace persistence"] subgraph Trace["Trace persistence"]
Session["diagnosis_session"] ChatSession["chat_session"]
Run["diagnosis_run"]
Step["agent_step"] Step["agent_step"]
Invocation["tool_invocation"] Invocation["tool_invocation"]
SelfEval["self_evaluation"] SelfEval["self_evaluation"]
end end
ChatService --> Session ChatService --> ChatSession
ChatService --> Run
ChatPlanner --> Step ChatPlanner --> Step
ChatExecutor --> Step ChatExecutor --> Step
ChatGatekeeper --> SelfEval
ChatVerifier --> Step ChatVerifier --> Step
ChatTools --> Invocation ChatTools --> Invocation
ChatDecision --> SelfEval ChatDecision --> SelfEval
ChatComposer --> Step
AiOpsService --> Session AiOpsService --> ChatSession
AiOpsService --> Run
AiOpsPlanner --> Step AiOpsPlanner --> Step
AiOpsExecutor --> Step AiOpsExecutor --> Step
AiOpsTools --> Invocation AiOpsTools --> Invocation
@@ -68,9 +75,13 @@ Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
chat_planner chat_planner
-> chat_executor -> chat_executor
-> lookup_knowledge / query_logs / query_metrics / date_time -> lookup_knowledge / query_logs / query_metrics / date_time
-> outputs executor_evidence_v2
-> VerifierInputHook / ExecutorGatekeeperService
-> validates source_invocation_id / raw_path / evidence_excerpt
-> chat_verifier -> chat_verifier
-> reads tool_trace_summary -> judges whether verified evidence can derive claims
-> outputs verifier JSON -> chat_composer
-> writes final user-facing answer
``` ```
关键行为: 关键行为:
@@ -78,8 +89,10 @@ chat_planner
| 角色 | 当前职责 | 输出 | | 角色 | 当前职责 | 输出 |
|---|---|---| |---|---|---|
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` | | `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
| `chat_executor` | 按计划调用证据工具,组合工具返回形成诊断答复 | `executor_feedback` | | `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
| `chat_verifier` | 只基于已有证据校验 Executor 答案,不做新检索 | `verifier_output` | | `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
Chat 链路最多支持两轮验证: Chat 链路最多支持两轮验证:
@@ -90,22 +103,28 @@ sequenceDiagram
participant P as chat_planner participant P as chat_planner
participant E as chat_executor participant E as chat_executor
participant T as tools participant T as tools
participant G as gatekeeper
participant V as chat_verifier participant V as chat_verifier
participant S as diagnosis_session participant M as chat_composer
participant R as diagnosis_run
C->>P: 原始问题 + history + retry_context C->>P: 原始问题 + history + retry_context
P-->>C: planner_plan P-->>C: planner_plan
C->>E: planner_plan + 上下文 C->>E: planner_plan + 上下文
E->>T: 调用证据工具 E->>T: 调用证据工具
T-->>E: 证据结果 T-->>E: 证据结果
E-->>C: executor_feedback E-->>C: executor_evidence_v2
C->>V: executor_final_answer + tool_trace_summary C->>G: executor_structured_output + tool_invocation.evidence_refs
G-->>C: gatekeeper_result
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
V-->>C: PASS / LOW_CONFID / REJECT V-->>C: PASS / LOW_CONFID / REJECT
C->>S: 写入 verifier_evaluation C->>R: 写入 verifier_evaluation
alt LOW_CONFID 且允许补证据 alt LOW_CONFID 且允许补证据
C->>P: retry_context: 仅补缺失证据 C->>P: retry_context: 仅补缺失证据
else PASS 或 REJECT else PASS 或 REJECT
C-->>S: 保存最终 answer C->>M: allowed_claims + missing_info + recommended_actions
M-->>C: composer_output
C->>R: 保存 Composer 最终 answer
end end
``` ```
@@ -113,7 +132,7 @@ sequenceDiagram
| Verdict | 行为 | | Verdict | 行为 |
|---|---| |---|---|
| `PASS` | 输出 Executor 答案 | | `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 | | `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 | | `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
@@ -177,21 +196,25 @@ flowchart LR
SkillBody --> Executor SkillBody --> Executor
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"] Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
EvidenceTools --> ToolTrace["tool_invocation evidence"] EvidenceTools --> ToolTrace["tool_invocation evidence"]
Executor --> Verifier["Verifier"] Executor --> Gatekeeper["Gatekeeper"]
Gatekeeper --> Verifier["Verifier"]
ToolTrace --> Verifier ToolTrace --> Verifier
Verifier --> Composer["Composer"]
``` ```
| 角色 | Skill 可见性 | 工具权限 | | 角色 | Skill 可见性 | 工具权限 |
|---|---|---| |---|---|---|
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` | | Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 | | Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 `tool_trace_summary` | | Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
## 7. 与旧版设计的差异 ## 7. 与旧版设计的差异
| 旧版设想 | 当前实现 | | 旧版设想 | 当前实现 |
|---|---| |---|---|
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Verifier;AIOps: Supervisor + Planner + Executor | | Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 | | ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 | | 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT | | Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
@@ -154,4 +154,4 @@ ALTER TABLE diagnosis_session ADD COLUMN answer LONGTEXT COMMENT 'Agent 返回
- **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑 - **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑
- **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL - **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../issues/ISS-001-duplicate-retrieval.md)) - **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../../../issues/archived/ISS-001-duplicate-retrieval.md))
+63 -28
View File
@@ -1,7 +1,7 @@
# 当前 MVP 架构 # 当前 MVP 架构
**更新日期**:2026-07-05 **更新日期**:2026-07-08
**状态**:当前可运行架构 **状态**:当前可运行架构
**适用范围**:Demo、面试讲解、后续迭代规划 **适用范围**:Demo、面试讲解、后续迭代规划
## 1. 系统定位 ## 1. 系统定位
@@ -38,7 +38,9 @@ flowchart TB
Supervisor["Supervisor"] Supervisor["Supervisor"]
Planner["Planner"] Planner["Planner"]
Executor["Executor"] Executor["Executor"]
Gatekeeper["Gatekeeper"]
Verifier["Verifier"] Verifier["Verifier"]
Composer["Composer"]
end end
subgraph Tools["Evidence Tools"] subgraph Tools["Evidence Tools"]
@@ -63,7 +65,8 @@ flowchart TB
end end
subgraph Store["Persistence and Trace"] subgraph Store["Persistence and Trace"]
Session["diagnosis_session"] ChatSession["chat_session"]
Run["diagnosis_run"]
Step["agent_step"] Step["agent_step"]
Invocation["tool_invocation"] Invocation["tool_invocation"]
ApiDoc["api_document"] ApiDoc["api_document"]
@@ -105,7 +108,9 @@ Agent Orchestration
-> Supervisor -> Supervisor
-> Planner -> Planner
-> Executor -> Executor
-> Gatekeeper
-> Verifier -> Verifier
-> Composer
Evidence Tools Evidence Tools
-> lookup_knowledge -> lookup_knowledge
@@ -126,13 +131,15 @@ RAG Retrieval
-> Milvus SDK fallback -> Milvus SDK fallback
Persistence Persistence
-> diagnosis_session -> chat_session
-> agent_step -> diagnosis_run
-> tool_invocation -> agent_step.run_id
-> tool_invocation.run_id
-> api_document -> api_document
-> Milvus/Zilliz collection -> Milvus/Zilliz collection
Quality Gates Quality Gates
-> executor gatekeeper
-> chat verifier -> chat verifier
-> AIOps rule evaluation -> AIOps rule evaluation
-> diagnosis eval baseline -> diagnosis eval baseline
@@ -150,23 +157,30 @@ sequenceDiagram
participant Planner as Planner Agent participant Planner as Planner Agent
participant Executor as Executor Agent participant Executor as Executor Agent
participant Tool as Evidence Tools participant Tool as Evidence Tools
participant Gatekeeper as Gatekeeper Hook
participant Verifier as Verifier Agent participant Verifier as Verifier Agent
participant Composer as Composer Agent
participant DB as Trace Tables participant DB as Trace Tables
participant Trace as Trace API participant Trace as Trace API
User->>API: 提交诊断问题 User->>API: 提交诊断问题
API->>Chat: execute chat strategy API->>Chat: execute chat strategy
Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
Chat->>Planner: 复杂问题进入规划 Chat->>Planner: 复杂问题进入规划
Planner->>DB: 写入 agent_step Planner->>DB: 写入 agent_step.run_id
Planner->>Executor: 下发排查方向 Planner->>Executor: 下发排查方向
Executor->>Tool: lookup_knowledge / logs / metrics Executor->>Tool: lookup_knowledge / logs / metrics
Tool->>DB: 写入 tool_invocation Tool->>DB: 写入 tool_invocation.run_id
Tool-->>Executor: 返回证据 Tool-->>Executor: 返回证据
Executor->>Verifier: 生成候选诊断并校验 Executor->>Gatekeeper: 输出 executor_evidence_v2
Verifier->>DB: 合并 self_evaluation.verifier_evaluation Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
Chat->>DB: 保存 diagnosis_session.answer Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
User->>Trace: GET /api/diagnosis/{sessionId}/trace Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
Trace->>DB: 聚合 session / step / tool Verifier->>Composer: 传入 allowed_claims / missing_info / actions
Composer->>Chat: 生成最终用户答复
Chat->>DB: 保存 diagnosis_run.answer
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
Trace->>DB: 聚合 run / step / tool
Trace-->>User: 返回可回放诊断链路 Trace-->>User: 返回可回放诊断链路
``` ```
@@ -180,14 +194,17 @@ POST /api/chat
-> lookup_knowledge -> lookup_knowledge
-> query_logs -> query_logs
-> query_metrics -> query_metrics
-> Verifier 校验最终诊断 -> Gatekeeper 校验 Executor 证据引用真实性
-> 保存 diagnosis_session -> Verifier 判断 claim 是否能由已核验证据推出
-> 保存 agent_step -> Composer 生成最终用户答复
-> 保存 tool_invocation -> 保存 chat_session metadata
-> 合并 self_evaluation.verifier_evaluation -> 保存 diagnosis_run
-> 保存 agent_step.run_id
-> 保存 tool_invocation.run_id
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
``` ```
Chat 链路的质量门禁是 LLM Verifier。Verifier 输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。 Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。 Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
@@ -242,7 +259,7 @@ POST /api/ai_ops
-> Prometheus / logs / knowledge tools -> Prometheus / logs / knowledge tools
-> 生成告警分析报告 -> 生成告警分析报告
-> AiOpsRuleEvaluationService -> AiOpsRuleEvaluationService
-> 合并 self_evaluation.aiops_rule_evaluation -> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
-> Trace API 可查看全链路 -> Trace API 可查看全链路
``` ```
@@ -304,32 +321,41 @@ RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细
## 6. 持久化模型 ## 6. 持久化模型
当前诊断持久化以三张表为核心: 当前诊断持久化以 session/run/trace 明细为核心:
```text ```text
diagnosis_session chat_session
-> 一次诊断会话的主记录 -> 多轮会话目录和元数据
-> session_id / status / message_pair_count
diagnosis_run
-> 一次诊断运行的主记录
-> run_id / session_id
-> query / status / agent_flow / answer -> query / status / agent_flow / answer
-> self_evaluation -> self_evaluation
-> step_count / tool_call_count / duration -> step_count / tool_call_count / duration
agent_step agent_step
-> Agent 模型调用步骤 -> Agent 模型调用步骤
-> session_id / run_id
-> step_index / agent_name -> step_index / agent_name
-> model_input / model_output / thought -> model_input / model_output / thought
-> duration / token_count -> duration / token_count
tool_invocation tool_invocation
-> 工具调用事实 -> 工具调用事实
-> session_id / run_id
-> tool_name / input_params / output_preview -> tool_name / input_params / output_preview
-> retrieval_layer / retrieval_details -> retrieval_layer / retrieval_details
-> retrieval_details.evidence_refs
-> relevance_level / dedup_reason -> relevance_level / dedup_reason
-> duration / success -> duration / success
``` ```
说明: 说明:
- 旧的 `diagnosis_record` 已不是当前主模型,迁移脚本中已经由 `diagnosis_session + agent_step + tool_invocation` 取代。 - 旧的 `diagnosis_record` 已不是当前主模型。
- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
- `api_document` 仍用于文档元数据管理。 - `api_document` 仍用于文档元数据管理。
- 文档向量内容存放在 Milvus/Zilliz collection 中。 - 文档向量内容存放在 Milvus/Zilliz collection 中。
@@ -339,19 +365,20 @@ tool_invocation
```text ```text
GET /api/diagnosis/{sessionId}/trace GET /api/diagnosis/{sessionId}/trace
GET /api/diagnosis/{sessionId}/trace?runId=run-...
``` ```
Trace API 聚合: Trace API 聚合:
- 会话状态和最终报告。 - 会话元数据、运行状态和最终报告。
- Agent step 序列。 - Agent step 序列。
- 工具调用和检索细节。 - 工具调用和检索细节。
- Chat verifier 结果。 - Chat Gatekeeper / Verifier / Composer 结果。
- AIOps rule evaluation 结果。 - AIOps rule evaluation 结果。
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。 Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。 Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
## 8. 质量门禁 ## 8. 质量门禁
@@ -359,7 +386,9 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
| 门禁 | 位置 | 作用 | | 门禁 | 位置 | 作用 |
|---|---|---| |---|---|---|
| Chat Verifier | `ChatService` | 校验普通诊断回答质量 | | Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 | | AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 | | Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 | | RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
@@ -379,6 +408,11 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
- `title`、`breadcrumb`、`content` 参与 embedding 文本。 - `title`、`breadcrumb`、`content` 参与 embedding 文本。
- `tool_invocation` 记录检索层、relevance level、dedup reason。 - `tool_invocation` 记录检索层、relevance level、dedup reason。
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。 - Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
- Verifier 只判断可推导性。
- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
- RAG offline baseline 和 live acceptance 脚本。 - RAG offline baseline 和 live acceptance 脚本。
暂不作为当前已完成能力声明: 暂不作为当前已完成能力声明:
@@ -406,4 +440,5 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` | | Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
| Trace 聚合 | `DiagnosisTraceService` | | Trace 聚合 | `DiagnosisTraceService` |
| 工具调用记录 | `ToolInvocationRecorder` | | 工具调用记录 | `ToolInvocationRecorder` |
| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
| self_evaluation 合并 | `SelfEvaluationMergeService` | | self_evaluation 合并 | `SelfEvaluationMergeService` |
+81 -151
View File
@@ -1,31 +1,44 @@
# 数据模型总览 # 数据模型总览
**更新日期**:2026-07-05 **更新日期**:2026-07-10
**状态**:当前可运行架构 **状态**:当前可运行架构
## 1. 定位 ## 1. 定位
本文从架构角度说明当前 MVP 的核心数据模型。详细字段仍以 Flyway migration 和 `mvp/tables/` 为准。 本文从架构角度说明当前 MVP 的核心数据模型。详细字段以 Flyway migration、实体类和 `mvp/tables/` 为准。
核心数据分三组: 核心数据分三组:
- 诊断 Trace:`diagnosis_session`、`agent_step`、`tool_invocation` - 会话与诊断 Trace:`chat_session`、`diagnosis_run`、`agent_step`、`tool_invocation`
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata - 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
- 反馈沉淀:`case_library` - 反馈沉淀:`case_library`
`diagnosis_session` 仍保留为历史兼容和回滚表,不再是新执行写入的主模型。
## 2. 总体关系 ## 2. 总体关系
```mermaid ```mermaid
erDiagram erDiagram
diagnosis_session ||--o{ agent_step : has chat_session ||--o{ diagnosis_run : owns
diagnosis_session ||--o{ tool_invocation : has diagnosis_run ||--o{ agent_step : has
diagnosis_session ||--o| case_library : creates_when_useful diagnosis_run ||--o{ tool_invocation : has
diagnosis_run ||--o| case_library : creates_when_useful
api_document ||--o{ milvus_chunk : indexed_as api_document ||--o{ milvus_chunk : indexed_as
knowledge_domain ||--o{ api_document : groups knowledge_domain ||--o{ api_document : groups
diagnosis_session { chat_session {
bigint id bigint id
varchar session_id varchar session_id
varchar status
int message_pair_count
datetime last_active_at
datetime expires_at
}
diagnosis_run {
bigint id
varchar run_id
varchar session_id
text query text query
varchar status varchar status
varchar agent_flow varchar agent_flow
@@ -37,42 +50,23 @@ erDiagram
agent_step { agent_step {
bigint id bigint id
varchar session_id varchar session_id
varchar run_id
int step_index int step_index
varchar agent_name varchar agent_name
text model_input text model_input
text model_output text model_output
text thought
boolean has_tool_call boolean has_tool_call
} }
tool_invocation { tool_invocation {
bigint id bigint id
varchar session_id varchar session_id
varchar run_id
bigint step_id
varchar tool_name varchar tool_name
json input_params json input_params
text output_preview text output_preview
varchar retrieval_layer
json retrieval_details json retrieval_details
varchar relevance_level
varchar dedup_reason
}
api_document {
bigint id
varchar doc_id
varchar file_name
varchar file_path
varchar status
int chunk_count
text metadata
}
knowledge_domain {
bigint id
varchar domain_id
varchar description
text when_to_retrieve
int document_count
} }
case_library { case_library {
@@ -84,123 +78,87 @@ erDiagram
text root_cause text root_cause
text solution text solution
} }
milvus_chunk {
varchar id
text content
json metadata
vector vector
}
``` ```
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。 说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
## 3. 诊断 Trace 模型 ## 3. 会话与运行模型
### diagnosis_session ### chat_session
会话级主记录。 `chat_session` 是会话目录表,保存 `sessionId` 的元数据:
关键字段:
| 字段 | 说明 | | 字段 | 说明 |
|---|---| |---|---|
| `session_id` | 外部关联键,Trace 和 Feedback 都使用它 | | `session_id` | 外部会话 ID,用于多轮上下文和 run 列表 |
| `query` | 用户原始问题或 AIOps 输入摘要 | | `status` | 会话目录状态 |
| `status` | 执行状态 | | `message_pair_count` | Redis 对话轮次数快照 |
| `last_active_at` | 最近活跃时间 |
| `expires_at` | 可为空的目录 TTL 元数据 |
它不保存完整对话历史,正文消息仍由 Redis `SessionContext.messageHistory` 管理。
### diagnosis_run
`diagnosis_run` 是一次可回放诊断执行的主记录:
| 字段 | 说明 |
|---|---|
| `run_id` | 运行 ID,格式为 `run-` + UUID |
| `session_id` | 所属 `chat_session.session_id` |
| `query` | 本次 Chat 问题或 AIOps 告警摘要 |
| `status` | 本次执行状态 |
| `agent_flow` | `CHAT` / `AI_OPS` | | `agent_flow` | `CHAT` / `AI_OPS` |
| `answer` | 最终答复或告警报告 | | `answer` | 本次运行最终答复或告警报告 |
| `self_evaluation` | rule/verifier/aiops 自评估容器 | | `self_evaluation` | 本次运行的 rule/verifier/aiops 自评估容器 |
| `feedback` | 用户反馈 | | `feedback` | 本次运行的用户反馈 |
同一个 `sessionId` 可以有多个 `runId`。Trace、反馈、评测和案例沉淀都应优先使用 `runId`,避免多轮同 session 下的数据混合。
## 4. Trace 明细模型
### agent_step ### agent_step
记录模型调用步骤。 `agent_step` 记录模型调用步骤。新写入同时保留 `session_id` 和 `run_id`,其中 `run_id` 是回放边界。Trace 页面和评测应先按 `run_id` 隔离取数,展示顺序以 Trace API 返回顺序为准。
用途:
- 回放 Agent 推理过程。
- 查看 Planner / Executor / Verifier 的输入输出摘要。
- 统计 step count、duration、token count。
### tool_invocation ### tool_invocation
记录工具调用事实。 `tool_invocation` 记录显式工具调用事实。`retrieval_details.evidence_refs` 是 Chat 证据链路的关键字段:
用途: ```json
{
- 给 Trace API 展示证据。 "evidence_status": "supported",
- 给 Verifier 构造 `tool_trace_summary`。 "evidence_refs": [
- 给 `EvaluationService` 计算 evidence score。 {
- 给 RAG eval 和人工排查提供检索细节。 "raw_path": "$.logs[0]",
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
## 4. 知识库模型 }
]
### api_document }
MySQL 中的文档元数据表。
职责:
- 管理上传文件。
- 保存 file hash,用于去重。
- 记录索引状态和 chunk 数量。
- 保存 frontmatter JSON。
### knowledge_domain
领域级元数据。
职责:
- 按 category 聚合文档。
- 存储领域描述。
- 存储 `when_to_retrieve`,辅助 Planner/Executor 判断什么时候检索该领域。
### Milvus/Zilliz metadata
向量 collection 中每个 chunk 的 metadata 主要包括:
```text
docId
_source
chunkIndex
totalChunks
title
breadcrumb
category
``` ```
这些字段支撑: `$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
- category filter。
- source 展示。
- breadcrumb 上下文。
- docId 删除和重建索引。
- evidence block 构造。
## 5. 反馈沉淀模型 ## 5. 反馈沉淀模型
### case_library `useful` 反馈会触发 `CaseLibraryService.createFromRun`。
`useful` 反馈会触发 `CaseLibraryService.createFromSession`。
当前自动映射: 当前自动映射:
| 字段 | 来源 | | 字段 | 来源 |
|---|---| |---|---|
| `case_id` | UUID | | `case_id` | UUID |
| `diagnosis_id` | `diagnosis_session.session_id` | | `diagnosis_id` | 新数据为 `diagnosis_run.run_id`;历史数据可能为 `diagnosis_session.session_id` |
| `source_type` | `AUTO` | | `source_type` | `AUTO` |
| `fault_category` | 当前默认 `GENERAL` | | `fault_category` | 当前默认 `GENERAL` |
| `title` | session query 前 100 字符 | | `title` | run query 前 100 字符 |
| `root_cause` | session answer | | `root_cause` | run answer |
| `solution` | session answer | | `solution` | run answer |
| `created_by` | `system` | | `created_by` | `system` |
## 6. self_evaluation 结构 ## 6. self_evaluation 结构
`diagnosis_session.self_evaluation` 是 JSON 容器: `diagnosis_run.self_evaluation` 是运行级 JSON 容器:
```json ```json
{ {
@@ -210,49 +168,21 @@ category
} }
``` ```
边界: Chat 通常写入 `rule_evaluation` 和 `verifier_evaluation`;AIOps 写入 `aiops_rule_evaluation`。
- `rule_evaluation` 评估证据收集充分度。 ## 7. 当前边界和后续
- `verifier_evaluation` 评估 Chat 答案关键事实是否有证据支撑。
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
## 7. 数据写入时序
```mermaid
sequenceDiagram
autonumber
participant API as API
participant Svc as ChatService/AiOpsService
participant Session as diagnosis_session
participant Agent as Agent
participant Step as agent_step
participant Tool as tool_invocation
participant Eval as self_evaluation
participant Feedback as case_library
API->>Svc: request
Svc->>Session: create/update RUNNING
Agent->>Step: before/after model
Agent->>Tool: tool call record
Svc->>Session: SUCCESS/FAILED + answer
Svc->>Eval: merge evaluation
API->>Svc: feedback useful
Svc->>Feedback: create case
```
## 8. 当前边界和后续
当前边界: 当前边界:
- `agent_step.session_id` 和 `tool_invocation.session_id` 通过 sessionId 关联,不强制外键。 - `chat_session` 只存会话元数据,不存完整正文历史。
- `tool_invocation.step_id` 可为空。 - `diagnosis_run` 存一次运行的长期审计状态。
- Milvus chunk 与 `api_document` 通过 metadata.docId 逻辑关联。 - `agent_step.run_id` 和 `tool_invocation.run_id` 是 Trace、Verifier、Eval 的运行边界。
- `case_library` 与 session 通过 `diagnosis_id=session_id` 关联。 - 当前实现主要使用逻辑关联,不依赖数据库外键。
- `case_library.diagnosis_id` 是过渡字段,新值按 `run_id` 解释,旧值可能按 `session_id` 解释。
- `diagnosis_session` 只作为历史兼容和回滚表保留。
后续可增强: 后续可增强:
1. 增加 run id,支持同 session 多次独立诊断。 1. 强化 `tool_invocation.step_id` 关联。
2. 强化 `tool_invocation.step_id` 关联。 2. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。
3. 将 evidence block 结构化保存。 3. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
@@ -0,0 +1,443 @@
# Chat Evidence Pipeline Contracts
**状态**:当前实现
**更新日期**:2026-07-08
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
当前 Chat 复杂诊断链路是:
```text
chat_planner
-> chat_executor
-> VerifierInputHook / ExecutorGatekeeperService
-> chat_verifier
-> chat_composer
-> final answer
```
设计原则:
- Planner 暂不输出 `scope_contract`。
- Executor 只做证据收集和微观事实提炼,不生成最终用户答案。
- Gatekeeper 在 Verifier 前做代码级引用真实性校验。
- Verifier 判断 claim 是否能由已核验证据推出。
- Composer 只表达 Verifier 允许输出的内容。
---
## 1. Planner
Planner 当前保持不变,输出 `planner_plan`:
```json
{
"selected_skill": "diagnose-mysql-connection-pool",
"selection_reason": "选择该 skill 的原因",
"plan": ["步骤1", "步骤2"],
"reasoning": "规划思路"
}
```
字段定义:
| 字段 | 类型 | 定义 |
|---|---|---|
| `selected_skill` | string/null | Planner 选择的诊断 skill 名称 |
| `selection_reason` | string | skill 选择理由 |
| `plan` | array | 给 Executor 的执行步骤 |
| `reasoning` | string | 规划思路说明 |
当前边界:
- 不新增 `scope_contract`。
- 不要求 Planner 显式列出 forbidden actions。
- 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。
---
## 2. Executor
Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。
### 2.1 输出结构
```json
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "observation",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "",
"tool_name": "query_metrics",
"source_invocation_id": 517,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
```
字段定义:
| 字段 | 类型 | 必填 | 定义 |
|---|---|---:|---|
| `answer_version` | string | 是 | 固定为 `executor_evidence_v2` |
| `claims` | array | 是 | Executor 提出的待验证事实断言 |
| `claims[].claim_id` | string | 是 | claim 标识 |
| `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 |
| `claims[].claim_text` | string | 是 | 事实断言文本 |
| `claims[].support_level` | string | 是 | `direct` 或 `indirect` |
| `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 |
| `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` |
| `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 |
| `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 |
| `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 |
| `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 |
| `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 |
| `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact |
| `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 |
| `missing_info` | array | 是 | 无法确认结论所缺少的证据 |
禁止字段:
- `diagnosis_summary`
- `user_facing_answer`
- `source_invocation_ids` 作为主引用字段
### 2.2 raw_path
当前支持的精确路径:
| 工具 | 正向证据路径 | 负向证据路径 |
|---|---|---|
| `query_metrics` | `$.alerts[i]` | `$.no_evidence` |
| `query_logs` | `$.logs[i]` | `$.no_evidence` |
| `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` |
约束:
- `raw_path` 必须指向数组条目或 `$.no_evidence`。
- 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。
- 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。
### 2.3 negative_observation
当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`:
```json
{
"claim_id": "claim-1",
"claim_type": "negative_observation",
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"support_level": "direct",
"evidence_bindings": [
{
"tool_name": "query_logs",
"source_invocation_id": 517,
"raw_path": "$.no_evidence",
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
}
]
}
```
语义边界:
- `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。
- 不表示“问题绝对不存在”。
- 不表示“根因被排除”。
- 不表示“系统已经健康”。
- `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。
### 2.4 窄范围任务
窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。
Executor 必须遵守:
- 只输出 `observation` / `negative_observation`。
- claim 数量通常 1 条,最多 2 条。
- claim 数量限制不限制 `evidence_bindings` 数量。
- 不输出根因、风险、修复建议、经验推断。
- 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。
- 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。
---
## 3. Tool Invocation Evidence Refs
工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。
### 3.1 正向证据
```json
{
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.logs[0]",
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
}
]
}
```
### 3.2 负向证据
```json
{
"evidence_status": "no_evidence",
"evidence_refs": [
{
"raw_path": "$.no_evidence",
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
}
]
}
```
字段定义:
| 字段 | 类型 | 定义 |
|---|---|---|
| `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` |
| `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 |
| `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 |
---
## 4. Gatekeeper
Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。
### 4.1 输入
- `sessionId`
- `executor_structured_output`
- 当前 session 的 `tool_invocation`
### 4.2 输出
```json
{
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_logs",
"source_invocation_id": 517,
"raw_path": "$.no_evidence",
"matched_text": "query_logs returned no evidence; ...",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
}
```
字段定义:
| 字段 | 类型 | 定义 |
|---|---|---|
| `status` | string | `pass` 或 `fail` |
| `severity` | string | `none`、`low_confid`、`reject` |
| `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 |
| `rules` | array | 已启用规则的轻量元数据摘要 |
| `checked_bindings` | array | 每条证据绑定的校验结果 |
| `failed_rules` | array | 失败规则 id |
| `warnings` | array | 自动回填等非阻断信息 |
| `errors` | array | 失败明细 |
校验规则:
- `answer_version` 必须是 `executor_evidence_v2`。
- 不允许 `diagnosis_summary` / `user_facing_answer`。
- 每个 claim 必须有非空 `evidence_bindings`。
- `tool_name` 必须和真实 invocation 对齐。
- `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。
- `raw_path` 必须存在于 `retrieval_details.evidence_refs`。
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
- `negative_observation` 只能绑定 `$.no_evidence`。
规则配置:
- 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。
- 规则实现仍是确定性 Java 代码,不执行动态脚本。
- 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。
失败分级:
| 场景 | severity |
|---|---|
| 伪造 invocation id | `reject` |
| tool_name 与 invocation 不匹配 | `reject` |
| raw_path 不存在 | `reject` |
| excerpt 与 matched_text 不匹配 | `reject` |
| negative_observation 绑定正向日志 | `reject` |
| 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` |
| 旧 invocation 没有 `evidence_refs` | `low_confid` |
---
## 5. Verifier
Verifier 输入由 `VerifierInputHook` 构造:
```json
{
"original_query": "用户原始问题",
"executor_final_answer": "{...executor raw text for debug/fallback only...}",
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": []
},
"executor_output_parse_status": {
"status": "valid",
"detail": "parsed executor evidence contract"
},
"tool_trace_summary": [],
"gatekeeper_result": {},
"retry_context": null
}
```
Verifier 职责:
- 不调用工具。
- 不读 skill。
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
- 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
输出:
```json
{
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"claim_checks": [],
"facts_checked": [],
"rationale": "..."
}
```
---
## 6. Composer
Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。
输入概念:
| 字段 | 定义 |
|---|---|
| `original_query` | 用户原始问题 |
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
| `allowed_claims` | Verifier 允许表达的 claims |
| `allowed_hypotheses` | Verifier 允许表达的假设 |
| `missing_info` | 证据缺口 |
| `recommended_actions` | 允许表达的建议动作 |
| `rationale` | Verifier 判定理由 |
输出:
```json
{
"answer_summary": "一句话概括",
"recommended_actions": [
{
"action_text": "下一步动作",
"reason": "原因"
}
],
"user_facing_answer": "最终给用户看的中文答案"
}
```
表达边界:
- Composer 不补事实、不补根因、不调用工具。
- 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。
- 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
- 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
---
## 7. Trace Persistence
`diagnosis_run.self_evaluation.verifier_evaluation` 持久化:
```json
{
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"claim_checks": [],
"facts_checked": [],
"rationale": "...",
"round": 1,
"traceability_version": "v1",
"executor_output_parse_status": {},
"executor_structured_output": {},
"gatekeeper_result": {
"rule_set_version": "gatekeeper-rules-v1"
},
"composer_output": {},
"tool_trace_summary": []
}
}
```
Trace API 可用于回放:
- Executor 输出了哪些 claim。
- 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。
- Gatekeeper 是否通过、是否自动回填。
- Verifier 如何判断可推导性。
- Composer 最终如何表达给用户。
---
## 8. 当前已验证样例
| 场景 | sessionId | 结果 |
|---|---|---|
| HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim |
| HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 |
---
## 9. 仍需记录或后续补强
当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。
3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
+67 -27
View File
@@ -1,15 +1,15 @@
# 反馈与自评估架构 # 反馈与自评估架构
**更新日期**:2026-07-05 **更新日期**:2026-07-10
**状态**:当前可运行架构 **状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md` **参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
## 1. 定位 ## 1. 定位
反馈架构包含两条闭环: 反馈架构包含两条闭环:
1. 系统自评估:基于工具调用、Verifier、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。 1. 系统自评估:基于当前 run 的工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_run.self_evaluation`。
2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。 2. 用户反馈:用户标记 `useful` 或 `not_useful`,优先写入 `diagnosis_run.feedback`,其中 `useful` 会沉淀案例。
当前重要边界: 当前重要边界:
@@ -21,13 +21,18 @@
```mermaid ```mermaid
flowchart TD flowchart TD
Answer["Chat / AIOps final answer"] --> Session["diagnosis_session.answer"] Answer["Chat / AIOps final answer"] --> Run["diagnosis_run.answer"]
subgraph SelfEval["Self evaluation"] subgraph SelfEval["Self evaluation"]
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"] Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
Invocation --> EvidenceRefs["evidence_refs"]
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
Invocation --> TraceSummary["ToolTraceSummaryService"] Invocation --> TraceSummary["ToolTraceSummaryService"]
TraceSummary --> Verifier["chat_verifier"] Gatekeeper --> Verifier["chat_verifier"]
TraceSummary --> Verifier
Verifier --> VerifierEval["verifier_evaluation"] Verifier --> VerifierEval["verifier_evaluation"]
Verifier --> Composer["chat_composer"]
Composer --> VerifierEval
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"] Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["aiops_rule_evaluation"] AiOpsRule --> AiOpsEval["aiops_rule_evaluation"]
end end
@@ -35,24 +40,24 @@ flowchart TD
RuleEval --> Merge["SelfEvaluationMergeService"] RuleEval --> Merge["SelfEvaluationMergeService"]
VerifierEval --> Merge VerifierEval --> Merge
AiOpsEval --> Merge AiOpsEval --> Merge
Merge --> SelfJson["diagnosis_session.self_evaluation"] Merge --> SelfJson["diagnosis_run.self_evaluation"]
subgraph UserFeedback["User feedback"] subgraph UserFeedback["User feedback"]
UI["Feedback bar"] --> API["POST /api/feedback"] UI["Feedback bar"] --> API["POST /api/feedback"]
API --> FeedbackService["FeedbackService"] API --> FeedbackService["FeedbackService"]
FeedbackService --> FeedbackField["diagnosis_session.feedback"] FeedbackService --> FeedbackField["diagnosis_run.feedback"]
FeedbackService --> Useful{"feedback == useful?"} FeedbackService --> Useful{"feedback == useful?"}
Useful -->|yes| CaseService["CaseLibraryService.createFromSession"] Useful -->|yes| CaseService["CaseLibraryService.createFromRun"]
CaseService --> Case["case_library"] CaseService --> Case["case_library"]
Useful -->|no| BadCase["Bad case by feedback=not_useful"] Useful -->|no| BadCase["Bad case by feedback=not_useful"]
end end
Session --> UI Run --> UI
``` ```
## 3. self_evaluation JSON ## 3. self_evaluation JSON
`SelfEvaluationMergeService` 统一维护 `diagnosis_session.self_evaluation`。 `SelfEvaluationMergeService` 统一维护当前运行的 `diagnosis_run.self_evaluation`。历史兼容数据可能仍存在于 `diagnosis_session.self_evaluation`,但新 Chat/AIOps 执行不再写旧表。
当前结构: 当前结构:
@@ -67,10 +72,15 @@ flowchart TD
"verdict": "PASS", "verdict": "PASS",
"groundedness_score": 0.8, "groundedness_score": 0.8,
"critical_fact_count": 2, "critical_fact_count": 2,
"claim_checks": [],
"facts_checked": [], "facts_checked": [],
"rationale": "...", "rationale": "...",
"round": 1, "round": 1,
"traceability_version": "v1", "traceability_version": "v1",
"executor_output_parse_status": {},
"executor_structured_output": {},
"gatekeeper_result": {},
"composer_output": {},
"tool_trace_summary": [] "tool_trace_summary": []
}, },
"aiops_rule_evaluation": { "aiops_rule_evaluation": {
@@ -117,17 +127,29 @@ flowchart TD
## 5. Chat Verifier 自评估 ## 5. Chat Verifier 自评估
Chat Verifier 校验 Executor 的最终答案是否被证据支撑。 Chat 自评估分三步:
1. Gatekeeper 用代码校验 Executor 的引用是否真实。
2. Verifier 判断已验真的 `evidence_excerpt` 是否能推出 `claim_text`。
3. Composer 只把 Verifier 允许表达的内容写成最终用户答复。
```mermaid ```mermaid
flowchart LR flowchart LR
Answer["executor_final_answer"] --> Verifier["chat_verifier"] ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"] Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
EvidenceRefs --> Gatekeeper
Gatekeeper --> GateResult["gatekeeper_result"]
Summary --> Evidence["tool_trace_summary"] Summary --> Evidence["tool_trace_summary"]
GateResult --> Verifier["chat_verifier"]
ExecutorOutput --> Verifier
Evidence --> Verifier Evidence --> Verifier
Verifier --> Output["verifier_output JSON"] Verifier --> Output["verifier_output JSON"]
Output --> Composer["chat_composer"]
Composer --> ComposerOutput["composer_output"]
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"] Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"] ComposerOutput --> Merge
Merge --> Run["diagnosis_run.self_evaluation.verifier_evaluation"]
``` ```
Verifier 输出: Verifier 输出:
@@ -137,16 +159,25 @@ Verifier 输出:
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` | | `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
| `groundedness_score` | 关键事实证据支撑度 | | `groundedness_score` | 关键事实证据支撑度 |
| `critical_fact_count` | 关键事实数量 | | `critical_fact_count` | 关键事实数量 |
| `claim_checks` | 对 Executor 结构化 claims 的逐条可推导性判断 |
| `facts_checked` | 逐条事实校验 | | `facts_checked` | 逐条事实校验 |
| `rationale` | 判定原因 | | `rationale` | 判定原因 |
| `tool_trace_summary` | 本次校验使用的证据索引 | | `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 |
| `gatekeeper_result` | 引用真实性校验结果 |
| `composer_output` | 最终表达的解析状态和摘要 |
| `tool_trace_summary` | 本次校验使用的工具调用导航索引 |
ChatService 根据 verdict 决定: ChatService 根据 verdict 决定:
- `PASS`:输出 Executor 答案。 - `PASS`:把允许表达的 claims 交给 Composer 输出。
- `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。 - `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。
- `REJECT`:降级输出,只保留已确认信息。 - `REJECT`:降级输出,只保留已确认信息。
边界:
- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。
- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。
## 6. AIOps 规则自评估 ## 6. AIOps 规则自评估
AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。 AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。
@@ -167,6 +198,7 @@ POST /api/feedback
Content-Type: application/json Content-Type: application/json
{ {
"runId": "run-xxx",
"sessionId": "xxx", "sessionId": "xxx",
"feedback": "useful" | "not_useful" "feedback": "useful" | "not_useful"
} }
@@ -178,6 +210,8 @@ Content-Type: application/json
{ {
"success": true, "success": true,
"message": "反馈已记录", "message": "反馈已记录",
"runId": "run-xxx",
"fallbackToLatestRun": false,
"caseId": "uuid 或 null" "caseId": "uuid 或 null"
} }
``` ```
@@ -186,10 +220,16 @@ Content-Type: application/json
| feedback | 行为 | | feedback | 行为 |
|---|---| |---|---|
| `useful` | 写入 `DiagnosisSession.feedback`,调用 `CaseLibraryService.createFromSession` | | `useful` | 写入 `DiagnosisRun.feedback`,调用 `CaseLibraryService.createFromRun` |
| `not_useful` | 写入 `DiagnosisSession.feedback`,不改变 session status | | `not_useful` | 写入 `DiagnosisRun.feedback`,不改变 run status |
| 其他值 | 返回 HTTP 400 | | 其他值 | 返回 HTTP 400 |
兼容行为:
- 请求带 `runId` 时,后端验证 `runId` 属于 `sessionId`。
- 请求缺少 `runId` 且存在 run-backed 数据时,后端绑定 latest run,并返回 `fallbackToLatestRun=true` 和实际 `runId`。
- 仅当没有 `diagnosis_run` 但存在历史 `diagnosis_session` 时,才使用历史 fallback;该路径不声明 latest-run fallback。
## 8. 案例沉淀 ## 8. 案例沉淀
`useful` 反馈会生成或复用 `case_library` 记录。 `useful` 反馈会生成或复用 `case_library` 记录。
@@ -199,7 +239,7 @@ Content-Type: application/json
| CaseLibrary 字段 | 来源 | | CaseLibrary 字段 | 来源 |
|---|---| |---|---|
| `caseId` | UUID | | `caseId` | UUID |
| `diagnosisId` | `DiagnosisSession.sessionId` | | `diagnosisId` | 新数据为 `DiagnosisRun.runId`;历史数据可能为 `DiagnosisSession.sessionId` |
| `sourceType` | `AUTO` | | `sourceType` | `AUTO` |
| `faultCategory` | 当前固定为 `GENERAL` | | `faultCategory` | 当前固定为 `GENERAL` |
| `title` | `query` 前 100 字符 | | `title` | `query` 前 100 字符 |
@@ -210,7 +250,7 @@ Content-Type: application/json
幂等性: 幂等性:
```text ```text
case_library.diagnosisId == sessionId case_library.diagnosisId == runId
-> existing case: return existing -> existing case: return existing
-> missing case: create new -> missing case: create new
``` ```
@@ -229,23 +269,23 @@ Trace API 会展示:
| 视角 | 数据来源 | | 视角 | 数据来源 |
|---|---| |---|---|
| 执行是否成功 | `diagnosis_session.status` | | 执行是否成功 | `diagnosis_run.status` |
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` | | 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
| 用户是否认可 | `diagnosis_session.feedback` | | 用户是否认可 | `diagnosis_run.feedback` |
## 10. 后续增强 ## 10. 后续增强
近期优先: 近期优先:
1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。 1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。
2. `not_useful` 反馈沉淀 bad case,而不是只写字段。 2. 将 ISS-008 / ISS-009 这类 E2E 通过样例固化进 diagnosis eval fixtures。
3. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。 3. `not_useful` 反馈沉淀 bad case,而不是只写字段。
4. AIOps 引入 LLM Verifier。 4. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
5. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。 5. AIOps 引入 LLM Verifier。
6. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
暂不优先: 暂不优先:
- 用用户反馈直接修改 session status。 - 用用户反馈直接修改 session status。
- 仅凭 `evidence_score` 判断答案正确。 - 仅凭 `evidence_score` 判断答案正确。
- 在没有人工审核时自动把 bad case 反向写入 Prompt。 - 在没有人工审核时自动把 bad case 反向写入 Prompt。
+76 -23
View File
@@ -1,7 +1,7 @@
# Harness 与质量门禁架构 # Harness 与质量门禁架构
**更新日期**:2026-07-05 **更新日期**:2026-07-08
**状态**:当前可运行架构 + 后续门禁规划 **状态**:当前可运行架构 + 后续门禁规划
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md` **参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 设计目标 ## 1. 设计目标
@@ -21,6 +21,7 @@ Prompt contract
+ Tool boundary + Tool boundary
+ Agent hooks + Agent hooks
+ Trace persistence + Trace persistence
+ Gatekeeper deterministic validation
+ Verifier / rule evaluation + Verifier / rule evaluation
+ Eval baseline + Eval baseline
``` ```
@@ -30,21 +31,25 @@ Prompt contract
```mermaid ```mermaid
flowchart TB flowchart TB
Input["User / AIOps input"] --> Prompt["Prompt contract"] Input["User / AIOps input"] --> Prompt["Prompt contract"]
Prompt --> Agent["Planner / Executor / Verifier"] Prompt --> Agent["Planner / Executor / Verifier / Composer"]
Agent --> Tools["Evidence tools"] Agent --> Tools["Evidence tools"]
Tools --> Invocation["tool_invocation"] Tools --> Invocation["tool_invocation"]
Agent --> StepHook["AgentLoggingHook"] Agent --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step"] StepHook --> Step["agent_step"]
Agent --> Session["diagnosis_session"] Agent --> Run["diagnosis_run"]
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
Agent --> Gatekeeper
Invocation --> TraceSummary["ToolTraceSummaryService"] Invocation --> TraceSummary["ToolTraceSummaryService"]
TraceSummary --> Verifier["chat_verifier"] Gatekeeper --> Verifier["chat_verifier"]
TraceSummary --> Verifier
Verifier --> SelfEval["self_evaluation.verifier_evaluation"] Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"] Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"] AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
Session --> TraceAPI["DiagnosisTraceService"] Run --> TraceAPI["DiagnosisTraceService"]
Step --> TraceAPI Step --> TraceAPI
Invocation --> TraceAPI Invocation --> TraceAPI
SelfEval --> TraceAPI SelfEval --> TraceAPI
@@ -63,17 +68,37 @@ flowchart TB
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 | | `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 | | `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
| `chat-planner-prompt.md` | Chat 复杂问题规划 | | `chat-planner-prompt.md` | Chat 复杂问题规划 |
| `chat-executor-prompt.md` | Chat 执行工具并形成诊断答复 | | `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
| `chat-verifier-prompt.md` | 校验 Executor 答案是否被工具证据支撑 | | `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
Prompt 层当前承担的门禁: Prompt 层当前承担的门禁:
- 禁止凭记忆回答错误码、接口定义、排障步骤。 - 禁止凭记忆回答错误码、接口定义、排障步骤。
- 需要外部信息时必须调用工具。 - 需要外部信息时必须调用工具。
- 工具连续失败或返回空结果时,最终报告必须诚实说明。 - 工具连续失败或返回空结果时,最终报告必须诚实说明。
- Chat Verifier 不允许做新检索,只能校验已有证据。 - Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
- AIOps payload 模式必须聚焦输入告警。 - AIOps payload 模式必须聚焦输入告警。
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
```json
{
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_executor",
"version": "chat-executor-v2",
"resource": "prompts/chat-executor-prompt.md"
}
]
}
```
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
## 4. Trace Hooks ## 4. Trace Hooks
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。 `AgentLoggingHook` 是当前 Agent step 可观测性的核心。
@@ -85,10 +110,10 @@ sequenceDiagram
participant H as AgentLoggingHook participant H as AgentLoggingHook
participant DB as agent_step participant DB as agent_step
A->>H: before_model(messages, sessionId) A->>H: before_model(messages, sessionId, runId)
H->>DB: 写入 model_input / step_index / agent_name H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
A-->>A: LLM 推理 A-->>A: LLM 推理
A->>H: after_model(messages, sessionId) A->>H: after_model(messages, sessionId, runId)
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
``` ```
@@ -101,6 +126,8 @@ sequenceDiagram
- token count。 - token count。
- Verifier 的 JSON 输出摘要。 - Verifier 的 JSON 输出摘要。
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
## 5. Tool Invocation 门禁 ## 5. Tool Invocation 门禁
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。 工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
@@ -115,6 +142,7 @@ retrieval_layer
l0_match_count l0_match_count
l1_match_count l1_match_count
retrieval_details retrieval_details
-> evidence_refs
relevance_level relevance_level
dedup_reason dedup_reason
duration_ms duration_ms
@@ -128,23 +156,44 @@ error_message
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。 - 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。 - 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。 - dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
## 6. Verifier 门禁 ## 6. Gatekeeper 与 Verifier 门禁
Chat Verifier 的输入不是原始工具日志,而是 `ToolTraceSummaryService` 构造的证据索引。 Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
```mermaid ```mermaid
flowchart LR flowchart LR
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"] Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
EvidenceRefs --> Gatekeeper
Gatekeeper --> GateResult["gatekeeper_result"]
Invocation --> Summary["ToolTraceSummaryService"]
Summary --> EvidenceIndex["tool_trace_summary"] Summary --> EvidenceIndex["tool_trace_summary"]
EvidenceIndex --> Verifier["chat_verifier"] GateResult --> Verifier["chat_verifier"]
ExecutorAnswer["executor_final_answer"] --> Verifier ExecutorOutput --> Verifier
EvidenceIndex --> Verifier
Verifier --> Verdict{"verdict"} Verifier --> Verdict{"verdict"}
Verdict -->|PASS| Pass["输出原答案"] Verdict -->|PASS| Composer["chat_composer"]
Composer --> Pass["输出最终答复"]
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"] Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
Verdict -->|REJECT| Reject["降级输出"] Verdict -->|REJECT| Reject["降级输出"]
``` ```
Gatekeeper 检查:
| 检查 | 失败语义 |
|---|---|
| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
Verifier 输出: Verifier 输出:
```json ```json
@@ -152,17 +201,22 @@ Verifier 输出:
"verdict": "PASS|LOW_CONFID|REJECT", "verdict": "PASS|LOW_CONFID|REJECT",
"groundedness_score": 0.8, "groundedness_score": 0.8,
"critical_fact_count": 2, "critical_fact_count": 2,
"claim_checks": [],
"facts_checked": [], "facts_checked": [],
"rationale": "..." "rationale": "..."
} }
``` ```
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
结果写入: 结果写入:
```text ```text
diagnosis_session.self_evaluation.verifier_evaluation diagnosis_run.self_evaluation.verifier_evaluation
``` ```
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
## 7. AIOps 规则门禁 ## 7. AIOps 规则门禁
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。 AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
@@ -177,7 +231,7 @@ AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻
结果写入: 结果写入:
```text ```text
diagnosis_session.self_evaluation.aiops_rule_evaluation diagnosis_run.self_evaluation.aiops_rule_evaluation
``` ```
## 8. Eval Baseline ## 8. Eval Baseline
@@ -197,9 +251,8 @@ diagnosis_session.self_evaluation.aiops_rule_evaluation
- 工具参数 schema 校验。 - 工具参数 schema 校验。
- 同一工具调用次数上限。 - 同一工具调用次数上限。
- 工具超时的统一熔断。 - 工具超时的统一熔断。
- 报告中的数值与工具返回值自动对齐校验。 - Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
- Prompt 版本记录和回滚。 - Prompt 版本回滚和更细粒度变更审计。
- Verifier 对 AIOps 报告的 LLM 级事实校验。 - Verifier 对 AIOps 报告的 LLM 级事实校验。
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。 这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
+13 -13
View File
@@ -5,7 +5,7 @@
## 1. 一句话 ## 1. 一句话
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Verifier 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。 SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行;所有工具证据、模型步骤、最终答案、自评估和用户反馈都能按 `sessionId + runId` 回放。
## 2. 一张图 ## 2. 一张图
@@ -16,7 +16,7 @@ flowchart TB
API --> Chat["ChatService"] API --> Chat["ChatService"]
API --> AiOps["AiOpsService"] API --> AiOps["AiOpsService"]
Chat --> ChatFlow["Chat: Planner -> Executor -> Verifier"] Chat --> ChatFlow["Chat: Planner -> Executor -> Gatekeeper -> Verifier -> Composer"]
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"] AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
ChatFlow --> Tools["Evidence Tools"] ChatFlow --> Tools["Evidence Tools"]
@@ -34,14 +34,15 @@ flowchart TB
AiOpsFlow --> Trace AiOpsFlow --> Trace
Tools --> Trace Tools --> Trace
Trace --> Session["diagnosis_session"] Trace --> ChatSession["chat_session"]
Trace --> Run["diagnosis_run"]
Trace --> Step["agent_step"] Trace --> Step["agent_step"]
Trace --> Invocation["tool_invocation"] Trace --> Invocation["tool_invocation"]
Invocation --> Verifier["Verifier / Rule Evaluation"] Invocation --> Verifier["Verifier / Rule Evaluation"]
Verifier --> SelfEval["self_evaluation"] Verifier --> SelfEval["self_evaluation"]
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"] Run --> TraceAPI["GET /api/diagnosis/{sessionId}/trace?runId=..."]
Step --> TraceAPI Step --> TraceAPI
Invocation --> TraceAPI Invocation --> TraceAPI
SelfEval --> TraceAPI SelfEval --> TraceAPI
@@ -55,24 +56,24 @@ flowchart TB
```text ```text
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。 这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
Chat 复杂问题走 Planner -> Executor -> Verifier: Chat 复杂问题走 Planner -> Executor -> Gatekeeper -> Verifier -> Composer:
Planner 负责拆解,Executor 负责调用知识库、日志和指标工具,Verifier 只基于已有工具证据校验最终答案。 Planner 负责拆解,Executor 只负责调用知识库、日志和指标工具并提炼带证据引用的微观事实;Gatekeeper 用代码核对 invocation、raw_path 和 excerpt 是否真实;Verifier 判断这些事实能否由已验真的证据推出;Composer 只把允许表达的结论写成最终答案。
AIOps 告警入口走 Supervisor 调度 Planner/Executor: AIOps 告警入口走 Supervisor 调度 Planner/Executor:
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。 如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
所有过程都会落到 diagnosis_session、agent_step、tool_invocation。 会话元数据会落到 chat_session,每次诊断运行会落到 diagnosis_run,步骤和工具明细通过 run_id 关联。
所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Verifier 怎么判定、用户最后是否反馈有用。 所以我可以用 sessionId + runId 精确回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。
``` ```
## 4. 五个亮点 ## 4. 五个亮点
| 亮点 | 怎么讲 | | 亮点 | 怎么讲 |
|---|---| |---|---|
| 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool | | 可追踪 Agent | 每次诊断都有 `runId`,Trace API 可以回放 run、step、tool;同一 `sessionId` 可有多次独立 run |
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` | | 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 | | RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
| 质量门禁 | Chat Verifier 校验 groundedness,AIOps rule evaluation 控制告警聚焦 | | 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
| 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 | | 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 |
## 5. 三个关键取舍 ## 5. 三个关键取舍
@@ -99,15 +100,14 @@ aiops_rule_evaluation -> AIOps 报告是否聚焦告警并使用证据
| 追问 | 回答方向 | | 追问 | 回答方向 |
|---|---| |---|---|
| 怎么防止幻觉? | Executor 必须用工具;Verifier 只基于 `tool_trace_summary` 校验;LOW_CONFID/REJECT 会降级输出 | | 怎么防止幻觉? | Executor 输出 `executor_evidence_v2`,每个 claim 绑定 `source_invocation_id + raw_path + evidence_excerpt`;Gatekeeper 用 `tool_invocation.retrieval_details.evidence_refs` 核验引用真实性;Verifier 只判断可推导性;Composer 防止把 no-evidence 说成已排除 |
| RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 | | RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 |
| 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter | | 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter |
| AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 | | AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 |
| 下一步怎么演进? | evidence block、邻居 chunk、Playbook、AIOps LLM Verifier、MCP 工具协议化 | | 下一步怎么演进? | 固化 E2E fixture、Prompt version、Gatekeeper 规则配置化、邻居 chunk、AIOps LLM Verifier、MCP 工具协议化 |
## 7. 现场演示入口 ## 7. 现场演示入口
- Demo 脚本:`mvp/demo/ten-minute-interview-demo.md` - Demo 脚本:`mvp/demo/ten-minute-interview-demo.md`
- 故事案例:`interview/story-cases.md` - 故事案例:`interview/story-cases.md`
- 架构细节:`mvp/architecture/README.md` - 架构细节:`mvp/architecture/README.md`
+1 -1
View File
@@ -2,7 +2,7 @@
**更新日期**:2026-07-06 **更新日期**:2026-07-06
**状态**:当前主架构 + 后续演进边界 **状态**:当前主架构 + 后续演进边界
**关联计划**:`mvp/issues/rag-refactor-plan.md` **关联计划**:[`mvp/issues/active/rag-refactor-plan.md`](../issues/active/rag-refactor-plan.md)
## 1. 架构目标 ## 1. 架构目标
+19 -1
View File
@@ -141,7 +141,8 @@ post-retrieval 层再把检索候选归一为:
- 给 Agent 输出 completeness hint。 - 给 Agent 输出 completeness hint。
- 写入 `tool_invocation.relevance_level`。 - 写入 `tool_invocation.relevance_level`。
- 给 Verifier 构造 `tool_trace_summary`。 - 给 Gatekeeper 提供 `evidence_refs` 引用验真源。
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
- 供 EvaluationService 计算 evidence score。 - 供 EvaluationService 计算 evidence score。
## 6. 文档切片和 metadata ## 6. 文档切片和 metadata
@@ -178,6 +179,7 @@ flowchart LR
LookupResult --> Recorder["ToolInvocationRecorder"] LookupResult --> Recorder["ToolInvocationRecorder"]
Recorder --> Invocation["tool_invocation"] Recorder --> Invocation["tool_invocation"]
Invocation --> Trace["DiagnosisTraceService"] Invocation --> Trace["DiagnosisTraceService"]
Invocation --> Gatekeeper["ExecutorGatekeeperService"]
Invocation --> Summary["ToolTraceSummaryService"] Invocation --> Summary["ToolTraceSummaryService"]
Summary --> Verifier["chat_verifier"] Summary --> Verifier["chat_verifier"]
Invocation --> Eval["EvaluationService / RAG eval"] Invocation --> Eval["EvaluationService / RAG eval"]
@@ -205,9 +207,25 @@ success
- evidence status。 - evidence status。
- dedup reason。 - dedup reason。
- evidence block summaries。 - evidence block summaries。
- evidence refs:`raw_path + text`,用于核对 Executor 的 `evidence_excerpt`。
- context pack summary。 - context pack summary。
- rerank trace。 - rerank trace。
其中 `evidence_refs` 是当前 Chat 证据链路的精确引用源:
```json
{
"evidence_refs": [
{
"raw_path": "$.evidence_blocks[0]",
"text": "最小证据文本"
}
]
}
```
如果检索返回 no evidence,应使用 `raw_path=$.no_evidence` 记录负向证据。它只能说明“本次检索没有匹配证据”,不能作为“问题不存在”的证明。
## 8. 去重与行动记忆 ## 8. 去重与行动记忆
当前 session 级去重由 `RetrievedDocTracker` 负责。 当前 session 级去重由 `RetrievedDocTracker` 负责。
+68 -104
View File
@@ -1,68 +1,68 @@
# 会话与 Trace 生命周期 # 会话与 Trace 生命周期
**更新日期**:2026-07-05 **更新日期**:2026-07-10
**状态**:当前可运行架构 **状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md` **参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
## 1. 定位 ## 1. 定位
旧版会话设计以 Redis 会话为主,MySQL 作为可选长期沉淀。当前 MVP 的可追踪诊断已经转为 MySQL Trace 三表为主: 当前 MVP 把“会话态”和“运行态”拆开:
```text ```text
diagnosis_session chat_session(sessionId)
-> agent_step -> diagnosis_run(runId)
-> tool_invocation -> agent_step(runId)
-> tool_invocation(runId)
``` ```
因此本文描述的是当前可运行链路: - `sessionId` 表示多轮会话目录和 Redis 上下文。
- `runId` 表示一次可回放诊断执行。
- `sessionId` 是一次诊断和后续 trace/feedback 的关联键。 - `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
- `diagnosis_session` 保存会话级状态、问题、答案、自评估和反馈。 - `diagnosis_session` 只保留为历史兼容和回滚表。
- `agent_step` 保存每个 Agent 模型调用。
- `tool_invocation` 保存工具调用事实。
- `DiagnosisTraceService` 聚合三类记录,形成可回放 trace。
## 2. 生命周期总图 ## 2. 生命周期总图
```mermaid ```mermaid
flowchart TD flowchart TD
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"] Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
Resolve --> Create["create or reset diagnosis_session"] Resolve --> Session["ensure chat_session metadata"]
Create --> Running["status = RUNNING"] Session --> Run["create diagnosis_run(runId)"]
Run --> Running["run.status = RUNNING"]
Running --> Agent["Agent workflow"] Running --> Agent["Agent workflow"]
Agent --> StepHook["AgentLoggingHook"] Agent --> Context["execution context(sessionId, runId)"]
StepHook --> Step["agent_step"] Context --> StepHook["AgentLoggingHook"]
Agent --> Tool["Evidence tools"] StepHook --> Step["agent_step(session_id, run_id)"]
Tool --> Invocation["tool_invocation"] Context --> Tool["Evidence tools"]
Tool --> Invocation["tool_invocation(session_id, run_id)"]
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
Agent --> Final{"workflow result"} Agent --> Final{"workflow result"}
Final -->|success| Success["status = SUCCESS, answer saved"] Final -->|success| Success["run.status = SUCCESS, answer saved"]
Final -->|failed| Failed["status = FAILED"] Final -->|failed| Failed["run.status = FAILED"]
Success --> Evaluation["self_evaluation merge"] Success --> Evaluation["diagnosis_run.self_evaluation merge"]
Failed --> Evaluation Failed --> Evaluation
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace"] Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
Success --> Feedback["POST /api/feedback"] Success --> Feedback["POST /api/feedback(sessionId, runId)"]
Feedback --> Case["useful -> case_library"] Feedback --> Case["useful -> case_library(run_id)"]
``` ```
## 3. sessionId 规则 ## 3. ID 规则
| 链路 | sessionId 来源 | | ID | 来源 | 含义 |
|---|---| |---|---|---|
| Chat | 如果请求带 sessionId,则复用;否则生成短 UUID | | `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
| AIOps | 如果 payload 带 sessionId,则复用;否则生成 UUID | | `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
| Trace | URL path 中的 `{sessionId}` |
| Feedback | request body 中的 `sessionId` |
设计含义: 设计含义:
- 同一个 `sessionId` 可以贯穿诊断、trace 查询和用户反馈。 - 同一个 `sessionId` 可以贯穿多轮 Chat。
- 当前诊断开始时会重置当前 session 的运行态字段,例如 answer、duration、step/tool count。 - 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
- `sessionId` 是业务关联键,不依赖数据库自增 ID 暴露给外部。 - Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
## 4. 状态流转 ## 4. 运行状态流转
```mermaid ```mermaid
stateDiagram-v2 stateDiagram-v2
@@ -76,14 +76,14 @@ stateDiagram-v2
字段边界: 字段边界:
| 字段 | 含义 | | 字段 | 所属表 | 含义 |
|---|---| |---|---|---|
| `status` | 执行状态:`PENDING` / `RUNNING` / `SUCCESS` / `FAILED` | | `status` | `diagnosis_run` | 单次运行执行状态 |
| `answer` | Agent 最终返回给用户的报告或答复 | | `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
| `self_evaluation` | 系统自评估 JSON | | `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
| `feedback` | 用户反馈:`useful` / `not_useful` / null | | `feedback` | `diagnosis_run` | 本次运行用户反馈 |
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 session,仍然应该是 `SUCCESS + feedback=not_useful`。 `feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
## 5. agent_step 写入 ## 5. agent_step 写入
@@ -92,105 +92,69 @@ stateDiagram-v2
```mermaid ```mermaid
sequenceDiagram sequenceDiagram
autonumber autonumber
participant Agent as ReactAgent participant Agent as Agent
participant Hook as AgentLoggingHook participant Hook as AgentLoggingHook
participant DB as agent_step participant DB as agent_step
Agent->>Hook: before_model(messages, sessionId) Agent->>Hook: before_model(messages, sessionId, runId)
Hook->>DB: insert step_index / agent_name / model_input Hook->>DB: insert step(session_id, run_id, model_input, step_index)
Agent-->>Agent: model call Agent->>Hook: after_model(output, sessionId, runId)
Agent->>Hook: after_model(messages, sessionId) Hook->>DB: update model_output, duration, token_count, has_tool_call
Hook->>DB: update model_output / thought / has_tool_call / duration / token_count
``` ```
当前记录: 新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
- `session_id`
- `step_index`
- `agent_name`
- `model_input`
- `model_output`
- `thought`
- `has_tool_call`
- `duration_ms`
- `token_count`
## 6. tool_invocation 写入 ## 6. tool_invocation 写入
工具调用记录真实工具事实,不记录模型猜测。 工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
关键字段:
```text ```text
session_id ToolInvocationRecorder
step_id -> tool_invocation.session_id
tool_name -> tool_invocation.run_id
input_params -> retrieval_details / evidence_refs
output_preview
output_length
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
relevance_level
dedup_reason
duration_ms
success
error_message
``` ```
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对非检索工具,检索字段可以为空。 Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
## 7. Trace API 聚合 ## 7. Trace API 聚合
```text ```text
GET /api/diagnosis/{sessionId}/trace GET /api/diagnosis/{sessionId}/trace
GET /api/diagnosis/{sessionId}/trace?runId=run-...
``` ```
聚合逻辑: 聚合逻辑:
```text ```text
diagnosis_session by sessionId diagnosis_run by sessionId + runId
+ agent_step ordered by step_index + chat_session metadata when available
+ tool_invocation ordered by id + agent_step where run_id = runId, ordered by the Trace API
+ tool_invocation where run_id = runId order by id
-> DiagnosisTraceResponse -> DiagnosisTraceResponse
``` ```
Trace 视图回答的问题: 当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
- 这次诊断是否成功?
- 哪些 Agent 参与了?
- 每一步模型输入输出是什么摘要?
- 调用了哪些工具?
- 工具返回了什么证据?
- Verifier / AIOps rule 是否通过?
- 用户是否反馈有用?
## 8. Chat 与 AIOps 差异 ## 8. Chat 与 AIOps 差异
| 维度 | Chat | AIOps | | 维度 | Chat | AIOps |
|---|---|---| |---|---|---|
| `agent_flow` | `CHAT` | `AI_OPS` | | `agent_flow` | `CHAT` | `AI_OPS` |
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Verifier | `SupervisorAgent`: Planner + Executor | | 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` | | 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
| 答案字段 | Chat 最终答复 | 告警分析报告 | | 答案字段 | Chat 最终答复 | 告警分析报告 |
| payload | 用户自然语言 + history | alert payload 或 auto-discovery | | runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
## 9. 清理与边界 ## 9. 清理与边界
当前会话持久化边界: - Redis 会话历史用于多轮上下文,不是长期审计记录。
- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
- MySQL Trace 记录是主要可回放来源。 - `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
- Chat 历史仍可作为请求上下文传入 Agent,但不是本文档的主持久化模型。 - `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
- Redis 主会话存储是历史设计,不作为当前架构事实。
- `RetrievedDocTracker` 是 session 级运行时去重状态,诊断结束后清理。
## 10. 后续增强 ## 10. 后续增强
可考虑:
1. Trace API 增加更结构化的 `self_evaluation` 展示。 1. Trace API 增加更结构化的 `self_evaluation` 展示。
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。 2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。 3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
4. 为 Trace 增加导出能力,服务面试演示和回归分析。
@@ -0,0 +1,17 @@
# 2026-07-09 MVP 文档清理归档
本目录保存本次清理中从当前入口移出的历史设计材料。这些文档仍有追溯价值,但不再代表当前可运行实现。
## 归档内容
| 目录 | 内容 | 归档原因 |
|---|---|---|
| `discuss/` | 早期 Executor Prompt、L0、RAG 讨论稿 | 已被当前 architecture、OpenSpec change 和 issue 取代 |
| `plan/` | `session-storage-design.md` | 会话存储已实现,当前表以 Flyway 和 `mvp/tables/` 为准 |
| `notes/` | 早期工程决策和 Demo Trace 验收笔记 | 相关内容已沉淀到 architecture、demo、eval 和 devflow |
## 使用原则
- 当前架构以 `mvp/architecture/` 为准。
- 当前表结构以 `mvp/tables/`、Flyway migration 和实体类为准。
- 当前问题入口以 `mvp/issues/README.md` 为准。
+45 -9
View File
@@ -6,9 +6,15 @@
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。 - `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
- `interview-walkthrough.md`:面试讲解话术。 - `interview-walkthrough.md`:面试讲解话术。
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
- `trace-inspection-checklist.md`:Trace 字段检查清单。 - `trace-inspection-checklist.md`:Trace 字段检查清单。
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。 - `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。 - `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
## 1. 前置条件 ## 1. 前置条件
@@ -33,7 +39,7 @@ http://localhost:9900
最快方式: 最快方式:
```powershell ```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
``` ```
脚本会生成: 脚本会生成:
@@ -42,6 +48,7 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
mvp/demo/output/chat-response.json mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json mvp/demo/output/feedback-response.json
mvp/demo/output/interview-demo-summary.json
``` ```
手动请求: 手动请求:
@@ -60,10 +67,23 @@ Invoke-RestMethod `
-Body $body -Body $body
``` ```
如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
```powershell
$chat = Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
$runId = $chat.data.runId
```
期望结果: 期望结果:
- `data.success = true` - `data.success = true`
- `data.sessionId = mvp-demo-payment-timeout-001` - `data.sessionId = mvp-demo-payment-timeout-001`
- `data.runId` 为本次诊断运行的唯一 ID
- `data.answer` 包含诊断答复 - `data.answer` 包含诊断答复
## 4. 查询 Trace ## 4. 查询 Trace
@@ -71,22 +91,27 @@ Invoke-RestMethod `
```powershell ```powershell
Invoke-RestMethod ` Invoke-RestMethod `
-Method Get ` -Method Get `
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace" -Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
``` ```
期望结果: 期望结果:
- `code = 200` - `code = 200`
- `data.runId` 等于 `$runId`
- `data.session.sessionId` 等于 Chat session id - `data.session.sessionId` 等于 Chat session id
- `data.run.runId` 等于 `$runId`
- `data.steps` 包含 planner / executor / verifier 等步骤 - `data.steps` 包含 planner / executor / verifier 等步骤
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 - `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation - `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
## 5. 提交反馈 ## 5. 提交反馈
```powershell ```powershell
$feedback = @{ $feedback = @{
sessionId = $sessionId sessionId = $sessionId
runId = $runId
feedback = "useful" feedback = "useful"
} | ConvertTo-Json } | ConvertTo-Json
@@ -100,7 +125,8 @@ Invoke-RestMethod `
期望结果: 期望结果:
- `success = true` - `success = true`
- 后续 Trace 中 `data.session.feedback = useful` - `runId = $runId`
- 后续精确 Trace 中 `data.session.feedback = useful`
- useful 反馈会尝试沉淀 `case_library` - useful 反馈会尝试沉淀 `case_library`
## 6. AIOps 告警诊断 Demo ## 6. AIOps 告警诊断 Demo
@@ -126,19 +152,19 @@ Invoke-WebRequest `
期望结果: 期望结果:
- SSE 首条包含 `session` 消息,sessionId 为 `mvp-demo-aiops-payment-cpu-001` - SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
- 后续流式输出包含 AIOps 告警分析报告 - 后续流式输出包含 AIOps 告警分析报告
- 报告聚焦输入的 `HighCPUUsage/payment-service` - 报告聚焦输入的 `HighCPUUsage/payment-service`
- 同一 session 的 Trace 中 `data.session.agentFlow = AI_OPS` - 精确 Trace 中 `data.session.agentFlow = AI_OPS`
- `data.session.answer` 包含最终告警报告 - `data.session.answer` 包含最终告警报告
- `data.toolInvocations` 包含证据工具调用 - `data.toolInvocations` 包含证据工具调用
查询 AIOps Trace: 查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
```powershell ```powershell
Invoke-RestMethod ` Invoke-RestMethod `
-Method Get ` -Method Get `
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace" -Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
``` ```
## 7. Demo 主线 ## 7. Demo 主线
@@ -146,7 +172,7 @@ Invoke-RestMethod `
Chat 主线: Chat 主线:
```text ```text
一个 session id 一个 session id + 一个 run id
-> 用户问题 -> 用户问题
-> 多 Agent 执行 -> 多 Agent 执行
-> 证据工具 -> 证据工具
@@ -159,7 +185,7 @@ Chat 主线:
AIOps 主线: AIOps 主线:
```text ```text
一个 session id 一个 session id + 一个 run id
-> 告警 payload -> 告警 payload
-> AIOps Planner / Executor -> AIOps Planner / Executor
-> 证据工具 -> 证据工具
@@ -167,3 +193,13 @@ AIOps 主线:
-> AIOps rule evaluation -> AIOps rule evaluation
-> Trace API 回放 -> Trace API 回放
``` ```
## 8. Evidence Pipeline 场景矩阵
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
这样可以同时展示真实链路和确定性回归能力。
+7 -7
View File
@@ -2,7 +2,7 @@
## 1. 目标 ## 1. 目标
验证旧版 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。 验证 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
## 2. 输入 ## 2. 输入
@@ -25,14 +25,14 @@
## 3. 验收标准 ## 3. 验收标准
1. SSE 流输出 `session` 消息,且包含请求中的 session id。 1. SSE 流首条输出 `type=metadata` 的 `message` 事件,且包含请求中的 session id 和本次 AIOps run id。
2. AIOps 执行创建或更新 `diagnosis_session`,并写入 `agent_flow = AI_OPS`。 2. AIOps 执行创建 `diagnosis_run`,并写入 `agent_flow = AI_OPS`。
3. 持久化的 session query 包含告警名、服务名、等级、时间范围和描述。 3. 持久化的 run query 包含告警名、服务名、等级、时间范围和描述。
4. 如果生成最终报告,`diagnosis_session.answer` 包含该报告。 4. 如果生成最终报告,`diagnosis_run.answer` 包含该报告。
5. `GET /api/diagnosis/{sessionId}/trace` 返回 AIOps session、按顺序排列的 agent steps 和 tool invocations。 5. `GET /api/diagnosis/{sessionId}/trace?runId=...` 返回 AIOps run、按顺序排列的 agent steps 和 tool invocations。
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。 6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。 7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
8. `self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。 8. `diagnosis_run.self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
## 4. 已知边界 ## 4. 已知边界
+63
View File
@@ -0,0 +1,63 @@
# Evidence Pipeline Demo Scenarios
这份清单用于面试时说明 Chat 证据链路如何覆盖 `PASS`、`LOW_CONFID`、`REJECT` 和 no-evidence 场景。
重点区别:
- Live demo 证明本地服务、工具、Trace、Feedback 主链路能跑通。
- Fixture-backed eval 证明固定安全场景可以确定性回归,不依赖 LLM 当场随机输出。
## Scenario Matrix
| 场景 | 类型 | 输入/证据 | 期望讲点 |
|---|---|---|---|
| Payment timeout | Live 主路径 | `requests/payment-timeout-chat.json` | 完整 Chat -> Trace -> Feedback 闭环 |
| Narrow HighCPU observation | Live 可尝试 + fixture-backed | `requests/narrow-highcpu-chat.json` / `mvp/eval/fixtures/narrow-highcpu-observation-pass.json` | Executor 只输出观察类 claim,Gatekeeper 验引用,Verifier PASS |
| Hikari no-evidence | Live 可尝试 + fixture-backed | `requests/hikari-no-evidence-chat.json` / `mvp/eval/fixtures/hikari-no-evidence-negative-observation-pass.json` | `$.no_evidence` 只表示本次查询无匹配证据,Composer 不说“已排除” |
| Unsupported claim filtering | Fixture-backed | `requests/safety-unsupported-claim-chat.json` / `mvp/eval/fixtures/unsupported-claim-filtering-low-confid.json` | Verifier 将 unsupported claim 降为 LOW_CONFID,最终答案不确认“主库故障” |
| Fabricated invocation reject | Fixture-backed | `mvp/eval/fixtures/gatekeeper-fabricated-invocation-reject.json` | Gatekeeper 拦截伪造 invocation,最终 REJECT/降级 |
| Composer fallback | Fixture-backed | `mvp/eval/fixtures/composer-fallback-no-raw-json-low-confid.json` | 即使 Composer 输出异常,也不能把 Executor JSON 泄漏给用户 |
## Trace Fields To Inspect
| 能力 | JSON path |
|---|---|
| Executor V2 输出 | `data.session.selfEvaluation.verifier_evaluation.executor_structured_output` |
| Gatekeeper 结果 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` |
| Gatekeeper 规则版本 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` |
| 证据绑定校验 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.checked_bindings` |
| Verifier claim checks | `data.session.selfEvaluation.verifier_evaluation.claim_checks` |
| Composer 输出 | `data.session.selfEvaluation.verifier_evaluation.composer_output` |
| 工具证据引用 | `data.toolInvocations[*].retrievalDetails.evidence_refs` |
## How To Present It
```text
我把现场 demo 和固定 eval 分开。
现场 demo 证明系统能跑通真实链路;
fixture-backed eval 证明反幻觉安全场景可以稳定回归。
Gatekeeper 的规则版本也进入 trace,所以后续调整阈值或规则时可以审计。
```
## Optional Live Requests
手动发送某个请求样例:
```powershell
$body = Get-Content -Raw -Encoding UTF8 "mvp/demo/requests/narrow-highcpu-chat.json"
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
然后保留响应里的 `runId`,查询同一 run 的 Trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace?runId=$runId"
```
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
+33
View File
@@ -0,0 +1,33 @@
# 面试追问 Q&A
## 为什么不用普通 Chatbot?
这个项目的重点不是生成一段诊断文本,而是把诊断拆成可审计链路:Planner 拆解问题,Executor 调工具拿证据,Gatekeeper 用代码核验证据引用,Verifier 判断可推导性,Composer 生成最终表达。`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行,Trace 和 Feedback 都可以按 `sessionId + runId` 回放和定位。
## 为什么 RAG 要做成显式工具?
`lookup_knowledge` 保持显式工具调用,才能在 `tool_invocation` 里看到 Agent 查了什么、命中了什么、相关性等级是什么,以及最终答案是否真的使用了这些证据。隐式 Advisor 更方便,但不利于审计 Agent 决策。
## 怎么防止 Executor 幻觉?
Executor 不直接负责最终用户答案,而是输出 `executor_evidence_v2` 的微观事实和证据引用。Gatekeeper 会校验 `source_invocation_id`、`raw_path`、`evidence_excerpt` 是否真实存在;Verifier 再判断 claim 是否能由已验真的证据推出;Composer 只表达 Verifier 允许输出的内容。
## LOW_CONFID 是失败吗?
不是。`LOW_CONFID` 表示当前证据不足以支撑强结论,但系统仍然可以安全表达已确认事实和缺失信息。面试时可以把它作为“没有证据就不强答”的质量门禁,而不是模型能力失败。
## Prompt 改了怎么审计?
Chat verifier evaluation 里会记录 `prompt_audit.version`,并列出 planner、executor、verifier、composer 的 Prompt 版本和资源路径。它不保存完整 Prompt 文本,只保留用于回放和回归解释的紧凑元数据。
## Gatekeeper 改了怎么审计?
Gatekeeper 结果里记录 `gatekeeper_result.rule_set_version` 和已启用规则元数据摘要。规则执行仍是确定性 Java 代码,版本和规则元数据用于解释“这次引用验真用的是哪套规则”。
## 为什么现在不拆 SubAgent?
当前 MVP 的主要风险不是 Agent 数量不够,而是证据、验证和回归是否稳定。文档里的演进路线把 SubAgent 放在 P2:等故障类型、工具权限和评测集足够明确后再拆,避免只是移动复杂度。
## 为什么 baseline 比 live demo 更重要?
live demo 证明链路在当前环境能跑通,但 LLM 和外部依赖会波动。`mvp/eval` 的固定 fixture baseline 是确定性回归来源,用来判断 Prompt、工具、Gatekeeper、Verifier 或 Composer 的改动有没有让系统退化。
+13 -5
View File
@@ -13,7 +13,7 @@
关键主张不是“模型回答了一次”,而是: 关键主张不是“模型回答了一次”,而是:
```text ```text
系统能展示用了什么证据、答案如何被检查、如何用 sessionId 回放整次诊断。 系统能展示用了什么证据、答案如何被检查、如何用 sessionId + runId 精确回放这次诊断。
``` ```
## 2. Demo 流程 ## 2. Demo 流程
@@ -23,7 +23,8 @@
3. 打开 `mvp/demo/output/chat-response.json`。 3. 打开 `mvp/demo/output/chat-response.json`。
4. 打开 `mvp/demo/output/trace-response.json`。 4. 打开 `mvp/demo/output/trace-response.json`。
5. 指出证据工具和 verifier evaluation。 5. 指出证据工具和 verifier evaluation。
6. 提交 feedback,并展示它挂在同一个 session 上。 6. 提交 feedback,并展示它挂在当前 run 上。
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
## 3. 命令 ## 3. 命令
@@ -58,7 +59,7 @@ mvp/demo/output/chat-response.json
话术: 话术:
```text ```text
这是用户看到的答案。这里的 sessionId 是稳定的,所以我后面可以追踪这一次回答是怎么来的。 这是用户看到的答案。这里的 sessionId 是稳定的,同时响应里会返回 runId,所以我后面可以精确追踪这一次回答是怎么来的。
``` ```
### 4.2 证据 Trace ### 4.2 证据 Trace
@@ -111,7 +112,7 @@ mvp/demo/output/feedback-response.json
话术: 话术:
```text ```text
feedback 会挂在同一个 diagnosis session 上。 feedback 会挂在当前 diagnosis run 上。
这让后续挖掘 useful case 或 not_useful bad case 成为可能。 这让后续挖掘 useful case 或 not_useful bad case 成为可能。
``` ```
@@ -125,6 +126,14 @@ Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。 这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
``` ```
如果被问到怎么防止证据归因幻觉,可以补充:
```text
Executor 的 claim 必须绑定 source_invocation_id、raw_path 和 evidence_excerpt。
Gatekeeper 用代码核验这些引用,并把 rule_set_version 写进 trace。
Verifier 只判断已核验证据能否推出 claim,Composer 只表达允许输出的内容。
```
## 5. 强面试表达 ## 5. 强面试表达
```text ```text
@@ -141,4 +150,3 @@ traceability、evidence persistence、verifier gating、feedback 和 regression
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。 mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。 密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
``` ```
+6 -4
View File
@@ -12,14 +12,16 @@
## 3. 验收标准 ## 3. 验收标准
1. Chat 返回成功答复,且 session id 与请求一致。 1. Chat 返回成功答复,且 session id 与请求一致,并返回本次诊断的 run id。
2. Trace API 返回 session 元数据、最终答案、按顺序排列的 agent steps 和 tool invocations。 2. Trace API 使用 `sessionId + runId` 返回会话元数据、运行摘要、最终答案、按顺序排列的 agent steps 和 tool invocations。
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。 3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
4. 可以使用同一个 session id 提交反馈。 4. 可以使用同一个 session id 和本次 run id 提交反馈。
5. 后续 Trace 查询能看到已持久化的 feedback 值。 5. 后续精确 Trace 查询能看到已持久化的 feedback 值。
## 4. 需要检查的 Trace 字段 ## 4. 需要检查的 Trace 字段
- `data.runId`
- `data.run.runId`
- `data.session.query` - `data.session.query`
- `data.session.answer` - `data.session.answer`
- `data.session.selfEvaluation` - `data.session.selfEvaluation`
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-hikari-no-evidence-001",
"Question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。"
}
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-narrow-highcpu-001",
"Question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。"
}
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-safety-unsupported-001",
"Question": "订单超时是否可以确认由数据库主库故障导致?请只基于当前证据回答。"
}
@@ -0,0 +1,168 @@
param(
[string]$BaseUrl = "http://localhost:9900",
[string]$SessionId = "mvp-demo-interview-payment-timeout-001",
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
[string]$OutputDir = "$PSScriptRoot/../output"
)
$ErrorActionPreference = "Stop"
function Test-ServiceReachable {
param([string]$Url)
try {
$request = [System.Net.WebRequest]::Create($Url)
$request.Method = "GET"
$request.Timeout = 5000
$response = $request.GetResponse()
$response.Close()
return $true
} catch [System.Net.WebException] {
if ($_.Exception.Response -ne $null) {
$_.Exception.Response.Close()
return $true
}
return $false
}
}
function Get-TraceData {
param($TraceResponse)
if ($TraceResponse.PSObject.Properties.Name -contains "data") {
return $TraceResponse.data
}
return $TraceResponse
}
function Get-SelfEvaluation {
param($TraceData)
if ($null -eq $TraceData -or $null -eq $TraceData.session) {
return $null
}
return $TraceData.session.selfEvaluation
}
function Get-ToolNames {
param($TraceData)
if ($null -eq $TraceData -or $null -eq $TraceData.toolInvocations) {
return @()
}
return @($TraceData.toolInvocations | ForEach-Object { $_.toolName } | Where-Object { $_ } | Sort-Object -Unique)
}
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
Write-Host "Running interview demo preflight..."
Write-Host "BaseUrl: $BaseUrl"
Write-Host "SessionId: $SessionId"
if (-not (Test-ServiceReachable -Url $BaseUrl)) {
throw "Service is not reachable: $BaseUrl. Start the app with mvp-demo profile first: mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo"
}
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
$request.Id = $SessionId
$body = $request | ConvertTo-Json -Depth 8
$chatRequest = @{
Method = "Post"
Uri = "$BaseUrl/api/chat"
ContentType = "application/json; charset=utf-8"
Body = $body
}
$chat = Invoke-RestMethod @chatRequest
$chatPath = Join-Path $OutputDir "chat-response.json"
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
$runId = $chat.data.runId
if (-not $runId) {
throw "Chat response did not include runId; exact trace verification cannot continue."
}
$traceRequest = @{
Method = "Get"
Uri = "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
}
$trace = Invoke-RestMethod @traceRequest
$tracePath = Join-Path $OutputDir "trace-response.json"
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
$feedbackBody = @{
sessionId = $SessionId
runId = $runId
feedback = "useful"
} | ConvertTo-Json
$feedbackRequest = @{
Method = "Post"
Uri = "$BaseUrl/api/feedback"
ContentType = "application/json; charset=utf-8"
Body = $feedbackBody
}
$feedback = Invoke-RestMethod @feedbackRequest
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
$traceData = Get-TraceData -TraceResponse $trace
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
$verifierEvaluation = $null
if ($null -ne $selfEvaluation) {
$verifierEvaluation = $selfEvaluation.verifier_evaluation
}
$gatekeeperResult = $null
$promptAudit = $null
if ($null -ne $verifierEvaluation) {
$gatekeeperResult = $verifierEvaluation.gatekeeper_result
$promptAudit = $verifierEvaluation.prompt_audit
}
$verdict = $null
$gatekeeperStatus = $null
$gatekeeperRuleSetVersion = $null
$promptAuditVersion = $null
if ($null -ne $verifierEvaluation) {
$verdict = $verifierEvaluation.verdict
}
if ($null -ne $gatekeeperResult) {
$gatekeeperStatus = $gatekeeperResult.status
$gatekeeperRuleSetVersion = $gatekeeperResult.rule_set_version
}
if ($null -ne $promptAudit) {
$promptAuditVersion = $promptAudit.version
}
$toolNames = Get-ToolNames -TraceData $traceData
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
$summary = [ordered]@{
sessionId = $SessionId
runId = $runId
baseUrl = $BaseUrl
chatSuccess = $chat.data.success
verdict = $verdict
gatekeeperStatus = $gatekeeperStatus
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
promptAuditVersion = $promptAuditVersion
toolNames = $toolNames
paths = [ordered]@{
chat = $chatPath
trace = $tracePath
feedback = $feedbackPath
summary = $summaryPath
}
}
$summary | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path $summaryPath
Write-Host ""
Write-Host "Interview demo preflight completed."
Write-Host "Verdict: $($summary.verdict)"
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
Write-Host "Summary: $summaryPath"
@@ -26,15 +26,22 @@ $chat = Invoke-RestMethod `
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json" $chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json" Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
$runId = $chat.data.runId
if (-not $runId) {
throw "Chat 响应缺少 runId,无法查询精确 Trace。"
}
Write-Host "RunId: $runId"
$trace = Invoke-RestMethod ` $trace = Invoke-RestMethod `
-Method Get ` -Method Get `
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace" -Uri "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json" $trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json" Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
$feedbackBody = @{ $feedbackBody = @{
sessionId = $SessionId sessionId = $SessionId
runId = $runId
feedback = "useful" feedback = "useful"
} | ConvertTo-Json } | ConvertTo-Json
+13 -4
View File
@@ -37,7 +37,7 @@ http://localhost:9900
推荐使用固定脚本: 推荐使用固定脚本:
```powershell ```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
``` ```
脚本会写出: 脚本会写出:
@@ -46,13 +46,14 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
mvp/demo/output/chat-response.json mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json mvp/demo/output/feedback-response.json
mvp/demo/output/interview-demo-summary.json
``` ```
现场话术: 现场话术:
```text ```text
这里我用固定 sessionId 跑一个支付接口超时问题。 这里我用固定 sessionId 跑一个支付接口超时问题。
固定 sessionId 的好处是,后面 trace 和 feedback 都能关联到同一次诊断。 固定 sessionId 的好处是保留多轮上下文;每次诊断还会返回 runId,后面 trace 和 feedback 都用这个 runId 精确关联到同一次运行。
``` ```
## 3. 展示用户答案 ## 3. 展示用户答案
@@ -67,6 +68,7 @@ mvp/demo/output/chat-response.json
```text ```text
data.sessionId data.sessionId
data.runId
data.answer data.answer
``` ```
@@ -75,7 +77,7 @@ data.answer
```text ```text
这是用户看到的答案。 这是用户看到的答案。
但这个项目的重点不是这段文字,而是这段文字是否有证据链。 但这个项目的重点不是这段文字,而是这段文字是否有证据链。
接下来我用同一个 sessionId 查 trace。 接下来我用同一个 sessionId 加 runId 查 trace。
``` ```
## 4. 展示 Trace ## 4. 展示 Trace
@@ -98,6 +100,8 @@ data.toolInvocations[*].outputPreview
data.toolInvocations[*].retrievalLayer data.toolInvocations[*].retrievalLayer
data.toolInvocations[*].relevanceLevel data.toolInvocations[*].relevanceLevel
data.summary.hasVerifierEvaluation data.summary.hasVerifierEvaluation
data.session.selfEvaluation.verifier_evaluation.prompt_audit.version
data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version
``` ```
现场话术: 现场话术:
@@ -155,6 +159,10 @@ Verifier 不做新检索,只看工具 trace 汇总。
如果 PASS,就输出原答案。 如果 PASS,就输出原答案。
如果 LOW_CONFID,可以补证据或加低置信提示。 如果 LOW_CONFID,可以补证据或加低置信提示。
如果 REJECT,就降级输出,只保留已确认信息。 如果 REJECT,就降级输出,只保留已确认信息。
Prompt 和 Gatekeeper 的版本也会进入 trace。
`prompt_audit.version` 用于说明本次 Chat 使用哪套 Prompt 契约,`gatekeeper_result.rule_set_version` 用于说明引用验真的规则版本。
固定 fixture baseline 是回归判断来源,live demo 主要证明当前环境链路可跑通。
``` ```
## 7. 展示反馈闭环 ## 7. 展示反馈闭环
@@ -175,7 +183,7 @@ caseId
现场话术: 现场话术:
```text ```text
用户反馈 useful 会写回同一个 diagnosis_session。 用户反馈 useful 会写回当前 diagnosis_run。
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。 后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
这里 status 和 feedback 是分开的: 这里 status 和 feedback 是分开的:
@@ -226,6 +234,7 @@ AIOps 有两个模式。
mvp/demo/output/chat-response.json mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json mvp/demo/output/feedback-response.json
mvp/demo/output/interview-demo-summary.json
``` ```
降级话术: 降级话术:
+10 -5
View File
@@ -1,16 +1,20 @@
# Trace 检查清单 # Trace 检查清单
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。 运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
## 1. Session ## 1. Session
| JSON path | 检查点 | 面试讲点 | | JSON path | 检查点 | 面试讲点 |
|---|---|---| |---|---|---|
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | 一个 session id 串起 chat、工具、verifier、feedback 和 trace | | `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 | | `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace | | `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 | | `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 | | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
## 2. Agent 步骤 ## 2. Agent 步骤
@@ -30,6 +34,7 @@
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload | | `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 | | `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 | | `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 | | `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
## 4. Summary ## 4. Summary
@@ -44,10 +49,10 @@
## 5. 好的结果长什么样 ## 5. 好的结果长什么样
```text ```text
同一个 session id 同一个 session id + run id
-> 最终答案 -> 最终答案
-> 持久化 agent steps -> 持久化 agent steps
-> 持久化 evidence tool calls -> 持久化 evidence tool calls
-> verifier / self-evaluation -> verifier / self-evaluation
-> feedback attached to the same session -> feedback attached to the same run
``` ```
+54 -14
View File
@@ -1,6 +1,16 @@
# Diagnosis Eval Harness # Diagnosis Eval Harness
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent. This folder contains the fixed offline regression set for the MVP diagnosis Agent.
## Background
The current diagnosis chain is:
```text
Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer
```
Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo.
## Scope ## Scope
@@ -14,7 +24,28 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
## Current Mode ## Current Mode
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM. The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
The committed baseline currently contains:
```text
12 fixed cases
12 passing fixture evaluations
5 PASS verdicts
6 LOW_CONFID verdicts
1 REJECT verdict
```
The V2 evidence-pipeline matrix covers:
- Positive supported evidence for a narrow HighCPU observation.
- No-evidence `negative_observation` using `$.no_evidence`.
- Gatekeeper failure for a fabricated tool invocation reference.
- Unsupported claim filtering before the final answer.
- Composer fallback rendering without raw Executor JSON leakage.
- Gatekeeper rule set version audit for new matrix fixtures.
- Prompt audit version checks for planner, executor, verifier, and composer prompts.
- Gatekeeper rule metadata checks for enabled rule id and default severity.
## Verification ## Verification
@@ -24,29 +55,38 @@ Run the focused evaluator test:
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
``` ```
The committed baseline report represents the current fixed fixture set: Run the broader phase-5 regression set:
```text ```powershell
5 fixed cases mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
5 passing fixture evaluations
2 PASS verdicts
3 LOW_CONFID verdicts
``` ```
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together. When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
## Interview Story ## Regression Signal
The harness gives the MVP a repeatable baseline: The harness is deterministic code, not an LLM judge:
```text ```text
fixed diagnosis case fixed diagnosis case
-> saved or runtime trace -> saved trace fixture
-> rule-based trace validation -> rule-based trace validation
-> JSON / Markdown report -> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, and verifier behavior -> regression signal for prompts, tools, retrieval, verifier, and composer behavior
``` ```
Stage 5 adds these V2 checks:
- `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`.
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
- `claim_checks` must be structurally auditable.
- Composer output must record whether normal parsing or fallback rendering was used.
- Gatekeeper rule set version can be asserted per fixture.
- Prompt audit version and per-prompt versions can be asserted per fixture.
- Gatekeeper rule metadata can be required per fixture.
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
## Baseline Diff ## Baseline Diff
Baseline diff compares a current report against `reports/baseline-report.json`. Baseline diff compares a current report against `reports/baseline-report.json`.
@@ -58,4 +98,4 @@ current report
-> regressions, improvements, and changed signals -> regressions, improvements, and changed signals
``` ```
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline? Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?
+141
View File
@@ -1,4 +1,67 @@
[ [
{
"id": "narrow-highcpu-observation",
"title": "Narrow HighCPU observation",
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
"traceFixture": "narrow-highcpu-observation-pass.json",
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics"],
"allowedVerdicts": ["PASS"],
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["数据库连接池"]
},
{
"id": "prompt-gatekeeper-audit-closure",
"title": "Prompt and Gatekeeper audit closure",
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
"traceFixture": "prompt-gatekeeper-audit-closure-pass.json",
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics"],
"allowedVerdicts": ["PASS"],
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["数据库连接池"],
"requirePromptAudit": true,
"expectedPromptAuditVersion": "chat-prompts-v1",
"expectedPromptVersions": {
"chat_planner": "chat-planner-v1",
"chat_executor": "chat-executor-v2",
"chat_verifier": "chat-verifier-v2",
"chat_composer": "chat-composer-v1"
},
"requireGatekeeperRules": true
},
{
"id": "hikari-no-evidence-negative-observation",
"title": "Hikari no-evidence negative observation",
"question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
"traceFixture": "hikari-no-evidence-negative-observation-pass.json",
"expectedRootCauseKeywords": ["未检索到", "HikariCP", "匹配证据"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["PASS"],
"forbiddenAnswerKeywords": ["已排除", "确认没有", "日志层面已排除"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["已排除 HikariCP"]
},
{ {
"id": "payment-timeout", "id": "payment-timeout",
"title": "Payment API timeout", "title": "Payment API timeout",
@@ -53,5 +116,83 @@
"requiredEvidenceTools": ["query_metrics", "query_logs"], "requiredEvidenceTools": ["query_metrics", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"], "allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["可以忽略"] "forbiddenAnswerKeywords": ["可以忽略"]
},
{
"id": "gatekeeper-fabricated-invocation",
"title": "Gatekeeper fabricated invocation",
"question": "支付失败是否能确认由日志中的连接池耗尽导致?",
"traceFixture": "gatekeeper-fabricated-invocation-reject.json",
"expectedRootCauseKeywords": ["证据", "引用", "失败"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["REJECT", "LOW_CONFID"],
"forbiddenAnswerKeywords": ["已经完全确认"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["fail"],
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["连接池耗尽导致支付失败"]
},
{
"id": "unsupported-claim-filtering",
"title": "Unsupported claim filtering",
"question": "订单超时是否可以确认由数据库主库故障导致?",
"traceFixture": "unsupported-claim-filtering-low-confid.json",
"expectedRootCauseKeywords": ["超时", "证据"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["LOW_CONFID"],
"forbiddenAnswerKeywords": ["已经确认"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["主库故障"]
},
{
"id": "audit-metadata-low-confid",
"title": "Audit metadata low confidence",
"question": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
"traceFixture": "audit-metadata-low-confid.json",
"expectedRootCauseKeywords": ["超时", "证据"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["LOW_CONFID"],
"forbiddenAnswerKeywords": ["已经确认"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["主库故障"],
"requirePromptAudit": true,
"expectedPromptAuditVersion": "chat-prompts-v1",
"expectedPromptVersions": {
"chat_planner": "chat-planner-v1",
"chat_executor": "chat-executor-v2",
"chat_verifier": "chat-verifier-v2",
"chat_composer": "chat-composer-v1"
},
"requireGatekeeperRules": true
},
{
"id": "composer-fallback-no-raw-json",
"title": "Composer fallback no raw JSON",
"question": "库存服务慢响应是否可以直接输出 Executor JSON?",
"traceFixture": "composer-fallback-no-raw-json-low-confid.json",
"expectedRootCauseKeywords": ["慢响应", "证据"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics"],
"allowedVerdicts": ["LOW_CONFID"],
"forbiddenAnswerKeywords": ["executor_evidence_v2", "answer_version", "claim_id"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedComposerStatuses": ["composer_malformed"],
"forbiddenConfirmedClaimKeywords": ["线程池已经耗尽"]
} }
] ]
@@ -0,0 +1,192 @@
{
"session": {
"sessionId": "eval-audit-metadata-low-confid",
"query": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 45000,
"toolCallCount": 1,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.42,
"critical_fact_count": 2,
"prompt_audit": {
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_planner",
"version": "chat-planner-v1",
"resource": "prompts/chat-planner-prompt.md"
},
{
"name": "chat_executor",
"version": "chat-executor-v2",
"resource": "prompts/chat-executor-prompt.md"
},
{
"name": "chat_verifier",
"version": "chat-verifier-v2",
"resource": "prompts/chat-verifier-prompt.md"
},
{
"name": "chat_composer",
"version": "chat-composer-v1",
"resource": "prompts/chat-composer-prompt.md"
}
]
},
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.invocation",
"description": "source_invocation_id must reference an existing tool invocation",
"enabled": true,
"default_severity": "reject"
},
{
"id": "evidence.excerpt",
"description": "evidence_excerpt must be supported by recorded evidence text",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-timeout",
"tool_name": "query_logs",
"source_invocation_id": 22,
"raw_path": "$.logs[0]",
"matched_text": "order api timeout",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-timeout",
"claim_type": "symptom",
"claim_text": "订单接口出现超时",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_logs",
"source_invocation_id": 22,
"raw_path": "$.logs[0]",
"evidence_excerpt": "order api timeout"
}
]
},
{
"claim_id": "claim-db-primary",
"claim_type": "root_cause",
"claim_text": "数据库主库故障导致订单超时",
"support_level": "weak",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_logs",
"source_invocation_id": 22,
"raw_path": "$.logs[0]",
"evidence_excerpt": "order api timeout"
}
]
}
],
"hypotheses": [],
"recommended_actions": [
{
"action_text": "补充查询数据库主库状态和错误日志",
"reason": "当前只有订单接口超时日志"
}
],
"missing_info": ["数据库主库状态", "数据库错误日志"]
},
"claim_checks": [
{
"claim_id": "claim-timeout",
"claim_text": "订单接口出现超时",
"claim_type": "symptom",
"verification": "direct_observation",
"detail": "日志直接记录 order api timeout",
"evidence_refs": [
{
"source_invocation_id": 22,
"raw_path": "$.logs[0]"
}
]
},
{
"claim_id": "claim-db-primary",
"claim_text": "数据库主库故障导致订单超时",
"claim_type": "root_cause",
"verification": "unsupported",
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
"evidence_refs": [
{
"source_invocation_id": 22,
"raw_path": "$.logs[0]"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "日志显示订单接口超时,但数据库方向证据不足。",
"recommended_actions": [
{
"action_text": "补充查询数据库主库状态和错误日志",
"reason": "当前只有订单接口超时日志"
}
],
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。"
},
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 22,
"sessionId": "eval-audit-metadata-low-confid",
"toolName": "query_logs",
"outputPreview": "order api timeout",
"retrievalDetails": {
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.logs[0]",
"text": "order api timeout"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,142 @@
{
"session": {
"sessionId": "eval-composer-fallback-no-raw-json",
"query": "库存服务慢响应是否可以直接输出 Executor JSON?",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 47000,
"toolCallCount": 1,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.45,
"critical_fact_count": 2,
"gatekeeper_result": {
"status": "pass",
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-slow-response",
"claim_type": "symptom",
"claim_text": "库存服务出现慢响应",
"support_level": "direct",
"evidence_bindings": [
{
"tool_invocation_id": 1,
"tool_name": "query_metrics",
"evidence_excerpt": "inventory p99 latency increased"
}
]
},
{
"claim_id": "claim-thread-pool",
"claim_type": "root_cause",
"claim_text": "线程池已经耗尽",
"support_level": "weak",
"evidence_bindings": [
{
"tool_invocation_id": 1,
"tool_name": "query_metrics",
"evidence_excerpt": "inventory p99 latency increased"
}
]
}
],
"hypotheses": [],
"recommended_actions": [
{
"action_text": "补充查询库存服务线程池指标",
"reason": "当前只有慢响应指标"
}
],
"missing_info": ["线程池队列长度", "活跃线程数"]
},
"claim_checks": [
{
"claim_id": "claim-slow-response",
"claim_text": "库存服务出现慢响应",
"claim_type": "symptom",
"verification": "direct_observation",
"detail": "指标显示 inventory p99 latency increased",
"evidence_refs": [
{
"tool_invocation_id": 1,
"tool_name": "query_metrics"
}
]
},
{
"claim_id": "claim-thread-pool",
"claim_text": "线程池已经耗尽",
"claim_type": "root_cause",
"verification": "external_unknown",
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
"evidence_refs": []
}
],
"facts_checked": [
{
"fact": "库存服务出现慢响应",
"verification": "direct_evidence",
"detail": "指标显示 inventory p99 latency increased",
"evidence_refs": [
{
"tool_invocation_id": 1,
"tool_name": "query_metrics"
}
]
},
{
"fact": "线程池已经耗尽",
"verification": "external_unknown",
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
"evidence_refs": []
}
],
"composer_output": {
"status": "composer_malformed",
"detail": "used safe fallback rendering",
"answer_summary": "指标显示库存服务出现慢响应。",
"recommended_actions": [
{
"action_text": "补充查询库存服务线程池指标",
"reason": "当前只有慢响应指标"
}
],
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。"
},
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-composer-fallback-no-raw-json",
"toolName": "query_metrics",
"outputPreview": "inventory p99 latency increased",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,110 @@
{
"session": {
"sessionId": "eval-gatekeeper-fabricated-invocation",
"query": "支付失败是否能确认由日志中的连接池耗尽导致?",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 39000,
"toolCallCount": 1,
"answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "REJECT",
"groundedness_score": 0.1,
"critical_fact_count": 1,
"gatekeeper_result": {
"status": "fail",
"failed_rules": ["evidence.invocation_ref"],
"warnings": [],
"errors": [
{
"rule_id": "evidence.invocation_ref",
"field": "claims[0].evidence_bindings[0].tool_invocation_id",
"message": "tool_invocation_id does not exist"
}
]
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "root_cause",
"claim_text": "连接池耗尽导致支付失败",
"support_level": "direct",
"evidence_bindings": [
{
"tool_invocation_id": 99,
"tool_name": "query_logs",
"evidence_excerpt": "connection pool exhausted"
}
]
}
],
"hypotheses": [],
"recommended_actions": [
{
"action_text": "重新查询支付服务错误日志",
"reason": "当前 Executor 证据引用无法回溯"
}
],
"missing_info": ["需要有效的日志工具调用记录"]
},
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "连接池耗尽导致支付失败",
"claim_type": "root_cause",
"verification": "unsupported",
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
"evidence_refs": []
}
],
"facts_checked": [
{
"fact": "连接池耗尽导致支付失败",
"verification": "unsupported",
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
"evidence_refs": []
}
],
"composer_output": {
"status": "valid",
"answer_summary": "证据引用校验失败,不能确认根因。",
"recommended_actions": [
{
"action_text": "重新查询支付服务错误日志",
"reason": "当前 Executor 证据引用无法回溯"
}
],
"user_facing_answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。"
},
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-gatekeeper-fabricated-invocation",
"toolName": "query_logs",
"outputPreview": "payment failed without matching connection pool exhaustion entry",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,126 @@
{
"session": {
"sessionId": "eval-hikari-no-evidence-negative-observation",
"query": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 21000,
"toolCallCount": 1,
"answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_logs",
"source_invocation_id": 12,
"raw_path": "$.no_evidence",
"matched_text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "negative_observation",
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_logs",
"source_invocation_id": 12,
"raw_path": "$.no_evidence",
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": [
"仅查询了 application-logs 中 inventory-service HikariCP 相关日志"
]
},
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"claim_type": "negative_observation",
"verification": "direct_observation",
"detail": "$.no_evidence 只支持本次查询未检索到匹配证据。",
"evidence_refs": [
{
"source_invocation_id": 12,
"raw_path": "$.no_evidence"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "本次查询未检索到匹配日志。",
"recommended_actions": [],
"user_facing_answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。"
},
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"source_invocation_ids": [12],
"evidence_level": "no_evidence"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 12,
"sessionId": "eval-hikari-no-evidence-negative-observation",
"toolName": "query_logs",
"outputPreview": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
"retrievalDetails": {
"evidence_status": "no_evidence",
"evidence_refs": [
{
"raw_path": "$.no_evidence",
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,124 @@
{
"session": {
"sessionId": "eval-narrow-highcpu-observation",
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 18000,
"toolCallCount": 1,
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_metrics",
"source_invocation_id": 11,
"raw_path": "$.alerts[0]",
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "observation",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_id": 11,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
},
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"claim_type": "observation",
"verification": "direct_observation",
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
"evidence_refs": [
{
"source_invocation_id": 11,
"raw_path": "$.alerts[0]"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
"recommended_actions": [],
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
},
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"source_invocation_ids": [11],
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 11,
"sessionId": "eval-narrow-highcpu-observation",
"toolName": "query_metrics",
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"retrievalDetails": {
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.alerts[0]",
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,155 @@
{
"session": {
"sessionId": "eval-prompt-gatekeeper-audit-closure",
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 19000,
"toolCallCount": 1,
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"prompt_audit": {
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_planner",
"version": "chat-planner-v1",
"resource": "prompts/chat-planner-prompt.md"
},
{
"name": "chat_executor",
"version": "chat-executor-v2",
"resource": "prompts/chat-executor-prompt.md"
},
{
"name": "chat_verifier",
"version": "chat-verifier-v2",
"resource": "prompts/chat-verifier-prompt.md"
},
{
"name": "chat_composer",
"version": "chat-composer-v1",
"resource": "prompts/chat-composer-prompt.md"
}
]
},
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.invocation",
"description": "source_invocation_id must reference an existing tool invocation",
"enabled": true,
"default_severity": "reject"
},
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_metrics",
"source_invocation_id": 21,
"raw_path": "$.alerts[0]",
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "observation",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_id": 21,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
},
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"claim_type": "observation",
"verification": "direct_observation",
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
"evidence_refs": [
{
"source_invocation_id": 21,
"raw_path": "$.alerts[0]"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
"recommended_actions": [],
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
},
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"source_invocation_ids": [21],
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 21,
"sessionId": "eval-prompt-gatekeeper-audit-closure",
"toolName": "query_metrics",
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"retrievalDetails": {
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.alerts[0]",
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,151 @@
{
"session": {
"sessionId": "eval-unsupported-claim-filtering",
"query": "订单超时是否可以确认由数据库主库故障导致?",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 44000,
"toolCallCount": 1,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.42,
"critical_fact_count": 2,
"gatekeeper_result": {
"status": "pass",
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-timeout",
"claim_type": "symptom",
"claim_text": "订单接口出现超时",
"support_level": "direct",
"evidence_bindings": [
{
"tool_invocation_id": 1,
"tool_name": "query_logs",
"evidence_excerpt": "order api timeout"
}
]
},
{
"claim_id": "claim-db-primary",
"claim_type": "root_cause",
"claim_text": "数据库主库故障导致订单超时",
"support_level": "weak",
"evidence_bindings": [
{
"tool_invocation_id": 1,
"tool_name": "query_logs",
"evidence_excerpt": "order api timeout"
}
]
}
],
"hypotheses": [],
"recommended_actions": [
{
"action_text": "补充查询数据库主库状态和错误日志",
"reason": "当前只有订单接口超时日志"
}
],
"missing_info": ["数据库主库状态", "数据库错误日志"]
},
"claim_checks": [
{
"claim_id": "claim-timeout",
"claim_text": "订单接口出现超时",
"claim_type": "symptom",
"verification": "direct_observation",
"detail": "日志直接记录 order api timeout",
"evidence_refs": [
{
"tool_invocation_id": 1,
"tool_name": "query_logs"
}
]
},
{
"claim_id": "claim-db-primary",
"claim_text": "数据库主库故障导致订单超时",
"claim_type": "root_cause",
"verification": "unsupported",
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
"evidence_refs": [
{
"tool_invocation_id": 1,
"tool_name": "query_logs"
}
]
}
],
"facts_checked": [
{
"fact": "订单接口出现超时",
"verification": "direct_evidence",
"detail": "日志直接记录 order api timeout",
"evidence_refs": [
{
"tool_invocation_id": 1,
"tool_name": "query_logs"
}
]
},
{
"fact": "数据库主库故障导致订单超时",
"verification": "unsupported",
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
"evidence_refs": [
{
"tool_invocation_id": 1,
"tool_name": "query_logs"
}
]
}
],
"composer_output": {
"status": "valid",
"answer_summary": "日志显示订单接口超时,但数据库方向证据不足。",
"recommended_actions": [
{
"action_text": "补充查询数据库主库状态和错误日志",
"reason": "当前只有订单接口超时日志"
}
],
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。"
},
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-unsupported-claim-filtering",
"toolName": "query_logs",
"outputPreview": "order api timeout",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
+171 -7
View File
@@ -1,14 +1,72 @@
{ {
"totalCases" : 5, "totalCases" : 12,
"passedCases" : 5, "passedCases" : 12,
"passRate" : 1.0, "passRate" : 1.0,
"verdictDistribution" : { "verdictDistribution" : {
"PASS" : 2, "PASS" : 5,
"LOW_CONFID" : 3 "LOW_CONFID" : 6,
"REJECT" : 1
}, },
"averageToolCallCount" : 2.0, "averageToolCallCount" : 1.4166666666666667,
"averageDurationMs" : 45800.0, "averageDurationMs" : 38500.0,
"results" : [ { "results" : [ {
"caseId" : "narrow-highcpu-observation",
"title" : "Narrow HighCPU observation",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_metrics" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"promptAuditVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 1,
"gatekeeperRuleCount" : 1,
"toolCallCount" : 1,
"durationMs" : 18000
}, {
"caseId" : "prompt-gatekeeper-audit-closure",
"title" : "Prompt and Gatekeeper audit closure",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_metrics" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"promptAuditVersion" : "chat-prompts-v1",
"composerStatus" : "valid",
"claimCheckCount" : 1,
"gatekeeperRuleCount" : 2,
"toolCallCount" : 1,
"durationMs" : 19000
}, {
"caseId" : "hikari-no-evidence-negative-observation",
"title" : "Hikari no-evidence negative observation",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_logs" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"promptAuditVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 1,
"gatekeeperRuleCount" : 1,
"toolCallCount" : 1,
"durationMs" : 21000
}, {
"caseId" : "payment-timeout", "caseId" : "payment-timeout",
"title" : "Payment API timeout", "title" : "Payment API timeout",
"passed" : true, "passed" : true,
@@ -21,6 +79,12 @@
"query_logs" : true, "query_logs" : true,
"query_metrics" : true "query_metrics" : true
}, },
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 3, "toolCallCount" : 3,
"durationMs" : 42000 "durationMs" : 42000
}, { }, {
@@ -35,6 +99,12 @@
"lookup_knowledge" : true, "lookup_knowledge" : true,
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 2, "toolCallCount" : 2,
"durationMs" : 51000 "durationMs" : 51000
}, { }, {
@@ -48,6 +118,12 @@
"evidenceCoverage" : { "evidenceCoverage" : {
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 1, "toolCallCount" : 1,
"durationMs" : 36000 "durationMs" : 36000
}, { }, {
@@ -62,6 +138,12 @@
"query_metrics" : true, "query_metrics" : true,
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 2, "toolCallCount" : 2,
"durationMs" : 47000 "durationMs" : 47000
}, { }, {
@@ -76,7 +158,89 @@
"query_metrics" : true, "query_metrics" : true,
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 2, "toolCallCount" : 2,
"durationMs" : 53000 "durationMs" : 53000
}, {
"caseId" : "gatekeeper-fabricated-invocation",
"title" : "Gatekeeper fabricated invocation",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "REJECT",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_logs" : true
},
"gatekeeperStatus" : "fail",
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 1,
"gatekeeperRuleCount" : null,
"toolCallCount" : 1,
"durationMs" : 39000
}, {
"caseId" : "unsupported-claim-filtering",
"title" : "Unsupported claim filtering",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_logs" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 2,
"gatekeeperRuleCount" : null,
"toolCallCount" : 1,
"durationMs" : 44000
}, {
"caseId" : "audit-metadata-low-confid",
"title" : "Audit metadata low confidence",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_logs" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"promptAuditVersion" : "chat-prompts-v1",
"composerStatus" : "valid",
"claimCheckCount" : 2,
"gatekeeperRuleCount" : 2,
"toolCallCount" : 1,
"durationMs" : 45000
}, {
"caseId" : "composer-fallback-no-raw-json",
"title" : "Composer fallback no raw JSON",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_metrics" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : "composer_malformed",
"claimCheckCount" : 2,
"gatekeeperRuleCount" : null,
"toolCallCount" : 1,
"durationMs" : 47000
} ] } ]
} }
+21 -13
View File
@@ -1,22 +1,30 @@
# Diagnosis Eval Report # Diagnosis Eval Report
- Total cases: 5 - Total cases: 12
- Passed cases: 5 - Passed cases: 12
- Pass rate: 100.00% - Pass rate: 100.00%
- Average tool calls: 2.00 - Average tool calls: 1.42
- Average duration ms: 45800.00 - Average duration ms: 38500.00
## Verdict Distribution ## Verdict Distribution
- PASS: 2 - PASS: 5
- LOW_CONFID: 3 - LOW_CONFID: 6
- REJECT: 1
## Cases ## Cases
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks | | Case | Result | Verdict | Gatekeeper | Rule Set | Prompt Audit | Composer | Claim Checks | Rules | Keywords | Tool Calls | Duration ms | Failed Checks |
| --- | --- | --- | --- | ---: | ---: | --- | | --- | --- | --- | --- | --- | --- | --- | ---: | ---: | --- | ---: | ---: | --- |
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - | | narrow-highcpu-observation | PASS | PASS | pass | gatekeeper-rules-v1 | - | valid | 1 | 1 | 3/3 | 1 | 18000 | - |
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - | | prompt-gatekeeper-audit-closure | PASS | PASS | pass | gatekeeper-rules-v1 | chat-prompts-v1 | valid | 1 | 2 | 3/3 | 1 | 19000 | - |
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - | | hikari-no-evidence-negative-observation | PASS | PASS | pass | gatekeeper-rules-v1 | - | valid | 1 | 1 | 3/3 | 1 | 21000 | - |
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - | | payment-timeout | PASS | PASS | - | - | - | - | - | - | 3/3 | 3 | 42000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - | | mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | - | - | - | 3/3 | 2 | 51000 | - |
| redis-timeout | PASS | LOW_CONFID | - | - | - | - | - | - | 2/2 | 1 | 36000 | - |
| slow-response | PASS | PASS | - | - | - | - | - | - | 2/2 | 2 | 47000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | - | - | - | 3/3 | 2 | 53000 | - |
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | - | - | valid | 1 | - | 3/3 | 1 | 39000 | - |
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | - | - | valid | 2 | - | 2/2 | 1 | 44000 | - |
| audit-metadata-low-confid | PASS | LOW_CONFID | pass | gatekeeper-rules-v1 | chat-prompts-v1 | valid | 2 | 2 | 2/2 | 1 | 45000 | - |
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | - | - | composer_malformed | 2 | - | 2/2 | 1 | 47000 | - |
+121 -146
View File
@@ -1,202 +1,177 @@
# Diagnosis Eval Data Schema # Diagnosis Eval Data Schema
这份文档记录评测基准里的数据结构。口语化理解就是: 这份文档记录 `mvp/eval` 固定评测集的数据结构。评测器读取保存好的 trace fixture,用确定性规则判断这次 Agent 运行是否满足预期。
```text ```text
用例文件说“我要考什么” case 文件:我要考什么
trace 文件说“Agent 实际做了什么” fixture 文件:Agent 实际做了什么
评测结果说“这次有没有跑偏” 评测结果:这条 case 是否通过,哪里失败
汇总报告说“整体稳定性怎么样” baseline report:整套固定集当前认可的结果
``` ```
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。 当前评测不调用 LLM 打分。
## 1. 用例定义 ## 1. Case 定义
文件:`mvp/eval/cases/diagnosis-cases.json` 文件:`mvp/eval/cases/diagnosis-cases.json`
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。 每条 case 定义一个固定诊断场景。
```json ```json
{ {
"id": "payment-timeout", "id": "unsupported-claim-filtering",
"title": "Payment API timeout", "title": "Unsupported claim filtering",
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。", "question": "订单超时是否可以确认由数据库主库故障导致?",
"traceFixture": "payment-timeout-pass.json", "traceFixture": "unsupported-claim-filtering-low-confid.json",
"expectedRootCauseKeywords": ["支付", "超时", "连接池"], "expectedRootCauseKeywords": ["超时", "证据"],
"minKeywordMatches": 2, "minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"], "requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["PASS", "LOW_CONFID"], "allowedVerdicts": ["LOW_CONFID"],
"forbiddenAnswerKeywords": ["无证据确定"] "forbiddenAnswerKeywords": ["已经确认"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["主库故障"],
"requirePromptAudit": true,
"expectedPromptAuditVersion": "chat-prompts-v1",
"expectedPromptVersions": {
"chat_planner": "chat-planner-v1",
"chat_executor": "chat-executor-v2",
"chat_verifier": "chat-verifier-v2",
"chat_composer": "chat-composer-v1"
},
"requireGatekeeperRules": true
} }
``` ```
字段说明: 字段说明:
| 字段 | 意思 | 评测器怎么用 | | 字段 | 含义 | 评测器怎么用 |
| --- | --- | --- | | --- | --- | --- |
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 | | `id` | case 唯一标识 | 出现在报告中 |
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 | | `title` | 可读标题 | 出现在报告中 |
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 | | `question` | 原始用户问题 | 用于说明场景,fixture 模式不会真实发送给 Agent |
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 | | `traceFixture` | 对应 fixture 文件名 | 从 `mvp/eval/fixtures` 加载 |
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 | | `expectedRootCauseKeywords` | 最终答案应覆盖的关键词 | 在 `session.answer` 中做包含判断 |
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 | | `minKeywordMatches` | 最少命中关键词数 | 低于该值则失败 |
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 | | `requiredEvidenceTools` | 必须出现的证据工具 | 从 `toolInvocations` 和 `tool_trace_summary` 中收集 |
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 | | `allowedVerdicts` | 允许的 Verifier verdict | verdict 不在列表中则失败 |
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 | | `forbiddenAnswerKeywords` | 最终答案禁止出现的词 | 用于拦截过度自信或危险表达 |
| `requireV2AuditClosure` | 是否要求 V2 审计闭环字段 | 要求 `gatekeeper_result`、`claim_checks`、`composer_output` 存在,并检查 raw JSON 泄漏 |
| `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 |
| `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` |
| `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 |
| `expectedGatekeeperRuleSetVersion` | 期望的 Gatekeeper 规则集版本 | 配置后校验 `gatekeeper_result.rule_set_version` |
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
| `requirePromptAudit` | 是否要求 Prompt 审计元数据 | 要求 `prompt_audit.version` 存在 |
| `expectedPromptAuditVersion` | 期望的 Prompt 审计目录版本 | 配置后校验 `prompt_audit.version` |
| `expectedPromptVersions` | 期望的各角色 Prompt 版本 | 校验 `prompt_audit.prompts[*].name/version` |
| `requireGatekeeperRules` | 是否要求 Gatekeeper 规则元数据 | 要求 `gatekeeper_result.rules` 非空,且每条规则有 `id`、`enabled`、`default_severity` |
## 2. Trace Fixture ## 2. Trace Fixture
目录:`mvp/eval/fixtures/*.json` 目录:`mvp/eval/fixtures/*.json`
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。 fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规则需要的字段。
当前会读取这些字段: | Trace 字段 | 含义 | 评测器怎么用 |
| Trace 字段 | 意思 | 评测器怎么用 |
| --- | --- | --- | | --- | --- | --- |
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 | | `session.answer` | 最终用户答案 | 检查关键词、禁用词、unsupported claim 泄漏、raw JSON 泄漏 |
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 | | `session.totalDurationMs` | 运行耗时 | 进入报告 |
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 | | `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 | | `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 | | `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | Gatekeeper 规则集版本 | 新矩阵 case 可显式断言该版本 |
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 | | `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rules` | Gatekeeper 规则元数据摘要 | 审计 case 可要求规则列表非空且字段完整 |
| `session.selfEvaluation.verifier_evaluation.prompt_audit.version` | Chat Prompt 审计目录版本 | 审计 case 可显式断言该版本 |
简单说,trace 里最重要的是三类信息: | `session.selfEvaluation.verifier_evaluation.prompt_audit.prompts` | 各 Chat Prompt 名称、版本和资源路径 | 审计 case 可断言 planner、executor、verifier、composer 版本 |
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
```text | `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
最终回答:它说了什么 | `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
工具证据:它查了什么 | `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 看到的工具证据 | 用于补充证据工具覆盖 |
Verifier:它自己有没有承认这个结论可靠 | `toolInvocations[*].toolName` | 实际调用工具名 | 用于检查 `requiredEvidenceTools` |
``` | `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 中 |
## 3. 单条评测结果 ## 3. 单条评测结果
Java 类型:`DiagnosisEvalResult` Java 类型:`DiagnosisEvalResult`
这是每条 case 跑完之后的判断结果。 | 字段 | 含义 |
| 字段 | 意思 |
| --- | --- | | --- | --- |
| `caseId` | 对应的 case id | | `caseId` | 对应 case id |
| `title` | case 标题 | | `title` | case 标题 |
| `passed` | 这条 case 是否通过 | | `passed` | 该 case 是否通过 |
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 | | `failedChecks` | 失败原因列表 |
| `verdict` | 从 trace 里读出来的 Verifier verdict | | `verdict` | 从 trace 中读到的 Verifier verdict |
| `matchedKeywordCount` | 最终回答命中的关键词数量 | | `matchedKeywordCount` | 最终答案命中的关键词数量 |
| `requiredKeywordCount` | case 定义里一共有多少个关键词 | | `requiredKeywordCount` | case 配置的关键词数量 |
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` | | `evidenceCoverage` | 每个必需工具是否出现 |
| `toolCallCount` | 本次 trace 里工具调用总数 | | `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
| `durationMs` | 本次 trace 的耗时 | | `gatekeeperRuleSetVersion` | 读到的 `gatekeeper_result.rule_set_version` |
| `promptAuditVersion` | 读到的 `prompt_audit.version` |
判断通过的口语化规则: | `composerStatus` | 读到的 `composer_output.status` |
| `claimCheckCount` | `claim_checks` 数量 |
```text | `gatekeeperRuleCount` | `gatekeeper_result.rules` 数量 |
回答要说到关键点 | `toolCallCount` | trace 中工具调用总数 |
该查的证据工具要查到 | `durationMs` | trace 总耗时 |
Verifier 的结论要在可接受范围内
回答不能出现危险的过度自信表达
如果是 REJECT,就必须走降级模板
```
## 4. 汇总报告 ## 4. 汇总报告
Java 类型:`DiagnosisEvalReport` Java 类型:`DiagnosisEvalReport`
这是整个基准集跑完之后的总结果。 | 字段 | 含义 |
| 字段 | 意思 |
| --- | --- | | --- | --- |
| `totalCases` | 总共评测了多少条 case | | `totalCases` | case 总数 |
| `passedCases` | 通过了多少条 | | `passedCases` | 通过数 |
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` | | `passRate` | 通过率,范围 `0.0` 到 `1.0` |
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` | | `verdictDistribution` | Verifier verdict 分布 |
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 | | `averageToolCallCount` | 平均工具调用数 |
| `averageDurationMs` | 平均耗时 | | `averageDurationMs` | 平均耗时 |
| `results` | 每条 case 的详细结果列表 | | `results` | 单条 case 结果列表 |
## 5. 怎么看这个基准 ## 5. Stage 5 V2 审计闭环规则
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答: 阶段 5 关注的是“前四段链路是否能被固定评测证明”:
```text ```text
以前能过的诊断题,现在还过不过? Executor structured output
它是不是少查了某些证据? -> Gatekeeper deterministic audit
它是不是变得更自信但证据不足? -> Verifier claim_checks
它是不是开始输出不该说的话? -> Composer filtered final answer
它是不是明显变慢了?
``` ```
所以面试里可以这样讲: 新增确定性规则:
```text - V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。 - `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。 - 配置 `expectedGatekeeperRuleSetVersion` 的 case 必须匹配 `gatekeeper_result.rule_set_version`。
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。 - 配置 `requirePromptAudit` 的 case 必须包含 `prompt_audit.version`。
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。 - 配置 `expectedPromptAuditVersion` 的 case 必须匹配 `prompt_audit.version`。
``` - 配置 `expectedPromptVersions` 的 case 必须能在 `prompt_audit.prompts` 中找到对应角色和版本。
- 配置 `requireGatekeeperRules` 的 case 必须包含非空 `gatekeeper_result.rules`,且每条规则有 `id`、`enabled`、`default_severity`。
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
- Composer 输出必须记录 `status`。
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
- case 配置的 `forbiddenConfirmedClaimKeywords` 不能出现在最终答案里。
## 6. Baseline Diff ## 6. Baseline Diff
Baseline diff 是拿两份 report 做对比: Baseline diff 比较两份 report:
```text ```text
baseline report:以前认可的基准结果 baseline report:已经认可的基准结果
current report:这次改动后跑出来的新结果 current report:当前代码/fixture 跑出的结果
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了 diff report:结构化列出退化、改善和普通变化
``` ```
Java 类型: 主要退化信号:
- `DiagnosisEvalDiffReport` - pass rate 下降。
- `DiagnosisEvalDiffItem` - case 从通过变失败。
- 必需证据工具从有变无。
`DiagnosisEvalDiffReport` 字段: - 关键词命中减少。
- 工具调用或耗时明显上升。
| 字段 | 意思 | - verdict 分布变化。
| --- | --- |
| `baselineTotalCases` | baseline 里有多少条 case |
| `currentTotalCases` | current 里有多少条 case |
| `baselinePassedCases` | baseline 通过了多少条 |
| `currentPassedCases` | current 通过了多少条 |
| `baselinePassRate` | baseline 通过率 |
| `currentPassRate` | current 通过率 |
| `regressionCount` | 退化项数量 |
| `improvementCount` | 改善项数量 |
| `changedCount` | 普通变化项数量 |
| `hasRegression` | 是否存在退化 |
| `items` | 具体 diff 明细 |
`DiagnosisEvalDiffItem` 字段:
| 字段 | 意思 |
| --- | --- |
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
| `caseId` | 如果是单条 case 变化,这里记录 case id |
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
| `baselineValue` | baseline 里的值 |
| `currentValue` | current 里的值 |
| `delta` | 数值变化量;非数值变化为空 |
| `message` | 给人看的变化说明 |
口语化判断规则:
```text
pass rate 下降:退化
case 从通过变失败:退化
证据工具从有变没有:退化
关键词命中变少:退化
工具调用或耗时升高:成本上升,记为退化信号
verdict 分布变化:记录变化,供人工判断是否符合预期
```
面试里可以这样讲:
```text
我把 baseline report 和当前 report 做结构化 diff。
它不是再问 LLM,而是用代码比较固定字段。
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
diff 会直接标成 regression。
这样 Agent 改动可以用固定基准做回归判断。
```
+56 -35
View File
@@ -1,44 +1,65 @@
# 已知问题记录 # MVP Issues 索引
| # | 标题 | 严重程度 | 状态 | 文件 | **更新日期**:2026-07-10
|---|---|---|---|---| **状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
| ISS-001 | Executor 重复召回同一文档 | 中 | 已修复 | [ISS-001-duplicate-retrieval.md](ISS-001-duplicate-retrieval.md) |
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) |
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) |
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 已归档 | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
## RAG 重构计划 ## 目录约定
| 目录 | 用途 |
|---|---|
| [active/](active/) | 仍需要规划或实现的问题 |
| [design-notes/](design-notes/) | 已形成方向、用于指导后续实现的设计记录 |
| [rag/](rag/) | RAG 子问题集合;多数已合并到 RAG 重构计划 |
| [archived/](archived/) | 已修复、已实施或已归档的问题 |
## 活跃问题
| 名称 | 标题 | 严重程度 | 状态 | 文件 | | 名称 | 标题 | 严重程度 | 状态 | 文件 |
|---|---|---|---|---| |---|---|---|---|---|
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [rag-refactor-plan.md](rag-refactor-plan.md) | | ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
## RAG 检索问题 ## 设计笔记
| 名称 | 标题 | 严重程度 | 状态 | 文件 | | 名称 | 标题 | 状态 | 文件 |
|---|---|---|---|---| |---|---|---|---|
| chunk-context-reconstruction | RAG 切片上下文重建缺失 | 高 | 已合并到重构计划 | [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md) | | executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 已形成方向 | [design-notes/executor-self-evidence-loop-design-note.md](design-notes/executor-self-evidence-loop-design-note.md) |
| breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 高 | 已合并到重构计划 | [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md) | | executor-structured-output-v2 | Executor 结构化输出 V2 阶段设计 | 部分已实施,保留为后续改造参考 | [design-notes/executor-structured-output-v2.md](design-notes/executor-structured-output-v2.md) |
| l0-l1-fusion-ranking | RAG L0 和 L1 未真正融合排序 | 中 | 已合并到重构计划 | [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md) |
| l0-keyword-matching-quality | RAG L0 关键词匹配质量不足 | 中 | 已合并到重构计划 | [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md) |
| l1-score-calibration | RAG L1 分数阈值未校准 | 中 | 已合并到重构计划 | [rag-l1-score-calibration.md](rag-l1-score-calibration.md) |
| context-packing-and-reranking | RAG 缺少上下文打包和 Rerank | 中 | 已合并到重构计划 | [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md) |
| upload-chunk-parameter-drift | RAG 上传切片参数未真正生效 | 低 | 已合并到重构计划 | [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md) |
| query-rewrite-gap | RAG 查询改写能力薄弱 | 中 | 已合并到重构计划 | [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md) |
## RAG 框架化改造 ## RAG 问题集
| 名称 | 标题 | 严重程度 | 状态 | 文件 | 这些问题已经收敛到 [active/rag-refactor-plan.md](active/rag-refactor-plan.md),单个文件保留用于追溯原始问题和设计背景。
|---|---|---|---|---|
| spring-ai-vectorstore-migration | RAG 迁移到 Spring AI VectorStore 检索抽象 | 高 | 已合并到重构计划 | [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md) | | 名称 | 标题 | 状态 | 文件 |
| spring-ai-query-transformer | RAG 接入 Spring AI Query Transformer | 中 | 已合并到重构计划 | [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md) | |---|---|---|---|
| spring-ai-document-postprocessor | RAG 使用 DocumentPostProcessor 做后处理 | 中 | 已合并到重构计划 | [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md) | | breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 已合并到重构计划 | [rag/rag-breadcrumb-embedding-gap.md](rag/rag-breadcrumb-embedding-gap.md) |
| l0-domain-entity-hint | RAG 将 L0 降级为领域和实体 Hint | 中 | 已合并到重构计划 | [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md) | | chunk-context-reconstruction | RAG 切片上下文重建缺失 | 已合并到重构计划 | [rag/rag-chunk-context-reconstruction.md](rag/rag-chunk-context-reconstruction.md) |
| spring-ai-advisor-boundary | RAG 明确 Spring AI Advisor 与 Agent Tool 的边界 | 中 | 已合并到重构计划 | [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md) | | context-packing-and-reranking | RAG 缺少上下文打包和 Rerank | 已合并到重构计划 | [rag/rag-context-packing-and-reranking.md](rag/rag-context-packing-and-reranking.md) |
| l0-domain-entity-hint | RAG 将 L0 降级为领域和实体 Hint | 已合并到重构计划 | [rag/rag-l0-domain-entity-hint.md](rag/rag-l0-domain-entity-hint.md) |
| l0-keyword-matching-quality | RAG L0 关键词匹配质量不足 | 已合并到重构计划 | [rag/rag-l0-keyword-matching-quality.md](rag/rag-l0-keyword-matching-quality.md) |
| l0-l1-fusion-ranking | RAG L0 和 L1 未真正融合排序 | 已合并到重构计划 | [rag/rag-l0-l1-fusion-ranking.md](rag/rag-l0-l1-fusion-ranking.md) |
| l1-score-calibration | RAG L1 分数阈值未校准 | 已合并到重构计划 | [rag/rag-l1-score-calibration.md](rag/rag-l1-score-calibration.md) |
| query-rewrite-gap | RAG 查询改写能力薄弱 | 已合并到重构计划 | [rag/rag-query-rewrite-gap.md](rag/rag-query-rewrite-gap.md) |
| spring-ai-advisor-boundary | Spring AI Advisor 与 Agent Tool 边界 | 已合并到重构计划 | [rag/rag-spring-ai-advisor-boundary.md](rag/rag-spring-ai-advisor-boundary.md) |
| spring-ai-document-postprocessor | 使用 DocumentPostProcessor 做后处理 | 已合并到重构计划 | [rag/rag-spring-ai-document-postprocessor.md](rag/rag-spring-ai-document-postprocessor.md) |
| spring-ai-query-transformer | 接入 Spring AI Query Transformer | 已合并到重构计划 | [rag/rag-spring-ai-query-transformer.md](rag/rag-spring-ai-query-transformer.md) |
| spring-ai-vectorstore-migration | 迁移到 Spring AI VectorStore 检索抽象 | 已合并到重构计划 | [rag/rag-spring-ai-vectorstore-migration.md](rag/rag-spring-ai-vectorstore-migration.md) |
| upload-chunk-parameter-drift | 上传切片参数未真正生效 | 已合并到重构计划 | [rag/rag-upload-chunk-parameter-drift.md](rag/rag-upload-chunk-parameter-drift.md) |
## 已归档问题
| 名称 | 标题 | 状态 | 文件 |
|---|---|---|---|
| ISS-001 | Executor 重复召回同一文档 | 已修复 | [archived/ISS-001-duplicate-retrieval.md](archived/ISS-001-duplicate-retrieval.md) |
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 已修复 | [archived/ISS-002-executor-unconstrained-lookup.md](archived/ISS-002-executor-unconstrained-lookup.md) |
| ISS-005 | 证据链补齐与降级契约收敛 | 已归档 | [archived/ISS-005-evidence-trace-hardening.md](archived/ISS-005-evidence-trace-hardening.md) |
| ISS-006 | 固定诊断评测集与回归 Harness | 已归档 | [archived/ISS-006-diagnosis-eval-harness.md](archived/ISS-006-diagnosis-eval-harness.md) |
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 已实施 | [archived/ISS-007-verifier-evidence-summary-fidelity.md](archived/ISS-007-verifier-evidence-summary-fidelity.md) |
| ISS-008 | Executor 窄范围查询越界 | 已修复 | [archived/ISS-008-executor-narrow-scope-overreach.md](archived/ISS-008-executor-narrow-scope-overreach.md) |
| ISS-009 | negative_observation 精确引用 no-evidence 结果 | 已修复 | [archived/ISS-009-negative-observation-no-evidence-reference.md](archived/ISS-009-negative-observation-no-evidence-reference.md) |
| ISS-010 | 同 session 多轮诊断 Trace 隔离 | 已归档 | [archived/ISS-010-session-run-trace-isolation.md](archived/ISS-010-session-run-trace-isolation.md) |
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 已归档 | [archived/diagnosis-eval-baseline-diff.md](archived/diagnosis-eval-baseline-diff.md) |
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 已归档 | [archived/expand-diagnosis-eval-fixtures.md](archived/expand-diagnosis-eval-fixtures.md) |
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 已归档 | [archived/mvp-demo-interview-runbook.md](archived/mvp-demo-interview-runbook.md) |
@@ -347,19 +347,19 @@ RAG、Agent、AIOps、数据库记录互相关联,必须分阶段推进,每
本计划合并以下问题和改造方向: 本计划合并以下问题和改造方向:
- [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md) - [rag-chunk-context-reconstruction.md](../rag/rag-chunk-context-reconstruction.md)
- [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md) - [rag-breadcrumb-embedding-gap.md](../rag/rag-breadcrumb-embedding-gap.md)
- [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md) - [rag-l0-l1-fusion-ranking.md](../rag/rag-l0-l1-fusion-ranking.md)
- [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md) - [rag-l0-keyword-matching-quality.md](../rag/rag-l0-keyword-matching-quality.md)
- [rag-l1-score-calibration.md](rag-l1-score-calibration.md) - [rag-l1-score-calibration.md](../rag/rag-l1-score-calibration.md)
- [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md) - [rag-context-packing-and-reranking.md](../rag/rag-context-packing-and-reranking.md)
- [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md) - [rag-upload-chunk-parameter-drift.md](../rag/rag-upload-chunk-parameter-drift.md)
- [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md) - [rag-query-rewrite-gap.md](../rag/rag-query-rewrite-gap.md)
- [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md) - [rag-spring-ai-vectorstore-migration.md](../rag/rag-spring-ai-vectorstore-migration.md)
- [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md) - [rag-spring-ai-query-transformer.md](../rag/rag-spring-ai-query-transformer.md)
- [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md) - [rag-spring-ai-document-postprocessor.md](../rag/rag-spring-ai-document-postprocessor.md)
- [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md) - [rag-l0-domain-entity-hint.md](../rag/rag-l0-domain-entity-hint.md)
- [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md) - [rag-spring-ai-advisor-boundary.md](../rag/rag-spring-ai-advisor-boundary.md)
--- ---
@@ -4,7 +4,7 @@
**严重程度**:中(影响 token 消耗和上下文质量,不影响功能正确性) **严重程度**:中(影响 token 消耗和上下文质量,不影响功能正确性)
**发现时间**:2026-06-30 **发现时间**:2026-06-30
**修复版本**:session-dedup-knowledge-map **修复版本**:session-dedup-knowledge-map
**历史架构文档**:[会话级去重与知识域地图](../architecture/archive/2026-07-05-legacy/session-dedup-knowledge-map.md) **历史架构文档**:[会话级去重与知识域地图](../../architecture/archive/2026-07-05-legacy/session-dedup-knowledge-map.md)
--- ---
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,308 @@
# ISS-008 Executor 窄范围查询越界
**严重程度**:中
**状态**:已修复
**发现时间**:2026-07-08
**关联**:
- `ISS-007-verifier-evidence-summary-fidelity`
- `executor-structured-output-v2`
- `executor-evidence-attribution-hallucination`
---
## 背景
当前 Chat 诊断链路已经演进为:
```text
Planner
-> Executor
-> VerifierInputHook / Gatekeeper
-> Verifier
-> Composer
```
其中 Executor 的定位已经从“生成最终诊断答案”收敛为:
```text
证据收集 + 微观事实提炼
```
但在窄范围问题中,Executor 仍可能把用户只要求确认的一件事扩展成多条 claim,例如用户只问 `HighCPUUsage`,Executor 可能顺手输出内存、连接池、数据库或修复建议相关内容。
这类问题不一定是证据伪造。很多时候工具返回里确实有其它信息,但它们不属于当前用户问题的范围。Gatekeeper 只能校验证据引用真假,不能完整承担“用户意图范围控制”;Verifier 虽然可以降级,但会增加链路负担。
因此本 issue 采用低成本的 Prompt-first 修复:先收紧 Executor prompt,不改 Planner,不引入 `scope_contract`。
---
## 问题类型
### 1. 窄范围查询越界
用户问题只要求确认一个服务、告警、日志、订单或时间窗口,但 Executor 输出了用户未要求的 claim。
示例:
```text
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单、OOM、数据库慢查询、连接池或 user-service。
```
错误输出包括:
- `HighMemoryUsage`
- `SlowResponse`
- `order-123`
- `HikariCP`
- `DB / database`
- `user-service`
### 2. Observation 变成 Diagnosis
Executor 本应输出观察事实,却输出根因、风险、修复建议或经验推断。
错误输出包括:
- “CPU 过高是请求超时的根因”
- “建议扩容”
- “通常这种情况是数据库慢查询导致”
- “存在内存泄漏风险”
### 3. Runbook 通用知识变成当前事实
Runbook、Skill、知识库可以指导要查什么,但不能直接变成本次环境已发生的事实。
错误输出包括:
```text
Runbook 中说 HighCPUUsage 常见原因是流量突增,所以当前环境发生了流量突增。
```
---
## 修复决策
本期只修 Executor prompt。
### 本期做
1. 强化 Executor 单一职责:证据收集 + 微观事实提炼。
2. 增加 `角色边界 HARD-GATE`。
3. 增加 `窄范围确认任务 HARD-GATE`。
4. 增加工具使用边界,避免为补全故事而扩展检索。
5. 增加输出前自检,要求输出 JSON 前删除越界 claim。
### 本期不做
1. 不改 Planner。
2. 不新增 `scope_contract`。
3. 不解析 Planner 输出中的 scope。
4. 不做 Gatekeeper scope 校验。
5. 不改多 Agent 编排。
---
## 设计原则
### Claim 要少,Evidence 可以多
窄范围任务下,Executor 应输出最少必要 claim,通常 1 条,最多 2 条。
但 claim 数量限制不限制 `evidence_bindings` 数量。一条核心 claim 可以绑定多条直接相关证据。
```text
正确:
1 条 claim + 多条 evidence_bindings
错误:
为了展示多条证据,把同一个观察事实拆成多条 claim
```
### 只输出当前问题范围内的 Observation
窄范围任务下,`claims` 只能使用:
- `observation`
- `negative_observation`
禁止使用:
- `root_cause`
- `risk`
- `recommendation`
- 其它建议类或诊断类 claim
### 证据不足时不要补故事
如果工具没有返回可被精确引用的证据:
```text
source_invocation_id + raw_path + evidence_excerpt
```
Executor 不应生成 confirmed claim,应写入 `missing_info`。
---
## Prompt 修复点
已更新:
- `src/main/resources/prompts/chat-executor-prompt.md`
核心新增约束:
1. `角色边界 HARD-GATE`
2. `窄范围确认任务 HARD-GATE`
3. `工具使用边界`
4. `输出前自检`
5. 条目级 `raw_path` 强约束:同一条工具数组项只能绑定一次,禁止输出 `$.alerts[0].alert_name`、`$.alerts[0].state` 等字段级子路径。
---
## 验证结果
### 2026-07-08 E2E 验证
输入:
```text
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。
```
第一次验证发现:
- Executor 已经只输出 `payment-service + HighCPUUsage` 相关 observation,没有输出越界 claim。
- 但 Executor 额外生成了字段级 `raw_path`:
- `$.alerts[0].alert_name`
- `$.alerts[0].state`
- 当前 Gatekeeper 只支持条目级路径 `$.alerts[i]` / `$.logs[i]` / `$.evidence_blocks[i]`,因此判定为 `REJECT`。
已追加 prompt 约束:
```text
同一条工具数组项只能绑定一次。
不要为了引用其中多个字段而拆成多个 evidence_bindings。
raw_path 禁止指向字段级子路径。
```
第二次验证结果:
```text
sessionId: iss008-narrow-highcpu-rerun-20260708-215510
verdict: PASS
groundedness_score: 1.0
gatekeeper_result.status: pass
gatekeeper_result.severity: none
claim_count: 1
claim_type: observation
forbidden_hits: none
```
Executor claim:
```text
payment-service 当前存在 HighCPUUsage 告警,CPU 使用率持续超过 80%,当前值为 92%,告警状态为 firing,已持续 25 分钟。
```
最终答案未出现以下排除项:
- `HighMemoryUsage`
- `SlowResponse`
- `order-123`
- `订单123`
- `OOM`
- `DB / database`
- `HikariCP`
- `connection pool / 连接池`
- `user-service`
---
## 验收标准
### 1. HighCPUUsage 窄范围
输入:
```text
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。
```
期望:
- `claims` 只围绕 `payment-service + HighCPUUsage`。
- 不出现 `HighMemoryUsage`。
- 不出现 `SlowResponse`。
- 不出现 `order-123`。
- 不出现 `OOM`。
- 不出现 `DB / database`。
- 不出现 `HikariCP / connection pool`。
- 不出现 `user-service`。
### 2. HighMemoryUsage 窄范围
输入:
```text
只确认 order-service 是否存在 HighMemoryUsage。
```
期望:
- 可以输出内存使用率、告警状态、持续时间等观察事实。
- 不输出“内存泄漏已确认”。
- 不输出扩容、重启、修改 JVM 参数等修复建议。
### 3. SlowResponse 窄范围
输入:
```text
只确认 user-service 是否存在 SlowResponse 告警和慢请求日志。
```
期望:
- 可以绑定 alert 和 logs 多条证据。
- 不推断数据库连接池耗尽。
- 不推断下游服务故障。
### 4. 用户明确排除项
输入:
```text
只看 order-service 支付失败日志,不要分析 HikariCP。
```
期望:
- `claim_text` 不出现 HikariCP 确认结论。
- 最终答案不出现 HikariCP 确认结论。
### 5. 证据不足
输入:
```text
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
```
期望:
- 如果工具返回 `logs=[]`,Executor 不编造 positive claim。
- 输出 `negative_observation` 或 `missing_info`。
- 不返回 `generic-service` 占位事实。
---
## 后续增强
如果 Prompt-first 后仍不稳定,再考虑:
1. Planner 输出 `scope_contract`。
2. Gatekeeper 增加 scope 校验。
3. eval fixture 增加 forbidden claim 自动断言。
本期暂不进入这些改造。
@@ -0,0 +1,239 @@
# ISS-009 negative_observation 精确引用 no-evidence 结果
**严重程度**:中
**状态**:已修复
**发现时间**:2026-07-08
**关联**:
- `ISS-007-verifier-evidence-summary-fidelity`
- `ISS-008-executor-narrow-scope-overreach`
- `executor-structured-output-v2`
---
## 背景
ISS-007 已经把正向证据引用收敛为:
```text
source_invocation_id + raw_path + evidence_excerpt
```
Gatekeeper 通过 `tool_invocation.retrieval_details.evidence_refs` 校验 Executor 引用是否真实存在。
但负向观察存在一个缺口:当工具明确返回“没查到”时,结果通常是空数组:
```json
{
"logs": [],
"total": 0,
"message": "未找到匹配的日志"
}
```
这时没有 `$.logs[0]`、`$.alerts[0]` 或 `$.evidence_blocks[0]` 可以引用。Executor 如果输出 `negative_observation`,Gatekeeper 无法稳定验证它引用的“无证据结果”,容易降级为 `LOW_CONFID` 或 `REJECT`。
---
## 问题
用户问:
```text
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
```
工具返回:
```json
{
"success": false,
"logs": [],
"total": 0,
"message": "未找到匹配的日志"
}
```
合理 claim 是:
```json
{
"claim_type": "negative_observation",
"claim_text": "未检索到 inventory-service 的 HikariCP 连接池耗尽日志。"
}
```
但旧设计只支持正向数组项:
```text
$.alerts[i]
$.logs[i]
$.evidence_blocks[i]
```
因此负向观察缺少可回溯的精确引用点。
---
## 修复决策
给 no-hit / no-evidence 结果增加一等证据引用:
```json
{
"raw_path": "$.no_evidence",
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
}
```
Executor 可以引用:
```json
{
"tool_name": "query_logs",
"source_invocation_id": 123,
"raw_path": "$.no_evidence",
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
}
```
### 语义边界
`$.no_evidence` 只表示:
```text
该工具对当前查询返回无匹配证据。
```
它不表示:
- 问题绝对不存在。
- 根因被排除。
- 系统已经健康。
- 没有必要继续排查。
---
## 实施范围
### 已修改
- `ToolInvocationRecorder`
- `query_logs` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
- `query_metrics` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
- `lookup_knowledge` no-hit 且无 evidence blocks 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
- `ExecutorGatekeeperService`
- 复用既有 `evidence_refs` 校验逻辑,无需新增特殊分支。
- `$.no_evidence` 和普通 raw_path 一样必须存在于 `retrieval_details.evidence_refs`。
- `chat-executor-prompt.md`
- 明确 `negative_observation` 必须引用 `$.no_evidence`。
- 明确没有实际工具调用时禁止使用 `$.no_evidence`。
### 未修改
- 不新增数据库表。
- 不新增复杂 metadata。
- 不改变 Planner。
- 不改变 Agent 编排。
---
## 验收标准
1. `query_logs` 返回 `logs=[] / total=0 / evidence_status=no_evidence` 时,`tool_invocation.retrieval_details.evidence_refs` 包含:
```json
{
"raw_path": "$.no_evidence"
}
```
2. Executor 输出 `negative_observation` 并引用 `$.no_evidence` 时,Gatekeeper 可以校验通过。
3. Executor 如果用 `$.no_evidence` 搭配正向证据文本,例如 `HikariCP active=50/50`,Gatekeeper 必须拒绝。
4. `$.no_evidence` 不得被解释为“问题绝对不存在”,只能表达“当前查询未检索到匹配证据”。
5. HikariCP negative E2E:
```text
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
```
期望:
- 工具返回 no-hit。
- 不返回 `generic-service`。
- Executor 输出 `negative_observation`。
- `raw_path="$.no_evidence"`。
- Gatekeeper `pass/none`。
- Verifier 不误判为正向 HikariCP 证据。
---
## 验证记录
### 2026-07-08 单元测试
命令:
```text
mvn '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,QueryLogsToolsTest' test
```
结果:
```text
Tests run: 23, Failures: 0, Errors: 0, Skipped: 0
BUILD SUCCESS
```
覆盖点:
- `query_logs` no-hit 生成 `$.no_evidence`。
- `query_metrics` no-hit 生成 `$.no_evidence`。
- `lookup_knowledge` no-hit 生成 `$.no_evidence`。
- Gatekeeper 可以校验 `$.no_evidence`。
- 多个 no-evidence 调用存在时,Gatekeeper 可按 `tool_name + raw_path + evidence_excerpt` 唯一回填 `source_invocation_id`。
- `negative_observation` 混绑正向 `$.logs[i]` 会被拒绝。
- HikariCP negative mock 不返回 `generic-service`。
### 2026-07-08 E2E 验证
输入:
```text
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
```
最终通过 session:
```text
sessionId: iss009-hikari-negative-latest-20260708-232428
verdict: PASS
groundedness_score: 1.0
gatekeeper_result.status: pass
gatekeeper_result.severity: none
claim_count: 1
claim_type: negative_observation
raw_path: $.no_evidence
generic_service_hit: false
overstate_hit: false
```
Executor claim:
```text
当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。
```
最终答案:
```text
本次查询在 inventory-service 中未发现 HikariCP 连接池耗尽的日志记录,检索结果未匹配到相关证据。
```
说明:
- Executor 仍可能输出多个 `$.no_evidence` binding。
- 如果 `source_invocation_id` 缺失,Gatekeeper 会按 `tool_name + raw_path + evidence_excerpt` 唯一匹配真实 invocation 并写入 warning。
- 最终答案不使用“排除”“确认没有”“不存在该问题”等过度表达。
@@ -0,0 +1,586 @@
# ISS-010 同 session 多轮诊断 Trace 隔离
**状态**:已归档
**严重程度**:高
**发现时间**:2026-07-10
**来源**:同一 `sessionId` 多轮 Chat E2E 验证
**归档日期**:2026-07-10
**OpenSpec**:`openspec/changes/archive/2026-07-10-session-run-trace-isolation`
**实现提交**:`52bf030`、`26d5529`、`027aed1`、`d928a19`、`78c1477`、`f9df943`
**归档提交**:`3578709`
---
## 背景
归档结论:当前 MVP 已将“会话态”和“运行态”拆开。`sessionId` 表示多轮会话目录和 Redis 上下文;`runId` 表示一次可回放诊断执行。Trace、Feedback、Evaluation、AIOps 和案例沉淀的新路径都按 `runId` 隔离。
原问题中 Chat 链路同时存在两类“会话”语义:
```text
Redis SessionContext
-> 保存同一 sessionId 的多轮对话历史
-> 用于下一轮模型上下文
MySQL diagnosis_session / agent_step / tool_invocation
-> 保存诊断 Trace
-> 用于 Trace API、Verifier、Evidence score、Feedback 和评测
```
多轮对话需要继续复用 `sessionId`,否则无法保留上下文。但一次诊断 Trace 应该是可独立回放、可独立评分、可独立反馈的执行单元。
当前实现只按 `sessionId` 关联 Trace,导致同一个 `sessionId` 下多轮诊断的 step/tool 记录混在一起。
当前实现已改为:
```text
chat_session(sessionId)
-> diagnosis_run(runId)
-> agent_step.run_id
-> tool_invocation.run_id
```
旧 `diagnosis_session` 保留为历史兼容和回滚表,新 Chat/AIOps 执行不再写入新的运行态。
## 归档结果
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-10-session-run-trace-isolation`。
- 主规格已同步到 `openspec/specs/session-run-trace-isolation/spec.md`。
- `mvp/architecture/` 和 `mvp/tables/` 已更新为 `chat_session -> diagnosis_run -> agent_step/tool_invocation(run_id)` 模型。
- Demo 脚本和 Trace UI 已支持 `sessionId + runId` 精确 Trace 和 Feedback。
- Maven E2E、`scripts/query_mysql.py` DB 检查、`logs/` 日志检查和 baseline drift 检查均已通过;未观察到 baseline drift。
- `devflow/projects/2026-07-10-session-run-trace-isolation/` 已保存 brief、evidence、decisions、acceptance。
---
## E2E 证据
本次使用 `mvp-demo` profile 通过 Maven 启动服务,并用同一个 `sessionId` 连续请求两轮 `/api/chat`:
```text
sessionId = e2e-multiturn-codex-20260710-1615
round 1 = 支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。
round 2 = 基于上一轮结论,只列出目前最缺的三类证据,以及下一步应该优先查哪个系统。
```
验证结果:
- 第一轮成功,走多 Agent:`planner -> executor -> verifier -> composer`。
- 第二轮成功,日志显示进入请求时 `会话历史消息对数: 1`,说明 Redis 历史上下文被复用。
- `/api/chat/session/{sessionId}` 返回 `messagePairCount=2`。
- `diagnosis_session` 只有一行,`query` 被第二轮问题覆盖。
- `agent_step` 返回 14 行,包含第一轮多 Agent step 和第二轮 `intelligent_assistant` step。
- `tool_invocation` 返回 19 行,包含两轮工具调用。
- `self_evaluation.verifier_evaluation` 仍保留第一轮 Verifier 结果;第二轮简单问答没有新的 Verifier,但 rule evaluation 会基于同 session 全部工具调用重新计算。
关键入库形态:
```text
diagnosis_session
session_id = e2e-multiturn-codex-20260710-1615
query = round 2 question
status = SUCCESS
step_count = 14
tool_call_count = 19
agent_step
round 1: planner, executor..., verifier, composer
round 2: intelligent_assistant...
tool_invocation
round 1 tools + round 2 tools all under same session_id
```
本次验证产物保存在:
- `target/e2e/request-round1.json`
- `target/e2e/response-round1.json`
- `target/e2e/request-round2.json`
- `target/e2e/response-round2.json`
- `target/e2e/trace-after-round2.json`
---
## 核心问题
### P0:Trace 不是单次诊断的稳定回放
`GET /api/diagnosis/{sessionId}/trace` 会聚合同一 `sessionId` 下所有 `agent_step` 和 `tool_invocation`。
多轮之后,Trace 不再表示某一轮诊断,而是混合历史执行轨迹。
### P0:Verifier 和评分可能读取跨轮证据
Verifier、Gatekeeper、`ToolTraceSummaryService` 和 `EvaluationService` 当前主要按 `sessionId` 查询工具调用。
如果上一轮和当前轮证据混在一起,当前轮可能引用或评分到历史工具结果。
### P1:反馈语义不清晰
`feedback` 当前在 `diagnosis_session` 上按 `sessionId` 保存。
多轮之后,用户反馈的是哪一轮答案不再明确。`useful` 反馈沉淀到 `case_library` 时也可能关联到最新主表答案,而不是用户实际评价的那一轮。
### P1:`diagnosis_session` 字段被覆盖但子表追加
主表 `query/answer/status/self_evaluation/step_count/tool_call_count` 表示最新运行或混合统计,子表却保留多轮历史。
这会让 Trace summary、数据库统计和人工排查产生歧义。
---
## 已确认决策
### D1:`runId` 是正式 API 字段
`/api/chat` 和 `/api/ai_ops` 的响应或 SSE 消息需要暴露本次执行的 `runId`。
```text
sessionId = 多轮对话上下文 ID
runId = 本轮诊断执行 ID
```
新客户端应优先用 `runId` 查询 Trace 和提交 Feedback。旧客户端只传 `sessionId` 时,服务端兼容解析该 session 的最新 run。
### D2:拆分会话态和运行态
不再把 session、run、trace 全部塞进 `diagnosis_session` 一张主表。
新增两张主表:
```text
chat_session
-> 多轮对话上下文主表
diagnosis_run
-> 单次诊断执行主表
```
Trace 继续使用现有明细表表达:
```text
agent_step
tool_invocation
```
暂不新增单独的 `diagnosis_trace` 或 `trace_event` 主表。
### D3:`runId` 格式
使用 `run-` + UUID 全量字符串。
```text
run-550e8400-e29b-41d4-a716-446655440000
```
### D4:Trace API 兼容旧路径
```text
GET /api/diagnosis/{sessionId}/trace
-> 查该 session 最新 run
GET /api/diagnosis/{sessionId}/trace?runId=run-xxx
-> 查指定 run
```
指定 `runId` 时必须校验该 run 属于 path 中的 `sessionId`。
最新 run 建议按 `diagnosis_run.created_at DESC, id DESC` 解析,避免旧 run 因反馈或异步评分更新 `updated_at` 后被误认为最新。
### D5:Feedback 优先绑定 run
Feedback request 支持 `runId`。
- 有 `runId`:绑定指定 run。
- 无 `runId`:短期兼容绑定该 `sessionId` 最新 run,并显式标记 fallback。
- `case_library.diagnosis_id` 新数据保存 `run_id`。
兼容语义:历史 `case_library.diagnosis_id` 可能保存 `diagnosis_session.session_id`;本 change 之后自动沉淀的新数据保存 `diagnosis_run.run_id`。查询、幂等和文档需要在过渡期识别两种来源,避免把旧案例误判为无效数据。
### D6:所有 `/api/chat` 执行请求都创建 run
只要请求通过参数校验并进入 `ChatService.executeChatWithStrategy`,就创建新的 diagnosis run。
- 简单问答也创建 run。
- 复杂诊断也创建 run。
- 空问题等参数校验失败不创建 run。
### D7:AIOps 同步纳入 run 隔离
每次 `/api/ai_ops` 执行也创建新的 diagnosis run。AIOps 的 step、tool invocation 和 rule evaluation 都按 `runId` 隔离。
阶段说明:AIOps 可作为独立实现切片排在 Chat 之后,但必须在本 change 整体完成前落地;Chat-only 的中间状态只能作为过渡验证状态,不能作为生产完成状态归档。
### D8:`chat_session` 第一阶段只保存会话元数据
`chat_session` 是会话目录/索引表,不保存完整对话历史正文。
建议保存:
```text
session_id
status
message_pair_count
created_at
last_active_at
expires_at
```
完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`,用于下一轮 prompt 上下文。
每轮需要长期审计的用户问题和最终回答保存到 `diagnosis_run.query` / `diagnosis_run.answer`。
`chat_session.expires_at` 只表示 MySQL 会话目录的过期/清理元数据;Redis TTL 到期后,`SessionContext.messageHistory` 可能不再存在,但已经持久化的 `diagnosis_run`、`agent_step` 和 `tool_invocation` 仍作为审计记录保留。
如果未来需要长期保存完整聊天历史,再单独设计 `chat_message` 表,不在本阶段引入。
### D9:旧 `diagnosis_session` 表保留但新代码不再写入
新增 `diagnosis_run` 后,旧 `diagnosis_session` 不立即删除、不立即改造成 view、不直接重命名。
迁移策略:
1. 新增 `chat_session` / `diagnosis_run`。
2. 为 `agent_step` / `tool_invocation` 新增 nullable `run_id`。
3. 将旧 `diagnosis_session` 数据迁移/复制为 `diagnosis_run` 兼容记录。
4. 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
5. 增加必要索引和查询方法,先保持兼容读取。
6. 新代码切换为只写 `chat_session` 和 `diagnosis_run`,并为新 step/tool 写入 `run_id`。
7. 验证新旧数据 `run_id` 覆盖情况后,再将新写路径要求 `run_id` 非空,并补充索引/约束。
8. 旧 `diagnosis_session` 暂时保留,用于历史核对和回滚窗口。
9. 后续确认无依赖后,再单独归档或删除旧表。
### D10:提供轻量 run 列表 API
新增轻量查询接口,用于查看一个 Chat Session 下有哪些 Diagnosis Run。
```text
GET /api/chat/session/{sessionId}/runs
```
建议返回字段:
```text
runId
sessionId
query
status
agentFlow
answerPreview
stepCount
toolCallCount
createdAt
updatedAt
```
该接口只读 `diagnosis_run` 主表,不展开 `agent_step` / `tool_invocation` 大字段。
### D11:Feedback 缺少 `runId` 时短期兼容,长期收紧
Feedback 新协议优先要求 `runId`。
短期兼容策略:
- 有 `runId`:绑定指定 run。
- 无 `runId`:绑定该 `sessionId` 最新 run。
- 无 `runId` fallback 时,在响应或日志中明确标记 `fallbackToLatestRun=true`,并返回实际绑定的 `runId`。
长期收紧策略:
- 当前端、demo 脚本和外部调用方都完成 `runId` 传递后,再评估是否将缺少 `runId` 改为参数错误。
### D12:同步更新 demo 脚本和 Trace UI 的 `runId` 最小支持
本 issue 实施范围包含 demo 脚本和 Trace UI 的最小协议适配。
范围:
- Demo 脚本读取 `/api/chat` 或 `/api/ai_ops` 返回的 `runId`。
- Demo 脚本查询 Trace 时传 `?runId=...`。
- Trace UI 支持 URL 参数 `?sessionId=...&runId=...`。
- Trace UI 查询时如果有 `runId`,带上 `runId`。
- 不在本阶段实现完整 run 列表 UI。
---
## 目标语义
引入明确的 `sessionId` / `runId` 分层:
```text
sessionId = 多轮对话上下文
runId = 单次诊断执行 / 单次可回放 Trace
```
目标关系:
```text
chat_session(sessionId)
-> one conversation context
-> conversation metadata / TTL / last active state
Redis SessionContext(sessionId)
-> hot messageHistory cache
-> supports prompt context window
diagnosis_run(runId, sessionId)
-> one diagnosis run
agent_step(runId, sessionId)
-> steps of one run
tool_invocation(runId, sessionId)
-> tool calls of one run
```
---
## 建议方案
采用“拆分主表 + 复用现有 Trace 明细表”的方案:
1. 新增 `chat_session`。
2. 新增 `diagnosis_run`。
3. 逐步迁移当前 `diagnosis_session` 语义到 `diagnosis_run`。
4. `agent_step` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
5. `tool_invocation` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
6. Trace API 聚合 `diagnosis_run + agent_step + tool_invocation`。
建议核心字段:
```text
chat_session
id
session_id unique
status
message_pair_count
created_at
last_active_at
expires_at
diagnosis_run
id
run_id unique
session_id
query
status
agent_flow
answer
self_evaluation
feedback
total_duration_ms
total_token_count
step_count
tool_call_count
created_at
updated_at
agent_step(run_id, step_index)
tool_invocation(run_id, id)
```
理由:
- `chat_session` 只表达会话态,避免会话上下文和诊断结果混在一起。
- `diagnosis_run` 只表达一次执行,天然隔离每轮 Trace、评分和反馈。
- `agent_step` / `tool_invocation` 已足够表达 Trace 明细,暂不需要额外 trace 主表。
- 后续如果需要统一时间线,再增加 `trace_event`,不阻塞本次隔离。
---
## 分阶段计划
### Phase 0:协议基线和数据边界
目标:先把语义定死,避免实现中反复。
已确认基线:
1. `/api/chat` 是否返回 `runId`。
2. `GET /api/diagnosis/{sessionId}/trace` 默认查最新 run 还是要求显式传 `runId`。
3. Feedback 是否优先绑定 `runId`,只有旧请求缺失 `runId` 时才回退最新 run。
4. AIOps 是否和 Chat 同步接入 `runId`。
5. 简单问答是否也创建 diagnosis run。
建议默认:
- `/api/chat` 返回 `sessionId + runId`。
- `GET /api/diagnosis/{sessionId}/trace` 兼容查最新 run。
- `GET /api/diagnosis/{sessionId}/trace?runId=...` 查指定 run。
- Feedback 优先按 `runId` 绑定。
- Chat 简单问答也创建 run。
- AIOps 同步接入 run 隔离。
### Phase 1:Schema 迁移和历史数据兼容
目标:引入 `chat_session` / `diagnosis_run`,并保留旧数据可查询。
任务:
- Flyway 新增 `chat_session`。
- Flyway 新增 `diagnosis_run`。
- 为旧 `diagnosis_session` 生成兼容 `diagnosis_run` 记录。
- 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
- 增加 `find latest run by sessionId` 查询。
- 增加 `find by runId` 查询。
- 保留旧 `diagnosis_session` 一段时间,新代码不再写入。
验收:
- 旧 session 的 Trace 仍可查。
- 新索引存在。
- 不改变旧 `/api/chat` 必需字段。
- `agent_step` / `tool_invocation` 支持 nullable `run_id` 并完成旧数据回填。
- 新增 repository 查询可以按 `sessionId` 找最新 run、按 `runId` 找指定 run。
### Phase 2:Chat 写入切到 runId
目标:每轮 `/api/chat` 创建一个新的 run,step/tool 按 run 隔离。
任务:
- `ChatService` 每次执行生成新的 `runId`。
- `ChatController` 确保 `chat_session` 存在并更新会话态。
- `ChatService` 按 `runId` 创建 `diagnosis_run`。
- `AgentLoggingHook` 写入 `agent_step.run_id`。
- `ToolInvocationRecorder` 写入 `tool_invocation.run_id`。
- `SessionContextHolder` 或新的上下文 holder 同时携带 `sessionId + runId`。
- `backfillSessionMetrics` 按 `runId` 统计。
- `EvaluationService` 按 `runId` 读取工具调用。
验收:
- 同一 `sessionId` 连续两轮后,`diagnosis_run` 有两行不同 `run_id`。
- `/api/chat` 响应增加正式字段 `runId`。
- 两轮 `agent_step` / `tool_invocation` 分别按各自 `run_id` 查询。
- Redis `messagePairCount` 仍为 2,证明上下文不被破坏。
### Phase 3:Trace API 兼容和精确查询
目标:Trace API 可查最新 run,也可查指定 run,并能列出一个 session 下的 run。
任务:
- `GET /api/diagnosis/{sessionId}/trace` 从 `diagnosis_run` 默认解析最新 run。
- 增加 `runId` query 参数。
- Trace response 增加 `runId`。
- 新增 `GET /api/chat/session/{sessionId}/runs`。
验收:
- 不传 `runId` 返回最新 run。
- 传第一轮 `runId` 只返回第一轮 step/tool。
- 传第二轮 `runId` 只返回第二轮 step/tool。
- run 列表 API 只返回轻量 run 摘要,不展开 trace 明细。
- Demo 脚本和 Trace UI 的 `runId` 最小适配按 OpenSpec tasks 放到 Phase 6,避免 Phase 3 同时混入前端/脚本范围。
### Phase 4:Feedback 和 CaseLibrary 绑定 run
目标:反馈明确评价哪一轮诊断。
任务:
- Feedback request 支持 `runId`。
- 旧请求只有 `sessionId` 时短期绑定最新 run,并显式标记 fallback。
- `case_library.diagnosis_id` 新数据保存 `run_id`。
- `CaseLibraryService` 以 run 为来源生成 case,并用 `run_id` 做新数据幂等键。
- 文档说明 `diagnosis_id` 的过渡语义:旧数据可能是 `session_id`,新数据是 `run_id`。
验收:
- 同 session 多轮后,对第一轮提交 feedback 不会覆盖第二轮。
- useful 生成 case 时能定位到对应 run 的 query/answer。
### Phase 5:AIOps 同步 run 隔离
目标:AIOps 使用同样的 run 语义,避免另一条入口继续混杂。
任务:
- `AiOpsService` 生成并返回/透出 `runId`。
- AIOps `agent_step` / `tool_invocation` 按 `runId` 隔离。
- AIOps rule evaluation 写入当前 run 的 `diagnosis_run.self_evaluation.aiops_rule_evaluation`。
- AIOps Trace 查询兼容 `sessionId + runId`。
验收:
- 同一 AIOps `sessionId` 重跑不会混合 step/tool。
- AIOps rule evaluation 只读取当前 run 工具调用。
---
## 暂不做
1. 暂不新增 `diagnosis_trace` 或 `trace_event` 主表。
2. 暂不做完整 run 列表 UI。
3. 暂不删除历史 Trace 数据。
4. 暂不改变 Redis 多轮上下文窗口策略。
5. 暂不立即物理删除旧 `diagnosis_session` 表。
---
## 风险
### 1. 兼容风险
现有脚本、Trace 页面和反馈接口可能只知道 `sessionId`。
缓解:保留 `sessionId` 默认查最新 run 的行为。
### 2. 异步上下文风险
工具调用和 Agent hook 依赖 ThreadLocal / RunnableConfig 传递上下文。
缓解:统一上下文对象,明确 `sessionId` 和 `runId` 必须同时传递。
### 3. 历史数据回填风险
旧数据没有真实 run 边界,只能按当前 `diagnosis_session` 生成一条兼容 `diagnosis_run`。
缓解:旧数据视为单 run,不尝试拆分历史混合数据。
### 4. 评分口径变化风险
按 `runId` 隔离后,工具调用数和 evidence score 可能下降,但语义更正确。
缓解:更新 eval fixture 和 baseline,记录这是预期行为变化。
---
## 已收敛问题
本 issue 当前已经收敛以下设计边界:
- `runId` 是正式 API 字段。
- `chat_session` 和 `diagnosis_run` 拆分为两张主表。
- Trace 明细继续由 `agent_step` / `tool_invocation` 承载。
- `chat_session` 只保存元数据,不保存完整对话历史。
- 旧 `diagnosis_session` 保留但新代码不再写入。
- 提供轻量 run 列表 API。
- Feedback 缺少 `runId` 时短期兼容、长期收紧。
- Demo 脚本和 Trace UI 做 `runId` 最小支持。
---
## 相关文件
- `mvp/architecture/session-trace-lifecycle.md`
- `mvp/architecture/data-model.md`
- `mvp/tables/聊天会话表-chat_session.md`
- `mvp/tables/诊断运行表-diagnosis_run.md`
- `mvp/tables/诊断会话表-diagnosis_session.md`
- `mvp/tables/Agent步骤表-agent_step.md`
- `mvp/tables/工具调用表-tool_invocation.md`
- `mvp/tables/案例库表-case_library.md`
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
- `src/main/java/com/superbiz/agent/service/EvaluationService.java`
- `src/main/java/com/superbiz/agent/service/FeedbackService.java`
- `src/main/java/com/superbiz/agent/service/CaseLibraryService.java`
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/util/SessionContextHolder.java`
- `src/main/resources/db/migration/V005__create_session_storage.sql`

Some files were not shown because too many files have changed in this diff Show More