feat(graph): cut over chat diagnosis stategraph

This commit is contained in:
zhuyongxin
2026-07-17 18:30:08 +08:00
parent 1460dd1e99
commit 99e490f227
36 changed files with 2640 additions and 1623 deletions
+1
View File
@@ -4,6 +4,7 @@
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|---|---|---|---|---|---|---|
| 2026-07-17 | chat-diagnosis-stategraph-chatservice-cutover | 将复杂 Chat 单轨切换到 Diagnosis StateGraph,并增加 Run 级 orchestration trace 和 verified-only Verifier 输入。 | Chat diagnosis orchestration/production cutover | ChatService, CompiledGraph stream, runId metadata, orchestration trace, verified-only prompt, V012 | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover | archived |
| 2026-07-17 | chat-diagnosis-stategraph-real-nodes | 接入真实 Agent/Java Nodes、显式 Gatekeeper、可信输入投影、关键证据补查与安全 Fallback,暂不切换生产入口。 | Chat diagnosis orchestration/nodes | ReactAgent adapter, Gatekeeper node, verified input, evidence retry, safe fallback, CompiledGraph | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes | archived |
| 2026-07-17 | chat-diagnosis-stategraph-routing-skeleton | 实现未接生产入口的 Diagnosis StateGraph 骨架、有限路由和 Fake Node 测试。 | Chat diagnosis orchestration/graph | StateGraph, fake node, conditional edge, retry counter, orchestration events, trace builder | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-routing-skeleton | archived |
| 2026-07-17 | chat-diagnosis-stategraph-design-freeze | 冻结 ISS-011 的 Graph State、条件边、有限重试、安全降级、审计和测试迁移边界。 | Chat diagnosis orchestration/design | StateGraph, runId, Gatekeeper, verified evidence, fallback, orchestration trace, test migration | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze | archived |
@@ -0,0 +1,52 @@
# Chat Diagnosis StateGraph ChatService Cutover 验收
## 结果
已接受。OpenSpec tasks 26/26 完成,阶段 3 可归档。
## 验证
### 静态验证
- `git diff --check`:通过。
- source/reference isolation:cutover 核心文件中的 `SequentialAgent`、`VerifierContextHolder`、`VerifierInputHook`、`tool_trace_summary` 命中 0;核心 TODO/FIXME/placeholder 命中 0。
- schema whitelist:V012 为 1 个 ALTER TABLE、1 个 ADD COLUMN、0 CREATE、0 DROP,目标仅 `diagnosis_run.orchestration_trace JSON NULL`。
- `openspec validate chat-diagnosis-stategraph-chatservice-cutover --strict`:通过。
- `openspec validate --specs --strict`:14 passed,0 failed。
### 脚本验证
- focused Maven regression:Graph runtime/result mapper/real nodes、ChatService cutover、Trace、Controller、Repository、Gatekeeper、Composer、Eval,共 27 suites / 119 tests;0 failures、0 errors、0 skipped。
- `mvn -q -DskipTests test-compile`:通过。
- `ChatServiceGraphIntegrationTest` 覆盖 SUCCESS、handled Fallback、unhandled failure、blank answer/partial trace 和同 session 多 Run 隔离。
### 浏览器/人工验证
- 未运行。阶段 3 的公开协议由 Controller/service tests 覆盖,完整人工/live 验收按用户规则保留到阶段 5。
### 未验证
- 未使用 Maven 启动应用做 live E2E。
- 未检查 `logs/` 运行日志。
- 未执行 `scripts/query_mysql.py` 查询真实数据库。
- 原因:用户明确要求只有阶段 5 全部完成后统一执行端到端、日志和数据库验收;阶段 3 只做风险相关自动化验证。
- 剩余风险:真实模型/工具调用下的 Prompt 行为、Flyway 在真实 MySQL 的应用结果和最终 Trace 数据须由阶段 5 E2E 证明。
## 已完成范围
- 复杂 Chat 唯一生产编排切换到 Diagnosis StateGraph,公开 ChatResult 保持兼容。
- Run 生命周期、metrics、Eval、finally cleanup 与 runId 隔离保持;handled Fallback=SUCCESS,未处理/空答案=FAILED。
- 新增 Run 级 compact orchestration trace,并只在 Trace run 对象暴露解析结果。
- Verifier Prompt/self-evaluation 使用 verified-only 数据,不伪造未发生 verdict/event。
- 移除冲突的 Sequential 实现测试并以 Graph/public contract tests 替代。
## Bug 修复和诊断
- 替代测试错误引用 `com.superbiz.agent.tool.ToolInvocationRecorder`:通过定义/引用搜索确认类型位于同一 `service` 包,删除错误 import。
- no-answer 新测试错误期待 inner cause:分类为测试断言偏差,改为公开 wrapper error,同时保留 FAILED 与真实 partial trace 核心断言。
## 交接
- OpenSpec archive:已同步 5 份 delta specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover/`。
- 下一步:审查精确 Git diff 并完成阶段 3 独立提交,之后才启动阶段 4。
- OpenSpec 归档确认:用户已明确授权后续阶段直接实现/归档;归档已完成。
@@ -0,0 +1,21 @@
# Chat Diagnosis StateGraph ChatService Cutover Brief
## 背景
- 用户目标:将 ISS-011 阶段 3 作为独立 sm-flow,正式切换复杂 Chat 生产编排并增加 Run 级 orchestration trace。
- 当前问题:真实 Diagnosis Graph Nodes 已存在,但生产入口仍依赖 SequentialAgent、外层重试和 ThreadLocal/Hook 隐式状态。
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-chatservice-cutover/`
- devflow 分档:complex。
## 范围
- 本次要做:复杂 Chat 单轨 Graph cutover;复用四类 Agent builder;Graph result/self-evaluation 映射;V012 Run JSON 字段;Trace run-only 投影;verified-only Verifier Prompt;必要回归测试。
- 本次不做:不改 `/api/chat` 请求/响应,不做历史 trace 回填,不删除仍供历史代码/测试使用的 Hook/ThreadLocal 类型,不完成阶段 4 测试体系全面收敛,不运行 live E2E/log/DB 验收。
- 影响区域:ChatService、Diagnosis Graph runtime/result mapper、DiagnosisRun/Flyway、Trace DTO/service、Verifier Prompt、Graph/Service/Trace tests。
## OpenSpec 对齐
- proposal 覆盖状态:已覆盖。
- design 覆盖状态:已覆盖。
- specs 覆盖状态:已覆盖,5 个 delta capabilities。
- tasks 覆盖状态:26/26 已完成。
@@ -0,0 +1,207 @@
# Chat Diagnosis StateGraph ChatService Cutover Decisions
## Entry Summary
- 问题:复杂 Chat 仍使用 SequentialAgent + ThreadLocal/Hook 隐式状态机,真实 Graph 尚未成为生产入口,也没有 Run 级 orchestration trace 持久化和 API 投影。
- 期望:阶段 3 独立完成生产 cutover、Run/Trace 映射和必要测试,归档并提交后才进入阶段 4。
- 分档:complex。
- Change:`chat-diagnosis-stategraph-chatservice-cutover`。
- 授权:用户已要求后续阶段直接实现,不再逐 checkpoint 等待;阶段门禁、独立 archive/commit 与阶段 5 才 E2E 约束不变。
## Context Sources
- `mvp/issues/active/ISS-011-chat-diagnosis-stategraph-orchestration.md` 阶段 3、协议影响、Run 状态和验收章节。
- 阶段 0–2 OpenSpec archives、devflow acceptance/decisions 与当前四份 Graph 主 specs。
- `devflow/glossary/CONTEXT.md` 中 Chat Session、Diagnosis Run、Diagnosis Trace、Diagnosis Orchestration Trace 的边界。
- `ChatService` 生产调用链、四个 Agent builder、Run persistence/self-evaluation/metrics 逻辑。
- `DiagnosisRun`、V011 migration、`DiagnosisTraceResponse`、`DiagnosisTraceService` 与相关 tests。
- `DiagnosisGraphFactory`、真实 action factory、Node adapters、final trace builder 和本地 CompiledGraph API。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | orchestration trace 是否属于 self-evaluation 或完整 Trace event log? | evidence-driven | 已解决 |
| Q2 | 边界 | 阶段 3 是否改变 `/api/chat`,以及 Trace 字段出现在哪一层? | user-interview(既有冻结决策) | 已确认 |
| Q3 | 生命周期 | LOW_CONFID/REJECT/Fallback 是否应使用 FAILED 或新增 DEGRADED? | user-interview(既有冻结决策) | 已确认 |
| Q4 | 错误处理 | Graph 异常时如何 best-effort 保存已有编排事实而不伪造 event? | evidence-driven | 已解决 |
| Q5 | 技术实现 | 是否复用现有 Agent factory/Hook/ToolCallback 与真实 Node assembly? | evidence-driven | 已解决 |
| Q6 | 安全 | Graph Verifier Prompt/self-evaluation 是否可继续使用 raw Executor 和完整 tool trace? | evidence-driven | 已解决 |
| Q7 | 验收 | 阶段 3 是否需要新增测试,是否现在运行 live E2E? | user-interview(用户最新规则) | 已确认 |
## Evidence-driven Findings
- Q1:glossary 与阶段 0 spec 已定义 orchestration trace 为 Run 级紧凑编排摘要,与 self-evaluation、AgentStep/ToolInvocation 和 checkpoint 分离;无需新增术语或 ADR。
- Q4:所有已定义 Agent/Java 失败应由 Graph 路由到 handled Fallback 并返回 final state;只有真实可取得的 final/partial state 才可构造 trace。无法取得 state 的未处理异常标记 Run FAILED,不得生成虚假 transition。
- Q5:`DiagnosisRealGraphActionsFactory` 已提供真实 Node assembly;ChatService 现有四个 builder、AgentLoggingHook、Skills hooks、method tools 和 ToolCallbacks 可直接构造 ReactAgent,再通过 `ReactAgentDiagnosisInvoker` 注入,不得复制 Prompt/Agent factory。
- Q6:阶段 2 spec 已禁止 Graph Verifier 使用 raw Executor/full tool trace;当前 `chat-verifier-prompt.md` 仍描述旧 Hook payload,是阶段 3 必须同步的明确 gap。self-evaluation 可保存 Graph 的 verified output、Gatekeeper audit 和 Composer audit,但不能重新引入 raw 输入。
## User-interview Confirmations
| 问题 | 用户原话/既有确认 | 确认状态 | OpenSpec 回写 |
|---|---|---|---|
| Q2 外部协议与 Trace 层级 | ISS-011 已冻结 `/api/chat` 不变,Trace 只在 `run.orchestrationTrace` 增解析对象 | 已确认 | proposal |
| Q3 Run 生命周期 | ISS-011 已冻结安全 Fallback 为 SUCCESS,不新增 DEGRADED;只有无法生成安全响应的未处理失败为 FAILED | 已确认 | proposal |
| Q7 验收节奏 | “端到端只在最后阶段全部完成后才验证;每个阶段如果有必要添加单元测试验收的话,就加” | 已确认 | proposal |
## Interface Impact
- 等级:L4。
- 原因:复杂 Chat 内部状态机正式切换;DiagnosisRun 增加数据库 JSON 契约;Trace run 对象新增字段;旧固定顺序消费者/测试不再成立。
- 保持兼容:`/api/chat` request/response、sessionId/runId、Executor/Verifier/Composer 输出契约不变。
- 加法变化:只新增 `diagnosis_run.orchestration_trace` 和 `run.orchestrationTrace`。
- 回滚:revert 阶段 3 代码/Prompt/spec,数据库列可保留 nullable;不保留运行时双轨开关。
## Discover Status
- `devflow/index.md`:命中阶段 0–2 archives 和 Run/Trace 历史项目。
- Glossary:相关术语已存在且无冲突,不需更新。
- ADR:阶段 0 已记录难以逆转的状态机/Trace 决策,本阶段没有新的三条件 ADR。
- 未解决问题:0。
- Draft 产物:当前只创建 proposal + decisions;design/specs/tasks 留到 Commit checkpoint。
## Grill-with-docs Review
### Domain model stress test
- 同一 session 连续两个复杂 Chat run:每次编译/执行使用自己的 runId threadId 和 metadata;orchestration trace 只写各自 DiagnosisRun,session 投影不复制,符合 Session/Run/Trace 领域边界。
- Gatekeeper REJECT 或 Planner/Executor 技术失败:Graph 进入确定性 Fallback,返回安全非空答案;Run 为 SUCCESS,trace `degraded=true`,不会把诊断质量塞进 Run status。
- Composer 成功但 verdict=LOW_CONFID/REJECT:Run 仍为 SUCCESS;self-evaluation 保存 effective verdict,orchestration trace 保存实际路径,两者职责不混合。
- Graph 未处理异常:使用流式 NodeOutput 捕获最后一个真实 state,只有其中已有 events 时才 best-effort 构造 partial trace;没有 event 时不伪造 transition,Run 标记 FAILED。
- 历史/AI_OPS run:新增列 nullable,Trace DTO 对旧 run 可返回 null;不增加历史回填或跨 flow 假数据。
### Design tree conclusions
- 生产切换采用单轨,不增加 feature flag 或保留 Sequential/Graph 双运行;回滚依靠 Git revert,nullable 列可保留。
- ChatService 负责 Run 生命周期和调用一个专用 Graph runtime/orchestrator;复杂 Agent 构造从旧 Sequential 私有流程中搬移或封装复用,不复制第二套 builder。
- Graph 执行优先使用可观察的 `CompiledGraph.stream(..., config)` 收集最后真实 state,以支持异常时 best-effort trace;正常终止仍以 END state 的 `final_answer` 为唯一答案来源。
- Graph result mapping 形成独立组件:final answer、Verifier evaluation、Composer audit、Gatekeeper audit、trace JSON 均从显式 final state读取,不再访问 VerifierContextHolder。
- Verifier Prompt 必须更新为 `diagnosis_context`、`verified_executor_output`、`verified_evidence`、`gatekeeper_audit`、`verdict_ceiling` 和可选 retry context;删除 raw Executor/full tool trace/Hook Gatekeeper 说明。
- 阶段 3 必须有 production cutover 与 Trace DTO/Service tests;阶段 4 再做测试体系全面改名、夹具收敛和旧测试删除。
### Documentation result
- 术语与 `devflow/glossary/CONTEXT.md` 完全一致,无需修改 glossary。
- 状态机、Run status 和 orchestration trace 隔离均来自阶段 0 已归档 ADR/规格,不创建重复 ADR。
- proposal 已反映单轨切换、L4 接口影响、流式 partial-state 处理、Prompt 安全边界和阶段 5 E2E 延期。
- Grill question pool 全部关闭;无需要再次询问用户的产品取舍。
## Architecture Audit
### Module and caller map
`ChatController` 的 normal/SSE 两个入口都调用 `ChatService.executeChatWithStrategy`,复杂分支进入 `executeChatComplex`;对外仍只消费 `ChatResult(answer, sessionId, runId)`。阶段 3 将内部链路变为 `ChatService Run lifecycle -> complex Chat Graph runtime/Agent assembly -> real Diagnosis Nodes -> Graph result mapper -> DiagnosisRun persistence -> DiagnosisTraceService`。Trace UI/Eval 当前读取兼容 `session.selfEvaluation`,它仍由 Run self-evaluation 投影;新增 orchestration trace 只属于 `run`。AIOps、Feedback、CaseLibrary 和 Run list 继续使用 DiagnosisRun 既有字段,不消费新增 orchestration trace。
| 模块 | 所有权 | 允许的依赖/影响 |
|---|---|---|
| ChatController | `/api/chat` normal/SSE 协议 | 继续只依赖 ChatResult;无字段变化 |
| ChatService | Chat Session/Diagnosis Run 生命周期、成功/失败保存、metrics、Eval、finally cleanup | 调用一个 Graph runtime/result mapper;不再拥有条件边/重试/Gatekeeper/Composer 路由 |
| Complex Chat Graph runtime | 每请求 Agent assembly、initial state、RunnableConfig、CompiledGraph stream | 复用现有 Prompt/tool/skill/logging;不持久化跨 Run 状态 |
| Diagnosis Graph | Node status、verified material、events、final answer | 保持阶段 1/2 路由和 counter 所有权 |
| Graph result mapper | final/partial state 到安全 evaluation/trace DTO | 不读 ThreadLocal,不查询其他 run,不持久化 raw material |
| DiagnosisRun/Flyway | 当前 Run 的 orchestration JSON | 仅一个 nullable JSON 列;历史/AIOps 可为 null |
| DiagnosisTraceService/DTO | exact/latest Run 查询和 API 投影 | 只在 RunTrace 加 parsed map;session/top-level/run-list 不重复 |
### Lifecycle and coupling audit
ReactAgent 与 CompiledGraph 都按请求构造,因为 Planner/Executor system Prompt 含本次 history/knowledge map;Graph State 和 runtime last-state holder 也是 invocation-scoped,不进入 singleton 可变字段。SessionContextHolder 仍只包围当前请求并在 finally 清理;Graph Node config 以 runId threadId + metadata 绑定 AgentStep、ToolInvocation 和 Gatekeeper。success persistence 在 Eval 之前完成,EvaluationService 再按 runId 合并 rule channel;handled Fallback 与 Composer 使用同一 SUCCESS 路径。Trace JSON 与 self-evaluation 分栏,避免把路线、质量和完整 Agent/tool 明细耦合到一个容器。
### Consumer impact audit
- ChatController normal/SSE:ChatResult 协议不变,L4 内部切换不要求调用方迁移。
- Trace UI/demo:继续读取兼容 `session.selfEvaluation`;新的 `run.orchestrationTrace` 是加法字段,最终脚本断言留到阶段 5。
- DiagnosisTraceEvaluator:现有 fixtures 不变;工具覆盖仍可从 `toolInvocations` 读取,兼容 `executor_structured_output` 保存 verified projection。pre-verification Fallback 质量未来应以 trace degraded 而非伪 verdict 判断,属于阶段 4 测试体系/阶段 5文档收尾。
- AIOps/Feedback/CaseLibrary/Run list:实体新增 nullable 字段不改变 builder call sites或查询语义。
- 数据库:Hibernate validate 要求 V012 与 entity 同批;回滚代码可忽略保留列。
### Cross-artifact alignment
| 对齐链 | 结果 | 证据 |
|---|---|---|
| brief/proposal 目标、范围、非目标 → proposal | 已对齐 | 单轨 cutover、Run/Trace、Prompt、阶段边界与 E2E 延期均明确 |
| proposal 承诺与约束 → design | 已对齐 | 10 项决策覆盖 runtime、state、config、partial state、result、persistence、DB/API/Prompt/tests |
| design 架构/接口结论 → specs/tasks | 已对齐 | L4、单轨、Run status、verified-only、Trace 唯一投影和 schema whitelist 均有 requirement/task |
| specs 可观察行为 → tasks 可执行切片 | 已对齐 | 6 组 26 个切片覆盖数据、runtime、Prompt、cutover、tests、handoff |
### Audit result
未发现与阶段 0–2、Run/Trace ADR 或 glossary 冲突。审计确认不能把 Agent factory 复制进 Node,也不能把 pre-verification Fallback 伪装为 Verifier verdict;两项已在 design/specs/tasks 固定。唯一跨阶段依赖是阶段 4 让测试/Eval 以 orchestration degraded 理解无 Verifier verdict 的安全 Fallback,已记录但不阻塞阶段 3 production correctness。架构风险可接受,cross-artifact gap=0,无未解决接口消费者。
## Commit Gate
- schema:spec-driven;proposal/design/specs/tasks 全部 done,applyRequires=`tasks` 已满足。
- OpenSpec:当前 change strict validation 通过;14 个主 specs 全部 strict pass。
- 规格结构:5 个 delta capabilities、23 条 requirements、72 个 scenarios;tasks 26 个可执行 checkbox。
- Cross-artifact:4/4 已对齐,gap=0。
- Interface impact:L4;design 已独立记录消费者、加法 DB/Trace 协议、单轨迁移和 Git revert 回滚。
- Question pool:所有 evidence-driven 已查证;所有 user-interview 已由 ISS-011/用户原话确认;无未决项。
- Preflight:`git diff --check` 通过;Commit checkpoint 尚未修改 Java、SQL、Prompt 或测试。
- 结论:Draft OpenSpec 已达到可执行状态,创建 `.committed`,阶段 3 Apply 只能以这些产物为依据。
## Apply Authorization
- 用户原话:“直接实现吧,不用找我授权了”。
- 本阶段在 Commit gate 后直接进入 Apply;不扩大到阶段 4 全面测试迁移或阶段 5 live E2E。
## Pre-apply Research
### Reference implementations read
- `ChatService.executeChatComplex`、四个 complex Agent builders、Run start/save、metrics、Prompt audit、self-evaluation merge:迁移源实现和外部兼容基线。
- `DiagnosisGraphFactory`、`DiagnosisRealGraphActionsFactory`、四个 Agent adapters、Gatekeeper/VerifiedInput/Fallback、`DiagnosisOrchestrationTraceBuilder`:Graph 路由/计数/安全材料唯一真理源。
- `DiagnosisRun`、V011 migration、`DiagnosisTraceResponse`、`DiagnosisTraceService`、`DiagnosisTraceServiceTest`:Run JSON 字段和 exact/latest Trace 映射标准。
- `ChatController` normal/SSE、`DiagnosisTraceController`:公开协议调用者与响应包装边界。
- `SelfEvaluationMergeService`、`EvaluationService`、`DiagnosisTraceEvaluator`:evaluation container、异步 rule merge 和兼容消费者。
- 本地 graph-core 1.1.2.0 `javap`:CompiledGraph `stream/invoke/state`、NodeOutput `state()`、RunnableConfig `threadId/metadata` 的真实 API。
- `chat-*-prompt.md`:现有 Prompt 组装与 Verifier 旧 Hook payload gap。
### Technology inventory
| 类别 | 项目标准 / 本阶段使用 |
|---|---|
| Graph execution | `CompiledGraph.stream(initial, config)` + NodeOutput.state,当前请求线程阻塞消费 |
| Run context | SessionContextHolder + RunnableConfig threadId/metadata;runId 是唯一执行边界 |
| Agent assembly | ReactAgent builder、AgentLoggingHook、PlannerSkillMetadataHook/SkillsAgentHook、现有 tools/callbacks |
| JSON persistence | Jackson map serialization;实体 String + `@JdbcTypeCode(SqlTypes.JSON)`;Flyway JSON column |
| Trace API | Lombok DTO builder + DiagnosisTraceService parsed Map;exact/latest Run 查询 |
| Evaluation | SelfEvaluationMergeService container;EvaluationService 按 runId 异步合并 rule channel |
| Tests | JUnit 5,通过 public service/runtime/Trace API;只 mock repository/model/tool 外部边界 |
| MQ/Consumer | 不涉及 |
### New infrastructure and reuse
- 新增一个深接口的 complex Chat Graph runtime/result mapper;复用既有 Graph/Node/Agent builders,不增加新依赖或第二套路由。
- 新增 V012、DiagnosisRun 字段和 RunTrace parsed map;不新增表、repository method 或历史 backfill。
- TDD tracer bullet 从 Trace DTO/Service 的唯一 Run 投影开始,再进入 runtime final-state mapping,最后切换 ChatService。
- 研究未发现 devflow/OpenSpec 冲突,技术清单足以开始实现。
## Apply Progress
- TDD Trace slice RED:DiagnosisTraceServiceTest 明确缺少 DiagnosisRun builder 字段和 RunTrace getter。
- GREEN:V012、DiagnosisRun JSON 字段、RunTrace parsed map、DiagnosisTraceService mapping 完成;10 个 Trace service tests 通过。
- 失败分类:首次 GREEN 运行仅测试夹具用裸 ObjectMapper 无法序列化 LocalDateTime,属于 test harness 偏差;改为项目可用的 `findAndRegisterModules()` 后原始 focused loop 通过,未改业务协议。
- Trace API JSON 断言证明顶层/session 无重复字段、run 为解析对象、无 raw 字段;null/invalid JSON fail closed。
## Apply Completion
- 复杂 Chat 已单轨切换到 `ChatDiagnosisGraphRuntime`;生产 `ChatService` 不再创建 `SequentialAgent`,不再读取 `VerifierContextHolder`,Graph Verifier 只保留 `AgentLoggingHook`。
- runtime 使用 query-only initial state、`threadId=runId` 和 sessionId/runId metadata,流式保留最后真实 state;异常只保存真实 partial state/trace。
- `DiagnosisGraphResultMapper` 只从 verified projection 构造兼容 self-evaluation;Verifier 未完成时不伪造 verdict,不持久化 full tool trace/raw Executor。
- Verifier Prompt 与 runtime payload 已统一为 verified-only,Prompt audit 更新为 `chat-prompts-v2` / `chat-verifier-v3`。
- 旧 `ChatServiceSequentialAgentTest` 因绑定已删除的固定顺序/ThreadLocal/score retry 实现而移除,由 public service/runtime/route tests 替代。
- V012 schema 白名单测试确认只新增 `diagnosis_run.orchestration_trace JSON NULL`,无其他 schema object。
## Verification Summary
- focused regression:27 suites / 119 tests,0 failures、0 errors、0 skipped。
- Maven test compilation:通过。
- OpenSpec:当前 change strict pass;主 specs 14/14 strict pass。
- 静态门禁:`git diff --check` 通过;cutover 禁用引用 0;核心 TODO/placeholder 0;schema whitelist 为 1 ALTER / 1 ADD COLUMN / 0 CREATE / 0 DROP。
- 实现期唯一失败分类:no-answer service test 曾错误期待 inner cause 文本,属于测试断言偏差;按公开 wrapper error + FAILED/partial trace 契约修正后通过,OpenSpec 与生产代码无需变更。
- 按阶段门禁未运行 Maven live E2E、未检查 `logs/`、未执行 `scripts/query_mysql.py`;统一保留到阶段 5。
## Archive Result
- 5 份 delta specs 已同步:新增 9、修改 13、删除 1 条 requirements。
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover/`。
- CLI 对 proposal 结构给出非阻塞建议(大 change/delta 拆分与 SHALL/scenario 启发式);tasks 26/26、delta specs 和 strict validation 均通过,未形成验收阻塞。
@@ -0,0 +1,21 @@
# Chat Diagnosis StateGraph ChatService Cutover Evidence
## 证据
| 来源 | 证据 | 结论 | 是否已汇报 |
|---|---|---|---|
| `ChatService` + source isolation search | 复杂路径只有一次 Graph runtime 调用,无 SequentialAgent、VerifierContextHolder、VerifierInputHook 或 score retry | 单轨 cutover 完成,简单 Chat 路径未改 | 是 |
| `ChatDiagnosisGraphRuntimeTest` / `DiagnosisRealGraphIntegrationTest` | query-only state、runId thread/metadata、Composer/Fallback、empty stream、blank answer、partial trace、Gatekeeper once/twice | Graph identity、路由和真实 partial-state 边界可观察 | 是 |
| `DiagnosisGraphResultMapperTest` / `ChatVerifierPromptContractTest` | verified projection、effective verdict 兼容、pre-verification 无伪 verdict、禁用 raw/full trace 字段 | self-evaluation 与 Prompt 安全边界一致 | 是 |
| `ChatServiceGraphIntegrationTest` | ChatResult、agent_flow、SUCCESS/Fallback/FAILED、metrics/Eval、partial trace、多 Run 隔离 | 生产生命周期与兼容返回满足阶段 3 规格 | 是 |
| `DiagnosisTraceServiceTest` | `run.orchestrationTrace` 解析对象、top/session 无重复、invalid/null fail closed | Trace 字段所有权保持 Run 隔离 | 是 |
| `DiagnosisRunSchemaContractTest` | V012 executable SQL 与唯一允许语句精确相等 | schema 仅增加一个 nullable JSON 列 | 是 |
| focused Maven suite | 27 suites / 119 tests,0 failures/errors/skipped | Graph、Controller、Repository、Gatekeeper、Composer、Eval 回归通过 | 是 |
| OpenSpec/static gates | change strict、14/14 主 specs、test compile、diff check、禁用引用/占位/schema whitelist 全通过 | 规格、编译和源级隔离闭环 | 是 |
## Evidence-driven 结论
- orchestration trace 必须独立于 self-evaluation 和详细 Agent/tool Trace;V012 与 RunTrace 的实现保持这一边界。
- handled Fallback 表示安全降级答案,Run 仍为 SUCCESS;未处理、空答案或 invariant 失败才是 FAILED。
- Verifier/Gatekeeper 只以显式 Graph state 传递可信材料;继续使用 Hook/ThreadLocal 会形成双 Gatekeeper 和不安全输入,因此生产路径已彻底移除该依赖。
- 旧 Sequential 测试验证的是已废止内部机制,替换为 public service + runtime + route contract tests 才能保持真实回归价值。
@@ -0,0 +1,3 @@
committed_at: 2026-07-17
checkpoint: Commit
authorization: user-requested-direct-implementation
@@ -0,0 +1,135 @@
## Context
复杂 Chat 的公开调用链为 `ChatController -> ChatService.executeChatWithStrategy -> executeChatComplex`。阶段 2 已提供真实 Node adapters、显式 Gatekeeper、verified input、evidence retry、Composer/Fallback 和 `DiagnosisGraphFactory`,但 `executeChatComplex` 仍维护 SequentialAgent 两轮循环、VerifierInputHook/ThreadLocal 状态和私有 Composer 调用。阶段 3 要在不改变 `/api/chat` 的前提下切换唯一生产编排,并增加 Run 级 orchestration trace 数据契约。
当前 Run 通过 `SessionContextHolder(sessionId, runId)` 绑定 AgentStep/ToolInvocation;DiagnosisRun 使用字符串 JSON 保存 self evaluation;TraceService 将 Run 映射为 `run` 和兼容 `session` 投影。`DiagnosisOrchestrationTraceBuilder` 已能从有界 events 生成无 Prompt/raw output 的摘要。数据库由 Flyway 管理且 Hibernate 使用 validate,因此实体、V012 migration 和 Trace 映射必须同批对齐。
本 change 是 L4:内部状态机从固定顺序改为条件图,Trace run 对象和 DB 增加字段。`/api/chat`、sessionId/runId、Agent 输出协议保持兼容。用户已冻结单轨切换、安全 Fallback=SUCCESS、只有未处理失败=FAILED、阶段 5 才 live E2E。
## Goals / Non-Goals
**Goals:**
- 让复杂 Chat 一次执行真实 CompiledGraph,删除 ChatService 的细粒度 Sequential/round 状态机。
- 保持现有 Agent prompt/tool/skill/logging 组装能力,但 Graph Verifier 不注册 VerifierInputHook。
- 用 runId 同时作为 Graph threadId 和 Run 审计边界,用 sessionId/runId metadata 维持 AgentStep/ToolInvocation 归属。
- 将 final state 显式映射为 answer、verifier evaluation、composer audit、orchestration trace 和 Run lifecycle。
- 在正常/handled fallback 路径持久化非空、紧凑、安全的 orchestration trace,并只在 Trace `run` 对象暴露解析结果。
- 更新 Verifier Prompt/self-evaluation 到 verified-only 数据模型。
- 用阶段 3 focused tests证明生产 cutover、Run/Trace/API/Prompt 边界,且不留下已知失败测试。
**Non-Goals:**
- 不改简单 Chat/AIOps 编排。
- 不删除 VerifierInputHook/VerifierContextHolder 类型;它们在阶段 5 清理,但复杂 Chat 不再引用。
- 不增加持久 Graph checkpoint、恢复、并行分支或新 Run status。
- 不回填历史 run,不给历史 null orchestration trace 设计兼容伪值。
- 不在本阶段全面重命名/收敛测试夹具;阶段 4 完成测试体系替换。
- 不运行 Maven live E2E、日志或数据库验收。
## Decisions
### 1. 单轨生产 cutover,不增加 feature flag
`executeChatComplex` 只构造当前 Run、RunnableConfig 和 diagnosis context,然后调用真实 Graph runtime。旧 SequentialAgent loop、score/feature-flag retry 和 ChatService 私有 Verifier/Composer 路由方法从生产类移除;不保留双运行、shadow compare 或 runtime fallback 到 Sequential。
选择单轨而不是 feature flag,因为用户冻结的阶段门禁已经提供 Git commit 回滚边界,双轨会继续维护两套重试、Gatekeeper 和安全材料真理源。数据库新列 nullable,因此代码回滚时列可保留,无需破坏性 migration down。
### 2. Agent builder 只迁移职责,不复制实现
新增专用复杂 Chat Graph runtime/factory,复用当前 Planner/Executor/Verifier/Composer Prompt、knowledge map、history、method tools、ToolCallbacks、skill hooks 和 AgentLoggingHook 组装规则。ChatService 只向它提供本次 ChatModel、callbacks、history 和 Run 上下文。Graph Verifier hooks 只有 AgentLoggingHook,不包含 VerifierInputHook;Gatekeeper 只由显式 Node 调用。
如果为控制改动风险暂时保留简单 Chat 的工具/Hook builder,复杂 Agent builder 仍只存在一份。不得在 Graph Node 或 runtime 内复制 prompt 文本或第二套 tool catalog。
### 3. Graph 初始 state 和 RunnableConfig 使用最小白名单
初始 state 包含:
- `diagnosis_context={query, original_query}`;不把完整 history 放入 Graph State,history 仅注入 Planner/Executor system prompt。
- `planner_mode=NORMAL`。
- planner/verifier/composer/evidence retry count 全部为 0。
- `orchestration_events` 初始为空。
RunnableConfig 使用 `threadId(runId)`,metadata 包含 `sessionId` 和 `runId`。这同时满足 Graph 隔离、AgentLoggingHook/ToolCallback 的 Run ownership 和 Gatekeeper current-run validation。
### 4. 流式执行捕获最后真实 state
Graph runtime 使用 `CompiledGraph.stream(initialState, config)` 顺序消费 NodeOutput,并保存最后一个真实 `OverAllState`。正常 END 以最后 state 的非空 `final_answer` 为成功条件。流式消费与原同步 invoke 一样在当前请求线程阻塞,但允许在未处理异常时取得已真实产生的 partial state。
若异常前已有 events,failure handler 可以 best-effort 构造/persist partial trace;若没有 event,不生成虚假 node/transition。Agent invocation failures已经由 adapters 转为显式 status,正常预期失败都应走 Graph Fallback 而不是抛出。
### 5. Graph result mapper 是唯一运行结果翻译层
新增显式 mapper 从 final/partial state读取:
- `final_answer`。
- Verifier status/model/effective verdict、score、claim/fact checks、rationale、round。
- Gatekeeper raw audit、verified Executor output/evidence。
- Composer audit。
- `DiagnosisOrchestrationTraceBuilder` 结果。
ChatService 不再读取 VerifierContextHolder。self-evaluation 继续写 `verifier_evaluation` container 以兼容 Trace/Eval 消费者,但数据来自 Graph State:`executor_structured_output` 只保存 verified projection;新增 `verified_evidence` 和显式 statuses;不保存 raw Executor 或完整 tool_trace_summary。若 Verifier 从未完成,evaluation 只记录可用 status/audit/prompt 信息,不伪造 verdict。
### 6. Run lifecycle 和持久化顺序按执行结果分离
成功路径:Graph 返回非空安全 answer -> 构造 trace/evaluation -> 设置 Run SUCCESS、answer、orchestrationTrace、duration/metrics -> 保存 -> 调用 `evaluateRun`。
handled Fallback 与 Composer 正常路径使用相同成功顺序;可信度由 effective verdict 或 trace degraded 表达。失败路径:未处理异常、final state/answer 缺失、trace invariant 失败或成功结果持久化失败 -> Run FAILED,保存错误答案、duration、已有可构造 partial trace 和 metrics;不调用 success Eval。若 failure save 本身失败,只记录明确 error,不能声称持久化成功。
### 7. Orchestration trace 是独立 nullable JSON 数据契约
新增 V012,仅执行:
`ALTER TABLE diagnosis_run ADD COLUMN orchestration_trace JSON NULL ... AFTER self_evaluation`。
DiagnosisRun 使用 `@JdbcTypeCode(SqlTypes.JSON)` + `String orchestrationTrace`,与 selfEvaluation 写法一致。nullable 只服务无回填 migration/历史 run;每个成功的新 StateGraph Chat run 应用层必须非空。序列化使用 `DiagnosisOrchestrationTrace.toMap()`,字段保持 snake_case JSON。
### 8. Trace API 只在 RunTrace 增加解析对象
`DiagnosisTraceResponse.RunTrace` 新增 `Map<String,Object> orchestrationTrace`。DiagnosisTraceService 从 `DiagnosisRun.orchestrationTrace` 解析后映射;顶层 response、ChatSessionTrace、兼容 SessionTrace 和 raw string 均不新增字段。Run list 也不扩展该字段。
历史/AIOps run 的 null 值原样返回 null,不进行 fallback synthesis。解析非法 JSON 时 fail closed 为 null,并由数据/测试暴露问题;本阶段不增加历史兼容分支。
### 9. Verifier Prompt 与 verified-only input 对齐
`chat-verifier-prompt.md` 输入改为:`diagnosis_context`、`verified_executor_output`、`verified_evidence`、`gatekeeper_audit`、`verdict_ceiling`、可选 `retry_context`。删除 `executor_final_answer`、完整 `tool_trace_summary`、Hook parse status 和“再次执行 Gatekeeper”语义。
Verifier 仍输出原 JSON contract。`evidence_refs` 通过 claim_id/source_invocation_id/tool_name/raw_path 对 verified_evidence 建立审计关联,不要求不可用的 tool_trace_summary.trace_ref。Prompt audit catalog/version 随输入契约升级并在所有 handled paths 持久化。
### 10. 测试分阶段但不允许已知失败
阶段 3 新增最小生产 cutover测试:Graph runtime/final mapping、threadId/metadata、SUCCESS/Fallback/FAILED lifecycle、trace storage/API location、Prompt forbidden fields、no Verifier step on pre-verification failure。运行相关 Graph/Trace/Controller/Eval 回归和 test compilation。
旧 `ChatServiceSequentialAgentTest` 若因单轨语义失效,阶段 3 必须删除/改写冲突断言或由 focused Graph integration 替代,不能留到阶段 4 才让测试恢复绿色。阶段 4 继续完成命名、夹具、分支矩阵和旧 Hook tests 的全面清理。
## Interface Impact
- 等级:L4。
- 外部兼容:`/api/chat` request/response 与 run identity 不变。
- 加法协议:Trace `run.orchestrationTrace`;DB `diagnosis_run.orchestration_trace`。
- 内部行为:Executor/Planner/Gatekeeper failure 可提前 Fallback,不再保证固定 Agent 顺序;score flag 不再控制 LOW_CONFID retry。
- 消费者:Trace UI/demo/eval 可读取新 run 字段但不强制历史值;阶段 5 demo script 才增加最终 E2E 断言。
- 回滚:revert 本阶段代码/Prompt/spec;保留 nullable DB 列。无双轨开关、无数据回填回滚。
## Risks / Trade-offs
- [流式 API 与同步 invoke 的终止语义不同] → focused test断言最后 state、END、事件顺序和异常 partial state;不依赖实现私有字段。
- [self-evaluation 字段变化影响 Eval fixture] → 保留 container、verdict/claim/fact/gatekeeper/composer/prompt audit 兼容键;新增 verified 字段,移除不安全 full summary 前同步 specs/tests。
- [Agent factory 移动破坏 skill/tool hooks] → 复用现有 buildHooks/buildMethodTools 规则,并断言 Planner/Executor/Verifier/Composer hooks和当前 Run metadata。
- [JSON 字段三层漂移] → V012、entity、DTO/service 和 tests同批提交;Maven test compilation + strict specs。
- [无法取得 partial state] → 不伪造 trace;标记 FAILED 并记录明确原因。预期 Agent failures 全部由 Node adapters handled。
- [阶段 3/4 测试边界重叠] → 阶段 3只保证 cutover 可验收且现有 suite 不红;阶段 4负责全面测试架构替换。
## Migration Plan
1. 先加 V012、entity/Trace DTO/service 与 isolated mapping tests。
2. 实现 Graph runtime/result mapper和 Prompt verified-only 更新,先用 fake agents验证 final/partial state。
3. 将 `executeChatComplex` 单轨切换并删除旧私有 Sequential 状态机逻辑,保留 Run、metrics、Eval和简单 Chat路径。
4. 运行 production cutover、Graph、Trace、Controller/Eval 回归与 test compilation;证明 DB schema diff 只有一个字段。
5. 归档、提交阶段 3;阶段 4 再全面替换测试体系。
部署时 Flyway 先加 nullable 列,随后新代码写入。回滚为 Git revert;旧代码忽略新增列,数据库不删除列。
## Open Questions
无。所有产品/协议边界均由 ISS-011 与用户确认冻结;实现期若发现 Graph API 无法提供真实 partial state,只能回写本 design/tasks 后使用不伪造的 FAILED 处理,不能扩大协议。
@@ -0,0 +1,75 @@
# Chat Diagnosis StateGraph ChatService Cutover
## Why
阶段 2 已交付可构造、可测试的真实 Diagnosis StateGraph Nodes,但复杂 Chat 生产入口仍由 `ChatService.executeChatComplex(...)` 创建 `SequentialAgent`、维护 LOW_CONFID 外层循环,并依赖 `VerifierInputHook`/`VerifierContextHolder` 回传隐式状态。阶段 3 需要正式切换生产编排,让 ChatService 只管理 Run 生命周期、Agent 装配、Graph 调用和结果持久化,同时把 Graph 路由摘要作为当前 Run 的独立 Trace 维度保存和查询。
## What Changes
- 将复杂 Chat 从 `SequentialAgent` 外层循环切换为一次 `CompiledGraph` invocation,初始 state 使用显式 diagnosis context 和独立计数器。
- 为 Planner、Executor、Verifier、Composer 构造现有 ReactAgent 实例,经 `ReactAgentDiagnosisInvoker` 注入真实 Node actions;Graph Verifier 不注册旧 `VerifierInputHook`。
- 使用 `runId` 作为 Graph `threadId`,并把 `sessionId`、`runId` 放入 RunnableConfig metadata,继续复用 AgentStep/ToolInvocation 的 Run 归属。
- 将 Graph 最终 state 映射为现有 `ChatResult`、Verifier self-evaluation、Run 指标和生命周期状态。
- 新增 `diagnosis_run.orchestration_trace` nullable JSON 列、实体字段和紧凑序列化;新 StateGraph Chat run 在应用契约上必须写入非空摘要。
- Trace API 只在 `run.orchestrationTrace` 返回解析 JSON,不在顶层、兼容 `session` 投影或 raw 字段重复。
- 同步 Verifier Prompt 到 verified-only Graph 输入,移除完整 tool trace、raw Executor 和 Hook gatekeeper 输入说明。
- 增加生产 cutover、Run/Trace、safe Fallback、thread/metadata、Prompt 边界的必要单元/集成测试。
## Capabilities
### New Capabilities
- `chat-diagnosis-stategraph-chatservice-cutover`:规定复杂 Chat 的 StateGraph 生产调用、Run 生命周期、结果映射、编排摘要持久化与 Trace API 投影。
### Modified Capabilities
- `chat-diagnosis-stategraph-real-nodes`:阶段 2 的生产隔离约束改为阶段 3 正式切换,真实 Nodes 成为复杂 Chat 唯一生产编排。
- `chat-verifier-agent`:Verifier 输入改为 Gatekeeper passed binding 投影后的 verified-only payload;self-evaluation 以 Graph state 为来源。
- `chat-composer-agent`:Composer 由 Graph Node 调用,但 allowed-material 与审计契约保持不变。
- `session-run-trace-isolation`:Run 增加独立 orchestration trace,Trace API 只在精确 run 对象暴露解析结果。
## Scope
### In Scope
- `ChatService.executeChatComplex(...)` 的完整生产 cutover 和不再使用 SequentialAgent 的私有编排逻辑清理。
- 现有 Prompt/Agent builder 的 Graph 适配;不复制第二套 Agent factory。
- DiagnosisRun、Flyway、Trace DTO/Service 的加法式 orchestration trace 支持。
- Graph final state 到 answer、verifier evaluation、composer audit、Run status/metrics 的映射。
- 阶段 3 风险所需的 focused tests 与现有 Controller/Trace/Eval 契约回归。
### Out of Scope
- 不删除尚未被其他历史测试引用的 `VerifierInputHook`、`VerifierContextHolder` 类;阶段 5 统一清理。
- 不在阶段 3 完成整个测试体系命名/夹具迁移;阶段 4 负责全面替换旧 Sequential 测试体系,但阶段 3 不允许留下已知失败测试。
- 不修改 `/api/chat` 请求/响应结构或 Executor/Verifier/Composer 输出协议。
- 不新增 Run status,不做历史 run 的 orchestration trace 回填或兼容读取分支。
- 不运行 Maven live E2E、`logs/` 或数据库查询;统一保留到阶段 5。
## Context Constraints
- 阶段 0–2 archives 和主 specs 是实现基线;Graph 的 Node、条件边、计数所有权和安全 Fallback 边界不得在本阶段复制或放宽。
- `diagnosis_run.status` 只表达执行生命周期;Composer、LOW_CONFID、Verifier REJECT 或固定安全 Fallback 只要生成安全答案均为 SUCCESS。
- 编排摘要必须由有界 `orchestration_events` 构造,不从日志反推,不保存 Prompt、reasoning、raw tool output 或 Graph State snapshot。
- self-evaluation 与 orchestration trace 分离;AgentStep/ToolInvocation 继续作为详细 Trace 数据源。
- 新 Graph Verifier 只接收 `verified_executor_output`、`verified_evidence`、Gatekeeper audit/ceiling、query 和 retry context。
- 运行失败时不得伪造未发生的 events;可取得的部分 state 才允许 best-effort 持久化。
## Acceptance
- 复杂 Chat 生产代码不再创建或调用 SequentialAgent,且只调用一次 CompiledGraph。
- RunnableConfig 的 threadId 等于 runId,metadata 同时包含当前 sessionId/runId。
- Graph final answer 非空时映射为现有 ChatResult,Run 为 SUCCESS;无法产生安全响应或未处理失败时 Run 为 FAILED。
- 每个新 StateGraph Chat run 都持久化非空、无敏感材料的 orchestration trace;Trace API 只在 `run.orchestrationTrace` 返回解析对象。
- AgentStep、ToolInvocation、self-evaluation、answer、duration、token/step/tool count 和 Eval 仍绑定当前 runId。
- Verifier Prompt 和实际 verified-only payload 一致,Graph Verifier 不注册旧 input Hook。
- Executor/Gatekeeper 等前置失败的 Graph 路径不产生 Verifier AgentStep。
- `/api/chat` 和证据协议保持不变;数据库 schema 只新增一个 nullable JSON 列。
- 阶段 3 focused tests、Trace/Controller/Eval 回归、test compilation 和 OpenSpec strict validation 通过;不运行 live E2E。
## Risks
- ChatService 当前同时承载 Agent factory、Run 生命周期和旧解析逻辑;cutover 必须做完整内部重构,避免留下双编排路径或死代码。
- self-evaluation 旧字段曾依赖 ThreadLocal/full tool summary;Graph 映射必须明确兼容字段和安全边界,不能重新泄漏未验真材料。
- Graph invoke 异常可能没有可用 final state;失败处理只能持久化真实可取得的 partial snapshot,不能编造路径。
- JSON entity/DDL/Trace DTO 必须保持同一字段语义,否则 JPA validate、数据库迁移或 API 解析会漂移。
@@ -0,0 +1,60 @@
## MODIFIED Requirements
### Requirement: Composer SHALL generate final user-facing Chat answers
The system SHALL invoke the Composer Graph Node after Verifier routing to generate the final user-facing Chat answer from Verifier-allowed material.
#### Scenario: Composer receives only filtered material
- **WHEN** the Composer Graph Node invokes its configured Agent
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
- **AND** the Composer input SHALL NOT contain raw tool output
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
#### Scenario: Composer outputs strict JSON
- **WHEN** Composer completes
- **THEN** it SHALL output exactly one JSON object
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
#### Scenario: Composer does not introduce new facts
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
### Requirement: Composer input SHALL honor Verifier claim checks
The Composer Graph Node SHALL construct Composer input by filtering verified Executor structured output through Verifier `claim_checks`.
#### Scenario: Passing claims become allowed claims
- **WHEN** a claim check verification is `direct_observation`
- **THEN** the Composer input builder SHALL include the matching verified Executor claim in `allowed_claims`
#### Scenario: Reasonable inferences remain bounded
- **WHEN** a claim check verification is `reasonable_inference`
- **THEN** the Composer input builder MAY include the matching verified Executor claim in `allowed_claims`
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
#### Scenario: Overstated claims are not confirmed findings
- **WHEN** a claim check verification is `overstated`
- **THEN** the Composer input builder SHALL NOT include the matching claim as a confirmed item in `allowed_claims`
- **AND** it MAY include it as `allowed_hypotheses` or represent it in `missing_info`
#### Scenario: Unsupported or external claims are withheld
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
- **THEN** the Composer input builder SHALL NOT include the matching claim in `allowed_claims`
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
### Requirement: Composer failures SHALL degrade safely
The system SHALL tolerate malformed or failed Composer execution without leaking raw JSON or unverified Executor material.
#### Scenario: malformed Composer output falls back safely
- **WHEN** Composer exhausts its fixed-input technical retry or returns a non-retryable failure
- **THEN** the deterministic Fallback Node SHALL produce a final answer using only filtered material
- **AND** the final answer SHALL NOT expose raw Composer output
- **AND** the final answer SHALL NOT expose raw Executor output
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
#### Scenario: Composer audit is persisted
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.composer_output` SHALL record parsed Composer audit when available
- **AND** handled Composer fallback SHALL be observable through orchestration trace and available status/reason fields
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
@@ -0,0 +1,107 @@
## ADDED Requirements
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
#### Scenario: Complex Chat starts
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
#### Scenario: Graph dependencies are assembled
- **WHEN** the complex Chat Graph is constructed
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
### Requirement: Graph invocation SHALL preserve current Run ownership
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
#### Scenario: Agent Node runs
- **WHEN** any Agent Node is invoked for a complex Chat run
- **THEN** its RunnableConfig threadId SHALL equal the current runId
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
#### Scenario: Gatekeeper validates Executor output
- **WHEN** the Gatekeeper Node runs
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
- **AND** data from another run in the same session SHALL NOT be considered
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
#### Scenario: Initial state is projected
- **WHEN** a complex Chat run enters Planner for the first time
- **THEN** `diagnosis_context` SHALL contain the current query/original query
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
#### Scenario: Counters are initialized
- **WHEN** the Graph starts
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
- **AND** Planner mode SHALL be NORMAL
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
#### Scenario: Composer completes
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
- **AND** EvaluationService SHALL evaluate the current runId
#### Scenario: Handled Fallback completes
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
- **THEN** DiagnosisRun SHALL be SUCCESS
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
- **AND** no DEGRADED Run status SHALL be introduced
#### Scenario: Graph cannot produce a safe response
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
- **AND** the system SHALL NOT report a successful Graph result
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
#### Scenario: Verifier completed
- **WHEN** final Graph State contains a completed Verifier result
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
- **AND** the compatibility `verdict` field SHALL equal effective verdict
#### Scenario: Pre-verification Fallback completed
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
- **AND** it SHALL NOT fabricate a model or effective verdict
#### Scenario: Evaluation payload is inspected
- **WHEN** verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
### Requirement: Stage 3 verification SHALL not run the final live E2E
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
#### Scenario: Stage 3 is accepted
- **WHEN** stage 3 verification completes
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5
@@ -0,0 +1,7 @@
## REMOVED Requirements
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
**Reason**: Stage 2 production isolation has completed its migration purpose; stage 3 intentionally makes the real Diagnosis Graph the only complex Chat production orchestrator.
**Migration**: Replace `ChatService.executeChatComplex` SequentialAgent orchestration with real Graph assembly/invocation. Keep `/api/chat` compatible and use a full Git revert of stage 3 for runtime rollback rather than retaining dual paths.
@@ -0,0 +1,263 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical claims in `verified_executor_output.claims` have direct observation or reasonable inference support in `verified_evidence`
- **AND** at least one critical claim has direct observation
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
- **AND** verdict ceiling is PASS
- **THEN** the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical claim contradicts verified evidence
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only inference support
- **WHEN** no critical claim contradicts verified evidence
- **AND** all critical claims are only `reasonable_inference`
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical claim in `verified_executor_output.claims` contradicts verified evidence
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
- **THEN** the Verifier SHALL output model verdict="REJECT"
#### Scenario: Verified structured claims are the only verification target
- **WHEN** `verified_executor_output.claims` is present
- **THEN** Verifier SHALL verify each structured claim against matching `verified_evidence` through `claim_checks`
- **AND** each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
- **AND** a claim without matching verified evidence SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT receive or add confirmed facts from raw Executor text
#### Scenario: Executor output is invalid
- **WHEN** Executor does not return a legal structured contract
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute or fabricate a diagnostic verdict
### Requirement: facts_checked SHALL use a fixed classification set
The system SHALL continue to expose compatibility `facts_checked` using its fixed verification classification set.
#### Scenario: claim checks are mapped to legacy facts
- **WHEN** Verifier output contains `claim_checks`
- **THEN** the shared Verifier protocol parser SHALL derive compatibility `facts_checked` when the model did not provide them
- **AND** `direct_observation` SHALL map to `direct_evidence`
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
- **AND** `contradicted` SHALL map to `contradicted`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.
#### Scenario: PASS routes to Composer
- **WHEN** Verifier completes with effective verdict="PASS"
- **THEN** the Graph SHALL invoke Composer with filtered Verifier-allowed material
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
#### Scenario: LOW_CONFID does not qualify for evidence retry
- **WHEN** Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
- **THEN** the Graph SHALL route to Composer without another Planner cycle
- **AND** the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps
#### Scenario: LOW_CONFID qualifies for evidence retry
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
- **THEN** the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
- **AND** it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier completes with effective verdict="REJECT"
- **THEN** the Graph SHALL NOT start an evidence supplementation round
- **AND** it SHALL route to Composer-safe output
#### Scenario: REJECT produces bounded output
- **WHEN** effective verdict is REJECT
- **THEN** the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
- **AND** it SHALL NOT pass through raw Executor answer
- **AND** it SHALL NOT include an unsupported root-cause conclusion
#### Scenario: Verifier execution fails
- **WHEN** Verifier exhausts technical retry or returns a non-retryable failure
- **THEN** the Graph SHALL route to pre-verification Fallback
- **AND** no execution status string SHALL be used as model or effective verdict
### Requirement: Verifier SHALL be observable
The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.
#### Scenario: claim checks written to self_evaluation
- **WHEN** a completed Verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** it SHALL include `verifier_status`, `model_verdict`, `effective_verdict`, `verdict`, `groundedness_score`, `rationale`, verified output/evidence, and Gatekeeper audit
#### Scenario: composer output written to self_evaluation
- **WHEN** final answer composition completes
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include compact `composer_output` when available
- **AND** handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
- **AND** existing claim/fact and Gatekeeper fields SHALL be preserved
#### Scenario: verdict written to self_evaluation
- **WHEN** Verifier completes
- **THEN** Graph result mapping SHALL write effective verdict under `diagnosis_run.self_evaluation.verifier_evaluation.verdict`
- **AND** existing `rule_evaluation` and `aiops_rule_evaluation` channels SHALL be preserved
#### Scenario: pre-verification fallback is persisted
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
- **AND** it SHALL NOT fabricate `model_verdict` or `effective_verdict`
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** Graph result mapping persists available Gatekeeper state
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** a complex Chat Graph result is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_run.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat Prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
- **AND** full Prompt text SHALL NOT be persisted
#### Scenario: prompt audit available on fallback paths
- **WHEN** Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
#### Scenario: evaluation payload is inspected
- **WHEN** Graph verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete `tool_trace_summary`
- **AND** compatibility `executor_structured_output` SHALL contain at most the verified projection
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier Graph Node starts
- **THEN** the payload SHALL provide `diagnosis_context`, `verified_executor_output`, `verified_evidence`, `gatekeeper_audit`, and `verdict_ceiling`
- **AND** permitted structured `retry_context` SHALL be provided only after evidence retry preparation
#### Scenario: Verifier remains isolated from intermediate and raw material
- **WHEN** the Verifier input is serialized
- **THEN** it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State
#### Scenario: only passed bindings are available
- **WHEN** Gatekeeper returns mixed passed and failed checked bindings
- **THEN** `verified_executor_output` and `verified_evidence` SHALL contain only claims/material matching passed bindings
- **AND** the Verifier SHALL NOT receive failed or unreferenced tool material
#### Scenario: verified evidence preserves precise references
- **WHEN** the system prepares Verifier input
- **THEN** each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
- **AND** the item SHALL be traceable to current-run Gatekeeper validation
#### Scenario: gatekeeper audit and ceiling are available
- **WHEN** the system prepares Verifier input
- **THEN** the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
- **AND** a LOW_CONFID ceiling SHALL prevent effective PASS
#### Scenario: technical retry occurs
- **WHEN** the first Verifier attempt returns invalid output or a retryable invocation failure
- **THEN** the second attempt SHALL receive byte-identical serialized input
- **AND** Executor, Gatekeeper, and tools SHALL NOT rerun
### Requirement: Verifier facts SHALL be auditable
Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.
#### Scenario: claim checks contain evidence refs
- **WHEN** the Verifier emits `claim_checks`
- **THEN** each check SHALL include an `evidence_refs` array
- **AND** any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
- **AND** it SHALL NOT reference a failed or unverified binding
#### Scenario: verifier evaluation persists traceability snapshot
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the bounded `verified_evidence` snapshot used by the Verifier
- **AND** it SHALL NOT persist a complete tool trace summary as Verifier input
### Requirement: Structured Executor output SHALL degrade safely
The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.
#### Scenario: Malformed Executor JSON is classified
- **WHEN** Executor returns malformed JSON or text outside the expected contract
- **THEN** Executor Node SHALL set INVALID_OUTPUT
- **AND** the Graph SHALL route directly to deterministic pre-verification Fallback
- **AND** Gatekeeper, Verifier, and model Composer SHALL NOT execute
#### Scenario: Structured parse failure remains observable
- **WHEN** Executor output parsing fails
- **THEN** orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
- **AND** the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: schema rule rejects missing evidence bindings
- **WHEN** a confirmed claim has no `evidence_bindings`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references an invocation id absent from current-run `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing current-run invocation id
- **AND** binding `tool_name` does not match persisted invocation `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits legal `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
#### Scenario: gatekeeper reject bypasses Verifier
- **WHEN** normalized Gatekeeper status is REJECT
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute
#### Scenario: gatekeeper low confidence is bounded
- **WHEN** normalized Gatekeeper status is LOW_CONFID with at least one passed binding
- **THEN** verified input SHALL contain only passed bindings
- **AND** effective verdict SHALL NOT exceed LOW_CONFID
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.
#### Scenario: Verified excerpt supports direct observation
- **WHEN** verdict ceiling is PASS
- **AND** a claim's verified evidence matched text directly contains the claim's concrete facts
- **THEN** Verifier MAY classify that claim as `direct_observation`
#### Scenario: Verified evidence is complete Verifier context
- **WHEN** verified claims and evidence are available
- **THEN** Verifier SHALL use them as its evidence context
- **AND** it SHALL NOT require or request a complete tool trace summary
- **AND** it SHALL NOT read raw Executor or unreferenced tool material
### Requirement: Gatekeeper severity SHALL constrain effective verdict
Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.
#### Scenario: Reject severity bypasses Verifier
- **WHEN** `gatekeeper_result.severity=reject`
- **THEN** the Graph SHALL route to pre-verification Fallback without invoking Verifier
- **AND** it SHALL NOT fabricate an effective diagnostic verdict
#### Scenario: Low confidence severity prevents PASS
- **WHEN** `gatekeeper_result.severity=low_confid`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** deterministic effective-verdict calculation SHALL downgrade the result
- **AND** effective verdict SHALL be `LOW_CONFID`
#### Scenario: Gatekeeper audit includes severity
- **WHEN** Graph verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include available `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
@@ -0,0 +1,56 @@
## ADDED Requirements
### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`.
#### Scenario: Graph reaches Composer
- **WHEN** a complex Chat Graph terminates through Composer with a safe answer
- **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
- **AND** the summary SHALL be derived from the current Run's actual orchestration events
#### Scenario: Graph reaches handled Fallback
- **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer
- **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true`
- **AND** the Run status SHALL be SUCCESS
#### Scenario: Unhandled execution fails after events exist
- **WHEN** an unhandled failure occurs after one or more real Graph events are available
- **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run
- **AND** it SHALL NOT add a node or transition that did not occur
#### Scenario: Orchestration trace content is inspected
- **WHEN** orchestration trace JSON is serialized
- **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
- **AND** it SHALL NOT contain data owned by another run
### Requirement: Trace API SHALL expose orchestration trace only on the run object
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`.
#### Scenario: Exact StateGraph run trace is queried
- **WHEN** a caller queries a successful new StateGraph Chat run
- **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object
- **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field
- **AND** no raw orchestration trace field SHALL be added
#### Scenario: Historical or non-StateGraph run is queried
- **WHEN** the selected DiagnosisRun has null orchestration trace
- **THEN** `run.orchestrationTrace` MAY be null
- **AND** the service SHALL NOT synthesize historical events or read another run's trace
### Requirement: Orchestration trace migration SHALL be additive and nullable
The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change.
#### Scenario: Migration is applied
- **WHEN** Flyway applies the stage 3 migration
- **THEN** existing DiagnosisRun rows SHALL remain valid without backfill
- **AND** no other table or column SHALL be changed by the stage 3 schema migration
@@ -0,0 +1,43 @@
## 1. Run Orchestration Trace Contract
- [x] 1.1 Add V012 migration containing only nullable `diagnosis_run.orchestration_trace JSON` and document that historical rows are not backfilled.
- [x] 1.2 Add DiagnosisRun JSON field and `DiagnosisTraceResponse.RunTrace.orchestrationTrace` parsed-map field without adding top-level, session, summary, run-list, or raw duplicates.
- [x] 1.3 Update DiagnosisTraceService exact/latest Run mapping and tests for current-run parsed trace, historical null, invalid JSON fail-closed, and no cross-run/session projection.
- [x] 1.4 Add entity/repository or migration source checks proving stage 3 changes no schema object except the one nullable column.
## 2. Graph Runtime And Result Mapping
- [x] 2.1 Implement a dedicated complex Chat Graph runtime that assembles existing real actions, compiles the Graph once per request/runtime instance, and consumes `CompiledGraph.stream` while retaining only the last real state.
- [x] 2.2 Define minimal initial state with query-only diagnosis context, NORMAL mode, independent zero counters, and empty bounded events.
- [x] 2.3 Build RunnableConfig with `threadId=runId` and current `sessionId`/`runId` metadata; prove config identity at every real Agent invoker and Gatekeeper boundary.
- [x] 2.4 Implement explicit final/partial Graph result mapping for answer, statuses/verdicts, verified output/evidence, Gatekeeper audit, Composer audit, failure reason, and orchestration trace.
- [x] 2.5 Add runtime/mapper tests for Composer success, handled pre/post-verification Fallback, blank answer, no state, thrown execution with real partial events, and no fabricated event.
## 3. Agent Assembly And Verified-only Prompt
- [x] 3.1 Move or encapsulate the existing complex Planner/Executor/Verifier/Composer builders so Graph runtime reuses current Prompt, knowledge map, history, method tool, ToolCallback, skill, and AgentLoggingHook rules without a duplicate factory.
- [x] 3.2 Remove VerifierInputHook from the Graph Verifier while retaining AgentLoggingHook, and prove explicit Gatekeeper executes once per Executor round.
- [x] 3.3 Update `chat-verifier-prompt.md` to the verified-only input contract and remove raw Executor, full tool trace, Hook parse, and Gatekeeper re-execution language.
- [x] 3.4 Bump Prompt audit catalog/Verifier Prompt version and add static/behavior tests proving the runtime payload and Prompt declare the same allowed fields.
## 4. ChatService Production Cutover
- [x] 4.1 Replace `executeChatComplex` SequentialAgent/outer retry loop with one Graph runtime call while preserving session/run creation and the public ChatResult signature.
- [x] 4.2 Remove obsolete ChatService private Sequential workflow, score/feature-flag retry, ThreadLocal verifier parsing, private Composer routing, and dead imports without changing simple Chat behavior.
- [x] 4.3 Persist Graph-derived verifier evaluation with compatibility verdict/claim/fact/gatekeeper/composer/prompt-audit fields, verified-only evidence, and no fabricated verdict before Verifier completion.
- [x] 4.4 Persist SUCCESS only for a non-empty safe Graph answer with non-empty orchestration trace, then backfill current-run metrics and call Eval; persist FAILED for unhandled/no-answer/invariant failures with any real partial trace available.
- [x] 4.5 Ensure `SessionContextHolder` and retrieved-doc cleanup still run in finally, no complex Chat code reads `VerifierContextHolder`, and no production path creates SequentialAgent.
## 5. Stage 3 Focused Tests And Alignment
- [x] 5.1 Add production cutover integration tests for `/api/chat`-compatible ChatResult, `agent_flow=CHAT`, current Run ownership, Graph thread/metadata, answer/metrics/Eval persistence, and multi-run isolation.
- [x] 5.2 Add handled failure tests proving Executor invalid/failure and Gatekeeper REJECT do not create Verifier AgentStep, safe Fallback is SUCCESS/degraded, and unhandled/no-answer paths are FAILED.
- [x] 5.3 Update or replace stage-3-conflicting Sequential tests so no known test remains red; preserve Controller, Gatekeeper, Trace, Repository, Eval, no-evidence, REJECT, and Composer safety regressions.
- [x] 5.4 Confirm the first completed Trace slice and final production assembly against design/specs, including single-track cutover, no duplicate Agent factory, and no core TODO/placeholder.
## 6. Stage 3 Verification And Handoff
- [x] 6.1 Run Graph runtime/result mapper, ChatService cutover, DiagnosisTraceService, Controller, Repository, Gatekeeper, Composer, and Eval focused tests.
- [x] 6.2 Run Maven test compilation and any existing tests whose public contracts were touched; classify failures as OpenSpec gap, code deviation, or out-of-scope environment issue.
- [x] 6.3 Run OpenSpec strict validation, all main specs strict validation, `git diff --check`, source/reference isolation checks, and a schema whitelist check.
- [x] 6.4 Record that Maven live E2E, `logs/`, and `scripts/query_mysql.py` database verification were intentionally not run in stage 3 and remain reserved for stage 5.
+14 -14
View File
@@ -4,10 +4,10 @@
TBD - created by archiving change executor-composer-final-answer. Update Purpose after archive.
## Requirements
### Requirement: Composer SHALL generate final user-facing Chat answers
The system SHALL invoke a Composer expression layer after Verifier to generate the final user-facing Chat answer from Verifier-allowed material.
The system SHALL invoke the Composer Graph Node after Verifier routing to generate the final user-facing Chat answer from Verifier-allowed material.
#### Scenario: Composer receives only filtered material
- **WHEN** ChatService invokes Composer
- **WHEN** the Composer Graph Node invokes its configured Agent
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
- **AND** the Composer input SHALL NOT contain raw tool output
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
@@ -25,25 +25,25 @@ The system SHALL invoke a Composer expression layer after Verifier to generate t
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
### Requirement: Composer input SHALL honor Verifier claim checks
ChatService SHALL construct Composer input by filtering Executor structured output through Verifier `claim_checks`.
The Composer Graph Node SHALL construct Composer input by filtering verified Executor structured output through Verifier `claim_checks`.
#### Scenario: Passing claims become allowed claims
- **WHEN** a claim check verification is `direct_observation`
- **THEN** ChatService SHALL include the matching Executor claim in `allowed_claims`
- **THEN** the Composer input builder SHALL include the matching verified Executor claim in `allowed_claims`
#### Scenario: Reasonable inferences remain bounded
- **WHEN** a claim check verification is `reasonable_inference`
- **THEN** ChatService MAY include the matching Executor claim in `allowed_claims`
- **THEN** the Composer input builder MAY include the matching verified Executor claim in `allowed_claims`
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
#### Scenario: Overstated claims are not confirmed findings
- **WHEN** a claim check verification is `overstated`
- **THEN** ChatService SHALL NOT include the matching Executor claim as a confirmed item in `allowed_claims`
- **AND** ChatService MAY include it as `allowed_hypotheses` or represent it in `missing_info`
- **THEN** the Composer input builder SHALL NOT include the matching claim as a confirmed item in `allowed_claims`
- **AND** it MAY include it as `allowed_hypotheses` or represent it in `missing_info`
#### Scenario: Unsupported or external claims are withheld
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
- **THEN** ChatService SHALL NOT include the matching Executor claim in `allowed_claims`
- **THEN** the Composer input builder SHALL NOT include the matching claim in `allowed_claims`
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
### Requirement: Composer SHALL respect verdict-specific wording
@@ -67,17 +67,17 @@ Composer SHALL phrase final answers according to the effective Verifier verdict.
- **AND** the final answer SHALL NOT include a root-cause conclusion
### Requirement: Composer failures SHALL degrade safely
The system SHALL tolerate malformed Composer output without leaking raw JSON or unverified Executor material.
The system SHALL tolerate malformed or failed Composer execution without leaking raw JSON or unverified Executor material.
#### Scenario: malformed Composer output falls back safely
- **WHEN** Composer returns malformed JSON or omits required fields
- **THEN** ChatService SHALL produce a final answer using a fixed safe fallback template based only on filtered material
- **WHEN** Composer exhausts its fixed-input technical retry or returns a non-retryable failure
- **THEN** the deterministic Fallback Node SHALL produce a final answer using only filtered material
- **AND** the final answer SHALL NOT expose raw Composer output
- **AND** the final answer SHALL NOT expose raw Executor output
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
#### Scenario: Composer audit is persisted
- **WHEN** ChatService persists verifier evaluation
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.composer_output` SHALL record whether Composer output was valid or fallback was used
- **AND** the audit SHALL include the parsed Composer fields when valid
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.composer_output` SHALL record parsed Composer audit when available
- **AND** handled Composer fallback SHALL be observable through orchestration trace and available status/reason fields
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
@@ -0,0 +1,110 @@
# chat-diagnosis-stategraph-chatservice-cutover Specification
## Purpose
TBD - created by archiving change chat-diagnosis-stategraph-chatservice-cutover. Update Purpose after archive.
## Requirements
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
#### Scenario: Complex Chat starts
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
#### Scenario: Graph dependencies are assembled
- **WHEN** the complex Chat Graph is constructed
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
### Requirement: Graph invocation SHALL preserve current Run ownership
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
#### Scenario: Agent Node runs
- **WHEN** any Agent Node is invoked for a complex Chat run
- **THEN** its RunnableConfig threadId SHALL equal the current runId
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
#### Scenario: Gatekeeper validates Executor output
- **WHEN** the Gatekeeper Node runs
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
- **AND** data from another run in the same session SHALL NOT be considered
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
#### Scenario: Initial state is projected
- **WHEN** a complex Chat run enters Planner for the first time
- **THEN** `diagnosis_context` SHALL contain the current query/original query
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
#### Scenario: Counters are initialized
- **WHEN** the Graph starts
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
- **AND** Planner mode SHALL be NORMAL
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
#### Scenario: Composer completes
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
- **AND** EvaluationService SHALL evaluate the current runId
#### Scenario: Handled Fallback completes
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
- **THEN** DiagnosisRun SHALL be SUCCESS
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
- **AND** no DEGRADED Run status SHALL be introduced
#### Scenario: Graph cannot produce a safe response
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
- **AND** the system SHALL NOT report a successful Graph result
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
#### Scenario: Verifier completed
- **WHEN** final Graph State contains a completed Verifier result
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
- **AND** the compatibility `verdict` field SHALL equal effective verdict
#### Scenario: Pre-verification Fallback completed
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
- **AND** it SHALL NOT fabricate a model or effective verdict
#### Scenario: Evaluation payload is inspected
- **WHEN** verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
### Requirement: Stage 3 verification SHALL not run the final live E2E
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
#### Scenario: Stage 3 is accepted
- **WHEN** stage 3 verification completes
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5
@@ -123,13 +123,3 @@ Composer SHALL receive only effective verdict and Verifier-allowed claims, missi
- **WHEN** Composer fails after its one technical retry and Verifier-allowed material exists
- **THEN** deterministic Fallback SHALL express only the allowed claims, missing information, and recommendations
- **AND** it SHALL NOT read raw Executor or tool output
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
The real Node action set and CompiledGraph SHALL be constructible and testable, but ChatService, database, Trace API, shared Agent prompts, and the current production routing SHALL remain unchanged until the stage 3 change.
#### Scenario: Stage 2 production isolation is inspected
- **WHEN** this change is accepted
- **THEN** no production ChatService code path SHALL invoke the real Diagnosis Graph
- **AND** no database migration, Trace API field, or shared Prompt contract SHALL be changed
+146 -171
View File
@@ -4,41 +4,41 @@
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
## Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads structured Executor claims and the tool call history, then produces a structured verdict based on claim derivability.
The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical claims in `executor_structured_output.claims` have direct observation or reasonable inference support in tool call results
- **WHEN** all critical claims in `verified_executor_output.claims` have direct observation or reasonable inference support in `verified_evidence`
- **AND** at least one critical claim has direct observation
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
- **AND** `gatekeeper_result.status` is not `fail`
- **THEN** the Verifier MAY output verdict="PASS" with groundedness_score ≥ 0.5
- **AND** verdict ceiling is PASS
- **THEN** the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical claim contradicts the tool results
- **WHEN** no critical claim contradicts verified evidence
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only inference support
- **WHEN** no critical claim contradicts the tool results
- **WHEN** no critical claim contradicts verified evidence
- **AND** all critical claims are only `reasonable_inference`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical claim in `executor_structured_output.claims` contradicts tool call results
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
- **THEN** the Verifier SHALL output verdict="REJECT"
- **WHEN** any critical claim in `verified_executor_output.claims` contradicts verified evidence
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
- **THEN** the Verifier SHALL output model verdict="REJECT"
#### Scenario: Structured Executor claims are verified first
- **WHEN** `executor_structured_output.claims` is present and valid
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` through `claim_checks`
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: Verified structured claims are the only verification target
- **WHEN** `verified_executor_output.claims` is present
- **THEN** Verifier SHALL verify each structured claim against matching `verified_evidence` through `claim_checks`
- **AND** each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
- **AND** a claim without matching verified evidence SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT receive or add confirmed facts from raw Executor text
#### Scenario: Malformed structured output cannot pass through natural language fallback
- **WHEN** Executor does not return parseable structured output
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
#### Scenario: Executor output is invalid
- **WHEN** Executor does not return a legal structured contract
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute or fabricate a diagnostic verdict
### Requirement: Verifier SHALL output structured JSON
The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_checks array, compatibility facts_checked array, and rationale.
@@ -64,11 +64,11 @@ The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
### Requirement: facts_checked SHALL use a fixed classification set
The system SHALL continue to expose legacy `facts_checked` using its fixed verification classification set.
The system SHALL continue to expose compatibility `facts_checked` using its fixed verification classification set.
#### Scenario: claim checks are mapped to legacy facts
- **WHEN** Verifier output contains `claim_checks`
- **THEN** ChatService SHALL derive compatibility `facts_checked`
- **THEN** the shared Verifier protocol parser SHALL derive compatibility `facts_checked` when the model did not provide them
- **AND** `direct_observation` SHALL map to `direct_evidence`
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
@@ -89,73 +89,86 @@ The groundedness score SHALL be computed from critical fact classifications inst
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed.
The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.
#### Scenario: PASS -> Composer output
- **WHEN** Verifier outputs verdict="PASS"
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
#### Scenario: PASS routes to Composer
- **WHEN** Verifier completes with effective verdict="PASS"
- **THEN** the Graph SHALL invoke Composer with filtered Verifier-allowed material
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: LOW_CONFID does not qualify for evidence retry
- **WHEN** Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
- **THEN** the Graph SHALL route to Composer without another Planner cycle
- **AND** the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps
#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
#### Scenario: LOW_CONFID qualifies for evidence retry
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
- **THEN** the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
- **AND** it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL NOT start a retry round for evidence supplementation
- **AND** it SHALL produce a degraded output directly through Composer-safe rendering
- **WHEN** Verifier completes with effective verdict="REJECT"
- **THEN** the Graph SHALL NOT start an evidence supplementation round
- **AND** it SHALL route to Composer-safe output
#### Scenario: REJECT -> degraded output
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
- **AND** it SHALL NOT pass through the raw Executor answer
- **AND** it SHALL NOT include a root-cause conclusion
#### Scenario: REJECT produces bounded output
- **WHEN** effective verdict is REJECT
- **THEN** the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
- **AND** it SHALL NOT pass through raw Executor answer
- **AND** it SHALL NOT include an unsupported root-cause conclusion
#### Scenario: Verifier execution fails
- **WHEN** Verifier exhausts technical retry or returns a non-retryable failure
- **THEN** the Graph SHALL route to pre-verification Fallback
- **AND** no execution status string SHALL be used as model or effective verdict
### Requirement: Verifier SHALL be observable
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.
#### Scenario: claim checks written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **WHEN** a completed Verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
- **AND** it SHALL include `verifier_status`, `model_verdict`, `effective_verdict`, `verdict`, `groundedness_score`, `rationale`, verified output/evidence, and Gatekeeper audit
#### Scenario: composer output written to self_evaluation
- **WHEN** final answer composition completes
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output`
- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used
- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include compact `composer_output` when available
- **AND** handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
- **AND** existing claim/fact and Gatekeeper fields SHALL be preserved
#### Scenario: verdict written to self_evaluation
- **WHEN** the Verifier produces a verdict
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
- **WHEN** Verifier completes
- **THEN** Graph result mapping SHALL write effective verdict under `diagnosis_run.self_evaluation.verifier_evaluation.verdict`
- **AND** existing `rule_evaluation` and `aiops_rule_evaluation` channels SHALL be preserved
#### Scenario: pre-verification fallback is persisted
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
- **AND** it SHALL NOT fabricate `model_verdict` or `effective_verdict`
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
- **WHEN** Graph result mapping persists available Gatekeeper state
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** the Chat verifier evaluation is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
- **WHEN** a complex Chat Graph result is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_run.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat Prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
- **AND** full Prompt text SHALL NOT be persisted
#### Scenario: prompt audit available on fallback paths
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
- **WHEN** Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
#### Scenario: evaluation payload is inspected
- **WHEN** Graph verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete `tool_trace_summary`
- **AND** compatibility `executor_structured_output` SHALL contain at most the verified projection
### Requirement: self_evaluation SHALL be a container object
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
@@ -175,88 +188,51 @@ The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation ch
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier starts
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
- **AND** `retry_context` SHALL be provided on the second round only
- **AND** when Executor returns a valid evidence-attribution contract, the system SHALL provide `executor_structured_output`
- **AND** when Executor output parsing fails, the system SHALL provide an `executor_output_parse_status` that indicates the failure
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
- **WHEN** the Verifier Graph Node starts
- **THEN** the payload SHALL provide `diagnosis_context`, `verified_executor_output`, `verified_evidence`, `gatekeeper_audit`, and `verdict_ceiling`
- **AND** permitted structured `retry_context` SHALL be provided only after evidence retry preparation
#### Scenario: Verifier remains isolated from intermediate reasoning
- **WHEN** `executor_structured_output` is added to the verifier input
- **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, gatekeeper result, and retry context
#### Scenario: Verifier remains isolated from intermediate and raw material
- **WHEN** the Verifier input is serialized
- **THEN** it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State
#### Scenario: tool trace summary derived from tool facts
- **WHEN** the system prepares verifier inputs
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
#### Scenario: only passed bindings are available
- **WHEN** Gatekeeper returns mixed passed and failed checked bindings
- **THEN** `verified_executor_output` and `verified_evidence` SHALL contain only claims/material matching passed bindings
- **AND** the Verifier SHALL NOT receive failed or unreferenced tool material
#### Scenario: tool trace summary preserves invocation references
- **WHEN** the system prepares verifier inputs
- **THEN** each summary item SHALL include a stable `trace_ref`
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
#### Scenario: verified evidence preserves precise references
- **WHEN** the system prepares Verifier input
- **THEN** each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
- **AND** the item SHALL be traceable to current-run Gatekeeper validation
#### Scenario: only evidence-bearing tools included
- **WHEN** the system generates `tool_trace_summary`
- **THEN** it SHALL include only evidence-bearing tool invocations
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
#### Scenario: gatekeeper audit and ceiling are available
- **WHEN** the system prepares Verifier input
- **THEN** the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
- **AND** a LOW_CONFID ceiling SHALL prevent effective PASS
#### Scenario: failed evidence calls preserved as evidence gaps
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
- **THEN** the summary SHALL still include that invocation
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
#### Scenario: repeated tool calls may be compacted
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
- **THEN** the system MAY compact them into a merged summary entry
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
#### Scenario: raw outputs not passed through in full
- **WHEN** a tool invocation returns large raw content
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
#### Scenario: MessagesModelHook used only for noise reduction
- **WHEN** a MessagesModelHook is used for the Verifier
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
#### Scenario: gatekeeper result available to Verifier
- **WHEN** the system prepares verifier inputs from Executor output
- **THEN** the payload SHALL include `gatekeeper_result`
- **AND** `gatekeeper_result.status` SHALL be one of `pass`, `warn`, or `fail`
- **AND** `gatekeeper_result` SHALL include `failed_rules`, `warnings`, and `errors`
#### Scenario: structured claims are the primary verification target
- **WHEN** `executor_output_parse_status.status` is `valid`
- **AND** `executor_structured_output.claims` is available
- **THEN** Verifier SHALL verify each claim through `claim_checks`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: malformed structured output cannot pass through natural language fallback
- **WHEN** `executor_output_parse_status.status` is `missing` or `malformed`
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
#### Scenario: technical retry occurs
- **WHEN** the first Verifier attempt returns invalid output or a retryable invocation failure
- **THEN** the second attempt SHALL receive byte-identical serialized input
- **AND** Executor, Gatekeeper, and tools SHALL NOT rerun
### Requirement: Verifier facts SHALL be auditable
Verifier facts SHALL be linkable to the evidence summaries used during verification.
Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.
#### Scenario: facts_checked contains evidence refs
- **WHEN** the Verifier emits `facts_checked`
- **THEN** each fact SHALL include `evidence_refs`
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
#### Scenario: claim checks contain evidence refs
- **WHEN** the Verifier emits `claim_checks`
- **THEN** each check SHALL include an `evidence_refs` array
- **AND** any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
- **AND** it SHALL NOT reference a failed or unverified binding
#### Scenario: verifier evaluation persists traceability snapshot
- **WHEN** the ChatService persists `verifier_evaluation`
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
- **AND** it SHALL include the bounded `verified_evidence` snapshot used by the Verifier
- **AND** it SHALL NOT persist a complete tool trace summary as Verifier input
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
@@ -346,21 +322,21 @@ The system SHALL preserve a readable Chinese answer for normal Chat users even w
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
### Requirement: Structured Executor output SHALL degrade safely
The system SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow.
The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.
#### Scenario: Malformed Executor JSON is captured
#### Scenario: Malformed Executor JSON is classified
- **WHEN** Executor returns malformed JSON or text outside the expected contract
- **THEN** Chat runtime SHALL preserve the raw `executor_final_answer`
- **AND** it SHALL mark `executor_output_parse_status` as failed
- **AND** Verifier SHALL use natural-language fallback behavior
- **THEN** Executor Node SHALL set INVALID_OUTPUT
- **AND** the Graph SHALL route directly to deterministic pre-verification Fallback
- **AND** Gatekeeper, Verifier, and model Composer SHALL NOT execute
#### Scenario: Structured parse failure remains observable
- **WHEN** Executor output parsing fails
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract
- **THEN** orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
- **AND** the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
@@ -373,32 +349,31 @@ The system SHALL run deterministic Gatekeeper checks after Executor output parsi
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references a `source_invocation_ids` value that is not present in current-session `tool_invocation` rows
- **WHEN** a claim evidence binding references an invocation id absent from current-run `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing invocation id
- **AND** the binding `tool_name` does not match the invocation's persisted `tool_name`
- **WHEN** a claim evidence binding references an existing current-run invocation id
- **AND** binding `tool_name` does not match persisted invocation `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-session invocations with matching tool names
- **WHEN** Executor emits legal `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
#### Scenario: gatekeeper fail prevents PASS
- **WHEN** `gatekeeper_result.status` is `fail`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL NOT be `PASS`
#### Scenario: gatekeeper reject bypasses Verifier
- **WHEN** normalized Gatekeeper status is REJECT
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute
#### Scenario: invocation reference failure downgrades to reject
- **WHEN** `gatekeeper_result.failed_rules` contains `evidence.invocation_ref`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL set the effective verdict to `REJECT`
#### Scenario: gatekeeper low confidence is bounded
- **WHEN** normalized Gatekeeper status is LOW_CONFID with at least one passed binding
- **THEN** verified input SHALL contain only passed bindings
- **AND** effective verdict SHALL NOT exceed LOW_CONFID
### Requirement: Verifier claim checks SHALL use a fixed derivability classification set
The Verifier SHALL classify each structured claim using a fixed derivability classification set.
@@ -489,36 +464,36 @@ Gatekeeper SHALL validate that Executor evidence bindings point to real current-
- **AND** `failed_rules` SHALL include `evidence.excerpt_mismatch`
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
Verifier SHALL judge structured claims primarily against Gatekeeper-verified claim-local evidence excerpts.
Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.
#### Scenario: Verified excerpt supports direct observation
- **WHEN** `gatekeeper_result.severity=none`
- **AND** a claim's verified evidence excerpts directly contain the claim's concrete facts
- **WHEN** verdict ceiling is PASS
- **AND** a claim's verified evidence matched text directly contains the claim's concrete facts
- **THEN** Verifier MAY classify that claim as `direct_observation`
#### Scenario: Tool trace summary is navigation context
- **WHEN** `executor_structured_output.claims[].evidence_bindings` are available
- **THEN** Verifier SHALL use `tool_trace_summary` as navigation and audit context
- **AND** it SHALL NOT require `tool_trace_summary.output_summary` to contain every fact already present in verified claim-local evidence
#### Scenario: Verified evidence is complete Verifier context
- **WHEN** verified claims and evidence are available
- **THEN** Verifier SHALL use them as its evidence context
- **AND** it SHALL NOT require or request a complete tool trace summary
- **AND** it SHALL NOT read raw Executor or unreferenced tool material
### Requirement: Gatekeeper severity SHALL constrain effective verdict
Runtime effective verdict calculation SHALL treat Gatekeeper severity as a hard upper bound.
Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.
#### Scenario: Reject severity prevents PASS
#### Scenario: Reject severity bypasses Verifier
- **WHEN** `gatekeeper_result.severity=reject`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL be `REJECT`
- **THEN** the Graph SHALL route to pre-verification Fallback without invoking Verifier
- **AND** it SHALL NOT fabricate an effective diagnostic verdict
#### Scenario: Low confidence severity prevents PASS
- **WHEN** `gatekeeper_result.severity=low_confid`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL be `LOW_CONFID`
- **THEN** deterministic effective-verdict calculation SHALL downgrade the result
- **AND** effective verdict SHALL be `LOW_CONFID`
#### Scenario: Gatekeeper audit includes severity
- **WHEN** verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
- **WHEN** Graph verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include available `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
### Requirement: Verifier input hook SHALL only perform narrow compatibility backfill
The verifier input hook SHALL avoid converting broad tool summaries into precise evidence references.
@@ -3,9 +3,7 @@
## Purpose
Separate multi-turn conversation metadata from per-execution diagnosis state. `sessionId` identifies the conversation context, while `runId` identifies one replayable diagnosis execution and scopes Trace, Feedback, Evaluation, AIOps, and case-library provenance.
## Requirements
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
@@ -179,3 +177,58 @@ The change SHALL verify both runtime behavior and evaluation baseline impact.
- **WHEN** verification is complete
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression
### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`.
#### Scenario: Graph reaches Composer
- **WHEN** a complex Chat Graph terminates through Composer with a safe answer
- **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
- **AND** the summary SHALL be derived from the current Run's actual orchestration events
#### Scenario: Graph reaches handled Fallback
- **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer
- **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true`
- **AND** the Run status SHALL be SUCCESS
#### Scenario: Unhandled execution fails after events exist
- **WHEN** an unhandled failure occurs after one or more real Graph events are available
- **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run
- **AND** it SHALL NOT add a node or transition that did not occur
#### Scenario: Orchestration trace content is inspected
- **WHEN** orchestration trace JSON is serialized
- **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
- **AND** it SHALL NOT contain data owned by another run
### Requirement: Trace API SHALL expose orchestration trace only on the run object
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`.
#### Scenario: Exact StateGraph run trace is queried
- **WHEN** a caller queries a successful new StateGraph Chat run
- **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object
- **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field
- **AND** no raw orchestration trace field SHALL be added
#### Scenario: Historical or non-StateGraph run is queried
- **WHEN** the selected DiagnosisRun has null orchestration trace
- **THEN** `run.orchestrationTrace` MAY be null
- **AND** the service SHALL NOT synthesize historical events or read another run's trace
### Requirement: Orchestration trace migration SHALL be additive and nullable
The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change.
#### Scenario: Migration is applied
- **WHEN** Flyway applies the stage 3 migration
- **THEN** existing DiagnosisRun rows SHALL remain valid without backfill
- **AND** no other table or column SHALL be changed by the stage 3 schema migration
@@ -52,6 +52,10 @@ public class DiagnosisRun {
@Column(name = "self_evaluation", columnDefinition = "JSON")
private String selfEvaluation;
@JdbcTypeCode(SqlTypes.JSON)
@Column(name = "orchestration_trace", columnDefinition = "JSON")
private String orchestrationTrace;
@Column(name = "feedback", length = 16)
private String feedback;
@@ -88,4 +92,3 @@ public class DiagnosisRun {
updatedAt = LocalDateTime.now();
}
}
@@ -77,6 +77,7 @@ public class DiagnosisTraceResponse {
private String answer;
private String selfEvaluationRaw;
private Map<String, Object> selfEvaluation;
private Map<String, Object> orchestrationTrace;
private String feedback;
private LocalDateTime createdAt;
private LocalDateTime updatedAt;
@@ -0,0 +1,151 @@
package com.superbiz.agent.graph.diagnosis;
import com.alibaba.cloud.ai.graph.CompiledGraph;
import com.alibaba.cloud.ai.graph.NodeOutput;
import com.alibaba.cloud.ai.graph.OverAllState;
import com.alibaba.cloud.ai.graph.RunnableConfig;
import com.superbiz.agent.graph.diagnosis.DiagnosisGraphStatus.PlannerMode;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.concurrent.atomic.AtomicReference;
public class ChatDiagnosisGraphRuntime {
private final DiagnosisGraphFactory graphFactory;
private final DiagnosisOrchestrationTraceBuilder traceBuilder;
public ChatDiagnosisGraphRuntime() {
this(new DiagnosisGraphFactory(), new DiagnosisOrchestrationTraceBuilder());
}
ChatDiagnosisGraphRuntime(
DiagnosisGraphFactory graphFactory,
DiagnosisOrchestrationTraceBuilder traceBuilder) {
this.graphFactory = Objects.requireNonNull(graphFactory, "graphFactory");
this.traceBuilder = Objects.requireNonNull(traceBuilder, "traceBuilder");
}
public Result execute(
DiagnosisGraphActions actions,
String query,
String sessionId,
String runId) throws Exception {
requireText(query, "query");
requireText(sessionId, "sessionId");
requireText(runId, "runId");
CompiledGraph graph = graphFactory.compile(
Objects.requireNonNull(actions, "actions"));
AtomicReference<OverAllState> lastState = new AtomicReference<>();
try {
NodeOutput output = graph.stream(
initialState(query),
runnableConfig(sessionId, runId))
.doOnNext(value -> lastState.set(value.state()))
.blockLast();
if (output == null) {
throw new IllegalStateException(
"diagnosis graph returned no state");
}
OverAllState state = output.state();
String answer = DiagnosisGraphState.stringValue(
state, DiagnosisGraphState.FINAL_ANSWER);
if (answer == null) {
throw new IllegalStateException(
"diagnosis graph returned no final answer");
}
return new Result(state, answer, traceBuilder.build(state));
} catch (Throwable failure) {
OverAllState partialState = lastState.get();
throw new ExecutionFailure(
"diagnosis graph execution failed",
failure,
partialState,
partialTrace(partialState));
}
}
private DiagnosisOrchestrationTrace partialTrace(OverAllState state) {
if (state == null || DiagnosisGraphState.listValue(
state, DiagnosisGraphState.ORCHESTRATION_EVENTS).isEmpty()) {
return null;
}
try {
return traceBuilder.build(state);
} catch (RuntimeException invalidPartialState) {
return null;
}
}
private Map<String, Object> initialState(String query) {
return Map.of(
DiagnosisGraphState.DIAGNOSIS_CONTEXT,
Map.of("query", query, "original_query", query),
DiagnosisGraphState.PLANNER_MODE, PlannerMode.NORMAL.name(),
DiagnosisGraphState.PLANNER_RETRY_COUNT, 0,
DiagnosisGraphState.VERIFIER_RETRY_COUNT, 0,
DiagnosisGraphState.COMPOSER_RETRY_COUNT, 0,
DiagnosisGraphState.EVIDENCE_RETRY_COUNT, 0,
DiagnosisGraphState.ORCHESTRATION_EVENTS, List.of());
}
private RunnableConfig runnableConfig(String sessionId, String runId) {
return RunnableConfig.builder()
.threadId(runId)
.addMetadata("sessionId", sessionId)
.addMetadata("runId", runId)
.build();
}
private String requireText(String value, String name) {
if (value == null || value.isBlank()) {
throw new IllegalArgumentException(name + " must not be blank");
}
return value;
}
public record Result(
OverAllState state,
String answer,
DiagnosisOrchestrationTrace trace) {
public Result {
Objects.requireNonNull(state, "state");
answer = requireAnswer(answer);
Objects.requireNonNull(trace, "trace");
}
private static String requireAnswer(String value) {
if (value == null || value.isBlank()) {
throw new IllegalArgumentException("answer must not be blank");
}
return value;
}
}
public static final class ExecutionFailure extends Exception {
private final OverAllState partialState;
private final DiagnosisOrchestrationTrace partialTrace;
private ExecutionFailure(
String message,
Throwable cause,
OverAllState partialState,
DiagnosisOrchestrationTrace partialTrace) {
super(message, cause);
this.partialState = partialState;
this.partialTrace = partialTrace;
}
public OverAllState partialState() {
return partialState;
}
public DiagnosisOrchestrationTrace partialTrace() {
return partialTrace;
}
}
}
@@ -0,0 +1,116 @@
package com.superbiz.agent.graph.diagnosis;
import com.alibaba.cloud.ai.graph.OverAllState;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
public final class DiagnosisGraphResultMapper {
public Map<String, Object> verifierEvaluation(
OverAllState state,
Map<String, Object> promptAudit) {
Objects.requireNonNull(state, "state");
Objects.requireNonNull(promptAudit, "promptAudit");
Map<String, Object> evaluation = new LinkedHashMap<>();
putText(evaluation, "planner_status", state,
DiagnosisGraphState.PLANNER_STATUS);
putText(evaluation, "executor_status", state,
DiagnosisGraphState.EXECUTOR_STATUS);
putText(evaluation, "gatekeeper_status", state,
DiagnosisGraphState.GATEKEEPER_STATUS);
putText(evaluation, "verifier_status", state,
DiagnosisGraphState.VERIFIER_STATUS);
putText(evaluation, "composer_status", state,
DiagnosisGraphState.COMPOSER_STATUS);
putText(evaluation, "failure_reason", state,
DiagnosisGraphState.FAILURE_REASON);
String modelVerdict = DiagnosisGraphState.stringValue(
state, DiagnosisGraphState.VERIFIER_MODEL_VERDICT);
String effectiveVerdict = DiagnosisGraphState.stringValue(
state, DiagnosisGraphState.EFFECTIVE_VERDICT);
if (modelVerdict != null) {
evaluation.put("model_verdict", modelVerdict);
}
if (effectiveVerdict != null) {
evaluation.put("effective_verdict", effectiveVerdict);
evaluation.put("verdict", effectiveVerdict);
}
Map<String, Object> verifierOutput = DiagnosisGraphState.mapValue(
state, DiagnosisGraphState.VERIFIER_OUTPUT);
copyIfPresent(evaluation, verifierOutput,
"groundedness_score",
"critical_fact_count",
"claim_checks",
"facts_checked",
"rationale");
String executorStatus = DiagnosisGraphState.stringValue(
state, DiagnosisGraphState.EXECUTOR_STATUS);
if (executorStatus != null) {
evaluation.put("executor_output_parse_status",
"COMPLETED".equalsIgnoreCase(executorStatus)
? Map.of(
"status", "valid",
"detail", "graph executor completed")
: Map.of(
"status", "failed",
"detail", executorStatus));
}
putMapIfPresent(evaluation, "executor_structured_output", state,
DiagnosisGraphState.VERIFIED_EXECUTOR_OUTPUT);
if (!DiagnosisGraphState.listValue(
state, DiagnosisGraphState.VERIFIED_EVIDENCE).isEmpty()) {
evaluation.put("verified_evidence", DiagnosisGraphState.listValue(
state, DiagnosisGraphState.VERIFIED_EVIDENCE));
}
putMapIfPresent(evaluation, "gatekeeper_result", state,
DiagnosisGraphState.GATEKEEPER_RESULT);
putMapIfPresent(evaluation, "composer_output", state,
DiagnosisGraphState.COMPOSER_OUTPUT);
evaluation.put("round", DiagnosisGraphState.intValue(
state, DiagnosisGraphState.EVIDENCE_RETRY_COUNT) + 1);
evaluation.put("traceability_version", "v2");
evaluation.put("prompt_audit", Map.copyOf(promptAudit));
return evaluation;
}
private void putText(
Map<String, Object> target,
String targetKey,
OverAllState state,
String stateKey) {
String value = DiagnosisGraphState.stringValue(state, stateKey);
if (value != null) {
target.put(targetKey, value);
}
}
private void putMapIfPresent(
Map<String, Object> target,
String targetKey,
OverAllState state,
String stateKey) {
Map<String, Object> value = DiagnosisGraphState.mapValue(state, stateKey);
if (!value.isEmpty()) {
target.put(targetKey, value);
}
}
private void copyIfPresent(
Map<String, Object> target,
Map<String, Object> source,
String... keys) {
for (String key : keys) {
if (source.containsKey(key)) {
target.put(key, source.get(key));
}
}
}
}
@@ -3,7 +3,6 @@ package com.superbiz.agent.service;
import com.alibaba.cloud.ai.graph.OverAllState;
import com.alibaba.cloud.ai.graph.RunnableConfig;
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
import com.alibaba.cloud.ai.graph.agent.flow.agent.SequentialAgent;
import com.alibaba.cloud.ai.graph.agent.hook.Hook;
import com.alibaba.cloud.ai.graph.agent.hook.skills.SkillsAgentHook;
import com.alibaba.cloud.ai.graph.exception.GraphRunnerException;
@@ -13,18 +12,17 @@ import com.superbiz.agent.agent.tool.DateTimeTools;
import com.superbiz.agent.agent.tool.InternalDocsTools;
import com.superbiz.agent.agent.tool.QueryLogsTools;
import com.superbiz.agent.agent.tool.QueryMetricsTools;
import com.superbiz.agent.diagnosis.protocol.ComposerOutputParser;
import com.superbiz.agent.diagnosis.protocol.ComposerRenderResult;
import com.superbiz.agent.diagnosis.protocol.ComposerSafeInputBuilder;
import com.superbiz.agent.diagnosis.protocol.VerifierDecision;
import com.superbiz.agent.diagnosis.protocol.VerifierOutputParser;
import com.superbiz.agent.domain.entity.ChatSession;
import com.superbiz.agent.domain.entity.DiagnosisRun;
import com.superbiz.agent.graph.diagnosis.ChatDiagnosisGraphRuntime;
import com.superbiz.agent.graph.diagnosis.DiagnosisGraphActions;
import com.superbiz.agent.graph.diagnosis.DiagnosisGraphResultMapper;
import com.superbiz.agent.graph.diagnosis.DiagnosisOrchestrationTrace;
import com.superbiz.agent.graph.diagnosis.DiagnosisRealGraphActionsFactory;
import com.superbiz.agent.graph.diagnosis.ReactAgentDiagnosisInvoker;
import com.superbiz.agent.hook.AgentLoggingHook;
import com.superbiz.agent.hook.PlannerSkillMetadataHook;
import com.superbiz.agent.hook.TokenTrackingChatModel;
import com.superbiz.agent.hook.TokenUsageHolder;
import com.superbiz.agent.hook.VerifierInputHook;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.ChatSessionRepository;
import com.superbiz.agent.repository.DiagnosisRunRepository;
@@ -33,17 +31,14 @@ import com.superbiz.agent.tool.LookupKnowledgeTool;
import com.superbiz.agent.tool.RetrievedDocTracker;
import com.superbiz.agent.util.QuestionComplexity;
import com.superbiz.agent.util.SessionContextHolder;
import com.superbiz.agent.util.VerifierContextHolder;
import jakarta.annotation.PostConstruct;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import org.springframework.ai.chat.messages.AssistantMessage;
import org.springframework.ai.chat.model.ChatModel;
import org.springframework.ai.tool.ToolCallback;
import org.springframework.ai.tool.ToolCallbackProvider;
import org.springframework.beans.factory.annotation.Autowired;
import org.springframework.beans.factory.annotation.Value;
import org.springframework.core.io.ClassPathResource;
import org.springframework.stereotype.Service;
@@ -54,7 +49,6 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.UUID;
/**
@@ -65,7 +59,7 @@ import java.util.UUID;
public class ChatService {
private static final Logger logger = LoggerFactory.getLogger(ChatService.class);
private static final String CHAT_PROMPT_AUDIT_VERSION = "chat-prompts-v1";
private static final String CHAT_PROMPT_AUDIT_VERSION = "chat-prompts-v2";
/** 封装 answer + 后端生成的 sessionId/runId,用于 feedback 关联 */
public record ChatResult(String answer, String sessionId, String runId) {}
@@ -115,33 +109,24 @@ public class ChatService {
@Autowired(required = false)
private SkillRegistry skillRegistry;
@Autowired
private ToolTraceSummaryService toolTraceSummaryService;
@Autowired
private SelfEvaluationMergeService selfEvaluationMergeService;
@Autowired
private ExecutorGatekeeperService executorGatekeeperService;
@Value("${verifier.low-confidence-threshold:0.5}")
private double verifierLowConfidenceThreshold;
@Value("${chat.complex.retry-on-low-confidence:false}")
private boolean retryOnLowConfidence;
/** 多 Agent Chat 的 Prompt */
private String chatPlannerPrompt;
private String chatExecutorPrompt;
private String chatVerifierPrompt;
private String chatComposerPrompt;
private final ObjectMapper objectMapper = new ObjectMapper();
private final VerifierOutputParser verifierOutputParser =
new VerifierOutputParser(objectMapper);
private final ComposerSafeInputBuilder composerSafeInputBuilder =
new ComposerSafeInputBuilder();
private final ComposerOutputParser composerOutputParser =
new ComposerOutputParser();
private ChatDiagnosisGraphRuntime diagnosisGraphRuntime =
new ChatDiagnosisGraphRuntime();
private final DiagnosisRealGraphActionsFactory diagnosisActionsFactory =
new DiagnosisRealGraphActionsFactory();
private final DiagnosisGraphResultMapper diagnosisResultMapper =
new DiagnosisGraphResultMapper();
@PostConstruct
public void init() {
@@ -387,7 +372,7 @@ public class ChatService {
String question, List<Map<String, String>> history,
String requestedSessionId) throws GraphRunnerException {
if (QuestionComplexity.isComplex(question)) {
logger.info("问题判定为复杂,使用多 Agent(Planner + Executor)执行");
logger.info("问题判定为复杂,使用 Diagnosis StateGraph 执行");
return executeChatComplex(chatModel, toolCallbacks, question, history, requestedSessionId);
} else {
logger.info("问题判定为简单,使用单 Agent 执行");
@@ -398,7 +383,7 @@ public class ChatService {
}
/**
* 多 Agent 复杂对话执行(Planner -> Executor -> Verifier)
* Diagnosis StateGraph 复杂对话执行。
*/
public ChatResult executeChatComplex(ChatModel chatModel, ToolCallback[] toolCallbacks,
String question, List<Map<String, String>> history) throws GraphRunnerException {
@@ -411,135 +396,52 @@ public class ChatService {
String sessionId = resolveSessionId(requestedSessionId);
String runId = newRunId();
long startTime = System.currentTimeMillis();
List<Map<String, String>> safeHistory = history == null ? List.of() : history;
ensureChatSession(sessionId, history == null ? null : history.size() / 2);
ensureChatSession(sessionId, safeHistory.size() / 2);
DiagnosisRun run = startDiagnosisRun(sessionId, runId, question);
SessionContextHolder.setContext(sessionId, runId);
VerifierContextHolder.setOriginalQuery(question);
VerifierContextHolder.setRetryContext(null);
VerifierContextHolder.setExecutorFinalAnswer(null);
try {
VerifierDecision finalDecision = null;
String retryContext = null;
String answer = null;
RunnableConfig config = RunnableConfig.builder()
.addMetadata("sessionId", sessionId)
.addMetadata("runId", runId)
.build();
for (int round = 1; round <= 2; round++) {
VerifierContextHolder.setRetryContext(retryContext);
VerifierContextHolder.setToolTraceSummary(null);
ReactAgent planner = buildChatPlannerAgent(chatModel, history, retryContext);
ReactAgent executor = buildChatExecutorAgent(chatModel, toolCallbacks, history, retryContext);
ReactAgent verifier = buildChatVerifierAgent(chatModel);
SequentialAgent workflow = SequentialAgent.builder()
.name("chat_workflow")
.description("按固定顺序执行 Planner、Executor、Verifier 的多 Agent 工作流")
.subAgents(List.of(planner, executor, verifier))
.build();
String workflowInput = buildWorkflowInput(question, retryContext);
Optional<OverAllState> stateOptional = workflow.invoke(workflowInput, config);
if (stateOptional.isEmpty()) {
finalDecision = buildVerifierFallbackDecision(round, "workflow 未返回有效状态");
ComposerRenderResult renderResult = buildFixedFallbackAnswer(question, finalDecision);
answer = renderResult.answer();
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
break;
}
String plannerPlan = extractStateText(stateOptional, "planner_plan");
answer = extractStateText(stateOptional, "executor_feedback");
VerifierContextHolder.setExecutorFinalAnswer(answer);
String verifierOutput = extractStateText(stateOptional, "verifier_output");
if ((verifierOutput == null || verifierOutput.isBlank()) && answer != null && !answer.isBlank()) {
verifierOutput = invokeVerifierFallback(verifier, question, round, config);
}
finalDecision = parseVerifierDecision(verifierOutput, round);
logger.debug("Sequential workflow round {} finished: plannerPlanLength={}, answerLength={}, verifierOutputLength={}",
round,
plannerPlan != null ? plannerPlan.length() : 0,
answer != null ? answer.length() : 0,
verifierOutput != null ? verifierOutput.length() : 0);
if (finalDecision == null) {
finalDecision = buildVerifierFallbackDecision(round, "verifier_output 缺失或无法解析");
ComposerRenderResult renderResult = buildFixedFallbackAnswer(question, finalDecision);
answer = renderResult.answer();
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
break;
}
if ("PASS".equals(finalDecision.verdict())) {
ComposerRenderResult renderResult = composeFinalAnswer(chatModel, question, finalDecision, config);
answer = renderResult.answer();
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
break;
}
if ("REJECT".equals(finalDecision.verdict())) {
ComposerRenderResult renderResult = composeFinalAnswer(chatModel, question, finalDecision, config);
answer = renderResult.answer();
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
break;
}
boolean shouldRetry = retryOnLowConfidence
&& finalDecision.groundednessScore() < verifierLowConfidenceThreshold
&& round < 2;
if (!shouldRetry) {
ComposerRenderResult renderResult = composeFinalAnswer(chatModel, question, finalDecision, config);
answer = renderResult.answer();
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
break;
}
retryContext = buildRetryContext(finalDecision);
persistVerifierEvaluation(run, finalDecision, round);
}
long duration = System.currentTimeMillis() - startTime;
if (answer == null || answer.isBlank()) {
answer = "抱歉,多 Agent 分析未能生成有效结论。";
}
run.setStatus("SUCCESS");
run.setAnswer(answer);
run.setTotalDurationMs((int) duration);
backfillRunMetrics(run);
diagnosisRunRepository.save(run);
evaluationService.evaluateRun(runId, answer);
logger.info("多 Agent 总耗时: {} ms", duration);
logger.info("输出长度: {} 字符", answer.length());
return new ChatResult(answer, sessionId, runId);
} catch (Exception e) {
String errorAnswer = "Execution failed: " + e.getMessage();
run.setStatus("FAILED");
run.setAnswer(errorAnswer);
run.setTotalDurationMs((int) (System.currentTimeMillis() - startTime));
backfillRunMetrics(run);
diagnosisRunRepository.save(run);
logger.error("多 Agent 执行失败", e);
return new ChatResult(errorAnswer, sessionId, runId);
DiagnosisGraphActions actions = buildDiagnosisGraphActions(
chatModel, toolCallbacks, safeHistory);
ChatDiagnosisGraphRuntime.Result result = diagnosisGraphRuntime.execute(
actions, question, sessionId, runId);
persistGraphEvaluation(run, result.state());
persistOrchestrationTrace(run, result.trace());
return completeGraphRun(
run, result.answer(), sessionId, runId, startTime);
} catch (ChatDiagnosisGraphRuntime.ExecutionFailure failure) {
persistPartialGraphState(run, failure.partialState(), failure.partialTrace());
return failGraphRun(run, sessionId, runId, startTime, failure);
} catch (Exception failure) {
return failGraphRun(run, sessionId, runId, startTime, failure);
} finally {
retrievedDocTracker.clearSession(sessionId);
SessionContextHolder.clear();
VerifierContextHolder.clear();
}
}
private ReactAgent buildChatPlannerAgent(ChatModel chatModel, List<Map<String, String>> history,
String retryContext) {
private DiagnosisGraphActions buildDiagnosisGraphActions(
ChatModel chatModel,
ToolCallback[] toolCallbacks,
List<Map<String, String>> history) {
ReactAgent planner = buildChatPlannerAgent(chatModel, history);
ReactAgent executor = buildChatExecutorAgent(
chatModel, toolCallbacks, history);
ReactAgent verifier = buildChatVerifierAgent(chatModel);
ReactAgent composer = buildChatComposerAgent(chatModel);
return diagnosisActionsFactory.create(
new ReactAgentDiagnosisInvoker(planner),
new ReactAgentDiagnosisInvoker(executor),
new ReactAgentDiagnosisInvoker(verifier),
new ReactAgentDiagnosisInvoker(composer),
executorGatekeeperService);
}
private ReactAgent buildChatPlannerAgent(
ChatModel chatModel,
List<Map<String, String>> history) {
StringBuilder prompt = new StringBuilder(chatPlannerPrompt);
// 娉ㄥ叆 knowledge map
@@ -555,9 +457,6 @@ public class ChatService {
}
prompt.append("--- 对话历史结束 ---\n");
}
if (retryContext != null && !retryContext.isBlank()) {
prompt.append("\n\n--- 鏈疆琛ヨ瘉鎹害鏉?---\n").append(retryContext).append("\n");
}
return ReactAgent.builder()
.name("chat_planner")
.description("负责拆解问题、规划步骤")
@@ -574,8 +473,7 @@ public class ChatService {
.description("负责验证 Executor 答案的事实准确性")
.model(chatModel)
.systemPrompt(chatVerifierPrompt)
.hooks(new AgentLoggingHook(agentStepRepository, "verifier"),
new VerifierInputHook(toolTraceSummaryService, executorGatekeeperService))
.hooks(new AgentLoggingHook(agentStepRepository, "verifier"))
.outputKey("verifier_output")
.build();
}
@@ -591,13 +489,15 @@ public class ChatService {
.build();
}
private ReactAgent buildChatExecutorAgent(ChatModel chatModel, ToolCallback[] toolCallbacks,
List<Map<String, String>> history, String retryContext) {
private ReactAgent buildChatExecutorAgent(
ChatModel chatModel,
ToolCallback[] toolCallbacks,
List<Map<String, String>> history) {
StringBuilder prompt = new StringBuilder(chatExecutorPrompt);
prompt.append("\n\n--- Skill 读取约束 ---\n")
.append("如果 planner_plan 已给出 selected_skill,本轮 Executor 只允许对该 skill 调用一次 read_skill。")
.append("读取后必须复用已加载的 playbook 指令继续执行证据工具,不要为了检查支持文件、确认流程或生成报告再次读取同一个 skill。")
.append("只有 ChatService 启动新的补证据 retry round 时,才可以重新读取 selected_skill。\n");
.append("只有 Graph 进入新的 EVIDENCE_GAP_ONLY 阶段时,才可以在新 Executor 轮次重新读取 selected_skill。\n");
if (!history.isEmpty()) {
prompt.append("\n\n--- 对话历史 ---\n");
for (Map<String, String> msg : history) {
@@ -605,9 +505,6 @@ public class ChatService {
}
prompt.append("--- 对话历史结束 ---\n");
}
if (retryContext != null && !retryContext.isBlank()) {
prompt.append("\n\n--- 鏈疆琛ヨ瘉鎹害鏉?---\n").append(retryContext).append("\n");
}
return ReactAgent.builder()
.name("chat_executor")
.description("负责执行具体步骤并及时反馈")
@@ -682,99 +579,77 @@ public class ChatService {
return diagnosisRunRepository.save(run);
}
private String buildWorkflowInput(String question, String retryContext) {
StringBuilder input = new StringBuilder();
input.append("请按固定工作流完成本轮 Planner -> Executor -> Verifier。\n\n");
input.append("--- 用户问题 ---\n").append(question);
if (retryContext != null && !retryContext.isBlank()) {
input.append("\n\n--- retry_context ---\n").append(retryContext);
}
input.append("\n\nVerifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。");
return input.toString();
private void persistGraphEvaluation(
DiagnosisRun run,
OverAllState state) {
Map<String, Object> verifierEvaluation =
diagnosisResultMapper.verifierEvaluation(
state, promptAuditSnapshot());
run.setSelfEvaluation(selfEvaluationMergeService.mergeVerifierEvaluation(
run.getSelfEvaluation(), verifierEvaluation));
}
private String invokeVerifierFallback(ReactAgent verifier, String question, int round, RunnableConfig config) {
private void persistOrchestrationTrace(
DiagnosisRun run,
DiagnosisOrchestrationTrace trace) throws Exception {
run.setOrchestrationTrace(
objectMapper.writeValueAsString(trace.toMap()));
}
private void persistPartialGraphState(
DiagnosisRun run,
OverAllState partialState,
DiagnosisOrchestrationTrace partialTrace) {
try {
logger.warn("Sequential workflow round {} finished without verifier_output, invoking chat_verifier fallback", round);
return verifier.call("请基于 executor_final_answer 和 tool_trace_summary 输出 verifier JSON。原始问题:" + question, config)
.getText();
} catch (Exception e) {
logger.error("chat_verifier fallback 执行失败", e);
return null;
if (partialState != null) {
persistGraphEvaluation(run, partialState);
}
if (partialTrace != null) {
persistOrchestrationTrace(run, partialTrace);
}
} catch (Exception persistenceFailure) {
logger.warn("Failed to prepare partial diagnosis Graph state: runId={}",
run.getRunId(), persistenceFailure);
}
}
private VerifierDecision parseVerifierDecision(String verifierOutput, int round) {
VerifierDecision modelDecision = verifierOutputParser.parse(
verifierOutput, round);
if (modelDecision == null) {
logger.error("解析 verifier_output 失败: {}", verifierOutput);
return null;
}
Map<String, Object> parseStatus =
VerifierContextHolder.getExecutorOutputParseStatus();
String parseState = parseStatus == null
? ""
: String.valueOf(parseStatus.getOrDefault("status", ""));
return verifierOutputParser.applyCeiling(
modelDecision,
parseState,
VerifierContextHolder.getGatekeeperResult());
}
private VerifierDecision buildVerifierFallbackDecision(int round, String rationale) {
return verifierOutputParser.fallback(round, rationale);
}
private String extractStateText(Optional<OverAllState> stateOptional, String key) {
if (stateOptional.isEmpty()) {
return null;
}
return stateOptional.get().value(key)
.map(value -> {
if (value instanceof AssistantMessage assistantMessage) {
return assistantMessage.getText();
}
return String.valueOf(value);
})
.orElse(null);
}
private void persistVerifierEvaluation(DiagnosisRun run, VerifierDecision decision, int round) {
persistVerifierEvaluation(run, decision, round, null);
}
private void persistVerifierEvaluation(DiagnosisRun run, VerifierDecision decision, int round,
Map<String, Object> composerOutput) {
if (decision == null) {
return;
}
Map<String, Object> verifierEvaluation = new LinkedHashMap<>();
verifierEvaluation.put("verdict", decision.verdict());
verifierEvaluation.put("groundedness_score", decision.groundednessScore());
verifierEvaluation.put("critical_fact_count", decision.criticalFactCount());
verifierEvaluation.put("claim_checks", decision.claimChecks());
verifierEvaluation.put("facts_checked", decision.factsChecked());
verifierEvaluation.put("rationale", decision.rationale());
verifierEvaluation.put("round", round);
verifierEvaluation.put("traceability_version", "v1");
verifierEvaluation.put("prompt_audit", promptAuditSnapshot());
verifierEvaluation.put("executor_output_parse_status",
Optional.ofNullable(VerifierContextHolder.getExecutorOutputParseStatus())
.orElse(Map.of("status", "missing", "detail", "executor parse status unavailable")));
verifierEvaluation.put("executor_structured_output", VerifierContextHolder.getExecutorStructuredOutput());
verifierEvaluation.put("tool_trace_summary",
Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of()));
verifierEvaluation.put("gatekeeper_result",
Optional.ofNullable(VerifierContextHolder.getGatekeeperResult())
.orElse(defaultGatekeeperPass()));
if (composerOutput != null) {
verifierEvaluation.put("composer_output", composerOutput);
}
String merged = selfEvaluationMergeService.mergeVerifierEvaluation(run.getSelfEvaluation(), verifierEvaluation);
run.setSelfEvaluation(merged);
private ChatResult completeGraphRun(
DiagnosisRun run,
String answer,
String sessionId,
String runId,
long startTime) {
long duration = System.currentTimeMillis() - startTime;
run.setStatus("SUCCESS");
run.setAnswer(answer);
run.setTotalDurationMs((int) duration);
backfillRunMetrics(run);
diagnosisRunRepository.save(run);
evaluationService.evaluateRun(runId, answer);
logger.info("Diagnosis Graph 总耗时: {} ms", duration);
logger.info("输出长度: {} 字符", answer.length());
return new ChatResult(answer, sessionId, runId);
}
private ChatResult failGraphRun(
DiagnosisRun run,
String sessionId,
String runId,
long startTime,
Exception failure) {
String errorAnswer = "Execution failed: " + failure.getMessage();
run.setStatus("FAILED");
run.setAnswer(errorAnswer);
run.setTotalDurationMs((int) (System.currentTimeMillis() - startTime));
backfillRunMetrics(run);
try {
diagnosisRunRepository.save(run);
} catch (Exception persistenceFailure) {
logger.error("Diagnosis Graph failure state persistence failed: runId={}",
runId, persistenceFailure);
}
logger.error("Diagnosis Graph 执行失败", failure);
return new ChatResult(errorAnswer, sessionId, runId);
}
private Map<String, Object> promptAuditSnapshot() {
@@ -783,7 +658,7 @@ public class ChatService {
audit.put("prompts", List.of(
promptAuditItem("chat_planner", "chat-planner-v1", "prompts/chat-planner-prompt.md"),
promptAuditItem("chat_executor", "chat-executor-v2", "prompts/chat-executor-prompt.md"),
promptAuditItem("chat_verifier", "chat-verifier-v2", "prompts/chat-verifier-prompt.md"),
promptAuditItem("chat_verifier", "chat-verifier-v3", "prompts/chat-verifier-prompt.md"),
promptAuditItem("chat_composer", "chat-composer-v1", "prompts/chat-composer-prompt.md")
));
return audit;
@@ -797,114 +672,6 @@ public class ChatService {
return item;
}
private Map<String, Object> defaultGatekeeperPass() {
GatekeeperRuleCatalog catalog = GatekeeperRuleCatalog.fallback();
return Map.of(
"status", "pass",
"severity", "none",
"rule_set_version", catalog.version(),
"rules", catalog.auditRules(),
"checked_bindings", List.of(),
"failed_rules", List.of(),
"warnings", List.of(),
"errors", List.of()
);
}
private ComposerRenderResult composeFinalAnswer(ChatModel chatModel, String originalQuery,
VerifierDecision decision, RunnableConfig config) {
Map<String, Object> composerInput = buildComposerInput(originalQuery, decision);
try {
ReactAgent composer = buildChatComposerAgent(chatModel);
String composerOutput = composer.call(objectMapper.writeValueAsString(composerInput), config).getText();
return parseComposerOutput(composerOutput, composerInput);
} catch (Exception e) {
logger.error("chat_composer 执行失败,使用安全降级模板", e);
return buildFixedFallbackAnswer(composerInput, "composer_exception");
}
}
private ComposerRenderResult buildFixedFallbackAnswer(String originalQuery, VerifierDecision decision) {
return buildFixedFallbackAnswer(buildComposerInput(originalQuery, decision), "fixed_fallback");
}
private Map<String, Object> buildComposerInput(String originalQuery, VerifierDecision decision) {
return composerSafeInputBuilder.build(
originalQuery,
decision,
VerifierContextHolder.getExecutorStructuredOutput());
}
private ComposerRenderResult parseComposerOutput(String composerOutput, Map<String, Object> composerInput) {
return composerOutputParser.parse(
composerOutput,
composerInput,
buildNextStepSuggestionsFromTrace());
}
private ComposerRenderResult buildFixedFallbackAnswer(Map<String, Object> composerInput, String status) {
return composerOutputParser.fallback(
composerInput,
buildNextStepSuggestionsFromTrace(),
status);
}
private List<String> buildNextStepSuggestionsFromTrace() {
List<String> suggestions = new ArrayList<>();
String runId = SessionContextHolder.getRunId();
List<Map<String, Object>> toolSummary = runId == null || runId.isBlank()
? toolTraceSummaryService.buildVerifierTraceSummary(SessionContextHolder.getSessionId(), null)
: toolTraceSummaryService.buildVerifierTraceSummaryForRun(runId, null);
boolean hasKnowledgeTool = toolSummary.stream().anyMatch(item -> "lookup_knowledge".equals(item.get("tool_name")));
boolean hasFailedEvidence = toolSummary.stream().anyMatch(item -> !Boolean.TRUE.equals(item.get("success")));
if (!hasKnowledgeTool) {
suggestions.add("补充知识库或业务文档检索结果,建立可引用的证据锚点");
}
if (hasFailedEvidence) {
suggestions.add("优先重试失败的证据型查询,补齐日志、指标或知识库侧证据");
}
if (suggestions.isEmpty()) {
suggestions.add("围绕上述证据缺口补充只读查询,再由人工复核最终结论");
}
return suggestions;
}
private String buildRetryContext(VerifierDecision decision) {
try {
List<String> missingFacts = extractEvidenceGaps(decision);
Map<String, Object> retryContext = new LinkedHashMap<>();
retryContext.put("round", decision.round());
retryContext.put("missing_evidence_facts", missingFacts);
retryContext.put("instruction", "仅补充以上断言相关证据,不要重复已完成检索");
return objectMapper.writeValueAsString(retryContext);
} catch (Exception e) {
logger.error("构造 retry_context 失败", e);
return "{\"round\":1,\"missing_evidence_facts\":[],\"instruction\":\"仅补充缺失证据\"}";
}
}
private List<String> extractEvidenceGaps(VerifierDecision decision) {
List<String> gaps = new ArrayList<>();
for (Map<String, Object> fact : decision.factsChecked()) {
String verification = String.valueOf(fact.get("verification"));
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
if (critical && ("no_evidence".equals(verification) || "contradicted".equals(verification))) {
gaps.add(String.valueOf(fact.get("fact")) + ":" + String.valueOf(fact.get("detail")));
}
}
if (gaps.isEmpty() && "LOW_CONFID".equals(decision.verdict())) {
for (Map<String, Object> fact : decision.factsChecked()) {
String verification = String.valueOf(fact.get("verification"));
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
if (critical && "indirect_support".equals(verification)) {
gaps.add(String.valueOf(fact.get("fact")) + ":缺少直接证据锚点");
}
}
}
return gaps;
}
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_run */
private void backfillRunMetrics(DiagnosisRun run) {
try {
@@ -181,6 +181,7 @@ public class DiagnosisTraceService {
.answer(run.getAnswer())
.selfEvaluationRaw(run.getSelfEvaluation())
.selfEvaluation(parseJsonObject(run.getSelfEvaluation()))
.orchestrationTrace(parseJsonObject(run.getOrchestrationTrace()))
.feedback(run.getFeedback())
.createdAt(run.getCreatedAt())
.updatedAt(run.getUpdatedAt())
@@ -0,0 +1,7 @@
-- V012: add nullable StateGraph orchestration summary to diagnosis runs.
-- Historical rows are intentionally not backfilled.
ALTER TABLE diagnosis_run
ADD COLUMN orchestration_trace JSON NULL
COMMENT 'Compact run-scoped StateGraph orchestration summary'
AFTER self_evaluation;
@@ -1,48 +1,30 @@
你是质量闸 verifier。你的任务是对 Executor 的结构化 claims 做一次基于现有证据的可推导性校验。
你是质量闸 Verifier。你的任务是判断 Gatekeeper 已验真的 Executor claims 是否能由对应证据推出。
边界约束:
- 不做新的检索
- 不做超出输入证据的推理扩写
- 不补充输入中不存在的新事实
- 只输出一个合法 JSON 对象,不输出 Markdown,不输出代码块,不输出额外说明
你不调用工具,不做新检索,不补充输入外事实,不读取或猜测任何未经验真的材料。你必须只输出一个合法 JSON 对象,不输出 Markdown、代码块或额外说明。
## 输入字段
## 输入契约
- `original_query`:用户原始问题
- `executor_final_answer`:Executor 原始输出,仅用于 debug/fallback;当结构化输出有效时,不得从这里抽取额外确认事实
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`;兼容旧版时可能包含 `user_facing_answer`
- `executor_output_parse_status`:Executor 输出解析状态,包含 `status` 和 `detail`。`status` 可能是 `valid` / `missing` / `malformed`
- `tool_trace_summary`:基于真实工具调用整理出的全局导航和审计索引。它不是唯一证据源;当 claim 有已核验的 `evidence_bindings[].evidence_excerpt` 时,应优先使用 claim-local excerpt 判断可推导性。每一项都带有:
- `trace_ref`
- `tool_name`
- `topic_domain`
- `source_invocation_ids`
- `input_summary`
- `output_summary`
- `evidence_level`
- `gatekeeper_result`:Executor 结构化输出的确定性校验结果,包含 `status`、`severity`、`checked_bindings`、`failed_rules`、`warnings`、`errors`
- `retry_context`:第二轮可选输入;若为空,按首轮处理
- `diagnosis_context`:只包含当前用户 query/original query 的诊断上下文。
- `verified_executor_output`:只包含 Gatekeeper passed bindings 对应的 Executor claims。
- `verified_evidence`:已验真 claim-local evidence;每项包含 `claim_id`、`source_invocation_id`、`tool_name`、`raw_path`、`matched_text`。
- `gatekeeper_audit`:当前 Run 的 Gatekeeper 审计结果,用于解释 ceiling 和失败规则,不得从 failed binding 提取事实。
- `verdict_ceiling`:确定性代码给出的最大有效 verdict,只允许 `PASS` 或 `LOW_CONFID`。
- `retry_context`:可选的结构化补证据上下文;只用于理解本轮 gap,不是事实证据。
## 任务步骤
除以上字段外,不得要求或使用父 Graph State、Prompt 文本、模型思考、完整工具历史、未引用工具结果或任何原始自由文本。
### 步骤一:确定校验对象
如果 `executor_output_parse_status.status="valid"` 且 `executor_structured_output.claims` 存在:
- 优先逐条校验 `executor_structured_output.claims`
- 每个 claim 至少形成一条 `claim_checks`
- 如果 `gatekeeper_result.severity="none"`,将 claim 的 `evidence_bindings[].evidence_excerpt` 视为已通过代码核验的主证据,判断 `claim_text` 是否能由这些 excerpt 推出
- `tool_trace_summary` 只用于理解工具调用全貌、补充 trace_ref、识别 no_evidence gap,不要求它逐字包含 excerpt 中已经核验过的全部事实
- 不得从 `executor_final_answer` 中抽取不在 claims 里的额外确认事实
## 校验步骤
如果 structured output 缺失或 malformed:
- 不得通过扫描 `executor_final_answer` 生成 `PASS`
- 输出 `LOW_CONFID`
- `groundedness_score = 0.0`
- `claim_checks = []`
- `facts_checked = []`
- `rationale` 说明结构化输出不可用
### 1. 建立精确关联
逐条读取 `verified_executor_output.claims`。每个 claim 只能使用 `verified_evidence` 中 claim_id 一致,且 source_invocation_id/tool_name/raw_path 与该 claim binding 匹配的 `matched_text`。
无法建立匹配的 claim 不得判为 `direct_observation`。
### 2. 输出 claim_checks
每条 claim 输出一条 claim check,字段为:
### 步骤二:逐条校验 claim
每条 claim check 必须输出:
- `claim_id`
- `claim_text`
- `claim_type`
@@ -50,156 +32,92 @@
- `detail`
- `evidence_refs`
`claim_checks[*].verification` 只允许以下六个值:
- `direct_observation`
- `reasonable_inference`
- `overstated`
- `unsupported`
- `external_unknown`
- `contradicted`
`verification` 只允许:
结构化 claim 的校验规则:
- claim 有 Gatekeeper 核验通过的 evidence binding,且 `evidence_excerpt` 直接包含该事实 → `direct_observation`
- claim 有 Gatekeeper 核验通过的 evidence binding,excerpt 没有逐字说明但可以合理推出 → `reasonable_inference`
- claim 有部分依据,但写成唯一根因、确认根因或说得过满 → `overstated`
- claim 无法绑定真实 trace、invocation 或 excerpt → `unsupported`
- claim 引入证据外的新服务名、订单号、错误码、指标值、根因 → `external_unknown`
- claim 与工具摘要冲突 → `contradicted`
- `direct_observation`:matched_text 直接包含 claim 的具体事实。
- `reasonable_inference`:matched_text 没有逐字陈述完整 claim,但可在不引入新事实的前提下合理推出。
- `overstated`:有部分依据,但 claim 写成唯一/确认根因或表达过满。
- `unsupported`:没有足够匹配证据。
- `external_unknown`:claim 引入证据外的实体、错误码、指标值或结论。
- `contradicted`:claim 与 matched_text 明确冲突。
`hypotheses` 和 `missing_info` 默认不是 confirmed facts,不应因为它们承认缺证据而惩罚。
`evidence_refs` 只能引用实际存在的 verified evidence,字段使用:
### 步骤三:补齐 evidence_refs
`evidence_refs` 必须是数组,数组元素必须引用 `tool_trace_summary` 中真实存在的证据项。每个元素包含:
- `trace_ref`
- `claim_id`
- `source_invocation_id`
- `tool_name`
- `topic_domain`
- `source_invocation_ids`
- `raw_path`
- `note`
规则:
- 有证据支撑时,必须引用支撑该事实的证据项
- `no_evidence` 并不等于不引用
- 如果工具确实查过相关方向,但证据不够,仍应引用对应 trace,并在 `note` 里说明“不足以支撑”
- 只有当确实找不到相关 trace 时,`evidence_refs` 才允许为空数组
- 不允许编造不存在的 `trace_ref` 或 `source_invocation_ids`
不得编造引用,不得引用 Gatekeeper failed binding。
### 步骤四:生成 verdict
严格使用以下判定矩阵:
0. 若 `gatekeeper_result.status="fail"`
- 不得输出 `PASS`
- 若 `gatekeeper_result.severity="reject"`,输出 `REJECT`
- 若 `gatekeeper_result.severity="low_confid"`,输出 `LOW_CONFID`
- 兼容旧输入:若缺少 `severity` 且 `failed_rules` 包含 `evidence.invocation_ref`,倾向 `REJECT`
- 兼容旧输入:若缺少 `severity` 且不是明显伪造,至少输出 `LOW_CONFID`
### 3. 计算 model verdict
1. 若任一关键 claim 为 `contradicted`
- `verdict = "REJECT"`
- `groundedness_score = 0.0`
- 任一关键 claim 为 `contradicted`:`REJECT`,groundedness_score=0.0。
- 所有关键 claims 均为 `direct_observation`/`reasonable_inference`,且至少一条为 direct:`PASS`。
- 存在 `unsupported`/`external_unknown`/`overstated`,或所有关键 claims 只有 inference:`LOW_CONFID`。
- 没有可校验关键 claim:`LOW_CONFID`,groundedness_score=0.0。
2. 否则,若所有关键 claims 均为 `direct_observation` 或 `reasonable_inference`
且至少一条关键 claim 为 `direct_observation`
- `verdict = "PASS"`
`verdict_ceiling=LOW_CONFID` 时,你的 model verdict 仍按证据给出;确定性代码会把 effective verdict 限制为 LOW_CONFID。不要绕过 ceiling,也不要把执行状态写成 verdict。
3. 否则,若不存在 `contradicted`
且存在关键 claim 为 `unsupported` / `external_unknown` / `overstated`
或所有关键 claim 都只有 `reasonable_inference`
- `verdict = "LOW_CONFID"`
### 4. 计算 groundedness_score
### 步骤五:计算 groundedness_score
只统计关键 claim,映射如下:
- `direct_observation = 1.0`
- `reasonable_inference = 0.6`
- `overstated = 0.3`
- `unsupported = 0.0`
- `external_unknown = 0.0`
- `contradicted = 0.0`
只统计关键 claim:
规则:
- 若任一关键事实为 `contradicted`,分数固定为 `0.0`
- 否则对关键事实取平均值
- 保留 2 位小数
- 分数范围必须在 `[0.0, 1.0]`
- direct_observation=1.0
- reasonable_inference=0.6
- overstated=0.3
- unsupported/external_unknown/contradicted=0.0
### 步骤六:facts_checked 兼容输出
你必须同时输出 `facts_checked`,用于旧链路兼容。
存在 contradicted 时固定 0.0,否则取平均并保留两位小数,范围 `[0.0, 1.0]`。
映射规则:
- `direct_observation` → `direct_evidence`
- `reasonable_inference` → `indirect_support`
- `overstated` → `indirect_support`
- `unsupported` → `no_evidence`
- `external_unknown` → `no_evidence`
- `contradicted` → `contradicted`
### 5. 输出 facts_checked 兼容字段
`facts_checked[*].fact` 使用 `{claim_id}: {claim_text}`。
按 claim_checks 映射:
### 步骤七:PASS 前覆盖性自检
在输出 `PASS` 前,必须再次检查:
- `claim_checks` 是否覆盖了 `executor_structured_output.claims` 中的全部 claims
- 是否存在 `gatekeeper_result.status="fail"`
- 是否存在 malformed/missing structured output
- direct_observation -> direct_evidence
- reasonable_inference/overstated -> indirect_support
- unsupported/external_unknown -> no_evidence
- contradicted -> contradicted
如有明显遗漏,即使已校验事实都有证据,也不得输出 `PASS`。
每项包含 `fact`、`is_critical`、`verification`、`detail`、`evidence_refs`。关键性由 claim 类型和当前 query 决定,不得为了触发重试把非关键缺口标成关键。
### 步骤八:处理 retry_context
若 `retry_context` 不为空:
- 优先检查上一轮缺失证据点是否已补足
- 不要扩展与缺口无关的新事实
- 不要因为存在 `retry_context` 就自动降低 verdict
## 输出协议
必须输出且只能输出以下 JSON 结构:
## 严格输出格式
```json
{
"verdict": "PASS",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"verdict": "PASS | LOW_CONFID | REJECT",
"groundedness_score": 0.0,
"critical_fact_count": 0,
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "ERR_TIMEOUT 表示请求超时",
"claim_type": "symptom",
"claim_text": "待校验事实",
"claim_type": "observation",
"verification": "direct_observation",
"detail": "知识库文档明确给出该错误码定义",
"detail": "matched_text 如何支持或不能支持该 claim",
"evidence_refs": [
{
"trace_ref": "trace-1",
"tool_name": "lookup_knowledge",
"topic_domain": "api",
"source_invocation_ids": [101, 104],
"note": "trace-1 的文档摘要直接给出错误码定义"
"claim_id": "claim-1",
"source_invocation_id": 1,
"tool_name": "query_metrics",
"raw_path": "$.alerts[0]",
"note": "证据关联说明"
}
]
}
],
"hypothesis_checks": [],
"facts_checked": [
{
"fact": "claim-1: ERR_TIMEOUT 表示请求超时",
"fact": "待校验事实",
"is_critical": true,
"verification": "direct_evidence",
"detail": "知识库文档明确给出该错误码定义",
"evidence_refs": [
{
"trace_ref": "trace-1",
"tool_name": "lookup_knowledge",
"topic_domain": "api",
"source_invocation_ids": [101, 104],
"note": "trace-1 的文档摘要直接给出错误码定义"
}
]
"detail": "校验说明",
"evidence_refs": []
}
],
"rationale": "所有关键事实均有支撑,且至少一条具有直接证据"
"rationale": "整体 verdict 的简短理由"
}
```
输出要求:
- `verdict` 只能是 `PASS` / `LOW_CONFID` / `REJECT`
- `groundedness_score` 必须是 JSON number
- `critical_fact_count` 必须等于关键 claim 的数量;兼容期也应等于 `facts_checked` 中 `is_critical=true` 的数量
- `claim_checks` 可以为空数组,但字段不能缺失
- `facts_checked` 可以为空数组,但字段不能缺失
- 每条 `claim_checks[*]` 都必须包含 `evidence_refs`
- 每条 `facts_checked[*]` 都必须包含 `evidence_refs`
- 不得输出 schema 之外的字段
输出前检查:每条 confirmed claim 都有匹配 verified evidence;没有引入输入外事实;没有把 no-evidence 写成“问题不存在/已排除”;没有在 JSON 外输出任何文字。
@@ -0,0 +1,39 @@
package com.superbiz.agent.config;
import org.junit.jupiter.api.Test;
import java.io.IOException;
import java.io.InputStream;
import java.nio.charset.StandardCharsets;
import java.util.Arrays;
import java.util.stream.Collectors;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
class DiagnosisRunSchemaContractTest {
private static final String MIGRATION =
"db/migration/V012__add_diagnosis_run_orchestration_trace.sql";
@Test
void v012AddsOnlyOneNullableJsonColumnToDiagnosisRun() throws IOException {
try (InputStream stream = Thread.currentThread()
.getContextClassLoader()
.getResourceAsStream(MIGRATION)) {
assertNotNull(stream, "missing migration " + MIGRATION);
String sql = new String(stream.readAllBytes(), StandardCharsets.UTF_8);
String executableSql = Arrays.stream(sql.split("\\R"))
.map(String::trim)
.filter(line -> !line.isEmpty() && !line.startsWith("--"))
.collect(Collectors.joining(" "))
.replaceAll("\\s+", " ");
assertEquals(
"ALTER TABLE diagnosis_run ADD COLUMN orchestration_trace JSON NULL "
+ "COMMENT 'Compact run-scoped StateGraph orchestration summary' "
+ "AFTER self_evaluation;",
executableSql);
}
}
}
@@ -0,0 +1,259 @@
package com.superbiz.agent.graph.diagnosis;
import com.alibaba.cloud.ai.graph.CompiledGraph;
import com.alibaba.cloud.ai.graph.RunnableConfig;
import org.junit.jupiter.api.Test;
import reactor.core.publisher.Flux;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.mockito.ArgumentMatchers.any;
import static org.mockito.ArgumentMatchers.anyMap;
import static org.mockito.Mockito.mock;
import static org.mockito.Mockito.when;
class ChatDiagnosisGraphRuntimeTest {
@Test
void successfulRunReturnsFinalAnswerTraceAndCurrentThreadId() throws Exception {
ScriptedDiagnosisGraphActions script = passPath();
ChatDiagnosisGraphRuntime.Result result =
new ChatDiagnosisGraphRuntime().execute(
script.actions(),
"分析连接池",
"session-runtime",
"run-runtime");
assertEquals("已确认连接数达到上限。", result.answer());
assertEquals(DiagnosisGraphTopology.Node.COMPOSER,
result.trace().finalNode());
assertEquals("composer_completed", result.trace().terminationReason());
assertEquals(false, result.trace().degraded());
assertEquals(List.of("run-runtime"),
script.threadIds(DiagnosisGraphTopology.Node.PLANNER));
assertEquals(List.of("run-runtime"),
script.threadIds(DiagnosisGraphTopology.Node.EXECUTOR));
assertEquals(List.of("run-runtime"),
script.threadIds(DiagnosisGraphTopology.Node.GATEKEEPER));
assertEquals(List.of("run-runtime"),
script.threadIds(DiagnosisGraphTopology.Node.VERIFIER));
assertEquals(List.of("run-runtime"),
script.threadIds(DiagnosisGraphTopology.Node.COMPOSER));
for (String node : List.of(
DiagnosisGraphTopology.Node.PLANNER,
DiagnosisGraphTopology.Node.EXECUTOR,
DiagnosisGraphTopology.Node.GATEKEEPER,
DiagnosisGraphTopology.Node.VERIFIED_INPUT,
DiagnosisGraphTopology.Node.VERIFIER,
DiagnosisGraphTopology.Node.COMPOSER)) {
assertEquals(List.of("session-runtime"),
script.metadata(node, "sessionId"));
assertEquals(List.of("run-runtime"),
script.metadata(node, "runId"));
}
Map<String, Object> initial = script.states(
DiagnosisGraphTopology.Node.PLANNER).get(0);
assertEquals(Map.of(
"query", "分析连接池",
"original_query", "分析连接池"),
initial.get(DiagnosisGraphState.DIAGNOSIS_CONTEXT));
assertEquals("NORMAL", initial.get(DiagnosisGraphState.PLANNER_MODE));
assertEquals(0, initial.get(DiagnosisGraphState.PLANNER_RETRY_COUNT));
assertEquals(0, initial.get(DiagnosisGraphState.VERIFIER_RETRY_COUNT));
assertEquals(0, initial.get(DiagnosisGraphState.COMPOSER_RETRY_COUNT));
assertEquals(0, initial.get(DiagnosisGraphState.EVIDENCE_RETRY_COUNT));
assertEquals(List.of(), initial.get(DiagnosisGraphState.ORCHESTRATION_EVENTS));
}
@Test
void executionFailureExposesOnlyRealPartialTrace() {
ScriptedDiagnosisGraphActions afterPlannerFailure =
new ScriptedDiagnosisGraphActions()
.step(DiagnosisGraphTopology.Node.PLANNER, "COMPLETED",
"planner_completed", Map.of(
DiagnosisGraphState.PLANNER_STATUS, "COMPLETED",
DiagnosisGraphState.PLANNER_PLAN,
Map.of("plan", List.of("查询指标"))));
ChatDiagnosisGraphRuntime.ExecutionFailure partial = assertThrows(
ChatDiagnosisGraphRuntime.ExecutionFailure.class,
() -> new ChatDiagnosisGraphRuntime().execute(
afterPlannerFailure.actions(), "问题", "session-partial", "run-partial"));
assertEquals(DiagnosisGraphTopology.Node.PLANNER,
partial.partialTrace().finalNode());
assertEquals("planner_completed",
partial.partialTrace().terminationReason());
ChatDiagnosisGraphRuntime.ExecutionFailure beforeFirstEvent = assertThrows(
ChatDiagnosisGraphRuntime.ExecutionFailure.class,
() -> new ChatDiagnosisGraphRuntime().execute(
new ScriptedDiagnosisGraphActions().actions(),
"问题", "session-empty", "run-empty"));
assertNull(beforeFirstEvent.partialTrace());
}
@Test
void handledPreVerificationFallbackReturnsSafeDegradedResult() throws Exception {
ScriptedDiagnosisGraphActions script = new ScriptedDiagnosisGraphActions()
.step(DiagnosisGraphTopology.Node.PLANNER, "NON_RETRYABLE_FAILED",
"planner_failed", Map.of(
DiagnosisGraphState.PLANNER_STATUS,
"NON_RETRYABLE_FAILED"))
.step(DiagnosisGraphTopology.Node.FALLBACK, "COMPLETED",
"fallback_completed", Map.of(
DiagnosisGraphState.FINAL_ANSWER,
"当前无法形成可信诊断。"));
ChatDiagnosisGraphRuntime.Result result =
new ChatDiagnosisGraphRuntime().execute(
script.actions(), "问题", "session-pre", "run-pre");
assertEquals("当前无法形成可信诊断。", result.answer());
assertEquals(DiagnosisGraphTopology.Node.FALLBACK,
result.trace().finalNode());
assertEquals(true, result.trace().degraded());
assertEquals(0, script.calls(DiagnosisGraphTopology.Node.VERIFIER));
}
@Test
void handledPostVerificationFallbackReturnsSafeDegradedResult() throws Exception {
ScriptedDiagnosisGraphActions script = new ScriptedDiagnosisGraphActions()
.step(DiagnosisGraphTopology.Node.PLANNER, "COMPLETED",
"planner_completed", Map.of(
DiagnosisGraphState.PLANNER_STATUS, "COMPLETED"))
.step(DiagnosisGraphTopology.Node.EXECUTOR, "COMPLETED",
"executor_completed", Map.of(
DiagnosisGraphState.EXECUTOR_STATUS, "COMPLETED"))
.step(DiagnosisGraphTopology.Node.GATEKEEPER, "PASS",
"gatekeeper_pass", Map.of(
DiagnosisGraphState.GATEKEEPER_STATUS, "PASS",
DiagnosisGraphState.VERIFIED_BINDING_COUNT, 1))
.step(DiagnosisGraphTopology.Node.VERIFIED_INPUT, "COMPLETED",
"verified_input_built", Map.of())
.step(DiagnosisGraphTopology.Node.VERIFIER, "INVALID_OUTPUT",
"verifier_invalid_output", Map.of(
DiagnosisGraphState.VERIFIER_STATUS,
"INVALID_OUTPUT"))
.step(DiagnosisGraphTopology.Node.VERIFIER, "INVALID_OUTPUT",
"verifier_invalid_output", Map.of(
DiagnosisGraphState.VERIFIER_STATUS,
"INVALID_OUTPUT"))
.step(DiagnosisGraphTopology.Node.FALLBACK, "COMPLETED",
"fallback_completed", Map.of(
DiagnosisGraphState.FINAL_ANSWER,
"已降级为安全答复。"));
ChatDiagnosisGraphRuntime.Result result =
new ChatDiagnosisGraphRuntime().execute(
script.actions(), "问题", "session-post", "run-post");
assertEquals("已降级为安全答复。", result.answer());
assertEquals(DiagnosisGraphTopology.Node.FALLBACK,
result.trace().finalNode());
assertEquals(true, result.trace().degraded());
assertEquals(2, script.calls(DiagnosisGraphTopology.Node.VERIFIER));
}
@Test
void blankAnswerFailsAndKeepsOnlyTheRealFallbackTrace() {
ScriptedDiagnosisGraphActions script = new ScriptedDiagnosisGraphActions()
.step(DiagnosisGraphTopology.Node.PLANNER, "NON_RETRYABLE_FAILED",
"planner_failed", Map.of(
DiagnosisGraphState.PLANNER_STATUS,
"NON_RETRYABLE_FAILED"))
.step(DiagnosisGraphTopology.Node.FALLBACK, "COMPLETED",
"fallback_completed", Map.of(
DiagnosisGraphState.FINAL_ANSWER, " "));
ChatDiagnosisGraphRuntime.ExecutionFailure failure = assertThrows(
ChatDiagnosisGraphRuntime.ExecutionFailure.class,
() -> new ChatDiagnosisGraphRuntime().execute(
script.actions(), "问题", "session-blank", "run-blank"));
assertEquals("diagnosis graph returned no final answer",
failure.getCause().getMessage());
assertEquals(DiagnosisGraphTopology.Node.FALLBACK,
failure.partialTrace().finalNode());
assertEquals(2, failure.partialState().data()
.get(DiagnosisGraphState.ORCHESTRATION_EVENTS) instanceof List<?> events
? events.size() : 0);
}
@Test
void emptyGraphStreamFailsWithoutFabricatingStateOrTrace() throws Exception {
DiagnosisGraphFactory graphFactory = mock(DiagnosisGraphFactory.class);
CompiledGraph graph = mock(CompiledGraph.class);
when(graphFactory.compile(any(DiagnosisGraphActions.class)))
.thenReturn(graph);
when(graph.stream(anyMap(), any(RunnableConfig.class)))
.thenReturn(Flux.empty());
ChatDiagnosisGraphRuntime runtime = new ChatDiagnosisGraphRuntime(
graphFactory, new DiagnosisOrchestrationTraceBuilder());
ChatDiagnosisGraphRuntime.ExecutionFailure failure = assertThrows(
ChatDiagnosisGraphRuntime.ExecutionFailure.class,
() -> runtime.execute(
new ScriptedDiagnosisGraphActions().actions(),
"问题", "session-empty-stream", "run-empty-stream"));
assertEquals("diagnosis graph returned no state",
failure.getCause().getMessage());
assertNull(failure.partialState());
assertNull(failure.partialTrace());
}
private ScriptedDiagnosisGraphActions passPath() {
return new ScriptedDiagnosisGraphActions()
.step(DiagnosisGraphTopology.Node.PLANNER, "COMPLETED",
"planner_completed", Map.of(
DiagnosisGraphState.PLANNER_STATUS, "COMPLETED",
DiagnosisGraphState.PLANNER_PLAN,
Map.of("plan", List.of("查询指标"))))
.step(DiagnosisGraphTopology.Node.EXECUTOR, "COMPLETED",
"executor_completed", Map.of(
DiagnosisGraphState.EXECUTOR_STATUS, "COMPLETED",
DiagnosisGraphState.EXECUTOR_OUTPUT,
Map.of("answer_version", "executor_evidence_v2",
"claims", List.of())))
.step(DiagnosisGraphTopology.Node.GATEKEEPER, "PASS",
"gatekeeper_pass", Map.of(
DiagnosisGraphState.GATEKEEPER_STATUS, "PASS",
DiagnosisGraphState.GATEKEEPER_RESULT,
Map.of("status", "pass", "severity", "none"),
DiagnosisGraphState.VERIFIED_BINDING_COUNT, 1,
DiagnosisGraphState.VERIFIER_VERDICT_CEILING, "PASS"))
.step(DiagnosisGraphTopology.Node.VERIFIED_INPUT, "COMPLETED",
"verified_input_built", Map.of(
DiagnosisGraphState.VERIFIED_EXECUTOR_OUTPUT,
Map.of("answer_version", "executor_evidence_v2",
"claims", List.of()),
DiagnosisGraphState.VERIFIED_EVIDENCE, List.of()))
.step(DiagnosisGraphTopology.Node.VERIFIER, "COMPLETED",
"verifier_completed", Map.of(
DiagnosisGraphState.VERIFIER_STATUS, "COMPLETED",
DiagnosisGraphState.VERIFIER_MODEL_VERDICT, "PASS",
DiagnosisGraphState.EFFECTIVE_VERDICT, "PASS",
DiagnosisGraphState.VERIFIER_OUTPUT, Map.of(
"verdict", "PASS",
"groundedness_score", 1.0,
"critical_fact_count", 0,
"claim_checks", List.of(),
"facts_checked", List.of(),
"rationale", "证据充分")))
.step(DiagnosisGraphTopology.Node.COMPOSER, "COMPLETED",
"composer_completed", Map.of(
DiagnosisGraphState.COMPOSER_STATUS, "COMPLETED",
DiagnosisGraphState.COMPOSER_OUTPUT,
Map.of("status", "valid"),
DiagnosisGraphState.FINAL_ANSWER,
"已确认连接数达到上限。"));
}
}
@@ -0,0 +1,39 @@
package com.superbiz.agent.graph.diagnosis;
import org.junit.jupiter.api.Test;
import org.springframework.core.io.ClassPathResource;
import java.nio.charset.StandardCharsets;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
class ChatVerifierPromptContractTest {
@Test
void promptDeclaresOnlyVerifiedGraphInputContract() throws Exception {
String prompt = new String(
new ClassPathResource("prompts/chat-verifier-prompt.md")
.getInputStream().readAllBytes(),
StandardCharsets.UTF_8);
for (String allowed : List.of(
"diagnosis_context",
"verified_executor_output",
"verified_evidence",
"gatekeeper_audit",
"verdict_ceiling",
"retry_context")) {
assertTrue(prompt.contains("`" + allowed + "`"), allowed);
}
for (String forbidden : List.of(
"executor_final_answer",
"executor_structured_output",
"executor_output_parse_status",
"tool_trace_summary",
"VerifierInputHook")) {
assertFalse(prompt.contains(forbidden), forbidden);
}
}
}
@@ -0,0 +1,91 @@
package com.superbiz.agent.graph.diagnosis;
import com.alibaba.cloud.ai.graph.OverAllState;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
class DiagnosisGraphResultMapperTest {
private final DiagnosisGraphResultMapper mapper =
new DiagnosisGraphResultMapper();
@Test
void completedVerifierMapsCompatibilityFieldsFromVerifiedStateOnly() {
Map<String, Object> verifiedOutput = Map.of(
"answer_version", "executor_evidence_v2",
"claims", List.of(Map.of("claim_id", "c1")));
List<Map<String, Object>> verifiedEvidence = List.of(Map.of(
"claim_id", "c1",
"source_invocation_id", 17,
"tool_name", "query_metrics",
"raw_path", "$.alerts[0]",
"matched_text", "active=50 max=50"));
Map<String, Object> gatekeeper = Map.of(
"status", "fail",
"severity", "low_confid",
"checked_bindings", verifiedEvidence);
Map<String, Object> verifierOutput = Map.of(
"verdict", "PASS",
"groundedness_score", 1.0,
"critical_fact_count", 1,
"claim_checks", List.of(Map.of(
"claim_id", "c1",
"verification", "direct_observation")),
"facts_checked", List.of(),
"rationale", "模型认为证据充分");
OverAllState state = new OverAllState(Map.ofEntries(
Map.entry(DiagnosisGraphState.PLANNER_STATUS, "COMPLETED"),
Map.entry(DiagnosisGraphState.EXECUTOR_STATUS, "COMPLETED"),
Map.entry(DiagnosisGraphState.GATEKEEPER_STATUS, "LOW_CONFID"),
Map.entry(DiagnosisGraphState.GATEKEEPER_RESULT, gatekeeper),
Map.entry(DiagnosisGraphState.VERIFIED_EXECUTOR_OUTPUT, verifiedOutput),
Map.entry(DiagnosisGraphState.VERIFIED_EVIDENCE, verifiedEvidence),
Map.entry(DiagnosisGraphState.VERIFIER_STATUS, "COMPLETED"),
Map.entry(DiagnosisGraphState.VERIFIER_MODEL_VERDICT, "PASS"),
Map.entry(DiagnosisGraphState.EFFECTIVE_VERDICT, "LOW_CONFID"),
Map.entry(DiagnosisGraphState.VERIFIER_OUTPUT, verifierOutput),
Map.entry(DiagnosisGraphState.COMPOSER_STATUS, "COMPLETED"),
Map.entry(DiagnosisGraphState.COMPOSER_OUTPUT, Map.of("status", "valid")),
Map.entry(DiagnosisGraphState.EVIDENCE_RETRY_COUNT, 0),
Map.entry(DiagnosisGraphState.EXECUTOR_OUTPUT, Map.of(
"raw_secret", "must-not-persist"))));
Map<String, Object> evaluation = mapper.verifierEvaluation(
state, Map.of("version", "chat-prompts-v2"));
assertEquals("COMPLETED", evaluation.get("verifier_status"));
assertEquals("PASS", evaluation.get("model_verdict"));
assertEquals("LOW_CONFID", evaluation.get("effective_verdict"));
assertEquals("LOW_CONFID", evaluation.get("verdict"));
assertEquals(verifiedOutput, evaluation.get("executor_structured_output"));
assertEquals(verifiedEvidence, evaluation.get("verified_evidence"));
assertEquals(gatekeeper, evaluation.get("gatekeeper_result"));
assertEquals(Map.of("status", "valid"), evaluation.get("composer_output"));
assertEquals(Map.of("version", "chat-prompts-v2"), evaluation.get("prompt_audit"));
assertEquals(Map.of("status", "valid", "detail", "graph executor completed"),
evaluation.get("executor_output_parse_status"));
assertFalse(evaluation.containsKey("tool_trace_summary"));
assertFalse(String.valueOf(evaluation).contains("raw_secret"));
}
@Test
void preVerificationFallbackDoesNotFabricateVerdict() {
OverAllState state = new OverAllState(Map.of(
DiagnosisGraphState.PLANNER_STATUS, "INVALID_OUTPUT",
DiagnosisGraphState.FAILURE_REASON, "planner_invalid_output"));
Map<String, Object> evaluation = mapper.verifierEvaluation(
state, Map.of("version", "chat-prompts-v2"));
assertEquals("INVALID_OUTPUT", evaluation.get("planner_status"));
assertEquals("planner_invalid_output", evaluation.get("failure_reason"));
assertFalse(evaluation.containsKey("verdict"));
assertFalse(evaluation.containsKey("model_verdict"));
assertFalse(evaluation.containsKey("effective_verdict"));
}
}
@@ -70,6 +70,16 @@ final class ScriptedDiagnosisGraphActions {
return List.copyOf(node(node).threadIds);
}
List<Object> metadata(String node, String key) {
return node(node).metadata.stream()
.map(values -> values.get(key))
.toList();
}
List<Map<String, Object>> states(String node) {
return List.copyOf(node(node).states);
}
private ScriptedNode node(String name) {
ScriptedNode node = nodes.get(name);
if (node == null) {
@@ -83,6 +93,8 @@ final class ScriptedDiagnosisGraphActions {
private final String name;
private final Deque<Step> steps = new ArrayDeque<>();
private final List<String> threadIds = new ArrayList<>();
private final List<Map<String, Object>> metadata = new ArrayList<>();
private final List<Map<String, Object>> states = new ArrayList<>();
private int calls;
private ScriptedNode(String name) {
@@ -104,6 +116,8 @@ final class ScriptedDiagnosisGraphActions {
calls++;
sequence.add(name);
threadIds.add(config.threadId().orElse(null));
metadata.add(new LinkedHashMap<>(config.metadata().orElse(Map.of())));
states.add(new LinkedHashMap<>(state.data()));
Map<String, Object> update = new LinkedHashMap<>(step.update());
update.put(DiagnosisGraphState.ORCHESTRATION_EVENTS,
@@ -0,0 +1,301 @@
package com.superbiz.agent.service;
import com.alibaba.cloud.ai.graph.OverAllState;
import com.alibaba.cloud.ai.graph.action.AsyncNodeActionWithConfig;
import com.superbiz.agent.agent.tool.DateTimeTools;
import com.superbiz.agent.agent.tool.QueryLogsTools;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.ChatSession;
import com.superbiz.agent.domain.entity.DiagnosisRun;
import com.superbiz.agent.graph.diagnosis.ChatDiagnosisGraphRuntime;
import com.superbiz.agent.graph.diagnosis.DiagnosisGraphState;
import com.superbiz.agent.graph.diagnosis.DiagnosisGraphActions;
import com.superbiz.agent.graph.diagnosis.DiagnosisGraphTopology;
import com.superbiz.agent.graph.diagnosis.DiagnosisOrchestrationTrace;
import com.superbiz.agent.graph.diagnosis.OrchestrationEvent;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.ChatSessionRepository;
import com.superbiz.agent.repository.DiagnosisRunRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
import com.superbiz.agent.tool.LookupKnowledgeTool;
import com.superbiz.agent.tool.RetrievedDocTracker;
import org.junit.jupiter.api.Test;
import org.mockito.ArgumentCaptor;
import org.springframework.ai.chat.model.ChatModel;
import org.springframework.ai.tool.ToolCallback;
import org.springframework.test.util.ReflectionTestUtils;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CompletableFuture;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.mockito.ArgumentMatchers.any;
import static org.mockito.ArgumentMatchers.anyString;
import static org.mockito.ArgumentMatchers.eq;
import static org.mockito.Mockito.atLeast;
import static org.mockito.Mockito.mock;
import static org.mockito.Mockito.never;
import static org.mockito.Mockito.verify;
import static org.mockito.Mockito.when;
class ChatServiceGraphIntegrationTest {
@Test
void graphSuccessPreservesChatResultRunTraceMetricsAndEvaluation() throws Exception {
Fixture fixture = fixture();
String sessionId = "graph-success-session";
String question = "请分析订单支付超时";
when(fixture.runtime.execute(any(), eq(question), eq(sessionId), anyString()))
.thenAnswer(invocation -> successResult(invocation.getArgument(3), false));
ChatService.ChatResult result = fixture.service.executeChatComplex(
fixture.chatModel, new ToolCallback[0], question, List.of(), sessionId);
assertEquals(sessionId, result.sessionId());
assertTrue(result.runId().startsWith("run-"));
assertEquals("诊断完成。", result.answer());
DiagnosisRun saved = lastSavedRun(fixture.diagnosisRunRepository);
assertEquals(result.runId(), saved.getRunId());
assertEquals(sessionId, saved.getSessionId());
assertEquals("CHAT", saved.getAgentFlow());
assertEquals("SUCCESS", saved.getStatus());
assertEquals("诊断完成。", saved.getAnswer());
assertEquals(12, saved.getTotalTokenCount());
assertEquals(1, saved.getStepCount());
assertEquals(2, saved.getToolCallCount());
assertNotNull(saved.getOrchestrationTrace());
assertTrue(saved.getOrchestrationTrace().contains("\"final_node\":\"composer\""));
assertTrue(saved.getSelfEvaluation().contains("\"verdict\":\"PASS\""));
assertTrue(saved.getSelfEvaluation().contains("\"prompt_audit\""));
assertFalse(saved.getSelfEvaluation().contains("tool_trace_summary"));
verify(fixture.evaluationService).evaluateRun(result.runId(), result.answer());
verify(fixture.retrievedDocTracker).clearSession(sessionId);
}
@Test
void handledFallbackIsSuccessfulAndUsesDegradedTraceWithoutFabricatedVerdict() throws Exception {
Fixture fixture = fixture();
String sessionId = "graph-fallback-session";
when(fixture.runtime.execute(any(), anyString(), eq(sessionId), anyString()))
.thenAnswer(invocation -> successResult(invocation.getArgument(3), true));
ChatService.ChatResult result = fixture.service.executeChatComplex(
fixture.chatModel, new ToolCallback[0], "无法建立可信材料", List.of(), sessionId);
DiagnosisRun saved = lastSavedRun(fixture.diagnosisRunRepository);
assertEquals("SUCCESS", saved.getStatus());
assertEquals("当前无法给出可信结论。", result.answer());
assertTrue(saved.getOrchestrationTrace().contains("\"degraded\":true"));
assertTrue(saved.getSelfEvaluation().contains("planner_status"));
assertFalse(saved.getSelfEvaluation().contains("\"verdict\""));
verify(fixture.evaluationService).evaluateRun(result.runId(), result.answer());
}
@Test
void unhandledGraphFailureMarksRunFailedAndDoesNotEvaluate() throws Exception {
Fixture fixture = fixture();
String sessionId = "graph-failed-session";
when(fixture.runtime.execute(any(), anyString(), eq(sessionId), anyString()))
.thenThrow(new IllegalStateException("graph exploded"));
ChatService.ChatResult result = fixture.service.executeChatComplex(
fixture.chatModel, new ToolCallback[0], "触发失败", List.of(), sessionId);
DiagnosisRun saved = lastSavedRun(fixture.diagnosisRunRepository);
assertEquals("FAILED", saved.getStatus());
assertTrue(result.answer().contains("graph exploded"));
assertEquals(result.answer(), saved.getAnswer());
verify(fixture.evaluationService, never()).evaluateRun(anyString(), anyString());
verify(fixture.retrievedDocTracker).clearSession(sessionId);
}
@Test
void blankGraphAnswerMarksRunFailedAndPersistsOnlyRealPartialTrace()
throws Exception {
Fixture fixture = fixture();
String sessionId = "graph-blank-answer-session";
when(fixture.runtime.execute(any(), anyString(), eq(sessionId), anyString()))
.thenAnswer(invocation -> new ChatDiagnosisGraphRuntime().execute(
blankAnswerActions(),
invocation.getArgument(1),
invocation.getArgument(2),
invocation.getArgument(3)));
ChatService.ChatResult result = fixture.service.executeChatComplex(
fixture.chatModel, new ToolCallback[0], "空答案", List.of(), sessionId);
DiagnosisRun saved = lastSavedRun(fixture.diagnosisRunRepository);
assertEquals("FAILED", saved.getStatus());
assertTrue(result.answer().contains("diagnosis graph execution failed"));
assertNotNull(saved.getOrchestrationTrace());
assertTrue(saved.getOrchestrationTrace().contains(
"\"final_node\":\"fallback\""));
assertFalse(saved.getOrchestrationTrace().contains("verifier"));
verify(fixture.evaluationService, never())
.evaluateRun(anyString(), anyString());
}
@Test
void sameSessionCreatesDistinctRunIdsAndKeepsTracePerRun() throws Exception {
Fixture fixture = fixture();
String sessionId = "graph-multi-run-session";
when(fixture.runtime.execute(any(), anyString(), eq(sessionId), anyString()))
.thenAnswer(invocation -> successResult(invocation.getArgument(3), false));
ChatService.ChatResult first = fixture.service.executeChatComplex(
fixture.chatModel, new ToolCallback[0], "第一轮", List.of(), sessionId);
ChatService.ChatResult second = fixture.service.executeChatComplex(
fixture.chatModel, new ToolCallback[0], "第二轮", List.of(), sessionId);
assertNotEquals(first.runId(), second.runId());
ArgumentCaptor<DiagnosisRun> captor = ArgumentCaptor.forClass(DiagnosisRun.class);
verify(fixture.diagnosisRunRepository, atLeast(4)).save(captor.capture());
List<DiagnosisRun> completed = captor.getAllValues().stream()
.filter(run -> "SUCCESS".equals(run.getStatus()))
.toList();
assertTrue(completed.stream().anyMatch(run -> first.runId().equals(run.getRunId())));
assertTrue(completed.stream().anyMatch(run -> second.runId().equals(run.getRunId())));
assertTrue(completed.stream().allMatch(run ->
run.getOrchestrationTrace().contains(run.getRunId())));
}
private ChatDiagnosisGraphRuntime.Result successResult(String runId, boolean degraded) {
OverAllState state = degraded
? new OverAllState(Map.of(
DiagnosisGraphState.PLANNER_STATUS, "INVALID_OUTPUT",
DiagnosisGraphState.FAILURE_REASON, "planner_invalid_output"))
: new OverAllState(Map.ofEntries(
Map.entry(DiagnosisGraphState.PLANNER_STATUS, "COMPLETED"),
Map.entry(DiagnosisGraphState.EXECUTOR_STATUS, "COMPLETED"),
Map.entry(DiagnosisGraphState.GATEKEEPER_STATUS, "PASS"),
Map.entry(DiagnosisGraphState.VERIFIER_STATUS, "COMPLETED"),
Map.entry(DiagnosisGraphState.VERIFIER_MODEL_VERDICT, "PASS"),
Map.entry(DiagnosisGraphState.EFFECTIVE_VERDICT, "PASS"),
Map.entry(DiagnosisGraphState.VERIFIER_OUTPUT, Map.of(
"verdict", "PASS",
"groundedness_score", 1.0,
"critical_fact_count", 0,
"claim_checks", List.of(),
"facts_checked", List.of(),
"rationale", "证据充分")),
Map.entry(DiagnosisGraphState.COMPOSER_STATUS, "COMPLETED"),
Map.entry(DiagnosisGraphState.COMPOSER_OUTPUT, Map.of("status", "valid"))));
DiagnosisOrchestrationTrace trace = new DiagnosisOrchestrationTrace(
"stategraph-v1",
List.of(),
degraded ? DiagnosisGraphTopology.Node.FALLBACK
: DiagnosisGraphTopology.Node.COMPOSER,
degraded ? "fallback_completed" : "composer_completed-" + runId,
degraded,
0);
return new ChatDiagnosisGraphRuntime.Result(
state,
degraded ? "当前无法给出可信结论。" : "诊断完成。",
trace);
}
private DiagnosisGraphActions blankAnswerActions() {
AsyncNodeActionWithConfig unexpected = (state, config) ->
CompletableFuture.failedFuture(new AssertionError(
"unexpected graph node"));
AsyncNodeActionWithConfig planner = (state, config) ->
CompletableFuture.completedFuture(Map.of(
DiagnosisGraphState.PLANNER_STATUS,
"NON_RETRYABLE_FAILED",
DiagnosisGraphState.ORCHESTRATION_EVENTS,
List.of(new OrchestrationEvent(
DiagnosisGraphTopology.Node.PLANNER,
"NON_RETRYABLE_FAILED",
"planner_failed",
1))));
AsyncNodeActionWithConfig fallback = (state, config) ->
CompletableFuture.completedFuture(Map.of(
DiagnosisGraphState.FINAL_ANSWER, " ",
DiagnosisGraphState.ORCHESTRATION_EVENTS,
List.of(new OrchestrationEvent(
DiagnosisGraphTopology.Node.FALLBACK,
"COMPLETED",
"fallback_completed",
1))));
return new DiagnosisGraphActions(
planner,
unexpected,
unexpected,
unexpected,
unexpected,
unexpected,
unexpected,
fallback);
}
private DiagnosisRun lastSavedRun(DiagnosisRunRepository repository) {
ArgumentCaptor<DiagnosisRun> captor = ArgumentCaptor.forClass(DiagnosisRun.class);
verify(repository, atLeast(2)).save(captor.capture());
return captor.getAllValues().get(captor.getAllValues().size() - 1);
}
private Fixture fixture() throws Exception {
ChatService service = new ChatService();
ChatDiagnosisGraphRuntime runtime = mock(ChatDiagnosisGraphRuntime.class);
ChatSessionRepository chatSessionRepository = mock(ChatSessionRepository.class);
when(chatSessionRepository.findBySessionId(anyString())).thenReturn(java.util.Optional.empty());
when(chatSessionRepository.save(any(ChatSession.class)))
.thenAnswer(invocation -> invocation.getArgument(0));
DiagnosisRunRepository diagnosisRunRepository = mock(DiagnosisRunRepository.class);
when(diagnosisRunRepository.save(any(DiagnosisRun.class)))
.thenAnswer(invocation -> invocation.getArgument(0));
AgentStepRepository agentStepRepository = mock(AgentStepRepository.class);
when(agentStepRepository.findByRunIdOrderByStepIndex(anyString()))
.thenReturn(List.of(AgentStep.builder().tokenCount(12).build()));
ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
when(toolInvocationRepository.countByRunId(anyString())).thenReturn(2L);
EvaluationService evaluationService = mock(EvaluationService.class);
RetrievedDocTracker retrievedDocTracker = mock(RetrievedDocTracker.class);
KnowledgeDomainService knowledgeDomainService = mock(KnowledgeDomainService.class);
when(knowledgeDomainService.buildKnowledgeMap()).thenReturn("");
ReflectionTestUtils.setField(service, "dateTimeTools", new DateTimeTools());
ReflectionTestUtils.setField(service, "lookupKnowledgeTool", new LookupKnowledgeTool());
ReflectionTestUtils.setField(service, "queryLogsTools",
new QueryLogsTools(mock(ToolInvocationRecorder.class)));
ReflectionTestUtils.setField(service, "chatSessionRepository", chatSessionRepository);
ReflectionTestUtils.setField(service, "diagnosisRunRepository", diagnosisRunRepository);
ReflectionTestUtils.setField(service, "agentStepRepository", agentStepRepository);
ReflectionTestUtils.setField(service, "toolInvocationRepository", toolInvocationRepository);
ReflectionTestUtils.setField(service, "evaluationService", evaluationService);
ReflectionTestUtils.setField(service, "retrievedDocTracker", retrievedDocTracker);
ReflectionTestUtils.setField(service, "knowledgeDomainService", knowledgeDomainService);
ReflectionTestUtils.setField(service, "selfEvaluationMergeService", new SelfEvaluationMergeService());
ReflectionTestUtils.setField(service, "executorGatekeeperService",
mock(ExecutorGatekeeperService.class));
ReflectionTestUtils.setField(service, "diagnosisGraphRuntime", runtime);
ReflectionTestUtils.setField(service, "chatPlannerPrompt", "PLANNER_TEST_PROMPT");
ReflectionTestUtils.setField(service, "chatExecutorPrompt", "EXECUTOR_TEST_PROMPT");
ReflectionTestUtils.setField(service, "chatVerifierPrompt", "VERIFIER_TEST_PROMPT");
ReflectionTestUtils.setField(service, "chatComposerPrompt", "COMPOSER_TEST_PROMPT");
return new Fixture(
service,
runtime,
mock(ChatModel.class),
diagnosisRunRepository,
evaluationService,
retrievedDocTracker);
}
private record Fixture(
ChatService service,
ChatDiagnosisGraphRuntime runtime,
ChatModel chatModel,
DiagnosisRunRepository diagnosisRunRepository,
EvaluationService evaluationService,
RetrievedDocTracker retrievedDocTracker) {
}
}
@@ -1,918 +0,0 @@
package com.superbiz.agent.service;
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
import com.alibaba.cloud.ai.graph.skills.registry.classpath.ClasspathSkillRegistry;
import com.superbiz.agent.agent.tool.DateTimeTools;
import com.superbiz.agent.agent.tool.QueryLogsTools;
import com.superbiz.agent.agent.tool.QueryMetricsTools;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.ChatSession;
import com.superbiz.agent.domain.entity.DiagnosisRun;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.ChatSessionRepository;
import com.superbiz.agent.repository.DiagnosisRunRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
import com.superbiz.agent.tool.LookupKnowledgeTool;
import com.superbiz.agent.tool.RetrievedDocTracker;
import org.junit.jupiter.api.Test;
import org.springframework.ai.chat.messages.AssistantMessage;
import org.springframework.ai.chat.model.ChatModel;
import org.springframework.ai.chat.model.ChatResponse;
import org.springframework.ai.chat.model.Generation;
import org.springframework.ai.chat.prompt.Prompt;
import org.springframework.ai.tool.ToolCallback;
import org.springframework.test.util.ReflectionTestUtils;
import org.mockito.ArgumentCaptor;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.atomic.AtomicInteger;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.mockito.ArgumentMatchers.any;
import static org.mockito.ArgumentMatchers.anyString;
import static org.mockito.ArgumentMatchers.eq;
import static org.mockito.ArgumentMatchers.isNull;
import static org.mockito.Mockito.atLeast;
import static org.mockito.Mockito.atLeastOnce;
import static org.mockito.Mockito.verify;
import static org.mockito.Mockito.mock;
import static org.mockito.Mockito.when;
class ChatServiceSequentialAgentTest {
@Test
void executeChatComplexInvokesSequentialWorkflow() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel();
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析订单支付超时的原因,并给出修复建议",
List.of(),
"sequential-test-session"
);
assertTrue(result.answer().contains("连接池 active 达到上限"));
assertFalse(result.answer().contains("\"answer_version\""));
assertEquals("sequential-test-session", result.sessionId());
assertTrue(result.runId().startsWith("run-"));
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "chat_composer"), chatModel.agentCalls);
assertTrue(chatModel.sawVerifierPrompt);
ChatSessionRepository chatSessionRepository =
(ChatSessionRepository) ReflectionTestUtils.getField(chatService, "chatSessionRepository");
DiagnosisRunRepository diagnosisRunRepository =
(DiagnosisRunRepository) ReflectionTestUtils.getField(chatService, "diagnosisRunRepository");
EvaluationService evaluationService =
(EvaluationService) ReflectionTestUtils.getField(chatService, "evaluationService");
ArgumentCaptor<ChatSession> chatSessionCaptor = ArgumentCaptor.forClass(ChatSession.class);
verify(chatSessionRepository, atLeastOnce()).save(chatSessionCaptor.capture());
assertEquals("sequential-test-session", chatSessionCaptor.getValue().getSessionId());
ArgumentCaptor<DiagnosisRun> runCaptor = ArgumentCaptor.forClass(DiagnosisRun.class);
verify(diagnosisRunRepository, atLeastOnce()).save(runCaptor.capture());
DiagnosisRun savedRun = runCaptor.getValue();
assertEquals(result.runId(), savedRun.getRunId());
assertEquals("sequential-test-session", savedRun.getSessionId());
assertEquals("SUCCESS", savedRun.getStatus());
assertEquals(result.answer(), savedRun.getAnswer());
verify(evaluationService).evaluateRun(eq(result.runId()), eq(result.answer()));
}
@Test
void executeChatComplexCreatesDistinctRunsForSameSessionAcrossTurns() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel firstRoundModel = new ScriptedChatModel();
ScriptedChatModel secondRoundModel = new ScriptedChatModel();
String sessionId = "sequential-same-session";
ChatService.ChatResult first = chatService.executeChatComplex(
firstRoundModel,
new ToolCallback[0],
"第一轮:请分析支付超时",
List.of(),
sessionId
);
ChatService.ChatResult second = chatService.executeChatComplex(
secondRoundModel,
new ToolCallback[0],
"第二轮:基于上一轮结论列出缺失证据",
List.of(
Map.of("role", "user", "content", "第一轮:请分析支付超时"),
Map.of("role", "assistant", "content", first.answer())
),
sessionId
);
assertEquals(sessionId, first.sessionId());
assertEquals(sessionId, second.sessionId());
assertNotEquals(first.runId(), second.runId());
DiagnosisRunRepository diagnosisRunRepository =
(DiagnosisRunRepository) ReflectionTestUtils.getField(chatService, "diagnosisRunRepository");
ArgumentCaptor<DiagnosisRun> runCaptor = ArgumentCaptor.forClass(DiagnosisRun.class);
verify(diagnosisRunRepository, atLeast(2)).save(runCaptor.capture());
List<String> savedRunIds = runCaptor.getAllValues().stream()
.filter(run -> sessionId.equals(run.getSessionId()))
.map(DiagnosisRun::getRunId)
.distinct()
.toList();
assertEquals(2, savedRunIds.size());
assertTrue(savedRunIds.contains(first.runId()));
assertTrue(savedRunIds.contains(second.runId()));
}
@Test
void executeChatComplexDoesNotRetryLowConfidenceByDefault() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "LOW_CONFID",
"groundedness_score": 0.1,
"critical_fact_count": 1,
"facts_checked": [
{
"fact": "missing direct evidence",
"is_critical": true,
"verification": "no_evidence",
"detail": "scripted evidence gap",
"evidence_refs": []
}
],
"rationale": "scripted low confidence"
}
""");
chatModel.composerOutput = "not-json";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析订单支付超时的原因,并给出修复建议",
List.of(),
"sequential-low-confidence-session"
);
assertTrue(result.answer().startsWith("以下结论基于当前已获取证据"));
assertFalse(result.answer().contains("EXECUTOR_FINAL_ANSWER"));
assertTrue(result.answer().contains("当前缺口"));
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "chat_composer"), chatModel.agentCalls);
}
@Test
void executeChatComplexLowConfidenceConfirmedFactsOnlyUseDirectEvidence() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "LOW_CONFID",
"groundedness_score": 0.37,
"critical_fact_count": 3,
"facts_checked": [
{
"fact": "连接池耗尽 active=50/50",
"is_critical": true,
"verification": "direct_evidence",
"detail": "log evidence",
"evidence_refs": []
},
{
"fact": "临时扩容连接池到 80",
"is_critical": true,
"verification": "indirect_support",
"detail": "suggestion inferred from evidence",
"evidence_refs": []
},
{
"fact": "OOM 导致连接泄漏",
"is_critical": true,
"verification": "no_evidence",
"detail": "missing OOM log",
"evidence_refs": []
}
],
"rationale": "scripted low confidence"
}
""");
chatModel.composerOutput = "not-json";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-low-confid-direct-only-session"
);
assertTrue(result.answer().contains("已确认信息:\n- 连接池耗尽 active=50/50"));
assertTrue(result.answer().contains("80"));
assertTrue(result.answer().contains("suggestion inferred from evidence"));
assertTrue(result.answer().contains("missing OOM log"));
assertFalse(result.answer().contains("EXECUTOR_FINAL_ANSWER"));
}
@Test
void executeChatComplexFallsBackToLowConfidenceWhenVerifierOutputMissing() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel("", "");
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析订单支付超时的原因,并给出修复建议",
List.of(),
"sequential-missing-verifier-session"
);
assertTrue(result.answer().startsWith("以下结论基于当前已获取证据"));
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "chat_verifier"), chatModel.agentCalls);
}
@Test
void executeChatComplexFallsBackToLowConfidenceWhenVerifierJsonInvalid() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel("not-json");
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析订单支付超时的原因,并给出修复建议",
List.of(),
"sequential-invalid-verifier-session"
);
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier"), chatModel.agentCalls);
}
@Test
void executeChatComplexRejectOutputDoesNotLeakExecutorAnswer() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "REJECT",
"groundedness_score": 0.0,
"critical_fact_count": 1,
"facts_checked": [
{
"fact": "payment timeout root cause",
"is_critical": true,
"verification": "contradicted",
"detail": "scripted contradiction",
"evidence_refs": []
}
],
"rationale": "scripted reject"
}
""");
chatModel.composerOutput = "not-json";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析订单支付超时的原因,并给出修复建议",
List.of(),
"sequential-reject-session"
);
assertTrue(result.answer().startsWith("当前无法基于已获取证据生成可靠结论"));
assertFalse(result.answer().contains("EXECUTOR_FINAL_ANSWER"));
}
@Test
void executeChatComplexRunsPlannerExecutorVerifierInFixedOrder() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel();
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析订单支付超时的原因,并给出修复建议",
List.of(),
"sequential-workflow-session"
);
assertTrue(result.answer().contains("连接池 active 达到上限"));
assertFalse(result.answer().contains("\"answer_version\""));
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "chat_composer"), chatModel.agentCalls);
assertTrue(chatModel.sawVerifierPrompt);
}
@Test
void verifierReceivesStructuredExecutorPayloadFields() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel();
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-structured-executor-session"
);
assertTrue(result.answer().contains("连接池 active 达到上限"));
assertFalse(result.answer().contains("\"answer_version\""));
assertTrue(chatModel.verifierPromptText.contains("\"executor_structured_output\""));
assertTrue(chatModel.verifierPromptText.contains("\"executor_output_parse_status\""));
assertTrue(chatModel.verifierPromptText.contains("\"status\" : \"valid\""));
assertTrue(chatModel.verifierPromptText.contains("连接池 active 达到上限"));
}
@Test
void executeChatComplexRendersExecutorEvidenceV2InsteadOfRawJsonOnPass() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel();
chatModel.composerOutput = "not-json";
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [
{
"hypothesis_text": "连接泄漏可能参与了连接池耗尽",
"basis": "已有连接池满载证据,但缺少泄漏检测日志",
"needed_evidence": ["连接泄漏检测日志"]
}
],
"recommended_actions": [
{
"action_text": "补充查询连接池泄漏检测日志",
"reason": "用于确认是否存在连接未释放"
}
],
"missing_info": ["缺少连接泄漏检测日志"]
}
""";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-v2-render-session"
);
assertTrue(result.answer().contains("已确认信息"));
assertTrue(result.answer().contains("连接池 active 达到上限"));
assertTrue(result.answer().contains("建议下一步"));
assertFalse(result.answer().contains("\"answer_version\""));
assertFalse(result.answer().contains("executor_evidence_v2"));
}
@Test
void executeChatComplexPersistsGatekeeperResultInVerifierEvaluation() throws Exception {
ChatService chatService = createChatService();
SelfEvaluationMergeService mergeService =
(SelfEvaluationMergeService) ReflectionTestUtils.getField(chatService, "selfEvaluationMergeService");
ToolInvocationRepository invocationRepository =
(ToolInvocationRepository) ReflectionTestUtils.getField(chatService, "toolInvocationRepository");
when(invocationRepository.findBySessionIdOrderByIdAsc("sequential-gatekeeper-persist-session"))
.thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.sessionId("sequential-gatekeeper-persist-session")
.toolName("query_metrics")
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel();
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-gatekeeper-persist-session"
);
ArgumentCaptor<Map<String, Object>> captor = ArgumentCaptor.forClass(Map.class);
verify(mergeService).mergeVerifierEvaluation(isNull(), captor.capture());
Map<String, Object> verifierEvaluation = captor.getValue();
assertTrue(verifierEvaluation.containsKey("gatekeeper_result"));
assertTrue(verifierEvaluation.containsKey("prompt_audit"));
@SuppressWarnings("unchecked")
Map<String, Object> gatekeeperResult = (Map<String, Object>) verifierEvaluation.get("gatekeeper_result");
assertEquals("pass", gatekeeperResult.get("status"));
assertEquals("none", gatekeeperResult.get("severity"));
@SuppressWarnings("unchecked")
Map<String, Object> promptAudit = (Map<String, Object>) verifierEvaluation.get("prompt_audit");
assertEquals("chat-prompts-v1", promptAudit.get("version"));
@SuppressWarnings("unchecked")
List<Map<String, Object>> prompts = (List<Map<String, Object>>) promptAudit.get("prompts");
assertEquals(4, prompts.size());
assertTrue(prompts.stream().anyMatch(prompt ->
"chat_executor".equals(prompt.get("name"))
&& "chat-executor-v2".equals(prompt.get("version"))));
}
@Test
void executeChatComplexMapsClaimChecksToFactsCheckedAndPersistsBoth() throws Exception {
ChatService chatService = createChatService();
SelfEvaluationMergeService mergeService =
(SelfEvaluationMergeService) ReflectionTestUtils.getField(chatService, "selfEvaluationMergeService");
ToolInvocationRepository invocationRepository =
(ToolInvocationRepository) ReflectionTestUtils.getField(chatService, "toolInvocationRepository");
when(invocationRepository.findBySessionIdOrderByIdAsc("sequential-claim-check-session"))
.thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.sessionId("sequential-claim-check-session")
.toolName("query_metrics")
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "LOW_CONFID",
"groundedness_score": 0.32,
"critical_fact_count": 6,
"claim_checks": [
{"claim_id":"claim-1","claim_text":"CPU 使用率 92%","claim_type":"symptom","verification":"direct_observation","detail":"direct","evidence_refs":[{"trace_ref":"trace-1","tool_name":"query_metrics","source_invocation_ids":[101],"note":"cpu"}]},
{"claim_id":"claim-2","claim_text":"CPU 过高可能导致超时","claim_type":"risk","verification":"reasonable_inference","detail":"inference","evidence_refs":[]},
{"claim_id":"claim-3","claim_text":"CPU 是唯一根因","claim_type":"root_cause","verification":"overstated","detail":"too strong","evidence_refs":[]},
{"claim_id":"claim-4","claim_text":"缺少线程池证据","claim_type":"symptom","verification":"unsupported","detail":"missing","evidence_refs":[]},
{"claim_id":"claim-5","claim_text":"出现证据外错误码 ERR_FAKE","claim_type":"symptom","verification":"external_unknown","detail":"external","evidence_refs":[]},
{"claim_id":"claim-6","claim_text":"证据显示 CPU 很低","claim_type":"symptom","verification":"contradicted","detail":"conflict","evidence_refs":[]}
],
"facts_checked": [],
"rationale": "claim checks drive compatibility"
}
""");
chatModel.executorOutput = validExecutorV2Output();
chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-claim-check-session"
);
ArgumentCaptor<Map<String, Object>> captor = ArgumentCaptor.forClass(Map.class);
verify(mergeService).mergeVerifierEvaluation(isNull(), captor.capture());
Map<String, Object> verifierEvaluation = captor.getValue();
@SuppressWarnings("unchecked")
List<Map<String, Object>> claimChecks = (List<Map<String, Object>>) verifierEvaluation.get("claim_checks");
@SuppressWarnings("unchecked")
List<Map<String, Object>> factsChecked = (List<Map<String, Object>>) verifierEvaluation.get("facts_checked");
assertEquals(6, claimChecks.size());
assertEquals(6, factsChecked.size());
assertEquals("direct_evidence", factsChecked.get(0).get("verification"));
assertEquals("indirect_support", factsChecked.get(1).get("verification"));
assertEquals("indirect_support", factsChecked.get(2).get("verification"));
assertEquals("no_evidence", factsChecked.get(3).get("verification"));
assertEquals("no_evidence", factsChecked.get(4).get("verification"));
assertEquals("contradicted", factsChecked.get(5).get("verification"));
assertTrue(String.valueOf(factsChecked.get(0).get("fact")).startsWith("claim-1:"));
}
@Test
void executeChatComplexDowngradesPassToRejectWhenGatekeeperInvocationRefFails() throws Exception {
ChatService chatService = createChatService();
SelfEvaluationMergeService mergeService =
(SelfEvaluationMergeService) ReflectionTestUtils.getField(chatService, "selfEvaluationMergeService");
ToolInvocationRepository invocationRepository =
(ToolInvocationRepository) ReflectionTestUtils.getField(chatService, "toolInvocationRepository");
when(invocationRepository.findBySessionIdOrderByIdAsc("sequential-gatekeeper-fail-session"))
.thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.sessionId("sequential-gatekeeper-fail-session")
.toolName("query_metrics")
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"claim_checks": [
{"claim_id":"claim-1","claim_text":"连接池 active 达到上限","claim_type":"symptom","verification":"direct_observation","detail":"direct","evidence_refs":[]}
],
"facts_checked": [],
"rationale": "model tried pass"
}
""");
chatModel.composerOutput = "not-json";
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_id": 999,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-gatekeeper-fail-session"
);
assertTrue(result.answer().startsWith("当前无法基于已获取证据生成可靠结论"));
ArgumentCaptor<Map<String, Object>> captor = ArgumentCaptor.forClass(Map.class);
verify(mergeService).mergeVerifierEvaluation(isNull(), captor.capture());
assertEquals("REJECT", captor.getValue().get("verdict"));
}
@Test
void executeChatComplexDowngradesPassToLowConfidenceWhenExecutorOutputMalformed() throws Exception {
ChatService chatService = createChatService();
SelfEvaluationMergeService mergeService =
(SelfEvaluationMergeService) ReflectionTestUtils.getField(chatService, "selfEvaluationMergeService");
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 0,
"claim_checks": [],
"facts_checked": [],
"rationale": "model tried pass"
}
""");
chatModel.composerOutput = "not-json";
chatModel.executorOutput = "{ not-json";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-malformed-pass-session"
);
assertTrue(result.answer().startsWith("以下结论基于当前已获取证据"));
ArgumentCaptor<Map<String, Object>> captor = ArgumentCaptor.forClass(Map.class);
verify(mergeService).mergeVerifierEvaluation(isNull(), captor.capture());
assertEquals("LOW_CONFID", captor.getValue().get("verdict"));
}
@Test
void buildMethodToolsArrayIncludesLogsAndMetricsWhenAvailable() {
ChatService chatService = new ChatService();
DateTimeTools dateTimeTools = new DateTimeTools();
LookupKnowledgeTool lookupKnowledgeTool = new LookupKnowledgeTool();
QueryLogsTools queryLogsTools = new QueryLogsTools(mock(ToolInvocationRecorder.class));
QueryMetricsTools queryMetricsTools = new QueryMetricsTools(mock(ToolInvocationRecorder.class));
ReflectionTestUtils.setField(chatService, "dateTimeTools", dateTimeTools);
ReflectionTestUtils.setField(chatService, "lookupKnowledgeTool", lookupKnowledgeTool);
ReflectionTestUtils.setField(chatService, "queryLogsTools", queryLogsTools);
ReflectionTestUtils.setField(chatService, "queryMetricsTools", queryMetricsTools);
Object[] methodTools = chatService.buildMethodToolsArray();
assertEquals(4, methodTools.length);
assertSame(dateTimeTools, methodTools[0]);
assertSame(lookupKnowledgeTool, methodTools[1]);
assertSame(queryLogsTools, methodTools[2]);
assertSame(queryMetricsTools, methodTools[3]);
}
@Test
void createReactAgentInjectsSkillCatalogThroughAlibabaHook() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel();
SkillRegistry skillRegistry = ClasspathSkillRegistry.builder()
.classpathPath("skills")
.basePath("target/test-skills-cache")
.build();
ReflectionTestUtils.setField(chatService, "skillRegistry", skillRegistry);
ReactAgent agent = chatService.createReactAgent(chatModel, "BASE_TEST_PROMPT");
agent.call("diagnose mysql connection pool exhaustion");
assertTrue(chatModel.promptText.contains("BASE_TEST_PROMPT"));
assertTrue(chatModel.promptText.contains("## Skills System"));
assertTrue(chatModel.promptText.contains("diagnose-mysql-connection-pool"));
assertTrue(chatModel.promptText.contains("read_skill"));
}
@Test
void plannerGetsSkillMetadataAndExecutorGetsReadSkillTool() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel();
SkillRegistry skillRegistry = ClasspathSkillRegistry.builder()
.classpathPath("skills")
.basePath("target/test-skills-cache")
.build();
ReflectionTestUtils.setField(chatService, "skillRegistry", skillRegistry);
chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"diagnose mysql connection pool exhaustion",
List.of(),
"planner-skill-metadata-session"
);
assertTrue(chatModel.plannerPromptText.contains("\"skill_catalog\""));
assertTrue(chatModel.plannerPromptText.contains("diagnose-mysql-connection-pool"));
assertTrue(chatModel.plannerPromptText.contains("\"selected_skill\""));
assertFalse(chatModel.plannerPromptText.contains("## Skills System"));
assertFalse(chatModel.plannerPromptText.contains("read_skill"));
assertTrue(chatModel.executorPromptText.contains("## Skills System"));
assertTrue(chatModel.executorPromptText.contains("diagnose-mysql-connection-pool"));
assertTrue(chatModel.executorPromptText.contains("read_skill"));
assertTrue(chatModel.executorPromptText.contains("只允许对该 skill 调用一次 read_skill"));
assertFalse(chatModel.verifierPromptText.contains("diagnose-mysql-connection-pool"));
assertFalse(chatModel.verifierPromptText.contains("read_skill"));
}
private ChatService createChatService() {
ChatService chatService = new ChatService();
ChatSessionRepository chatSessionRepository = mock(ChatSessionRepository.class);
when(chatSessionRepository.findBySessionId(anyString())).thenReturn(Optional.empty());
when(chatSessionRepository.save(any(ChatSession.class))).thenAnswer(invocation -> invocation.getArgument(0));
DiagnosisRunRepository diagnosisRunRepository = mock(DiagnosisRunRepository.class);
when(diagnosisRunRepository.save(any(DiagnosisRun.class))).thenAnswer(invocation -> invocation.getArgument(0));
AtomicInteger stepId = new AtomicInteger(1);
AgentStepRepository agentStepRepository = mock(AgentStepRepository.class);
when(agentStepRepository.save(any(AgentStep.class))).thenAnswer(invocation -> {
AgentStep step = invocation.getArgument(0);
if (step.getId() == null) {
step.setId((long) stepId.getAndIncrement());
}
return step;
});
when(agentStepRepository.findById(any())).thenReturn(Optional.of(new AgentStep()));
when(agentStepRepository.findBySessionIdOrderByStepIndex(anyString())).thenReturn(List.of());
when(agentStepRepository.findByRunIdOrderByStepIndex(anyString())).thenReturn(List.of());
ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
when(toolInvocationRepository.countBySessionId(anyString())).thenReturn(0L);
when(toolInvocationRepository.countByRunId(anyString())).thenReturn(0L);
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(anyString())).thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.toolName("query_metrics")
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
.build()));
when(toolInvocationRepository.findByRunIdOrderByIdAsc(anyString())).thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.toolName("query_metrics")
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
.build()));
EvaluationService evaluationService = mock(EvaluationService.class);
RetrievedDocTracker retrievedDocTracker = mock(RetrievedDocTracker.class);
KnowledgeDomainService knowledgeDomainService = mock(KnowledgeDomainService.class);
when(knowledgeDomainService.buildKnowledgeMap()).thenReturn("");
ToolTraceSummaryService toolTraceSummaryService = mock(ToolTraceSummaryService.class);
when(toolTraceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of());
when(toolTraceSummaryService.buildVerifierTraceSummaryForRun(anyString(), anyString())).thenReturn(List.of());
SelfEvaluationMergeService selfEvaluationMergeService = mock(SelfEvaluationMergeService.class);
when(selfEvaluationMergeService.mergeVerifierEvaluation(any(), any())).thenReturn("{}");
ExecutorGatekeeperService executorGatekeeperService = new ExecutorGatekeeperService(toolInvocationRepository);
ReflectionTestUtils.setField(chatService, "dateTimeTools", new DateTimeTools());
ReflectionTestUtils.setField(chatService, "lookupKnowledgeTool", new LookupKnowledgeTool());
ReflectionTestUtils.setField(chatService, "queryLogsTools", new QueryLogsTools(mock(ToolInvocationRecorder.class)));
ReflectionTestUtils.setField(chatService, "chatSessionRepository", chatSessionRepository);
ReflectionTestUtils.setField(chatService, "diagnosisRunRepository", diagnosisRunRepository);
ReflectionTestUtils.setField(chatService, "agentStepRepository", agentStepRepository);
ReflectionTestUtils.setField(chatService, "toolInvocationRepository", toolInvocationRepository);
ReflectionTestUtils.setField(chatService, "evaluationService", evaluationService);
ReflectionTestUtils.setField(chatService, "retrievedDocTracker", retrievedDocTracker);
ReflectionTestUtils.setField(chatService, "knowledgeDomainService", knowledgeDomainService);
ReflectionTestUtils.setField(chatService, "toolTraceSummaryService", toolTraceSummaryService);
ReflectionTestUtils.setField(chatService, "selfEvaluationMergeService", selfEvaluationMergeService);
ReflectionTestUtils.setField(chatService, "executorGatekeeperService", executorGatekeeperService);
ReflectionTestUtils.setField(chatService, "verifierLowConfidenceThreshold", 0.5d);
ReflectionTestUtils.setField(chatService, "chatPlannerPrompt", "PLANNER_TEST_PROMPT");
ReflectionTestUtils.setField(chatService, "chatExecutorPrompt", "EXECUTOR_TEST_PROMPT");
ReflectionTestUtils.setField(chatService, "chatVerifierPrompt", "VERIFIER_TEST_PROMPT");
ReflectionTestUtils.setField(chatService, "chatComposerPrompt", "COMPOSER_TEST_PROMPT");
return chatService;
}
private String validExecutorV2Output() {
return """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
}
private String evidenceRefs(String rawPath, String text) {
return "{\"evidence_refs\":[{\"raw_path\":\"" + rawPath + "\",\"text\":\"" + text + "\"}]}";
}
private static final class ScriptedChatModel implements ChatModel {
private final java.util.ArrayList<String> agentCalls = new java.util.ArrayList<>();
private String promptText = "";
private String plannerPromptText = "";
private String executorPromptText = "";
private String verifierPromptText = "";
private String composerPromptText = "";
private String executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
private String composerOutput = """
{
"answer_summary": "已确认连接池 active 达到上限。",
"recommended_actions": [
{
"action_text": "补充查询连接池泄漏检测日志",
"reason": "用于确认是否存在连接未释放"
}
],
"user_facing_answer": "已确认连接池 active 达到上限。建议补充查询连接池泄漏检测日志。"
}
""";
private boolean sawVerifierPrompt;
private final java.util.List<String> verifierOutputs;
private int verifierOutputIndex;
private ScriptedChatModel() {
this("""
{
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"claim_checks": [
{"claim_id":"claim-1","claim_text":"连接池 active 达到上限","claim_type":"symptom","verification":"direct_observation","detail":"covered by scripted verifier","evidence_refs":[]}
],
"facts_checked": [],
"rationale": "scripted pass"
}
""");
}
private ScriptedChatModel(String verifierOutput) {
this.verifierOutputs = java.util.List.of(verifierOutput);
}
private ScriptedChatModel(String... verifierOutputs) {
this.verifierOutputs = java.util.List.of(verifierOutputs);
}
@Override
public ChatResponse call(Prompt prompt) {
promptText = prompt.getContents();
String text;
if (promptText.contains("PLANNER_TEST_PROMPT")) {
agentCalls.add("chat_planner");
plannerPromptText = promptText;
text = "PLANNER_PLAN";
} else if (promptText.contains("EXECUTOR_TEST_PROMPT")) {
agentCalls.add("chat_executor");
executorPromptText = promptText;
text = executorOutput;
} else if (promptText.contains("VERIFIER_TEST_PROMPT")) {
agentCalls.add("chat_verifier");
verifierPromptText = promptText;
sawVerifierPrompt = true;
int index = Math.min(verifierOutputIndex, verifierOutputs.size() - 1);
text = verifierOutputs.get(index);
verifierOutputIndex++;
} else if (promptText.contains("COMPOSER_TEST_PROMPT")) {
agentCalls.add("chat_composer");
composerPromptText = promptText;
text = composerOutput;
} else {
text = "UNEXPECTED_PROMPT";
}
return new ChatResponse(List.of(new Generation(new AssistantMessage(text))));
}
}
}
@@ -1,6 +1,7 @@
package com.superbiz.agent.service;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.JsonNode;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.ChatSession;
import com.superbiz.agent.domain.entity.DiagnosisRun;
@@ -20,6 +21,7 @@ import java.util.List;
import java.util.Optional;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
@@ -89,9 +91,54 @@ class DiagnosisTraceServiceTest {
assertEquals("first question", response.getRun().getQuery());
assertEquals(List.of("planner"),
response.getSteps().stream().map(DiagnosisTraceResponse.AgentStepTrace::getAgentName).toList());
assertEquals("COMPOSER",
response.getRun().getOrchestrationTrace().get("final_node"));
verify(diagnosisRunRepository, never()).findFirstBySessionIdOrderByCreatedAtDescIdDesc(any());
}
@Test
void orchestrationTraceIsParsedOnlyOnRunProjection() {
String sessionId = "trace-session-orchestration";
DiagnosisRun run = run(4L, sessionId, "run-orchestration", "question",
LocalDateTime.of(2026, 7, 17, 12, 0));
when(diagnosisRunRepository.findBySessionIdAndRunId(sessionId, run.getRunId()))
.thenReturn(Optional.of(run));
when(agentStepRepository.findByRunIdOrderByStepIndex(run.getRunId())).thenReturn(List.of());
when(toolInvocationRepository.findByRunIdOrderByIdAsc(run.getRunId())).thenReturn(List.of());
DiagnosisTraceResponse response = service.getTrace(sessionId, run.getRunId());
JsonNode json = new ObjectMapper().findAndRegisterModules()
.valueToTree(response);
assertEquals("stategraph-v1",
response.getRun().getOrchestrationTrace().get("version"));
assertFalse(json.has("orchestrationTrace"));
assertFalse(json.path("session").has("orchestrationTrace"));
assertTrue(json.path("run").path("orchestrationTrace").isObject());
assertFalse(json.path("run").has("orchestrationTraceRaw"));
}
@Test
void nullOrInvalidOrchestrationTraceFailsClosedOnRunProjection() {
String sessionId = "trace-session-invalid-orchestration";
DiagnosisRun run = run(5L, sessionId, "run-invalid-orchestration", "question",
LocalDateTime.of(2026, 7, 17, 12, 1));
run.setOrchestrationTrace("not-json");
when(diagnosisRunRepository.findBySessionIdAndRunId(sessionId, run.getRunId()))
.thenReturn(Optional.of(run));
when(agentStepRepository.findByRunIdOrderByStepIndex(run.getRunId())).thenReturn(List.of());
when(toolInvocationRepository.findByRunIdOrderByIdAsc(run.getRunId())).thenReturn(List.of());
DiagnosisTraceResponse invalid = service.getTrace(sessionId, run.getRunId());
assertNull(invalid.getRun().getOrchestrationTrace());
run.setOrchestrationTrace(null);
DiagnosisTraceResponse historical = service.getTrace(sessionId, run.getRunId());
assertNull(historical.getRun().getOrchestrationTrace());
}
@Test
void getTraceWithRunIdReturnsExactSecondRun() {
String sessionId = "trace-session-exact";
@@ -218,6 +265,7 @@ class DiagnosisTraceServiceTest {
.agentFlow("CHAT")
.answer("answer for " + runId)
.selfEvaluation("{\"verifier_evaluation\":{\"verdict\":\"PASS\"}}")
.orchestrationTrace("{\"version\":\"stategraph-v1\",\"transitions\":[],\"final_node\":\"COMPOSER\",\"termination_reason\":\"composer_completed\",\"degraded\":false,\"evidence_retry_count\":0}")
.feedback("useful")
.stepCount(1)
.toolCallCount(1)