feat(graph): cut over chat diagnosis stategraph

This commit is contained in:
zhuyongxin
2026-07-17 18:30:08 +08:00
parent 1460dd1e99
commit 99e490f227
36 changed files with 2640 additions and 1623 deletions
@@ -0,0 +1,3 @@
committed_at: 2026-07-17
checkpoint: Commit
authorization: user-requested-direct-implementation
@@ -0,0 +1,135 @@
## Context
复杂 Chat 的公开调用链为 `ChatController -> ChatService.executeChatWithStrategy -> executeChatComplex`。阶段 2 已提供真实 Node adapters、显式 Gatekeeper、verified input、evidence retry、Composer/Fallback 和 `DiagnosisGraphFactory`,但 `executeChatComplex` 仍维护 SequentialAgent 两轮循环、VerifierInputHook/ThreadLocal 状态和私有 Composer 调用。阶段 3 要在不改变 `/api/chat` 的前提下切换唯一生产编排,并增加 Run 级 orchestration trace 数据契约。
当前 Run 通过 `SessionContextHolder(sessionId, runId)` 绑定 AgentStep/ToolInvocation;DiagnosisRun 使用字符串 JSON 保存 self evaluation;TraceService 将 Run 映射为 `run` 和兼容 `session` 投影。`DiagnosisOrchestrationTraceBuilder` 已能从有界 events 生成无 Prompt/raw output 的摘要。数据库由 Flyway 管理且 Hibernate 使用 validate,因此实体、V012 migration 和 Trace 映射必须同批对齐。
本 change 是 L4:内部状态机从固定顺序改为条件图,Trace run 对象和 DB 增加字段。`/api/chat`、sessionId/runId、Agent 输出协议保持兼容。用户已冻结单轨切换、安全 Fallback=SUCCESS、只有未处理失败=FAILED、阶段 5 才 live E2E。
## Goals / Non-Goals
**Goals:**
- 让复杂 Chat 一次执行真实 CompiledGraph,删除 ChatService 的细粒度 Sequential/round 状态机。
- 保持现有 Agent prompt/tool/skill/logging 组装能力,但 Graph Verifier 不注册 VerifierInputHook。
- 用 runId 同时作为 Graph threadId 和 Run 审计边界,用 sessionId/runId metadata 维持 AgentStep/ToolInvocation 归属。
- 将 final state 显式映射为 answer、verifier evaluation、composer audit、orchestration trace 和 Run lifecycle。
- 在正常/handled fallback 路径持久化非空、紧凑、安全的 orchestration trace,并只在 Trace `run` 对象暴露解析结果。
- 更新 Verifier Prompt/self-evaluation 到 verified-only 数据模型。
- 用阶段 3 focused tests证明生产 cutover、Run/Trace/API/Prompt 边界,且不留下已知失败测试。
**Non-Goals:**
- 不改简单 Chat/AIOps 编排。
- 不删除 VerifierInputHook/VerifierContextHolder 类型;它们在阶段 5 清理,但复杂 Chat 不再引用。
- 不增加持久 Graph checkpoint、恢复、并行分支或新 Run status。
- 不回填历史 run,不给历史 null orchestration trace 设计兼容伪值。
- 不在本阶段全面重命名/收敛测试夹具;阶段 4 完成测试体系替换。
- 不运行 Maven live E2E、日志或数据库验收。
## Decisions
### 1. 单轨生产 cutover,不增加 feature flag
`executeChatComplex` 只构造当前 Run、RunnableConfig 和 diagnosis context,然后调用真实 Graph runtime。旧 SequentialAgent loop、score/feature-flag retry 和 ChatService 私有 Verifier/Composer 路由方法从生产类移除;不保留双运行、shadow compare 或 runtime fallback 到 Sequential。
选择单轨而不是 feature flag,因为用户冻结的阶段门禁已经提供 Git commit 回滚边界,双轨会继续维护两套重试、Gatekeeper 和安全材料真理源。数据库新列 nullable,因此代码回滚时列可保留,无需破坏性 migration down。
### 2. Agent builder 只迁移职责,不复制实现
新增专用复杂 Chat Graph runtime/factory,复用当前 Planner/Executor/Verifier/Composer Prompt、knowledge map、history、method tools、ToolCallbacks、skill hooks 和 AgentLoggingHook 组装规则。ChatService 只向它提供本次 ChatModel、callbacks、history 和 Run 上下文。Graph Verifier hooks 只有 AgentLoggingHook,不包含 VerifierInputHook;Gatekeeper 只由显式 Node 调用。
如果为控制改动风险暂时保留简单 Chat 的工具/Hook builder,复杂 Agent builder 仍只存在一份。不得在 Graph Node 或 runtime 内复制 prompt 文本或第二套 tool catalog。
### 3. Graph 初始 state 和 RunnableConfig 使用最小白名单
初始 state 包含:
- `diagnosis_context={query, original_query}`;不把完整 history 放入 Graph State,history 仅注入 Planner/Executor system prompt。
- `planner_mode=NORMAL`。
- planner/verifier/composer/evidence retry count 全部为 0。
- `orchestration_events` 初始为空。
RunnableConfig 使用 `threadId(runId)`,metadata 包含 `sessionId` 和 `runId`。这同时满足 Graph 隔离、AgentLoggingHook/ToolCallback 的 Run ownership 和 Gatekeeper current-run validation。
### 4. 流式执行捕获最后真实 state
Graph runtime 使用 `CompiledGraph.stream(initialState, config)` 顺序消费 NodeOutput,并保存最后一个真实 `OverAllState`。正常 END 以最后 state 的非空 `final_answer` 为成功条件。流式消费与原同步 invoke 一样在当前请求线程阻塞,但允许在未处理异常时取得已真实产生的 partial state。
若异常前已有 events,failure handler 可以 best-effort 构造/persist partial trace;若没有 event,不生成虚假 node/transition。Agent invocation failures已经由 adapters 转为显式 status,正常预期失败都应走 Graph Fallback 而不是抛出。
### 5. Graph result mapper 是唯一运行结果翻译层
新增显式 mapper 从 final/partial state读取:
- `final_answer`。
- Verifier status/model/effective verdict、score、claim/fact checks、rationale、round。
- Gatekeeper raw audit、verified Executor output/evidence。
- Composer audit。
- `DiagnosisOrchestrationTraceBuilder` 结果。
ChatService 不再读取 VerifierContextHolder。self-evaluation 继续写 `verifier_evaluation` container 以兼容 Trace/Eval 消费者,但数据来自 Graph State:`executor_structured_output` 只保存 verified projection;新增 `verified_evidence` 和显式 statuses;不保存 raw Executor 或完整 tool_trace_summary。若 Verifier 从未完成,evaluation 只记录可用 status/audit/prompt 信息,不伪造 verdict。
### 6. Run lifecycle 和持久化顺序按执行结果分离
成功路径:Graph 返回非空安全 answer -> 构造 trace/evaluation -> 设置 Run SUCCESS、answer、orchestrationTrace、duration/metrics -> 保存 -> 调用 `evaluateRun`。
handled Fallback 与 Composer 正常路径使用相同成功顺序;可信度由 effective verdict 或 trace degraded 表达。失败路径:未处理异常、final state/answer 缺失、trace invariant 失败或成功结果持久化失败 -> Run FAILED,保存错误答案、duration、已有可构造 partial trace 和 metrics;不调用 success Eval。若 failure save 本身失败,只记录明确 error,不能声称持久化成功。
### 7. Orchestration trace 是独立 nullable JSON 数据契约
新增 V012,仅执行:
`ALTER TABLE diagnosis_run ADD COLUMN orchestration_trace JSON NULL ... AFTER self_evaluation`。
DiagnosisRun 使用 `@JdbcTypeCode(SqlTypes.JSON)` + `String orchestrationTrace`,与 selfEvaluation 写法一致。nullable 只服务无回填 migration/历史 run;每个成功的新 StateGraph Chat run 应用层必须非空。序列化使用 `DiagnosisOrchestrationTrace.toMap()`,字段保持 snake_case JSON。
### 8. Trace API 只在 RunTrace 增加解析对象
`DiagnosisTraceResponse.RunTrace` 新增 `Map<String,Object> orchestrationTrace`。DiagnosisTraceService 从 `DiagnosisRun.orchestrationTrace` 解析后映射;顶层 response、ChatSessionTrace、兼容 SessionTrace 和 raw string 均不新增字段。Run list 也不扩展该字段。
历史/AIOps run 的 null 值原样返回 null,不进行 fallback synthesis。解析非法 JSON 时 fail closed 为 null,并由数据/测试暴露问题;本阶段不增加历史兼容分支。
### 9. Verifier Prompt 与 verified-only input 对齐
`chat-verifier-prompt.md` 输入改为:`diagnosis_context`、`verified_executor_output`、`verified_evidence`、`gatekeeper_audit`、`verdict_ceiling`、可选 `retry_context`。删除 `executor_final_answer`、完整 `tool_trace_summary`、Hook parse status 和“再次执行 Gatekeeper”语义。
Verifier 仍输出原 JSON contract。`evidence_refs` 通过 claim_id/source_invocation_id/tool_name/raw_path 对 verified_evidence 建立审计关联,不要求不可用的 tool_trace_summary.trace_ref。Prompt audit catalog/version 随输入契约升级并在所有 handled paths 持久化。
### 10. 测试分阶段但不允许已知失败
阶段 3 新增最小生产 cutover测试:Graph runtime/final mapping、threadId/metadata、SUCCESS/Fallback/FAILED lifecycle、trace storage/API location、Prompt forbidden fields、no Verifier step on pre-verification failure。运行相关 Graph/Trace/Controller/Eval 回归和 test compilation。
旧 `ChatServiceSequentialAgentTest` 若因单轨语义失效,阶段 3 必须删除/改写冲突断言或由 focused Graph integration 替代,不能留到阶段 4 才让测试恢复绿色。阶段 4 继续完成命名、夹具、分支矩阵和旧 Hook tests 的全面清理。
## Interface Impact
- 等级:L4。
- 外部兼容:`/api/chat` request/response 与 run identity 不变。
- 加法协议:Trace `run.orchestrationTrace`;DB `diagnosis_run.orchestration_trace`。
- 内部行为:Executor/Planner/Gatekeeper failure 可提前 Fallback,不再保证固定 Agent 顺序;score flag 不再控制 LOW_CONFID retry。
- 消费者:Trace UI/demo/eval 可读取新 run 字段但不强制历史值;阶段 5 demo script 才增加最终 E2E 断言。
- 回滚:revert 本阶段代码/Prompt/spec;保留 nullable DB 列。无双轨开关、无数据回填回滚。
## Risks / Trade-offs
- [流式 API 与同步 invoke 的终止语义不同] → focused test断言最后 state、END、事件顺序和异常 partial state;不依赖实现私有字段。
- [self-evaluation 字段变化影响 Eval fixture] → 保留 container、verdict/claim/fact/gatekeeper/composer/prompt audit 兼容键;新增 verified 字段,移除不安全 full summary 前同步 specs/tests。
- [Agent factory 移动破坏 skill/tool hooks] → 复用现有 buildHooks/buildMethodTools 规则,并断言 Planner/Executor/Verifier/Composer hooks和当前 Run metadata。
- [JSON 字段三层漂移] → V012、entity、DTO/service 和 tests同批提交;Maven test compilation + strict specs。
- [无法取得 partial state] → 不伪造 trace;标记 FAILED 并记录明确原因。预期 Agent failures 全部由 Node adapters handled。
- [阶段 3/4 测试边界重叠] → 阶段 3只保证 cutover 可验收且现有 suite 不红;阶段 4负责全面测试架构替换。
## Migration Plan
1. 先加 V012、entity/Trace DTO/service 与 isolated mapping tests。
2. 实现 Graph runtime/result mapper和 Prompt verified-only 更新,先用 fake agents验证 final/partial state。
3. 将 `executeChatComplex` 单轨切换并删除旧私有 Sequential 状态机逻辑,保留 Run、metrics、Eval和简单 Chat路径。
4. 运行 production cutover、Graph、Trace、Controller/Eval 回归与 test compilation;证明 DB schema diff 只有一个字段。
5. 归档、提交阶段 3;阶段 4 再全面替换测试体系。
部署时 Flyway 先加 nullable 列,随后新代码写入。回滚为 Git revert;旧代码忽略新增列,数据库不删除列。
## Open Questions
无。所有产品/协议边界均由 ISS-011 与用户确认冻结;实现期若发现 Graph API 无法提供真实 partial state,只能回写本 design/tasks 后使用不伪造的 FAILED 处理,不能扩大协议。
@@ -0,0 +1,75 @@
# Chat Diagnosis StateGraph ChatService Cutover
## Why
阶段 2 已交付可构造、可测试的真实 Diagnosis StateGraph Nodes,但复杂 Chat 生产入口仍由 `ChatService.executeChatComplex(...)` 创建 `SequentialAgent`、维护 LOW_CONFID 外层循环,并依赖 `VerifierInputHook`/`VerifierContextHolder` 回传隐式状态。阶段 3 需要正式切换生产编排,让 ChatService 只管理 Run 生命周期、Agent 装配、Graph 调用和结果持久化,同时把 Graph 路由摘要作为当前 Run 的独立 Trace 维度保存和查询。
## What Changes
- 将复杂 Chat 从 `SequentialAgent` 外层循环切换为一次 `CompiledGraph` invocation,初始 state 使用显式 diagnosis context 和独立计数器。
- 为 Planner、Executor、Verifier、Composer 构造现有 ReactAgent 实例,经 `ReactAgentDiagnosisInvoker` 注入真实 Node actions;Graph Verifier 不注册旧 `VerifierInputHook`。
- 使用 `runId` 作为 Graph `threadId`,并把 `sessionId`、`runId` 放入 RunnableConfig metadata,继续复用 AgentStep/ToolInvocation 的 Run 归属。
- 将 Graph 最终 state 映射为现有 `ChatResult`、Verifier self-evaluation、Run 指标和生命周期状态。
- 新增 `diagnosis_run.orchestration_trace` nullable JSON 列、实体字段和紧凑序列化;新 StateGraph Chat run 在应用契约上必须写入非空摘要。
- Trace API 只在 `run.orchestrationTrace` 返回解析 JSON,不在顶层、兼容 `session` 投影或 raw 字段重复。
- 同步 Verifier Prompt 到 verified-only Graph 输入,移除完整 tool trace、raw Executor 和 Hook gatekeeper 输入说明。
- 增加生产 cutover、Run/Trace、safe Fallback、thread/metadata、Prompt 边界的必要单元/集成测试。
## Capabilities
### New Capabilities
- `chat-diagnosis-stategraph-chatservice-cutover`:规定复杂 Chat 的 StateGraph 生产调用、Run 生命周期、结果映射、编排摘要持久化与 Trace API 投影。
### Modified Capabilities
- `chat-diagnosis-stategraph-real-nodes`:阶段 2 的生产隔离约束改为阶段 3 正式切换,真实 Nodes 成为复杂 Chat 唯一生产编排。
- `chat-verifier-agent`:Verifier 输入改为 Gatekeeper passed binding 投影后的 verified-only payload;self-evaluation 以 Graph state 为来源。
- `chat-composer-agent`:Composer 由 Graph Node 调用,但 allowed-material 与审计契约保持不变。
- `session-run-trace-isolation`:Run 增加独立 orchestration trace,Trace API 只在精确 run 对象暴露解析结果。
## Scope
### In Scope
- `ChatService.executeChatComplex(...)` 的完整生产 cutover 和不再使用 SequentialAgent 的私有编排逻辑清理。
- 现有 Prompt/Agent builder 的 Graph 适配;不复制第二套 Agent factory。
- DiagnosisRun、Flyway、Trace DTO/Service 的加法式 orchestration trace 支持。
- Graph final state 到 answer、verifier evaluation、composer audit、Run status/metrics 的映射。
- 阶段 3 风险所需的 focused tests 与现有 Controller/Trace/Eval 契约回归。
### Out of Scope
- 不删除尚未被其他历史测试引用的 `VerifierInputHook`、`VerifierContextHolder` 类;阶段 5 统一清理。
- 不在阶段 3 完成整个测试体系命名/夹具迁移;阶段 4 负责全面替换旧 Sequential 测试体系,但阶段 3 不允许留下已知失败测试。
- 不修改 `/api/chat` 请求/响应结构或 Executor/Verifier/Composer 输出协议。
- 不新增 Run status,不做历史 run 的 orchestration trace 回填或兼容读取分支。
- 不运行 Maven live E2E、`logs/` 或数据库查询;统一保留到阶段 5。
## Context Constraints
- 阶段 0–2 archives 和主 specs 是实现基线;Graph 的 Node、条件边、计数所有权和安全 Fallback 边界不得在本阶段复制或放宽。
- `diagnosis_run.status` 只表达执行生命周期;Composer、LOW_CONFID、Verifier REJECT 或固定安全 Fallback 只要生成安全答案均为 SUCCESS。
- 编排摘要必须由有界 `orchestration_events` 构造,不从日志反推,不保存 Prompt、reasoning、raw tool output 或 Graph State snapshot。
- self-evaluation 与 orchestration trace 分离;AgentStep/ToolInvocation 继续作为详细 Trace 数据源。
- 新 Graph Verifier 只接收 `verified_executor_output`、`verified_evidence`、Gatekeeper audit/ceiling、query 和 retry context。
- 运行失败时不得伪造未发生的 events;可取得的部分 state 才允许 best-effort 持久化。
## Acceptance
- 复杂 Chat 生产代码不再创建或调用 SequentialAgent,且只调用一次 CompiledGraph。
- RunnableConfig 的 threadId 等于 runId,metadata 同时包含当前 sessionId/runId。
- Graph final answer 非空时映射为现有 ChatResult,Run 为 SUCCESS;无法产生安全响应或未处理失败时 Run 为 FAILED。
- 每个新 StateGraph Chat run 都持久化非空、无敏感材料的 orchestration trace;Trace API 只在 `run.orchestrationTrace` 返回解析对象。
- AgentStep、ToolInvocation、self-evaluation、answer、duration、token/step/tool count 和 Eval 仍绑定当前 runId。
- Verifier Prompt 和实际 verified-only payload 一致,Graph Verifier 不注册旧 input Hook。
- Executor/Gatekeeper 等前置失败的 Graph 路径不产生 Verifier AgentStep。
- `/api/chat` 和证据协议保持不变;数据库 schema 只新增一个 nullable JSON 列。
- 阶段 3 focused tests、Trace/Controller/Eval 回归、test compilation 和 OpenSpec strict validation 通过;不运行 live E2E。
## Risks
- ChatService 当前同时承载 Agent factory、Run 生命周期和旧解析逻辑;cutover 必须做完整内部重构,避免留下双编排路径或死代码。
- self-evaluation 旧字段曾依赖 ThreadLocal/full tool summary;Graph 映射必须明确兼容字段和安全边界,不能重新泄漏未验真材料。
- Graph invoke 异常可能没有可用 final state;失败处理只能持久化真实可取得的 partial snapshot,不能编造路径。
- JSON entity/DDL/Trace DTO 必须保持同一字段语义,否则 JPA validate、数据库迁移或 API 解析会漂移。
@@ -0,0 +1,60 @@
## MODIFIED Requirements
### Requirement: Composer SHALL generate final user-facing Chat answers
The system SHALL invoke the Composer Graph Node after Verifier routing to generate the final user-facing Chat answer from Verifier-allowed material.
#### Scenario: Composer receives only filtered material
- **WHEN** the Composer Graph Node invokes its configured Agent
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
- **AND** the Composer input SHALL NOT contain raw tool output
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
#### Scenario: Composer outputs strict JSON
- **WHEN** Composer completes
- **THEN** it SHALL output exactly one JSON object
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
#### Scenario: Composer does not introduce new facts
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
### Requirement: Composer input SHALL honor Verifier claim checks
The Composer Graph Node SHALL construct Composer input by filtering verified Executor structured output through Verifier `claim_checks`.
#### Scenario: Passing claims become allowed claims
- **WHEN** a claim check verification is `direct_observation`
- **THEN** the Composer input builder SHALL include the matching verified Executor claim in `allowed_claims`
#### Scenario: Reasonable inferences remain bounded
- **WHEN** a claim check verification is `reasonable_inference`
- **THEN** the Composer input builder MAY include the matching verified Executor claim in `allowed_claims`
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
#### Scenario: Overstated claims are not confirmed findings
- **WHEN** a claim check verification is `overstated`
- **THEN** the Composer input builder SHALL NOT include the matching claim as a confirmed item in `allowed_claims`
- **AND** it MAY include it as `allowed_hypotheses` or represent it in `missing_info`
#### Scenario: Unsupported or external claims are withheld
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
- **THEN** the Composer input builder SHALL NOT include the matching claim in `allowed_claims`
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
### Requirement: Composer failures SHALL degrade safely
The system SHALL tolerate malformed or failed Composer execution without leaking raw JSON or unverified Executor material.
#### Scenario: malformed Composer output falls back safely
- **WHEN** Composer exhausts its fixed-input technical retry or returns a non-retryable failure
- **THEN** the deterministic Fallback Node SHALL produce a final answer using only filtered material
- **AND** the final answer SHALL NOT expose raw Composer output
- **AND** the final answer SHALL NOT expose raw Executor output
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
#### Scenario: Composer audit is persisted
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.composer_output` SHALL record parsed Composer audit when available
- **AND** handled Composer fallback SHALL be observable through orchestration trace and available status/reason fields
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
@@ -0,0 +1,107 @@
## ADDED Requirements
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
#### Scenario: Complex Chat starts
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
#### Scenario: Graph dependencies are assembled
- **WHEN** the complex Chat Graph is constructed
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
### Requirement: Graph invocation SHALL preserve current Run ownership
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
#### Scenario: Agent Node runs
- **WHEN** any Agent Node is invoked for a complex Chat run
- **THEN** its RunnableConfig threadId SHALL equal the current runId
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
#### Scenario: Gatekeeper validates Executor output
- **WHEN** the Gatekeeper Node runs
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
- **AND** data from another run in the same session SHALL NOT be considered
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
#### Scenario: Initial state is projected
- **WHEN** a complex Chat run enters Planner for the first time
- **THEN** `diagnosis_context` SHALL contain the current query/original query
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
#### Scenario: Counters are initialized
- **WHEN** the Graph starts
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
- **AND** Planner mode SHALL be NORMAL
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
#### Scenario: Composer completes
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
- **AND** EvaluationService SHALL evaluate the current runId
#### Scenario: Handled Fallback completes
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
- **THEN** DiagnosisRun SHALL be SUCCESS
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
- **AND** no DEGRADED Run status SHALL be introduced
#### Scenario: Graph cannot produce a safe response
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
- **AND** the system SHALL NOT report a successful Graph result
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
#### Scenario: Verifier completed
- **WHEN** final Graph State contains a completed Verifier result
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
- **AND** the compatibility `verdict` field SHALL equal effective verdict
#### Scenario: Pre-verification Fallback completed
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
- **AND** it SHALL NOT fabricate a model or effective verdict
#### Scenario: Evaluation payload is inspected
- **WHEN** verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
### Requirement: Stage 3 verification SHALL not run the final live E2E
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
#### Scenario: Stage 3 is accepted
- **WHEN** stage 3 verification completes
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5
@@ -0,0 +1,7 @@
## REMOVED Requirements
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
**Reason**: Stage 2 production isolation has completed its migration purpose; stage 3 intentionally makes the real Diagnosis Graph the only complex Chat production orchestrator.
**Migration**: Replace `ChatService.executeChatComplex` SequentialAgent orchestration with real Graph assembly/invocation. Keep `/api/chat` compatible and use a full Git revert of stage 3 for runtime rollback rather than retaining dual paths.
@@ -0,0 +1,263 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical claims in `verified_executor_output.claims` have direct observation or reasonable inference support in `verified_evidence`
- **AND** at least one critical claim has direct observation
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
- **AND** verdict ceiling is PASS
- **THEN** the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical claim contradicts verified evidence
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only inference support
- **WHEN** no critical claim contradicts verified evidence
- **AND** all critical claims are only `reasonable_inference`
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical claim in `verified_executor_output.claims` contradicts verified evidence
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
- **THEN** the Verifier SHALL output model verdict="REJECT"
#### Scenario: Verified structured claims are the only verification target
- **WHEN** `verified_executor_output.claims` is present
- **THEN** Verifier SHALL verify each structured claim against matching `verified_evidence` through `claim_checks`
- **AND** each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
- **AND** a claim without matching verified evidence SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT receive or add confirmed facts from raw Executor text
#### Scenario: Executor output is invalid
- **WHEN** Executor does not return a legal structured contract
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute or fabricate a diagnostic verdict
### Requirement: facts_checked SHALL use a fixed classification set
The system SHALL continue to expose compatibility `facts_checked` using its fixed verification classification set.
#### Scenario: claim checks are mapped to legacy facts
- **WHEN** Verifier output contains `claim_checks`
- **THEN** the shared Verifier protocol parser SHALL derive compatibility `facts_checked` when the model did not provide them
- **AND** `direct_observation` SHALL map to `direct_evidence`
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
- **AND** `contradicted` SHALL map to `contradicted`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.
#### Scenario: PASS routes to Composer
- **WHEN** Verifier completes with effective verdict="PASS"
- **THEN** the Graph SHALL invoke Composer with filtered Verifier-allowed material
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
#### Scenario: LOW_CONFID does not qualify for evidence retry
- **WHEN** Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
- **THEN** the Graph SHALL route to Composer without another Planner cycle
- **AND** the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps
#### Scenario: LOW_CONFID qualifies for evidence retry
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
- **THEN** the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
- **AND** it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier completes with effective verdict="REJECT"
- **THEN** the Graph SHALL NOT start an evidence supplementation round
- **AND** it SHALL route to Composer-safe output
#### Scenario: REJECT produces bounded output
- **WHEN** effective verdict is REJECT
- **THEN** the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
- **AND** it SHALL NOT pass through raw Executor answer
- **AND** it SHALL NOT include an unsupported root-cause conclusion
#### Scenario: Verifier execution fails
- **WHEN** Verifier exhausts technical retry or returns a non-retryable failure
- **THEN** the Graph SHALL route to pre-verification Fallback
- **AND** no execution status string SHALL be used as model or effective verdict
### Requirement: Verifier SHALL be observable
The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.
#### Scenario: claim checks written to self_evaluation
- **WHEN** a completed Verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** it SHALL include `verifier_status`, `model_verdict`, `effective_verdict`, `verdict`, `groundedness_score`, `rationale`, verified output/evidence, and Gatekeeper audit
#### Scenario: composer output written to self_evaluation
- **WHEN** final answer composition completes
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include compact `composer_output` when available
- **AND** handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
- **AND** existing claim/fact and Gatekeeper fields SHALL be preserved
#### Scenario: verdict written to self_evaluation
- **WHEN** Verifier completes
- **THEN** Graph result mapping SHALL write effective verdict under `diagnosis_run.self_evaluation.verifier_evaluation.verdict`
- **AND** existing `rule_evaluation` and `aiops_rule_evaluation` channels SHALL be preserved
#### Scenario: pre-verification fallback is persisted
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
- **AND** it SHALL NOT fabricate `model_verdict` or `effective_verdict`
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** Graph result mapping persists available Gatekeeper state
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** a complex Chat Graph result is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_run.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat Prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
- **AND** full Prompt text SHALL NOT be persisted
#### Scenario: prompt audit available on fallback paths
- **WHEN** Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
#### Scenario: evaluation payload is inspected
- **WHEN** Graph verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete `tool_trace_summary`
- **AND** compatibility `executor_structured_output` SHALL contain at most the verified projection
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier Graph Node starts
- **THEN** the payload SHALL provide `diagnosis_context`, `verified_executor_output`, `verified_evidence`, `gatekeeper_audit`, and `verdict_ceiling`
- **AND** permitted structured `retry_context` SHALL be provided only after evidence retry preparation
#### Scenario: Verifier remains isolated from intermediate and raw material
- **WHEN** the Verifier input is serialized
- **THEN** it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State
#### Scenario: only passed bindings are available
- **WHEN** Gatekeeper returns mixed passed and failed checked bindings
- **THEN** `verified_executor_output` and `verified_evidence` SHALL contain only claims/material matching passed bindings
- **AND** the Verifier SHALL NOT receive failed or unreferenced tool material
#### Scenario: verified evidence preserves precise references
- **WHEN** the system prepares Verifier input
- **THEN** each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
- **AND** the item SHALL be traceable to current-run Gatekeeper validation
#### Scenario: gatekeeper audit and ceiling are available
- **WHEN** the system prepares Verifier input
- **THEN** the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
- **AND** a LOW_CONFID ceiling SHALL prevent effective PASS
#### Scenario: technical retry occurs
- **WHEN** the first Verifier attempt returns invalid output or a retryable invocation failure
- **THEN** the second attempt SHALL receive byte-identical serialized input
- **AND** Executor, Gatekeeper, and tools SHALL NOT rerun
### Requirement: Verifier facts SHALL be auditable
Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.
#### Scenario: claim checks contain evidence refs
- **WHEN** the Verifier emits `claim_checks`
- **THEN** each check SHALL include an `evidence_refs` array
- **AND** any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
- **AND** it SHALL NOT reference a failed or unverified binding
#### Scenario: verifier evaluation persists traceability snapshot
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the bounded `verified_evidence` snapshot used by the Verifier
- **AND** it SHALL NOT persist a complete tool trace summary as Verifier input
### Requirement: Structured Executor output SHALL degrade safely
The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.
#### Scenario: Malformed Executor JSON is classified
- **WHEN** Executor returns malformed JSON or text outside the expected contract
- **THEN** Executor Node SHALL set INVALID_OUTPUT
- **AND** the Graph SHALL route directly to deterministic pre-verification Fallback
- **AND** Gatekeeper, Verifier, and model Composer SHALL NOT execute
#### Scenario: Structured parse failure remains observable
- **WHEN** Executor output parsing fails
- **THEN** orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
- **AND** the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: schema rule rejects missing evidence bindings
- **WHEN** a confirmed claim has no `evidence_bindings`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references an invocation id absent from current-run `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing current-run invocation id
- **AND** binding `tool_name` does not match persisted invocation `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits legal `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
#### Scenario: gatekeeper reject bypasses Verifier
- **WHEN** normalized Gatekeeper status is REJECT
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute
#### Scenario: gatekeeper low confidence is bounded
- **WHEN** normalized Gatekeeper status is LOW_CONFID with at least one passed binding
- **THEN** verified input SHALL contain only passed bindings
- **AND** effective verdict SHALL NOT exceed LOW_CONFID
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.
#### Scenario: Verified excerpt supports direct observation
- **WHEN** verdict ceiling is PASS
- **AND** a claim's verified evidence matched text directly contains the claim's concrete facts
- **THEN** Verifier MAY classify that claim as `direct_observation`
#### Scenario: Verified evidence is complete Verifier context
- **WHEN** verified claims and evidence are available
- **THEN** Verifier SHALL use them as its evidence context
- **AND** it SHALL NOT require or request a complete tool trace summary
- **AND** it SHALL NOT read raw Executor or unreferenced tool material
### Requirement: Gatekeeper severity SHALL constrain effective verdict
Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.
#### Scenario: Reject severity bypasses Verifier
- **WHEN** `gatekeeper_result.severity=reject`
- **THEN** the Graph SHALL route to pre-verification Fallback without invoking Verifier
- **AND** it SHALL NOT fabricate an effective diagnostic verdict
#### Scenario: Low confidence severity prevents PASS
- **WHEN** `gatekeeper_result.severity=low_confid`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** deterministic effective-verdict calculation SHALL downgrade the result
- **AND** effective verdict SHALL be `LOW_CONFID`
#### Scenario: Gatekeeper audit includes severity
- **WHEN** Graph verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include available `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
@@ -0,0 +1,56 @@
## ADDED Requirements
### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`.
#### Scenario: Graph reaches Composer
- **WHEN** a complex Chat Graph terminates through Composer with a safe answer
- **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
- **AND** the summary SHALL be derived from the current Run's actual orchestration events
#### Scenario: Graph reaches handled Fallback
- **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer
- **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true`
- **AND** the Run status SHALL be SUCCESS
#### Scenario: Unhandled execution fails after events exist
- **WHEN** an unhandled failure occurs after one or more real Graph events are available
- **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run
- **AND** it SHALL NOT add a node or transition that did not occur
#### Scenario: Orchestration trace content is inspected
- **WHEN** orchestration trace JSON is serialized
- **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
- **AND** it SHALL NOT contain data owned by another run
### Requirement: Trace API SHALL expose orchestration trace only on the run object
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`.
#### Scenario: Exact StateGraph run trace is queried
- **WHEN** a caller queries a successful new StateGraph Chat run
- **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object
- **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field
- **AND** no raw orchestration trace field SHALL be added
#### Scenario: Historical or non-StateGraph run is queried
- **WHEN** the selected DiagnosisRun has null orchestration trace
- **THEN** `run.orchestrationTrace` MAY be null
- **AND** the service SHALL NOT synthesize historical events or read another run's trace
### Requirement: Orchestration trace migration SHALL be additive and nullable
The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change.
#### Scenario: Migration is applied
- **WHEN** Flyway applies the stage 3 migration
- **THEN** existing DiagnosisRun rows SHALL remain valid without backfill
- **AND** no other table or column SHALL be changed by the stage 3 schema migration
@@ -0,0 +1,43 @@
## 1. Run Orchestration Trace Contract
- [x] 1.1 Add V012 migration containing only nullable `diagnosis_run.orchestration_trace JSON` and document that historical rows are not backfilled.
- [x] 1.2 Add DiagnosisRun JSON field and `DiagnosisTraceResponse.RunTrace.orchestrationTrace` parsed-map field without adding top-level, session, summary, run-list, or raw duplicates.
- [x] 1.3 Update DiagnosisTraceService exact/latest Run mapping and tests for current-run parsed trace, historical null, invalid JSON fail-closed, and no cross-run/session projection.
- [x] 1.4 Add entity/repository or migration source checks proving stage 3 changes no schema object except the one nullable column.
## 2. Graph Runtime And Result Mapping
- [x] 2.1 Implement a dedicated complex Chat Graph runtime that assembles existing real actions, compiles the Graph once per request/runtime instance, and consumes `CompiledGraph.stream` while retaining only the last real state.
- [x] 2.2 Define minimal initial state with query-only diagnosis context, NORMAL mode, independent zero counters, and empty bounded events.
- [x] 2.3 Build RunnableConfig with `threadId=runId` and current `sessionId`/`runId` metadata; prove config identity at every real Agent invoker and Gatekeeper boundary.
- [x] 2.4 Implement explicit final/partial Graph result mapping for answer, statuses/verdicts, verified output/evidence, Gatekeeper audit, Composer audit, failure reason, and orchestration trace.
- [x] 2.5 Add runtime/mapper tests for Composer success, handled pre/post-verification Fallback, blank answer, no state, thrown execution with real partial events, and no fabricated event.
## 3. Agent Assembly And Verified-only Prompt
- [x] 3.1 Move or encapsulate the existing complex Planner/Executor/Verifier/Composer builders so Graph runtime reuses current Prompt, knowledge map, history, method tool, ToolCallback, skill, and AgentLoggingHook rules without a duplicate factory.
- [x] 3.2 Remove VerifierInputHook from the Graph Verifier while retaining AgentLoggingHook, and prove explicit Gatekeeper executes once per Executor round.
- [x] 3.3 Update `chat-verifier-prompt.md` to the verified-only input contract and remove raw Executor, full tool trace, Hook parse, and Gatekeeper re-execution language.
- [x] 3.4 Bump Prompt audit catalog/Verifier Prompt version and add static/behavior tests proving the runtime payload and Prompt declare the same allowed fields.
## 4. ChatService Production Cutover
- [x] 4.1 Replace `executeChatComplex` SequentialAgent/outer retry loop with one Graph runtime call while preserving session/run creation and the public ChatResult signature.
- [x] 4.2 Remove obsolete ChatService private Sequential workflow, score/feature-flag retry, ThreadLocal verifier parsing, private Composer routing, and dead imports without changing simple Chat behavior.
- [x] 4.3 Persist Graph-derived verifier evaluation with compatibility verdict/claim/fact/gatekeeper/composer/prompt-audit fields, verified-only evidence, and no fabricated verdict before Verifier completion.
- [x] 4.4 Persist SUCCESS only for a non-empty safe Graph answer with non-empty orchestration trace, then backfill current-run metrics and call Eval; persist FAILED for unhandled/no-answer/invariant failures with any real partial trace available.
- [x] 4.5 Ensure `SessionContextHolder` and retrieved-doc cleanup still run in finally, no complex Chat code reads `VerifierContextHolder`, and no production path creates SequentialAgent.
## 5. Stage 3 Focused Tests And Alignment
- [x] 5.1 Add production cutover integration tests for `/api/chat`-compatible ChatResult, `agent_flow=CHAT`, current Run ownership, Graph thread/metadata, answer/metrics/Eval persistence, and multi-run isolation.
- [x] 5.2 Add handled failure tests proving Executor invalid/failure and Gatekeeper REJECT do not create Verifier AgentStep, safe Fallback is SUCCESS/degraded, and unhandled/no-answer paths are FAILED.
- [x] 5.3 Update or replace stage-3-conflicting Sequential tests so no known test remains red; preserve Controller, Gatekeeper, Trace, Repository, Eval, no-evidence, REJECT, and Composer safety regressions.
- [x] 5.4 Confirm the first completed Trace slice and final production assembly against design/specs, including single-track cutover, no duplicate Agent factory, and no core TODO/placeholder.
## 6. Stage 3 Verification And Handoff
- [x] 6.1 Run Graph runtime/result mapper, ChatService cutover, DiagnosisTraceService, Controller, Repository, Gatekeeper, Composer, and Eval focused tests.
- [x] 6.2 Run Maven test compilation and any existing tests whose public contracts were touched; classify failures as OpenSpec gap, code deviation, or out-of-scope environment issue.
- [x] 6.3 Run OpenSpec strict validation, all main specs strict validation, `git diff --check`, source/reference isolation checks, and a schema whitelist check.
- [x] 6.4 Record that Maven live E2E, `logs/`, and `scripts/query_mysql.py` database verification were intentionally not run in stage 3 and remain reserved for stage 5.