feat(graph): cut over chat diagnosis stategraph

This commit is contained in:
zhuyongxin
2026-07-17 18:30:08 +08:00
parent 1460dd1e99
commit 99e490f227
36 changed files with 2640 additions and 1623 deletions
@@ -0,0 +1,3 @@
committed_at: 2026-07-17
checkpoint: Commit
authorization: user-requested-direct-implementation
@@ -0,0 +1,135 @@
## Context
复杂 Chat 的公开调用链为 `ChatController -> ChatService.executeChatWithStrategy -> executeChatComplex`。阶段 2 已提供真实 Node adapters、显式 Gatekeeper、verified input、evidence retry、Composer/Fallback 和 `DiagnosisGraphFactory`,但 `executeChatComplex` 仍维护 SequentialAgent 两轮循环、VerifierInputHook/ThreadLocal 状态和私有 Composer 调用。阶段 3 要在不改变 `/api/chat` 的前提下切换唯一生产编排,并增加 Run 级 orchestration trace 数据契约。
当前 Run 通过 `SessionContextHolder(sessionId, runId)` 绑定 AgentStep/ToolInvocation;DiagnosisRun 使用字符串 JSON 保存 self evaluation;TraceService 将 Run 映射为 `run` 和兼容 `session` 投影。`DiagnosisOrchestrationTraceBuilder` 已能从有界 events 生成无 Prompt/raw output 的摘要。数据库由 Flyway 管理且 Hibernate 使用 validate,因此实体、V012 migration 和 Trace 映射必须同批对齐。
本 change 是 L4:内部状态机从固定顺序改为条件图,Trace run 对象和 DB 增加字段。`/api/chat`、sessionId/runId、Agent 输出协议保持兼容。用户已冻结单轨切换、安全 Fallback=SUCCESS、只有未处理失败=FAILED、阶段 5 才 live E2E。
## Goals / Non-Goals
**Goals:**
- 让复杂 Chat 一次执行真实 CompiledGraph,删除 ChatService 的细粒度 Sequential/round 状态机。
- 保持现有 Agent prompt/tool/skill/logging 组装能力,但 Graph Verifier 不注册 VerifierInputHook。
- 用 runId 同时作为 Graph threadId 和 Run 审计边界,用 sessionId/runId metadata 维持 AgentStep/ToolInvocation 归属。
- 将 final state 显式映射为 answer、verifier evaluation、composer audit、orchestration trace 和 Run lifecycle。
- 在正常/handled fallback 路径持久化非空、紧凑、安全的 orchestration trace,并只在 Trace `run` 对象暴露解析结果。
- 更新 Verifier Prompt/self-evaluation 到 verified-only 数据模型。
- 用阶段 3 focused tests证明生产 cutover、Run/Trace/API/Prompt 边界,且不留下已知失败测试。
**Non-Goals:**
- 不改简单 Chat/AIOps 编排。
- 不删除 VerifierInputHook/VerifierContextHolder 类型;它们在阶段 5 清理,但复杂 Chat 不再引用。
- 不增加持久 Graph checkpoint、恢复、并行分支或新 Run status。
- 不回填历史 run,不给历史 null orchestration trace 设计兼容伪值。
- 不在本阶段全面重命名/收敛测试夹具;阶段 4 完成测试体系替换。
- 不运行 Maven live E2E、日志或数据库验收。
## Decisions
### 1. 单轨生产 cutover,不增加 feature flag
`executeChatComplex` 只构造当前 Run、RunnableConfig 和 diagnosis context,然后调用真实 Graph runtime。旧 SequentialAgent loop、score/feature-flag retry 和 ChatService 私有 Verifier/Composer 路由方法从生产类移除;不保留双运行、shadow compare 或 runtime fallback 到 Sequential。
选择单轨而不是 feature flag,因为用户冻结的阶段门禁已经提供 Git commit 回滚边界,双轨会继续维护两套重试、Gatekeeper 和安全材料真理源。数据库新列 nullable,因此代码回滚时列可保留,无需破坏性 migration down。
### 2. Agent builder 只迁移职责,不复制实现
新增专用复杂 Chat Graph runtime/factory,复用当前 Planner/Executor/Verifier/Composer Prompt、knowledge map、history、method tools、ToolCallbacks、skill hooks 和 AgentLoggingHook 组装规则。ChatService 只向它提供本次 ChatModel、callbacks、history 和 Run 上下文。Graph Verifier hooks 只有 AgentLoggingHook,不包含 VerifierInputHook;Gatekeeper 只由显式 Node 调用。
如果为控制改动风险暂时保留简单 Chat 的工具/Hook builder,复杂 Agent builder 仍只存在一份。不得在 Graph Node 或 runtime 内复制 prompt 文本或第二套 tool catalog。
### 3. Graph 初始 state 和 RunnableConfig 使用最小白名单
初始 state 包含:
- `diagnosis_context={query, original_query}`;不把完整 history 放入 Graph State,history 仅注入 Planner/Executor system prompt。
- `planner_mode=NORMAL`。
- planner/verifier/composer/evidence retry count 全部为 0。
- `orchestration_events` 初始为空。
RunnableConfig 使用 `threadId(runId)`,metadata 包含 `sessionId` 和 `runId`。这同时满足 Graph 隔离、AgentLoggingHook/ToolCallback 的 Run ownership 和 Gatekeeper current-run validation。
### 4. 流式执行捕获最后真实 state
Graph runtime 使用 `CompiledGraph.stream(initialState, config)` 顺序消费 NodeOutput,并保存最后一个真实 `OverAllState`。正常 END 以最后 state 的非空 `final_answer` 为成功条件。流式消费与原同步 invoke 一样在当前请求线程阻塞,但允许在未处理异常时取得已真实产生的 partial state。
若异常前已有 events,failure handler 可以 best-effort 构造/persist partial trace;若没有 event,不生成虚假 node/transition。Agent invocation failures已经由 adapters 转为显式 status,正常预期失败都应走 Graph Fallback 而不是抛出。
### 5. Graph result mapper 是唯一运行结果翻译层
新增显式 mapper 从 final/partial state读取:
- `final_answer`。
- Verifier status/model/effective verdict、score、claim/fact checks、rationale、round。
- Gatekeeper raw audit、verified Executor output/evidence。
- Composer audit。
- `DiagnosisOrchestrationTraceBuilder` 结果。
ChatService 不再读取 VerifierContextHolder。self-evaluation 继续写 `verifier_evaluation` container 以兼容 Trace/Eval 消费者,但数据来自 Graph State:`executor_structured_output` 只保存 verified projection;新增 `verified_evidence` 和显式 statuses;不保存 raw Executor 或完整 tool_trace_summary。若 Verifier 从未完成,evaluation 只记录可用 status/audit/prompt 信息,不伪造 verdict。
### 6. Run lifecycle 和持久化顺序按执行结果分离
成功路径:Graph 返回非空安全 answer -> 构造 trace/evaluation -> 设置 Run SUCCESS、answer、orchestrationTrace、duration/metrics -> 保存 -> 调用 `evaluateRun`。
handled Fallback 与 Composer 正常路径使用相同成功顺序;可信度由 effective verdict 或 trace degraded 表达。失败路径:未处理异常、final state/answer 缺失、trace invariant 失败或成功结果持久化失败 -> Run FAILED,保存错误答案、duration、已有可构造 partial trace 和 metrics;不调用 success Eval。若 failure save 本身失败,只记录明确 error,不能声称持久化成功。
### 7. Orchestration trace 是独立 nullable JSON 数据契约
新增 V012,仅执行:
`ALTER TABLE diagnosis_run ADD COLUMN orchestration_trace JSON NULL ... AFTER self_evaluation`。
DiagnosisRun 使用 `@JdbcTypeCode(SqlTypes.JSON)` + `String orchestrationTrace`,与 selfEvaluation 写法一致。nullable 只服务无回填 migration/历史 run;每个成功的新 StateGraph Chat run 应用层必须非空。序列化使用 `DiagnosisOrchestrationTrace.toMap()`,字段保持 snake_case JSON。
### 8. Trace API 只在 RunTrace 增加解析对象
`DiagnosisTraceResponse.RunTrace` 新增 `Map<String,Object> orchestrationTrace`。DiagnosisTraceService 从 `DiagnosisRun.orchestrationTrace` 解析后映射;顶层 response、ChatSessionTrace、兼容 SessionTrace 和 raw string 均不新增字段。Run list 也不扩展该字段。
历史/AIOps run 的 null 值原样返回 null,不进行 fallback synthesis。解析非法 JSON 时 fail closed 为 null,并由数据/测试暴露问题;本阶段不增加历史兼容分支。
### 9. Verifier Prompt 与 verified-only input 对齐
`chat-verifier-prompt.md` 输入改为:`diagnosis_context`、`verified_executor_output`、`verified_evidence`、`gatekeeper_audit`、`verdict_ceiling`、可选 `retry_context`。删除 `executor_final_answer`、完整 `tool_trace_summary`、Hook parse status 和“再次执行 Gatekeeper”语义。
Verifier 仍输出原 JSON contract。`evidence_refs` 通过 claim_id/source_invocation_id/tool_name/raw_path 对 verified_evidence 建立审计关联,不要求不可用的 tool_trace_summary.trace_ref。Prompt audit catalog/version 随输入契约升级并在所有 handled paths 持久化。
### 10. 测试分阶段但不允许已知失败
阶段 3 新增最小生产 cutover测试:Graph runtime/final mapping、threadId/metadata、SUCCESS/Fallback/FAILED lifecycle、trace storage/API location、Prompt forbidden fields、no Verifier step on pre-verification failure。运行相关 Graph/Trace/Controller/Eval 回归和 test compilation。
旧 `ChatServiceSequentialAgentTest` 若因单轨语义失效,阶段 3 必须删除/改写冲突断言或由 focused Graph integration 替代,不能留到阶段 4 才让测试恢复绿色。阶段 4 继续完成命名、夹具、分支矩阵和旧 Hook tests 的全面清理。
## Interface Impact
- 等级:L4。
- 外部兼容:`/api/chat` request/response 与 run identity 不变。
- 加法协议:Trace `run.orchestrationTrace`;DB `diagnosis_run.orchestration_trace`。
- 内部行为:Executor/Planner/Gatekeeper failure 可提前 Fallback,不再保证固定 Agent 顺序;score flag 不再控制 LOW_CONFID retry。
- 消费者:Trace UI/demo/eval 可读取新 run 字段但不强制历史值;阶段 5 demo script 才增加最终 E2E 断言。
- 回滚:revert 本阶段代码/Prompt/spec;保留 nullable DB 列。无双轨开关、无数据回填回滚。
## Risks / Trade-offs
- [流式 API 与同步 invoke 的终止语义不同] → focused test断言最后 state、END、事件顺序和异常 partial state;不依赖实现私有字段。
- [self-evaluation 字段变化影响 Eval fixture] → 保留 container、verdict/claim/fact/gatekeeper/composer/prompt audit 兼容键;新增 verified 字段,移除不安全 full summary 前同步 specs/tests。
- [Agent factory 移动破坏 skill/tool hooks] → 复用现有 buildHooks/buildMethodTools 规则,并断言 Planner/Executor/Verifier/Composer hooks和当前 Run metadata。
- [JSON 字段三层漂移] → V012、entity、DTO/service 和 tests同批提交;Maven test compilation + strict specs。
- [无法取得 partial state] → 不伪造 trace;标记 FAILED 并记录明确原因。预期 Agent failures 全部由 Node adapters handled。
- [阶段 3/4 测试边界重叠] → 阶段 3只保证 cutover 可验收且现有 suite 不红;阶段 4负责全面测试架构替换。
## Migration Plan
1. 先加 V012、entity/Trace DTO/service 与 isolated mapping tests。
2. 实现 Graph runtime/result mapper和 Prompt verified-only 更新,先用 fake agents验证 final/partial state。
3. 将 `executeChatComplex` 单轨切换并删除旧私有 Sequential 状态机逻辑,保留 Run、metrics、Eval和简单 Chat路径。
4. 运行 production cutover、Graph、Trace、Controller/Eval 回归与 test compilation;证明 DB schema diff 只有一个字段。
5. 归档、提交阶段 3;阶段 4 再全面替换测试体系。
部署时 Flyway 先加 nullable 列,随后新代码写入。回滚为 Git revert;旧代码忽略新增列,数据库不删除列。
## Open Questions
无。所有产品/协议边界均由 ISS-011 与用户确认冻结;实现期若发现 Graph API 无法提供真实 partial state,只能回写本 design/tasks 后使用不伪造的 FAILED 处理,不能扩大协议。
@@ -0,0 +1,75 @@
# Chat Diagnosis StateGraph ChatService Cutover
## Why
阶段 2 已交付可构造、可测试的真实 Diagnosis StateGraph Nodes,但复杂 Chat 生产入口仍由 `ChatService.executeChatComplex(...)` 创建 `SequentialAgent`、维护 LOW_CONFID 外层循环,并依赖 `VerifierInputHook`/`VerifierContextHolder` 回传隐式状态。阶段 3 需要正式切换生产编排,让 ChatService 只管理 Run 生命周期、Agent 装配、Graph 调用和结果持久化,同时把 Graph 路由摘要作为当前 Run 的独立 Trace 维度保存和查询。
## What Changes
- 将复杂 Chat 从 `SequentialAgent` 外层循环切换为一次 `CompiledGraph` invocation,初始 state 使用显式 diagnosis context 和独立计数器。
- 为 Planner、Executor、Verifier、Composer 构造现有 ReactAgent 实例,经 `ReactAgentDiagnosisInvoker` 注入真实 Node actions;Graph Verifier 不注册旧 `VerifierInputHook`。
- 使用 `runId` 作为 Graph `threadId`,并把 `sessionId`、`runId` 放入 RunnableConfig metadata,继续复用 AgentStep/ToolInvocation 的 Run 归属。
- 将 Graph 最终 state 映射为现有 `ChatResult`、Verifier self-evaluation、Run 指标和生命周期状态。
- 新增 `diagnosis_run.orchestration_trace` nullable JSON 列、实体字段和紧凑序列化;新 StateGraph Chat run 在应用契约上必须写入非空摘要。
- Trace API 只在 `run.orchestrationTrace` 返回解析 JSON,不在顶层、兼容 `session` 投影或 raw 字段重复。
- 同步 Verifier Prompt 到 verified-only Graph 输入,移除完整 tool trace、raw Executor 和 Hook gatekeeper 输入说明。
- 增加生产 cutover、Run/Trace、safe Fallback、thread/metadata、Prompt 边界的必要单元/集成测试。
## Capabilities
### New Capabilities
- `chat-diagnosis-stategraph-chatservice-cutover`:规定复杂 Chat 的 StateGraph 生产调用、Run 生命周期、结果映射、编排摘要持久化与 Trace API 投影。
### Modified Capabilities
- `chat-diagnosis-stategraph-real-nodes`:阶段 2 的生产隔离约束改为阶段 3 正式切换,真实 Nodes 成为复杂 Chat 唯一生产编排。
- `chat-verifier-agent`:Verifier 输入改为 Gatekeeper passed binding 投影后的 verified-only payload;self-evaluation 以 Graph state 为来源。
- `chat-composer-agent`:Composer 由 Graph Node 调用,但 allowed-material 与审计契约保持不变。
- `session-run-trace-isolation`:Run 增加独立 orchestration trace,Trace API 只在精确 run 对象暴露解析结果。
## Scope
### In Scope
- `ChatService.executeChatComplex(...)` 的完整生产 cutover 和不再使用 SequentialAgent 的私有编排逻辑清理。
- 现有 Prompt/Agent builder 的 Graph 适配;不复制第二套 Agent factory。
- DiagnosisRun、Flyway、Trace DTO/Service 的加法式 orchestration trace 支持。
- Graph final state 到 answer、verifier evaluation、composer audit、Run status/metrics 的映射。
- 阶段 3 风险所需的 focused tests 与现有 Controller/Trace/Eval 契约回归。
### Out of Scope
- 不删除尚未被其他历史测试引用的 `VerifierInputHook`、`VerifierContextHolder` 类;阶段 5 统一清理。
- 不在阶段 3 完成整个测试体系命名/夹具迁移;阶段 4 负责全面替换旧 Sequential 测试体系,但阶段 3 不允许留下已知失败测试。
- 不修改 `/api/chat` 请求/响应结构或 Executor/Verifier/Composer 输出协议。
- 不新增 Run status,不做历史 run 的 orchestration trace 回填或兼容读取分支。
- 不运行 Maven live E2E、`logs/` 或数据库查询;统一保留到阶段 5。
## Context Constraints
- 阶段 0–2 archives 和主 specs 是实现基线;Graph 的 Node、条件边、计数所有权和安全 Fallback 边界不得在本阶段复制或放宽。
- `diagnosis_run.status` 只表达执行生命周期;Composer、LOW_CONFID、Verifier REJECT 或固定安全 Fallback 只要生成安全答案均为 SUCCESS。
- 编排摘要必须由有界 `orchestration_events` 构造,不从日志反推,不保存 Prompt、reasoning、raw tool output 或 Graph State snapshot。
- self-evaluation 与 orchestration trace 分离;AgentStep/ToolInvocation 继续作为详细 Trace 数据源。
- 新 Graph Verifier 只接收 `verified_executor_output`、`verified_evidence`、Gatekeeper audit/ceiling、query 和 retry context。
- 运行失败时不得伪造未发生的 events;可取得的部分 state 才允许 best-effort 持久化。
## Acceptance
- 复杂 Chat 生产代码不再创建或调用 SequentialAgent,且只调用一次 CompiledGraph。
- RunnableConfig 的 threadId 等于 runId,metadata 同时包含当前 sessionId/runId。
- Graph final answer 非空时映射为现有 ChatResult,Run 为 SUCCESS;无法产生安全响应或未处理失败时 Run 为 FAILED。
- 每个新 StateGraph Chat run 都持久化非空、无敏感材料的 orchestration trace;Trace API 只在 `run.orchestrationTrace` 返回解析对象。
- AgentStep、ToolInvocation、self-evaluation、answer、duration、token/step/tool count 和 Eval 仍绑定当前 runId。
- Verifier Prompt 和实际 verified-only payload 一致,Graph Verifier 不注册旧 input Hook。
- Executor/Gatekeeper 等前置失败的 Graph 路径不产生 Verifier AgentStep。
- `/api/chat` 和证据协议保持不变;数据库 schema 只新增一个 nullable JSON 列。
- 阶段 3 focused tests、Trace/Controller/Eval 回归、test compilation 和 OpenSpec strict validation 通过;不运行 live E2E。
## Risks
- ChatService 当前同时承载 Agent factory、Run 生命周期和旧解析逻辑;cutover 必须做完整内部重构,避免留下双编排路径或死代码。
- self-evaluation 旧字段曾依赖 ThreadLocal/full tool summary;Graph 映射必须明确兼容字段和安全边界,不能重新泄漏未验真材料。
- Graph invoke 异常可能没有可用 final state;失败处理只能持久化真实可取得的 partial snapshot,不能编造路径。
- JSON entity/DDL/Trace DTO 必须保持同一字段语义,否则 JPA validate、数据库迁移或 API 解析会漂移。
@@ -0,0 +1,60 @@
## MODIFIED Requirements
### Requirement: Composer SHALL generate final user-facing Chat answers
The system SHALL invoke the Composer Graph Node after Verifier routing to generate the final user-facing Chat answer from Verifier-allowed material.
#### Scenario: Composer receives only filtered material
- **WHEN** the Composer Graph Node invokes its configured Agent
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
- **AND** the Composer input SHALL NOT contain raw tool output
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
#### Scenario: Composer outputs strict JSON
- **WHEN** Composer completes
- **THEN** it SHALL output exactly one JSON object
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
#### Scenario: Composer does not introduce new facts
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
### Requirement: Composer input SHALL honor Verifier claim checks
The Composer Graph Node SHALL construct Composer input by filtering verified Executor structured output through Verifier `claim_checks`.
#### Scenario: Passing claims become allowed claims
- **WHEN** a claim check verification is `direct_observation`
- **THEN** the Composer input builder SHALL include the matching verified Executor claim in `allowed_claims`
#### Scenario: Reasonable inferences remain bounded
- **WHEN** a claim check verification is `reasonable_inference`
- **THEN** the Composer input builder MAY include the matching verified Executor claim in `allowed_claims`
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
#### Scenario: Overstated claims are not confirmed findings
- **WHEN** a claim check verification is `overstated`
- **THEN** the Composer input builder SHALL NOT include the matching claim as a confirmed item in `allowed_claims`
- **AND** it MAY include it as `allowed_hypotheses` or represent it in `missing_info`
#### Scenario: Unsupported or external claims are withheld
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
- **THEN** the Composer input builder SHALL NOT include the matching claim in `allowed_claims`
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
### Requirement: Composer failures SHALL degrade safely
The system SHALL tolerate malformed or failed Composer execution without leaking raw JSON or unverified Executor material.
#### Scenario: malformed Composer output falls back safely
- **WHEN** Composer exhausts its fixed-input technical retry or returns a non-retryable failure
- **THEN** the deterministic Fallback Node SHALL produce a final answer using only filtered material
- **AND** the final answer SHALL NOT expose raw Composer output
- **AND** the final answer SHALL NOT expose raw Executor output
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
#### Scenario: Composer audit is persisted
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.composer_output` SHALL record parsed Composer audit when available
- **AND** handled Composer fallback SHALL be observable through orchestration trace and available status/reason fields
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
@@ -0,0 +1,107 @@
## ADDED Requirements
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
#### Scenario: Complex Chat starts
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
#### Scenario: Graph dependencies are assembled
- **WHEN** the complex Chat Graph is constructed
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
### Requirement: Graph invocation SHALL preserve current Run ownership
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
#### Scenario: Agent Node runs
- **WHEN** any Agent Node is invoked for a complex Chat run
- **THEN** its RunnableConfig threadId SHALL equal the current runId
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
#### Scenario: Gatekeeper validates Executor output
- **WHEN** the Gatekeeper Node runs
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
- **AND** data from another run in the same session SHALL NOT be considered
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
#### Scenario: Initial state is projected
- **WHEN** a complex Chat run enters Planner for the first time
- **THEN** `diagnosis_context` SHALL contain the current query/original query
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
#### Scenario: Counters are initialized
- **WHEN** the Graph starts
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
- **AND** Planner mode SHALL be NORMAL
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
#### Scenario: Composer completes
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
- **AND** EvaluationService SHALL evaluate the current runId
#### Scenario: Handled Fallback completes
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
- **THEN** DiagnosisRun SHALL be SUCCESS
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
- **AND** no DEGRADED Run status SHALL be introduced
#### Scenario: Graph cannot produce a safe response
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
- **AND** the system SHALL NOT report a successful Graph result
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
#### Scenario: Verifier completed
- **WHEN** final Graph State contains a completed Verifier result
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
- **AND** the compatibility `verdict` field SHALL equal effective verdict
#### Scenario: Pre-verification Fallback completed
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
- **AND** it SHALL NOT fabricate a model or effective verdict
#### Scenario: Evaluation payload is inspected
- **WHEN** verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
### Requirement: Stage 3 verification SHALL not run the final live E2E
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
#### Scenario: Stage 3 is accepted
- **WHEN** stage 3 verification completes
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5
@@ -0,0 +1,7 @@
## REMOVED Requirements
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
**Reason**: Stage 2 production isolation has completed its migration purpose; stage 3 intentionally makes the real Diagnosis Graph the only complex Chat production orchestrator.
**Migration**: Replace `ChatService.executeChatComplex` SequentialAgent orchestration with real Graph assembly/invocation. Keep `/api/chat` compatible and use a full Git revert of stage 3 for runtime rollback rather than retaining dual paths.
@@ -0,0 +1,263 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical claims in `verified_executor_output.claims` have direct observation or reasonable inference support in `verified_evidence`
- **AND** at least one critical claim has direct observation
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
- **AND** verdict ceiling is PASS
- **THEN** the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical claim contradicts verified evidence
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only inference support
- **WHEN** no critical claim contradicts verified evidence
- **AND** all critical claims are only `reasonable_inference`
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical claim in `verified_executor_output.claims` contradicts verified evidence
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
- **THEN** the Verifier SHALL output model verdict="REJECT"
#### Scenario: Verified structured claims are the only verification target
- **WHEN** `verified_executor_output.claims` is present
- **THEN** Verifier SHALL verify each structured claim against matching `verified_evidence` through `claim_checks`
- **AND** each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
- **AND** a claim without matching verified evidence SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT receive or add confirmed facts from raw Executor text
#### Scenario: Executor output is invalid
- **WHEN** Executor does not return a legal structured contract
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute or fabricate a diagnostic verdict
### Requirement: facts_checked SHALL use a fixed classification set
The system SHALL continue to expose compatibility `facts_checked` using its fixed verification classification set.
#### Scenario: claim checks are mapped to legacy facts
- **WHEN** Verifier output contains `claim_checks`
- **THEN** the shared Verifier protocol parser SHALL derive compatibility `facts_checked` when the model did not provide them
- **AND** `direct_observation` SHALL map to `direct_evidence`
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
- **AND** `contradicted` SHALL map to `contradicted`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.
#### Scenario: PASS routes to Composer
- **WHEN** Verifier completes with effective verdict="PASS"
- **THEN** the Graph SHALL invoke Composer with filtered Verifier-allowed material
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
#### Scenario: LOW_CONFID does not qualify for evidence retry
- **WHEN** Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
- **THEN** the Graph SHALL route to Composer without another Planner cycle
- **AND** the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps
#### Scenario: LOW_CONFID qualifies for evidence retry
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
- **THEN** the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
- **AND** it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier completes with effective verdict="REJECT"
- **THEN** the Graph SHALL NOT start an evidence supplementation round
- **AND** it SHALL route to Composer-safe output
#### Scenario: REJECT produces bounded output
- **WHEN** effective verdict is REJECT
- **THEN** the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
- **AND** it SHALL NOT pass through raw Executor answer
- **AND** it SHALL NOT include an unsupported root-cause conclusion
#### Scenario: Verifier execution fails
- **WHEN** Verifier exhausts technical retry or returns a non-retryable failure
- **THEN** the Graph SHALL route to pre-verification Fallback
- **AND** no execution status string SHALL be used as model or effective verdict
### Requirement: Verifier SHALL be observable
The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.
#### Scenario: claim checks written to self_evaluation
- **WHEN** a completed Verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** it SHALL include `verifier_status`, `model_verdict`, `effective_verdict`, `verdict`, `groundedness_score`, `rationale`, verified output/evidence, and Gatekeeper audit
#### Scenario: composer output written to self_evaluation
- **WHEN** final answer composition completes
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include compact `composer_output` when available
- **AND** handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
- **AND** existing claim/fact and Gatekeeper fields SHALL be preserved
#### Scenario: verdict written to self_evaluation
- **WHEN** Verifier completes
- **THEN** Graph result mapping SHALL write effective verdict under `diagnosis_run.self_evaluation.verifier_evaluation.verdict`
- **AND** existing `rule_evaluation` and `aiops_rule_evaluation` channels SHALL be preserved
#### Scenario: pre-verification fallback is persisted
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
- **AND** it SHALL NOT fabricate `model_verdict` or `effective_verdict`
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** Graph result mapping persists available Gatekeeper state
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** a complex Chat Graph result is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_run.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat Prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
- **AND** full Prompt text SHALL NOT be persisted
#### Scenario: prompt audit available on fallback paths
- **WHEN** Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
#### Scenario: evaluation payload is inspected
- **WHEN** Graph verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete `tool_trace_summary`
- **AND** compatibility `executor_structured_output` SHALL contain at most the verified projection
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier Graph Node starts
- **THEN** the payload SHALL provide `diagnosis_context`, `verified_executor_output`, `verified_evidence`, `gatekeeper_audit`, and `verdict_ceiling`
- **AND** permitted structured `retry_context` SHALL be provided only after evidence retry preparation
#### Scenario: Verifier remains isolated from intermediate and raw material
- **WHEN** the Verifier input is serialized
- **THEN** it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State
#### Scenario: only passed bindings are available
- **WHEN** Gatekeeper returns mixed passed and failed checked bindings
- **THEN** `verified_executor_output` and `verified_evidence` SHALL contain only claims/material matching passed bindings
- **AND** the Verifier SHALL NOT receive failed or unreferenced tool material
#### Scenario: verified evidence preserves precise references
- **WHEN** the system prepares Verifier input
- **THEN** each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
- **AND** the item SHALL be traceable to current-run Gatekeeper validation
#### Scenario: gatekeeper audit and ceiling are available
- **WHEN** the system prepares Verifier input
- **THEN** the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
- **AND** a LOW_CONFID ceiling SHALL prevent effective PASS
#### Scenario: technical retry occurs
- **WHEN** the first Verifier attempt returns invalid output or a retryable invocation failure
- **THEN** the second attempt SHALL receive byte-identical serialized input
- **AND** Executor, Gatekeeper, and tools SHALL NOT rerun
### Requirement: Verifier facts SHALL be auditable
Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.
#### Scenario: claim checks contain evidence refs
- **WHEN** the Verifier emits `claim_checks`
- **THEN** each check SHALL include an `evidence_refs` array
- **AND** any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
- **AND** it SHALL NOT reference a failed or unverified binding
#### Scenario: verifier evaluation persists traceability snapshot
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the bounded `verified_evidence` snapshot used by the Verifier
- **AND** it SHALL NOT persist a complete tool trace summary as Verifier input
### Requirement: Structured Executor output SHALL degrade safely
The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.
#### Scenario: Malformed Executor JSON is classified
- **WHEN** Executor returns malformed JSON or text outside the expected contract
- **THEN** Executor Node SHALL set INVALID_OUTPUT
- **AND** the Graph SHALL route directly to deterministic pre-verification Fallback
- **AND** Gatekeeper, Verifier, and model Composer SHALL NOT execute
#### Scenario: Structured parse failure remains observable
- **WHEN** Executor output parsing fails
- **THEN** orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
- **AND** the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: schema rule rejects missing evidence bindings
- **WHEN** a confirmed claim has no `evidence_bindings`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references an invocation id absent from current-run `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing current-run invocation id
- **AND** binding `tool_name` does not match persisted invocation `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits legal `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
#### Scenario: gatekeeper reject bypasses Verifier
- **WHEN** normalized Gatekeeper status is REJECT
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute
#### Scenario: gatekeeper low confidence is bounded
- **WHEN** normalized Gatekeeper status is LOW_CONFID with at least one passed binding
- **THEN** verified input SHALL contain only passed bindings
- **AND** effective verdict SHALL NOT exceed LOW_CONFID
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.
#### Scenario: Verified excerpt supports direct observation
- **WHEN** verdict ceiling is PASS
- **AND** a claim's verified evidence matched text directly contains the claim's concrete facts
- **THEN** Verifier MAY classify that claim as `direct_observation`
#### Scenario: Verified evidence is complete Verifier context
- **WHEN** verified claims and evidence are available
- **THEN** Verifier SHALL use them as its evidence context
- **AND** it SHALL NOT require or request a complete tool trace summary
- **AND** it SHALL NOT read raw Executor or unreferenced tool material
### Requirement: Gatekeeper severity SHALL constrain effective verdict
Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.
#### Scenario: Reject severity bypasses Verifier
- **WHEN** `gatekeeper_result.severity=reject`
- **THEN** the Graph SHALL route to pre-verification Fallback without invoking Verifier
- **AND** it SHALL NOT fabricate an effective diagnostic verdict
#### Scenario: Low confidence severity prevents PASS
- **WHEN** `gatekeeper_result.severity=low_confid`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** deterministic effective-verdict calculation SHALL downgrade the result
- **AND** effective verdict SHALL be `LOW_CONFID`
#### Scenario: Gatekeeper audit includes severity
- **WHEN** Graph verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include available `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
@@ -0,0 +1,56 @@
## ADDED Requirements
### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`.
#### Scenario: Graph reaches Composer
- **WHEN** a complex Chat Graph terminates through Composer with a safe answer
- **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
- **AND** the summary SHALL be derived from the current Run's actual orchestration events
#### Scenario: Graph reaches handled Fallback
- **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer
- **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true`
- **AND** the Run status SHALL be SUCCESS
#### Scenario: Unhandled execution fails after events exist
- **WHEN** an unhandled failure occurs after one or more real Graph events are available
- **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run
- **AND** it SHALL NOT add a node or transition that did not occur
#### Scenario: Orchestration trace content is inspected
- **WHEN** orchestration trace JSON is serialized
- **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
- **AND** it SHALL NOT contain data owned by another run
### Requirement: Trace API SHALL expose orchestration trace only on the run object
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`.
#### Scenario: Exact StateGraph run trace is queried
- **WHEN** a caller queries a successful new StateGraph Chat run
- **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object
- **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field
- **AND** no raw orchestration trace field SHALL be added
#### Scenario: Historical or non-StateGraph run is queried
- **WHEN** the selected DiagnosisRun has null orchestration trace
- **THEN** `run.orchestrationTrace` MAY be null
- **AND** the service SHALL NOT synthesize historical events or read another run's trace
### Requirement: Orchestration trace migration SHALL be additive and nullable
The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change.
#### Scenario: Migration is applied
- **WHEN** Flyway applies the stage 3 migration
- **THEN** existing DiagnosisRun rows SHALL remain valid without backfill
- **AND** no other table or column SHALL be changed by the stage 3 schema migration
@@ -0,0 +1,43 @@
## 1. Run Orchestration Trace Contract
- [x] 1.1 Add V012 migration containing only nullable `diagnosis_run.orchestration_trace JSON` and document that historical rows are not backfilled.
- [x] 1.2 Add DiagnosisRun JSON field and `DiagnosisTraceResponse.RunTrace.orchestrationTrace` parsed-map field without adding top-level, session, summary, run-list, or raw duplicates.
- [x] 1.3 Update DiagnosisTraceService exact/latest Run mapping and tests for current-run parsed trace, historical null, invalid JSON fail-closed, and no cross-run/session projection.
- [x] 1.4 Add entity/repository or migration source checks proving stage 3 changes no schema object except the one nullable column.
## 2. Graph Runtime And Result Mapping
- [x] 2.1 Implement a dedicated complex Chat Graph runtime that assembles existing real actions, compiles the Graph once per request/runtime instance, and consumes `CompiledGraph.stream` while retaining only the last real state.
- [x] 2.2 Define minimal initial state with query-only diagnosis context, NORMAL mode, independent zero counters, and empty bounded events.
- [x] 2.3 Build RunnableConfig with `threadId=runId` and current `sessionId`/`runId` metadata; prove config identity at every real Agent invoker and Gatekeeper boundary.
- [x] 2.4 Implement explicit final/partial Graph result mapping for answer, statuses/verdicts, verified output/evidence, Gatekeeper audit, Composer audit, failure reason, and orchestration trace.
- [x] 2.5 Add runtime/mapper tests for Composer success, handled pre/post-verification Fallback, blank answer, no state, thrown execution with real partial events, and no fabricated event.
## 3. Agent Assembly And Verified-only Prompt
- [x] 3.1 Move or encapsulate the existing complex Planner/Executor/Verifier/Composer builders so Graph runtime reuses current Prompt, knowledge map, history, method tool, ToolCallback, skill, and AgentLoggingHook rules without a duplicate factory.
- [x] 3.2 Remove VerifierInputHook from the Graph Verifier while retaining AgentLoggingHook, and prove explicit Gatekeeper executes once per Executor round.
- [x] 3.3 Update `chat-verifier-prompt.md` to the verified-only input contract and remove raw Executor, full tool trace, Hook parse, and Gatekeeper re-execution language.
- [x] 3.4 Bump Prompt audit catalog/Verifier Prompt version and add static/behavior tests proving the runtime payload and Prompt declare the same allowed fields.
## 4. ChatService Production Cutover
- [x] 4.1 Replace `executeChatComplex` SequentialAgent/outer retry loop with one Graph runtime call while preserving session/run creation and the public ChatResult signature.
- [x] 4.2 Remove obsolete ChatService private Sequential workflow, score/feature-flag retry, ThreadLocal verifier parsing, private Composer routing, and dead imports without changing simple Chat behavior.
- [x] 4.3 Persist Graph-derived verifier evaluation with compatibility verdict/claim/fact/gatekeeper/composer/prompt-audit fields, verified-only evidence, and no fabricated verdict before Verifier completion.
- [x] 4.4 Persist SUCCESS only for a non-empty safe Graph answer with non-empty orchestration trace, then backfill current-run metrics and call Eval; persist FAILED for unhandled/no-answer/invariant failures with any real partial trace available.
- [x] 4.5 Ensure `SessionContextHolder` and retrieved-doc cleanup still run in finally, no complex Chat code reads `VerifierContextHolder`, and no production path creates SequentialAgent.
## 5. Stage 3 Focused Tests And Alignment
- [x] 5.1 Add production cutover integration tests for `/api/chat`-compatible ChatResult, `agent_flow=CHAT`, current Run ownership, Graph thread/metadata, answer/metrics/Eval persistence, and multi-run isolation.
- [x] 5.2 Add handled failure tests proving Executor invalid/failure and Gatekeeper REJECT do not create Verifier AgentStep, safe Fallback is SUCCESS/degraded, and unhandled/no-answer paths are FAILED.
- [x] 5.3 Update or replace stage-3-conflicting Sequential tests so no known test remains red; preserve Controller, Gatekeeper, Trace, Repository, Eval, no-evidence, REJECT, and Composer safety regressions.
- [x] 5.4 Confirm the first completed Trace slice and final production assembly against design/specs, including single-track cutover, no duplicate Agent factory, and no core TODO/placeholder.
## 6. Stage 3 Verification And Handoff
- [x] 6.1 Run Graph runtime/result mapper, ChatService cutover, DiagnosisTraceService, Controller, Repository, Gatekeeper, Composer, and Eval focused tests.
- [x] 6.2 Run Maven test compilation and any existing tests whose public contracts were touched; classify failures as OpenSpec gap, code deviation, or out-of-scope environment issue.
- [x] 6.3 Run OpenSpec strict validation, all main specs strict validation, `git diff --check`, source/reference isolation checks, and a schema whitelist check.
- [x] 6.4 Record that Maven live E2E, `logs/`, and `scripts/query_mysql.py` database verification were intentionally not run in stage 3 and remain reserved for stage 5.
+14 -14
View File
@@ -4,10 +4,10 @@
TBD - created by archiving change executor-composer-final-answer. Update Purpose after archive.
## Requirements
### Requirement: Composer SHALL generate final user-facing Chat answers
The system SHALL invoke a Composer expression layer after Verifier to generate the final user-facing Chat answer from Verifier-allowed material.
The system SHALL invoke the Composer Graph Node after Verifier routing to generate the final user-facing Chat answer from Verifier-allowed material.
#### Scenario: Composer receives only filtered material
- **WHEN** ChatService invokes Composer
- **WHEN** the Composer Graph Node invokes its configured Agent
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
- **AND** the Composer input SHALL NOT contain raw tool output
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
@@ -25,25 +25,25 @@ The system SHALL invoke a Composer expression layer after Verifier to generate t
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
### Requirement: Composer input SHALL honor Verifier claim checks
ChatService SHALL construct Composer input by filtering Executor structured output through Verifier `claim_checks`.
The Composer Graph Node SHALL construct Composer input by filtering verified Executor structured output through Verifier `claim_checks`.
#### Scenario: Passing claims become allowed claims
- **WHEN** a claim check verification is `direct_observation`
- **THEN** ChatService SHALL include the matching Executor claim in `allowed_claims`
- **THEN** the Composer input builder SHALL include the matching verified Executor claim in `allowed_claims`
#### Scenario: Reasonable inferences remain bounded
- **WHEN** a claim check verification is `reasonable_inference`
- **THEN** ChatService MAY include the matching Executor claim in `allowed_claims`
- **THEN** the Composer input builder MAY include the matching verified Executor claim in `allowed_claims`
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
#### Scenario: Overstated claims are not confirmed findings
- **WHEN** a claim check verification is `overstated`
- **THEN** ChatService SHALL NOT include the matching Executor claim as a confirmed item in `allowed_claims`
- **AND** ChatService MAY include it as `allowed_hypotheses` or represent it in `missing_info`
- **THEN** the Composer input builder SHALL NOT include the matching claim as a confirmed item in `allowed_claims`
- **AND** it MAY include it as `allowed_hypotheses` or represent it in `missing_info`
#### Scenario: Unsupported or external claims are withheld
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
- **THEN** ChatService SHALL NOT include the matching Executor claim in `allowed_claims`
- **THEN** the Composer input builder SHALL NOT include the matching claim in `allowed_claims`
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
### Requirement: Composer SHALL respect verdict-specific wording
@@ -67,17 +67,17 @@ Composer SHALL phrase final answers according to the effective Verifier verdict.
- **AND** the final answer SHALL NOT include a root-cause conclusion
### Requirement: Composer failures SHALL degrade safely
The system SHALL tolerate malformed Composer output without leaking raw JSON or unverified Executor material.
The system SHALL tolerate malformed or failed Composer execution without leaking raw JSON or unverified Executor material.
#### Scenario: malformed Composer output falls back safely
- **WHEN** Composer returns malformed JSON or omits required fields
- **THEN** ChatService SHALL produce a final answer using a fixed safe fallback template based only on filtered material
- **WHEN** Composer exhausts its fixed-input technical retry or returns a non-retryable failure
- **THEN** the deterministic Fallback Node SHALL produce a final answer using only filtered material
- **AND** the final answer SHALL NOT expose raw Composer output
- **AND** the final answer SHALL NOT expose raw Executor output
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
#### Scenario: Composer audit is persisted
- **WHEN** ChatService persists verifier evaluation
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.composer_output` SHALL record whether Composer output was valid or fallback was used
- **AND** the audit SHALL include the parsed Composer fields when valid
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.composer_output` SHALL record parsed Composer audit when available
- **AND** handled Composer fallback SHALL be observable through orchestration trace and available status/reason fields
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
@@ -0,0 +1,110 @@
# chat-diagnosis-stategraph-chatservice-cutover Specification
## Purpose
TBD - created by archiving change chat-diagnosis-stategraph-chatservice-cutover. Update Purpose after archive.
## Requirements
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
#### Scenario: Complex Chat starts
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
#### Scenario: Graph dependencies are assembled
- **WHEN** the complex Chat Graph is constructed
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
### Requirement: Graph invocation SHALL preserve current Run ownership
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
#### Scenario: Agent Node runs
- **WHEN** any Agent Node is invoked for a complex Chat run
- **THEN** its RunnableConfig threadId SHALL equal the current runId
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
#### Scenario: Gatekeeper validates Executor output
- **WHEN** the Gatekeeper Node runs
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
- **AND** data from another run in the same session SHALL NOT be considered
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
#### Scenario: Initial state is projected
- **WHEN** a complex Chat run enters Planner for the first time
- **THEN** `diagnosis_context` SHALL contain the current query/original query
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
#### Scenario: Counters are initialized
- **WHEN** the Graph starts
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
- **AND** Planner mode SHALL be NORMAL
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
#### Scenario: Composer completes
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
- **AND** EvaluationService SHALL evaluate the current runId
#### Scenario: Handled Fallback completes
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
- **THEN** DiagnosisRun SHALL be SUCCESS
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
- **AND** no DEGRADED Run status SHALL be introduced
#### Scenario: Graph cannot produce a safe response
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
- **AND** the system SHALL NOT report a successful Graph result
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
#### Scenario: Verifier completed
- **WHEN** final Graph State contains a completed Verifier result
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
- **AND** the compatibility `verdict` field SHALL equal effective verdict
#### Scenario: Pre-verification Fallback completed
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
- **AND** it SHALL NOT fabricate a model or effective verdict
#### Scenario: Evaluation payload is inspected
- **WHEN** verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
### Requirement: Stage 3 verification SHALL not run the final live E2E
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
#### Scenario: Stage 3 is accepted
- **WHEN** stage 3 verification completes
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5
@@ -123,13 +123,3 @@ Composer SHALL receive only effective verdict and Verifier-allowed claims, missi
- **WHEN** Composer fails after its one technical retry and Verifier-allowed material exists
- **THEN** deterministic Fallback SHALL express only the allowed claims, missing information, and recommendations
- **AND** it SHALL NOT read raw Executor or tool output
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
The real Node action set and CompiledGraph SHALL be constructible and testable, but ChatService, database, Trace API, shared Agent prompts, and the current production routing SHALL remain unchanged until the stage 3 change.
#### Scenario: Stage 2 production isolation is inspected
- **WHEN** this change is accepted
- **THEN** no production ChatService code path SHALL invoke the real Diagnosis Graph
- **AND** no database migration, Trace API field, or shared Prompt contract SHALL be changed
+146 -171
View File
@@ -4,41 +4,41 @@
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
## Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads structured Executor claims and the tool call history, then produces a structured verdict based on claim derivability.
The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical claims in `executor_structured_output.claims` have direct observation or reasonable inference support in tool call results
- **WHEN** all critical claims in `verified_executor_output.claims` have direct observation or reasonable inference support in `verified_evidence`
- **AND** at least one critical claim has direct observation
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
- **AND** `gatekeeper_result.status` is not `fail`
- **THEN** the Verifier MAY output verdict="PASS" with groundedness_score ≥ 0.5
- **AND** verdict ceiling is PASS
- **THEN** the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical claim contradicts the tool results
- **WHEN** no critical claim contradicts verified evidence
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only inference support
- **WHEN** no critical claim contradicts the tool results
- **WHEN** no critical claim contradicts verified evidence
- **AND** all critical claims are only `reasonable_inference`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical claim in `executor_structured_output.claims` contradicts tool call results
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
- **THEN** the Verifier SHALL output verdict="REJECT"
- **WHEN** any critical claim in `verified_executor_output.claims` contradicts verified evidence
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
- **THEN** the Verifier SHALL output model verdict="REJECT"
#### Scenario: Structured Executor claims are verified first
- **WHEN** `executor_structured_output.claims` is present and valid
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` through `claim_checks`
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: Verified structured claims are the only verification target
- **WHEN** `verified_executor_output.claims` is present
- **THEN** Verifier SHALL verify each structured claim against matching `verified_evidence` through `claim_checks`
- **AND** each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
- **AND** a claim without matching verified evidence SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT receive or add confirmed facts from raw Executor text
#### Scenario: Malformed structured output cannot pass through natural language fallback
- **WHEN** Executor does not return parseable structured output
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
#### Scenario: Executor output is invalid
- **WHEN** Executor does not return a legal structured contract
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute or fabricate a diagnostic verdict
### Requirement: Verifier SHALL output structured JSON
The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_checks array, compatibility facts_checked array, and rationale.
@@ -64,11 +64,11 @@ The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
### Requirement: facts_checked SHALL use a fixed classification set
The system SHALL continue to expose legacy `facts_checked` using its fixed verification classification set.
The system SHALL continue to expose compatibility `facts_checked` using its fixed verification classification set.
#### Scenario: claim checks are mapped to legacy facts
- **WHEN** Verifier output contains `claim_checks`
- **THEN** ChatService SHALL derive compatibility `facts_checked`
- **THEN** the shared Verifier protocol parser SHALL derive compatibility `facts_checked` when the model did not provide them
- **AND** `direct_observation` SHALL map to `direct_evidence`
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
@@ -89,73 +89,86 @@ The groundedness score SHALL be computed from critical fact classifications inst
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed.
The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.
#### Scenario: PASS -> Composer output
- **WHEN** Verifier outputs verdict="PASS"
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
#### Scenario: PASS routes to Composer
- **WHEN** Verifier completes with effective verdict="PASS"
- **THEN** the Graph SHALL invoke Composer with filtered Verifier-allowed material
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: LOW_CONFID does not qualify for evidence retry
- **WHEN** Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
- **THEN** the Graph SHALL route to Composer without another Planner cycle
- **AND** the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps
#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
#### Scenario: LOW_CONFID qualifies for evidence retry
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
- **THEN** the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
- **AND** it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL NOT start a retry round for evidence supplementation
- **AND** it SHALL produce a degraded output directly through Composer-safe rendering
- **WHEN** Verifier completes with effective verdict="REJECT"
- **THEN** the Graph SHALL NOT start an evidence supplementation round
- **AND** it SHALL route to Composer-safe output
#### Scenario: REJECT -> degraded output
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
- **AND** it SHALL NOT pass through the raw Executor answer
- **AND** it SHALL NOT include a root-cause conclusion
#### Scenario: REJECT produces bounded output
- **WHEN** effective verdict is REJECT
- **THEN** the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
- **AND** it SHALL NOT pass through raw Executor answer
- **AND** it SHALL NOT include an unsupported root-cause conclusion
#### Scenario: Verifier execution fails
- **WHEN** Verifier exhausts technical retry or returns a non-retryable failure
- **THEN** the Graph SHALL route to pre-verification Fallback
- **AND** no execution status string SHALL be used as model or effective verdict
### Requirement: Verifier SHALL be observable
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.
#### Scenario: claim checks written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **WHEN** a completed Verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
- **AND** it SHALL include `verifier_status`, `model_verdict`, `effective_verdict`, `verdict`, `groundedness_score`, `rationale`, verified output/evidence, and Gatekeeper audit
#### Scenario: composer output written to self_evaluation
- **WHEN** final answer composition completes
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output`
- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used
- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include compact `composer_output` when available
- **AND** handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
- **AND** existing claim/fact and Gatekeeper fields SHALL be preserved
#### Scenario: verdict written to self_evaluation
- **WHEN** the Verifier produces a verdict
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
- **WHEN** Verifier completes
- **THEN** Graph result mapping SHALL write effective verdict under `diagnosis_run.self_evaluation.verifier_evaluation.verdict`
- **AND** existing `rule_evaluation` and `aiops_rule_evaluation` channels SHALL be preserved
#### Scenario: pre-verification fallback is persisted
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
- **AND** it SHALL NOT fabricate `model_verdict` or `effective_verdict`
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
- **WHEN** Graph result mapping persists available Gatekeeper state
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** the Chat verifier evaluation is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
- **WHEN** a complex Chat Graph result is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_run.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat Prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
- **AND** full Prompt text SHALL NOT be persisted
#### Scenario: prompt audit available on fallback paths
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
- **WHEN** Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
#### Scenario: evaluation payload is inspected
- **WHEN** Graph verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete `tool_trace_summary`
- **AND** compatibility `executor_structured_output` SHALL contain at most the verified projection
### Requirement: self_evaluation SHALL be a container object
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
@@ -175,88 +188,51 @@ The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation ch
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier starts
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
- **AND** `retry_context` SHALL be provided on the second round only
- **AND** when Executor returns a valid evidence-attribution contract, the system SHALL provide `executor_structured_output`
- **AND** when Executor output parsing fails, the system SHALL provide an `executor_output_parse_status` that indicates the failure
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
- **WHEN** the Verifier Graph Node starts
- **THEN** the payload SHALL provide `diagnosis_context`, `verified_executor_output`, `verified_evidence`, `gatekeeper_audit`, and `verdict_ceiling`
- **AND** permitted structured `retry_context` SHALL be provided only after evidence retry preparation
#### Scenario: Verifier remains isolated from intermediate reasoning
- **WHEN** `executor_structured_output` is added to the verifier input
- **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, gatekeeper result, and retry context
#### Scenario: Verifier remains isolated from intermediate and raw material
- **WHEN** the Verifier input is serialized
- **THEN** it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State
#### Scenario: tool trace summary derived from tool facts
- **WHEN** the system prepares verifier inputs
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
#### Scenario: only passed bindings are available
- **WHEN** Gatekeeper returns mixed passed and failed checked bindings
- **THEN** `verified_executor_output` and `verified_evidence` SHALL contain only claims/material matching passed bindings
- **AND** the Verifier SHALL NOT receive failed or unreferenced tool material
#### Scenario: tool trace summary preserves invocation references
- **WHEN** the system prepares verifier inputs
- **THEN** each summary item SHALL include a stable `trace_ref`
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
#### Scenario: verified evidence preserves precise references
- **WHEN** the system prepares Verifier input
- **THEN** each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
- **AND** the item SHALL be traceable to current-run Gatekeeper validation
#### Scenario: only evidence-bearing tools included
- **WHEN** the system generates `tool_trace_summary`
- **THEN** it SHALL include only evidence-bearing tool invocations
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
#### Scenario: gatekeeper audit and ceiling are available
- **WHEN** the system prepares Verifier input
- **THEN** the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
- **AND** a LOW_CONFID ceiling SHALL prevent effective PASS
#### Scenario: failed evidence calls preserved as evidence gaps
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
- **THEN** the summary SHALL still include that invocation
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
#### Scenario: repeated tool calls may be compacted
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
- **THEN** the system MAY compact them into a merged summary entry
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
#### Scenario: raw outputs not passed through in full
- **WHEN** a tool invocation returns large raw content
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
#### Scenario: MessagesModelHook used only for noise reduction
- **WHEN** a MessagesModelHook is used for the Verifier
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
#### Scenario: gatekeeper result available to Verifier
- **WHEN** the system prepares verifier inputs from Executor output
- **THEN** the payload SHALL include `gatekeeper_result`
- **AND** `gatekeeper_result.status` SHALL be one of `pass`, `warn`, or `fail`
- **AND** `gatekeeper_result` SHALL include `failed_rules`, `warnings`, and `errors`
#### Scenario: structured claims are the primary verification target
- **WHEN** `executor_output_parse_status.status` is `valid`
- **AND** `executor_structured_output.claims` is available
- **THEN** Verifier SHALL verify each claim through `claim_checks`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: malformed structured output cannot pass through natural language fallback
- **WHEN** `executor_output_parse_status.status` is `missing` or `malformed`
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
#### Scenario: technical retry occurs
- **WHEN** the first Verifier attempt returns invalid output or a retryable invocation failure
- **THEN** the second attempt SHALL receive byte-identical serialized input
- **AND** Executor, Gatekeeper, and tools SHALL NOT rerun
### Requirement: Verifier facts SHALL be auditable
Verifier facts SHALL be linkable to the evidence summaries used during verification.
Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.
#### Scenario: facts_checked contains evidence refs
- **WHEN** the Verifier emits `facts_checked`
- **THEN** each fact SHALL include `evidence_refs`
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
#### Scenario: claim checks contain evidence refs
- **WHEN** the Verifier emits `claim_checks`
- **THEN** each check SHALL include an `evidence_refs` array
- **AND** any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
- **AND** it SHALL NOT reference a failed or unverified binding
#### Scenario: verifier evaluation persists traceability snapshot
- **WHEN** the ChatService persists `verifier_evaluation`
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
- **AND** it SHALL include the bounded `verified_evidence` snapshot used by the Verifier
- **AND** it SHALL NOT persist a complete tool trace summary as Verifier input
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
@@ -346,21 +322,21 @@ The system SHALL preserve a readable Chinese answer for normal Chat users even w
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
### Requirement: Structured Executor output SHALL degrade safely
The system SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow.
The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.
#### Scenario: Malformed Executor JSON is captured
#### Scenario: Malformed Executor JSON is classified
- **WHEN** Executor returns malformed JSON or text outside the expected contract
- **THEN** Chat runtime SHALL preserve the raw `executor_final_answer`
- **AND** it SHALL mark `executor_output_parse_status` as failed
- **AND** Verifier SHALL use natural-language fallback behavior
- **THEN** Executor Node SHALL set INVALID_OUTPUT
- **AND** the Graph SHALL route directly to deterministic pre-verification Fallback
- **AND** Gatekeeper, Verifier, and model Composer SHALL NOT execute
#### Scenario: Structured parse failure remains observable
- **WHEN** Executor output parsing fails
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract
- **THEN** orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
- **AND** the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
@@ -373,32 +349,31 @@ The system SHALL run deterministic Gatekeeper checks after Executor output parsi
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references a `source_invocation_ids` value that is not present in current-session `tool_invocation` rows
- **WHEN** a claim evidence binding references an invocation id absent from current-run `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing invocation id
- **AND** the binding `tool_name` does not match the invocation's persisted `tool_name`
- **WHEN** a claim evidence binding references an existing current-run invocation id
- **AND** binding `tool_name` does not match persisted invocation `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-session invocations with matching tool names
- **WHEN** Executor emits legal `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
#### Scenario: gatekeeper fail prevents PASS
- **WHEN** `gatekeeper_result.status` is `fail`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL NOT be `PASS`
#### Scenario: gatekeeper reject bypasses Verifier
- **WHEN** normalized Gatekeeper status is REJECT
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute
#### Scenario: invocation reference failure downgrades to reject
- **WHEN** `gatekeeper_result.failed_rules` contains `evidence.invocation_ref`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL set the effective verdict to `REJECT`
#### Scenario: gatekeeper low confidence is bounded
- **WHEN** normalized Gatekeeper status is LOW_CONFID with at least one passed binding
- **THEN** verified input SHALL contain only passed bindings
- **AND** effective verdict SHALL NOT exceed LOW_CONFID
### Requirement: Verifier claim checks SHALL use a fixed derivability classification set
The Verifier SHALL classify each structured claim using a fixed derivability classification set.
@@ -489,36 +464,36 @@ Gatekeeper SHALL validate that Executor evidence bindings point to real current-
- **AND** `failed_rules` SHALL include `evidence.excerpt_mismatch`
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
Verifier SHALL judge structured claims primarily against Gatekeeper-verified claim-local evidence excerpts.
Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.
#### Scenario: Verified excerpt supports direct observation
- **WHEN** `gatekeeper_result.severity=none`
- **AND** a claim's verified evidence excerpts directly contain the claim's concrete facts
- **WHEN** verdict ceiling is PASS
- **AND** a claim's verified evidence matched text directly contains the claim's concrete facts
- **THEN** Verifier MAY classify that claim as `direct_observation`
#### Scenario: Tool trace summary is navigation context
- **WHEN** `executor_structured_output.claims[].evidence_bindings` are available
- **THEN** Verifier SHALL use `tool_trace_summary` as navigation and audit context
- **AND** it SHALL NOT require `tool_trace_summary.output_summary` to contain every fact already present in verified claim-local evidence
#### Scenario: Verified evidence is complete Verifier context
- **WHEN** verified claims and evidence are available
- **THEN** Verifier SHALL use them as its evidence context
- **AND** it SHALL NOT require or request a complete tool trace summary
- **AND** it SHALL NOT read raw Executor or unreferenced tool material
### Requirement: Gatekeeper severity SHALL constrain effective verdict
Runtime effective verdict calculation SHALL treat Gatekeeper severity as a hard upper bound.
Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.
#### Scenario: Reject severity prevents PASS
#### Scenario: Reject severity bypasses Verifier
- **WHEN** `gatekeeper_result.severity=reject`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL be `REJECT`
- **THEN** the Graph SHALL route to pre-verification Fallback without invoking Verifier
- **AND** it SHALL NOT fabricate an effective diagnostic verdict
#### Scenario: Low confidence severity prevents PASS
- **WHEN** `gatekeeper_result.severity=low_confid`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL be `LOW_CONFID`
- **THEN** deterministic effective-verdict calculation SHALL downgrade the result
- **AND** effective verdict SHALL be `LOW_CONFID`
#### Scenario: Gatekeeper audit includes severity
- **WHEN** verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
- **WHEN** Graph verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include available `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
### Requirement: Verifier input hook SHALL only perform narrow compatibility backfill
The verifier input hook SHALL avoid converting broad tool summaries into precise evidence references.
@@ -3,9 +3,7 @@
## Purpose
Separate multi-turn conversation metadata from per-execution diagnosis state. `sessionId` identifies the conversation context, while `runId` identifies one replayable diagnosis execution and scopes Trace, Feedback, Evaluation, AIOps, and case-library provenance.
## Requirements
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
@@ -179,3 +177,58 @@ The change SHALL verify both runtime behavior and evaluation baseline impact.
- **WHEN** verification is complete
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression
### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`.
#### Scenario: Graph reaches Composer
- **WHEN** a complex Chat Graph terminates through Composer with a safe answer
- **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
- **AND** the summary SHALL be derived from the current Run's actual orchestration events
#### Scenario: Graph reaches handled Fallback
- **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer
- **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true`
- **AND** the Run status SHALL be SUCCESS
#### Scenario: Unhandled execution fails after events exist
- **WHEN** an unhandled failure occurs after one or more real Graph events are available
- **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run
- **AND** it SHALL NOT add a node or transition that did not occur
#### Scenario: Orchestration trace content is inspected
- **WHEN** orchestration trace JSON is serialized
- **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
- **AND** it SHALL NOT contain data owned by another run
### Requirement: Trace API SHALL expose orchestration trace only on the run object
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`.
#### Scenario: Exact StateGraph run trace is queried
- **WHEN** a caller queries a successful new StateGraph Chat run
- **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object
- **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field
- **AND** no raw orchestration trace field SHALL be added
#### Scenario: Historical or non-StateGraph run is queried
- **WHEN** the selected DiagnosisRun has null orchestration trace
- **THEN** `run.orchestrationTrace` MAY be null
- **AND** the service SHALL NOT synthesize historical events or read another run's trace
### Requirement: Orchestration trace migration SHALL be additive and nullable
The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change.
#### Scenario: Migration is applied
- **WHEN** Flyway applies the stage 3 migration
- **THEN** existing DiagnosisRun rows SHALL remain valid without backfill
- **AND** no other table or column SHALL be changed by the stage 3 schema migration