feat(graph): add diagnosis real nodes
This commit is contained in:
+3
@@ -0,0 +1,3 @@
|
||||
ready_at: 2026-07-17
|
||||
devflow: devflow/projects/2026-07-17-chat-diagnosis-stategraph-real-nodes
|
||||
authorization: user-requested-per-stage-archive
|
||||
@@ -0,0 +1 @@
|
||||
Committed by sm-flow after stage 2 Commit checkpoint validation on 2026-07-17.
|
||||
@@ -0,0 +1,106 @@
|
||||
## Context
|
||||
|
||||
阶段 1 已交付未接生产入口的 Diagnosis StateGraph 状态、拓扑、路由和 trace builder。当前复杂 Chat 仍由 `ChatService`、`SequentialAgent`、`VerifierInputHook` 和 `VerifierContextHolder` 编排;Executor 解析在 Hook 内,Verifier/Composer 解析与安全渲染在 ChatService 内。阶段 2 只把真实 ReactAgent 和确定性 Java 服务接入 Graph action ports,并抽出可复用协议组件,生产切换留到阶段 3。
|
||||
|
||||
本 change 涉及 `graph.diagnosis`、Hook 和 ChatService 内部协作,接口影响为 L2。`/api/chat`、数据库、Trace API、Prompt 业务协议和当前生产路由均不改变。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 提供显式透传 `RunnableConfig` 的 Planner、Executor、Verifier、Composer adapters。
|
||||
- 将 Executor、Verifier、Composer 协议解析和安全渲染抽为单一共享实现,并让旧路径委托以保持行为。
|
||||
- 将 Gatekeeper、可信输入投影、evidence retry prepare 和两类 Fallback 实现为确定性 Java Node。
|
||||
- 保证 Verifier 只接收通过 binding 投影的材料,并将执行状态、模型 verdict 和 effective verdict 分离。
|
||||
- 修正 LOW_CONFID evidence retry 仅接受 critical gap 的阶段 1 偏差。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不切换 `ChatService.executeChatComplex` 到 StateGraph。
|
||||
- 不修改 `/api/chat`、数据库、Run/Trace DTO 或持久化。
|
||||
- 不删除 SequentialAgent、VerifierInputHook 或 VerifierContextHolder。
|
||||
- 不修改共享 Verifier Prompt;其输入说明在阶段 3 切换时同步。
|
||||
- 不运行 Maven live E2E、日志或数据库验收。
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. Agent 调用经最小 port 隔离
|
||||
|
||||
`DiagnosisAgentInvoker` 只暴露 `invoke(String, RunnableConfig) -> String`;`ReactAgentDiagnosisInvoker` 包装 `ReactAgent.call(input, config)` 并返回消息文本。各 Agent adapter 自己负责白名单输入序列化、输出解析、状态映射和 orchestration event,不读取完整父 Graph State。
|
||||
|
||||
选择最小 port 而不是直接使用 `ReactAgent.asNode(...)`,因为显式字符串输入便于证明输入白名单、固定技术重试输入和 run config 透传,也避免父 Graph messages/private state 泄漏。ReactAgent/invoker 通过构造注入;本阶段不复制 ChatService 的 Prompt 或 Agent factory,生产实例装配属于阶段 3。
|
||||
|
||||
### 2. 共享协议组件替代复制
|
||||
|
||||
在中立的 `com.superbiz.agent.diagnosis.protocol` 包抽取无状态的 `JsonPayloadSupport`、Executor/Verifier/Composer parser、Composer safe input builder 和 fallback renderer。旧 `VerifierInputHook`/`ChatService` 委托这些组件,新 Nodes 使用同一实现;抽取本身不得改变旧可观察行为。共享组件不得依赖 `graph.diagnosis`、Hook、ChatService 或 ThreadLocal。
|
||||
|
||||
选择共享组件而不是在 Graph 包复制旧私有方法,避免旧 Sequential 与新 Graph 对同一 JSON 契约产生两个真理源。阶段 2 保留旧控制流只是迁移顺序,不形成长期双轨。
|
||||
|
||||
### 3. Node failure 分类 fail closed
|
||||
|
||||
Adapter 将合法结构映射为 COMPLETED,将可识别的临时调用失败映射为 RETRYABLE_FAILED,将非法 JSON/结构映射为 INVALID_OUTPUT;未知或明确不可重试异常映射为 NON_RETRYABLE_FAILED/FAILED。Executor 不论失败类型都不重试,合法 no-evidence 结构仍为 COMPLETED。
|
||||
|
||||
分类器允许注入以便测试。无法确定的异常不猜测为可重试,防止无限或扩大副作用。
|
||||
|
||||
### 4. Gatekeeper 是 Graph 唯一验证入口
|
||||
|
||||
Graph Verifier ReactAgent 不注册旧 `VerifierInputHook`。显式 Gatekeeper Node 从 `RunnableConfig.metadata.runId` 和 `executor_output` 调用 `ExecutorGatekeeperService.validateRun` 恰好一次,同时保存 raw result 与 normalized status:pass 为 PASS;fail/low_confid 为 LOW_CONFID;fail/reject 为 REJECT;缺失、unknown、异常或自相矛盾均为 REJECT。
|
||||
|
||||
旧 Sequential 路径在阶段 3 前仍由 Hook 调用 Gatekeeper。两个入口服务于互斥的控制流,不允许同一次 Graph run 双执行。
|
||||
|
||||
### 5. Verified Input 按 binding 精确投影
|
||||
|
||||
Builder 只接受 Gatekeeper `checked_bindings.status=pass`。它以 `claim_id + source_invocation_id + tool_name + raw_path` 精确关联 Executor claim binding,输出过滤后的 `verified_executor_output` 和只含 `claim_id/source_invocation_id/tool_name/raw_path/matched_text` 的 `verified_evidence`。
|
||||
|
||||
失败 binding、hypotheses、未引用工具结果、完整 `tool_trace_summary` 和 raw Executor 文本不得进入 Verifier。合法零 claim 输出保留空列表,但 LOW_CONFID 零可信 binding 已由 Router 在 Builder 前阻断。
|
||||
|
||||
### 6. Verifier 状态与 verdict 分离
|
||||
|
||||
Verifier parser 产出执行状态和模型 verdict;Adapter 再应用 Gatekeeper ceiling 得到 effective verdict。技术失败只写 `verifier_status`,不得伪造诊断 verdict。ceiling=LOW_CONFID 时模型 PASS 最高只能得到 LOW_CONFID。
|
||||
|
||||
Verifier technical retry 复用首次生成的完全相同输入字符串;Composer technical retry 同样复用首次 allowed-material 输入。重试不得读取变化后的前序 raw state。
|
||||
|
||||
### 7. Evidence retry 使用共享 critical-gap extractor
|
||||
|
||||
`EvidenceGapExtractor` 同时被 Router guard 与 Retry Prepare Node 使用,唯一资格为 `is_critical=true` 且 `verification` 为 `no_evidence` 或 `indirect_support`。这样修复阶段 1 Router 漏检 critical 标记,又避免路由判断与 retry payload 漂移。
|
||||
|
||||
Retry context 包含 prior verified output/evidence、结构化 gaps、从 verified evidence 去重得到的 completed query refs,以及固定约束:最多一次、不重复成功查询、只做增量查询、保留 prior verified claims。第二轮 Executor 被要求执行增量查询但输出完整 `executor_evidence_v2` 快照;Java 不合并 claim 文本,完整快照重新经过 Gatekeeper。
|
||||
|
||||
### 8. 两类 Fallback 使用不同材料边界
|
||||
|
||||
前置验证 Fallback 只使用 reason code、校验状态、工具概况和人工查看 Trace 建议,绝不输出 Executor claim。Composer 后置 Fallback 只使用 Verifier 已允许的 claims、missing info 和 recommendations,不读取 raw Executor/tool output。
|
||||
|
||||
Fallback 由确定性 renderer 生成;若连安全答案都无法生成,异常交给阶段 3 的外层 Run failure handling。
|
||||
|
||||
### 9. 阶段内装配与测试边界
|
||||
|
||||
新增真实 Node action set/factory 装配入口,但不让 ChatService 成为消费者。单元测试使用 fake invoker 验证 Node 契约,使用真实 `ExecutorGatekeeperService` mock 边界验证单次调用,并以 CompiledGraph 场景验证 adapters、critical evidence retry 和 Fallback 路由。旧 Sequential/Hook/Gatekeeper focused tests作为共享抽取回归门禁。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- 等级:L2 内部接口。
|
||||
- 新内部消费者:阶段 3 的 Graph orchestrator/Agent factory。
|
||||
- 旧内部消费者:VerifierInputHook 与 ChatService 改为委托共享 parser/renderer,公开方法与外部响应不变。
|
||||
- 外部 API、DB、Trace、Prompt 输出协议:无变化。
|
||||
- 回滚:revert 本阶段提交;生产仍走旧 Sequential 路径。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [共享解析器抽取造成旧行为漂移] → 保留原输入/输出形态并运行旧 focused tests。
|
||||
- [binding 关联不精确导致证据泄漏] → 使用四元键精确匹配并覆盖 mixed pass/fail binding。
|
||||
- [Gatekeeper 双执行] → Graph Verifier 不注册旧 Hook,并以调用次数测试锁定。
|
||||
- [异常误判为可重试] → 未知异常默认 fail closed;分类器用显式用例覆盖。
|
||||
- [阶段 2 Prompt 与 Node payload 说明暂不一致] → 本阶段不接生产入口;阶段 3 切换时同一 change 更新 Prompt。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 先抽共享协议组件并让旧 Hook/ChatService 委托,运行旧 focused tests。
|
||||
2. 接入 invoker、Agent adapters、Gatekeeper、Verified Input、Retry Prepare 和 Fallback Nodes。
|
||||
3. 修正 Router critical-gap guard,装配未接生产入口的真实 Graph 并运行 Node/Graph tests。
|
||||
4. 归档并提交本阶段;阶段 3 再切换 ChatService、Prompt、DB 和 Trace。
|
||||
|
||||
回滚为完整 revert 本阶段提交;不需要数据库回滚,不存在外部协议迁移。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。阶段 3 前不得扩大范围到生产入口或 Prompt 切换。
|
||||
@@ -0,0 +1,82 @@
|
||||
# Chat Diagnosis StateGraph Real Nodes
|
||||
|
||||
## Why
|
||||
|
||||
阶段 1 已提供可编译的 StateGraph 骨架和 35 个 Fake Node 路由测试,但所有 action ports 仍是测试脚本,尚不能调用现有 ReactAgent、Executor Gatekeeper 或构造安全的 Verifier/Composer 输入。阶段 2 需要把现有 Agent 协议和确定性服务接到 Graph ports,同时保持 ChatService 生产入口仍走旧链路,避免把真实 Node 接入与生产切换混成一个不可回滚阶段。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增 `DiagnosisAgentInvoker` 及 ReactAgent adapter,所有 Agent Node 显式传递 `RunnableConfig`。
|
||||
- 新增 Planner、Executor、Verifier、Composer Node Adapters 和可注入失败分类器。
|
||||
- 抽取 Executor/Verifier/Composer 共享解析组件,旧 Hook/ChatService 委托它们,避免复制协议逻辑。
|
||||
- 新增显式 Gatekeeper Node,调用 `ExecutorGatekeeperService.validateRun` 并 fail-closed 标准化状态。
|
||||
- 新增 Verified Input Builder,只投影 Gatekeeper passed bindings 对应的 claims 与 `matched_text` evidence。
|
||||
- 新增 Evidence Gap Extractor / Retry Prepare Node,仅处理关键 `no_evidence` / `indirect_support` facts。
|
||||
- 新增 Composer Safe Input Builder 和两类固定 Fallback Node 表达。
|
||||
- 修正阶段 1 Router:evidence retry guard 必须要求关键 evidence gap。
|
||||
- 新增 Node 契约和真实 CompiledGraph 集成测试,并保留旧 Sequential/Hook/Gatekeeper 回归。
|
||||
|
||||
本 change 不切换 ChatService 到 StateGraph,不修改 DB/Trace API,也不删除旧 Hook/Sequential 结构。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `chat-diagnosis-stategraph-real-nodes`:提供真实 Agent/Java Node adapters、可信输入投影、关键证据补查上下文和安全 Fallback。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `chat-diagnosis-stategraph-routing-skeleton`:将 LOW_CONFID evidence retry guard 收紧为仅接受 `is_critical=true` 的 `no_evidence` 或 `indirect_support` facts,与阶段 0 冻结基线一致。
|
||||
|
||||
## Scope
|
||||
|
||||
### In Scope
|
||||
|
||||
- `com.superbiz.agent.graph.diagnosis` 下的 Agent adapters、Gatekeeper、Verified Input、Evidence Retry、Composer Input 和 Fallback。
|
||||
- 可被旧 Hook/ChatService 与新 Nodes 共用的输出解析/安全渲染组件。
|
||||
- 旧 Hook/ChatService 的行为保持型委托重构。
|
||||
- `DiagnosisGraphRouter` 的 critical gap guard 修正。
|
||||
- Node 单元测试、真实 CompiledGraph adapter tests、旧路径 focused regressions。
|
||||
|
||||
### Out of Scope
|
||||
|
||||
- 将 `ChatService.executeChatComplex` 替换为 Graph invocation。
|
||||
- 修改 Agent 基础 Prompt 业务语义或输出协议。
|
||||
- 移除 `VerifierInputHook` / `VerifierContextHolder` / SequentialAgent。
|
||||
- Flyway、DiagnosisRun orchestration trace 字段、Trace DTO/API。
|
||||
- 最终全量 Graph 测试体系替换。
|
||||
- Maven E2E、`logs/` 和数据库核验。
|
||||
|
||||
## Context Constraints
|
||||
|
||||
- 阶段 0/1 archives 是执行基线;Node 不得改变条件边和计数所有权。
|
||||
- Graph Verifier 只接收 `verified_executor_output`、`verified_evidence`、gatekeeper audit/ceiling、query 和 retry context,不接收完整 tool trace 或 Executor raw text。
|
||||
- Gatekeeper validation 使用当前 `runId`,缺失/未知/异常一律 REJECT。
|
||||
- Gatekeeper raw result 与 normalized status 分开保存。
|
||||
- Verifier execution status、model verdict 和 effective verdict 分离;status 永远不写进 verdict。
|
||||
- Planner/Verifier/Composer retry 输入由相同 state 白名单投影构造,技术 retry 不重跑前序节点。
|
||||
- Executor 不重试;合法 no-evidence 仍为 COMPLETED。
|
||||
- Evidence Retry Context 保留 prior verified output/evidence、关键 gaps、completed query refs 和固定约束;Java 不合并 claim 文本。
|
||||
- 旧 Sequential path 在阶段 3 前继续工作;共享解析器重构必须通过旧 focused tests。
|
||||
- 本阶段不修改 `/api/chat`、DB 或当前运行路由。
|
||||
|
||||
## Acceptance
|
||||
|
||||
- ReactAgent invoker 透传输入与 RunnableConfig,现有 Agent Hook/ToolCallback 可继续工作。
|
||||
- Planner 合法 JSON 为 COMPLETED;非法 JSON/结构为 INVALID_OUTPUT;临时失败与不可重试失败分开。
|
||||
- Executor 合法 `executor_evidence_v2`(包括 no-evidence)为 COMPLETED;非法结构、工具阻断、其他失败正确映射且不重试。
|
||||
- Gatekeeper Node 每轮只调用一次 `validateRun`,unknown/error fail closed。
|
||||
- PASS 与可继续 LOW_CONFID 都经过 Verified Input Builder;失败 binding、未引用工具结果和 raw Executor text 不进入 Verifier。
|
||||
- Verifier input 只有白名单字段;输出分离 status/model/effective verdict,ceiling 正确限制 PASS。
|
||||
- Verifier/Composer technical retry 使用相同序列化输入。
|
||||
- Evidence retry 只从关键 gap 生成,包含 prior verified data/completed refs/约束,最多一次。
|
||||
- 第二轮 Executor 输入明确要求增量查询和完整 `executor_evidence_v2` 快照,不由 Java 合并 claims。
|
||||
- Pre-verification Fallback 不输出 Executor claim;post-verification Composer fallback 只使用 allowed material。
|
||||
- 新 Node/Graph tests 和旧路径 focused regressions 通过;不运行 Maven E2E。
|
||||
|
||||
## Risks
|
||||
|
||||
- 共享解析器抽取可能改变旧 Sequential 行为;旧 ChatService/Hook tests 必须同批通过。
|
||||
- Gatekeeper checked binding 与 claim binding 关联错误可能泄漏失败 evidence;使用 claim_id + invocation/tool/path 精确匹配并测试 mixed binding。
|
||||
- 异常分类依赖 cause/message;未知异常默认不可重试或 FAILED,安全优先。
|
||||
- Prompt 当前仍描述旧 Hook payload;本阶段只验证 Node input contract,阶段 3 在生产切换时同步 Verifier Prompt 输入说明,避免旧生产路径提前不兼容。
|
||||
+132
@@ -0,0 +1,132 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Agent adapters SHALL invoke real agents through explicit run config
|
||||
|
||||
The system SHALL provide Planner, Executor, Verifier, and Composer adapters that invoke their configured ReactAgent through a minimal invoker port with an explicit `RunnableConfig`. Each adapter SHALL serialize only its declared input fields and append one terminal orchestration event per attempt.
|
||||
|
||||
#### Scenario: Adapter invokes an agent
|
||||
|
||||
- **WHEN** an adapter receives valid Graph State and a RunnableConfig containing the current run metadata
|
||||
- **THEN** it SHALL pass its projected input and the same RunnableConfig to the configured invoker
|
||||
- **AND** existing Agent hooks and ToolCallbacks SHALL be able to observe the current run metadata
|
||||
|
||||
#### Scenario: Parent state is inspected
|
||||
|
||||
- **WHEN** an adapter builds an Agent input
|
||||
- **THEN** it SHALL NOT serialize undeclared Graph State, raw prompts, other Agent private state, or orchestration events
|
||||
|
||||
### Requirement: Agent outputs SHALL be parsed into explicit execution statuses
|
||||
|
||||
The system SHALL use shared Executor, Verifier, and Composer protocol parsers for both Graph Nodes and the legacy path. Legal structured output SHALL map to COMPLETED; invalid JSON or contract shape SHALL map to INVALID_OUTPUT; recognized transient invocation failure SHALL map to RETRYABLE_FAILED where that Agent supports technical retry; unknown or permanent failure SHALL fail closed. A legal Executor no-evidence result SHALL be COMPLETED.
|
||||
|
||||
#### Scenario: Legal no-evidence Executor output
|
||||
|
||||
- **WHEN** Executor returns a valid `executor_evidence_v2` document containing a legal no-evidence result
|
||||
- **THEN** Executor status SHALL be COMPLETED
|
||||
- **AND** the Graph SHALL continue to Gatekeeper
|
||||
|
||||
#### Scenario: Invalid structured output
|
||||
|
||||
- **WHEN** an Agent returns malformed JSON or violates its required output structure
|
||||
- **THEN** its adapter SHALL set INVALID_OUTPUT
|
||||
- **AND** it SHALL NOT fabricate a diagnostic verdict or evidence
|
||||
|
||||
#### Scenario: Legacy parser behavior is exercised
|
||||
|
||||
- **WHEN** the existing Sequential path parses the same payloads after shared component extraction
|
||||
- **THEN** its externally observable parser and safe-rendering behavior SHALL remain unchanged
|
||||
|
||||
### Requirement: Gatekeeper Node SHALL validate exactly once and fail closed
|
||||
|
||||
The Gatekeeper Node SHALL call `ExecutorGatekeeperService.validateRun` exactly once for the current run and Executor structured output. It SHALL preserve the raw result separately from normalized status. Raw pass SHALL normalize to PASS; fail with low-confid severity SHALL normalize to LOW_CONFID; fail with reject severity SHALL normalize to REJECT; missing, unknown, inconsistent, or exceptional results SHALL normalize to REJECT.
|
||||
|
||||
#### Scenario: Gatekeeper passes output
|
||||
|
||||
- **WHEN** `validateRun` returns a valid pass result
|
||||
- **THEN** the Node SHALL store the raw result and status PASS
|
||||
- **AND** validation SHALL have been called exactly once with the current runId
|
||||
|
||||
#### Scenario: Gatekeeper result cannot be trusted
|
||||
|
||||
- **WHEN** validation throws or returns a missing, unknown, or inconsistent result
|
||||
- **THEN** the Node SHALL normalize status to REJECT
|
||||
- **AND** Verifier SHALL NOT receive unverified Executor material
|
||||
|
||||
### Requirement: Verified Input Builder SHALL project only passed bindings
|
||||
|
||||
PASS and continuable LOW_CONFID results SHALL pass through a Verified Input Builder. The Builder SHALL match passed checked bindings to Executor claims by `claim_id`, `source_invocation_id`, `tool_name`, and `raw_path`, and SHALL produce only filtered `verified_executor_output` plus `verified_evidence` entries containing the matched binding fields and `matched_text`.
|
||||
|
||||
#### Scenario: Mixed checked bindings are projected
|
||||
|
||||
- **WHEN** Gatekeeper returns both passed and failed checked bindings
|
||||
- **THEN** only claims and evidence matching passed bindings SHALL be projected
|
||||
- **AND** failed bindings, hypotheses, unreferenced tool results, and raw Executor text SHALL be absent
|
||||
|
||||
#### Scenario: Verifier input is serialized
|
||||
|
||||
- **WHEN** the Verifier adapter builds its input
|
||||
- **THEN** it SHALL include only diagnosis query context, verified Executor output, verified evidence, Gatekeeper audit/ceiling, and permitted retry context
|
||||
- **AND** it SHALL NOT include complete `tool_trace_summary` or raw Executor output
|
||||
|
||||
### Requirement: Verifier SHALL separate execution status from diagnostic verdict
|
||||
|
||||
The Verifier adapter SHALL store `verifier_status`, `verifier_model_verdict`, and `effective_verdict` as separate values. A Gatekeeper LOW_CONFID ceiling SHALL prevent model PASS from producing effective PASS. Technical retry SHALL reuse the exact same serialized verified input and SHALL NOT rerun any preceding Node.
|
||||
|
||||
#### Scenario: Ceiling limits model verdict
|
||||
|
||||
- **WHEN** Gatekeeper ceiling is LOW_CONFID and the model verdict is PASS
|
||||
- **THEN** effective verdict SHALL be LOW_CONFID
|
||||
- **AND** verifier execution status SHALL remain COMPLETED
|
||||
|
||||
#### Scenario: Verifier technical retry occurs
|
||||
|
||||
- **WHEN** the first Verifier attempt returns INVALID_OUTPUT or RETRYABLE_FAILED
|
||||
- **THEN** its single retry SHALL receive the same serialized input
|
||||
- **AND** Executor, Gatekeeper, Verified Input, and tools SHALL NOT rerun
|
||||
|
||||
### Requirement: Evidence retry SHALL contain only structured critical gaps and incremental constraints
|
||||
|
||||
The system SHALL extract evidence gaps only from facts with `is_critical=true` and verification `no_evidence` or `indirect_support`. Retry Prepare SHALL include prior verified output/evidence, structured gaps, deduplicated completed query references, and fixed constraints requiring at most one incremental retry without repeating successful queries. The second Executor invocation SHALL be instructed to return a complete `executor_evidence_v2` snapshot; Java code SHALL NOT merge claim text.
|
||||
|
||||
#### Scenario: Critical gaps prepare a retry
|
||||
|
||||
- **WHEN** effective verdict is LOW_CONFID, ceiling is PASS, evidence retry count is zero, and at least one qualifying critical gap exists
|
||||
- **THEN** Retry Prepare SHALL build the bounded retry context and increment evidence retry count once
|
||||
- **AND** the new Planner stage SHALL use EVIDENCE_GAP_ONLY mode
|
||||
|
||||
#### Scenario: Non-critical gap is present
|
||||
|
||||
- **WHEN** facts contain only non-critical no-evidence or indirect-support items
|
||||
- **THEN** no evidence retry SHALL occur
|
||||
- **AND** the Graph SHALL continue to Composer
|
||||
|
||||
#### Scenario: Second Executor input is built
|
||||
|
||||
- **WHEN** Planner produces an evidence-gap-only incremental plan
|
||||
- **THEN** Executor input SHALL prohibit repeating completed queries and require a complete output snapshot preserving prior verified claims
|
||||
- **AND** the Java layer SHALL NOT semantically merge old and new claims
|
||||
|
||||
### Requirement: Composer and Fallback SHALL use only allowed material
|
||||
|
||||
Composer SHALL receive only effective verdict and Verifier-allowed claims, missing information, and recommendations. Composer technical retry SHALL reuse the exact same serialized input. A pre-verification Fallback SHALL never output Executor claims; a post-verification Composer Fallback SHALL use only Verifier-allowed material.
|
||||
|
||||
#### Scenario: Pre-verification path degrades
|
||||
|
||||
- **WHEN** Planner, Executor, Gatekeeper, Verified Input, or Verifier cannot establish trusted material
|
||||
- **THEN** deterministic Fallback output SHALL contain no Executor claim or raw tool output
|
||||
|
||||
#### Scenario: Composer retry is exhausted
|
||||
|
||||
- **WHEN** Composer fails after its one technical retry and Verifier-allowed material exists
|
||||
- **THEN** deterministic Fallback SHALL express only the allowed claims, missing information, and recommendations
|
||||
- **AND** it SHALL NOT read raw Executor or tool output
|
||||
|
||||
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
|
||||
|
||||
The real Node action set and CompiledGraph SHALL be constructible and testable, but ChatService, database, Trace API, shared Agent prompts, and the current production routing SHALL remain unchanged until the stage 3 change.
|
||||
|
||||
#### Scenario: Stage 2 production isolation is inspected
|
||||
|
||||
- **WHEN** this change is accepted
|
||||
- **THEN** no production ChatService code path SHALL invoke the real Diagnosis Graph
|
||||
- **AND** no database migration, Trace API field, or shared Prompt contract SHALL be changed
|
||||
+33
@@ -0,0 +1,33 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Verifier routing SHALL separate technical retry from evidence retry
|
||||
|
||||
Verifier INVALID_OUTPUT and RETRYABLE_FAILED SHALL self-retry once with the same verified input. COMPLETED PASS or REJECT SHALL route to Composer. COMPLETED LOW_CONFID SHALL route to one Evidence Retry only when all frozen guards are true, including at least one critical evidence gap; otherwise it SHALL route to Composer. Other outcomes SHALL fail closed.
|
||||
|
||||
#### Scenario: Verifier first technical failure
|
||||
|
||||
- **WHEN** Verifier first returns INVALID_OUTPUT or RETRYABLE_FAILED
|
||||
- **THEN** only Verifier SHALL run again
|
||||
- **AND** Gatekeeper, Executor, and tools SHALL NOT rerun
|
||||
|
||||
#### Scenario: Verifier retry is exhausted
|
||||
|
||||
- **WHEN** Verifier returns a technical failure after `verifier_retry_count=1`
|
||||
- **THEN** the Graph SHALL route to Fallback
|
||||
|
||||
#### Scenario: Verifier verdict reaches Composer
|
||||
|
||||
- **WHEN** Verifier completes with effective PASS or REJECT
|
||||
- **THEN** Composer SHALL run
|
||||
|
||||
#### Scenario: LOW_CONFID qualifies for evidence retry
|
||||
|
||||
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one fact having `is_critical=true` and verification `no_evidence` or `indirect_support`, and `evidence_retry_count=0`
|
||||
- **THEN** Evidence Retry SHALL run once and return to a new Planner stage
|
||||
- **AND** `planner_retry_count` SHALL reset to 0
|
||||
- **AND** `evidence_retry_count` SHALL become 1
|
||||
|
||||
#### Scenario: LOW_CONFID does not qualify for evidence retry
|
||||
|
||||
- **WHEN** ceiling is LOW_CONFID, facts contain no critical valid gap, or evidence retry count is already 1
|
||||
- **THEN** Composer SHALL run without another Planner cycle
|
||||
@@ -0,0 +1,44 @@
|
||||
## 1. Shared Protocol Components
|
||||
|
||||
- [x] 1.1 Extract stateless JSON payload and Executor evidence parsing components into neutral `com.superbiz.agent.diagnosis.protocol`, preserving legal no-evidence behavior and introducing no dependency on Graph, Hook, ChatService, or ThreadLocal.
|
||||
- [x] 1.2 Extract Verifier and Composer output parsers, effective-verdict calculation, Composer safe-input construction, and deterministic safe rendering from ChatService.
|
||||
- [x] 1.3 Make VerifierInputHook and ChatService delegate to the shared components without changing their public methods or legacy payload/output behavior.
|
||||
- [x] 1.4 Run `VerifierInputHookTest`, `ChatServiceSequentialAgentTest`, and `ExecutorGatekeeperServiceTest` after the extraction and fix only behavior regressions in scope.
|
||||
|
||||
## 2. Agent Invocation And Adapters
|
||||
|
||||
- [x] 2.1 Implement constructor-injected `DiagnosisAgentInvoker` and a ReactAgent wrapper that forwards the exact input and RunnableConfig and returns AssistantMessage text, without adding a second Prompt/Agent factory.
|
||||
- [x] 2.2 Implement a fail-closed, injectable Node failure classifier for invalid output, retryable invocation failure, and non-retryable failure.
|
||||
- [x] 2.3 Implement Planner adapter with NORMAL/EVIDENCE_GAP_ONLY input projection, structured plan parsing, status mapping, and one event per attempt.
|
||||
- [x] 2.4 Implement Executor adapter with incremental retry instructions, legal no-evidence completion, no technical retry, and complete-snapshot output parsing.
|
||||
- [x] 2.5 Add focused adapter tests proving input whitelists, RunnableConfig identity, failure classification, event shape, legal no-evidence handling, and absence of parent-state leakage.
|
||||
|
||||
## 3. Gatekeeper And Verified Projection
|
||||
|
||||
- [x] 3.1 Implement the explicit Gatekeeper Node using current runId and `ExecutorGatekeeperService.validateRun`, preserving raw result separately from normalized status.
|
||||
- [x] 3.2 Normalize missing, unknown, inconsistent, and exceptional Gatekeeper results to REJECT and record a deterministic reason code.
|
||||
- [x] 3.3 Implement Verified Input Builder matching passed bindings by claim/invocation/tool/path and projecting only filtered claims plus matched-text evidence.
|
||||
- [x] 3.4 Add mixed-binding tests proving one Gatekeeper call, correct PASS/LOW_CONFID/REJECT normalization, precise pass projection, and exclusion of failed, unreferenced, hypothesis, tool-summary, and raw materials.
|
||||
|
||||
## 4. Verifier, Evidence Retry, Composer And Fallback
|
||||
|
||||
- [x] 4.1 Implement Verifier adapter with whitelisted verified input, separate execution/model/effective fields, ceiling enforcement, and byte-identical technical retry input.
|
||||
- [x] 4.2 Implement shared `EvidenceGapExtractor` and update `DiagnosisGraphRouter` so only critical no-evidence/indirect-support facts qualify.
|
||||
- [x] 4.3 Implement Evidence Retry Prepare Node with prior verified data, structured gaps, deduplicated completed query refs, bounded constraints, and planner retry reset.
|
||||
- [x] 4.4 Implement Composer adapter and safe-input builder with byte-identical technical retry input and no raw Executor/tool access.
|
||||
- [x] 4.5 Implement distinct deterministic pre-verification and post-verification Fallback inputs/rendering with their allowed-material boundaries.
|
||||
- [x] 4.6 Add focused tests for ceiling, fixed retry inputs, non-critical gap rejection, critical gap context, no Java claim merge, and both Fallback safety boundaries.
|
||||
|
||||
## 5. Real Graph Assembly
|
||||
|
||||
- [x] 5.1 Extend the diagnosis action set/factory to construct a CompiledGraph from the real Agent and deterministic Node dependencies without registering legacy VerifierInputHook on the Graph Verifier.
|
||||
- [x] 5.2 Add CompiledGraph integration tests for PASS, legal no-evidence, Gatekeeper REJECT/LOW_CONFID, critical evidence retry, exhausted Agent technical retry, and Composer fallback paths.
|
||||
- [x] 5.3 Verify Graph events/transitions, per-node invocation counts, same-input retries, one full-snapshot Gatekeeper revalidation after evidence retry, and bounded termination.
|
||||
- [x] 5.4 Confirm the first completed Node module and final assembly against design/specs, with no core TODO or placeholder implementation.
|
||||
|
||||
## 6. Stage 2 Verification And Handoff
|
||||
|
||||
- [x] 6.1 Run new Node/Graph focused tests plus `DiagnosisGraphRoutingTest` and `DiagnosisOrchestrationTraceBuilderTest`.
|
||||
- [x] 6.2 Run legacy focused regressions `ChatServiceSequentialAgentTest`, `VerifierInputHookTest`, and `ExecutorGatekeeperServiceTest`, then run Maven test compilation.
|
||||
- [x] 6.3 Run OpenSpec strict validation, `git diff --check`, and source/reference checks proving ChatService does not invoke the real Graph and no DB/Trace/Prompt contract changed.
|
||||
- [x] 6.4 Record that Maven live E2E, `logs/`, and database verification were intentionally not run in stage 2 and are reserved for stage 5.
|
||||
Reference in New Issue
Block a user