feat(graph): add diagnosis real nodes

This commit is contained in:
zhuyongxin
2026-07-17 16:44:05 +08:00
parent 42ba204532
commit 1460dd1e99
57 changed files with 4380 additions and 710 deletions
@@ -0,0 +1,135 @@
# chat-diagnosis-stategraph-real-nodes Specification
## Purpose
TBD - created by archiving change chat-diagnosis-stategraph-real-nodes. Update Purpose after archive.
## Requirements
### Requirement: Agent adapters SHALL invoke real agents through explicit run config
The system SHALL provide Planner, Executor, Verifier, and Composer adapters that invoke their configured ReactAgent through a minimal invoker port with an explicit `RunnableConfig`. Each adapter SHALL serialize only its declared input fields and append one terminal orchestration event per attempt.
#### Scenario: Adapter invokes an agent
- **WHEN** an adapter receives valid Graph State and a RunnableConfig containing the current run metadata
- **THEN** it SHALL pass its projected input and the same RunnableConfig to the configured invoker
- **AND** existing Agent hooks and ToolCallbacks SHALL be able to observe the current run metadata
#### Scenario: Parent state is inspected
- **WHEN** an adapter builds an Agent input
- **THEN** it SHALL NOT serialize undeclared Graph State, raw prompts, other Agent private state, or orchestration events
### Requirement: Agent outputs SHALL be parsed into explicit execution statuses
The system SHALL use shared Executor, Verifier, and Composer protocol parsers for both Graph Nodes and the legacy path. Legal structured output SHALL map to COMPLETED; invalid JSON or contract shape SHALL map to INVALID_OUTPUT; recognized transient invocation failure SHALL map to RETRYABLE_FAILED where that Agent supports technical retry; unknown or permanent failure SHALL fail closed. A legal Executor no-evidence result SHALL be COMPLETED.
#### Scenario: Legal no-evidence Executor output
- **WHEN** Executor returns a valid `executor_evidence_v2` document containing a legal no-evidence result
- **THEN** Executor status SHALL be COMPLETED
- **AND** the Graph SHALL continue to Gatekeeper
#### Scenario: Invalid structured output
- **WHEN** an Agent returns malformed JSON or violates its required output structure
- **THEN** its adapter SHALL set INVALID_OUTPUT
- **AND** it SHALL NOT fabricate a diagnostic verdict or evidence
#### Scenario: Legacy parser behavior is exercised
- **WHEN** the existing Sequential path parses the same payloads after shared component extraction
- **THEN** its externally observable parser and safe-rendering behavior SHALL remain unchanged
### Requirement: Gatekeeper Node SHALL validate exactly once and fail closed
The Gatekeeper Node SHALL call `ExecutorGatekeeperService.validateRun` exactly once for the current run and Executor structured output. It SHALL preserve the raw result separately from normalized status. Raw pass SHALL normalize to PASS; fail with low-confid severity SHALL normalize to LOW_CONFID; fail with reject severity SHALL normalize to REJECT; missing, unknown, inconsistent, or exceptional results SHALL normalize to REJECT.
#### Scenario: Gatekeeper passes output
- **WHEN** `validateRun` returns a valid pass result
- **THEN** the Node SHALL store the raw result and status PASS
- **AND** validation SHALL have been called exactly once with the current runId
#### Scenario: Gatekeeper result cannot be trusted
- **WHEN** validation throws or returns a missing, unknown, or inconsistent result
- **THEN** the Node SHALL normalize status to REJECT
- **AND** Verifier SHALL NOT receive unverified Executor material
### Requirement: Verified Input Builder SHALL project only passed bindings
PASS and continuable LOW_CONFID results SHALL pass through a Verified Input Builder. The Builder SHALL match passed checked bindings to Executor claims by `claim_id`, `source_invocation_id`, `tool_name`, and `raw_path`, and SHALL produce only filtered `verified_executor_output` plus `verified_evidence` entries containing the matched binding fields and `matched_text`.
#### Scenario: Mixed checked bindings are projected
- **WHEN** Gatekeeper returns both passed and failed checked bindings
- **THEN** only claims and evidence matching passed bindings SHALL be projected
- **AND** failed bindings, hypotheses, unreferenced tool results, and raw Executor text SHALL be absent
#### Scenario: Verifier input is serialized
- **WHEN** the Verifier adapter builds its input
- **THEN** it SHALL include only diagnosis query context, verified Executor output, verified evidence, Gatekeeper audit/ceiling, and permitted retry context
- **AND** it SHALL NOT include complete `tool_trace_summary` or raw Executor output
### Requirement: Verifier SHALL separate execution status from diagnostic verdict
The Verifier adapter SHALL store `verifier_status`, `verifier_model_verdict`, and `effective_verdict` as separate values. A Gatekeeper LOW_CONFID ceiling SHALL prevent model PASS from producing effective PASS. Technical retry SHALL reuse the exact same serialized verified input and SHALL NOT rerun any preceding Node.
#### Scenario: Ceiling limits model verdict
- **WHEN** Gatekeeper ceiling is LOW_CONFID and the model verdict is PASS
- **THEN** effective verdict SHALL be LOW_CONFID
- **AND** verifier execution status SHALL remain COMPLETED
#### Scenario: Verifier technical retry occurs
- **WHEN** the first Verifier attempt returns INVALID_OUTPUT or RETRYABLE_FAILED
- **THEN** its single retry SHALL receive the same serialized input
- **AND** Executor, Gatekeeper, Verified Input, and tools SHALL NOT rerun
### Requirement: Evidence retry SHALL contain only structured critical gaps and incremental constraints
The system SHALL extract evidence gaps only from facts with `is_critical=true` and verification `no_evidence` or `indirect_support`. Retry Prepare SHALL include prior verified output/evidence, structured gaps, deduplicated completed query references, and fixed constraints requiring at most one incremental retry without repeating successful queries. The second Executor invocation SHALL be instructed to return a complete `executor_evidence_v2` snapshot; Java code SHALL NOT merge claim text.
#### Scenario: Critical gaps prepare a retry
- **WHEN** effective verdict is LOW_CONFID, ceiling is PASS, evidence retry count is zero, and at least one qualifying critical gap exists
- **THEN** Retry Prepare SHALL build the bounded retry context and increment evidence retry count once
- **AND** the new Planner stage SHALL use EVIDENCE_GAP_ONLY mode
#### Scenario: Non-critical gap is present
- **WHEN** facts contain only non-critical no-evidence or indirect-support items
- **THEN** no evidence retry SHALL occur
- **AND** the Graph SHALL continue to Composer
#### Scenario: Second Executor input is built
- **WHEN** Planner produces an evidence-gap-only incremental plan
- **THEN** Executor input SHALL prohibit repeating completed queries and require a complete output snapshot preserving prior verified claims
- **AND** the Java layer SHALL NOT semantically merge old and new claims
### Requirement: Composer and Fallback SHALL use only allowed material
Composer SHALL receive only effective verdict and Verifier-allowed claims, missing information, and recommendations. Composer technical retry SHALL reuse the exact same serialized input. A pre-verification Fallback SHALL never output Executor claims; a post-verification Composer Fallback SHALL use only Verifier-allowed material.
#### Scenario: Pre-verification path degrades
- **WHEN** Planner, Executor, Gatekeeper, Verified Input, or Verifier cannot establish trusted material
- **THEN** deterministic Fallback output SHALL contain no Executor claim or raw tool output
#### Scenario: Composer retry is exhausted
- **WHEN** Composer fails after its one technical retry and Verifier-allowed material exists
- **THEN** deterministic Fallback SHALL express only the allowed claims, missing information, and recommendations
- **AND** it SHALL NOT read raw Executor or tool output
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
The real Node action set and CompiledGraph SHALL be constructible and testable, but ChatService, database, Trace API, shared Agent prompts, and the current production routing SHALL remain unchanged until the stage 3 change.
#### Scenario: Stage 2 production isolation is inspected
- **WHEN** this change is accepted
- **THEN** no production ChatService code path SHALL invoke the real Diagnosis Graph
- **AND** no database migration, Trace API field, or shared Prompt contract SHALL be changed
@@ -74,7 +74,7 @@ Executor SHALL route only COMPLETED output to Gatekeeper and SHALL never retry.
### Requirement: Verifier routing SHALL separate technical retry from evidence retry
Verifier INVALID_OUTPUT and RETRYABLE_FAILED SHALL self-retry once with the same verified input. COMPLETED PASS or REJECT SHALL route to Composer. COMPLETED LOW_CONFID SHALL route to one Evidence Retry only when all frozen guards are true; otherwise it SHALL route to Composer. Other outcomes SHALL fail closed.
Verifier INVALID_OUTPUT and RETRYABLE_FAILED SHALL self-retry once with the same verified input. COMPLETED PASS or REJECT SHALL route to Composer. COMPLETED LOW_CONFID SHALL route to one Evidence Retry only when all frozen guards are true, including at least one critical evidence gap; otherwise it SHALL route to Composer. Other outcomes SHALL fail closed.
#### Scenario: Verifier first technical failure
@@ -94,14 +94,14 @@ Verifier INVALID_OUTPUT and RETRYABLE_FAILED SHALL self-retry once with the same
#### Scenario: LOW_CONFID qualifies for evidence retry
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, valid no_evidence or indirect_support facts, and `evidence_retry_count=0`
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one fact having `is_critical=true` and verification `no_evidence` or `indirect_support`, and `evidence_retry_count=0`
- **THEN** Evidence Retry SHALL run once and return to a new Planner stage
- **AND** `planner_retry_count` SHALL reset to 0
- **AND** `evidence_retry_count` SHALL become 1
#### Scenario: LOW_CONFID does not qualify for evidence retry
- **WHEN** ceiling is LOW_CONFID, facts contain no valid gap, or evidence retry count is already 1
- **WHEN** ceiling is LOW_CONFID, facts contain no critical valid gap, or evidence retry count is already 1
- **THEN** Composer SHALL run without another Planner cycle
### Requirement: Composer routing SHALL allow one technical retry and then terminate safely