126 lines
7.7 KiB
Markdown
126 lines
7.7 KiB
Markdown
# chat-diagnosis-stategraph-real-nodes Specification
|
|
|
|
## Purpose
|
|
TBD - created by archiving change chat-diagnosis-stategraph-real-nodes. Update Purpose after archive.
|
|
## Requirements
|
|
### Requirement: Agent adapters SHALL invoke real agents through explicit run config
|
|
|
|
The system SHALL provide Planner, Executor, Verifier, and Composer adapters that invoke their configured ReactAgent through a minimal invoker port with an explicit `RunnableConfig`. Each adapter SHALL serialize only its declared input fields and append one terminal orchestration event per attempt.
|
|
|
|
#### Scenario: Adapter invokes an agent
|
|
|
|
- **WHEN** an adapter receives valid Graph State and a RunnableConfig containing the current run metadata
|
|
- **THEN** it SHALL pass its projected input and the same RunnableConfig to the configured invoker
|
|
- **AND** existing Agent hooks and ToolCallbacks SHALL be able to observe the current run metadata
|
|
|
|
#### Scenario: Parent state is inspected
|
|
|
|
- **WHEN** an adapter builds an Agent input
|
|
- **THEN** it SHALL NOT serialize undeclared Graph State, raw prompts, other Agent private state, or orchestration events
|
|
|
|
### Requirement: Agent outputs SHALL be parsed into explicit execution statuses
|
|
|
|
The system SHALL use shared Executor, Verifier, and Composer protocol parsers for both Graph Nodes and the legacy path. Legal structured output SHALL map to COMPLETED; invalid JSON or contract shape SHALL map to INVALID_OUTPUT; recognized transient invocation failure SHALL map to RETRYABLE_FAILED where that Agent supports technical retry; unknown or permanent failure SHALL fail closed. A legal Executor no-evidence result SHALL be COMPLETED.
|
|
|
|
#### Scenario: Legal no-evidence Executor output
|
|
|
|
- **WHEN** Executor returns a valid `executor_evidence_v2` document containing a legal no-evidence result
|
|
- **THEN** Executor status SHALL be COMPLETED
|
|
- **AND** the Graph SHALL continue to Gatekeeper
|
|
|
|
#### Scenario: Invalid structured output
|
|
|
|
- **WHEN** an Agent returns malformed JSON or violates its required output structure
|
|
- **THEN** its adapter SHALL set INVALID_OUTPUT
|
|
- **AND** it SHALL NOT fabricate a diagnostic verdict or evidence
|
|
|
|
#### Scenario: Legacy parser behavior is exercised
|
|
|
|
- **WHEN** the existing Sequential path parses the same payloads after shared component extraction
|
|
- **THEN** its externally observable parser and safe-rendering behavior SHALL remain unchanged
|
|
|
|
### Requirement: Gatekeeper Node SHALL validate exactly once and fail closed
|
|
|
|
The Gatekeeper Node SHALL call `ExecutorGatekeeperService.validateRun` exactly once for the current run and Executor structured output. It SHALL preserve the raw result separately from normalized status. Raw pass SHALL normalize to PASS; fail with low-confid severity SHALL normalize to LOW_CONFID; fail with reject severity SHALL normalize to REJECT; missing, unknown, inconsistent, or exceptional results SHALL normalize to REJECT.
|
|
|
|
#### Scenario: Gatekeeper passes output
|
|
|
|
- **WHEN** `validateRun` returns a valid pass result
|
|
- **THEN** the Node SHALL store the raw result and status PASS
|
|
- **AND** validation SHALL have been called exactly once with the current runId
|
|
|
|
#### Scenario: Gatekeeper result cannot be trusted
|
|
|
|
- **WHEN** validation throws or returns a missing, unknown, or inconsistent result
|
|
- **THEN** the Node SHALL normalize status to REJECT
|
|
- **AND** Verifier SHALL NOT receive unverified Executor material
|
|
|
|
### Requirement: Verified Input Builder SHALL project only passed bindings
|
|
|
|
PASS and continuable LOW_CONFID results SHALL pass through a Verified Input Builder. The Builder SHALL match passed checked bindings to Executor claims by `claim_id`, `source_invocation_id`, `tool_name`, and `raw_path`, and SHALL produce only filtered `verified_executor_output` plus `verified_evidence` entries containing the matched binding fields and `matched_text`.
|
|
|
|
#### Scenario: Mixed checked bindings are projected
|
|
|
|
- **WHEN** Gatekeeper returns both passed and failed checked bindings
|
|
- **THEN** only claims and evidence matching passed bindings SHALL be projected
|
|
- **AND** failed bindings, hypotheses, unreferenced tool results, and raw Executor text SHALL be absent
|
|
|
|
#### Scenario: Verifier input is serialized
|
|
|
|
- **WHEN** the Verifier adapter builds its input
|
|
- **THEN** it SHALL include only diagnosis query context, verified Executor output, verified evidence, Gatekeeper audit/ceiling, and permitted retry context
|
|
- **AND** it SHALL NOT include complete `tool_trace_summary` or raw Executor output
|
|
|
|
### Requirement: Verifier SHALL separate execution status from diagnostic verdict
|
|
|
|
The Verifier adapter SHALL store `verifier_status`, `verifier_model_verdict`, and `effective_verdict` as separate values. A Gatekeeper LOW_CONFID ceiling SHALL prevent model PASS from producing effective PASS. Technical retry SHALL reuse the exact same serialized verified input and SHALL NOT rerun any preceding Node.
|
|
|
|
#### Scenario: Ceiling limits model verdict
|
|
|
|
- **WHEN** Gatekeeper ceiling is LOW_CONFID and the model verdict is PASS
|
|
- **THEN** effective verdict SHALL be LOW_CONFID
|
|
- **AND** verifier execution status SHALL remain COMPLETED
|
|
|
|
#### Scenario: Verifier technical retry occurs
|
|
|
|
- **WHEN** the first Verifier attempt returns INVALID_OUTPUT or RETRYABLE_FAILED
|
|
- **THEN** its single retry SHALL receive the same serialized input
|
|
- **AND** Executor, Gatekeeper, Verified Input, and tools SHALL NOT rerun
|
|
|
|
### Requirement: Evidence retry SHALL contain only structured critical gaps and incremental constraints
|
|
|
|
The system SHALL extract evidence gaps only from facts with `is_critical=true` and verification `no_evidence` or `indirect_support`. Retry Prepare SHALL include prior verified output/evidence, structured gaps, deduplicated completed query references, and fixed constraints requiring at most one incremental retry without repeating successful queries. The second Executor invocation SHALL be instructed to return a complete `executor_evidence_v2` snapshot; Java code SHALL NOT merge claim text.
|
|
|
|
#### Scenario: Critical gaps prepare a retry
|
|
|
|
- **WHEN** effective verdict is LOW_CONFID, ceiling is PASS, evidence retry count is zero, and at least one qualifying critical gap exists
|
|
- **THEN** Retry Prepare SHALL build the bounded retry context and increment evidence retry count once
|
|
- **AND** the new Planner stage SHALL use EVIDENCE_GAP_ONLY mode
|
|
|
|
#### Scenario: Non-critical gap is present
|
|
|
|
- **WHEN** facts contain only non-critical no-evidence or indirect-support items
|
|
- **THEN** no evidence retry SHALL occur
|
|
- **AND** the Graph SHALL continue to Composer
|
|
|
|
#### Scenario: Second Executor input is built
|
|
|
|
- **WHEN** Planner produces an evidence-gap-only incremental plan
|
|
- **THEN** Executor input SHALL prohibit repeating completed queries and require a complete output snapshot preserving prior verified claims
|
|
- **AND** the Java layer SHALL NOT semantically merge old and new claims
|
|
|
|
### Requirement: Composer and Fallback SHALL use only allowed material
|
|
|
|
Composer SHALL receive only effective verdict and Verifier-allowed claims, missing information, and recommendations. Composer technical retry SHALL reuse the exact same serialized input. A pre-verification Fallback SHALL never output Executor claims; a post-verification Composer Fallback SHALL use only Verifier-allowed material.
|
|
|
|
#### Scenario: Pre-verification path degrades
|
|
|
|
- **WHEN** Planner, Executor, Gatekeeper, Verified Input, or Verifier cannot establish trusted material
|
|
- **THEN** deterministic Fallback output SHALL contain no Executor claim or raw tool output
|
|
|
|
#### Scenario: Composer retry is exhausted
|
|
|
|
- **WHEN** Composer fails after its one technical retry and Verifier-allowed material exists
|
|
- **THEN** deterministic Fallback SHALL express only the allowed claims, missing information, and recommendations
|
|
- **AND** it SHALL NOT read raw Executor or tool output
|