test(graph): establish diagnosis stategraph suite

This commit is contained in:
zhuyongxin
2026-07-17 18:53:29 +08:00
parent 99e490f227
commit 208a231113
19 changed files with 868 additions and 395 deletions
@@ -0,0 +1,124 @@
# chat-diagnosis-stategraph-test-suite Specification
## Purpose
TBD - created by archiving change chat-diagnosis-stategraph-test-suite. Update Purpose after archive.
## Requirements
### Requirement: Diagnosis Graph tests SHALL have three authoritative layers
The test suite SHALL use Workflow, Node Contract, and ChatService Integration layers as the authoritative ISS-011 verification structure while retaining focused component tests for failure localization.
#### Scenario: Test architecture is inspected
- **WHEN** the Diagnosis Graph test sources are listed
- **THEN** `DiagnosisGraphWorkflowTest`, `DiagnosisGraphNodeContractTest`, and `ChatServiceGraphIntegrationTest` SHALL exist
- **AND** no `ChatServiceSequentialAgentTest` SHALL exist
#### Scenario: Focused unit tests are retained
- **WHEN** a single parser, adapter, Gatekeeper, projection, retry preparer, trace builder, or Fallback contract fails
- **THEN** a focused component test SHALL be able to identify that boundary
- **AND** the suite SHALL NOT require all component assertions to be duplicated in the three authoritative classes
### Requirement: Workflow tests SHALL cover every bounded routing class
`DiagnosisGraphWorkflowTest` SHALL execute the real compiled topology with scripted nodes and SHALL verify path order, retry ownership, bounded termination, and orchestration events without invoking models or tools.
#### Scenario: Normal and Planner paths are tested
- **WHEN** the Workflow suite runs
- **THEN** it SHALL cover PASS, Planner INVALID_OUTPUT/RETRYABLE_FAILED one-time retry success, second technical failure, NON_RETRYABLE_FAILED, and fail-closed unknown status
- **AND** Planner terminal failures SHALL skip Executor and reach Fallback
#### Scenario: Executor and Gatekeeper paths are tested
- **WHEN** the Workflow suite runs
- **THEN** it SHALL cover Executor FAILED, TOOL_BLOCKED, INVALID_OUTPUT, legal no-evidence, Gatekeeper PASS, LOW_CONFID with and without verified binding, REJECT, and unknown result
- **AND** unsafe pre-verification outcomes SHALL skip Verifier
#### Scenario: Verifier evidence retry paths are tested
- **WHEN** the Workflow suite runs
- **THEN** it SHALL cover Verifier one-time technical retry, retry exhaustion, NON_RETRYABLE_FAILED, REJECT, critical evidence retry, no valid gap, non-critical gap, ceiling-driven LOW_CONFID, and second LOW_CONFID termination
- **AND** evidence retry SHALL occur at most once
#### Scenario: Composer paths are tested
- **WHEN** the Workflow suite runs
- **THEN** it SHALL cover Composer one-time technical retry, retry exhaustion, NON_RETRYABLE_FAILED, normal completion, and deterministic post-verification Fallback
- **AND** every terminal path SHALL have events matching the executed node sequence
### Requirement: Node Contract tests SHALL enforce explicit safe state projection
`DiagnosisGraphNodeContractTest` and focused component tests SHALL verify that Nodes receive only allowed state, use the current RunnableConfig identity, return standardized statuses, and never promote unverified material.
#### Scenario: Agent and Gatekeeper config is inspected
- **WHEN** Planner, Executor, Verifier, Composer, or Gatekeeper is invoked
- **THEN** the exact current RunnableConfig SHALL be preserved
- **AND** Gatekeeper SHALL validate the current run exactly once per Executor round
#### Scenario: Executor returns a legal limited result
- **WHEN** tool data is empty or a tool failed but Executor still returns a legal `executor_evidence_v2` with limitations
- **THEN** Executor status SHALL be COMPLETED
- **AND** the workflow SHALL continue to Gatekeeper
#### Scenario: Gatekeeper returns partial or unsafe evidence
- **WHEN** Gatekeeper is REJECT with any partial passed binding, LOW_CONFID with zero passed binding, missing, or unknown
- **THEN** the path SHALL fail closed to pre-verification Fallback
- **AND** no Executor claim from a passed or failed binding SHALL appear in the answer
#### Scenario: Verified-only input is projected
- **WHEN** Gatekeeper PASS or continuable LOW_CONFID reaches Verified Input
- **THEN** Verifier SHALL receive only claims/bindings matched to passed checked bindings and their `matched_text`
- **AND** it SHALL NOT receive unreferenced tool results, raw Executor text, or full tool trace
#### Scenario: Verifier execution fails
- **WHEN** Verifier output is invalid or invocation fails
- **THEN** only `verifier_status` SHALL express the execution failure
- **AND** no model/effective diagnostic verdict SHALL be fabricated
#### Scenario: Evidence retry revalidates the full snapshot
- **WHEN** one critical evidence retry occurs
- **THEN** the second Executor input SHALL contain prior verified material and incremental-query constraints
- **AND** the second complete Executor snapshot SHALL pass through Gatekeeper again without reusing the first verdict
### Requirement: Chat integration tests SHALL verify Run-owned public behavior
`ChatServiceGraphIntegrationTest` SHALL verify the public ChatResult and current DiagnosisRun lifecycle without binding to internal Agent call order.
#### Scenario: Safe answer completes
- **WHEN** Graph returns a Composer or handled Fallback answer
- **THEN** ChatResult SHALL preserve answer/sessionId/runId
- **AND** the current Run SHALL persist SUCCESS, agent_flow, metrics, self-evaluation, non-empty orchestration trace, and Eval invocation
#### Scenario: Unsafe completion fails
- **WHEN** Graph throws an unhandled failure, returns no state, or has a blank final answer
- **THEN** the current Run SHALL persist FAILED when possible
- **AND** Eval SHALL NOT run
- **AND** only a real partial trace SHALL be retained
#### Scenario: Multiple runs share one session
- **WHEN** two complex Chat requests use the same sessionId
- **THEN** each SHALL receive a distinct runId
- **AND** trace/evaluation/metrics SHALL remain owned by their current run
### Requirement: Legacy implementation tests SHALL retire without losing safety regressions
The stage 4 suite SHALL remove tests that make Sequential or Hook payload internals correctness criteria while preserving equivalent public and security contracts.
#### Scenario: Hook implementation test is retired
- **WHEN** explicit Gatekeeper and Verified Input Nodes are authoritative
- **THEN** `VerifierInputHookTest` SHALL NOT exist
- **AND** parser sanitization, Gatekeeper reference fidelity, passed-binding projection, no-evidence, REJECT, and safe Composer behavior SHALL remain covered by independent tests
#### Scenario: Retained regression suite runs
- **WHEN** stage 4 is accepted
- **THEN** Controller, Trace, Repository, Gatekeeper service, Composer/protocol, Eval, Workflow, Node Contract, and Chat integration tests SHALL pass
- **AND** Maven test compilation SHALL pass
### Requirement: Stage 4 SHALL remain a test-only change
The change SHALL reorganize and strengthen automated tests without changing production runtime behavior and SHALL defer final live verification to stage 5.
#### Scenario: Source diff is inspected
- **WHEN** stage 4 implementation completes
- **THEN** no file under `src/main` SHALL be changed by this stage
- **AND** OpenSpec/devflow/test files MAY change
#### Scenario: Stage 4 verification completes
- **WHEN** the test suite and static gates pass
- **THEN** Maven live startup, `logs/` inspection, and `scripts/query_mysql.py` database verification SHALL remain not run
- **AND** acceptance SHALL record them as reserved for stage 5