122 lines
7.2 KiB
Markdown
122 lines
7.2 KiB
Markdown
## ADDED Requirements
|
|
|
|
### Requirement: Diagnosis Graph tests SHALL have three authoritative layers
|
|
|
|
The test suite SHALL use Workflow, Node Contract, and ChatService Integration layers as the authoritative ISS-011 verification structure while retaining focused component tests for failure localization.
|
|
|
|
#### Scenario: Test architecture is inspected
|
|
- **WHEN** the Diagnosis Graph test sources are listed
|
|
- **THEN** `DiagnosisGraphWorkflowTest`, `DiagnosisGraphNodeContractTest`, and `ChatServiceGraphIntegrationTest` SHALL exist
|
|
- **AND** no `ChatServiceSequentialAgentTest` SHALL exist
|
|
|
|
#### Scenario: Focused unit tests are retained
|
|
- **WHEN** a single parser, adapter, Gatekeeper, projection, retry preparer, trace builder, or Fallback contract fails
|
|
- **THEN** a focused component test SHALL be able to identify that boundary
|
|
- **AND** the suite SHALL NOT require all component assertions to be duplicated in the three authoritative classes
|
|
|
|
### Requirement: Workflow tests SHALL cover every bounded routing class
|
|
|
|
`DiagnosisGraphWorkflowTest` SHALL execute the real compiled topology with scripted nodes and SHALL verify path order, retry ownership, bounded termination, and orchestration events without invoking models or tools.
|
|
|
|
#### Scenario: Normal and Planner paths are tested
|
|
- **WHEN** the Workflow suite runs
|
|
- **THEN** it SHALL cover PASS, Planner INVALID_OUTPUT/RETRYABLE_FAILED one-time retry success, second technical failure, NON_RETRYABLE_FAILED, and fail-closed unknown status
|
|
- **AND** Planner terminal failures SHALL skip Executor and reach Fallback
|
|
|
|
#### Scenario: Executor and Gatekeeper paths are tested
|
|
- **WHEN** the Workflow suite runs
|
|
- **THEN** it SHALL cover Executor FAILED, TOOL_BLOCKED, INVALID_OUTPUT, legal no-evidence, Gatekeeper PASS, LOW_CONFID with and without verified binding, REJECT, and unknown result
|
|
- **AND** unsafe pre-verification outcomes SHALL skip Verifier
|
|
|
|
#### Scenario: Verifier evidence retry paths are tested
|
|
- **WHEN** the Workflow suite runs
|
|
- **THEN** it SHALL cover Verifier one-time technical retry, retry exhaustion, NON_RETRYABLE_FAILED, REJECT, critical evidence retry, no valid gap, non-critical gap, ceiling-driven LOW_CONFID, and second LOW_CONFID termination
|
|
- **AND** evidence retry SHALL occur at most once
|
|
|
|
#### Scenario: Composer paths are tested
|
|
- **WHEN** the Workflow suite runs
|
|
- **THEN** it SHALL cover Composer one-time technical retry, retry exhaustion, NON_RETRYABLE_FAILED, normal completion, and deterministic post-verification Fallback
|
|
- **AND** every terminal path SHALL have events matching the executed node sequence
|
|
|
|
### Requirement: Node Contract tests SHALL enforce explicit safe state projection
|
|
|
|
`DiagnosisGraphNodeContractTest` and focused component tests SHALL verify that Nodes receive only allowed state, use the current RunnableConfig identity, return standardized statuses, and never promote unverified material.
|
|
|
|
#### Scenario: Agent and Gatekeeper config is inspected
|
|
- **WHEN** Planner, Executor, Verifier, Composer, or Gatekeeper is invoked
|
|
- **THEN** the exact current RunnableConfig SHALL be preserved
|
|
- **AND** Gatekeeper SHALL validate the current run exactly once per Executor round
|
|
|
|
#### Scenario: Executor returns a legal limited result
|
|
- **WHEN** tool data is empty or a tool failed but Executor still returns a legal `executor_evidence_v2` with limitations
|
|
- **THEN** Executor status SHALL be COMPLETED
|
|
- **AND** the workflow SHALL continue to Gatekeeper
|
|
|
|
#### Scenario: Gatekeeper returns partial or unsafe evidence
|
|
- **WHEN** Gatekeeper is REJECT with any partial passed binding, LOW_CONFID with zero passed binding, missing, or unknown
|
|
- **THEN** the path SHALL fail closed to pre-verification Fallback
|
|
- **AND** no Executor claim from a passed or failed binding SHALL appear in the answer
|
|
|
|
#### Scenario: Verified-only input is projected
|
|
- **WHEN** Gatekeeper PASS or continuable LOW_CONFID reaches Verified Input
|
|
- **THEN** Verifier SHALL receive only claims/bindings matched to passed checked bindings and their `matched_text`
|
|
- **AND** it SHALL NOT receive unreferenced tool results, raw Executor text, or full tool trace
|
|
|
|
#### Scenario: Verifier execution fails
|
|
- **WHEN** Verifier output is invalid or invocation fails
|
|
- **THEN** only `verifier_status` SHALL express the execution failure
|
|
- **AND** no model/effective diagnostic verdict SHALL be fabricated
|
|
|
|
#### Scenario: Evidence retry revalidates the full snapshot
|
|
- **WHEN** one critical evidence retry occurs
|
|
- **THEN** the second Executor input SHALL contain prior verified material and incremental-query constraints
|
|
- **AND** the second complete Executor snapshot SHALL pass through Gatekeeper again without reusing the first verdict
|
|
|
|
### Requirement: Chat integration tests SHALL verify Run-owned public behavior
|
|
|
|
`ChatServiceGraphIntegrationTest` SHALL verify the public ChatResult and current DiagnosisRun lifecycle without binding to internal Agent call order.
|
|
|
|
#### Scenario: Safe answer completes
|
|
- **WHEN** Graph returns a Composer or handled Fallback answer
|
|
- **THEN** ChatResult SHALL preserve answer/sessionId/runId
|
|
- **AND** the current Run SHALL persist SUCCESS, agent_flow, metrics, self-evaluation, non-empty orchestration trace, and Eval invocation
|
|
|
|
#### Scenario: Unsafe completion fails
|
|
- **WHEN** Graph throws an unhandled failure, returns no state, or has a blank final answer
|
|
- **THEN** the current Run SHALL persist FAILED when possible
|
|
- **AND** Eval SHALL NOT run
|
|
- **AND** only a real partial trace SHALL be retained
|
|
|
|
#### Scenario: Multiple runs share one session
|
|
- **WHEN** two complex Chat requests use the same sessionId
|
|
- **THEN** each SHALL receive a distinct runId
|
|
- **AND** trace/evaluation/metrics SHALL remain owned by their current run
|
|
|
|
### Requirement: Legacy implementation tests SHALL retire without losing safety regressions
|
|
|
|
The stage 4 suite SHALL remove tests that make Sequential or Hook payload internals correctness criteria while preserving equivalent public and security contracts.
|
|
|
|
#### Scenario: Hook implementation test is retired
|
|
- **WHEN** explicit Gatekeeper and Verified Input Nodes are authoritative
|
|
- **THEN** `VerifierInputHookTest` SHALL NOT exist
|
|
- **AND** parser sanitization, Gatekeeper reference fidelity, passed-binding projection, no-evidence, REJECT, and safe Composer behavior SHALL remain covered by independent tests
|
|
|
|
#### Scenario: Retained regression suite runs
|
|
- **WHEN** stage 4 is accepted
|
|
- **THEN** Controller, Trace, Repository, Gatekeeper service, Composer/protocol, Eval, Workflow, Node Contract, and Chat integration tests SHALL pass
|
|
- **AND** Maven test compilation SHALL pass
|
|
|
|
### Requirement: Stage 4 SHALL remain a test-only change
|
|
|
|
The change SHALL reorganize and strengthen automated tests without changing production runtime behavior and SHALL defer final live verification to stage 5.
|
|
|
|
#### Scenario: Source diff is inspected
|
|
- **WHEN** stage 4 implementation completes
|
|
- **THEN** no file under `src/main` SHALL be changed by this stage
|
|
- **AND** OpenSpec/devflow/test files MAY change
|
|
|
|
#### Scenario: Stage 4 verification completes
|
|
- **WHEN** the test suite and static gates pass
|
|
- **THEN** Maven live startup, `logs/` inspection, and `scripts/query_mysql.py` database verification SHALL remain not run
|
|
- **AND** acceptance SHALL record them as reserved for stage 5
|