122 lines
6.0 KiB
Markdown
122 lines
6.0 KiB
Markdown
## ADDED Requirements
|
||
|
||
### Requirement: StateGraph design baseline SHALL be versioned and authoritative
|
||
|
||
The project SHALL maintain an archived ISS-011 design baseline that defines Graph State, routing, retry limits, fallback classes, audit boundaries, migration stages, and test replacement scope without claiming runtime cutover is complete.
|
||
|
||
#### Scenario: Later stage starts implementation
|
||
|
||
- **WHEN** an ISS-011 stage 1–5 OpenSpec change is proposed
|
||
- **THEN** its context and design SHALL reference the archived stage 0 baseline
|
||
- **AND** any intentional deviation SHALL be resolved through that stage's OpenSpec before implementation
|
||
|
||
#### Scenario: Stage zero is accepted
|
||
|
||
- **WHEN** the design-freeze change is archived
|
||
- **THEN** no Java, SQL, Prompt, configuration, or runtime behavior SHALL have been changed by this change
|
||
- **AND** runtime specs SHALL NOT claim StateGraph cutover is already implemented
|
||
|
||
### Requirement: Graph State design SHALL use explicit bounded control state
|
||
|
||
Every cross-node field SHALL have one purpose, owner, and update strategy. Fields SHALL use Replace semantics except bounded `orchestration_events`, which SHALL use Append.
|
||
|
||
#### Scenario: Node state is designed
|
||
|
||
- **WHEN** an Agent or Java Node reads or writes parent Graph State
|
||
- **THEN** its allowed input projection and output fields SHALL be explicit
|
||
- **AND** Prompt text, model reasoning, complete tool output, and complete State snapshots SHALL NOT be control state
|
||
|
||
#### Scenario: Orchestration event is designed
|
||
|
||
- **WHEN** a node attempt reaches a handled terminal outcome
|
||
- **THEN** at most one event SHALL be appended for that attempt
|
||
- **AND** it SHALL contain only stable node, outcome, reason code, and attempt data
|
||
|
||
### Requirement: Routing design SHALL be complete and terminating
|
||
|
||
Every Planner, Executor, Gatekeeper, Verifier, Composer, evidence-retry, and Fallback outcome SHALL map to one next node or terminal result. Every loop SHALL have an explicit business limit and the Graph SHALL use a recursion limit.
|
||
|
||
#### Scenario: Technical retry is eligible
|
||
|
||
- **WHEN** Planner, Verifier, or Composer first returns INVALID_OUTPUT or RETRYABLE_FAILED with the same allowed input
|
||
- **THEN** only that node SHALL be retried once
|
||
- **AND** no preceding Agent, Gatekeeper, or tool SHALL be rerun
|
||
|
||
#### Scenario: Technical retry is exhausted
|
||
|
||
- **WHEN** Planner, Verifier, or Composer returns NON_RETRYABLE_FAILED or exhausts its retry
|
||
- **THEN** the route SHALL terminate through the defined safe Fallback
|
||
- **AND** no unbounded loop SHALL remain
|
||
|
||
#### Scenario: Executor cannot produce a legal contract
|
||
|
||
- **WHEN** Executor returns INVALID_OUTPUT, TOOL_BLOCKED, or FAILED without legal `executor_evidence_v2`
|
||
- **THEN** the route SHALL go directly to pre-verification Fallback
|
||
- **AND** Executor SHALL NOT be retried
|
||
|
||
#### Scenario: Evidence retry is eligible
|
||
|
||
- **WHEN** effective verdict is LOW_CONFID, it was not caused by Gatekeeper ceiling, valid evidence gaps exist, and the Run has not retried evidence
|
||
- **THEN** one EVIDENCE_GAP_ONLY retry SHALL return to a new Planner stage
|
||
- **AND** technical retry counters SHALL remain independent from evidence retry count
|
||
|
||
### Requirement: Verified evidence and fallback boundaries SHALL fail closed
|
||
|
||
PASS and eligible LOW_CONFID Gatekeeper outcomes SHALL pass through a verified-input builder. Verifier SHALL receive only claims and evidence projected from passed checked bindings, not complete `tool_trace_summary` or unverified Executor text.
|
||
|
||
#### Scenario: Gatekeeper result is unsafe or unknown
|
||
|
||
- **WHEN** Gatekeeper returns REJECT, unknown, or LOW_CONFID with zero verified bindings
|
||
- **THEN** the route SHALL enter pre-verification Fallback without Verifier
|
||
- **AND** the final answer SHALL NOT contain any Executor claim
|
||
|
||
#### Scenario: Gatekeeper permits verification
|
||
|
||
- **WHEN** Gatekeeper returns PASS or eligible LOW_CONFID
|
||
- **THEN** the builder SHALL project evidence only from passed bindings
|
||
- **AND** a LOW_CONFID ceiling SHALL NOT be upgraded to PASS
|
||
|
||
#### Scenario: Composer fails after verification
|
||
|
||
- **WHEN** Composer exhausts retry after receiving Verifier-allowed material
|
||
- **THEN** deterministic fallback MAY use only that allowed material
|
||
- **AND** it SHALL NOT read raw Executor or tool output
|
||
|
||
### Requirement: Orchestration audit design SHALL preserve Run ownership
|
||
|
||
Orchestration audit SHALL remain separate from self evaluation, AgentStep, ToolInvocation, and Graph checkpoint data. A compact summary SHALL be derived from bounded events and persisted only to the current Run.
|
||
|
||
#### Scenario: A safe response is produced
|
||
|
||
- **WHEN** Graph reaches Composer or handled Fallback and produces a safe answer
|
||
- **THEN** Run status SHALL be SUCCESS
|
||
- **AND** degradation SHALL be represented by verdict or `orchestrationTrace.degraded`
|
||
|
||
#### Scenario: Trace is exposed
|
||
|
||
- **WHEN** an exact new StateGraph Chat Run is queried
|
||
- **THEN** parsed audit SHALL appear only at `run.orchestrationTrace`
|
||
- **AND** it SHALL NOT be duplicated at top level, session projection, or raw field
|
||
|
||
#### Scenario: Audit ownership is evaluated
|
||
|
||
- **WHEN** events or summaries are persisted
|
||
- **THEN** `runId` SHALL be the ownership and Graph thread boundary
|
||
- **AND** no data SHALL include Prompt, reasoning, complete tool output, or another Run
|
||
|
||
### Requirement: Test migration design SHALL preserve safety behavior
|
||
|
||
The baseline SHALL identify Sequential/Hook implementation tests to replace and public/security contract tests to retain or extend. Fixed Agent call order SHALL NOT remain a correctness criterion.
|
||
|
||
#### Scenario: Old tests are replaced
|
||
|
||
- **WHEN** StateGraph tests become authoritative
|
||
- **THEN** `ChatServiceSequentialAgentTest` SHALL be replaced by route, node-contract, and Chat integration coverage
|
||
- **AND** `VerifierInputHookTest` SHALL be removed or rewritten for explicit nodes
|
||
|
||
#### Scenario: Safety tests are retained
|
||
|
||
- **WHEN** the new suite is assembled
|
||
- **THEN** Gatekeeper, Controller, Run/Trace, Repository, Composer, no-evidence, REJECT, and Eval safety contracts SHALL remain covered
|
||
- **AND** fixed-order-only assertions SHALL be removed
|