6.0 KiB
ADDED Requirements
Requirement: StateGraph design baseline SHALL be versioned and authoritative
The project SHALL maintain an archived ISS-011 design baseline that defines Graph State, routing, retry limits, fallback classes, audit boundaries, migration stages, and test replacement scope without claiming runtime cutover is complete.
Scenario: Later stage starts implementation
- WHEN an ISS-011 stage 1–5 OpenSpec change is proposed
- THEN its context and design SHALL reference the archived stage 0 baseline
- AND any intentional deviation SHALL be resolved through that stage's OpenSpec before implementation
Scenario: Stage zero is accepted
- WHEN the design-freeze change is archived
- THEN no Java, SQL, Prompt, configuration, or runtime behavior SHALL have been changed by this change
- AND runtime specs SHALL NOT claim StateGraph cutover is already implemented
Requirement: Graph State design SHALL use explicit bounded control state
Every cross-node field SHALL have one purpose, owner, and update strategy. Fields SHALL use Replace semantics except bounded orchestration_events, which SHALL use Append.
Scenario: Node state is designed
- WHEN an Agent or Java Node reads or writes parent Graph State
- THEN its allowed input projection and output fields SHALL be explicit
- AND Prompt text, model reasoning, complete tool output, and complete State snapshots SHALL NOT be control state
Scenario: Orchestration event is designed
- WHEN a node attempt reaches a handled terminal outcome
- THEN at most one event SHALL be appended for that attempt
- AND it SHALL contain only stable node, outcome, reason code, and attempt data
Requirement: Routing design SHALL be complete and terminating
Every Planner, Executor, Gatekeeper, Verifier, Composer, evidence-retry, and Fallback outcome SHALL map to one next node or terminal result. Every loop SHALL have an explicit business limit and the Graph SHALL use a recursion limit.
Scenario: Technical retry is eligible
- WHEN Planner, Verifier, or Composer first returns INVALID_OUTPUT or RETRYABLE_FAILED with the same allowed input
- THEN only that node SHALL be retried once
- AND no preceding Agent, Gatekeeper, or tool SHALL be rerun
Scenario: Technical retry is exhausted
- WHEN Planner, Verifier, or Composer returns NON_RETRYABLE_FAILED or exhausts its retry
- THEN the route SHALL terminate through the defined safe Fallback
- AND no unbounded loop SHALL remain
Scenario: Executor cannot produce a legal contract
- WHEN Executor returns INVALID_OUTPUT, TOOL_BLOCKED, or FAILED without legal
executor_evidence_v2 - THEN the route SHALL go directly to pre-verification Fallback
- AND Executor SHALL NOT be retried
Scenario: Evidence retry is eligible
- WHEN effective verdict is LOW_CONFID, it was not caused by Gatekeeper ceiling, valid evidence gaps exist, and the Run has not retried evidence
- THEN one EVIDENCE_GAP_ONLY retry SHALL return to a new Planner stage
- AND technical retry counters SHALL remain independent from evidence retry count
Requirement: Verified evidence and fallback boundaries SHALL fail closed
PASS and eligible LOW_CONFID Gatekeeper outcomes SHALL pass through a verified-input builder. Verifier SHALL receive only claims and evidence projected from passed checked bindings, not complete tool_trace_summary or unverified Executor text.
Scenario: Gatekeeper result is unsafe or unknown
- WHEN Gatekeeper returns REJECT, unknown, or LOW_CONFID with zero verified bindings
- THEN the route SHALL enter pre-verification Fallback without Verifier
- AND the final answer SHALL NOT contain any Executor claim
Scenario: Gatekeeper permits verification
- WHEN Gatekeeper returns PASS or eligible LOW_CONFID
- THEN the builder SHALL project evidence only from passed bindings
- AND a LOW_CONFID ceiling SHALL NOT be upgraded to PASS
Scenario: Composer fails after verification
- WHEN Composer exhausts retry after receiving Verifier-allowed material
- THEN deterministic fallback MAY use only that allowed material
- AND it SHALL NOT read raw Executor or tool output
Requirement: Orchestration audit design SHALL preserve Run ownership
Orchestration audit SHALL remain separate from self evaluation, AgentStep, ToolInvocation, and Graph checkpoint data. A compact summary SHALL be derived from bounded events and persisted only to the current Run.
Scenario: A safe response is produced
- WHEN Graph reaches Composer or handled Fallback and produces a safe answer
- THEN Run status SHALL be SUCCESS
- AND degradation SHALL be represented by verdict or
orchestrationTrace.degraded
Scenario: Trace is exposed
- WHEN an exact new StateGraph Chat Run is queried
- THEN parsed audit SHALL appear only at
run.orchestrationTrace - AND it SHALL NOT be duplicated at top level, session projection, or raw field
Scenario: Audit ownership is evaluated
- WHEN events or summaries are persisted
- THEN
runIdSHALL be the ownership and Graph thread boundary - AND no data SHALL include Prompt, reasoning, complete tool output, or another Run
Requirement: Test migration design SHALL preserve safety behavior
The baseline SHALL identify Sequential/Hook implementation tests to replace and public/security contract tests to retain or extend. Fixed Agent call order SHALL NOT remain a correctness criterion.
Scenario: Old tests are replaced
- WHEN StateGraph tests become authoritative
- THEN
ChatServiceSequentialAgentTestSHALL be replaced by route, node-contract, and Chat integration coverage - AND
VerifierInputHookTestSHALL be removed or rewritten for explicit nodes
Scenario: Safety tests are retained
- WHEN the new suite is assembled
- THEN Gatekeeper, Controller, Run/Trace, Repository, Composer, no-evidence, REJECT, and Eval safety contracts SHALL remain covered
- AND fixed-order-only assertions SHALL be removed