Files
SuperBizAgent-java/openspec/specs/chat-diagnosis-stategraph-design-freeze/spec.md
T

133 lines
7.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# chat-diagnosis-stategraph-design-freeze Specification
## Purpose
为 ISS-011 阶段 1–5 提供已归档、可审计的 StateGraph 设计基线,规范跨节点状态、条件边、有限重试、安全降级、Run 级编排审计和测试迁移边界,同时明确阶段 0 不代表运行时已经切换。
## Requirements
### Requirement: StateGraph design baseline SHALL be versioned and authoritative
The project SHALL maintain an archived ISS-011 design baseline that defines Graph State, routing, retry limits, fallback classes, audit boundaries, migration stages, and test replacement scope without claiming runtime cutover is complete.
#### Scenario: Later stage starts implementation
- **WHEN** an ISS-011 stage 1–5 OpenSpec change is proposed
- **THEN** its context and design SHALL reference the archived stage 0 baseline
- **AND** any intentional deviation SHALL be resolved through that stage's OpenSpec before implementation
#### Scenario: Stage zero is accepted
- **WHEN** the design-freeze change is archived
- **THEN** no Java, SQL, Prompt, configuration, or runtime behavior SHALL have been changed by this change
- **AND** runtime specs SHALL NOT claim StateGraph cutover is already implemented
### Requirement: Graph State design SHALL use explicit bounded control state
Every cross-node field SHALL have one purpose, owner, and update strategy. Fields SHALL use Replace semantics except bounded `orchestration_events`, which SHALL use Append.
#### Scenario: Node state is designed
- **WHEN** an Agent or Java Node reads or writes parent Graph State
- **THEN** its allowed input projection and output fields SHALL be explicit
- **AND** Prompt text, model reasoning, complete tool output, and complete State snapshots SHALL NOT be control state
#### Scenario: Orchestration event is designed
- **WHEN** a node attempt reaches a handled terminal outcome
- **THEN** at most one event SHALL be appended for that attempt
- **AND** it SHALL contain only stable node, outcome, reason code, and attempt data
### Requirement: Routing design SHALL be complete and terminating
Every Planner, Executor, Gatekeeper, Verifier, Composer, evidence-retry, and Fallback outcome SHALL map to one next node or terminal result. Every loop SHALL have an explicit business limit and the Graph SHALL use a recursion limit.
#### Scenario: Technical retry is eligible
- **WHEN** Planner, Verifier, or Composer first returns INVALID_OUTPUT or RETRYABLE_FAILED with the same allowed input
- **THEN** only that node SHALL be retried once
- **AND** no preceding Agent, Gatekeeper, or tool SHALL be rerun
#### Scenario: Technical retry is exhausted
- **WHEN** Planner, Verifier, or Composer returns NON_RETRYABLE_FAILED or exhausts its retry
- **THEN** the route SHALL terminate through the defined safe Fallback
- **AND** no unbounded loop SHALL remain
#### Scenario: Executor cannot produce a legal contract
- **WHEN** Executor returns INVALID_OUTPUT, TOOL_BLOCKED, or FAILED without legal `executor_evidence_v2`
- **THEN** the route SHALL go directly to pre-verification Fallback
- **AND** Executor SHALL NOT be retried
#### Scenario: Evidence retry is eligible
- **WHEN** effective verdict is LOW_CONFID, it was not caused by Gatekeeper ceiling, valid evidence gaps exist, and the Run has not retried evidence
- **THEN** one EVIDENCE_GAP_ONLY retry SHALL return to a new Planner stage
- **AND** technical retry counters SHALL remain independent from evidence retry count
### Requirement: Verified evidence and fallback boundaries SHALL fail closed
PASS and eligible LOW_CONFID Gatekeeper outcomes SHALL pass through a verified-input builder. Verifier SHALL receive only claims and evidence projected from passed checked bindings, not complete `tool_trace_summary` or unverified Executor text.
#### Scenario: Gatekeeper result is unsafe or unknown
- **WHEN** Gatekeeper returns REJECT, unknown, or LOW_CONFID with zero verified bindings
- **THEN** the route SHALL enter pre-verification Fallback without Verifier
- **AND** the final answer SHALL NOT contain any Executor claim
#### Scenario: Gatekeeper permits verification
- **WHEN** Gatekeeper returns PASS or eligible LOW_CONFID
- **THEN** the builder SHALL project evidence only from passed bindings
- **AND** a LOW_CONFID ceiling SHALL NOT be upgraded to PASS
#### Scenario: Composer fails after verification
- **WHEN** Composer exhausts retry after receiving Verifier-allowed material
- **THEN** deterministic fallback MAY use only that allowed material
- **AND** it SHALL NOT read raw Executor or tool output
### Requirement: Orchestration audit design SHALL preserve Run ownership
Orchestration audit SHALL remain separate from self evaluation, AgentStep, ToolInvocation, and Graph checkpoint data. A compact summary SHALL be derived from bounded events and persisted only to the current Run.
#### Scenario: A safe response is produced
- **WHEN** Graph reaches Composer or handled Fallback and produces a safe answer
- **THEN** Run status SHALL be SUCCESS
- **AND** degradation SHALL be represented by verdict or `orchestrationTrace.degraded`
#### Scenario: Trace is exposed
- **WHEN** an exact new StateGraph Chat Run is queried
- **THEN** parsed audit SHALL appear only at `run.orchestrationTrace`
- **AND** it SHALL NOT be duplicated at top level, session projection, or raw field
#### Scenario: Audit ownership is evaluated
- **WHEN** events or summaries are persisted
- **THEN** `runId` SHALL be the ownership and Graph thread boundary
- **AND** no data SHALL include Prompt, reasoning, complete tool output, or another Run
### Requirement: Test migration design SHALL preserve safety behavior
The project SHALL replace Sequential/Hook implementation tests with authoritative `DiagnosisGraphWorkflowTest`, `DiagnosisGraphNodeContractTest`, and `ChatServiceGraphIntegrationTest` layers while retaining focused public/security contract tests. Fixed Agent call order and legacy Hook payload shape SHALL NOT remain correctness criteria.
#### Scenario: Old tests are replaced
- **WHEN** StateGraph tests become authoritative
- **THEN** `ChatServiceSequentialAgentTest` SHALL NOT exist
- **AND** `VerifierInputHookTest` SHALL NOT exist after explicit Gatekeeper/Verified Input coverage is established
- **AND** no replacement test SHALL restore fixed Sequential Agent order as a public requirement
#### Scenario: Safety tests are retained
- **WHEN** the new suite is assembled
- **THEN** Gatekeeper, Controller, Run/Trace, Repository, Composer, parser/projection, no-evidence, REJECT, Eval, and multi-run safety contracts SHALL remain covered
- **AND** Workflow and Node Contract matrices SHALL cover bounded retry, fail-closed routing, verified-only inputs, and deterministic Fallback
#### Scenario: Authoritative test layers are inspected
- **WHEN** a maintainer needs to locate ISS-011 orchestration verification
- **THEN** Workflow tests SHALL own route/loop/event behavior
- **AND** Node Contract tests SHALL own input/output/security projection behavior
- **AND** Chat integration tests SHALL own Run lifecycle and public result behavior