Files
SuperBizAgent-java/openspec/specs/diagnosis-harness-run-context/spec.md
T

94 lines
7.3 KiB
Markdown

# diagnosis-harness-run-context Specification
## Purpose
定义 Diagnosis Harness 的显式 RunContext、deadline、协作式取消、线程安全预算、唯一 Run 终态、类型化重试、Tool Call Key 和单 Run 容量基础,供后续 Tool boundary、Agent、Guard 和应用入口复用。
## Requirements
### Requirement: RunContext SHALL be explicit and structurally immutable
The Harness SHALL create a structurally immutable RunContext containing non-blank `sessionId`, non-blank `runId`, deadline, cancellation, budget, retry policies, and lifecycle. Every new Model, Tool, Guard, and application boundary SHALL receive the RunContext explicitly by parameter or framework-controlled context and SHALL NOT read or write ThreadLocal state.
#### Scenario: Context crosses an asynchronous boundary
- **WHEN** a Fake Tool runs on another thread with an explicitly supplied RunContext
- **THEN** it observes the same session ID, run ID, deadline, cancellation, budget, retry policies, and lifecycle handles without copying ThreadLocal state
#### Scenario: Invalid identity starts a Run
- **WHEN** Run creation receives a blank session ID or run ID
- **THEN** the Core rejects creation before any budget, cancellation, or lifecycle state is published
### Requirement: Deadline and cancellation SHALL stop subsequent controlled work
The Core SHALL compare the current Clock instant with the Run deadline before each controlled Model, Tool, retry, and capacity boundary. At or after the deadline it SHALL produce `TIMED_OUT`, signal `DEADLINE_EXCEEDED`, and reject subsequent work. Explicit cancellation SHALL preserve the first cancellation reason, notify all registered resource callbacks, and produce `CANCELLED` unless an earlier terminal state exists.
#### Scenario: Deadline expires before a Tool call
- **WHEN** the controlled Clock reaches the Run deadline before the Fake Tool boundary
- **THEN** the Tool is not invoked, the Run terminates as `TIMED_OUT`, and later checks return the same terminal outcome
#### Scenario: Client disconnects
- **WHEN** the Core receives `CLIENT_DISCONNECTED`
- **THEN** registered cancellation callbacks run, subsequent Model/Tool boundaries are rejected, and the Run has one `CANCELLED` terminal outcome
### Requirement: Run budgets SHALL be explicit, thread-safe, and observable
The Core SHALL require caller-supplied positive limits for Model calls, total Tool calls, per-Tool calls, input Tokens, output Tokens, total Tokens, and Run bytes. Budget updates SHALL be thread-safe and preserve a usage snapshot. Exceeding any limit SHALL record the actual applicable usage, produce `BUDGET_EXHAUSTED`, signal cancellation, and prevent subsequent controlled work.
#### Scenario: Per-Tool budget is exhausted
- **WHEN** a Fake Tool attempts one call beyond its configured per-Tool limit
- **THEN** the extra call is not reserved, the total and per-Tool usage remain internally consistent, and the Run terminates as `BUDGET_EXHAUSTED`
#### Scenario: Actual Token usage exceeds the limit
- **WHEN** a completed Fake Model response reports Token usage above the configured limit
- **THEN** the actual input/output/total usage remains recorded and the Core rejects the next controlled operation
### Requirement: A Run SHALL have exactly one terminal outcome
The lifecycle SHALL start as `RUNNING` and accept only the first terminal outcome from `SUCCESS`, `FAILED`, `CANCELLED`, `TIMED_OUT`, or `BUDGET_EXHAUSTED`. Later success, failure, cancellation, deadline, or budget events SHALL NOT replace the first `RunTermination` state, reason, or timestamp.
#### Scenario: Success wins a terminal race
- **WHEN** success is recorded before a later failure and cancellation
- **THEN** the lifecycle remains `SUCCESS` and both later finish attempts report that they did not change the terminal outcome
#### Scenario: Budget exhaustion wins a terminal race
- **WHEN** budget exhaustion is recorded before an exception handler reports failure
- **THEN** the lifecycle remains `BUDGET_EXHAUSTED` with its original reason and completion time
### Requirement: Harness retry SHALL be typed, bounded, and recorded
The Harness SHALL provide immutable policies and an executor that records every actual attempt. Intent Router and SemanticGuard SHALL allow at most two attempts for their listed technical failures; Diagnosis Agent, Tool calls, and Evidence repair SHALL allow one attempt. Cancellation, deadline, budget exhaustion, `NO_EVIDENCE`, business rejection, and unknown failures SHALL NOT be retried.
#### Scenario: Router transport failure succeeds on retry
- **WHEN** a Fake Router fails its first attempt with a retryable transport failure and succeeds on the second
- **THEN** the executor returns the second result and records exactly one failed and one successful attempt
#### Scenario: Tool call fails
- **WHEN** a Fake Tool operation fails on its first attempt
- **THEN** the Tool policy records one failed attempt and throws without invoking the Tool again
#### Scenario: Run is cancelled between attempts
- **WHEN** cancellation is signalled after a retryable first failure
- **THEN** the executor rejects the next attempt and does not invoke the operation again
### Requirement: Spring AI hidden retry SHALL be disabled
Application configuration SHALL set `spring.ai.retry.max-attempts=1`, overriding the Spring AI 1.1.7 default of 10. Harness-permitted retries SHALL occur only in HarnessRetryExecutor and SHALL be observable as separate attempts.
#### Scenario: Retry configuration is loaded
- **WHEN** the repository application YAML is parsed in a focused configuration test
- **THEN** `spring.ai.retry.max-attempts` equals 1
### Requirement: Tool invocation key and Run capacity primitives SHALL be safe and store-independent
The Harness SHALL provide a ToolCallKeyFactory that constructs `{prefix}:{runId}:{toolCallId}` only from valid safe segments and never generates or rewrites the framework Tool Call ID. The Run capacity counter SHALL atomically reserve positive bytes up to its configured limit and SHALL leave usage unchanged when a reservation is rejected.
#### Scenario: Framework Tool Call ID is used in a key
- **WHEN** a valid run ID and framework Tool Call ID are supplied
- **THEN** the factory returns the configured prefix followed by the exact run ID and exact Tool Call ID
#### Scenario: Unsafe key segment is supplied
- **WHEN** either ID is blank, too long, or contains a separator or unsafe character
- **THEN** the factory rejects it and does not produce a Redis key
#### Scenario: Run byte limit would be exceeded
- **WHEN** a reservation would exceed max Run bytes
- **THEN** the counter rejects it and retains the usage from previously successful reservations
### Requirement: Harness Core SHALL remain independent from current runtime and persistence
This change SHALL NOT modify Controller/HTTP/SSE contracts, connect RunContext to the old ChatService or AiOpsService, use current ThreadLocal holders, persist `diagnosis_run`, access Redis, or implement Tool-specific projection. The new Core SHALL compile and pass focused Fake Model/Tool tests as a zero-consumer foundation for stage 3A.
#### Scenario: Stage 2 implementation completes
- **WHEN** focused Core, retry, budget, key, capacity, and configuration tests pass
- **THEN** current Chat/AIOps call sites remain unchanged and stage 3A can depend directly on the new Harness types