Files

7.3 KiB

diagnosis-harness-run-context Specification

Purpose

定义 Diagnosis Harness 的显式 RunContext、deadline、协作式取消、线程安全预算、唯一 Run 终态、类型化重试、Tool Call Key 和单 Run 容量基础,供后续 Tool boundary、Agent、Guard 和应用入口复用。

Requirements

Requirement: RunContext SHALL be explicit and structurally immutable

The Harness SHALL create a structurally immutable RunContext containing non-blank sessionId, non-blank runId, deadline, cancellation, budget, retry policies, and lifecycle. Every new Model, Tool, Guard, and application boundary SHALL receive the RunContext explicitly by parameter or framework-controlled context and SHALL NOT read or write ThreadLocal state.

Scenario: Context crosses an asynchronous boundary

  • WHEN a Fake Tool runs on another thread with an explicitly supplied RunContext
  • THEN it observes the same session ID, run ID, deadline, cancellation, budget, retry policies, and lifecycle handles without copying ThreadLocal state

Scenario: Invalid identity starts a Run

  • WHEN Run creation receives a blank session ID or run ID
  • THEN the Core rejects creation before any budget, cancellation, or lifecycle state is published

Requirement: Deadline and cancellation SHALL stop subsequent controlled work

The Core SHALL compare the current Clock instant with the Run deadline before each controlled Model, Tool, retry, and capacity boundary. At or after the deadline it SHALL produce TIMED_OUT, signal DEADLINE_EXCEEDED, and reject subsequent work. Explicit cancellation SHALL preserve the first cancellation reason, notify all registered resource callbacks, and produce CANCELLED unless an earlier terminal state exists.

Scenario: Deadline expires before a Tool call

  • WHEN the controlled Clock reaches the Run deadline before the Fake Tool boundary
  • THEN the Tool is not invoked, the Run terminates as TIMED_OUT, and later checks return the same terminal outcome

Scenario: Client disconnects

  • WHEN the Core receives CLIENT_DISCONNECTED
  • THEN registered cancellation callbacks run, subsequent Model/Tool boundaries are rejected, and the Run has one CANCELLED terminal outcome

Requirement: Run budgets SHALL be explicit, thread-safe, and observable

The Core SHALL require caller-supplied positive limits for Model calls, total Tool calls, per-Tool calls, input Tokens, output Tokens, total Tokens, and Run bytes. Budget updates SHALL be thread-safe and preserve a usage snapshot. Exceeding any limit SHALL record the actual applicable usage, produce BUDGET_EXHAUSTED, signal cancellation, and prevent subsequent controlled work.

Scenario: Per-Tool budget is exhausted

  • WHEN a Fake Tool attempts one call beyond its configured per-Tool limit
  • THEN the extra call is not reserved, the total and per-Tool usage remain internally consistent, and the Run terminates as BUDGET_EXHAUSTED

Scenario: Actual Token usage exceeds the limit

  • WHEN a completed Fake Model response reports Token usage above the configured limit
  • THEN the actual input/output/total usage remains recorded and the Core rejects the next controlled operation

Requirement: A Run SHALL have exactly one terminal outcome

The lifecycle SHALL start as RUNNING and accept only the first terminal outcome from SUCCESS, FAILED, CANCELLED, TIMED_OUT, or BUDGET_EXHAUSTED. Later success, failure, cancellation, deadline, or budget events SHALL NOT replace the first RunTermination state, reason, or timestamp.

Scenario: Success wins a terminal race

  • WHEN success is recorded before a later failure and cancellation
  • THEN the lifecycle remains SUCCESS and both later finish attempts report that they did not change the terminal outcome

Scenario: Budget exhaustion wins a terminal race

  • WHEN budget exhaustion is recorded before an exception handler reports failure
  • THEN the lifecycle remains BUDGET_EXHAUSTED with its original reason and completion time

Requirement: Harness retry SHALL be typed, bounded, and recorded

The Harness SHALL provide immutable policies and an executor that records every actual attempt. Intent Router and SemanticGuard SHALL allow at most two attempts for their listed technical failures; Diagnosis Agent, Tool calls, and Evidence repair SHALL allow one attempt. Cancellation, deadline, budget exhaustion, NO_EVIDENCE, business rejection, and unknown failures SHALL NOT be retried.

Scenario: Router transport failure succeeds on retry

  • WHEN a Fake Router fails its first attempt with a retryable transport failure and succeeds on the second
  • THEN the executor returns the second result and records exactly one failed and one successful attempt

Scenario: Tool call fails

  • WHEN a Fake Tool operation fails on its first attempt
  • THEN the Tool policy records one failed attempt and throws without invoking the Tool again

Scenario: Run is cancelled between attempts

  • WHEN cancellation is signalled after a retryable first failure
  • THEN the executor rejects the next attempt and does not invoke the operation again

Requirement: Spring AI hidden retry SHALL be disabled

Application configuration SHALL set spring.ai.retry.max-attempts=1, overriding the Spring AI 1.1.7 default of 10. Harness-permitted retries SHALL occur only in HarnessRetryExecutor and SHALL be observable as separate attempts.

Scenario: Retry configuration is loaded

  • WHEN the repository application YAML is parsed in a focused configuration test
  • THEN spring.ai.retry.max-attempts equals 1

Requirement: Tool invocation key and Run capacity primitives SHALL be safe and store-independent

The Harness SHALL provide a ToolCallKeyFactory that constructs {prefix}:{runId}:{toolCallId} only from valid safe segments and never generates or rewrites the framework Tool Call ID. The Run capacity counter SHALL atomically reserve positive bytes up to its configured limit and SHALL leave usage unchanged when a reservation is rejected.

Scenario: Framework Tool Call ID is used in a key

  • WHEN a valid run ID and framework Tool Call ID are supplied
  • THEN the factory returns the configured prefix followed by the exact run ID and exact Tool Call ID

Scenario: Unsafe key segment is supplied

  • WHEN either ID is blank, too long, or contains a separator or unsafe character
  • THEN the factory rejects it and does not produce a Redis key

Scenario: Run byte limit would be exceeded

  • WHEN a reservation would exceed max Run bytes
  • THEN the counter rejects it and retains the usage from previously successful reservations

Requirement: Harness Core SHALL remain independent from current runtime and persistence

This change SHALL NOT modify Controller/HTTP/SSE contracts, connect RunContext to the old ChatService or AiOpsService, use current ThreadLocal holders, persist diagnosis_run, access Redis, or implement Tool-specific projection. The new Core SHALL compile and pass focused Fake Model/Tool tests as a zero-consumer foundation for stage 3A.

Scenario: Stage 2 implementation completes

  • WHEN focused Core, retry, budget, key, capacity, and configuration tests pass
  • THEN current Chat/AIOps call sites remain unchanged and stage 3A can depend directly on the new Harness types