feat(harness): add run context and retry core

This commit is contained in:
zhuyongxin
2026-07-21 18:36:19 +08:00
parent 4274f3350b
commit 6b74990f86
47 changed files with 1910 additions and 1 deletions
@@ -0,0 +1 @@
Archive-ready after implementation and focused verification on 2026-07-21.
@@ -0,0 +1 @@
Committed after strict validation on 2026-07-21.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-21
@@ -0,0 +1,109 @@
## Context
现有 Chat/AIOps 在进入 Agent 前自行生成 session/run 标识,通过多个 ThreadLocal 和 RunnableConfig metadata 混合传播,并在业务方法内直接处理 retry loop、运行持久化和成功/失败终态。`TokenTrackingChatModel` 也只把 total token 写到 ThreadLocal,无法在异步或并发边界可靠归属 Run。后续 Tool interceptor、ResultProjector、Diagnosis Agent、EvidenceGuard、SemanticGuard 和 Chat application use case 需要共同使用一个不依赖旧 ChatService 的显式执行上下文。
Spring AI 1.1.7 的全局 retry properties 前缀为 `spring.ai.retry`,默认 `maxAttempts=10`。ISS-014 已确认所有底层模型 attempt 必须压为一次,由 Harness 对 Router/SemanticGuard 的特定技术失败显式执行最多一次重试并记录每次 attempt。
## Goals / Non-Goals
**Goals:**
- 提供显式、结构不可变、可跨同步/异步边界传递的 RunContext。
- 在内存执行层统一 deadline、取消、预算和唯一终态语义。
- 提供无隐藏循环的类型化 retry policy/executor 和 attempt 记录。
- 提供阶段 3A 可直接复用的 Tool Call Key Factory 与单 Run 字节容量计数器。
- 关闭 Spring AI 默认隐藏重试,保证实际模型 attempt 可由 Harness 观察。
**Non-Goals:**
- 不接入 Controller、SSE、旧 ChatService、AiOpsService、ReactAgent 或现有 Tool。
- 不替换或删除旧 ThreadLocal;阶段 7 在新入口切换后清理。
- 不修改 `diagnosis_run` 实体、表结构或 Repository,不写数据库终态。
- 不实现 Redis invocation store、ToolInterceptor、ToolResultProjector 或 Tool Call ID 校验全套策略。
- 不实现定时调度器、线程中断、工作流引擎、动态配置中心、Retry DSL 或退避算法。
## Decisions
### 1. RunContext 是结构不可变的共享状态句柄集合
`RunContext` 使用 Java 17 record,字段固定为 `sessionId`、`runId`、`deadline`、`RunCancellation`、`RunBudget`、`HarnessRetryPolicies` 和 `RunLifecycle`。record 不提供 setter;同一 Run 的同步/异步消费者必须显式传递同一个 context,因此共享相同的取消、预算和生命周期句柄。
替代方案是把所有字段做成纯值并在每次变化时复制 context;这会造成并发分支状态分叉,无法保证唯一终态和原子预算,故不采用。ThreadLocal/RunnableConfig fallback 也不进入新 Core。
### 2. DiagnosisHarnessCore 是唯一执行门禁与状态转换入口
Core 由 `Clock`、Run ID supplier、最大 Run duration、显式 `RunBudgetLimits` 和 `HarnessRetryPolicies` 构造。它负责创建 context、在每个模型/Tool/容量边界检查 active/deadline、预留预算、记录 Token、显式取消以及完成成功/失败终态。
Core 不保存全局 Run map;生命周期完全随 RunContext 所有,避免跨请求泄漏。调用方可以传入既有 runId 进行确定性测试/恢复,但 Core 不生成 sessionId,也不修改业务持久化。
### 3. Deadline 与取消采用协作式、可观察语义
`RunCancellation` 用 atomic first-reason-wins 保存 `CLIENT_DISCONNECTED`、`USER_REQUESTED`、`DEADLINE_EXCEEDED`、`BUDGET_EXHAUSTED` 或 `INTERNAL_FAILURE`,并允许注册资源取消回调。Core 在每个受控边界比较 `Clock.instant()` 与 deadline;到达或超过 deadline 时先写 `TIMED_OUT` 终态,再触发取消信号并拒绝后续执行。
取消回调异常会记录并继续通知其他回调,不能阻止取消传播。此模型不承诺已进入的同步第三方调用立即停止;具体 HTTP/JDBC future 取消由后续 adapter 注册回调实现。
### 4. RunLifecycle 使用 first-terminal-wins
`RunLifecycle` 初始为 `RUNNING`,终态固定为 `SUCCESS`、`FAILED`、`CANCELLED`、`TIMED_OUT`、`BUDGET_EXHAUSTED`。内部使用 atomic compare-and-set 保存唯一 `RunTermination(state, reason, completedAt)`;任何后续 finish 都返回 false 且不得覆盖首个终态。
这只是执行内存真理;阶段 6A 的应用用例负责把终态映射到 `diagnosis_run.status/release_outcome`。本阶段不创建第二套数据库状态机。
### 5. RunBudget 使用显式 limits 和一致性更新
`RunBudgetLimits` 必须由调用方显式提供,包含模型调用、总 Tool 调用、单 Tool 调用、input/output/total Token 和 max Run bytes,所有值必须为正。`RunBudget` 使用同步临界区保证总 Tool/单 Tool 计数和三类 Token 计数要么一致提交,要么明确抛出 `BudgetExceededException`;模型返回后的实际 Token 即使超限也会先记入 usage,再终止后续执行。
`RunCapacityCounter` 用 CAS 原子预留 UTF-8 JSON 字节,失败时不部分增长。它被 RunBudget 持有并可由阶段 3A 直接调用。Core 在预算失败时先写 `BUDGET_EXHAUSTED` 终态,再触发取消。
### 6. Retry policy 只描述 attempt 上限与类型
`RetryPolicy(maxAttempts, retryableFailures)` 只允许 1 或 2 attempts。`HarnessRetryPolicies.strict()` 固定:Router=2(timeout/transport/invalid output)、Diagnosis=1、Tool=1、SemanticGuard=2(timeout/transport/parse/schema)、EvidenceRepair=1。
`HarnessRetryExecutor` 在每次 attempt 前调用 Core active check,操作成功/失败都发送 `RetryAttempt` 给 recorder。只有 policy 允许且尚有 attempt 时继续;取消、deadline、预算、`NO_EVIDENCE`、业务拒绝和未知失败不得重试。Executor 不 sleep、不退避、不递归、不调用 Agent。
### 7. 底层 Spring AI retry 固定为一次
`application.yml` 设置 `spring.ai.retry.max-attempts: 1`,并通过直接读取 YAML 的单元测试锁定,避免启动完整外部基础设施。该配置会立即影响旧运行路径:瞬时模型错误不再由 SDK 隐式重试;这是保证 attempt 可观测的有意内部行为变化。
### 8. ToolCallKeyFactory 不拥有 ID
Factory 接收可配置前缀,验证 `runId/toolCallId` 为非空、安全长度和安全字符 segment 后构造 `{prefix}:{runId}:{toolCallId}`。它不生成、规范化或哈希框架 Tool Call ID,不访问 Redis。严格的 duplicate/cross-run invocation 语义属于阶段 3A store。
## Module Flow
```text
future Chat application use case
-> DiagnosisHarnessCore.startRun(...)
-> RunContext
-> RunCancellation
-> RunBudget -> RunCapacityCounter
-> HarnessRetryPolicies
-> RunLifecycle
-> future Model/Tool/Guard boundary explicitly receives RunContext
-> Core check/reserve/record/finish
-> HarnessRetryExecutor for allowed technical operations only
-> future application use case maps the one RunTermination to diagnosis_run
```
阶段 3A 直接依赖 RunContext、Core、Key Factory 和 capacity counter;它不需要调用旧 ChatService 或读取 ThreadLocal。
## Risks / Trade-offs
- [结构不可变但句柄可变容易被误用] -> 命名、Javadoc 和并发/异步测试明确同一 Run 必须共享同一 context 实例。
- [关闭 SDK retry 后旧路径瞬时失败率可能上升] -> 明确列为行为变化,配置测试锁定;后续 Router/SemanticGuard 只按已确认策略补回可观测重试。
- [实际 Token 在响应后才知道,可能超预算] -> usage 保留实际值并立即进入预算终态,不丢失消耗、不再执行下一步。
- [取消回调由 adapter 决定能否硬取消] -> Core 只承诺协作式取消和后续门禁,后续 JDBC/HTTP adapter 必须注册资源回调。
- [纯内存终态无法替代审计] -> 阶段 6A 持久化映射;本阶段 focused tests 只验证执行语义。
## Migration Plan
1. 本阶段新增零消费者 Core 类型和 tests,同时把 Spring AI retry 压为一次。
2. 阶段 3A 的 Tool interceptor/store 显式接收 RunContext,复用 Key/容量门禁。
3. 阶段 4/5 的 Agent/Guard 使用 Core model/retry/deadline 边界。
4. 阶段 6A 的 Chat application use case 创建 RunContext 并映射终态到数据库。
5. 阶段 6B 切换入口;阶段 7 删除 ThreadLocal 和旧重试循环。
回滚本阶段代码可删除新包;`spring.ai.retry.max-attempts` 若回滚到默认 10 会恢复隐藏重试,但会再次失去 attempt 可观测性,因此只允许在整体重构回滚时明确执行。
## Open Questions
无。具体预算数值、Run ID 格式和持久化状态映射由后续 wiring/应用用例在不改变本 Core 语义的前提下配置。
@@ -0,0 +1,46 @@
## Why
当前 Chat/AIOps 通过 `SessionContextHolder`、`VerifierContextHolder` 和 `TokenUsageHolder` 等 ThreadLocal 隐式传播 session/run/token/retry 状态,且 Run 终态、取消、预算和重试分散在业务循环与 SDK 默认行为中。阶段 3A 的 Tool boundary、后续 Agent/Guard 和最终入口需要先共同依赖一个显式、可跨同步/异步边界传递的 Harness RunContext,否则会继续耦合旧 ChatService 并产生重复状态机。
## What Changes
- 新增结构不可变的 `RunContext`,显式携带 `sessionId`、`runId`、deadline、取消信号、线程安全预算、类型化重试策略和 first-terminal-wins 生命周期。
- 新增 `DiagnosisHarnessCore` 负责 Run 创建、deadline 检查、模型/Tool/Token/容量预算门禁、取消传播和唯一终态,不承担业务推理或持久化。
- 新增 caller-supplied `RunBudgetLimits`、线程安全 `RunBudget` 与单 Run 字节容量计数器;Core 不硬编码尚未校准的预算默认值。
- 新增 `HarnessRetryPolicies` 与 `HarnessRetryExecutor`,只允许 Router/SemanticGuard 的类型化技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 固定一次 attempt。
- 将 Spring AI 1.1.7 全局底层重试 `spring.ai.retry.max-attempts` 从默认 10 压为 1,避免与 Harness 形成隐藏嵌套重试。
- 新增 Redis Tool Call Key Factory,只负责安全构造固定前缀下的 `runId + toolCallId` Key,不访问 Redis、不生成 Tool Call ID。
- 使用 Fake Model/Tool 和可控 Clock 验证显式异步传播、deadline、客户端取消、预算耗尽、重试记录、Key/容量边界和唯一 Run 终态。
- 本阶段不改 Controller/HTTP/SSE,不接入或临时适配旧 ChatService,不修改 `diagnosis_run` 持久化,也不实现 Tool-specific 投影或 Redis store。
## Capabilities
### New Capabilities
- `diagnosis-harness-run-context`: 提供显式 RunContext、Harness Core、预算、取消、deadline、类型化重试、唯一终态和后续 Tool store 所需的 Key/容量基础。
### Modified Capabilities
- None. 本阶段不声明旧 Chat/AIOps 运行路径已迁移;公开运行和数据库状态映射在后续 change 接入。
## Context Constraints
- 新代码不得读取或写入任何 ThreadLocal;RunContext 只能通过参数或框架受控 context 显式传播。
- RunContext 的结构不可变,但其 cancellation、budget 和 lifecycle 是线程安全的单 Run 状态句柄。
- 同一 Run 只允许第一个终态生效;取消、deadline、预算和异常之间的竞态不得覆盖先到终态。
- 同步模型/Tool 调用的取消只保证阻止后续执行并触发已注册资源取消回调,不虚假承诺无法证明的立即线程中断。
- `NO_EVIDENCE`、业务拒绝、取消和预算耗尽不是可重试技术失败。
- Tool Call Key 使用框架 `tool_call_id`;Key Factory 不生成、替换或回退到其他 ID。
## Interface Impact
- 等级:L2(内部 Harness 基础接口)。新增 Core 类型将被后续 3A-7 阶段消费,但本 change 不修改现有调用方。
- `application.yml` 的 Spring AI retry 从默认 10 attempts 变为 1 attempt,属于有意的内部运行配置变化;当前旧调用若遇到瞬时模型失败将不再由 SDK 隐式重试,避免未记录 attempts。Harness 允许的 Router/SemanticGuard 重试要到后续接入后显式执行和记录。
- 不改变 Controller、SSE、DTO、数据库 Schema 或公开错误码。
## Risks
- 旧运行路径在切换前会失去 SDK 隐式重试但尚未使用 Harness retry;这是为保证“所有 attempt 可观测”接受的短期行为变化,focused tests 需证明启动配置正确。
- RunContext 内含线程安全可变句柄,若被误解为纯值对象可能错误复制;设计和测试必须明确同一 Run 共享同一状态句柄。
- Budget 在模型调用后才能取得实际 Token,可能在记录实际消耗时才发现超限;Core 必须保留实际 usage 并立即终止后续执行。
- 本阶段不写数据库,内存生命周期只服务一次调用链;持久化终态映射由 Chat application use case 阶段负责。
@@ -0,0 +1,89 @@
## ADDED Requirements
### Requirement: RunContext SHALL be explicit and structurally immutable
The Harness SHALL create a structurally immutable RunContext containing non-blank `sessionId`, non-blank `runId`, deadline, cancellation, budget, retry policies, and lifecycle. Every new Model, Tool, Guard, and application boundary SHALL receive the RunContext explicitly by parameter or framework-controlled context and SHALL NOT read or write ThreadLocal state.
#### Scenario: Context crosses an asynchronous boundary
- **WHEN** a Fake Tool runs on another thread with an explicitly supplied RunContext
- **THEN** it observes the same session ID, run ID, deadline, cancellation, budget, retry policies, and lifecycle handles without copying ThreadLocal state
#### Scenario: Invalid identity starts a Run
- **WHEN** Run creation receives a blank session ID or run ID
- **THEN** the Core rejects creation before any budget, cancellation, or lifecycle state is published
### Requirement: Deadline and cancellation SHALL stop subsequent controlled work
The Core SHALL compare the current Clock instant with the Run deadline before each controlled Model, Tool, retry, and capacity boundary. At or after the deadline it SHALL produce `TIMED_OUT`, signal `DEADLINE_EXCEEDED`, and reject subsequent work. Explicit cancellation SHALL preserve the first cancellation reason, notify all registered resource callbacks, and produce `CANCELLED` unless an earlier terminal state exists.
#### Scenario: Deadline expires before a Tool call
- **WHEN** the controlled Clock reaches the Run deadline before the Fake Tool boundary
- **THEN** the Tool is not invoked, the Run terminates as `TIMED_OUT`, and later checks return the same terminal outcome
#### Scenario: Client disconnects
- **WHEN** the Core receives `CLIENT_DISCONNECTED`
- **THEN** registered cancellation callbacks run, subsequent Model/Tool boundaries are rejected, and the Run has one `CANCELLED` terminal outcome
### Requirement: Run budgets SHALL be explicit, thread-safe, and observable
The Core SHALL require caller-supplied positive limits for Model calls, total Tool calls, per-Tool calls, input Tokens, output Tokens, total Tokens, and Run bytes. Budget updates SHALL be thread-safe and preserve a usage snapshot. Exceeding any limit SHALL record the actual applicable usage, produce `BUDGET_EXHAUSTED`, signal cancellation, and prevent subsequent controlled work.
#### Scenario: Per-Tool budget is exhausted
- **WHEN** a Fake Tool attempts one call beyond its configured per-Tool limit
- **THEN** the extra call is not reserved, the total and per-Tool usage remain internally consistent, and the Run terminates as `BUDGET_EXHAUSTED`
#### Scenario: Actual Token usage exceeds the limit
- **WHEN** a completed Fake Model response reports Token usage above the configured limit
- **THEN** the actual input/output/total usage remains recorded and the Core rejects the next controlled operation
### Requirement: A Run SHALL have exactly one terminal outcome
The lifecycle SHALL start as `RUNNING` and accept only the first terminal outcome from `SUCCESS`, `FAILED`, `CANCELLED`, `TIMED_OUT`, or `BUDGET_EXHAUSTED`. Later success, failure, cancellation, deadline, or budget events SHALL NOT replace the first `RunTermination` state, reason, or timestamp.
#### Scenario: Success wins a terminal race
- **WHEN** success is recorded before a later failure and cancellation
- **THEN** the lifecycle remains `SUCCESS` and both later finish attempts report that they did not change the terminal outcome
#### Scenario: Budget exhaustion wins a terminal race
- **WHEN** budget exhaustion is recorded before an exception handler reports failure
- **THEN** the lifecycle remains `BUDGET_EXHAUSTED` with its original reason and completion time
### Requirement: Harness retry SHALL be typed, bounded, and recorded
The Harness SHALL provide immutable policies and an executor that records every actual attempt. Intent Router and SemanticGuard SHALL allow at most two attempts for their listed technical failures; Diagnosis Agent, Tool calls, and Evidence repair SHALL allow one attempt. Cancellation, deadline, budget exhaustion, `NO_EVIDENCE`, business rejection, and unknown failures SHALL NOT be retried.
#### Scenario: Router transport failure succeeds on retry
- **WHEN** a Fake Router fails its first attempt with a retryable transport failure and succeeds on the second
- **THEN** the executor returns the second result and records exactly one failed and one successful attempt
#### Scenario: Tool call fails
- **WHEN** a Fake Tool operation fails on its first attempt
- **THEN** the Tool policy records one failed attempt and throws without invoking the Tool again
#### Scenario: Run is cancelled between attempts
- **WHEN** cancellation is signalled after a retryable first failure
- **THEN** the executor rejects the next attempt and does not invoke the operation again
### Requirement: Spring AI hidden retry SHALL be disabled
Application configuration SHALL set `spring.ai.retry.max-attempts=1`, overriding the Spring AI 1.1.7 default of 10. Harness-permitted retries SHALL occur only in HarnessRetryExecutor and SHALL be observable as separate attempts.
#### Scenario: Retry configuration is loaded
- **WHEN** the repository application YAML is parsed in a focused configuration test
- **THEN** `spring.ai.retry.max-attempts` equals 1
### Requirement: Tool invocation key and Run capacity primitives SHALL be safe and store-independent
The Harness SHALL provide a ToolCallKeyFactory that constructs `{prefix}:{runId}:{toolCallId}` only from valid safe segments and never generates or rewrites the framework Tool Call ID. The Run capacity counter SHALL atomically reserve positive bytes up to its configured limit and SHALL leave usage unchanged when a reservation is rejected.
#### Scenario: Framework Tool Call ID is used in a key
- **WHEN** a valid run ID and framework Tool Call ID are supplied
- **THEN** the factory returns the configured prefix followed by the exact run ID and exact Tool Call ID
#### Scenario: Unsafe key segment is supplied
- **WHEN** either ID is blank, too long, or contains a separator or unsafe character
- **THEN** the factory rejects it and does not produce a Redis key
#### Scenario: Run byte limit would be exceeded
- **WHEN** a reservation would exceed max Run bytes
- **THEN** the counter rejects it and retains the usage from previously successful reservations
### Requirement: Harness Core SHALL remain independent from current runtime and persistence
This change SHALL NOT modify Controller/HTTP/SSE contracts, connect RunContext to the old ChatService or AiOpsService, use current ThreadLocal holders, persist `diagnosis_run`, access Redis, or implement Tool-specific projection. The new Core SHALL compile and pass focused Fake Model/Tool tests as a zero-consumer foundation for stage 3A.
#### Scenario: Stage 2 implementation completes
- **WHEN** focused Core, retry, budget, key, capacity, and configuration tests pass
- **THEN** current Chat/AIOps call sites remain unchanged and stage 3A can depend directly on the new Harness types
@@ -0,0 +1,38 @@
## 1. Run State Primitives
- [x] 1.1 Implement structurally immutable RunContext with explicit identity, deadline, cancellation, budget, retry policy, and lifecycle handles.
- [x] 1.2 Implement first-reason-wins cancellation with resource callbacks and first-terminal-wins Run lifecycle.
- [x] 1.3 Add Run identity, cancellation, and terminal race tests with a controlled Clock.
## 2. Budget and Capacity
- [x] 2.1 Implement validated caller-supplied RunBudgetLimits and thread-safe Model/Tool/per-Tool/Token accounting.
- [x] 2.2 Implement atomic single-Run byte capacity reservation with no partial update on rejection.
- [x] 2.3 Add budget tests for consistent counters, actual Token recording, concurrent reservations, and exhaustion details.
## 3. Harness Core
- [x] 3.1 Implement DiagnosisHarnessCore Run creation, active/deadline checks, budget gates, explicit cancellation, and success/failure completion.
- [x] 3.2 Add Fake Model/Tool tests proving explicit synchronous/asynchronous RunContext propagation and cancellation/deadline/budget terminal outcomes.
## 4. Retry Boundary
- [x] 4.1 Implement typed RetryFailure, bounded RetryPolicy, strict HarnessRetryPolicies, RetryAttempt, and RetryExecutionException.
- [x] 4.2 Implement HarnessRetryExecutor with per-attempt active checks and attempt recording, without backoff, recursion, or hidden loops.
- [x] 4.3 Add Fake Router/Tool/SemanticGuard retry tests for allowed technical retry, one-attempt policies, cancellation between attempts, and non-retryable outcomes.
## 5. Tool Store Foundations
- [x] 5.1 Implement configurable ToolCallKeyFactory with exact framework ID preservation and safe-segment validation.
- [x] 5.2 Add key factory tests covering exact key format, blank/unsafe/oversized segments, and no generated fallback ID.
## 6. Hidden Retry Configuration
- [x] 6.1 Set `spring.ai.retry.max-attempts` to 1 without changing Model routing or provider configuration.
- [x] 6.2 Add a focused YAML configuration test proving the hidden Spring AI retry override.
## 7. Verification
- [x] 7.1 Run focused Harness Core, retry, budget, key/capacity, and retry configuration tests.
- [x] 7.2 Run stage 0/1 Harness and ACI contract tests plus existing ChatController test as regression coverage.
- [x] 7.3 Verify new production code has no ThreadLocal/current-holder usage and current Chat/AIOps/Controller/JPA/Redis call sites remain unchanged.
@@ -0,0 +1,93 @@
# diagnosis-harness-run-context Specification
## Purpose
定义 Diagnosis Harness 的显式 RunContext、deadline、协作式取消、线程安全预算、唯一 Run 终态、类型化重试、Tool Call Key 和单 Run 容量基础,供后续 Tool boundary、Agent、Guard 和应用入口复用。
## Requirements
### Requirement: RunContext SHALL be explicit and structurally immutable
The Harness SHALL create a structurally immutable RunContext containing non-blank `sessionId`, non-blank `runId`, deadline, cancellation, budget, retry policies, and lifecycle. Every new Model, Tool, Guard, and application boundary SHALL receive the RunContext explicitly by parameter or framework-controlled context and SHALL NOT read or write ThreadLocal state.
#### Scenario: Context crosses an asynchronous boundary
- **WHEN** a Fake Tool runs on another thread with an explicitly supplied RunContext
- **THEN** it observes the same session ID, run ID, deadline, cancellation, budget, retry policies, and lifecycle handles without copying ThreadLocal state
#### Scenario: Invalid identity starts a Run
- **WHEN** Run creation receives a blank session ID or run ID
- **THEN** the Core rejects creation before any budget, cancellation, or lifecycle state is published
### Requirement: Deadline and cancellation SHALL stop subsequent controlled work
The Core SHALL compare the current Clock instant with the Run deadline before each controlled Model, Tool, retry, and capacity boundary. At or after the deadline it SHALL produce `TIMED_OUT`, signal `DEADLINE_EXCEEDED`, and reject subsequent work. Explicit cancellation SHALL preserve the first cancellation reason, notify all registered resource callbacks, and produce `CANCELLED` unless an earlier terminal state exists.
#### Scenario: Deadline expires before a Tool call
- **WHEN** the controlled Clock reaches the Run deadline before the Fake Tool boundary
- **THEN** the Tool is not invoked, the Run terminates as `TIMED_OUT`, and later checks return the same terminal outcome
#### Scenario: Client disconnects
- **WHEN** the Core receives `CLIENT_DISCONNECTED`
- **THEN** registered cancellation callbacks run, subsequent Model/Tool boundaries are rejected, and the Run has one `CANCELLED` terminal outcome
### Requirement: Run budgets SHALL be explicit, thread-safe, and observable
The Core SHALL require caller-supplied positive limits for Model calls, total Tool calls, per-Tool calls, input Tokens, output Tokens, total Tokens, and Run bytes. Budget updates SHALL be thread-safe and preserve a usage snapshot. Exceeding any limit SHALL record the actual applicable usage, produce `BUDGET_EXHAUSTED`, signal cancellation, and prevent subsequent controlled work.
#### Scenario: Per-Tool budget is exhausted
- **WHEN** a Fake Tool attempts one call beyond its configured per-Tool limit
- **THEN** the extra call is not reserved, the total and per-Tool usage remain internally consistent, and the Run terminates as `BUDGET_EXHAUSTED`
#### Scenario: Actual Token usage exceeds the limit
- **WHEN** a completed Fake Model response reports Token usage above the configured limit
- **THEN** the actual input/output/total usage remains recorded and the Core rejects the next controlled operation
### Requirement: A Run SHALL have exactly one terminal outcome
The lifecycle SHALL start as `RUNNING` and accept only the first terminal outcome from `SUCCESS`, `FAILED`, `CANCELLED`, `TIMED_OUT`, or `BUDGET_EXHAUSTED`. Later success, failure, cancellation, deadline, or budget events SHALL NOT replace the first `RunTermination` state, reason, or timestamp.
#### Scenario: Success wins a terminal race
- **WHEN** success is recorded before a later failure and cancellation
- **THEN** the lifecycle remains `SUCCESS` and both later finish attempts report that they did not change the terminal outcome
#### Scenario: Budget exhaustion wins a terminal race
- **WHEN** budget exhaustion is recorded before an exception handler reports failure
- **THEN** the lifecycle remains `BUDGET_EXHAUSTED` with its original reason and completion time
### Requirement: Harness retry SHALL be typed, bounded, and recorded
The Harness SHALL provide immutable policies and an executor that records every actual attempt. Intent Router and SemanticGuard SHALL allow at most two attempts for their listed technical failures; Diagnosis Agent, Tool calls, and Evidence repair SHALL allow one attempt. Cancellation, deadline, budget exhaustion, `NO_EVIDENCE`, business rejection, and unknown failures SHALL NOT be retried.
#### Scenario: Router transport failure succeeds on retry
- **WHEN** a Fake Router fails its first attempt with a retryable transport failure and succeeds on the second
- **THEN** the executor returns the second result and records exactly one failed and one successful attempt
#### Scenario: Tool call fails
- **WHEN** a Fake Tool operation fails on its first attempt
- **THEN** the Tool policy records one failed attempt and throws without invoking the Tool again
#### Scenario: Run is cancelled between attempts
- **WHEN** cancellation is signalled after a retryable first failure
- **THEN** the executor rejects the next attempt and does not invoke the operation again
### Requirement: Spring AI hidden retry SHALL be disabled
Application configuration SHALL set `spring.ai.retry.max-attempts=1`, overriding the Spring AI 1.1.7 default of 10. Harness-permitted retries SHALL occur only in HarnessRetryExecutor and SHALL be observable as separate attempts.
#### Scenario: Retry configuration is loaded
- **WHEN** the repository application YAML is parsed in a focused configuration test
- **THEN** `spring.ai.retry.max-attempts` equals 1
### Requirement: Tool invocation key and Run capacity primitives SHALL be safe and store-independent
The Harness SHALL provide a ToolCallKeyFactory that constructs `{prefix}:{runId}:{toolCallId}` only from valid safe segments and never generates or rewrites the framework Tool Call ID. The Run capacity counter SHALL atomically reserve positive bytes up to its configured limit and SHALL leave usage unchanged when a reservation is rejected.
#### Scenario: Framework Tool Call ID is used in a key
- **WHEN** a valid run ID and framework Tool Call ID are supplied
- **THEN** the factory returns the configured prefix followed by the exact run ID and exact Tool Call ID
#### Scenario: Unsafe key segment is supplied
- **WHEN** either ID is blank, too long, or contains a separator or unsafe character
- **THEN** the factory rejects it and does not produce a Redis key
#### Scenario: Run byte limit would be exceeded
- **WHEN** a reservation would exceed max Run bytes
- **THEN** the counter rejects it and retains the usage from previously successful reservations
### Requirement: Harness Core SHALL remain independent from current runtime and persistence
This change SHALL NOT modify Controller/HTTP/SSE contracts, connect RunContext to the old ChatService or AiOpsService, use current ThreadLocal holders, persist `diagnosis_run`, access Redis, or implement Tool-specific projection. The new Core SHALL compile and pass focused Fake Model/Tool tests as a zero-consumer foundation for stage 3A.
#### Scenario: Stage 2 implementation completes
- **WHEN** focused Core, retry, budget, key, capacity, and configuration tests pass
- **THEN** current Chat/AIOps call sites remain unchanged and stage 3A can depend directly on the new Harness types