feat(harness): add run context and retry core
This commit is contained in:
@@ -175,6 +175,22 @@
|
|||||||
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
|
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
|
||||||
- 边界:`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
|
- 边界:`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
|
||||||
|
|
||||||
|
### RunContext
|
||||||
|
- 定义:一次 Diagnosis Run 的显式执行上下文,结构不可变地携带 `sessionId`、`runId`、deadline,以及该 Run 独占的取消、预算、重试策略和生命周期状态句柄。
|
||||||
|
- 边界:RunContext 通过方法参数或框架受控 context 显式传播,不依赖 ThreadLocal;结构不可变不等于内部计数和取消状态不能变化,这些变化由线程安全句柄管理。
|
||||||
|
|
||||||
|
### Run Lifecycle
|
||||||
|
- 定义:Diagnosis Harness 对单次 Run 执行状态的内存控制,采用 first-terminal-wins 规则保证成功、失败、取消、超时和预算耗尽只能产生一个最终终态。
|
||||||
|
- 边界:Run Lifecycle 不直接等同于数据库实体写入;应用用例负责把最终状态映射到 `diagnosis_run` 持久化。
|
||||||
|
|
||||||
|
### Run Budget
|
||||||
|
- 定义:单次 Run 的模型调用、Tool 调用、单 Tool 调用、输入/输出/总 Token 和 canonical invocation 字节容量的线程安全消耗计数与门禁。
|
||||||
|
- 边界:预算上限由 Harness 配置显式提供;实际 Token 在模型响应后记录,超限后保留真实消耗并阻止后续执行。
|
||||||
|
|
||||||
|
### Harness Retry Policy
|
||||||
|
- 定义:Harness 对同一技术操作 attempt 数和可重试失败类型的显式策略。
|
||||||
|
- 边界:Router 与 SemanticGuard 的技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 只有一次 attempt。Agent 正常 ReAct 轮次不是 retry,`NO_EVIDENCE`、业务拒绝、取消和预算耗尽不可重试。
|
||||||
|
|
||||||
### Verifier Skill Isolation
|
### Verifier Skill Isolation
|
||||||
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
|
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
|
||||||
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
||||||
|
|||||||
@@ -4,6 +4,7 @@
|
|||||||
|
|
||||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||||
|---|---|---|---|---|---|---|
|
|---|---|---|---|---|---|---|
|
||||||
|
| 2026-07-21 | single-react-harness-run-context | 建立显式 RunContext、Harness Core、预算、取消、类型化重试和 Tool Store 基础。 | Harness/Run lifecycle/Budget | ISS-014, RunContext, deadline, cancellation, budget, retry, ToolCallKey | openspec/changes/archive/2026-07-21-single-react-harness-run-context | archived |
|
||||||
| 2026-07-21 | single-react-aci-tool-contracts | 冻结 RAG、日志和 MySQL evidence Tool 的 Agent-facing ACI Schema、状态、框架调用引用和描述边界。 | Harness/Agent Tool contract | ISS-014, ACI, tool_call_id, evidence_status, RAG, query_logs, query_mysql, MOCK | openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts | archived |
|
| 2026-07-21 | single-react-aci-tool-contracts | 冻结 RAG、日志和 MySQL evidence Tool 的 Agent-facing ACI Schema、状态、框架调用引用和描述边界。 | Harness/Agent Tool contract | ISS-014, ACI, tool_call_id, evidence_status, RAG, query_logs, query_mysql, MOCK | openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts | archived |
|
||||||
| 2026-07-21 | single-react-design-freeze | 冻结单体 Diagnosis Agent、Harness、Guard、工具证据与阶段门禁契约。 | Chat/Harness/Agent contract | ISS-014, single ReactAgent, Harness, EvidenceGuard, SemanticGuard, tool_call_id, evidence_status | openspec/changes/archive/2026-07-21-single-react-design-freeze | archived |
|
| 2026-07-21 | single-react-design-freeze | 冻结单体 Diagnosis Agent、Harness、Guard、工具证据与阶段门禁契约。 | Chat/Harness/Agent contract | ISS-014, single ReactAgent, Harness, EvidenceGuard, SemanticGuard, tool_call_id, evidence_status | openspec/changes/archive/2026-07-21-single-react-design-freeze | archived |
|
||||||
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
||||||
|
|||||||
@@ -0,0 +1,45 @@
|
|||||||
|
# Acceptance: single-react-harness-run-context
|
||||||
|
|
||||||
|
## 实现结果
|
||||||
|
|
||||||
|
- 新增显式 `RunContext`、取消信号、deadline 检查、Run Lifecycle 和 first-terminal-wins。
|
||||||
|
- 新增 caller-supplied limits、线程安全 Model/Tool/Token/Run bytes 预算和容量 CAS。
|
||||||
|
- 新增 strict typed retry policies/executor,记录每个实际 attempt 并阻止取消/预算异常误重试。
|
||||||
|
- 新增无 Redis 依赖的 ToolCallKeyFactory,精确保留框架 Tool Call ID。
|
||||||
|
- Spring AI `spring.ai.retry.max-attempts` 设置为 1;未修改 provider/model routing。
|
||||||
|
- 未接入旧 Chat/AIOps/Controller、ThreadLocal、JPA、Redis、Agent 或公开协议。
|
||||||
|
|
||||||
|
## 静态验证
|
||||||
|
|
||||||
|
- `openspec validate single-react-harness-run-context --strict`:通过。
|
||||||
|
- 新 Harness 包 `rg`:无 ThreadLocal/current-holder/Redis 引用。
|
||||||
|
- 旧 Chat/AIOps/Controller/JPA 调用链 diff:为空。
|
||||||
|
- 受保护配置检查:仅新增 `spring.ai.retry.max-attempts: 1`,model routing/provider 保持不变。
|
||||||
|
|
||||||
|
## 脚本验证
|
||||||
|
|
||||||
|
- `mvn -q -DskipTests compile`:通过。
|
||||||
|
- Core focused suite:通过。
|
||||||
|
- 综合回归 suite(阶段 0/1 契约 + 阶段 2 Core + ChatController):通过。
|
||||||
|
|
||||||
|
## 浏览器/人工验证
|
||||||
|
|
||||||
|
- 不适用。本阶段没有 UI、Controller、SSE 或公开协议变化。
|
||||||
|
|
||||||
|
## 未验证
|
||||||
|
|
||||||
|
- 未执行真实模型调用、Redis、CLS、MySQL 或客户端断开 live E2E;这些属于后续 Tool store/adapter/最终 E2E 门禁。
|
||||||
|
- 未将 RunContext 接入旧 ChatService;这是本阶段明确非目标,阶段 6A 才接入。
|
||||||
|
- `RunContext` 取消对已进入的同步第三方调用仍是协作式;实际 HTTP/JDBC future 取消留给后续 adapter。
|
||||||
|
- Provider 侧凭据轮换状态不由仓库证明。
|
||||||
|
|
||||||
|
## 剩余风险与后续门禁
|
||||||
|
|
||||||
|
- 关闭 SDK 隐式 retry 会使旧路径瞬时错误不再自动重试,直到后续 Router/SemanticGuard 接入 Harness;该变化已记录并可通过恢复配置回滚。
|
||||||
|
- 下一阶段 3A 必须在 ToolInterceptor 接收 RunContext 和框架 Tool Call ID,直接复用 Key/Capacity/Cancellation 门禁。
|
||||||
|
|
||||||
|
## 状态
|
||||||
|
|
||||||
|
- Stage acceptance: accepted
|
||||||
|
- OpenSpec archive: archived at `openspec/changes/archive/2026-07-21-single-react-harness-run-context`
|
||||||
|
- Main spec sync: `openspec/specs/diagnosis-harness-run-context/spec.md`(8 added requirements)
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# Brief: single-react-harness-run-context
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
旧 Chat/AIOps 通过多个 ThreadLocal 和业务方法内状态机传播 session/run/token/retry,无法成为后续 Tool、Agent、Guard 和新入口的稳定共同边界。Spring AI 默认 10 attempts 还会制造未被 Harness 记录的隐藏重试。
|
||||||
|
|
||||||
|
## 目标
|
||||||
|
|
||||||
|
- 建立显式、结构不可变、可异步传播的 RunContext。
|
||||||
|
- 集中实现 deadline、协作式取消、线程安全预算和唯一 Run 终态。
|
||||||
|
- 实现类型化、最多两次 attempt、逐 attempt 记录的 Harness retry。
|
||||||
|
- 提供阶段 3A 可直接使用的 Tool Call Key Factory 和单 Run 容量计数器。
|
||||||
|
- 将 Spring AI 底层 retry 压为一次。
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- 新增 Harness Core/Retry/Tool Store foundation 类型和 focused Fake tests。
|
||||||
|
- 更新 `application.yml` 的 Spring AI retry 配置。
|
||||||
|
- 更新 glossary 和 OpenSpec/devflow 档案。
|
||||||
|
|
||||||
|
## 非目标
|
||||||
|
|
||||||
|
- 不接入旧 ChatService/AiOpsService/Controller/SSE。
|
||||||
|
- 不读写现有 ThreadLocal,不删除旧实现。
|
||||||
|
- 不写 `diagnosis_run` 或 Redis,不实现 Tool projection。
|
||||||
|
|
||||||
|
## 元数据
|
||||||
|
|
||||||
|
- 分档:complex
|
||||||
|
- 接口影响:L2;另有关闭旧 SDK 隐式 retry 的有意内部行为变化
|
||||||
|
- 关联 Issue:ISS-014 阶段 2
|
||||||
|
- 关联 OpenSpec:`openspec/changes/single-react-harness-run-context`
|
||||||
@@ -0,0 +1,110 @@
|
|||||||
|
# Decisions: single-react-harness-run-context
|
||||||
|
|
||||||
|
## 规模与入口
|
||||||
|
|
||||||
|
- 分档:complex。
|
||||||
|
- 入口:ISS-014 阶段 2;阶段 0/1 已分别由 Git commit `58c3910`、`4274f33` 固化并 Archive。
|
||||||
|
- 目标:建立后续 Tool、Agent、Guard 和入口共同依赖的显式 Harness 运行边界,不接旧 ChatService。
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
- `devflow/index.md` 命中 design freeze、ACI contracts 和 session/run trace isolation。
|
||||||
|
- 当前 `SessionContextHolder`、`VerifierContextHolder`、`TokenUsageHolder` 使用 ThreadLocal;新 Harness 禁止复用。
|
||||||
|
- 当前 ChatService 自行生成 runId、写 `diagnosis_run`、执行两轮 retry loop 并在 finally 清理 ThreadLocal,职责与新 Harness 边界冲突。
|
||||||
|
- `DiagnosisRun.status` 当前是字符串 `PENDING/RUNNING/SUCCESS/FAILED`;本阶段不修改实体或迁移,由后续应用用例映射。
|
||||||
|
- Spring AI 1.1.7 `SpringAiRetryProperties` 的配置前缀是 `spring.ai.retry`,默认 `maxAttempts=10`。
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| 维度 | 问题 | 模式 | 证据与结论 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| 术语 | “不可变 RunContext”是否意味着预算/取消也不能变化? | evidence-driven | ISS 要求上下文不可变同时要求计数/取消/终态;采用结构不可变 record + 线程安全单 Run 状态句柄。 | 已解决并汇报 |
|
||||||
|
| 术语 | Run 与 DiagnosisRun 是否在阶段 2 直接持久化绑定? | evidence-driven | 阶段 2 只建 Core,阶段 6A 才接应用用例;本阶段生命周期为内存执行真理,不改 JPA。 | 已解决并汇报 |
|
||||||
|
| 边界 | 是否迁移旧 ThreadLocal/ChatService 调用? | evidence-driven | ISS 1109-1111 明确禁止 ThreadLocal 和旧 ChatService 临时适配;只新增零消费者 Core。 | 已解决并汇报 |
|
||||||
|
| 边界 | 取消是否承诺立即中断同步模型调用? | evidence-driven | 阶段 0 已确认分层取消;阻止后续边界并执行资源回调,不承诺不可证明的硬中断。 | 已解决并汇报 |
|
||||||
|
| 验收 | 唯一终态如何证明? | evidence-driven | atomic first-terminal-wins lifecycle,并用取消/预算/异常/成功竞态测试证明后续终态不能覆盖。 | 已解决并汇报 |
|
||||||
|
| 验收 | 异步传播如何证明不依赖 ThreadLocal? | evidence-driven | Fake Tool 在 `CompletableFuture` 线程只接收显式 RunContext,并验证相同 session/run/state handles。 | 已解决并汇报 |
|
||||||
|
| 技术 | 隐藏重试如何关闭? | evidence-driven | 本地依赖 `SpringAiRetryProperties` 证实 `spring.ai.retry.max-attempts` 默认 10;配置改为 1 并加配置测试。 | 已解决并汇报 |
|
||||||
|
| 技术 | 尚未校准的预算默认值如何处理? | evidence-driven | ISS 要求集中配置且不伪造数值;Core 接受显式 limits,不内置默认预算,后续 Spring wiring 决定配置值。 | 已解决并汇报 |
|
||||||
|
| 技术 | Tool Call Key 是否生成 Tool ID? | evidence-driven | 阶段 0/1 冻结框架 ID;Factory 仅验证安全 segment 并拼接,不生成或改写。 | 已解决并汇报 |
|
||||||
|
|
||||||
|
## Grill 结论
|
||||||
|
|
||||||
|
- 术语、边界、验收和技术问题均可由已确认 ISS、现有代码和本地依赖 API 证明。
|
||||||
|
- 没有新的产品偏好、公开协议或风险接受度问题需要 `user-interview`;短期关闭旧 SDK retry 的影响已在 proposal 明示。
|
||||||
|
- `grill-with-docs` 的代码可证问题已先查证并向用户汇报;所有结论均已回写 proposal。
|
||||||
|
|
||||||
|
## 能力与工具限制
|
||||||
|
|
||||||
|
- Discover 能力来源:`sm-flow` + `grill-with-docs`。
|
||||||
|
- 当前工具集没有 `codebase-retrieval` 和 LSP;使用 `rg`、源码阅读、本地依赖 `jar/javap`、编译和 focused tests 补足调用链与 API 核对。
|
||||||
|
|
||||||
|
## Cross-artifact 对齐
|
||||||
|
|
||||||
|
| 链路 | 状态 | 结论 |
|
||||||
|
|---|---|---|
|
||||||
|
| brief 目标/范围/非目标 -> proposal | 已对齐 | 显式 context、Core、预算/取消/终态、retry、key/capacity、隐藏 retry 和不接旧 runtime 全部覆盖。 |
|
||||||
|
| proposal 范围/约束/承诺 -> design | 已对齐 | 类职责、并发语义、first-wins、配置变化、迁移和回滚均有明确设计。 |
|
||||||
|
| design 决策/接口影响/风险 -> specs/tasks | 已对齐 | 每个状态边界都有 scenario,配置行为变化和 zero-consumer 边界有独立任务与验证。 |
|
||||||
|
| specs 可观察行为 -> tasks | 已对齐 | 8 条 requirements 分解为 primitives、budget、Core、retry、key、配置和三层验证。 |
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
- 能力来源:`zoom-out`,使用 glossary 的 RunContext、Run Lifecycle、Run Budget、Harness Retry Policy 和 Diagnosis Run 术语。
|
||||||
|
- 输入链路为未来 application use case 创建 RunContext,处理链路由 Core/Retry/Key/Capacity 通过显式参数消费,输出为唯一 RunTermination;当前 runtime 不接入。
|
||||||
|
- RunContext 拥有单 Run 内存状态,应用用例拥有数据库映射,Tool store 拥有 Redis invocation;数据所有权没有重叠。
|
||||||
|
- Core 不保存全局 Run map、不调用 Agent/数据库/Redis、不实现循环编排,因此不会演变为工作流引擎。
|
||||||
|
- 最大风险是关闭 SDK retry 对旧路径的短期行为影响;已作为显式配置变更进入 proposal/spec/test 和回滚说明,无 ADR 冲突。
|
||||||
|
|
||||||
|
## Commit Gate Preflight
|
||||||
|
|
||||||
|
- proposal、design、specs、tasks 完整,OpenSpec status complete,strict validation 通过。
|
||||||
|
- question pool 无未汇报 evidence-driven 结论、无未确认 user-interview 问题。
|
||||||
|
- 接口影响 L2;`spring.ai.retry.max-attempts=1` 的有意内部行为变化已明确影响和回滚边界。
|
||||||
|
- cross-artifact 四段对齐无 gap,架构审计约束已进入 design/spec/tasks。
|
||||||
|
- Apply 持续授权已存在;执行范围严格限制为新 Harness foundation、配置 override 和 focused tests。
|
||||||
|
|
||||||
|
## Pre-apply Research
|
||||||
|
|
||||||
|
### 参考实现与反例
|
||||||
|
|
||||||
|
- `SessionContextHolder`:ThreadLocal session/run fallback,新 Harness 明确禁止复用。
|
||||||
|
- `TokenUsageHolder` / `TokenTrackingChatModel`:当前只记录 total token 且依赖 ThreadLocal,后续 model boundary 应改为显式 RunBudget;本阶段不改旧类。
|
||||||
|
- `ChatService.executeChatComplex`:当前业务方法内创建 Run、两轮 retry、写终态和清理 ThreadLocal,是后续替换对象,不是 Core 参考实现。
|
||||||
|
- `DiagnosisRun`:现有持久化字段和字符串状态;本阶段只确认映射边界,不修改实体或 Repository。
|
||||||
|
- Spring AI 1.1.7 `SpringAiRetryProperties`:`spring.ai.retry` 前缀、默认 `maxAttempts=10`,支持精确配置 override。
|
||||||
|
|
||||||
|
### 技术栈清单
|
||||||
|
|
||||||
|
- Java 17 record 表达结构不可变 context/limits/snapshot/policy/attempt。
|
||||||
|
- `AtomicReference` 实现 first-reason/first-terminal-wins;`AtomicLong` 实现 capacity CAS;同步临界区维护复合预算一致性。
|
||||||
|
- `Clock` 和 `Supplier<String>` 注入保证 deadline/ID 可测试,不引入 scheduler 或全局 registry。
|
||||||
|
- SLF4J 只记录取消 callback 异常,不记录用户输入、Tool payload 或凭据。
|
||||||
|
- SnakeYAML 直接解析 classpath `application.yml` 验证 retry override,不启动外部 MySQL/Redis/Milvus/模型。
|
||||||
|
|
||||||
|
### 新建基础设施
|
||||||
|
|
||||||
|
- `harness.core`:RunContext、Cancellation、Lifecycle、Budget、Capacity、Core 和类型化异常/状态。
|
||||||
|
- `harness.retry`:RetryFailure/Policy/Policies/Attempt/Executor/Exception 与函数接口。
|
||||||
|
- `harness.tool.store.ToolCallKeyFactory`:纯 Key 构造,不访问 Redis。
|
||||||
|
- focused unit tests 与 Fake Model/Tool;无需新 Maven 依赖。
|
||||||
|
|
||||||
|
### 影响半径
|
||||||
|
|
||||||
|
- 新生产包在本阶段保持零消费者。
|
||||||
|
- 唯一现有运行配置变化为 `spring.ai.retry.max-attempts=1`;Model 路由、provider、Controller、JPA 和 Redis 配置保持不变。
|
||||||
|
|
||||||
|
## Apply 结果
|
||||||
|
|
||||||
|
- 冲突分类:未发现 OpenSpec 遗漏、代码偏离或方向不确定项;一次自审发现 RetryExecutor 需要无条件拦截预算/取消异常,已回写代码并通过回归测试。
|
||||||
|
- 新增 `RunContext`、Cancellation、Lifecycle、Budget、Capacity、DiagnosisHarnessCore、typed Retry 和 ToolCallKeyFactory;未接旧 Chat/AIOps/Controller/Redis/JPA。
|
||||||
|
- Spring AI 全局 retry 已由默认 10 压为 1;Harness strict policies 只允许 Router/SemanticGuard 技术失败一次显式重试。
|
||||||
|
- 首模块对齐:Run state/budget/Core/retry/key/config 与 design/tasks 全部完成;Tool interceptor/store/Agent/应用用例仍留给后续阶段。
|
||||||
|
|
||||||
|
## Apply 验证
|
||||||
|
|
||||||
|
- 编译:`mvn -q -DskipTests compile`:通过。
|
||||||
|
- Core focused:`mvn -q '-Dtest=RunContextTest,RunBudgetTest,DiagnosisHarnessCoreTest,HarnessRetryExecutorTest,ToolCallKeyFactoryTest,SpringAiRetryConfigurationTest' test`:通过(加固后复跑通过)。
|
||||||
|
- 综合回归:`mvn -q '-Dtest=HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest,RunContextTest,RunBudgetTest,DiagnosisHarnessCoreTest,HarnessRetryExecutorTest,ToolCallKeyFactoryTest,SpringAiRetryConfigurationTest,ChatControllerTest' test`:通过。
|
||||||
|
- 静态 scope:新 Harness 包无 ThreadLocal/current-holder/Redis 引用;旧 Chat/AIOps/Controller/JPA 调用链 diff 为空;模型路由/provider 未改。
|
||||||
|
- OpenSpec:`openspec validate single-react-harness-run-context --strict`:通过。
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# Evidence: single-react-harness-run-context
|
||||||
|
|
||||||
|
## 文档与依赖证据
|
||||||
|
|
||||||
|
- ISS-014 4.2/阶段 2 要求显式 `RunContext`、deadline、取消、预算、retry、Key Factory 和 no-ThreadLocal 边界。
|
||||||
|
- 阶段 0/1 OpenSpec 已冻结框架 `tool_call_id`、两套状态语义和后续阶段串行门禁。
|
||||||
|
- 本地 Spring AI 1.1.7 `SpringAiRetryProperties` 的 `@ConfigurationProperties("spring.ai.retry")` 默认 `maxAttempts=10`;配置已覆盖为 1。
|
||||||
|
|
||||||
|
## 代码证据
|
||||||
|
|
||||||
|
- `SessionContextHolder`、`TokenUsageHolder`、`VerifierContextHolder` 当前是旧链路 ThreadLocal;新 `com.superbiz.agent.harness` 包无任何 holder/ThreadLocal/Redis 引用。
|
||||||
|
- `ChatService.executeChatComplex` 当前自行创建 run、执行两轮 retry、写 `diagnosis_run` 和清理 ThreadLocal;新 Core 不接入该方法,后续应用用例负责迁移。
|
||||||
|
- `DiagnosisRun` 仍保留现有字符串状态和 JPA Schema;阶段 2 未修改实体、Repository 或数据库。
|
||||||
|
- `DiagnosisHarnessCore` 不保存全局 Run map;RunContext 结构不可变,Cancellation/Budget/Lifecycle 为同一 Run 的线程安全句柄。
|
||||||
|
|
||||||
|
## Evidence-driven 结论
|
||||||
|
|
||||||
|
- first-reason-wins cancellation + first-terminal-wins lifecycle 可以用 AtomicReference 实现并跨异步边界共享。
|
||||||
|
- 复合 Tool/Token 预算需要同步一致性;Run bytes 用 CAS 预留避免并发超限或部分增长。
|
||||||
|
- SDK 隐藏 retry 压为一次后,Harness 才能记录 Router/SemanticGuard 的显式 attempt;Tool/Diagnosis/Evidence repair 固定一次。
|
||||||
|
- Key Factory 可直接复用阶段 3A,但不生成或改写框架 Tool Call ID,也不访问 Redis。
|
||||||
|
|
||||||
|
## 工具限制
|
||||||
|
|
||||||
|
- `codebase-retrieval` 和 LSP 不在当前工具集中;使用 `rg`、源码阅读、`jar/javap`、Maven 编译、YAML 解析和 focused tests 补足核对。
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
# ISS-014 单体 ReAct Agent、Harness 与 ACI 工具瘦身
|
# ISS-014 单体 ReAct Agent、Harness 与 ACI 工具瘦身
|
||||||
|
|
||||||
**状态**:实施中(阶段 0-1 已归档,下一阶段 2)
|
**状态**:实施中(阶段 0-2 已归档,下一阶段 3)
|
||||||
**严重程度**:高
|
**严重程度**:高
|
||||||
**发现时间**:2026-07-20
|
**发现时间**:2026-07-20
|
||||||
**目标分支**:`refactor/chat-single-react-harness`
|
**目标分支**:`refactor/chat-single-react-harness`
|
||||||
|
|||||||
@@ -0,0 +1 @@
|
|||||||
|
Archive-ready after implementation and focused verification on 2026-07-21.
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
Committed after strict validation on 2026-07-21.
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
schema: spec-driven
|
||||||
|
created: 2026-07-21
|
||||||
@@ -0,0 +1,109 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
现有 Chat/AIOps 在进入 Agent 前自行生成 session/run 标识,通过多个 ThreadLocal 和 RunnableConfig metadata 混合传播,并在业务方法内直接处理 retry loop、运行持久化和成功/失败终态。`TokenTrackingChatModel` 也只把 total token 写到 ThreadLocal,无法在异步或并发边界可靠归属 Run。后续 Tool interceptor、ResultProjector、Diagnosis Agent、EvidenceGuard、SemanticGuard 和 Chat application use case 需要共同使用一个不依赖旧 ChatService 的显式执行上下文。
|
||||||
|
|
||||||
|
Spring AI 1.1.7 的全局 retry properties 前缀为 `spring.ai.retry`,默认 `maxAttempts=10`。ISS-014 已确认所有底层模型 attempt 必须压为一次,由 Harness 对 Router/SemanticGuard 的特定技术失败显式执行最多一次重试并记录每次 attempt。
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- 提供显式、结构不可变、可跨同步/异步边界传递的 RunContext。
|
||||||
|
- 在内存执行层统一 deadline、取消、预算和唯一终态语义。
|
||||||
|
- 提供无隐藏循环的类型化 retry policy/executor 和 attempt 记录。
|
||||||
|
- 提供阶段 3A 可直接复用的 Tool Call Key Factory 与单 Run 字节容量计数器。
|
||||||
|
- 关闭 Spring AI 默认隐藏重试,保证实际模型 attempt 可由 Harness 观察。
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- 不接入 Controller、SSE、旧 ChatService、AiOpsService、ReactAgent 或现有 Tool。
|
||||||
|
- 不替换或删除旧 ThreadLocal;阶段 7 在新入口切换后清理。
|
||||||
|
- 不修改 `diagnosis_run` 实体、表结构或 Repository,不写数据库终态。
|
||||||
|
- 不实现 Redis invocation store、ToolInterceptor、ToolResultProjector 或 Tool Call ID 校验全套策略。
|
||||||
|
- 不实现定时调度器、线程中断、工作流引擎、动态配置中心、Retry DSL 或退避算法。
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
### 1. RunContext 是结构不可变的共享状态句柄集合
|
||||||
|
|
||||||
|
`RunContext` 使用 Java 17 record,字段固定为 `sessionId`、`runId`、`deadline`、`RunCancellation`、`RunBudget`、`HarnessRetryPolicies` 和 `RunLifecycle`。record 不提供 setter;同一 Run 的同步/异步消费者必须显式传递同一个 context,因此共享相同的取消、预算和生命周期句柄。
|
||||||
|
|
||||||
|
替代方案是把所有字段做成纯值并在每次变化时复制 context;这会造成并发分支状态分叉,无法保证唯一终态和原子预算,故不采用。ThreadLocal/RunnableConfig fallback 也不进入新 Core。
|
||||||
|
|
||||||
|
### 2. DiagnosisHarnessCore 是唯一执行门禁与状态转换入口
|
||||||
|
|
||||||
|
Core 由 `Clock`、Run ID supplier、最大 Run duration、显式 `RunBudgetLimits` 和 `HarnessRetryPolicies` 构造。它负责创建 context、在每个模型/Tool/容量边界检查 active/deadline、预留预算、记录 Token、显式取消以及完成成功/失败终态。
|
||||||
|
|
||||||
|
Core 不保存全局 Run map;生命周期完全随 RunContext 所有,避免跨请求泄漏。调用方可以传入既有 runId 进行确定性测试/恢复,但 Core 不生成 sessionId,也不修改业务持久化。
|
||||||
|
|
||||||
|
### 3. Deadline 与取消采用协作式、可观察语义
|
||||||
|
|
||||||
|
`RunCancellation` 用 atomic first-reason-wins 保存 `CLIENT_DISCONNECTED`、`USER_REQUESTED`、`DEADLINE_EXCEEDED`、`BUDGET_EXHAUSTED` 或 `INTERNAL_FAILURE`,并允许注册资源取消回调。Core 在每个受控边界比较 `Clock.instant()` 与 deadline;到达或超过 deadline 时先写 `TIMED_OUT` 终态,再触发取消信号并拒绝后续执行。
|
||||||
|
|
||||||
|
取消回调异常会记录并继续通知其他回调,不能阻止取消传播。此模型不承诺已进入的同步第三方调用立即停止;具体 HTTP/JDBC future 取消由后续 adapter 注册回调实现。
|
||||||
|
|
||||||
|
### 4. RunLifecycle 使用 first-terminal-wins
|
||||||
|
|
||||||
|
`RunLifecycle` 初始为 `RUNNING`,终态固定为 `SUCCESS`、`FAILED`、`CANCELLED`、`TIMED_OUT`、`BUDGET_EXHAUSTED`。内部使用 atomic compare-and-set 保存唯一 `RunTermination(state, reason, completedAt)`;任何后续 finish 都返回 false 且不得覆盖首个终态。
|
||||||
|
|
||||||
|
这只是执行内存真理;阶段 6A 的应用用例负责把终态映射到 `diagnosis_run.status/release_outcome`。本阶段不创建第二套数据库状态机。
|
||||||
|
|
||||||
|
### 5. RunBudget 使用显式 limits 和一致性更新
|
||||||
|
|
||||||
|
`RunBudgetLimits` 必须由调用方显式提供,包含模型调用、总 Tool 调用、单 Tool 调用、input/output/total Token 和 max Run bytes,所有值必须为正。`RunBudget` 使用同步临界区保证总 Tool/单 Tool 计数和三类 Token 计数要么一致提交,要么明确抛出 `BudgetExceededException`;模型返回后的实际 Token 即使超限也会先记入 usage,再终止后续执行。
|
||||||
|
|
||||||
|
`RunCapacityCounter` 用 CAS 原子预留 UTF-8 JSON 字节,失败时不部分增长。它被 RunBudget 持有并可由阶段 3A 直接调用。Core 在预算失败时先写 `BUDGET_EXHAUSTED` 终态,再触发取消。
|
||||||
|
|
||||||
|
### 6. Retry policy 只描述 attempt 上限与类型
|
||||||
|
|
||||||
|
`RetryPolicy(maxAttempts, retryableFailures)` 只允许 1 或 2 attempts。`HarnessRetryPolicies.strict()` 固定:Router=2(timeout/transport/invalid output)、Diagnosis=1、Tool=1、SemanticGuard=2(timeout/transport/parse/schema)、EvidenceRepair=1。
|
||||||
|
|
||||||
|
`HarnessRetryExecutor` 在每次 attempt 前调用 Core active check,操作成功/失败都发送 `RetryAttempt` 给 recorder。只有 policy 允许且尚有 attempt 时继续;取消、deadline、预算、`NO_EVIDENCE`、业务拒绝和未知失败不得重试。Executor 不 sleep、不退避、不递归、不调用 Agent。
|
||||||
|
|
||||||
|
### 7. 底层 Spring AI retry 固定为一次
|
||||||
|
|
||||||
|
`application.yml` 设置 `spring.ai.retry.max-attempts: 1`,并通过直接读取 YAML 的单元测试锁定,避免启动完整外部基础设施。该配置会立即影响旧运行路径:瞬时模型错误不再由 SDK 隐式重试;这是保证 attempt 可观测的有意内部行为变化。
|
||||||
|
|
||||||
|
### 8. ToolCallKeyFactory 不拥有 ID
|
||||||
|
|
||||||
|
Factory 接收可配置前缀,验证 `runId/toolCallId` 为非空、安全长度和安全字符 segment 后构造 `{prefix}:{runId}:{toolCallId}`。它不生成、规范化或哈希框架 Tool Call ID,不访问 Redis。严格的 duplicate/cross-run invocation 语义属于阶段 3A store。
|
||||||
|
|
||||||
|
## Module Flow
|
||||||
|
|
||||||
|
```text
|
||||||
|
future Chat application use case
|
||||||
|
-> DiagnosisHarnessCore.startRun(...)
|
||||||
|
-> RunContext
|
||||||
|
-> RunCancellation
|
||||||
|
-> RunBudget -> RunCapacityCounter
|
||||||
|
-> HarnessRetryPolicies
|
||||||
|
-> RunLifecycle
|
||||||
|
-> future Model/Tool/Guard boundary explicitly receives RunContext
|
||||||
|
-> Core check/reserve/record/finish
|
||||||
|
-> HarnessRetryExecutor for allowed technical operations only
|
||||||
|
-> future application use case maps the one RunTermination to diagnosis_run
|
||||||
|
```
|
||||||
|
|
||||||
|
阶段 3A 直接依赖 RunContext、Core、Key Factory 和 capacity counter;它不需要调用旧 ChatService 或读取 ThreadLocal。
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [结构不可变但句柄可变容易被误用] -> 命名、Javadoc 和并发/异步测试明确同一 Run 必须共享同一 context 实例。
|
||||||
|
- [关闭 SDK retry 后旧路径瞬时失败率可能上升] -> 明确列为行为变化,配置测试锁定;后续 Router/SemanticGuard 只按已确认策略补回可观测重试。
|
||||||
|
- [实际 Token 在响应后才知道,可能超预算] -> usage 保留实际值并立即进入预算终态,不丢失消耗、不再执行下一步。
|
||||||
|
- [取消回调由 adapter 决定能否硬取消] -> Core 只承诺协作式取消和后续门禁,后续 JDBC/HTTP adapter 必须注册资源回调。
|
||||||
|
- [纯内存终态无法替代审计] -> 阶段 6A 持久化映射;本阶段 focused tests 只验证执行语义。
|
||||||
|
|
||||||
|
## Migration Plan
|
||||||
|
|
||||||
|
1. 本阶段新增零消费者 Core 类型和 tests,同时把 Spring AI retry 压为一次。
|
||||||
|
2. 阶段 3A 的 Tool interceptor/store 显式接收 RunContext,复用 Key/容量门禁。
|
||||||
|
3. 阶段 4/5 的 Agent/Guard 使用 Core model/retry/deadline 边界。
|
||||||
|
4. 阶段 6A 的 Chat application use case 创建 RunContext 并映射终态到数据库。
|
||||||
|
5. 阶段 6B 切换入口;阶段 7 删除 ThreadLocal 和旧重试循环。
|
||||||
|
|
||||||
|
回滚本阶段代码可删除新包;`spring.ai.retry.max-attempts` 若回滚到默认 10 会恢复隐藏重试,但会再次失去 attempt 可观测性,因此只允许在整体重构回滚时明确执行。
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
无。具体预算数值、Run ID 格式和持久化状态映射由后续 wiring/应用用例在不改变本 Core 语义的前提下配置。
|
||||||
@@ -0,0 +1,46 @@
|
|||||||
|
## Why
|
||||||
|
|
||||||
|
当前 Chat/AIOps 通过 `SessionContextHolder`、`VerifierContextHolder` 和 `TokenUsageHolder` 等 ThreadLocal 隐式传播 session/run/token/retry 状态,且 Run 终态、取消、预算和重试分散在业务循环与 SDK 默认行为中。阶段 3A 的 Tool boundary、后续 Agent/Guard 和最终入口需要先共同依赖一个显式、可跨同步/异步边界传递的 Harness RunContext,否则会继续耦合旧 ChatService 并产生重复状态机。
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- 新增结构不可变的 `RunContext`,显式携带 `sessionId`、`runId`、deadline、取消信号、线程安全预算、类型化重试策略和 first-terminal-wins 生命周期。
|
||||||
|
- 新增 `DiagnosisHarnessCore` 负责 Run 创建、deadline 检查、模型/Tool/Token/容量预算门禁、取消传播和唯一终态,不承担业务推理或持久化。
|
||||||
|
- 新增 caller-supplied `RunBudgetLimits`、线程安全 `RunBudget` 与单 Run 字节容量计数器;Core 不硬编码尚未校准的预算默认值。
|
||||||
|
- 新增 `HarnessRetryPolicies` 与 `HarnessRetryExecutor`,只允许 Router/SemanticGuard 的类型化技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 固定一次 attempt。
|
||||||
|
- 将 Spring AI 1.1.7 全局底层重试 `spring.ai.retry.max-attempts` 从默认 10 压为 1,避免与 Harness 形成隐藏嵌套重试。
|
||||||
|
- 新增 Redis Tool Call Key Factory,只负责安全构造固定前缀下的 `runId + toolCallId` Key,不访问 Redis、不生成 Tool Call ID。
|
||||||
|
- 使用 Fake Model/Tool 和可控 Clock 验证显式异步传播、deadline、客户端取消、预算耗尽、重试记录、Key/容量边界和唯一 Run 终态。
|
||||||
|
- 本阶段不改 Controller/HTTP/SSE,不接入或临时适配旧 ChatService,不修改 `diagnosis_run` 持久化,也不实现 Tool-specific 投影或 Redis store。
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `diagnosis-harness-run-context`: 提供显式 RunContext、Harness Core、预算、取消、deadline、类型化重试、唯一终态和后续 Tool store 所需的 Key/容量基础。
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- None. 本阶段不声明旧 Chat/AIOps 运行路径已迁移;公开运行和数据库状态映射在后续 change 接入。
|
||||||
|
|
||||||
|
## Context Constraints
|
||||||
|
|
||||||
|
- 新代码不得读取或写入任何 ThreadLocal;RunContext 只能通过参数或框架受控 context 显式传播。
|
||||||
|
- RunContext 的结构不可变,但其 cancellation、budget 和 lifecycle 是线程安全的单 Run 状态句柄。
|
||||||
|
- 同一 Run 只允许第一个终态生效;取消、deadline、预算和异常之间的竞态不得覆盖先到终态。
|
||||||
|
- 同步模型/Tool 调用的取消只保证阻止后续执行并触发已注册资源取消回调,不虚假承诺无法证明的立即线程中断。
|
||||||
|
- `NO_EVIDENCE`、业务拒绝、取消和预算耗尽不是可重试技术失败。
|
||||||
|
- Tool Call Key 使用框架 `tool_call_id`;Key Factory 不生成、替换或回退到其他 ID。
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- 等级:L2(内部 Harness 基础接口)。新增 Core 类型将被后续 3A-7 阶段消费,但本 change 不修改现有调用方。
|
||||||
|
- `application.yml` 的 Spring AI retry 从默认 10 attempts 变为 1 attempt,属于有意的内部运行配置变化;当前旧调用若遇到瞬时模型失败将不再由 SDK 隐式重试,避免未记录 attempts。Harness 允许的 Router/SemanticGuard 重试要到后续接入后显式执行和记录。
|
||||||
|
- 不改变 Controller、SSE、DTO、数据库 Schema 或公开错误码。
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- 旧运行路径在切换前会失去 SDK 隐式重试但尚未使用 Harness retry;这是为保证“所有 attempt 可观测”接受的短期行为变化,focused tests 需证明启动配置正确。
|
||||||
|
- RunContext 内含线程安全可变句柄,若被误解为纯值对象可能错误复制;设计和测试必须明确同一 Run 共享同一状态句柄。
|
||||||
|
- Budget 在模型调用后才能取得实际 Token,可能在记录实际消耗时才发现超限;Core 必须保留实际 usage 并立即终止后续执行。
|
||||||
|
- 本阶段不写数据库,内存生命周期只服务一次调用链;持久化终态映射由 Chat application use case 阶段负责。
|
||||||
+89
@@ -0,0 +1,89 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: RunContext SHALL be explicit and structurally immutable
|
||||||
|
The Harness SHALL create a structurally immutable RunContext containing non-blank `sessionId`, non-blank `runId`, deadline, cancellation, budget, retry policies, and lifecycle. Every new Model, Tool, Guard, and application boundary SHALL receive the RunContext explicitly by parameter or framework-controlled context and SHALL NOT read or write ThreadLocal state.
|
||||||
|
|
||||||
|
#### Scenario: Context crosses an asynchronous boundary
|
||||||
|
- **WHEN** a Fake Tool runs on another thread with an explicitly supplied RunContext
|
||||||
|
- **THEN** it observes the same session ID, run ID, deadline, cancellation, budget, retry policies, and lifecycle handles without copying ThreadLocal state
|
||||||
|
|
||||||
|
#### Scenario: Invalid identity starts a Run
|
||||||
|
- **WHEN** Run creation receives a blank session ID or run ID
|
||||||
|
- **THEN** the Core rejects creation before any budget, cancellation, or lifecycle state is published
|
||||||
|
|
||||||
|
### Requirement: Deadline and cancellation SHALL stop subsequent controlled work
|
||||||
|
The Core SHALL compare the current Clock instant with the Run deadline before each controlled Model, Tool, retry, and capacity boundary. At or after the deadline it SHALL produce `TIMED_OUT`, signal `DEADLINE_EXCEEDED`, and reject subsequent work. Explicit cancellation SHALL preserve the first cancellation reason, notify all registered resource callbacks, and produce `CANCELLED` unless an earlier terminal state exists.
|
||||||
|
|
||||||
|
#### Scenario: Deadline expires before a Tool call
|
||||||
|
- **WHEN** the controlled Clock reaches the Run deadline before the Fake Tool boundary
|
||||||
|
- **THEN** the Tool is not invoked, the Run terminates as `TIMED_OUT`, and later checks return the same terminal outcome
|
||||||
|
|
||||||
|
#### Scenario: Client disconnects
|
||||||
|
- **WHEN** the Core receives `CLIENT_DISCONNECTED`
|
||||||
|
- **THEN** registered cancellation callbacks run, subsequent Model/Tool boundaries are rejected, and the Run has one `CANCELLED` terminal outcome
|
||||||
|
|
||||||
|
### Requirement: Run budgets SHALL be explicit, thread-safe, and observable
|
||||||
|
The Core SHALL require caller-supplied positive limits for Model calls, total Tool calls, per-Tool calls, input Tokens, output Tokens, total Tokens, and Run bytes. Budget updates SHALL be thread-safe and preserve a usage snapshot. Exceeding any limit SHALL record the actual applicable usage, produce `BUDGET_EXHAUSTED`, signal cancellation, and prevent subsequent controlled work.
|
||||||
|
|
||||||
|
#### Scenario: Per-Tool budget is exhausted
|
||||||
|
- **WHEN** a Fake Tool attempts one call beyond its configured per-Tool limit
|
||||||
|
- **THEN** the extra call is not reserved, the total and per-Tool usage remain internally consistent, and the Run terminates as `BUDGET_EXHAUSTED`
|
||||||
|
|
||||||
|
#### Scenario: Actual Token usage exceeds the limit
|
||||||
|
- **WHEN** a completed Fake Model response reports Token usage above the configured limit
|
||||||
|
- **THEN** the actual input/output/total usage remains recorded and the Core rejects the next controlled operation
|
||||||
|
|
||||||
|
### Requirement: A Run SHALL have exactly one terminal outcome
|
||||||
|
The lifecycle SHALL start as `RUNNING` and accept only the first terminal outcome from `SUCCESS`, `FAILED`, `CANCELLED`, `TIMED_OUT`, or `BUDGET_EXHAUSTED`. Later success, failure, cancellation, deadline, or budget events SHALL NOT replace the first `RunTermination` state, reason, or timestamp.
|
||||||
|
|
||||||
|
#### Scenario: Success wins a terminal race
|
||||||
|
- **WHEN** success is recorded before a later failure and cancellation
|
||||||
|
- **THEN** the lifecycle remains `SUCCESS` and both later finish attempts report that they did not change the terminal outcome
|
||||||
|
|
||||||
|
#### Scenario: Budget exhaustion wins a terminal race
|
||||||
|
- **WHEN** budget exhaustion is recorded before an exception handler reports failure
|
||||||
|
- **THEN** the lifecycle remains `BUDGET_EXHAUSTED` with its original reason and completion time
|
||||||
|
|
||||||
|
### Requirement: Harness retry SHALL be typed, bounded, and recorded
|
||||||
|
The Harness SHALL provide immutable policies and an executor that records every actual attempt. Intent Router and SemanticGuard SHALL allow at most two attempts for their listed technical failures; Diagnosis Agent, Tool calls, and Evidence repair SHALL allow one attempt. Cancellation, deadline, budget exhaustion, `NO_EVIDENCE`, business rejection, and unknown failures SHALL NOT be retried.
|
||||||
|
|
||||||
|
#### Scenario: Router transport failure succeeds on retry
|
||||||
|
- **WHEN** a Fake Router fails its first attempt with a retryable transport failure and succeeds on the second
|
||||||
|
- **THEN** the executor returns the second result and records exactly one failed and one successful attempt
|
||||||
|
|
||||||
|
#### Scenario: Tool call fails
|
||||||
|
- **WHEN** a Fake Tool operation fails on its first attempt
|
||||||
|
- **THEN** the Tool policy records one failed attempt and throws without invoking the Tool again
|
||||||
|
|
||||||
|
#### Scenario: Run is cancelled between attempts
|
||||||
|
- **WHEN** cancellation is signalled after a retryable first failure
|
||||||
|
- **THEN** the executor rejects the next attempt and does not invoke the operation again
|
||||||
|
|
||||||
|
### Requirement: Spring AI hidden retry SHALL be disabled
|
||||||
|
Application configuration SHALL set `spring.ai.retry.max-attempts=1`, overriding the Spring AI 1.1.7 default of 10. Harness-permitted retries SHALL occur only in HarnessRetryExecutor and SHALL be observable as separate attempts.
|
||||||
|
|
||||||
|
#### Scenario: Retry configuration is loaded
|
||||||
|
- **WHEN** the repository application YAML is parsed in a focused configuration test
|
||||||
|
- **THEN** `spring.ai.retry.max-attempts` equals 1
|
||||||
|
|
||||||
|
### Requirement: Tool invocation key and Run capacity primitives SHALL be safe and store-independent
|
||||||
|
The Harness SHALL provide a ToolCallKeyFactory that constructs `{prefix}:{runId}:{toolCallId}` only from valid safe segments and never generates or rewrites the framework Tool Call ID. The Run capacity counter SHALL atomically reserve positive bytes up to its configured limit and SHALL leave usage unchanged when a reservation is rejected.
|
||||||
|
|
||||||
|
#### Scenario: Framework Tool Call ID is used in a key
|
||||||
|
- **WHEN** a valid run ID and framework Tool Call ID are supplied
|
||||||
|
- **THEN** the factory returns the configured prefix followed by the exact run ID and exact Tool Call ID
|
||||||
|
|
||||||
|
#### Scenario: Unsafe key segment is supplied
|
||||||
|
- **WHEN** either ID is blank, too long, or contains a separator or unsafe character
|
||||||
|
- **THEN** the factory rejects it and does not produce a Redis key
|
||||||
|
|
||||||
|
#### Scenario: Run byte limit would be exceeded
|
||||||
|
- **WHEN** a reservation would exceed max Run bytes
|
||||||
|
- **THEN** the counter rejects it and retains the usage from previously successful reservations
|
||||||
|
|
||||||
|
### Requirement: Harness Core SHALL remain independent from current runtime and persistence
|
||||||
|
This change SHALL NOT modify Controller/HTTP/SSE contracts, connect RunContext to the old ChatService or AiOpsService, use current ThreadLocal holders, persist `diagnosis_run`, access Redis, or implement Tool-specific projection. The new Core SHALL compile and pass focused Fake Model/Tool tests as a zero-consumer foundation for stage 3A.
|
||||||
|
|
||||||
|
#### Scenario: Stage 2 implementation completes
|
||||||
|
- **WHEN** focused Core, retry, budget, key, capacity, and configuration tests pass
|
||||||
|
- **THEN** current Chat/AIOps call sites remain unchanged and stage 3A can depend directly on the new Harness types
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
## 1. Run State Primitives
|
||||||
|
|
||||||
|
- [x] 1.1 Implement structurally immutable RunContext with explicit identity, deadline, cancellation, budget, retry policy, and lifecycle handles.
|
||||||
|
- [x] 1.2 Implement first-reason-wins cancellation with resource callbacks and first-terminal-wins Run lifecycle.
|
||||||
|
- [x] 1.3 Add Run identity, cancellation, and terminal race tests with a controlled Clock.
|
||||||
|
|
||||||
|
## 2. Budget and Capacity
|
||||||
|
|
||||||
|
- [x] 2.1 Implement validated caller-supplied RunBudgetLimits and thread-safe Model/Tool/per-Tool/Token accounting.
|
||||||
|
- [x] 2.2 Implement atomic single-Run byte capacity reservation with no partial update on rejection.
|
||||||
|
- [x] 2.3 Add budget tests for consistent counters, actual Token recording, concurrent reservations, and exhaustion details.
|
||||||
|
|
||||||
|
## 3. Harness Core
|
||||||
|
|
||||||
|
- [x] 3.1 Implement DiagnosisHarnessCore Run creation, active/deadline checks, budget gates, explicit cancellation, and success/failure completion.
|
||||||
|
- [x] 3.2 Add Fake Model/Tool tests proving explicit synchronous/asynchronous RunContext propagation and cancellation/deadline/budget terminal outcomes.
|
||||||
|
|
||||||
|
## 4. Retry Boundary
|
||||||
|
|
||||||
|
- [x] 4.1 Implement typed RetryFailure, bounded RetryPolicy, strict HarnessRetryPolicies, RetryAttempt, and RetryExecutionException.
|
||||||
|
- [x] 4.2 Implement HarnessRetryExecutor with per-attempt active checks and attempt recording, without backoff, recursion, or hidden loops.
|
||||||
|
- [x] 4.3 Add Fake Router/Tool/SemanticGuard retry tests for allowed technical retry, one-attempt policies, cancellation between attempts, and non-retryable outcomes.
|
||||||
|
|
||||||
|
## 5. Tool Store Foundations
|
||||||
|
|
||||||
|
- [x] 5.1 Implement configurable ToolCallKeyFactory with exact framework ID preservation and safe-segment validation.
|
||||||
|
- [x] 5.2 Add key factory tests covering exact key format, blank/unsafe/oversized segments, and no generated fallback ID.
|
||||||
|
|
||||||
|
## 6. Hidden Retry Configuration
|
||||||
|
|
||||||
|
- [x] 6.1 Set `spring.ai.retry.max-attempts` to 1 without changing Model routing or provider configuration.
|
||||||
|
- [x] 6.2 Add a focused YAML configuration test proving the hidden Spring AI retry override.
|
||||||
|
|
||||||
|
## 7. Verification
|
||||||
|
|
||||||
|
- [x] 7.1 Run focused Harness Core, retry, budget, key/capacity, and retry configuration tests.
|
||||||
|
- [x] 7.2 Run stage 0/1 Harness and ACI contract tests plus existing ChatController test as regression coverage.
|
||||||
|
- [x] 7.3 Verify new production code has no ThreadLocal/current-holder usage and current Chat/AIOps/Controller/JPA/Redis call sites remain unchanged.
|
||||||
@@ -0,0 +1,93 @@
|
|||||||
|
# diagnosis-harness-run-context Specification
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
定义 Diagnosis Harness 的显式 RunContext、deadline、协作式取消、线程安全预算、唯一 Run 终态、类型化重试、Tool Call Key 和单 Run 容量基础,供后续 Tool boundary、Agent、Guard 和应用入口复用。
|
||||||
|
|
||||||
|
## Requirements
|
||||||
|
### Requirement: RunContext SHALL be explicit and structurally immutable
|
||||||
|
The Harness SHALL create a structurally immutable RunContext containing non-blank `sessionId`, non-blank `runId`, deadline, cancellation, budget, retry policies, and lifecycle. Every new Model, Tool, Guard, and application boundary SHALL receive the RunContext explicitly by parameter or framework-controlled context and SHALL NOT read or write ThreadLocal state.
|
||||||
|
|
||||||
|
#### Scenario: Context crosses an asynchronous boundary
|
||||||
|
- **WHEN** a Fake Tool runs on another thread with an explicitly supplied RunContext
|
||||||
|
- **THEN** it observes the same session ID, run ID, deadline, cancellation, budget, retry policies, and lifecycle handles without copying ThreadLocal state
|
||||||
|
|
||||||
|
#### Scenario: Invalid identity starts a Run
|
||||||
|
- **WHEN** Run creation receives a blank session ID or run ID
|
||||||
|
- **THEN** the Core rejects creation before any budget, cancellation, or lifecycle state is published
|
||||||
|
|
||||||
|
### Requirement: Deadline and cancellation SHALL stop subsequent controlled work
|
||||||
|
The Core SHALL compare the current Clock instant with the Run deadline before each controlled Model, Tool, retry, and capacity boundary. At or after the deadline it SHALL produce `TIMED_OUT`, signal `DEADLINE_EXCEEDED`, and reject subsequent work. Explicit cancellation SHALL preserve the first cancellation reason, notify all registered resource callbacks, and produce `CANCELLED` unless an earlier terminal state exists.
|
||||||
|
|
||||||
|
#### Scenario: Deadline expires before a Tool call
|
||||||
|
- **WHEN** the controlled Clock reaches the Run deadline before the Fake Tool boundary
|
||||||
|
- **THEN** the Tool is not invoked, the Run terminates as `TIMED_OUT`, and later checks return the same terminal outcome
|
||||||
|
|
||||||
|
#### Scenario: Client disconnects
|
||||||
|
- **WHEN** the Core receives `CLIENT_DISCONNECTED`
|
||||||
|
- **THEN** registered cancellation callbacks run, subsequent Model/Tool boundaries are rejected, and the Run has one `CANCELLED` terminal outcome
|
||||||
|
|
||||||
|
### Requirement: Run budgets SHALL be explicit, thread-safe, and observable
|
||||||
|
The Core SHALL require caller-supplied positive limits for Model calls, total Tool calls, per-Tool calls, input Tokens, output Tokens, total Tokens, and Run bytes. Budget updates SHALL be thread-safe and preserve a usage snapshot. Exceeding any limit SHALL record the actual applicable usage, produce `BUDGET_EXHAUSTED`, signal cancellation, and prevent subsequent controlled work.
|
||||||
|
|
||||||
|
#### Scenario: Per-Tool budget is exhausted
|
||||||
|
- **WHEN** a Fake Tool attempts one call beyond its configured per-Tool limit
|
||||||
|
- **THEN** the extra call is not reserved, the total and per-Tool usage remain internally consistent, and the Run terminates as `BUDGET_EXHAUSTED`
|
||||||
|
|
||||||
|
#### Scenario: Actual Token usage exceeds the limit
|
||||||
|
- **WHEN** a completed Fake Model response reports Token usage above the configured limit
|
||||||
|
- **THEN** the actual input/output/total usage remains recorded and the Core rejects the next controlled operation
|
||||||
|
|
||||||
|
### Requirement: A Run SHALL have exactly one terminal outcome
|
||||||
|
The lifecycle SHALL start as `RUNNING` and accept only the first terminal outcome from `SUCCESS`, `FAILED`, `CANCELLED`, `TIMED_OUT`, or `BUDGET_EXHAUSTED`. Later success, failure, cancellation, deadline, or budget events SHALL NOT replace the first `RunTermination` state, reason, or timestamp.
|
||||||
|
|
||||||
|
#### Scenario: Success wins a terminal race
|
||||||
|
- **WHEN** success is recorded before a later failure and cancellation
|
||||||
|
- **THEN** the lifecycle remains `SUCCESS` and both later finish attempts report that they did not change the terminal outcome
|
||||||
|
|
||||||
|
#### Scenario: Budget exhaustion wins a terminal race
|
||||||
|
- **WHEN** budget exhaustion is recorded before an exception handler reports failure
|
||||||
|
- **THEN** the lifecycle remains `BUDGET_EXHAUSTED` with its original reason and completion time
|
||||||
|
|
||||||
|
### Requirement: Harness retry SHALL be typed, bounded, and recorded
|
||||||
|
The Harness SHALL provide immutable policies and an executor that records every actual attempt. Intent Router and SemanticGuard SHALL allow at most two attempts for their listed technical failures; Diagnosis Agent, Tool calls, and Evidence repair SHALL allow one attempt. Cancellation, deadline, budget exhaustion, `NO_EVIDENCE`, business rejection, and unknown failures SHALL NOT be retried.
|
||||||
|
|
||||||
|
#### Scenario: Router transport failure succeeds on retry
|
||||||
|
- **WHEN** a Fake Router fails its first attempt with a retryable transport failure and succeeds on the second
|
||||||
|
- **THEN** the executor returns the second result and records exactly one failed and one successful attempt
|
||||||
|
|
||||||
|
#### Scenario: Tool call fails
|
||||||
|
- **WHEN** a Fake Tool operation fails on its first attempt
|
||||||
|
- **THEN** the Tool policy records one failed attempt and throws without invoking the Tool again
|
||||||
|
|
||||||
|
#### Scenario: Run is cancelled between attempts
|
||||||
|
- **WHEN** cancellation is signalled after a retryable first failure
|
||||||
|
- **THEN** the executor rejects the next attempt and does not invoke the operation again
|
||||||
|
|
||||||
|
### Requirement: Spring AI hidden retry SHALL be disabled
|
||||||
|
Application configuration SHALL set `spring.ai.retry.max-attempts=1`, overriding the Spring AI 1.1.7 default of 10. Harness-permitted retries SHALL occur only in HarnessRetryExecutor and SHALL be observable as separate attempts.
|
||||||
|
|
||||||
|
#### Scenario: Retry configuration is loaded
|
||||||
|
- **WHEN** the repository application YAML is parsed in a focused configuration test
|
||||||
|
- **THEN** `spring.ai.retry.max-attempts` equals 1
|
||||||
|
|
||||||
|
### Requirement: Tool invocation key and Run capacity primitives SHALL be safe and store-independent
|
||||||
|
The Harness SHALL provide a ToolCallKeyFactory that constructs `{prefix}:{runId}:{toolCallId}` only from valid safe segments and never generates or rewrites the framework Tool Call ID. The Run capacity counter SHALL atomically reserve positive bytes up to its configured limit and SHALL leave usage unchanged when a reservation is rejected.
|
||||||
|
|
||||||
|
#### Scenario: Framework Tool Call ID is used in a key
|
||||||
|
- **WHEN** a valid run ID and framework Tool Call ID are supplied
|
||||||
|
- **THEN** the factory returns the configured prefix followed by the exact run ID and exact Tool Call ID
|
||||||
|
|
||||||
|
#### Scenario: Unsafe key segment is supplied
|
||||||
|
- **WHEN** either ID is blank, too long, or contains a separator or unsafe character
|
||||||
|
- **THEN** the factory rejects it and does not produce a Redis key
|
||||||
|
|
||||||
|
#### Scenario: Run byte limit would be exceeded
|
||||||
|
- **WHEN** a reservation would exceed max Run bytes
|
||||||
|
- **THEN** the counter rejects it and retains the usage from previously successful reservations
|
||||||
|
|
||||||
|
### Requirement: Harness Core SHALL remain independent from current runtime and persistence
|
||||||
|
This change SHALL NOT modify Controller/HTTP/SSE contracts, connect RunContext to the old ChatService or AiOpsService, use current ThreadLocal holders, persist `diagnosis_run`, access Redis, or implement Tool-specific projection. The new Core SHALL compile and pass focused Fake Model/Tool tests as a zero-consumer foundation for stage 3A.
|
||||||
|
|
||||||
|
#### Scenario: Stage 2 implementation completes
|
||||||
|
- **WHEN** focused Core, retry, budget, key, capacity, and configuration tests pass
|
||||||
|
- **THEN** current Chat/AIOps call sites remain unchanged and stage 3A can depend directly on the new Harness types
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
public final class BudgetExceededException extends RuntimeException {
|
||||||
|
|
||||||
|
private final BudgetKind kind;
|
||||||
|
private final long limit;
|
||||||
|
private final long attempted;
|
||||||
|
|
||||||
|
public BudgetExceededException(BudgetKind kind, long limit, long attempted) {
|
||||||
|
super("Budget exceeded: kind=" + kind + ", limit=" + limit + ", attempted=" + attempted);
|
||||||
|
this.kind = kind;
|
||||||
|
this.limit = limit;
|
||||||
|
this.attempted = attempted;
|
||||||
|
}
|
||||||
|
|
||||||
|
public BudgetKind kind() {
|
||||||
|
return kind;
|
||||||
|
}
|
||||||
|
|
||||||
|
public long limit() {
|
||||||
|
return limit;
|
||||||
|
}
|
||||||
|
|
||||||
|
public long attempted() {
|
||||||
|
return attempted;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
public enum BudgetKind {
|
||||||
|
MODEL_CALLS,
|
||||||
|
TOOL_CALLS,
|
||||||
|
TOOL_CALLS_PER_TOOL,
|
||||||
|
INPUT_TOKENS,
|
||||||
|
OUTPUT_TOKENS,
|
||||||
|
TOTAL_TOKENS,
|
||||||
|
RUN_BYTES
|
||||||
|
}
|
||||||
@@ -0,0 +1,147 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
|
||||||
|
|
||||||
|
import java.time.Clock;
|
||||||
|
import java.time.Duration;
|
||||||
|
import java.time.Instant;
|
||||||
|
import java.util.Objects;
|
||||||
|
import java.util.function.Supplier;
|
||||||
|
|
||||||
|
public final class DiagnosisHarnessCore {
|
||||||
|
|
||||||
|
private final Clock clock;
|
||||||
|
private final Supplier<String> runIdSupplier;
|
||||||
|
private final Duration maxRunDuration;
|
||||||
|
private final RunBudgetLimits budgetLimits;
|
||||||
|
private final HarnessRetryPolicies retryPolicies;
|
||||||
|
|
||||||
|
public DiagnosisHarnessCore(Clock clock,
|
||||||
|
Supplier<String> runIdSupplier,
|
||||||
|
Duration maxRunDuration,
|
||||||
|
RunBudgetLimits budgetLimits,
|
||||||
|
HarnessRetryPolicies retryPolicies) {
|
||||||
|
this.clock = Objects.requireNonNull(clock, "clock must not be null");
|
||||||
|
this.runIdSupplier = Objects.requireNonNull(runIdSupplier, "runIdSupplier must not be null");
|
||||||
|
this.maxRunDuration = requirePositive(maxRunDuration, "maxRunDuration");
|
||||||
|
this.budgetLimits = Objects.requireNonNull(budgetLimits, "budgetLimits must not be null");
|
||||||
|
this.retryPolicies = Objects.requireNonNull(retryPolicies, "retryPolicies must not be null");
|
||||||
|
}
|
||||||
|
|
||||||
|
public RunContext startRun(String sessionId) {
|
||||||
|
return startRun(sessionId, runIdSupplier.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
public RunContext startRun(String sessionId, String runId) {
|
||||||
|
Instant deadline = clock.instant().plus(maxRunDuration);
|
||||||
|
RunCancellation cancellation = new RunCancellation();
|
||||||
|
RunLifecycle lifecycle = new RunLifecycle(clock);
|
||||||
|
RunContext context = new RunContext(
|
||||||
|
sessionId,
|
||||||
|
runId,
|
||||||
|
deadline,
|
||||||
|
cancellation,
|
||||||
|
new RunBudget(budgetLimits),
|
||||||
|
retryPolicies,
|
||||||
|
lifecycle);
|
||||||
|
cancellation.onCancel(reason -> lifecycle.finish(terminalState(reason), reason.name()));
|
||||||
|
return context;
|
||||||
|
}
|
||||||
|
|
||||||
|
public void checkActive(RunContext context) {
|
||||||
|
Objects.requireNonNull(context, "context must not be null");
|
||||||
|
context.lifecycle().termination().ifPresent(termination -> {
|
||||||
|
throw new RunAbortedException(termination);
|
||||||
|
});
|
||||||
|
|
||||||
|
if (!clock.instant().isBefore(context.deadline())) {
|
||||||
|
context.lifecycle().finish(RunState.TIMED_OUT, RunCancellationReason.DEADLINE_EXCEEDED.name());
|
||||||
|
context.cancellation().cancel(RunCancellationReason.DEADLINE_EXCEEDED);
|
||||||
|
throw new RunAbortedException(context.lifecycle().termination().orElseThrow());
|
||||||
|
}
|
||||||
|
|
||||||
|
if (context.cancellation().isCancelled()) {
|
||||||
|
throw new RunAbortedException(context.lifecycle().termination().orElseThrow());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
public void beforeModelCall(RunContext context) {
|
||||||
|
checkActive(context);
|
||||||
|
applyBudget(context, context.budget()::reserveModelCall);
|
||||||
|
}
|
||||||
|
|
||||||
|
public void beforeToolCall(RunContext context, String toolName) {
|
||||||
|
checkActive(context);
|
||||||
|
applyBudget(context, () -> context.budget().reserveToolCall(toolName));
|
||||||
|
}
|
||||||
|
|
||||||
|
public void recordTokens(RunContext context, long inputTokens, long outputTokens) {
|
||||||
|
Objects.requireNonNull(context, "context must not be null");
|
||||||
|
applyBudget(context, () -> context.budget().recordTokens(inputTokens, outputTokens));
|
||||||
|
}
|
||||||
|
|
||||||
|
public long reserveRunBytes(RunContext context, long bytes) {
|
||||||
|
checkActive(context);
|
||||||
|
try {
|
||||||
|
return context.budget().reserveRunBytes(bytes);
|
||||||
|
} catch (BudgetExceededException e) {
|
||||||
|
exhaustBudget(context, e);
|
||||||
|
throw e;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
public boolean cancel(RunContext context, RunCancellationReason reason) {
|
||||||
|
Objects.requireNonNull(context, "context must not be null");
|
||||||
|
Objects.requireNonNull(reason, "reason must not be null");
|
||||||
|
if (context.lifecycle().termination().isPresent()) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
return context.cancellation().cancel(reason);
|
||||||
|
}
|
||||||
|
|
||||||
|
public boolean completeSuccess(RunContext context) {
|
||||||
|
Objects.requireNonNull(context, "context must not be null");
|
||||||
|
return context.lifecycle().finish(RunState.SUCCESS, "COMPLETED");
|
||||||
|
}
|
||||||
|
|
||||||
|
public boolean completeFailure(RunContext context, String reason) {
|
||||||
|
Objects.requireNonNull(context, "context must not be null");
|
||||||
|
String safeReason = reason == null || reason.isBlank() ? "INTERNAL_FAILURE" : reason;
|
||||||
|
boolean completed = context.lifecycle().finish(RunState.FAILED, safeReason);
|
||||||
|
if (completed) {
|
||||||
|
context.cancellation().cancel(RunCancellationReason.INTERNAL_FAILURE);
|
||||||
|
}
|
||||||
|
return completed;
|
||||||
|
}
|
||||||
|
|
||||||
|
private void applyBudget(RunContext context, Runnable operation) {
|
||||||
|
try {
|
||||||
|
operation.run();
|
||||||
|
} catch (BudgetExceededException e) {
|
||||||
|
exhaustBudget(context, e);
|
||||||
|
throw e;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private void exhaustBudget(RunContext context, BudgetExceededException exception) {
|
||||||
|
context.lifecycle().finish(RunState.BUDGET_EXHAUSTED, exception.getMessage());
|
||||||
|
context.cancellation().cancel(RunCancellationReason.BUDGET_EXHAUSTED);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static RunState terminalState(RunCancellationReason reason) {
|
||||||
|
return switch (reason) {
|
||||||
|
case DEADLINE_EXCEEDED -> RunState.TIMED_OUT;
|
||||||
|
case BUDGET_EXHAUSTED -> RunState.BUDGET_EXHAUSTED;
|
||||||
|
case INTERNAL_FAILURE -> RunState.FAILED;
|
||||||
|
case CLIENT_DISCONNECTED, USER_REQUESTED -> RunState.CANCELLED;
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
private static Duration requirePositive(Duration duration, String name) {
|
||||||
|
Objects.requireNonNull(duration, name + " must not be null");
|
||||||
|
if (duration.isZero() || duration.isNegative()) {
|
||||||
|
throw new IllegalArgumentException(name + " must be positive");
|
||||||
|
}
|
||||||
|
return duration;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
public final class RunAbortedException extends RuntimeException {
|
||||||
|
|
||||||
|
private final RunTermination termination;
|
||||||
|
|
||||||
|
public RunAbortedException(RunTermination termination) {
|
||||||
|
super("Run is no longer active: state=" + Objects.requireNonNull(termination).state()
|
||||||
|
+ ", reason=" + termination.reason());
|
||||||
|
this.termination = termination;
|
||||||
|
}
|
||||||
|
|
||||||
|
public RunTermination termination() {
|
||||||
|
return termination;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,98 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import java.util.HashMap;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
public final class RunBudget {
|
||||||
|
|
||||||
|
private final RunBudgetLimits limits;
|
||||||
|
private final RunCapacityCounter capacity;
|
||||||
|
private final Map<String, Integer> toolCallsByName = new HashMap<>();
|
||||||
|
|
||||||
|
private int modelCalls;
|
||||||
|
private int toolCalls;
|
||||||
|
private long inputTokens;
|
||||||
|
private long outputTokens;
|
||||||
|
private long totalTokens;
|
||||||
|
|
||||||
|
public RunBudget(RunBudgetLimits limits) {
|
||||||
|
this.limits = Objects.requireNonNull(limits, "limits must not be null");
|
||||||
|
this.capacity = new RunCapacityCounter(limits.maxRunBytes());
|
||||||
|
}
|
||||||
|
|
||||||
|
public synchronized void reserveModelCall() {
|
||||||
|
int attempted = modelCalls + 1;
|
||||||
|
if (attempted > limits.maxModelCalls()) {
|
||||||
|
throw new BudgetExceededException(BudgetKind.MODEL_CALLS, limits.maxModelCalls(), attempted);
|
||||||
|
}
|
||||||
|
modelCalls = attempted;
|
||||||
|
}
|
||||||
|
|
||||||
|
public synchronized void reserveToolCall(String toolName) {
|
||||||
|
if (toolName == null || toolName.isBlank()) {
|
||||||
|
throw new IllegalArgumentException("toolName must not be blank");
|
||||||
|
}
|
||||||
|
int attemptedTotal = toolCalls + 1;
|
||||||
|
int attemptedForTool = toolCallsByName.getOrDefault(toolName, 0) + 1;
|
||||||
|
if (attemptedTotal > limits.maxToolCalls()) {
|
||||||
|
throw new BudgetExceededException(BudgetKind.TOOL_CALLS, limits.maxToolCalls(), attemptedTotal);
|
||||||
|
}
|
||||||
|
if (attemptedForTool > limits.maxCallsPerTool()) {
|
||||||
|
throw new BudgetExceededException(
|
||||||
|
BudgetKind.TOOL_CALLS_PER_TOOL, limits.maxCallsPerTool(), attemptedForTool);
|
||||||
|
}
|
||||||
|
toolCalls = attemptedTotal;
|
||||||
|
toolCallsByName.put(toolName, attemptedForTool);
|
||||||
|
}
|
||||||
|
|
||||||
|
public synchronized void recordTokens(long input, long output) {
|
||||||
|
if (input < 0 || output < 0) {
|
||||||
|
throw new IllegalArgumentException("token counts must not be negative");
|
||||||
|
}
|
||||||
|
inputTokens = safeAdd(inputTokens, input);
|
||||||
|
outputTokens = safeAdd(outputTokens, output);
|
||||||
|
totalTokens = safeAdd(totalTokens, safeAdd(input, output));
|
||||||
|
|
||||||
|
if (inputTokens > limits.maxInputTokens()) {
|
||||||
|
throw new BudgetExceededException(BudgetKind.INPUT_TOKENS, limits.maxInputTokens(), inputTokens);
|
||||||
|
}
|
||||||
|
if (outputTokens > limits.maxOutputTokens()) {
|
||||||
|
throw new BudgetExceededException(BudgetKind.OUTPUT_TOKENS, limits.maxOutputTokens(), outputTokens);
|
||||||
|
}
|
||||||
|
if (totalTokens > limits.maxTotalTokens()) {
|
||||||
|
throw new BudgetExceededException(BudgetKind.TOTAL_TOKENS, limits.maxTotalTokens(), totalTokens);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
public long reserveRunBytes(long bytes) {
|
||||||
|
return capacity.reserve(bytes);
|
||||||
|
}
|
||||||
|
|
||||||
|
public synchronized RunBudgetUsage snapshot() {
|
||||||
|
return new RunBudgetUsage(
|
||||||
|
modelCalls,
|
||||||
|
toolCalls,
|
||||||
|
toolCallsByName,
|
||||||
|
inputTokens,
|
||||||
|
outputTokens,
|
||||||
|
totalTokens,
|
||||||
|
capacity.usedBytes());
|
||||||
|
}
|
||||||
|
|
||||||
|
public RunBudgetLimits limits() {
|
||||||
|
return limits;
|
||||||
|
}
|
||||||
|
|
||||||
|
public RunCapacityCounter capacity() {
|
||||||
|
return capacity;
|
||||||
|
}
|
||||||
|
|
||||||
|
private static long safeAdd(long left, long right) {
|
||||||
|
try {
|
||||||
|
return Math.addExact(left, right);
|
||||||
|
} catch (ArithmeticException e) {
|
||||||
|
return Long.MAX_VALUE;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
public record RunBudgetLimits(
|
||||||
|
int maxModelCalls,
|
||||||
|
int maxToolCalls,
|
||||||
|
int maxCallsPerTool,
|
||||||
|
long maxInputTokens,
|
||||||
|
long maxOutputTokens,
|
||||||
|
long maxTotalTokens,
|
||||||
|
long maxRunBytes) {
|
||||||
|
|
||||||
|
public RunBudgetLimits {
|
||||||
|
requirePositive(maxModelCalls, "maxModelCalls");
|
||||||
|
requirePositive(maxToolCalls, "maxToolCalls");
|
||||||
|
requirePositive(maxCallsPerTool, "maxCallsPerTool");
|
||||||
|
requirePositive(maxInputTokens, "maxInputTokens");
|
||||||
|
requirePositive(maxOutputTokens, "maxOutputTokens");
|
||||||
|
requirePositive(maxTotalTokens, "maxTotalTokens");
|
||||||
|
requirePositive(maxRunBytes, "maxRunBytes");
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void requirePositive(long value, String name) {
|
||||||
|
if (value <= 0) {
|
||||||
|
throw new IllegalArgumentException(name + " must be positive");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
public record RunBudgetUsage(
|
||||||
|
int modelCalls,
|
||||||
|
int toolCalls,
|
||||||
|
Map<String, Integer> toolCallsByName,
|
||||||
|
long inputTokens,
|
||||||
|
long outputTokens,
|
||||||
|
long totalTokens,
|
||||||
|
long runBytes) {
|
||||||
|
|
||||||
|
public RunBudgetUsage {
|
||||||
|
toolCallsByName = toolCallsByName == null ? Map.of() : Map.copyOf(toolCallsByName);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,64 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import org.slf4j.Logger;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.util.Objects;
|
||||||
|
import java.util.Optional;
|
||||||
|
import java.util.concurrent.CopyOnWriteArrayList;
|
||||||
|
import java.util.concurrent.atomic.AtomicBoolean;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
import java.util.function.Consumer;
|
||||||
|
|
||||||
|
public final class RunCancellation {
|
||||||
|
|
||||||
|
private static final Logger log = LoggerFactory.getLogger(RunCancellation.class);
|
||||||
|
|
||||||
|
private final AtomicReference<RunCancellationReason> reason = new AtomicReference<>();
|
||||||
|
private final CopyOnWriteArrayList<Consumer<RunCancellationReason>> callbacks =
|
||||||
|
new CopyOnWriteArrayList<>();
|
||||||
|
|
||||||
|
public boolean cancel(RunCancellationReason cancellationReason) {
|
||||||
|
Objects.requireNonNull(cancellationReason, "cancellationReason must not be null");
|
||||||
|
if (!reason.compareAndSet(null, cancellationReason)) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
callbacks.forEach(callback -> notifyCallback(callback, cancellationReason));
|
||||||
|
callbacks.clear();
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
public boolean isCancelled() {
|
||||||
|
return reason.get() != null;
|
||||||
|
}
|
||||||
|
|
||||||
|
public Optional<RunCancellationReason> reason() {
|
||||||
|
return Optional.ofNullable(reason.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
public void onCancel(Consumer<RunCancellationReason> callback) {
|
||||||
|
Objects.requireNonNull(callback, "callback must not be null");
|
||||||
|
AtomicBoolean invoked = new AtomicBoolean();
|
||||||
|
Consumer<RunCancellationReason> once = cancellationReason -> {
|
||||||
|
if (invoked.compareAndSet(false, true)) {
|
||||||
|
callback.accept(cancellationReason);
|
||||||
|
}
|
||||||
|
};
|
||||||
|
callbacks.add(once);
|
||||||
|
|
||||||
|
RunCancellationReason current = reason.get();
|
||||||
|
if (current != null) {
|
||||||
|
notifyCallback(once, current);
|
||||||
|
callbacks.remove(once);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private void notifyCallback(Consumer<RunCancellationReason> callback,
|
||||||
|
RunCancellationReason cancellationReason) {
|
||||||
|
try {
|
||||||
|
callback.accept(cancellationReason);
|
||||||
|
} catch (RuntimeException e) {
|
||||||
|
log.warn("Run cancellation callback failed: reason={}", cancellationReason, e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
public enum RunCancellationReason {
|
||||||
|
CLIENT_DISCONNECTED,
|
||||||
|
USER_REQUESTED,
|
||||||
|
DEADLINE_EXCEEDED,
|
||||||
|
BUDGET_EXHAUSTED,
|
||||||
|
INTERNAL_FAILURE
|
||||||
|
}
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
|
||||||
|
public final class RunCapacityCounter {
|
||||||
|
|
||||||
|
private final long maxBytes;
|
||||||
|
private final AtomicLong usedBytes = new AtomicLong();
|
||||||
|
|
||||||
|
public RunCapacityCounter(long maxBytes) {
|
||||||
|
if (maxBytes <= 0) {
|
||||||
|
throw new IllegalArgumentException("maxBytes must be positive");
|
||||||
|
}
|
||||||
|
this.maxBytes = maxBytes;
|
||||||
|
}
|
||||||
|
|
||||||
|
public long reserve(long bytes) {
|
||||||
|
if (bytes <= 0) {
|
||||||
|
throw new IllegalArgumentException("bytes must be positive");
|
||||||
|
}
|
||||||
|
while (true) {
|
||||||
|
long current = usedBytes.get();
|
||||||
|
long attempted = safeAdd(current, bytes);
|
||||||
|
if (attempted > maxBytes) {
|
||||||
|
throw new BudgetExceededException(BudgetKind.RUN_BYTES, maxBytes, attempted);
|
||||||
|
}
|
||||||
|
if (usedBytes.compareAndSet(current, attempted)) {
|
||||||
|
return attempted;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
public long usedBytes() {
|
||||||
|
return usedBytes.get();
|
||||||
|
}
|
||||||
|
|
||||||
|
public long maxBytes() {
|
||||||
|
return maxBytes;
|
||||||
|
}
|
||||||
|
|
||||||
|
private static long safeAdd(long left, long right) {
|
||||||
|
try {
|
||||||
|
return Math.addExact(left, right);
|
||||||
|
} catch (ArithmeticException e) {
|
||||||
|
return Long.MAX_VALUE;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
|
||||||
|
|
||||||
|
import java.time.Instant;
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Structurally immutable context. Mutable per-run state lives in thread-safe handles.
|
||||||
|
*/
|
||||||
|
public record RunContext(
|
||||||
|
String sessionId,
|
||||||
|
String runId,
|
||||||
|
Instant deadline,
|
||||||
|
RunCancellation cancellation,
|
||||||
|
RunBudget budget,
|
||||||
|
HarnessRetryPolicies retryPolicies,
|
||||||
|
RunLifecycle lifecycle) {
|
||||||
|
|
||||||
|
public RunContext {
|
||||||
|
requireText(sessionId, "sessionId");
|
||||||
|
requireText(runId, "runId");
|
||||||
|
Objects.requireNonNull(deadline, "deadline must not be null");
|
||||||
|
Objects.requireNonNull(cancellation, "cancellation must not be null");
|
||||||
|
Objects.requireNonNull(budget, "budget must not be null");
|
||||||
|
Objects.requireNonNull(retryPolicies, "retryPolicies must not be null");
|
||||||
|
Objects.requireNonNull(lifecycle, "lifecycle must not be null");
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void requireText(String value, String name) {
|
||||||
|
if (value == null || value.isBlank()) {
|
||||||
|
throw new IllegalArgumentException(name + " must not be blank");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import java.time.Clock;
|
||||||
|
import java.util.Objects;
|
||||||
|
import java.util.Optional;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
|
||||||
|
public final class RunLifecycle {
|
||||||
|
|
||||||
|
private final Clock clock;
|
||||||
|
private final AtomicReference<RunTermination> termination = new AtomicReference<>();
|
||||||
|
|
||||||
|
public RunLifecycle(Clock clock) {
|
||||||
|
this.clock = Objects.requireNonNull(clock, "clock must not be null");
|
||||||
|
}
|
||||||
|
|
||||||
|
public RunState state() {
|
||||||
|
RunTermination current = termination.get();
|
||||||
|
return current == null ? RunState.RUNNING : current.state();
|
||||||
|
}
|
||||||
|
|
||||||
|
public Optional<RunTermination> termination() {
|
||||||
|
return Optional.ofNullable(termination.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
public boolean finish(RunState state, String reason) {
|
||||||
|
return termination.compareAndSet(null, new RunTermination(state, reason, clock.instant()));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,20 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
public enum RunState {
|
||||||
|
RUNNING(false),
|
||||||
|
SUCCESS(true),
|
||||||
|
FAILED(true),
|
||||||
|
CANCELLED(true),
|
||||||
|
TIMED_OUT(true),
|
||||||
|
BUDGET_EXHAUSTED(true);
|
||||||
|
|
||||||
|
private final boolean terminal;
|
||||||
|
|
||||||
|
RunState(boolean terminal) {
|
||||||
|
this.terminal = terminal;
|
||||||
|
}
|
||||||
|
|
||||||
|
public boolean isTerminal() {
|
||||||
|
return terminal;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import java.time.Instant;
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
public record RunTermination(RunState state, String reason, Instant completedAt) {
|
||||||
|
|
||||||
|
public RunTermination {
|
||||||
|
Objects.requireNonNull(state, "state must not be null");
|
||||||
|
Objects.requireNonNull(completedAt, "completedAt must not be null");
|
||||||
|
if (!state.isTerminal()) {
|
||||||
|
throw new IllegalArgumentException("state must be terminal");
|
||||||
|
}
|
||||||
|
if (reason == null || reason.isBlank()) {
|
||||||
|
throw new IllegalArgumentException("reason must not be blank");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,60 @@
|
|||||||
|
package com.superbiz.agent.harness.retry;
|
||||||
|
|
||||||
|
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
|
||||||
|
import com.superbiz.agent.harness.core.BudgetExceededException;
|
||||||
|
import com.superbiz.agent.harness.core.RunAbortedException;
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import com.superbiz.agent.harness.core.RunState;
|
||||||
|
|
||||||
|
import java.util.Objects;
|
||||||
|
import java.util.function.Consumer;
|
||||||
|
|
||||||
|
public final class HarnessRetryExecutor {
|
||||||
|
|
||||||
|
private final DiagnosisHarnessCore core;
|
||||||
|
|
||||||
|
public HarnessRetryExecutor(DiagnosisHarnessCore core) {
|
||||||
|
this.core = Objects.requireNonNull(core, "core must not be null");
|
||||||
|
}
|
||||||
|
|
||||||
|
public <T> T execute(RunContext context,
|
||||||
|
RetryPolicy policy,
|
||||||
|
RetryOperation<T> operation,
|
||||||
|
RetryFailureClassifier classifier,
|
||||||
|
Consumer<RetryAttempt> recorder) {
|
||||||
|
Objects.requireNonNull(context, "context must not be null");
|
||||||
|
Objects.requireNonNull(policy, "policy must not be null");
|
||||||
|
Objects.requireNonNull(operation, "operation must not be null");
|
||||||
|
Objects.requireNonNull(classifier, "classifier must not be null");
|
||||||
|
Objects.requireNonNull(recorder, "recorder must not be null");
|
||||||
|
|
||||||
|
for (int attempt = 1; attempt <= policy.maxAttempts(); attempt++) {
|
||||||
|
core.checkActive(context);
|
||||||
|
try {
|
||||||
|
T result = operation.execute();
|
||||||
|
recorder.accept(RetryAttempt.succeeded(attempt));
|
||||||
|
return result;
|
||||||
|
} catch (RunAbortedException exception) {
|
||||||
|
RetryFailure failure = exception.termination().state() == RunState.BUDGET_EXHAUSTED
|
||||||
|
? RetryFailure.BUDGET_EXHAUSTED
|
||||||
|
: RetryFailure.CANCELLED;
|
||||||
|
recorder.accept(RetryAttempt.failed(attempt, failure));
|
||||||
|
throw new RetryExecutionException(attempt, failure, exception);
|
||||||
|
} catch (BudgetExceededException exception) {
|
||||||
|
recorder.accept(RetryAttempt.failed(attempt, RetryFailure.BUDGET_EXHAUSTED));
|
||||||
|
throw new RetryExecutionException(
|
||||||
|
attempt, RetryFailure.BUDGET_EXHAUSTED, exception);
|
||||||
|
} catch (Exception exception) {
|
||||||
|
RetryFailure failure = classifier.classify(exception);
|
||||||
|
if (failure == null) {
|
||||||
|
failure = RetryFailure.UNKNOWN;
|
||||||
|
}
|
||||||
|
recorder.accept(RetryAttempt.failed(attempt, failure));
|
||||||
|
if (!policy.allowsRetry(attempt, failure)) {
|
||||||
|
throw new RetryExecutionException(attempt, failure, exception);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
throw new IllegalStateException("retry loop exited without a result");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
package com.superbiz.agent.harness.retry;
|
||||||
|
|
||||||
|
import java.util.Objects;
|
||||||
|
import java.util.Set;
|
||||||
|
|
||||||
|
public record HarnessRetryPolicies(
|
||||||
|
RetryPolicy intentRouter,
|
||||||
|
RetryPolicy diagnosisAgent,
|
||||||
|
RetryPolicy toolCall,
|
||||||
|
RetryPolicy semanticGuard,
|
||||||
|
RetryPolicy evidenceRepair) {
|
||||||
|
|
||||||
|
public HarnessRetryPolicies {
|
||||||
|
Objects.requireNonNull(intentRouter, "intentRouter must not be null");
|
||||||
|
Objects.requireNonNull(diagnosisAgent, "diagnosisAgent must not be null");
|
||||||
|
Objects.requireNonNull(toolCall, "toolCall must not be null");
|
||||||
|
Objects.requireNonNull(semanticGuard, "semanticGuard must not be null");
|
||||||
|
Objects.requireNonNull(evidenceRepair, "evidenceRepair must not be null");
|
||||||
|
}
|
||||||
|
|
||||||
|
public static HarnessRetryPolicies strict() {
|
||||||
|
RetryPolicy oneAttempt = new RetryPolicy(1, Set.of());
|
||||||
|
return new HarnessRetryPolicies(
|
||||||
|
new RetryPolicy(2, Set.of(
|
||||||
|
RetryFailure.TIMEOUT,
|
||||||
|
RetryFailure.TRANSPORT,
|
||||||
|
RetryFailure.INVALID_OUTPUT)),
|
||||||
|
oneAttempt,
|
||||||
|
oneAttempt,
|
||||||
|
new RetryPolicy(2, Set.of(
|
||||||
|
RetryFailure.TIMEOUT,
|
||||||
|
RetryFailure.TRANSPORT,
|
||||||
|
RetryFailure.PARSE_ERROR,
|
||||||
|
RetryFailure.SCHEMA_INVALID)),
|
||||||
|
oneAttempt);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
package com.superbiz.agent.harness.retry;
|
||||||
|
|
||||||
|
public record RetryAttempt(int attemptNumber, boolean success, RetryFailure failure) {
|
||||||
|
|
||||||
|
public RetryAttempt {
|
||||||
|
if (attemptNumber <= 0) {
|
||||||
|
throw new IllegalArgumentException("attemptNumber must be positive");
|
||||||
|
}
|
||||||
|
if (success && failure != null) {
|
||||||
|
throw new IllegalArgumentException("successful attempt must not contain a failure");
|
||||||
|
}
|
||||||
|
if (!success && failure == null) {
|
||||||
|
throw new IllegalArgumentException("failed attempt must contain a failure");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
public static RetryAttempt succeeded(int attemptNumber) {
|
||||||
|
return new RetryAttempt(attemptNumber, true, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
public static RetryAttempt failed(int attemptNumber, RetryFailure failure) {
|
||||||
|
return new RetryAttempt(attemptNumber, false, failure);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
package com.superbiz.agent.harness.retry;
|
||||||
|
|
||||||
|
public final class RetryExecutionException extends RuntimeException {
|
||||||
|
|
||||||
|
private final int attempts;
|
||||||
|
private final RetryFailure failure;
|
||||||
|
|
||||||
|
public RetryExecutionException(int attempts, RetryFailure failure, Exception cause) {
|
||||||
|
super("Operation failed after " + attempts + " attempt(s): " + failure, cause);
|
||||||
|
this.attempts = attempts;
|
||||||
|
this.failure = failure;
|
||||||
|
}
|
||||||
|
|
||||||
|
public int attempts() {
|
||||||
|
return attempts;
|
||||||
|
}
|
||||||
|
|
||||||
|
public RetryFailure failure() {
|
||||||
|
return failure;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
package com.superbiz.agent.harness.retry;
|
||||||
|
|
||||||
|
public enum RetryFailure {
|
||||||
|
TIMEOUT,
|
||||||
|
TRANSPORT,
|
||||||
|
INVALID_OUTPUT,
|
||||||
|
PARSE_ERROR,
|
||||||
|
SCHEMA_INVALID,
|
||||||
|
NO_EVIDENCE,
|
||||||
|
BUSINESS_REJECTION,
|
||||||
|
CANCELLED,
|
||||||
|
BUDGET_EXHAUSTED,
|
||||||
|
UNKNOWN
|
||||||
|
}
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
package com.superbiz.agent.harness.retry;
|
||||||
|
|
||||||
|
@FunctionalInterface
|
||||||
|
public interface RetryFailureClassifier {
|
||||||
|
RetryFailure classify(Exception exception);
|
||||||
|
}
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
package com.superbiz.agent.harness.retry;
|
||||||
|
|
||||||
|
@FunctionalInterface
|
||||||
|
public interface RetryOperation<T> {
|
||||||
|
T execute() throws Exception;
|
||||||
|
}
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
package com.superbiz.agent.harness.retry;
|
||||||
|
|
||||||
|
import java.util.Set;
|
||||||
|
|
||||||
|
public record RetryPolicy(int maxAttempts, Set<RetryFailure> retryableFailures) {
|
||||||
|
|
||||||
|
public RetryPolicy {
|
||||||
|
if (maxAttempts < 1 || maxAttempts > 2) {
|
||||||
|
throw new IllegalArgumentException("maxAttempts must be 1 or 2");
|
||||||
|
}
|
||||||
|
retryableFailures = retryableFailures == null ? Set.of() : Set.copyOf(retryableFailures);
|
||||||
|
}
|
||||||
|
|
||||||
|
public boolean allowsRetry(int completedAttempts, RetryFailure failure) {
|
||||||
|
return completedAttempts < maxAttempts && retryableFailures.contains(failure);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
package com.superbiz.agent.harness.tool.store;
|
||||||
|
|
||||||
|
import java.util.regex.Pattern;
|
||||||
|
|
||||||
|
public final class ToolCallKeyFactory {
|
||||||
|
|
||||||
|
private static final int MAX_SEGMENT_LENGTH = 128;
|
||||||
|
private static final Pattern SAFE_SEGMENT = Pattern.compile("[A-Za-z0-9][A-Za-z0-9._-]*");
|
||||||
|
|
||||||
|
private final String keyPrefix;
|
||||||
|
|
||||||
|
public ToolCallKeyFactory(String keyPrefix) {
|
||||||
|
if (keyPrefix == null || keyPrefix.isBlank()) {
|
||||||
|
throw new IllegalArgumentException("keyPrefix must not be blank");
|
||||||
|
}
|
||||||
|
if (keyPrefix.startsWith(":") || keyPrefix.endsWith(":") || keyPrefix.contains("::")) {
|
||||||
|
throw new IllegalArgumentException("keyPrefix contains an empty segment");
|
||||||
|
}
|
||||||
|
this.keyPrefix = keyPrefix;
|
||||||
|
}
|
||||||
|
|
||||||
|
public String create(String runId, String toolCallId) {
|
||||||
|
requireSafeSegment(runId, "runId");
|
||||||
|
requireSafeSegment(toolCallId, "toolCallId");
|
||||||
|
return keyPrefix + ":" + runId + ":" + toolCallId;
|
||||||
|
}
|
||||||
|
|
||||||
|
public String keyPrefix() {
|
||||||
|
return keyPrefix;
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void requireSafeSegment(String value, String name) {
|
||||||
|
if (value == null || value.isBlank()) {
|
||||||
|
throw new IllegalArgumentException(name + " must not be blank");
|
||||||
|
}
|
||||||
|
if (value.length() > MAX_SEGMENT_LENGTH || !SAFE_SEGMENT.matcher(value).matches()) {
|
||||||
|
throw new IllegalArgumentException(name + " contains unsafe characters or is too long");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -93,6 +93,10 @@ spring:
|
|||||||
min-idle: 0
|
min-idle: 0
|
||||||
|
|
||||||
ai:
|
ai:
|
||||||
|
# Harness owns allowed retries; provider SDK calls execute once per recorded attempt.
|
||||||
|
retry:
|
||||||
|
max-attempts: 1
|
||||||
|
|
||||||
vectorstore:
|
vectorstore:
|
||||||
type: milvus
|
type: milvus
|
||||||
milvus:
|
milvus:
|
||||||
|
|||||||
@@ -0,0 +1,33 @@
|
|||||||
|
package com.superbiz.agent.config;
|
||||||
|
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.yaml.snakeyaml.Yaml;
|
||||||
|
|
||||||
|
import java.io.InputStream;
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
|
||||||
|
class SpringAiRetryConfigurationTest {
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void disablesHiddenSpringAiRetries() {
|
||||||
|
try (InputStream input = getClass().getClassLoader().getResourceAsStream("application.yml")) {
|
||||||
|
assertNotNull(input);
|
||||||
|
Map<String, Object> root = new Yaml().load(input);
|
||||||
|
Map<String, Object> spring = map(root.get("spring"));
|
||||||
|
Map<String, Object> ai = map(spring.get("ai"));
|
||||||
|
Map<String, Object> retry = map(ai.get("retry"));
|
||||||
|
|
||||||
|
assertEquals(1, ((Number) retry.get("max-attempts")).intValue());
|
||||||
|
} catch (Exception e) {
|
||||||
|
throw new AssertionError("Failed to read application.yml", e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
private Map<String, Object> map(Object value) {
|
||||||
|
return (Map<String, Object>) value;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,92 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.time.Duration;
|
||||||
|
import java.time.Instant;
|
||||||
|
import java.util.concurrent.atomic.AtomicBoolean;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
class DiagnosisHarnessCoreTest {
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void deadlinePreventsFakeToolInvocation() {
|
||||||
|
MutableClock clock = new MutableClock(Instant.parse("2026-07-21T10:00:00Z"));
|
||||||
|
DiagnosisHarnessCore core = HarnessCoreFixtures.core(clock);
|
||||||
|
RunContext context = core.startRun("session-1", "run-1");
|
||||||
|
AtomicInteger toolCalls = new AtomicInteger();
|
||||||
|
clock.advance(Duration.ofMinutes(2));
|
||||||
|
|
||||||
|
RunAbortedException failure = assertThrows(RunAbortedException.class, () -> {
|
||||||
|
core.beforeToolCall(context, "query_logs");
|
||||||
|
toolCalls.incrementAndGet();
|
||||||
|
});
|
||||||
|
|
||||||
|
assertEquals(0, toolCalls.get());
|
||||||
|
assertEquals(RunState.TIMED_OUT, failure.termination().state());
|
||||||
|
assertEquals(RunCancellationReason.DEADLINE_EXCEEDED, context.cancellation().reason().orElseThrow());
|
||||||
|
assertThrows(RunAbortedException.class, () -> core.beforeModelCall(context));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void clientCancellationStopsSubsequentFakeModel() {
|
||||||
|
DiagnosisHarnessCore core = HarnessCoreFixtures.core(new MutableClock(Instant.parse("2026-07-21T10:00:00Z")));
|
||||||
|
RunContext context = core.startRun("session-1", "run-1");
|
||||||
|
AtomicBoolean resourceCancelled = new AtomicBoolean();
|
||||||
|
context.cancellation().onCancel(reason -> resourceCancelled.set(true));
|
||||||
|
|
||||||
|
assertTrue(core.cancel(context, RunCancellationReason.CLIENT_DISCONNECTED));
|
||||||
|
|
||||||
|
assertTrue(resourceCancelled.get());
|
||||||
|
assertThrows(RunAbortedException.class, () -> core.beforeModelCall(context));
|
||||||
|
assertEquals(RunState.CANCELLED, context.lifecycle().state());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void toolBudgetExhaustionWinsOverLaterFailure() {
|
||||||
|
MutableClock clock = new MutableClock(Instant.parse("2026-07-21T10:00:00Z"));
|
||||||
|
DiagnosisHarnessCore core = new DiagnosisHarnessCore(
|
||||||
|
clock,
|
||||||
|
() -> "run-1",
|
||||||
|
Duration.ofMinutes(1),
|
||||||
|
new RunBudgetLimits(2, 1, 1, 100, 100, 200, 100),
|
||||||
|
HarnessRetryPolicies.strict());
|
||||||
|
RunContext context = core.startRun("session-1");
|
||||||
|
core.beforeToolCall(context, "query_logs");
|
||||||
|
|
||||||
|
BudgetExceededException failure = assertThrows(
|
||||||
|
BudgetExceededException.class,
|
||||||
|
() -> core.beforeToolCall(context, "query_mysql"));
|
||||||
|
RunTermination termination = context.lifecycle().termination().orElseThrow();
|
||||||
|
|
||||||
|
assertEquals(BudgetKind.TOOL_CALLS, failure.kind());
|
||||||
|
assertEquals(RunState.BUDGET_EXHAUSTED, termination.state());
|
||||||
|
assertFalse(core.completeFailure(context, "late failure"));
|
||||||
|
assertEquals(termination, context.lifecycle().termination().orElseThrow());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void actualTokenExhaustionIsRecordedAndStopsNextOperation() {
|
||||||
|
MutableClock clock = new MutableClock(Instant.parse("2026-07-21T10:00:00Z"));
|
||||||
|
DiagnosisHarnessCore core = new DiagnosisHarnessCore(
|
||||||
|
clock,
|
||||||
|
() -> "run-1",
|
||||||
|
Duration.ofMinutes(1),
|
||||||
|
new RunBudgetLimits(2, 2, 2, 10, 10, 15, 100),
|
||||||
|
HarnessRetryPolicies.strict());
|
||||||
|
RunContext context = core.startRun("session-1");
|
||||||
|
core.beforeModelCall(context);
|
||||||
|
|
||||||
|
assertThrows(BudgetExceededException.class, () -> core.recordTokens(context, 9, 8));
|
||||||
|
|
||||||
|
assertEquals(17, context.budget().snapshot().totalTokens());
|
||||||
|
assertEquals(RunState.BUDGET_EXHAUSTED, context.lifecycle().state());
|
||||||
|
assertThrows(RunAbortedException.class, () -> core.beforeToolCall(context, "query_logs"));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
|
||||||
|
|
||||||
|
import java.time.Duration;
|
||||||
|
|
||||||
|
public final class HarnessCoreFixtures {
|
||||||
|
|
||||||
|
private HarnessCoreFixtures() {
|
||||||
|
}
|
||||||
|
|
||||||
|
public static RunBudgetLimits generousLimits() {
|
||||||
|
return new RunBudgetLimits(10, 10, 5, 10_000, 10_000, 20_000, 1_000_000);
|
||||||
|
}
|
||||||
|
|
||||||
|
public static DiagnosisHarnessCore core(MutableClock clock) {
|
||||||
|
return new DiagnosisHarnessCore(
|
||||||
|
clock,
|
||||||
|
() -> "run-generated-1",
|
||||||
|
Duration.ofMinutes(2),
|
||||||
|
generousLimits(),
|
||||||
|
HarnessRetryPolicies.strict());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,42 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import java.time.Clock;
|
||||||
|
import java.time.Duration;
|
||||||
|
import java.time.Instant;
|
||||||
|
import java.time.ZoneId;
|
||||||
|
import java.util.Objects;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
|
||||||
|
public final class MutableClock extends Clock {
|
||||||
|
|
||||||
|
private final AtomicReference<Instant> instant;
|
||||||
|
private final ZoneId zone;
|
||||||
|
|
||||||
|
public MutableClock(Instant instant) {
|
||||||
|
this(instant, ZoneId.of("UTC"));
|
||||||
|
}
|
||||||
|
|
||||||
|
private MutableClock(Instant instant, ZoneId zone) {
|
||||||
|
this.instant = new AtomicReference<>(Objects.requireNonNull(instant));
|
||||||
|
this.zone = Objects.requireNonNull(zone);
|
||||||
|
}
|
||||||
|
|
||||||
|
public void advance(Duration duration) {
|
||||||
|
instant.updateAndGet(current -> current.plus(duration));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ZoneId getZone() {
|
||||||
|
return zone;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Clock withZone(ZoneId zone) {
|
||||||
|
return new MutableClock(instant(), zone);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Instant instant() {
|
||||||
|
return instant.get();
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,74 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.concurrent.ExecutorService;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.Future;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
|
||||||
|
class RunBudgetTest {
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void rejectsInvalidLimits() {
|
||||||
|
assertThrows(IllegalArgumentException.class,
|
||||||
|
() -> new RunBudgetLimits(0, 1, 1, 1, 1, 1, 1));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void perToolExhaustionDoesNotPartiallyIncrementCounters() {
|
||||||
|
RunBudget budget = new RunBudget(new RunBudgetLimits(2, 3, 1, 100, 100, 200, 100));
|
||||||
|
budget.reserveToolCall("query_logs");
|
||||||
|
|
||||||
|
BudgetExceededException failure = assertThrows(
|
||||||
|
BudgetExceededException.class,
|
||||||
|
() -> budget.reserveToolCall("query_logs"));
|
||||||
|
|
||||||
|
RunBudgetUsage usage = budget.snapshot();
|
||||||
|
assertEquals(BudgetKind.TOOL_CALLS_PER_TOOL, failure.kind());
|
||||||
|
assertEquals(1, usage.toolCalls());
|
||||||
|
assertEquals(1, usage.toolCallsByName().get("query_logs"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void tokenExhaustionKeepsActualUsage() {
|
||||||
|
RunBudget budget = new RunBudget(new RunBudgetLimits(2, 2, 2, 10, 10, 15, 100));
|
||||||
|
|
||||||
|
BudgetExceededException failure = assertThrows(
|
||||||
|
BudgetExceededException.class,
|
||||||
|
() -> budget.recordTokens(9, 8));
|
||||||
|
|
||||||
|
assertEquals(BudgetKind.TOTAL_TOKENS, failure.kind());
|
||||||
|
assertEquals(9, budget.snapshot().inputTokens());
|
||||||
|
assertEquals(8, budget.snapshot().outputTokens());
|
||||||
|
assertEquals(17, budget.snapshot().totalTokens());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void concurrentCapacityReservationsAreAtomic() throws Exception {
|
||||||
|
RunCapacityCounter counter = new RunCapacityCounter(100);
|
||||||
|
ExecutorService executor = Executors.newFixedThreadPool(5);
|
||||||
|
try {
|
||||||
|
List<Future<Long>> reservations = new ArrayList<>();
|
||||||
|
for (int i = 0; i < 10; i++) {
|
||||||
|
reservations.add(executor.submit(() -> counter.reserve(10)));
|
||||||
|
}
|
||||||
|
for (Future<Long> reservation : reservations) {
|
||||||
|
reservation.get();
|
||||||
|
}
|
||||||
|
} finally {
|
||||||
|
executor.shutdownNow();
|
||||||
|
}
|
||||||
|
|
||||||
|
assertEquals(100, counter.usedBytes());
|
||||||
|
BudgetExceededException failure = assertThrows(
|
||||||
|
BudgetExceededException.class,
|
||||||
|
() -> counter.reserve(1));
|
||||||
|
assertEquals(BudgetKind.RUN_BYTES, failure.kind());
|
||||||
|
assertEquals(100, counter.usedBytes());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,79 @@
|
|||||||
|
package com.superbiz.agent.harness.core;
|
||||||
|
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.time.Instant;
|
||||||
|
import java.util.concurrent.CompletableFuture;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
class RunContextTest {
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void rejectsBlankIdentityBeforePublishingContext() {
|
||||||
|
DiagnosisHarnessCore core = HarnessCoreFixtures.core(new MutableClock(Instant.parse("2026-07-21T10:00:00Z")));
|
||||||
|
|
||||||
|
assertThrows(IllegalArgumentException.class, () -> core.startRun(" ", "run-1"));
|
||||||
|
assertThrows(IllegalArgumentException.class, () -> core.startRun("session-1", " "));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void explicitlyPropagatesSameHandlesAcrossAsyncBoundary() throws Exception {
|
||||||
|
DiagnosisHarnessCore core = HarnessCoreFixtures.core(new MutableClock(Instant.parse("2026-07-21T10:00:00Z")));
|
||||||
|
RunContext context = core.startRun("session-1", "run-1");
|
||||||
|
AtomicReference<Thread> worker = new AtomicReference<>();
|
||||||
|
|
||||||
|
RunContext observed = CompletableFuture.supplyAsync(() -> {
|
||||||
|
worker.set(Thread.currentThread());
|
||||||
|
return fakeTool(context);
|
||||||
|
}).get(5, TimeUnit.SECONDS);
|
||||||
|
|
||||||
|
assertEquals("session-1", observed.sessionId());
|
||||||
|
assertEquals("run-1", observed.runId());
|
||||||
|
assertSame(context.cancellation(), observed.cancellation());
|
||||||
|
assertSame(context.budget(), observed.budget());
|
||||||
|
assertSame(context.lifecycle(), observed.lifecycle());
|
||||||
|
assertFalse(worker.get().equals(Thread.currentThread()));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void cancellationUsesFirstReasonAndNotifiesRegisteredCallbacks() {
|
||||||
|
DiagnosisHarnessCore core = HarnessCoreFixtures.core(new MutableClock(Instant.parse("2026-07-21T10:00:00Z")));
|
||||||
|
RunContext context = core.startRun("session-1", "run-1");
|
||||||
|
AtomicInteger callbacks = new AtomicInteger();
|
||||||
|
context.cancellation().onCancel(reason -> callbacks.incrementAndGet());
|
||||||
|
context.cancellation().onCancel(reason -> callbacks.incrementAndGet());
|
||||||
|
|
||||||
|
assertTrue(core.cancel(context, RunCancellationReason.CLIENT_DISCONNECTED));
|
||||||
|
assertFalse(context.cancellation().cancel(RunCancellationReason.USER_REQUESTED));
|
||||||
|
|
||||||
|
assertEquals(RunCancellationReason.CLIENT_DISCONNECTED, context.cancellation().reason().orElseThrow());
|
||||||
|
assertEquals(2, callbacks.get());
|
||||||
|
assertEquals(RunState.CANCELLED, context.lifecycle().state());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void firstTerminalOutcomeCannotBeOverwritten() {
|
||||||
|
DiagnosisHarnessCore core = HarnessCoreFixtures.core(new MutableClock(Instant.parse("2026-07-21T10:00:00Z")));
|
||||||
|
RunContext context = core.startRun("session-1", "run-1");
|
||||||
|
|
||||||
|
assertTrue(core.completeSuccess(context));
|
||||||
|
RunTermination success = context.lifecycle().termination().orElseThrow();
|
||||||
|
assertFalse(core.completeFailure(context, "late failure"));
|
||||||
|
assertFalse(core.cancel(context, RunCancellationReason.CLIENT_DISCONNECTED));
|
||||||
|
|
||||||
|
assertSame(success, context.lifecycle().termination().orElseThrow());
|
||||||
|
assertEquals(RunState.SUCCESS, context.lifecycle().state());
|
||||||
|
}
|
||||||
|
|
||||||
|
private RunContext fakeTool(RunContext context) {
|
||||||
|
return context;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,134 @@
|
|||||||
|
package com.superbiz.agent.harness.retry;
|
||||||
|
|
||||||
|
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
|
||||||
|
import com.superbiz.agent.harness.core.BudgetExceededException;
|
||||||
|
import com.superbiz.agent.harness.core.BudgetKind;
|
||||||
|
import com.superbiz.agent.harness.core.HarnessCoreFixtures;
|
||||||
|
import com.superbiz.agent.harness.core.MutableClock;
|
||||||
|
import com.superbiz.agent.harness.core.RunAbortedException;
|
||||||
|
import com.superbiz.agent.harness.core.RunCancellationReason;
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import org.junit.jupiter.api.BeforeEach;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.io.IOException;
|
||||||
|
import java.time.Instant;
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
|
||||||
|
class HarnessRetryExecutorTest {
|
||||||
|
|
||||||
|
private DiagnosisHarnessCore core;
|
||||||
|
private RunContext context;
|
||||||
|
private HarnessRetryExecutor executor;
|
||||||
|
|
||||||
|
@BeforeEach
|
||||||
|
void setUp() {
|
||||||
|
core = HarnessCoreFixtures.core(new MutableClock(Instant.parse("2026-07-21T10:00:00Z")));
|
||||||
|
context = core.startRun("session-1", "run-1");
|
||||||
|
executor = new HarnessRetryExecutor(core);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void retriesRouterTransportFailureOnceAndRecordsBothAttempts() {
|
||||||
|
AtomicInteger calls = new AtomicInteger();
|
||||||
|
List<RetryAttempt> attempts = new ArrayList<>();
|
||||||
|
|
||||||
|
String result = executor.execute(
|
||||||
|
context,
|
||||||
|
context.retryPolicies().intentRouter(),
|
||||||
|
() -> {
|
||||||
|
if (calls.incrementAndGet() == 1) {
|
||||||
|
throw new IOException("temporary transport failure");
|
||||||
|
}
|
||||||
|
return "DIAGNOSIS";
|
||||||
|
},
|
||||||
|
exception -> RetryFailure.TRANSPORT,
|
||||||
|
attempts::add);
|
||||||
|
|
||||||
|
assertEquals("DIAGNOSIS", result);
|
||||||
|
assertEquals(2, calls.get());
|
||||||
|
assertEquals(List.of(
|
||||||
|
RetryAttempt.failed(1, RetryFailure.TRANSPORT),
|
||||||
|
RetryAttempt.succeeded(2)), attempts);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void toolPolicyRunsOnlyOneAttempt() {
|
||||||
|
AtomicInteger calls = new AtomicInteger();
|
||||||
|
List<RetryAttempt> attempts = new ArrayList<>();
|
||||||
|
|
||||||
|
RetryExecutionException failure = assertThrows(RetryExecutionException.class, () -> executor.execute(
|
||||||
|
context,
|
||||||
|
context.retryPolicies().toolCall(),
|
||||||
|
() -> {
|
||||||
|
calls.incrementAndGet();
|
||||||
|
throw new IOException("tool failed");
|
||||||
|
},
|
||||||
|
exception -> RetryFailure.TRANSPORT,
|
||||||
|
attempts::add));
|
||||||
|
|
||||||
|
assertEquals(1, calls.get());
|
||||||
|
assertEquals(1, failure.attempts());
|
||||||
|
assertEquals(List.of(RetryAttempt.failed(1, RetryFailure.TRANSPORT)), attempts);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void cancellationBetweenAttemptsPreventsSecondInvocation() {
|
||||||
|
AtomicInteger calls = new AtomicInteger();
|
||||||
|
|
||||||
|
assertThrows(RunAbortedException.class, () -> executor.execute(
|
||||||
|
context,
|
||||||
|
context.retryPolicies().semanticGuard(),
|
||||||
|
() -> {
|
||||||
|
calls.incrementAndGet();
|
||||||
|
throw new IOException("temporary transport failure");
|
||||||
|
},
|
||||||
|
exception -> RetryFailure.TRANSPORT,
|
||||||
|
attempt -> core.cancel(context, RunCancellationReason.CLIENT_DISCONNECTED)));
|
||||||
|
|
||||||
|
assertEquals(1, calls.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void noEvidenceAndBusinessRejectionAreNotRetryable() {
|
||||||
|
AtomicInteger calls = new AtomicInteger();
|
||||||
|
|
||||||
|
RetryExecutionException failure = assertThrows(RetryExecutionException.class, () -> executor.execute(
|
||||||
|
context,
|
||||||
|
context.retryPolicies().intentRouter(),
|
||||||
|
() -> {
|
||||||
|
calls.incrementAndGet();
|
||||||
|
throw new IllegalStateException("valid empty result");
|
||||||
|
},
|
||||||
|
exception -> RetryFailure.NO_EVIDENCE,
|
||||||
|
attempt -> { }));
|
||||||
|
|
||||||
|
assertEquals(1, calls.get());
|
||||||
|
assertEquals(RetryFailure.NO_EVIDENCE, failure.failure());
|
||||||
|
assertEquals(1, context.retryPolicies().diagnosisAgent().maxAttempts());
|
||||||
|
assertEquals(1, context.retryPolicies().evidenceRepair().maxAttempts());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void budgetFailureCannotBeMisclassifiedAsRetryable() {
|
||||||
|
AtomicInteger calls = new AtomicInteger();
|
||||||
|
|
||||||
|
RetryExecutionException failure = assertThrows(RetryExecutionException.class, () -> executor.execute(
|
||||||
|
context,
|
||||||
|
context.retryPolicies().intentRouter(),
|
||||||
|
() -> {
|
||||||
|
calls.incrementAndGet();
|
||||||
|
throw new BudgetExceededException(BudgetKind.MODEL_CALLS, 1, 2);
|
||||||
|
},
|
||||||
|
exception -> RetryFailure.TRANSPORT,
|
||||||
|
attempt -> { }));
|
||||||
|
|
||||||
|
assertEquals(1, calls.get());
|
||||||
|
assertEquals(RetryFailure.BUDGET_EXHAUSTED, failure.failure());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
package com.superbiz.agent.harness.tool.store;
|
||||||
|
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
|
||||||
|
class ToolCallKeyFactoryTest {
|
||||||
|
|
||||||
|
private final ToolCallKeyFactory factory = new ToolCallKeyFactory("superbiz:harness:tool-call");
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void preservesExactFrameworkToolCallId() {
|
||||||
|
assertEquals(
|
||||||
|
"superbiz:harness:tool-call:run-123:call_ABC-123",
|
||||||
|
factory.create("run-123", "call_ABC-123"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void rejectsBlankUnsafeAndOversizedSegmentsWithoutFallbackId() {
|
||||||
|
assertThrows(IllegalArgumentException.class, () -> factory.create("run-1", ""));
|
||||||
|
assertThrows(IllegalArgumentException.class, () -> factory.create("run:1", "call-1"));
|
||||||
|
assertThrows(IllegalArgumentException.class, () -> factory.create("run-1", "call/1"));
|
||||||
|
assertThrows(IllegalArgumentException.class, () -> factory.create("run-1", "a".repeat(129)));
|
||||||
|
}
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user