Files
SuperBizAgent-java/mvp/engineering/harness/RunBudget预算流程-一次Run的资源门禁时序图.md
T

252 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RunBudget 预算流程:一次 Run 的资源门禁时序图
**用途**:面试讲解 RunBudget 用的聚焦时序图,回答"一次 Run 的资源消耗是如何被门禁控制的"。
**代码基线**:`RunContext` → `RunBudget` + `RunBudgetLimits` + `RunCapacityCounter`
## 1. 在完整 Harness 中的位置(极简上下文)
RunBudget 是 `RunContext` 里的一个执行控制句柄,两个门禁经过它:
```mermaid
flowchart LR
MI["ModelInterceptor<br/>每轮模型调用前"] -->|"beforeModelCall"| CORE["DiagnosisHarnessCore"]
TI["ToolInterceptor / ToolBoundary<br/>每次 Tool 执行前"] -->|"beforeToolCall"| CORE
CORE --> B["RunBudget<br/>预扣 + 超限升级"]
T["Canonical 写入前"] -->|"reserveRunBytes"| CORE
B --> E["BudgetExceededException →<br/>finish(BUDGET_EXHAUSTED) + cancel"]
```
## 2. RunBudget 流程时序图(核心)
```mermaid
sequenceDiagram
participant APP as ChatApplication
participant CORE as DiagnosisHarnessCore
participant MI as ModelInterceptor
participant TI as ToolInterceptor
participant T as ToolBoundary
participant B as RunBudget
participant C as RunCapacityCounter(CAS)
Note over APP,C: 启动:startRun(sessionId) → new RunBudget(RunBudgetLimits)<br/>句柄挂到 RunContext,随请求显式传递
loop 每一轮模型调用
MI->>CORE: beforeModelCall(context)
CORE->>CORE: checkActive() 先确认 Run 还能跑
CORE->>B: reserveModelCall() 预扣 1 轮
alt 超限
B-->>CORE: BudgetExceededException(MODEL_CALLS)
CORE->>CORE: exhaustBudget:finish(BUDGET_EXHAUSTED) + cancel()
CORE-->>MI: 抛异常,之后所有 checkActive 拒绝
end
CORE->>B: recordTokens(input, output) 调用后记账
B->>B: 三档检查 input / output / total
end
loop 每一次 Tool 调用
TI->>CORE: beforeToolCall(context, toolName)
CORE->>CORE: checkActive()
CORE->>B: reserveToolCall(toolName)
alt 超限(总量 或 单 Tool 独立限额)
B-->>CORE: BudgetExceededException(TOOL_CALLS / TOOL_CALLS_PER_TOOL)
CORE->>CORE: exhaustBudget(...)
end
T->>CORE: reserveRunBytes(bytes) canonical 体积预扣
CORE->>C: capacity.reserve(bytes) AtomicLong CAS 自旋
alt 超限
C-->>CORE: BudgetExceededException(RUN_BYTES)
CORE->>CORE: exhaustBudget(...)
end
end
APP->>B: snapshot() → RunBudgetUsage
Note over APP,B: Run 结束对账:各维度实际用量(模型轮数/工具次数/Token/字节)
```
## 3. 五个流程节点
1. **创建**:`startRun` 里 `new RunBudget(RunBudgetLimits)`,限额不可变,消耗状态可变,句柄随 RunContext 显式传递。
2. **模型调用前**:`reserveModelCall()` synchronized 预扣,超限抛异常。
3. **Tool 调用前**:`reserveToolCall(toolName)` 双重限额——总次数 + 单 Tool 次数。
4. **Canonical 写入前**:`reserveRunBytes(bytes)` 走 CAS 计数器。
5. **调用后**:`recordTokens` 三档 Token 上限记账。
**统一超限出口**:任何维度超限都抛带 `BudgetKind` 的 `BudgetExceededException` → Core `exhaustBudget`(固化终态 + 广播取消)→ 后续所有调用被 checkActive 拒绝。预算失败是 Run 级事实,不是局部异常。
## 4. 四个维度的时机对照表
| 维度 | 时机 | 预扣/记账 | 超限 BudgetKind |
|---|---|---|---|
| 模型轮数 | 模型调用前 | 预扣 | `MODEL_CALLS` |
| Tool 次数 | Tool 调用前 | 预扣 | `TOOL_CALLS` / `TOOL_CALLS_PER_TOOL` |
| Token | 模型调用后 | 记账(实际用量) | `INPUT_TOKENS` / `OUTPUT_TOKENS` / `TOTAL_TOKENS` |
| 字节 | canonical 写入前 | 预扣 | `RUN_BYTES` |
## 5. 自测:对着图能回答这四个问题吗
1. 第 7 轮模型调用时 `reserveModelCall` 超限——哪个组件抛异常、Run 变成什么终态、后续调用为什么全部被拒?
(ModelInterceptor 调 beforeModelCall → Core reserveModelCall 抛 BudgetExceededException → exhaustBudget 写 BUDGET_EXHAUSTED + cancel → 之后 checkActive 见终态直接抛 RunAbortedException)
2. 为什么 Tool 需要"总次数 + 单 Tool 次数"双重限额?
(总量防"调用太多",单 Tool 限额防"死磕同一个工具",比如反复查同一份日志)
3. 为什么字节预算用 CAS 自旋,计数预算用 synchronized?
(字节是高频原子累加,AtomicLong + compareAndSet 无锁乐观并发;计数是"读-判-写"复合操作,synchronized 保证原子性)
4. RunBudget 和 ModelCallLedger 都是记消耗,区别在哪?
(Budget 调用前预扣、管允不允许、超限会停止 Run;Ledger 调用后记账、管记了多少、幂等去重,只服务审计对账)
## 6. RunBudget 字段与组件速查
### 6.1 RunBudget 状态字段
| 字段 | 类型 | 含义 |
|---|---|---|
| `limits` | `RunBudgetLimits` | 限额定义(不可变) |
| `capacity` | `RunCapacityCounter` | 字节 CAS 计数器 |
| `modelCalls` | int | 累计模型调用轮数 |
| `toolCalls` | int | 累计工具调用次数 |
| `toolCallsByName` | `Map<String,Integer>` | 单工具名次数(防死磕) |
| `inputTokens` / `outputTokens` / `totalTokens` | long | 三档累计 token |
### 6.2 RunBudgetLimits:限额定义(7 字段 + 默认值)
| 字段 | 默认值 | 来源 |
|---|---|---|
| `maxModelCalls` | 24 | 配置 `harness.chat.*` |
| `maxToolCalls` | 24 | 配置 |
| `maxCallsPerTool` | 8 | 配置 |
| `maxInputTokens` | 100_000 | 配置 |
| `maxOutputTokens` | 100_000 | 配置 |
| `maxTotalTokens` | 200_000 | 配置 |
| `maxRunBytes` | 1_000_000(1MB) | 配置 |
### 6.3 支撑类型
| 类型 | 字段/枚举 | 作用 |
|---|---|---|
| `RunCapacityCounter` | `maxBytes` + `usedBytes`(AtomicLong) | 字节 CAS 自旋计数 |
| `RunBudgetUsage` | modelCalls / toolCalls / toolCallsByName / 三档 token / runBytes | `snapshot()` 只读快照,Run 结束对账 |
| `BudgetExceededException` | `kind` / `limit` / `attempted` | 超限异常,携带具体维度 |
| `BudgetKind` | 7 个枚举 | `MODEL_CALLS` / `TOOL_CALLS` / `TOOL_CALLS_PER_TOOL` / `INPUT_TOKENS` / `OUTPUT_TOKENS` / `TOTAL_TOKENS` / `RUN_BYTES` |
### 6.4 持有与消费组件
| 组件 | 与 budget 的关系 |
|---|---|
| `DiagnosisHarnessCore` | **门禁枢纽**:持有 `RunBudgetLimits`,创建 `RunBudget`;`beforeModelCall` / `beforeToolCall` / `recordTokens` / `reserveRunBytes` 统一入口;超限走 `exhaustBudget` 升级终态 |
| `RunContext` | 持有 `RunBudget` 句柄,随请求显式传递 |
| `HarnessModelInterceptor` | 每轮模型调用:`beforeModelCall`(预扣)+ 调用后 `ModelCallAuditor.recordUsage`(→ `core.recordTokens`) |
| `HarnessToolInterceptor` | 工具请求:`beforeToolCall` 预扣 |
| `ToolBoundary` | `beforeToolCall` + **3 处** `reserveRunBytes`:request 字节、raw response 字节、agent_result 字节 |
| `GuardModelCall` / `DiagnosisAgentUseCase` / `EvidenceRepair` | 各自 `beforeModelCall` + `reserveRunBytes`(输入/输出/Draft/Repair 输入) |
Token 记账链路:`HarnessModelInterceptor.recordUsage` → `ModelCallAuditor.recordUsage` → `core.recordTokens` → `RunBudget.recordTokens`(三档检查)。
### 6.5 配置来源:两层字节控制
字节控制有两层,作用域不同:
```text
单次 payload 上限(每类内容独立闸门,不在 RunBudgetLimits 内):
diagnosis-max-query-bytes: 16384 查询输入
diagnosis-max-previous-turn-bytes: 16384 历史轮次
diagnosis-max-input-bytes: 49152 诊断输入合计
diagnosis-max-draft-bytes: 49152 Draft
semantic-max-input/output-bytes: 100000 / 10000
repair-max-input/output-bytes: 100000 / 48000
canonical-max-record-bytes: 1048576 单条 canonical 记录
canonical-max-agent-result-bytes: 65536 agent_result
Run 累计上限:
max-run-bytes: 1000000 整个 Run 累计预扣
```
**注意**:`canonicalMaxRecordBytes(1MB)` 是"单条记录"上限,`maxRunBytes(1MB)` 是整个 Run 累计上限——两者都是 1MB 但作用域不同,一条记录就能占满 Run 预算的一半以上。
## 7. 预算异常与终态对应
### 7.1 异常 → 终态对应表
| 异常 | 抛出处 | 携带信息 | 写入/对应终态 |
|---|---|---|---|
| `BudgetExceededException` | `RunBudget`(reserve/record) | `kind` / `limit` / `attempted` | **BUDGET_EXHAUSTED**(写入) |
| `RunAbortedException`(deadline 超时) | `checkActive` 第二道闸 | `RunTermination(TIMED_OUT, ...)` | **TIMED_OUT**(主动写入后抛出) |
| `RunAbortedException`(已取消) | `checkActive` 第三道闸 | 已有终态(如 CANCELLED) | **读取**已有终态,不新写 |
| `RunAbortedException`(终态已存在) | `checkActive` 第一道闸 | 已有终态(可能是任何终态) | **读取**已有终态,不新写 |
| `IllegalArgumentException` | `RunBudgetLimits` 构造 / `RunBudget` 参数 | 校验信息 | **无终态**(Run 开始前 fail fast) |
### 7.2 两阶段异常:预算超限后的完整路径
```text
第一次(reserve/record 超限):
RunBudget 抛 BudgetExceededException(kind/limit/attempted)
→ Core 捕获 → exhaustBudget:finish(BUDGET_EXHAUSTED) + cancel
→ 异常继续向上抛(可审计"哪个维度爆了")
之后(任何 checkActive):
termination 已存在 → 抛 RunAbortedException(携带 BUDGET_EXHAUSTED 终态)
```
第一枪是预算异常(带维度),之后所有拦截是终止异常(带终态)——两者配合。
### 7.3 各消费组件的异常处理
| 组件 | 处理 |
|---|---|
| `DiagnosisHarnessCore.applyBudget` | catch `BudgetExceededException` → `exhaustBudget` → 再 throw |
| `GuardModelCall` | `ExecutionException` 的 cause 判断:`instanceof BudgetExceededException` → 原样 rethrow;`CancellationException` → `checkActive` → 可能抛 `RunAbortedException` |
| `HarnessRetryExecutor` | `BudgetExceededException` / `RunAbortedException` 属于**从不重试**类(取消、预算耗尽、协议错误不能被重试吞掉) |
## 8. 异常三要素:kind / limit / attempted
### 8.1 三个字段
| 字段 | 含义 | 例子 |
|---|---|---|
| `kind` | **哪个资源维度**超限(BudgetKind 枚举) | `MODEL_CALLS` |
| `limit` | 该维度的**限额**(来自 RunBudgetLimits) | `maxModelCalls=24` |
| `attempted` | 本次**试图达到的值**(尝试后的总量,不是超出差额) | `25` |
### 8.2 attempted 是"尝试后的总量",不是"超出的部分"
```java
int attempted = modelCalls + 1; // 尝试让计数变成多少
if (attempted > limits.maxModelCalls()) {
throw new BudgetExceededException(BudgetKind.MODEL_CALLS,
limits.maxModelCalls(), attempted);
}
modelCalls = attempted;
```
```text
已调用 24 次(正好达到上限)→ 第 25 次尝试:attempted=25 > 24
→ 抛异常:kind=MODEL_CALLS, limit=24, attempted=25
```
### 8.3 各维度实际值举例
| kind | limit(配置) | attempted | 含义 |
|---|---|---|---|
| `MODEL_CALLS` | 24 | 25 | 第 25 轮模型调用被拒 |
| `TOOL_CALLS` | 24 | 25 | 第 25 次工具调用被拒 |
| `TOOL_CALLS_PER_TOOL` | 8 | 9 | 某个工具第 9 次调用被拒(死磕拦截) |
| `INPUT_TOKENS` | 100_000 | 100_003 | 累计输入 token 超出 3 个 |
| `TOTAL_TOKENS` | 200_000 | 200_500 | 累计总 token 超出 |
| `RUN_BYTES` | 1_000_000 | 1_000_001 | Run 累计字节超出 1 字节 |
### 8.4 审计价值
- `kind` → 定位哪一类资源(token / 次数 / 字节)
- `limit` → 知道配置上限(是否配置太紧)
- `attempted` → 知道差多少爆的(贴线超限说明要调配置,暴涨说明有失控路径)
## 9. 关联文档
| 文档 | 用途 |
|---|---|
| [Harness 面试速查-一张图讲清设计](Harness面试速查-一张图讲清设计.md) | 面试主叙事 + 常见追问 |
| [Harness 设计-非确定性 Agent 的确定性控制边界](Harness设计-非确定性Agent的确定性控制边界.md) | 决策五:预算与信息增益双机制 |
| [Harness 信息增益停止-让无证据诊断正常收敛](Harness信息增益停止-让无证据诊断正常收敛.md) | 预算之外的第二套停止机制 |
| [Harness 异常处理-Loop 内外与状态流](Harness异常处理-Loop内外与状态流.md) | 预算耗尽如何落终态 |