diff --git a/devflow/glossary/CONTEXT.md b/devflow/glossary/CONTEXT.md
index 45b8951..c5a0228 100644
--- a/devflow/glossary/CONTEXT.md
+++ b/devflow/glossary/CONTEXT.md
@@ -181,7 +181,39 @@
### Evidence Status
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
-- 边界:`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
+- 边界:`EVIDENCE_FOUND` 只表示存在候选内容,不保证内容能够支持当前诊断;`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
+
+### Information Gain
+- 定义:一次 Tool 结果是否推进当前 Diagnosis Run 的语义评价,固定为 `GAINED` 或 `NO_GAIN`。
+- 边界:它评价的是结果对当前诊断的作用,不评价 Tool 产品质量;`NO_EVIDENCE` 和重复的规范化 `tool + scope` 可由 Harness 机械标记为 `NO_GAIN`,其他成功非空结果(包括 RAG `REFERENCE`)由模型评价。
+
+### Collection State
+- 定义:Diagnosis Harness 对当前 Run 是否允许继续收集证据的控制状态,固定为 `COLLECTING` 或 `SATURATED`。
+- 边界:状态由 Harness 维护;`SATURATED` 只表示连续 `NO_GAIN` 达到配置阈值,不包括硬预算耗尽。模型可以请求新的 Tool 调用,但不能绕过 `SATURATED`。
+
+### Diagnosis Stop Reason
+- 定义:Harness 停止当前 Run 继续调用 Tool 的内部原因,首版区分 `INFORMATION_SATURATED` 与 `BUDGET_LIMIT_REACHED`。
+- 边界:它用于控制、Trace 和 Release 输入,不是用户可见生命周期状态,也不进入模型上下文;真正的不可恢复技术故障走失败通道。
+
+### Progress Snapshot
+- 定义:Tool Loop 结束时,从当前 Run 的 Canonical Tool Result 一次性投影出的有界发布视图,用于生成已检查范围和客观结果。
+- 边界:Canonical Tool Result 是真理源;Progress Snapshot 不逐轮维护、不保存原始 Tool Response、Prompt 或内部 thought,也不直接进入模型上下文。
+
+### Safe Fallback Type
+- 定义:`SafeFallback.type` 对没有发布诊断结论的业务原因分类,例如 `INSUFFICIENT_EVIDENCE`、`MISSING_REQUIRED_CONTEXT`、`BUDGET_EXHAUSTED` 或安全校验失败。
+- 边界:它是 `ReleaseOutcome.FALLBACK` 的原因字段,不是与 `SUCCESS / FALLBACK / FAILED / CANCELLED` 平行的第二套生命周期状态。
+
+### Diagnosis Release Use Case
+- 定义:诊断业务发布的唯一决策入口,接收 DiagnosisDraft 和/或 Harness `stop_reason + ProgressSnapshot`,生成安全的 `SUCCESS / FALLBACK` 结果。
+- 边界:`conclusion=null` 不触发 EvidenceRepair;只有存在结论时才执行完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard 链路。不可形成安全业务内容的技术故障由 Chat Application Use Case 映射为 `FAILED / CANCELLED`。
+
+### Diagnosis Draft Contract Failure
+- 定义:Diagnosis Agent 最终文本为空、不是严格 JSON,或不满足 `DiagnosisDraft` Schema 时产生的 Agent 输出合同失败。
+- 边界:非法文本始终丢弃,不做 Markdown/自然语言提取,也不调用模型修复;仅当当前 Run 的 `ProgressSnapshot` 含已验真 observed facts 时,Release 才能确定性发布 `INSUFFICIENT_EVIDENCE`,否则保持 `FAILED`。它不是 `Diagnosis Stop Reason`,不得伪装成信息饱和或预算终止。
+
+### Model Observation
+- 定义:Tool 内部标准化结果经过白名单投影后,作为 Tool Response 进入 Diagnosis Agent 上下文的有界视图。
+- 边界:只包含模型完成语义判断和证据引用所需的信息;预算、阈值、重复指纹、原始相似度、原始 Tool Response 和完整 Harness 控制状态不得进入该视图。
### RunContext
- 定义:一次 Diagnosis Run 的显式执行上下文,结构不可变地携带 `sessionId`、`runId`、deadline,以及该 Run 独占的取消、预算、重试策略和生命周期状态句柄。
@@ -201,7 +233,7 @@
### Chat Application Use Case
- 定义:一次 Chat 请求的唯一业务入口,拥有 Session/Run、意图路由、固定执行器、PreviousTurn 和最终持久化。
-- 边界:不拥有 HTTP/SSE 连接,也不把 ChatModel 或 Tool 选择权交给 Controller。
+- 边界:不拥有 HTTP/SSE 连接,不把 ChatModel 或 Tool 选择权交给 Controller,也不在 Diagnosis Release Use Case 之外单独决定预算 Fallback 的业务内容。
### Chat SSE Contract
- 定义:Chat 公开入口的五事件协议,顺序固定为 `metadata -> status* -> content|failure -> done`。
diff --git a/devflow/projects/2026-07-26-diagnosis-information-gain-stop-contract/decisions.md b/devflow/projects/2026-07-26-diagnosis-information-gain-stop-contract/decisions.md
new file mode 100644
index 0000000..2cc2a70
--- /dev/null
+++ b/devflow/projects/2026-07-26-diagnosis-information-gain-stop-contract/decisions.md
@@ -0,0 +1,157 @@
+# Diagnosis 信息增益停止契约 Decisions
+
+## Discover Status
+
+- Checkpoint:Discover。
+- Capability source:`sm-flow` 内置 Discover 协议;grill 使用 `grill-with-docs`,代码可证问题通过源码、测试和引用搜索处理。
+- Scale:`complex`。变更跨越 Agent、Tool 协议、Run 生命周期、Release、Guard、配置、Trace 和 E2E。
+- 接口影响:L3。模型可见 Tool Schema 发生有意协议变更,公开 HTTP/SSE 和业务 Tool backend 协议不变。
+- 工具降级:当前没有 `codebase-retrieval` 和 LSP 工具;以 `rg`、源码和测试引用核查替代。用户已明确“可以忽略gitnexus”。
+
+## Question Pool
+
+| # | 维度 | 问题 | 模式 | 状态 |
+|---|---|---|---|---|
+| Q1 | 术语 | Tool 客观状态、信息增益、收集状态、停止原因和最终发布状态是否应合并为一个枚举? | user-interview | 已解决 |
+| Q2 | 语义 | Tool 返回价值由谁判断,是否需要多级质量分数? | user-interview | 已解决 |
+| Q3 | 协议 | 模型继续调用 Tool 时如何回传上一轮信息增益,是否需要 `next_action`? | user-interview | 已解决 |
+| Q4 | 边界 | Tool Schema 应来自 Prompt 还是服务端原生 Tool Calling 注册? | user-interview | 已解决 |
+| Q5 | 边界 | Harness 能确定性判定哪些 `NO_GAIN`,RAG `REFERENCE` 由谁判定? | user-interview | 已解决 |
+| Q6 | 范围 | 首版是否需要 `new_count` 或自然语言语义去重? | user-interview | 已解决 |
+| Q7 | 配置 | 连续无增益阈值是否可配置,默认值与生效时机是什么? | user-interview | 已解决 |
+| Q8 | 发布 | 信息饱和、预算终止和 `conclusion=null` 由谁转换为用户可见结果? | user-interview | 已解决 |
+| Q9 | 验收 | 如何证明未知问题不再以通用内部错误结束,同时不放过无证据结论? | evidence-driven | 已解决 |
+| Q10 | 技术 | 当前 Tool schema 是否能直接容纳 `previous_observation`? | evidence-driven | 已解决 |
+| Q11 | 技术 | 进展控制状态应扩展 Redis Store 还是放入 RunContext handle? | evidence-driven | 已解决 |
+| Q12 | 技术 | RAG `relevance_level` 在哪一层丢失,前端是否已有过程展示能力? | evidence-driven | 已解决 |
+
+## Evidence-driven
+
+| 结论 | 证据来源 | 是否已汇报用户 |
+|---|---|---|
+| 当前三个 Agent-facing Tool 直接使用 `RagToolRequest`、`QueryLogsRequest`、`MysqlToolRequest` 生成 Schema;要增加 `previous_observation + input` 必须显式演进 Tool Schema,不能只改 interceptor。 | `HarnessEvidenceTools`、三个 request records、`DiagnosisAgentFactory` | 已汇报 |
+| `HarnessToolInterceptor` 当前把完整 `agentResult` 放入 `ToolCallResponse.content`,控制视图与模型观察尚未分离。 | `HarnessToolInterceptor` | 已汇报 |
+| `RunContext` 已采用结构不可变、可变状态存在线程安全 handle 的模式;进展 tracker 放入 RunContext 比扩展 Redis 按 Run 枚举更符合现有所有权。 | `RunContext`、`DiagnosisHarnessCore.startRun` | 已汇报 |
+| `CanonicalInvocationStore` 只有 begin/find/markReady/markError,扩展按 Run 枚举会影响 Redis 实现和多组 fake store;首版可由 tracker 保存完成调用 key,在结束时按 key 读取 canonical 记录。 | `CanonicalInvocationStore` 及其引用测试 | 已汇报 |
+| `RagResultProjector` 只按 evidence 是否为空生成 `EVIDENCE_FOUND / NO_EVIDENCE`,没有读取上游 `relevanceLevel / relevance_level`。 | `RagResultProjector`、`LookupResult`、`KnowledgeEvidencePostProcessor` | 已汇报 |
+| `DiagnosisReleaseUseCase.execute` 当前强制 draft 非空并对所有 Draft 运行 EvidenceGuard;`EvidenceGuard` 又把空 analysis 判为 `ANALYSIS_MISSING`,与合法无结论结果冲突。 | `DiagnosisReleaseUseCase`、`EvidenceGuard` | 已汇报 |
+| `ChatApplicationUseCase.recoverBudgetExhaustion` 已有未提交预算 Fallback,但它绕过 Diagnosis Release,需迁移而不是丢弃用户价值。 | `ChatApplicationUseCase`、`SafeFallbackFactory`、现有测试 diff | 已汇报 |
+| 前端已渲染 `observed_facts / verified_sources / limitations / next_steps`,不需要新增公开展示协议。 | `src/main/resources/static/app.js` | 已汇报 |
+| 验收必须同时覆盖主动无结论、Harness 饱和、预算终止、无证据结论被 Guard 拦截,以及原始未知 Query 的 live SSE、日志和 exact run 数据。 | 当前事故现象、ISS-016 验收项、现有 E2E 工具 | 已汇报 |
+
+## User-interview
+
+| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
+|---|---|---|---|
+| 是否精简状态而不建立第二套生命周期? | “这里我觉得设计得太混乱了,怎么简化”以及对最终简化架构“我觉得可以” | 已确认 | 已回写 |
+| 是否只保留 `GAINED / NO_GAIN`? | “质量状态只要 information_gain = GAINED \| NO_GAIN 就够了?”后确认“我觉得可以” | 已确认 | 已回写 |
+| 是否需要 `new_count`? | “好那就去掉new_count” | 已确认 | 已回写 |
+| 是否需要 `next_action`? | “那就去掉next_action,我觉得由llm自己去判断就好了,不用显示的指定” | 已确认 | 已回写 |
+| Tool 如何注入? | “首先 tool注入,由服务端注入,而不是写死在提示词中” | 已确认 | 已回写 |
+| Prompt 是否强调合法放弃和正确但无用的内容? | “需要说明 模型不必须要给出一个答案”以及“如果你发现工具返回的是正确但对推导无用的废话,请停止调用” | 已确认 | 已回写 |
+| 阈值是否可配置? | “我觉得这个可以暴露出一个配置来控制” | 已确认 | 已回写 |
+| 首版重复检测做到什么程度? | “也就是说这一版只是做参数的去重校验”后确认“可以” | 已确认 | 已回写 |
+| 是否按当前简化方案进入实施? | “可以用更简单的方式”“可以。修正一下文档”“用sm-flow开始实施把” | 已确认 | 已回写,并授权完成 Commit 后进入 Apply |
+
+## 关键取舍
+
+- 决策:停止权归 Harness,语义价值判断由模型与确定性规则共同产生。
+ - 原因:Tool 只能知道客观返回,模型才能判断内容是否推进当前假设;但空结果和完全重复 scope 可由代码零 Token 判定。
+ - 影响:Harness 只消费二值信息增益,不引入独立 Judge 或质量分数。
+- 决策:使用下一次 Tool Call Envelope 回传上一轮模型评价。
+ - 原因:模型只有看到 Tool Observation 后才能评价,下一次真实行为正好提供受 Schema 约束的回传边界。
+ - 影响:这是 L3 Agent-facing Tool Schema 变更,业务 request 在 interceptor 内解包后保持不变。
+- 决策:进展 tracker 是 RunContext handle,Canonical Store 保持 Tool 真相源。
+ - 原因:停止决策需要低延迟 Run 内状态,完整证据仍应由 canonical 记录提供;两者职责不同。
+ - 影响:tracker 保存计数、scope、待评价调用和 canonical keys,不复制 raw payload。
+- 决策:不创建 ADR。
+ - 原因:这些是 ISS-016 范围内可通过 OpenSpec 回滚的内部协议演进,已有架构文档详细记录取舍,尚不满足独立 ADR 的必要性。
+
+## OpenSpec 回写
+
+- 必须进入 proposal/design/spec/tasks:Tool Envelope、L3 影响、Run tracker、确定性 `NO_GAIN`、RAG `relevance_level`、双视图、STOP_REQUIRED、ProgressSnapshot、统一 Release、Prompt、配置和 E2E。
+- 必须保留非目标:无 `new_count`、无 `next_action`、无 Judge、无语义去重、无第二套诊断生命周期、无公开 SSE 协议新增。
+- 当前没有未确认的 user-interview 问题,也没有 devflow/OpenSpec 冲突。
+
+## Cross-Artifact 对齐检查
+
+| 上游 → 下游 | 检查内容 | 状态 |
+|---|---|---|
+| ISS-016/架构文档 → proposal | 未知问题、合法放弃、信息增益、饱和停止、双视图、统一 Release、非目标和 E2E | 已对齐 |
+| proposal → design | L3 Envelope、Run tracker、scope、STOP_REQUIRED、ProgressSnapshot、预算终态、Prompt 和迁移方案 | 已对齐 |
+| design → specs/tasks | 状态机、Tool 门禁、RAG relevance、无结论 Guard、Release 所有权、Trace 和兼容边界 | 已对齐 |
+| specs → tasks | 每个可观察行为均有 contract/state/loop/release/config/E2E 可执行切片 | 已对齐 |
+
+### Gap 详情
+
+- 无。
+
+## Architecture Audit
+
+- Capability source:`zoom-out`。按 glossary 的 Diagnosis Agent、Diagnosis Harness、RunContext、Evidence Status、Invocation Status、Release Outcome 术语审计。
+- 顶层链路:`ChatApplicationUseCase -> DiagnosisChatExecutor -> DiagnosisAgentUseCase -> ReactAgent/interceptors -> ToolBoundary/canonical store -> ProgressSnapshot -> DiagnosisReleaseUseCase -> SSE/persistence`。
+- 所有权:Agent 负责诊断语义;Tool/Projector 负责客观结果;Run tracker 负责停止控制;Canonical Store 负责 Tool 真相;Guard 负责引用与结论安全;Release 负责用户可见 SUCCESS/FALLBACK;Application 只负责编排和持久化。
+- `RunContext` 的生产代码构造点只有 `DiagnosisHarnessCore.startRun`,大量测试通过该工厂获取;新增 tracker 不需要扩散手工构造。
+- `DiagnosisAgentUseCase`、`HarnessToolInterceptor`、`HarnessEvidenceTools` 和 `DiagnosisReleaseUseCase` 的直接消费者均已由配置类和 focused tests 覆盖,任务清单包含所有构造调用更新。
+- `FallbackType` 新语义只通过通用 SafeFallback JSON/前端渲染消费,没有前端枚举 switch;公开协议不新增字段。
+- 最大框架风险是 Tool Envelope Schema 和 STOP_REQUIRED 后异常传播;design 要求三个具体 record、真实 callback schema 测试和 scripted framework-loop 测试在 Release 迁移前锁定行为。
+- 最大生命周期风险是预算已把 RunLifecycle 置为 `BUDGET_EXHAUSTED` 后 Application 再次 `checkActive`;design 将其限制为“Diagnosis Release 已处理的预算 Fallback”窄分支,并禁止 Application 重建业务内容。
+- 审计结论:模块职责没有形成新的循环依赖或第二真相源;L3 风险已进入 specs 和 tasks,可进入 commit gate。
+
+## Commit Gate Preflight
+
+- `proposal.md`、`design.md`、六份 capability delta specs 和 `tasks.md` 均存在。
+- `openspec status --change diagnosis-information-gain-stop-contract --json` 返回 `isComplete=true`。
+- `openspec validate diagnosis-information-gain-stop-contract --strict` 通过。
+- question pool 全部已解决;evidence-driven 结论已汇报;user-interview 决策均有用户原话和确认状态。
+- 接口影响已判为 L3,并有独立 Interface Impact、兼容、迁移、回滚和验收说明。
+- Cross-artifact 检查无 gap;架构风险均已进入 design/tasks。
+- 用户已通过“用sm-flow开始实施把”明确授权 Commit 后进入 Apply。
+
+## Pre-apply Research
+
+### 参考实现
+
+- `HarnessEvidenceTools`:现有三类 `FunctionToolCallback` 注册点和 adapter bridge,继续作为 Agent-facing Schema 唯一入口。
+- `HarnessToolInterceptor`:可获得 exact framework Tool Call ID,适合消费 Envelope 和执行 progress gate。
+- `ToolBoundary`:Tool 预算、Run 校验、canonical 写入和安全错误的单点,不在 interceptor 重复 reserve。
+- `RunContext` / `DiagnosisHarnessCore.startRun`:结构不可变 + 可变 handle 模式和唯一生产构造点。
+- `RagResultProjector` / `QueryLogsResultProjector` / `MysqlResultProjector`:bounded canonical agent result 的现有标准化模式。
+- `EvidenceGuard` / `DiagnosisReleaseUseCase`:当前结论验证链和无结论冲突位置。
+- `ChatApplicationUseCase.recoverBudgetExhaustion`:保留用户价值、需要迁移所有权的临时预算 Fallback。
+- `DiagnosisAgentUseCaseTest.ScriptedChatModel`:真实框架 model -> Tool -> model loop 回归模式。
+
+### 技术栈清单
+
+- Tool Schema:三个具体 record 交给 Spring AI `FunctionToolCallback.inputType`,共享 `PreviousObservation`,不使用泛型擦除或 JsonNode Schema。
+- JSON:继续使用项目 `ObjectMapper` 严格解析/序列化;控制字段在 interceptor 消费后只传业务 input。
+- Run 状态:新增线程安全 tracker handle,由 `DiagnosisHarnessCore.startRun` 创建,不使用 ThreadLocal。
+- Canonical 真相:继续使用 `ToolCallKeyFactory + CanonicalInvocationStore.find`;tracker 只记录 identity。
+- 视图:从 bounded canonical `agent_result` 白名单投影 Model Observation,不读取 raw response。
+- 异常:受控停止使用专用异常和 cause-chain 分类;未知异常保持 fail closed。
+- 测试:JUnit 5、scripted ChatModel、现有 fake store/adapter fixture;不增加 Maven 依赖。
+
+### 新建基础设施
+
+- `harness.progress`:信息增益、收集状态、停止原因、tracker、scope、snapshot/projector。
+- `harness.agent`:三个 Agent-facing Envelope、白名单 observation projector、受控停止异常和执行结果。
+- 不新增数据库表、Redis 数据结构、HTTP DTO、SSE event 或外部依赖。
+
+## Apply 期间设计补充:Draft 合同失败
+
+- 真实 E2E `runId=4e667111-524e-4407-87ab-b4b262952017` 已完成一次 READY RAG 调用,第二轮模型返回文本后在 Draft/Release 边界失败;后续三次同 Query 均走零 Tool 的 `MISSING_REQUIRED_CONTEXT`,证明模型输出存在随机分支。
+- 用户确认采用窄化降级:非法 Draft 自身不被接受;已有当前 Run 的安全 ProgressSnapshot 时发布 `INSUFFICIENT_EVIDENCE`,没有安全过程时继续 `FAILED`。
+- 这是有意行为变更:从“所有非法 Draft 都发布技术失败”调整为“非法 Draft + 已验真过程可发布过程型 Fallback”;公开 SSE 字段、Tool 协议和最终生命周期枚举不变。
+- 不新增 stop reason,不把 Draft 解析失败伪装成 `INFORMATION_SATURATED` 或 `BUDGET_LIMIT_REACHED`;使用 Agent 输出异常携带有界 snapshot,并以脱敏 Trace 区分输出合同失败。
+
+## Apply Verification
+
+- Focused tests:`DiagnosisAgentUseCaseTest`、`DiagnosisReleaseUseCaseTest`、`DiagnosisChatExecutorTest`、`HarnessChatConfigurationTest` 通过。
+- 完整回归:`mvn -q -Dtest='!MilvusConnectionTest' test` 退出码为 `0`;本轮 Surefire 报告汇总 `Tests=292, Failures=0, Errors=0, Skipped=3`。
+- 外部凭据边界:未排除时唯一失败为 `MilvusConnectionTest.connect`,原因是当前测试进程未设置 `MILVUS_TOKEN`;这不是本变更回归。
+- OpenSpec:`openspec.cmd validate diagnosis-information-gain-stop-contract --strict` 通过。
+- 格式与清理:`git diff --check` 通过;未发现临时 E2E JSON、DEBUG 或 tmp 文件。
+- named SSE E2E:Query `诊断切换企业失败的问题`,`sessionId=iss016-final-20260726-a`,`runId=3ab22ed7-d0ed-45d8-b928-dce5790c0542`;SSE 返回 `SAFE_FALLBACK`、`type=MISSING_REQUIRED_CONTEXT`、`done.outcome=FALLBACK`。
+- 数据库核对:`status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=0`、`total_token_count=2890`、answer 非空。
+- Trace 核对:`RUN_STARTED -> ROUTING_ATTEMPT -> ROUTING_DECISION -> AGENT_MODEL_STEP -> EVIDENCE_GUARD_INITIAL -> RELEASE_DECISION/FALLBACK -> RUN_FINISHED/FALLBACK`。
+- 兼容性:公开 HTTP/SSE 字段、前端 SafeFallback 消费结构、数据库表和业务 Tool request 均未新增字段;模型侧 Tool Envelope 是本变更已确认的 L3 协议变更。
diff --git a/mvp/architecture/README.md b/mvp/architecture/README.md
index 6fa924e..be5159b 100644
--- a/mvp/architecture/README.md
+++ b/mvp/architecture/README.md
@@ -1,6 +1,6 @@
# MVP 架构文档
-**更新日期**:2026-07-23
+**更新日期**:2026-07-26
**状态**:当前单 Diagnosis Agent + Harness 架构
当前文档入口:
@@ -11,6 +11,7 @@
| [agent-orchestration.md](agent-orchestration.md) | 单 Diagnosis ReAct Agent 的职责、执行方式和 reasoning 采集边界 |
| [harness-quality-gates.md](harness-quality-gates.md) | Run、Tool、Evidence、Semantic、Release 与 Trace Recorder 门禁 |
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | sessionId/runId、SSE、统一 Timeline 和 reasoning audit 生命周期 |
+| [diagnosis-information-gain-stop-architecture.md](diagnosis-information-gain-stop-architecture.md) | 已实施的信息增益评价、Harness 饱和检测、Draft 合同失败降级与证据不足停止设计 |
2026-07-22 前的多角色编排、双入口和旧证据链文档已移动到 `archive/2026-07-22-legacy/`,仅用于历史决策追溯,不代表当前运行时。
diff --git a/mvp/architecture/diagnosis-information-gain-stop-architecture.md b/mvp/architecture/diagnosis-information-gain-stop-architecture.md
new file mode 100644
index 0000000..89211ae
--- /dev/null
+++ b/mvp/architecture/diagnosis-information-gain-stop-architecture.md
@@ -0,0 +1,608 @@
+# Diagnosis Agent 信息增益与停止控制架构
+
+**更新日期**:2026-07-27
+**状态**:已实施,待归档
+**关联 Issue**:[ISS-016](../issues/active/ISS-016-diagnosis-information-gain-stop-contract.md)
+
+## 1. 背景
+
+当前 Diagnosis ReAct Agent 在知识库没有直接答案、日志持续为空或查询条件不足时,可能继续改写查询并反复调用 Tool,直到资源预算耗尽。此时 Run 被发布为技术失败,用户只能看到通用错误,而不是“已完成有限排查,但证据不足”。
+
+本设计解决两个不同问题:
+
+1. 查询何时已经不再产生有效信息,必须停止继续调用 Tool。
+2. 停止后如何避免模型在证据不足时生成无依据结论。
+
+## 2. 核心原则
+
+```text
+Tool:返回客观数据
+Harness:判断查询过程是否还允许继续
+Model:判断数据是否推进了当前诊断,并通过实际 Tool Call 或 Draft 表达行为
+SemanticGuard:判断最终结论是否有证据支撑
+Release:统一把 Draft 或 Harness 停止投影为安全的 SUCCESS / FALLBACK
+```
+
+停止权属于 Harness。模型可以建议继续,但不能绕过 Harness 的饱和状态和硬资源边界。
+
+“没有找到根因”是合法业务结果,不等于执行失败。模型被允许输出 `conclusion=null` 的 `DiagnosisDraft`;最终仍只使用 `SUCCESS / FALLBACK / FAILED / CANCELLED`,不增加第二套诊断终态。
+
+## 3. 总体架构
+
+```mermaid
+flowchart TD
+ U["用户提出诊断问题"] --> M["Diagnosis Agent"]
+
+ P["ReAct Prompt
允许合法放弃"] -.-> M
+ C["会话上下文
有界历史和已检查范围"] -.-> M
+
+ M -->|"Tool Call Envelope"| B["Harness 调用门禁"]
+ B --> A["应用 previous_observation
更新连续 NO_GAIN"]
+ A --> GATE{"是否允许继续"}
+
+ GATE -->|"INFORMATION_SATURATED"| STOP["拒绝调用
注入 STOP_REQUIRED"]
+ GATE -->|"BUDGET_LIMIT_REACHED"| BSTOP["停止调用
记录资源保护原因"]
+ GATE -->|"允许调用"| T["剥离控制字段
执行业务 Tool"]
+
+ T -->|"执行失败"| ERR["Tool 状态:FAILED"]
+ ERR --> EP["技术故障策略
有限重试或 FAILED 结束"]
+
+ T -->|"执行成功"| CR["Canonical Tool Result
完整内部结果"]
+ CR --> CV["Harness Control View"]
+ CR --> AP["Agent Observation Projector"]
+ CR --> STORE["Canonical Store"]
+ AP --> MO["Model Observation
白名单有界结果"]
+ MO --> SM["模型阅读返回内容"]
+
+ CV --> O{"Harness 客观判断"}
+ O -->|"evidence_status=NO_EVIDENCE"| NG1["Harness 标记 NO_GAIN"]
+ O -->|"重复 Tool + 相同 scope"| NG1
+
+ O -->|"其他成功非空结果
包括 REFERENCE"| SM
+ SM --> SD{"语义信息增益"}
+ SD -->|"推进或排除诊断假设"| G["GAINED"]
+ SD -->|"没有可验证的新事实"| NG2["NO_GAIN"]
+
+ G --> M
+ NG1 --> M
+ NG2 --> M
+
+ M -->|"DiagnosisDraft"| SNAP["结束时投影 ProgressSnapshot"]
+ M -->|"空或非法 Draft"| DRAFTERR["丢弃非法内容
投影当前进展"]
+ DRAFTERR -->|"有已验真 observed facts"| SNAP
+ DRAFTERR -->|"无安全进展"| EP
+ STOP -->|"给模型一次合法完成机会"| M
+ STOP -->|"仍无合法 Draft"| SNAP
+ BSTOP --> SNAP
+ STORE --> SNAP
+
+ SNAP --> RELEASE["DiagnosisReleaseUseCase"]
+ RELEASE -->|"conclusion 非空"| GUARD["EvidenceGuard / Repair / SemanticGuard"]
+ RELEASE -->|"conclusion 为空或无 Draft"| SAFE["确定性 SafeFallback
不修复结论"]
+ GUARD -->|"证据支持"| SUCCESS["SUCCESS"]
+ GUARD -->|"无证据推断"| SAFE
+ SAFE --> FALLBACK["FALLBACK"]
+ EP --> FAILED["FAILED"]
+
+ SUCCESS --> APP["ChatApplicationUseCase
持久化"]
+ FALLBACK --> APP
+ FAILED --> APP
+ APP --> SSE["ChatSseSession
发送 content/failure + done"]
+```
+
+## 4. 状态模型
+
+状态按职责分开,不使用一个枚举表达整个生命周期。
+
+| 维度 | 状态 | 负责人 | 含义 |
+|---|---|---|---|
+| Tool 执行 | `SUCCEEDED / FAILED` | ToolBoundary | Tool 是否完成技术执行 |
+| 信息增益 | `GAINED / NO_GAIN` | Harness 或模型 | 本次结果是否推进当前诊断 |
+| 收集生命周期 | `COLLECTING / SATURATED` | Harness | 是否允许继续调用 Tool |
+| 内部停止原因 | `INFORMATION_SATURATED / BUDGET_LIMIT_REACHED` | Harness | 为什么停止继续调用 Tool |
+| Fallback 原因 | `SafeFallback.type` | DiagnosisReleaseUseCase | 为什么没有发布诊断结论 |
+| 最终发布 | `SUCCESS / FALLBACK / FAILED / CANCELLED` | Release | 对外发布结果 |
+
+错误码、停止原因、Fallback 原因和 RAG 相关度是附带字段,不提升为全局生命周期状态。`CONFIRMED / INSUFFICIENT_EVIDENCE / NEED_MORE_INFO` 不作为第二套诊断状态;其中后两种语义由 `SafeFallback.type` 表达。
+
+当前实现中的兼容映射是:内部 `PROJECTING` 不对外暴露,`READY` 对应目标语义 `SUCCEEDED`,`ERROR` 对应目标语义 `FAILED`。
+
+不定义独立的模型行动状态。模型发起 Tool Call 表示继续收集,输出正常 `DiagnosisDraft` 表示完成诊断,输出证据不足 Draft 表示主动停止;Harness 直接从实际输出推导行为。
+
+## 5. Tool 结果与上下文边界
+
+Tool 不判断业务结论,只提供事实和可机械计算的元信息。
+
+```text
+tool_call_id 本次调用的唯一标识
+execution_status SUCCEEDED | FAILED
+returned_count 本次实际返回的记录数量
+scope 本次查询实际覆盖的结构化范围
+metadata Tool 特有的有限元信息
+result 提供给模型的有界结果
+```
+
+字段说明:
+
+- `returned_count` 只表达“返回了多少条”,不表示这些内容有诊断价值。
+- `scope` 用于识别相同 Tool、相同查询范围的重复调用,例如企业、时间窗、服务和过滤条件。
+- `metadata` 保存 Tool 特有信息,例如 RAG 的 `relevance_level`、是否截断和分页信息。
+- 第一版不引入 `new_count`。它要求为不同 Tool 建立稳定的结果指纹,而且“新数据”也不等于“有效数据”。
+
+### 5.1 客观信号来源
+
+这些信号不是模型推导的:
+
+| 信号 | 来源 | 含义 |
+|---|---|---|
+| `PROJECTING / READY / ERROR` | ToolBoundary | 当前实现的调用生命周期;目标外部语义映射为 `SUCCEEDED / FAILED` |
+| `EVIDENCE_FOUND / NO_EVIDENCE / ERROR` | Tool Result Projector / ToolBoundary | 是否存在候选结果或发生技术错误 |
+| `PRECISE / HIGHLY_RELEVANT / REFERENCE / null` | RAG 后处理规则 | RAG 候选内容的相关度 |
+| 重复 `tool + scope` | Harness | 本次查询范围是否与历史调用等价 |
+
+`NO_EVIDENCE` 由各 Projector 根据结果集合是否为空确定。`REFERENCE` 由 RAG 根据归一化相似度及查询提示命中情况计算。两者都不是 LLM 生成的状态,但职责不同:`NO_EVIDENCE` 可由 Harness 机械判定为 `NO_GAIN`;`REFERENCE` 只表示候选内容相关度一般,仍由模型判断是否推进当前诊断。
+
+改造前 `RagResultProjector` 只根据 `evidence` 是否为空生成 `EVIDENCE_FOUND / NO_EVIDENCE`,没有保留上游 `relevanceLevel`,因此会出现:
+
+```text
+relevanceLevel = REFERENCE
+evidenceBlocks 非空
+ -> RagResultProjector
+ -> evidence_status = EVIDENCE_FOUND
+```
+
+这不是两个判断冲突,而是 Projector 丢失了相关度维度。当前实现已兼容读取上游 `relevanceLevel` 或 `relevance_level`,并统一保留为 `relevance_level`。
+
+### 5.2 Tool 双视图架构
+
+内部控制信息和模型观察不能继续共用一个无差别 JSON。标准化结果产生两个明确视图:
+
+```mermaid
+flowchart TD
+ A["Tool 原始返回"] --> B["结果标准化"]
+ B --> C["Canonical Tool Result
完整内部结果"]
+
+ C --> D["Harness Control View"]
+ C --> E["Agent Observation Projector"]
+
+ D --> F["Harness
重复检测、低收益计数、预算、饱和状态"]
+ D --> G["Canonical Store"]
+
+ E --> H["Model Observation
最小必要字段"]
+ H --> I["ToolCallResponse.content"]
+ I --> J["模型上下文"]
+```
+
+改造前会把完整 `agent_result` 作为 `ToolCallResponse.content` 交给模型。当前实现已改为白名单投影,Harness 控制字段默认不进入模型上下文。
+
+字段边界:
+
+| 字段 | Harness | 模型 | 说明 |
+|---|---:|---:|---|
+| `tool_call_id` | 是 | 是 | 引用与归属校验 |
+| 实际查询 `scope` | 是 | 是 | 重复检测与有界负向观察 |
+| 有界 `evidence` | 是 | 是 | 模型语义判断所需内容 |
+| `evidence_status` | 是 | 是 | 候选结果是否为空或失败 |
+| `relevance_level` | 是 | 是 | 只暴露粗粒度标签,不暴露原始分数 |
+| `truncated` | 是 | 是 | 提醒模型结果并不完整 |
+| `returned_count` | 是 | 否 | Harness 客观统计 |
+| 原始相似度和检索轨迹 | 是 | 否 | 内部质量与审计信息 |
+| 规范化 scope、重复指纹 | 是 | 否 | Harness 控制信息 |
+| 低收益计数、阈值、预算 | 是 | 否 | Run 内部状态 |
+| 原始 Tool Response | 是 | 否 | 不进入模型上下文 |
+
+只有当 Harness 必须改变模型行为时,才注入有界控制指令,例如:
+
+```json
+{
+ "stop_required": true,
+ "reason": "INFORMATION_SATURATED",
+ "checked_scopes": ["本 Run 已检查范围的有界摘要"]
+}
+```
+
+计数器、阈值和剩余预算不随控制指令进入模型上下文。
+
+### 5.3 Tool 调用与投影流程
+
+```mermaid
+sequenceDiagram
+ participant Model as Diagnosis Agent
+ participant Interceptor as Harness Tool Interceptor
+ participant Gate as Harness Gate
+ participant Adapter as Tool Adapter
+ participant Backend as Tool Backend
+ participant Projector as Result Normalizer / Projector
+ participant Progress as Progress Tracker
+ participant Store as Canonical Store
+
+ Model->>Interceptor: Tool Call Envelope
+ Interceptor->>Progress: 校验并应用 previous_observation(如有)
+ Progress-->>Interceptor: 更新后的收集状态
+ Interceptor->>Gate: 校验 Run、授权、预算和饱和状态
+ alt 已经 SATURATED
+ Gate-->>Interceptor: STOP_REQUIRED
+ Interceptor-->>Model: 有界停止指令
+ else 达到硬预算
+ Gate-->>Interceptor: BUDGET_LIMIT_REACHED
+ else 允许调用
+ Gate-->>Interceptor: ALLOW
+ Interceptor->>Interceptor: 剥离 previous_observation
+ Interceptor->>Adapter: input 中的 typed request
+ Adapter->>Backend: 执行只读查询
+ Backend-->>Adapter: raw response
+ Adapter->>Projector: 标准化原始结果
+ Projector-->>Progress: Harness Control View
+ Projector-->>Store: Canonical Tool Result
+ Projector-->>Interceptor: Model Observation
+ Interceptor-->>Model: 有界 Tool Observation
+ end
+```
+
+重复查询由 Harness 根据 `tool_name + normalized_scope` 判断,不属于 Tool 返回状态。规范化只处理确定性的业务参数,例如时间格式、无序数组和非业务字段;第一版不判断自然语言改写是否语义等价。
+
+### 5.4 结束时进展投影
+
+每次 Tool 调用完成后只追加 Canonical Tool Result,不在每轮维护另一份用户可见摘要。Tool Loop 因 Agent 输出 Draft、信息饱和或预算限制结束时,一次性生成内部 `ProgressSnapshot`:
+
+```mermaid
+flowchart LR
+ T["每次 Tool 完成"] --> C["Canonical Tool Result"]
+ C --> S["Canonical Store"]
+ S -->|"Tool Loop 结束时只读投影"| P["ProgressSnapshot"]
+ P --> R["DiagnosisReleaseUseCase"]
+ R --> O["SafeFallback.observed_facts"]
+```
+
+`ProgressSnapshot` 不是新的真理源,也不进入模型上下文。它只包含来源、实际 scope、客观结果摘要、截断标记和 Harness `stop_reason`;原始 Tool Response、Prompt、内部 thought、计数器和剩余预算不进入发布内容。
+
+现有 `SafeFallback` 和前端已经支持 `observed_facts / verified_sources / limitations / next_steps`,因此不新增前端协议。`FallbackType` 增加 `INSUFFICIENT_EVIDENCE / MISSING_REQUIRED_CONTEXT`,分别表达有限排查后证据不足和缺少有效查询条件。
+
+## 6. 信息增益契约
+
+信息增益只保留两个值:
+
+```text
+information_gain = GAINED | NO_GAIN
+```
+
+### 6.1 GAINED
+
+满足以下任一条件:
+
+- 返回了与当前问题直接相关、可验证的新事实。
+- 新事实确认了一个当前诊断假设。
+- 新事实排除了一个当前诊断假设。
+- 新事实实质性缩小了故障范围。
+
+“排除假设”也是信息增益。信息增益不要求得到最终根因。
+
+### 6.2 NO_GAIN
+
+满足以下任一条件:
+
+- 内容为空、重复或只有通用参考资料。
+- 数据虽然非空,但与当前问题没有直接关系。
+- 没有新增可验证事实,也没有改变任何诊断假设。
+- 只是建议继续查询另一个位置,没有提供新的诊断事实。
+- 模型无法判断结果是否推进诊断。
+
+不增加 `UNKNOWN`。无法判断时归入 `NO_GAIN`,避免它成为无限继续查询的出口。
+
+### 6.3 谁负责赋值
+
+```text
+FAILED
+ -> 不产生 information_gain,进入技术故障流程
+
+SUCCEEDED + evidence_status=NO_EVIDENCE
+或重复的 tool_name + normalized_scope
+ -> Harness 直接赋 NO_GAIN
+
+其他 SUCCEEDED 非空结果,包括 relevance_level=REFERENCE
+ -> 模型必须赋 GAINED 或 NO_GAIN
+```
+
+模型只判断“本次结果是否推进当前诊断”,不评价 Tool 产品质量,也不输出 0 到 100 的主观分数。
+
+## 7. 模型步骤契约
+
+模型收到未被 Harness 客观标记为 `NO_GAIN` 的成功非空结果后,如果要继续调用 Tool,必须在下一次 Tool Call 中给出上一调用的信息增益。该设计不要求模型显式声明下一步行动。
+
+```json
+{
+ "previous_observation": {
+ "tool_call_id": "call-123",
+ "information_gain": "NO_GAIN"
+ },
+ "input": {
+ "query": "下一次查询参数"
+ }
+}
+```
+
+约束:
+
+1. `previous_observation` 只在上一轮存在待模型评价的 Tool Observation 时必填;首次调用以及上一轮已由 Harness 客观赋值时省略。
+2. `tool_call_id` 必须指向当前 Run 中最后一个尚未评价的成功 Tool 调用。
+3. Harness 必须先校验并应用 `information_gain`,再判断是否允许执行本次 Tool 调用。
+4. Harness 消费并剥离 `previous_observation`,业务 Tool 只接收 `input` 中的原有业务参数。
+5. 上一个待语义评价的结果未被评价时,Harness 不接受新的 Tool 调用。
+6. Harness 已经进入 `SATURATED` 时,新的 Tool Call 被拒绝,并向模型注入 `STOP_REQUIRED`。
+7. 模型主动判断没有合理查询方向时,应直接输出证据不足 Draft,不必等待 Harness 强制停止,也无需为放行下一次调用而回传最后一轮评价。
+
+模型行为直接从实际输出推导:
+
+```text
+Tool Call -> 继续收集
+正常 DiagnosisDraft -> 完成诊断
+证据不足 DiagnosisDraft -> 主动停止
+```
+
+该 Envelope 是模型侧 Tool 调用协议。服务端为动态注册的 Tool Schema 统一增加 Envelope,Harness 在边界处消费控制字段,业务 Tool 的入参协议保持不变。这是一项有意的协议变更,影响模型 Tool Schema 生成、Tool Call 解析、Harness 拦截和相关测试,不影响 Tool Backend。
+
+## 8. Harness 饱和规则
+
+第一版使用连续低收益计数,不维护复杂进展账本。硬调用上限是独立资源保护,不作为信息饱和条件。
+
+```text
+evidence_status=NO_EVIDENCE -> consecutiveLowYield + 1
+重复 Tool + 相同 scope -> consecutiveLowYield + 1
+成功非空结果(包括 REFERENCE)+ 模型 NO_GAIN
+ -> consecutiveLowYield + 1
+其他成功非空结果 + 模型 GAINED -> consecutiveLowYield = 0
+FAILED -> 技术故障流程,不计入低收益
+
+consecutiveLowYield >= stopAfterConsecutiveNoGain
+ -> SATURATED
+ -> stop_reason=INFORMATION_SATURATED
+达到单 Tool 或 Run 硬调用上限 -> stop_reason=BUDGET_LIMIT_REACHED
+```
+
+配置项 `stop-after-consecutive-no-gain` 默认值为 `2`、必须大于等于 `1`,只在 Run 启动时读取并固定,不进入模型上下文。默认值需要通过固定 E2E 和评测集校准,不作为不可调整的业务真理。
+
+伪代码:
+
+```text
+onToolResult(result):
+ if result.execution_status == FAILED:
+ handleTechnicalFailure(result)
+ return
+
+ if isEmpty(result) or isRepeatedScope(result):
+ applyInformationGain(NO_GAIN)
+ return
+
+ requireModelAssessment(result.tool_call_id)
+
+applyInformationGain(gain):
+ if gain == GAINED:
+ consecutiveLowYield = 0
+ else:
+ consecutiveLowYield += 1
+
+ if consecutiveLowYield >= stopAfterConsecutiveNoGain:
+ collectionState = SATURATED
+ stopReason = INFORMATION_SATURATED
+
+onHardLimitReached():
+ stopReason = BUDGET_LIMIT_REACHED
+ stopToolCollection()
+```
+
+`SATURATED` 是当前 Run 的 Tool 收集终态。模型不能在同一 Run 中把它恢复为 `COLLECTING`。用户补充新范围或新事实后,应开启新的诊断 Run。
+
+## 9. 控制与发布边界
+
+### 9.1 Harness 层:物理停止
+
+Harness 使用 `NO_EVIDENCE`、重复的 `tool_name + normalized_scope` 和连续 `NO_GAIN` 检测信息饱和。RAG `REFERENCE` 与其他成功非空结果一样,由 Diagnosis Agent 通过下一次 Tool Call Envelope 回传语义评价。硬调用上限只产生 `BUDGET_LIMIT_REACHED`,不伪装成信息饱和。
+
+### 9.2 SemanticGuard 层:无证据推断兜底
+
+只有 `DiagnosisDraft.conclusion` 非空时才进入完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard 链路。SemanticGuard 在最终发布前检查:
+
+- 正常诊断中的关键结论是否绑定已验证证据。
+- `REFERENCE` 和通用资料是否被错误当作当前故障事实。
+- scoped `NO_EVIDENCE` 是否被错误解释为“故障不存在”。
+- 证据不足时是否生成了确定性根因。
+
+发现无证据推断时,不继续 Tool Loop,而是转换为有界的证据不足结果。
+
+`conclusion=null` 是合法业务结果,不触发 EvidenceRepair,也不要求额外调用 SemanticGuard 来证明“没有结论”。已有 analysis 时只校验其 Tool 引用真实性;Release 从 `ProgressSnapshot` 和 `limitations` 生成安全 Fallback。零次 Tool 调用且 `limitations.missing_info` 非空时,直接生成 `MISSING_REQUIRED_CONTEXT` Fallback。
+
+### 9.3 ReAct Prompt 层:合法放弃许可证
+
+Prompt 必须明确:
+
+- 不要求模型必须得出诊断结论或根因,但必须输出一个诚实、有界的完整 Draft。
+- 无法得到足够证据时,`conclusion=null` 是成功完成,不是失败。
+- Tool 调用次数可以为零。缺少企业、时间范围、服务或错误信息,且不存在预期能产生新诊断信息的明确查询时,直接在 `limitations.missing_info` 中列出缺口。
+- 不得为了表现“已经排查”而调用没有明确目标的 Tool。
+- Tool 返回内容事实正确、表述完整或结果非空,不代表它对当前诊断有信息增益。
+- 通用知识、背景说明、重复内容或不能改变当前判断的内容属于 `NO_GAIN`。
+- `NO_GAIN` 后不得通过改写相似关键词或重复相同 scope 继续尝试。
+- 如果仍存在范围明确、可能产生新诊断信息的不同查询,可以继续;否则应主动停止。
+- 不得为了表现“尽力”而重复或扩大无明确目标的查询。
+- 收到 `STOP_REQUIRED` 后不得继续调用 Tool。
+
+Prompt 不包含 Tool 名称、Tool Schema 或需要展开 Schema 的文本占位符。Tool 只通过服务端的模型原生 Tool Calling 通道注册;通用 Envelope 由服务端包装动态 Tool Schema。
+
+### 9.4 会话上下文层:历史边界认知
+
+每条 Tool 结果只以白名单 `Model Observation` 进入模型上下文一次。Harness 不在后续步骤中重复注入完整 Tool 结果、原始响应或内部进展账本。
+
+只有 Harness 必须改变模型行为时,才注入有界停止指令:
+
+```text
+stop_required
+stop_reason
+已检查 Tool 和 scope 的有界摘要
+```
+
+不向模型注入 `consecutiveLowYield`、阈值、剩余预算、重复指纹、原始相似度、完整 Tool 原始载荷或内部 thought。
+
+### 9.5 Release 层:统一安全发布
+
+`DiagnosisReleaseUseCase` 是诊断业务发布的唯一决策入口:
+
+```text
+DiagnosisDraft.conclusion 非空
+ -> EvidenceGuard
+ -> 必要时 EvidenceRepair
+ -> SemanticGuard
+ -> SUCCESS 或 FALLBACK
+
+DiagnosisDraft.conclusion 为空
+ -> 不执行 EvidenceRepair
+ -> 校验已有 Tool 引用(如有)
+ -> ProgressSnapshot + limitations
+ -> FALLBACK
+
+Harness 已停止且没有 DiagnosisDraft
+ -> stop_reason + ProgressSnapshot
+ -> 确定性 FALLBACK
+
+最终 Draft 为空或违反严格 JSON/Schema 合同
+ -> 丢弃非法 Draft,不做宽松提取或模型修复
+ -> ProgressSnapshot 有已验真 observed facts:INSUFFICIENT_EVIDENCE
+ -> 无安全进展:FAILED
+```
+
+Draft 合同失败不是新的 `stop_reason`。Trace 只记录固定失败类别、输出字节数和是否存在可发布进展,不记录模型原文、字段值或解析异常文本。
+
+`ChatApplicationUseCase` 只负责编排、持久化,以及把无法形成安全业务内容的不可恢复故障映射为 `FAILED / CANCELLED`。预算 Fallback 不再由它单独构造。`ChatSseSession` 只发送 `content|failure + done`,不判断诊断语义。
+
+## 10. 生命周期
+
+```mermaid
+stateDiagram-v2
+ [*] --> COLLECTING
+
+ COLLECTING --> COLLECTING: GAINED / 低收益清零
+ COLLECTING --> COLLECTING: NO_GAIN / 未达到阈值
+ COLLECTING --> SATURATED: 连续 NO_GAIN 达到阈值
+ COLLECTING --> BUDGET_STOP: 达到硬调用上限
+ COLLECTING --> TECHNICAL_FAILURE: 不可恢复的 Tool 或基础设施故障
+ COLLECTING --> DRAFT_READY: 模型输出 DiagnosisDraft
+ COLLECTING --> DRAFT_INVALID: 最终 Draft 为空或违反合同
+ COLLECTING --> CANCELLED: 用户取消
+
+ SATURATED --> RELEASE_INPUT: INFORMATION_SATURATED + ProgressSnapshot
+ BUDGET_STOP --> RELEASE_INPUT: BUDGET_LIMIT_REACHED + ProgressSnapshot
+ DRAFT_READY --> RELEASE_INPUT: Draft + ProgressSnapshot
+ DRAFT_INVALID --> RELEASE_INPUT: 有已验真 observed facts
+ DRAFT_INVALID --> FAILED: 无安全进展
+
+ RELEASE_INPUT --> EVIDENCE_GUARD: conclusion 非空
+ RELEASE_INPUT --> FALLBACK: conclusion 为空或无 Draft
+
+ EVIDENCE_GUARD --> SEMANTIC_GUARD: 引用有效或修复成功
+ EVIDENCE_GUARD --> FALLBACK: 引用无法安全验证
+ SEMANTIC_GUARD --> SUCCESS: 证据支持最终结论
+ SEMANTIC_GUARD --> FALLBACK: 无证据推断
+ TECHNICAL_FAILURE --> FAILED
+
+ SUCCESS --> [*]
+ FALLBACK --> [*]
+ FAILED --> [*]
+ CANCELLED --> [*]
+```
+
+## 11. 示例
+
+查询:`诊断切换企业失败的问题`
+
+```text
+第 1 次 lookup_knowledge
+ execution_status = SUCCEEDED
+ returned_count = 5
+ relevance_level = REFERENCE
+ information_gain = NO_GAIN(模型)
+ consecutiveLowYield = 1
+
+第 2 次 query_logs
+ execution_status = SUCCEEDED
+ returned_count = 0
+ information_gain = NO_GAIN(Harness)
+ consecutiveLowYield = 2
+
+Harness
+ collectionState = SATURATED
+ 拒绝新的 Tool 调用
+ 注入 STOP_REQUIRED
+
+最终结果
+ stop_reason = INFORMATION_SATURATED
+ release_outcome = FALLBACK
+ SafeFallback.type = INSUFFICIENT_EVIDENCE
+ 展示已检查范围、客观结果、限制和需要补充的信息
+ 不发布 INTERNAL_FAILURE
+```
+
+如果第二次查询返回了能够排除某个假设的日志,模型应标记 `GAINED`,低收益计数清零,允许继续进行有目标的诊断。
+
+如果用户首次请求即缺少企业、时间范围等有效查询条件,模型可以零次 Tool 调用直接输出 `conclusion=null`,把缺口写入 `limitations.missing_info`。Release 将其发布为 `FALLBACK + MISSING_REQUIRED_CONTEXT`,不执行 EvidenceRepair 或 SemanticGuard。
+
+## 12. 非目标
+
+- 不通过简单增加 Tool、Token 或超时预算解决空转。
+- 不让 Tool 或 Harness 判断业务根因。
+- 不增加多级质量分数、`UNKNOWN` 或复杂状态矩阵。
+- 不引入 `new_count` 和跨 Tool 通用内容指纹。
+- 不新增独立 Progress Judge 模型调用。
+- 不重新引入 Planner/Executor/Verifier/Composer 多角色 Graph。
+
+## 13. 模型 Token 与 Tool 拒绝审计
+
+Run 使用 `RunBudget` 作为 Token 总账,`diagnosis_trace_event` 作为模型调用明细账,`AgentStep.token_count` 只作为 Diagnosis Agent 轮次摘要,不新增独立 Token 表。
+
+```mermaid
+flowchart LR
+ R["Router / System / Knowledge / Repair / Semantic"] --> G["GuardModelCall"]
+ A["Diagnosis ReAct round"] --> I["HarnessModelInterceptor"]
+ G --> U["Provider Usage"]
+ I --> U
+ U --> L["Run ModelCallLedger"]
+ U --> B["RunBudget Token 总账"]
+ U --> T["MODEL_TOKEN_USAGE Trace"]
+ I --> S["AgentStep.token_count"]
+ L --> F["RUN_FINISHED 对账摘要"]
+ B --> F
+```
+
+每个 `MODEL_TOKEN_USAGE` 只记录:
+
+- `component` 与 `component_round`;
+- `usage_available`;
+- Usage 可用时的 `input_tokens / output_tokens / total_tokens`。
+
+Usage 不可用时不写 Token 字段,不把未知值记成零;Run 结束以 `usage_unavailable_count` 和 `tokens_reconciled=false` 暴露缺口。`RUN_FINISHED` 同时记录预算总量与审计明细合计,便于 exact Run 对账。
+
+Tool 请求进入 `HarnessToolInterceptor` 后若被进展协议、重复 scope、信息饱和或观察合同门禁拒绝,写入 `TOOL_REQUEST_REJECTED`。事件只含安全 Tool Call ID、Tool name 和稳定错误码,不含业务参数、normalized scope 正文、原始响应或内部异常。实际执行仍只由 `TOOL_INVOCATION` 表达,因此两类数量不能混用。
+
+真实审计发现:`INVALID_PROGRESS_PROTOCOL` 已可完整观察,但当前不会增加 `NO_GAIN` 或触发信息饱和。连续协议拒绝可能产生高 Token 空转,这属于待确认的停止行为改造,不属于 Token 审计本身。
+
+## 14. 验收标准
+
+- [x] Tool 执行状态、信息增益、收集状态和发布状态职责分离。
+- [x] RAG `REFERENCE` 不再自动等同于诊断证据。
+- [x] `NO_EVIDENCE` 和重复的 `tool_name + normalized_scope` 被 Harness 确定性标记为 `NO_GAIN`;`REFERENCE` 由模型评价。
+- [x] 模型继续调用 Tool 前,对上一轮待评价结果给出 `GAINED` 或 `NO_GAIN`。
+- [x] 下一次 Tool Call Envelope 能回传上一轮语义评价;Harness 应用评价后才决定是否放行,并在调用业务 Tool 前剥离控制字段。
+- [x] 不引入独立 `next_action`;Harness 从 Tool Call 或 Draft 等实际输出推导模型行为。
+- [x] `stop-after-consecutive-no-gain` 可配置且默认值为 `2`;连续低收益达到阈值后 Harness 阻止新的 Tool 调用。
+- [x] 只对 `tool_name + normalized_scope` 做确定性去重,第一版不承诺自然语言语义去重。
+- [x] `INFORMATION_SATURATED` 与 `BUDGET_LIMIT_REACHED` 分开,硬预算不进入 `SATURATED`。
+- [x] `FAILED` 不被统计为 `NO_GAIN`,技术故障与证据不足保持区分。
+- [x] 模型可以零次 Tool 调用输出 `conclusion=null + limitations.missing_info`,收到 `STOP_REQUIRED` 后必须停止。
+- [x] `conclusion=null` 不触发 EvidenceRepair;只有存在结论时才执行完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard 链路。
+- [x] SemanticGuard 能把无证据确定性结论转换为安全的证据不足结果。
+- [x] `DiagnosisReleaseUseCase` 统一处理 Draft、信息饱和和预算终止,并复用 `SafeFallback.observed_facts` 展示 `ProgressSnapshot`。
+- [x] 空或非法 Draft 仅在有已验真进展时降级为 `INSUFFICIENT_EVIDENCE`,无安全进展继续 fail closed。
+- [x] 不引入 `CONFIRMED / INSUFFICIENT_EVIDENCE / NEED_MORE_INFO` 第二套生命周期状态;后两种语义由 `SafeFallback.type` 表达。
+- [x] Tool Schema 只通过模型原生 Tool Calling 通道注册,不拼接进 Prompt。
+- [x] 原始复现 Query 在预算耗尽前正常收敛,不再发布通用 `INTERNAL_FAILURE`。
+- [x] Trace 能解释每轮信息增益和停止原因,但不保存原始 Tool 载荷或内部 thought。
+- [x] 模型调用 Token 可按组件和轮次对账,Run 总账与明细账差异显式可见。
+- [x] Tool 请求拒绝与实际 Tool invocation 分开审计,且拒绝事件不泄露 payload。
diff --git a/mvp/architecture/harness-quality-gates.md b/mvp/architecture/harness-quality-gates.md
index d2a433d..2ad82d9 100644
--- a/mvp/architecture/harness-quality-gates.md
+++ b/mvp/architecture/harness-quality-gates.md
@@ -1,6 +1,6 @@
# Harness 与质量门禁
-**更新日期**:2026-07-23
+**更新日期**:2026-07-27
**状态**:当前可运行架构
## 1. Harness 定位
@@ -57,6 +57,16 @@ SemanticGuard 使用隔离的单轮模型调用,只接收原始 query、完整
Trace 写入失败只记录警告,不应改变业务执行结果;`details` 禁止包含 Prompt、Thought、Draft 正文或 raw Tool payload。普通 Trace API 按 `sequence_no, id` 返回 Timeline。
+### 6.1 Token 对账
+
+所有 Harness 模型入口统一从 Provider Usage 记录 `MODEL_TOKEN_USAGE`:Router、System Chat、Knowledge Answer、Diagnosis Agent、Evidence Repair 和 SemanticGuard。事件按组件和组件轮次记录 input/output/total Token;Usage 不可用时只记录 unavailable,不估算消耗。
+
+`RUN_FINISHED` 同时记录 RunBudget 总账、模型调用明细合计、不可用 Usage 数量和 `tokens_reconciled`。Diagnosis Agent 的每轮 total Token 还会回填 `AgentStep.token_count`;其他模型组件不创建伪 AgentStep。
+
+### 6.2 Tool 拒绝
+
+`TOOL_INVOCATION` 只表示进入业务 ToolBoundary 的实际执行。协议错误、重复 scope、信息饱和或观察合同拒绝使用独立 `TOOL_REQUEST_REJECTED`,只记录 Tool Call ID、Tool name 和稳定错误码。预算 Tool 计数、实际执行次数和 Harness 拒绝次数是三个不同观察维度。
+
## 7. Audit 安全
AgentStep 不保存 Prompt、消息正文、模型正文、Tool arguments 或 Thought。ToolInvocation 不保存完整 request、SQL/日志 query、raw response 或 Agent projection。Provider reasoning 仅写入独立 `agent_reasoning_audit`,不进入普通 Trace 或发布结果;无 Provider 内容时必须记录 unavailable,不能伪造。应用日志不得打印这些字段。
diff --git a/mvp/issues/README.md b/mvp/issues/README.md
index e512a94..05ff083 100644
--- a/mvp/issues/README.md
+++ b/mvp/issues/README.md
@@ -1,6 +1,6 @@
# MVP Issues 索引
-**更新日期**:2026-07-23
+**更新日期**:2026-07-26
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
## 目录约定
@@ -17,6 +17,7 @@
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|---|---|---|---|---|
| ISS-015 | 诊断运行质量与 Reasoning 审计收敛 | 高 | 待实施 | [active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md](active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md) |
+| ISS-016 | 诊断 Agent 缺少基于信息增益的停止契约 | 高 | 已实施,待归档 | [active/ISS-016-diagnosis-information-gain-stop-contract.md](active/ISS-016-diagnosis-information-gain-stop-contract.md) |
## 设计笔记
diff --git a/mvp/issues/active/ISS-016-diagnosis-information-gain-stop-contract.md b/mvp/issues/active/ISS-016-diagnosis-information-gain-stop-contract.md
new file mode 100644
index 0000000..6b4344d
--- /dev/null
+++ b/mvp/issues/active/ISS-016-diagnosis-information-gain-stop-contract.md
@@ -0,0 +1,358 @@
+# ISS-016 诊断 Agent 缺少基于信息增益的停止契约
+
+**状态**:已实施;审计发现协议拒绝空转,暂不归档
+**严重程度**:高
+**发现时间**:2026-07-26
+**来源**:知识库无直接答案场景的真实 E2E 与停止策略复盘
+**关联**:ISS-015、ISS-012、ISS-002、ISS-009
+**架构设计**:[Diagnosis Agent 信息增益与停止控制架构](../../architecture/diagnosis-information-gain-stop-architecture.md)
+
+---
+
+## 1. 背景
+
+当前单体 Diagnosis ReAct Agent 已具备 Tool 调用、Run 预算、EvidenceGuard、SemanticGuard、Trace 和安全发布边界,但停止条件仍主要依赖模型自行结束或 Harness 资源预算耗尽。
+
+这使系统能够限制一次 Run 最多消耗多少资源,却不能稳定判断“继续查询是否还可能增加有效诊断信息”。当知识库只有通用参考资料、日志查询持续为空或用户缺少必要查询条件时,Agent 可能不断改写关键词和扩大尝试,最终由预算被动终止。
+
+该问题不能简化为“Prompt 没写好”或“预算太大”:Prompt、模型终态契约、Tool 客观结果、模型语义评价和 Harness 强制停止之间缺少完整闭环。
+
+## 2. 真实复现
+
+用户 Query:
+
+```text
+诊断切换企业失败的问题
+```
+
+真实 E2E:
+
+- `sessionId=iss015-enterprise-switch-e2e-20260726000222`
+- `runId=054afbb1-0bb3-48d2-8cb8-ddb31d20b2b6`
+- SSE:`metadata -> ROUTING -> DIAGNOSIS_RUNNING -> failure(INTERNAL_FAILURE) -> done(FAILED)`
+- Run 耗时约 48 秒。
+- Agent 已完成多轮模型和 Tool 调用。
+- `query_logs` 多次返回 `NO_EVIDENCE`。
+- `lookup_knowledge` 的检索层结果主要为 `REFERENCE`,但 Harness 投影仍表现为 `EVIDENCE_FOUND`。
+- 最后一次 Tool 调用触发 `BUDGET_EXHAUSTED`,Run 未生成最终 Draft,也未进入安全校验阶段。
+- 数据库和 Trace 中存在执行过程,但用户只能看到“当前暂时无法处理该请求,请稍后重试”。
+
+这不是“没有执行过程”,而是 Agent 未能在无新增信息时主动结束,预算终止又被发布链路映射成了技术失败。
+
+### 2.1 实施后验证
+
+同一 Query 的最终 named SSE E2E:
+
+- `sessionId=iss016-final-20260726-a`
+- `runId=3ab22ed7-d0ed-45d8-b928-dce5790c0542`
+- SSE 发布 `SAFE_FALLBACK`,`type=MISSING_REQUIRED_CONTEXT`,最终 `done.outcome=FALLBACK`。
+- 数据库记录 `status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=0`、`total_token_count=2890`,answer 非空。
+- Trace 为 `RUN_STARTED -> ROUTING_ATTEMPT -> ROUTING_DECISION -> AGENT_MODEL_STEP -> EVIDENCE_GUARD_INITIAL -> RELEASE_DECISION/FALLBACK -> RUN_FINISHED/FALLBACK`。
+
+本次零 Tool 是合法行为:原 Query 缺少企业、时间范围和错误信息,模型直接报告缺失上下文。另有 focused tests 固定“非法 Draft + 已验真 ProgressSnapshot -> INSUFFICIENT_EVIDENCE”和“非法 Draft + 无安全进展 -> FAILED”边界,避免模型随机返回非 JSON 时再次丢失已完成过程。
+
+### 2.2 Token 与拒绝审计补充验证
+
+2026-07-27 增加按模型组件/轮次的 Token 明细、AgentStep Token 回填、Run 对账摘要和 Tool 请求拒绝 Trace。审计事件不保存 Prompt、模型正文、Tool 参数或原始响应。
+
+缺少上下文 E2E:
+
+- `sessionId=audit-e2e-20260727-001`
+- `runId=87f38bea-108b-482a-8b72-d0890cc825f2`
+- Router Token `248`,Diagnosis Agent Token `2635`,Run 总 Token `2883`。
+- `step_count=1`,对应 `AgentStep.token_count=2635`;Tool 预算计数和实际执行均为 `0`。
+- `RUN_FINISHED.tokens_reconciled=true`,Trace 序列 `1..9` 连续。
+
+有界 Tool E2E:
+
+- `sessionId=audit-e2e-20260727-003`
+- `runId=3f8f4a0c-3b96-4942-be1e-61cc1431b337`
+- 5 次 Tool 实际执行:`lookup_knowledge=4`、`query_logs=1`。
+- 9 次 Tool 请求被 `INVALID_PROGRESS_PROTOCOL` 拒绝,均有独立 `TOOL_REQUEST_REJECTED`,没有伪装成 canonical invocation。
+- 13 个 Diagnosis Agent 轮次加 1 个 Router 调用,总 Token `68469`;各轮 AgentStep Token 之和加 Router Token与 Run 总量一致。
+- `RUN_FINISHED.tokens_reconciled=true`,Trace 序列 `1..52` 连续。
+
+该 E2E 证明审计可观察、可对账、可按组件和轮次重放,也暴露了新的停止缺口:连续协议拒绝当前不产生 `NO_GAIN`,模型可在 Harness 拒绝后继续消耗模型轮次。这个 Run 的 Token 消耗不合理;后续需要单独确认“连续 `INVALID_PROGRESS_PROTOCOL` 是否进入确定性停止”的行为设计,不能仅靠提高或降低预算掩盖。
+
+## 3. 核心问题
+
+### 3.1 成功定义只有“完成诊断”
+
+如果模型的唯一合法输出是完整 `DiagnosisDraft`,那么“证据不足”和“需要用户补充信息”不是一等终态。即使模型已经知道当前信息无法支持根因,它仍可能认为任务尚未完成,并继续调用 Tool。
+
+### 3.2 Tool 客观结果与诊断价值混淆
+
+Tool 可以可靠说明查询是否成功、范围、数量、检索分数、是否为空和是否截断,但不能判断结果是否支持当前诊断假设。
+
+例如:
+
+- 检索到三篇相似文档,只能说明存在候选内容,即当前兼容状态 `EVIDENCE_FOUND`;
+- 文档是直接证据、背景材料还是无关内容,需要模型结合当前假设判断;
+- 当前范围没有日志,只能说明 scoped `NO_EVIDENCE`,不能说明故障不存在。
+
+把检索层 `REFERENCE` 直接表达为诊断层 `EVIDENCE_FOUND`,会向模型发送“仍在取得进展”的错误信号。
+
+### 3.3 模型没有显式的信息价值判断义务
+
+当前系统主要通过“模型是否继续调用 Tool”推测它是否认为查询有价值。对于无法由代码确定质量的成功非空结果,模型没有被要求给出最小、机器可读的信息增益判断:
+
+```text
+information_gain = GAINED | NO_GAIN
+```
+
+因此 Harness 无法区分“合理继续收集证据”和“换关键词重复尝试”。
+
+### 3.4 Harness 只有资源预算,没有无进展契约
+
+模型调用次数、Tool 调用次数、Token、字节和超时属于资源边界。它们是最后的安全保护,不应承担正常停止策略。
+
+Harness 当前缺少以下可观察状态:
+
+- 连续 `NO_GAIN` 次数;
+- 重复或等价查询;
+- 当前收集状态 `COLLECTING / SATURATED`;
+- 模型对需要语义评价的结果给出的 `GAINED / NO_GAIN`。
+
+## 4. 顶层设计原则
+
+### 4.1 不要求模型必须找到根因
+
+诊断任务的首要目标是避免无依据结论,而不是每次都输出根因。模型只输出 `DiagnosisDraft`,不产生额外的诊断生命周期状态:
+
+```text
+有受证据支持的 conclusion
+ -> ReleaseOutcome.SUCCESS
+
+conclusion=null,且形成了诚实、有界的结果
+ -> ReleaseOutcome.FALLBACK
+
+不可恢复的 Tool、模型或基础设施故障
+ -> ReleaseOutcome.FAILED
+```
+
+证据不足和缺少查询条件是 `FALLBACK` 的不同原因,不是与 `SUCCESS / FALLBACK / FAILED / CANCELLED` 平行的第二套状态机。现有 `SafeFallback.type` 负责表达 `INSUFFICIENT_EVIDENCE / MISSING_REQUIRED_CONTEXT` 等发布原因。
+
+### 4.2 Tool 提供事实,模型判断价值
+
+| 组件 | 职责 |
+|---|---|
+| Tool / Adapter | 返回客观执行状态、查询范围、数量、来源和有界结果 |
+| Diagnosis Agent | 判断结果对当前假设的语义价值 |
+| Harness | 记录进展、识别重复和不一致、执行确定性停止 |
+| EvidenceGuard / SemanticGuard | 校验最终引用真实性和结论支持度 |
+| Release | 把正常停止与技术失败发布为不同用户结果 |
+
+不应让 Tool 声称业务证据价值,也不应让 Harness 通过规则替代模型完成诊断推理。
+
+### 4.3 模型主动停止,Harness 确定性兜底
+
+模型应优先根据 Prompt 和显式进展状态主动选择结束;模型未能收敛时,Harness 必须根据可验证规则停止后续 Tool 调用。
+
+只加强 Prompt 不足以形成系统保证;只增加硬预算则会继续产生高成本、低信息量的失败 Run。
+
+## 5. 目标交互契约
+
+### 5.1 已确认的 Tool 侧决策
+
+1. Tool 由服务端通过模型原生 Tool Calling 通道动态注册;Prompt 不写死 Tool 名称或 Schema,也不重复插入 Tool Schema 占位符。
+2. Tool 和 Projector 只产生客观结果,不判断业务根因和语义信息增益。
+3. `NO_EVIDENCE` 由 Projector 根据结果集合为空确定,不是 LLM 推导。
+4. `REFERENCE` 由 RAG 后处理规则根据归一化相似度及查询提示命中情况确定,不是 LLM 推导。
+5. `RagResultProjector` 必须兼容读取上游 `relevanceLevel` 或 `relevance_level`,并统一保留为 `relevance_level`。
+6. `EVIDENCE_FOUND` 暂时保留,但只表示存在候选内容,不表示存在能够支持诊断结论的证据。
+7. 第一版不引入 `new_count`,不建立跨 Tool 通用内容指纹。
+8. 不引入显式 `next_action`。模型发起 Tool Call 表示继续,输出 Draft 表示结束。
+
+### 5.2 Tool 结果双视图
+
+Tool 原始结果经过标准化后产生内部 Canonical Tool Result,再分别提供:
+
+| 视图 | 消费者 | 内容 |
+|---|---|---|
+| Harness Control View | Harness / Canonical Store | 执行状态、数量、规范化 scope、相关度、重复指纹、预算与饱和控制信息 |
+| Model Observation | Diagnosis Agent | `tool_call_id`、实际 scope、有界 evidence、`evidence_status`、粗粒度 `relevance_level`、`truncated` |
+
+原始相似度、检索轨迹、重复指纹、低收益计数、阈值、预算和原始 Tool Response 不进入模型上下文。只有 Harness 必须改变模型行为时,才注入有界 `STOP_REQUIRED` 指令和已检查范围摘要。
+
+详细 Tool 架构图和调用流程图见架构文档的“Tool 结果与上下文边界”章节。
+
+### 5.3 信息增益与停止
+
+信息增益只保留:
+
+```text
+information_gain = GAINED | NO_GAIN
+```
+
+赋值规则:
+
+```text
+FAILED
+ -> 技术故障流程,不产生 information_gain
+
+evidence_status=NO_EVIDENCE
+或重复的规范化 tool + scope
+ -> Harness 直接赋 NO_GAIN
+
+其他成功非空结果,包括 relevance_level=REFERENCE
+ -> Diagnosis Agent 判断 GAINED / NO_GAIN
+```
+
+Harness 使用连续 `NO_GAIN` 维护 `COLLECTING / SATURATED`。连续次数达到配置项 `stop-after-consecutive-no-gain` 后进入 `SATURATED`;默认值为 `2`,只对新 Run 生效,且不进入模型上下文。
+
+硬调用上限不等于信息饱和。Harness 分别记录:
+
+```text
+连续 NO_GAIN 达到阈值 -> stop_reason=INFORMATION_SATURATED
+达到 Tool 或 Run 硬预算 -> stop_reason=BUDGET_LIMIT_REACHED
+```
+
+`SATURATED` 只表示信息饱和;预算限制属于独立的资源保护终止原因。
+
+模型不显式声明行动状态。Harness 从模型实际发起 Tool Call 或输出 Draft 推导其行为。
+
+### 5.4 Tool Call Envelope 协议
+
+模型通过下一次 Tool Call 的通用 Envelope 回传上一轮语义评价:
+
+```json
+{
+ "previous_observation": {
+ "tool_call_id": "call-123",
+ "information_gain": "NO_GAIN"
+ },
+ "input": {
+ "query": "下一次查询参数"
+ }
+}
+```
+
+协议约束:
+
+- `previous_observation` 只在上一轮存在待模型评价的 Tool Observation 时必填;首次调用以及上一轮已由 Harness 客观赋值时省略。
+- Harness 在执行新 Tool 前校验 `tool_call_id`,应用 `information_gain` 并更新 `COLLECTING / SATURATED`。
+- 如果应用评价后进入 `SATURATED`,Harness 拒绝本次 Tool 调用并返回 `STOP_REQUIRED`。
+- Harness 消费并剥离 `previous_observation`;业务 Tool 只接收 `input` 中原有的业务参数。
+- 模型选择直接输出 DiagnosisDraft 时,不存在需要放行的下一次 Tool Call,因此无需额外回传最后一轮评价;Harness 将其记录为模型主动结束。
+- 上一个待评价结果未完成评价时,Harness 不接受新的 Tool 调用。
+
+这是模型侧 Tool 调用协议变更,但不改变业务 Tool 的入参协议,也不增加独立 Progress Judge 调用,不把 Harness 内部状态暴露给模型,不把 `information_gain` 混入最终用户可见的 DiagnosisDraft。
+
+### 5.5 已确认的 Prompt 原则
+
+Prompt 使用中文,并保持最小职责,不解释 Projector、计数器、阈值、预算或 Harness 状态机。Tool 名称和 Schema 不写死在正文中,也不拼接进 Prompt,由服务端通过模型原生 Tool Calling 通道动态注册。
+
+Prompt 必须明确:
+
+- 模型不必须得出诊断结论或根因,但必须完成一个诚实、有界的 `DiagnosisDraft`;
+- `conclusion=null` 的证据不足结果是合法完成,不是失败;
+- Tool 调用次数可以为零;缺少有效查询所需的企业、时间、服务或错误信息时,应直接在 `limitations.missing_info` 中列出缺口;
+- 不得为了表现“已经排查”而执行没有明确范围、预期不会产生新诊断信息的 Tool 调用;
+- Tool 返回内容事实正确、表述完整或结果非空,不代表它对当前诊断有信息增益;
+- 只有新增事实确认、排除或缩小当前诊断假设时,才属于 `GAINED`;
+- 通用知识、背景说明、重复内容或不能改变当前判断的内容属于 `NO_GAIN`;
+- `NO_GAIN` 后不得通过改写相似关键词或重复相同 scope 继续尝试;
+- 如果仍存在范围明确、并可能产生新诊断信息的不同查询,可以继续;否则应主动停止并完成证据不足结果;
+- 收到服务端 `STOP_REQUIRED` 后必须停止调用 Tool。
+
+Prompt 不要求模型输出 `next_action`。模型发起 Tool Call 或输出 Draft 即表达其实际选择。
+
+## 6. 实施阶段(已完成)
+
+### 阶段 1:终态与 Prompt 契约
+
+- 保持 `conclusion=null` 作为合法的证据不足 Draft,并明确“有依据地停止”属于成功完成。
+- 按 5.5 的最小原则重写中文 Prompt,Tool 定义只通过模型原生 Tool Calling 通道动态注册。
+- 允许零次 Tool 调用;查询条件不足时使用现有 `limitations.missing_info`,不新增 `required_context`。
+- 明确正确但不能推进当前诊断的 Tool 内容属于 `NO_GAIN`,不得触发等价重试。
+- 保持 Tool 预算作为最终安全保护。
+
+### 阶段 2:Tool 状态与模型评价解耦
+
+- 标准化 Canonical Tool Result,并拆分 Harness Control View 与 Model Observation。
+- `RagResultProjector` 兼容并保留 `relevance_level`。
+- 不把检索命中、候选文档或 `REFERENCE` 自动等同于诊断证据。
+- 通过白名单控制进入模型上下文的 Tool 字段。
+
+### 阶段 3:Run 级最小进展状态
+
+- 在 RunContext 中维护已检查 `tool + scope`、连续 `NO_GAIN` 和 `COLLECTING / SATURATED`。
+- 只对 `tool_name + normalized_scope` 做确定性参数去重;第一版不做自然语言语义去重。
+- 使用配置项 `stop-after-consecutive-no-gain` 控制连续低收益阈值,默认值为 `2`。
+- 实现通用 Tool Call Envelope,由下一次 Tool Call 回传上一轮 `information_gain`,并在进入业务 Tool 前剥离控制字段。
+- 不把 Prompt、内部 thought 或原始 Tool 载荷混入普通 Trace 和用户结果。
+
+### 阶段 4:Harness 确定性停止
+
+- 拒绝相同 `tool_name + normalized_scope` 的重复查询。
+- 达到无进展条件时阻止新的 Tool 调用。
+- 分别产生 `INFORMATION_SATURATED / BUDGET_LIMIT_REACHED`,不把硬预算终止伪装成信息饱和。
+
+### 阶段 5:安全发布与 E2E
+
+- 每次 Tool 完成后只保存 Canonical Tool Result;Tool Loop 结束时一次性投影有界 `ProgressSnapshot`。
+- `DiagnosisReleaseUseCase` 同时接受正常 Draft 和 `stop_reason + ProgressSnapshot`,统一产生 `SUCCESS / FALLBACK`;`ChatApplicationUseCase` 不再单独决定预算 Fallback 的业务内容。
+- 复用现有 `SafeFallback.observed_facts / limitations / next_steps` 和前端渲染协议,并为 `FallbackType` 增加 `INSUFFICIENT_EVIDENCE / MISSING_REQUIRED_CONTEXT`。
+- `conclusion=null` 不进入 EvidenceRepair;有 Tool 引用时只验证引用真实性,并从 `ProgressSnapshot` 生成 Fallback。只有存在结论时才执行完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard 链路。
+- 不依赖预算耗尽后的额外模型调用来生成 Fallback。
+- 前端明确区分证据不足、需要补充信息和技术故障。
+
+## 7. 非目标
+
+- 不通过简单增加 Tool/Token 预算解决问题。
+- 不把“降低最大 Tool 调用次数”当作信息增益策略。
+- 不引入 `new_count`、多级质量评分或显式 `next_action`。
+- 不增加独立 Progress Judge 模型调用。
+- 不依赖 Prompt 作为唯一控制手段。
+- 不让 Tool 或规则代码判断业务根因。
+- 不重新引入已由 ISS-014 删除的 Planner/Executor/Verifier/Composer 旧业务 Graph。
+- 不在本 Issue 中解决 Reasoning 原文审计治理;该内容仍由 ISS-015 跟踪。
+
+## 8. 验收标准
+
+- [x] 不引入 `CONFIRMED / INSUFFICIENT_EVIDENCE / NEED_MORE_INFO` 第二套诊断生命周期状态;最终状态只使用 `SUCCESS / FALLBACK / FAILED / CANCELLED`。
+- [x] Tool 客观命中状态与模型诊断价值评价在契约和 Trace 中明确分离。
+- [x] Tool 结果拆分为 Harness Control View 和白名单 Model Observation,Harness 内部计数、阈值和预算不进入模型上下文。
+- [x] `RagResultProjector` 保留 `relevance_level`,且模型与 Harness 都能获得各自所需的有界视图。
+- [x] `NO_EVIDENCE` 和重复的规范化 `tool + scope` 被 Harness 确定性标记为 `NO_GAIN`;RAG `REFERENCE` 由模型判断 `GAINED / NO_GAIN`。
+- [x] 模型继续调用 Tool 前,对上一轮待评价结果返回 `GAINED / NO_GAIN`,不返回 `next_action`。
+- [x] 下一次 Tool Call Envelope 能回传上一轮语义评价;Harness 应用评价后才决定是否放行,并在调用业务 Tool 前剥离控制字段。
+- [x] Prompt 明确不要求模型必须得出根因,`conclusion=null` 是合法完成结果。
+- [x] Prompt 允许零次 Tool 调用;缺少必要查询条件时使用 `limitations.missing_info`,不强制执行无明确目标的查询。
+- [x] Prompt 明确正确但不能推进诊断的 Tool 内容属于 `NO_GAIN`,并禁止相似关键词或相同 scope 的等价重试。
+- [x] `REFERENCE`、候选文档或非空结果不会自动成为支持根因的证据。
+- [x] 相同 `tool_name + normalized_scope` 的重复查询能够被 Harness 识别并阻止;第一版不承诺自然语言语义去重。
+- [x] `stop-after-consecutive-no-gain` 可配置且默认值为 `2`;`GAINED` 清零连续计数。
+- [x] `INFORMATION_SATURATED` 与 `BUDGET_LIMIT_REACHED` 在 Trace 和 Release 输入中保持区分。
+- [x] 连续无新增信息时,Run 在资源预算耗尽前收敛为正常业务结果。
+- [x] 模型主动停止和 Harness 强制停止都能输出已检查来源、查询范围、客观结果、信息缺口和下一步。
+- [x] 无证据不被解释为故障不存在,通用参考资料不被解释为当前故障事实。
+- [x] 真正取得新证据的多轮诊断不会被无进展策略过早终止。
+- [x] 原始复现 Query 不再以 `INTERNAL_FAILURE` 结束,也不再通过大量改写查询运行到 `BUDGET_EXHAUSTED`。
+- [x] Tool Loop 结束时由 Canonical Tool Result 一次性投影 `ProgressSnapshot`,并映射到现有 `SafeFallback.observed_facts`,无需新增前端协议。
+- [x] `DiagnosisReleaseUseCase` 统一处理正常 Draft、信息饱和和预算终止;只有不可形成安全业务结果的技术故障发布为 `FAILED`。
+- [x] `conclusion=null` 不触发 EvidenceRepair;有结论时才执行完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard 链路。
+- [x] 最终 Draft 合同失败时丢弃非法内容;仅在当前 Run 有已验真 ProgressSnapshot 时发布 `INSUFFICIENT_EVIDENCE`,否则保持 `FAILED`。
+- [x] 原始复现 Query 在缺少企业、时间和错误信息时返回 `FALLBACK + MISSING_REQUIRED_CONTEXT`,或在有限排查后返回 `FALLBACK + INSUFFICIENT_EVIDENCE`,并展示已经完成的检查。
+- [x] exact `sessionId + runId` Trace 能解释每轮继续或停止的原因,但不泄露 Prompt、内部 thought、原始 Tool 载荷或凭据。
+- [x] 每个 Provider Usage 可按模型组件和组件轮次审计,Diagnosis Agent Usage 回填 AgentStep,Run 结束显式记录 Token 是否对账。
+- [x] Harness 拒绝的 Tool 请求与实际 `TOOL_INVOCATION` 分开记录,拒绝 Trace 不包含 Tool 参数或原始响应。
+- [ ] 连续 `INVALID_PROGRESS_PROTOCOL` 拒绝应在硬模型/Token 预算前确定性停止;当前真实 E2E 仍出现 9 次拒绝和 13 个 Agent 轮次。
+
+## 9. 已确认的首版边界
+
+1. 连续 `NO_GAIN` 阈值由 `stop-after-consecutive-no-gain` 配置,默认值为 `2`,后续通过固定 E2E 和评测集校准。
+2. 重复检测只比较 `tool_name + normalized_scope`,不做自然语言语义去重。
+3. 模型可以在首次 Tool 调用前以 `conclusion=null + limitations.missing_info` 合法结束。
+4. Harness 产生内部 `stop_reason`;Release 产生 `SafeFallback.type` 和最终 `release_outcome`。
+5. Tool Schema 只通过模型原生 Tool Calling 通道注册,不拼接进 Prompt。
+
+## 10. 相关文件
+
+- `mvp/issues/active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md`
+- `mvp/issues/archived/ISS-002-executor-unconstrained-lookup.md`
+- `mvp/issues/archived/ISS-012-executor-token-budget-and-context-growth.md`
+- `mvp/issues/archived/ISS-009-negative-observation-no-evidence-reference.md`
+- `mvp/architecture/agent-orchestration.md`
+- `mvp/architecture/harness-quality-gates.md`
+- `mvp/architecture/diagnosis-information-gain-stop-architecture.md`
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/.committed b/openspec/changes/diagnosis-information-gain-stop-contract/.committed
new file mode 100644
index 0000000..b2681a5
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/.committed
@@ -0,0 +1,5 @@
+Committed OpenSpec
+
+Validated: 2026-07-26
+Scale: complex
+Interface impact: L3
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/.openspec.yaml b/openspec/changes/diagnosis-information-gain-stop-contract/.openspec.yaml
new file mode 100644
index 0000000..2bc06e0
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/.openspec.yaml
@@ -0,0 +1,2 @@
+schema: spec-driven
+created: 2026-07-26
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/design.md b/openspec/changes/diagnosis-information-gain-stop-contract/design.md
new file mode 100644
index 0000000..00476e0
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/design.md
@@ -0,0 +1,172 @@
+## Context
+
+当前 Diagnosis ReAct loop 只有模型、Tool、Token、字节和时间预算,没有“查询是否仍在产生信息”的状态。`HarnessToolInterceptor` 将完整 canonical `agent_result` 直接放入模型上下文;三个 Agent-facing Tool 使用裸业务 request 生成 Schema;`DiagnosisAgentUseCase` 只返回非空 `DiagnosisDraft`;`DiagnosisReleaseUseCase` 要求 Draft 非空并对所有 Draft 运行 EvidenceGuard/Repair/SemanticGuard。预算耗尽的临时 Fallback 位于 `ChatApplicationUseCase`,导致业务 Release 决策分散。
+
+本变更是 L3 模型协作接口演进。公开 HTTP/SSE、业务 Tool backend、数据库表和 `SafeFallback` JSON 结构保持兼容。当前环境没有 semantic retrieval/LSP,调用链通过源码、测试和 `rg` 引用核查完成。
+
+## Goals / Non-Goals
+
+**Goals:**
+
+- 在硬预算耗尽前确定性停止连续无增益 Tool 调用。
+- 让模型在看到成功非空结果后,以 `GAINED / NO_GAIN` 表达该结果是否推进当前诊断。
+- 允许模型零次调用 Tool 或以 `conclusion=null` 合法结束。
+- 信息饱和、预算终止和主动无结论均由 Diagnosis Release 发布为有过程的安全 Fallback。
+- 有结论 Draft 继续经过完整 EvidenceGuard、单次 EvidenceRepair 和 SemanticGuard。
+- 控制数据、raw payload、内部 thought 和预算不进入模型上下文或用户结果。
+
+**Non-Goals:**
+
+- 不引入 Judge、`UNKNOWN`、分数、`new_count`、`next_action` 或第二套诊断生命周期。
+- 不做自然语言语义去重或跨 Run 进展继承。
+- 不改变 Tool backend 参数、公开 SSE 事件或前端字段协议。
+- 不把候选文档、`REFERENCE` 或非空结果自动认定为诊断证据。
+
+## Decisions
+
+### 1. RunContext 持有最小线程安全 ProgressTracker
+
+新增 `DiagnosisProgressTracker` handle,并在 `DiagnosisHarnessCore.startRun` 时用固定阈值创建。tracker 维护:连续 `NO_GAIN`、`COLLECTING / SATURATED`、可选 `stop_reason`、最后一个待模型评价的成功 Tool Call、已成功检查的规范化 scope,以及已完成 canonical key/Tool Call ID 的有序索引。
+
+tracker 不保存 raw response、完整 Agent result 或用户可见摘要。Canonical Store 仍是 Tool 真相源;结束投影器按 tracker 的 key 索引逐条读取 canonical 记录。
+
+替代方案是给 `CanonicalInvocationStore` 增加按 Run 枚举。拒绝,因为停止状态是短生命周期控制数据,且该方案会扩大 Redis 接口及所有 fake store,实现另一种事实索引。
+
+### 2. 使用三个强类型 Envelope,共享上一轮评价类型
+
+Agent-facing Schema 使用三个具体输入类型:`RagToolCall`、`QueryLogsToolCall`、`MysqlToolCall`。每个类型包含可选 `previous_observation` 和必填 `input`;`input` 继续使用现有业务 request,`previous_observation` 使用共享的 `PreviousObservation(tool_call_id, information_gain)`。
+
+不使用泛型 `ToolCallEnvelope` 直接生成 Schema,因为运行时类型擦除可能使嵌套 `input` 丢失具体字段;不使用 `JsonNode`,因为它无法给模型提供强 Schema。interceptor 严格解析对应 Envelope,先消费控制字段,再把 `input` 序列化为原业务 JSON 交给 adapter。
+
+首个 Tool Call 以及上一结果已由 Harness 确定性评价时允许省略 `previous_observation`。存在待评价调用时,新 Tool Call 必须携带完全匹配的上一 Tool Call ID 和二值评价;缺失、错序或跨 Run 引用返回有界协议错误且不执行 Tool。
+
+### 3. 先应用上一轮评价,再决定本轮是否放行
+
+interceptor 的固定顺序为:
+
+1. 严格解析 Envelope 并校验 `previous_observation`。
+2. 将上一轮 `GAINED` 清零计数,或将 `NO_GAIN` 加一。
+3. 若达到阈值,拒绝当前业务 Tool,返回一次 `STOP_REQUIRED/INFORMATION_SATURATED`。
+4. 规范化本次业务 request 的 scope;若与同 Tool 的成功历史 scope 重复,则不执行 Tool、记一次确定性 `NO_GAIN`,并返回有界重复提示或 STOP_REQUIRED。
+5. 其余请求进入现有 adapter/ToolBoundary。
+6. 成功后记录 canonical key;`NO_EVIDENCE` 由 Harness 立即记为 `NO_GAIN`,`EVIDENCE_FOUND` 标记为待模型评价。
+
+Tool `ERROR` 不产生 information gain,也不把失败 scope 写入成功去重集合;它沿用技术故障和现有预算/重试边界。
+
+替代方案是让 Tool 自行报告质量。拒绝,因为 Tool 只知道客观返回,无法判断对当前诊断假设的价值。
+
+### 4. 只做确定性 scope 规范化
+
+每个 Tool 提供一个无副作用 scope projector:RAG 使用规范化 query;日志使用 topic、query 和实际 lookback;MySQL 使用 logical datasource、规范化 SQL 文本和参数。仅规范化空白、大小写明确不敏感的枚举/标识、确定性默认值和结构化集合,不判断自然语言改写是否语义等价。
+
+重复只比较 `tool_name + normalized_scope`,且只基于已成功执行的历史 scope。被拒绝的重复调用不进入 ToolBoundary、canonical store 或 Tool 调用预算,但会记录安全 Trace 并推进无增益计数。
+
+### 5. Canonical 结果产生控制视图和模型白名单视图
+
+ToolBoundary 继续保存完整 request、raw response 和 bounded `agent_result`。Tool 完成后:
+
+- Harness Control View 读取 execution/evidence status、returned count、normalized scope、RAG relevance、truncated 和 canonical identity。
+- Model Observation 只包含 Tool Call ID、实际 scope、有界 evidence/rows/events、`evidence_status`、可选 `relevance_level` 和 `truncated`。
+
+`HarnessToolInterceptor` 不再直接返回完整 `agent_result`,而是通过按 Tool 类型的 `AgentObservationProjector` 白名单序列化。RAG projector 兼容读取 `relevanceLevel` 和 `relevance_level`,在 canonical RAG result 中统一为 `relevance_level`。`REFERENCE` 仍交给模型判断信息增益。
+
+### 6. 饱和后只有一次正常收尾机会
+
+当 Tool 结果使 tracker 直接饱和时,当前 Model Observation 同时携带 `stop_required=true` 和 `reason=INFORMATION_SATURATED`。当模型在下一次 Tool Call 中回传 `NO_GAIN` 后达到阈值时,该本轮 Tool 被拒绝并返回同样的 STOP_REQUIRED observation。
+
+tracker 记录控制指令已交付。下一轮模型仍发起 Tool Call时,interceptor 抛出可识别的 `DiagnosisCollectionStoppedException`;`DiagnosisAgentUseCase` 不把它包装为内部故障,而是返回“无 Draft + stop reason”的执行结果。这样模型获得一次生成合法 Draft 的机会,同时无法靠重复 Tool Call继续空转。
+
+替代方案是无限返回 STOP_REQUIRED。拒绝,因为模型无视指令时仍会消耗模型预算并重现原问题。
+
+### 7. Agent 执行返回 Draft 与停止投影,而不是只返回 Draft
+
+`DiagnosisAgentUseCase` 返回内部 `DiagnosisAgentExecution`:可选 Draft、`ProgressSnapshot` 和可选 `DiagnosisStopReason`。正常 Draft、强制饱和停止和预算异常都在 Diagnosis executor 边界形成 Release 输入。
+
+`DiagnosisProgressProjector` 在 Tool loop 结束时读取 tracker 索引和 canonical store,一次性生成有界 snapshot;缺失、过期、ERROR 或不属于当前 Run 的记录不会成为 verified fact,并产生稳定 limitation/trace。snapshot 不进入模型上下文。
+
+### 8. DiagnosisReleaseUseCase 统一业务发布决策
+
+Release 接收 query、可选 Draft、ProgressSnapshot 和可选 stop reason:
+
+- `conclusion != null`:执行现有完整 EvidenceGuard、一次 Repair、重验和 SemanticGuard。
+- `conclusion == null`:不执行 Repair/SemanticGuard;若 Draft 有 Tool 引用,仅验证当前 Run canonical 引用真实性和负向语义,不要求正常结论结构。
+- 无 Draft且 `INFORMATION_SATURATED` 或 `BUDGET_LIMIT_REACHED`:从 snapshot 确定性生成 Fallback,不调用额外模型。
+- 零 Tool 且 Draft 的 `limitations.missing_info` 非空:发布 `MISSING_REQUIRED_CONTEXT`。
+- 有有限排查但无可支持结论:发布 `INSUFFICIENT_EVIDENCE`。
+- 真正 Tool/模型/基础设施故障且无法形成安全过程:继续 `FAILED`。
+
+最终模型文本为空或不能严格解析为 `DiagnosisDraft` 时,非法内容本身始终被丢弃,不做 Markdown/自然语言 JSON 抽取,也不调用额外模型修复。Agent 输出异常携带一次有界 `ProgressSnapshot`;Executor 仅在 snapshot 含当前 Run 已验真的 observed facts 时交给 Release 生成 `INSUFFICIENT_EVIDENCE`,否则保持 `FAILED`。Trace 只记录固定失败类别、输出字节数和是否存在可发布进展,不记录模型原文、字段值或解析异常文本。
+
+`ChatApplicationUseCase` 删除业务内容级 `recoverBudgetExhaustion`。Diagnosis executor 捕获可识别预算终止并调用 Release;Application 只允许该已处理 Diagnosis Fallback 跳过 `core.checkActive/completeSuccess`,持久化 `release_outcome=FALLBACK`。Run 内部仍保留 `BUDGET_EXHAUSTED`,数据库公开运行状态继续按既有规则记录为成功发布的 Fallback。
+
+### 9. Prompt 只约束模型职责,不复制 Tool Schema
+
+中文 Prompt 明确:无需强行得出根因;`conclusion=null` 是合法完成;缺少企业、时间、服务或错误信息时允许零 Tool 并填写 `limitations.missing_info`;只有新增可验证事实确认、排除或缩小假设才是 `GAINED`;正确但无用、通用或重复内容是 `NO_GAIN`;没有明确不同且可能产生新信息的 scope 时停止;收到 STOP_REQUIRED 后不得继续调用 Tool。
+
+Prompt 不写 Tool 名、Schema、阈值、计数器、Projector、预算或 `next_action`。
+
+### 10. Trace 记录决策,不泄露推理
+
+Trace 增加有界事件或字段,记录 Tool Call ID、Tool name、scope 摘要、information gain 的生产者(Harness/Model)、连续计数变化、collection state 和 stop reason。不得记录 Prompt、模型 thought、raw Tool response、完整 SQL 参数、预算余量或模型评价理由。
+
+### 11. Run 总账与模型调用明细使用同一份 Provider Usage
+
+`RunContext` 增加最小线程安全模型调用账本,只维护组件轮次、已审计调用数、Usage 不可用调用数和 Token 合计。每次实际模型调用在预算放行后取得组件轮次;`HarnessModelInterceptor` 负责 Diagnosis Agent,`GuardModelCall` 负责 Router、System Chat、Knowledge Answer、Evidence Repair 和 Semantic Guard。两条入口都从 Spring AI `Usage` 读取同一组 input/output Token,先登记调用明细,再交给现有 `RunBudget` 累加总账。
+
+每个模型调用 Trace 只包含 `component`、`component_round`、`usage_available`,并在 Usage 可用时包含 `input_tokens`、`output_tokens` 和 `total_tokens`;Usage 不可用时不写 Token 字段。Diagnosis Agent 的对应 `AgentStep.token_count` 回填 total Token;其他组件不伪装成 AgentStep。Run 结束事件同时写入预算总账、审计明细合计、Usage 不可用数量和 `tokens_reconciled`,从而显式暴露缺口而不是把未知 Token 当作零消耗。
+
+进入 `HarnessToolInterceptor` 的 Tool 请求若因进展协议、重复 scope、信息饱和或观察合同失败而未进入/未成功交付业务边界,记录 `TOOL_REQUEST_REJECTED`。事件只保留安全 Tool Call ID、Tool name 和稳定 `error_code`;业务参数、原始响应、内部异常和预算余量均不进入 Trace。
+
+不新增模型审计表:`diagnosis_trace_event` 是调用明细账,`RunBudget`/`diagnosis_run.total_token_count` 是 Run 总账,`AgentStep.token_count` 是 Diagnosis Agent 轮次摘要。这样避免三套可独立漂移的 Token 真相源。
+
+## Module Map
+
+```text
+ChatApplicationUseCase
+ -> DiagnosisChatExecutor
+ -> DiagnosisAgentUseCase
+ -> ReactAgent
+ -> HarnessModelInterceptor
+ -> HarnessToolInterceptor
+ -> DiagnosisProgressTracker (RunContext handle)
+ -> HarnessEvidenceTools -> ToolBoundary -> CanonicalInvocationStore
+ -> AgentObservationProjector
+ -> DiagnosisProgressProjector -> CanonicalInvocationStore
+ -> DiagnosisReleaseUseCase
+ -> no-conclusion reference validation / SafeFallbackFactory
+ -> EvidenceGuard -> EvidenceRepair -> SemanticGuard (conclusion only)
+ -> ChatRunStore -> named SSE content
+```
+
+## Interface Impact
+
+- 级别:L3 协作接口。
+- Agent-facing input 从裸业务 request 改为 `{previous_observation?, input}`。
+- 业务 adapter、backend、公开 HTTP/SSE、数据库和前端消费字段保持兼容。
+- 所有 Tool loop scripted tests 必须使用新 Envelope;KnowledgeQueryExecutor 若直接调用 registry bridge,继续走业务 request,不使用 Agent-facing Envelope。
+- 不提供旧/新 Schema 双轨;回滚以整个 change 为单位。
+
+## Risks / Trade-offs
+
+- [框架不能按预期生成嵌套强类型 Schema] -> 使用三个具体 Envelope record,并增加真实 callback schema 测试。
+- [STOP_REQUIRED 后异常被框架包装] -> 使用 cause-chain 分类测试,只有专用受控停止异常可转换为 Release 输入。
+- [预算终止与 Application active check 冲突] -> DiagnosisExecutionResult 显式标记已处理终止 Fallback,Application 仅对此窄分支跳过 success transition。
+- [canonical TTL 到期导致过程不完整] -> Run timeout 小于 canonical TTL;投影缺失 fail closed 为 limitation,不伪造事实。
+- [scope 规范化误判不同查询为重复] -> 首版只规范确定性字段,测试每个 Tool 的相同/不同 scope。
+- [模型伪造上一轮评价 ID] -> tracker 只接受当前 Run 最后一个待评价 ID,错序/重复消费均拒绝。
+- [无结论 Draft 绕过安全检查] -> 只跳过结论 Repair/SemanticGuard;引用真实性、当前 Run 所有权和负向语义仍确定性验证。
+- [模型完成排查后输出非法 Draft 导致过程丢失] -> 丢弃非法 Draft;仅当 ProgressSnapshot 含已验真 observed facts 时由 Release 确定性降级,无进展仍 fail closed。
+- [脏工作区行为丢失] -> 迁移预算 Fallback 的测试意图,实施前后用 scoped diff 核对,不覆盖无关修改。
+
+## Migration Plan
+
+1. 先加入状态/Envelope/scope/projector 类型和 focused contract tests,不切换 Release。
+2. 接入 interceptor、tracker、STOP_REQUIRED 和执行结果,固定真实框架 loop 行为。
+3. 接入 ProgressSnapshot 与统一 Release,迁移 Application 预算 Fallback。
+4. 更新中文 Prompt、Trace、配置和文档。
+5. 运行 focused、Harness 回归和全量测试;再用 Maven 启动项目执行原始未知 Query 的 SSE、日志、数据库 exact-run E2E。
+6. 回滚时整体回滚本 change;不单独恢复旧 Tool Schema 或 Application 预算分支。
+
+## Open Questions
+
+无。阈值、状态、Prompt、协议、去重边界、Release 所有权和兼容范围均已确认。
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/proposal.md b/openspec/changes/diagnosis-information-gain-stop-contract/proposal.md
new file mode 100644
index 0000000..9e74bef
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/proposal.md
@@ -0,0 +1,79 @@
+## Why
+
+Diagnosis Agent 面对知识库未知、日志为空或查询条件不足的问题时,当前只能依赖 Tool、Token 和轮次预算停止。模型可能不断改写查询继续调用 Tool,最终以 `BUDGET_EXHAUSTED` 或 `INTERNAL_FAILURE` 结束;用户只能看到通用错误,无法看到已经完成的检查和证据缺口。
+
+系统需要把“没有足够证据得出结论”视为正常、可发布的诊断结果,同时用确定性的 Harness 规则阻止无信息增益的空转,并继续由 EvidenceGuard 和 SemanticGuard 拦截无证据结论。
+
+## What Changes
+
+- 在 Run 内增加最小进展状态:`GAINED / NO_GAIN`、连续无增益计数、`COLLECTING / SATURATED` 和内部 `stop_reason`。
+- 通过模型原生 Tool Calling 注册统一 Tool Call Envelope;模型继续调用 Tool 时,在下一次调用中回传上一轮 `information_gain`,Harness 校验并消费控制字段,业务 Tool 请求保持原结构。
+- Harness 对 `NO_EVIDENCE` 和重复的 `tool_name + normalized_scope` 确定性赋值 `NO_GAIN`;其他成功非空结果由模型判断 `GAINED / NO_GAIN`。
+- 配置 `harness.chat.stop-after-consecutive-no-gain`,默认值为 `2`,达到阈值后拒绝新的业务 Tool 执行并要求结束。
+- 将 Canonical Tool Result 分为 Harness Control View 和白名单 Model Observation;保留 RAG `relevance_level`,不把内部计数、预算、检索轨迹或 raw response 放进模型上下文。
+- Tool Loop 结束时从当前 Run 已完成的 canonical 调用一次性投影 `ProgressSnapshot`,用于发布已检查来源、scope、客观结果和证据缺口。
+- 统一 `DiagnosisReleaseUseCase` 对正常 Draft、信息饱和和预算终止的发布决策;`conclusion=null` 不触发 EvidenceRepair,有结论时才执行完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard。
+- 最终 Draft 违反结构化输出契约时继续拒绝该 Draft;若当前 Run 已有可验真的 ProgressSnapshot,则仅从 canonical 过程确定性发布 `INSUFFICIENT_EVIDENCE`,没有安全进展时仍 fail closed。
+- 复用现有 `SafeFallback` 和 SSE/前端协议,增加或规范 `INSUFFICIENT_EVIDENCE`、`MISSING_REQUIRED_CONTEXT`,迁移当前位于 `ChatApplicationUseCase` 的预算兜底意图。
+- 精简中文 Diagnosis Prompt,明确模型不必须得出根因、允许零次 Tool 调用、正确但对当前推导无用的内容属于 `NO_GAIN`,且合法放弃是成功完成。
+- 增加最小模型调用审计:按组件和轮次记录 input/output/total Token,回填 Diagnosis AgentStep Token,并在 Run 结束时与预算总账对账;不记录 Prompt、模型正文或推理内容。
+- 对进入 Harness 后被协议、饱和或重复 scope 门禁拒绝的 Tool 请求记录安全 Trace,使预算 Tool 计数、实际执行和拒绝决策可区分;不记录 Tool 参数或原始响应。
+
+## Capabilities
+
+### New Capabilities
+
+- `diagnosis-information-gain-stop-contract`: 定义 Tool 信息增益回传、Harness 饱和停止、进展投影和安全发布行为。
+
+### Modified Capabilities
+
+- `single-react-diagnosis-agent`: Tool Schema 改为服务端注册的统一 Envelope,Prompt 和 Agent 结束行为支持合法放弃。
+- `canonical-tool-invocation-store`: canonical 结果继续作为 Tool 真相源,并支持当前 Run 在结束时投影进展,不新增第二套持久化真相。
+- `aci-evidence-tool-contracts`: RAG 投影保留 `relevance_level`,Tool 结果拆分控制视图和模型白名单视图。
+- `single-react-evidence-semantic-guards`: 无结论 Draft 不进入结论修复链;有结论仍执行完整证据与语义保护。
+- `single-react-chat-application-usecase`: 预算与饱和的业务 Fallback 由 Diagnosis Release 统一决策。
+
+## Scope
+
+- Diagnosis Agent Prompt、Tool Schema/Interceptor、三个 Tool contract 的 Agent-facing 包装。
+- RunContext 进展 tracker、确定性 scope 规范化和连续无增益停止门禁。
+- RAG/日志/MySQL canonical 控制视图与模型观察投影。
+- Diagnosis Agent 执行结果、ProgressSnapshot、Release、Fallback 和 Trace。
+- Draft 合同失败的脱敏 Trace 与“有安全进展才允许降级”的 Executor/Release 边界。
+- 模型调用 Token 明细、Run 对账摘要和 Tool 请求拒绝 Trace。
+- Spring 配置绑定、focused tests、回归测试和真实 SSE/日志/数据库 E2E。
+
+## Non-goals
+
+- 不引入 `new_count`、`next_action`、多级质量分数、`UNKNOWN` 或独立 Judge 模型。
+- 不做自然语言语义去重,只比较确定性的 `tool_name + normalized_scope`。
+- 不让 Tool 或 Harness 判断业务根因,不把 `REFERENCE` 自动等同于 `NO_GAIN`。
+- 不增加 `CONFIRMED / INSUFFICIENT_EVIDENCE / NEED_MORE_INFO` 第二套诊断生命周期状态。
+- 不修改公开 HTTP/SSE 事件结构,不新增前端页面或新的用户可见进度协议。
+- 不重新引入 Planner/Executor/Verifier/Composer 多 Agent 链路,不处理 ISS-015 的 Reasoning 原文审计治理。
+
+## Context Constraints
+
+- Tool 通过 `DiagnosisAgentFactory.tools(...)` 的原生 Tool Calling 通道注册,Prompt 不写 Tool 名称、Schema 或 Schema 占位符。
+- `CanonicalInvocationStore` 是当前 Run 完整 Tool 调用真相源;Run 内 tracker 只维护控制状态和完成调用索引,不复制 raw response。
+- `RunContext` 继续使用“结构不可变 + 线程安全可变 handle”的既有模式;阈值在 Run 启动时固定。
+- 现有 `SafeFallback.observed_facts / verified_sources / limitations / next_steps` 和前端渲染能力必须复用。
+- 当前未提交的预算 Fallback 修改保留用户价值,但业务 Release 决策需要从 `ChatApplicationUseCase` 迁移到 `DiagnosisReleaseUseCase`。
+- 当前环境缺少 `codebase-retrieval` 和 LSP;本次影响核查使用 `rg` 引用搜索、源码和测试阅读降级完成。用户已明确允许忽略 GitNexus。
+
+## Interface Impact
+
+- 级别:L3 协作接口。
+- 变更对象:模型可见的三个 Tool input schema、Tool Call 解析、Diagnosis Agent 到 Release 的内部执行结果。
+- 兼容性:业务 Tool request、Tool backend、公开 HTTP/SSE 和持久化表结构保持不变;旧的模型 Tool 参数形状不再被 Agent-facing schema 接受。
+- 消费者:`DiagnosisAgentFactory`、`HarnessEvidenceTools`、`HarnessToolInterceptor`、三个 adapter 及 scripted Tool-loop tests。
+- 回滚:整体回滚本 change;不提供双 Tool Schema 或兼容分支。
+
+## Risks
+
+- 框架 Tool Schema 生成或 ToolInterceptor 参数处理不符合 Envelope 假设,导致模型无法正确回传或业务 request 未被剥离。
+- STOP_REQUIRED 若未形成受控结束,模型可能继续请求 Tool,或 Agent 调用异常绕过统一 Release。
+- 最后一轮结果可能没有下一次 Tool Call 来回传模型评价;该情况只能表示模型主动结束,不能伪造 `GAINED / NO_GAIN`。
+- ProgressSnapshot 若读取不完整或混入 raw payload,会造成过程缺失或上下文/隐私边界倒退。
+- `conclusion=null` 与现有 EvidenceGuard 的 `ANALYSIS_MISSING` 规则冲突,需要明确区分“验证引用真实性”和“验证结论完整性”。
+- 现有脏工作区包含相关预算 Fallback 改动,实施时必须迁移其意图并避免覆盖其他历史修改。
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/specs/aci-evidence-tool-contracts/spec.md b/openspec/changes/diagnosis-information-gain-stop-contract/specs/aci-evidence-tool-contracts/spec.md
new file mode 100644
index 0000000..dd4c9df
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/specs/aci-evidence-tool-contracts/spec.md
@@ -0,0 +1,25 @@
+## MODIFIED Requirements
+
+### Requirement: RAG Tool contract SHALL expose only bounded document evidence
+The RAG business Request SHALL contain only `query`. The canonical RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, optional normalized `relevance_level`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`. The RAG projector SHALL accept upstream `relevanceLevel` or `relevance_level` and normalize recognized values without exposing raw relevance scores or retrieval traces.
+
+#### Scenario: RAG evidence is serialized
+- **WHEN** a RAG result contains a matching document excerpt and an upstream relevance level
+- **THEN** its canonical JSON preserves the bounded evidence and normalized `relevance_level` while excluding ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
+
+#### Scenario: RAG query has no evidence
+- **WHEN** RAG executes successfully without a usable document excerpt
+- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, returns an empty evidence list, and does not upgrade relevance into evidence
+
+## ADDED Requirements
+
+### Requirement: Tool results SHALL have separate Harness and model views
+Each successful evidence Tool result SHALL provide a Harness Control View and a bounded Model Observation derived from the same canonical result. The control view MAY contain returned counts, normalized scope, relevance, truncation and duplicate identity. The Model Observation SHALL contain only fields needed to understand and cite the result and SHALL NOT contain raw responses, internal scores, retrieval traces, duplicate fingerprints, counters, thresholds, budgets or store identities.
+
+#### Scenario: Model receives RAG observation
+- **WHEN** a canonical RAG result is READY
+- **THEN** the model receives Tool Call ID, actual query scope, bounded evidence, evidence status, optional coarse relevance and truncation, but not raw scores or Harness counters
+
+#### Scenario: Harness evaluates duplicate scope
+- **WHEN** the same normalized Tool scope is requested again
+- **THEN** the Harness can compare its control view identity without exposing that fingerprint to the model
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/specs/canonical-tool-invocation-store/spec.md b/openspec/changes/diagnosis-information-gain-stop-contract/specs/canonical-tool-invocation-store/spec.md
new file mode 100644
index 0000000..682cf23
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/specs/canonical-tool-invocation-store/spec.md
@@ -0,0 +1,12 @@
+## ADDED Requirements
+
+### Requirement: Run progress projection SHALL reference canonical records without duplicating truth
+The Run progress tracker SHALL retain only ordered canonical identities for completed Tool calls. At Tool-loop completion, a projector SHALL resolve those identities through the existing canonical store and SHALL accept only READY records owned by the current Run. The tracker SHALL NOT store or reconstruct raw Tool responses, complete Agent results, or a second durable evidence record.
+
+#### Scenario: Completed calls are projected
+- **WHEN** a Run ends after multiple READY canonical Tool invocations
+- **THEN** the progress projector reads each indexed canonical record in execution order and creates bounded observed facts
+
+#### Scenario: Indexed identity is invalid
+- **WHEN** an indexed canonical identity is missing, expired, cross-Run, PROJECTING, or ERROR
+- **THEN** it is excluded from observed facts and cannot become verified evidence
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/specs/diagnosis-information-gain-stop-contract/spec.md b/openspec/changes/diagnosis-information-gain-stop-contract/specs/diagnosis-information-gain-stop-contract/spec.md
new file mode 100644
index 0000000..f5c1f7b
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/specs/diagnosis-information-gain-stop-contract/spec.md
@@ -0,0 +1,110 @@
+## ADDED Requirements
+
+### Requirement: Harness SHALL track binary information gain per Run
+Each Diagnosis Run SHALL own a thread-safe progress tracker with `GAINED` and `NO_GAIN` as the only information-gain values. `GAINED` SHALL reset the consecutive no-gain count; `NO_GAIN` SHALL increment it. The tracker SHALL expose only `COLLECTING` or `SATURATED` as collection state and SHALL NOT create a second diagnosis lifecycle.
+
+#### Scenario: New evidence advances diagnosis
+- **WHEN** a valid pending Tool observation is evaluated as `GAINED`
+- **THEN** the Run remains `COLLECTING` and its consecutive no-gain count becomes zero
+
+#### Scenario: Consecutive observations do not advance diagnosis
+- **WHEN** valid `NO_GAIN` observations reach the Run's configured threshold
+- **THEN** the tracker becomes `SATURATED` with stop reason `INFORMATION_SATURATED`
+
+### Requirement: Harness SHALL assign only deterministic no-gain signals
+The Harness SHALL assign `NO_GAIN` when a successful Tool result has `evidence_status=NO_EVIDENCE` or when a requested `tool_name + normalized_scope` duplicates a successfully completed scope in the same Run. Other successful non-empty results, including RAG `REFERENCE`, SHALL require a model-provided `GAINED` or `NO_GAIN` before another Tool executes.
+
+#### Scenario: Empty scoped result
+- **WHEN** a Tool completes READY with `NO_EVIDENCE`
+- **THEN** the Harness records `NO_GAIN` without asking a model to judge Tool quality
+
+#### Scenario: Reference material is non-empty
+- **WHEN** RAG returns bounded evidence with `relevance_level=REFERENCE`
+- **THEN** the Harness leaves it pending for model evaluation and does not automatically mark it `NO_GAIN`
+
+#### Scenario: Equivalent structured scope repeats
+- **WHEN** the model requests the same Tool with the same deterministically normalized successful scope
+- **THEN** the Harness does not execute the Tool, records `NO_GAIN`, and does not create a second canonical invocation
+
+### Requirement: Consecutive no-gain threshold SHALL be fixed per Run
+The system SHALL bind `harness.chat.stop-after-consecutive-no-gain`, require a value of at least one, and default it to `2`. The value SHALL be copied into each new Run's tracker and SHALL NOT enter model context or change an active Run.
+
+#### Scenario: Default configuration is used
+- **WHEN** no external value is configured
+- **THEN** a new Run becomes saturated after two consecutive `NO_GAIN` decisions
+
+#### Scenario: Invalid threshold is configured
+- **WHEN** the configured threshold is zero or negative
+- **THEN** Harness configuration validation fails before serving Chat requests
+
+### Requirement: Saturated collection SHALL stop further Tool execution
+When collection becomes saturated, the Harness SHALL reject the pending or next business Tool execution and deliver one bounded `STOP_REQUIRED` observation with `reason=INFORMATION_SATURATED`. If the next model round requests another Tool, the Agent execution SHALL terminate through a typed controlled-stop path without consuming another Tool budget or publishing an internal failure.
+
+#### Scenario: Model evaluation reaches threshold
+- **WHEN** the next Tool Call reports `NO_GAIN` and that evaluation reaches the threshold
+- **THEN** the requested business Tool is not executed and the model receives one STOP_REQUIRED observation
+
+#### Scenario: Model ignores stop instruction
+- **WHEN** the model requests another Tool after STOP_REQUIRED was delivered
+- **THEN** the loop ends as controlled information saturation and proceeds to safe release
+
+### Requirement: Tool loop completion SHALL project bounded progress once
+At normal Draft completion, information saturation, or budget termination, the Harness SHALL use the current Run's completed canonical invocation keys to create one bounded `ProgressSnapshot`. The snapshot SHALL contain only safe source, actual scope, objective result summary, truncation and stop reason; it SHALL NOT contain Prompt, thought, raw Tool response, internal counters, remaining budget, Redis keys, or model evaluation rationale.
+
+#### Scenario: Unknown problem has completed checks
+- **WHEN** multiple Tools completed but no supported conclusion exists
+- **THEN** the release input contains their bounded checked scopes and objective results in stable execution order
+
+#### Scenario: Canonical record cannot be verified
+- **WHEN** an indexed record is missing, expired, incomplete, ERROR, or not owned by the current Run
+- **THEN** it is not projected as an observed fact and the snapshot records a bounded limitation
+
+### Requirement: Information stop reasons SHALL remain distinct from release outcomes
+The Harness SHALL distinguish `INFORMATION_SATURATED` from `BUDGET_LIMIT_REACHED`. Release SHALL continue to expose only `SUCCESS`, `FALLBACK`, `FAILED`, or `CANCELLED`, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
+
+#### Scenario: Low gain stops before budget exhaustion
+- **WHEN** consecutive no-gain reaches the configured threshold while hard budget remains
+- **THEN** Trace records `INFORMATION_SATURATED` and public release is a normal `FALLBACK`
+
+#### Scenario: Hard budget terminates collection
+- **WHEN** model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
+- **THEN** Trace retains `BUDGET_LIMIT_REACHED` and Release attempts a deterministic `FALLBACK` from the existing progress without an extra model call
+
+### Requirement: Diagnosis Prompt SHALL license bounded abandonment
+The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, `conclusion=null` is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is `NO_GAIN`. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, `next_action`, or Harness implementation details.
+
+#### Scenario: Required context is missing before any Tool call
+- **WHEN** the Query lacks the enterprise, time, service, error, or other context needed for a bounded query
+- **THEN** the model may return `conclusion=null` with `limitations.missing_info` without calling a Tool
+
+#### Scenario: Correct content is diagnostically useless
+- **WHEN** a Tool response is factually correct but only generic, repeated, or unable to change a current hypothesis
+- **THEN** the model treats it as `NO_GAIN` and does not continue with an equivalent query
+
+### Requirement: Model token audit SHALL be component-scoped and reconcilable
+Every Harness model call admitted by the Run budget SHALL receive a bounded component and component round. When Provider Usage is available, the same non-negative input and output Token counts SHALL update both the Run budget total and a `MODEL_TOKEN_USAGE` Trace event. Diagnosis Agent usage SHALL also update the matching `AgentStep.token_count`. At Run completion, Trace SHALL expose whether audited Token totals reconcile with the Run budget total and SHALL expose unavailable Usage counts without fabricating Token values.
+
+The audit SHALL NOT persist Prompt content, user or model text, reasoning content, Tool arguments, raw model responses, credentials, or provider-specific metadata.
+
+#### Scenario: Diagnosis Agent round returns Usage
+- **WHEN** a Diagnosis Agent model round returns input and output Token Usage
+- **THEN** its component round Trace and matching AgentStep contain the same total Token count and the Run total increases by that amount
+
+#### Scenario: Multiple model components execute
+- **WHEN** Router, Diagnosis Agent, Evidence Repair or Semantic Guard model calls execute in one Run
+- **THEN** each call is distinguishable by bounded component and component round and their audited Token sum can be compared with the Run total
+
+#### Scenario: Provider Usage is unavailable
+- **WHEN** a model attempt completes or fails without Provider Usage
+- **THEN** the audit marks Usage unavailable and Run reconciliation exposes the gap without estimating Token counts
+
+### Requirement: Rejected Tool requests SHALL remain observable without payload disclosure
+Every supported Tool request rejected by the Harness before a usable business observation is delivered SHALL emit a `TOOL_REQUEST_REJECTED` Trace event containing only safe Tool Call ID, Tool name and stable error code. Rejected requests SHALL remain distinguishable from canonical `TOOL_INVOCATION` events and SHALL NOT include Tool arguments, normalized scope content, raw responses, internal exception messages, credentials or budget values.
+
+#### Scenario: Progress protocol is invalid
+- **WHEN** a supported Tool request omits or misorders a required previous observation
+- **THEN** no business Tool executes and Trace records `INVALID_PROGRESS_PROTOCOL` for that Tool request
+
+#### Scenario: Duplicate or saturated request is blocked
+- **WHEN** a supported Tool request repeats a successful normalized scope or arrives after collection saturation
+- **THEN** Trace records the stable rejection reason while canonical invocation count remains unchanged
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/specs/single-react-chat-application-usecase/spec.md b/openspec/changes/diagnosis-information-gain-stop-contract/specs/single-react-chat-application-usecase/spec.md
new file mode 100644
index 0000000..de68cfa
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/specs/single-react-chat-application-usecase/spec.md
@@ -0,0 +1,33 @@
+## MODIFIED Requirements
+
+### Requirement: Fixed isolated executors
+The Application Use Case SHALL map SYSTEM_CHAT to one no-Tool model response, KNOWLEDGE_QUERY to exactly one lookup-knowledge invocation plus one bounded answer model call, and DIAGNOSIS to the single Diagnosis Agent followed by the Diagnosis Release boundary. Executors MUST NOT call one another or rewrite the Query. Information saturation and budget termination in Diagnosis SHALL be converted to safe content by Diagnosis Release, not by ChatApplicationUseCase.
+
+#### Scenario: System Chat
+- **WHEN** intent is SYSTEM_CHAT
+- **THEN** no evidence Tool, Diagnosis Agent, EvidenceGuard, or SemanticGuard is invoked
+
+#### Scenario: Knowledge Query
+- **WHEN** intent is KNOWLEDGE_QUERY
+- **THEN** only lookup_knowledge is invoked once and query_logs/query_mysql/Diagnosis ReAct are unavailable
+
+#### Scenario: Diagnosis succeeds with a conclusion
+- **WHEN** intent is DIAGNOSIS and the Draft passes the release guards
+- **THEN** the original Query and bounded PreviousTurn enter DiagnosisAgentUseCase and the Draft publishes only after DiagnosisReleaseUseCase
+
+#### Scenario: Diagnosis stops without a conclusion
+- **WHEN** Diagnosis collection is saturated, required context is missing, or a handled budget limit is reached
+- **THEN** DiagnosisReleaseUseCase returns bounded Fallback content and ChatApplicationUseCase only persists and transports that decision
+
+## ADDED Requirements
+
+### Requirement: Handled Diagnosis budget termination SHALL remain a Fallback release
+When Diagnosis Release has converted a recognized budget termination and existing safe progress into a Fallback, ChatApplicationUseCase SHALL persist that result exactly once without reclassifying it as `INTERNAL_FAILURE`, invoking another model, or rebuilding business fallback content. The internal Run lifecycle MAY retain `BUDGET_EXHAUSTED`, while the persisted public release outcome SHALL be `FALLBACK` and no PublishedResult SHALL be stored.
+
+#### Scenario: Tool budget ends after finite checks
+- **WHEN** Diagnosis reaches a hard Tool budget after at least one canonical safe observation and Release creates an insufficient-evidence fallback
+- **THEN** the Run persists status SUCCESS, release outcome FALLBACK, safe content and actual budget usage, and the SSE sends content followed by done
+
+#### Scenario: Budget ends without safe publishable progress
+- **WHEN** budget termination occurs before Diagnosis Release can form a safe bounded result
+- **THEN** the existing failure path remains fail closed and does not fabricate observed facts
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/specs/single-react-diagnosis-agent/spec.md b/openspec/changes/diagnosis-information-gain-stop-contract/specs/single-react-diagnosis-agent/spec.md
new file mode 100644
index 0000000..550af3b
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/specs/single-react-diagnosis-agent/spec.md
@@ -0,0 +1,56 @@
+## MODIFIED Requirements
+
+### Requirement: Evidence Tools SHALL execute through the Harness boundary
+The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions through native Tool Calling. Each Agent-facing Tool input SHALL be a typed Envelope containing optional `previous_observation` and required business `input`. A Tool interceptor SHALL validate and consume the previous observation, propagate the exact framework Tool Call ID, Run ID, Tool name, and unwrapped business JSON into the corresponding adapter and `ToolBoundary`, and SHALL enforce Run progress before execution. Successful observations SHALL use a bounded per-Tool whitelist projection; failed or control observations SHALL contain only stable safe semantics and SHALL NOT contain raw responses, internal exceptions, credentials, invocation lifecycle internals, counters, thresholds, or remaining budget.
+
+#### Scenario: Framework requests first RAG evidence
+- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`, no pending evaluation and a typed business input
+- **THEN** the RAG adapter receives only the unwrapped business request, uses exactly `call-1`, and the canonical invocation is owned by the current Run
+
+#### Scenario: Model continues after a non-empty result
+- **WHEN** the last successful Tool result is pending semantic evaluation and the model requests another Tool
+- **THEN** the Envelope must identify that exact prior Tool Call and contain `GAINED` or `NO_GAIN` before the new business Tool can execute
+
+#### Scenario: Tool execution fails
+- **WHEN** a registered adapter returns an error result
+- **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail
+
+#### Scenario: Unknown Tool is requested
+- **WHEN** a model requests a Tool outside the three registered definitions
+- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
+
+### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
+The single Chinese Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. The Agent SHALL NOT be required to find a root cause. If current evidence cannot support a diagnosis, no required context exists for a bounded Tool call, or available results are correct but do not advance any diagnosis hypothesis, the Agent SHALL stop with `conclusion=null`, describe actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist, fabricate a root cause, or make equivalent Tool calls merely to show activity.
+
+#### Scenario: Tool finds no evidence
+- **WHEN** completed Tool observations have no diagnostic information gain
+- **THEN** the final Draft has no confirmed Conclusion, records bounded checked scope and missing information, and does not make an equivalent retry
+
+#### Scenario: No bounded Tool query is possible
+- **WHEN** the Query lacks required enterprise, time, service or error context
+- **THEN** the Agent may perform zero Tool calls and returns a no-conclusion Draft whose `limitations.missing_info` identifies the required context
+
+#### Scenario: Harness requires stop
+- **WHEN** the Agent receives `STOP_REQUIRED`
+- **THEN** it emits a bounded final Draft without another Tool call
+
+## ADDED Requirements
+
+### Requirement: Diagnosis execution SHALL preserve controlled stop outcomes
+The internal Agent use case SHALL distinguish a valid Draft, controlled information saturation, and budget termination from an unclassified Agent failure. It SHALL return a bounded internal execution result containing optional Draft, ProgressSnapshot and stop reason, and SHALL NOT convert a recognized controlled stop into `DiagnosisAgentOutputException`.
+
+#### Scenario: Model ignores STOP_REQUIRED
+- **WHEN** the framework surfaces the typed collection-stopped signal after the final completion opportunity
+- **THEN** the Agent use case returns no Draft with `INFORMATION_SATURATED` and the current ProgressSnapshot
+
+#### Scenario: Unclassified framework failure
+- **WHEN** Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination
+- **THEN** execution still fails closed and no safe progress is fabricated
+
+#### Scenario: Invalid final Draft after verified checks
+- **WHEN** the final model text is empty or violates the strict DiagnosisDraft contract after the current Run has completed READY canonical Tool checks
+- **THEN** the invalid text is discarded, the output failure carries only the bounded ProgressSnapshot and safe failure metadata, and no model repair or loose JSON extraction occurs
+
+#### Scenario: Invalid final Draft without verified checks
+- **WHEN** the final model text violates the strict DiagnosisDraft contract before any publishable ProgressSnapshot exists
+- **THEN** execution remains failed and MUST NOT fabricate missing context, observed facts or a no-conclusion Draft
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/specs/single-react-evidence-semantic-guards/spec.md b/openspec/changes/diagnosis-information-gain-stop-contract/specs/single-react-evidence-semantic-guards/spec.md
new file mode 100644
index 0000000..1d4d0a7
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/specs/single-react-evidence-semantic-guards/spec.md
@@ -0,0 +1,51 @@
+## MODIFIED Requirements
+
+### Requirement: Deterministic Draft and evidence validation
+For a Draft with a non-null Conclusion, the Harness SHALL deterministically reject it unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs. For a Draft with `conclusion=null`, Release SHALL NOT require normal conclusion structure or invoke EvidenceRepair; any supplied Tool references SHALL still resolve to current-Run READY canonical invocations and SHALL obey positive/negative evidence semantics.
+
+#### Scenario: Duplicate or missing Analysis ID in concluded Draft
+- **WHEN** a Draft with a Conclusion contains a blank or duplicate Analysis ID
+- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
+
+#### Scenario: Broken report reference in concluded Draft
+- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
+- **THEN** EvidenceGuard rejects the Draft before semantic review
+
+#### Scenario: No-conclusion Draft has valid negative observation
+- **WHEN** a `conclusion=null` Draft cites a current-Run READY `NO_EVIDENCE` call as `NEGATIVE_OBSERVATION`
+- **THEN** Release accepts the reference authenticity without running EvidenceRepair or SemanticGuard
+
+#### Scenario: No-conclusion Draft fabricates a Tool reference
+- **WHEN** a `conclusion=null` Draft cites a missing, cross-Run, incomplete or ERROR Tool call
+- **THEN** the reference is excluded and cannot be published as an observed fact
+
+### Requirement: Fail-closed release policy
+The release use case SHALL publish the unchanged verified Draft only when a non-null Conclusion passes EvidenceGuard and SemanticGuard returns `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, final semantic technical failure, a valid no-conclusion Draft, information saturation, or handled budget termination. No-conclusion and controlled-stop release SHALL be deterministic from verified references and ProgressSnapshot and MUST NOT invoke a repair or semantic model call. A fallback result MUST NOT include an unsupported Draft, full verified snapshot, internal stop counters, or SemanticGuard reason.
+
+#### Scenario: Supported report release
+- **WHEN** EvidenceGuard succeeds for a concluded Draft and SemanticGuard returns `SUPPORTED`
+- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
+
+#### Scenario: Unsupported report fallback
+- **WHEN** SemanticGuard returns `UNSUPPORTED`
+- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot
+
+#### Scenario: Semantic review remains unavailable
+- **WHEN** all permitted technical attempts fail
+- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
+
+#### Scenario: Evidence validation fallback sources
+- **WHEN** evidence repair fails or the second EvidenceGuard rejects a concluded Draft
+- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
+
+#### Scenario: Missing context ends without Tool calls
+- **WHEN** a valid no-conclusion Draft has no Tool calls and identifies required missing context
+- **THEN** release outcome is `FALLBACK`, type is `MISSING_REQUIRED_CONTEXT`, and no guard model call occurs
+
+#### Scenario: Finite checks do not support a conclusion
+- **WHEN** a no-conclusion Draft or controlled stop has a non-empty verified ProgressSnapshot
+- **THEN** release outcome is `FALLBACK`, type is `INSUFFICIENT_EVIDENCE`, and observed facts describe only actual completed checks
+
+#### Scenario: Invalid Draft has publishable progress
+- **WHEN** the Agent's final Draft is rejected by strict parsing but its bounded ProgressSnapshot contains current-Run verified observed facts
+- **THEN** Release publishes `FALLBACK` with type `INSUFFICIENT_EVIDENCE` using only that snapshot and MUST NOT use any content from the invalid Draft
diff --git a/openspec/changes/diagnosis-information-gain-stop-contract/tasks.md b/openspec/changes/diagnosis-information-gain-stop-contract/tasks.md
new file mode 100644
index 0000000..20c40a8
--- /dev/null
+++ b/openspec/changes/diagnosis-information-gain-stop-contract/tasks.md
@@ -0,0 +1,51 @@
+## 1. Tool Envelope 与进展状态契约
+
+- [x] 1.1 新增 `InformationGain`、collection state、stop reason、上一轮评价和三个强类型 Agent-facing Tool Envelope,并用序列化/Schema 测试固定 snake_case、必填项和非法枚举拒绝行为
+- [x] 1.2 在 `ChatHarnessProperties` 增加 `stopAfterConsecutiveNoGain=2`、正数校验和配置绑定测试,并在创建 Run 时固定阈值
+- [x] 1.3 实现线程安全 `DiagnosisProgressTracker`,覆盖 GAINED 清零、NO_GAIN 累加、待评价 ID 单次消费、错序拒绝、饱和状态和一次 STOP_REQUIRED 交付
+- [x] 1.4 为 RAG、日志和 MySQL 实现确定性 scope projector,覆盖相同参数去重、不同结构化 scope 放行和失败 scope 不进入成功去重集合
+
+## 2. Canonical 双视图与 Tool 门禁
+
+- [x] 2.1 扩展 canonical RAG result 以兼容读取 `relevanceLevel/relevance_level` 并输出规范化 `relevance_level`,补充 REFERENCE 与空 evidence 回归测试
+- [x] 2.2 实现按 Tool 类型的 Harness Control View 和白名单 Model Observation projector,证明 returned count、内部指纹、raw score、预算和 raw response 不进入模型观察
+- [x] 2.3 将 `HarnessEvidenceTools` 的模型可见 Schema 切换为三个强类型 Envelope,同时保持 adapter bridge 只接收原业务 request JSON
+- [x] 2.4 重构 `HarnessToolInterceptor`,按“消费上一轮评价 -> 饱和门禁 -> scope 去重 -> ToolBoundary -> 记录结果”顺序执行,并覆盖 NO_EVIDENCE、REFERENCE、重复 scope、协议错误和 Tool ERROR
+- [x] 2.5 实现 STOP_REQUIRED 一次收尾机会和 `DiagnosisCollectionStoppedException`,用真实 scripted framework loop 证明忽略停止指令不会继续执行 Tool 或耗尽 Tool 预算
+
+## 3. ProgressSnapshot 与 Agent 执行结果
+
+- [x] 3.1 让 tracker 有序索引当前 Run 已完成的 canonical identity,并实现 `DiagnosisProgressProjector` 逐条验证 READY/Run ownership 后生成有界 snapshot
+- [x] 3.2 覆盖 canonical 记录缺失、过期、PROJECTING、ERROR、跨 Run 和投影超限,确保它们只产生 limitation 而不成为 observed fact
+- [x] 3.3 将 `DiagnosisAgentUseCase` 返回值改为可选 Draft、ProgressSnapshot 和 stop reason,更新全部调用方并保留未分类框架异常 fail closed
+- [x] 3.4 让预算/取消异常在 cause chain 中保持可分类:预算可进入统一 Release,取消保持 CANCELLED,其他异常不得伪装为正常证据不足
+
+## 4. 统一 Diagnosis Release
+
+- [x] 4.1 为无结论 Draft 增加确定性引用真实性/负向语义校验路径,证明它不调用 EvidenceRepair 或 SemanticGuard
+- [x] 4.2 扩展 `SafeFallbackFactory`,从 ProgressSnapshot 生成 `INSUFFICIENT_EVIDENCE` 或从 `limitations.missing_info` 生成 `MISSING_REQUIRED_CONTEXT`,复用现有 observed facts、sources、limitations 和 next steps
+- [x] 4.3 重构 `DiagnosisReleaseUseCase`,统一处理有结论 Draft、主动无结论、信息饱和、预算终止,并保留有结论时完整 EvidenceGuard/Repair/SemanticGuard 行为
+- [x] 4.4 更新 `DiagnosisChatExecutor` 与内部 result contract,使已处理的饱和/预算 Fallback 返回 Application 而不创建 PublishedResult
+- [x] 4.5 从 `ChatApplicationUseCase` 移除内容级预算 Fallback 决策,保留窄化的已处理预算终态持久化,并迁移现有预算耗尽测试意图
+- [x] 4.6 对最终 Draft 合同失败增加窄化降级:丢弃非法内容,仅在 ProgressSnapshot 含已验真 observed facts 时由 Release 生成 `INSUFFICIENT_EVIDENCE`,无安全进展仍 fail closed,并记录脱敏失败 Trace
+
+## 5. Prompt、Trace 与配置装配
+
+- [x] 5.1 按最小中文契约重写 Diagnosis Prompt:合法放弃、零 Tool、missing_info、GAINED/NO_GAIN、禁止等价重试和服从 STOP_REQUIRED,且不写 Tool 名/Schema/阈值/next_action
+- [x] 5.2 增加有界 progress/stop Trace 事件,记录信息增益生产者、scope 摘要、collection state 和 stop reason,并用测试证明不记录 Prompt、thought、raw payload、SQL 参数或预算余量
+- [x] 5.3 更新 `HarnessChatConfiguration` 装配 tracker factory、scope/projector、ProgressSnapshot 和统一 Release 依赖,修复全部构造调用与 Spring context 测试
+
+## 6. 验证与文档收口
+
+- [x] 6.1 运行 Tool contract、tracker、projector、interceptor、Agent loop、Guard、Release 和 Application focused tests,修复所有回归
+- [x] 6.2 运行 Harness 综合回归、`mvn test` 和 `openspec validate diagnosis-information-gain-stop-contract --strict`,记录命令与结果
+- [x] 6.3 用 Maven 启动真实应用,对“诊断切换企业失败的问题”执行 named SSE E2E,并从 `logs/` 与 `scripts/query_mysql.py` 按 exact sessionId/runId 核对过程、停止原因、Tool 次数和 FALLBACK 内容
+- [x] 6.4 更新 ISS-016、架构文档、glossary 和 OpenSpec task 状态,核对公开 SSE/前端协议无新增字段且工作区原有无关修改未被覆盖
+
+## 7. 模型 Token 与 Tool 拒绝审计补充
+
+- [x] 7.1 新增 Run 内模型调用组件/轮次账本和脱敏 Trace contract 测试,覆盖 Usage 可用、不可用、组件轮次和 Token 合计
+- [x] 7.2 将 Router、System Chat、Knowledge Answer、Diagnosis Agent、Evidence Repair 和 Semantic Guard 接入统一 Token 审计,并回填对应 AgentStep token_count
+- [x] 7.3 为进展协议错误、重复 scope、信息饱和和观察合同拒绝增加安全 `TOOL_REQUEST_REJECTED` Trace,证明不记录参数或原始响应
+- [x] 7.4 在 Run 完成时持久化 step_count 并记录 Token 总账/明细对账结果,补充 Store 和 Trace 回归测试
+- [x] 7.5 运行 focused、Harness 全量回归、strict validate 和真实 SSE/数据库 E2E,核对 Tool 执行/拒绝数量、Token 对账和 Trace 连续性
diff --git a/src/main/java/com/superbiz/agent/config/ChatHarnessProperties.java b/src/main/java/com/superbiz/agent/config/ChatHarnessProperties.java
index d545a1f..929c8cf 100644
--- a/src/main/java/com/superbiz/agent/config/ChatHarnessProperties.java
+++ b/src/main/java/com/superbiz/agent/config/ChatHarnessProperties.java
@@ -27,6 +27,7 @@ public class ChatHarnessProperties {
private int maxModelCalls = 24;
private int maxToolCalls = 24;
private int maxCallsPerTool = 8;
+ private int stopAfterConsecutiveNoGain = 2;
private long maxInputTokens = 100_000;
private long maxOutputTokens = 100_000;
private long maxTotalTokens = 200_000;
@@ -71,6 +72,7 @@ public class ChatHarnessProperties {
positive(maxModelCalls, "maxModelCalls");
positive(maxToolCalls, "maxToolCalls");
positive(maxCallsPerTool, "maxCallsPerTool");
+ positive(stopAfterConsecutiveNoGain, "stopAfterConsecutiveNoGain");
positive(maxInputTokens, "maxInputTokens");
positive(maxOutputTokens, "maxOutputTokens");
positive(maxTotalTokens, "maxTotalTokens");
diff --git a/src/main/java/com/superbiz/agent/config/HarnessChatConfiguration.java b/src/main/java/com/superbiz/agent/config/HarnessChatConfiguration.java
index 5c4cb9f..99dc733 100644
--- a/src/main/java/com/superbiz/agent/config/HarnessChatConfiguration.java
+++ b/src/main/java/com/superbiz/agent/config/HarnessChatConfiguration.java
@@ -27,6 +27,7 @@ import com.superbiz.agent.harness.release.DiagnosisReleaseUseCase;
import com.superbiz.agent.harness.release.EvidenceRepair;
import com.superbiz.agent.harness.release.EvidenceRepairLimits;
import com.superbiz.agent.harness.release.SafeFallbackFactory;
+import com.superbiz.agent.harness.progress.DiagnosisProgressProjector;
import com.superbiz.agent.harness.retry.HarnessRetryExecutor;
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
import com.superbiz.agent.harness.tool.adapter.MysqlToolAdapter;
@@ -50,6 +51,7 @@ import com.superbiz.agent.harness.application.KnowledgeQueryOperation;
import com.superbiz.agent.harness.application.SystemChatOperation;
import com.superbiz.agent.harness.application.IntentRouting;
import com.superbiz.agent.harness.audit.HarnessAgentAuditHook;
+import com.superbiz.agent.harness.audit.ModelCallAuditor;
import com.superbiz.agent.harness.audit.DiagnosisTraceRecorder;
import com.superbiz.agent.harness.audit.ToolInvocationAuditSink;
import com.superbiz.agent.repository.AgentStepRepository;
@@ -106,7 +108,8 @@ public class HarnessChatConfiguration {
new RunBudgetLimits(properties.getMaxModelCalls(), properties.getMaxToolCalls(),
properties.getMaxCallsPerTool(), properties.getMaxInputTokens(),
properties.getMaxOutputTokens(), properties.getMaxTotalTokens(), properties.getMaxRunBytes()),
- HarnessRetryPolicies.strict());
+ HarnessRetryPolicies.strict(),
+ properties.getStopAfterConsecutiveNoGain());
}
@Bean
@@ -116,8 +119,16 @@ public class HarnessChatConfiguration {
@Bean
public GuardModelCall guardModelCall(DiagnosisHarnessCore core, ChatModel chatModel,
- @Qualifier("harnessModelExecutor") ThreadPoolExecutor executor) {
- return new GuardModelCall(core, chatModel, executor);
+ @Qualifier("harnessModelExecutor") ThreadPoolExecutor executor,
+ ModelCallAuditor modelCallAuditor) {
+ return new GuardModelCall(core, chatModel, executor, modelCallAuditor);
+ }
+
+ @Bean
+ public ModelCallAuditor modelCallAuditor(DiagnosisHarnessCore core,
+ DiagnosisTraceRecorder traceRecorder,
+ AgentStepRepository steps) {
+ return new ModelCallAuditor(core, traceRecorder, steps);
}
@Bean
@@ -237,20 +248,30 @@ public class HarnessChatConfiguration {
HarnessEvidenceTools tools, ObjectMapper mapper,
AgentStepRepository steps,
AgentReasoningAuditRepository reasoningAudits,
- DiagnosisTraceRecorder traceRecorder) {
+ DiagnosisTraceRecorder traceRecorder,
+ ModelCallAuditor modelCallAuditor) {
return new DiagnosisAgentFactory(chatModel, core, tools, mapper,
List.of(new HarnessAgentAuditHook(
- steps, mapper, DiagnosisAgentFactory.AGENT_NAME, traceRecorder, reasoningAudits)));
+ steps, mapper, DiagnosisAgentFactory.AGENT_NAME, traceRecorder, reasoningAudits)),
+ traceRecorder, modelCallAuditor);
}
@Bean
public DiagnosisAgentUseCase diagnosisAgentUseCase(DiagnosisHarnessCore core,
DiagnosisAgentFactory factory,
ObjectMapper mapper,
- ChatHarnessProperties properties) {
+ ChatHarnessProperties properties,
+ DiagnosisProgressProjector progressProjector) {
return new DiagnosisAgentUseCase(core, factory, mapper, new DiagnosisAgentLimits(
properties.getDiagnosisMaxQueryBytes(), properties.getDiagnosisMaxPreviousTurnBytes(),
- properties.getDiagnosisMaxInputBytes(), properties.getDiagnosisMaxDraftBytes()));
+ properties.getDiagnosisMaxInputBytes(), properties.getDiagnosisMaxDraftBytes()),
+ progressProjector);
+ }
+
+ @Bean
+ public DiagnosisProgressProjector diagnosisProgressProjector(
+ CanonicalInvocationStore store, ToolCallKeyFactory keyFactory, ObjectMapper mapper) {
+ return new DiagnosisProgressProjector(store, keyFactory, mapper);
}
@Bean
@@ -281,13 +302,19 @@ public class HarnessChatConfiguration {
attempt -> { }, traceRecorder);
}
+ @Bean
+ public SafeFallbackFactory safeFallbackFactory() {
+ return new SafeFallbackFactory();
+ }
+
@Bean
public DiagnosisReleaseUseCase diagnosisReleaseUseCase(EvidenceGuard evidenceGuard,
EvidenceRepair repair,
SemanticGuard semanticGuard,
- DiagnosisTraceRecorder traceRecorder) {
+ DiagnosisTraceRecorder traceRecorder,
+ SafeFallbackFactory fallbackFactory) {
return new DiagnosisReleaseUseCase(
- evidenceGuard, repair, semanticGuard, new SafeFallbackFactory(), traceRecorder);
+ evidenceGuard, repair, semanticGuard, fallbackFactory, traceRecorder);
}
@Bean
@@ -327,9 +354,10 @@ public class HarnessChatConfiguration {
@Bean
public DiagnosisOperation diagnosisOperation(DiagnosisAgentUseCase agent,
- DiagnosisReleaseUseCase release,
- PublishedResultPolicy policy) {
- return new DiagnosisChatExecutor(agent, release, policy);
+ DiagnosisReleaseUseCase release,
+ PublishedResultPolicy policy,
+ DiagnosisTraceRecorder traceRecorder) {
+ return new DiagnosisChatExecutor(agent, release, policy, traceRecorder);
}
@Bean
diff --git a/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentExecution.java b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentExecution.java
new file mode 100644
index 0000000..6a68a22
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentExecution.java
@@ -0,0 +1,32 @@
+package com.superbiz.agent.harness.agent;
+
+import com.superbiz.agent.harness.contract.DiagnosisDraft;
+import com.superbiz.agent.harness.progress.DiagnosisProgressSnapshot;
+import com.superbiz.agent.harness.progress.DiagnosisStopReason;
+
+import java.util.Objects;
+
+public record DiagnosisAgentExecution(
+ DiagnosisDraft draft,
+ DiagnosisProgressSnapshot progress,
+ DiagnosisStopReason stopReason) {
+
+ public DiagnosisAgentExecution {
+ Objects.requireNonNull(progress, "progress must not be null");
+ if (draft == null && stopReason == null) {
+ throw new IllegalArgumentException("Execution without Draft requires a stop reason");
+ }
+ }
+
+ public static DiagnosisAgentExecution completed(
+ DiagnosisDraft draft, DiagnosisProgressSnapshot progress) {
+ return new DiagnosisAgentExecution(
+ Objects.requireNonNull(draft, "draft must not be null"), progress, null);
+ }
+
+ public static DiagnosisAgentExecution stopped(
+ DiagnosisProgressSnapshot progress, DiagnosisStopReason stopReason) {
+ return new DiagnosisAgentExecution(
+ null, progress, Objects.requireNonNull(stopReason, "stopReason must not be null"));
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentFactory.java b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentFactory.java
index d26ca18..5633a4d 100644
--- a/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentFactory.java
+++ b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentFactory.java
@@ -3,6 +3,8 @@ package com.superbiz.agent.harness.agent;
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
import com.alibaba.cloud.ai.graph.agent.hook.Hook;
import com.fasterxml.jackson.databind.ObjectMapper;
+import com.superbiz.agent.harness.audit.DiagnosisTraceRecorder;
+import com.superbiz.agent.harness.audit.ModelCallAuditor;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunContext;
import org.springframework.ai.chat.model.ChatModel;
@@ -19,6 +21,8 @@ public final class DiagnosisAgentFactory {
private final HarnessEvidenceTools evidenceTools;
private final ObjectMapper objectMapper;
private final List auditHooks;
+ private final DiagnosisTraceRecorder traceRecorder;
+ private final ModelCallAuditor modelCallAuditor;
private final String prompt;
public DiagnosisAgentFactory(ChatModel chatModel,
@@ -26,11 +30,35 @@ public final class DiagnosisAgentFactory {
HarnessEvidenceTools evidenceTools,
ObjectMapper objectMapper,
List extends Hook> auditHooks) {
+ this(chatModel, core, evidenceTools, objectMapper, auditHooks,
+ DiagnosisTraceRecorder.noop(), new ModelCallAuditor(core));
+ }
+
+ public DiagnosisAgentFactory(ChatModel chatModel,
+ DiagnosisHarnessCore core,
+ HarnessEvidenceTools evidenceTools,
+ ObjectMapper objectMapper,
+ List extends Hook> auditHooks,
+ DiagnosisTraceRecorder traceRecorder) {
+ this(chatModel, core, evidenceTools, objectMapper, auditHooks,
+ traceRecorder, new ModelCallAuditor(core));
+ }
+
+ public DiagnosisAgentFactory(ChatModel chatModel,
+ DiagnosisHarnessCore core,
+ HarnessEvidenceTools evidenceTools,
+ ObjectMapper objectMapper,
+ List extends Hook> auditHooks,
+ DiagnosisTraceRecorder traceRecorder,
+ ModelCallAuditor modelCallAuditor) {
this.chatModel = Objects.requireNonNull(chatModel, "chatModel must not be null");
this.core = Objects.requireNonNull(core, "core must not be null");
this.evidenceTools = Objects.requireNonNull(evidenceTools, "evidenceTools must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
this.auditHooks = List.copyOf(Objects.requireNonNull(auditHooks, "auditHooks must not be null"));
+ this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
+ this.modelCallAuditor = Objects.requireNonNull(
+ modelCallAuditor, "modelCallAuditor must not be null");
this.prompt = DiagnosisAgentPrompt.load();
}
@@ -43,8 +71,9 @@ public final class DiagnosisAgentFactory {
.systemPrompt(prompt)
.tools(evidenceTools.callbacks())
.interceptors(
- new HarnessModelInterceptor(core, context),
- new HarnessToolInterceptor(context, evidenceTools, objectMapper))
+ new HarnessModelInterceptor(core, context, modelCallAuditor),
+ new HarnessToolInterceptor(
+ context, evidenceTools, objectMapper, traceRecorder))
.hooks(auditHooks)
.outputSchema(new DiagnosisDraftOutputSchema(objectMapper).getFormat())
.returnReasoningContents(true)
diff --git a/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentOutputException.java b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentOutputException.java
index 71054a6..9d01e60 100644
--- a/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentOutputException.java
+++ b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentOutputException.java
@@ -1,12 +1,59 @@
package com.superbiz.agent.harness.agent;
+import com.superbiz.agent.harness.progress.DiagnosisProgressSnapshot;
+
+import java.util.Objects;
+
public final class DiagnosisAgentOutputException extends RuntimeException {
+ public enum Kind {
+ EXECUTION_FAILED,
+ EMPTY_DRAFT,
+ INVALID_JSON,
+ SCHEMA_INVALID
+ }
+
+ private final Kind kind;
+ private final long outputBytes;
+ private final DiagnosisProgressSnapshot progress;
+
public DiagnosisAgentOutputException(String message) {
- super(message);
+ this(message, null);
}
public DiagnosisAgentOutputException(String message, Throwable cause) {
+ this(message, cause, Kind.EXECUTION_FAILED, 0L, DiagnosisProgressSnapshot.empty());
+ }
+
+ public DiagnosisAgentOutputException(String message,
+ Throwable cause,
+ Kind kind,
+ long outputBytes,
+ DiagnosisProgressSnapshot progress) {
super(message, cause);
+ if (outputBytes < 0) {
+ throw new IllegalArgumentException("outputBytes must not be negative");
+ }
+ this.kind = Objects.requireNonNull(kind, "kind must not be null");
+ this.outputBytes = outputBytes;
+ this.progress = Objects.requireNonNull(progress, "progress must not be null");
+ }
+
+ public Kind kind() {
+ return kind;
+ }
+
+ public long outputBytes() {
+ return outputBytes;
+ }
+
+ public DiagnosisProgressSnapshot progress() {
+ return progress;
+ }
+
+ public boolean isDraftContractFailure() {
+ return kind == Kind.EMPTY_DRAFT
+ || kind == Kind.INVALID_JSON
+ || kind == Kind.SCHEMA_INVALID;
}
}
diff --git a/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentUseCase.java b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentUseCase.java
index 3c5f6a0..b41eecb 100644
--- a/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentUseCase.java
+++ b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentUseCase.java
@@ -2,13 +2,20 @@ package com.superbiz.agent.harness.agent;
import com.alibaba.cloud.ai.graph.RunnableConfig;
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
+import com.fasterxml.jackson.core.JsonParseException;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.DeserializationFeature;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.ObjectReader;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
+import com.superbiz.agent.harness.core.BudgetExceededException;
+import com.superbiz.agent.harness.core.RunAbortedException;
import com.superbiz.agent.harness.core.RunContext;
+import com.superbiz.agent.harness.core.RunState;
+import com.superbiz.agent.harness.progress.DiagnosisProgressProjection;
+import com.superbiz.agent.harness.progress.DiagnosisProgressSnapshot;
+import com.superbiz.agent.harness.progress.DiagnosisStopReason;
import org.springframework.ai.chat.messages.AssistantMessage;
import java.nio.charset.StandardCharsets;
@@ -23,11 +30,20 @@ public final class DiagnosisAgentUseCase {
private final ObjectMapper objectMapper;
private final ObjectReader draftReader;
private final DiagnosisAgentLimits limits;
+ private final DiagnosisProgressProjection progressProjection;
public DiagnosisAgentUseCase(DiagnosisHarnessCore core,
DiagnosisAgentFactory agentFactory,
ObjectMapper objectMapper,
DiagnosisAgentLimits limits) {
+ this(core, agentFactory, objectMapper, limits, DiagnosisProgressProjection.empty());
+ }
+
+ public DiagnosisAgentUseCase(DiagnosisHarnessCore core,
+ DiagnosisAgentFactory agentFactory,
+ ObjectMapper objectMapper,
+ DiagnosisAgentLimits limits,
+ DiagnosisProgressProjection progressProjection) {
this.core = Objects.requireNonNull(core, "core must not be null");
this.agentFactory = Objects.requireNonNull(agentFactory, "agentFactory must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
@@ -35,9 +51,11 @@ public final class DiagnosisAgentUseCase {
.with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
.with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
this.limits = Objects.requireNonNull(limits, "limits must not be null");
+ this.progressProjection = Objects.requireNonNull(
+ progressProjection, "progressProjection must not be null");
}
- public DiagnosisDraft execute(RunContext context, DiagnosisAgentInput input) {
+ public DiagnosisAgentExecution execute(RunContext context, DiagnosisAgentInput input) {
Objects.requireNonNull(context, "context must not be null");
Objects.requireNonNull(input, "input must not be null");
core.checkActive(context);
@@ -64,23 +82,85 @@ public final class DiagnosisAgentUseCase {
AssistantMessage response;
try {
response = agent.call(inputJson, config);
+ core.checkActive(context);
} catch (Exception e) {
+ DiagnosisAgentExecution controlled = controlledExecution(context, e);
+ if (controlled != null) {
+ return controlled;
+ }
throw new DiagnosisAgentOutputException("Diagnosis Agent execution failed", e);
}
- core.checkActive(context);
String output = response == null ? null : response.getText();
if (output == null || output.isBlank()) {
- throw new DiagnosisAgentOutputException("Diagnosis Agent returned an empty Draft");
+ throw new DiagnosisAgentOutputException(
+ "Diagnosis Agent returned an empty Draft",
+ null,
+ DiagnosisAgentOutputException.Kind.EMPTY_DRAFT,
+ 0L,
+ progressProjection.project(context));
}
long draftBytes = utf8Bytes(output);
checkLimit("draft", draftBytes, limits.maxDraftBytes());
- core.reserveRunBytes(context, draftBytes);
try {
- return draftReader.readValue(output);
- } catch (JsonProcessingException e) {
- throw new DiagnosisAgentOutputException("Diagnosis Agent returned an invalid Draft", e);
+ core.reserveRunBytes(context, draftBytes);
+ } catch (RuntimeException exception) {
+ DiagnosisAgentExecution controlled = controlledExecution(context, exception);
+ if (controlled != null) {
+ return controlled;
+ }
+ throw exception;
}
+ try {
+ DiagnosisDraft draft = draftReader.readValue(output);
+ return DiagnosisAgentExecution.completed(draft, progressProjection.project(context));
+ } catch (JsonProcessingException e) {
+ DiagnosisAgentOutputException.Kind kind = e instanceof JsonParseException
+ ? DiagnosisAgentOutputException.Kind.INVALID_JSON
+ : DiagnosisAgentOutputException.Kind.SCHEMA_INVALID;
+ throw new DiagnosisAgentOutputException(
+ "Diagnosis Agent returned an invalid Draft",
+ e,
+ kind,
+ draftBytes,
+ progressProjection.project(context));
+ }
+ }
+
+ private DiagnosisAgentExecution controlledExecution(RunContext context, Throwable failure) {
+ DiagnosisCollectionStoppedException stopped = findCause(
+ failure, DiagnosisCollectionStoppedException.class);
+ if (stopped != null) {
+ return DiagnosisAgentExecution.stopped(
+ progressProjection.project(context), stopped.stopReason());
+ }
+ RunAbortedException aborted = findCause(failure, RunAbortedException.class);
+ if (aborted != null) {
+ if (aborted.termination().state() == RunState.BUDGET_EXHAUSTED) {
+ context.progress().markBudgetLimitReached();
+ return DiagnosisAgentExecution.stopped(
+ progressProjection.project(context), DiagnosisStopReason.BUDGET_LIMIT_REACHED);
+ }
+ throw aborted;
+ }
+ BudgetExceededException budget = findCause(failure, BudgetExceededException.class);
+ if (budget != null || context.lifecycle().state() == RunState.BUDGET_EXHAUSTED) {
+ context.progress().markBudgetLimitReached();
+ return DiagnosisAgentExecution.stopped(
+ progressProjection.project(context), DiagnosisStopReason.BUDGET_LIMIT_REACHED);
+ }
+ return null;
+ }
+
+ private static T findCause(Throwable failure, Class type) {
+ Throwable current = failure;
+ while (current != null) {
+ if (type.isInstance(current)) {
+ return type.cast(current);
+ }
+ current = current.getCause();
+ }
+ return null;
}
private String writeJson(Object value) {
diff --git a/src/main/java/com/superbiz/agent/harness/agent/DiagnosisCollectionStoppedException.java b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisCollectionStoppedException.java
new file mode 100644
index 0000000..a5c0234
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/agent/DiagnosisCollectionStoppedException.java
@@ -0,0 +1,17 @@
+package com.superbiz.agent.harness.agent;
+
+import com.superbiz.agent.harness.progress.DiagnosisStopReason;
+
+public final class DiagnosisCollectionStoppedException extends RuntimeException {
+
+ private final DiagnosisStopReason stopReason;
+
+ public DiagnosisCollectionStoppedException(DiagnosisStopReason stopReason) {
+ super("Diagnosis collection stopped: " + stopReason);
+ this.stopReason = stopReason;
+ }
+
+ public DiagnosisStopReason stopReason() {
+ return stopReason;
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/agent/HarnessEvidenceTools.java b/src/main/java/com/superbiz/agent/harness/agent/HarnessEvidenceTools.java
index a9e3100..a22ec8c 100644
--- a/src/main/java/com/superbiz/agent/harness/agent/HarnessEvidenceTools.java
+++ b/src/main/java/com/superbiz/agent/harness/agent/HarnessEvidenceTools.java
@@ -1,5 +1,8 @@
package com.superbiz.agent.harness.agent;
+import com.fasterxml.jackson.databind.DeserializationFeature;
+import com.fasterxml.jackson.databind.ObjectMapper;
+import com.fasterxml.jackson.databind.ObjectReader;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.tool.adapter.MysqlToolAdapter;
import com.superbiz.agent.harness.tool.adapter.QueryLogsToolAdapter;
@@ -8,8 +11,11 @@ import com.superbiz.agent.harness.tool.boundary.ToolBoundaryResult;
import com.superbiz.agent.harness.tool.boundary.ToolCallRequestEnvelope;
import com.superbiz.agent.harness.tool.contract.AgentToolContracts;
import com.superbiz.agent.harness.tool.contract.MysqlToolRequest;
+import com.superbiz.agent.harness.tool.contract.MysqlToolCall;
import com.superbiz.agent.harness.tool.contract.QueryLogsRequest;
+import com.superbiz.agent.harness.tool.contract.QueryLogsToolCall;
import com.superbiz.agent.harness.tool.contract.RagToolRequest;
+import com.superbiz.agent.harness.tool.contract.RagToolCall;
import org.springframework.ai.tool.ToolCallback;
import org.springframework.ai.tool.function.FunctionToolCallback;
@@ -36,11 +42,11 @@ public final class HarnessEvidenceTools {
this.invokers = Map.copyOf(registered);
this.callbacks = List.of(
definition(AgentToolContracts.LOOKUP_KNOWLEDGE,
- AgentToolContracts.LOOKUP_KNOWLEDGE_DESCRIPTION, RagToolRequest.class),
+ AgentToolContracts.LOOKUP_KNOWLEDGE_DESCRIPTION, RagToolCall.class),
definition(AgentToolContracts.QUERY_LOGS,
- AgentToolContracts.QUERY_LOGS_DESCRIPTION, QueryLogsRequest.class),
+ AgentToolContracts.QUERY_LOGS_DESCRIPTION, QueryLogsToolCall.class),
definition(AgentToolContracts.QUERY_MYSQL,
- AgentToolContracts.QUERY_MYSQL_DESCRIPTION, MysqlToolRequest.class));
+ AgentToolContracts.QUERY_MYSQL_DESCRIPTION, MysqlToolCall.class));
}
public static HarnessEvidenceTools fromAdapters(RagToolAdapter ragAdapter,
@@ -72,6 +78,51 @@ public final class HarnessEvidenceTools {
return invoker.invoke(context, toolCallId, arguments);
}
+ public ParsedAgentToolCall parse(String toolName, String arguments, ObjectMapper objectMapper) {
+ Objects.requireNonNull(objectMapper, "objectMapper must not be null");
+ try {
+ ObjectReader reader = objectMapper.reader()
+ .with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
+ .with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
+ return switch (toolName) {
+ case AgentToolContracts.LOOKUP_KNOWLEDGE -> {
+ RagToolCall call = reader.forType(RagToolCall.class).readValue(arguments);
+ requireInput(call == null ? null : call.input());
+ yield parsed(call.previousObservation(), call.input(), objectMapper);
+ }
+ case AgentToolContracts.QUERY_LOGS -> {
+ QueryLogsToolCall call = reader.forType(QueryLogsToolCall.class).readValue(arguments);
+ requireInput(call == null ? null : call.input());
+ yield parsed(call.previousObservation(), call.input(), objectMapper);
+ }
+ case AgentToolContracts.QUERY_MYSQL -> {
+ MysqlToolCall call = reader.forType(MysqlToolCall.class).readValue(arguments);
+ requireInput(call == null ? null : call.input());
+ yield parsed(call.previousObservation(), call.input(), objectMapper);
+ }
+ default -> throw new IllegalArgumentException("Unsupported evidence Tool: " + toolName);
+ };
+ } catch (IllegalArgumentException exception) {
+ throw exception;
+ } catch (Exception exception) {
+ throw new IllegalArgumentException("Tool Call Envelope is invalid", exception);
+ }
+ }
+
+ private static ParsedAgentToolCall parsed(
+ com.superbiz.agent.harness.progress.PreviousObservation previousObservation,
+ Object input,
+ ObjectMapper objectMapper) throws Exception {
+ return new ParsedAgentToolCall(
+ previousObservation, input, objectMapper.writeValueAsString(input));
+ }
+
+ private static void requireInput(Object input) {
+ if (input == null) {
+ throw new IllegalArgumentException("Tool Call Envelope input is required");
+ }
+ }
+
private static EvidenceToolInvoker bridge(String toolName, AdapterCall adapter) {
return (context, toolCallId, arguments) -> adapter.execute(
context,
diff --git a/src/main/java/com/superbiz/agent/harness/agent/HarnessModelInterceptor.java b/src/main/java/com/superbiz/agent/harness/agent/HarnessModelInterceptor.java
index 8baa9fd..f81df18 100644
--- a/src/main/java/com/superbiz/agent/harness/agent/HarnessModelInterceptor.java
+++ b/src/main/java/com/superbiz/agent/harness/agent/HarnessModelInterceptor.java
@@ -6,6 +6,9 @@ import com.alibaba.cloud.ai.graph.agent.interceptor.ModelRequest;
import com.alibaba.cloud.ai.graph.agent.interceptor.ModelResponse;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunContext;
+import com.superbiz.agent.harness.audit.ModelCallAuditor;
+import com.superbiz.agent.harness.audit.ModelCallComponent;
+import com.superbiz.agent.harness.audit.ModelCallLedger;
import org.springframework.ai.chat.metadata.Usage;
import org.springframework.ai.chat.model.ChatResponse;
@@ -15,10 +18,17 @@ public final class HarnessModelInterceptor extends ModelInterceptor {
private final DiagnosisHarnessCore core;
private final RunContext context;
+ private final ModelCallAuditor auditor;
public HarnessModelInterceptor(DiagnosisHarnessCore core, RunContext context) {
+ this(core, context, new ModelCallAuditor(core));
+ }
+
+ public HarnessModelInterceptor(DiagnosisHarnessCore core, RunContext context,
+ ModelCallAuditor auditor) {
this.core = Objects.requireNonNull(core, "core must not be null");
this.context = Objects.requireNonNull(context, "context must not be null");
+ this.auditor = Objects.requireNonNull(auditor, "auditor must not be null");
}
@Override
@@ -31,23 +41,32 @@ public final class HarnessModelInterceptor extends ModelInterceptor {
Objects.requireNonNull(request, "request must not be null");
Objects.requireNonNull(handler, "handler must not be null");
core.beforeModelCall(context);
- ModelResponse response = handler.call(request);
- recordUsage(response == null ? null : response.getChatResponse());
+ ModelCallLedger.Call call = auditor.begin(context, ModelCallComponent.DIAGNOSIS_AGENT);
+ ModelResponse response;
+ try {
+ response = handler.call(request);
+ } catch (RuntimeException exception) {
+ auditor.recordUsage(context, call, 0, 0, false);
+ throw exception;
+ }
+ recordUsage(call, response == null ? null : response.getChatResponse());
core.checkActive(context);
return response;
}
- private void recordUsage(ChatResponse response) {
+ private void recordUsage(ModelCallLedger.Call call, ChatResponse response) {
if (response == null || response.getMetadata() == null) {
+ auditor.recordUsage(context, call, 0, 0, false);
return;
}
Usage usage = response.getMetadata().getUsage();
if (usage == null) {
+ auditor.recordUsage(context, call, 0, 0, false);
return;
}
long inputTokens = nonNegative(usage.getPromptTokens());
long outputTokens = nonNegative(usage.getCompletionTokens());
- core.recordTokens(context, inputTokens, outputTokens);
+ auditor.recordUsage(context, call, inputTokens, outputTokens, true);
}
private static long nonNegative(Integer value) {
diff --git a/src/main/java/com/superbiz/agent/harness/agent/HarnessToolInterceptor.java b/src/main/java/com/superbiz/agent/harness/agent/HarnessToolInterceptor.java
index 5170750..14b5484 100644
--- a/src/main/java/com/superbiz/agent/harness/agent/HarnessToolInterceptor.java
+++ b/src/main/java/com/superbiz/agent/harness/agent/HarnessToolInterceptor.java
@@ -6,8 +6,16 @@ import com.alibaba.cloud.ai.graph.agent.interceptor.ToolCallResponse;
import com.alibaba.cloud.ai.graph.agent.interceptor.ToolInterceptor;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.ObjectMapper;
+import com.superbiz.agent.harness.audit.DiagnosisTraceRecorder;
+import com.superbiz.agent.harness.audit.TraceAuditEvents;
import com.superbiz.agent.harness.contract.InvocationStatus;
import com.superbiz.agent.harness.core.RunContext;
+import com.superbiz.agent.harness.progress.CompletedToolCall;
+import com.superbiz.agent.harness.progress.DiagnosisCollectionState;
+import com.superbiz.agent.harness.progress.DiagnosisProgressSnapshotState;
+import com.superbiz.agent.harness.progress.DiagnosisStopReason;
+import com.superbiz.agent.harness.progress.InformationGain;
+import com.superbiz.agent.harness.progress.ToolScopeNormalizer;
import com.superbiz.agent.harness.tool.boundary.ToolBoundaryResult;
import java.util.LinkedHashMap;
@@ -19,13 +27,26 @@ public final class HarnessToolInterceptor extends ToolInterceptor {
private final RunContext context;
private final HarnessEvidenceTools evidenceTools;
private final ObjectMapper objectMapper;
+ private final ToolScopeNormalizer scopeNormalizer;
+ private final ToolResultViewProjector viewProjector;
+ private final DiagnosisTraceRecorder traceRecorder;
public HarnessToolInterceptor(RunContext context,
HarnessEvidenceTools evidenceTools,
ObjectMapper objectMapper) {
+ this(context, evidenceTools, objectMapper, DiagnosisTraceRecorder.noop());
+ }
+
+ public HarnessToolInterceptor(RunContext context,
+ HarnessEvidenceTools evidenceTools,
+ ObjectMapper objectMapper,
+ DiagnosisTraceRecorder traceRecorder) {
this.context = Objects.requireNonNull(context, "context must not be null");
this.evidenceTools = Objects.requireNonNull(evidenceTools, "evidenceTools must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
+ this.scopeNormalizer = new ToolScopeNormalizer(objectMapper);
+ this.viewProjector = new ToolResultViewProjector(objectMapper);
+ this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
}
@Override
@@ -41,10 +62,76 @@ public final class HarnessToolInterceptor extends ToolInterceptor {
return handler.call(request);
}
+ DiagnosisProgressSnapshotState before = context.progress().snapshot();
+ if (before.collectionState() == DiagnosisCollectionState.SATURATED
+ && before.stopInstructionDelivered()) {
+ recordRejection(request, "COLLECTION_STOPPED");
+ throw new DiagnosisCollectionStoppedException(before.stopReason());
+ }
+
+ ParsedAgentToolCall call;
+ String normalizedScope;
+ try {
+ call = evidenceTools.parse(request.getToolName(), request.getArguments(), objectMapper);
+ context.progress().applyPreviousObservation(call.previousObservation());
+ DiagnosisProgressSnapshotState evaluated = context.progress().snapshot();
+ recordModelProgress(before, call, evaluated);
+ if (evaluated.collectionState() == DiagnosisCollectionState.SATURATED) {
+ recordRejection(request, "INFORMATION_SATURATED");
+ return stopRequired(request, evaluated.stopReason());
+ }
+ normalizedScope = scopeNormalizer.normalize(request.getToolName(), call.businessInput());
+ } catch (IllegalArgumentException | IllegalStateException exception) {
+ recordRejection(request, "INVALID_PROGRESS_PROTOCOL");
+ return safeError(request, "INVALID_PROGRESS_PROTOCOL");
+ }
+
+ if (context.progress().isDuplicate(request.getToolName(), normalizedScope)) {
+ recordRejection(request, "DUPLICATE_SCOPE");
+ context.progress().recordDuplicateScope();
+ DiagnosisProgressSnapshotState duplicate = context.progress().snapshot();
+ recordProgress(request.getToolCallId(), request.getToolName(), normalizedScope,
+ InformationGain.NO_GAIN, "HARNESS", duplicate);
+ if (duplicate.collectionState() == DiagnosisCollectionState.SATURATED) {
+ return stopRequired(request, duplicate.stopReason());
+ }
+ return duplicateScope(request);
+ }
+
ToolBoundaryResult result = evidenceTools.invoke(
- context, request.getToolName(), request.getToolCallId(), request.getArguments());
+ context, request.getToolName(), request.getToolCallId(), call.businessArguments());
if (result.status() == InvocationStatus.READY) {
- return ToolCallResponse.of(request.getToolCallId(), request.getToolName(), result.agentResult());
+ ToolControlView control = viewProjector.controlView(result.agentResult());
+ if (control.evidenceStatus() != result.evidenceStatus()) {
+ recordRejection(request, "OBSERVATION_CONTRACT_MISMATCH");
+ return safeError(request, "OBSERVATION_CONTRACT_MISMATCH");
+ }
+ context.progress().recordCompleted(
+ new CompletedToolCall(
+ request.getToolCallId(), request.getToolName(), normalizedScope),
+ result.evidenceStatus());
+ DiagnosisProgressSnapshotState completed = context.progress().snapshot();
+ if (result.evidenceStatus()
+ == com.superbiz.agent.harness.contract.EvidenceStatus.NO_EVIDENCE) {
+ recordProgress(request.getToolCallId(), request.getToolName(), normalizedScope,
+ InformationGain.NO_GAIN, "HARNESS", completed);
+ }
+ boolean stopRequired = completed.collectionState() == DiagnosisCollectionState.SATURATED
+ && context.progress().claimStopInstruction();
+ if (stopRequired) {
+ traceRecorder.record(TraceAuditEvents.collectionStop(
+ context, request.getToolCallId(), request.getToolName(), completed));
+ }
+ String observation = viewProjector.modelObservation(
+ request.getToolName(), result.agentResult(), normalizedScope,
+ stopRequired, completed.stopReason());
+ return ToolCallResponse.of(request.getToolCallId(), request.getToolName(), observation);
+ }
+ if ("BUDGET_EXHAUSTED".equals(result.errorCode())) {
+ context.progress().markBudgetLimitReached();
+ traceRecorder.record(TraceAuditEvents.collectionStop(
+ context, request.getToolCallId(), request.getToolName(),
+ context.progress().snapshot()));
}
return ToolCallResponse.builder()
.toolCallId(request.getToolCallId())
@@ -55,6 +142,90 @@ public final class HarnessToolInterceptor extends ToolInterceptor {
.build();
}
+ private ToolCallResponse stopRequired(ToolCallRequest request, DiagnosisStopReason reason) {
+ if (!context.progress().claimStopInstruction()) {
+ throw new DiagnosisCollectionStoppedException(reason);
+ }
+ traceRecorder.record(TraceAuditEvents.collectionStop(
+ context, request.getToolCallId(), request.getToolName(),
+ context.progress().snapshot()));
+ Map observation = new LinkedHashMap<>();
+ observation.put("tool_call_id", request.getToolCallId());
+ observation.put("stop_required", true);
+ observation.put("reason", reason.name());
+ return ToolCallResponse.of(
+ request.getToolCallId(), request.getToolName(), writeObservation(observation));
+ }
+
+ private ToolCallResponse duplicateScope(ToolCallRequest request) {
+ Map observation = new LinkedHashMap<>();
+ observation.put("tool_call_id", request.getToolCallId());
+ observation.put("observation_status", "DUPLICATE_SCOPE");
+ observation.put("information_gain", "NO_GAIN");
+ observation.put("stop_required", false);
+ return ToolCallResponse.of(
+ request.getToolCallId(), request.getToolName(), writeObservation(observation));
+ }
+
+ private ToolCallResponse safeError(ToolCallRequest request, String errorCode) {
+ Map observation = new LinkedHashMap<>();
+ observation.put("evidence_status", "ERROR");
+ observation.put("tool_call_id", request.getToolCallId());
+ observation.put("error_code", errorCode);
+ return ToolCallResponse.builder()
+ .toolCallId(request.getToolCallId())
+ .toolName(request.getToolName())
+ .content(writeObservation(observation))
+ .status("error")
+ .metadata(Map.of("error", true))
+ .build();
+ }
+
+ private void recordRejection(ToolCallRequest request, String errorCode) {
+ traceRecorder.record(TraceAuditEvents.toolRequestRejected(
+ context, request.getToolCallId(), request.getToolName(), errorCode));
+ }
+
+ private String writeObservation(Map observation) {
+ try {
+ return objectMapper.writeValueAsString(observation);
+ } catch (JsonProcessingException exception) {
+ return "{\"evidence_status\":\"ERROR\",\"error_code\":\"SERIALIZATION_ERROR\"}";
+ }
+ }
+
+ private void recordModelProgress(
+ DiagnosisProgressSnapshotState before,
+ ParsedAgentToolCall call,
+ DiagnosisProgressSnapshotState after) {
+ if (call.previousObservation() == null) {
+ return;
+ }
+ before.completedToolCalls().stream()
+ .filter(completed -> completed.toolCallId()
+ .equals(call.previousObservation().toolCallId()))
+ .findFirst()
+ .ifPresent(completed -> recordProgress(
+ completed.toolCallId(), completed.toolName(), completed.normalizedScope(),
+ call.previousObservation().informationGain(), "MODEL", after));
+ }
+
+ private void recordProgress(
+ String toolCallId,
+ String toolName,
+ String normalizedScope,
+ InformationGain informationGain,
+ String producer,
+ DiagnosisProgressSnapshotState state) {
+ traceRecorder.record(TraceAuditEvents.toolProgress(
+ context, toolCallId, toolName, scopeSummary(toolName, normalizedScope),
+ informationGain, producer, state));
+ }
+
+ private String scopeSummary(String toolName, String normalizedScope) {
+ return toolName + "#" + String.format("%08x", normalizedScope.hashCode());
+ }
+
private String errorObservation(ToolBoundaryResult result) {
Map observation = new LinkedHashMap<>();
observation.put("evidence_status", result.evidenceStatus());
diff --git a/src/main/java/com/superbiz/agent/harness/agent/ParsedAgentToolCall.java b/src/main/java/com/superbiz/agent/harness/agent/ParsedAgentToolCall.java
new file mode 100644
index 0000000..057efc9
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/agent/ParsedAgentToolCall.java
@@ -0,0 +1,18 @@
+package com.superbiz.agent.harness.agent;
+
+import com.superbiz.agent.harness.progress.PreviousObservation;
+
+public record ParsedAgentToolCall(
+ PreviousObservation previousObservation,
+ Object businessInput,
+ String businessArguments) {
+
+ public ParsedAgentToolCall {
+ if (businessInput == null) {
+ throw new IllegalArgumentException("businessInput must not be null");
+ }
+ if (businessArguments == null || businessArguments.isBlank()) {
+ throw new IllegalArgumentException("businessArguments must not be blank");
+ }
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/agent/ToolControlView.java b/src/main/java/com/superbiz/agent/harness/agent/ToolControlView.java
new file mode 100644
index 0000000..aa1bc08
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/agent/ToolControlView.java
@@ -0,0 +1,11 @@
+package com.superbiz.agent.harness.agent;
+
+import com.superbiz.agent.harness.contract.EvidenceStatus;
+import com.superbiz.agent.harness.tool.contract.RagRelevanceLevel;
+
+public record ToolControlView(
+ EvidenceStatus evidenceStatus,
+ int returnedCount,
+ RagRelevanceLevel relevanceLevel,
+ boolean truncated) {
+}
diff --git a/src/main/java/com/superbiz/agent/harness/agent/ToolResultViewProjector.java b/src/main/java/com/superbiz/agent/harness/agent/ToolResultViewProjector.java
new file mode 100644
index 0000000..0029da4
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/agent/ToolResultViewProjector.java
@@ -0,0 +1,120 @@
+package com.superbiz.agent.harness.agent;
+
+import com.fasterxml.jackson.core.JsonProcessingException;
+import com.fasterxml.jackson.databind.JsonNode;
+import com.fasterxml.jackson.databind.ObjectMapper;
+import com.fasterxml.jackson.databind.node.ObjectNode;
+import com.superbiz.agent.harness.contract.EvidenceStatus;
+import com.superbiz.agent.harness.progress.DiagnosisStopReason;
+import com.superbiz.agent.harness.tool.contract.AgentToolContracts;
+import com.superbiz.agent.harness.tool.contract.RagRelevanceLevel;
+
+import java.util.List;
+import java.util.Objects;
+
+public final class ToolResultViewProjector {
+
+ private static final List COMMON_FIELDS = List.of(
+ "evidence_status", "tool_call_id", "truncated");
+
+ private final ObjectMapper objectMapper;
+
+ public ToolResultViewProjector(ObjectMapper objectMapper) {
+ this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
+ }
+
+ public ToolControlView controlView(String agentResult) {
+ JsonNode root = readObject(agentResult);
+ EvidenceStatus status;
+ try {
+ status = EvidenceStatus.valueOf(requiredText(root, "evidence_status"));
+ } catch (IllegalArgumentException exception) {
+ throw new IllegalArgumentException("Invalid evidence status", exception);
+ }
+ RagRelevanceLevel relevance = null;
+ String relevanceText = text(root, "relevance_level");
+ if (!relevanceText.isBlank()) {
+ try {
+ relevance = RagRelevanceLevel.valueOf(relevanceText);
+ } catch (IllegalArgumentException exception) {
+ throw new IllegalArgumentException("Invalid relevance level", exception);
+ }
+ }
+ return new ToolControlView(
+ status,
+ Math.max(0, root.path("returned_count").asInt(0)),
+ relevance,
+ root.path("truncated").asBoolean(false));
+ }
+
+ public String modelObservation(String toolName,
+ String agentResult,
+ String normalizedScope,
+ boolean stopRequired,
+ DiagnosisStopReason stopReason) {
+ JsonNode root = readObject(agentResult);
+ ObjectNode observation = objectMapper.createObjectNode();
+ copy(root, observation, COMMON_FIELDS);
+ switch (toolName) {
+ case AgentToolContracts.LOOKUP_KNOWLEDGE -> {
+ ObjectNode scope = observation.putObject("scope");
+ scope.put("query", text(root, "query"));
+ copy(root, observation, List.of("evidence", "relevance_level"));
+ }
+ case AgentToolContracts.QUERY_LOGS ->
+ copy(root, observation, List.of("source_kind", "scope", "patterns", "events"));
+ case AgentToolContracts.QUERY_MYSQL -> {
+ try {
+ observation.set("scope", objectMapper.readTree(normalizedScope));
+ } catch (JsonProcessingException exception) {
+ throw new IllegalArgumentException("Normalized scope is invalid", exception);
+ }
+ copy(root, observation, List.of("columns", "rows"));
+ }
+ default -> throw new IllegalArgumentException("Unsupported evidence Tool: " + toolName);
+ }
+ if (stopRequired) {
+ observation.put("stop_required", true);
+ observation.put("reason", Objects.requireNonNull(stopReason, "stopReason must not be null").name());
+ }
+ try {
+ return objectMapper.writeValueAsString(observation);
+ } catch (JsonProcessingException exception) {
+ throw new IllegalArgumentException("Model observation is not serializable", exception);
+ }
+ }
+
+ private JsonNode readObject(String value) {
+ try {
+ JsonNode root = objectMapper.readTree(value);
+ if (root == null || !root.isObject()) {
+ throw new IllegalArgumentException("Tool result must be a JSON object");
+ }
+ return root;
+ } catch (JsonProcessingException exception) {
+ throw new IllegalArgumentException("Tool result must be valid JSON", exception);
+ }
+ }
+
+ private static void copy(JsonNode source, ObjectNode target, List fields) {
+ for (String field : fields) {
+ JsonNode value = source.get(field);
+ if (value != null && !value.isNull()) {
+ target.set(field, value);
+ }
+ }
+ }
+
+ private static String requiredText(JsonNode root, String field) {
+ String value = text(root, field);
+ if (value.isBlank()) {
+ throw new IllegalArgumentException(field + " must not be blank");
+ }
+ return value;
+ }
+
+ private static String text(JsonNode root, String field) {
+ JsonNode value = root.get(field);
+ return value == null || value.isNull() ? "" : value.asText("");
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/application/ChatApplicationUseCase.java b/src/main/java/com/superbiz/agent/harness/application/ChatApplicationUseCase.java
index 2d5320a..59e6224 100644
--- a/src/main/java/com/superbiz/agent/harness/application/ChatApplicationUseCase.java
+++ b/src/main/java/com/superbiz/agent/harness/application/ChatApplicationUseCase.java
@@ -107,8 +107,7 @@ public final class ChatApplicationUseCase {
PathResult path = executePath(
intent, context, request.query(), previousTurn.orElse(null), observer);
- core.checkActive(context);
- core.completeSuccess(context);
+ completePath(context, intent, path);
String safeJson = write(path.content());
persistFinish(context, intent, path.outcome(), safeJson,
path.publishedResult(), durationMillis(startedNanos));
@@ -131,6 +130,20 @@ public final class ChatApplicationUseCase {
}
}
+ private void completePath(RunContext context, IntentType intent, PathResult path) {
+ if (path.handledBudgetTermination()) {
+ if (intent != IntentType.DIAGNOSIS
+ || path.outcome() != ReleaseOutcome.FALLBACK
+ || path.publishedResult() != null
+ || context.lifecycle().state() != RunState.BUDGET_EXHAUSTED) {
+ throw new IllegalStateException("Invalid handled Diagnosis budget fallback");
+ }
+ return;
+ }
+ core.checkActive(context);
+ core.completeSuccess(context);
+ }
+
private PathResult executePath(IntentType intent,
RunContext context,
String query,
@@ -140,18 +153,19 @@ public final class ChatApplicationUseCase {
case SYSTEM_CHAT -> {
observer.onStatus(ChatApplicationStatus.SYSTEM_RESPONDING);
yield new PathResult(
- ReleaseOutcome.SUCCESS, systemChat.execute(context, query), null);
+ ReleaseOutcome.SUCCESS, systemChat.execute(context, query), null, false);
}
case KNOWLEDGE_QUERY -> {
observer.onStatus(ChatApplicationStatus.KNOWLEDGE_SEARCHING);
observer.onStatus(ChatApplicationStatus.KNOWLEDGE_ANSWERING);
yield new PathResult(
- ReleaseOutcome.SUCCESS, knowledgeQuery.execute(context, query), null);
+ ReleaseOutcome.SUCCESS, knowledgeQuery.execute(context, query), null, false);
}
case DIAGNOSIS -> {
DiagnosisExecutionResult result = diagnosis.execute(
context, query, previousTurn, observer::onStatus);
- yield new PathResult(result.outcome(), result.content(), result.publishedResult());
+ yield new PathResult(result.outcome(), result.content(), result.publishedResult(),
+ result.handledBudgetTermination());
}
};
}
@@ -251,7 +265,8 @@ public final class ChatApplicationUseCase {
private record PathResult(
ReleaseOutcome outcome,
ChatApplicationContent content,
- PublishedResult publishedResult) {
+ PublishedResult publishedResult,
+ boolean handledBudgetTermination) {
}
private static final class CoreRunControl implements ChatRunControl {
diff --git a/src/main/java/com/superbiz/agent/harness/application/DiagnosisExecutionResult.java b/src/main/java/com/superbiz/agent/harness/application/DiagnosisExecutionResult.java
index 81f249e..90ba56e 100644
--- a/src/main/java/com/superbiz/agent/harness/application/DiagnosisExecutionResult.java
+++ b/src/main/java/com/superbiz/agent/harness/application/DiagnosisExecutionResult.java
@@ -8,7 +8,15 @@ import java.util.Objects;
public record DiagnosisExecutionResult(
ReleaseOutcome outcome,
ChatApplicationContent content,
- PublishedResult publishedResult) {
+ PublishedResult publishedResult,
+ boolean handledBudgetTermination) {
+
+ public DiagnosisExecutionResult(
+ ReleaseOutcome outcome,
+ ChatApplicationContent content,
+ PublishedResult publishedResult) {
+ this(outcome, content, publishedResult, false);
+ }
public DiagnosisExecutionResult {
Objects.requireNonNull(outcome, "outcome must not be null");
@@ -25,5 +33,9 @@ public record DiagnosisExecutionResult(
if (outcome != ReleaseOutcome.SUCCESS && publishedResult != null) {
throw new IllegalArgumentException("only success can contain PublishedResult");
}
+ if (handledBudgetTermination && outcome != ReleaseOutcome.FALLBACK) {
+ throw new IllegalArgumentException(
+ "handled budget termination requires fallback release");
+ }
}
}
diff --git a/src/main/java/com/superbiz/agent/harness/application/executor/DiagnosisChatExecutor.java b/src/main/java/com/superbiz/agent/harness/application/executor/DiagnosisChatExecutor.java
index b6506b7..0ce8e2b 100644
--- a/src/main/java/com/superbiz/agent/harness/application/executor/DiagnosisChatExecutor.java
+++ b/src/main/java/com/superbiz/agent/harness/application/executor/DiagnosisChatExecutor.java
@@ -1,6 +1,8 @@
package com.superbiz.agent.harness.application.executor;
import com.superbiz.agent.harness.agent.DiagnosisAgentInput;
+import com.superbiz.agent.harness.agent.DiagnosisAgentExecution;
+import com.superbiz.agent.harness.agent.DiagnosisAgentOutputException;
import com.superbiz.agent.harness.agent.DiagnosisAgentUseCase;
import com.superbiz.agent.harness.application.ChatApplicationStatus;
import com.superbiz.agent.harness.application.DiagnosisContent;
@@ -8,7 +10,8 @@ import com.superbiz.agent.harness.application.DiagnosisExecutionResult;
import com.superbiz.agent.harness.application.DiagnosisOperation;
import com.superbiz.agent.harness.application.FallbackContent;
import com.superbiz.agent.harness.application.persistence.PublishedResultPolicy;
-import com.superbiz.agent.harness.contract.DiagnosisDraft;
+import com.superbiz.agent.harness.audit.DiagnosisTraceRecorder;
+import com.superbiz.agent.harness.audit.TraceAuditEvents;
import com.superbiz.agent.harness.contract.PreviousTurn;
import com.superbiz.agent.harness.contract.PublishedResult;
import com.superbiz.agent.harness.contract.ReleaseOutcome;
@@ -25,13 +28,22 @@ public final class DiagnosisChatExecutor implements DiagnosisOperation {
private final DiagnosisAgentUseCase diagnosisAgent;
private final DiagnosisReleaseUseCase releaseUseCase;
private final PublishedResultPolicy publishedPolicy;
+ private final DiagnosisTraceRecorder traceRecorder;
public DiagnosisChatExecutor(DiagnosisAgentUseCase diagnosisAgent,
DiagnosisReleaseUseCase releaseUseCase,
PublishedResultPolicy publishedPolicy) {
+ this(diagnosisAgent, releaseUseCase, publishedPolicy, DiagnosisTraceRecorder.noop());
+ }
+
+ public DiagnosisChatExecutor(DiagnosisAgentUseCase diagnosisAgent,
+ DiagnosisReleaseUseCase releaseUseCase,
+ PublishedResultPolicy publishedPolicy,
+ DiagnosisTraceRecorder traceRecorder) {
this.diagnosisAgent = Objects.requireNonNull(diagnosisAgent, "diagnosisAgent must not be null");
this.releaseUseCase = Objects.requireNonNull(releaseUseCase, "releaseUseCase must not be null");
this.publishedPolicy = Objects.requireNonNull(publishedPolicy, "publishedPolicy must not be null");
+ this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
}
@Override
@@ -41,15 +53,22 @@ public final class DiagnosisChatExecutor implements DiagnosisOperation {
Consumer statusSink) {
Objects.requireNonNull(statusSink, "statusSink must not be null");
statusSink.accept(ChatApplicationStatus.DIAGNOSIS_RUNNING);
- DiagnosisDraft draft = diagnosisAgent.execute(
- context, new DiagnosisAgentInput(query, previousTurn));
+ DiagnosisAgentExecution execution;
+ try {
+ execution = diagnosisAgent.execute(
+ context, new DiagnosisAgentInput(query, previousTurn));
+ } catch (DiagnosisAgentOutputException exception) {
+ return recoverInvalidDraft(context, exception, statusSink);
+ }
statusSink.accept(ChatApplicationStatus.SAFETY_VALIDATING);
- DiagnosisReleaseResult released = releaseUseCase.execute(context, query, draft);
+ DiagnosisReleaseResult released = releaseUseCase.execute(context, query, execution);
if (released.outcome() == ReleaseOutcome.FALLBACK) {
return new DiagnosisExecutionResult(
ReleaseOutcome.FALLBACK,
new FallbackContent(released.fallback()),
- null);
+ null,
+ execution.stopReason()
+ == com.superbiz.agent.harness.progress.DiagnosisStopReason.BUDGET_LIMIT_REACHED);
}
PublishedResult published = publishedPolicy.create(
query, released.draft(), released.verifiedEvidence()).orElse(null);
@@ -60,4 +79,26 @@ public final class DiagnosisChatExecutor implements DiagnosisOperation {
released.verifiedEvidence().verifiedSources()),
published);
}
+
+ private DiagnosisExecutionResult recoverInvalidDraft(
+ RunContext context,
+ DiagnosisAgentOutputException exception,
+ Consumer statusSink) {
+ if (!exception.isDraftContractFailure()) {
+ throw exception;
+ }
+ boolean hasProgress = exception.progress().hasObservedFacts();
+ traceRecorder.record(TraceAuditEvents.agentDraftInvalid(
+ context, exception.kind(), exception.outputBytes(), hasProgress));
+ if (!hasProgress) {
+ throw exception;
+ }
+ statusSink.accept(ChatApplicationStatus.SAFETY_VALIDATING);
+ DiagnosisReleaseResult released = releaseUseCase.releaseInvalidDraft(
+ context, exception.progress());
+ return new DiagnosisExecutionResult(
+ ReleaseOutcome.FALLBACK,
+ new FallbackContent(released.fallback()),
+ null);
+ }
}
diff --git a/src/main/java/com/superbiz/agent/harness/application/executor/KnowledgeQueryExecutor.java b/src/main/java/com/superbiz/agent/harness/application/executor/KnowledgeQueryExecutor.java
index f3be88c..1675ba4 100644
--- a/src/main/java/com/superbiz/agent/harness/application/executor/KnowledgeQueryExecutor.java
+++ b/src/main/java/com/superbiz/agent/harness/application/executor/KnowledgeQueryExecutor.java
@@ -16,6 +16,7 @@ import com.superbiz.agent.harness.contract.SourceDocument;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.guard.semantic.GuardModelCall;
+import com.superbiz.agent.harness.audit.ModelCallComponent;
import com.superbiz.agent.harness.tool.boundary.ToolBoundaryResult;
import com.superbiz.agent.harness.tool.contract.AgentToolContracts;
import com.superbiz.agent.harness.tool.contract.RagEvidence;
@@ -106,7 +107,7 @@ public final class KnowledgeQueryExecutor implements KnowledgeQueryOperation {
}
core.reserveRunBytes(context, bytes);
String output = modelCall.call(
- context,
+ context, ModelCallComponent.KNOWLEDGE_ANSWER,
new Prompt(List.of(new SystemMessage(prompt), new UserMessage(modelInput))),
limits.modelTimeout(),
limits.maxModelOutputBytes());
diff --git a/src/main/java/com/superbiz/agent/harness/application/executor/SystemChatExecutor.java b/src/main/java/com/superbiz/agent/harness/application/executor/SystemChatExecutor.java
index ce5a79e..aeacd11 100644
--- a/src/main/java/com/superbiz/agent/harness/application/executor/SystemChatExecutor.java
+++ b/src/main/java/com/superbiz/agent/harness/application/executor/SystemChatExecutor.java
@@ -7,6 +7,7 @@ import com.superbiz.agent.harness.application.SystemChatOperation;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.guard.semantic.GuardModelCall;
+import com.superbiz.agent.harness.audit.ModelCallComponent;
import org.springframework.ai.chat.messages.SystemMessage;
import org.springframework.ai.chat.messages.UserMessage;
import org.springframework.ai.chat.prompt.Prompt;
@@ -47,7 +48,7 @@ public final class SystemChatExecutor implements SystemChatOperation {
}
core.reserveRunBytes(context, bytes);
String answer = modelCall.call(
- context,
+ context, ModelCallComponent.SYSTEM_CHAT,
new Prompt(List.of(new SystemMessage(SYSTEM_PROMPT), new UserMessage(query))),
limits.timeout(),
limits.maxOutputBytes());
diff --git a/src/main/java/com/superbiz/agent/harness/application/persistence/JpaChatRunStore.java b/src/main/java/com/superbiz/agent/harness/application/persistence/JpaChatRunStore.java
index 237af32..103ae4b 100644
--- a/src/main/java/com/superbiz/agent/harness/application/persistence/JpaChatRunStore.java
+++ b/src/main/java/com/superbiz/agent/harness/application/persistence/JpaChatRunStore.java
@@ -12,6 +12,7 @@ import com.superbiz.agent.harness.contract.PublishedResult;
import com.superbiz.agent.harness.contract.ReleaseOutcome;
import com.superbiz.agent.harness.core.RunBudgetUsage;
import com.superbiz.agent.harness.core.RunContext;
+import com.superbiz.agent.harness.audit.ModelCallComponent;
import com.superbiz.agent.repository.ChatSessionRepository;
import com.superbiz.agent.repository.DiagnosisRunRepository;
import org.springframework.stereotype.Component;
@@ -119,6 +120,8 @@ public class JpaChatRunStore implements ChatRunStore {
run.setTotalDurationMs(Math.max(0, durationMs));
RunBudgetUsage usage = context.budget().snapshot();
run.setTotalTokenCount(saturatingInt(usage.totalTokens()));
+ run.setStepCount(context.modelCalls().componentCallCount(
+ ModelCallComponent.DIAGNOSIS_AGENT));
run.setToolCallCount(saturatingInt(usage.toolCalls()));
runs.save(run);
if (outcome == ReleaseOutcome.SUCCESS || outcome == ReleaseOutcome.FALLBACK) {
diff --git a/src/main/java/com/superbiz/agent/harness/application/routing/IntentRouter.java b/src/main/java/com/superbiz/agent/harness/application/routing/IntentRouter.java
index 00c0a98..da8663f 100644
--- a/src/main/java/com/superbiz/agent/harness/application/routing/IntentRouter.java
+++ b/src/main/java/com/superbiz/agent/harness/application/routing/IntentRouter.java
@@ -6,6 +6,7 @@ import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.harness.contract.IntentType;
import com.superbiz.agent.harness.audit.DiagnosisTraceRecorder;
import com.superbiz.agent.harness.audit.TraceAuditEvents;
+import com.superbiz.agent.harness.audit.ModelCallComponent;
import com.superbiz.agent.harness.application.IntentRouting;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunContext;
@@ -86,7 +87,8 @@ public final class IntentRouter implements IntentRouting {
context,
context.retryPolicies().intentRouter(),
() -> parse(modelCall.call(
- context, prompt, remaining(started), limits.maxOutputBytes())),
+ context, ModelCallComponent.INTENT_ROUTER,
+ prompt, remaining(started), limits.maxOutputBytes())),
this::classify,
attempt -> {
attemptRecorder.accept(attempt);
diff --git a/src/main/java/com/superbiz/agent/harness/audit/ModelCallAuditor.java b/src/main/java/com/superbiz/agent/harness/audit/ModelCallAuditor.java
new file mode 100644
index 0000000..5da6187
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/audit/ModelCallAuditor.java
@@ -0,0 +1,85 @@
+package com.superbiz.agent.harness.audit;
+
+import com.superbiz.agent.domain.entity.AgentStep;
+import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
+import com.superbiz.agent.harness.core.RunContext;
+import com.superbiz.agent.repository.AgentStepRepository;
+import org.slf4j.Logger;
+import org.slf4j.LoggerFactory;
+
+import java.util.Objects;
+
+public final class ModelCallAuditor {
+
+ private static final Logger log = LoggerFactory.getLogger(ModelCallAuditor.class);
+
+ private final DiagnosisHarnessCore core;
+ private final DiagnosisTraceRecorder traceRecorder;
+ private final AgentStepRepository agentSteps;
+
+ public ModelCallAuditor(DiagnosisHarnessCore core) {
+ this(core, DiagnosisTraceRecorder.noop(), null);
+ }
+
+ public ModelCallAuditor(DiagnosisHarnessCore core,
+ DiagnosisTraceRecorder traceRecorder,
+ AgentStepRepository agentSteps) {
+ this.core = Objects.requireNonNull(core, "core must not be null");
+ this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
+ this.agentSteps = agentSteps;
+ }
+
+ public ModelCallLedger.Call begin(RunContext context, ModelCallComponent component) {
+ Objects.requireNonNull(context, "context must not be null");
+ return context.modelCalls().begin(component);
+ }
+
+ public void recordUsage(RunContext context, ModelCallLedger.Call call,
+ long inputTokens, long outputTokens, boolean usageAvailable) {
+ Objects.requireNonNull(context, "context must not be null");
+ if (!context.modelCalls().record(call, inputTokens, outputTokens, usageAvailable)) {
+ return;
+ }
+ try {
+ if (usageAvailable) {
+ core.recordTokens(context, inputTokens, outputTokens);
+ }
+ } finally {
+ traceRecorder.record(TraceAuditEvents.modelTokenUsage(
+ context, call, inputTokens, outputTokens, usageAvailable));
+ updateAgentStep(context, call, inputTokens, outputTokens, usageAvailable);
+ }
+ }
+
+ private void updateAgentStep(RunContext context, ModelCallLedger.Call call,
+ long inputTokens, long outputTokens, boolean usageAvailable) {
+ if (!usageAvailable || agentSteps == null
+ || call.component() != ModelCallComponent.DIAGNOSIS_AGENT) {
+ return;
+ }
+ try {
+ agentSteps.findByRunIdAndStepIndex(context.runId(), call.componentRound() - 1)
+ .ifPresent(step -> saveTokens(step, inputTokens, outputTokens));
+ } catch (RuntimeException exception) {
+ log.warn("Failed to persist AgentStep token audit: runId={}, round={}",
+ context.runId(), call.componentRound());
+ }
+ }
+
+ private void saveTokens(AgentStep step, long inputTokens, long outputTokens) {
+ step.setTokenCount(saturatingInt(safeAdd(inputTokens, outputTokens)));
+ agentSteps.save(step);
+ }
+
+ private static long safeAdd(long left, long right) {
+ try {
+ return Math.addExact(left, right);
+ } catch (ArithmeticException exception) {
+ return Long.MAX_VALUE;
+ }
+ }
+
+ private static int saturatingInt(long value) {
+ return value >= Integer.MAX_VALUE ? Integer.MAX_VALUE : (int) Math.max(0, value);
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/audit/ModelCallComponent.java b/src/main/java/com/superbiz/agent/harness/audit/ModelCallComponent.java
new file mode 100644
index 0000000..4f9f5fb
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/audit/ModelCallComponent.java
@@ -0,0 +1,20 @@
+package com.superbiz.agent.harness.audit;
+
+public enum ModelCallComponent {
+ INTENT_ROUTER(TracePhase.ROUTING),
+ SYSTEM_CHAT(TracePhase.AGENT),
+ KNOWLEDGE_ANSWER(TracePhase.AGENT),
+ DIAGNOSIS_AGENT(TracePhase.AGENT),
+ EVIDENCE_REPAIR(TracePhase.EVIDENCE),
+ SEMANTIC_GUARD(TracePhase.SEMANTIC);
+
+ private final TracePhase tracePhase;
+
+ ModelCallComponent(TracePhase tracePhase) {
+ this.tracePhase = tracePhase;
+ }
+
+ public TracePhase tracePhase() {
+ return tracePhase;
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/audit/ModelCallLedger.java b/src/main/java/com/superbiz/agent/harness/audit/ModelCallLedger.java
new file mode 100644
index 0000000..b502938
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/audit/ModelCallLedger.java
@@ -0,0 +1,86 @@
+package com.superbiz.agent.harness.audit;
+
+import java.util.EnumMap;
+import java.util.HashSet;
+import java.util.Map;
+import java.util.Objects;
+import java.util.Set;
+
+public final class ModelCallLedger {
+
+ private final Map componentRounds =
+ new EnumMap<>(ModelCallComponent.class);
+ private final Set auditedCalls = new HashSet<>();
+
+ private int startedCallCount;
+ private int usageUnavailableCount;
+ private long inputTokens;
+ private long outputTokens;
+ private long totalTokens;
+
+ public synchronized Call begin(ModelCallComponent component) {
+ Objects.requireNonNull(component, "component must not be null");
+ int round = componentRounds.merge(component, 1, Integer::sum);
+ startedCallCount++;
+ return new Call(component, round);
+ }
+
+ public synchronized boolean record(Call call, long input, long output, boolean usageAvailable) {
+ Objects.requireNonNull(call, "call must not be null");
+ if (input < 0 || output < 0) {
+ throw new IllegalArgumentException("token counts must not be negative");
+ }
+ if (!auditedCalls.add(call)) {
+ return false;
+ }
+ if (!usageAvailable) {
+ usageUnavailableCount++;
+ return true;
+ }
+ inputTokens = safeAdd(inputTokens, input);
+ outputTokens = safeAdd(outputTokens, output);
+ totalTokens = safeAdd(totalTokens, safeAdd(input, output));
+ return true;
+ }
+
+ public synchronized int componentCallCount(ModelCallComponent component) {
+ Objects.requireNonNull(component, "component must not be null");
+ return componentRounds.getOrDefault(component, 0);
+ }
+
+ public synchronized Snapshot snapshot() {
+ return new Snapshot(
+ startedCallCount,
+ auditedCalls.size(),
+ usageUnavailableCount,
+ inputTokens,
+ outputTokens,
+ totalTokens);
+ }
+
+ private static long safeAdd(long left, long right) {
+ try {
+ return Math.addExact(left, right);
+ } catch (ArithmeticException exception) {
+ return Long.MAX_VALUE;
+ }
+ }
+
+ public record Call(ModelCallComponent component, int componentRound) {
+ public Call {
+ Objects.requireNonNull(component, "component must not be null");
+ if (componentRound < 1) {
+ throw new IllegalArgumentException("componentRound must be positive");
+ }
+ }
+ }
+
+ public record Snapshot(
+ int startedCallCount,
+ int auditedCallCount,
+ int usageUnavailableCount,
+ long inputTokens,
+ long outputTokens,
+ long totalTokens) {
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/audit/TraceAuditEvents.java b/src/main/java/com/superbiz/agent/harness/audit/TraceAuditEvents.java
index 39ba100..b1cee31 100644
--- a/src/main/java/com/superbiz/agent/harness/audit/TraceAuditEvents.java
+++ b/src/main/java/com/superbiz/agent/harness/audit/TraceAuditEvents.java
@@ -1,6 +1,7 @@
package com.superbiz.agent.harness.audit;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
+import com.superbiz.agent.harness.agent.DiagnosisAgentOutputException;
import com.superbiz.agent.harness.contract.FallbackType;
import com.superbiz.agent.harness.contract.IntentType;
import com.superbiz.agent.harness.contract.ReleaseOutcome;
@@ -8,6 +9,8 @@ import com.superbiz.agent.harness.contract.SemanticVerdict;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.guard.evidence.EvidenceGuardResult;
import com.superbiz.agent.harness.guard.evidence.EvidenceViolation;
+import com.superbiz.agent.harness.progress.DiagnosisProgressSnapshotState;
+import com.superbiz.agent.harness.progress.InformationGain;
import com.superbiz.agent.harness.retry.RetryAttempt;
import java.util.ArrayList;
@@ -33,11 +36,27 @@ public final class TraceAuditEvents {
public static DiagnosisTraceAuditEvent runFinished(
RunContext context, IntentType intent, ReleaseOutcome outcome, int durationMs) {
+ var budget = context.budget().snapshot();
+ ModelCallLedger.Snapshot modelCalls = context.modelCalls().snapshot();
Map details = new LinkedHashMap<>();
if (intent != null) {
details.put("intent", intent.name());
}
details.put("release_outcome", outcome.name());
+ details.put("model_call_count", budget.modelCalls());
+ details.put("run_input_tokens", budget.inputTokens());
+ details.put("run_output_tokens", budget.outputTokens());
+ details.put("run_total_tokens", budget.totalTokens());
+ details.put("audited_model_call_count", modelCalls.auditedCallCount());
+ details.put("audited_input_tokens", modelCalls.inputTokens());
+ details.put("audited_output_tokens", modelCalls.outputTokens());
+ details.put("audited_total_tokens", modelCalls.totalTokens());
+ details.put("usage_unavailable_count", modelCalls.usageUnavailableCount());
+ details.put("tokens_reconciled", budget.modelCalls() == modelCalls.auditedCallCount()
+ && modelCalls.usageUnavailableCount() == 0
+ && budget.inputTokens() == modelCalls.inputTokens()
+ && budget.outputTokens() == modelCalls.outputTokens()
+ && budget.totalTokens() == modelCalls.totalTokens());
return event(context, TracePhase.RUN, TraceEventType.RUN_FINISHED,
terminalStatus(outcome), null, durationMs, details);
}
@@ -66,6 +85,41 @@ public final class TraceAuditEvents {
TraceEventStatus.SUCCEEDED, stepIndex + 1, durationMs, details);
}
+ public static DiagnosisTraceAuditEvent modelTokenUsage(
+ RunContext context,
+ ModelCallLedger.Call call,
+ long inputTokens,
+ long outputTokens,
+ boolean usageAvailable) {
+ Map details = new LinkedHashMap<>();
+ details.put("component", call.component().name());
+ details.put("component_round", call.componentRound());
+ details.put("usage_available", usageAvailable);
+ if (usageAvailable) {
+ details.put("input_tokens", inputTokens);
+ details.put("output_tokens", outputTokens);
+ details.put("total_tokens", safeAdd(inputTokens, outputTokens));
+ }
+ return new DiagnosisTraceAuditEvent(
+ context.sessionId(), context.runId(), call.component().tracePhase(),
+ TraceEventType.MODEL_TOKEN_USAGE,
+ usageAvailable ? TraceEventStatus.SUCCEEDED : TraceEventStatus.UNAVAILABLE,
+ call.componentRound(), null, details);
+ }
+
+ public static DiagnosisTraceAuditEvent agentDraftInvalid(
+ RunContext context,
+ DiagnosisAgentOutputException.Kind kind,
+ long outputBytes,
+ boolean hasPublishableProgress) {
+ Map details = new LinkedHashMap<>();
+ details.put("failure_kind", kind.name());
+ details.put("output_bytes", outputBytes);
+ details.put("has_publishable_progress", hasPublishableProgress);
+ return event(context, TracePhase.AGENT, TraceEventType.AGENT_DRAFT_INVALID,
+ TraceEventStatus.REJECTED, null, null, details);
+ }
+
public static DiagnosisTraceAuditEvent toolInvocation(ToolInvocationAuditEvent tool) {
Map details = new LinkedHashMap<>();
details.put("tool_call_id", tool.toolCallId());
@@ -84,6 +138,55 @@ public final class TraceAuditEvents {
null, tool.durationMs(), details);
}
+ public static DiagnosisTraceAuditEvent toolRequestRejected(
+ RunContext context, String toolCallId, String toolName, String errorCode) {
+ Map details = new LinkedHashMap<>();
+ details.put("tool_call_id", safeIdentifier(toolCallId));
+ details.put("tool_name", safeIdentifier(toolName));
+ details.put("error_code", safeIdentifier(errorCode));
+ return event(context, TracePhase.TOOL, TraceEventType.TOOL_REQUEST_REJECTED,
+ TraceEventStatus.REJECTED, null, null, details);
+ }
+
+ public static DiagnosisTraceAuditEvent toolProgress(
+ RunContext context,
+ String toolCallId,
+ String toolName,
+ String scopeSummary,
+ InformationGain informationGain,
+ String producer,
+ DiagnosisProgressSnapshotState state) {
+ Map details = new LinkedHashMap<>();
+ details.put("tool_call_id", toolCallId);
+ details.put("tool_name", toolName);
+ details.put("scope_summary", scopeSummary);
+ details.put("information_gain", informationGain.name());
+ details.put("producer", producer);
+ details.put("consecutive_no_gain", state.consecutiveNoGain());
+ details.put("collection_state", state.collectionState().name());
+ if (state.stopReason() != null) {
+ details.put("stop_reason", state.stopReason().name());
+ }
+ return event(context, TracePhase.TOOL, TraceEventType.TOOL_PROGRESS,
+ TraceEventStatus.SUCCEEDED, null, null, details);
+ }
+
+ public static DiagnosisTraceAuditEvent collectionStop(
+ RunContext context,
+ String toolCallId,
+ String toolName,
+ DiagnosisProgressSnapshotState state) {
+ Map details = new LinkedHashMap<>();
+ details.put("tool_call_id", toolCallId);
+ details.put("tool_name", toolName);
+ details.put("collection_state", state.collectionState().name());
+ if (state.stopReason() != null) {
+ details.put("stop_reason", state.stopReason().name());
+ }
+ return event(context, TracePhase.TOOL, TraceEventType.COLLECTION_STOP,
+ TraceEventStatus.SUCCEEDED, null, null, details);
+ }
+
public static DiagnosisTraceAuditEvent evidenceValidation(
RunContext context, TraceEventType type,
EvidenceGuardResult result, DiagnosisDraft draft) {
@@ -198,6 +301,18 @@ public final class TraceAuditEvents {
return new ToolReferences(List.copyOf(safeIds), invalid);
}
+ private static String safeIdentifier(String value) {
+ return value != null && SAFE_TOOL_CALL_ID.matcher(value).matches() ? value : "INVALID";
+ }
+
+ private static long safeAdd(long left, long right) {
+ try {
+ return Math.addExact(left, right);
+ } catch (ArithmeticException exception) {
+ return Long.MAX_VALUE;
+ }
+ }
+
private record ToolReferences(List safeIds, int invalidCount) {
}
}
diff --git a/src/main/java/com/superbiz/agent/harness/audit/TraceEventType.java b/src/main/java/com/superbiz/agent/harness/audit/TraceEventType.java
index 129e393..a03beca 100644
--- a/src/main/java/com/superbiz/agent/harness/audit/TraceEventType.java
+++ b/src/main/java/com/superbiz/agent/harness/audit/TraceEventType.java
@@ -4,8 +4,13 @@ public enum TraceEventType {
RUN_STARTED,
ROUTING_ATTEMPT,
ROUTING_DECISION,
+ MODEL_TOKEN_USAGE,
AGENT_MODEL_STEP,
+ AGENT_DRAFT_INVALID,
TOOL_INVOCATION,
+ TOOL_REQUEST_REJECTED,
+ TOOL_PROGRESS,
+ COLLECTION_STOP,
EVIDENCE_GUARD_INITIAL,
EVIDENCE_REPAIR_ATTEMPT,
EVIDENCE_GUARD_RECHECK,
diff --git a/src/main/java/com/superbiz/agent/harness/contract/FallbackType.java b/src/main/java/com/superbiz/agent/harness/contract/FallbackType.java
index d382d1c..79d29d8 100644
--- a/src/main/java/com/superbiz/agent/harness/contract/FallbackType.java
+++ b/src/main/java/com/superbiz/agent/harness/contract/FallbackType.java
@@ -3,5 +3,8 @@ package com.superbiz.agent.harness.contract;
public enum FallbackType {
EVIDENCE_VALIDATION_FAILED,
SEMANTIC_UNSUPPORTED,
- SEMANTIC_UNAVAILABLE
+ SEMANTIC_UNAVAILABLE,
+ BUDGET_EXHAUSTED,
+ INSUFFICIENT_EVIDENCE,
+ MISSING_REQUIRED_CONTEXT
}
diff --git a/src/main/java/com/superbiz/agent/harness/core/DiagnosisHarnessCore.java b/src/main/java/com/superbiz/agent/harness/core/DiagnosisHarnessCore.java
index 29d00fb..7ee5305 100644
--- a/src/main/java/com/superbiz/agent/harness/core/DiagnosisHarnessCore.java
+++ b/src/main/java/com/superbiz/agent/harness/core/DiagnosisHarnessCore.java
@@ -1,6 +1,8 @@
package com.superbiz.agent.harness.core;
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
+import com.superbiz.agent.harness.progress.DiagnosisProgressTracker;
+import com.superbiz.agent.harness.audit.ModelCallLedger;
import java.time.Clock;
import java.time.Duration;
@@ -15,17 +17,31 @@ public final class DiagnosisHarnessCore {
private final Duration maxRunDuration;
private final RunBudgetLimits budgetLimits;
private final HarnessRetryPolicies retryPolicies;
+ private final int stopAfterConsecutiveNoGain;
public DiagnosisHarnessCore(Clock clock,
Supplier runIdSupplier,
Duration maxRunDuration,
RunBudgetLimits budgetLimits,
HarnessRetryPolicies retryPolicies) {
+ this(clock, runIdSupplier, maxRunDuration, budgetLimits, retryPolicies, 2);
+ }
+
+ public DiagnosisHarnessCore(Clock clock,
+ Supplier runIdSupplier,
+ Duration maxRunDuration,
+ RunBudgetLimits budgetLimits,
+ HarnessRetryPolicies retryPolicies,
+ int stopAfterConsecutiveNoGain) {
this.clock = Objects.requireNonNull(clock, "clock must not be null");
this.runIdSupplier = Objects.requireNonNull(runIdSupplier, "runIdSupplier must not be null");
this.maxRunDuration = requirePositive(maxRunDuration, "maxRunDuration");
this.budgetLimits = Objects.requireNonNull(budgetLimits, "budgetLimits must not be null");
this.retryPolicies = Objects.requireNonNull(retryPolicies, "retryPolicies must not be null");
+ if (stopAfterConsecutiveNoGain <= 0) {
+ throw new IllegalArgumentException("stopAfterConsecutiveNoGain must be positive");
+ }
+ this.stopAfterConsecutiveNoGain = stopAfterConsecutiveNoGain;
}
public RunContext startRun(String sessionId) {
@@ -42,8 +58,10 @@ public final class DiagnosisHarnessCore {
deadline,
cancellation,
new RunBudget(budgetLimits),
+ new ModelCallLedger(),
retryPolicies,
- lifecycle);
+ lifecycle,
+ new DiagnosisProgressTracker(stopAfterConsecutiveNoGain));
cancellation.onCancel(reason -> lifecycle.finish(terminalState(reason), reason.name()));
return context;
}
diff --git a/src/main/java/com/superbiz/agent/harness/core/RunContext.java b/src/main/java/com/superbiz/agent/harness/core/RunContext.java
index 0579f36..d175d22 100644
--- a/src/main/java/com/superbiz/agent/harness/core/RunContext.java
+++ b/src/main/java/com/superbiz/agent/harness/core/RunContext.java
@@ -1,6 +1,8 @@
package com.superbiz.agent.harness.core;
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
+import com.superbiz.agent.harness.progress.DiagnosisProgressTracker;
+import com.superbiz.agent.harness.audit.ModelCallLedger;
import java.time.Instant;
import java.util.Objects;
@@ -14,8 +16,10 @@ public record RunContext(
Instant deadline,
RunCancellation cancellation,
RunBudget budget,
+ ModelCallLedger modelCalls,
HarnessRetryPolicies retryPolicies,
- RunLifecycle lifecycle) {
+ RunLifecycle lifecycle,
+ DiagnosisProgressTracker progress) {
public RunContext {
requireText(sessionId, "sessionId");
@@ -23,8 +27,10 @@ public record RunContext(
Objects.requireNonNull(deadline, "deadline must not be null");
Objects.requireNonNull(cancellation, "cancellation must not be null");
Objects.requireNonNull(budget, "budget must not be null");
+ Objects.requireNonNull(modelCalls, "modelCalls must not be null");
Objects.requireNonNull(retryPolicies, "retryPolicies must not be null");
Objects.requireNonNull(lifecycle, "lifecycle must not be null");
+ Objects.requireNonNull(progress, "progress must not be null");
}
private static void requireText(String value, String name) {
diff --git a/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuard.java b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuard.java
index 3becc55..c72cbc6 100644
--- a/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuard.java
+++ b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuard.java
@@ -79,6 +79,40 @@ public final class EvidenceGuard {
: EvidenceGuardResult.invalid(violations);
}
+ public EvidenceGuardResult validateNoConclusionReferences(
+ RunContext context, DiagnosisDraft draft) {
+ Objects.requireNonNull(context, "context must not be null");
+ if (draft == null) {
+ return EvidenceGuardResult.invalid(List.of(
+ violation(EvidenceViolationCode.DRAFT_MISSING, "draft")));
+ }
+ if (draft.conclusion() != null) {
+ throw new IllegalArgumentException("no-conclusion validation requires conclusion=null");
+ }
+
+ List violations = new ArrayList<>();
+ for (int index = 0; index < draft.analysis().size(); index++) {
+ DiagnosisDraft.AnalysisItem analysis = draft.analysis().get(index);
+ if (analysis == null || analysis.toolCallIds().isEmpty()) {
+ continue;
+ }
+ if (analysis.kind() == null) {
+ violations.add(violation(
+ EvidenceViolationCode.ANALYSIS_KIND_MISSING,
+ "analysis[" + index + "].kind"));
+ continue;
+ }
+ List ignoredEvidence = new ArrayList<>();
+ for (String toolCallId : analysis.toolCallIds()) {
+ verifyInvocation(
+ context, analysis, index, toolCallId, ignoredEvidence, violations);
+ }
+ }
+ return violations.isEmpty()
+ ? EvidenceGuardResult.valid(VerifiedEvidenceSnapshot.empty())
+ : EvidenceGuardResult.invalid(violations);
+ }
+
private List validateDraft(DiagnosisDraft draft) {
List violations = new ArrayList<>();
if (draft == null) {
diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCall.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCall.java
index 0b1d2d8..a4e876b 100644
--- a/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCall.java
+++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCall.java
@@ -4,6 +4,9 @@ import com.superbiz.agent.harness.core.BudgetExceededException;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunAbortedException;
import com.superbiz.agent.harness.core.RunContext;
+import com.superbiz.agent.harness.audit.ModelCallAuditor;
+import com.superbiz.agent.harness.audit.ModelCallComponent;
+import com.superbiz.agent.harness.audit.ModelCallLedger;
import com.superbiz.agent.harness.retry.RetryFailure;
import org.springframework.ai.chat.messages.AssistantMessage;
import org.springframework.ai.chat.metadata.Usage;
@@ -26,15 +29,23 @@ public final class GuardModelCall {
private final DiagnosisHarnessCore core;
private final ChatModel chatModel;
private final ExecutorService executor;
+ private final ModelCallAuditor auditor;
public GuardModelCall(DiagnosisHarnessCore core, ChatModel chatModel,
ExecutorService executor) {
+ this(core, chatModel, executor, new ModelCallAuditor(core));
+ }
+
+ public GuardModelCall(DiagnosisHarnessCore core, ChatModel chatModel,
+ ExecutorService executor, ModelCallAuditor auditor) {
this.core = Objects.requireNonNull(core, "core must not be null");
this.chatModel = Objects.requireNonNull(chatModel, "chatModel must not be null");
this.executor = Objects.requireNonNull(executor, "executor must not be null");
+ this.auditor = Objects.requireNonNull(auditor, "auditor must not be null");
}
- public String call(RunContext context, Prompt prompt, Duration timeout, long maxOutputBytes) {
+ public String call(RunContext context, ModelCallComponent component,
+ Prompt prompt, Duration timeout, long maxOutputBytes) {
Objects.requireNonNull(context, "context must not be null");
Objects.requireNonNull(prompt, "prompt must not be null");
Objects.requireNonNull(timeout, "timeout must not be null");
@@ -42,20 +53,24 @@ public final class GuardModelCall {
throw new IllegalArgumentException("timeout and output limit must be positive");
}
core.beforeModelCall(context);
- Future future = executor.submit(() -> invoke(context, prompt, maxOutputBytes));
+ ModelCallLedger.Call call = auditor.begin(context, component);
+ Future future = executor.submit(() -> invoke(context, call, prompt, maxOutputBytes));
context.cancellation().onCancel(ignored -> future.cancel(true));
try {
return future.get(timeout.toNanos(), TimeUnit.NANOSECONDS);
} catch (TimeoutException exception) {
future.cancel(true);
+ auditor.recordUsage(context, call, 0, 0, false);
throw new GuardModelCallException(
RetryFailure.TIMEOUT, "Guard model attempt timed out", exception);
} catch (CancellationException exception) {
+ auditor.recordUsage(context, call, 0, 0, false);
core.checkActive(context);
throw new GuardModelCallException(
RetryFailure.TRANSPORT, "Guard model attempt was cancelled", exception);
} catch (InterruptedException exception) {
future.cancel(true);
+ auditor.recordUsage(context, call, 0, 0, false);
Thread.currentThread().interrupt();
core.checkActive(context);
throw new GuardModelCallException(
@@ -76,15 +91,17 @@ public final class GuardModelCall {
}
}
- private String invoke(RunContext context, Prompt prompt, long maxOutputBytes) {
+ private String invoke(RunContext context, ModelCallLedger.Call call,
+ Prompt prompt, long maxOutputBytes) {
ChatResponse response;
try {
response = chatModel.call(prompt);
} catch (RuntimeException exception) {
+ auditor.recordUsage(context, call, 0, 0, false);
throw new GuardModelCallException(
RetryFailure.TRANSPORT, "Guard model transport failed", exception);
}
- recordUsage(context, response);
+ recordUsage(context, call, response);
core.checkActive(context);
AssistantMessage output = response == null || response.getResult() == null
? null : response.getResult().getOutput();
@@ -102,16 +119,18 @@ public final class GuardModelCall {
return output.getText();
}
- private void recordUsage(RunContext context, ChatResponse response) {
+ private void recordUsage(RunContext context, ModelCallLedger.Call call, ChatResponse response) {
if (response == null || response.getMetadata() == null) {
+ auditor.recordUsage(context, call, 0, 0, false);
return;
}
Usage usage = response.getMetadata().getUsage();
if (usage == null) {
+ auditor.recordUsage(context, call, 0, 0, false);
return;
}
- core.recordTokens(context, nonNegative(usage.getPromptTokens()),
- nonNegative(usage.getCompletionTokens()));
+ auditor.recordUsage(context, call, nonNegative(usage.getPromptTokens()),
+ nonNegative(usage.getCompletionTokens()), true);
}
private static long nonNegative(Integer value) {
diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuard.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuard.java
index f684004..4dec111 100644
--- a/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuard.java
+++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuard.java
@@ -6,6 +6,7 @@ import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.harness.contract.SemanticVerdict;
import com.superbiz.agent.harness.audit.DiagnosisTraceRecorder;
import com.superbiz.agent.harness.audit.TraceAuditEvents;
+import com.superbiz.agent.harness.audit.ModelCallComponent;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.retry.HarnessRetryExecutor;
@@ -81,7 +82,8 @@ public final class SemanticGuard {
context,
context.retryPolicies().semanticGuard(),
() -> parse(modelCall.call(
- context, modelPrompt, remainingTimeout(startedNanos), limits.maxOutputBytes())),
+ context, ModelCallComponent.SEMANTIC_GUARD,
+ modelPrompt, remainingTimeout(startedNanos), limits.maxOutputBytes())),
this::classify,
attempt -> {
attemptRecorder.accept(attempt);
diff --git a/src/main/java/com/superbiz/agent/harness/progress/CompletedToolCall.java b/src/main/java/com/superbiz/agent/harness/progress/CompletedToolCall.java
new file mode 100644
index 0000000..79e6ddd
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/CompletedToolCall.java
@@ -0,0 +1,19 @@
+package com.superbiz.agent.harness.progress;
+
+public record CompletedToolCall(
+ String toolCallId,
+ String toolName,
+ String normalizedScope) {
+
+ public CompletedToolCall {
+ requireText(toolCallId, "toolCallId");
+ requireText(toolName, "toolName");
+ requireText(normalizedScope, "normalizedScope");
+ }
+
+ private static void requireText(String value, String name) {
+ if (value == null || value.isBlank()) {
+ throw new IllegalArgumentException(name + " must not be blank");
+ }
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/DiagnosisCollectionState.java b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisCollectionState.java
new file mode 100644
index 0000000..9a903b9
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisCollectionState.java
@@ -0,0 +1,6 @@
+package com.superbiz.agent.harness.progress;
+
+public enum DiagnosisCollectionState {
+ COLLECTING,
+ SATURATED
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressProjection.java b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressProjection.java
new file mode 100644
index 0000000..2e94883
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressProjection.java
@@ -0,0 +1,13 @@
+package com.superbiz.agent.harness.progress;
+
+import com.superbiz.agent.harness.core.RunContext;
+
+@FunctionalInterface
+public interface DiagnosisProgressProjection {
+
+ DiagnosisProgressSnapshot project(RunContext context);
+
+ static DiagnosisProgressProjection empty() {
+ return ignored -> DiagnosisProgressSnapshot.empty();
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressProjector.java b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressProjector.java
new file mode 100644
index 0000000..b3cccd9
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressProjector.java
@@ -0,0 +1,228 @@
+package com.superbiz.agent.harness.progress;
+
+import com.fasterxml.jackson.core.JsonProcessingException;
+import com.fasterxml.jackson.databind.JsonNode;
+import com.fasterxml.jackson.databind.ObjectMapper;
+import com.superbiz.agent.harness.contract.SafeFallback;
+import com.superbiz.agent.harness.core.RunContext;
+import com.superbiz.agent.harness.tool.contract.AgentToolContracts;
+import com.superbiz.agent.harness.tool.store.CanonicalInvocationStore;
+import com.superbiz.agent.harness.tool.store.CanonicalToolInvocation;
+import com.superbiz.agent.harness.tool.store.ToolCallKeyFactory;
+
+import java.util.ArrayList;
+import java.util.LinkedHashMap;
+import java.util.List;
+import java.util.Map;
+import java.util.Objects;
+
+public final class DiagnosisProgressProjector implements DiagnosisProgressProjection {
+
+ private static final int MAX_FACTS = 12;
+ private static final int MAX_SUMMARY_CHARS = 320;
+ private static final int MAX_SCOPE_CHARS = 320;
+
+ private final CanonicalInvocationStore store;
+ private final ToolCallKeyFactory keyFactory;
+ private final ObjectMapper objectMapper;
+
+ public DiagnosisProgressProjector(CanonicalInvocationStore store,
+ ToolCallKeyFactory keyFactory,
+ ObjectMapper objectMapper) {
+ this.store = Objects.requireNonNull(store, "store must not be null");
+ this.keyFactory = Objects.requireNonNull(keyFactory, "keyFactory must not be null");
+ this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
+ }
+
+ @Override
+ public DiagnosisProgressSnapshot project(RunContext context) {
+ Objects.requireNonNull(context, "context must not be null");
+ DiagnosisProgressSnapshotState state = context.progress().snapshot();
+ Map sources = new LinkedHashMap<>();
+ Map facts = new LinkedHashMap<>();
+ List limitations = new ArrayList<>();
+ for (CompletedToolCall completed : state.completedToolCalls()) {
+ CanonicalToolInvocation invocation = resolve(context, completed, limitations);
+ if (invocation == null) {
+ continue;
+ }
+ try {
+ projectInvocation(completed, invocation, sources, facts);
+ } catch (RuntimeException exception) {
+ addLimitation(limitations, "部分已完成的工具结果格式无法验证,未纳入已检查事实");
+ }
+ if (facts.size() >= MAX_FACTS) {
+ addLimitation(limitations, "已检查事实较多,展示内容已截断");
+ break;
+ }
+ }
+ return new DiagnosisProgressSnapshot(
+ List.copyOf(sources.values()),
+ List.copyOf(facts.values()),
+ List.copyOf(limitations),
+ state.stopReason());
+ }
+
+ private CanonicalToolInvocation resolve(RunContext context,
+ CompletedToolCall completed,
+ List limitations) {
+ try {
+ String key = keyFactory.create(context.runId(), completed.toolCallId());
+ CanonicalToolInvocation invocation = store.find(key).orElse(null);
+ if (invocation == null
+ || !invocation.isReferencableBy(context.runId())
+ || !completed.toolCallId().equals(invocation.toolCallId())
+ || !completed.toolName().equals(invocation.toolName())) {
+ addLimitation(limitations, "部分已完成的工具记录无法验证,未纳入已检查事实");
+ return null;
+ }
+ return invocation;
+ } catch (RuntimeException exception) {
+ addLimitation(limitations, "部分已完成的工具记录暂时不可读取,未纳入已检查事实");
+ return null;
+ }
+ }
+
+ private void projectInvocation(CompletedToolCall completed,
+ CanonicalToolInvocation invocation,
+ Map sources,
+ Map facts) {
+ JsonNode root = readObject(invocation.agentResult());
+ String scope = publicScope(completed.toolName(), completed.normalizedScope(), root);
+ switch (completed.toolName()) {
+ case AgentToolContracts.LOOKUP_KNOWLEDGE -> projectRag(root, scope, sources, facts);
+ case AgentToolContracts.QUERY_LOGS -> projectLogs(root, scope, sources, facts);
+ case AgentToolContracts.QUERY_MYSQL -> projectMysql(root, scope, sources, facts);
+ default -> throw new IllegalArgumentException("Unsupported evidence Tool");
+ }
+ }
+
+ private void projectRag(JsonNode root,
+ String scope,
+ Map sources,
+ Map facts) {
+ JsonNode evidence = root.path("evidence");
+ if (!evidence.isArray() || evidence.isEmpty()) {
+ addFact(sources, facts, "RAG", "knowledge_base", scope,
+ "该知识检索范围内未发现可用文档证据");
+ return;
+ }
+ for (JsonNode item : evidence) {
+ String source = firstNonBlank(text(item, "source"), text(item, "title"),
+ text(item, "document_id"), "knowledge_base");
+ addFact(sources, facts, "RAG", source, scope,
+ firstNonBlank(text(item, "excerpt"), "已找到候选知识片段"));
+ }
+ }
+
+ private void projectLogs(JsonNode root,
+ String scope,
+ Map sources,
+ Map facts) {
+ String source = firstNonBlank(text(root, "source_kind"), "logs");
+ JsonNode events = root.path("events");
+ if (!events.isArray() || events.isEmpty()) {
+ addFact(sources, facts, "LOG", source, scope,
+ "该日志查询范围内未发现匹配事件");
+ return;
+ }
+ for (JsonNode event : events) {
+ addFact(sources, facts, "LOG", source, scope,
+ firstNonBlank(text(event, "message"), "已找到匹配日志事件"));
+ }
+ }
+
+ private void projectMysql(JsonNode root,
+ String scope,
+ Map sources,
+ Map facts) {
+ String source = mysqlSource(scope);
+ JsonNode rows = root.path("rows");
+ if (!rows.isArray() || rows.isEmpty()) {
+ addFact(sources, facts, "MYSQL", source, scope,
+ "该只读数据查询范围内未发现匹配记录");
+ return;
+ }
+ for (JsonNode row : rows) {
+ addFact(sources, facts, "MYSQL", source, scope, row.toString());
+ }
+ }
+
+ private String publicScope(String toolName, String normalizedScope, JsonNode result) {
+ if (AgentToolContracts.LOOKUP_KNOWLEDGE.equals(toolName)) {
+ return bounded("query=" + text(result, "query"), MAX_SCOPE_CHARS);
+ }
+ if (AgentToolContracts.QUERY_LOGS.equals(toolName)) {
+ JsonNode scope = result.path("scope");
+ return bounded(scope.isObject() ? scope.toString() : normalizedScope, MAX_SCOPE_CHARS);
+ }
+ try {
+ JsonNode scope = objectMapper.readTree(normalizedScope);
+ return bounded("data_source=" + text(scope, "data_source"), MAX_SCOPE_CHARS);
+ } catch (JsonProcessingException exception) {
+ return "data_source=unknown";
+ }
+ }
+
+ private String mysqlSource(String scope) {
+ int separator = scope.indexOf('=');
+ return separator < 0 ? "mysql" : scope.substring(separator + 1);
+ }
+
+ private void addFact(Map sources,
+ Map facts,
+ String sourceType,
+ String source,
+ String scope,
+ String summary) {
+ if (facts.size() >= MAX_FACTS) {
+ return;
+ }
+ String safeSource = bounded(source, 160);
+ String safeScope = bounded(scope, MAX_SCOPE_CHARS);
+ String safeSummary = bounded(summary, MAX_SUMMARY_CHARS);
+ String sourceKey = sourceType + '\u0000' + safeSource + '\u0000' + safeScope;
+ sources.putIfAbsent(sourceKey,
+ new SafeFallback.VerifiedSource(sourceType, safeSource, safeScope));
+ String factKey = sourceKey + '\u0000' + safeSummary;
+ facts.putIfAbsent(factKey,
+ new SafeFallback.ObservedFact(sourceType, safeSource, safeScope, safeSummary));
+ }
+
+ private JsonNode readObject(String value) {
+ try {
+ JsonNode root = objectMapper.readTree(value);
+ if (root == null || !root.isObject()) {
+ throw new IllegalArgumentException("Canonical Agent result must be an object");
+ }
+ return root;
+ } catch (JsonProcessingException exception) {
+ throw new IllegalArgumentException("Canonical Agent result is invalid", exception);
+ }
+ }
+
+ private static void addLimitation(List limitations, String value) {
+ if (!limitations.contains(value)) {
+ limitations.add(value);
+ }
+ }
+
+ private static String firstNonBlank(String... values) {
+ for (String value : values) {
+ if (value != null && !value.isBlank()) {
+ return value;
+ }
+ }
+ return "unknown";
+ }
+
+ private static String text(JsonNode node, String field) {
+ JsonNode value = node == null ? null : node.get(field);
+ return value == null || value.isNull() ? "" : value.asText("");
+ }
+
+ private static String bounded(String value, int max) {
+ String safe = value == null ? "" : value;
+ return safe.length() <= max ? safe : safe.substring(0, max);
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressSnapshot.java b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressSnapshot.java
new file mode 100644
index 0000000..3e078db
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressSnapshot.java
@@ -0,0 +1,26 @@
+package com.superbiz.agent.harness.progress;
+
+import com.superbiz.agent.harness.contract.SafeFallback;
+
+import java.util.List;
+
+public record DiagnosisProgressSnapshot(
+ List verifiedSources,
+ List observedFacts,
+ List limitations,
+ DiagnosisStopReason stopReason) {
+
+ public DiagnosisProgressSnapshot {
+ verifiedSources = verifiedSources == null ? List.of() : List.copyOf(verifiedSources);
+ observedFacts = observedFacts == null ? List.of() : List.copyOf(observedFacts);
+ limitations = limitations == null ? List.of() : List.copyOf(limitations);
+ }
+
+ public static DiagnosisProgressSnapshot empty() {
+ return new DiagnosisProgressSnapshot(List.of(), List.of(), List.of(), null);
+ }
+
+ public boolean hasObservedFacts() {
+ return !observedFacts.isEmpty();
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressSnapshotState.java b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressSnapshotState.java
new file mode 100644
index 0000000..18d5cbe
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressSnapshotState.java
@@ -0,0 +1,16 @@
+package com.superbiz.agent.harness.progress;
+
+import java.util.List;
+
+public record DiagnosisProgressSnapshotState(
+ int consecutiveNoGain,
+ DiagnosisCollectionState collectionState,
+ DiagnosisStopReason stopReason,
+ String pendingToolCallId,
+ boolean stopInstructionDelivered,
+ List completedToolCalls) {
+
+ public DiagnosisProgressSnapshotState {
+ completedToolCalls = completedToolCalls == null ? List.of() : List.copyOf(completedToolCalls);
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressTracker.java b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressTracker.java
new file mode 100644
index 0000000..7300b08
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisProgressTracker.java
@@ -0,0 +1,118 @@
+package com.superbiz.agent.harness.progress;
+
+import com.superbiz.agent.harness.contract.EvidenceStatus;
+
+import java.util.ArrayList;
+import java.util.LinkedHashSet;
+import java.util.List;
+import java.util.Set;
+
+public final class DiagnosisProgressTracker {
+
+ private final int stopAfterConsecutiveNoGain;
+ private final Set completedScopes = new LinkedHashSet<>();
+ private final List completedToolCalls = new ArrayList<>();
+ private int consecutiveNoGain;
+ private DiagnosisCollectionState collectionState = DiagnosisCollectionState.COLLECTING;
+ private DiagnosisStopReason stopReason;
+ private String pendingToolCallId;
+ private boolean stopInstructionDelivered;
+
+ public DiagnosisProgressTracker(int stopAfterConsecutiveNoGain) {
+ if (stopAfterConsecutiveNoGain <= 0) {
+ throw new IllegalArgumentException("stopAfterConsecutiveNoGain must be positive");
+ }
+ this.stopAfterConsecutiveNoGain = stopAfterConsecutiveNoGain;
+ }
+
+ public synchronized void applyPreviousObservation(PreviousObservation observation) {
+ if (pendingToolCallId == null) {
+ if (observation != null) {
+ throw new IllegalArgumentException("No Tool observation is pending evaluation");
+ }
+ return;
+ }
+ if (observation == null) {
+ throw new IllegalArgumentException("Previous Tool observation must be evaluated");
+ }
+ if (!pendingToolCallId.equals(observation.toolCallId())) {
+ throw new IllegalArgumentException("Previous Tool observation ID is out of order");
+ }
+ pendingToolCallId = null;
+ applyGain(observation.informationGain());
+ }
+
+ public synchronized boolean isDuplicate(String toolName, String normalizedScope) {
+ return completedScopes.contains(new ToolScopeIdentity(toolName, normalizedScope));
+ }
+
+ public synchronized void recordDuplicateScope() {
+ applyGain(InformationGain.NO_GAIN);
+ }
+
+ public synchronized void recordCompleted(CompletedToolCall call, EvidenceStatus evidenceStatus) {
+ if (collectionState == DiagnosisCollectionState.SATURATED) {
+ throw new IllegalStateException("Cannot record Tool completion after saturation");
+ }
+ if (evidenceStatus != EvidenceStatus.EVIDENCE_FOUND
+ && evidenceStatus != EvidenceStatus.NO_EVIDENCE) {
+ throw new IllegalArgumentException("Completed Tool requires a successful evidence status");
+ }
+ ToolScopeIdentity scope = new ToolScopeIdentity(call.toolName(), call.normalizedScope());
+ if (!completedScopes.add(scope)) {
+ throw new IllegalStateException("Completed Tool scope was already recorded");
+ }
+ completedToolCalls.add(call);
+ if (evidenceStatus == EvidenceStatus.NO_EVIDENCE) {
+ applyGain(InformationGain.NO_GAIN);
+ } else {
+ pendingToolCallId = call.toolCallId();
+ }
+ }
+
+ public synchronized boolean claimStopInstruction() {
+ if (collectionState != DiagnosisCollectionState.SATURATED) {
+ return false;
+ }
+ if (stopInstructionDelivered) {
+ return false;
+ }
+ stopInstructionDelivered = true;
+ return true;
+ }
+
+ public synchronized void markBudgetLimitReached() {
+ if (stopReason == null) {
+ stopReason = DiagnosisStopReason.BUDGET_LIMIT_REACHED;
+ }
+ }
+
+ public synchronized DiagnosisProgressSnapshotState snapshot() {
+ return new DiagnosisProgressSnapshotState(
+ consecutiveNoGain,
+ collectionState,
+ stopReason,
+ pendingToolCallId,
+ stopInstructionDelivered,
+ completedToolCalls);
+ }
+
+ public int stopAfterConsecutiveNoGain() {
+ return stopAfterConsecutiveNoGain;
+ }
+
+ private void applyGain(InformationGain gain) {
+ if (collectionState == DiagnosisCollectionState.SATURATED) {
+ throw new IllegalStateException("Collection is already saturated");
+ }
+ if (gain == InformationGain.GAINED) {
+ consecutiveNoGain = 0;
+ return;
+ }
+ consecutiveNoGain++;
+ if (consecutiveNoGain >= stopAfterConsecutiveNoGain) {
+ collectionState = DiagnosisCollectionState.SATURATED;
+ stopReason = DiagnosisStopReason.INFORMATION_SATURATED;
+ }
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/DiagnosisStopReason.java b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisStopReason.java
new file mode 100644
index 0000000..de5eef0
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/DiagnosisStopReason.java
@@ -0,0 +1,6 @@
+package com.superbiz.agent.harness.progress;
+
+public enum DiagnosisStopReason {
+ INFORMATION_SATURATED,
+ BUDGET_LIMIT_REACHED
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/InformationGain.java b/src/main/java/com/superbiz/agent/harness/progress/InformationGain.java
new file mode 100644
index 0000000..6a8a934
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/InformationGain.java
@@ -0,0 +1,6 @@
+package com.superbiz.agent.harness.progress;
+
+public enum InformationGain {
+ GAINED,
+ NO_GAIN
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/PreviousObservation.java b/src/main/java/com/superbiz/agent/harness/progress/PreviousObservation.java
new file mode 100644
index 0000000..de5a409
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/PreviousObservation.java
@@ -0,0 +1,17 @@
+package com.superbiz.agent.harness.progress;
+
+import com.fasterxml.jackson.annotation.JsonProperty;
+
+public record PreviousObservation(
+ @JsonProperty("tool_call_id") String toolCallId,
+ @JsonProperty("information_gain") InformationGain informationGain) {
+
+ public PreviousObservation {
+ if (toolCallId == null || toolCallId.isBlank()) {
+ throw new IllegalArgumentException("toolCallId must not be blank");
+ }
+ if (informationGain == null) {
+ throw new IllegalArgumentException("informationGain must not be null");
+ }
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/ToolScopeIdentity.java b/src/main/java/com/superbiz/agent/harness/progress/ToolScopeIdentity.java
new file mode 100644
index 0000000..253677c
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/ToolScopeIdentity.java
@@ -0,0 +1,13 @@
+package com.superbiz.agent.harness.progress;
+
+public record ToolScopeIdentity(String toolName, String normalizedScope) {
+
+ public ToolScopeIdentity {
+ if (toolName == null || toolName.isBlank()) {
+ throw new IllegalArgumentException("toolName must not be blank");
+ }
+ if (normalizedScope == null || normalizedScope.isBlank()) {
+ throw new IllegalArgumentException("normalizedScope must not be blank");
+ }
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/progress/ToolScopeNormalizer.java b/src/main/java/com/superbiz/agent/harness/progress/ToolScopeNormalizer.java
new file mode 100644
index 0000000..0d05619
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/progress/ToolScopeNormalizer.java
@@ -0,0 +1,76 @@
+package com.superbiz.agent.harness.progress;
+
+import com.fasterxml.jackson.core.JsonProcessingException;
+import com.fasterxml.jackson.databind.ObjectMapper;
+import com.superbiz.agent.harness.tool.contract.AgentToolContracts;
+import com.superbiz.agent.harness.tool.contract.MysqlToolRequest;
+import com.superbiz.agent.harness.tool.contract.QueryLogsRequest;
+import com.superbiz.agent.harness.tool.contract.RagToolRequest;
+
+import java.util.LinkedHashMap;
+import java.util.Map;
+import java.util.Objects;
+
+public final class ToolScopeNormalizer {
+
+ private static final int DEFAULT_LOG_LOOKBACK_MINUTES = 30;
+
+ private final ObjectMapper objectMapper;
+
+ public ToolScopeNormalizer(ObjectMapper objectMapper) {
+ this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
+ }
+
+ public String normalize(String toolName, Object input) {
+ Objects.requireNonNull(input, "input must not be null");
+ Map scope = switch (toolName) {
+ case AgentToolContracts.LOOKUP_KNOWLEDGE -> ragScope(requireType(input, RagToolRequest.class));
+ case AgentToolContracts.QUERY_LOGS -> logScope(requireType(input, QueryLogsRequest.class));
+ case AgentToolContracts.QUERY_MYSQL -> mysqlScope(requireType(input, MysqlToolRequest.class));
+ default -> throw new IllegalArgumentException("Unsupported evidence Tool: " + toolName);
+ };
+ try {
+ return objectMapper.writeValueAsString(scope);
+ } catch (JsonProcessingException exception) {
+ throw new IllegalArgumentException("Tool scope is not serializable", exception);
+ }
+ }
+
+ private static Map ragScope(RagToolRequest request) {
+ return ordered("query", normalizeText(request.query()));
+ }
+
+ private static Map logScope(QueryLogsRequest request) {
+ Map scope = new LinkedHashMap<>();
+ scope.put("topic", request.topic() == null ? null : request.topic().name());
+ scope.put("query", normalizeText(request.query()));
+ scope.put("lookback_minutes", request.lookbackMinutes() == null
+ ? DEFAULT_LOG_LOOKBACK_MINUTES : request.lookbackMinutes());
+ return scope;
+ }
+
+ private static Map mysqlScope(MysqlToolRequest request) {
+ Map scope = new LinkedHashMap<>();
+ scope.put("data_source", normalizeText(request.dataSource()));
+ scope.put("sql", normalizeText(request.sql()));
+ scope.put("params", request.params());
+ return scope;
+ }
+
+ private static Map ordered(String name, Object value) {
+ Map result = new LinkedHashMap<>();
+ result.put(name, value);
+ return result;
+ }
+
+ private static String normalizeText(String value) {
+ return value == null ? "" : value.trim();
+ }
+
+ private static T requireType(Object input, Class type) {
+ if (!type.isInstance(input)) {
+ throw new IllegalArgumentException("Unexpected Tool input type: " + input.getClass().getName());
+ }
+ return type.cast(input);
+ }
+}
diff --git a/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCase.java b/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCase.java
index d7eaa84..a7b0a70 100644
--- a/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCase.java
+++ b/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCase.java
@@ -1,5 +1,6 @@
package com.superbiz.agent.harness.release;
+import com.superbiz.agent.harness.agent.DiagnosisAgentExecution;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.audit.DiagnosisTraceRecorder;
import com.superbiz.agent.harness.audit.TraceAuditEvents;
@@ -15,10 +16,13 @@ import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot;
import com.superbiz.agent.harness.guard.semantic.SemanticGuard;
import com.superbiz.agent.harness.guard.semantic.SemanticGuardDecision;
import com.superbiz.agent.harness.guard.semantic.SemanticGuardInput;
+import com.superbiz.agent.harness.progress.DiagnosisProgressSnapshot;
+import com.superbiz.agent.harness.progress.DiagnosisStopReason;
import com.superbiz.agent.harness.retry.RetryExecutionException;
import com.superbiz.agent.harness.retry.RetryFailure;
import java.util.Objects;
+import java.util.List;
public final class DiagnosisReleaseUseCase {
@@ -50,12 +54,44 @@ public final class DiagnosisReleaseUseCase {
}
public DiagnosisReleaseResult execute(RunContext context, String query, DiagnosisDraft draft) {
+ return execute(context, query, DiagnosisAgentExecution.completed(
+ Objects.requireNonNull(draft, "draft must not be null"),
+ DiagnosisProgressSnapshot.empty()));
+ }
+
+ public DiagnosisReleaseResult execute(
+ RunContext context, String query, DiagnosisAgentExecution execution) {
Objects.requireNonNull(context, "context must not be null");
if (query == null || query.isBlank()) {
throw new IllegalArgumentException("query must not be blank");
}
- Objects.requireNonNull(draft, "draft must not be null");
+ Objects.requireNonNull(execution, "execution must not be null");
+ DiagnosisDraft draft = execution.draft();
+ if (draft == null) {
+ return releaseControlledStop(context, execution.progress(), execution.stopReason());
+ }
+ if (draft.conclusion() == null) {
+ return releaseNoConclusion(context, draft, execution.progress());
+ }
+ return releaseConclusion(context, query, draft);
+ }
+
+ public DiagnosisReleaseResult releaseInvalidDraft(
+ RunContext context, DiagnosisProgressSnapshot progress) {
+ Objects.requireNonNull(context, "context must not be null");
+ Objects.requireNonNull(progress, "progress must not be null");
+ if (!progress.hasObservedFacts()) {
+ throw new IllegalStateException(
+ "Invalid Diagnosis Draft has no verified publishable progress");
+ }
+ return progressFallback(context,
+ fallbackFactory.insufficientEvidence(progress, List.of()),
+ FallbackType.INSUFFICIENT_EVIDENCE);
+ }
+
+ private DiagnosisReleaseResult releaseConclusion(
+ RunContext context, String query, DiagnosisDraft draft) {
DiagnosisDraft candidate = draft;
EvidenceGuardResult evidence = evidenceGuard.validate(context, candidate);
traceRecorder.record(TraceAuditEvents.evidenceValidation(
@@ -101,6 +137,61 @@ public final class DiagnosisReleaseUseCase {
return DiagnosisReleaseResult.fallback(fallbackFactory.semanticUnsupported(snapshot));
}
+ private DiagnosisReleaseResult releaseNoConclusion(
+ RunContext context, DiagnosisDraft draft, DiagnosisProgressSnapshot progress) {
+ EvidenceGuardResult evidence = evidenceGuard.validateNoConclusionReferences(context, draft);
+ traceRecorder.record(TraceAuditEvents.evidenceValidation(
+ context, TraceEventType.EVIDENCE_GUARD_INITIAL, evidence, draft));
+ if (!evidence.valid()) {
+ return evidenceFailure(context, evidence);
+ }
+
+ List missingInfo = missingInfo(draft);
+ if (progress.hasObservedFacts()) {
+ return progressFallback(context,
+ fallbackFactory.insufficientEvidence(progress, missingInfo),
+ FallbackType.INSUFFICIENT_EVIDENCE);
+ }
+ if (!missingInfo.isEmpty()) {
+ return progressFallback(context,
+ fallbackFactory.missingRequiredContext(missingInfo),
+ FallbackType.MISSING_REQUIRED_CONTEXT);
+ }
+ throw new IllegalStateException(
+ "No-conclusion Diagnosis has neither verified progress nor missing context");
+ }
+
+ private DiagnosisReleaseResult releaseControlledStop(
+ RunContext context,
+ DiagnosisProgressSnapshot progress,
+ DiagnosisStopReason stopReason) {
+ if (stopReason != DiagnosisStopReason.INFORMATION_SATURATED
+ && stopReason != DiagnosisStopReason.BUDGET_LIMIT_REACHED) {
+ throw new IllegalStateException("Unsupported Diagnosis stop reason");
+ }
+ if (!progress.hasObservedFacts()) {
+ throw new IllegalStateException(
+ "Controlled Diagnosis stop has no verified publishable progress");
+ }
+ return progressFallback(context,
+ fallbackFactory.insufficientEvidence(progress, List.of()),
+ FallbackType.INSUFFICIENT_EVIDENCE);
+ }
+
+ private DiagnosisReleaseResult progressFallback(
+ RunContext context,
+ com.superbiz.agent.harness.contract.SafeFallback fallback,
+ FallbackType type) {
+ traceRecorder.record(TraceAuditEvents.releaseDecision(
+ context, com.superbiz.agent.harness.contract.ReleaseOutcome.FALLBACK, type));
+ return DiagnosisReleaseResult.fallback(fallback);
+ }
+
+ private List missingInfo(DiagnosisDraft draft) {
+ return draft.limitations() == null
+ ? List.of() : draft.limitations().missingInfo();
+ }
+
private DiagnosisReleaseResult evidenceFailure(
RunContext context, EvidenceGuardResult evidence) {
traceRecorder.record(TraceAuditEvents.releaseDecision(
diff --git a/src/main/java/com/superbiz/agent/harness/release/EvidenceRepair.java b/src/main/java/com/superbiz/agent/harness/release/EvidenceRepair.java
index 396ee33..d014510 100644
--- a/src/main/java/com/superbiz/agent/harness/release/EvidenceRepair.java
+++ b/src/main/java/com/superbiz/agent/harness/release/EvidenceRepair.java
@@ -7,6 +7,7 @@ import com.fasterxml.jackson.databind.ObjectReader;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.audit.DiagnosisTraceRecorder;
import com.superbiz.agent.harness.audit.TraceAuditEvents;
+import com.superbiz.agent.harness.audit.ModelCallComponent;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.guard.evidence.EvidenceViolation;
@@ -92,7 +93,8 @@ public final class EvidenceRepair {
context.retryPolicies().evidenceRepair(),
() -> {
DiagnosisDraft repaired = parse(modelCall.call(
- context, modelPrompt, limits.timeout(), limits.maxOutputBytes()));
+ context, ModelCallComponent.EVIDENCE_REPAIR,
+ modelPrompt, limits.timeout(), limits.maxOutputBytes()));
if (!originalSemantics.hasSameUserVisibleSemantics(
SemanticDraftView.from(repaired))) {
throw new GuardModelCallException(
diff --git a/src/main/java/com/superbiz/agent/harness/release/SafeFallbackFactory.java b/src/main/java/com/superbiz/agent/harness/release/SafeFallbackFactory.java
index 604951a..9416491 100644
--- a/src/main/java/com/superbiz/agent/harness/release/SafeFallbackFactory.java
+++ b/src/main/java/com/superbiz/agent/harness/release/SafeFallbackFactory.java
@@ -5,6 +5,7 @@ import com.superbiz.agent.harness.contract.SafeFallback;
import com.superbiz.agent.harness.guard.evidence.EvidenceViolation;
import com.superbiz.agent.harness.guard.evidence.VerifiedEvidence;
import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot;
+import com.superbiz.agent.harness.progress.DiagnosisProgressSnapshot;
import java.util.ArrayList;
import java.util.LinkedHashMap;
@@ -53,6 +54,47 @@ public final class SafeFallbackFactory {
List.of());
}
+ public SafeFallback insufficientEvidence(
+ DiagnosisProgressSnapshot progress, List missingInfo) {
+ Objects.requireNonNull(progress, "progress must not be null");
+ if (!progress.hasObservedFacts()) {
+ throw new IllegalArgumentException(
+ "insufficient-evidence fallback requires verified progress");
+ }
+ List safeMissingInfo = boundedMissingInfo(missingInfo);
+ List limitations = new ArrayList<>(progress.limitations());
+ limitations.add("当前已检查范围不足以支持根因结论");
+ safeMissingInfo.forEach(item -> limitations.add("仍缺少:" + item));
+ return fallback(
+ FallbackType.INSUFFICIENT_EVIDENCE,
+ progress.verifiedSources(),
+ "已完成有限范围的检查,但现有证据不足以确认根因",
+ List.copyOf(limitations),
+ List.of(safeMissingInfo.isEmpty()
+ ? "补充故障对象、发生时间、错误信息或新的可查询范围后重试"
+ : "补充缺失信息后,在新的明确范围内继续诊断"),
+ "DIAGNOSIS_COLLECTION",
+ progress.observedFacts(),
+ List.of());
+ }
+
+ public SafeFallback missingRequiredContext(List missingInfo) {
+ List safeMissingInfo = boundedMissingInfo(missingInfo);
+ if (safeMissingInfo.isEmpty()) {
+ throw new IllegalArgumentException(
+ "missing-context fallback requires missing information");
+ }
+ return fallback(
+ FallbackType.MISSING_REQUIRED_CONTEXT,
+ List.of(),
+ "当前缺少执行定向诊断所需的信息,暂时无法开始有效查询或确认根因",
+ safeMissingInfo.stream().map(item -> "缺少:" + item).toList(),
+ List.of("补充上述故障上下文后重试"),
+ "DIAGNOSIS_INPUT",
+ List.of(),
+ List.of());
+ }
+
private SafeFallback fallback(FallbackType type,
List sources,
String message,
@@ -100,6 +142,20 @@ public final class SafeFallbackFactory {
return List.copyOf(result);
}
+ private List boundedMissingInfo(List missingInfo) {
+ List result = new ArrayList<>();
+ for (String item : missingInfo == null ? List.of() : missingInfo) {
+ String safe = bounded(item);
+ if (safe != null && !result.contains(safe)) {
+ result.add(safe);
+ }
+ if (result.size() >= 8) {
+ break;
+ }
+ }
+ return List.copyOf(result);
+ }
+
private String bounded(String value) {
if (value == null || value.isBlank()) {
return null;
diff --git a/src/main/java/com/superbiz/agent/harness/tool/contract/MysqlToolCall.java b/src/main/java/com/superbiz/agent/harness/tool/contract/MysqlToolCall.java
new file mode 100644
index 0000000..c292679
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/tool/contract/MysqlToolCall.java
@@ -0,0 +1,9 @@
+package com.superbiz.agent.harness.tool.contract;
+
+import com.fasterxml.jackson.annotation.JsonProperty;
+import com.superbiz.agent.harness.progress.PreviousObservation;
+
+public record MysqlToolCall(
+ @JsonProperty("previous_observation") PreviousObservation previousObservation,
+ @JsonProperty("input") MysqlToolRequest input) {
+}
diff --git a/src/main/java/com/superbiz/agent/harness/tool/contract/QueryLogsToolCall.java b/src/main/java/com/superbiz/agent/harness/tool/contract/QueryLogsToolCall.java
new file mode 100644
index 0000000..df4673a
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/tool/contract/QueryLogsToolCall.java
@@ -0,0 +1,9 @@
+package com.superbiz.agent.harness.tool.contract;
+
+import com.fasterxml.jackson.annotation.JsonProperty;
+import com.superbiz.agent.harness.progress.PreviousObservation;
+
+public record QueryLogsToolCall(
+ @JsonProperty("previous_observation") PreviousObservation previousObservation,
+ @JsonProperty("input") QueryLogsRequest input) {
+}
diff --git a/src/main/java/com/superbiz/agent/harness/tool/contract/RagRelevanceLevel.java b/src/main/java/com/superbiz/agent/harness/tool/contract/RagRelevanceLevel.java
new file mode 100644
index 0000000..ed27052
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/tool/contract/RagRelevanceLevel.java
@@ -0,0 +1,7 @@
+package com.superbiz.agent.harness.tool.contract;
+
+public enum RagRelevanceLevel {
+ PRECISE,
+ HIGHLY_RELEVANT,
+ REFERENCE
+}
diff --git a/src/main/java/com/superbiz/agent/harness/tool/contract/RagToolCall.java b/src/main/java/com/superbiz/agent/harness/tool/contract/RagToolCall.java
new file mode 100644
index 0000000..9c9cb3d
--- /dev/null
+++ b/src/main/java/com/superbiz/agent/harness/tool/contract/RagToolCall.java
@@ -0,0 +1,9 @@
+package com.superbiz.agent.harness.tool.contract;
+
+import com.fasterxml.jackson.annotation.JsonProperty;
+import com.superbiz.agent.harness.progress.PreviousObservation;
+
+public record RagToolCall(
+ @JsonProperty("previous_observation") PreviousObservation previousObservation,
+ @JsonProperty("input") RagToolRequest input) {
+}
diff --git a/src/main/java/com/superbiz/agent/harness/tool/contract/RagToolResult.java b/src/main/java/com/superbiz/agent/harness/tool/contract/RagToolResult.java
index e608e5f..563b03b 100644
--- a/src/main/java/com/superbiz/agent/harness/tool/contract/RagToolResult.java
+++ b/src/main/java/com/superbiz/agent/harness/tool/contract/RagToolResult.java
@@ -1,6 +1,7 @@
package com.superbiz.agent.harness.tool.contract;
import com.fasterxml.jackson.annotation.JsonProperty;
+import com.fasterxml.jackson.annotation.JsonInclude;
import com.superbiz.agent.harness.contract.EvidenceStatus;
import java.util.List;
@@ -11,9 +12,20 @@ public record RagToolResult(
@JsonProperty("query") String query,
@JsonProperty("evidence") List evidence,
@JsonProperty("returned_count") int returnedCount,
+ @JsonProperty("relevance_level") @JsonInclude(JsonInclude.Include.NON_NULL)
+ RagRelevanceLevel relevanceLevel,
@JsonProperty("truncated") boolean truncated) {
public RagToolResult {
evidence = ToolContractCollections.immutable(evidence);
}
+
+ public RagToolResult(EvidenceStatus evidenceStatus,
+ String toolCallId,
+ String query,
+ List evidence,
+ int returnedCount,
+ boolean truncated) {
+ this(evidenceStatus, toolCallId, query, evidence, returnedCount, null, truncated);
+ }
}
diff --git a/src/main/java/com/superbiz/agent/harness/tool/projection/RagResultProjector.java b/src/main/java/com/superbiz/agent/harness/tool/projection/RagResultProjector.java
index c0990b2..263b49f 100644
--- a/src/main/java/com/superbiz/agent/harness/tool/projection/RagResultProjector.java
+++ b/src/main/java/com/superbiz/agent/harness/tool/projection/RagResultProjector.java
@@ -5,6 +5,7 @@ import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.harness.contract.EvidenceStatus;
import com.superbiz.agent.harness.tool.boundary.ProjectedToolResult;
import com.superbiz.agent.harness.tool.contract.RagEvidence;
+import com.superbiz.agent.harness.tool.contract.RagRelevanceLevel;
import com.superbiz.agent.harness.tool.contract.RagToolRequest;
import com.superbiz.agent.harness.tool.contract.RagToolResult;
@@ -81,7 +82,10 @@ public final class RagResultProjector {
EvidenceStatus status = evidence.isEmpty()
? EvidenceStatus.NO_EVIDENCE
: EvidenceStatus.EVIDENCE_FOUND;
- RagToolResult result = new RagToolResult(status, toolCallId, query, evidence, evidence.size(), truncated);
+ RagRelevanceLevel relevanceLevel = status == EvidenceStatus.EVIDENCE_FOUND
+ ? relevanceLevel(root) : null;
+ RagToolResult result = new RagToolResult(
+ status, toolCallId, query, evidence, evidence.size(), relevanceLevel, truncated);
result = fitBudget(result, truncated);
return new ProjectedToolResult(objectMapper.writeValueAsString(result), result.evidenceStatus());
}
@@ -94,7 +98,8 @@ public final class RagResultProjector {
reduced.remove(reduced.size() - 1);
current = new RagToolResult(
reduced.isEmpty() ? EvidenceStatus.NO_EVIDENCE : EvidenceStatus.EVIDENCE_FOUND,
- current.toolCallId(), current.query(), reduced, reduced.size(), true);
+ current.toolCallId(), current.query(), reduced, reduced.size(),
+ reduced.isEmpty() ? null : current.relevanceLevel(), true);
}
if (bytes(objectMapper.writeValueAsString(current)) > limits.maxAgentUtf8Bytes()) {
throw new IllegalArgumentException("RAG projection exceeds total budget");
@@ -120,6 +125,21 @@ public final class RagResultProjector {
return value == null || value.isBlank() ? null : value;
}
+ private static RagRelevanceLevel relevanceLevel(JsonNode root) {
+ String value = text(root, "relevance_level");
+ if (value.isBlank()) {
+ value = text(root, "relevanceLevel");
+ }
+ if (value.isBlank()) {
+ return null;
+ }
+ try {
+ return RagRelevanceLevel.valueOf(value.trim().toUpperCase(java.util.Locale.ROOT));
+ } catch (IllegalArgumentException ignored) {
+ return null;
+ }
+ }
+
private static String bounded(String value, int max) {
if (value == null) {
return "";
diff --git a/src/main/java/com/superbiz/agent/repository/AgentStepRepository.java b/src/main/java/com/superbiz/agent/repository/AgentStepRepository.java
index aea34e9..1bf3c79 100644
--- a/src/main/java/com/superbiz/agent/repository/AgentStepRepository.java
+++ b/src/main/java/com/superbiz/agent/repository/AgentStepRepository.java
@@ -5,6 +5,7 @@ import org.springframework.data.jpa.repository.JpaRepository;
import org.springframework.stereotype.Repository;
import java.util.List;
+import java.util.Optional;
/**
* Agent 决策步骤 Repository
@@ -27,6 +28,8 @@ public interface AgentStepRepository extends JpaRepository {
*/
List findByRunIdOrderByStepIndex(String runId);
+ Optional findByRunIdAndStepIndex(String runId, Integer stepIndex);
+
/**
* 统计某个会话的步骤数
*/
diff --git a/src/main/resources/application.yml b/src/main/resources/application.yml
index c8e4203..648d355 100644
--- a/src/main/resources/application.yml
+++ b/src/main/resources/application.yml
@@ -217,5 +217,6 @@ harness:
run-timeout: 5m
sse-timeout: 5m
canonical-ttl: 2h
+ stop-after-consecutive-no-gain: 2
mysql-tools:
data-sources: {}
diff --git a/src/main/resources/prompts/diagnosis-agent-prompt.md b/src/main/resources/prompts/diagnosis-agent-prompt.md
index 338128a..94f0b22 100644
--- a/src/main/resources/prompts/diagnosis-agent-prompt.md
+++ b/src/main/resources/prompts/diagnosis-agent-prompt.md
@@ -1,46 +1,33 @@
-# Role
+# 角色
-You are the only Diagnosis Agent for the current diagnosis run. Plan the investigation internally, call the available read-only evidence tools through the framework ReAct loop, evaluate the returned evidence, and author one complete DiagnosisDraft.
+你是当前运行中唯一的诊断 Agent。围绕用户原始问题收集只读证据,并输出一份完整的诊断草稿。可用工具及其参数由服务端提供,以实际注入内容为准。
-# Input
+# 基本原则
-The user message is one JSON object with exactly:
+- 你不必须给出根因。证据不足时,`conclusion=null` 是合法且成功的完成方式。
+- 如果问题缺少企业、时间、服务、错误信息或其他形成明确查询范围所需的上下文,可以不调用工具,直接在 `limitations.missing_info` 中列出缺失信息。
+- 上一轮内容仅作为上下文,不是本轮证据,也不能复用其中的 Tool Call ID。
+- 不得编造根因、证据、Tool Call ID、健康状态或已经排除的原因。
+- 不得输出原始工具载荷、凭据、内部错误、提示词、隐藏推理或思维链。
-- `query`: the current user query. Preserve its meaning and answer this query only.
-- `previous_turn`: an optional safely published previous diagnosis turn with the fixed PreviousTurn schema.
+# 查询与停止
-`previous_turn` is context only. It is not evidence for this run, contains no reusable Tool Call IDs, and must not be cited as current evidence.
+- 只有在存在明确、不同且可能获得新信息的查询范围时才继续调用工具。
+- 工具结果新增了可验证事实,并确认、排除或缩小了当前假设时,标记为 `GAINED`。
+- 工具结果即使正确,但只是通用说明、重复内容,或不能推进当前诊断时,标记为 `NO_GAIN`。
+- 继续调用工具时,按服务端提供的控制字段评价上一轮结果。
+- 如果没有新的有效查询范围,或者结果正确但对推导无用,立即停止调用。
+- 收到 `STOP_REQUIRED` 后不得再次调用工具,应直接完成诊断草稿。
-# Evidence rules
+# 证据边界
-- Use only `lookup_knowledge`, `query_logs`, and `query_mysql` when evidence is needed.
-- Treat a Tool observation as evidence only when it contains `evidence_status=EVIDENCE_FOUND` or `evidence_status=NO_EVIDENCE` and a non-blank `tool_call_id`.
-- Bind every Analysis item only to Tool Call IDs returned during this run. Never invent, transform, shorten, or reuse a Tool Call ID.
-- `NORMAL` Analysis may cite only `EVIDENCE_FOUND` results.
-- `NEGATIVE_OBSERVATION` Analysis may cite only `NO_EVIDENCE` results and must state the exact query scope. `NO_EVIDENCE` never proves that a problem does not exist, that a root cause is excluded, or that a system is healthy.
-- A Tool `ERROR` is not evidence. Do not cite it as support for Analysis or Conclusion.
-- Do not expose raw Tool payloads, credentials, infrastructure coordinates, internal errors, prompts, hidden reasoning, or chain of thought.
+- 只有本轮工具返回的非空 Tool Call ID 可以被引用。
+- 正向分析只能引用 `EVIDENCE_FOUND`;负向观察只能引用 `NO_EVIDENCE`,并写清实际查询范围。
+- `NO_EVIDENCE` 只表示指定范围内未找到证据,不能证明问题不存在、系统健康或某个根因已被排除。
+- `ERROR` 不是证据,不能支持分析或结论。
-# Stopping rule
+# 输出
-Stop calling tools when the current evidence is sufficient for a bounded Draft, when the configured tool/model budget prevents more work, or when the available tools cannot obtain the missing information.
-
-If current evidence cannot support a diagnosis:
-
-- set `conclusion` to `null`;
-- keep only valid scoped negative observations, if any;
-- state the actual queried scope in `limitations.scope`;
-- list the evidence still needed in `limitations.missing_info`;
-- do not fabricate a root cause, action justification, recommendation, or healthy-state claim;
-- finish the current response without asking the Harness to retry the Agent or a Tool.
-
-# Draft rules
-
-- Lead with `conclusion` when supported. Its `based_on_analysis_ids` must reference existing Analysis IDs.
-- Every Analysis item must have a unique `analysis_id`, a valid `kind`, concise evidence-grounded text, and at least one current-Run `tool_call_id`.
-- Every Action Plan and Recommendation item must reference existing Analysis IDs.
-- Mark any dangerous or side-effecting action with `requires_human_confirmation=true`.
-- `limitations.scope` must not exceed the actual Tool query scopes.
-- `limitations.missing_info` must name material gaps that constrain the conclusion.
-
-Return exactly one JSON object matching the supplied DiagnosisDraft schema. Do not use Markdown fences, headings, prose before or after JSON, or a separate answer field.
+- 有充分证据时,结论、分析、行动和建议必须通过分析 ID 与本轮 Tool Call ID 建立完整引用关系。
+- 证据不足时,使用 `conclusion=null`,只保留可验证的范围和观察,并在 `limitations` 中说明实际检查范围与缺失信息。
+- 返回一个符合服务端输出约束的 JSON 对象。不要使用 Markdown 代码块,也不要添加 JSON 之外的文字。
diff --git a/src/main/resources/static/app.js b/src/main/resources/static/app.js
index 19648b4..cc3bfb4 100644
--- a/src/main/resources/static/app.js
+++ b/src/main/resources/static/app.js
@@ -703,7 +703,33 @@ class SuperBizAgentApp {
return [payload.answer || '', references ? `\n\n参考资料:\n${references}` : ''].join('');
}
if (contentType === 'SAFE_FALLBACK') {
- return payload.fallback?.message || '当前证据不足,无法确认根因';
+ const fallback = payload.fallback || {};
+ const sections = [];
+ sections.push(`## 当前判断\n${fallback.message || '当前证据不足,无法确认根因'}`);
+ if (fallback.observed_facts?.length) {
+ const facts = fallback.observed_facts.map(item =>
+ `- **${item.source_type || '证据'} / ${item.source || '未知来源'}**(${item.scope || '范围未提供'}):${item.summary || '已完成检查'}`
+ ).join('\n');
+ sections.push(`## 已完成排查\n${facts}`);
+ } else if (fallback.verified_sources?.length) {
+ const sources = fallback.verified_sources.map(item =>
+ `- **${item.source_type || '证据'} / ${item.source || '未知来源'}**:${item.scope || '范围未提供'}`
+ ).join('\n');
+ sections.push(`## 已检查范围\n${sources}`);
+ }
+ if (fallback.validation_issues?.length) {
+ const issues = fallback.validation_issues.map(item =>
+ `- ${item.code || 'VALIDATION_FAILED'}${item.target ? `:${item.target}` : ''}`
+ ).join('\n');
+ sections.push(`## 校验问题\n${issues}`);
+ }
+ if (fallback.limitations?.length) {
+ sections.push(`## 限制\n${fallback.limitations.map(item => `- ${item}`).join('\n')}`);
+ }
+ if (fallback.next_steps?.length) {
+ sections.push(`## 下一步\n${fallback.next_steps.map(item => `- ${item}`).join('\n')}`);
+ }
+ return sections.join('\n\n');
}
if (contentType === 'DIAGNOSIS_REPORT') {
const report = payload.report || {};
@@ -713,6 +739,15 @@ class SuperBizAgentApp {
if (report.action_plan?.length) sections.push(`## 行动计划\n${report.action_plan.map(item => `- ${item.action}`).join('\n')}`);
if (report.recommendations?.length) sections.push(`## 建议\n${report.recommendations.map(item => `- ${item.text}`).join('\n')}`);
if (report.limitations?.scope) sections.push(`## 限制\n${report.limitations.scope}`);
+ if (report.limitations?.missing_info?.length) {
+ sections.push(`## 仍缺信息\n${report.limitations.missing_info.map(item => `- ${item}`).join('\n')}`);
+ }
+ if (payload.references?.length) {
+ const references = payload.references.map(item =>
+ `- **${item.source_type || '证据'} / ${item.source || '未知来源'}**:${item.scope || '范围未提供'}`
+ ).join('\n');
+ sections.push(`## 已检查范围\n${references}`);
+ }
return sections.join('\n\n');
}
throw new Error(`未知 content_type: ${contentType}`);
diff --git a/src/main/resources/static/index.html b/src/main/resources/static/index.html
index 1e60135..e649f5d 100644
--- a/src/main/resources/static/index.html
+++ b/src/main/resources/static/index.html
@@ -111,6 +111,6 @@
-
+