feat(harness): add evidence and semantic guards

This commit is contained in:
zhuyongxin
2026-07-21 23:50:51 +08:00
parent 2362665519
commit ee0949d464
40 changed files with 2980 additions and 2 deletions
+1
View File
@@ -40,3 +40,4 @@
| 2026-07-21 | single-react-rag-log-projections | RAG/log projection adapters through ToolBoundary | Harness/Tool projection | ISS-014, RAG, query_logs, projection, scope, redaction, MOCK, NO_EVIDENCE | openspec/changes/archive/2026-07-21-single-react-rag-log-projections | archived | | 2026-07-21 | single-react-rag-log-projections | RAG/log projection adapters through ToolBoundary | Harness/Tool projection | ISS-014, RAG, query_logs, projection, scope, redaction, MOCK, NO_EVIDENCE | openspec/changes/archive/2026-07-21-single-react-rag-log-projections | archived |
| 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived | | 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived |
| 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived | | 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived |
| 2026-07-21 | single-react-evidence-semantic-guards | Deterministic evidence validation, isolated semantic review and fail-closed diagnosis release | Harness/EvidenceGuard/SemanticGuard/Release | ISS-014, EvidenceGuard, verified snapshot, SemanticGuard, repair, fallback, release policy | openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards | archived |
@@ -0,0 +1,44 @@
# Acceptance: single-react-evidence-semantic-guards
## Result
- Status: archived
- OpenSpec tasks: 14/14 complete
- Interface impact: L2 internal
- Public protocol: unchanged
## Static Verification
- `git diff --check`:阶段文件无 whitespace error;工作区用户已有 `AGENTS.md`/`CLAUDE.md` 仅报告 line-ending warning,未纳入阶段范围。
- 公开 `controller`、`ChatService.java`、`AiOpsService.java` diff 为空。
- SemanticGuard 包和 prompt 不含 Tool Call ID、raw response、ReactAgent、StateGraph、ThreadLocal 或手写 while。
- 新 guard/release 包不包含业务 Agent loop。
## Script Verification
- `mvn -q -DskipTests compile`:通过。
- Stage-focused `EvidenceGuardTest,SemanticGuardTest,DiagnosisReleaseUseCaseTest`:19 tests,0 failure/error。
- Harness/Tool/Agent regression selection:18 suites / 70 tests,0 failure/error/skipped。
- `openspec validate single-react-evidence-semantic-guards --strict`:通过。
- `openspec validate --specs --strict`:16 passed / 2 failed;失败是前序 `mysql-readonly-tool`、`rag-log-projections` scenario 标题格式,不阻塞本 change,计划阶段 7 统一修复。
## Browser or Manual Verification
- Not applicable。阶段 5 没有 UI 或公开入口变化。
## Not Verified
- 未运行真实 LLM、Redis、日志和 MySQL live E2E;按 ISS-014 串行门禁统一留到阶段 7。
- Provider 是否立即响应 Java thread interrupt 取决于 SDK;Harness 已保证 Future cancel、迟到 Token 记账和迟到结果不释放,阶段 7 需用真实模型观察取消延迟。
## Remaining Work
- 阶段 6A:Chat Application Use Case、Intent Router、PreviousTurn 和 Run persistence。
- 阶段 6B:公开 SSE 原子切换。
- 阶段 7:旧链路清理、全局 spec 格式修复和最终 live E2E。
## Archive
- `.archive-ready`: created
- OpenSpec archive: `openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards`
- Main spec: `openspec/specs/single-react-evidence-semantic-guards/spec.md`
@@ -0,0 +1,32 @@
# Brief: single-react-evidence-semantic-guards
## Background
阶段 4 已能生成带框架 Tool Call ID 的 `DiagnosisDraft`,但草稿在发布前还缺少当前 Run 证据验真、独立语义审查和 fail-closed 释放边界。
## Goal
在不恢复业务 Graph、不切换公开入口的前提下,实现确定性 EvidenceGuard、一次语义不变的结构修复、无 Tool/无记忆的单轮 SemanticGuard,以及只发布原 Draft 或固定 SafeFallback 的内部 release use case。
## Scope
- Draft 结构、Analysis/报告引用和 canonical Tool invocation 验真。
- RAG/log/MySQL verified evidence snapshot。
- Evidence repair 单次调用与语义不变约束。
- SemanticGuard 模型/Token/字节预算、超时、取消、技术重试和严格二元输出。
- SUPPORTED/UNSUPPORTED/unavailable/evidence-failed release policy。
- focused tests、Harness/Tool/Agent 回归和静态范围验证。
## Non-goals
- 不切换 Chat/AiOps/SSE 公开协议。
- 不实现 Intent Router、PreviousTurn、Run persistence 或最终报告渲染。
- 不删除旧 Gatekeeper/Verifier/Composer/多 Agent 链路。
- 不运行 live model/Redis/MySQL E2E;统一留到阶段 7。
## Metadata
- Scale: complex
- Interface impact: L2 internal
- OpenSpec: `single-react-evidence-semantic-guards`
- Parent issue: `ISS-014`
@@ -0,0 +1,144 @@
# Decisions: single-react-evidence-semantic-guards
## Discover Status
- Checkpoint: Discover
- Capability source: `sm-flow`,使用 `grill-with-docs` 做 evidence-driven 澄清;`codebase-retrieval`、LSP 和 GitNexus MCP 在当前会话不可用,降级为既有 GitNexus 研究结论、`rg` 调用点核对和逐文件源码阅读。
- Scale: complex。变更跨 canonical store、三类 Tool projection、模型预算、超时/取消、结构修复、语义审查和释放边界,但不切换公开协议。
- `devflow/index.md` 命中阶段 0-4 的设计冻结、RunContext、canonical store、Tool projection 和 Diagnosis Agent;没有与当前 OpenSpec 冲突的 ADR。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | EvidenceGuard 是否判断证据足以推出结论? | evidence-driven | 已解决 |
| Q2 | 边界 | verified snapshot 能读取哪些 canonical 字段,并向 SemanticGuard 暴露哪些内容? | evidence-driven | 已解决 |
| Q3 | 边界 | `NO_EVIDENCE` 能支持哪类 Analysis? | evidence-driven | 已解决 |
| Q4 | 修复 | 首次 EvidenceGuard 失败允许如何修复,是否可重跑 Agent 或 Tool? | evidence-driven | 已解决 |
| Q5 | 语义 | SemanticGuard 是否拥有 Tool、记忆、ReAct loop 或报告写作能力? | evidence-driven | 已解决 |
| Q6 | 重试 | 哪些 SemanticGuard 结果允许第二次 attempt? | evidence-driven | 已解决 |
| Q7 | 超时 | 单轮模型超时和 Run 取消如何终止后台调用? | evidence-driven | 已解决 |
| Q8 | 释放 | 三种 Fallback 能否包含 Draft、reason 或未验真来源? | evidence-driven | 已解决 |
| Q9 | 接口 | 阶段 5 是否切换公开入口或修改 SSE/持久化协议? | evidence-driven | 已解决 |
| Q10 | 验收 | 如何证明不存在隐式 Agent/repair/retry loop? | evidence-driven | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| EvidenceGuard 只验证结构、引用和 canonical invocation 真实性,不判断 Analysis/Conclusion 的语义充分性。 | ISS-014 4.3、阶段 5;glossary `EvidenceGuard` | 已汇报 |
| Store 只能按 `ToolCallKeyFactory.create(runId, toolCallId)` 查找;可引用记录必须是当前 Run、`READY`、含 `agent_result` 且状态为 `EVIDENCE_FOUND|NO_EVIDENCE`。 | `CanonicalInvocationStore`、`CanonicalToolInvocation.isReferencableBy` | 已汇报 |
| Snapshot 解析 `agent_result` 和必要的有界 request 字段,不读取 `raw_response`;输出按 `analysis_id` 分组且不包含 Tool Call ID。 | ISS-014 4.4、10.2、已冻结 Tool contracts | 已汇报 |
| `NORMAL` 只绑定 `EVIDENCE_FOUND`;`NEGATIVE_OBSERVATION` 只绑定 `NO_EVIDENCE`,且零结果范围来自投影/请求。 | ISS-014 4.3、10.2;`EvidenceStatus` glossary | 已汇报 |
| Evidence repair 只在首次物理验真失败后执行一次无 Tool单轮模型调用,不重跑 Diagnosis Agent 或 Tool loop。 | ISS-014 10.3、阶段 5、重试边界决策 | 已汇报 |
| SemanticGuard 复用同一 `ChatModel`,使用全新 `Prompt` 单轮调用,无 Tool/记忆/ReAct;只返回二元 verdict 和审计 reason。 | ISS-014 4.4、阶段 5;Spring AI `ChatModel.call(Prompt)` | 已汇报 |
| 仅超时、传输、解析或 Schema 技术失败可重试一次;`UNSUPPORTED` 是有效业务结果,不重试。 | `HarnessRetryPolicies.strict().semanticGuard()`、ISS-014 重试策略 | 已汇报 |
| `ChatResponseMetadata.Usage` 可复用 Core 的 model/Token 预算;受控 `Future.get(timeout)` 可在超时或 Run 取消时取消任务。 | `HarnessModelInterceptor`、`DiagnosisHarnessCore`、`RunCancellation` | 已汇报 |
| `EVIDENCE_VALIDATION_FAILED` 的来源必须为空;其余 Fallback 只能包含已验真来源,不包含 Draft 或 SemanticGuard reason。 | ISS-014 10.3、阶段 5;`SafeFallback`/`FallbackType` | 已汇报 |
| 阶段 5 只新增内部用例,公开入口切换留到阶段 6B。 | ISS-014 阶段顺序和 6B 门禁 | 已汇报 |
## User-interview
- 本阶段没有新增 user-interview 问题。上述方向、范围、修复次数、模型复用、超时/重试、Fallback 和阶段串行规则均已在 ISS-014 评审及前序对话中由用户确认。
## Key Decisions
- Snapshot 使用 Tool-specific adapter 严格解析冻结 projection;未知 Tool、ID 不一致、projection 反序列化失败或 projection 的 EvidenceStatus 与 canonical record 不一致均 fail closed。
- MySQL 的逻辑数据源和查询范围来自 canonical `request`,行/列和值来自 `agent_result`;RAG/log 以 `agent_result` 自带的稳定来源和范围为准。
- SemanticGuard 使用直接 `ChatModel.call(Prompt)`,不创建第二个 `ReactAgent`;模型调用由受控 Executor 执行,超时/取消时调用 `Future.cancel(true)`。
- Evidence repair 与 SemanticGuard 共用同一模型抽象,但各自拥有独立 prompt、严格输出类型和预算;repair 策略固定一次 attempt,不能嵌套 retry executor。
- Release use case 只在 EvidenceGuard 成功后调用 SemanticGuard;SUPPORTED 返回原 Draft 对象,所有其他路径只返回固定 `SafeFallback`。
- 不创建 ADR:这些是 ISS-014 已冻结架构的阶段实现,不产生新的难逆转跨项目决策。
## OpenSpec Backfill
- 需进入 proposal/design/spec/tasks:EvidenceGuard 规则、Tool-specific snapshot、无 raw/ID 泄漏、repair 单次边界、SemanticGuard 隔离/预算/超时/重试、原样发布与三类 Fallback、L2/公开隔离。
- 不进入本阶段:公开 Chat/SSE 装配、Run 持久化、旧多 Agent 删除和 live E2E。
## Cross-artifact Alignment
| 上游 -> 下游 | 检查内容 | 状态 |
|---|---|---|
| ISS-014/brief -> proposal | 物理验真、结构修复、隔离语义审查、固定 Fallback、阶段边界 | 已对齐 |
| proposal -> design | Store ownership、Tool-specific snapshot、timeout/cancel/retry、唯一报告作者和 L2 影响 | 已对齐 |
| design -> specs/tasks | 每个关键决策均有可观察 requirement 和对应实现/测试切片 | 已对齐 |
| specs -> tasks | 9 组 requirements 覆盖为 Evidence、model boundary、repair/release 和回归验证任务 | 已对齐 |
## Architecture Audit
- 能力来源:`zoom-out`,使用 glossary 的 Diagnosis Harness、EvidenceGuard、SemanticGuard、RunContext、Invocation Status 和 Evidence Status 术语。
- 链路为 `query + Draft + RunContext -> EvidenceGuard/store -> optional repair/shared model boundary -> SemanticGuard/shared model boundary -> original Draft or fixed Fallback`,没有外层 Graph 或第二个 Tool loop。
- Store 拥有物理调用记录,EvidenceGuard 只读并生成 snapshot;Diagnosis Agent 拥有报告语义,repair 只修 ID,SemanticGuard 只审查,release 只做二选一,数据所有权没有重叠。
- 最大运行风险是 provider 忽略 interrupt;design 要求 Future cancel、迟到 Token 记录和 `core.checkActive` 丢弃迟到结果,阶段 6A 再负责 Run 终态。
- 架构审计未发现与阶段 0-4 或 ADR 冲突;所有风险缓解已回写 design/spec/tasks。
## Interface Impact
- 级别:L2 内部接口。
- 变更对象:新增 guard/release records、interfaces、use cases、prompts 和 tests;现有方法签名不变。
- 消费者:本阶段只有 focused tests,阶段 6A 才接入 production application use case。
- 兼容性:公开 HTTP/SSE、Controller DTO、数据库、Redis key schema 和旧 Chat/AiOps 路径不变。
## Commit Gate Preflight
- proposal、design、specs、tasks 完整,`openspec status` 为 complete,`openspec validate single-react-evidence-semantic-guards --strict` 通过。
- Question pool 全部为已汇报的 evidence-driven 结论;无未确认 user-interview、接口等级或风险接受问题。
- Cross-artifact 四段对齐无 gap;架构审计的数据所有权、模型取消和唯一报告作者约束已进入 design/spec/tasks。
- 接口影响为 L2,仅新增内部 Java API;公开 Controller/SSE/JPA/Redis schema/旧 ChatService 保持不变。
- Apply、Archive 和阶段 Git commit 使用用户对 ISS-014 各阶段的持续授权;实现必须严格限制为 Committed OpenSpec。
- `.committed` 已创建,Committed OpenSpec 可进入 Apply。
## Pre-apply Research
### Reference Implementations
- `harness/tool/store/CanonicalInvocationStore.java`、`CanonicalToolInvocation.java`、`ToolCallKeyFactory.java`:当前 Run 物理验真和 lifecycle 真理源。
- `harness/tool/projection/RagResultProjector.java`、`QueryLogsResultProjector.java`、`harness/tool/mysql/MysqlResultProjector.java`:三类 `agent_result` 的唯一生产者和有界字段来源。
- `harness/agent/HarnessModelInterceptor.java`:模型调用前预算、响应 Usage 记账和 late-result active check 模式。
- `harness/retry/HarnessRetryExecutor.java`、`HarnessRetryPolicies.java`:SemanticGuard 两次技术 attempt 与 repair 一次 attempt 的装配边界。
- `harness/core/RunCancellation.java`:取消 callback 注册和 first-cancel 语义。
- `service/ExecutorGatekeeperService.java`:仅参考报告内部引用检查思想,不复用旧 Map/JPA/session DTO 或规则目录。
### Technology Stack
- Jackson strict `ObjectReader` 解析 frozen Draft/Tool projections;Tool-specific adapter 生成 typed snapshot,不透传 JsonNode/raw response。
- Spring AI `ChatModel.call(Prompt)` 执行无 Tool 单轮 guard;`ChatResponseMetadata.Usage` 接入现有 Core Token 预算。
- Java 17 `ExecutorService/Future.get(timeout)` 提供 per-attempt timeout,`Future.cancel(true)` 接入 Run cancellation;Semantic 总时限由 monotonic elapsed time 控制。
- JUnit 5 scripted `ChatModel` 和 in-memory canonical store 作为外部边界 fake;测试只通过 EvidenceGuard/SemanticGuard/Release use case 公共接口断言行为。
- 无 Controller、MQ、JPA、Flyway 或新 Maven dependency;不需要新共享基础设施。
## Apply Progress
- 首模块对齐:EvidenceGuard typed contracts、Draft/reference checks、current-Run canonical validation 和 RAG/log/MySQL snapshot adapters 已完成,tasks 1.1-1.4 完成。
- 8 个 `EvidenceGuardTest` 行为测试通过;快照序列化不含 Tool Call ID 或 `raw_response`,`NO_EVIDENCE` 仅在 `NEGATIVE_OBSERVATION` 下进入带 scope/zero-match 的快照。
- TODO:tasks 2.1-4.2,尚未实现模型边界、SemanticGuard、repair/release 和综合回归。
- REVIEW 发现 corrupted Store key 下 record ID 可能与 Draft reference 不同;分类为代码偏离,已增加 exact canonical ID 校验和回归测试,无需改变规格方向。
- REVIEW 发现 fallback `DiagnosisReleaseResult` 携带完整 snapshot 会通过 `analysis_text` 间接泄漏 Draft;分类为规格安全边界细化,已回写 design/spec,并让 fallback result 强制使用空 snapshot,安全来源只保留在 `SafeFallback.verified_sources`。
## Apply Result
- 新增确定性 EvidenceGuard:校验 typed Draft、唯一 Analysis ID、报告引用、exact current-Run Tool ID、READY lifecycle、agent result、Evidence Status 与 Analysis Kind,并严格展开 RAG/log/MySQL projection。
- verified snapshot 按 Analysis ID 分组,只含稳定 source/scope/timestamp/excerpt/values;SemanticGuard 输入不含 Tool Call ID、Redis key 或 raw response。
- 新增共享 `GuardModelCall`:复用系统 ChatModel,执行 Core model/Token/Run byte budget、per-attempt timeout、total timeout、Future cancellation 和 late-result active check。
- 新增无 Tool/无记忆/无 ReAct 的单轮 SemanticGuard,严格解析 `SUPPORTED|UNSUPPORTED + reason`;仅技术失败使用既有两次 attempt,业务 UNSUPPORTED 不重试。
- 新增一次 EvidenceRepair,只允许 ID/reference 修复;任何可见文本、kind、顺序、human-confirmation 或 limitations 变化都直接 Evidence fallback。
- 新增 DiagnosisReleaseUseCase 和固定 SafeFallbackFactory:SUPPORTED 原样返回 Draft;evidence failed、semantic unsupported/unavailable 均不返回 Draft、完整 snapshot 或审计 reason。
- 公开 Controller、ChatService、AiOpsService、SSE、JPA/Flyway 和旧多 Agent 链路无修改。
## Apply Verification
- 编译:`mvn -q -DskipTests compile` 通过。
- 阶段 5 focused:`EvidenceGuardTest` 9、`SemanticGuardTest` 4、`DiagnosisReleaseUseCaseTest` 6,共 19 tests,0 failure/error。
- 综合回归:阶段 5 + Harness Core/Retry/Tool boundary/store/projections/contracts + Diagnosis Agent,共 18 suites / 70 tests,0 failure/error/skipped。
- OpenSpec:`openspec validate single-react-evidence-semantic-guards --strict` 通过。
- 静态范围:公开 Controller/ChatService/AiOpsService diff 为空;SemanticGuard 无 Tool ID/raw response/ReactAgent/StateGraph/ThreadLocal/手写 while;新 guard/release 包无业务 Agent loop。
- 仓库级 `openspec validate --specs --strict` 为 16 passed / 2 failed;失败仍是前序 `mysql-readonly-tool`、`rag-log-projections` 的 scenario 标题格式,按计划阶段 7 统一修复。
- 未执行 live E2E:按 ISS-014 阶段门禁统一留到阶段 7。
## Archive Result
- 14/14 OpenSpec tasks 完成,`.archive-ready` 已创建。
- 主规格已同步至 `openspec/specs/single-react-evidence-semantic-guards/spec.md`,新增 9 个 requirements。
- Change 已归档至 `openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards`。
- `devflow/index.md` 和 ISS-014 阶段表已更新为阶段 0-5 archived,下一阶段为 6A。
- 未创建 ADR/compound knowledge:本阶段落实 ISS-014 已冻结边界,没有新的跨项目难逆转决策。
@@ -0,0 +1,29 @@
# Evidence: single-react-evidence-semantic-guards
## Code and Contract Evidence
- `CanonicalInvocationStore` 只提供 key lookup;`ToolCallKeyFactory` 冻结 `runId + toolCallId` 隔离,`CanonicalToolInvocation` 冻结 READY/agent_result/evidence status 可引用条件。
- `RagToolResult`、`QueryLogsToolResult`、`MysqlToolResult` 是 Agent-facing 有界 projection,足以构造 snapshot;只有 MySQL 的逻辑数据源/SQL/params 需要从 canonical request 补足。
- `AnalysisKind.accepts` 已冻结 `NORMAL -> EVIDENCE_FOUND`、`NEGATIVE_OBSERVATION -> NO_EVIDENCE`。
- `HarnessRetryPolicies.strict()` 已冻结 SemanticGuard 两次技术 attempt 和 Evidence repair 一次 attempt。
- `ChatModel.call(Prompt)`、`ChatResponseMetadata.Usage` 和 `RunCancellation.onCancel` 支持直接单轮模型调用、Token 记账与 Future cancellation,无需 ReactAgent。
## Confirmed Boundaries
- EvidenceGuard 不判断证据是否支持结论,只验证结构、引用和物理真实性。
- SemanticGuard 只接收原始 Query、移除 Tool ID 的完整 Draft view 和 verified snapshot,不访问 Redis/raw response。
- Evidence repair 只修改 ID/reference;用户可见语义发生任何变化即失败。
- `UNSUPPORTED` 是有效业务结果,不重试;只有 timeout/transport/parse/schema 技术失败可进行第二次 attempt。
- Fallback 不含 Draft、完整 snapshot 或 SemanticGuard reason;Evidence failure 的 verified sources 为空。
## Review Findings
- 增加 canonical record ID 与 Draft reference 的 exact match,防止 corrupted key 映射被误信任。
- Fallback release result 强制使用空 snapshot,防止 `analysis_text` 通过误序列化泄漏 Draft;安全来源只保留在 `SafeFallback.verified_sources`。
## Verification Evidence
- Stage-focused: 19 tests,覆盖 EvidenceGuard 9、SemanticGuard 4、DiagnosisReleaseUseCase 6。
- Regression: 18 suites / 70 tests,0 failure/error/skipped。
- Maven compile 和 change strict validation 通过。
- 公开 Controller/ChatService/AiOpsService 零 diff;SemanticGuard 无 Tool ID/raw/ReactAgent/Graph/ThreadLocal/手写 loop。
@@ -1,6 +1,6 @@
# ISS-014 单体 ReAct Agent、Harness 与 ACI 工具瘦身 # ISS-014 单体 ReAct Agent、Harness 与 ACI 工具瘦身
**状态**:实施中(阶段 0-4 已归档,下一阶段 5) **状态**:实施中(阶段 0-5 已归档,下一阶段 6A)
**严重程度**:高 **严重程度**:高
**发现时间**:2026-07-20 **发现时间**:2026-07-20
**目标分支**:`refactor/chat-single-react-harness` **目标分支**:`refactor/chat-single-react-harness`
@@ -1045,7 +1045,7 @@ ISS-014 是总设计 Issue,不创建跨阶段共享的 OpenSpec change。以
| 3B | `single-react-rag-log-projections` | Completed;已归档 | | 3B | `single-react-rag-log-projections` | Completed;已归档 |
| 3C | `single-react-mysql-readonly-tool` | Completed;已归档 | | 3C | `single-react-mysql-readonly-tool` | Completed;已归档 |
| 4 | `single-react-diagnosis-agent` | Completed;已归档 | | 4 | `single-react-diagnosis-agent` | Completed;已归档 |
| 5 | `single-react-evidence-semantic-guards` | Pending | | 5 | `single-react-evidence-semantic-guards` | Completed;已归档 |
| 6A | `single-react-chat-application-usecase` | Pending | | 6A | `single-react-chat-application-usecase` | Pending |
| 6B | `single-react-chat-sse-cutover` | Pending | | 6B | `single-react-chat-sse-cutover` | Pending |
| 7 | `single-react-cleanup-e2e` | Pending | | 7 | `single-react-cleanup-e2e` | Pending |
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-21
@@ -0,0 +1,118 @@
## Context
阶段 2-4 已提供显式 `RunContext`、类型化重试策略、canonical invocation store、三类有界 Tool projection 和内部 `DiagnosisAgentUseCase`。当前 `DiagnosisDraft` 仍只是未发布草稿:类型反序列化不保证 ID 唯一、报告内部引用闭合,也不证明 Tool Call 属于当前 Run 或 projection 可引用;即使物理证据真实,也还需要独立判断整份报告是否被证据支持。
本阶段只建立内部 release boundary。公开 `ChatController`、SSE、Run persistence 和旧 Planner/Executor/Verifier/Composer 链路由阶段 6A/6B/7 处理。接口影响为 L2,新增内部 Java use case 供后续阶段消费,不改变当前外部协议。
## Goals / Non-Goals
**Goals:**
- 用确定性代码验证 Draft 结构、Analysis 引用和当前 Run canonical invocation。
- 只从 `agent_result` 和必要的 canonical request 生成无 Tool ID、无 raw response 的 verified evidence snapshot。
- 首次验真失败时允许一次保持用户可见语义不变的结构修复。
- 使用同一 `ChatModel` 执行无 Tool、无记忆、无 ReAct 的隔离 SemanticGuard,并由 Harness 控制预算、超时、取消和技术重试。
- 只发布通过两层门禁的原 Draft;所有其他可降级路径返回固定 `SafeFallback`。
**Non-Goals:**
- 不切换 `/api/chat`、`/api/chat_stream` 或 `/api/ai_ops`。
- 不实现 SSE 事件、Run 持久化、PreviousTurn、Intent Router 或最终报告渲染。
- 不删除或改造旧 Gatekeeper、Verifier、Composer 和多 Agent 生产链路。
- 不访问 canonical `raw_response`,不新增 Redis 索引或历史证据召回。
- 不执行 live model/Redis/MySQL E2E;最终 live 验收留到阶段 7。
## Decisions
### 1. EvidenceGuard 是 fail-closed 的纯确定性边界
`EvidenceGuard.validate(RunContext, DiagnosisDraft)` 分两步执行:先验证 Draft 顶层/嵌套结构、非空文本、Analysis ID 唯一性和 `based_on_analysis_ids` 闭合;再对每个 Analysis 的非空 Tool Call ID 使用 `ToolCallKeyFactory.create(runId, id)` 查询 `CanonicalInvocationStore`。
可引用记录必须同时满足:record `runId` 等于当前 Run、`status=READY`、`agent_result` 非空、`evidence_status=EVIDENCE_FOUND|NO_EVIDENCE`。`AnalysisKind.accepts` 固定 `NORMAL -> EVIDENCE_FOUND`、`NEGATIVE_OBSERVATION -> NO_EVIDENCE`。任何非法 key、缺失记录、跨 Run、错误/投影中记录、空 projection、状态不一致或解析错误都产生稳定 `EvidenceViolation`,不会抛出原始存储内容或进入 SemanticGuard。
替代方案是复用旧 `ExecutorGatekeeperService`;拒绝该方案,因为旧实现绑定 JPA `ToolInvocation`、Session/Run 双入口和 Map DTO,不能表达 Redis canonical lifecycle、`NO_EVIDENCE` kind 约束或新的 typed Draft。
### 2. verified snapshot 使用 Tool-specific 严格 projection adapter
Guard 根据冻结的 `AgentToolContracts` 严格解析 `RagToolResult`、`QueryLogsToolResult`、`MysqlToolResult`,并校验 projection 内 `tool_call_id`、`evidence_status` 与 canonical record 一致。未知 Tool 一律拒绝。
快照结构为 `analysis_id + analysis_text + analysis_kind + verified_evidence[]`。单条 evidence 只包含 `source_type/source/scope/timestamp/excerpt/values`:RAG 展开稳定 document metadata/excerpt;日志展开 pattern/event 或带精确 query/time window 的零结果;MySQL 从 canonical request 取得逻辑 data source/SQL/params scope,从 projection 取得 columns/有界 rows 或零结果。快照不包含 Tool Call ID、Redis key、raw response 或未被 Draft 引用的调用。
Snapshot 同时可确定性提取去重的 `SafeFallback.VerifiedSource`,因此 release policy 无需读取 Store。替代通用 JsonNode 透传;拒绝该方案,因为它会扩大模型输入面并削弱字段级测试。
### 3. Evidence repair 只能改引用结构,不能成为报告作者
首次 Guard 失败后,`EvidenceRepair` 使用直接单轮 `ChatModel.call(Prompt)`,输入为原始 query、完整 Draft 和稳定 violation code/target;没有 Tool callback、memory 或 ReAct Agent。调用由 `context.retryPolicies().evidenceRepair()` 执行,其 `maxAttempts=1`,因此不能形成隐式 retry loop。
修复输出经过严格 `DiagnosisDraft` 解析后,转换为移除所有标识/引用字段的 `SemanticDraftView` 并与原 Draft 比较。Conclusion/Analysis/Action/Recommendation/Limitations 的文本、kind、顺序和人工确认标志有任何变化都拒绝;只允许 `analysis_id`、`tool_call_ids` 和 `based_on_analysis_ids` 修正。修复后完整重跑 EvidenceGuard,第二次仍失败直接 `EVIDENCE_VALIDATION_FAILED`。
替代方案是允许模型删除无效 Analysis 或改写结论;拒绝该方案,因为 ISS-014 固定 Diagnosis Agent 是唯一报告作者,Harness/repair 不能改变用户报告语义。
### 4. SemanticGuard 使用直接 ChatModel 单轮调用
`SemanticGuardInput` 包含原始 query、`SemanticDraftView` 和 verified snapshot。View 保留完整用户可见 Draft 语义和 Analysis IDs,但删除所有 Tool Call IDs。每个 attempt 都构造只含一个 system message 和一个 user JSON message 的新 `Prompt`,直接调用系统注入的同一 `ChatModel`;不构建 `ReactAgent`,因此没有 Tool、memory、checkpoint、Graph 或 Agent 回调。
输出只允许精确 JSON object `{verdict, reason}`,字段集合固定,verdict 只允许 `SUPPORTED|UNSUPPORTED`,reason 必须非空。SemanticGuard 不返回 corrected report。`UNSUPPORTED` 正常返回给 release policy,不能进入 retry classifier。
### 5. 单轮模型执行由共享受控边界管理
`GuardModelCall` 在每次调用前执行 `DiagnosisHarnessCore.beforeModelCall`,在模型实际返回后从 `ChatResponseMetadata.Usage` 记录 input/output Token,并再次检查 Run active。调用提交到注入的 `ExecutorService`,等待时间为 `min(singleAttemptTimeout, totalTimeoutRemaining)`;attempt 超时或 Run cancellation 时执行 `Future.cancel(true)`。模型即使忽略中断而迟到返回,后台任务仍记录真实 Token,并在 `core.checkActive` 处阻止迟到结果被使用。
SemanticGuard 通过既有 `HarnessRetryExecutor` 和 `context.retryPolicies().semanticGuard()` 执行。只有 timeout、transport、parse error、schema invalid 可进行第二次 attempt;两次使用完全相同的序列化输入。Run cancelled/budget exhausted 不转换为 Fallback,而是继续抛给阶段 6A 生命周期边界;其他最终技术失败映射 `SEMANTIC_UNAVAILABLE`。
### 6. Release policy 不携带未验证语义
`DiagnosisReleaseUseCase.execute(context, query, draft)` 固定顺序为 `EvidenceGuard -> optional repair -> EvidenceGuard -> SemanticGuard -> decision`。成功结果保留与 Guard 通过的同一个 Draft 实例/修复实例,不总结、裁剪或重排。
- `SUPPORTED`:`ReleaseOutcome.SUCCESS`,返回完整 Draft 和 verified snapshot。
- `UNSUPPORTED`:`ReleaseOutcome.FALLBACK` + `SEMANTIC_UNSUPPORTED`。
- SemanticGuard 最终技术失败:`ReleaseOutcome.FALLBACK` + `SEMANTIC_UNAVAILABLE`。
- 首次修复调用失败、语义被改写或第二次 Guard 失败:`ReleaseOutcome.FALLBACK` + `EVIDENCE_VALIDATION_FAILED`。
`SafeFallbackFactory` 使用固定模板。Evidence failure 的 `verified_sources=[]`;Semantic fallback 只从 snapshot 提取来源。所有 Fallback 的 `conclusion=null`,且不接收 Draft 文本、SemanticGuard reason、供应商错误或异常作为模板参数。
Fallback release result 不保留完整 snapshot,因为 snapshot 含 Analysis text;已验真来源只以 `SafeFallback.verified_sources` 的安全子集返回,防止误序列化内部 result 泄漏草稿。
## Module and Ownership Audit
```text
query + DiagnosisDraft + RunContext
-> DiagnosisReleaseUseCase
-> EvidenceGuard -> CanonicalInvocationStore + Tool projection parsers
-> optional EvidenceRepair -> GuardModelCall -> shared ChatModel
-> SemanticGuardInput(snapshot + ID-free Draft)
-> SemanticGuard -> HarnessRetryExecutor -> GuardModelCall -> shared ChatModel
-> original Draft OR SafeFallbackFactory
```
- `RunContext` owns deadline、cancellation、budget、retry policy 和 lifecycle;本阶段不创建第二套 Run 状态。
- canonical store owns Tool 调用真实性;EvidenceGuard 只读并生成隔离 snapshot,不持有 Redis client。
- Diagnosis Agent owns report semantics;repair 只修标识结构,SemanticGuard 只审查,release 只选择 Draft 或模板。
- 最大耦合风险是阻塞式 `ChatModel` 的取消能力;受控 Future 可发出 interrupt,但 provider 是否立即终止取决于 SDK,故迟到结果必须由 Core active check 丢弃。
- 无新 ADR:模块边界均来自 ISS-014 已冻结设计,且公开消费者尚未切换。
## Interface Impact
- Level: L2 internal interface。
- 新增内部 records/interfaces/use cases;不修改现有 Java 方法签名。
- 当前消费者仅 focused tests;阶段 6A 将装配真实 `DiagnosisAgentUseCase -> DiagnosisReleaseUseCase`。
- HTTP/SSE、Controller DTO、JPA/Flyway、Redis key schema、旧 Chat/AiOps 行为均不变,无迁移或回滚数据操作。
## Risks / Trade-offs
- [Projection 字段遗漏导致误判] -> 三类显式 adapter、ID/status 双重校验和 contract fixture tests。
- [Repair 改写报告] -> `SemanticDraftView` 确定性相等检查,任何可见差异直接 Evidence fallback。
- [模型忽略 interrupt] -> Future cancel、迟到 Token 记录、Core active check;不允许迟到结果进入 release。
- [业务 `UNSUPPORTED` 被重试] -> verdict 作为正常返回值,classifier 只处理异常。
- [Fallback 泄漏草稿或内部 reason] -> 固定 factory 不接受这些参数,并对序列化结果做负向断言。
- [输入快照过大] -> repair/SemanticGuard 独立 UTF-8 input/output limits,并占用 Run byte/model/token budget。
## Migration Plan
1. 新增内部 guard/release 包和聚焦测试,不连接公开入口。
2. 阶段 6A 将内部 Chat application use case 装配到这些接口。
3. 阶段 6B 在新 release boundary 通过后原子切换公开 SSE。
4. 如阶段 5 回滚,只需移除新增内部包;阶段 4 和旧生产链路不受影响。
## Open Questions
- None。单次/总超时与字节预算由构造配置提供可调默认值,具体生产数值按阶段 7 Trace 校准,不构成本阶段方向问题。
@@ -0,0 +1,30 @@
## Why
阶段 4 已能生成带框架 Tool Call ID 的 `DiagnosisDraft`,但草稿尚未经过当前 Run 证据验真和隔离语义审查,不能安全发布。阶段 5 需要补齐确定性 EvidenceGuard、单轮 SemanticGuard 和 fail-closed 释放策略,作为阶段 6A/6B 接入公开链路的前置门禁。
## What Changes
- 新增确定性 EvidenceGuard,校验 Draft 结构、Analysis ID、报告内部引用,以及当前 Run canonical Tool invocation 的生命周期、证据状态和 `agent_result`。
- 从三类冻结 Tool projection 确定性生成按 Analysis 分组的 verified evidence snapshot;不读取 `raw_response`,不向 SemanticGuard 暴露 Tool Call ID 或 Redis 细节。
- 首次 EvidenceGuard 失败时允许一次无 Tool、无 ReAct 的结构修复;修复后仍失败直接返回 `EVIDENCE_VALIDATION_FAILED`。
- 新增复用系统同一 `ChatModel` 的隔离单轮 SemanticGuard,校验完整报告语义,只输出 `SUPPORTED|UNSUPPORTED + reason`,不生成或修改报告。
- Harness 对 SemanticGuard 执行输入预算、模型/Token 预算、单次超时、取消、严格输出解析和最多两次技术 attempt;`UNSUPPORTED` 不重试。
- 新增 release use case:`SUPPORTED` 原样释放 Draft,业务不支持或技术不可用返回固定 SafeFallback,任何失败均不泄漏 Draft 或 SemanticGuard reason。
- 不切换公开 Chat/AiOps,不删除旧 Gatekeeper/Verifier/Composer,不修改 SSE 或持久化协议。
## Capabilities
### New Capabilities
- `single-react-evidence-semantic-guards`: 定义 DiagnosisDraft 的物理验真、verified snapshot、单次结构修复、隔离语义审查和安全释放行为。
### Modified Capabilities
- None. 既有 Agent、Tool、Harness Core 和公开 Chat capability 的需求语义不在本阶段改变。
## Impact
- 新增 `com.superbiz.agent.harness.guard.evidence`、`guard.semantic` 和 `release` 内部包,以及 SemanticGuard prompt 和 focused tests。
- 复用 `CanonicalInvocationStore`、`ToolCallKeyFactory`、三类 Tool contract、`DiagnosisHarnessCore`、`HarnessRetryExecutor`、`RunContext`、`DiagnosisDraft`、`SafeFallback` 和系统 `ChatModel`。
- 内部接口影响为 L2:阶段 6A 将消费新的 release use case;本阶段无公开 Controller、SSE、DTO、数据库或旧 ChatService 行为变化。
- 主要风险是快照字段遗漏、超时任务残留、错误重试业务 `UNSUPPORTED`,以及 Fallback 意外携带未验证内容;设计和测试必须逐项 fail closed。
@@ -0,0 +1,112 @@
## ADDED Requirements
### Requirement: Deterministic Draft and evidence validation
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
#### Scenario: Duplicate or missing Analysis ID
- **WHEN** a Draft contains a blank or duplicate Analysis ID
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
#### Scenario: Broken report reference
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
- **THEN** EvidenceGuard rejects the Draft before semantic review
### Requirement: Current Run canonical invocation ownership
EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`.
#### Scenario: Fabricated or cross-Run Tool Call
- **WHEN** a Draft references a missing Tool Call or the resolved record belongs to another Run
- **THEN** EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard
#### Scenario: Failed or incomplete invocation
- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR`
- **THEN** EvidenceGuard rejects the Draft
### Requirement: Analysis kind matches evidence semantics
EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls.
#### Scenario: Valid negative observation
- **WHEN** a `NEGATIVE_OBSERVATION` references a READY `NO_EVIDENCE` projection with its query scope and zero-match data
- **THEN** EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion
#### Scenario: Positive claim uses no-evidence result
- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation
- **THEN** EvidenceGuard rejects the binding
### Requirement: Verified evidence snapshot is minimal and deterministic
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
#### Scenario: Supported RAG, log, and MySQL projections
- **WHEN** a Draft references valid RAG, log, or MySQL calls
- **THEN** the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis
#### Scenario: Projection contract mismatch
- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
- **THEN** EvidenceGuard fails closed
### Requirement: Evidence repair is single-turn and semantics-preserving
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
#### Scenario: Structural repair succeeds
- **WHEN** the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard
- **THEN** the repaired Draft proceeds to SemanticGuard
#### Scenario: Repair changes report text
- **WHEN** the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED`
#### Scenario: Second validation fails
- **WHEN** the repaired Draft still fails EvidenceGuard
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard
### Requirement: Isolated single-turn SemanticGuard
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
#### Scenario: Semantic input isolation
- **WHEN** a verified Draft enters SemanticGuard
- **THEN** the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition
#### Scenario: Binary review output
- **WHEN** SemanticGuard completes normally
- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report
### Requirement: Semantic model budgets timeout cancellation and retry
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
#### Scenario: Technical failure then success
- **WHEN** the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict
- **THEN** exactly two model attempts are recorded and the second verdict controls release
#### Scenario: Unsupported is not retried
- **WHEN** SemanticGuard returns valid `UNSUPPORTED`
- **THEN** the Harness records one attempt and immediately applies the unsupported fallback
#### Scenario: Run cancellation during model call
- **WHEN** the Run is cancelled while a guard model call is pending
- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
### Requirement: Fail-closed release policy
The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
#### Scenario: Supported report release
- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED`
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
#### Scenario: Unsupported report fallback
- **WHEN** SemanticGuard returns `UNSUPPORTED`
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot
#### Scenario: Semantic review remains unavailable
- **WHEN** all permitted technical attempts fail
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
#### Scenario: Evidence validation fallback sources
- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
### Requirement: Stage-five public isolation
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.
#### Scenario: Focused implementation scope
- **WHEN** stage-five changes are inspected
- **THEN** only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed
@@ -0,0 +1,25 @@
## 1. Evidence validation and snapshot
- [x] 1.1 Add typed EvidenceViolation, EvidenceGuardResult and verified snapshot contracts with immutable source extraction.
- [x] 1.2 Implement Draft structure/reference validation and current-Run canonical invocation checks.
- [x] 1.3 Implement strict RAG/log/MySQL projection adapters that build ID-free, raw-free evidence grouped by Analysis ID.
- [x] 1.4 Add focused tests for duplicate/missing IDs, empty/broken references, fabricated/cross-Run calls, invalid lifecycle/result, kind/status mismatch and valid negative observations.
## 2. Guard model boundary and semantic review
- [x] 2.1 Add the shared single-turn GuardModelCall with Core model/Token budgets, UTF-8 limits, per-attempt timeout, total timeout support and Future cancellation.
- [x] 2.2 Add SemanticDraftView and SemanticGuardInput that preserve all user-visible report semantics while excluding Tool Call IDs.
- [x] 2.3 Implement strict binary SemanticGuard output parsing and HarnessRetryExecutor integration with stable attempt auditing.
- [x] 2.4 Add tests proving input isolation, same-model direct calls, retryable technical failures, non-retried UNSUPPORTED, timeout cancellation and no Tool/ReAct loop.
## 3. Evidence repair and release policy
- [x] 3.1 Implement one-attempt no-Tool EvidenceRepair and reject any repair that changes SemanticDraftView.
- [x] 3.2 Implement fixed SafeFallbackFactory and release result contracts without Draft/reason parameters on fallback paths.
- [x] 3.3 Implement DiagnosisReleaseUseCase sequencing Guard, optional repair, SemanticGuard and fail-closed outcomes while propagating cancellation/budget exhaustion.
- [x] 3.4 Add release tests for successful repair, second validation failure, semantic support/unsupported/unavailable, fallback types, empty evidence-failure sources and no Draft/reason leakage.
## 4. Verification and scope
- [x] 4.1 Run stage-five focused tests plus Harness/Tool/Agent regression tests and Maven compile.
- [x] 4.2 Verify strict OpenSpec validation, no public Controller/ChatService/AiOps diff, no raw response/Tool ID in semantic input, and no hidden retry or handwritten Agent loop.
@@ -0,0 +1,115 @@
# single-react-evidence-semantic-guards Specification
## Purpose
TBD - created by archiving change single-react-evidence-semantic-guards. Update Purpose after archive.
## Requirements
### Requirement: Deterministic Draft and evidence validation
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
#### Scenario: Duplicate or missing Analysis ID
- **WHEN** a Draft contains a blank or duplicate Analysis ID
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
#### Scenario: Broken report reference
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
- **THEN** EvidenceGuard rejects the Draft before semantic review
### Requirement: Current Run canonical invocation ownership
EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`.
#### Scenario: Fabricated or cross-Run Tool Call
- **WHEN** a Draft references a missing Tool Call or the resolved record belongs to another Run
- **THEN** EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard
#### Scenario: Failed or incomplete invocation
- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR`
- **THEN** EvidenceGuard rejects the Draft
### Requirement: Analysis kind matches evidence semantics
EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls.
#### Scenario: Valid negative observation
- **WHEN** a `NEGATIVE_OBSERVATION` references a READY `NO_EVIDENCE` projection with its query scope and zero-match data
- **THEN** EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion
#### Scenario: Positive claim uses no-evidence result
- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation
- **THEN** EvidenceGuard rejects the binding
### Requirement: Verified evidence snapshot is minimal and deterministic
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
#### Scenario: Supported RAG, log, and MySQL projections
- **WHEN** a Draft references valid RAG, log, or MySQL calls
- **THEN** the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis
#### Scenario: Projection contract mismatch
- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
- **THEN** EvidenceGuard fails closed
### Requirement: Evidence repair is single-turn and semantics-preserving
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
#### Scenario: Structural repair succeeds
- **WHEN** the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard
- **THEN** the repaired Draft proceeds to SemanticGuard
#### Scenario: Repair changes report text
- **WHEN** the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED`
#### Scenario: Second validation fails
- **WHEN** the repaired Draft still fails EvidenceGuard
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard
### Requirement: Isolated single-turn SemanticGuard
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
#### Scenario: Semantic input isolation
- **WHEN** a verified Draft enters SemanticGuard
- **THEN** the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition
#### Scenario: Binary review output
- **WHEN** SemanticGuard completes normally
- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report
### Requirement: Semantic model budgets timeout cancellation and retry
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
#### Scenario: Technical failure then success
- **WHEN** the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict
- **THEN** exactly two model attempts are recorded and the second verdict controls release
#### Scenario: Unsupported is not retried
- **WHEN** SemanticGuard returns valid `UNSUPPORTED`
- **THEN** the Harness records one attempt and immediately applies the unsupported fallback
#### Scenario: Run cancellation during model call
- **WHEN** the Run is cancelled while a guard model call is pending
- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
### Requirement: Fail-closed release policy
The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
#### Scenario: Supported report release
- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED`
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
#### Scenario: Unsupported report fallback
- **WHEN** SemanticGuard returns `UNSUPPORTED`
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot
#### Scenario: Semantic review remains unavailable
- **WHEN** all permitted technical attempts fail
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
#### Scenario: Evidence validation fallback sources
- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
### Requirement: Stage-five public isolation
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.
#### Scenario: Focused implementation scope
- **WHEN** stage-five changes are inspected
- **THEN** only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed
@@ -0,0 +1,435 @@
package com.superbiz.agent.harness.guard.evidence;
import com.fasterxml.jackson.databind.DeserializationFeature;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.ObjectReader;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.tool.contract.AgentToolContracts;
import com.superbiz.agent.harness.tool.contract.LogEvent;
import com.superbiz.agent.harness.tool.contract.LogPattern;
import com.superbiz.agent.harness.tool.contract.LogQueryScope;
import com.superbiz.agent.harness.tool.contract.MysqlToolRequest;
import com.superbiz.agent.harness.tool.contract.MysqlToolResult;
import com.superbiz.agent.harness.tool.contract.QueryLogsToolResult;
import com.superbiz.agent.harness.tool.contract.RagEvidence;
import com.superbiz.agent.harness.tool.contract.RagToolResult;
import com.superbiz.agent.harness.tool.store.CanonicalInvocationStore;
import com.superbiz.agent.harness.tool.store.CanonicalToolInvocation;
import com.superbiz.agent.harness.tool.store.ToolCallKeyFactory;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
public final class EvidenceGuard {
private final CanonicalInvocationStore store;
private final ToolCallKeyFactory keyFactory;
private final ObjectReader ragReader;
private final ObjectReader logsReader;
private final ObjectReader mysqlRequestReader;
private final ObjectReader mysqlResultReader;
public EvidenceGuard(CanonicalInvocationStore store,
ToolCallKeyFactory keyFactory,
ObjectMapper objectMapper) {
this.store = Objects.requireNonNull(store, "store must not be null");
this.keyFactory = Objects.requireNonNull(keyFactory, "keyFactory must not be null");
ObjectMapper mapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
this.ragReader = mapper.readerFor(RagToolResult.class)
.with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
.with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
this.logsReader = mapper.readerFor(QueryLogsToolResult.class)
.with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
.with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
this.mysqlRequestReader = mapper.readerFor(MysqlToolRequest.class)
.with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
.with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
this.mysqlResultReader = mapper.readerFor(MysqlToolResult.class)
.with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
.with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
}
public EvidenceGuardResult validate(RunContext context, DiagnosisDraft draft) {
Objects.requireNonNull(context, "context must not be null");
List<EvidenceViolation> violations = validateDraft(draft);
if (!violations.isEmpty()) {
return EvidenceGuardResult.invalid(violations);
}
List<VerifiedAnalysisEvidence> verifiedAnalyses = new ArrayList<>();
for (int index = 0; index < draft.analysis().size(); index++) {
DiagnosisDraft.AnalysisItem analysis = draft.analysis().get(index);
List<VerifiedEvidence> evidence = new ArrayList<>();
for (String toolCallId : analysis.toolCallIds()) {
verifyInvocation(context, analysis, index, toolCallId, evidence, violations);
}
if (!evidence.isEmpty()) {
verifiedAnalyses.add(new VerifiedAnalysisEvidence(
analysis.analysisId(), analysis.text(), analysis.kind(), evidence));
}
}
return violations.isEmpty()
? EvidenceGuardResult.valid(new VerifiedEvidenceSnapshot(verifiedAnalyses))
: EvidenceGuardResult.invalid(violations);
}
private List<EvidenceViolation> validateDraft(DiagnosisDraft draft) {
List<EvidenceViolation> violations = new ArrayList<>();
if (draft == null) {
violations.add(violation(EvidenceViolationCode.DRAFT_MISSING, "draft"));
return violations;
}
if (draft.analysis().isEmpty()) {
violations.add(violation(EvidenceViolationCode.ANALYSIS_MISSING, "analysis"));
}
Set<String> ids = new HashSet<>();
for (int index = 0; index < draft.analysis().size(); index++) {
DiagnosisDraft.AnalysisItem item = draft.analysis().get(index);
String target = "analysis[" + index + "]";
if (item == null) {
violations.add(violation(EvidenceViolationCode.ANALYSIS_MISSING, target));
continue;
}
if (isBlank(item.analysisId())) {
violations.add(violation(EvidenceViolationCode.ANALYSIS_ID_MISSING, target + ".analysis_id"));
} else if (!ids.add(item.analysisId())) {
violations.add(violation(EvidenceViolationCode.ANALYSIS_ID_DUPLICATE, target + ".analysis_id"));
}
if (item.kind() == null) {
violations.add(violation(EvidenceViolationCode.ANALYSIS_KIND_MISSING, target + ".kind"));
}
if (isBlank(item.text())) {
violations.add(violation(EvidenceViolationCode.ANALYSIS_TEXT_MISSING, target + ".text"));
}
if (item.toolCallIds().isEmpty()) {
violations.add(violation(EvidenceViolationCode.TOOL_REFERENCE_MISSING, target + ".tool_call_ids"));
}
}
validateReportReferences(draft, ids, violations);
return violations;
}
private void validateReportReferences(DiagnosisDraft draft, Set<String> ids,
List<EvidenceViolation> violations) {
if (draft.conclusion() != null) {
validateTextAndReferences("conclusion", draft.conclusion().text(),
draft.conclusion().basedOnAnalysisIds(), ids, violations);
}
for (int index = 0; index < draft.actionPlan().size(); index++) {
DiagnosisDraft.ActionPlanItem item = draft.actionPlan().get(index);
String target = "action_plan[" + index + "]";
if (item == null) {
violations.add(violation(EvidenceViolationCode.REPORT_TEXT_MISSING, target));
} else {
validateTextAndReferences(target, item.action(), item.basedOnAnalysisIds(), ids, violations);
}
}
for (int index = 0; index < draft.recommendations().size(); index++) {
DiagnosisDraft.Recommendation item = draft.recommendations().get(index);
String target = "recommendations[" + index + "]";
if (item == null) {
violations.add(violation(EvidenceViolationCode.REPORT_TEXT_MISSING, target));
} else {
validateTextAndReferences(target, item.text(), item.basedOnAnalysisIds(), ids, violations);
}
}
if (draft.limitations() == null || isBlank(draft.limitations().scope())) {
violations.add(violation(EvidenceViolationCode.LIMITATIONS_MISSING, "limitations"));
}
}
private void validateTextAndReferences(String target, String text, List<String> references,
Set<String> ids, List<EvidenceViolation> violations) {
if (isBlank(text)) {
violations.add(violation(EvidenceViolationCode.REPORT_TEXT_MISSING, target));
}
if (references == null || references.isEmpty()) {
violations.add(violation(EvidenceViolationCode.ANALYSIS_REFERENCE_MISSING,
target + ".based_on_analysis_ids"));
return;
}
for (int index = 0; index < references.size(); index++) {
String reference = references.get(index);
if (isBlank(reference)) {
violations.add(violation(EvidenceViolationCode.ANALYSIS_REFERENCE_MISSING,
target + ".based_on_analysis_ids[" + index + "]"));
} else if (!ids.contains(reference)) {
violations.add(violation(EvidenceViolationCode.ANALYSIS_REFERENCE_UNKNOWN,
target + ".based_on_analysis_ids[" + index + "]"));
}
}
}
private void verifyInvocation(RunContext context, DiagnosisDraft.AnalysisItem analysis,
int analysisIndex, String toolCallId,
List<VerifiedEvidence> evidence,
List<EvidenceViolation> violations) {
String target = "analysis[" + analysisIndex + "].tool_call_ids";
if (isBlank(toolCallId)) {
violations.add(violation(EvidenceViolationCode.TOOL_REFERENCE_INVALID, target));
return;
}
String key;
try {
key = keyFactory.create(context.runId(), toolCallId);
} catch (IllegalArgumentException exception) {
violations.add(violation(EvidenceViolationCode.TOOL_REFERENCE_INVALID, target));
return;
}
CanonicalToolInvocation invocation;
try {
invocation = store.find(key).orElse(null);
} catch (RuntimeException exception) {
violations.add(violation(EvidenceViolationCode.CANONICAL_LOOKUP_FAILED, target));
return;
}
if (invocation == null) {
violations.add(violation(EvidenceViolationCode.INVOCATION_MISSING, target));
return;
}
if (!Objects.equals(toolCallId, invocation.toolCallId())) {
violations.add(violation(EvidenceViolationCode.INVOCATION_ID_MISMATCH, target));
return;
}
if (!invocation.isReferencableBy(context.runId()) || invocation.agentResult().isBlank()) {
violations.add(violation(EvidenceViolationCode.INVOCATION_NOT_REFERENCABLE, target));
return;
}
if (!analysis.kind().accepts(invocation.evidenceStatus())) {
violations.add(violation(EvidenceViolationCode.EVIDENCE_KIND_MISMATCH, target));
return;
}
switch (invocation.toolName()) {
case AgentToolContracts.LOOKUP_KNOWLEDGE ->
readRag(invocation, evidence, target, violations);
case AgentToolContracts.QUERY_LOGS ->
readLogs(invocation, evidence, target, violations);
case AgentToolContracts.QUERY_MYSQL ->
readMysql(invocation, evidence, target, violations);
default -> violations.add(violation(EvidenceViolationCode.TOOL_UNSUPPORTED, target));
}
}
private void readRag(CanonicalToolInvocation invocation, List<VerifiedEvidence> evidence,
String target, List<EvidenceViolation> violations) {
RagToolResult result;
try {
result = ragReader.readValue(invocation.agentResult());
} catch (Exception exception) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
if (!Objects.equals(result.toolCallId(), invocation.toolCallId())) {
violations.add(violation(EvidenceViolationCode.PROJECTION_ID_MISMATCH, target));
return;
}
if (result.evidenceStatus() != invocation.evidenceStatus()) {
violations.add(violation(EvidenceViolationCode.PROJECTION_STATUS_MISMATCH, target));
return;
}
if (result.returnedCount() != result.evidence().size()) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
if (result.evidenceStatus() == com.superbiz.agent.harness.contract.EvidenceStatus.NO_EVIDENCE) {
if (!result.evidence().isEmpty() || result.returnedCount() != 0) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
evidence.add(new VerifiedEvidence(
"RAG", "KNOWLEDGE_BASE", firstText(result.query(), "knowledge query"),
null, "No evidence matched the query scope",
Map.of("match_count", 0, "truncated", result.truncated())));
return;
}
if (result.evidence().isEmpty()) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
for (RagEvidence item : result.evidence()) {
if (item == null || isBlank(item.documentId()) || isBlank(item.excerpt())) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
Map<String, Object> values = new LinkedHashMap<>();
values.put("document_id", item.documentId());
putIfText(values, "title", item.title());
putIfText(values, "breadcrumb", item.breadcrumb());
evidence.add(new VerifiedEvidence(
"RAG", firstText(item.source(), item.title(), item.documentId()),
firstText(result.query(), "knowledge query"), null, item.excerpt(), values));
}
}
private void readLogs(CanonicalToolInvocation invocation, List<VerifiedEvidence> evidence,
String target, List<EvidenceViolation> violations) {
QueryLogsToolResult result;
try {
result = logsReader.readValue(invocation.agentResult());
} catch (Exception exception) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
if (!projectionMatches(invocation, result.toolCallId(), result.evidenceStatus(),
target, violations)) {
return;
}
if (result.sourceKind() == null || result.scope() == null
|| result.scope().topic() == null || isBlank(result.scope().query())
|| isBlank(result.scope().startTime()) || isBlank(result.scope().endTime())
|| result.matchCount() < 0 || result.returnedCount() != result.events().size()) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
String source = result.scope().topic().name() + " (" + result.sourceKind().name() + ")";
String scope = logScope(result.scope());
if (result.evidenceStatus() == com.superbiz.agent.harness.contract.EvidenceStatus.NO_EVIDENCE) {
if (result.matchCount() != 0 || !result.patterns().isEmpty() || !result.events().isEmpty()) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
evidence.add(new VerifiedEvidence(
"LOG", source, scope, null, "No log events matched the query scope",
Map.of("match_count", 0, "truncated", result.truncated())));
return;
}
if (result.patterns().isEmpty() && result.events().isEmpty()) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
for (LogPattern pattern : result.patterns()) {
if (pattern == null || pattern.count() <= 0 || isBlank(pattern.example())) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
Map<String, Object> values = new LinkedHashMap<>();
values.put("count", pattern.count());
putIfText(values, "first_seen", pattern.firstSeen());
putIfText(values, "last_seen", pattern.lastSeen());
putIfText(values, "level", pattern.level());
putIfText(values, "service", pattern.service());
evidence.add(new VerifiedEvidence(
"LOG", source, scope, pattern.lastSeen(), pattern.example(), values));
}
for (LogEvent event : result.events()) {
if (event == null || isBlank(event.message())) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
Map<String, Object> values = new LinkedHashMap<>();
putIfText(values, "level", event.level());
putIfText(values, "service", event.service());
evidence.add(new VerifiedEvidence(
"LOG", source, scope, event.timestamp(), event.message(), values));
}
}
private boolean projectionMatches(CanonicalToolInvocation invocation, String toolCallId,
com.superbiz.agent.harness.contract.EvidenceStatus evidenceStatus,
String target, List<EvidenceViolation> violations) {
if (!Objects.equals(toolCallId, invocation.toolCallId())) {
violations.add(violation(EvidenceViolationCode.PROJECTION_ID_MISMATCH, target));
return false;
}
if (evidenceStatus != invocation.evidenceStatus()) {
violations.add(violation(EvidenceViolationCode.PROJECTION_STATUS_MISMATCH, target));
return false;
}
return true;
}
private void readMysql(CanonicalToolInvocation invocation, List<VerifiedEvidence> evidence,
String target, List<EvidenceViolation> violations) {
MysqlToolRequest request;
MysqlToolResult result;
try {
request = mysqlRequestReader.readValue(invocation.request());
result = mysqlResultReader.readValue(invocation.agentResult());
} catch (Exception exception) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
if (!projectionMatches(invocation, result.toolCallId(), result.evidenceStatus(),
target, violations)) {
return;
}
if (isBlank(request.dataSource()) || isBlank(request.sql())
|| result.returnedCount() != result.rows().size()
|| result.columns().isEmpty() || hasInvalidColumns(result.columns())) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
String scope = "sql=" + request.sql() + " params=" + request.params();
if (result.evidenceStatus() == com.superbiz.agent.harness.contract.EvidenceStatus.NO_EVIDENCE) {
if (!result.rows().isEmpty() || result.returnedCount() != 0) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
evidence.add(new VerifiedEvidence(
"MYSQL", request.dataSource(), scope, null,
"No rows matched the query scope",
Map.of("match_count", 0, "columns", result.columns(),
"truncated", result.truncated())));
return;
}
if (result.rows().isEmpty()) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
for (int index = 0; index < result.rows().size(); index++) {
Map<String, Object> row = result.rows().get(index);
if (row == null || !row.keySet().equals(new java.util.LinkedHashSet<>(result.columns()))) {
violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target));
return;
}
Map<String, Object> values = new LinkedHashMap<>(row);
values.put("_row_number", index + 1);
evidence.add(new VerifiedEvidence(
"MYSQL", request.dataSource(), scope, null, null, values));
}
}
private static boolean hasInvalidColumns(List<String> columns) {
Set<String> unique = new HashSet<>();
for (String column : columns) {
if (isBlank(column) || !unique.add(column)) {
return true;
}
}
return false;
}
private static String logScope(LogQueryScope scope) {
return scope.topic().name() + " query=" + scope.query()
+ " from=" + scope.startTime() + " to=" + scope.endTime();
}
private static EvidenceViolation violation(EvidenceViolationCode code, String target) {
return new EvidenceViolation(code, target);
}
private static void putIfText(Map<String, Object> values, String name, String value) {
if (!isBlank(value)) {
values.put(name, value);
}
}
private static String firstText(String... values) {
for (String value : values) {
if (!isBlank(value)) {
return value;
}
}
return "unknown";
}
private static boolean isBlank(String value) {
return value == null || value.isBlank();
}
}
@@ -0,0 +1,33 @@
package com.superbiz.agent.harness.guard.evidence;
import java.util.List;
import java.util.Optional;
public record EvidenceGuardResult(
List<EvidenceViolation> violations,
VerifiedEvidenceSnapshot snapshot) {
public EvidenceGuardResult {
violations = violations == null ? List.of() : List.copyOf(violations);
if (violations.isEmpty() == (snapshot == null)) {
throw new IllegalArgumentException(
"valid result requires snapshot and invalid result requires violations");
}
}
public static EvidenceGuardResult valid(VerifiedEvidenceSnapshot snapshot) {
return new EvidenceGuardResult(List.of(), snapshot);
}
public static EvidenceGuardResult invalid(List<EvidenceViolation> violations) {
return new EvidenceGuardResult(violations, null);
}
public boolean valid() {
return violations.isEmpty();
}
public Optional<VerifiedEvidenceSnapshot> verifiedSnapshot() {
return Optional.ofNullable(snapshot);
}
}
@@ -0,0 +1,17 @@
package com.superbiz.agent.harness.guard.evidence;
import com.fasterxml.jackson.annotation.JsonProperty;
import java.util.Objects;
public record EvidenceViolation(
@JsonProperty("code") EvidenceViolationCode code,
@JsonProperty("target") String target) {
public EvidenceViolation {
Objects.requireNonNull(code, "code must not be null");
if (target == null || target.isBlank()) {
throw new IllegalArgumentException("target must not be blank");
}
}
}
@@ -0,0 +1,25 @@
package com.superbiz.agent.harness.guard.evidence;
public enum EvidenceViolationCode {
DRAFT_MISSING,
ANALYSIS_MISSING,
ANALYSIS_ID_MISSING,
ANALYSIS_ID_DUPLICATE,
ANALYSIS_KIND_MISSING,
ANALYSIS_TEXT_MISSING,
TOOL_REFERENCE_MISSING,
REPORT_TEXT_MISSING,
ANALYSIS_REFERENCE_MISSING,
ANALYSIS_REFERENCE_UNKNOWN,
LIMITATIONS_MISSING,
TOOL_REFERENCE_INVALID,
INVOCATION_MISSING,
INVOCATION_ID_MISMATCH,
INVOCATION_NOT_REFERENCABLE,
EVIDENCE_KIND_MISMATCH,
TOOL_UNSUPPORTED,
PROJECTION_INVALID,
PROJECTION_ID_MISMATCH,
PROJECTION_STATUS_MISMATCH,
CANONICAL_LOOKUP_FAILED
}
@@ -0,0 +1,30 @@
package com.superbiz.agent.harness.guard.evidence;
import com.fasterxml.jackson.annotation.JsonProperty;
import com.superbiz.agent.harness.contract.AnalysisKind;
import java.util.List;
import java.util.Objects;
public record VerifiedAnalysisEvidence(
@JsonProperty("analysis_id") String analysisId,
@JsonProperty("analysis_text") String analysisText,
@JsonProperty("analysis_kind") AnalysisKind analysisKind,
@JsonProperty("verified_evidence") List<VerifiedEvidence> evidence) {
public VerifiedAnalysisEvidence {
requireText(analysisId, "analysisId");
requireText(analysisText, "analysisText");
Objects.requireNonNull(analysisKind, "analysisKind must not be null");
evidence = evidence == null ? List.of() : List.copyOf(evidence);
if (evidence.isEmpty()) {
throw new IllegalArgumentException("evidence must not be empty");
}
}
private static void requireText(String value, String name) {
if (value == null || value.isBlank()) {
throw new IllegalArgumentException(name + " must not be blank");
}
}
}
@@ -0,0 +1,35 @@
package com.superbiz.agent.harness.guard.evidence;
import com.fasterxml.jackson.annotation.JsonProperty;
import java.util.LinkedHashMap;
import java.util.Map;
public record VerifiedEvidence(
@JsonProperty("source_type") String sourceType,
@JsonProperty("source") String source,
@JsonProperty("scope") String scope,
@JsonProperty("timestamp") String timestamp,
@JsonProperty("excerpt") String excerpt,
@JsonProperty("values") Map<String, Object> values) {
public VerifiedEvidence {
requireText(sourceType, "sourceType");
requireText(source, "source");
requireText(scope, "scope");
values = immutableValues(values);
}
private static Map<String, Object> immutableValues(Map<String, Object> values) {
if (values == null || values.isEmpty()) {
return Map.of();
}
return java.util.Collections.unmodifiableMap(new LinkedHashMap<>(values));
}
private static void requireText(String value, String name) {
if (value == null || value.isBlank()) {
throw new IllegalArgumentException(name + " must not be blank");
}
}
}
@@ -0,0 +1,33 @@
package com.superbiz.agent.harness.guard.evidence;
import com.fasterxml.jackson.annotation.JsonProperty;
import com.superbiz.agent.harness.contract.SafeFallback;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
public record VerifiedEvidenceSnapshot(
@JsonProperty("analyses") List<VerifiedAnalysisEvidence> analyses) {
public VerifiedEvidenceSnapshot {
analyses = analyses == null ? List.of() : List.copyOf(analyses);
}
public static VerifiedEvidenceSnapshot empty() {
return new VerifiedEvidenceSnapshot(List.of());
}
public List<SafeFallback.VerifiedSource> verifiedSources() {
Map<String, SafeFallback.VerifiedSource> unique = new LinkedHashMap<>();
for (VerifiedAnalysisEvidence analysis : analyses) {
for (VerifiedEvidence evidence : analysis.evidence()) {
String key = evidence.sourceType() + "\u0000" + evidence.source()
+ "\u0000" + evidence.scope();
unique.putIfAbsent(key, new SafeFallback.VerifiedSource(
evidence.sourceType(), evidence.source(), evidence.scope()));
}
}
return List.copyOf(unique.values());
}
}
@@ -0,0 +1,120 @@
package com.superbiz.agent.harness.guard.semantic;
import com.superbiz.agent.harness.core.BudgetExceededException;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunAbortedException;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.retry.RetryFailure;
import org.springframework.ai.chat.messages.AssistantMessage;
import org.springframework.ai.chat.metadata.Usage;
import org.springframework.ai.chat.model.ChatModel;
import org.springframework.ai.chat.model.ChatResponse;
import org.springframework.ai.chat.prompt.Prompt;
import java.nio.charset.StandardCharsets;
import java.time.Duration;
import java.util.Objects;
import java.util.concurrent.CancellationException;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Future;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
public final class GuardModelCall {
private final DiagnosisHarnessCore core;
private final ChatModel chatModel;
private final ExecutorService executor;
public GuardModelCall(DiagnosisHarnessCore core, ChatModel chatModel,
ExecutorService executor) {
this.core = Objects.requireNonNull(core, "core must not be null");
this.chatModel = Objects.requireNonNull(chatModel, "chatModel must not be null");
this.executor = Objects.requireNonNull(executor, "executor must not be null");
}
public String call(RunContext context, Prompt prompt, Duration timeout, long maxOutputBytes) {
Objects.requireNonNull(context, "context must not be null");
Objects.requireNonNull(prompt, "prompt must not be null");
Objects.requireNonNull(timeout, "timeout must not be null");
if (timeout.isZero() || timeout.isNegative() || maxOutputBytes <= 0) {
throw new IllegalArgumentException("timeout and output limit must be positive");
}
core.beforeModelCall(context);
Future<String> future = executor.submit(() -> invoke(context, prompt, maxOutputBytes));
context.cancellation().onCancel(ignored -> future.cancel(true));
try {
return future.get(timeout.toNanos(), TimeUnit.NANOSECONDS);
} catch (TimeoutException exception) {
future.cancel(true);
throw new GuardModelCallException(
RetryFailure.TIMEOUT, "Guard model attempt timed out", exception);
} catch (CancellationException exception) {
core.checkActive(context);
throw new GuardModelCallException(
RetryFailure.TRANSPORT, "Guard model attempt was cancelled", exception);
} catch (InterruptedException exception) {
future.cancel(true);
Thread.currentThread().interrupt();
core.checkActive(context);
throw new GuardModelCallException(
RetryFailure.TRANSPORT, "Interrupted while waiting for guard model", exception);
} catch (ExecutionException exception) {
Throwable cause = exception.getCause();
if (cause instanceof RunAbortedException aborted) {
throw aborted;
}
if (cause instanceof BudgetExceededException exceeded) {
throw exceeded;
}
if (cause instanceof GuardModelCallException guardFailure) {
throw guardFailure;
}
throw new GuardModelCallException(
RetryFailure.TRANSPORT, "Guard model call failed", cause);
}
}
private String invoke(RunContext context, Prompt prompt, long maxOutputBytes) {
ChatResponse response;
try {
response = chatModel.call(prompt);
} catch (RuntimeException exception) {
throw new GuardModelCallException(
RetryFailure.TRANSPORT, "Guard model transport failed", exception);
}
recordUsage(context, response);
core.checkActive(context);
AssistantMessage output = response == null || response.getResult() == null
? null : response.getResult().getOutput();
if (output == null || output.getText() == null || output.getText().isBlank()
|| (output.getToolCalls() != null && !output.getToolCalls().isEmpty())) {
throw new GuardModelCallException(
RetryFailure.SCHEMA_INVALID, "Guard model returned invalid message shape");
}
long bytes = output.getText().getBytes(StandardCharsets.UTF_8).length;
if (bytes > maxOutputBytes) {
throw new GuardModelCallException(
RetryFailure.SCHEMA_INVALID, "Guard model output exceeded limit");
}
core.reserveRunBytes(context, bytes);
return output.getText();
}
private void recordUsage(RunContext context, ChatResponse response) {
if (response == null || response.getMetadata() == null) {
return;
}
Usage usage = response.getMetadata().getUsage();
if (usage == null) {
return;
}
core.recordTokens(context, nonNegative(usage.getPromptTokens()),
nonNegative(usage.getCompletionTokens()));
}
private static long nonNegative(Integer value) {
return value == null || value < 0 ? 0L : value.longValue();
}
}
@@ -0,0 +1,24 @@
package com.superbiz.agent.harness.guard.semantic;
import com.superbiz.agent.harness.retry.RetryFailure;
import java.util.Objects;
public final class GuardModelCallException extends RuntimeException {
private final RetryFailure failure;
public GuardModelCallException(RetryFailure failure, String message) {
super(message);
this.failure = Objects.requireNonNull(failure, "failure must not be null");
}
public GuardModelCallException(RetryFailure failure, String message, Throwable cause) {
super(message, cause);
this.failure = Objects.requireNonNull(failure, "failure must not be null");
}
public RetryFailure failure() {
return failure;
}
}
@@ -0,0 +1,126 @@
package com.superbiz.agent.harness.guard.semantic;
import com.fasterxml.jackson.annotation.JsonProperty;
import com.superbiz.agent.harness.contract.AnalysisKind;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import java.util.List;
import java.util.Objects;
public record SemanticDraftView(
@JsonProperty("conclusion") Conclusion conclusion,
@JsonProperty("analysis") List<Analysis> analysis,
@JsonProperty("action_plan") List<Action> actionPlan,
@JsonProperty("recommendations") List<Recommendation> recommendations,
@JsonProperty("limitations") Limitations limitations) {
public SemanticDraftView {
analysis = immutable(analysis);
actionPlan = immutable(actionPlan);
recommendations = immutable(recommendations);
}
public static SemanticDraftView from(DiagnosisDraft draft) {
Objects.requireNonNull(draft, "draft must not be null");
Conclusion conclusion = draft.conclusion() == null ? null : new Conclusion(
draft.conclusion().text(), draft.conclusion().basedOnAnalysisIds());
List<Analysis> analysis = draft.analysis().stream()
.map(item -> new Analysis(item.analysisId(), item.kind(), item.text()))
.toList();
List<Action> actions = draft.actionPlan().stream()
.map(item -> new Action(item.action(), item.basedOnAnalysisIds(),
item.requiresHumanConfirmation()))
.toList();
List<Recommendation> recommendations = draft.recommendations().stream()
.map(item -> new Recommendation(item.text(), item.basedOnAnalysisIds()))
.toList();
Limitations limitations = draft.limitations() == null ? null : new Limitations(
draft.limitations().scope(), draft.limitations().missingInfo());
return new SemanticDraftView(conclusion, analysis, actions, recommendations, limitations);
}
public boolean hasSameUserVisibleSemantics(SemanticDraftView other) {
if (other == null || !Objects.equals(text(conclusion), text(other.conclusion))
|| analysis.size() != other.analysis.size()
|| actionPlan.size() != other.actionPlan.size()
|| recommendations.size() != other.recommendations.size()) {
return false;
}
for (int i = 0; i < analysis.size(); i++) {
Analysis left = analysis.get(i);
Analysis right = other.analysis.get(i);
if (left.kind() != right.kind() || !Objects.equals(left.text(), right.text())) {
return false;
}
}
for (int i = 0; i < actionPlan.size(); i++) {
Action left = actionPlan.get(i);
Action right = other.actionPlan.get(i);
if (!Objects.equals(left.action(), right.action())
|| left.requiresHumanConfirmation() != right.requiresHumanConfirmation()) {
return false;
}
}
for (int i = 0; i < recommendations.size(); i++) {
if (!Objects.equals(recommendations.get(i).text(), other.recommendations.get(i).text())) {
return false;
}
}
return limitationsEqual(limitations, other.limitations);
}
private static String text(Conclusion value) {
return value == null ? null : value.text();
}
private static boolean limitationsEqual(Limitations left, Limitations right) {
if (left == null || right == null) {
return left == right;
}
return Objects.equals(left.scope(), right.scope())
&& Objects.equals(left.missingInfo(), right.missingInfo());
}
private static <T> List<T> immutable(List<T> values) {
return values == null ? List.of() : List.copyOf(values);
}
public record Conclusion(
@JsonProperty("text") String text,
@JsonProperty("based_on_analysis_ids") List<String> basedOnAnalysisIds) {
public Conclusion {
basedOnAnalysisIds = immutable(basedOnAnalysisIds);
}
}
public record Analysis(
@JsonProperty("analysis_id") String analysisId,
@JsonProperty("kind") AnalysisKind kind,
@JsonProperty("text") String text) {
}
public record Action(
@JsonProperty("action") String action,
@JsonProperty("based_on_analysis_ids") List<String> basedOnAnalysisIds,
@JsonProperty("requires_human_confirmation") boolean requiresHumanConfirmation) {
public Action {
basedOnAnalysisIds = immutable(basedOnAnalysisIds);
}
}
public record Recommendation(
@JsonProperty("text") String text,
@JsonProperty("based_on_analysis_ids") List<String> basedOnAnalysisIds) {
public Recommendation {
basedOnAnalysisIds = immutable(basedOnAnalysisIds);
}
}
public record Limitations(
@JsonProperty("scope") String scope,
@JsonProperty("missing_info") List<String> missingInfo) {
public Limitations {
missingInfo = immutable(missingInfo);
}
}
}
@@ -0,0 +1,134 @@
package com.superbiz.agent.harness.guard.semantic;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.harness.contract.SemanticVerdict;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.retry.HarnessRetryExecutor;
import com.superbiz.agent.harness.retry.RetryAttempt;
import com.superbiz.agent.harness.retry.RetryFailure;
import org.springframework.ai.chat.messages.SystemMessage;
import org.springframework.ai.chat.messages.UserMessage;
import org.springframework.ai.chat.prompt.Prompt;
import java.nio.charset.StandardCharsets;
import java.time.Duration;
import java.util.HashSet;
import java.util.Iterator;
import java.util.Objects;
import java.util.Set;
import java.util.function.Consumer;
public final class SemanticGuard {
private static final Set<String> OUTPUT_FIELDS = Set.of("verdict", "reason");
private final DiagnosisHarnessCore core;
private final HarnessRetryExecutor retryExecutor;
private final GuardModelCall modelCall;
private final ObjectMapper objectMapper;
private final SemanticGuardLimits limits;
private final Consumer<RetryAttempt> attemptRecorder;
private final String prompt;
public SemanticGuard(DiagnosisHarnessCore core,
HarnessRetryExecutor retryExecutor,
GuardModelCall modelCall,
ObjectMapper objectMapper,
SemanticGuardLimits limits,
Consumer<RetryAttempt> attemptRecorder) {
this.core = Objects.requireNonNull(core, "core must not be null");
this.retryExecutor = Objects.requireNonNull(retryExecutor, "retryExecutor must not be null");
this.modelCall = Objects.requireNonNull(modelCall, "modelCall must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
this.limits = Objects.requireNonNull(limits, "limits must not be null");
this.attemptRecorder = Objects.requireNonNull(
attemptRecorder, "attemptRecorder must not be null");
this.prompt = SemanticGuardPrompt.load();
}
public SemanticGuardDecision review(RunContext context, SemanticGuardInput input) {
Objects.requireNonNull(context, "context must not be null");
Objects.requireNonNull(input, "input must not be null");
String inputJson = serialize(input);
long inputBytes = utf8Bytes(inputJson);
if (inputBytes > limits.maxInputBytes()) {
throw new GuardModelCallException(
RetryFailure.SCHEMA_INVALID, "SemanticGuard input exceeded limit");
}
core.reserveRunBytes(context, inputBytes);
Prompt modelPrompt = new Prompt(java.util.List.of(
new SystemMessage(prompt), new UserMessage(inputJson)));
long startedNanos = System.nanoTime();
return retryExecutor.execute(
context,
context.retryPolicies().semanticGuard(),
() -> parse(modelCall.call(
context, modelPrompt, remainingTimeout(startedNanos), limits.maxOutputBytes())),
this::classify,
attemptRecorder);
}
private SemanticGuardDecision parse(String output) {
JsonNode root;
try {
root = objectMapper.readTree(output);
} catch (JsonProcessingException exception) {
throw new GuardModelCallException(
RetryFailure.PARSE_ERROR, "SemanticGuard output is not JSON", exception);
}
if (root == null || !root.isObject() || !fieldNames(root).equals(OUTPUT_FIELDS)
|| !root.path("verdict").isTextual() || !root.path("reason").isTextual()
|| root.path("reason").asText().isBlank()) {
throw new GuardModelCallException(
RetryFailure.SCHEMA_INVALID, "SemanticGuard output schema is invalid");
}
SemanticVerdict verdict;
try {
verdict = SemanticVerdict.valueOf(root.path("verdict").asText());
} catch (IllegalArgumentException exception) {
throw new GuardModelCallException(
RetryFailure.SCHEMA_INVALID, "SemanticGuard verdict is invalid", exception);
}
return new SemanticGuardDecision(verdict, root.path("reason").asText());
}
private Duration remainingTimeout(long startedNanos) {
long elapsed = Math.max(0L, System.nanoTime() - startedNanos);
long remaining = limits.totalTimeout().toNanos() - elapsed;
if (remaining <= 0) {
throw new GuardModelCallException(
RetryFailure.TIMEOUT, "SemanticGuard total timeout exhausted");
}
return Duration.ofNanos(Math.min(remaining, limits.perAttemptTimeout().toNanos()));
}
private RetryFailure classify(Exception exception) {
if (exception instanceof GuardModelCallException guardFailure) {
return guardFailure.failure();
}
return RetryFailure.UNKNOWN;
}
private String serialize(SemanticGuardInput input) {
try {
return objectMapper.writeValueAsString(input);
} catch (JsonProcessingException exception) {
throw new GuardModelCallException(
RetryFailure.SCHEMA_INVALID, "SemanticGuard input is not serializable", exception);
}
}
private static Set<String> fieldNames(JsonNode node) {
Set<String> names = new HashSet<>();
Iterator<String> fields = node.fieldNames();
fields.forEachRemaining(names::add);
return names;
}
private static long utf8Bytes(String value) {
return value.getBytes(StandardCharsets.UTF_8).length;
}
}
@@ -0,0 +1,18 @@
package com.superbiz.agent.harness.guard.semantic;
import com.fasterxml.jackson.annotation.JsonProperty;
import com.superbiz.agent.harness.contract.SemanticVerdict;
import java.util.Objects;
public record SemanticGuardDecision(
@JsonProperty("verdict") SemanticVerdict verdict,
@JsonProperty("reason") String reason) {
public SemanticGuardDecision {
Objects.requireNonNull(verdict, "verdict must not be null");
if (reason == null || reason.isBlank()) {
throw new IllegalArgumentException("reason must not be blank");
}
}
}
@@ -0,0 +1,26 @@
package com.superbiz.agent.harness.guard.semantic;
import com.fasterxml.jackson.annotation.JsonProperty;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot;
import java.util.Objects;
public record SemanticGuardInput(
@JsonProperty("query") String query,
@JsonProperty("draft") SemanticDraftView draft,
@JsonProperty("verified_evidence") VerifiedEvidenceSnapshot verifiedEvidence) {
public SemanticGuardInput {
if (query == null || query.isBlank()) {
throw new IllegalArgumentException("query must not be blank");
}
Objects.requireNonNull(draft, "draft must not be null");
Objects.requireNonNull(verifiedEvidence, "verifiedEvidence must not be null");
}
public static SemanticGuardInput from(String query, DiagnosisDraft draft,
VerifiedEvidenceSnapshot verifiedEvidence) {
return new SemanticGuardInput(query, SemanticDraftView.from(draft), verifiedEvidence);
}
}
@@ -0,0 +1,30 @@
package com.superbiz.agent.harness.guard.semantic;
import java.time.Duration;
import java.util.Objects;
public record SemanticGuardLimits(
long maxInputBytes,
long maxOutputBytes,
Duration perAttemptTimeout,
Duration totalTimeout) {
public SemanticGuardLimits {
if (maxInputBytes <= 0 || maxOutputBytes <= 0) {
throw new IllegalArgumentException("byte limits must be positive");
}
requirePositive(perAttemptTimeout, "perAttemptTimeout");
requirePositive(totalTimeout, "totalTimeout");
if (totalTimeout.compareTo(perAttemptTimeout) < 0) {
throw new IllegalArgumentException("totalTimeout must not be shorter than perAttemptTimeout");
}
}
private static void requirePositive(Duration value, String name) {
Objects.requireNonNull(value, name + " must not be null");
if (value.isZero() || value.isNegative()) {
throw new IllegalArgumentException(name + " must be positive");
}
value.toNanos();
}
}
@@ -0,0 +1,24 @@
package com.superbiz.agent.harness.guard.semantic;
import org.springframework.core.io.ClassPathResource;
import java.io.IOException;
import java.io.InputStream;
import java.nio.charset.StandardCharsets;
final class SemanticGuardPrompt {
private static final String RESOURCE_PATH = "prompts/semantic-guard-prompt.md";
private SemanticGuardPrompt() {
}
static String load() {
ClassPathResource resource = new ClassPathResource(RESOURCE_PATH);
try (InputStream input = resource.getInputStream()) {
return new String(input.readAllBytes(), StandardCharsets.UTF_8);
} catch (IOException exception) {
throw new IllegalStateException("Failed to load SemanticGuard prompt", exception);
}
}
}
@@ -0,0 +1,44 @@
package com.superbiz.agent.harness.release;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.contract.ReleaseOutcome;
import com.superbiz.agent.harness.contract.SafeFallback;
import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot;
import java.util.Objects;
public record DiagnosisReleaseResult(
ReleaseOutcome outcome,
DiagnosisDraft draft,
SafeFallback fallback,
VerifiedEvidenceSnapshot verifiedEvidence) {
public DiagnosisReleaseResult {
Objects.requireNonNull(outcome, "outcome must not be null");
Objects.requireNonNull(verifiedEvidence, "verifiedEvidence must not be null");
if (outcome == ReleaseOutcome.SUCCESS) {
Objects.requireNonNull(draft, "successful release requires draft");
if (fallback != null) {
throw new IllegalArgumentException("successful release must not contain fallback");
}
} else if (outcome == ReleaseOutcome.FALLBACK) {
Objects.requireNonNull(fallback, "fallback release requires fallback");
if (draft != null) {
throw new IllegalArgumentException("fallback release must not contain draft");
}
} else {
throw new IllegalArgumentException("release use case supports SUCCESS or FALLBACK only");
}
}
public static DiagnosisReleaseResult success(
DiagnosisDraft draft, VerifiedEvidenceSnapshot verifiedEvidence) {
return new DiagnosisReleaseResult(
ReleaseOutcome.SUCCESS, draft, null, verifiedEvidence);
}
public static DiagnosisReleaseResult fallback(SafeFallback fallback) {
return new DiagnosisReleaseResult(
ReleaseOutcome.FALLBACK, null, fallback, VerifiedEvidenceSnapshot.empty());
}
}
@@ -0,0 +1,91 @@
package com.superbiz.agent.harness.release;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.contract.SemanticVerdict;
import com.superbiz.agent.harness.core.BudgetExceededException;
import com.superbiz.agent.harness.core.RunAbortedException;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.guard.evidence.EvidenceGuard;
import com.superbiz.agent.harness.guard.evidence.EvidenceGuardResult;
import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot;
import com.superbiz.agent.harness.guard.semantic.SemanticGuard;
import com.superbiz.agent.harness.guard.semantic.SemanticGuardDecision;
import com.superbiz.agent.harness.guard.semantic.SemanticGuardInput;
import com.superbiz.agent.harness.retry.RetryExecutionException;
import com.superbiz.agent.harness.retry.RetryFailure;
import java.util.Objects;
public final class DiagnosisReleaseUseCase {
private final EvidenceGuard evidenceGuard;
private final EvidenceRepair evidenceRepair;
private final SemanticGuard semanticGuard;
private final SafeFallbackFactory fallbackFactory;
public DiagnosisReleaseUseCase(EvidenceGuard evidenceGuard,
EvidenceRepair evidenceRepair,
SemanticGuard semanticGuard,
SafeFallbackFactory fallbackFactory) {
this.evidenceGuard = Objects.requireNonNull(evidenceGuard, "evidenceGuard must not be null");
this.evidenceRepair = Objects.requireNonNull(evidenceRepair, "evidenceRepair must not be null");
this.semanticGuard = Objects.requireNonNull(semanticGuard, "semanticGuard must not be null");
this.fallbackFactory = Objects.requireNonNull(
fallbackFactory, "fallbackFactory must not be null");
}
public DiagnosisReleaseResult execute(RunContext context, String query, DiagnosisDraft draft) {
Objects.requireNonNull(context, "context must not be null");
if (query == null || query.isBlank()) {
throw new IllegalArgumentException("query must not be blank");
}
Objects.requireNonNull(draft, "draft must not be null");
DiagnosisDraft candidate = draft;
EvidenceGuardResult evidence = evidenceGuard.validate(context, candidate);
if (!evidence.valid()) {
try {
candidate = evidenceRepair.repair(context, query, draft, evidence.violations());
evidence = evidenceGuard.validate(context, candidate);
} catch (RuntimeException exception) {
propagateTerminal(exception);
return evidenceFailure();
}
if (!evidence.valid()) {
return evidenceFailure();
}
}
VerifiedEvidenceSnapshot snapshot = evidence.verifiedSnapshot().orElseThrow();
SemanticGuardDecision decision;
try {
decision = semanticGuard.review(
context, SemanticGuardInput.from(query, candidate, snapshot));
} catch (RuntimeException exception) {
propagateTerminal(exception);
return DiagnosisReleaseResult.fallback(
fallbackFactory.semanticUnavailable(snapshot));
}
return decision.verdict() == SemanticVerdict.SUPPORTED
? DiagnosisReleaseResult.success(candidate, snapshot)
: DiagnosisReleaseResult.fallback(
fallbackFactory.semanticUnsupported(snapshot));
}
private DiagnosisReleaseResult evidenceFailure() {
return DiagnosisReleaseResult.fallback(
fallbackFactory.evidenceValidationFailed());
}
private void propagateTerminal(RuntimeException exception) {
if (exception instanceof RunAbortedException
|| exception instanceof BudgetExceededException) {
throw exception;
}
if (exception instanceof RetryExecutionException retry
&& (retry.failure() == RetryFailure.CANCELLED
|| retry.failure() == RetryFailure.BUDGET_EXHAUSTED)) {
throw retry;
}
}
}
@@ -0,0 +1,125 @@
package com.superbiz.agent.harness.release;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.DeserializationFeature;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.ObjectReader;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.guard.evidence.EvidenceViolation;
import com.superbiz.agent.harness.guard.semantic.GuardModelCall;
import com.superbiz.agent.harness.guard.semantic.GuardModelCallException;
import com.superbiz.agent.harness.guard.semantic.SemanticDraftView;
import com.superbiz.agent.harness.retry.HarnessRetryExecutor;
import com.superbiz.agent.harness.retry.RetryAttempt;
import com.superbiz.agent.harness.retry.RetryFailure;
import org.springframework.ai.chat.messages.SystemMessage;
import org.springframework.ai.chat.messages.UserMessage;
import org.springframework.ai.chat.prompt.Prompt;
import java.nio.charset.StandardCharsets;
import java.util.List;
import java.util.Objects;
import java.util.function.Consumer;
public final class EvidenceRepair {
private final DiagnosisHarnessCore core;
private final HarnessRetryExecutor retryExecutor;
private final GuardModelCall modelCall;
private final ObjectMapper objectMapper;
private final ObjectReader draftReader;
private final EvidenceRepairLimits limits;
private final Consumer<RetryAttempt> attemptRecorder;
private final String prompt;
public EvidenceRepair(DiagnosisHarnessCore core,
HarnessRetryExecutor retryExecutor,
GuardModelCall modelCall,
ObjectMapper objectMapper,
EvidenceRepairLimits limits,
Consumer<RetryAttempt> attemptRecorder) {
this.core = Objects.requireNonNull(core, "core must not be null");
this.retryExecutor = Objects.requireNonNull(retryExecutor, "retryExecutor must not be null");
this.modelCall = Objects.requireNonNull(modelCall, "modelCall must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
this.draftReader = objectMapper.readerFor(DiagnosisDraft.class)
.with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
.with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
this.limits = Objects.requireNonNull(limits, "limits must not be null");
this.attemptRecorder = Objects.requireNonNull(
attemptRecorder, "attemptRecorder must not be null");
this.prompt = EvidenceRepairPrompt.load();
}
public DiagnosisDraft repair(RunContext context, String query, DiagnosisDraft original,
List<EvidenceViolation> violations) {
Objects.requireNonNull(context, "context must not be null");
if (query == null || query.isBlank()) {
throw new IllegalArgumentException("query must not be blank");
}
Objects.requireNonNull(original, "original must not be null");
List<EvidenceViolation> safeViolations = violations == null
? List.of() : List.copyOf(violations);
String inputJson = serialize(new RepairInput(query, original, safeViolations));
long inputBytes = utf8Bytes(inputJson);
if (inputBytes > limits.maxInputBytes()) {
throw new GuardModelCallException(
RetryFailure.SCHEMA_INVALID, "Evidence repair input exceeded limit");
}
core.reserveRunBytes(context, inputBytes);
Prompt modelPrompt = new Prompt(List.of(
new SystemMessage(prompt), new UserMessage(inputJson)));
SemanticDraftView originalSemantics = SemanticDraftView.from(original);
return retryExecutor.execute(
context,
context.retryPolicies().evidenceRepair(),
() -> {
DiagnosisDraft repaired = parse(modelCall.call(
context, modelPrompt, limits.timeout(), limits.maxOutputBytes()));
if (!originalSemantics.hasSameUserVisibleSemantics(
SemanticDraftView.from(repaired))) {
throw new GuardModelCallException(
RetryFailure.SCHEMA_INVALID,
"Evidence repair changed user-visible semantics");
}
return repaired;
},
this::classify,
attemptRecorder);
}
private DiagnosisDraft parse(String output) {
try {
return draftReader.readValue(output);
} catch (JsonProcessingException exception) {
throw new GuardModelCallException(
RetryFailure.PARSE_ERROR, "Evidence repair output is invalid", exception);
}
}
private RetryFailure classify(Exception exception) {
return exception instanceof GuardModelCallException failure
? failure.failure() : RetryFailure.UNKNOWN;
}
private String serialize(RepairInput input) {
try {
return objectMapper.writeValueAsString(input);
} catch (JsonProcessingException exception) {
throw new GuardModelCallException(
RetryFailure.SCHEMA_INVALID, "Evidence repair input is not serializable", exception);
}
}
private static long utf8Bytes(String value) {
return value.getBytes(StandardCharsets.UTF_8).length;
}
private record RepairInput(
String query,
DiagnosisDraft draft,
List<EvidenceViolation> violations) {
}
}
@@ -0,0 +1,21 @@
package com.superbiz.agent.harness.release;
import java.time.Duration;
import java.util.Objects;
public record EvidenceRepairLimits(
long maxInputBytes,
long maxOutputBytes,
Duration timeout) {
public EvidenceRepairLimits {
if (maxInputBytes <= 0 || maxOutputBytes <= 0) {
throw new IllegalArgumentException("byte limits must be positive");
}
Objects.requireNonNull(timeout, "timeout must not be null");
if (timeout.isZero() || timeout.isNegative()) {
throw new IllegalArgumentException("timeout must be positive");
}
timeout.toNanos();
}
}
@@ -0,0 +1,24 @@
package com.superbiz.agent.harness.release;
import org.springframework.core.io.ClassPathResource;
import java.io.IOException;
import java.io.InputStream;
import java.nio.charset.StandardCharsets;
final class EvidenceRepairPrompt {
private static final String RESOURCE_PATH = "prompts/evidence-repair-prompt.md";
private EvidenceRepairPrompt() {
}
static String load() {
ClassPathResource resource = new ClassPathResource(RESOURCE_PATH);
try (InputStream input = resource.getInputStream()) {
return new String(input.readAllBytes(), StandardCharsets.UTF_8);
} catch (IOException exception) {
throw new IllegalStateException("Failed to load evidence repair prompt", exception);
}
}
}
@@ -0,0 +1,51 @@
package com.superbiz.agent.harness.release;
import com.superbiz.agent.harness.contract.FallbackType;
import com.superbiz.agent.harness.contract.SafeFallback;
import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot;
import java.util.List;
import java.util.Objects;
public final class SafeFallbackFactory {
public SafeFallback evidenceValidationFailed() {
return fallback(
FallbackType.EVIDENCE_VALIDATION_FAILED,
List.of(),
"当前证据无法完成真实性校验,无法确认根因",
"证据引用校验未通过",
"重新收集当前诊断范围内的证据后再发起诊断");
}
public SafeFallback semanticUnsupported(VerifiedEvidenceSnapshot snapshot) {
return fallback(
FallbackType.SEMANTIC_UNSUPPORTED,
sources(snapshot),
"当前证据不足,无法确认根因",
"语义校验未通过",
"补充当前缺失的数据后重新发起诊断");
}
public SafeFallback semanticUnavailable(VerifiedEvidenceSnapshot snapshot) {
return fallback(
FallbackType.SEMANTIC_UNAVAILABLE,
sources(snapshot),
"当前证据暂时无法完成语义校验,无法确认根因",
"语义校验暂不可用",
"稍后重试或补充当前缺失的数据");
}
private SafeFallback fallback(FallbackType type,
List<SafeFallback.VerifiedSource> sources,
String message,
String limitation,
String nextStep) {
return new SafeFallback(
type, null, message, sources, List.of(limitation), List.of(nextStep));
}
private List<SafeFallback.VerifiedSource> sources(VerifiedEvidenceSnapshot snapshot) {
return Objects.requireNonNull(snapshot, "snapshot must not be null").verifiedSources();
}
}
@@ -0,0 +1,7 @@
You repair only the identifier and reference structure of one DiagnosisDraft.
Use the supplied violation codes and targets. You may change only analysis_id, tool_call_ids, and based_on_analysis_ids.
Do not add, remove, reorder, summarize, or rewrite any conclusion, analysis, action, recommendation, limitation, kind, or human-confirmation flag.
You have no tools, memory, evidence store, or permission to call another agent.
Return exactly the complete repaired DiagnosisDraft JSON with no markdown or explanation.
@@ -0,0 +1,10 @@
You are the isolated SemanticGuard for one diagnosis report.
Review only the supplied original query, complete user-visible draft, and verified evidence snapshot.
Check whether evidence supports each analysis, analyses support the conclusion, actions and recommendations stay within supported findings, and limitations accurately state the observed scope.
You have no tools, memory, Redis access, diagnosis history, or permission to call another agent.
Do not rewrite, correct, summarize, extend, or partially approve the report.
Return exactly one JSON object with no markdown and no extra fields:
{"verdict":"SUPPORTED|UNSUPPORTED","reason":"non-empty audit reason"}
@@ -0,0 +1,296 @@
package com.superbiz.agent.harness.guard.evidence;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.harness.contract.AnalysisKind;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.contract.EvidenceStatus;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunBudgetLimits;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
import com.superbiz.agent.harness.tool.store.CanonicalInvocationLimits;
import com.superbiz.agent.harness.tool.store.CanonicalInvocationStore;
import com.superbiz.agent.harness.tool.store.CanonicalToolInvocation;
import com.superbiz.agent.harness.tool.store.ToolCallKeyFactory;
import org.junit.jupiter.api.Test;
import java.time.Clock;
import java.time.Duration;
import java.time.Instant;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
class EvidenceGuardTest {
private static final String PREFIX = "superbiz:harness:tool-call";
private final ObjectMapper objectMapper = new ObjectMapper();
private final InMemoryStore store = new InMemoryStore();
private final EvidenceGuard guard = new EvidenceGuard(
store, new ToolCallKeyFactory(PREFIX), objectMapper);
private final RunContext context = context("run-guard");
@Test
void validCurrentRunRagReferenceProducesInternalIdFreeSnapshot() throws Exception {
String callId = "call-rag-1";
ready(callId, "lookup_knowledge", "{\"query\":\"pool timeout\"}", """
{"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-rag-1",
"query":"pool timeout","evidence":[{"document_id":"doc-1",
"source":"runbook.md","title":"Pool guide","breadcrumb":"DB > Pool",
"excerpt":"active=50 max=50"}],"returned_count":1,"truncated":false}
""", EvidenceStatus.EVIDENCE_FOUND);
EvidenceGuardResult result = guard.validate(context, draft(callId, AnalysisKind.NORMAL));
assertTrue(result.valid());
VerifiedEvidenceSnapshot snapshot = result.verifiedSnapshot().orElseThrow();
assertEquals("a-1", snapshot.analyses().get(0).analysisId());
assertEquals("runbook.md", snapshot.analyses().get(0).evidence().get(0).source());
assertEquals("active=50 max=50", snapshot.analyses().get(0).evidence().get(0).excerpt());
String json = objectMapper.writeValueAsString(snapshot);
assertFalse(json.contains(callId));
assertFalse(json.contains("raw_response"));
assertEquals(1, snapshot.verifiedSources().size());
}
@Test
void noEvidenceSupportsOnlyScopedNegativeObservation() {
String callId = "call-rag-empty";
ready(callId, "lookup_knowledge", "{\"query\":\"unknown failure\"}", """
{"evidence_status":"NO_EVIDENCE","tool_call_id":"call-rag-empty",
"query":"unknown failure","evidence":[],"returned_count":0,"truncated":false}
""", EvidenceStatus.NO_EVIDENCE);
EvidenceGuardResult result = guard.validate(
context, draft(callId, AnalysisKind.NEGATIVE_OBSERVATION));
assertTrue(result.valid());
VerifiedEvidence evidence = result.verifiedSnapshot().orElseThrow()
.analyses().get(0).evidence().get(0);
assertEquals("unknown failure", evidence.scope());
assertEquals(0, evidence.values().get("match_count"));
}
@Test
void validLogProjectionPreservesSourceScopeTimelineAndExactMessage() {
String callId = "call-log-1";
ready(callId, "query_logs", """
{"topic":"APPLICATION","query":"pool timeout","lookback_minutes":30}
""", """
{"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-log-1",
"source_kind":"MOCK","scope":{"topic":"APPLICATION","query":"pool timeout",
"start_time":"2026-07-21T10:00:00Z","end_time":"2026-07-21T10:30:00Z"},
"match_count":1,"returned_count":1,"patterns":[],
"events":[{"timestamp":"2026-07-21T10:29:00Z","level":"ERROR",
"service":"order-service","message":"active=50 max=50"}],"truncated":false}
""", EvidenceStatus.EVIDENCE_FOUND);
EvidenceGuardResult result = guard.validate(context, draft(callId, AnalysisKind.NORMAL));
assertTrue(result.valid());
VerifiedEvidence evidence = result.verifiedSnapshot().orElseThrow()
.analyses().get(0).evidence().get(0);
assertEquals("LOG", evidence.sourceType());
assertEquals("APPLICATION (MOCK)", evidence.source());
assertEquals("2026-07-21T10:29:00Z", evidence.timestamp());
assertEquals("active=50 max=50", evidence.excerpt());
assertTrue(evidence.scope().contains("pool timeout"));
}
@Test
void validMysqlProjectionCombinesBoundedRequestScopeAndProjectedRows() {
String callId = "call-mysql-1";
ready(callId, "query_mysql", """
{"data_source":"order_readonly","sql":"SELECT status FROM biz_order WHERE id = ?",
"params":["order-1"]}
""", """
{"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-mysql-1",
"columns":["status"],"rows":[{"status":"FAILED"}],
"returned_count":1,"truncated":false}
""", EvidenceStatus.EVIDENCE_FOUND);
EvidenceGuardResult result = guard.validate(context, draft(callId, AnalysisKind.NORMAL));
assertTrue(result.valid());
VerifiedEvidence evidence = result.verifiedSnapshot().orElseThrow()
.analyses().get(0).evidence().get(0);
assertEquals("MYSQL", evidence.sourceType());
assertEquals("order_readonly", evidence.source());
assertTrue(evidence.scope().contains("SELECT status"));
assertTrue(evidence.scope().contains("order-1"));
assertEquals("FAILED", evidence.values().get("status"));
}
@Test
void duplicateAnalysisAndBrokenReportReferenceFailBeforeStoreLookup() {
DiagnosisDraft invalid = new DiagnosisDraft(
new DiagnosisDraft.Conclusion("Pool exhausted", List.of("missing")),
List.of(
new DiagnosisDraft.AnalysisItem(
"a-1", AnalysisKind.NORMAL, "first", List.of("call-1")),
new DiagnosisDraft.AnalysisItem(
"a-1", AnalysisKind.NORMAL, "second", List.of("call-2"))),
List.of(new DiagnosisDraft.ActionPlanItem("inspect", List.of(), false)),
List.of(), new DiagnosisDraft.Limitations("order-service", List.of()));
EvidenceGuardResult result = guard.validate(context, invalid);
assertFalse(result.valid());
assertTrue(hasViolation(result, EvidenceViolationCode.ANALYSIS_ID_DUPLICATE));
assertTrue(hasViolation(result, EvidenceViolationCode.ANALYSIS_REFERENCE_UNKNOWN));
assertTrue(hasViolation(result, EvidenceViolationCode.ANALYSIS_REFERENCE_MISSING));
assertTrue(result.verifiedSnapshot().isEmpty());
}
@Test
void fabricatedAndCrossRunReferencesAreRejected() {
EvidenceGuardResult missing = guard.validate(
context, draft("fabricated-call", AnalysisKind.NORMAL));
String callId = "call-cross-run";
String currentKey = PREFIX + ":" + context.runId() + ":" + callId;
store.records.put(currentKey, CanonicalToolInvocation.projecting(
callId, "another-run", "lookup_knowledge", "{\"query\":\"pool\"}",
Instant.parse("2026-07-21T10:00:00Z"))
.markReady("raw", """
{"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-cross-run",
"query":"pool","evidence":[{"document_id":"doc-1","source":"guide",
"title":"guide","breadcrumb":"pool","excerpt":"active=50"}],
"returned_count":1,"truncated":false}
""", EvidenceStatus.EVIDENCE_FOUND,
Instant.parse("2026-07-21T10:00:01Z")));
EvidenceGuardResult crossRun = guard.validate(
context, draft(callId, AnalysisKind.NORMAL));
assertTrue(hasViolation(missing, EvidenceViolationCode.INVOCATION_MISSING));
assertTrue(hasViolation(crossRun, EvidenceViolationCode.INVOCATION_NOT_REFERENCABLE));
}
@Test
void projectingInvocationAndEvidenceKindMismatchAreRejected() {
String projectingId = "call-projecting";
store.records.put(PREFIX + ":" + context.runId() + ":" + projectingId,
CanonicalToolInvocation.projecting(
projectingId, context.runId(), "lookup_knowledge", "{\"query\":\"pool\"}",
Instant.parse("2026-07-21T10:00:00Z")));
EvidenceGuardResult projecting = guard.validate(
context, draft(projectingId, AnalysisKind.NORMAL));
String emptyId = "call-empty-normal";
ready(emptyId, "lookup_knowledge", "{\"query\":\"pool\"}", """
{"evidence_status":"NO_EVIDENCE","tool_call_id":"call-empty-normal",
"query":"pool","evidence":[],"returned_count":0,"truncated":false}
""", EvidenceStatus.NO_EVIDENCE);
EvidenceGuardResult mismatch = guard.validate(
context, draft(emptyId, AnalysisKind.NORMAL));
assertTrue(hasViolation(projecting, EvidenceViolationCode.INVOCATION_NOT_REFERENCABLE));
assertTrue(hasViolation(mismatch, EvidenceViolationCode.EVIDENCE_KIND_MISMATCH));
}
@Test
void projectionIdOrStatusMismatchFailsClosed() {
String callId = "call-mismatch";
ready(callId, "lookup_knowledge", "{\"query\":\"pool\"}", """
{"evidence_status":"NO_EVIDENCE","tool_call_id":"another-call",
"query":"pool","evidence":[],"returned_count":0,"truncated":false}
""", EvidenceStatus.EVIDENCE_FOUND);
EvidenceGuardResult result = guard.validate(context, draft(callId, AnalysisKind.NORMAL));
assertTrue(hasViolation(result, EvidenceViolationCode.PROJECTION_ID_MISMATCH));
}
@Test
void canonicalRecordIdMustMatchTheDraftReferenceEvenUnderCorruptedKey() {
String referencedId = "call-requested";
CanonicalToolInvocation wrongRecord = CanonicalToolInvocation.projecting(
"call-other", context.runId(), "lookup_knowledge", "{\"query\":\"pool\"}",
Instant.parse("2026-07-21T10:00:00Z"))
.markReady("raw", """
{"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-other",
"query":"pool","evidence":[{"document_id":"doc-1","source":"guide",
"title":"guide","breadcrumb":"pool","excerpt":"active=50"}],
"returned_count":1,"truncated":false}
""", EvidenceStatus.EVIDENCE_FOUND,
Instant.parse("2026-07-21T10:00:01Z"));
store.records.put(PREFIX + ":" + context.runId() + ":" + referencedId, wrongRecord);
EvidenceGuardResult result = guard.validate(
context, draft(referencedId, AnalysisKind.NORMAL));
assertTrue(hasViolation(result, EvidenceViolationCode.INVOCATION_ID_MISMATCH));
}
private boolean hasViolation(EvidenceGuardResult result, EvidenceViolationCode code) {
return result.violations().stream().anyMatch(violation -> violation.code() == code);
}
private DiagnosisDraft draft(String callId, AnalysisKind kind) {
return new DiagnosisDraft(
new DiagnosisDraft.Conclusion("Pool exhausted", List.of("a-1")),
List.of(new DiagnosisDraft.AnalysisItem(
"a-1", kind, "Pool reached its limit", List.of(callId))),
List.of(new DiagnosisDraft.ActionPlanItem(
"Inspect long transactions", List.of("a-1"), false)),
List.of(new DiagnosisDraft.Recommendation(
"Add saturation alert", List.of("a-1"))),
new DiagnosisDraft.Limitations("order-service", List.of()));
}
private void ready(String callId, String toolName, String request,
String agentResult, EvidenceStatus evidenceStatus) {
String key = PREFIX + ":" + context.runId() + ":" + callId;
CanonicalToolInvocation invocation = CanonicalToolInvocation.projecting(
callId, context.runId(), toolName, request, Instant.parse("2026-07-21T10:00:00Z"))
.markReady("raw-must-not-be-read", agentResult, evidenceStatus,
Instant.parse("2026-07-21T10:00:01Z"));
store.records.put(key, invocation);
}
private RunContext context(String runId) {
return new DiagnosisHarnessCore(
Clock.systemUTC(), () -> runId, Duration.ofMinutes(5),
new RunBudgetLimits(10, 10, 10, 10_000, 10_000, 20_000, 1_000_000),
HarnessRetryPolicies.strict()).startRun("session-guard", runId);
}
private static final class InMemoryStore implements CanonicalInvocationStore {
private final Map<String, CanonicalToolInvocation> records = new HashMap<>();
private final CanonicalInvocationLimits limits =
new CanonicalInvocationLimits(Duration.ofHours(2), 1_000_000, 64_000);
@Override
public CanonicalInvocationLimits limits() {
return limits;
}
@Override
public void begin(String key, CanonicalToolInvocation invocation) {
records.put(key, invocation);
}
@Override
public Optional<CanonicalToolInvocation> find(String key) {
return Optional.ofNullable(records.get(key));
}
@Override
public CanonicalToolInvocation markReady(String key, String rawResponse,
String agentResult, EvidenceStatus evidenceStatus,
Instant completedAt) {
throw new UnsupportedOperationException();
}
@Override
public CanonicalToolInvocation markError(String key, String rawResponse,
String errorCode, Instant completedAt) {
throw new UnsupportedOperationException();
}
}
}
@@ -0,0 +1,231 @@
package com.superbiz.agent.harness.guard.semantic;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.harness.contract.AnalysisKind;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.contract.SemanticVerdict;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunBudgetLimits;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.guard.evidence.VerifiedAnalysisEvidence;
import com.superbiz.agent.harness.guard.evidence.VerifiedEvidence;
import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot;
import com.superbiz.agent.harness.retry.HarnessRetryExecutor;
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
import com.superbiz.agent.harness.retry.RetryAttempt;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.springframework.ai.chat.messages.AssistantMessage;
import org.springframework.ai.chat.metadata.ChatResponseMetadata;
import org.springframework.ai.chat.metadata.DefaultUsage;
import org.springframework.ai.chat.model.ChatModel;
import org.springframework.ai.chat.model.ChatResponse;
import org.springframework.ai.chat.model.Generation;
import org.springframework.ai.chat.prompt.Prompt;
import java.time.Clock;
import java.time.Duration;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.CompletionException;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicInteger;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assertions.assertThrows;
class SemanticGuardTest {
private final ObjectMapper objectMapper = new ObjectMapper();
private final ExecutorService executor = Executors.newCachedThreadPool();
@AfterEach
void shutdownExecutor() {
executor.shutdownNow();
}
@Test
void unsupportedIsAValidBusinessDecisionAndIsNotRetried() {
DiagnosisHarnessCore core = core();
RunContext context = core.startRun("session-semantic", "run-semantic");
ScriptedChatModel model = new ScriptedChatModel(
"{\"verdict\":\"UNSUPPORTED\",\"reason\":\"evidence does not prove the root cause\"}");
List<RetryAttempt> attempts = new ArrayList<>();
SemanticGuard guard = new SemanticGuard(
core,
new HarnessRetryExecutor(core),
new GuardModelCall(core, model, executor),
objectMapper,
new SemanticGuardLimits(100_000, 10_000,
Duration.ofSeconds(2), Duration.ofSeconds(3)),
attempts::add);
SemanticGuardDecision decision = guard.review(
context, SemanticGuardInput.from("Why did payment fail?", draft(), snapshot()));
assertEquals(SemanticVerdict.UNSUPPORTED, decision.verdict());
assertEquals(1, model.calls.get());
assertEquals(1, attempts.size());
assertTrue(attempts.get(0).success());
String prompt = model.prompts.get(0);
assertTrue(prompt.contains("Why did payment fail?"));
assertTrue(prompt.contains("active=50 max=50"));
assertFalse(prompt.contains("call-rag-1"));
assertFalse(prompt.contains("tool_call_id"));
assertFalse(prompt.contains("raw_response"));
}
@Test
void parseFailureRetriesOnceWithTheExactSameInput() {
DiagnosisHarnessCore core = core();
RunContext context = core.startRun("session-retry", "run-retry");
ScriptedChatModel model = new ScriptedChatModel(
"not-json", "{\"verdict\":\"SUPPORTED\",\"reason\":\"all claims are grounded\"}");
List<RetryAttempt> attempts = new ArrayList<>();
SemanticGuard guard = new SemanticGuard(
core, new HarnessRetryExecutor(core),
new GuardModelCall(core, model, executor), objectMapper,
new SemanticGuardLimits(100_000, 10_000,
Duration.ofSeconds(2), Duration.ofSeconds(3)), attempts::add);
SemanticGuardDecision decision = guard.review(
context, SemanticGuardInput.from("Why did payment fail?", draft(), snapshot()));
assertEquals(SemanticVerdict.SUPPORTED, decision.verdict());
assertEquals(2, model.calls.get());
assertEquals(model.prompts.get(0), model.prompts.get(1));
assertEquals(com.superbiz.agent.harness.retry.RetryFailure.PARSE_ERROR,
attempts.get(0).failure());
assertTrue(attempts.get(1).success());
assertEquals(2, context.budget().snapshot().modelCalls());
assertEquals(16, context.budget().snapshot().totalTokens());
}
@Test
void attemptTimeoutCancelsBothPermittedModelCalls() throws Exception {
DiagnosisHarnessCore core = core();
RunContext context = core.startRun("session-timeout", "run-timeout");
BlockingChatModel model = new BlockingChatModel(2);
List<RetryAttempt> attempts = new ArrayList<>();
SemanticGuard guard = new SemanticGuard(
core, new HarnessRetryExecutor(core),
new GuardModelCall(core, model, executor), objectMapper,
new SemanticGuardLimits(100_000, 10_000,
Duration.ofMillis(50), Duration.ofMillis(500)), attempts::add);
com.superbiz.agent.harness.retry.RetryExecutionException failure = assertThrows(
com.superbiz.agent.harness.retry.RetryExecutionException.class,
() -> guard.review(context,
SemanticGuardInput.from("Why?", draft(), snapshot())));
assertEquals(com.superbiz.agent.harness.retry.RetryFailure.TIMEOUT, failure.failure());
assertEquals(2, failure.attempts());
assertTrue(model.interrupted.await(2, TimeUnit.SECONDS));
assertEquals(2, model.calls.get());
assertEquals(2, attempts.size());
}
@Test
void runCancellationInterruptsPendingModelAndDoesNotRetry() throws Exception {
DiagnosisHarnessCore core = core();
RunContext context = core.startRun("session-cancel", "run-cancel");
BlockingChatModel model = new BlockingChatModel(1);
SemanticGuard guard = new SemanticGuard(
core, new HarnessRetryExecutor(core),
new GuardModelCall(core, model, executor), objectMapper,
new SemanticGuardLimits(100_000, 10_000,
Duration.ofSeconds(2), Duration.ofSeconds(3)), ignored -> { });
CompletableFuture<SemanticGuardDecision> execution = CompletableFuture.supplyAsync(
() -> guard.review(context,
SemanticGuardInput.from("Why?", draft(), snapshot())));
assertTrue(model.started.await(2, TimeUnit.SECONDS));
core.cancel(context,
com.superbiz.agent.harness.core.RunCancellationReason.USER_REQUESTED);
CompletionException thrown = assertThrows(CompletionException.class, execution::join);
com.superbiz.agent.harness.retry.RetryExecutionException failure =
(com.superbiz.agent.harness.retry.RetryExecutionException) thrown.getCause();
assertEquals(com.superbiz.agent.harness.retry.RetryFailure.CANCELLED, failure.failure());
assertTrue(model.interrupted.await(2, TimeUnit.SECONDS));
assertEquals(1, model.calls.get());
}
private DiagnosisDraft draft() {
return new DiagnosisDraft(
new DiagnosisDraft.Conclusion("Pool exhausted", List.of("a-1")),
List.of(new DiagnosisDraft.AnalysisItem(
"a-1", AnalysisKind.NORMAL, "Pool reached its limit", List.of("call-rag-1"))),
List.of(new DiagnosisDraft.ActionPlanItem(
"Inspect long transactions", List.of("a-1"), false)),
List.of(new DiagnosisDraft.Recommendation(
"Add saturation alert", List.of("a-1"))),
new DiagnosisDraft.Limitations("order-service, last 30 minutes", List.of("No slow SQL")));
}
private VerifiedEvidenceSnapshot snapshot() {
return new VerifiedEvidenceSnapshot(List.of(new VerifiedAnalysisEvidence(
"a-1", "Pool reached its limit", AnalysisKind.NORMAL,
List.of(new VerifiedEvidence(
"LOG", "APPLICATION (MOCK)", "order-service, last 30 minutes",
"2026-07-21T10:29:00Z", "active=50 max=50", Map.of("count", 1))))));
}
private DiagnosisHarnessCore core() {
return new DiagnosisHarnessCore(
Clock.systemUTC(), () -> "unused", Duration.ofMinutes(5),
new RunBudgetLimits(10, 10, 10, 100_000, 100_000, 200_000, 1_000_000),
HarnessRetryPolicies.strict());
}
private static final class ScriptedChatModel implements ChatModel {
private final List<String> responses;
private final List<String> prompts = new ArrayList<>();
private final AtomicInteger calls = new AtomicInteger();
private ScriptedChatModel(String... responses) {
this.responses = List.of(responses);
}
@Override
public ChatResponse call(Prompt prompt) {
prompts.add(prompt.getContents());
int index = calls.getAndIncrement();
ChatResponseMetadata metadata = ChatResponseMetadata.builder()
.usage(new DefaultUsage(5, 3)).build();
return new ChatResponse(
List.of(new Generation(new AssistantMessage(responses.get(index)))), metadata);
}
}
private static final class BlockingChatModel implements ChatModel {
private final AtomicInteger calls = new AtomicInteger();
private final CountDownLatch started = new CountDownLatch(1);
private final CountDownLatch interrupted;
private BlockingChatModel(int expectedInterruptions) {
this.interrupted = new CountDownLatch(expectedInterruptions);
}
@Override
public ChatResponse call(Prompt prompt) {
calls.incrementAndGet();
started.countDown();
try {
Thread.sleep(10_000);
throw new AssertionError("blocking model should be interrupted");
} catch (InterruptedException exception) {
interrupted.countDown();
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted", exception);
}
}
}
}
@@ -0,0 +1,314 @@
package com.superbiz.agent.harness.release;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.harness.contract.AnalysisKind;
import com.superbiz.agent.harness.contract.DiagnosisDraft;
import com.superbiz.agent.harness.contract.EvidenceStatus;
import com.superbiz.agent.harness.contract.FallbackType;
import com.superbiz.agent.harness.contract.ReleaseOutcome;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunBudgetLimits;
import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.guard.evidence.EvidenceGuard;
import com.superbiz.agent.harness.guard.semantic.GuardModelCall;
import com.superbiz.agent.harness.guard.semantic.SemanticGuard;
import com.superbiz.agent.harness.guard.semantic.SemanticGuardLimits;
import com.superbiz.agent.harness.retry.HarnessRetryExecutor;
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
import com.superbiz.agent.harness.tool.store.CanonicalInvocationLimits;
import com.superbiz.agent.harness.tool.store.CanonicalInvocationStore;
import com.superbiz.agent.harness.tool.store.CanonicalToolInvocation;
import com.superbiz.agent.harness.tool.store.ToolCallKeyFactory;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.springframework.ai.chat.messages.AssistantMessage;
import org.springframework.ai.chat.metadata.ChatResponseMetadata;
import org.springframework.ai.chat.metadata.DefaultUsage;
import org.springframework.ai.chat.model.ChatModel;
import org.springframework.ai.chat.model.ChatResponse;
import org.springframework.ai.chat.model.Generation;
import org.springframework.ai.chat.prompt.Prompt;
import java.time.Clock;
import java.time.Duration;
import java.time.Instant;
import java.util.ArrayList;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicInteger;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertSame;
class DiagnosisReleaseUseCaseTest {
private static final String PREFIX = "superbiz:harness:tool-call";
private final ObjectMapper objectMapper = new ObjectMapper();
private final ExecutorService executor = Executors.newCachedThreadPool();
@AfterEach
void shutdownExecutor() {
executor.shutdownNow();
}
@Test
void supportedVerifiedDraftIsReleasedUnchangedWithoutRepair() {
DiagnosisHarnessCore core = core();
RunContext context = core.startRun("session-release", "run-release");
InMemoryStore store = new InMemoryStore();
ready(store, context, "call-rag-1");
ScriptedChatModel model = new ScriptedChatModel(
"{\"verdict\":\"SUPPORTED\",\"reason\":\"all claims are grounded\"}");
HarnessRetryExecutor retries = new HarnessRetryExecutor(core);
GuardModelCall modelCall = new GuardModelCall(core, model, executor);
DiagnosisReleaseUseCase useCase = new DiagnosisReleaseUseCase(
new EvidenceGuard(store, new ToolCallKeyFactory(PREFIX), objectMapper),
new EvidenceRepair(core, retries, modelCall, objectMapper,
new EvidenceRepairLimits(100_000, 20_000, Duration.ofSeconds(2)), ignored -> { }),
new SemanticGuard(core, retries, modelCall, objectMapper,
new SemanticGuardLimits(100_000, 20_000,
Duration.ofSeconds(2), Duration.ofSeconds(3)), ignored -> { }),
new SafeFallbackFactory());
DiagnosisDraft draft = draft("a-1", "call-rag-1", "Pool exhausted");
DiagnosisReleaseResult result = useCase.execute(context, "Why did payment fail?", draft);
assertEquals(ReleaseOutcome.SUCCESS, result.outcome());
assertSame(draft, result.draft());
assertNull(result.fallback());
assertEquals(1, result.verifiedEvidence().analyses().size());
assertEquals(1, model.calls.get());
}
@Test
void oneStructuralRepairCanFixIdsWithoutChangingReportSemantics() {
Fixture fixture = fixture(
repairedDuplicateDraft("Pool exhausted", "a-2"),
"{\"verdict\":\"SUPPORTED\",\"reason\":\"grounded\"}");
DiagnosisDraft original = duplicateDraft();
DiagnosisReleaseResult result = fixture.useCase.execute(
fixture.context, "Why did payment fail?", original);
assertEquals(ReleaseOutcome.SUCCESS, result.outcome());
assertEquals(List.of("a-1", "a-2"), result.draft().analysis().stream()
.map(DiagnosisDraft.AnalysisItem::analysisId).toList());
assertEquals("Pool exhausted", result.draft().conclusion().text());
assertEquals(2, fixture.model.calls.get());
assertFalse(fixture.model.prompts.get(1).contains("call-rag-1"));
}
@Test
void repairThatChangesVisibleSemanticsReturnsEvidenceFallback() throws Exception {
Fixture fixture = fixture(repairedDuplicateDraft("Different conclusion", "a-2"));
DiagnosisReleaseResult result = fixture.useCase.execute(
fixture.context, "Why did payment fail?", duplicateDraft());
assertFallback(result, FallbackType.EVIDENCE_VALIDATION_FAILED, 0);
String json = objectMapper.writeValueAsString(result);
assertFalse(json.contains("Pool exhausted"));
assertFalse(json.contains("Different conclusion"));
assertEquals(1, fixture.model.calls.get());
}
@Test
void secondEvidenceValidationFailureHasNoVerifiedSources() {
Fixture fixture = fixture(repairedDuplicateDraft("Pool exhausted", "a-1"));
DiagnosisReleaseResult result = fixture.useCase.execute(
fixture.context, "Why did payment fail?", duplicateDraft());
assertFallback(result, FallbackType.EVIDENCE_VALIDATION_FAILED, 0);
assertEquals(1, fixture.model.calls.get());
}
@Test
void unsupportedNeverReleasesDraftOrAuditReason() throws Exception {
Fixture fixture = fixture(
"{\"verdict\":\"UNSUPPORTED\",\"reason\":\"secret-audit-reason\"}");
DiagnosisDraft draft = draft("a-1", "call-rag-1", "Pool exhausted");
DiagnosisReleaseResult result = fixture.useCase.execute(
fixture.context, "Why did payment fail?", draft);
assertFallback(result, FallbackType.SEMANTIC_UNSUPPORTED, 1);
String json = objectMapper.writeValueAsString(result);
assertFalse(json.contains("Pool exhausted"));
assertFalse(json.contains("secret-audit-reason"));
assertEquals(1, fixture.model.calls.get());
}
@Test
void twoInvalidSemanticOutputsReturnUnavailableWithoutDraft() throws Exception {
Fixture fixture = fixture("not-json", "still-not-json");
DiagnosisReleaseResult result = fixture.useCase.execute(
fixture.context, "Why did payment fail?",
draft("a-1", "call-rag-1", "Pool exhausted"));
assertFallback(result, FallbackType.SEMANTIC_UNAVAILABLE, 1);
assertFalse(objectMapper.writeValueAsString(result).contains("Pool exhausted"));
assertEquals(2, fixture.model.calls.get());
}
private Fixture fixture(String... responses) {
DiagnosisHarnessCore core = core();
RunContext context = core.startRun("session-fixture", "run-fixture");
InMemoryStore store = new InMemoryStore();
ready(store, context, "call-rag-1");
ScriptedChatModel model = new ScriptedChatModel(responses);
HarnessRetryExecutor retries = new HarnessRetryExecutor(core);
GuardModelCall modelCall = new GuardModelCall(core, model, executor);
DiagnosisReleaseUseCase useCase = new DiagnosisReleaseUseCase(
new EvidenceGuard(store, new ToolCallKeyFactory(PREFIX), objectMapper),
new EvidenceRepair(core, retries, modelCall, objectMapper,
new EvidenceRepairLimits(100_000, 20_000, Duration.ofSeconds(2)), ignored -> { }),
new SemanticGuard(core, retries, modelCall, objectMapper,
new SemanticGuardLimits(100_000, 20_000,
Duration.ofSeconds(2), Duration.ofSeconds(3)), ignored -> { }),
new SafeFallbackFactory());
return new Fixture(context, model, useCase);
}
private DiagnosisDraft duplicateDraft() {
return new DiagnosisDraft(
new DiagnosisDraft.Conclusion("Pool exhausted", List.of("dup")),
List.of(
new DiagnosisDraft.AnalysisItem(
"dup", AnalysisKind.NORMAL, "Pool reached its limit", List.of("call-rag-1")),
new DiagnosisDraft.AnalysisItem(
"dup", AnalysisKind.NORMAL, "Requests are waiting", List.of("call-rag-1"))),
List.of(new DiagnosisDraft.ActionPlanItem(
"Inspect long transactions", List.of("dup"), false)),
List.of(new DiagnosisDraft.Recommendation(
"Add saturation alert", List.of("dup"))),
new DiagnosisDraft.Limitations("order-service", List.of()));
}
private String repairedDuplicateDraft(String conclusion, String secondAnalysisId) {
return """
{
"conclusion":{"text":"%s","based_on_analysis_ids":["a-1"]},
"analysis":[
{"analysis_id":"a-1","kind":"NORMAL","text":"Pool reached its limit","tool_call_ids":["call-rag-1"]},
{"analysis_id":"%s","kind":"NORMAL","text":"Requests are waiting","tool_call_ids":["call-rag-1"]}
],
"action_plan":[{"action":"Inspect long transactions","based_on_analysis_ids":["a-1"],"requires_human_confirmation":false}],
"recommendations":[{"text":"Add saturation alert","based_on_analysis_ids":["a-1"]}],
"limitations":{"scope":"order-service","missing_info":[]}
}
""".formatted(conclusion, secondAnalysisId);
}
private void assertFallback(DiagnosisReleaseResult result,
FallbackType type, int expectedSources) {
assertEquals(ReleaseOutcome.FALLBACK, result.outcome());
assertNull(result.draft());
assertNotNull(result.fallback());
assertEquals(type, result.fallback().type());
assertNull(result.fallback().conclusion());
assertEquals(expectedSources, result.fallback().verifiedSources().size());
assertEquals(0, result.verifiedEvidence().analyses().size());
}
private DiagnosisDraft draft(String analysisId, String callId, String conclusion) {
return new DiagnosisDraft(
new DiagnosisDraft.Conclusion(conclusion, List.of(analysisId)),
List.of(new DiagnosisDraft.AnalysisItem(
analysisId, AnalysisKind.NORMAL, "Pool reached its limit", List.of(callId))),
List.of(new DiagnosisDraft.ActionPlanItem(
"Inspect long transactions", List.of(analysisId), false)),
List.of(new DiagnosisDraft.Recommendation(
"Add saturation alert", List.of(analysisId))),
new DiagnosisDraft.Limitations("order-service", List.of()));
}
private void ready(InMemoryStore store, RunContext context, String callId) {
CanonicalToolInvocation invocation = CanonicalToolInvocation.projecting(
callId, context.runId(), "lookup_knowledge", "{\"query\":\"pool timeout\"}",
Instant.parse("2026-07-21T10:00:00Z"))
.markReady("raw-must-not-be-read", """
{"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-rag-1",
"query":"pool timeout","evidence":[{"document_id":"doc-1",
"source":"runbook.md","title":"Pool guide","breadcrumb":"DB > Pool",
"excerpt":"active=50 max=50"}],"returned_count":1,"truncated":false}
""", EvidenceStatus.EVIDENCE_FOUND,
Instant.parse("2026-07-21T10:00:01Z"));
store.records.put(PREFIX + ":" + context.runId() + ":" + callId, invocation);
}
private DiagnosisHarnessCore core() {
return new DiagnosisHarnessCore(
Clock.systemUTC(), () -> "unused", Duration.ofMinutes(5),
new RunBudgetLimits(10, 10, 10, 100_000, 100_000, 200_000, 1_000_000),
HarnessRetryPolicies.strict());
}
private static final class ScriptedChatModel implements ChatModel {
private final List<String> responses;
private final List<String> prompts = new ArrayList<>();
private final AtomicInteger calls = new AtomicInteger();
private ScriptedChatModel(String... responses) {
this.responses = List.of(responses);
}
@Override
public ChatResponse call(Prompt prompt) {
prompts.add(prompt.getContents());
int index = calls.getAndIncrement();
ChatResponseMetadata metadata = ChatResponseMetadata.builder()
.usage(new DefaultUsage(5, 3)).build();
return new ChatResponse(
List.of(new Generation(new AssistantMessage(responses.get(index)))), metadata);
}
}
private static final class InMemoryStore implements CanonicalInvocationStore {
private final Map<String, CanonicalToolInvocation> records = new HashMap<>();
private final CanonicalInvocationLimits limits =
new CanonicalInvocationLimits(Duration.ofHours(2), 1_000_000, 64_000);
@Override
public CanonicalInvocationLimits limits() {
return limits;
}
@Override
public void begin(String key, CanonicalToolInvocation invocation) {
records.put(key, invocation);
}
@Override
public Optional<CanonicalToolInvocation> find(String key) {
return Optional.ofNullable(records.get(key));
}
@Override
public CanonicalToolInvocation markReady(String key, String rawResponse,
String agentResult, EvidenceStatus evidenceStatus,
Instant completedAt) {
throw new UnsupportedOperationException();
}
@Override
public CanonicalToolInvocation markError(String key, String rawResponse,
String errorCode, Instant completedAt) {
throw new UnsupportedOperationException();
}
}
private record Fixture(
RunContext context,
ScriptedChatModel model,
DiagnosisReleaseUseCase useCase) {
}
}