diff --git a/devflow/index.md b/devflow/index.md index 29ad6a3..ddd299d 100644 --- a/devflow/index.md +++ b/devflow/index.md @@ -40,3 +40,4 @@ | 2026-07-21 | single-react-rag-log-projections | RAG/log projection adapters through ToolBoundary | Harness/Tool projection | ISS-014, RAG, query_logs, projection, scope, redaction, MOCK, NO_EVIDENCE | openspec/changes/archive/2026-07-21-single-react-rag-log-projections | archived | | 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived | | 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived | +| 2026-07-21 | single-react-evidence-semantic-guards | Deterministic evidence validation, isolated semantic review and fail-closed diagnosis release | Harness/EvidenceGuard/SemanticGuard/Release | ISS-014, EvidenceGuard, verified snapshot, SemanticGuard, repair, fallback, release policy | openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards | archived | diff --git a/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/acceptance.md b/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/acceptance.md new file mode 100644 index 0000000..36c770b --- /dev/null +++ b/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/acceptance.md @@ -0,0 +1,44 @@ +# Acceptance: single-react-evidence-semantic-guards + +## Result + +- Status: archived +- OpenSpec tasks: 14/14 complete +- Interface impact: L2 internal +- Public protocol: unchanged + +## Static Verification + +- `git diff --check`:阶段文件无 whitespace error;工作区用户已有 `AGENTS.md`/`CLAUDE.md` 仅报告 line-ending warning,未纳入阶段范围。 +- 公开 `controller`、`ChatService.java`、`AiOpsService.java` diff 为空。 +- SemanticGuard 包和 prompt 不含 Tool Call ID、raw response、ReactAgent、StateGraph、ThreadLocal 或手写 while。 +- 新 guard/release 包不包含业务 Agent loop。 + +## Script Verification + +- `mvn -q -DskipTests compile`:通过。 +- Stage-focused `EvidenceGuardTest,SemanticGuardTest,DiagnosisReleaseUseCaseTest`:19 tests,0 failure/error。 +- Harness/Tool/Agent regression selection:18 suites / 70 tests,0 failure/error/skipped。 +- `openspec validate single-react-evidence-semantic-guards --strict`:通过。 +- `openspec validate --specs --strict`:16 passed / 2 failed;失败是前序 `mysql-readonly-tool`、`rag-log-projections` scenario 标题格式,不阻塞本 change,计划阶段 7 统一修复。 + +## Browser or Manual Verification + +- Not applicable。阶段 5 没有 UI 或公开入口变化。 + +## Not Verified + +- 未运行真实 LLM、Redis、日志和 MySQL live E2E;按 ISS-014 串行门禁统一留到阶段 7。 +- Provider 是否立即响应 Java thread interrupt 取决于 SDK;Harness 已保证 Future cancel、迟到 Token 记账和迟到结果不释放,阶段 7 需用真实模型观察取消延迟。 + +## Remaining Work + +- 阶段 6A:Chat Application Use Case、Intent Router、PreviousTurn 和 Run persistence。 +- 阶段 6B:公开 SSE 原子切换。 +- 阶段 7:旧链路清理、全局 spec 格式修复和最终 live E2E。 + +## Archive + +- `.archive-ready`: created +- OpenSpec archive: `openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards` +- Main spec: `openspec/specs/single-react-evidence-semantic-guards/spec.md` diff --git a/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/brief.md b/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/brief.md new file mode 100644 index 0000000..7ccab00 --- /dev/null +++ b/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/brief.md @@ -0,0 +1,32 @@ +# Brief: single-react-evidence-semantic-guards + +## Background + +阶段 4 已能生成带框架 Tool Call ID 的 `DiagnosisDraft`,但草稿在发布前还缺少当前 Run 证据验真、独立语义审查和 fail-closed 释放边界。 + +## Goal + +在不恢复业务 Graph、不切换公开入口的前提下,实现确定性 EvidenceGuard、一次语义不变的结构修复、无 Tool/无记忆的单轮 SemanticGuard,以及只发布原 Draft 或固定 SafeFallback 的内部 release use case。 + +## Scope + +- Draft 结构、Analysis/报告引用和 canonical Tool invocation 验真。 +- RAG/log/MySQL verified evidence snapshot。 +- Evidence repair 单次调用与语义不变约束。 +- SemanticGuard 模型/Token/字节预算、超时、取消、技术重试和严格二元输出。 +- SUPPORTED/UNSUPPORTED/unavailable/evidence-failed release policy。 +- focused tests、Harness/Tool/Agent 回归和静态范围验证。 + +## Non-goals + +- 不切换 Chat/AiOps/SSE 公开协议。 +- 不实现 Intent Router、PreviousTurn、Run persistence 或最终报告渲染。 +- 不删除旧 Gatekeeper/Verifier/Composer/多 Agent 链路。 +- 不运行 live model/Redis/MySQL E2E;统一留到阶段 7。 + +## Metadata + +- Scale: complex +- Interface impact: L2 internal +- OpenSpec: `single-react-evidence-semantic-guards` +- Parent issue: `ISS-014` diff --git a/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/decisions.md b/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/decisions.md new file mode 100644 index 0000000..907d632 --- /dev/null +++ b/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/decisions.md @@ -0,0 +1,144 @@ +# Decisions: single-react-evidence-semantic-guards + +## Discover Status + +- Checkpoint: Discover +- Capability source: `sm-flow`,使用 `grill-with-docs` 做 evidence-driven 澄清;`codebase-retrieval`、LSP 和 GitNexus MCP 在当前会话不可用,降级为既有 GitNexus 研究结论、`rg` 调用点核对和逐文件源码阅读。 +- Scale: complex。变更跨 canonical store、三类 Tool projection、模型预算、超时/取消、结构修复、语义审查和释放边界,但不切换公开协议。 +- `devflow/index.md` 命中阶段 0-4 的设计冻结、RunContext、canonical store、Tool projection 和 Diagnosis Agent;没有与当前 OpenSpec 冲突的 ADR。 + +## Question Pool + +| # | 维度 | 问题 | 模式 | 状态 | +|---|---|---|---|---| +| Q1 | 术语 | EvidenceGuard 是否判断证据足以推出结论? | evidence-driven | 已解决 | +| Q2 | 边界 | verified snapshot 能读取哪些 canonical 字段,并向 SemanticGuard 暴露哪些内容? | evidence-driven | 已解决 | +| Q3 | 边界 | `NO_EVIDENCE` 能支持哪类 Analysis? | evidence-driven | 已解决 | +| Q4 | 修复 | 首次 EvidenceGuard 失败允许如何修复,是否可重跑 Agent 或 Tool? | evidence-driven | 已解决 | +| Q5 | 语义 | SemanticGuard 是否拥有 Tool、记忆、ReAct loop 或报告写作能力? | evidence-driven | 已解决 | +| Q6 | 重试 | 哪些 SemanticGuard 结果允许第二次 attempt? | evidence-driven | 已解决 | +| Q7 | 超时 | 单轮模型超时和 Run 取消如何终止后台调用? | evidence-driven | 已解决 | +| Q8 | 释放 | 三种 Fallback 能否包含 Draft、reason 或未验真来源? | evidence-driven | 已解决 | +| Q9 | 接口 | 阶段 5 是否切换公开入口或修改 SSE/持久化协议? | evidence-driven | 已解决 | +| Q10 | 验收 | 如何证明不存在隐式 Agent/repair/retry loop? | evidence-driven | 已解决 | + +## Evidence-driven + +| 结论 | 证据来源 | 是否已汇报用户 | +|---|---|---| +| EvidenceGuard 只验证结构、引用和 canonical invocation 真实性,不判断 Analysis/Conclusion 的语义充分性。 | ISS-014 4.3、阶段 5;glossary `EvidenceGuard` | 已汇报 | +| Store 只能按 `ToolCallKeyFactory.create(runId, toolCallId)` 查找;可引用记录必须是当前 Run、`READY`、含 `agent_result` 且状态为 `EVIDENCE_FOUND|NO_EVIDENCE`。 | `CanonicalInvocationStore`、`CanonicalToolInvocation.isReferencableBy` | 已汇报 | +| Snapshot 解析 `agent_result` 和必要的有界 request 字段,不读取 `raw_response`;输出按 `analysis_id` 分组且不包含 Tool Call ID。 | ISS-014 4.4、10.2、已冻结 Tool contracts | 已汇报 | +| `NORMAL` 只绑定 `EVIDENCE_FOUND`;`NEGATIVE_OBSERVATION` 只绑定 `NO_EVIDENCE`,且零结果范围来自投影/请求。 | ISS-014 4.3、10.2;`EvidenceStatus` glossary | 已汇报 | +| Evidence repair 只在首次物理验真失败后执行一次无 Tool单轮模型调用,不重跑 Diagnosis Agent 或 Tool loop。 | ISS-014 10.3、阶段 5、重试边界决策 | 已汇报 | +| SemanticGuard 复用同一 `ChatModel`,使用全新 `Prompt` 单轮调用,无 Tool/记忆/ReAct;只返回二元 verdict 和审计 reason。 | ISS-014 4.4、阶段 5;Spring AI `ChatModel.call(Prompt)` | 已汇报 | +| 仅超时、传输、解析或 Schema 技术失败可重试一次;`UNSUPPORTED` 是有效业务结果,不重试。 | `HarnessRetryPolicies.strict().semanticGuard()`、ISS-014 重试策略 | 已汇报 | +| `ChatResponseMetadata.Usage` 可复用 Core 的 model/Token 预算;受控 `Future.get(timeout)` 可在超时或 Run 取消时取消任务。 | `HarnessModelInterceptor`、`DiagnosisHarnessCore`、`RunCancellation` | 已汇报 | +| `EVIDENCE_VALIDATION_FAILED` 的来源必须为空;其余 Fallback 只能包含已验真来源,不包含 Draft 或 SemanticGuard reason。 | ISS-014 10.3、阶段 5;`SafeFallback`/`FallbackType` | 已汇报 | +| 阶段 5 只新增内部用例,公开入口切换留到阶段 6B。 | ISS-014 阶段顺序和 6B 门禁 | 已汇报 | + +## User-interview + +- 本阶段没有新增 user-interview 问题。上述方向、范围、修复次数、模型复用、超时/重试、Fallback 和阶段串行规则均已在 ISS-014 评审及前序对话中由用户确认。 + +## Key Decisions + +- Snapshot 使用 Tool-specific adapter 严格解析冻结 projection;未知 Tool、ID 不一致、projection 反序列化失败或 projection 的 EvidenceStatus 与 canonical record 不一致均 fail closed。 +- MySQL 的逻辑数据源和查询范围来自 canonical `request`,行/列和值来自 `agent_result`;RAG/log 以 `agent_result` 自带的稳定来源和范围为准。 +- SemanticGuard 使用直接 `ChatModel.call(Prompt)`,不创建第二个 `ReactAgent`;模型调用由受控 Executor 执行,超时/取消时调用 `Future.cancel(true)`。 +- Evidence repair 与 SemanticGuard 共用同一模型抽象,但各自拥有独立 prompt、严格输出类型和预算;repair 策略固定一次 attempt,不能嵌套 retry executor。 +- Release use case 只在 EvidenceGuard 成功后调用 SemanticGuard;SUPPORTED 返回原 Draft 对象,所有其他路径只返回固定 `SafeFallback`。 +- 不创建 ADR:这些是 ISS-014 已冻结架构的阶段实现,不产生新的难逆转跨项目决策。 + +## OpenSpec Backfill + +- 需进入 proposal/design/spec/tasks:EvidenceGuard 规则、Tool-specific snapshot、无 raw/ID 泄漏、repair 单次边界、SemanticGuard 隔离/预算/超时/重试、原样发布与三类 Fallback、L2/公开隔离。 +- 不进入本阶段:公开 Chat/SSE 装配、Run 持久化、旧多 Agent 删除和 live E2E。 + +## Cross-artifact Alignment + +| 上游 -> 下游 | 检查内容 | 状态 | +|---|---|---| +| ISS-014/brief -> proposal | 物理验真、结构修复、隔离语义审查、固定 Fallback、阶段边界 | 已对齐 | +| proposal -> design | Store ownership、Tool-specific snapshot、timeout/cancel/retry、唯一报告作者和 L2 影响 | 已对齐 | +| design -> specs/tasks | 每个关键决策均有可观察 requirement 和对应实现/测试切片 | 已对齐 | +| specs -> tasks | 9 组 requirements 覆盖为 Evidence、model boundary、repair/release 和回归验证任务 | 已对齐 | + +## Architecture Audit + +- 能力来源:`zoom-out`,使用 glossary 的 Diagnosis Harness、EvidenceGuard、SemanticGuard、RunContext、Invocation Status 和 Evidence Status 术语。 +- 链路为 `query + Draft + RunContext -> EvidenceGuard/store -> optional repair/shared model boundary -> SemanticGuard/shared model boundary -> original Draft or fixed Fallback`,没有外层 Graph 或第二个 Tool loop。 +- Store 拥有物理调用记录,EvidenceGuard 只读并生成 snapshot;Diagnosis Agent 拥有报告语义,repair 只修 ID,SemanticGuard 只审查,release 只做二选一,数据所有权没有重叠。 +- 最大运行风险是 provider 忽略 interrupt;design 要求 Future cancel、迟到 Token 记录和 `core.checkActive` 丢弃迟到结果,阶段 6A 再负责 Run 终态。 +- 架构审计未发现与阶段 0-4 或 ADR 冲突;所有风险缓解已回写 design/spec/tasks。 + +## Interface Impact + +- 级别:L2 内部接口。 +- 变更对象:新增 guard/release records、interfaces、use cases、prompts 和 tests;现有方法签名不变。 +- 消费者:本阶段只有 focused tests,阶段 6A 才接入 production application use case。 +- 兼容性:公开 HTTP/SSE、Controller DTO、数据库、Redis key schema 和旧 Chat/AiOps 路径不变。 + +## Commit Gate Preflight + +- proposal、design、specs、tasks 完整,`openspec status` 为 complete,`openspec validate single-react-evidence-semantic-guards --strict` 通过。 +- Question pool 全部为已汇报的 evidence-driven 结论;无未确认 user-interview、接口等级或风险接受问题。 +- Cross-artifact 四段对齐无 gap;架构审计的数据所有权、模型取消和唯一报告作者约束已进入 design/spec/tasks。 +- 接口影响为 L2,仅新增内部 Java API;公开 Controller/SSE/JPA/Redis schema/旧 ChatService 保持不变。 +- Apply、Archive 和阶段 Git commit 使用用户对 ISS-014 各阶段的持续授权;实现必须严格限制为 Committed OpenSpec。 +- `.committed` 已创建,Committed OpenSpec 可进入 Apply。 + +## Pre-apply Research + +### Reference Implementations + +- `harness/tool/store/CanonicalInvocationStore.java`、`CanonicalToolInvocation.java`、`ToolCallKeyFactory.java`:当前 Run 物理验真和 lifecycle 真理源。 +- `harness/tool/projection/RagResultProjector.java`、`QueryLogsResultProjector.java`、`harness/tool/mysql/MysqlResultProjector.java`:三类 `agent_result` 的唯一生产者和有界字段来源。 +- `harness/agent/HarnessModelInterceptor.java`:模型调用前预算、响应 Usage 记账和 late-result active check 模式。 +- `harness/retry/HarnessRetryExecutor.java`、`HarnessRetryPolicies.java`:SemanticGuard 两次技术 attempt 与 repair 一次 attempt 的装配边界。 +- `harness/core/RunCancellation.java`:取消 callback 注册和 first-cancel 语义。 +- `service/ExecutorGatekeeperService.java`:仅参考报告内部引用检查思想,不复用旧 Map/JPA/session DTO 或规则目录。 + +### Technology Stack + +- Jackson strict `ObjectReader` 解析 frozen Draft/Tool projections;Tool-specific adapter 生成 typed snapshot,不透传 JsonNode/raw response。 +- Spring AI `ChatModel.call(Prompt)` 执行无 Tool 单轮 guard;`ChatResponseMetadata.Usage` 接入现有 Core Token 预算。 +- Java 17 `ExecutorService/Future.get(timeout)` 提供 per-attempt timeout,`Future.cancel(true)` 接入 Run cancellation;Semantic 总时限由 monotonic elapsed time 控制。 +- JUnit 5 scripted `ChatModel` 和 in-memory canonical store 作为外部边界 fake;测试只通过 EvidenceGuard/SemanticGuard/Release use case 公共接口断言行为。 +- 无 Controller、MQ、JPA、Flyway 或新 Maven dependency;不需要新共享基础设施。 + +## Apply Progress + +- 首模块对齐:EvidenceGuard typed contracts、Draft/reference checks、current-Run canonical validation 和 RAG/log/MySQL snapshot adapters 已完成,tasks 1.1-1.4 完成。 +- 8 个 `EvidenceGuardTest` 行为测试通过;快照序列化不含 Tool Call ID 或 `raw_response`,`NO_EVIDENCE` 仅在 `NEGATIVE_OBSERVATION` 下进入带 scope/zero-match 的快照。 +- TODO:tasks 2.1-4.2,尚未实现模型边界、SemanticGuard、repair/release 和综合回归。 +- REVIEW 发现 corrupted Store key 下 record ID 可能与 Draft reference 不同;分类为代码偏离,已增加 exact canonical ID 校验和回归测试,无需改变规格方向。 +- REVIEW 发现 fallback `DiagnosisReleaseResult` 携带完整 snapshot 会通过 `analysis_text` 间接泄漏 Draft;分类为规格安全边界细化,已回写 design/spec,并让 fallback result 强制使用空 snapshot,安全来源只保留在 `SafeFallback.verified_sources`。 + +## Apply Result + +- 新增确定性 EvidenceGuard:校验 typed Draft、唯一 Analysis ID、报告引用、exact current-Run Tool ID、READY lifecycle、agent result、Evidence Status 与 Analysis Kind,并严格展开 RAG/log/MySQL projection。 +- verified snapshot 按 Analysis ID 分组,只含稳定 source/scope/timestamp/excerpt/values;SemanticGuard 输入不含 Tool Call ID、Redis key 或 raw response。 +- 新增共享 `GuardModelCall`:复用系统 ChatModel,执行 Core model/Token/Run byte budget、per-attempt timeout、total timeout、Future cancellation 和 late-result active check。 +- 新增无 Tool/无记忆/无 ReAct 的单轮 SemanticGuard,严格解析 `SUPPORTED|UNSUPPORTED + reason`;仅技术失败使用既有两次 attempt,业务 UNSUPPORTED 不重试。 +- 新增一次 EvidenceRepair,只允许 ID/reference 修复;任何可见文本、kind、顺序、human-confirmation 或 limitations 变化都直接 Evidence fallback。 +- 新增 DiagnosisReleaseUseCase 和固定 SafeFallbackFactory:SUPPORTED 原样返回 Draft;evidence failed、semantic unsupported/unavailable 均不返回 Draft、完整 snapshot 或审计 reason。 +- 公开 Controller、ChatService、AiOpsService、SSE、JPA/Flyway 和旧多 Agent 链路无修改。 + +## Apply Verification + +- 编译:`mvn -q -DskipTests compile` 通过。 +- 阶段 5 focused:`EvidenceGuardTest` 9、`SemanticGuardTest` 4、`DiagnosisReleaseUseCaseTest` 6,共 19 tests,0 failure/error。 +- 综合回归:阶段 5 + Harness Core/Retry/Tool boundary/store/projections/contracts + Diagnosis Agent,共 18 suites / 70 tests,0 failure/error/skipped。 +- OpenSpec:`openspec validate single-react-evidence-semantic-guards --strict` 通过。 +- 静态范围:公开 Controller/ChatService/AiOpsService diff 为空;SemanticGuard 无 Tool ID/raw response/ReactAgent/StateGraph/ThreadLocal/手写 while;新 guard/release 包无业务 Agent loop。 +- 仓库级 `openspec validate --specs --strict` 为 16 passed / 2 failed;失败仍是前序 `mysql-readonly-tool`、`rag-log-projections` 的 scenario 标题格式,按计划阶段 7 统一修复。 +- 未执行 live E2E:按 ISS-014 阶段门禁统一留到阶段 7。 + +## Archive Result + +- 14/14 OpenSpec tasks 完成,`.archive-ready` 已创建。 +- 主规格已同步至 `openspec/specs/single-react-evidence-semantic-guards/spec.md`,新增 9 个 requirements。 +- Change 已归档至 `openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards`。 +- `devflow/index.md` 和 ISS-014 阶段表已更新为阶段 0-5 archived,下一阶段为 6A。 +- 未创建 ADR/compound knowledge:本阶段落实 ISS-014 已冻结边界,没有新的跨项目难逆转决策。 diff --git a/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/evidence.md b/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/evidence.md new file mode 100644 index 0000000..5cc35a9 --- /dev/null +++ b/devflow/projects/2026-07-21-single-react-evidence-semantic-guards/evidence.md @@ -0,0 +1,29 @@ +# Evidence: single-react-evidence-semantic-guards + +## Code and Contract Evidence + +- `CanonicalInvocationStore` 只提供 key lookup;`ToolCallKeyFactory` 冻结 `runId + toolCallId` 隔离,`CanonicalToolInvocation` 冻结 READY/agent_result/evidence status 可引用条件。 +- `RagToolResult`、`QueryLogsToolResult`、`MysqlToolResult` 是 Agent-facing 有界 projection,足以构造 snapshot;只有 MySQL 的逻辑数据源/SQL/params 需要从 canonical request 补足。 +- `AnalysisKind.accepts` 已冻结 `NORMAL -> EVIDENCE_FOUND`、`NEGATIVE_OBSERVATION -> NO_EVIDENCE`。 +- `HarnessRetryPolicies.strict()` 已冻结 SemanticGuard 两次技术 attempt 和 Evidence repair 一次 attempt。 +- `ChatModel.call(Prompt)`、`ChatResponseMetadata.Usage` 和 `RunCancellation.onCancel` 支持直接单轮模型调用、Token 记账与 Future cancellation,无需 ReactAgent。 + +## Confirmed Boundaries + +- EvidenceGuard 不判断证据是否支持结论,只验证结构、引用和物理真实性。 +- SemanticGuard 只接收原始 Query、移除 Tool ID 的完整 Draft view 和 verified snapshot,不访问 Redis/raw response。 +- Evidence repair 只修改 ID/reference;用户可见语义发生任何变化即失败。 +- `UNSUPPORTED` 是有效业务结果,不重试;只有 timeout/transport/parse/schema 技术失败可进行第二次 attempt。 +- Fallback 不含 Draft、完整 snapshot 或 SemanticGuard reason;Evidence failure 的 verified sources 为空。 + +## Review Findings + +- 增加 canonical record ID 与 Draft reference 的 exact match,防止 corrupted key 映射被误信任。 +- Fallback release result 强制使用空 snapshot,防止 `analysis_text` 通过误序列化泄漏 Draft;安全来源只保留在 `SafeFallback.verified_sources`。 + +## Verification Evidence + +- Stage-focused: 19 tests,覆盖 EvidenceGuard 9、SemanticGuard 4、DiagnosisReleaseUseCase 6。 +- Regression: 18 suites / 70 tests,0 failure/error/skipped。 +- Maven compile 和 change strict validation 通过。 +- 公开 Controller/ChatService/AiOpsService 零 diff;SemanticGuard 无 Tool ID/raw/ReactAgent/Graph/ThreadLocal/手写 loop。 diff --git a/mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md b/mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md index a1ec620..1c8ecb5 100644 --- a/mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md +++ b/mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md @@ -1,6 +1,6 @@ # ISS-014 单体 ReAct Agent、Harness 与 ACI 工具瘦身 -**状态**:实施中(阶段 0-4 已归档,下一阶段 5) +**状态**:实施中(阶段 0-5 已归档,下一阶段 6A) **严重程度**:高 **发现时间**:2026-07-20 **目标分支**:`refactor/chat-single-react-harness` @@ -1045,7 +1045,7 @@ ISS-014 是总设计 Issue,不创建跨阶段共享的 OpenSpec change。以 | 3B | `single-react-rag-log-projections` | Completed;已归档 | | 3C | `single-react-mysql-readonly-tool` | Completed;已归档 | | 4 | `single-react-diagnosis-agent` | Completed;已归档 | -| 5 | `single-react-evidence-semantic-guards` | Pending | +| 5 | `single-react-evidence-semantic-guards` | Completed;已归档 | | 6A | `single-react-chat-application-usecase` | Pending | | 6B | `single-react-chat-sse-cutover` | Pending | | 7 | `single-react-cleanup-e2e` | Pending | diff --git a/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/.archive-ready b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/.archive-ready new file mode 100644 index 0000000..395527d --- /dev/null +++ b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/.archive-ready @@ -0,0 +1 @@ +ready diff --git a/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/.committed b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/.committed new file mode 100644 index 0000000..d0fe822 --- /dev/null +++ b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/.committed @@ -0,0 +1 @@ +committed diff --git a/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/.openspec.yaml b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/.openspec.yaml new file mode 100644 index 0000000..c0a8162 --- /dev/null +++ b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/.openspec.yaml @@ -0,0 +1,2 @@ +schema: spec-driven +created: 2026-07-21 diff --git a/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/design.md b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/design.md new file mode 100644 index 0000000..d4e0097 --- /dev/null +++ b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/design.md @@ -0,0 +1,118 @@ +## Context + +阶段 2-4 已提供显式 `RunContext`、类型化重试策略、canonical invocation store、三类有界 Tool projection 和内部 `DiagnosisAgentUseCase`。当前 `DiagnosisDraft` 仍只是未发布草稿:类型反序列化不保证 ID 唯一、报告内部引用闭合,也不证明 Tool Call 属于当前 Run 或 projection 可引用;即使物理证据真实,也还需要独立判断整份报告是否被证据支持。 + +本阶段只建立内部 release boundary。公开 `ChatController`、SSE、Run persistence 和旧 Planner/Executor/Verifier/Composer 链路由阶段 6A/6B/7 处理。接口影响为 L2,新增内部 Java use case 供后续阶段消费,不改变当前外部协议。 + +## Goals / Non-Goals + +**Goals:** + +- 用确定性代码验证 Draft 结构、Analysis 引用和当前 Run canonical invocation。 +- 只从 `agent_result` 和必要的 canonical request 生成无 Tool ID、无 raw response 的 verified evidence snapshot。 +- 首次验真失败时允许一次保持用户可见语义不变的结构修复。 +- 使用同一 `ChatModel` 执行无 Tool、无记忆、无 ReAct 的隔离 SemanticGuard,并由 Harness 控制预算、超时、取消和技术重试。 +- 只发布通过两层门禁的原 Draft;所有其他可降级路径返回固定 `SafeFallback`。 + +**Non-Goals:** + +- 不切换 `/api/chat`、`/api/chat_stream` 或 `/api/ai_ops`。 +- 不实现 SSE 事件、Run 持久化、PreviousTurn、Intent Router 或最终报告渲染。 +- 不删除或改造旧 Gatekeeper、Verifier、Composer 和多 Agent 生产链路。 +- 不访问 canonical `raw_response`,不新增 Redis 索引或历史证据召回。 +- 不执行 live model/Redis/MySQL E2E;最终 live 验收留到阶段 7。 + +## Decisions + +### 1. EvidenceGuard 是 fail-closed 的纯确定性边界 + +`EvidenceGuard.validate(RunContext, DiagnosisDraft)` 分两步执行:先验证 Draft 顶层/嵌套结构、非空文本、Analysis ID 唯一性和 `based_on_analysis_ids` 闭合;再对每个 Analysis 的非空 Tool Call ID 使用 `ToolCallKeyFactory.create(runId, id)` 查询 `CanonicalInvocationStore`。 + +可引用记录必须同时满足:record `runId` 等于当前 Run、`status=READY`、`agent_result` 非空、`evidence_status=EVIDENCE_FOUND|NO_EVIDENCE`。`AnalysisKind.accepts` 固定 `NORMAL -> EVIDENCE_FOUND`、`NEGATIVE_OBSERVATION -> NO_EVIDENCE`。任何非法 key、缺失记录、跨 Run、错误/投影中记录、空 projection、状态不一致或解析错误都产生稳定 `EvidenceViolation`,不会抛出原始存储内容或进入 SemanticGuard。 + +替代方案是复用旧 `ExecutorGatekeeperService`;拒绝该方案,因为旧实现绑定 JPA `ToolInvocation`、Session/Run 双入口和 Map DTO,不能表达 Redis canonical lifecycle、`NO_EVIDENCE` kind 约束或新的 typed Draft。 + +### 2. verified snapshot 使用 Tool-specific 严格 projection adapter + +Guard 根据冻结的 `AgentToolContracts` 严格解析 `RagToolResult`、`QueryLogsToolResult`、`MysqlToolResult`,并校验 projection 内 `tool_call_id`、`evidence_status` 与 canonical record 一致。未知 Tool 一律拒绝。 + +快照结构为 `analysis_id + analysis_text + analysis_kind + verified_evidence[]`。单条 evidence 只包含 `source_type/source/scope/timestamp/excerpt/values`:RAG 展开稳定 document metadata/excerpt;日志展开 pattern/event 或带精确 query/time window 的零结果;MySQL 从 canonical request 取得逻辑 data source/SQL/params scope,从 projection 取得 columns/有界 rows 或零结果。快照不包含 Tool Call ID、Redis key、raw response 或未被 Draft 引用的调用。 + +Snapshot 同时可确定性提取去重的 `SafeFallback.VerifiedSource`,因此 release policy 无需读取 Store。替代通用 JsonNode 透传;拒绝该方案,因为它会扩大模型输入面并削弱字段级测试。 + +### 3. Evidence repair 只能改引用结构,不能成为报告作者 + +首次 Guard 失败后,`EvidenceRepair` 使用直接单轮 `ChatModel.call(Prompt)`,输入为原始 query、完整 Draft 和稳定 violation code/target;没有 Tool callback、memory 或 ReAct Agent。调用由 `context.retryPolicies().evidenceRepair()` 执行,其 `maxAttempts=1`,因此不能形成隐式 retry loop。 + +修复输出经过严格 `DiagnosisDraft` 解析后,转换为移除所有标识/引用字段的 `SemanticDraftView` 并与原 Draft 比较。Conclusion/Analysis/Action/Recommendation/Limitations 的文本、kind、顺序和人工确认标志有任何变化都拒绝;只允许 `analysis_id`、`tool_call_ids` 和 `based_on_analysis_ids` 修正。修复后完整重跑 EvidenceGuard,第二次仍失败直接 `EVIDENCE_VALIDATION_FAILED`。 + +替代方案是允许模型删除无效 Analysis 或改写结论;拒绝该方案,因为 ISS-014 固定 Diagnosis Agent 是唯一报告作者,Harness/repair 不能改变用户报告语义。 + +### 4. SemanticGuard 使用直接 ChatModel 单轮调用 + +`SemanticGuardInput` 包含原始 query、`SemanticDraftView` 和 verified snapshot。View 保留完整用户可见 Draft 语义和 Analysis IDs,但删除所有 Tool Call IDs。每个 attempt 都构造只含一个 system message 和一个 user JSON message 的新 `Prompt`,直接调用系统注入的同一 `ChatModel`;不构建 `ReactAgent`,因此没有 Tool、memory、checkpoint、Graph 或 Agent 回调。 + +输出只允许精确 JSON object `{verdict, reason}`,字段集合固定,verdict 只允许 `SUPPORTED|UNSUPPORTED`,reason 必须非空。SemanticGuard 不返回 corrected report。`UNSUPPORTED` 正常返回给 release policy,不能进入 retry classifier。 + +### 5. 单轮模型执行由共享受控边界管理 + +`GuardModelCall` 在每次调用前执行 `DiagnosisHarnessCore.beforeModelCall`,在模型实际返回后从 `ChatResponseMetadata.Usage` 记录 input/output Token,并再次检查 Run active。调用提交到注入的 `ExecutorService`,等待时间为 `min(singleAttemptTimeout, totalTimeoutRemaining)`;attempt 超时或 Run cancellation 时执行 `Future.cancel(true)`。模型即使忽略中断而迟到返回,后台任务仍记录真实 Token,并在 `core.checkActive` 处阻止迟到结果被使用。 + +SemanticGuard 通过既有 `HarnessRetryExecutor` 和 `context.retryPolicies().semanticGuard()` 执行。只有 timeout、transport、parse error、schema invalid 可进行第二次 attempt;两次使用完全相同的序列化输入。Run cancelled/budget exhausted 不转换为 Fallback,而是继续抛给阶段 6A 生命周期边界;其他最终技术失败映射 `SEMANTIC_UNAVAILABLE`。 + +### 6. Release policy 不携带未验证语义 + +`DiagnosisReleaseUseCase.execute(context, query, draft)` 固定顺序为 `EvidenceGuard -> optional repair -> EvidenceGuard -> SemanticGuard -> decision`。成功结果保留与 Guard 通过的同一个 Draft 实例/修复实例,不总结、裁剪或重排。 + +- `SUPPORTED`:`ReleaseOutcome.SUCCESS`,返回完整 Draft 和 verified snapshot。 +- `UNSUPPORTED`:`ReleaseOutcome.FALLBACK` + `SEMANTIC_UNSUPPORTED`。 +- SemanticGuard 最终技术失败:`ReleaseOutcome.FALLBACK` + `SEMANTIC_UNAVAILABLE`。 +- 首次修复调用失败、语义被改写或第二次 Guard 失败:`ReleaseOutcome.FALLBACK` + `EVIDENCE_VALIDATION_FAILED`。 + +`SafeFallbackFactory` 使用固定模板。Evidence failure 的 `verified_sources=[]`;Semantic fallback 只从 snapshot 提取来源。所有 Fallback 的 `conclusion=null`,且不接收 Draft 文本、SemanticGuard reason、供应商错误或异常作为模板参数。 +Fallback release result 不保留完整 snapshot,因为 snapshot 含 Analysis text;已验真来源只以 `SafeFallback.verified_sources` 的安全子集返回,防止误序列化内部 result 泄漏草稿。 + +## Module and Ownership Audit + +```text +query + DiagnosisDraft + RunContext + -> DiagnosisReleaseUseCase + -> EvidenceGuard -> CanonicalInvocationStore + Tool projection parsers + -> optional EvidenceRepair -> GuardModelCall -> shared ChatModel + -> SemanticGuardInput(snapshot + ID-free Draft) + -> SemanticGuard -> HarnessRetryExecutor -> GuardModelCall -> shared ChatModel + -> original Draft OR SafeFallbackFactory +``` + +- `RunContext` owns deadline、cancellation、budget、retry policy 和 lifecycle;本阶段不创建第二套 Run 状态。 +- canonical store owns Tool 调用真实性;EvidenceGuard 只读并生成隔离 snapshot,不持有 Redis client。 +- Diagnosis Agent owns report semantics;repair 只修标识结构,SemanticGuard 只审查,release 只选择 Draft 或模板。 +- 最大耦合风险是阻塞式 `ChatModel` 的取消能力;受控 Future 可发出 interrupt,但 provider 是否立即终止取决于 SDK,故迟到结果必须由 Core active check 丢弃。 +- 无新 ADR:模块边界均来自 ISS-014 已冻结设计,且公开消费者尚未切换。 + +## Interface Impact + +- Level: L2 internal interface。 +- 新增内部 records/interfaces/use cases;不修改现有 Java 方法签名。 +- 当前消费者仅 focused tests;阶段 6A 将装配真实 `DiagnosisAgentUseCase -> DiagnosisReleaseUseCase`。 +- HTTP/SSE、Controller DTO、JPA/Flyway、Redis key schema、旧 Chat/AiOps 行为均不变,无迁移或回滚数据操作。 + +## Risks / Trade-offs + +- [Projection 字段遗漏导致误判] -> 三类显式 adapter、ID/status 双重校验和 contract fixture tests。 +- [Repair 改写报告] -> `SemanticDraftView` 确定性相等检查,任何可见差异直接 Evidence fallback。 +- [模型忽略 interrupt] -> Future cancel、迟到 Token 记录、Core active check;不允许迟到结果进入 release。 +- [业务 `UNSUPPORTED` 被重试] -> verdict 作为正常返回值,classifier 只处理异常。 +- [Fallback 泄漏草稿或内部 reason] -> 固定 factory 不接受这些参数,并对序列化结果做负向断言。 +- [输入快照过大] -> repair/SemanticGuard 独立 UTF-8 input/output limits,并占用 Run byte/model/token budget。 + +## Migration Plan + +1. 新增内部 guard/release 包和聚焦测试,不连接公开入口。 +2. 阶段 6A 将内部 Chat application use case 装配到这些接口。 +3. 阶段 6B 在新 release boundary 通过后原子切换公开 SSE。 +4. 如阶段 5 回滚,只需移除新增内部包;阶段 4 和旧生产链路不受影响。 + +## Open Questions + +- None。单次/总超时与字节预算由构造配置提供可调默认值,具体生产数值按阶段 7 Trace 校准,不构成本阶段方向问题。 diff --git a/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/proposal.md b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/proposal.md new file mode 100644 index 0000000..d363a19 --- /dev/null +++ b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/proposal.md @@ -0,0 +1,30 @@ +## Why + +阶段 4 已能生成带框架 Tool Call ID 的 `DiagnosisDraft`,但草稿尚未经过当前 Run 证据验真和隔离语义审查,不能安全发布。阶段 5 需要补齐确定性 EvidenceGuard、单轮 SemanticGuard 和 fail-closed 释放策略,作为阶段 6A/6B 接入公开链路的前置门禁。 + +## What Changes + +- 新增确定性 EvidenceGuard,校验 Draft 结构、Analysis ID、报告内部引用,以及当前 Run canonical Tool invocation 的生命周期、证据状态和 `agent_result`。 +- 从三类冻结 Tool projection 确定性生成按 Analysis 分组的 verified evidence snapshot;不读取 `raw_response`,不向 SemanticGuard 暴露 Tool Call ID 或 Redis 细节。 +- 首次 EvidenceGuard 失败时允许一次无 Tool、无 ReAct 的结构修复;修复后仍失败直接返回 `EVIDENCE_VALIDATION_FAILED`。 +- 新增复用系统同一 `ChatModel` 的隔离单轮 SemanticGuard,校验完整报告语义,只输出 `SUPPORTED|UNSUPPORTED + reason`,不生成或修改报告。 +- Harness 对 SemanticGuard 执行输入预算、模型/Token 预算、单次超时、取消、严格输出解析和最多两次技术 attempt;`UNSUPPORTED` 不重试。 +- 新增 release use case:`SUPPORTED` 原样释放 Draft,业务不支持或技术不可用返回固定 SafeFallback,任何失败均不泄漏 Draft 或 SemanticGuard reason。 +- 不切换公开 Chat/AiOps,不删除旧 Gatekeeper/Verifier/Composer,不修改 SSE 或持久化协议。 + +## Capabilities + +### New Capabilities + +- `single-react-evidence-semantic-guards`: 定义 DiagnosisDraft 的物理验真、verified snapshot、单次结构修复、隔离语义审查和安全释放行为。 + +### Modified Capabilities + +- None. 既有 Agent、Tool、Harness Core 和公开 Chat capability 的需求语义不在本阶段改变。 + +## Impact + +- 新增 `com.superbiz.agent.harness.guard.evidence`、`guard.semantic` 和 `release` 内部包,以及 SemanticGuard prompt 和 focused tests。 +- 复用 `CanonicalInvocationStore`、`ToolCallKeyFactory`、三类 Tool contract、`DiagnosisHarnessCore`、`HarnessRetryExecutor`、`RunContext`、`DiagnosisDraft`、`SafeFallback` 和系统 `ChatModel`。 +- 内部接口影响为 L2:阶段 6A 将消费新的 release use case;本阶段无公开 Controller、SSE、DTO、数据库或旧 ChatService 行为变化。 +- 主要风险是快照字段遗漏、超时任务残留、错误重试业务 `UNSUPPORTED`,以及 Fallback 意外携带未验证内容;设计和测试必须逐项 fail closed。 diff --git a/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/specs/single-react-evidence-semantic-guards/spec.md b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/specs/single-react-evidence-semantic-guards/spec.md new file mode 100644 index 0000000..edda602 --- /dev/null +++ b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/specs/single-react-evidence-semantic-guards/spec.md @@ -0,0 +1,112 @@ +## ADDED Requirements + +### Requirement: Deterministic Draft and evidence validation +The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs. + +#### Scenario: Duplicate or missing Analysis ID +- **WHEN** a Draft contains a blank or duplicate Analysis ID +- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked + +#### Scenario: Broken report reference +- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference +- **THEN** EvidenceGuard rejects the Draft before semantic review + +### Requirement: Current Run canonical invocation ownership +EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`. + +#### Scenario: Fabricated or cross-Run Tool Call +- **WHEN** a Draft references a missing Tool Call or the resolved record belongs to another Run +- **THEN** EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard + +#### Scenario: Failed or incomplete invocation +- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR` +- **THEN** EvidenceGuard rejects the Draft + +### Requirement: Analysis kind matches evidence semantics +EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls. + +#### Scenario: Valid negative observation +- **WHEN** a `NEGATIVE_OBSERVATION` references a READY `NO_EVIDENCE` projection with its query scope and zero-match data +- **THEN** EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion + +#### Scenario: Positive claim uses no-evidence result +- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation +- **THEN** EvidenceGuard rejects the binding + +### Requirement: Verified evidence snapshot is minimal and deterministic +The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations. + +#### Scenario: Supported RAG, log, and MySQL projections +- **WHEN** a Draft references valid RAG, log, or MySQL calls +- **THEN** the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis + +#### Scenario: Projection contract mismatch +- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record +- **THEN** EvidenceGuard fails closed + +### Requirement: Evidence repair is single-turn and semantics-preserving +On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics. + +#### Scenario: Structural repair succeeds +- **WHEN** the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard +- **THEN** the repaired Draft proceeds to SemanticGuard + +#### Scenario: Repair changes report text +- **WHEN** the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag +- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` + +#### Scenario: Second validation fails +- **WHEN** the repaired Draft still fails EvidenceGuard +- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard + +### Requirement: Isolated single-turn SemanticGuard +SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent. + +#### Scenario: Semantic input isolation +- **WHEN** a verified Draft enters SemanticGuard +- **THEN** the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition + +#### Scenario: Binary review output +- **WHEN** SemanticGuard completes normally +- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report + +### Requirement: Semantic model budgets timeout cancellation and retry +The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts. + +#### Scenario: Technical failure then success +- **WHEN** the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict +- **THEN** exactly two model attempts are recorded and the second verdict controls release + +#### Scenario: Unsupported is not retried +- **WHEN** SemanticGuard returns valid `UNSUPPORTED` +- **THEN** the Harness records one attempt and immediately applies the unsupported fallback + +#### Scenario: Run cancellation during model call +- **WHEN** the Run is cancelled while a guard model call is pending +- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback + +### Requirement: Fail-closed release policy +The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result. + +#### Scenario: Supported report release +- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED` +- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing + +#### Scenario: Unsupported report fallback +- **WHEN** SemanticGuard returns `UNSUPPORTED` +- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot + +#### Scenario: Semantic review remains unavailable +- **WHEN** all permitted technical attempts fail +- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed + +#### Scenario: Evidence validation fallback sources +- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft +- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty + +### Requirement: Stage-five public isolation +The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior. + +#### Scenario: Focused implementation scope +- **WHEN** stage-five changes are inspected +- **THEN** only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed diff --git a/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/tasks.md b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/tasks.md new file mode 100644 index 0000000..0cf564b --- /dev/null +++ b/openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards/tasks.md @@ -0,0 +1,25 @@ +## 1. Evidence validation and snapshot + +- [x] 1.1 Add typed EvidenceViolation, EvidenceGuardResult and verified snapshot contracts with immutable source extraction. +- [x] 1.2 Implement Draft structure/reference validation and current-Run canonical invocation checks. +- [x] 1.3 Implement strict RAG/log/MySQL projection adapters that build ID-free, raw-free evidence grouped by Analysis ID. +- [x] 1.4 Add focused tests for duplicate/missing IDs, empty/broken references, fabricated/cross-Run calls, invalid lifecycle/result, kind/status mismatch and valid negative observations. + +## 2. Guard model boundary and semantic review + +- [x] 2.1 Add the shared single-turn GuardModelCall with Core model/Token budgets, UTF-8 limits, per-attempt timeout, total timeout support and Future cancellation. +- [x] 2.2 Add SemanticDraftView and SemanticGuardInput that preserve all user-visible report semantics while excluding Tool Call IDs. +- [x] 2.3 Implement strict binary SemanticGuard output parsing and HarnessRetryExecutor integration with stable attempt auditing. +- [x] 2.4 Add tests proving input isolation, same-model direct calls, retryable technical failures, non-retried UNSUPPORTED, timeout cancellation and no Tool/ReAct loop. + +## 3. Evidence repair and release policy + +- [x] 3.1 Implement one-attempt no-Tool EvidenceRepair and reject any repair that changes SemanticDraftView. +- [x] 3.2 Implement fixed SafeFallbackFactory and release result contracts without Draft/reason parameters on fallback paths. +- [x] 3.3 Implement DiagnosisReleaseUseCase sequencing Guard, optional repair, SemanticGuard and fail-closed outcomes while propagating cancellation/budget exhaustion. +- [x] 3.4 Add release tests for successful repair, second validation failure, semantic support/unsupported/unavailable, fallback types, empty evidence-failure sources and no Draft/reason leakage. + +## 4. Verification and scope + +- [x] 4.1 Run stage-five focused tests plus Harness/Tool/Agent regression tests and Maven compile. +- [x] 4.2 Verify strict OpenSpec validation, no public Controller/ChatService/AiOps diff, no raw response/Tool ID in semantic input, and no hidden retry or handwritten Agent loop. diff --git a/openspec/specs/single-react-evidence-semantic-guards/spec.md b/openspec/specs/single-react-evidence-semantic-guards/spec.md new file mode 100644 index 0000000..67c1d72 --- /dev/null +++ b/openspec/specs/single-react-evidence-semantic-guards/spec.md @@ -0,0 +1,115 @@ +# single-react-evidence-semantic-guards Specification + +## Purpose +TBD - created by archiving change single-react-evidence-semantic-guards. Update Purpose after archive. +## Requirements +### Requirement: Deterministic Draft and evidence validation +The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs. + +#### Scenario: Duplicate or missing Analysis ID +- **WHEN** a Draft contains a blank or duplicate Analysis ID +- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked + +#### Scenario: Broken report reference +- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference +- **THEN** EvidenceGuard rejects the Draft before semantic review + +### Requirement: Current Run canonical invocation ownership +EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`. + +#### Scenario: Fabricated or cross-Run Tool Call +- **WHEN** a Draft references a missing Tool Call or the resolved record belongs to another Run +- **THEN** EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard + +#### Scenario: Failed or incomplete invocation +- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR` +- **THEN** EvidenceGuard rejects the Draft + +### Requirement: Analysis kind matches evidence semantics +EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls. + +#### Scenario: Valid negative observation +- **WHEN** a `NEGATIVE_OBSERVATION` references a READY `NO_EVIDENCE` projection with its query scope and zero-match data +- **THEN** EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion + +#### Scenario: Positive claim uses no-evidence result +- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation +- **THEN** EvidenceGuard rejects the binding + +### Requirement: Verified evidence snapshot is minimal and deterministic +The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations. + +#### Scenario: Supported RAG, log, and MySQL projections +- **WHEN** a Draft references valid RAG, log, or MySQL calls +- **THEN** the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis + +#### Scenario: Projection contract mismatch +- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record +- **THEN** EvidenceGuard fails closed + +### Requirement: Evidence repair is single-turn and semantics-preserving +On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics. + +#### Scenario: Structural repair succeeds +- **WHEN** the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard +- **THEN** the repaired Draft proceeds to SemanticGuard + +#### Scenario: Repair changes report text +- **WHEN** the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag +- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` + +#### Scenario: Second validation fails +- **WHEN** the repaired Draft still fails EvidenceGuard +- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard + +### Requirement: Isolated single-turn SemanticGuard +SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent. + +#### Scenario: Semantic input isolation +- **WHEN** a verified Draft enters SemanticGuard +- **THEN** the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition + +#### Scenario: Binary review output +- **WHEN** SemanticGuard completes normally +- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report + +### Requirement: Semantic model budgets timeout cancellation and retry +The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts. + +#### Scenario: Technical failure then success +- **WHEN** the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict +- **THEN** exactly two model attempts are recorded and the second verdict controls release + +#### Scenario: Unsupported is not retried +- **WHEN** SemanticGuard returns valid `UNSUPPORTED` +- **THEN** the Harness records one attempt and immediately applies the unsupported fallback + +#### Scenario: Run cancellation during model call +- **WHEN** the Run is cancelled while a guard model call is pending +- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback + +### Requirement: Fail-closed release policy +The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result. + +#### Scenario: Supported report release +- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED` +- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing + +#### Scenario: Unsupported report fallback +- **WHEN** SemanticGuard returns `UNSUPPORTED` +- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot + +#### Scenario: Semantic review remains unavailable +- **WHEN** all permitted technical attempts fail +- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed + +#### Scenario: Evidence validation fallback sources +- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft +- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty + +### Requirement: Stage-five public isolation +The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior. + +#### Scenario: Focused implementation scope +- **WHEN** stage-five changes are inspected +- **THEN** only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed diff --git a/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuard.java b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuard.java new file mode 100644 index 0000000..3becc55 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuard.java @@ -0,0 +1,435 @@ +package com.superbiz.agent.harness.guard.evidence; + +import com.fasterxml.jackson.databind.DeserializationFeature; +import com.fasterxml.jackson.databind.ObjectMapper; +import com.fasterxml.jackson.databind.ObjectReader; +import com.superbiz.agent.harness.contract.DiagnosisDraft; +import com.superbiz.agent.harness.core.RunContext; +import com.superbiz.agent.harness.tool.contract.AgentToolContracts; +import com.superbiz.agent.harness.tool.contract.LogEvent; +import com.superbiz.agent.harness.tool.contract.LogPattern; +import com.superbiz.agent.harness.tool.contract.LogQueryScope; +import com.superbiz.agent.harness.tool.contract.MysqlToolRequest; +import com.superbiz.agent.harness.tool.contract.MysqlToolResult; +import com.superbiz.agent.harness.tool.contract.QueryLogsToolResult; +import com.superbiz.agent.harness.tool.contract.RagEvidence; +import com.superbiz.agent.harness.tool.contract.RagToolResult; +import com.superbiz.agent.harness.tool.store.CanonicalInvocationStore; +import com.superbiz.agent.harness.tool.store.CanonicalToolInvocation; +import com.superbiz.agent.harness.tool.store.ToolCallKeyFactory; + +import java.util.ArrayList; +import java.util.HashSet; +import java.util.LinkedHashMap; +import java.util.List; +import java.util.Map; +import java.util.Objects; +import java.util.Set; + +public final class EvidenceGuard { + + private final CanonicalInvocationStore store; + private final ToolCallKeyFactory keyFactory; + private final ObjectReader ragReader; + private final ObjectReader logsReader; + private final ObjectReader mysqlRequestReader; + private final ObjectReader mysqlResultReader; + + public EvidenceGuard(CanonicalInvocationStore store, + ToolCallKeyFactory keyFactory, + ObjectMapper objectMapper) { + this.store = Objects.requireNonNull(store, "store must not be null"); + this.keyFactory = Objects.requireNonNull(keyFactory, "keyFactory must not be null"); + ObjectMapper mapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null"); + this.ragReader = mapper.readerFor(RagToolResult.class) + .with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES) + .with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS); + this.logsReader = mapper.readerFor(QueryLogsToolResult.class) + .with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES) + .with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS); + this.mysqlRequestReader = mapper.readerFor(MysqlToolRequest.class) + .with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES) + .with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS); + this.mysqlResultReader = mapper.readerFor(MysqlToolResult.class) + .with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES) + .with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS); + } + + public EvidenceGuardResult validate(RunContext context, DiagnosisDraft draft) { + Objects.requireNonNull(context, "context must not be null"); + List violations = validateDraft(draft); + if (!violations.isEmpty()) { + return EvidenceGuardResult.invalid(violations); + } + + List verifiedAnalyses = new ArrayList<>(); + for (int index = 0; index < draft.analysis().size(); index++) { + DiagnosisDraft.AnalysisItem analysis = draft.analysis().get(index); + List evidence = new ArrayList<>(); + for (String toolCallId : analysis.toolCallIds()) { + verifyInvocation(context, analysis, index, toolCallId, evidence, violations); + } + if (!evidence.isEmpty()) { + verifiedAnalyses.add(new VerifiedAnalysisEvidence( + analysis.analysisId(), analysis.text(), analysis.kind(), evidence)); + } + } + return violations.isEmpty() + ? EvidenceGuardResult.valid(new VerifiedEvidenceSnapshot(verifiedAnalyses)) + : EvidenceGuardResult.invalid(violations); + } + + private List validateDraft(DiagnosisDraft draft) { + List violations = new ArrayList<>(); + if (draft == null) { + violations.add(violation(EvidenceViolationCode.DRAFT_MISSING, "draft")); + return violations; + } + if (draft.analysis().isEmpty()) { + violations.add(violation(EvidenceViolationCode.ANALYSIS_MISSING, "analysis")); + } + Set ids = new HashSet<>(); + for (int index = 0; index < draft.analysis().size(); index++) { + DiagnosisDraft.AnalysisItem item = draft.analysis().get(index); + String target = "analysis[" + index + "]"; + if (item == null) { + violations.add(violation(EvidenceViolationCode.ANALYSIS_MISSING, target)); + continue; + } + if (isBlank(item.analysisId())) { + violations.add(violation(EvidenceViolationCode.ANALYSIS_ID_MISSING, target + ".analysis_id")); + } else if (!ids.add(item.analysisId())) { + violations.add(violation(EvidenceViolationCode.ANALYSIS_ID_DUPLICATE, target + ".analysis_id")); + } + if (item.kind() == null) { + violations.add(violation(EvidenceViolationCode.ANALYSIS_KIND_MISSING, target + ".kind")); + } + if (isBlank(item.text())) { + violations.add(violation(EvidenceViolationCode.ANALYSIS_TEXT_MISSING, target + ".text")); + } + if (item.toolCallIds().isEmpty()) { + violations.add(violation(EvidenceViolationCode.TOOL_REFERENCE_MISSING, target + ".tool_call_ids")); + } + } + validateReportReferences(draft, ids, violations); + return violations; + } + + private void validateReportReferences(DiagnosisDraft draft, Set ids, + List violations) { + if (draft.conclusion() != null) { + validateTextAndReferences("conclusion", draft.conclusion().text(), + draft.conclusion().basedOnAnalysisIds(), ids, violations); + } + for (int index = 0; index < draft.actionPlan().size(); index++) { + DiagnosisDraft.ActionPlanItem item = draft.actionPlan().get(index); + String target = "action_plan[" + index + "]"; + if (item == null) { + violations.add(violation(EvidenceViolationCode.REPORT_TEXT_MISSING, target)); + } else { + validateTextAndReferences(target, item.action(), item.basedOnAnalysisIds(), ids, violations); + } + } + for (int index = 0; index < draft.recommendations().size(); index++) { + DiagnosisDraft.Recommendation item = draft.recommendations().get(index); + String target = "recommendations[" + index + "]"; + if (item == null) { + violations.add(violation(EvidenceViolationCode.REPORT_TEXT_MISSING, target)); + } else { + validateTextAndReferences(target, item.text(), item.basedOnAnalysisIds(), ids, violations); + } + } + if (draft.limitations() == null || isBlank(draft.limitations().scope())) { + violations.add(violation(EvidenceViolationCode.LIMITATIONS_MISSING, "limitations")); + } + } + + private void validateTextAndReferences(String target, String text, List references, + Set ids, List violations) { + if (isBlank(text)) { + violations.add(violation(EvidenceViolationCode.REPORT_TEXT_MISSING, target)); + } + if (references == null || references.isEmpty()) { + violations.add(violation(EvidenceViolationCode.ANALYSIS_REFERENCE_MISSING, + target + ".based_on_analysis_ids")); + return; + } + for (int index = 0; index < references.size(); index++) { + String reference = references.get(index); + if (isBlank(reference)) { + violations.add(violation(EvidenceViolationCode.ANALYSIS_REFERENCE_MISSING, + target + ".based_on_analysis_ids[" + index + "]")); + } else if (!ids.contains(reference)) { + violations.add(violation(EvidenceViolationCode.ANALYSIS_REFERENCE_UNKNOWN, + target + ".based_on_analysis_ids[" + index + "]")); + } + } + } + + private void verifyInvocation(RunContext context, DiagnosisDraft.AnalysisItem analysis, + int analysisIndex, String toolCallId, + List evidence, + List violations) { + String target = "analysis[" + analysisIndex + "].tool_call_ids"; + if (isBlank(toolCallId)) { + violations.add(violation(EvidenceViolationCode.TOOL_REFERENCE_INVALID, target)); + return; + } + String key; + try { + key = keyFactory.create(context.runId(), toolCallId); + } catch (IllegalArgumentException exception) { + violations.add(violation(EvidenceViolationCode.TOOL_REFERENCE_INVALID, target)); + return; + } + CanonicalToolInvocation invocation; + try { + invocation = store.find(key).orElse(null); + } catch (RuntimeException exception) { + violations.add(violation(EvidenceViolationCode.CANONICAL_LOOKUP_FAILED, target)); + return; + } + if (invocation == null) { + violations.add(violation(EvidenceViolationCode.INVOCATION_MISSING, target)); + return; + } + if (!Objects.equals(toolCallId, invocation.toolCallId())) { + violations.add(violation(EvidenceViolationCode.INVOCATION_ID_MISMATCH, target)); + return; + } + if (!invocation.isReferencableBy(context.runId()) || invocation.agentResult().isBlank()) { + violations.add(violation(EvidenceViolationCode.INVOCATION_NOT_REFERENCABLE, target)); + return; + } + if (!analysis.kind().accepts(invocation.evidenceStatus())) { + violations.add(violation(EvidenceViolationCode.EVIDENCE_KIND_MISMATCH, target)); + return; + } + switch (invocation.toolName()) { + case AgentToolContracts.LOOKUP_KNOWLEDGE -> + readRag(invocation, evidence, target, violations); + case AgentToolContracts.QUERY_LOGS -> + readLogs(invocation, evidence, target, violations); + case AgentToolContracts.QUERY_MYSQL -> + readMysql(invocation, evidence, target, violations); + default -> violations.add(violation(EvidenceViolationCode.TOOL_UNSUPPORTED, target)); + } + } + + private void readRag(CanonicalToolInvocation invocation, List evidence, + String target, List violations) { + RagToolResult result; + try { + result = ragReader.readValue(invocation.agentResult()); + } catch (Exception exception) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + if (!Objects.equals(result.toolCallId(), invocation.toolCallId())) { + violations.add(violation(EvidenceViolationCode.PROJECTION_ID_MISMATCH, target)); + return; + } + if (result.evidenceStatus() != invocation.evidenceStatus()) { + violations.add(violation(EvidenceViolationCode.PROJECTION_STATUS_MISMATCH, target)); + return; + } + if (result.returnedCount() != result.evidence().size()) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + if (result.evidenceStatus() == com.superbiz.agent.harness.contract.EvidenceStatus.NO_EVIDENCE) { + if (!result.evidence().isEmpty() || result.returnedCount() != 0) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + evidence.add(new VerifiedEvidence( + "RAG", "KNOWLEDGE_BASE", firstText(result.query(), "knowledge query"), + null, "No evidence matched the query scope", + Map.of("match_count", 0, "truncated", result.truncated()))); + return; + } + if (result.evidence().isEmpty()) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + for (RagEvidence item : result.evidence()) { + if (item == null || isBlank(item.documentId()) || isBlank(item.excerpt())) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + Map values = new LinkedHashMap<>(); + values.put("document_id", item.documentId()); + putIfText(values, "title", item.title()); + putIfText(values, "breadcrumb", item.breadcrumb()); + evidence.add(new VerifiedEvidence( + "RAG", firstText(item.source(), item.title(), item.documentId()), + firstText(result.query(), "knowledge query"), null, item.excerpt(), values)); + } + } + + private void readLogs(CanonicalToolInvocation invocation, List evidence, + String target, List violations) { + QueryLogsToolResult result; + try { + result = logsReader.readValue(invocation.agentResult()); + } catch (Exception exception) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + if (!projectionMatches(invocation, result.toolCallId(), result.evidenceStatus(), + target, violations)) { + return; + } + if (result.sourceKind() == null || result.scope() == null + || result.scope().topic() == null || isBlank(result.scope().query()) + || isBlank(result.scope().startTime()) || isBlank(result.scope().endTime()) + || result.matchCount() < 0 || result.returnedCount() != result.events().size()) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + String source = result.scope().topic().name() + " (" + result.sourceKind().name() + ")"; + String scope = logScope(result.scope()); + if (result.evidenceStatus() == com.superbiz.agent.harness.contract.EvidenceStatus.NO_EVIDENCE) { + if (result.matchCount() != 0 || !result.patterns().isEmpty() || !result.events().isEmpty()) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + evidence.add(new VerifiedEvidence( + "LOG", source, scope, null, "No log events matched the query scope", + Map.of("match_count", 0, "truncated", result.truncated()))); + return; + } + if (result.patterns().isEmpty() && result.events().isEmpty()) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + for (LogPattern pattern : result.patterns()) { + if (pattern == null || pattern.count() <= 0 || isBlank(pattern.example())) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + Map values = new LinkedHashMap<>(); + values.put("count", pattern.count()); + putIfText(values, "first_seen", pattern.firstSeen()); + putIfText(values, "last_seen", pattern.lastSeen()); + putIfText(values, "level", pattern.level()); + putIfText(values, "service", pattern.service()); + evidence.add(new VerifiedEvidence( + "LOG", source, scope, pattern.lastSeen(), pattern.example(), values)); + } + for (LogEvent event : result.events()) { + if (event == null || isBlank(event.message())) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + Map values = new LinkedHashMap<>(); + putIfText(values, "level", event.level()); + putIfText(values, "service", event.service()); + evidence.add(new VerifiedEvidence( + "LOG", source, scope, event.timestamp(), event.message(), values)); + } + } + + private boolean projectionMatches(CanonicalToolInvocation invocation, String toolCallId, + com.superbiz.agent.harness.contract.EvidenceStatus evidenceStatus, + String target, List violations) { + if (!Objects.equals(toolCallId, invocation.toolCallId())) { + violations.add(violation(EvidenceViolationCode.PROJECTION_ID_MISMATCH, target)); + return false; + } + if (evidenceStatus != invocation.evidenceStatus()) { + violations.add(violation(EvidenceViolationCode.PROJECTION_STATUS_MISMATCH, target)); + return false; + } + return true; + } + + private void readMysql(CanonicalToolInvocation invocation, List evidence, + String target, List violations) { + MysqlToolRequest request; + MysqlToolResult result; + try { + request = mysqlRequestReader.readValue(invocation.request()); + result = mysqlResultReader.readValue(invocation.agentResult()); + } catch (Exception exception) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + if (!projectionMatches(invocation, result.toolCallId(), result.evidenceStatus(), + target, violations)) { + return; + } + if (isBlank(request.dataSource()) || isBlank(request.sql()) + || result.returnedCount() != result.rows().size() + || result.columns().isEmpty() || hasInvalidColumns(result.columns())) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + String scope = "sql=" + request.sql() + " params=" + request.params(); + if (result.evidenceStatus() == com.superbiz.agent.harness.contract.EvidenceStatus.NO_EVIDENCE) { + if (!result.rows().isEmpty() || result.returnedCount() != 0) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + evidence.add(new VerifiedEvidence( + "MYSQL", request.dataSource(), scope, null, + "No rows matched the query scope", + Map.of("match_count", 0, "columns", result.columns(), + "truncated", result.truncated()))); + return; + } + if (result.rows().isEmpty()) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + for (int index = 0; index < result.rows().size(); index++) { + Map row = result.rows().get(index); + if (row == null || !row.keySet().equals(new java.util.LinkedHashSet<>(result.columns()))) { + violations.add(violation(EvidenceViolationCode.PROJECTION_INVALID, target)); + return; + } + Map values = new LinkedHashMap<>(row); + values.put("_row_number", index + 1); + evidence.add(new VerifiedEvidence( + "MYSQL", request.dataSource(), scope, null, null, values)); + } + } + + private static boolean hasInvalidColumns(List columns) { + Set unique = new HashSet<>(); + for (String column : columns) { + if (isBlank(column) || !unique.add(column)) { + return true; + } + } + return false; + } + + private static String logScope(LogQueryScope scope) { + return scope.topic().name() + " query=" + scope.query() + + " from=" + scope.startTime() + " to=" + scope.endTime(); + } + + private static EvidenceViolation violation(EvidenceViolationCode code, String target) { + return new EvidenceViolation(code, target); + } + + private static void putIfText(Map values, String name, String value) { + if (!isBlank(value)) { + values.put(name, value); + } + } + + private static String firstText(String... values) { + for (String value : values) { + if (!isBlank(value)) { + return value; + } + } + return "unknown"; + } + + private static boolean isBlank(String value) { + return value == null || value.isBlank(); + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuardResult.java b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuardResult.java new file mode 100644 index 0000000..0c777e2 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuardResult.java @@ -0,0 +1,33 @@ +package com.superbiz.agent.harness.guard.evidence; + +import java.util.List; +import java.util.Optional; + +public record EvidenceGuardResult( + List violations, + VerifiedEvidenceSnapshot snapshot) { + + public EvidenceGuardResult { + violations = violations == null ? List.of() : List.copyOf(violations); + if (violations.isEmpty() == (snapshot == null)) { + throw new IllegalArgumentException( + "valid result requires snapshot and invalid result requires violations"); + } + } + + public static EvidenceGuardResult valid(VerifiedEvidenceSnapshot snapshot) { + return new EvidenceGuardResult(List.of(), snapshot); + } + + public static EvidenceGuardResult invalid(List violations) { + return new EvidenceGuardResult(violations, null); + } + + public boolean valid() { + return violations.isEmpty(); + } + + public Optional verifiedSnapshot() { + return Optional.ofNullable(snapshot); + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceViolation.java b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceViolation.java new file mode 100644 index 0000000..f5a3fbe --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceViolation.java @@ -0,0 +1,17 @@ +package com.superbiz.agent.harness.guard.evidence; + +import com.fasterxml.jackson.annotation.JsonProperty; + +import java.util.Objects; + +public record EvidenceViolation( + @JsonProperty("code") EvidenceViolationCode code, + @JsonProperty("target") String target) { + + public EvidenceViolation { + Objects.requireNonNull(code, "code must not be null"); + if (target == null || target.isBlank()) { + throw new IllegalArgumentException("target must not be blank"); + } + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceViolationCode.java b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceViolationCode.java new file mode 100644 index 0000000..4894627 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/evidence/EvidenceViolationCode.java @@ -0,0 +1,25 @@ +package com.superbiz.agent.harness.guard.evidence; + +public enum EvidenceViolationCode { + DRAFT_MISSING, + ANALYSIS_MISSING, + ANALYSIS_ID_MISSING, + ANALYSIS_ID_DUPLICATE, + ANALYSIS_KIND_MISSING, + ANALYSIS_TEXT_MISSING, + TOOL_REFERENCE_MISSING, + REPORT_TEXT_MISSING, + ANALYSIS_REFERENCE_MISSING, + ANALYSIS_REFERENCE_UNKNOWN, + LIMITATIONS_MISSING, + TOOL_REFERENCE_INVALID, + INVOCATION_MISSING, + INVOCATION_ID_MISMATCH, + INVOCATION_NOT_REFERENCABLE, + EVIDENCE_KIND_MISMATCH, + TOOL_UNSUPPORTED, + PROJECTION_INVALID, + PROJECTION_ID_MISMATCH, + PROJECTION_STATUS_MISMATCH, + CANONICAL_LOOKUP_FAILED +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/evidence/VerifiedAnalysisEvidence.java b/src/main/java/com/superbiz/agent/harness/guard/evidence/VerifiedAnalysisEvidence.java new file mode 100644 index 0000000..8f32bb5 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/evidence/VerifiedAnalysisEvidence.java @@ -0,0 +1,30 @@ +package com.superbiz.agent.harness.guard.evidence; + +import com.fasterxml.jackson.annotation.JsonProperty; +import com.superbiz.agent.harness.contract.AnalysisKind; + +import java.util.List; +import java.util.Objects; + +public record VerifiedAnalysisEvidence( + @JsonProperty("analysis_id") String analysisId, + @JsonProperty("analysis_text") String analysisText, + @JsonProperty("analysis_kind") AnalysisKind analysisKind, + @JsonProperty("verified_evidence") List evidence) { + + public VerifiedAnalysisEvidence { + requireText(analysisId, "analysisId"); + requireText(analysisText, "analysisText"); + Objects.requireNonNull(analysisKind, "analysisKind must not be null"); + evidence = evidence == null ? List.of() : List.copyOf(evidence); + if (evidence.isEmpty()) { + throw new IllegalArgumentException("evidence must not be empty"); + } + } + + private static void requireText(String value, String name) { + if (value == null || value.isBlank()) { + throw new IllegalArgumentException(name + " must not be blank"); + } + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/evidence/VerifiedEvidence.java b/src/main/java/com/superbiz/agent/harness/guard/evidence/VerifiedEvidence.java new file mode 100644 index 0000000..98c61cf --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/evidence/VerifiedEvidence.java @@ -0,0 +1,35 @@ +package com.superbiz.agent.harness.guard.evidence; + +import com.fasterxml.jackson.annotation.JsonProperty; + +import java.util.LinkedHashMap; +import java.util.Map; + +public record VerifiedEvidence( + @JsonProperty("source_type") String sourceType, + @JsonProperty("source") String source, + @JsonProperty("scope") String scope, + @JsonProperty("timestamp") String timestamp, + @JsonProperty("excerpt") String excerpt, + @JsonProperty("values") Map values) { + + public VerifiedEvidence { + requireText(sourceType, "sourceType"); + requireText(source, "source"); + requireText(scope, "scope"); + values = immutableValues(values); + } + + private static Map immutableValues(Map values) { + if (values == null || values.isEmpty()) { + return Map.of(); + } + return java.util.Collections.unmodifiableMap(new LinkedHashMap<>(values)); + } + + private static void requireText(String value, String name) { + if (value == null || value.isBlank()) { + throw new IllegalArgumentException(name + " must not be blank"); + } + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/evidence/VerifiedEvidenceSnapshot.java b/src/main/java/com/superbiz/agent/harness/guard/evidence/VerifiedEvidenceSnapshot.java new file mode 100644 index 0000000..7025c61 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/evidence/VerifiedEvidenceSnapshot.java @@ -0,0 +1,33 @@ +package com.superbiz.agent.harness.guard.evidence; + +import com.fasterxml.jackson.annotation.JsonProperty; +import com.superbiz.agent.harness.contract.SafeFallback; + +import java.util.LinkedHashMap; +import java.util.List; +import java.util.Map; + +public record VerifiedEvidenceSnapshot( + @JsonProperty("analyses") List analyses) { + + public VerifiedEvidenceSnapshot { + analyses = analyses == null ? List.of() : List.copyOf(analyses); + } + + public static VerifiedEvidenceSnapshot empty() { + return new VerifiedEvidenceSnapshot(List.of()); + } + + public List verifiedSources() { + Map unique = new LinkedHashMap<>(); + for (VerifiedAnalysisEvidence analysis : analyses) { + for (VerifiedEvidence evidence : analysis.evidence()) { + String key = evidence.sourceType() + "\u0000" + evidence.source() + + "\u0000" + evidence.scope(); + unique.putIfAbsent(key, new SafeFallback.VerifiedSource( + evidence.sourceType(), evidence.source(), evidence.scope())); + } + } + return List.copyOf(unique.values()); + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCall.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCall.java new file mode 100644 index 0000000..0b1d2d8 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCall.java @@ -0,0 +1,120 @@ +package com.superbiz.agent.harness.guard.semantic; + +import com.superbiz.agent.harness.core.BudgetExceededException; +import com.superbiz.agent.harness.core.DiagnosisHarnessCore; +import com.superbiz.agent.harness.core.RunAbortedException; +import com.superbiz.agent.harness.core.RunContext; +import com.superbiz.agent.harness.retry.RetryFailure; +import org.springframework.ai.chat.messages.AssistantMessage; +import org.springframework.ai.chat.metadata.Usage; +import org.springframework.ai.chat.model.ChatModel; +import org.springframework.ai.chat.model.ChatResponse; +import org.springframework.ai.chat.prompt.Prompt; + +import java.nio.charset.StandardCharsets; +import java.time.Duration; +import java.util.Objects; +import java.util.concurrent.CancellationException; +import java.util.concurrent.ExecutionException; +import java.util.concurrent.ExecutorService; +import java.util.concurrent.Future; +import java.util.concurrent.TimeUnit; +import java.util.concurrent.TimeoutException; + +public final class GuardModelCall { + + private final DiagnosisHarnessCore core; + private final ChatModel chatModel; + private final ExecutorService executor; + + public GuardModelCall(DiagnosisHarnessCore core, ChatModel chatModel, + ExecutorService executor) { + this.core = Objects.requireNonNull(core, "core must not be null"); + this.chatModel = Objects.requireNonNull(chatModel, "chatModel must not be null"); + this.executor = Objects.requireNonNull(executor, "executor must not be null"); + } + + public String call(RunContext context, Prompt prompt, Duration timeout, long maxOutputBytes) { + Objects.requireNonNull(context, "context must not be null"); + Objects.requireNonNull(prompt, "prompt must not be null"); + Objects.requireNonNull(timeout, "timeout must not be null"); + if (timeout.isZero() || timeout.isNegative() || maxOutputBytes <= 0) { + throw new IllegalArgumentException("timeout and output limit must be positive"); + } + core.beforeModelCall(context); + Future future = executor.submit(() -> invoke(context, prompt, maxOutputBytes)); + context.cancellation().onCancel(ignored -> future.cancel(true)); + try { + return future.get(timeout.toNanos(), TimeUnit.NANOSECONDS); + } catch (TimeoutException exception) { + future.cancel(true); + throw new GuardModelCallException( + RetryFailure.TIMEOUT, "Guard model attempt timed out", exception); + } catch (CancellationException exception) { + core.checkActive(context); + throw new GuardModelCallException( + RetryFailure.TRANSPORT, "Guard model attempt was cancelled", exception); + } catch (InterruptedException exception) { + future.cancel(true); + Thread.currentThread().interrupt(); + core.checkActive(context); + throw new GuardModelCallException( + RetryFailure.TRANSPORT, "Interrupted while waiting for guard model", exception); + } catch (ExecutionException exception) { + Throwable cause = exception.getCause(); + if (cause instanceof RunAbortedException aborted) { + throw aborted; + } + if (cause instanceof BudgetExceededException exceeded) { + throw exceeded; + } + if (cause instanceof GuardModelCallException guardFailure) { + throw guardFailure; + } + throw new GuardModelCallException( + RetryFailure.TRANSPORT, "Guard model call failed", cause); + } + } + + private String invoke(RunContext context, Prompt prompt, long maxOutputBytes) { + ChatResponse response; + try { + response = chatModel.call(prompt); + } catch (RuntimeException exception) { + throw new GuardModelCallException( + RetryFailure.TRANSPORT, "Guard model transport failed", exception); + } + recordUsage(context, response); + core.checkActive(context); + AssistantMessage output = response == null || response.getResult() == null + ? null : response.getResult().getOutput(); + if (output == null || output.getText() == null || output.getText().isBlank() + || (output.getToolCalls() != null && !output.getToolCalls().isEmpty())) { + throw new GuardModelCallException( + RetryFailure.SCHEMA_INVALID, "Guard model returned invalid message shape"); + } + long bytes = output.getText().getBytes(StandardCharsets.UTF_8).length; + if (bytes > maxOutputBytes) { + throw new GuardModelCallException( + RetryFailure.SCHEMA_INVALID, "Guard model output exceeded limit"); + } + core.reserveRunBytes(context, bytes); + return output.getText(); + } + + private void recordUsage(RunContext context, ChatResponse response) { + if (response == null || response.getMetadata() == null) { + return; + } + Usage usage = response.getMetadata().getUsage(); + if (usage == null) { + return; + } + core.recordTokens(context, nonNegative(usage.getPromptTokens()), + nonNegative(usage.getCompletionTokens())); + } + + private static long nonNegative(Integer value) { + return value == null || value < 0 ? 0L : value.longValue(); + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCallException.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCallException.java new file mode 100644 index 0000000..b2cb8b9 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/GuardModelCallException.java @@ -0,0 +1,24 @@ +package com.superbiz.agent.harness.guard.semantic; + +import com.superbiz.agent.harness.retry.RetryFailure; + +import java.util.Objects; + +public final class GuardModelCallException extends RuntimeException { + + private final RetryFailure failure; + + public GuardModelCallException(RetryFailure failure, String message) { + super(message); + this.failure = Objects.requireNonNull(failure, "failure must not be null"); + } + + public GuardModelCallException(RetryFailure failure, String message, Throwable cause) { + super(message, cause); + this.failure = Objects.requireNonNull(failure, "failure must not be null"); + } + + public RetryFailure failure() { + return failure; + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticDraftView.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticDraftView.java new file mode 100644 index 0000000..e9d134a --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticDraftView.java @@ -0,0 +1,126 @@ +package com.superbiz.agent.harness.guard.semantic; + +import com.fasterxml.jackson.annotation.JsonProperty; +import com.superbiz.agent.harness.contract.AnalysisKind; +import com.superbiz.agent.harness.contract.DiagnosisDraft; + +import java.util.List; +import java.util.Objects; + +public record SemanticDraftView( + @JsonProperty("conclusion") Conclusion conclusion, + @JsonProperty("analysis") List analysis, + @JsonProperty("action_plan") List actionPlan, + @JsonProperty("recommendations") List recommendations, + @JsonProperty("limitations") Limitations limitations) { + + public SemanticDraftView { + analysis = immutable(analysis); + actionPlan = immutable(actionPlan); + recommendations = immutable(recommendations); + } + + public static SemanticDraftView from(DiagnosisDraft draft) { + Objects.requireNonNull(draft, "draft must not be null"); + Conclusion conclusion = draft.conclusion() == null ? null : new Conclusion( + draft.conclusion().text(), draft.conclusion().basedOnAnalysisIds()); + List analysis = draft.analysis().stream() + .map(item -> new Analysis(item.analysisId(), item.kind(), item.text())) + .toList(); + List actions = draft.actionPlan().stream() + .map(item -> new Action(item.action(), item.basedOnAnalysisIds(), + item.requiresHumanConfirmation())) + .toList(); + List recommendations = draft.recommendations().stream() + .map(item -> new Recommendation(item.text(), item.basedOnAnalysisIds())) + .toList(); + Limitations limitations = draft.limitations() == null ? null : new Limitations( + draft.limitations().scope(), draft.limitations().missingInfo()); + return new SemanticDraftView(conclusion, analysis, actions, recommendations, limitations); + } + + public boolean hasSameUserVisibleSemantics(SemanticDraftView other) { + if (other == null || !Objects.equals(text(conclusion), text(other.conclusion)) + || analysis.size() != other.analysis.size() + || actionPlan.size() != other.actionPlan.size() + || recommendations.size() != other.recommendations.size()) { + return false; + } + for (int i = 0; i < analysis.size(); i++) { + Analysis left = analysis.get(i); + Analysis right = other.analysis.get(i); + if (left.kind() != right.kind() || !Objects.equals(left.text(), right.text())) { + return false; + } + } + for (int i = 0; i < actionPlan.size(); i++) { + Action left = actionPlan.get(i); + Action right = other.actionPlan.get(i); + if (!Objects.equals(left.action(), right.action()) + || left.requiresHumanConfirmation() != right.requiresHumanConfirmation()) { + return false; + } + } + for (int i = 0; i < recommendations.size(); i++) { + if (!Objects.equals(recommendations.get(i).text(), other.recommendations.get(i).text())) { + return false; + } + } + return limitationsEqual(limitations, other.limitations); + } + + private static String text(Conclusion value) { + return value == null ? null : value.text(); + } + + private static boolean limitationsEqual(Limitations left, Limitations right) { + if (left == null || right == null) { + return left == right; + } + return Objects.equals(left.scope(), right.scope()) + && Objects.equals(left.missingInfo(), right.missingInfo()); + } + + private static List immutable(List values) { + return values == null ? List.of() : List.copyOf(values); + } + + public record Conclusion( + @JsonProperty("text") String text, + @JsonProperty("based_on_analysis_ids") List basedOnAnalysisIds) { + public Conclusion { + basedOnAnalysisIds = immutable(basedOnAnalysisIds); + } + } + + public record Analysis( + @JsonProperty("analysis_id") String analysisId, + @JsonProperty("kind") AnalysisKind kind, + @JsonProperty("text") String text) { + } + + public record Action( + @JsonProperty("action") String action, + @JsonProperty("based_on_analysis_ids") List basedOnAnalysisIds, + @JsonProperty("requires_human_confirmation") boolean requiresHumanConfirmation) { + public Action { + basedOnAnalysisIds = immutable(basedOnAnalysisIds); + } + } + + public record Recommendation( + @JsonProperty("text") String text, + @JsonProperty("based_on_analysis_ids") List basedOnAnalysisIds) { + public Recommendation { + basedOnAnalysisIds = immutable(basedOnAnalysisIds); + } + } + + public record Limitations( + @JsonProperty("scope") String scope, + @JsonProperty("missing_info") List missingInfo) { + public Limitations { + missingInfo = immutable(missingInfo); + } + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuard.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuard.java new file mode 100644 index 0000000..707424d --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuard.java @@ -0,0 +1,134 @@ +package com.superbiz.agent.harness.guard.semantic; + +import com.fasterxml.jackson.core.JsonProcessingException; +import com.fasterxml.jackson.databind.JsonNode; +import com.fasterxml.jackson.databind.ObjectMapper; +import com.superbiz.agent.harness.contract.SemanticVerdict; +import com.superbiz.agent.harness.core.DiagnosisHarnessCore; +import com.superbiz.agent.harness.core.RunContext; +import com.superbiz.agent.harness.retry.HarnessRetryExecutor; +import com.superbiz.agent.harness.retry.RetryAttempt; +import com.superbiz.agent.harness.retry.RetryFailure; +import org.springframework.ai.chat.messages.SystemMessage; +import org.springframework.ai.chat.messages.UserMessage; +import org.springframework.ai.chat.prompt.Prompt; + +import java.nio.charset.StandardCharsets; +import java.time.Duration; +import java.util.HashSet; +import java.util.Iterator; +import java.util.Objects; +import java.util.Set; +import java.util.function.Consumer; + +public final class SemanticGuard { + + private static final Set OUTPUT_FIELDS = Set.of("verdict", "reason"); + + private final DiagnosisHarnessCore core; + private final HarnessRetryExecutor retryExecutor; + private final GuardModelCall modelCall; + private final ObjectMapper objectMapper; + private final SemanticGuardLimits limits; + private final Consumer attemptRecorder; + private final String prompt; + + public SemanticGuard(DiagnosisHarnessCore core, + HarnessRetryExecutor retryExecutor, + GuardModelCall modelCall, + ObjectMapper objectMapper, + SemanticGuardLimits limits, + Consumer attemptRecorder) { + this.core = Objects.requireNonNull(core, "core must not be null"); + this.retryExecutor = Objects.requireNonNull(retryExecutor, "retryExecutor must not be null"); + this.modelCall = Objects.requireNonNull(modelCall, "modelCall must not be null"); + this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null"); + this.limits = Objects.requireNonNull(limits, "limits must not be null"); + this.attemptRecorder = Objects.requireNonNull( + attemptRecorder, "attemptRecorder must not be null"); + this.prompt = SemanticGuardPrompt.load(); + } + + public SemanticGuardDecision review(RunContext context, SemanticGuardInput input) { + Objects.requireNonNull(context, "context must not be null"); + Objects.requireNonNull(input, "input must not be null"); + String inputJson = serialize(input); + long inputBytes = utf8Bytes(inputJson); + if (inputBytes > limits.maxInputBytes()) { + throw new GuardModelCallException( + RetryFailure.SCHEMA_INVALID, "SemanticGuard input exceeded limit"); + } + core.reserveRunBytes(context, inputBytes); + Prompt modelPrompt = new Prompt(java.util.List.of( + new SystemMessage(prompt), new UserMessage(inputJson))); + long startedNanos = System.nanoTime(); + return retryExecutor.execute( + context, + context.retryPolicies().semanticGuard(), + () -> parse(modelCall.call( + context, modelPrompt, remainingTimeout(startedNanos), limits.maxOutputBytes())), + this::classify, + attemptRecorder); + } + + private SemanticGuardDecision parse(String output) { + JsonNode root; + try { + root = objectMapper.readTree(output); + } catch (JsonProcessingException exception) { + throw new GuardModelCallException( + RetryFailure.PARSE_ERROR, "SemanticGuard output is not JSON", exception); + } + if (root == null || !root.isObject() || !fieldNames(root).equals(OUTPUT_FIELDS) + || !root.path("verdict").isTextual() || !root.path("reason").isTextual() + || root.path("reason").asText().isBlank()) { + throw new GuardModelCallException( + RetryFailure.SCHEMA_INVALID, "SemanticGuard output schema is invalid"); + } + SemanticVerdict verdict; + try { + verdict = SemanticVerdict.valueOf(root.path("verdict").asText()); + } catch (IllegalArgumentException exception) { + throw new GuardModelCallException( + RetryFailure.SCHEMA_INVALID, "SemanticGuard verdict is invalid", exception); + } + return new SemanticGuardDecision(verdict, root.path("reason").asText()); + } + + private Duration remainingTimeout(long startedNanos) { + long elapsed = Math.max(0L, System.nanoTime() - startedNanos); + long remaining = limits.totalTimeout().toNanos() - elapsed; + if (remaining <= 0) { + throw new GuardModelCallException( + RetryFailure.TIMEOUT, "SemanticGuard total timeout exhausted"); + } + return Duration.ofNanos(Math.min(remaining, limits.perAttemptTimeout().toNanos())); + } + + private RetryFailure classify(Exception exception) { + if (exception instanceof GuardModelCallException guardFailure) { + return guardFailure.failure(); + } + return RetryFailure.UNKNOWN; + } + + private String serialize(SemanticGuardInput input) { + try { + return objectMapper.writeValueAsString(input); + } catch (JsonProcessingException exception) { + throw new GuardModelCallException( + RetryFailure.SCHEMA_INVALID, "SemanticGuard input is not serializable", exception); + } + } + + private static Set fieldNames(JsonNode node) { + Set names = new HashSet<>(); + Iterator fields = node.fieldNames(); + fields.forEachRemaining(names::add); + return names; + } + + private static long utf8Bytes(String value) { + return value.getBytes(StandardCharsets.UTF_8).length; + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardDecision.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardDecision.java new file mode 100644 index 0000000..c9e1ad4 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardDecision.java @@ -0,0 +1,18 @@ +package com.superbiz.agent.harness.guard.semantic; + +import com.fasterxml.jackson.annotation.JsonProperty; +import com.superbiz.agent.harness.contract.SemanticVerdict; + +import java.util.Objects; + +public record SemanticGuardDecision( + @JsonProperty("verdict") SemanticVerdict verdict, + @JsonProperty("reason") String reason) { + + public SemanticGuardDecision { + Objects.requireNonNull(verdict, "verdict must not be null"); + if (reason == null || reason.isBlank()) { + throw new IllegalArgumentException("reason must not be blank"); + } + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardInput.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardInput.java new file mode 100644 index 0000000..7b47cbc --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardInput.java @@ -0,0 +1,26 @@ +package com.superbiz.agent.harness.guard.semantic; + +import com.fasterxml.jackson.annotation.JsonProperty; +import com.superbiz.agent.harness.contract.DiagnosisDraft; +import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot; + +import java.util.Objects; + +public record SemanticGuardInput( + @JsonProperty("query") String query, + @JsonProperty("draft") SemanticDraftView draft, + @JsonProperty("verified_evidence") VerifiedEvidenceSnapshot verifiedEvidence) { + + public SemanticGuardInput { + if (query == null || query.isBlank()) { + throw new IllegalArgumentException("query must not be blank"); + } + Objects.requireNonNull(draft, "draft must not be null"); + Objects.requireNonNull(verifiedEvidence, "verifiedEvidence must not be null"); + } + + public static SemanticGuardInput from(String query, DiagnosisDraft draft, + VerifiedEvidenceSnapshot verifiedEvidence) { + return new SemanticGuardInput(query, SemanticDraftView.from(draft), verifiedEvidence); + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardLimits.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardLimits.java new file mode 100644 index 0000000..dc652a0 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardLimits.java @@ -0,0 +1,30 @@ +package com.superbiz.agent.harness.guard.semantic; + +import java.time.Duration; +import java.util.Objects; + +public record SemanticGuardLimits( + long maxInputBytes, + long maxOutputBytes, + Duration perAttemptTimeout, + Duration totalTimeout) { + + public SemanticGuardLimits { + if (maxInputBytes <= 0 || maxOutputBytes <= 0) { + throw new IllegalArgumentException("byte limits must be positive"); + } + requirePositive(perAttemptTimeout, "perAttemptTimeout"); + requirePositive(totalTimeout, "totalTimeout"); + if (totalTimeout.compareTo(perAttemptTimeout) < 0) { + throw new IllegalArgumentException("totalTimeout must not be shorter than perAttemptTimeout"); + } + } + + private static void requirePositive(Duration value, String name) { + Objects.requireNonNull(value, name + " must not be null"); + if (value.isZero() || value.isNegative()) { + throw new IllegalArgumentException(name + " must be positive"); + } + value.toNanos(); + } +} diff --git a/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardPrompt.java b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardPrompt.java new file mode 100644 index 0000000..a1200d3 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardPrompt.java @@ -0,0 +1,24 @@ +package com.superbiz.agent.harness.guard.semantic; + +import org.springframework.core.io.ClassPathResource; + +import java.io.IOException; +import java.io.InputStream; +import java.nio.charset.StandardCharsets; + +final class SemanticGuardPrompt { + + private static final String RESOURCE_PATH = "prompts/semantic-guard-prompt.md"; + + private SemanticGuardPrompt() { + } + + static String load() { + ClassPathResource resource = new ClassPathResource(RESOURCE_PATH); + try (InputStream input = resource.getInputStream()) { + return new String(input.readAllBytes(), StandardCharsets.UTF_8); + } catch (IOException exception) { + throw new IllegalStateException("Failed to load SemanticGuard prompt", exception); + } + } +} diff --git a/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseResult.java b/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseResult.java new file mode 100644 index 0000000..16a85ce --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseResult.java @@ -0,0 +1,44 @@ +package com.superbiz.agent.harness.release; + +import com.superbiz.agent.harness.contract.DiagnosisDraft; +import com.superbiz.agent.harness.contract.ReleaseOutcome; +import com.superbiz.agent.harness.contract.SafeFallback; +import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot; + +import java.util.Objects; + +public record DiagnosisReleaseResult( + ReleaseOutcome outcome, + DiagnosisDraft draft, + SafeFallback fallback, + VerifiedEvidenceSnapshot verifiedEvidence) { + + public DiagnosisReleaseResult { + Objects.requireNonNull(outcome, "outcome must not be null"); + Objects.requireNonNull(verifiedEvidence, "verifiedEvidence must not be null"); + if (outcome == ReleaseOutcome.SUCCESS) { + Objects.requireNonNull(draft, "successful release requires draft"); + if (fallback != null) { + throw new IllegalArgumentException("successful release must not contain fallback"); + } + } else if (outcome == ReleaseOutcome.FALLBACK) { + Objects.requireNonNull(fallback, "fallback release requires fallback"); + if (draft != null) { + throw new IllegalArgumentException("fallback release must not contain draft"); + } + } else { + throw new IllegalArgumentException("release use case supports SUCCESS or FALLBACK only"); + } + } + + public static DiagnosisReleaseResult success( + DiagnosisDraft draft, VerifiedEvidenceSnapshot verifiedEvidence) { + return new DiagnosisReleaseResult( + ReleaseOutcome.SUCCESS, draft, null, verifiedEvidence); + } + + public static DiagnosisReleaseResult fallback(SafeFallback fallback) { + return new DiagnosisReleaseResult( + ReleaseOutcome.FALLBACK, null, fallback, VerifiedEvidenceSnapshot.empty()); + } +} diff --git a/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCase.java b/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCase.java new file mode 100644 index 0000000..28d7b3d --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCase.java @@ -0,0 +1,91 @@ +package com.superbiz.agent.harness.release; + +import com.superbiz.agent.harness.contract.DiagnosisDraft; +import com.superbiz.agent.harness.contract.SemanticVerdict; +import com.superbiz.agent.harness.core.BudgetExceededException; +import com.superbiz.agent.harness.core.RunAbortedException; +import com.superbiz.agent.harness.core.RunContext; +import com.superbiz.agent.harness.guard.evidence.EvidenceGuard; +import com.superbiz.agent.harness.guard.evidence.EvidenceGuardResult; +import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot; +import com.superbiz.agent.harness.guard.semantic.SemanticGuard; +import com.superbiz.agent.harness.guard.semantic.SemanticGuardDecision; +import com.superbiz.agent.harness.guard.semantic.SemanticGuardInput; +import com.superbiz.agent.harness.retry.RetryExecutionException; +import com.superbiz.agent.harness.retry.RetryFailure; + +import java.util.Objects; + +public final class DiagnosisReleaseUseCase { + + private final EvidenceGuard evidenceGuard; + private final EvidenceRepair evidenceRepair; + private final SemanticGuard semanticGuard; + private final SafeFallbackFactory fallbackFactory; + + public DiagnosisReleaseUseCase(EvidenceGuard evidenceGuard, + EvidenceRepair evidenceRepair, + SemanticGuard semanticGuard, + SafeFallbackFactory fallbackFactory) { + this.evidenceGuard = Objects.requireNonNull(evidenceGuard, "evidenceGuard must not be null"); + this.evidenceRepair = Objects.requireNonNull(evidenceRepair, "evidenceRepair must not be null"); + this.semanticGuard = Objects.requireNonNull(semanticGuard, "semanticGuard must not be null"); + this.fallbackFactory = Objects.requireNonNull( + fallbackFactory, "fallbackFactory must not be null"); + } + + public DiagnosisReleaseResult execute(RunContext context, String query, DiagnosisDraft draft) { + Objects.requireNonNull(context, "context must not be null"); + if (query == null || query.isBlank()) { + throw new IllegalArgumentException("query must not be blank"); + } + Objects.requireNonNull(draft, "draft must not be null"); + + DiagnosisDraft candidate = draft; + EvidenceGuardResult evidence = evidenceGuard.validate(context, candidate); + if (!evidence.valid()) { + try { + candidate = evidenceRepair.repair(context, query, draft, evidence.violations()); + evidence = evidenceGuard.validate(context, candidate); + } catch (RuntimeException exception) { + propagateTerminal(exception); + return evidenceFailure(); + } + if (!evidence.valid()) { + return evidenceFailure(); + } + } + + VerifiedEvidenceSnapshot snapshot = evidence.verifiedSnapshot().orElseThrow(); + SemanticGuardDecision decision; + try { + decision = semanticGuard.review( + context, SemanticGuardInput.from(query, candidate, snapshot)); + } catch (RuntimeException exception) { + propagateTerminal(exception); + return DiagnosisReleaseResult.fallback( + fallbackFactory.semanticUnavailable(snapshot)); + } + return decision.verdict() == SemanticVerdict.SUPPORTED + ? DiagnosisReleaseResult.success(candidate, snapshot) + : DiagnosisReleaseResult.fallback( + fallbackFactory.semanticUnsupported(snapshot)); + } + + private DiagnosisReleaseResult evidenceFailure() { + return DiagnosisReleaseResult.fallback( + fallbackFactory.evidenceValidationFailed()); + } + + private void propagateTerminal(RuntimeException exception) { + if (exception instanceof RunAbortedException + || exception instanceof BudgetExceededException) { + throw exception; + } + if (exception instanceof RetryExecutionException retry + && (retry.failure() == RetryFailure.CANCELLED + || retry.failure() == RetryFailure.BUDGET_EXHAUSTED)) { + throw retry; + } + } +} diff --git a/src/main/java/com/superbiz/agent/harness/release/EvidenceRepair.java b/src/main/java/com/superbiz/agent/harness/release/EvidenceRepair.java new file mode 100644 index 0000000..8df5855 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/release/EvidenceRepair.java @@ -0,0 +1,125 @@ +package com.superbiz.agent.harness.release; + +import com.fasterxml.jackson.core.JsonProcessingException; +import com.fasterxml.jackson.databind.DeserializationFeature; +import com.fasterxml.jackson.databind.ObjectMapper; +import com.fasterxml.jackson.databind.ObjectReader; +import com.superbiz.agent.harness.contract.DiagnosisDraft; +import com.superbiz.agent.harness.core.DiagnosisHarnessCore; +import com.superbiz.agent.harness.core.RunContext; +import com.superbiz.agent.harness.guard.evidence.EvidenceViolation; +import com.superbiz.agent.harness.guard.semantic.GuardModelCall; +import com.superbiz.agent.harness.guard.semantic.GuardModelCallException; +import com.superbiz.agent.harness.guard.semantic.SemanticDraftView; +import com.superbiz.agent.harness.retry.HarnessRetryExecutor; +import com.superbiz.agent.harness.retry.RetryAttempt; +import com.superbiz.agent.harness.retry.RetryFailure; +import org.springframework.ai.chat.messages.SystemMessage; +import org.springframework.ai.chat.messages.UserMessage; +import org.springframework.ai.chat.prompt.Prompt; + +import java.nio.charset.StandardCharsets; +import java.util.List; +import java.util.Objects; +import java.util.function.Consumer; + +public final class EvidenceRepair { + + private final DiagnosisHarnessCore core; + private final HarnessRetryExecutor retryExecutor; + private final GuardModelCall modelCall; + private final ObjectMapper objectMapper; + private final ObjectReader draftReader; + private final EvidenceRepairLimits limits; + private final Consumer attemptRecorder; + private final String prompt; + + public EvidenceRepair(DiagnosisHarnessCore core, + HarnessRetryExecutor retryExecutor, + GuardModelCall modelCall, + ObjectMapper objectMapper, + EvidenceRepairLimits limits, + Consumer attemptRecorder) { + this.core = Objects.requireNonNull(core, "core must not be null"); + this.retryExecutor = Objects.requireNonNull(retryExecutor, "retryExecutor must not be null"); + this.modelCall = Objects.requireNonNull(modelCall, "modelCall must not be null"); + this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null"); + this.draftReader = objectMapper.readerFor(DiagnosisDraft.class) + .with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES) + .with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS); + this.limits = Objects.requireNonNull(limits, "limits must not be null"); + this.attemptRecorder = Objects.requireNonNull( + attemptRecorder, "attemptRecorder must not be null"); + this.prompt = EvidenceRepairPrompt.load(); + } + + public DiagnosisDraft repair(RunContext context, String query, DiagnosisDraft original, + List violations) { + Objects.requireNonNull(context, "context must not be null"); + if (query == null || query.isBlank()) { + throw new IllegalArgumentException("query must not be blank"); + } + Objects.requireNonNull(original, "original must not be null"); + List safeViolations = violations == null + ? List.of() : List.copyOf(violations); + String inputJson = serialize(new RepairInput(query, original, safeViolations)); + long inputBytes = utf8Bytes(inputJson); + if (inputBytes > limits.maxInputBytes()) { + throw new GuardModelCallException( + RetryFailure.SCHEMA_INVALID, "Evidence repair input exceeded limit"); + } + core.reserveRunBytes(context, inputBytes); + Prompt modelPrompt = new Prompt(List.of( + new SystemMessage(prompt), new UserMessage(inputJson))); + SemanticDraftView originalSemantics = SemanticDraftView.from(original); + return retryExecutor.execute( + context, + context.retryPolicies().evidenceRepair(), + () -> { + DiagnosisDraft repaired = parse(modelCall.call( + context, modelPrompt, limits.timeout(), limits.maxOutputBytes())); + if (!originalSemantics.hasSameUserVisibleSemantics( + SemanticDraftView.from(repaired))) { + throw new GuardModelCallException( + RetryFailure.SCHEMA_INVALID, + "Evidence repair changed user-visible semantics"); + } + return repaired; + }, + this::classify, + attemptRecorder); + } + + private DiagnosisDraft parse(String output) { + try { + return draftReader.readValue(output); + } catch (JsonProcessingException exception) { + throw new GuardModelCallException( + RetryFailure.PARSE_ERROR, "Evidence repair output is invalid", exception); + } + } + + private RetryFailure classify(Exception exception) { + return exception instanceof GuardModelCallException failure + ? failure.failure() : RetryFailure.UNKNOWN; + } + + private String serialize(RepairInput input) { + try { + return objectMapper.writeValueAsString(input); + } catch (JsonProcessingException exception) { + throw new GuardModelCallException( + RetryFailure.SCHEMA_INVALID, "Evidence repair input is not serializable", exception); + } + } + + private static long utf8Bytes(String value) { + return value.getBytes(StandardCharsets.UTF_8).length; + } + + private record RepairInput( + String query, + DiagnosisDraft draft, + List violations) { + } +} diff --git a/src/main/java/com/superbiz/agent/harness/release/EvidenceRepairLimits.java b/src/main/java/com/superbiz/agent/harness/release/EvidenceRepairLimits.java new file mode 100644 index 0000000..333a1bb --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/release/EvidenceRepairLimits.java @@ -0,0 +1,21 @@ +package com.superbiz.agent.harness.release; + +import java.time.Duration; +import java.util.Objects; + +public record EvidenceRepairLimits( + long maxInputBytes, + long maxOutputBytes, + Duration timeout) { + + public EvidenceRepairLimits { + if (maxInputBytes <= 0 || maxOutputBytes <= 0) { + throw new IllegalArgumentException("byte limits must be positive"); + } + Objects.requireNonNull(timeout, "timeout must not be null"); + if (timeout.isZero() || timeout.isNegative()) { + throw new IllegalArgumentException("timeout must be positive"); + } + timeout.toNanos(); + } +} diff --git a/src/main/java/com/superbiz/agent/harness/release/EvidenceRepairPrompt.java b/src/main/java/com/superbiz/agent/harness/release/EvidenceRepairPrompt.java new file mode 100644 index 0000000..f53eec3 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/release/EvidenceRepairPrompt.java @@ -0,0 +1,24 @@ +package com.superbiz.agent.harness.release; + +import org.springframework.core.io.ClassPathResource; + +import java.io.IOException; +import java.io.InputStream; +import java.nio.charset.StandardCharsets; + +final class EvidenceRepairPrompt { + + private static final String RESOURCE_PATH = "prompts/evidence-repair-prompt.md"; + + private EvidenceRepairPrompt() { + } + + static String load() { + ClassPathResource resource = new ClassPathResource(RESOURCE_PATH); + try (InputStream input = resource.getInputStream()) { + return new String(input.readAllBytes(), StandardCharsets.UTF_8); + } catch (IOException exception) { + throw new IllegalStateException("Failed to load evidence repair prompt", exception); + } + } +} diff --git a/src/main/java/com/superbiz/agent/harness/release/SafeFallbackFactory.java b/src/main/java/com/superbiz/agent/harness/release/SafeFallbackFactory.java new file mode 100644 index 0000000..ed37cf9 --- /dev/null +++ b/src/main/java/com/superbiz/agent/harness/release/SafeFallbackFactory.java @@ -0,0 +1,51 @@ +package com.superbiz.agent.harness.release; + +import com.superbiz.agent.harness.contract.FallbackType; +import com.superbiz.agent.harness.contract.SafeFallback; +import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot; + +import java.util.List; +import java.util.Objects; + +public final class SafeFallbackFactory { + + public SafeFallback evidenceValidationFailed() { + return fallback( + FallbackType.EVIDENCE_VALIDATION_FAILED, + List.of(), + "当前证据无法完成真实性校验,无法确认根因", + "证据引用校验未通过", + "重新收集当前诊断范围内的证据后再发起诊断"); + } + + public SafeFallback semanticUnsupported(VerifiedEvidenceSnapshot snapshot) { + return fallback( + FallbackType.SEMANTIC_UNSUPPORTED, + sources(snapshot), + "当前证据不足,无法确认根因", + "语义校验未通过", + "补充当前缺失的数据后重新发起诊断"); + } + + public SafeFallback semanticUnavailable(VerifiedEvidenceSnapshot snapshot) { + return fallback( + FallbackType.SEMANTIC_UNAVAILABLE, + sources(snapshot), + "当前证据暂时无法完成语义校验,无法确认根因", + "语义校验暂不可用", + "稍后重试或补充当前缺失的数据"); + } + + private SafeFallback fallback(FallbackType type, + List sources, + String message, + String limitation, + String nextStep) { + return new SafeFallback( + type, null, message, sources, List.of(limitation), List.of(nextStep)); + } + + private List sources(VerifiedEvidenceSnapshot snapshot) { + return Objects.requireNonNull(snapshot, "snapshot must not be null").verifiedSources(); + } +} diff --git a/src/main/resources/prompts/evidence-repair-prompt.md b/src/main/resources/prompts/evidence-repair-prompt.md new file mode 100644 index 0000000..2fe1d22 --- /dev/null +++ b/src/main/resources/prompts/evidence-repair-prompt.md @@ -0,0 +1,7 @@ +You repair only the identifier and reference structure of one DiagnosisDraft. + +Use the supplied violation codes and targets. You may change only analysis_id, tool_call_ids, and based_on_analysis_ids. +Do not add, remove, reorder, summarize, or rewrite any conclusion, analysis, action, recommendation, limitation, kind, or human-confirmation flag. +You have no tools, memory, evidence store, or permission to call another agent. + +Return exactly the complete repaired DiagnosisDraft JSON with no markdown or explanation. diff --git a/src/main/resources/prompts/semantic-guard-prompt.md b/src/main/resources/prompts/semantic-guard-prompt.md new file mode 100644 index 0000000..e6aae2f --- /dev/null +++ b/src/main/resources/prompts/semantic-guard-prompt.md @@ -0,0 +1,10 @@ +You are the isolated SemanticGuard for one diagnosis report. + +Review only the supplied original query, complete user-visible draft, and verified evidence snapshot. +Check whether evidence supports each analysis, analyses support the conclusion, actions and recommendations stay within supported findings, and limitations accurately state the observed scope. + +You have no tools, memory, Redis access, diagnosis history, or permission to call another agent. +Do not rewrite, correct, summarize, extend, or partially approve the report. + +Return exactly one JSON object with no markdown and no extra fields: +{"verdict":"SUPPORTED|UNSUPPORTED","reason":"non-empty audit reason"} diff --git a/src/test/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuardTest.java b/src/test/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuardTest.java new file mode 100644 index 0000000..d7ecc3c --- /dev/null +++ b/src/test/java/com/superbiz/agent/harness/guard/evidence/EvidenceGuardTest.java @@ -0,0 +1,296 @@ +package com.superbiz.agent.harness.guard.evidence; + +import com.fasterxml.jackson.databind.ObjectMapper; +import com.superbiz.agent.harness.contract.AnalysisKind; +import com.superbiz.agent.harness.contract.DiagnosisDraft; +import com.superbiz.agent.harness.contract.EvidenceStatus; +import com.superbiz.agent.harness.core.DiagnosisHarnessCore; +import com.superbiz.agent.harness.core.RunBudgetLimits; +import com.superbiz.agent.harness.core.RunContext; +import com.superbiz.agent.harness.retry.HarnessRetryPolicies; +import com.superbiz.agent.harness.tool.store.CanonicalInvocationLimits; +import com.superbiz.agent.harness.tool.store.CanonicalInvocationStore; +import com.superbiz.agent.harness.tool.store.CanonicalToolInvocation; +import com.superbiz.agent.harness.tool.store.ToolCallKeyFactory; +import org.junit.jupiter.api.Test; + +import java.time.Clock; +import java.time.Duration; +import java.time.Instant; +import java.util.HashMap; +import java.util.List; +import java.util.Map; +import java.util.Optional; + +import static org.junit.jupiter.api.Assertions.assertEquals; +import static org.junit.jupiter.api.Assertions.assertFalse; +import static org.junit.jupiter.api.Assertions.assertTrue; + +class EvidenceGuardTest { + + private static final String PREFIX = "superbiz:harness:tool-call"; + private final ObjectMapper objectMapper = new ObjectMapper(); + private final InMemoryStore store = new InMemoryStore(); + private final EvidenceGuard guard = new EvidenceGuard( + store, new ToolCallKeyFactory(PREFIX), objectMapper); + private final RunContext context = context("run-guard"); + + @Test + void validCurrentRunRagReferenceProducesInternalIdFreeSnapshot() throws Exception { + String callId = "call-rag-1"; + ready(callId, "lookup_knowledge", "{\"query\":\"pool timeout\"}", """ + {"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-rag-1", + "query":"pool timeout","evidence":[{"document_id":"doc-1", + "source":"runbook.md","title":"Pool guide","breadcrumb":"DB > Pool", + "excerpt":"active=50 max=50"}],"returned_count":1,"truncated":false} + """, EvidenceStatus.EVIDENCE_FOUND); + + EvidenceGuardResult result = guard.validate(context, draft(callId, AnalysisKind.NORMAL)); + + assertTrue(result.valid()); + VerifiedEvidenceSnapshot snapshot = result.verifiedSnapshot().orElseThrow(); + assertEquals("a-1", snapshot.analyses().get(0).analysisId()); + assertEquals("runbook.md", snapshot.analyses().get(0).evidence().get(0).source()); + assertEquals("active=50 max=50", snapshot.analyses().get(0).evidence().get(0).excerpt()); + String json = objectMapper.writeValueAsString(snapshot); + assertFalse(json.contains(callId)); + assertFalse(json.contains("raw_response")); + assertEquals(1, snapshot.verifiedSources().size()); + } + + @Test + void noEvidenceSupportsOnlyScopedNegativeObservation() { + String callId = "call-rag-empty"; + ready(callId, "lookup_knowledge", "{\"query\":\"unknown failure\"}", """ + {"evidence_status":"NO_EVIDENCE","tool_call_id":"call-rag-empty", + "query":"unknown failure","evidence":[],"returned_count":0,"truncated":false} + """, EvidenceStatus.NO_EVIDENCE); + + EvidenceGuardResult result = guard.validate( + context, draft(callId, AnalysisKind.NEGATIVE_OBSERVATION)); + + assertTrue(result.valid()); + VerifiedEvidence evidence = result.verifiedSnapshot().orElseThrow() + .analyses().get(0).evidence().get(0); + assertEquals("unknown failure", evidence.scope()); + assertEquals(0, evidence.values().get("match_count")); + } + + @Test + void validLogProjectionPreservesSourceScopeTimelineAndExactMessage() { + String callId = "call-log-1"; + ready(callId, "query_logs", """ + {"topic":"APPLICATION","query":"pool timeout","lookback_minutes":30} + """, """ + {"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-log-1", + "source_kind":"MOCK","scope":{"topic":"APPLICATION","query":"pool timeout", + "start_time":"2026-07-21T10:00:00Z","end_time":"2026-07-21T10:30:00Z"}, + "match_count":1,"returned_count":1,"patterns":[], + "events":[{"timestamp":"2026-07-21T10:29:00Z","level":"ERROR", + "service":"order-service","message":"active=50 max=50"}],"truncated":false} + """, EvidenceStatus.EVIDENCE_FOUND); + + EvidenceGuardResult result = guard.validate(context, draft(callId, AnalysisKind.NORMAL)); + + assertTrue(result.valid()); + VerifiedEvidence evidence = result.verifiedSnapshot().orElseThrow() + .analyses().get(0).evidence().get(0); + assertEquals("LOG", evidence.sourceType()); + assertEquals("APPLICATION (MOCK)", evidence.source()); + assertEquals("2026-07-21T10:29:00Z", evidence.timestamp()); + assertEquals("active=50 max=50", evidence.excerpt()); + assertTrue(evidence.scope().contains("pool timeout")); + } + + @Test + void validMysqlProjectionCombinesBoundedRequestScopeAndProjectedRows() { + String callId = "call-mysql-1"; + ready(callId, "query_mysql", """ + {"data_source":"order_readonly","sql":"SELECT status FROM biz_order WHERE id = ?", + "params":["order-1"]} + """, """ + {"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-mysql-1", + "columns":["status"],"rows":[{"status":"FAILED"}], + "returned_count":1,"truncated":false} + """, EvidenceStatus.EVIDENCE_FOUND); + + EvidenceGuardResult result = guard.validate(context, draft(callId, AnalysisKind.NORMAL)); + + assertTrue(result.valid()); + VerifiedEvidence evidence = result.verifiedSnapshot().orElseThrow() + .analyses().get(0).evidence().get(0); + assertEquals("MYSQL", evidence.sourceType()); + assertEquals("order_readonly", evidence.source()); + assertTrue(evidence.scope().contains("SELECT status")); + assertTrue(evidence.scope().contains("order-1")); + assertEquals("FAILED", evidence.values().get("status")); + } + + @Test + void duplicateAnalysisAndBrokenReportReferenceFailBeforeStoreLookup() { + DiagnosisDraft invalid = new DiagnosisDraft( + new DiagnosisDraft.Conclusion("Pool exhausted", List.of("missing")), + List.of( + new DiagnosisDraft.AnalysisItem( + "a-1", AnalysisKind.NORMAL, "first", List.of("call-1")), + new DiagnosisDraft.AnalysisItem( + "a-1", AnalysisKind.NORMAL, "second", List.of("call-2"))), + List.of(new DiagnosisDraft.ActionPlanItem("inspect", List.of(), false)), + List.of(), new DiagnosisDraft.Limitations("order-service", List.of())); + + EvidenceGuardResult result = guard.validate(context, invalid); + + assertFalse(result.valid()); + assertTrue(hasViolation(result, EvidenceViolationCode.ANALYSIS_ID_DUPLICATE)); + assertTrue(hasViolation(result, EvidenceViolationCode.ANALYSIS_REFERENCE_UNKNOWN)); + assertTrue(hasViolation(result, EvidenceViolationCode.ANALYSIS_REFERENCE_MISSING)); + assertTrue(result.verifiedSnapshot().isEmpty()); + } + + @Test + void fabricatedAndCrossRunReferencesAreRejected() { + EvidenceGuardResult missing = guard.validate( + context, draft("fabricated-call", AnalysisKind.NORMAL)); + String callId = "call-cross-run"; + String currentKey = PREFIX + ":" + context.runId() + ":" + callId; + store.records.put(currentKey, CanonicalToolInvocation.projecting( + callId, "another-run", "lookup_knowledge", "{\"query\":\"pool\"}", + Instant.parse("2026-07-21T10:00:00Z")) + .markReady("raw", """ + {"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-cross-run", + "query":"pool","evidence":[{"document_id":"doc-1","source":"guide", + "title":"guide","breadcrumb":"pool","excerpt":"active=50"}], + "returned_count":1,"truncated":false} + """, EvidenceStatus.EVIDENCE_FOUND, + Instant.parse("2026-07-21T10:00:01Z"))); + + EvidenceGuardResult crossRun = guard.validate( + context, draft(callId, AnalysisKind.NORMAL)); + + assertTrue(hasViolation(missing, EvidenceViolationCode.INVOCATION_MISSING)); + assertTrue(hasViolation(crossRun, EvidenceViolationCode.INVOCATION_NOT_REFERENCABLE)); + } + + @Test + void projectingInvocationAndEvidenceKindMismatchAreRejected() { + String projectingId = "call-projecting"; + store.records.put(PREFIX + ":" + context.runId() + ":" + projectingId, + CanonicalToolInvocation.projecting( + projectingId, context.runId(), "lookup_knowledge", "{\"query\":\"pool\"}", + Instant.parse("2026-07-21T10:00:00Z"))); + EvidenceGuardResult projecting = guard.validate( + context, draft(projectingId, AnalysisKind.NORMAL)); + + String emptyId = "call-empty-normal"; + ready(emptyId, "lookup_knowledge", "{\"query\":\"pool\"}", """ + {"evidence_status":"NO_EVIDENCE","tool_call_id":"call-empty-normal", + "query":"pool","evidence":[],"returned_count":0,"truncated":false} + """, EvidenceStatus.NO_EVIDENCE); + EvidenceGuardResult mismatch = guard.validate( + context, draft(emptyId, AnalysisKind.NORMAL)); + + assertTrue(hasViolation(projecting, EvidenceViolationCode.INVOCATION_NOT_REFERENCABLE)); + assertTrue(hasViolation(mismatch, EvidenceViolationCode.EVIDENCE_KIND_MISMATCH)); + } + + @Test + void projectionIdOrStatusMismatchFailsClosed() { + String callId = "call-mismatch"; + ready(callId, "lookup_knowledge", "{\"query\":\"pool\"}", """ + {"evidence_status":"NO_EVIDENCE","tool_call_id":"another-call", + "query":"pool","evidence":[],"returned_count":0,"truncated":false} + """, EvidenceStatus.EVIDENCE_FOUND); + + EvidenceGuardResult result = guard.validate(context, draft(callId, AnalysisKind.NORMAL)); + + assertTrue(hasViolation(result, EvidenceViolationCode.PROJECTION_ID_MISMATCH)); + } + + @Test + void canonicalRecordIdMustMatchTheDraftReferenceEvenUnderCorruptedKey() { + String referencedId = "call-requested"; + CanonicalToolInvocation wrongRecord = CanonicalToolInvocation.projecting( + "call-other", context.runId(), "lookup_knowledge", "{\"query\":\"pool\"}", + Instant.parse("2026-07-21T10:00:00Z")) + .markReady("raw", """ + {"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-other", + "query":"pool","evidence":[{"document_id":"doc-1","source":"guide", + "title":"guide","breadcrumb":"pool","excerpt":"active=50"}], + "returned_count":1,"truncated":false} + """, EvidenceStatus.EVIDENCE_FOUND, + Instant.parse("2026-07-21T10:00:01Z")); + store.records.put(PREFIX + ":" + context.runId() + ":" + referencedId, wrongRecord); + + EvidenceGuardResult result = guard.validate( + context, draft(referencedId, AnalysisKind.NORMAL)); + + assertTrue(hasViolation(result, EvidenceViolationCode.INVOCATION_ID_MISMATCH)); + } + + private boolean hasViolation(EvidenceGuardResult result, EvidenceViolationCode code) { + return result.violations().stream().anyMatch(violation -> violation.code() == code); + } + + private DiagnosisDraft draft(String callId, AnalysisKind kind) { + return new DiagnosisDraft( + new DiagnosisDraft.Conclusion("Pool exhausted", List.of("a-1")), + List.of(new DiagnosisDraft.AnalysisItem( + "a-1", kind, "Pool reached its limit", List.of(callId))), + List.of(new DiagnosisDraft.ActionPlanItem( + "Inspect long transactions", List.of("a-1"), false)), + List.of(new DiagnosisDraft.Recommendation( + "Add saturation alert", List.of("a-1"))), + new DiagnosisDraft.Limitations("order-service", List.of())); + } + + private void ready(String callId, String toolName, String request, + String agentResult, EvidenceStatus evidenceStatus) { + String key = PREFIX + ":" + context.runId() + ":" + callId; + CanonicalToolInvocation invocation = CanonicalToolInvocation.projecting( + callId, context.runId(), toolName, request, Instant.parse("2026-07-21T10:00:00Z")) + .markReady("raw-must-not-be-read", agentResult, evidenceStatus, + Instant.parse("2026-07-21T10:00:01Z")); + store.records.put(key, invocation); + } + + private RunContext context(String runId) { + return new DiagnosisHarnessCore( + Clock.systemUTC(), () -> runId, Duration.ofMinutes(5), + new RunBudgetLimits(10, 10, 10, 10_000, 10_000, 20_000, 1_000_000), + HarnessRetryPolicies.strict()).startRun("session-guard", runId); + } + + private static final class InMemoryStore implements CanonicalInvocationStore { + private final Map records = new HashMap<>(); + private final CanonicalInvocationLimits limits = + new CanonicalInvocationLimits(Duration.ofHours(2), 1_000_000, 64_000); + + @Override + public CanonicalInvocationLimits limits() { + return limits; + } + + @Override + public void begin(String key, CanonicalToolInvocation invocation) { + records.put(key, invocation); + } + + @Override + public Optional find(String key) { + return Optional.ofNullable(records.get(key)); + } + + @Override + public CanonicalToolInvocation markReady(String key, String rawResponse, + String agentResult, EvidenceStatus evidenceStatus, + Instant completedAt) { + throw new UnsupportedOperationException(); + } + + @Override + public CanonicalToolInvocation markError(String key, String rawResponse, + String errorCode, Instant completedAt) { + throw new UnsupportedOperationException(); + } + } +} diff --git a/src/test/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardTest.java b/src/test/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardTest.java new file mode 100644 index 0000000..6053201 --- /dev/null +++ b/src/test/java/com/superbiz/agent/harness/guard/semantic/SemanticGuardTest.java @@ -0,0 +1,231 @@ +package com.superbiz.agent.harness.guard.semantic; + +import com.fasterxml.jackson.databind.ObjectMapper; +import com.superbiz.agent.harness.contract.AnalysisKind; +import com.superbiz.agent.harness.contract.DiagnosisDraft; +import com.superbiz.agent.harness.contract.SemanticVerdict; +import com.superbiz.agent.harness.core.DiagnosisHarnessCore; +import com.superbiz.agent.harness.core.RunBudgetLimits; +import com.superbiz.agent.harness.core.RunContext; +import com.superbiz.agent.harness.guard.evidence.VerifiedAnalysisEvidence; +import com.superbiz.agent.harness.guard.evidence.VerifiedEvidence; +import com.superbiz.agent.harness.guard.evidence.VerifiedEvidenceSnapshot; +import com.superbiz.agent.harness.retry.HarnessRetryExecutor; +import com.superbiz.agent.harness.retry.HarnessRetryPolicies; +import com.superbiz.agent.harness.retry.RetryAttempt; +import org.junit.jupiter.api.AfterEach; +import org.junit.jupiter.api.Test; +import org.springframework.ai.chat.messages.AssistantMessage; +import org.springframework.ai.chat.metadata.ChatResponseMetadata; +import org.springframework.ai.chat.metadata.DefaultUsage; +import org.springframework.ai.chat.model.ChatModel; +import org.springframework.ai.chat.model.ChatResponse; +import org.springframework.ai.chat.model.Generation; +import org.springframework.ai.chat.prompt.Prompt; + +import java.time.Clock; +import java.time.Duration; +import java.util.ArrayList; +import java.util.List; +import java.util.Map; +import java.util.concurrent.ExecutorService; +import java.util.concurrent.Executors; +import java.util.concurrent.CompletableFuture; +import java.util.concurrent.CompletionException; +import java.util.concurrent.CountDownLatch; +import java.util.concurrent.TimeUnit; +import java.util.concurrent.atomic.AtomicInteger; + +import static org.junit.jupiter.api.Assertions.assertEquals; +import static org.junit.jupiter.api.Assertions.assertFalse; +import static org.junit.jupiter.api.Assertions.assertTrue; +import static org.junit.jupiter.api.Assertions.assertThrows; + +class SemanticGuardTest { + + private final ObjectMapper objectMapper = new ObjectMapper(); + private final ExecutorService executor = Executors.newCachedThreadPool(); + + @AfterEach + void shutdownExecutor() { + executor.shutdownNow(); + } + + @Test + void unsupportedIsAValidBusinessDecisionAndIsNotRetried() { + DiagnosisHarnessCore core = core(); + RunContext context = core.startRun("session-semantic", "run-semantic"); + ScriptedChatModel model = new ScriptedChatModel( + "{\"verdict\":\"UNSUPPORTED\",\"reason\":\"evidence does not prove the root cause\"}"); + List attempts = new ArrayList<>(); + SemanticGuard guard = new SemanticGuard( + core, + new HarnessRetryExecutor(core), + new GuardModelCall(core, model, executor), + objectMapper, + new SemanticGuardLimits(100_000, 10_000, + Duration.ofSeconds(2), Duration.ofSeconds(3)), + attempts::add); + + SemanticGuardDecision decision = guard.review( + context, SemanticGuardInput.from("Why did payment fail?", draft(), snapshot())); + + assertEquals(SemanticVerdict.UNSUPPORTED, decision.verdict()); + assertEquals(1, model.calls.get()); + assertEquals(1, attempts.size()); + assertTrue(attempts.get(0).success()); + String prompt = model.prompts.get(0); + assertTrue(prompt.contains("Why did payment fail?")); + assertTrue(prompt.contains("active=50 max=50")); + assertFalse(prompt.contains("call-rag-1")); + assertFalse(prompt.contains("tool_call_id")); + assertFalse(prompt.contains("raw_response")); + } + + @Test + void parseFailureRetriesOnceWithTheExactSameInput() { + DiagnosisHarnessCore core = core(); + RunContext context = core.startRun("session-retry", "run-retry"); + ScriptedChatModel model = new ScriptedChatModel( + "not-json", "{\"verdict\":\"SUPPORTED\",\"reason\":\"all claims are grounded\"}"); + List attempts = new ArrayList<>(); + SemanticGuard guard = new SemanticGuard( + core, new HarnessRetryExecutor(core), + new GuardModelCall(core, model, executor), objectMapper, + new SemanticGuardLimits(100_000, 10_000, + Duration.ofSeconds(2), Duration.ofSeconds(3)), attempts::add); + + SemanticGuardDecision decision = guard.review( + context, SemanticGuardInput.from("Why did payment fail?", draft(), snapshot())); + + assertEquals(SemanticVerdict.SUPPORTED, decision.verdict()); + assertEquals(2, model.calls.get()); + assertEquals(model.prompts.get(0), model.prompts.get(1)); + assertEquals(com.superbiz.agent.harness.retry.RetryFailure.PARSE_ERROR, + attempts.get(0).failure()); + assertTrue(attempts.get(1).success()); + assertEquals(2, context.budget().snapshot().modelCalls()); + assertEquals(16, context.budget().snapshot().totalTokens()); + } + + @Test + void attemptTimeoutCancelsBothPermittedModelCalls() throws Exception { + DiagnosisHarnessCore core = core(); + RunContext context = core.startRun("session-timeout", "run-timeout"); + BlockingChatModel model = new BlockingChatModel(2); + List attempts = new ArrayList<>(); + SemanticGuard guard = new SemanticGuard( + core, new HarnessRetryExecutor(core), + new GuardModelCall(core, model, executor), objectMapper, + new SemanticGuardLimits(100_000, 10_000, + Duration.ofMillis(50), Duration.ofMillis(500)), attempts::add); + + com.superbiz.agent.harness.retry.RetryExecutionException failure = assertThrows( + com.superbiz.agent.harness.retry.RetryExecutionException.class, + () -> guard.review(context, + SemanticGuardInput.from("Why?", draft(), snapshot()))); + + assertEquals(com.superbiz.agent.harness.retry.RetryFailure.TIMEOUT, failure.failure()); + assertEquals(2, failure.attempts()); + assertTrue(model.interrupted.await(2, TimeUnit.SECONDS)); + assertEquals(2, model.calls.get()); + assertEquals(2, attempts.size()); + } + + @Test + void runCancellationInterruptsPendingModelAndDoesNotRetry() throws Exception { + DiagnosisHarnessCore core = core(); + RunContext context = core.startRun("session-cancel", "run-cancel"); + BlockingChatModel model = new BlockingChatModel(1); + SemanticGuard guard = new SemanticGuard( + core, new HarnessRetryExecutor(core), + new GuardModelCall(core, model, executor), objectMapper, + new SemanticGuardLimits(100_000, 10_000, + Duration.ofSeconds(2), Duration.ofSeconds(3)), ignored -> { }); + + CompletableFuture execution = CompletableFuture.supplyAsync( + () -> guard.review(context, + SemanticGuardInput.from("Why?", draft(), snapshot()))); + assertTrue(model.started.await(2, TimeUnit.SECONDS)); + core.cancel(context, + com.superbiz.agent.harness.core.RunCancellationReason.USER_REQUESTED); + + CompletionException thrown = assertThrows(CompletionException.class, execution::join); + com.superbiz.agent.harness.retry.RetryExecutionException failure = + (com.superbiz.agent.harness.retry.RetryExecutionException) thrown.getCause(); + assertEquals(com.superbiz.agent.harness.retry.RetryFailure.CANCELLED, failure.failure()); + assertTrue(model.interrupted.await(2, TimeUnit.SECONDS)); + assertEquals(1, model.calls.get()); + } + + private DiagnosisDraft draft() { + return new DiagnosisDraft( + new DiagnosisDraft.Conclusion("Pool exhausted", List.of("a-1")), + List.of(new DiagnosisDraft.AnalysisItem( + "a-1", AnalysisKind.NORMAL, "Pool reached its limit", List.of("call-rag-1"))), + List.of(new DiagnosisDraft.ActionPlanItem( + "Inspect long transactions", List.of("a-1"), false)), + List.of(new DiagnosisDraft.Recommendation( + "Add saturation alert", List.of("a-1"))), + new DiagnosisDraft.Limitations("order-service, last 30 minutes", List.of("No slow SQL"))); + } + + private VerifiedEvidenceSnapshot snapshot() { + return new VerifiedEvidenceSnapshot(List.of(new VerifiedAnalysisEvidence( + "a-1", "Pool reached its limit", AnalysisKind.NORMAL, + List.of(new VerifiedEvidence( + "LOG", "APPLICATION (MOCK)", "order-service, last 30 minutes", + "2026-07-21T10:29:00Z", "active=50 max=50", Map.of("count", 1)))))); + } + + private DiagnosisHarnessCore core() { + return new DiagnosisHarnessCore( + Clock.systemUTC(), () -> "unused", Duration.ofMinutes(5), + new RunBudgetLimits(10, 10, 10, 100_000, 100_000, 200_000, 1_000_000), + HarnessRetryPolicies.strict()); + } + + private static final class ScriptedChatModel implements ChatModel { + private final List responses; + private final List prompts = new ArrayList<>(); + private final AtomicInteger calls = new AtomicInteger(); + + private ScriptedChatModel(String... responses) { + this.responses = List.of(responses); + } + + @Override + public ChatResponse call(Prompt prompt) { + prompts.add(prompt.getContents()); + int index = calls.getAndIncrement(); + ChatResponseMetadata metadata = ChatResponseMetadata.builder() + .usage(new DefaultUsage(5, 3)).build(); + return new ChatResponse( + List.of(new Generation(new AssistantMessage(responses.get(index)))), metadata); + } + } + + private static final class BlockingChatModel implements ChatModel { + private final AtomicInteger calls = new AtomicInteger(); + private final CountDownLatch started = new CountDownLatch(1); + private final CountDownLatch interrupted; + + private BlockingChatModel(int expectedInterruptions) { + this.interrupted = new CountDownLatch(expectedInterruptions); + } + + @Override + public ChatResponse call(Prompt prompt) { + calls.incrementAndGet(); + started.countDown(); + try { + Thread.sleep(10_000); + throw new AssertionError("blocking model should be interrupted"); + } catch (InterruptedException exception) { + interrupted.countDown(); + Thread.currentThread().interrupt(); + throw new IllegalStateException("interrupted", exception); + } + } + } +} diff --git a/src/test/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCaseTest.java b/src/test/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCaseTest.java new file mode 100644 index 0000000..8d537d1 --- /dev/null +++ b/src/test/java/com/superbiz/agent/harness/release/DiagnosisReleaseUseCaseTest.java @@ -0,0 +1,314 @@ +package com.superbiz.agent.harness.release; + +import com.fasterxml.jackson.databind.ObjectMapper; +import com.superbiz.agent.harness.contract.AnalysisKind; +import com.superbiz.agent.harness.contract.DiagnosisDraft; +import com.superbiz.agent.harness.contract.EvidenceStatus; +import com.superbiz.agent.harness.contract.FallbackType; +import com.superbiz.agent.harness.contract.ReleaseOutcome; +import com.superbiz.agent.harness.core.DiagnosisHarnessCore; +import com.superbiz.agent.harness.core.RunBudgetLimits; +import com.superbiz.agent.harness.core.RunContext; +import com.superbiz.agent.harness.guard.evidence.EvidenceGuard; +import com.superbiz.agent.harness.guard.semantic.GuardModelCall; +import com.superbiz.agent.harness.guard.semantic.SemanticGuard; +import com.superbiz.agent.harness.guard.semantic.SemanticGuardLimits; +import com.superbiz.agent.harness.retry.HarnessRetryExecutor; +import com.superbiz.agent.harness.retry.HarnessRetryPolicies; +import com.superbiz.agent.harness.tool.store.CanonicalInvocationLimits; +import com.superbiz.agent.harness.tool.store.CanonicalInvocationStore; +import com.superbiz.agent.harness.tool.store.CanonicalToolInvocation; +import com.superbiz.agent.harness.tool.store.ToolCallKeyFactory; +import org.junit.jupiter.api.AfterEach; +import org.junit.jupiter.api.Test; +import org.springframework.ai.chat.messages.AssistantMessage; +import org.springframework.ai.chat.metadata.ChatResponseMetadata; +import org.springframework.ai.chat.metadata.DefaultUsage; +import org.springframework.ai.chat.model.ChatModel; +import org.springframework.ai.chat.model.ChatResponse; +import org.springframework.ai.chat.model.Generation; +import org.springframework.ai.chat.prompt.Prompt; + +import java.time.Clock; +import java.time.Duration; +import java.time.Instant; +import java.util.ArrayList; +import java.util.HashMap; +import java.util.List; +import java.util.Map; +import java.util.Optional; +import java.util.concurrent.ExecutorService; +import java.util.concurrent.Executors; +import java.util.concurrent.atomic.AtomicInteger; + +import static org.junit.jupiter.api.Assertions.assertEquals; +import static org.junit.jupiter.api.Assertions.assertFalse; +import static org.junit.jupiter.api.Assertions.assertNotNull; +import static org.junit.jupiter.api.Assertions.assertNull; +import static org.junit.jupiter.api.Assertions.assertSame; + +class DiagnosisReleaseUseCaseTest { + + private static final String PREFIX = "superbiz:harness:tool-call"; + private final ObjectMapper objectMapper = new ObjectMapper(); + private final ExecutorService executor = Executors.newCachedThreadPool(); + + @AfterEach + void shutdownExecutor() { + executor.shutdownNow(); + } + + @Test + void supportedVerifiedDraftIsReleasedUnchangedWithoutRepair() { + DiagnosisHarnessCore core = core(); + RunContext context = core.startRun("session-release", "run-release"); + InMemoryStore store = new InMemoryStore(); + ready(store, context, "call-rag-1"); + ScriptedChatModel model = new ScriptedChatModel( + "{\"verdict\":\"SUPPORTED\",\"reason\":\"all claims are grounded\"}"); + HarnessRetryExecutor retries = new HarnessRetryExecutor(core); + GuardModelCall modelCall = new GuardModelCall(core, model, executor); + DiagnosisReleaseUseCase useCase = new DiagnosisReleaseUseCase( + new EvidenceGuard(store, new ToolCallKeyFactory(PREFIX), objectMapper), + new EvidenceRepair(core, retries, modelCall, objectMapper, + new EvidenceRepairLimits(100_000, 20_000, Duration.ofSeconds(2)), ignored -> { }), + new SemanticGuard(core, retries, modelCall, objectMapper, + new SemanticGuardLimits(100_000, 20_000, + Duration.ofSeconds(2), Duration.ofSeconds(3)), ignored -> { }), + new SafeFallbackFactory()); + DiagnosisDraft draft = draft("a-1", "call-rag-1", "Pool exhausted"); + + DiagnosisReleaseResult result = useCase.execute(context, "Why did payment fail?", draft); + + assertEquals(ReleaseOutcome.SUCCESS, result.outcome()); + assertSame(draft, result.draft()); + assertNull(result.fallback()); + assertEquals(1, result.verifiedEvidence().analyses().size()); + assertEquals(1, model.calls.get()); + } + + @Test + void oneStructuralRepairCanFixIdsWithoutChangingReportSemantics() { + Fixture fixture = fixture( + repairedDuplicateDraft("Pool exhausted", "a-2"), + "{\"verdict\":\"SUPPORTED\",\"reason\":\"grounded\"}"); + DiagnosisDraft original = duplicateDraft(); + + DiagnosisReleaseResult result = fixture.useCase.execute( + fixture.context, "Why did payment fail?", original); + + assertEquals(ReleaseOutcome.SUCCESS, result.outcome()); + assertEquals(List.of("a-1", "a-2"), result.draft().analysis().stream() + .map(DiagnosisDraft.AnalysisItem::analysisId).toList()); + assertEquals("Pool exhausted", result.draft().conclusion().text()); + assertEquals(2, fixture.model.calls.get()); + assertFalse(fixture.model.prompts.get(1).contains("call-rag-1")); + } + + @Test + void repairThatChangesVisibleSemanticsReturnsEvidenceFallback() throws Exception { + Fixture fixture = fixture(repairedDuplicateDraft("Different conclusion", "a-2")); + + DiagnosisReleaseResult result = fixture.useCase.execute( + fixture.context, "Why did payment fail?", duplicateDraft()); + + assertFallback(result, FallbackType.EVIDENCE_VALIDATION_FAILED, 0); + String json = objectMapper.writeValueAsString(result); + assertFalse(json.contains("Pool exhausted")); + assertFalse(json.contains("Different conclusion")); + assertEquals(1, fixture.model.calls.get()); + } + + @Test + void secondEvidenceValidationFailureHasNoVerifiedSources() { + Fixture fixture = fixture(repairedDuplicateDraft("Pool exhausted", "a-1")); + + DiagnosisReleaseResult result = fixture.useCase.execute( + fixture.context, "Why did payment fail?", duplicateDraft()); + + assertFallback(result, FallbackType.EVIDENCE_VALIDATION_FAILED, 0); + assertEquals(1, fixture.model.calls.get()); + } + + @Test + void unsupportedNeverReleasesDraftOrAuditReason() throws Exception { + Fixture fixture = fixture( + "{\"verdict\":\"UNSUPPORTED\",\"reason\":\"secret-audit-reason\"}"); + DiagnosisDraft draft = draft("a-1", "call-rag-1", "Pool exhausted"); + + DiagnosisReleaseResult result = fixture.useCase.execute( + fixture.context, "Why did payment fail?", draft); + + assertFallback(result, FallbackType.SEMANTIC_UNSUPPORTED, 1); + String json = objectMapper.writeValueAsString(result); + assertFalse(json.contains("Pool exhausted")); + assertFalse(json.contains("secret-audit-reason")); + assertEquals(1, fixture.model.calls.get()); + } + + @Test + void twoInvalidSemanticOutputsReturnUnavailableWithoutDraft() throws Exception { + Fixture fixture = fixture("not-json", "still-not-json"); + + DiagnosisReleaseResult result = fixture.useCase.execute( + fixture.context, "Why did payment fail?", + draft("a-1", "call-rag-1", "Pool exhausted")); + + assertFallback(result, FallbackType.SEMANTIC_UNAVAILABLE, 1); + assertFalse(objectMapper.writeValueAsString(result).contains("Pool exhausted")); + assertEquals(2, fixture.model.calls.get()); + } + + private Fixture fixture(String... responses) { + DiagnosisHarnessCore core = core(); + RunContext context = core.startRun("session-fixture", "run-fixture"); + InMemoryStore store = new InMemoryStore(); + ready(store, context, "call-rag-1"); + ScriptedChatModel model = new ScriptedChatModel(responses); + HarnessRetryExecutor retries = new HarnessRetryExecutor(core); + GuardModelCall modelCall = new GuardModelCall(core, model, executor); + DiagnosisReleaseUseCase useCase = new DiagnosisReleaseUseCase( + new EvidenceGuard(store, new ToolCallKeyFactory(PREFIX), objectMapper), + new EvidenceRepair(core, retries, modelCall, objectMapper, + new EvidenceRepairLimits(100_000, 20_000, Duration.ofSeconds(2)), ignored -> { }), + new SemanticGuard(core, retries, modelCall, objectMapper, + new SemanticGuardLimits(100_000, 20_000, + Duration.ofSeconds(2), Duration.ofSeconds(3)), ignored -> { }), + new SafeFallbackFactory()); + return new Fixture(context, model, useCase); + } + + private DiagnosisDraft duplicateDraft() { + return new DiagnosisDraft( + new DiagnosisDraft.Conclusion("Pool exhausted", List.of("dup")), + List.of( + new DiagnosisDraft.AnalysisItem( + "dup", AnalysisKind.NORMAL, "Pool reached its limit", List.of("call-rag-1")), + new DiagnosisDraft.AnalysisItem( + "dup", AnalysisKind.NORMAL, "Requests are waiting", List.of("call-rag-1"))), + List.of(new DiagnosisDraft.ActionPlanItem( + "Inspect long transactions", List.of("dup"), false)), + List.of(new DiagnosisDraft.Recommendation( + "Add saturation alert", List.of("dup"))), + new DiagnosisDraft.Limitations("order-service", List.of())); + } + + private String repairedDuplicateDraft(String conclusion, String secondAnalysisId) { + return """ + { + "conclusion":{"text":"%s","based_on_analysis_ids":["a-1"]}, + "analysis":[ + {"analysis_id":"a-1","kind":"NORMAL","text":"Pool reached its limit","tool_call_ids":["call-rag-1"]}, + {"analysis_id":"%s","kind":"NORMAL","text":"Requests are waiting","tool_call_ids":["call-rag-1"]} + ], + "action_plan":[{"action":"Inspect long transactions","based_on_analysis_ids":["a-1"],"requires_human_confirmation":false}], + "recommendations":[{"text":"Add saturation alert","based_on_analysis_ids":["a-1"]}], + "limitations":{"scope":"order-service","missing_info":[]} + } + """.formatted(conclusion, secondAnalysisId); + } + + private void assertFallback(DiagnosisReleaseResult result, + FallbackType type, int expectedSources) { + assertEquals(ReleaseOutcome.FALLBACK, result.outcome()); + assertNull(result.draft()); + assertNotNull(result.fallback()); + assertEquals(type, result.fallback().type()); + assertNull(result.fallback().conclusion()); + assertEquals(expectedSources, result.fallback().verifiedSources().size()); + assertEquals(0, result.verifiedEvidence().analyses().size()); + } + + private DiagnosisDraft draft(String analysisId, String callId, String conclusion) { + return new DiagnosisDraft( + new DiagnosisDraft.Conclusion(conclusion, List.of(analysisId)), + List.of(new DiagnosisDraft.AnalysisItem( + analysisId, AnalysisKind.NORMAL, "Pool reached its limit", List.of(callId))), + List.of(new DiagnosisDraft.ActionPlanItem( + "Inspect long transactions", List.of(analysisId), false)), + List.of(new DiagnosisDraft.Recommendation( + "Add saturation alert", List.of(analysisId))), + new DiagnosisDraft.Limitations("order-service", List.of())); + } + + private void ready(InMemoryStore store, RunContext context, String callId) { + CanonicalToolInvocation invocation = CanonicalToolInvocation.projecting( + callId, context.runId(), "lookup_knowledge", "{\"query\":\"pool timeout\"}", + Instant.parse("2026-07-21T10:00:00Z")) + .markReady("raw-must-not-be-read", """ + {"evidence_status":"EVIDENCE_FOUND","tool_call_id":"call-rag-1", + "query":"pool timeout","evidence":[{"document_id":"doc-1", + "source":"runbook.md","title":"Pool guide","breadcrumb":"DB > Pool", + "excerpt":"active=50 max=50"}],"returned_count":1,"truncated":false} + """, EvidenceStatus.EVIDENCE_FOUND, + Instant.parse("2026-07-21T10:00:01Z")); + store.records.put(PREFIX + ":" + context.runId() + ":" + callId, invocation); + } + + private DiagnosisHarnessCore core() { + return new DiagnosisHarnessCore( + Clock.systemUTC(), () -> "unused", Duration.ofMinutes(5), + new RunBudgetLimits(10, 10, 10, 100_000, 100_000, 200_000, 1_000_000), + HarnessRetryPolicies.strict()); + } + + private static final class ScriptedChatModel implements ChatModel { + private final List responses; + private final List prompts = new ArrayList<>(); + private final AtomicInteger calls = new AtomicInteger(); + + private ScriptedChatModel(String... responses) { + this.responses = List.of(responses); + } + + @Override + public ChatResponse call(Prompt prompt) { + prompts.add(prompt.getContents()); + int index = calls.getAndIncrement(); + ChatResponseMetadata metadata = ChatResponseMetadata.builder() + .usage(new DefaultUsage(5, 3)).build(); + return new ChatResponse( + List.of(new Generation(new AssistantMessage(responses.get(index)))), metadata); + } + } + + private static final class InMemoryStore implements CanonicalInvocationStore { + private final Map records = new HashMap<>(); + private final CanonicalInvocationLimits limits = + new CanonicalInvocationLimits(Duration.ofHours(2), 1_000_000, 64_000); + + @Override + public CanonicalInvocationLimits limits() { + return limits; + } + + @Override + public void begin(String key, CanonicalToolInvocation invocation) { + records.put(key, invocation); + } + + @Override + public Optional find(String key) { + return Optional.ofNullable(records.get(key)); + } + + @Override + public CanonicalToolInvocation markReady(String key, String rawResponse, + String agentResult, EvidenceStatus evidenceStatus, + Instant completedAt) { + throw new UnsupportedOperationException(); + } + + @Override + public CanonicalToolInvocation markError(String key, String rawResponse, + String errorCode, Instant completedAt) { + throw new UnsupportedOperationException(); + } + } + + private record Fixture( + RunContext context, + ScriptedChatModel model, + DiagnosisReleaseUseCase useCase) { + } +}