feat(harness): add evidence and semantic guards
This commit is contained in:
+1
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,118 @@
|
||||
## Context
|
||||
|
||||
阶段 2-4 已提供显式 `RunContext`、类型化重试策略、canonical invocation store、三类有界 Tool projection 和内部 `DiagnosisAgentUseCase`。当前 `DiagnosisDraft` 仍只是未发布草稿:类型反序列化不保证 ID 唯一、报告内部引用闭合,也不证明 Tool Call 属于当前 Run 或 projection 可引用;即使物理证据真实,也还需要独立判断整份报告是否被证据支持。
|
||||
|
||||
本阶段只建立内部 release boundary。公开 `ChatController`、SSE、Run persistence 和旧 Planner/Executor/Verifier/Composer 链路由阶段 6A/6B/7 处理。接口影响为 L2,新增内部 Java use case 供后续阶段消费,不改变当前外部协议。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 用确定性代码验证 Draft 结构、Analysis 引用和当前 Run canonical invocation。
|
||||
- 只从 `agent_result` 和必要的 canonical request 生成无 Tool ID、无 raw response 的 verified evidence snapshot。
|
||||
- 首次验真失败时允许一次保持用户可见语义不变的结构修复。
|
||||
- 使用同一 `ChatModel` 执行无 Tool、无记忆、无 ReAct 的隔离 SemanticGuard,并由 Harness 控制预算、超时、取消和技术重试。
|
||||
- 只发布通过两层门禁的原 Draft;所有其他可降级路径返回固定 `SafeFallback`。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不切换 `/api/chat`、`/api/chat_stream` 或 `/api/ai_ops`。
|
||||
- 不实现 SSE 事件、Run 持久化、PreviousTurn、Intent Router 或最终报告渲染。
|
||||
- 不删除或改造旧 Gatekeeper、Verifier、Composer 和多 Agent 生产链路。
|
||||
- 不访问 canonical `raw_response`,不新增 Redis 索引或历史证据召回。
|
||||
- 不执行 live model/Redis/MySQL E2E;最终 live 验收留到阶段 7。
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. EvidenceGuard 是 fail-closed 的纯确定性边界
|
||||
|
||||
`EvidenceGuard.validate(RunContext, DiagnosisDraft)` 分两步执行:先验证 Draft 顶层/嵌套结构、非空文本、Analysis ID 唯一性和 `based_on_analysis_ids` 闭合;再对每个 Analysis 的非空 Tool Call ID 使用 `ToolCallKeyFactory.create(runId, id)` 查询 `CanonicalInvocationStore`。
|
||||
|
||||
可引用记录必须同时满足:record `runId` 等于当前 Run、`status=READY`、`agent_result` 非空、`evidence_status=EVIDENCE_FOUND|NO_EVIDENCE`。`AnalysisKind.accepts` 固定 `NORMAL -> EVIDENCE_FOUND`、`NEGATIVE_OBSERVATION -> NO_EVIDENCE`。任何非法 key、缺失记录、跨 Run、错误/投影中记录、空 projection、状态不一致或解析错误都产生稳定 `EvidenceViolation`,不会抛出原始存储内容或进入 SemanticGuard。
|
||||
|
||||
替代方案是复用旧 `ExecutorGatekeeperService`;拒绝该方案,因为旧实现绑定 JPA `ToolInvocation`、Session/Run 双入口和 Map DTO,不能表达 Redis canonical lifecycle、`NO_EVIDENCE` kind 约束或新的 typed Draft。
|
||||
|
||||
### 2. verified snapshot 使用 Tool-specific 严格 projection adapter
|
||||
|
||||
Guard 根据冻结的 `AgentToolContracts` 严格解析 `RagToolResult`、`QueryLogsToolResult`、`MysqlToolResult`,并校验 projection 内 `tool_call_id`、`evidence_status` 与 canonical record 一致。未知 Tool 一律拒绝。
|
||||
|
||||
快照结构为 `analysis_id + analysis_text + analysis_kind + verified_evidence[]`。单条 evidence 只包含 `source_type/source/scope/timestamp/excerpt/values`:RAG 展开稳定 document metadata/excerpt;日志展开 pattern/event 或带精确 query/time window 的零结果;MySQL 从 canonical request 取得逻辑 data source/SQL/params scope,从 projection 取得 columns/有界 rows 或零结果。快照不包含 Tool Call ID、Redis key、raw response 或未被 Draft 引用的调用。
|
||||
|
||||
Snapshot 同时可确定性提取去重的 `SafeFallback.VerifiedSource`,因此 release policy 无需读取 Store。替代通用 JsonNode 透传;拒绝该方案,因为它会扩大模型输入面并削弱字段级测试。
|
||||
|
||||
### 3. Evidence repair 只能改引用结构,不能成为报告作者
|
||||
|
||||
首次 Guard 失败后,`EvidenceRepair` 使用直接单轮 `ChatModel.call(Prompt)`,输入为原始 query、完整 Draft 和稳定 violation code/target;没有 Tool callback、memory 或 ReAct Agent。调用由 `context.retryPolicies().evidenceRepair()` 执行,其 `maxAttempts=1`,因此不能形成隐式 retry loop。
|
||||
|
||||
修复输出经过严格 `DiagnosisDraft` 解析后,转换为移除所有标识/引用字段的 `SemanticDraftView` 并与原 Draft 比较。Conclusion/Analysis/Action/Recommendation/Limitations 的文本、kind、顺序和人工确认标志有任何变化都拒绝;只允许 `analysis_id`、`tool_call_ids` 和 `based_on_analysis_ids` 修正。修复后完整重跑 EvidenceGuard,第二次仍失败直接 `EVIDENCE_VALIDATION_FAILED`。
|
||||
|
||||
替代方案是允许模型删除无效 Analysis 或改写结论;拒绝该方案,因为 ISS-014 固定 Diagnosis Agent 是唯一报告作者,Harness/repair 不能改变用户报告语义。
|
||||
|
||||
### 4. SemanticGuard 使用直接 ChatModel 单轮调用
|
||||
|
||||
`SemanticGuardInput` 包含原始 query、`SemanticDraftView` 和 verified snapshot。View 保留完整用户可见 Draft 语义和 Analysis IDs,但删除所有 Tool Call IDs。每个 attempt 都构造只含一个 system message 和一个 user JSON message 的新 `Prompt`,直接调用系统注入的同一 `ChatModel`;不构建 `ReactAgent`,因此没有 Tool、memory、checkpoint、Graph 或 Agent 回调。
|
||||
|
||||
输出只允许精确 JSON object `{verdict, reason}`,字段集合固定,verdict 只允许 `SUPPORTED|UNSUPPORTED`,reason 必须非空。SemanticGuard 不返回 corrected report。`UNSUPPORTED` 正常返回给 release policy,不能进入 retry classifier。
|
||||
|
||||
### 5. 单轮模型执行由共享受控边界管理
|
||||
|
||||
`GuardModelCall` 在每次调用前执行 `DiagnosisHarnessCore.beforeModelCall`,在模型实际返回后从 `ChatResponseMetadata.Usage` 记录 input/output Token,并再次检查 Run active。调用提交到注入的 `ExecutorService`,等待时间为 `min(singleAttemptTimeout, totalTimeoutRemaining)`;attempt 超时或 Run cancellation 时执行 `Future.cancel(true)`。模型即使忽略中断而迟到返回,后台任务仍记录真实 Token,并在 `core.checkActive` 处阻止迟到结果被使用。
|
||||
|
||||
SemanticGuard 通过既有 `HarnessRetryExecutor` 和 `context.retryPolicies().semanticGuard()` 执行。只有 timeout、transport、parse error、schema invalid 可进行第二次 attempt;两次使用完全相同的序列化输入。Run cancelled/budget exhausted 不转换为 Fallback,而是继续抛给阶段 6A 生命周期边界;其他最终技术失败映射 `SEMANTIC_UNAVAILABLE`。
|
||||
|
||||
### 6. Release policy 不携带未验证语义
|
||||
|
||||
`DiagnosisReleaseUseCase.execute(context, query, draft)` 固定顺序为 `EvidenceGuard -> optional repair -> EvidenceGuard -> SemanticGuard -> decision`。成功结果保留与 Guard 通过的同一个 Draft 实例/修复实例,不总结、裁剪或重排。
|
||||
|
||||
- `SUPPORTED`:`ReleaseOutcome.SUCCESS`,返回完整 Draft 和 verified snapshot。
|
||||
- `UNSUPPORTED`:`ReleaseOutcome.FALLBACK` + `SEMANTIC_UNSUPPORTED`。
|
||||
- SemanticGuard 最终技术失败:`ReleaseOutcome.FALLBACK` + `SEMANTIC_UNAVAILABLE`。
|
||||
- 首次修复调用失败、语义被改写或第二次 Guard 失败:`ReleaseOutcome.FALLBACK` + `EVIDENCE_VALIDATION_FAILED`。
|
||||
|
||||
`SafeFallbackFactory` 使用固定模板。Evidence failure 的 `verified_sources=[]`;Semantic fallback 只从 snapshot 提取来源。所有 Fallback 的 `conclusion=null`,且不接收 Draft 文本、SemanticGuard reason、供应商错误或异常作为模板参数。
|
||||
Fallback release result 不保留完整 snapshot,因为 snapshot 含 Analysis text;已验真来源只以 `SafeFallback.verified_sources` 的安全子集返回,防止误序列化内部 result 泄漏草稿。
|
||||
|
||||
## Module and Ownership Audit
|
||||
|
||||
```text
|
||||
query + DiagnosisDraft + RunContext
|
||||
-> DiagnosisReleaseUseCase
|
||||
-> EvidenceGuard -> CanonicalInvocationStore + Tool projection parsers
|
||||
-> optional EvidenceRepair -> GuardModelCall -> shared ChatModel
|
||||
-> SemanticGuardInput(snapshot + ID-free Draft)
|
||||
-> SemanticGuard -> HarnessRetryExecutor -> GuardModelCall -> shared ChatModel
|
||||
-> original Draft OR SafeFallbackFactory
|
||||
```
|
||||
|
||||
- `RunContext` owns deadline、cancellation、budget、retry policy 和 lifecycle;本阶段不创建第二套 Run 状态。
|
||||
- canonical store owns Tool 调用真实性;EvidenceGuard 只读并生成隔离 snapshot,不持有 Redis client。
|
||||
- Diagnosis Agent owns report semantics;repair 只修标识结构,SemanticGuard 只审查,release 只选择 Draft 或模板。
|
||||
- 最大耦合风险是阻塞式 `ChatModel` 的取消能力;受控 Future 可发出 interrupt,但 provider 是否立即终止取决于 SDK,故迟到结果必须由 Core active check 丢弃。
|
||||
- 无新 ADR:模块边界均来自 ISS-014 已冻结设计,且公开消费者尚未切换。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Level: L2 internal interface。
|
||||
- 新增内部 records/interfaces/use cases;不修改现有 Java 方法签名。
|
||||
- 当前消费者仅 focused tests;阶段 6A 将装配真实 `DiagnosisAgentUseCase -> DiagnosisReleaseUseCase`。
|
||||
- HTTP/SSE、Controller DTO、JPA/Flyway、Redis key schema、旧 Chat/AiOps 行为均不变,无迁移或回滚数据操作。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Projection 字段遗漏导致误判] -> 三类显式 adapter、ID/status 双重校验和 contract fixture tests。
|
||||
- [Repair 改写报告] -> `SemanticDraftView` 确定性相等检查,任何可见差异直接 Evidence fallback。
|
||||
- [模型忽略 interrupt] -> Future cancel、迟到 Token 记录、Core active check;不允许迟到结果进入 release。
|
||||
- [业务 `UNSUPPORTED` 被重试] -> verdict 作为正常返回值,classifier 只处理异常。
|
||||
- [Fallback 泄漏草稿或内部 reason] -> 固定 factory 不接受这些参数,并对序列化结果做负向断言。
|
||||
- [输入快照过大] -> repair/SemanticGuard 独立 UTF-8 input/output limits,并占用 Run byte/model/token budget。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 新增内部 guard/release 包和聚焦测试,不连接公开入口。
|
||||
2. 阶段 6A 将内部 Chat application use case 装配到这些接口。
|
||||
3. 阶段 6B 在新 release boundary 通过后原子切换公开 SSE。
|
||||
4. 如阶段 5 回滚,只需移除新增内部包;阶段 4 和旧生产链路不受影响。
|
||||
|
||||
## Open Questions
|
||||
|
||||
- None。单次/总超时与字节预算由构造配置提供可调默认值,具体生产数值按阶段 7 Trace 校准,不构成本阶段方向问题。
|
||||
@@ -0,0 +1,30 @@
|
||||
## Why
|
||||
|
||||
阶段 4 已能生成带框架 Tool Call ID 的 `DiagnosisDraft`,但草稿尚未经过当前 Run 证据验真和隔离语义审查,不能安全发布。阶段 5 需要补齐确定性 EvidenceGuard、单轮 SemanticGuard 和 fail-closed 释放策略,作为阶段 6A/6B 接入公开链路的前置门禁。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增确定性 EvidenceGuard,校验 Draft 结构、Analysis ID、报告内部引用,以及当前 Run canonical Tool invocation 的生命周期、证据状态和 `agent_result`。
|
||||
- 从三类冻结 Tool projection 确定性生成按 Analysis 分组的 verified evidence snapshot;不读取 `raw_response`,不向 SemanticGuard 暴露 Tool Call ID 或 Redis 细节。
|
||||
- 首次 EvidenceGuard 失败时允许一次无 Tool、无 ReAct 的结构修复;修复后仍失败直接返回 `EVIDENCE_VALIDATION_FAILED`。
|
||||
- 新增复用系统同一 `ChatModel` 的隔离单轮 SemanticGuard,校验完整报告语义,只输出 `SUPPORTED|UNSUPPORTED + reason`,不生成或修改报告。
|
||||
- Harness 对 SemanticGuard 执行输入预算、模型/Token 预算、单次超时、取消、严格输出解析和最多两次技术 attempt;`UNSUPPORTED` 不重试。
|
||||
- 新增 release use case:`SUPPORTED` 原样释放 Draft,业务不支持或技术不可用返回固定 SafeFallback,任何失败均不泄漏 Draft 或 SemanticGuard reason。
|
||||
- 不切换公开 Chat/AiOps,不删除旧 Gatekeeper/Verifier/Composer,不修改 SSE 或持久化协议。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `single-react-evidence-semantic-guards`: 定义 DiagnosisDraft 的物理验真、verified snapshot、单次结构修复、隔离语义审查和安全释放行为。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 既有 Agent、Tool、Harness Core 和公开 Chat capability 的需求语义不在本阶段改变。
|
||||
|
||||
## Impact
|
||||
|
||||
- 新增 `com.superbiz.agent.harness.guard.evidence`、`guard.semantic` 和 `release` 内部包,以及 SemanticGuard prompt 和 focused tests。
|
||||
- 复用 `CanonicalInvocationStore`、`ToolCallKeyFactory`、三类 Tool contract、`DiagnosisHarnessCore`、`HarnessRetryExecutor`、`RunContext`、`DiagnosisDraft`、`SafeFallback` 和系统 `ChatModel`。
|
||||
- 内部接口影响为 L2:阶段 6A 将消费新的 release use case;本阶段无公开 Controller、SSE、DTO、数据库或旧 ChatService 行为变化。
|
||||
- 主要风险是快照字段遗漏、超时任务残留、错误重试业务 `UNSUPPORTED`,以及 Fallback 意外携带未验证内容;设计和测试必须逐项 fail closed。
|
||||
+112
@@ -0,0 +1,112 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Deterministic Draft and evidence validation
|
||||
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
|
||||
|
||||
#### Scenario: Duplicate or missing Analysis ID
|
||||
- **WHEN** a Draft contains a blank or duplicate Analysis ID
|
||||
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
|
||||
|
||||
#### Scenario: Broken report reference
|
||||
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
|
||||
- **THEN** EvidenceGuard rejects the Draft before semantic review
|
||||
|
||||
### Requirement: Current Run canonical invocation ownership
|
||||
EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`.
|
||||
|
||||
#### Scenario: Fabricated or cross-Run Tool Call
|
||||
- **WHEN** a Draft references a missing Tool Call or the resolved record belongs to another Run
|
||||
- **THEN** EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard
|
||||
|
||||
#### Scenario: Failed or incomplete invocation
|
||||
- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR`
|
||||
- **THEN** EvidenceGuard rejects the Draft
|
||||
|
||||
### Requirement: Analysis kind matches evidence semantics
|
||||
EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls.
|
||||
|
||||
#### Scenario: Valid negative observation
|
||||
- **WHEN** a `NEGATIVE_OBSERVATION` references a READY `NO_EVIDENCE` projection with its query scope and zero-match data
|
||||
- **THEN** EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion
|
||||
|
||||
#### Scenario: Positive claim uses no-evidence result
|
||||
- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation
|
||||
- **THEN** EvidenceGuard rejects the binding
|
||||
|
||||
### Requirement: Verified evidence snapshot is minimal and deterministic
|
||||
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
|
||||
|
||||
#### Scenario: Supported RAG, log, and MySQL projections
|
||||
- **WHEN** a Draft references valid RAG, log, or MySQL calls
|
||||
- **THEN** the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis
|
||||
|
||||
#### Scenario: Projection contract mismatch
|
||||
- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
|
||||
- **THEN** EvidenceGuard fails closed
|
||||
|
||||
### Requirement: Evidence repair is single-turn and semantics-preserving
|
||||
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
|
||||
|
||||
#### Scenario: Structural repair succeeds
|
||||
- **WHEN** the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard
|
||||
- **THEN** the repaired Draft proceeds to SemanticGuard
|
||||
|
||||
#### Scenario: Repair changes report text
|
||||
- **WHEN** the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag
|
||||
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED`
|
||||
|
||||
#### Scenario: Second validation fails
|
||||
- **WHEN** the repaired Draft still fails EvidenceGuard
|
||||
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard
|
||||
|
||||
### Requirement: Isolated single-turn SemanticGuard
|
||||
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
|
||||
|
||||
#### Scenario: Semantic input isolation
|
||||
- **WHEN** a verified Draft enters SemanticGuard
|
||||
- **THEN** the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition
|
||||
|
||||
#### Scenario: Binary review output
|
||||
- **WHEN** SemanticGuard completes normally
|
||||
- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report
|
||||
|
||||
### Requirement: Semantic model budgets timeout cancellation and retry
|
||||
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
|
||||
|
||||
#### Scenario: Technical failure then success
|
||||
- **WHEN** the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict
|
||||
- **THEN** exactly two model attempts are recorded and the second verdict controls release
|
||||
|
||||
#### Scenario: Unsupported is not retried
|
||||
- **WHEN** SemanticGuard returns valid `UNSUPPORTED`
|
||||
- **THEN** the Harness records one attempt and immediately applies the unsupported fallback
|
||||
|
||||
#### Scenario: Run cancellation during model call
|
||||
- **WHEN** the Run is cancelled while a guard model call is pending
|
||||
- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
|
||||
|
||||
### Requirement: Fail-closed release policy
|
||||
The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
|
||||
|
||||
#### Scenario: Supported report release
|
||||
- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED`
|
||||
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
|
||||
|
||||
#### Scenario: Unsupported report fallback
|
||||
- **WHEN** SemanticGuard returns `UNSUPPORTED`
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot
|
||||
|
||||
#### Scenario: Semantic review remains unavailable
|
||||
- **WHEN** all permitted technical attempts fail
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
|
||||
|
||||
#### Scenario: Evidence validation fallback sources
|
||||
- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft
|
||||
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
|
||||
|
||||
### Requirement: Stage-five public isolation
|
||||
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.
|
||||
|
||||
#### Scenario: Focused implementation scope
|
||||
- **WHEN** stage-five changes are inspected
|
||||
- **THEN** only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed
|
||||
@@ -0,0 +1,25 @@
|
||||
## 1. Evidence validation and snapshot
|
||||
|
||||
- [x] 1.1 Add typed EvidenceViolation, EvidenceGuardResult and verified snapshot contracts with immutable source extraction.
|
||||
- [x] 1.2 Implement Draft structure/reference validation and current-Run canonical invocation checks.
|
||||
- [x] 1.3 Implement strict RAG/log/MySQL projection adapters that build ID-free, raw-free evidence grouped by Analysis ID.
|
||||
- [x] 1.4 Add focused tests for duplicate/missing IDs, empty/broken references, fabricated/cross-Run calls, invalid lifecycle/result, kind/status mismatch and valid negative observations.
|
||||
|
||||
## 2. Guard model boundary and semantic review
|
||||
|
||||
- [x] 2.1 Add the shared single-turn GuardModelCall with Core model/Token budgets, UTF-8 limits, per-attempt timeout, total timeout support and Future cancellation.
|
||||
- [x] 2.2 Add SemanticDraftView and SemanticGuardInput that preserve all user-visible report semantics while excluding Tool Call IDs.
|
||||
- [x] 2.3 Implement strict binary SemanticGuard output parsing and HarnessRetryExecutor integration with stable attempt auditing.
|
||||
- [x] 2.4 Add tests proving input isolation, same-model direct calls, retryable technical failures, non-retried UNSUPPORTED, timeout cancellation and no Tool/ReAct loop.
|
||||
|
||||
## 3. Evidence repair and release policy
|
||||
|
||||
- [x] 3.1 Implement one-attempt no-Tool EvidenceRepair and reject any repair that changes SemanticDraftView.
|
||||
- [x] 3.2 Implement fixed SafeFallbackFactory and release result contracts without Draft/reason parameters on fallback paths.
|
||||
- [x] 3.3 Implement DiagnosisReleaseUseCase sequencing Guard, optional repair, SemanticGuard and fail-closed outcomes while propagating cancellation/budget exhaustion.
|
||||
- [x] 3.4 Add release tests for successful repair, second validation failure, semantic support/unsupported/unavailable, fallback types, empty evidence-failure sources and no Draft/reason leakage.
|
||||
|
||||
## 4. Verification and scope
|
||||
|
||||
- [x] 4.1 Run stage-five focused tests plus Harness/Tool/Agent regression tests and Maven compile.
|
||||
- [x] 4.2 Verify strict OpenSpec validation, no public Controller/ChatService/AiOps diff, no raw response/Tool ID in semantic input, and no hidden retry or handwritten Agent loop.
|
||||
@@ -0,0 +1,115 @@
|
||||
# single-react-evidence-semantic-guards Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change single-react-evidence-semantic-guards. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Deterministic Draft and evidence validation
|
||||
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
|
||||
|
||||
#### Scenario: Duplicate or missing Analysis ID
|
||||
- **WHEN** a Draft contains a blank or duplicate Analysis ID
|
||||
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
|
||||
|
||||
#### Scenario: Broken report reference
|
||||
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
|
||||
- **THEN** EvidenceGuard rejects the Draft before semantic review
|
||||
|
||||
### Requirement: Current Run canonical invocation ownership
|
||||
EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`.
|
||||
|
||||
#### Scenario: Fabricated or cross-Run Tool Call
|
||||
- **WHEN** a Draft references a missing Tool Call or the resolved record belongs to another Run
|
||||
- **THEN** EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard
|
||||
|
||||
#### Scenario: Failed or incomplete invocation
|
||||
- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR`
|
||||
- **THEN** EvidenceGuard rejects the Draft
|
||||
|
||||
### Requirement: Analysis kind matches evidence semantics
|
||||
EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls.
|
||||
|
||||
#### Scenario: Valid negative observation
|
||||
- **WHEN** a `NEGATIVE_OBSERVATION` references a READY `NO_EVIDENCE` projection with its query scope and zero-match data
|
||||
- **THEN** EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion
|
||||
|
||||
#### Scenario: Positive claim uses no-evidence result
|
||||
- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation
|
||||
- **THEN** EvidenceGuard rejects the binding
|
||||
|
||||
### Requirement: Verified evidence snapshot is minimal and deterministic
|
||||
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
|
||||
|
||||
#### Scenario: Supported RAG, log, and MySQL projections
|
||||
- **WHEN** a Draft references valid RAG, log, or MySQL calls
|
||||
- **THEN** the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis
|
||||
|
||||
#### Scenario: Projection contract mismatch
|
||||
- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
|
||||
- **THEN** EvidenceGuard fails closed
|
||||
|
||||
### Requirement: Evidence repair is single-turn and semantics-preserving
|
||||
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
|
||||
|
||||
#### Scenario: Structural repair succeeds
|
||||
- **WHEN** the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard
|
||||
- **THEN** the repaired Draft proceeds to SemanticGuard
|
||||
|
||||
#### Scenario: Repair changes report text
|
||||
- **WHEN** the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag
|
||||
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED`
|
||||
|
||||
#### Scenario: Second validation fails
|
||||
- **WHEN** the repaired Draft still fails EvidenceGuard
|
||||
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard
|
||||
|
||||
### Requirement: Isolated single-turn SemanticGuard
|
||||
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
|
||||
|
||||
#### Scenario: Semantic input isolation
|
||||
- **WHEN** a verified Draft enters SemanticGuard
|
||||
- **THEN** the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition
|
||||
|
||||
#### Scenario: Binary review output
|
||||
- **WHEN** SemanticGuard completes normally
|
||||
- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report
|
||||
|
||||
### Requirement: Semantic model budgets timeout cancellation and retry
|
||||
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
|
||||
|
||||
#### Scenario: Technical failure then success
|
||||
- **WHEN** the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict
|
||||
- **THEN** exactly two model attempts are recorded and the second verdict controls release
|
||||
|
||||
#### Scenario: Unsupported is not retried
|
||||
- **WHEN** SemanticGuard returns valid `UNSUPPORTED`
|
||||
- **THEN** the Harness records one attempt and immediately applies the unsupported fallback
|
||||
|
||||
#### Scenario: Run cancellation during model call
|
||||
- **WHEN** the Run is cancelled while a guard model call is pending
|
||||
- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
|
||||
|
||||
### Requirement: Fail-closed release policy
|
||||
The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
|
||||
|
||||
#### Scenario: Supported report release
|
||||
- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED`
|
||||
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
|
||||
|
||||
#### Scenario: Unsupported report fallback
|
||||
- **WHEN** SemanticGuard returns `UNSUPPORTED`
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot
|
||||
|
||||
#### Scenario: Semantic review remains unavailable
|
||||
- **WHEN** all permitted technical attempts fail
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
|
||||
|
||||
#### Scenario: Evidence validation fallback sources
|
||||
- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft
|
||||
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
|
||||
|
||||
### Requirement: Stage-five public isolation
|
||||
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.
|
||||
|
||||
#### Scenario: Focused implementation scope
|
||||
- **WHEN** stage-five changes are inspected
|
||||
- **THEN** only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed
|
||||
Reference in New Issue
Block a user