Files
SuperBizAgent-java/openspec/specs/single-react-evidence-semantic-guards/spec.md
T
zhuyongxin 38f781b157 feat(harness): complete protocol repair stop and archive ISS-016
Add repairable INVALID_PROGRESS_PROTOCOL observations, independent
PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths.
Archive the OpenSpec change after syncing main specs and devflow.
2026-07-27 19:10:07 +08:00

9.8 KiB

single-react-evidence-semantic-guards Specification

Purpose

TBD - created by archiving change single-react-evidence-semantic-guards. Update Purpose after archive.

Requirements

Requirement: Deterministic Draft and evidence validation

For a Draft with a non-null Conclusion, the Harness SHALL deterministically reject it unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs. For a Draft with conclusion=null, Release SHALL NOT require normal conclusion structure or invoke EvidenceRepair; any supplied Tool references SHALL still resolve to current-Run READY canonical invocations and SHALL obey positive/negative evidence semantics.

Scenario: Duplicate or missing Analysis ID in concluded Draft

  • WHEN a Draft with a Conclusion contains a blank or duplicate Analysis ID
  • THEN EvidenceGuard returns violations and SemanticGuard is not invoked

Scenario: Broken report reference in concluded Draft

  • WHEN a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
  • THEN EvidenceGuard rejects the Draft before semantic review

Scenario: No-conclusion Draft has valid negative observation

  • WHEN a conclusion=null Draft cites a current-Run READY NO_EVIDENCE call as NEGATIVE_OBSERVATION
  • THEN Release accepts the reference authenticity without running EvidenceRepair or SemanticGuard

Scenario: No-conclusion Draft fabricates a Tool reference

  • WHEN a conclusion=null Draft cites a missing, cross-Run, incomplete or ERROR Tool call
  • THEN the reference is excluded and cannot be published as an observed fact

Requirement: Current Run canonical invocation ownership

EvidenceGuard SHALL resolve each referenced Tool Call through runId + toolCallId and SHALL accept only an invocation owned by the current Run with lifecycle READY, a non-empty agent_result, and evidence status EVIDENCE_FOUND or NO_EVIDENCE.

Scenario: Fabricated or cross-Run Tool Call

  • WHEN a Draft references a missing Tool Call or the resolved record belongs to another Run
  • THEN EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard

Scenario: Failed or incomplete invocation

  • WHEN a referenced invocation is PROJECTING, ERROR, lacks agent_result, or has evidence_status=ERROR
  • THEN EvidenceGuard rejects the Draft

Requirement: Analysis kind matches evidence semantics

EvidenceGuard SHALL permit NORMAL Analysis only with EVIDENCE_FOUND calls and SHALL permit NEGATIVE_OBSERVATION Analysis only with NO_EVIDENCE calls.

Scenario: Valid negative observation

  • WHEN a NEGATIVE_OBSERVATION references a READY NO_EVIDENCE projection with its query scope and zero-match data
  • THEN EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion

Scenario: Positive claim uses no-evidence result

  • WHEN a NORMAL Analysis references a NO_EVIDENCE invocation
  • THEN EvidenceGuard rejects the binding

Requirement: Verified evidence snapshot is minimal and deterministic

The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced agent_result and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.

Scenario: Supported RAG, log, and MySQL projections

  • WHEN a Draft references valid RAG, log, or MySQL calls
  • THEN the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis

Scenario: Projection contract mismatch

  • WHEN a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
  • THEN EvidenceGuard fails closed

Requirement: Evidence repair is single-turn and semantics-preserving

On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.

Scenario: Structural repair succeeds

  • WHEN the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard
  • THEN the repaired Draft proceeds to SemanticGuard

Scenario: Repair changes report text

  • WHEN the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag
  • THEN the Harness returns EVIDENCE_VALIDATION_FAILED

Scenario: Second validation fails

  • WHEN the repaired Draft still fails EvidenceGuard
  • THEN the Harness returns EVIDENCE_VALIDATION_FAILED with empty verified sources and does not invoke SemanticGuard

Requirement: Isolated single-turn SemanticGuard

SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.

Scenario: Semantic input isolation

  • WHEN a verified Draft enters SemanticGuard
  • THEN the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition

Scenario: Binary review output

  • WHEN SemanticGuard completes normally
  • THEN it returns only SUPPORTED or UNSUPPORTED with a non-blank audit reason and cannot return a corrected report

Requirement: Semantic model budgets timeout cancellation and retry

The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.

Scenario: Technical failure then success

  • WHEN the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict
  • THEN exactly two model attempts are recorded and the second verdict controls release

Scenario: Unsupported is not retried

  • WHEN SemanticGuard returns valid UNSUPPORTED
  • THEN the Harness records one attempt and immediately applies the unsupported fallback

Scenario: Run cancellation during model call

  • WHEN the Run is cancelled while a guard model call is pending
  • THEN the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback

Requirement: Fail-closed release policy

The release use case SHALL publish the unchanged verified Draft only when a non-null Conclusion passes EvidenceGuard and SemanticGuard returns SUPPORTED. It SHALL publish fixed SafeFallback content for evidence failure, semantic unsupported, final semantic technical failure, a valid no-conclusion Draft, information saturation, or handled budget termination. No-conclusion and controlled-stop release SHALL be deterministic from verified references and ProgressSnapshot and MUST NOT invoke a repair or semantic model call. A fallback result MUST NOT include an unsupported Draft, full verified snapshot, internal stop counters, or SemanticGuard reason.

Scenario: Supported report release

  • WHEN EvidenceGuard succeeds for a concluded Draft and SemanticGuard returns SUPPORTED
  • THEN release outcome is SUCCESS and the same verified Draft semantics are returned without summarization or partial editing

Scenario: Unsupported report fallback

  • WHEN SemanticGuard returns UNSUPPORTED
  • THEN release outcome is FALLBACK, type is SEMANTIC_UNSUPPORTED, and verified sources are derived only from the snapshot

Scenario: Semantic review remains unavailable

  • WHEN all permitted technical attempts fail
  • THEN release outcome is FALLBACK, type is SEMANTIC_UNAVAILABLE, and no Draft or internal failure reason is exposed

Scenario: Evidence validation fallback sources

  • WHEN evidence repair fails or the second EvidenceGuard rejects a concluded Draft
  • THEN release outcome is FALLBACK, type is EVIDENCE_VALIDATION_FAILED, and verified_sources is empty

Scenario: Missing context ends without Tool calls

  • WHEN a valid no-conclusion Draft has no Tool calls and identifies required missing context
  • THEN release outcome is FALLBACK, type is MISSING_REQUIRED_CONTEXT, and no guard model call occurs

Scenario: Finite checks do not support a conclusion

  • WHEN a no-conclusion Draft or controlled stop has a non-empty verified ProgressSnapshot
  • THEN release outcome is FALLBACK, type is INSUFFICIENT_EVIDENCE, and observed facts describe only actual completed checks

Scenario: Invalid Draft has publishable progress

  • WHEN the Agent's final Draft is rejected by strict parsing but its bounded ProgressSnapshot contains current-Run verified observed facts
  • THEN Release publishes FALLBACK with type INSUFFICIENT_EVIDENCE using only that snapshot and MUST NOT use any content from the invalid Draft

Requirement: Stage-five public isolation

The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.

Scenario: Focused implementation scope

  • WHEN stage-five changes are inspected
  • THEN only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed