7.7 KiB
single-react-evidence-semantic-guards Specification
Purpose
TBD - created by archiving change single-react-evidence-semantic-guards. Update Purpose after archive.
Requirements
Requirement: Deterministic Draft and evidence validation
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
Scenario: Duplicate or missing Analysis ID
- WHEN a Draft contains a blank or duplicate Analysis ID
- THEN EvidenceGuard returns violations and SemanticGuard is not invoked
Scenario: Broken report reference
- WHEN a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
- THEN EvidenceGuard rejects the Draft before semantic review
Requirement: Current Run canonical invocation ownership
EvidenceGuard SHALL resolve each referenced Tool Call through runId + toolCallId and SHALL accept only an invocation owned by the current Run with lifecycle READY, a non-empty agent_result, and evidence status EVIDENCE_FOUND or NO_EVIDENCE.
Scenario: Fabricated or cross-Run Tool Call
- WHEN a Draft references a missing Tool Call or the resolved record belongs to another Run
- THEN EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard
Scenario: Failed or incomplete invocation
- WHEN a referenced invocation is
PROJECTING,ERROR, lacksagent_result, or hasevidence_status=ERROR - THEN EvidenceGuard rejects the Draft
Requirement: Analysis kind matches evidence semantics
EvidenceGuard SHALL permit NORMAL Analysis only with EVIDENCE_FOUND calls and SHALL permit NEGATIVE_OBSERVATION Analysis only with NO_EVIDENCE calls.
Scenario: Valid negative observation
- WHEN a
NEGATIVE_OBSERVATIONreferences a READYNO_EVIDENCEprojection with its query scope and zero-match data - THEN EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion
Scenario: Positive claim uses no-evidence result
- WHEN a
NORMALAnalysis references aNO_EVIDENCEinvocation - THEN EvidenceGuard rejects the binding
Requirement: Verified evidence snapshot is minimal and deterministic
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced agent_result and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
Scenario: Supported RAG, log, and MySQL projections
- WHEN a Draft references valid RAG, log, or MySQL calls
- THEN the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis
Scenario: Projection contract mismatch
- WHEN a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
- THEN EvidenceGuard fails closed
Requirement: Evidence repair is single-turn and semantics-preserving
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
Scenario: Structural repair succeeds
- WHEN the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard
- THEN the repaired Draft proceeds to SemanticGuard
Scenario: Repair changes report text
- WHEN the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag
- THEN the Harness returns
EVIDENCE_VALIDATION_FAILED
Scenario: Second validation fails
- WHEN the repaired Draft still fails EvidenceGuard
- THEN the Harness returns
EVIDENCE_VALIDATION_FAILEDwith empty verified sources and does not invoke SemanticGuard
Requirement: Isolated single-turn SemanticGuard
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
Scenario: Semantic input isolation
- WHEN a verified Draft enters SemanticGuard
- THEN the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition
Scenario: Binary review output
- WHEN SemanticGuard completes normally
- THEN it returns only
SUPPORTEDorUNSUPPORTEDwith a non-blank audit reason and cannot return a corrected report
Requirement: Semantic model budgets timeout cancellation and retry
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
Scenario: Technical failure then success
- WHEN the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict
- THEN exactly two model attempts are recorded and the second verdict controls release
Scenario: Unsupported is not retried
- WHEN SemanticGuard returns valid
UNSUPPORTED - THEN the Harness records one attempt and immediately applies the unsupported fallback
Scenario: Run cancellation during model call
- WHEN the Run is cancelled while a guard model call is pending
- THEN the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
Requirement: Fail-closed release policy
The release use case SHALL publish the unchanged verified Draft only for SUPPORTED. It SHALL publish fixed SafeFallback content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
Scenario: Supported report release
- WHEN EvidenceGuard succeeds and SemanticGuard returns
SUPPORTED - THEN release outcome is
SUCCESSand the same verified Draft semantics are returned without summarization or partial editing
Scenario: Unsupported report fallback
- WHEN SemanticGuard returns
UNSUPPORTED - THEN release outcome is
FALLBACK, type isSEMANTIC_UNSUPPORTED, and verified sources are derived only from the snapshot
Scenario: Semantic review remains unavailable
- WHEN all permitted technical attempts fail
- THEN release outcome is
FALLBACK, type isSEMANTIC_UNAVAILABLE, and no Draft or internal failure reason is exposed
Scenario: Evidence validation fallback sources
- WHEN evidence repair fails or the second EvidenceGuard rejects the Draft
- THEN release outcome is
FALLBACK, type isEVIDENCE_VALIDATION_FAILED, andverified_sourcesis empty
Requirement: Stage-five public isolation
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.
Scenario: Focused implementation scope
- WHEN stage-five changes are inspected
- THEN only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed