116 lines
7.7 KiB
Markdown
116 lines
7.7 KiB
Markdown
# single-react-evidence-semantic-guards Specification
|
|
|
|
## Purpose
|
|
TBD - created by archiving change single-react-evidence-semantic-guards. Update Purpose after archive.
|
|
## Requirements
|
|
### Requirement: Deterministic Draft and evidence validation
|
|
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
|
|
|
|
#### Scenario: Duplicate or missing Analysis ID
|
|
- **WHEN** a Draft contains a blank or duplicate Analysis ID
|
|
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
|
|
|
|
#### Scenario: Broken report reference
|
|
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
|
|
- **THEN** EvidenceGuard rejects the Draft before semantic review
|
|
|
|
### Requirement: Current Run canonical invocation ownership
|
|
EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`.
|
|
|
|
#### Scenario: Fabricated or cross-Run Tool Call
|
|
- **WHEN** a Draft references a missing Tool Call or the resolved record belongs to another Run
|
|
- **THEN** EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard
|
|
|
|
#### Scenario: Failed or incomplete invocation
|
|
- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR`
|
|
- **THEN** EvidenceGuard rejects the Draft
|
|
|
|
### Requirement: Analysis kind matches evidence semantics
|
|
EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls.
|
|
|
|
#### Scenario: Valid negative observation
|
|
- **WHEN** a `NEGATIVE_OBSERVATION` references a READY `NO_EVIDENCE` projection with its query scope and zero-match data
|
|
- **THEN** EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion
|
|
|
|
#### Scenario: Positive claim uses no-evidence result
|
|
- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation
|
|
- **THEN** EvidenceGuard rejects the binding
|
|
|
|
### Requirement: Verified evidence snapshot is minimal and deterministic
|
|
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
|
|
|
|
#### Scenario: Supported RAG, log, and MySQL projections
|
|
- **WHEN** a Draft references valid RAG, log, or MySQL calls
|
|
- **THEN** the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis
|
|
|
|
#### Scenario: Projection contract mismatch
|
|
- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
|
|
- **THEN** EvidenceGuard fails closed
|
|
|
|
### Requirement: Evidence repair is single-turn and semantics-preserving
|
|
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
|
|
|
|
#### Scenario: Structural repair succeeds
|
|
- **WHEN** the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard
|
|
- **THEN** the repaired Draft proceeds to SemanticGuard
|
|
|
|
#### Scenario: Repair changes report text
|
|
- **WHEN** the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag
|
|
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED`
|
|
|
|
#### Scenario: Second validation fails
|
|
- **WHEN** the repaired Draft still fails EvidenceGuard
|
|
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard
|
|
|
|
### Requirement: Isolated single-turn SemanticGuard
|
|
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
|
|
|
|
#### Scenario: Semantic input isolation
|
|
- **WHEN** a verified Draft enters SemanticGuard
|
|
- **THEN** the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition
|
|
|
|
#### Scenario: Binary review output
|
|
- **WHEN** SemanticGuard completes normally
|
|
- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report
|
|
|
|
### Requirement: Semantic model budgets timeout cancellation and retry
|
|
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
|
|
|
|
#### Scenario: Technical failure then success
|
|
- **WHEN** the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict
|
|
- **THEN** exactly two model attempts are recorded and the second verdict controls release
|
|
|
|
#### Scenario: Unsupported is not retried
|
|
- **WHEN** SemanticGuard returns valid `UNSUPPORTED`
|
|
- **THEN** the Harness records one attempt and immediately applies the unsupported fallback
|
|
|
|
#### Scenario: Run cancellation during model call
|
|
- **WHEN** the Run is cancelled while a guard model call is pending
|
|
- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
|
|
|
|
### Requirement: Fail-closed release policy
|
|
The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
|
|
|
|
#### Scenario: Supported report release
|
|
- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED`
|
|
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
|
|
|
|
#### Scenario: Unsupported report fallback
|
|
- **WHEN** SemanticGuard returns `UNSUPPORTED`
|
|
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot
|
|
|
|
#### Scenario: Semantic review remains unavailable
|
|
- **WHEN** all permitted technical attempts fail
|
|
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
|
|
|
|
#### Scenario: Evidence validation fallback sources
|
|
- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft
|
|
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
|
|
|
|
### Requirement: Stage-five public isolation
|
|
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.
|
|
|
|
#### Scenario: Focused implementation scope
|
|
- **WHEN** stage-five changes are inspected
|
|
- **THEN** only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed
|