feat(harness): complete protocol repair stop and archive ISS-016

Add repairable INVALID_PROGRESS_PROTOCOL observations, independent
PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths.
Archive the OpenSpec change after syncing main specs and devflow.
This commit is contained in:
zhuyongxin
2026-07-27 19:10:07 +08:00
parent 5c369f3b6c
commit 38f781b157
44 changed files with 977 additions and 80 deletions
@@ -4,16 +4,23 @@
TBD - created by archiving change single-react-evidence-semantic-guards. Update Purpose after archive.
## Requirements
### Requirement: Deterministic Draft and evidence validation
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
For a Draft with a non-null Conclusion, the Harness SHALL deterministically reject it unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs. For a Draft with `conclusion=null`, Release SHALL NOT require normal conclusion structure or invoke EvidenceRepair; any supplied Tool references SHALL still resolve to current-Run READY canonical invocations and SHALL obey positive/negative evidence semantics.
#### Scenario: Duplicate or missing Analysis ID
- **WHEN** a Draft contains a blank or duplicate Analysis ID
#### Scenario: Duplicate or missing Analysis ID in concluded Draft
- **WHEN** a Draft with a Conclusion contains a blank or duplicate Analysis ID
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
#### Scenario: Broken report reference
#### Scenario: Broken report reference in concluded Draft
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
- **THEN** EvidenceGuard rejects the Draft before semantic review
#### Scenario: No-conclusion Draft has valid negative observation
- **WHEN** a `conclusion=null` Draft cites a current-Run READY `NO_EVIDENCE` call as `NEGATIVE_OBSERVATION`
- **THEN** Release accepts the reference authenticity without running EvidenceRepair or SemanticGuard
#### Scenario: No-conclusion Draft fabricates a Tool reference
- **WHEN** a `conclusion=null` Draft cites a missing, cross-Run, incomplete or ERROR Tool call
- **THEN** the reference is excluded and cannot be published as an observed fact
### Requirement: Current Run canonical invocation ownership
EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`.
@@ -24,7 +31,6 @@ EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallI
#### Scenario: Failed or incomplete invocation
- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR`
- **THEN** EvidenceGuard rejects the Draft
### Requirement: Analysis kind matches evidence semantics
EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls.
@@ -35,7 +41,6 @@ EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls an
#### Scenario: Positive claim uses no-evidence result
- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation
- **THEN** EvidenceGuard rejects the binding
### Requirement: Verified evidence snapshot is minimal and deterministic
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
@@ -46,7 +51,6 @@ The Harness SHALL strictly parse only supported Tool projections and SHALL const
#### Scenario: Projection contract mismatch
- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
- **THEN** EvidenceGuard fails closed
### Requirement: Evidence repair is single-turn and semantics-preserving
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
@@ -61,7 +65,6 @@ On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct n
#### Scenario: Second validation fails
- **WHEN** the repaired Draft still fails EvidenceGuard
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard
### Requirement: Isolated single-turn SemanticGuard
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
@@ -72,7 +75,6 @@ SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Promp
#### Scenario: Binary review output
- **WHEN** SemanticGuard completes normally
- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report
### Requirement: Semantic model budgets timeout cancellation and retry
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
@@ -87,12 +89,11 @@ The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets
#### Scenario: Run cancellation during model call
- **WHEN** the Run is cancelled while a guard model call is pending
- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
### Requirement: Fail-closed release policy
The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
The release use case SHALL publish the unchanged verified Draft only when a non-null Conclusion passes EvidenceGuard and SemanticGuard returns `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, final semantic technical failure, a valid no-conclusion Draft, information saturation, or handled budget termination. No-conclusion and controlled-stop release SHALL be deterministic from verified references and ProgressSnapshot and MUST NOT invoke a repair or semantic model call. A fallback result MUST NOT include an unsupported Draft, full verified snapshot, internal stop counters, or SemanticGuard reason.
#### Scenario: Supported report release
- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED`
- **WHEN** EvidenceGuard succeeds for a concluded Draft and SemanticGuard returns `SUPPORTED`
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
#### Scenario: Unsupported report fallback
@@ -104,9 +105,20 @@ The release use case SHALL publish the unchanged verified Draft only for `SUPPOR
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
#### Scenario: Evidence validation fallback sources
- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft
- **WHEN** evidence repair fails or the second EvidenceGuard rejects a concluded Draft
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
#### Scenario: Missing context ends without Tool calls
- **WHEN** a valid no-conclusion Draft has no Tool calls and identifies required missing context
- **THEN** release outcome is `FALLBACK`, type is `MISSING_REQUIRED_CONTEXT`, and no guard model call occurs
#### Scenario: Finite checks do not support a conclusion
- **WHEN** a no-conclusion Draft or controlled stop has a non-empty verified ProgressSnapshot
- **THEN** release outcome is `FALLBACK`, type is `INSUFFICIENT_EVIDENCE`, and observed facts describe only actual completed checks
#### Scenario: Invalid Draft has publishable progress
- **WHEN** the Agent's final Draft is rejected by strict parsing but its bounded ProgressSnapshot contains current-Run verified observed facts
- **THEN** Release publishes `FALLBACK` with type `INSUFFICIENT_EVIDENCE` using only that snapshot and MUST NOT use any content from the invalid Draft
### Requirement: Stage-five public isolation
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.