feat(harness): complete protocol repair stop and archive ISS-016
Add repairable INVALID_PROGRESS_PROTOCOL observations, independent PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths. Archive the OpenSpec change after syncing main specs and devflow.
This commit is contained in:
@@ -0,0 +1,134 @@
|
||||
# diagnosis-information-gain-stop-contract Specification
|
||||
|
||||
## Purpose
|
||||
Diagnosis information-gain stop contract for Harness progress control, protocol repair, and safe release.
|
||||
## Requirements
|
||||
### Requirement: Harness SHALL track binary information gain per Run
|
||||
Each Diagnosis Run SHALL own a thread-safe progress tracker with `GAINED` and `NO_GAIN` as the only information-gain values. `GAINED` SHALL reset the consecutive no-gain count; `NO_GAIN` SHALL increment it. The tracker SHALL expose only `COLLECTING` or `SATURATED` as collection state and SHALL NOT create a second diagnosis lifecycle.
|
||||
|
||||
#### Scenario: New evidence advances diagnosis
|
||||
- **WHEN** a valid pending Tool observation is evaluated as `GAINED`
|
||||
- **THEN** the Run remains `COLLECTING` and its consecutive no-gain count becomes zero
|
||||
|
||||
#### Scenario: Consecutive observations do not advance diagnosis
|
||||
- **WHEN** valid `NO_GAIN` observations reach the Run's configured threshold
|
||||
- **THEN** the tracker becomes `SATURATED` with stop reason `INFORMATION_SATURATED`
|
||||
|
||||
### Requirement: Harness SHALL assign only deterministic no-gain signals
|
||||
The Harness SHALL assign `NO_GAIN` when a successful Tool result has `evidence_status=NO_EVIDENCE` or when a requested `tool_name + normalized_scope` duplicates a successfully completed scope in the same Run. Other successful non-empty results, including RAG `REFERENCE`, SHALL require a model-provided `GAINED` or `NO_GAIN` before another Tool executes.
|
||||
|
||||
#### Scenario: Empty scoped result
|
||||
- **WHEN** a Tool completes READY with `NO_EVIDENCE`
|
||||
- **THEN** the Harness records `NO_GAIN` without asking a model to judge Tool quality
|
||||
|
||||
#### Scenario: Reference material is non-empty
|
||||
- **WHEN** RAG returns bounded evidence with `relevance_level=REFERENCE`
|
||||
- **THEN** the Harness leaves it pending for model evaluation and does not automatically mark it `NO_GAIN`
|
||||
|
||||
#### Scenario: Equivalent structured scope repeats
|
||||
- **WHEN** the model requests the same Tool with the same deterministically normalized successful scope
|
||||
- **THEN** the Harness does not execute the Tool, records `NO_GAIN`, and does not create a second canonical invocation
|
||||
|
||||
### Requirement: Consecutive no-gain threshold SHALL be fixed per Run
|
||||
The system SHALL bind `harness.chat.stop-after-consecutive-no-gain`, require a value of at least one, and default it to `2`. The value SHALL be copied into each new Run's tracker and SHALL NOT enter model context or change an active Run.
|
||||
|
||||
#### Scenario: Default configuration is used
|
||||
- **WHEN** no external value is configured
|
||||
- **THEN** a new Run becomes saturated after two consecutive `NO_GAIN` decisions
|
||||
|
||||
#### Scenario: Invalid threshold is configured
|
||||
- **WHEN** the configured threshold is zero or negative
|
||||
- **THEN** Harness configuration validation fails before serving Chat requests
|
||||
|
||||
### Requirement: Saturated collection SHALL stop further Tool execution
|
||||
When collection becomes saturated, the Harness SHALL reject the pending or next business Tool execution and deliver one bounded `STOP_REQUIRED` observation with `reason=INFORMATION_SATURATED`. If the next model round requests another Tool, the Agent execution SHALL terminate through a typed controlled-stop path without consuming another Tool budget or publishing an internal failure.
|
||||
|
||||
#### Scenario: Model evaluation reaches threshold
|
||||
- **WHEN** the next Tool Call reports `NO_GAIN` and that evaluation reaches the threshold
|
||||
- **THEN** the requested business Tool is not executed and the model receives one STOP_REQUIRED observation
|
||||
|
||||
#### Scenario: Model ignores stop instruction
|
||||
- **WHEN** the model requests another Tool after STOP_REQUIRED was delivered
|
||||
- **THEN** the loop ends as controlled information saturation and proceeds to safe release
|
||||
|
||||
### Requirement: Tool loop completion SHALL project bounded progress once
|
||||
At normal Draft completion, information saturation, or budget termination, the Harness SHALL use the current Run's completed canonical invocation keys to create one bounded `ProgressSnapshot`. The snapshot SHALL contain only safe source, actual scope, objective result summary, truncation and stop reason; it SHALL NOT contain Prompt, thought, raw Tool response, internal counters, remaining budget, Redis keys, or model evaluation rationale.
|
||||
|
||||
#### Scenario: Unknown problem has completed checks
|
||||
- **WHEN** multiple Tools completed but no supported conclusion exists
|
||||
- **THEN** the release input contains their bounded checked scopes and objective results in stable execution order
|
||||
|
||||
#### Scenario: Canonical record cannot be verified
|
||||
- **WHEN** an indexed record is missing, expired, incomplete, ERROR, or not owned by the current Run
|
||||
- **THEN** it is not projected as an observed fact and the snapshot records a bounded limitation
|
||||
|
||||
### Requirement: Information stop reasons SHALL remain distinct from release outcomes
|
||||
The Harness SHALL distinguish `INFORMATION_SATURATED`, `BUDGET_LIMIT_REACHED`, and `PROGRESS_PROTOCOL_VIOLATED`. Release SHALL continue to expose only `SUCCESS`, `FALLBACK`, `FAILED`, or `CANCELLED`, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
|
||||
|
||||
#### Scenario: Low gain stops before budget exhaustion
|
||||
- **WHEN** consecutive no-gain reaches the configured threshold while hard budget remains
|
||||
- **THEN** Trace records `INFORMATION_SATURATED` and public release is a normal `FALLBACK`
|
||||
|
||||
#### Scenario: Hard budget terminates collection
|
||||
- **WHEN** model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
|
||||
- **THEN** Trace retains `BUDGET_LIMIT_REACHED` and Release attempts a deterministic `FALLBACK` from the existing progress without an extra model call
|
||||
|
||||
#### Scenario: Progress protocol violations exceed threshold
|
||||
- **WHEN** consecutive invalid progress protocol requests reach the configured threshold
|
||||
- **THEN** Trace records `PROGRESS_PROTOCOL_VIOLATED`, the model receives one STOP_REQUIRED observation, and public release remains a normal Fallback only if safe progress exists
|
||||
|
||||
### Requirement: Diagnosis Prompt SHALL license bounded abandonment
|
||||
The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, `conclusion=null` is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is `NO_GAIN`. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, `next_action`, or Harness implementation details.
|
||||
|
||||
#### Scenario: Required context is missing before any Tool call
|
||||
- **WHEN** the Query lacks the enterprise, time, service, error, or other context needed for a bounded query
|
||||
- **THEN** the model may return `conclusion=null` with `limitations.missing_info` without calling a Tool
|
||||
|
||||
#### Scenario: Correct content is diagnostically useless
|
||||
- **WHEN** a Tool response is factually correct but only generic, repeated, or unable to change a current hypothesis
|
||||
- **THEN** the model treats it as `NO_GAIN` and does not continue with an equivalent query
|
||||
|
||||
### Requirement: Model token audit SHALL be component-scoped and reconcilable
|
||||
Every Harness model call admitted by the Run budget SHALL receive a bounded component and component round. When Provider Usage is available, the same non-negative input and output Token counts SHALL update both the Run budget total and a `MODEL_TOKEN_USAGE` Trace event. Diagnosis Agent usage SHALL also update the matching `AgentStep.token_count`. At Run completion, Trace SHALL expose whether audited Token totals reconcile with the Run budget total and SHALL expose unavailable Usage counts without fabricating Token values.
|
||||
|
||||
The audit SHALL NOT persist Prompt content, user or model text, reasoning content, Tool arguments, raw model responses, credentials, or provider-specific metadata.
|
||||
|
||||
#### Scenario: Diagnosis Agent round returns Usage
|
||||
- **WHEN** a Diagnosis Agent model round returns input and output Token Usage
|
||||
- **THEN** its component round Trace and matching AgentStep contain the same total Token count and the Run total increases by that amount
|
||||
|
||||
#### Scenario: Multiple model components execute
|
||||
- **WHEN** Router, Diagnosis Agent, Evidence Repair or Semantic Guard model calls execute in one Run
|
||||
- **THEN** each call is distinguishable by bounded component and component round and their audited Token sum can be compared with the Run total
|
||||
|
||||
#### Scenario: Provider Usage is unavailable
|
||||
- **WHEN** a model attempt completes or fails without Provider Usage
|
||||
- **THEN** the audit marks Usage unavailable and Run reconciliation exposes the gap without estimating Token counts
|
||||
|
||||
### Requirement: Rejected Tool requests SHALL remain observable without payload disclosure
|
||||
Every supported Tool request rejected by the Harness before a usable business observation is delivered SHALL emit a `TOOL_REQUEST_REJECTED` Trace event containing only safe Tool Call ID, Tool name and stable error code. Rejected requests SHALL remain distinguishable from canonical `TOOL_INVOCATION` events and SHALL NOT include Tool arguments, normalized scope content, raw responses, internal exception messages, credentials or budget values.
|
||||
|
||||
#### Scenario: Progress protocol is invalid
|
||||
- **WHEN** a supported Tool request omits or misorders a required previous observation
|
||||
- **THEN** no business Tool executes, Trace records `INVALID_PROGRESS_PROTOCOL` with a safe violation type, and the model receives a repairable observation naming the missing or expected protocol field
|
||||
|
||||
#### Scenario: Duplicate or saturated request is blocked
|
||||
- **WHEN** a supported Tool request repeats a successful normalized scope or arrives after collection saturation
|
||||
- **THEN** Trace records the stable rejection reason while canonical invocation count remains unchanged
|
||||
|
||||
### Requirement: Invalid progress protocol SHALL be repairable before bounded stop
|
||||
When a supported Tool request violates the Tool Envelope progress protocol, the Harness SHALL return a bounded error observation that helps the model repair the next request. The observation MAY include safe protocol fields such as `repair_required`, `violation_type`, `missing_field`, `expected_previous_tool_call_id`, and allowed `information_gain` values. It SHALL NOT include Tool arguments, normalized scope, raw responses, Prompt, model text, budget values, counters except the bounded consecutive protocol violation count, or internal exception text.
|
||||
|
||||
Consecutive invalid progress protocol requests SHALL be counted independently from `NO_GAIN`. Reaching the Run's configured protocol-violation threshold SHALL set stop reason `PROGRESS_PROTOCOL_VIOLATED` and deliver one `STOP_REQUIRED` observation. A later Tool request after that instruction SHALL terminate through the controlled-stop path.
|
||||
|
||||
#### Scenario: Missing previous observation is repairable
|
||||
- **WHEN** a non-empty Tool observation is pending semantic evaluation and the next Tool request omits `previous_observation`
|
||||
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `missing_field=previous_observation` and the expected previous Tool Call ID
|
||||
|
||||
#### Scenario: Wrong previous observation id is repairable
|
||||
- **WHEN** a pending Tool observation exists and the next Tool request references a different `previous_observation.tool_call_id`
|
||||
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `violation_type=OUT_OF_ORDER_PREVIOUS_OBSERVATION`
|
||||
|
||||
#### Scenario: Repeated repair failure stops collection
|
||||
- **WHEN** the model repeats invalid progress protocol requests until the configured threshold is reached
|
||||
- **THEN** the current Tool is not executed, the model receives `STOP_REQUIRED` with `reason=PROGRESS_PROTOCOL_VIOLATED`, and any further Tool request ends as a controlled stop
|
||||
Reference in New Issue
Block a user