Files
SuperBizAgent-java/openspec/specs/diagnosis-information-gain-stop-contract/spec.md
T
zhuyongxin 38f781b157 feat(harness): complete protocol repair stop and archive ISS-016
Add repairable INVALID_PROGRESS_PROTOCOL observations, independent
PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths.
Archive the OpenSpec change after syncing main specs and devflow.
2026-07-27 19:10:07 +08:00

135 lines
12 KiB
Markdown

# diagnosis-information-gain-stop-contract Specification
## Purpose
Diagnosis information-gain stop contract for Harness progress control, protocol repair, and safe release.
## Requirements
### Requirement: Harness SHALL track binary information gain per Run
Each Diagnosis Run SHALL own a thread-safe progress tracker with `GAINED` and `NO_GAIN` as the only information-gain values. `GAINED` SHALL reset the consecutive no-gain count; `NO_GAIN` SHALL increment it. The tracker SHALL expose only `COLLECTING` or `SATURATED` as collection state and SHALL NOT create a second diagnosis lifecycle.
#### Scenario: New evidence advances diagnosis
- **WHEN** a valid pending Tool observation is evaluated as `GAINED`
- **THEN** the Run remains `COLLECTING` and its consecutive no-gain count becomes zero
#### Scenario: Consecutive observations do not advance diagnosis
- **WHEN** valid `NO_GAIN` observations reach the Run's configured threshold
- **THEN** the tracker becomes `SATURATED` with stop reason `INFORMATION_SATURATED`
### Requirement: Harness SHALL assign only deterministic no-gain signals
The Harness SHALL assign `NO_GAIN` when a successful Tool result has `evidence_status=NO_EVIDENCE` or when a requested `tool_name + normalized_scope` duplicates a successfully completed scope in the same Run. Other successful non-empty results, including RAG `REFERENCE`, SHALL require a model-provided `GAINED` or `NO_GAIN` before another Tool executes.
#### Scenario: Empty scoped result
- **WHEN** a Tool completes READY with `NO_EVIDENCE`
- **THEN** the Harness records `NO_GAIN` without asking a model to judge Tool quality
#### Scenario: Reference material is non-empty
- **WHEN** RAG returns bounded evidence with `relevance_level=REFERENCE`
- **THEN** the Harness leaves it pending for model evaluation and does not automatically mark it `NO_GAIN`
#### Scenario: Equivalent structured scope repeats
- **WHEN** the model requests the same Tool with the same deterministically normalized successful scope
- **THEN** the Harness does not execute the Tool, records `NO_GAIN`, and does not create a second canonical invocation
### Requirement: Consecutive no-gain threshold SHALL be fixed per Run
The system SHALL bind `harness.chat.stop-after-consecutive-no-gain`, require a value of at least one, and default it to `2`. The value SHALL be copied into each new Run's tracker and SHALL NOT enter model context or change an active Run.
#### Scenario: Default configuration is used
- **WHEN** no external value is configured
- **THEN** a new Run becomes saturated after two consecutive `NO_GAIN` decisions
#### Scenario: Invalid threshold is configured
- **WHEN** the configured threshold is zero or negative
- **THEN** Harness configuration validation fails before serving Chat requests
### Requirement: Saturated collection SHALL stop further Tool execution
When collection becomes saturated, the Harness SHALL reject the pending or next business Tool execution and deliver one bounded `STOP_REQUIRED` observation with `reason=INFORMATION_SATURATED`. If the next model round requests another Tool, the Agent execution SHALL terminate through a typed controlled-stop path without consuming another Tool budget or publishing an internal failure.
#### Scenario: Model evaluation reaches threshold
- **WHEN** the next Tool Call reports `NO_GAIN` and that evaluation reaches the threshold
- **THEN** the requested business Tool is not executed and the model receives one STOP_REQUIRED observation
#### Scenario: Model ignores stop instruction
- **WHEN** the model requests another Tool after STOP_REQUIRED was delivered
- **THEN** the loop ends as controlled information saturation and proceeds to safe release
### Requirement: Tool loop completion SHALL project bounded progress once
At normal Draft completion, information saturation, or budget termination, the Harness SHALL use the current Run's completed canonical invocation keys to create one bounded `ProgressSnapshot`. The snapshot SHALL contain only safe source, actual scope, objective result summary, truncation and stop reason; it SHALL NOT contain Prompt, thought, raw Tool response, internal counters, remaining budget, Redis keys, or model evaluation rationale.
#### Scenario: Unknown problem has completed checks
- **WHEN** multiple Tools completed but no supported conclusion exists
- **THEN** the release input contains their bounded checked scopes and objective results in stable execution order
#### Scenario: Canonical record cannot be verified
- **WHEN** an indexed record is missing, expired, incomplete, ERROR, or not owned by the current Run
- **THEN** it is not projected as an observed fact and the snapshot records a bounded limitation
### Requirement: Information stop reasons SHALL remain distinct from release outcomes
The Harness SHALL distinguish `INFORMATION_SATURATED`, `BUDGET_LIMIT_REACHED`, and `PROGRESS_PROTOCOL_VIOLATED`. Release SHALL continue to expose only `SUCCESS`, `FALLBACK`, `FAILED`, or `CANCELLED`, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
#### Scenario: Low gain stops before budget exhaustion
- **WHEN** consecutive no-gain reaches the configured threshold while hard budget remains
- **THEN** Trace records `INFORMATION_SATURATED` and public release is a normal `FALLBACK`
#### Scenario: Hard budget terminates collection
- **WHEN** model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
- **THEN** Trace retains `BUDGET_LIMIT_REACHED` and Release attempts a deterministic `FALLBACK` from the existing progress without an extra model call
#### Scenario: Progress protocol violations exceed threshold
- **WHEN** consecutive invalid progress protocol requests reach the configured threshold
- **THEN** Trace records `PROGRESS_PROTOCOL_VIOLATED`, the model receives one STOP_REQUIRED observation, and public release remains a normal Fallback only if safe progress exists
### Requirement: Diagnosis Prompt SHALL license bounded abandonment
The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, `conclusion=null` is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is `NO_GAIN`. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, `next_action`, or Harness implementation details.
#### Scenario: Required context is missing before any Tool call
- **WHEN** the Query lacks the enterprise, time, service, error, or other context needed for a bounded query
- **THEN** the model may return `conclusion=null` with `limitations.missing_info` without calling a Tool
#### Scenario: Correct content is diagnostically useless
- **WHEN** a Tool response is factually correct but only generic, repeated, or unable to change a current hypothesis
- **THEN** the model treats it as `NO_GAIN` and does not continue with an equivalent query
### Requirement: Model token audit SHALL be component-scoped and reconcilable
Every Harness model call admitted by the Run budget SHALL receive a bounded component and component round. When Provider Usage is available, the same non-negative input and output Token counts SHALL update both the Run budget total and a `MODEL_TOKEN_USAGE` Trace event. Diagnosis Agent usage SHALL also update the matching `AgentStep.token_count`. At Run completion, Trace SHALL expose whether audited Token totals reconcile with the Run budget total and SHALL expose unavailable Usage counts without fabricating Token values.
The audit SHALL NOT persist Prompt content, user or model text, reasoning content, Tool arguments, raw model responses, credentials, or provider-specific metadata.
#### Scenario: Diagnosis Agent round returns Usage
- **WHEN** a Diagnosis Agent model round returns input and output Token Usage
- **THEN** its component round Trace and matching AgentStep contain the same total Token count and the Run total increases by that amount
#### Scenario: Multiple model components execute
- **WHEN** Router, Diagnosis Agent, Evidence Repair or Semantic Guard model calls execute in one Run
- **THEN** each call is distinguishable by bounded component and component round and their audited Token sum can be compared with the Run total
#### Scenario: Provider Usage is unavailable
- **WHEN** a model attempt completes or fails without Provider Usage
- **THEN** the audit marks Usage unavailable and Run reconciliation exposes the gap without estimating Token counts
### Requirement: Rejected Tool requests SHALL remain observable without payload disclosure
Every supported Tool request rejected by the Harness before a usable business observation is delivered SHALL emit a `TOOL_REQUEST_REJECTED` Trace event containing only safe Tool Call ID, Tool name and stable error code. Rejected requests SHALL remain distinguishable from canonical `TOOL_INVOCATION` events and SHALL NOT include Tool arguments, normalized scope content, raw responses, internal exception messages, credentials or budget values.
#### Scenario: Progress protocol is invalid
- **WHEN** a supported Tool request omits or misorders a required previous observation
- **THEN** no business Tool executes, Trace records `INVALID_PROGRESS_PROTOCOL` with a safe violation type, and the model receives a repairable observation naming the missing or expected protocol field
#### Scenario: Duplicate or saturated request is blocked
- **WHEN** a supported Tool request repeats a successful normalized scope or arrives after collection saturation
- **THEN** Trace records the stable rejection reason while canonical invocation count remains unchanged
### Requirement: Invalid progress protocol SHALL be repairable before bounded stop
When a supported Tool request violates the Tool Envelope progress protocol, the Harness SHALL return a bounded error observation that helps the model repair the next request. The observation MAY include safe protocol fields such as `repair_required`, `violation_type`, `missing_field`, `expected_previous_tool_call_id`, and allowed `information_gain` values. It SHALL NOT include Tool arguments, normalized scope, raw responses, Prompt, model text, budget values, counters except the bounded consecutive protocol violation count, or internal exception text.
Consecutive invalid progress protocol requests SHALL be counted independently from `NO_GAIN`. Reaching the Run's configured protocol-violation threshold SHALL set stop reason `PROGRESS_PROTOCOL_VIOLATED` and deliver one `STOP_REQUIRED` observation. A later Tool request after that instruction SHALL terminate through the controlled-stop path.
#### Scenario: Missing previous observation is repairable
- **WHEN** a non-empty Tool observation is pending semantic evaluation and the next Tool request omits `previous_observation`
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `missing_field=previous_observation` and the expected previous Tool Call ID
#### Scenario: Wrong previous observation id is repairable
- **WHEN** a pending Tool observation exists and the next Tool request references a different `previous_observation.tool_call_id`
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `violation_type=OUT_OF_ORDER_PREVIOUS_OBSERVATION`
#### Scenario: Repeated repair failure stops collection
- **WHEN** the model repeats invalid progress protocol requests until the configured threshold is reached
- **THEN** the current Tool is not executed, the model receives `STOP_REQUIRED` with `reason=PROGRESS_PROTOCOL_VIOLATED`, and any further Tool request ends as a controlled stop