Files
zhuyongxin 38f781b157 feat(harness): complete protocol repair stop and archive ISS-016
Add repairable INVALID_PROGRESS_PROTOCOL observations, independent
PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths.
Archive the OpenSpec change after syncing main specs and devflow.
2026-07-27 19:10:07 +08:00

11 KiB

ADDED Requirements

Requirement: Harness SHALL track binary information gain per Run

Each Diagnosis Run SHALL own a thread-safe progress tracker with GAINED and NO_GAIN as the only information-gain values. GAINED SHALL reset the consecutive no-gain count; NO_GAIN SHALL increment it. The tracker SHALL expose only COLLECTING or SATURATED as collection state and SHALL NOT create a second diagnosis lifecycle.

Scenario: New evidence advances diagnosis

  • WHEN a valid pending Tool observation is evaluated as GAINED
  • THEN the Run remains COLLECTING and its consecutive no-gain count becomes zero

Scenario: Consecutive observations do not advance diagnosis

  • WHEN valid NO_GAIN observations reach the Run's configured threshold
  • THEN the tracker becomes SATURATED with stop reason INFORMATION_SATURATED

Requirement: Harness SHALL assign only deterministic no-gain signals

The Harness SHALL assign NO_GAIN when a successful Tool result has evidence_status=NO_EVIDENCE or when a requested tool_name + normalized_scope duplicates a successfully completed scope in the same Run. Other successful non-empty results, including RAG REFERENCE, SHALL require a model-provided GAINED or NO_GAIN before another Tool executes.

Scenario: Empty scoped result

  • WHEN a Tool completes READY with NO_EVIDENCE
  • THEN the Harness records NO_GAIN without asking a model to judge Tool quality

Scenario: Reference material is non-empty

  • WHEN RAG returns bounded evidence with relevance_level=REFERENCE
  • THEN the Harness leaves it pending for model evaluation and does not automatically mark it NO_GAIN

Scenario: Equivalent structured scope repeats

  • WHEN the model requests the same Tool with the same deterministically normalized successful scope
  • THEN the Harness does not execute the Tool, records NO_GAIN, and does not create a second canonical invocation

Requirement: Consecutive no-gain threshold SHALL be fixed per Run

The system SHALL bind harness.chat.stop-after-consecutive-no-gain, require a value of at least one, and default it to 2. The value SHALL be copied into each new Run's tracker and SHALL NOT enter model context or change an active Run.

Scenario: Default configuration is used

  • WHEN no external value is configured
  • THEN a new Run becomes saturated after two consecutive NO_GAIN decisions

Scenario: Invalid threshold is configured

  • WHEN the configured threshold is zero or negative
  • THEN Harness configuration validation fails before serving Chat requests

Requirement: Saturated collection SHALL stop further Tool execution

When collection becomes saturated, the Harness SHALL reject the pending or next business Tool execution and deliver one bounded STOP_REQUIRED observation with reason=INFORMATION_SATURATED. If the next model round requests another Tool, the Agent execution SHALL terminate through a typed controlled-stop path without consuming another Tool budget or publishing an internal failure.

Scenario: Model evaluation reaches threshold

  • WHEN the next Tool Call reports NO_GAIN and that evaluation reaches the threshold
  • THEN the requested business Tool is not executed and the model receives one STOP_REQUIRED observation

Scenario: Model ignores stop instruction

  • WHEN the model requests another Tool after STOP_REQUIRED was delivered
  • THEN the loop ends as controlled information saturation and proceeds to safe release

Requirement: Tool loop completion SHALL project bounded progress once

At normal Draft completion, information saturation, or budget termination, the Harness SHALL use the current Run's completed canonical invocation keys to create one bounded ProgressSnapshot. The snapshot SHALL contain only safe source, actual scope, objective result summary, truncation and stop reason; it SHALL NOT contain Prompt, thought, raw Tool response, internal counters, remaining budget, Redis keys, or model evaluation rationale.

Scenario: Unknown problem has completed checks

  • WHEN multiple Tools completed but no supported conclusion exists
  • THEN the release input contains their bounded checked scopes and objective results in stable execution order

Scenario: Canonical record cannot be verified

  • WHEN an indexed record is missing, expired, incomplete, ERROR, or not owned by the current Run
  • THEN it is not projected as an observed fact and the snapshot records a bounded limitation

Requirement: Information stop reasons SHALL remain distinct from release outcomes

The Harness SHALL distinguish INFORMATION_SATURATED, BUDGET_LIMIT_REACHED, and PROGRESS_PROTOCOL_VIOLATED. Release SHALL continue to expose only SUCCESS, FALLBACK, FAILED, or CANCELLED, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.

Scenario: Low gain stops before budget exhaustion

  • WHEN consecutive no-gain reaches the configured threshold while hard budget remains
  • THEN Trace records INFORMATION_SATURATED and public release is a normal FALLBACK

Scenario: Hard budget terminates collection

  • WHEN model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
  • THEN Trace retains BUDGET_LIMIT_REACHED and Release attempts a deterministic FALLBACK from the existing progress without an extra model call

Scenario: Progress protocol violations exceed threshold

  • WHEN consecutive invalid progress protocol requests reach the configured threshold
  • THEN Trace records PROGRESS_PROTOCOL_VIOLATED, the model receives one STOP_REQUIRED observation, and public release remains a normal Fallback only if safe progress exists

Requirement: Diagnosis Prompt SHALL license bounded abandonment

The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, conclusion=null is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is NO_GAIN. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, next_action, or Harness implementation details.

Scenario: Required context is missing before any Tool call

  • WHEN the Query lacks the enterprise, time, service, error, or other context needed for a bounded query
  • THEN the model may return conclusion=null with limitations.missing_info without calling a Tool

Scenario: Correct content is diagnostically useless

  • WHEN a Tool response is factually correct but only generic, repeated, or unable to change a current hypothesis
  • THEN the model treats it as NO_GAIN and does not continue with an equivalent query

Requirement: Model token audit SHALL be component-scoped and reconcilable

Every Harness model call admitted by the Run budget SHALL receive a bounded component and component round. When Provider Usage is available, the same non-negative input and output Token counts SHALL update both the Run budget total and a MODEL_TOKEN_USAGE Trace event. Diagnosis Agent usage SHALL also update the matching AgentStep.token_count. At Run completion, Trace SHALL expose whether audited Token totals reconcile with the Run budget total and SHALL expose unavailable Usage counts without fabricating Token values.

The audit SHALL NOT persist Prompt content, user or model text, reasoning content, Tool arguments, raw model responses, credentials, or provider-specific metadata.

Scenario: Diagnosis Agent round returns Usage

  • WHEN a Diagnosis Agent model round returns input and output Token Usage
  • THEN its component round Trace and matching AgentStep contain the same total Token count and the Run total increases by that amount

Scenario: Multiple model components execute

  • WHEN Router, Diagnosis Agent, Evidence Repair or Semantic Guard model calls execute in one Run
  • THEN each call is distinguishable by bounded component and component round and their audited Token sum can be compared with the Run total

Scenario: Provider Usage is unavailable

  • WHEN a model attempt completes or fails without Provider Usage
  • THEN the audit marks Usage unavailable and Run reconciliation exposes the gap without estimating Token counts

Requirement: Rejected Tool requests SHALL remain observable without payload disclosure

Every supported Tool request rejected by the Harness before a usable business observation is delivered SHALL emit a TOOL_REQUEST_REJECTED Trace event containing only safe Tool Call ID, Tool name and stable error code. Rejected requests SHALL remain distinguishable from canonical TOOL_INVOCATION events and SHALL NOT include Tool arguments, normalized scope content, raw responses, internal exception messages, credentials or budget values.

Scenario: Progress protocol is invalid

  • WHEN a supported Tool request omits or misorders a required previous observation
  • THEN no business Tool executes, Trace records INVALID_PROGRESS_PROTOCOL with a safe violation type, and the model receives a repairable observation naming the missing or expected protocol field

Scenario: Duplicate or saturated request is blocked

  • WHEN a supported Tool request repeats a successful normalized scope or arrives after collection saturation
  • THEN Trace records the stable rejection reason while canonical invocation count remains unchanged

Requirement: Invalid progress protocol SHALL be repairable before bounded stop

When a supported Tool request violates the Tool Envelope progress protocol, the Harness SHALL return a bounded error observation that helps the model repair the next request. The observation MAY include safe protocol fields such as repair_required, violation_type, missing_field, expected_previous_tool_call_id, and allowed information_gain values. It SHALL NOT include Tool arguments, normalized scope, raw responses, Prompt, model text, budget values, counters except the bounded consecutive protocol violation count, or internal exception text.

Consecutive invalid progress protocol requests SHALL be counted independently from NO_GAIN. Reaching the Run's configured protocol-violation threshold SHALL set stop reason PROGRESS_PROTOCOL_VIOLATED and deliver one STOP_REQUIRED observation. A later Tool request after that instruction SHALL terminate through the controlled-stop path.

Scenario: Missing previous observation is repairable

  • WHEN a non-empty Tool observation is pending semantic evaluation and the next Tool request omits previous_observation
  • THEN the Tool is not executed and the model receives a repairable INVALID_PROGRESS_PROTOCOL observation containing missing_field=previous_observation and the expected previous Tool Call ID

Scenario: Wrong previous observation id is repairable

  • WHEN a pending Tool observation exists and the next Tool request references a different previous_observation.tool_call_id
  • THEN the Tool is not executed and the model receives a repairable INVALID_PROGRESS_PROTOCOL observation containing violation_type=OUT_OF_ORDER_PREVIOUS_OBSERVATION

Scenario: Repeated repair failure stops collection

  • WHEN the model repeats invalid progress protocol requests until the configured threshold is reached
  • THEN the current Tool is not executed, the model receives STOP_REQUIRED with reason=PROGRESS_PROTOCOL_VIOLATED, and any further Tool request ends as a controlled stop