Add repairable INVALID_PROGRESS_PROTOCOL observations, independent PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths. Archive the OpenSpec change after syncing main specs and devflow.
11 KiB
ADDED Requirements
Requirement: Harness SHALL track binary information gain per Run
Each Diagnosis Run SHALL own a thread-safe progress tracker with GAINED and NO_GAIN as the only information-gain values. GAINED SHALL reset the consecutive no-gain count; NO_GAIN SHALL increment it. The tracker SHALL expose only COLLECTING or SATURATED as collection state and SHALL NOT create a second diagnosis lifecycle.
Scenario: New evidence advances diagnosis
- WHEN a valid pending Tool observation is evaluated as
GAINED - THEN the Run remains
COLLECTINGand its consecutive no-gain count becomes zero
Scenario: Consecutive observations do not advance diagnosis
- WHEN valid
NO_GAINobservations reach the Run's configured threshold - THEN the tracker becomes
SATURATEDwith stop reasonINFORMATION_SATURATED
Requirement: Harness SHALL assign only deterministic no-gain signals
The Harness SHALL assign NO_GAIN when a successful Tool result has evidence_status=NO_EVIDENCE or when a requested tool_name + normalized_scope duplicates a successfully completed scope in the same Run. Other successful non-empty results, including RAG REFERENCE, SHALL require a model-provided GAINED or NO_GAIN before another Tool executes.
Scenario: Empty scoped result
- WHEN a Tool completes READY with
NO_EVIDENCE - THEN the Harness records
NO_GAINwithout asking a model to judge Tool quality
Scenario: Reference material is non-empty
- WHEN RAG returns bounded evidence with
relevance_level=REFERENCE - THEN the Harness leaves it pending for model evaluation and does not automatically mark it
NO_GAIN
Scenario: Equivalent structured scope repeats
- WHEN the model requests the same Tool with the same deterministically normalized successful scope
- THEN the Harness does not execute the Tool, records
NO_GAIN, and does not create a second canonical invocation
Requirement: Consecutive no-gain threshold SHALL be fixed per Run
The system SHALL bind harness.chat.stop-after-consecutive-no-gain, require a value of at least one, and default it to 2. The value SHALL be copied into each new Run's tracker and SHALL NOT enter model context or change an active Run.
Scenario: Default configuration is used
- WHEN no external value is configured
- THEN a new Run becomes saturated after two consecutive
NO_GAINdecisions
Scenario: Invalid threshold is configured
- WHEN the configured threshold is zero or negative
- THEN Harness configuration validation fails before serving Chat requests
Requirement: Saturated collection SHALL stop further Tool execution
When collection becomes saturated, the Harness SHALL reject the pending or next business Tool execution and deliver one bounded STOP_REQUIRED observation with reason=INFORMATION_SATURATED. If the next model round requests another Tool, the Agent execution SHALL terminate through a typed controlled-stop path without consuming another Tool budget or publishing an internal failure.
Scenario: Model evaluation reaches threshold
- WHEN the next Tool Call reports
NO_GAINand that evaluation reaches the threshold - THEN the requested business Tool is not executed and the model receives one STOP_REQUIRED observation
Scenario: Model ignores stop instruction
- WHEN the model requests another Tool after STOP_REQUIRED was delivered
- THEN the loop ends as controlled information saturation and proceeds to safe release
Requirement: Tool loop completion SHALL project bounded progress once
At normal Draft completion, information saturation, or budget termination, the Harness SHALL use the current Run's completed canonical invocation keys to create one bounded ProgressSnapshot. The snapshot SHALL contain only safe source, actual scope, objective result summary, truncation and stop reason; it SHALL NOT contain Prompt, thought, raw Tool response, internal counters, remaining budget, Redis keys, or model evaluation rationale.
Scenario: Unknown problem has completed checks
- WHEN multiple Tools completed but no supported conclusion exists
- THEN the release input contains their bounded checked scopes and objective results in stable execution order
Scenario: Canonical record cannot be verified
- WHEN an indexed record is missing, expired, incomplete, ERROR, or not owned by the current Run
- THEN it is not projected as an observed fact and the snapshot records a bounded limitation
Requirement: Information stop reasons SHALL remain distinct from release outcomes
The Harness SHALL distinguish INFORMATION_SATURATED, BUDGET_LIMIT_REACHED, and PROGRESS_PROTOCOL_VIOLATED. Release SHALL continue to expose only SUCCESS, FALLBACK, FAILED, or CANCELLED, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
Scenario: Low gain stops before budget exhaustion
- WHEN consecutive no-gain reaches the configured threshold while hard budget remains
- THEN Trace records
INFORMATION_SATURATEDand public release is a normalFALLBACK
Scenario: Hard budget terminates collection
- WHEN model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
- THEN Trace retains
BUDGET_LIMIT_REACHEDand Release attempts a deterministicFALLBACKfrom the existing progress without an extra model call
Scenario: Progress protocol violations exceed threshold
- WHEN consecutive invalid progress protocol requests reach the configured threshold
- THEN Trace records
PROGRESS_PROTOCOL_VIOLATED, the model receives one STOP_REQUIRED observation, and public release remains a normal Fallback only if safe progress exists
Requirement: Diagnosis Prompt SHALL license bounded abandonment
The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, conclusion=null is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is NO_GAIN. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, next_action, or Harness implementation details.
Scenario: Required context is missing before any Tool call
- WHEN the Query lacks the enterprise, time, service, error, or other context needed for a bounded query
- THEN the model may return
conclusion=nullwithlimitations.missing_infowithout calling a Tool
Scenario: Correct content is diagnostically useless
- WHEN a Tool response is factually correct but only generic, repeated, or unable to change a current hypothesis
- THEN the model treats it as
NO_GAINand does not continue with an equivalent query
Requirement: Model token audit SHALL be component-scoped and reconcilable
Every Harness model call admitted by the Run budget SHALL receive a bounded component and component round. When Provider Usage is available, the same non-negative input and output Token counts SHALL update both the Run budget total and a MODEL_TOKEN_USAGE Trace event. Diagnosis Agent usage SHALL also update the matching AgentStep.token_count. At Run completion, Trace SHALL expose whether audited Token totals reconcile with the Run budget total and SHALL expose unavailable Usage counts without fabricating Token values.
The audit SHALL NOT persist Prompt content, user or model text, reasoning content, Tool arguments, raw model responses, credentials, or provider-specific metadata.
Scenario: Diagnosis Agent round returns Usage
- WHEN a Diagnosis Agent model round returns input and output Token Usage
- THEN its component round Trace and matching AgentStep contain the same total Token count and the Run total increases by that amount
Scenario: Multiple model components execute
- WHEN Router, Diagnosis Agent, Evidence Repair or Semantic Guard model calls execute in one Run
- THEN each call is distinguishable by bounded component and component round and their audited Token sum can be compared with the Run total
Scenario: Provider Usage is unavailable
- WHEN a model attempt completes or fails without Provider Usage
- THEN the audit marks Usage unavailable and Run reconciliation exposes the gap without estimating Token counts
Requirement: Rejected Tool requests SHALL remain observable without payload disclosure
Every supported Tool request rejected by the Harness before a usable business observation is delivered SHALL emit a TOOL_REQUEST_REJECTED Trace event containing only safe Tool Call ID, Tool name and stable error code. Rejected requests SHALL remain distinguishable from canonical TOOL_INVOCATION events and SHALL NOT include Tool arguments, normalized scope content, raw responses, internal exception messages, credentials or budget values.
Scenario: Progress protocol is invalid
- WHEN a supported Tool request omits or misorders a required previous observation
- THEN no business Tool executes, Trace records
INVALID_PROGRESS_PROTOCOLwith a safe violation type, and the model receives a repairable observation naming the missing or expected protocol field
Scenario: Duplicate or saturated request is blocked
- WHEN a supported Tool request repeats a successful normalized scope or arrives after collection saturation
- THEN Trace records the stable rejection reason while canonical invocation count remains unchanged
Requirement: Invalid progress protocol SHALL be repairable before bounded stop
When a supported Tool request violates the Tool Envelope progress protocol, the Harness SHALL return a bounded error observation that helps the model repair the next request. The observation MAY include safe protocol fields such as repair_required, violation_type, missing_field, expected_previous_tool_call_id, and allowed information_gain values. It SHALL NOT include Tool arguments, normalized scope, raw responses, Prompt, model text, budget values, counters except the bounded consecutive protocol violation count, or internal exception text.
Consecutive invalid progress protocol requests SHALL be counted independently from NO_GAIN. Reaching the Run's configured protocol-violation threshold SHALL set stop reason PROGRESS_PROTOCOL_VIOLATED and deliver one STOP_REQUIRED observation. A later Tool request after that instruction SHALL terminate through the controlled-stop path.
Scenario: Missing previous observation is repairable
- WHEN a non-empty Tool observation is pending semantic evaluation and the next Tool request omits
previous_observation - THEN the Tool is not executed and the model receives a repairable
INVALID_PROGRESS_PROTOCOLobservation containingmissing_field=previous_observationand the expected previous Tool Call ID
Scenario: Wrong previous observation id is repairable
- WHEN a pending Tool observation exists and the next Tool request references a different
previous_observation.tool_call_id - THEN the Tool is not executed and the model receives a repairable
INVALID_PROGRESS_PROTOCOLobservation containingviolation_type=OUT_OF_ORDER_PREVIOUS_OBSERVATION
Scenario: Repeated repair failure stops collection
- WHEN the model repeats invalid progress protocol requests until the configured threshold is reached
- THEN the current Tool is not executed, the model receives
STOP_REQUIREDwithreason=PROGRESS_PROTOCOL_VIOLATED, and any further Tool request ends as a controlled stop