feat(harness): complete protocol repair stop and archive ISS-016
Add repairable INVALID_PROGRESS_PROTOCOL observations, independent PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths. Archive the OpenSpec change after syncing main specs and devflow.
This commit is contained in:
@@ -13,7 +13,6 @@ The internal diagnosis path SHALL create exactly one `diagnosis_agent` with the
|
||||
#### Scenario: Agent execution fails
|
||||
- **WHEN** the single Diagnosis Agent invocation throws or returns invalid structured output
|
||||
- **THEN** the internal use case fails closed without automatically invoking the Agent or model again
|
||||
|
||||
### Requirement: Diagnosis input SHALL contain only current Query and optional PreviousTurn
|
||||
The internal use case SHALL accept a non-blank current Query and an optional frozen `PreviousTurn`, serialize them as the fixed `query` and `previous_turn` input fields, and SHALL NOT load or accept complete Session history, Redis memory, prior raw Tool results, or model-generated history summaries. The current Query SHALL retain its original text and SHALL NOT be rewritten or silently truncated.
|
||||
|
||||
@@ -24,7 +23,6 @@ The internal use case SHALL accept a non-blank current Query and an optional fro
|
||||
#### Scenario: Query exceeds configured context budget
|
||||
- **WHEN** the current Query exceeds its UTF-8 byte limit
|
||||
- **THEN** the use case rejects it before any model call instead of truncating or rewriting it
|
||||
|
||||
### Requirement: Every model round SHALL be controlled by RunContext
|
||||
The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model interceptor to call `DiagnosisHarnessCore.beforeModelCall` for every framework ReAct model round. Non-streaming model response Usage SHALL be recorded into the same Run budget when available. Cancellation, deadline, model-call exhaustion, or Token exhaustion SHALL prevent subsequent controlled work.
|
||||
|
||||
@@ -35,13 +33,20 @@ The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model
|
||||
#### Scenario: Model-call budget is exhausted
|
||||
- **WHEN** the framework attempts a model round beyond the configured maximum
|
||||
- **THEN** the call is rejected before reaching ChatModel and the Run records budget exhaustion
|
||||
|
||||
### Requirement: Evidence Tools SHALL execute through the Harness boundary
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions. A Tool interceptor SHALL propagate the exact framework Tool Call ID, Run ID, Tool name, and raw JSON arguments into the corresponding stage 3B/3C adapter and `ToolBoundary`. Successful observations SHALL contain only the bounded `agent_result`; failed observations SHALL contain only stable error semantics and SHALL NOT contain raw responses, internal exceptions, credentials, or invocation lifecycle internals.
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions through native Tool Calling. Each Agent-facing Tool input SHALL be a typed Envelope containing optional `previous_observation` and required business `input`. A Tool interceptor SHALL validate and consume the previous observation, propagate the exact framework Tool Call ID, Run ID, Tool name, and unwrapped business JSON into the corresponding adapter and `ToolBoundary`, and SHALL enforce Run progress before execution. Successful observations SHALL use a bounded per-Tool whitelist projection; failed or control observations SHALL contain only stable safe semantics and SHALL NOT contain raw responses, internal exceptions, credentials, invocation lifecycle internals, counters, thresholds, or remaining budget.
|
||||
|
||||
#### Scenario: Framework requests RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`
|
||||
- **THEN** the RAG adapter and Agent observation use exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
#### Scenario: Framework requests first RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`, no pending evaluation and a typed business input
|
||||
- **THEN** the RAG adapter receives only the unwrapped business request, uses exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
|
||||
#### Scenario: Model continues after a non-empty result
|
||||
- **WHEN** the last successful Tool result is pending semantic evaluation and the model requests another Tool
|
||||
- **THEN** the Envelope must identify that exact prior Tool Call and contain `GAINED` or `NO_GAIN` before the new business Tool can execute
|
||||
|
||||
#### Scenario: Model omits required progress field
|
||||
- **WHEN** a prior non-empty Tool observation is pending and the model requests another Tool without `previous_observation`
|
||||
- **THEN** the business Tool does not execute and the Agent receives a bounded repair observation explaining the missing `previous_observation` field and expected prior Tool Call ID
|
||||
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** a registered adapter returns an error result
|
||||
@@ -50,7 +55,6 @@ The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs
|
||||
#### Scenario: Unknown Tool is requested
|
||||
- **WHEN** a model requests a Tool outside the three registered definitions
|
||||
- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
|
||||
|
||||
### Requirement: Diagnosis output SHALL be a bounded DiagnosisDraft
|
||||
The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL return JSON that the internal use case strictly parses into the frozen record. The use case SHALL enforce configured UTF-8 limits for query, previous turn, total input and Draft output and account accepted input/output bytes against the Run capacity. It SHALL reject blank, fenced, prefixed, malformed, oversized or schema-incompatible output without repair or retry.
|
||||
|
||||
@@ -61,14 +65,20 @@ The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL retu
|
||||
#### Scenario: Model returns prose around JSON
|
||||
- **WHEN** the final response contains a Markdown fence or explanatory prefix around an otherwise valid object
|
||||
- **THEN** strict parsing fails and the Diagnosis Agent is not invoked a second time
|
||||
|
||||
### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
|
||||
The single Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. If current evidence cannot support a diagnosis, the Agent SHALL stop the current ReAct execution with `conclusion=null`, describe the actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist or fabricate a root cause.
|
||||
The single Chinese Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. The Agent SHALL NOT be required to find a root cause. If current evidence cannot support a diagnosis, no required context exists for a bounded Tool call, or available results are correct but do not advance any diagnosis hypothesis, the Agent SHALL stop with `conclusion=null`, describe actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist, fabricate a root cause, or make equivalent Tool calls merely to show activity.
|
||||
|
||||
#### Scenario: Tool finds no evidence
|
||||
- **WHEN** the only completed Tool observation has `evidence_status=NO_EVIDENCE`
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records the bounded negative observation and missing information, and makes no additional automatic retry
|
||||
- **WHEN** completed Tool observations have no diagnostic information gain
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records bounded checked scope and missing information, and does not make an equivalent retry
|
||||
|
||||
#### Scenario: No bounded Tool query is possible
|
||||
- **WHEN** the Query lacks required enterprise, time, service or error context
|
||||
- **THEN** the Agent may perform zero Tool calls and returns a no-conclusion Draft whose `limitations.missing_info` identifies the required context
|
||||
|
||||
#### Scenario: Harness requires stop
|
||||
- **WHEN** the Agent receives `STOP_REQUIRED`
|
||||
- **THEN** it emits a bounded final Draft without another Tool call
|
||||
### Requirement: Stage-four execution SHALL remain internal and auditable
|
||||
The new use case SHALL be callable only as an internal Java/test entry in this stage and SHALL NOT be wired into public Chat or AIOps controllers. It SHALL propagate the same sessionId/runId through RunnableConfig metadata, allow the existing AgentStep Hook to be injected, retain ToolBoundary canonical invocation audit, and leave final Run success/persistence ownership to later Guard/application stages.
|
||||
|
||||
@@ -79,3 +89,21 @@ The new use case SHALL be callable only as an internal Java/test entry in this s
|
||||
#### Scenario: Stage four is archived
|
||||
- **WHEN** focused Agent tests pass and the change is archived
|
||||
- **THEN** `ChatController`, public SSE behavior and the old multi-Agent implementation remain available for stages 5, 6A and 6B
|
||||
### Requirement: Diagnosis execution SHALL preserve controlled stop outcomes
|
||||
The internal Agent use case SHALL distinguish a valid Draft, controlled information saturation, and budget termination from an unclassified Agent failure. It SHALL return a bounded internal execution result containing optional Draft, ProgressSnapshot and stop reason, and SHALL NOT convert a recognized controlled stop into `DiagnosisAgentOutputException`.
|
||||
|
||||
#### Scenario: Model ignores STOP_REQUIRED
|
||||
- **WHEN** the framework surfaces the typed collection-stopped signal after the final completion opportunity
|
||||
- **THEN** the Agent use case returns no Draft with the controlled stop reason and the current ProgressSnapshot
|
||||
|
||||
#### Scenario: Unclassified framework failure
|
||||
- **WHEN** Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination
|
||||
- **THEN** execution still fails closed and no safe progress is fabricated
|
||||
|
||||
#### Scenario: Invalid final Draft after verified checks
|
||||
- **WHEN** the final model text is empty or violates the strict DiagnosisDraft contract after the current Run has completed READY canonical Tool checks
|
||||
- **THEN** the invalid text is discarded, the output failure carries only the bounded ProgressSnapshot and safe failure metadata, and no model repair or loose JSON extraction occurs
|
||||
|
||||
#### Scenario: Invalid final Draft without verified checks
|
||||
- **WHEN** the final model text violates the strict DiagnosisDraft contract before any publishable ProgressSnapshot exists
|
||||
- **THEN** execution remains failed and MUST NOT fabricate missing context, observed facts or a no-conclusion Draft
|
||||
|
||||
Reference in New Issue
Block a user