feat(harness): complete protocol repair stop and archive ISS-016
Add repairable INVALID_PROGRESS_PROTOCOL observations, independent PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths. Archive the OpenSpec change after syncing main specs and devflow.
This commit is contained in:
@@ -14,7 +14,6 @@ The system SHALL expose `EVIDENCE_FOUND`, `NO_EVIDENCE`, and `ERROR` as Agent-fa
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** schema, authorization, execution, or projection fails
|
||||
- **THEN** the Agent-facing result uses `evidence_status=ERROR` and SHALL NOT report `NO_EVIDENCE`
|
||||
|
||||
### Requirement: Referencable Tool results SHALL use the framework Tool Call ID
|
||||
Each `EVIDENCE_FOUND` or `NO_EVIDENCE` result SHALL contain the non-blank `tool_call_id` supplied by the framework Tool Call request. Agent inputs SHALL NOT contain `tool_call_id`, and contract code SHALL NOT generate, replace, or derive a second call ID.
|
||||
|
||||
@@ -25,18 +24,16 @@ Each `EVIDENCE_FOUND` or `NO_EVIDENCE` result SHALL contain the non-blank `tool_
|
||||
#### Scenario: Framework ID is invalid
|
||||
- **WHEN** the Tool request has a missing or invalid framework Tool Call ID
|
||||
- **THEN** the call returns an `ERROR` result without inventing a referencable ID
|
||||
|
||||
### Requirement: RAG Tool contract SHALL expose only bounded document evidence
|
||||
The RAG Request SHALL contain only `query`. The RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`.
|
||||
The RAG business Request SHALL contain only `query`. The canonical RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, optional normalized `relevance_level`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`. The RAG projector SHALL accept upstream `relevanceLevel` or `relevance_level` and normalize recognized values without exposing raw relevance scores or retrieval traces.
|
||||
|
||||
#### Scenario: RAG evidence is serialized
|
||||
- **WHEN** a RAG result contains a matching document excerpt
|
||||
- **THEN** its JSON matches the frozen snake_case fields and excludes ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
|
||||
- **WHEN** a RAG result contains a matching document excerpt and an upstream relevance level
|
||||
- **THEN** its canonical JSON preserves the bounded evidence and normalized `relevance_level` while excluding ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
|
||||
|
||||
#### Scenario: RAG query has no evidence
|
||||
- **WHEN** RAG executes successfully without a usable document excerpt
|
||||
- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, and returns an empty evidence list
|
||||
|
||||
- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, returns an empty evidence list, and does not upgrade relevance into evidence
|
||||
### Requirement: Log Tool contract SHALL use logical scope and retain Mock provenance
|
||||
The log Request SHALL contain logical `topic`, `query`, and optional `lookback_minutes` only. The log Result SHALL contain `evidence_status`, `tool_call_id`, `source_kind`, complete query `scope`, `match_count`, `returned_count`, bounded `patterns`, bounded timeline `events`, and `truncated`. The initial logical topics SHALL be `APPLICATION`, `DATABASE_SLOW_QUERY`, and `SYSTEM_EVENTS`, and the initial source kind SHALL be `MOCK`.
|
||||
|
||||
@@ -47,7 +44,6 @@ The log Request SHALL contain logical `topic`, `query`, and optional `lookback_m
|
||||
#### Scenario: Agent creates a log request
|
||||
- **WHEN** the Agent requests log evidence
|
||||
- **THEN** it selects a logical topic and lookback window without supplying region, physical TopicId, credentials, or result limit and without calling a Topic discovery Tool first
|
||||
|
||||
### Requirement: MySQL Tool contract SHALL expose a logical read-only query interface
|
||||
The MySQL Request SHALL contain only logical `data_source`, parameterized `sql`, and `params`. The MySQL Result SHALL contain `evidence_status`, `tool_call_id`, `columns`, bounded structured `rows`, `returned_count`, and `truncated`; it SHALL NOT expose connection details, credentials, internal stack traces, or resource-limit controls.
|
||||
|
||||
@@ -58,17 +54,25 @@ The MySQL Request SHALL contain only logical `data_source`, parameterized `sql`,
|
||||
#### Scenario: Agent creates a MySQL request
|
||||
- **WHEN** the Agent requests business database evidence
|
||||
- **THEN** it supplies a logical data source, SQL placeholders, and parameter values without supplying JDBC connection information or security policy
|
||||
|
||||
### Requirement: Tool descriptions SHALL be concise and implementation-neutral
|
||||
The system SHALL define stable snake_case names and concise descriptions for `lookup_knowledge`, `query_logs`, and `query_mysql`. Each description SHALL state when to call the Tool, its minimal input, and what it cannot query, and SHALL NOT describe retrieval internals, infrastructure coordinates, credentials, audit storage, retries, result limits, or ranking implementation.
|
||||
|
||||
#### Scenario: Agent receives Tool definitions
|
||||
- **WHEN** a future Agent adapter registers the frozen Tool definitions
|
||||
- **THEN** the definitions describe available actions and boundaries without exposing Milvus, L0/L1, rerank, CLS region/TopicId, Redis, JDBC credentials, topK, limit, or Trace internals
|
||||
|
||||
### Requirement: Contract freeze SHALL NOT cut over the current runtime
|
||||
This change SHALL add contract types, descriptions, and tests without changing the current `LookupKnowledgeTool`, `QueryLogsTools`, ChatService, AiOpsService, Controller, Tool registration, or Agent-visible runtime results.
|
||||
|
||||
#### Scenario: Stage 1 tests pass
|
||||
- **WHEN** all ACI contract tests pass
|
||||
- **THEN** the current public Chat and AIOps paths still execute the old Tool implementations until their later projector and cutover changes
|
||||
### Requirement: Tool results SHALL have separate Harness and model views
|
||||
Each successful evidence Tool result SHALL provide a Harness Control View and a bounded Model Observation derived from the same canonical result. The control view MAY contain returned counts, normalized scope, relevance, truncation and duplicate identity. The Model Observation SHALL contain only fields needed to understand and cite the result and SHALL NOT contain raw responses, internal scores, retrieval traces, duplicate fingerprints, counters, thresholds, budgets or store identities.
|
||||
|
||||
#### Scenario: Model receives RAG observation
|
||||
- **WHEN** a canonical RAG result is READY
|
||||
- **THEN** the model receives Tool Call ID, actual query scope, bounded evidence, evidence status, optional coarse relevance and truncation, but not raw scores or Harness counters
|
||||
|
||||
#### Scenario: Harness evaluates duplicate scope
|
||||
- **WHEN** the same normalized Tool scope is requested again
|
||||
- **THEN** the Harness can compare its control view identity without exposing that fingerprint to the model
|
||||
|
||||
@@ -14,7 +14,6 @@ The Harness SHALL reject a Tool call before invoking the executor when the envel
|
||||
#### Scenario: Unauthorized or writable Tool is rejected
|
||||
- **WHEN** `authorized=false` or `readOnly=false`
|
||||
- **THEN** ToolBoundary returns `ERROR` before execution and does not expose the request to an Agent
|
||||
|
||||
### Requirement: Canonical invocation SHALL keep complete data in one record
|
||||
The store SHALL save the complete request JSON, raw Tool response, bounded Agent result, framework `tool_call_id`, run ID, tool name, lifecycle status, evidence status, timestamps, and error code in one canonical record. The raw response SHALL NOT be replaced by a preview, and SHALL NOT be returned to the Agent.
|
||||
|
||||
@@ -25,7 +24,6 @@ The store SHALL save the complete request JSON, raw Tool response, bounded Agent
|
||||
#### Scenario: Projector succeeds
|
||||
- **WHEN** the executor returns raw data and the projector returns a bounded result
|
||||
- **THEN** the same record contains raw data and Agent result with `status=READY`, and ToolBoundary returns only the bounded Agent result
|
||||
|
||||
### Requirement: Invocation lifecycle and evidence semantics SHALL be enforced centrally
|
||||
The store SHALL allow only `PROJECTING -> READY` or `PROJECTING -> ERROR`. READY SHALL require `EVIDENCE_FOUND` or `NO_EVIDENCE`; ERROR SHALL use `evidence_status=ERROR` and SHALL NOT be referencable. `NO_EVIDENCE` SHALL remain a scoped negative observation and SHALL NOT be upgraded to success or retried by the boundary.
|
||||
|
||||
@@ -36,7 +34,6 @@ The store SHALL allow only `PROJECTING -> READY` or `PROJECTING -> ERROR`. READY
|
||||
#### Scenario: Projection fails
|
||||
- **WHEN** the projector throws or returns an invalid evidence status
|
||||
- **THEN** the same record becomes `ERROR`, stores a safe error code, and no Agent result is returned
|
||||
|
||||
### Requirement: Tool Call ID and Run ownership SHALL be preserved
|
||||
The boundary SHALL use the exact framework `tool_call_id` with the RunContext run ID to create the canonical key. It SHALL reject missing, unsafe, duplicate, and cross-Run references and SHALL never generate or replace a fallback ID.
|
||||
|
||||
@@ -47,7 +44,6 @@ The boundary SHALL use the exact framework `tool_call_id` with the RunContext ru
|
||||
#### Scenario: Framework ID is preserved
|
||||
- **WHEN** a valid envelope passes preflight
|
||||
- **THEN** the key and canonical record contain the exact supplied `tool_call_id`
|
||||
|
||||
### Requirement: TTL and size limits SHALL fail closed
|
||||
The Redis implementation SHALL set TTL only when a canonical record is created. Reads SHALL NOT refresh TTL; updates SHALL preserve only the current remaining TTL. Request/raw/record and Agent-result byte limits SHALL be enforced in UTF-8. Raw overflow SHALL produce `ERROR/RESULT_TOO_LARGE` without silent truncation; Agent projection overflow SHALL produce the same error.
|
||||
|
||||
@@ -62,17 +58,25 @@ The Redis implementation SHALL set TTL only when a canonical record is created.
|
||||
#### Scenario: Agent result is too large
|
||||
- **WHEN** a projector returns a result beyond the Agent projection limit
|
||||
- **THEN** the record becomes `ERROR/RESULT_TOO_LARGE` and the oversized result is not returned to the Agent
|
||||
|
||||
### Requirement: Tool and store failures SHALL return safe errors
|
||||
Execution, serialization, Redis, duplicate, preflight, and projection failures SHALL return a bounded error code without raw payload, credentials, internal stack trace, or Redis key to the Agent. If raw data was safely stored before a projection failure, it remains Harness-only canonical data.
|
||||
|
||||
#### Scenario: Tool execution throws
|
||||
- **WHEN** the executor raises an exception after PROJECTING begins
|
||||
- **THEN** the record becomes `ERROR` with a stable execution error code and ToolBoundary returns no raw response
|
||||
|
||||
### Requirement: Stage 3A SHALL remain reusable and independent from legacy audit
|
||||
The boundary and canonical store SHALL be usable by later RAG/log/MySQL projectors without modifying legacy `ToolInvocationRecorder`, JPA `ToolInvocation`, Chat/AIOps, Controller, or public protocol. Redis access SHALL remain inside the store implementation and no Agent-facing API SHALL expose the Redis client.
|
||||
|
||||
#### Scenario: Fake projector tests pass
|
||||
- **WHEN** Fake Tool/Projector tests cover success, no evidence, error, duplicate, TTL, and capacity cases
|
||||
- **THEN** later Tool-specific changes can depend on the same boundary without copying lifecycle or Redis logic, while legacy call sites remain unchanged
|
||||
### Requirement: Run progress projection SHALL reference canonical records without duplicating truth
|
||||
The Run progress tracker SHALL retain only ordered canonical identities for completed Tool calls. At Tool-loop completion, a projector SHALL resolve those identities through the existing canonical store and SHALL accept only READY records owned by the current Run. The tracker SHALL NOT store or reconstruct raw Tool responses, complete Agent results, or a second durable evidence record.
|
||||
|
||||
#### Scenario: Completed calls are projected
|
||||
- **WHEN** a Run ends after multiple READY canonical Tool invocations
|
||||
- **THEN** the progress projector reads each indexed canonical record in execution order and creates bounded observed facts
|
||||
|
||||
#### Scenario: Indexed identity is invalid
|
||||
- **WHEN** an indexed canonical identity is missing, expired, cross-Run, PROJECTING, or ERROR
|
||||
- **THEN** it is excluded from observed facts and cannot become verified evidence
|
||||
|
||||
@@ -0,0 +1,134 @@
|
||||
# diagnosis-information-gain-stop-contract Specification
|
||||
|
||||
## Purpose
|
||||
Diagnosis information-gain stop contract for Harness progress control, protocol repair, and safe release.
|
||||
## Requirements
|
||||
### Requirement: Harness SHALL track binary information gain per Run
|
||||
Each Diagnosis Run SHALL own a thread-safe progress tracker with `GAINED` and `NO_GAIN` as the only information-gain values. `GAINED` SHALL reset the consecutive no-gain count; `NO_GAIN` SHALL increment it. The tracker SHALL expose only `COLLECTING` or `SATURATED` as collection state and SHALL NOT create a second diagnosis lifecycle.
|
||||
|
||||
#### Scenario: New evidence advances diagnosis
|
||||
- **WHEN** a valid pending Tool observation is evaluated as `GAINED`
|
||||
- **THEN** the Run remains `COLLECTING` and its consecutive no-gain count becomes zero
|
||||
|
||||
#### Scenario: Consecutive observations do not advance diagnosis
|
||||
- **WHEN** valid `NO_GAIN` observations reach the Run's configured threshold
|
||||
- **THEN** the tracker becomes `SATURATED` with stop reason `INFORMATION_SATURATED`
|
||||
|
||||
### Requirement: Harness SHALL assign only deterministic no-gain signals
|
||||
The Harness SHALL assign `NO_GAIN` when a successful Tool result has `evidence_status=NO_EVIDENCE` or when a requested `tool_name + normalized_scope` duplicates a successfully completed scope in the same Run. Other successful non-empty results, including RAG `REFERENCE`, SHALL require a model-provided `GAINED` or `NO_GAIN` before another Tool executes.
|
||||
|
||||
#### Scenario: Empty scoped result
|
||||
- **WHEN** a Tool completes READY with `NO_EVIDENCE`
|
||||
- **THEN** the Harness records `NO_GAIN` without asking a model to judge Tool quality
|
||||
|
||||
#### Scenario: Reference material is non-empty
|
||||
- **WHEN** RAG returns bounded evidence with `relevance_level=REFERENCE`
|
||||
- **THEN** the Harness leaves it pending for model evaluation and does not automatically mark it `NO_GAIN`
|
||||
|
||||
#### Scenario: Equivalent structured scope repeats
|
||||
- **WHEN** the model requests the same Tool with the same deterministically normalized successful scope
|
||||
- **THEN** the Harness does not execute the Tool, records `NO_GAIN`, and does not create a second canonical invocation
|
||||
|
||||
### Requirement: Consecutive no-gain threshold SHALL be fixed per Run
|
||||
The system SHALL bind `harness.chat.stop-after-consecutive-no-gain`, require a value of at least one, and default it to `2`. The value SHALL be copied into each new Run's tracker and SHALL NOT enter model context or change an active Run.
|
||||
|
||||
#### Scenario: Default configuration is used
|
||||
- **WHEN** no external value is configured
|
||||
- **THEN** a new Run becomes saturated after two consecutive `NO_GAIN` decisions
|
||||
|
||||
#### Scenario: Invalid threshold is configured
|
||||
- **WHEN** the configured threshold is zero or negative
|
||||
- **THEN** Harness configuration validation fails before serving Chat requests
|
||||
|
||||
### Requirement: Saturated collection SHALL stop further Tool execution
|
||||
When collection becomes saturated, the Harness SHALL reject the pending or next business Tool execution and deliver one bounded `STOP_REQUIRED` observation with `reason=INFORMATION_SATURATED`. If the next model round requests another Tool, the Agent execution SHALL terminate through a typed controlled-stop path without consuming another Tool budget or publishing an internal failure.
|
||||
|
||||
#### Scenario: Model evaluation reaches threshold
|
||||
- **WHEN** the next Tool Call reports `NO_GAIN` and that evaluation reaches the threshold
|
||||
- **THEN** the requested business Tool is not executed and the model receives one STOP_REQUIRED observation
|
||||
|
||||
#### Scenario: Model ignores stop instruction
|
||||
- **WHEN** the model requests another Tool after STOP_REQUIRED was delivered
|
||||
- **THEN** the loop ends as controlled information saturation and proceeds to safe release
|
||||
|
||||
### Requirement: Tool loop completion SHALL project bounded progress once
|
||||
At normal Draft completion, information saturation, or budget termination, the Harness SHALL use the current Run's completed canonical invocation keys to create one bounded `ProgressSnapshot`. The snapshot SHALL contain only safe source, actual scope, objective result summary, truncation and stop reason; it SHALL NOT contain Prompt, thought, raw Tool response, internal counters, remaining budget, Redis keys, or model evaluation rationale.
|
||||
|
||||
#### Scenario: Unknown problem has completed checks
|
||||
- **WHEN** multiple Tools completed but no supported conclusion exists
|
||||
- **THEN** the release input contains their bounded checked scopes and objective results in stable execution order
|
||||
|
||||
#### Scenario: Canonical record cannot be verified
|
||||
- **WHEN** an indexed record is missing, expired, incomplete, ERROR, or not owned by the current Run
|
||||
- **THEN** it is not projected as an observed fact and the snapshot records a bounded limitation
|
||||
|
||||
### Requirement: Information stop reasons SHALL remain distinct from release outcomes
|
||||
The Harness SHALL distinguish `INFORMATION_SATURATED`, `BUDGET_LIMIT_REACHED`, and `PROGRESS_PROTOCOL_VIOLATED`. Release SHALL continue to expose only `SUCCESS`, `FALLBACK`, `FAILED`, or `CANCELLED`, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
|
||||
|
||||
#### Scenario: Low gain stops before budget exhaustion
|
||||
- **WHEN** consecutive no-gain reaches the configured threshold while hard budget remains
|
||||
- **THEN** Trace records `INFORMATION_SATURATED` and public release is a normal `FALLBACK`
|
||||
|
||||
#### Scenario: Hard budget terminates collection
|
||||
- **WHEN** model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
|
||||
- **THEN** Trace retains `BUDGET_LIMIT_REACHED` and Release attempts a deterministic `FALLBACK` from the existing progress without an extra model call
|
||||
|
||||
#### Scenario: Progress protocol violations exceed threshold
|
||||
- **WHEN** consecutive invalid progress protocol requests reach the configured threshold
|
||||
- **THEN** Trace records `PROGRESS_PROTOCOL_VIOLATED`, the model receives one STOP_REQUIRED observation, and public release remains a normal Fallback only if safe progress exists
|
||||
|
||||
### Requirement: Diagnosis Prompt SHALL license bounded abandonment
|
||||
The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, `conclusion=null` is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is `NO_GAIN`. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, `next_action`, or Harness implementation details.
|
||||
|
||||
#### Scenario: Required context is missing before any Tool call
|
||||
- **WHEN** the Query lacks the enterprise, time, service, error, or other context needed for a bounded query
|
||||
- **THEN** the model may return `conclusion=null` with `limitations.missing_info` without calling a Tool
|
||||
|
||||
#### Scenario: Correct content is diagnostically useless
|
||||
- **WHEN** a Tool response is factually correct but only generic, repeated, or unable to change a current hypothesis
|
||||
- **THEN** the model treats it as `NO_GAIN` and does not continue with an equivalent query
|
||||
|
||||
### Requirement: Model token audit SHALL be component-scoped and reconcilable
|
||||
Every Harness model call admitted by the Run budget SHALL receive a bounded component and component round. When Provider Usage is available, the same non-negative input and output Token counts SHALL update both the Run budget total and a `MODEL_TOKEN_USAGE` Trace event. Diagnosis Agent usage SHALL also update the matching `AgentStep.token_count`. At Run completion, Trace SHALL expose whether audited Token totals reconcile with the Run budget total and SHALL expose unavailable Usage counts without fabricating Token values.
|
||||
|
||||
The audit SHALL NOT persist Prompt content, user or model text, reasoning content, Tool arguments, raw model responses, credentials, or provider-specific metadata.
|
||||
|
||||
#### Scenario: Diagnosis Agent round returns Usage
|
||||
- **WHEN** a Diagnosis Agent model round returns input and output Token Usage
|
||||
- **THEN** its component round Trace and matching AgentStep contain the same total Token count and the Run total increases by that amount
|
||||
|
||||
#### Scenario: Multiple model components execute
|
||||
- **WHEN** Router, Diagnosis Agent, Evidence Repair or Semantic Guard model calls execute in one Run
|
||||
- **THEN** each call is distinguishable by bounded component and component round and their audited Token sum can be compared with the Run total
|
||||
|
||||
#### Scenario: Provider Usage is unavailable
|
||||
- **WHEN** a model attempt completes or fails without Provider Usage
|
||||
- **THEN** the audit marks Usage unavailable and Run reconciliation exposes the gap without estimating Token counts
|
||||
|
||||
### Requirement: Rejected Tool requests SHALL remain observable without payload disclosure
|
||||
Every supported Tool request rejected by the Harness before a usable business observation is delivered SHALL emit a `TOOL_REQUEST_REJECTED` Trace event containing only safe Tool Call ID, Tool name and stable error code. Rejected requests SHALL remain distinguishable from canonical `TOOL_INVOCATION` events and SHALL NOT include Tool arguments, normalized scope content, raw responses, internal exception messages, credentials or budget values.
|
||||
|
||||
#### Scenario: Progress protocol is invalid
|
||||
- **WHEN** a supported Tool request omits or misorders a required previous observation
|
||||
- **THEN** no business Tool executes, Trace records `INVALID_PROGRESS_PROTOCOL` with a safe violation type, and the model receives a repairable observation naming the missing or expected protocol field
|
||||
|
||||
#### Scenario: Duplicate or saturated request is blocked
|
||||
- **WHEN** a supported Tool request repeats a successful normalized scope or arrives after collection saturation
|
||||
- **THEN** Trace records the stable rejection reason while canonical invocation count remains unchanged
|
||||
|
||||
### Requirement: Invalid progress protocol SHALL be repairable before bounded stop
|
||||
When a supported Tool request violates the Tool Envelope progress protocol, the Harness SHALL return a bounded error observation that helps the model repair the next request. The observation MAY include safe protocol fields such as `repair_required`, `violation_type`, `missing_field`, `expected_previous_tool_call_id`, and allowed `information_gain` values. It SHALL NOT include Tool arguments, normalized scope, raw responses, Prompt, model text, budget values, counters except the bounded consecutive protocol violation count, or internal exception text.
|
||||
|
||||
Consecutive invalid progress protocol requests SHALL be counted independently from `NO_GAIN`. Reaching the Run's configured protocol-violation threshold SHALL set stop reason `PROGRESS_PROTOCOL_VIOLATED` and deliver one `STOP_REQUIRED` observation. A later Tool request after that instruction SHALL terminate through the controlled-stop path.
|
||||
|
||||
#### Scenario: Missing previous observation is repairable
|
||||
- **WHEN** a non-empty Tool observation is pending semantic evaluation and the next Tool request omits `previous_observation`
|
||||
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `missing_field=previous_observation` and the expected previous Tool Call ID
|
||||
|
||||
#### Scenario: Wrong previous observation id is repairable
|
||||
- **WHEN** a pending Tool observation exists and the next Tool request references a different `previous_observation.tool_call_id`
|
||||
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `violation_type=OUT_OF_ORDER_PREVIOUS_OBSERVATION`
|
||||
|
||||
#### Scenario: Repeated repair failure stops collection
|
||||
- **WHEN** the model repeats invalid progress protocol requests until the configured threshold is reached
|
||||
- **THEN** the current Tool is not executed, the model receives `STOP_REQUIRED` with `reason=PROGRESS_PROTOCOL_VIOLATED`, and any further Tool request ends as a controlled stop
|
||||
@@ -13,7 +13,6 @@ The Chat Application Use Case SHALL resolve or generate the session ID, create e
|
||||
#### Scenario: New session request
|
||||
- **WHEN** an internal request omits the session ID
|
||||
- **THEN** the use case generates one valid session ID and uses it for all Run operations
|
||||
|
||||
### Requirement: Minimal isolated Intent Router
|
||||
The Router SHALL execute a direct no-Tool, no-memory, no-ReAct ChatModel call whose input contains only the unchanged Query, optional last intent, and optional last user Query. It SHALL accept only `SYSTEM_CHAT`, `KNOWLEDGE_QUERY`, or `DIAGNOSIS`.
|
||||
|
||||
@@ -24,7 +23,6 @@ The Router SHALL execute a direct no-Tool, no-memory, no-ReAct ChatModel call wh
|
||||
#### Scenario: New topic overrides history
|
||||
- **WHEN** the current Query identifies a new topic while prior routing context exists
|
||||
- **THEN** the model input still preserves the original current Query and prior fields are only optional context
|
||||
|
||||
### Requirement: Router technical retry fails closed
|
||||
The Router SHALL use `HarnessRetryPolicies.intentRouter()` and SHALL retry timeout, transport, or invalid output at most once with identical input. A second failure MUST produce `ROUTING_UNAVAILABLE` and MUST NOT dispatch Diagnosis.
|
||||
|
||||
@@ -35,9 +33,8 @@ The Router SHALL use `HarnessRetryPolicies.intentRouter()` and SHALL retry timeo
|
||||
#### Scenario: Two invalid outputs
|
||||
- **WHEN** both permitted attempts return invalid output
|
||||
- **THEN** the Run ends FAILED and no intent executor is called
|
||||
|
||||
### Requirement: Fixed isolated executors
|
||||
The Application Use Case SHALL map SYSTEM_CHAT to one no-Tool model response, KNOWLEDGE_QUERY to exactly one lookup-knowledge invocation plus one bounded answer model call, and DIAGNOSIS to the single Diagnosis Agent followed by the release boundary. Executors MUST NOT call one another or rewrite the Query.
|
||||
The Application Use Case SHALL map SYSTEM_CHAT to one no-Tool model response, KNOWLEDGE_QUERY to exactly one lookup-knowledge invocation plus one bounded answer model call, and DIAGNOSIS to the single Diagnosis Agent followed by the Diagnosis Release boundary. Executors MUST NOT call one another or rewrite the Query. Information saturation and budget termination in Diagnosis SHALL be converted to safe content by Diagnosis Release, not by ChatApplicationUseCase.
|
||||
|
||||
#### Scenario: System Chat
|
||||
- **WHEN** intent is SYSTEM_CHAT
|
||||
@@ -47,10 +44,13 @@ The Application Use Case SHALL map SYSTEM_CHAT to one no-Tool model response, KN
|
||||
- **WHEN** intent is KNOWLEDGE_QUERY
|
||||
- **THEN** only lookup_knowledge is invoked once and query_logs/query_mysql/Diagnosis ReAct are unavailable
|
||||
|
||||
#### Scenario: Diagnosis
|
||||
- **WHEN** intent is DIAGNOSIS
|
||||
- **THEN** the original Query and bounded PreviousTurn enter DiagnosisAgentUseCase and the Draft cannot publish before DiagnosisReleaseUseCase
|
||||
#### Scenario: Diagnosis succeeds with a conclusion
|
||||
- **WHEN** intent is DIAGNOSIS and the Draft passes the release guards
|
||||
- **THEN** the original Query and bounded PreviousTurn enter DiagnosisAgentUseCase and the Draft publishes only after DiagnosisReleaseUseCase
|
||||
|
||||
#### Scenario: Diagnosis stops without a conclusion
|
||||
- **WHEN** Diagnosis collection is saturated, required context is missing, or a handled budget limit is reached
|
||||
- **THEN** DiagnosisReleaseUseCase returns bounded Fallback content and ChatApplicationUseCase only persists and transports that decision
|
||||
### Requirement: Knowledge references are physically validated
|
||||
The Knowledge executor SHALL accept only a bounded READY RAG projection, SHALL validate every model answer item against the exact direct invocation ID and returned document ID set, and SHALL remove Tool Call IDs from public content.
|
||||
|
||||
@@ -65,7 +65,6 @@ The Knowledge executor SHALL accept only a bounded READY RAG projection, SHALL v
|
||||
#### Scenario: No knowledge evidence
|
||||
- **WHEN** lookup returns READY `NO_EVIDENCE`
|
||||
- **THEN** the executor returns a fixed bounded no-evidence answer without calling the answer model
|
||||
|
||||
### Requirement: Safe bounded PreviousTurn
|
||||
The system SHALL load PreviousTurn only from the same Session's most recent Run with `intent=DIAGNOSIS`, `release_outcome=SUCCESS`, and non-null valid `published_result`. It SHALL deterministically bound fields and source documents without model summarization.
|
||||
|
||||
@@ -80,7 +79,6 @@ The system SHALL load PreviousTurn only from the same Session's most recent Run
|
||||
#### Scenario: Published result is corrupt
|
||||
- **WHEN** stored JSON is invalid or required safe fields are blank
|
||||
- **THEN** PreviousTurn is null and no raw stored value reaches a model
|
||||
|
||||
### Requirement: Diagnosis Run persistence contract
|
||||
The `diagnosis_run` schema SHALL add nullable `intent`, `release_outcome`, and JSON `published_result`, and the JPA entity/repository/store SHALL write and query them consistently. PublishedResult MUST NOT contain Tool IDs, raw evidence, full Draft, or SemanticGuard reasons.
|
||||
|
||||
@@ -95,7 +93,6 @@ The `diagnosis_run` schema SHALL add nullable `intent`, `release_outcome`, and J
|
||||
#### Scenario: Failure or cancellation
|
||||
- **WHEN** routing/execution fails or the client cancels the Run
|
||||
- **THEN** the Run stores exactly one FAILED or CANCELLED release outcome and no PublishedResult
|
||||
|
||||
### Requirement: Protocol-neutral progress and cancellation
|
||||
The use case SHALL notify a protocol-neutral observer after the Run is persisted, expose only session/run identifiers and client-disconnect cancellation, and emit only fixed safe application status codes.
|
||||
|
||||
@@ -106,10 +103,19 @@ The use case SHALL notify a protocol-neutral observer after the Run is persisted
|
||||
#### Scenario: Client disconnect control
|
||||
- **WHEN** the observer invokes client-disconnect cancellation
|
||||
- **THEN** the same RunContext is cancelled and late executor results cannot complete successfully
|
||||
|
||||
### Requirement: Stage-six-A public isolation
|
||||
Stage 6A SHALL NOT modify or switch public Chat Controller endpoints, SSE contracts, frontend consumers, or legacy ChatService behavior.
|
||||
|
||||
#### Scenario: Internal-only delivery
|
||||
- **WHEN** stage 6A changes are inspected
|
||||
- **THEN** Controller/frontend/public endpoint behavior has zero diff and stage 6B can consume the completed use case without rewriting it
|
||||
### Requirement: Handled Diagnosis budget termination SHALL remain a Fallback release
|
||||
When Diagnosis Release has converted a recognized budget termination and existing safe progress into a Fallback, ChatApplicationUseCase SHALL persist that result exactly once without reclassifying it as `INTERNAL_FAILURE`, invoking another model, or rebuilding business fallback content. The internal Run lifecycle MAY retain `BUDGET_EXHAUSTED`, while the persisted public release outcome SHALL be `FALLBACK` and no PublishedResult SHALL be stored.
|
||||
|
||||
#### Scenario: Tool budget ends after finite checks
|
||||
- **WHEN** Diagnosis reaches a hard Tool budget after at least one canonical safe observation and Release creates an insufficient-evidence fallback
|
||||
- **THEN** the Run persists status SUCCESS, release outcome FALLBACK, safe content and actual budget usage, and the SSE sends content followed by done
|
||||
|
||||
#### Scenario: Budget ends without safe publishable progress
|
||||
- **WHEN** budget termination occurs before Diagnosis Release can form a safe bounded result
|
||||
- **THEN** the existing failure path remains fail closed and does not fabricate observed facts
|
||||
|
||||
@@ -13,7 +13,6 @@ The internal diagnosis path SHALL create exactly one `diagnosis_agent` with the
|
||||
#### Scenario: Agent execution fails
|
||||
- **WHEN** the single Diagnosis Agent invocation throws or returns invalid structured output
|
||||
- **THEN** the internal use case fails closed without automatically invoking the Agent or model again
|
||||
|
||||
### Requirement: Diagnosis input SHALL contain only current Query and optional PreviousTurn
|
||||
The internal use case SHALL accept a non-blank current Query and an optional frozen `PreviousTurn`, serialize them as the fixed `query` and `previous_turn` input fields, and SHALL NOT load or accept complete Session history, Redis memory, prior raw Tool results, or model-generated history summaries. The current Query SHALL retain its original text and SHALL NOT be rewritten or silently truncated.
|
||||
|
||||
@@ -24,7 +23,6 @@ The internal use case SHALL accept a non-blank current Query and an optional fro
|
||||
#### Scenario: Query exceeds configured context budget
|
||||
- **WHEN** the current Query exceeds its UTF-8 byte limit
|
||||
- **THEN** the use case rejects it before any model call instead of truncating or rewriting it
|
||||
|
||||
### Requirement: Every model round SHALL be controlled by RunContext
|
||||
The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model interceptor to call `DiagnosisHarnessCore.beforeModelCall` for every framework ReAct model round. Non-streaming model response Usage SHALL be recorded into the same Run budget when available. Cancellation, deadline, model-call exhaustion, or Token exhaustion SHALL prevent subsequent controlled work.
|
||||
|
||||
@@ -35,13 +33,20 @@ The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model
|
||||
#### Scenario: Model-call budget is exhausted
|
||||
- **WHEN** the framework attempts a model round beyond the configured maximum
|
||||
- **THEN** the call is rejected before reaching ChatModel and the Run records budget exhaustion
|
||||
|
||||
### Requirement: Evidence Tools SHALL execute through the Harness boundary
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions. A Tool interceptor SHALL propagate the exact framework Tool Call ID, Run ID, Tool name, and raw JSON arguments into the corresponding stage 3B/3C adapter and `ToolBoundary`. Successful observations SHALL contain only the bounded `agent_result`; failed observations SHALL contain only stable error semantics and SHALL NOT contain raw responses, internal exceptions, credentials, or invocation lifecycle internals.
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions through native Tool Calling. Each Agent-facing Tool input SHALL be a typed Envelope containing optional `previous_observation` and required business `input`. A Tool interceptor SHALL validate and consume the previous observation, propagate the exact framework Tool Call ID, Run ID, Tool name, and unwrapped business JSON into the corresponding adapter and `ToolBoundary`, and SHALL enforce Run progress before execution. Successful observations SHALL use a bounded per-Tool whitelist projection; failed or control observations SHALL contain only stable safe semantics and SHALL NOT contain raw responses, internal exceptions, credentials, invocation lifecycle internals, counters, thresholds, or remaining budget.
|
||||
|
||||
#### Scenario: Framework requests RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`
|
||||
- **THEN** the RAG adapter and Agent observation use exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
#### Scenario: Framework requests first RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`, no pending evaluation and a typed business input
|
||||
- **THEN** the RAG adapter receives only the unwrapped business request, uses exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
|
||||
#### Scenario: Model continues after a non-empty result
|
||||
- **WHEN** the last successful Tool result is pending semantic evaluation and the model requests another Tool
|
||||
- **THEN** the Envelope must identify that exact prior Tool Call and contain `GAINED` or `NO_GAIN` before the new business Tool can execute
|
||||
|
||||
#### Scenario: Model omits required progress field
|
||||
- **WHEN** a prior non-empty Tool observation is pending and the model requests another Tool without `previous_observation`
|
||||
- **THEN** the business Tool does not execute and the Agent receives a bounded repair observation explaining the missing `previous_observation` field and expected prior Tool Call ID
|
||||
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** a registered adapter returns an error result
|
||||
@@ -50,7 +55,6 @@ The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs
|
||||
#### Scenario: Unknown Tool is requested
|
||||
- **WHEN** a model requests a Tool outside the three registered definitions
|
||||
- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
|
||||
|
||||
### Requirement: Diagnosis output SHALL be a bounded DiagnosisDraft
|
||||
The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL return JSON that the internal use case strictly parses into the frozen record. The use case SHALL enforce configured UTF-8 limits for query, previous turn, total input and Draft output and account accepted input/output bytes against the Run capacity. It SHALL reject blank, fenced, prefixed, malformed, oversized or schema-incompatible output without repair or retry.
|
||||
|
||||
@@ -61,14 +65,20 @@ The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL retu
|
||||
#### Scenario: Model returns prose around JSON
|
||||
- **WHEN** the final response contains a Markdown fence or explanatory prefix around an otherwise valid object
|
||||
- **THEN** strict parsing fails and the Diagnosis Agent is not invoked a second time
|
||||
|
||||
### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
|
||||
The single Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. If current evidence cannot support a diagnosis, the Agent SHALL stop the current ReAct execution with `conclusion=null`, describe the actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist or fabricate a root cause.
|
||||
The single Chinese Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. The Agent SHALL NOT be required to find a root cause. If current evidence cannot support a diagnosis, no required context exists for a bounded Tool call, or available results are correct but do not advance any diagnosis hypothesis, the Agent SHALL stop with `conclusion=null`, describe actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist, fabricate a root cause, or make equivalent Tool calls merely to show activity.
|
||||
|
||||
#### Scenario: Tool finds no evidence
|
||||
- **WHEN** the only completed Tool observation has `evidence_status=NO_EVIDENCE`
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records the bounded negative observation and missing information, and makes no additional automatic retry
|
||||
- **WHEN** completed Tool observations have no diagnostic information gain
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records bounded checked scope and missing information, and does not make an equivalent retry
|
||||
|
||||
#### Scenario: No bounded Tool query is possible
|
||||
- **WHEN** the Query lacks required enterprise, time, service or error context
|
||||
- **THEN** the Agent may perform zero Tool calls and returns a no-conclusion Draft whose `limitations.missing_info` identifies the required context
|
||||
|
||||
#### Scenario: Harness requires stop
|
||||
- **WHEN** the Agent receives `STOP_REQUIRED`
|
||||
- **THEN** it emits a bounded final Draft without another Tool call
|
||||
### Requirement: Stage-four execution SHALL remain internal and auditable
|
||||
The new use case SHALL be callable only as an internal Java/test entry in this stage and SHALL NOT be wired into public Chat or AIOps controllers. It SHALL propagate the same sessionId/runId through RunnableConfig metadata, allow the existing AgentStep Hook to be injected, retain ToolBoundary canonical invocation audit, and leave final Run success/persistence ownership to later Guard/application stages.
|
||||
|
||||
@@ -79,3 +89,21 @@ The new use case SHALL be callable only as an internal Java/test entry in this s
|
||||
#### Scenario: Stage four is archived
|
||||
- **WHEN** focused Agent tests pass and the change is archived
|
||||
- **THEN** `ChatController`, public SSE behavior and the old multi-Agent implementation remain available for stages 5, 6A and 6B
|
||||
### Requirement: Diagnosis execution SHALL preserve controlled stop outcomes
|
||||
The internal Agent use case SHALL distinguish a valid Draft, controlled information saturation, and budget termination from an unclassified Agent failure. It SHALL return a bounded internal execution result containing optional Draft, ProgressSnapshot and stop reason, and SHALL NOT convert a recognized controlled stop into `DiagnosisAgentOutputException`.
|
||||
|
||||
#### Scenario: Model ignores STOP_REQUIRED
|
||||
- **WHEN** the framework surfaces the typed collection-stopped signal after the final completion opportunity
|
||||
- **THEN** the Agent use case returns no Draft with the controlled stop reason and the current ProgressSnapshot
|
||||
|
||||
#### Scenario: Unclassified framework failure
|
||||
- **WHEN** Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination
|
||||
- **THEN** execution still fails closed and no safe progress is fabricated
|
||||
|
||||
#### Scenario: Invalid final Draft after verified checks
|
||||
- **WHEN** the final model text is empty or violates the strict DiagnosisDraft contract after the current Run has completed READY canonical Tool checks
|
||||
- **THEN** the invalid text is discarded, the output failure carries only the bounded ProgressSnapshot and safe failure metadata, and no model repair or loose JSON extraction occurs
|
||||
|
||||
#### Scenario: Invalid final Draft without verified checks
|
||||
- **WHEN** the final model text violates the strict DiagnosisDraft contract before any publishable ProgressSnapshot exists
|
||||
- **THEN** execution remains failed and MUST NOT fabricate missing context, observed facts or a no-conclusion Draft
|
||||
|
||||
@@ -4,16 +4,23 @@
|
||||
TBD - created by archiving change single-react-evidence-semantic-guards. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Deterministic Draft and evidence validation
|
||||
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
|
||||
For a Draft with a non-null Conclusion, the Harness SHALL deterministically reject it unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs. For a Draft with `conclusion=null`, Release SHALL NOT require normal conclusion structure or invoke EvidenceRepair; any supplied Tool references SHALL still resolve to current-Run READY canonical invocations and SHALL obey positive/negative evidence semantics.
|
||||
|
||||
#### Scenario: Duplicate or missing Analysis ID
|
||||
- **WHEN** a Draft contains a blank or duplicate Analysis ID
|
||||
#### Scenario: Duplicate or missing Analysis ID in concluded Draft
|
||||
- **WHEN** a Draft with a Conclusion contains a blank or duplicate Analysis ID
|
||||
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
|
||||
|
||||
#### Scenario: Broken report reference
|
||||
#### Scenario: Broken report reference in concluded Draft
|
||||
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
|
||||
- **THEN** EvidenceGuard rejects the Draft before semantic review
|
||||
|
||||
#### Scenario: No-conclusion Draft has valid negative observation
|
||||
- **WHEN** a `conclusion=null` Draft cites a current-Run READY `NO_EVIDENCE` call as `NEGATIVE_OBSERVATION`
|
||||
- **THEN** Release accepts the reference authenticity without running EvidenceRepair or SemanticGuard
|
||||
|
||||
#### Scenario: No-conclusion Draft fabricates a Tool reference
|
||||
- **WHEN** a `conclusion=null` Draft cites a missing, cross-Run, incomplete or ERROR Tool call
|
||||
- **THEN** the reference is excluded and cannot be published as an observed fact
|
||||
### Requirement: Current Run canonical invocation ownership
|
||||
EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`.
|
||||
|
||||
@@ -24,7 +31,6 @@ EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallI
|
||||
#### Scenario: Failed or incomplete invocation
|
||||
- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR`
|
||||
- **THEN** EvidenceGuard rejects the Draft
|
||||
|
||||
### Requirement: Analysis kind matches evidence semantics
|
||||
EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls.
|
||||
|
||||
@@ -35,7 +41,6 @@ EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls an
|
||||
#### Scenario: Positive claim uses no-evidence result
|
||||
- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation
|
||||
- **THEN** EvidenceGuard rejects the binding
|
||||
|
||||
### Requirement: Verified evidence snapshot is minimal and deterministic
|
||||
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
|
||||
|
||||
@@ -46,7 +51,6 @@ The Harness SHALL strictly parse only supported Tool projections and SHALL const
|
||||
#### Scenario: Projection contract mismatch
|
||||
- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
|
||||
- **THEN** EvidenceGuard fails closed
|
||||
|
||||
### Requirement: Evidence repair is single-turn and semantics-preserving
|
||||
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
|
||||
|
||||
@@ -61,7 +65,6 @@ On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct n
|
||||
#### Scenario: Second validation fails
|
||||
- **WHEN** the repaired Draft still fails EvidenceGuard
|
||||
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard
|
||||
|
||||
### Requirement: Isolated single-turn SemanticGuard
|
||||
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
|
||||
|
||||
@@ -72,7 +75,6 @@ SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Promp
|
||||
#### Scenario: Binary review output
|
||||
- **WHEN** SemanticGuard completes normally
|
||||
- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report
|
||||
|
||||
### Requirement: Semantic model budgets timeout cancellation and retry
|
||||
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
|
||||
|
||||
@@ -87,12 +89,11 @@ The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets
|
||||
#### Scenario: Run cancellation during model call
|
||||
- **WHEN** the Run is cancelled while a guard model call is pending
|
||||
- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
|
||||
|
||||
### Requirement: Fail-closed release policy
|
||||
The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
|
||||
The release use case SHALL publish the unchanged verified Draft only when a non-null Conclusion passes EvidenceGuard and SemanticGuard returns `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, final semantic technical failure, a valid no-conclusion Draft, information saturation, or handled budget termination. No-conclusion and controlled-stop release SHALL be deterministic from verified references and ProgressSnapshot and MUST NOT invoke a repair or semantic model call. A fallback result MUST NOT include an unsupported Draft, full verified snapshot, internal stop counters, or SemanticGuard reason.
|
||||
|
||||
#### Scenario: Supported report release
|
||||
- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED`
|
||||
- **WHEN** EvidenceGuard succeeds for a concluded Draft and SemanticGuard returns `SUPPORTED`
|
||||
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
|
||||
|
||||
#### Scenario: Unsupported report fallback
|
||||
@@ -104,9 +105,20 @@ The release use case SHALL publish the unchanged verified Draft only for `SUPPOR
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
|
||||
|
||||
#### Scenario: Evidence validation fallback sources
|
||||
- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft
|
||||
- **WHEN** evidence repair fails or the second EvidenceGuard rejects a concluded Draft
|
||||
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
|
||||
|
||||
#### Scenario: Missing context ends without Tool calls
|
||||
- **WHEN** a valid no-conclusion Draft has no Tool calls and identifies required missing context
|
||||
- **THEN** release outcome is `FALLBACK`, type is `MISSING_REQUIRED_CONTEXT`, and no guard model call occurs
|
||||
|
||||
#### Scenario: Finite checks do not support a conclusion
|
||||
- **WHEN** a no-conclusion Draft or controlled stop has a non-empty verified ProgressSnapshot
|
||||
- **THEN** release outcome is `FALLBACK`, type is `INSUFFICIENT_EVIDENCE`, and observed facts describe only actual completed checks
|
||||
|
||||
#### Scenario: Invalid Draft has publishable progress
|
||||
- **WHEN** the Agent's final Draft is rejected by strict parsing but its bounded ProgressSnapshot contains current-Run verified observed facts
|
||||
- **THEN** Release publishes `FALLBACK` with type `INSUFFICIENT_EVIDENCE` using only that snapshot and MUST NOT use any content from the invalid Draft
|
||||
### Requirement: Stage-five public isolation
|
||||
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user