feat(harness): add information gain stop and audit
This commit is contained in:
+25
@@ -0,0 +1,25 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: RAG Tool contract SHALL expose only bounded document evidence
|
||||
The RAG business Request SHALL contain only `query`. The canonical RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, optional normalized `relevance_level`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`. The RAG projector SHALL accept upstream `relevanceLevel` or `relevance_level` and normalize recognized values without exposing raw relevance scores or retrieval traces.
|
||||
|
||||
#### Scenario: RAG evidence is serialized
|
||||
- **WHEN** a RAG result contains a matching document excerpt and an upstream relevance level
|
||||
- **THEN** its canonical JSON preserves the bounded evidence and normalized `relevance_level` while excluding ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
|
||||
|
||||
#### Scenario: RAG query has no evidence
|
||||
- **WHEN** RAG executes successfully without a usable document excerpt
|
||||
- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, returns an empty evidence list, and does not upgrade relevance into evidence
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Tool results SHALL have separate Harness and model views
|
||||
Each successful evidence Tool result SHALL provide a Harness Control View and a bounded Model Observation derived from the same canonical result. The control view MAY contain returned counts, normalized scope, relevance, truncation and duplicate identity. The Model Observation SHALL contain only fields needed to understand and cite the result and SHALL NOT contain raw responses, internal scores, retrieval traces, duplicate fingerprints, counters, thresholds, budgets or store identities.
|
||||
|
||||
#### Scenario: Model receives RAG observation
|
||||
- **WHEN** a canonical RAG result is READY
|
||||
- **THEN** the model receives Tool Call ID, actual query scope, bounded evidence, evidence status, optional coarse relevance and truncation, but not raw scores or Harness counters
|
||||
|
||||
#### Scenario: Harness evaluates duplicate scope
|
||||
- **WHEN** the same normalized Tool scope is requested again
|
||||
- **THEN** the Harness can compare its control view identity without exposing that fingerprint to the model
|
||||
+12
@@ -0,0 +1,12 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Run progress projection SHALL reference canonical records without duplicating truth
|
||||
The Run progress tracker SHALL retain only ordered canonical identities for completed Tool calls. At Tool-loop completion, a projector SHALL resolve those identities through the existing canonical store and SHALL accept only READY records owned by the current Run. The tracker SHALL NOT store or reconstruct raw Tool responses, complete Agent results, or a second durable evidence record.
|
||||
|
||||
#### Scenario: Completed calls are projected
|
||||
- **WHEN** a Run ends after multiple READY canonical Tool invocations
|
||||
- **THEN** the progress projector reads each indexed canonical record in execution order and creates bounded observed facts
|
||||
|
||||
#### Scenario: Indexed identity is invalid
|
||||
- **WHEN** an indexed canonical identity is missing, expired, cross-Run, PROJECTING, or ERROR
|
||||
- **THEN** it is excluded from observed facts and cannot become verified evidence
|
||||
+110
@@ -0,0 +1,110 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Harness SHALL track binary information gain per Run
|
||||
Each Diagnosis Run SHALL own a thread-safe progress tracker with `GAINED` and `NO_GAIN` as the only information-gain values. `GAINED` SHALL reset the consecutive no-gain count; `NO_GAIN` SHALL increment it. The tracker SHALL expose only `COLLECTING` or `SATURATED` as collection state and SHALL NOT create a second diagnosis lifecycle.
|
||||
|
||||
#### Scenario: New evidence advances diagnosis
|
||||
- **WHEN** a valid pending Tool observation is evaluated as `GAINED`
|
||||
- **THEN** the Run remains `COLLECTING` and its consecutive no-gain count becomes zero
|
||||
|
||||
#### Scenario: Consecutive observations do not advance diagnosis
|
||||
- **WHEN** valid `NO_GAIN` observations reach the Run's configured threshold
|
||||
- **THEN** the tracker becomes `SATURATED` with stop reason `INFORMATION_SATURATED`
|
||||
|
||||
### Requirement: Harness SHALL assign only deterministic no-gain signals
|
||||
The Harness SHALL assign `NO_GAIN` when a successful Tool result has `evidence_status=NO_EVIDENCE` or when a requested `tool_name + normalized_scope` duplicates a successfully completed scope in the same Run. Other successful non-empty results, including RAG `REFERENCE`, SHALL require a model-provided `GAINED` or `NO_GAIN` before another Tool executes.
|
||||
|
||||
#### Scenario: Empty scoped result
|
||||
- **WHEN** a Tool completes READY with `NO_EVIDENCE`
|
||||
- **THEN** the Harness records `NO_GAIN` without asking a model to judge Tool quality
|
||||
|
||||
#### Scenario: Reference material is non-empty
|
||||
- **WHEN** RAG returns bounded evidence with `relevance_level=REFERENCE`
|
||||
- **THEN** the Harness leaves it pending for model evaluation and does not automatically mark it `NO_GAIN`
|
||||
|
||||
#### Scenario: Equivalent structured scope repeats
|
||||
- **WHEN** the model requests the same Tool with the same deterministically normalized successful scope
|
||||
- **THEN** the Harness does not execute the Tool, records `NO_GAIN`, and does not create a second canonical invocation
|
||||
|
||||
### Requirement: Consecutive no-gain threshold SHALL be fixed per Run
|
||||
The system SHALL bind `harness.chat.stop-after-consecutive-no-gain`, require a value of at least one, and default it to `2`. The value SHALL be copied into each new Run's tracker and SHALL NOT enter model context or change an active Run.
|
||||
|
||||
#### Scenario: Default configuration is used
|
||||
- **WHEN** no external value is configured
|
||||
- **THEN** a new Run becomes saturated after two consecutive `NO_GAIN` decisions
|
||||
|
||||
#### Scenario: Invalid threshold is configured
|
||||
- **WHEN** the configured threshold is zero or negative
|
||||
- **THEN** Harness configuration validation fails before serving Chat requests
|
||||
|
||||
### Requirement: Saturated collection SHALL stop further Tool execution
|
||||
When collection becomes saturated, the Harness SHALL reject the pending or next business Tool execution and deliver one bounded `STOP_REQUIRED` observation with `reason=INFORMATION_SATURATED`. If the next model round requests another Tool, the Agent execution SHALL terminate through a typed controlled-stop path without consuming another Tool budget or publishing an internal failure.
|
||||
|
||||
#### Scenario: Model evaluation reaches threshold
|
||||
- **WHEN** the next Tool Call reports `NO_GAIN` and that evaluation reaches the threshold
|
||||
- **THEN** the requested business Tool is not executed and the model receives one STOP_REQUIRED observation
|
||||
|
||||
#### Scenario: Model ignores stop instruction
|
||||
- **WHEN** the model requests another Tool after STOP_REQUIRED was delivered
|
||||
- **THEN** the loop ends as controlled information saturation and proceeds to safe release
|
||||
|
||||
### Requirement: Tool loop completion SHALL project bounded progress once
|
||||
At normal Draft completion, information saturation, or budget termination, the Harness SHALL use the current Run's completed canonical invocation keys to create one bounded `ProgressSnapshot`. The snapshot SHALL contain only safe source, actual scope, objective result summary, truncation and stop reason; it SHALL NOT contain Prompt, thought, raw Tool response, internal counters, remaining budget, Redis keys, or model evaluation rationale.
|
||||
|
||||
#### Scenario: Unknown problem has completed checks
|
||||
- **WHEN** multiple Tools completed but no supported conclusion exists
|
||||
- **THEN** the release input contains their bounded checked scopes and objective results in stable execution order
|
||||
|
||||
#### Scenario: Canonical record cannot be verified
|
||||
- **WHEN** an indexed record is missing, expired, incomplete, ERROR, or not owned by the current Run
|
||||
- **THEN** it is not projected as an observed fact and the snapshot records a bounded limitation
|
||||
|
||||
### Requirement: Information stop reasons SHALL remain distinct from release outcomes
|
||||
The Harness SHALL distinguish `INFORMATION_SATURATED` from `BUDGET_LIMIT_REACHED`. Release SHALL continue to expose only `SUCCESS`, `FALLBACK`, `FAILED`, or `CANCELLED`, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
|
||||
|
||||
#### Scenario: Low gain stops before budget exhaustion
|
||||
- **WHEN** consecutive no-gain reaches the configured threshold while hard budget remains
|
||||
- **THEN** Trace records `INFORMATION_SATURATED` and public release is a normal `FALLBACK`
|
||||
|
||||
#### Scenario: Hard budget terminates collection
|
||||
- **WHEN** model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
|
||||
- **THEN** Trace retains `BUDGET_LIMIT_REACHED` and Release attempts a deterministic `FALLBACK` from the existing progress without an extra model call
|
||||
|
||||
### Requirement: Diagnosis Prompt SHALL license bounded abandonment
|
||||
The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, `conclusion=null` is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is `NO_GAIN`. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, `next_action`, or Harness implementation details.
|
||||
|
||||
#### Scenario: Required context is missing before any Tool call
|
||||
- **WHEN** the Query lacks the enterprise, time, service, error, or other context needed for a bounded query
|
||||
- **THEN** the model may return `conclusion=null` with `limitations.missing_info` without calling a Tool
|
||||
|
||||
#### Scenario: Correct content is diagnostically useless
|
||||
- **WHEN** a Tool response is factually correct but only generic, repeated, or unable to change a current hypothesis
|
||||
- **THEN** the model treats it as `NO_GAIN` and does not continue with an equivalent query
|
||||
|
||||
### Requirement: Model token audit SHALL be component-scoped and reconcilable
|
||||
Every Harness model call admitted by the Run budget SHALL receive a bounded component and component round. When Provider Usage is available, the same non-negative input and output Token counts SHALL update both the Run budget total and a `MODEL_TOKEN_USAGE` Trace event. Diagnosis Agent usage SHALL also update the matching `AgentStep.token_count`. At Run completion, Trace SHALL expose whether audited Token totals reconcile with the Run budget total and SHALL expose unavailable Usage counts without fabricating Token values.
|
||||
|
||||
The audit SHALL NOT persist Prompt content, user or model text, reasoning content, Tool arguments, raw model responses, credentials, or provider-specific metadata.
|
||||
|
||||
#### Scenario: Diagnosis Agent round returns Usage
|
||||
- **WHEN** a Diagnosis Agent model round returns input and output Token Usage
|
||||
- **THEN** its component round Trace and matching AgentStep contain the same total Token count and the Run total increases by that amount
|
||||
|
||||
#### Scenario: Multiple model components execute
|
||||
- **WHEN** Router, Diagnosis Agent, Evidence Repair or Semantic Guard model calls execute in one Run
|
||||
- **THEN** each call is distinguishable by bounded component and component round and their audited Token sum can be compared with the Run total
|
||||
|
||||
#### Scenario: Provider Usage is unavailable
|
||||
- **WHEN** a model attempt completes or fails without Provider Usage
|
||||
- **THEN** the audit marks Usage unavailable and Run reconciliation exposes the gap without estimating Token counts
|
||||
|
||||
### Requirement: Rejected Tool requests SHALL remain observable without payload disclosure
|
||||
Every supported Tool request rejected by the Harness before a usable business observation is delivered SHALL emit a `TOOL_REQUEST_REJECTED` Trace event containing only safe Tool Call ID, Tool name and stable error code. Rejected requests SHALL remain distinguishable from canonical `TOOL_INVOCATION` events and SHALL NOT include Tool arguments, normalized scope content, raw responses, internal exception messages, credentials or budget values.
|
||||
|
||||
#### Scenario: Progress protocol is invalid
|
||||
- **WHEN** a supported Tool request omits or misorders a required previous observation
|
||||
- **THEN** no business Tool executes and Trace records `INVALID_PROGRESS_PROTOCOL` for that Tool request
|
||||
|
||||
#### Scenario: Duplicate or saturated request is blocked
|
||||
- **WHEN** a supported Tool request repeats a successful normalized scope or arrives after collection saturation
|
||||
- **THEN** Trace records the stable rejection reason while canonical invocation count remains unchanged
|
||||
+33
@@ -0,0 +1,33 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Fixed isolated executors
|
||||
The Application Use Case SHALL map SYSTEM_CHAT to one no-Tool model response, KNOWLEDGE_QUERY to exactly one lookup-knowledge invocation plus one bounded answer model call, and DIAGNOSIS to the single Diagnosis Agent followed by the Diagnosis Release boundary. Executors MUST NOT call one another or rewrite the Query. Information saturation and budget termination in Diagnosis SHALL be converted to safe content by Diagnosis Release, not by ChatApplicationUseCase.
|
||||
|
||||
#### Scenario: System Chat
|
||||
- **WHEN** intent is SYSTEM_CHAT
|
||||
- **THEN** no evidence Tool, Diagnosis Agent, EvidenceGuard, or SemanticGuard is invoked
|
||||
|
||||
#### Scenario: Knowledge Query
|
||||
- **WHEN** intent is KNOWLEDGE_QUERY
|
||||
- **THEN** only lookup_knowledge is invoked once and query_logs/query_mysql/Diagnosis ReAct are unavailable
|
||||
|
||||
#### Scenario: Diagnosis succeeds with a conclusion
|
||||
- **WHEN** intent is DIAGNOSIS and the Draft passes the release guards
|
||||
- **THEN** the original Query and bounded PreviousTurn enter DiagnosisAgentUseCase and the Draft publishes only after DiagnosisReleaseUseCase
|
||||
|
||||
#### Scenario: Diagnosis stops without a conclusion
|
||||
- **WHEN** Diagnosis collection is saturated, required context is missing, or a handled budget limit is reached
|
||||
- **THEN** DiagnosisReleaseUseCase returns bounded Fallback content and ChatApplicationUseCase only persists and transports that decision
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Handled Diagnosis budget termination SHALL remain a Fallback release
|
||||
When Diagnosis Release has converted a recognized budget termination and existing safe progress into a Fallback, ChatApplicationUseCase SHALL persist that result exactly once without reclassifying it as `INTERNAL_FAILURE`, invoking another model, or rebuilding business fallback content. The internal Run lifecycle MAY retain `BUDGET_EXHAUSTED`, while the persisted public release outcome SHALL be `FALLBACK` and no PublishedResult SHALL be stored.
|
||||
|
||||
#### Scenario: Tool budget ends after finite checks
|
||||
- **WHEN** Diagnosis reaches a hard Tool budget after at least one canonical safe observation and Release creates an insufficient-evidence fallback
|
||||
- **THEN** the Run persists status SUCCESS, release outcome FALLBACK, safe content and actual budget usage, and the SSE sends content followed by done
|
||||
|
||||
#### Scenario: Budget ends without safe publishable progress
|
||||
- **WHEN** budget termination occurs before Diagnosis Release can form a safe bounded result
|
||||
- **THEN** the existing failure path remains fail closed and does not fabricate observed facts
|
||||
+56
@@ -0,0 +1,56 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Evidence Tools SHALL execute through the Harness boundary
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions through native Tool Calling. Each Agent-facing Tool input SHALL be a typed Envelope containing optional `previous_observation` and required business `input`. A Tool interceptor SHALL validate and consume the previous observation, propagate the exact framework Tool Call ID, Run ID, Tool name, and unwrapped business JSON into the corresponding adapter and `ToolBoundary`, and SHALL enforce Run progress before execution. Successful observations SHALL use a bounded per-Tool whitelist projection; failed or control observations SHALL contain only stable safe semantics and SHALL NOT contain raw responses, internal exceptions, credentials, invocation lifecycle internals, counters, thresholds, or remaining budget.
|
||||
|
||||
#### Scenario: Framework requests first RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`, no pending evaluation and a typed business input
|
||||
- **THEN** the RAG adapter receives only the unwrapped business request, uses exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
|
||||
#### Scenario: Model continues after a non-empty result
|
||||
- **WHEN** the last successful Tool result is pending semantic evaluation and the model requests another Tool
|
||||
- **THEN** the Envelope must identify that exact prior Tool Call and contain `GAINED` or `NO_GAIN` before the new business Tool can execute
|
||||
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** a registered adapter returns an error result
|
||||
- **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail
|
||||
|
||||
#### Scenario: Unknown Tool is requested
|
||||
- **WHEN** a model requests a Tool outside the three registered definitions
|
||||
- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
|
||||
|
||||
### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
|
||||
The single Chinese Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. The Agent SHALL NOT be required to find a root cause. If current evidence cannot support a diagnosis, no required context exists for a bounded Tool call, or available results are correct but do not advance any diagnosis hypothesis, the Agent SHALL stop with `conclusion=null`, describe actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist, fabricate a root cause, or make equivalent Tool calls merely to show activity.
|
||||
|
||||
#### Scenario: Tool finds no evidence
|
||||
- **WHEN** completed Tool observations have no diagnostic information gain
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records bounded checked scope and missing information, and does not make an equivalent retry
|
||||
|
||||
#### Scenario: No bounded Tool query is possible
|
||||
- **WHEN** the Query lacks required enterprise, time, service or error context
|
||||
- **THEN** the Agent may perform zero Tool calls and returns a no-conclusion Draft whose `limitations.missing_info` identifies the required context
|
||||
|
||||
#### Scenario: Harness requires stop
|
||||
- **WHEN** the Agent receives `STOP_REQUIRED`
|
||||
- **THEN** it emits a bounded final Draft without another Tool call
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Diagnosis execution SHALL preserve controlled stop outcomes
|
||||
The internal Agent use case SHALL distinguish a valid Draft, controlled information saturation, and budget termination from an unclassified Agent failure. It SHALL return a bounded internal execution result containing optional Draft, ProgressSnapshot and stop reason, and SHALL NOT convert a recognized controlled stop into `DiagnosisAgentOutputException`.
|
||||
|
||||
#### Scenario: Model ignores STOP_REQUIRED
|
||||
- **WHEN** the framework surfaces the typed collection-stopped signal after the final completion opportunity
|
||||
- **THEN** the Agent use case returns no Draft with `INFORMATION_SATURATED` and the current ProgressSnapshot
|
||||
|
||||
#### Scenario: Unclassified framework failure
|
||||
- **WHEN** Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination
|
||||
- **THEN** execution still fails closed and no safe progress is fabricated
|
||||
|
||||
#### Scenario: Invalid final Draft after verified checks
|
||||
- **WHEN** the final model text is empty or violates the strict DiagnosisDraft contract after the current Run has completed READY canonical Tool checks
|
||||
- **THEN** the invalid text is discarded, the output failure carries only the bounded ProgressSnapshot and safe failure metadata, and no model repair or loose JSON extraction occurs
|
||||
|
||||
#### Scenario: Invalid final Draft without verified checks
|
||||
- **WHEN** the final model text violates the strict DiagnosisDraft contract before any publishable ProgressSnapshot exists
|
||||
- **THEN** execution remains failed and MUST NOT fabricate missing context, observed facts or a no-conclusion Draft
|
||||
+51
@@ -0,0 +1,51 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Deterministic Draft and evidence validation
|
||||
For a Draft with a non-null Conclusion, the Harness SHALL deterministically reject it unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs. For a Draft with `conclusion=null`, Release SHALL NOT require normal conclusion structure or invoke EvidenceRepair; any supplied Tool references SHALL still resolve to current-Run READY canonical invocations and SHALL obey positive/negative evidence semantics.
|
||||
|
||||
#### Scenario: Duplicate or missing Analysis ID in concluded Draft
|
||||
- **WHEN** a Draft with a Conclusion contains a blank or duplicate Analysis ID
|
||||
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
|
||||
|
||||
#### Scenario: Broken report reference in concluded Draft
|
||||
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
|
||||
- **THEN** EvidenceGuard rejects the Draft before semantic review
|
||||
|
||||
#### Scenario: No-conclusion Draft has valid negative observation
|
||||
- **WHEN** a `conclusion=null` Draft cites a current-Run READY `NO_EVIDENCE` call as `NEGATIVE_OBSERVATION`
|
||||
- **THEN** Release accepts the reference authenticity without running EvidenceRepair or SemanticGuard
|
||||
|
||||
#### Scenario: No-conclusion Draft fabricates a Tool reference
|
||||
- **WHEN** a `conclusion=null` Draft cites a missing, cross-Run, incomplete or ERROR Tool call
|
||||
- **THEN** the reference is excluded and cannot be published as an observed fact
|
||||
|
||||
### Requirement: Fail-closed release policy
|
||||
The release use case SHALL publish the unchanged verified Draft only when a non-null Conclusion passes EvidenceGuard and SemanticGuard returns `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, final semantic technical failure, a valid no-conclusion Draft, information saturation, or handled budget termination. No-conclusion and controlled-stop release SHALL be deterministic from verified references and ProgressSnapshot and MUST NOT invoke a repair or semantic model call. A fallback result MUST NOT include an unsupported Draft, full verified snapshot, internal stop counters, or SemanticGuard reason.
|
||||
|
||||
#### Scenario: Supported report release
|
||||
- **WHEN** EvidenceGuard succeeds for a concluded Draft and SemanticGuard returns `SUPPORTED`
|
||||
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
|
||||
|
||||
#### Scenario: Unsupported report fallback
|
||||
- **WHEN** SemanticGuard returns `UNSUPPORTED`
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot
|
||||
|
||||
#### Scenario: Semantic review remains unavailable
|
||||
- **WHEN** all permitted technical attempts fail
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
|
||||
|
||||
#### Scenario: Evidence validation fallback sources
|
||||
- **WHEN** evidence repair fails or the second EvidenceGuard rejects a concluded Draft
|
||||
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
|
||||
|
||||
#### Scenario: Missing context ends without Tool calls
|
||||
- **WHEN** a valid no-conclusion Draft has no Tool calls and identifies required missing context
|
||||
- **THEN** release outcome is `FALLBACK`, type is `MISSING_REQUIRED_CONTEXT`, and no guard model call occurs
|
||||
|
||||
#### Scenario: Finite checks do not support a conclusion
|
||||
- **WHEN** a no-conclusion Draft or controlled stop has a non-empty verified ProgressSnapshot
|
||||
- **THEN** release outcome is `FALLBACK`, type is `INSUFFICIENT_EVIDENCE`, and observed facts describe only actual completed checks
|
||||
|
||||
#### Scenario: Invalid Draft has publishable progress
|
||||
- **WHEN** the Agent's final Draft is rejected by strict parsing but its bounded ProgressSnapshot contains current-Run verified observed facts
|
||||
- **THEN** Release publishes `FALLBACK` with type `INSUFFICIENT_EVIDENCE` using only that snapshot and MUST NOT use any content from the invalid Draft
|
||||
Reference in New Issue
Block a user