feat(harness): complete protocol repair stop and archive ISS-016
Add repairable INVALID_PROGRESS_PROTOCOL observations, independent PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths. Archive the OpenSpec change after syncing main specs and devflow.
This commit is contained in:
+10
@@ -40,6 +40,8 @@ Agent-facing Schema 使用三个具体输入类型:`RagToolCall`、`QueryLogsT
|
||||
|
||||
首个 Tool Call 以及上一结果已由 Harness 确定性评价时允许省略 `previous_observation`。存在待评价调用时,新 Tool Call 必须携带完全匹配的上一 Tool Call ID 和二值评价;缺失、错序或跨 Run 引用返回有界协议错误且不执行 Tool。
|
||||
|
||||
协议错误返回可修正 Model Observation,而不是只返回稳定错误码。Observation 可包含 `repair_required=true`、`violation_type`、`missing_field`、`expected_previous_tool_call_id` 和允许的 `information_gain` 值;这些字段均来自 Harness 控制状态和安全 Tool Call ID,不包含业务参数、scope、预算、Prompt、raw response 或内部异常。
|
||||
|
||||
### 3. 先应用上一轮评价,再决定本轮是否放行
|
||||
|
||||
interceptor 的固定顺序为:
|
||||
@@ -78,6 +80,12 @@ tracker 记录控制指令已交付。下一轮模型仍发起 Tool Call时,in
|
||||
|
||||
替代方案是无限返回 STOP_REQUIRED。拒绝,因为模型无视指令时仍会消耗模型预算并重现原问题。
|
||||
|
||||
### 6.1 连续协议错误也进入受控停止
|
||||
|
||||
`INVALID_PROGRESS_PROTOCOL` 表示模型没有遵守 Tool Call Envelope 进展协议,不表示 Tool 内容无信息增益。因此它不累计 `NO_GAIN`,而是单独维护连续协议错误计数。首次或未达阈值的协议错误返回可修正 observation,给模型一次按字段修复的机会;达到阈值后 tracker 进入 `SATURATED`,`stop_reason=PROGRESS_PROTOCOL_VIOLATED`,当前请求收到一次 `STOP_REQUIRED`。
|
||||
|
||||
如果模型在 `STOP_REQUIRED` 后继续请求 Tool,沿用既有 `DiagnosisCollectionStoppedException` 受控停止路径。Release 仅在 ProgressSnapshot 有当前 Run 已验真 observed facts 时发布 `INSUFFICIENT_EVIDENCE`;没有安全进展时继续 fail closed,不伪造用户可见事实。
|
||||
|
||||
### 7. Agent 执行返回 Draft 与停止投影,而不是只返回 Draft
|
||||
|
||||
`DiagnosisAgentUseCase` 返回内部 `DiagnosisAgentExecution`:可选 Draft、`ProgressSnapshot` 和可选 `DiagnosisStopReason`。正常 Draft、强制饱和停止和预算异常都在 Diagnosis executor 边界形成 Release 输入。
|
||||
@@ -109,6 +117,8 @@ Prompt 不写 Tool 名、Schema、阈值、计数器、Projector、预算或 `ne
|
||||
|
||||
Trace 增加有界事件或字段,记录 Tool Call ID、Tool name、scope 摘要、information gain 的生产者(Harness/Model)、连续计数变化、collection state 和 stop reason。不得记录 Prompt、模型 thought、raw Tool response、完整 SQL 参数、预算余量或模型评价理由。
|
||||
|
||||
协议拒绝 Trace 在稳定 `error_code` 之外记录安全字段:`violation_type`、`repair_prompt_delivered`、`consecutive_protocol_violations` 和达到阈值时的 `stop_reason`。Trace 不记录缺失字段对应的业务值、上一轮观察正文、Tool 参数或异常文本。
|
||||
|
||||
### 11. Run 总账与模型调用明细使用同一份 Provider Usage
|
||||
|
||||
`RunContext` 增加最小线程安全模型调用账本,只维护组件轮次、已审计调用数、Usage 不可用调用数和 Token 合计。每次实际模型调用在预算放行后取得组件轮次;`HarnessModelInterceptor` 负责 Diagnosis Agent,`GuardModelCall` 负责 Router、System Chat、Knowledge Answer、Evidence Repair 和 Semantic Guard。两条入口都从 Spring AI `Usage` 读取同一组 input/output Token,先登记调用明细,再交给现有 `RunBudget` 累加总账。
|
||||
+3
@@ -18,6 +18,7 @@ Diagnosis Agent 面对知识库未知、日志为空或查询条件不足的问
|
||||
- 精简中文 Diagnosis Prompt,明确模型不必须得出根因、允许零次 Tool 调用、正确但对当前推导无用的内容属于 `NO_GAIN`,且合法放弃是成功完成。
|
||||
- 增加最小模型调用审计:按组件和轮次记录 input/output/total Token,回填 Diagnosis AgentStep Token,并在 Run 结束时与预算总账对账;不记录 Prompt、模型正文或推理内容。
|
||||
- 对进入 Harness 后被协议、饱和或重复 scope 门禁拒绝的 Tool 请求记录安全 Trace,使预算 Tool 计数、实际执行和拒绝决策可区分;不记录 Tool 参数或原始响应。
|
||||
- 对 `INVALID_PROGRESS_PROTOCOL` 增加可修正的模型观察:明确缺失/错序的协议字段、期望上一轮 Tool Call ID 和允许的 `information_gain` 值;连续协议错误达到阈值后转为受控停止,避免模型反复空转。
|
||||
|
||||
## Capabilities
|
||||
|
||||
@@ -58,6 +59,7 @@ Diagnosis Agent 面对知识库未知、日志为空或查询条件不足的问
|
||||
- `CanonicalInvocationStore` 是当前 Run 完整 Tool 调用真相源;Run 内 tracker 只维护控制状态和完成调用索引,不复制 raw response。
|
||||
- `RunContext` 继续使用“结构不可变 + 线程安全可变 handle”的既有模式;阈值在 Run 启动时固定。
|
||||
- 现有 `SafeFallback.observed_facts / verified_sources / limitations / next_steps` 和前端渲染能力必须复用。
|
||||
- 协议错误不等同于 Tool 内容 `NO_GAIN`;它使用独立 `PROGRESS_PROTOCOL_VIOLATED` stop reason,并且只在有可发布 ProgressSnapshot 时进入安全 Fallback。
|
||||
- 当前未提交的预算 Fallback 修改保留用户价值,但业务 Release 决策需要从 `ChatApplicationUseCase` 迁移到 `DiagnosisReleaseUseCase`。
|
||||
- 当前环境缺少 `codebase-retrieval` 和 LSP;本次影响核查使用 `rg` 引用搜索、源码和测试阅读降级完成。用户已明确允许忽略 GitNexus。
|
||||
|
||||
@@ -73,6 +75,7 @@ Diagnosis Agent 面对知识库未知、日志为空或查询条件不足的问
|
||||
|
||||
- 框架 Tool Schema 生成或 ToolInterceptor 参数处理不符合 Envelope 假设,导致模型无法正确回传或业务 request 未被剥离。
|
||||
- STOP_REQUIRED 若未形成受控结束,模型可能继续请求 Tool,或 Agent 调用异常绕过统一 Release。
|
||||
- 仅返回通用 `INVALID_PROGRESS_PROTOCOL` 可能不足以让模型自修复;必须返回字段级 repair hint,同时设置连续错误兜底。
|
||||
- 最后一轮结果可能没有下一次 Tool Call 来回传模型评价;该情况只能表示模型主动结束,不能伪造 `GAINED / NO_GAIN`。
|
||||
- ProgressSnapshot 若读取不完整或混入 raw payload,会造成过程缺失或上下文/隐私边界倒退。
|
||||
- `conclusion=null` 与现有 EvidenceGuard 的 `ANALYSIS_MISSING` 规则冲突,需要明确区分“验证引用真实性”和“验证结论完整性”。
|
||||
+23
-2
@@ -60,7 +60,7 @@ At normal Draft completion, information saturation, or budget termination, the H
|
||||
- **THEN** it is not projected as an observed fact and the snapshot records a bounded limitation
|
||||
|
||||
### Requirement: Information stop reasons SHALL remain distinct from release outcomes
|
||||
The Harness SHALL distinguish `INFORMATION_SATURATED` from `BUDGET_LIMIT_REACHED`. Release SHALL continue to expose only `SUCCESS`, `FALLBACK`, `FAILED`, or `CANCELLED`, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
|
||||
The Harness SHALL distinguish `INFORMATION_SATURATED`, `BUDGET_LIMIT_REACHED`, and `PROGRESS_PROTOCOL_VIOLATED`. Release SHALL continue to expose only `SUCCESS`, `FALLBACK`, `FAILED`, or `CANCELLED`, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
|
||||
|
||||
#### Scenario: Low gain stops before budget exhaustion
|
||||
- **WHEN** consecutive no-gain reaches the configured threshold while hard budget remains
|
||||
@@ -70,6 +70,10 @@ The Harness SHALL distinguish `INFORMATION_SATURATED` from `BUDGET_LIMIT_REACHED
|
||||
- **WHEN** model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
|
||||
- **THEN** Trace retains `BUDGET_LIMIT_REACHED` and Release attempts a deterministic `FALLBACK` from the existing progress without an extra model call
|
||||
|
||||
#### Scenario: Progress protocol violations exceed threshold
|
||||
- **WHEN** consecutive invalid progress protocol requests reach the configured threshold
|
||||
- **THEN** Trace records `PROGRESS_PROTOCOL_VIOLATED`, the model receives one STOP_REQUIRED observation, and public release remains a normal Fallback only if safe progress exists
|
||||
|
||||
### Requirement: Diagnosis Prompt SHALL license bounded abandonment
|
||||
The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, `conclusion=null` is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is `NO_GAIN`. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, `next_action`, or Harness implementation details.
|
||||
|
||||
@@ -103,8 +107,25 @@ Every supported Tool request rejected by the Harness before a usable business ob
|
||||
|
||||
#### Scenario: Progress protocol is invalid
|
||||
- **WHEN** a supported Tool request omits or misorders a required previous observation
|
||||
- **THEN** no business Tool executes and Trace records `INVALID_PROGRESS_PROTOCOL` for that Tool request
|
||||
- **THEN** no business Tool executes, Trace records `INVALID_PROGRESS_PROTOCOL` with a safe violation type, and the model receives a repairable observation naming the missing or expected protocol field
|
||||
|
||||
#### Scenario: Duplicate or saturated request is blocked
|
||||
- **WHEN** a supported Tool request repeats a successful normalized scope or arrives after collection saturation
|
||||
- **THEN** Trace records the stable rejection reason while canonical invocation count remains unchanged
|
||||
|
||||
### Requirement: Invalid progress protocol SHALL be repairable before bounded stop
|
||||
When a supported Tool request violates the Tool Envelope progress protocol, the Harness SHALL return a bounded error observation that helps the model repair the next request. The observation MAY include safe protocol fields such as `repair_required`, `violation_type`, `missing_field`, `expected_previous_tool_call_id`, and allowed `information_gain` values. It SHALL NOT include Tool arguments, normalized scope, raw responses, Prompt, model text, budget values, counters except the bounded consecutive protocol violation count, or internal exception text.
|
||||
|
||||
Consecutive invalid progress protocol requests SHALL be counted independently from `NO_GAIN`. Reaching the Run's configured protocol-violation threshold SHALL set stop reason `PROGRESS_PROTOCOL_VIOLATED` and deliver one `STOP_REQUIRED` observation. A later Tool request after that instruction SHALL terminate through the controlled-stop path.
|
||||
|
||||
#### Scenario: Missing previous observation is repairable
|
||||
- **WHEN** a non-empty Tool observation is pending semantic evaluation and the next Tool request omits `previous_observation`
|
||||
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `missing_field=previous_observation` and the expected previous Tool Call ID
|
||||
|
||||
#### Scenario: Wrong previous observation id is repairable
|
||||
- **WHEN** a pending Tool observation exists and the next Tool request references a different `previous_observation.tool_call_id`
|
||||
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `violation_type=OUT_OF_ORDER_PREVIOUS_OBSERVATION`
|
||||
|
||||
#### Scenario: Repeated repair failure stops collection
|
||||
- **WHEN** the model repeats invalid progress protocol requests until the configured threshold is reached
|
||||
- **THEN** the current Tool is not executed, the model receives `STOP_REQUIRED` with `reason=PROGRESS_PROTOCOL_VIOLATED`, and any further Tool request ends as a controlled stop
|
||||
+5
-1
@@ -11,6 +11,10 @@ The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs
|
||||
- **WHEN** the last successful Tool result is pending semantic evaluation and the model requests another Tool
|
||||
- **THEN** the Envelope must identify that exact prior Tool Call and contain `GAINED` or `NO_GAIN` before the new business Tool can execute
|
||||
|
||||
#### Scenario: Model omits required progress field
|
||||
- **WHEN** a prior non-empty Tool observation is pending and the model requests another Tool without `previous_observation`
|
||||
- **THEN** the business Tool does not execute and the Agent receives a bounded repair observation explaining the missing `previous_observation` field and expected prior Tool Call ID
|
||||
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** a registered adapter returns an error result
|
||||
- **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail
|
||||
@@ -41,7 +45,7 @@ The internal Agent use case SHALL distinguish a valid Draft, controlled informat
|
||||
|
||||
#### Scenario: Model ignores STOP_REQUIRED
|
||||
- **WHEN** the framework surfaces the typed collection-stopped signal after the final completion opportunity
|
||||
- **THEN** the Agent use case returns no Draft with `INFORMATION_SATURATED` and the current ProgressSnapshot
|
||||
- **THEN** the Agent use case returns no Draft with the controlled stop reason and the current ProgressSnapshot
|
||||
|
||||
#### Scenario: Unclassified framework failure
|
||||
- **WHEN** Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination
|
||||
+9
@@ -49,3 +49,12 @@
|
||||
- [x] 7.3 为进展协议错误、重复 scope、信息饱和和观察合同拒绝增加安全 `TOOL_REQUEST_REJECTED` Trace,证明不记录参数或原始响应
|
||||
- [x] 7.4 在 Run 完成时持久化 step_count 并记录 Token 总账/明细对账结果,补充 Store 和 Trace 回归测试
|
||||
- [x] 7.5 运行 focused、Harness 全量回归、strict validate 和真实 SSE/数据库 E2E,核对 Tool 执行/拒绝数量、Token 对账和 Trace 连续性
|
||||
|
||||
## 8. 协议修复反馈与兜底停止
|
||||
|
||||
- [x] 8.1 为进展协议错误建模安全 `violation_type`,覆盖缺失 `previous_observation`、错序 Tool Call ID、意外 previous observation、缺失 `input` 和非法 Envelope
|
||||
- [x] 8.2 将 `INVALID_PROGRESS_PROTOCOL` observation 改为可修正响应,包含安全 repair hint、期望上一轮 Tool Call ID 和允许的 `information_gain` 值,不泄露业务参数或原始响应
|
||||
- [x] 8.3 为 Run tracker 增加连续协议错误计数和 `PROGRESS_PROTOCOL_VIOLATED` stop reason;未达阈值打回修正,达到阈值返回一次 `STOP_REQUIRED`
|
||||
- [x] 8.4 Release/Agent 受控停止路径支持 `PROGRESS_PROTOCOL_VIOLATED`:有安全 ProgressSnapshot 时发布 `INSUFFICIENT_EVIDENCE`,无进展时继续 fail closed
|
||||
- [x] 8.5 扩展 `TOOL_REQUEST_REJECTED` 审计字段,记录安全 `violation_type`、`repair_prompt_delivered`、连续协议错误次数和 stop reason,并用测试证明不含 Tool 参数、上一轮观察正文或内部异常
|
||||
- [x] 8.6 运行 tracker、interceptor、release、agent-loop focused tests 和 strict OpenSpec validate,更新 ISS-015/ISS-016 进度
|
||||
@@ -14,7 +14,6 @@ The system SHALL expose `EVIDENCE_FOUND`, `NO_EVIDENCE`, and `ERROR` as Agent-fa
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** schema, authorization, execution, or projection fails
|
||||
- **THEN** the Agent-facing result uses `evidence_status=ERROR` and SHALL NOT report `NO_EVIDENCE`
|
||||
|
||||
### Requirement: Referencable Tool results SHALL use the framework Tool Call ID
|
||||
Each `EVIDENCE_FOUND` or `NO_EVIDENCE` result SHALL contain the non-blank `tool_call_id` supplied by the framework Tool Call request. Agent inputs SHALL NOT contain `tool_call_id`, and contract code SHALL NOT generate, replace, or derive a second call ID.
|
||||
|
||||
@@ -25,18 +24,16 @@ Each `EVIDENCE_FOUND` or `NO_EVIDENCE` result SHALL contain the non-blank `tool_
|
||||
#### Scenario: Framework ID is invalid
|
||||
- **WHEN** the Tool request has a missing or invalid framework Tool Call ID
|
||||
- **THEN** the call returns an `ERROR` result without inventing a referencable ID
|
||||
|
||||
### Requirement: RAG Tool contract SHALL expose only bounded document evidence
|
||||
The RAG Request SHALL contain only `query`. The RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`.
|
||||
The RAG business Request SHALL contain only `query`. The canonical RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, optional normalized `relevance_level`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`. The RAG projector SHALL accept upstream `relevanceLevel` or `relevance_level` and normalize recognized values without exposing raw relevance scores or retrieval traces.
|
||||
|
||||
#### Scenario: RAG evidence is serialized
|
||||
- **WHEN** a RAG result contains a matching document excerpt
|
||||
- **THEN** its JSON matches the frozen snake_case fields and excludes ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
|
||||
- **WHEN** a RAG result contains a matching document excerpt and an upstream relevance level
|
||||
- **THEN** its canonical JSON preserves the bounded evidence and normalized `relevance_level` while excluding ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
|
||||
|
||||
#### Scenario: RAG query has no evidence
|
||||
- **WHEN** RAG executes successfully without a usable document excerpt
|
||||
- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, and returns an empty evidence list
|
||||
|
||||
- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, returns an empty evidence list, and does not upgrade relevance into evidence
|
||||
### Requirement: Log Tool contract SHALL use logical scope and retain Mock provenance
|
||||
The log Request SHALL contain logical `topic`, `query`, and optional `lookback_minutes` only. The log Result SHALL contain `evidence_status`, `tool_call_id`, `source_kind`, complete query `scope`, `match_count`, `returned_count`, bounded `patterns`, bounded timeline `events`, and `truncated`. The initial logical topics SHALL be `APPLICATION`, `DATABASE_SLOW_QUERY`, and `SYSTEM_EVENTS`, and the initial source kind SHALL be `MOCK`.
|
||||
|
||||
@@ -47,7 +44,6 @@ The log Request SHALL contain logical `topic`, `query`, and optional `lookback_m
|
||||
#### Scenario: Agent creates a log request
|
||||
- **WHEN** the Agent requests log evidence
|
||||
- **THEN** it selects a logical topic and lookback window without supplying region, physical TopicId, credentials, or result limit and without calling a Topic discovery Tool first
|
||||
|
||||
### Requirement: MySQL Tool contract SHALL expose a logical read-only query interface
|
||||
The MySQL Request SHALL contain only logical `data_source`, parameterized `sql`, and `params`. The MySQL Result SHALL contain `evidence_status`, `tool_call_id`, `columns`, bounded structured `rows`, `returned_count`, and `truncated`; it SHALL NOT expose connection details, credentials, internal stack traces, or resource-limit controls.
|
||||
|
||||
@@ -58,17 +54,25 @@ The MySQL Request SHALL contain only logical `data_source`, parameterized `sql`,
|
||||
#### Scenario: Agent creates a MySQL request
|
||||
- **WHEN** the Agent requests business database evidence
|
||||
- **THEN** it supplies a logical data source, SQL placeholders, and parameter values without supplying JDBC connection information or security policy
|
||||
|
||||
### Requirement: Tool descriptions SHALL be concise and implementation-neutral
|
||||
The system SHALL define stable snake_case names and concise descriptions for `lookup_knowledge`, `query_logs`, and `query_mysql`. Each description SHALL state when to call the Tool, its minimal input, and what it cannot query, and SHALL NOT describe retrieval internals, infrastructure coordinates, credentials, audit storage, retries, result limits, or ranking implementation.
|
||||
|
||||
#### Scenario: Agent receives Tool definitions
|
||||
- **WHEN** a future Agent adapter registers the frozen Tool definitions
|
||||
- **THEN** the definitions describe available actions and boundaries without exposing Milvus, L0/L1, rerank, CLS region/TopicId, Redis, JDBC credentials, topK, limit, or Trace internals
|
||||
|
||||
### Requirement: Contract freeze SHALL NOT cut over the current runtime
|
||||
This change SHALL add contract types, descriptions, and tests without changing the current `LookupKnowledgeTool`, `QueryLogsTools`, ChatService, AiOpsService, Controller, Tool registration, or Agent-visible runtime results.
|
||||
|
||||
#### Scenario: Stage 1 tests pass
|
||||
- **WHEN** all ACI contract tests pass
|
||||
- **THEN** the current public Chat and AIOps paths still execute the old Tool implementations until their later projector and cutover changes
|
||||
### Requirement: Tool results SHALL have separate Harness and model views
|
||||
Each successful evidence Tool result SHALL provide a Harness Control View and a bounded Model Observation derived from the same canonical result. The control view MAY contain returned counts, normalized scope, relevance, truncation and duplicate identity. The Model Observation SHALL contain only fields needed to understand and cite the result and SHALL NOT contain raw responses, internal scores, retrieval traces, duplicate fingerprints, counters, thresholds, budgets or store identities.
|
||||
|
||||
#### Scenario: Model receives RAG observation
|
||||
- **WHEN** a canonical RAG result is READY
|
||||
- **THEN** the model receives Tool Call ID, actual query scope, bounded evidence, evidence status, optional coarse relevance and truncation, but not raw scores or Harness counters
|
||||
|
||||
#### Scenario: Harness evaluates duplicate scope
|
||||
- **WHEN** the same normalized Tool scope is requested again
|
||||
- **THEN** the Harness can compare its control view identity without exposing that fingerprint to the model
|
||||
|
||||
@@ -14,7 +14,6 @@ The Harness SHALL reject a Tool call before invoking the executor when the envel
|
||||
#### Scenario: Unauthorized or writable Tool is rejected
|
||||
- **WHEN** `authorized=false` or `readOnly=false`
|
||||
- **THEN** ToolBoundary returns `ERROR` before execution and does not expose the request to an Agent
|
||||
|
||||
### Requirement: Canonical invocation SHALL keep complete data in one record
|
||||
The store SHALL save the complete request JSON, raw Tool response, bounded Agent result, framework `tool_call_id`, run ID, tool name, lifecycle status, evidence status, timestamps, and error code in one canonical record. The raw response SHALL NOT be replaced by a preview, and SHALL NOT be returned to the Agent.
|
||||
|
||||
@@ -25,7 +24,6 @@ The store SHALL save the complete request JSON, raw Tool response, bounded Agent
|
||||
#### Scenario: Projector succeeds
|
||||
- **WHEN** the executor returns raw data and the projector returns a bounded result
|
||||
- **THEN** the same record contains raw data and Agent result with `status=READY`, and ToolBoundary returns only the bounded Agent result
|
||||
|
||||
### Requirement: Invocation lifecycle and evidence semantics SHALL be enforced centrally
|
||||
The store SHALL allow only `PROJECTING -> READY` or `PROJECTING -> ERROR`. READY SHALL require `EVIDENCE_FOUND` or `NO_EVIDENCE`; ERROR SHALL use `evidence_status=ERROR` and SHALL NOT be referencable. `NO_EVIDENCE` SHALL remain a scoped negative observation and SHALL NOT be upgraded to success or retried by the boundary.
|
||||
|
||||
@@ -36,7 +34,6 @@ The store SHALL allow only `PROJECTING -> READY` or `PROJECTING -> ERROR`. READY
|
||||
#### Scenario: Projection fails
|
||||
- **WHEN** the projector throws or returns an invalid evidence status
|
||||
- **THEN** the same record becomes `ERROR`, stores a safe error code, and no Agent result is returned
|
||||
|
||||
### Requirement: Tool Call ID and Run ownership SHALL be preserved
|
||||
The boundary SHALL use the exact framework `tool_call_id` with the RunContext run ID to create the canonical key. It SHALL reject missing, unsafe, duplicate, and cross-Run references and SHALL never generate or replace a fallback ID.
|
||||
|
||||
@@ -47,7 +44,6 @@ The boundary SHALL use the exact framework `tool_call_id` with the RunContext ru
|
||||
#### Scenario: Framework ID is preserved
|
||||
- **WHEN** a valid envelope passes preflight
|
||||
- **THEN** the key and canonical record contain the exact supplied `tool_call_id`
|
||||
|
||||
### Requirement: TTL and size limits SHALL fail closed
|
||||
The Redis implementation SHALL set TTL only when a canonical record is created. Reads SHALL NOT refresh TTL; updates SHALL preserve only the current remaining TTL. Request/raw/record and Agent-result byte limits SHALL be enforced in UTF-8. Raw overflow SHALL produce `ERROR/RESULT_TOO_LARGE` without silent truncation; Agent projection overflow SHALL produce the same error.
|
||||
|
||||
@@ -62,17 +58,25 @@ The Redis implementation SHALL set TTL only when a canonical record is created.
|
||||
#### Scenario: Agent result is too large
|
||||
- **WHEN** a projector returns a result beyond the Agent projection limit
|
||||
- **THEN** the record becomes `ERROR/RESULT_TOO_LARGE` and the oversized result is not returned to the Agent
|
||||
|
||||
### Requirement: Tool and store failures SHALL return safe errors
|
||||
Execution, serialization, Redis, duplicate, preflight, and projection failures SHALL return a bounded error code without raw payload, credentials, internal stack trace, or Redis key to the Agent. If raw data was safely stored before a projection failure, it remains Harness-only canonical data.
|
||||
|
||||
#### Scenario: Tool execution throws
|
||||
- **WHEN** the executor raises an exception after PROJECTING begins
|
||||
- **THEN** the record becomes `ERROR` with a stable execution error code and ToolBoundary returns no raw response
|
||||
|
||||
### Requirement: Stage 3A SHALL remain reusable and independent from legacy audit
|
||||
The boundary and canonical store SHALL be usable by later RAG/log/MySQL projectors without modifying legacy `ToolInvocationRecorder`, JPA `ToolInvocation`, Chat/AIOps, Controller, or public protocol. Redis access SHALL remain inside the store implementation and no Agent-facing API SHALL expose the Redis client.
|
||||
|
||||
#### Scenario: Fake projector tests pass
|
||||
- **WHEN** Fake Tool/Projector tests cover success, no evidence, error, duplicate, TTL, and capacity cases
|
||||
- **THEN** later Tool-specific changes can depend on the same boundary without copying lifecycle or Redis logic, while legacy call sites remain unchanged
|
||||
### Requirement: Run progress projection SHALL reference canonical records without duplicating truth
|
||||
The Run progress tracker SHALL retain only ordered canonical identities for completed Tool calls. At Tool-loop completion, a projector SHALL resolve those identities through the existing canonical store and SHALL accept only READY records owned by the current Run. The tracker SHALL NOT store or reconstruct raw Tool responses, complete Agent results, or a second durable evidence record.
|
||||
|
||||
#### Scenario: Completed calls are projected
|
||||
- **WHEN** a Run ends after multiple READY canonical Tool invocations
|
||||
- **THEN** the progress projector reads each indexed canonical record in execution order and creates bounded observed facts
|
||||
|
||||
#### Scenario: Indexed identity is invalid
|
||||
- **WHEN** an indexed canonical identity is missing, expired, cross-Run, PROJECTING, or ERROR
|
||||
- **THEN** it is excluded from observed facts and cannot become verified evidence
|
||||
|
||||
@@ -0,0 +1,134 @@
|
||||
# diagnosis-information-gain-stop-contract Specification
|
||||
|
||||
## Purpose
|
||||
Diagnosis information-gain stop contract for Harness progress control, protocol repair, and safe release.
|
||||
## Requirements
|
||||
### Requirement: Harness SHALL track binary information gain per Run
|
||||
Each Diagnosis Run SHALL own a thread-safe progress tracker with `GAINED` and `NO_GAIN` as the only information-gain values. `GAINED` SHALL reset the consecutive no-gain count; `NO_GAIN` SHALL increment it. The tracker SHALL expose only `COLLECTING` or `SATURATED` as collection state and SHALL NOT create a second diagnosis lifecycle.
|
||||
|
||||
#### Scenario: New evidence advances diagnosis
|
||||
- **WHEN** a valid pending Tool observation is evaluated as `GAINED`
|
||||
- **THEN** the Run remains `COLLECTING` and its consecutive no-gain count becomes zero
|
||||
|
||||
#### Scenario: Consecutive observations do not advance diagnosis
|
||||
- **WHEN** valid `NO_GAIN` observations reach the Run's configured threshold
|
||||
- **THEN** the tracker becomes `SATURATED` with stop reason `INFORMATION_SATURATED`
|
||||
|
||||
### Requirement: Harness SHALL assign only deterministic no-gain signals
|
||||
The Harness SHALL assign `NO_GAIN` when a successful Tool result has `evidence_status=NO_EVIDENCE` or when a requested `tool_name + normalized_scope` duplicates a successfully completed scope in the same Run. Other successful non-empty results, including RAG `REFERENCE`, SHALL require a model-provided `GAINED` or `NO_GAIN` before another Tool executes.
|
||||
|
||||
#### Scenario: Empty scoped result
|
||||
- **WHEN** a Tool completes READY with `NO_EVIDENCE`
|
||||
- **THEN** the Harness records `NO_GAIN` without asking a model to judge Tool quality
|
||||
|
||||
#### Scenario: Reference material is non-empty
|
||||
- **WHEN** RAG returns bounded evidence with `relevance_level=REFERENCE`
|
||||
- **THEN** the Harness leaves it pending for model evaluation and does not automatically mark it `NO_GAIN`
|
||||
|
||||
#### Scenario: Equivalent structured scope repeats
|
||||
- **WHEN** the model requests the same Tool with the same deterministically normalized successful scope
|
||||
- **THEN** the Harness does not execute the Tool, records `NO_GAIN`, and does not create a second canonical invocation
|
||||
|
||||
### Requirement: Consecutive no-gain threshold SHALL be fixed per Run
|
||||
The system SHALL bind `harness.chat.stop-after-consecutive-no-gain`, require a value of at least one, and default it to `2`. The value SHALL be copied into each new Run's tracker and SHALL NOT enter model context or change an active Run.
|
||||
|
||||
#### Scenario: Default configuration is used
|
||||
- **WHEN** no external value is configured
|
||||
- **THEN** a new Run becomes saturated after two consecutive `NO_GAIN` decisions
|
||||
|
||||
#### Scenario: Invalid threshold is configured
|
||||
- **WHEN** the configured threshold is zero or negative
|
||||
- **THEN** Harness configuration validation fails before serving Chat requests
|
||||
|
||||
### Requirement: Saturated collection SHALL stop further Tool execution
|
||||
When collection becomes saturated, the Harness SHALL reject the pending or next business Tool execution and deliver one bounded `STOP_REQUIRED` observation with `reason=INFORMATION_SATURATED`. If the next model round requests another Tool, the Agent execution SHALL terminate through a typed controlled-stop path without consuming another Tool budget or publishing an internal failure.
|
||||
|
||||
#### Scenario: Model evaluation reaches threshold
|
||||
- **WHEN** the next Tool Call reports `NO_GAIN` and that evaluation reaches the threshold
|
||||
- **THEN** the requested business Tool is not executed and the model receives one STOP_REQUIRED observation
|
||||
|
||||
#### Scenario: Model ignores stop instruction
|
||||
- **WHEN** the model requests another Tool after STOP_REQUIRED was delivered
|
||||
- **THEN** the loop ends as controlled information saturation and proceeds to safe release
|
||||
|
||||
### Requirement: Tool loop completion SHALL project bounded progress once
|
||||
At normal Draft completion, information saturation, or budget termination, the Harness SHALL use the current Run's completed canonical invocation keys to create one bounded `ProgressSnapshot`. The snapshot SHALL contain only safe source, actual scope, objective result summary, truncation and stop reason; it SHALL NOT contain Prompt, thought, raw Tool response, internal counters, remaining budget, Redis keys, or model evaluation rationale.
|
||||
|
||||
#### Scenario: Unknown problem has completed checks
|
||||
- **WHEN** multiple Tools completed but no supported conclusion exists
|
||||
- **THEN** the release input contains their bounded checked scopes and objective results in stable execution order
|
||||
|
||||
#### Scenario: Canonical record cannot be verified
|
||||
- **WHEN** an indexed record is missing, expired, incomplete, ERROR, or not owned by the current Run
|
||||
- **THEN** it is not projected as an observed fact and the snapshot records a bounded limitation
|
||||
|
||||
### Requirement: Information stop reasons SHALL remain distinct from release outcomes
|
||||
The Harness SHALL distinguish `INFORMATION_SATURATED`, `BUDGET_LIMIT_REACHED`, and `PROGRESS_PROTOCOL_VIOLATED`. Release SHALL continue to expose only `SUCCESS`, `FALLBACK`, `FAILED`, or `CANCELLED`, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
|
||||
|
||||
#### Scenario: Low gain stops before budget exhaustion
|
||||
- **WHEN** consecutive no-gain reaches the configured threshold while hard budget remains
|
||||
- **THEN** Trace records `INFORMATION_SATURATED` and public release is a normal `FALLBACK`
|
||||
|
||||
#### Scenario: Hard budget terminates collection
|
||||
- **WHEN** model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
|
||||
- **THEN** Trace retains `BUDGET_LIMIT_REACHED` and Release attempts a deterministic `FALLBACK` from the existing progress without an extra model call
|
||||
|
||||
#### Scenario: Progress protocol violations exceed threshold
|
||||
- **WHEN** consecutive invalid progress protocol requests reach the configured threshold
|
||||
- **THEN** Trace records `PROGRESS_PROTOCOL_VIOLATED`, the model receives one STOP_REQUIRED observation, and public release remains a normal Fallback only if safe progress exists
|
||||
|
||||
### Requirement: Diagnosis Prompt SHALL license bounded abandonment
|
||||
The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, `conclusion=null` is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is `NO_GAIN`. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, `next_action`, or Harness implementation details.
|
||||
|
||||
#### Scenario: Required context is missing before any Tool call
|
||||
- **WHEN** the Query lacks the enterprise, time, service, error, or other context needed for a bounded query
|
||||
- **THEN** the model may return `conclusion=null` with `limitations.missing_info` without calling a Tool
|
||||
|
||||
#### Scenario: Correct content is diagnostically useless
|
||||
- **WHEN** a Tool response is factually correct but only generic, repeated, or unable to change a current hypothesis
|
||||
- **THEN** the model treats it as `NO_GAIN` and does not continue with an equivalent query
|
||||
|
||||
### Requirement: Model token audit SHALL be component-scoped and reconcilable
|
||||
Every Harness model call admitted by the Run budget SHALL receive a bounded component and component round. When Provider Usage is available, the same non-negative input and output Token counts SHALL update both the Run budget total and a `MODEL_TOKEN_USAGE` Trace event. Diagnosis Agent usage SHALL also update the matching `AgentStep.token_count`. At Run completion, Trace SHALL expose whether audited Token totals reconcile with the Run budget total and SHALL expose unavailable Usage counts without fabricating Token values.
|
||||
|
||||
The audit SHALL NOT persist Prompt content, user or model text, reasoning content, Tool arguments, raw model responses, credentials, or provider-specific metadata.
|
||||
|
||||
#### Scenario: Diagnosis Agent round returns Usage
|
||||
- **WHEN** a Diagnosis Agent model round returns input and output Token Usage
|
||||
- **THEN** its component round Trace and matching AgentStep contain the same total Token count and the Run total increases by that amount
|
||||
|
||||
#### Scenario: Multiple model components execute
|
||||
- **WHEN** Router, Diagnosis Agent, Evidence Repair or Semantic Guard model calls execute in one Run
|
||||
- **THEN** each call is distinguishable by bounded component and component round and their audited Token sum can be compared with the Run total
|
||||
|
||||
#### Scenario: Provider Usage is unavailable
|
||||
- **WHEN** a model attempt completes or fails without Provider Usage
|
||||
- **THEN** the audit marks Usage unavailable and Run reconciliation exposes the gap without estimating Token counts
|
||||
|
||||
### Requirement: Rejected Tool requests SHALL remain observable without payload disclosure
|
||||
Every supported Tool request rejected by the Harness before a usable business observation is delivered SHALL emit a `TOOL_REQUEST_REJECTED` Trace event containing only safe Tool Call ID, Tool name and stable error code. Rejected requests SHALL remain distinguishable from canonical `TOOL_INVOCATION` events and SHALL NOT include Tool arguments, normalized scope content, raw responses, internal exception messages, credentials or budget values.
|
||||
|
||||
#### Scenario: Progress protocol is invalid
|
||||
- **WHEN** a supported Tool request omits or misorders a required previous observation
|
||||
- **THEN** no business Tool executes, Trace records `INVALID_PROGRESS_PROTOCOL` with a safe violation type, and the model receives a repairable observation naming the missing or expected protocol field
|
||||
|
||||
#### Scenario: Duplicate or saturated request is blocked
|
||||
- **WHEN** a supported Tool request repeats a successful normalized scope or arrives after collection saturation
|
||||
- **THEN** Trace records the stable rejection reason while canonical invocation count remains unchanged
|
||||
|
||||
### Requirement: Invalid progress protocol SHALL be repairable before bounded stop
|
||||
When a supported Tool request violates the Tool Envelope progress protocol, the Harness SHALL return a bounded error observation that helps the model repair the next request. The observation MAY include safe protocol fields such as `repair_required`, `violation_type`, `missing_field`, `expected_previous_tool_call_id`, and allowed `information_gain` values. It SHALL NOT include Tool arguments, normalized scope, raw responses, Prompt, model text, budget values, counters except the bounded consecutive protocol violation count, or internal exception text.
|
||||
|
||||
Consecutive invalid progress protocol requests SHALL be counted independently from `NO_GAIN`. Reaching the Run's configured protocol-violation threshold SHALL set stop reason `PROGRESS_PROTOCOL_VIOLATED` and deliver one `STOP_REQUIRED` observation. A later Tool request after that instruction SHALL terminate through the controlled-stop path.
|
||||
|
||||
#### Scenario: Missing previous observation is repairable
|
||||
- **WHEN** a non-empty Tool observation is pending semantic evaluation and the next Tool request omits `previous_observation`
|
||||
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `missing_field=previous_observation` and the expected previous Tool Call ID
|
||||
|
||||
#### Scenario: Wrong previous observation id is repairable
|
||||
- **WHEN** a pending Tool observation exists and the next Tool request references a different `previous_observation.tool_call_id`
|
||||
- **THEN** the Tool is not executed and the model receives a repairable `INVALID_PROGRESS_PROTOCOL` observation containing `violation_type=OUT_OF_ORDER_PREVIOUS_OBSERVATION`
|
||||
|
||||
#### Scenario: Repeated repair failure stops collection
|
||||
- **WHEN** the model repeats invalid progress protocol requests until the configured threshold is reached
|
||||
- **THEN** the current Tool is not executed, the model receives `STOP_REQUIRED` with `reason=PROGRESS_PROTOCOL_VIOLATED`, and any further Tool request ends as a controlled stop
|
||||
@@ -13,7 +13,6 @@ The Chat Application Use Case SHALL resolve or generate the session ID, create e
|
||||
#### Scenario: New session request
|
||||
- **WHEN** an internal request omits the session ID
|
||||
- **THEN** the use case generates one valid session ID and uses it for all Run operations
|
||||
|
||||
### Requirement: Minimal isolated Intent Router
|
||||
The Router SHALL execute a direct no-Tool, no-memory, no-ReAct ChatModel call whose input contains only the unchanged Query, optional last intent, and optional last user Query. It SHALL accept only `SYSTEM_CHAT`, `KNOWLEDGE_QUERY`, or `DIAGNOSIS`.
|
||||
|
||||
@@ -24,7 +23,6 @@ The Router SHALL execute a direct no-Tool, no-memory, no-ReAct ChatModel call wh
|
||||
#### Scenario: New topic overrides history
|
||||
- **WHEN** the current Query identifies a new topic while prior routing context exists
|
||||
- **THEN** the model input still preserves the original current Query and prior fields are only optional context
|
||||
|
||||
### Requirement: Router technical retry fails closed
|
||||
The Router SHALL use `HarnessRetryPolicies.intentRouter()` and SHALL retry timeout, transport, or invalid output at most once with identical input. A second failure MUST produce `ROUTING_UNAVAILABLE` and MUST NOT dispatch Diagnosis.
|
||||
|
||||
@@ -35,9 +33,8 @@ The Router SHALL use `HarnessRetryPolicies.intentRouter()` and SHALL retry timeo
|
||||
#### Scenario: Two invalid outputs
|
||||
- **WHEN** both permitted attempts return invalid output
|
||||
- **THEN** the Run ends FAILED and no intent executor is called
|
||||
|
||||
### Requirement: Fixed isolated executors
|
||||
The Application Use Case SHALL map SYSTEM_CHAT to one no-Tool model response, KNOWLEDGE_QUERY to exactly one lookup-knowledge invocation plus one bounded answer model call, and DIAGNOSIS to the single Diagnosis Agent followed by the release boundary. Executors MUST NOT call one another or rewrite the Query.
|
||||
The Application Use Case SHALL map SYSTEM_CHAT to one no-Tool model response, KNOWLEDGE_QUERY to exactly one lookup-knowledge invocation plus one bounded answer model call, and DIAGNOSIS to the single Diagnosis Agent followed by the Diagnosis Release boundary. Executors MUST NOT call one another or rewrite the Query. Information saturation and budget termination in Diagnosis SHALL be converted to safe content by Diagnosis Release, not by ChatApplicationUseCase.
|
||||
|
||||
#### Scenario: System Chat
|
||||
- **WHEN** intent is SYSTEM_CHAT
|
||||
@@ -47,10 +44,13 @@ The Application Use Case SHALL map SYSTEM_CHAT to one no-Tool model response, KN
|
||||
- **WHEN** intent is KNOWLEDGE_QUERY
|
||||
- **THEN** only lookup_knowledge is invoked once and query_logs/query_mysql/Diagnosis ReAct are unavailable
|
||||
|
||||
#### Scenario: Diagnosis
|
||||
- **WHEN** intent is DIAGNOSIS
|
||||
- **THEN** the original Query and bounded PreviousTurn enter DiagnosisAgentUseCase and the Draft cannot publish before DiagnosisReleaseUseCase
|
||||
#### Scenario: Diagnosis succeeds with a conclusion
|
||||
- **WHEN** intent is DIAGNOSIS and the Draft passes the release guards
|
||||
- **THEN** the original Query and bounded PreviousTurn enter DiagnosisAgentUseCase and the Draft publishes only after DiagnosisReleaseUseCase
|
||||
|
||||
#### Scenario: Diagnosis stops without a conclusion
|
||||
- **WHEN** Diagnosis collection is saturated, required context is missing, or a handled budget limit is reached
|
||||
- **THEN** DiagnosisReleaseUseCase returns bounded Fallback content and ChatApplicationUseCase only persists and transports that decision
|
||||
### Requirement: Knowledge references are physically validated
|
||||
The Knowledge executor SHALL accept only a bounded READY RAG projection, SHALL validate every model answer item against the exact direct invocation ID and returned document ID set, and SHALL remove Tool Call IDs from public content.
|
||||
|
||||
@@ -65,7 +65,6 @@ The Knowledge executor SHALL accept only a bounded READY RAG projection, SHALL v
|
||||
#### Scenario: No knowledge evidence
|
||||
- **WHEN** lookup returns READY `NO_EVIDENCE`
|
||||
- **THEN** the executor returns a fixed bounded no-evidence answer without calling the answer model
|
||||
|
||||
### Requirement: Safe bounded PreviousTurn
|
||||
The system SHALL load PreviousTurn only from the same Session's most recent Run with `intent=DIAGNOSIS`, `release_outcome=SUCCESS`, and non-null valid `published_result`. It SHALL deterministically bound fields and source documents without model summarization.
|
||||
|
||||
@@ -80,7 +79,6 @@ The system SHALL load PreviousTurn only from the same Session's most recent Run
|
||||
#### Scenario: Published result is corrupt
|
||||
- **WHEN** stored JSON is invalid or required safe fields are blank
|
||||
- **THEN** PreviousTurn is null and no raw stored value reaches a model
|
||||
|
||||
### Requirement: Diagnosis Run persistence contract
|
||||
The `diagnosis_run` schema SHALL add nullable `intent`, `release_outcome`, and JSON `published_result`, and the JPA entity/repository/store SHALL write and query them consistently. PublishedResult MUST NOT contain Tool IDs, raw evidence, full Draft, or SemanticGuard reasons.
|
||||
|
||||
@@ -95,7 +93,6 @@ The `diagnosis_run` schema SHALL add nullable `intent`, `release_outcome`, and J
|
||||
#### Scenario: Failure or cancellation
|
||||
- **WHEN** routing/execution fails or the client cancels the Run
|
||||
- **THEN** the Run stores exactly one FAILED or CANCELLED release outcome and no PublishedResult
|
||||
|
||||
### Requirement: Protocol-neutral progress and cancellation
|
||||
The use case SHALL notify a protocol-neutral observer after the Run is persisted, expose only session/run identifiers and client-disconnect cancellation, and emit only fixed safe application status codes.
|
||||
|
||||
@@ -106,10 +103,19 @@ The use case SHALL notify a protocol-neutral observer after the Run is persisted
|
||||
#### Scenario: Client disconnect control
|
||||
- **WHEN** the observer invokes client-disconnect cancellation
|
||||
- **THEN** the same RunContext is cancelled and late executor results cannot complete successfully
|
||||
|
||||
### Requirement: Stage-six-A public isolation
|
||||
Stage 6A SHALL NOT modify or switch public Chat Controller endpoints, SSE contracts, frontend consumers, or legacy ChatService behavior.
|
||||
|
||||
#### Scenario: Internal-only delivery
|
||||
- **WHEN** stage 6A changes are inspected
|
||||
- **THEN** Controller/frontend/public endpoint behavior has zero diff and stage 6B can consume the completed use case without rewriting it
|
||||
### Requirement: Handled Diagnosis budget termination SHALL remain a Fallback release
|
||||
When Diagnosis Release has converted a recognized budget termination and existing safe progress into a Fallback, ChatApplicationUseCase SHALL persist that result exactly once without reclassifying it as `INTERNAL_FAILURE`, invoking another model, or rebuilding business fallback content. The internal Run lifecycle MAY retain `BUDGET_EXHAUSTED`, while the persisted public release outcome SHALL be `FALLBACK` and no PublishedResult SHALL be stored.
|
||||
|
||||
#### Scenario: Tool budget ends after finite checks
|
||||
- **WHEN** Diagnosis reaches a hard Tool budget after at least one canonical safe observation and Release creates an insufficient-evidence fallback
|
||||
- **THEN** the Run persists status SUCCESS, release outcome FALLBACK, safe content and actual budget usage, and the SSE sends content followed by done
|
||||
|
||||
#### Scenario: Budget ends without safe publishable progress
|
||||
- **WHEN** budget termination occurs before Diagnosis Release can form a safe bounded result
|
||||
- **THEN** the existing failure path remains fail closed and does not fabricate observed facts
|
||||
|
||||
@@ -13,7 +13,6 @@ The internal diagnosis path SHALL create exactly one `diagnosis_agent` with the
|
||||
#### Scenario: Agent execution fails
|
||||
- **WHEN** the single Diagnosis Agent invocation throws or returns invalid structured output
|
||||
- **THEN** the internal use case fails closed without automatically invoking the Agent or model again
|
||||
|
||||
### Requirement: Diagnosis input SHALL contain only current Query and optional PreviousTurn
|
||||
The internal use case SHALL accept a non-blank current Query and an optional frozen `PreviousTurn`, serialize them as the fixed `query` and `previous_turn` input fields, and SHALL NOT load or accept complete Session history, Redis memory, prior raw Tool results, or model-generated history summaries. The current Query SHALL retain its original text and SHALL NOT be rewritten or silently truncated.
|
||||
|
||||
@@ -24,7 +23,6 @@ The internal use case SHALL accept a non-blank current Query and an optional fro
|
||||
#### Scenario: Query exceeds configured context budget
|
||||
- **WHEN** the current Query exceeds its UTF-8 byte limit
|
||||
- **THEN** the use case rejects it before any model call instead of truncating or rewriting it
|
||||
|
||||
### Requirement: Every model round SHALL be controlled by RunContext
|
||||
The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model interceptor to call `DiagnosisHarnessCore.beforeModelCall` for every framework ReAct model round. Non-streaming model response Usage SHALL be recorded into the same Run budget when available. Cancellation, deadline, model-call exhaustion, or Token exhaustion SHALL prevent subsequent controlled work.
|
||||
|
||||
@@ -35,13 +33,20 @@ The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model
|
||||
#### Scenario: Model-call budget is exhausted
|
||||
- **WHEN** the framework attempts a model round beyond the configured maximum
|
||||
- **THEN** the call is rejected before reaching ChatModel and the Run records budget exhaustion
|
||||
|
||||
### Requirement: Evidence Tools SHALL execute through the Harness boundary
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions. A Tool interceptor SHALL propagate the exact framework Tool Call ID, Run ID, Tool name, and raw JSON arguments into the corresponding stage 3B/3C adapter and `ToolBoundary`. Successful observations SHALL contain only the bounded `agent_result`; failed observations SHALL contain only stable error semantics and SHALL NOT contain raw responses, internal exceptions, credentials, or invocation lifecycle internals.
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions through native Tool Calling. Each Agent-facing Tool input SHALL be a typed Envelope containing optional `previous_observation` and required business `input`. A Tool interceptor SHALL validate and consume the previous observation, propagate the exact framework Tool Call ID, Run ID, Tool name, and unwrapped business JSON into the corresponding adapter and `ToolBoundary`, and SHALL enforce Run progress before execution. Successful observations SHALL use a bounded per-Tool whitelist projection; failed or control observations SHALL contain only stable safe semantics and SHALL NOT contain raw responses, internal exceptions, credentials, invocation lifecycle internals, counters, thresholds, or remaining budget.
|
||||
|
||||
#### Scenario: Framework requests RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`
|
||||
- **THEN** the RAG adapter and Agent observation use exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
#### Scenario: Framework requests first RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`, no pending evaluation and a typed business input
|
||||
- **THEN** the RAG adapter receives only the unwrapped business request, uses exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
|
||||
#### Scenario: Model continues after a non-empty result
|
||||
- **WHEN** the last successful Tool result is pending semantic evaluation and the model requests another Tool
|
||||
- **THEN** the Envelope must identify that exact prior Tool Call and contain `GAINED` or `NO_GAIN` before the new business Tool can execute
|
||||
|
||||
#### Scenario: Model omits required progress field
|
||||
- **WHEN** a prior non-empty Tool observation is pending and the model requests another Tool without `previous_observation`
|
||||
- **THEN** the business Tool does not execute and the Agent receives a bounded repair observation explaining the missing `previous_observation` field and expected prior Tool Call ID
|
||||
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** a registered adapter returns an error result
|
||||
@@ -50,7 +55,6 @@ The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs
|
||||
#### Scenario: Unknown Tool is requested
|
||||
- **WHEN** a model requests a Tool outside the three registered definitions
|
||||
- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
|
||||
|
||||
### Requirement: Diagnosis output SHALL be a bounded DiagnosisDraft
|
||||
The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL return JSON that the internal use case strictly parses into the frozen record. The use case SHALL enforce configured UTF-8 limits for query, previous turn, total input and Draft output and account accepted input/output bytes against the Run capacity. It SHALL reject blank, fenced, prefixed, malformed, oversized or schema-incompatible output without repair or retry.
|
||||
|
||||
@@ -61,14 +65,20 @@ The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL retu
|
||||
#### Scenario: Model returns prose around JSON
|
||||
- **WHEN** the final response contains a Markdown fence or explanatory prefix around an otherwise valid object
|
||||
- **THEN** strict parsing fails and the Diagnosis Agent is not invoked a second time
|
||||
|
||||
### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
|
||||
The single Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. If current evidence cannot support a diagnosis, the Agent SHALL stop the current ReAct execution with `conclusion=null`, describe the actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist or fabricate a root cause.
|
||||
The single Chinese Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. The Agent SHALL NOT be required to find a root cause. If current evidence cannot support a diagnosis, no required context exists for a bounded Tool call, or available results are correct but do not advance any diagnosis hypothesis, the Agent SHALL stop with `conclusion=null`, describe actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist, fabricate a root cause, or make equivalent Tool calls merely to show activity.
|
||||
|
||||
#### Scenario: Tool finds no evidence
|
||||
- **WHEN** the only completed Tool observation has `evidence_status=NO_EVIDENCE`
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records the bounded negative observation and missing information, and makes no additional automatic retry
|
||||
- **WHEN** completed Tool observations have no diagnostic information gain
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records bounded checked scope and missing information, and does not make an equivalent retry
|
||||
|
||||
#### Scenario: No bounded Tool query is possible
|
||||
- **WHEN** the Query lacks required enterprise, time, service or error context
|
||||
- **THEN** the Agent may perform zero Tool calls and returns a no-conclusion Draft whose `limitations.missing_info` identifies the required context
|
||||
|
||||
#### Scenario: Harness requires stop
|
||||
- **WHEN** the Agent receives `STOP_REQUIRED`
|
||||
- **THEN** it emits a bounded final Draft without another Tool call
|
||||
### Requirement: Stage-four execution SHALL remain internal and auditable
|
||||
The new use case SHALL be callable only as an internal Java/test entry in this stage and SHALL NOT be wired into public Chat or AIOps controllers. It SHALL propagate the same sessionId/runId through RunnableConfig metadata, allow the existing AgentStep Hook to be injected, retain ToolBoundary canonical invocation audit, and leave final Run success/persistence ownership to later Guard/application stages.
|
||||
|
||||
@@ -79,3 +89,21 @@ The new use case SHALL be callable only as an internal Java/test entry in this s
|
||||
#### Scenario: Stage four is archived
|
||||
- **WHEN** focused Agent tests pass and the change is archived
|
||||
- **THEN** `ChatController`, public SSE behavior and the old multi-Agent implementation remain available for stages 5, 6A and 6B
|
||||
### Requirement: Diagnosis execution SHALL preserve controlled stop outcomes
|
||||
The internal Agent use case SHALL distinguish a valid Draft, controlled information saturation, and budget termination from an unclassified Agent failure. It SHALL return a bounded internal execution result containing optional Draft, ProgressSnapshot and stop reason, and SHALL NOT convert a recognized controlled stop into `DiagnosisAgentOutputException`.
|
||||
|
||||
#### Scenario: Model ignores STOP_REQUIRED
|
||||
- **WHEN** the framework surfaces the typed collection-stopped signal after the final completion opportunity
|
||||
- **THEN** the Agent use case returns no Draft with the controlled stop reason and the current ProgressSnapshot
|
||||
|
||||
#### Scenario: Unclassified framework failure
|
||||
- **WHEN** Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination
|
||||
- **THEN** execution still fails closed and no safe progress is fabricated
|
||||
|
||||
#### Scenario: Invalid final Draft after verified checks
|
||||
- **WHEN** the final model text is empty or violates the strict DiagnosisDraft contract after the current Run has completed READY canonical Tool checks
|
||||
- **THEN** the invalid text is discarded, the output failure carries only the bounded ProgressSnapshot and safe failure metadata, and no model repair or loose JSON extraction occurs
|
||||
|
||||
#### Scenario: Invalid final Draft without verified checks
|
||||
- **WHEN** the final model text violates the strict DiagnosisDraft contract before any publishable ProgressSnapshot exists
|
||||
- **THEN** execution remains failed and MUST NOT fabricate missing context, observed facts or a no-conclusion Draft
|
||||
|
||||
@@ -4,16 +4,23 @@
|
||||
TBD - created by archiving change single-react-evidence-semantic-guards. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Deterministic Draft and evidence validation
|
||||
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
|
||||
For a Draft with a non-null Conclusion, the Harness SHALL deterministically reject it unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs. For a Draft with `conclusion=null`, Release SHALL NOT require normal conclusion structure or invoke EvidenceRepair; any supplied Tool references SHALL still resolve to current-Run READY canonical invocations and SHALL obey positive/negative evidence semantics.
|
||||
|
||||
#### Scenario: Duplicate or missing Analysis ID
|
||||
- **WHEN** a Draft contains a blank or duplicate Analysis ID
|
||||
#### Scenario: Duplicate or missing Analysis ID in concluded Draft
|
||||
- **WHEN** a Draft with a Conclusion contains a blank or duplicate Analysis ID
|
||||
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
|
||||
|
||||
#### Scenario: Broken report reference
|
||||
#### Scenario: Broken report reference in concluded Draft
|
||||
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
|
||||
- **THEN** EvidenceGuard rejects the Draft before semantic review
|
||||
|
||||
#### Scenario: No-conclusion Draft has valid negative observation
|
||||
- **WHEN** a `conclusion=null` Draft cites a current-Run READY `NO_EVIDENCE` call as `NEGATIVE_OBSERVATION`
|
||||
- **THEN** Release accepts the reference authenticity without running EvidenceRepair or SemanticGuard
|
||||
|
||||
#### Scenario: No-conclusion Draft fabricates a Tool reference
|
||||
- **WHEN** a `conclusion=null` Draft cites a missing, cross-Run, incomplete or ERROR Tool call
|
||||
- **THEN** the reference is excluded and cannot be published as an observed fact
|
||||
### Requirement: Current Run canonical invocation ownership
|
||||
EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`.
|
||||
|
||||
@@ -24,7 +31,6 @@ EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallI
|
||||
#### Scenario: Failed or incomplete invocation
|
||||
- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR`
|
||||
- **THEN** EvidenceGuard rejects the Draft
|
||||
|
||||
### Requirement: Analysis kind matches evidence semantics
|
||||
EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls.
|
||||
|
||||
@@ -35,7 +41,6 @@ EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls an
|
||||
#### Scenario: Positive claim uses no-evidence result
|
||||
- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation
|
||||
- **THEN** EvidenceGuard rejects the binding
|
||||
|
||||
### Requirement: Verified evidence snapshot is minimal and deterministic
|
||||
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
|
||||
|
||||
@@ -46,7 +51,6 @@ The Harness SHALL strictly parse only supported Tool projections and SHALL const
|
||||
#### Scenario: Projection contract mismatch
|
||||
- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
|
||||
- **THEN** EvidenceGuard fails closed
|
||||
|
||||
### Requirement: Evidence repair is single-turn and semantics-preserving
|
||||
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
|
||||
|
||||
@@ -61,7 +65,6 @@ On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct n
|
||||
#### Scenario: Second validation fails
|
||||
- **WHEN** the repaired Draft still fails EvidenceGuard
|
||||
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard
|
||||
|
||||
### Requirement: Isolated single-turn SemanticGuard
|
||||
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
|
||||
|
||||
@@ -72,7 +75,6 @@ SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Promp
|
||||
#### Scenario: Binary review output
|
||||
- **WHEN** SemanticGuard completes normally
|
||||
- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report
|
||||
|
||||
### Requirement: Semantic model budgets timeout cancellation and retry
|
||||
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
|
||||
|
||||
@@ -87,12 +89,11 @@ The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets
|
||||
#### Scenario: Run cancellation during model call
|
||||
- **WHEN** the Run is cancelled while a guard model call is pending
|
||||
- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
|
||||
|
||||
### Requirement: Fail-closed release policy
|
||||
The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
|
||||
The release use case SHALL publish the unchanged verified Draft only when a non-null Conclusion passes EvidenceGuard and SemanticGuard returns `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, final semantic technical failure, a valid no-conclusion Draft, information saturation, or handled budget termination. No-conclusion and controlled-stop release SHALL be deterministic from verified references and ProgressSnapshot and MUST NOT invoke a repair or semantic model call. A fallback result MUST NOT include an unsupported Draft, full verified snapshot, internal stop counters, or SemanticGuard reason.
|
||||
|
||||
#### Scenario: Supported report release
|
||||
- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED`
|
||||
- **WHEN** EvidenceGuard succeeds for a concluded Draft and SemanticGuard returns `SUPPORTED`
|
||||
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
|
||||
|
||||
#### Scenario: Unsupported report fallback
|
||||
@@ -104,9 +105,20 @@ The release use case SHALL publish the unchanged verified Draft only for `SUPPOR
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
|
||||
|
||||
#### Scenario: Evidence validation fallback sources
|
||||
- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft
|
||||
- **WHEN** evidence repair fails or the second EvidenceGuard rejects a concluded Draft
|
||||
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
|
||||
|
||||
#### Scenario: Missing context ends without Tool calls
|
||||
- **WHEN** a valid no-conclusion Draft has no Tool calls and identifies required missing context
|
||||
- **THEN** release outcome is `FALLBACK`, type is `MISSING_REQUIRED_CONTEXT`, and no guard model call occurs
|
||||
|
||||
#### Scenario: Finite checks do not support a conclusion
|
||||
- **WHEN** a no-conclusion Draft or controlled stop has a non-empty verified ProgressSnapshot
|
||||
- **THEN** release outcome is `FALLBACK`, type is `INSUFFICIENT_EVIDENCE`, and observed facts describe only actual completed checks
|
||||
|
||||
#### Scenario: Invalid Draft has publishable progress
|
||||
- **WHEN** the Agent's final Draft is rejected by strict parsing but its bounded ProgressSnapshot contains current-Run verified observed facts
|
||||
- **THEN** Release publishes `FALLBACK` with type `INSUFFICIENT_EVIDENCE` using only that snapshot and MUST NOT use any content from the invalid Draft
|
||||
### Requirement: Stage-five public isolation
|
||||
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user