feat(harness): add information gain stop and audit
This commit is contained in:
@@ -0,0 +1,5 @@
|
||||
Committed OpenSpec
|
||||
|
||||
Validated: 2026-07-26
|
||||
Scale: complex
|
||||
Interface impact: L3
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-26
|
||||
@@ -0,0 +1,172 @@
|
||||
## Context
|
||||
|
||||
当前 Diagnosis ReAct loop 只有模型、Tool、Token、字节和时间预算,没有“查询是否仍在产生信息”的状态。`HarnessToolInterceptor` 将完整 canonical `agent_result` 直接放入模型上下文;三个 Agent-facing Tool 使用裸业务 request 生成 Schema;`DiagnosisAgentUseCase` 只返回非空 `DiagnosisDraft`;`DiagnosisReleaseUseCase` 要求 Draft 非空并对所有 Draft 运行 EvidenceGuard/Repair/SemanticGuard。预算耗尽的临时 Fallback 位于 `ChatApplicationUseCase`,导致业务 Release 决策分散。
|
||||
|
||||
本变更是 L3 模型协作接口演进。公开 HTTP/SSE、业务 Tool backend、数据库表和 `SafeFallback` JSON 结构保持兼容。当前环境没有 semantic retrieval/LSP,调用链通过源码、测试和 `rg` 引用核查完成。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 在硬预算耗尽前确定性停止连续无增益 Tool 调用。
|
||||
- 让模型在看到成功非空结果后,以 `GAINED / NO_GAIN` 表达该结果是否推进当前诊断。
|
||||
- 允许模型零次调用 Tool 或以 `conclusion=null` 合法结束。
|
||||
- 信息饱和、预算终止和主动无结论均由 Diagnosis Release 发布为有过程的安全 Fallback。
|
||||
- 有结论 Draft 继续经过完整 EvidenceGuard、单次 EvidenceRepair 和 SemanticGuard。
|
||||
- 控制数据、raw payload、内部 thought 和预算不进入模型上下文或用户结果。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不引入 Judge、`UNKNOWN`、分数、`new_count`、`next_action` 或第二套诊断生命周期。
|
||||
- 不做自然语言语义去重或跨 Run 进展继承。
|
||||
- 不改变 Tool backend 参数、公开 SSE 事件或前端字段协议。
|
||||
- 不把候选文档、`REFERENCE` 或非空结果自动认定为诊断证据。
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. RunContext 持有最小线程安全 ProgressTracker
|
||||
|
||||
新增 `DiagnosisProgressTracker` handle,并在 `DiagnosisHarnessCore.startRun` 时用固定阈值创建。tracker 维护:连续 `NO_GAIN`、`COLLECTING / SATURATED`、可选 `stop_reason`、最后一个待模型评价的成功 Tool Call、已成功检查的规范化 scope,以及已完成 canonical key/Tool Call ID 的有序索引。
|
||||
|
||||
tracker 不保存 raw response、完整 Agent result 或用户可见摘要。Canonical Store 仍是 Tool 真相源;结束投影器按 tracker 的 key 索引逐条读取 canonical 记录。
|
||||
|
||||
替代方案是给 `CanonicalInvocationStore` 增加按 Run 枚举。拒绝,因为停止状态是短生命周期控制数据,且该方案会扩大 Redis 接口及所有 fake store,实现另一种事实索引。
|
||||
|
||||
### 2. 使用三个强类型 Envelope,共享上一轮评价类型
|
||||
|
||||
Agent-facing Schema 使用三个具体输入类型:`RagToolCall`、`QueryLogsToolCall`、`MysqlToolCall`。每个类型包含可选 `previous_observation` 和必填 `input`;`input` 继续使用现有业务 request,`previous_observation` 使用共享的 `PreviousObservation(tool_call_id, information_gain)`。
|
||||
|
||||
不使用泛型 `ToolCallEnvelope<T>` 直接生成 Schema,因为运行时类型擦除可能使嵌套 `input` 丢失具体字段;不使用 `JsonNode`,因为它无法给模型提供强 Schema。interceptor 严格解析对应 Envelope,先消费控制字段,再把 `input` 序列化为原业务 JSON 交给 adapter。
|
||||
|
||||
首个 Tool Call 以及上一结果已由 Harness 确定性评价时允许省略 `previous_observation`。存在待评价调用时,新 Tool Call 必须携带完全匹配的上一 Tool Call ID 和二值评价;缺失、错序或跨 Run 引用返回有界协议错误且不执行 Tool。
|
||||
|
||||
### 3. 先应用上一轮评价,再决定本轮是否放行
|
||||
|
||||
interceptor 的固定顺序为:
|
||||
|
||||
1. 严格解析 Envelope 并校验 `previous_observation`。
|
||||
2. 将上一轮 `GAINED` 清零计数,或将 `NO_GAIN` 加一。
|
||||
3. 若达到阈值,拒绝当前业务 Tool,返回一次 `STOP_REQUIRED/INFORMATION_SATURATED`。
|
||||
4. 规范化本次业务 request 的 scope;若与同 Tool 的成功历史 scope 重复,则不执行 Tool、记一次确定性 `NO_GAIN`,并返回有界重复提示或 STOP_REQUIRED。
|
||||
5. 其余请求进入现有 adapter/ToolBoundary。
|
||||
6. 成功后记录 canonical key;`NO_EVIDENCE` 由 Harness 立即记为 `NO_GAIN`,`EVIDENCE_FOUND` 标记为待模型评价。
|
||||
|
||||
Tool `ERROR` 不产生 information gain,也不把失败 scope 写入成功去重集合;它沿用技术故障和现有预算/重试边界。
|
||||
|
||||
替代方案是让 Tool 自行报告质量。拒绝,因为 Tool 只知道客观返回,无法判断对当前诊断假设的价值。
|
||||
|
||||
### 4. 只做确定性 scope 规范化
|
||||
|
||||
每个 Tool 提供一个无副作用 scope projector:RAG 使用规范化 query;日志使用 topic、query 和实际 lookback;MySQL 使用 logical datasource、规范化 SQL 文本和参数。仅规范化空白、大小写明确不敏感的枚举/标识、确定性默认值和结构化集合,不判断自然语言改写是否语义等价。
|
||||
|
||||
重复只比较 `tool_name + normalized_scope`,且只基于已成功执行的历史 scope。被拒绝的重复调用不进入 ToolBoundary、canonical store 或 Tool 调用预算,但会记录安全 Trace 并推进无增益计数。
|
||||
|
||||
### 5. Canonical 结果产生控制视图和模型白名单视图
|
||||
|
||||
ToolBoundary 继续保存完整 request、raw response 和 bounded `agent_result`。Tool 完成后:
|
||||
|
||||
- Harness Control View 读取 execution/evidence status、returned count、normalized scope、RAG relevance、truncated 和 canonical identity。
|
||||
- Model Observation 只包含 Tool Call ID、实际 scope、有界 evidence/rows/events、`evidence_status`、可选 `relevance_level` 和 `truncated`。
|
||||
|
||||
`HarnessToolInterceptor` 不再直接返回完整 `agent_result`,而是通过按 Tool 类型的 `AgentObservationProjector` 白名单序列化。RAG projector 兼容读取 `relevanceLevel` 和 `relevance_level`,在 canonical RAG result 中统一为 `relevance_level`。`REFERENCE` 仍交给模型判断信息增益。
|
||||
|
||||
### 6. 饱和后只有一次正常收尾机会
|
||||
|
||||
当 Tool 结果使 tracker 直接饱和时,当前 Model Observation 同时携带 `stop_required=true` 和 `reason=INFORMATION_SATURATED`。当模型在下一次 Tool Call 中回传 `NO_GAIN` 后达到阈值时,该本轮 Tool 被拒绝并返回同样的 STOP_REQUIRED observation。
|
||||
|
||||
tracker 记录控制指令已交付。下一轮模型仍发起 Tool Call时,interceptor 抛出可识别的 `DiagnosisCollectionStoppedException`;`DiagnosisAgentUseCase` 不把它包装为内部故障,而是返回“无 Draft + stop reason”的执行结果。这样模型获得一次生成合法 Draft 的机会,同时无法靠重复 Tool Call继续空转。
|
||||
|
||||
替代方案是无限返回 STOP_REQUIRED。拒绝,因为模型无视指令时仍会消耗模型预算并重现原问题。
|
||||
|
||||
### 7. Agent 执行返回 Draft 与停止投影,而不是只返回 Draft
|
||||
|
||||
`DiagnosisAgentUseCase` 返回内部 `DiagnosisAgentExecution`:可选 Draft、`ProgressSnapshot` 和可选 `DiagnosisStopReason`。正常 Draft、强制饱和停止和预算异常都在 Diagnosis executor 边界形成 Release 输入。
|
||||
|
||||
`DiagnosisProgressProjector` 在 Tool loop 结束时读取 tracker 索引和 canonical store,一次性生成有界 snapshot;缺失、过期、ERROR 或不属于当前 Run 的记录不会成为 verified fact,并产生稳定 limitation/trace。snapshot 不进入模型上下文。
|
||||
|
||||
### 8. DiagnosisReleaseUseCase 统一业务发布决策
|
||||
|
||||
Release 接收 query、可选 Draft、ProgressSnapshot 和可选 stop reason:
|
||||
|
||||
- `conclusion != null`:执行现有完整 EvidenceGuard、一次 Repair、重验和 SemanticGuard。
|
||||
- `conclusion == null`:不执行 Repair/SemanticGuard;若 Draft 有 Tool 引用,仅验证当前 Run canonical 引用真实性和负向语义,不要求正常结论结构。
|
||||
- 无 Draft且 `INFORMATION_SATURATED` 或 `BUDGET_LIMIT_REACHED`:从 snapshot 确定性生成 Fallback,不调用额外模型。
|
||||
- 零 Tool 且 Draft 的 `limitations.missing_info` 非空:发布 `MISSING_REQUIRED_CONTEXT`。
|
||||
- 有有限排查但无可支持结论:发布 `INSUFFICIENT_EVIDENCE`。
|
||||
- 真正 Tool/模型/基础设施故障且无法形成安全过程:继续 `FAILED`。
|
||||
|
||||
最终模型文本为空或不能严格解析为 `DiagnosisDraft` 时,非法内容本身始终被丢弃,不做 Markdown/自然语言 JSON 抽取,也不调用额外模型修复。Agent 输出异常携带一次有界 `ProgressSnapshot`;Executor 仅在 snapshot 含当前 Run 已验真的 observed facts 时交给 Release 生成 `INSUFFICIENT_EVIDENCE`,否则保持 `FAILED`。Trace 只记录固定失败类别、输出字节数和是否存在可发布进展,不记录模型原文、字段值或解析异常文本。
|
||||
|
||||
`ChatApplicationUseCase` 删除业务内容级 `recoverBudgetExhaustion`。Diagnosis executor 捕获可识别预算终止并调用 Release;Application 只允许该已处理 Diagnosis Fallback 跳过 `core.checkActive/completeSuccess`,持久化 `release_outcome=FALLBACK`。Run 内部仍保留 `BUDGET_EXHAUSTED`,数据库公开运行状态继续按既有规则记录为成功发布的 Fallback。
|
||||
|
||||
### 9. Prompt 只约束模型职责,不复制 Tool Schema
|
||||
|
||||
中文 Prompt 明确:无需强行得出根因;`conclusion=null` 是合法完成;缺少企业、时间、服务或错误信息时允许零 Tool 并填写 `limitations.missing_info`;只有新增可验证事实确认、排除或缩小假设才是 `GAINED`;正确但无用、通用或重复内容是 `NO_GAIN`;没有明确不同且可能产生新信息的 scope 时停止;收到 STOP_REQUIRED 后不得继续调用 Tool。
|
||||
|
||||
Prompt 不写 Tool 名、Schema、阈值、计数器、Projector、预算或 `next_action`。
|
||||
|
||||
### 10. Trace 记录决策,不泄露推理
|
||||
|
||||
Trace 增加有界事件或字段,记录 Tool Call ID、Tool name、scope 摘要、information gain 的生产者(Harness/Model)、连续计数变化、collection state 和 stop reason。不得记录 Prompt、模型 thought、raw Tool response、完整 SQL 参数、预算余量或模型评价理由。
|
||||
|
||||
### 11. Run 总账与模型调用明细使用同一份 Provider Usage
|
||||
|
||||
`RunContext` 增加最小线程安全模型调用账本,只维护组件轮次、已审计调用数、Usage 不可用调用数和 Token 合计。每次实际模型调用在预算放行后取得组件轮次;`HarnessModelInterceptor` 负责 Diagnosis Agent,`GuardModelCall` 负责 Router、System Chat、Knowledge Answer、Evidence Repair 和 Semantic Guard。两条入口都从 Spring AI `Usage` 读取同一组 input/output Token,先登记调用明细,再交给现有 `RunBudget` 累加总账。
|
||||
|
||||
每个模型调用 Trace 只包含 `component`、`component_round`、`usage_available`,并在 Usage 可用时包含 `input_tokens`、`output_tokens` 和 `total_tokens`;Usage 不可用时不写 Token 字段。Diagnosis Agent 的对应 `AgentStep.token_count` 回填 total Token;其他组件不伪装成 AgentStep。Run 结束事件同时写入预算总账、审计明细合计、Usage 不可用数量和 `tokens_reconciled`,从而显式暴露缺口而不是把未知 Token 当作零消耗。
|
||||
|
||||
进入 `HarnessToolInterceptor` 的 Tool 请求若因进展协议、重复 scope、信息饱和或观察合同失败而未进入/未成功交付业务边界,记录 `TOOL_REQUEST_REJECTED`。事件只保留安全 Tool Call ID、Tool name 和稳定 `error_code`;业务参数、原始响应、内部异常和预算余量均不进入 Trace。
|
||||
|
||||
不新增模型审计表:`diagnosis_trace_event` 是调用明细账,`RunBudget`/`diagnosis_run.total_token_count` 是 Run 总账,`AgentStep.token_count` 是 Diagnosis Agent 轮次摘要。这样避免三套可独立漂移的 Token 真相源。
|
||||
|
||||
## Module Map
|
||||
|
||||
```text
|
||||
ChatApplicationUseCase
|
||||
-> DiagnosisChatExecutor
|
||||
-> DiagnosisAgentUseCase
|
||||
-> ReactAgent
|
||||
-> HarnessModelInterceptor
|
||||
-> HarnessToolInterceptor
|
||||
-> DiagnosisProgressTracker (RunContext handle)
|
||||
-> HarnessEvidenceTools -> ToolBoundary -> CanonicalInvocationStore
|
||||
-> AgentObservationProjector
|
||||
-> DiagnosisProgressProjector -> CanonicalInvocationStore
|
||||
-> DiagnosisReleaseUseCase
|
||||
-> no-conclusion reference validation / SafeFallbackFactory
|
||||
-> EvidenceGuard -> EvidenceRepair -> SemanticGuard (conclusion only)
|
||||
-> ChatRunStore -> named SSE content
|
||||
```
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- 级别:L3 协作接口。
|
||||
- Agent-facing input 从裸业务 request 改为 `{previous_observation?, input}`。
|
||||
- 业务 adapter、backend、公开 HTTP/SSE、数据库和前端消费字段保持兼容。
|
||||
- 所有 Tool loop scripted tests 必须使用新 Envelope;KnowledgeQueryExecutor 若直接调用 registry bridge,继续走业务 request,不使用 Agent-facing Envelope。
|
||||
- 不提供旧/新 Schema 双轨;回滚以整个 change 为单位。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [框架不能按预期生成嵌套强类型 Schema] -> 使用三个具体 Envelope record,并增加真实 callback schema 测试。
|
||||
- [STOP_REQUIRED 后异常被框架包装] -> 使用 cause-chain 分类测试,只有专用受控停止异常可转换为 Release 输入。
|
||||
- [预算终止与 Application active check 冲突] -> DiagnosisExecutionResult 显式标记已处理终止 Fallback,Application 仅对此窄分支跳过 success transition。
|
||||
- [canonical TTL 到期导致过程不完整] -> Run timeout 小于 canonical TTL;投影缺失 fail closed 为 limitation,不伪造事实。
|
||||
- [scope 规范化误判不同查询为重复] -> 首版只规范确定性字段,测试每个 Tool 的相同/不同 scope。
|
||||
- [模型伪造上一轮评价 ID] -> tracker 只接受当前 Run 最后一个待评价 ID,错序/重复消费均拒绝。
|
||||
- [无结论 Draft 绕过安全检查] -> 只跳过结论 Repair/SemanticGuard;引用真实性、当前 Run 所有权和负向语义仍确定性验证。
|
||||
- [模型完成排查后输出非法 Draft 导致过程丢失] -> 丢弃非法 Draft;仅当 ProgressSnapshot 含已验真 observed facts 时由 Release 确定性降级,无进展仍 fail closed。
|
||||
- [脏工作区行为丢失] -> 迁移预算 Fallback 的测试意图,实施前后用 scoped diff 核对,不覆盖无关修改。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 先加入状态/Envelope/scope/projector 类型和 focused contract tests,不切换 Release。
|
||||
2. 接入 interceptor、tracker、STOP_REQUIRED 和执行结果,固定真实框架 loop 行为。
|
||||
3. 接入 ProgressSnapshot 与统一 Release,迁移 Application 预算 Fallback。
|
||||
4. 更新中文 Prompt、Trace、配置和文档。
|
||||
5. 运行 focused、Harness 回归和全量测试;再用 Maven 启动项目执行原始未知 Query 的 SSE、日志、数据库 exact-run E2E。
|
||||
6. 回滚时整体回滚本 change;不单独恢复旧 Tool Schema 或 Application 预算分支。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。阈值、状态、Prompt、协议、去重边界、Release 所有权和兼容范围均已确认。
|
||||
@@ -0,0 +1,79 @@
|
||||
## Why
|
||||
|
||||
Diagnosis Agent 面对知识库未知、日志为空或查询条件不足的问题时,当前只能依赖 Tool、Token 和轮次预算停止。模型可能不断改写查询继续调用 Tool,最终以 `BUDGET_EXHAUSTED` 或 `INTERNAL_FAILURE` 结束;用户只能看到通用错误,无法看到已经完成的检查和证据缺口。
|
||||
|
||||
系统需要把“没有足够证据得出结论”视为正常、可发布的诊断结果,同时用确定性的 Harness 规则阻止无信息增益的空转,并继续由 EvidenceGuard 和 SemanticGuard 拦截无证据结论。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 在 Run 内增加最小进展状态:`GAINED / NO_GAIN`、连续无增益计数、`COLLECTING / SATURATED` 和内部 `stop_reason`。
|
||||
- 通过模型原生 Tool Calling 注册统一 Tool Call Envelope;模型继续调用 Tool 时,在下一次调用中回传上一轮 `information_gain`,Harness 校验并消费控制字段,业务 Tool 请求保持原结构。
|
||||
- Harness 对 `NO_EVIDENCE` 和重复的 `tool_name + normalized_scope` 确定性赋值 `NO_GAIN`;其他成功非空结果由模型判断 `GAINED / NO_GAIN`。
|
||||
- 配置 `harness.chat.stop-after-consecutive-no-gain`,默认值为 `2`,达到阈值后拒绝新的业务 Tool 执行并要求结束。
|
||||
- 将 Canonical Tool Result 分为 Harness Control View 和白名单 Model Observation;保留 RAG `relevance_level`,不把内部计数、预算、检索轨迹或 raw response 放进模型上下文。
|
||||
- Tool Loop 结束时从当前 Run 已完成的 canonical 调用一次性投影 `ProgressSnapshot`,用于发布已检查来源、scope、客观结果和证据缺口。
|
||||
- 统一 `DiagnosisReleaseUseCase` 对正常 Draft、信息饱和和预算终止的发布决策;`conclusion=null` 不触发 EvidenceRepair,有结论时才执行完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard。
|
||||
- 最终 Draft 违反结构化输出契约时继续拒绝该 Draft;若当前 Run 已有可验真的 ProgressSnapshot,则仅从 canonical 过程确定性发布 `INSUFFICIENT_EVIDENCE`,没有安全进展时仍 fail closed。
|
||||
- 复用现有 `SafeFallback` 和 SSE/前端协议,增加或规范 `INSUFFICIENT_EVIDENCE`、`MISSING_REQUIRED_CONTEXT`,迁移当前位于 `ChatApplicationUseCase` 的预算兜底意图。
|
||||
- 精简中文 Diagnosis Prompt,明确模型不必须得出根因、允许零次 Tool 调用、正确但对当前推导无用的内容属于 `NO_GAIN`,且合法放弃是成功完成。
|
||||
- 增加最小模型调用审计:按组件和轮次记录 input/output/total Token,回填 Diagnosis AgentStep Token,并在 Run 结束时与预算总账对账;不记录 Prompt、模型正文或推理内容。
|
||||
- 对进入 Harness 后被协议、饱和或重复 scope 门禁拒绝的 Tool 请求记录安全 Trace,使预算 Tool 计数、实际执行和拒绝决策可区分;不记录 Tool 参数或原始响应。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `diagnosis-information-gain-stop-contract`: 定义 Tool 信息增益回传、Harness 饱和停止、进展投影和安全发布行为。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `single-react-diagnosis-agent`: Tool Schema 改为服务端注册的统一 Envelope,Prompt 和 Agent 结束行为支持合法放弃。
|
||||
- `canonical-tool-invocation-store`: canonical 结果继续作为 Tool 真相源,并支持当前 Run 在结束时投影进展,不新增第二套持久化真相。
|
||||
- `aci-evidence-tool-contracts`: RAG 投影保留 `relevance_level`,Tool 结果拆分控制视图和模型白名单视图。
|
||||
- `single-react-evidence-semantic-guards`: 无结论 Draft 不进入结论修复链;有结论仍执行完整证据与语义保护。
|
||||
- `single-react-chat-application-usecase`: 预算与饱和的业务 Fallback 由 Diagnosis Release 统一决策。
|
||||
|
||||
## Scope
|
||||
|
||||
- Diagnosis Agent Prompt、Tool Schema/Interceptor、三个 Tool contract 的 Agent-facing 包装。
|
||||
- RunContext 进展 tracker、确定性 scope 规范化和连续无增益停止门禁。
|
||||
- RAG/日志/MySQL canonical 控制视图与模型观察投影。
|
||||
- Diagnosis Agent 执行结果、ProgressSnapshot、Release、Fallback 和 Trace。
|
||||
- Draft 合同失败的脱敏 Trace 与“有安全进展才允许降级”的 Executor/Release 边界。
|
||||
- 模型调用 Token 明细、Run 对账摘要和 Tool 请求拒绝 Trace。
|
||||
- Spring 配置绑定、focused tests、回归测试和真实 SSE/日志/数据库 E2E。
|
||||
|
||||
## Non-goals
|
||||
|
||||
- 不引入 `new_count`、`next_action`、多级质量分数、`UNKNOWN` 或独立 Judge 模型。
|
||||
- 不做自然语言语义去重,只比较确定性的 `tool_name + normalized_scope`。
|
||||
- 不让 Tool 或 Harness 判断业务根因,不把 `REFERENCE` 自动等同于 `NO_GAIN`。
|
||||
- 不增加 `CONFIRMED / INSUFFICIENT_EVIDENCE / NEED_MORE_INFO` 第二套诊断生命周期状态。
|
||||
- 不修改公开 HTTP/SSE 事件结构,不新增前端页面或新的用户可见进度协议。
|
||||
- 不重新引入 Planner/Executor/Verifier/Composer 多 Agent 链路,不处理 ISS-015 的 Reasoning 原文审计治理。
|
||||
|
||||
## Context Constraints
|
||||
|
||||
- Tool 通过 `DiagnosisAgentFactory.tools(...)` 的原生 Tool Calling 通道注册,Prompt 不写 Tool 名称、Schema 或 Schema 占位符。
|
||||
- `CanonicalInvocationStore` 是当前 Run 完整 Tool 调用真相源;Run 内 tracker 只维护控制状态和完成调用索引,不复制 raw response。
|
||||
- `RunContext` 继续使用“结构不可变 + 线程安全可变 handle”的既有模式;阈值在 Run 启动时固定。
|
||||
- 现有 `SafeFallback.observed_facts / verified_sources / limitations / next_steps` 和前端渲染能力必须复用。
|
||||
- 当前未提交的预算 Fallback 修改保留用户价值,但业务 Release 决策需要从 `ChatApplicationUseCase` 迁移到 `DiagnosisReleaseUseCase`。
|
||||
- 当前环境缺少 `codebase-retrieval` 和 LSP;本次影响核查使用 `rg` 引用搜索、源码和测试阅读降级完成。用户已明确允许忽略 GitNexus。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- 级别:L3 协作接口。
|
||||
- 变更对象:模型可见的三个 Tool input schema、Tool Call 解析、Diagnosis Agent 到 Release 的内部执行结果。
|
||||
- 兼容性:业务 Tool request、Tool backend、公开 HTTP/SSE 和持久化表结构保持不变;旧的模型 Tool 参数形状不再被 Agent-facing schema 接受。
|
||||
- 消费者:`DiagnosisAgentFactory`、`HarnessEvidenceTools`、`HarnessToolInterceptor`、三个 adapter 及 scripted Tool-loop tests。
|
||||
- 回滚:整体回滚本 change;不提供双 Tool Schema 或兼容分支。
|
||||
|
||||
## Risks
|
||||
|
||||
- 框架 Tool Schema 生成或 ToolInterceptor 参数处理不符合 Envelope 假设,导致模型无法正确回传或业务 request 未被剥离。
|
||||
- STOP_REQUIRED 若未形成受控结束,模型可能继续请求 Tool,或 Agent 调用异常绕过统一 Release。
|
||||
- 最后一轮结果可能没有下一次 Tool Call 来回传模型评价;该情况只能表示模型主动结束,不能伪造 `GAINED / NO_GAIN`。
|
||||
- ProgressSnapshot 若读取不完整或混入 raw payload,会造成过程缺失或上下文/隐私边界倒退。
|
||||
- `conclusion=null` 与现有 EvidenceGuard 的 `ANALYSIS_MISSING` 规则冲突,需要明确区分“验证引用真实性”和“验证结论完整性”。
|
||||
- 现有脏工作区包含相关预算 Fallback 改动,实施时必须迁移其意图并避免覆盖其他历史修改。
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: RAG Tool contract SHALL expose only bounded document evidence
|
||||
The RAG business Request SHALL contain only `query`. The canonical RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, optional normalized `relevance_level`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`. The RAG projector SHALL accept upstream `relevanceLevel` or `relevance_level` and normalize recognized values without exposing raw relevance scores or retrieval traces.
|
||||
|
||||
#### Scenario: RAG evidence is serialized
|
||||
- **WHEN** a RAG result contains a matching document excerpt and an upstream relevance level
|
||||
- **THEN** its canonical JSON preserves the bounded evidence and normalized `relevance_level` while excluding ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
|
||||
|
||||
#### Scenario: RAG query has no evidence
|
||||
- **WHEN** RAG executes successfully without a usable document excerpt
|
||||
- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, returns an empty evidence list, and does not upgrade relevance into evidence
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Tool results SHALL have separate Harness and model views
|
||||
Each successful evidence Tool result SHALL provide a Harness Control View and a bounded Model Observation derived from the same canonical result. The control view MAY contain returned counts, normalized scope, relevance, truncation and duplicate identity. The Model Observation SHALL contain only fields needed to understand and cite the result and SHALL NOT contain raw responses, internal scores, retrieval traces, duplicate fingerprints, counters, thresholds, budgets or store identities.
|
||||
|
||||
#### Scenario: Model receives RAG observation
|
||||
- **WHEN** a canonical RAG result is READY
|
||||
- **THEN** the model receives Tool Call ID, actual query scope, bounded evidence, evidence status, optional coarse relevance and truncation, but not raw scores or Harness counters
|
||||
|
||||
#### Scenario: Harness evaluates duplicate scope
|
||||
- **WHEN** the same normalized Tool scope is requested again
|
||||
- **THEN** the Harness can compare its control view identity without exposing that fingerprint to the model
|
||||
+12
@@ -0,0 +1,12 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Run progress projection SHALL reference canonical records without duplicating truth
|
||||
The Run progress tracker SHALL retain only ordered canonical identities for completed Tool calls. At Tool-loop completion, a projector SHALL resolve those identities through the existing canonical store and SHALL accept only READY records owned by the current Run. The tracker SHALL NOT store or reconstruct raw Tool responses, complete Agent results, or a second durable evidence record.
|
||||
|
||||
#### Scenario: Completed calls are projected
|
||||
- **WHEN** a Run ends after multiple READY canonical Tool invocations
|
||||
- **THEN** the progress projector reads each indexed canonical record in execution order and creates bounded observed facts
|
||||
|
||||
#### Scenario: Indexed identity is invalid
|
||||
- **WHEN** an indexed canonical identity is missing, expired, cross-Run, PROJECTING, or ERROR
|
||||
- **THEN** it is excluded from observed facts and cannot become verified evidence
|
||||
+110
@@ -0,0 +1,110 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Harness SHALL track binary information gain per Run
|
||||
Each Diagnosis Run SHALL own a thread-safe progress tracker with `GAINED` and `NO_GAIN` as the only information-gain values. `GAINED` SHALL reset the consecutive no-gain count; `NO_GAIN` SHALL increment it. The tracker SHALL expose only `COLLECTING` or `SATURATED` as collection state and SHALL NOT create a second diagnosis lifecycle.
|
||||
|
||||
#### Scenario: New evidence advances diagnosis
|
||||
- **WHEN** a valid pending Tool observation is evaluated as `GAINED`
|
||||
- **THEN** the Run remains `COLLECTING` and its consecutive no-gain count becomes zero
|
||||
|
||||
#### Scenario: Consecutive observations do not advance diagnosis
|
||||
- **WHEN** valid `NO_GAIN` observations reach the Run's configured threshold
|
||||
- **THEN** the tracker becomes `SATURATED` with stop reason `INFORMATION_SATURATED`
|
||||
|
||||
### Requirement: Harness SHALL assign only deterministic no-gain signals
|
||||
The Harness SHALL assign `NO_GAIN` when a successful Tool result has `evidence_status=NO_EVIDENCE` or when a requested `tool_name + normalized_scope` duplicates a successfully completed scope in the same Run. Other successful non-empty results, including RAG `REFERENCE`, SHALL require a model-provided `GAINED` or `NO_GAIN` before another Tool executes.
|
||||
|
||||
#### Scenario: Empty scoped result
|
||||
- **WHEN** a Tool completes READY with `NO_EVIDENCE`
|
||||
- **THEN** the Harness records `NO_GAIN` without asking a model to judge Tool quality
|
||||
|
||||
#### Scenario: Reference material is non-empty
|
||||
- **WHEN** RAG returns bounded evidence with `relevance_level=REFERENCE`
|
||||
- **THEN** the Harness leaves it pending for model evaluation and does not automatically mark it `NO_GAIN`
|
||||
|
||||
#### Scenario: Equivalent structured scope repeats
|
||||
- **WHEN** the model requests the same Tool with the same deterministically normalized successful scope
|
||||
- **THEN** the Harness does not execute the Tool, records `NO_GAIN`, and does not create a second canonical invocation
|
||||
|
||||
### Requirement: Consecutive no-gain threshold SHALL be fixed per Run
|
||||
The system SHALL bind `harness.chat.stop-after-consecutive-no-gain`, require a value of at least one, and default it to `2`. The value SHALL be copied into each new Run's tracker and SHALL NOT enter model context or change an active Run.
|
||||
|
||||
#### Scenario: Default configuration is used
|
||||
- **WHEN** no external value is configured
|
||||
- **THEN** a new Run becomes saturated after two consecutive `NO_GAIN` decisions
|
||||
|
||||
#### Scenario: Invalid threshold is configured
|
||||
- **WHEN** the configured threshold is zero or negative
|
||||
- **THEN** Harness configuration validation fails before serving Chat requests
|
||||
|
||||
### Requirement: Saturated collection SHALL stop further Tool execution
|
||||
When collection becomes saturated, the Harness SHALL reject the pending or next business Tool execution and deliver one bounded `STOP_REQUIRED` observation with `reason=INFORMATION_SATURATED`. If the next model round requests another Tool, the Agent execution SHALL terminate through a typed controlled-stop path without consuming another Tool budget or publishing an internal failure.
|
||||
|
||||
#### Scenario: Model evaluation reaches threshold
|
||||
- **WHEN** the next Tool Call reports `NO_GAIN` and that evaluation reaches the threshold
|
||||
- **THEN** the requested business Tool is not executed and the model receives one STOP_REQUIRED observation
|
||||
|
||||
#### Scenario: Model ignores stop instruction
|
||||
- **WHEN** the model requests another Tool after STOP_REQUIRED was delivered
|
||||
- **THEN** the loop ends as controlled information saturation and proceeds to safe release
|
||||
|
||||
### Requirement: Tool loop completion SHALL project bounded progress once
|
||||
At normal Draft completion, information saturation, or budget termination, the Harness SHALL use the current Run's completed canonical invocation keys to create one bounded `ProgressSnapshot`. The snapshot SHALL contain only safe source, actual scope, objective result summary, truncation and stop reason; it SHALL NOT contain Prompt, thought, raw Tool response, internal counters, remaining budget, Redis keys, or model evaluation rationale.
|
||||
|
||||
#### Scenario: Unknown problem has completed checks
|
||||
- **WHEN** multiple Tools completed but no supported conclusion exists
|
||||
- **THEN** the release input contains their bounded checked scopes and objective results in stable execution order
|
||||
|
||||
#### Scenario: Canonical record cannot be verified
|
||||
- **WHEN** an indexed record is missing, expired, incomplete, ERROR, or not owned by the current Run
|
||||
- **THEN** it is not projected as an observed fact and the snapshot records a bounded limitation
|
||||
|
||||
### Requirement: Information stop reasons SHALL remain distinct from release outcomes
|
||||
The Harness SHALL distinguish `INFORMATION_SATURATED` from `BUDGET_LIMIT_REACHED`. Release SHALL continue to expose only `SUCCESS`, `FALLBACK`, `FAILED`, or `CANCELLED`, and SHALL use SafeFallback type to distinguish insufficient evidence from missing required context.
|
||||
|
||||
#### Scenario: Low gain stops before budget exhaustion
|
||||
- **WHEN** consecutive no-gain reaches the configured threshold while hard budget remains
|
||||
- **THEN** Trace records `INFORMATION_SATURATED` and public release is a normal `FALLBACK`
|
||||
|
||||
#### Scenario: Hard budget terminates collection
|
||||
- **WHEN** model, Tool, Token, byte, or time protection stops the Diagnosis after at least one safe observation
|
||||
- **THEN** Trace retains `BUDGET_LIMIT_REACHED` and Release attempts a deterministic `FALLBACK` from the existing progress without an extra model call
|
||||
|
||||
### Requirement: Diagnosis Prompt SHALL license bounded abandonment
|
||||
The Chinese Diagnosis Prompt SHALL state that a root cause is not mandatory, `conclusion=null` is valid completion, zero Tool calls are allowed when required query context is missing, and correct but non-advancing content is `NO_GAIN`. It SHALL require the model to stop when no distinct bounded query can produce new diagnostic information and to obey STOP_REQUIRED. It SHALL NOT embed Tool names, Tool schemas, thresholds, counters, `next_action`, or Harness implementation details.
|
||||
|
||||
#### Scenario: Required context is missing before any Tool call
|
||||
- **WHEN** the Query lacks the enterprise, time, service, error, or other context needed for a bounded query
|
||||
- **THEN** the model may return `conclusion=null` with `limitations.missing_info` without calling a Tool
|
||||
|
||||
#### Scenario: Correct content is diagnostically useless
|
||||
- **WHEN** a Tool response is factually correct but only generic, repeated, or unable to change a current hypothesis
|
||||
- **THEN** the model treats it as `NO_GAIN` and does not continue with an equivalent query
|
||||
|
||||
### Requirement: Model token audit SHALL be component-scoped and reconcilable
|
||||
Every Harness model call admitted by the Run budget SHALL receive a bounded component and component round. When Provider Usage is available, the same non-negative input and output Token counts SHALL update both the Run budget total and a `MODEL_TOKEN_USAGE` Trace event. Diagnosis Agent usage SHALL also update the matching `AgentStep.token_count`. At Run completion, Trace SHALL expose whether audited Token totals reconcile with the Run budget total and SHALL expose unavailable Usage counts without fabricating Token values.
|
||||
|
||||
The audit SHALL NOT persist Prompt content, user or model text, reasoning content, Tool arguments, raw model responses, credentials, or provider-specific metadata.
|
||||
|
||||
#### Scenario: Diagnosis Agent round returns Usage
|
||||
- **WHEN** a Diagnosis Agent model round returns input and output Token Usage
|
||||
- **THEN** its component round Trace and matching AgentStep contain the same total Token count and the Run total increases by that amount
|
||||
|
||||
#### Scenario: Multiple model components execute
|
||||
- **WHEN** Router, Diagnosis Agent, Evidence Repair or Semantic Guard model calls execute in one Run
|
||||
- **THEN** each call is distinguishable by bounded component and component round and their audited Token sum can be compared with the Run total
|
||||
|
||||
#### Scenario: Provider Usage is unavailable
|
||||
- **WHEN** a model attempt completes or fails without Provider Usage
|
||||
- **THEN** the audit marks Usage unavailable and Run reconciliation exposes the gap without estimating Token counts
|
||||
|
||||
### Requirement: Rejected Tool requests SHALL remain observable without payload disclosure
|
||||
Every supported Tool request rejected by the Harness before a usable business observation is delivered SHALL emit a `TOOL_REQUEST_REJECTED` Trace event containing only safe Tool Call ID, Tool name and stable error code. Rejected requests SHALL remain distinguishable from canonical `TOOL_INVOCATION` events and SHALL NOT include Tool arguments, normalized scope content, raw responses, internal exception messages, credentials or budget values.
|
||||
|
||||
#### Scenario: Progress protocol is invalid
|
||||
- **WHEN** a supported Tool request omits or misorders a required previous observation
|
||||
- **THEN** no business Tool executes and Trace records `INVALID_PROGRESS_PROTOCOL` for that Tool request
|
||||
|
||||
#### Scenario: Duplicate or saturated request is blocked
|
||||
- **WHEN** a supported Tool request repeats a successful normalized scope or arrives after collection saturation
|
||||
- **THEN** Trace records the stable rejection reason while canonical invocation count remains unchanged
|
||||
+33
@@ -0,0 +1,33 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Fixed isolated executors
|
||||
The Application Use Case SHALL map SYSTEM_CHAT to one no-Tool model response, KNOWLEDGE_QUERY to exactly one lookup-knowledge invocation plus one bounded answer model call, and DIAGNOSIS to the single Diagnosis Agent followed by the Diagnosis Release boundary. Executors MUST NOT call one another or rewrite the Query. Information saturation and budget termination in Diagnosis SHALL be converted to safe content by Diagnosis Release, not by ChatApplicationUseCase.
|
||||
|
||||
#### Scenario: System Chat
|
||||
- **WHEN** intent is SYSTEM_CHAT
|
||||
- **THEN** no evidence Tool, Diagnosis Agent, EvidenceGuard, or SemanticGuard is invoked
|
||||
|
||||
#### Scenario: Knowledge Query
|
||||
- **WHEN** intent is KNOWLEDGE_QUERY
|
||||
- **THEN** only lookup_knowledge is invoked once and query_logs/query_mysql/Diagnosis ReAct are unavailable
|
||||
|
||||
#### Scenario: Diagnosis succeeds with a conclusion
|
||||
- **WHEN** intent is DIAGNOSIS and the Draft passes the release guards
|
||||
- **THEN** the original Query and bounded PreviousTurn enter DiagnosisAgentUseCase and the Draft publishes only after DiagnosisReleaseUseCase
|
||||
|
||||
#### Scenario: Diagnosis stops without a conclusion
|
||||
- **WHEN** Diagnosis collection is saturated, required context is missing, or a handled budget limit is reached
|
||||
- **THEN** DiagnosisReleaseUseCase returns bounded Fallback content and ChatApplicationUseCase only persists and transports that decision
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Handled Diagnosis budget termination SHALL remain a Fallback release
|
||||
When Diagnosis Release has converted a recognized budget termination and existing safe progress into a Fallback, ChatApplicationUseCase SHALL persist that result exactly once without reclassifying it as `INTERNAL_FAILURE`, invoking another model, or rebuilding business fallback content. The internal Run lifecycle MAY retain `BUDGET_EXHAUSTED`, while the persisted public release outcome SHALL be `FALLBACK` and no PublishedResult SHALL be stored.
|
||||
|
||||
#### Scenario: Tool budget ends after finite checks
|
||||
- **WHEN** Diagnosis reaches a hard Tool budget after at least one canonical safe observation and Release creates an insufficient-evidence fallback
|
||||
- **THEN** the Run persists status SUCCESS, release outcome FALLBACK, safe content and actual budget usage, and the SSE sends content followed by done
|
||||
|
||||
#### Scenario: Budget ends without safe publishable progress
|
||||
- **WHEN** budget termination occurs before Diagnosis Release can form a safe bounded result
|
||||
- **THEN** the existing failure path remains fail closed and does not fabricate observed facts
|
||||
+56
@@ -0,0 +1,56 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Evidence Tools SHALL execute through the Harness boundary
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions through native Tool Calling. Each Agent-facing Tool input SHALL be a typed Envelope containing optional `previous_observation` and required business `input`. A Tool interceptor SHALL validate and consume the previous observation, propagate the exact framework Tool Call ID, Run ID, Tool name, and unwrapped business JSON into the corresponding adapter and `ToolBoundary`, and SHALL enforce Run progress before execution. Successful observations SHALL use a bounded per-Tool whitelist projection; failed or control observations SHALL contain only stable safe semantics and SHALL NOT contain raw responses, internal exceptions, credentials, invocation lifecycle internals, counters, thresholds, or remaining budget.
|
||||
|
||||
#### Scenario: Framework requests first RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`, no pending evaluation and a typed business input
|
||||
- **THEN** the RAG adapter receives only the unwrapped business request, uses exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
|
||||
#### Scenario: Model continues after a non-empty result
|
||||
- **WHEN** the last successful Tool result is pending semantic evaluation and the model requests another Tool
|
||||
- **THEN** the Envelope must identify that exact prior Tool Call and contain `GAINED` or `NO_GAIN` before the new business Tool can execute
|
||||
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** a registered adapter returns an error result
|
||||
- **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail
|
||||
|
||||
#### Scenario: Unknown Tool is requested
|
||||
- **WHEN** a model requests a Tool outside the three registered definitions
|
||||
- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
|
||||
|
||||
### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
|
||||
The single Chinese Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. The Agent SHALL NOT be required to find a root cause. If current evidence cannot support a diagnosis, no required context exists for a bounded Tool call, or available results are correct but do not advance any diagnosis hypothesis, the Agent SHALL stop with `conclusion=null`, describe actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist, fabricate a root cause, or make equivalent Tool calls merely to show activity.
|
||||
|
||||
#### Scenario: Tool finds no evidence
|
||||
- **WHEN** completed Tool observations have no diagnostic information gain
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records bounded checked scope and missing information, and does not make an equivalent retry
|
||||
|
||||
#### Scenario: No bounded Tool query is possible
|
||||
- **WHEN** the Query lacks required enterprise, time, service or error context
|
||||
- **THEN** the Agent may perform zero Tool calls and returns a no-conclusion Draft whose `limitations.missing_info` identifies the required context
|
||||
|
||||
#### Scenario: Harness requires stop
|
||||
- **WHEN** the Agent receives `STOP_REQUIRED`
|
||||
- **THEN** it emits a bounded final Draft without another Tool call
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Diagnosis execution SHALL preserve controlled stop outcomes
|
||||
The internal Agent use case SHALL distinguish a valid Draft, controlled information saturation, and budget termination from an unclassified Agent failure. It SHALL return a bounded internal execution result containing optional Draft, ProgressSnapshot and stop reason, and SHALL NOT convert a recognized controlled stop into `DiagnosisAgentOutputException`.
|
||||
|
||||
#### Scenario: Model ignores STOP_REQUIRED
|
||||
- **WHEN** the framework surfaces the typed collection-stopped signal after the final completion opportunity
|
||||
- **THEN** the Agent use case returns no Draft with `INFORMATION_SATURATED` and the current ProgressSnapshot
|
||||
|
||||
#### Scenario: Unclassified framework failure
|
||||
- **WHEN** Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination
|
||||
- **THEN** execution still fails closed and no safe progress is fabricated
|
||||
|
||||
#### Scenario: Invalid final Draft after verified checks
|
||||
- **WHEN** the final model text is empty or violates the strict DiagnosisDraft contract after the current Run has completed READY canonical Tool checks
|
||||
- **THEN** the invalid text is discarded, the output failure carries only the bounded ProgressSnapshot and safe failure metadata, and no model repair or loose JSON extraction occurs
|
||||
|
||||
#### Scenario: Invalid final Draft without verified checks
|
||||
- **WHEN** the final model text violates the strict DiagnosisDraft contract before any publishable ProgressSnapshot exists
|
||||
- **THEN** execution remains failed and MUST NOT fabricate missing context, observed facts or a no-conclusion Draft
|
||||
+51
@@ -0,0 +1,51 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Deterministic Draft and evidence validation
|
||||
For a Draft with a non-null Conclusion, the Harness SHALL deterministically reject it unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs. For a Draft with `conclusion=null`, Release SHALL NOT require normal conclusion structure or invoke EvidenceRepair; any supplied Tool references SHALL still resolve to current-Run READY canonical invocations and SHALL obey positive/negative evidence semantics.
|
||||
|
||||
#### Scenario: Duplicate or missing Analysis ID in concluded Draft
|
||||
- **WHEN** a Draft with a Conclusion contains a blank or duplicate Analysis ID
|
||||
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
|
||||
|
||||
#### Scenario: Broken report reference in concluded Draft
|
||||
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
|
||||
- **THEN** EvidenceGuard rejects the Draft before semantic review
|
||||
|
||||
#### Scenario: No-conclusion Draft has valid negative observation
|
||||
- **WHEN** a `conclusion=null` Draft cites a current-Run READY `NO_EVIDENCE` call as `NEGATIVE_OBSERVATION`
|
||||
- **THEN** Release accepts the reference authenticity without running EvidenceRepair or SemanticGuard
|
||||
|
||||
#### Scenario: No-conclusion Draft fabricates a Tool reference
|
||||
- **WHEN** a `conclusion=null` Draft cites a missing, cross-Run, incomplete or ERROR Tool call
|
||||
- **THEN** the reference is excluded and cannot be published as an observed fact
|
||||
|
||||
### Requirement: Fail-closed release policy
|
||||
The release use case SHALL publish the unchanged verified Draft only when a non-null Conclusion passes EvidenceGuard and SemanticGuard returns `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, final semantic technical failure, a valid no-conclusion Draft, information saturation, or handled budget termination. No-conclusion and controlled-stop release SHALL be deterministic from verified references and ProgressSnapshot and MUST NOT invoke a repair or semantic model call. A fallback result MUST NOT include an unsupported Draft, full verified snapshot, internal stop counters, or SemanticGuard reason.
|
||||
|
||||
#### Scenario: Supported report release
|
||||
- **WHEN** EvidenceGuard succeeds for a concluded Draft and SemanticGuard returns `SUPPORTED`
|
||||
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
|
||||
|
||||
#### Scenario: Unsupported report fallback
|
||||
- **WHEN** SemanticGuard returns `UNSUPPORTED`
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot
|
||||
|
||||
#### Scenario: Semantic review remains unavailable
|
||||
- **WHEN** all permitted technical attempts fail
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
|
||||
|
||||
#### Scenario: Evidence validation fallback sources
|
||||
- **WHEN** evidence repair fails or the second EvidenceGuard rejects a concluded Draft
|
||||
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
|
||||
|
||||
#### Scenario: Missing context ends without Tool calls
|
||||
- **WHEN** a valid no-conclusion Draft has no Tool calls and identifies required missing context
|
||||
- **THEN** release outcome is `FALLBACK`, type is `MISSING_REQUIRED_CONTEXT`, and no guard model call occurs
|
||||
|
||||
#### Scenario: Finite checks do not support a conclusion
|
||||
- **WHEN** a no-conclusion Draft or controlled stop has a non-empty verified ProgressSnapshot
|
||||
- **THEN** release outcome is `FALLBACK`, type is `INSUFFICIENT_EVIDENCE`, and observed facts describe only actual completed checks
|
||||
|
||||
#### Scenario: Invalid Draft has publishable progress
|
||||
- **WHEN** the Agent's final Draft is rejected by strict parsing but its bounded ProgressSnapshot contains current-Run verified observed facts
|
||||
- **THEN** Release publishes `FALLBACK` with type `INSUFFICIENT_EVIDENCE` using only that snapshot and MUST NOT use any content from the invalid Draft
|
||||
@@ -0,0 +1,51 @@
|
||||
## 1. Tool Envelope 与进展状态契约
|
||||
|
||||
- [x] 1.1 新增 `InformationGain`、collection state、stop reason、上一轮评价和三个强类型 Agent-facing Tool Envelope,并用序列化/Schema 测试固定 snake_case、必填项和非法枚举拒绝行为
|
||||
- [x] 1.2 在 `ChatHarnessProperties` 增加 `stopAfterConsecutiveNoGain=2`、正数校验和配置绑定测试,并在创建 Run 时固定阈值
|
||||
- [x] 1.3 实现线程安全 `DiagnosisProgressTracker`,覆盖 GAINED 清零、NO_GAIN 累加、待评价 ID 单次消费、错序拒绝、饱和状态和一次 STOP_REQUIRED 交付
|
||||
- [x] 1.4 为 RAG、日志和 MySQL 实现确定性 scope projector,覆盖相同参数去重、不同结构化 scope 放行和失败 scope 不进入成功去重集合
|
||||
|
||||
## 2. Canonical 双视图与 Tool 门禁
|
||||
|
||||
- [x] 2.1 扩展 canonical RAG result 以兼容读取 `relevanceLevel/relevance_level` 并输出规范化 `relevance_level`,补充 REFERENCE 与空 evidence 回归测试
|
||||
- [x] 2.2 实现按 Tool 类型的 Harness Control View 和白名单 Model Observation projector,证明 returned count、内部指纹、raw score、预算和 raw response 不进入模型观察
|
||||
- [x] 2.3 将 `HarnessEvidenceTools` 的模型可见 Schema 切换为三个强类型 Envelope,同时保持 adapter bridge 只接收原业务 request JSON
|
||||
- [x] 2.4 重构 `HarnessToolInterceptor`,按“消费上一轮评价 -> 饱和门禁 -> scope 去重 -> ToolBoundary -> 记录结果”顺序执行,并覆盖 NO_EVIDENCE、REFERENCE、重复 scope、协议错误和 Tool ERROR
|
||||
- [x] 2.5 实现 STOP_REQUIRED 一次收尾机会和 `DiagnosisCollectionStoppedException`,用真实 scripted framework loop 证明忽略停止指令不会继续执行 Tool 或耗尽 Tool 预算
|
||||
|
||||
## 3. ProgressSnapshot 与 Agent 执行结果
|
||||
|
||||
- [x] 3.1 让 tracker 有序索引当前 Run 已完成的 canonical identity,并实现 `DiagnosisProgressProjector` 逐条验证 READY/Run ownership 后生成有界 snapshot
|
||||
- [x] 3.2 覆盖 canonical 记录缺失、过期、PROJECTING、ERROR、跨 Run 和投影超限,确保它们只产生 limitation 而不成为 observed fact
|
||||
- [x] 3.3 将 `DiagnosisAgentUseCase` 返回值改为可选 Draft、ProgressSnapshot 和 stop reason,更新全部调用方并保留未分类框架异常 fail closed
|
||||
- [x] 3.4 让预算/取消异常在 cause chain 中保持可分类:预算可进入统一 Release,取消保持 CANCELLED,其他异常不得伪装为正常证据不足
|
||||
|
||||
## 4. 统一 Diagnosis Release
|
||||
|
||||
- [x] 4.1 为无结论 Draft 增加确定性引用真实性/负向语义校验路径,证明它不调用 EvidenceRepair 或 SemanticGuard
|
||||
- [x] 4.2 扩展 `SafeFallbackFactory`,从 ProgressSnapshot 生成 `INSUFFICIENT_EVIDENCE` 或从 `limitations.missing_info` 生成 `MISSING_REQUIRED_CONTEXT`,复用现有 observed facts、sources、limitations 和 next steps
|
||||
- [x] 4.3 重构 `DiagnosisReleaseUseCase`,统一处理有结论 Draft、主动无结论、信息饱和、预算终止,并保留有结论时完整 EvidenceGuard/Repair/SemanticGuard 行为
|
||||
- [x] 4.4 更新 `DiagnosisChatExecutor` 与内部 result contract,使已处理的饱和/预算 Fallback 返回 Application 而不创建 PublishedResult
|
||||
- [x] 4.5 从 `ChatApplicationUseCase` 移除内容级预算 Fallback 决策,保留窄化的已处理预算终态持久化,并迁移现有预算耗尽测试意图
|
||||
- [x] 4.6 对最终 Draft 合同失败增加窄化降级:丢弃非法内容,仅在 ProgressSnapshot 含已验真 observed facts 时由 Release 生成 `INSUFFICIENT_EVIDENCE`,无安全进展仍 fail closed,并记录脱敏失败 Trace
|
||||
|
||||
## 5. Prompt、Trace 与配置装配
|
||||
|
||||
- [x] 5.1 按最小中文契约重写 Diagnosis Prompt:合法放弃、零 Tool、missing_info、GAINED/NO_GAIN、禁止等价重试和服从 STOP_REQUIRED,且不写 Tool 名/Schema/阈值/next_action
|
||||
- [x] 5.2 增加有界 progress/stop Trace 事件,记录信息增益生产者、scope 摘要、collection state 和 stop reason,并用测试证明不记录 Prompt、thought、raw payload、SQL 参数或预算余量
|
||||
- [x] 5.3 更新 `HarnessChatConfiguration` 装配 tracker factory、scope/projector、ProgressSnapshot 和统一 Release 依赖,修复全部构造调用与 Spring context 测试
|
||||
|
||||
## 6. 验证与文档收口
|
||||
|
||||
- [x] 6.1 运行 Tool contract、tracker、projector、interceptor、Agent loop、Guard、Release 和 Application focused tests,修复所有回归
|
||||
- [x] 6.2 运行 Harness 综合回归、`mvn test` 和 `openspec validate diagnosis-information-gain-stop-contract --strict`,记录命令与结果
|
||||
- [x] 6.3 用 Maven 启动真实应用,对“诊断切换企业失败的问题”执行 named SSE E2E,并从 `logs/` 与 `scripts/query_mysql.py` 按 exact sessionId/runId 核对过程、停止原因、Tool 次数和 FALLBACK 内容
|
||||
- [x] 6.4 更新 ISS-016、架构文档、glossary 和 OpenSpec task 状态,核对公开 SSE/前端协议无新增字段且工作区原有无关修改未被覆盖
|
||||
|
||||
## 7. 模型 Token 与 Tool 拒绝审计补充
|
||||
|
||||
- [x] 7.1 新增 Run 内模型调用组件/轮次账本和脱敏 Trace contract 测试,覆盖 Usage 可用、不可用、组件轮次和 Token 合计
|
||||
- [x] 7.2 将 Router、System Chat、Knowledge Answer、Diagnosis Agent、Evidence Repair 和 Semantic Guard 接入统一 Token 审计,并回填对应 AgentStep token_count
|
||||
- [x] 7.3 为进展协议错误、重复 scope、信息饱和和观察合同拒绝增加安全 `TOOL_REQUEST_REJECTED` Trace,证明不记录参数或原始响应
|
||||
- [x] 7.4 在 Run 完成时持久化 step_count 并记录 Token 总账/明细对账结果,补充 Store 和 Trace 回归测试
|
||||
- [x] 7.5 运行 focused、Harness 全量回归、strict validate 和真实 SSE/数据库 E2E,核对 Tool 执行/拒绝数量、Token 对账和 Trace 连续性
|
||||
Reference in New Issue
Block a user