Files
SuperBizAgent-java/mvp/architecture/executor-evidence-pipeline-refactor.md

446 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chat Evidence Pipeline Contracts
**状态**:当前实现
**更新日期**:2026-07-17
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
当前 Chat 复杂诊断链路是:
```text
chat_planner
-> chat_executor
-> GatekeeperNode / ExecutorGatekeeperService
-> VerifiedInputNode
-> chat_verifier
-> chat_composer
-> final answer
```
设计原则:
- Planner 暂不输出 `scope_contract`。
- Executor 只做证据收集和微观事实提炼,不生成最终用户答案。
- Gatekeeper 在 Verifier 前做代码级引用真实性校验。
- Verifier 判断 claim 是否能由已核验证据推出。
- Composer 只表达 Verifier 允许输出的内容。
---
## 1. Planner
Planner 当前保持不变,输出 `planner_plan`:
```json
{
"selected_skill": "diagnose-mysql-connection-pool",
"selection_reason": "选择该 skill 的原因",
"plan": ["步骤1", "步骤2"],
"reasoning": "规划思路"
}
```
字段定义:
| 字段 | 类型 | 定义 |
|---|---|---|
| `selected_skill` | string/null | Planner 选择的诊断 skill 名称 |
| `selection_reason` | string | skill 选择理由 |
| `plan` | array | 给 Executor 的执行步骤 |
| `reasoning` | string | 规划思路说明 |
当前边界:
- 不新增 `scope_contract`。
- 不要求 Planner 显式列出 forbidden actions。
- 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。
---
## 2. Executor
Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。
### 2.1 输出结构
```json
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "observation",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "",
"tool_name": "query_metrics",
"source_invocation_id": 517,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
```
字段定义:
| 字段 | 类型 | 必填 | 定义 |
|---|---|---:|---|
| `answer_version` | string | 是 | 固定为 `executor_evidence_v2` |
| `claims` | array | 是 | Executor 提出的待验证事实断言 |
| `claims[].claim_id` | string | 是 | claim 标识 |
| `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 |
| `claims[].claim_text` | string | 是 | 事实断言文本 |
| `claims[].support_level` | string | 是 | `direct` 或 `indirect` |
| `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 |
| `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` |
| `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 |
| `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 |
| `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 |
| `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 |
| `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 |
| `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact |
| `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 |
| `missing_info` | array | 是 | 无法确认结论所缺少的证据 |
禁止字段:
- `diagnosis_summary`
- `user_facing_answer`
- `source_invocation_ids` 作为主引用字段
### 2.2 raw_path
当前支持的精确路径:
| 工具 | 正向证据路径 | 负向证据路径 |
|---|---|---|
| `query_metrics` | `$.alerts[i]` | `$.no_evidence` |
| `query_logs` | `$.logs[i]` | `$.no_evidence` |
| `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` |
约束:
- `raw_path` 必须指向数组条目或 `$.no_evidence`。
- 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。
- 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。
### 2.3 negative_observation
当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`:
```json
{
"claim_id": "claim-1",
"claim_type": "negative_observation",
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"support_level": "direct",
"evidence_bindings": [
{
"tool_name": "query_logs",
"source_invocation_id": 517,
"raw_path": "$.no_evidence",
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
}
]
}
```
语义边界:
- `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。
- 不表示“问题绝对不存在”。
- 不表示“根因被排除”。
- 不表示“系统已经健康”。
- `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。
### 2.4 窄范围任务
窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。
Executor 必须遵守:
- 只输出 `observation` / `negative_observation`。
- claim 数量通常 1 条,最多 2 条。
- claim 数量限制不限制 `evidence_bindings` 数量。
- 不输出根因、风险、修复建议、经验推断。
- 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。
- 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。
---
## 3. Tool Invocation Evidence Refs
工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。
### 3.1 正向证据
```json
{
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.logs[0]",
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
}
]
}
```
### 3.2 负向证据
```json
{
"evidence_status": "no_evidence",
"evidence_refs": [
{
"raw_path": "$.no_evidence",
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
}
]
}
```
字段定义:
| 字段 | 类型 | 定义 |
|---|---|---|
| `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` |
| `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 |
| `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 |
---
## 4. Gatekeeper
Gatekeeper 是 StateGraph 中的显式 Node,调用 `ExecutorGatekeeperService` 对当前 Run 的证据引用做代码级真实性校验。
### 4.1 输入
- `sessionId + runId`(来自 `RunnableConfig`,工具查询以 `runId` 为边界)
- `executor_structured_output`
- 当前 run 的 `tool_invocation`
### 4.2 输出
```json
{
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_logs",
"source_invocation_id": 517,
"raw_path": "$.no_evidence",
"matched_text": "query_logs returned no evidence; ...",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
}
```
字段定义:
| 字段 | 类型 | 定义 |
|---|---|---|
| `status` | string | `pass` 或 `fail` |
| `severity` | string | `none`、`low_confid`、`reject` |
| `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 |
| `rules` | array | 已启用规则的轻量元数据摘要 |
| `checked_bindings` | array | 每条证据绑定的校验结果 |
| `failed_rules` | array | 失败规则 id |
| `warnings` | array | 自动回填等非阻断信息 |
| `errors` | array | 失败明细 |
校验规则:
- `answer_version` 必须是 `executor_evidence_v2`。
- 不允许 `diagnosis_summary` / `user_facing_answer`。
- 每个 claim 必须有非空 `evidence_bindings`。
- `tool_name` 必须和真实 invocation 对齐。
- `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。
- `raw_path` 必须存在于 `retrieval_details.evidence_refs`。
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
- `negative_observation` 只能绑定 `$.no_evidence`。
规则配置:
- 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。
- 规则实现仍是确定性 Java 代码,不执行动态脚本。
- 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。
失败分级:
| 场景 | severity |
|---|---|
| 伪造 invocation id | `reject` |
| tool_name 与 invocation 不匹配 | `reject` |
| raw_path 不存在 | `reject` |
| excerpt 与 matched_text 不匹配 | `reject` |
| negative_observation 绑定正向日志 | `reject` |
| 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` |
| 旧 invocation 没有 `evidence_refs` | `low_confid` |
---
## 5. Verifier
`VerifiedInputNode` 只保留 Gatekeeper 检查通过的 claim/binding,并为 Verifier 构造最小输入:
```json
{
"diagnosis_context": {
"query": "用户原始问题"
},
"verified_executor_output": {
"answer_version": "executor_evidence_v2",
"claims": []
},
"verified_evidence": [],
"gatekeeper_audit": {},
"verdict_ceiling": "PASS",
"retry_context": null
}
```
Verifier 职责:
- 不调用工具。
- 不读 skill。
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
- 不读取 Executor 原始答复或完整工具 Trace,只读取 verified projection。
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
输出:
```json
{
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"claim_checks": [],
"facts_checked": [],
"rationale": "..."
}
```
---
## 6. Composer
Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。
输入概念:
| 字段 | 定义 |
|---|---|
| `original_query` | 用户原始问题 |
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
| `allowed_claims` | Verifier 允许表达的 claims |
| `allowed_hypotheses` | Verifier 允许表达的假设 |
| `missing_info` | 证据缺口 |
| `recommended_actions` | 允许表达的建议动作 |
| `rationale` | Verifier 判定理由 |
输出:
```json
{
"answer_summary": "一句话概括",
"recommended_actions": [
{
"action_text": "下一步动作",
"reason": "原因"
}
],
"user_facing_answer": "最终给用户看的中文答案"
}
```
表达边界:
- Composer 不补事实、不补根因、不调用工具。
- 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。
- 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
- 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
---
## 7. Trace Persistence
`diagnosis_run.self_evaluation.verifier_evaluation` 持久化:
```json
{
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"claim_checks": [],
"facts_checked": [],
"rationale": "...",
"round": 1,
"traceability_version": "v1",
"executor_output_parse_status": {},
"executor_structured_output": {},
"gatekeeper_result": {
"rule_set_version": "gatekeeper-rules-v1"
},
"composer_output": {},
"verified_evidence": []
}
}
```
Trace API 可用于回放:
- Executor 输出了哪些 claim。
- 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。
- Gatekeeper 是否通过、是否自动回填。
- Verifier 如何判断可推导性。
- Composer 最终如何表达给用户。
- `run.orchestrationTrace` 如何经过条件边、有限重试并终止。
当前版本不生成、读取或展示 `verifier_evaluation.tool_trace_summary`;Verifier 只消费 verified projection。
---
## 8. 当前已验证样例
| 场景 | sessionId | 结果 |
|---|---|---|
| HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim |
| HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 |
---
## 9. 仍需记录或后续补强
当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。
3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。