446 lines
14 KiB
Markdown
446 lines
14 KiB
Markdown
# Chat Evidence Pipeline Contracts
|
||
|
||
**状态**:当前实现
|
||
**更新日期**:2026-07-17
|
||
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
|
||
|
||
当前 Chat 复杂诊断链路是:
|
||
|
||
```text
|
||
chat_planner
|
||
-> chat_executor
|
||
-> GatekeeperNode / ExecutorGatekeeperService
|
||
-> VerifiedInputNode
|
||
-> chat_verifier
|
||
-> chat_composer
|
||
-> final answer
|
||
```
|
||
|
||
设计原则:
|
||
|
||
- Planner 暂不输出 `scope_contract`。
|
||
- Executor 只做证据收集和微观事实提炼,不生成最终用户答案。
|
||
- Gatekeeper 在 Verifier 前做代码级引用真实性校验。
|
||
- Verifier 判断 claim 是否能由已核验证据推出。
|
||
- Composer 只表达 Verifier 允许输出的内容。
|
||
|
||
---
|
||
|
||
## 1. Planner
|
||
|
||
Planner 当前保持不变,输出 `planner_plan`:
|
||
|
||
```json
|
||
{
|
||
"selected_skill": "diagnose-mysql-connection-pool",
|
||
"selection_reason": "选择该 skill 的原因",
|
||
"plan": ["步骤1", "步骤2"],
|
||
"reasoning": "规划思路"
|
||
}
|
||
```
|
||
|
||
字段定义:
|
||
|
||
| 字段 | 类型 | 定义 |
|
||
|---|---|---|
|
||
| `selected_skill` | string/null | Planner 选择的诊断 skill 名称 |
|
||
| `selection_reason` | string | skill 选择理由 |
|
||
| `plan` | array | 给 Executor 的执行步骤 |
|
||
| `reasoning` | string | 规划思路说明 |
|
||
|
||
当前边界:
|
||
|
||
- 不新增 `scope_contract`。
|
||
- 不要求 Planner 显式列出 forbidden actions。
|
||
- 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。
|
||
|
||
---
|
||
|
||
## 2. Executor
|
||
|
||
Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。
|
||
|
||
### 2.1 输出结构
|
||
|
||
```json
|
||
{
|
||
"answer_version": "executor_evidence_v2",
|
||
"claims": [
|
||
{
|
||
"claim_id": "claim-1",
|
||
"claim_type": "observation",
|
||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||
"support_level": "direct",
|
||
"evidence_bindings": [
|
||
{
|
||
"source_type": "tool_trace",
|
||
"source_id": "",
|
||
"tool_name": "query_metrics",
|
||
"source_invocation_id": 517,
|
||
"raw_path": "$.alerts[0]",
|
||
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"hypotheses": [],
|
||
"recommended_actions": [],
|
||
"missing_info": []
|
||
}
|
||
```
|
||
|
||
字段定义:
|
||
|
||
| 字段 | 类型 | 必填 | 定义 |
|
||
|---|---|---:|---|
|
||
| `answer_version` | string | 是 | 固定为 `executor_evidence_v2` |
|
||
| `claims` | array | 是 | Executor 提出的待验证事实断言 |
|
||
| `claims[].claim_id` | string | 是 | claim 标识 |
|
||
| `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 |
|
||
| `claims[].claim_text` | string | 是 | 事实断言文本 |
|
||
| `claims[].support_level` | string | 是 | `direct` 或 `indirect` |
|
||
| `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 |
|
||
| `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` |
|
||
| `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 |
|
||
| `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 |
|
||
| `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 |
|
||
| `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 |
|
||
| `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 |
|
||
| `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact |
|
||
| `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 |
|
||
| `missing_info` | array | 是 | 无法确认结论所缺少的证据 |
|
||
|
||
禁止字段:
|
||
|
||
- `diagnosis_summary`
|
||
- `user_facing_answer`
|
||
- `source_invocation_ids` 作为主引用字段
|
||
|
||
### 2.2 raw_path
|
||
|
||
当前支持的精确路径:
|
||
|
||
| 工具 | 正向证据路径 | 负向证据路径 |
|
||
|---|---|---|
|
||
| `query_metrics` | `$.alerts[i]` | `$.no_evidence` |
|
||
| `query_logs` | `$.logs[i]` | `$.no_evidence` |
|
||
| `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` |
|
||
|
||
约束:
|
||
|
||
- `raw_path` 必须指向数组条目或 `$.no_evidence`。
|
||
- 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。
|
||
- 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。
|
||
|
||
### 2.3 negative_observation
|
||
|
||
当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`:
|
||
|
||
```json
|
||
{
|
||
"claim_id": "claim-1",
|
||
"claim_type": "negative_observation",
|
||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||
"support_level": "direct",
|
||
"evidence_bindings": [
|
||
{
|
||
"tool_name": "query_logs",
|
||
"source_invocation_id": 517,
|
||
"raw_path": "$.no_evidence",
|
||
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
语义边界:
|
||
|
||
- `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。
|
||
- 不表示“问题绝对不存在”。
|
||
- 不表示“根因被排除”。
|
||
- 不表示“系统已经健康”。
|
||
- `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。
|
||
|
||
### 2.4 窄范围任务
|
||
|
||
窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。
|
||
|
||
Executor 必须遵守:
|
||
|
||
- 只输出 `observation` / `negative_observation`。
|
||
- claim 数量通常 1 条,最多 2 条。
|
||
- claim 数量限制不限制 `evidence_bindings` 数量。
|
||
- 不输出根因、风险、修复建议、经验推断。
|
||
- 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。
|
||
- 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。
|
||
|
||
---
|
||
|
||
## 3. Tool Invocation Evidence Refs
|
||
|
||
工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。
|
||
|
||
### 3.1 正向证据
|
||
|
||
```json
|
||
{
|
||
"evidence_status": "supported",
|
||
"evidence_refs": [
|
||
{
|
||
"raw_path": "$.logs[0]",
|
||
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
### 3.2 负向证据
|
||
|
||
```json
|
||
{
|
||
"evidence_status": "no_evidence",
|
||
"evidence_refs": [
|
||
{
|
||
"raw_path": "$.no_evidence",
|
||
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
字段定义:
|
||
|
||
| 字段 | 类型 | 定义 |
|
||
|---|---|---|
|
||
| `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` |
|
||
| `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 |
|
||
| `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 |
|
||
|
||
---
|
||
|
||
## 4. Gatekeeper
|
||
|
||
Gatekeeper 是 StateGraph 中的显式 Node,调用 `ExecutorGatekeeperService` 对当前 Run 的证据引用做代码级真实性校验。
|
||
|
||
### 4.1 输入
|
||
|
||
- `sessionId + runId`(来自 `RunnableConfig`,工具查询以 `runId` 为边界)
|
||
- `executor_structured_output`
|
||
- 当前 run 的 `tool_invocation`
|
||
|
||
### 4.2 输出
|
||
|
||
```json
|
||
{
|
||
"status": "pass",
|
||
"severity": "none",
|
||
"rule_set_version": "gatekeeper-rules-v1",
|
||
"rules": [
|
||
{
|
||
"id": "evidence.raw_path",
|
||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||
"enabled": true,
|
||
"default_severity": "reject"
|
||
}
|
||
],
|
||
"checked_bindings": [
|
||
{
|
||
"claim_id": "claim-1",
|
||
"tool_name": "query_logs",
|
||
"source_invocation_id": 517,
|
||
"raw_path": "$.no_evidence",
|
||
"matched_text": "query_logs returned no evidence; ...",
|
||
"status": "pass"
|
||
}
|
||
],
|
||
"failed_rules": [],
|
||
"warnings": [],
|
||
"errors": []
|
||
}
|
||
```
|
||
|
||
字段定义:
|
||
|
||
| 字段 | 类型 | 定义 |
|
||
|---|---|---|
|
||
| `status` | string | `pass` 或 `fail` |
|
||
| `severity` | string | `none`、`low_confid`、`reject` |
|
||
| `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 |
|
||
| `rules` | array | 已启用规则的轻量元数据摘要 |
|
||
| `checked_bindings` | array | 每条证据绑定的校验结果 |
|
||
| `failed_rules` | array | 失败规则 id |
|
||
| `warnings` | array | 自动回填等非阻断信息 |
|
||
| `errors` | array | 失败明细 |
|
||
|
||
校验规则:
|
||
|
||
- `answer_version` 必须是 `executor_evidence_v2`。
|
||
- 不允许 `diagnosis_summary` / `user_facing_answer`。
|
||
- 每个 claim 必须有非空 `evidence_bindings`。
|
||
- `tool_name` 必须和真实 invocation 对齐。
|
||
- `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。
|
||
- `raw_path` 必须存在于 `retrieval_details.evidence_refs`。
|
||
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
|
||
- `negative_observation` 只能绑定 `$.no_evidence`。
|
||
|
||
规则配置:
|
||
|
||
- 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。
|
||
- 规则实现仍是确定性 Java 代码,不执行动态脚本。
|
||
- 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。
|
||
|
||
失败分级:
|
||
|
||
| 场景 | severity |
|
||
|---|---|
|
||
| 伪造 invocation id | `reject` |
|
||
| tool_name 与 invocation 不匹配 | `reject` |
|
||
| raw_path 不存在 | `reject` |
|
||
| excerpt 与 matched_text 不匹配 | `reject` |
|
||
| negative_observation 绑定正向日志 | `reject` |
|
||
| 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` |
|
||
| 旧 invocation 没有 `evidence_refs` | `low_confid` |
|
||
|
||
---
|
||
|
||
## 5. Verifier
|
||
|
||
`VerifiedInputNode` 只保留 Gatekeeper 检查通过的 claim/binding,并为 Verifier 构造最小输入:
|
||
|
||
```json
|
||
{
|
||
"diagnosis_context": {
|
||
"query": "用户原始问题"
|
||
},
|
||
"verified_executor_output": {
|
||
"answer_version": "executor_evidence_v2",
|
||
"claims": []
|
||
},
|
||
"verified_evidence": [],
|
||
"gatekeeper_audit": {},
|
||
"verdict_ceiling": "PASS",
|
||
"retry_context": null
|
||
}
|
||
```
|
||
|
||
Verifier 职责:
|
||
|
||
- 不调用工具。
|
||
- 不读 skill。
|
||
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
|
||
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
|
||
- 不读取 Executor 原始答复或完整工具 Trace,只读取 verified projection。
|
||
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
|
||
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
|
||
|
||
输出:
|
||
|
||
```json
|
||
{
|
||
"verdict": "PASS",
|
||
"groundedness_score": 1.0,
|
||
"critical_fact_count": 1,
|
||
"claim_checks": [],
|
||
"facts_checked": [],
|
||
"rationale": "..."
|
||
}
|
||
```
|
||
|
||
---
|
||
|
||
## 6. Composer
|
||
|
||
Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。
|
||
|
||
输入概念:
|
||
|
||
| 字段 | 定义 |
|
||
|---|---|
|
||
| `original_query` | 用户原始问题 |
|
||
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
|
||
| `allowed_claims` | Verifier 允许表达的 claims |
|
||
| `allowed_hypotheses` | Verifier 允许表达的假设 |
|
||
| `missing_info` | 证据缺口 |
|
||
| `recommended_actions` | 允许表达的建议动作 |
|
||
| `rationale` | Verifier 判定理由 |
|
||
|
||
输出:
|
||
|
||
```json
|
||
{
|
||
"answer_summary": "一句话概括",
|
||
"recommended_actions": [
|
||
{
|
||
"action_text": "下一步动作",
|
||
"reason": "原因"
|
||
}
|
||
],
|
||
"user_facing_answer": "最终给用户看的中文答案"
|
||
}
|
||
```
|
||
|
||
表达边界:
|
||
|
||
- Composer 不补事实、不补根因、不调用工具。
|
||
- 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。
|
||
- 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
|
||
- 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
|
||
|
||
---
|
||
|
||
## 7. Trace Persistence
|
||
|
||
`diagnosis_run.self_evaluation.verifier_evaluation` 持久化:
|
||
|
||
```json
|
||
{
|
||
"verifier_evaluation": {
|
||
"verdict": "PASS",
|
||
"groundedness_score": 1.0,
|
||
"critical_fact_count": 1,
|
||
"claim_checks": [],
|
||
"facts_checked": [],
|
||
"rationale": "...",
|
||
"round": 1,
|
||
"traceability_version": "v1",
|
||
"executor_output_parse_status": {},
|
||
"executor_structured_output": {},
|
||
"gatekeeper_result": {
|
||
"rule_set_version": "gatekeeper-rules-v1"
|
||
},
|
||
"composer_output": {},
|
||
"verified_evidence": []
|
||
}
|
||
}
|
||
```
|
||
|
||
Trace API 可用于回放:
|
||
|
||
- Executor 输出了哪些 claim。
|
||
- 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。
|
||
- Gatekeeper 是否通过、是否自动回填。
|
||
- Verifier 如何判断可推导性。
|
||
- Composer 最终如何表达给用户。
|
||
- `run.orchestrationTrace` 如何经过条件边、有限重试并终止。
|
||
|
||
历史 Run/fixture 的 `verifier_evaluation.tool_trace_summary` 仍可被 Trace UI 或离线评测只读解析,但它是旧链路兼容字段,不是当前 Verifier 输入,也不再由生产链路生成。
|
||
|
||
---
|
||
|
||
## 8. 当前已验证样例
|
||
|
||
| 场景 | sessionId | 结果 |
|
||
|---|---|---|
|
||
| HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim |
|
||
| HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 |
|
||
|
||
---
|
||
|
||
## 9. 仍需记录或后续补强
|
||
|
||
当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
|
||
|
||
1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
|
||
2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。
|
||
3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
|