# Chat Evidence Pipeline Contracts **状态**:当前实现 **更新日期**:2026-07-08 **范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约 当前 Chat 复杂诊断链路是: ```text chat_planner -> chat_executor -> VerifierInputHook / ExecutorGatekeeperService -> chat_verifier -> chat_composer -> final answer ``` 设计原则: - Planner 暂不输出 `scope_contract`。 - Executor 只做证据收集和微观事实提炼,不生成最终用户答案。 - Gatekeeper 在 Verifier 前做代码级引用真实性校验。 - Verifier 判断 claim 是否能由已核验证据推出。 - Composer 只表达 Verifier 允许输出的内容。 --- ## 1. Planner Planner 当前保持不变,输出 `planner_plan`: ```json { "selected_skill": "diagnose-mysql-connection-pool", "selection_reason": "选择该 skill 的原因", "plan": ["步骤1", "步骤2"], "reasoning": "规划思路" } ``` 字段定义: | 字段 | 类型 | 定义 | |---|---|---| | `selected_skill` | string/null | Planner 选择的诊断 skill 名称 | | `selection_reason` | string | skill 选择理由 | | `plan` | array | 给 Executor 的执行步骤 | | `reasoning` | string | 规划思路说明 | 当前边界: - 不新增 `scope_contract`。 - 不要求 Planner 显式列出 forbidden actions。 - 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。 --- ## 2. Executor Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。 ### 2.1 输出结构 ```json { "answer_version": "executor_evidence_v2", "claims": [ { "claim_id": "claim-1", "claim_type": "observation", "claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。", "support_level": "direct", "evidence_bindings": [ { "source_type": "tool_trace", "source_id": "", "tool_name": "query_metrics", "source_invocation_id": 517, "raw_path": "$.alerts[0]", "evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m" } ] } ], "hypotheses": [], "recommended_actions": [], "missing_info": [] } ``` 字段定义: | 字段 | 类型 | 必填 | 定义 | |---|---|---:|---| | `answer_version` | string | 是 | 固定为 `executor_evidence_v2` | | `claims` | array | 是 | Executor 提出的待验证事实断言 | | `claims[].claim_id` | string | 是 | claim 标识 | | `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 | | `claims[].claim_text` | string | 是 | 事实断言文本 | | `claims[].support_level` | string | 是 | `direct` 或 `indirect` | | `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 | | `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` | | `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 | | `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 | | `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 | | `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 | | `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 | | `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact | | `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 | | `missing_info` | array | 是 | 无法确认结论所缺少的证据 | 禁止字段: - `diagnosis_summary` - `user_facing_answer` - `source_invocation_ids` 作为主引用字段 ### 2.2 raw_path 当前支持的精确路径: | 工具 | 正向证据路径 | 负向证据路径 | |---|---|---| | `query_metrics` | `$.alerts[i]` | `$.no_evidence` | | `query_logs` | `$.logs[i]` | `$.no_evidence` | | `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` | 约束: - `raw_path` 必须指向数组条目或 `$.no_evidence`。 - 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。 - 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。 ### 2.3 negative_observation 当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`: ```json { "claim_id": "claim-1", "claim_type": "negative_observation", "claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。", "support_level": "direct", "evidence_bindings": [ { "tool_name": "query_logs", "source_invocation_id": 517, "raw_path": "$.no_evidence", "evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence" } ] } ``` 语义边界: - `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。 - 不表示“问题绝对不存在”。 - 不表示“根因被排除”。 - 不表示“系统已经健康”。 - `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。 ### 2.4 窄范围任务 窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。 Executor 必须遵守: - 只输出 `observation` / `negative_observation`。 - claim 数量通常 1 条,最多 2 条。 - claim 数量限制不限制 `evidence_bindings` 数量。 - 不输出根因、风险、修复建议、经验推断。 - 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。 - 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。 --- ## 3. Tool Invocation Evidence Refs 工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。 ### 3.1 正向证据 ```json { "evidence_status": "supported", "evidence_refs": [ { "raw_path": "$.logs[0]", "text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..." } ] } ``` ### 3.2 负向证据 ```json { "evidence_status": "no_evidence", "evidence_refs": [ { "raw_path": "$.no_evidence", "text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志" } ] } ``` 字段定义: | 字段 | 类型 | 定义 | |---|---|---| | `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` | | `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 | | `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 | --- ## 4. Gatekeeper Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。 ### 4.1 输入 - `sessionId` - `executor_structured_output` - 当前 session 的 `tool_invocation` ### 4.2 输出 ```json { "status": "pass", "severity": "none", "rule_set_version": "gatekeeper-rules-v1", "rules": [ { "id": "evidence.raw_path", "description": "raw_path must exist in retrieval_details.evidence_refs", "enabled": true, "default_severity": "reject" } ], "checked_bindings": [ { "claim_id": "claim-1", "tool_name": "query_logs", "source_invocation_id": 517, "raw_path": "$.no_evidence", "matched_text": "query_logs returned no evidence; ...", "status": "pass" } ], "failed_rules": [], "warnings": [], "errors": [] } ``` 字段定义: | 字段 | 类型 | 定义 | |---|---|---| | `status` | string | `pass` 或 `fail` | | `severity` | string | `none`、`low_confid`、`reject` | | `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 | | `rules` | array | 已启用规则的轻量元数据摘要 | | `checked_bindings` | array | 每条证据绑定的校验结果 | | `failed_rules` | array | 失败规则 id | | `warnings` | array | 自动回填等非阻断信息 | | `errors` | array | 失败明细 | 校验规则: - `answer_version` 必须是 `executor_evidence_v2`。 - 不允许 `diagnosis_summary` / `user_facing_answer`。 - 每个 claim 必须有非空 `evidence_bindings`。 - `tool_name` 必须和真实 invocation 对齐。 - `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。 - `raw_path` 必须存在于 `retrieval_details.evidence_refs`。 - `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。 - `negative_observation` 只能绑定 `$.no_evidence`。 规则配置: - 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。 - 规则实现仍是确定性 Java 代码,不执行动态脚本。 - 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。 失败分级: | 场景 | severity | |---|---| | 伪造 invocation id | `reject` | | tool_name 与 invocation 不匹配 | `reject` | | raw_path 不存在 | `reject` | | excerpt 与 matched_text 不匹配 | `reject` | | negative_observation 绑定正向日志 | `reject` | | 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` | | 旧 invocation 没有 `evidence_refs` | `low_confid` | --- ## 5. Verifier Verifier 输入由 `VerifierInputHook` 构造: ```json { "original_query": "用户原始问题", "executor_final_answer": "{...executor raw text for debug/fallback only...}", "executor_structured_output": { "answer_version": "executor_evidence_v2", "claims": [] }, "executor_output_parse_status": { "status": "valid", "detail": "parsed executor evidence contract" }, "tool_trace_summary": [], "gatekeeper_result": {}, "retry_context": null } ``` Verifier 职责: - 不调用工具。 - 不读 skill。 - 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。 - 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。 - 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。 - 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。 - 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。 输出: ```json { "verdict": "PASS", "groundedness_score": 1.0, "critical_fact_count": 1, "claim_checks": [], "facts_checked": [], "rationale": "..." } ``` --- ## 6. Composer Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。 输入概念: | 字段 | 定义 | |---|---| | `original_query` | 用户原始问题 | | `verdict` | `PASS` / `LOW_CONFID` / `REJECT` | | `allowed_claims` | Verifier 允许表达的 claims | | `allowed_hypotheses` | Verifier 允许表达的假设 | | `missing_info` | 证据缺口 | | `recommended_actions` | 允许表达的建议动作 | | `rationale` | Verifier 判定理由 | 输出: ```json { "answer_summary": "一句话概括", "recommended_actions": [ { "action_text": "下一步动作", "reason": "原因" } ], "user_facing_answer": "最终给用户看的中文答案" } ``` 表达边界: - Composer 不补事实、不补根因、不调用工具。 - 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。 - 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。 - 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。 --- ## 7. Trace Persistence `diagnosis_session.self_evaluation.verifier_evaluation` 持久化: ```json { "verifier_evaluation": { "verdict": "PASS", "groundedness_score": 1.0, "critical_fact_count": 1, "claim_checks": [], "facts_checked": [], "rationale": "...", "round": 1, "traceability_version": "v1", "executor_output_parse_status": {}, "executor_structured_output": {}, "gatekeeper_result": { "rule_set_version": "gatekeeper-rules-v1" }, "composer_output": {}, "tool_trace_summary": [] } } ``` Trace API 可用于回放: - Executor 输出了哪些 claim。 - 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。 - Gatekeeper 是否通过、是否自动回填。 - Verifier 如何判断可推导性。 - Composer 最终如何表达给用户。 --- ## 8. 当前已验证样例 | 场景 | sessionId | 结果 | |---|---|---| | HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim | | HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 | --- ## 9. 仍需记录或后续补强 当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补: 1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。 2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。 3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。