docs(architecture): align evidence pipeline design
This commit is contained in:
@@ -1,78 +1,44 @@
|
||||
# Current Chat Agent Data Contracts
|
||||
# Chat Evidence Pipeline Contracts
|
||||
|
||||
**状态**:当前实现
|
||||
**日期**:2026-07-07
|
||||
**范围**:当前 Chat 复杂诊断链路的数据结构定义
|
||||
**状态**:当前实现
|
||||
**更新日期**:2026-07-08
|
||||
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
|
||||
|
||||
当前代码实现是三 Agent 顺序链路:
|
||||
当前 Chat 复杂诊断链路是:
|
||||
|
||||
```text
|
||||
chat_planner -> chat_executor -> chat_verifier
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> chat_verifier
|
||||
-> chat_composer
|
||||
-> final answer
|
||||
```
|
||||
|
||||
对应 `ChatService.executeChatComplex(...)` 中的 `SequentialAgent`。
|
||||
设计原则:
|
||||
|
||||
- Planner 暂不输出 `scope_contract`。
|
||||
- Executor 只做证据收集和微观事实提炼,不生成最终用户答案。
|
||||
- Gatekeeper 在 Verifier 前做代码级引用真实性校验。
|
||||
- Verifier 判断 claim 是否能由已核验证据推出。
|
||||
- Composer 只表达 Verifier 允许输出的内容。
|
||||
|
||||
---
|
||||
|
||||
## 1. Workflow Input
|
||||
## 1. Planner
|
||||
|
||||
由 `ChatService.buildWorkflowInput(...)` 构造,传给 `chat_workflow`。
|
||||
|
||||
```text
|
||||
请按固定工作流完成本轮 Planner -> Executor -> Verifier。
|
||||
|
||||
--- 用户问题 ---
|
||||
{question}
|
||||
|
||||
--- retry_context ---
|
||||
{retry_context}
|
||||
|
||||
Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `question` | 用户输入 | 用户本轮原始问题 |
|
||||
| `retry_context` | ChatService | 第二轮补证据约束;首轮为空 |
|
||||
|
||||
---
|
||||
|
||||
## 2. chat_planner
|
||||
|
||||
### 2.1 Input
|
||||
|
||||
`chat_planner` 的输入来自 workflow input 和 system prompt 追加上下文。
|
||||
Planner 当前保持不变,输出 `planner_plan`:
|
||||
|
||||
```json
|
||||
{
|
||||
"question": "用户原始问题",
|
||||
"history": [],
|
||||
"available_knowledge_domains": "...",
|
||||
"skill_catalog": {},
|
||||
"retry_context": null
|
||||
"selected_skill": "diagnose-mysql-connection-pool",
|
||||
"selection_reason": "选择该 skill 的原因",
|
||||
"plan": ["步骤1", "步骤2"],
|
||||
"reasoning": "规划思路"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `question` | workflow input | 用户原始问题 |
|
||||
| `history` | `ChatService.buildChatPlannerAgent(...)` | 对话历史,拼接到 planner system prompt |
|
||||
| `available_knowledge_domains` | `KnowledgeDomainService.buildKnowledgeMap()` | 可用知识域地图,拼接到 planner system prompt |
|
||||
| `skill_catalog` | `PlannerSkillMetadataHook` | Planner 可见的 skill name/description 元数据 |
|
||||
| `retry_context` | `ChatService` | Verifier 低置信后构造的补证据上下文 |
|
||||
|
||||
### 2.2 Output:`planner_plan`
|
||||
|
||||
当前 prompt 要求输出 JSON:
|
||||
|
||||
```json
|
||||
{
|
||||
"selected_skill": "匹配的 skill 名称;如果没有匹配则为 null",
|
||||
"selection_reason": "选择该 skill 的原因;如果没有匹配则说明不使用 skill",
|
||||
"plan": ["步骤1描述", "步骤2描述", "步骤3描述"],
|
||||
"reasoning": "规划思路说明"
|
||||
}
|
||||
```
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
@@ -81,434 +47,379 @@ Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输
|
||||
| `plan` | array | 给 Executor 的执行步骤 |
|
||||
| `reasoning` | string | 规划思路说明 |
|
||||
|
||||
运行态输出 key:
|
||||
当前边界:
|
||||
|
||||
```text
|
||||
planner_plan
|
||||
```
|
||||
- 不新增 `scope_contract`。
|
||||
- 不要求 Planner 显式列出 forbidden actions。
|
||||
- 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。
|
||||
|
||||
---
|
||||
|
||||
## 3. chat_executor
|
||||
## 2. Executor
|
||||
|
||||
### 3.1 Input
|
||||
Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。
|
||||
|
||||
`chat_executor` 接收前序 `planner_plan`,并通过 system prompt 获得历史、skill 读取约束、retry 约束和工具权限。
|
||||
### 2.1 输出结构
|
||||
|
||||
```json
|
||||
{
|
||||
"planner_plan": {},
|
||||
"history": [],
|
||||
"retry_context": null,
|
||||
"tool_permissions": {
|
||||
"method_tools": ["dateTimeTools", "lookupKnowledgeTool", "queryMetricsTools", "queryLogsTools"],
|
||||
"tool_callbacks": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `planner_plan` | `chat_planner` | Planner 输出的计划 |
|
||||
| `history` | `ChatService.buildChatExecutorAgent(...)` | 对话历史,拼接到 executor system prompt |
|
||||
| `retry_context` | `ChatService` | 本轮补证据约束 |
|
||||
| `method_tools` | `ChatService.buildMethodToolsArray()` | Executor 可直接调用的本地工具 |
|
||||
| `tool_callbacks` | `ToolCallback[]` | 框架发现或外部注入工具 |
|
||||
| `read_skill` | `SkillsAgentHook` | 当存在 skillRegistry 时,Executor 可读取 Planner 选中的 skill |
|
||||
|
||||
### 3.2 Output:`executor_feedback`
|
||||
|
||||
当前 `chat-executor-prompt.md` 要求输出一个 JSON 对象,即 `executor_evidence_v1`。
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_version": "executor_evidence_v1",
|
||||
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "事实断言或有限结论",
|
||||
"claim_type": "observation",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识",
|
||||
"tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等",
|
||||
"source_invocation_ids": [],
|
||||
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
|
||||
"source_id": "",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [
|
||||
{
|
||||
"hypothesis_text": "未被证实但值得排查的方向",
|
||||
"basis": "它基于哪些已知证据或为什么只是推测",
|
||||
"needed_evidence": ["需要补充的证据"]
|
||||
}
|
||||
],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "建议动作",
|
||||
"reason": "为什么建议做这个动作",
|
||||
"evidence_bindings": []
|
||||
}
|
||||
],
|
||||
"missing_info": [
|
||||
"导致无法确认完整根因的证据缺口"
|
||||
],
|
||||
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致。"
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 必填 | 定义 |
|
||||
|---|---|---:|---|
|
||||
| `answer_version` | string | 是 | 固定为 `executor_evidence_v2` |
|
||||
| `claims` | array | 是 | Executor 提出的待验证事实断言 |
|
||||
| `claims[].claim_id` | string | 是 | claim 标识 |
|
||||
| `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 |
|
||||
| `claims[].claim_text` | string | 是 | 事实断言文本 |
|
||||
| `claims[].support_level` | string | 是 | `direct` 或 `indirect` |
|
||||
| `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 |
|
||||
| `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` |
|
||||
| `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 |
|
||||
| `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 |
|
||||
| `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 |
|
||||
| `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 |
|
||||
| `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 |
|
||||
| `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact |
|
||||
| `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 |
|
||||
| `missing_info` | array | 是 | 无法确认结论所缺少的证据 |
|
||||
|
||||
禁止字段:
|
||||
|
||||
- `diagnosis_summary`
|
||||
- `user_facing_answer`
|
||||
- `source_invocation_ids` 作为主引用字段
|
||||
|
||||
### 2.2 raw_path
|
||||
|
||||
当前支持的精确路径:
|
||||
|
||||
| 工具 | 正向证据路径 | 负向证据路径 |
|
||||
|---|---|---|
|
||||
| `answer_version` | string | 当前固定为 `executor_evidence_v1` |
|
||||
| `diagnosis_summary` | string | 有证据边界的简短诊断摘要 |
|
||||
| `claims` | array | 已证实或有明确间接支撑的事实断言 |
|
||||
| `claims[].claim_id` | string | claim 标识 |
|
||||
| `claims[].claim_type` | string | claim 类型,例如 `root_cause`、`symptom`、`impact` |
|
||||
| `claims[].claim_text` | string | 事实断言文本 |
|
||||
| `claims[].support_level` | string | `direct` 或 `indirect` |
|
||||
| `claims[].evidence_bindings` | array | 支撑 claim 的证据绑定,不能为空 |
|
||||
| `evidence_bindings[].source_type` | string | 证据来源类型,例如 `tool_trace` |
|
||||
| `evidence_bindings[].source_id` | string | evidence block id、trace_ref 或其它定位标识 |
|
||||
| `evidence_bindings[].tool_name` | string | 来源工具名 |
|
||||
| `evidence_bindings[].source_invocation_ids` | array | 来源 `tool_invocation.id` |
|
||||
| `evidence_bindings[].evidence_excerpt` | string | 工具返回中的原话、指标值、日志片段或关键数据 |
|
||||
| `hypotheses` | array | 未证实但值得排查的方向 |
|
||||
| `hypotheses[].hypothesis_text` | string | 假设文本 |
|
||||
| `hypotheses[].basis` | string | 假设依据和未证实原因 |
|
||||
| `hypotheses[].needed_evidence` | array | 确认该假设还需要的证据 |
|
||||
| `recommended_actions` | array | 建议动作 |
|
||||
| `recommended_actions[].action_text` | string | 建议动作文本 |
|
||||
| `recommended_actions[].reason` | string | 建议原因 |
|
||||
| `recommended_actions[].evidence_bindings` | array | 建议动作关联证据,可为空 |
|
||||
| `missing_info` | array | 证据缺口 |
|
||||
| `user_facing_answer` | string | 候选用户答案,PASS 时由 ChatService 提取输出 |
|
||||
| `query_metrics` | `$.alerts[i]` | `$.no_evidence` |
|
||||
| `query_logs` | `$.logs[i]` | `$.no_evidence` |
|
||||
| `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` |
|
||||
|
||||
运行态输出 key:
|
||||
约束:
|
||||
|
||||
```text
|
||||
executor_feedback
|
||||
- `raw_path` 必须指向数组条目或 `$.no_evidence`。
|
||||
- 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。
|
||||
- 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。
|
||||
|
||||
### 2.3 negative_observation
|
||||
|
||||
当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`:
|
||||
|
||||
```json
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "negative_observation",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.no_evidence",
|
||||
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
语义边界:
|
||||
|
||||
- `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。
|
||||
- 不表示“问题绝对不存在”。
|
||||
- 不表示“根因被排除”。
|
||||
- 不表示“系统已经健康”。
|
||||
- `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。
|
||||
|
||||
### 2.4 窄范围任务
|
||||
|
||||
窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。
|
||||
|
||||
Executor 必须遵守:
|
||||
|
||||
- 只输出 `observation` / `negative_observation`。
|
||||
- claim 数量通常 1 条,最多 2 条。
|
||||
- claim 数量限制不限制 `evidence_bindings` 数量。
|
||||
- 不输出根因、风险、修复建议、经验推断。
|
||||
- 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。
|
||||
- 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。
|
||||
|
||||
---
|
||||
|
||||
## 4. chat_verifier
|
||||
## 3. Tool Invocation Evidence Refs
|
||||
|
||||
### 4.1 Input
|
||||
工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。
|
||||
|
||||
`VerifierInputHook` 会在 Verifier 调用前替换消息历史,构造显式 JSON payload。
|
||||
### 3.1 正向证据
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 3.2 负向证据
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "no_evidence",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.no_evidence",
|
||||
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` |
|
||||
| `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 |
|
||||
| `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 |
|
||||
|
||||
---
|
||||
|
||||
## 4. Gatekeeper
|
||||
|
||||
Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。
|
||||
|
||||
### 4.1 输入
|
||||
|
||||
- `sessionId`
|
||||
- `executor_structured_output`
|
||||
- 当前 session 的 `tool_invocation`
|
||||
|
||||
### 4.2 输出
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.no_evidence",
|
||||
"matched_text": "query_logs returned no evidence; ...",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `status` | string | `pass` 或 `fail` |
|
||||
| `severity` | string | `none`、`low_confid`、`reject` |
|
||||
| `checked_bindings` | array | 每条证据绑定的校验结果 |
|
||||
| `failed_rules` | array | 失败规则 id |
|
||||
| `warnings` | array | 自动回填等非阻断信息 |
|
||||
| `errors` | array | 失败明细 |
|
||||
|
||||
校验规则:
|
||||
|
||||
- `answer_version` 必须是 `executor_evidence_v2`。
|
||||
- 不允许 `diagnosis_summary` / `user_facing_answer`。
|
||||
- 每个 claim 必须有非空 `evidence_bindings`。
|
||||
- `tool_name` 必须和真实 invocation 对齐。
|
||||
- `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。
|
||||
- `raw_path` 必须存在于 `retrieval_details.evidence_refs`。
|
||||
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
|
||||
- `negative_observation` 只能绑定 `$.no_evidence`。
|
||||
|
||||
失败分级:
|
||||
|
||||
| 场景 | severity |
|
||||
|---|---|
|
||||
| 伪造 invocation id | `reject` |
|
||||
| tool_name 与 invocation 不匹配 | `reject` |
|
||||
| raw_path 不存在 | `reject` |
|
||||
| excerpt 与 matched_text 不匹配 | `reject` |
|
||||
| negative_observation 绑定正向日志 | `reject` |
|
||||
| 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` |
|
||||
| 旧 invocation 没有 `evidence_refs` | `low_confid` |
|
||||
|
||||
---
|
||||
|
||||
## 5. Verifier
|
||||
|
||||
Verifier 输入由 `VerifierInputHook` 构造:
|
||||
|
||||
```json
|
||||
{
|
||||
"original_query": "用户原始问题",
|
||||
"executor_final_answer": "{...executor_feedback raw text...}",
|
||||
"executor_structured_output": {},
|
||||
"executor_final_answer": "{...executor raw text for debug/fallback only...}",
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": []
|
||||
},
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
},
|
||||
"tool_trace_summary": [],
|
||||
"gatekeeper_result": {},
|
||||
"retry_context": null
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `original_query` | `VerifierContextHolder` | 用户原始问题 |
|
||||
| `executor_final_answer` | `VerifierContextHolder` 或上一条 AssistantMessage | Executor 原始输出文本 |
|
||||
| `executor_structured_output` | `VerifierInputHook.parseExecutorOutput(...)` | Executor 输出可解析且包含 `claims` 时的 JSON 对象;否则为 null |
|
||||
| `executor_output_parse_status.status` | `VerifierInputHook` | `valid` / `missing` / `malformed` |
|
||||
| `executor_output_parse_status.detail` | `VerifierInputHook` | 解析状态说明 |
|
||||
| `tool_trace_summary` | `ToolTraceSummaryService.buildVerifierTraceSummary(...)` | 基于真实 `tool_invocation` 构建的证据索引 |
|
||||
| `retry_context` | `VerifierContextHolder` | 当前补证据上下文 |
|
||||
Verifier 职责:
|
||||
|
||||
### 4.2 `tool_trace_summary`
|
||||
- 不调用工具。
|
||||
- 不读 skill。
|
||||
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
|
||||
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
|
||||
- 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。
|
||||
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
|
||||
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
|
||||
|
||||
`ToolTraceSummaryService` 聚合 evidence tools:
|
||||
|
||||
```text
|
||||
lookup_knowledge, query_logs, query_metrics, query_order
|
||||
```
|
||||
|
||||
输出项结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"input_summary": "query=payment-service timeout",
|
||||
"output_summary": "log_evidence: ...",
|
||||
"evidence_level": "direct",
|
||||
"topic_domain": "general",
|
||||
"source_invocation_ids": [394],
|
||||
"invocation_count": 1,
|
||||
"failed_invocation_count": 0,
|
||||
"no_hit_invocation_count": 0,
|
||||
"query_samples": ["payment-service timeout"],
|
||||
"retrieval_layers": [],
|
||||
"relevance_levels": [],
|
||||
"source_documents": []
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `trace_ref` | string | Verifier 可引用的证据摘要编号 |
|
||||
| `tool_name` | string | 聚合后的工具名 |
|
||||
| `success` | boolean | 是否存在可用证据 |
|
||||
| `input_summary` | string | 工具输入摘要 |
|
||||
| `output_summary` | string | 工具输出摘要 |
|
||||
| `evidence_level` | string | `direct` / `indirect` / `none` |
|
||||
| `topic_domain` | string | 主题域,优先来自 `retrieval_details.retrieved_domains` |
|
||||
| `source_invocation_ids` | array | 聚合的 `tool_invocation.id` |
|
||||
| `invocation_count` | number | 聚合调用次数 |
|
||||
| `failed_invocation_count` | number | 失败调用次数 |
|
||||
| `no_hit_invocation_count` | number | 无证据或去重调用次数 |
|
||||
| `query_samples` | array | 查询样例 |
|
||||
| `retrieval_layers` | array | 检索层级 |
|
||||
| `relevance_levels` | array | 相关性等级 |
|
||||
| `source_documents` | array | 来源文档标签 |
|
||||
|
||||
### 4.3 Output:`verifier_output`
|
||||
|
||||
当前 `chat-verifier-prompt.md` 要求输出:
|
||||
输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"facts_checked": [
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Composer
|
||||
|
||||
Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。
|
||||
|
||||
输入概念:
|
||||
|
||||
| 字段 | 定义 |
|
||||
|---|---|
|
||||
| `original_query` | 用户原始问题 |
|
||||
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `allowed_claims` | Verifier 允许表达的 claims |
|
||||
| `allowed_hypotheses` | Verifier 允许表达的假设 |
|
||||
| `missing_info` | 证据缺口 |
|
||||
| `recommended_actions` | 允许表达的建议动作 |
|
||||
| `rationale` | Verifier 判定理由 |
|
||||
|
||||
输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_summary": "一句话概括",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"fact": "ERR_TIMEOUT 表示请求超时",
|
||||
"is_critical": true,
|
||||
"verification": "direct_evidence",
|
||||
"detail": "知识库文档明确给出该错误码定义",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "lookup_knowledge",
|
||||
"topic_domain": "api",
|
||||
"source_invocation_ids": [101, 104],
|
||||
"note": "trace-1 的文档摘要直接给出错误码定义"
|
||||
}
|
||||
]
|
||||
"action_text": "下一步动作",
|
||||
"reason": "原因"
|
||||
}
|
||||
],
|
||||
"rationale": "所有关键事实均有支撑,且至少一条具有直接证据"
|
||||
"user_facing_answer": "最终给用户看的中文答案"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verdict` | string | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `groundedness_score` | number | 关键事实证据支撑评分 |
|
||||
| `critical_fact_count` | number | `facts_checked` 中 `is_critical=true` 的数量 |
|
||||
| `facts_checked` | array | Verifier 校验过的事实列表 |
|
||||
| `facts_checked[].fact` | string | 被校验事实 |
|
||||
| `facts_checked[].is_critical` | boolean | 是否关键事实 |
|
||||
| `facts_checked[].verification` | string | `direct_evidence` / `indirect_support` / `no_evidence` / `contradicted` |
|
||||
| `facts_checked[].detail` | string | 校验说明 |
|
||||
| `facts_checked[].evidence_refs` | array | 证据引用 |
|
||||
| `evidence_refs[].trace_ref` | string | 引用的 `tool_trace_summary.trace_ref` |
|
||||
| `evidence_refs[].tool_name` | string | 引用工具 |
|
||||
| `evidence_refs[].topic_domain` | string | 引用主题域 |
|
||||
| `evidence_refs[].source_invocation_ids` | array | 引用的 `tool_invocation.id` |
|
||||
| `evidence_refs[].note` | string | 引用说明 |
|
||||
| `rationale` | string | verdict 判定理由 |
|
||||
表达边界:
|
||||
|
||||
运行态输出 key:
|
||||
|
||||
```text
|
||||
verifier_output
|
||||
```
|
||||
- Composer 不补事实、不补根因、不调用工具。
|
||||
- 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。
|
||||
- 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
|
||||
- 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
|
||||
|
||||
---
|
||||
|
||||
## 5. VerifierDecision
|
||||
## 7. Trace Persistence
|
||||
|
||||
`ChatService.parseVerifierDecision(...)` 将 `verifier_output` 解析为内部 record:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundednessScore": 0.5,
|
||||
"criticalFactCount": 2,
|
||||
"factsChecked": [],
|
||||
"rationale": "证据不足",
|
||||
"round": 1
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verdict` | string | Verifier verdict |
|
||||
| `groundednessScore` | number | groundedness score |
|
||||
| `criticalFactCount` | number | 关键事实数量 |
|
||||
| `factsChecked` | array | 解析后的 facts_checked |
|
||||
| `rationale` | string | 判定理由 |
|
||||
| `round` | number | 当前验证轮次 |
|
||||
|
||||
---
|
||||
|
||||
## 6. retry_context
|
||||
|
||||
当 `LOW_CONFID` 且满足重试条件时,`ChatService.buildRetryContext(...)` 构造:
|
||||
|
||||
```json
|
||||
{
|
||||
"round": 1,
|
||||
"missing_evidence_facts": [
|
||||
"某关键事实:缺少直接证据"
|
||||
],
|
||||
"instruction": "仅补充以上断言相关证据,不要重复已完成检索"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `round` | number | 触发 retry 的轮次 |
|
||||
| `missing_evidence_facts` | array | 来自 Verifier 的证据缺口 |
|
||||
| `instruction` | string | 补证据约束 |
|
||||
|
||||
---
|
||||
|
||||
## 7. diagnosis_session.self_evaluation.verifier_evaluation
|
||||
|
||||
`ChatService.persistVerifierEvaluation(...)` 将 Verifier 结果合并进 `diagnosis_session.self_evaluation`。
|
||||
`diagnosis_session.self_evaluation.verifier_evaluation` 持久化:
|
||||
|
||||
```json
|
||||
{
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.5,
|
||||
"critical_fact_count": 2,
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "证据不足",
|
||||
"rationale": "...",
|
||||
"round": 1,
|
||||
"traceability_version": "v1",
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
},
|
||||
"executor_output_parse_status": {},
|
||||
"executor_structured_output": {},
|
||||
"gatekeeper_result": {},
|
||||
"composer_output": {},
|
||||
"tool_trace_summary": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verifier_evaluation.verdict` | string | Verifier verdict |
|
||||
| `verifier_evaluation.groundedness_score` | number | groundedness score |
|
||||
| `verifier_evaluation.critical_fact_count` | number | 关键事实数量 |
|
||||
| `verifier_evaluation.facts_checked` | array | 校验事实列表 |
|
||||
| `verifier_evaluation.rationale` | string | 判定理由 |
|
||||
| `verifier_evaluation.round` | number | 验证轮次 |
|
||||
| `verifier_evaluation.traceability_version` | string | 当前固定为 `v1` |
|
||||
| `verifier_evaluation.executor_output_parse_status` | object | Executor 输出解析状态 |
|
||||
| `verifier_evaluation.executor_structured_output` | object/null | 解析后的 Executor 结构化输出 |
|
||||
| `verifier_evaluation.tool_trace_summary` | array | Verifier 使用的工具证据索引 |
|
||||
Trace API 可用于回放:
|
||||
|
||||
- Executor 输出了哪些 claim。
|
||||
- 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。
|
||||
- Gatekeeper 是否通过、是否自动回填。
|
||||
- Verifier 如何判断可推导性。
|
||||
- Composer 最终如何表达给用户。
|
||||
|
||||
---
|
||||
|
||||
## 8. Final Answer Rendering
|
||||
## 8. 当前已验证样例
|
||||
|
||||
ChatService 根据 Verifier verdict 决定最终 `diagnosis_session.answer`。
|
||||
|
||||
| Verdict | 当前行为 |
|
||||
|---|---|
|
||||
| `PASS` | 优先提取 `executor_feedback.user_facing_answer`;提取失败则使用 executor 原文 |
|
||||
| `LOW_CONFID` | 输出低置信模板:已确认信息、当前缺口、建议下一步 |
|
||||
| `REJECT` | 输出降级模板:已确认信息、证据缺口、建议下一步 |
|
||||
|
||||
低置信模板使用:
|
||||
|
||||
```text
|
||||
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
当前缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
|
||||
拒绝模板使用:
|
||||
|
||||
```text
|
||||
当前无法基于已获取证据生成可靠结论。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
证据缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
| 场景 | sessionId | 结果 |
|
||||
|---|---|---|
|
||||
| HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim |
|
||||
| HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 |
|
||||
|
||||
---
|
||||
|
||||
## 9. Trace Persistence Data
|
||||
## 9. 仍需记录或后续补强
|
||||
|
||||
### 9.1 diagnosis_session
|
||||
当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `session_id` | string | 会话 id |
|
||||
| `query` | text | 用户问题 |
|
||||
| `status` | string | 会话状态 |
|
||||
| `agent_flow` | string | 当前 Chat 链路为 `CHAT` |
|
||||
| `total_duration_ms` | number | 总耗时 |
|
||||
| `total_token_count` | number | 总 token |
|
||||
| `step_count` | number | agent step 数 |
|
||||
| `tool_call_count` | number | tool invocation 数 |
|
||||
| `answer` | longtext | 最终用户答案 |
|
||||
| `self_evaluation` | json | 包含 verifier_evaluation |
|
||||
| `feedback` | string | 用户反馈 |
|
||||
|
||||
### 9.2 agent_step
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `session_id` | string | 会话 id |
|
||||
| `step_index` | number | 步骤序号 |
|
||||
| `agent_name` | string | `planner` / `executor` / `verifier` |
|
||||
| `model_input` | text | 模型输入摘要 |
|
||||
| `model_output` | text | 模型输出摘要 |
|
||||
| `thought` | text | hook 记录的摘要信息 |
|
||||
| `has_tool_call` | boolean | 是否包含工具调用 |
|
||||
| `duration_ms` | number | 模型调用耗时 |
|
||||
| `token_count` | number | token 数 |
|
||||
|
||||
### 9.3 tool_invocation
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `id` | number | 工具调用 id |
|
||||
| `session_id` | string | 会话 id |
|
||||
| `step_id` | number | 对应 agent_step id |
|
||||
| `tool_name` | string | 工具名 |
|
||||
| `input_params` | json | 工具输入参数 |
|
||||
| `output_preview` | text | 工具输出预览 |
|
||||
| `output_length` | number | 原始输出长度 |
|
||||
| `retrieval_layer` | string | 检索层 |
|
||||
| `l0_match_count` | number | L0 命中数 |
|
||||
| `l1_match_count` | number | L1 命中数 |
|
||||
| `is_truncated` | boolean | 输出是否截断 |
|
||||
| `relevance_level` | string | 相关性等级 |
|
||||
| `dedup_reason` | string | 去重原因 |
|
||||
| `retrieval_details` | json | 检索细节 |
|
||||
| `duration_ms` | number | 工具耗时 |
|
||||
| `success` | boolean | 是否成功 |
|
||||
| `error_message` | text | 错误信息 |
|
||||
1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
|
||||
2. Gatekeeper 规则配置化文档:如果后续把规则做成索引层、元数据层、规则层,需要单独记录加载顺序和审计字段。
|
||||
3. E2E fixture 矩阵:把 ISS-008/ISS-009 的用例固化到诊断评测集,而不是只存在 issue 验证记录。
|
||||
4. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
|
||||
|
||||
Reference in New Issue
Block a user