From a77c947cd4469f598654530984866a30adb9aa11 Mon Sep 17 00:00:00 2001 From: aruo <40362743+zyongxin@users.noreply.github.com> Date: Wed, 8 Jul 2026 23:53:21 +0800 Subject: [PATCH] docs(architecture): align evidence pipeline design --- mvp/architecture/README.md | 26 +- mvp/architecture/agent-orchestration.md | 52 +- mvp/architecture/current-mvp-architecture.md | 37 +- mvp/architecture/data-model.md | 64 +- .../executor-evidence-pipeline-refactor.md | 723 ++++++++---------- mvp/architecture/feedback-architecture.md | 57 +- mvp/architecture/harness-quality-gates.md | 62 +- mvp/architecture/interview-one-pager.md | 17 +- mvp/architecture/retrieval-observability.md | 20 +- mvp/architecture/session-trace-lifecycle.md | 26 +- 10 files changed, 593 insertions(+), 491 deletions(-) diff --git a/mvp/architecture/README.md b/mvp/architecture/README.md index 05446ae..8cb507c 100644 --- a/mvp/architecture/README.md +++ b/mvp/architecture/README.md @@ -1,6 +1,6 @@ # MVP 架构文档 -**更新日期**:2026-07-06 +**更新日期**:2026-07-08 这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到: @@ -15,7 +15,8 @@ | [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 | | [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 | | [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 | -| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 | +| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` | +| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 | | [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 | | [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace | | [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance | @@ -28,19 +29,20 @@ ## 当前架构一句话 -SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。 +SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。 ## 阅读顺序 1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。 2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。 3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。 -4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。 -5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。 -6. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。 -7. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。 -8. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。 -9. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。 -10. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。 -11. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。 -12. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。 +4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。 +5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。 +6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。 +7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。 +8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。 +9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。 +10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。 +11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。 +12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。 +13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。 diff --git a/mvp/architecture/agent-orchestration.md b/mvp/architecture/agent-orchestration.md index 22ab6e1..17dbbdd 100644 --- a/mvp/architecture/agent-orchestration.md +++ b/mvp/architecture/agent-orchestration.md @@ -1,14 +1,14 @@ # Agent 编排架构 -**更新日期**:2026-07-05 -**状态**:当前可运行架构 +**更新日期**:2026-07-08 +**状态**:当前可运行架构 **参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md` ## 1. 设计定位 旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛: -- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Verifier`。 +- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。 - AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。 - 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。 - 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。 @@ -23,9 +23,11 @@ flowchart TB ChatPlanner --> ChatExecutor["chat_executor"] ChatExecutor --> ChatTools["evidence tools"] ChatTools --> ChatExecutor - ChatExecutor --> ChatVerifier["chat_verifier"] + ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"] + ChatGatekeeper --> ChatVerifier["chat_verifier"] ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"} - ChatDecision --> ChatAnswer["final answer"] + ChatDecision --> ChatComposer["chat_composer"] + ChatComposer --> ChatAnswer["final answer"] end subgraph AiOps["AIOps diagnosis"] @@ -49,9 +51,11 @@ flowchart TB ChatService --> Session ChatPlanner --> Step ChatExecutor --> Step + ChatGatekeeper --> SelfEval ChatVerifier --> Step ChatTools --> Invocation ChatDecision --> SelfEval + ChatComposer --> Step AiOpsService --> Session AiOpsPlanner --> Step @@ -68,9 +72,13 @@ Chat 复杂诊断采用 `SequentialAgent`,顺序固定: chat_planner -> chat_executor -> lookup_knowledge / query_logs / query_metrics / date_time + -> outputs executor_evidence_v2 + -> VerifierInputHook / ExecutorGatekeeperService + -> validates source_invocation_id / raw_path / evidence_excerpt -> chat_verifier - -> reads tool_trace_summary - -> outputs verifier JSON + -> judges whether verified evidence can derive claims + -> chat_composer + -> writes final user-facing answer ``` 关键行为: @@ -78,8 +86,10 @@ chat_planner | 角色 | 当前职责 | 输出 | |---|---|---| | `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` | -| `chat_executor` | 按计划调用证据工具,组合工具返回形成诊断答复 | `executor_feedback` | -| `chat_verifier` | 只基于已有证据校验 Executor 答案,不做新检索 | `verifier_output` | +| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` | +| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` | +| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` | +| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` | Chat 链路最多支持两轮验证: @@ -90,7 +100,9 @@ sequenceDiagram participant P as chat_planner participant E as chat_executor participant T as tools + participant G as gatekeeper participant V as chat_verifier + participant M as chat_composer participant S as diagnosis_session C->>P: 原始问题 + history + retry_context @@ -98,14 +110,18 @@ sequenceDiagram C->>E: planner_plan + 上下文 E->>T: 调用证据工具 T-->>E: 证据结果 - E-->>C: executor_feedback - C->>V: executor_final_answer + tool_trace_summary + E-->>C: executor_evidence_v2 + C->>G: executor_structured_output + tool_invocation.evidence_refs + G-->>C: gatekeeper_result + C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary V-->>C: PASS / LOW_CONFID / REJECT C->>S: 写入 verifier_evaluation alt LOW_CONFID 且允许补证据 C->>P: retry_context: 仅补缺失证据 else PASS 或 REJECT - C-->>S: 保存最终 answer + C->>M: allowed_claims + missing_info + recommended_actions + M-->>C: composer_output + C->>S: 保存 Composer 最终 answer end ``` @@ -113,7 +129,7 @@ sequenceDiagram | Verdict | 行为 | |---|---| -| `PASS` | 输出 Executor 答案 | +| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 | | `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 | | `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 | @@ -177,21 +193,25 @@ flowchart LR SkillBody --> Executor Executor --> EvidenceTools["lookup_knowledge / logs / metrics"] EvidenceTools --> ToolTrace["tool_invocation evidence"] - Executor --> Verifier["Verifier"] + Executor --> Gatekeeper["Gatekeeper"] + Gatekeeper --> Verifier["Verifier"] ToolTrace --> Verifier + Verifier --> Composer["Composer"] ``` | 角色 | Skill 可见性 | 工具权限 | |---|---|---| | Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` | | Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 | -| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 `tool_trace_summary` | +| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` | +| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary | +| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 | ## 7. 与旧版设计的差异 | 旧版设想 | 当前实现 | |---|---| -| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Verifier;AIOps: Supervisor + Planner + Executor | +| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor | | ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 | | 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 | | Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT | diff --git a/mvp/architecture/current-mvp-architecture.md b/mvp/architecture/current-mvp-architecture.md index 03355de..ad36e84 100644 --- a/mvp/architecture/current-mvp-architecture.md +++ b/mvp/architecture/current-mvp-architecture.md @@ -1,7 +1,7 @@ # 当前 MVP 架构 -**更新日期**:2026-07-05 -**状态**:当前可运行架构 +**更新日期**:2026-07-08 +**状态**:当前可运行架构 **适用范围**:Demo、面试讲解、后续迭代规划 ## 1. 系统定位 @@ -38,7 +38,9 @@ flowchart TB Supervisor["Supervisor"] Planner["Planner"] Executor["Executor"] + Gatekeeper["Gatekeeper"] Verifier["Verifier"] + Composer["Composer"] end subgraph Tools["Evidence Tools"] @@ -105,7 +107,9 @@ Agent Orchestration -> Supervisor -> Planner -> Executor + -> Gatekeeper -> Verifier + -> Composer Evidence Tools -> lookup_knowledge @@ -133,6 +137,7 @@ Persistence -> Milvus/Zilliz collection Quality Gates + -> executor gatekeeper -> chat verifier -> AIOps rule evaluation -> diagnosis eval baseline @@ -150,7 +155,9 @@ sequenceDiagram participant Planner as Planner Agent participant Executor as Executor Agent participant Tool as Evidence Tools + participant Gatekeeper as Gatekeeper Hook participant Verifier as Verifier Agent + participant Composer as Composer Agent participant DB as Trace Tables participant Trace as Trace API @@ -162,8 +169,12 @@ sequenceDiagram Executor->>Tool: lookup_knowledge / logs / metrics Tool->>DB: 写入 tool_invocation Tool-->>Executor: 返回证据 - Executor->>Verifier: 生成候选诊断并校验 + Executor->>Gatekeeper: 输出 executor_evidence_v2 + Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用 + Gatekeeper->>Verifier: 传入已验真的 claims / excerpts Verifier->>DB: 合并 self_evaluation.verifier_evaluation + Verifier->>Composer: 传入 allowed_claims / missing_info / actions + Composer->>Chat: 生成最终用户答复 Chat->>DB: 保存 diagnosis_session.answer User->>Trace: GET /api/diagnosis/{sessionId}/trace Trace->>DB: 聚合 session / step / tool @@ -180,14 +191,16 @@ POST /api/chat -> lookup_knowledge -> query_logs -> query_metrics - -> Verifier 校验最终诊断 + -> Gatekeeper 校验 Executor 证据引用真实性 + -> Verifier 判断 claim 是否能由已核验证据推出 + -> Composer 生成最终用户答复 -> 保存 diagnosis_session -> 保存 agent_step -> 保存 tool_invocation -> 合并 self_evaluation.verifier_evaluation ``` -Chat 链路的质量门禁是 LLM Verifier。Verifier 输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。 +Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。 Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。 @@ -323,6 +336,7 @@ tool_invocation -> 工具调用事实 -> tool_name / input_params / output_preview -> retrieval_layer / retrieval_details + -> retrieval_details.evidence_refs -> relevance_level / dedup_reason -> duration / success ``` @@ -346,12 +360,12 @@ Trace API 聚合: - 会话状态和最终报告。 - Agent step 序列。 - 工具调用和检索细节。 -- Chat verifier 结果。 +- Chat Gatekeeper / Verifier / Composer 结果。 - AIOps rule evaluation 结果。 Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。 -Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。 +Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。 ## 8. 质量门禁 @@ -359,7 +373,9 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate | 门禁 | 位置 | 作用 | |---|---|---| -| Chat Verifier | `ChatService` | 校验普通诊断回答质量 | +| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 | +| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims | +| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 | | AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 | | Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 | | RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 | @@ -379,6 +395,10 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate - `title`、`breadcrumb`、`content` 参与 embedding 文本。 - `tool_invocation` 记录检索层、relevance level、dedup reason。 - Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。 +- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。 +- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。 +- Gatekeeper 对 Executor 引用做代码级验真,Verifier 只判断可推导性。 +- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。 - RAG offline baseline 和 live acceptance 脚本。 暂不作为当前已完成能力声明: @@ -406,4 +426,5 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate | Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` | | Trace 聚合 | `DiagnosisTraceService` | | 工具调用记录 | `ToolInvocationRecorder` | +| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` | | self_evaluation 合并 | `SelfEvaluationMergeService` | diff --git a/mvp/architecture/data-model.md b/mvp/architecture/data-model.md index 36bed82..672dff7 100644 --- a/mvp/architecture/data-model.md +++ b/mvp/architecture/data-model.md @@ -1,6 +1,6 @@ # 数据模型总览 -**更新日期**:2026-07-05 +**更新日期**:2026-07-08 **状态**:当前可运行架构 ## 1. 定位 @@ -130,10 +130,35 @@ erDiagram 用途: - 给 Trace API 展示证据。 -- 给 Verifier 构造 `tool_trace_summary`。 +- 给 Gatekeeper 提供 `retrieval_details.evidence_refs` 引用验真源。 +- 给 Verifier 构造 `tool_trace_summary` 审计导航。 - 给 `EvaluationService` 计算 evidence score。 - 给 RAG eval 和人工排查提供检索细节。 +`retrieval_details.evidence_refs` 是当前 Chat 证据链路的关键字段: + +```json +{ + "evidence_status": "supported", + "evidence_refs": [ + { + "raw_path": "$.logs[0]", + "text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..." + } + ] +} +``` + +字段边界: + +| 字段 | 说明 | +|---|---| +| `evidence_status` | 工具证据状态,例如 `supported`、`no_evidence`、`deduped`、`failed` | +| `evidence_refs[].raw_path` | Executor 可引用的稳定路径,例如 `$.logs[0]`、`$.alerts[0]`、`$.evidence_blocks[0]`、`$.no_evidence` | +| `evidence_refs[].text` | 系统抽取的最小证据文本,Gatekeeper 用它核对 `evidence_excerpt` | + +`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。 + ## 4. 知识库模型 ### api_document @@ -213,9 +238,39 @@ category 边界: - `rule_evaluation` 评估证据收集充分度。 -- `verifier_evaluation` 评估 Chat 答案关键事实是否有证据支撑。 +- `verifier_evaluation` 评估 Chat 结构化 claims 是否能由已验真证据推出,并保存 Gatekeeper、Verifier、Composer 的审计数据。 - `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。 +当前 `verifier_evaluation` 关键结构: + +```json +{ + "verdict": "PASS", + "groundedness_score": 1.0, + "critical_fact_count": 1, + "claim_checks": [], + "facts_checked": [], + "rationale": "...", + "round": 1, + "traceability_version": "v1", + "executor_output_parse_status": {}, + "executor_structured_output": {}, + "gatekeeper_result": {}, + "composer_output": {}, + "tool_trace_summary": [] +} +``` + +必要审计字段: + +| 字段 | 说明 | +|---|---| +| `executor_output_parse_status` | Executor 输出是否能解析为 `executor_evidence_v2` | +| `executor_structured_output` | Executor 结构化 claims、hypotheses、recommended_actions、missing_info | +| `gatekeeper_result` | 引用真实性校验结果,包括 checked bindings、failed rules、warnings、errors | +| `composer_output` | Composer 最终表达及解析状态 | +| `tool_trace_summary` | Verifier 调用时使用的工具调用导航索引,不是唯一证据源 | + ## 7. 数据写入时序 ```mermaid @@ -253,6 +308,5 @@ sequenceDiagram 1. 增加 run id,支持同 session 多次独立诊断。 2. 强化 `tool_invocation.step_id` 关联。 -3. 将 evidence block 结构化保存。 +3. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。 4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。 - diff --git a/mvp/architecture/executor-evidence-pipeline-refactor.md b/mvp/architecture/executor-evidence-pipeline-refactor.md index d34f5ac..e75f9fd 100644 --- a/mvp/architecture/executor-evidence-pipeline-refactor.md +++ b/mvp/architecture/executor-evidence-pipeline-refactor.md @@ -1,78 +1,44 @@ -# Current Chat Agent Data Contracts +# Chat Evidence Pipeline Contracts -**状态**:当前实现 -**日期**:2026-07-07 -**范围**:当前 Chat 复杂诊断链路的数据结构定义 +**状态**:当前实现 +**更新日期**:2026-07-08 +**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约 -当前代码实现是三 Agent 顺序链路: +当前 Chat 复杂诊断链路是: ```text -chat_planner -> chat_executor -> chat_verifier +chat_planner + -> chat_executor + -> VerifierInputHook / ExecutorGatekeeperService + -> chat_verifier + -> chat_composer + -> final answer ``` -对应 `ChatService.executeChatComplex(...)` 中的 `SequentialAgent`。 +设计原则: + +- Planner 暂不输出 `scope_contract`。 +- Executor 只做证据收集和微观事实提炼,不生成最终用户答案。 +- Gatekeeper 在 Verifier 前做代码级引用真实性校验。 +- Verifier 判断 claim 是否能由已核验证据推出。 +- Composer 只表达 Verifier 允许输出的内容。 --- -## 1. Workflow Input +## 1. Planner -由 `ChatService.buildWorkflowInput(...)` 构造,传给 `chat_workflow`。 - -```text -请按固定工作流完成本轮 Planner -> Executor -> Verifier。 - ---- 用户问题 --- -{question} - ---- retry_context --- -{retry_context} - -Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。 -``` - -| 字段 | 来源 | 定义 | -|---|---|---| -| `question` | 用户输入 | 用户本轮原始问题 | -| `retry_context` | ChatService | 第二轮补证据约束;首轮为空 | - ---- - -## 2. chat_planner - -### 2.1 Input - -`chat_planner` 的输入来自 workflow input 和 system prompt 追加上下文。 +Planner 当前保持不变,输出 `planner_plan`: ```json { - "question": "用户原始问题", - "history": [], - "available_knowledge_domains": "...", - "skill_catalog": {}, - "retry_context": null + "selected_skill": "diagnose-mysql-connection-pool", + "selection_reason": "选择该 skill 的原因", + "plan": ["步骤1", "步骤2"], + "reasoning": "规划思路" } ``` -| 字段 | 来源 | 定义 | -|---|---|---| -| `question` | workflow input | 用户原始问题 | -| `history` | `ChatService.buildChatPlannerAgent(...)` | 对话历史,拼接到 planner system prompt | -| `available_knowledge_domains` | `KnowledgeDomainService.buildKnowledgeMap()` | 可用知识域地图,拼接到 planner system prompt | -| `skill_catalog` | `PlannerSkillMetadataHook` | Planner 可见的 skill name/description 元数据 | -| `retry_context` | `ChatService` | Verifier 低置信后构造的补证据上下文 | - -### 2.2 Output:`planner_plan` - -当前 prompt 要求输出 JSON: - -```json -{ - "selected_skill": "匹配的 skill 名称;如果没有匹配则为 null", - "selection_reason": "选择该 skill 的原因;如果没有匹配则说明不使用 skill", - "plan": ["步骤1描述", "步骤2描述", "步骤3描述"], - "reasoning": "规划思路说明" -} -``` +字段定义: | 字段 | 类型 | 定义 | |---|---|---| @@ -81,434 +47,379 @@ Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输 | `plan` | array | 给 Executor 的执行步骤 | | `reasoning` | string | 规划思路说明 | -运行态输出 key: +当前边界: -```text -planner_plan -``` +- 不新增 `scope_contract`。 +- 不要求 Planner 显式列出 forbidden actions。 +- 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。 --- -## 3. chat_executor +## 2. Executor -### 3.1 Input +Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。 -`chat_executor` 接收前序 `planner_plan`,并通过 system prompt 获得历史、skill 读取约束、retry 约束和工具权限。 +### 2.1 输出结构 ```json { - "planner_plan": {}, - "history": [], - "retry_context": null, - "tool_permissions": { - "method_tools": ["dateTimeTools", "lookupKnowledgeTool", "queryMetricsTools", "queryLogsTools"], - "tool_callbacks": [] - } -} -``` - -| 字段 | 来源 | 定义 | -|---|---|---| -| `planner_plan` | `chat_planner` | Planner 输出的计划 | -| `history` | `ChatService.buildChatExecutorAgent(...)` | 对话历史,拼接到 executor system prompt | -| `retry_context` | `ChatService` | 本轮补证据约束 | -| `method_tools` | `ChatService.buildMethodToolsArray()` | Executor 可直接调用的本地工具 | -| `tool_callbacks` | `ToolCallback[]` | 框架发现或外部注入工具 | -| `read_skill` | `SkillsAgentHook` | 当存在 skillRegistry 时,Executor 可读取 Planner 选中的 skill | - -### 3.2 Output:`executor_feedback` - -当前 `chat-executor-prompt.md` 要求输出一个 JSON 对象,即 `executor_evidence_v1`。 - -```json -{ - "answer_version": "executor_evidence_v1", - "diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界", + "answer_version": "executor_evidence_v2", "claims": [ { "claim_id": "claim-1", - "claim_type": "root_cause", - "claim_text": "事实断言或有限结论", + "claim_type": "observation", + "claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。", "support_level": "direct", "evidence_bindings": [ { "source_type": "tool_trace", - "source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识", - "tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等", - "source_invocation_ids": [], - "evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据" + "source_id": "", + "tool_name": "query_metrics", + "source_invocation_id": 517, + "raw_path": "$.alerts[0]", + "evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m" } ] } ], - "hypotheses": [ - { - "hypothesis_text": "未被证实但值得排查的方向", - "basis": "它基于哪些已知证据或为什么只是推测", - "needed_evidence": ["需要补充的证据"] - } - ], - "recommended_actions": [ - { - "action_text": "建议动作", - "reason": "为什么建议做这个动作", - "evidence_bindings": [] - } - ], - "missing_info": [ - "导致无法确认完整根因的证据缺口" - ], - "user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致。" + "hypotheses": [], + "recommended_actions": [], + "missing_info": [] } ``` -| 字段 | 类型 | 定义 | +字段定义: + +| 字段 | 类型 | 必填 | 定义 | +|---|---|---:|---| +| `answer_version` | string | 是 | 固定为 `executor_evidence_v2` | +| `claims` | array | 是 | Executor 提出的待验证事实断言 | +| `claims[].claim_id` | string | 是 | claim 标识 | +| `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 | +| `claims[].claim_text` | string | 是 | 事实断言文本 | +| `claims[].support_level` | string | 是 | `direct` 或 `indirect` | +| `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 | +| `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` | +| `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 | +| `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 | +| `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 | +| `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 | +| `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 | +| `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact | +| `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 | +| `missing_info` | array | 是 | 无法确认结论所缺少的证据 | + +禁止字段: + +- `diagnosis_summary` +- `user_facing_answer` +- `source_invocation_ids` 作为主引用字段 + +### 2.2 raw_path + +当前支持的精确路径: + +| 工具 | 正向证据路径 | 负向证据路径 | |---|---|---| -| `answer_version` | string | 当前固定为 `executor_evidence_v1` | -| `diagnosis_summary` | string | 有证据边界的简短诊断摘要 | -| `claims` | array | 已证实或有明确间接支撑的事实断言 | -| `claims[].claim_id` | string | claim 标识 | -| `claims[].claim_type` | string | claim 类型,例如 `root_cause`、`symptom`、`impact` | -| `claims[].claim_text` | string | 事实断言文本 | -| `claims[].support_level` | string | `direct` 或 `indirect` | -| `claims[].evidence_bindings` | array | 支撑 claim 的证据绑定,不能为空 | -| `evidence_bindings[].source_type` | string | 证据来源类型,例如 `tool_trace` | -| `evidence_bindings[].source_id` | string | evidence block id、trace_ref 或其它定位标识 | -| `evidence_bindings[].tool_name` | string | 来源工具名 | -| `evidence_bindings[].source_invocation_ids` | array | 来源 `tool_invocation.id` | -| `evidence_bindings[].evidence_excerpt` | string | 工具返回中的原话、指标值、日志片段或关键数据 | -| `hypotheses` | array | 未证实但值得排查的方向 | -| `hypotheses[].hypothesis_text` | string | 假设文本 | -| `hypotheses[].basis` | string | 假设依据和未证实原因 | -| `hypotheses[].needed_evidence` | array | 确认该假设还需要的证据 | -| `recommended_actions` | array | 建议动作 | -| `recommended_actions[].action_text` | string | 建议动作文本 | -| `recommended_actions[].reason` | string | 建议原因 | -| `recommended_actions[].evidence_bindings` | array | 建议动作关联证据,可为空 | -| `missing_info` | array | 证据缺口 | -| `user_facing_answer` | string | 候选用户答案,PASS 时由 ChatService 提取输出 | +| `query_metrics` | `$.alerts[i]` | `$.no_evidence` | +| `query_logs` | `$.logs[i]` | `$.no_evidence` | +| `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` | -运行态输出 key: +约束: -```text -executor_feedback +- `raw_path` 必须指向数组条目或 `$.no_evidence`。 +- 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。 +- 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。 + +### 2.3 negative_observation + +当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`: + +```json +{ + "claim_id": "claim-1", + "claim_type": "negative_observation", + "claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。", + "support_level": "direct", + "evidence_bindings": [ + { + "tool_name": "query_logs", + "source_invocation_id": 517, + "raw_path": "$.no_evidence", + "evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence" + } + ] +} ``` +语义边界: + +- `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。 +- 不表示“问题绝对不存在”。 +- 不表示“根因被排除”。 +- 不表示“系统已经健康”。 +- `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。 + +### 2.4 窄范围任务 + +窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。 + +Executor 必须遵守: + +- 只输出 `observation` / `negative_observation`。 +- claim 数量通常 1 条,最多 2 条。 +- claim 数量限制不限制 `evidence_bindings` 数量。 +- 不输出根因、风险、修复建议、经验推断。 +- 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。 +- 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。 + --- -## 4. chat_verifier +## 3. Tool Invocation Evidence Refs -### 4.1 Input +工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。 -`VerifierInputHook` 会在 Verifier 调用前替换消息历史,构造显式 JSON payload。 +### 3.1 正向证据 + +```json +{ + "evidence_status": "supported", + "evidence_refs": [ + { + "raw_path": "$.logs[0]", + "text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..." + } + ] +} +``` + +### 3.2 负向证据 + +```json +{ + "evidence_status": "no_evidence", + "evidence_refs": [ + { + "raw_path": "$.no_evidence", + "text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志" + } + ] +} +``` + +字段定义: + +| 字段 | 类型 | 定义 | +|---|---|---| +| `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` | +| `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 | +| `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 | + +--- + +## 4. Gatekeeper + +Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。 + +### 4.1 输入 + +- `sessionId` +- `executor_structured_output` +- 当前 session 的 `tool_invocation` + +### 4.2 输出 + +```json +{ + "status": "pass", + "severity": "none", + "checked_bindings": [ + { + "claim_id": "claim-1", + "tool_name": "query_logs", + "source_invocation_id": 517, + "raw_path": "$.no_evidence", + "matched_text": "query_logs returned no evidence; ...", + "status": "pass" + } + ], + "failed_rules": [], + "warnings": [], + "errors": [] +} +``` + +字段定义: + +| 字段 | 类型 | 定义 | +|---|---|---| +| `status` | string | `pass` 或 `fail` | +| `severity` | string | `none`、`low_confid`、`reject` | +| `checked_bindings` | array | 每条证据绑定的校验结果 | +| `failed_rules` | array | 失败规则 id | +| `warnings` | array | 自动回填等非阻断信息 | +| `errors` | array | 失败明细 | + +校验规则: + +- `answer_version` 必须是 `executor_evidence_v2`。 +- 不允许 `diagnosis_summary` / `user_facing_answer`。 +- 每个 claim 必须有非空 `evidence_bindings`。 +- `tool_name` 必须和真实 invocation 对齐。 +- `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。 +- `raw_path` 必须存在于 `retrieval_details.evidence_refs`。 +- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。 +- `negative_observation` 只能绑定 `$.no_evidence`。 + +失败分级: + +| 场景 | severity | +|---|---| +| 伪造 invocation id | `reject` | +| tool_name 与 invocation 不匹配 | `reject` | +| raw_path 不存在 | `reject` | +| excerpt 与 matched_text 不匹配 | `reject` | +| negative_observation 绑定正向日志 | `reject` | +| 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` | +| 旧 invocation 没有 `evidence_refs` | `low_confid` | + +--- + +## 5. Verifier + +Verifier 输入由 `VerifierInputHook` 构造: ```json { "original_query": "用户原始问题", - "executor_final_answer": "{...executor_feedback raw text...}", - "executor_structured_output": {}, + "executor_final_answer": "{...executor raw text for debug/fallback only...}", + "executor_structured_output": { + "answer_version": "executor_evidence_v2", + "claims": [] + }, "executor_output_parse_status": { "status": "valid", "detail": "parsed executor evidence contract" }, "tool_trace_summary": [], + "gatekeeper_result": {}, "retry_context": null } ``` -| 字段 | 来源 | 定义 | -|---|---|---| -| `original_query` | `VerifierContextHolder` | 用户原始问题 | -| `executor_final_answer` | `VerifierContextHolder` 或上一条 AssistantMessage | Executor 原始输出文本 | -| `executor_structured_output` | `VerifierInputHook.parseExecutorOutput(...)` | Executor 输出可解析且包含 `claims` 时的 JSON 对象;否则为 null | -| `executor_output_parse_status.status` | `VerifierInputHook` | `valid` / `missing` / `malformed` | -| `executor_output_parse_status.detail` | `VerifierInputHook` | 解析状态说明 | -| `tool_trace_summary` | `ToolTraceSummaryService.buildVerifierTraceSummary(...)` | 基于真实 `tool_invocation` 构建的证据索引 | -| `retry_context` | `VerifierContextHolder` | 当前补证据上下文 | +Verifier 职责: -### 4.2 `tool_trace_summary` +- 不调用工具。 +- 不读 skill。 +- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。 +- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。 +- 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。 +- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。 +- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。 -`ToolTraceSummaryService` 聚合 evidence tools: - -```text -lookup_knowledge, query_logs, query_metrics, query_order -``` - -输出项结构: - -```json -{ - "trace_ref": "trace-1", - "tool_name": "query_logs", - "success": true, - "input_summary": "query=payment-service timeout", - "output_summary": "log_evidence: ...", - "evidence_level": "direct", - "topic_domain": "general", - "source_invocation_ids": [394], - "invocation_count": 1, - "failed_invocation_count": 0, - "no_hit_invocation_count": 0, - "query_samples": ["payment-service timeout"], - "retrieval_layers": [], - "relevance_levels": [], - "source_documents": [] -} -``` - -| 字段 | 类型 | 定义 | -|---|---|---| -| `trace_ref` | string | Verifier 可引用的证据摘要编号 | -| `tool_name` | string | 聚合后的工具名 | -| `success` | boolean | 是否存在可用证据 | -| `input_summary` | string | 工具输入摘要 | -| `output_summary` | string | 工具输出摘要 | -| `evidence_level` | string | `direct` / `indirect` / `none` | -| `topic_domain` | string | 主题域,优先来自 `retrieval_details.retrieved_domains` | -| `source_invocation_ids` | array | 聚合的 `tool_invocation.id` | -| `invocation_count` | number | 聚合调用次数 | -| `failed_invocation_count` | number | 失败调用次数 | -| `no_hit_invocation_count` | number | 无证据或去重调用次数 | -| `query_samples` | array | 查询样例 | -| `retrieval_layers` | array | 检索层级 | -| `relevance_levels` | array | 相关性等级 | -| `source_documents` | array | 来源文档标签 | - -### 4.3 Output:`verifier_output` - -当前 `chat-verifier-prompt.md` 要求输出: +输出: ```json { "verdict": "PASS", - "groundedness_score": 0.8, - "critical_fact_count": 2, - "facts_checked": [ + "groundedness_score": 1.0, + "critical_fact_count": 1, + "claim_checks": [], + "facts_checked": [], + "rationale": "..." +} +``` + +--- + +## 6. Composer + +Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。 + +输入概念: + +| 字段 | 定义 | +|---|---| +| `original_query` | 用户原始问题 | +| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` | +| `allowed_claims` | Verifier 允许表达的 claims | +| `allowed_hypotheses` | Verifier 允许表达的假设 | +| `missing_info` | 证据缺口 | +| `recommended_actions` | 允许表达的建议动作 | +| `rationale` | Verifier 判定理由 | + +输出: + +```json +{ + "answer_summary": "一句话概括", + "recommended_actions": [ { - "fact": "ERR_TIMEOUT 表示请求超时", - "is_critical": true, - "verification": "direct_evidence", - "detail": "知识库文档明确给出该错误码定义", - "evidence_refs": [ - { - "trace_ref": "trace-1", - "tool_name": "lookup_knowledge", - "topic_domain": "api", - "source_invocation_ids": [101, 104], - "note": "trace-1 的文档摘要直接给出错误码定义" - } - ] + "action_text": "下一步动作", + "reason": "原因" } ], - "rationale": "所有关键事实均有支撑,且至少一条具有直接证据" + "user_facing_answer": "最终给用户看的中文答案" } ``` -| 字段 | 类型 | 定义 | -|---|---|---| -| `verdict` | string | `PASS` / `LOW_CONFID` / `REJECT` | -| `groundedness_score` | number | 关键事实证据支撑评分 | -| `critical_fact_count` | number | `facts_checked` 中 `is_critical=true` 的数量 | -| `facts_checked` | array | Verifier 校验过的事实列表 | -| `facts_checked[].fact` | string | 被校验事实 | -| `facts_checked[].is_critical` | boolean | 是否关键事实 | -| `facts_checked[].verification` | string | `direct_evidence` / `indirect_support` / `no_evidence` / `contradicted` | -| `facts_checked[].detail` | string | 校验说明 | -| `facts_checked[].evidence_refs` | array | 证据引用 | -| `evidence_refs[].trace_ref` | string | 引用的 `tool_trace_summary.trace_ref` | -| `evidence_refs[].tool_name` | string | 引用工具 | -| `evidence_refs[].topic_domain` | string | 引用主题域 | -| `evidence_refs[].source_invocation_ids` | array | 引用的 `tool_invocation.id` | -| `evidence_refs[].note` | string | 引用说明 | -| `rationale` | string | verdict 判定理由 | +表达边界: -运行态输出 key: - -```text -verifier_output -``` +- Composer 不补事实、不补根因、不调用工具。 +- 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。 +- 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。 +- 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。 --- -## 5. VerifierDecision +## 7. Trace Persistence -`ChatService.parseVerifierDecision(...)` 将 `verifier_output` 解析为内部 record: - -```json -{ - "verdict": "LOW_CONFID", - "groundednessScore": 0.5, - "criticalFactCount": 2, - "factsChecked": [], - "rationale": "证据不足", - "round": 1 -} -``` - -| 字段 | 类型 | 定义 | -|---|---|---| -| `verdict` | string | Verifier verdict | -| `groundednessScore` | number | groundedness score | -| `criticalFactCount` | number | 关键事实数量 | -| `factsChecked` | array | 解析后的 facts_checked | -| `rationale` | string | 判定理由 | -| `round` | number | 当前验证轮次 | - ---- - -## 6. retry_context - -当 `LOW_CONFID` 且满足重试条件时,`ChatService.buildRetryContext(...)` 构造: - -```json -{ - "round": 1, - "missing_evidence_facts": [ - "某关键事实:缺少直接证据" - ], - "instruction": "仅补充以上断言相关证据,不要重复已完成检索" -} -``` - -| 字段 | 类型 | 定义 | -|---|---|---| -| `round` | number | 触发 retry 的轮次 | -| `missing_evidence_facts` | array | 来自 Verifier 的证据缺口 | -| `instruction` | string | 补证据约束 | - ---- - -## 7. diagnosis_session.self_evaluation.verifier_evaluation - -`ChatService.persistVerifierEvaluation(...)` 将 Verifier 结果合并进 `diagnosis_session.self_evaluation`。 +`diagnosis_session.self_evaluation.verifier_evaluation` 持久化: ```json { "verifier_evaluation": { - "verdict": "LOW_CONFID", - "groundedness_score": 0.5, - "critical_fact_count": 2, + "verdict": "PASS", + "groundedness_score": 1.0, + "critical_fact_count": 1, + "claim_checks": [], "facts_checked": [], - "rationale": "证据不足", + "rationale": "...", "round": 1, "traceability_version": "v1", - "executor_output_parse_status": { - "status": "valid", - "detail": "parsed executor evidence contract" - }, + "executor_output_parse_status": {}, "executor_structured_output": {}, + "gatekeeper_result": {}, + "composer_output": {}, "tool_trace_summary": [] } } ``` -| 字段 | 类型 | 定义 | -|---|---|---| -| `verifier_evaluation.verdict` | string | Verifier verdict | -| `verifier_evaluation.groundedness_score` | number | groundedness score | -| `verifier_evaluation.critical_fact_count` | number | 关键事实数量 | -| `verifier_evaluation.facts_checked` | array | 校验事实列表 | -| `verifier_evaluation.rationale` | string | 判定理由 | -| `verifier_evaluation.round` | number | 验证轮次 | -| `verifier_evaluation.traceability_version` | string | 当前固定为 `v1` | -| `verifier_evaluation.executor_output_parse_status` | object | Executor 输出解析状态 | -| `verifier_evaluation.executor_structured_output` | object/null | 解析后的 Executor 结构化输出 | -| `verifier_evaluation.tool_trace_summary` | array | Verifier 使用的工具证据索引 | +Trace API 可用于回放: + +- Executor 输出了哪些 claim。 +- 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。 +- Gatekeeper 是否通过、是否自动回填。 +- Verifier 如何判断可推导性。 +- Composer 最终如何表达给用户。 --- -## 8. Final Answer Rendering +## 8. 当前已验证样例 -ChatService 根据 Verifier verdict 决定最终 `diagnosis_session.answer`。 - -| Verdict | 当前行为 | -|---|---| -| `PASS` | 优先提取 `executor_feedback.user_facing_answer`;提取失败则使用 executor 原文 | -| `LOW_CONFID` | 输出低置信模板:已确认信息、当前缺口、建议下一步 | -| `REJECT` | 输出降级模板:已确认信息、证据缺口、建议下一步 | - -低置信模板使用: - -```text -以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。 - -已确认信息: -- ... - -当前缺口: -- ... - -建议下一步: -- ... -``` - -拒绝模板使用: - -```text -当前无法基于已获取证据生成可靠结论。 - -已确认信息: -- ... - -证据缺口: -- ... - -建议下一步: -- ... -``` +| 场景 | sessionId | 结果 | +|---|---|---| +| HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim | +| HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 | --- -## 9. Trace Persistence Data +## 9. 仍需记录或后续补强 -### 9.1 diagnosis_session +当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补: -| 字段 | 类型 | 定义 | -|---|---|---| -| `session_id` | string | 会话 id | -| `query` | text | 用户问题 | -| `status` | string | 会话状态 | -| `agent_flow` | string | 当前 Chat 链路为 `CHAT` | -| `total_duration_ms` | number | 总耗时 | -| `total_token_count` | number | 总 token | -| `step_count` | number | agent step 数 | -| `tool_call_count` | number | tool invocation 数 | -| `answer` | longtext | 最终用户答案 | -| `self_evaluation` | json | 包含 verifier_evaluation | -| `feedback` | string | 用户反馈 | - -### 9.2 agent_step - -| 字段 | 类型 | 定义 | -|---|---|---| -| `session_id` | string | 会话 id | -| `step_index` | number | 步骤序号 | -| `agent_name` | string | `planner` / `executor` / `verifier` | -| `model_input` | text | 模型输入摘要 | -| `model_output` | text | 模型输出摘要 | -| `thought` | text | hook 记录的摘要信息 | -| `has_tool_call` | boolean | 是否包含工具调用 | -| `duration_ms` | number | 模型调用耗时 | -| `token_count` | number | token 数 | - -### 9.3 tool_invocation - -| 字段 | 类型 | 定义 | -|---|---|---| -| `id` | number | 工具调用 id | -| `session_id` | string | 会话 id | -| `step_id` | number | 对应 agent_step id | -| `tool_name` | string | 工具名 | -| `input_params` | json | 工具输入参数 | -| `output_preview` | text | 工具输出预览 | -| `output_length` | number | 原始输出长度 | -| `retrieval_layer` | string | 检索层 | -| `l0_match_count` | number | L0 命中数 | -| `l1_match_count` | number | L1 命中数 | -| `is_truncated` | boolean | 输出是否截断 | -| `relevance_level` | string | 相关性等级 | -| `dedup_reason` | string | 去重原因 | -| `retrieval_details` | json | 检索细节 | -| `duration_ms` | number | 工具耗时 | -| `success` | boolean | 是否成功 | -| `error_message` | text | 错误信息 | +1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。 +2. Gatekeeper 规则配置化文档:如果后续把规则做成索引层、元数据层、规则层,需要单独记录加载顺序和审计字段。 +3. E2E fixture 矩阵:把 ISS-008/ISS-009 的用例固化到诊断评测集,而不是只存在 issue 验证记录。 +4. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。 diff --git a/mvp/architecture/feedback-architecture.md b/mvp/architecture/feedback-architecture.md index 5615986..21df1d7 100644 --- a/mvp/architecture/feedback-architecture.md +++ b/mvp/architecture/feedback-architecture.md @@ -1,14 +1,14 @@ # 反馈与自评估架构 -**更新日期**:2026-07-05 -**状态**:当前可运行架构 +**更新日期**:2026-07-08 +**状态**:当前可运行架构 **参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md` ## 1. 定位 反馈架构包含两条闭环: -1. 系统自评估:基于工具调用、Verifier、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。 +1. 系统自评估:基于工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。 2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。 当前重要边界: @@ -25,9 +25,14 @@ flowchart TD subgraph SelfEval["Self evaluation"] Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"] + Invocation --> EvidenceRefs["evidence_refs"] + EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"] Invocation --> TraceSummary["ToolTraceSummaryService"] - TraceSummary --> Verifier["chat_verifier"] + Gatekeeper --> Verifier["chat_verifier"] + TraceSummary --> Verifier Verifier --> VerifierEval["verifier_evaluation"] + Verifier --> Composer["chat_composer"] + Composer --> VerifierEval Invocation --> AiOpsRule["AiOpsRuleEvaluationService"] AiOpsRule --> AiOpsEval["aiops_rule_evaluation"] end @@ -67,10 +72,15 @@ flowchart TD "verdict": "PASS", "groundedness_score": 0.8, "critical_fact_count": 2, + "claim_checks": [], "facts_checked": [], "rationale": "...", "round": 1, "traceability_version": "v1", + "executor_output_parse_status": {}, + "executor_structured_output": {}, + "gatekeeper_result": {}, + "composer_output": {}, "tool_trace_summary": [] }, "aiops_rule_evaluation": { @@ -117,16 +127,28 @@ flowchart TD ## 5. Chat Verifier 自评估 -Chat Verifier 校验 Executor 的最终答案是否被证据支撑。 +Chat 自评估分三步: + +1. Gatekeeper 用代码校验 Executor 的引用是否真实。 +2. Verifier 判断已验真的 `evidence_excerpt` 是否能推出 `claim_text`。 +3. Composer 只把 Verifier 允许表达的内容写成最终用户答复。 ```mermaid flowchart LR - Answer["executor_final_answer"] --> Verifier["chat_verifier"] + ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"] Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"] + Invocation --> EvidenceRefs["retrieval_details.evidence_refs"] + EvidenceRefs --> Gatekeeper + Gatekeeper --> GateResult["gatekeeper_result"] Summary --> Evidence["tool_trace_summary"] + GateResult --> Verifier["chat_verifier"] + ExecutorOutput --> Verifier Evidence --> Verifier Verifier --> Output["verifier_output JSON"] + Output --> Composer["chat_composer"] + Composer --> ComposerOutput["composer_output"] Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"] + ComposerOutput --> Merge Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"] ``` @@ -137,16 +159,25 @@ Verifier 输出: | `verdict` | `PASS` / `LOW_CONFID` / `REJECT` | | `groundedness_score` | 关键事实证据支撑度 | | `critical_fact_count` | 关键事实数量 | +| `claim_checks` | 对 Executor 结构化 claims 的逐条可推导性判断 | | `facts_checked` | 逐条事实校验 | | `rationale` | 判定原因 | -| `tool_trace_summary` | 本次校验使用的证据索引 | +| `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 | +| `gatekeeper_result` | 引用真实性校验结果 | +| `composer_output` | 最终表达的解析状态和摘要 | +| `tool_trace_summary` | 本次校验使用的工具调用导航索引 | ChatService 根据 verdict 决定: -- `PASS`:输出 Executor 答案。 +- `PASS`:把允许表达的 claims 交给 Composer 输出。 - `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。 - `REJECT`:降级输出,只保留已确认信息。 +边界: + +- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。 +- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。 + ## 6. AIOps 规则自评估 AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。 @@ -238,14 +269,14 @@ Trace API 会展示: 近期优先: 1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。 -2. `not_useful` 反馈沉淀 bad case,而不是只写字段。 -3. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。 -4. AIOps 引入 LLM Verifier。 -5. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。 +2. 将 ISS-008 / ISS-009 这类 E2E 通过样例固化进 diagnosis eval fixtures。 +3. `not_useful` 反馈沉淀 bad case,而不是只写字段。 +4. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。 +5. AIOps 引入 LLM Verifier。 +6. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。 暂不优先: - 用用户反馈直接修改 session status。 - 仅凭 `evidence_score` 判断答案正确。 - 在没有人工审核时自动把 bad case 反向写入 Prompt。 - diff --git a/mvp/architecture/harness-quality-gates.md b/mvp/architecture/harness-quality-gates.md index ef9a3a4..9f53022 100644 --- a/mvp/architecture/harness-quality-gates.md +++ b/mvp/architecture/harness-quality-gates.md @@ -1,7 +1,7 @@ # Harness 与质量门禁架构 -**更新日期**:2026-07-05 -**状态**:当前可运行架构 + 后续门禁规划 +**更新日期**:2026-07-08 +**状态**:当前可运行架构 + 后续门禁规划 **参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md` ## 1. 设计目标 @@ -21,6 +21,7 @@ Prompt contract + Tool boundary + Agent hooks + Trace persistence + + Gatekeeper deterministic validation + Verifier / rule evaluation + Eval baseline ``` @@ -30,15 +31,19 @@ Prompt contract ```mermaid flowchart TB Input["User / AIOps input"] --> Prompt["Prompt contract"] - Prompt --> Agent["Planner / Executor / Verifier"] + Prompt --> Agent["Planner / Executor / Verifier / Composer"] Agent --> Tools["Evidence tools"] Tools --> Invocation["tool_invocation"] Agent --> StepHook["AgentLoggingHook"] StepHook --> Step["agent_step"] Agent --> Session["diagnosis_session"] + Invocation --> EvidenceRefs["retrieval_details.evidence_refs"] + EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"] + Agent --> Gatekeeper Invocation --> TraceSummary["ToolTraceSummaryService"] - TraceSummary --> Verifier["chat_verifier"] + Gatekeeper --> Verifier["chat_verifier"] + TraceSummary --> Verifier Verifier --> SelfEval["self_evaluation.verifier_evaluation"] Invocation --> AiOpsRule["AiOpsRuleEvaluationService"] @@ -63,15 +68,18 @@ flowchart TB | `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 | | `executor-prompt.md` | AIOps Executor 按步骤调用工具 | | `chat-planner-prompt.md` | Chat 复杂问题规划 | -| `chat-executor-prompt.md` | Chat 执行工具并形成诊断答复 | -| `chat-verifier-prompt.md` | 校验 Executor 答案是否被工具证据支撑 | +| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 | +| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 | +| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 | Prompt 层当前承担的门禁: - 禁止凭记忆回答错误码、接口定义、排障步骤。 - 需要外部信息时必须调用工具。 - 工具连续失败或返回空结果时,最终报告必须诚实说明。 -- Chat Verifier 不允许做新检索,只能校验已有证据。 +- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。 +- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。 +- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。 - AIOps payload 模式必须聚焦输入告警。 ## 4. Trace Hooks @@ -115,6 +123,7 @@ retrieval_layer l0_match_count l1_match_count retrieval_details + -> evidence_refs relevance_level dedup_reason duration_ms @@ -128,23 +137,42 @@ error_message - 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。 - 同 session 内重复文档会被 `RetrievedDocTracker` 去重。 - dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。 +- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。 +- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。 -## 6. Verifier 门禁 +## 6. Gatekeeper 与 Verifier 门禁 -Chat Verifier 的输入不是原始工具日志,而是 `ToolTraceSummaryService` 构造的证据索引。 +Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。 ```mermaid flowchart LR - Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"] + Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"] + ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"] + EvidenceRefs --> Gatekeeper + Gatekeeper --> GateResult["gatekeeper_result"] + Invocation --> Summary["ToolTraceSummaryService"] Summary --> EvidenceIndex["tool_trace_summary"] - EvidenceIndex --> Verifier["chat_verifier"] - ExecutorAnswer["executor_final_answer"] --> Verifier + GateResult --> Verifier["chat_verifier"] + ExecutorOutput --> Verifier + EvidenceIndex --> Verifier Verifier --> Verdict{"verdict"} - Verdict -->|PASS| Pass["输出原答案"] + Verdict -->|PASS| Composer["chat_composer"] + Composer --> Pass["输出最终答复"] Verdict -->|LOW_CONFID| Low["补证据或低置信输出"] Verdict -->|REJECT| Reject["降级输出"] ``` +Gatekeeper 检查: + +| 检查 | 失败语义 | +|---|---| +| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 | +| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 | +| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 | +| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 | +| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 | +| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 | + Verifier 输出: ```json @@ -152,17 +180,22 @@ Verifier 输出: "verdict": "PASS|LOW_CONFID|REJECT", "groundedness_score": 0.8, "critical_fact_count": 2, + "claim_checks": [], "facts_checked": [], "rationale": "..." } ``` +Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。 + 结果写入: ```text diagnosis_session.self_evaluation.verifier_evaluation ``` +其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary` 和 `composer_output`,用于 Trace 回放。 + ## 7. AIOps 规则门禁 AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。 @@ -197,9 +230,8 @@ diagnosis_session.self_evaluation.aiops_rule_evaluation - 工具参数 schema 校验。 - 同一工具调用次数上限。 - 工具超时的统一熔断。 -- 报告中的数值与工具返回值自动对齐校验。 +- Gatekeeper 规则三层分离:索引层、元数据层、规则实现层。 - Prompt 版本记录和回滚。 - Verifier 对 AIOps 报告的 LLM 级事实校验。 这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。 - diff --git a/mvp/architecture/interview-one-pager.md b/mvp/architecture/interview-one-pager.md index 667f259..bfca187 100644 --- a/mvp/architecture/interview-one-pager.md +++ b/mvp/architecture/interview-one-pager.md @@ -5,7 +5,7 @@ ## 1. 一句话 -SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Verifier 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。 +SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。 ## 2. 一张图 @@ -16,7 +16,7 @@ flowchart TB API --> Chat["ChatService"] API --> AiOps["AiOpsService"] - Chat --> ChatFlow["Chat: Planner -> Executor -> Verifier"] + Chat --> ChatFlow["Chat: Planner -> Executor -> Gatekeeper -> Verifier -> Composer"] AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"] ChatFlow --> Tools["Evidence Tools"] @@ -55,14 +55,14 @@ flowchart TB ```text 这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。 -Chat 复杂问题走 Planner -> Executor -> Verifier: -Planner 负责拆解,Executor 负责调用知识库、日志和指标工具,Verifier 只基于已有工具证据校验最终答案。 +Chat 复杂问题走 Planner -> Executor -> Gatekeeper -> Verifier -> Composer: +Planner 负责拆解,Executor 只负责调用知识库、日志和指标工具并提炼带证据引用的微观事实;Gatekeeper 用代码核对 invocation、raw_path 和 excerpt 是否真实;Verifier 判断这些事实能否由已验真的证据推出;Composer 只把允许表达的结论写成最终答案。 AIOps 告警入口走 Supervisor 调度 Planner/Executor: 如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。 所有过程都会落到 diagnosis_session、agent_step、tool_invocation。 -所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Verifier 怎么判定、用户最后是否反馈有用。 +所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。 ``` ## 4. 五个亮点 @@ -72,7 +72,7 @@ AIOps 告警入口走 Supervisor 调度 Planner/Executor: | 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool | | 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` | | RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 | -| 质量门禁 | Chat Verifier 校验 groundedness,AIOps rule evaluation 控制告警聚焦 | +| 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 | | 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 | ## 5. 三个关键取舍 @@ -99,15 +99,14 @@ aiops_rule_evaluation -> AIOps 报告是否聚焦告警并使用证据 | 追问 | 回答方向 | |---|---| -| 怎么防止幻觉? | Executor 必须用工具;Verifier 只基于 `tool_trace_summary` 校验;LOW_CONFID/REJECT 会降级输出 | +| 怎么防止幻觉? | Executor 输出 `executor_evidence_v2`,每个 claim 绑定 `source_invocation_id + raw_path + evidence_excerpt`;Gatekeeper 用 `tool_invocation.retrieval_details.evidence_refs` 核验引用真实性;Verifier 只判断可推导性;Composer 防止把 no-evidence 说成已排除 | | RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 | | 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter | | AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 | -| 下一步怎么演进? | evidence block、邻居 chunk、Playbook、AIOps LLM Verifier、MCP 工具协议化 | +| 下一步怎么演进? | 固化 E2E fixture、Prompt version、Gatekeeper 规则配置化、邻居 chunk、AIOps LLM Verifier、MCP 工具协议化 | ## 7. 现场演示入口 - Demo 脚本:`mvp/demo/ten-minute-interview-demo.md` - 故事案例:`interview/story-cases.md` - 架构细节:`mvp/architecture/README.md` - diff --git a/mvp/architecture/retrieval-observability.md b/mvp/architecture/retrieval-observability.md index d06f088..b6f00a6 100644 --- a/mvp/architecture/retrieval-observability.md +++ b/mvp/architecture/retrieval-observability.md @@ -141,7 +141,8 @@ post-retrieval 层再把检索候选归一为: - 给 Agent 输出 completeness hint。 - 写入 `tool_invocation.relevance_level`。 -- 给 Verifier 构造 `tool_trace_summary`。 +- 给 Gatekeeper 提供 `evidence_refs` 引用验真源。 +- 给 Verifier 构造 `tool_trace_summary` 审计导航。 - 供 EvaluationService 计算 evidence score。 ## 6. 文档切片和 metadata @@ -178,6 +179,7 @@ flowchart LR LookupResult --> Recorder["ToolInvocationRecorder"] Recorder --> Invocation["tool_invocation"] Invocation --> Trace["DiagnosisTraceService"] + Invocation --> Gatekeeper["ExecutorGatekeeperService"] Invocation --> Summary["ToolTraceSummaryService"] Summary --> Verifier["chat_verifier"] Invocation --> Eval["EvaluationService / RAG eval"] @@ -205,9 +207,25 @@ success - evidence status。 - dedup reason。 - evidence block summaries。 +- evidence refs:`raw_path + text`,用于核对 Executor 的 `evidence_excerpt`。 - context pack summary。 - rerank trace。 +其中 `evidence_refs` 是当前 Chat 证据链路的精确引用源: + +```json +{ + "evidence_refs": [ + { + "raw_path": "$.evidence_blocks[0]", + "text": "最小证据文本" + } + ] +} +``` + +如果检索返回 no evidence,应使用 `raw_path=$.no_evidence` 记录负向证据。它只能说明“本次检索没有匹配证据”,不能作为“问题不存在”的证明。 + ## 8. 去重与行动记忆 当前 session 级去重由 `RetrievedDocTracker` 负责。 diff --git a/mvp/architecture/session-trace-lifecycle.md b/mvp/architecture/session-trace-lifecycle.md index 7121086..c48d798 100644 --- a/mvp/architecture/session-trace-lifecycle.md +++ b/mvp/architecture/session-trace-lifecycle.md @@ -1,7 +1,7 @@ # 会话与 Trace 生命周期 -**更新日期**:2026-07-05 -**状态**:当前可运行架构 +**更新日期**:2026-07-08 +**状态**:当前可运行架构 **参考历史文档**:`archive/2026-07-05-legacy/session-management.md` ## 1. 定位 @@ -35,6 +35,7 @@ flowchart TD StepHook --> Step["agent_step"] Agent --> Tool["Evidence tools"] Tool --> Invocation["tool_invocation"] + Invocation --> Gatekeeper["Gatekeeper evidence validation"] Agent --> Final{"workflow result"} Final -->|success| Success["status = SUCCESS, answer saved"] @@ -132,6 +133,7 @@ retrieval_layer l0_match_count l1_match_count retrieval_details + -> evidence_refs relevance_level dedup_reason duration_ms @@ -139,7 +141,20 @@ success error_message ``` -对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对非检索工具,检索字段可以为空。 +对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对日志、指标和知识库工具,`retrieval_details.evidence_refs` 会记录 Gatekeeper 可核验的最小证据引用: + +```json +{ + "evidence_refs": [ + { + "raw_path": "$.logs[0]", + "text": "最小证据文本" + } + ] +} +``` + +当工具明确没有返回匹配证据时,可以记录 `raw_path=$.no_evidence`。该路径只表示“本次工具查询未检索到匹配证据”,不表示问题被排除。 ## 7. Trace API 聚合 @@ -163,7 +178,7 @@ Trace 视图回答的问题: - 每一步模型输入输出是什么摘要? - 调用了哪些工具? - 工具返回了什么证据? -- Verifier / AIOps rule 是否通过? +- Gatekeeper / Verifier / Composer / AIOps rule 是否通过? - 用户是否反馈有用? ## 8. Chat 与 AIOps 差异 @@ -171,7 +186,7 @@ Trace 视图回答的问题: | 维度 | Chat | AIOps | |---|---|---| | `agent_flow` | `CHAT` | `AI_OPS` | -| 编排方式 | `SequentialAgent`: Planner -> Executor -> Verifier | `SupervisorAgent`: Planner + Executor | +| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor | | 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` | | 答案字段 | Chat 最终答复 | 告警分析报告 | | payload | 用户自然语言 + history | alert payload 或 auto-discovery | @@ -193,4 +208,3 @@ Trace 视图回答的问题: 2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。 3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。 4. 为 Trace 增加导出能力,服务面试演示和回归分析。 -