refactor(harness): remove legacy agent architecture
This commit is contained in:
@@ -0,0 +1,48 @@
|
||||
# MVP 架构文档
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
|
||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||
|
||||
- `mvp/architecture/archive/2026-07-05-legacy/`
|
||||
|
||||
归档材料只作为设计历史阅读,不再作为当前实现依据。
|
||||
|
||||
## 当前文档
|
||||
|
||||
| 文档 | 用途 |
|
||||
|---|---|
|
||||
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
|
||||
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
||||
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
|
||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
||||
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
||||
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
|
||||
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
|
||||
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
|
||||
| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库文档编写与维护规范,覆盖 frontmatter、category、chunk、reindex |
|
||||
| [data-model.md](data-model.md) | 数据模型总览,覆盖 Trace、知识库、反馈沉淀和 Milvus metadata |
|
||||
| [evolution-roadmap.md](evolution-roadmap.md) | 从旧版 Agent 蓝图继承的后续演进路线,不代表当前已实现 |
|
||||
|
||||
## 当前架构一句话
|
||||
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
|
||||
## 阅读顺序
|
||||
|
||||
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
|
||||
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
|
||||
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
||||
4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
|
||||
5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||
6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||
7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||
8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
||||
9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
@@ -0,0 +1,3 @@
|
||||
# Archive Note
|
||||
|
||||
本目录保存 2026-07-22 单 Diagnosis Agent + Harness 切换前的当前架构文档。内容用于历史决策追溯,不代表现行 runtime、API 或验收口径。
|
||||
@@ -0,0 +1,237 @@
|
||||
# Agent 编排架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计定位
|
||||
|
||||
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
||||
|
||||
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
|
||||
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
||||
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
||||
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
||||
|
||||
## 2. 当前 Agent 全景
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph Chat["Chat diagnosis"]
|
||||
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
|
||||
ChatService --> ChatPlanner["chat_planner"]
|
||||
ChatPlanner --> ChatExecutor["chat_executor"]
|
||||
ChatExecutor --> ChatTools["evidence tools"]
|
||||
ChatTools --> ChatExecutor
|
||||
ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
|
||||
ChatGatekeeper --> ChatVerifier["chat_verifier"]
|
||||
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
||||
ChatDecision --> ChatComposer["chat_composer"]
|
||||
ChatComposer --> ChatAnswer["final answer"]
|
||||
end
|
||||
|
||||
subgraph AiOps["AIOps diagnosis"]
|
||||
AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
|
||||
AiOpsService --> Supervisor["ai_ops_supervisor"]
|
||||
Supervisor --> AiOpsPlanner["planner_agent"]
|
||||
Supervisor --> AiOpsExecutor["executor_agent"]
|
||||
AiOpsPlanner --> AiOpsExecutor
|
||||
AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
|
||||
AiOpsTools --> AiOpsReport["alert report"]
|
||||
AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
end
|
||||
|
||||
subgraph Trace["Trace persistence"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
SelfEval["self_evaluation"]
|
||||
end
|
||||
|
||||
ChatService --> ChatSession
|
||||
ChatService --> Run
|
||||
ChatPlanner --> Step
|
||||
ChatExecutor --> Step
|
||||
ChatGatekeeper --> SelfEval
|
||||
ChatVerifier --> Step
|
||||
ChatTools --> Invocation
|
||||
ChatDecision --> SelfEval
|
||||
ChatComposer --> Step
|
||||
|
||||
AiOpsService --> ChatSession
|
||||
AiOpsService --> Run
|
||||
AiOpsPlanner --> Step
|
||||
AiOpsExecutor --> Step
|
||||
AiOpsTools --> Invocation
|
||||
AiOpsRule --> SelfEval
|
||||
```
|
||||
|
||||
## 3. Chat 编排
|
||||
|
||||
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> lookup_knowledge / query_logs / query_metrics / date_time
|
||||
-> outputs executor_evidence_v2
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> validates source_invocation_id / raw_path / evidence_excerpt
|
||||
-> chat_verifier
|
||||
-> judges whether verified evidence can derive claims
|
||||
-> chat_composer
|
||||
-> writes final user-facing answer
|
||||
```
|
||||
|
||||
关键行为:
|
||||
|
||||
| 角色 | 当前职责 | 输出 |
|
||||
|---|---|---|
|
||||
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
||||
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
|
||||
| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
||||
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
|
||||
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
|
||||
|
||||
Chat 链路最多支持两轮验证:
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant C as ChatService
|
||||
participant P as chat_planner
|
||||
participant E as chat_executor
|
||||
participant T as tools
|
||||
participant G as gatekeeper
|
||||
participant V as chat_verifier
|
||||
participant M as chat_composer
|
||||
participant R as diagnosis_run
|
||||
|
||||
C->>P: 原始问题 + history + retry_context
|
||||
P-->>C: planner_plan
|
||||
C->>E: planner_plan + 上下文
|
||||
E->>T: 调用证据工具
|
||||
T-->>E: 证据结果
|
||||
E-->>C: executor_evidence_v2
|
||||
C->>G: executor_structured_output + tool_invocation.evidence_refs
|
||||
G-->>C: gatekeeper_result
|
||||
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
||||
V-->>C: PASS / LOW_CONFID / REJECT
|
||||
C->>R: 写入 verifier_evaluation
|
||||
alt LOW_CONFID 且允许补证据
|
||||
C->>P: retry_context: 仅补缺失证据
|
||||
else PASS 或 REJECT
|
||||
C->>M: allowed_claims + missing_info + recommended_actions
|
||||
M-->>C: composer_output
|
||||
C->>R: 保存 Composer 最终 answer
|
||||
end
|
||||
```
|
||||
|
||||
决策语义:
|
||||
|
||||
| Verdict | 行为 |
|
||||
|---|---|
|
||||
| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
|
||||
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
|
||||
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
|
||||
|
||||
## 4. AIOps 编排
|
||||
|
||||
AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
|
||||
|
||||
```text
|
||||
ai_ops_supervisor
|
||||
-> planner_agent
|
||||
-> executor_agent
|
||||
-> final report
|
||||
-> AiOpsRuleEvaluationService
|
||||
```
|
||||
|
||||
与 Chat 的差异:
|
||||
|
||||
- AIOps 的输入可能是结构化告警 payload。
|
||||
- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
|
||||
- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
|
||||
- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
|
||||
|
||||
## 5. 工具边界
|
||||
|
||||
当前 Executor 可用工具来自两类:
|
||||
|
||||
```text
|
||||
methodTools
|
||||
-> dateTimeTools
|
||||
-> lookupKnowledgeTool
|
||||
-> queryMetricsTools
|
||||
-> queryLogsTools when mock enabled
|
||||
|
||||
ToolCallbackProvider
|
||||
-> framework-discovered tools
|
||||
```
|
||||
|
||||
工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
|
||||
|
||||
- L0/L1 命中数量。
|
||||
- 检索层。
|
||||
- relevance level。
|
||||
- retrieved domains。
|
||||
- dedup reason。
|
||||
|
||||
## 6. Skill / Playbook 流程
|
||||
|
||||
当前 Skill 是诊断流程编排提示,不是事实证据来源。Planner 只能看到 `SkillRegistry.listAll()` 暴露的 name/description 元数据;Executor 才能通过 Spring AI Alibaba 官方 `SkillsAgentHook` 使用 `read_skill` 读取完整 `SKILL.md`。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Registry["SkillRegistry<br/>active skill metadata"] --> PlannerHook["PlannerSkillMetadataHook"]
|
||||
PlannerHook --> Planner["Planner<br/>metadata only"]
|
||||
Planner --> Plan["planner_plan<br/>selected_skill + steps"]
|
||||
|
||||
Registry --> ExecutorHook["SkillsAgentHook"]
|
||||
ExecutorHook --> ReadSkill["read_skill"]
|
||||
Plan --> Executor["Executor"]
|
||||
Executor --> ReadSkill
|
||||
ReadSkill --> SkillBody["SKILL.md workflow"]
|
||||
SkillBody --> Executor
|
||||
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
|
||||
EvidenceTools --> ToolTrace["tool_invocation evidence"]
|
||||
Executor --> Gatekeeper["Gatekeeper"]
|
||||
Gatekeeper --> Verifier["Verifier"]
|
||||
ToolTrace --> Verifier
|
||||
Verifier --> Composer["Composer"]
|
||||
```
|
||||
|
||||
| 角色 | Skill 可见性 | 工具权限 |
|
||||
|---|---|---|
|
||||
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
||||
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
||||
| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
|
||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
|
||||
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
|
||||
|
||||
## 7. 与旧版设计的差异
|
||||
|
||||
| 旧版设想 | 当前实现 |
|
||||
|---|---|
|
||||
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
|
||||
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
||||
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
||||
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
||||
| Skill 驱动不同诊断流程 | 当前以 Planner 元数据选择 + Executor 读取 playbook 的方式接入 |
|
||||
|
||||
## 8. 后续演进
|
||||
|
||||
当诊断场景和工具复杂度继续上升时,再考虑拆分:
|
||||
|
||||
- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
|
||||
- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
|
||||
- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
|
||||
- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
|
||||
|
||||
拆分前提:
|
||||
|
||||
- 当前 Executor prompt 已难以维护。
|
||||
- 不同故障类型的工具权限明显不同。
|
||||
- Trace 能证明某类问题需要独立的推理策略。
|
||||
- 评测集能覆盖拆分前后的行为差异。
|
||||
@@ -0,0 +1,444 @@
|
||||
# 当前 MVP 架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构
|
||||
**适用范围**:Demo、面试讲解、后续迭代规划
|
||||
|
||||
## 1. 系统定位
|
||||
|
||||
SuperBizAgent MVP 不是通用 Chatbot,而是面向故障诊断的 Agent 工程项目。
|
||||
|
||||
核心目标:
|
||||
|
||||
- 支持用户主动发起的 Chat 诊断。
|
||||
- 支持 AIOps 告警触发的自动诊断。
|
||||
- 保留 Agent 的规划、执行、验证过程。
|
||||
- 工具调用必须显式、可追踪、可回放。
|
||||
- RAG 检索必须通过 `lookup_knowledge` 暴露证据链。
|
||||
- 每次诊断都沉淀 session、step、tool invocation 和 self evaluation。
|
||||
|
||||
## 2. 总体分层
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph API["API Layer"]
|
||||
ChatController["ChatController"]
|
||||
TraceController["DiagnosisTraceController"]
|
||||
SearchController["SearchController"]
|
||||
DocumentController["DocumentController"]
|
||||
end
|
||||
|
||||
subgraph App["Application Service"]
|
||||
ChatService["ChatService"]
|
||||
AiOpsService["AiOpsService"]
|
||||
TraceService["DiagnosisTraceService"]
|
||||
end
|
||||
|
||||
subgraph Agent["Agent Orchestration"]
|
||||
Supervisor["Supervisor"]
|
||||
Planner["Planner"]
|
||||
Executor["Executor"]
|
||||
Gatekeeper["Gatekeeper"]
|
||||
Verifier["Verifier"]
|
||||
Composer["Composer"]
|
||||
end
|
||||
|
||||
subgraph Tools["Evidence Tools"]
|
||||
KnowledgeTool["lookup_knowledge"]
|
||||
LogsTool["query_logs"]
|
||||
MetricsTool["query_metrics"]
|
||||
AlertsTool["queryPrometheusAlerts"]
|
||||
end
|
||||
|
||||
subgraph Skills["Skill / Playbook"]
|
||||
SkillRegistry["SkillRegistry"]
|
||||
PlannerSkillHook["PlannerSkillMetadataHook"]
|
||||
SkillsHook["SkillsAgentHook"]
|
||||
ReadSkill["read_skill"]
|
||||
end
|
||||
|
||||
subgraph RAG["RAG Retrieval"]
|
||||
L0["KnowledgeIndexService"]
|
||||
VectorSearch["VectorSearchService"]
|
||||
VectorStore["Spring AI VectorStore"]
|
||||
SdkFallback["Milvus SDK fallback"]
|
||||
end
|
||||
|
||||
subgraph Store["Persistence and Trace"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
ApiDoc["api_document"]
|
||||
Milvus["Milvus/Zilliz"]
|
||||
end
|
||||
|
||||
API --> App
|
||||
ChatService --> Agent
|
||||
AiOpsService --> Agent
|
||||
SkillRegistry --> PlannerSkillHook
|
||||
PlannerSkillHook --> Planner
|
||||
SkillRegistry --> SkillsHook
|
||||
SkillsHook --> Executor
|
||||
Executor --> ReadSkill
|
||||
Agent --> Tools
|
||||
KnowledgeTool --> RAG
|
||||
RAG --> Store
|
||||
Tools --> Invocation
|
||||
Agent --> Step
|
||||
App --> Session
|
||||
TraceService --> Session
|
||||
TraceService --> Step
|
||||
TraceService --> Invocation
|
||||
```
|
||||
|
||||
```text
|
||||
API Layer
|
||||
-> ChatController
|
||||
-> DiagnosisTraceController
|
||||
-> SearchController
|
||||
-> DocumentController
|
||||
|
||||
Application Service
|
||||
-> ChatService
|
||||
-> AiOpsService
|
||||
-> DiagnosisTraceService
|
||||
|
||||
Agent Orchestration
|
||||
-> Supervisor
|
||||
-> Planner
|
||||
-> Executor
|
||||
-> Gatekeeper
|
||||
-> Verifier
|
||||
-> Composer
|
||||
|
||||
Evidence Tools
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> queryPrometheusAlerts
|
||||
|
||||
Skill / Playbook
|
||||
-> SkillRegistry
|
||||
-> PlannerSkillMetadataHook gives Planner name/description only
|
||||
-> SkillsAgentHook gives Executor read_skill
|
||||
-> Verifier is isolated from skills
|
||||
|
||||
RAG Retrieval
|
||||
-> KnowledgeIndexService
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
|
||||
Persistence
|
||||
-> chat_session
|
||||
-> diagnosis_run
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
-> api_document
|
||||
-> Milvus/Zilliz collection
|
||||
|
||||
Quality Gates
|
||||
-> executor gatekeeper
|
||||
-> chat verifier
|
||||
-> AIOps rule evaluation
|
||||
-> diagnosis eval baseline
|
||||
-> RAG retrieval baseline
|
||||
```
|
||||
|
||||
## 3. Chat 诊断链路
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
actor User as 用户
|
||||
participant API as POST /api/chat
|
||||
participant Chat as ChatService
|
||||
participant Planner as Planner Agent
|
||||
participant Executor as Executor Agent
|
||||
participant Tool as Evidence Tools
|
||||
participant Gatekeeper as Gatekeeper Hook
|
||||
participant Verifier as Verifier Agent
|
||||
participant Composer as Composer Agent
|
||||
participant DB as Trace Tables
|
||||
participant Trace as Trace API
|
||||
|
||||
User->>API: 提交诊断问题
|
||||
API->>Chat: execute chat strategy
|
||||
Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
|
||||
Chat->>Planner: 复杂问题进入规划
|
||||
Planner->>DB: 写入 agent_step.run_id
|
||||
Planner->>Executor: 下发排查方向
|
||||
Executor->>Tool: lookup_knowledge / logs / metrics
|
||||
Tool->>DB: 写入 tool_invocation.run_id
|
||||
Tool-->>Executor: 返回证据
|
||||
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
||||
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
||||
Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
|
||||
Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
||||
Composer->>Chat: 生成最终用户答复
|
||||
Chat->>DB: 保存 diagnosis_run.answer
|
||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
|
||||
Trace->>DB: 聚合 run / step / tool
|
||||
Trace-->>User: 返回可回放诊断链路
|
||||
```
|
||||
|
||||
```text
|
||||
POST /api/chat
|
||||
-> ChatService
|
||||
-> 简单问题:轻量回答
|
||||
-> 复杂诊断:Agent 编排
|
||||
-> Planner 制定排查方向
|
||||
-> Executor 调用证据工具
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> Gatekeeper 校验 Executor 证据引用真实性
|
||||
-> Verifier 判断 claim 是否能由已核验证据推出
|
||||
-> Composer 生成最终用户答复
|
||||
-> 保存 chat_session metadata
|
||||
-> 保存 diagnosis_run
|
||||
-> 保存 agent_step.run_id
|
||||
-> 保存 tool_invocation.run_id
|
||||
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||
|
||||
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
||||
|
||||
关键代码:
|
||||
|
||||
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
|
||||
## 4. AIOps 诊断链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Request["POST /api/ai_ops"] --> Payload{"包含告警 payload?"}
|
||||
Payload -->|是| Targeted["PAYLOAD_TARGETED"]
|
||||
Payload -->|否| Discovery["AUTO_DISCOVERY"]
|
||||
|
||||
Targeted --> BuildPrompt["构造聚焦 payload 的诊断 prompt"]
|
||||
Targeted --> QueryAug["生成 recommended lookup_knowledge query"]
|
||||
Discovery --> DiscoverAlert["通过 queryPrometheusAlerts 发现活跃告警"]
|
||||
|
||||
BuildPrompt --> Plan["Planner 规划排查"]
|
||||
QueryAug --> Plan
|
||||
DiscoverAlert --> Plan
|
||||
|
||||
Plan --> Execute["Executor 收集证据"]
|
||||
Execute --> Knowledge["lookup_knowledge"]
|
||||
Execute --> Metrics["query_metrics / Prometheus"]
|
||||
Execute --> Logs["query_logs"]
|
||||
|
||||
Knowledge --> Report["告警分析报告"]
|
||||
Metrics --> Report
|
||||
Logs --> Report
|
||||
|
||||
Report --> RuleEval["AiOpsRuleEvaluationService"]
|
||||
RuleEval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
|
||||
Report --> Trace["DiagnosisTraceService"]
|
||||
SelfEval --> Trace
|
||||
```
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> AiOpsService
|
||||
-> 判断是否有告警 payload
|
||||
-> PAYLOAD_TARGETED
|
||||
-> AUTO_DISCOVERY
|
||||
-> 构造 AIOps 诊断 prompt
|
||||
-> payload 模式补充 recommended lookup_knowledge query
|
||||
-> Agent 编排
|
||||
-> Planner / Executor
|
||||
-> Prometheus / logs / knowledge tools
|
||||
-> 生成告警分析报告
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
-> Trace API 可查看全链路
|
||||
```
|
||||
|
||||
AIOps 保留两种模式:
|
||||
|
||||
| 模式 | 触发条件 | 行为 |
|
||||
|---|---|---|
|
||||
| `PAYLOAD_TARGETED` | 请求包含 alertName、service、severity、description、timeRange 等字段 | 以 payload 为唯一主诊断对象,并生成推荐知识库 query |
|
||||
| `AUTO_DISCOVERY` | 请求没有明确告警 payload | 先查询当前活跃告警,再选择目标排查 |
|
||||
|
||||
AIOps 当前使用轻量规则验证器,重点检查:
|
||||
|
||||
- 最终报告是否存在。
|
||||
- payload 模式是否聚焦输入告警。
|
||||
- 是否使用关键证据工具,例如 `lookup_knowledge`、日志、指标。
|
||||
|
||||
关键代码:
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
|
||||
|
||||
## 5. RAG 位置
|
||||
|
||||
RAG 不是隐藏在 Chat Advisor 里的隐式能力,而是 Executor 可以显式调用的工具:
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Executor["Executor Agent"] --> Tool["lookup_knowledge Tool"]
|
||||
Tool --> L0["L0 domain/entity hint"]
|
||||
Tool --> Search["VectorSearchService"]
|
||||
L0 --> Search
|
||||
Search --> VectorStore["Spring AI VectorStore"]
|
||||
Search --> Fallback["Milvus SDK fallback"]
|
||||
VectorStore --> Normalize["score/rawScore/scoreLabel"]
|
||||
Fallback --> Normalize
|
||||
Normalize --> Evidence["evidence output"]
|
||||
Evidence --> Invocation["tool_invocation"]
|
||||
Evidence --> Executor
|
||||
```
|
||||
|
||||
```text
|
||||
Executor
|
||||
-> lookup_knowledge(query)
|
||||
-> L0 domain/entity hint
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
-> evidence shaping
|
||||
-> tool_invocation
|
||||
```
|
||||
|
||||
保留显式工具的原因:
|
||||
|
||||
- Agent 何时检索、检索什么、证据是什么,必须能在 trace 中解释。
|
||||
- AIOps payload 到 query 的业务映射需要项目内控制。
|
||||
- `tool_invocation` 是后续评测、回放和面试讲解的核心材料。
|
||||
|
||||
RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细节见 [retrieval-observability.md](retrieval-observability.md)。
|
||||
|
||||
## 6. 持久化模型
|
||||
|
||||
当前诊断持久化以 session/run/trace 明细为核心:
|
||||
|
||||
```text
|
||||
chat_session
|
||||
-> 多轮会话目录和元数据
|
||||
-> session_id / status / message_pair_count
|
||||
|
||||
diagnosis_run
|
||||
-> 一次诊断运行的主记录
|
||||
-> run_id / session_id
|
||||
-> query / status / agent_flow / answer
|
||||
-> self_evaluation
|
||||
-> step_count / tool_call_count / duration
|
||||
|
||||
agent_step
|
||||
-> Agent 模型调用步骤
|
||||
-> session_id / run_id
|
||||
-> step_index / agent_name
|
||||
-> model_input / model_output / thought
|
||||
-> duration / token_count
|
||||
|
||||
tool_invocation
|
||||
-> 工具调用事实
|
||||
-> session_id / run_id
|
||||
-> tool_name / input_params / output_preview
|
||||
-> retrieval_layer / retrieval_details
|
||||
-> retrieval_details.evidence_refs
|
||||
-> relevance_level / dedup_reason
|
||||
-> duration / success
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- 旧的 `diagnosis_record` 已不是当前主模型。
|
||||
- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
|
||||
- `api_document` 仍用于文档元数据管理。
|
||||
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
||||
|
||||
会话和 Trace 生命周期见 [session-trace-lifecycle.md](session-trace-lifecycle.md),完整数据关系见 [data-model.md](data-model.md)。
|
||||
|
||||
## 7. Trace API
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
Trace API 聚合:
|
||||
|
||||
- 会话元数据、运行状态和最终报告。
|
||||
- Agent step 序列。
|
||||
- 工具调用和检索细节。
|
||||
- Chat Gatekeeper / Verifier / Composer 结果。
|
||||
- AIOps rule evaluation 结果。
|
||||
|
||||
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
|
||||
|
||||
Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||
|
||||
## 8. 质量门禁
|
||||
|
||||
当前质量门禁分层如下:
|
||||
|
||||
| 门禁 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
|
||||
| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
|
||||
| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
|
||||
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
|
||||
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
|
||||
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
|
||||
| Live RAG Acceptance | `scripts/eval_rag_live_acceptance.py` | 在运行环境中验证重建索引后的真实检索 |
|
||||
|
||||
## 9. 当前完成状态
|
||||
|
||||
已经完成:
|
||||
|
||||
- Chat 和 AIOps 两条入口链路。
|
||||
- 显式 `lookup_knowledge` Agent Tool。
|
||||
- L0 从最终决策降级为 domain/entity hint。
|
||||
- `VectorSearchService` 作为稳定检索门面。
|
||||
- Spring AI VectorStore 读取路径。
|
||||
- Milvus SDK fallback。
|
||||
- `score` / `rawScore` / `scoreLabel` 分数语义拆分。
|
||||
- `title`、`breadcrumb`、`content` 参与 embedding 文本。
|
||||
- `tool_invocation` 记录检索层、relevance level、dedup reason。
|
||||
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
|
||||
- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
|
||||
- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
|
||||
- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
|
||||
- Verifier 只判断可推导性。
|
||||
- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
|
||||
- RAG offline baseline 和 live acceptance 脚本。
|
||||
|
||||
暂不作为当前已完成能力声明:
|
||||
|
||||
- 完整 QueryTransformer / MultiQuery。
|
||||
- BM25、RRF、cross-encoder rerank。
|
||||
- 完整邻居 chunk / section context expansion。
|
||||
- VectorStore 写入路径全面迁移。
|
||||
- 完整 LLM-based AIOps verifier。
|
||||
|
||||
后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
|
||||
|
||||
## 10. 关键代码索引
|
||||
|
||||
| 能力 | 代码 |
|
||||
|---|---|
|
||||
| Chat 入口与编排 | `ChatController`, `ChatService` |
|
||||
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
|
||||
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
|
||||
| 知识库工具 | `LookupKnowledgeTool` |
|
||||
| L0 hint | `KnowledgeIndexService` |
|
||||
| 向量检索门面 | `VectorSearchService` |
|
||||
| 文档切片 | `DocumentChunkService` |
|
||||
| 向量写入 | `VectorIndexService` |
|
||||
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
|
||||
| Trace 聚合 | `DiagnosisTraceService` |
|
||||
| 工具调用记录 | `ToolInvocationRecorder` |
|
||||
| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
|
||||
| self_evaluation 合并 | `SelfEvaluationMergeService` |
|
||||
@@ -0,0 +1,188 @@
|
||||
# 数据模型总览
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
|
||||
## 1. 定位
|
||||
|
||||
本文从架构角度说明当前 MVP 的核心数据模型。详细字段以 Flyway migration、实体类和 `mvp/tables/` 为准。
|
||||
|
||||
核心数据分三组:
|
||||
|
||||
- 会话与诊断 Trace:`chat_session`、`diagnosis_run`、`agent_step`、`tool_invocation`
|
||||
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
|
||||
- 反馈沉淀:`case_library`
|
||||
|
||||
`diagnosis_session` 仍保留为历史兼容和回滚表,不再是新执行写入的主模型。
|
||||
|
||||
## 2. 总体关系
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
chat_session ||--o{ diagnosis_run : owns
|
||||
diagnosis_run ||--o{ agent_step : has
|
||||
diagnosis_run ||--o{ tool_invocation : has
|
||||
diagnosis_run ||--o| case_library : creates_when_useful
|
||||
api_document ||--o{ milvus_chunk : indexed_as
|
||||
knowledge_domain ||--o{ api_document : groups
|
||||
|
||||
chat_session {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar status
|
||||
int message_pair_count
|
||||
datetime last_active_at
|
||||
datetime expires_at
|
||||
}
|
||||
|
||||
diagnosis_run {
|
||||
bigint id
|
||||
varchar run_id
|
||||
varchar session_id
|
||||
text query
|
||||
varchar status
|
||||
varchar agent_flow
|
||||
longtext answer
|
||||
json self_evaluation
|
||||
varchar feedback
|
||||
}
|
||||
|
||||
agent_step {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar run_id
|
||||
int step_index
|
||||
varchar agent_name
|
||||
text model_input
|
||||
text model_output
|
||||
boolean has_tool_call
|
||||
}
|
||||
|
||||
tool_invocation {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar run_id
|
||||
bigint step_id
|
||||
varchar tool_name
|
||||
json input_params
|
||||
text output_preview
|
||||
json retrieval_details
|
||||
}
|
||||
|
||||
case_library {
|
||||
bigint id
|
||||
varchar case_id
|
||||
varchar diagnosis_id
|
||||
varchar source_type
|
||||
varchar fault_category
|
||||
text root_cause
|
||||
text solution
|
||||
}
|
||||
```
|
||||
|
||||
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
|
||||
|
||||
## 3. 会话与运行模型
|
||||
|
||||
### chat_session
|
||||
|
||||
`chat_session` 是会话目录表,保存 `sessionId` 的元数据:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `session_id` | 外部会话 ID,用于多轮上下文和 run 列表 |
|
||||
| `status` | 会话目录状态 |
|
||||
| `message_pair_count` | Redis 对话轮次数快照 |
|
||||
| `last_active_at` | 最近活跃时间 |
|
||||
| `expires_at` | 可为空的目录 TTL 元数据 |
|
||||
|
||||
它不保存完整对话历史,正文消息仍由 Redis `SessionContext.messageHistory` 管理。
|
||||
|
||||
### diagnosis_run
|
||||
|
||||
`diagnosis_run` 是一次可回放诊断执行的主记录:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `run_id` | 运行 ID,格式为 `run-` + UUID |
|
||||
| `session_id` | 所属 `chat_session.session_id` |
|
||||
| `query` | 本次 Chat 问题或 AIOps 告警摘要 |
|
||||
| `status` | 本次执行状态 |
|
||||
| `agent_flow` | `CHAT` / `AI_OPS` |
|
||||
| `answer` | 本次运行最终答复或告警报告 |
|
||||
| `self_evaluation` | 本次运行的 rule/verifier/aiops 自评估容器 |
|
||||
| `feedback` | 本次运行的用户反馈 |
|
||||
|
||||
同一个 `sessionId` 可以有多个 `runId`。Trace、反馈、评测和案例沉淀都应优先使用 `runId`,避免多轮同 session 下的数据混合。
|
||||
|
||||
## 4. Trace 明细模型
|
||||
|
||||
### agent_step
|
||||
|
||||
`agent_step` 记录模型调用步骤。新写入同时保留 `session_id` 和 `run_id`,其中 `run_id` 是回放边界。Trace 页面和评测应先按 `run_id` 隔离取数,展示顺序以 Trace API 返回顺序为准。
|
||||
|
||||
### tool_invocation
|
||||
|
||||
`tool_invocation` 记录显式工具调用事实。`retrieval_details.evidence_refs` 是 Chat 证据链路的关键字段:
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
|
||||
|
||||
## 5. 反馈沉淀模型
|
||||
|
||||
`useful` 反馈会触发 `CaseLibraryService.createFromRun`。
|
||||
|
||||
当前自动映射:
|
||||
|
||||
| 字段 | 来源 |
|
||||
|---|---|
|
||||
| `case_id` | UUID |
|
||||
| `diagnosis_id` | 新数据为 `diagnosis_run.run_id`;历史数据可能为 `diagnosis_session.session_id` |
|
||||
| `source_type` | `AUTO` |
|
||||
| `fault_category` | 当前默认 `GENERAL` |
|
||||
| `title` | run query 前 100 字符 |
|
||||
| `root_cause` | run answer |
|
||||
| `solution` | run answer |
|
||||
| `created_by` | `system` |
|
||||
|
||||
## 6. self_evaluation 结构
|
||||
|
||||
`diagnosis_run.self_evaluation` 是运行级 JSON 容器:
|
||||
|
||||
```json
|
||||
{
|
||||
"rule_evaluation": {},
|
||||
"verifier_evaluation": {},
|
||||
"aiops_rule_evaluation": {}
|
||||
}
|
||||
```
|
||||
|
||||
Chat 通常写入 `rule_evaluation` 和 `verifier_evaluation`;AIOps 写入 `aiops_rule_evaluation`。
|
||||
|
||||
## 7. 当前边界和后续
|
||||
|
||||
当前边界:
|
||||
|
||||
- `chat_session` 只存会话元数据,不存完整正文历史。
|
||||
- `diagnosis_run` 存一次运行的长期审计状态。
|
||||
- `agent_step.run_id` 和 `tool_invocation.run_id` 是 Trace、Verifier、Eval 的运行边界。
|
||||
- 当前实现主要使用逻辑关联,不依赖数据库外键。
|
||||
- `case_library.diagnosis_id` 是过渡字段,新值按 `run_id` 解释,旧值可能按 `session_id` 解释。
|
||||
- `diagnosis_session` 只作为历史兼容和回滚表保留。
|
||||
|
||||
后续可增强:
|
||||
|
||||
1. 强化 `tool_invocation.step_id` 关联。
|
||||
2. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。
|
||||
3. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
||||
@@ -0,0 +1,179 @@
|
||||
# Agent 架构演进路线
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:后续演进设计,不代表当前已实现
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 为什么需要演进路线
|
||||
|
||||
旧版 `agent-architecture.md` 包含很多生产级设想:专科 SubAgent、Skill 体系、进程隔离、回退路由、MCP 工具协议化、进化引擎。它们不应作为当前 MVP 事实写入主架构,但可以作为后续扩展路线。
|
||||
|
||||
当前原则:
|
||||
|
||||
- 当前文档只声明已经可运行或明确落地的能力。
|
||||
- 演进路线记录未来方向和触发条件。
|
||||
- 每个演进项必须有可验证收益,不能只因为“架构更炫”就拆。
|
||||
|
||||
## 2. 演进总图
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
MVP["Current MVP: Planner + Executor + Verifier"] --> Split{"Executor 是否过载?"}
|
||||
Split -->|是| SubAgents["专科 SubAgent"]
|
||||
Split -->|否| Keep["继续强化通用 Executor"]
|
||||
|
||||
SubAgents --> Skills["Skill / Playbook 体系"]
|
||||
Skills --> Fallback["回退路由"]
|
||||
Fallback --> Isolation["进程或 Pod 隔离"]
|
||||
|
||||
MVP --> ToolGrowth{"工具数量和来源是否增长?"}
|
||||
ToolGrowth -->|是| MCP["MCP / Tool Server 协议化"]
|
||||
ToolGrowth -->|否| ToolCallbacks["继续使用 @Tool / ToolCallback"]
|
||||
|
||||
MVP --> EvalGrowth{"评测数据是否足够?"}
|
||||
EvalGrowth -->|是| Evolution["Prompt / Skill 进化引擎"]
|
||||
EvalGrowth -->|否| Baseline["先扩大 baseline"]
|
||||
```
|
||||
|
||||
## 3. 专科 SubAgent
|
||||
|
||||
### 触发条件
|
||||
|
||||
- Executor prompt 变得臃肿,难以同时覆盖接口、数据库、缓存、网络等场景。
|
||||
- 不同故障类型需要明显不同的工具权限。
|
||||
- Trace 显示某些场景经常走错排查路径。
|
||||
- 评测集已经能衡量拆分前后的收益。
|
||||
|
||||
### 候选 SubAgent
|
||||
|
||||
| SubAgent | 场景 | 工具倾向 |
|
||||
|---|---|---|
|
||||
| `ExternalApiSubAgent` | 错误码、接口参数、第三方调用失败 | `lookup_knowledge`, logs, trace |
|
||||
| `DatabaseSubAgent` | 连接池、慢 SQL、死锁、数据库不可用 | metrics, logs, knowledge |
|
||||
| `CacheSubAgent` | Redis 超时、热点 key、内存风险 | metrics, logs, knowledge |
|
||||
| `GenericDiagnosisSubAgent` | 兜底诊断 | 全量只读证据工具 |
|
||||
|
||||
### 不立即拆分的原因
|
||||
|
||||
- 当前 MVP 的工具规模还可由通用 Executor 管理。
|
||||
- 过早拆分会增加 Prompt、评测和 trace 分析成本。
|
||||
- 没有足够分类评测前,拆分可能只是移动复杂度。
|
||||
|
||||
## 4. Skill / Playbook 体系
|
||||
|
||||
旧版设计中的 Skill 可以在当前项目中演进为可版本化的诊断 Playbook。
|
||||
|
||||
```text
|
||||
fault_category
|
||||
-> playbook
|
||||
-> required evidence
|
||||
-> tool sequence
|
||||
-> stop condition
|
||||
-> report template
|
||||
-> evaluation checks
|
||||
```
|
||||
|
||||
优先落地方向:
|
||||
|
||||
- AIOps 告警处理 Playbook。
|
||||
- 支付超时 Playbook。
|
||||
- MySQL 连接池风险 Playbook。
|
||||
- Redis timeout Playbook。
|
||||
|
||||
落地前提:
|
||||
|
||||
- 每个 Playbook 至少有 3-5 个 eval case。
|
||||
- Playbook 失败时可以回退到通用 Executor。
|
||||
- Trace 中能标记使用了哪个 Playbook 和哪个版本。
|
||||
|
||||
## 5. 回退路由
|
||||
|
||||
当前 Chat 已有低置信补证据和 REJECT 降级输出。后续如果引入 SubAgent,可扩展为:
|
||||
|
||||
```text
|
||||
Specialized SubAgent
|
||||
-> failed / low confidence
|
||||
-> another specialized SubAgent
|
||||
-> GenericDiagnosisSubAgent
|
||||
-> degraded answer with confirmed facts only
|
||||
```
|
||||
|
||||
回退依据:
|
||||
|
||||
- 工具连续失败。
|
||||
- Verifier `REJECT`。
|
||||
- Verifier `LOW_CONFID` 且补证据失败。
|
||||
- Agent 输出缺失关键报告字段。
|
||||
|
||||
## 6. 进程隔离
|
||||
|
||||
当前所有 Agent 在同一 JVM 内运行。生产级隔离可以考虑:
|
||||
|
||||
```text
|
||||
API service
|
||||
-> Supervisor service
|
||||
-> Planner service
|
||||
-> SubAgent services
|
||||
-> Verifier service
|
||||
```
|
||||
|
||||
触发条件:
|
||||
|
||||
- 某类 Agent 需要独立扩缩容。
|
||||
- 某类工具依赖不稳定,可能拖垮主应用。
|
||||
- 不同 Agent 需要不同权限和网络访问策略。
|
||||
- 单 JVM 内资源隔离不足。
|
||||
|
||||
MVP 阶段暂不拆分进程,优先保证 trace、评测和工具边界清晰。
|
||||
|
||||
## 7. MCP / Tool Server 协议化
|
||||
|
||||
当前工具主要通过 `@Tool`、`methodTools` 和 `ToolCallbackProvider` 暴露。工具数量增加后,可演进为:
|
||||
|
||||
```text
|
||||
Agent
|
||||
-> Tool registry
|
||||
-> MCP / tool server
|
||||
-> log server
|
||||
-> metrics server
|
||||
-> knowledge server
|
||||
-> ticket/change server
|
||||
```
|
||||
|
||||
收益:
|
||||
|
||||
- 工具独立部署。
|
||||
- 新工具上线不必重发主应用。
|
||||
- 不同 Agent 可获得不同工具子集。
|
||||
- 工具调用协议统一,更利于审计。
|
||||
|
||||
风险:
|
||||
|
||||
- 调用链更长。
|
||||
- 权限和超时治理更复杂。
|
||||
- 本地开发和 Demo 成本上升。
|
||||
|
||||
## 8. 进化引擎
|
||||
|
||||
旧版文档提到从诊断中学习。当前可以拆成更务实的步骤:
|
||||
|
||||
1. 先扩大 diagnosis eval 和 RAG eval。
|
||||
2. 从失败 trace 中标注 bad case。
|
||||
3. 将高频失败沉淀为 Playbook 或 Prompt 规则。
|
||||
4. 对 Prompt 版本做离线对比。
|
||||
5. 足够稳定后再考虑线上 A/B。
|
||||
|
||||
不建议 MVP 直接做自动 Prompt 自优化。没有可靠评测和回滚机制时,自动优化更容易引入不可解释变化。
|
||||
|
||||
## 9. 演进优先级
|
||||
|
||||
| 优先级 | 项目 | 原因 |
|
||||
|---|---|---|
|
||||
| P0 | 扩大 eval baseline | 没有评测,拆任何架构都难以证明收益 |
|
||||
| P1 | Playbook 化高频故障 | 可控、可解释、比拆 SubAgent 更轻 |
|
||||
| P1 | 完整 evidence block | 提升 Verifier 和 Trace 质量 |
|
||||
| P2 | 专科 SubAgent | 等问题类型和工具权限差异足够明显 |
|
||||
| P2 | AIOps LLM Verifier | 规则门禁不足时再引入 |
|
||||
| P3 | MCP 工具协议化 | 工具来源复杂后再做 |
|
||||
| P3 | 进程隔离 | 生产负载和权限隔离需要明确后再做 |
|
||||
|
||||
@@ -0,0 +1,443 @@
|
||||
# Chat Evidence Pipeline Contracts
|
||||
|
||||
**状态**:当前实现
|
||||
**更新日期**:2026-07-08
|
||||
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
|
||||
|
||||
当前 Chat 复杂诊断链路是:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> chat_verifier
|
||||
-> chat_composer
|
||||
-> final answer
|
||||
```
|
||||
|
||||
设计原则:
|
||||
|
||||
- Planner 暂不输出 `scope_contract`。
|
||||
- Executor 只做证据收集和微观事实提炼,不生成最终用户答案。
|
||||
- Gatekeeper 在 Verifier 前做代码级引用真实性校验。
|
||||
- Verifier 判断 claim 是否能由已核验证据推出。
|
||||
- Composer 只表达 Verifier 允许输出的内容。
|
||||
|
||||
---
|
||||
|
||||
## 1. Planner
|
||||
|
||||
Planner 当前保持不变,输出 `planner_plan`:
|
||||
|
||||
```json
|
||||
{
|
||||
"selected_skill": "diagnose-mysql-connection-pool",
|
||||
"selection_reason": "选择该 skill 的原因",
|
||||
"plan": ["步骤1", "步骤2"],
|
||||
"reasoning": "规划思路"
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `selected_skill` | string/null | Planner 选择的诊断 skill 名称 |
|
||||
| `selection_reason` | string | skill 选择理由 |
|
||||
| `plan` | array | 给 Executor 的执行步骤 |
|
||||
| `reasoning` | string | 规划思路说明 |
|
||||
|
||||
当前边界:
|
||||
|
||||
- 不新增 `scope_contract`。
|
||||
- 不要求 Planner 显式列出 forbidden actions。
|
||||
- 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。
|
||||
|
||||
---
|
||||
|
||||
## 2. Executor
|
||||
|
||||
Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。
|
||||
|
||||
### 2.1 输出结构
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "observation",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 必填 | 定义 |
|
||||
|---|---|---:|---|
|
||||
| `answer_version` | string | 是 | 固定为 `executor_evidence_v2` |
|
||||
| `claims` | array | 是 | Executor 提出的待验证事实断言 |
|
||||
| `claims[].claim_id` | string | 是 | claim 标识 |
|
||||
| `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 |
|
||||
| `claims[].claim_text` | string | 是 | 事实断言文本 |
|
||||
| `claims[].support_level` | string | 是 | `direct` 或 `indirect` |
|
||||
| `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 |
|
||||
| `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` |
|
||||
| `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 |
|
||||
| `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 |
|
||||
| `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 |
|
||||
| `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 |
|
||||
| `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 |
|
||||
| `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact |
|
||||
| `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 |
|
||||
| `missing_info` | array | 是 | 无法确认结论所缺少的证据 |
|
||||
|
||||
禁止字段:
|
||||
|
||||
- `diagnosis_summary`
|
||||
- `user_facing_answer`
|
||||
- `source_invocation_ids` 作为主引用字段
|
||||
|
||||
### 2.2 raw_path
|
||||
|
||||
当前支持的精确路径:
|
||||
|
||||
| 工具 | 正向证据路径 | 负向证据路径 |
|
||||
|---|---|---|
|
||||
| `query_metrics` | `$.alerts[i]` | `$.no_evidence` |
|
||||
| `query_logs` | `$.logs[i]` | `$.no_evidence` |
|
||||
| `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` |
|
||||
|
||||
约束:
|
||||
|
||||
- `raw_path` 必须指向数组条目或 `$.no_evidence`。
|
||||
- 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。
|
||||
- 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。
|
||||
|
||||
### 2.3 negative_observation
|
||||
|
||||
当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`:
|
||||
|
||||
```json
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "negative_observation",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.no_evidence",
|
||||
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
语义边界:
|
||||
|
||||
- `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。
|
||||
- 不表示“问题绝对不存在”。
|
||||
- 不表示“根因被排除”。
|
||||
- 不表示“系统已经健康”。
|
||||
- `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。
|
||||
|
||||
### 2.4 窄范围任务
|
||||
|
||||
窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。
|
||||
|
||||
Executor 必须遵守:
|
||||
|
||||
- 只输出 `observation` / `negative_observation`。
|
||||
- claim 数量通常 1 条,最多 2 条。
|
||||
- claim 数量限制不限制 `evidence_bindings` 数量。
|
||||
- 不输出根因、风险、修复建议、经验推断。
|
||||
- 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。
|
||||
- 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。
|
||||
|
||||
---
|
||||
|
||||
## 3. Tool Invocation Evidence Refs
|
||||
|
||||
工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。
|
||||
|
||||
### 3.1 正向证据
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 3.2 负向证据
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "no_evidence",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.no_evidence",
|
||||
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` |
|
||||
| `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 |
|
||||
| `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 |
|
||||
|
||||
---
|
||||
|
||||
## 4. Gatekeeper
|
||||
|
||||
Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。
|
||||
|
||||
### 4.1 输入
|
||||
|
||||
- `sessionId`
|
||||
- `executor_structured_output`
|
||||
- 当前 session 的 `tool_invocation`
|
||||
|
||||
### 4.2 输出
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.no_evidence",
|
||||
"matched_text": "query_logs returned no evidence; ...",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `status` | string | `pass` 或 `fail` |
|
||||
| `severity` | string | `none`、`low_confid`、`reject` |
|
||||
| `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 |
|
||||
| `rules` | array | 已启用规则的轻量元数据摘要 |
|
||||
| `checked_bindings` | array | 每条证据绑定的校验结果 |
|
||||
| `failed_rules` | array | 失败规则 id |
|
||||
| `warnings` | array | 自动回填等非阻断信息 |
|
||||
| `errors` | array | 失败明细 |
|
||||
|
||||
校验规则:
|
||||
|
||||
- `answer_version` 必须是 `executor_evidence_v2`。
|
||||
- 不允许 `diagnosis_summary` / `user_facing_answer`。
|
||||
- 每个 claim 必须有非空 `evidence_bindings`。
|
||||
- `tool_name` 必须和真实 invocation 对齐。
|
||||
- `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。
|
||||
- `raw_path` 必须存在于 `retrieval_details.evidence_refs`。
|
||||
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
|
||||
- `negative_observation` 只能绑定 `$.no_evidence`。
|
||||
|
||||
规则配置:
|
||||
|
||||
- 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。
|
||||
- 规则实现仍是确定性 Java 代码,不执行动态脚本。
|
||||
- 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。
|
||||
|
||||
失败分级:
|
||||
|
||||
| 场景 | severity |
|
||||
|---|---|
|
||||
| 伪造 invocation id | `reject` |
|
||||
| tool_name 与 invocation 不匹配 | `reject` |
|
||||
| raw_path 不存在 | `reject` |
|
||||
| excerpt 与 matched_text 不匹配 | `reject` |
|
||||
| negative_observation 绑定正向日志 | `reject` |
|
||||
| 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` |
|
||||
| 旧 invocation 没有 `evidence_refs` | `low_confid` |
|
||||
|
||||
---
|
||||
|
||||
## 5. Verifier
|
||||
|
||||
Verifier 输入由 `VerifierInputHook` 构造:
|
||||
|
||||
```json
|
||||
{
|
||||
"original_query": "用户原始问题",
|
||||
"executor_final_answer": "{...executor raw text for debug/fallback only...}",
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": []
|
||||
},
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
},
|
||||
"tool_trace_summary": [],
|
||||
"gatekeeper_result": {},
|
||||
"retry_context": null
|
||||
}
|
||||
```
|
||||
|
||||
Verifier 职责:
|
||||
|
||||
- 不调用工具。
|
||||
- 不读 skill。
|
||||
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
|
||||
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
|
||||
- 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。
|
||||
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
|
||||
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
|
||||
|
||||
输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Composer
|
||||
|
||||
Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。
|
||||
|
||||
输入概念:
|
||||
|
||||
| 字段 | 定义 |
|
||||
|---|---|
|
||||
| `original_query` | 用户原始问题 |
|
||||
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `allowed_claims` | Verifier 允许表达的 claims |
|
||||
| `allowed_hypotheses` | Verifier 允许表达的假设 |
|
||||
| `missing_info` | 证据缺口 |
|
||||
| `recommended_actions` | 允许表达的建议动作 |
|
||||
| `rationale` | Verifier 判定理由 |
|
||||
|
||||
输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_summary": "一句话概括",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "下一步动作",
|
||||
"reason": "原因"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "最终给用户看的中文答案"
|
||||
}
|
||||
```
|
||||
|
||||
表达边界:
|
||||
|
||||
- Composer 不补事实、不补根因、不调用工具。
|
||||
- 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。
|
||||
- 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
|
||||
- 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
|
||||
|
||||
---
|
||||
|
||||
## 7. Trace Persistence
|
||||
|
||||
`diagnosis_run.self_evaluation.verifier_evaluation` 持久化:
|
||||
|
||||
```json
|
||||
{
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "...",
|
||||
"round": 1,
|
||||
"traceability_version": "v1",
|
||||
"executor_output_parse_status": {},
|
||||
"executor_structured_output": {},
|
||||
"gatekeeper_result": {
|
||||
"rule_set_version": "gatekeeper-rules-v1"
|
||||
},
|
||||
"composer_output": {},
|
||||
"tool_trace_summary": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Trace API 可用于回放:
|
||||
|
||||
- Executor 输出了哪些 claim。
|
||||
- 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。
|
||||
- Gatekeeper 是否通过、是否自动回填。
|
||||
- Verifier 如何判断可推导性。
|
||||
- Composer 最终如何表达给用户。
|
||||
|
||||
---
|
||||
|
||||
## 8. 当前已验证样例
|
||||
|
||||
| 场景 | sessionId | 结果 |
|
||||
|---|---|---|
|
||||
| HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim |
|
||||
| HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 |
|
||||
|
||||
---
|
||||
|
||||
## 9. 仍需记录或后续补强
|
||||
|
||||
当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
|
||||
|
||||
1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
|
||||
2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。
|
||||
3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
|
||||
@@ -0,0 +1,291 @@
|
||||
# 反馈与自评估架构
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
反馈架构包含两条闭环:
|
||||
|
||||
1. 系统自评估:基于当前 run 的工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_run.self_evaluation`。
|
||||
2. 用户反馈:用户标记 `useful` 或 `not_useful`,优先写入 `diagnosis_run.feedback`,其中 `useful` 会沉淀案例。
|
||||
|
||||
当前重要边界:
|
||||
|
||||
- `status` 表示执行状态,不表示答案质量。
|
||||
- `feedback` 表示用户反馈,不覆盖 `status`。
|
||||
- `self_evaluation` 是 JSON 容器,内部按来源分层,不再把所有评分字段平铺在根节点。
|
||||
|
||||
## 2. 总体闭环
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Answer["Chat / AIOps final answer"] --> Run["diagnosis_run.answer"]
|
||||
|
||||
subgraph SelfEval["Self evaluation"]
|
||||
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
||||
Invocation --> EvidenceRefs["evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
Gatekeeper --> Verifier["chat_verifier"]
|
||||
TraceSummary --> Verifier
|
||||
Verifier --> VerifierEval["verifier_evaluation"]
|
||||
Verifier --> Composer["chat_composer"]
|
||||
Composer --> VerifierEval
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["aiops_rule_evaluation"]
|
||||
end
|
||||
|
||||
RuleEval --> Merge["SelfEvaluationMergeService"]
|
||||
VerifierEval --> Merge
|
||||
AiOpsEval --> Merge
|
||||
Merge --> SelfJson["diagnosis_run.self_evaluation"]
|
||||
|
||||
subgraph UserFeedback["User feedback"]
|
||||
UI["Feedback bar"] --> API["POST /api/feedback"]
|
||||
API --> FeedbackService["FeedbackService"]
|
||||
FeedbackService --> FeedbackField["diagnosis_run.feedback"]
|
||||
FeedbackService --> Useful{"feedback == useful?"}
|
||||
Useful -->|yes| CaseService["CaseLibraryService.createFromRun"]
|
||||
CaseService --> Case["case_library"]
|
||||
Useful -->|no| BadCase["Bad case by feedback=not_useful"]
|
||||
end
|
||||
|
||||
Run --> UI
|
||||
```
|
||||
|
||||
## 3. self_evaluation JSON
|
||||
|
||||
`SelfEvaluationMergeService` 统一维护当前运行的 `diagnosis_run.self_evaluation`。历史兼容数据可能仍存在于 `diagnosis_session.self_evaluation`,但新 Chat/AIOps 执行不再写旧表。
|
||||
|
||||
当前结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"rule_evaluation": {
|
||||
"evidence_score": 65,
|
||||
"source": "rule",
|
||||
"factors": []
|
||||
},
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "...",
|
||||
"round": 1,
|
||||
"traceability_version": "v1",
|
||||
"executor_output_parse_status": {},
|
||||
"executor_structured_output": {},
|
||||
"gatekeeper_result": {},
|
||||
"composer_output": {},
|
||||
"tool_trace_summary": []
|
||||
},
|
||||
"aiops_rule_evaluation": {
|
||||
"verdict": "...",
|
||||
"checks": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
兼容逻辑:
|
||||
|
||||
- 如果旧 JSON 根节点包含 `evidence_score`,会被包进 `rule_evaluation`。
|
||||
- 如果旧 JSON 根节点包含 `verdict` / `groundedness_score`,会被包进 `verifier_evaluation`。
|
||||
|
||||
## 4. 规则评分
|
||||
|
||||
`EvaluationService` 只消费 `tool_invocation` 和 session 状态,输出 `rule_evaluation`。
|
||||
|
||||
定位:
|
||||
|
||||
- 衡量证据收集充分度。
|
||||
- 不直接证明答案是否推理正确。
|
||||
- 不依赖 LLM。
|
||||
|
||||
规则:
|
||||
|
||||
| 规则名 | 条件 | 分数变化 |
|
||||
|---|---|---|
|
||||
| `execution_failed` | session status = `FAILED` | 直接 0 |
|
||||
| `no_tool_call` | 没有工具调用 | 直接 0 |
|
||||
| `has_successful_tool_call` | 至少一次工具成功 | +30 |
|
||||
| `l0_exact_match` | 任意工具调用有 L0 命中 | +35 |
|
||||
| `l1_semantic_match` | 无 L0 命中但有 L1 命中 | +20 |
|
||||
| `retrieval_no_hit` | 有检索调用但无命中 | -10 |
|
||||
| `all_tool_calls_failed` | 工具全部失败 | -20 |
|
||||
|
||||
最终分数裁剪到 `[0, 100]`。
|
||||
|
||||
说明:
|
||||
|
||||
- 当前 `rule_evaluation` 是异步写入,失败时 `self_evaluation` 可能暂时为空或缺少该节点。
|
||||
- L0/L1 分支互斥:有 L0 命中时优先记 L0。
|
||||
- 更强的答案真实性校验由 Chat Verifier 承担。
|
||||
|
||||
## 5. Chat Verifier 自评估
|
||||
|
||||
Chat 自评估分三步:
|
||||
|
||||
1. Gatekeeper 用代码校验 Executor 的引用是否真实。
|
||||
2. Verifier 判断已验真的 `evidence_excerpt` 是否能推出 `claim_text`。
|
||||
3. Composer 只把 Verifier 允许表达的内容写成最终用户答复。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
Summary --> Evidence["tool_trace_summary"]
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
ExecutorOutput --> Verifier
|
||||
Evidence --> Verifier
|
||||
Verifier --> Output["verifier_output JSON"]
|
||||
Output --> Composer["chat_composer"]
|
||||
Composer --> ComposerOutput["composer_output"]
|
||||
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
|
||||
ComposerOutput --> Merge
|
||||
Merge --> Run["diagnosis_run.self_evaluation.verifier_evaluation"]
|
||||
```
|
||||
|
||||
Verifier 输出:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `groundedness_score` | 关键事实证据支撑度 |
|
||||
| `critical_fact_count` | 关键事实数量 |
|
||||
| `claim_checks` | 对 Executor 结构化 claims 的逐条可推导性判断 |
|
||||
| `facts_checked` | 逐条事实校验 |
|
||||
| `rationale` | 判定原因 |
|
||||
| `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 |
|
||||
| `gatekeeper_result` | 引用真实性校验结果 |
|
||||
| `composer_output` | 最终表达的解析状态和摘要 |
|
||||
| `tool_trace_summary` | 本次校验使用的工具调用导航索引 |
|
||||
|
||||
ChatService 根据 verdict 决定:
|
||||
|
||||
- `PASS`:把允许表达的 claims 交给 Composer 输出。
|
||||
- `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。
|
||||
- `REJECT`:降级输出,只保留已确认信息。
|
||||
|
||||
边界:
|
||||
|
||||
- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。
|
||||
- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。
|
||||
|
||||
## 6. AIOps 规则自评估
|
||||
|
||||
AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。
|
||||
|
||||
检查重点:
|
||||
|
||||
- 是否有最终报告。
|
||||
- payload 模式是否聚焦输入告警。
|
||||
- 是否调用 `lookup_knowledge`、日志、指标等证据工具。
|
||||
- 是否把无关活跃告警扩展成主诊断对象。
|
||||
|
||||
这是轻量规则检查,不等价于完整 LLM Verifier。完整 AIOps Verifier 是后续增强项。
|
||||
|
||||
## 7. 用户反馈 API
|
||||
|
||||
```text
|
||||
POST /api/feedback
|
||||
Content-Type: application/json
|
||||
|
||||
{
|
||||
"runId": "run-xxx",
|
||||
"sessionId": "xxx",
|
||||
"feedback": "useful" | "not_useful"
|
||||
}
|
||||
```
|
||||
|
||||
响应:
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"message": "反馈已记录",
|
||||
"runId": "run-xxx",
|
||||
"fallbackToLatestRun": false,
|
||||
"caseId": "uuid 或 null"
|
||||
}
|
||||
```
|
||||
|
||||
后端行为:
|
||||
|
||||
| feedback | 行为 |
|
||||
|---|---|
|
||||
| `useful` | 写入 `DiagnosisRun.feedback`,调用 `CaseLibraryService.createFromRun` |
|
||||
| `not_useful` | 写入 `DiagnosisRun.feedback`,不改变 run status |
|
||||
| 其他值 | 返回 HTTP 400 |
|
||||
|
||||
兼容行为:
|
||||
|
||||
- 请求带 `runId` 时,后端验证 `runId` 属于 `sessionId`。
|
||||
- 请求缺少 `runId` 且存在 run-backed 数据时,后端绑定 latest run,并返回 `fallbackToLatestRun=true` 和实际 `runId`。
|
||||
- 仅当没有 `diagnosis_run` 但存在历史 `diagnosis_session` 时,才使用历史 fallback;该路径不声明 latest-run fallback。
|
||||
|
||||
## 8. 案例沉淀
|
||||
|
||||
`useful` 反馈会生成或复用 `case_library` 记录。
|
||||
|
||||
字段映射:
|
||||
|
||||
| CaseLibrary 字段 | 来源 |
|
||||
|---|---|
|
||||
| `caseId` | UUID |
|
||||
| `diagnosisId` | 新数据为 `DiagnosisRun.runId`;历史数据可能为 `DiagnosisSession.sessionId` |
|
||||
| `sourceType` | `AUTO` |
|
||||
| `faultCategory` | 当前固定为 `GENERAL` |
|
||||
| `title` | `query` 前 100 字符 |
|
||||
| `rootCause` | `answer` |
|
||||
| `solution` | `answer` |
|
||||
| `createdBy` | `system` |
|
||||
|
||||
幂等性:
|
||||
|
||||
```text
|
||||
case_library.diagnosisId == runId
|
||||
-> existing case: return existing
|
||||
-> missing case: create new
|
||||
```
|
||||
|
||||
## 9. Trace 呈现
|
||||
|
||||
Trace API 会展示:
|
||||
|
||||
- `feedback`
|
||||
- `hasFeedback`
|
||||
- `hasVerifierEvaluation`
|
||||
- `hasAiOpsRuleEvaluation`
|
||||
- session、step、tool invocation 明细
|
||||
|
||||
这让一次诊断可以被分成三种视角查看:
|
||||
|
||||
| 视角 | 数据来源 |
|
||||
|---|---|
|
||||
| 执行是否成功 | `diagnosis_run.status` |
|
||||
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
|
||||
| 用户是否认可 | `diagnosis_run.feedback` |
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
近期优先:
|
||||
|
||||
1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。
|
||||
2. 将 ISS-008 / ISS-009 这类 E2E 通过样例固化进 diagnosis eval fixtures。
|
||||
3. `not_useful` 反馈沉淀 bad case,而不是只写字段。
|
||||
4. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
|
||||
5. AIOps 引入 LLM Verifier。
|
||||
6. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
|
||||
|
||||
暂不优先:
|
||||
|
||||
- 用用户反馈直接修改 session status。
|
||||
- 仅凭 `evidence_score` 判断答案正确。
|
||||
- 在没有人工审核时自动把 bad case 反向写入 Prompt。
|
||||
@@ -0,0 +1,258 @@
|
||||
# Harness 与质量门禁架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构 + 后续门禁规划
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计目标
|
||||
|
||||
Agent 系统的核心风险不是“没有答案”,而是:
|
||||
|
||||
- 答案引用了不存在的证据。
|
||||
- 工具调用失败后仍然编造结论。
|
||||
- 检索结果相关性不足但被当作强证据。
|
||||
- 多轮诊断重复检索同一文档,浪费上下文。
|
||||
- 最终报告无法回放执行过程。
|
||||
|
||||
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
|
||||
|
||||
```text
|
||||
Prompt contract
|
||||
+ Tool boundary
|
||||
+ Agent hooks
|
||||
+ Trace persistence
|
||||
+ Gatekeeper deterministic validation
|
||||
+ Verifier / rule evaluation
|
||||
+ Eval baseline
|
||||
```
|
||||
|
||||
## 2. Harness 总图
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||||
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
|
||||
Agent --> Tools["Evidence tools"]
|
||||
Tools --> Invocation["tool_invocation"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Run["diagnosis_run"]
|
||||
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Agent --> Gatekeeper
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
Gatekeeper --> Verifier["chat_verifier"]
|
||||
TraceSummary --> Verifier
|
||||
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||||
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||
|
||||
Run --> TraceAPI["DiagnosisTraceService"]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
AiOpsEval --> TraceAPI
|
||||
|
||||
TraceAPI --> Eval["diagnosis eval / RAG eval"]
|
||||
```
|
||||
|
||||
## 3. Prompt Contract
|
||||
|
||||
当前 Prompt 按角色拆分:
|
||||
|
||||
| Prompt | 用途 |
|
||||
|---|---|
|
||||
| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
|
||||
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
|
||||
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
|
||||
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
|
||||
| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
|
||||
| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
|
||||
| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
|
||||
|
||||
Prompt 层当前承担的门禁:
|
||||
|
||||
- 禁止凭记忆回答错误码、接口定义、排障步骤。
|
||||
- 需要外部信息时必须调用工具。
|
||||
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
|
||||
- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
|
||||
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
|
||||
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
|
||||
- AIOps payload 模式必须聚焦输入告警。
|
||||
|
||||
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
|
||||
|
||||
## 4. Trace Hooks
|
||||
|
||||
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant A as Agent
|
||||
participant H as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
A->>H: before_model(messages, sessionId, runId)
|
||||
H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
|
||||
A-->>A: LLM 推理
|
||||
A->>H: after_model(messages, sessionId, runId)
|
||||
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||||
```
|
||||
|
||||
记录内容:
|
||||
|
||||
- 最近输入消息摘要。
|
||||
- Agent 输出摘要。
|
||||
- 是否包含 tool call。
|
||||
- duration。
|
||||
- token count。
|
||||
- Verifier 的 JSON 输出摘要。
|
||||
|
||||
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
|
||||
|
||||
## 5. Tool Invocation 门禁
|
||||
|
||||
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||||
|
||||
核心记录:
|
||||
|
||||
```text
|
||||
tool_name
|
||||
input_params
|
||||
output_preview
|
||||
retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
-> evidence_refs
|
||||
relevance_level
|
||||
dedup_reason
|
||||
duration_ms
|
||||
success
|
||||
error_message
|
||||
```
|
||||
|
||||
对 `lookup_knowledge` 的质量约束:
|
||||
|
||||
- L0 只作为 hint,不绕过 L1。
|
||||
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
|
||||
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
|
||||
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
|
||||
- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
|
||||
- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
|
||||
|
||||
## 6. Gatekeeper 与 Verifier 门禁
|
||||
|
||||
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> EvidenceIndex["tool_trace_summary"]
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
ExecutorOutput --> Verifier
|
||||
EvidenceIndex --> Verifier
|
||||
Verifier --> Verdict{"verdict"}
|
||||
Verdict -->|PASS| Composer["chat_composer"]
|
||||
Composer --> Pass["输出最终答复"]
|
||||
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
|
||||
Verdict -->|REJECT| Reject["降级输出"]
|
||||
```
|
||||
|
||||
Gatekeeper 检查:
|
||||
|
||||
| 检查 | 失败语义 |
|
||||
|---|---|
|
||||
| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
|
||||
| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
|
||||
| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
|
||||
| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
|
||||
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
|
||||
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
|
||||
|
||||
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
|
||||
|
||||
Verifier 输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS|LOW_CONFID|REJECT",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
||||
|
||||
## 7. AIOps 规则门禁
|
||||
|
||||
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
|
||||
|
||||
检查重点:
|
||||
|
||||
- 最终报告是否存在。
|
||||
- payload 模式是否围绕输入告警展开。
|
||||
- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
|
||||
- 是否把无关活跃告警扩展成主诊断对象。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
## 8. Eval Baseline
|
||||
|
||||
当前质量门禁还包括离线评测资产:
|
||||
|
||||
| 评测 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
|
||||
| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
|
||||
| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
|
||||
|
||||
## 9. 后续门禁规划
|
||||
|
||||
从旧版设计继承但尚未完整实现的门禁:
|
||||
|
||||
- 工具参数 schema 校验。
|
||||
- 同一工具调用次数上限。
|
||||
- 工具超时的统一熔断。
|
||||
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
||||
- Prompt 版本回滚和更细粒度变更审计。
|
||||
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||||
|
||||
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|
||||
@@ -0,0 +1,113 @@
|
||||
# 面试一页式架构讲解
|
||||
|
||||
**用途**:面试现场 2-5 分钟讲清项目
|
||||
**适合场景**:开场介绍、架构追问、Demo 前铺垫
|
||||
|
||||
## 1. 一句话
|
||||
|
||||
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行;所有工具证据、模型步骤、最终答案、自评估和用户反馈都能按 `sessionId + runId` 回放。
|
||||
|
||||
## 2. 一张图
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
User["用户问题 / AIOps 告警"] --> API["API Layer"]
|
||||
|
||||
API --> Chat["ChatService"]
|
||||
API --> AiOps["AiOpsService"]
|
||||
|
||||
Chat --> ChatFlow["Chat: Planner -> Executor -> Gatekeeper -> Verifier -> Composer"]
|
||||
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
|
||||
|
||||
ChatFlow --> Tools["Evidence Tools"]
|
||||
AiOpsFlow --> Tools
|
||||
|
||||
Tools --> Knowledge["lookup_knowledge"]
|
||||
Tools --> Logs["query_logs"]
|
||||
Tools --> Metrics["query_metrics / Prometheus"]
|
||||
|
||||
Knowledge --> RAG["RAG: L0 hint + VectorSearchService"]
|
||||
RAG --> VectorStore["Spring AI VectorStore"]
|
||||
RAG --> SDK["Milvus SDK fallback"]
|
||||
|
||||
ChatFlow --> Trace["Trace Persistence"]
|
||||
AiOpsFlow --> Trace
|
||||
Tools --> Trace
|
||||
|
||||
Trace --> ChatSession["chat_session"]
|
||||
Trace --> Run["diagnosis_run"]
|
||||
Trace --> Step["agent_step"]
|
||||
Trace --> Invocation["tool_invocation"]
|
||||
|
||||
Invocation --> Verifier["Verifier / Rule Evaluation"]
|
||||
Verifier --> SelfEval["self_evaluation"]
|
||||
|
||||
Run --> TraceAPI["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
|
||||
TraceAPI --> Feedback["POST /api/feedback"]
|
||||
Feedback --> Case["useful -> case_library"]
|
||||
```
|
||||
|
||||
## 3. 面试讲法
|
||||
|
||||
```text
|
||||
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
|
||||
|
||||
Chat 复杂问题走 Planner -> Executor -> Gatekeeper -> Verifier -> Composer:
|
||||
Planner 负责拆解,Executor 只负责调用知识库、日志和指标工具并提炼带证据引用的微观事实;Gatekeeper 用代码核对 invocation、raw_path 和 excerpt 是否真实;Verifier 判断这些事实能否由已验真的证据推出;Composer 只把允许表达的结论写成最终答案。
|
||||
|
||||
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
|
||||
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
|
||||
|
||||
会话元数据会落到 chat_session,每次诊断运行会落到 diagnosis_run,步骤和工具明细通过 run_id 关联。
|
||||
所以我可以用 sessionId + runId 精确回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。
|
||||
```
|
||||
|
||||
## 4. 五个亮点
|
||||
|
||||
| 亮点 | 怎么讲 |
|
||||
|---|---|
|
||||
| 可追踪 Agent | 每次诊断都有 `runId`,Trace API 可以回放 run、step、tool;同一 `sessionId` 可有多次独立 run |
|
||||
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
|
||||
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
|
||||
| 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
|
||||
| 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 |
|
||||
|
||||
## 5. 三个关键取舍
|
||||
|
||||
### 取舍 1:为什么不用隐式 Advisor 做 RAG?
|
||||
|
||||
因为这个项目强调 Agent 决策可见性。`lookup_knowledge` 必须作为显式工具调用被记录,这样才能解释“什么时候检索、检索了什么、证据如何支撑结论”。
|
||||
|
||||
### 取舍 2:为什么保留 Milvus SDK fallback?
|
||||
|
||||
因为迁移到 Spring AI VectorStore 期间,schema、collection、score 语义都可能变化。`auto` 模式先走 VectorStore,失败时 fallback 到 SDK,保证 MVP 主链路可运行,也方便对比新旧检索质量。
|
||||
|
||||
### 取舍 3:为什么 self_evaluation 分三层?
|
||||
|
||||
因为三类评估回答的问题不同:
|
||||
|
||||
```text
|
||||
rule_evaluation -> 工具证据是否充分
|
||||
verifier_evaluation -> Chat 答案关键事实是否有证据支撑
|
||||
aiops_rule_evaluation -> AIOps 报告是否聚焦告警并使用证据
|
||||
```
|
||||
|
||||
## 6. 面试官可能追问
|
||||
|
||||
| 追问 | 回答方向 |
|
||||
|---|---|
|
||||
| 怎么防止幻觉? | Executor 输出 `executor_evidence_v2`,每个 claim 绑定 `source_invocation_id + raw_path + evidence_excerpt`;Gatekeeper 用 `tool_invocation.retrieval_details.evidence_refs` 核验引用真实性;Verifier 只判断可推导性;Composer 防止把 no-evidence 说成已排除 |
|
||||
| RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 |
|
||||
| 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter |
|
||||
| AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 |
|
||||
| 下一步怎么演进? | 固化 E2E fixture、Prompt version、Gatekeeper 规则配置化、邻居 chunk、AIOps LLM Verifier、MCP 工具协议化 |
|
||||
|
||||
## 7. 现场演示入口
|
||||
|
||||
- Demo 脚本:`mvp/demo/ten-minute-interview-demo.md`
|
||||
- 故事案例:`interview/story-cases.md`
|
||||
- 架构细节:`mvp/architecture/README.md`
|
||||
@@ -0,0 +1,240 @@
|
||||
# 知识库文档编写与维护
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前建议规范
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-usage.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
知识库文档不是普通 Markdown 资料堆叠,而是 RAG 检索的输入资产。写得好的文档会提升:
|
||||
|
||||
- L0 hint 的关键词和领域识别。
|
||||
- L1 向量召回质量。
|
||||
- `breadcrumb` 上下文恢复能力。
|
||||
- Verifier 可引用的证据质量。
|
||||
|
||||
当前推荐写法:结构化 Markdown + frontmatter + 明确分类 + 可检索关键词。
|
||||
|
||||
## 2. 文档进入系统的链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Markdown["Markdown file"] --> Upload["POST /api/documents/upload"]
|
||||
Upload --> Parse["FrontmatterParser"]
|
||||
Parse --> Enrich["DocumentFieldEnricher"]
|
||||
Enrich --> Metadata["api_document.metadata"]
|
||||
Upload --> Chunk["DocumentChunkService"]
|
||||
Chunk --> Breadcrumb["title / breadcrumb / chunkIndex"]
|
||||
Breadcrumb --> Embedding["VectorIndexService embedding text"]
|
||||
Embedding --> Milvus["Milvus/Zilliz"]
|
||||
Metadata --> L0["KnowledgeIndexService L0 index"]
|
||||
Milvus --> L1["VectorSearchService L1 retrieval"]
|
||||
```
|
||||
|
||||
## 3. Frontmatter
|
||||
|
||||
推荐模板:
|
||||
|
||||
```markdown
|
||||
---
|
||||
title: 支付网关错误码定义
|
||||
keywords: [ERR_TIMEOUT, 支付超时, payment timeout, 支付网关]
|
||||
summary: 记录支付网关核心错误码的含义、常见原因和排查步骤
|
||||
category: api
|
||||
version: 1.0
|
||||
author: sre-team
|
||||
---
|
||||
|
||||
# 支付网关错误码定义
|
||||
|
||||
...
|
||||
```
|
||||
|
||||
字段说明:
|
||||
|
||||
| 字段 | 必填 | 用途 |
|
||||
|---|---:|---|
|
||||
| `title` | 是 | 文档标题,进入 L0 索引和 embedding 上下文 |
|
||||
| `keywords` | 是 | L0 hint 的主要来源 |
|
||||
| `summary` | 是 | 文档摘要,进入知识域描述和 Agent 上下文 |
|
||||
| `category` | 建议 | 知识域、metadata filter、上传目录 |
|
||||
| `version` | 可选 | 文档版本 |
|
||||
| `author` | 可选 | 维护人 |
|
||||
|
||||
当前解析器会提示缺少 `title`、`keywords`、`summary` 的情况;缺失不一定阻断上传,但会降低检索质量。
|
||||
|
||||
## 4. category 建议
|
||||
|
||||
`category` 会影响:
|
||||
|
||||
- 上传文件本地目录。
|
||||
- Milvus metadata。
|
||||
- L0 domain hint。
|
||||
- `knowledge_domain` 聚合。
|
||||
- VectorStore / SDK category filter。
|
||||
|
||||
推荐保持稳定,不要频繁换名。
|
||||
|
||||
| category | 用途 |
|
||||
|---|---|
|
||||
| `api` | 接口、错误码、请求/响应协议 |
|
||||
| `infrastructure` | MySQL、Redis、JVM、网络、中间件 |
|
||||
| `troubleshooting` | 通用排障流程、Runbook |
|
||||
| `domain` | 业务领域规则 |
|
||||
| `spring-ai` | Spring AI / Agent / 工具最佳实践 |
|
||||
|
||||
注意:分类过细会导致 filter 召回不足;分类过粗会降低 L0 hint 解释力。
|
||||
|
||||
## 5. 关键词写法
|
||||
|
||||
好的关键词应该覆盖:
|
||||
|
||||
- 精确实体:错误码、服务名、指标名。
|
||||
- 常用中文说法。
|
||||
- 英文别名。
|
||||
- 组合词。
|
||||
|
||||
示例:
|
||||
|
||||
```yaml
|
||||
keywords: [ERR_TIMEOUT, timeout, 支付超时, 支付网关超时, payment-service, gateway timeout]
|
||||
```
|
||||
|
||||
避免:
|
||||
|
||||
```yaml
|
||||
keywords: [错误, 问题, 系统]
|
||||
```
|
||||
|
||||
原因:过宽关键词会让 L0 hint 变脏,多个文档同时命中,影响解释性和 category filter。
|
||||
|
||||
## 6. Markdown 结构
|
||||
|
||||
推荐结构:
|
||||
|
||||
```markdown
|
||||
# 文档总标题
|
||||
|
||||
## 场景或错误码
|
||||
|
||||
### 含义
|
||||
|
||||
### 常见原因
|
||||
|
||||
### 排查步骤
|
||||
|
||||
### 处理方案
|
||||
|
||||
### 日志示例
|
||||
```
|
||||
|
||||
为什么这样写:
|
||||
|
||||
- `DocumentChunkService` 会按 Markdown 标题切分。
|
||||
- 标题层级会生成 `breadcrumb`。
|
||||
- `title + breadcrumb + content` 会一起进入 embedding 文本。
|
||||
- 命中 chunk 时,Agent 更容易知道证据属于哪个章节。
|
||||
|
||||
## 7. 内容建议
|
||||
|
||||
每个可诊断条目尽量包含:
|
||||
|
||||
- 现象。
|
||||
- 判断条件。
|
||||
- 可能原因。
|
||||
- 证据来源。
|
||||
- 排查步骤。
|
||||
- 处理建议。
|
||||
- 日志或配置示例。
|
||||
|
||||
示例:
|
||||
|
||||
```markdown
|
||||
## ERR_TIMEOUT
|
||||
|
||||
### 含义
|
||||
|
||||
支付网关请求超过本地或上游超时时间。
|
||||
|
||||
### 常见原因
|
||||
|
||||
1. 第三方支付服务响应慢。
|
||||
2. 本地 timeout 配置过短。
|
||||
3. 网络链路抖动。
|
||||
|
||||
### 排查步骤
|
||||
|
||||
1. 查询 payment-service 日志中的请求耗时。
|
||||
2. 查看网关 5xx 和 timeout 指标。
|
||||
3. 对比当前 timeout 配置。
|
||||
|
||||
### 处理建议
|
||||
|
||||
- 短期:重试受影响订单。
|
||||
- 长期:调整 timeout 和重试策略,并监控上游延迟。
|
||||
```
|
||||
|
||||
## 8. 上传与索引
|
||||
|
||||
上传接口:
|
||||
|
||||
```text
|
||||
POST /api/documents/upload
|
||||
Content-Type: multipart/form-data
|
||||
|
||||
file=<markdown file>
|
||||
category=<category>
|
||||
```
|
||||
|
||||
系统处理:
|
||||
|
||||
1. 计算文件 hash,避免重复上传。
|
||||
2. 保存原始文件。
|
||||
3. 解析 frontmatter。
|
||||
4. 补全文档字段。
|
||||
5. 写入 `api_document`。
|
||||
6. Markdown-aware chunking。
|
||||
7. 写入 Milvus/Zilliz。
|
||||
8. 更新 L0 索引和 `knowledge_domain`。
|
||||
|
||||
## 9. 重建索引注意事项
|
||||
|
||||
当以下内容变化时,需要重新索引:
|
||||
|
||||
- 正文内容。
|
||||
- 标题层级。
|
||||
- `category`。
|
||||
- `title`、`summary`、`keywords`。
|
||||
- embedding 输入策略,例如加入 `breadcrumb`。
|
||||
|
||||
特别注意:
|
||||
|
||||
```text
|
||||
修改 Markdown 或 embedding 输入策略,不会自动改变已有向量。
|
||||
必须重新上传或重建索引后,live retrieval 才能体现变化。
|
||||
```
|
||||
|
||||
可用 live 验收:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py
|
||||
```
|
||||
|
||||
## 10. 维护 checklist
|
||||
|
||||
新增文档前检查:
|
||||
|
||||
- frontmatter 是否包含 `title`、`keywords`、`summary`。
|
||||
- `category` 是否属于现有稳定分类。
|
||||
- 关键词是否既有精确词也有常用表达。
|
||||
- Markdown 标题层级是否清晰。
|
||||
- 每个故障条目是否包含可执行排查步骤。
|
||||
- 日志/配置示例是否脱敏。
|
||||
|
||||
更新文档后检查:
|
||||
|
||||
- `api_document.status` 是否为 `INDEXED`。
|
||||
- `/api/search/similar` 是否能搜到目标文档。
|
||||
- `eval/rag-retrieval` 是否需要新增 golden case。
|
||||
- Trace 中 `tool_invocation` 是否记录到正确 source 和 breadcrumb。
|
||||
|
||||
@@ -0,0 +1,201 @@
|
||||
# 模块化 RAG Pipeline 架构
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前已实现架构
|
||||
**关联 OpenSpec**:`openspec/changes/archive/2026-07-06-modular-rag-pipeline`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
本文记录 `lookup_knowledge` 的当前模块化 RAG 实现。它是 [rag-architecture.md](rag-architecture.md) 的落地版,重点说明代码模块、数据契约、降级策略和可观测性边界。
|
||||
|
||||
核心目标:
|
||||
|
||||
- 保留显式 Agent Tool:`lookup_knowledge(query)`。
|
||||
- L0 只作为 query understanding / filter / rerank / trace hint。
|
||||
- L1 向量检索作为事实证据来源。
|
||||
- filtered L1 低质量时,降级为 raw query unfiltered L1 retry。
|
||||
- 输出 evidence-first contract,替代旧 `primary/supplement`。
|
||||
|
||||
## 2. 当前链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Agent["Executor Agent"] --> Tool["LookupKnowledgeTool.lookupKnowledge(query)"]
|
||||
|
||||
Tool --> Transform["KnowledgeQueryTransformer"]
|
||||
Transform --> KQ["KnowledgeQuery"]
|
||||
|
||||
KQ --> Retriever["KnowledgeDocumentRetriever"]
|
||||
Retriever --> Attempt1["FILTERED_VECTOR or UNFILTERED_VECTOR"]
|
||||
Attempt1 --> Post1["KnowledgeEvidencePostProcessor"]
|
||||
Post1 --> Quality{"usable evidence?"}
|
||||
Quality -->|yes| Pack
|
||||
Quality -->|no and categoryFilter exists| Retry["UNFILTERED_VECTOR_RETRY"]
|
||||
Retry --> Post2["KnowledgeEvidencePostProcessor"]
|
||||
Post2 --> Pack["KnowledgeContextPacker"]
|
||||
|
||||
Pack --> Assemble["LookupResultAssembler"]
|
||||
Assemble --> Result["LookupResult"]
|
||||
Result --> Dedup["RetrievedDocTracker session dedup"]
|
||||
Dedup --> Recorder["ToolInvocationRecorder"]
|
||||
Recorder --> Trace["tool_invocation.retrieval_details"]
|
||||
Result --> Agent
|
||||
```
|
||||
|
||||
对应代码:
|
||||
|
||||
| 阶段 | 类 | 职责 |
|
||||
|---|---|---|
|
||||
| Tool Boundary | `LookupKnowledgeTool` | 接收 Agent 工具调用,编排 pipeline,处理 session dedup 和 recorder |
|
||||
| Query Transformation | `KnowledgeQueryTransformer` | 复用 L0,输出 query hints 和可选 category filter |
|
||||
| Retrieval | `KnowledgeDocumentRetriever` | 调用 `VectorSearchService`,统一 filtered / unfiltered attempt |
|
||||
| Post-Retrieval | `KnowledgeEvidencePostProcessor` | L2 归一化、证据块构建、source dedup、规则 rerank |
|
||||
| Context Packing | `KnowledgeContextPacker` | 按字符预算打包 Agent 可消费 context |
|
||||
| Result Assembly | `LookupResultAssembler` | 统一 evidence result、no-evidence result、dedup result |
|
||||
| Observability | `ToolInvocationRecorder` | 写入 query transform、retrieval trace、rerank trace、context pack summary |
|
||||
|
||||
## 3. L0 与 L1 边界
|
||||
|
||||
L0 来源于 `KnowledgeIndexService.analyzeQuery`,输出进入 `KnowledgeQuery`:
|
||||
|
||||
```text
|
||||
originalQuery
|
||||
rewrittenQuery
|
||||
domainHints
|
||||
matchedKeywords
|
||||
entities
|
||||
categoryFilter
|
||||
l0Titles
|
||||
l0MatchCount
|
||||
```
|
||||
|
||||
L0 可以做:
|
||||
|
||||
- 给 L1 提供单一 category filter。
|
||||
- 给 rerank 提供 domain / keyword / entity boost 信号。
|
||||
- 给 trace 提供解释信息。
|
||||
|
||||
L0 不再做:
|
||||
|
||||
- 不因唯一命中直接返回文档正文。
|
||||
- 不在 L1 无结果时作为事实证据兜底。
|
||||
- 不进入 `evidenceBlocks`,除非未来明确引入新的 evidence source 规则。
|
||||
|
||||
L1 通过 `VectorSearchService.searchSimilarDocuments(query, topK, category)` 执行,内部仍保留 Spring AI VectorStore 优先和 Milvus SDK fallback。
|
||||
|
||||
## 4. 降级策略
|
||||
|
||||
MVP 降级策略保持简单:
|
||||
|
||||
```text
|
||||
if categoryFilter exists:
|
||||
run FILTERED_VECTOR
|
||||
if empty / no evidence / top similarity < referenceThreshold:
|
||||
run UNFILTERED_VECTOR_RETRY with original query
|
||||
else:
|
||||
run UNFILTERED_VECTOR
|
||||
```
|
||||
|
||||
fallback reason:
|
||||
|
||||
| reason | 含义 |
|
||||
|---|---|
|
||||
| `filtered_vector_no_evidence` | filtered L1 无候选或 post-processing 后无 evidence |
|
||||
| `filtered_vector_low_quality` | filtered L1 有候选,但 top normalized similarity 低于 `retrieval.normalization.reference-threshold` |
|
||||
|
||||
当 retry 后仍无证据:
|
||||
|
||||
- `found=false`
|
||||
- `evidenceStatus=no_evidence`
|
||||
- 保留 `retrievalTrace`
|
||||
- 不返回 L0 文档作为事实证据
|
||||
|
||||
## 5. Evidence-First Contract
|
||||
|
||||
`LookupResult` 当前核心字段:
|
||||
|
||||
```text
|
||||
found
|
||||
evidenceBlocks
|
||||
contextPack
|
||||
retrievalTrace
|
||||
rerankTrace
|
||||
relevanceLevel
|
||||
completenessHint
|
||||
retrievedDomainsThisSession
|
||||
message
|
||||
```
|
||||
|
||||
旧字段已删除:
|
||||
|
||||
```text
|
||||
primary
|
||||
supplement
|
||||
```
|
||||
|
||||
这是一项 L4 breaking interface change。项目内已同步迁移:
|
||||
|
||||
- `LookupKnowledgeTool`
|
||||
- `ToolInvocationRecorder`
|
||||
- executor prompts
|
||||
- lookup / recorder tests
|
||||
- RAG architecture docs
|
||||
- OpenSpec 主 spec
|
||||
|
||||
## 6. Trace 结构
|
||||
|
||||
`retrieval_details` 保持 JSON 扩展,不改表结构。关键内容:
|
||||
|
||||
```json
|
||||
{
|
||||
"query_transform": {},
|
||||
"retrieval_trace": {
|
||||
"selected_attempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallback_reason": "filtered_vector_no_evidence",
|
||||
"attempts": []
|
||||
},
|
||||
"context_pack_summary": {},
|
||||
"rerank_trace": {},
|
||||
"evidence_blocks": []
|
||||
}
|
||||
```
|
||||
|
||||
这样 Trace API、Verifier、Eval 可以继续从 `tool_invocation` 读取证据链。
|
||||
|
||||
## 7. Review 修正
|
||||
|
||||
归档前 review 发现:session dedup 命中时返回 `found=false`,但仍携带 `evidenceBlocks/contextPack`,可能导致 Agent 重复消费证据。
|
||||
|
||||
当前行为已修正:
|
||||
|
||||
- dedup result 不再返回可消费 evidence/context。
|
||||
- 保留 message、retrieval trace、relevance hint 和 retrieved domains。
|
||||
- 测试覆盖:`LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`。
|
||||
|
||||
## 8. 验证
|
||||
|
||||
已执行:
|
||||
|
||||
```powershell
|
||||
mvn -q -DskipTests compile
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
|
||||
$env:MILVUS_TOKEN = <application.yml 中的 milvus.token>; mvn -q test
|
||||
openspec validate --all --strict
|
||||
git diff --check
|
||||
```
|
||||
|
||||
结果:全部通过。
|
||||
|
||||
注意:
|
||||
|
||||
- `MilvusConnectionTest` 直接读 `MILVUS_TOKEN` 环境变量,不读 Spring 配置。
|
||||
- 完整测试需要在 Maven 进程里注入该环境变量。
|
||||
|
||||
## 9. 后续演进
|
||||
|
||||
建议后续按评测结果推进,而不是先堆复杂能力:
|
||||
|
||||
- 增加 RAG eval cases:固定 query、期望 source、期望 fallback path。
|
||||
- 引入更严格的 evidence grounding 检查。
|
||||
- 当规则 rerank 不足时,再考虑 model-based rerank。
|
||||
- 当召回覆盖率不足时,再考虑 BM25/RRF/hybrid retrieval。
|
||||
@@ -0,0 +1,415 @@
|
||||
# RAG 新架构
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前主架构 + 后续演进边界
|
||||
**关联计划**:[`mvp/issues/active/rag-refactor-plan.md`](../issues/active/rag-refactor-plan.md)
|
||||
|
||||
## 1. 架构目标
|
||||
|
||||
RAG 重构的目标不是把所有能力交给框架,也不是继续维护一套完全自研检索框架,而是形成:
|
||||
|
||||
```text
|
||||
成熟框架能力 + 业务可观测编排
|
||||
```
|
||||
|
||||
具体原则:
|
||||
|
||||
- 通用向量检索能力交给 Spring AI `VectorStore`。
|
||||
- 项目保留 Agent Tool 入口、AIOps 业务 query 映射、证据打包、trace 记录。
|
||||
- `lookup_knowledge` 继续是显式工具,不替换成隐式 Advisor。
|
||||
- Spring AI 读取路径作为主路径,Milvus SDK 作为 fallback。
|
||||
- 所有检索行为必须可评测、可回放、可解释。
|
||||
|
||||
## 2. 当前主链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Agent["Agent Executor"] --> Tool["lookup_knowledge(query)"]
|
||||
|
||||
Tool --> L0["KnowledgeIndexService.analyzeQuery"]
|
||||
L0 --> Hint["L0 hint: domain / entities / matchedKeywords"]
|
||||
Hint --> Filter["category filter candidate"]
|
||||
|
||||
Tool --> Search["VectorSearchService.searchSimilarDocuments"]
|
||||
Filter --> Search
|
||||
|
||||
Search --> Mode{"retrieval.vector-store.mode"}
|
||||
Mode -->|auto| SpringTry["try Spring AI VectorStore"]
|
||||
SpringTry -->|success| Results["SearchResult list"]
|
||||
SpringTry -->|failure| SdkFallback["Milvus SDK fallback"]
|
||||
Mode -->|spring / spring-ai| SpringOnly["Spring AI VectorStore only"]
|
||||
Mode -->|sdk| SdkOnly["Milvus SDK only"]
|
||||
|
||||
SpringOnly --> Results
|
||||
SdkFallback --> Results
|
||||
SdkOnly --> Results
|
||||
|
||||
Results --> Retry{"filtered result usable?"}
|
||||
Retry -->|no| RetryL1["raw query unfiltered L1 retry"]
|
||||
Retry -->|yes| Post["post-retrieval processing"]
|
||||
RetryL1 --> Post
|
||||
Post --> Pack["context packing"]
|
||||
Pack --> Dedup["session dedup: RetrievedDocTracker"]
|
||||
Dedup --> Output["LookupResult: evidenceBlocks / contextPack / traces"]
|
||||
Output --> Record["tool_invocation record"]
|
||||
Output --> Agent
|
||||
```
|
||||
|
||||
```text
|
||||
Agent Executor
|
||||
-> lookup_knowledge(query)
|
||||
-> KnowledgeIndexService.analyzeQuery
|
||||
-> L0 domain/entity hint
|
||||
-> matchedKeywords
|
||||
-> category filter candidate
|
||||
-> VectorSearchService.searchSimilarDocuments
|
||||
-> mode=auto
|
||||
-> Spring AI VectorStore
|
||||
-> fallback: Milvus SDK
|
||||
-> mode=spring / spring-ai
|
||||
-> Spring AI VectorStore only
|
||||
-> mode=sdk
|
||||
-> Milvus SDK only
|
||||
-> post-retrieval processing
|
||||
-> relevanceLevel
|
||||
-> completenessHint
|
||||
-> evidenceBlocks
|
||||
-> rerankTrace
|
||||
-> context packing
|
||||
-> contextPack
|
||||
-> session dedup
|
||||
-> RetrievedDocTracker
|
||||
-> tool_invocation record
|
||||
```
|
||||
|
||||
运行配置:
|
||||
|
||||
```properties
|
||||
retrieval.vector-store.mode=auto
|
||||
retrieval.normalization.max-l2-distance=2.0
|
||||
retrieval.normalization.highly-relevant-threshold=0.75
|
||||
retrieval.normalization.reference-threshold=0.5
|
||||
```
|
||||
|
||||
## 3. 稳定边界
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph AgentBoundary["Agent boundary"]
|
||||
Executor["Executor Agent"]
|
||||
Tool["LookupKnowledgeTool"]
|
||||
end
|
||||
|
||||
subgraph RetrievalBoundary["Retrieval boundary"]
|
||||
Search["VectorSearchService"]
|
||||
Spring["Spring AI VectorStore"]
|
||||
SDK["Milvus SDK"]
|
||||
end
|
||||
|
||||
subgraph ObservabilityBoundary["Observability boundary"]
|
||||
Invocation["tool_invocation"]
|
||||
Eval["RAG baseline / trace inspection"]
|
||||
end
|
||||
|
||||
Executor --> Tool
|
||||
Tool --> Search
|
||||
Search --> Spring
|
||||
Search --> SDK
|
||||
Tool --> Invocation
|
||||
Invocation --> Eval
|
||||
```
|
||||
|
||||
### 3.1 Agent 边界
|
||||
|
||||
Agent 只知道自己可以调用 `lookup_knowledge`,不直接关心底层是 Spring AI VectorStore 还是 Milvus SDK。
|
||||
|
||||
```text
|
||||
Executor -> LookupKnowledgeTool -> VectorSearchService
|
||||
```
|
||||
|
||||
这个边界让 RAG 底层迁移不影响 Agent prompt、工具声明和 trace 数据结构。
|
||||
|
||||
### 3.2 检索边界
|
||||
|
||||
`VectorSearchService` 是当前检索门面:
|
||||
|
||||
- `auto`:优先 Spring AI VectorStore,失败后 fallback 到 SDK。
|
||||
- `spring` / `spring-ai`:只走 Spring AI VectorStore。
|
||||
- `sdk`:只走原 Milvus SDK。
|
||||
|
||||
这样可以在不改 Agent 工具的情况下切换检索实现,并支持线上验证和回退。
|
||||
|
||||
### 3.3 可观测边界
|
||||
|
||||
无论底层检索路径如何变化,都必须写入 `tool_invocation`:
|
||||
|
||||
```text
|
||||
sessionId
|
||||
toolName
|
||||
inputParams
|
||||
outputPreview
|
||||
retrievalLayer
|
||||
l0MatchCount
|
||||
l1MatchCount
|
||||
retrievalDetails
|
||||
relevanceLevel
|
||||
dedupReason
|
||||
duration
|
||||
success
|
||||
```
|
||||
|
||||
## 4. L0 的新职责
|
||||
|
||||
旧版 L0 容易承担过重职责,例如唯一匹配后直接跳过 L1。当前架构中 L0 被降级为 hint 层。
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Input["query / AIOps payload"] --> L0["L0 hint analysis"]
|
||||
L0 --> Domain["domain detector"]
|
||||
L0 --> Entity["entity extractor"]
|
||||
L0 --> Keyword["matched keyword explanation"]
|
||||
L0 --> Filter["metadata/category filter candidate"]
|
||||
|
||||
Domain --> Retrieval["L1 semantic retrieval"]
|
||||
Entity --> Retrieval
|
||||
Keyword --> Trace["hit reason in tool_invocation"]
|
||||
Filter --> Retrieval
|
||||
|
||||
Retrieval --> Normalize["relevance normalization"]
|
||||
Normalize --> Evidence["evidence returned to Agent"]
|
||||
```
|
||||
|
||||
L0 负责:
|
||||
|
||||
- domain detector
|
||||
- entity extractor
|
||||
- matched keyword explanation
|
||||
- metadata/category filter candidate
|
||||
- trace 中的 hit reason
|
||||
|
||||
L0 不再负责:
|
||||
|
||||
```text
|
||||
L0 unique hit -> 直接作为最终检索结果
|
||||
L1 no result -> 返回 L0 文档作为事实证据
|
||||
```
|
||||
|
||||
当前职责是:
|
||||
|
||||
```text
|
||||
query / AIOps payload
|
||||
-> L0 matched keywords / domains / entities
|
||||
-> category filter candidate
|
||||
-> filtered L1 semantic retrieval
|
||||
-> low-quality? raw query unfiltered L1 retry
|
||||
-> post-retrieval processing
|
||||
-> context packing
|
||||
```
|
||||
|
||||
这样既保留精确关键词和领域 hint 的价值,也避免 L0 误召回直接污染最终证据。
|
||||
|
||||
## 5. L1 向量检索
|
||||
|
||||
L1 语义检索通过 `VectorSearchService` 调度。
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Search["VectorSearchService"] --> Request["SearchRequest: query / topK / threshold / filter"]
|
||||
Request --> VectorStore["Spring AI VectorStore"]
|
||||
VectorStore --> Docs["Document results"]
|
||||
Docs --> Map["map to SearchResult"]
|
||||
Map --> Score["score compatibility mapping"]
|
||||
|
||||
Search --> SDK["Milvus SDK fallback"]
|
||||
SDK --> SdkRows["id / content / metadata / L2 distance"]
|
||||
SdkRows --> Map
|
||||
|
||||
Score --> Output["id / content / metadata / score / rawScore / scoreLabel"]
|
||||
```
|
||||
|
||||
### Spring AI VectorStore 路径
|
||||
|
||||
```text
|
||||
SearchRequest
|
||||
-> query
|
||||
-> topK
|
||||
-> similarityThresholdAll
|
||||
-> optional filterExpression: category == '...'
|
||||
-> VectorStore.similaritySearch
|
||||
```
|
||||
|
||||
返回结果会映射为项目兼容结构:
|
||||
|
||||
```text
|
||||
id
|
||||
content
|
||||
metadata
|
||||
score
|
||||
rawScore
|
||||
scoreLabel
|
||||
```
|
||||
|
||||
### Milvus SDK fallback
|
||||
|
||||
SDK 路径仍保留:
|
||||
|
||||
- 用于 `auto` 模式兜底。
|
||||
- 用于与旧链路对比。
|
||||
- 用于 VectorStore 配置或 collection schema 异常时保证 MVP 可运行。
|
||||
|
||||
## 6. 分数语义
|
||||
|
||||
旧 SDK 使用 L2 distance,Spring AI 返回 similarity。两者不能混用为同一个含义。
|
||||
|
||||
当前统一输出:
|
||||
|
||||
| 字段 | 含义 |
|
||||
|---|---|
|
||||
| `score` | 兼容旧逻辑的距离型分数,越小越近 |
|
||||
| `rawScore` | 底层实现的原始分数 |
|
||||
| `scoreLabel` | `l2_distance` 或 `similarity` |
|
||||
|
||||
SDK 路径:
|
||||
|
||||
```text
|
||||
score = L2 distance
|
||||
rawScore = L2 distance
|
||||
scoreLabel = l2_distance
|
||||
```
|
||||
|
||||
VectorStore 路径:
|
||||
|
||||
```text
|
||||
rawScore = Spring AI similarity
|
||||
scoreLabel = similarity
|
||||
score = metadata.distance if available else compatible distance
|
||||
```
|
||||
|
||||
## 7. 文档切片和 embedding 输入
|
||||
|
||||
当前保留 Markdown-aware chunking:
|
||||
|
||||
- 识别 Markdown 标题层级。
|
||||
- 生成 `title`。
|
||||
- 生成 `breadcrumb`。
|
||||
- 保留 `chunkIndex`。
|
||||
- 使用 token 估算和软/硬上限控制 chunk 大小。
|
||||
- 尽量不打断列表和代码块。
|
||||
|
||||
embedding 输入中已经加强:
|
||||
|
||||
```text
|
||||
title + breadcrumb + content
|
||||
```
|
||||
|
||||
这样可以降低单个 chunk 脱离章节上下文后的召回损失。
|
||||
|
||||
## 8. AIOps query 增强
|
||||
|
||||
AIOps payload 中的业务字段不能完全交给通用检索框架隐式理解。
|
||||
|
||||
payload 模式会把以下字段拼成推荐知识库 query:
|
||||
|
||||
- `alertName`
|
||||
- `service`
|
||||
- `severity`
|
||||
- `description`
|
||||
- `timeRange`
|
||||
- `userRequest`
|
||||
|
||||
Prompt 会明确要求 Agent 在需要知识库证据时,优先使用推荐 query 或保留 alertName/service 的更窄 query。
|
||||
|
||||
```text
|
||||
AIOps payload
|
||||
-> buildKnowledgeRetrievalQuery
|
||||
-> Recommended lookup_knowledge query
|
||||
-> lookup_knowledge
|
||||
-> tool_invocation
|
||||
```
|
||||
|
||||
## 9. Evidence 与去重
|
||||
|
||||
当前 evidence 输出已从旧 `primary/supplement` 迁移为 evidence-first contract,核心字段包括:
|
||||
|
||||
- `evidenceBlocks`
|
||||
- `contextPack`
|
||||
- `retrievalTrace`
|
||||
- `rerankTrace`
|
||||
- `relevanceLevel`
|
||||
- `completenessHint`
|
||||
- `retrievedDomainsThisSession`
|
||||
- `tool_invocation.retrieval_details`
|
||||
|
||||
evidence block 结构:
|
||||
|
||||
```text
|
||||
source / title / breadcrumb / retrievalLayer / content / score / hitReasons
|
||||
```
|
||||
|
||||
context pack 会按重排后的证据顺序生成 Agent 可消费的紧凑上下文,并保留 included/omitted sources 供 trace 检查。
|
||||
|
||||
## 10. 评测与验收
|
||||
|
||||
RAG 架构变更必须先过评测,再认为可合入主链路。
|
||||
|
||||
当前评测资产:
|
||||
|
||||
- `eval/rag-retrieval/cases/golden-cases.json`
|
||||
- `eval/rag-retrieval/fixtures/`
|
||||
- `eval/rag-retrieval/reports/baseline.json`
|
||||
- `eval/rag-retrieval/reports/baseline.md`
|
||||
- `scripts/eval_rag_retrieval.py`
|
||||
- `scripts/eval_rag_live_acceptance.py`
|
||||
|
||||
评测层次:
|
||||
|
||||
| 层次 | 作用 |
|
||||
|---|---|
|
||||
| Offline baseline | 不依赖 MySQL、Redis、Milvus、LLM,用固定 fixtures 检查召回行为 |
|
||||
| Live acceptance | 应用运行并重建索引后,调用 `/api/search/similar` 验证真实检索 |
|
||||
| Trace inspection | 通过 `tool_invocation` 检查 Agent 是否真的使用了证据 |
|
||||
|
||||
## 11. 当前已完成
|
||||
|
||||
- `lookup_knowledge` 保持显式 Agent Tool。
|
||||
- L0 降级为 domain/entity hint。
|
||||
- L1 默认执行语义检索。
|
||||
- `VectorSearchService` 支持 `auto`、`spring`/`spring-ai`、`sdk` 三种模式。
|
||||
- Spring AI VectorStore 成为读取主路径。
|
||||
- Milvus SDK fallback 保留。
|
||||
- 分数语义拆成 `score`、`rawScore`、`scoreLabel`。
|
||||
- Markdown chunk 保留 `title` 和 `breadcrumb`。
|
||||
- embedding 输入包含 `title`、`breadcrumb` 和 `content`。
|
||||
- AIOps payload 生成推荐知识库 query。
|
||||
- `tool_invocation` 记录 relevance level、dedup reason、evidence summaries、retrieval trace、rerank trace 和 context pack summary。
|
||||
- `lookup_knowledge` 输出使用 evidence-first contract,不再暴露旧 `primary/supplement` 字段。
|
||||
- RAG offline baseline 和 live acceptance 脚本已补齐。
|
||||
|
||||
## 12. 后续演进
|
||||
|
||||
近期优先:
|
||||
|
||||
1. 命中 chunk 的相邻 chunk / 同章节上下文扩展。
|
||||
2. metadata taxonomy 清理,例如 `database` 与 `infrastructure` 的分类边界。
|
||||
3. Query Transformer / MultiQuery 的可回退接入。
|
||||
4. VectorStore 写入路径评估。
|
||||
|
||||
暂不优先:
|
||||
|
||||
- 把 `lookup_knowledge` 替换成隐式 Advisor。
|
||||
- 完整自研 RRF 框架。
|
||||
- 立即引入 Elasticsearch / OpenSearch。
|
||||
- 立即引入 cross-encoder 或 LLM rerank。
|
||||
|
||||
## 13. 关键代码索引
|
||||
|
||||
| 能力 | 代码 |
|
||||
|---|---|
|
||||
| Agent 工具入口 | `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java` |
|
||||
| L0 hint | `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java` |
|
||||
| 向量检索门面 | `src/main/java/com/superbiz/agent/service/VectorSearchService.java` |
|
||||
| 文档切片 | `src/main/java/com/superbiz/agent/service/DocumentChunkService.java` |
|
||||
| 文档管理 | `src/main/java/com/superbiz/agent/service/DocumentManagementService.java` |
|
||||
| 向量写入 | `src/main/java/com/superbiz/agent/service/VectorIndexService.java` |
|
||||
| AIOps query 增强 | `src/main/java/com/superbiz/agent/service/AiOpsService.java` |
|
||||
| 工具调用记录 | `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java` |
|
||||
@@ -0,0 +1,183 @@
|
||||
# RAG 评测闭环架构
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
|
||||
本文记录当前 RAG 质量闭环。它的目标不是证明检索“永远正确”,而是让每次改 `lookup_knowledge`、L0 hint、向量召回、post-retrieval、rerank 或 context packing 时,都能得到可重复的回归信号。
|
||||
|
||||
## 1. 闭环分层
|
||||
|
||||
```text
|
||||
RAG pipeline change
|
||||
-> LookupKnowledgeTool snapshot generation
|
||||
-> offline RAG retrieval baseline
|
||||
-> RAG baseline diff
|
||||
-> diagnosis eval baseline
|
||||
-> diagnosis baseline diff
|
||||
-> accept / fix / archive
|
||||
```
|
||||
|
||||
| 层级 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| RAG retrieval baseline | `eval/rag-retrieval/` | 检查固定 query 是否命中期望证据、路径和 fallback |
|
||||
| RAG baseline diff | `scripts/eval_rag_retrieval.py --compare-to ...` | 对比当前报告和旧基线,输出 regression/change |
|
||||
| Diagnosis eval baseline | `mvp/eval/` | 检查 Agent 最终诊断 trace、报告和证据行为 |
|
||||
| Live acceptance | `scripts/eval_rag_live_acceptance.py` | 在应用和向量库运行后做真实环境 smoke check |
|
||||
|
||||
## 2. Offline RAG Baseline
|
||||
|
||||
核心资产:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/cases/golden-cases.json
|
||||
eval/rag-retrieval/fixtures/*.json
|
||||
eval/rag-retrieval/reports/baseline.json
|
||||
eval/rag-retrieval/reports/baseline.md
|
||||
scripts/eval_rag_retrieval.py
|
||||
scripts/generate_rag_lookup_snapshots.ps1
|
||||
src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java
|
||||
```
|
||||
|
||||
运行:
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
该 baseline 完全离线,不依赖 MySQL、Redis、Milvus、LLM 或 Spring Boot。它适合在改 RAG 代码后快速判断:
|
||||
|
||||
- 期望 source 是否仍在 topK 内。
|
||||
- breadcrumb 和 evidence keyword 是否仍能覆盖。
|
||||
- `LookupResult` 是否仍包含 `evidenceBlocks/contextPack/retrievalTrace/rerankTrace`。
|
||||
- selected attempt 是否符合预期。
|
||||
- fallback reason 是否符合预期。
|
||||
- context pack 是否包含期望 source。
|
||||
- rerank top source 是否稳定。
|
||||
|
||||
## 3. 模块化输出契约
|
||||
|
||||
fixture 必须使用当前模块化格式:
|
||||
|
||||
```json
|
||||
{
|
||||
"lookupResult": {
|
||||
"evidenceBlocks": [],
|
||||
"contextPack": {},
|
||||
"retrievalTrace": {},
|
||||
"rerankTrace": {}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
当前 golden cases 直接以模块化格式为唯一契约,因为这个版本的目标是验证完整 RAG pipeline,而不只是验证候选召回。
|
||||
|
||||
## 4. Fallback Case
|
||||
|
||||
当前 baseline 增加了 `chat-l0-filter-fallback`:
|
||||
|
||||
```text
|
||||
FILTERED_VECTOR low quality or no evidence
|
||||
-> UNFILTERED_VECTOR_RETRY
|
||||
-> fallbackReason = filtered_vector_low_quality | filtered_vector_no_evidence
|
||||
```
|
||||
|
||||
这个 case 固化了 MVP 版本的降级策略:如果经过 L0 filter 后 L1 低质量或没有证据,就跳过 L0 filter,用原始 query 再做一次无过滤向量检索。不同向量后端对“低质量候选”和“无候选”的边界可能不同,所以 golden case 允许两个 fallback reason,但强制要求 retry 行为和最终证据正确。
|
||||
|
||||
## 5. Diff 闭环
|
||||
|
||||
生成当前报告并与旧基线对比:
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py `
|
||||
--json-report eval\rag-retrieval\reports\current.json `
|
||||
--markdown-report eval\rag-retrieval\reports\current.md `
|
||||
--compare-to eval\rag-retrieval\reports\baseline.json `
|
||||
--diff-json-report eval\rag-retrieval\reports\baseline-diff.json `
|
||||
--diff-markdown-report eval\rag-retrieval\reports\baseline-diff.md
|
||||
```
|
||||
|
||||
diff 会检查:
|
||||
|
||||
- pass rate
|
||||
- recall@K
|
||||
- strong hit rate
|
||||
- miss count
|
||||
- case pass state
|
||||
- hit level
|
||||
- first expected rank
|
||||
- selected attempt
|
||||
- fallback reason
|
||||
- evidence status
|
||||
- rerank top source
|
||||
|
||||
当 case 失败或 diff 出现 regression 时,脚本会返回非 0 退出码,可作为本地质量门禁或 CI 门禁。
|
||||
|
||||
## 6. 与 Diagnosis Eval 的关系
|
||||
|
||||
RAG baseline 解决的是“证据有没有被正确检索、处理和打包”。
|
||||
|
||||
Diagnosis eval 解决的是“Agent 有没有把证据用于最终诊断,并保持 trace 可解释”。
|
||||
|
||||
两者不是替代关系:
|
||||
|
||||
- 改 RAG pipeline:先跑 RAG baseline,再跑相关 Agent 测试。
|
||||
- 改 prompt、Agent 编排、Verifier:重点跑 diagnosis eval。
|
||||
- 改 embedding 输入、reindex、向量库配置:跑 RAG baseline + live acceptance。
|
||||
|
||||
## 7. 面试表达
|
||||
|
||||
可以概括为:
|
||||
|
||||
> 我没有只做一个 RAG 调用,而是把 RAG 拆成 Query Transform、Retrieval、Post-Retrieval、Rerank、Context Packing,并为它建设了离线 golden cases、baseline report、baseline diff 和上层 diagnosis eval,形成可回放、可对比、可回归的 Agent 质量闭环。
|
||||
|
||||
## 8. LookupKnowledgeTool Snapshot
|
||||
|
||||
真实工具快照生成命令:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
```
|
||||
|
||||
该命令默认使用 `retrieval.vector-store.mode=spring`,通过 `RagLookupSnapshotGeneratorTest` 启动 Spring test context,注入真实 `LookupKnowledgeTool` bean,对 `golden-cases.json` 中每个 query 调用 `lookupKnowledge(query)`,并把返回的 `LookupResult` 写入 `eval/rag-retrieval/fixtures/{caseId}.json`。
|
||||
|
||||
普通测试不会执行快照生成器;只有显式传入 `rag.snapshot.enabled=true` 时才会写 fixture。
|
||||
|
||||
## 9. Seed Docs And Scope Isolation
|
||||
|
||||
Live `LookupKnowledgeTool` snapshots are only stable if the expected documents
|
||||
exist in the real knowledge base and vector index. The eval loop therefore adds
|
||||
a canonical seed layer:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/seed-docs/*.md
|
||||
-> scripts/prepare_rag_eval_seed.ps1
|
||||
-> RagEvalSeedImporterTest
|
||||
-> DocumentManagementService.uploadDocument
|
||||
-> api_document metadata + L0 index + Milvus chunks
|
||||
```
|
||||
|
||||
仓库内还保留一份 `knowledge_base/rag-eval/` 镜像,方便直接查看和提交 eval 知识库文档。它们放在单独目录下,避免和 `knowledge_base/api`、`knowledge_base/infrastructure` 等业务知识目录混在一起;检索 category 仍由 frontmatter 中的 `category` 决定。
|
||||
|
||||
Seed frontmatter includes:
|
||||
|
||||
```yaml
|
||||
source: mysql-connection-pool
|
||||
breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
```
|
||||
|
||||
`source` becomes the stable `docId` when it fits the DB column, and is also
|
||||
written to vector metadata as `_source` and `source`. `breadcrumb` is copied into
|
||||
chunk metadata so evidence blocks can keep a stable path. `kb_scope` isolates
|
||||
eval documents from local production documents.
|
||||
|
||||
Default runtime behavior keeps `retrieval.kb-scope` empty, so existing documents
|
||||
without `kb_scope` are still searchable. Eval scripts pass
|
||||
`-Dretrieval.kb-scope=rag-eval`, so L0 query hints, the filtered attempt, and
|
||||
the unfiltered retry stay inside the eval corpus while the retry still skips the
|
||||
L0 category filter.
|
||||
|
||||
Frontmatter is used for DB metadata, L0 hints, and vector metadata. It is
|
||||
stripped before document chunking so embedding content represents the Markdown
|
||||
body, not the YAML control plane. This is important for fallback eval: a decoy
|
||||
document may intentionally match L0 keywords, but its body should remain low
|
||||
quality evidence so the retry path can be exercised.
|
||||
@@ -0,0 +1,292 @@
|
||||
# 检索与可观测性架构
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-architecture.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
本文补充 [rag-architecture.md](rag-architecture.md) 中的检索细节,重点回答:
|
||||
|
||||
- 查询如何进入 `lookup_knowledge`。
|
||||
- L0 和 L1 当前分别承担什么职责。
|
||||
- 检索结果如何归一化、去重、记录。
|
||||
- 如何通过 trace 和 eval 判断检索质量。
|
||||
|
||||
当前架构与旧版最大的差异是:L0 不再因为唯一命中而默认跳过 L1,也不在 L1 失败时作为事实证据兜底。L0 是 hint 和解释信号,L1 语义检索是默认召回路径。
|
||||
|
||||
## 2. 检索总图
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Query["Agent query / AIOps recommended query"] --> Tool["LookupKnowledgeTool"]
|
||||
|
||||
Tool --> L0["KnowledgeIndexService.analyzeQuery"]
|
||||
L0 --> L0Result["L0 hint: matches / domains / keywords"]
|
||||
L0Result --> Filter["singleDomainOrNull -> category filter"]
|
||||
|
||||
Tool --> L1["VectorSearchService.searchSimilarDocuments"]
|
||||
Filter --> L1
|
||||
L1 --> Mode{"retrieval.vector-store.mode"}
|
||||
Mode -->|auto| Spring["Spring AI VectorStore"]
|
||||
Spring -->|failure| SDK["Milvus SDK fallback"]
|
||||
Mode -->|spring / spring-ai| Spring
|
||||
Mode -->|sdk| SDK
|
||||
|
||||
Spring --> Candidates["L1 candidates"]
|
||||
SDK --> Candidates
|
||||
Candidates --> Quality{"filtered L1 usable?"}
|
||||
Quality -->|no| Retry["raw query unfiltered L1 retry"]
|
||||
Quality -->|yes| Post["post-retrieval processing"]
|
||||
Retry --> Post
|
||||
L0Result --> Post
|
||||
Post --> Pack["context packing"]
|
||||
Pack --> Result["LookupResult evidenceBlocks/contextPack/traces"]
|
||||
|
||||
Result --> Dedup["RetrievedDocTracker session dedup"]
|
||||
Dedup --> Final["final tool output"]
|
||||
Final --> Invocation["tool_invocation"]
|
||||
Final --> Agent["Agent Executor"]
|
||||
```
|
||||
|
||||
## 3. L0 Hint 层
|
||||
|
||||
L0 的输入是原始 query,输出是解释性结构:
|
||||
|
||||
```text
|
||||
matches
|
||||
matchedKeywords
|
||||
domains
|
||||
singleDomainOrNull
|
||||
```
|
||||
|
||||
当前职责:
|
||||
|
||||
| 职责 | 说明 |
|
||||
|---|---|
|
||||
| domain hint | 判断 query 可能属于哪个知识域 |
|
||||
| entity / keyword hint | 记录命中的关键词、错误码、服务名等 |
|
||||
| category filter candidate | 当只有单一领域时,给 L1 一个 metadata filter 候选 |
|
||||
| trace explanation | 写入 `tool_invocation.retrieval_details`,用于解释检索为什么这么走 |
|
||||
|
||||
不再承担:
|
||||
|
||||
```text
|
||||
matches=1 -> skip L1 -> 直接返回 L0 文档正文
|
||||
L1 无可用证据 -> 返回 L0 文档正文
|
||||
```
|
||||
|
||||
原因:
|
||||
|
||||
- 子串命中不等价于最终相关性。
|
||||
- L0 没有稳定排序和语义相似度。
|
||||
- AIOps query 往往包含多个字段,单点关键词命中容易误导。
|
||||
|
||||
## 4. L1 语义检索层
|
||||
|
||||
L1 通过 `VectorSearchService` 调度,支持三种模式:
|
||||
|
||||
| 模式 | 行为 | 用途 |
|
||||
|---|---|---|
|
||||
| `auto` | 优先 Spring AI VectorStore,失败 fallback 到 SDK | 默认运行模式 |
|
||||
| `spring` / `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
|
||||
| `sdk` | 只走 Milvus SDK | 对比旧链路或临时回退 |
|
||||
|
||||
### Spring AI VectorStore 路径
|
||||
|
||||
```text
|
||||
SearchRequest
|
||||
-> query
|
||||
-> topK
|
||||
-> similarityThresholdAll
|
||||
-> optional filterExpression
|
||||
-> VectorStore.similaritySearch
|
||||
```
|
||||
|
||||
### Milvus SDK fallback
|
||||
|
||||
```text
|
||||
query
|
||||
-> VectorEmbeddingService.generateQueryVector
|
||||
-> Milvus search(vector, topK, L2)
|
||||
-> id / content / metadata
|
||||
```
|
||||
|
||||
SDK fallback 保留的价值:
|
||||
|
||||
- VectorStore bean 缺失时不让 MVP 主链路中断。
|
||||
- Spring AI collection/schema 配置异常时可回退。
|
||||
- 便于 SDK 与 VectorStore 的结果对比。
|
||||
|
||||
## 5. 分数与相关性归一化
|
||||
|
||||
检索结果输出三类分数字段:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `score` | 兼容旧逻辑的距离型分数 |
|
||||
| `rawScore` | 底层检索实现原始分数 |
|
||||
| `scoreLabel` | 原始分数语义,例如 `similarity` 或 `l2_distance` |
|
||||
|
||||
post-retrieval 层再把检索候选归一为:
|
||||
|
||||
| relevanceLevel | 含义 |
|
||||
|---|---|
|
||||
| `PRECISE` | L1 相似度高且 query hint 与候选证据互相支撑 |
|
||||
| `HIGHLY_RELEVANT` | L1 相似度高 |
|
||||
| `REFERENCE` | 可作为参考,但不足以声明强证据 |
|
||||
| `DEDUPED` | 同 session 中已检索过,不重复注入上下文 |
|
||||
|
||||
归一化结果用于:
|
||||
|
||||
- 给 Agent 输出 completeness hint。
|
||||
- 写入 `tool_invocation.relevance_level`。
|
||||
- 给 Gatekeeper 提供 `evidence_refs` 引用验真源。
|
||||
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
|
||||
- 供 EvaluationService 计算 evidence score。
|
||||
|
||||
## 6. 文档切片和 metadata
|
||||
|
||||
当前保留 Markdown-aware chunking。
|
||||
|
||||
关键 metadata:
|
||||
|
||||
```text
|
||||
docId
|
||||
chunkIndex
|
||||
totalChunks
|
||||
title
|
||||
breadcrumb
|
||||
category
|
||||
source
|
||||
```
|
||||
|
||||
embedding 输入已经增强为:
|
||||
|
||||
```text
|
||||
title + breadcrumb + content
|
||||
```
|
||||
|
||||
这解决旧版检索中的一个主要问题:单个 chunk 被召回后,LLM 不知道它属于哪个文档、哪个章节。
|
||||
|
||||
## 7. 输出和记录
|
||||
|
||||
`lookup_knowledge` 的输出会进入两条路径:
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
LookupResult["LookupResult: evidenceBlocks/contextPack/traces"] --> Agent["Agent context"]
|
||||
LookupResult --> Recorder["ToolInvocationRecorder"]
|
||||
Recorder --> Invocation["tool_invocation"]
|
||||
Invocation --> Trace["DiagnosisTraceService"]
|
||||
Invocation --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> Verifier["chat_verifier"]
|
||||
Invocation --> Eval["EvaluationService / RAG eval"]
|
||||
```
|
||||
|
||||
`tool_invocation` 中与检索相关的字段:
|
||||
|
||||
```text
|
||||
retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
relevance_level
|
||||
dedup_reason
|
||||
output_preview
|
||||
duration_ms
|
||||
success
|
||||
```
|
||||
|
||||
`retrieval_details` 承载更细信息,例如:
|
||||
|
||||
- L0 命中文档标题和路径。
|
||||
- L1 attempts、fallback reason、分数和 similarity。
|
||||
- retrieved domains。
|
||||
- evidence status。
|
||||
- dedup reason。
|
||||
- evidence block summaries。
|
||||
- evidence refs:`raw_path + text`,用于核对 Executor 的 `evidence_excerpt`。
|
||||
- context pack summary。
|
||||
- rerank trace。
|
||||
|
||||
其中 `evidence_refs` 是当前 Chat 证据链路的精确引用源:
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.evidence_blocks[0]",
|
||||
"text": "最小证据文本"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
如果检索返回 no evidence,应使用 `raw_path=$.no_evidence` 记录负向证据。它只能说明“本次检索没有匹配证据”,不能作为“问题不存在”的证明。
|
||||
|
||||
## 8. 去重与行动记忆
|
||||
|
||||
当前 session 级去重由 `RetrievedDocTracker` 负责。
|
||||
|
||||
```text
|
||||
sessionId + docKey
|
||||
-> already retrieved?
|
||||
-> yes: return dedup message and record dedup_reason
|
||||
-> no: mark retrieved and return evidence
|
||||
```
|
||||
|
||||
去重目的:
|
||||
|
||||
- 避免同一文档反复进入上下文。
|
||||
- 降低 token 浪费。
|
||||
- 给 Executor 一个“这个方向已经查过”的行动记忆。
|
||||
|
||||
注意:去重不是全局缓存,只在当前诊断 session 内生效。
|
||||
|
||||
## 9. 检索质量评测
|
||||
|
||||
检索质量不能只看一次接口返回,需要用固定 query 回归。
|
||||
|
||||
当前评测资产:
|
||||
|
||||
| 资产 | 用途 |
|
||||
|---|---|
|
||||
| `eval/rag-retrieval/cases/golden-cases.json` | 固定 query 和期望证据 |
|
||||
| `eval/rag-retrieval/fixtures/` | 离线候选结果 |
|
||||
| `eval/rag-retrieval/reports/baseline.md` | 人类可读基线 |
|
||||
| `scripts/eval_rag_retrieval.py` | 离线回归 |
|
||||
| `scripts/eval_rag_live_acceptance.py` | 运行环境验收 |
|
||||
|
||||
评测层次:
|
||||
|
||||
```text
|
||||
offline baseline
|
||||
-> 不依赖服务和外部组件
|
||||
|
||||
live acceptance
|
||||
-> 调用 /api/search/similar
|
||||
-> 验证重建索引后的真实检索
|
||||
|
||||
trace inspection
|
||||
-> 检查 Agent 是否真的调用 lookup_knowledge
|
||||
-> 检查 tool_invocation 证据是否完整
|
||||
```
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
近期优先:
|
||||
|
||||
1. 邻居 chunk / 同章节上下文扩展。
|
||||
2. metadata taxonomy 清理。
|
||||
3. Query Transformer / MultiQuery 可回退接入。
|
||||
4. 更完整的 Recall@K、MRR、nDCG 报告。
|
||||
|
||||
暂不优先:
|
||||
|
||||
- 重新引入 L0 直接返回。
|
||||
- 重新引入 L0 文档作为 L1 失败时的事实证据兜底。
|
||||
- 一次性迁移所有写入路径。
|
||||
- 在没有评测收益前引入模型 rerank / RRF / BM25。
|
||||
|
||||
@@ -0,0 +1,160 @@
|
||||
# 会话与 Trace 生命周期
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
当前 MVP 把“会话态”和“运行态”拆开:
|
||||
|
||||
```text
|
||||
chat_session(sessionId)
|
||||
-> diagnosis_run(runId)
|
||||
-> agent_step(runId)
|
||||
-> tool_invocation(runId)
|
||||
```
|
||||
|
||||
- `sessionId` 表示多轮会话目录和 Redis 上下文。
|
||||
- `runId` 表示一次可回放诊断执行。
|
||||
- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
|
||||
- `diagnosis_session` 只保留为历史兼容和回滚表。
|
||||
|
||||
## 2. 生命周期总图
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
|
||||
Resolve --> Session["ensure chat_session metadata"]
|
||||
Session --> Run["create diagnosis_run(runId)"]
|
||||
Run --> Running["run.status = RUNNING"]
|
||||
|
||||
Running --> Agent["Agent workflow"]
|
||||
Agent --> Context["execution context(sessionId, runId)"]
|
||||
Context --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step(session_id, run_id)"]
|
||||
Context --> Tool["Evidence tools"]
|
||||
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
||||
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
||||
|
||||
Agent --> Final{"workflow result"}
|
||||
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["run.status = FAILED"]
|
||||
|
||||
Success --> Evaluation["diagnosis_run.self_evaluation merge"]
|
||||
Failed --> Evaluation
|
||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
Success --> Feedback["POST /api/feedback(sessionId, runId)"]
|
||||
Feedback --> Case["useful -> case_library(run_id)"]
|
||||
```
|
||||
|
||||
## 3. ID 规则
|
||||
|
||||
| ID | 来源 | 含义 |
|
||||
|---|---|---|
|
||||
| `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
|
||||
| `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
|
||||
|
||||
设计含义:
|
||||
|
||||
- 同一个 `sessionId` 可以贯穿多轮 Chat。
|
||||
- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
|
||||
- Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
|
||||
- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
|
||||
|
||||
## 4. 运行状态流转
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> PENDING
|
||||
PENDING --> RUNNING: start diagnosis
|
||||
RUNNING --> SUCCESS: workflow completed
|
||||
RUNNING --> FAILED: exception / empty state
|
||||
SUCCESS --> SUCCESS: feedback submitted
|
||||
FAILED --> FAILED: feedback submitted
|
||||
```
|
||||
|
||||
字段边界:
|
||||
|
||||
| 字段 | 所属表 | 含义 |
|
||||
|---|---|---|
|
||||
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
||||
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
||||
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
||||
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
||||
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
|
||||
## 5. agent_step 写入
|
||||
|
||||
`AgentLoggingHook` 在模型调用前后写入和回填 `agent_step`。
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant Agent as Agent
|
||||
participant Hook as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
Agent->>Hook: before_model(messages, sessionId, runId)
|
||||
Hook->>DB: insert step(session_id, run_id, model_input, step_index)
|
||||
Agent->>Hook: after_model(output, sessionId, runId)
|
||||
Hook->>DB: update model_output, duration, token_count, has_tool_call
|
||||
```
|
||||
|
||||
新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
|
||||
|
||||
## 6. tool_invocation 写入
|
||||
|
||||
工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
|
||||
|
||||
```text
|
||||
ToolInvocationRecorder
|
||||
-> tool_invocation.session_id
|
||||
-> tool_invocation.run_id
|
||||
-> retrieval_details / evidence_refs
|
||||
```
|
||||
|
||||
Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
|
||||
|
||||
## 7. Trace API 聚合
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
聚合逻辑:
|
||||
|
||||
```text
|
||||
diagnosis_run by sessionId + runId
|
||||
+ chat_session metadata when available
|
||||
+ agent_step where run_id = runId, ordered by the Trace API
|
||||
+ tool_invocation where run_id = runId order by id
|
||||
-> DiagnosisTraceResponse
|
||||
```
|
||||
|
||||
当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
|
||||
|
||||
## 8. Chat 与 AIOps 差异
|
||||
|
||||
| 维度 | Chat | AIOps |
|
||||
|---|---|---|
|
||||
| `agent_flow` | `CHAT` | `AI_OPS` |
|
||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
|
||||
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
||||
|
||||
## 9. 清理与边界
|
||||
|
||||
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
||||
- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
|
||||
- `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
|
||||
- `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
||||
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
||||
3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
|
||||
Reference in New Issue
Block a user