docs: reorganize MVP interview documentation

This commit is contained in:
aruo
2026-07-05 15:29:28 +08:00
parent b22f2d22c8
commit 88e0a6c944
51 changed files with 4352 additions and 1318 deletions
+42
View File
@@ -0,0 +1,42 @@
# MVP 架构文档
**更新日期**:2026-07-05
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
- `mvp/architecture/archive/2026-07-05-legacy/`
归档材料只作为设计历史阅读,不再作为当前实现依据。
## 当前文档
| 文档 | 用途 |
|---|---|
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库文档编写与维护规范,覆盖 frontmatter、category、chunk、reindex |
| [data-model.md](data-model.md) | 数据模型总览,覆盖 Trace、知识库、反馈沉淀和 Milvus metadata |
| [evolution-roadmap.md](evolution-roadmap.md) | 从旧版 Agent 蓝图继承的后续演进路线,不代表当前已实现 |
## 当前架构一句话
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
## 阅读顺序
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
6. 继续读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
7. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
8. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
9. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
10. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
+187
View File
@@ -0,0 +1,187 @@
# Agent 编排架构
**更新日期**:2026-07-05
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 设计定位
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Verifier`。
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
## 2. 当前 Agent 全景
```mermaid
flowchart TB
subgraph Chat["Chat diagnosis"]
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
ChatService --> ChatPlanner["chat_planner"]
ChatPlanner --> ChatExecutor["chat_executor"]
ChatExecutor --> ChatTools["evidence tools"]
ChatTools --> ChatExecutor
ChatExecutor --> ChatVerifier["chat_verifier"]
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
ChatDecision --> ChatAnswer["final answer"]
end
subgraph AiOps["AIOps diagnosis"]
AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
AiOpsService --> Supervisor["ai_ops_supervisor"]
Supervisor --> AiOpsPlanner["planner_agent"]
Supervisor --> AiOpsExecutor["executor_agent"]
AiOpsPlanner --> AiOpsExecutor
AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
AiOpsTools --> AiOpsReport["alert report"]
AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
end
subgraph Trace["Trace persistence"]
Session["diagnosis_session"]
Step["agent_step"]
Invocation["tool_invocation"]
SelfEval["self_evaluation"]
end
ChatService --> Session
ChatPlanner --> Step
ChatExecutor --> Step
ChatVerifier --> Step
ChatTools --> Invocation
ChatDecision --> SelfEval
AiOpsService --> Session
AiOpsPlanner --> Step
AiOpsExecutor --> Step
AiOpsTools --> Invocation
AiOpsRule --> SelfEval
```
## 3. Chat 编排
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
```text
chat_planner
-> chat_executor
-> lookup_knowledge / query_logs / query_metrics / date_time
-> chat_verifier
-> reads tool_trace_summary
-> outputs verifier JSON
```
关键行为:
| 角色 | 当前职责 | 输出 |
|---|---|---|
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
| `chat_executor` | 按计划调用证据工具,组合工具返回形成诊断答复 | `executor_feedback` |
| `chat_verifier` | 只基于已有证据校验 Executor 答案,不做新检索 | `verifier_output` |
Chat 链路最多支持两轮验证:
```mermaid
sequenceDiagram
autonumber
participant C as ChatService
participant P as chat_planner
participant E as chat_executor
participant T as tools
participant V as chat_verifier
participant S as diagnosis_session
C->>P: 原始问题 + history + retry_context
P-->>C: planner_plan
C->>E: planner_plan + 上下文
E->>T: 调用证据工具
T-->>E: 证据结果
E-->>C: executor_feedback
C->>V: executor_final_answer + tool_trace_summary
V-->>C: PASS / LOW_CONFID / REJECT
C->>S: 写入 verifier_evaluation
alt LOW_CONFID 且允许补证据
C->>P: retry_context: 仅补缺失证据
else PASS 或 REJECT
C-->>S: 保存最终 answer
end
```
决策语义:
| Verdict | 行为 |
|---|---|
| `PASS` | 输出 Executor 答案 |
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
## 4. AIOps 编排
AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
```text
ai_ops_supervisor
-> planner_agent
-> executor_agent
-> final report
-> AiOpsRuleEvaluationService
```
与 Chat 的差异:
- AIOps 的输入可能是结构化告警 payload。
- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
## 5. 工具边界
当前 Executor 可用工具来自两类:
```text
methodTools
-> dateTimeTools
-> lookupKnowledgeTool
-> queryMetricsTools
-> queryLogsTools when mock enabled
ToolCallbackProvider
-> framework-discovered tools
```
工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
- L0/L1 命中数量。
- 检索层。
- relevance level。
- retrieved domains。
- dedup reason。
## 6. 与旧版设计的差异
| 旧版设想 | 当前实现 |
|---|---|
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Verifier;AIOps: Supervisor + Planner + Executor |
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
| Skill 驱动不同诊断流程 | 当前以 Prompt、知识域地图、工具调用和评测 baseline 控制 |
## 7. 后续演进
当诊断场景和工具复杂度继续上升时,再考虑拆分:
- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
拆分前提:
- 当前 Executor prompt 已难以维护。
- 不同故障类型的工具权限明显不同。
- Trace 能证明某类问题需要独立的推理策略。
- 评测集能覆盖拆分前后的行为差异。
@@ -0,0 +1,28 @@
# 旧版架构文档归档
**归档日期**:2026-07-05
本目录保存 `mvp/architecture` 下的旧版架构文档。它们包含早期 MVP 设计、旧 RAG 方案、会话存储设计、行动记忆和实施计划等历史材料。
这些文档不再作为当前实现依据。当前架构请阅读:
- `mvp/architecture/README.md`
- `mvp/architecture/current-mvp-architecture.md`
- `mvp/architecture/rag-architecture.md`
## 归档文件
| 文件 | 说明 |
|---|---|
| `agent-architecture.md` | 早期完整 Agent 设想,包含较多超出当前 MVP 的 SubAgent 设计 |
| `agent-architecture-mvp.md` | 早期 MVP Agent 设计 |
| `knowledge-retrieval-architecture.md` | 旧版 L0 + L1 检索架构,包含 L0 唯一命中跳过 L1 的旧逻辑 |
| `knowledge-retrieval-usage.md` | 旧版知识库检索使用说明 |
| `current-mvp-architecture.md` | 归档前的当前架构快照 |
| `implementation-plan.md` | 早期实施计划 |
| `implementation-detail.md` | 早期完整实施计划 |
| `session-management.md` | 会话管理旧设计 |
| `session-dedup-knowledge-map.md` | 会话去重和知识域地图设计 |
| `confidence-feedback.md` | 证据评分和用户反馈旧设计 |
| `action-memory-relevance.md` | 行动记忆和检索质量归一化旧设计 |
@@ -0,0 +1,220 @@
# Current MVP Architecture Snapshot
**Updated**: 2026-07-05
This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.
## 1. Positioning
The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.
Core goals:
- Support normal chat-based diagnosis.
- Support AIOps alert-triggered diagnosis.
- Keep tool calls explicit and traceable.
- Keep RAG retrieval observable through `lookup_knowledge`.
- Persist enough execution evidence for replay, evaluation, and interview explanation.
## 2. Runtime Architecture
```text
HTTP API
-> ChatService / AiOpsService
-> Agent orchestration
-> Supervisor / Planner / Executor / Verifier
-> Tools
-> lookup_knowledge
-> query_logs
-> query_metrics
-> other diagnosis tools
-> Persistence
-> diagnosis_session
-> agent_step
-> tool_invocation
-> Trace API
-> DiagnosisTraceService
```
Current entry points:
- `ChatService`: user-driven troubleshooting and follow-up diagnosis.
- `AiOpsService`: alert-driven diagnosis, including payload mode and auto-discovery mode.
- `DiagnosisTraceService`: trace view of session, steps, tool calls, and self-evaluation.
## 3. Chat Diagnosis Flow
```text
User question
-> ChatService
-> simple response or diagnosis flow
-> Planner creates investigation direction
-> Executor calls tools for evidence
-> lookup_knowledge
-> query_logs
-> query_metrics
-> Verifier checks final diagnosis quality
-> self_evaluation.verifier_evaluation
-> diagnosis trace
```
The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under `diagnosis_session.self_evaluation.verifier_evaluation`.
## 4. AIOps Diagnosis Flow
```text
AIOps request
-> AiOpsService
-> payload mode or auto-discovery mode
-> build alert-focused diagnosis prompt
-> append recommended lookup_knowledge query when payload exists
-> Agent diagnosis flow
-> Supervisor / Planner / Executor
-> evidence tools
-> final report
-> AiOpsRuleEvaluationService
-> self_evaluation.aiops_rule_evaluation
-> diagnosis trace
```
AIOps keeps two modes:
- Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
- Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.
The AIOps verifier is currently lightweight and rule-based. It checks:
- Whether the final report exists.
- Whether the result stays focused on the alert payload when payload exists.
- Whether evidence tools were used, especially `lookup_knowledge`, `query_logs`, and `query_metrics`.
## 5. RAG Architecture
```text
lookup_knowledge
-> L0 domain/entity hint
-> matched domain
-> matched keywords/entities
-> metadata filter signal
-> VectorSearchService
-> Spring AI VectorStore path
-> Milvus SDK fallback path
-> evidence post-processing
-> score / rawScore / scoreLabel
-> source metadata
-> title / breadcrumb / content evidence block
-> tool_invocation record
```
Important decisions:
- `lookup_knowledge` remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
- L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
- L1 retrieval now goes through `VectorSearchService`.
- Spring AI `VectorStore` is the preferred retrieval path.
- The original Milvus SDK path is retained as fallback and compatibility path.
- `title`, `breadcrumb`, and `content` participate in embedding text so chunk context is less likely to be lost.
- Retrieval output keeps compatibility fields: `score`, `rawScore`, and `scoreLabel`.
Vector retrieval modes:
```text
retrieval.vector-store.mode=auto # Prefer Spring AI VectorStore, fallback to SDK
retrieval.vector-store.mode=spring-ai # Use Spring AI VectorStore only
retrieval.vector-store.mode=sdk # Use original Milvus SDK path
```
## 6. Persistence And Trace
Current trace-related persistence:
```text
diagnosis_session
-> final_report
-> self_evaluation
-> verifier_evaluation
-> aiops_rule_evaluation
agent_step
-> role
-> step input/output
-> execution order
tool_invocation
-> tool_name
-> query
-> retrieval_layer
-> retrieval_details
-> evidence blocks
-> duration
```
Trace API aggregates these records into a session-level view:
- Agent step sequence.
- Tool calls and retrieval details.
- Final diagnosis report.
- Chat verifier status.
- AIOps rule verifier status.
## 7. Quality Gates
Current quality gates:
- Chat verifier: LLM-based final answer verification for normal diagnosis.
- AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
- Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
- RAG retrieval baseline: golden query set with offline baseline report.
- Live RAG acceptance: post-reindex script for validating retrieval against the running stack.
These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.
## 8. Current Completion State
Completed for the current MVP stage:
- Explicit `lookup_knowledge` Agent tool.
- L0 + L1 retrieval shape retained.
- L0 downgraded to domain/entity hint.
- Spring AI VectorStore retrieval path integrated.
- Milvus SDK fallback retained.
- RAG evidence post-processing added.
- Breadcrumb/title/content embedding text improved.
- RAG offline baseline and live acceptance script added.
- AIOps payload query augmentation added.
- AIOps lightweight verifier added.
- Trace summary includes both chat verifier and AIOps verifier signals.
Deferred future enhancements:
- LLM QueryTransformer / MultiQuery.
- BM25, RRF, and reranker.
- Neighbor chunk or section-level context expansion.
- VectorStore write path migration.
- Full LLM-based AIOps verifier.
- More complete golden set for recall, MRR, and nDCG metrics.
## 9. Key Code References
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
## 10. Supporting Materials
- `mvp/issues/rag-refactor-plan.md`
- `eval/rag-retrieval/README.md`
- `scripts/eval_rag_live_acceptance.py`
- `interview/rag-refactor-story.md`
- `interview/rag-vectorstore-interview-notes.md`
- `interview/rag-retrieval-quality-report.md`
- `interview/rag-breadcrumb-embedding-acceptance.md`
- `interview/aiops-query-augmentation.md`
- `interview/aiops-lightweight-verifier.md`
+341 -170
View File
@@ -1,220 +1,391 @@
# Current MVP Architecture Snapshot
# 当前 MVP 架构
**Updated**: 2026-07-05
**更新日期**:2026-07-05
**状态**:当前可运行架构
**适用范围**:Demo、面试讲解、后续迭代规划
This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.
## 1. 系统定位
## 1. Positioning
SuperBizAgent MVP 不是通用 Chatbot,而是面向故障诊断的 Agent 工程项目。
The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.
核心目标:
Core goals:
- 支持用户主动发起的 Chat 诊断。
- 支持 AIOps 告警触发的自动诊断。
- 保留 Agent 的规划、执行、验证过程。
- 工具调用必须显式、可追踪、可回放。
- RAG 检索必须通过 `lookup_knowledge` 暴露证据链。
- 每次诊断都沉淀 session、step、tool invocation 和 self evaluation。
- Support normal chat-based diagnosis.
- Support AIOps alert-triggered diagnosis.
- Keep tool calls explicit and traceable.
- Keep RAG retrieval observable through `lookup_knowledge`.
- Persist enough execution evidence for replay, evaluation, and interview explanation.
## 2. 总体分层
## 2. Runtime Architecture
```mermaid
flowchart TB
subgraph API["API Layer"]
ChatController["ChatController"]
TraceController["DiagnosisTraceController"]
SearchController["SearchController"]
DocumentController["DocumentController"]
end
```text
HTTP API
-> ChatService / AiOpsService
-> Agent orchestration
-> Supervisor / Planner / Executor / Verifier
-> Tools
-> lookup_knowledge
-> query_logs
-> query_metrics
-> other diagnosis tools
-> Persistence
-> diagnosis_session
-> agent_step
-> tool_invocation
-> Trace API
-> DiagnosisTraceService
subgraph App["Application Service"]
ChatService["ChatService"]
AiOpsService["AiOpsService"]
TraceService["DiagnosisTraceService"]
end
subgraph Agent["Agent Orchestration"]
Supervisor["Supervisor"]
Planner["Planner"]
Executor["Executor"]
Verifier["Verifier"]
end
subgraph Tools["Evidence Tools"]
KnowledgeTool["lookup_knowledge"]
LogsTool["query_logs"]
MetricsTool["query_metrics"]
AlertsTool["queryPrometheusAlerts"]
end
subgraph RAG["RAG Retrieval"]
L0["KnowledgeIndexService"]
VectorSearch["VectorSearchService"]
VectorStore["Spring AI VectorStore"]
SdkFallback["Milvus SDK fallback"]
end
subgraph Store["Persistence and Trace"]
Session["diagnosis_session"]
Step["agent_step"]
Invocation["tool_invocation"]
ApiDoc["api_document"]
Milvus["Milvus/Zilliz"]
end
API --> App
ChatService --> Agent
AiOpsService --> Agent
Agent --> Tools
KnowledgeTool --> RAG
RAG --> Store
Tools --> Invocation
Agent --> Step
App --> Session
TraceService --> Session
TraceService --> Step
TraceService --> Invocation
```
Current entry points:
- `ChatService`: user-driven troubleshooting and follow-up diagnosis.
- `AiOpsService`: alert-driven diagnosis, including payload mode and auto-discovery mode.
- `DiagnosisTraceService`: trace view of session, steps, tool calls, and self-evaluation.
## 3. Chat Diagnosis Flow
```text
User question
API Layer
-> ChatController
-> DiagnosisTraceController
-> SearchController
-> DocumentController
Application Service
-> ChatService
-> simple response or diagnosis flow
-> Planner creates investigation direction
-> Executor calls tools for evidence
-> lookup_knowledge
-> query_logs
-> query_metrics
-> Verifier checks final diagnosis quality
-> self_evaluation.verifier_evaluation
-> diagnosis trace
```
The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under `diagnosis_session.self_evaluation.verifier_evaluation`.
## 4. AIOps Diagnosis Flow
```text
AIOps request
-> AiOpsService
-> payload mode or auto-discovery mode
-> build alert-focused diagnosis prompt
-> append recommended lookup_knowledge query when payload exists
-> Agent diagnosis flow
-> Supervisor / Planner / Executor
-> evidence tools
-> final report
-> AiOpsRuleEvaluationService
-> self_evaluation.aiops_rule_evaluation
-> diagnosis trace
```
-> DiagnosisTraceService
AIOps keeps two modes:
Agent Orchestration
-> Supervisor
-> Planner
-> Executor
-> Verifier
- Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
- Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.
Evidence Tools
-> lookup_knowledge
-> query_logs
-> query_metrics
-> queryPrometheusAlerts
The AIOps verifier is currently lightweight and rule-based. It checks:
- Whether the final report exists.
- Whether the result stays focused on the alert payload when payload exists.
- Whether evidence tools were used, especially `lookup_knowledge`, `query_logs`, and `query_metrics`.
## 5. RAG Architecture
```text
lookup_knowledge
-> L0 domain/entity hint
-> matched domain
-> matched keywords/entities
-> metadata filter signal
RAG Retrieval
-> KnowledgeIndexService
-> VectorSearchService
-> Spring AI VectorStore path
-> Milvus SDK fallback path
-> evidence post-processing
-> score / rawScore / scoreLabel
-> source metadata
-> title / breadcrumb / content evidence block
-> tool_invocation record
-> Spring AI VectorStore
-> Milvus SDK fallback
Persistence
-> diagnosis_session
-> agent_step
-> tool_invocation
-> api_document
-> Milvus/Zilliz collection
Quality Gates
-> chat verifier
-> AIOps rule evaluation
-> diagnosis eval baseline
-> RAG retrieval baseline
```
Important decisions:
## 3. Chat 诊断链路
- `lookup_knowledge` remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
- L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
- L1 retrieval now goes through `VectorSearchService`.
- Spring AI `VectorStore` is the preferred retrieval path.
- The original Milvus SDK path is retained as fallback and compatibility path.
- `title`, `breadcrumb`, and `content` participate in embedding text so chunk context is less likely to be lost.
- Retrieval output keeps compatibility fields: `score`, `rawScore`, and `scoreLabel`.
```mermaid
sequenceDiagram
autonumber
actor User as 用户
participant API as POST /api/chat
participant Chat as ChatService
participant Planner as Planner Agent
participant Executor as Executor Agent
participant Tool as Evidence Tools
participant Verifier as Verifier Agent
participant DB as Trace Tables
participant Trace as Trace API
Vector retrieval modes:
User->>API: 提交诊断问题
API->>Chat: execute chat strategy
Chat->>Planner: 复杂问题进入规划
Planner->>DB: 写入 agent_step
Planner->>Executor: 下发排查方向
Executor->>Tool: lookup_knowledge / logs / metrics
Tool->>DB: 写入 tool_invocation
Tool-->>Executor: 返回证据
Executor->>Verifier: 生成候选诊断并校验
Verifier->>DB: 合并 self_evaluation.verifier_evaluation
Chat->>DB: 保存 diagnosis_session.answer
User->>Trace: GET /api/diagnosis/{sessionId}/trace
Trace->>DB: 聚合 session / step / tool
Trace-->>User: 返回可回放诊断链路
```
```text
retrieval.vector-store.mode=auto # Prefer Spring AI VectorStore, fallback to SDK
retrieval.vector-store.mode=spring-ai # Use Spring AI VectorStore only
retrieval.vector-store.mode=sdk # Use original Milvus SDK path
POST /api/chat
-> ChatService
-> 简单问题:轻量回答
-> 复杂诊断:Agent 编排
-> Planner 制定排查方向
-> Executor 调用证据工具
-> lookup_knowledge
-> query_logs
-> query_metrics
-> Verifier 校验最终诊断
-> 保存 diagnosis_session
-> 保存 agent_step
-> 保存 tool_invocation
-> 合并 self_evaluation.verifier_evaluation
```
## 6. Persistence And Trace
Chat 链路的质量门禁是 LLM Verifier。Verifier 输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
Current trace-related persistence:
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
关键代码:
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
## 4. AIOps 诊断链路
```mermaid
flowchart TD
Request["POST /api/ai_ops"] --> Payload{"包含告警 payload?"}
Payload -->|是| Targeted["PAYLOAD_TARGETED"]
Payload -->|否| Discovery["AUTO_DISCOVERY"]
Targeted --> BuildPrompt["构造聚焦 payload 的诊断 prompt"]
Targeted --> QueryAug["生成 recommended lookup_knowledge query"]
Discovery --> DiscoverAlert["通过 queryPrometheusAlerts 发现活跃告警"]
BuildPrompt --> Plan["Planner 规划排查"]
QueryAug --> Plan
DiscoverAlert --> Plan
Plan --> Execute["Executor 收集证据"]
Execute --> Knowledge["lookup_knowledge"]
Execute --> Metrics["query_metrics / Prometheus"]
Execute --> Logs["query_logs"]
Knowledge --> Report["告警分析报告"]
Metrics --> Report
Logs --> Report
Report --> RuleEval["AiOpsRuleEvaluationService"]
RuleEval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
Report --> Trace["DiagnosisTraceService"]
SelfEval --> Trace
```
```text
POST /api/ai_ops
-> AiOpsService
-> 判断是否有告警 payload
-> PAYLOAD_TARGETED
-> AUTO_DISCOVERY
-> 构造 AIOps 诊断 prompt
-> payload 模式补充 recommended lookup_knowledge query
-> Agent 编排
-> Planner / Executor
-> Prometheus / logs / knowledge tools
-> 生成告警分析报告
-> AiOpsRuleEvaluationService
-> 合并 self_evaluation.aiops_rule_evaluation
-> Trace API 可查看全链路
```
AIOps 保留两种模式:
| 模式 | 触发条件 | 行为 |
|---|---|---|
| `PAYLOAD_TARGETED` | 请求包含 alertName、service、severity、description、timeRange 等字段 | 以 payload 为唯一主诊断对象,并生成推荐知识库 query |
| `AUTO_DISCOVERY` | 请求没有明确告警 payload | 先查询当前活跃告警,再选择目标排查 |
AIOps 当前使用轻量规则验证器,重点检查:
- 最终报告是否存在。
- payload 模式是否聚焦输入告警。
- 是否使用关键证据工具,例如 `lookup_knowledge`、日志、指标。
关键代码:
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
## 5. RAG 位置
RAG 不是隐藏在 Chat Advisor 里的隐式能力,而是 Executor 可以显式调用的工具:
```mermaid
flowchart LR
Executor["Executor Agent"] --> Tool["lookup_knowledge Tool"]
Tool --> L0["L0 domain/entity hint"]
Tool --> Search["VectorSearchService"]
L0 --> Search
Search --> VectorStore["Spring AI VectorStore"]
Search --> Fallback["Milvus SDK fallback"]
VectorStore --> Normalize["score/rawScore/scoreLabel"]
Fallback --> Normalize
Normalize --> Evidence["evidence output"]
Evidence --> Invocation["tool_invocation"]
Evidence --> Executor
```
```text
Executor
-> lookup_knowledge(query)
-> L0 domain/entity hint
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
-> evidence shaping
-> tool_invocation
```
保留显式工具的原因:
- Agent 何时检索、检索什么、证据是什么,必须能在 trace 中解释。
- AIOps payload 到 query 的业务映射需要项目内控制。
- `tool_invocation` 是后续评测、回放和面试讲解的核心材料。
RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细节见 [retrieval-observability.md](retrieval-observability.md)。
## 6. 持久化模型
当前诊断持久化以三张表为核心:
```text
diagnosis_session
-> final_report
-> 一次诊断会话的主记录
-> query / status / agent_flow / answer
-> self_evaluation
-> verifier_evaluation
-> aiops_rule_evaluation
-> step_count / tool_call_count / duration
agent_step
-> role
-> step input/output
-> execution order
-> Agent 模型调用步骤
-> step_index / agent_name
-> model_input / model_output / thought
-> duration / token_count
tool_invocation
-> tool_name
-> query
-> retrieval_layer
-> retrieval_details
-> evidence blocks
-> duration
-> 工具调用事实
-> tool_name / input_params / output_preview
-> retrieval_layer / retrieval_details
-> relevance_level / dedup_reason
-> duration / success
```
Trace API aggregates these records into a session-level view:
说明:
- Agent step sequence.
- Tool calls and retrieval details.
- Final diagnosis report.
- Chat verifier status.
- AIOps rule verifier status.
- 旧的 `diagnosis_record` 已不是当前主模型,迁移脚本中已经由 `diagnosis_session + agent_step + tool_invocation` 取代。
- `api_document` 仍用于文档元数据管理。
- 文档向量内容存放在 Milvus/Zilliz collection 中。
## 7. Quality Gates
会话和 Trace 生命周期见 [session-trace-lifecycle.md](session-trace-lifecycle.md),完整数据关系见 [data-model.md](data-model.md)。
Current quality gates:
## 7. Trace API
- Chat verifier: LLM-based final answer verification for normal diagnosis.
- AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
- Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
- RAG retrieval baseline: golden query set with offline baseline report.
- Live RAG acceptance: post-reindex script for validating retrieval against the running stack.
```text
GET /api/diagnosis/{sessionId}/trace
```
These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.
Trace API 聚合:
## 8. Current Completion State
- 会话状态和最终报告。
- Agent step 序列。
- 工具调用和检索细节。
- Chat verifier 结果。
- AIOps rule evaluation 结果。
Completed for the current MVP stage:
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
- Explicit `lookup_knowledge` Agent tool.
- L0 + L1 retrieval shape retained.
- L0 downgraded to domain/entity hint.
- Spring AI VectorStore retrieval path integrated.
- Milvus SDK fallback retained.
- RAG evidence post-processing added.
- Breadcrumb/title/content embedding text improved.
- RAG offline baseline and live acceptance script added.
- AIOps payload query augmentation added.
- AIOps lightweight verifier added.
- Trace summary includes both chat verifier and AIOps verifier signals.
Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
Deferred future enhancements:
## 8. 质量门禁
- LLM QueryTransformer / MultiQuery.
- BM25, RRF, and reranker.
- Neighbor chunk or section-level context expansion.
- VectorStore write path migration.
- Full LLM-based AIOps verifier.
- More complete golden set for recall, MRR, and nDCG metrics.
当前质量门禁分层如下:
## 9. Key Code References
| 门禁 | 位置 | 作用 |
|---|---|---|
| Chat Verifier | `ChatService` | 校验普通诊断回答质量 |
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
| Live RAG Acceptance | `scripts/eval_rag_live_acceptance.py` | 在运行环境中验证重建索引后的真实检索 |
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
## 9. 当前完成状态
## 10. Supporting Materials
已经完成:
- `mvp/issues/rag-refactor-plan.md`
- `eval/rag-retrieval/README.md`
- `scripts/eval_rag_live_acceptance.py`
- `interview/rag-refactor-story.md`
- `interview/rag-vectorstore-interview-notes.md`
- `interview/rag-retrieval-quality-report.md`
- `interview/rag-breadcrumb-embedding-acceptance.md`
- `interview/aiops-query-augmentation.md`
- `interview/aiops-lightweight-verifier.md`
- Chat 和 AIOps 两条入口链路。
- 显式 `lookup_knowledge` Agent Tool。
- L0 从最终决策降级为 domain/entity hint。
- `VectorSearchService` 作为稳定检索门面。
- Spring AI VectorStore 读取路径。
- Milvus SDK fallback。
- `score` / `rawScore` / `scoreLabel` 分数语义拆分。
- `title`、`breadcrumb`、`content` 参与 embedding 文本。
- `tool_invocation` 记录检索层、relevance level、dedup reason。
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
- RAG offline baseline 和 live acceptance 脚本。
暂不作为当前已完成能力声明:
- 完整 QueryTransformer / MultiQuery。
- BM25、RRF、cross-encoder rerank。
- 完整邻居 chunk / section context expansion。
- VectorStore 写入路径全面迁移。
- 完整 LLM-based AIOps verifier。
后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
## 10. 关键代码索引
| 能力 | 代码 |
|---|---|
| Chat 入口与编排 | `ChatController`, `ChatService` |
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
| 知识库工具 | `LookupKnowledgeTool` |
| L0 hint | `KnowledgeIndexService` |
| 向量检索门面 | `VectorSearchService` |
| 文档切片 | `DocumentChunkService` |
| 向量写入 | `VectorIndexService` |
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
| Trace 聚合 | `DiagnosisTraceService` |
| 工具调用记录 | `ToolInvocationRecorder` |
| self_evaluation 合并 | `SelfEvaluationMergeService` |
+258
View File
@@ -0,0 +1,258 @@
# 数据模型总览
**更新日期**:2026-07-05
**状态**:当前可运行架构
## 1. 定位
本文从架构角度说明当前 MVP 的核心数据模型。详细字段仍以 Flyway migration 和 `mvp/tables/` 为准。
核心数据分三组:
- 诊断 Trace:`diagnosis_session`、`agent_step`、`tool_invocation`
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
- 反馈沉淀:`case_library`
## 2. 总体关系
```mermaid
erDiagram
diagnosis_session ||--o{ agent_step : has
diagnosis_session ||--o{ tool_invocation : has
diagnosis_session ||--o| case_library : creates_when_useful
api_document ||--o{ milvus_chunk : indexed_as
knowledge_domain ||--o{ api_document : groups
diagnosis_session {
bigint id
varchar session_id
text query
varchar status
varchar agent_flow
longtext answer
json self_evaluation
varchar feedback
}
agent_step {
bigint id
varchar session_id
int step_index
varchar agent_name
text model_input
text model_output
text thought
boolean has_tool_call
}
tool_invocation {
bigint id
varchar session_id
varchar tool_name
json input_params
text output_preview
varchar retrieval_layer
json retrieval_details
varchar relevance_level
varchar dedup_reason
}
api_document {
bigint id
varchar doc_id
varchar file_name
varchar file_path
varchar status
int chunk_count
text metadata
}
knowledge_domain {
bigint id
varchar domain_id
varchar description
text when_to_retrieve
int document_count
}
case_library {
bigint id
varchar case_id
varchar diagnosis_id
varchar source_type
varchar fault_category
text root_cause
text solution
}
milvus_chunk {
varchar id
text content
json metadata
vector vector
}
```
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
## 3. 诊断 Trace 模型
### diagnosis_session
会话级主记录。
关键字段:
| 字段 | 说明 |
|---|---|
| `session_id` | 外部关联键,Trace 和 Feedback 都使用它 |
| `query` | 用户原始问题或 AIOps 输入摘要 |
| `status` | 执行状态 |
| `agent_flow` | `CHAT` / `AI_OPS` |
| `answer` | 最终答复或告警报告 |
| `self_evaluation` | rule/verifier/aiops 自评估容器 |
| `feedback` | 用户反馈 |
### agent_step
记录模型调用步骤。
用途:
- 回放 Agent 推理过程。
- 查看 Planner / Executor / Verifier 的输入输出摘要。
- 统计 step count、duration、token count。
### tool_invocation
记录工具调用事实。
用途:
- 给 Trace API 展示证据。
- 给 Verifier 构造 `tool_trace_summary`。
- 给 `EvaluationService` 计算 evidence score。
- 给 RAG eval 和人工排查提供检索细节。
## 4. 知识库模型
### api_document
MySQL 中的文档元数据表。
职责:
- 管理上传文件。
- 保存 file hash,用于去重。
- 记录索引状态和 chunk 数量。
- 保存 frontmatter JSON。
### knowledge_domain
领域级元数据。
职责:
- 按 category 聚合文档。
- 存储领域描述。
- 存储 `when_to_retrieve`,辅助 Planner/Executor 判断什么时候检索该领域。
### Milvus/Zilliz metadata
向量 collection 中每个 chunk 的 metadata 主要包括:
```text
docId
_source
chunkIndex
totalChunks
title
breadcrumb
category
```
这些字段支撑:
- category filter。
- source 展示。
- breadcrumb 上下文。
- docId 删除和重建索引。
- evidence block 构造。
## 5. 反馈沉淀模型
### case_library
`useful` 反馈会触发 `CaseLibraryService.createFromSession`。
当前自动映射:
| 字段 | 来源 |
|---|---|
| `case_id` | UUID |
| `diagnosis_id` | `diagnosis_session.session_id` |
| `source_type` | `AUTO` |
| `fault_category` | 当前默认 `GENERAL` |
| `title` | session query 前 100 字符 |
| `root_cause` | session answer |
| `solution` | session answer |
| `created_by` | `system` |
## 6. self_evaluation 结构
`diagnosis_session.self_evaluation` 是 JSON 容器:
```json
{
"rule_evaluation": {},
"verifier_evaluation": {},
"aiops_rule_evaluation": {}
}
```
边界:
- `rule_evaluation` 评估证据收集充分度。
- `verifier_evaluation` 评估 Chat 答案关键事实是否有证据支撑。
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
## 7. 数据写入时序
```mermaid
sequenceDiagram
autonumber
participant API as API
participant Svc as ChatService/AiOpsService
participant Session as diagnosis_session
participant Agent as Agent
participant Step as agent_step
participant Tool as tool_invocation
participant Eval as self_evaluation
participant Feedback as case_library
API->>Svc: request
Svc->>Session: create/update RUNNING
Agent->>Step: before/after model
Agent->>Tool: tool call record
Svc->>Session: SUCCESS/FAILED + answer
Svc->>Eval: merge evaluation
API->>Svc: feedback useful
Svc->>Feedback: create case
```
## 8. 当前边界和后续
当前边界:
- `agent_step.session_id` 和 `tool_invocation.session_id` 通过 sessionId 关联,不强制外键。
- `tool_invocation.step_id` 可为空。
- Milvus chunk 与 `api_document` 通过 metadata.docId 逻辑关联。
- `case_library` 与 session 通过 `diagnosis_id=session_id` 关联。
后续可增强:
1. 增加 run id,支持同 session 多次独立诊断。
2. 强化 `tool_invocation.step_id` 关联。
3. 将 evidence block 结构化保存。
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
+179
View File
@@ -0,0 +1,179 @@
# Agent 架构演进路线
**更新日期**:2026-07-05
**状态**:后续演进设计,不代表当前已实现
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 为什么需要演进路线
旧版 `agent-architecture.md` 包含很多生产级设想:专科 SubAgent、Skill 体系、进程隔离、回退路由、MCP 工具协议化、进化引擎。它们不应作为当前 MVP 事实写入主架构,但可以作为后续扩展路线。
当前原则:
- 当前文档只声明已经可运行或明确落地的能力。
- 演进路线记录未来方向和触发条件。
- 每个演进项必须有可验证收益,不能只因为“架构更炫”就拆。
## 2. 演进总图
```mermaid
flowchart TD
MVP["Current MVP: Planner + Executor + Verifier"] --> Split{"Executor 是否过载?"}
Split -->|是| SubAgents["专科 SubAgent"]
Split -->|否| Keep["继续强化通用 Executor"]
SubAgents --> Skills["Skill / Playbook 体系"]
Skills --> Fallback["回退路由"]
Fallback --> Isolation["进程或 Pod 隔离"]
MVP --> ToolGrowth{"工具数量和来源是否增长?"}
ToolGrowth -->|是| MCP["MCP / Tool Server 协议化"]
ToolGrowth -->|否| ToolCallbacks["继续使用 @Tool / ToolCallback"]
MVP --> EvalGrowth{"评测数据是否足够?"}
EvalGrowth -->|是| Evolution["Prompt / Skill 进化引擎"]
EvalGrowth -->|否| Baseline["先扩大 baseline"]
```
## 3. 专科 SubAgent
### 触发条件
- Executor prompt 变得臃肿,难以同时覆盖接口、数据库、缓存、网络等场景。
- 不同故障类型需要明显不同的工具权限。
- Trace 显示某些场景经常走错排查路径。
- 评测集已经能衡量拆分前后的收益。
### 候选 SubAgent
| SubAgent | 场景 | 工具倾向 |
|---|---|---|
| `ExternalApiSubAgent` | 错误码、接口参数、第三方调用失败 | `lookup_knowledge`, logs, trace |
| `DatabaseSubAgent` | 连接池、慢 SQL、死锁、数据库不可用 | metrics, logs, knowledge |
| `CacheSubAgent` | Redis 超时、热点 key、内存风险 | metrics, logs, knowledge |
| `GenericDiagnosisSubAgent` | 兜底诊断 | 全量只读证据工具 |
### 不立即拆分的原因
- 当前 MVP 的工具规模还可由通用 Executor 管理。
- 过早拆分会增加 Prompt、评测和 trace 分析成本。
- 没有足够分类评测前,拆分可能只是移动复杂度。
## 4. Skill / Playbook 体系
旧版设计中的 Skill 可以在当前项目中演进为可版本化的诊断 Playbook。
```text
fault_category
-> playbook
-> required evidence
-> tool sequence
-> stop condition
-> report template
-> evaluation checks
```
优先落地方向:
- AIOps 告警处理 Playbook。
- 支付超时 Playbook。
- MySQL 连接池风险 Playbook。
- Redis timeout Playbook。
落地前提:
- 每个 Playbook 至少有 3-5 个 eval case。
- Playbook 失败时可以回退到通用 Executor。
- Trace 中能标记使用了哪个 Playbook 和哪个版本。
## 5. 回退路由
当前 Chat 已有低置信补证据和 REJECT 降级输出。后续如果引入 SubAgent,可扩展为:
```text
Specialized SubAgent
-> failed / low confidence
-> another specialized SubAgent
-> GenericDiagnosisSubAgent
-> degraded answer with confirmed facts only
```
回退依据:
- 工具连续失败。
- Verifier `REJECT`。
- Verifier `LOW_CONFID` 且补证据失败。
- Agent 输出缺失关键报告字段。
## 6. 进程隔离
当前所有 Agent 在同一 JVM 内运行。生产级隔离可以考虑:
```text
API service
-> Supervisor service
-> Planner service
-> SubAgent services
-> Verifier service
```
触发条件:
- 某类 Agent 需要独立扩缩容。
- 某类工具依赖不稳定,可能拖垮主应用。
- 不同 Agent 需要不同权限和网络访问策略。
- 单 JVM 内资源隔离不足。
MVP 阶段暂不拆分进程,优先保证 trace、评测和工具边界清晰。
## 7. MCP / Tool Server 协议化
当前工具主要通过 `@Tool`、`methodTools` 和 `ToolCallbackProvider` 暴露。工具数量增加后,可演进为:
```text
Agent
-> Tool registry
-> MCP / tool server
-> log server
-> metrics server
-> knowledge server
-> ticket/change server
```
收益:
- 工具独立部署。
- 新工具上线不必重发主应用。
- 不同 Agent 可获得不同工具子集。
- 工具调用协议统一,更利于审计。
风险:
- 调用链更长。
- 权限和超时治理更复杂。
- 本地开发和 Demo 成本上升。
## 8. 进化引擎
旧版文档提到从诊断中学习。当前可以拆成更务实的步骤:
1. 先扩大 diagnosis eval 和 RAG eval。
2. 从失败 trace 中标注 bad case。
3. 将高频失败沉淀为 Playbook 或 Prompt 规则。
4. 对 Prompt 版本做离线对比。
5. 足够稳定后再考虑线上 A/B。
不建议 MVP 直接做自动 Prompt 自优化。没有可靠评测和回滚机制时,自动优化更容易引入不可解释变化。
## 9. 演进优先级
| 优先级 | 项目 | 原因 |
|---|---|---|
| P0 | 扩大 eval baseline | 没有评测,拆任何架构都难以证明收益 |
| P1 | Playbook 化高频故障 | 可控、可解释、比拆 SubAgent 更轻 |
| P1 | 完整 evidence block | 提升 Verifier 和 Trace 质量 |
| P2 | 专科 SubAgent | 等问题类型和工具权限差异足够明显 |
| P2 | AIOps LLM Verifier | 规则门禁不足时再引入 |
| P3 | MCP 工具协议化 | 工具来源复杂后再做 |
| P3 | 进程隔离 | 生产负载和权限隔离需要明确后再做 |
+251
View File
@@ -0,0 +1,251 @@
# 反馈与自评估架构
**更新日期**:2026-07-05
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
## 1. 定位
反馈架构包含两条闭环:
1. 系统自评估:基于工具调用、Verifier、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。
2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。
当前重要边界:
- `status` 表示执行状态,不表示答案质量。
- `feedback` 表示用户反馈,不覆盖 `status`。
- `self_evaluation` 是 JSON 容器,内部按来源分层,不再把所有评分字段平铺在根节点。
## 2. 总体闭环
```mermaid
flowchart TD
Answer["Chat / AIOps final answer"] --> Session["diagnosis_session.answer"]
subgraph SelfEval["Self evaluation"]
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
Invocation --> TraceSummary["ToolTraceSummaryService"]
TraceSummary --> Verifier["chat_verifier"]
Verifier --> VerifierEval["verifier_evaluation"]
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["aiops_rule_evaluation"]
end
RuleEval --> Merge["SelfEvaluationMergeService"]
VerifierEval --> Merge
AiOpsEval --> Merge
Merge --> SelfJson["diagnosis_session.self_evaluation"]
subgraph UserFeedback["User feedback"]
UI["Feedback bar"] --> API["POST /api/feedback"]
API --> FeedbackService["FeedbackService"]
FeedbackService --> FeedbackField["diagnosis_session.feedback"]
FeedbackService --> Useful{"feedback == useful?"}
Useful -->|yes| CaseService["CaseLibraryService.createFromSession"]
CaseService --> Case["case_library"]
Useful -->|no| BadCase["Bad case by feedback=not_useful"]
end
Session --> UI
```
## 3. self_evaluation JSON
`SelfEvaluationMergeService` 统一维护 `diagnosis_session.self_evaluation`。
当前结构:
```json
{
"rule_evaluation": {
"evidence_score": 65,
"source": "rule",
"factors": []
},
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"facts_checked": [],
"rationale": "...",
"round": 1,
"traceability_version": "v1",
"tool_trace_summary": []
},
"aiops_rule_evaluation": {
"verdict": "...",
"checks": []
}
}
```
兼容逻辑:
- 如果旧 JSON 根节点包含 `evidence_score`,会被包进 `rule_evaluation`。
- 如果旧 JSON 根节点包含 `verdict` / `groundedness_score`,会被包进 `verifier_evaluation`。
## 4. 规则评分
`EvaluationService` 只消费 `tool_invocation` 和 session 状态,输出 `rule_evaluation`。
定位:
- 衡量证据收集充分度。
- 不直接证明答案是否推理正确。
- 不依赖 LLM。
规则:
| 规则名 | 条件 | 分数变化 |
|---|---|---|
| `execution_failed` | session status = `FAILED` | 直接 0 |
| `no_tool_call` | 没有工具调用 | 直接 0 |
| `has_successful_tool_call` | 至少一次工具成功 | +30 |
| `l0_exact_match` | 任意工具调用有 L0 命中 | +35 |
| `l1_semantic_match` | 无 L0 命中但有 L1 命中 | +20 |
| `retrieval_no_hit` | 有检索调用但无命中 | -10 |
| `all_tool_calls_failed` | 工具全部失败 | -20 |
最终分数裁剪到 `[0, 100]`。
说明:
- 当前 `rule_evaluation` 是异步写入,失败时 `self_evaluation` 可能暂时为空或缺少该节点。
- L0/L1 分支互斥:有 L0 命中时优先记 L0。
- 更强的答案真实性校验由 Chat Verifier 承担。
## 5. Chat Verifier 自评估
Chat Verifier 校验 Executor 的最终答案是否被证据支撑。
```mermaid
flowchart LR
Answer["executor_final_answer"] --> Verifier["chat_verifier"]
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
Summary --> Evidence["tool_trace_summary"]
Evidence --> Verifier
Verifier --> Output["verifier_output JSON"]
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"]
```
Verifier 输出:
| 字段 | 说明 |
|---|---|
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
| `groundedness_score` | 关键事实证据支撑度 |
| `critical_fact_count` | 关键事实数量 |
| `facts_checked` | 逐条事实校验 |
| `rationale` | 判定原因 |
| `tool_trace_summary` | 本次校验使用的证据索引 |
ChatService 根据 verdict 决定:
- `PASS`:输出 Executor 答案。
- `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。
- `REJECT`:降级输出,只保留已确认信息。
## 6. AIOps 规则自评估
AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。
检查重点:
- 是否有最终报告。
- payload 模式是否聚焦输入告警。
- 是否调用 `lookup_knowledge`、日志、指标等证据工具。
- 是否把无关活跃告警扩展成主诊断对象。
这是轻量规则检查,不等价于完整 LLM Verifier。完整 AIOps Verifier 是后续增强项。
## 7. 用户反馈 API
```text
POST /api/feedback
Content-Type: application/json
{
"sessionId": "xxx",
"feedback": "useful" | "not_useful"
}
```
响应:
```json
{
"success": true,
"message": "反馈已记录",
"caseId": "uuid 或 null"
}
```
后端行为:
| feedback | 行为 |
|---|---|
| `useful` | 写入 `DiagnosisSession.feedback`,调用 `CaseLibraryService.createFromSession` |
| `not_useful` | 写入 `DiagnosisSession.feedback`,不改变 session status |
| 其他值 | 返回 HTTP 400 |
## 8. 案例沉淀
`useful` 反馈会生成或复用 `case_library` 记录。
字段映射:
| CaseLibrary 字段 | 来源 |
|---|---|
| `caseId` | UUID |
| `diagnosisId` | `DiagnosisSession.sessionId` |
| `sourceType` | `AUTO` |
| `faultCategory` | 当前固定为 `GENERAL` |
| `title` | `query` 前 100 字符 |
| `rootCause` | `answer` |
| `solution` | `answer` |
| `createdBy` | `system` |
幂等性:
```text
case_library.diagnosisId == sessionId
-> existing case: return existing
-> missing case: create new
```
## 9. Trace 呈现
Trace API 会展示:
- `feedback`
- `hasFeedback`
- `hasVerifierEvaluation`
- `hasAiOpsRuleEvaluation`
- session、step、tool invocation 明细
这让一次诊断可以被分成三种视角查看:
| 视角 | 数据来源 |
|---|---|
| 执行是否成功 | `diagnosis_session.status` |
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
| 用户是否认可 | `diagnosis_session.feedback` |
## 10. 后续增强
近期优先:
1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。
2. `not_useful` 反馈沉淀 bad case,而不是只写字段。
3. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
4. AIOps 引入 LLM Verifier。
5. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
暂不优先:
- 用用户反馈直接修改 session status。
- 仅凭 `evidence_score` 判断答案正确。
- 在没有人工审核时自动把 bad case 反向写入 Prompt。
+205
View File
@@ -0,0 +1,205 @@
# Harness 与质量门禁架构
**更新日期**:2026-07-05
**状态**:当前可运行架构 + 后续门禁规划
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 设计目标
Agent 系统的核心风险不是“没有答案”,而是:
- 答案引用了不存在的证据。
- 工具调用失败后仍然编造结论。
- 检索结果相关性不足但被当作强证据。
- 多轮诊断重复检索同一文档,浪费上下文。
- 最终报告无法回放执行过程。
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
```text
Prompt contract
+ Tool boundary
+ Agent hooks
+ Trace persistence
+ Verifier / rule evaluation
+ Eval baseline
```
## 2. Harness 总图
```mermaid
flowchart TB
Input["User / AIOps input"] --> Prompt["Prompt contract"]
Prompt --> Agent["Planner / Executor / Verifier"]
Agent --> Tools["Evidence tools"]
Tools --> Invocation["tool_invocation"]
Agent --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step"]
Agent --> Session["diagnosis_session"]
Invocation --> TraceSummary["ToolTraceSummaryService"]
TraceSummary --> Verifier["chat_verifier"]
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
Session --> TraceAPI["DiagnosisTraceService"]
Step --> TraceAPI
Invocation --> TraceAPI
SelfEval --> TraceAPI
AiOpsEval --> TraceAPI
TraceAPI --> Eval["diagnosis eval / RAG eval"]
```
## 3. Prompt Contract
当前 Prompt 按角色拆分:
| Prompt | 用途 |
|---|---|
| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
| `chat-executor-prompt.md` | Chat 执行工具并形成诊断答复 |
| `chat-verifier-prompt.md` | 校验 Executor 答案是否被工具证据支撑 |
Prompt 层当前承担的门禁:
- 禁止凭记忆回答错误码、接口定义、排障步骤。
- 需要外部信息时必须调用工具。
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
- Chat Verifier 不允许做新检索,只能校验已有证据。
- AIOps payload 模式必须聚焦输入告警。
## 4. Trace Hooks
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
```mermaid
sequenceDiagram
autonumber
participant A as Agent
participant H as AgentLoggingHook
participant DB as agent_step
A->>H: before_model(messages, sessionId)
H->>DB: 写入 model_input / step_index / agent_name
A-->>A: LLM 推理
A->>H: after_model(messages, sessionId)
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
```
记录内容:
- 最近输入消息摘要。
- Agent 输出摘要。
- 是否包含 tool call。
- duration。
- token count。
- Verifier 的 JSON 输出摘要。
## 5. Tool Invocation 门禁
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
核心记录:
```text
tool_name
input_params
output_preview
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
relevance_level
dedup_reason
duration_ms
success
error_message
```
对 `lookup_knowledge` 的质量约束:
- L0 只作为 hint,不绕过 L1。
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
## 6. Verifier 门禁
Chat Verifier 的输入不是原始工具日志,而是 `ToolTraceSummaryService` 构造的证据索引。
```mermaid
flowchart LR
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
Summary --> EvidenceIndex["tool_trace_summary"]
EvidenceIndex --> Verifier["chat_verifier"]
ExecutorAnswer["executor_final_answer"] --> Verifier
Verifier --> Verdict{"verdict"}
Verdict -->|PASS| Pass["输出原答案"]
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
Verdict -->|REJECT| Reject["降级输出"]
```
Verifier 输出:
```json
{
"verdict": "PASS|LOW_CONFID|REJECT",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"facts_checked": [],
"rationale": "..."
}
```
结果写入:
```text
diagnosis_session.self_evaluation.verifier_evaluation
```
## 7. AIOps 规则门禁
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
检查重点:
- 最终报告是否存在。
- payload 模式是否围绕输入告警展开。
- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
- 是否把无关活跃告警扩展成主诊断对象。
结果写入:
```text
diagnosis_session.self_evaluation.aiops_rule_evaluation
```
## 8. Eval Baseline
当前质量门禁还包括离线评测资产:
| 评测 | 位置 | 作用 |
|---|---|---|
| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
## 9. 后续门禁规划
从旧版设计继承但尚未完整实现的门禁:
- 工具参数 schema 校验。
- 同一工具调用次数上限。
- 工具超时的统一熔断。
- 报告中的数值与工具返回值自动对齐校验。
- Prompt 版本记录和回滚。
- Verifier 对 AIOps 报告的 LLM 级事实校验。
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
+113
View File
@@ -0,0 +1,113 @@
# 面试一页式架构讲解
**用途**:面试现场 2-5 分钟讲清项目
**适合场景**:开场介绍、架构追问、Demo 前铺垫
## 1. 一句话
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Verifier 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。
## 2. 一张图
```mermaid
flowchart TB
User["用户问题 / AIOps 告警"] --> API["API Layer"]
API --> Chat["ChatService"]
API --> AiOps["AiOpsService"]
Chat --> ChatFlow["Chat: Planner -> Executor -> Verifier"]
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
ChatFlow --> Tools["Evidence Tools"]
AiOpsFlow --> Tools
Tools --> Knowledge["lookup_knowledge"]
Tools --> Logs["query_logs"]
Tools --> Metrics["query_metrics / Prometheus"]
Knowledge --> RAG["RAG: L0 hint + VectorSearchService"]
RAG --> VectorStore["Spring AI VectorStore"]
RAG --> SDK["Milvus SDK fallback"]
ChatFlow --> Trace["Trace Persistence"]
AiOpsFlow --> Trace
Tools --> Trace
Trace --> Session["diagnosis_session"]
Trace --> Step["agent_step"]
Trace --> Invocation["tool_invocation"]
Invocation --> Verifier["Verifier / Rule Evaluation"]
Verifier --> SelfEval["self_evaluation"]
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
Step --> TraceAPI
Invocation --> TraceAPI
SelfEval --> TraceAPI
TraceAPI --> Feedback["POST /api/feedback"]
Feedback --> Case["useful -> case_library"]
```
## 3. 面试讲法
```text
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
Chat 复杂问题走 Planner -> Executor -> Verifier:
Planner 负责拆解,Executor 负责调用知识库、日志和指标工具,Verifier 只基于已有工具证据校验最终答案。
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
所有过程都会落到 diagnosis_session、agent_step、tool_invocation。
所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Verifier 怎么判定、用户最后是否反馈有用。
```
## 4. 五个亮点
| 亮点 | 怎么讲 |
|---|---|
| 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool |
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
| 质量门禁 | Chat Verifier 校验 groundedness,AIOps rule evaluation 控制告警聚焦 |
| 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 |
## 5. 三个关键取舍
### 取舍 1:为什么不用隐式 Advisor 做 RAG?
因为这个项目强调 Agent 决策可见性。`lookup_knowledge` 必须作为显式工具调用被记录,这样才能解释“什么时候检索、检索了什么、证据如何支撑结论”。
### 取舍 2:为什么保留 Milvus SDK fallback?
因为迁移到 Spring AI VectorStore 期间,schema、collection、score 语义都可能变化。`auto` 模式先走 VectorStore,失败时 fallback 到 SDK,保证 MVP 主链路可运行,也方便对比新旧检索质量。
### 取舍 3:为什么 self_evaluation 分三层?
因为三类评估回答的问题不同:
```text
rule_evaluation -> 工具证据是否充分
verifier_evaluation -> Chat 答案关键事实是否有证据支撑
aiops_rule_evaluation -> AIOps 报告是否聚焦告警并使用证据
```
## 6. 面试官可能追问
| 追问 | 回答方向 |
|---|---|
| 怎么防止幻觉? | Executor 必须用工具;Verifier 只基于 `tool_trace_summary` 校验;LOW_CONFID/REJECT 会降级输出 |
| RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 |
| 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter |
| AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 |
| 下一步怎么演进? | evidence block、邻居 chunk、Playbook、AIOps LLM Verifier、MCP 工具协议化 |
## 7. 现场演示入口
- Demo 脚本:`mvp/demo/ten-minute-interview-demo.md`
- 故事案例:`interview/story-cases.md`
- 架构细节:`mvp/architecture/README.md`
@@ -0,0 +1,240 @@
# 知识库文档编写与维护
**更新日期**:2026-07-05
**状态**:当前建议规范
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-usage.md`
## 1. 定位
知识库文档不是普通 Markdown 资料堆叠,而是 RAG 检索的输入资产。写得好的文档会提升:
- L0 hint 的关键词和领域识别。
- L1 向量召回质量。
- `breadcrumb` 上下文恢复能力。
- Verifier 可引用的证据质量。
当前推荐写法:结构化 Markdown + frontmatter + 明确分类 + 可检索关键词。
## 2. 文档进入系统的链路
```mermaid
flowchart TD
Markdown["Markdown file"] --> Upload["POST /api/documents/upload"]
Upload --> Parse["FrontmatterParser"]
Parse --> Enrich["DocumentFieldEnricher"]
Enrich --> Metadata["api_document.metadata"]
Upload --> Chunk["DocumentChunkService"]
Chunk --> Breadcrumb["title / breadcrumb / chunkIndex"]
Breadcrumb --> Embedding["VectorIndexService embedding text"]
Embedding --> Milvus["Milvus/Zilliz"]
Metadata --> L0["KnowledgeIndexService L0 index"]
Milvus --> L1["VectorSearchService L1 retrieval"]
```
## 3. Frontmatter
推荐模板:
```markdown
---
title: 支付网关错误码定义
keywords: [ERR_TIMEOUT, 支付超时, payment timeout, 支付网关]
summary: 记录支付网关核心错误码的含义、常见原因和排查步骤
category: api
version: 1.0
author: sre-team
---
# 支付网关错误码定义
...
```
字段说明:
| 字段 | 必填 | 用途 |
|---|---:|---|
| `title` | 是 | 文档标题,进入 L0 索引和 embedding 上下文 |
| `keywords` | 是 | L0 hint 的主要来源 |
| `summary` | 是 | 文档摘要,进入知识域描述和 Agent 上下文 |
| `category` | 建议 | 知识域、metadata filter、上传目录 |
| `version` | 可选 | 文档版本 |
| `author` | 可选 | 维护人 |
当前解析器会提示缺少 `title`、`keywords`、`summary` 的情况;缺失不一定阻断上传,但会降低检索质量。
## 4. category 建议
`category` 会影响:
- 上传文件本地目录。
- Milvus metadata。
- L0 domain hint。
- `knowledge_domain` 聚合。
- VectorStore / SDK category filter。
推荐保持稳定,不要频繁换名。
| category | 用途 |
|---|---|
| `api` | 接口、错误码、请求/响应协议 |
| `infrastructure` | MySQL、Redis、JVM、网络、中间件 |
| `troubleshooting` | 通用排障流程、Runbook |
| `domain` | 业务领域规则 |
| `spring-ai` | Spring AI / Agent / 工具最佳实践 |
注意:分类过细会导致 filter 召回不足;分类过粗会降低 L0 hint 解释力。
## 5. 关键词写法
好的关键词应该覆盖:
- 精确实体:错误码、服务名、指标名。
- 常用中文说法。
- 英文别名。
- 组合词。
示例:
```yaml
keywords: [ERR_TIMEOUT, timeout, 支付超时, 支付网关超时, payment-service, gateway timeout]
```
避免:
```yaml
keywords: [错误, 问题, 系统]
```
原因:过宽关键词会让 L0 hint 变脏,多个文档同时命中,影响解释性和 category filter。
## 6. Markdown 结构
推荐结构:
```markdown
# 文档总标题
## 场景或错误码
### 含义
### 常见原因
### 排查步骤
### 处理方案
### 日志示例
```
为什么这样写:
- `DocumentChunkService` 会按 Markdown 标题切分。
- 标题层级会生成 `breadcrumb`。
- `title + breadcrumb + content` 会一起进入 embedding 文本。
- 命中 chunk 时,Agent 更容易知道证据属于哪个章节。
## 7. 内容建议
每个可诊断条目尽量包含:
- 现象。
- 判断条件。
- 可能原因。
- 证据来源。
- 排查步骤。
- 处理建议。
- 日志或配置示例。
示例:
```markdown
## ERR_TIMEOUT
### 含义
支付网关请求超过本地或上游超时时间。
### 常见原因
1. 第三方支付服务响应慢。
2. 本地 timeout 配置过短。
3. 网络链路抖动。
### 排查步骤
1. 查询 payment-service 日志中的请求耗时。
2. 查看网关 5xx 和 timeout 指标。
3. 对比当前 timeout 配置。
### 处理建议
- 短期:重试受影响订单。
- 长期:调整 timeout 和重试策略,并监控上游延迟。
```
## 8. 上传与索引
上传接口:
```text
POST /api/documents/upload
Content-Type: multipart/form-data
file=<markdown file>
category=<category>
```
系统处理:
1. 计算文件 hash,避免重复上传。
2. 保存原始文件。
3. 解析 frontmatter。
4. 补全文档字段。
5. 写入 `api_document`。
6. Markdown-aware chunking。
7. 写入 Milvus/Zilliz。
8. 更新 L0 索引和 `knowledge_domain`。
## 9. 重建索引注意事项
当以下内容变化时,需要重新索引:
- 正文内容。
- 标题层级。
- `category`。
- `title`、`summary`、`keywords`。
- embedding 输入策略,例如加入 `breadcrumb`。
特别注意:
```text
修改 Markdown 或 embedding 输入策略,不会自动改变已有向量。
必须重新上传或重建索引后,live retrieval 才能体现变化。
```
可用 live 验收:
```bash
python scripts/eval_rag_live_acceptance.py
```
## 10. 维护 checklist
新增文档前检查:
- frontmatter 是否包含 `title`、`keywords`、`summary`。
- `category` 是否属于现有稳定分类。
- 关键词是否既有精确词也有常用表达。
- Markdown 标题层级是否清晰。
- 每个故障条目是否包含可执行排查步骤。
- 日志/配置示例是否脱敏。
更新文档后检查:
- `api_document.status` 是否为 `INDEXED`。
- `/api/search/similar` 是否能搜到目标文档。
- `eval/rag-retrieval` 是否需要新增 golden case。
- Trace 中 `tool_invocation` 是否记录到正确 source 和 breadcrumb。
+414
View File
@@ -0,0 +1,414 @@
# RAG 新架构
**更新日期**:2026-07-05
**状态**:当前主架构 + 后续演进边界
**关联计划**:`mvp/issues/rag-refactor-plan.md`
## 1. 架构目标
RAG 重构的目标不是把所有能力交给框架,也不是继续维护一套完全自研检索框架,而是形成:
```text
成熟框架能力 + 业务可观测编排
```
具体原则:
- 通用向量检索能力交给 Spring AI `VectorStore`。
- 项目保留 Agent Tool 入口、AIOps 业务 query 映射、证据打包、trace 记录。
- `lookup_knowledge` 继续是显式工具,不替换成隐式 Advisor。
- Spring AI 读取路径作为主路径,Milvus SDK 作为 fallback。
- 所有检索行为必须可评测、可回放、可解释。
## 2. 当前主链路
```mermaid
flowchart TD
Agent["Agent Executor"] --> Tool["lookup_knowledge(query)"]
Tool --> L0["KnowledgeIndexService.analyzeQuery"]
L0 --> Hint["L0 hint: domain / entities / matchedKeywords"]
Hint --> Filter["category filter candidate"]
Tool --> Search["VectorSearchService.searchSimilarDocuments"]
Filter --> Search
Search --> Mode{"retrieval.vector-store.mode"}
Mode -->|auto| SpringTry["try Spring AI VectorStore"]
SpringTry -->|success| Results["SearchResult list"]
SpringTry -->|failure| SdkFallback["Milvus SDK fallback"]
Mode -->|spring-ai| SpringOnly["Spring AI VectorStore only"]
Mode -->|sdk| SdkOnly["Milvus SDK only"]
SpringOnly --> Results
SdkFallback --> Results
SdkOnly --> Results
Results --> Normalize["relevance normalization"]
Normalize --> Dedup["session dedup: RetrievedDocTracker"]
Dedup --> Output["LookupResult"]
Output --> Record["tool_invocation record"]
Output --> Agent
```
```text
Agent Executor
-> lookup_knowledge(query)
-> KnowledgeIndexService.analyzeQuery
-> L0 domain/entity hint
-> matchedKeywords
-> category filter candidate
-> VectorSearchService.searchSimilarDocuments
-> mode=auto
-> Spring AI VectorStore
-> fallback: Milvus SDK
-> mode=spring-ai
-> Spring AI VectorStore only
-> mode=sdk
-> Milvus SDK only
-> result normalization
-> relevanceLevel
-> completenessHint
-> score/rawScore/scoreLabel
-> session dedup
-> RetrievedDocTracker
-> tool_invocation record
```
运行配置:
```properties
retrieval.vector-store.mode=auto
retrieval.normalization.max-l2-distance=2.0
retrieval.normalization.highly-relevant-threshold=0.75
retrieval.normalization.reference-threshold=0.5
```
## 3. 稳定边界
```mermaid
flowchart LR
subgraph AgentBoundary["Agent boundary"]
Executor["Executor Agent"]
Tool["LookupKnowledgeTool"]
end
subgraph RetrievalBoundary["Retrieval boundary"]
Search["VectorSearchService"]
Spring["Spring AI VectorStore"]
SDK["Milvus SDK"]
end
subgraph ObservabilityBoundary["Observability boundary"]
Invocation["tool_invocation"]
Eval["RAG baseline / trace inspection"]
end
Executor --> Tool
Tool --> Search
Search --> Spring
Search --> SDK
Tool --> Invocation
Invocation --> Eval
```
### 3.1 Agent 边界
Agent 只知道自己可以调用 `lookup_knowledge`,不直接关心底层是 Spring AI VectorStore 还是 Milvus SDK。
```text
Executor -> LookupKnowledgeTool -> VectorSearchService
```
这个边界让 RAG 底层迁移不影响 Agent prompt、工具声明和 trace 数据结构。
### 3.2 检索边界
`VectorSearchService` 是当前检索门面:
- `auto`:优先 Spring AI VectorStore,失败后 fallback 到 SDK。
- `spring-ai`:只走 Spring AI VectorStore。
- `sdk`:只走原 Milvus SDK。
这样可以在不改 Agent 工具的情况下切换检索实现,并支持线上验证和回退。
### 3.3 可观测边界
无论底层检索路径如何变化,都必须写入 `tool_invocation`:
```text
sessionId
toolName
inputParams
outputPreview
retrievalLayer
l0MatchCount
l1MatchCount
retrievalDetails
relevanceLevel
dedupReason
duration
success
```
## 4. L0 的新职责
旧版 L0 容易承担过重职责,例如唯一匹配后直接跳过 L1。当前架构中 L0 被降级为 hint 层。
```mermaid
flowchart TD
Input["query / AIOps payload"] --> L0["L0 hint analysis"]
L0 --> Domain["domain detector"]
L0 --> Entity["entity extractor"]
L0 --> Keyword["matched keyword explanation"]
L0 --> Filter["metadata/category filter candidate"]
Domain --> Retrieval["L1 semantic retrieval"]
Entity --> Retrieval
Keyword --> Trace["hit reason in tool_invocation"]
Filter --> Retrieval
Retrieval --> Normalize["relevance normalization"]
Normalize --> Evidence["evidence returned to Agent"]
```
L0 负责:
- domain detector
- entity extractor
- matched keyword explanation
- metadata/category filter candidate
- trace 中的 hit reason
L0 不再默认负责:
```text
L0 unique hit -> 直接作为最终检索结果
```
当前职责是:
```text
query / AIOps payload
-> L0 matched keywords / domains / entities
-> category filter candidate
-> L1 semantic retrieval
-> relevance normalization
```
这样既保留精确关键词和领域 hint 的价值,也避免 L0 误召回直接污染最终证据。
## 5. L1 向量检索
L1 语义检索通过 `VectorSearchService` 调度。
```mermaid
flowchart TD
Search["VectorSearchService"] --> Request["SearchRequest: query / topK / threshold / filter"]
Request --> VectorStore["Spring AI VectorStore"]
VectorStore --> Docs["Document results"]
Docs --> Map["map to SearchResult"]
Map --> Score["score compatibility mapping"]
Search --> SDK["Milvus SDK fallback"]
SDK --> SdkRows["id / content / metadata / L2 distance"]
SdkRows --> Map
Score --> Output["id / content / metadata / score / rawScore / scoreLabel"]
```
### Spring AI VectorStore 路径
```text
SearchRequest
-> query
-> topK
-> similarityThresholdAll
-> optional filterExpression: category == '...'
-> VectorStore.similaritySearch
```
返回结果会映射为项目兼容结构:
```text
id
content
metadata
score
rawScore
scoreLabel
```
### Milvus SDK fallback
SDK 路径仍保留:
- 用于 `auto` 模式兜底。
- 用于与旧链路对比。
- 用于 VectorStore 配置或 collection schema 异常时保证 MVP 可运行。
## 6. 分数语义
旧 SDK 使用 L2 distance,Spring AI 返回 similarity。两者不能混用为同一个含义。
当前统一输出:
| 字段 | 含义 |
|---|---|
| `score` | 兼容旧逻辑的距离型分数,越小越近 |
| `rawScore` | 底层实现的原始分数 |
| `scoreLabel` | `l2_distance` 或 `similarity` |
SDK 路径:
```text
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
```
VectorStore 路径:
```text
rawScore = Spring AI similarity
scoreLabel = similarity
score = metadata.distance if available else compatible distance
```
## 7. 文档切片和 embedding 输入
当前保留 Markdown-aware chunking:
- 识别 Markdown 标题层级。
- 生成 `title`。
- 生成 `breadcrumb`。
- 保留 `chunkIndex`。
- 使用 token 估算和软/硬上限控制 chunk 大小。
- 尽量不打断列表和代码块。
embedding 输入中已经加强:
```text
title + breadcrumb + content
```
这样可以降低单个 chunk 脱离章节上下文后的召回损失。
## 8. AIOps query 增强
AIOps payload 中的业务字段不能完全交给通用检索框架隐式理解。
payload 模式会把以下字段拼成推荐知识库 query:
- `alertName`
- `service`
- `severity`
- `description`
- `timeRange`
- `userRequest`
Prompt 会明确要求 Agent 在需要知识库证据时,优先使用推荐 query 或保留 alertName/service 的更窄 query。
```text
AIOps payload
-> buildKnowledgeRetrievalQuery
-> Recommended lookup_knowledge query
-> lookup_knowledge
-> tool_invocation
```
## 9. Evidence 与去重
当前 evidence 输出仍以 `LookupResult` 和工具返回文本为主,已经具备:
- L0/L1 命中数量。
- 检索层记录。
- relevance level。
- completeness hint。
- session 级文档去重。
- domain 行动记忆。
- `tool_invocation` 明细记录。
后续更完整的 evidence block 目标:
```text
source
docId
chunkIndex
title
breadcrumb
score
rawScore
scoreLabel
hitReason
content
expandedFrom
```
这部分应作为下一阶段增强,而不是当前已完全完成能力。
## 10. 评测与验收
RAG 架构变更必须先过评测,再认为可合入主链路。
当前评测资产:
- `eval/rag-retrieval/cases/golden-cases.json`
- `eval/rag-retrieval/fixtures/`
- `eval/rag-retrieval/reports/baseline.json`
- `eval/rag-retrieval/reports/baseline.md`
- `scripts/eval_rag_retrieval.py`
- `scripts/eval_rag_live_acceptance.py`
评测层次:
| 层次 | 作用 |
|---|---|
| Offline baseline | 不依赖 MySQL、Redis、Milvus、LLM,用固定 fixtures 检查召回行为 |
| Live acceptance | 应用运行并重建索引后,调用 `/api/search/similar` 验证真实检索 |
| Trace inspection | 通过 `tool_invocation` 检查 Agent 是否真的使用了证据 |
## 11. 当前已完成
- `lookup_knowledge` 保持显式 Agent Tool。
- L0 降级为 domain/entity hint。
- L1 默认执行语义检索。
- `VectorSearchService` 支持 `auto`、`spring-ai`、`sdk` 三种模式。
- Spring AI VectorStore 成为读取主路径。
- Milvus SDK fallback 保留。
- 分数语义拆成 `score`、`rawScore`、`scoreLabel`。
- Markdown chunk 保留 `title` 和 `breadcrumb`。
- embedding 输入包含 `title`、`breadcrumb` 和 `content`。
- AIOps payload 生成推荐知识库 query。
- `tool_invocation` 记录 relevance level 和 dedup reason。
- RAG offline baseline 和 live acceptance 脚本已补齐。
## 12. 后续演进
近期优先:
1. 完整 evidence block 结构化输出。
2. 命中 chunk 的相邻 chunk / 同章节上下文扩展。
3. metadata taxonomy 清理,例如 `database` 与 `infrastructure` 的分类边界。
4. Query Transformer / MultiQuery 的可回退接入。
5. VectorStore 写入路径评估。
暂不优先:
- 把 `lookup_knowledge` 替换成隐式 Advisor。
- 完整自研 RRF 框架。
- 立即引入 Elasticsearch / OpenSearch。
- 立即引入 cross-encoder 或 LLM rerank。
## 13. 关键代码索引
| 能力 | 代码 |
|---|---|
| Agent 工具入口 | `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java` |
| L0 hint | `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java` |
| 向量检索门面 | `src/main/java/com/superbiz/agent/service/VectorSearchService.java` |
| 文档切片 | `src/main/java/com/superbiz/agent/service/DocumentChunkService.java` |
| 文档管理 | `src/main/java/com/superbiz/agent/service/DocumentManagementService.java` |
| 向量写入 | `src/main/java/com/superbiz/agent/service/VectorIndexService.java` |
| AIOps query 增强 | `src/main/java/com/superbiz/agent/service/AiOpsService.java` |
| 工具调用记录 | `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java` |
+266
View File
@@ -0,0 +1,266 @@
# 检索与可观测性架构
**更新日期**:2026-07-05
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-architecture.md`
## 1. 定位
本文补充 [rag-architecture.md](rag-architecture.md) 中的检索细节,重点回答:
- 查询如何进入 `lookup_knowledge`。
- L0 和 L1 当前分别承担什么职责。
- 检索结果如何归一化、去重、记录。
- 如何通过 trace 和 eval 判断检索质量。
当前架构与旧版最大的差异是:L0 不再因为唯一命中而默认跳过 L1。L0 是 hint 和解释信号,L1 语义检索是默认召回路径。
## 2. 检索总图
```mermaid
flowchart TD
Query["Agent query / AIOps recommended query"] --> Tool["LookupKnowledgeTool"]
Tool --> L0["KnowledgeIndexService.analyzeQuery"]
L0 --> L0Result["L0 hint: matches / domains / keywords"]
L0Result --> Filter["singleDomainOrNull -> category filter"]
Tool --> L1["VectorSearchService.searchSimilarDocuments"]
Filter --> L1
L1 --> Mode{"retrieval.vector-store.mode"}
Mode -->|auto| Spring["Spring AI VectorStore"]
Spring -->|failure| SDK["Milvus SDK fallback"]
Mode -->|spring-ai| Spring
Mode -->|sdk| SDK
Spring --> Candidates["L1 candidates"]
SDK --> Candidates
Candidates --> Normalize["relevance normalization"]
L0Result --> Normalize
Normalize --> Result["LookupResult"]
Result --> Dedup["RetrievedDocTracker session dedup"]
Dedup --> Final["final tool output"]
Final --> Invocation["tool_invocation"]
Final --> Agent["Agent Executor"]
```
## 3. L0 Hint 层
L0 的输入是原始 query,输出是解释性结构:
```text
matches
matchedKeywords
domains
singleDomainOrNull
```
当前职责:
| 职责 | 说明 |
|---|---|
| domain hint | 判断 query 可能属于哪个知识域 |
| entity / keyword hint | 记录命中的关键词、错误码、服务名等 |
| category filter candidate | 当只有单一领域时,给 L1 一个 metadata filter 候选 |
| trace explanation | 写入 `tool_invocation.retrieval_details`,用于解释检索为什么这么走 |
不再承担:
```text
matches=1 -> skip L1 -> 直接返回 L0 文档正文
```
原因:
- 子串命中不等价于最终相关性。
- L0 没有稳定排序和语义相似度。
- AIOps query 往往包含多个字段,单点关键词命中容易误导。
## 4. L1 语义检索层
L1 通过 `VectorSearchService` 调度,支持三种模式:
| 模式 | 行为 | 用途 |
|---|---|---|
| `auto` | 优先 Spring AI VectorStore,失败 fallback 到 SDK | 默认运行模式 |
| `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
| `sdk` | 只走 Milvus SDK | 对比旧链路或临时回退 |
### Spring AI VectorStore 路径
```text
SearchRequest
-> query
-> topK
-> similarityThresholdAll
-> optional filterExpression
-> VectorStore.similaritySearch
```
### Milvus SDK fallback
```text
query
-> VectorEmbeddingService.generateQueryVector
-> Milvus search(vector, topK, L2)
-> id / content / metadata
```
SDK fallback 保留的价值:
- VectorStore bean 缺失时不让 MVP 主链路中断。
- Spring AI collection/schema 配置异常时可回退。
- 便于 SDK 与 VectorStore 的结果对比。
## 5. 分数与相关性归一化
检索结果输出三类分数字段:
| 字段 | 说明 |
|---|---|
| `score` | 兼容旧逻辑的距离型分数 |
| `rawScore` | 底层检索实现原始分数 |
| `scoreLabel` | 原始分数语义,例如 `similarity` 或 `l2_distance` |
工具层再把 L0/L1 情况归一为:
| relevanceLevel | 含义 |
|---|---|
| `PRECISE` | L0 单命中且 L1 相似度高 |
| `HIGHLY_RELEVANT` | L1 相似度高,或 L0 多命中且 L1 支撑强 |
| `REFERENCE` | 可作为参考,但不足以声明强证据 |
| `DEDUPED` | 同 session 中已检索过,不重复注入上下文 |
归一化结果用于:
- 给 Agent 输出 completeness hint。
- 写入 `tool_invocation.relevance_level`。
- 给 Verifier 构造 `tool_trace_summary`。
- 供 EvaluationService 计算 evidence score。
## 6. 文档切片和 metadata
当前保留 Markdown-aware chunking。
关键 metadata:
```text
docId
chunkIndex
totalChunks
title
breadcrumb
category
source
```
embedding 输入已经增强为:
```text
title + breadcrumb + content
```
这解决旧版检索中的一个主要问题:单个 chunk 被召回后,LLM 不知道它属于哪个文档、哪个章节。
## 7. 输出和记录
`lookup_knowledge` 的输出会进入两条路径:
```mermaid
flowchart LR
LookupResult["LookupResult"] --> Agent["Agent context"]
LookupResult --> Recorder["ToolInvocationRecorder"]
Recorder --> Invocation["tool_invocation"]
Invocation --> Trace["DiagnosisTraceService"]
Invocation --> Summary["ToolTraceSummaryService"]
Summary --> Verifier["chat_verifier"]
Invocation --> Eval["EvaluationService / RAG eval"]
```
`tool_invocation` 中与检索相关的字段:
```text
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
relevance_level
dedup_reason
output_preview
duration_ms
success
```
`retrieval_details` 承载更细信息,例如:
- L0 命中文档标题和路径。
- L1 分数。
- retrieved domains。
- evidence status。
- dedup reason。
## 8. 去重与行动记忆
当前 session 级去重由 `RetrievedDocTracker` 负责。
```text
sessionId + docKey
-> already retrieved?
-> yes: return dedup message and record dedup_reason
-> no: mark retrieved and return evidence
```
去重目的:
- 避免同一文档反复进入上下文。
- 降低 token 浪费。
- 给 Executor 一个“这个方向已经查过”的行动记忆。
注意:去重不是全局缓存,只在当前诊断 session 内生效。
## 9. 检索质量评测
检索质量不能只看一次接口返回,需要用固定 query 回归。
当前评测资产:
| 资产 | 用途 |
|---|---|
| `eval/rag-retrieval/cases/golden-cases.json` | 固定 query 和期望证据 |
| `eval/rag-retrieval/fixtures/` | 离线候选结果 |
| `eval/rag-retrieval/reports/baseline.md` | 人类可读基线 |
| `scripts/eval_rag_retrieval.py` | 离线回归 |
| `scripts/eval_rag_live_acceptance.py` | 运行环境验收 |
评测层次:
```text
offline baseline
-> 不依赖服务和外部组件
live acceptance
-> 调用 /api/search/similar
-> 验证重建索引后的真实检索
trace inspection
-> 检查 Agent 是否真的调用 lookup_knowledge
-> 检查 tool_invocation 证据是否完整
```
## 10. 后续增强
近期优先:
1. 完整 evidence block 输出。
2. 邻居 chunk / 同章节上下文扩展。
3. metadata taxonomy 清理。
4. Query Transformer / MultiQuery 可回退接入。
5. 更完整的 Recall@K、MRR、nDCG 报告。
暂不优先:
- 重新引入 L0 直接返回。
- 一次性迁移所有写入路径。
- 在没有评测收益前引入 rerank / RRF / BM25。
+196
View File
@@ -0,0 +1,196 @@
# 会话与 Trace 生命周期
**更新日期**:2026-07-05
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
## 1. 定位
旧版会话设计以 Redis 会话为主,MySQL 作为可选长期沉淀。当前 MVP 的可追踪诊断已经转为 MySQL Trace 三表为主:
```text
diagnosis_session
-> agent_step
-> tool_invocation
```
因此本文描述的是当前可运行链路:
- `sessionId` 是一次诊断和后续 trace/feedback 的关联键。
- `diagnosis_session` 保存会话级状态、问题、答案、自评估和反馈。
- `agent_step` 保存每个 Agent 模型调用。
- `tool_invocation` 保存工具调用事实。
- `DiagnosisTraceService` 聚合三类记录,形成可回放 trace。
## 2. 生命周期总图
```mermaid
flowchart TD
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
Resolve --> Create["create or reset diagnosis_session"]
Create --> Running["status = RUNNING"]
Running --> Agent["Agent workflow"]
Agent --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step"]
Agent --> Tool["Evidence tools"]
Tool --> Invocation["tool_invocation"]
Agent --> Final{"workflow result"}
Final -->|success| Success["status = SUCCESS, answer saved"]
Final -->|failed| Failed["status = FAILED"]
Success --> Evaluation["self_evaluation merge"]
Failed --> Evaluation
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace"]
Success --> Feedback["POST /api/feedback"]
Feedback --> Case["useful -> case_library"]
```
## 3. sessionId 规则
| 链路 | sessionId 来源 |
|---|---|
| Chat | 如果请求带 sessionId,则复用;否则生成短 UUID |
| AIOps | 如果 payload 带 sessionId,则复用;否则生成 UUID |
| Trace | URL path 中的 `{sessionId}` |
| Feedback | request body 中的 `sessionId` |
设计含义:
- 同一个 `sessionId` 可以贯穿诊断、trace 查询和用户反馈。
- 当前诊断开始时会重置当前 session 的运行态字段,例如 answer、duration、step/tool count。
- `sessionId` 是业务关联键,不依赖数据库自增 ID 暴露给外部。
## 4. 状态流转
```mermaid
stateDiagram-v2
[*] --> PENDING
PENDING --> RUNNING: start diagnosis
RUNNING --> SUCCESS: workflow completed
RUNNING --> FAILED: exception / empty state
SUCCESS --> SUCCESS: feedback submitted
FAILED --> FAILED: feedback submitted
```
字段边界:
| 字段 | 含义 |
|---|---|
| `status` | 执行状态:`PENDING` / `RUNNING` / `SUCCESS` / `FAILED` |
| `answer` | Agent 最终返回给用户的报告或答复 |
| `self_evaluation` | 系统自评估 JSON |
| `feedback` | 用户反馈:`useful` / `not_useful` / null |
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 session,仍然应该是 `SUCCESS + feedback=not_useful`。
## 5. agent_step 写入
`AgentLoggingHook` 在模型调用前后写入和回填 `agent_step`。
```mermaid
sequenceDiagram
autonumber
participant Agent as ReactAgent
participant Hook as AgentLoggingHook
participant DB as agent_step
Agent->>Hook: before_model(messages, sessionId)
Hook->>DB: insert step_index / agent_name / model_input
Agent-->>Agent: model call
Agent->>Hook: after_model(messages, sessionId)
Hook->>DB: update model_output / thought / has_tool_call / duration / token_count
```
当前记录:
- `session_id`
- `step_index`
- `agent_name`
- `model_input`
- `model_output`
- `thought`
- `has_tool_call`
- `duration_ms`
- `token_count`
## 6. tool_invocation 写入
工具调用记录真实工具事实,不记录模型猜测。
关键字段:
```text
session_id
step_id
tool_name
input_params
output_preview
output_length
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
relevance_level
dedup_reason
duration_ms
success
error_message
```
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对非检索工具,检索字段可以为空。
## 7. Trace API 聚合
```text
GET /api/diagnosis/{sessionId}/trace
```
聚合逻辑:
```text
diagnosis_session by sessionId
+ agent_step ordered by step_index
+ tool_invocation ordered by id
-> DiagnosisTraceResponse
```
Trace 视图回答的问题:
- 这次诊断是否成功?
- 哪些 Agent 参与了?
- 每一步模型输入输出是什么摘要?
- 调用了哪些工具?
- 工具返回了什么证据?
- Verifier / AIOps rule 是否通过?
- 用户是否反馈有用?
## 8. Chat 与 AIOps 差异
| 维度 | Chat | AIOps |
|---|---|---|
| `agent_flow` | `CHAT` | `AI_OPS` |
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Verifier | `SupervisorAgent`: Planner + Executor |
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
| 答案字段 | Chat 最终答复 | 告警分析报告 |
| payload | 用户自然语言 + history | alert payload 或 auto-discovery |
## 9. 清理与边界
当前会话持久化边界:
- MySQL Trace 记录是主要可回放来源。
- Chat 历史仍可作为请求上下文传入 Agent,但不是本文档的主持久化模型。
- Redis 主会话存储是历史设计,不作为当前架构事实。
- `RetrievedDocTracker` 是 session 级运行时去重状态,诊断结束后清理。
## 10. 后续增强
可考虑:
1. Trace API 增加更结构化的 `self_evaluation` 展示。
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。
4. 为 Trace 增加导出能力,服务面试演示和回归分析。