docs: reorganize MVP interview documentation

This commit is contained in:
aruo
2026-07-05 15:29:28 +08:00
parent b22f2d22c8
commit 88e0a6c944
51 changed files with 4352 additions and 1318 deletions
+258
View File
@@ -0,0 +1,258 @@
# 数据模型总览
**更新日期**:2026-07-05
**状态**:当前可运行架构
## 1. 定位
本文从架构角度说明当前 MVP 的核心数据模型。详细字段仍以 Flyway migration 和 `mvp/tables/` 为准。
核心数据分三组:
- 诊断 Trace:`diagnosis_session`、`agent_step`、`tool_invocation`
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
- 反馈沉淀:`case_library`
## 2. 总体关系
```mermaid
erDiagram
diagnosis_session ||--o{ agent_step : has
diagnosis_session ||--o{ tool_invocation : has
diagnosis_session ||--o| case_library : creates_when_useful
api_document ||--o{ milvus_chunk : indexed_as
knowledge_domain ||--o{ api_document : groups
diagnosis_session {
bigint id
varchar session_id
text query
varchar status
varchar agent_flow
longtext answer
json self_evaluation
varchar feedback
}
agent_step {
bigint id
varchar session_id
int step_index
varchar agent_name
text model_input
text model_output
text thought
boolean has_tool_call
}
tool_invocation {
bigint id
varchar session_id
varchar tool_name
json input_params
text output_preview
varchar retrieval_layer
json retrieval_details
varchar relevance_level
varchar dedup_reason
}
api_document {
bigint id
varchar doc_id
varchar file_name
varchar file_path
varchar status
int chunk_count
text metadata
}
knowledge_domain {
bigint id
varchar domain_id
varchar description
text when_to_retrieve
int document_count
}
case_library {
bigint id
varchar case_id
varchar diagnosis_id
varchar source_type
varchar fault_category
text root_cause
text solution
}
milvus_chunk {
varchar id
text content
json metadata
vector vector
}
```
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
## 3. 诊断 Trace 模型
### diagnosis_session
会话级主记录。
关键字段:
| 字段 | 说明 |
|---|---|
| `session_id` | 外部关联键,Trace 和 Feedback 都使用它 |
| `query` | 用户原始问题或 AIOps 输入摘要 |
| `status` | 执行状态 |
| `agent_flow` | `CHAT` / `AI_OPS` |
| `answer` | 最终答复或告警报告 |
| `self_evaluation` | rule/verifier/aiops 自评估容器 |
| `feedback` | 用户反馈 |
### agent_step
记录模型调用步骤。
用途:
- 回放 Agent 推理过程。
- 查看 Planner / Executor / Verifier 的输入输出摘要。
- 统计 step count、duration、token count。
### tool_invocation
记录工具调用事实。
用途:
- 给 Trace API 展示证据。
- 给 Verifier 构造 `tool_trace_summary`。
- 给 `EvaluationService` 计算 evidence score。
- 给 RAG eval 和人工排查提供检索细节。
## 4. 知识库模型
### api_document
MySQL 中的文档元数据表。
职责:
- 管理上传文件。
- 保存 file hash,用于去重。
- 记录索引状态和 chunk 数量。
- 保存 frontmatter JSON。
### knowledge_domain
领域级元数据。
职责:
- 按 category 聚合文档。
- 存储领域描述。
- 存储 `when_to_retrieve`,辅助 Planner/Executor 判断什么时候检索该领域。
### Milvus/Zilliz metadata
向量 collection 中每个 chunk 的 metadata 主要包括:
```text
docId
_source
chunkIndex
totalChunks
title
breadcrumb
category
```
这些字段支撑:
- category filter。
- source 展示。
- breadcrumb 上下文。
- docId 删除和重建索引。
- evidence block 构造。
## 5. 反馈沉淀模型
### case_library
`useful` 反馈会触发 `CaseLibraryService.createFromSession`。
当前自动映射:
| 字段 | 来源 |
|---|---|
| `case_id` | UUID |
| `diagnosis_id` | `diagnosis_session.session_id` |
| `source_type` | `AUTO` |
| `fault_category` | 当前默认 `GENERAL` |
| `title` | session query 前 100 字符 |
| `root_cause` | session answer |
| `solution` | session answer |
| `created_by` | `system` |
## 6. self_evaluation 结构
`diagnosis_session.self_evaluation` 是 JSON 容器:
```json
{
"rule_evaluation": {},
"verifier_evaluation": {},
"aiops_rule_evaluation": {}
}
```
边界:
- `rule_evaluation` 评估证据收集充分度。
- `verifier_evaluation` 评估 Chat 答案关键事实是否有证据支撑。
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
## 7. 数据写入时序
```mermaid
sequenceDiagram
autonumber
participant API as API
participant Svc as ChatService/AiOpsService
participant Session as diagnosis_session
participant Agent as Agent
participant Step as agent_step
participant Tool as tool_invocation
participant Eval as self_evaluation
participant Feedback as case_library
API->>Svc: request
Svc->>Session: create/update RUNNING
Agent->>Step: before/after model
Agent->>Tool: tool call record
Svc->>Session: SUCCESS/FAILED + answer
Svc->>Eval: merge evaluation
API->>Svc: feedback useful
Svc->>Feedback: create case
```
## 8. 当前边界和后续
当前边界:
- `agent_step.session_id` 和 `tool_invocation.session_id` 通过 sessionId 关联,不强制外键。
- `tool_invocation.step_id` 可为空。
- Milvus chunk 与 `api_document` 通过 metadata.docId 逻辑关联。
- `case_library` 与 session 通过 `diagnosis_id=session_id` 关联。
后续可增强:
1. 增加 run id,支持同 session 多次独立诊断。
2. 强化 `tool_invocation.step_id` 关联。
3. 将 evidence block 结构化保存。
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。