Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a1876286fd | ||
|
|
934d8eee29 | ||
|
|
7c8758d7fa | ||
|
|
b3ea6e202d | ||
|
|
8890cd2806 | ||
|
|
f4f0c63325 | ||
|
|
f02a1389c8 | ||
|
|
91931363d4 | ||
|
|
4e3502a51b | ||
|
|
b01f133efb | ||
|
|
3ed48e38cd | ||
|
|
dec587959c | ||
|
|
f002571629 | ||
|
|
553d1d1faf | ||
|
|
c88b287f83 | ||
|
|
463d8b817b | ||
|
|
363767d3e7 | ||
|
|
c4d23c3bd8 |
@@ -0,0 +1,233 @@
|
||||
# AI Ops Prompt 配置化 & LookupKnowledgeTool 集成
|
||||
|
||||
**日期**: 2026-06-24
|
||||
**类型**: 功能增强 + 架构优化
|
||||
**影响范围**: AI Ops 服务
|
||||
|
||||
---
|
||||
|
||||
## 一、变更背景
|
||||
|
||||
### 1.1 问题
|
||||
|
||||
- **硬编码 Prompt**:Planner、Executor、Supervisor 的系统提示词硬编码在 `AiOpsService.java` 中,难以维护和版本控制
|
||||
- **缺少知识库精确检索**:现有 `InternalDocsTools` 只支持 L1 语义检索(200-500ms),对于错误码、配置项等精确关键词查询效率较低
|
||||
|
||||
### 1.2 解决方案
|
||||
|
||||
1. **Prompt 配置化**:将所有 Agent 的 Prompt 抽取到 `prompts/ai-ops-prompts.yml` 配置文件
|
||||
2. **集成 L0+L1 混合检索**:引入 `LookupKnowledgeTool`,支持精确关键词匹配(< 10ms)+ 语义检索补充
|
||||
|
||||
---
|
||||
|
||||
## 二、架构变更
|
||||
|
||||
### 2.1 Prompt 配置化架构
|
||||
|
||||
```
|
||||
AiOpsService
|
||||
↓ 注入
|
||||
AiOpsPromptProperties (配置类)
|
||||
↓ @PostConstruct 加载
|
||||
ClassPathResource 读取 Markdown 文件
|
||||
↓ 读取
|
||||
prompts/
|
||||
├── planner-prompt.md
|
||||
├── executor-prompt.md
|
||||
└── supervisor-prompt.md
|
||||
```
|
||||
|
||||
**优点**:
|
||||
- 易于维护:Prompt 修改不需要重新编译
|
||||
- 格式友好:Markdown 格式支持代码块、表格,无 YAML 转义问题
|
||||
- 版本控制:配置文件独立管理
|
||||
- 易于扩展:后续可按环境区分(dev/prod)
|
||||
|
||||
### 2.2 工具层增强
|
||||
|
||||
```
|
||||
原有工具:
|
||||
- queryInternalDocs (纯 L1 语义检索,200-500ms)
|
||||
|
||||
新增工具:
|
||||
- lookup_knowledge (L0 精确匹配 + L1 补充,< 10ms 高置信度)
|
||||
```
|
||||
|
||||
**使用策略**:
|
||||
- 精确关键词(错误码、配置项)→ `lookup_knowledge`,未找到时降级到 `queryInternalDocs`
|
||||
- 模糊概念、故障流程 → 直接使用 `queryInternalDocs`
|
||||
|
||||
---
|
||||
|
||||
## 三、核心改动
|
||||
|
||||
### 3.1 新增文件
|
||||
|
||||
#### `AiOpsPromptProperties.java`
|
||||
```java
|
||||
@Configuration
|
||||
public class AiOpsPromptProperties {
|
||||
private String planner;
|
||||
private String executor;
|
||||
private String supervisor;
|
||||
|
||||
@PostConstruct
|
||||
public void loadPrompts() {
|
||||
planner = loadPromptFromFile("prompts/planner-prompt.md");
|
||||
executor = loadPromptFromFile("prompts/executor-prompt.md");
|
||||
supervisor = loadPromptFromFile("prompts/supervisor-prompt.md");
|
||||
}
|
||||
|
||||
private String loadPromptFromFile(String path) throws IOException {
|
||||
ClassPathResource resource = new ClassPathResource(path);
|
||||
return new String(resource.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### `prompts/*.md`
|
||||
三个独立的 Markdown 文件,包含 Agent 的完整系统提示词:
|
||||
- `planner-prompt.md` - Planner Agent 系统提示词
|
||||
- `executor-prompt.md` - Executor Agent 系统提示词(含工具选择指南)
|
||||
- `supervisor-prompt.md` - Supervisor Agent 系统提示词
|
||||
|
||||
### 3.2 修改文件
|
||||
|
||||
#### `AiOpsService.java`
|
||||
|
||||
**注入新组件**:
|
||||
```java
|
||||
@Autowired
|
||||
private LookupKnowledgeTool lookupKnowledgeTool;
|
||||
|
||||
@Autowired
|
||||
private AiOpsPromptProperties promptProperties;
|
||||
```
|
||||
|
||||
**使用配置化 Prompt**:
|
||||
```java
|
||||
// 原来
|
||||
.systemPrompt(buildPlannerPrompt())
|
||||
|
||||
// 改为
|
||||
.systemPrompt(promptProperties.getPlanner())
|
||||
```
|
||||
|
||||
**添加工具到工具数组**:
|
||||
```java
|
||||
return new Object[]{
|
||||
dateTimeTools,
|
||||
internalDocsTools,
|
||||
queryMetricsTools,
|
||||
lookupKnowledgeTool // 新增
|
||||
};
|
||||
```
|
||||
|
||||
**删除方法**:
|
||||
- `buildPlannerPrompt()`
|
||||
- `buildExecutorPrompt()`
|
||||
- `buildSupervisorSystemPrompt()`
|
||||
|
||||
---
|
||||
|
||||
## 四、Executor Prompt 变更详情
|
||||
|
||||
### 4.1 新增工具选择指南
|
||||
|
||||
```yaml
|
||||
- 根据查询内容选择合适的工具:
|
||||
* 精确关键词(错误码、配置项名称)→ 优先使用 lookup_knowledge,未找到时降级到 queryInternalDocs
|
||||
* 模糊概念、故障流程 → 直接使用 queryInternalDocs
|
||||
* 告警数据 → queryPrometheusAlerts
|
||||
* 日志数据 → queryLogs
|
||||
```
|
||||
|
||||
### 4.2 降级策略
|
||||
|
||||
关键改进:明确了 `lookup_knowledge` 未找到时的降级策略。
|
||||
|
||||
**流程**:
|
||||
```
|
||||
1. Planner: "查询 ERR_TIMEOUT 定义"
|
||||
2. Executor: 调用 lookup_knowledge("ERR_TIMEOUT")
|
||||
3a. 如果 found=true, confidence=high → 使用 primary.content
|
||||
3b. 如果 found=false → 自动降级到 queryInternalDocs("ERR_TIMEOUT 超时错误")
|
||||
4. 返回 feedback 给 Planner
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 五、兼容性说明
|
||||
|
||||
### 5.1 向后兼容
|
||||
|
||||
✅ **完全兼容**:
|
||||
- 现有工具调用逻辑不变
|
||||
- 3-Agent 协同模式不变
|
||||
- Planner/Executor/Supervisor 的职责边界不变
|
||||
|
||||
### 5.2 新增依赖
|
||||
|
||||
- `LookupKnowledgeTool` 依赖 `KnowledgeIndexService` 和 `VectorSearchService`
|
||||
- 需要 `knowledge_base/` 目录存在(已在 `application.yml` 中配置)
|
||||
|
||||
---
|
||||
|
||||
## 六、验证清单
|
||||
|
||||
### 6.1 编译验证
|
||||
|
||||
```bash
|
||||
mvn clean compile -DskipTests
|
||||
```
|
||||
|
||||
✅ **结果**: BUILD SUCCESS
|
||||
|
||||
### 6.2 运行时验证(待完成)
|
||||
|
||||
- [ ] 启动应用,验证 Prompt 配置加载成功
|
||||
- [ ] 触发 AI Ops 流程,验证 `lookup_knowledge` 工具可调用
|
||||
- [ ] 测试精确关键词查询(如 "ERR_TIMEOUT")
|
||||
- [ ] 测试降级策略(查询不存在的关键词)
|
||||
|
||||
---
|
||||
|
||||
## 七、后续工作
|
||||
|
||||
### 7.1 知识库内容准备
|
||||
|
||||
当前 `knowledge_base/` 目录需要补充文档:
|
||||
- 错误码定义(支付网关、订单系统等)
|
||||
- 配置最佳实践(Redis、HikariCP、Flyway 等)
|
||||
- 故障排查流程
|
||||
|
||||
**文档格式示例**:
|
||||
```markdown
|
||||
---
|
||||
title: 支付网关错误码定义
|
||||
keywords: [ERR_TIMEOUT, 超时, 支付网关]
|
||||
summary: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
category: api
|
||||
---
|
||||
|
||||
# 支付网关错误码定义
|
||||
|
||||
## ERR_TIMEOUT
|
||||
...
|
||||
```
|
||||
|
||||
### 7.2 Prompt 优化
|
||||
|
||||
基于实际运行反馈,持续优化 `prompts/ai-ops-prompts.yml` 中的提示词。
|
||||
|
||||
### 7.3 可观测性增强
|
||||
|
||||
- 监控 `lookup_knowledge` 的调用频率和命中率
|
||||
- 记录降级场景(L0 未找到 → L1 补充)
|
||||
|
||||
---
|
||||
|
||||
## 八、参考文档
|
||||
|
||||
- [知识库检索架构说明](../mvp/architecture/knowledge-retrieval-architecture.md)
|
||||
- [AI Ops 核心设计 Essence 报告](../docs/learning/01-AI-Ops-核心设计-Essence报告.md)
|
||||
@@ -0,0 +1,100 @@
|
||||
# Prompt 配置化改进总结
|
||||
|
||||
**日期**: 2026-06-24
|
||||
**改进**: 从 YAML 配置改为 Markdown 文件
|
||||
|
||||
---
|
||||
|
||||
## 改进原因
|
||||
|
||||
YAML 格式存在以下问题:
|
||||
1. **多行字符串缩进敏感**:容易出现格式错误
|
||||
2. **转义字符复杂**:代码块、表格需要转义处理
|
||||
3. **可读性差**:长文本在 YAML 中难以阅读和维护
|
||||
|
||||
Markdown 格式优势:
|
||||
- ✅ 原生支持代码块、表格、列表
|
||||
- ✅ 无需转义,所见即所得
|
||||
- ✅ 版本控制 diff 更清晰
|
||||
- ✅ 编辑器语法高亮支持好
|
||||
|
||||
---
|
||||
|
||||
## 最终方案
|
||||
|
||||
### 文件结构
|
||||
```
|
||||
src/main/resources/prompts/
|
||||
├── planner-prompt.md # Planner Agent 系统提示词
|
||||
├── executor-prompt.md # Executor Agent 系统提示词
|
||||
└── supervisor-prompt.md # Supervisor Agent 系统提示词
|
||||
```
|
||||
|
||||
### 加载方式
|
||||
```java
|
||||
@Configuration
|
||||
public class AiOpsPromptProperties {
|
||||
|
||||
@PostConstruct
|
||||
public void loadPrompts() {
|
||||
planner = loadPromptFromFile("prompts/planner-prompt.md");
|
||||
executor = loadPromptFromFile("prompts/executor-prompt.md");
|
||||
supervisor = loadPromptFromFile("prompts/supervisor-prompt.md");
|
||||
}
|
||||
|
||||
private String loadPromptFromFile(String path) throws IOException {
|
||||
ClassPathResource resource = new ClassPathResource(path);
|
||||
return new String(resource.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 使用方式
|
||||
```java
|
||||
@Autowired
|
||||
private AiOpsPromptProperties promptProperties;
|
||||
|
||||
// 直接使用
|
||||
.systemPrompt(promptProperties.getPlanner())
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 编译验证
|
||||
|
||||
```bash
|
||||
mvn clean compile -DskipTests
|
||||
```
|
||||
|
||||
✅ **结果**: BUILD SUCCESS
|
||||
|
||||
---
|
||||
|
||||
## 完整改动清单
|
||||
|
||||
| 文件 | 改动 |
|
||||
|------|------|
|
||||
| `AiOpsService.java` | 注入 `LookupKnowledgeTool` + `AiOpsPromptProperties` |
|
||||
| `AiOpsPromptProperties.java` | 从 Markdown 文件加载 Prompt(使用 `@PostConstruct`)|
|
||||
| `prompts/planner-prompt.md` | 新增:Planner 系统提示词 |
|
||||
| `prompts/executor-prompt.md` | 新增:Executor 系统提示词(含工具选择指南)|
|
||||
| `prompts/supervisor-prompt.md` | 新增:Supervisor 系统提示词 |
|
||||
| ~~`YamlPropertySourceFactory.java`~~ | 已删除(不再需要)|
|
||||
| ~~`prompts/ai-ops-prompts.yml`~~ | 已删除(改用 Markdown)|
|
||||
|
||||
---
|
||||
|
||||
## Executor Prompt 关键改进
|
||||
|
||||
新增工具选择指南:
|
||||
```markdown
|
||||
- 根据查询内容选择合适的工具:
|
||||
* 精确关键词(错误码、配置项名称)→ 优先使用 lookup_knowledge,未找到时降级到 queryInternalDocs
|
||||
* 模糊概念、故障流程 → 直接使用 queryInternalDocs
|
||||
* 告警数据 → queryPrometheusAlerts
|
||||
* 日志数据 → queryLogs
|
||||
```
|
||||
|
||||
降级策略:
|
||||
- `lookup_knowledge` 未找到 → 自动降级到 `queryInternalDocs`
|
||||
- 确保查询不会因为知识库缺少内容而失败
|
||||
@@ -0,0 +1,469 @@
|
||||
# 知识库初始化 API 使用文档
|
||||
|
||||
## 概述
|
||||
|
||||
提供了知识库批量初始化接口,用于将 `knowledge_base` 目录下的所有 Markdown 文档导入到数据库和向量索引(L0 + L1)。
|
||||
|
||||
**功能特点**:
|
||||
1. ✅ **批量扫描**:递归扫描 knowledge_base 目录下所有 .md 文件
|
||||
2. ✅ **自动去重**:基于文件路径检查,避免重复导入
|
||||
3. ✅ **数据入库**:保存文档元数据到 MySQL
|
||||
4. ✅ **L0 索引**:自动加入内存精确匹配索引
|
||||
5. ✅ **L1 索引**:文档分块并上传到 Milvus 向量数据库
|
||||
|
||||
---
|
||||
|
||||
## API 接口
|
||||
|
||||
### 1. 初始化知识库
|
||||
|
||||
**端点**:
|
||||
```
|
||||
POST /api/knowledge/init?force=false
|
||||
```
|
||||
|
||||
**参数**:
|
||||
- `force`(可选):是否强制重新导入,跳过去重检查
|
||||
- `false`(默认):跳过已存在的文档
|
||||
- `true`:强制重新导入所有文档
|
||||
|
||||
**请求示例**:
|
||||
```bash
|
||||
# 首次导入(去重模式)
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
|
||||
# 强制重新导入
|
||||
curl -X POST http://localhost:9900/api/knowledge/init?force=true
|
||||
```
|
||||
|
||||
**响应示例**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"message": "知识库初始化完成",
|
||||
"scanned": 6,
|
||||
"skipped": 0,
|
||||
"inserted": 6,
|
||||
"failed": 0,
|
||||
"details": {
|
||||
"api/payment-errors.md": "导入成功(L0+L1)",
|
||||
"domain/spring-ai-tool-best-practices.md": "导入成功(L0+L1)",
|
||||
"infrastructure/flyway-best-practices.md": "导入成功(L0+L1)",
|
||||
"infrastructure/mysql-connection-pool.md": "导入成功(L0+L1)",
|
||||
"infrastructure/redis-config.md": "导入成功(L0+L1)",
|
||||
"troubleshooting/fault-diagnosis-process.md": "导入成功(L0+L1)"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**字段说明**:
|
||||
- `scanned`:扫描到的文件总数
|
||||
- `skipped`:跳过的文件数量(已存在)
|
||||
- `inserted`:成功导入的文件数量
|
||||
- `failed`:失败的文件数量
|
||||
- `details`:每个文件的处理结果详情
|
||||
|
||||
---
|
||||
|
||||
### 2. 查询知识库统计
|
||||
|
||||
**端点**:
|
||||
```
|
||||
GET /api/knowledge/stats
|
||||
```
|
||||
|
||||
**请求示例**:
|
||||
```bash
|
||||
curl http://localhost:9900/api/knowledge/stats
|
||||
```
|
||||
|
||||
**响应示例**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"totalDocuments": 6,
|
||||
"totalVectors": 48,
|
||||
"categories": {
|
||||
"api": 1,
|
||||
"domain": 1,
|
||||
"infrastructure": 3,
|
||||
"troubleshooting": 1
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**字段说明**:
|
||||
- `totalDocuments`:数据库中的文档总数
|
||||
- `totalVectors`:Milvus 中的向量总数(chunk 数量)
|
||||
- `categories`:按分类统计的文档数量
|
||||
|
||||
---
|
||||
|
||||
## 使用场景
|
||||
|
||||
### 场景 1:项目启动时初始化
|
||||
|
||||
```bash
|
||||
# 1. 启动应用
|
||||
mvn spring-boot:run
|
||||
|
||||
# 2. 等待应用启动完成(约 10 秒)
|
||||
|
||||
# 3. 调用初始化接口
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
|
||||
# 4. 查看结果
|
||||
# 日志输出:知识库初始化完成: 扫描=6, 跳过=0, 新增=6, 失败=0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 场景 2:添加新文档后重新初始化
|
||||
|
||||
```bash
|
||||
# 1. 添加新文档到 knowledge_base 目录
|
||||
echo "---
|
||||
title: 新文档
|
||||
keywords: [测试, test]
|
||||
summary: 这是一个测试文档
|
||||
category: test
|
||||
---
|
||||
|
||||
# 新文档内容
|
||||
" > knowledge_base/test/new-doc.md
|
||||
|
||||
# 2. 调用初始化接口(去重模式)
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
|
||||
# 3. 查看结果
|
||||
# 只会导入新文档,跳过已存在的 6 个文档
|
||||
# 响应: scanned=7, skipped=6, inserted=1, failed=0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 场景 3:强制重新导入所有文档
|
||||
|
||||
```bash
|
||||
# 适用场景:
|
||||
# - 数据库被清空,需要重新导入
|
||||
# - 文档内容有更新,需要刷新
|
||||
# - 索引损坏,需要重建
|
||||
|
||||
curl -X POST http://localhost:9900/api/knowledge/init?force=true
|
||||
|
||||
# 响应: scanned=6, skipped=0, inserted=6, failed=0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 去重机制
|
||||
|
||||
### 去重依据
|
||||
- **文件路径**:相对于 `knowledge_base` 目录的相对路径
|
||||
- 示例:`api/payment-errors.md`
|
||||
|
||||
### 去重逻辑
|
||||
```
|
||||
if (!force && existingFilePaths.contains(relativePath)) {
|
||||
跳过该文档
|
||||
} else {
|
||||
导入该文档
|
||||
}
|
||||
```
|
||||
|
||||
### 注意事项
|
||||
1. **文件移动会被视为新文档**:
|
||||
```bash
|
||||
# 移动前:api/payment-errors.md
|
||||
# 移动后:errors/payment-errors.md
|
||||
# 结果:会被当作两个不同的文档
|
||||
```
|
||||
|
||||
2. **文件重命名会被视为新文档**:
|
||||
```bash
|
||||
# 重命名前:payment-errors.md
|
||||
# 重命名后:payment-error-codes.md
|
||||
# 结果:会被当作两个不同的文档
|
||||
```
|
||||
|
||||
3. **内容更新不触发重新导入**(非 force 模式):
|
||||
```bash
|
||||
# 修改文件内容后调用 init(非 force)
|
||||
# 结果:跳过该文档,数据库中仍是旧内容
|
||||
# 解决:使用 force=true 强制重新导入
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 数据存储
|
||||
|
||||
### 完整的数据流
|
||||
|
||||
```
|
||||
knowledge_base/*.md
|
||||
↓ 1. 扫描
|
||||
KnowledgeBaseInitService
|
||||
↓ 2. 解析 frontmatter
|
||||
Frontmatter (title, keywords, summary)
|
||||
↓ 3. 保存到数据库
|
||||
MySQL (api_document)
|
||||
↓ 4. 提取正文 & 分块
|
||||
DocumentChunkService
|
||||
↓ 5. 生成向量
|
||||
VectorEmbeddingService
|
||||
↓ 6. 索引到 Milvus
|
||||
Milvus (L1 向量索引)
|
||||
↓ 7. 加入内存索引
|
||||
KnowledgeIndexService (L0)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 数据库表结构(api_document)
|
||||
|
||||
| 字段 | 类型 | 说明 | 示例 |
|
||||
|------|------|------|------|
|
||||
| `id` | BIGINT | 主键 | 1 |
|
||||
| `doc_id` | VARCHAR(64) | 文档唯一标识 | uuid |
|
||||
| `file_name` | VARCHAR(256) | 文件名 | payment-errors.md |
|
||||
| `file_path` | VARCHAR(512) | 相对路径 | api/payment-errors.md |
|
||||
| `api_name` | VARCHAR(128) | 文档标题 | 支付网关错误码定义 |
|
||||
| `status` | VARCHAR(16) | 状态 | INDEXED / FAILED |
|
||||
| `chunk_count` | INT | 分块数量 | 8 |
|
||||
| `error_message` | TEXT | 错误信息 | null |
|
||||
| `metadata` | TEXT | Frontmatter JSON | {"title":"...","keywords":[...]} |
|
||||
| `file_size` | BIGINT | 文件大小(字节) | 2048 |
|
||||
| `indexed_at` | DATETIME | 索引时间 | 2026-06-25 10:00:00 |
|
||||
|
||||
### metadata JSON 结构
|
||||
|
||||
```json
|
||||
{
|
||||
"title": "支付网关错误码定义",
|
||||
"summary": "记录了支付网关所有核心错误码的含义及排查方向",
|
||||
"category": "api",
|
||||
"keywords": ["ERR_TIMEOUT","超时","支付网关"]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Milvus 向量索引
|
||||
|
||||
每个文档会被分块(chunk)并生成向量,存储到 Milvus 集合中:
|
||||
|
||||
**Collection**: `knowledge_base_collection`
|
||||
|
||||
**字段**:
|
||||
- `doc_id`:文档 ID
|
||||
- `chunk_id`:分块 ID
|
||||
- `chunk_text`:分块文本内容
|
||||
- `embedding`:768 维向量
|
||||
- `category`:文档分类
|
||||
- `file_path`:文件路径
|
||||
|
||||
**分块策略**:
|
||||
- Chunk Size:根据 `DocumentChunkConfig` 配置(默认 500 token)
|
||||
- Overlap:重叠区域(默认 50 token)
|
||||
|
||||
---
|
||||
|
||||
## L0 内存索引
|
||||
|
||||
导入过程会自动将文档加入 `KnowledgeIndexService` 的内存索引:
|
||||
|
||||
```java
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath(relativePath)
|
||||
.title(title)
|
||||
.keywords(keywords)
|
||||
.summary(summary)
|
||||
.category(category)
|
||||
.build();
|
||||
knowledgeIndexService.addToIndex(entry);
|
||||
```
|
||||
|
||||
**验证 L0 索引**:
|
||||
```bash
|
||||
# 应用启动后查看日志
|
||||
grep "知识库索引加载完成" logs/application.log
|
||||
|
||||
# 输出示例:
|
||||
# [INFO] 知识库索引加载完成,共 6 个文档
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 错误处理
|
||||
|
||||
### 常见错误
|
||||
|
||||
#### 1. 目录不存在
|
||||
```json
|
||||
{
|
||||
"success": false,
|
||||
"message": "初始化失败: 知识库目录不存在: knowledge_base"
|
||||
}
|
||||
```
|
||||
|
||||
**解决**:
|
||||
```bash
|
||||
mkdir -p knowledge_base/api
|
||||
mkdir -p knowledge_base/infrastructure
|
||||
mkdir -p knowledge_base/domain
|
||||
mkdir -p knowledge_base/troubleshooting
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### 2. 文档格式无效
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"scanned": 6,
|
||||
"inserted": 5,
|
||||
"failed": 1,
|
||||
"details": {
|
||||
"test/invalid.md": "格式无效: frontmatter 解析失败"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**原因**:
|
||||
- 缺少 frontmatter
|
||||
- YAML 格式错误
|
||||
- 缺少必填字段(title, keywords, summary)
|
||||
|
||||
**解决**:
|
||||
```markdown
|
||||
---
|
||||
title: 文档标题
|
||||
keywords: [关键词1, 关键词2]
|
||||
summary: 文档摘要
|
||||
category: api
|
||||
---
|
||||
|
||||
# 正文内容
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 问题 4: Milvus 连接失败
|
||||
|
||||
**症状**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"scanned": 6,
|
||||
"inserted": 0,
|
||||
"failed": 6,
|
||||
"details": {
|
||||
"api/payment-errors.md": "Milvus 索引失败: Connection refused"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**原因**:
|
||||
- Milvus 服务未启动
|
||||
- 网络连接问题
|
||||
- 配置错误
|
||||
|
||||
**解决**:
|
||||
```bash
|
||||
# 检查 Milvus 是否运行
|
||||
docker ps | grep milvus
|
||||
|
||||
# 检查配置
|
||||
grep milvus application.yml
|
||||
|
||||
# 启动 Milvus
|
||||
docker-compose up -d milvus-standalone
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 问题 5: 文档分块失败
|
||||
|
||||
**症状**:
|
||||
```json
|
||||
{
|
||||
"details": {
|
||||
"test/large-doc.md": "Milvus 索引失败: Document too large"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**原因**:
|
||||
- 文档内容过大
|
||||
- 分块配置不当
|
||||
|
||||
**解决**:
|
||||
- 检查 `DocumentChunkConfig` 配置
|
||||
- 调整 chunk size 和 overlap
|
||||
|
||||
---
|
||||
|
||||
#### 3. 文档缺少标题
|
||||
```json
|
||||
{
|
||||
"details": {
|
||||
"test/no-title.md": "缺少标题"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**解决**:在 frontmatter 中添加 `title` 字段。
|
||||
|
||||
---
|
||||
|
||||
## 最佳实践
|
||||
|
||||
### ✅ 推荐做法
|
||||
|
||||
1. **首次启动后立即初始化**:
|
||||
```bash
|
||||
mvn spring-boot:run
|
||||
sleep 15 # 等待启动完成
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
```
|
||||
|
||||
2. **新增文档后增量导入**:
|
||||
```bash
|
||||
# 不使用 force,只导入新文档
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
```
|
||||
|
||||
3. **定期检查统计信息**:
|
||||
```bash
|
||||
curl http://localhost:9900/api/knowledge/stats
|
||||
```
|
||||
|
||||
4. **更新文档内容后强制刷新**:
|
||||
```bash
|
||||
curl -X POST http://localhost:9900/api/knowledge/init?force=true
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### ❌ 避免做法
|
||||
|
||||
1. **不检查响应就认为成功**:
|
||||
- 始终检查 `failed` 字段
|
||||
- 查看 `details` 了解具体失败原因
|
||||
|
||||
2. **频繁使用 force=true**:
|
||||
- 会重复插入数据(违反唯一约束)
|
||||
- 建议先清理数据库,再使用 force
|
||||
|
||||
3. **不检查文档格式就导入**:
|
||||
- 先手动验证 frontmatter 格式
|
||||
- 确保必填字段完整
|
||||
|
||||
---
|
||||
|
||||
## 相关文档
|
||||
|
||||
- **知识库使用指南**:`mvp/architecture/knowledge-retrieval-usage.md`
|
||||
- **知识库架构**:`mvp/architecture/knowledge-retrieval-architecture.md`
|
||||
- **Executor Prompt**:`src/main/resources/prompts/executor-prompt.md`
|
||||
@@ -0,0 +1,79 @@
|
||||
# 执行者 System Prompt
|
||||
|
||||
## 角色定位
|
||||
|
||||
你是诊断流程的**执行者**。你的任务非常明确:严格遵循规划者下发的任务清单,按步骤调用工具完成任务,并输出最终结果。
|
||||
|
||||
---
|
||||
|
||||
## 核心行为准则
|
||||
|
||||
### 1. 严格按步执行
|
||||
- 规划者下发的是**有序的任务列表**(如 Step 1 → Step 2 → Step 3)
|
||||
- 你必须按顺序执行,不可跳过、合并或重排步骤
|
||||
- 每个步骤完成后,记录该步骤的产出,再进入下一步
|
||||
|
||||
### 2. 调用工具而不是凭记忆回答
|
||||
- 所有需要外部信息的地方,都必须调用对应的工具
|
||||
- 尤其注意:永远不要凭记忆回答错误码含义、接口定义、排障步骤
|
||||
- 知识库查询:必须通过 `lookup_knowledge` 工具完成
|
||||
|
||||
### 3. 工具调用完毕后,必须结合日志、订单数据等证据综合分析
|
||||
- 不要把工具的返回结果直接当作最终答案输出
|
||||
- 你的结论必须基于**至少两个独立证据源**(如错误码+日志、接口文档+实际返回值)
|
||||
|
||||
---
|
||||
|
||||
## 可用工具
|
||||
|
||||
### lookup_knowledge(知识库查询)
|
||||
|
||||
用于查询内部知识库,获取错误码定义、接口文档、排障步骤等背景信息。
|
||||
|
||||
| 参数 | 说明 |
|
||||
|------|------|
|
||||
| `query_text` | 查询关键词。可以是错误码(ERR_TIMEOUT)、服务名(payment-gateway)、模糊问题(支付为什么失败) |
|
||||
|
||||
**内部机制**:
|
||||
工具内部自动执行「先精确匹配(L0),未命中则语义检索(L1)」的两阶段检索逻辑,你无需关心哪一层。返回结果中包含 `match_type` 字段标记来源类型。
|
||||
|
||||
**返回字段**:
|
||||
- `primary`:主要信息(L0 命中文档内容 或 L1 返回的 Top-1 片段)
|
||||
- `primary.match_type`:`exact_l0`(精确匹配)或 `semantic_l1`(语义搜索)
|
||||
- `primary.source`:信息来源的文件路径
|
||||
|
||||
**使用规则**:
|
||||
- 当你查到了错误码、接口名、服务名时:**必须**调用此工具
|
||||
- 当需要查排障步骤、业务流程、最佳实践时:**必须**调用此工具
|
||||
- 对当前结果没有十足把握时:**建议**调用此工具验证
|
||||
|
||||
---
|
||||
|
||||
## 任务执行规范
|
||||
|
||||
### 1. 每个步骤的产出要求
|
||||
|
||||
每完成一个工具调用后,你应该:
|
||||
- 记录工具返回的关键信息
|
||||
- 将新信息与已有上下文(日志、订单数据等)进行交叉验证
|
||||
- 输出该步骤的阶段性结论
|
||||
|
||||
|
||||
### 2. 最终输出的报告格式
|
||||
|
||||
```yaml
|
||||
## 诊断结论
|
||||
|
||||
**问题根因**:XXX
|
||||
|
||||
**证据链**:
|
||||
1. 订单状态返回错误码 ERR_TIMEOUT
|
||||
2. 知识库 lookup_knowledge("ERR_TIMEOUT") 返回:支付网关响应超时(>5秒)
|
||||
3. 日志确认:14:32:15 请求耗时 5.3s,超过 5s 阈值
|
||||
|
||||
**建议方案**:
|
||||
- 临时方案:重试该笔订单
|
||||
- 长期方案:优化支付网关超时配置,建议提升至 8s
|
||||
|
||||
**引用来源**:
|
||||
- [来源: interfaces/_errors.md]
|
||||
@@ -15,7 +15,12 @@ import java.util.List;
|
||||
/**
|
||||
* 内部文档查询工具
|
||||
* 使用 RAG (Retrieval-Augmented Generation) 从内部知识库检索相关文档
|
||||
*
|
||||
* @deprecated 请使用 {@link com.superbiz.agent.tool.LookupKnowledgeTool} 替代。
|
||||
* lookup_knowledge 支持 L0 精确匹配 + L1 语义检索,性能更优且功能更全面。
|
||||
* 计划在下一个版本中移除此工具。
|
||||
*/
|
||||
@Deprecated
|
||||
@Component
|
||||
public class InternalDocsTools {
|
||||
|
||||
@@ -45,7 +50,9 @@ public class InternalDocsTools {
|
||||
*
|
||||
* @param query 搜索查询,描述您要查找的信息
|
||||
* @return JSON 格式的搜索结果,包含相关文档内容、相似度分数和元数据
|
||||
* @deprecated 请使用 {@link com.superbiz.agent.tool.LookupKnowledgeTool#lookupKnowledge(String)} 替代
|
||||
*/
|
||||
@Deprecated
|
||||
@Tool(description = "Use this tool to search internal documentation and knowledge base for relevant information. " +
|
||||
"It performs RAG (Retrieval-Augmented Generation) to find similar documents and extract processing steps. " +
|
||||
"This is useful when you need to understand internal procedures, best practices, or step-by-step guides " +
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
package com.superbiz.agent.config;
|
||||
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.context.annotation.Configuration;
|
||||
import org.springframework.core.io.ClassPathResource;
|
||||
|
||||
import jakarta.annotation.PostConstruct;
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
|
||||
/**
|
||||
* AI Ops Agent Prompt 配置
|
||||
* 从独立的 Markdown 文件加载 Prompt 模板
|
||||
*/
|
||||
@Slf4j
|
||||
@Configuration
|
||||
public class AiOpsPromptProperties {
|
||||
|
||||
private String planner;
|
||||
private String executor;
|
||||
private String supervisor;
|
||||
|
||||
@PostConstruct
|
||||
public void loadPrompts() {
|
||||
try {
|
||||
planner = loadPromptFromFile("prompts/planner-prompt.md");
|
||||
executor = loadPromptFromFile("prompts/executor-prompt.md");
|
||||
supervisor = loadPromptFromFile("prompts/supervisor-prompt.md");
|
||||
|
||||
log.info("AI Ops Prompts 加载成功");
|
||||
log.debug("Planner Prompt 长度: {} 字符", planner.length());
|
||||
log.debug("Executor Prompt 长度: {} 字符", executor.length());
|
||||
log.debug("Supervisor Prompt 长度: {} 字符", supervisor.length());
|
||||
} catch (IOException e) {
|
||||
log.error("加载 Prompt 文件失败", e);
|
||||
throw new RuntimeException("Failed to load AI Ops prompts", e);
|
||||
}
|
||||
}
|
||||
|
||||
private String loadPromptFromFile(String path) throws IOException {
|
||||
ClassPathResource resource = new ClassPathResource(path);
|
||||
return new String(resource.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
|
||||
}
|
||||
|
||||
public String getPlanner() {
|
||||
return planner;
|
||||
}
|
||||
|
||||
public String getExecutor() {
|
||||
return executor;
|
||||
}
|
||||
|
||||
public String getSupervisor() {
|
||||
return supervisor;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,94 @@
|
||||
package com.superbiz.agent.controller;
|
||||
|
||||
import com.superbiz.agent.service.KnowledgeBaseInitService;
|
||||
import lombok.Data;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.http.ResponseEntity;
|
||||
import org.springframework.web.bind.annotation.*;
|
||||
|
||||
import java.util.HashMap;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* 知识库管理控制器
|
||||
* 提供知识库初始化、查询等接口
|
||||
*/
|
||||
@RestController
|
||||
@RequestMapping("/api/knowledge")
|
||||
public class KnowledgeBaseController {
|
||||
|
||||
private static final Logger logger = LoggerFactory.getLogger(KnowledgeBaseController.class);
|
||||
|
||||
@Autowired
|
||||
private KnowledgeBaseInitService initService;
|
||||
|
||||
/**
|
||||
* 初始化知识库
|
||||
* 扫描 knowledge_base 目录下的所有文档,去重后批量导入到数据库和 Milvus
|
||||
*
|
||||
* @param force 是否强制重新导入(跳过去重检查)
|
||||
* @return 初始化结果
|
||||
*/
|
||||
@PostMapping("/init")
|
||||
public ResponseEntity<?> initKnowledgeBase(@RequestParam(defaultValue = "false") boolean force) {
|
||||
logger.info("收到知识库初始化请求, force={}", force);
|
||||
|
||||
try {
|
||||
KnowledgeBaseInitService.InitResult result = initService.initializeKnowledgeBase(force);
|
||||
|
||||
Map<String, Object> response = new HashMap<>();
|
||||
response.put("success", true);
|
||||
response.put("message", "知识库初始化完成");
|
||||
response.put("scanned", result.getScanned());
|
||||
response.put("skipped", result.getSkipped());
|
||||
response.put("inserted", result.getInserted());
|
||||
response.put("failed", result.getFailed());
|
||||
response.put("details", result.getDetails());
|
||||
|
||||
logger.info("知识库初始化成功: 扫描={}, 跳过={}, 新增={}, 失败={}",
|
||||
result.getScanned(), result.getSkipped(), result.getInserted(), result.getFailed());
|
||||
|
||||
return ResponseEntity.ok(response);
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("知识库初始化失败", e);
|
||||
|
||||
Map<String, Object> response = new HashMap<>();
|
||||
response.put("success", false);
|
||||
response.put("message", "初始化失败: " + e.getMessage());
|
||||
|
||||
return ResponseEntity.internalServerError().body(response);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 查询知识库统计信息
|
||||
*
|
||||
* @return 统计信息
|
||||
*/
|
||||
@GetMapping("/stats")
|
||||
public ResponseEntity<?> getStats() {
|
||||
try {
|
||||
KnowledgeBaseInitService.Stats stats = initService.getStats();
|
||||
|
||||
Map<String, Object> response = new HashMap<>();
|
||||
response.put("success", true);
|
||||
response.put("totalDocuments", stats.getTotalDocuments());
|
||||
response.put("totalVectors", stats.getTotalVectors());
|
||||
response.put("categories", stats.getCategoryCount());
|
||||
|
||||
return ResponseEntity.ok(response);
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("查询统计信息失败", e);
|
||||
|
||||
Map<String, Object> response = new HashMap<>();
|
||||
response.put("success", false);
|
||||
response.put("message", "查询失败: " + e.getMessage());
|
||||
|
||||
return ResponseEntity.internalServerError().body(response);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -38,7 +38,7 @@ public class ApiDocument {
|
||||
// 文档分类
|
||||
@Enumerated(EnumType.STRING)
|
||||
@Column(name = "fault_category", length = 32, columnDefinition = "VARCHAR(32)")
|
||||
private FaultCategory faultCategory = FaultCategory.EXTERNAL_API;
|
||||
private FaultCategory faultCategory = FaultCategory.GENERAL;
|
||||
|
||||
@Column(name = "fault_source", length = 128)
|
||||
private String faultSource;
|
||||
|
||||
@@ -1,17 +1,14 @@
|
||||
package com.superbiz.agent.domain.enums;
|
||||
|
||||
/**
|
||||
* 故障类别枚举
|
||||
* 文档分类枚举
|
||||
*/
|
||||
public enum FaultCategory {
|
||||
EXTERNAL_API("外部接口调用失败"),
|
||||
INTERNAL_ERROR("系统内部错误"),
|
||||
DATABASE("数据库问题"),
|
||||
CACHE("缓存问题"),
|
||||
NETWORK("网络问题"),
|
||||
THREAD("线程问题"),
|
||||
MEMORY("内存问题"),
|
||||
CONFIG("配置问题");
|
||||
API("API 接口文档"),
|
||||
INFRASTRUCTURE("基础设施文档"),
|
||||
DOMAIN("领域业务文档"),
|
||||
TROUBLESHOOTING("故障排查文档"),
|
||||
GENERAL("通用文档");
|
||||
|
||||
private final String description;
|
||||
|
||||
@@ -22,4 +19,26 @@ public enum FaultCategory {
|
||||
public String getDescription() {
|
||||
return description;
|
||||
}
|
||||
|
||||
/**
|
||||
* 从字符串映射到枚举
|
||||
*/
|
||||
public static FaultCategory fromString(String category) {
|
||||
if (category == null || category.isEmpty()) {
|
||||
return GENERAL;
|
||||
}
|
||||
|
||||
switch (category.toLowerCase()) {
|
||||
case "api":
|
||||
return API;
|
||||
case "infrastructure":
|
||||
return INFRASTRUCTURE;
|
||||
case "domain":
|
||||
return DOMAIN;
|
||||
case "troubleshooting":
|
||||
return TROUBLESHOOTING;
|
||||
default:
|
||||
return GENERAL;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,208 @@
|
||||
package com.superbiz.agent.hook;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPosition;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPositions;
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.ai.chat.messages.Message;
|
||||
import org.springframework.ai.chat.messages.AssistantMessage;
|
||||
import org.springframework.ai.chat.messages.UserMessage;
|
||||
import org.springframework.ai.chat.messages.ToolResponseMessage;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Agent 日志 Hook
|
||||
* 用于记录 Agent 的思考过程、消息流转
|
||||
*/
|
||||
@Slf4j
|
||||
@HookPositions({HookPosition.BEFORE_MODEL, HookPosition.AFTER_MODEL})
|
||||
public class AgentLoggingHook extends MessagesModelHook {
|
||||
|
||||
private int modelCallCount = 0;
|
||||
|
||||
@Override
|
||||
public String getName() {
|
||||
return "agent_logging_hook";
|
||||
}
|
||||
|
||||
@Override
|
||||
public AgentCommand beforeModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
modelCallCount++;
|
||||
log.info("========================================");
|
||||
log.info("*** [Agent 思考] 第 {} 轮思考开始", modelCallCount);
|
||||
log.info("*** [Agent 思考] 当前消息数量: {}", previousMessages.size());
|
||||
|
||||
// 打印最后几条消息
|
||||
int lastN = Math.min(3, previousMessages.size());
|
||||
if (lastN > 0) {
|
||||
log.info("*** [Agent 思考] 最近 {} 条消息:", lastN);
|
||||
List<Message> recentMessages = previousMessages.subList(previousMessages.size() - lastN, previousMessages.size());
|
||||
|
||||
for (int i = 0; i < recentMessages.size(); i++) {
|
||||
Message msg = recentMessages.get(i);
|
||||
String role = getMessageRole(msg);
|
||||
|
||||
log.info(" [{}] 角色: {}, 类型: {}", i + 1, role, msg.getClass().getSimpleName());
|
||||
// Message 接口可能没有直接的 getContent() 方法,跳过内容打印
|
||||
// 具体内容会在工具调用日志中体现
|
||||
}
|
||||
}
|
||||
|
||||
log.info("*** [Agent 思考] 准备调用模型...");
|
||||
log.info("========================================");
|
||||
|
||||
// 不修改消息,直接返回
|
||||
return new AgentCommand(previousMessages);
|
||||
}
|
||||
|
||||
@Override
|
||||
public AgentCommand afterModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
log.info("========================================");
|
||||
log.info("*** [Agent 思考] 第 {} 轮思考完成", modelCallCount);
|
||||
|
||||
// 查找最后一条 AssistantMessage(模型的回复)
|
||||
AssistantMessage lastAssistant = null;
|
||||
for (int i = previousMessages.size() - 1; i >= 0; i--) {
|
||||
if (previousMessages.get(i) instanceof AssistantMessage) {
|
||||
lastAssistant = (AssistantMessage) previousMessages.get(i);
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
if (lastAssistant != null) {
|
||||
// 打印模型返回的文本内容
|
||||
String textContent = extractTextContent(lastAssistant);
|
||||
if (textContent != null && !textContent.isEmpty()) {
|
||||
log.info("*** [Agent 思考] 模型返回文本: {}",
|
||||
textContent.length() > 500
|
||||
? textContent.substring(0, 500) + "... (已截断,总长度: " + textContent.length() + ")"
|
||||
: textContent);
|
||||
}
|
||||
|
||||
// 检查是否有工具调用
|
||||
if (lastAssistant.getToolCalls() != null && !lastAssistant.getToolCalls().isEmpty()) {
|
||||
log.info("*** [Agent 思考] 模型决定调用 {} 个工具:",
|
||||
lastAssistant.getToolCalls().size());
|
||||
lastAssistant.getToolCalls().forEach(toolCall -> {
|
||||
log.info(" - 工具: {}, 参数: {}",
|
||||
toolCall.name(),
|
||||
toolCall.arguments());
|
||||
});
|
||||
log.info("*** [Agent 思考] 等待工具执行结果...");
|
||||
} else {
|
||||
log.info("*** [Agent 思考] 模型决定不调用工具");
|
||||
log.info("*** [Agent 思考] 这是最终答案,准备返回给用户");
|
||||
}
|
||||
}
|
||||
|
||||
log.info("========================================");
|
||||
|
||||
// 不修改消息,直接返回
|
||||
return new AgentCommand(previousMessages);
|
||||
}
|
||||
|
||||
/**
|
||||
* 提取 AssistantMessage 的文本内容
|
||||
*/
|
||||
private String extractTextContent(AssistantMessage message) {
|
||||
try {
|
||||
// 方法 1: 尝试通过反射获取 text 字段
|
||||
try {
|
||||
java.lang.reflect.Field textField = message.getClass().getDeclaredField("text");
|
||||
textField.setAccessible(true);
|
||||
Object value = textField.get(message);
|
||||
if (value != null) {
|
||||
String text = value.toString();
|
||||
log.debug("通过 text 字段提取成功");
|
||||
return text;
|
||||
}
|
||||
} catch (NoSuchFieldException e) {
|
||||
// text 字段不存在,尝试下一种方法
|
||||
}
|
||||
|
||||
// 方法 2: 尝试 content 字段
|
||||
try {
|
||||
java.lang.reflect.Field contentField = message.getClass().getDeclaredField("content");
|
||||
contentField.setAccessible(true);
|
||||
Object value = contentField.get(message);
|
||||
if (value != null) {
|
||||
String text = value.toString();
|
||||
log.debug("通过 content 字段提取成功");
|
||||
return text;
|
||||
}
|
||||
} catch (NoSuchFieldException e) {
|
||||
// content 字段不存在,尝试下一种方法
|
||||
}
|
||||
|
||||
// 方法 3: 尝试调用 getText() 方法
|
||||
try {
|
||||
java.lang.reflect.Method getTextMethod = message.getClass().getMethod("getText");
|
||||
Object value = getTextMethod.invoke(message);
|
||||
if (value != null) {
|
||||
String text = value.toString();
|
||||
log.debug("通过 getText() 方法提取成功");
|
||||
return text;
|
||||
}
|
||||
} catch (NoSuchMethodException e) {
|
||||
// getText() 方法不存在,尝试下一种方法
|
||||
}
|
||||
|
||||
// 方法 4: 尝试调用 getContent() 方法
|
||||
try {
|
||||
java.lang.reflect.Method getContentMethod = message.getClass().getMethod("getContent");
|
||||
Object value = getContentMethod.invoke(message);
|
||||
if (value != null) {
|
||||
String text = value.toString();
|
||||
log.debug("通过 getContent() 方法提取成功");
|
||||
return text;
|
||||
}
|
||||
} catch (NoSuchMethodException e) {
|
||||
// getContent() 方法不存在
|
||||
}
|
||||
|
||||
// 方法 5: 打印所有字段和方法,帮助调试
|
||||
log.warn("无法提取 AssistantMessage 文本内容,打印类信息:");
|
||||
log.warn("类名: {}", message.getClass().getName());
|
||||
log.warn("字段列表:");
|
||||
for (java.lang.reflect.Field field : message.getClass().getDeclaredFields()) {
|
||||
log.warn(" - {}: {}", field.getName(), field.getType().getSimpleName());
|
||||
}
|
||||
log.warn("方法列表:");
|
||||
for (java.lang.reflect.Method method : message.getClass().getMethods()) {
|
||||
if (method.getName().startsWith("get") && method.getParameterCount() == 0) {
|
||||
log.warn(" - {}(): {}", method.getName(), method.getReturnType().getSimpleName());
|
||||
}
|
||||
}
|
||||
|
||||
// 方法 6: 最后尝试 toString()
|
||||
String toString = message.toString();
|
||||
if (toString != null && !toString.startsWith("AssistantMessage@")) {
|
||||
log.debug("通过 toString() 提取");
|
||||
return toString;
|
||||
}
|
||||
|
||||
return null;
|
||||
} catch (Exception e) {
|
||||
log.error("提取 AssistantMessage 文本内容时出错", e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取消息角色
|
||||
*/
|
||||
private String getMessageRole(Message message) {
|
||||
if (message instanceof UserMessage) {
|
||||
return "User(用户)";
|
||||
} else if (message instanceof AssistantMessage) {
|
||||
return "Assistant(模型)";
|
||||
} else if (message instanceof ToolResponseMessage) {
|
||||
return "Tool(工具返回)";
|
||||
} else {
|
||||
return message.getClass().getSimpleName();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -15,6 +15,8 @@ import org.springframework.ai.chat.messages.AssistantMessage;
|
||||
import org.springframework.ai.tool.ToolCallback;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.stereotype.Service;
|
||||
import com.superbiz.agent.config.AiOpsPromptProperties;
|
||||
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Optional;
|
||||
@@ -40,6 +42,12 @@ public class AiOpsService {
|
||||
@Autowired(required = false) // Mock 模式下才注册
|
||||
private QueryLogsTools queryLogsTools;
|
||||
|
||||
@Autowired
|
||||
private LookupKnowledgeTool lookupKnowledgeTool;
|
||||
|
||||
@Autowired
|
||||
private AiOpsPromptProperties promptProperties;
|
||||
|
||||
/**
|
||||
* 执行 AI Ops 告警分析流程
|
||||
*
|
||||
@@ -60,7 +68,7 @@ public class AiOpsService {
|
||||
.name("ai_ops_supervisor")
|
||||
.description("负责调度 Planner 与 Executor 的多 Agent 控制器")
|
||||
.model(chatModel)
|
||||
.systemPrompt(buildSupervisorSystemPrompt())
|
||||
.systemPrompt(promptProperties.getSupervisor())
|
||||
.subAgents(List.of(plannerAgent, executorAgent))
|
||||
.build();
|
||||
|
||||
@@ -113,7 +121,7 @@ public class AiOpsService {
|
||||
.name("planner_agent")
|
||||
.description("负责拆解告警、规划与再规划步骤")
|
||||
.model(chatModel)
|
||||
.systemPrompt(buildPlannerPrompt())
|
||||
.systemPrompt(promptProperties.getPlanner())
|
||||
.methodTools(buildMethodToolsArray())
|
||||
.tools(toolCallbacks)
|
||||
.outputKey("planner_plan")
|
||||
@@ -128,7 +136,7 @@ public class AiOpsService {
|
||||
.name("executor_agent")
|
||||
.description("负责执行 Planner 的首个步骤并及时反馈")
|
||||
.model(chatModel)
|
||||
.systemPrompt(buildExecutorPrompt())
|
||||
.systemPrompt(promptProperties.getExecutor())
|
||||
.methodTools(buildMethodToolsArray())
|
||||
.tools(toolCallbacks)
|
||||
.outputKey("executor_feedback")
|
||||
@@ -138,152 +146,15 @@ public class AiOpsService {
|
||||
/**
|
||||
* 动态构建方法工具数组
|
||||
* 根据 cls.mock-enabled 决定是否包含 QueryLogsTools
|
||||
* 工具顺序:知识库查询优先,日志查询次之,弃用工具最后
|
||||
*/
|
||||
private Object[] buildMethodToolsArray() {
|
||||
if (queryLogsTools != null) {
|
||||
// Mock 模式:包含 QueryLogsTools
|
||||
return new Object[]{dateTimeTools, internalDocsTools, queryMetricsTools, queryLogsTools};
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools, queryLogsTools};
|
||||
} else {
|
||||
// 真实模式:不包含 QueryLogsTools(由 MCP 提供日志查询功能)
|
||||
return new Object[]{dateTimeTools, internalDocsTools, queryMetricsTools};
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools};
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建 Planner Agent 系统提示词
|
||||
*/
|
||||
private String buildPlannerPrompt() {
|
||||
return """
|
||||
你是 Planner Agent,同时承担 Replanner 角色,负责:
|
||||
1. 读取当前输入任务 {input} 以及 Executor 的最近反馈 {executor_feedback}。
|
||||
2. 分析 Prometheus 告警、日志、内部文档等信息,制定可执行的下一步步骤。
|
||||
3. 在执行阶段,输出 JSON,包含 decision (PLAN|EXECUTE|FINISH)、step 描述、预期要调用的工具、以及必要的上下文。
|
||||
4. 调用任何腾讯云日志/主题相关工具时,region 参数必须使用连字符格式(如 ap-guangzhou),若不确定请省略以使用默认值。
|
||||
5. 严格禁止编造数据,只能引用工具返回的真实内容;如果连续 3 次调用同一工具仍失败或返回空结果,需停止该方向并在最终报告的结论部分说明"无法完成"的原因。
|
||||
|
||||
## 最终报告输出要求(CRITICAL)
|
||||
|
||||
当 decision=FINISH 时,你必须:
|
||||
1. **不要输出 JSON 格式**
|
||||
2. **直接输出完整的 Markdown 格式报告文本**
|
||||
3. **报告必须严格遵循以下模板**:
|
||||
|
||||
```
|
||||
# 告警分析报告
|
||||
|
||||
---
|
||||
|
||||
## 📋 活跃告警清单
|
||||
|
||||
| 告警名称 | 级别 | 目标服务 | 首次触发时间 | 最新触发时间 | 状态 |
|
||||
|---------|------|----------|-------------|-------------|------|
|
||||
| [告警1名称] | [级别] | [服务名] | [时间] | [时间] | 活跃 |
|
||||
| [告警2名称] | [级别] | [服务名] | [时间] | [时间] | 活跃 |
|
||||
|
||||
---
|
||||
|
||||
## 🔍 告警根因分析1 - [告警名称]
|
||||
|
||||
### 告警详情
|
||||
- **告警级别**: [级别]
|
||||
- **受影响服务**: [服务名]
|
||||
- **持续时间**: [X分钟]
|
||||
|
||||
### 症状描述
|
||||
[根据监控指标描述症状]
|
||||
|
||||
### 日志证据
|
||||
[引用查询到的关键日志]
|
||||
|
||||
### 根因结论
|
||||
[基于证据得出的根本原因]
|
||||
|
||||
---
|
||||
|
||||
## 🛠️ 处理方案执行1 - [告警名称]
|
||||
|
||||
### 已执行的排查步骤
|
||||
1. [步骤1]
|
||||
2. [步骤2]
|
||||
|
||||
### 处理建议
|
||||
[给出具体的处理建议]
|
||||
|
||||
### 预期效果
|
||||
[说明预期的效果]
|
||||
|
||||
---
|
||||
|
||||
## 🔍 告警根因分析2 - [告警名称]
|
||||
[如果有第2个告警,重复上述格式]
|
||||
|
||||
---
|
||||
|
||||
## 📊 结论
|
||||
|
||||
### 整体评估
|
||||
[总结所有告警的整体情况]
|
||||
|
||||
### 关键发现
|
||||
- [发现1]
|
||||
- [发现2]
|
||||
|
||||
### 后续建议
|
||||
1. [建议1]
|
||||
2. [建议2]
|
||||
|
||||
### 风险评估
|
||||
[评估当前风险等级和影响范围]
|
||||
```
|
||||
|
||||
**重要提醒**:
|
||||
- 最终输出必须是纯 Markdown 文本,不要包含 JSON 结构
|
||||
- 不要使用 "finalReport": "..." 这样的格式
|
||||
- 直接从 "# 告警分析报告" 开始输出
|
||||
- 所有内容必须基于工具查询的真实数据,严禁编造
|
||||
- 如果某个步骤失败,在结论中如实说明,不要跳过
|
||||
|
||||
""";
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建 Executor Agent 系统提示词
|
||||
*/
|
||||
private String buildExecutorPrompt() {
|
||||
return """
|
||||
你是 Executor Agent,负责读取 Planner 最新输出 {planner_plan},只执行其中的第一步。
|
||||
- 确认步骤所需的工具与参数,尤其是 region 参数要使用连字符格式(ap-guangzhou);若 Planner 未给出则使用默认区域。
|
||||
- 调用相应的工具并收集结果,如工具返回错误或空数据,需要将失败原因、请求参数一并记录,并停止进一步调用该工具(同一工具失败达到 3 次时应直接返回 FAILED)。
|
||||
- 将日志、指标、文档等证据整理成结构化摘要,标注对应的告警名称或资源,方便 Planner 填充"告警根因分析 / 处理方案执行"章节。
|
||||
- 以 JSON 形式返回执行状态、证据以及给 Planner 的建议,写入 executor_feedback,严禁编造未实际查询到的内容。
|
||||
|
||||
|
||||
输出示例:
|
||||
{
|
||||
"status": "SUCCESS",
|
||||
"summary": "近1小时未见 error 日志,仅有 info",
|
||||
"evidence": "...",
|
||||
"nextHint": "建议转向高占用进程"
|
||||
}
|
||||
""";
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建 Supervisor Agent 系统提示词
|
||||
*/
|
||||
private String buildSupervisorSystemPrompt() {
|
||||
return """
|
||||
你是 AI Ops Supervisor,负责调度 planner_agent 与 executor_agent:
|
||||
1. 当需要拆解任务或重新制定策略时,调用 planner_agent。
|
||||
2. 当 planner_agent 输出 decision=EXECUTE 时,调用 executor_agent 执行第一步。
|
||||
3. 根据 executor_agent 的反馈,评估是否需要再次调用 planner_agent,直到 decision=FINISH。
|
||||
4. FINISH 后,确保向最终用户输出完整的《告警分析报告》,格式必须严格为:
|
||||
告警分析报告\n---\n# 告警处理详情\n## 活跃告警清单\n## 告警根因分析N\n## 处理方案执行N\n## 结论。
|
||||
5. 若步骤涉及腾讯云日志/主题工具,请确保使用连字符区域 ID(ap-guangzhou 等),或省略 region 以采用默认值。
|
||||
6. 如果发现 Planner/Executor 在同一方向连续 3 次调用工具仍失败或没有数据,必须终止流程,直接输出"任务无法完成"的报告,明确告知失败原因,严禁凭空编造结果。
|
||||
|
||||
只允许在 planner_agent、executor_agent 与 FINISH 之间做出选择。
|
||||
|
||||
""";
|
||||
}
|
||||
}
|
||||
|
||||
@@ -6,6 +6,9 @@ import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||
import com.superbiz.agent.agent.tool.InternalDocsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
||||
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
import org.springframework.ai.chat.model.ChatModel;
|
||||
@@ -44,6 +47,9 @@ public class ChatService {
|
||||
@Autowired
|
||||
private ChatModel chatModel;
|
||||
|
||||
@Autowired
|
||||
private LookupKnowledgeTool lookupKnowledgeTool;
|
||||
|
||||
/**
|
||||
* 获取注入的 ChatModel
|
||||
*/
|
||||
@@ -62,7 +68,7 @@ public class ChatService {
|
||||
// 基础系统提示
|
||||
systemPromptBuilder.append("你是一个专业的智能助手,可以获取当前时间、查询天气信息、搜索内部文档知识库,以及查询 Prometheus 告警信息。\n");
|
||||
systemPromptBuilder.append("当用户询问时间相关问题时,**必须每次都调用 getCurrentDateTime 工具**,因为时间会不断变化。即使历史消息中有时间信息,也不要直接复用,必须重新查询最新时间。\n");
|
||||
systemPromptBuilder.append("当用户需要查询公司内部文档、流程、最佳实践或技术指南时,使用 queryInternalDocs 工具。\n");
|
||||
systemPromptBuilder.append("当用户需要查询公司内部文档、流程、最佳实践或技术指南时,使用 lookupKnowledgeTool 工具。\n");
|
||||
systemPromptBuilder.append("当用户需要查询 Prometheus 告警、监控指标或系统告警状态时,使用 queryPrometheusAlerts 工具。\n");
|
||||
systemPromptBuilder.append("当用户需要查询腾讯云日志时,请调用腾讯云mcp服务查询,默认查询地域ap-guangzhou,查询时间范围为近一个月。\n\n");
|
||||
|
||||
@@ -126,10 +132,10 @@ public class ChatService {
|
||||
public Object[] buildMethodToolsArray() {
|
||||
if (queryLogsTools != null) {
|
||||
// Mock 模式:包含 QueryLogsTools
|
||||
return new Object[]{dateTimeTools, internalDocsTools, queryMetricsTools, queryLogsTools};
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool};
|
||||
} else {
|
||||
// 真实模式:不包含 QueryLogsTools(由 MCP 提供日志查询功能)
|
||||
return new Object[]{dateTimeTools, internalDocsTools, queryMetricsTools};
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools};
|
||||
}
|
||||
}
|
||||
|
||||
@@ -171,6 +177,7 @@ public class ChatService {
|
||||
.systemPrompt(systemPrompt)
|
||||
.methodTools(buildMethodToolsArray())
|
||||
.tools(getToolCallbacks())
|
||||
.hooks(new AgentLoggingHook()) // 添加日志 Hook
|
||||
.build();
|
||||
}
|
||||
|
||||
@@ -181,10 +188,19 @@ public class ChatService {
|
||||
* @return AI 回复
|
||||
*/
|
||||
public String executeChat(ReactAgent agent, String question) throws GraphRunnerException {
|
||||
logger.info("执行 ReactAgent.call() - 自动处理工具调用");
|
||||
logger.info("========================================");
|
||||
logger.info("📝 用户问题: {}", question);
|
||||
|
||||
long startTime = System.currentTimeMillis();
|
||||
var response = agent.call(question);
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
|
||||
String answer = response.getText();
|
||||
logger.info("ReactAgent 对话完成,答案长度: {}", answer.length());
|
||||
|
||||
logger.info("⏱️ 总耗时: {} ms", duration);
|
||||
logger.info("📏 输出长度: {} 字符", answer.length());
|
||||
logger.info("========================================");
|
||||
|
||||
return answer;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -287,12 +287,12 @@ public class DocumentManagementService {
|
||||
*/
|
||||
private FaultCategory parseFaultCategory(String category) {
|
||||
if (category == null || category.isBlank()) {
|
||||
return FaultCategory.EXTERNAL_API;
|
||||
return FaultCategory.GENERAL;
|
||||
}
|
||||
try {
|
||||
return FaultCategory.valueOf(category.toUpperCase());
|
||||
} catch (IllegalArgumentException e) {
|
||||
return FaultCategory.EXTERNAL_API;
|
||||
return FaultCategory.GENERAL;
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,347 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.domain.entity.ApiDocument;
|
||||
import com.superbiz.agent.domain.enums.FaultCategory;
|
||||
import com.superbiz.agent.repository.ApiDocumentRepository;
|
||||
import com.superbiz.agent.dto.KnowledgeEntry;
|
||||
import com.superbiz.agent.dto.Frontmatter;
|
||||
import com.superbiz.agent.dto.DocumentChunk;
|
||||
import lombok.Data;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Service;
|
||||
import org.springframework.transaction.annotation.Transactional;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.file.*;
|
||||
import java.nio.file.attribute.BasicFileAttributes;
|
||||
import java.time.LocalDateTime;
|
||||
import java.util.*;
|
||||
import java.util.stream.Collectors;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* 知识库初始化服务
|
||||
* 负责批量导入 knowledge_base 目录下的文档到数据库和 Milvus
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeBaseInitService {
|
||||
|
||||
private static final Logger logger = LoggerFactory.getLogger(KnowledgeBaseInitService.class);
|
||||
|
||||
@Value("${knowledge.base-path:knowledge_base}")
|
||||
private String knowledgeBasePath;
|
||||
|
||||
@Autowired
|
||||
private ApiDocumentRepository apiDocumentRepository;
|
||||
|
||||
@Autowired
|
||||
private FrontmatterParser frontmatterParser;
|
||||
|
||||
@Autowired
|
||||
private DocumentChunkService documentChunkService;
|
||||
|
||||
@Autowired
|
||||
private VectorIndexService vectorIndexService;
|
||||
|
||||
@Autowired
|
||||
private VectorEmbeddingService vectorEmbeddingService;
|
||||
|
||||
@Autowired
|
||||
private KnowledgeIndexService knowledgeIndexService;
|
||||
|
||||
/**
|
||||
* 初始化知识库
|
||||
*
|
||||
* @param force 是否强制重新导入(跳过去重检查)
|
||||
* @return 初始化结果
|
||||
*/
|
||||
@Transactional(rollbackFor = Exception.class)
|
||||
public InitResult initializeKnowledgeBase(boolean force) {
|
||||
logger.info("开始初始化知识库: basePath={}, force={}", knowledgeBasePath, force);
|
||||
|
||||
InitResult result = new InitResult();
|
||||
Path baseDir = Paths.get(knowledgeBasePath);
|
||||
|
||||
if (!Files.exists(baseDir)) {
|
||||
logger.error("知识库目录不存在: {}", knowledgeBasePath);
|
||||
throw new RuntimeException("知识库目录不存在: " + knowledgeBasePath);
|
||||
}
|
||||
|
||||
// 1. 扫描所有 Markdown 文件
|
||||
List<Path> markdownFiles = scanMarkdownFiles(baseDir);
|
||||
result.setScanned(markdownFiles.size());
|
||||
logger.info("扫描到 {} 个 Markdown 文件", markdownFiles.size());
|
||||
|
||||
// 2. 如果非强制模式,获取已存在的文档(用于去重)
|
||||
Set<String> existingFilePaths = new HashSet<>();
|
||||
if (!force) {
|
||||
existingFilePaths = apiDocumentRepository.findAll().stream()
|
||||
.map(ApiDocument::getFilePath)
|
||||
.collect(Collectors.toSet());
|
||||
logger.info("已存在 个文档记录", existingFilePaths.size());
|
||||
}
|
||||
|
||||
// 3. 逐个处理文档
|
||||
for (Path file : markdownFiles) {
|
||||
String relativePath = baseDir.relativize(file).toString().replace("\\", "/");
|
||||
|
||||
try {
|
||||
// 去重检查
|
||||
if (!force && existingFilePaths.contains(relativePath)) {
|
||||
logger.debug("跳过已存在的文档: {}", relativePath);
|
||||
result.incrementSkipped();
|
||||
result.addDetail(relativePath, "已存在,跳过");
|
||||
continue;
|
||||
}
|
||||
|
||||
// 解析文档
|
||||
String content = Files.readString(file);
|
||||
Frontmatter frontmatter = frontmatterParser.parse(content);
|
||||
|
||||
if (frontmatter == null) {
|
||||
logger.warn("文档格式无效: {}, frontmatter 解析失败", relativePath);
|
||||
result.incrementFailed();
|
||||
result.addDetail(relativePath, "格式无效: frontmatter 解析失败");
|
||||
continue;
|
||||
}
|
||||
|
||||
// 提取字段
|
||||
String title = frontmatter.getTitle();
|
||||
String summary = frontmatter.getSummary();
|
||||
String category = frontmatter.getCategory() != null ? frontmatter.getCategory() : "general";
|
||||
List<String> keywords = frontmatter.getKeywords();
|
||||
|
||||
if (title == null || title.isBlank()) {
|
||||
logger.warn("文档缺少标题: {}", relativePath);
|
||||
result.incrementFailed();
|
||||
result.addDetail(relativePath, "缺少标题");
|
||||
continue;
|
||||
}
|
||||
|
||||
// 保存到数据库
|
||||
ApiDocument document = saveToDatabase(relativePath, title, summary, category, content, keywords);
|
||||
|
||||
// 提取文档正文(去除 frontmatter)
|
||||
String body = extractBody(content);
|
||||
|
||||
// 文档分块
|
||||
List<DocumentChunk> chunks = documentChunkService.chunkDocument(body, relativePath);
|
||||
logger.debug("文档分块完成: {} -> {} 个 chunk", relativePath, chunks.size());
|
||||
|
||||
// 上传到 Milvus
|
||||
try {
|
||||
vectorIndexService.indexDocumentChunks(document.getDocId(), chunks, category);
|
||||
|
||||
document.setStatus("INDEXED");
|
||||
document.setChunkCount(chunks.size());
|
||||
document.setIndexedAt(LocalDateTime.now());
|
||||
apiDocumentRepository.save(document);
|
||||
|
||||
logger.info("文档已索引到 Milvus: {} (docId={}, chunks={})",
|
||||
title, document.getDocId(), chunks.size());
|
||||
} catch (Exception e) {
|
||||
logger.error("上传到 Milvus 失败: {}", relativePath, e);
|
||||
|
||||
document.setStatus("FAILED");
|
||||
document.setErrorMessage(e.getMessage());
|
||||
apiDocumentRepository.save(document);
|
||||
|
||||
result.incrementFailed();
|
||||
result.addDetail(relativePath, "Milvus 索引失败: " + e.getMessage());
|
||||
continue; // 跳过该文档,继续处理下一个
|
||||
}
|
||||
|
||||
// 添加到 L0 内存索引
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath(relativePath)
|
||||
.title(title)
|
||||
.keywords(keywords)
|
||||
.summary(summary)
|
||||
.category(category)
|
||||
.build();
|
||||
knowledgeIndexService.addToIndex(entry);
|
||||
|
||||
result.incrementInserted();
|
||||
result.addDetail(relativePath, "导入成功(L0+L1)");
|
||||
logger.info("文档导入成功: {} -> {} (L0+L1 索引已更新)", relativePath, title);
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("处理文档失败: {}", relativePath, e);
|
||||
result.incrementFailed();
|
||||
result.addDetail(relativePath, "处理失败: " + e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
logger.info("知识库初始化完成: 扫描={}, 跳过={}, 新增={}, 失败={}",
|
||||
result.getScanned(), result.getSkipped(), result.getInserted(), result.getFailed());
|
||||
|
||||
return result;
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取知识库统计信息
|
||||
*/
|
||||
public Stats getStats() {
|
||||
Stats stats = new Stats();
|
||||
|
||||
// 数据库中的文档数量
|
||||
long totalDocuments = apiDocumentRepository.count();
|
||||
stats.setTotalDocuments(totalDocuments);
|
||||
|
||||
// L0 索引中的文档数量
|
||||
int indexSize = knowledgeIndexService.getIndexSize();
|
||||
logger.debug("L0 索引大小: {}", indexSize);
|
||||
|
||||
// 按分类统计(从 fault_category 字段读取)
|
||||
Map<String, Long> categoryCount = apiDocumentRepository.findAll().stream()
|
||||
.collect(Collectors.groupingBy(
|
||||
doc -> doc.getFaultCategory() != null ? doc.getFaultCategory().name() : "GENERAL",
|
||||
Collectors.counting()
|
||||
));
|
||||
stats.setCategoryCount(categoryCount);
|
||||
|
||||
// Milvus 中的向量数量(需要实现)
|
||||
// TODO: 查询 Milvus collection 的实体数量
|
||||
stats.setTotalVectors(0L);
|
||||
|
||||
return stats;
|
||||
}
|
||||
|
||||
/**
|
||||
* 扫描目录下所有 Markdown 文件
|
||||
*/
|
||||
private List<Path> scanMarkdownFiles(Path baseDir) {
|
||||
List<Path> files = new ArrayList<>();
|
||||
|
||||
try {
|
||||
Files.walkFileTree(baseDir, new SimpleFileVisitor<Path>() {
|
||||
@Override
|
||||
public FileVisitResult visitFile(Path file, BasicFileAttributes attrs) {
|
||||
if (file.toString().endsWith(".md")) {
|
||||
files.add(file);
|
||||
}
|
||||
return FileVisitResult.CONTINUE;
|
||||
}
|
||||
|
||||
@Override
|
||||
public FileVisitResult visitFileFailed(Path file, IOException exc) {
|
||||
logger.warn("访问文件失败: {}", file, exc);
|
||||
return FileVisitResult.CONTINUE;
|
||||
}
|
||||
});
|
||||
} catch (IOException e) {
|
||||
logger.error("扫描目录失败: {}", baseDir, e);
|
||||
throw new RuntimeException("扫描目录失败", e);
|
||||
}
|
||||
|
||||
return files;
|
||||
}
|
||||
|
||||
/**
|
||||
* 保存文档到数据库
|
||||
*/
|
||||
private ApiDocument saveToDatabase(String filePath, String title, String summary,
|
||||
String category, String content, List<String> keywords) {
|
||||
ApiDocument document = new ApiDocument();
|
||||
document.setDocId(UUID.randomUUID().toString());
|
||||
document.setFileName(Paths.get(filePath).getFileName().toString());
|
||||
document.setFilePath(filePath);
|
||||
document.setApiName(title); // 使用 title 作为 apiName
|
||||
document.setStatus("PENDING"); // 初始状态为 PENDING,索引成功后更新为 INDEXED
|
||||
|
||||
// 映射 category 到 FaultCategory 枚举
|
||||
FaultCategory faultCategory = FaultCategory.fromString(category);
|
||||
document.setFaultCategory(faultCategory);
|
||||
|
||||
// 将 frontmatter 信息保存到 metadata(JSON 格式)
|
||||
String metadataJson = String.format(
|
||||
"{\"title\":\"%s\",\"summary\":\"%s\",\"category\":\"%s\",\"keywords\":%s}",
|
||||
escapeJson(title),
|
||||
escapeJson(summary),
|
||||
escapeJson(category),
|
||||
"[\"" + String.join("\",\"", keywords.stream().map(this::escapeJson).toArray(String[]::new)) + "\"]"
|
||||
);
|
||||
document.setMetadata(metadataJson);
|
||||
|
||||
document.setFileSize((long) content.length());
|
||||
|
||||
return apiDocumentRepository.save(document);
|
||||
}
|
||||
|
||||
/**
|
||||
* JSON 转义
|
||||
*/
|
||||
private String escapeJson(String str) {
|
||||
if (str == null) {
|
||||
return "";
|
||||
}
|
||||
return str.replace("\\", "\\\\")
|
||||
.replace("\"", "\\\"")
|
||||
.replace("\n", "\\n")
|
||||
.replace("\r", "\\r");
|
||||
}
|
||||
|
||||
/**
|
||||
* 提取文档正文(去除 frontmatter)
|
||||
*/
|
||||
private String extractBody(String content) {
|
||||
if (!content.trim().startsWith("---")) {
|
||||
return content;
|
||||
}
|
||||
|
||||
int firstEnd = content.indexOf("---", 3);
|
||||
if (firstEnd == -1) {
|
||||
return content;
|
||||
}
|
||||
|
||||
int secondEnd = content.indexOf("---", firstEnd + 3);
|
||||
if (secondEnd == -1) {
|
||||
return content.substring(firstEnd + 3).trim();
|
||||
}
|
||||
|
||||
return content.substring(secondEnd + 3).trim();
|
||||
}
|
||||
|
||||
// ==================== 数据模型 ====================
|
||||
|
||||
/**
|
||||
* 初始化结果
|
||||
*/
|
||||
@Data
|
||||
public static class InitResult {
|
||||
private int scanned; // 扫描到的文件数量
|
||||
private int skipped; // 跳过的文件数量(已存在)
|
||||
private int inserted; // 成功导入的文件数量
|
||||
private int failed; // 失败的文件数量
|
||||
private Map<String, String> details = new LinkedHashMap<>(); // 详细信息
|
||||
|
||||
public void incrementSkipped() {
|
||||
this.skipped++;
|
||||
}
|
||||
|
||||
public void incrementInserted() {
|
||||
this.inserted++;
|
||||
}
|
||||
|
||||
public void incrementFailed() {
|
||||
this.failed++;
|
||||
}
|
||||
|
||||
public void addDetail(String filePath, String message) {
|
||||
this.details.put(filePath, message);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 统计信息
|
||||
*/
|
||||
@Data
|
||||
public static class Stats {
|
||||
private long totalDocuments; // 数据库中的文档总数
|
||||
private long totalVectors; // Milvus 中的向量总数
|
||||
private Map<String, Long> categoryCount; // 按分类统计
|
||||
}
|
||||
}
|
||||
@@ -1,5 +1,7 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.domain.entity.ApiDocument;
|
||||
import com.superbiz.agent.repository.ApiDocumentRepository;
|
||||
import com.superbiz.agent.dto.Frontmatter;
|
||||
import com.superbiz.agent.dto.KnowledgeEntry;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
@@ -12,6 +14,8 @@ import java.io.IOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.nio.file.Paths;
|
||||
import java.util.Arrays;
|
||||
import java.util.Collections;
|
||||
import java.util.List;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.stream.Collectors;
|
||||
@@ -25,11 +29,11 @@ import java.util.stream.Stream;
|
||||
@Service
|
||||
public class KnowledgeIndexService {
|
||||
|
||||
@Value("${knowledge.base-path}")
|
||||
@Value("${knowledge.base-path:knowledge_base}")
|
||||
private String knowledgeBasePath;
|
||||
|
||||
@Autowired
|
||||
private FrontmatterParser frontmatterParser;
|
||||
private ApiDocumentRepository apiDocumentRepository;
|
||||
|
||||
/**
|
||||
* 内存索引(线程安全)
|
||||
@@ -37,88 +41,108 @@ public class KnowledgeIndexService {
|
||||
private final List<KnowledgeEntry> knowledgeIndex = new CopyOnWriteArrayList<>();
|
||||
|
||||
/**
|
||||
* 启动时扫描知识库目录,构建索引
|
||||
* 启动时从数据库加载索引
|
||||
*/
|
||||
@PostConstruct
|
||||
public void loadIndex() {
|
||||
log.info("开始扫描知识库目录: {}", knowledgeBasePath);
|
||||
log.info("开始从数据库加载知识库索引");
|
||||
|
||||
try {
|
||||
Path basePath = Paths.get(knowledgeBasePath);
|
||||
// 从数据库读取所有已索引的文档
|
||||
List<ApiDocument> documents = apiDocumentRepository.findAll();
|
||||
|
||||
// 目录不存在时自动创建
|
||||
if (!Files.exists(basePath)) {
|
||||
Files.createDirectories(basePath);
|
||||
log.info("知识库目录已创建: {}", basePath.toAbsolutePath());
|
||||
int loaded = 0;
|
||||
for (ApiDocument doc : documents) {
|
||||
try {
|
||||
// 从 metadata JSON 中提取信息
|
||||
KnowledgeEntry entry = parseDocumentToEntry(doc);
|
||||
if (entry != null) {
|
||||
knowledgeIndex.add(entry);
|
||||
loaded++;
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.warn("解析文档失败: docId={}, error={}", doc.getDocId(), e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
// 递归扫描 .md 文件
|
||||
try (Stream<Path> paths = Files.walk(basePath)) {
|
||||
paths.filter(p -> p.toString().endsWith(".md"))
|
||||
.forEach(this::indexFile);
|
||||
}
|
||||
log.info("知识库索引加载完成,共 {} 个文档", loaded);
|
||||
|
||||
log.info("知识库索引加载完成,共 {} 个文档", knowledgeIndex.size());
|
||||
|
||||
} catch (IOException e) {
|
||||
} catch (Exception e) {
|
||||
log.error("知识库索引加载失败", e);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 索引单个文件
|
||||
*
|
||||
* @param filePath 文件路径
|
||||
* 将 ApiDocument 转换为 KnowledgeEntry
|
||||
*/
|
||||
private void indexFile(Path filePath) {
|
||||
private KnowledgeEntry parseDocumentToEntry(ApiDocument doc) {
|
||||
if (doc.getMetadata() == null || doc.getMetadata().isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
|
||||
try {
|
||||
// 读取文件内容
|
||||
String content = Files.readString(filePath);
|
||||
// 简单的 JSON 解析
|
||||
String metadata = doc.getMetadata();
|
||||
|
||||
// 解析 frontmatter
|
||||
Frontmatter frontmatter = frontmatterParser.parse(content);
|
||||
if (frontmatter == null) {
|
||||
log.debug("跳过文件(无有效 frontmatter): {}", filePath);
|
||||
return;
|
||||
}
|
||||
String title = extractJsonValue(metadata, "title");
|
||||
String summary = extractJsonValue(metadata, "summary");
|
||||
String category = extractJsonValue(metadata, "category");
|
||||
List<String> keywords = extractJsonArray(metadata, "keywords");
|
||||
|
||||
// 提取 category(从路径中获取)
|
||||
String category = extractCategoryFromPath(filePath.toString());
|
||||
|
||||
// 构建索引条目
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath(filePath.toString())
|
||||
.title(frontmatter.getTitle())
|
||||
.keywords(frontmatter.getKeywords())
|
||||
.summary(frontmatter.getSummary())
|
||||
return KnowledgeEntry.builder()
|
||||
.filePath(doc.getFilePath())
|
||||
.title(title != null ? title : doc.getApiName())
|
||||
.keywords(keywords)
|
||||
.summary(summary)
|
||||
.category(category)
|
||||
.sections(frontmatter.getSections())
|
||||
.build();
|
||||
|
||||
knowledgeIndex.add(entry);
|
||||
log.debug("文档已加入索引: title={}, filePath={}", entry.getTitle(), filePath);
|
||||
|
||||
} catch (IOException e) {
|
||||
log.warn("读取文件失败: {}", filePath, e);
|
||||
} catch (Exception e) {
|
||||
log.warn("解析 metadata 失败: {}", doc.getDocId(), e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 从文件路径中提取 category
|
||||
* 例如:knowledge_base/api/test.md -> api
|
||||
* 从 JSON 字符串中提取值
|
||||
*/
|
||||
private String extractCategoryFromPath(String filePath) {
|
||||
String normalized = filePath.replace("\\", "/");
|
||||
String[] parts = normalized.split("/");
|
||||
|
||||
// 查找 knowledge_base 后的第一个目录
|
||||
for (int i = 0; i < parts.length - 1; i++) {
|
||||
if (parts[i].equals("knowledge_base") && i + 1 < parts.length) {
|
||||
return parts[i + 1];
|
||||
}
|
||||
private String extractJsonValue(String json, String key) {
|
||||
String pattern = "\"" + key + "\":\"";
|
||||
int startIndex = json.indexOf(pattern);
|
||||
if (startIndex == -1) {
|
||||
return null;
|
||||
}
|
||||
|
||||
return "default";
|
||||
startIndex += pattern.length();
|
||||
int endIndex = json.indexOf("\"", startIndex);
|
||||
if (endIndex == -1) {
|
||||
return null;
|
||||
}
|
||||
|
||||
return json.substring(startIndex, endIndex);
|
||||
}
|
||||
|
||||
/**
|
||||
* 从 JSON 字符串中提取数组
|
||||
*/
|
||||
private List<String> extractJsonArray(String json, String key) {
|
||||
String pattern = "\"" + key + "\":[";
|
||||
int startIndex = json.indexOf(pattern);
|
||||
if (startIndex == -1) {
|
||||
return Collections.emptyList();
|
||||
}
|
||||
|
||||
startIndex += pattern.length();
|
||||
int endIndex = json.indexOf("]", startIndex);
|
||||
if (endIndex == -1) {
|
||||
return Collections.emptyList();
|
||||
}
|
||||
|
||||
String arrayContent = json.substring(startIndex, endIndex);
|
||||
return Arrays.stream(arrayContent.split(","))
|
||||
.map(s -> s.trim().replaceAll("^\"|\"$", ""))
|
||||
.filter(s -> !s.isEmpty())
|
||||
.collect(Collectors.toList());
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -174,13 +198,15 @@ public class KnowledgeIndexService {
|
||||
/**
|
||||
* 读取文档内容
|
||||
*
|
||||
* @param filePath 文件路径
|
||||
* @param filePath 文件相对路径(如 api/payment-errors.md)
|
||||
* @param maxChars 最大字符数
|
||||
* @return 文档内容(前 maxChars 字符),失败返回 null
|
||||
*/
|
||||
public String readDocument(String filePath, int maxChars) {
|
||||
try {
|
||||
String content = Files.readString(Paths.get(filePath));
|
||||
// 拼接完整路径:knowledge_base + 相对路径
|
||||
Path fullPath = Paths.get(knowledgeBasePath, filePath);
|
||||
String content = Files.readString(fullPath);
|
||||
|
||||
if (content.length() > maxChars) {
|
||||
return content.substring(0, maxChars) + "...";
|
||||
@@ -189,7 +215,7 @@ public class KnowledgeIndexService {
|
||||
return content;
|
||||
|
||||
} catch (IOException e) {
|
||||
log.error("读取文档失败: {}", filePath, e);
|
||||
log.error("读取文档失败: {}/{}", knowledgeBasePath, filePath, e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -30,37 +30,62 @@ public class LookupKnowledgeTool {
|
||||
* @param query 查询关键词
|
||||
* @return 查询结果
|
||||
*/
|
||||
@Tool(description = "查询知识库文档。优先精确匹配关键词,未命中或多个匹配时自动补充语义相关片段。" +
|
||||
"参数 query: 查询关键词,例如 'ERR_TIMEOUT'、'支付网关超时'")
|
||||
@Tool(description = "查询内部知识库文档,获取错误码定义、接口文档、排障步骤、配置说明等背景信息。" +
|
||||
"采用两阶段检索:L0 精确匹配关键词(< 10ms),L1 语义检索补充(200-500ms)。" +
|
||||
"IMPORTANT: 遇到错误码、接口名、配置项、排障问题时,优先使用此工具。" +
|
||||
"支持的查询场景:" +
|
||||
"1) 错误码定义 - 查询错误码的含义和处理方法,例如 'ERR_TIMEOUT'、'ERR_CONNECTION_REFUSED';" +
|
||||
"2) 接口文档 - 查询 API 接口定义、参数说明、返回格式,例如 'payment-gateway'、'/api/v1/orders';" +
|
||||
"3) 排障步骤 - 查询故障诊断流程、最佳实践,例如 '支付超时排查'、'数据库连接池配置';" +
|
||||
"4) 配置说明 - 查询系统配置、中间件参数,例如 'HikariCP'、'Redis 集群配置'。" +
|
||||
"参数 query: 查询关键词或描述")
|
||||
public LookupResult lookupKnowledge(String query) {
|
||||
// 生成请求ID用于追踪
|
||||
String requestId = java.util.UUID.randomUUID().toString().substring(0, 8);
|
||||
long startTime = System.currentTimeMillis();
|
||||
|
||||
log.info("[{}] 收到知识库查询请求: query={}", requestId, query);
|
||||
log.info("========================================");
|
||||
log.info(">>> [工具调用] lookup_knowledge");
|
||||
log.info(">>> 参数: query = \"{}\"", query);
|
||||
log.info(">>> RequestId: {}", requestId);
|
||||
log.info("----------------------------------------");
|
||||
|
||||
// Step 1: L0 精确匹配
|
||||
long l0Start = System.currentTimeMillis();
|
||||
List<KnowledgeEntry> l0Matches = knowledgeIndexService.exactMatch(query);
|
||||
long l0Time = System.currentTimeMillis() - l0Start;
|
||||
log.info("[{}] L0精确匹配完成: matches={}, time={}ms", requestId, l0Matches.size(), l0Time);
|
||||
log.info("[L0 精确匹配] 完成: matches={}, time={}ms", l0Matches.size(), l0Time);
|
||||
if (!l0Matches.isEmpty()) {
|
||||
log.info("[L0 精确匹配] 找到文档:");
|
||||
for (int i = 0; i < Math.min(3, l0Matches.size()); i++) {
|
||||
KnowledgeEntry entry = l0Matches.get(i);
|
||||
log.info(" - [{}] 标题: {}, 路径: {}", i+1, entry.getTitle(), entry.getFilePath());
|
||||
}
|
||||
}
|
||||
|
||||
// Step 2: 判断是否高置信度(唯一匹配)
|
||||
boolean highConfidence = (l0Matches.size() == 1);
|
||||
log.debug("[{}] 置信度判断: highConfidence={}, reason={}",
|
||||
requestId, highConfidence, highConfidence ? "唯一匹配" : "多个或零个匹配");
|
||||
log.info("[置信度判断] highConfidence={}, reason={}",
|
||||
highConfidence, highConfidence ? "唯一匹配" : "多个或零个匹配");
|
||||
|
||||
// Step 3: L1 条件调用
|
||||
List<VectorSearchService.SearchResult> l1Results = null;
|
||||
if (!highConfidence) {
|
||||
log.info("[{}] L0非唯一匹配,触发L1语义检索", requestId);
|
||||
log.info("[L1 语义检索] L0非唯一匹配,触发L1语义检索...");
|
||||
long l1Start = System.currentTimeMillis();
|
||||
l1Results = vectorSearchService.searchSimilarDocuments(query, 3, null);
|
||||
long l1Time = System.currentTimeMillis() - l1Start;
|
||||
log.info("[{}] L1语义检索完成: matches={}, time={}ms",
|
||||
requestId, l1Results != null ? l1Results.size() : 0, l1Time);
|
||||
log.info("[L1 语义检索] 完成: matches={}, time={}ms",
|
||||
l1Results != null ? l1Results.size() : 0, l1Time);
|
||||
if (l1Results != null && !l1Results.isEmpty()) {
|
||||
log.info("[L1 语义检索] 找到文档:");
|
||||
for (int i = 0; i < Math.min(3, l1Results.size()); i++) {
|
||||
VectorSearchService.SearchResult result = l1Results.get(i);
|
||||
log.info(" - [{}] 文档ID: {}, 相似度得分: {}", i+1, result.getId(), result.getScore());
|
||||
}
|
||||
}
|
||||
} else {
|
||||
log.debug("[{}] L0唯一匹配,跳过L1检索", requestId);
|
||||
log.info("[L1 语义检索] L0唯一匹配,跳过L1检索");
|
||||
}
|
||||
|
||||
// Step 4: 组装结果
|
||||
@@ -68,13 +93,22 @@ public class LookupKnowledgeTool {
|
||||
|
||||
// 记录完整结果
|
||||
long totalTime = System.currentTimeMillis() - startTime;
|
||||
log.info("[{}] 查询完成: found={}, hasL0={}, hasL1={}, confidence={}, totalTime={}ms",
|
||||
requestId,
|
||||
log.info("----------------------------------------");
|
||||
log.info("<<< [工具返回] lookup_knowledge");
|
||||
log.info("<<< 结果: found={}, matchType={}, confidence={}",
|
||||
result.isFound(),
|
||||
result.getPrimary() != null,
|
||||
result.getSupplement() != null,
|
||||
result.getPrimary() != null ? result.getPrimary().getConfidence() : "N/A",
|
||||
totalTime);
|
||||
result.getPrimary() != null ? result.getPrimary().getMatchType() : "N/A",
|
||||
result.getPrimary() != null ? result.getPrimary().getConfidence() : "N/A");
|
||||
log.info("<<< 总耗时: {}ms (L0={}ms, L1={}ms)",
|
||||
totalTime, l0Time, l1Results != null ? (totalTime - l0Time) : 0);
|
||||
if (result.isFound() && result.getPrimary() != null) {
|
||||
String content = result.getPrimary().getContent();
|
||||
log.info("<<< 返回内容长度: {} 字符", content != null ? content.length() : 0);
|
||||
if (content != null && content.length() > 200) {
|
||||
log.info("<<< 内容预览: {}", content.substring(0, 200) + "...");
|
||||
}
|
||||
}
|
||||
log.info("========================================");
|
||||
|
||||
return result;
|
||||
}
|
||||
|
||||
@@ -0,0 +1,76 @@
|
||||
# 执行者 System Prompt
|
||||
|
||||
## 角色定位
|
||||
|
||||
你是诊断流程的**执行者**。你的任务非常明确:严格遵循规划者下发的任务清单,按步骤调用工具完成任务,并输出最终结果。
|
||||
|
||||
---
|
||||
|
||||
## 核心行为准则
|
||||
|
||||
### 1. 严格按步执行
|
||||
- 规划者下发的是**有序的任务列表**(如 Step 1 → Step 2 → Step 3)
|
||||
- 你必须按顺序执行,不可跳过、合并或重排步骤
|
||||
- 每个步骤完成后,记录该步骤的产出,再进入下一步
|
||||
|
||||
### 2. 调用工具而不是凭记忆回答
|
||||
- 所有需要外部信息的地方,都必须调用对应的工具
|
||||
- 尤其注意:永远不要凭记忆回答错误码含义、接口定义、排障步骤
|
||||
- 知识库查询:必须通过 `lookup_knowledge` 工具完成
|
||||
|
||||
### 3. 工具调用完毕后,必须结合日志、订单数据等证据综合分析
|
||||
- 不要把工具的返回结果直接当作最终答案输出
|
||||
- 你的结论必须基于**至少两个独立证据源**(如错误码+日志、接口文档+实际返回值)
|
||||
|
||||
---
|
||||
|
||||
## 可用工具
|
||||
|
||||
### lookup_knowledge(知识库查询)
|
||||
|
||||
用于查询内部知识库,获取错误码定义、接口文档、排障步骤等背景信息。
|
||||
|
||||
| 参数 | 说明 |
|
||||
|------|------|
|
||||
| `query` | 查询关键词或描述。例如:`ERR_TIMEOUT`、`payment-gateway`、`支付为什么失败` |
|
||||
|
||||
**内部机制**:
|
||||
工具内部自动执行「先精确匹配(L0),未命中则语义检索(L1)」的两阶段检索逻辑,你无需关心哪一层。
|
||||
|
||||
**返回结果**:包含 `found`(是否找到)、`primary.content`(文档内容)、`primary.match_type`(来源标记:`exact_L0` 或 `semantic_L1`)等字段。
|
||||
|
||||
**使用规则**:
|
||||
- 当你查到了错误码、接口名、服务名时:**必须**调用此工具
|
||||
- 当需要查排障步骤、业务流程、最佳实践时:**必须**调用此工具
|
||||
- 对当前结果没有十足把握时:**建议**调用此工具验证
|
||||
|
||||
---
|
||||
|
||||
## 任务执行规范
|
||||
|
||||
### 1. 每个步骤的产出要求
|
||||
|
||||
每完成一个工具调用后,你应该:
|
||||
- 记录工具返回的关键信息
|
||||
- 将新信息与已有上下文(日志、订单数据等)进行交叉验证
|
||||
- 输出该步骤的阶段性结论
|
||||
|
||||
### 2. 最终输出的报告格式
|
||||
|
||||
```yaml
|
||||
## 诊断结论
|
||||
|
||||
**问题根因**:XXX
|
||||
|
||||
**证据链**:
|
||||
1. 订单状态返回错误码 ERR_TIMEOUT
|
||||
2. 知识库 lookup_knowledge("ERR_TIMEOUT") 返回:支付网关响应超时(>5秒)
|
||||
3. 日志确认:14:32:15 请求耗时 5.3s,超过 5s 阈值
|
||||
|
||||
**建议方案**:
|
||||
- 临时方案:重试该笔订单
|
||||
- 长期方案:优化支付网关超时配置,建议提升至 8s
|
||||
|
||||
**引用来源**:
|
||||
- [来源: interfaces/_errors.md]
|
||||
```
|
||||
@@ -0,0 +1,88 @@
|
||||
你是 Planner Agent,同时承担 Replanner 角色,负责:
|
||||
1. 读取当前输入任务 {input} 以及 Executor 的最近反馈 {executor_feedback}。
|
||||
2. 分析 Prometheus 告警、日志、内部文档等信息,制定可执行的下一步步骤。
|
||||
3. 在执行阶段,输出 JSON,包含 decision (PLAN|EXECUTE|FINISH)、step 描述、预期要调用的工具、以及必要的上下文。
|
||||
4. 调用任何腾讯云日志/主题相关工具时,region 参数必须使用连字符格式(如 ap-guangzhou),若不确定请省略以使用默认值。
|
||||
5. 严格禁止编造数据,只能引用工具返回的真实内容;如果连续 3 次调用同一工具仍失败或返回空结果,需停止该方向并在最终报告的结论部分说明"无法完成"的原因。
|
||||
|
||||
## 最终报告输出要求(CRITICAL)
|
||||
|
||||
当 decision=FINISH 时,你必须:
|
||||
1. **不要输出 JSON 格式**
|
||||
2. **直接输出完整的 Markdown 格式报告文本**
|
||||
3. **报告必须严格遵循以下模板**:
|
||||
|
||||
```
|
||||
# 告警分析报告
|
||||
|
||||
---
|
||||
|
||||
## 📋 活跃告警清单
|
||||
|
||||
| 告警名称 | 级别 | 目标服务 | 首次触发时间 | 最新触发时间 | 状态 |
|
||||
|---------|------|----------|-------------|-------------|------|
|
||||
| [告警1名称] | [级别] | [服务名] | [时间] | [时间] | 活跃 |
|
||||
| [告警2名称] | [级别] | [服务名] | [时间] | [时间] | 活跃 |
|
||||
|
||||
---
|
||||
|
||||
## 🔍 告警根因分析1 - [告警名称]
|
||||
|
||||
### 告警详情
|
||||
- **告警级别**: [级别]
|
||||
- **受影响服务**: [服务名]
|
||||
- **持续时间**: [X分钟]
|
||||
|
||||
### 症状描述
|
||||
[根据监控指标描述症状]
|
||||
|
||||
### 日志证据
|
||||
[引用查询到的关键日志]
|
||||
|
||||
### 根因结论
|
||||
[基于证据得出的根本原因]
|
||||
|
||||
---
|
||||
|
||||
## 🛠️ 处理方案执行1 - [告警名称]
|
||||
|
||||
### 已执行的排查步骤
|
||||
1. [步骤1]
|
||||
2. [步骤2]
|
||||
|
||||
### 处理建议
|
||||
[给出具体的处理建议]
|
||||
|
||||
### 预期效果
|
||||
[说明预期的效果]
|
||||
|
||||
---
|
||||
|
||||
## 🔍 告警根因分析2 - [告警名称]
|
||||
[如果有第2个告警,重复上述格式]
|
||||
|
||||
---
|
||||
|
||||
## 📊 结论
|
||||
|
||||
### 整体评估
|
||||
[总结所有告警的整体情况]
|
||||
|
||||
### 关键发现
|
||||
- [发现1]
|
||||
- [发现2]
|
||||
|
||||
### 后续建议
|
||||
1. [建议1]
|
||||
2. [建议2]
|
||||
|
||||
### 风险评估
|
||||
[评估当前风险等级和影响范围]
|
||||
```
|
||||
|
||||
**重要提醒**:
|
||||
- 最终输出必须是纯 Markdown 文本,不要包含 JSON 结构
|
||||
- 不要使用 "finalReport": "..." 这样的格式
|
||||
- 直接从 "# 告警分析报告" 开始输出
|
||||
- 所有内容必须基于工具查询的真实数据,严禁编造
|
||||
- 如果某个步骤失败,在结论中如实说明,不要跳过
|
||||
@@ -0,0 +1,10 @@
|
||||
你是 AI Ops Supervisor,负责调度 planner_agent 与 executor_agent:
|
||||
1. 当需要拆解任务或重新制定策略时,调用 planner_agent。
|
||||
2. 当 planner_agent 输出 decision=EXECUTE 时,调用 executor_agent 执行第一步。
|
||||
3. 根据 executor_agent 的反馈,评估是否需要再次调用 planner_agent,直到 decision=FINISH。
|
||||
4. FINISH 后,确保向最终用户输出完整的《告警分析报告》,格式必须严格为:
|
||||
告警分析报告\n---\n# 告警处理详情\n## 活跃告警清单\n## 告警根因分析N\n## 处理方案执行N\n## 结论。
|
||||
5. 若步骤涉及腾讯云日志/主题工具,请确保使用连字符区域 ID(ap-guangzhou 等),或省略 region 以采用默认值。
|
||||
6. 如果发现 Planner/Executor 在同一方向连续 3 次调用工具仍失败或没有数据,必须终止流程,直接输出"任务无法完成"的报告,明确告知失败原因,严禁凭空编造结果。
|
||||
|
||||
只允许在 planner_agent、executor_agent 与 FINISH 之间做出选择。
|
||||
@@ -40,7 +40,7 @@ class ApiDocumentRepositoryTest {
|
||||
.filePath("/uploads/api-spec.md")
|
||||
.fileHash("abc123hash")
|
||||
.fileSize(1024L)
|
||||
.faultCategory(FaultCategory.EXTERNAL_API)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.faultSource("广东")
|
||||
.apiName("查询接口")
|
||||
.status("PENDING")
|
||||
@@ -167,7 +167,7 @@ class ApiDocumentRepositoryTest {
|
||||
.docId(UUID.randomUUID().toString())
|
||||
.fileName("guangdong-api.md")
|
||||
.faultSource("广东")
|
||||
.faultCategory(FaultCategory.EXTERNAL_API)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.build();
|
||||
|
||||
repository.save(doc);
|
||||
|
||||
@@ -39,7 +39,7 @@ class CaseLibraryRepositoryTest {
|
||||
.title("接口超时案例")
|
||||
.rootCause("网络延迟导致接口超时")
|
||||
.solution("增加超时时间和重试机制")
|
||||
.faultCategory(FaultCategory.EXTERNAL_API)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.errorCode("40003")
|
||||
.sourceType(SourceType.AUTO)
|
||||
.build();
|
||||
@@ -63,7 +63,7 @@ class CaseLibraryRepositoryTest {
|
||||
.title("数据库死锁案例")
|
||||
.rootCause("并发更新导致死锁")
|
||||
.solution("优化事务粒度")
|
||||
.faultCategory(FaultCategory.DATABASE)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.build();
|
||||
|
||||
repository.save(caseLib);
|
||||
@@ -81,7 +81,7 @@ class CaseLibraryRepositoryTest {
|
||||
.title("案例1")
|
||||
.rootCause("原因1")
|
||||
.solution("方案1")
|
||||
.faultCategory(FaultCategory.EXTERNAL_API)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.errorCode("40003")
|
||||
.build();
|
||||
|
||||
@@ -90,7 +90,7 @@ class CaseLibraryRepositoryTest {
|
||||
.title("案例2")
|
||||
.rootCause("原因2")
|
||||
.solution("方案2")
|
||||
.faultCategory(FaultCategory.EXTERNAL_API)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.errorCode("40003")
|
||||
.build();
|
||||
|
||||
@@ -98,7 +98,7 @@ class CaseLibraryRepositoryTest {
|
||||
repository.save(case2);
|
||||
|
||||
List<CaseLibrary> results = repository.findByFaultCategoryAndErrorCode(
|
||||
FaultCategory.EXTERNAL_API, "40003");
|
||||
FaultCategory.API, "40003");
|
||||
|
||||
assertFalse(results.isEmpty());
|
||||
assertTrue(results.size() >= 2);
|
||||
|
||||
@@ -38,7 +38,7 @@ class DiagnosisRecordRepositoryTest {
|
||||
.sessionId("session-001")
|
||||
.businessId("order-12345")
|
||||
.traceId("trace-abc123")
|
||||
.faultCategory(FaultCategory.EXTERNAL_API)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.faultSource("广东")
|
||||
.faultTarget("http://api.example.com/query")
|
||||
.errorCode("40003")
|
||||
@@ -67,7 +67,7 @@ class DiagnosisRecordRepositoryTest {
|
||||
DiagnosisRecord record = DiagnosisRecord.builder()
|
||||
.diagnosisId(diagnosisId)
|
||||
.businessId("order-test-001")
|
||||
.faultCategory(FaultCategory.DATABASE)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.status(DiagnosisStatus.PENDING)
|
||||
.build();
|
||||
|
||||
@@ -84,14 +84,14 @@ class DiagnosisRecordRepositoryTest {
|
||||
// 创建测试数据
|
||||
DiagnosisRecord record1 = DiagnosisRecord.builder()
|
||||
.diagnosisId(UUID.randomUUID().toString())
|
||||
.faultCategory(FaultCategory.EXTERNAL_API)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.errorCode("40003")
|
||||
.status(DiagnosisStatus.SUCCESS)
|
||||
.build();
|
||||
|
||||
DiagnosisRecord record2 = DiagnosisRecord.builder()
|
||||
.diagnosisId(UUID.randomUUID().toString())
|
||||
.faultCategory(FaultCategory.EXTERNAL_API)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.errorCode("40003")
|
||||
.status(DiagnosisStatus.FAILED)
|
||||
.build();
|
||||
@@ -101,7 +101,7 @@ class DiagnosisRecordRepositoryTest {
|
||||
|
||||
// 查询
|
||||
List<DiagnosisRecord> results = repository.findByFaultCategoryAndErrorCode(
|
||||
FaultCategory.EXTERNAL_API, "40003");
|
||||
FaultCategory.API, "40003");
|
||||
|
||||
assertFalse(results.isEmpty());
|
||||
assertTrue(results.size() >= 2);
|
||||
@@ -113,7 +113,7 @@ class DiagnosisRecordRepositoryTest {
|
||||
DiagnosisRecord record = DiagnosisRecord.builder()
|
||||
.diagnosisId(UUID.randomUUID().toString())
|
||||
.status(DiagnosisStatus.RUNNING)
|
||||
.faultCategory(FaultCategory.CACHE)
|
||||
.faultCategory(FaultCategory.API)
|
||||
.build();
|
||||
|
||||
repository.save(record);
|
||||
|
||||
Reference in New Issue
Block a user