docs: add interview project materials
This commit is contained in:
@@ -0,0 +1,59 @@
|
|||||||
|
# SuperBizAgent Interview Guide
|
||||||
|
|
||||||
|
## 一句话定位
|
||||||
|
|
||||||
|
SuperBizAgent 是一个面向企业故障诊断场景的 Agent Engineering 项目:它把用户问题或告警事件转成可追踪的多 Agent 执行链路,并把工具证据、模型步骤、最终答案和反馈统一落到诊断 trace 中。
|
||||||
|
|
||||||
|
## 面试重点
|
||||||
|
|
||||||
|
- **多 Agent 编排**:普通 Chat 的复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 Supervisor 调度 Planner/Executor。
|
||||||
|
- **工具证据链**:知识库、日志、指标和 Prometheus 告警都通过工具调用进入链路,并记录到 `tool_invocation`。
|
||||||
|
- **可追踪诊断**:一次会话对应一个 `sessionId`,最终可以通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
|
||||||
|
- **质量门**:Chat 链路包含 Verifier,把 groundedness、facts checked 和 evidence refs 写回 `diagnosis_session.self_evaluation`。
|
||||||
|
- **AIOps 产品边界**:有告警 payload 时聚焦该告警;没有 payload 时先自动发现 active alerts。
|
||||||
|
- **可复现 Demo**:`mvp-demo` profile 使用 mock Prometheus 和 mock CLS,让面试演示不依赖真实线上故障。
|
||||||
|
|
||||||
|
## 推荐阅读顺序
|
||||||
|
|
||||||
|
1. `interview/demo-script.md`:面试现场怎么讲、怎么演示。
|
||||||
|
2. `interview/architecture.md`:系统架构和两条主链路。
|
||||||
|
3. `interview/design-tradeoffs.md`:关键设计取舍和可被追问的问题。
|
||||||
|
4. `interview/acceptance-checklist.md`:面试前验证清单。
|
||||||
|
5. `mvp/demo/README.md`:更细的 MVP 可执行 runbook。
|
||||||
|
|
||||||
|
## 核心 Demo
|
||||||
|
|
||||||
|
### Chat Diagnosis
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/chat
|
||||||
|
-> ChatService.executeChatWithStrategy(...)
|
||||||
|
-> simple ReactAgent or Planner -> Executor -> Verifier
|
||||||
|
-> lookup_knowledge / query_logs / query_metrics
|
||||||
|
-> diagnosis_session + agent_step + tool_invocation
|
||||||
|
-> GET /api/diagnosis/{sessionId}/trace
|
||||||
|
```
|
||||||
|
|
||||||
|
### AIOps Alert Diagnosis
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/ai_ops
|
||||||
|
-> AiOpsService.executeAiOpsAnalysis(...)
|
||||||
|
-> ai_ops_supervisor
|
||||||
|
-> planner_agent / executor_agent
|
||||||
|
-> queryPrometheusAlerts + logs + knowledge
|
||||||
|
-> scoped alert report
|
||||||
|
-> GET /api/diagnosis/{sessionId}/trace
|
||||||
|
```
|
||||||
|
|
||||||
|
## 当前完成度
|
||||||
|
|
||||||
|
- Chat 诊断链路:可运行、可追踪、有 Verifier。
|
||||||
|
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
|
||||||
|
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
|
||||||
|
- Demo 文档:`mvp/demo/README.md` 和 `mvp/demo/aiops-alert-acceptance.md`。
|
||||||
|
- Devflow 沉淀:`devflow/index.md` 记录了 MVP、Verifier、AIOps trace 和 AIOps scope-control 的演进。
|
||||||
|
|
||||||
|
## 面试时的主叙事
|
||||||
|
|
||||||
|
这个项目不是简单调用大模型,而是在做一个可审计的 Agent 诊断系统。核心价值是:模型可以规划和推理,但每一步工具证据、最终结论和质量评估都能被 trace API 回放。面试时重点展示“从问题到证据到答案到验证”的完整闭环。
|
||||||
@@ -0,0 +1,170 @@
|
|||||||
|
# Acceptance Checklist
|
||||||
|
|
||||||
|
## 面试前环境检查
|
||||||
|
|
||||||
|
- 当前分支包含最新 AIOps trace/scope 变更。
|
||||||
|
- MySQL 可连接。
|
||||||
|
- Redis 可连接。
|
||||||
|
- Milvus/Zilliz 可连接。
|
||||||
|
- 模型 API key 可用。
|
||||||
|
- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。
|
||||||
|
|
||||||
|
启动:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||||
|
```
|
||||||
|
|
||||||
|
编译检查:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn -q -DskipTests compile
|
||||||
|
```
|
||||||
|
|
||||||
|
目标测试:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test
|
||||||
|
```
|
||||||
|
|
||||||
|
## Chat Demo 验收
|
||||||
|
|
||||||
|
请求:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
$sessionId = "interview-chat-payment-timeout-001"
|
||||||
|
$body = @{
|
||||||
|
Id = $sessionId
|
||||||
|
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "http://localhost:9900/api/chat" `
|
||||||
|
-ContentType "application/json" `
|
||||||
|
-Body $body
|
||||||
|
```
|
||||||
|
|
||||||
|
验收:
|
||||||
|
|
||||||
|
- 返回 `data.success = true`。
|
||||||
|
- 返回 `data.sessionId = interview-chat-payment-timeout-001`。
|
||||||
|
- `diagnosis_session.agent_flow = CHAT`。
|
||||||
|
- trace API 返回 session、steps、toolInvocations。
|
||||||
|
- 复杂问题下 trace 中能看到 verifier 相关数据。
|
||||||
|
|
||||||
|
SQL:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'"
|
||||||
|
```
|
||||||
|
|
||||||
|
## AIOps Demo 验收
|
||||||
|
|
||||||
|
请求:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
$aiopsSessionId = "interview-aiops-payment-cpu-001"
|
||||||
|
$aiopsBody = @{
|
||||||
|
sessionId = $aiopsSessionId
|
||||||
|
alertName = "HighCPUUsage"
|
||||||
|
service = "payment-service"
|
||||||
|
severity = "P1"
|
||||||
|
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||||
|
timeRange = "last_15m"
|
||||||
|
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
Invoke-WebRequest `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "http://localhost:9900/api/ai_ops" `
|
||||||
|
-ContentType "application/json" `
|
||||||
|
-Body $aiopsBody
|
||||||
|
```
|
||||||
|
|
||||||
|
验收:
|
||||||
|
|
||||||
|
- SSE 首条包含 `type=session`。
|
||||||
|
- SSE 最后包含 `type=done`。
|
||||||
|
- `diagnosis_session.agent_flow = AI_OPS`。
|
||||||
|
- `diagnosis_session.status = SUCCESS`。
|
||||||
|
- `diagnosis_session.answer` 有最终报告。
|
||||||
|
- trace API 返回 AIOps steps 和 tool invocations。
|
||||||
|
- 报告主章节聚焦 `HighCPUUsage/payment-service`。
|
||||||
|
- 无 `告警根因分析 - HighMemoryUsage` 独立章节。
|
||||||
|
- 无 `告警根因分析 - SlowResponse` 独立章节。
|
||||||
|
- 有“相关风险告警”或类似上下文说明。
|
||||||
|
|
||||||
|
SQL:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, total_duration_ms, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
||||||
|
```
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name ORDER BY tool_name"
|
||||||
|
```
|
||||||
|
|
||||||
|
Scope 检查:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
python scripts/query_mysql.py "SELECT (answer LIKE '%告警根因分析 - HighCPUUsage%') AS has_main_root_cause, (answer LIKE '%告警根因分析 - HighMemoryUsage%') AS has_memory_root_cause, (answer LIKE '%告警根因分析 - SlowResponse%') AS has_slow_root_cause, (answer LIKE '%相关风险告警%') AS has_related_risk FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
||||||
|
```
|
||||||
|
|
||||||
|
期望:
|
||||||
|
|
||||||
|
```text
|
||||||
|
has_main_root_cause = 1
|
||||||
|
has_memory_root_cause = 0
|
||||||
|
has_slow_root_cause = 0
|
||||||
|
has_related_risk = 1
|
||||||
|
```
|
||||||
|
|
||||||
|
## Trace API 验收
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Get `
|
||||||
|
-Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||||
|
```
|
||||||
|
|
||||||
|
若 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||||
|
```
|
||||||
|
|
||||||
|
## 常见问题
|
||||||
|
|
||||||
|
### MySQL stale connection
|
||||||
|
|
||||||
|
现象:
|
||||||
|
|
||||||
|
```text
|
||||||
|
HikariPool - Connection is not available
|
||||||
|
No operations allowed after connection closed
|
||||||
|
```
|
||||||
|
|
||||||
|
当前已在 `application.yml` 配置:
|
||||||
|
|
||||||
|
- `maximum-pool-size: 5`
|
||||||
|
- `minimum-idle: 1`
|
||||||
|
- `connection-timeout: 10000`
|
||||||
|
- `validation-timeout: 5000`
|
||||||
|
- `idle-timeout: 60000`
|
||||||
|
- `max-lifetime: 120000`
|
||||||
|
- `keepalive-time: 30000`
|
||||||
|
|
||||||
|
处理:
|
||||||
|
|
||||||
|
- 重新编译或重启服务。
|
||||||
|
- 确认日志中新的 HikariPool 启动成功。
|
||||||
|
- 再跑 trace 或 AIOps 请求。
|
||||||
|
|
||||||
|
### SSE 客户端显示异常
|
||||||
|
|
||||||
|
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 trace API 验证结果。
|
||||||
|
|
||||||
|
### OpenSpec 全量校验失败
|
||||||
|
|
||||||
|
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试材料主要依赖已归档的 AIOps spec 和 MVP trace spec,可以单独验证相关 spec。
|
||||||
@@ -0,0 +1,147 @@
|
|||||||
|
# Architecture
|
||||||
|
|
||||||
|
## 系统分层
|
||||||
|
|
||||||
|
```text
|
||||||
|
API Layer
|
||||||
|
-> ChatController / DiagnosisTraceController
|
||||||
|
|
||||||
|
Agent Orchestration
|
||||||
|
-> ChatService / AiOpsService
|
||||||
|
|
||||||
|
Tools
|
||||||
|
-> lookupKnowledgeTool / queryLogs / queryMetrics / queryPrometheusAlerts
|
||||||
|
|
||||||
|
Persistence
|
||||||
|
-> diagnosis_session / agent_step / tool_invocation
|
||||||
|
|
||||||
|
Trace
|
||||||
|
-> GET /api/diagnosis/{sessionId}/trace
|
||||||
|
```
|
||||||
|
|
||||||
|
## Chat 链路
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
User[User Question] --> ChatAPI[POST /api/chat]
|
||||||
|
ChatAPI --> Strategy[ChatService.executeChatWithStrategy]
|
||||||
|
Strategy --> Complexity{QuestionComplexity}
|
||||||
|
Complexity -->|simple| Single[ReactAgent]
|
||||||
|
Complexity -->|complex| Planner[Planner Agent]
|
||||||
|
Planner --> Executor[Executor Agent]
|
||||||
|
Executor --> Tools[Evidence Tools]
|
||||||
|
Tools --> Executor
|
||||||
|
Executor --> Verifier[Verifier Agent]
|
||||||
|
Verifier --> Answer[Final Answer]
|
||||||
|
Answer --> Session[diagnosis_session]
|
||||||
|
Planner --> Steps[agent_step]
|
||||||
|
Executor --> Steps
|
||||||
|
Verifier --> Steps
|
||||||
|
Tools --> Invocations[tool_invocation]
|
||||||
|
Session --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
||||||
|
Steps --> Trace
|
||||||
|
Invocations --> Trace
|
||||||
|
```
|
||||||
|
|
||||||
|
关键代码:
|
||||||
|
|
||||||
|
- `ChatController.chat(...)`
|
||||||
|
- `ChatService.executeChatWithStrategy(...)`
|
||||||
|
- `ChatService.executeChatComplex(...)`
|
||||||
|
- `AgentLoggingHook`
|
||||||
|
- `ToolInvocationRecorder`
|
||||||
|
- `DiagnosisTraceService.getTrace(...)`
|
||||||
|
|
||||||
|
## AIOps 链路
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Alert[Alert Payload or Empty Request] --> AiOpsAPI[POST /api/ai_ops]
|
||||||
|
AiOpsAPI --> SessionEvent[SSE session event]
|
||||||
|
AiOpsAPI --> AiOpsService[AiOpsService.executeAiOpsAnalysis]
|
||||||
|
AiOpsService --> PromptMode{Payload?}
|
||||||
|
PromptMode -->|yes| Targeted[PAYLOAD_TARGETED]
|
||||||
|
PromptMode -->|no| Discovery[AUTO_DISCOVERY]
|
||||||
|
Targeted --> Supervisor[ai_ops_supervisor]
|
||||||
|
Discovery --> Supervisor
|
||||||
|
Supervisor --> Planner[planner_agent]
|
||||||
|
Supervisor --> Executor[executor_agent]
|
||||||
|
Planner --> Tools[Prometheus / Logs / Knowledge]
|
||||||
|
Executor --> Tools
|
||||||
|
Tools --> Report[Alert Report]
|
||||||
|
Report --> Persist[diagnosis_session.answer]
|
||||||
|
Planner --> Steps[agent_step]
|
||||||
|
Executor --> Steps
|
||||||
|
Tools --> Invocations[tool_invocation]
|
||||||
|
Persist --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
||||||
|
Steps --> Trace
|
||||||
|
Invocations --> Trace
|
||||||
|
```
|
||||||
|
|
||||||
|
关键代码:
|
||||||
|
|
||||||
|
- `ChatController.aiOps(...)`
|
||||||
|
- `AIOpsRequest`
|
||||||
|
- `AiOpsService.resolveSessionId(...)`
|
||||||
|
- `AiOpsService.buildTaskPrompt(...)`
|
||||||
|
- `AiOpsService.hasAlertPayload(...)`
|
||||||
|
- `AiOpsService.persistFinalReport(...)`
|
||||||
|
|
||||||
|
## Trace 数据模型
|
||||||
|
|
||||||
|
### `diagnosis_session`
|
||||||
|
|
||||||
|
记录一次诊断会话的主信息:
|
||||||
|
|
||||||
|
- `session_id`
|
||||||
|
- `query`
|
||||||
|
- `status`
|
||||||
|
- `agent_flow`
|
||||||
|
- `total_duration_ms`
|
||||||
|
- `total_token_count`
|
||||||
|
- `step_count`
|
||||||
|
- `tool_call_count`
|
||||||
|
- `answer`
|
||||||
|
- `self_evaluation`
|
||||||
|
- `feedback`
|
||||||
|
|
||||||
|
### `agent_step`
|
||||||
|
|
||||||
|
记录 Agent 模型调用过程:
|
||||||
|
|
||||||
|
- `session_id`
|
||||||
|
- `step_index`
|
||||||
|
- `agent_name`
|
||||||
|
- `model_input`
|
||||||
|
- `model_output`
|
||||||
|
- `thought`
|
||||||
|
- `has_tool_call`
|
||||||
|
- `duration_ms`
|
||||||
|
- `token_count`
|
||||||
|
|
||||||
|
### `tool_invocation`
|
||||||
|
|
||||||
|
记录真实工具调用:
|
||||||
|
|
||||||
|
- `session_id`
|
||||||
|
- `tool_name`
|
||||||
|
- `input_params`
|
||||||
|
- `output_preview`
|
||||||
|
- `output_length`
|
||||||
|
- `retrieval_layer`
|
||||||
|
- `relevance_level`
|
||||||
|
- `duration_ms`
|
||||||
|
- `success`
|
||||||
|
- `error_message`
|
||||||
|
|
||||||
|
## 为什么 trace 是核心
|
||||||
|
|
||||||
|
Agent 系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
|
||||||
|
|
||||||
|
- 模型为什么这么答
|
||||||
|
- 调了哪些工具
|
||||||
|
- 工具返回了什么证据
|
||||||
|
- Verifier 如何判断答案可信度
|
||||||
|
- 用户反馈如何回写到同一个 session
|
||||||
|
|
||||||
|
这就是项目区别于普通 Chatbot 的地方。
|
||||||
@@ -0,0 +1,132 @@
|
|||||||
|
# Interview Demo Script
|
||||||
|
|
||||||
|
## 30 秒开场
|
||||||
|
|
||||||
|
这是一个 Agent Engineering 项目,场景是企业故障诊断。它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。项目重点不是单次回答,而是把多 Agent 执行、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 trace。
|
||||||
|
|
||||||
|
## Demo 准备
|
||||||
|
|
||||||
|
启动服务:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||||
|
```
|
||||||
|
|
||||||
|
确认服务地址:
|
||||||
|
|
||||||
|
```text
|
||||||
|
http://localhost:9900
|
||||||
|
```
|
||||||
|
|
||||||
|
`mvp-demo` profile 下:
|
||||||
|
|
||||||
|
- Prometheus 告警使用 mock 数据。
|
||||||
|
- CLS 日志使用 mock 数据。
|
||||||
|
- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。
|
||||||
|
|
||||||
|
## Demo 1: Chat 诊断
|
||||||
|
|
||||||
|
目标:展示普通用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 trace。
|
||||||
|
|
||||||
|
请求:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
$sessionId = "interview-chat-payment-timeout-001"
|
||||||
|
$body = @{
|
||||||
|
Id = $sessionId
|
||||||
|
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "http://localhost:9900/api/chat" `
|
||||||
|
-ContentType "application/json" `
|
||||||
|
-Body $body
|
||||||
|
```
|
||||||
|
|
||||||
|
讲解点:
|
||||||
|
|
||||||
|
- `ChatController` 把请求交给 `ChatService.executeChatWithStrategy(...)`。
|
||||||
|
- 简单问题走单 ReactAgent,复杂问题走 `Planner -> Executor -> Verifier`。
|
||||||
|
- Executor 可以调用知识库、日志、指标等工具。
|
||||||
|
- Verifier 会基于工具证据生成 groundedness 评估。
|
||||||
|
- 最终会写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
|
||||||
|
|
||||||
|
查询 trace:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Get `
|
||||||
|
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||||
|
```
|
||||||
|
|
||||||
|
展示点:
|
||||||
|
|
||||||
|
- `data.session.agentFlow = CHAT`
|
||||||
|
- `data.steps` 中能看到 planner/executor/verifier
|
||||||
|
- `data.toolInvocations` 中能看到证据工具
|
||||||
|
- `data.session.selfEvaluation` 中有 verifier 结果
|
||||||
|
|
||||||
|
## Demo 2: AIOps 告警诊断
|
||||||
|
|
||||||
|
目标:展示告警 payload 如何触发 AIOps 入口,并且报告只聚焦目标告警。
|
||||||
|
|
||||||
|
请求:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
$aiopsSessionId = "interview-aiops-payment-cpu-001"
|
||||||
|
$aiopsBody = @{
|
||||||
|
sessionId = $aiopsSessionId
|
||||||
|
alertName = "HighCPUUsage"
|
||||||
|
service = "payment-service"
|
||||||
|
severity = "P1"
|
||||||
|
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||||
|
timeRange = "last_15m"
|
||||||
|
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
Invoke-WebRequest `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "http://localhost:9900/api/ai_ops" `
|
||||||
|
-ContentType "application/json" `
|
||||||
|
-Body $aiopsBody
|
||||||
|
```
|
||||||
|
|
||||||
|
讲解点:
|
||||||
|
|
||||||
|
- `/api/ai_ops` 接受可选 `AIOpsRequest`。
|
||||||
|
- 首条 SSE 消息会返回 `type=session`。
|
||||||
|
- `AiOpsService` 根据 payload 判断模式:
|
||||||
|
- `PAYLOAD_TARGETED`:聚焦传入告警。
|
||||||
|
- `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。
|
||||||
|
- AIOps 暂时不加 Verifier,先保证告警入口、证据工具和 trace 可用。
|
||||||
|
|
||||||
|
查询 trace:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Get `
|
||||||
|
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||||
|
```
|
||||||
|
|
||||||
|
展示点:
|
||||||
|
|
||||||
|
- `data.session.agentFlow = AI_OPS`
|
||||||
|
- `data.session.answer` 有最终告警报告
|
||||||
|
- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge`
|
||||||
|
- 报告有 `HighCPUUsage/payment-service` 的完整根因分析
|
||||||
|
- 其他 active alerts 只作为相关风险出现,不展开成独立根因章节
|
||||||
|
|
||||||
|
## MySQL 验证
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session ORDER BY id DESC LIMIT 5"
|
||||||
|
```
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name"
|
||||||
|
```
|
||||||
|
|
||||||
|
## 收尾总结
|
||||||
|
|
||||||
|
这套 Demo 展示的是一个完整 Agent 系统,而不是一次模型问答:入口有明确场景边界,Agent 负责规划和执行,工具提供证据,Verifier 提供质量门,trace API 提供审计和复盘能力。AIOps 入口进一步证明它可以从用户问答扩展到事件驱动诊断。
|
||||||
@@ -0,0 +1,99 @@
|
|||||||
|
# Design Tradeoffs
|
||||||
|
|
||||||
|
## 1. 为什么要做 trace,而不是只返回答案
|
||||||
|
|
||||||
|
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答三件事:
|
||||||
|
|
||||||
|
- 结论是什么
|
||||||
|
- 证据来自哪里
|
||||||
|
- 哪些步骤由哪个 Agent 完成
|
||||||
|
|
||||||
|
因此项目把一次会话拆成:
|
||||||
|
|
||||||
|
- `diagnosis_session`:会话级摘要、最终答案、质量评估、反馈。
|
||||||
|
- `agent_step`:Agent 模型输入输出、耗时、token 和工具调用标记。
|
||||||
|
- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。
|
||||||
|
|
||||||
|
这个设计牺牲了一些实现复杂度,但换来了可回放、可调试、可演示。
|
||||||
|
|
||||||
|
## 2. 为什么 Chat 有 Verifier,AIOps 暂时没有
|
||||||
|
|
||||||
|
Chat 入口的问题更开放,用户可能要求复杂推理或跨领域结论,所以 Verifier 是必要的质量门。当前 Chat 链路通过 `Planner -> Executor -> Verifier` 固定流程,把 groundedness 和 facts checked 写入 `self_evaluation`。
|
||||||
|
|
||||||
|
AIOps 当前阶段先不加 Verifier,原因是:
|
||||||
|
|
||||||
|
- AIOps 刚完成从“自动跑告警”到“可追踪告警入口”的改造。
|
||||||
|
- 先要确认告警 payload、工具证据、最终报告和 trace 能闭环。
|
||||||
|
- AIOps Verifier 的规则不同于 Chat Verifier,需要检查告警 scope、证据覆盖和处置建议,不宜直接复用。
|
||||||
|
|
||||||
|
后续可以做 lightweight AIOps Verifier,检查报告是否聚焦 payload、是否引用工具证据、是否误展开无关告警。
|
||||||
|
|
||||||
|
## 3. 为什么 AIOps payload scope 先用 prompt 控制
|
||||||
|
|
||||||
|
运行验证发现:传入 `HighCPUUsage/payment-service` 后,Agent 仍可能把 mock Prometheus 返回的所有 active alerts 都展开分析。这个问题的本质是任务边界不清晰。
|
||||||
|
|
||||||
|
当前选择 prompt-level scope control:
|
||||||
|
|
||||||
|
- 有 payload:`PAYLOAD_TARGETED`,最终报告围绕传入告警。
|
||||||
|
- 无 payload:`AUTO_DISCOVERY`,先调用 `queryPrometheusAlerts` 自动发现告警。
|
||||||
|
|
||||||
|
没有先做 Java 侧过滤,是因为:
|
||||||
|
|
||||||
|
- 过滤工具结果会降低 Agent 发现关联风险的能力。
|
||||||
|
- 目前需要的是报告主线聚焦,而不是完全屏蔽上下文。
|
||||||
|
- Prompt 改动小,风险低,能保留 Agent 灵活性。
|
||||||
|
|
||||||
|
已验证结果:主报告有 `HighCPUUsage/payment-service` 的完整根因分析,`HighMemoryUsage` 和 `SlowResponse` 只作为相关风险出现。
|
||||||
|
|
||||||
|
## 4. 为什么用 `tool_invocation` 统计真实工具调用次数
|
||||||
|
|
||||||
|
早期可以通过 `agent_step.hasToolCall` 粗略判断是否调用工具,但它统计的是“哪些模型步骤包含工具调用”,不是“真实调用了几次工具”。
|
||||||
|
|
||||||
|
现在 `tool_call_count` 来自:
|
||||||
|
|
||||||
|
```text
|
||||||
|
ToolInvocationRepository.countBySessionId(sessionId)
|
||||||
|
```
|
||||||
|
|
||||||
|
这样更符合 trace 语义:
|
||||||
|
|
||||||
|
- 一个 step 可能调用多个工具。
|
||||||
|
- 工具可能来自不同来源:知识库、日志、指标、Prometheus。
|
||||||
|
- 面试时可以把 `tool_call_count` 和 trace 中返回的工具明细对上。
|
||||||
|
|
||||||
|
## 5. 为什么保留 mock Prometheus 和 mock CLS
|
||||||
|
|
||||||
|
面试 Demo 最怕不稳定。真实 Prometheus、日志平台和线上故障都有不可控因素,所以 MVP profile 保留 mock 工具:
|
||||||
|
|
||||||
|
- `prometheus.mock-enabled=true`
|
||||||
|
- `cls.mock-enabled=true`
|
||||||
|
|
||||||
|
这样可以稳定复现:
|
||||||
|
|
||||||
|
- `HighCPUUsage/payment-service`
|
||||||
|
- `HighMemoryUsage/order-service`
|
||||||
|
- `SlowResponse/user-service`
|
||||||
|
- system-metrics、application-logs、database-slow-query 等日志证据
|
||||||
|
|
||||||
|
这不是逃避真实集成,而是把“Agent 编排和证据追踪”作为面试演示的主目标。
|
||||||
|
|
||||||
|
## 6. 为什么把面试材料单独放 `interview/`
|
||||||
|
|
||||||
|
`mvp/` 是持续迭代现场,包含过程文档、验收记录和 runbook。面试材料的目标不同,它应该是可讲、可演示、可评估的展示层。
|
||||||
|
|
||||||
|
因此:
|
||||||
|
|
||||||
|
- `mvp/` 保留真实演进材料。
|
||||||
|
- `devflow/` 保留决策沉淀。
|
||||||
|
- `interview/` 只组织面试叙事和演示脚本。
|
||||||
|
|
||||||
|
这样后续继续做 AIOps Verifier、UI、更多工具集成时,不会污染面试讲稿。
|
||||||
|
|
||||||
|
## 7. 可以主动承认的限制
|
||||||
|
|
||||||
|
- AIOps 还没有 Verifier。
|
||||||
|
- Prompt-level scope control 不能做到强约束,只能通过 trace 和测试观察遵循情况。
|
||||||
|
- 当前 mock 数据适合 demo,不代表生产接入已经完成。
|
||||||
|
- Hikari 连接池已经加了短生命周期和 keepalive,但真实生产还需要按数据库 wait_timeout 和连接数预算调优。
|
||||||
|
|
||||||
|
主动讲清这些限制,反而能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
|
||||||
Reference in New Issue
Block a user