diff --git a/interview/README.md b/interview/README.md new file mode 100644 index 0000000..38daf47 --- /dev/null +++ b/interview/README.md @@ -0,0 +1,59 @@ +# SuperBizAgent Interview Guide + +## 一句话定位 + +SuperBizAgent 是一个面向企业故障诊断场景的 Agent Engineering 项目:它把用户问题或告警事件转成可追踪的多 Agent 执行链路,并把工具证据、模型步骤、最终答案和反馈统一落到诊断 trace 中。 + +## 面试重点 + +- **多 Agent 编排**:普通 Chat 的复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 Supervisor 调度 Planner/Executor。 +- **工具证据链**:知识库、日志、指标和 Prometheus 告警都通过工具调用进入链路,并记录到 `tool_invocation`。 +- **可追踪诊断**:一次会话对应一个 `sessionId`,最终可以通过 `GET /api/diagnosis/{sessionId}/trace` 回放。 +- **质量门**:Chat 链路包含 Verifier,把 groundedness、facts checked 和 evidence refs 写回 `diagnosis_session.self_evaluation`。 +- **AIOps 产品边界**:有告警 payload 时聚焦该告警;没有 payload 时先自动发现 active alerts。 +- **可复现 Demo**:`mvp-demo` profile 使用 mock Prometheus 和 mock CLS,让面试演示不依赖真实线上故障。 + +## 推荐阅读顺序 + +1. `interview/demo-script.md`:面试现场怎么讲、怎么演示。 +2. `interview/architecture.md`:系统架构和两条主链路。 +3. `interview/design-tradeoffs.md`:关键设计取舍和可被追问的问题。 +4. `interview/acceptance-checklist.md`:面试前验证清单。 +5. `mvp/demo/README.md`:更细的 MVP 可执行 runbook。 + +## 核心 Demo + +### Chat Diagnosis + +```text +POST /api/chat +-> ChatService.executeChatWithStrategy(...) +-> simple ReactAgent or Planner -> Executor -> Verifier +-> lookup_knowledge / query_logs / query_metrics +-> diagnosis_session + agent_step + tool_invocation +-> GET /api/diagnosis/{sessionId}/trace +``` + +### AIOps Alert Diagnosis + +```text +POST /api/ai_ops +-> AiOpsService.executeAiOpsAnalysis(...) +-> ai_ops_supervisor +-> planner_agent / executor_agent +-> queryPrometheusAlerts + logs + knowledge +-> scoped alert report +-> GET /api/diagnosis/{sessionId}/trace +``` + +## 当前完成度 + +- Chat 诊断链路:可运行、可追踪、有 Verifier。 +- AIOps 告警链路:可运行、可追踪、支持 payload scope control。 +- Trace API:统一返回 session、agent steps、tool invocations 和 summary。 +- Demo 文档:`mvp/demo/README.md` 和 `mvp/demo/aiops-alert-acceptance.md`。 +- Devflow 沉淀:`devflow/index.md` 记录了 MVP、Verifier、AIOps trace 和 AIOps scope-control 的演进。 + +## 面试时的主叙事 + +这个项目不是简单调用大模型,而是在做一个可审计的 Agent 诊断系统。核心价值是:模型可以规划和推理,但每一步工具证据、最终结论和质量评估都能被 trace API 回放。面试时重点展示“从问题到证据到答案到验证”的完整闭环。 diff --git a/interview/acceptance-checklist.md b/interview/acceptance-checklist.md new file mode 100644 index 0000000..e1344ea --- /dev/null +++ b/interview/acceptance-checklist.md @@ -0,0 +1,170 @@ +# Acceptance Checklist + +## 面试前环境检查 + +- 当前分支包含最新 AIOps trace/scope 变更。 +- MySQL 可连接。 +- Redis 可连接。 +- Milvus/Zilliz 可连接。 +- 模型 API key 可用。 +- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。 + +启动: + +```powershell +mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo" +``` + +编译检查: + +```powershell +mvn -q -DskipTests compile +``` + +目标测试: + +```powershell +mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test +``` + +## Chat Demo 验收 + +请求: + +```powershell +$sessionId = "interview-chat-payment-timeout-001" +$body = @{ + Id = $sessionId + Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。" +} | ConvertTo-Json + +Invoke-RestMethod ` + -Method Post ` + -Uri "http://localhost:9900/api/chat" ` + -ContentType "application/json" ` + -Body $body +``` + +验收: + +- 返回 `data.success = true`。 +- 返回 `data.sessionId = interview-chat-payment-timeout-001`。 +- `diagnosis_session.agent_flow = CHAT`。 +- trace API 返回 session、steps、toolInvocations。 +- 复杂问题下 trace 中能看到 verifier 相关数据。 + +SQL: + +```powershell +python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'" +``` + +## AIOps Demo 验收 + +请求: + +```powershell +$aiopsSessionId = "interview-aiops-payment-cpu-001" +$aiopsBody = @{ + sessionId = $aiopsSessionId + alertName = "HighCPUUsage" + service = "payment-service" + severity = "P1" + description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。" + timeRange = "last_15m" + userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。" +} | ConvertTo-Json + +Invoke-WebRequest ` + -Method Post ` + -Uri "http://localhost:9900/api/ai_ops" ` + -ContentType "application/json" ` + -Body $aiopsBody +``` + +验收: + +- SSE 首条包含 `type=session`。 +- SSE 最后包含 `type=done`。 +- `diagnosis_session.agent_flow = AI_OPS`。 +- `diagnosis_session.status = SUCCESS`。 +- `diagnosis_session.answer` 有最终报告。 +- trace API 返回 AIOps steps 和 tool invocations。 +- 报告主章节聚焦 `HighCPUUsage/payment-service`。 +- 无 `告警根因分析 - HighMemoryUsage` 独立章节。 +- 无 `告警根因分析 - SlowResponse` 独立章节。 +- 有“相关风险告警”或类似上下文说明。 + +SQL: + +```powershell +python scripts/query_mysql.py "SELECT session_id, agent_flow, status, total_duration_ms, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'" +``` + +```powershell +python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name ORDER BY tool_name" +``` + +Scope 检查: + +```powershell +python scripts/query_mysql.py "SELECT (answer LIKE '%告警根因分析 - HighCPUUsage%') AS has_main_root_cause, (answer LIKE '%告警根因分析 - HighMemoryUsage%') AS has_memory_root_cause, (answer LIKE '%告警根因分析 - SlowResponse%') AS has_slow_root_cause, (answer LIKE '%相关风险告警%') AS has_related_risk FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'" +``` + +期望: + +```text +has_main_root_cause = 1 +has_memory_root_cause = 0 +has_slow_root_cause = 0 +has_related_risk = 1 +``` + +## Trace API 验收 + +```powershell +Invoke-RestMethod ` + -Method Get ` + -Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace" +``` + +若 PowerShell 对长 JSON 或特殊字符不稳定,可以用: + +```powershell +curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace" +``` + +## 常见问题 + +### MySQL stale connection + +现象: + +```text +HikariPool - Connection is not available +No operations allowed after connection closed +``` + +当前已在 `application.yml` 配置: + +- `maximum-pool-size: 5` +- `minimum-idle: 1` +- `connection-timeout: 10000` +- `validation-timeout: 5000` +- `idle-timeout: 60000` +- `max-lifetime: 120000` +- `keepalive-time: 30000` + +处理: + +- 重新编译或重启服务。 +- 确认日志中新的 HikariPool 启动成功。 +- 再跑 trace 或 AIOps 请求。 + +### SSE 客户端显示异常 + +PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 trace API 验证结果。 + +### OpenSpec 全量校验失败 + +`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试材料主要依赖已归档的 AIOps spec 和 MVP trace spec,可以单独验证相关 spec。 diff --git a/interview/architecture.md b/interview/architecture.md new file mode 100644 index 0000000..4afd2b7 --- /dev/null +++ b/interview/architecture.md @@ -0,0 +1,147 @@ +# Architecture + +## 系统分层 + +```text +API Layer +-> ChatController / DiagnosisTraceController + +Agent Orchestration +-> ChatService / AiOpsService + +Tools +-> lookupKnowledgeTool / queryLogs / queryMetrics / queryPrometheusAlerts + +Persistence +-> diagnosis_session / agent_step / tool_invocation + +Trace +-> GET /api/diagnosis/{sessionId}/trace +``` + +## Chat 链路 + +```mermaid +flowchart TD + User[User Question] --> ChatAPI[POST /api/chat] + ChatAPI --> Strategy[ChatService.executeChatWithStrategy] + Strategy --> Complexity{QuestionComplexity} + Complexity -->|simple| Single[ReactAgent] + Complexity -->|complex| Planner[Planner Agent] + Planner --> Executor[Executor Agent] + Executor --> Tools[Evidence Tools] + Tools --> Executor + Executor --> Verifier[Verifier Agent] + Verifier --> Answer[Final Answer] + Answer --> Session[diagnosis_session] + Planner --> Steps[agent_step] + Executor --> Steps + Verifier --> Steps + Tools --> Invocations[tool_invocation] + Session --> Trace[GET /api/diagnosis/{sessionId}/trace] + Steps --> Trace + Invocations --> Trace +``` + +关键代码: + +- `ChatController.chat(...)` +- `ChatService.executeChatWithStrategy(...)` +- `ChatService.executeChatComplex(...)` +- `AgentLoggingHook` +- `ToolInvocationRecorder` +- `DiagnosisTraceService.getTrace(...)` + +## AIOps 链路 + +```mermaid +flowchart TD + Alert[Alert Payload or Empty Request] --> AiOpsAPI[POST /api/ai_ops] + AiOpsAPI --> SessionEvent[SSE session event] + AiOpsAPI --> AiOpsService[AiOpsService.executeAiOpsAnalysis] + AiOpsService --> PromptMode{Payload?} + PromptMode -->|yes| Targeted[PAYLOAD_TARGETED] + PromptMode -->|no| Discovery[AUTO_DISCOVERY] + Targeted --> Supervisor[ai_ops_supervisor] + Discovery --> Supervisor + Supervisor --> Planner[planner_agent] + Supervisor --> Executor[executor_agent] + Planner --> Tools[Prometheus / Logs / Knowledge] + Executor --> Tools + Tools --> Report[Alert Report] + Report --> Persist[diagnosis_session.answer] + Planner --> Steps[agent_step] + Executor --> Steps + Tools --> Invocations[tool_invocation] + Persist --> Trace[GET /api/diagnosis/{sessionId}/trace] + Steps --> Trace + Invocations --> Trace +``` + +关键代码: + +- `ChatController.aiOps(...)` +- `AIOpsRequest` +- `AiOpsService.resolveSessionId(...)` +- `AiOpsService.buildTaskPrompt(...)` +- `AiOpsService.hasAlertPayload(...)` +- `AiOpsService.persistFinalReport(...)` + +## Trace 数据模型 + +### `diagnosis_session` + +记录一次诊断会话的主信息: + +- `session_id` +- `query` +- `status` +- `agent_flow` +- `total_duration_ms` +- `total_token_count` +- `step_count` +- `tool_call_count` +- `answer` +- `self_evaluation` +- `feedback` + +### `agent_step` + +记录 Agent 模型调用过程: + +- `session_id` +- `step_index` +- `agent_name` +- `model_input` +- `model_output` +- `thought` +- `has_tool_call` +- `duration_ms` +- `token_count` + +### `tool_invocation` + +记录真实工具调用: + +- `session_id` +- `tool_name` +- `input_params` +- `output_preview` +- `output_length` +- `retrieval_layer` +- `relevance_level` +- `duration_ms` +- `success` +- `error_message` + +## 为什么 trace 是核心 + +Agent 系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到: + +- 模型为什么这么答 +- 调了哪些工具 +- 工具返回了什么证据 +- Verifier 如何判断答案可信度 +- 用户反馈如何回写到同一个 session + +这就是项目区别于普通 Chatbot 的地方。 diff --git a/interview/demo-script.md b/interview/demo-script.md new file mode 100644 index 0000000..44de57d --- /dev/null +++ b/interview/demo-script.md @@ -0,0 +1,132 @@ +# Interview Demo Script + +## 30 秒开场 + +这是一个 Agent Engineering 项目,场景是企业故障诊断。它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。项目重点不是单次回答,而是把多 Agent 执行、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 trace。 + +## Demo 准备 + +启动服务: + +```powershell +mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo" +``` + +确认服务地址: + +```text +http://localhost:9900 +``` + +`mvp-demo` profile 下: + +- Prometheus 告警使用 mock 数据。 +- CLS 日志使用 mock 数据。 +- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。 + +## Demo 1: Chat 诊断 + +目标:展示普通用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 trace。 + +请求: + +```powershell +$sessionId = "interview-chat-payment-timeout-001" +$body = @{ + Id = $sessionId + Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。" +} | ConvertTo-Json + +Invoke-RestMethod ` + -Method Post ` + -Uri "http://localhost:9900/api/chat" ` + -ContentType "application/json" ` + -Body $body +``` + +讲解点: + +- `ChatController` 把请求交给 `ChatService.executeChatWithStrategy(...)`。 +- 简单问题走单 ReactAgent,复杂问题走 `Planner -> Executor -> Verifier`。 +- Executor 可以调用知识库、日志、指标等工具。 +- Verifier 会基于工具证据生成 groundedness 评估。 +- 最终会写入 `diagnosis_session`、`agent_step`、`tool_invocation`。 + +查询 trace: + +```powershell +Invoke-RestMethod ` + -Method Get ` + -Uri "http://localhost:9900/api/diagnosis/$sessionId/trace" +``` + +展示点: + +- `data.session.agentFlow = CHAT` +- `data.steps` 中能看到 planner/executor/verifier +- `data.toolInvocations` 中能看到证据工具 +- `data.session.selfEvaluation` 中有 verifier 结果 + +## Demo 2: AIOps 告警诊断 + +目标:展示告警 payload 如何触发 AIOps 入口,并且报告只聚焦目标告警。 + +请求: + +```powershell +$aiopsSessionId = "interview-aiops-payment-cpu-001" +$aiopsBody = @{ + sessionId = $aiopsSessionId + alertName = "HighCPUUsage" + service = "payment-service" + severity = "P1" + description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。" + timeRange = "last_15m" + userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。" +} | ConvertTo-Json + +Invoke-WebRequest ` + -Method Post ` + -Uri "http://localhost:9900/api/ai_ops" ` + -ContentType "application/json" ` + -Body $aiopsBody +``` + +讲解点: + +- `/api/ai_ops` 接受可选 `AIOpsRequest`。 +- 首条 SSE 消息会返回 `type=session`。 +- `AiOpsService` 根据 payload 判断模式: + - `PAYLOAD_TARGETED`:聚焦传入告警。 + - `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。 +- AIOps 暂时不加 Verifier,先保证告警入口、证据工具和 trace 可用。 + +查询 trace: + +```powershell +Invoke-RestMethod ` + -Method Get ` + -Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace" +``` + +展示点: + +- `data.session.agentFlow = AI_OPS` +- `data.session.answer` 有最终告警报告 +- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge` +- 报告有 `HighCPUUsage/payment-service` 的完整根因分析 +- 其他 active alerts 只作为相关风险出现,不展开成独立根因章节 + +## MySQL 验证 + +```powershell +python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session ORDER BY id DESC LIMIT 5" +``` + +```powershell +python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name" +``` + +## 收尾总结 + +这套 Demo 展示的是一个完整 Agent 系统,而不是一次模型问答:入口有明确场景边界,Agent 负责规划和执行,工具提供证据,Verifier 提供质量门,trace API 提供审计和复盘能力。AIOps 入口进一步证明它可以从用户问答扩展到事件驱动诊断。 diff --git a/interview/design-tradeoffs.md b/interview/design-tradeoffs.md new file mode 100644 index 0000000..5edb801 --- /dev/null +++ b/interview/design-tradeoffs.md @@ -0,0 +1,99 @@ +# Design Tradeoffs + +## 1. 为什么要做 trace,而不是只返回答案 + +普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答三件事: + +- 结论是什么 +- 证据来自哪里 +- 哪些步骤由哪个 Agent 完成 + +因此项目把一次会话拆成: + +- `diagnosis_session`:会话级摘要、最终答案、质量评估、反馈。 +- `agent_step`:Agent 模型输入输出、耗时、token 和工具调用标记。 +- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。 + +这个设计牺牲了一些实现复杂度,但换来了可回放、可调试、可演示。 + +## 2. 为什么 Chat 有 Verifier,AIOps 暂时没有 + +Chat 入口的问题更开放,用户可能要求复杂推理或跨领域结论,所以 Verifier 是必要的质量门。当前 Chat 链路通过 `Planner -> Executor -> Verifier` 固定流程,把 groundedness 和 facts checked 写入 `self_evaluation`。 + +AIOps 当前阶段先不加 Verifier,原因是: + +- AIOps 刚完成从“自动跑告警”到“可追踪告警入口”的改造。 +- 先要确认告警 payload、工具证据、最终报告和 trace 能闭环。 +- AIOps Verifier 的规则不同于 Chat Verifier,需要检查告警 scope、证据覆盖和处置建议,不宜直接复用。 + +后续可以做 lightweight AIOps Verifier,检查报告是否聚焦 payload、是否引用工具证据、是否误展开无关告警。 + +## 3. 为什么 AIOps payload scope 先用 prompt 控制 + +运行验证发现:传入 `HighCPUUsage/payment-service` 后,Agent 仍可能把 mock Prometheus 返回的所有 active alerts 都展开分析。这个问题的本质是任务边界不清晰。 + +当前选择 prompt-level scope control: + +- 有 payload:`PAYLOAD_TARGETED`,最终报告围绕传入告警。 +- 无 payload:`AUTO_DISCOVERY`,先调用 `queryPrometheusAlerts` 自动发现告警。 + +没有先做 Java 侧过滤,是因为: + +- 过滤工具结果会降低 Agent 发现关联风险的能力。 +- 目前需要的是报告主线聚焦,而不是完全屏蔽上下文。 +- Prompt 改动小,风险低,能保留 Agent 灵活性。 + +已验证结果:主报告有 `HighCPUUsage/payment-service` 的完整根因分析,`HighMemoryUsage` 和 `SlowResponse` 只作为相关风险出现。 + +## 4. 为什么用 `tool_invocation` 统计真实工具调用次数 + +早期可以通过 `agent_step.hasToolCall` 粗略判断是否调用工具,但它统计的是“哪些模型步骤包含工具调用”,不是“真实调用了几次工具”。 + +现在 `tool_call_count` 来自: + +```text +ToolInvocationRepository.countBySessionId(sessionId) +``` + +这样更符合 trace 语义: + +- 一个 step 可能调用多个工具。 +- 工具可能来自不同来源:知识库、日志、指标、Prometheus。 +- 面试时可以把 `tool_call_count` 和 trace 中返回的工具明细对上。 + +## 5. 为什么保留 mock Prometheus 和 mock CLS + +面试 Demo 最怕不稳定。真实 Prometheus、日志平台和线上故障都有不可控因素,所以 MVP profile 保留 mock 工具: + +- `prometheus.mock-enabled=true` +- `cls.mock-enabled=true` + +这样可以稳定复现: + +- `HighCPUUsage/payment-service` +- `HighMemoryUsage/order-service` +- `SlowResponse/user-service` +- system-metrics、application-logs、database-slow-query 等日志证据 + +这不是逃避真实集成,而是把“Agent 编排和证据追踪”作为面试演示的主目标。 + +## 6. 为什么把面试材料单独放 `interview/` + +`mvp/` 是持续迭代现场,包含过程文档、验收记录和 runbook。面试材料的目标不同,它应该是可讲、可演示、可评估的展示层。 + +因此: + +- `mvp/` 保留真实演进材料。 +- `devflow/` 保留决策沉淀。 +- `interview/` 只组织面试叙事和演示脚本。 + +这样后续继续做 AIOps Verifier、UI、更多工具集成时,不会污染面试讲稿。 + +## 7. 可以主动承认的限制 + +- AIOps 还没有 Verifier。 +- Prompt-level scope control 不能做到强约束,只能通过 trace 和测试观察遵循情况。 +- 当前 mock 数据适合 demo,不代表生产接入已经完成。 +- Hikari 连接池已经加了短生命周期和 keepalive,但真实生产还需要按数据库 wait_timeout 和连接数预算调优。 + +主动讲清这些限制,反而能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。