3.8 KiB
3.8 KiB
Decisions: aiops-traceable-diagnosis-entry
Clarify
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
- Slug:
aiops-traceable-diagnosis-entry - Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
Context
mvp-demo-trace-acceptancealready addedGET /api/diagnosis/{sessionId}/trace.chat-verifier-agentmade the chat path stronger than the older AIOps path.- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
| Q2 | API | Should we keep /api/ai_ops or add a new endpoint? |
evidence-driven | Resolved: keep existing endpoint and extend optional body |
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: sessionId, alertName, service, severity, description, timeRange, plus userRequest fallback |
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type session |
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| AIOps is currently isolated from request-driven trace replay. | ChatController.aiOps() has no request body; AiOpsService creates its own random session id. |
Extend endpoint and service. |
| No schema change is needed. | DiagnosisSession already has query, agentFlow, answer, counts, and status. |
Reuse existing table. |
| Trace API can already replay AIOps if session id and answer are persisted. | DiagnosisTraceService loads by session id and is flow-agnostic. |
Keep trace API unchanged. |
| Blast radius is moderate and local. | rg shows only ChatController calls executeAiOpsAnalysis and extractFinalReport. |
Change service/controller carefully and add tests. |
User-Interview Confirmations
| Topic | User Words | Decision |
|---|---|---|
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
Key Decisions
- Keep
/api/ai_opsand make its body optional. - Emit
SseMessage.type=sessionbefore long-running analysis starts. - Store AIOps request summary in
diagnosis_session.query. - Store final report in
diagnosis_session.answer. - Defer AIOps Verifier to a later change so this slice stays focused.
Architecture Audit
POST /api/ai_ops
-> optional AIOpsRequest
-> resolve sessionId
-> create diagnosis_session(agentFlow=AI_OPS)
-> run ai_ops_supervisor(planner, executor)
-> AgentLoggingHook persists steps
-> tools persist invocations under SessionContextHolder
-> extract final report
-> persist answer
-> GET /api/diagnosis/{sessionId}/trace replays the run
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.