Files

3.8 KiB

Decisions: aiops-traceable-diagnosis-entry

Clarify

  • Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
  • Slug: aiops-traceable-diagnosis-entry
  • Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.

Context

  • mvp-demo-trace-acceptance already added GET /api/diagnosis/{sessionId}/trace.
  • chat-verifier-agent made the chat path stronger than the older AIOps path.
  • Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.

Grill Question Pool

# Dimension Question Mode Status
Q1 Positioning Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? user-interview Resolved: sibling entry, unified trace story
Q2 API Should we keep /api/ai_ops or add a new endpoint? evidence-driven Resolved: keep existing endpoint and extend optional body
Q3 Input What is the minimum alert payload? user-interview Resolved: sessionId, alertName, service, severity, description, timeRange, plus userRequest fallback
Q4 Output How does the caller learn the trace session id? evidence-driven Resolved: first SSE event uses type session
Q5 Trace Must AIOps be replayable with existing trace API? evidence-driven Resolved: yes, this is the main acceptance criterion
Q6 Verifier Must this slice add AIOps Verifier? user-interview Resolved: no, defer as follow-up
Q7 Compatibility Should no-body calls still work? evidence-driven Resolved: yes, preserve old demo behavior
Q8 GitNexus Should unavailable GitNexus block implementation? user-interview Resolved: skip GitNexus by user decision

Evidence-Driven Conclusions

Conclusion Evidence Source Result
AIOps is currently isolated from request-driven trace replay. ChatController.aiOps() has no request body; AiOpsService creates its own random session id. Extend endpoint and service.
No schema change is needed. DiagnosisSession already has query, agentFlow, answer, counts, and status. Reuse existing table.
Trace API can already replay AIOps if session id and answer are persisted. DiagnosisTraceService loads by session id and is flow-agnostic. Keep trace API unchanged.
Blast radius is moderate and local. rg shows only ChatController calls executeAiOpsAnalysis and extractFinalReport. Change service/controller carefully and add tests.

User-Interview Confirmations

Topic User Words Decision
Use sm-flow "可以,改造一下AIOps 接口,用sm-flow流程看看" Use OpenSpec + devflow.
GitNexus "跳过gitnexus把" Record skip and use local impact analysis.
Proceed after Grill "可以" Continue with lightweight Grill conclusions.

Key Decisions

  • Keep /api/ai_ops and make its body optional.
  • Emit SseMessage.type=session before long-running analysis starts.
  • Store AIOps request summary in diagnosis_session.query.
  • Store final report in diagnosis_session.answer.
  • Defer AIOps Verifier to a later change so this slice stays focused.

Architecture Audit

POST /api/ai_ops
  -> optional AIOpsRequest
  -> resolve sessionId
  -> create diagnosis_session(agentFlow=AI_OPS)
  -> run ai_ops_supervisor(planner, executor)
  -> AgentLoggingHook persists steps
  -> tools persist invocations under SessionContextHolder
  -> extract final report
  -> persist answer
  -> GET /api/diagnosis/{sessionId}/trace replays the run

Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.