Files

7.4 KiB

Current MVP Architecture Snapshot

Updated: 2026-07-05

This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.

1. Positioning

The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.

Core goals:

  • Support normal chat-based diagnosis.
  • Support AIOps alert-triggered diagnosis.
  • Keep tool calls explicit and traceable.
  • Keep RAG retrieval observable through lookup_knowledge.
  • Persist enough execution evidence for replay, evaluation, and interview explanation.

2. Runtime Architecture

HTTP API
  -> ChatService / AiOpsService
      -> Agent orchestration
          -> Supervisor / Planner / Executor / Verifier
      -> Tools
          -> lookup_knowledge
          -> query_logs
          -> query_metrics
          -> other diagnosis tools
      -> Persistence
          -> diagnosis_session
          -> agent_step
          -> tool_invocation
      -> Trace API
          -> DiagnosisTraceService

Current entry points:

  • ChatService: user-driven troubleshooting and follow-up diagnosis.
  • AiOpsService: alert-driven diagnosis, including payload mode and auto-discovery mode.
  • DiagnosisTraceService: trace view of session, steps, tool calls, and self-evaluation.

3. Chat Diagnosis Flow

User question
  -> ChatService
  -> simple response or diagnosis flow
  -> Planner creates investigation direction
  -> Executor calls tools for evidence
      -> lookup_knowledge
      -> query_logs
      -> query_metrics
  -> Verifier checks final diagnosis quality
  -> self_evaluation.verifier_evaluation
  -> diagnosis trace

The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under diagnosis_session.self_evaluation.verifier_evaluation.

4. AIOps Diagnosis Flow

AIOps request
  -> AiOpsService
  -> payload mode or auto-discovery mode
  -> build alert-focused diagnosis prompt
  -> append recommended lookup_knowledge query when payload exists
  -> Agent diagnosis flow
      -> Supervisor / Planner / Executor
      -> evidence tools
  -> final report
  -> AiOpsRuleEvaluationService
  -> self_evaluation.aiops_rule_evaluation
  -> diagnosis trace

AIOps keeps two modes:

  • Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
  • Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.

The AIOps verifier is currently lightweight and rule-based. It checks:

  • Whether the final report exists.
  • Whether the result stays focused on the alert payload when payload exists.
  • Whether evidence tools were used, especially lookup_knowledge, query_logs, and query_metrics.

5. RAG Architecture

lookup_knowledge
  -> L0 domain/entity hint
      -> matched domain
      -> matched keywords/entities
      -> metadata filter signal
  -> VectorSearchService
      -> Spring AI VectorStore path
      -> Milvus SDK fallback path
  -> evidence post-processing
      -> score / rawScore / scoreLabel
      -> source metadata
      -> title / breadcrumb / content evidence block
  -> tool_invocation record

Important decisions:

  • lookup_knowledge remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
  • L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
  • L1 retrieval now goes through VectorSearchService.
  • Spring AI VectorStore is the preferred retrieval path.
  • The original Milvus SDK path is retained as fallback and compatibility path.
  • title, breadcrumb, and content participate in embedding text so chunk context is less likely to be lost.
  • Retrieval output keeps compatibility fields: score, rawScore, and scoreLabel.

Vector retrieval modes:

retrieval.vector-store.mode=auto       # Prefer Spring AI VectorStore, fallback to SDK
retrieval.vector-store.mode=spring-ai  # Use Spring AI VectorStore only
retrieval.vector-store.mode=sdk        # Use original Milvus SDK path

6. Persistence And Trace

Current trace-related persistence:

diagnosis_session
  -> final_report
  -> self_evaluation
      -> verifier_evaluation
      -> aiops_rule_evaluation

agent_step
  -> role
  -> step input/output
  -> execution order

tool_invocation
  -> tool_name
  -> query
  -> retrieval_layer
  -> retrieval_details
  -> evidence blocks
  -> duration

Trace API aggregates these records into a session-level view:

  • Agent step sequence.
  • Tool calls and retrieval details.
  • Final diagnosis report.
  • Chat verifier status.
  • AIOps rule verifier status.

7. Quality Gates

Current quality gates:

  • Chat verifier: LLM-based final answer verification for normal diagnosis.
  • AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
  • Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
  • RAG retrieval baseline: golden query set with offline baseline report.
  • Live RAG acceptance: post-reindex script for validating retrieval against the running stack.

These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.

8. Current Completion State

Completed for the current MVP stage:

  • Explicit lookup_knowledge Agent tool.
  • L0 + L1 retrieval shape retained.
  • L0 downgraded to domain/entity hint.
  • Spring AI VectorStore retrieval path integrated.
  • Milvus SDK fallback retained.
  • RAG evidence post-processing added.
  • Breadcrumb/title/content embedding text improved.
  • RAG offline baseline and live acceptance script added.
  • AIOps payload query augmentation added.
  • AIOps lightweight verifier added.
  • Trace summary includes both chat verifier and AIOps verifier signals.

Deferred future enhancements:

  • LLM QueryTransformer / MultiQuery.
  • BM25, RRF, and reranker.
  • Neighbor chunk or section-level context expansion.
  • VectorStore write path migration.
  • Full LLM-based AIOps verifier.
  • More complete golden set for recall, MRR, and nDCG metrics.

9. Key Code References

  • src/main/java/com/superbiz/agent/service/ChatService.java
  • src/main/java/com/superbiz/agent/service/AiOpsService.java
  • src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java
  • src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java
  • src/main/java/com/superbiz/agent/service/VectorSearchService.java
  • src/main/java/com/superbiz/agent/service/VectorIndexService.java
  • src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java
  • src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java
  • src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java
  • src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java

10. Supporting Materials

  • mvp/issues/rag-refactor-plan.md
  • eval/rag-retrieval/README.md
  • scripts/eval_rag_live_acceptance.py
  • interview/rag-refactor-story.md
  • interview/rag-vectorstore-interview-notes.md
  • interview/rag-retrieval-quality-report.md
  • interview/rag-breadcrumb-embedding-acceptance.md
  • interview/aiops-query-augmentation.md
  • interview/aiops-lightweight-verifier.md