Compare commits
3
Commits
ed267d753d
...
b22f2d22c8
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b22f2d22c8 | ||
|
|
63b62b28a2 | ||
|
|
d5902a0499 |
@@ -1,5 +1,7 @@
|
||||
# 数据库设计文档
|
||||
|
||||
> 当前架构快照:[mvp/architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md)
|
||||
|
||||
## 📚 文档导航
|
||||
|
||||
### 核心表设计
|
||||
|
||||
@@ -0,0 +1,220 @@
|
||||
# Current MVP Architecture Snapshot
|
||||
|
||||
**Updated**: 2026-07-05
|
||||
|
||||
This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.
|
||||
|
||||
## 1. Positioning
|
||||
|
||||
The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.
|
||||
|
||||
Core goals:
|
||||
|
||||
- Support normal chat-based diagnosis.
|
||||
- Support AIOps alert-triggered diagnosis.
|
||||
- Keep tool calls explicit and traceable.
|
||||
- Keep RAG retrieval observable through `lookup_knowledge`.
|
||||
- Persist enough execution evidence for replay, evaluation, and interview explanation.
|
||||
|
||||
## 2. Runtime Architecture
|
||||
|
||||
```text
|
||||
HTTP API
|
||||
-> ChatService / AiOpsService
|
||||
-> Agent orchestration
|
||||
-> Supervisor / Planner / Executor / Verifier
|
||||
-> Tools
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> other diagnosis tools
|
||||
-> Persistence
|
||||
-> diagnosis_session
|
||||
-> agent_step
|
||||
-> tool_invocation
|
||||
-> Trace API
|
||||
-> DiagnosisTraceService
|
||||
```
|
||||
|
||||
Current entry points:
|
||||
|
||||
- `ChatService`: user-driven troubleshooting and follow-up diagnosis.
|
||||
- `AiOpsService`: alert-driven diagnosis, including payload mode and auto-discovery mode.
|
||||
- `DiagnosisTraceService`: trace view of session, steps, tool calls, and self-evaluation.
|
||||
|
||||
## 3. Chat Diagnosis Flow
|
||||
|
||||
```text
|
||||
User question
|
||||
-> ChatService
|
||||
-> simple response or diagnosis flow
|
||||
-> Planner creates investigation direction
|
||||
-> Executor calls tools for evidence
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> Verifier checks final diagnosis quality
|
||||
-> self_evaluation.verifier_evaluation
|
||||
-> diagnosis trace
|
||||
```
|
||||
|
||||
The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||
|
||||
## 4. AIOps Diagnosis Flow
|
||||
|
||||
```text
|
||||
AIOps request
|
||||
-> AiOpsService
|
||||
-> payload mode or auto-discovery mode
|
||||
-> build alert-focused diagnosis prompt
|
||||
-> append recommended lookup_knowledge query when payload exists
|
||||
-> Agent diagnosis flow
|
||||
-> Supervisor / Planner / Executor
|
||||
-> evidence tools
|
||||
-> final report
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> self_evaluation.aiops_rule_evaluation
|
||||
-> diagnosis trace
|
||||
```
|
||||
|
||||
AIOps keeps two modes:
|
||||
|
||||
- Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
|
||||
- Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.
|
||||
|
||||
The AIOps verifier is currently lightweight and rule-based. It checks:
|
||||
|
||||
- Whether the final report exists.
|
||||
- Whether the result stays focused on the alert payload when payload exists.
|
||||
- Whether evidence tools were used, especially `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
|
||||
## 5. RAG Architecture
|
||||
|
||||
```text
|
||||
lookup_knowledge
|
||||
-> L0 domain/entity hint
|
||||
-> matched domain
|
||||
-> matched keywords/entities
|
||||
-> metadata filter signal
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore path
|
||||
-> Milvus SDK fallback path
|
||||
-> evidence post-processing
|
||||
-> score / rawScore / scoreLabel
|
||||
-> source metadata
|
||||
-> title / breadcrumb / content evidence block
|
||||
-> tool_invocation record
|
||||
```
|
||||
|
||||
Important decisions:
|
||||
|
||||
- `lookup_knowledge` remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
|
||||
- L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
|
||||
- L1 retrieval now goes through `VectorSearchService`.
|
||||
- Spring AI `VectorStore` is the preferred retrieval path.
|
||||
- The original Milvus SDK path is retained as fallback and compatibility path.
|
||||
- `title`, `breadcrumb`, and `content` participate in embedding text so chunk context is less likely to be lost.
|
||||
- Retrieval output keeps compatibility fields: `score`, `rawScore`, and `scoreLabel`.
|
||||
|
||||
Vector retrieval modes:
|
||||
|
||||
```text
|
||||
retrieval.vector-store.mode=auto # Prefer Spring AI VectorStore, fallback to SDK
|
||||
retrieval.vector-store.mode=spring-ai # Use Spring AI VectorStore only
|
||||
retrieval.vector-store.mode=sdk # Use original Milvus SDK path
|
||||
```
|
||||
|
||||
## 6. Persistence And Trace
|
||||
|
||||
Current trace-related persistence:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
-> final_report
|
||||
-> self_evaluation
|
||||
-> verifier_evaluation
|
||||
-> aiops_rule_evaluation
|
||||
|
||||
agent_step
|
||||
-> role
|
||||
-> step input/output
|
||||
-> execution order
|
||||
|
||||
tool_invocation
|
||||
-> tool_name
|
||||
-> query
|
||||
-> retrieval_layer
|
||||
-> retrieval_details
|
||||
-> evidence blocks
|
||||
-> duration
|
||||
```
|
||||
|
||||
Trace API aggregates these records into a session-level view:
|
||||
|
||||
- Agent step sequence.
|
||||
- Tool calls and retrieval details.
|
||||
- Final diagnosis report.
|
||||
- Chat verifier status.
|
||||
- AIOps rule verifier status.
|
||||
|
||||
## 7. Quality Gates
|
||||
|
||||
Current quality gates:
|
||||
|
||||
- Chat verifier: LLM-based final answer verification for normal diagnosis.
|
||||
- AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
|
||||
- Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
|
||||
- RAG retrieval baseline: golden query set with offline baseline report.
|
||||
- Live RAG acceptance: post-reindex script for validating retrieval against the running stack.
|
||||
|
||||
These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.
|
||||
|
||||
## 8. Current Completion State
|
||||
|
||||
Completed for the current MVP stage:
|
||||
|
||||
- Explicit `lookup_knowledge` Agent tool.
|
||||
- L0 + L1 retrieval shape retained.
|
||||
- L0 downgraded to domain/entity hint.
|
||||
- Spring AI VectorStore retrieval path integrated.
|
||||
- Milvus SDK fallback retained.
|
||||
- RAG evidence post-processing added.
|
||||
- Breadcrumb/title/content embedding text improved.
|
||||
- RAG offline baseline and live acceptance script added.
|
||||
- AIOps payload query augmentation added.
|
||||
- AIOps lightweight verifier added.
|
||||
- Trace summary includes both chat verifier and AIOps verifier signals.
|
||||
|
||||
Deferred future enhancements:
|
||||
|
||||
- LLM QueryTransformer / MultiQuery.
|
||||
- BM25, RRF, and reranker.
|
||||
- Neighbor chunk or section-level context expansion.
|
||||
- VectorStore write path migration.
|
||||
- Full LLM-based AIOps verifier.
|
||||
- More complete golden set for recall, MRR, and nDCG metrics.
|
||||
|
||||
## 9. Key Code References
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||
|
||||
## 10. Supporting Materials
|
||||
|
||||
- `mvp/issues/rag-refactor-plan.md`
|
||||
- `eval/rag-retrieval/README.md`
|
||||
- `scripts/eval_rag_live_acceptance.py`
|
||||
- `interview/rag-refactor-story.md`
|
||||
- `interview/rag-vectorstore-interview-notes.md`
|
||||
- `interview/rag-retrieval-quality-report.md`
|
||||
- `interview/rag-breadcrumb-embedding-acceptance.md`
|
||||
- `interview/aiops-query-augmentation.md`
|
||||
- `interview/aiops-lightweight-verifier.md`
|
||||
@@ -59,3 +59,24 @@ When an AIOps request includes alert payload fields, the system SHALL include a
|
||||
- **WHEN** an AIOps request does not include alert payload fields
|
||||
- **THEN** the generated task prompt SHALL remain in auto-discovery mode
|
||||
- **AND** it SHALL not include a payload-derived recommended knowledge query
|
||||
|
||||
### Requirement: AIOps sessions SHALL persist lightweight rule evaluation
|
||||
When an AIOps final report is persisted, the system SHALL evaluate it with deterministic AIOps-specific quality rules and store the result in session self-evaluation.
|
||||
|
||||
#### Scenario: Payload-focused report is evaluated
|
||||
- **WHEN** an AIOps session has alert payload fields and a final report is persisted
|
||||
- **THEN** the system SHALL evaluate whether the report mentions the supplied alert and service
|
||||
- **AND** it SHALL store the result under `self_evaluation.aiops_rule_evaluation`
|
||||
|
||||
#### Scenario: Evidence coverage is evaluated
|
||||
- **WHEN** an AIOps final report is evaluated
|
||||
- **THEN** the system SHALL check whether evidence tool invocations such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted for the session
|
||||
|
||||
#### Scenario: Evaluation is traceable
|
||||
- **WHEN** the diagnosis trace API returns an AIOps session
|
||||
- **THEN** the session self-evaluation payload SHALL include `aiops_rule_evaluation` when it has been generated
|
||||
|
||||
#### Scenario: Evaluation uses stable verdicts
|
||||
- **WHEN** AIOps rule evaluation completes
|
||||
- **THEN** it SHALL produce a verdict from `PASS`, `WARN`, or `FAIL`
|
||||
- **AND** it SHALL include check details and a human-readable rationale
|
||||
|
||||
Reference in New Issue
Block a user