Compare commits
3
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b22f2d22c8 | ||
|
|
63b62b28a2 | ||
|
|
d5902a0499 |
@@ -1,5 +1,7 @@
|
|||||||
# 数据库设计文档
|
# 数据库设计文档
|
||||||
|
|
||||||
|
> 当前架构快照:[mvp/architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md)
|
||||||
|
|
||||||
## 📚 文档导航
|
## 📚 文档导航
|
||||||
|
|
||||||
### 核心表设计
|
### 核心表设计
|
||||||
|
|||||||
@@ -0,0 +1,220 @@
|
|||||||
|
# Current MVP Architecture Snapshot
|
||||||
|
|
||||||
|
**Updated**: 2026-07-05
|
||||||
|
|
||||||
|
This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.
|
||||||
|
|
||||||
|
## 1. Positioning
|
||||||
|
|
||||||
|
The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.
|
||||||
|
|
||||||
|
Core goals:
|
||||||
|
|
||||||
|
- Support normal chat-based diagnosis.
|
||||||
|
- Support AIOps alert-triggered diagnosis.
|
||||||
|
- Keep tool calls explicit and traceable.
|
||||||
|
- Keep RAG retrieval observable through `lookup_knowledge`.
|
||||||
|
- Persist enough execution evidence for replay, evaluation, and interview explanation.
|
||||||
|
|
||||||
|
## 2. Runtime Architecture
|
||||||
|
|
||||||
|
```text
|
||||||
|
HTTP API
|
||||||
|
-> ChatService / AiOpsService
|
||||||
|
-> Agent orchestration
|
||||||
|
-> Supervisor / Planner / Executor / Verifier
|
||||||
|
-> Tools
|
||||||
|
-> lookup_knowledge
|
||||||
|
-> query_logs
|
||||||
|
-> query_metrics
|
||||||
|
-> other diagnosis tools
|
||||||
|
-> Persistence
|
||||||
|
-> diagnosis_session
|
||||||
|
-> agent_step
|
||||||
|
-> tool_invocation
|
||||||
|
-> Trace API
|
||||||
|
-> DiagnosisTraceService
|
||||||
|
```
|
||||||
|
|
||||||
|
Current entry points:
|
||||||
|
|
||||||
|
- `ChatService`: user-driven troubleshooting and follow-up diagnosis.
|
||||||
|
- `AiOpsService`: alert-driven diagnosis, including payload mode and auto-discovery mode.
|
||||||
|
- `DiagnosisTraceService`: trace view of session, steps, tool calls, and self-evaluation.
|
||||||
|
|
||||||
|
## 3. Chat Diagnosis Flow
|
||||||
|
|
||||||
|
```text
|
||||||
|
User question
|
||||||
|
-> ChatService
|
||||||
|
-> simple response or diagnosis flow
|
||||||
|
-> Planner creates investigation direction
|
||||||
|
-> Executor calls tools for evidence
|
||||||
|
-> lookup_knowledge
|
||||||
|
-> query_logs
|
||||||
|
-> query_metrics
|
||||||
|
-> Verifier checks final diagnosis quality
|
||||||
|
-> self_evaluation.verifier_evaluation
|
||||||
|
-> diagnosis trace
|
||||||
|
```
|
||||||
|
|
||||||
|
The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||||
|
|
||||||
|
## 4. AIOps Diagnosis Flow
|
||||||
|
|
||||||
|
```text
|
||||||
|
AIOps request
|
||||||
|
-> AiOpsService
|
||||||
|
-> payload mode or auto-discovery mode
|
||||||
|
-> build alert-focused diagnosis prompt
|
||||||
|
-> append recommended lookup_knowledge query when payload exists
|
||||||
|
-> Agent diagnosis flow
|
||||||
|
-> Supervisor / Planner / Executor
|
||||||
|
-> evidence tools
|
||||||
|
-> final report
|
||||||
|
-> AiOpsRuleEvaluationService
|
||||||
|
-> self_evaluation.aiops_rule_evaluation
|
||||||
|
-> diagnosis trace
|
||||||
|
```
|
||||||
|
|
||||||
|
AIOps keeps two modes:
|
||||||
|
|
||||||
|
- Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
|
||||||
|
- Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.
|
||||||
|
|
||||||
|
The AIOps verifier is currently lightweight and rule-based. It checks:
|
||||||
|
|
||||||
|
- Whether the final report exists.
|
||||||
|
- Whether the result stays focused on the alert payload when payload exists.
|
||||||
|
- Whether evidence tools were used, especially `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||||
|
|
||||||
|
## 5. RAG Architecture
|
||||||
|
|
||||||
|
```text
|
||||||
|
lookup_knowledge
|
||||||
|
-> L0 domain/entity hint
|
||||||
|
-> matched domain
|
||||||
|
-> matched keywords/entities
|
||||||
|
-> metadata filter signal
|
||||||
|
-> VectorSearchService
|
||||||
|
-> Spring AI VectorStore path
|
||||||
|
-> Milvus SDK fallback path
|
||||||
|
-> evidence post-processing
|
||||||
|
-> score / rawScore / scoreLabel
|
||||||
|
-> source metadata
|
||||||
|
-> title / breadcrumb / content evidence block
|
||||||
|
-> tool_invocation record
|
||||||
|
```
|
||||||
|
|
||||||
|
Important decisions:
|
||||||
|
|
||||||
|
- `lookup_knowledge` remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
|
||||||
|
- L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
|
||||||
|
- L1 retrieval now goes through `VectorSearchService`.
|
||||||
|
- Spring AI `VectorStore` is the preferred retrieval path.
|
||||||
|
- The original Milvus SDK path is retained as fallback and compatibility path.
|
||||||
|
- `title`, `breadcrumb`, and `content` participate in embedding text so chunk context is less likely to be lost.
|
||||||
|
- Retrieval output keeps compatibility fields: `score`, `rawScore`, and `scoreLabel`.
|
||||||
|
|
||||||
|
Vector retrieval modes:
|
||||||
|
|
||||||
|
```text
|
||||||
|
retrieval.vector-store.mode=auto # Prefer Spring AI VectorStore, fallback to SDK
|
||||||
|
retrieval.vector-store.mode=spring-ai # Use Spring AI VectorStore only
|
||||||
|
retrieval.vector-store.mode=sdk # Use original Milvus SDK path
|
||||||
|
```
|
||||||
|
|
||||||
|
## 6. Persistence And Trace
|
||||||
|
|
||||||
|
Current trace-related persistence:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session
|
||||||
|
-> final_report
|
||||||
|
-> self_evaluation
|
||||||
|
-> verifier_evaluation
|
||||||
|
-> aiops_rule_evaluation
|
||||||
|
|
||||||
|
agent_step
|
||||||
|
-> role
|
||||||
|
-> step input/output
|
||||||
|
-> execution order
|
||||||
|
|
||||||
|
tool_invocation
|
||||||
|
-> tool_name
|
||||||
|
-> query
|
||||||
|
-> retrieval_layer
|
||||||
|
-> retrieval_details
|
||||||
|
-> evidence blocks
|
||||||
|
-> duration
|
||||||
|
```
|
||||||
|
|
||||||
|
Trace API aggregates these records into a session-level view:
|
||||||
|
|
||||||
|
- Agent step sequence.
|
||||||
|
- Tool calls and retrieval details.
|
||||||
|
- Final diagnosis report.
|
||||||
|
- Chat verifier status.
|
||||||
|
- AIOps rule verifier status.
|
||||||
|
|
||||||
|
## 7. Quality Gates
|
||||||
|
|
||||||
|
Current quality gates:
|
||||||
|
|
||||||
|
- Chat verifier: LLM-based final answer verification for normal diagnosis.
|
||||||
|
- AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
|
||||||
|
- Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
|
||||||
|
- RAG retrieval baseline: golden query set with offline baseline report.
|
||||||
|
- Live RAG acceptance: post-reindex script for validating retrieval against the running stack.
|
||||||
|
|
||||||
|
These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.
|
||||||
|
|
||||||
|
## 8. Current Completion State
|
||||||
|
|
||||||
|
Completed for the current MVP stage:
|
||||||
|
|
||||||
|
- Explicit `lookup_knowledge` Agent tool.
|
||||||
|
- L0 + L1 retrieval shape retained.
|
||||||
|
- L0 downgraded to domain/entity hint.
|
||||||
|
- Spring AI VectorStore retrieval path integrated.
|
||||||
|
- Milvus SDK fallback retained.
|
||||||
|
- RAG evidence post-processing added.
|
||||||
|
- Breadcrumb/title/content embedding text improved.
|
||||||
|
- RAG offline baseline and live acceptance script added.
|
||||||
|
- AIOps payload query augmentation added.
|
||||||
|
- AIOps lightweight verifier added.
|
||||||
|
- Trace summary includes both chat verifier and AIOps verifier signals.
|
||||||
|
|
||||||
|
Deferred future enhancements:
|
||||||
|
|
||||||
|
- LLM QueryTransformer / MultiQuery.
|
||||||
|
- BM25, RRF, and reranker.
|
||||||
|
- Neighbor chunk or section-level context expansion.
|
||||||
|
- VectorStore write path migration.
|
||||||
|
- Full LLM-based AIOps verifier.
|
||||||
|
- More complete golden set for recall, MRR, and nDCG metrics.
|
||||||
|
|
||||||
|
## 9. Key Code References
|
||||||
|
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||||
|
|
||||||
|
## 10. Supporting Materials
|
||||||
|
|
||||||
|
- `mvp/issues/rag-refactor-plan.md`
|
||||||
|
- `eval/rag-retrieval/README.md`
|
||||||
|
- `scripts/eval_rag_live_acceptance.py`
|
||||||
|
- `interview/rag-refactor-story.md`
|
||||||
|
- `interview/rag-vectorstore-interview-notes.md`
|
||||||
|
- `interview/rag-retrieval-quality-report.md`
|
||||||
|
- `interview/rag-breadcrumb-embedding-acceptance.md`
|
||||||
|
- `interview/aiops-query-augmentation.md`
|
||||||
|
- `interview/aiops-lightweight-verifier.md`
|
||||||
@@ -59,3 +59,24 @@ When an AIOps request includes alert payload fields, the system SHALL include a
|
|||||||
- **WHEN** an AIOps request does not include alert payload fields
|
- **WHEN** an AIOps request does not include alert payload fields
|
||||||
- **THEN** the generated task prompt SHALL remain in auto-discovery mode
|
- **THEN** the generated task prompt SHALL remain in auto-discovery mode
|
||||||
- **AND** it SHALL not include a payload-derived recommended knowledge query
|
- **AND** it SHALL not include a payload-derived recommended knowledge query
|
||||||
|
|
||||||
|
### Requirement: AIOps sessions SHALL persist lightweight rule evaluation
|
||||||
|
When an AIOps final report is persisted, the system SHALL evaluate it with deterministic AIOps-specific quality rules and store the result in session self-evaluation.
|
||||||
|
|
||||||
|
#### Scenario: Payload-focused report is evaluated
|
||||||
|
- **WHEN** an AIOps session has alert payload fields and a final report is persisted
|
||||||
|
- **THEN** the system SHALL evaluate whether the report mentions the supplied alert and service
|
||||||
|
- **AND** it SHALL store the result under `self_evaluation.aiops_rule_evaluation`
|
||||||
|
|
||||||
|
#### Scenario: Evidence coverage is evaluated
|
||||||
|
- **WHEN** an AIOps final report is evaluated
|
||||||
|
- **THEN** the system SHALL check whether evidence tool invocations such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted for the session
|
||||||
|
|
||||||
|
#### Scenario: Evaluation is traceable
|
||||||
|
- **WHEN** the diagnosis trace API returns an AIOps session
|
||||||
|
- **THEN** the session self-evaluation payload SHALL include `aiops_rule_evaluation` when it has been generated
|
||||||
|
|
||||||
|
#### Scenario: Evaluation uses stable verdicts
|
||||||
|
- **WHEN** AIOps rule evaluation completes
|
||||||
|
- **THEN** it SHALL produce a verdict from `PASS`, `WARN`, or `FAIL`
|
||||||
|
- **AND** it SHALL include check details and a human-readable rationale
|
||||||
|
|||||||
Reference in New Issue
Block a user