docs: add rag retrieval quality report

This commit is contained in:
aruo
2026-07-05 11:29:16 +08:00
parent 9376448804
commit 9dd6823fe7
+211
View File
@@ -0,0 +1,211 @@
# RAG Retrieval Quality Report
## Purpose
This report compares the live retrieval behavior of the original Milvus SDK path and the new Spring AI VectorStore path.
The goal is to answer an interview-critical question:
> After moving retrieval to Spring AI VectorStore, how do we know retrieval quality did not regress?
This is not a full benchmark yet. It is a focused live smoke comparison using representative RAG queries against the current Milvus/Zilliz collection.
## Setup
Service endpoint:
```text
GET http://127.0.0.1:9900/api/search/similar
```
Collection:
```text
biz
```
Compared modes:
```text
retrieval.vector-store.mode=sdk
retrieval.vector-store.mode=spring-ai
```
Each case used:
```text
topK=3
```
The application was restarted once per mode using command-line configuration so no repository config file had to be changed.
## Cases
| Case | Query | Purpose |
| --- | --- | --- |
| `err-timeout` | `ERR_TIMEOUT` | Exact error-code retrieval |
| `payment-service-timeout` | `payment-service timeout` | Service timeout troubleshooting |
| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | Database troubleshooting |
| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps alert-style retrieval |
| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | Abstract RAG design query |
| `database-filter` | `mysql timeout`, category=`database` | Metadata filter behavior |
## Summary
| Case | SDK Count | VectorStore Count | Top1 Same | TopK Overlap | Notes |
| --- | ---: | ---: | --- | ---: | --- |
| `err-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
| `payment-service-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
| `mysql-connection-pool` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
| `high-cpu-payment` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
| `rag-l0-l1` | 3 | 1 | Yes | 1/3 | VectorStore returned only the strongest candidate |
| `database-filter` | 0 | 0 | N/A | N/A | Both paths applied the filter consistently; no live docs matched `category=database` |
## Representative Results
### `ERR_TIMEOUT`
SDK:
```text
1. ERR_TIMEOUT score=0.5659486 label=l2_distance
2. ERR_GATEWAY_TIMEOUT score=0.6048740 label=l2_distance
3. Error handling score=0.7735061 label=l2_distance
```
VectorStore:
```text
1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity
2. ERR_GATEWAY_TIMEOUT score=0.6048740 rawScore=0.3951259 label=similarity
3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity
```
Interpretation:
- Document ordering is identical.
- Compatibility `score` is identical to SDK L2 distance.
- VectorStore `rawScore` exposes Spring AI similarity separately.
### `MySQL connection pool`
Both paths returned:
```text
1. MySQL connection pool config
2. wait_timeout timeout
3. idle-timeout
```
Interpretation:
- The migration preserves a precise infrastructure troubleshooting retrieval case.
- Metadata fields such as title, category, and source remain available.
### `HighCPUUsage payment-service`
Both paths returned:
```text
1. 3. HighCPUUsage / payment-service troubleshooting steps
2. evidence mapping table row for HighCPUUsage/payment-service
3. 3.1 Symptom confirmation
```
Interpretation:
- AIOps-style alert terms still retrieve the expected troubleshooting document.
- This is important because AIOps diagnosis depends on knowledge retrieval plus metrics/log evidence.
### `rag-l0-l1`
SDK returned three results, while VectorStore returned one:
```text
Top1: Return error information
```
Interpretation:
- Top1 did not regress.
- VectorStore appears stricter for low-similarity tail results because the Spring AI path uses `similarityThresholdAll()`.
- This is acceptable for current read-path migration, but it is worth tracking because abstract design questions may need query rewriting, better indexed docs, or adjusted threshold behavior.
### `database-filter`
Both paths returned zero results for:
```text
query=mysql timeout
category=database
```
Interpretation:
- The filter path is consistent.
- The live indexed MySQL docs are categorized as `infrastructure`, not `database`.
- This highlights a metadata taxonomy issue rather than a VectorStore migration regression.
## Score Compatibility
The comparison validates the score design:
```text
SDK:
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
VectorStore:
score = Milvus metadata.distance
rawScore = Spring AI similarity
scoreLabel = similarity
```
This keeps `lookup_knowledge` relevance normalization stable while still exposing the VectorStore score semantics for trace/debugging.
## Findings
### Finding 1: Main live cases are equivalent
For exact error code, service timeout, MySQL troubleshooting, and AIOps alert-style retrieval, SDK and VectorStore returned identical top3 documents in identical order.
This is strong evidence that the read-path migration did not regress the most important demo and troubleshooting cases.
### Finding 2: Abstract RAG design queries need better retrieval support
The `rag-l0-l1` query only returned one VectorStore candidate. The top result matched SDK top1, but the tail differed.
This suggests the next quality work should focus on:
- Query transformation for abstract design questions.
- Better indexing of interview/devflow RAG design docs.
- Context expansion around same-section chunks.
- Possibly tuning VectorStore threshold behavior.
### Finding 3: Metadata taxonomy matters
The category filter case returned zero results in both modes because the relevant MySQL docs are categorized as `infrastructure`, not `database`.
This supports a previous RAG issue: category/domain metadata should be normalized before it is used as a hard filter.
## Acceptance Decision
The Spring AI VectorStore read path is accepted for current MVP/interview use:
- Core troubleshooting cases match SDK behavior.
- Score compatibility is preserved.
- The VectorStore path exposes better score semantics without changing the `lookup_knowledge` API.
- SDK fallback remains available for runtime safety.
The next retrieval-quality improvements should not block this migration. They should be handled as separate RAG quality work.
## Next Work
Recommended next steps:
- Add a small automated live comparison script if repeated validation becomes common.
- Add topK overlap and top1 hit metrics to the offline evaluator.
- Normalize metadata categories such as `database` vs `infrastructure`.
- Add query rewriting for abstract RAG questions.
- Decide later whether to migrate indexing writes to Spring AI `VectorStore.add(...)`.