6.6 KiB
RAG Retrieval Quality Report
Purpose
This report compares the live retrieval behavior of the original Milvus SDK path and the new Spring AI VectorStore path.
The goal is to answer an interview-critical question:
After moving retrieval to Spring AI VectorStore, how do we know retrieval quality did not regress?
This is not a full benchmark yet. It is a focused live smoke comparison using representative RAG queries against the current Milvus/Zilliz collection.
Setup
Service endpoint:
GET http://127.0.0.1:9900/api/search/similar
Collection:
biz
Compared modes:
retrieval.vector-store.mode=sdk
retrieval.vector-store.mode=spring-ai
Each case used:
topK=3
The application was restarted once per mode using command-line configuration so no repository config file had to be changed.
Cases
| Case | Query | Purpose |
|---|---|---|
err-timeout |
ERR_TIMEOUT |
Exact error-code retrieval |
payment-service-timeout |
payment-service timeout |
Service timeout troubleshooting |
mysql-connection-pool |
MySQL connection pool is exhausted. How should I diagnose it? |
Database troubleshooting |
high-cpu-payment |
HighCPUUsage payment-service |
AIOps alert-style retrieval |
rag-l0-l1 |
Should L0 keyword matching decide the final retrieval result? |
Abstract RAG design query |
database-filter |
mysql timeout, category=database |
Metadata filter behavior |
Summary
| Case | SDK Count | VectorStore Count | Top1 Same | TopK Overlap | Notes |
|---|---|---|---|---|---|
err-timeout |
3 | 3 | Yes | 3/3 | Same ordering and same documents |
payment-service-timeout |
3 | 3 | Yes | 3/3 | Same ordering and same documents |
mysql-connection-pool |
3 | 3 | Yes | 3/3 | Same ordering and same documents |
high-cpu-payment |
3 | 3 | Yes | 3/3 | Same ordering and same documents |
rag-l0-l1 |
3 | 1 | Yes | 1/3 | VectorStore returned only the strongest candidate |
database-filter |
0 | 0 | N/A | N/A | Both paths applied the filter consistently; no live docs matched category=database |
Representative Results
ERR_TIMEOUT
SDK:
1. ERR_TIMEOUT score=0.5659486 label=l2_distance
2. ERR_GATEWAY_TIMEOUT score=0.6048740 label=l2_distance
3. Error handling score=0.7735061 label=l2_distance
VectorStore:
1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity
2. ERR_GATEWAY_TIMEOUT score=0.6048740 rawScore=0.3951259 label=similarity
3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity
Interpretation:
- Document ordering is identical.
- Compatibility
scoreis identical to SDK L2 distance. - VectorStore
rawScoreexposes Spring AI similarity separately.
MySQL connection pool
Both paths returned:
1. MySQL connection pool config
2. wait_timeout timeout
3. idle-timeout
Interpretation:
- The migration preserves a precise infrastructure troubleshooting retrieval case.
- Metadata fields such as title, category, and source remain available.
HighCPUUsage payment-service
Both paths returned:
1. 3. HighCPUUsage / payment-service troubleshooting steps
2. evidence mapping table row for HighCPUUsage/payment-service
3. 3.1 Symptom confirmation
Interpretation:
- AIOps-style alert terms still retrieve the expected troubleshooting document.
- This is important because AIOps diagnosis depends on knowledge retrieval plus metrics/log evidence.
rag-l0-l1
SDK returned three results, while VectorStore returned one:
Top1: Return error information
Interpretation:
- Top1 did not regress.
- VectorStore appears stricter for low-similarity tail results because the Spring AI path uses
similarityThresholdAll(). - This is acceptable for current read-path migration, but it is worth tracking because abstract design questions may need query rewriting, better indexed docs, or adjusted threshold behavior.
database-filter
Both paths returned zero results for:
query=mysql timeout
category=database
Interpretation:
- The filter path is consistent.
- The live indexed MySQL docs are categorized as
infrastructure, notdatabase. - This highlights a metadata taxonomy issue rather than a VectorStore migration regression.
Score Compatibility
The comparison validates the score design:
SDK:
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
VectorStore:
score = Milvus metadata.distance
rawScore = Spring AI similarity
scoreLabel = similarity
This keeps lookup_knowledge relevance normalization stable while still exposing the VectorStore score semantics for trace/debugging.
Findings
Finding 1: Main live cases are equivalent
For exact error code, service timeout, MySQL troubleshooting, and AIOps alert-style retrieval, SDK and VectorStore returned identical top3 documents in identical order.
This is strong evidence that the read-path migration did not regress the most important demo and troubleshooting cases.
Finding 2: Abstract RAG design queries need better retrieval support
The rag-l0-l1 query only returned one VectorStore candidate. The top result matched SDK top1, but the tail differed.
This suggests the next quality work should focus on:
- Query transformation for abstract design questions.
- Better indexing of interview/devflow RAG design docs.
- Context expansion around same-section chunks.
- Possibly tuning VectorStore threshold behavior.
Finding 3: Metadata taxonomy matters
The category filter case returned zero results in both modes because the relevant MySQL docs are categorized as infrastructure, not database.
This supports a previous RAG issue: category/domain metadata should be normalized before it is used as a hard filter.
Acceptance Decision
The Spring AI VectorStore read path is accepted for current MVP/interview use:
- Core troubleshooting cases match SDK behavior.
- Score compatibility is preserved.
- The VectorStore path exposes better score semantics without changing the
lookup_knowledgeAPI. - SDK fallback remains available for runtime safety.
The next retrieval-quality improvements should not block this migration. They should be handled as separate RAG quality work.
Next Work
Recommended next steps:
- Add a small automated live comparison script if repeated validation becomes common.
- Add topK overlap and top1 hit metrics to the offline evaluator.
- Normalize metadata categories such as
databasevsinfrastructure. - Add query rewriting for abstract RAG questions.
- Decide later whether to migrate indexing writes to Spring AI
VectorStore.add(...).