# RAG Retrieval Quality Report ## Purpose This report compares the live retrieval behavior of the original Milvus SDK path and the new Spring AI VectorStore path. The goal is to answer an interview-critical question: > After moving retrieval to Spring AI VectorStore, how do we know retrieval quality did not regress? This is not a full benchmark yet. It is a focused live smoke comparison using representative RAG queries against the current Milvus/Zilliz collection. ## Setup Service endpoint: ```text GET http://127.0.0.1:9900/api/search/similar ``` Collection: ```text biz ``` Compared modes: ```text retrieval.vector-store.mode=sdk retrieval.vector-store.mode=spring-ai ``` Each case used: ```text topK=3 ``` The application was restarted once per mode using command-line configuration so no repository config file had to be changed. ## Cases | Case | Query | Purpose | | --- | --- | --- | | `err-timeout` | `ERR_TIMEOUT` | Exact error-code retrieval | | `payment-service-timeout` | `payment-service timeout` | Service timeout troubleshooting | | `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | Database troubleshooting | | `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps alert-style retrieval | | `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | Abstract RAG design query | | `database-filter` | `mysql timeout`, category=`database` | Metadata filter behavior | ## Summary | Case | SDK Count | VectorStore Count | Top1 Same | TopK Overlap | Notes | | --- | ---: | ---: | --- | ---: | --- | | `err-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents | | `payment-service-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents | | `mysql-connection-pool` | 3 | 3 | Yes | 3/3 | Same ordering and same documents | | `high-cpu-payment` | 3 | 3 | Yes | 3/3 | Same ordering and same documents | | `rag-l0-l1` | 3 | 1 | Yes | 1/3 | VectorStore returned only the strongest candidate | | `database-filter` | 0 | 0 | N/A | N/A | Both paths applied the filter consistently; no live docs matched `category=database` | ## Representative Results ### `ERR_TIMEOUT` SDK: ```text 1. ERR_TIMEOUT score=0.5659486 label=l2_distance 2. ERR_GATEWAY_TIMEOUT score=0.6048740 label=l2_distance 3. Error handling score=0.7735061 label=l2_distance ``` VectorStore: ```text 1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity 2. ERR_GATEWAY_TIMEOUT score=0.6048740 rawScore=0.3951259 label=similarity 3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity ``` Interpretation: - Document ordering is identical. - Compatibility `score` is identical to SDK L2 distance. - VectorStore `rawScore` exposes Spring AI similarity separately. ### `MySQL connection pool` Both paths returned: ```text 1. MySQL connection pool config 2. wait_timeout timeout 3. idle-timeout ``` Interpretation: - The migration preserves a precise infrastructure troubleshooting retrieval case. - Metadata fields such as title, category, and source remain available. ### `HighCPUUsage payment-service` Both paths returned: ```text 1. 3. HighCPUUsage / payment-service troubleshooting steps 2. evidence mapping table row for HighCPUUsage/payment-service 3. 3.1 Symptom confirmation ``` Interpretation: - AIOps-style alert terms still retrieve the expected troubleshooting document. - This is important because AIOps diagnosis depends on knowledge retrieval plus metrics/log evidence. ### `rag-l0-l1` SDK returned three results, while VectorStore returned one: ```text Top1: Return error information ``` Interpretation: - Top1 did not regress. - VectorStore appears stricter for low-similarity tail results because the Spring AI path uses `similarityThresholdAll()`. - This is acceptable for current read-path migration, but it is worth tracking because abstract design questions may need query rewriting, better indexed docs, or adjusted threshold behavior. ### `database-filter` Both paths returned zero results for: ```text query=mysql timeout category=database ``` Interpretation: - The filter path is consistent. - The live indexed MySQL docs are categorized as `infrastructure`, not `database`. - This highlights a metadata taxonomy issue rather than a VectorStore migration regression. ## Score Compatibility The comparison validates the score design: ```text SDK: score = L2 distance rawScore = L2 distance scoreLabel = l2_distance VectorStore: score = Milvus metadata.distance rawScore = Spring AI similarity scoreLabel = similarity ``` This keeps `lookup_knowledge` relevance normalization stable while still exposing the VectorStore score semantics for trace/debugging. ## Findings ### Finding 1: Main live cases are equivalent For exact error code, service timeout, MySQL troubleshooting, and AIOps alert-style retrieval, SDK and VectorStore returned identical top3 documents in identical order. This is strong evidence that the read-path migration did not regress the most important demo and troubleshooting cases. ### Finding 2: Abstract RAG design queries need better retrieval support The `rag-l0-l1` query only returned one VectorStore candidate. The top result matched SDK top1, but the tail differed. This suggests the next quality work should focus on: - Query transformation for abstract design questions. - Better indexing of interview/devflow RAG design docs. - Context expansion around same-section chunks. - Possibly tuning VectorStore threshold behavior. ### Finding 3: Metadata taxonomy matters The category filter case returned zero results in both modes because the relevant MySQL docs are categorized as `infrastructure`, not `database`. This supports a previous RAG issue: category/domain metadata should be normalized before it is used as a hard filter. ## Acceptance Decision The Spring AI VectorStore read path is accepted for current MVP/interview use: - Core troubleshooting cases match SDK behavior. - Score compatibility is preserved. - The VectorStore path exposes better score semantics without changing the `lookup_knowledge` API. - SDK fallback remains available for runtime safety. The next retrieval-quality improvements should not block this migration. They should be handled as separate RAG quality work. ## Next Work Recommended next steps: - Add a small automated live comparison script if repeated validation becomes common. - Add topK overlap and top1 hit metrics to the offline evaluator. - Normalize metadata categories such as `database` vs `infrastructure`. - Add query rewriting for abstract RAG questions. - Decide later whether to migrate indexing writes to Spring AI `VectorStore.add(...)`.