From 9dd6823fe7db875900f29f4fee2b737071038064 Mon Sep 17 00:00:00 2001 From: aruo <40362743+zyongxin@users.noreply.github.com> Date: Sun, 5 Jul 2026 11:29:16 +0800 Subject: [PATCH] docs: add rag retrieval quality report --- interview/rag-retrieval-quality-report.md | 211 ++++++++++++++++++++++ 1 file changed, 211 insertions(+) create mode 100644 interview/rag-retrieval-quality-report.md diff --git a/interview/rag-retrieval-quality-report.md b/interview/rag-retrieval-quality-report.md new file mode 100644 index 0000000..0b6eca4 --- /dev/null +++ b/interview/rag-retrieval-quality-report.md @@ -0,0 +1,211 @@ +# RAG Retrieval Quality Report + +## Purpose + +This report compares the live retrieval behavior of the original Milvus SDK path and the new Spring AI VectorStore path. + +The goal is to answer an interview-critical question: + +> After moving retrieval to Spring AI VectorStore, how do we know retrieval quality did not regress? + +This is not a full benchmark yet. It is a focused live smoke comparison using representative RAG queries against the current Milvus/Zilliz collection. + +## Setup + +Service endpoint: + +```text +GET http://127.0.0.1:9900/api/search/similar +``` + +Collection: + +```text +biz +``` + +Compared modes: + +```text +retrieval.vector-store.mode=sdk +retrieval.vector-store.mode=spring-ai +``` + +Each case used: + +```text +topK=3 +``` + +The application was restarted once per mode using command-line configuration so no repository config file had to be changed. + +## Cases + +| Case | Query | Purpose | +| --- | --- | --- | +| `err-timeout` | `ERR_TIMEOUT` | Exact error-code retrieval | +| `payment-service-timeout` | `payment-service timeout` | Service timeout troubleshooting | +| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | Database troubleshooting | +| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps alert-style retrieval | +| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | Abstract RAG design query | +| `database-filter` | `mysql timeout`, category=`database` | Metadata filter behavior | + +## Summary + +| Case | SDK Count | VectorStore Count | Top1 Same | TopK Overlap | Notes | +| --- | ---: | ---: | --- | ---: | --- | +| `err-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents | +| `payment-service-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents | +| `mysql-connection-pool` | 3 | 3 | Yes | 3/3 | Same ordering and same documents | +| `high-cpu-payment` | 3 | 3 | Yes | 3/3 | Same ordering and same documents | +| `rag-l0-l1` | 3 | 1 | Yes | 1/3 | VectorStore returned only the strongest candidate | +| `database-filter` | 0 | 0 | N/A | N/A | Both paths applied the filter consistently; no live docs matched `category=database` | + +## Representative Results + +### `ERR_TIMEOUT` + +SDK: + +```text +1. ERR_TIMEOUT score=0.5659486 label=l2_distance +2. ERR_GATEWAY_TIMEOUT score=0.6048740 label=l2_distance +3. Error handling score=0.7735061 label=l2_distance +``` + +VectorStore: + +```text +1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity +2. ERR_GATEWAY_TIMEOUT score=0.6048740 rawScore=0.3951259 label=similarity +3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity +``` + +Interpretation: + +- Document ordering is identical. +- Compatibility `score` is identical to SDK L2 distance. +- VectorStore `rawScore` exposes Spring AI similarity separately. + +### `MySQL connection pool` + +Both paths returned: + +```text +1. MySQL connection pool config +2. wait_timeout timeout +3. idle-timeout +``` + +Interpretation: + +- The migration preserves a precise infrastructure troubleshooting retrieval case. +- Metadata fields such as title, category, and source remain available. + +### `HighCPUUsage payment-service` + +Both paths returned: + +```text +1. 3. HighCPUUsage / payment-service troubleshooting steps +2. evidence mapping table row for HighCPUUsage/payment-service +3. 3.1 Symptom confirmation +``` + +Interpretation: + +- AIOps-style alert terms still retrieve the expected troubleshooting document. +- This is important because AIOps diagnosis depends on knowledge retrieval plus metrics/log evidence. + +### `rag-l0-l1` + +SDK returned three results, while VectorStore returned one: + +```text +Top1: Return error information +``` + +Interpretation: + +- Top1 did not regress. +- VectorStore appears stricter for low-similarity tail results because the Spring AI path uses `similarityThresholdAll()`. +- This is acceptable for current read-path migration, but it is worth tracking because abstract design questions may need query rewriting, better indexed docs, or adjusted threshold behavior. + +### `database-filter` + +Both paths returned zero results for: + +```text +query=mysql timeout +category=database +``` + +Interpretation: + +- The filter path is consistent. +- The live indexed MySQL docs are categorized as `infrastructure`, not `database`. +- This highlights a metadata taxonomy issue rather than a VectorStore migration regression. + +## Score Compatibility + +The comparison validates the score design: + +```text +SDK: + score = L2 distance + rawScore = L2 distance + scoreLabel = l2_distance + +VectorStore: + score = Milvus metadata.distance + rawScore = Spring AI similarity + scoreLabel = similarity +``` + +This keeps `lookup_knowledge` relevance normalization stable while still exposing the VectorStore score semantics for trace/debugging. + +## Findings + +### Finding 1: Main live cases are equivalent + +For exact error code, service timeout, MySQL troubleshooting, and AIOps alert-style retrieval, SDK and VectorStore returned identical top3 documents in identical order. + +This is strong evidence that the read-path migration did not regress the most important demo and troubleshooting cases. + +### Finding 2: Abstract RAG design queries need better retrieval support + +The `rag-l0-l1` query only returned one VectorStore candidate. The top result matched SDK top1, but the tail differed. + +This suggests the next quality work should focus on: + +- Query transformation for abstract design questions. +- Better indexing of interview/devflow RAG design docs. +- Context expansion around same-section chunks. +- Possibly tuning VectorStore threshold behavior. + +### Finding 3: Metadata taxonomy matters + +The category filter case returned zero results in both modes because the relevant MySQL docs are categorized as `infrastructure`, not `database`. + +This supports a previous RAG issue: category/domain metadata should be normalized before it is used as a hard filter. + +## Acceptance Decision + +The Spring AI VectorStore read path is accepted for current MVP/interview use: + +- Core troubleshooting cases match SDK behavior. +- Score compatibility is preserved. +- The VectorStore path exposes better score semantics without changing the `lookup_knowledge` API. +- SDK fallback remains available for runtime safety. + +The next retrieval-quality improvements should not block this migration. They should be handled as separate RAG quality work. + +## Next Work + +Recommended next steps: + +- Add a small automated live comparison script if repeated validation becomes common. +- Add topK overlap and top1 hit metrics to the offline evaluator. +- Normalize metadata categories such as `database` vs `infrastructure`. +- Add query rewriting for abstract RAG questions. +- Decide later whether to migrate indexing writes to Spring AI `VectorStore.add(...)`.