Files
SuperBizAgent-java/interview/rag-retrieval-quality-report.md
T

6.6 KiB

RAG Retrieval Quality Report

Purpose

This report compares the live retrieval behavior of the original Milvus SDK path and the new Spring AI VectorStore path.

The goal is to answer an interview-critical question:

After moving retrieval to Spring AI VectorStore, how do we know retrieval quality did not regress?

This is not a full benchmark yet. It is a focused live smoke comparison using representative RAG queries against the current Milvus/Zilliz collection.

Setup

Service endpoint:

GET http://127.0.0.1:9900/api/search/similar

Collection:

biz

Compared modes:

retrieval.vector-store.mode=sdk
retrieval.vector-store.mode=spring-ai

Each case used:

topK=3

The application was restarted once per mode using command-line configuration so no repository config file had to be changed.

Cases

Case Query Purpose
err-timeout ERR_TIMEOUT Exact error-code retrieval
payment-service-timeout payment-service timeout Service timeout troubleshooting
mysql-connection-pool MySQL connection pool is exhausted. How should I diagnose it? Database troubleshooting
high-cpu-payment HighCPUUsage payment-service AIOps alert-style retrieval
rag-l0-l1 Should L0 keyword matching decide the final retrieval result? Abstract RAG design query
database-filter mysql timeout, category=database Metadata filter behavior

Summary

Case SDK Count VectorStore Count Top1 Same TopK Overlap Notes
err-timeout 3 3 Yes 3/3 Same ordering and same documents
payment-service-timeout 3 3 Yes 3/3 Same ordering and same documents
mysql-connection-pool 3 3 Yes 3/3 Same ordering and same documents
high-cpu-payment 3 3 Yes 3/3 Same ordering and same documents
rag-l0-l1 3 1 Yes 1/3 VectorStore returned only the strongest candidate
database-filter 0 0 N/A N/A Both paths applied the filter consistently; no live docs matched category=database

Representative Results

ERR_TIMEOUT

SDK:

1. ERR_TIMEOUT              score=0.5659486 label=l2_distance
2. ERR_GATEWAY_TIMEOUT      score=0.6048740 label=l2_distance
3. Error handling           score=0.7735061 label=l2_distance

VectorStore:

1. ERR_TIMEOUT              score=0.5659486 rawScore=0.4340513 label=similarity
2. ERR_GATEWAY_TIMEOUT      score=0.6048740 rawScore=0.3951259 label=similarity
3. Error handling           score=0.7735061 rawScore=0.2264938 label=similarity

Interpretation:

  • Document ordering is identical.
  • Compatibility score is identical to SDK L2 distance.
  • VectorStore rawScore exposes Spring AI similarity separately.

MySQL connection pool

Both paths returned:

1. MySQL connection pool config
2. wait_timeout timeout
3. idle-timeout

Interpretation:

  • The migration preserves a precise infrastructure troubleshooting retrieval case.
  • Metadata fields such as title, category, and source remain available.

HighCPUUsage payment-service

Both paths returned:

1. 3. HighCPUUsage / payment-service troubleshooting steps
2. evidence mapping table row for HighCPUUsage/payment-service
3. 3.1 Symptom confirmation

Interpretation:

  • AIOps-style alert terms still retrieve the expected troubleshooting document.
  • This is important because AIOps diagnosis depends on knowledge retrieval plus metrics/log evidence.

rag-l0-l1

SDK returned three results, while VectorStore returned one:

Top1: Return error information

Interpretation:

  • Top1 did not regress.
  • VectorStore appears stricter for low-similarity tail results because the Spring AI path uses similarityThresholdAll().
  • This is acceptable for current read-path migration, but it is worth tracking because abstract design questions may need query rewriting, better indexed docs, or adjusted threshold behavior.

database-filter

Both paths returned zero results for:

query=mysql timeout
category=database

Interpretation:

  • The filter path is consistent.
  • The live indexed MySQL docs are categorized as infrastructure, not database.
  • This highlights a metadata taxonomy issue rather than a VectorStore migration regression.

Score Compatibility

The comparison validates the score design:

SDK:
  score = L2 distance
  rawScore = L2 distance
  scoreLabel = l2_distance

VectorStore:
  score = Milvus metadata.distance
  rawScore = Spring AI similarity
  scoreLabel = similarity

This keeps lookup_knowledge relevance normalization stable while still exposing the VectorStore score semantics for trace/debugging.

Findings

Finding 1: Main live cases are equivalent

For exact error code, service timeout, MySQL troubleshooting, and AIOps alert-style retrieval, SDK and VectorStore returned identical top3 documents in identical order.

This is strong evidence that the read-path migration did not regress the most important demo and troubleshooting cases.

Finding 2: Abstract RAG design queries need better retrieval support

The rag-l0-l1 query only returned one VectorStore candidate. The top result matched SDK top1, but the tail differed.

This suggests the next quality work should focus on:

  • Query transformation for abstract design questions.
  • Better indexing of interview/devflow RAG design docs.
  • Context expansion around same-section chunks.
  • Possibly tuning VectorStore threshold behavior.

Finding 3: Metadata taxonomy matters

The category filter case returned zero results in both modes because the relevant MySQL docs are categorized as infrastructure, not database.

This supports a previous RAG issue: category/domain metadata should be normalized before it is used as a hard filter.

Acceptance Decision

The Spring AI VectorStore read path is accepted for current MVP/interview use:

  • Core troubleshooting cases match SDK behavior.
  • Score compatibility is preserved.
  • The VectorStore path exposes better score semantics without changing the lookup_knowledge API.
  • SDK fallback remains available for runtime safety.

The next retrieval-quality improvements should not block this migration. They should be handled as separate RAG quality work.

Next Work

Recommended next steps:

  • Add a small automated live comparison script if repeated validation becomes common.
  • Add topK overlap and top1 hit metrics to the offline evaluator.
  • Normalize metadata categories such as database vs infrastructure.
  • Add query rewriting for abstract RAG questions.
  • Decide later whether to migrate indexing writes to Spring AI VectorStore.add(...).