168 lines
7.0 KiB
Markdown
168 lines
7.0 KiB
Markdown
# RAG VectorStore Interview Notes
|
|
|
|
## 60-Second Explanation
|
|
|
|
I refactored the RAG retrieval path from a direct Milvus SDK-only implementation to a Spring AI `VectorStore` main path, while keeping the SDK path as a fallback.
|
|
|
|
The important part is not just the dependency change. I kept `VectorSearchService` as the boundary, so `lookup_knowledge` and the Agent workflow did not need to change. The system now supports three modes:
|
|
|
|
```text
|
|
auto -> try Spring AI VectorStore, fallback to SDK
|
|
spring-ai -> force VectorStore
|
|
sdk -> force SDK
|
|
```
|
|
|
|
During live verification, the first run found a real config mismatch: VectorStore was pointed at `business_knowledge`, but the real Zilliz collection was `biz`. The fallback worked, so the system still returned results through SDK. After aligning the collection name, the same query went through Spring AI VectorStore successfully.
|
|
|
|
I also fixed score compatibility. Spring AI Milvus exposes similarity as the document score, but the old `lookup_knowledge` logic expects L2 distance. So I preserve `rawScore` and `scoreLabel`, and use Milvus `metadata.distance` as the compatibility `score` when available.
|
|
|
|
## Architecture Answer
|
|
|
|
```text
|
|
Agent / API
|
|
-> lookup_knowledge or /api/search/similar
|
|
-> VectorSearchService
|
|
-> Spring AI VectorStore
|
|
-> Milvus SDK fallback
|
|
-> Milvus/Zilliz collection: biz
|
|
```
|
|
|
|
The key design choice is that `VectorSearchService` remains the retrieval facade. This avoids spreading framework-specific code into the Agent tool layer.
|
|
|
|
## Why Keep The SDK Path?
|
|
|
|
I kept SDK fallback for three reasons:
|
|
|
|
- Migration safety: the existing SDK path was already proven against the live collection.
|
|
- Runtime resilience: if VectorStore schema mapping or filtering fails, retrieval still works.
|
|
- Interview/demo stability: a retrieval abstraction change should not break the main Agent diagnosis demo.
|
|
|
|
This was validated in practice. When VectorStore pointed at the wrong collection, `auto` mode fell back to SDK and still returned results.
|
|
|
|
## Why Use Spring AI VectorStore At All?
|
|
|
|
Using Spring AI `VectorStore` moves the project closer to a standard RAG abstraction:
|
|
|
|
- Retrieval code no longer needs to own all Milvus-specific search details.
|
|
- Later features such as query transformers, document postprocessors, advisors, or retrievers can be introduced more naturally.
|
|
- The code becomes easier to compare with common Spring AI RAG patterns in an interview.
|
|
|
|
But I did not blindly replace everything. Writes/indexing still use SDK because changing read and write paths at the same time would make failures harder to isolate.
|
|
|
|
## Why Keep L0?
|
|
|
|
L0 is no longer treated as the final source of truth. It is a deterministic hint layer:
|
|
|
|
- It extracts domain/entity hints from indexed metadata.
|
|
- It helps constrain L1 retrieval by category when possible.
|
|
- It gives the Agent a stable clue even when semantic retrieval is weak.
|
|
|
|
The current design is:
|
|
|
|
```text
|
|
L0 = domain/entity hint
|
|
L1 = semantic retrieval through VectorStore/SDK
|
|
postprocess = evidence trace and relevance normalization
|
|
```
|
|
|
|
This is easier to defend than saying "we only use vector search." Real incident diagnosis often has exact identifiers, error codes, service names, and alert names. L0 is useful for those.
|
|
|
|
## Why Not Use Hidden Spring AI Advisors Directly?
|
|
|
|
For this project, `lookup_knowledge` remains an explicit tool.
|
|
|
|
Reason:
|
|
|
|
- The Agent trace needs to show when knowledge was retrieved.
|
|
- `tool_invocation` records input, output preview, relevance level, and evidence metadata.
|
|
- The interview story is about auditable Agent execution, not only answer quality.
|
|
|
|
Spring AI Advisors may be useful later, but hiding retrieval inside an advisor would make the evidence chain less visible unless we rebuild trace hooks around it.
|
|
|
|
## Score Design
|
|
|
|
The result object intentionally separates these fields:
|
|
|
|
```text
|
|
score -> compatibility score used by old relevance normalization
|
|
rawScore -> raw score from the retrieval implementation
|
|
scoreLabel -> semantic label for rawScore
|
|
```
|
|
|
|
For SDK:
|
|
|
|
```text
|
|
score = L2 distance
|
|
rawScore = L2 distance
|
|
scoreLabel = l2_distance
|
|
```
|
|
|
|
For VectorStore:
|
|
|
|
```text
|
|
score = metadata.distance if present
|
|
rawScore = Spring AI document score
|
|
scoreLabel = similarity
|
|
```
|
|
|
|
This prevents a subtle bug: if we treat Spring AI similarity as L2 distance, relevance becomes wrong. If we only expose distance, we lose the ability to compare Spring AI behavior. Keeping both makes the migration inspectable.
|
|
|
|
## How I Verified It
|
|
|
|
I verified at three levels:
|
|
|
|
- Unit tests: SDK mode, auto VectorStore mode, fallback mode, category filter, distance metadata mapping.
|
|
- Live API: `/api/search/similar?query=ERR_TIMEOUT&topK=3`.
|
|
- Logs: confirmed whether the path was VectorStore success or SDK fallback.
|
|
|
|
The live API returned:
|
|
|
|
```text
|
|
scoreLabel = similarity
|
|
rawScore = Spring AI similarity
|
|
score = Milvus distance metadata
|
|
```
|
|
|
|
That means the main path was Spring AI VectorStore and compatibility scoring remained stable.
|
|
|
|
## What I Would Do Next
|
|
|
|
I would not immediately migrate indexing writes. The next responsible steps are:
|
|
|
|
- Add a small live acceptance report for several golden queries.
|
|
- Compare `sdk` and `spring-ai` mode side by side for topK overlap.
|
|
- Decide whether `VectorIndexService` should move to `VectorStore.add(...)`.
|
|
- Add query transformation or hybrid retrieval only after we have baseline metrics.
|
|
|
|
This staged approach is intentional: first stabilize the read path, then evaluate retrieval quality, then migrate writes if the abstraction proves reliable.
|
|
|
|
## Interview Questions And Short Answers
|
|
|
|
### Why did you not remove the SDK?
|
|
|
|
Because this is a migration, not a rewrite. SDK fallback gives rollback safety and proved useful when VectorStore config was initially wrong.
|
|
|
|
### What changed for `lookup_knowledge`?
|
|
|
|
The public contract did not change. It still calls `VectorSearchService.searchSimilarDocuments(...)`. The implementation behind that facade changed.
|
|
|
|
### How do you know VectorStore is actually used?
|
|
|
|
The logs show `Starting Spring AI VectorStore search` followed by `Spring AI VectorStore search complete`. The API response also has `scoreLabel=similarity`, which only comes from the VectorStore path.
|
|
|
|
### What was the main bug found during live validation?
|
|
|
|
The configured collection name was wrong. Spring AI looked for `business_knowledge`, but the actual Milvus collection was `biz`.
|
|
|
|
### What did fallback prove?
|
|
|
|
It proved that `auto` mode is resilient: VectorStore failed, SDK search still returned valid results, and the API did not fail.
|
|
|
|
### Why is `metadata.distance` important?
|
|
|
|
Because `lookup_knowledge` uses L2 distance normalization. Spring AI returns similarity as the main document score, but the Milvus distance is available in metadata. Using it preserves old relevance behavior.
|
|
|
|
### Is this full Spring AI RAG now?
|
|
|
|
Not yet. It uses Spring AI VectorStore for the main read path, but keeps explicit tools, custom evidence trace, L0 hints, and SDK indexing. That is deliberate because the project values auditability and staged migration.
|