Files
SuperBizAgent-java/interview/rag-refactor-story.md
T
2026-07-05 12:09:32 +08:00

210 lines
9.4 KiB
Markdown

# RAG Refactor Story
## The Starting Point
The original RAG implementation was already usable for the MVP:
- Documents could be uploaded, chunked, embedded, and written to Milvus/Zilliz.
- The Agent could call `lookup_knowledge` as an explicit tool.
- AIOps diagnosis could retrieve troubleshooting knowledge during an alert workflow.
- Tool invocations were persisted, so the retrieval step was visible in the execution trace.
But the design had several engineering problems:
- Retrieval was too SDK-specific. The business code directly owned many Milvus search details.
- L0 and L1 responsibilities were blurry. L0 keyword matching could look like a final retrieval decision instead of a hint.
- Chunk-level retrieval could lose section context when one section was split into multiple chunks.
- Metadata such as `breadcrumb` existed, but it was not fully used in retrieval, filtering, or context reconstruction.
- Retrieval quality was mostly checked by manual API calls and logs, not by repeatable cases.
So the refactor goal was not "replace everything with a framework." The goal was to move generic RAG infrastructure toward Spring AI while keeping the project-specific Agent evidence chain.
## How I Broke The Problem Down
I treated this as a staged migration, because RAG touches the Agent tool layer, AIOps diagnosis, vector retrieval, evidence packing, and database traces.
The first step was to establish a baseline. I added retrieval evaluation cases under `eval/rag-retrieval/` so future changes could be compared against known queries instead of judged only by intuition.
Then I clarified the retrieval roles:
```text
L0 = domain/entity hint
L1 = semantic retrieval
postprocess = evidence shaping and trace-friendly output
```
That means L0 is still valuable, but it should not bypass semantic retrieval as the default path. It is better used to extract service names, alert names, error codes, domains, and metadata hints.
After that, I added evidence postprocessing. The Agent should not just receive raw chunks; it should receive structured evidence with source, title, breadcrumb, score, hit reason, and content. This makes the result easier to inspect and easier to explain in an interview.
Finally, I integrated Spring AI `VectorStore` as the main read path while preserving the original Milvus SDK implementation as fallback.
## Current Architecture
The current retrieval path is:
```text
Agent / API
-> lookup_knowledge or /api/search/similar
-> L0 domain/entity hint
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
-> evidence postprocess
-> tool_invocation trace
```
`VectorSearchService` is still the public retrieval facade. This is deliberate: the Agent tool layer does not need to know whether the underlying retrieval engine is SDK-based or Spring AI-based.
The supported retrieval modes are:
```text
auto -> try Spring AI VectorStore, fallback to SDK
spring-ai -> force Spring AI VectorStore
sdk -> force Milvus SDK
```
This keeps the migration reversible and testable.
## Key Tradeoffs
### Keep The Explicit Tool
I did not hide retrieval inside a Spring AI Advisor.
For this project, `lookup_knowledge` is part of the Agent execution story. It records what query was used, which evidence was retrieved, how relevant it looked, and how it supported diagnosis. If retrieval is hidden inside an advisor, the answer may still work, but the audit trail becomes harder to show.
### Keep SDK Fallback
The SDK path is not dead code. It is a safety net during migration.
This proved useful during live validation. The first VectorStore run pointed at the wrong collection name, but `auto` mode fell back to SDK and still returned results. After the collection was corrected to `biz`, the Spring AI path worked as the main path.
### Keep L0, But Reduce Its Authority
L0 is worth keeping because production incidents often contain exact identifiers:
- error code
- alert name
- service name
- metric name
- domain tag
But L0 should not be the final judge of retrieval quality. Its role is now closer to domain hint, entity extraction, metadata filtering, and explainability signal.
### Split Score Semantics
The old SDK path used L2 distance. Spring AI exposes similarity. Treating those as the same number would quietly break relevance normalization.
So the result separates:
```text
score -> compatibility score used by existing logic
rawScore -> raw score from the retrieval implementation
scoreLabel -> semantic label for rawScore
```
For SDK:
```text
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
```
For VectorStore:
```text
score = Milvus metadata.distance when available
rawScore = Spring AI similarity
scoreLabel = similarity
```
This makes the migration inspectable instead of hiding score changes behind one overloaded field.
### Do Not Migrate Writes Yet
Writes and indexing still use the SDK path.
That is intentional. Migrating reads and writes at the same time would make debugging harder. The read path can be validated first; write-path migration can happen later if Spring AI `VectorStore.add(...)` fits the existing metadata and chunk model.
## Validation Story
I validated the refactor at multiple levels.
Unit tests cover:
- SDK mode.
- Spring AI mode.
- `auto` fallback.
- category filter behavior.
- distance metadata mapping.
Live API verification used:
```text
GET /api/search/similar?query=ERR_TIMEOUT&topK=3
```
Logs confirmed when the Spring AI VectorStore path was used and when fallback happened.
Then I compared SDK and VectorStore retrieval quality on representative queries:
| Query Type | Result |
| --- | --- |
| exact error code | same top3 |
| payment-service timeout | same top3 |
| MySQL connection pool | same top3 |
| AIOps alert-style query | same top3 |
| abstract RAG design query | same top1, VectorStore returned fewer tail results |
| category filter | both returned zero because metadata taxonomy did not match |
The acceptance decision was that Spring AI VectorStore is good enough for the current MVP read path, with SDK fallback preserved.
## Known Gaps
The refactor improved the architecture, but it did not solve every retrieval-quality problem.
Known gaps:
- Metadata taxonomy still needs cleanup, for example `database` vs `infrastructure`.
- Abstract design questions may need query rewriting or better indexed interview/devflow documents.
- Chunk context reconstruction is still limited when one logical section spans multiple chunks.
- `breadcrumb` should be used more strongly in embedding text, retrieval metadata, and evidence packing.
- Rerank, RRF, BM25, and hybrid retrieval are not implemented yet.
- Indexing writes still use SDK.
These are good follow-up issues because they are retrieval-quality improvements, not blockers for the VectorStore migration.
## How I Present This In An Interview
My short version would be:
> This RAG system started as a self-built MVP around Milvus SDK retrieval. It worked, but too much infrastructure logic lived in business code, and L0/L1 responsibilities were unclear. I refactored it in stages: first I added baseline retrieval cases, then made L0 a domain/entity hint instead of a final decision layer, then added evidence postprocessing, and finally moved the main read path to Spring AI VectorStore with SDK fallback. I kept `lookup_knowledge` as an explicit Agent tool because the project values traceability: the interviewer can see when retrieval happened, what evidence was found, and how it supported the diagnosis. The result is closer to standard Spring AI RAG while still preserving business-specific observability.
If asked why this is not a full framework migration:
> I intentionally did not migrate everything at once. Reads moved first because they are easier to compare using golden queries. Writes/indexing stayed on SDK to avoid mixing schema and retrieval behavior changes in one step. Advisors were not used as the main interface because hidden retrieval would weaken the Agent trace.
If asked what I would improve next:
> I would add query transformation for AIOps payloads, improve metadata taxonomy, use breadcrumb and section metadata for context expansion, and then evaluate whether hybrid retrieval or rerank is necessary based on measured recall and topK overlap.
## Interview Follow-Up Questions
### Why introduce Spring AI VectorStore if the SDK path already worked?
Because SDK-only retrieval made the project own too much low-level RAG infrastructure. `VectorStore` gives a standard abstraction for retrieval and makes future Spring AI features easier to adopt, while the facade keeps the Agent layer stable.
### Why keep custom code at all?
The custom code is where the Agent engineering value lives: AIOps payload mapping, L0 hints, evidence packing, score compatibility, and tool invocation tracing. Those are domain-specific and should remain visible.
### How do you know quality did not regress?
I compared SDK and VectorStore modes on representative live queries. Core troubleshooting and AIOps cases returned the same top3 documents in the same order. The differences were isolated to abstract design queries and metadata taxonomy, which are documented follow-up work.
### What is the most important design decision?
Keeping a stable boundary: `lookup_knowledge` calls `VectorSearchService`, and `VectorSearchService` decides whether to use Spring AI or SDK. That boundary made the migration small enough to validate and explain.