docs: reorganize MVP interview documentation
This commit is contained in:
+72
-154
@@ -1,47 +1,50 @@
|
||||
# RAG Refactor Story
|
||||
# RAG 重构故事
|
||||
|
||||
## The Starting Point
|
||||
## 1. 起点
|
||||
|
||||
The original RAG implementation was already usable for the MVP:
|
||||
原始 RAG 实现已经能支撑 MVP:
|
||||
|
||||
- Documents could be uploaded, chunked, embedded, and written to Milvus/Zilliz.
|
||||
- The Agent could call `lookup_knowledge` as an explicit tool.
|
||||
- AIOps diagnosis could retrieve troubleshooting knowledge during an alert workflow.
|
||||
- Tool invocations were persisted, so the retrieval step was visible in the execution trace.
|
||||
- 文档可以上传、切片、向量化,并写入 Milvus/Zilliz。
|
||||
- Agent 可以显式调用 `lookup_knowledge`。
|
||||
- AIOps 诊断能在告警流程里检索排障知识。
|
||||
- 工具调用会落到 `tool_invocation`,检索步骤可见。
|
||||
|
||||
But the design had several engineering problems:
|
||||
但它有几个工程问题:
|
||||
|
||||
- Retrieval was too SDK-specific. The business code directly owned many Milvus search details.
|
||||
- L0 and L1 responsibilities were blurry. L0 keyword matching could look like a final retrieval decision instead of a hint.
|
||||
- Chunk-level retrieval could lose section context when one section was split into multiple chunks.
|
||||
- Metadata such as `breadcrumb` existed, but it was not fully used in retrieval, filtering, or context reconstruction.
|
||||
- Retrieval quality was mostly checked by manual API calls and logs, not by repeatable cases.
|
||||
- 检索实现过于依赖 Milvus SDK,业务代码承担了太多底层搜索细节。
|
||||
- L0 和 L1 职责不清,L0 关键词命中容易被当作最终召回决策。
|
||||
- chunk 级检索容易丢失章节上下文。
|
||||
- `breadcrumb` 存在 metadata 中,但没有充分参与 embedding、filter 和上下文重建。
|
||||
- 检索质量主要靠手工接口和日志判断,缺少可重复的 golden cases。
|
||||
|
||||
So the refactor goal was not "replace everything with a framework." The goal was to move generic RAG infrastructure toward Spring AI while keeping the project-specific Agent evidence chain.
|
||||
所以重构目标不是“全盘替换成框架”,而是:
|
||||
|
||||
## How I Broke The Problem Down
|
||||
```text
|
||||
通用 RAG 基础设施交给 Spring AI,
|
||||
业务可观测链路保留在项目里。
|
||||
```
|
||||
|
||||
I treated this as a staged migration, because RAG touches the Agent tool layer, AIOps diagnosis, vector retrieval, evidence packing, and database traces.
|
||||
## 2. 我如何拆解问题
|
||||
|
||||
The first step was to establish a baseline. I added retrieval evaluation cases under `eval/rag-retrieval/` so future changes could be compared against known queries instead of judged only by intuition.
|
||||
我把迁移拆成几个阶段,因为 RAG 同时影响 Agent 工具层、AIOps、向量检索、证据打包和 Trace。
|
||||
|
||||
Then I clarified the retrieval roles:
|
||||
第一步是建立 baseline。`eval/rag-retrieval/` 中的 golden cases 用来对比后续改动,而不是只靠直觉判断检索有没有变好。
|
||||
|
||||
第二步是明确职责:
|
||||
|
||||
```text
|
||||
L0 = domain/entity hint
|
||||
L1 = semantic retrieval
|
||||
postprocess = evidence shaping and trace-friendly output
|
||||
postprocess = evidence shaping + trace-friendly output
|
||||
```
|
||||
|
||||
That means L0 is still valuable, but it should not bypass semantic retrieval as the default path. It is better used to extract service names, alert names, error codes, domains, and metadata hints.
|
||||
L0 仍然有价值,但不再默认绕过语义检索。它更适合提取服务名、告警名、错误码、领域和 metadata filter。
|
||||
|
||||
After that, I added evidence postprocessing. The Agent should not just receive raw chunks; it should receive structured evidence with source, title, breadcrumb, score, hit reason, and content. This makes the result easier to inspect and easier to explain in an interview.
|
||||
第三步是增强 evidence 输出。Agent 不应该只拿到 raw chunk,而应该拿到带 source、title、breadcrumb、score、hit reason 的证据块。
|
||||
|
||||
Finally, I integrated Spring AI `VectorStore` as the main read path while preserving the original Milvus SDK implementation as fallback.
|
||||
最后,我把 Spring AI `VectorStore` 接入为读取主路径,同时保留原 Milvus SDK 作为 fallback。
|
||||
|
||||
## Current Architecture
|
||||
|
||||
The current retrieval path is:
|
||||
## 3. 当前架构
|
||||
|
||||
```text
|
||||
Agent / API
|
||||
@@ -50,160 +53,75 @@ Agent / API
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
-> evidence postprocess
|
||||
-> relevance normalization
|
||||
-> tool_invocation trace
|
||||
```
|
||||
|
||||
`VectorSearchService` is still the public retrieval facade. This is deliberate: the Agent tool layer does not need to know whether the underlying retrieval engine is SDK-based or Spring AI-based.
|
||||
`VectorSearchService` 仍然是公共检索门面。Agent 工具层不需要知道底层是 SDK 还是 Spring AI。
|
||||
|
||||
The supported retrieval modes are:
|
||||
支持三种模式:
|
||||
|
||||
```text
|
||||
auto -> try Spring AI VectorStore, fallback to SDK
|
||||
spring-ai -> force Spring AI VectorStore
|
||||
sdk -> force Milvus SDK
|
||||
auto -> 优先 Spring AI VectorStore,失败后 fallback 到 SDK
|
||||
spring-ai -> 强制 Spring AI VectorStore
|
||||
sdk -> 强制 Milvus SDK
|
||||
```
|
||||
|
||||
This keeps the migration reversible and testable.
|
||||
## 4. 关键取舍
|
||||
|
||||
## Key Tradeoffs
|
||||
### 保留显式工具
|
||||
|
||||
### Keep The Explicit Tool
|
||||
我没有把检索藏进 Spring AI Advisor。原因是这个项目强调 Agent 执行可见性:`lookup_knowledge` 的 query、命中文档、相关性和证据预览都要进入 Trace。
|
||||
|
||||
I did not hide retrieval inside a Spring AI Advisor.
|
||||
### 保留 SDK fallback
|
||||
|
||||
For this project, `lookup_knowledge` is part of the Agent execution story. It records what query was used, which evidence was retrieved, how relevant it looked, and how it supported diagnosis. If retrieval is hidden inside an advisor, the answer may still work, but the audit trail becomes harder to show.
|
||||
SDK fallback 不是废代码,而是迁移安全网。实际验证时,第一次 VectorStore 指向了错误 collection,`auto` 模式 fallback 到 SDK 后仍能返回结果。修正 collection 后,Spring AI 路径成为主路径。
|
||||
|
||||
### Keep SDK Fallback
|
||||
### L0 降权
|
||||
|
||||
The SDK path is not dead code. It is a safety net during migration.
|
||||
生产事故中经常有精确标识:错误码、告警名、服务名、指标名。L0 适合做 hint,但不应该做最终裁判。
|
||||
|
||||
This proved useful during live validation. The first VectorStore run pointed at the wrong collection name, but `auto` mode fell back to SDK and still returned results. After the collection was corrected to `biz`, the Spring AI path worked as the main path.
|
||||
### 分数语义拆开
|
||||
|
||||
### Keep L0, But Reduce Its Authority
|
||||
SDK 使用 L2 distance,Spring AI 暴露 similarity。混在一个字段里会让 relevance normalization 出错。
|
||||
|
||||
L0 is worth keeping because production incidents often contain exact identifiers:
|
||||
|
||||
- error code
|
||||
- alert name
|
||||
- service name
|
||||
- metric name
|
||||
- domain tag
|
||||
|
||||
But L0 should not be the final judge of retrieval quality. Its role is now closer to domain hint, entity extraction, metadata filtering, and explainability signal.
|
||||
|
||||
### Split Score Semantics
|
||||
|
||||
The old SDK path used L2 distance. Spring AI exposes similarity. Treating those as the same number would quietly break relevance normalization.
|
||||
|
||||
So the result separates:
|
||||
当前拆成:
|
||||
|
||||
```text
|
||||
score -> compatibility score used by existing logic
|
||||
rawScore -> raw score from the retrieval implementation
|
||||
scoreLabel -> semantic label for rawScore
|
||||
score -> 兼容旧逻辑的距离型分数
|
||||
rawScore -> 底层原始分数
|
||||
scoreLabel -> rawScore 的语义
|
||||
```
|
||||
|
||||
For SDK:
|
||||
### 暂不迁移写入
|
||||
|
||||
写入和索引仍走 SDK。这是有意分阶段:先验证读路径,再评估 `VectorStore.add(...)` 是否适合现有 metadata 和 chunk 模型。
|
||||
|
||||
## 5. 验证方式
|
||||
|
||||
我用了三层验证:
|
||||
|
||||
- 单元测试:SDK mode、Spring AI mode、auto fallback、category filter、distance metadata mapping。
|
||||
- Live API:`GET /api/search/similar?query=ERR_TIMEOUT&topK=3`。
|
||||
- 代表性 query 对比:错误码、支付超时、MySQL 连接池、AIOps 告警式 query、抽象 RAG 设计问题。
|
||||
|
||||
核心排障和 AIOps query 在 SDK 与 VectorStore 下 top3 一致。差异主要集中在抽象设计类问题和 metadata taxonomy,这些被记录为后续质量工作。
|
||||
|
||||
## 6. 面试短版
|
||||
|
||||
```text
|
||||
score = L2 distance
|
||||
rawScore = L2 distance
|
||||
scoreLabel = l2_distance
|
||||
这个 RAG 系统最初是基于 Milvus SDK 的自研 MVP。它能跑,但底层检索细节过多地散落在业务代码里,L0/L1 职责也不够清晰。
|
||||
我按阶段重构:先加 retrieval baseline,再把 L0 降级为 domain/entity hint,再增强 evidence postprocess,最后把读取主路径切到 Spring AI VectorStore,并保留 SDK fallback。
|
||||
我没有把 lookup_knowledge 替换成隐式 Advisor,因为这个项目的核心是可追踪 Agent:面试官可以看到什么时候检索、检索了什么、证据如何支撑诊断。
|
||||
```
|
||||
|
||||
For VectorStore:
|
||||
## 7. 可主动承认的不足
|
||||
|
||||
```text
|
||||
score = Milvus metadata.distance when available
|
||||
rawScore = Spring AI similarity
|
||||
scoreLabel = similarity
|
||||
```
|
||||
- metadata taxonomy 还需要清理,例如 `database` 与 `infrastructure`。
|
||||
- 抽象设计问题可能需要 query rewrite 或更好的文档索引。
|
||||
- 邻居 chunk / 同章节上下文扩展还不完整。
|
||||
- rerank、RRF、BM25、hybrid retrieval 还没有接入。
|
||||
- 写入路径仍使用 SDK。
|
||||
|
||||
This makes the migration inspectable instead of hiding score changes behind one overloaded field.
|
||||
这些不是当前迁移阻塞项,而是后续检索质量优化方向。
|
||||
|
||||
### Do Not Migrate Writes Yet
|
||||
|
||||
Writes and indexing still use the SDK path.
|
||||
|
||||
That is intentional. Migrating reads and writes at the same time would make debugging harder. The read path can be validated first; write-path migration can happen later if Spring AI `VectorStore.add(...)` fits the existing metadata and chunk model.
|
||||
|
||||
## Validation Story
|
||||
|
||||
I validated the refactor at multiple levels.
|
||||
|
||||
Unit tests cover:
|
||||
|
||||
- SDK mode.
|
||||
- Spring AI mode.
|
||||
- `auto` fallback.
|
||||
- category filter behavior.
|
||||
- distance metadata mapping.
|
||||
|
||||
Live API verification used:
|
||||
|
||||
```text
|
||||
GET /api/search/similar?query=ERR_TIMEOUT&topK=3
|
||||
```
|
||||
|
||||
Logs confirmed when the Spring AI VectorStore path was used and when fallback happened.
|
||||
|
||||
Then I compared SDK and VectorStore retrieval quality on representative queries:
|
||||
|
||||
| Query Type | Result |
|
||||
| --- | --- |
|
||||
| exact error code | same top3 |
|
||||
| payment-service timeout | same top3 |
|
||||
| MySQL connection pool | same top3 |
|
||||
| AIOps alert-style query | same top3 |
|
||||
| abstract RAG design query | same top1, VectorStore returned fewer tail results |
|
||||
| category filter | both returned zero because metadata taxonomy did not match |
|
||||
|
||||
The acceptance decision was that Spring AI VectorStore is good enough for the current MVP read path, with SDK fallback preserved.
|
||||
|
||||
## Known Gaps
|
||||
|
||||
The refactor improved the architecture, but it did not solve every retrieval-quality problem.
|
||||
|
||||
Known gaps:
|
||||
|
||||
- Metadata taxonomy still needs cleanup, for example `database` vs `infrastructure`.
|
||||
- Abstract design questions may need query rewriting or better indexed interview/devflow documents.
|
||||
- Chunk context reconstruction is still limited when one logical section spans multiple chunks.
|
||||
- `breadcrumb` now participates in embedding text, but it can still be used more strongly in context expansion, rerank, and evidence packing.
|
||||
- Rerank, RRF, BM25, and hybrid retrieval are not implemented yet.
|
||||
- Indexing writes still use SDK.
|
||||
|
||||
These are good follow-up issues because they are retrieval-quality improvements, not blockers for the VectorStore migration.
|
||||
|
||||
## How I Present This In An Interview
|
||||
|
||||
My short version would be:
|
||||
|
||||
> This RAG system started as a self-built MVP around Milvus SDK retrieval. It worked, but too much infrastructure logic lived in business code, and L0/L1 responsibilities were unclear. I refactored it in stages: first I added baseline retrieval cases, then made L0 a domain/entity hint instead of a final decision layer, then added evidence postprocessing, and finally moved the main read path to Spring AI VectorStore with SDK fallback. I kept `lookup_knowledge` as an explicit Agent tool because the project values traceability: the interviewer can see when retrieval happened, what evidence was found, and how it supported the diagnosis. The result is closer to standard Spring AI RAG while still preserving business-specific observability.
|
||||
|
||||
If asked why this is not a full framework migration:
|
||||
|
||||
> I intentionally did not migrate everything at once. Reads moved first because they are easier to compare using golden queries. Writes/indexing stayed on SDK to avoid mixing schema and retrieval behavior changes in one step. Advisors were not used as the main interface because hidden retrieval would weaken the Agent trace.
|
||||
|
||||
If asked what I would improve next:
|
||||
|
||||
> I would add query transformation for AIOps payloads, improve metadata taxonomy, use breadcrumb and section metadata for context expansion, and then evaluate whether hybrid retrieval or rerank is necessary based on measured recall and topK overlap.
|
||||
|
||||
## Interview Follow-Up Questions
|
||||
|
||||
### Why introduce Spring AI VectorStore if the SDK path already worked?
|
||||
|
||||
Because SDK-only retrieval made the project own too much low-level RAG infrastructure. `VectorStore` gives a standard abstraction for retrieval and makes future Spring AI features easier to adopt, while the facade keeps the Agent layer stable.
|
||||
|
||||
### Why keep custom code at all?
|
||||
|
||||
The custom code is where the Agent engineering value lives: AIOps payload mapping, L0 hints, evidence packing, score compatibility, and tool invocation tracing. Those are domain-specific and should remain visible.
|
||||
|
||||
### How do you know quality did not regress?
|
||||
|
||||
I compared SDK and VectorStore modes on representative live queries. Core troubleshooting and AIOps cases returned the same top3 documents in the same order. The differences were isolated to abstract design queries and metadata taxonomy, which are documented follow-up work.
|
||||
|
||||
### What is the most important design decision?
|
||||
|
||||
Keeping a stable boundary: `lookup_knowledge` calls `VectorSearchService`, and `VectorSearchService` decides whether to use Spring AI or SDK. That boundary made the migration small enough to validate and explain.
|
||||
|
||||
Reference in New Issue
Block a user