Archive pre-refactor interview notes and add current deep-dives on architecture evolution, issue-derived stories, and evidence gates.
120 lines
3.8 KiB
Markdown
120 lines
3.8 KiB
Markdown
# RAG VectorStore 面试要点
|
||
|
||
## 1. 60 秒讲法
|
||
|
||
```text
|
||
我把 RAG 检索从 Milvus SDK-only 重构为 Spring AI VectorStore 主路径,同时保留 SDK fallback。
|
||
关键不是换了一个依赖,而是保留 VectorSearchService 作为边界,所以 lookup_knowledge 和 Agent workflow 不需要改。
|
||
现在支持 auto、spring-ai、sdk 三种模式。auto 会优先尝试 VectorStore,失败后 fallback 到 SDK。
|
||
```
|
||
|
||
现场验证时,第一次发现 VectorStore 指向了错误 collection:`business_knowledge`,而实际 Zilliz collection 是 `biz`。fallback 生效,所以系统仍能通过 SDK 返回结果。修正 collection 后,同一个 query 成功走 Spring AI VectorStore。
|
||
|
||
## 2. 架构回答
|
||
|
||
```text
|
||
Agent / API
|
||
-> lookup_knowledge or /api/search/similar
|
||
-> VectorSearchService
|
||
-> Spring AI VectorStore
|
||
-> Milvus SDK fallback
|
||
-> Milvus/Zilliz collection: biz
|
||
```
|
||
|
||
关键设计:`VectorSearchService` 是检索门面,避免 Spring AI 或 SDK 细节扩散到 Agent 工具层。
|
||
|
||
## 3. 为什么保留 SDK
|
||
|
||
- 迁移安全:原 SDK 路径已验证可用。
|
||
- 运行韧性:VectorStore schema、filter 或配置失败时,检索仍可用。
|
||
- Demo 稳定:检索抽象变化不应该破坏主诊断演示。
|
||
|
||
这在实际验证中发挥了作用:VectorStore 配置错时,`auto` 模式 fallback 到 SDK,API 没有失败。
|
||
|
||
## 4. 为什么引入 Spring AI VectorStore
|
||
|
||
使用 `VectorStore` 可以让项目更接近标准 RAG 抽象:
|
||
|
||
- 业务代码不再持有全部 Milvus search 细节。
|
||
- 后续 QueryTransformer、DocumentPostProcessor、Retriever 等能力更容易接入。
|
||
- 面试中也更容易解释和 Spring AI 生态的关系。
|
||
|
||
但我没有一次性迁移写入,因为读写同时迁移会让问题难定位。当前先稳定读路径。
|
||
|
||
## 5. 为什么保留 L0
|
||
|
||
L0 现在不是最终答案来源,而是确定性 hint 层:
|
||
|
||
- 提取 domain/entity。
|
||
- 在可能时生成 category filter。
|
||
- 给 trace 提供解释信号。
|
||
|
||
当前职责:
|
||
|
||
```text
|
||
L0 = domain/entity hint
|
||
L1 = semantic retrieval through VectorStore/SDK
|
||
postprocess = evidence trace + relevance normalization
|
||
```
|
||
|
||
真实故障诊断里有很多精确标识,完全只靠向量检索并不稳。
|
||
|
||
## 6. 为什么不用隐藏 Advisor
|
||
|
||
`lookup_knowledge` 保持显式工具,因为:
|
||
|
||
- Trace 要展示什么时候检索。
|
||
- `tool_invocation` 要记录输入、输出预览、相关性和 metadata。
|
||
- 面试故事是可审计 Agent 执行,而不只是答案质量。
|
||
|
||
Advisor 后续可以接入,但需要先解决可观测性。
|
||
|
||
## 7. 分数设计
|
||
|
||
当前结果故意拆成:
|
||
|
||
```text
|
||
score -> 兼容旧 relevance normalization 的分数
|
||
rawScore -> 当前检索实现原始分数
|
||
scoreLabel -> rawScore 的语义
|
||
```
|
||
|
||
SDK:
|
||
|
||
```text
|
||
score = L2 distance
|
||
rawScore = L2 distance
|
||
scoreLabel = l2_distance
|
||
```
|
||
|
||
VectorStore:
|
||
|
||
```text
|
||
score = metadata.distance if present
|
||
rawScore = Spring AI document score
|
||
scoreLabel = similarity
|
||
```
|
||
|
||
这样避免把 similarity 当成 L2 distance 的隐蔽 bug。
|
||
|
||
## 8. 如何证明 VectorStore 被使用
|
||
|
||
- 日志出现 `Starting Spring AI VectorStore search` 和 `Spring AI VectorStore search complete`。
|
||
- API 响应中 `scoreLabel=similarity`。
|
||
- `rawScore` 是 Spring AI similarity,`score` 仍是兼容 distance。
|
||
|
||
## 9. 常见追问
|
||
|
||
### 为什么不删 SDK?
|
||
|
||
这是迁移,不是重写。fallback 提供回滚安全,并且已经证明配置错误时仍能保证主链路可用。
|
||
|
||
### `lookup_knowledge` 变了吗?
|
||
|
||
外部契约没变。它仍然调用 `VectorSearchService.searchSimilarDocuments(...)`,变化在门面背后的实现。
|
||
|
||
### 这是完整 Spring AI RAG 了吗?
|
||
|
||
还不是。当前是 Spring AI VectorStore 读路径 + 显式工具 + 自定义 evidence trace + SDK 写入。这样做是为了保留审计能力和分阶段迁移安全。
|
||
|