docs: reorganize MVP interview documentation
This commit is contained in:
@@ -1,8 +1,8 @@
|
||||
# RAG Breadcrumb Embedding Acceptance
|
||||
# RAG Breadcrumb Embedding 验收说明
|
||||
|
||||
## What Changed
|
||||
## 1. 改动是什么
|
||||
|
||||
The indexing path now builds embedding text from chunk structure plus content:
|
||||
索引路径现在构造 embedding 文本时,不只使用 chunk 内容,还会把结构上下文拼进去:
|
||||
|
||||
```text
|
||||
Title: {title}
|
||||
@@ -11,56 +11,81 @@ Content:
|
||||
{content}
|
||||
```
|
||||
|
||||
The stored Milvus `content` field remains the original chunk content. This keeps display and evidence output clean while allowing the vector to carry section-level semantics.
|
||||
Milvus 中存储的 `content` 字段仍然保留原始 chunk 内容。这样展示和证据输出保持干净,而向量本身携带章节语义。
|
||||
|
||||
## Why Reindex Is Required
|
||||
## 2. 为什么必须重新索引
|
||||
|
||||
Embeddings are materialized at index time. Existing vectors were generated from the previous content-only text, so they cannot benefit from `title` and `breadcrumb` until the knowledge base is reindexed.
|
||||
Embedding 是索引时物化的。已有向量是用旧的 content-only 文本生成的,所以只有代码变化并不会改变线上检索结果。
|
||||
|
||||
This is the key acceptance point:
|
||||
验收关键点:
|
||||
|
||||
```text
|
||||
code change alone != live retrieval changed
|
||||
code change + reindex + live query report = accepted behavior
|
||||
只改代码 != live retrieval 已变化
|
||||
代码改动 + 重新索引 + live query report = 行为验收完成
|
||||
```
|
||||
|
||||
## How To Validate
|
||||
## 3. 如何验证
|
||||
|
||||
1. Start the Spring Boot application.
|
||||
2. Reindex the knowledge base through the existing indexing path.
|
||||
3. Run:
|
||||
1. 启动 Spring Boot 应用。
|
||||
2. 通过现有索引路径重新索引知识库。
|
||||
3. 运行:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py
|
||||
```
|
||||
|
||||
The script writes:
|
||||
脚本输出:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/reports/live-post-reindex.json
|
||||
eval/rag-retrieval/reports/live-post-reindex.md
|
||||
```
|
||||
|
||||
The default cases cover:
|
||||
默认覆盖:
|
||||
|
||||
- RAG chunk context questions where breadcrumb matters.
|
||||
- Diagnosis flow questions where section path matters.
|
||||
- `ERR_TIMEOUT` exact error-code retrieval.
|
||||
- MySQL connection pool troubleshooting.
|
||||
- AIOps payment-service latency alert retrieval.
|
||||
- breadcrumb 敏感的 RAG chunk context query。
|
||||
- 需要章节路径的诊断流程问题。
|
||||
- `ERR_TIMEOUT` 精确错误码检索。
|
||||
- MySQL 连接池排障。
|
||||
- AIOps payment-service 延迟告警检索。
|
||||
|
||||
## What To Look For
|
||||
## 4. 看什么结果
|
||||
|
||||
For breadcrumb-sensitive cases, inspect whether top candidates expose expected `title` and `breadcrumb` values in the report.
|
||||
对 breadcrumb 敏感 case:
|
||||
|
||||
For core troubleshooting cases, check that result counts and top candidates remain stable. The goal is not to prove a full benchmark; it is to prove that reindexing did not obviously break important demo retrieval paths.
|
||||
- top candidates 是否暴露预期 `title`。
|
||||
- top candidates 是否暴露预期 `breadcrumb`。
|
||||
- 命中内容是否能看出所属章节。
|
||||
|
||||
## Interview Answer
|
||||
对核心排障 case:
|
||||
|
||||
If asked how I verified the breadcrumb embedding change:
|
||||
- 结果数量是否稳定。
|
||||
- top candidates 是否仍然命中核心文档。
|
||||
- 没有因为拼接 title/breadcrumb 导致核心检索退化。
|
||||
|
||||
> I separated deterministic regression from live acceptance. The offline fixture baseline still runs without services. But because embedding changes only affect newly indexed vectors, I added a live post-reindex acceptance script. It calls the real `/api/search/similar` endpoint against representative breadcrumb-sensitive, troubleshooting, and AIOps queries, then writes JSON and Markdown reports. This lets me prove both that the code changed and that the live vector collection was refreshed.
|
||||
## 5. 面试回答
|
||||
|
||||
If asked why the script does not reindex automatically:
|
||||
如果被问:你怎么验证 breadcrumb 参与 embedding 后真的生效?
|
||||
|
||||
```text
|
||||
我把 deterministic regression 和 live acceptance 分开。
|
||||
离线 fixture baseline 不依赖服务,可以做稳定回归。
|
||||
但 embedding 改动只会影响新生成的向量,所以我另外加了 live post-reindex acceptance 脚本。
|
||||
脚本会调用真实 /api/search/similar,对 breadcrumb 敏感、排障和 AIOps query 生成 JSON/Markdown 报告。
|
||||
这样能证明代码改了,也能证明 live vector collection 已经刷新。
|
||||
```
|
||||
|
||||
如果被问:为什么脚本不自动 reindex?
|
||||
|
||||
```text
|
||||
reindex 会修改向量库,而且依赖环境中的知识库数据。
|
||||
我把 reindex 保持为显式动作,验收脚本只做读取验证。
|
||||
这样如果检索没有改善,我能区分是代码问题、索引未刷新,还是运行时检索行为问题。
|
||||
```
|
||||
|
||||
## 6. 后续增强
|
||||
|
||||
- 将 live acceptance 结果加入面试 Demo 输出。
|
||||
- 增加 breadcrumb hit rate 统计。
|
||||
- 对同章节 chunk 做邻居扩展,进一步利用 breadcrumb。
|
||||
|
||||
> Reindexing mutates the vector store and depends on environment-specific data. I kept mutation explicit and made the script validation-only. That makes failures easier to diagnose: if retrieval does not improve, I can distinguish code changes, reindex state, and runtime retrieval behavior.
|
||||
|
||||
Reference in New Issue
Block a user