chore(rag): add eval knowledge base mirror

This commit is contained in:
zhuyongxin
2026-07-06 21:48:21 +08:00
parent ed7efc58b7
commit 64adb998cf
10 changed files with 231 additions and 0 deletions
@@ -0,0 +1,28 @@
---
title: RAG Chunk Context Reconstruction
keywords: [split into multiple chunks, retrieval context, neighbor chunk, same section, breadcrumb context]
summary: Preserve context when long RAG sections are split into multiple retrievable chunks.
category: rag
source: rag-chunk-context-reconstruction
breadcrumb: RAG > Chunking > Context Reconstruction
kb_scope: rag-eval
covers: [rag chunking, context packing, breadcrumbs]
when_to_retrieve: Use when a retrieval question asks how to preserve context across split chunks.
---
# RAG
## Chunking
### Context Reconstruction
When a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.
Recommended behavior:
1. Store the breadcrumb with every chunk.
2. Preserve the same section identity across adjacent chunks.
3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.
4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.
The key concepts are neighbor chunk, same section, and breadcrumb.
@@ -0,0 +1,29 @@
---
title: RAG L0 Domain Entity Hint
keywords: [L0 keyword matching, final retrieval result, domain detector, entity extractor, metadata filter]
summary: Define L0 as a query transformation hint layer instead of final retrieval evidence.
category: rag
source: rag-l0-domain-entity-hint
breadcrumb: RAG > L0 > Domain Entity Hint
kb_scope: rag-eval
covers: [l0 hint, query transformation, metadata filter]
when_to_retrieve: Use when a question asks whether L0 should decide final retrieval or only provide hints.
---
# RAG
## L0
### Domain Entity Hint
L0 keyword matching should not decide the final retrieval result.
In the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.
The output can provide:
1. Candidate domain hints.
2. Matched entities and keywords.
3. An optional metadata filter for the first vector retrieval attempt.
Final evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.
The important terms are domain detector, entity extractor, and metadata filter.