docs(rag): add comments on knowledge retrieval pipeline

Document the lookup_knowledge flow from L0 hints through L1 retrieval,
post-processing, packing, and Agent projection so the boundaries and
current limitations are easier to follow.
This commit is contained in:
zhuyongxin
2026-07-27 16:26:06 +08:00
parent edb6153fd6
commit 99d4f6f216
19 changed files with 480 additions and 37 deletions
@@ -13,8 +13,21 @@ import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* 文档分片服务
* 负责将长文档切分为多个有语义完整性的小片段
* 文档切片服务(RAG 入库前处理)。
*
* <p>把长 Markdown/文本切成带 title/breadcrumb 的 {@link com.superbiz.agent.dto.DocumentChunk},
* 供 {@link VectorIndexService} 向量化。</p>
*
* <h3>策略摘要</h3>
* <ol>
* <li>先按 Markdown 标题分 section,并维护 breadcrumb 层级</li>
* <li>section 过长再按段落累积;用 token 估算做软边界 / 硬上限</li>
* <li>尽量不在有序/无序列表或未闭合代码块中间切断</li>
* <li>相邻 chunk 保留 overlap,减轻边界语义断裂</li>
* </ol>
*
* <p>检索命中单个 chunk 后,当前主链路不会自动回补同章节相邻 chunk
* (上下文重建仍是后续增强点)。</p>
*/
@Service
public class DocumentChunkService {