# Design: RAG chunk evidence identity and dedup ## Context - Modular RAG pipeline already exists (`KnowledgeQueryTransformer` → retriever → post-processor → packer → assembler → projector). - Delivery baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1. - Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary. - Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic. ## Goals / Non-Goals **Goals** - Preserve multiple relevant chunks from the same document in one lookup. - Make evidence identity stable enough for future hybrid hits. - Separate retrieve width from return width. - Keep Agent tool name/input unchanged (`lookup_knowledge(query)`). **Non-Goals** - Hybrid / BM25 / sparse schema. - Cross-encoder rerank. - Deleting SDK mode. - Session dedup tracker. ## Decisions ### D1. Evidence identity key ```text evidenceKey = if docId != null && chunkIndex != null: docId + "#chunk-" + chunkIndex else if vectorId != null: "vector:" + vectorId else: "rank:" + originalRank ``` Rationale: works with current metadata (`docId`, `chunkIndex`) and degrades safely for older rows. ### D2. Dedup granularity - Dedup key = `evidenceKey` only. - True duplicates (same key) merge hitReasons; keep higher-ranked content (first after score sort). - Do **not** merge different chunks of the same source into one content block. ### D3. Per-document cap - After score sort, accept at most `rag.max-chunks-per-document` (default `2`) evidence blocks per `docId`. - If `docId` missing, treat each evidenceKey as its own document bucket. ### D4. retrieve-k / return-n ```properties rag.retrieve-k=20 # vector recall width rag.return-n=5 # max evidence blocks after post-process (before projector budget) rag.max-chunks-per-document=2 ``` Compatibility: - If only legacy `rag.top-k` is set, use it as fallback for both until removed. - Prefer explicit retrieve-k/return-n when present. ### D5. Agent projection identity (behavior change) - Projected `document_id` SHOULD be `evidenceKey` (chunk-scoped), not raw source. - This is intentional so EvidenceGuard references remain 1:1 with returned excerpts. - `source` remains human-readable document path/id and MAY repeat across chunks. - Marked as **observable behavior change** for Agent consumers and report references. ### D6. KnowledgeSearchPort (thin) ```text KnowledgeSearchPort.search(KnowledgeSearchRequest) -> List ``` - Default adapter delegates to existing `VectorSearchService.searchSimilarDocuments`. - Request carries query, topK, categoryFilter, mode placeholder (`DENSE` only in this change). - Hit carries id, content, score fields, metadata map, and extracted identity fields when available. - No hybrid implementation in this change. ## Module flow ```text LookupKnowledge -> transform (L0 hints) -> KnowledgeSearchPort.search(retrieveK, filter) -> map to RetrievedEvidenceCandidate (+ identity) -> post-process: score/sort evidenceKey dedup maxChunksPerDocument return-n truncate -> pack + assemble -> RagResultProjector (dedupe by evidence document_id=evidenceKey) ``` ## Interface impact | Level | What | |---|---| | L2 | Internal DTO fields on candidate/EvidenceBlock | | L3 | Agent-facing `document_id` becomes chunk-scoped evidence id | Migration/compat: - In-repo guards/tests updated to accept chunk-scoped ids. - External human readers still see `source`/`title`. ## Risks / Trade-offs | Risk | Mitigation | |---|---| | Larger Agent payload | return-n + maxChunksPerDocument + existing projector budgets | | Missing chunkIndex in old data | vector id fallback keeps chunks distinct | | document_id semantic shift | documented; projector tests updated | ## Open questions None remaining for Delivery 1. Hybrid belongs to next change.