Preserve same-document multi-chunk evidence with evidenceKey identity, per-document caps, retrieve-k/return-n split, and a dense KnowledgeSearchPort. Archives Delivery 1 OpenSpec change as the foundation for hybrid retrieval.
3.9 KiB
3.9 KiB
Design: RAG chunk evidence identity and dedup
Context
- Modular RAG pipeline already exists (
KnowledgeQueryTransformer→ retriever → post-processor → packer → assembler → projector). - Delivery baseline:
docs/milvus-hybrid-search-integration-checklist.md§1.1 Delivery 1. - Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary.
- Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic.
Goals / Non-Goals
Goals
- Preserve multiple relevant chunks from the same document in one lookup.
- Make evidence identity stable enough for future hybrid hits.
- Separate retrieve width from return width.
- Keep Agent tool name/input unchanged (
lookup_knowledge(query)).
Non-Goals
- Hybrid / BM25 / sparse schema.
- Cross-encoder rerank.
- Deleting SDK mode.
- Session dedup tracker.
Decisions
D1. Evidence identity key
evidenceKey =
if docId != null && chunkIndex != null:
docId + "#chunk-" + chunkIndex
else if vectorId != null:
"vector:" + vectorId
else:
"rank:" + originalRank
Rationale: works with current metadata (docId, chunkIndex) and degrades safely for older rows.
D2. Dedup granularity
- Dedup key =
evidenceKeyonly. - True duplicates (same key) merge hitReasons; keep higher-ranked content (first after score sort).
- Do not merge different chunks of the same source into one content block.
D3. Per-document cap
- After score sort, accept at most
rag.max-chunks-per-document(default2) evidence blocks perdocId. - If
docIdmissing, treat each evidenceKey as its own document bucket.
D4. retrieve-k / return-n
rag.retrieve-k=20 # vector recall width
rag.return-n=5 # max evidence blocks after post-process (before projector budget)
rag.max-chunks-per-document=2
Compatibility:
- If only legacy
rag.top-kis set, use it as fallback for both until removed. - Prefer explicit retrieve-k/return-n when present.
D5. Agent projection identity (behavior change)
- Projected
document_idSHOULD beevidenceKey(chunk-scoped), not raw source. - This is intentional so EvidenceGuard references remain 1:1 with returned excerpts.
sourceremains human-readable document path/id and MAY repeat across chunks.- Marked as observable behavior change for Agent consumers and report references.
D6. KnowledgeSearchPort (thin)
KnowledgeSearchPort.search(KnowledgeSearchRequest) -> List<KnowledgeSearchHit>
- Default adapter delegates to existing
VectorSearchService.searchSimilarDocuments. - Request carries query, topK, categoryFilter, mode placeholder (
DENSEonly in this change). - Hit carries id, content, score fields, metadata map, and extracted identity fields when available.
- No hybrid implementation in this change.
Module flow
LookupKnowledge
-> transform (L0 hints)
-> KnowledgeSearchPort.search(retrieveK, filter)
-> map to RetrievedEvidenceCandidate (+ identity)
-> post-process:
score/sort
evidenceKey dedup
maxChunksPerDocument
return-n truncate
-> pack + assemble
-> RagResultProjector (dedupe by evidence document_id=evidenceKey)
Interface impact
| Level | What |
|---|---|
| L2 | Internal DTO fields on candidate/EvidenceBlock |
| L3 | Agent-facing document_id becomes chunk-scoped evidence id |
Migration/compat:
- In-repo guards/tests updated to accept chunk-scoped ids.
- External human readers still see
source/title.
Risks / Trade-offs
| Risk | Mitigation |
|---|---|
| Larger Agent payload | return-n + maxChunksPerDocument + existing projector budgets |
| Missing chunkIndex in old data | vector id fallback keeps chunks distinct |
| document_id semantic shift | documented; projector tests updated |
Open questions
None remaining for Delivery 1. Hybrid belongs to next change.