Files
zhuyongxin 584639fa2a docs(mvp): move engineering notes under mvp/engineering
Relocate RAG and diagnosis decision/E2E writeups from docs/ root into
mvp/engineering so architecture, issues, and engineering narrative stay
together. Update indexes and cross-links; leave docs/learning as legacy.
2026-07-29 10:49:45 +08:00

3.9 KiB

Design: RAG chunk evidence identity and dedup

Context

  • Modular RAG pipeline already exists (KnowledgeQueryTransformer → retriever → post-processor → packer → assembler → projector).
  • Delivery baseline: mvp/engineering/rag/Milvus-Hybrid接入清单.md §1.1 Delivery 1.
  • Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary.
  • Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic.

Goals / Non-Goals

Goals

  • Preserve multiple relevant chunks from the same document in one lookup.
  • Make evidence identity stable enough for future hybrid hits.
  • Separate retrieve width from return width.
  • Keep Agent tool name/input unchanged (lookup_knowledge(query)).

Non-Goals

  • Hybrid / BM25 / sparse schema.
  • Cross-encoder rerank.
  • Deleting SDK mode.
  • Session dedup tracker.

Decisions

D1. Evidence identity key

evidenceKey =
  if docId != null && chunkIndex != null:
    docId + "#chunk-" + chunkIndex
  else if vectorId != null:
    "vector:" + vectorId
  else:
    "rank:" + originalRank

Rationale: works with current metadata (docId, chunkIndex) and degrades safely for older rows.

D2. Dedup granularity

  • Dedup key = evidenceKey only.
  • True duplicates (same key) merge hitReasons; keep higher-ranked content (first after score sort).
  • Do not merge different chunks of the same source into one content block.

D3. Per-document cap

  • After score sort, accept at most rag.max-chunks-per-document (default 2) evidence blocks per docId.
  • If docId missing, treat each evidenceKey as its own document bucket.

D4. retrieve-k / return-n

rag.retrieve-k=20   # vector recall width
rag.return-n=5      # max evidence blocks after post-process (before projector budget)
rag.max-chunks-per-document=2

Compatibility:

  • If only legacy rag.top-k is set, use it as fallback for both until removed.
  • Prefer explicit retrieve-k/return-n when present.

D5. Agent projection identity (behavior change)

  • Projected document_id SHOULD be evidenceKey (chunk-scoped), not raw source.
  • This is intentional so EvidenceGuard references remain 1:1 with returned excerpts.
  • source remains human-readable document path/id and MAY repeat across chunks.
  • Marked as observable behavior change for Agent consumers and report references.

D6. KnowledgeSearchPort (thin)

KnowledgeSearchPort.search(KnowledgeSearchRequest) -> List<KnowledgeSearchHit>
  • Default adapter delegates to existing VectorSearchService.searchSimilarDocuments.
  • Request carries query, topK, categoryFilter, mode placeholder (DENSE only in this change).
  • Hit carries id, content, score fields, metadata map, and extracted identity fields when available.
  • No hybrid implementation in this change.

Module flow

LookupKnowledge
  -> transform (L0 hints)
  -> KnowledgeSearchPort.search(retrieveK, filter)
  -> map to RetrievedEvidenceCandidate (+ identity)
  -> post-process:
       score/sort
       evidenceKey dedup
       maxChunksPerDocument
       return-n truncate
  -> pack + assemble
  -> RagResultProjector (dedupe by evidence document_id=evidenceKey)

Interface impact

Level What
L2 Internal DTO fields on candidate/EvidenceBlock
L3 Agent-facing document_id becomes chunk-scoped evidence id

Migration/compat:

  • In-repo guards/tests updated to accept chunk-scoped ids.
  • External human readers still see source/title.

Risks / Trade-offs

Risk Mitigation
Larger Agent payload return-n + maxChunksPerDocument + existing projector budgets
Missing chunkIndex in old data vector id fallback keeps chunks distinct
document_id semantic shift documented; projector tests updated

Open questions

None remaining for Delivery 1. Hybrid belongs to next change.