Files
SuperBizAgent-java/openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup/proposal.md
T
zhuyongxin ac1f831903 feat(rag): chunk evidence identity, dedup, and search port
Preserve same-document multi-chunk evidence with evidenceKey identity,
per-document caps, retrieve-k/return-n split, and a dense KnowledgeSearchPort.
Archives Delivery 1 OpenSpec change as the foundation for hybrid retrieval.
2026-07-27 18:26:15 +08:00

2.5 KiB

Change: RAG chunk evidence identity and dedup

Why

lookup_knowledge currently collapses same-document chunks at two layers:

  1. KnowledgeEvidencePostProcessor dedupes by source/title/breadcrumb
  2. RagResultProjector dedupes by document_id, which usually falls back to source

As a result, L1 can recall multiple useful chunks from one document, but Agent often sees only one. This blocks multi-path / hybrid retrieval benefits and hurts long runbook completeness.

This change is Delivery 1 from docs/milvus-hybrid-search-integration-checklist.md. Hybrid schema/search is out of scope and will be a separate change after this is archived.

What Changes

  • Add chunk-level evidence identity: docId, chunkIndex, evidenceKey
  • Extract identity in retrieval adapter from vector metadata
  • Deduplicate by evidenceKey (chunk identity), not document source
  • Cap chunks per document (maxChunksPerDocument, default 2)
  • Split retrieve-k and return-n (stop overloading single top-k)
  • Align Agent projection so same-source different chunks can both appear
  • Introduce a thin KnowledgeSearchPort so later hybrid can swap implementation without rewriting the pipeline
  • Keep L0 as hint-only; no BM25/sparse schema; no legacy SDK deletion in this change

Non-goals

  • Milvus sparse/BM25 schema or reindex
  • Enabling hybrid search mode
  • Removing Milvus SDK path
  • Session-level RetrievedDocTracker restore
  • Neighbor chunk context reconstruction
  • Model reranker

Capabilities

New Capabilities

  • rag-chunk-evidence-identity: chunk-level identity, dedup, retrieve/return split, search port boundary

Modified Capabilities

  • rag-knowledge-retrieval: replace document-level evidence dedup requirement with chunk-level identity
  • rag-log-projections (RAG portion only): projection identity may be chunk-scoped document_id

Impact

  • Code: retrieval DTO/services, post-processor, lookup tool config, projector, tests
  • Agent-visible: more evidence items possible for same logical document when multiple chunks are relevant
  • Interface level: L2/L3 — Agent document_id semantics become chunk-scoped evidence id (often docId#chunk-N); EvidenceGuard still validates against tool projection ids
  • Docs baseline: docs/milvus-hybrid-search-integration-checklist.md §1.1 Delivery 1

Risks

  • Agent context grows if many chunks pass; mitigated by return-n and maxChunksPerDocument
  • Existing tests assume source-level dedup; must update intentionally
  • Old indexes without chunkIndex need stable fallback keys (vector:{id})