Files
SuperBizAgent-java/openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup/proposal.md
T
zhuyongxin 584639fa2a docs(mvp): move engineering notes under mvp/engineering
Relocate RAG and diagnosis decision/E2E writeups from docs/ root into
mvp/engineering so architecture, issues, and engineering narrative stay
together. Update indexes and cross-links; leave docs/learning as legacy.
2026-07-29 10:49:45 +08:00

2.5 KiB

Change: RAG chunk evidence identity and dedup

Why

lookup_knowledge currently collapses same-document chunks at two layers:

  1. KnowledgeEvidencePostProcessor dedupes by source/title/breadcrumb
  2. RagResultProjector dedupes by document_id, which usually falls back to source

As a result, L1 can recall multiple useful chunks from one document, but Agent often sees only one. This blocks multi-path / hybrid retrieval benefits and hurts long runbook completeness.

This change is Delivery 1 from mvp/engineering/rag/Milvus-Hybrid接入清单.md. Hybrid schema/search is out of scope and will be a separate change after this is archived.

What Changes

  • Add chunk-level evidence identity: docId, chunkIndex, evidenceKey
  • Extract identity in retrieval adapter from vector metadata
  • Deduplicate by evidenceKey (chunk identity), not document source
  • Cap chunks per document (maxChunksPerDocument, default 2)
  • Split retrieve-k and return-n (stop overloading single top-k)
  • Align Agent projection so same-source different chunks can both appear
  • Introduce a thin KnowledgeSearchPort so later hybrid can swap implementation without rewriting the pipeline
  • Keep L0 as hint-only; no BM25/sparse schema; no legacy SDK deletion in this change

Non-goals

  • Milvus sparse/BM25 schema or reindex
  • Enabling hybrid search mode
  • Removing Milvus SDK path
  • Session-level RetrievedDocTracker restore
  • Neighbor chunk context reconstruction
  • Model reranker

Capabilities

New Capabilities

  • rag-chunk-evidence-identity: chunk-level identity, dedup, retrieve/return split, search port boundary

Modified Capabilities

  • rag-knowledge-retrieval: replace document-level evidence dedup requirement with chunk-level identity
  • rag-log-projections (RAG portion only): projection identity may be chunk-scoped document_id

Impact

  • Code: retrieval DTO/services, post-processor, lookup tool config, projector, tests
  • Agent-visible: more evidence items possible for same logical document when multiple chunks are relevant
  • Interface level: L2/L3 — Agent document_id semantics become chunk-scoped evidence id (often docId#chunk-N); EvidenceGuard still validates against tool projection ids
  • Docs baseline: mvp/engineering/rag/Milvus-Hybrid接入清单.md §1.1 Delivery 1

Risks

  • Agent context grows if many chunks pass; mitigated by return-n and maxChunksPerDocument
  • Existing tests assume source-level dedup; must update intentionally
  • Old indexes without chunkIndex need stable fallback keys (vector:{id})