Files
zhuyongxin ac1f831903 feat(rag): chunk evidence identity, dedup, and search port
Preserve same-document multi-chunk evidence with evidenceKey identity,
per-document caps, retrieve-k/return-n split, and a dense KnowledgeSearchPort.
Archives Delivery 1 OpenSpec change as the foundation for hybrid retrieval.
2026-07-27 18:26:15 +08:00

4.0 KiB

rag-chunk-evidence-identity Specification

Purpose

Define chunk-level evidence identity, deduplication, retrieve/return separation, and the thin search port boundary used by lookup_knowledge before hybrid retrieval is introduced.

ADDED Requirements

Requirement: Retrieval candidates SHALL carry chunk-level identity

The knowledge retrieval pipeline SHALL attach stable identity fields to each retrieved candidate and resulting evidence block: document id when available, chunk index when available, and an evidenceKey derived by the identity rules in design.

Scenario: Identity extracted from metadata

  • WHEN a vector hit includes metadata docId (or doc_id) and chunkIndex (or chunk_index)
  • THEN the candidate SHALL set docId, chunkIndex, and evidenceKey to docId#chunk-{chunkIndex}

Scenario: Fallback identity without chunk index

  • WHEN chunk index is missing but vector id is present
  • THEN the candidate SHALL use evidenceKey of the form vector:{id}

Requirement: Evidence deduplication SHALL be chunk-scoped

Post-processing SHALL treat two candidates as duplicates only when they share the same evidenceKey. Different chunks of the same source SHALL remain separate evidence blocks subject to per-document caps.

Scenario: Same document different chunks are kept

  • WHEN two candidates share the same source/docId but different chunk indexes
  • THEN post-processing SHALL keep both as separate evidence blocks unless a per-document cap removes the lower-ranked one

Scenario: True duplicate keys merge without replacing higher-ranked content

  • WHEN two candidates share the same evidenceKey
  • THEN post-processing SHALL keep a single block
  • AND SHALL preserve the higher-ranked content
  • AND MAY merge hit reasons

Requirement: Post-processing SHALL enforce max chunks per document

The system SHALL limit accepted evidence blocks per document id using configuration rag.max-chunks-per-document with default 2.

Scenario: Excess chunks from one document are dropped

  • WHEN more than N ranked chunks belong to the same docId and N equals the configured max
  • THEN only the top N by ranking score SHALL remain in evidence blocks

Requirement: Retrieval width and return width SHALL be separate

The lookup pipeline SHALL use rag.retrieve-k for vector recall width and rag.return-n for maximum evidence blocks after post-processing. A legacy rag.top-k MAY act as fallback when the new keys are absent.

Scenario: retrieve-k widens recall without unbounded return

  • WHEN retrieve-k is 20 and return-n is 5
  • THEN the search port is asked for up to 20 candidates
  • AND the assembled result contains at most 5 evidence blocks before Agent projection budgets

Requirement: Agent projection SHALL preserve distinct chunk evidence

The RAG projector SHALL deduplicate Agent-facing evidence by chunk-scoped evidence identity. It SHALL NOT drop a second chunk solely because source matches a previous block.

Scenario: Same source different evidence keys both project

  • WHEN two evidence blocks have different evidence keys (or chunk-scoped document ids) and the same source
  • AND projection budgets still allow both
  • THEN both SHALL appear in the projected evidence list

Scenario: Projected document_id is chunk-scoped

  • WHEN an evidence block has evidenceKey docA#chunk-2
  • THEN the projected document_id SHALL be that evidenceKey (or an equivalent chunk-scoped id)
  • AND source MAY still be the document path or doc id

Requirement: Lookup pipeline SHALL call a KnowledgeSearchPort boundary

Semantic candidate fetch SHALL go through a KnowledgeSearchPort abstraction rather than embedding new long-term SDK-specific hybrid logic into the tool orchestrator.

Scenario: Default dense adapter

  • WHEN search mode is dense-only (this change)
  • THEN the port implementation MAY delegate to the existing vector search facade
  • AND the retriever consumes port hits normalized into candidates with identity fields