# rag-chunk-evidence-identity Specification ## Purpose TBD - created by archiving change rag-chunk-evidence-identity-dedup. Update Purpose after archive. ## Requirements ### Requirement: Retrieval candidates SHALL carry chunk-level identity The knowledge retrieval pipeline SHALL attach stable identity fields to each retrieved candidate and resulting evidence block: document id when available, chunk index when available, and an `evidenceKey` derived by the identity rules in design. #### Scenario: Identity extracted from metadata - **WHEN** a vector hit includes metadata `docId` (or `doc_id`) and `chunkIndex` (or `chunk_index`) - **THEN** the candidate SHALL set `docId`, `chunkIndex`, and `evidenceKey` to `docId#chunk-{chunkIndex}` #### Scenario: Fallback identity without chunk index - **WHEN** chunk index is missing but vector id is present - **THEN** the candidate SHALL use `evidenceKey` of the form `vector:{id}` ### Requirement: Evidence deduplication SHALL be chunk-scoped Post-processing SHALL treat two candidates as duplicates only when they share the same `evidenceKey`. Different chunks of the same source SHALL remain separate evidence blocks subject to per-document caps. #### Scenario: Same document different chunks are kept - **WHEN** two candidates share the same source/docId but different chunk indexes - **THEN** post-processing SHALL keep both as separate evidence blocks unless a per-document cap removes the lower-ranked one #### Scenario: True duplicate keys merge without replacing higher-ranked content - **WHEN** two candidates share the same `evidenceKey` - **THEN** post-processing SHALL keep a single block - **AND** SHALL preserve the higher-ranked content - **AND** MAY merge hit reasons ### Requirement: Post-processing SHALL enforce max chunks per document The system SHALL limit accepted evidence blocks per document id using configuration `rag.max-chunks-per-document` with default 2. #### Scenario: Excess chunks from one document are dropped - **WHEN** more than N ranked chunks belong to the same docId and N equals the configured max - **THEN** only the top N by ranking score SHALL remain in evidence blocks ### Requirement: Retrieval width and return width SHALL be separate The lookup pipeline SHALL use `rag.retrieve-k` for vector recall width and `rag.return-n` for maximum evidence blocks after post-processing. A legacy `rag.top-k` MAY act as fallback when the new keys are absent. #### Scenario: retrieve-k widens recall without unbounded return - **WHEN** `retrieve-k` is 20 and `return-n` is 5 - **THEN** the search port is asked for up to 20 candidates - **AND** the assembled result contains at most 5 evidence blocks before Agent projection budgets ### Requirement: Agent projection SHALL preserve distinct chunk evidence The RAG projector SHALL deduplicate Agent-facing evidence by chunk-scoped evidence identity. It SHALL NOT drop a second chunk solely because `source` matches a previous block. #### Scenario: Same source different evidence keys both project - **WHEN** two evidence blocks have different evidence keys (or chunk-scoped document ids) and the same source - **AND** projection budgets still allow both - **THEN** both SHALL appear in the projected evidence list #### Scenario: Projected document_id is chunk-scoped - **WHEN** an evidence block has evidenceKey `docA#chunk-2` - **THEN** the projected `document_id` SHALL be that evidenceKey (or an equivalent chunk-scoped id) - **AND** `source` MAY still be the document path or doc id ### Requirement: Lookup pipeline SHALL call a KnowledgeSearchPort boundary Semantic candidate fetch SHALL go through a `KnowledgeSearchPort` abstraction rather than embedding new long-term SDK-specific hybrid logic into the tool orchestrator. #### Scenario: Default dense adapter - **WHEN** search mode is dense-only (this change) - **THEN** the port implementation MAY delegate to the existing vector search facade - **AND** the retriever consumes port hits normalized into candidates with identity fields