feat(rag): chunk evidence identity, dedup, and search port

Preserve same-document multi-chunk evidence with evidenceKey identity,
per-document caps, retrieve-k/return-n split, and a dense KnowledgeSearchPort.
Archives Delivery 1 OpenSpec change as the foundation for hybrid retrieval.
This commit is contained in:
zhuyongxin
2026-07-27 18:26:15 +08:00
parent 99d4f6f216
commit ac1f831903
34 changed files with 2880 additions and 298 deletions
@@ -0,0 +1,82 @@
# rag-chunk-evidence-identity Specification
## Purpose
Define chunk-level evidence identity, deduplication, retrieve/return separation, and the thin search port boundary used by `lookup_knowledge` before hybrid retrieval is introduced.
## ADDED Requirements
### Requirement: Retrieval candidates SHALL carry chunk-level identity
The knowledge retrieval pipeline SHALL attach stable identity fields to each retrieved candidate and resulting evidence block: document id when available, chunk index when available, and an `evidenceKey` derived by the identity rules in design.
#### Scenario: Identity extracted from metadata
- **WHEN** a vector hit includes metadata `docId` (or `doc_id`) and `chunkIndex` (or `chunk_index`)
- **THEN** the candidate SHALL set `docId`, `chunkIndex`, and `evidenceKey` to `docId#chunk-{chunkIndex}`
#### Scenario: Fallback identity without chunk index
- **WHEN** chunk index is missing but vector id is present
- **THEN** the candidate SHALL use `evidenceKey` of the form `vector:{id}`
### Requirement: Evidence deduplication SHALL be chunk-scoped
Post-processing SHALL treat two candidates as duplicates only when they share the same `evidenceKey`. Different chunks of the same source SHALL remain separate evidence blocks subject to per-document caps.
#### Scenario: Same document different chunks are kept
- **WHEN** two candidates share the same source/docId but different chunk indexes
- **THEN** post-processing SHALL keep both as separate evidence blocks unless a per-document cap removes the lower-ranked one
#### Scenario: True duplicate keys merge without replacing higher-ranked content
- **WHEN** two candidates share the same `evidenceKey`
- **THEN** post-processing SHALL keep a single block
- **AND** SHALL preserve the higher-ranked content
- **AND** MAY merge hit reasons
### Requirement: Post-processing SHALL enforce max chunks per document
The system SHALL limit accepted evidence blocks per document id using configuration `rag.max-chunks-per-document` with default 2.
#### Scenario: Excess chunks from one document are dropped
- **WHEN** more than N ranked chunks belong to the same docId and N equals the configured max
- **THEN** only the top N by ranking score SHALL remain in evidence blocks
### Requirement: Retrieval width and return width SHALL be separate
The lookup pipeline SHALL use `rag.retrieve-k` for vector recall width and `rag.return-n` for maximum evidence blocks after post-processing. A legacy `rag.top-k` MAY act as fallback when the new keys are absent.
#### Scenario: retrieve-k widens recall without unbounded return
- **WHEN** `retrieve-k` is 20 and `return-n` is 5
- **THEN** the search port is asked for up to 20 candidates
- **AND** the assembled result contains at most 5 evidence blocks before Agent projection budgets
### Requirement: Agent projection SHALL preserve distinct chunk evidence
The RAG projector SHALL deduplicate Agent-facing evidence by chunk-scoped evidence identity. It SHALL NOT drop a second chunk solely because `source` matches a previous block.
#### Scenario: Same source different evidence keys both project
- **WHEN** two evidence blocks have different evidence keys (or chunk-scoped document ids) and the same source
- **AND** projection budgets still allow both
- **THEN** both SHALL appear in the projected evidence list
#### Scenario: Projected document_id is chunk-scoped
- **WHEN** an evidence block has evidenceKey `docA#chunk-2`
- **THEN** the projected `document_id` SHALL be that evidenceKey (or an equivalent chunk-scoped id)
- **AND** `source` MAY still be the document path or doc id
### Requirement: Lookup pipeline SHALL call a KnowledgeSearchPort boundary
Semantic candidate fetch SHALL go through a `KnowledgeSearchPort` abstraction rather than embedding new long-term SDK-specific hybrid logic into the tool orchestrator.
#### Scenario: Default dense adapter
- **WHEN** search mode is dense-only (this change)
- **THEN** the port implementation MAY delegate to the existing vector search facade
- **AND** the retriever consumes port hits normalized into candidates with identity fields
@@ -0,0 +1,24 @@
# rag-knowledge-retrieval Delta
## MODIFIED Requirements
### Requirement: Knowledge retrieval SHALL deduplicate evidence blocks
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent. Duplicates are defined by chunk-level `evidenceKey` identity, not by document source alone.
#### Scenario: Chunk-level deduplication keeps distinct chunks
- **WHEN** L1 produces multiple candidates with the same source but different chunk identities
- **THEN** the retrieval flow SHALL keep separate evidence blocks for those chunk identities subject to per-document caps
- **AND** SHALL NOT collapse them solely because source is equal
#### Scenario: Same evidenceKey collapses
- **WHEN** two candidates share the same evidenceKey
- **THEN** the retrieval flow SHALL keep a single evidence block for that identity
- **AND** the evidence block SHALL preserve hit reasons from both paths when available
#### Scenario: Postprocess count tracking
- **WHEN** evidence post-processing completes
- **THEN** the tool invocation details SHALL record candidate count and final evidence block count
@@ -0,0 +1,19 @@
# rag-log-projections Delta (RAG portion)
## MODIFIED Requirements
### Requirement: RAG projection SHALL expose bounded document evidence only
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result. Projected evidence identity SHALL be chunk-scoped when chunk identity is available, allowing multiple excerpts from one logical source document.
#### Scenario: Project only bounded RAG evidence
- **WHEN** the adapter receives a logical query and the knowledge backend returns a result
- **THEN** it invokes through ToolBoundary and exposes only the bounded `RagToolResult` evidence fields, excluding retrieval internals and full document bodies
#### Scenario: Multiple chunks from one source may project
- **WHEN** the backend returns multiple evidence blocks with the same source and different chunk-scoped identities
- **AND** projection budgets allow them
- **THEN** the projected evidence list SHALL include more than one item for that source
- **AND** each item SHALL have a distinct `document_id`