feat(rag): chunk evidence identity, dedup, and search port

Preserve same-document multi-chunk evidence with evidenceKey identity,
per-document caps, retrieve-k/return-n split, and a dense KnowledgeSearchPort.
Archives Delivery 1 OpenSpec change as the foundation for hybrid retrieval.
This commit is contained in:
zhuyongxin
2026-07-27 18:26:15 +08:00
parent 99d4f6f216
commit ac1f831903
34 changed files with 2880 additions and 298 deletions
@@ -0,0 +1,80 @@
# rag-chunk-evidence-identity Specification
## Purpose
TBD - created by archiving change rag-chunk-evidence-identity-dedup. Update Purpose after archive.
## Requirements
### Requirement: Retrieval candidates SHALL carry chunk-level identity
The knowledge retrieval pipeline SHALL attach stable identity fields to each retrieved candidate and resulting evidence block: document id when available, chunk index when available, and an `evidenceKey` derived by the identity rules in design.
#### Scenario: Identity extracted from metadata
- **WHEN** a vector hit includes metadata `docId` (or `doc_id`) and `chunkIndex` (or `chunk_index`)
- **THEN** the candidate SHALL set `docId`, `chunkIndex`, and `evidenceKey` to `docId#chunk-{chunkIndex}`
#### Scenario: Fallback identity without chunk index
- **WHEN** chunk index is missing but vector id is present
- **THEN** the candidate SHALL use `evidenceKey` of the form `vector:{id}`
### Requirement: Evidence deduplication SHALL be chunk-scoped
Post-processing SHALL treat two candidates as duplicates only when they share the same `evidenceKey`. Different chunks of the same source SHALL remain separate evidence blocks subject to per-document caps.
#### Scenario: Same document different chunks are kept
- **WHEN** two candidates share the same source/docId but different chunk indexes
- **THEN** post-processing SHALL keep both as separate evidence blocks unless a per-document cap removes the lower-ranked one
#### Scenario: True duplicate keys merge without replacing higher-ranked content
- **WHEN** two candidates share the same `evidenceKey`
- **THEN** post-processing SHALL keep a single block
- **AND** SHALL preserve the higher-ranked content
- **AND** MAY merge hit reasons
### Requirement: Post-processing SHALL enforce max chunks per document
The system SHALL limit accepted evidence blocks per document id using configuration `rag.max-chunks-per-document` with default 2.
#### Scenario: Excess chunks from one document are dropped
- **WHEN** more than N ranked chunks belong to the same docId and N equals the configured max
- **THEN** only the top N by ranking score SHALL remain in evidence blocks
### Requirement: Retrieval width and return width SHALL be separate
The lookup pipeline SHALL use `rag.retrieve-k` for vector recall width and `rag.return-n` for maximum evidence blocks after post-processing. A legacy `rag.top-k` MAY act as fallback when the new keys are absent.
#### Scenario: retrieve-k widens recall without unbounded return
- **WHEN** `retrieve-k` is 20 and `return-n` is 5
- **THEN** the search port is asked for up to 20 candidates
- **AND** the assembled result contains at most 5 evidence blocks before Agent projection budgets
### Requirement: Agent projection SHALL preserve distinct chunk evidence
The RAG projector SHALL deduplicate Agent-facing evidence by chunk-scoped evidence identity. It SHALL NOT drop a second chunk solely because `source` matches a previous block.
#### Scenario: Same source different evidence keys both project
- **WHEN** two evidence blocks have different evidence keys (or chunk-scoped document ids) and the same source
- **AND** projection budgets still allow both
- **THEN** both SHALL appear in the projected evidence list
#### Scenario: Projected document_id is chunk-scoped
- **WHEN** an evidence block has evidenceKey `docA#chunk-2`
- **THEN** the projected `document_id` SHALL be that evidenceKey (or an equivalent chunk-scoped id)
- **AND** `source` MAY still be the document path or doc id
### Requirement: Lookup pipeline SHALL call a KnowledgeSearchPort boundary
Semantic candidate fetch SHALL go through a `KnowledgeSearchPort` abstraction rather than embedding new long-term SDK-specific hybrid logic into the tool orchestrator.
#### Scenario: Default dense adapter
- **WHEN** search mode is dense-only (this change)
- **THEN** the port implementation MAY delegate to the existing vector search facade
- **AND** the retriever consumes port hits normalized into candidates with identity fields
+15 -5
View File
@@ -54,14 +54,23 @@ The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structu
- **AND** the result SHALL NOT rely on legacy `primary` or `supplement` fields for L0/L1 meaning
### Requirement: Knowledge retrieval SHALL deduplicate evidence blocks
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent.
#### Scenario: Duplicate source deduplication
- **WHEN** L0 and L1 produce evidence with the same source identity
- **THEN** the retrieval flow SHALL keep a single evidence block for that source
- **AND** the evidence block SHALL preserve hit reasons from both retrieval paths when available
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent. Duplicates are defined by chunk-level `evidenceKey` identity, not by document source alone.
#### Scenario: Chunk-level deduplication keeps distinct chunks
- **WHEN** L1 produces multiple candidates with the same source but different chunk identities
- **THEN** the retrieval flow SHALL keep separate evidence blocks for those chunk identities subject to per-document caps
- **AND** SHALL NOT collapse them solely because source is equal
#### Scenario: Same evidenceKey collapses
- **WHEN** two candidates share the same evidenceKey
- **THEN** the retrieval flow SHALL keep a single evidence block for that identity
- **AND** the evidence block SHALL preserve hit reasons from both paths when available
#### Scenario: Postprocess count tracking
- **WHEN** evidence post-processing completes
- **THEN** the tool invocation details SHALL record candidate count and final evidence block count
@@ -248,3 +257,4 @@ The `lookup_knowledge` tool SHALL persist modular RAG pipeline details in `tool_
- **WHEN** `lookup_knowledge` returns no usable evidence
- **THEN** `retrieval_details` SHALL still include retrieval trace information for attempted retrieval paths
- **AND** it SHALL not include full evidence content as duplicated trace data
+9 -3
View File
@@ -3,18 +3,23 @@
## Purpose
Define bounded RAG and Mock query-log projections that execute through the stage 3A ToolBoundary and expose only the frozen ACI contracts to the Agent.
## Requirements
### Requirement: RAG projection SHALL expose bounded document evidence only
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result.
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result. Projected evidence identity SHALL be chunk-scoped when chunk identity is available, allowing multiple excerpts from one logical source document.
#### Scenario: Project only bounded RAG evidence
- **WHEN** the adapter receives a logical query and the knowledge backend returns a result
- **THEN** it invokes through ToolBoundary and exposes only the bounded `RagToolResult` evidence fields, excluding retrieval internals and full document bodies
#### Scenario: Multiple chunks from one source may project
- **WHEN** the backend returns multiple evidence blocks with the same source and different chunk-scoped identities
- **AND** projection budgets allow them
- **THEN** the projected evidence list SHALL include more than one item for that source
- **AND** each item SHALL have a distinct `document_id`
### Requirement: Query-log projection SHALL preserve logical scope and Mock provenance
The query-log adapter SHALL accept only logical topic, query and optional lookback minutes, execute the existing Mock source through ToolBoundary, and project `source_kind=MOCK`, complete scope, match count, returned count, bounded patterns, bounded timeline events and truncation.
@@ -41,3 +46,4 @@ Both adapters SHALL pass the framework `tool_call_id` and RunContext to the exis
- **WHEN** either the RAG or query-log adapter executes
- **THEN** it passes the exact framework `tool_call_id` and RunContext through ToolBoundary without creating a second identifier, parallel store, raw response path, or legacy audit side effect