feat(rag): chunk evidence identity, dedup, and search port

Preserve same-document multi-chunk evidence with evidenceKey identity,
per-document caps, retrieve-k/return-n split, and a dense KnowledgeSearchPort.
Archives Delivery 1 OpenSpec change as the foundation for hybrid retrieval.
This commit is contained in:
zhuyongxin
2026-07-27 18:26:15 +08:00
parent 99d4f6f216
commit ac1f831903
34 changed files with 2880 additions and 298 deletions
@@ -0,0 +1,2 @@
ready: 2026-07-27
change: rag-chunk-evidence-identity-dedup
@@ -0,0 +1,4 @@
committed: 2026-07-27
change: rag-chunk-evidence-identity-dedup
scale: standard
authorized-apply: user-preauthorized-sm-flow
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-27
@@ -0,0 +1,122 @@
# Design: RAG chunk evidence identity and dedup
## Context
- Modular RAG pipeline already exists (`KnowledgeQueryTransformer` → retriever → post-processor → packer → assembler → projector).
- Delivery baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1.
- Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary.
- Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic.
## Goals / Non-Goals
**Goals**
- Preserve multiple relevant chunks from the same document in one lookup.
- Make evidence identity stable enough for future hybrid hits.
- Separate retrieve width from return width.
- Keep Agent tool name/input unchanged (`lookup_knowledge(query)`).
**Non-Goals**
- Hybrid / BM25 / sparse schema.
- Cross-encoder rerank.
- Deleting SDK mode.
- Session dedup tracker.
## Decisions
### D1. Evidence identity key
```text
evidenceKey =
if docId != null && chunkIndex != null:
docId + "#chunk-" + chunkIndex
else if vectorId != null:
"vector:" + vectorId
else:
"rank:" + originalRank
```
Rationale: works with current metadata (`docId`, `chunkIndex`) and degrades safely for older rows.
### D2. Dedup granularity
- Dedup key = `evidenceKey` only.
- True duplicates (same key) merge hitReasons; keep higher-ranked content (first after score sort).
- Do **not** merge different chunks of the same source into one content block.
### D3. Per-document cap
- After score sort, accept at most `rag.max-chunks-per-document` (default `2`) evidence blocks per `docId`.
- If `docId` missing, treat each evidenceKey as its own document bucket.
### D4. retrieve-k / return-n
```properties
rag.retrieve-k=20 # vector recall width
rag.return-n=5 # max evidence blocks after post-process (before projector budget)
rag.max-chunks-per-document=2
```
Compatibility:
- If only legacy `rag.top-k` is set, use it as fallback for both until removed.
- Prefer explicit retrieve-k/return-n when present.
### D5. Agent projection identity (behavior change)
- Projected `document_id` SHOULD be `evidenceKey` (chunk-scoped), not raw source.
- This is intentional so EvidenceGuard references remain 1:1 with returned excerpts.
- `source` remains human-readable document path/id and MAY repeat across chunks.
- Marked as **observable behavior change** for Agent consumers and report references.
### D6. KnowledgeSearchPort (thin)
```text
KnowledgeSearchPort.search(KnowledgeSearchRequest) -> List<KnowledgeSearchHit>
```
- Default adapter delegates to existing `VectorSearchService.searchSimilarDocuments`.
- Request carries query, topK, categoryFilter, mode placeholder (`DENSE` only in this change).
- Hit carries id, content, score fields, metadata map, and extracted identity fields when available.
- No hybrid implementation in this change.
## Module flow
```text
LookupKnowledge
-> transform (L0 hints)
-> KnowledgeSearchPort.search(retrieveK, filter)
-> map to RetrievedEvidenceCandidate (+ identity)
-> post-process:
score/sort
evidenceKey dedup
maxChunksPerDocument
return-n truncate
-> pack + assemble
-> RagResultProjector (dedupe by evidence document_id=evidenceKey)
```
## Interface impact
| Level | What |
|---|---|
| L2 | Internal DTO fields on candidate/EvidenceBlock |
| L3 | Agent-facing `document_id` becomes chunk-scoped evidence id |
Migration/compat:
- In-repo guards/tests updated to accept chunk-scoped ids.
- External human readers still see `source`/`title`.
## Risks / Trade-offs
| Risk | Mitigation |
|---|---|
| Larger Agent payload | return-n + maxChunksPerDocument + existing projector budgets |
| Missing chunkIndex in old data | vector id fallback keeps chunks distinct |
| document_id semantic shift | documented; projector tests updated |
## Open questions
None remaining for Delivery 1. Hybrid belongs to next change.
@@ -0,0 +1,56 @@
# Change: RAG chunk evidence identity and dedup
## Why
`lookup_knowledge` currently collapses same-document chunks at two layers:
1. `KnowledgeEvidencePostProcessor` dedupes by `source/title/breadcrumb`
2. `RagResultProjector` dedupes by `document_id`, which usually falls back to `source`
As a result, L1 can recall multiple useful chunks from one document, but Agent often sees only one. This blocks multi-path / hybrid retrieval benefits and hurts long runbook completeness.
This change is **Delivery 1** from `docs/milvus-hybrid-search-integration-checklist.md`. Hybrid schema/search is **out of scope** and will be a separate change after this is archived.
## What Changes
- Add chunk-level evidence identity: `docId`, `chunkIndex`, `evidenceKey`
- Extract identity in retrieval adapter from vector metadata
- Deduplicate by `evidenceKey` (chunk identity), not document source
- Cap chunks per document (`maxChunksPerDocument`, default 2)
- Split `retrieve-k` and `return-n` (stop overloading single `top-k`)
- Align Agent projection so same-source different chunks can both appear
- Introduce a thin `KnowledgeSearchPort` so later hybrid can swap implementation without rewriting the pipeline
- Keep L0 as hint-only; no BM25/sparse schema; no legacy SDK deletion in this change
## Non-goals
- Milvus sparse/BM25 schema or reindex
- Enabling hybrid search mode
- Removing Milvus SDK path
- Session-level RetrievedDocTracker restore
- Neighbor chunk context reconstruction
- Model reranker
## Capabilities
### New Capabilities
- `rag-chunk-evidence-identity`: chunk-level identity, dedup, retrieve/return split, search port boundary
### Modified Capabilities
- `rag-knowledge-retrieval`: replace document-level evidence dedup requirement with chunk-level identity
- `rag-log-projections` (RAG portion only): projection identity may be chunk-scoped `document_id`
## Impact
- Code: retrieval DTO/services, post-processor, lookup tool config, projector, tests
- Agent-visible: more evidence items possible for same logical document when multiple chunks are relevant
- Interface level: **L2/L3** — Agent `document_id` semantics become chunk-scoped evidence id (often `docId#chunk-N`); EvidenceGuard still validates against tool projection ids
- Docs baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1
## Risks
- Agent context grows if many chunks pass; mitigated by `return-n` and `maxChunksPerDocument`
- Existing tests assume source-level dedup; must update intentionally
- Old indexes without `chunkIndex` need stable fallback keys (`vector:{id}`)
@@ -0,0 +1,82 @@
# rag-chunk-evidence-identity Specification
## Purpose
Define chunk-level evidence identity, deduplication, retrieve/return separation, and the thin search port boundary used by `lookup_knowledge` before hybrid retrieval is introduced.
## ADDED Requirements
### Requirement: Retrieval candidates SHALL carry chunk-level identity
The knowledge retrieval pipeline SHALL attach stable identity fields to each retrieved candidate and resulting evidence block: document id when available, chunk index when available, and an `evidenceKey` derived by the identity rules in design.
#### Scenario: Identity extracted from metadata
- **WHEN** a vector hit includes metadata `docId` (or `doc_id`) and `chunkIndex` (or `chunk_index`)
- **THEN** the candidate SHALL set `docId`, `chunkIndex`, and `evidenceKey` to `docId#chunk-{chunkIndex}`
#### Scenario: Fallback identity without chunk index
- **WHEN** chunk index is missing but vector id is present
- **THEN** the candidate SHALL use `evidenceKey` of the form `vector:{id}`
### Requirement: Evidence deduplication SHALL be chunk-scoped
Post-processing SHALL treat two candidates as duplicates only when they share the same `evidenceKey`. Different chunks of the same source SHALL remain separate evidence blocks subject to per-document caps.
#### Scenario: Same document different chunks are kept
- **WHEN** two candidates share the same source/docId but different chunk indexes
- **THEN** post-processing SHALL keep both as separate evidence blocks unless a per-document cap removes the lower-ranked one
#### Scenario: True duplicate keys merge without replacing higher-ranked content
- **WHEN** two candidates share the same `evidenceKey`
- **THEN** post-processing SHALL keep a single block
- **AND** SHALL preserve the higher-ranked content
- **AND** MAY merge hit reasons
### Requirement: Post-processing SHALL enforce max chunks per document
The system SHALL limit accepted evidence blocks per document id using configuration `rag.max-chunks-per-document` with default 2.
#### Scenario: Excess chunks from one document are dropped
- **WHEN** more than N ranked chunks belong to the same docId and N equals the configured max
- **THEN** only the top N by ranking score SHALL remain in evidence blocks
### Requirement: Retrieval width and return width SHALL be separate
The lookup pipeline SHALL use `rag.retrieve-k` for vector recall width and `rag.return-n` for maximum evidence blocks after post-processing. A legacy `rag.top-k` MAY act as fallback when the new keys are absent.
#### Scenario: retrieve-k widens recall without unbounded return
- **WHEN** `retrieve-k` is 20 and `return-n` is 5
- **THEN** the search port is asked for up to 20 candidates
- **AND** the assembled result contains at most 5 evidence blocks before Agent projection budgets
### Requirement: Agent projection SHALL preserve distinct chunk evidence
The RAG projector SHALL deduplicate Agent-facing evidence by chunk-scoped evidence identity. It SHALL NOT drop a second chunk solely because `source` matches a previous block.
#### Scenario: Same source different evidence keys both project
- **WHEN** two evidence blocks have different evidence keys (or chunk-scoped document ids) and the same source
- **AND** projection budgets still allow both
- **THEN** both SHALL appear in the projected evidence list
#### Scenario: Projected document_id is chunk-scoped
- **WHEN** an evidence block has evidenceKey `docA#chunk-2`
- **THEN** the projected `document_id` SHALL be that evidenceKey (or an equivalent chunk-scoped id)
- **AND** `source` MAY still be the document path or doc id
### Requirement: Lookup pipeline SHALL call a KnowledgeSearchPort boundary
Semantic candidate fetch SHALL go through a `KnowledgeSearchPort` abstraction rather than embedding new long-term SDK-specific hybrid logic into the tool orchestrator.
#### Scenario: Default dense adapter
- **WHEN** search mode is dense-only (this change)
- **THEN** the port implementation MAY delegate to the existing vector search facade
- **AND** the retriever consumes port hits normalized into candidates with identity fields
@@ -0,0 +1,24 @@
# rag-knowledge-retrieval Delta
## MODIFIED Requirements
### Requirement: Knowledge retrieval SHALL deduplicate evidence blocks
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent. Duplicates are defined by chunk-level `evidenceKey` identity, not by document source alone.
#### Scenario: Chunk-level deduplication keeps distinct chunks
- **WHEN** L1 produces multiple candidates with the same source but different chunk identities
- **THEN** the retrieval flow SHALL keep separate evidence blocks for those chunk identities subject to per-document caps
- **AND** SHALL NOT collapse them solely because source is equal
#### Scenario: Same evidenceKey collapses
- **WHEN** two candidates share the same evidenceKey
- **THEN** the retrieval flow SHALL keep a single evidence block for that identity
- **AND** the evidence block SHALL preserve hit reasons from both paths when available
#### Scenario: Postprocess count tracking
- **WHEN** evidence post-processing completes
- **THEN** the tool invocation details SHALL record candidate count and final evidence block count
@@ -0,0 +1,19 @@
# rag-log-projections Delta (RAG portion)
## MODIFIED Requirements
### Requirement: RAG projection SHALL expose bounded document evidence only
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result. Projected evidence identity SHALL be chunk-scoped when chunk identity is available, allowing multiple excerpts from one logical source document.
#### Scenario: Project only bounded RAG evidence
- **WHEN** the adapter receives a logical query and the knowledge backend returns a result
- **THEN** it invokes through ToolBoundary and exposes only the bounded `RagToolResult` evidence fields, excluding retrieval internals and full document bodies
#### Scenario: Multiple chunks from one source may project
- **WHEN** the backend returns multiple evidence blocks with the same source and different chunk-scoped identities
- **AND** projection budgets allow them
- **THEN** the projected evidence list SHALL include more than one item for that source
- **AND** each item SHALL have a distinct `document_id`
@@ -0,0 +1,43 @@
# Tasks: rag-chunk-evidence-identity-dedup
## 1. Identity model
- [x] 1.1 Add `docId`, `chunkIndex`, `evidenceKey` to `RetrievedEvidenceCandidate` and `EvidenceBlock`
- [x] 1.2 Implement shared identity helper (`docId#chunk-N` / `vector:{id}` / `rank:{n}`)
- [x] 1.3 Extract identity in `KnowledgeDocumentRetriever` (or search-port mapper) from metadata
## 2. Search port boundary
- [x] 2.1 Add `KnowledgeSearchPort`, `KnowledgeSearchRequest`, `KnowledgeSearchHit`
- [x] 2.2 Implement dense adapter delegating to `VectorSearchService`
- [x] 2.3 Wire retriever to port; keep tool orchestration free of SDK details
## 3. Post-process dedup and caps
- [x] 3.1 Change dedup key to `evidenceKey`
- [x] 3.2 Add `rag.max-chunks-per-document` (default 2)
- [x] 3.3 Apply `rag.return-n` truncation after ranking/dedup/cap
- [x] 3.4 Keep score-threshold / relevance behavior unchanged except ordering inputs
## 4. Lookup config
- [x] 4.1 Add `rag.retrieve-k` and `rag.return-n` with legacy `rag.top-k` fallback
- [x] 4.2 Use retrieve-k for search port calls in `LookupKnowledgeTool`
## 5. Projector alignment
- [x] 5.1 Prefer evidenceKey / chunk-scoped id as projected `document_id`
- [x] 5.2 Stop dropping second evidence solely because `source` matches
- [x] 5.3 Keep existing budget/truncation behavior
## 6. Tests
- [x] 6.1 Update `LookupKnowledgeToolTest` source-dedup case to chunk-preserving behavior
- [x] 6.2 Add post-processor tests: multi-chunk keep, true-dup merge, maxChunksPerDocument
- [x] 6.3 Update/add `RagResultProjectorTest` for same-source multi-chunk projection
- [x] 6.4 Add search-port adapter smoke test if practical
## 7. Verification
- [x] 7.1 Run targeted unit tests for lookup / post-processor / projector
- [x] 7.2 Mark tasks complete and note any residual risks for Delivery 2
@@ -0,0 +1,80 @@
# rag-chunk-evidence-identity Specification
## Purpose
TBD - created by archiving change rag-chunk-evidence-identity-dedup. Update Purpose after archive.
## Requirements
### Requirement: Retrieval candidates SHALL carry chunk-level identity
The knowledge retrieval pipeline SHALL attach stable identity fields to each retrieved candidate and resulting evidence block: document id when available, chunk index when available, and an `evidenceKey` derived by the identity rules in design.
#### Scenario: Identity extracted from metadata
- **WHEN** a vector hit includes metadata `docId` (or `doc_id`) and `chunkIndex` (or `chunk_index`)
- **THEN** the candidate SHALL set `docId`, `chunkIndex`, and `evidenceKey` to `docId#chunk-{chunkIndex}`
#### Scenario: Fallback identity without chunk index
- **WHEN** chunk index is missing but vector id is present
- **THEN** the candidate SHALL use `evidenceKey` of the form `vector:{id}`
### Requirement: Evidence deduplication SHALL be chunk-scoped
Post-processing SHALL treat two candidates as duplicates only when they share the same `evidenceKey`. Different chunks of the same source SHALL remain separate evidence blocks subject to per-document caps.
#### Scenario: Same document different chunks are kept
- **WHEN** two candidates share the same source/docId but different chunk indexes
- **THEN** post-processing SHALL keep both as separate evidence blocks unless a per-document cap removes the lower-ranked one
#### Scenario: True duplicate keys merge without replacing higher-ranked content
- **WHEN** two candidates share the same `evidenceKey`
- **THEN** post-processing SHALL keep a single block
- **AND** SHALL preserve the higher-ranked content
- **AND** MAY merge hit reasons
### Requirement: Post-processing SHALL enforce max chunks per document
The system SHALL limit accepted evidence blocks per document id using configuration `rag.max-chunks-per-document` with default 2.
#### Scenario: Excess chunks from one document are dropped
- **WHEN** more than N ranked chunks belong to the same docId and N equals the configured max
- **THEN** only the top N by ranking score SHALL remain in evidence blocks
### Requirement: Retrieval width and return width SHALL be separate
The lookup pipeline SHALL use `rag.retrieve-k` for vector recall width and `rag.return-n` for maximum evidence blocks after post-processing. A legacy `rag.top-k` MAY act as fallback when the new keys are absent.
#### Scenario: retrieve-k widens recall without unbounded return
- **WHEN** `retrieve-k` is 20 and `return-n` is 5
- **THEN** the search port is asked for up to 20 candidates
- **AND** the assembled result contains at most 5 evidence blocks before Agent projection budgets
### Requirement: Agent projection SHALL preserve distinct chunk evidence
The RAG projector SHALL deduplicate Agent-facing evidence by chunk-scoped evidence identity. It SHALL NOT drop a second chunk solely because `source` matches a previous block.
#### Scenario: Same source different evidence keys both project
- **WHEN** two evidence blocks have different evidence keys (or chunk-scoped document ids) and the same source
- **AND** projection budgets still allow both
- **THEN** both SHALL appear in the projected evidence list
#### Scenario: Projected document_id is chunk-scoped
- **WHEN** an evidence block has evidenceKey `docA#chunk-2`
- **THEN** the projected `document_id` SHALL be that evidenceKey (or an equivalent chunk-scoped id)
- **AND** `source` MAY still be the document path or doc id
### Requirement: Lookup pipeline SHALL call a KnowledgeSearchPort boundary
Semantic candidate fetch SHALL go through a `KnowledgeSearchPort` abstraction rather than embedding new long-term SDK-specific hybrid logic into the tool orchestrator.
#### Scenario: Default dense adapter
- **WHEN** search mode is dense-only (this change)
- **THEN** the port implementation MAY delegate to the existing vector search facade
- **AND** the retriever consumes port hits normalized into candidates with identity fields
+15 -5
View File
@@ -54,14 +54,23 @@ The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structu
- **AND** the result SHALL NOT rely on legacy `primary` or `supplement` fields for L0/L1 meaning
### Requirement: Knowledge retrieval SHALL deduplicate evidence blocks
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent.
#### Scenario: Duplicate source deduplication
- **WHEN** L0 and L1 produce evidence with the same source identity
- **THEN** the retrieval flow SHALL keep a single evidence block for that source
- **AND** the evidence block SHALL preserve hit reasons from both retrieval paths when available
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent. Duplicates are defined by chunk-level `evidenceKey` identity, not by document source alone.
#### Scenario: Chunk-level deduplication keeps distinct chunks
- **WHEN** L1 produces multiple candidates with the same source but different chunk identities
- **THEN** the retrieval flow SHALL keep separate evidence blocks for those chunk identities subject to per-document caps
- **AND** SHALL NOT collapse them solely because source is equal
#### Scenario: Same evidenceKey collapses
- **WHEN** two candidates share the same evidenceKey
- **THEN** the retrieval flow SHALL keep a single evidence block for that identity
- **AND** the evidence block SHALL preserve hit reasons from both paths when available
#### Scenario: Postprocess count tracking
- **WHEN** evidence post-processing completes
- **THEN** the tool invocation details SHALL record candidate count and final evidence block count
@@ -248,3 +257,4 @@ The `lookup_knowledge` tool SHALL persist modular RAG pipeline details in `tool_
- **WHEN** `lookup_knowledge` returns no usable evidence
- **THEN** `retrieval_details` SHALL still include retrieval trace information for attempted retrieval paths
- **AND** it SHALL not include full evidence content as duplicated trace data
+9 -3
View File
@@ -3,18 +3,23 @@
## Purpose
Define bounded RAG and Mock query-log projections that execute through the stage 3A ToolBoundary and expose only the frozen ACI contracts to the Agent.
## Requirements
### Requirement: RAG projection SHALL expose bounded document evidence only
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result.
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result. Projected evidence identity SHALL be chunk-scoped when chunk identity is available, allowing multiple excerpts from one logical source document.
#### Scenario: Project only bounded RAG evidence
- **WHEN** the adapter receives a logical query and the knowledge backend returns a result
- **THEN** it invokes through ToolBoundary and exposes only the bounded `RagToolResult` evidence fields, excluding retrieval internals and full document bodies
#### Scenario: Multiple chunks from one source may project
- **WHEN** the backend returns multiple evidence blocks with the same source and different chunk-scoped identities
- **AND** projection budgets allow them
- **THEN** the projected evidence list SHALL include more than one item for that source
- **AND** each item SHALL have a distinct `document_id`
### Requirement: Query-log projection SHALL preserve logical scope and Mock provenance
The query-log adapter SHALL accept only logical topic, query and optional lookback minutes, execute the existing Mock source through ToolBoundary, and project `source_kind=MOCK`, complete scope, match count, returned count, bounded patterns, bounded timeline events and truncation.
@@ -41,3 +46,4 @@ Both adapters SHALL pass the framework `tool_call_id` and RunContext to the exis
- **WHEN** either the RAG or query-log adapter executes
- **THEN** it passes the exact framework `tool_call_id` and RunContext through ToolBoundary without creating a second identifier, parallel store, raw response path, or legacy audit side effect