Files
SuperBizAgent-java/openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup/design.md
T
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

123 lines
3.8 KiB
Markdown

# Design: RAG chunk evidence identity and dedup
## Context
- Modular RAG pipeline already exists (`KnowledgeQueryTransformer` → retriever → post-processor → packer → assembler → projector).
- Delivery baseline: `docs/Milvus-Hybrid接入清单.md` §1.1 Delivery 1.
- Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary.
- Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic.
## Goals / Non-Goals
**Goals**
- Preserve multiple relevant chunks from the same document in one lookup.
- Make evidence identity stable enough for future hybrid hits.
- Separate retrieve width from return width.
- Keep Agent tool name/input unchanged (`lookup_knowledge(query)`).
**Non-Goals**
- Hybrid / BM25 / sparse schema.
- Cross-encoder rerank.
- Deleting SDK mode.
- Session dedup tracker.
## Decisions
### D1. Evidence identity key
```text
evidenceKey =
if docId != null && chunkIndex != null:
docId + "#chunk-" + chunkIndex
else if vectorId != null:
"vector:" + vectorId
else:
"rank:" + originalRank
```
Rationale: works with current metadata (`docId`, `chunkIndex`) and degrades safely for older rows.
### D2. Dedup granularity
- Dedup key = `evidenceKey` only.
- True duplicates (same key) merge hitReasons; keep higher-ranked content (first after score sort).
- Do **not** merge different chunks of the same source into one content block.
### D3. Per-document cap
- After score sort, accept at most `rag.max-chunks-per-document` (default `2`) evidence blocks per `docId`.
- If `docId` missing, treat each evidenceKey as its own document bucket.
### D4. retrieve-k / return-n
```properties
rag.retrieve-k=20 # vector recall width
rag.return-n=5 # max evidence blocks after post-process (before projector budget)
rag.max-chunks-per-document=2
```
Compatibility:
- If only legacy `rag.top-k` is set, use it as fallback for both until removed.
- Prefer explicit retrieve-k/return-n when present.
### D5. Agent projection identity (behavior change)
- Projected `document_id` SHOULD be `evidenceKey` (chunk-scoped), not raw source.
- This is intentional so EvidenceGuard references remain 1:1 with returned excerpts.
- `source` remains human-readable document path/id and MAY repeat across chunks.
- Marked as **observable behavior change** for Agent consumers and report references.
### D6. KnowledgeSearchPort (thin)
```text
KnowledgeSearchPort.search(KnowledgeSearchRequest) -> List<KnowledgeSearchHit>
```
- Default adapter delegates to existing `VectorSearchService.searchSimilarDocuments`.
- Request carries query, topK, categoryFilter, mode placeholder (`DENSE` only in this change).
- Hit carries id, content, score fields, metadata map, and extracted identity fields when available.
- No hybrid implementation in this change.
## Module flow
```text
LookupKnowledge
-> transform (L0 hints)
-> KnowledgeSearchPort.search(retrieveK, filter)
-> map to RetrievedEvidenceCandidate (+ identity)
-> post-process:
score/sort
evidenceKey dedup
maxChunksPerDocument
return-n truncate
-> pack + assemble
-> RagResultProjector (dedupe by evidence document_id=evidenceKey)
```
## Interface impact
| Level | What |
|---|---|
| L2 | Internal DTO fields on candidate/EvidenceBlock |
| L3 | Agent-facing `document_id` becomes chunk-scoped evidence id |
Migration/compat:
- In-repo guards/tests updated to accept chunk-scoped ids.
- External human readers still see `source`/`title`.
## Risks / Trade-offs
| Risk | Mitigation |
|---|---|
| Larger Agent payload | return-n + maxChunksPerDocument + existing projector budgets |
| Missing chunkIndex in old data | vector id fallback keeps chunks distinct |
| document_id semantic shift | documented; projector tests updated |
## Open questions
None remaining for Delivery 1. Hybrid belongs to next change.