Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
123 lines
3.8 KiB
Markdown
123 lines
3.8 KiB
Markdown
# Design: RAG chunk evidence identity and dedup
|
|
|
|
## Context
|
|
|
|
- Modular RAG pipeline already exists (`KnowledgeQueryTransformer` → retriever → post-processor → packer → assembler → projector).
|
|
- Delivery baseline: `docs/Milvus-Hybrid接入清单.md` §1.1 Delivery 1.
|
|
- Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary.
|
|
- Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals**
|
|
|
|
- Preserve multiple relevant chunks from the same document in one lookup.
|
|
- Make evidence identity stable enough for future hybrid hits.
|
|
- Separate retrieve width from return width.
|
|
- Keep Agent tool name/input unchanged (`lookup_knowledge(query)`).
|
|
|
|
**Non-Goals**
|
|
|
|
- Hybrid / BM25 / sparse schema.
|
|
- Cross-encoder rerank.
|
|
- Deleting SDK mode.
|
|
- Session dedup tracker.
|
|
|
|
## Decisions
|
|
|
|
### D1. Evidence identity key
|
|
|
|
```text
|
|
evidenceKey =
|
|
if docId != null && chunkIndex != null:
|
|
docId + "#chunk-" + chunkIndex
|
|
else if vectorId != null:
|
|
"vector:" + vectorId
|
|
else:
|
|
"rank:" + originalRank
|
|
```
|
|
|
|
Rationale: works with current metadata (`docId`, `chunkIndex`) and degrades safely for older rows.
|
|
|
|
### D2. Dedup granularity
|
|
|
|
- Dedup key = `evidenceKey` only.
|
|
- True duplicates (same key) merge hitReasons; keep higher-ranked content (first after score sort).
|
|
- Do **not** merge different chunks of the same source into one content block.
|
|
|
|
### D3. Per-document cap
|
|
|
|
- After score sort, accept at most `rag.max-chunks-per-document` (default `2`) evidence blocks per `docId`.
|
|
- If `docId` missing, treat each evidenceKey as its own document bucket.
|
|
|
|
### D4. retrieve-k / return-n
|
|
|
|
```properties
|
|
rag.retrieve-k=20 # vector recall width
|
|
rag.return-n=5 # max evidence blocks after post-process (before projector budget)
|
|
rag.max-chunks-per-document=2
|
|
```
|
|
|
|
Compatibility:
|
|
|
|
- If only legacy `rag.top-k` is set, use it as fallback for both until removed.
|
|
- Prefer explicit retrieve-k/return-n when present.
|
|
|
|
### D5. Agent projection identity (behavior change)
|
|
|
|
- Projected `document_id` SHOULD be `evidenceKey` (chunk-scoped), not raw source.
|
|
- This is intentional so EvidenceGuard references remain 1:1 with returned excerpts.
|
|
- `source` remains human-readable document path/id and MAY repeat across chunks.
|
|
- Marked as **observable behavior change** for Agent consumers and report references.
|
|
|
|
### D6. KnowledgeSearchPort (thin)
|
|
|
|
```text
|
|
KnowledgeSearchPort.search(KnowledgeSearchRequest) -> List<KnowledgeSearchHit>
|
|
```
|
|
|
|
- Default adapter delegates to existing `VectorSearchService.searchSimilarDocuments`.
|
|
- Request carries query, topK, categoryFilter, mode placeholder (`DENSE` only in this change).
|
|
- Hit carries id, content, score fields, metadata map, and extracted identity fields when available.
|
|
- No hybrid implementation in this change.
|
|
|
|
## Module flow
|
|
|
|
```text
|
|
LookupKnowledge
|
|
-> transform (L0 hints)
|
|
-> KnowledgeSearchPort.search(retrieveK, filter)
|
|
-> map to RetrievedEvidenceCandidate (+ identity)
|
|
-> post-process:
|
|
score/sort
|
|
evidenceKey dedup
|
|
maxChunksPerDocument
|
|
return-n truncate
|
|
-> pack + assemble
|
|
-> RagResultProjector (dedupe by evidence document_id=evidenceKey)
|
|
```
|
|
|
|
## Interface impact
|
|
|
|
| Level | What |
|
|
|---|---|
|
|
| L2 | Internal DTO fields on candidate/EvidenceBlock |
|
|
| L3 | Agent-facing `document_id` becomes chunk-scoped evidence id |
|
|
|
|
Migration/compat:
|
|
|
|
- In-repo guards/tests updated to accept chunk-scoped ids.
|
|
- External human readers still see `source`/`title`.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
| Risk | Mitigation |
|
|
|---|---|
|
|
| Larger Agent payload | return-n + maxChunksPerDocument + existing projector budgets |
|
|
| Missing chunkIndex in old data | vector id fallback keeps chunks distinct |
|
|
| document_id semantic shift | documented; projector tests updated |
|
|
|
|
## Open questions
|
|
|
|
None remaining for Delivery 1. Hybrid belongs to next change.
|