feat(rag): dense+BM25 hybrid on MilvusClientV2, drop SDK path

Replace legacy MilvusServiceClient knowledge search/write with a single
MilvusClientV2 hybrid store (BM25 function + dense ANN + RRFRanker).
Use collection biz_hybrid and require knowledge reindex.
This commit is contained in:
zhuyongxin
2026-07-27 18:49:20 +08:00
parent 376ad0c241
commit f035538531
22 changed files with 813 additions and 936 deletions
@@ -0,0 +1,2 @@
committed: 2026-07-27
authorized-apply: user-preauthorized
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-27
@@ -0,0 +1,57 @@
# Design: rag-bm25-hybrid-drop-sdk
## Decision: single backend
```text
KnowledgeSearchPort
-> MilvusHybridKnowledgeStore (MilvusClientV2 only)
Write
-> VectorIndexService -> MilvusHybridKnowledgeStore
```
No `retrieval.vector-store.mode=sdk|spring|auto` for knowledge lookup.
## Schema (`milvus.collection`, default `biz_hybrid`)
| field | type | notes |
|---|---|---|
| id | VarChar PK | chunk id |
| content | VarChar | original chunk body returned to agent |
| search_text | VarChar + analyzer | BM25 input (title/path/content) |
| sparse_vector | SparseFloatVector | BM25 function output |
| vector | FloatVector | dense embedding |
| metadata | JSON | docId, chunkIndex, category, kb_scope, ... |
Function: `BM25(search_text -> sparse_vector)`
Indexes:
- `vector`: IVF_FLAT / L2 (or COSINE if configured later)
- `sparse_vector`: SPARSE_INVERTED_INDEX / BM25
## Hybrid query
```text
AnnSearchReq(vector, FloatVec(queryEmbedding), topK, filter?)
AnnSearchReq(sparse_vector, EmbeddedText(query), topK, filter?)
HybridSearchReq + RRFRanker(k)
```
Dense-compatible score for thresholds: if entity lacks dense distance, map fused score conservatively or re-use dense-only probe. Preferred: keep dense distance when available from a parallel dense search metadata; for hybrid hits use inverse-rank placeholder only in metadata and set score from optional dense sub-hit map.
Implementation approach for threshold stability:
1. Run hybridSearch for ordering/identity
2. Build map evidenceKey -> dense L2 from a concurrent dense search (same filter/topK)
3. Attach dense score onto fused hits when present; else maxL2 (low similarity) so weak BM25-only hits don't fake PRECISE
## Migration
- New collection name avoids mutating legacy `biz`
- Operators re-run knowledge init / document reindex
- Document in acceptance
## Drop SDK search
- Delete SDK search methods usage from knowledge path
- `MilvusServiceClient` bean may remain temporarily only if other non-search utilities need it; prefer migrate write/delete fully to V2 and stop creating V1 client if unused
@@ -0,0 +1,34 @@
# Change: Real BM25 hybrid search and drop legacy SDK path
## Why
Delivery 2 shipped application-layer multi-path + sparse-lite lexical ranking. Product requirement is **true dense + BM25 hybrid retrieval**. Legacy `MilvusServiceClient` search path will be abandoned in this change; keep a single knowledge vector backend based on Milvus Java SDK v2 (`MilvusClientV2`) with BM25 function + `hybridSearch` + RRF.
## What Changes
- New hybrid collection schema: dense float vector + analyzed text + sparse BM25 output
- Write path inserts text for BM25 auto-sparsification and dense embedding
- Search path: single backend
- dense mode: dense ANN only
- hybrid mode: dense ANN + BM25 sparse ANN fused by RRFRanker
- Remove/disable legacy SDK `search` / `searchSimilarDocumentsWithSdk` knowledge path and mode routing (`sdk|spring|auto`)
- KnowledgeSearchPort adapter calls the single backend
- Config: collection name, search mode, rrf-k, path weights optional
- Tests for store mapping/fusion wiring; no dependency on live Zilliz in unit tests
## Non-goals
- Automatic full production reindex job UI
- Cross-encoder model rerank
- Keeping dual Spring-AI + SDK knowledge search modes
## Impact
- **Breaking for existing `biz` collection**: requires new hybrid collection + reindex
- Agent tool contract unchanged
- Interface: L2 internal retrieval backend
## Depends on
- Delivery 1 chunk identity
- Delivery 2 SearchPort (will replace sparse-lite hybrid implementation)
@@ -0,0 +1,31 @@
# rag-bm25-hybrid Specification
## ADDED Requirements
### Requirement: Knowledge vector backend SHALL be a single Milvus V2 store
Knowledge indexing and retrieval SHALL use one MilvusClientV2-backed store. Legacy MilvusServiceClient SDK search modes (`sdk`, `auto` fallback to SDK) SHALL NOT be used for `lookup_knowledge`.
#### Scenario: No SDK search mode
- **WHEN** knowledge retrieval executes
- **THEN** it SHALL NOT call legacy SDK `search` APIs for candidate generation
### Requirement: Hybrid mode SHALL fuse dense ANN and BM25 sparse ANN
When search mode is hybrid, the store SHALL query dense vectors and BM25 sparse vectors and fuse results with RRF (or equivalent ranker) before returning hits.
#### Scenario: Hybrid uses BM25 text query
- **WHEN** hybrid search runs with a text query
- **THEN** one search leg SHALL use the BM25/sparse field with the raw query text
- **AND** another leg SHALL use the dense embedding of the query
### Requirement: Write path SHALL populate BM25 input text and dense vectors
Document chunk indexing SHALL write original content, BM25 input text, dense embedding, and metadata required for chunk identity.
#### Scenario: Index writes search_text and vector
- **WHEN** a chunk is indexed
- **THEN** the store row SHALL include searchable text for BM25 and a dense vector field
@@ -0,0 +1,10 @@
# Tasks
- [x] 1. Add MilvusClientV2 factory + hybrid collection bootstrap
- [x] 2. Implement MilvusHybridKnowledgeStore (insert/delete/dense/hybrid)
- [x] 3. Point VectorIndexService writes at V2 store
- [x] 4. Replace VectorSearchService knowledge path with V2 store (drop SDK/spring mode routing)
- [x] 5. Update VectorKnowledgeSearchAdapter to delegate single backend
- [x] 6. Config + constants for collection/mode
- [x] 7. Unit tests (store routing / adapter / no SDK search)
- [x] 8. Archive + commit
+33
View File
@@ -0,0 +1,33 @@
# rag-bm25-hybrid Specification
## Purpose
TBD - created by archiving change rag-bm25-hybrid-drop-sdk. Update Purpose after archive.
## Requirements
### Requirement: Knowledge vector backend SHALL be a single Milvus V2 store
Knowledge indexing and retrieval SHALL use one MilvusClientV2-backed store. Legacy MilvusServiceClient SDK search modes (`sdk`, `auto` fallback to SDK) SHALL NOT be used for `lookup_knowledge`.
#### Scenario: No SDK search mode
- **WHEN** knowledge retrieval executes
- **THEN** it SHALL NOT call legacy SDK `search` APIs for candidate generation
### Requirement: Hybrid mode SHALL fuse dense ANN and BM25 sparse ANN
When search mode is hybrid, the store SHALL query dense vectors and BM25 sparse vectors and fuse results with RRF (or equivalent ranker) before returning hits.
#### Scenario: Hybrid uses BM25 text query
- **WHEN** hybrid search runs with a text query
- **THEN** one search leg SHALL use the BM25/sparse field with the raw query text
- **AND** another leg SHALL use the dense embedding of the query
### Requirement: Write path SHALL populate BM25 input text and dense vectors
Document chunk indexing SHALL write original content, BM25 input text, dense embedding, and metadata required for chunk identity.
#### Scenario: Index writes search_text and vector
- **WHEN** a chunk is indexed
- **THEN** the store row SHALL include searchable text for BM25 and a dense vector field