feat(rag): dense+BM25 hybrid on MilvusClientV2, drop SDK path
Replace legacy MilvusServiceClient knowledge search/write with a single MilvusClientV2 hybrid store (BM25 function + dense ANN + RRFRanker). Use collection biz_hybrid and require knowledge reindex.
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
committed: 2026-07-27
|
||||
authorized-apply: user-preauthorized
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-27
|
||||
@@ -0,0 +1,57 @@
|
||||
# Design: rag-bm25-hybrid-drop-sdk
|
||||
|
||||
## Decision: single backend
|
||||
|
||||
```text
|
||||
KnowledgeSearchPort
|
||||
-> MilvusHybridKnowledgeStore (MilvusClientV2 only)
|
||||
Write
|
||||
-> VectorIndexService -> MilvusHybridKnowledgeStore
|
||||
```
|
||||
|
||||
No `retrieval.vector-store.mode=sdk|spring|auto` for knowledge lookup.
|
||||
|
||||
## Schema (`milvus.collection`, default `biz_hybrid`)
|
||||
|
||||
| field | type | notes |
|
||||
|---|---|---|
|
||||
| id | VarChar PK | chunk id |
|
||||
| content | VarChar | original chunk body returned to agent |
|
||||
| search_text | VarChar + analyzer | BM25 input (title/path/content) |
|
||||
| sparse_vector | SparseFloatVector | BM25 function output |
|
||||
| vector | FloatVector | dense embedding |
|
||||
| metadata | JSON | docId, chunkIndex, category, kb_scope, ... |
|
||||
|
||||
Function: `BM25(search_text -> sparse_vector)`
|
||||
|
||||
Indexes:
|
||||
|
||||
- `vector`: IVF_FLAT / L2 (or COSINE if configured later)
|
||||
- `sparse_vector`: SPARSE_INVERTED_INDEX / BM25
|
||||
|
||||
## Hybrid query
|
||||
|
||||
```text
|
||||
AnnSearchReq(vector, FloatVec(queryEmbedding), topK, filter?)
|
||||
AnnSearchReq(sparse_vector, EmbeddedText(query), topK, filter?)
|
||||
HybridSearchReq + RRFRanker(k)
|
||||
```
|
||||
|
||||
Dense-compatible score for thresholds: if entity lacks dense distance, map fused score conservatively or re-use dense-only probe. Preferred: keep dense distance when available from a parallel dense search metadata; for hybrid hits use inverse-rank placeholder only in metadata and set score from optional dense sub-hit map.
|
||||
|
||||
Implementation approach for threshold stability:
|
||||
|
||||
1. Run hybridSearch for ordering/identity
|
||||
2. Build map evidenceKey -> dense L2 from a concurrent dense search (same filter/topK)
|
||||
3. Attach dense score onto fused hits when present; else maxL2 (low similarity) so weak BM25-only hits don't fake PRECISE
|
||||
|
||||
## Migration
|
||||
|
||||
- New collection name avoids mutating legacy `biz`
|
||||
- Operators re-run knowledge init / document reindex
|
||||
- Document in acceptance
|
||||
|
||||
## Drop SDK search
|
||||
|
||||
- Delete SDK search methods usage from knowledge path
|
||||
- `MilvusServiceClient` bean may remain temporarily only if other non-search utilities need it; prefer migrate write/delete fully to V2 and stop creating V1 client if unused
|
||||
@@ -0,0 +1,34 @@
|
||||
# Change: Real BM25 hybrid search and drop legacy SDK path
|
||||
|
||||
## Why
|
||||
|
||||
Delivery 2 shipped application-layer multi-path + sparse-lite lexical ranking. Product requirement is **true dense + BM25 hybrid retrieval**. Legacy `MilvusServiceClient` search path will be abandoned in this change; keep a single knowledge vector backend based on Milvus Java SDK v2 (`MilvusClientV2`) with BM25 function + `hybridSearch` + RRF.
|
||||
|
||||
## What Changes
|
||||
|
||||
- New hybrid collection schema: dense float vector + analyzed text + sparse BM25 output
|
||||
- Write path inserts text for BM25 auto-sparsification and dense embedding
|
||||
- Search path: single backend
|
||||
- dense mode: dense ANN only
|
||||
- hybrid mode: dense ANN + BM25 sparse ANN fused by RRFRanker
|
||||
- Remove/disable legacy SDK `search` / `searchSimilarDocumentsWithSdk` knowledge path and mode routing (`sdk|spring|auto`)
|
||||
- KnowledgeSearchPort adapter calls the single backend
|
||||
- Config: collection name, search mode, rrf-k, path weights optional
|
||||
- Tests for store mapping/fusion wiring; no dependency on live Zilliz in unit tests
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Automatic full production reindex job UI
|
||||
- Cross-encoder model rerank
|
||||
- Keeping dual Spring-AI + SDK knowledge search modes
|
||||
|
||||
## Impact
|
||||
|
||||
- **Breaking for existing `biz` collection**: requires new hybrid collection + reindex
|
||||
- Agent tool contract unchanged
|
||||
- Interface: L2 internal retrieval backend
|
||||
|
||||
## Depends on
|
||||
|
||||
- Delivery 1 chunk identity
|
||||
- Delivery 2 SearchPort (will replace sparse-lite hybrid implementation)
|
||||
+31
@@ -0,0 +1,31 @@
|
||||
# rag-bm25-hybrid Specification
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Knowledge vector backend SHALL be a single Milvus V2 store
|
||||
|
||||
Knowledge indexing and retrieval SHALL use one MilvusClientV2-backed store. Legacy MilvusServiceClient SDK search modes (`sdk`, `auto` fallback to SDK) SHALL NOT be used for `lookup_knowledge`.
|
||||
|
||||
#### Scenario: No SDK search mode
|
||||
|
||||
- **WHEN** knowledge retrieval executes
|
||||
- **THEN** it SHALL NOT call legacy SDK `search` APIs for candidate generation
|
||||
|
||||
### Requirement: Hybrid mode SHALL fuse dense ANN and BM25 sparse ANN
|
||||
|
||||
When search mode is hybrid, the store SHALL query dense vectors and BM25 sparse vectors and fuse results with RRF (or equivalent ranker) before returning hits.
|
||||
|
||||
#### Scenario: Hybrid uses BM25 text query
|
||||
|
||||
- **WHEN** hybrid search runs with a text query
|
||||
- **THEN** one search leg SHALL use the BM25/sparse field with the raw query text
|
||||
- **AND** another leg SHALL use the dense embedding of the query
|
||||
|
||||
### Requirement: Write path SHALL populate BM25 input text and dense vectors
|
||||
|
||||
Document chunk indexing SHALL write original content, BM25 input text, dense embedding, and metadata required for chunk identity.
|
||||
|
||||
#### Scenario: Index writes search_text and vector
|
||||
|
||||
- **WHEN** a chunk is indexed
|
||||
- **THEN** the store row SHALL include searchable text for BM25 and a dense vector field
|
||||
@@ -0,0 +1,10 @@
|
||||
# Tasks
|
||||
|
||||
- [x] 1. Add MilvusClientV2 factory + hybrid collection bootstrap
|
||||
- [x] 2. Implement MilvusHybridKnowledgeStore (insert/delete/dense/hybrid)
|
||||
- [x] 3. Point VectorIndexService writes at V2 store
|
||||
- [x] 4. Replace VectorSearchService knowledge path with V2 store (drop SDK/spring mode routing)
|
||||
- [x] 5. Update VectorKnowledgeSearchAdapter to delegate single backend
|
||||
- [x] 6. Config + constants for collection/mode
|
||||
- [x] 7. Unit tests (store routing / adapter / no SDK search)
|
||||
- [x] 8. Archive + commit
|
||||
@@ -0,0 +1,33 @@
|
||||
# rag-bm25-hybrid Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change rag-bm25-hybrid-drop-sdk. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Knowledge vector backend SHALL be a single Milvus V2 store
|
||||
|
||||
Knowledge indexing and retrieval SHALL use one MilvusClientV2-backed store. Legacy MilvusServiceClient SDK search modes (`sdk`, `auto` fallback to SDK) SHALL NOT be used for `lookup_knowledge`.
|
||||
|
||||
#### Scenario: No SDK search mode
|
||||
|
||||
- **WHEN** knowledge retrieval executes
|
||||
- **THEN** it SHALL NOT call legacy SDK `search` APIs for candidate generation
|
||||
|
||||
### Requirement: Hybrid mode SHALL fuse dense ANN and BM25 sparse ANN
|
||||
|
||||
When search mode is hybrid, the store SHALL query dense vectors and BM25 sparse vectors and fuse results with RRF (or equivalent ranker) before returning hits.
|
||||
|
||||
#### Scenario: Hybrid uses BM25 text query
|
||||
|
||||
- **WHEN** hybrid search runs with a text query
|
||||
- **THEN** one search leg SHALL use the BM25/sparse field with the raw query text
|
||||
- **AND** another leg SHALL use the dense embedding of the query
|
||||
|
||||
### Requirement: Write path SHALL populate BM25 input text and dense vectors
|
||||
|
||||
Document chunk indexing SHALL write original content, BM25 input text, dense embedding, and metadata required for chunk identity.
|
||||
|
||||
#### Scenario: Index writes search_text and vector
|
||||
|
||||
- **WHEN** a chunk is indexed
|
||||
- **THEN** the store row SHALL include searchable text for BM25 and a dense vector field
|
||||
|
||||
Reference in New Issue
Block a user