feat: integrate spring ai vectorstore fallback
This commit is contained in:
+2
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-05
|
||||
+72
@@ -0,0 +1,72 @@
|
||||
## Context
|
||||
|
||||
The project currently uses `VectorSearchService` to call Milvus directly through the Java SDK. A previous change added a Spring AI `VectorStore` sidecar and normalized its results, but the production path still uses SDK-only retrieval. Spring AI provides an official Milvus VectorStore starter, so the main path can now move to the framework abstraction without deleting the proven SDK implementation.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Use official Spring AI Milvus VectorStore integration.
|
||||
- Keep existing Milvus collection compatibility by configuring field names and embedding dimension.
|
||||
- Preserve current `VectorSearchService` public API.
|
||||
- Support retrieval mode selection:
|
||||
- `spring-ai`: use VectorStore and fail if unavailable.
|
||||
- `sdk`: use existing SDK path.
|
||||
- `auto`: try VectorStore, then fall back to SDK.
|
||||
- Preserve existing score semantics in normalized results by labeling Spring AI scores as `similarity` and SDK scores as `l2_distance`.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not remove Milvus SDK code.
|
||||
- Do not migrate document writes/indexing to Spring AI in this change.
|
||||
- Do not change chunking, metadata shape, or evidence post-processing behavior.
|
||||
- Do not introduce QueryTransformer, MultiQuery, rerank, or neighbor chunk expansion.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Decision: VectorSearchService remains the boundary
|
||||
|
||||
`LookupKnowledgeTool` will keep calling `VectorSearchService.searchSimilarDocuments(...)`.
|
||||
|
||||
Rationale: this protects Agent and AIOps behavior from retrieval implementation churn and keeps the refactor testable.
|
||||
|
||||
### Decision: Auto fallback is the default
|
||||
|
||||
Configure `retrieval.vector-store.mode=auto` so the system prefers Spring AI VectorStore when available but falls back to the existing SDK path on missing beans or runtime errors.
|
||||
|
||||
Rationale: official VectorStore integration may expose schema or scoring differences; fallback keeps the MVP runnable.
|
||||
|
||||
### Decision: Existing collection is reused
|
||||
|
||||
Spring AI Milvus configuration will map to the current collection:
|
||||
|
||||
- id field: `id`
|
||||
- content field: `content`
|
||||
- embedding field: `vector`
|
||||
- metadata field: `metadata`
|
||||
- embedding dimension: `1024`
|
||||
- metric type: `L2`
|
||||
|
||||
Rationale: this avoids reindexing as part of this change and lets golden cases reveal behavior differences first.
|
||||
|
||||
### Decision: SDK indexing remains for now
|
||||
|
||||
`VectorIndexService` continues writing to Milvus using SDK.
|
||||
|
||||
Rationale: replacing both read and write paths at once would make failures harder to isolate. The current change is read-path migration.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Spring AI filter syntax may not map perfectly to Milvus JSON metadata filters -> keep SDK fallback and add tests for filter expression creation.
|
||||
- Spring AI score may be similarity while SDK score is L2 distance -> keep score label explicit.
|
||||
- Auto-configuration could create a VectorStore bean against an incompatible collection -> make retrieval mode configurable and validate with golden cases.
|
||||
- Keeping two paths adds temporary complexity -> isolate SDK and VectorStore code paths inside `VectorSearchService`.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add Spring AI Milvus starter dependency and configuration.
|
||||
2. Add retrieval mode properties.
|
||||
3. Refactor `VectorSearchService` to prefer VectorStore based on mode.
|
||||
4. Preserve and test SDK fallback.
|
||||
5. Run targeted tests and the offline RAG baseline.
|
||||
6. In a later change, decide whether to migrate indexing/writes after read-path behavior is stable.
|
||||
+28
@@ -0,0 +1,28 @@
|
||||
## Why
|
||||
|
||||
The previous sidecar change proved the project can normalize Spring AI `VectorStore` results without changing the Agent tool boundary. The next step is to integrate the official Spring AI Milvus VectorStore into the main retrieval service while preserving the existing Milvus SDK path as a fallback.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add the official Spring AI Milvus VectorStore starter dependency.
|
||||
- Configure Spring AI Milvus to reuse the existing collection, field names, embedding dimension, metric type, and connection settings.
|
||||
- Update `VectorSearchService` to support selectable retrieval modes: Spring AI VectorStore, SDK, or automatic fallback.
|
||||
- Preserve the existing `searchSimilarDocuments(query, topK, category)` API used by `lookup_knowledge`.
|
||||
- Keep the current Milvus SDK implementation available and covered by tests.
|
||||
- Add tests proving SDK fallback is used when VectorStore is unavailable or fails.
|
||||
- No breaking API changes.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
- `rag-knowledge-retrieval`: Add requirements for using Spring AI VectorStore as the preferred retrieval abstraction while preserving SDK fallback and traceable score semantics.
|
||||
- `rag-retrieval-evaluation`: Add requirements that the offline baseline remains stable after the retrieval implementation changes.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects `pom.xml`, RAG/Milvus configuration, `VectorSearchService`, and related tests.
|
||||
- Does not change `lookup_knowledge` tool signature, evidence block format, document chunking, upload API, or `tool_invocation` schema.
|
||||
- Uses Spring AI Milvus integration but keeps the existing Milvus SDK code path for rollback and compatibility.
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL prefer Spring AI VectorStore when configured
|
||||
The retrieval service SHALL support Spring AI VectorStore as the preferred vector retrieval abstraction without changing the `lookup_knowledge` tool contract.
|
||||
|
||||
#### Scenario: VectorStore mode uses Spring AI
|
||||
- **WHEN** retrieval vector store mode is configured as `spring-ai`
|
||||
- **THEN** semantic retrieval SHALL query through Spring AI `VectorStore`
|
||||
- **AND** the returned candidates SHALL be normalized into the existing vector search result shape
|
||||
|
||||
#### Scenario: Auto mode prefers VectorStore
|
||||
- **WHEN** retrieval vector store mode is configured as `auto`
|
||||
- **AND** a Spring AI `VectorStore` bean is available
|
||||
- **THEN** semantic retrieval SHALL attempt Spring AI `VectorStore` before the SDK path
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL preserve SDK fallback
|
||||
The retrieval service SHALL keep the existing Milvus SDK retrieval implementation available.
|
||||
|
||||
#### Scenario: SDK mode bypasses VectorStore
|
||||
- **WHEN** retrieval vector store mode is configured as `sdk`
|
||||
- **THEN** semantic retrieval SHALL use the existing Milvus SDK path
|
||||
|
||||
#### Scenario: Auto fallback uses SDK
|
||||
- **WHEN** retrieval vector store mode is `auto`
|
||||
- **AND** Spring AI `VectorStore` is unavailable or fails
|
||||
- **THEN** semantic retrieval SHALL fall back to the existing Milvus SDK path
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL keep score semantics explicit
|
||||
The retrieval service SHALL preserve score semantics when results come from different retrieval implementations.
|
||||
|
||||
#### Scenario: SDK score remains L2 distance
|
||||
- **WHEN** a candidate is returned by the SDK path
|
||||
- **THEN** its score semantics SHALL remain compatible with existing L2 distance normalization
|
||||
|
||||
#### Scenario: VectorStore score is mapped without changing tool contract
|
||||
- **WHEN** a candidate is returned by Spring AI `VectorStore`
|
||||
- **THEN** it SHALL be mapped into the existing result shape
|
||||
- **AND** trace or comparison code SHALL be able to distinguish it as a VectorStore similarity score when needed
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL reuse the existing Milvus collection
|
||||
Spring AI Milvus integration SHALL be configured to use the existing collection schema unless explicitly changed.
|
||||
|
||||
#### Scenario: Existing field mapping
|
||||
- **WHEN** Spring AI Milvus VectorStore is configured
|
||||
- **THEN** it SHALL use the existing id, content, vector, and metadata field names
|
||||
- **AND** it SHALL use the configured embedding dimension and metric type compatible with existing vectors
|
||||
+12
@@ -0,0 +1,12 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL remain stable after VectorStore migration
|
||||
The offline RAG retrieval baseline SHALL remain runnable after the main retrieval service gains Spring AI VectorStore support.
|
||||
|
||||
#### Scenario: Offline evaluator remains service-free
|
||||
- **WHEN** the offline baseline evaluator is run
|
||||
- **THEN** it SHALL not require Spring Boot, live Milvus, Spring AI VectorStore, or the SDK path
|
||||
|
||||
#### Scenario: Baseline is checked during migration
|
||||
- **WHEN** the VectorStore integration change is implemented
|
||||
- **THEN** the existing offline baseline evaluator SHALL be run and its generated report noise SHALL not be committed unless the baseline intentionally changes
|
||||
+26
@@ -0,0 +1,26 @@
|
||||
## 1. Spring AI Milvus Setup
|
||||
|
||||
- [x] 1.1 Add the official Spring AI Milvus VectorStore starter dependency.
|
||||
- [x] 1.2 Configure Spring AI Milvus to reuse the existing collection field names, dimension, metric type, and connection settings.
|
||||
- [x] 1.3 Add retrieval mode configuration for `auto`, `spring-ai`, and `sdk`.
|
||||
|
||||
## 2. Retrieval Service Refactor
|
||||
|
||||
- [x] 2.1 Refactor `VectorSearchService` to inject optional Spring AI `VectorStore`.
|
||||
- [x] 2.2 Implement VectorStore search and map Spring AI `Document` results to existing `SearchResult`.
|
||||
- [x] 2.3 Preserve the existing SDK search path as a dedicated fallback method.
|
||||
- [x] 2.4 Route retrieval by configured mode and fallback rules.
|
||||
|
||||
## 3. Tests
|
||||
|
||||
- [x] 3.1 Add tests for SDK mode bypassing VectorStore.
|
||||
- [x] 3.2 Add tests for auto mode using VectorStore when available.
|
||||
- [x] 3.3 Add tests for auto mode falling back to SDK when VectorStore fails or is unavailable.
|
||||
- [x] 3.4 Add tests for category metadata filter behavior.
|
||||
|
||||
## 4. Validation And Archive
|
||||
|
||||
- [x] 4.1 Run targeted retrieval tests.
|
||||
- [x] 4.2 Run the offline RAG retrieval baseline evaluator.
|
||||
- [x] 4.3 Run OpenSpec validation.
|
||||
- [x] 4.4 Review git diff to confirm `lookup_knowledge` API and evidence output remain compatible.
|
||||
@@ -106,3 +106,48 @@ The sidecar retrieval path SHALL expose results in a comparable structure aligne
|
||||
#### Scenario: Score semantics are explicit
|
||||
- **WHEN** current retrieval and sidecar retrieval scores are compared
|
||||
- **THEN** the report SHALL label score semantics by path instead of assuming direct numeric equivalence
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL prefer Spring AI VectorStore when configured
|
||||
The retrieval service SHALL support Spring AI VectorStore as the preferred vector retrieval abstraction without changing the `lookup_knowledge` tool contract.
|
||||
|
||||
#### Scenario: VectorStore mode uses Spring AI
|
||||
- **WHEN** retrieval vector store mode is configured as `spring-ai`
|
||||
- **THEN** semantic retrieval SHALL query through Spring AI `VectorStore`
|
||||
- **AND** the returned candidates SHALL be normalized into the existing vector search result shape
|
||||
|
||||
#### Scenario: Auto mode prefers VectorStore
|
||||
- **WHEN** retrieval vector store mode is configured as `auto`
|
||||
- **AND** a Spring AI `VectorStore` bean is available
|
||||
- **THEN** semantic retrieval SHALL attempt Spring AI `VectorStore` before the SDK path
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL preserve SDK fallback
|
||||
The retrieval service SHALL keep the existing Milvus SDK retrieval implementation available.
|
||||
|
||||
#### Scenario: SDK mode bypasses VectorStore
|
||||
- **WHEN** retrieval vector store mode is configured as `sdk`
|
||||
- **THEN** semantic retrieval SHALL use the existing Milvus SDK path
|
||||
|
||||
#### Scenario: Auto fallback uses SDK
|
||||
- **WHEN** retrieval vector store mode is `auto`
|
||||
- **AND** Spring AI `VectorStore` is unavailable or fails
|
||||
- **THEN** semantic retrieval SHALL fall back to the existing Milvus SDK path
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL keep score semantics explicit
|
||||
The retrieval service SHALL preserve score semantics when results come from different retrieval implementations.
|
||||
|
||||
#### Scenario: SDK score remains L2 distance
|
||||
- **WHEN** a candidate is returned by the SDK path
|
||||
- **THEN** its score semantics SHALL remain compatible with existing L2 distance normalization
|
||||
|
||||
#### Scenario: VectorStore score is mapped without changing tool contract
|
||||
- **WHEN** a candidate is returned by Spring AI `VectorStore`
|
||||
- **THEN** it SHALL be mapped into the existing result shape
|
||||
- **AND** trace or comparison code SHALL be able to distinguish it as a VectorStore similarity score when needed
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL reuse the existing Milvus collection
|
||||
Spring AI Milvus integration SHALL be configured to use the existing collection schema unless explicitly changed.
|
||||
|
||||
#### Scenario: Existing field mapping
|
||||
- **WHEN** Spring AI Milvus VectorStore is configured
|
||||
- **THEN** it SHALL use the existing id, content, vector, and metadata field names
|
||||
- **AND** it SHALL use the configured embedding dimension and metric type compatible with existing vectors
|
||||
|
||||
@@ -82,3 +82,14 @@ The sidecar comparison report SHALL show whether the Spring AI sidecar was runna
|
||||
- **WHEN** sidecar comparison is requested but the sidecar is disabled or unavailable
|
||||
- **THEN** the report SHALL mark sidecar status as unavailable
|
||||
- **AND** it SHALL keep current-path baseline results available for review
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL remain stable after VectorStore migration
|
||||
The offline RAG retrieval baseline SHALL remain runnable after the main retrieval service gains Spring AI VectorStore support.
|
||||
|
||||
#### Scenario: Offline evaluator remains service-free
|
||||
- **WHEN** the offline baseline evaluator is run
|
||||
- **THEN** it SHALL not require Spring Boot, live Milvus, Spring AI VectorStore, or the SDK path
|
||||
|
||||
#### Scenario: Baseline is checked during migration
|
||||
- **WHEN** the VectorStore integration change is implemented
|
||||
- **THEN** the existing offline baseline evaluator SHALL be run and its generated report noise SHALL not be committed unless the baseline intentionally changes
|
||||
|
||||
Reference in New Issue
Block a user