feat: add rag post-reindex acceptance
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-05
|
||||
@@ -0,0 +1,72 @@
|
||||
## Context
|
||||
|
||||
The indexing path now builds embeddings from structured text:
|
||||
|
||||
```text
|
||||
Title: {title}
|
||||
Path: {breadcrumb}
|
||||
Content:
|
||||
{content}
|
||||
```
|
||||
|
||||
The persisted Milvus `content` field remains the raw chunk content. This improves semantic recall for section-aware questions, but only after documents are reindexed. Existing vectors were generated from the previous content-only input and cannot reflect the new breadcrumb signal.
|
||||
|
||||
The repository already has an offline fixture-based retrieval baseline. That baseline is useful for deterministic regression checks, but it does not prove that the live Milvus/Zilliz collection has been reindexed or that the running service returns breadcrumb-aware results.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Provide an explicit post-reindex live acceptance flow.
|
||||
- Make the reindex prerequisite visible in documentation.
|
||||
- Add a small script that calls the live retrieval endpoint with representative queries and writes reviewable reports.
|
||||
- Keep the live flow optional so unit tests and offline evaluation remain service-free.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not add a new reindex API in this change.
|
||||
- Do not automatically mutate live Milvus/Zilliz data from the acceptance script.
|
||||
- Do not change `lookup_knowledge`, VectorStore retrieval, or Milvus schema.
|
||||
- Do not commit environment-specific live results unless they were intentionally captured for interview evidence.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Decision 1: Keep Reindex Manual And Explicit
|
||||
|
||||
The acceptance flow documents that reindexing must happen before live validation, but it does not perform the reindex itself.
|
||||
|
||||
Rationale:
|
||||
|
||||
- Reindexing is a data mutation and can be slow or environment-specific.
|
||||
- The existing project already has indexing paths through upload, document management, and knowledge-base initialization.
|
||||
- Keeping mutation separate from validation makes failures easier to diagnose.
|
||||
|
||||
Alternative considered: add a script that triggers reindex and then validates. This was rejected for now because it would need environment-specific credentials, source selection, and safety controls.
|
||||
|
||||
### Decision 2: Use HTTP Endpoint Validation
|
||||
|
||||
The script calls `/api/search/similar` instead of invoking Java services directly.
|
||||
|
||||
Rationale:
|
||||
|
||||
- It validates the same runtime path used in demos.
|
||||
- It works across SDK, Spring AI, and auto retrieval modes.
|
||||
- It produces a simple artifact that can be shown in interview material.
|
||||
|
||||
Alternative considered: add a Java integration test. This was rejected because live Milvus and Spring Boot availability should remain optional.
|
||||
|
||||
### Decision 3: Preserve Offline Baseline Separately
|
||||
|
||||
The existing fixture-based evaluator remains the deterministic baseline. The new live acceptance flow is a smoke/regression companion, not a replacement.
|
||||
|
||||
Rationale:
|
||||
|
||||
- Offline reports are stable and CI-friendly.
|
||||
- Live reports prove environment readiness and post-reindex behavior.
|
||||
- Keeping both avoids mixing deterministic fixture checks with external-service validation.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Live results vary by environment, indexed documents, and retrieval mode. -> Mitigation: report the base URL, query set, result count, top candidates, score labels, and timestamp.
|
||||
- [Risk] A developer may run live validation before reindexing. -> Mitigation: document the prerequisite clearly and include a report note.
|
||||
- [Risk] The script could be mistaken for a benchmark. -> Mitigation: position it as acceptance smoke coverage; keep offline baseline for deterministic metrics.
|
||||
@@ -0,0 +1,26 @@
|
||||
## Why
|
||||
|
||||
`title` and `breadcrumb` now participate in embedding text, but that improvement only affects newly indexed vectors. We need a repeatable acceptance path that tells us how to reindex the knowledge base and verify live retrieval after the embedding input changes.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a live RAG retrieval acceptance flow for breadcrumb-aware embedding changes.
|
||||
- Document the reindex prerequisite so reviewers understand old vectors do not change automatically.
|
||||
- Provide a small repeatable script for calling live retrieval cases and writing JSON/Markdown reports.
|
||||
- Add interview-facing acceptance notes that explain what was verified and what remains manual or environment-dependent.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `rag-retrieval-evaluation`: Extend retrieval evaluation with an opt-in live acceptance flow for post-reindex verification.
|
||||
|
||||
## Impact
|
||||
|
||||
- Adds scripts and documentation under the retrieval evaluation/interview areas.
|
||||
- Does not change the Agent runtime path, `lookup_knowledge`, VectorStore search logic, or Milvus schema.
|
||||
- Live verification depends on a running Spring Boot service and a reindexed Milvus/Zilliz collection.
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL provide live post-reindex acceptance
|
||||
The retrieval evaluation system SHALL provide an opt-in live acceptance flow for validating retrieval behavior after embedding input changes require a knowledge-base reindex.
|
||||
|
||||
#### Scenario: Live acceptance requires a running service
|
||||
- **WHEN** live retrieval acceptance is run
|
||||
- **THEN** it SHALL call the configured Spring Boot retrieval endpoint
|
||||
- **AND** it SHALL not be required by the offline fixture baseline
|
||||
|
||||
#### Scenario: Live acceptance records retrieval evidence
|
||||
- **WHEN** a live retrieval case is executed
|
||||
- **THEN** the report SHALL include the query, requested topK, result count, top candidate titles or sources, score labels, and raw response fields needed for review
|
||||
|
||||
#### Scenario: Reindex prerequisite is documented
|
||||
- **WHEN** a developer prepares to validate breadcrumb-aware embedding behavior
|
||||
- **THEN** the repository SHALL explain that existing vectors must be reindexed before live validation can reflect the new embedding text
|
||||
|
||||
#### Scenario: Live report is reviewable
|
||||
- **WHEN** the live acceptance script completes
|
||||
- **THEN** it SHALL write JSON and Markdown outputs that can be inspected or attached to interview evidence
|
||||
@@ -0,0 +1,14 @@
|
||||
## 1. Live Acceptance Tooling
|
||||
|
||||
- [x] 1.1 Add a script that runs representative live `/api/search/similar` queries and writes JSON/Markdown reports.
|
||||
- [x] 1.2 Include breadcrumb-sensitive and core troubleshooting cases in the default live query set.
|
||||
|
||||
## 2. Documentation
|
||||
|
||||
- [x] 2.1 Document the post-reindex validation flow under `eval/rag-retrieval`.
|
||||
- [x] 2.2 Add interview-facing acceptance notes for breadcrumb-aware embedding validation.
|
||||
|
||||
## 3. Verification
|
||||
|
||||
- [x] 3.1 Run targeted tests or syntax checks for the new script.
|
||||
- [x] 3.2 Validate the OpenSpec change and confirm the working tree only contains expected files.
|
||||
Reference in New Issue
Block a user