feat: add rag post-reindex acceptance

This commit is contained in:
aruo
2026-07-05 12:29:24 +08:00
parent c7e2fc2ee2
commit 674dd27a48
8 changed files with 528 additions and 0 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-05
@@ -0,0 +1,72 @@
## Context
The indexing path now builds embeddings from structured text:
```text
Title: {title}
Path: {breadcrumb}
Content:
{content}
```
The persisted Milvus `content` field remains the raw chunk content. This improves semantic recall for section-aware questions, but only after documents are reindexed. Existing vectors were generated from the previous content-only input and cannot reflect the new breadcrumb signal.
The repository already has an offline fixture-based retrieval baseline. That baseline is useful for deterministic regression checks, but it does not prove that the live Milvus/Zilliz collection has been reindexed or that the running service returns breadcrumb-aware results.
## Goals / Non-Goals
**Goals:**
- Provide an explicit post-reindex live acceptance flow.
- Make the reindex prerequisite visible in documentation.
- Add a small script that calls the live retrieval endpoint with representative queries and writes reviewable reports.
- Keep the live flow optional so unit tests and offline evaluation remain service-free.
**Non-Goals:**
- Do not add a new reindex API in this change.
- Do not automatically mutate live Milvus/Zilliz data from the acceptance script.
- Do not change `lookup_knowledge`, VectorStore retrieval, or Milvus schema.
- Do not commit environment-specific live results unless they were intentionally captured for interview evidence.
## Decisions
### Decision 1: Keep Reindex Manual And Explicit
The acceptance flow documents that reindexing must happen before live validation, but it does not perform the reindex itself.
Rationale:
- Reindexing is a data mutation and can be slow or environment-specific.
- The existing project already has indexing paths through upload, document management, and knowledge-base initialization.
- Keeping mutation separate from validation makes failures easier to diagnose.
Alternative considered: add a script that triggers reindex and then validates. This was rejected for now because it would need environment-specific credentials, source selection, and safety controls.
### Decision 2: Use HTTP Endpoint Validation
The script calls `/api/search/similar` instead of invoking Java services directly.
Rationale:
- It validates the same runtime path used in demos.
- It works across SDK, Spring AI, and auto retrieval modes.
- It produces a simple artifact that can be shown in interview material.
Alternative considered: add a Java integration test. This was rejected because live Milvus and Spring Boot availability should remain optional.
### Decision 3: Preserve Offline Baseline Separately
The existing fixture-based evaluator remains the deterministic baseline. The new live acceptance flow is a smoke/regression companion, not a replacement.
Rationale:
- Offline reports are stable and CI-friendly.
- Live reports prove environment readiness and post-reindex behavior.
- Keeping both avoids mixing deterministic fixture checks with external-service validation.
## Risks / Trade-offs
- [Risk] Live results vary by environment, indexed documents, and retrieval mode. -> Mitigation: report the base URL, query set, result count, top candidates, score labels, and timestamp.
- [Risk] A developer may run live validation before reindexing. -> Mitigation: document the prerequisite clearly and include a report note.
- [Risk] The script could be mistaken for a benchmark. -> Mitigation: position it as acceptance smoke coverage; keep offline baseline for deterministic metrics.
@@ -0,0 +1,26 @@
## Why
`title` and `breadcrumb` now participate in embedding text, but that improvement only affects newly indexed vectors. We need a repeatable acceptance path that tells us how to reindex the knowledge base and verify live retrieval after the embedding input changes.
## What Changes
- Add a live RAG retrieval acceptance flow for breadcrumb-aware embedding changes.
- Document the reindex prerequisite so reviewers understand old vectors do not change automatically.
- Provide a small repeatable script for calling live retrieval cases and writing JSON/Markdown reports.
- Add interview-facing acceptance notes that explain what was verified and what remains manual or environment-dependent.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `rag-retrieval-evaluation`: Extend retrieval evaluation with an opt-in live acceptance flow for post-reindex verification.
## Impact
- Adds scripts and documentation under the retrieval evaluation/interview areas.
- Does not change the Agent runtime path, `lookup_knowledge`, VectorStore search logic, or Milvus schema.
- Live verification depends on a running Spring Boot service and a reindexed Milvus/Zilliz collection.
@@ -0,0 +1,21 @@
## ADDED Requirements
### Requirement: Retrieval evaluation SHALL provide live post-reindex acceptance
The retrieval evaluation system SHALL provide an opt-in live acceptance flow for validating retrieval behavior after embedding input changes require a knowledge-base reindex.
#### Scenario: Live acceptance requires a running service
- **WHEN** live retrieval acceptance is run
- **THEN** it SHALL call the configured Spring Boot retrieval endpoint
- **AND** it SHALL not be required by the offline fixture baseline
#### Scenario: Live acceptance records retrieval evidence
- **WHEN** a live retrieval case is executed
- **THEN** the report SHALL include the query, requested topK, result count, top candidate titles or sources, score labels, and raw response fields needed for review
#### Scenario: Reindex prerequisite is documented
- **WHEN** a developer prepares to validate breadcrumb-aware embedding behavior
- **THEN** the repository SHALL explain that existing vectors must be reindexed before live validation can reflect the new embedding text
#### Scenario: Live report is reviewable
- **WHEN** the live acceptance script completes
- **THEN** it SHALL write JSON and Markdown outputs that can be inspected or attached to interview evidence
@@ -0,0 +1,14 @@
## 1. Live Acceptance Tooling
- [x] 1.1 Add a script that runs representative live `/api/search/similar` queries and writes JSON/Markdown reports.
- [x] 1.2 Include breadcrumb-sensitive and core troubleshooting cases in the default live query set.
## 2. Documentation
- [x] 2.1 Document the post-reindex validation flow under `eval/rag-retrieval`.
- [x] 2.2 Add interview-facing acceptance notes for breadcrumb-aware embedding validation.
## 3. Verification
- [x] 3.1 Run targeted tests or syntax checks for the new script.
- [x] 3.2 Validate the OpenSpec change and confirm the working tree only contains expected files.