# rag-retrieval-evaluation Specification ## Purpose Provide a repeatable offline evaluation baseline for RAG retrieval behavior, so L0, query augmentation, evidence post-processing, and vector store changes can be checked against fixed golden retrieval cases before they affect Agent diagnosis quality. ## Requirements ### Requirement: Retrieval evaluation SHALL define fixed golden cases The system SHALL provide a fixed set of RAG retrieval golden cases that can be evaluated deterministically. #### Scenario: Golden case includes expected retrieval evidence - **WHEN** a retrieval golden case is defined - **THEN** it SHALL include a case id, query, expected document identifiers or labels, and expected evidence keywords or breadcrumbs #### Scenario: Golden case distinguishes scenario type - **WHEN** a retrieval golden case is defined - **THEN** it SHALL indicate whether it covers Chat-style knowledge lookup, AIOps alert diagnosis retrieval, or another explicit retrieval scenario type ### Requirement: Retrieval evaluation SHALL run offline against fixtures The evaluator SHALL run without requiring live MySQL, Redis, Milvus, LLM, or Spring Boot services. #### Scenario: Fixture evaluation - **WHEN** the evaluator is run with a golden case file and retrieval fixture directory - **THEN** it SHALL evaluate each case against its matching fixture file - **AND** it SHALL not call external services #### Scenario: Missing fixture is reported - **WHEN** a golden case has no matching retrieval fixture - **THEN** the evaluator SHALL report the case as failed or not run with a clear reason ### Requirement: Retrieval evaluation SHALL classify hit quality The evaluator SHALL classify each case into a deterministic hit level. #### Scenario: Strong hit classification - **WHEN** retrieved candidates include an expected document and satisfy expected breadcrumb or evidence keyword coverage - **THEN** the evaluator SHALL classify the case as `strong` #### Scenario: Medium hit classification - **WHEN** retrieved candidates include an expected document but do not satisfy expected breadcrumb or evidence keyword coverage - **THEN** the evaluator SHALL classify the case as `medium` #### Scenario: Miss classification - **WHEN** retrieved candidates do not include expected documents or expected evidence - **THEN** the evaluator SHALL classify the case as `miss` ### Requirement: Retrieval evaluation SHALL report ranking signals The evaluator SHALL report ranking and aggregate retrieval signals suitable for future regression checks. #### Scenario: Per-case ranking output - **WHEN** a case is evaluated - **THEN** the report SHALL include hit level, first expected document rank when available, top candidate labels, and failed checks #### Scenario: Aggregate metrics output - **WHEN** multiple cases are evaluated - **THEN** the report SHALL include case count, strong hit count, medium hit count, miss count, recall at configured K, and average first hit rank when available ### Requirement: Retrieval evaluation SHALL preserve baseline reports The system SHALL preserve generated baseline reports in JSON and Markdown formats. #### Scenario: Baseline report generation - **WHEN** the baseline evaluator is run for the fixed golden case set - **THEN** it SHALL write a JSON report and a Markdown report under the retrieval evaluation documentation area #### Scenario: Baseline regeneration is documented - **WHEN** a developer changes golden cases, fixtures, or evaluator logic - **THEN** the repository SHALL explain how to regenerate the retrieval baseline reports ### Requirement: Retrieval evaluation SHALL compare current and sidecar retrieval paths The retrieval evaluation system SHALL provide an opt-in comparison between the existing retrieval path and the Spring AI sidecar retrieval path. #### Scenario: Sidecar comparison report - **WHEN** sidecar comparison is run for the golden case set - **THEN** the report SHALL include per-case current-path top candidates and sidecar top candidates - **AND** it SHALL highlight source, breadcrumb, category, rank, and score-label differences #### Scenario: Offline baseline remains unchanged - **WHEN** the fixture-based offline baseline evaluator is run - **THEN** it SHALL not require live Milvus, Spring Boot, or Spring AI sidecar configuration ### Requirement: Retrieval evaluation SHALL make sidecar readiness visible The sidecar comparison report SHALL show whether the Spring AI sidecar was runnable for the current environment. #### Scenario: Sidecar unavailable - **WHEN** sidecar comparison is requested but the sidecar is disabled or unavailable - **THEN** the report SHALL mark sidecar status as unavailable - **AND** it SHALL keep current-path baseline results available for review ### Requirement: Retrieval evaluation SHALL remain stable after VectorStore migration The offline RAG retrieval baseline SHALL remain runnable after the main retrieval service gains Spring AI VectorStore support. #### Scenario: Offline evaluator remains service-free - **WHEN** the offline baseline evaluator is run - **THEN** it SHALL not require Spring Boot, live Milvus, Spring AI VectorStore, or the SDK path #### Scenario: Baseline is checked during migration - **WHEN** the VectorStore integration change is implemented - **THEN** the existing offline baseline evaluator SHALL be run and its generated report noise SHALL not be committed unless the baseline intentionally changes ### Requirement: Retrieval evaluation SHALL provide live post-reindex acceptance The retrieval evaluation system SHALL provide an opt-in live acceptance flow for validating retrieval behavior after embedding input changes require a knowledge-base reindex. #### Scenario: Live acceptance requires a running service - **WHEN** live retrieval acceptance is run - **THEN** it SHALL call the configured Spring Boot retrieval endpoint - **AND** it SHALL not be required by the offline fixture baseline #### Scenario: Live acceptance records retrieval evidence - **WHEN** a live retrieval case is executed - **THEN** the report SHALL include the query, requested topK, result count, top candidate titles or sources, score labels, and raw response fields needed for review #### Scenario: Reindex prerequisite is documented - **WHEN** a developer prepares to validate breadcrumb-aware embedding behavior - **THEN** the repository SHALL explain that existing vectors must be reindexed before live validation can reflect the new embedding text #### Scenario: Live report is reviewable - **WHEN** the live acceptance script completes - **THEN** it SHALL write JSON and Markdown outputs that can be inspected or attached to interview evidence