4.7 KiB
rag-retrieval-evaluation Specification
Purpose
Provide a repeatable offline evaluation baseline for RAG retrieval behavior, so L0, query augmentation, evidence post-processing, and vector store changes can be checked against fixed golden retrieval cases before they affect Agent diagnosis quality.
Requirements
Requirement: Retrieval evaluation SHALL define fixed golden cases
The system SHALL provide a fixed set of RAG retrieval golden cases that can be evaluated deterministically.
Scenario: Golden case includes expected retrieval evidence
- WHEN a retrieval golden case is defined
- THEN it SHALL include a case id, query, expected document identifiers or labels, and expected evidence keywords or breadcrumbs
Scenario: Golden case distinguishes scenario type
- WHEN a retrieval golden case is defined
- THEN it SHALL indicate whether it covers Chat-style knowledge lookup, AIOps alert diagnosis retrieval, or another explicit retrieval scenario type
Requirement: Retrieval evaluation SHALL run offline against fixtures
The evaluator SHALL run without requiring live MySQL, Redis, Milvus, LLM, or Spring Boot services.
Scenario: Fixture evaluation
- WHEN the evaluator is run with a golden case file and retrieval fixture directory
- THEN it SHALL evaluate each case against its matching fixture file
- AND it SHALL not call external services
Scenario: Missing fixture is reported
- WHEN a golden case has no matching retrieval fixture
- THEN the evaluator SHALL report the case as failed or not run with a clear reason
Requirement: Retrieval evaluation SHALL classify hit quality
The evaluator SHALL classify each case into a deterministic hit level.
Scenario: Strong hit classification
- WHEN retrieved candidates include an expected document and satisfy expected breadcrumb or evidence keyword coverage
- THEN the evaluator SHALL classify the case as
strong
Scenario: Medium hit classification
- WHEN retrieved candidates include an expected document but do not satisfy expected breadcrumb or evidence keyword coverage
- THEN the evaluator SHALL classify the case as
medium
Scenario: Miss classification
- WHEN retrieved candidates do not include expected documents or expected evidence
- THEN the evaluator SHALL classify the case as
miss
Requirement: Retrieval evaluation SHALL report ranking signals
The evaluator SHALL report ranking and aggregate retrieval signals suitable for future regression checks.
Scenario: Per-case ranking output
- WHEN a case is evaluated
- THEN the report SHALL include hit level, first expected document rank when available, top candidate labels, and failed checks
Scenario: Aggregate metrics output
- WHEN multiple cases are evaluated
- THEN the report SHALL include case count, strong hit count, medium hit count, miss count, recall at configured K, and average first hit rank when available
Requirement: Retrieval evaluation SHALL preserve baseline reports
The system SHALL preserve generated baseline reports in JSON and Markdown formats.
Scenario: Baseline report generation
- WHEN the baseline evaluator is run for the fixed golden case set
- THEN it SHALL write a JSON report and a Markdown report under the retrieval evaluation documentation area
Scenario: Baseline regeneration is documented
- WHEN a developer changes golden cases, fixtures, or evaluator logic
- THEN the repository SHALL explain how to regenerate the retrieval baseline reports
Requirement: Retrieval evaluation SHALL compare current and sidecar retrieval paths
The retrieval evaluation system SHALL provide an opt-in comparison between the existing retrieval path and the Spring AI sidecar retrieval path.
Scenario: Sidecar comparison report
- WHEN sidecar comparison is run for the golden case set
- THEN the report SHALL include per-case current-path top candidates and sidecar top candidates
- AND it SHALL highlight source, breadcrumb, category, rank, and score-label differences
Scenario: Offline baseline remains unchanged
- WHEN the fixture-based offline baseline evaluator is run
- THEN it SHALL not require live Milvus, Spring Boot, or Spring AI sidecar configuration
Requirement: Retrieval evaluation SHALL make sidecar readiness visible
The sidecar comparison report SHALL show whether the Spring AI sidecar was runnable for the current environment.
Scenario: Sidecar unavailable
- WHEN sidecar comparison is requested but the sidecar is disabled or unavailable
- THEN the report SHALL mark sidecar status as unavailable
- AND it SHALL keep current-path baseline results available for review