test: add rag retrieval baseline
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,62 @@
|
||||
## Context
|
||||
|
||||
The repository already has diagnosis-level evaluation specs and trace observability, but the RAG refactor plan needs a narrower retrieval baseline. The upcoming changes will alter L0 responsibilities, query augmentation, evidence post-processing, and eventually the vector store implementation. Those changes need a fixed set of retrieval cases and deterministic scoring before production retrieval behavior changes.
|
||||
|
||||
The first baseline must be offline. It should not require MySQL, Milvus, Redis, LLM calls, or a running Spring Boot application. It can evaluate saved retrieval result fixtures that represent the current behavior and produce reports that future changes can compare against.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Add a small fixed golden query set for RAG retrieval.
|
||||
- Evaluate retrieval fixtures against expected documents, breadcrumbs, evidence keywords, and hit levels.
|
||||
- Produce JSON and Markdown baseline reports.
|
||||
- Document how to regenerate the reports.
|
||||
- Keep the evaluator simple enough to run from the repository with Python.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not change `lookup_knowledge`, `VectorSearchService`, Milvus schema, L0 matching, or Agent prompts.
|
||||
- Do not require live services.
|
||||
- Do not implement Spring AI VectorStore migration in this change.
|
||||
- Do not implement RRF, BM25, rerank, or evidence packing in this change.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Decision: Use offline retrieval fixtures first
|
||||
|
||||
The evaluator will read saved retrieval fixtures rather than calling the live application.
|
||||
|
||||
Rationale: the first change should establish a stable measurement surface before the RAG internals change. Live retrieval depends on embeddings, Milvus state, and service configuration, which makes it a poor first baseline.
|
||||
|
||||
Alternative considered: call `SearchController` or `lookup_knowledge` directly. That is useful later, but it would require a running app and seeded knowledge base.
|
||||
|
||||
### Decision: Score by hit level, not exact chunk id only
|
||||
|
||||
The evaluator will classify each case as:
|
||||
|
||||
- `strong`: expected document plus expected breadcrumb or key evidence coverage.
|
||||
- `medium`: expected document found, but breadcrumb or evidence coverage is incomplete.
|
||||
- `weak`: related evidence is present but the expected document is missing.
|
||||
- `miss`: no expected document or expected evidence is found.
|
||||
|
||||
Rationale: chunk indexes can change after splitter changes, so exact chunk-only scoring would make later refactors look worse even when evidence quality is preserved.
|
||||
|
||||
### Decision: Keep case format explicit and reviewable
|
||||
|
||||
Golden cases will be stored as JSON with fields such as `caseId`, `query`, `expectedDocIds`, `expectedBreadcrumbs`, `expectedKeywords`, and optional `notes`.
|
||||
|
||||
Rationale: the case file should be easy to inspect in code review and easy to extend during interviews or later refactors.
|
||||
|
||||
### Decision: Preserve both machine and human reports
|
||||
|
||||
The evaluator will write JSON for automation and Markdown for review.
|
||||
|
||||
Rationale: future changes can compare JSON, while the Markdown report is easier to use during design review and interview preparation.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Offline fixtures can drift from real runtime behavior -> add documentation that this is a baseline harness, not a live retrieval accuracy claim.
|
||||
- Keyword-based evidence checks are approximate -> use them only as deterministic guardrails, not as a replacement for human review.
|
||||
- Small golden set may underrepresent production queries -> start with 10-20 cases and expand as new RAG issues are found.
|
||||
- Fixture schema may not match future retrieval outputs -> normalize fixtures into a simple candidate shape and keep raw fields optional.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The RAG refactor needs a repeatable baseline before changing L0, metadata filtering, post-processing, or Spring AI retriever integration. Without fixed retrieval cases and measurable output, later changes can look cleaner architecturally while silently degrading recall or evidence quality.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a retrieval evaluation baseline for RAG queries, separate from full diagnosis evaluation.
|
||||
- Define golden retrieval cases covering Chat-style knowledge lookup and AIOps-style alert diagnosis retrieval.
|
||||
- Add a lightweight offline evaluator that compares retrieved candidates against expected documents, breadcrumbs, and evidence keywords.
|
||||
- Preserve baseline JSON and Markdown reports so future changes can compare retrieval behavior.
|
||||
- No production retrieval behavior changes in this change.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `rag-retrieval-evaluation`: Defines fixed retrieval golden cases, deterministic retrieval evaluation, and baseline report preservation.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Adds retrieval evaluation fixtures, documentation, and scripts.
|
||||
- May read existing retrieval/tool trace output or saved fixtures, but does not require live LLM calls.
|
||||
- Does not change the `lookup_knowledge` runtime behavior, Milvus schema, document upload API, or Agent flow.
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL define fixed golden cases
|
||||
The system SHALL provide a fixed set of RAG retrieval golden cases that can be evaluated deterministically.
|
||||
|
||||
#### Scenario: Golden case includes expected retrieval evidence
|
||||
- **WHEN** a retrieval golden case is defined
|
||||
- **THEN** it SHALL include a case id, query, expected document identifiers or labels, and expected evidence keywords or breadcrumbs
|
||||
|
||||
#### Scenario: Golden case distinguishes scenario type
|
||||
- **WHEN** a retrieval golden case is defined
|
||||
- **THEN** it SHALL indicate whether it covers Chat-style knowledge lookup, AIOps alert diagnosis retrieval, or another explicit retrieval scenario type
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL run offline against fixtures
|
||||
The evaluator SHALL run without requiring live MySQL, Redis, Milvus, LLM, or Spring Boot services.
|
||||
|
||||
#### Scenario: Fixture evaluation
|
||||
- **WHEN** the evaluator is run with a golden case file and retrieval fixture directory
|
||||
- **THEN** it SHALL evaluate each case against its matching fixture file
|
||||
- **AND** it SHALL not call external services
|
||||
|
||||
#### Scenario: Missing fixture is reported
|
||||
- **WHEN** a golden case has no matching retrieval fixture
|
||||
- **THEN** the evaluator SHALL report the case as failed or not run with a clear reason
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL classify hit quality
|
||||
The evaluator SHALL classify each case into a deterministic hit level.
|
||||
|
||||
#### Scenario: Strong hit classification
|
||||
- **WHEN** retrieved candidates include an expected document and satisfy expected breadcrumb or evidence keyword coverage
|
||||
- **THEN** the evaluator SHALL classify the case as `strong`
|
||||
|
||||
#### Scenario: Medium hit classification
|
||||
- **WHEN** retrieved candidates include an expected document but do not satisfy expected breadcrumb or evidence keyword coverage
|
||||
- **THEN** the evaluator SHALL classify the case as `medium`
|
||||
|
||||
#### Scenario: Miss classification
|
||||
- **WHEN** retrieved candidates do not include expected documents or expected evidence
|
||||
- **THEN** the evaluator SHALL classify the case as `miss`
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL report ranking signals
|
||||
The evaluator SHALL report ranking and aggregate retrieval signals suitable for future regression checks.
|
||||
|
||||
#### Scenario: Per-case ranking output
|
||||
- **WHEN** a case is evaluated
|
||||
- **THEN** the report SHALL include hit level, first expected document rank when available, top candidate labels, and failed checks
|
||||
|
||||
#### Scenario: Aggregate metrics output
|
||||
- **WHEN** multiple cases are evaluated
|
||||
- **THEN** the report SHALL include case count, strong hit count, medium hit count, miss count, recall at configured K, and average first hit rank when available
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL preserve baseline reports
|
||||
The system SHALL preserve generated baseline reports in JSON and Markdown formats.
|
||||
|
||||
#### Scenario: Baseline report generation
|
||||
- **WHEN** the baseline evaluator is run for the fixed golden case set
|
||||
- **THEN** it SHALL write a JSON report and a Markdown report under the retrieval evaluation documentation area
|
||||
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes golden cases, fixtures, or evaluator logic
|
||||
- **THEN** the repository SHALL explain how to regenerate the retrieval baseline reports
|
||||
@@ -0,0 +1,26 @@
|
||||
## 1. Golden Cases
|
||||
|
||||
- [x] 1.1 Create retrieval evaluation directory structure.
|
||||
- [x] 1.2 Add fixed golden retrieval cases covering Chat and AIOps retrieval scenarios.
|
||||
- [x] 1.3 Add matching offline retrieval fixtures for every golden case.
|
||||
|
||||
## 2. Evaluator
|
||||
|
||||
- [x] 2.1 Implement an offline retrieval evaluator script.
|
||||
- [x] 2.2 Support hit-level classification and first expected document rank.
|
||||
- [x] 2.3 Support JSON and Markdown report output.
|
||||
|
||||
## 3. Baseline Report
|
||||
|
||||
- [x] 3.1 Generate the baseline JSON report from the fixed cases and fixtures.
|
||||
- [x] 3.2 Generate the baseline Markdown report from the fixed cases and fixtures.
|
||||
|
||||
## 4. Documentation
|
||||
|
||||
- [x] 4.1 Document the retrieval baseline purpose, file layout, and regeneration command.
|
||||
- [x] 4.2 Link the retrieval baseline from the RAG refactor issue or related MVP documentation.
|
||||
|
||||
## 5. Verification
|
||||
|
||||
- [x] 5.1 Run the evaluator successfully against the fixed baseline cases.
|
||||
- [x] 5.2 Run OpenSpec status/validation for the change and confirm tasks are complete.
|
||||
@@ -0,0 +1,64 @@
|
||||
# rag-retrieval-evaluation Specification
|
||||
|
||||
## Purpose
|
||||
Provide a repeatable offline evaluation baseline for RAG retrieval behavior, so L0, query augmentation, evidence post-processing, and vector store changes can be checked against fixed golden retrieval cases before they affect Agent diagnosis quality.
|
||||
## Requirements
|
||||
### Requirement: Retrieval evaluation SHALL define fixed golden cases
|
||||
The system SHALL provide a fixed set of RAG retrieval golden cases that can be evaluated deterministically.
|
||||
|
||||
#### Scenario: Golden case includes expected retrieval evidence
|
||||
- **WHEN** a retrieval golden case is defined
|
||||
- **THEN** it SHALL include a case id, query, expected document identifiers or labels, and expected evidence keywords or breadcrumbs
|
||||
|
||||
#### Scenario: Golden case distinguishes scenario type
|
||||
- **WHEN** a retrieval golden case is defined
|
||||
- **THEN** it SHALL indicate whether it covers Chat-style knowledge lookup, AIOps alert diagnosis retrieval, or another explicit retrieval scenario type
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL run offline against fixtures
|
||||
The evaluator SHALL run without requiring live MySQL, Redis, Milvus, LLM, or Spring Boot services.
|
||||
|
||||
#### Scenario: Fixture evaluation
|
||||
- **WHEN** the evaluator is run with a golden case file and retrieval fixture directory
|
||||
- **THEN** it SHALL evaluate each case against its matching fixture file
|
||||
- **AND** it SHALL not call external services
|
||||
|
||||
#### Scenario: Missing fixture is reported
|
||||
- **WHEN** a golden case has no matching retrieval fixture
|
||||
- **THEN** the evaluator SHALL report the case as failed or not run with a clear reason
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL classify hit quality
|
||||
The evaluator SHALL classify each case into a deterministic hit level.
|
||||
|
||||
#### Scenario: Strong hit classification
|
||||
- **WHEN** retrieved candidates include an expected document and satisfy expected breadcrumb or evidence keyword coverage
|
||||
- **THEN** the evaluator SHALL classify the case as `strong`
|
||||
|
||||
#### Scenario: Medium hit classification
|
||||
- **WHEN** retrieved candidates include an expected document but do not satisfy expected breadcrumb or evidence keyword coverage
|
||||
- **THEN** the evaluator SHALL classify the case as `medium`
|
||||
|
||||
#### Scenario: Miss classification
|
||||
- **WHEN** retrieved candidates do not include expected documents or expected evidence
|
||||
- **THEN** the evaluator SHALL classify the case as `miss`
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL report ranking signals
|
||||
The evaluator SHALL report ranking and aggregate retrieval signals suitable for future regression checks.
|
||||
|
||||
#### Scenario: Per-case ranking output
|
||||
- **WHEN** a case is evaluated
|
||||
- **THEN** the report SHALL include hit level, first expected document rank when available, top candidate labels, and failed checks
|
||||
|
||||
#### Scenario: Aggregate metrics output
|
||||
- **WHEN** multiple cases are evaluated
|
||||
- **THEN** the report SHALL include case count, strong hit count, medium hit count, miss count, recall at configured K, and average first hit rank when available
|
||||
|
||||
### Requirement: Retrieval evaluation SHALL preserve baseline reports
|
||||
The system SHALL preserve generated baseline reports in JSON and Markdown formats.
|
||||
|
||||
#### Scenario: Baseline report generation
|
||||
- **WHEN** the baseline evaluator is run for the fixed golden case set
|
||||
- **THEN** it SHALL write a JSON report and a Markdown report under the retrieval evaluation documentation area
|
||||
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes golden cases, fixtures, or evaluator logic
|
||||
- **THEN** the repository SHALL explain how to regenerate the retrieval baseline reports
|
||||
Reference in New Issue
Block a user