test: add rag retrieval baseline

This commit is contained in:
aruo
2026-07-05 02:02:27 +08:00
parent 79feed3314
commit 9a2a44d1b5
19 changed files with 946 additions and 0 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,62 @@
## Context
The repository already has diagnosis-level evaluation specs and trace observability, but the RAG refactor plan needs a narrower retrieval baseline. The upcoming changes will alter L0 responsibilities, query augmentation, evidence post-processing, and eventually the vector store implementation. Those changes need a fixed set of retrieval cases and deterministic scoring before production retrieval behavior changes.
The first baseline must be offline. It should not require MySQL, Milvus, Redis, LLM calls, or a running Spring Boot application. It can evaluate saved retrieval result fixtures that represent the current behavior and produce reports that future changes can compare against.
## Goals / Non-Goals
**Goals:**
- Add a small fixed golden query set for RAG retrieval.
- Evaluate retrieval fixtures against expected documents, breadcrumbs, evidence keywords, and hit levels.
- Produce JSON and Markdown baseline reports.
- Document how to regenerate the reports.
- Keep the evaluator simple enough to run from the repository with Python.
**Non-Goals:**
- Do not change `lookup_knowledge`, `VectorSearchService`, Milvus schema, L0 matching, or Agent prompts.
- Do not require live services.
- Do not implement Spring AI VectorStore migration in this change.
- Do not implement RRF, BM25, rerank, or evidence packing in this change.
## Decisions
### Decision: Use offline retrieval fixtures first
The evaluator will read saved retrieval fixtures rather than calling the live application.
Rationale: the first change should establish a stable measurement surface before the RAG internals change. Live retrieval depends on embeddings, Milvus state, and service configuration, which makes it a poor first baseline.
Alternative considered: call `SearchController` or `lookup_knowledge` directly. That is useful later, but it would require a running app and seeded knowledge base.
### Decision: Score by hit level, not exact chunk id only
The evaluator will classify each case as:
- `strong`: expected document plus expected breadcrumb or key evidence coverage.
- `medium`: expected document found, but breadcrumb or evidence coverage is incomplete.
- `weak`: related evidence is present but the expected document is missing.
- `miss`: no expected document or expected evidence is found.
Rationale: chunk indexes can change after splitter changes, so exact chunk-only scoring would make later refactors look worse even when evidence quality is preserved.
### Decision: Keep case format explicit and reviewable
Golden cases will be stored as JSON with fields such as `caseId`, `query`, `expectedDocIds`, `expectedBreadcrumbs`, `expectedKeywords`, and optional `notes`.
Rationale: the case file should be easy to inspect in code review and easy to extend during interviews or later refactors.
### Decision: Preserve both machine and human reports
The evaluator will write JSON for automation and Markdown for review.
Rationale: future changes can compare JSON, while the Markdown report is easier to use during design review and interview preparation.
## Risks / Trade-offs
- Offline fixtures can drift from real runtime behavior -> add documentation that this is a baseline harness, not a live retrieval accuracy claim.
- Keyword-based evidence checks are approximate -> use them only as deterministic guardrails, not as a replacement for human review.
- Small golden set may underrepresent production queries -> start with 10-20 cases and expand as new RAG issues are found.
- Fixture schema may not match future retrieval outputs -> normalize fixtures into a simple candidate shape and keep raw fields optional.
@@ -0,0 +1,27 @@
## Why
The RAG refactor needs a repeatable baseline before changing L0, metadata filtering, post-processing, or Spring AI retriever integration. Without fixed retrieval cases and measurable output, later changes can look cleaner architecturally while silently degrading recall or evidence quality.
## What Changes
- Add a retrieval evaluation baseline for RAG queries, separate from full diagnosis evaluation.
- Define golden retrieval cases covering Chat-style knowledge lookup and AIOps-style alert diagnosis retrieval.
- Add a lightweight offline evaluator that compares retrieved candidates against expected documents, breadcrumbs, and evidence keywords.
- Preserve baseline JSON and Markdown reports so future changes can compare retrieval behavior.
- No production retrieval behavior changes in this change.
## Capabilities
### New Capabilities
- `rag-retrieval-evaluation`: Defines fixed retrieval golden cases, deterministic retrieval evaluation, and baseline report preservation.
### Modified Capabilities
- None.
## Impact
- Adds retrieval evaluation fixtures, documentation, and scripts.
- May read existing retrieval/tool trace output or saved fixtures, but does not require live LLM calls.
- Does not change the `lookup_knowledge` runtime behavior, Milvus schema, document upload API, or Agent flow.
@@ -0,0 +1,61 @@
## ADDED Requirements
### Requirement: Retrieval evaluation SHALL define fixed golden cases
The system SHALL provide a fixed set of RAG retrieval golden cases that can be evaluated deterministically.
#### Scenario: Golden case includes expected retrieval evidence
- **WHEN** a retrieval golden case is defined
- **THEN** it SHALL include a case id, query, expected document identifiers or labels, and expected evidence keywords or breadcrumbs
#### Scenario: Golden case distinguishes scenario type
- **WHEN** a retrieval golden case is defined
- **THEN** it SHALL indicate whether it covers Chat-style knowledge lookup, AIOps alert diagnosis retrieval, or another explicit retrieval scenario type
### Requirement: Retrieval evaluation SHALL run offline against fixtures
The evaluator SHALL run without requiring live MySQL, Redis, Milvus, LLM, or Spring Boot services.
#### Scenario: Fixture evaluation
- **WHEN** the evaluator is run with a golden case file and retrieval fixture directory
- **THEN** it SHALL evaluate each case against its matching fixture file
- **AND** it SHALL not call external services
#### Scenario: Missing fixture is reported
- **WHEN** a golden case has no matching retrieval fixture
- **THEN** the evaluator SHALL report the case as failed or not run with a clear reason
### Requirement: Retrieval evaluation SHALL classify hit quality
The evaluator SHALL classify each case into a deterministic hit level.
#### Scenario: Strong hit classification
- **WHEN** retrieved candidates include an expected document and satisfy expected breadcrumb or evidence keyword coverage
- **THEN** the evaluator SHALL classify the case as `strong`
#### Scenario: Medium hit classification
- **WHEN** retrieved candidates include an expected document but do not satisfy expected breadcrumb or evidence keyword coverage
- **THEN** the evaluator SHALL classify the case as `medium`
#### Scenario: Miss classification
- **WHEN** retrieved candidates do not include expected documents or expected evidence
- **THEN** the evaluator SHALL classify the case as `miss`
### Requirement: Retrieval evaluation SHALL report ranking signals
The evaluator SHALL report ranking and aggregate retrieval signals suitable for future regression checks.
#### Scenario: Per-case ranking output
- **WHEN** a case is evaluated
- **THEN** the report SHALL include hit level, first expected document rank when available, top candidate labels, and failed checks
#### Scenario: Aggregate metrics output
- **WHEN** multiple cases are evaluated
- **THEN** the report SHALL include case count, strong hit count, medium hit count, miss count, recall at configured K, and average first hit rank when available
### Requirement: Retrieval evaluation SHALL preserve baseline reports
The system SHALL preserve generated baseline reports in JSON and Markdown formats.
#### Scenario: Baseline report generation
- **WHEN** the baseline evaluator is run for the fixed golden case set
- **THEN** it SHALL write a JSON report and a Markdown report under the retrieval evaluation documentation area
#### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes golden cases, fixtures, or evaluator logic
- **THEN** the repository SHALL explain how to regenerate the retrieval baseline reports
@@ -0,0 +1,26 @@
## 1. Golden Cases
- [x] 1.1 Create retrieval evaluation directory structure.
- [x] 1.2 Add fixed golden retrieval cases covering Chat and AIOps retrieval scenarios.
- [x] 1.3 Add matching offline retrieval fixtures for every golden case.
## 2. Evaluator
- [x] 2.1 Implement an offline retrieval evaluator script.
- [x] 2.2 Support hit-level classification and first expected document rank.
- [x] 2.3 Support JSON and Markdown report output.
## 3. Baseline Report
- [x] 3.1 Generate the baseline JSON report from the fixed cases and fixtures.
- [x] 3.2 Generate the baseline Markdown report from the fixed cases and fixtures.
## 4. Documentation
- [x] 4.1 Document the retrieval baseline purpose, file layout, and regeneration command.
- [x] 4.2 Link the retrieval baseline from the RAG refactor issue or related MVP documentation.
## 5. Verification
- [x] 5.1 Run the evaluator successfully against the fixed baseline cases.
- [x] 5.2 Run OpenSpec status/validation for the change and confirm tasks are complete.
@@ -0,0 +1,64 @@
# rag-retrieval-evaluation Specification
## Purpose
Provide a repeatable offline evaluation baseline for RAG retrieval behavior, so L0, query augmentation, evidence post-processing, and vector store changes can be checked against fixed golden retrieval cases before they affect Agent diagnosis quality.
## Requirements
### Requirement: Retrieval evaluation SHALL define fixed golden cases
The system SHALL provide a fixed set of RAG retrieval golden cases that can be evaluated deterministically.
#### Scenario: Golden case includes expected retrieval evidence
- **WHEN** a retrieval golden case is defined
- **THEN** it SHALL include a case id, query, expected document identifiers or labels, and expected evidence keywords or breadcrumbs
#### Scenario: Golden case distinguishes scenario type
- **WHEN** a retrieval golden case is defined
- **THEN** it SHALL indicate whether it covers Chat-style knowledge lookup, AIOps alert diagnosis retrieval, or another explicit retrieval scenario type
### Requirement: Retrieval evaluation SHALL run offline against fixtures
The evaluator SHALL run without requiring live MySQL, Redis, Milvus, LLM, or Spring Boot services.
#### Scenario: Fixture evaluation
- **WHEN** the evaluator is run with a golden case file and retrieval fixture directory
- **THEN** it SHALL evaluate each case against its matching fixture file
- **AND** it SHALL not call external services
#### Scenario: Missing fixture is reported
- **WHEN** a golden case has no matching retrieval fixture
- **THEN** the evaluator SHALL report the case as failed or not run with a clear reason
### Requirement: Retrieval evaluation SHALL classify hit quality
The evaluator SHALL classify each case into a deterministic hit level.
#### Scenario: Strong hit classification
- **WHEN** retrieved candidates include an expected document and satisfy expected breadcrumb or evidence keyword coverage
- **THEN** the evaluator SHALL classify the case as `strong`
#### Scenario: Medium hit classification
- **WHEN** retrieved candidates include an expected document but do not satisfy expected breadcrumb or evidence keyword coverage
- **THEN** the evaluator SHALL classify the case as `medium`
#### Scenario: Miss classification
- **WHEN** retrieved candidates do not include expected documents or expected evidence
- **THEN** the evaluator SHALL classify the case as `miss`
### Requirement: Retrieval evaluation SHALL report ranking signals
The evaluator SHALL report ranking and aggregate retrieval signals suitable for future regression checks.
#### Scenario: Per-case ranking output
- **WHEN** a case is evaluated
- **THEN** the report SHALL include hit level, first expected document rank when available, top candidate labels, and failed checks
#### Scenario: Aggregate metrics output
- **WHEN** multiple cases are evaluated
- **THEN** the report SHALL include case count, strong hit count, medium hit count, miss count, recall at configured K, and average first hit rank when available
### Requirement: Retrieval evaluation SHALL preserve baseline reports
The system SHALL preserve generated baseline reports in JSON and Markdown formats.
#### Scenario: Baseline report generation
- **WHEN** the baseline evaluator is run for the fixed golden case set
- **THEN** it SHALL write a JSON report and a Markdown report under the retrieval evaluation documentation area
#### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes golden cases, fixtures, or evaluator logic
- **THEN** the repository SHALL explain how to regenerate the retrieval baseline reports