Files
SuperBizAgent-java/openspec/specs/rag-retrieval-evaluation/spec.md
T

6.5 KiB

rag-retrieval-evaluation Specification

Purpose

Provide a repeatable offline evaluation baseline for RAG retrieval behavior, so L0, query augmentation, evidence post-processing, and vector store changes can be checked against fixed golden retrieval cases before they affect Agent diagnosis quality.

Requirements

Requirement: Retrieval evaluation SHALL define fixed golden cases

The system SHALL provide a fixed set of RAG retrieval golden cases that can be evaluated deterministically.

Scenario: Golden case includes expected retrieval evidence

  • WHEN a retrieval golden case is defined
  • THEN it SHALL include a case id, query, expected document identifiers or labels, and expected evidence keywords or breadcrumbs

Scenario: Golden case distinguishes scenario type

  • WHEN a retrieval golden case is defined
  • THEN it SHALL indicate whether it covers Chat-style knowledge lookup, AIOps alert diagnosis retrieval, or another explicit retrieval scenario type

Requirement: Retrieval evaluation SHALL run offline against fixtures

The evaluator SHALL run without requiring live MySQL, Redis, Milvus, LLM, or Spring Boot services.

Scenario: Fixture evaluation

  • WHEN the evaluator is run with a golden case file and retrieval fixture directory
  • THEN it SHALL evaluate each case against its matching fixture file
  • AND it SHALL not call external services

Scenario: Missing fixture is reported

  • WHEN a golden case has no matching retrieval fixture
  • THEN the evaluator SHALL report the case as failed or not run with a clear reason

Requirement: Retrieval evaluation SHALL classify hit quality

The evaluator SHALL classify each case into a deterministic hit level.

Scenario: Strong hit classification

  • WHEN retrieved candidates include an expected document and satisfy expected breadcrumb or evidence keyword coverage
  • THEN the evaluator SHALL classify the case as strong

Scenario: Medium hit classification

  • WHEN retrieved candidates include an expected document but do not satisfy expected breadcrumb or evidence keyword coverage
  • THEN the evaluator SHALL classify the case as medium

Scenario: Miss classification

  • WHEN retrieved candidates do not include expected documents or expected evidence
  • THEN the evaluator SHALL classify the case as miss

Requirement: Retrieval evaluation SHALL report ranking signals

The evaluator SHALL report ranking and aggregate retrieval signals suitable for future regression checks.

Scenario: Per-case ranking output

  • WHEN a case is evaluated
  • THEN the report SHALL include hit level, first expected document rank when available, top candidate labels, and failed checks

Scenario: Aggregate metrics output

  • WHEN multiple cases are evaluated
  • THEN the report SHALL include case count, strong hit count, medium hit count, miss count, recall at configured K, and average first hit rank when available

Requirement: Retrieval evaluation SHALL preserve baseline reports

The system SHALL preserve generated baseline reports in JSON and Markdown formats.

Scenario: Baseline report generation

  • WHEN the baseline evaluator is run for the fixed golden case set
  • THEN it SHALL write a JSON report and a Markdown report under the retrieval evaluation documentation area

Scenario: Baseline regeneration is documented

  • WHEN a developer changes golden cases, fixtures, or evaluator logic
  • THEN the repository SHALL explain how to regenerate the retrieval baseline reports

Requirement: Retrieval evaluation SHALL compare current and sidecar retrieval paths

The retrieval evaluation system SHALL provide an opt-in comparison between the existing retrieval path and the Spring AI sidecar retrieval path.

Scenario: Sidecar comparison report

  • WHEN sidecar comparison is run for the golden case set
  • THEN the report SHALL include per-case current-path top candidates and sidecar top candidates
  • AND it SHALL highlight source, breadcrumb, category, rank, and score-label differences

Scenario: Offline baseline remains unchanged

  • WHEN the fixture-based offline baseline evaluator is run
  • THEN it SHALL not require live Milvus, Spring Boot, or Spring AI sidecar configuration

Requirement: Retrieval evaluation SHALL make sidecar readiness visible

The sidecar comparison report SHALL show whether the Spring AI sidecar was runnable for the current environment.

Scenario: Sidecar unavailable

  • WHEN sidecar comparison is requested but the sidecar is disabled or unavailable
  • THEN the report SHALL mark sidecar status as unavailable
  • AND it SHALL keep current-path baseline results available for review

Requirement: Retrieval evaluation SHALL remain stable after VectorStore migration

The offline RAG retrieval baseline SHALL remain runnable after the main retrieval service gains Spring AI VectorStore support.

Scenario: Offline evaluator remains service-free

  • WHEN the offline baseline evaluator is run
  • THEN it SHALL not require Spring Boot, live Milvus, Spring AI VectorStore, or the SDK path

Scenario: Baseline is checked during migration

  • WHEN the VectorStore integration change is implemented
  • THEN the existing offline baseline evaluator SHALL be run and its generated report noise SHALL not be committed unless the baseline intentionally changes

Requirement: Retrieval evaluation SHALL provide live post-reindex acceptance

The retrieval evaluation system SHALL provide an opt-in live acceptance flow for validating retrieval behavior after embedding input changes require a knowledge-base reindex.

Scenario: Live acceptance requires a running service

  • WHEN live retrieval acceptance is run
  • THEN it SHALL call the configured Spring Boot retrieval endpoint
  • AND it SHALL not be required by the offline fixture baseline

Scenario: Live acceptance records retrieval evidence

  • WHEN a live retrieval case is executed
  • THEN the report SHALL include the query, requested topK, result count, top candidate titles or sources, score labels, and raw response fields needed for review

Scenario: Reindex prerequisite is documented

  • WHEN a developer prepares to validate breadcrumb-aware embedding behavior
  • THEN the repository SHALL explain that existing vectors must be reindexed before live validation can reflect the new embedding text

Scenario: Live report is reviewable

  • WHEN the live acceptance script completes
  • THEN it SHALL write JSON and Markdown outputs that can be inspected or attached to interview evidence