Files
SuperBizAgent-java/openspec/changes/archive/2026-07-04-rag-retrieval-baseline/design.md
T
2026-07-05 02:02:27 +08:00

3.6 KiB

Context

The repository already has diagnosis-level evaluation specs and trace observability, but the RAG refactor plan needs a narrower retrieval baseline. The upcoming changes will alter L0 responsibilities, query augmentation, evidence post-processing, and eventually the vector store implementation. Those changes need a fixed set of retrieval cases and deterministic scoring before production retrieval behavior changes.

The first baseline must be offline. It should not require MySQL, Milvus, Redis, LLM calls, or a running Spring Boot application. It can evaluate saved retrieval result fixtures that represent the current behavior and produce reports that future changes can compare against.

Goals / Non-Goals

Goals:

  • Add a small fixed golden query set for RAG retrieval.
  • Evaluate retrieval fixtures against expected documents, breadcrumbs, evidence keywords, and hit levels.
  • Produce JSON and Markdown baseline reports.
  • Document how to regenerate the reports.
  • Keep the evaluator simple enough to run from the repository with Python.

Non-Goals:

  • Do not change lookup_knowledge, VectorSearchService, Milvus schema, L0 matching, or Agent prompts.
  • Do not require live services.
  • Do not implement Spring AI VectorStore migration in this change.
  • Do not implement RRF, BM25, rerank, or evidence packing in this change.

Decisions

Decision: Use offline retrieval fixtures first

The evaluator will read saved retrieval fixtures rather than calling the live application.

Rationale: the first change should establish a stable measurement surface before the RAG internals change. Live retrieval depends on embeddings, Milvus state, and service configuration, which makes it a poor first baseline.

Alternative considered: call SearchController or lookup_knowledge directly. That is useful later, but it would require a running app and seeded knowledge base.

Decision: Score by hit level, not exact chunk id only

The evaluator will classify each case as:

  • strong: expected document plus expected breadcrumb or key evidence coverage.
  • medium: expected document found, but breadcrumb or evidence coverage is incomplete.
  • weak: related evidence is present but the expected document is missing.
  • miss: no expected document or expected evidence is found.

Rationale: chunk indexes can change after splitter changes, so exact chunk-only scoring would make later refactors look worse even when evidence quality is preserved.

Decision: Keep case format explicit and reviewable

Golden cases will be stored as JSON with fields such as caseId, query, expectedDocIds, expectedBreadcrumbs, expectedKeywords, and optional notes.

Rationale: the case file should be easy to inspect in code review and easy to extend during interviews or later refactors.

Decision: Preserve both machine and human reports

The evaluator will write JSON for automation and Markdown for review.

Rationale: future changes can compare JSON, while the Markdown report is easier to use during design review and interview preparation.

Risks / Trade-offs

  • Offline fixtures can drift from real runtime behavior -> add documentation that this is a baseline harness, not a live retrieval accuracy claim.
  • Keyword-based evidence checks are approximate -> use them only as deterministic guardrails, not as a replacement for human review.
  • Small golden set may underrepresent production queries -> start with 10-20 cases and expand as new RAG issues are found.
  • Fixture schema may not match future retrieval outputs -> normalize fixtures into a simple candidate shape and keep raw fields optional.