3.6 KiB
Context
The repository already has diagnosis-level evaluation specs and trace observability, but the RAG refactor plan needs a narrower retrieval baseline. The upcoming changes will alter L0 responsibilities, query augmentation, evidence post-processing, and eventually the vector store implementation. Those changes need a fixed set of retrieval cases and deterministic scoring before production retrieval behavior changes.
The first baseline must be offline. It should not require MySQL, Milvus, Redis, LLM calls, or a running Spring Boot application. It can evaluate saved retrieval result fixtures that represent the current behavior and produce reports that future changes can compare against.
Goals / Non-Goals
Goals:
- Add a small fixed golden query set for RAG retrieval.
- Evaluate retrieval fixtures against expected documents, breadcrumbs, evidence keywords, and hit levels.
- Produce JSON and Markdown baseline reports.
- Document how to regenerate the reports.
- Keep the evaluator simple enough to run from the repository with Python.
Non-Goals:
- Do not change
lookup_knowledge,VectorSearchService, Milvus schema, L0 matching, or Agent prompts. - Do not require live services.
- Do not implement Spring AI VectorStore migration in this change.
- Do not implement RRF, BM25, rerank, or evidence packing in this change.
Decisions
Decision: Use offline retrieval fixtures first
The evaluator will read saved retrieval fixtures rather than calling the live application.
Rationale: the first change should establish a stable measurement surface before the RAG internals change. Live retrieval depends on embeddings, Milvus state, and service configuration, which makes it a poor first baseline.
Alternative considered: call SearchController or lookup_knowledge directly. That is useful later, but it would require a running app and seeded knowledge base.
Decision: Score by hit level, not exact chunk id only
The evaluator will classify each case as:
strong: expected document plus expected breadcrumb or key evidence coverage.medium: expected document found, but breadcrumb or evidence coverage is incomplete.weak: related evidence is present but the expected document is missing.miss: no expected document or expected evidence is found.
Rationale: chunk indexes can change after splitter changes, so exact chunk-only scoring would make later refactors look worse even when evidence quality is preserved.
Decision: Keep case format explicit and reviewable
Golden cases will be stored as JSON with fields such as caseId, query, expectedDocIds, expectedBreadcrumbs, expectedKeywords, and optional notes.
Rationale: the case file should be easy to inspect in code review and easy to extend during interviews or later refactors.
Decision: Preserve both machine and human reports
The evaluator will write JSON for automation and Markdown for review.
Rationale: future changes can compare JSON, while the Markdown report is easier to use during design review and interview preparation.
Risks / Trade-offs
- Offline fixtures can drift from real runtime behavior -> add documentation that this is a baseline harness, not a live retrieval accuracy claim.
- Keyword-based evidence checks are approximate -> use them only as deterministic guardrails, not as a replacement for human review.
- Small golden set may underrepresent production queries -> start with 10-20 cases and expand as new RAG issues are found.
- Fixture schema may not match future retrieval outputs -> normalize fixtures into a simple candidate shape and keep raw fields optional.