## Context The repository already has diagnosis-level evaluation specs and trace observability, but the RAG refactor plan needs a narrower retrieval baseline. The upcoming changes will alter L0 responsibilities, query augmentation, evidence post-processing, and eventually the vector store implementation. Those changes need a fixed set of retrieval cases and deterministic scoring before production retrieval behavior changes. The first baseline must be offline. It should not require MySQL, Milvus, Redis, LLM calls, or a running Spring Boot application. It can evaluate saved retrieval result fixtures that represent the current behavior and produce reports that future changes can compare against. ## Goals / Non-Goals **Goals:** - Add a small fixed golden query set for RAG retrieval. - Evaluate retrieval fixtures against expected documents, breadcrumbs, evidence keywords, and hit levels. - Produce JSON and Markdown baseline reports. - Document how to regenerate the reports. - Keep the evaluator simple enough to run from the repository with Python. **Non-Goals:** - Do not change `lookup_knowledge`, `VectorSearchService`, Milvus schema, L0 matching, or Agent prompts. - Do not require live services. - Do not implement Spring AI VectorStore migration in this change. - Do not implement RRF, BM25, rerank, or evidence packing in this change. ## Decisions ### Decision: Use offline retrieval fixtures first The evaluator will read saved retrieval fixtures rather than calling the live application. Rationale: the first change should establish a stable measurement surface before the RAG internals change. Live retrieval depends on embeddings, Milvus state, and service configuration, which makes it a poor first baseline. Alternative considered: call `SearchController` or `lookup_knowledge` directly. That is useful later, but it would require a running app and seeded knowledge base. ### Decision: Score by hit level, not exact chunk id only The evaluator will classify each case as: - `strong`: expected document plus expected breadcrumb or key evidence coverage. - `medium`: expected document found, but breadcrumb or evidence coverage is incomplete. - `weak`: related evidence is present but the expected document is missing. - `miss`: no expected document or expected evidence is found. Rationale: chunk indexes can change after splitter changes, so exact chunk-only scoring would make later refactors look worse even when evidence quality is preserved. ### Decision: Keep case format explicit and reviewable Golden cases will be stored as JSON with fields such as `caseId`, `query`, `expectedDocIds`, `expectedBreadcrumbs`, `expectedKeywords`, and optional `notes`. Rationale: the case file should be easy to inspect in code review and easy to extend during interviews or later refactors. ### Decision: Preserve both machine and human reports The evaluator will write JSON for automation and Markdown for review. Rationale: future changes can compare JSON, while the Markdown report is easier to use during design review and interview preparation. ## Risks / Trade-offs - Offline fixtures can drift from real runtime behavior -> add documentation that this is a baseline harness, not a live retrieval accuracy claim. - Keyword-based evidence checks are approximate -> use them only as deterministic guardrails, not as a replacement for human review. - Small golden set may underrepresent production queries -> start with 10-20 cases and expand as new RAG issues are found. - Fixture schema may not match future retrieval outputs -> normalize fixtures into a simple candidate shape and keep raw fields optional.