63 lines
3.6 KiB
Markdown
63 lines
3.6 KiB
Markdown
## Context
|
|
|
|
The repository already has diagnosis-level evaluation specs and trace observability, but the RAG refactor plan needs a narrower retrieval baseline. The upcoming changes will alter L0 responsibilities, query augmentation, evidence post-processing, and eventually the vector store implementation. Those changes need a fixed set of retrieval cases and deterministic scoring before production retrieval behavior changes.
|
|
|
|
The first baseline must be offline. It should not require MySQL, Milvus, Redis, LLM calls, or a running Spring Boot application. It can evaluate saved retrieval result fixtures that represent the current behavior and produce reports that future changes can compare against.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals:**
|
|
|
|
- Add a small fixed golden query set for RAG retrieval.
|
|
- Evaluate retrieval fixtures against expected documents, breadcrumbs, evidence keywords, and hit levels.
|
|
- Produce JSON and Markdown baseline reports.
|
|
- Document how to regenerate the reports.
|
|
- Keep the evaluator simple enough to run from the repository with Python.
|
|
|
|
**Non-Goals:**
|
|
|
|
- Do not change `lookup_knowledge`, `VectorSearchService`, Milvus schema, L0 matching, or Agent prompts.
|
|
- Do not require live services.
|
|
- Do not implement Spring AI VectorStore migration in this change.
|
|
- Do not implement RRF, BM25, rerank, or evidence packing in this change.
|
|
|
|
## Decisions
|
|
|
|
### Decision: Use offline retrieval fixtures first
|
|
|
|
The evaluator will read saved retrieval fixtures rather than calling the live application.
|
|
|
|
Rationale: the first change should establish a stable measurement surface before the RAG internals change. Live retrieval depends on embeddings, Milvus state, and service configuration, which makes it a poor first baseline.
|
|
|
|
Alternative considered: call `SearchController` or `lookup_knowledge` directly. That is useful later, but it would require a running app and seeded knowledge base.
|
|
|
|
### Decision: Score by hit level, not exact chunk id only
|
|
|
|
The evaluator will classify each case as:
|
|
|
|
- `strong`: expected document plus expected breadcrumb or key evidence coverage.
|
|
- `medium`: expected document found, but breadcrumb or evidence coverage is incomplete.
|
|
- `weak`: related evidence is present but the expected document is missing.
|
|
- `miss`: no expected document or expected evidence is found.
|
|
|
|
Rationale: chunk indexes can change after splitter changes, so exact chunk-only scoring would make later refactors look worse even when evidence quality is preserved.
|
|
|
|
### Decision: Keep case format explicit and reviewable
|
|
|
|
Golden cases will be stored as JSON with fields such as `caseId`, `query`, `expectedDocIds`, `expectedBreadcrumbs`, `expectedKeywords`, and optional `notes`.
|
|
|
|
Rationale: the case file should be easy to inspect in code review and easy to extend during interviews or later refactors.
|
|
|
|
### Decision: Preserve both machine and human reports
|
|
|
|
The evaluator will write JSON for automation and Markdown for review.
|
|
|
|
Rationale: future changes can compare JSON, while the Markdown report is easier to use during design review and interview preparation.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
- Offline fixtures can drift from real runtime behavior -> add documentation that this is a baseline harness, not a live retrieval accuracy claim.
|
|
- Keyword-based evidence checks are approximate -> use them only as deterministic guardrails, not as a replacement for human review.
|
|
- Small golden set may underrepresent production queries -> start with 10-20 cases and expand as new RAG issues are found.
|
|
- Fixture schema may not match future retrieval outputs -> normalize fixtures into a simple candidate shape and keep raw fields optional.
|