Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
2.2 KiB
2.2 KiB
rag-eval-offline-baseline Specification
Purpose
Keep the offline RAG retrieval baseline aligned with the hybrid search main path and fixture metadata needed for reproducible regression.
ADDED Requirements
Requirement: Snapshot generation SHALL use retrieval search mode
RAG lookup fixture generation SHALL configure retrieval.search.mode and SHALL NOT require retrieval.vector-store.mode for knowledge snapshot generation.
Scenario: Default hybrid generation
- WHEN the snapshot generator is invoked with default parameters
- THEN it SHALL run with
retrieval.search.mode=hybrid(or equivalent default) - AND it SHALL NOT pass
retrieval.vector-store.modeas a required generation setting
Scenario: Dense mode override for comparison runs
- WHEN the operator sets search mode to
dense - THEN fixture generation SHALL use dense retrieval for that run
Requirement: Generated fixtures SHALL record search meta
Each generated fixture file SHALL include stable identity and generation meta in addition to the lookup payload.
Scenario: Meta fields present
- WHEN a fixture is written for a golden case
- THEN the fixture SHALL contain
caseId,query,retrievedAt, andsearchMode - AND when kb scope is configured non-empty, the fixture SHOULD contain
kbScope
Requirement: Offline evaluation SHALL remain dependency-free
The offline baseline checker SHALL evaluate golden cases against fixture files without calling Milvus, embedding APIs, or starting the full application.
Scenario: Offline eval without live stack
- WHEN
eval_rag_retrieval.py(or successor) runs against cases and fixtures - THEN it SHALL produce pass/fail results using fixture contents only
Requirement: Eval documentation SHALL describe the hybrid-era loop
Project eval README SHALL document seed import, snapshot generation with search.mode, offline eval, and baseline diff, without presenting Spring vector-store mode as the knowledge main path.
Scenario: README main path
- WHEN an engineer follows the eval README happy path
- THEN the documented default generation mode SHALL be hybrid search mode