## 1. Golden Cases - [x] 1.1 Create retrieval evaluation directory structure. - [x] 1.2 Add fixed golden retrieval cases covering Chat and AIOps retrieval scenarios. - [x] 1.3 Add matching offline retrieval fixtures for every golden case. ## 2. Evaluator - [x] 2.1 Implement an offline retrieval evaluator script. - [x] 2.2 Support hit-level classification and first expected document rank. - [x] 2.3 Support JSON and Markdown report output. ## 3. Baseline Report - [x] 3.1 Generate the baseline JSON report from the fixed cases and fixtures. - [x] 3.2 Generate the baseline Markdown report from the fixed cases and fixtures. ## 4. Documentation - [x] 4.1 Document the retrieval baseline purpose, file layout, and regeneration command. - [x] 4.2 Link the retrieval baseline from the RAG refactor issue or related MVP documentation. ## 5. Verification - [x] 5.1 Run the evaluator successfully against the fixed baseline cases. - [x] 5.2 Run OpenSpec status/validation for the change and confirm tasks are complete.