feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,39 @@
|
||||
# Acceptance: rag-eval-hybrid-baseline
|
||||
|
||||
## Tasks
|
||||
|
||||
All tasks in OpenSpec `tasks.md` checked, including apply-discovered 6.x quality-gate fix.
|
||||
|
||||
## 静态验证
|
||||
|
||||
- Snapshot generator path: no required `retrieval.vector-store.mode`.
|
||||
- README documents hybrid generation and offline/live split.
|
||||
|
||||
## 脚本验证
|
||||
|
||||
```text
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
|
||||
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
# Result: Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
|
||||
|
||||
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
|
||||
# exit 0
|
||||
```
|
||||
|
||||
Fixture sample meta: `searchMode=hybrid`, `kbScope=rag-eval`.
|
||||
Fallback case: `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`.
|
||||
|
||||
## 浏览器/人工
|
||||
|
||||
- 未做 UI 验证。
|
||||
|
||||
## 未验证 / 后续
|
||||
|
||||
- Dense vs hybrid dual-directory comparison report (knife-2).
|
||||
- CI wiring of offline eval as required gate (optional process).
|
||||
- Long-term calibration of hybrid PRECISE distribution under denseDistance quality.
|
||||
|
||||
## Specs
|
||||
|
||||
Main spec synced: `openspec/specs/rag-eval-offline-baseline/spec.md`.
|
||||
Reference in New Issue
Block a user