feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,39 @@
|
||||
# Acceptance: rag-eval-hybrid-baseline
|
||||
|
||||
## Tasks
|
||||
|
||||
All tasks in OpenSpec `tasks.md` checked, including apply-discovered 6.x quality-gate fix.
|
||||
|
||||
## 静态验证
|
||||
|
||||
- Snapshot generator path: no required `retrieval.vector-store.mode`.
|
||||
- README documents hybrid generation and offline/live split.
|
||||
|
||||
## 脚本验证
|
||||
|
||||
```text
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
|
||||
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
# Result: Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
|
||||
|
||||
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
|
||||
# exit 0
|
||||
```
|
||||
|
||||
Fixture sample meta: `searchMode=hybrid`, `kbScope=rag-eval`.
|
||||
Fallback case: `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`.
|
||||
|
||||
## 浏览器/人工
|
||||
|
||||
- 未做 UI 验证。
|
||||
|
||||
## 未验证 / 后续
|
||||
|
||||
- Dense vs hybrid dual-directory comparison report (knife-2).
|
||||
- CI wiring of offline eval as required gate (optional process).
|
||||
- Long-term calibration of hybrid PRECISE distribution under denseDistance quality.
|
||||
|
||||
## Specs
|
||||
|
||||
Main spec synced: `openspec/specs/rag-eval-offline-baseline/spec.md`.
|
||||
@@ -0,0 +1,16 @@
|
||||
# Brief: rag-eval-hybrid-baseline
|
||||
|
||||
## Background
|
||||
|
||||
Offline RAG eval (golden × fixture × key-field baseline) existed but generator/docs still used dead `retrieval.vector-store.mode=spring`. Fixtures lacked search meta and did not reflect hybrid main path.
|
||||
|
||||
## Goals (knife-1 only)
|
||||
|
||||
- Snapshot generation uses `retrieval.search.mode` (default hybrid; dense override).
|
||||
- Fixtures record `searchMode` / `kbScope`.
|
||||
- README documents hybrid-era offline vs live loop.
|
||||
- Best-effort live seed + regenerate fixtures + update baseline.
|
||||
|
||||
## Non-goals
|
||||
|
||||
Dense/hybrid dual fixture trees; golden mustNot/chunk/level hard gates; new eval frameworks.
|
||||
@@ -0,0 +1,18 @@
|
||||
# Decisions: rag-eval-hybrid-baseline(最终版)
|
||||
|
||||
## Process
|
||||
|
||||
sm-flow standard-lean: Discover → Commit → Apply → Archive.
|
||||
|
||||
## Key decisions
|
||||
|
||||
1. Replace eval generator `vector-store.mode` with `retrieval.search.mode` (default hybrid).
|
||||
2. Fixture meta: `searchMode`, `kbScope` when set.
|
||||
3. Knife-2 (dual fixtures / mustNot golden) deferred.
|
||||
4. Live refresh succeeded in apply env; baseline updated to hybrid snapshots.
|
||||
5. **Quality gate refinement (apply-found):** hybrid absolute quality for `isLowQuality` / relevance uses optional dense L2 (`denseDistance`); does not overwrite hybrid scoreLabel or RRF order. Rank mapping remains fallback when dense missing.
|
||||
|
||||
## Trade-offs
|
||||
|
||||
- Extra dense ANN on hybrid path for gate calibration (latency) vs correct filter-fallback behavior.
|
||||
- relevance_level still not a hard golden assertion (ordinal vs absolute mix).
|
||||
@@ -0,0 +1,17 @@
|
||||
# Evidence: rag-eval-hybrid-baseline
|
||||
|
||||
## Pre-change
|
||||
|
||||
- `generate_rag_lookup_snapshots.ps1` passed `-Dretrieval.vector-store.mode=spring`.
|
||||
- Fixtures had `caseId/query/retrievedAt/lookupResult` only.
|
||||
- Offline eval already supported Hit levels, recall@K, baseline diff.
|
||||
|
||||
## User decisions
|
||||
|
||||
- Scope: knife-1 only (no dual fixture dirs).
|
||||
- Acceptance: wiring required; fixture refresh best-effort (env allowed full refresh).
|
||||
|
||||
## Apply-discovered
|
||||
|
||||
- After hybrid refresh, `chat-l0-filter-fallback` failed: pure rank→quality made topSimilarity=1.0 on decoy-only filtered hits → no unfiltered retry.
|
||||
- Fix: optional `denseDistance` on hybrid hits; quality gate uses L2 when present; sort order remains RRF.
|
||||
Reference in New Issue
Block a user