feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,111 @@
|
||||
# Decisions — rag-eval-hybrid-baseline
|
||||
|
||||
## sm-flow meta
|
||||
|
||||
- **Checkpoint**: Discover (in progress)
|
||||
- **Scale**: standard (lean) — eval harness alignment, multi-file, low prod risk
|
||||
- **Capability**: sm-flow built-in; openspec CLI `new change`; grill fallback (no external grill-with-docs runner)
|
||||
- **Slug**: `rag-eval-hybrid-baseline`
|
||||
- **Path**: `openspec/changes/rag-eval-hybrid-baseline/`
|
||||
|
||||
## Clarify summary
|
||||
|
||||
| Item | Content |
|
||||
|------|---------|
|
||||
| Problem | Offline eval model OK but wiring/fixtures/docs pre-hybrid; cannot gate current main path |
|
||||
| Goal | Knife-1: hybrid generator + meta + docs (+ refresh). Optional knife-2: dense/hybrid dual fixtures |
|
||||
| Touch | `scripts/generate_rag_lookup_snapshots.ps1`, snapshot test, eval README, fixtures/baseline, maybe `eval_rag_retrieval.py` |
|
||||
| Non-goals | New framework, LLM judge, prod retrieval redesign |
|
||||
|
||||
## Context summary
|
||||
|
||||
| Source | Conclusion | Into OpenSpec |
|
||||
|--------|------------|---------------|
|
||||
| Conversation design | Golden×fixture×key fields; not full JSON diff | Yes |
|
||||
| Current eval audit | ~70% aligned; dead spring mode; old fixtures | Yes |
|
||||
| `rag-quality-score-unify` | hybrid quality rank-based; don't hard-lock PRECISE | Yes |
|
||||
| `eval/rag-retrieval/README` | seed + kb_scope good; generator props stale | Yes |
|
||||
| Generator ps1 | `VectorStoreMode=spring` → must replace with search.mode | Yes |
|
||||
|
||||
**index**: hit rag-quality-score-unify / bm25-hybrid / chunk-identity archives.
|
||||
|
||||
## Question pool (grill)
|
||||
|
||||
| ID | Dim | Mode | Question | Status |
|
||||
|----|-----|------|----------|--------|
|
||||
| Q1 | 边界 | user-interview | 本 change 范围:仅第一刀,还是第一刀+第二刀(dense/hybrid 双目录对照)? | **已确认:仅第一刀** |
|
||||
| Q2 | 验收 | user-interview | Apply 时若本机无法连 embedding/Milvus 重刷 fixture,是否允许「只交接线+文档,fixture 刷新记未验证」? | **已确认:接线优先,刷新可未验证** |
|
||||
| Q3 | 术语 | evidence-driven | 生成器是否仍传 `vector-store.mode`? | **已查证:是** |
|
||||
| Q4 | 验收 | evidence-driven | 离线脚本是否已支持 Hit 分层与 baseline diff? | **已查证:是** |
|
||||
| Q5 | 边界 | evidence-driven | Golden 是否已有 mustNot/chunk key? | **已查证:无** |
|
||||
|
||||
### Q3–Q5 evidence
|
||||
|
||||
- `scripts/generate_rag_lookup_snapshots.ps1`: `-Dretrieval.vector-store.mode=$VectorStoreMode` default spring.
|
||||
- `eval_rag_retrieval.py`: strong/medium/weak/miss, recall, firstExpectedRank, compare-to diff.
|
||||
- `golden-cases.json`: doc/source/keyword/attempt/fallback; no mustNot, no evidenceKey expectations.
|
||||
|
||||
### Q1 用户确认
|
||||
|
||||
- **选择**: 仅第一刀(推荐)
|
||||
- **含义**: 生成器 `search.mode=hybrid`;fixture meta;README;能连环境则重刷。不做 dense/hybrid 双目录对照。
|
||||
|
||||
### Q2 用户确认
|
||||
|
||||
- **选择**: 接线优先,刷新可记未验证
|
||||
- **含义**: 脚本/测试/README/meta 必交付;fixtures/baseline 能刷则刷,不能刷则 acceptance 记未验证与补跑命令,不阻塞 apply 完成。
|
||||
|
||||
---
|
||||
|
||||
## Discover status
|
||||
|
||||
- [x] clarify
|
||||
- [x] context
|
||||
- [x] propose (`proposal.md`)
|
||||
- [x] grill (Q1–Q5 closed)
|
||||
|
||||
**Discover checkpoint: 完成。**
|
||||
|
||||
---
|
||||
|
||||
## Commit checkpoint
|
||||
|
||||
- **Capability**: sm-flow built-in specify/audit/commit; openspec status 4/4
|
||||
- **Cross-artifact**: brief/proposal → design → specs → tasks aligned (knife-1 only; Q1/Q2 reflected)
|
||||
- **Audit**: eval-only; L1 impact; no Agent ACI; live refresh best-effort per Q2
|
||||
- **Gate**: `.committed` written
|
||||
|
||||
**Commit checkpoint: 完成。Committed OpenSpec 就绪。**
|
||||
|
||||
**Next**: wait for explicit **Apply** authorization (e.g.「开始 apply / 实现」).
|
||||
|
||||
---
|
||||
|
||||
## Apply checkpoint
|
||||
|
||||
- **Capability**: openspec-apply-change + Committed tasks
|
||||
- **Authorization**: user「实现」
|
||||
|
||||
### Delivered (knife-1)
|
||||
|
||||
- `scripts/generate_rag_lookup_snapshots.ps1`: `-SearchMode hybrid|dense`, no `vector-store.mode`
|
||||
- `RagLookupSnapshotGeneratorTest`: `@DynamicPropertySource` for search.mode/kb-scope; fixture meta `searchMode`/`kbScope`
|
||||
- `eval/rag-retrieval/README.md` hybrid-era docs
|
||||
- Live refresh: seed OK → hybrid generate OK → offline **7/7 pass**, baseline updated
|
||||
|
||||
### Apply-discovered regression + fix
|
||||
|
||||
- **Issue**: pure rank→quality made hybrid rank1 always quality=1.0 → `isLowQuality` never true → L0 filter fallback case stuck on decoy (`FILTERED_VECTOR`).
|
||||
- **Fix**: hybrid still sorts by RRF order; optional parallel dense L2 stored as `denseDistance`; `toQualityScore(hybrid)` uses dense L2 for absolute gates when present (rank fallback if missing). Does **not** restore scoreLabel overwrite / boost re-rank.
|
||||
- **Verify**: fixture `chat-l0-filter-fallback` → `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`; offline passRate=1.0
|
||||
|
||||
### Commands run
|
||||
|
||||
```text
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
|
||||
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
|
||||
```
|
||||
|
||||
**Apply checkpoint: 完成。** Ready for Archive when user requests.
|
||||
Reference in New Issue
Block a user