Files
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

112 lines
5.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Decisions — rag-eval-hybrid-baseline
## sm-flow meta
- **Checkpoint**: Discover (in progress)
- **Scale**: standard (lean) — eval harness alignment, multi-file, low prod risk
- **Capability**: sm-flow built-in; openspec CLI `new change`; grill fallback (no external grill-with-docs runner)
- **Slug**: `rag-eval-hybrid-baseline`
- **Path**: `openspec/changes/rag-eval-hybrid-baseline/`
## Clarify summary
| Item | Content |
|------|---------|
| Problem | Offline eval model OK but wiring/fixtures/docs pre-hybrid; cannot gate current main path |
| Goal | Knife-1: hybrid generator + meta + docs (+ refresh). Optional knife-2: dense/hybrid dual fixtures |
| Touch | `scripts/generate_rag_lookup_snapshots.ps1`, snapshot test, eval README, fixtures/baseline, maybe `eval_rag_retrieval.py` |
| Non-goals | New framework, LLM judge, prod retrieval redesign |
## Context summary
| Source | Conclusion | Into OpenSpec |
|--------|------------|---------------|
| Conversation design | Golden×fixture×key fields; not full JSON diff | Yes |
| Current eval audit | ~70% aligned; dead spring mode; old fixtures | Yes |
| `rag-quality-score-unify` | hybrid quality rank-based; don't hard-lock PRECISE | Yes |
| `eval/rag-retrieval/README` | seed + kb_scope good; generator props stale | Yes |
| Generator ps1 | `VectorStoreMode=spring` → must replace with search.mode | Yes |
**index**: hit rag-quality-score-unify / bm25-hybrid / chunk-identity archives.
## Question pool (grill)
| ID | Dim | Mode | Question | Status |
|----|-----|------|----------|--------|
| Q1 | 边界 | user-interview | 本 change 范围:仅第一刀,还是第一刀+第二刀(dense/hybrid 双目录对照)? | **已确认:仅第一刀** |
| Q2 | 验收 | user-interview | Apply 时若本机无法连 embedding/Milvus 重刷 fixture,是否允许「只交接线+文档,fixture 刷新记未验证」? | **已确认:接线优先,刷新可未验证** |
| Q3 | 术语 | evidence-driven | 生成器是否仍传 `vector-store.mode`? | **已查证:是** |
| Q4 | 验收 | evidence-driven | 离线脚本是否已支持 Hit 分层与 baseline diff? | **已查证:是** |
| Q5 | 边界 | evidence-driven | Golden 是否已有 mustNot/chunk key? | **已查证:无** |
### Q3–Q5 evidence
- `scripts/generate_rag_lookup_snapshots.ps1`: `-Dretrieval.vector-store.mode=$VectorStoreMode` default spring.
- `eval_rag_retrieval.py`: strong/medium/weak/miss, recall, firstExpectedRank, compare-to diff.
- `golden-cases.json`: doc/source/keyword/attempt/fallback; no mustNot, no evidenceKey expectations.
### Q1 用户确认
- **选择**: 仅第一刀(推荐)
- **含义**: 生成器 `search.mode=hybrid`;fixture meta;README;能连环境则重刷。不做 dense/hybrid 双目录对照。
### Q2 用户确认
- **选择**: 接线优先,刷新可记未验证
- **含义**: 脚本/测试/README/meta 必交付;fixtures/baseline 能刷则刷,不能刷则 acceptance 记未验证与补跑命令,不阻塞 apply 完成。
---
## Discover status
- [x] clarify
- [x] context
- [x] propose (`proposal.md`)
- [x] grill (Q1–Q5 closed)
**Discover checkpoint: 完成。**
---
## Commit checkpoint
- **Capability**: sm-flow built-in specify/audit/commit; openspec status 4/4
- **Cross-artifact**: brief/proposal → design → specs → tasks aligned (knife-1 only; Q1/Q2 reflected)
- **Audit**: eval-only; L1 impact; no Agent ACI; live refresh best-effort per Q2
- **Gate**: `.committed` written
**Commit checkpoint: 完成。Committed OpenSpec 就绪。**
**Next**: wait for explicit **Apply** authorization (e.g.「开始 apply / 实现」).
---
## Apply checkpoint
- **Capability**: openspec-apply-change + Committed tasks
- **Authorization**: user「实现」
### Delivered (knife-1)
- `scripts/generate_rag_lookup_snapshots.ps1`: `-SearchMode hybrid|dense`, no `vector-store.mode`
- `RagLookupSnapshotGeneratorTest`: `@DynamicPropertySource` for search.mode/kb-scope; fixture meta `searchMode`/`kbScope`
- `eval/rag-retrieval/README.md` hybrid-era docs
- Live refresh: seed OK → hybrid generate OK → offline **7/7 pass**, baseline updated
### Apply-discovered regression + fix
- **Issue**: pure rank→quality made hybrid rank1 always quality=1.0 → `isLowQuality` never true → L0 filter fallback case stuck on decoy (`FILTERED_VECTOR`).
- **Fix**: hybrid still sorts by RRF order; optional parallel dense L2 stored as `denseDistance`; `toQualityScore(hybrid)` uses dense L2 for absolute gates when present (rank fallback if missing). Does **not** restore scoreLabel overwrite / boost re-rank.
- **Verify**: fixture `chat-l0-filter-fallback` → `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`; offline passRate=1.0
### Commands run
```text
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
```
**Apply checkpoint: 完成。** Ready for Archive when user requests.