Files
SuperBizAgent-java/openspec/changes/archive/2026-07-28-rag-eval-hybrid-baseline/decisions.md
T
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

5.2 KiB
Raw Blame History

Decisions — rag-eval-hybrid-baseline

sm-flow meta

  • Checkpoint: Discover (in progress)
  • Scale: standard (lean) — eval harness alignment, multi-file, low prod risk
  • Capability: sm-flow built-in; openspec CLI new change; grill fallback (no external grill-with-docs runner)
  • Slug: rag-eval-hybrid-baseline
  • Path: openspec/changes/rag-eval-hybrid-baseline/

Clarify summary

Item Content
Problem Offline eval model OK but wiring/fixtures/docs pre-hybrid; cannot gate current main path
Goal Knife-1: hybrid generator + meta + docs (+ refresh). Optional knife-2: dense/hybrid dual fixtures
Touch scripts/generate_rag_lookup_snapshots.ps1, snapshot test, eval README, fixtures/baseline, maybe eval_rag_retrieval.py
Non-goals New framework, LLM judge, prod retrieval redesign

Context summary

Source Conclusion Into OpenSpec
Conversation design Golden×fixture×key fields; not full JSON diff Yes
Current eval audit ~70% aligned; dead spring mode; old fixtures Yes
rag-quality-score-unify hybrid quality rank-based; don't hard-lock PRECISE Yes
eval/rag-retrieval/README seed + kb_scope good; generator props stale Yes
Generator ps1 VectorStoreMode=spring → must replace with search.mode Yes

index: hit rag-quality-score-unify / bm25-hybrid / chunk-identity archives.

Question pool (grill)

ID Dim Mode Question Status
Q1 边界 user-interview 本 change 范围:仅第一刀,还是第一刀+第二刀(dense/hybrid 双目录对照)? 已确认:仅第一刀
Q2 验收 user-interview Apply 时若本机无法连 embedding/Milvus 重刷 fixture,是否允许「只交接线+文档,fixture 刷新记未验证」? 已确认:接线优先,刷新可未验证
Q3 术语 evidence-driven 生成器是否仍传 vector-store.mode? 已查证:是
Q4 验收 evidence-driven 离线脚本是否已支持 Hit 分层与 baseline diff? 已查证:是
Q5 边界 evidence-driven Golden 是否已有 mustNot/chunk key? 已查证:无

Q3–Q5 evidence

  • scripts/generate_rag_lookup_snapshots.ps1: -Dretrieval.vector-store.mode=$VectorStoreMode default spring.
  • eval_rag_retrieval.py: strong/medium/weak/miss, recall, firstExpectedRank, compare-to diff.
  • golden-cases.json: doc/source/keyword/attempt/fallback; no mustNot, no evidenceKey expectations.

Q1 用户确认

  • 选择: 仅第一刀(推荐)
  • 含义: 生成器 search.mode=hybrid;fixture meta;README;能连环境则重刷。不做 dense/hybrid 双目录对照。

Q2 用户确认

  • 选择: 接线优先,刷新可记未验证
  • 含义: 脚本/测试/README/meta 必交付;fixtures/baseline 能刷则刷,不能刷则 acceptance 记未验证与补跑命令,不阻塞 apply 完成。

Discover status

  • clarify
  • context
  • propose (proposal.md)
  • grill (Q1–Q5 closed)

Discover checkpoint: 完成。


Commit checkpoint

  • Capability: sm-flow built-in specify/audit/commit; openspec status 4/4
  • Cross-artifact: brief/proposal → design → specs → tasks aligned (knife-1 only; Q1/Q2 reflected)
  • Audit: eval-only; L1 impact; no Agent ACI; live refresh best-effort per Q2
  • Gate: .committed written

Commit checkpoint: 完成。Committed OpenSpec 就绪。

Next: wait for explicit Apply authorization (e.g.「开始 apply / 实现」).


Apply checkpoint

  • Capability: openspec-apply-change + Committed tasks
  • Authorization: user「实现」

Delivered (knife-1)

  • scripts/generate_rag_lookup_snapshots.ps1: -SearchMode hybrid|dense, no vector-store.mode
  • RagLookupSnapshotGeneratorTest: @DynamicPropertySource for search.mode/kb-scope; fixture meta searchMode/kbScope
  • eval/rag-retrieval/README.md hybrid-era docs
  • Live refresh: seed OK → hybrid generate OK → offline 7/7 pass, baseline updated

Apply-discovered regression + fix

  • Issue: pure rank→quality made hybrid rank1 always quality=1.0 → isLowQuality never true → L0 filter fallback case stuck on decoy (FILTERED_VECTOR).
  • Fix: hybrid still sorts by RRF order; optional parallel dense L2 stored as denseDistance; toQualityScore(hybrid) uses dense L2 for absolute gates when present (rank fallback if missing). Does not restore scoreLabel overwrite / boost re-rank.
  • Verify: fixture chat-l0-filter-fallback → UNFILTERED_VECTOR_RETRY + filtered_vector_low_quality; offline passRate=1.0

Commands run

.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test

Apply checkpoint: 完成。 Ready for Archive when user requests.