29 lines
1.1 KiB
Markdown
29 lines
1.1 KiB
Markdown
# Diagnosis Eval Harness Decisions
|
|
|
|
## Clarify
|
|
|
|
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
|
|
- Slug: `diagnosis-eval-harness`
|
|
- Devflow scale: standard-light
|
|
|
|
## Context
|
|
|
|
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
|
|
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
|
|
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
|
|
|
|
## Key Decisions
|
|
|
|
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
|
|
- Reason: The first regression signal should be deterministic and tied to trace contracts.
|
|
|
|
- Decision: Support offline fixture traces first.
|
|
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
|
|
|
|
- Decision: Output both JSON and Markdown.
|
|
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
|
|
|
|
## Open Questions
|
|
|
|
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
|