Add diagnosis eval harness
This commit is contained in:
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The project now has the pieces needed for trace-based evaluation:
|
||||
|
||||
- `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation`
|
||||
- `agent_step` stores ordered agent execution records
|
||||
- `tool_invocation` stores evidence tool calls with normalized evidence semantics
|
||||
- `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review
|
||||
- `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics
|
||||
|
||||
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
|
||||
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
|
||||
- Produce JSON and Markdown reports with pass/fail status and key metrics.
|
||||
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
|
||||
- Leave room for a later runtime mode that queries the trace API after a demo run.
|
||||
|
||||
**Non-Goals:**
|
||||
- No LLM-as-judge in this slice.
|
||||
- No automatic prompt optimization.
|
||||
- No new production API.
|
||||
- No change to chat, verifier, retrieval, upload, or feedback behavior.
|
||||
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
|
||||
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
|
||||
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
|
||||
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
|
||||
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
|
||||
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
|
||||
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
|
||||
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required.
|
||||
- The harness is additive and can be run locally as a test or script.
|
||||
- Rollback is deleting the eval case files, runner, and report docs.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?
|
||||
Reference in New Issue
Block a user