Files
SuperBizAgent-java/openspec/changes/archive/2026-07-04-diagnosis-eval-harness/design.md
T
2026-07-04 23:51:43 +08:00

55 lines
3.3 KiB
Markdown

## Context
The project now has the pieces needed for trace-based evaluation:
- `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation`
- `agent_step` stores ordered agent execution records
- `tool_invocation` stores evidence tool calls with normalized evidence semantics
- `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review
- `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
## Goals / Non-Goals
**Goals:**
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
- Produce JSON and Markdown reports with pass/fail status and key metrics.
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
- Leave room for a later runtime mode that queries the trace API after a demo run.
**Non-Goals:**
- No LLM-as-judge in this slice.
- No automatic prompt optimization.
- No new production API.
- No change to chat, verifier, retrieval, upload, or feedback behavior.
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
## Risks / Trade-offs
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
## Migration Plan
- No deployment migration is required.
- The harness is additive and can be run locally as a test or script.
- Rollback is deleting the eval case files, runner, and report docs.
## Open Questions
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?