3.3 KiB
3.3 KiB
Context
The project now has the pieces needed for trace-based evaluation:
diagnosis_sessionstores final answer, status, duration, counts, feedback, andself_evaluationagent_stepstores ordered agent execution recordstool_invocationstores evidence tool calls with normalized evidence semanticsGET /api/diagnosis/{sessionId}/tracecan aggregate one diagnosis trace for demo reviewevidence-trace-hardeningdefined stablesupported,no_evidence,deduped, andfailedsemantics
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
Goals / Non-Goals
Goals:
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
- Produce JSON and Markdown reports with pass/fail status and key metrics.
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
- Leave room for a later runtime mode that queries the trace API after a demo run.
Non-Goals:
- No LLM-as-judge in this slice.
- No automatic prompt optimization.
- No new production API.
- No change to chat, verifier, retrieval, upload, or feedback behavior.
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live /api/chat first |
Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
Risks / Trade-offs
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
Migration Plan
- No deployment migration is required.
- The harness is additive and can be run locally as a test or script.
- Rollback is deleting the eval case files, runner, and report docs.
Open Questions
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?