Files
SuperBizAgent-java/openspec/changes/archive/2026-07-04-diagnosis-eval-harness/design.md
T
2026-07-04 23:51:43 +08:00

3.3 KiB

Context

The project now has the pieces needed for trace-based evaluation:

  • diagnosis_session stores final answer, status, duration, counts, feedback, and self_evaluation
  • agent_step stores ordered agent execution records
  • tool_invocation stores evidence tool calls with normalized evidence semantics
  • GET /api/diagnosis/{sessionId}/trace can aggregate one diagnosis trace for demo review
  • evidence-trace-hardening defined stable supported, no_evidence, deduped, and failed semantics

P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.

Goals / Non-Goals

Goals:

  • Define fixed MVP diagnosis cases with expected evidence and verdict rules.
  • Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
  • Produce JSON and Markdown reports with pass/fail status and key metrics.
  • Keep the first version usable without a real LLM by allowing fixture trace inputs.
  • Leave room for a later runtime mode that queries the trace API after a demo run.

Non-Goals:

  • No LLM-as-judge in this slice.
  • No automatic prompt optimization.
  • No new production API.
  • No change to chat, verifier, retrieval, upload, or feedback behavior.
  • No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.

Decisions

Decision Choice Alternative Considered Rationale
Evaluation source Start with fixture / persisted trace JSON input Always run live /api/chat first Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability.
Judging strategy Rule-based trace validation LLM-as-judge The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring.
Case format Static JSON/YAML case definitions Hard-coded Java tests only Case files are easier to inspect and explain in interviews.
Report format JSON plus Markdown Console-only output JSON supports automation; Markdown supports quick human review.
Metrics Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage Full semantic correctness These metrics are available from existing trace data and align with the MVP's observable contract.

Risks / Trade-offs

  • [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
  • [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
  • [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
  • [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.

Migration Plan

  • No deployment migration is required.
  • The harness is additive and can be run locally as a test or script.
  • Rollback is deleting the eval case files, runner, and report docs.

Open Questions

  • Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?