Files
2026-07-04 23:51:43 +08:00

1.7 KiB

Why

The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.

What Changes

  • Add a small fixed evaluation set for representative MVP diagnosis scenarios.
  • Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
  • Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
  • Add JSON and Markdown report output for quick review after a run.
  • Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.

Capabilities

New Capabilities

  • diagnosis-eval-harness: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.

Modified Capabilities

  • None.

Impact

  • Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
  • Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
  • Affected APIs: none.
  • Dependencies: relies on the evidence semantics from evidence-trace-hardening, especially tool_invocation, tool_trace_summary, verifier_evaluation, and evidence status conventions.
  • Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.