## Why The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo. ## What Changes - Add a small fixed evaluation set for representative MVP diagnosis scenarios. - Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior. - Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration. - Add JSON and Markdown report output for quick review after a run. - Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration. ## Capabilities ### New Capabilities - `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks. ### Modified Capabilities - None. ## Impact - Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records. - Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path. - Affected APIs: none. - Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions. - Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.