1.7 KiB
1.7 KiB
Why
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
What Changes
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
- Add JSON and Markdown report output for quick review after a run.
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
Capabilities
New Capabilities
diagnosis-eval-harness: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
Modified Capabilities
- None.
Impact
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
- Affected APIs: none.
- Dependencies: relies on the evidence semantics from
evidence-trace-hardening, especiallytool_invocation,tool_trace_summary,verifier_evaluation, and evidence status conventions. - Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.