28 lines
1.7 KiB
Markdown
28 lines
1.7 KiB
Markdown
## Why
|
|
|
|
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
|
|
|
|
## What Changes
|
|
|
|
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
|
|
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
|
|
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
|
|
- Add JSON and Markdown report output for quick review after a run.
|
|
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
|
|
|
|
## Capabilities
|
|
|
|
### New Capabilities
|
|
- `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
|
|
|
|
### Modified Capabilities
|
|
- None.
|
|
|
|
## Impact
|
|
|
|
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
|
|
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
|
|
- Affected APIs: none.
|
|
- Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions.
|
|
- Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.
|