Add diagnosis eval harness
This commit is contained in:
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
|
||||
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
|
||||
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
|
||||
- Add JSON and Markdown report output for quick review after a run.
|
||||
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
|
||||
|
||||
### Modified Capabilities
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
|
||||
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
|
||||
- Affected APIs: none.
|
||||
- Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions.
|
||||
- Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.
|
||||
Reference in New Issue
Block a user