Add diagnosis eval harness

This commit is contained in:
aruo
2026-07-04 23:51:43 +08:00
parent dc6cd32a67
commit ca5c61fabf
24 changed files with 1251 additions and 2 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,54 @@
## Context
The project now has the pieces needed for trace-based evaluation:
- `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation`
- `agent_step` stores ordered agent execution records
- `tool_invocation` stores evidence tool calls with normalized evidence semantics
- `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review
- `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
## Goals / Non-Goals
**Goals:**
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
- Produce JSON and Markdown reports with pass/fail status and key metrics.
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
- Leave room for a later runtime mode that queries the trace API after a demo run.
**Non-Goals:**
- No LLM-as-judge in this slice.
- No automatic prompt optimization.
- No new production API.
- No change to chat, verifier, retrieval, upload, or feedback behavior.
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
## Risks / Trade-offs
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
## Migration Plan
- No deployment migration is required.
- The harness is additive and can be run locally as a test or script.
- Rollback is deleting the eval case files, runner, and report docs.
## Open Questions
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?
@@ -0,0 +1,27 @@
## Why
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
## What Changes
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
- Add JSON and Markdown report output for quick review after a run.
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
## Capabilities
### New Capabilities
- `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
### Modified Capabilities
- None.
## Impact
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
- Affected APIs: none.
- Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions.
- Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.
@@ -0,0 +1,59 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
#### Scenario: Case definition includes expected evidence
- **WHEN** an evaluation case is defined
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
#### Scenario: Case definition can express forbidden behavior
- **WHEN** a case has known unsafe behavior
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
### Requirement: Evaluation harness SHALL validate diagnosis traces
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
#### Scenario: Evidence coverage validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
#### Scenario: Verifier evaluation validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
- **AND** the verdict SHALL be one of the case's allowed verdicts
#### Scenario: Answer keyword validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
#### Scenario: Degraded output validation
- **WHEN** a trace verdict is `REJECT`
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
### Requirement: Evaluation harness SHALL report quality and cost signals
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
#### Scenario: JSON report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
#### Scenario: Markdown report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
#### Scenario: Aggregate metrics
- **WHEN** multiple cases are evaluated
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
### Requirement: Evaluation harness SHALL support offline fixture mode
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
#### Scenario: Fixture trace evaluation
- **WHEN** the evaluator is run against a directory of trace fixture files
- **THEN** it SHALL evaluate each trace file against its matching case definition
- **AND** it SHALL not require a running application service
#### Scenario: Missing fixture is reported clearly
- **WHEN** a case has no matching trace fixture
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
@@ -0,0 +1,25 @@
## 1. Case Definitions
- [x] 1.1 Add evaluation case definition format for fixed MVP diagnosis scenarios.
- [x] 1.2 Add the first 5 case definitions: payment timeout, MySQL pool exhausted, Redis timeout, slow response, and JVM memory risk.
- [x] 1.3 Document the meaning of expected keywords, required evidence tools, allowed verdicts, and forbidden behavior.
## 2. Trace Fixtures
- [x] 2.1 Add fixture trace schema or DTOs that match `DiagnosisTraceResponse` enough for offline evaluation.
- [x] 2.2 Add at least one representative trace fixture for a passing case.
- [x] 2.3 Add at least one fixture covering low-confidence or degraded behavior.
## 3. Evaluator
- [x] 3.1 Implement trace validation rules for evidence coverage, verifier verdict, answer keyword coverage, and degraded-output contract.
- [x] 3.2 Implement aggregate metrics: pass rate, verdict distribution, average tool-call count, and average duration.
- [x] 3.3 Implement JSON report output.
- [x] 3.4 Implement Markdown report output.
## 4. Tests And Documentation
- [x] 4.1 Add focused offline tests for the evaluator.
- [x] 4.2 Add run instructions under `mvp/demo` or `mvp/notes`.
- [x] 4.3 Run targeted tests for the evaluator.
- [x] 4.4 Run compile verification.
@@ -0,0 +1,64 @@
# diagnosis-eval-harness Specification
## Purpose
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
## Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
#### Scenario: Case definition includes expected evidence
- **WHEN** an evaluation case is defined
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
#### Scenario: Case definition can express forbidden behavior
- **WHEN** a case has known unsafe behavior
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
### Requirement: Evaluation harness SHALL validate diagnosis traces
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
#### Scenario: Evidence coverage validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
#### Scenario: Verifier evaluation validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
- **AND** the verdict SHALL be one of the case's allowed verdicts
#### Scenario: Answer keyword validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
#### Scenario: Degraded output validation
- **WHEN** a trace verdict is `REJECT`
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
### Requirement: Evaluation harness SHALL report quality and cost signals
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
#### Scenario: JSON report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
#### Scenario: Markdown report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
#### Scenario: Aggregate metrics
- **WHEN** multiple cases are evaluated
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
### Requirement: Evaluation harness SHALL support offline fixture mode
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
#### Scenario: Fixture trace evaluation
- **WHEN** the evaluator is run against a directory of trace fixture files
- **THEN** it SHALL evaluate each trace file against its matching case definition
- **AND** it SHALL not require a running application service
#### Scenario: Missing fixture is reported clearly
- **WHEN** a case has no matching trace fixture
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason