Add diagnosis eval harness
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The project now has the pieces needed for trace-based evaluation:
|
||||
|
||||
- `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation`
|
||||
- `agent_step` stores ordered agent execution records
|
||||
- `tool_invocation` stores evidence tool calls with normalized evidence semantics
|
||||
- `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review
|
||||
- `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics
|
||||
|
||||
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
|
||||
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
|
||||
- Produce JSON and Markdown reports with pass/fail status and key metrics.
|
||||
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
|
||||
- Leave room for a later runtime mode that queries the trace API after a demo run.
|
||||
|
||||
**Non-Goals:**
|
||||
- No LLM-as-judge in this slice.
|
||||
- No automatic prompt optimization.
|
||||
- No new production API.
|
||||
- No change to chat, verifier, retrieval, upload, or feedback behavior.
|
||||
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
|
||||
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
|
||||
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
|
||||
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
|
||||
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
|
||||
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
|
||||
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
|
||||
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required.
|
||||
- The harness is additive and can be run locally as a test or script.
|
||||
- Rollback is deleting the eval case files, runner, and report docs.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
|
||||
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
|
||||
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
|
||||
- Add JSON and Markdown report output for quick review after a run.
|
||||
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
|
||||
|
||||
### Modified Capabilities
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
|
||||
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
|
||||
- Affected APIs: none.
|
||||
- Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions.
|
||||
- Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
|
||||
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
|
||||
|
||||
#### Scenario: Case definition includes expected evidence
|
||||
- **WHEN** an evaluation case is defined
|
||||
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
|
||||
|
||||
#### Scenario: Case definition can express forbidden behavior
|
||||
- **WHEN** a case has known unsafe behavior
|
||||
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate diagnosis traces
|
||||
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
|
||||
|
||||
#### Scenario: Evidence coverage validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
|
||||
|
||||
#### Scenario: Verifier evaluation validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
|
||||
- **AND** the verdict SHALL be one of the case's allowed verdicts
|
||||
|
||||
#### Scenario: Answer keyword validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
|
||||
|
||||
#### Scenario: Degraded output validation
|
||||
- **WHEN** a trace verdict is `REJECT`
|
||||
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
|
||||
|
||||
### Requirement: Evaluation harness SHALL report quality and cost signals
|
||||
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
|
||||
|
||||
#### Scenario: JSON report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
|
||||
|
||||
#### Scenario: Markdown report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
|
||||
|
||||
#### Scenario: Aggregate metrics
|
||||
- **WHEN** multiple cases are evaluated
|
||||
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
|
||||
|
||||
### Requirement: Evaluation harness SHALL support offline fixture mode
|
||||
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
|
||||
|
||||
#### Scenario: Fixture trace evaluation
|
||||
- **WHEN** the evaluator is run against a directory of trace fixture files
|
||||
- **THEN** it SHALL evaluate each trace file against its matching case definition
|
||||
- **AND** it SHALL not require a running application service
|
||||
|
||||
#### Scenario: Missing fixture is reported clearly
|
||||
- **WHEN** a case has no matching trace fixture
|
||||
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
|
||||
@@ -0,0 +1,25 @@
|
||||
## 1. Case Definitions
|
||||
|
||||
- [x] 1.1 Add evaluation case definition format for fixed MVP diagnosis scenarios.
|
||||
- [x] 1.2 Add the first 5 case definitions: payment timeout, MySQL pool exhausted, Redis timeout, slow response, and JVM memory risk.
|
||||
- [x] 1.3 Document the meaning of expected keywords, required evidence tools, allowed verdicts, and forbidden behavior.
|
||||
|
||||
## 2. Trace Fixtures
|
||||
|
||||
- [x] 2.1 Add fixture trace schema or DTOs that match `DiagnosisTraceResponse` enough for offline evaluation.
|
||||
- [x] 2.2 Add at least one representative trace fixture for a passing case.
|
||||
- [x] 2.3 Add at least one fixture covering low-confidence or degraded behavior.
|
||||
|
||||
## 3. Evaluator
|
||||
|
||||
- [x] 3.1 Implement trace validation rules for evidence coverage, verifier verdict, answer keyword coverage, and degraded-output contract.
|
||||
- [x] 3.2 Implement aggregate metrics: pass rate, verdict distribution, average tool-call count, and average duration.
|
||||
- [x] 3.3 Implement JSON report output.
|
||||
- [x] 3.4 Implement Markdown report output.
|
||||
|
||||
## 4. Tests And Documentation
|
||||
|
||||
- [x] 4.1 Add focused offline tests for the evaluator.
|
||||
- [x] 4.2 Add run instructions under `mvp/demo` or `mvp/notes`.
|
||||
- [x] 4.3 Run targeted tests for the evaluator.
|
||||
- [x] 4.4 Run compile verification.
|
||||
@@ -0,0 +1,64 @@
|
||||
# diagnosis-eval-harness Specification
|
||||
|
||||
## Purpose
|
||||
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
|
||||
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
|
||||
|
||||
#### Scenario: Case definition includes expected evidence
|
||||
- **WHEN** an evaluation case is defined
|
||||
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
|
||||
|
||||
#### Scenario: Case definition can express forbidden behavior
|
||||
- **WHEN** a case has known unsafe behavior
|
||||
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate diagnosis traces
|
||||
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
|
||||
|
||||
#### Scenario: Evidence coverage validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
|
||||
|
||||
#### Scenario: Verifier evaluation validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
|
||||
- **AND** the verdict SHALL be one of the case's allowed verdicts
|
||||
|
||||
#### Scenario: Answer keyword validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
|
||||
|
||||
#### Scenario: Degraded output validation
|
||||
- **WHEN** a trace verdict is `REJECT`
|
||||
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
|
||||
|
||||
### Requirement: Evaluation harness SHALL report quality and cost signals
|
||||
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
|
||||
|
||||
#### Scenario: JSON report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
|
||||
|
||||
#### Scenario: Markdown report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
|
||||
|
||||
#### Scenario: Aggregate metrics
|
||||
- **WHEN** multiple cases are evaluated
|
||||
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
|
||||
|
||||
### Requirement: Evaluation harness SHALL support offline fixture mode
|
||||
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
|
||||
|
||||
#### Scenario: Fixture trace evaluation
|
||||
- **WHEN** the evaluator is run against a directory of trace fixture files
|
||||
- **THEN** it SHALL evaluate each trace file against its matching case definition
|
||||
- **AND** it SHALL not require a running application service
|
||||
|
||||
#### Scenario: Missing fixture is reported clearly
|
||||
- **WHEN** a case has no matching trace fixture
|
||||
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
|
||||
Reference in New Issue
Block a user