Add diagnosis eval baseline diff

This commit is contained in:
aruo
2026-07-05 00:59:53 +08:00
parent 4c7c53b024
commit 69deb15330
22 changed files with 1069 additions and 0 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,39 @@
## Context
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
## Goals / Non-Goals
**Goals:**
- Compare two `DiagnosisEvalReport` objects without requiring external services.
- Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
- Write JSON and Markdown diff outputs for review.
**Non-Goals:**
- Do not run the Agent or regenerate traces.
- Do not introduce LLM-as-judge.
- Do not change evaluator scoring rules.
- Do not block on performance thresholds beyond simple numeric diff signals.
## Decisions
- Decision: Compare report DTOs instead of raw traces.
- Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
- Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
- Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`.
- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
- Decision: Keep thresholds explicit and conservative.
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
## Risks / Trade-offs
- Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed.
- Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
- Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.
@@ -0,0 +1,28 @@
## Why
The evaluation baseline is now complete, but developers still need a repeatable way to decide whether a new Agent run regressed against that baseline. A deterministic baseline diff turns saved reports into an actionable regression signal instead of a static artifact.
## What Changes
- Add a baseline diff model that compares two `DiagnosisEvalReport` objects.
- Detect aggregate changes such as pass-rate drops, verdict distribution shifts, tool-call cost changes, and duration changes.
- Detect per-case changes such as pass/fail regression, verdict changes, keyword coverage changes, evidence coverage loss, and missing/new cases.
- Add JSON and Markdown diff output suitable for review.
- Document how to interpret the diff in the eval docs.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `diagnosis-eval-harness`: Extend the existing evaluation harness so a current report can be compared against the saved baseline report.
## Impact
- Affects eval-only Java code under `src/main/java/com/superbiz/agent/eval`.
- Adds focused tests under `src/test/java/com/superbiz/agent/eval`.
- Updates `mvp/eval` documentation and may add sample diff output.
- No production Agent runtime, API, database schema, or external dependency changes are expected.
@@ -0,0 +1,46 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
#### Scenario: Aggregate regression detection
- **WHEN** the current report has a lower pass rate than the baseline report
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
#### Scenario: Cost signal detection
- **WHEN** average tool-call count or average duration changes between reports
- **THEN** the diff SHALL record the baseline value, current value, and delta
#### Scenario: Verdict distribution comparison
- **WHEN** verdict counts differ between reports
- **THEN** the diff SHALL record the verdict distribution changes
### Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
#### Scenario: Case pass/fail regression
- **WHEN** a case changes from passing in the baseline to failing in the current report
- **THEN** the diff SHALL record a regression for that case
#### Scenario: Evidence coverage regression
- **WHEN** a required evidence tool changes from covered to uncovered for a case
- **THEN** the diff SHALL record a regression naming the case and tool
#### Scenario: Missing case detection
- **WHEN** a baseline case is absent from the current report
- **THEN** the diff SHALL record a regression for the missing case
#### Scenario: New case detection
- **WHEN** a current report contains a case absent from the baseline
- **THEN** the diff SHALL record the case as a non-regression change
### Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
#### Scenario: JSON diff output
- **WHEN** a baseline diff is written as JSON
- **THEN** it SHALL include aggregate summary fields and detailed diff items
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items
@@ -0,0 +1,23 @@
## 1. OpenSpec And Issue Setup
- [x] 1.1 Create slug-based issue and devflow tracking files.
- [x] 1.2 Create OpenSpec proposal, design, delta spec, and tasks.
## 2. Baseline Diff Implementation
- [x] 2.1 Add diff result data structures for summary and per-item changes.
- [x] 2.2 Implement deterministic report comparison rules.
- [x] 2.3 Implement JSON and Markdown diff report writing.
## 3. Documentation
- [x] 3.1 Document baseline diff inputs, outputs, and interpretation in eval docs.
- [x] 3.2 Add sample diff output for a representative regression.
## 4. Tests And Validation
- [x] 4.1 Add focused tests for aggregate and case-level diff behavior.
- [x] 4.2 Add focused tests for JSON and Markdown diff output.
- [x] 4.3 Run evaluator/diff tests.
- [x] 4.4 Run compile verification.
- [x] 4.5 Run OpenSpec validation.
@@ -88,3 +88,48 @@ The system SHALL preserve a generated baseline report for the full fixed fixture
#### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes fixtures or evaluator rules
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
### Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
#### Scenario: Aggregate regression detection
- **WHEN** the current report has a lower pass rate than the baseline report
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
#### Scenario: Cost signal detection
- **WHEN** average tool-call count or average duration changes between reports
- **THEN** the diff SHALL record the baseline value, current value, and delta
#### Scenario: Verdict distribution comparison
- **WHEN** verdict counts differ between reports
- **THEN** the diff SHALL record the verdict distribution changes
### Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
#### Scenario: Case pass/fail regression
- **WHEN** a case changes from passing in the baseline to failing in the current report
- **THEN** the diff SHALL record a regression for that case
#### Scenario: Evidence coverage regression
- **WHEN** a required evidence tool changes from covered to uncovered for a case
- **THEN** the diff SHALL record a regression naming the case and tool
#### Scenario: Missing case detection
- **WHEN** a baseline case is absent from the current report
- **THEN** the diff SHALL record a regression for the missing case
#### Scenario: New case detection
- **WHEN** a current report contains a case absent from the baseline
- **THEN** the diff SHALL record the case as a non-regression change
### Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
#### Scenario: JSON diff output
- **WHEN** a baseline diff is written as JSON
- **THEN** it SHALL include aggregate summary fields and detailed diff items
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items