Add diagnosis eval baseline diff
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,39 @@
|
||||
## Context
|
||||
|
||||
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Compare two `DiagnosisEvalReport` objects without requiring external services.
|
||||
- Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
|
||||
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
|
||||
- Write JSON and Markdown diff outputs for review.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not run the Agent or regenerate traces.
|
||||
- Do not introduce LLM-as-judge.
|
||||
- Do not change evaluator scoring rules.
|
||||
- Do not block on performance thresholds beyond simple numeric diff signals.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Decision: Compare report DTOs instead of raw traces.
|
||||
- Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
|
||||
- Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
|
||||
|
||||
- Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`.
|
||||
- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
|
||||
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
|
||||
|
||||
- Decision: Keep thresholds explicit and conservative.
|
||||
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
|
||||
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed.
|
||||
- Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
|
||||
- Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.
|
||||
@@ -0,0 +1,28 @@
|
||||
## Why
|
||||
|
||||
The evaluation baseline is now complete, but developers still need a repeatable way to decide whether a new Agent run regressed against that baseline. A deterministic baseline diff turns saved reports into an actionable regression signal instead of a static artifact.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a baseline diff model that compares two `DiagnosisEvalReport` objects.
|
||||
- Detect aggregate changes such as pass-rate drops, verdict distribution shifts, tool-call cost changes, and duration changes.
|
||||
- Detect per-case changes such as pass/fail regression, verdict changes, keyword coverage changes, evidence coverage loss, and missing/new cases.
|
||||
- Add JSON and Markdown diff output suitable for review.
|
||||
- Document how to interpret the diff in the eval docs.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness so a current report can be compared against the saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects eval-only Java code under `src/main/java/com/superbiz/agent/eval`.
|
||||
- Adds focused tests under `src/test/java/com/superbiz/agent/eval`.
|
||||
- Updates `mvp/eval` documentation and may add sample diff output.
|
||||
- No production Agent runtime, API, database schema, or external dependency changes are expected.
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
||||
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
||||
|
||||
#### Scenario: Aggregate regression detection
|
||||
- **WHEN** the current report has a lower pass rate than the baseline report
|
||||
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
||||
|
||||
#### Scenario: Cost signal detection
|
||||
- **WHEN** average tool-call count or average duration changes between reports
|
||||
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
||||
|
||||
#### Scenario: Verdict distribution comparison
|
||||
- **WHEN** verdict counts differ between reports
|
||||
- **THEN** the diff SHALL record the verdict distribution changes
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare case-level report results
|
||||
The system SHALL compare case results by case id and report actionable per-case changes.
|
||||
|
||||
#### Scenario: Case pass/fail regression
|
||||
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
||||
- **THEN** the diff SHALL record a regression for that case
|
||||
|
||||
#### Scenario: Evidence coverage regression
|
||||
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
||||
- **THEN** the diff SHALL record a regression naming the case and tool
|
||||
|
||||
#### Scenario: Missing case detection
|
||||
- **WHEN** a baseline case is absent from the current report
|
||||
- **THEN** the diff SHALL record a regression for the missing case
|
||||
|
||||
#### Scenario: New case detection
|
||||
- **WHEN** a current report contains a case absent from the baseline
|
||||
- **THEN** the diff SHALL record the case as a non-regression change
|
||||
|
||||
### Requirement: Evaluation harness SHALL report baseline diff results
|
||||
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
||||
|
||||
#### Scenario: JSON diff output
|
||||
- **WHEN** a baseline diff is written as JSON
|
||||
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
||||
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
@@ -0,0 +1,23 @@
|
||||
## 1. OpenSpec And Issue Setup
|
||||
|
||||
- [x] 1.1 Create slug-based issue and devflow tracking files.
|
||||
- [x] 1.2 Create OpenSpec proposal, design, delta spec, and tasks.
|
||||
|
||||
## 2. Baseline Diff Implementation
|
||||
|
||||
- [x] 2.1 Add diff result data structures for summary and per-item changes.
|
||||
- [x] 2.2 Implement deterministic report comparison rules.
|
||||
- [x] 2.3 Implement JSON and Markdown diff report writing.
|
||||
|
||||
## 3. Documentation
|
||||
|
||||
- [x] 3.1 Document baseline diff inputs, outputs, and interpretation in eval docs.
|
||||
- [x] 3.2 Add sample diff output for a representative regression.
|
||||
|
||||
## 4. Tests And Validation
|
||||
|
||||
- [x] 4.1 Add focused tests for aggregate and case-level diff behavior.
|
||||
- [x] 4.2 Add focused tests for JSON and Markdown diff output.
|
||||
- [x] 4.3 Run evaluator/diff tests.
|
||||
- [x] 4.4 Run compile verification.
|
||||
- [x] 4.5 Run OpenSpec validation.
|
||||
@@ -88,3 +88,48 @@ The system SHALL preserve a generated baseline report for the full fixed fixture
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes fixtures or evaluator rules
|
||||
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
||||
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
||||
|
||||
#### Scenario: Aggregate regression detection
|
||||
- **WHEN** the current report has a lower pass rate than the baseline report
|
||||
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
||||
|
||||
#### Scenario: Cost signal detection
|
||||
- **WHEN** average tool-call count or average duration changes between reports
|
||||
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
||||
|
||||
#### Scenario: Verdict distribution comparison
|
||||
- **WHEN** verdict counts differ between reports
|
||||
- **THEN** the diff SHALL record the verdict distribution changes
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare case-level report results
|
||||
The system SHALL compare case results by case id and report actionable per-case changes.
|
||||
|
||||
#### Scenario: Case pass/fail regression
|
||||
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
||||
- **THEN** the diff SHALL record a regression for that case
|
||||
|
||||
#### Scenario: Evidence coverage regression
|
||||
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
||||
- **THEN** the diff SHALL record a regression naming the case and tool
|
||||
|
||||
#### Scenario: Missing case detection
|
||||
- **WHEN** a baseline case is absent from the current report
|
||||
- **THEN** the diff SHALL record a regression for the missing case
|
||||
|
||||
#### Scenario: New case detection
|
||||
- **WHEN** a current report contains a case absent from the baseline
|
||||
- **THEN** the diff SHALL record the case as a non-regression change
|
||||
|
||||
### Requirement: Evaluation harness SHALL report baseline diff results
|
||||
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
||||
|
||||
#### Scenario: JSON diff output
|
||||
- **WHEN** a baseline diff is written as JSON
|
||||
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
||||
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
|
||||
Reference in New Issue
Block a user