Add diagnosis eval baseline diff

This commit is contained in:
aruo
2026-07-05 00:59:53 +08:00
parent 4c7c53b024
commit 69deb15330
22 changed files with 1069 additions and 0 deletions
+1
View File
@@ -4,6 +4,7 @@
| 日期 | slug | 领域 | 关键词 | 状态 | | 日期 | slug | 领域 | 关键词 | 状态 |
|---|---|---|---|---| |---|---|---|---|---|
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived | | 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived | | 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived | | 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
@@ -0,0 +1,37 @@
# Acceptance: diagnosis-eval-baseline-diff
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
- JSON and Markdown diff output are available.
- No production runtime behavior has been changed.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
- Result: passed
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
- Result: passed
@@ -0,0 +1,30 @@
# Brief: diagnosis-eval-baseline-diff
## Background
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
## Goals
1. Compare baseline and current `DiagnosisEvalReport` objects.
2. Detect aggregate and per-case regressions.
3. Output JSON and Markdown diff reports.
4. Document how to read the diff in interview and engineering terms.
## Scope
- Diff data structures
- Deterministic report comparison
- JSON / Markdown diff output
- Focused tests and eval docs
## Non-Goals
- No live Agent execution
- No LLM-as-judge
- No evaluator scoring rule changes
- No production API changes
## Related OpenSpec
`openspec/changes/diagnosis-eval-baseline-diff/`
@@ -0,0 +1,28 @@
# Diagnosis Eval Baseline Diff Decisions
## Clarify
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
- Slug: `diagnosis-eval-baseline-diff`
- Devflow scale: standard-light
## Context
- `diagnosis-eval-harness` created deterministic fixture evaluation.
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
- This change compares new reports against that baseline.
## Key Decisions
- Decision: Diff report DTOs instead of raw traces.
- Reason: the report is the stable contract for regression review.
- Decision: Use deterministic code rules instead of LLM-as-judge.
- Reason: baseline regression checks should be repeatable and explainable.
- Decision: Output both JSON and Markdown.
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
## Open Questions
- Whether a future change should expose this through a CLI or Maven goal.
@@ -0,0 +1,11 @@
# Evidence: diagnosis-eval-baseline-diff
## Evidence Log
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
+14
View File
@@ -8,6 +8,7 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
- Offline trace fixtures: `fixtures/*.json` - Offline trace fixtures: `fixtures/*.json`
- Field definitions: `schema.md` - Field definitions: `schema.md`
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md` - Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
- Evaluator implementation: `DiagnosisTraceEvaluator` - Evaluator implementation: `DiagnosisTraceEvaluator`
- Report writer: `DiagnosisEvalReportWriter` - Report writer: `DiagnosisEvalReportWriter`
@@ -45,3 +46,16 @@ fixed diagnosis case
-> JSON / Markdown report -> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, and verifier behavior -> regression signal for prompts, tools, retrieval, and verifier behavior
``` ```
## Baseline Diff
Baseline diff compares a current report against `reports/baseline-report.json`.
```text
baseline report
current report
-> deterministic diff
-> regressions, improvements, and changed signals
```
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
@@ -0,0 +1,85 @@
{
"baselineTotalCases" : 5,
"currentTotalCases" : 5,
"baselinePassedCases" : 5,
"currentPassedCases" : 4,
"baselinePassRate" : 1.0,
"currentPassRate" : 0.8,
"regressionCount" : 6,
"improvementCount" : 0,
"changedCount" : 2,
"hasRegression" : true,
"items" : [ {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "passRate",
"baselineValue" : "1.0",
"currentValue" : "0.8",
"delta" : -0.19999999999999996,
"message" : "passRate changed"
}, {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "averageToolCallCount",
"baselineValue" : "2.0",
"currentValue" : "3.0",
"delta" : 1.0,
"message" : "averageToolCallCount changed"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.LOW_CONFID",
"baselineValue" : "3",
"currentValue" : "2",
"delta" : -1.0,
"message" : "verdict count changed for LOW_CONFID"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.REJECT",
"baselineValue" : "0",
"currentValue" : "1",
"delta" : 1.0,
"message" : "verdict count changed for REJECT"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "passed",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout pass state changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "verdict",
"baselineValue" : "LOW_CONFID",
"currentValue" : "REJECT",
"delta" : -1.0,
"message" : "redis-timeout verdict changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "matchedKeywordCount",
"baselineValue" : "2",
"currentValue" : "1",
"delta" : -1.0,
"message" : "redis-timeout matchedKeywordCount changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "evidenceCoverage.query_logs",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout evidence coverage changed for query_logs"
} ]
}
+22
View File
@@ -0,0 +1,22 @@
# Diagnosis Eval Baseline Diff
- Baseline pass rate: 100.00%
- Current pass rate: 80.00%
- Baseline passed cases: 5/5
- Current passed cases: 4/5
- Regressions: 6
- Improvements: 0
- Other changes: 2
## Diff Items
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
| --- | --- | --- | --- | --- | --- | ---: | --- |
| REGRESSION | aggregate | - | passRate | 1.0 | 0.8 | -0.200 | passRate changed |
| REGRESSION | aggregate | - | averageToolCallCount | 2.0 | 3.0 | 1.000 | averageToolCallCount changed |
| CHANGED | aggregate | - | verdictDistribution.LOW_CONFID | 3 | 2 | -1.000 | verdict count changed for LOW_CONFID |
| CHANGED | aggregate | - | verdictDistribution.REJECT | 0 | 1 | 1.000 | verdict count changed for REJECT |
| REGRESSION | case | redis-timeout | passed | true | false | - | redis-timeout pass state changed |
| REGRESSION | case | redis-timeout | verdict | LOW_CONFID | REJECT | -1.000 | redis-timeout verdict changed |
| REGRESSION | case | redis-timeout | matchedKeywordCount | 2 | 1 | -1.000 | redis-timeout matchedKeywordCount changed |
| REGRESSION | case | redis-timeout | evidenceCoverage.query_logs | true | false | - | redis-timeout evidence coverage changed for query_logs |
+65
View File
@@ -135,3 +135,68 @@ Java 类型:`DiagnosisEvalReport`
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。 Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。 这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
``` ```
## 6. Baseline Diff
Baseline diff 是拿两份 report 做对比:
```text
baseline report:以前认可的基准结果
current report:这次改动后跑出来的新结果
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
```
Java 类型:
- `DiagnosisEvalDiffReport`
- `DiagnosisEvalDiffItem`
`DiagnosisEvalDiffReport` 字段:
| 字段 | 意思 |
| --- | --- |
| `baselineTotalCases` | baseline 里有多少条 case |
| `currentTotalCases` | current 里有多少条 case |
| `baselinePassedCases` | baseline 通过了多少条 |
| `currentPassedCases` | current 通过了多少条 |
| `baselinePassRate` | baseline 通过率 |
| `currentPassRate` | current 通过率 |
| `regressionCount` | 退化项数量 |
| `improvementCount` | 改善项数量 |
| `changedCount` | 普通变化项数量 |
| `hasRegression` | 是否存在退化 |
| `items` | 具体 diff 明细 |
`DiagnosisEvalDiffItem` 字段:
| 字段 | 意思 |
| --- | --- |
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
| `caseId` | 如果是单条 case 变化,这里记录 case id |
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
| `baselineValue` | baseline 里的值 |
| `currentValue` | current 里的值 |
| `delta` | 数值变化量;非数值变化为空 |
| `message` | 给人看的变化说明 |
口语化判断规则:
```text
pass rate 下降:退化
case 从通过变失败:退化
证据工具从有变没有:退化
关键词命中变少:退化
工具调用或耗时升高:成本上升,记为退化信号
verdict 分布变化:记录变化,供人工判断是否符合预期
```
面试里可以这样讲:
```text
我把 baseline report 和当前 report 做结构化 diff。
它不是再问 LLM,而是用代码比较固定字段。
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
diff 会直接标成 regression。
这样 Agent 改动可以用固定基准做回归判断。
```
+1
View File
@@ -9,3 +9,4 @@
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) | | ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) | | ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) | | expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
@@ -0,0 +1,73 @@
# Diagnosis Eval Baseline Diff
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-05
**来源**:P1-B follow-up
**依赖**:`diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`
---
## 背景
现在项目已经有固定诊断 case、完整 fixture 和 baseline report。下一步需要把 baseline 真正用起来:每次改 Agent 后,把新的 report 和 baseline report 做对比。
---
## 问题
当前 baseline 只能告诉我们“标准状态是什么”,但还不能自动告诉我们“这次改动有没有变差”。
典型问题包括:
- pass rate 是否下降。
- 某个 case 是否从通过变失败。
- 某个 evidence tool 是否从覆盖变成缺失。
- verifier verdict 分布是否异常变化。
- 平均工具调用数和耗时是否明显上升。
---
## 目标
新增一个 deterministic baseline diff 能力,用代码比较两份 `DiagnosisEvalReport`。
完成后应该做到:
- 输入 baseline report 和 current report。
- 输出结构化 diff。
- 标出 regression、improvement 和普通 changed。
- 支持 JSON 和 Markdown 输出。
- 文档说明面试时怎么解释这套回归判断。
---
## 范围
### In scope
- report-level diff 数据结构。
- aggregate 指标比较。
- case-level 指标比较。
- JSON / Markdown diff writer。
- focused tests 和 eval 文档。
### Out of scope
- 不运行真实 Agent。
- 不生成新 trace。
- 不引入 LLM-as-judge。
- 不改现有 evaluator 评分规则。
---
## 面试表达
可以这样讲:
```text
我不是只保存了一份 baseline,而是加了 baseline diff。
每次改 prompt、tool、retrieval 或 verifier 后,
我都能把新 report 和 baseline 比较,
直接看到哪些 case 退化、哪些证据缺失、成本有没有上升。
```
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,39 @@
## Context
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
## Goals / Non-Goals
**Goals:**
- Compare two `DiagnosisEvalReport` objects without requiring external services.
- Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
- Write JSON and Markdown diff outputs for review.
**Non-Goals:**
- Do not run the Agent or regenerate traces.
- Do not introduce LLM-as-judge.
- Do not change evaluator scoring rules.
- Do not block on performance thresholds beyond simple numeric diff signals.
## Decisions
- Decision: Compare report DTOs instead of raw traces.
- Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
- Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
- Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`.
- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
- Decision: Keep thresholds explicit and conservative.
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
## Risks / Trade-offs
- Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed.
- Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
- Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.
@@ -0,0 +1,28 @@
## Why
The evaluation baseline is now complete, but developers still need a repeatable way to decide whether a new Agent run regressed against that baseline. A deterministic baseline diff turns saved reports into an actionable regression signal instead of a static artifact.
## What Changes
- Add a baseline diff model that compares two `DiagnosisEvalReport` objects.
- Detect aggregate changes such as pass-rate drops, verdict distribution shifts, tool-call cost changes, and duration changes.
- Detect per-case changes such as pass/fail regression, verdict changes, keyword coverage changes, evidence coverage loss, and missing/new cases.
- Add JSON and Markdown diff output suitable for review.
- Document how to interpret the diff in the eval docs.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `diagnosis-eval-harness`: Extend the existing evaluation harness so a current report can be compared against the saved baseline report.
## Impact
- Affects eval-only Java code under `src/main/java/com/superbiz/agent/eval`.
- Adds focused tests under `src/test/java/com/superbiz/agent/eval`.
- Updates `mvp/eval` documentation and may add sample diff output.
- No production Agent runtime, API, database schema, or external dependency changes are expected.
@@ -0,0 +1,46 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
#### Scenario: Aggregate regression detection
- **WHEN** the current report has a lower pass rate than the baseline report
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
#### Scenario: Cost signal detection
- **WHEN** average tool-call count or average duration changes between reports
- **THEN** the diff SHALL record the baseline value, current value, and delta
#### Scenario: Verdict distribution comparison
- **WHEN** verdict counts differ between reports
- **THEN** the diff SHALL record the verdict distribution changes
### Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
#### Scenario: Case pass/fail regression
- **WHEN** a case changes from passing in the baseline to failing in the current report
- **THEN** the diff SHALL record a regression for that case
#### Scenario: Evidence coverage regression
- **WHEN** a required evidence tool changes from covered to uncovered for a case
- **THEN** the diff SHALL record a regression naming the case and tool
#### Scenario: Missing case detection
- **WHEN** a baseline case is absent from the current report
- **THEN** the diff SHALL record a regression for the missing case
#### Scenario: New case detection
- **WHEN** a current report contains a case absent from the baseline
- **THEN** the diff SHALL record the case as a non-regression change
### Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
#### Scenario: JSON diff output
- **WHEN** a baseline diff is written as JSON
- **THEN** it SHALL include aggregate summary fields and detailed diff items
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items
@@ -0,0 +1,23 @@
## 1. OpenSpec And Issue Setup
- [x] 1.1 Create slug-based issue and devflow tracking files.
- [x] 1.2 Create OpenSpec proposal, design, delta spec, and tasks.
## 2. Baseline Diff Implementation
- [x] 2.1 Add diff result data structures for summary and per-item changes.
- [x] 2.2 Implement deterministic report comparison rules.
- [x] 2.3 Implement JSON and Markdown diff report writing.
## 3. Documentation
- [x] 3.1 Document baseline diff inputs, outputs, and interpretation in eval docs.
- [x] 3.2 Add sample diff output for a representative regression.
## 4. Tests And Validation
- [x] 4.1 Add focused tests for aggregate and case-level diff behavior.
- [x] 4.2 Add focused tests for JSON and Markdown diff output.
- [x] 4.3 Run evaluator/diff tests.
- [x] 4.4 Run compile verification.
- [x] 4.5 Run OpenSpec validation.
@@ -88,3 +88,48 @@ The system SHALL preserve a generated baseline report for the full fixed fixture
#### Scenario: Baseline regeneration is documented #### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes fixtures or evaluator rules - **WHEN** a developer changes fixtures or evaluator rules
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report - **THEN** the eval documentation SHALL explain how to regenerate the baseline report
### Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
#### Scenario: Aggregate regression detection
- **WHEN** the current report has a lower pass rate than the baseline report
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
#### Scenario: Cost signal detection
- **WHEN** average tool-call count or average duration changes between reports
- **THEN** the diff SHALL record the baseline value, current value, and delta
#### Scenario: Verdict distribution comparison
- **WHEN** verdict counts differ between reports
- **THEN** the diff SHALL record the verdict distribution changes
### Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
#### Scenario: Case pass/fail regression
- **WHEN** a case changes from passing in the baseline to failing in the current report
- **THEN** the diff SHALL record a regression for that case
#### Scenario: Evidence coverage regression
- **WHEN** a required evidence tool changes from covered to uncovered for a case
- **THEN** the diff SHALL record a regression naming the case and tool
#### Scenario: Missing case detection
- **WHEN** a baseline case is absent from the current report
- **THEN** the diff SHALL record a regression for the missing case
#### Scenario: New case detection
- **WHEN** a current report contains a case absent from the baseline
- **THEN** the diff SHALL record the case as a non-regression change
### Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
#### Scenario: JSON diff output
- **WHEN** a baseline diff is written as JSON
- **THEN** it SHALL include aggregate summary fields and detailed diff items
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items
@@ -0,0 +1,252 @@
package com.superbiz.agent.eval;
import java.util.ArrayList;
import java.util.Comparator;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.function.Function;
import java.util.stream.Collectors;
public class DiagnosisEvalBaselineDiffer {
private static final String REGRESSION = "REGRESSION";
private static final String IMPROVEMENT = "IMPROVEMENT";
private static final String CHANGED = "CHANGED";
public DiagnosisEvalDiffReport compare(DiagnosisEvalReport baseline, DiagnosisEvalReport current) {
List<DiagnosisEvalDiffItem> items = new ArrayList<>();
compareDouble(items, "aggregate", null, "passRate",
baseline.getPassRate(), current.getPassRate(), true);
compareDouble(items, "aggregate", null, "averageToolCallCount",
baseline.getAverageToolCallCount(), current.getAverageToolCallCount(), false);
compareDouble(items, "aggregate", null, "averageDurationMs",
baseline.getAverageDurationMs(), current.getAverageDurationMs(), false);
compareVerdictDistribution(items, baseline.getVerdictDistribution(), current.getVerdictDistribution());
compareCases(items, safeResults(baseline), safeResults(current));
int regressionCount = countType(items, REGRESSION);
int improvementCount = countType(items, IMPROVEMENT);
int changedCount = countType(items, CHANGED);
return DiagnosisEvalDiffReport.builder()
.baselineTotalCases(baseline.getTotalCases())
.currentTotalCases(current.getTotalCases())
.baselinePassedCases(baseline.getPassedCases())
.currentPassedCases(current.getPassedCases())
.baselinePassRate(baseline.getPassRate())
.currentPassRate(current.getPassRate())
.regressionCount(regressionCount)
.improvementCount(improvementCount)
.changedCount(changedCount)
.hasRegression(regressionCount > 0)
.items(items)
.build();
}
private void compareVerdictDistribution(List<DiagnosisEvalDiffItem> items,
Map<String, Long> baseline,
Map<String, Long> current) {
Set<String> verdicts = new LinkedHashSet<>();
verdicts.addAll(safeMap(baseline).keySet());
verdicts.addAll(safeMap(current).keySet());
for (String verdict : verdicts) {
long baselineCount = safeMap(baseline).getOrDefault(verdict, 0L);
long currentCount = safeMap(current).getOrDefault(verdict, 0L);
if (baselineCount != currentCount) {
items.add(item(CHANGED, "aggregate", null, "verdictDistribution." + verdict,
String.valueOf(baselineCount), String.valueOf(currentCount),
(double) currentCount - baselineCount,
"verdict count changed for " + verdict));
}
}
}
private void compareCases(List<DiagnosisEvalDiffItem> items,
List<DiagnosisEvalResult> baselineResults,
List<DiagnosisEvalResult> currentResults) {
Map<String, DiagnosisEvalResult> baselineById = byCaseId(baselineResults);
Map<String, DiagnosisEvalResult> currentById = byCaseId(currentResults);
Set<String> caseIds = new LinkedHashSet<>();
caseIds.addAll(baselineById.keySet());
caseIds.addAll(currentById.keySet());
for (String caseId : caseIds) {
DiagnosisEvalResult baseline = baselineById.get(caseId);
DiagnosisEvalResult current = currentById.get(caseId);
if (baseline == null) {
items.add(item(CHANGED, "case", caseId, "casePresence",
"missing", "present", null, "new case appears in current report"));
continue;
}
if (current == null) {
items.add(item(REGRESSION, "case", caseId, "casePresence",
"present", "missing", null, "baseline case is missing from current report"));
continue;
}
comparePassState(items, baseline, current);
compareVerdict(items, baseline, current);
compareInteger(items, caseId, "matchedKeywordCount",
baseline.getMatchedKeywordCount(), current.getMatchedKeywordCount(), true);
compareInteger(items, caseId, "toolCallCount",
baseline.getToolCallCount(), current.getToolCallCount(), false);
compareInteger(items, caseId, "durationMs",
baseline.getDurationMs(), current.getDurationMs(), false);
compareEvidenceCoverage(items, baseline, current);
}
}
private void comparePassState(List<DiagnosisEvalDiffItem> items,
DiagnosisEvalResult baseline,
DiagnosisEvalResult current) {
if (baseline.isPassed() == current.isPassed()) {
return;
}
String type = baseline.isPassed() ? REGRESSION : IMPROVEMENT;
items.add(item(type, "case", baseline.getCaseId(), "passed",
String.valueOf(baseline.isPassed()), String.valueOf(current.isPassed()), null,
baseline.getCaseId() + " pass state changed"));
}
private void compareVerdict(List<DiagnosisEvalDiffItem> items,
DiagnosisEvalResult baseline,
DiagnosisEvalResult current) {
if (Objects.equals(baseline.getVerdict(), current.getVerdict())) {
return;
}
int baselineRank = verdictRank(baseline.getVerdict());
int currentRank = verdictRank(current.getVerdict());
String type = currentRank < baselineRank ? REGRESSION : currentRank > baselineRank ? IMPROVEMENT : CHANGED;
items.add(item(type, "case", baseline.getCaseId(), "verdict",
value(baseline.getVerdict()), value(current.getVerdict()), (double) currentRank - baselineRank,
baseline.getCaseId() + " verdict changed"));
}
private void compareEvidenceCoverage(List<DiagnosisEvalDiffItem> items,
DiagnosisEvalResult baseline,
DiagnosisEvalResult current) {
Set<String> tools = new LinkedHashSet<>();
tools.addAll(safeMap(baseline.getEvidenceCoverage()).keySet());
tools.addAll(safeMap(current.getEvidenceCoverage()).keySet());
for (String tool : tools) {
boolean baselineCovered = Boolean.TRUE.equals(safeMap(baseline.getEvidenceCoverage()).get(tool));
boolean currentCovered = Boolean.TRUE.equals(safeMap(current.getEvidenceCoverage()).get(tool));
if (baselineCovered == currentCovered) {
continue;
}
String type = baselineCovered ? REGRESSION : IMPROVEMENT;
items.add(item(type, "case", baseline.getCaseId(), "evidenceCoverage." + tool,
String.valueOf(baselineCovered), String.valueOf(currentCovered), null,
baseline.getCaseId() + " evidence coverage changed for " + tool));
}
}
private void compareDouble(List<DiagnosisEvalDiffItem> items,
String scope,
String caseId,
String metric,
double baseline,
double current,
boolean higherIsBetter) {
if (Double.compare(baseline, current) == 0) {
return;
}
double delta = current - baseline;
String type = classifyDelta(delta, higherIsBetter);
items.add(item(type, scope, caseId, metric,
String.valueOf(baseline), String.valueOf(current), delta,
metric + " changed"));
}
private void compareInteger(List<DiagnosisEvalDiffItem> items,
String caseId,
String metric,
Integer baseline,
Integer current,
boolean higherIsBetter) {
if (Objects.equals(baseline, current)) {
return;
}
if (baseline == null || current == null) {
items.add(item(CHANGED, "case", caseId, metric,
value(baseline), value(current), null, caseId + " " + metric + " changed"));
return;
}
int delta = current - baseline;
items.add(item(classifyDelta(delta, higherIsBetter), "case", caseId, metric,
String.valueOf(baseline), String.valueOf(current), (double) delta,
caseId + " " + metric + " changed"));
}
private String classifyDelta(double delta, boolean higherIsBetter) {
if (delta == 0.0) {
return CHANGED;
}
boolean improved = higherIsBetter ? delta > 0 : delta < 0;
return improved ? IMPROVEMENT : REGRESSION;
}
private DiagnosisEvalDiffItem item(String type,
String scope,
String caseId,
String metric,
String baselineValue,
String currentValue,
Double delta,
String message) {
return DiagnosisEvalDiffItem.builder()
.type(type)
.scope(scope)
.caseId(caseId)
.metric(metric)
.baselineValue(baselineValue)
.currentValue(currentValue)
.delta(delta)
.message(message)
.build();
}
private Map<String, DiagnosisEvalResult> byCaseId(List<DiagnosisEvalResult> results) {
return results.stream()
.sorted(Comparator.comparing(DiagnosisEvalResult::getCaseId))
.collect(Collectors.toMap(
DiagnosisEvalResult::getCaseId,
Function.identity(),
(left, right) -> right,
LinkedHashMap::new));
}
private List<DiagnosisEvalResult> safeResults(DiagnosisEvalReport report) {
return report.getResults() == null ? List.of() : report.getResults();
}
private <T> Map<String, T> safeMap(Map<String, T> value) {
return value == null ? Map.of() : value;
}
private int countType(List<DiagnosisEvalDiffItem> items, String type) {
return (int) items.stream().filter(item -> type.equals(item.getType())).count();
}
private int verdictRank(String verdict) {
if ("PASS".equals(verdict)) {
return 3;
}
if ("LOW_CONFID".equals(verdict)) {
return 2;
}
if ("REJECT".equals(verdict)) {
return 1;
}
return 0;
}
private String value(Object value) {
return value == null ? "-" : String.valueOf(value);
}
}
@@ -0,0 +1,22 @@
package com.superbiz.agent.eval;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.NoArgsConstructor;
@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
public class DiagnosisEvalDiffItem {
private String type;
private String scope;
private String caseId;
private String metric;
private String baselineValue;
private String currentValue;
private Double delta;
private String message;
}
@@ -0,0 +1,27 @@
package com.superbiz.agent.eval;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.NoArgsConstructor;
import java.util.List;
@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
public class DiagnosisEvalDiffReport {
private int baselineTotalCases;
private int currentTotalCases;
private int baselinePassedCases;
private int currentPassedCases;
private double baselinePassRate;
private double currentPassRate;
private int regressionCount;
private int improvementCount;
private int changedCount;
private boolean hasRegression;
private List<DiagnosisEvalDiffItem> items;
}
@@ -0,0 +1,78 @@
package com.superbiz.agent.eval;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
public class DiagnosisEvalDiffReportWriter {
private final ObjectMapper objectMapper;
public DiagnosisEvalDiffReportWriter(ObjectMapper objectMapper) {
this.objectMapper = objectMapper;
}
public void writeJson(DiagnosisEvalDiffReport report, Path outputFile) throws IOException {
Files.createDirectories(outputFile.getParent());
objectMapper.writerWithDefaultPrettyPrinter().writeValue(outputFile.toFile(), report);
}
public void writeMarkdown(DiagnosisEvalDiffReport report, Path outputFile) throws IOException {
Files.createDirectories(outputFile.getParent());
Files.writeString(outputFile, toMarkdown(report), StandardCharsets.UTF_8);
}
public String toMarkdown(DiagnosisEvalDiffReport report) {
StringBuilder builder = new StringBuilder();
builder.append("# Diagnosis Eval Baseline Diff\n\n");
builder.append("- Baseline pass rate: ").append(formatPercent(report.getBaselinePassRate())).append("\n");
builder.append("- Current pass rate: ").append(formatPercent(report.getCurrentPassRate())).append("\n");
builder.append("- Baseline passed cases: ").append(report.getBaselinePassedCases()).append("/")
.append(report.getBaselineTotalCases()).append("\n");
builder.append("- Current passed cases: ").append(report.getCurrentPassedCases()).append("/")
.append(report.getCurrentTotalCases()).append("\n");
builder.append("- Regressions: ").append(report.getRegressionCount()).append("\n");
builder.append("- Improvements: ").append(report.getImprovementCount()).append("\n");
builder.append("- Other changes: ").append(report.getChangedCount()).append("\n\n");
builder.append("## Diff Items\n\n");
if (report.getItems() == null || report.getItems().isEmpty()) {
builder.append("- No differences\n");
return builder.toString();
}
builder.append("| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |\n");
builder.append("| --- | --- | --- | --- | --- | --- | ---: | --- |\n");
for (DiagnosisEvalDiffItem item : report.getItems()) {
builder.append("| ")
.append(valueOrDash(item.getType()))
.append(" | ")
.append(valueOrDash(item.getScope()))
.append(" | ")
.append(valueOrDash(item.getCaseId()))
.append(" | ")
.append(valueOrDash(item.getMetric()))
.append(" | ")
.append(valueOrDash(item.getBaselineValue()))
.append(" | ")
.append(valueOrDash(item.getCurrentValue()))
.append(" | ")
.append(item.getDelta() == null ? "-" : String.format("%.3f", item.getDelta()))
.append(" | ")
.append(valueOrDash(item.getMessage()))
.append(" |\n");
}
return builder.toString();
}
private String formatPercent(double value) {
return String.format("%.2f%%", value * 100);
}
private String valueOrDash(String value) {
return value == null || value.isBlank() ? "-" : value;
}
}
@@ -0,0 +1,140 @@
package com.superbiz.agent.eval;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
class DiagnosisEvalBaselineDiffTest {
private final ObjectMapper objectMapper = new ObjectMapper();
private final DiagnosisEvalBaselineDiffer differ = new DiagnosisEvalBaselineDiffer();
@Test
void compareReportsDetectsAggregateAndCaseRegressions() throws Exception {
DiagnosisEvalReport baseline = readBaselineReport();
DiagnosisEvalReport current = readBaselineReport();
degradeRedisCase(current);
DiagnosisEvalDiffReport diff = differ.compare(baseline, current);
assertTrue(diff.isHasRegression());
assertEquals(6, diff.getRegressionCount());
assertEquals(2, diff.getChangedCount());
assertTrue(hasItem(diff, "REGRESSION", "aggregate", null, "passRate"));
assertTrue(hasItem(diff, "REGRESSION", "aggregate", null, "averageToolCallCount"));
assertTrue(hasItem(diff, "REGRESSION", "case", "redis-timeout", "passed"));
assertTrue(hasItem(diff, "REGRESSION", "case", "redis-timeout", "verdict"));
assertTrue(hasItem(diff, "REGRESSION", "case", "redis-timeout", "matchedKeywordCount"));
assertTrue(hasItem(diff, "REGRESSION", "case", "redis-timeout", "evidenceCoverage.query_logs"));
assertTrue(hasItem(diff, "CHANGED", "aggregate", null, "verdictDistribution.LOW_CONFID"));
assertTrue(hasItem(diff, "CHANGED", "aggregate", null, "verdictDistribution.REJECT"));
}
@Test
void compareReportsDetectsMissingAndNewCases() throws Exception {
DiagnosisEvalReport baseline = readBaselineReport();
DiagnosisEvalReport current = readBaselineReport();
DiagnosisEvalResult removed = current.getResults().remove(0);
current.getResults().add(DiagnosisEvalResult.builder()
.caseId("new-case")
.title("New case")
.passed(true)
.failedChecks(List.of())
.verdict("PASS")
.matchedKeywordCount(1)
.requiredKeywordCount(1)
.evidenceCoverage(new LinkedHashMap<>())
.toolCallCount(1)
.durationMs(1000)
.build());
DiagnosisEvalDiffReport diff = differ.compare(baseline, current);
assertTrue(hasItem(diff, "REGRESSION", "case", removed.getCaseId(), "casePresence"));
assertTrue(hasItem(diff, "CHANGED", "case", "new-case", "casePresence"));
}
@Test
void compareSameReportHasNoDiff() throws Exception {
DiagnosisEvalReport baseline = readBaselineReport();
DiagnosisEvalDiffReport diff = differ.compare(baseline, readBaselineReport());
assertFalse(diff.isHasRegression());
assertEquals(0, diff.getRegressionCount());
assertTrue(diff.getItems().isEmpty());
}
@Test
void writerOutputsJsonAndMarkdown(@TempDir Path tempDir) throws Exception {
DiagnosisEvalReport baseline = readBaselineReport();
DiagnosisEvalReport current = readBaselineReport();
degradeRedisCase(current);
DiagnosisEvalDiffReport diff = differ.compare(baseline, current);
DiagnosisEvalDiffReportWriter writer = new DiagnosisEvalDiffReportWriter(objectMapper);
Path json = tempDir.resolve("baseline-diff.json");
Path markdown = tempDir.resolve("baseline-diff.md");
writer.writeJson(diff, json);
writer.writeMarkdown(diff, markdown);
assertTrue(Files.exists(json));
assertTrue(Files.readString(json).contains("\"hasRegression\" : true"));
assertTrue(Files.readString(markdown).contains("# Diagnosis Eval Baseline Diff"));
assertTrue(Files.readString(markdown).contains("redis-timeout"));
}
private DiagnosisEvalReport readBaselineReport() throws Exception {
return objectMapper.readValue(Path.of("mvp/eval/reports/baseline-report.json").toFile(),
DiagnosisEvalReport.class);
}
private void degradeRedisCase(DiagnosisEvalReport report) {
report.setPassedCases(4);
report.setPassRate(0.8);
report.setAverageToolCallCount(3.0);
report.setAverageDurationMs(45800.0);
report.setVerdictDistribution(new LinkedHashMap<>());
report.getVerdictDistribution().put("PASS", 2L);
report.getVerdictDistribution().put("LOW_CONFID", 2L);
report.getVerdictDistribution().put("REJECT", 1L);
DiagnosisEvalResult redis = result(report, "redis-timeout");
redis.setPassed(false);
redis.setFailedChecks(new ArrayList<>(List.of("missing required evidence tool: query_logs")));
redis.setVerdict("REJECT");
redis.setMatchedKeywordCount(1);
redis.getEvidenceCoverage().put("query_logs", false);
redis.setToolCallCount(1);
redis.setDurationMs(36000);
}
private DiagnosisEvalResult result(DiagnosisEvalReport report, String caseId) {
return report.getResults().stream()
.filter(item -> caseId.equals(item.getCaseId()))
.findFirst()
.orElseThrow();
}
private boolean hasItem(DiagnosisEvalDiffReport diff,
String type,
String scope,
String caseId,
String metric) {
return diff.getItems().stream().anyMatch(item ->
type.equals(item.getType())
&& scope.equals(item.getScope())
&& java.util.Objects.equals(caseId, item.getCaseId())
&& metric.equals(item.getMetric()));
}
}