Add diagnosis eval baseline diff
This commit is contained in:
@@ -4,6 +4,7 @@
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: diagnosis-eval-baseline-diff
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
|
||||
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
|
||||
- JSON and Markdown diff output are available.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
|
||||
- Result: passed
|
||||
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,30 @@
|
||||
# Brief: diagnosis-eval-baseline-diff
|
||||
|
||||
## Background
|
||||
|
||||
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Compare baseline and current `DiagnosisEvalReport` objects.
|
||||
2. Detect aggregate and per-case regressions.
|
||||
3. Output JSON and Markdown diff reports.
|
||||
4. Document how to read the diff in interview and engineering terms.
|
||||
|
||||
## Scope
|
||||
|
||||
- Diff data structures
|
||||
- Deterministic report comparison
|
||||
- JSON / Markdown diff output
|
||||
- Focused tests and eval docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No live Agent execution
|
||||
- No LLM-as-judge
|
||||
- No evaluator scoring rule changes
|
||||
- No production API changes
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-baseline-diff/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Baseline Diff Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
|
||||
- Slug: `diagnosis-eval-baseline-diff`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created deterministic fixture evaluation.
|
||||
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
|
||||
- This change compares new reports against that baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Diff report DTOs instead of raw traces.
|
||||
- Reason: the report is the stable contract for regression review.
|
||||
|
||||
- Decision: Use deterministic code rules instead of LLM-as-judge.
|
||||
- Reason: baseline regression checks should be repeatable and explainable.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should expose this through a CLI or Maven goal.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: diagnosis-eval-baseline-diff
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
|
||||
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
|
||||
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
|
||||
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
|
||||
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
|
||||
@@ -8,6 +8,7 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
|
||||
- Offline trace fixtures: `fixtures/*.json`
|
||||
- Field definitions: `schema.md`
|
||||
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
|
||||
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
|
||||
- Evaluator implementation: `DiagnosisTraceEvaluator`
|
||||
- Report writer: `DiagnosisEvalReportWriter`
|
||||
|
||||
@@ -45,3 +46,16 @@ fixed diagnosis case
|
||||
-> JSON / Markdown report
|
||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
||||
```
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||
|
||||
```text
|
||||
baseline report
|
||||
current report
|
||||
-> deterministic diff
|
||||
-> regressions, improvements, and changed signals
|
||||
```
|
||||
|
||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
||||
|
||||
@@ -0,0 +1,85 @@
|
||||
{
|
||||
"baselineTotalCases" : 5,
|
||||
"currentTotalCases" : 5,
|
||||
"baselinePassedCases" : 5,
|
||||
"currentPassedCases" : 4,
|
||||
"baselinePassRate" : 1.0,
|
||||
"currentPassRate" : 0.8,
|
||||
"regressionCount" : 6,
|
||||
"improvementCount" : 0,
|
||||
"changedCount" : 2,
|
||||
"hasRegression" : true,
|
||||
"items" : [ {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "passRate",
|
||||
"baselineValue" : "1.0",
|
||||
"currentValue" : "0.8",
|
||||
"delta" : -0.19999999999999996,
|
||||
"message" : "passRate changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "averageToolCallCount",
|
||||
"baselineValue" : "2.0",
|
||||
"currentValue" : "3.0",
|
||||
"delta" : 1.0,
|
||||
"message" : "averageToolCallCount changed"
|
||||
}, {
|
||||
"type" : "CHANGED",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "verdictDistribution.LOW_CONFID",
|
||||
"baselineValue" : "3",
|
||||
"currentValue" : "2",
|
||||
"delta" : -1.0,
|
||||
"message" : "verdict count changed for LOW_CONFID"
|
||||
}, {
|
||||
"type" : "CHANGED",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "verdictDistribution.REJECT",
|
||||
"baselineValue" : "0",
|
||||
"currentValue" : "1",
|
||||
"delta" : 1.0,
|
||||
"message" : "verdict count changed for REJECT"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "passed",
|
||||
"baselineValue" : "true",
|
||||
"currentValue" : "false",
|
||||
"delta" : null,
|
||||
"message" : "redis-timeout pass state changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "verdict",
|
||||
"baselineValue" : "LOW_CONFID",
|
||||
"currentValue" : "REJECT",
|
||||
"delta" : -1.0,
|
||||
"message" : "redis-timeout verdict changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "matchedKeywordCount",
|
||||
"baselineValue" : "2",
|
||||
"currentValue" : "1",
|
||||
"delta" : -1.0,
|
||||
"message" : "redis-timeout matchedKeywordCount changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "evidenceCoverage.query_logs",
|
||||
"baselineValue" : "true",
|
||||
"currentValue" : "false",
|
||||
"delta" : null,
|
||||
"message" : "redis-timeout evidence coverage changed for query_logs"
|
||||
} ]
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
# Diagnosis Eval Baseline Diff
|
||||
|
||||
- Baseline pass rate: 100.00%
|
||||
- Current pass rate: 80.00%
|
||||
- Baseline passed cases: 5/5
|
||||
- Current passed cases: 4/5
|
||||
- Regressions: 6
|
||||
- Improvements: 0
|
||||
- Other changes: 2
|
||||
|
||||
## Diff Items
|
||||
|
||||
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
|
||||
| --- | --- | --- | --- | --- | --- | ---: | --- |
|
||||
| REGRESSION | aggregate | - | passRate | 1.0 | 0.8 | -0.200 | passRate changed |
|
||||
| REGRESSION | aggregate | - | averageToolCallCount | 2.0 | 3.0 | 1.000 | averageToolCallCount changed |
|
||||
| CHANGED | aggregate | - | verdictDistribution.LOW_CONFID | 3 | 2 | -1.000 | verdict count changed for LOW_CONFID |
|
||||
| CHANGED | aggregate | - | verdictDistribution.REJECT | 0 | 1 | 1.000 | verdict count changed for REJECT |
|
||||
| REGRESSION | case | redis-timeout | passed | true | false | - | redis-timeout pass state changed |
|
||||
| REGRESSION | case | redis-timeout | verdict | LOW_CONFID | REJECT | -1.000 | redis-timeout verdict changed |
|
||||
| REGRESSION | case | redis-timeout | matchedKeywordCount | 2 | 1 | -1.000 | redis-timeout matchedKeywordCount changed |
|
||||
| REGRESSION | case | redis-timeout | evidenceCoverage.query_logs | true | false | - | redis-timeout evidence coverage changed for query_logs |
|
||||
@@ -135,3 +135,68 @@ Java 类型:`DiagnosisEvalReport`
|
||||
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
|
||||
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
|
||||
```
|
||||
|
||||
## 6. Baseline Diff
|
||||
|
||||
Baseline diff 是拿两份 report 做对比:
|
||||
|
||||
```text
|
||||
baseline report:以前认可的基准结果
|
||||
current report:这次改动后跑出来的新结果
|
||||
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
|
||||
```
|
||||
|
||||
Java 类型:
|
||||
|
||||
- `DiagnosisEvalDiffReport`
|
||||
- `DiagnosisEvalDiffItem`
|
||||
|
||||
`DiagnosisEvalDiffReport` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `baselineTotalCases` | baseline 里有多少条 case |
|
||||
| `currentTotalCases` | current 里有多少条 case |
|
||||
| `baselinePassedCases` | baseline 通过了多少条 |
|
||||
| `currentPassedCases` | current 通过了多少条 |
|
||||
| `baselinePassRate` | baseline 通过率 |
|
||||
| `currentPassRate` | current 通过率 |
|
||||
| `regressionCount` | 退化项数量 |
|
||||
| `improvementCount` | 改善项数量 |
|
||||
| `changedCount` | 普通变化项数量 |
|
||||
| `hasRegression` | 是否存在退化 |
|
||||
| `items` | 具体 diff 明细 |
|
||||
|
||||
`DiagnosisEvalDiffItem` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
|
||||
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
|
||||
| `caseId` | 如果是单条 case 变化,这里记录 case id |
|
||||
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
|
||||
| `baselineValue` | baseline 里的值 |
|
||||
| `currentValue` | current 里的值 |
|
||||
| `delta` | 数值变化量;非数值变化为空 |
|
||||
| `message` | 给人看的变化说明 |
|
||||
|
||||
口语化判断规则:
|
||||
|
||||
```text
|
||||
pass rate 下降:退化
|
||||
case 从通过变失败:退化
|
||||
证据工具从有变没有:退化
|
||||
关键词命中变少:退化
|
||||
工具调用或耗时升高:成本上升,记为退化信号
|
||||
verdict 分布变化:记录变化,供人工判断是否符合预期
|
||||
```
|
||||
|
||||
面试里可以这样讲:
|
||||
|
||||
```text
|
||||
我把 baseline report 和当前 report 做结构化 diff。
|
||||
它不是再问 LLM,而是用代码比较固定字段。
|
||||
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
|
||||
diff 会直接标成 regression。
|
||||
这样 Agent 改动可以用固定基准做回归判断。
|
||||
```
|
||||
|
||||
@@ -9,3 +9,4 @@
|
||||
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
# Diagnosis Eval Baseline Diff
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**来源**:P1-B follow-up
|
||||
**依赖**:`diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
现在项目已经有固定诊断 case、完整 fixture 和 baseline report。下一步需要把 baseline 真正用起来:每次改 Agent 后,把新的 report 和 baseline report 做对比。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 baseline 只能告诉我们“标准状态是什么”,但还不能自动告诉我们“这次改动有没有变差”。
|
||||
|
||||
典型问题包括:
|
||||
|
||||
- pass rate 是否下降。
|
||||
- 某个 case 是否从通过变失败。
|
||||
- 某个 evidence tool 是否从覆盖变成缺失。
|
||||
- verifier verdict 分布是否异常变化。
|
||||
- 平均工具调用数和耗时是否明显上升。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
新增一个 deterministic baseline diff 能力,用代码比较两份 `DiagnosisEvalReport`。
|
||||
|
||||
完成后应该做到:
|
||||
|
||||
- 输入 baseline report 和 current report。
|
||||
- 输出结构化 diff。
|
||||
- 标出 regression、improvement 和普通 changed。
|
||||
- 支持 JSON 和 Markdown 输出。
|
||||
- 文档说明面试时怎么解释这套回归判断。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- report-level diff 数据结构。
|
||||
- aggregate 指标比较。
|
||||
- case-level 指标比较。
|
||||
- JSON / Markdown diff writer。
|
||||
- focused tests 和 eval 文档。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不运行真实 Agent。
|
||||
- 不生成新 trace。
|
||||
- 不引入 LLM-as-judge。
|
||||
- 不改现有 evaluator 评分规则。
|
||||
|
||||
---
|
||||
|
||||
## 面试表达
|
||||
|
||||
可以这样讲:
|
||||
|
||||
```text
|
||||
我不是只保存了一份 baseline,而是加了 baseline diff。
|
||||
每次改 prompt、tool、retrieval 或 verifier 后,
|
||||
我都能把新 report 和 baseline 比较,
|
||||
直接看到哪些 case 退化、哪些证据缺失、成本有没有上升。
|
||||
```
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,39 @@
|
||||
## Context
|
||||
|
||||
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Compare two `DiagnosisEvalReport` objects without requiring external services.
|
||||
- Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
|
||||
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
|
||||
- Write JSON and Markdown diff outputs for review.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not run the Agent or regenerate traces.
|
||||
- Do not introduce LLM-as-judge.
|
||||
- Do not change evaluator scoring rules.
|
||||
- Do not block on performance thresholds beyond simple numeric diff signals.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Decision: Compare report DTOs instead of raw traces.
|
||||
- Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
|
||||
- Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
|
||||
|
||||
- Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`.
|
||||
- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
|
||||
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
|
||||
|
||||
- Decision: Keep thresholds explicit and conservative.
|
||||
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
|
||||
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed.
|
||||
- Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
|
||||
- Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.
|
||||
@@ -0,0 +1,28 @@
|
||||
## Why
|
||||
|
||||
The evaluation baseline is now complete, but developers still need a repeatable way to decide whether a new Agent run regressed against that baseline. A deterministic baseline diff turns saved reports into an actionable regression signal instead of a static artifact.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a baseline diff model that compares two `DiagnosisEvalReport` objects.
|
||||
- Detect aggregate changes such as pass-rate drops, verdict distribution shifts, tool-call cost changes, and duration changes.
|
||||
- Detect per-case changes such as pass/fail regression, verdict changes, keyword coverage changes, evidence coverage loss, and missing/new cases.
|
||||
- Add JSON and Markdown diff output suitable for review.
|
||||
- Document how to interpret the diff in the eval docs.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness so a current report can be compared against the saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects eval-only Java code under `src/main/java/com/superbiz/agent/eval`.
|
||||
- Adds focused tests under `src/test/java/com/superbiz/agent/eval`.
|
||||
- Updates `mvp/eval` documentation and may add sample diff output.
|
||||
- No production Agent runtime, API, database schema, or external dependency changes are expected.
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
||||
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
||||
|
||||
#### Scenario: Aggregate regression detection
|
||||
- **WHEN** the current report has a lower pass rate than the baseline report
|
||||
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
||||
|
||||
#### Scenario: Cost signal detection
|
||||
- **WHEN** average tool-call count or average duration changes between reports
|
||||
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
||||
|
||||
#### Scenario: Verdict distribution comparison
|
||||
- **WHEN** verdict counts differ between reports
|
||||
- **THEN** the diff SHALL record the verdict distribution changes
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare case-level report results
|
||||
The system SHALL compare case results by case id and report actionable per-case changes.
|
||||
|
||||
#### Scenario: Case pass/fail regression
|
||||
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
||||
- **THEN** the diff SHALL record a regression for that case
|
||||
|
||||
#### Scenario: Evidence coverage regression
|
||||
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
||||
- **THEN** the diff SHALL record a regression naming the case and tool
|
||||
|
||||
#### Scenario: Missing case detection
|
||||
- **WHEN** a baseline case is absent from the current report
|
||||
- **THEN** the diff SHALL record a regression for the missing case
|
||||
|
||||
#### Scenario: New case detection
|
||||
- **WHEN** a current report contains a case absent from the baseline
|
||||
- **THEN** the diff SHALL record the case as a non-regression change
|
||||
|
||||
### Requirement: Evaluation harness SHALL report baseline diff results
|
||||
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
||||
|
||||
#### Scenario: JSON diff output
|
||||
- **WHEN** a baseline diff is written as JSON
|
||||
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
||||
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
@@ -0,0 +1,23 @@
|
||||
## 1. OpenSpec And Issue Setup
|
||||
|
||||
- [x] 1.1 Create slug-based issue and devflow tracking files.
|
||||
- [x] 1.2 Create OpenSpec proposal, design, delta spec, and tasks.
|
||||
|
||||
## 2. Baseline Diff Implementation
|
||||
|
||||
- [x] 2.1 Add diff result data structures for summary and per-item changes.
|
||||
- [x] 2.2 Implement deterministic report comparison rules.
|
||||
- [x] 2.3 Implement JSON and Markdown diff report writing.
|
||||
|
||||
## 3. Documentation
|
||||
|
||||
- [x] 3.1 Document baseline diff inputs, outputs, and interpretation in eval docs.
|
||||
- [x] 3.2 Add sample diff output for a representative regression.
|
||||
|
||||
## 4. Tests And Validation
|
||||
|
||||
- [x] 4.1 Add focused tests for aggregate and case-level diff behavior.
|
||||
- [x] 4.2 Add focused tests for JSON and Markdown diff output.
|
||||
- [x] 4.3 Run evaluator/diff tests.
|
||||
- [x] 4.4 Run compile verification.
|
||||
- [x] 4.5 Run OpenSpec validation.
|
||||
@@ -88,3 +88,48 @@ The system SHALL preserve a generated baseline report for the full fixed fixture
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes fixtures or evaluator rules
|
||||
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
||||
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
||||
|
||||
#### Scenario: Aggregate regression detection
|
||||
- **WHEN** the current report has a lower pass rate than the baseline report
|
||||
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
||||
|
||||
#### Scenario: Cost signal detection
|
||||
- **WHEN** average tool-call count or average duration changes between reports
|
||||
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
||||
|
||||
#### Scenario: Verdict distribution comparison
|
||||
- **WHEN** verdict counts differ between reports
|
||||
- **THEN** the diff SHALL record the verdict distribution changes
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare case-level report results
|
||||
The system SHALL compare case results by case id and report actionable per-case changes.
|
||||
|
||||
#### Scenario: Case pass/fail regression
|
||||
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
||||
- **THEN** the diff SHALL record a regression for that case
|
||||
|
||||
#### Scenario: Evidence coverage regression
|
||||
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
||||
- **THEN** the diff SHALL record a regression naming the case and tool
|
||||
|
||||
#### Scenario: Missing case detection
|
||||
- **WHEN** a baseline case is absent from the current report
|
||||
- **THEN** the diff SHALL record a regression for the missing case
|
||||
|
||||
#### Scenario: New case detection
|
||||
- **WHEN** a current report contains a case absent from the baseline
|
||||
- **THEN** the diff SHALL record the case as a non-regression change
|
||||
|
||||
### Requirement: Evaluation harness SHALL report baseline diff results
|
||||
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
||||
|
||||
#### Scenario: JSON diff output
|
||||
- **WHEN** a baseline diff is written as JSON
|
||||
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
||||
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
|
||||
@@ -0,0 +1,252 @@
|
||||
package com.superbiz.agent.eval;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Comparator;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Set;
|
||||
import java.util.function.Function;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
public class DiagnosisEvalBaselineDiffer {
|
||||
|
||||
private static final String REGRESSION = "REGRESSION";
|
||||
private static final String IMPROVEMENT = "IMPROVEMENT";
|
||||
private static final String CHANGED = "CHANGED";
|
||||
|
||||
public DiagnosisEvalDiffReport compare(DiagnosisEvalReport baseline, DiagnosisEvalReport current) {
|
||||
List<DiagnosisEvalDiffItem> items = new ArrayList<>();
|
||||
|
||||
compareDouble(items, "aggregate", null, "passRate",
|
||||
baseline.getPassRate(), current.getPassRate(), true);
|
||||
compareDouble(items, "aggregate", null, "averageToolCallCount",
|
||||
baseline.getAverageToolCallCount(), current.getAverageToolCallCount(), false);
|
||||
compareDouble(items, "aggregate", null, "averageDurationMs",
|
||||
baseline.getAverageDurationMs(), current.getAverageDurationMs(), false);
|
||||
compareVerdictDistribution(items, baseline.getVerdictDistribution(), current.getVerdictDistribution());
|
||||
compareCases(items, safeResults(baseline), safeResults(current));
|
||||
|
||||
int regressionCount = countType(items, REGRESSION);
|
||||
int improvementCount = countType(items, IMPROVEMENT);
|
||||
int changedCount = countType(items, CHANGED);
|
||||
|
||||
return DiagnosisEvalDiffReport.builder()
|
||||
.baselineTotalCases(baseline.getTotalCases())
|
||||
.currentTotalCases(current.getTotalCases())
|
||||
.baselinePassedCases(baseline.getPassedCases())
|
||||
.currentPassedCases(current.getPassedCases())
|
||||
.baselinePassRate(baseline.getPassRate())
|
||||
.currentPassRate(current.getPassRate())
|
||||
.regressionCount(regressionCount)
|
||||
.improvementCount(improvementCount)
|
||||
.changedCount(changedCount)
|
||||
.hasRegression(regressionCount > 0)
|
||||
.items(items)
|
||||
.build();
|
||||
}
|
||||
|
||||
private void compareVerdictDistribution(List<DiagnosisEvalDiffItem> items,
|
||||
Map<String, Long> baseline,
|
||||
Map<String, Long> current) {
|
||||
Set<String> verdicts = new LinkedHashSet<>();
|
||||
verdicts.addAll(safeMap(baseline).keySet());
|
||||
verdicts.addAll(safeMap(current).keySet());
|
||||
for (String verdict : verdicts) {
|
||||
long baselineCount = safeMap(baseline).getOrDefault(verdict, 0L);
|
||||
long currentCount = safeMap(current).getOrDefault(verdict, 0L);
|
||||
if (baselineCount != currentCount) {
|
||||
items.add(item(CHANGED, "aggregate", null, "verdictDistribution." + verdict,
|
||||
String.valueOf(baselineCount), String.valueOf(currentCount),
|
||||
(double) currentCount - baselineCount,
|
||||
"verdict count changed for " + verdict));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private void compareCases(List<DiagnosisEvalDiffItem> items,
|
||||
List<DiagnosisEvalResult> baselineResults,
|
||||
List<DiagnosisEvalResult> currentResults) {
|
||||
Map<String, DiagnosisEvalResult> baselineById = byCaseId(baselineResults);
|
||||
Map<String, DiagnosisEvalResult> currentById = byCaseId(currentResults);
|
||||
Set<String> caseIds = new LinkedHashSet<>();
|
||||
caseIds.addAll(baselineById.keySet());
|
||||
caseIds.addAll(currentById.keySet());
|
||||
|
||||
for (String caseId : caseIds) {
|
||||
DiagnosisEvalResult baseline = baselineById.get(caseId);
|
||||
DiagnosisEvalResult current = currentById.get(caseId);
|
||||
if (baseline == null) {
|
||||
items.add(item(CHANGED, "case", caseId, "casePresence",
|
||||
"missing", "present", null, "new case appears in current report"));
|
||||
continue;
|
||||
}
|
||||
if (current == null) {
|
||||
items.add(item(REGRESSION, "case", caseId, "casePresence",
|
||||
"present", "missing", null, "baseline case is missing from current report"));
|
||||
continue;
|
||||
}
|
||||
|
||||
comparePassState(items, baseline, current);
|
||||
compareVerdict(items, baseline, current);
|
||||
compareInteger(items, caseId, "matchedKeywordCount",
|
||||
baseline.getMatchedKeywordCount(), current.getMatchedKeywordCount(), true);
|
||||
compareInteger(items, caseId, "toolCallCount",
|
||||
baseline.getToolCallCount(), current.getToolCallCount(), false);
|
||||
compareInteger(items, caseId, "durationMs",
|
||||
baseline.getDurationMs(), current.getDurationMs(), false);
|
||||
compareEvidenceCoverage(items, baseline, current);
|
||||
}
|
||||
}
|
||||
|
||||
private void comparePassState(List<DiagnosisEvalDiffItem> items,
|
||||
DiagnosisEvalResult baseline,
|
||||
DiagnosisEvalResult current) {
|
||||
if (baseline.isPassed() == current.isPassed()) {
|
||||
return;
|
||||
}
|
||||
String type = baseline.isPassed() ? REGRESSION : IMPROVEMENT;
|
||||
items.add(item(type, "case", baseline.getCaseId(), "passed",
|
||||
String.valueOf(baseline.isPassed()), String.valueOf(current.isPassed()), null,
|
||||
baseline.getCaseId() + " pass state changed"));
|
||||
}
|
||||
|
||||
private void compareVerdict(List<DiagnosisEvalDiffItem> items,
|
||||
DiagnosisEvalResult baseline,
|
||||
DiagnosisEvalResult current) {
|
||||
if (Objects.equals(baseline.getVerdict(), current.getVerdict())) {
|
||||
return;
|
||||
}
|
||||
int baselineRank = verdictRank(baseline.getVerdict());
|
||||
int currentRank = verdictRank(current.getVerdict());
|
||||
String type = currentRank < baselineRank ? REGRESSION : currentRank > baselineRank ? IMPROVEMENT : CHANGED;
|
||||
items.add(item(type, "case", baseline.getCaseId(), "verdict",
|
||||
value(baseline.getVerdict()), value(current.getVerdict()), (double) currentRank - baselineRank,
|
||||
baseline.getCaseId() + " verdict changed"));
|
||||
}
|
||||
|
||||
private void compareEvidenceCoverage(List<DiagnosisEvalDiffItem> items,
|
||||
DiagnosisEvalResult baseline,
|
||||
DiagnosisEvalResult current) {
|
||||
Set<String> tools = new LinkedHashSet<>();
|
||||
tools.addAll(safeMap(baseline.getEvidenceCoverage()).keySet());
|
||||
tools.addAll(safeMap(current.getEvidenceCoverage()).keySet());
|
||||
for (String tool : tools) {
|
||||
boolean baselineCovered = Boolean.TRUE.equals(safeMap(baseline.getEvidenceCoverage()).get(tool));
|
||||
boolean currentCovered = Boolean.TRUE.equals(safeMap(current.getEvidenceCoverage()).get(tool));
|
||||
if (baselineCovered == currentCovered) {
|
||||
continue;
|
||||
}
|
||||
String type = baselineCovered ? REGRESSION : IMPROVEMENT;
|
||||
items.add(item(type, "case", baseline.getCaseId(), "evidenceCoverage." + tool,
|
||||
String.valueOf(baselineCovered), String.valueOf(currentCovered), null,
|
||||
baseline.getCaseId() + " evidence coverage changed for " + tool));
|
||||
}
|
||||
}
|
||||
|
||||
private void compareDouble(List<DiagnosisEvalDiffItem> items,
|
||||
String scope,
|
||||
String caseId,
|
||||
String metric,
|
||||
double baseline,
|
||||
double current,
|
||||
boolean higherIsBetter) {
|
||||
if (Double.compare(baseline, current) == 0) {
|
||||
return;
|
||||
}
|
||||
double delta = current - baseline;
|
||||
String type = classifyDelta(delta, higherIsBetter);
|
||||
items.add(item(type, scope, caseId, metric,
|
||||
String.valueOf(baseline), String.valueOf(current), delta,
|
||||
metric + " changed"));
|
||||
}
|
||||
|
||||
private void compareInteger(List<DiagnosisEvalDiffItem> items,
|
||||
String caseId,
|
||||
String metric,
|
||||
Integer baseline,
|
||||
Integer current,
|
||||
boolean higherIsBetter) {
|
||||
if (Objects.equals(baseline, current)) {
|
||||
return;
|
||||
}
|
||||
if (baseline == null || current == null) {
|
||||
items.add(item(CHANGED, "case", caseId, metric,
|
||||
value(baseline), value(current), null, caseId + " " + metric + " changed"));
|
||||
return;
|
||||
}
|
||||
int delta = current - baseline;
|
||||
items.add(item(classifyDelta(delta, higherIsBetter), "case", caseId, metric,
|
||||
String.valueOf(baseline), String.valueOf(current), (double) delta,
|
||||
caseId + " " + metric + " changed"));
|
||||
}
|
||||
|
||||
private String classifyDelta(double delta, boolean higherIsBetter) {
|
||||
if (delta == 0.0) {
|
||||
return CHANGED;
|
||||
}
|
||||
boolean improved = higherIsBetter ? delta > 0 : delta < 0;
|
||||
return improved ? IMPROVEMENT : REGRESSION;
|
||||
}
|
||||
|
||||
private DiagnosisEvalDiffItem item(String type,
|
||||
String scope,
|
||||
String caseId,
|
||||
String metric,
|
||||
String baselineValue,
|
||||
String currentValue,
|
||||
Double delta,
|
||||
String message) {
|
||||
return DiagnosisEvalDiffItem.builder()
|
||||
.type(type)
|
||||
.scope(scope)
|
||||
.caseId(caseId)
|
||||
.metric(metric)
|
||||
.baselineValue(baselineValue)
|
||||
.currentValue(currentValue)
|
||||
.delta(delta)
|
||||
.message(message)
|
||||
.build();
|
||||
}
|
||||
|
||||
private Map<String, DiagnosisEvalResult> byCaseId(List<DiagnosisEvalResult> results) {
|
||||
return results.stream()
|
||||
.sorted(Comparator.comparing(DiagnosisEvalResult::getCaseId))
|
||||
.collect(Collectors.toMap(
|
||||
DiagnosisEvalResult::getCaseId,
|
||||
Function.identity(),
|
||||
(left, right) -> right,
|
||||
LinkedHashMap::new));
|
||||
}
|
||||
|
||||
private List<DiagnosisEvalResult> safeResults(DiagnosisEvalReport report) {
|
||||
return report.getResults() == null ? List.of() : report.getResults();
|
||||
}
|
||||
|
||||
private <T> Map<String, T> safeMap(Map<String, T> value) {
|
||||
return value == null ? Map.of() : value;
|
||||
}
|
||||
|
||||
private int countType(List<DiagnosisEvalDiffItem> items, String type) {
|
||||
return (int) items.stream().filter(item -> type.equals(item.getType())).count();
|
||||
}
|
||||
|
||||
private int verdictRank(String verdict) {
|
||||
if ("PASS".equals(verdict)) {
|
||||
return 3;
|
||||
}
|
||||
if ("LOW_CONFID".equals(verdict)) {
|
||||
return 2;
|
||||
}
|
||||
if ("REJECT".equals(verdict)) {
|
||||
return 1;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
private String value(Object value) {
|
||||
return value == null ? "-" : String.valueOf(value);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
package com.superbiz.agent.eval;
|
||||
|
||||
import lombok.AllArgsConstructor;
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
import lombok.NoArgsConstructor;
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public class DiagnosisEvalDiffItem {
|
||||
|
||||
private String type;
|
||||
private String scope;
|
||||
private String caseId;
|
||||
private String metric;
|
||||
private String baselineValue;
|
||||
private String currentValue;
|
||||
private Double delta;
|
||||
private String message;
|
||||
}
|
||||
@@ -0,0 +1,27 @@
|
||||
package com.superbiz.agent.eval;
|
||||
|
||||
import lombok.AllArgsConstructor;
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
import lombok.NoArgsConstructor;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public class DiagnosisEvalDiffReport {
|
||||
|
||||
private int baselineTotalCases;
|
||||
private int currentTotalCases;
|
||||
private int baselinePassedCases;
|
||||
private int currentPassedCases;
|
||||
private double baselinePassRate;
|
||||
private double currentPassRate;
|
||||
private int regressionCount;
|
||||
private int improvementCount;
|
||||
private int changedCount;
|
||||
private boolean hasRegression;
|
||||
private List<DiagnosisEvalDiffItem> items;
|
||||
}
|
||||
@@ -0,0 +1,78 @@
|
||||
package com.superbiz.agent.eval;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
public class DiagnosisEvalDiffReportWriter {
|
||||
|
||||
private final ObjectMapper objectMapper;
|
||||
|
||||
public DiagnosisEvalDiffReportWriter(ObjectMapper objectMapper) {
|
||||
this.objectMapper = objectMapper;
|
||||
}
|
||||
|
||||
public void writeJson(DiagnosisEvalDiffReport report, Path outputFile) throws IOException {
|
||||
Files.createDirectories(outputFile.getParent());
|
||||
objectMapper.writerWithDefaultPrettyPrinter().writeValue(outputFile.toFile(), report);
|
||||
}
|
||||
|
||||
public void writeMarkdown(DiagnosisEvalDiffReport report, Path outputFile) throws IOException {
|
||||
Files.createDirectories(outputFile.getParent());
|
||||
Files.writeString(outputFile, toMarkdown(report), StandardCharsets.UTF_8);
|
||||
}
|
||||
|
||||
public String toMarkdown(DiagnosisEvalDiffReport report) {
|
||||
StringBuilder builder = new StringBuilder();
|
||||
builder.append("# Diagnosis Eval Baseline Diff\n\n");
|
||||
builder.append("- Baseline pass rate: ").append(formatPercent(report.getBaselinePassRate())).append("\n");
|
||||
builder.append("- Current pass rate: ").append(formatPercent(report.getCurrentPassRate())).append("\n");
|
||||
builder.append("- Baseline passed cases: ").append(report.getBaselinePassedCases()).append("/")
|
||||
.append(report.getBaselineTotalCases()).append("\n");
|
||||
builder.append("- Current passed cases: ").append(report.getCurrentPassedCases()).append("/")
|
||||
.append(report.getCurrentTotalCases()).append("\n");
|
||||
builder.append("- Regressions: ").append(report.getRegressionCount()).append("\n");
|
||||
builder.append("- Improvements: ").append(report.getImprovementCount()).append("\n");
|
||||
builder.append("- Other changes: ").append(report.getChangedCount()).append("\n\n");
|
||||
|
||||
builder.append("## Diff Items\n\n");
|
||||
if (report.getItems() == null || report.getItems().isEmpty()) {
|
||||
builder.append("- No differences\n");
|
||||
return builder.toString();
|
||||
}
|
||||
|
||||
builder.append("| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |\n");
|
||||
builder.append("| --- | --- | --- | --- | --- | --- | ---: | --- |\n");
|
||||
for (DiagnosisEvalDiffItem item : report.getItems()) {
|
||||
builder.append("| ")
|
||||
.append(valueOrDash(item.getType()))
|
||||
.append(" | ")
|
||||
.append(valueOrDash(item.getScope()))
|
||||
.append(" | ")
|
||||
.append(valueOrDash(item.getCaseId()))
|
||||
.append(" | ")
|
||||
.append(valueOrDash(item.getMetric()))
|
||||
.append(" | ")
|
||||
.append(valueOrDash(item.getBaselineValue()))
|
||||
.append(" | ")
|
||||
.append(valueOrDash(item.getCurrentValue()))
|
||||
.append(" | ")
|
||||
.append(item.getDelta() == null ? "-" : String.format("%.3f", item.getDelta()))
|
||||
.append(" | ")
|
||||
.append(valueOrDash(item.getMessage()))
|
||||
.append(" |\n");
|
||||
}
|
||||
return builder.toString();
|
||||
}
|
||||
|
||||
private String formatPercent(double value) {
|
||||
return String.format("%.2f%%", value * 100);
|
||||
}
|
||||
|
||||
private String valueOrDash(String value) {
|
||||
return value == null || value.isBlank() ? "-" : value;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,140 @@
|
||||
package com.superbiz.agent.eval;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
class DiagnosisEvalBaselineDiffTest {
|
||||
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
private final DiagnosisEvalBaselineDiffer differ = new DiagnosisEvalBaselineDiffer();
|
||||
|
||||
@Test
|
||||
void compareReportsDetectsAggregateAndCaseRegressions() throws Exception {
|
||||
DiagnosisEvalReport baseline = readBaselineReport();
|
||||
DiagnosisEvalReport current = readBaselineReport();
|
||||
degradeRedisCase(current);
|
||||
|
||||
DiagnosisEvalDiffReport diff = differ.compare(baseline, current);
|
||||
|
||||
assertTrue(diff.isHasRegression());
|
||||
assertEquals(6, diff.getRegressionCount());
|
||||
assertEquals(2, diff.getChangedCount());
|
||||
assertTrue(hasItem(diff, "REGRESSION", "aggregate", null, "passRate"));
|
||||
assertTrue(hasItem(diff, "REGRESSION", "aggregate", null, "averageToolCallCount"));
|
||||
assertTrue(hasItem(diff, "REGRESSION", "case", "redis-timeout", "passed"));
|
||||
assertTrue(hasItem(diff, "REGRESSION", "case", "redis-timeout", "verdict"));
|
||||
assertTrue(hasItem(diff, "REGRESSION", "case", "redis-timeout", "matchedKeywordCount"));
|
||||
assertTrue(hasItem(diff, "REGRESSION", "case", "redis-timeout", "evidenceCoverage.query_logs"));
|
||||
assertTrue(hasItem(diff, "CHANGED", "aggregate", null, "verdictDistribution.LOW_CONFID"));
|
||||
assertTrue(hasItem(diff, "CHANGED", "aggregate", null, "verdictDistribution.REJECT"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void compareReportsDetectsMissingAndNewCases() throws Exception {
|
||||
DiagnosisEvalReport baseline = readBaselineReport();
|
||||
DiagnosisEvalReport current = readBaselineReport();
|
||||
DiagnosisEvalResult removed = current.getResults().remove(0);
|
||||
current.getResults().add(DiagnosisEvalResult.builder()
|
||||
.caseId("new-case")
|
||||
.title("New case")
|
||||
.passed(true)
|
||||
.failedChecks(List.of())
|
||||
.verdict("PASS")
|
||||
.matchedKeywordCount(1)
|
||||
.requiredKeywordCount(1)
|
||||
.evidenceCoverage(new LinkedHashMap<>())
|
||||
.toolCallCount(1)
|
||||
.durationMs(1000)
|
||||
.build());
|
||||
|
||||
DiagnosisEvalDiffReport diff = differ.compare(baseline, current);
|
||||
|
||||
assertTrue(hasItem(diff, "REGRESSION", "case", removed.getCaseId(), "casePresence"));
|
||||
assertTrue(hasItem(diff, "CHANGED", "case", "new-case", "casePresence"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void compareSameReportHasNoDiff() throws Exception {
|
||||
DiagnosisEvalReport baseline = readBaselineReport();
|
||||
|
||||
DiagnosisEvalDiffReport diff = differ.compare(baseline, readBaselineReport());
|
||||
|
||||
assertFalse(diff.isHasRegression());
|
||||
assertEquals(0, diff.getRegressionCount());
|
||||
assertTrue(diff.getItems().isEmpty());
|
||||
}
|
||||
|
||||
@Test
|
||||
void writerOutputsJsonAndMarkdown(@TempDir Path tempDir) throws Exception {
|
||||
DiagnosisEvalReport baseline = readBaselineReport();
|
||||
DiagnosisEvalReport current = readBaselineReport();
|
||||
degradeRedisCase(current);
|
||||
DiagnosisEvalDiffReport diff = differ.compare(baseline, current);
|
||||
DiagnosisEvalDiffReportWriter writer = new DiagnosisEvalDiffReportWriter(objectMapper);
|
||||
|
||||
Path json = tempDir.resolve("baseline-diff.json");
|
||||
Path markdown = tempDir.resolve("baseline-diff.md");
|
||||
writer.writeJson(diff, json);
|
||||
writer.writeMarkdown(diff, markdown);
|
||||
|
||||
assertTrue(Files.exists(json));
|
||||
assertTrue(Files.readString(json).contains("\"hasRegression\" : true"));
|
||||
assertTrue(Files.readString(markdown).contains("# Diagnosis Eval Baseline Diff"));
|
||||
assertTrue(Files.readString(markdown).contains("redis-timeout"));
|
||||
}
|
||||
|
||||
private DiagnosisEvalReport readBaselineReport() throws Exception {
|
||||
return objectMapper.readValue(Path.of("mvp/eval/reports/baseline-report.json").toFile(),
|
||||
DiagnosisEvalReport.class);
|
||||
}
|
||||
|
||||
private void degradeRedisCase(DiagnosisEvalReport report) {
|
||||
report.setPassedCases(4);
|
||||
report.setPassRate(0.8);
|
||||
report.setAverageToolCallCount(3.0);
|
||||
report.setAverageDurationMs(45800.0);
|
||||
report.setVerdictDistribution(new LinkedHashMap<>());
|
||||
report.getVerdictDistribution().put("PASS", 2L);
|
||||
report.getVerdictDistribution().put("LOW_CONFID", 2L);
|
||||
report.getVerdictDistribution().put("REJECT", 1L);
|
||||
|
||||
DiagnosisEvalResult redis = result(report, "redis-timeout");
|
||||
redis.setPassed(false);
|
||||
redis.setFailedChecks(new ArrayList<>(List.of("missing required evidence tool: query_logs")));
|
||||
redis.setVerdict("REJECT");
|
||||
redis.setMatchedKeywordCount(1);
|
||||
redis.getEvidenceCoverage().put("query_logs", false);
|
||||
redis.setToolCallCount(1);
|
||||
redis.setDurationMs(36000);
|
||||
}
|
||||
|
||||
private DiagnosisEvalResult result(DiagnosisEvalReport report, String caseId) {
|
||||
return report.getResults().stream()
|
||||
.filter(item -> caseId.equals(item.getCaseId()))
|
||||
.findFirst()
|
||||
.orElseThrow();
|
||||
}
|
||||
|
||||
private boolean hasItem(DiagnosisEvalDiffReport diff,
|
||||
String type,
|
||||
String scope,
|
||||
String caseId,
|
||||
String metric) {
|
||||
return diff.getItems().stream().anyMatch(item ->
|
||||
type.equals(item.getType())
|
||||
&& scope.equals(item.getScope())
|
||||
&& java.util.Objects.equals(caseId, item.getCaseId())
|
||||
&& metric.equals(item.getMetric()));
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user