Expand diagnosis eval fixtures
This commit is contained in:
@@ -4,6 +4,7 @@
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added remaining fixtures, full baseline reports, and documentation updates. |
|
||||
| Verification | Done | Evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Fixture coverage is complete for the five fixed diagnosis cases.
|
||||
- Baseline reports are saved under `mvp/eval/reports`.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers full fixture coverage, baseline report matching, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate expand-diagnosis-eval-fixtures --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Background
|
||||
|
||||
The diagnosis eval harness is implemented and archived, but the fixed baseline is incomplete because three of the five diagnosis cases still reference missing fixtures.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Add representative trace fixtures for all remaining fixed diagnosis cases.
|
||||
2. Save a reproducible baseline report in JSON and Markdown.
|
||||
3. Document how to regenerate and interpret the baseline.
|
||||
4. Keep evaluation offline and deterministic.
|
||||
|
||||
## Scope
|
||||
|
||||
- Redis timeout fixture
|
||||
- Slow response fixture
|
||||
- JVM memory risk fixture
|
||||
- Baseline reports under `mvp/eval/reports`
|
||||
- Focused tests for full fixture coverage and report generation
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new diagnosis cases
|
||||
- No production Agent runtime changes
|
||||
- No LLM-as-judge
|
||||
- No live infrastructure requirement
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/expand-diagnosis-eval-fixtures/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Expand Diagnosis Eval Fixtures Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: complete the fixed diagnosis eval baseline after the harness is in place.
|
||||
- Slug: `expand-diagnosis-eval-fixtures`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created the evaluator, case file, fixture mode, and report writer.
|
||||
- The first baseline still has missing fixtures by design.
|
||||
- This follow-up turns that partial baseline into a full fixed-case baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Keep this change data-focused.
|
||||
- Reason: the evaluator rules already landed; this change should not blur fixture expansion with harness behavior changes.
|
||||
|
||||
- Decision: Save baseline reports in the repository.
|
||||
- Reason: interview review and future diffs are easier when the expected baseline is visible.
|
||||
|
||||
- Decision: Use deterministic fixture traces instead of live trace generation.
|
||||
- Reason: this baseline should run without infrastructure or external model calls.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should add a CLI or Maven goal for report regeneration.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-04: Created slug-based issue `expand-diagnosis-eval-fixtures.md`.
|
||||
- 2026-07-04: Created OpenSpec change `expand-diagnosis-eval-fixtures`.
|
||||
- 2026-07-04: Added Redis timeout, slow response, and JVM memory risk fixtures.
|
||||
- 2026-07-04: Added baseline JSON and Markdown reports under `mvp/eval/reports`.
|
||||
- 2026-07-04: Verification passed with `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-04: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-04: Verification passed with `openspec validate expand-diagnosis-eval-fixtures --strict`.
|
||||
@@ -7,6 +7,7 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
|
||||
- Case definitions: `cases/diagnosis-cases.json`
|
||||
- Offline trace fixtures: `fixtures/*.json`
|
||||
- Field definitions: `schema.md`
|
||||
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
|
||||
- Evaluator implementation: `DiagnosisTraceEvaluator`
|
||||
- Report writer: `DiagnosisEvalReportWriter`
|
||||
|
||||
@@ -22,6 +23,17 @@ Run the focused evaluator test:
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||
```
|
||||
|
||||
The committed baseline report represents the current fixed fixture set:
|
||||
|
||||
```text
|
||||
5 fixed cases
|
||||
5 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
3 LOW_CONFID verdicts
|
||||
```
|
||||
|
||||
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
||||
|
||||
## Interview Story
|
||||
|
||||
The harness gives the MVP a repeatable baseline:
|
||||
|
||||
@@ -25,7 +25,7 @@
|
||||
"id": "redis-timeout",
|
||||
"title": "Redis timeout",
|
||||
"question": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
||||
"traceFixture": "redis-timeout-missing.json",
|
||||
"traceFixture": "redis-timeout-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["redis", "超时"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
@@ -36,7 +36,7 @@
|
||||
"id": "slow-response",
|
||||
"title": "Slow response",
|
||||
"question": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
||||
"traceFixture": "slow-response-missing.json",
|
||||
"traceFixture": "slow-response-pass.json",
|
||||
"expectedRootCauseKeywords": ["p99", "慢响应"],
|
||||
"minKeywordMatches": 1,
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
@@ -47,7 +47,7 @@
|
||||
"id": "jvm-memory-risk",
|
||||
"title": "JVM memory risk",
|
||||
"question": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||
"traceFixture": "jvm-memory-risk-missing.json",
|
||||
"traceFixture": "jvm-memory-risk-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["jvm", "内存", "oom"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
|
||||
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 53000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n订单服务存在 JVM 内存风险,但还不能完全确认会发生 OOM。指标显示 heap 使用率持续高于 88%,日志出现多次 Full GC 和 allocation pressure,需要继续观察对象增长来源并检查最近发布。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.52,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "indirect"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,41 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-redis-timeout",
|
||||
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 36000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nRedis 连接超时可能和支付服务到 Redis 的网络抖动或连接池等待有关。日志中出现 redis timeout 和 command timeout 记录,但当前缺少指标侧证据,因此只能作为低置信结论处理。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.46,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-redis-timeout",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 2,
|
||||
"returnedStepCount": 2,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-slow-response",
|
||||
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 47000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "用户服务 P99 升高主要表现为慢响应。指标显示 P99 latency 从 280ms 上升到 1800ms,日志中同时出现 slow request 和 downstream timeout,因此优先排查下游依赖耗时和线程池排队。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.78,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-slow-response",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-slow-response",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,82 @@
|
||||
{
|
||||
"totalCases" : 5,
|
||||
"passedCases" : 5,
|
||||
"passRate" : 1.0,
|
||||
"verdictDistribution" : {
|
||||
"PASS" : 2,
|
||||
"LOW_CONFID" : 3
|
||||
},
|
||||
"averageToolCallCount" : 2.0,
|
||||
"averageDurationMs" : 45800.0,
|
||||
"results" : [ {
|
||||
"caseId" : "payment-timeout",
|
||||
"title" : "Payment API timeout",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true,
|
||||
"query_metrics" : true
|
||||
},
|
||||
"toolCallCount" : 3,
|
||||
"durationMs" : 42000
|
||||
}, {
|
||||
"caseId" : "mysql-pool-exhausted",
|
||||
"title" : "MySQL connection pool exhausted",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 51000
|
||||
}, {
|
||||
"caseId" : "redis-timeout",
|
||||
"title" : "Redis timeout",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 36000
|
||||
}, {
|
||||
"caseId" : "slow-response",
|
||||
"title" : "Slow response",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 47000
|
||||
}, {
|
||||
"caseId" : "jvm-memory-risk",
|
||||
"title" : "JVM memory risk",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 53000
|
||||
} ]
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
# Diagnosis Eval Report
|
||||
|
||||
- Total cases: 5
|
||||
- Passed cases: 5
|
||||
- Pass rate: 100.00%
|
||||
- Average tool calls: 2.00
|
||||
- Average duration ms: 45800.00
|
||||
|
||||
## Verdict Distribution
|
||||
|
||||
- PASS: 2
|
||||
- LOW_CONFID: 3
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
|
||||
@@ -8,3 +8,4 @@
|
||||
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
|
||||
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
||||
|
||||
@@ -0,0 +1,83 @@
|
||||
# Expand Diagnosis Eval Fixtures
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-B follow-up
|
||||
**依赖**:`diagnosis-eval-harness`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
`diagnosis-eval-harness` 已经把固定 case、trace evaluator、JSON / Markdown report 和字段文档搭起来了。
|
||||
|
||||
现在还差一步:5 条固定诊断 case 里,只有 2 条有 fixture,另外 3 条还是 missing 状态。这个状态可以验证 evaluator 的错误报告能力,但还不能作为完整 baseline 展示。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 baseline 还不够完整:
|
||||
|
||||
- `redis-timeout` 没有对应 trace fixture。
|
||||
- `slow-response` 没有对应 trace fixture。
|
||||
- `jvm-memory-risk` 没有对应 trace fixture。
|
||||
- 仓库里还没有一份固定的 baseline JSON / Markdown 报告可供对比。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
补齐固定诊断评测集,让它从“框架可跑”变成“基准可用”。
|
||||
|
||||
完成后应该做到:
|
||||
|
||||
- 5 条固定 case 都能加载到对应 fixture。
|
||||
- evaluator 能输出完整 baseline report。
|
||||
- baseline report 被保存到仓库,后续 Agent 改动可以拿它做对比。
|
||||
- 文档说明怎么重新生成和怎么看报告。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- 补齐 3 个缺失 fixture。
|
||||
- 保存 baseline JSON / Markdown 报告。
|
||||
- 更新 eval 文档。
|
||||
- 补充测试,确保 case 文件引用的 fixture 都存在。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不新增 case 数量。
|
||||
- 不改生产 Agent 主链路。
|
||||
- 不引入 LLM-as-judge。
|
||||
- 不启动真实 MySQL、Redis、Milvus 或 LLM。
|
||||
|
||||
---
|
||||
|
||||
## 面试表达
|
||||
|
||||
可以这样讲:
|
||||
|
||||
```text
|
||||
我先搭了评测 harness,然后把固定 case 的 trace fixture 补齐,
|
||||
生成一份可复现的 baseline report。
|
||||
这样以后每次改 prompt、tool 或 verifier,
|
||||
都能看固定诊断集有没有行为回退,而不是只靠人工感觉。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- `mvp/eval/fixtures/`
|
||||
- `mvp/eval/reports/`
|
||||
- `mvp/eval/README.md`
|
||||
- `mvp/eval/schema.md`
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||
- `src/test/java/com/superbiz/agent/eval/DiagnosisTraceEvaluatorTest.java`
|
||||
- `openspec/specs/diagnosis-eval-harness/spec.md`
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,39 @@
|
||||
## Context
|
||||
|
||||
`diagnosis-eval-harness` already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Add representative trace fixtures for every fixed diagnosis case.
|
||||
- Save a baseline report that can be reviewed and compared after future Agent changes.
|
||||
- Keep the baseline reproducible in offline mode.
|
||||
- Document how to regenerate the baseline.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not change production Agent runtime behavior.
|
||||
- Do not require live infrastructure or a real LLM.
|
||||
- Do not introduce a new LLM-based grader.
|
||||
- Do not expand the case set beyond the existing five fixed MVP diagnosis cases.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Use checked-in fixture traces instead of live service calls.
|
||||
- Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
|
||||
- Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
|
||||
|
||||
- Save baseline reports under `mvp/eval/reports`.
|
||||
- Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
|
||||
- Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
|
||||
|
||||
- Keep fixture outcomes representative rather than forcing every case to pass.
|
||||
- Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
|
||||
- Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like `DiagnosisTraceResponse` and add tests that load every referenced fixture.
|
||||
- A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
|
||||
- Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The diagnosis evaluation harness is implemented, but the baseline is still incomplete because only two of the five fixed cases have trace fixtures. Completing the fixture set and saving a baseline report makes the harness useful as a practical regression signal for interview demos and future Agent changes.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add trace fixtures for the remaining fixed diagnosis cases: Redis timeout, slow response, and JVM memory risk.
|
||||
- Add a reproducible baseline report generated from the full fixture set.
|
||||
- Document how to regenerate and interpret the baseline.
|
||||
- Keep the evaluator deterministic and offline; no live MySQL, Redis, Milvus, or LLM service is required.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness requirement so the fixed MVP case set has complete fixture coverage and a saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects `mvp/eval/cases`, `mvp/eval/fixtures`, and eval documentation.
|
||||
- May add baseline output files under `mvp/eval/reports`.
|
||||
- May add or update focused evaluator tests to assert full fixture coverage and report generation.
|
||||
- No production runtime API or database schema changes are expected.
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
|
||||
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
|
||||
|
||||
#### Scenario: Every case resolves to a fixture file
|
||||
- **WHEN** the evaluator loads the fixed case definition file
|
||||
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
|
||||
|
||||
#### Scenario: Fixture files are loadable as diagnosis traces
|
||||
- **WHEN** each referenced fixture is loaded
|
||||
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
|
||||
|
||||
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
|
||||
The system SHALL preserve a generated baseline report for the full fixed fixture set.
|
||||
|
||||
#### Scenario: Baseline report includes all fixed cases
|
||||
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
|
||||
- **THEN** the report SHALL include one result for every fixed case
|
||||
|
||||
#### Scenario: Baseline report is reviewable
|
||||
- **WHEN** the baseline report is written
|
||||
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
|
||||
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes fixtures or evaluator rules
|
||||
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
||||
@@ -0,0 +1,19 @@
|
||||
## 1. Fixture Coverage
|
||||
|
||||
- [x] 1.1 Add Redis timeout trace fixture referenced by the fixed case file.
|
||||
- [x] 1.2 Add slow response trace fixture referenced by the fixed case file.
|
||||
- [x] 1.3 Add JVM memory risk trace fixture referenced by the fixed case file.
|
||||
- [x] 1.4 Verify every `traceFixture` in `diagnosis-cases.json` resolves to an existing fixture file.
|
||||
|
||||
## 2. Baseline Reports
|
||||
|
||||
- [x] 2.1 Generate a full baseline JSON report for all fixed cases.
|
||||
- [x] 2.2 Generate a full baseline Markdown report for review.
|
||||
- [x] 2.3 Document how to regenerate and interpret the baseline reports.
|
||||
|
||||
## 3. Tests And Validation
|
||||
|
||||
- [x] 3.1 Add or update focused tests for full fixture coverage and baseline report generation.
|
||||
- [x] 3.2 Run evaluator tests.
|
||||
- [x] 3.3 Run compile verification.
|
||||
- [x] 3.4 Run OpenSpec validation.
|
||||
@@ -62,3 +62,29 @@ The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus,
|
||||
#### Scenario: Missing fixture is reported clearly
|
||||
- **WHEN** a case has no matching trace fixture
|
||||
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
|
||||
|
||||
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
|
||||
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
|
||||
|
||||
#### Scenario: Every case resolves to a fixture file
|
||||
- **WHEN** the evaluator loads the fixed case definition file
|
||||
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
|
||||
|
||||
#### Scenario: Fixture files are loadable as diagnosis traces
|
||||
- **WHEN** each referenced fixture is loaded
|
||||
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
|
||||
|
||||
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
|
||||
The system SHALL preserve a generated baseline report for the full fixed fixture set.
|
||||
|
||||
#### Scenario: Baseline report includes all fixed cases
|
||||
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
|
||||
- **THEN** the report SHALL include one result for every fixed case
|
||||
|
||||
#### Scenario: Baseline report is reviewable
|
||||
- **WHEN** the baseline report is written
|
||||
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
|
||||
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes fixtures or evaluator rules
|
||||
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
||||
|
||||
@@ -19,16 +19,16 @@ class DiagnosisTraceEvaluatorTest {
|
||||
private final DiagnosisTraceEvaluator evaluator = new DiagnosisTraceEvaluator(objectMapper);
|
||||
|
||||
@Test
|
||||
void evaluateFixtureReportsPassingAndMissingCases() {
|
||||
void evaluateFixtureReportsFullBaseline() {
|
||||
List<DiagnosisEvalCase> cases = readCases();
|
||||
|
||||
DiagnosisEvalReport report = evaluator.evaluate(cases, Path.of("mvp/eval/fixtures"));
|
||||
|
||||
assertEquals(5, report.getTotalCases());
|
||||
assertEquals(2, report.getPassedCases());
|
||||
assertEquals(0.4, report.getPassRate(), 0.001);
|
||||
assertEquals(1L, report.getVerdictDistribution().get("PASS"));
|
||||
assertEquals(1L, report.getVerdictDistribution().get("LOW_CONFID"));
|
||||
assertEquals(5, report.getPassedCases());
|
||||
assertEquals(1.0, report.getPassRate(), 0.001);
|
||||
assertEquals(2L, report.getVerdictDistribution().get("PASS"));
|
||||
assertEquals(3L, report.getVerdictDistribution().get("LOW_CONFID"));
|
||||
|
||||
DiagnosisEvalResult payment = result(report, "payment-timeout");
|
||||
assertTrue(payment.isPassed());
|
||||
@@ -36,9 +36,17 @@ class DiagnosisTraceEvaluatorTest {
|
||||
assertTrue(payment.getEvidenceCoverage().get("query_logs"));
|
||||
assertTrue(payment.getEvidenceCoverage().get("query_metrics"));
|
||||
|
||||
DiagnosisEvalResult missing = result(report, "redis-timeout");
|
||||
assertFalse(missing.isPassed());
|
||||
assertTrue(missing.getFailedChecks().get(0).contains("trace fixture unavailable"));
|
||||
DiagnosisEvalResult redis = result(report, "redis-timeout");
|
||||
assertTrue(redis.isPassed());
|
||||
assertTrue(redis.getEvidenceCoverage().get("query_logs"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyFixedCaseReferencesExistingFixture() {
|
||||
for (DiagnosisEvalCase evalCase : readCases()) {
|
||||
Path fixture = Path.of("mvp/eval/fixtures").resolve(evalCase.getTraceFixture());
|
||||
assertTrue(Files.exists(fixture), "missing fixture: " + fixture);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -78,6 +86,12 @@ class DiagnosisTraceEvaluatorTest {
|
||||
assertTrue(Files.exists(json));
|
||||
assertTrue(Files.readString(markdown).contains("# Diagnosis Eval Report"));
|
||||
assertTrue(Files.readString(markdown).contains("payment-timeout"));
|
||||
assertEquals(
|
||||
comparableReportText(Files.readString(Path.of("mvp/eval/reports/baseline-report.json"))),
|
||||
comparableReportText(Files.readString(json)));
|
||||
assertEquals(
|
||||
comparableReportText(Files.readString(Path.of("mvp/eval/reports/baseline-report.md"))),
|
||||
comparableReportText(Files.readString(markdown)));
|
||||
}
|
||||
|
||||
private List<DiagnosisEvalCase> readCases() {
|
||||
@@ -94,4 +108,8 @@ class DiagnosisTraceEvaluatorTest {
|
||||
.findFirst()
|
||||
.orElseThrow();
|
||||
}
|
||||
|
||||
private String comparableReportText(String value) {
|
||||
return value.replace("\r\n", "\n").stripTrailing();
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user