Add diagnosis eval harness
This commit is contained in:
@@ -0,0 +1,36 @@
|
||||
# Acceptance: diagnosis-eval-harness
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | `ISS-006` and initial OpenSpec artifacts were created. |
|
||||
| Implementation | Done | Added fixed cases, fixture-mode trace evaluation, aggregate metrics, and JSON / Markdown report writer. |
|
||||
| Verification | Done | Targeted evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- First implementation uses fixture-mode evaluation.
|
||||
- Live trace API polling remains a follow-up option.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers fixed case loading, fixture evaluation, missing fixture reporting, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-harness --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: diagnosis-eval-harness
|
||||
|
||||
## Background
|
||||
|
||||
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Define fixed diagnosis cases for the MVP demo domain.
|
||||
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
|
||||
3. Produce JSON and Markdown reports for interview and regression use.
|
||||
4. Keep the first version offline by supporting trace fixtures.
|
||||
|
||||
## Scope
|
||||
|
||||
- Evaluation case definitions
|
||||
- Trace fixture shape
|
||||
- Rule-based evaluator
|
||||
- JSON / Markdown report output
|
||||
- Focused offline tests and docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No LLM-as-judge
|
||||
- No live end-to-end runtime requirement
|
||||
- No production API
|
||||
- No chat or verifier runtime change
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-harness/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Harness Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
|
||||
- Slug: `diagnosis-eval-harness`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
|
||||
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
|
||||
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
|
||||
- Reason: The first regression signal should be deterministic and tied to trace contracts.
|
||||
|
||||
- Decision: Support offline fixture traces first.
|
||||
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
|
||||
@@ -0,0 +1,10 @@
|
||||
# Diagnosis Eval Harness Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `openspec/specs/evidence-trace-hardening/spec.md` | Defines stable evidence states and summary behavior | Evaluation can rely on trace semantics rather than ad hoc log parsing | Yes |
|
||||
| `mvp/demo/README.md` | Documents an end-to-end demo flow with chat, trace, and feedback | Existing demo flow provides the runtime story, but not a reusable evaluation baseline | Yes |
|
||||
| `DiagnosisTraceService` | Aggregates session, steps, tools, and self-evaluation | Trace response shape can be reused as evaluation input | Yes |
|
||||
| `ToolTraceSummaryService` | Builds verifier-facing evidence summaries from persisted tool rows | Evaluator can check evidence coverage through persisted trace artifacts | Yes |
|
||||
Reference in New Issue
Block a user