Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0
# Conflicts: # mvp/issues/README.md
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The project now has the pieces needed for trace-based evaluation:
|
||||
|
||||
- `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation`
|
||||
- `agent_step` stores ordered agent execution records
|
||||
- `tool_invocation` stores evidence tool calls with normalized evidence semantics
|
||||
- `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review
|
||||
- `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics
|
||||
|
||||
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
|
||||
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
|
||||
- Produce JSON and Markdown reports with pass/fail status and key metrics.
|
||||
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
|
||||
- Leave room for a later runtime mode that queries the trace API after a demo run.
|
||||
|
||||
**Non-Goals:**
|
||||
- No LLM-as-judge in this slice.
|
||||
- No automatic prompt optimization.
|
||||
- No new production API.
|
||||
- No change to chat, verifier, retrieval, upload, or feedback behavior.
|
||||
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
|
||||
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
|
||||
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
|
||||
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
|
||||
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
|
||||
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
|
||||
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
|
||||
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required.
|
||||
- The harness is additive and can be run locally as a test or script.
|
||||
- Rollback is deleting the eval case files, runner, and report docs.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
|
||||
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
|
||||
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
|
||||
- Add JSON and Markdown report output for quick review after a run.
|
||||
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
|
||||
|
||||
### Modified Capabilities
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
|
||||
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
|
||||
- Affected APIs: none.
|
||||
- Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions.
|
||||
- Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
|
||||
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
|
||||
|
||||
#### Scenario: Case definition includes expected evidence
|
||||
- **WHEN** an evaluation case is defined
|
||||
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
|
||||
|
||||
#### Scenario: Case definition can express forbidden behavior
|
||||
- **WHEN** a case has known unsafe behavior
|
||||
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate diagnosis traces
|
||||
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
|
||||
|
||||
#### Scenario: Evidence coverage validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
|
||||
|
||||
#### Scenario: Verifier evaluation validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
|
||||
- **AND** the verdict SHALL be one of the case's allowed verdicts
|
||||
|
||||
#### Scenario: Answer keyword validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
|
||||
|
||||
#### Scenario: Degraded output validation
|
||||
- **WHEN** a trace verdict is `REJECT`
|
||||
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
|
||||
|
||||
### Requirement: Evaluation harness SHALL report quality and cost signals
|
||||
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
|
||||
|
||||
#### Scenario: JSON report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
|
||||
|
||||
#### Scenario: Markdown report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
|
||||
|
||||
#### Scenario: Aggregate metrics
|
||||
- **WHEN** multiple cases are evaluated
|
||||
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
|
||||
|
||||
### Requirement: Evaluation harness SHALL support offline fixture mode
|
||||
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
|
||||
|
||||
#### Scenario: Fixture trace evaluation
|
||||
- **WHEN** the evaluator is run against a directory of trace fixture files
|
||||
- **THEN** it SHALL evaluate each trace file against its matching case definition
|
||||
- **AND** it SHALL not require a running application service
|
||||
|
||||
#### Scenario: Missing fixture is reported clearly
|
||||
- **WHEN** a case has no matching trace fixture
|
||||
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
|
||||
@@ -0,0 +1,25 @@
|
||||
## 1. Case Definitions
|
||||
|
||||
- [x] 1.1 Add evaluation case definition format for fixed MVP diagnosis scenarios.
|
||||
- [x] 1.2 Add the first 5 case definitions: payment timeout, MySQL pool exhausted, Redis timeout, slow response, and JVM memory risk.
|
||||
- [x] 1.3 Document the meaning of expected keywords, required evidence tools, allowed verdicts, and forbidden behavior.
|
||||
|
||||
## 2. Trace Fixtures
|
||||
|
||||
- [x] 2.1 Add fixture trace schema or DTOs that match `DiagnosisTraceResponse` enough for offline evaluation.
|
||||
- [x] 2.2 Add at least one representative trace fixture for a passing case.
|
||||
- [x] 2.3 Add at least one fixture covering low-confidence or degraded behavior.
|
||||
|
||||
## 3. Evaluator
|
||||
|
||||
- [x] 3.1 Implement trace validation rules for evidence coverage, verifier verdict, answer keyword coverage, and degraded-output contract.
|
||||
- [x] 3.2 Implement aggregate metrics: pass rate, verdict distribution, average tool-call count, and average duration.
|
||||
- [x] 3.3 Implement JSON report output.
|
||||
- [x] 3.4 Implement Markdown report output.
|
||||
|
||||
## 4. Tests And Documentation
|
||||
|
||||
- [x] 4.1 Add focused offline tests for the evaluator.
|
||||
- [x] 4.2 Add run instructions under `mvp/demo` or `mvp/notes`.
|
||||
- [x] 4.3 Run targeted tests for the evaluator.
|
||||
- [x] 4.4 Run compile verification.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,64 @@
|
||||
## Context
|
||||
|
||||
The current MVP already has the core pieces required for traceable agent execution:
|
||||
|
||||
- `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation`
|
||||
- `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder`
|
||||
- `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries
|
||||
- `ChatService` already contains fallback behavior for missing or invalid `verifier_output`
|
||||
|
||||
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
|
||||
|
||||
- `LookupKnowledgeTool` builds `ToolInvocation` rows itself
|
||||
- the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)`
|
||||
- “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
|
||||
- degraded output behavior exists in code but is only lightly covered by tests
|
||||
|
||||
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Centralize the common persistence contract for evidence-bearing tools.
|
||||
- Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
|
||||
- Define stable summarization semantics for:
|
||||
- successful evidence
|
||||
- no-hit / no-usable-evidence
|
||||
- deduped retrievals
|
||||
- failed evidence queries
|
||||
- Make `ChatService` fallback and degraded-output paths testable as explicit product behavior.
|
||||
- Keep the scope small enough to unblock the next P1-B evaluation harness.
|
||||
|
||||
**Non-Goals:**
|
||||
- No new table, column, or Flyway migration.
|
||||
- No new public API.
|
||||
- No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`.
|
||||
- No attempt to redesign planner/executor routing.
|
||||
- No full offline runtime or end-to-end benchmark harness in this slice.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. |
|
||||
| `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
|
||||
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
|
||||
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
|
||||
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
|
||||
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
|
||||
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
|
||||
- [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required beyond shipping the code changes.
|
||||
- Existing `tool_invocation` rows remain valid because this change reuses the same schema.
|
||||
- Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.
|
||||
@@ -0,0 +1,26 @@
|
||||
## Why
|
||||
|
||||
The MVP already persists evidence tool invocations and uses a Verifier to judge answer quality, but the current evidence trace semantics are still only partially standardized. For interview-grade agent engineering, the system needs a tighter contract for evidence persistence, no-evidence/failure states, and degraded output behavior, plus focused tests that prove those paths work offline.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
- Align how evidence tools represent success, no-hit, deduped, and failed calls so `ToolTraceSummaryService` can summarize them consistently.
|
||||
- Harden `ChatService` fallback behavior for invalid or missing verifier output and make the degraded-output paths explicitly testable.
|
||||
- Add focused offline tests for evidence recording, trace summarization, and verifier fallback / degraded output behavior.
|
||||
- Record this slice as a dedicated P1-A change tied to the interview-focused MVP hardening track.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `evidence-trace-hardening`: Covers standardized evidence invocation persistence, verifier-facing evidence summary semantics, and explicit degraded-output contracts for evidence gaps and verifier failures.
|
||||
|
||||
### Modified Capabilities
|
||||
- `chat-verifier-agent`: Tightens verifier input evidence semantics and fallback guarantees without changing the high-level planner/executor/verifier workflow.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `ToolInvocationRecorder`, `LookupKnowledgeTool`, `QueryLogsTools`, `QueryMetricsTools`, `ToolTraceSummaryService`, `ChatService`, and focused test classes.
|
||||
- Affected runtime behavior: evidence-bearing tools will persist more consistent invocation semantics; verifier fallback and degraded outputs remain additive hardening, not a product-flow rewrite.
|
||||
- Affected APIs: none. No new endpoint or schema is introduced.
|
||||
- Non-goals: no new evidence tools, no database migration, no evaluation harness, no trace UI, no security/config cleanup in this slice.
|
||||
+14
@@ -0,0 +1,14 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
|
||||
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
|
||||
|
||||
#### Scenario: Failed evidence remains a verifier-visible gap
|
||||
- **WHEN** an evidence-bearing tool invocation fails
|
||||
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
|
||||
- **AND** the verifier flow SHALL continue without crashing
|
||||
|
||||
#### Scenario: Deduped retrievals do not count as fresh support
|
||||
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
|
||||
- **THEN** those entries SHALL be treated as no-new-evidence
|
||||
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
|
||||
+82
@@ -0,0 +1,82 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
|
||||
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
|
||||
#### Scenario: Common evidence fields are always persisted
|
||||
- **WHEN** an evidence-bearing tool finishes a call
|
||||
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
|
||||
|
||||
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
|
||||
- **WHEN** `lookup_knowledge` persists a tool invocation
|
||||
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
|
||||
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
|
||||
|
||||
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
|
||||
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
|
||||
|
||||
#### Scenario: Tool failure is preserved as failure
|
||||
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
|
||||
- **THEN** the persisted row SHALL set `success=false`
|
||||
- **AND** it SHALL preserve an `error_message` explaining the failure
|
||||
|
||||
#### Scenario: No usable evidence is preserved without pretending success
|
||||
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
|
||||
- **THEN** the persisted contract SHALL preserve that the call completed
|
||||
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
|
||||
|
||||
#### Scenario: Deduped retrieval remains auditable
|
||||
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
|
||||
- **THEN** the persisted row SHALL preserve the dedup reason
|
||||
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
|
||||
|
||||
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
|
||||
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
|
||||
|
||||
#### Scenario: Failed evidence calls remain visible in the summary
|
||||
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
|
||||
- **THEN** the summary SHALL retain them
|
||||
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
|
||||
|
||||
#### Scenario: No-hit and deduped calls do not upgrade evidence level
|
||||
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
|
||||
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
|
||||
- **AND** their counts SHALL still be reflected in the merged summary entry
|
||||
|
||||
#### Scenario: Successful evidence keeps the strongest available support
|
||||
- **WHEN** multiple rows for the same tool and topic domain are merged
|
||||
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
|
||||
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
|
||||
|
||||
### Requirement: ChatService SHALL degrade predictably on verifier output failures
|
||||
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
|
||||
|
||||
#### Scenario: Missing verifier output falls back to LOW_CONFID
|
||||
- **WHEN** the verifier step completes without a usable `verifier_output`
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
|
||||
|
||||
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
|
||||
- **WHEN** the verifier returns malformed or non-parseable JSON
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the fallback SHALL still persist a verifier-evaluation record
|
||||
|
||||
#### Scenario: REJECT output hides unverified raw answer text
|
||||
- **WHEN** the final verifier decision is `REJECT`
|
||||
- **THEN** the user-facing output SHALL use the degraded template
|
||||
- **AND** it SHALL NOT pass through the raw executor answer
|
||||
|
||||
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
|
||||
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
|
||||
|
||||
#### Scenario: Evidence recorder contract is tested offline
|
||||
- **WHEN** the test suite runs the focused recorder tests
|
||||
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
|
||||
|
||||
#### Scenario: Trace summary hardening is tested offline
|
||||
- **WHEN** the test suite runs the focused trace-summary tests
|
||||
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
|
||||
|
||||
#### Scenario: Verifier fallback behavior is tested offline
|
||||
- **WHEN** the test suite runs the focused `ChatService` fallback tests
|
||||
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Evidence Persistence Contract
|
||||
|
||||
- [x] 1.1 Extend `ToolInvocationRecorder` with a richer evidence-recording path that can preserve retrieval-aware fields as well as common evidence fields.
|
||||
- [x] 1.2 Refactor `LookupKnowledgeTool` to persist `tool_invocation` rows through `ToolInvocationRecorder` instead of its own ad hoc row-construction path.
|
||||
- [x] 1.3 Align `QueryLogsTools` and `QueryMetricsTools` no-hit / failure payloads with the hardened evidence contract.
|
||||
|
||||
## 2. Verifier-Facing Summary Semantics
|
||||
|
||||
- [x] 2.1 Harden `ToolTraceSummaryService` so failed, no-hit, and deduped evidence rows are summarized with stable no-evidence semantics.
|
||||
- [x] 2.2 Preserve merged-call counts for repeated hits, failures, and no-new-evidence rows without overstating evidence strength.
|
||||
|
||||
## 3. Chat Degraded Paths
|
||||
|
||||
- [x] 3.1 Add focused `ChatService` tests for missing verifier output fallback to `LOW_CONFID`.
|
||||
- [x] 3.2 Add focused `ChatService` tests for invalid verifier JSON fallback to `LOW_CONFID`.
|
||||
- [x] 3.3 Add focused `ChatService` tests that `REJECT` output uses the degraded template and does not leak raw executor answer content.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Add focused offline tests for the recorder contract and `ToolTraceSummaryService`.
|
||||
- [x] 4.2 Run targeted test commands for the new/updated offline tests.
|
||||
- [x] 4.3 Run compile verification.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,39 @@
|
||||
## Context
|
||||
|
||||
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Compare two `DiagnosisEvalReport` objects without requiring external services.
|
||||
- Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
|
||||
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
|
||||
- Write JSON and Markdown diff outputs for review.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not run the Agent or regenerate traces.
|
||||
- Do not introduce LLM-as-judge.
|
||||
- Do not change evaluator scoring rules.
|
||||
- Do not block on performance thresholds beyond simple numeric diff signals.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Decision: Compare report DTOs instead of raw traces.
|
||||
- Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
|
||||
- Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
|
||||
|
||||
- Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`.
|
||||
- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
|
||||
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
|
||||
|
||||
- Decision: Keep thresholds explicit and conservative.
|
||||
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
|
||||
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed.
|
||||
- Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
|
||||
- Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.
|
||||
@@ -0,0 +1,28 @@
|
||||
## Why
|
||||
|
||||
The evaluation baseline is now complete, but developers still need a repeatable way to decide whether a new Agent run regressed against that baseline. A deterministic baseline diff turns saved reports into an actionable regression signal instead of a static artifact.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a baseline diff model that compares two `DiagnosisEvalReport` objects.
|
||||
- Detect aggregate changes such as pass-rate drops, verdict distribution shifts, tool-call cost changes, and duration changes.
|
||||
- Detect per-case changes such as pass/fail regression, verdict changes, keyword coverage changes, evidence coverage loss, and missing/new cases.
|
||||
- Add JSON and Markdown diff output suitable for review.
|
||||
- Document how to interpret the diff in the eval docs.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness so a current report can be compared against the saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects eval-only Java code under `src/main/java/com/superbiz/agent/eval`.
|
||||
- Adds focused tests under `src/test/java/com/superbiz/agent/eval`.
|
||||
- Updates `mvp/eval` documentation and may add sample diff output.
|
||||
- No production Agent runtime, API, database schema, or external dependency changes are expected.
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
||||
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
||||
|
||||
#### Scenario: Aggregate regression detection
|
||||
- **WHEN** the current report has a lower pass rate than the baseline report
|
||||
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
||||
|
||||
#### Scenario: Cost signal detection
|
||||
- **WHEN** average tool-call count or average duration changes between reports
|
||||
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
||||
|
||||
#### Scenario: Verdict distribution comparison
|
||||
- **WHEN** verdict counts differ between reports
|
||||
- **THEN** the diff SHALL record the verdict distribution changes
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare case-level report results
|
||||
The system SHALL compare case results by case id and report actionable per-case changes.
|
||||
|
||||
#### Scenario: Case pass/fail regression
|
||||
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
||||
- **THEN** the diff SHALL record a regression for that case
|
||||
|
||||
#### Scenario: Evidence coverage regression
|
||||
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
||||
- **THEN** the diff SHALL record a regression naming the case and tool
|
||||
|
||||
#### Scenario: Missing case detection
|
||||
- **WHEN** a baseline case is absent from the current report
|
||||
- **THEN** the diff SHALL record a regression for the missing case
|
||||
|
||||
#### Scenario: New case detection
|
||||
- **WHEN** a current report contains a case absent from the baseline
|
||||
- **THEN** the diff SHALL record the case as a non-regression change
|
||||
|
||||
### Requirement: Evaluation harness SHALL report baseline diff results
|
||||
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
||||
|
||||
#### Scenario: JSON diff output
|
||||
- **WHEN** a baseline diff is written as JSON
|
||||
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
||||
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
@@ -0,0 +1,23 @@
|
||||
## 1. OpenSpec And Issue Setup
|
||||
|
||||
- [x] 1.1 Create slug-based issue and devflow tracking files.
|
||||
- [x] 1.2 Create OpenSpec proposal, design, delta spec, and tasks.
|
||||
|
||||
## 2. Baseline Diff Implementation
|
||||
|
||||
- [x] 2.1 Add diff result data structures for summary and per-item changes.
|
||||
- [x] 2.2 Implement deterministic report comparison rules.
|
||||
- [x] 2.3 Implement JSON and Markdown diff report writing.
|
||||
|
||||
## 3. Documentation
|
||||
|
||||
- [x] 3.1 Document baseline diff inputs, outputs, and interpretation in eval docs.
|
||||
- [x] 3.2 Add sample diff output for a representative regression.
|
||||
|
||||
## 4. Tests And Validation
|
||||
|
||||
- [x] 4.1 Add focused tests for aggregate and case-level diff behavior.
|
||||
- [x] 4.2 Add focused tests for JSON and Markdown diff output.
|
||||
- [x] 4.3 Run evaluator/diff tests.
|
||||
- [x] 4.4 Run compile verification.
|
||||
- [x] 4.5 Run OpenSpec validation.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,39 @@
|
||||
## Context
|
||||
|
||||
`diagnosis-eval-harness` already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Add representative trace fixtures for every fixed diagnosis case.
|
||||
- Save a baseline report that can be reviewed and compared after future Agent changes.
|
||||
- Keep the baseline reproducible in offline mode.
|
||||
- Document how to regenerate the baseline.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not change production Agent runtime behavior.
|
||||
- Do not require live infrastructure or a real LLM.
|
||||
- Do not introduce a new LLM-based grader.
|
||||
- Do not expand the case set beyond the existing five fixed MVP diagnosis cases.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Use checked-in fixture traces instead of live service calls.
|
||||
- Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
|
||||
- Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
|
||||
|
||||
- Save baseline reports under `mvp/eval/reports`.
|
||||
- Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
|
||||
- Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
|
||||
|
||||
- Keep fixture outcomes representative rather than forcing every case to pass.
|
||||
- Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
|
||||
- Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like `DiagnosisTraceResponse` and add tests that load every referenced fixture.
|
||||
- A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
|
||||
- Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The diagnosis evaluation harness is implemented, but the baseline is still incomplete because only two of the five fixed cases have trace fixtures. Completing the fixture set and saving a baseline report makes the harness useful as a practical regression signal for interview demos and future Agent changes.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add trace fixtures for the remaining fixed diagnosis cases: Redis timeout, slow response, and JVM memory risk.
|
||||
- Add a reproducible baseline report generated from the full fixture set.
|
||||
- Document how to regenerate and interpret the baseline.
|
||||
- Keep the evaluator deterministic and offline; no live MySQL, Redis, Milvus, or LLM service is required.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness requirement so the fixed MVP case set has complete fixture coverage and a saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects `mvp/eval/cases`, `mvp/eval/fixtures`, and eval documentation.
|
||||
- May add baseline output files under `mvp/eval/reports`.
|
||||
- May add or update focused evaluator tests to assert full fixture coverage and report generation.
|
||||
- No production runtime API or database schema changes are expected.
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
|
||||
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
|
||||
|
||||
#### Scenario: Every case resolves to a fixture file
|
||||
- **WHEN** the evaluator loads the fixed case definition file
|
||||
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
|
||||
|
||||
#### Scenario: Fixture files are loadable as diagnosis traces
|
||||
- **WHEN** each referenced fixture is loaded
|
||||
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
|
||||
|
||||
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
|
||||
The system SHALL preserve a generated baseline report for the full fixed fixture set.
|
||||
|
||||
#### Scenario: Baseline report includes all fixed cases
|
||||
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
|
||||
- **THEN** the report SHALL include one result for every fixed case
|
||||
|
||||
#### Scenario: Baseline report is reviewable
|
||||
- **WHEN** the baseline report is written
|
||||
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
|
||||
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes fixtures or evaluator rules
|
||||
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
||||
@@ -0,0 +1,19 @@
|
||||
## 1. Fixture Coverage
|
||||
|
||||
- [x] 1.1 Add Redis timeout trace fixture referenced by the fixed case file.
|
||||
- [x] 1.2 Add slow response trace fixture referenced by the fixed case file.
|
||||
- [x] 1.3 Add JVM memory risk trace fixture referenced by the fixed case file.
|
||||
- [x] 1.4 Verify every `traceFixture` in `diagnosis-cases.json` resolves to an existing fixture file.
|
||||
|
||||
## 2. Baseline Reports
|
||||
|
||||
- [x] 2.1 Generate a full baseline JSON report for all fixed cases.
|
||||
- [x] 2.2 Generate a full baseline Markdown report for review.
|
||||
- [x] 2.3 Document how to regenerate and interpret the baseline reports.
|
||||
|
||||
## 3. Tests And Validation
|
||||
|
||||
- [x] 3.1 Add or update focused tests for full fixture coverage and baseline report generation.
|
||||
- [x] 3.2 Run evaluator tests.
|
||||
- [x] 3.3 Run compile verification.
|
||||
- [x] 3.4 Run OpenSpec validation.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,35 @@
|
||||
## Context
|
||||
|
||||
The current `mvp/demo` folder documents the core flow, but the steps are embedded in prose. For an interview, the demo needs a sharper entry point: what to start, what to run, what files get produced, and what to point at when explaining Agent engineering quality.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make the payment-timeout demo runnable through a small script.
|
||||
- Save chat, trace, and feedback responses for review.
|
||||
- Provide a short interview walkthrough that connects runtime evidence to the engineering story.
|
||||
- Keep the demo focused on existing APIs and existing `mvp-demo` profile behavior.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not add new backend endpoints.
|
||||
- Do not modify Agent prompts or runtime orchestration.
|
||||
- Do not solve secret cleanup or full offline test isolation in this change.
|
||||
- Do not expand the eval harness.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Decision: Use PowerShell scripts.
|
||||
- Reason: the current runbook already uses PowerShell and the user environment is Windows.
|
||||
|
||||
- Decision: Save outputs under `mvp/demo/output`.
|
||||
- Reason: interview review is easier when chat, trace, and feedback responses are persisted as files.
|
||||
|
||||
- Decision: Keep the walkthrough separate from the low-level runbook.
|
||||
- Reason: `README.md` should tell how to run; `interview-walkthrough.md` should tell how to explain.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- The demo still depends on configured MySQL, Redis, Milvus, and model keys. Mitigation: document this explicitly and keep mock log/metric providers enabled through `mvp-demo`.
|
||||
- Script assertions are intentionally lightweight. Mitigation: use the trace checklist for human review and keep automated regression in `mvp/eval`.
|
||||
@@ -0,0 +1,26 @@
|
||||
## Why
|
||||
|
||||
The MVP already has trace, evidence hardening, and evaluation artifacts, but the interview demo path is still too scattered. This change packages the existing capabilities into a repeatable demo runbook that can be executed and explained in a short interview window.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a focused interview walkthrough for the payment-timeout MVP demo.
|
||||
- Add reusable request payloads and PowerShell scripts under `mvp/demo`.
|
||||
- Add a trace inspection checklist that maps runtime output to the engineering story.
|
||||
- Keep the change documentation-only and script-only; no backend runtime behavior changes.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `mvp-demo-trace-acceptance`: Extend the demo acceptance surface with a repeatable interview runbook and executable local demo scripts.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects `mvp/demo` documentation and scripts.
|
||||
- Adds issue and devflow tracking files.
|
||||
- No Java production code, API contract, database schema, or dependency changes are expected.
|
||||
+30
@@ -0,0 +1,30 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: MVP demo SHALL provide an interview runbook
|
||||
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
|
||||
|
||||
#### Scenario: Walkthrough explains the demo story
|
||||
- **WHEN** a developer opens the interview walkthrough
|
||||
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
|
||||
|
||||
#### Scenario: Walkthrough stays scoped to existing capabilities
|
||||
- **WHEN** the walkthrough describes the demo
|
||||
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
|
||||
|
||||
### Requirement: MVP demo SHALL provide executable local demo scripts
|
||||
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
|
||||
|
||||
#### Scenario: Demo script sends the fixed diagnosis request
|
||||
- **WHEN** the demo script is executed against a running local service
|
||||
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
|
||||
|
||||
#### Scenario: Demo script captures review artifacts
|
||||
- **WHEN** the demo script finishes successfully
|
||||
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
|
||||
|
||||
### Requirement: MVP demo SHALL provide a trace inspection checklist
|
||||
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
|
||||
|
||||
#### Scenario: Checklist maps fields to interview claims
|
||||
- **WHEN** a developer reviews a trace response
|
||||
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
|
||||
@@ -0,0 +1,16 @@
|
||||
## 1. Demo Artifacts
|
||||
|
||||
- [x] 1.1 Add fixed payment-timeout request payload.
|
||||
- [x] 1.2 Add PowerShell script to run chat, trace, and feedback steps.
|
||||
- [x] 1.3 Add output directory documentation without committing generated outputs.
|
||||
|
||||
## 2. Interview Documentation
|
||||
|
||||
- [x] 2.1 Add interview walkthrough for the demo story.
|
||||
- [x] 2.2 Add trace inspection checklist.
|
||||
- [x] 2.3 Update `mvp/demo/README.md` to link the runnable demo package.
|
||||
|
||||
## 3. Tracking And Validation
|
||||
|
||||
- [x] 3.1 Add slug-based issue and devflow tracking files.
|
||||
- [x] 3.2 Run OpenSpec validation.
|
||||
Reference in New Issue
Block a user