Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0

# Conflicts:
#	mvp/issues/README.md
This commit is contained in:
aruo
2026-07-05 01:42:42 +08:00
96 changed files with 4602 additions and 141 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,54 @@
## Context
The project now has the pieces needed for trace-based evaluation:
- `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation`
- `agent_step` stores ordered agent execution records
- `tool_invocation` stores evidence tool calls with normalized evidence semantics
- `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review
- `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
## Goals / Non-Goals
**Goals:**
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
- Produce JSON and Markdown reports with pass/fail status and key metrics.
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
- Leave room for a later runtime mode that queries the trace API after a demo run.
**Non-Goals:**
- No LLM-as-judge in this slice.
- No automatic prompt optimization.
- No new production API.
- No change to chat, verifier, retrieval, upload, or feedback behavior.
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
## Risks / Trade-offs
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
## Migration Plan
- No deployment migration is required.
- The harness is additive and can be run locally as a test or script.
- Rollback is deleting the eval case files, runner, and report docs.
## Open Questions
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?
@@ -0,0 +1,27 @@
## Why
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
## What Changes
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
- Add JSON and Markdown report output for quick review after a run.
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
## Capabilities
### New Capabilities
- `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
### Modified Capabilities
- None.
## Impact
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
- Affected APIs: none.
- Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions.
- Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.
@@ -0,0 +1,59 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
#### Scenario: Case definition includes expected evidence
- **WHEN** an evaluation case is defined
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
#### Scenario: Case definition can express forbidden behavior
- **WHEN** a case has known unsafe behavior
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
### Requirement: Evaluation harness SHALL validate diagnosis traces
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
#### Scenario: Evidence coverage validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
#### Scenario: Verifier evaluation validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
- **AND** the verdict SHALL be one of the case's allowed verdicts
#### Scenario: Answer keyword validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
#### Scenario: Degraded output validation
- **WHEN** a trace verdict is `REJECT`
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
### Requirement: Evaluation harness SHALL report quality and cost signals
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
#### Scenario: JSON report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
#### Scenario: Markdown report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
#### Scenario: Aggregate metrics
- **WHEN** multiple cases are evaluated
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
### Requirement: Evaluation harness SHALL support offline fixture mode
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
#### Scenario: Fixture trace evaluation
- **WHEN** the evaluator is run against a directory of trace fixture files
- **THEN** it SHALL evaluate each trace file against its matching case definition
- **AND** it SHALL not require a running application service
#### Scenario: Missing fixture is reported clearly
- **WHEN** a case has no matching trace fixture
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
@@ -0,0 +1,25 @@
## 1. Case Definitions
- [x] 1.1 Add evaluation case definition format for fixed MVP diagnosis scenarios.
- [x] 1.2 Add the first 5 case definitions: payment timeout, MySQL pool exhausted, Redis timeout, slow response, and JVM memory risk.
- [x] 1.3 Document the meaning of expected keywords, required evidence tools, allowed verdicts, and forbidden behavior.
## 2. Trace Fixtures
- [x] 2.1 Add fixture trace schema or DTOs that match `DiagnosisTraceResponse` enough for offline evaluation.
- [x] 2.2 Add at least one representative trace fixture for a passing case.
- [x] 2.3 Add at least one fixture covering low-confidence or degraded behavior.
## 3. Evaluator
- [x] 3.1 Implement trace validation rules for evidence coverage, verifier verdict, answer keyword coverage, and degraded-output contract.
- [x] 3.2 Implement aggregate metrics: pass rate, verdict distribution, average tool-call count, and average duration.
- [x] 3.3 Implement JSON report output.
- [x] 3.4 Implement Markdown report output.
## 4. Tests And Documentation
- [x] 4.1 Add focused offline tests for the evaluator.
- [x] 4.2 Add run instructions under `mvp/demo` or `mvp/notes`.
- [x] 4.3 Run targeted tests for the evaluator.
- [x] 4.4 Run compile verification.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,64 @@
## Context
The current MVP already has the core pieces required for traceable agent execution:
- `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation`
- `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder`
- `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries
- `ChatService` already contains fallback behavior for missing or invalid `verifier_output`
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
- `LookupKnowledgeTool` builds `ToolInvocation` rows itself
- the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)`
- “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
- degraded output behavior exists in code but is only lightly covered by tests
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
## Goals / Non-Goals
**Goals:**
- Centralize the common persistence contract for evidence-bearing tools.
- Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
- Define stable summarization semantics for:
- successful evidence
- no-hit / no-usable-evidence
- deduped retrievals
- failed evidence queries
- Make `ChatService` fallback and degraded-output paths testable as explicit product behavior.
- Keep the scope small enough to unblock the next P1-B evaluation harness.
**Non-Goals:**
- No new table, column, or Flyway migration.
- No new public API.
- No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`.
- No attempt to redesign planner/executor routing.
- No full offline runtime or end-to-end benchmark harness in this slice.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. |
| `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
## Risks / Trade-offs
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
- [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic.
## Migration Plan
- No deployment migration is required beyond shipping the code changes.
- Existing `tool_invocation` rows remain valid because this change reuses the same schema.
- Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
## Open Questions
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.
@@ -0,0 +1,26 @@
## Why
The MVP already persists evidence tool invocations and uses a Verifier to judge answer quality, but the current evidence trace semantics are still only partially standardized. For interview-grade agent engineering, the system needs a tighter contract for evidence persistence, no-evidence/failure states, and degraded output behavior, plus focused tests that prove those paths work offline.
## What Changes
- Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
- Align how evidence tools represent success, no-hit, deduped, and failed calls so `ToolTraceSummaryService` can summarize them consistently.
- Harden `ChatService` fallback behavior for invalid or missing verifier output and make the degraded-output paths explicitly testable.
- Add focused offline tests for evidence recording, trace summarization, and verifier fallback / degraded output behavior.
- Record this slice as a dedicated P1-A change tied to the interview-focused MVP hardening track.
## Capabilities
### New Capabilities
- `evidence-trace-hardening`: Covers standardized evidence invocation persistence, verifier-facing evidence summary semantics, and explicit degraded-output contracts for evidence gaps and verifier failures.
### Modified Capabilities
- `chat-verifier-agent`: Tightens verifier input evidence semantics and fallback guarantees without changing the high-level planner/executor/verifier workflow.
## Impact
- Affected code: `ToolInvocationRecorder`, `LookupKnowledgeTool`, `QueryLogsTools`, `QueryMetricsTools`, `ToolTraceSummaryService`, `ChatService`, and focused test classes.
- Affected runtime behavior: evidence-bearing tools will persist more consistent invocation semantics; verifier fallback and degraded outputs remain additive hardening, not a product-flow rewrite.
- Affected APIs: none. No new endpoint or schema is introduced.
- Non-goals: no new evidence tools, no database migration, no evaluation harness, no trace UI, no security/config cleanup in this slice.
@@ -0,0 +1,14 @@
## ADDED Requirements
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
#### Scenario: Failed evidence remains a verifier-visible gap
- **WHEN** an evidence-bearing tool invocation fails
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
- **AND** the verifier flow SHALL continue without crashing
#### Scenario: Deduped retrievals do not count as fresh support
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
- **THEN** those entries SHALL be treated as no-new-evidence
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
@@ -0,0 +1,82 @@
## ADDED Requirements
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
#### Scenario: Common evidence fields are always persisted
- **WHEN** an evidence-bearing tool finishes a call
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
- **WHEN** `lookup_knowledge` persists a tool invocation
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
#### Scenario: Tool failure is preserved as failure
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
- **THEN** the persisted row SHALL set `success=false`
- **AND** it SHALL preserve an `error_message` explaining the failure
#### Scenario: No usable evidence is preserved without pretending success
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
- **THEN** the persisted contract SHALL preserve that the call completed
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
#### Scenario: Deduped retrieval remains auditable
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
- **THEN** the persisted row SHALL preserve the dedup reason
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
#### Scenario: Failed evidence calls remain visible in the summary
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
- **THEN** the summary SHALL retain them
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
#### Scenario: No-hit and deduped calls do not upgrade evidence level
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
- **AND** their counts SHALL still be reflected in the merged summary entry
#### Scenario: Successful evidence keeps the strongest available support
- **WHEN** multiple rows for the same tool and topic domain are merged
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
### Requirement: ChatService SHALL degrade predictably on verifier output failures
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
#### Scenario: Missing verifier output falls back to LOW_CONFID
- **WHEN** the verifier step completes without a usable `verifier_output`
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
- **WHEN** the verifier returns malformed or non-parseable JSON
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the fallback SHALL still persist a verifier-evaluation record
#### Scenario: REJECT output hides unverified raw answer text
- **WHEN** the final verifier decision is `REJECT`
- **THEN** the user-facing output SHALL use the degraded template
- **AND** it SHALL NOT pass through the raw executor answer
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
#### Scenario: Evidence recorder contract is tested offline
- **WHEN** the test suite runs the focused recorder tests
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
#### Scenario: Trace summary hardening is tested offline
- **WHEN** the test suite runs the focused trace-summary tests
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
#### Scenario: Verifier fallback behavior is tested offline
- **WHEN** the test suite runs the focused `ChatService` fallback tests
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
@@ -0,0 +1,22 @@
## 1. Evidence Persistence Contract
- [x] 1.1 Extend `ToolInvocationRecorder` with a richer evidence-recording path that can preserve retrieval-aware fields as well as common evidence fields.
- [x] 1.2 Refactor `LookupKnowledgeTool` to persist `tool_invocation` rows through `ToolInvocationRecorder` instead of its own ad hoc row-construction path.
- [x] 1.3 Align `QueryLogsTools` and `QueryMetricsTools` no-hit / failure payloads with the hardened evidence contract.
## 2. Verifier-Facing Summary Semantics
- [x] 2.1 Harden `ToolTraceSummaryService` so failed, no-hit, and deduped evidence rows are summarized with stable no-evidence semantics.
- [x] 2.2 Preserve merged-call counts for repeated hits, failures, and no-new-evidence rows without overstating evidence strength.
## 3. Chat Degraded Paths
- [x] 3.1 Add focused `ChatService` tests for missing verifier output fallback to `LOW_CONFID`.
- [x] 3.2 Add focused `ChatService` tests for invalid verifier JSON fallback to `LOW_CONFID`.
- [x] 3.3 Add focused `ChatService` tests that `REJECT` output uses the degraded template and does not leak raw executor answer content.
## 4. Verification
- [x] 4.1 Add focused offline tests for the recorder contract and `ToolTraceSummaryService`.
- [x] 4.2 Run targeted test commands for the new/updated offline tests.
- [x] 4.3 Run compile verification.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,39 @@
## Context
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
## Goals / Non-Goals
**Goals:**
- Compare two `DiagnosisEvalReport` objects without requiring external services.
- Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
- Write JSON and Markdown diff outputs for review.
**Non-Goals:**
- Do not run the Agent or regenerate traces.
- Do not introduce LLM-as-judge.
- Do not change evaluator scoring rules.
- Do not block on performance thresholds beyond simple numeric diff signals.
## Decisions
- Decision: Compare report DTOs instead of raw traces.
- Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
- Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
- Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`.
- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
- Decision: Keep thresholds explicit and conservative.
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
## Risks / Trade-offs
- Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed.
- Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
- Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.
@@ -0,0 +1,28 @@
## Why
The evaluation baseline is now complete, but developers still need a repeatable way to decide whether a new Agent run regressed against that baseline. A deterministic baseline diff turns saved reports into an actionable regression signal instead of a static artifact.
## What Changes
- Add a baseline diff model that compares two `DiagnosisEvalReport` objects.
- Detect aggregate changes such as pass-rate drops, verdict distribution shifts, tool-call cost changes, and duration changes.
- Detect per-case changes such as pass/fail regression, verdict changes, keyword coverage changes, evidence coverage loss, and missing/new cases.
- Add JSON and Markdown diff output suitable for review.
- Document how to interpret the diff in the eval docs.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `diagnosis-eval-harness`: Extend the existing evaluation harness so a current report can be compared against the saved baseline report.
## Impact
- Affects eval-only Java code under `src/main/java/com/superbiz/agent/eval`.
- Adds focused tests under `src/test/java/com/superbiz/agent/eval`.
- Updates `mvp/eval` documentation and may add sample diff output.
- No production Agent runtime, API, database schema, or external dependency changes are expected.
@@ -0,0 +1,46 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
#### Scenario: Aggregate regression detection
- **WHEN** the current report has a lower pass rate than the baseline report
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
#### Scenario: Cost signal detection
- **WHEN** average tool-call count or average duration changes between reports
- **THEN** the diff SHALL record the baseline value, current value, and delta
#### Scenario: Verdict distribution comparison
- **WHEN** verdict counts differ between reports
- **THEN** the diff SHALL record the verdict distribution changes
### Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
#### Scenario: Case pass/fail regression
- **WHEN** a case changes from passing in the baseline to failing in the current report
- **THEN** the diff SHALL record a regression for that case
#### Scenario: Evidence coverage regression
- **WHEN** a required evidence tool changes from covered to uncovered for a case
- **THEN** the diff SHALL record a regression naming the case and tool
#### Scenario: Missing case detection
- **WHEN** a baseline case is absent from the current report
- **THEN** the diff SHALL record a regression for the missing case
#### Scenario: New case detection
- **WHEN** a current report contains a case absent from the baseline
- **THEN** the diff SHALL record the case as a non-regression change
### Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
#### Scenario: JSON diff output
- **WHEN** a baseline diff is written as JSON
- **THEN** it SHALL include aggregate summary fields and detailed diff items
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items
@@ -0,0 +1,23 @@
## 1. OpenSpec And Issue Setup
- [x] 1.1 Create slug-based issue and devflow tracking files.
- [x] 1.2 Create OpenSpec proposal, design, delta spec, and tasks.
## 2. Baseline Diff Implementation
- [x] 2.1 Add diff result data structures for summary and per-item changes.
- [x] 2.2 Implement deterministic report comparison rules.
- [x] 2.3 Implement JSON and Markdown diff report writing.
## 3. Documentation
- [x] 3.1 Document baseline diff inputs, outputs, and interpretation in eval docs.
- [x] 3.2 Add sample diff output for a representative regression.
## 4. Tests And Validation
- [x] 4.1 Add focused tests for aggregate and case-level diff behavior.
- [x] 4.2 Add focused tests for JSON and Markdown diff output.
- [x] 4.3 Run evaluator/diff tests.
- [x] 4.4 Run compile verification.
- [x] 4.5 Run OpenSpec validation.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,39 @@
## Context
`diagnosis-eval-harness` already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.
## Goals / Non-Goals
**Goals:**
- Add representative trace fixtures for every fixed diagnosis case.
- Save a baseline report that can be reviewed and compared after future Agent changes.
- Keep the baseline reproducible in offline mode.
- Document how to regenerate the baseline.
**Non-Goals:**
- Do not change production Agent runtime behavior.
- Do not require live infrastructure or a real LLM.
- Do not introduce a new LLM-based grader.
- Do not expand the case set beyond the existing five fixed MVP diagnosis cases.
## Decisions
- Use checked-in fixture traces instead of live service calls.
- Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
- Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
- Save baseline reports under `mvp/eval/reports`.
- Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
- Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
- Keep fixture outcomes representative rather than forcing every case to pass.
- Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
- Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.
## Risks / Trade-offs
- Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like `DiagnosisTraceResponse` and add tests that load every referenced fixture.
- A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
- Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.
@@ -0,0 +1,27 @@
## Why
The diagnosis evaluation harness is implemented, but the baseline is still incomplete because only two of the five fixed cases have trace fixtures. Completing the fixture set and saving a baseline report makes the harness useful as a practical regression signal for interview demos and future Agent changes.
## What Changes
- Add trace fixtures for the remaining fixed diagnosis cases: Redis timeout, slow response, and JVM memory risk.
- Add a reproducible baseline report generated from the full fixture set.
- Document how to regenerate and interpret the baseline.
- Keep the evaluator deterministic and offline; no live MySQL, Redis, Milvus, or LLM service is required.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `diagnosis-eval-harness`: Extend the existing evaluation harness requirement so the fixed MVP case set has complete fixture coverage and a saved baseline report.
## Impact
- Affects `mvp/eval/cases`, `mvp/eval/fixtures`, and eval documentation.
- May add baseline output files under `mvp/eval/reports`.
- May add or update focused evaluator tests to assert full fixture coverage and report generation.
- No production runtime API or database schema changes are expected.
@@ -0,0 +1,27 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
#### Scenario: Every case resolves to a fixture file
- **WHEN** the evaluator loads the fixed case definition file
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
#### Scenario: Fixture files are loadable as diagnosis traces
- **WHEN** each referenced fixture is loaded
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
The system SHALL preserve a generated baseline report for the full fixed fixture set.
#### Scenario: Baseline report includes all fixed cases
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
- **THEN** the report SHALL include one result for every fixed case
#### Scenario: Baseline report is reviewable
- **WHEN** the baseline report is written
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
#### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes fixtures or evaluator rules
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
@@ -0,0 +1,19 @@
## 1. Fixture Coverage
- [x] 1.1 Add Redis timeout trace fixture referenced by the fixed case file.
- [x] 1.2 Add slow response trace fixture referenced by the fixed case file.
- [x] 1.3 Add JVM memory risk trace fixture referenced by the fixed case file.
- [x] 1.4 Verify every `traceFixture` in `diagnosis-cases.json` resolves to an existing fixture file.
## 2. Baseline Reports
- [x] 2.1 Generate a full baseline JSON report for all fixed cases.
- [x] 2.2 Generate a full baseline Markdown report for review.
- [x] 2.3 Document how to regenerate and interpret the baseline reports.
## 3. Tests And Validation
- [x] 3.1 Add or update focused tests for full fixture coverage and baseline report generation.
- [x] 3.2 Run evaluator tests.
- [x] 3.3 Run compile verification.
- [x] 3.4 Run OpenSpec validation.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,35 @@
## Context
The current `mvp/demo` folder documents the core flow, but the steps are embedded in prose. For an interview, the demo needs a sharper entry point: what to start, what to run, what files get produced, and what to point at when explaining Agent engineering quality.
## Goals / Non-Goals
**Goals:**
- Make the payment-timeout demo runnable through a small script.
- Save chat, trace, and feedback responses for review.
- Provide a short interview walkthrough that connects runtime evidence to the engineering story.
- Keep the demo focused on existing APIs and existing `mvp-demo` profile behavior.
**Non-Goals:**
- Do not add new backend endpoints.
- Do not modify Agent prompts or runtime orchestration.
- Do not solve secret cleanup or full offline test isolation in this change.
- Do not expand the eval harness.
## Decisions
- Decision: Use PowerShell scripts.
- Reason: the current runbook already uses PowerShell and the user environment is Windows.
- Decision: Save outputs under `mvp/demo/output`.
- Reason: interview review is easier when chat, trace, and feedback responses are persisted as files.
- Decision: Keep the walkthrough separate from the low-level runbook.
- Reason: `README.md` should tell how to run; `interview-walkthrough.md` should tell how to explain.
## Risks / Trade-offs
- The demo still depends on configured MySQL, Redis, Milvus, and model keys. Mitigation: document this explicitly and keep mock log/metric providers enabled through `mvp-demo`.
- Script assertions are intentionally lightweight. Mitigation: use the trace checklist for human review and keep automated regression in `mvp/eval`.
@@ -0,0 +1,26 @@
## Why
The MVP already has trace, evidence hardening, and evaluation artifacts, but the interview demo path is still too scattered. This change packages the existing capabilities into a repeatable demo runbook that can be executed and explained in a short interview window.
## What Changes
- Add a focused interview walkthrough for the payment-timeout MVP demo.
- Add reusable request payloads and PowerShell scripts under `mvp/demo`.
- Add a trace inspection checklist that maps runtime output to the engineering story.
- Keep the change documentation-only and script-only; no backend runtime behavior changes.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `mvp-demo-trace-acceptance`: Extend the demo acceptance surface with a repeatable interview runbook and executable local demo scripts.
## Impact
- Affects `mvp/demo` documentation and scripts.
- Adds issue and devflow tracking files.
- No Java production code, API contract, database schema, or dependency changes are expected.
@@ -0,0 +1,30 @@
## ADDED Requirements
### Requirement: MVP demo SHALL provide an interview runbook
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
#### Scenario: Walkthrough explains the demo story
- **WHEN** a developer opens the interview walkthrough
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
#### Scenario: Walkthrough stays scoped to existing capabilities
- **WHEN** the walkthrough describes the demo
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
### Requirement: MVP demo SHALL provide executable local demo scripts
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
#### Scenario: Demo script sends the fixed diagnosis request
- **WHEN** the demo script is executed against a running local service
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
#### Scenario: Demo script captures review artifacts
- **WHEN** the demo script finishes successfully
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
### Requirement: MVP demo SHALL provide a trace inspection checklist
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
#### Scenario: Checklist maps fields to interview claims
- **WHEN** a developer reviews a trace response
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
@@ -0,0 +1,16 @@
## 1. Demo Artifacts
- [x] 1.1 Add fixed payment-timeout request payload.
- [x] 1.2 Add PowerShell script to run chat, trace, and feedback steps.
- [x] 1.3 Add output directory documentation without committing generated outputs.
## 2. Interview Documentation
- [x] 2.1 Add interview walkthrough for the demo story.
- [x] 2.2 Add trace inspection checklist.
- [x] 2.3 Update `mvp/demo/README.md` to link the runnable demo package.
## 3. Tracking And Validation
- [x] 3.1 Add slug-based issue and devflow tracking files.
- [x] 3.2 Run OpenSpec validation.
@@ -190,3 +190,17 @@ Verifier facts SHALL be linkable to the evidence summaries used during verificat
- **WHEN** the ChatService persists `verifier_evaluation`
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
#### Scenario: Failed evidence remains a verifier-visible gap
- **WHEN** an evidence-bearing tool invocation fails
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
- **AND** the verifier flow SHALL continue without crashing
#### Scenario: Deduped retrievals do not count as fresh support
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
- **THEN** those entries SHALL be treated as no-new-evidence
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
@@ -0,0 +1,135 @@
# diagnosis-eval-harness Specification
## Purpose
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
## Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
#### Scenario: Case definition includes expected evidence
- **WHEN** an evaluation case is defined
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
#### Scenario: Case definition can express forbidden behavior
- **WHEN** a case has known unsafe behavior
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
### Requirement: Evaluation harness SHALL validate diagnosis traces
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
#### Scenario: Evidence coverage validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
#### Scenario: Verifier evaluation validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
- **AND** the verdict SHALL be one of the case's allowed verdicts
#### Scenario: Answer keyword validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
#### Scenario: Degraded output validation
- **WHEN** a trace verdict is `REJECT`
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
### Requirement: Evaluation harness SHALL report quality and cost signals
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
#### Scenario: JSON report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
#### Scenario: Markdown report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
#### Scenario: Aggregate metrics
- **WHEN** multiple cases are evaluated
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
### Requirement: Evaluation harness SHALL support offline fixture mode
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
#### Scenario: Fixture trace evaluation
- **WHEN** the evaluator is run against a directory of trace fixture files
- **THEN** it SHALL evaluate each trace file against its matching case definition
- **AND** it SHALL not require a running application service
#### Scenario: Missing fixture is reported clearly
- **WHEN** a case has no matching trace fixture
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
#### Scenario: Every case resolves to a fixture file
- **WHEN** the evaluator loads the fixed case definition file
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
#### Scenario: Fixture files are loadable as diagnosis traces
- **WHEN** each referenced fixture is loaded
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
The system SHALL preserve a generated baseline report for the full fixed fixture set.
#### Scenario: Baseline report includes all fixed cases
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
- **THEN** the report SHALL include one result for every fixed case
#### Scenario: Baseline report is reviewable
- **WHEN** the baseline report is written
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
#### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes fixtures or evaluator rules
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
### Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
#### Scenario: Aggregate regression detection
- **WHEN** the current report has a lower pass rate than the baseline report
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
#### Scenario: Cost signal detection
- **WHEN** average tool-call count or average duration changes between reports
- **THEN** the diff SHALL record the baseline value, current value, and delta
#### Scenario: Verdict distribution comparison
- **WHEN** verdict counts differ between reports
- **THEN** the diff SHALL record the verdict distribution changes
### Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
#### Scenario: Case pass/fail regression
- **WHEN** a case changes from passing in the baseline to failing in the current report
- **THEN** the diff SHALL record a regression for that case
#### Scenario: Evidence coverage regression
- **WHEN** a required evidence tool changes from covered to uncovered for a case
- **THEN** the diff SHALL record a regression naming the case and tool
#### Scenario: Missing case detection
- **WHEN** a baseline case is absent from the current report
- **THEN** the diff SHALL record a regression for the missing case
#### Scenario: New case detection
- **WHEN** a current report contains a case absent from the baseline
- **THEN** the diff SHALL record the case as a non-regression change
### Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
#### Scenario: JSON diff output
- **WHEN** a baseline diff is written as JSON
- **THEN** it SHALL include aggregate summary fields and detailed diff items
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items
@@ -0,0 +1,86 @@
# evidence-trace-hardening Specification
## Purpose
TBD - created by archiving change evidence-trace-hardening. Update Purpose after archive.
## Requirements
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
#### Scenario: Common evidence fields are always persisted
- **WHEN** an evidence-bearing tool finishes a call
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
- **WHEN** `lookup_knowledge` persists a tool invocation
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
#### Scenario: Tool failure is preserved as failure
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
- **THEN** the persisted row SHALL set `success=false`
- **AND** it SHALL preserve an `error_message` explaining the failure
#### Scenario: No usable evidence is preserved without pretending success
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
- **THEN** the persisted contract SHALL preserve that the call completed
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
#### Scenario: Deduped retrieval remains auditable
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
- **THEN** the persisted row SHALL preserve the dedup reason
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
#### Scenario: Failed evidence calls remain visible in the summary
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
- **THEN** the summary SHALL retain them
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
#### Scenario: No-hit and deduped calls do not upgrade evidence level
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
- **AND** their counts SHALL still be reflected in the merged summary entry
#### Scenario: Successful evidence keeps the strongest available support
- **WHEN** multiple rows for the same tool and topic domain are merged
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
### Requirement: ChatService SHALL degrade predictably on verifier output failures
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
#### Scenario: Missing verifier output falls back to LOW_CONFID
- **WHEN** the verifier step completes without a usable `verifier_output`
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
- **WHEN** the verifier returns malformed or non-parseable JSON
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the fallback SHALL still persist a verifier-evaluation record
#### Scenario: REJECT output hides unverified raw answer text
- **WHEN** the final verifier decision is `REJECT`
- **THEN** the user-facing output SHALL use the degraded template
- **AND** it SHALL NOT pass through the raw executor answer
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
#### Scenario: Evidence recorder contract is tested offline
- **WHEN** the test suite runs the focused recorder tests
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
#### Scenario: Trace summary hardening is tested offline
- **WHEN** the test suite runs the focused trace-summary tests
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
#### Scenario: Verifier fallback behavior is tested offline
- **WHEN** the test suite runs the focused `ChatService` fallback tests
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
@@ -35,3 +35,32 @@ The project SHALL include an end-to-end acceptance case that demonstrates start-
#### Scenario: Reviewer follows the acceptance case
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
### Requirement: MVP demo SHALL provide an interview runbook
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
#### Scenario: Walkthrough explains the demo story
- **WHEN** a developer opens the interview walkthrough
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
#### Scenario: Walkthrough stays scoped to existing capabilities
- **WHEN** the walkthrough describes the demo
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
### Requirement: MVP demo SHALL provide executable local demo scripts
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
#### Scenario: Demo script sends the fixed diagnosis request
- **WHEN** the demo script is executed against a running local service
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
#### Scenario: Demo script captures review artifacts
- **WHEN** the demo script finishes successfully
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
### Requirement: MVP demo SHALL provide a trace inspection checklist
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
#### Scenario: Checklist maps fields to interview claims
- **WHEN** a developer reviews a trace response
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough