Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0
# Conflicts: # mvp/issues/README.md
This commit is contained in:
@@ -190,3 +190,17 @@ Verifier facts SHALL be linkable to the evidence summaries used during verificat
|
||||
- **WHEN** the ChatService persists `verifier_evaluation`
|
||||
- **THEN** it SHALL include `traceability_version`
|
||||
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
|
||||
|
||||
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
|
||||
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
|
||||
|
||||
#### Scenario: Failed evidence remains a verifier-visible gap
|
||||
- **WHEN** an evidence-bearing tool invocation fails
|
||||
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
|
||||
- **AND** the verifier flow SHALL continue without crashing
|
||||
|
||||
#### Scenario: Deduped retrievals do not count as fresh support
|
||||
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
|
||||
- **THEN** those entries SHALL be treated as no-new-evidence
|
||||
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
|
||||
|
||||
|
||||
@@ -0,0 +1,135 @@
|
||||
# diagnosis-eval-harness Specification
|
||||
|
||||
## Purpose
|
||||
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
|
||||
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
|
||||
|
||||
#### Scenario: Case definition includes expected evidence
|
||||
- **WHEN** an evaluation case is defined
|
||||
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
|
||||
|
||||
#### Scenario: Case definition can express forbidden behavior
|
||||
- **WHEN** a case has known unsafe behavior
|
||||
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate diagnosis traces
|
||||
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
|
||||
|
||||
#### Scenario: Evidence coverage validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
|
||||
|
||||
#### Scenario: Verifier evaluation validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
|
||||
- **AND** the verdict SHALL be one of the case's allowed verdicts
|
||||
|
||||
#### Scenario: Answer keyword validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
|
||||
|
||||
#### Scenario: Degraded output validation
|
||||
- **WHEN** a trace verdict is `REJECT`
|
||||
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
|
||||
|
||||
### Requirement: Evaluation harness SHALL report quality and cost signals
|
||||
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
|
||||
|
||||
#### Scenario: JSON report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
|
||||
|
||||
#### Scenario: Markdown report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
|
||||
|
||||
#### Scenario: Aggregate metrics
|
||||
- **WHEN** multiple cases are evaluated
|
||||
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
|
||||
|
||||
### Requirement: Evaluation harness SHALL support offline fixture mode
|
||||
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
|
||||
|
||||
#### Scenario: Fixture trace evaluation
|
||||
- **WHEN** the evaluator is run against a directory of trace fixture files
|
||||
- **THEN** it SHALL evaluate each trace file against its matching case definition
|
||||
- **AND** it SHALL not require a running application service
|
||||
|
||||
#### Scenario: Missing fixture is reported clearly
|
||||
- **WHEN** a case has no matching trace fixture
|
||||
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
|
||||
|
||||
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
|
||||
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
|
||||
|
||||
#### Scenario: Every case resolves to a fixture file
|
||||
- **WHEN** the evaluator loads the fixed case definition file
|
||||
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
|
||||
|
||||
#### Scenario: Fixture files are loadable as diagnosis traces
|
||||
- **WHEN** each referenced fixture is loaded
|
||||
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
|
||||
|
||||
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
|
||||
The system SHALL preserve a generated baseline report for the full fixed fixture set.
|
||||
|
||||
#### Scenario: Baseline report includes all fixed cases
|
||||
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
|
||||
- **THEN** the report SHALL include one result for every fixed case
|
||||
|
||||
#### Scenario: Baseline report is reviewable
|
||||
- **WHEN** the baseline report is written
|
||||
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
|
||||
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes fixtures or evaluator rules
|
||||
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
||||
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
||||
|
||||
#### Scenario: Aggregate regression detection
|
||||
- **WHEN** the current report has a lower pass rate than the baseline report
|
||||
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
||||
|
||||
#### Scenario: Cost signal detection
|
||||
- **WHEN** average tool-call count or average duration changes between reports
|
||||
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
||||
|
||||
#### Scenario: Verdict distribution comparison
|
||||
- **WHEN** verdict counts differ between reports
|
||||
- **THEN** the diff SHALL record the verdict distribution changes
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare case-level report results
|
||||
The system SHALL compare case results by case id and report actionable per-case changes.
|
||||
|
||||
#### Scenario: Case pass/fail regression
|
||||
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
||||
- **THEN** the diff SHALL record a regression for that case
|
||||
|
||||
#### Scenario: Evidence coverage regression
|
||||
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
||||
- **THEN** the diff SHALL record a regression naming the case and tool
|
||||
|
||||
#### Scenario: Missing case detection
|
||||
- **WHEN** a baseline case is absent from the current report
|
||||
- **THEN** the diff SHALL record a regression for the missing case
|
||||
|
||||
#### Scenario: New case detection
|
||||
- **WHEN** a current report contains a case absent from the baseline
|
||||
- **THEN** the diff SHALL record the case as a non-regression change
|
||||
|
||||
### Requirement: Evaluation harness SHALL report baseline diff results
|
||||
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
||||
|
||||
#### Scenario: JSON diff output
|
||||
- **WHEN** a baseline diff is written as JSON
|
||||
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
||||
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
@@ -0,0 +1,86 @@
|
||||
# evidence-trace-hardening Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change evidence-trace-hardening. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
|
||||
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
|
||||
#### Scenario: Common evidence fields are always persisted
|
||||
- **WHEN** an evidence-bearing tool finishes a call
|
||||
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
|
||||
|
||||
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
|
||||
- **WHEN** `lookup_knowledge` persists a tool invocation
|
||||
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
|
||||
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
|
||||
|
||||
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
|
||||
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
|
||||
|
||||
#### Scenario: Tool failure is preserved as failure
|
||||
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
|
||||
- **THEN** the persisted row SHALL set `success=false`
|
||||
- **AND** it SHALL preserve an `error_message` explaining the failure
|
||||
|
||||
#### Scenario: No usable evidence is preserved without pretending success
|
||||
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
|
||||
- **THEN** the persisted contract SHALL preserve that the call completed
|
||||
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
|
||||
|
||||
#### Scenario: Deduped retrieval remains auditable
|
||||
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
|
||||
- **THEN** the persisted row SHALL preserve the dedup reason
|
||||
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
|
||||
|
||||
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
|
||||
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
|
||||
|
||||
#### Scenario: Failed evidence calls remain visible in the summary
|
||||
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
|
||||
- **THEN** the summary SHALL retain them
|
||||
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
|
||||
|
||||
#### Scenario: No-hit and deduped calls do not upgrade evidence level
|
||||
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
|
||||
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
|
||||
- **AND** their counts SHALL still be reflected in the merged summary entry
|
||||
|
||||
#### Scenario: Successful evidence keeps the strongest available support
|
||||
- **WHEN** multiple rows for the same tool and topic domain are merged
|
||||
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
|
||||
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
|
||||
|
||||
### Requirement: ChatService SHALL degrade predictably on verifier output failures
|
||||
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
|
||||
|
||||
#### Scenario: Missing verifier output falls back to LOW_CONFID
|
||||
- **WHEN** the verifier step completes without a usable `verifier_output`
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
|
||||
|
||||
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
|
||||
- **WHEN** the verifier returns malformed or non-parseable JSON
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the fallback SHALL still persist a verifier-evaluation record
|
||||
|
||||
#### Scenario: REJECT output hides unverified raw answer text
|
||||
- **WHEN** the final verifier decision is `REJECT`
|
||||
- **THEN** the user-facing output SHALL use the degraded template
|
||||
- **AND** it SHALL NOT pass through the raw executor answer
|
||||
|
||||
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
|
||||
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
|
||||
|
||||
#### Scenario: Evidence recorder contract is tested offline
|
||||
- **WHEN** the test suite runs the focused recorder tests
|
||||
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
|
||||
|
||||
#### Scenario: Trace summary hardening is tested offline
|
||||
- **WHEN** the test suite runs the focused trace-summary tests
|
||||
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
|
||||
|
||||
#### Scenario: Verifier fallback behavior is tested offline
|
||||
- **WHEN** the test suite runs the focused `ChatService` fallback tests
|
||||
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
|
||||
|
||||
@@ -35,3 +35,32 @@ The project SHALL include an end-to-end acceptance case that demonstrates start-
|
||||
#### Scenario: Reviewer follows the acceptance case
|
||||
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
|
||||
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
|
||||
|
||||
### Requirement: MVP demo SHALL provide an interview runbook
|
||||
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
|
||||
|
||||
#### Scenario: Walkthrough explains the demo story
|
||||
- **WHEN** a developer opens the interview walkthrough
|
||||
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
|
||||
|
||||
#### Scenario: Walkthrough stays scoped to existing capabilities
|
||||
- **WHEN** the walkthrough describes the demo
|
||||
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
|
||||
|
||||
### Requirement: MVP demo SHALL provide executable local demo scripts
|
||||
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
|
||||
|
||||
#### Scenario: Demo script sends the fixed diagnosis request
|
||||
- **WHEN** the demo script is executed against a running local service
|
||||
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
|
||||
|
||||
#### Scenario: Demo script captures review artifacts
|
||||
- **WHEN** the demo script finishes successfully
|
||||
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
|
||||
|
||||
### Requirement: MVP demo SHALL provide a trace inspection checklist
|
||||
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
|
||||
|
||||
#### Scenario: Checklist maps fields to interview claims
|
||||
- **WHEN** a developer reviews a trace response
|
||||
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
|
||||
|
||||
Reference in New Issue
Block a user