101 lines
5.8 KiB
Markdown
101 lines
5.8 KiB
Markdown
## Purpose
|
|
|
|
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same session id.
|
|
## Requirements
|
|
### Requirement: Diagnosis trace can be queried by session id
|
|
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
|
|
|
|
#### Scenario: Existing session trace is returned
|
|
- **WHEN** a caller requests trace data for a session id that exists in `diagnosis_session`
|
|
- **THEN** the system returns a success response containing the session summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback
|
|
|
|
#### Scenario: Missing session returns not found
|
|
- **WHEN** a caller requests trace data for a session id that does not exist in `diagnosis_session`
|
|
- **THEN** the system returns a 404 response using the existing session-not-found error contract
|
|
|
|
### Requirement: Trace aggregation is read-only
|
|
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate diagnosis sessions, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
|
|
|
#### Scenario: Trace query does not change persisted state
|
|
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
|
|
- **THEN** the system reads `diagnosis_session`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
|
|
|
### Requirement: MVP demo profile is available
|
|
The system SHALL provide an `mvp-demo` Spring profile that documents the demo runtime intent and keeps mock log and metric providers enabled for repeatable diagnosis demonstrations.
|
|
|
|
#### Scenario: Demo profile loads mock evidence providers
|
|
- **WHEN** the application starts with `--spring.profiles.active=mvp-demo`
|
|
- **THEN** `prometheus.mock-enabled` and `cls.mock-enabled` are enabled by profile configuration
|
|
|
|
### Requirement: End-to-end MVP acceptance case is documented
|
|
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same session id.
|
|
|
|
#### Scenario: Reviewer follows the acceptance case
|
|
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
|
|
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
|
|
|
|
### Requirement: MVP demo SHALL provide an interview runbook
|
|
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
|
|
|
|
#### Scenario: Walkthrough explains the demo story
|
|
- **WHEN** a developer opens the interview walkthrough
|
|
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
|
|
|
|
#### Scenario: Walkthrough stays scoped to existing capabilities
|
|
- **WHEN** the walkthrough describes the demo
|
|
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
|
|
|
|
### Requirement: MVP demo SHALL provide executable local demo scripts
|
|
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
|
|
|
|
#### Scenario: Demo script sends the fixed diagnosis request
|
|
- **WHEN** the demo script is executed against a running local service
|
|
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
|
|
|
|
#### Scenario: Demo script captures review artifacts
|
|
- **WHEN** the demo script finishes successfully
|
|
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
|
|
|
|
### Requirement: MVP demo SHALL provide a trace inspection checklist
|
|
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
|
|
|
|
#### Scenario: Checklist maps fields to interview claims
|
|
- **WHEN** a developer reviews a trace response
|
|
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
|
|
|
|
### Requirement: MVP demo SHALL provide a browser trace workbench
|
|
The MVP demo SHALL provide a browser-accessible static page for inspecting one
|
|
diagnosis trace by session id using the existing read-only Trace API.
|
|
|
|
#### Scenario: Existing trace renders in the workbench
|
|
- **WHEN** a reviewer opens the Trace workbench with a session id that exists
|
|
- **THEN** the page SHALL request `GET /api/diagnosis/{sessionId}/trace`
|
|
- **AND** it SHALL render session summary, ordered agent steps, ordered tool
|
|
invocations, verifier evaluation, and final answer when present
|
|
|
|
#### Scenario: Trace workbench handles missing or failed traces
|
|
- **WHEN** the Trace API returns an error or the session id is empty
|
|
- **THEN** the page SHALL show a clear error or empty state without mutating any
|
|
diagnosis data
|
|
|
|
### Requirement: Trace workbench SHALL expose evidence and skill boundaries
|
|
The Trace workbench SHALL make tool evidence and skill-loading boundaries visible
|
|
without requiring raw JSON inspection first.
|
|
|
|
#### Scenario: Tool and verifier evidence are inspectable
|
|
- **WHEN** a trace includes tool invocations and verifier facts
|
|
- **THEN** the page SHALL show tool invocation counts, success status, tool
|
|
filters, verifier facts, and evidence references
|
|
|
|
#### Scenario: RAG details are inspectable for lookup knowledge calls
|
|
- **WHEN** a `lookup_knowledge` invocation includes retrieval details
|
|
- **THEN** the page SHALL show query transform, retrieval trace, context pack,
|
|
rerank trace, and evidence block data where available
|
|
|
|
#### Scenario: Skill boundary checks are visible
|
|
- **WHEN** a trace includes planner, executor, or verifier steps
|
|
- **THEN** the page SHALL show best-effort indicators for selected skill,
|
|
planner `read_skill` text mentions, executor `read_skill` text mentions, and
|
|
verifier `read_skill` text mentions
|
|
|