9.2 KiB
Purpose
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same diagnosis run.
Requirements
Requirement: Diagnosis trace can be queried by session id
The system SHALL expose a read-only HTTP endpoint GET /api/diagnosis/{sessionId}/trace that returns the persisted diagnosis trace for the requested session id. When runId is omitted, the endpoint SHALL return the latest diagnosis run for compatibility. When runId is provided, the endpoint SHALL return that exact run after validating it belongs to the path sessionId.
Scenario: Existing session latest trace is returned
- WHEN a caller requests trace data for a session id that has at least one
diagnosis_run - THEN the system returns a success response containing the resolved run id, session summary, run summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback for the latest run
Scenario: Existing session exact trace is returned
- WHEN a caller requests trace data with
GET /api/diagnosis/{sessionId}/trace?runId=run-xxx - THEN the system validates that
runIdbelongs tosessionId - AND it returns a success response containing only the trace data for that run
Scenario: Missing session returns not found
- WHEN a caller requests trace data for a session id that does not exist in
chat_session,diagnosis_run, or historical compatibility data - THEN the system returns a 404 response using the existing session-not-found error contract
Requirement: Trace aggregation is read-only
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate chat sessions, diagnosis runs, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
Scenario: Trace query does not change persisted state
- WHEN a caller requests
GET /api/diagnosis/{sessionId}/trace - THEN the system reads
diagnosis_run,agent_step, andtool_invocationrecords and returns an aggregate without saving any of those records
Scenario: Exact trace query does not change persisted state
- WHEN a caller requests
GET /api/diagnosis/{sessionId}/trace?runId=run-xxx - THEN the system reads the specified run and its trace detail records without saving any of those records
Requirement: MVP demo profile is available
The system SHALL provide an mvp-demo Spring profile that documents the demo runtime intent and keeps mock log and metric providers enabled for repeatable diagnosis demonstrations.
Scenario: Demo profile loads mock evidence providers
- WHEN the application starts with
--spring.profiles.active=mvp-demo - THEN
prometheus.mock-enabledandcls.mock-enabledare enabled by profile configuration
Requirement: End-to-end MVP acceptance case is documented
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same sessionId + runId.
Scenario: Reviewer follows the acceptance case
- WHEN a reviewer follows the documented MVP demo acceptance steps
- THEN they can run the application, submit a diagnosis question, query the exact trace endpoint, and submit feedback for the same run
Requirement: MVP demo SHALL provide an interview runbook
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
Scenario: Walkthrough explains the demo story
- WHEN a developer opens the interview walkthrough
- THEN it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
Scenario: Walkthrough stays scoped to existing capabilities
- WHEN the walkthrough describes the demo
- THEN it SHALL avoid claiming unsupported runtime behavior or new production features
Requirement: MVP demo SHALL provide executable local demo scripts
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
Scenario: Demo script sends the fixed diagnosis request
- WHEN the demo script is executed against a running local service
- THEN it SHALL send the fixed payment-timeout chat request with a stable session id
- AND it SHALL read the returned run id for exact trace and feedback calls
Scenario: Demo script captures review artifacts
- WHEN the demo script finishes successfully
- THEN it SHALL write chat, exact trace, and feedback responses under a demo output directory
Requirement: MVP demo SHALL be reproducible for interviews
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
Scenario: interview demo check script records an evidence bundle
- WHEN the user runs the interview demo check script against a running
mvp-demoservice - THEN the script SHALL submit a fixed Chat diagnosis request
- AND it SHALL fetch the trace for the same
sessionId + runId - AND it SHALL submit useful feedback for that run
- AND it SHALL write chat, trace, feedback, and summary outputs under
mvp/demo/output/
Scenario: interview demo check fails with actionable readiness output
- WHEN the target service is not reachable
- THEN the script SHALL fail before issuing diagnosis requests
- AND the failure message SHALL name the base URL and the expected startup profile
Scenario: interview documentation explains audit fields
- WHEN an interviewer asks how prompt or Gatekeeper changes are audited
- THEN the demo documentation SHALL point to
prompt_audit.versionandgatekeeper_result.rule_set_version - AND it SHALL explain that deterministic eval fixtures are the regression source of truth
Requirement: MVP demo SHALL provide a trace inspection checklist
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and run-level auditability.
Scenario: Checklist maps fields to interview claims
- WHEN a developer reviews a trace response
- THEN the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
Requirement: MVP demo SHALL provide a browser trace workbench
The MVP demo SHALL provide a browser-accessible static page for inspecting one diagnosis trace by session id and optional run id using the existing read-only Trace API.
Scenario: Existing trace renders in the workbench
- WHEN a reviewer opens the Trace workbench with a session id that exists
- THEN the page SHALL request
GET /api/diagnosis/{sessionId}/trace - AND it SHALL render session summary, ordered agent steps, ordered tool invocations, verifier evaluation, and final answer when present
Scenario: Exact trace renders in the workbench
- WHEN a reviewer opens the Trace workbench with
?sessionId=...&runId=... - THEN the page SHALL request
GET /api/diagnosis/{sessionId}/trace?runId=... - AND it SHALL render only that run's trace data
Scenario: Trace workbench handles missing or failed traces
- WHEN the Trace API returns an error or the session id is empty
- THEN the page SHALL show a clear error or empty state without mutating any diagnosis data
Requirement: Trace workbench SHALL expose evidence and skill boundaries
The Trace workbench SHALL make tool evidence and skill-loading boundaries visible without requiring raw JSON inspection first.
Scenario: Tool and verifier evidence are inspectable
- WHEN a trace includes tool invocations and verifier facts
- THEN the page SHALL show tool invocation counts, success status, tool filters, verifier facts, and evidence references
Scenario: RAG details are inspectable for lookup knowledge calls
- WHEN a
lookup_knowledgeinvocation includes retrieval details - THEN the page SHALL show query transform, retrieval trace, context pack, rerank trace, and evidence block data where available
Scenario: Skill boundary checks are visible
- WHEN a trace includes planner, executor, or verifier steps
- THEN the page SHALL show best-effort indicators for selected skill,
planner
read_skilltext mentions, executorread_skilltext mentions, and verifierread_skilltext mentions
Requirement: MVP demo SHALL provide stable evidence-pipeline scenarios
The MVP demo SHALL provide stable scenarios that explain how to demonstrate positive evidence, no-evidence, and safety/reject behavior for an Agent engineering interview.
Scenario: Scenario guide maps demo inputs to evidence claims
- WHEN a reviewer opens the demo scenario guide
- THEN it SHALL list the supported positive, no-evidence, and safety/reject scenarios
- AND it SHALL map each scenario to a request payload or fixture id
- AND it SHALL describe the expected Gatekeeper, Verifier, Composer, and trace fields to inspect
Scenario: Live and fixture-backed scenarios are distinguished
- WHEN a demo scenario is fixture-backed rather than live-scripted
- THEN the documentation SHALL say so explicitly
- AND it SHALL avoid promising deterministic live LLM output for that scenario