Add MVP demo interview runbook

This commit is contained in:
aruo
2026-07-05 01:25:20 +08:00
parent 69deb15330
commit cbef3ddd3c
19 changed files with 548 additions and 0 deletions
+4
View File
@@ -60,3 +60,7 @@ uploads/
### Windows / Runtime Artifacts ### Windows / Runtime Artifacts
*.stackdump *.stackdump
NUL NUL
### MVP Demo Generated Outputs
mvp/demo/output/*.json
!mvp/demo/output/README.md
+1
View File
@@ -4,6 +4,7 @@
| 日期 | slug | 领域 | 关键词 | 状态 | | 日期 | slug | 领域 | 关键词 | 状态 |
|---|---|---|---|---| |---|---|---|---|---|
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/mvp-demo-interview-runbook | active |
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived | | 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived | | 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived | | 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
@@ -0,0 +1,25 @@
# Acceptance: mvp-demo-interview-runbook
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
| Verification | Done | OpenSpec validation passed. |
## Current State
- No backend runtime behavior has been changed.
- Demo is packaged under `mvp/demo` for interview use.
## Verification
### OpenSpec Verification
- Command: `openspec validate mvp-demo-interview-runbook --strict`
- Result: passed
@@ -0,0 +1,29 @@
# Brief: mvp-demo-interview-runbook
## Background
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
## Goals
1. Provide a fixed payment-timeout request payload.
2. Provide a PowerShell script that runs chat, trace, and feedback.
3. Save demo responses under `mvp/demo/output`.
4. Add interview walkthrough and trace checklist.
## Scope
- Demo docs and scripts only
- Existing local APIs only
- Existing `mvp-demo` profile only
## Non-Goals
- No backend code changes
- No eval extension
- No secret cleanup
- No full offline runtime
## Related OpenSpec
`openspec/changes/mvp-demo-interview-runbook/`
@@ -0,0 +1,27 @@
# MVP Demo Interview Runbook Decisions
## Clarify
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
- Slug: `mvp-demo-interview-runbook`
- Devflow scale: standard-light
## Context
- Evidence trace and eval baseline work are already done.
- The next useful step is not more eval tooling, but a runnable demo path.
## Key Decisions
- Decision: Keep this change documentation/script-only.
- Reason: Plan C is about demo packaging, not new runtime capability.
- Decision: Use a stable session id.
- Reason: it makes trace lookup and saved output predictable.
- Decision: Save outputs to `mvp/demo/output`.
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
## Open Questions
- Whether a later change should add a truly offline stubbed demo mode.
@@ -0,0 +1,9 @@
# Evidence: mvp-demo-interview-runbook
## Evidence Log
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
- 2026-07-05: Added fixed payment-timeout request payload.
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
+23
View File
@@ -2,6 +2,13 @@
This demo proves the MVP flow from user question to persisted diagnosis trace. This demo proves the MVP flow from user question to persisted diagnosis trace.
For interview use, start with:
- `interview-walkthrough.md` for the talk track
- `trace-inspection-checklist.md` for fields to inspect
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
- `requests/payment-timeout-chat.json` for the fixed request payload
## Prerequisites ## Prerequisites
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration. - MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
@@ -22,6 +29,22 @@ http://localhost:9900
## 1. Run Chat Diagnosis ## 1. Run Chat Diagnosis
Fast path:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
This writes:
```text
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
Manual path:
```powershell ```powershell
$sessionId = "mvp-demo-payment-timeout-001" $sessionId = "mvp-demo-payment-timeout-001"
$body = @{ $body = @{
+146
View File
@@ -0,0 +1,146 @@
# Interview Walkthrough: MVP Diagnosis Agent
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
## 30-Second Summary
```text
This is an enterprise diagnosis Agent MVP.
It takes a payment-timeout question, plans the investigation, calls evidence tools,
checks the answer through a verifier, persists the full trace, and accepts feedback.
```
The important claim is not "the model answered once." The claim is:
```text
The system can show what evidence was used, how the answer was checked, and how to replay the session.
```
## Demo Flow
1. Start the service with the `mvp-demo` profile.
2. Run the fixed payment-timeout request.
3. Open `mvp/demo/output/chat-response.json`.
4. Open `mvp/demo/output/trace-response.json`.
5. Point to evidence tools and verifier evaluation.
6. Submit feedback and show it is attached to the same session.
## Commands
Start service:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
Run the demo from another terminal:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
Optional custom session:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
```
## What To Show
### 1. User-Facing Answer
File:
```text
mvp/demo/output/chat-response.json
```
Say:
```text
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
```
### 2. Evidence Trace
File:
```text
mvp/demo/output/trace-response.json
```
Say:
```text
This is the important Agent engineering part.
I can inspect which tools were called, what inputs they received,
whether they succeeded, and what evidence preview was persisted.
```
Point to:
- `data.toolInvocations[*].toolName`
- `data.toolInvocations[*].inputParams`
- `data.toolInvocations[*].outputPreview`
- `data.toolInvocations[*].success`
### 3. Verifier / Self-Evaluation
Point to:
- `data.session.selfEvaluation`
- `data.summary.hasVerifierEvaluation`
Say:
```text
The final answer is not just raw Executor output.
It is checked by a verifier or self-evaluation layer using the persisted trace.
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
```
### 4. Feedback Loop
File:
```text
mvp/demo/output/feedback-response.json
```
Then re-query trace if needed.
Say:
```text
Feedback is attached to the same diagnosis session.
That makes it possible to mine useful / not useful cases later.
```
### 5. Regression Story
Mention, do not deep dive unless asked:
```text
For repeatability, I also built an offline eval baseline.
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
The two are separate on purpose: demo for human review, eval for automated signal.
```
## Strong Interview Framing
Use this phrasing:
```text
I focused on the Agent engineering surface:
traceability, evidence persistence, verifier gating, feedback, and regression checks.
The model answer is only one part of the system.
The more important part is whether we can audit and improve the answer after it is produced.
```
## Known Limits To Say Proactively
```text
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
```
+11
View File
@@ -0,0 +1,11 @@
# Demo Output
This directory is the default output location for local demo responses.
Generated files are intentionally ignored by Git:
- `chat-response.json`
- `trace-response.json`
- `feedback-response.json`
Keep this README so the directory exists in the repository.
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-payment-timeout-001",
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
}
@@ -0,0 +1,54 @@
param(
[string]$BaseUrl = "http://localhost:9900",
[string]$SessionId = "mvp-demo-payment-timeout-001",
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
[string]$OutputDir = "$PSScriptRoot/../output"
)
$ErrorActionPreference = "Stop"
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
$request.Id = $SessionId
$body = $request | ConvertTo-Json -Depth 8
Write-Host "Running payment-timeout chat demo..."
Write-Host "BaseUrl: $BaseUrl"
Write-Host "SessionId: $SessionId"
$chat = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/chat" `
-ContentType "application/json; charset=utf-8" `
-Body $body
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
Write-Host "Saved chat response: $OutputDir/chat-response.json"
$trace = Invoke-RestMethod `
-Method Get `
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
Write-Host "Saved trace response: $OutputDir/trace-response.json"
$feedbackBody = @{
sessionId = $SessionId
feedback = "useful"
} | ConvertTo-Json
$feedback = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/feedback" `
-ContentType "application/json; charset=utf-8" `
-Body $feedbackBody
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
Write-Host ""
Write-Host "Demo completed. Review:"
Write-Host "- mvp/demo/output/chat-response.json"
Write-Host "- mvp/demo/output/trace-response.json"
Write-Host "- mvp/demo/output/feedback-response.json"
+52
View File
@@ -0,0 +1,52 @@
# Trace Inspection Checklist
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
## Session
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
## Agent Steps
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
## Tool Evidence
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
## Summary
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
## What Good Looks Like
```text
same session id
-> final answer
-> persisted agent steps
-> persisted evidence tool calls
-> verifier/self-evaluation
-> feedback attached to the same session
```
+1
View File
@@ -10,3 +10,4 @@
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) | | ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) | | expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) | | diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 进行中(sm-flow) | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
+53
View File
@@ -0,0 +1,53 @@
# MVP Demo Interview Runbook
**状态**:进行中(sm-flow)
**严重程度**:中
**发现时间**:2026-07-05
**来源**:Plan C
**依赖**:`mvp-demo-trace-acceptance`, `evidence-trace-hardening`, `diagnosis-eval-harness`
---
## 背景
项目已经有 Agent 主链路、证据 trace、Verifier、反馈、eval baseline,但这些材料分散在不同目录。面试时真正需要的是一个能快速跑、快速讲清楚的 demo 入口。
---
## 问题
当前 demo 还不够“面试友好”:
- 启动、请求、trace、反馈步骤分散在文档里。
- 没有固定请求 payload 文件。
- 没有一键跑 payment-timeout demo 的脚本。
- 没有把 trace 字段和面试讲法对应起来的 walkthrough。
---
## 目标
把 Plan C 落地成 `mvp/demo` 下的可复现 demo 包:
- 固定支付超时请求。
- 一键执行 chat、trace、feedback。
- 保存 demo 输出,便于复盘。
- 提供面试讲解稿和 trace 检查清单。
---
## 范围
### In scope
- `mvp/demo` 文档。
- `mvp/demo/requests` 请求文件。
- `mvp/demo/scripts` PowerShell 脚本。
- `mvp/demo/output` 目录说明。
### Out of scope
- 不新增后端 API。
- 不改 Agent prompt。
- 不扩 eval harness。
- 不处理密钥外置和完整离线化。
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,35 @@
## Context
The current `mvp/demo` folder documents the core flow, but the steps are embedded in prose. For an interview, the demo needs a sharper entry point: what to start, what to run, what files get produced, and what to point at when explaining Agent engineering quality.
## Goals / Non-Goals
**Goals:**
- Make the payment-timeout demo runnable through a small script.
- Save chat, trace, and feedback responses for review.
- Provide a short interview walkthrough that connects runtime evidence to the engineering story.
- Keep the demo focused on existing APIs and existing `mvp-demo` profile behavior.
**Non-Goals:**
- Do not add new backend endpoints.
- Do not modify Agent prompts or runtime orchestration.
- Do not solve secret cleanup or full offline test isolation in this change.
- Do not expand the eval harness.
## Decisions
- Decision: Use PowerShell scripts.
- Reason: the current runbook already uses PowerShell and the user environment is Windows.
- Decision: Save outputs under `mvp/demo/output`.
- Reason: interview review is easier when chat, trace, and feedback responses are persisted as files.
- Decision: Keep the walkthrough separate from the low-level runbook.
- Reason: `README.md` should tell how to run; `interview-walkthrough.md` should tell how to explain.
## Risks / Trade-offs
- The demo still depends on configured MySQL, Redis, Milvus, and model keys. Mitigation: document this explicitly and keep mock log/metric providers enabled through `mvp-demo`.
- Script assertions are intentionally lightweight. Mitigation: use the trace checklist for human review and keep automated regression in `mvp/eval`.
@@ -0,0 +1,26 @@
## Why
The MVP already has trace, evidence hardening, and evaluation artifacts, but the interview demo path is still too scattered. This change packages the existing capabilities into a repeatable demo runbook that can be executed and explained in a short interview window.
## What Changes
- Add a focused interview walkthrough for the payment-timeout MVP demo.
- Add reusable request payloads and PowerShell scripts under `mvp/demo`.
- Add a trace inspection checklist that maps runtime output to the engineering story.
- Keep the change documentation-only and script-only; no backend runtime behavior changes.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `mvp-demo-trace-acceptance`: Extend the demo acceptance surface with a repeatable interview runbook and executable local demo scripts.
## Impact
- Affects `mvp/demo` documentation and scripts.
- Adds issue and devflow tracking files.
- No Java production code, API contract, database schema, or dependency changes are expected.
@@ -0,0 +1,30 @@
## ADDED Requirements
### Requirement: MVP demo SHALL provide an interview runbook
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
#### Scenario: Walkthrough explains the demo story
- **WHEN** a developer opens the interview walkthrough
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
#### Scenario: Walkthrough stays scoped to existing capabilities
- **WHEN** the walkthrough describes the demo
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
### Requirement: MVP demo SHALL provide executable local demo scripts
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
#### Scenario: Demo script sends the fixed diagnosis request
- **WHEN** the demo script is executed against a running local service
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
#### Scenario: Demo script captures review artifacts
- **WHEN** the demo script finishes successfully
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
### Requirement: MVP demo SHALL provide a trace inspection checklist
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
#### Scenario: Checklist maps fields to interview claims
- **WHEN** a developer reviews a trace response
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
@@ -0,0 +1,16 @@
## 1. Demo Artifacts
- [x] 1.1 Add fixed payment-timeout request payload.
- [x] 1.2 Add PowerShell script to run chat, trace, and feedback steps.
- [x] 1.3 Add output directory documentation without committing generated outputs.
## 2. Interview Documentation
- [x] 2.1 Add interview walkthrough for the demo story.
- [x] 2.2 Add trace inspection checklist.
- [x] 2.3 Update `mvp/demo/README.md` to link the runnable demo package.
## 3. Tracking And Validation
- [x] 3.1 Add slug-based issue and devflow tracking files.
- [x] 3.2 Run OpenSpec validation.