Add MVP demo interview runbook
This commit is contained in:
@@ -60,3 +60,7 @@ uploads/
|
|||||||
### Windows / Runtime Artifacts
|
### Windows / Runtime Artifacts
|
||||||
*.stackdump
|
*.stackdump
|
||||||
NUL
|
NUL
|
||||||
|
|
||||||
|
### MVP Demo Generated Outputs
|
||||||
|
mvp/demo/output/*.json
|
||||||
|
!mvp/demo/output/README.md
|
||||||
|
|||||||
@@ -4,6 +4,7 @@
|
|||||||
|
|
||||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
|
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/mvp-demo-interview-runbook | active |
|
||||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||||
|
|||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# Acceptance: mvp-demo-interview-runbook
|
||||||
|
|
||||||
|
## Classification
|
||||||
|
|
||||||
|
standard-light
|
||||||
|
|
||||||
|
## Task Status
|
||||||
|
|
||||||
|
| Task | Status | Notes |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
|
||||||
|
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
|
||||||
|
| Verification | Done | OpenSpec validation passed. |
|
||||||
|
|
||||||
|
## Current State
|
||||||
|
|
||||||
|
- No backend runtime behavior has been changed.
|
||||||
|
- Demo is packaged under `mvp/demo` for interview use.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
### OpenSpec Verification
|
||||||
|
|
||||||
|
- Command: `openspec validate mvp-demo-interview-runbook --strict`
|
||||||
|
- Result: passed
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
# Brief: mvp-demo-interview-runbook
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
|
||||||
|
|
||||||
|
## Goals
|
||||||
|
|
||||||
|
1. Provide a fixed payment-timeout request payload.
|
||||||
|
2. Provide a PowerShell script that runs chat, trace, and feedback.
|
||||||
|
3. Save demo responses under `mvp/demo/output`.
|
||||||
|
4. Add interview walkthrough and trace checklist.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- Demo docs and scripts only
|
||||||
|
- Existing local APIs only
|
||||||
|
- Existing `mvp-demo` profile only
|
||||||
|
|
||||||
|
## Non-Goals
|
||||||
|
|
||||||
|
- No backend code changes
|
||||||
|
- No eval extension
|
||||||
|
- No secret cleanup
|
||||||
|
- No full offline runtime
|
||||||
|
|
||||||
|
## Related OpenSpec
|
||||||
|
|
||||||
|
`openspec/changes/mvp-demo-interview-runbook/`
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
# MVP Demo Interview Runbook Decisions
|
||||||
|
|
||||||
|
## Clarify
|
||||||
|
|
||||||
|
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
|
||||||
|
- Slug: `mvp-demo-interview-runbook`
|
||||||
|
- Devflow scale: standard-light
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
- Evidence trace and eval baseline work are already done.
|
||||||
|
- The next useful step is not more eval tooling, but a runnable demo path.
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
- Decision: Keep this change documentation/script-only.
|
||||||
|
- Reason: Plan C is about demo packaging, not new runtime capability.
|
||||||
|
|
||||||
|
- Decision: Use a stable session id.
|
||||||
|
- Reason: it makes trace lookup and saved output predictable.
|
||||||
|
|
||||||
|
- Decision: Save outputs to `mvp/demo/output`.
|
||||||
|
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
- Whether a later change should add a truly offline stubbed demo mode.
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
# Evidence: mvp-demo-interview-runbook
|
||||||
|
|
||||||
|
## Evidence Log
|
||||||
|
|
||||||
|
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
|
||||||
|
- 2026-07-05: Added fixed payment-timeout request payload.
|
||||||
|
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
|
||||||
|
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
|
||||||
|
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
|
||||||
@@ -2,6 +2,13 @@
|
|||||||
|
|
||||||
This demo proves the MVP flow from user question to persisted diagnosis trace.
|
This demo proves the MVP flow from user question to persisted diagnosis trace.
|
||||||
|
|
||||||
|
For interview use, start with:
|
||||||
|
|
||||||
|
- `interview-walkthrough.md` for the talk track
|
||||||
|
- `trace-inspection-checklist.md` for fields to inspect
|
||||||
|
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
|
||||||
|
- `requests/payment-timeout-chat.json` for the fixed request payload
|
||||||
|
|
||||||
## Prerequisites
|
## Prerequisites
|
||||||
|
|
||||||
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
|
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
|
||||||
@@ -22,6 +29,22 @@ http://localhost:9900
|
|||||||
|
|
||||||
## 1. Run Chat Diagnosis
|
## 1. Run Chat Diagnosis
|
||||||
|
|
||||||
|
Fast path:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||||
|
```
|
||||||
|
|
||||||
|
This writes:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/output/chat-response.json
|
||||||
|
mvp/demo/output/trace-response.json
|
||||||
|
mvp/demo/output/feedback-response.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Manual path:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
$sessionId = "mvp-demo-payment-timeout-001"
|
$sessionId = "mvp-demo-payment-timeout-001"
|
||||||
$body = @{
|
$body = @{
|
||||||
|
|||||||
@@ -0,0 +1,146 @@
|
|||||||
|
# Interview Walkthrough: MVP Diagnosis Agent
|
||||||
|
|
||||||
|
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
|
||||||
|
|
||||||
|
## 30-Second Summary
|
||||||
|
|
||||||
|
```text
|
||||||
|
This is an enterprise diagnosis Agent MVP.
|
||||||
|
It takes a payment-timeout question, plans the investigation, calls evidence tools,
|
||||||
|
checks the answer through a verifier, persists the full trace, and accepts feedback.
|
||||||
|
```
|
||||||
|
|
||||||
|
The important claim is not "the model answered once." The claim is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
The system can show what evidence was used, how the answer was checked, and how to replay the session.
|
||||||
|
```
|
||||||
|
|
||||||
|
## Demo Flow
|
||||||
|
|
||||||
|
1. Start the service with the `mvp-demo` profile.
|
||||||
|
2. Run the fixed payment-timeout request.
|
||||||
|
3. Open `mvp/demo/output/chat-response.json`.
|
||||||
|
4. Open `mvp/demo/output/trace-response.json`.
|
||||||
|
5. Point to evidence tools and verifier evaluation.
|
||||||
|
6. Submit feedback and show it is attached to the same session.
|
||||||
|
|
||||||
|
## Commands
|
||||||
|
|
||||||
|
Start service:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||||
|
```
|
||||||
|
|
||||||
|
Run the demo from another terminal:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||||
|
```
|
||||||
|
|
||||||
|
Optional custom session:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
||||||
|
```
|
||||||
|
|
||||||
|
## What To Show
|
||||||
|
|
||||||
|
### 1. User-Facing Answer
|
||||||
|
|
||||||
|
File:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/output/chat-response.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Say:
|
||||||
|
|
||||||
|
```text
|
||||||
|
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Evidence Trace
|
||||||
|
|
||||||
|
File:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/output/trace-response.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Say:
|
||||||
|
|
||||||
|
```text
|
||||||
|
This is the important Agent engineering part.
|
||||||
|
I can inspect which tools were called, what inputs they received,
|
||||||
|
whether they succeeded, and what evidence preview was persisted.
|
||||||
|
```
|
||||||
|
|
||||||
|
Point to:
|
||||||
|
|
||||||
|
- `data.toolInvocations[*].toolName`
|
||||||
|
- `data.toolInvocations[*].inputParams`
|
||||||
|
- `data.toolInvocations[*].outputPreview`
|
||||||
|
- `data.toolInvocations[*].success`
|
||||||
|
|
||||||
|
### 3. Verifier / Self-Evaluation
|
||||||
|
|
||||||
|
Point to:
|
||||||
|
|
||||||
|
- `data.session.selfEvaluation`
|
||||||
|
- `data.summary.hasVerifierEvaluation`
|
||||||
|
|
||||||
|
Say:
|
||||||
|
|
||||||
|
```text
|
||||||
|
The final answer is not just raw Executor output.
|
||||||
|
It is checked by a verifier or self-evaluation layer using the persisted trace.
|
||||||
|
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. Feedback Loop
|
||||||
|
|
||||||
|
File:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/output/feedback-response.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Then re-query trace if needed.
|
||||||
|
|
||||||
|
Say:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Feedback is attached to the same diagnosis session.
|
||||||
|
That makes it possible to mine useful / not useful cases later.
|
||||||
|
```
|
||||||
|
|
||||||
|
### 5. Regression Story
|
||||||
|
|
||||||
|
Mention, do not deep dive unless asked:
|
||||||
|
|
||||||
|
```text
|
||||||
|
For repeatability, I also built an offline eval baseline.
|
||||||
|
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
|
||||||
|
The two are separate on purpose: demo for human review, eval for automated signal.
|
||||||
|
```
|
||||||
|
|
||||||
|
## Strong Interview Framing
|
||||||
|
|
||||||
|
Use this phrasing:
|
||||||
|
|
||||||
|
```text
|
||||||
|
I focused on the Agent engineering surface:
|
||||||
|
traceability, evidence persistence, verifier gating, feedback, and regression checks.
|
||||||
|
The model answer is only one part of the system.
|
||||||
|
The more important part is whether we can audit and improve the answer after it is produced.
|
||||||
|
```
|
||||||
|
|
||||||
|
## Known Limits To Say Proactively
|
||||||
|
|
||||||
|
```text
|
||||||
|
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
|
||||||
|
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
|
||||||
|
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
|
||||||
|
```
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
# Demo Output
|
||||||
|
|
||||||
|
This directory is the default output location for local demo responses.
|
||||||
|
|
||||||
|
Generated files are intentionally ignored by Git:
|
||||||
|
|
||||||
|
- `chat-response.json`
|
||||||
|
- `trace-response.json`
|
||||||
|
- `feedback-response.json`
|
||||||
|
|
||||||
|
Keep this README so the directory exists in the repository.
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
{
|
||||||
|
"Id": "mvp-demo-payment-timeout-001",
|
||||||
|
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||||
|
}
|
||||||
@@ -0,0 +1,54 @@
|
|||||||
|
param(
|
||||||
|
[string]$BaseUrl = "http://localhost:9900",
|
||||||
|
[string]$SessionId = "mvp-demo-payment-timeout-001",
|
||||||
|
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
|
||||||
|
[string]$OutputDir = "$PSScriptRoot/../output"
|
||||||
|
)
|
||||||
|
|
||||||
|
$ErrorActionPreference = "Stop"
|
||||||
|
|
||||||
|
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||||
|
|
||||||
|
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||||
|
$request.Id = $SessionId
|
||||||
|
$body = $request | ConvertTo-Json -Depth 8
|
||||||
|
|
||||||
|
Write-Host "Running payment-timeout chat demo..."
|
||||||
|
Write-Host "BaseUrl: $BaseUrl"
|
||||||
|
Write-Host "SessionId: $SessionId"
|
||||||
|
|
||||||
|
$chat = Invoke-RestMethod `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "$BaseUrl/api/chat" `
|
||||||
|
-ContentType "application/json; charset=utf-8" `
|
||||||
|
-Body $body
|
||||||
|
|
||||||
|
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||||
|
Write-Host "Saved chat response: $OutputDir/chat-response.json"
|
||||||
|
|
||||||
|
$trace = Invoke-RestMethod `
|
||||||
|
-Method Get `
|
||||||
|
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||||
|
|
||||||
|
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||||
|
Write-Host "Saved trace response: $OutputDir/trace-response.json"
|
||||||
|
|
||||||
|
$feedbackBody = @{
|
||||||
|
sessionId = $SessionId
|
||||||
|
feedback = "useful"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
$feedback = Invoke-RestMethod `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "$BaseUrl/api/feedback" `
|
||||||
|
-ContentType "application/json; charset=utf-8" `
|
||||||
|
-Body $feedbackBody
|
||||||
|
|
||||||
|
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
|
||||||
|
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
|
||||||
|
|
||||||
|
Write-Host ""
|
||||||
|
Write-Host "Demo completed. Review:"
|
||||||
|
Write-Host "- mvp/demo/output/chat-response.json"
|
||||||
|
Write-Host "- mvp/demo/output/trace-response.json"
|
||||||
|
Write-Host "- mvp/demo/output/feedback-response.json"
|
||||||
@@ -0,0 +1,52 @@
|
|||||||
|
# Trace Inspection Checklist
|
||||||
|
|
||||||
|
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
|
||||||
|
|
||||||
|
## Session
|
||||||
|
|
||||||
|
| JSON path | What to check | Interview point |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
|
||||||
|
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
|
||||||
|
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
|
||||||
|
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
|
||||||
|
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
|
||||||
|
|
||||||
|
## Agent Steps
|
||||||
|
|
||||||
|
| JSON path | What to check | Interview point |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
|
||||||
|
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
|
||||||
|
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
|
||||||
|
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
|
||||||
|
|
||||||
|
## Tool Evidence
|
||||||
|
|
||||||
|
| JSON path | What to check | Interview point |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
|
||||||
|
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
|
||||||
|
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
|
||||||
|
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
|
||||||
|
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
| JSON path | What to check | Interview point |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
|
||||||
|
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
|
||||||
|
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
|
||||||
|
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
|
||||||
|
|
||||||
|
## What Good Looks Like
|
||||||
|
|
||||||
|
```text
|
||||||
|
same session id
|
||||||
|
-> final answer
|
||||||
|
-> persisted agent steps
|
||||||
|
-> persisted evidence tool calls
|
||||||
|
-> verifier/self-evaluation
|
||||||
|
-> feedback attached to the same session
|
||||||
|
```
|
||||||
@@ -10,3 +10,4 @@
|
|||||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
||||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
||||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
|
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
|
||||||
|
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 进行中(sm-flow) | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
|
||||||
|
|||||||
@@ -0,0 +1,53 @@
|
|||||||
|
# MVP Demo Interview Runbook
|
||||||
|
|
||||||
|
**状态**:进行中(sm-flow)
|
||||||
|
**严重程度**:中
|
||||||
|
**发现时间**:2026-07-05
|
||||||
|
**来源**:Plan C
|
||||||
|
**依赖**:`mvp-demo-trace-acceptance`, `evidence-trace-hardening`, `diagnosis-eval-harness`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
项目已经有 Agent 主链路、证据 trace、Verifier、反馈、eval baseline,但这些材料分散在不同目录。面试时真正需要的是一个能快速跑、快速讲清楚的 demo 入口。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 问题
|
||||||
|
|
||||||
|
当前 demo 还不够“面试友好”:
|
||||||
|
|
||||||
|
- 启动、请求、trace、反馈步骤分散在文档里。
|
||||||
|
- 没有固定请求 payload 文件。
|
||||||
|
- 没有一键跑 payment-timeout demo 的脚本。
|
||||||
|
- 没有把 trace 字段和面试讲法对应起来的 walkthrough。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 目标
|
||||||
|
|
||||||
|
把 Plan C 落地成 `mvp/demo` 下的可复现 demo 包:
|
||||||
|
|
||||||
|
- 固定支付超时请求。
|
||||||
|
- 一键执行 chat、trace、feedback。
|
||||||
|
- 保存 demo 输出,便于复盘。
|
||||||
|
- 提供面试讲解稿和 trace 检查清单。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
### In scope
|
||||||
|
|
||||||
|
- `mvp/demo` 文档。
|
||||||
|
- `mvp/demo/requests` 请求文件。
|
||||||
|
- `mvp/demo/scripts` PowerShell 脚本。
|
||||||
|
- `mvp/demo/output` 目录说明。
|
||||||
|
|
||||||
|
### Out of scope
|
||||||
|
|
||||||
|
- 不新增后端 API。
|
||||||
|
- 不改 Agent prompt。
|
||||||
|
- 不扩 eval harness。
|
||||||
|
- 不处理密钥外置和完整离线化。
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
schema: spec-driven
|
||||||
|
created: 2026-07-04
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
The current `mvp/demo` folder documents the core flow, but the steps are embedded in prose. For an interview, the demo needs a sharper entry point: what to start, what to run, what files get produced, and what to point at when explaining Agent engineering quality.
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- Make the payment-timeout demo runnable through a small script.
|
||||||
|
- Save chat, trace, and feedback responses for review.
|
||||||
|
- Provide a short interview walkthrough that connects runtime evidence to the engineering story.
|
||||||
|
- Keep the demo focused on existing APIs and existing `mvp-demo` profile behavior.
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- Do not add new backend endpoints.
|
||||||
|
- Do not modify Agent prompts or runtime orchestration.
|
||||||
|
- Do not solve secret cleanup or full offline test isolation in this change.
|
||||||
|
- Do not expand the eval harness.
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
- Decision: Use PowerShell scripts.
|
||||||
|
- Reason: the current runbook already uses PowerShell and the user environment is Windows.
|
||||||
|
|
||||||
|
- Decision: Save outputs under `mvp/demo/output`.
|
||||||
|
- Reason: interview review is easier when chat, trace, and feedback responses are persisted as files.
|
||||||
|
|
||||||
|
- Decision: Keep the walkthrough separate from the low-level runbook.
|
||||||
|
- Reason: `README.md` should tell how to run; `interview-walkthrough.md` should tell how to explain.
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- The demo still depends on configured MySQL, Redis, Milvus, and model keys. Mitigation: document this explicitly and keep mock log/metric providers enabled through `mvp-demo`.
|
||||||
|
- Script assertions are intentionally lightweight. Mitigation: use the trace checklist for human review and keep automated regression in `mvp/eval`.
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
## Why
|
||||||
|
|
||||||
|
The MVP already has trace, evidence hardening, and evaluation artifacts, but the interview demo path is still too scattered. This change packages the existing capabilities into a repeatable demo runbook that can be executed and explained in a short interview window.
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- Add a focused interview walkthrough for the payment-timeout MVP demo.
|
||||||
|
- Add reusable request payloads and PowerShell scripts under `mvp/demo`.
|
||||||
|
- Add a trace inspection checklist that maps runtime output to the engineering story.
|
||||||
|
- Keep the change documentation-only and script-only; no backend runtime behavior changes.
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- None.
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- `mvp-demo-trace-acceptance`: Extend the demo acceptance surface with a repeatable interview runbook and executable local demo scripts.
|
||||||
|
|
||||||
|
## Impact
|
||||||
|
|
||||||
|
- Affects `mvp/demo` documentation and scripts.
|
||||||
|
- Adds issue and devflow tracking files.
|
||||||
|
- No Java production code, API contract, database schema, or dependency changes are expected.
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: MVP demo SHALL provide an interview runbook
|
||||||
|
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
|
||||||
|
|
||||||
|
#### Scenario: Walkthrough explains the demo story
|
||||||
|
- **WHEN** a developer opens the interview walkthrough
|
||||||
|
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
|
||||||
|
|
||||||
|
#### Scenario: Walkthrough stays scoped to existing capabilities
|
||||||
|
- **WHEN** the walkthrough describes the demo
|
||||||
|
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
|
||||||
|
|
||||||
|
### Requirement: MVP demo SHALL provide executable local demo scripts
|
||||||
|
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
|
||||||
|
|
||||||
|
#### Scenario: Demo script sends the fixed diagnosis request
|
||||||
|
- **WHEN** the demo script is executed against a running local service
|
||||||
|
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
|
||||||
|
|
||||||
|
#### Scenario: Demo script captures review artifacts
|
||||||
|
- **WHEN** the demo script finishes successfully
|
||||||
|
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
|
||||||
|
|
||||||
|
### Requirement: MVP demo SHALL provide a trace inspection checklist
|
||||||
|
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
|
||||||
|
|
||||||
|
#### Scenario: Checklist maps fields to interview claims
|
||||||
|
- **WHEN** a developer reviews a trace response
|
||||||
|
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
## 1. Demo Artifacts
|
||||||
|
|
||||||
|
- [x] 1.1 Add fixed payment-timeout request payload.
|
||||||
|
- [x] 1.2 Add PowerShell script to run chat, trace, and feedback steps.
|
||||||
|
- [x] 1.3 Add output directory documentation without committing generated outputs.
|
||||||
|
|
||||||
|
## 2. Interview Documentation
|
||||||
|
|
||||||
|
- [x] 2.1 Add interview walkthrough for the demo story.
|
||||||
|
- [x] 2.2 Add trace inspection checklist.
|
||||||
|
- [x] 2.3 Update `mvp/demo/README.md` to link the runnable demo package.
|
||||||
|
|
||||||
|
## 3. Tracking And Validation
|
||||||
|
|
||||||
|
- [x] 3.1 Add slug-based issue and devflow tracking files.
|
||||||
|
- [x] 3.2 Run OpenSpec validation.
|
||||||
Reference in New Issue
Block a user