Add MVP demo interview runbook
This commit is contained in:
@@ -0,0 +1,146 @@
|
||||
# Interview Walkthrough: MVP Diagnosis Agent
|
||||
|
||||
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
|
||||
|
||||
## 30-Second Summary
|
||||
|
||||
```text
|
||||
This is an enterprise diagnosis Agent MVP.
|
||||
It takes a payment-timeout question, plans the investigation, calls evidence tools,
|
||||
checks the answer through a verifier, persists the full trace, and accepts feedback.
|
||||
```
|
||||
|
||||
The important claim is not "the model answered once." The claim is:
|
||||
|
||||
```text
|
||||
The system can show what evidence was used, how the answer was checked, and how to replay the session.
|
||||
```
|
||||
|
||||
## Demo Flow
|
||||
|
||||
1. Start the service with the `mvp-demo` profile.
|
||||
2. Run the fixed payment-timeout request.
|
||||
3. Open `mvp/demo/output/chat-response.json`.
|
||||
4. Open `mvp/demo/output/trace-response.json`.
|
||||
5. Point to evidence tools and verifier evaluation.
|
||||
6. Submit feedback and show it is attached to the same session.
|
||||
|
||||
## Commands
|
||||
|
||||
Start service:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
Run the demo from another terminal:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
Optional custom session:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
||||
```
|
||||
|
||||
## What To Show
|
||||
|
||||
### 1. User-Facing Answer
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
|
||||
```
|
||||
|
||||
### 2. Evidence Trace
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/trace-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the important Agent engineering part.
|
||||
I can inspect which tools were called, what inputs they received,
|
||||
whether they succeeded, and what evidence preview was persisted.
|
||||
```
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.toolInvocations[*].toolName`
|
||||
- `data.toolInvocations[*].inputParams`
|
||||
- `data.toolInvocations[*].outputPreview`
|
||||
- `data.toolInvocations[*].success`
|
||||
|
||||
### 3. Verifier / Self-Evaluation
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.session.selfEvaluation`
|
||||
- `data.summary.hasVerifierEvaluation`
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
The final answer is not just raw Executor output.
|
||||
It is checked by a verifier or self-evaluation layer using the persisted trace.
|
||||
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
|
||||
```
|
||||
|
||||
### 4. Feedback Loop
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
Then re-query trace if needed.
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
Feedback is attached to the same diagnosis session.
|
||||
That makes it possible to mine useful / not useful cases later.
|
||||
```
|
||||
|
||||
### 5. Regression Story
|
||||
|
||||
Mention, do not deep dive unless asked:
|
||||
|
||||
```text
|
||||
For repeatability, I also built an offline eval baseline.
|
||||
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
|
||||
The two are separate on purpose: demo for human review, eval for automated signal.
|
||||
```
|
||||
|
||||
## Strong Interview Framing
|
||||
|
||||
Use this phrasing:
|
||||
|
||||
```text
|
||||
I focused on the Agent engineering surface:
|
||||
traceability, evidence persistence, verifier gating, feedback, and regression checks.
|
||||
The model answer is only one part of the system.
|
||||
The more important part is whether we can audit and improve the answer after it is produced.
|
||||
```
|
||||
|
||||
## Known Limits To Say Proactively
|
||||
|
||||
```text
|
||||
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
|
||||
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
|
||||
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
|
||||
```
|
||||
Reference in New Issue
Block a user