147 lines
3.5 KiB
Markdown
147 lines
3.5 KiB
Markdown
# Interview Walkthrough: MVP Diagnosis Agent
|
|
|
|
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
|
|
|
|
## 30-Second Summary
|
|
|
|
```text
|
|
This is an enterprise diagnosis Agent MVP.
|
|
It takes a payment-timeout question, plans the investigation, calls evidence tools,
|
|
checks the answer through a verifier, persists the full trace, and accepts feedback.
|
|
```
|
|
|
|
The important claim is not "the model answered once." The claim is:
|
|
|
|
```text
|
|
The system can show what evidence was used, how the answer was checked, and how to replay the session.
|
|
```
|
|
|
|
## Demo Flow
|
|
|
|
1. Start the service with the `mvp-demo` profile.
|
|
2. Run the fixed payment-timeout request.
|
|
3. Open `mvp/demo/output/chat-response.json`.
|
|
4. Open `mvp/demo/output/trace-response.json`.
|
|
5. Point to evidence tools and verifier evaluation.
|
|
6. Submit feedback and show it is attached to the same session.
|
|
|
|
## Commands
|
|
|
|
Start service:
|
|
|
|
```powershell
|
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
|
```
|
|
|
|
Run the demo from another terminal:
|
|
|
|
```powershell
|
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
|
```
|
|
|
|
Optional custom session:
|
|
|
|
```powershell
|
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
|
```
|
|
|
|
## What To Show
|
|
|
|
### 1. User-Facing Answer
|
|
|
|
File:
|
|
|
|
```text
|
|
mvp/demo/output/chat-response.json
|
|
```
|
|
|
|
Say:
|
|
|
|
```text
|
|
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
|
|
```
|
|
|
|
### 2. Evidence Trace
|
|
|
|
File:
|
|
|
|
```text
|
|
mvp/demo/output/trace-response.json
|
|
```
|
|
|
|
Say:
|
|
|
|
```text
|
|
This is the important Agent engineering part.
|
|
I can inspect which tools were called, what inputs they received,
|
|
whether they succeeded, and what evidence preview was persisted.
|
|
```
|
|
|
|
Point to:
|
|
|
|
- `data.toolInvocations[*].toolName`
|
|
- `data.toolInvocations[*].inputParams`
|
|
- `data.toolInvocations[*].outputPreview`
|
|
- `data.toolInvocations[*].success`
|
|
|
|
### 3. Verifier / Self-Evaluation
|
|
|
|
Point to:
|
|
|
|
- `data.session.selfEvaluation`
|
|
- `data.summary.hasVerifierEvaluation`
|
|
|
|
Say:
|
|
|
|
```text
|
|
The final answer is not just raw Executor output.
|
|
It is checked by a verifier or self-evaluation layer using the persisted trace.
|
|
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
|
|
```
|
|
|
|
### 4. Feedback Loop
|
|
|
|
File:
|
|
|
|
```text
|
|
mvp/demo/output/feedback-response.json
|
|
```
|
|
|
|
Then re-query trace if needed.
|
|
|
|
Say:
|
|
|
|
```text
|
|
Feedback is attached to the same diagnosis session.
|
|
That makes it possible to mine useful / not useful cases later.
|
|
```
|
|
|
|
### 5. Regression Story
|
|
|
|
Mention, do not deep dive unless asked:
|
|
|
|
```text
|
|
For repeatability, I also built an offline eval baseline.
|
|
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
|
|
The two are separate on purpose: demo for human review, eval for automated signal.
|
|
```
|
|
|
|
## Strong Interview Framing
|
|
|
|
Use this phrasing:
|
|
|
|
```text
|
|
I focused on the Agent engineering surface:
|
|
traceability, evidence persistence, verifier gating, feedback, and regression checks.
|
|
The model answer is only one part of the system.
|
|
The more important part is whether we can audit and improve the answer after it is produced.
|
|
```
|
|
|
|
## Known Limits To Say Proactively
|
|
|
|
```text
|
|
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
|
|
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
|
|
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
|
|
```
|