2.8 KiB
2.8 KiB
Trace Inspection Checklist
Use this checklist after running scripts/run-payment-timeout-demo.ps1.
Session
| JSON path | What to check | Interview point |
|---|---|---|
data.session.sessionId |
Matches mvp-demo-payment-timeout-001 |
One session id connects chat, tools, verifier, feedback, and trace. |
data.session.query |
Contains the payment-timeout question | The trace records the original user intent. |
data.session.answer |
Contains the final diagnosis answer | The final answer is not detached from the trace. |
data.session.selfEvaluation |
Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
data.session.feedback |
Becomes useful after feedback submission |
User feedback is attached to the same diagnosis session. |
Agent Steps
| JSON path | What to check | Interview point |
|---|---|---|
data.steps[*].agentName |
Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
data.steps[*].thought |
High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
data.steps[*].durationMs |
Step duration | The trace can support cost and latency review. |
data.steps[*].tokenCount |
Token count where available | The trace can support model-cost review. |
Tool Evidence
| JSON path | What to check | Interview point |
|---|---|---|
data.toolInvocations[*].toolName |
Includes evidence tools such as lookup_knowledge, query_logs, query_metrics |
The Agent uses tools, not unsupported guesses. |
data.toolInvocations[*].inputParams |
Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
data.toolInvocations[*].outputPreview |
Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
data.toolInvocations[*].success |
Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
data.toolInvocations[*].retrievalDetails |
Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
Summary
| JSON path | What to check | Interview point |
|---|---|---|
data.summary.persistedStepCount |
Step rows were persisted | The trace is backed by storage, not only response memory. |
data.summary.persistedToolCallCount |
Tool rows were persisted | Evidence survives the request. |
data.summary.hasVerifierEvaluation |
Verifier evaluation exists | The final answer passed through a quality gate. |
data.summary.hasFeedback |
Feedback exists after feedback step | Human feedback closes the loop. |
What Good Looks Like
same session id
-> final answer
-> persisted agent steps
-> persisted evidence tool calls
-> verifier/self-evaluation
-> feedback attached to the same session