docs: reorganize MVP interview documentation
This commit is contained in:
@@ -1,52 +1,53 @@
|
||||
# Trace Inspection Checklist
|
||||
# Trace 检查清单
|
||||
|
||||
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
|
||||
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。
|
||||
|
||||
## Session
|
||||
## 1. Session
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
|
||||
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
|
||||
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
|
||||
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
|
||||
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | 一个 session id 串起 chat、工具、verifier、feedback 和 trace |
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
|
||||
|
||||
## Agent Steps
|
||||
## 2. Agent 步骤
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
|
||||
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
|
||||
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
|
||||
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.steps[*].agentName` | 是否有 Planner / Executor / Verifier 或等价步骤 | 流程被拆成可检查的 Agent 步骤 |
|
||||
| `data.steps[*].thought` | 是否有高层步骤摘要 | 内部过程可审计,不只看最终文本 |
|
||||
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
||||
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
||||
|
||||
## Tool Evidence
|
||||
## 3. 工具证据
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
|
||||
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
|
||||
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
|
||||
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
|
||||
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.toolInvocations[*].toolName` | 是否包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 | Agent 通过工具收集证据,而不是无依据猜测 |
|
||||
| `data.toolInvocations[*].inputParams` | 是否能看到每个工具的入参 | 工具输入可审计、可调试 |
|
||||
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
|
||||
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
|
||||
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
|
||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||
|
||||
## Summary
|
||||
## 4. Summary
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
|
||||
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
|
||||
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
|
||||
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.summary.persistedStepCount` | step 行是否持久化 | Trace 来自存储,不是响应内存 |
|
||||
| `data.summary.persistedToolCallCount` | tool 行是否持久化 | 工具证据在请求结束后仍可回放 |
|
||||
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
||||
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
||||
|
||||
## What Good Looks Like
|
||||
## 5. 好的结果长什么样
|
||||
|
||||
```text
|
||||
same session id
|
||||
-> final answer
|
||||
-> persisted agent steps
|
||||
-> persisted evidence tool calls
|
||||
-> verifier/self-evaluation
|
||||
同一个 session id
|
||||
-> 最终答案
|
||||
-> 持久化 agent steps
|
||||
-> 持久化 evidence tool calls
|
||||
-> verifier / self-evaluation
|
||||
-> feedback attached to the same session
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user