feat(graph): complete stategraph cleanup and acceptance
This commit is contained in:
+9
-2
@@ -8,7 +8,7 @@
|
||||
- `interview-walkthrough.md`:面试讲解话术。
|
||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,绑定 exact runId,强制校验 Run orchestration trace,并输出 Chat、Trace、反馈和 summary。
|
||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||
@@ -51,6 +51,8 @@ mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
自动化验收应传入唯一 `-SessionId`,并用 `-OutputDir target/...` 避免覆盖仓库样例。脚本从 Chat 响应取得 exact `runId`,缺少 `data.run.orchestrationTrace` 或 version/final node/termination reason/transitions/degraded/evidence retry count 时会立即失败。summary 额外包含 `orchestrationVersion`、`finalNode`、`terminationReason`、`degraded`、`transitionCount` 和 `evidenceRetryCount`。
|
||||
|
||||
手动请求:
|
||||
|
||||
```powershell
|
||||
@@ -100,6 +102,10 @@ Invoke-RestMethod `
|
||||
- `data.runId` 等于 `$runId`
|
||||
- `data.session.sessionId` 等于 Chat session id
|
||||
- `data.run.runId` 等于 `$runId`
|
||||
- `data.run.orchestrationTrace.version` 非空
|
||||
- `data.run.orchestrationTrace.final_node` 和 `termination_reason` 非空
|
||||
- `data.run.orchestrationTrace.transitions` 是本次 Graph 的条件边记录
|
||||
- `data.run.orchestrationTrace.degraded` 和 `evidence_retry_count` 记录安全降级与补证据次数
|
||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||
@@ -174,9 +180,10 @@ Chat 主线:
|
||||
```text
|
||||
一个 session id + 一个 run id
|
||||
-> 用户问题
|
||||
-> 多 Agent 执行
|
||||
-> bounded StateGraph(Planner / Executor / Gatekeeper / Verified Input / Verifier / Composer / Fallback)
|
||||
-> 证据工具
|
||||
-> Verifier / self_evaluation
|
||||
-> run.orchestrationTrace 路由摘要
|
||||
-> 最终答案
|
||||
-> 用户反馈
|
||||
-> Trace API 回放
|
||||
|
||||
@@ -79,6 +79,15 @@ $chatPath = Join-Path $OutputDir "chat-response.json"
|
||||
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if ($chat.data.success -ne $true) {
|
||||
throw "Chat response was not successful."
|
||||
}
|
||||
if ([string]::IsNullOrWhiteSpace([string]$chat.data.answer)) {
|
||||
throw "Chat response did not include a non-empty answer."
|
||||
}
|
||||
if ($chat.data.sessionId -ne $SessionId) {
|
||||
throw "Chat response sessionId '$($chat.data.sessionId)' did not match requested sessionId '$SessionId'."
|
||||
}
|
||||
if (-not $runId) {
|
||||
throw "Chat response did not include runId; exact trace verification cannot continue."
|
||||
}
|
||||
@@ -92,6 +101,51 @@ $trace = Invoke-RestMethod @traceRequest
|
||||
$tracePath = Join-Path $OutputDir "trace-response.json"
|
||||
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
||||
|
||||
$traceData = Get-TraceData -TraceResponse $trace
|
||||
if ($null -eq $traceData -or $null -eq $traceData.run) {
|
||||
throw "Exact Trace response did not include data.run."
|
||||
}
|
||||
if ($traceData.runId -ne $runId -or $traceData.run.runId -ne $runId) {
|
||||
throw "Exact Trace runId did not match Chat runId '$runId'."
|
||||
}
|
||||
if ($traceData.run.sessionId -ne $SessionId) {
|
||||
throw "Exact Trace run did not belong to requested sessionId '$SessionId'."
|
||||
}
|
||||
|
||||
$orchestrationTrace = $traceData.run.orchestrationTrace
|
||||
if ($null -eq $orchestrationTrace) {
|
||||
throw "Exact Trace data.run.orchestrationTrace is missing."
|
||||
}
|
||||
foreach ($field in @("version", "final_node", "termination_reason")) {
|
||||
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field) -or
|
||||
[string]::IsNullOrWhiteSpace([string]$orchestrationTrace.$field)) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
|
||||
}
|
||||
}
|
||||
foreach ($field in @("transitions", "degraded", "evidence_retry_count")) {
|
||||
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field)) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
|
||||
}
|
||||
}
|
||||
if ($null -eq $orchestrationTrace.transitions) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.transitions must be an array."
|
||||
}
|
||||
if ([int]$orchestrationTrace.evidence_retry_count -lt 0) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.evidence_retry_count must not be negative."
|
||||
}
|
||||
if ($traceData.run.status -ne "SUCCESS" -or $traceData.run.agentFlow -ne "CHAT") {
|
||||
throw "Exact Trace run must be CHAT/SUCCESS."
|
||||
}
|
||||
if ([string]::IsNullOrWhiteSpace([string]$traceData.run.answer)) {
|
||||
throw "Exact Trace run did not include a non-empty answer."
|
||||
}
|
||||
if (@($traceData.steps).Count -eq 0 -or @($traceData.toolInvocations).Count -eq 0) {
|
||||
throw "Exact Trace did not include both Agent steps and tool invocation evidence."
|
||||
}
|
||||
if ($null -eq $traceData.run.selfEvaluation) {
|
||||
throw "Exact Trace run did not include selfEvaluation."
|
||||
}
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
@@ -105,11 +159,13 @@ $feedbackRequest = @{
|
||||
Body = $feedbackBody
|
||||
}
|
||||
$feedback = Invoke-RestMethod @feedbackRequest
|
||||
if ($feedback.success -ne $true) {
|
||||
throw "Feedback request was not successful for runId '$runId'."
|
||||
}
|
||||
|
||||
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
|
||||
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
|
||||
|
||||
$traceData = Get-TraceData -TraceResponse $trace
|
||||
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
|
||||
$verifierEvaluation = $null
|
||||
if ($null -ne $selfEvaluation) {
|
||||
@@ -138,6 +194,7 @@ if ($null -ne $promptAudit) {
|
||||
$promptAuditVersion = $promptAudit.version
|
||||
}
|
||||
$toolNames = Get-ToolNames -TraceData $traceData
|
||||
$transitionCount = @($orchestrationTrace.transitions).Count
|
||||
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
||||
|
||||
$summary = [ordered]@{
|
||||
@@ -149,6 +206,12 @@ $summary = [ordered]@{
|
||||
gatekeeperStatus = $gatekeeperStatus
|
||||
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
|
||||
promptAuditVersion = $promptAuditVersion
|
||||
orchestrationVersion = $orchestrationTrace.version
|
||||
finalNode = $orchestrationTrace.final_node
|
||||
terminationReason = $orchestrationTrace.termination_reason
|
||||
degraded = [bool]$orchestrationTrace.degraded
|
||||
transitionCount = $transitionCount
|
||||
evidenceRetryCount = [int]$orchestrationTrace.evidence_retry_count
|
||||
toolNames = $toolNames
|
||||
paths = [ordered]@{
|
||||
chat = $chatPath
|
||||
@@ -165,4 +228,6 @@ Write-Host "Interview demo preflight completed."
|
||||
Write-Host "Verdict: $($summary.verdict)"
|
||||
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
|
||||
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
|
||||
Write-Host "Graph final node: $($summary.finalNode)"
|
||||
Write-Host "Graph termination: $($summary.terminationReason)"
|
||||
Write-Host "Summary: $summaryPath"
|
||||
|
||||
@@ -7,16 +7,30 @@
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
| `data.run.sessionId` | 是否等于本次 Chat 请求的唯一 sessionId | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.run.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.run.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.run.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.run.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.version` | 是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.run.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
## 2. StateGraph 路由
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.run.orchestrationTrace.version` | 是否存在当前 trace contract 版本 | 路由摘要可演进、可兼容 |
|
||||
| `data.run.orchestrationTrace.transitions[*]` | 是否记录实际经过的 Node 和 route | Graph 条件边不是从日志推断 |
|
||||
| `data.run.orchestrationTrace.final_node` | 最终是 Composer 还是 Fallback | 正常输出与安全降级明确区分 |
|
||||
| `data.run.orchestrationTrace.termination_reason` | 是否给出终止原因 | 每次 Run 都有可解释终点 |
|
||||
| `data.run.orchestrationTrace.degraded` | 是否发生安全降级 | fallback 是可审计行为 |
|
||||
| `data.run.orchestrationTrace.evidence_retry_count` | 是否为 0 或 1 | 补证据循环有硬上限 |
|
||||
| `interview-demo-summary.json.finalNode` 等摘要字段 | 是否与 exact Trace 一致 | summary 只消费 Run 路由真理源 |
|
||||
|
||||
`orchestrationTrace` 负责路由;`selfEvaluation` 负责证据和答案质量;AgentStep/ToolInvocation 负责详细执行与工具证据。三者不能互相替代。
|
||||
|
||||
## 3. Agent 步骤
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
@@ -25,7 +39,7 @@
|
||||
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
||||
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
||||
|
||||
## 3. 工具证据
|
||||
## 4. 工具证据
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
@@ -37,7 +51,7 @@
|
||||
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||
|
||||
## 4. Summary
|
||||
## 5. Summary
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
@@ -46,7 +60,7 @@
|
||||
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
||||
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
||||
|
||||
## 5. 好的结果长什么样
|
||||
## 6. 好的结果长什么样
|
||||
|
||||
```text
|
||||
同一个 session id + run id
|
||||
@@ -54,5 +68,6 @@
|
||||
-> 持久化 agent steps
|
||||
-> 持久化 evidence tool calls
|
||||
-> verifier / self-evaluation
|
||||
-> run.orchestrationTrace routing summary
|
||||
-> feedback attached to the same run
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user