feat(demo): add interview quality audit
This commit is contained in:
+8
-3
@@ -8,7 +8,9 @@
|
||||
- `interview-walkthrough.md`:面试讲解话术。
|
||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
|
||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
|
||||
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
|
||||
@@ -37,7 +39,7 @@ http://localhost:9900
|
||||
最快方式:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会生成:
|
||||
@@ -46,6 +48,7 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
手动请求:
|
||||
@@ -85,6 +88,8 @@ Invoke-RestMethod `
|
||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
|
||||
|
||||
## 5. 提交反馈
|
||||
|
||||
@@ -176,8 +181,8 @@ AIOps 主线:
|
||||
|
||||
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
|
||||
|
||||
- `scripts/run-payment-timeout-demo.ps1` 跑主路径。
|
||||
- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
|
||||
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
|
||||
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 10/10 通过。
|
||||
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
|
||||
|
||||
这样可以同时展示真实链路和确定性回归能力。
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
# 面试追问 Q&A
|
||||
|
||||
## 为什么不用普通 Chatbot?
|
||||
|
||||
这个项目的重点不是生成一段诊断文本,而是把诊断拆成可审计链路:Planner 拆解问题,Executor 调工具拿证据,Gatekeeper 用代码核验证据引用,Verifier 判断可推导性,Composer 生成最终表达。每次运行都能通过同一个 `sessionId` 回放。
|
||||
|
||||
## 为什么 RAG 要做成显式工具?
|
||||
|
||||
`lookup_knowledge` 保持显式工具调用,才能在 `tool_invocation` 里看到 Agent 查了什么、命中了什么、相关性等级是什么,以及最终答案是否真的使用了这些证据。隐式 Advisor 更方便,但不利于审计 Agent 决策。
|
||||
|
||||
## 怎么防止 Executor 幻觉?
|
||||
|
||||
Executor 不直接负责最终用户答案,而是输出 `executor_evidence_v2` 的微观事实和证据引用。Gatekeeper 会校验 `source_invocation_id`、`raw_path`、`evidence_excerpt` 是否真实存在;Verifier 再判断 claim 是否能由已验真的证据推出;Composer 只表达 Verifier 允许输出的内容。
|
||||
|
||||
## LOW_CONFID 是失败吗?
|
||||
|
||||
不是。`LOW_CONFID` 表示当前证据不足以支撑强结论,但系统仍然可以安全表达已确认事实和缺失信息。面试时可以把它作为“没有证据就不强答”的质量门禁,而不是模型能力失败。
|
||||
|
||||
## Prompt 改了怎么审计?
|
||||
|
||||
Chat verifier evaluation 里会记录 `prompt_audit.version`,并列出 planner、executor、verifier、composer 的 Prompt 版本和资源路径。它不保存完整 Prompt 文本,只保留用于回放和回归解释的紧凑元数据。
|
||||
|
||||
## Gatekeeper 改了怎么审计?
|
||||
|
||||
Gatekeeper 结果里记录 `gatekeeper_result.rule_set_version` 和已启用规则元数据摘要。规则执行仍是确定性 Java 代码,版本和规则元数据用于解释“这次引用验真用的是哪套规则”。
|
||||
|
||||
## 为什么现在不拆 SubAgent?
|
||||
|
||||
当前 MVP 的主要风险不是 Agent 数量不够,而是证据、验证和回归是否稳定。文档里的演进路线把 SubAgent 放在 P2:等故障类型、工具权限和评测集足够明确后再拆,避免只是移动复杂度。
|
||||
|
||||
## 为什么 baseline 比 live demo 更重要?
|
||||
|
||||
live demo 证明链路在当前环境能跑通,但 LLM 和外部依赖会波动。`mvp/eval` 的固定 fixture baseline 是确定性回归来源,用来判断 Prompt、工具、Gatekeeper、Verifier 或 Composer 的改动有没有让系统退化。
|
||||
@@ -0,0 +1,161 @@
|
||||
param(
|
||||
[string]$BaseUrl = "http://localhost:9900",
|
||||
[string]$SessionId = "mvp-demo-interview-payment-timeout-001",
|
||||
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
|
||||
[string]$OutputDir = "$PSScriptRoot/../output"
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
function Test-ServiceReachable {
|
||||
param([string]$Url)
|
||||
|
||||
try {
|
||||
$request = [System.Net.WebRequest]::Create($Url)
|
||||
$request.Method = "GET"
|
||||
$request.Timeout = 5000
|
||||
$response = $request.GetResponse()
|
||||
$response.Close()
|
||||
return $true
|
||||
} catch [System.Net.WebException] {
|
||||
if ($_.Exception.Response -ne $null) {
|
||||
$_.Exception.Response.Close()
|
||||
return $true
|
||||
}
|
||||
return $false
|
||||
}
|
||||
}
|
||||
|
||||
function Get-TraceData {
|
||||
param($TraceResponse)
|
||||
|
||||
if ($TraceResponse.PSObject.Properties.Name -contains "data") {
|
||||
return $TraceResponse.data
|
||||
}
|
||||
return $TraceResponse
|
||||
}
|
||||
|
||||
function Get-SelfEvaluation {
|
||||
param($TraceData)
|
||||
|
||||
if ($null -eq $TraceData -or $null -eq $TraceData.session) {
|
||||
return $null
|
||||
}
|
||||
return $TraceData.session.selfEvaluation
|
||||
}
|
||||
|
||||
function Get-ToolNames {
|
||||
param($TraceData)
|
||||
|
||||
if ($null -eq $TraceData -or $null -eq $TraceData.toolInvocations) {
|
||||
return @()
|
||||
}
|
||||
return @($TraceData.toolInvocations | ForEach-Object { $_.toolName } | Where-Object { $_ } | Sort-Object -Unique)
|
||||
}
|
||||
|
||||
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||
|
||||
Write-Host "Running interview demo preflight..."
|
||||
Write-Host "BaseUrl: $BaseUrl"
|
||||
Write-Host "SessionId: $SessionId"
|
||||
|
||||
if (-not (Test-ServiceReachable -Url $BaseUrl)) {
|
||||
throw "Service is not reachable: $BaseUrl. Start the app with mvp-demo profile first: mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo"
|
||||
}
|
||||
|
||||
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||
$request.Id = $SessionId
|
||||
$body = $request | ConvertTo-Json -Depth 8
|
||||
|
||||
$chatRequest = @{
|
||||
Method = "Post"
|
||||
Uri = "$BaseUrl/api/chat"
|
||||
ContentType = "application/json; charset=utf-8"
|
||||
Body = $body
|
||||
}
|
||||
$chat = Invoke-RestMethod @chatRequest
|
||||
|
||||
$chatPath = Join-Path $OutputDir "chat-response.json"
|
||||
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
||||
|
||||
$traceRequest = @{
|
||||
Method = "Get"
|
||||
Uri = "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||
}
|
||||
$trace = Invoke-RestMethod @traceRequest
|
||||
|
||||
$tracePath = Join-Path $OutputDir "trace-response.json"
|
||||
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
$feedbackRequest = @{
|
||||
Method = "Post"
|
||||
Uri = "$BaseUrl/api/feedback"
|
||||
ContentType = "application/json; charset=utf-8"
|
||||
Body = $feedbackBody
|
||||
}
|
||||
$feedback = Invoke-RestMethod @feedbackRequest
|
||||
|
||||
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
|
||||
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
|
||||
|
||||
$traceData = Get-TraceData -TraceResponse $trace
|
||||
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
|
||||
$verifierEvaluation = $null
|
||||
if ($null -ne $selfEvaluation) {
|
||||
$verifierEvaluation = $selfEvaluation.verifier_evaluation
|
||||
}
|
||||
|
||||
$gatekeeperResult = $null
|
||||
$promptAudit = $null
|
||||
if ($null -ne $verifierEvaluation) {
|
||||
$gatekeeperResult = $verifierEvaluation.gatekeeper_result
|
||||
$promptAudit = $verifierEvaluation.prompt_audit
|
||||
}
|
||||
|
||||
$verdict = $null
|
||||
$gatekeeperStatus = $null
|
||||
$gatekeeperRuleSetVersion = $null
|
||||
$promptAuditVersion = $null
|
||||
if ($null -ne $verifierEvaluation) {
|
||||
$verdict = $verifierEvaluation.verdict
|
||||
}
|
||||
if ($null -ne $gatekeeperResult) {
|
||||
$gatekeeperStatus = $gatekeeperResult.status
|
||||
$gatekeeperRuleSetVersion = $gatekeeperResult.rule_set_version
|
||||
}
|
||||
if ($null -ne $promptAudit) {
|
||||
$promptAuditVersion = $promptAudit.version
|
||||
}
|
||||
$toolNames = Get-ToolNames -TraceData $traceData
|
||||
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
||||
|
||||
$summary = [ordered]@{
|
||||
sessionId = $SessionId
|
||||
baseUrl = $BaseUrl
|
||||
chatSuccess = $chat.data.success
|
||||
verdict = $verdict
|
||||
gatekeeperStatus = $gatekeeperStatus
|
||||
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
|
||||
promptAuditVersion = $promptAuditVersion
|
||||
toolNames = $toolNames
|
||||
paths = [ordered]@{
|
||||
chat = $chatPath
|
||||
trace = $tracePath
|
||||
feedback = $feedbackPath
|
||||
summary = $summaryPath
|
||||
}
|
||||
}
|
||||
|
||||
$summary | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path $summaryPath
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "Interview demo preflight completed."
|
||||
Write-Host "Verdict: $($summary.verdict)"
|
||||
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
|
||||
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
|
||||
Write-Host "Summary: $summaryPath"
|
||||
@@ -37,7 +37,7 @@ http://localhost:9900
|
||||
推荐使用固定脚本:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会写出:
|
||||
@@ -46,6 +46,7 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
现场话术:
|
||||
@@ -98,6 +99,8 @@ data.toolInvocations[*].outputPreview
|
||||
data.toolInvocations[*].retrievalLayer
|
||||
data.toolInvocations[*].relevanceLevel
|
||||
data.summary.hasVerifierEvaluation
|
||||
data.session.selfEvaluation.verifier_evaluation.prompt_audit.version
|
||||
data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version
|
||||
```
|
||||
|
||||
现场话术:
|
||||
@@ -155,6 +158,10 @@ Verifier 不做新检索,只看工具 trace 汇总。
|
||||
如果 PASS,就输出原答案。
|
||||
如果 LOW_CONFID,可以补证据或加低置信提示。
|
||||
如果 REJECT,就降级输出,只保留已确认信息。
|
||||
|
||||
Prompt 和 Gatekeeper 的版本也会进入 trace。
|
||||
`prompt_audit.version` 用于说明本次 Chat 使用哪套 Prompt 契约,`gatekeeper_result.rule_set_version` 用于说明引用验真的规则版本。
|
||||
固定 fixture baseline 是回归判断来源,live demo 主要证明当前环境链路可跑通。
|
||||
```
|
||||
|
||||
## 7. 展示反馈闭环
|
||||
@@ -226,6 +233,7 @@ AIOps 有两个模式。
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
降级话术:
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Trace 检查清单
|
||||
|
||||
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。
|
||||
运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
|
||||
|
||||
## 1. Session
|
||||
|
||||
@@ -11,6 +11,8 @@
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
|
||||
Reference in New Issue
Block a user