feat(demo): add interview quality audit

This commit is contained in:
zhuyongxin
2026-07-09 11:18:49 +08:00
parent a6c2d4459c
commit 9c9a0024d4
37 changed files with 1162 additions and 86 deletions
+19 -2
View File
@@ -82,6 +82,23 @@ Prompt 层当前承担的门禁:
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
- AIOps payload 模式必须聚焦输入告警。
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
```json
{
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_executor",
"version": "chat-executor-v2",
"resource": "prompts/chat-executor-prompt.md"
}
]
}
```
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
## 4. Trace Hooks
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
@@ -196,7 +213,7 @@ Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier
diagnosis_session.self_evaluation.verifier_evaluation
```
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary` 和 `composer_output`,用于 Trace 回放。
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
## 7. AIOps 规则门禁
@@ -233,7 +250,7 @@ diagnosis_session.self_evaluation.aiops_rule_evaluation
- 同一工具调用次数上限。
- 工具超时的统一熔断。
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
- Prompt 版本记录和回滚。
- Prompt 版本回滚和更细粒度变更审计。
- Verifier 对 AIOps 报告的 LLM 级事实校验。
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
+8 -3
View File
@@ -8,7 +8,9 @@
- `interview-walkthrough.md`:面试讲解话术。
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
- `trace-inspection-checklist.md`:Trace 字段检查清单。
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
@@ -37,7 +39,7 @@ http://localhost:9900
最快方式:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
```
脚本会生成:
@@ -46,6 +48,7 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
mvp/demo/output/interview-demo-summary.json
```
手动请求:
@@ -85,6 +88,8 @@ Invoke-RestMethod `
- `data.steps` 包含 planner / executor / verifier 等步骤
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
## 5. 提交反馈
@@ -176,8 +181,8 @@ AIOps 主线:
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
- `scripts/run-payment-timeout-demo.ps1` 跑主路径。
- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 10/10 通过。
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
这样可以同时展示真实链路和确定性回归能力。
+33
View File
@@ -0,0 +1,33 @@
# 面试追问 Q&A
## 为什么不用普通 Chatbot?
这个项目的重点不是生成一段诊断文本,而是把诊断拆成可审计链路:Planner 拆解问题,Executor 调工具拿证据,Gatekeeper 用代码核验证据引用,Verifier 判断可推导性,Composer 生成最终表达。每次运行都能通过同一个 `sessionId` 回放。
## 为什么 RAG 要做成显式工具?
`lookup_knowledge` 保持显式工具调用,才能在 `tool_invocation` 里看到 Agent 查了什么、命中了什么、相关性等级是什么,以及最终答案是否真的使用了这些证据。隐式 Advisor 更方便,但不利于审计 Agent 决策。
## 怎么防止 Executor 幻觉?
Executor 不直接负责最终用户答案,而是输出 `executor_evidence_v2` 的微观事实和证据引用。Gatekeeper 会校验 `source_invocation_id`、`raw_path`、`evidence_excerpt` 是否真实存在;Verifier 再判断 claim 是否能由已验真的证据推出;Composer 只表达 Verifier 允许输出的内容。
## LOW_CONFID 是失败吗?
不是。`LOW_CONFID` 表示当前证据不足以支撑强结论,但系统仍然可以安全表达已确认事实和缺失信息。面试时可以把它作为“没有证据就不强答”的质量门禁,而不是模型能力失败。
## Prompt 改了怎么审计?
Chat verifier evaluation 里会记录 `prompt_audit.version`,并列出 planner、executor、verifier、composer 的 Prompt 版本和资源路径。它不保存完整 Prompt 文本,只保留用于回放和回归解释的紧凑元数据。
## Gatekeeper 改了怎么审计?
Gatekeeper 结果里记录 `gatekeeper_result.rule_set_version` 和已启用规则元数据摘要。规则执行仍是确定性 Java 代码,版本和规则元数据用于解释“这次引用验真用的是哪套规则”。
## 为什么现在不拆 SubAgent?
当前 MVP 的主要风险不是 Agent 数量不够,而是证据、验证和回归是否稳定。文档里的演进路线把 SubAgent 放在 P2:等故障类型、工具权限和评测集足够明确后再拆,避免只是移动复杂度。
## 为什么 baseline 比 live demo 更重要?
live demo 证明链路在当前环境能跑通,但 LLM 和外部依赖会波动。`mvp/eval` 的固定 fixture baseline 是确定性回归来源,用来判断 Prompt、工具、Gatekeeper、Verifier 或 Composer 的改动有没有让系统退化。
@@ -0,0 +1,161 @@
param(
[string]$BaseUrl = "http://localhost:9900",
[string]$SessionId = "mvp-demo-interview-payment-timeout-001",
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
[string]$OutputDir = "$PSScriptRoot/../output"
)
$ErrorActionPreference = "Stop"
function Test-ServiceReachable {
param([string]$Url)
try {
$request = [System.Net.WebRequest]::Create($Url)
$request.Method = "GET"
$request.Timeout = 5000
$response = $request.GetResponse()
$response.Close()
return $true
} catch [System.Net.WebException] {
if ($_.Exception.Response -ne $null) {
$_.Exception.Response.Close()
return $true
}
return $false
}
}
function Get-TraceData {
param($TraceResponse)
if ($TraceResponse.PSObject.Properties.Name -contains "data") {
return $TraceResponse.data
}
return $TraceResponse
}
function Get-SelfEvaluation {
param($TraceData)
if ($null -eq $TraceData -or $null -eq $TraceData.session) {
return $null
}
return $TraceData.session.selfEvaluation
}
function Get-ToolNames {
param($TraceData)
if ($null -eq $TraceData -or $null -eq $TraceData.toolInvocations) {
return @()
}
return @($TraceData.toolInvocations | ForEach-Object { $_.toolName } | Where-Object { $_ } | Sort-Object -Unique)
}
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
Write-Host "Running interview demo preflight..."
Write-Host "BaseUrl: $BaseUrl"
Write-Host "SessionId: $SessionId"
if (-not (Test-ServiceReachable -Url $BaseUrl)) {
throw "Service is not reachable: $BaseUrl. Start the app with mvp-demo profile first: mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo"
}
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
$request.Id = $SessionId
$body = $request | ConvertTo-Json -Depth 8
$chatRequest = @{
Method = "Post"
Uri = "$BaseUrl/api/chat"
ContentType = "application/json; charset=utf-8"
Body = $body
}
$chat = Invoke-RestMethod @chatRequest
$chatPath = Join-Path $OutputDir "chat-response.json"
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
$traceRequest = @{
Method = "Get"
Uri = "$BaseUrl/api/diagnosis/$SessionId/trace"
}
$trace = Invoke-RestMethod @traceRequest
$tracePath = Join-Path $OutputDir "trace-response.json"
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
$feedbackBody = @{
sessionId = $SessionId
feedback = "useful"
} | ConvertTo-Json
$feedbackRequest = @{
Method = "Post"
Uri = "$BaseUrl/api/feedback"
ContentType = "application/json; charset=utf-8"
Body = $feedbackBody
}
$feedback = Invoke-RestMethod @feedbackRequest
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
$traceData = Get-TraceData -TraceResponse $trace
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
$verifierEvaluation = $null
if ($null -ne $selfEvaluation) {
$verifierEvaluation = $selfEvaluation.verifier_evaluation
}
$gatekeeperResult = $null
$promptAudit = $null
if ($null -ne $verifierEvaluation) {
$gatekeeperResult = $verifierEvaluation.gatekeeper_result
$promptAudit = $verifierEvaluation.prompt_audit
}
$verdict = $null
$gatekeeperStatus = $null
$gatekeeperRuleSetVersion = $null
$promptAuditVersion = $null
if ($null -ne $verifierEvaluation) {
$verdict = $verifierEvaluation.verdict
}
if ($null -ne $gatekeeperResult) {
$gatekeeperStatus = $gatekeeperResult.status
$gatekeeperRuleSetVersion = $gatekeeperResult.rule_set_version
}
if ($null -ne $promptAudit) {
$promptAuditVersion = $promptAudit.version
}
$toolNames = Get-ToolNames -TraceData $traceData
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
$summary = [ordered]@{
sessionId = $SessionId
baseUrl = $BaseUrl
chatSuccess = $chat.data.success
verdict = $verdict
gatekeeperStatus = $gatekeeperStatus
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
promptAuditVersion = $promptAuditVersion
toolNames = $toolNames
paths = [ordered]@{
chat = $chatPath
trace = $tracePath
feedback = $feedbackPath
summary = $summaryPath
}
}
$summary | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path $summaryPath
Write-Host ""
Write-Host "Interview demo preflight completed."
Write-Host "Verdict: $($summary.verdict)"
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
Write-Host "Summary: $summaryPath"
+9 -1
View File
@@ -37,7 +37,7 @@ http://localhost:9900
推荐使用固定脚本:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
```
脚本会写出:
@@ -46,6 +46,7 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
mvp/demo/output/interview-demo-summary.json
```
现场话术:
@@ -98,6 +99,8 @@ data.toolInvocations[*].outputPreview
data.toolInvocations[*].retrievalLayer
data.toolInvocations[*].relevanceLevel
data.summary.hasVerifierEvaluation
data.session.selfEvaluation.verifier_evaluation.prompt_audit.version
data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version
```
现场话术:
@@ -155,6 +158,10 @@ Verifier 不做新检索,只看工具 trace 汇总。
如果 PASS,就输出原答案。
如果 LOW_CONFID,可以补证据或加低置信提示。
如果 REJECT,就降级输出,只保留已确认信息。
Prompt 和 Gatekeeper 的版本也会进入 trace。
`prompt_audit.version` 用于说明本次 Chat 使用哪套 Prompt 契约,`gatekeeper_result.rule_set_version` 用于说明引用验真的规则版本。
固定 fixture baseline 是回归判断来源,live demo 主要证明当前环境链路可跑通。
```
## 7. 展示反馈闭环
@@ -226,6 +233,7 @@ AIOps 有两个模式。
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
mvp/demo/output/interview-demo-summary.json
```
降级话术:
+3 -1
View File
@@ -1,6 +1,6 @@
# Trace 检查清单
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。
运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
## 1. Session
@@ -11,6 +11,8 @@
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
## 2. Agent 步骤
+8 -4
View File
@@ -29,10 +29,10 @@ The baseline evaluates saved trace fixtures. It does not start the application a
The committed baseline currently contains:
```text
10 fixed cases
10 passing fixture evaluations
4 PASS verdicts
5 LOW_CONFID verdicts
12 fixed cases
12 passing fixture evaluations
5 PASS verdicts
6 LOW_CONFID verdicts
1 REJECT verdict
```
@@ -44,6 +44,8 @@ The V2 evidence-pipeline matrix covers:
- Unsupported claim filtering before the final answer.
- Composer fallback rendering without raw Executor JSON leakage.
- Gatekeeper rule set version audit for new matrix fixtures.
- Prompt audit version checks for planner, executor, verifier, and composer prompts.
- Gatekeeper rule metadata checks for enabled rule id and default severity.
## Verification
@@ -80,6 +82,8 @@ Stage 5 adds these V2 checks:
- `claim_checks` must be structurally auditable.
- Composer output must record whether normal parsing or fallback rendering was used.
- Gatekeeper rule set version can be asserted per fixture.
- Prompt audit version and per-prompt versions can be asserted per fixture.
- Gatekeeper rule metadata can be required per fixture.
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
+54
View File
@@ -17,6 +17,33 @@
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["数据库连接池"]
},
{
"id": "prompt-gatekeeper-audit-closure",
"title": "Prompt and Gatekeeper audit closure",
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
"traceFixture": "prompt-gatekeeper-audit-closure-pass.json",
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics"],
"allowedVerdicts": ["PASS"],
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["数据库连接池"],
"requirePromptAudit": true,
"expectedPromptAuditVersion": "chat-prompts-v1",
"expectedPromptVersions": {
"chat_planner": "chat-planner-v1",
"chat_executor": "chat-executor-v2",
"chat_verifier": "chat-verifier-v2",
"chat_composer": "chat-composer-v1"
},
"requireGatekeeperRules": true
},
{
"id": "hikari-no-evidence-negative-observation",
"title": "Hikari no-evidence negative observation",
@@ -124,6 +151,33 @@
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["主库故障"]
},
{
"id": "audit-metadata-low-confid",
"title": "Audit metadata low confidence",
"question": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
"traceFixture": "audit-metadata-low-confid.json",
"expectedRootCauseKeywords": ["超时", "证据"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["LOW_CONFID"],
"forbiddenAnswerKeywords": ["已经确认"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["主库故障"],
"requirePromptAudit": true,
"expectedPromptAuditVersion": "chat-prompts-v1",
"expectedPromptVersions": {
"chat_planner": "chat-planner-v1",
"chat_executor": "chat-executor-v2",
"chat_verifier": "chat-verifier-v2",
"chat_composer": "chat-composer-v1"
},
"requireGatekeeperRules": true
},
{
"id": "composer-fallback-no-raw-json",
"title": "Composer fallback no raw JSON",
@@ -0,0 +1,192 @@
{
"session": {
"sessionId": "eval-audit-metadata-low-confid",
"query": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 45000,
"toolCallCount": 1,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.42,
"critical_fact_count": 2,
"prompt_audit": {
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_planner",
"version": "chat-planner-v1",
"resource": "prompts/chat-planner-prompt.md"
},
{
"name": "chat_executor",
"version": "chat-executor-v2",
"resource": "prompts/chat-executor-prompt.md"
},
{
"name": "chat_verifier",
"version": "chat-verifier-v2",
"resource": "prompts/chat-verifier-prompt.md"
},
{
"name": "chat_composer",
"version": "chat-composer-v1",
"resource": "prompts/chat-composer-prompt.md"
}
]
},
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.invocation",
"description": "source_invocation_id must reference an existing tool invocation",
"enabled": true,
"default_severity": "reject"
},
{
"id": "evidence.excerpt",
"description": "evidence_excerpt must be supported by recorded evidence text",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-timeout",
"tool_name": "query_logs",
"source_invocation_id": 22,
"raw_path": "$.logs[0]",
"matched_text": "order api timeout",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-timeout",
"claim_type": "symptom",
"claim_text": "订单接口出现超时",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_logs",
"source_invocation_id": 22,
"raw_path": "$.logs[0]",
"evidence_excerpt": "order api timeout"
}
]
},
{
"claim_id": "claim-db-primary",
"claim_type": "root_cause",
"claim_text": "数据库主库故障导致订单超时",
"support_level": "weak",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_logs",
"source_invocation_id": 22,
"raw_path": "$.logs[0]",
"evidence_excerpt": "order api timeout"
}
]
}
],
"hypotheses": [],
"recommended_actions": [
{
"action_text": "补充查询数据库主库状态和错误日志",
"reason": "当前只有订单接口超时日志"
}
],
"missing_info": ["数据库主库状态", "数据库错误日志"]
},
"claim_checks": [
{
"claim_id": "claim-timeout",
"claim_text": "订单接口出现超时",
"claim_type": "symptom",
"verification": "direct_observation",
"detail": "日志直接记录 order api timeout",
"evidence_refs": [
{
"source_invocation_id": 22,
"raw_path": "$.logs[0]"
}
]
},
{
"claim_id": "claim-db-primary",
"claim_text": "数据库主库故障导致订单超时",
"claim_type": "root_cause",
"verification": "unsupported",
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
"evidence_refs": [
{
"source_invocation_id": 22,
"raw_path": "$.logs[0]"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "日志显示订单接口超时,但数据库方向证据不足。",
"recommended_actions": [
{
"action_text": "补充查询数据库主库状态和错误日志",
"reason": "当前只有订单接口超时日志"
}
],
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。"
},
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 22,
"sessionId": "eval-audit-metadata-low-confid",
"toolName": "query_logs",
"outputPreview": "order api timeout",
"retrievalDetails": {
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.logs[0]",
"text": "order api timeout"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,155 @@
{
"session": {
"sessionId": "eval-prompt-gatekeeper-audit-closure",
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 19000,
"toolCallCount": 1,
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"prompt_audit": {
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_planner",
"version": "chat-planner-v1",
"resource": "prompts/chat-planner-prompt.md"
},
{
"name": "chat_executor",
"version": "chat-executor-v2",
"resource": "prompts/chat-executor-prompt.md"
},
{
"name": "chat_verifier",
"version": "chat-verifier-v2",
"resource": "prompts/chat-verifier-prompt.md"
},
{
"name": "chat_composer",
"version": "chat-composer-v1",
"resource": "prompts/chat-composer-prompt.md"
}
]
},
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.invocation",
"description": "source_invocation_id must reference an existing tool invocation",
"enabled": true,
"default_severity": "reject"
},
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_metrics",
"source_invocation_id": 21,
"raw_path": "$.alerts[0]",
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "observation",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_id": 21,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
},
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"claim_type": "observation",
"verification": "direct_observation",
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
"evidence_refs": [
{
"source_invocation_id": 21,
"raw_path": "$.alerts[0]"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
"recommended_actions": [],
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
},
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"source_invocation_ids": [21],
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 21,
"sessionId": "eval-prompt-gatekeeper-audit-closure",
"toolName": "query_metrics",
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"retrievalDetails": {
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.alerts[0]",
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
+64 -6
View File
@@ -1,14 +1,14 @@
{
"totalCases" : 10,
"passedCases" : 10,
"totalCases" : 12,
"passedCases" : 12,
"passRate" : 1.0,
"verdictDistribution" : {
"PASS" : 4,
"LOW_CONFID" : 5,
"PASS" : 5,
"LOW_CONFID" : 6,
"REJECT" : 1
},
"averageToolCallCount" : 1.5,
"averageDurationMs" : 39800.0,
"averageToolCallCount" : 1.4166666666666667,
"averageDurationMs" : 38500.0,
"results" : [ {
"caseId" : "narrow-highcpu-observation",
"title" : "Narrow HighCPU observation",
@@ -22,10 +22,31 @@
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"promptAuditVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 1,
"gatekeeperRuleCount" : 1,
"toolCallCount" : 1,
"durationMs" : 18000
}, {
"caseId" : "prompt-gatekeeper-audit-closure",
"title" : "Prompt and Gatekeeper audit closure",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_metrics" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"promptAuditVersion" : "chat-prompts-v1",
"composerStatus" : "valid",
"claimCheckCount" : 1,
"gatekeeperRuleCount" : 2,
"toolCallCount" : 1,
"durationMs" : 19000
}, {
"caseId" : "hikari-no-evidence-negative-observation",
"title" : "Hikari no-evidence negative observation",
@@ -39,8 +60,10 @@
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"promptAuditVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 1,
"gatekeeperRuleCount" : 1,
"toolCallCount" : 1,
"durationMs" : 21000
}, {
@@ -58,8 +81,10 @@
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 3,
"durationMs" : 42000
}, {
@@ -76,8 +101,10 @@
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 2,
"durationMs" : 51000
}, {
@@ -93,8 +120,10 @@
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 1,
"durationMs" : 36000
}, {
@@ -111,8 +140,10 @@
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 2,
"durationMs" : 47000
}, {
@@ -129,8 +160,10 @@
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"gatekeeperRuleCount" : null,
"toolCallCount" : 2,
"durationMs" : 53000
}, {
@@ -146,8 +179,10 @@
},
"gatekeeperStatus" : "fail",
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 1,
"gatekeeperRuleCount" : null,
"toolCallCount" : 1,
"durationMs" : 39000
}, {
@@ -163,10 +198,31 @@
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 2,
"gatekeeperRuleCount" : null,
"toolCallCount" : 1,
"durationMs" : 44000
}, {
"caseId" : "audit-metadata-low-confid",
"title" : "Audit metadata low confidence",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_logs" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"promptAuditVersion" : "chat-prompts-v1",
"composerStatus" : "valid",
"claimCheckCount" : 2,
"gatekeeperRuleCount" : 2,
"toolCallCount" : 1,
"durationMs" : 45000
}, {
"caseId" : "composer-fallback-no-raw-json",
"title" : "Composer fallback no raw JSON",
@@ -180,8 +236,10 @@
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : null,
"promptAuditVersion" : null,
"composerStatus" : "composer_malformed",
"claimCheckCount" : 2,
"gatekeeperRuleCount" : null,
"toolCallCount" : 1,
"durationMs" : 47000
} ]
+20 -18
View File
@@ -1,28 +1,30 @@
# Diagnosis Eval Report
- Total cases: 10
- Passed cases: 10
- Total cases: 12
- Passed cases: 12
- Pass rate: 100.00%
- Average tool calls: 1.50
- Average duration ms: 39800.00
- Average tool calls: 1.42
- Average duration ms: 38500.00
## Verdict Distribution
- PASS: 4
- LOW_CONFID: 5
- PASS: 5
- LOW_CONFID: 6
- REJECT: 1
## Cases
| Case | Result | Verdict | Gatekeeper | Rule Set | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
| --- | --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
| narrow-highcpu-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 18000 | - |
| hikari-no-evidence-negative-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 21000 | - |
| payment-timeout | PASS | PASS | - | - | - | - | 3/3 | 3 | 42000 | - |
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 51000 | - |
| redis-timeout | PASS | LOW_CONFID | - | - | - | - | 2/2 | 1 | 36000 | - |
| slow-response | PASS | PASS | - | - | - | - | 2/2 | 2 | 47000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 53000 | - |
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | - | valid | 1 | 3/3 | 1 | 39000 | - |
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | - | valid | 2 | 2/2 | 1 | 44000 | - |
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | - | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
| Case | Result | Verdict | Gatekeeper | Rule Set | Prompt Audit | Composer | Claim Checks | Rules | Keywords | Tool Calls | Duration ms | Failed Checks |
| --- | --- | --- | --- | --- | --- | --- | ---: | ---: | --- | ---: | ---: | --- |
| narrow-highcpu-observation | PASS | PASS | pass | gatekeeper-rules-v1 | - | valid | 1 | 1 | 3/3 | 1 | 18000 | - |
| prompt-gatekeeper-audit-closure | PASS | PASS | pass | gatekeeper-rules-v1 | chat-prompts-v1 | valid | 1 | 2 | 3/3 | 1 | 19000 | - |
| hikari-no-evidence-negative-observation | PASS | PASS | pass | gatekeeper-rules-v1 | - | valid | 1 | 1 | 3/3 | 1 | 21000 | - |
| payment-timeout | PASS | PASS | - | - | - | - | - | - | 3/3 | 3 | 42000 | - |
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | - | - | - | 3/3 | 2 | 51000 | - |
| redis-timeout | PASS | LOW_CONFID | - | - | - | - | - | - | 2/2 | 1 | 36000 | - |
| slow-response | PASS | PASS | - | - | - | - | - | - | 2/2 | 2 | 47000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | - | - | - | 3/3 | 2 | 53000 | - |
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | - | - | valid | 1 | - | 3/3 | 1 | 39000 | - |
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | - | - | valid | 2 | - | 2/2 | 1 | 44000 | - |
| audit-metadata-low-confid | PASS | LOW_CONFID | pass | gatekeeper-rules-v1 | chat-prompts-v1 | valid | 2 | 2 | 2/2 | 1 | 45000 | - |
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | - | - | composer_malformed | 2 | - | 2/2 | 1 | 47000 | - |
+23 -1
View File
@@ -34,7 +34,16 @@ baseline report:整套固定集当前认可的结果
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["主库故障"]
"forbiddenConfirmedClaimKeywords": ["主库故障"],
"requirePromptAudit": true,
"expectedPromptAuditVersion": "chat-prompts-v1",
"expectedPromptVersions": {
"chat_planner": "chat-planner-v1",
"chat_executor": "chat-executor-v2",
"chat_verifier": "chat-verifier-v2",
"chat_composer": "chat-composer-v1"
},
"requireGatekeeperRules": true
}
```
@@ -58,6 +67,10 @@ baseline report:整套固定集当前认可的结果
| `expectedGatekeeperRuleSetVersion` | 期望的 Gatekeeper 规则集版本 | 配置后校验 `gatekeeper_result.rule_set_version` |
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
| `requirePromptAudit` | 是否要求 Prompt 审计元数据 | 要求 `prompt_audit.version` 存在 |
| `expectedPromptAuditVersion` | 期望的 Prompt 审计目录版本 | 配置后校验 `prompt_audit.version` |
| `expectedPromptVersions` | 期望的各角色 Prompt 版本 | 校验 `prompt_audit.prompts[*].name/version` |
| `requireGatekeeperRules` | 是否要求 Gatekeeper 规则元数据 | 要求 `gatekeeper_result.rules` 非空,且每条规则有 `id`、`enabled`、`default_severity` |
## 2. Trace Fixture
@@ -72,6 +85,9 @@ fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | Gatekeeper 规则集版本 | 新矩阵 case 可显式断言该版本 |
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rules` | Gatekeeper 规则元数据摘要 | 审计 case 可要求规则列表非空且字段完整 |
| `session.selfEvaluation.verifier_evaluation.prompt_audit.version` | Chat Prompt 审计目录版本 | 审计 case 可显式断言该版本 |
| `session.selfEvaluation.verifier_evaluation.prompt_audit.prompts` | 各 Chat Prompt 名称、版本和资源路径 | 审计 case 可断言 planner、executor、verifier、composer 版本 |
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
@@ -95,8 +111,10 @@ Java 类型:`DiagnosisEvalResult`
| `evidenceCoverage` | 每个必需工具是否出现 |
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
| `gatekeeperRuleSetVersion` | 读到的 `gatekeeper_result.rule_set_version` |
| `promptAuditVersion` | 读到的 `prompt_audit.version` |
| `composerStatus` | 读到的 `composer_output.status` |
| `claimCheckCount` | `claim_checks` 数量 |
| `gatekeeperRuleCount` | `gatekeeper_result.rules` 数量 |
| `toolCallCount` | trace 中工具调用总数 |
| `durationMs` | trace 总耗时 |
@@ -130,6 +148,10 @@ Executor structured output
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
- 配置 `expectedGatekeeperRuleSetVersion` 的 case 必须匹配 `gatekeeper_result.rule_set_version`。
- 配置 `requirePromptAudit` 的 case 必须包含 `prompt_audit.version`。
- 配置 `expectedPromptAuditVersion` 的 case 必须匹配 `prompt_audit.version`。
- 配置 `expectedPromptVersions` 的 case 必须能在 `prompt_audit.prompts` 中找到对应角色和版本。
- 配置 `requireGatekeeperRules` 的 case 必须包含非空 `gatekeeper_result.rules`,且每条规则有 `id`、`enabled`、`default_severity`。
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
- Composer 输出必须记录 `status`。
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。