feat(eval): add evidence pipeline acceptance closure

This commit is contained in:
aruo
2026-07-09 00:47:48 +08:00
parent a77c947cd4
commit db0f229285
46 changed files with 1434 additions and 53 deletions
+2 -1
View File
@@ -397,7 +397,8 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
- Gatekeeper 对 Executor 引用做代码级验真,Verifier 只判断可推导性。
- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
- Verifier 只判断可推导性。
- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
- RAG offline baseline 和 live acceptance 脚本。
+1 -1
View File
@@ -267,7 +267,7 @@ category
|---|---|
| `executor_output_parse_status` | Executor 输出是否能解析为 `executor_evidence_v2` |
| `executor_structured_output` | Executor 结构化 claims、hypotheses、recommended_actions、missing_info |
| `gatekeeper_result` | 引用真实性校验结果,包括 checked bindings、failed rules、warnings、errors |
| `gatekeeper_result` | 引用真实性校验结果,包括 rule set version、checked bindings、failed rules、warnings、errors |
| `composer_output` | Composer 最终表达及解析状态 |
| `tool_trace_summary` | Verifier 调用时使用的工具调用导航索引,不是唯一证据源 |
@@ -233,6 +233,15 @@ Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码
{
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
@@ -255,6 +264,8 @@ Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码
|---|---|---|
| `status` | string | `pass` 或 `fail` |
| `severity` | string | `none`、`low_confid`、`reject` |
| `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 |
| `rules` | array | 已启用规则的轻量元数据摘要 |
| `checked_bindings` | array | 每条证据绑定的校验结果 |
| `failed_rules` | array | 失败规则 id |
| `warnings` | array | 自动回填等非阻断信息 |
@@ -271,6 +282,12 @@ Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
- `negative_observation` 只能绑定 `$.no_evidence`。
规则配置:
- 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。
- 规则实现仍是确定性 Java 代码,不执行动态脚本。
- 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。
失败分级:
| 场景 | severity |
@@ -389,7 +406,9 @@ Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达
"traceability_version": "v1",
"executor_output_parse_status": {},
"executor_structured_output": {},
"gatekeeper_result": {},
"gatekeeper_result": {
"rule_set_version": "gatekeeper-rules-v1"
},
"composer_output": {},
"tool_trace_summary": []
}
@@ -420,6 +439,5 @@ Trace API 可用于回放:
当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
2. Gatekeeper 规则配置化文档:如果后续把规则做成索引层、元数据层、规则层,需要单独记录加载顺序和审计字段。
3. E2E fixture 矩阵:把 ISS-008/ISS-009 的用例固化到诊断评测集,而不是只存在 issue 验证记录。
4. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。
3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
+3 -1
View File
@@ -173,6 +173,8 @@ Gatekeeper 检查:
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
Verifier 输出:
```json
@@ -230,7 +232,7 @@ diagnosis_session.self_evaluation.aiops_rule_evaluation
- 工具参数 schema 校验。
- 同一工具调用次数上限。
- 工具超时的统一熔断。
- Gatekeeper 规则三层分离:索引层、元数据层、规则实现层。
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
- Prompt 版本记录和回滚。
- Verifier 对 AIOps 报告的 LLM 级事实校验。
+14
View File
@@ -6,9 +6,13 @@
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
- `interview-walkthrough.md`:面试讲解话术。
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
- `trace-inspection-checklist.md`:Trace 字段检查清单。
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
## 1. 前置条件
@@ -167,3 +171,13 @@ AIOps 主线:
-> AIOps rule evaluation
-> Trace API 回放
```
## 8. Evidence Pipeline 场景矩阵
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
- `scripts/run-payment-timeout-demo.ps1` 跑主路径。
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 10/10 通过。
这样可以同时展示真实链路和确定性回归能力。
+63
View File
@@ -0,0 +1,63 @@
# Evidence Pipeline Demo Scenarios
这份清单用于面试时说明 Chat 证据链路如何覆盖 `PASS`、`LOW_CONFID`、`REJECT` 和 no-evidence 场景。
重点区别:
- Live demo 证明本地服务、工具、Trace、Feedback 主链路能跑通。
- Fixture-backed eval 证明固定安全场景可以确定性回归,不依赖 LLM 当场随机输出。
## Scenario Matrix
| 场景 | 类型 | 输入/证据 | 期望讲点 |
|---|---|---|---|
| Payment timeout | Live 主路径 | `requests/payment-timeout-chat.json` | 完整 Chat -> Trace -> Feedback 闭环 |
| Narrow HighCPU observation | Live 可尝试 + fixture-backed | `requests/narrow-highcpu-chat.json` / `mvp/eval/fixtures/narrow-highcpu-observation-pass.json` | Executor 只输出观察类 claim,Gatekeeper 验引用,Verifier PASS |
| Hikari no-evidence | Live 可尝试 + fixture-backed | `requests/hikari-no-evidence-chat.json` / `mvp/eval/fixtures/hikari-no-evidence-negative-observation-pass.json` | `$.no_evidence` 只表示本次查询无匹配证据,Composer 不说“已排除” |
| Unsupported claim filtering | Fixture-backed | `requests/safety-unsupported-claim-chat.json` / `mvp/eval/fixtures/unsupported-claim-filtering-low-confid.json` | Verifier 将 unsupported claim 降为 LOW_CONFID,最终答案不确认“主库故障” |
| Fabricated invocation reject | Fixture-backed | `mvp/eval/fixtures/gatekeeper-fabricated-invocation-reject.json` | Gatekeeper 拦截伪造 invocation,最终 REJECT/降级 |
| Composer fallback | Fixture-backed | `mvp/eval/fixtures/composer-fallback-no-raw-json-low-confid.json` | 即使 Composer 输出异常,也不能把 Executor JSON 泄漏给用户 |
## Trace Fields To Inspect
| 能力 | JSON path |
|---|---|
| Executor V2 输出 | `data.session.selfEvaluation.verifier_evaluation.executor_structured_output` |
| Gatekeeper 结果 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` |
| Gatekeeper 规则版本 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` |
| 证据绑定校验 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.checked_bindings` |
| Verifier claim checks | `data.session.selfEvaluation.verifier_evaluation.claim_checks` |
| Composer 输出 | `data.session.selfEvaluation.verifier_evaluation.composer_output` |
| 工具证据引用 | `data.toolInvocations[*].retrievalDetails.evidence_refs` |
## How To Present It
```text
我把现场 demo 和固定 eval 分开。
现场 demo 证明系统能跑通真实链路;
fixture-backed eval 证明反幻觉安全场景可以稳定回归。
Gatekeeper 的规则版本也进入 trace,所以后续调整阈值或规则时可以审计。
```
## Optional Live Requests
手动发送某个请求样例:
```powershell
$body = Get-Content -Raw -Encoding UTF8 "mvp/demo/requests/narrow-highcpu-chat.json"
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
然后查询同一 session:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace"
```
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
+9 -1
View File
@@ -24,6 +24,7 @@
4. 打开 `mvp/demo/output/trace-response.json`。
5. 指出证据工具和 verifier evaluation。
6. 提交 feedback,并展示它挂在同一个 session 上。
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
## 3. 命令
@@ -125,6 +126,14 @@ Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
```
如果被问到怎么防止证据归因幻觉,可以补充:
```text
Executor 的 claim 必须绑定 source_invocation_id、raw_path 和 evidence_excerpt。
Gatekeeper 用代码核验这些引用,并把 rule_set_version 写进 trace。
Verifier 只判断已核验证据能否推出 claim,Composer 只表达允许输出的内容。
```
## 5. 强面试表达
```text
@@ -141,4 +150,3 @@ traceability、evidence persistence、verifier gating、feedback 和 regression
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
```
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-hikari-no-evidence-001",
"Question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。"
}
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-narrow-highcpu-001",
"Question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。"
}
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-safety-unsupported-001",
"Question": "订单超时是否可以确认由数据库主库故障导致?请只基于当前证据回答。"
}
+2
View File
@@ -10,6 +10,7 @@
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
## 2. Agent 步骤
@@ -30,6 +31,7 @@
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
## 4. Summary
+8 -4
View File
@@ -29,18 +29,21 @@ The baseline evaluates saved trace fixtures. It does not start the application a
The committed baseline currently contains:
```text
8 fixed cases
8 passing fixture evaluations
2 PASS verdicts
10 fixed cases
10 passing fixture evaluations
4 PASS verdicts
5 LOW_CONFID verdicts
1 REJECT verdict
```
The three V2 audit-closure cases cover:
The V2 evidence-pipeline matrix covers:
- Positive supported evidence for a narrow HighCPU observation.
- No-evidence `negative_observation` using `$.no_evidence`.
- Gatekeeper failure for a fabricated tool invocation reference.
- Unsupported claim filtering before the final answer.
- Composer fallback rendering without raw Executor JSON leakage.
- Gatekeeper rule set version audit for new matrix fixtures.
## Verification
@@ -76,6 +79,7 @@ Stage 5 adds these V2 checks:
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
- `claim_checks` must be structurally auditable.
- Composer output must record whether normal parsing or fallback rendering was used.
- Gatekeeper rule set version can be asserted per fixture.
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
+36
View File
@@ -1,4 +1,40 @@
[
{
"id": "narrow-highcpu-observation",
"title": "Narrow HighCPU observation",
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
"traceFixture": "narrow-highcpu-observation-pass.json",
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics"],
"allowedVerdicts": ["PASS"],
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["数据库连接池"]
},
{
"id": "hikari-no-evidence-negative-observation",
"title": "Hikari no-evidence negative observation",
"question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
"traceFixture": "hikari-no-evidence-negative-observation-pass.json",
"expectedRootCauseKeywords": ["未检索到", "HikariCP", "匹配证据"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["PASS"],
"forbiddenAnswerKeywords": ["已排除", "确认没有", "日志层面已排除"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["已排除 HikariCP"]
},
{
"id": "payment-timeout",
"title": "Payment API timeout",
@@ -0,0 +1,126 @@
{
"session": {
"sessionId": "eval-hikari-no-evidence-negative-observation",
"query": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 21000,
"toolCallCount": 1,
"answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_logs",
"source_invocation_id": 12,
"raw_path": "$.no_evidence",
"matched_text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "negative_observation",
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_logs",
"source_invocation_id": 12,
"raw_path": "$.no_evidence",
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": [
"仅查询了 application-logs 中 inventory-service HikariCP 相关日志"
]
},
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"claim_type": "negative_observation",
"verification": "direct_observation",
"detail": "$.no_evidence 只支持本次查询未检索到匹配证据。",
"evidence_refs": [
{
"source_invocation_id": 12,
"raw_path": "$.no_evidence"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "本次查询未检索到匹配日志。",
"recommended_actions": [],
"user_facing_answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。"
},
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"source_invocation_ids": [12],
"evidence_level": "no_evidence"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 12,
"sessionId": "eval-hikari-no-evidence-negative-observation",
"toolName": "query_logs",
"outputPreview": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
"retrievalDetails": {
"evidence_status": "no_evidence",
"evidence_refs": [
{
"raw_path": "$.no_evidence",
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,124 @@
{
"session": {
"sessionId": "eval-narrow-highcpu-observation",
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 18000,
"toolCallCount": 1,
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_metrics",
"source_invocation_id": 11,
"raw_path": "$.alerts[0]",
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "observation",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_id": 11,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
},
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"claim_type": "observation",
"verification": "direct_observation",
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
"evidence_refs": [
{
"source_invocation_id": 11,
"raw_path": "$.alerts[0]"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
"recommended_actions": [],
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
},
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"source_invocation_ids": [11],
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 11,
"sessionId": "eval-narrow-highcpu-observation",
"toolName": "query_metrics",
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"retrievalDetails": {
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.alerts[0]",
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
+48 -6
View File
@@ -1,15 +1,49 @@
{
"totalCases" : 8,
"passedCases" : 8,
"totalCases" : 10,
"passedCases" : 10,
"passRate" : 1.0,
"verdictDistribution" : {
"PASS" : 2,
"PASS" : 4,
"LOW_CONFID" : 5,
"REJECT" : 1
},
"averageToolCallCount" : 1.625,
"averageDurationMs" : 44875.0,
"averageToolCallCount" : 1.5,
"averageDurationMs" : 39800.0,
"results" : [ {
"caseId" : "narrow-highcpu-observation",
"title" : "Narrow HighCPU observation",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_metrics" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"composerStatus" : "valid",
"claimCheckCount" : 1,
"toolCallCount" : 1,
"durationMs" : 18000
}, {
"caseId" : "hikari-no-evidence-negative-observation",
"title" : "Hikari no-evidence negative observation",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_logs" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"composerStatus" : "valid",
"claimCheckCount" : 1,
"toolCallCount" : 1,
"durationMs" : 21000
}, {
"caseId" : "payment-timeout",
"title" : "Payment API timeout",
"passed" : true,
@@ -23,6 +57,7 @@
"query_metrics" : true
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"toolCallCount" : 3,
@@ -40,6 +75,7 @@
"query_logs" : true
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"toolCallCount" : 2,
@@ -56,6 +92,7 @@
"query_logs" : true
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"toolCallCount" : 1,
@@ -73,6 +110,7 @@
"query_logs" : true
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"toolCallCount" : 2,
@@ -90,6 +128,7 @@
"query_logs" : true
},
"gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null,
"claimCheckCount" : null,
"toolCallCount" : 2,
@@ -106,6 +145,7 @@
"query_logs" : true
},
"gatekeeperStatus" : "fail",
"gatekeeperRuleSetVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 1,
"toolCallCount" : 1,
@@ -122,6 +162,7 @@
"query_logs" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : null,
"composerStatus" : "valid",
"claimCheckCount" : 2,
"toolCallCount" : 1,
@@ -138,9 +179,10 @@
"query_metrics" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : null,
"composerStatus" : "composer_malformed",
"claimCheckCount" : 2,
"toolCallCount" : 1,
"durationMs" : 47000
} ]
}
}
+17 -15
View File
@@ -1,26 +1,28 @@
# Diagnosis Eval Report
- Total cases: 8
- Passed cases: 8
- Total cases: 10
- Passed cases: 10
- Pass rate: 100.00%
- Average tool calls: 1.63
- Average duration ms: 44875.00
- Average tool calls: 1.50
- Average duration ms: 39800.00
## Verdict Distribution
- PASS: 2
- PASS: 4
- LOW_CONFID: 5
- REJECT: 1
## Cases
| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
| payment-timeout | PASS | PASS | - | - | - | 3/3 | 3 | 42000 | - |
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 51000 | - |
| redis-timeout | PASS | LOW_CONFID | - | - | - | 2/2 | 1 | 36000 | - |
| slow-response | PASS | PASS | - | - | - | 2/2 | 2 | 47000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 53000 | - |
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | valid | 1 | 3/3 | 1 | 39000 | - |
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | valid | 2 | 2/2 | 1 | 44000 | - |
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
| Case | Result | Verdict | Gatekeeper | Rule Set | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
| --- | --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
| narrow-highcpu-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 18000 | - |
| hikari-no-evidence-negative-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 21000 | - |
| payment-timeout | PASS | PASS | - | - | - | - | 3/3 | 3 | 42000 | - |
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 51000 | - |
| redis-timeout | PASS | LOW_CONFID | - | - | - | - | 2/2 | 1 | 36000 | - |
| slow-response | PASS | PASS | - | - | - | - | 2/2 | 2 | 47000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 53000 | - |
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | - | valid | 1 | 3/3 | 1 | 39000 | - |
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | - | valid | 2 | 2/2 | 1 | 44000 | - |
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | - | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
+5
View File
@@ -32,6 +32,7 @@ baseline report:整套固定集当前认可的结果
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["主库故障"]
}
@@ -54,6 +55,7 @@ baseline report:整套固定集当前认可的结果
| `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 |
| `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` |
| `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 |
| `expectedGatekeeperRuleSetVersion` | 期望的 Gatekeeper 规则集版本 | 配置后校验 `gatekeeper_result.rule_set_version` |
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
@@ -69,6 +71,7 @@ fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规
| `session.totalDurationMs` | 运行耗时 | 进入报告 |
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | Gatekeeper 规则集版本 | 新矩阵 case 可显式断言该版本 |
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
@@ -91,6 +94,7 @@ Java 类型:`DiagnosisEvalResult`
| `requiredKeywordCount` | case 配置的关键词数量 |
| `evidenceCoverage` | 每个必需工具是否出现 |
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
| `gatekeeperRuleSetVersion` | 读到的 `gatekeeper_result.rule_set_version` |
| `composerStatus` | 读到的 `composer_output.status` |
| `claimCheckCount` | `claim_checks` 数量 |
| `toolCallCount` | trace 中工具调用总数 |
@@ -125,6 +129,7 @@ Executor structured output
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
- 配置 `expectedGatekeeperRuleSetVersion` 的 case 必须匹配 `gatekeeper_result.rule_set_version`。
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
- Composer 输出必须记录 `status`。
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。