feat(eval): add evidence pipeline acceptance closure
This commit is contained in:
@@ -397,7 +397,8 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
|
||||
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
|
||||
- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
|
||||
- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
|
||||
- Gatekeeper 对 Executor 引用做代码级验真,Verifier 只判断可推导性。
|
||||
- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
|
||||
- Verifier 只判断可推导性。
|
||||
- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
|
||||
- RAG offline baseline 和 live acceptance 脚本。
|
||||
|
||||
|
||||
@@ -267,7 +267,7 @@ category
|
||||
|---|---|
|
||||
| `executor_output_parse_status` | Executor 输出是否能解析为 `executor_evidence_v2` |
|
||||
| `executor_structured_output` | Executor 结构化 claims、hypotheses、recommended_actions、missing_info |
|
||||
| `gatekeeper_result` | 引用真实性校验结果,包括 checked bindings、failed rules、warnings、errors |
|
||||
| `gatekeeper_result` | 引用真实性校验结果,包括 rule set version、checked bindings、failed rules、warnings、errors |
|
||||
| `composer_output` | Composer 最终表达及解析状态 |
|
||||
| `tool_trace_summary` | Verifier 调用时使用的工具调用导航索引,不是唯一证据源 |
|
||||
|
||||
|
||||
@@ -233,6 +233,15 @@ Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码
|
||||
{
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
@@ -255,6 +264,8 @@ Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码
|
||||
|---|---|---|
|
||||
| `status` | string | `pass` 或 `fail` |
|
||||
| `severity` | string | `none`、`low_confid`、`reject` |
|
||||
| `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 |
|
||||
| `rules` | array | 已启用规则的轻量元数据摘要 |
|
||||
| `checked_bindings` | array | 每条证据绑定的校验结果 |
|
||||
| `failed_rules` | array | 失败规则 id |
|
||||
| `warnings` | array | 自动回填等非阻断信息 |
|
||||
@@ -271,6 +282,12 @@ Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码
|
||||
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
|
||||
- `negative_observation` 只能绑定 `$.no_evidence`。
|
||||
|
||||
规则配置:
|
||||
|
||||
- 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。
|
||||
- 规则实现仍是确定性 Java 代码,不执行动态脚本。
|
||||
- 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。
|
||||
|
||||
失败分级:
|
||||
|
||||
| 场景 | severity |
|
||||
@@ -389,7 +406,9 @@ Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达
|
||||
"traceability_version": "v1",
|
||||
"executor_output_parse_status": {},
|
||||
"executor_structured_output": {},
|
||||
"gatekeeper_result": {},
|
||||
"gatekeeper_result": {
|
||||
"rule_set_version": "gatekeeper-rules-v1"
|
||||
},
|
||||
"composer_output": {},
|
||||
"tool_trace_summary": []
|
||||
}
|
||||
@@ -420,6 +439,5 @@ Trace API 可用于回放:
|
||||
当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
|
||||
|
||||
1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
|
||||
2. Gatekeeper 规则配置化文档:如果后续把规则做成索引层、元数据层、规则层,需要单独记录加载顺序和审计字段。
|
||||
3. E2E fixture 矩阵:把 ISS-008/ISS-009 的用例固化到诊断评测集,而不是只存在 issue 验证记录。
|
||||
4. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
|
||||
2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。
|
||||
3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
|
||||
|
||||
@@ -173,6 +173,8 @@ Gatekeeper 检查:
|
||||
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
|
||||
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
|
||||
|
||||
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
|
||||
|
||||
Verifier 输出:
|
||||
|
||||
```json
|
||||
@@ -230,7 +232,7 @@ diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||
- 工具参数 schema 校验。
|
||||
- 同一工具调用次数上限。
|
||||
- 工具超时的统一熔断。
|
||||
- Gatekeeper 规则三层分离:索引层、元数据层、规则实现层。
|
||||
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
||||
- Prompt 版本记录和回滚。
|
||||
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||||
|
||||
|
||||
@@ -6,9 +6,13 @@
|
||||
|
||||
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||
- `interview-walkthrough.md`:面试讲解话术。
|
||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
|
||||
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
|
||||
- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
|
||||
|
||||
## 1. 前置条件
|
||||
|
||||
@@ -167,3 +171,13 @@ AIOps 主线:
|
||||
-> AIOps rule evaluation
|
||||
-> Trace API 回放
|
||||
```
|
||||
|
||||
## 8. Evidence Pipeline 场景矩阵
|
||||
|
||||
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
|
||||
|
||||
- `scripts/run-payment-timeout-demo.ps1` 跑主路径。
|
||||
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
|
||||
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 10/10 通过。
|
||||
|
||||
这样可以同时展示真实链路和确定性回归能力。
|
||||
|
||||
@@ -0,0 +1,63 @@
|
||||
# Evidence Pipeline Demo Scenarios
|
||||
|
||||
这份清单用于面试时说明 Chat 证据链路如何覆盖 `PASS`、`LOW_CONFID`、`REJECT` 和 no-evidence 场景。
|
||||
|
||||
重点区别:
|
||||
|
||||
- Live demo 证明本地服务、工具、Trace、Feedback 主链路能跑通。
|
||||
- Fixture-backed eval 证明固定安全场景可以确定性回归,不依赖 LLM 当场随机输出。
|
||||
|
||||
## Scenario Matrix
|
||||
|
||||
| 场景 | 类型 | 输入/证据 | 期望讲点 |
|
||||
|---|---|---|---|
|
||||
| Payment timeout | Live 主路径 | `requests/payment-timeout-chat.json` | 完整 Chat -> Trace -> Feedback 闭环 |
|
||||
| Narrow HighCPU observation | Live 可尝试 + fixture-backed | `requests/narrow-highcpu-chat.json` / `mvp/eval/fixtures/narrow-highcpu-observation-pass.json` | Executor 只输出观察类 claim,Gatekeeper 验引用,Verifier PASS |
|
||||
| Hikari no-evidence | Live 可尝试 + fixture-backed | `requests/hikari-no-evidence-chat.json` / `mvp/eval/fixtures/hikari-no-evidence-negative-observation-pass.json` | `$.no_evidence` 只表示本次查询无匹配证据,Composer 不说“已排除” |
|
||||
| Unsupported claim filtering | Fixture-backed | `requests/safety-unsupported-claim-chat.json` / `mvp/eval/fixtures/unsupported-claim-filtering-low-confid.json` | Verifier 将 unsupported claim 降为 LOW_CONFID,最终答案不确认“主库故障” |
|
||||
| Fabricated invocation reject | Fixture-backed | `mvp/eval/fixtures/gatekeeper-fabricated-invocation-reject.json` | Gatekeeper 拦截伪造 invocation,最终 REJECT/降级 |
|
||||
| Composer fallback | Fixture-backed | `mvp/eval/fixtures/composer-fallback-no-raw-json-low-confid.json` | 即使 Composer 输出异常,也不能把 Executor JSON 泄漏给用户 |
|
||||
|
||||
## Trace Fields To Inspect
|
||||
|
||||
| 能力 | JSON path |
|
||||
|---|---|
|
||||
| Executor V2 输出 | `data.session.selfEvaluation.verifier_evaluation.executor_structured_output` |
|
||||
| Gatekeeper 结果 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` |
|
||||
| Gatekeeper 规则版本 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` |
|
||||
| 证据绑定校验 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.checked_bindings` |
|
||||
| Verifier claim checks | `data.session.selfEvaluation.verifier_evaluation.claim_checks` |
|
||||
| Composer 输出 | `data.session.selfEvaluation.verifier_evaluation.composer_output` |
|
||||
| 工具证据引用 | `data.toolInvocations[*].retrievalDetails.evidence_refs` |
|
||||
|
||||
## How To Present It
|
||||
|
||||
```text
|
||||
我把现场 demo 和固定 eval 分开。
|
||||
现场 demo 证明系统能跑通真实链路;
|
||||
fixture-backed eval 证明反幻觉安全场景可以稳定回归。
|
||||
Gatekeeper 的规则版本也进入 trace,所以后续调整阈值或规则时可以审计。
|
||||
```
|
||||
|
||||
## Optional Live Requests
|
||||
|
||||
手动发送某个请求样例:
|
||||
|
||||
```powershell
|
||||
$body = Get-Content -Raw -Encoding UTF8 "mvp/demo/requests/narrow-highcpu-chat.json"
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
然后查询同一 session:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace"
|
||||
```
|
||||
|
||||
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
|
||||
@@ -24,6 +24,7 @@
|
||||
4. 打开 `mvp/demo/output/trace-response.json`。
|
||||
5. 指出证据工具和 verifier evaluation。
|
||||
6. 提交 feedback,并展示它挂在同一个 session 上。
|
||||
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
|
||||
|
||||
## 3. 命令
|
||||
|
||||
@@ -125,6 +126,14 @@ Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可
|
||||
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
|
||||
```
|
||||
|
||||
如果被问到怎么防止证据归因幻觉,可以补充:
|
||||
|
||||
```text
|
||||
Executor 的 claim 必须绑定 source_invocation_id、raw_path 和 evidence_excerpt。
|
||||
Gatekeeper 用代码核验这些引用,并把 rule_set_version 写进 trace。
|
||||
Verifier 只判断已核验证据能否推出 claim,Composer 只表达允许输出的内容。
|
||||
```
|
||||
|
||||
## 5. 强面试表达
|
||||
|
||||
```text
|
||||
@@ -141,4 +150,3 @@ traceability、evidence persistence、verifier gating、feedback 和 regression
|
||||
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
|
||||
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
|
||||
```
|
||||
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-hikari-no-evidence-001",
|
||||
"Question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。"
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-narrow-highcpu-001",
|
||||
"Question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。"
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-safety-unsupported-001",
|
||||
"Question": "订单超时是否可以确认由数据库主库故障导致?请只基于当前证据回答。"
|
||||
}
|
||||
@@ -10,6 +10,7 @@
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
@@ -30,6 +31,7 @@
|
||||
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
|
||||
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
|
||||
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
|
||||
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||
|
||||
## 4. Summary
|
||||
|
||||
+8
-4
@@ -29,18 +29,21 @@ The baseline evaluates saved trace fixtures. It does not start the application a
|
||||
The committed baseline currently contains:
|
||||
|
||||
```text
|
||||
8 fixed cases
|
||||
8 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
10 fixed cases
|
||||
10 passing fixture evaluations
|
||||
4 PASS verdicts
|
||||
5 LOW_CONFID verdicts
|
||||
1 REJECT verdict
|
||||
```
|
||||
|
||||
The three V2 audit-closure cases cover:
|
||||
The V2 evidence-pipeline matrix covers:
|
||||
|
||||
- Positive supported evidence for a narrow HighCPU observation.
|
||||
- No-evidence `negative_observation` using `$.no_evidence`.
|
||||
- Gatekeeper failure for a fabricated tool invocation reference.
|
||||
- Unsupported claim filtering before the final answer.
|
||||
- Composer fallback rendering without raw Executor JSON leakage.
|
||||
- Gatekeeper rule set version audit for new matrix fixtures.
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -76,6 +79,7 @@ Stage 5 adds these V2 checks:
|
||||
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
|
||||
- `claim_checks` must be structurally auditable.
|
||||
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||
- Gatekeeper rule set version can be asserted per fixture.
|
||||
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||
|
||||
|
||||
@@ -1,4 +1,40 @@
|
||||
[
|
||||
{
|
||||
"id": "narrow-highcpu-observation",
|
||||
"title": "Narrow HighCPU observation",
|
||||
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
|
||||
"traceFixture": "narrow-highcpu-observation-pass.json",
|
||||
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics"],
|
||||
"allowedVerdicts": ["PASS"],
|
||||
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["数据库连接池"]
|
||||
},
|
||||
{
|
||||
"id": "hikari-no-evidence-negative-observation",
|
||||
"title": "Hikari no-evidence negative observation",
|
||||
"question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
|
||||
"traceFixture": "hikari-no-evidence-negative-observation-pass.json",
|
||||
"expectedRootCauseKeywords": ["未检索到", "HikariCP", "匹配证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["PASS"],
|
||||
"forbiddenAnswerKeywords": ["已排除", "确认没有", "日志层面已排除"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["已排除 HikariCP"]
|
||||
},
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
|
||||
@@ -0,0 +1,126 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-hikari-no-evidence-negative-observation",
|
||||
"query": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 21000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 12,
|
||||
"raw_path": "$.no_evidence",
|
||||
"matched_text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "negative_observation",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 12,
|
||||
"raw_path": "$.no_evidence",
|
||||
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": [
|
||||
"仅查询了 application-logs 中 inventory-service HikariCP 相关日志"
|
||||
]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"claim_type": "negative_observation",
|
||||
"verification": "direct_observation",
|
||||
"detail": "$.no_evidence 只支持本次查询未检索到匹配证据。",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 12,
|
||||
"raw_path": "$.no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "本次查询未检索到匹配日志。",
|
||||
"recommended_actions": [],
|
||||
"user_facing_answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"source_invocation_ids": [12],
|
||||
"evidence_level": "no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 12,
|
||||
"sessionId": "eval-hikari-no-evidence-negative-observation",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "no_evidence",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.no_evidence",
|
||||
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,124 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-narrow-highcpu-observation",
|
||||
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 18000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 11,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "observation",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 11,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"claim_type": "observation",
|
||||
"verification": "direct_observation",
|
||||
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 11,
|
||||
"raw_path": "$.alerts[0]"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
|
||||
"recommended_actions": [],
|
||||
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"source_invocation_ids": [11],
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 11,
|
||||
"sessionId": "eval-narrow-highcpu-observation",
|
||||
"toolName": "query_metrics",
|
||||
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.alerts[0]",
|
||||
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -1,15 +1,49 @@
|
||||
{
|
||||
"totalCases" : 8,
|
||||
"passedCases" : 8,
|
||||
"totalCases" : 10,
|
||||
"passedCases" : 10,
|
||||
"passRate" : 1.0,
|
||||
"verdictDistribution" : {
|
||||
"PASS" : 2,
|
||||
"PASS" : 4,
|
||||
"LOW_CONFID" : 5,
|
||||
"REJECT" : 1
|
||||
},
|
||||
"averageToolCallCount" : 1.625,
|
||||
"averageDurationMs" : 44875.0,
|
||||
"averageToolCallCount" : 1.5,
|
||||
"averageDurationMs" : 39800.0,
|
||||
"results" : [ {
|
||||
"caseId" : "narrow-highcpu-observation",
|
||||
"title" : "Narrow HighCPU observation",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 1,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 18000
|
||||
}, {
|
||||
"caseId" : "hikari-no-evidence-negative-observation",
|
||||
"title" : "Hikari no-evidence negative observation",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 1,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 21000
|
||||
}, {
|
||||
"caseId" : "payment-timeout",
|
||||
"title" : "Payment API timeout",
|
||||
"passed" : true,
|
||||
@@ -23,6 +57,7 @@
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 3,
|
||||
@@ -40,6 +75,7 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
@@ -56,6 +92,7 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 1,
|
||||
@@ -73,6 +110,7 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
@@ -90,6 +128,7 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
@@ -106,6 +145,7 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "fail",
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 1,
|
||||
"toolCallCount" : 1,
|
||||
@@ -122,6 +162,7 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 2,
|
||||
"toolCallCount" : 1,
|
||||
@@ -138,9 +179,10 @@
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"composerStatus" : "composer_malformed",
|
||||
"claimCheckCount" : 2,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 47000
|
||||
} ]
|
||||
}
|
||||
}
|
||||
@@ -1,26 +1,28 @@
|
||||
# Diagnosis Eval Report
|
||||
|
||||
- Total cases: 8
|
||||
- Passed cases: 8
|
||||
- Total cases: 10
|
||||
- Passed cases: 10
|
||||
- Pass rate: 100.00%
|
||||
- Average tool calls: 1.63
|
||||
- Average duration ms: 44875.00
|
||||
- Average tool calls: 1.50
|
||||
- Average duration ms: 39800.00
|
||||
|
||||
## Verdict Distribution
|
||||
|
||||
- PASS: 2
|
||||
- PASS: 4
|
||||
- LOW_CONFID: 5
|
||||
- REJECT: 1
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | - | - | - | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | - | - | - | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | - | - | - | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 53000 | - |
|
||||
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | valid | 1 | 3/3 | 1 | 39000 | - |
|
||||
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | valid | 2 | 2/2 | 1 | 44000 | - |
|
||||
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
|
||||
| Case | Result | Verdict | Gatekeeper | Rule Set | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
|
||||
| narrow-highcpu-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 18000 | - |
|
||||
| hikari-no-evidence-negative-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 21000 | - |
|
||||
| payment-timeout | PASS | PASS | - | - | - | - | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | - | - | - | - | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | - | - | - | - | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 53000 | - |
|
||||
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | - | valid | 1 | 3/3 | 1 | 39000 | - |
|
||||
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | - | valid | 2 | 2/2 | 1 | 44000 | - |
|
||||
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | - | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
|
||||
|
||||
@@ -32,6 +32,7 @@ baseline report:整套固定集当前认可的结果
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||
}
|
||||
@@ -54,6 +55,7 @@ baseline report:整套固定集当前认可的结果
|
||||
| `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 |
|
||||
| `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` |
|
||||
| `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 |
|
||||
| `expectedGatekeeperRuleSetVersion` | 期望的 Gatekeeper 规则集版本 | 配置后校验 `gatekeeper_result.rule_set_version` |
|
||||
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
|
||||
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
|
||||
|
||||
@@ -69,6 +71,7 @@ fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规
|
||||
| `session.totalDurationMs` | 运行耗时 | 进入报告 |
|
||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
|
||||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
|
||||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | Gatekeeper 规则集版本 | 新矩阵 case 可显式断言该版本 |
|
||||
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
|
||||
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
|
||||
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
|
||||
@@ -91,6 +94,7 @@ Java 类型:`DiagnosisEvalResult`
|
||||
| `requiredKeywordCount` | case 配置的关键词数量 |
|
||||
| `evidenceCoverage` | 每个必需工具是否出现 |
|
||||
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
|
||||
| `gatekeeperRuleSetVersion` | 读到的 `gatekeeper_result.rule_set_version` |
|
||||
| `composerStatus` | 读到的 `composer_output.status` |
|
||||
| `claimCheckCount` | `claim_checks` 数量 |
|
||||
| `toolCallCount` | trace 中工具调用总数 |
|
||||
@@ -125,6 +129,7 @@ Executor structured output
|
||||
|
||||
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
|
||||
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
|
||||
- 配置 `expectedGatekeeperRuleSetVersion` 的 case 必须匹配 `gatekeeper_result.rule_set_version`。
|
||||
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
|
||||
- Composer 输出必须记录 `status`。
|
||||
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
|
||||
|
||||
Reference in New Issue
Block a user