156 lines
7.1 KiB
Markdown
156 lines
7.1 KiB
Markdown
# Diagnosis Eval Data Schema
|
||
|
||
这份文档记录 `mvp/eval` 固定评测集的数据结构。评测器读取保存好的 trace fixture,用确定性规则判断这次 Agent 运行是否满足预期。
|
||
|
||
```text
|
||
case 文件:我要考什么
|
||
fixture 文件:Agent 实际做了什么
|
||
评测结果:这条 case 是否通过,哪里失败
|
||
baseline report:整套固定集当前认可的结果
|
||
```
|
||
|
||
当前评测不调用 LLM 打分。
|
||
|
||
## 1. Case 定义
|
||
|
||
文件:`mvp/eval/cases/diagnosis-cases.json`
|
||
|
||
每条 case 定义一个固定诊断场景。
|
||
|
||
```json
|
||
{
|
||
"id": "unsupported-claim-filtering",
|
||
"title": "Unsupported claim filtering",
|
||
"question": "订单超时是否可以确认由数据库主库故障导致?",
|
||
"traceFixture": "unsupported-claim-filtering-low-confid.json",
|
||
"expectedRootCauseKeywords": ["超时", "证据"],
|
||
"minKeywordMatches": 2,
|
||
"requiredEvidenceTools": ["query_logs"],
|
||
"allowedVerdicts": ["LOW_CONFID"],
|
||
"forbiddenAnswerKeywords": ["已经确认"],
|
||
"requireV2AuditClosure": true,
|
||
"requireClaimChecks": true,
|
||
"requireComposerOutput": true,
|
||
"expectedGatekeeperStatuses": ["pass"],
|
||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||
"expectedComposerStatuses": ["valid"],
|
||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||
}
|
||
```
|
||
|
||
字段说明:
|
||
|
||
| 字段 | 含义 | 评测器怎么用 |
|
||
| --- | --- | --- |
|
||
| `id` | case 唯一标识 | 出现在报告中 |
|
||
| `title` | 可读标题 | 出现在报告中 |
|
||
| `question` | 原始用户问题 | 用于说明场景,fixture 模式不会真实发送给 Agent |
|
||
| `traceFixture` | 对应 fixture 文件名 | 从 `mvp/eval/fixtures` 加载 |
|
||
| `expectedRootCauseKeywords` | 最终答案应覆盖的关键词 | 在 `session.answer` 中做包含判断 |
|
||
| `minKeywordMatches` | 最少命中关键词数 | 低于该值则失败 |
|
||
| `requiredEvidenceTools` | 必须出现的证据工具 | 从 `toolInvocations` 和 `tool_trace_summary` 中收集 |
|
||
| `allowedVerdicts` | 允许的 Verifier verdict | verdict 不在列表中则失败 |
|
||
| `forbiddenAnswerKeywords` | 最终答案禁止出现的词 | 用于拦截过度自信或危险表达 |
|
||
| `requireV2AuditClosure` | 是否要求 V2 审计闭环字段 | 要求 `gatekeeper_result`、`claim_checks`、`composer_output` 存在,并检查 raw JSON 泄漏 |
|
||
| `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 |
|
||
| `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` |
|
||
| `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 |
|
||
| `expectedGatekeeperRuleSetVersion` | 期望的 Gatekeeper 规则集版本 | 配置后校验 `gatekeeper_result.rule_set_version` |
|
||
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
|
||
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
|
||
|
||
## 2. Trace Fixture
|
||
|
||
目录:`mvp/eval/fixtures/*.json`
|
||
|
||
fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规则需要的字段。
|
||
|
||
| Trace 字段 | 含义 | 评测器怎么用 |
|
||
| --- | --- | --- |
|
||
| `session.answer` | 最终用户答案 | 检查关键词、禁用词、unsupported claim 泄漏、raw JSON 泄漏 |
|
||
| `session.totalDurationMs` | 运行耗时 | 进入报告 |
|
||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
|
||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
|
||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | Gatekeeper 规则集版本 | 新矩阵 case 可显式断言该版本 |
|
||
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
|
||
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
|
||
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
|
||
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 看到的工具证据 | 用于补充证据工具覆盖 |
|
||
| `toolInvocations[*].toolName` | 实际调用工具名 | 用于检查 `requiredEvidenceTools` |
|
||
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 中 |
|
||
|
||
## 3. 单条评测结果
|
||
|
||
Java 类型:`DiagnosisEvalResult`
|
||
|
||
| 字段 | 含义 |
|
||
| --- | --- |
|
||
| `caseId` | 对应 case id |
|
||
| `title` | case 标题 |
|
||
| `passed` | 该 case 是否通过 |
|
||
| `failedChecks` | 失败原因列表 |
|
||
| `verdict` | 从 trace 中读到的 Verifier verdict |
|
||
| `matchedKeywordCount` | 最终答案命中的关键词数量 |
|
||
| `requiredKeywordCount` | case 配置的关键词数量 |
|
||
| `evidenceCoverage` | 每个必需工具是否出现 |
|
||
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
|
||
| `gatekeeperRuleSetVersion` | 读到的 `gatekeeper_result.rule_set_version` |
|
||
| `composerStatus` | 读到的 `composer_output.status` |
|
||
| `claimCheckCount` | `claim_checks` 数量 |
|
||
| `toolCallCount` | trace 中工具调用总数 |
|
||
| `durationMs` | trace 总耗时 |
|
||
|
||
## 4. 汇总报告
|
||
|
||
Java 类型:`DiagnosisEvalReport`
|
||
|
||
| 字段 | 含义 |
|
||
| --- | --- |
|
||
| `totalCases` | case 总数 |
|
||
| `passedCases` | 通过数 |
|
||
| `passRate` | 通过率,范围 `0.0` 到 `1.0` |
|
||
| `verdictDistribution` | Verifier verdict 分布 |
|
||
| `averageToolCallCount` | 平均工具调用数 |
|
||
| `averageDurationMs` | 平均耗时 |
|
||
| `results` | 单条 case 结果列表 |
|
||
|
||
## 5. Stage 5 V2 审计闭环规则
|
||
|
||
阶段 5 关注的是“前四段链路是否能被固定评测证明”:
|
||
|
||
```text
|
||
Executor structured output
|
||
-> Gatekeeper deterministic audit
|
||
-> Verifier claim_checks
|
||
-> Composer filtered final answer
|
||
```
|
||
|
||
新增确定性规则:
|
||
|
||
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
|
||
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
|
||
- 配置 `expectedGatekeeperRuleSetVersion` 的 case 必须匹配 `gatekeeper_result.rule_set_version`。
|
||
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
|
||
- Composer 输出必须记录 `status`。
|
||
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
|
||
- case 配置的 `forbiddenConfirmedClaimKeywords` 不能出现在最终答案里。
|
||
|
||
## 6. Baseline Diff
|
||
|
||
Baseline diff 比较两份 report:
|
||
|
||
```text
|
||
baseline report:已经认可的基准结果
|
||
current report:当前代码/fixture 跑出的结果
|
||
diff report:结构化列出退化、改善和普通变化
|
||
```
|
||
|
||
主要退化信号:
|
||
|
||
- pass rate 下降。
|
||
- case 从通过变失败。
|
||
- 必需证据工具从有变无。
|
||
- 关键词命中减少。
|
||
- 工具调用或耗时明显上升。
|
||
- verdict 分布变化。
|