feat(eval): add executor audit closure checks
This commit is contained in:
+46
-14
@@ -1,6 +1,16 @@
|
||||
# Diagnosis Eval Harness
|
||||
|
||||
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
|
||||
This folder contains the fixed offline regression set for the MVP diagnosis Agent.
|
||||
|
||||
## Background
|
||||
|
||||
The current diagnosis chain is:
|
||||
|
||||
```text
|
||||
Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer
|
||||
```
|
||||
|
||||
Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo.
|
||||
|
||||
## Scope
|
||||
|
||||
@@ -14,7 +24,23 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
|
||||
|
||||
## Current Mode
|
||||
|
||||
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
The committed baseline currently contains:
|
||||
|
||||
```text
|
||||
8 fixed cases
|
||||
8 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
5 LOW_CONFID verdicts
|
||||
1 REJECT verdict
|
||||
```
|
||||
|
||||
The three V2 audit-closure cases cover:
|
||||
|
||||
- Gatekeeper failure for a fabricated tool invocation reference.
|
||||
- Unsupported claim filtering before the final answer.
|
||||
- Composer fallback rendering without raw Executor JSON leakage.
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -24,29 +50,35 @@ Run the focused evaluator test:
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||
```
|
||||
|
||||
The committed baseline report represents the current fixed fixture set:
|
||||
Run the broader phase-5 regression set:
|
||||
|
||||
```text
|
||||
5 fixed cases
|
||||
5 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
3 LOW_CONFID verdicts
|
||||
```powershell
|
||||
mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
||||
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
|
||||
|
||||
## Interview Story
|
||||
## Regression Signal
|
||||
|
||||
The harness gives the MVP a repeatable baseline:
|
||||
The harness is deterministic code, not an LLM judge:
|
||||
|
||||
```text
|
||||
fixed diagnosis case
|
||||
-> saved or runtime trace
|
||||
-> saved trace fixture
|
||||
-> rule-based trace validation
|
||||
-> JSON / Markdown report
|
||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
||||
-> regression signal for prompts, tools, retrieval, verifier, and composer behavior
|
||||
```
|
||||
|
||||
Stage 5 adds these V2 checks:
|
||||
|
||||
- `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`.
|
||||
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
|
||||
- `claim_checks` must be structurally auditable.
|
||||
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||
@@ -58,4 +90,4 @@ current report
|
||||
-> regressions, improvements, and changed signals
|
||||
```
|
||||
|
||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
||||
Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?
|
||||
|
||||
@@ -53,5 +53,56 @@
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["可以忽略"]
|
||||
},
|
||||
{
|
||||
"id": "gatekeeper-fabricated-invocation",
|
||||
"title": "Gatekeeper fabricated invocation",
|
||||
"question": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
||||
"traceFixture": "gatekeeper-fabricated-invocation-reject.json",
|
||||
"expectedRootCauseKeywords": ["证据", "引用", "失败"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["REJECT", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经完全确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["fail"],
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["连接池耗尽导致支付失败"]
|
||||
},
|
||||
{
|
||||
"id": "unsupported-claim-filtering",
|
||||
"title": "Unsupported claim filtering",
|
||||
"question": "订单超时是否可以确认由数据库主库故障导致?",
|
||||
"traceFixture": "unsupported-claim-filtering-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||
},
|
||||
{
|
||||
"id": "composer-fallback-no-raw-json",
|
||||
"title": "Composer fallback no raw JSON",
|
||||
"question": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
||||
"traceFixture": "composer-fallback-no-raw-json-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["慢响应", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["executor_evidence_v2", "answer_version", "claim_id"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedComposerStatuses": ["composer_malformed"],
|
||||
"forbiddenConfirmedClaimKeywords": ["线程池已经耗尽"]
|
||||
}
|
||||
]
|
||||
|
||||
@@ -0,0 +1,142 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-composer-fallback-no-raw-json",
|
||||
"query": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 47000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.45,
|
||||
"critical_fact_count": 2,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-slow-response",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "库存服务出现慢响应",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics",
|
||||
"evidence_excerpt": "inventory p99 latency increased"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-thread-pool",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "线程池已经耗尽",
|
||||
"support_level": "weak",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics",
|
||||
"evidence_excerpt": "inventory p99 latency increased"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询库存服务线程池指标",
|
||||
"reason": "当前只有慢响应指标"
|
||||
}
|
||||
],
|
||||
"missing_info": ["线程池队列长度", "活跃线程数"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-slow-response",
|
||||
"claim_text": "库存服务出现慢响应",
|
||||
"claim_type": "symptom",
|
||||
"verification": "direct_observation",
|
||||
"detail": "指标显示 inventory p99 latency increased",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-thread-pool",
|
||||
"claim_text": "线程池已经耗尽",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "external_unknown",
|
||||
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "库存服务出现慢响应",
|
||||
"verification": "direct_evidence",
|
||||
"detail": "指标显示 inventory p99 latency increased",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"fact": "线程池已经耗尽",
|
||||
"verification": "external_unknown",
|
||||
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"composer_output": {
|
||||
"status": "composer_malformed",
|
||||
"detail": "used safe fallback rendering",
|
||||
"answer_summary": "指标显示库存服务出现慢响应。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询库存服务线程池指标",
|
||||
"reason": "当前只有慢响应指标"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-composer-fallback-no-raw-json",
|
||||
"toolName": "query_metrics",
|
||||
"outputPreview": "inventory p99 latency increased",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,110 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
||||
"query": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 39000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "REJECT",
|
||||
"groundedness_score": 0.1,
|
||||
"critical_fact_count": 1,
|
||||
"gatekeeper_result": {
|
||||
"status": "fail",
|
||||
"failed_rules": ["evidence.invocation_ref"],
|
||||
"warnings": [],
|
||||
"errors": [
|
||||
{
|
||||
"rule_id": "evidence.invocation_ref",
|
||||
"field": "claims[0].evidence_bindings[0].tool_invocation_id",
|
||||
"message": "tool_invocation_id does not exist"
|
||||
}
|
||||
]
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "连接池耗尽导致支付失败",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 99,
|
||||
"tool_name": "query_logs",
|
||||
"evidence_excerpt": "connection pool exhausted"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "重新查询支付服务错误日志",
|
||||
"reason": "当前 Executor 证据引用无法回溯"
|
||||
}
|
||||
],
|
||||
"missing_info": ["需要有效的日志工具调用记录"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "连接池耗尽导致支付失败",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "unsupported",
|
||||
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "连接池耗尽导致支付失败",
|
||||
"verification": "unsupported",
|
||||
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "证据引用校验失败,不能确认根因。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "重新查询支付服务错误日志",
|
||||
"reason": "当前 Executor 证据引用无法回溯"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "payment failed without matching connection pool exhaustion entry",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,151 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-unsupported-claim-filtering",
|
||||
"query": "订单超时是否可以确认由数据库主库故障导致?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 44000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.42,
|
||||
"critical_fact_count": 2,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "订单接口出现超时",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs",
|
||||
"evidence_excerpt": "order api timeout"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-db-primary",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "数据库主库故障导致订单超时",
|
||||
"support_level": "weak",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs",
|
||||
"evidence_excerpt": "order api timeout"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询数据库主库状态和错误日志",
|
||||
"reason": "当前只有订单接口超时日志"
|
||||
}
|
||||
],
|
||||
"missing_info": ["数据库主库状态", "数据库错误日志"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"claim_text": "订单接口出现超时",
|
||||
"claim_type": "symptom",
|
||||
"verification": "direct_observation",
|
||||
"detail": "日志直接记录 order api timeout",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-db-primary",
|
||||
"claim_text": "数据库主库故障导致订单超时",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "unsupported",
|
||||
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "订单接口出现超时",
|
||||
"verification": "direct_evidence",
|
||||
"detail": "日志直接记录 order api timeout",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"fact": "数据库主库故障导致订单超时",
|
||||
"verification": "unsupported",
|
||||
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "日志显示订单接口超时,但数据库方向证据不足。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询数据库主库状态和错误日志",
|
||||
"reason": "当前只有订单接口超时日志"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-unsupported-claim-filtering",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "order api timeout",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -1,13 +1,14 @@
|
||||
{
|
||||
"totalCases" : 5,
|
||||
"passedCases" : 5,
|
||||
"totalCases" : 8,
|
||||
"passedCases" : 8,
|
||||
"passRate" : 1.0,
|
||||
"verdictDistribution" : {
|
||||
"PASS" : 2,
|
||||
"LOW_CONFID" : 3
|
||||
"LOW_CONFID" : 5,
|
||||
"REJECT" : 1
|
||||
},
|
||||
"averageToolCallCount" : 2.0,
|
||||
"averageDurationMs" : 45800.0,
|
||||
"averageToolCallCount" : 1.625,
|
||||
"averageDurationMs" : 44875.0,
|
||||
"results" : [ {
|
||||
"caseId" : "payment-timeout",
|
||||
"title" : "Payment API timeout",
|
||||
@@ -21,6 +22,9 @@
|
||||
"query_logs" : true,
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 3,
|
||||
"durationMs" : 42000
|
||||
}, {
|
||||
@@ -35,6 +39,9 @@
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 51000
|
||||
}, {
|
||||
@@ -48,6 +55,9 @@
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 36000
|
||||
}, {
|
||||
@@ -62,6 +72,9 @@
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 47000
|
||||
}, {
|
||||
@@ -76,7 +89,58 @@
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 53000
|
||||
}, {
|
||||
"caseId" : "gatekeeper-fabricated-invocation",
|
||||
"title" : "Gatekeeper fabricated invocation",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "REJECT",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "fail",
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 1,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 39000
|
||||
}, {
|
||||
"caseId" : "unsupported-claim-filtering",
|
||||
"title" : "Unsupported claim filtering",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 2,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 44000
|
||||
}, {
|
||||
"caseId" : "composer-fallback-no-raw-json",
|
||||
"title" : "Composer fallback no raw JSON",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"composerStatus" : "composer_malformed",
|
||||
"claimCheckCount" : 2,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 47000
|
||||
} ]
|
||||
}
|
||||
|
||||
@@ -1,22 +1,26 @@
|
||||
# Diagnosis Eval Report
|
||||
|
||||
- Total cases: 5
|
||||
- Passed cases: 5
|
||||
- Total cases: 8
|
||||
- Passed cases: 8
|
||||
- Pass rate: 100.00%
|
||||
- Average tool calls: 2.00
|
||||
- Average duration ms: 45800.00
|
||||
- Average tool calls: 1.63
|
||||
- Average duration ms: 44875.00
|
||||
|
||||
## Verdict Distribution
|
||||
|
||||
- PASS: 2
|
||||
- LOW_CONFID: 3
|
||||
- LOW_CONFID: 5
|
||||
- REJECT: 1
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
|
||||
| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | - | - | - | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | - | - | - | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | - | - | - | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 53000 | - |
|
||||
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | valid | 1 | 3/3 | 1 | 39000 | - |
|
||||
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | valid | 2 | 2/2 | 1 | 44000 | - |
|
||||
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
|
||||
|
||||
+94
-147
@@ -1,203 +1,150 @@
|
||||
# Diagnosis Eval Data Schema
|
||||
|
||||
这份文档记录评测基准里的数据结构。口语化理解就是:
|
||||
这份文档记录 `mvp/eval` 固定评测集的数据结构。评测器读取保存好的 trace fixture,用确定性规则判断这次 Agent 运行是否满足预期。
|
||||
|
||||
```text
|
||||
用例文件说“我要考什么”
|
||||
trace 文件说“Agent 实际做了什么”
|
||||
评测结果说“这次有没有跑偏”
|
||||
汇总报告说“整体稳定性怎么样”
|
||||
case 文件:我要考什么
|
||||
fixture 文件:Agent 实际做了什么
|
||||
评测结果:这条 case 是否通过,哪里失败
|
||||
baseline report:整套固定集当前认可的结果
|
||||
```
|
||||
|
||||
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。
|
||||
当前评测不调用 LLM 打分。
|
||||
|
||||
## 1. 用例定义
|
||||
## 1. Case 定义
|
||||
|
||||
文件:`mvp/eval/cases/diagnosis-cases.json`
|
||||
|
||||
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。
|
||||
每条 case 定义一个固定诊断场景。
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"traceFixture": "payment-timeout-pass.json",
|
||||
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
|
||||
"id": "unsupported-claim-filtering",
|
||||
"title": "Unsupported claim filtering",
|
||||
"question": "订单超时是否可以确认由数据库主库故障导致?",
|
||||
"traceFixture": "unsupported-claim-filtering-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
|
||||
"allowedVerdicts": ["PASS", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["无证据确定"]
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||
}
|
||||
```
|
||||
|
||||
字段说明:
|
||||
|
||||
| 字段 | 意思 | 评测器怎么用 |
|
||||
| 字段 | 含义 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 |
|
||||
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 |
|
||||
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 |
|
||||
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 |
|
||||
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 |
|
||||
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 |
|
||||
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 |
|
||||
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 |
|
||||
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 |
|
||||
| `id` | case 唯一标识 | 出现在报告中 |
|
||||
| `title` | 可读标题 | 出现在报告中 |
|
||||
| `question` | 原始用户问题 | 用于说明场景,fixture 模式不会真实发送给 Agent |
|
||||
| `traceFixture` | 对应 fixture 文件名 | 从 `mvp/eval/fixtures` 加载 |
|
||||
| `expectedRootCauseKeywords` | 最终答案应覆盖的关键词 | 在 `session.answer` 中做包含判断 |
|
||||
| `minKeywordMatches` | 最少命中关键词数 | 低于该值则失败 |
|
||||
| `requiredEvidenceTools` | 必须出现的证据工具 | 从 `toolInvocations` 和 `tool_trace_summary` 中收集 |
|
||||
| `allowedVerdicts` | 允许的 Verifier verdict | verdict 不在列表中则失败 |
|
||||
| `forbiddenAnswerKeywords` | 最终答案禁止出现的词 | 用于拦截过度自信或危险表达 |
|
||||
| `requireV2AuditClosure` | 是否要求 V2 审计闭环字段 | 要求 `gatekeeper_result`、`claim_checks`、`composer_output` 存在,并检查 raw JSON 泄漏 |
|
||||
| `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 |
|
||||
| `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` |
|
||||
| `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 |
|
||||
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
|
||||
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
|
||||
|
||||
## 2. Trace Fixture
|
||||
|
||||
目录:`mvp/eval/fixtures/*.json`
|
||||
|
||||
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。
|
||||
fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规则需要的字段。
|
||||
|
||||
当前会读取这些字段:
|
||||
|
||||
| Trace 字段 | 意思 | 评测器怎么用 |
|
||||
| Trace 字段 | 含义 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 |
|
||||
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 |
|
||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 |
|
||||
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor 结构化输出里的确认事实证据绑定 | 如果 trace 中存在结构化 Executor 输出,每条 confirmed claim 必须有证据绑定 |
|
||||
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 |
|
||||
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 |
|
||||
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 |
|
||||
|
||||
简单说,trace 里最重要的是三类信息:
|
||||
|
||||
```text
|
||||
最终回答:它说了什么
|
||||
工具证据:它查了什么
|
||||
Verifier:它自己有没有承认这个结论可靠
|
||||
```
|
||||
| `session.answer` | 最终用户答案 | 检查关键词、禁用词、unsupported claim 泄漏、raw JSON 泄漏 |
|
||||
| `session.totalDurationMs` | 运行耗时 | 进入报告 |
|
||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
|
||||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
|
||||
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
|
||||
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
|
||||
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
|
||||
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 看到的工具证据 | 用于补充证据工具覆盖 |
|
||||
| `toolInvocations[*].toolName` | 实际调用工具名 | 用于检查 `requiredEvidenceTools` |
|
||||
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 中 |
|
||||
|
||||
## 3. 单条评测结果
|
||||
|
||||
Java 类型:`DiagnosisEvalResult`
|
||||
|
||||
这是每条 case 跑完之后的判断结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| 字段 | 含义 |
|
||||
| --- | --- |
|
||||
| `caseId` | 对应的 case id |
|
||||
| `caseId` | 对应 case id |
|
||||
| `title` | case 标题 |
|
||||
| `passed` | 这条 case 是否通过 |
|
||||
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 |
|
||||
| `verdict` | 从 trace 里读出来的 Verifier verdict |
|
||||
| `matchedKeywordCount` | 最终回答命中的关键词数量 |
|
||||
| `requiredKeywordCount` | case 定义里一共有多少个关键词 |
|
||||
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` |
|
||||
| `toolCallCount` | 本次 trace 里工具调用总数 |
|
||||
| `durationMs` | 本次 trace 的耗时 |
|
||||
|
||||
判断通过的口语化规则:
|
||||
|
||||
```text
|
||||
回答要说到关键点
|
||||
该查的证据工具要查到
|
||||
Verifier 的结论要在可接受范围内
|
||||
回答不能出现危险的过度自信表达
|
||||
如果是 REJECT,就必须走降级模板
|
||||
```
|
||||
| `passed` | 该 case 是否通过 |
|
||||
| `failedChecks` | 失败原因列表 |
|
||||
| `verdict` | 从 trace 中读到的 Verifier verdict |
|
||||
| `matchedKeywordCount` | 最终答案命中的关键词数量 |
|
||||
| `requiredKeywordCount` | case 配置的关键词数量 |
|
||||
| `evidenceCoverage` | 每个必需工具是否出现 |
|
||||
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
|
||||
| `composerStatus` | 读到的 `composer_output.status` |
|
||||
| `claimCheckCount` | `claim_checks` 数量 |
|
||||
| `toolCallCount` | trace 中工具调用总数 |
|
||||
| `durationMs` | trace 总耗时 |
|
||||
|
||||
## 4. 汇总报告
|
||||
|
||||
Java 类型:`DiagnosisEvalReport`
|
||||
|
||||
这是整个基准集跑完之后的总结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| 字段 | 含义 |
|
||||
| --- | --- |
|
||||
| `totalCases` | 总共评测了多少条 case |
|
||||
| `passedCases` | 通过了多少条 |
|
||||
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` |
|
||||
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` |
|
||||
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 |
|
||||
| `totalCases` | case 总数 |
|
||||
| `passedCases` | 通过数 |
|
||||
| `passRate` | 通过率,范围 `0.0` 到 `1.0` |
|
||||
| `verdictDistribution` | Verifier verdict 分布 |
|
||||
| `averageToolCallCount` | 平均工具调用数 |
|
||||
| `averageDurationMs` | 平均耗时 |
|
||||
| `results` | 每条 case 的详细结果列表 |
|
||||
| `results` | 单条 case 结果列表 |
|
||||
|
||||
## 5. 怎么看这个基准
|
||||
## 5. Stage 5 V2 审计闭环规则
|
||||
|
||||
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答:
|
||||
阶段 5 关注的是“前四段链路是否能被固定评测证明”:
|
||||
|
||||
```text
|
||||
以前能过的诊断题,现在还过不过?
|
||||
它是不是少查了某些证据?
|
||||
它是不是变得更自信但证据不足?
|
||||
它是不是开始输出不该说的话?
|
||||
它是不是明显变慢了?
|
||||
Executor structured output
|
||||
-> Gatekeeper deterministic audit
|
||||
-> Verifier claim_checks
|
||||
-> Composer filtered final answer
|
||||
```
|
||||
|
||||
所以面试里可以这样讲:
|
||||
新增确定性规则:
|
||||
|
||||
```text
|
||||
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。
|
||||
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。
|
||||
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
|
||||
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
|
||||
```
|
||||
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
|
||||
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
|
||||
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
|
||||
- Composer 输出必须记录 `status`。
|
||||
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
|
||||
- case 配置的 `forbiddenConfirmedClaimKeywords` 不能出现在最终答案里。
|
||||
|
||||
## 6. Baseline Diff
|
||||
|
||||
Baseline diff 是拿两份 report 做对比:
|
||||
Baseline diff 比较两份 report:
|
||||
|
||||
```text
|
||||
baseline report:以前认可的基准结果
|
||||
current report:这次改动后跑出来的新结果
|
||||
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
|
||||
baseline report:已经认可的基准结果
|
||||
current report:当前代码/fixture 跑出的结果
|
||||
diff report:结构化列出退化、改善和普通变化
|
||||
```
|
||||
|
||||
Java 类型:
|
||||
主要退化信号:
|
||||
|
||||
- `DiagnosisEvalDiffReport`
|
||||
- `DiagnosisEvalDiffItem`
|
||||
|
||||
`DiagnosisEvalDiffReport` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `baselineTotalCases` | baseline 里有多少条 case |
|
||||
| `currentTotalCases` | current 里有多少条 case |
|
||||
| `baselinePassedCases` | baseline 通过了多少条 |
|
||||
| `currentPassedCases` | current 通过了多少条 |
|
||||
| `baselinePassRate` | baseline 通过率 |
|
||||
| `currentPassRate` | current 通过率 |
|
||||
| `regressionCount` | 退化项数量 |
|
||||
| `improvementCount` | 改善项数量 |
|
||||
| `changedCount` | 普通变化项数量 |
|
||||
| `hasRegression` | 是否存在退化 |
|
||||
| `items` | 具体 diff 明细 |
|
||||
|
||||
`DiagnosisEvalDiffItem` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
|
||||
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
|
||||
| `caseId` | 如果是单条 case 变化,这里记录 case id |
|
||||
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
|
||||
| `baselineValue` | baseline 里的值 |
|
||||
| `currentValue` | current 里的值 |
|
||||
| `delta` | 数值变化量;非数值变化为空 |
|
||||
| `message` | 给人看的变化说明 |
|
||||
|
||||
口语化判断规则:
|
||||
|
||||
```text
|
||||
pass rate 下降:退化
|
||||
case 从通过变失败:退化
|
||||
证据工具从有变没有:退化
|
||||
关键词命中变少:退化
|
||||
工具调用或耗时升高:成本上升,记为退化信号
|
||||
verdict 分布变化:记录变化,供人工判断是否符合预期
|
||||
```
|
||||
|
||||
面试里可以这样讲:
|
||||
|
||||
```text
|
||||
我把 baseline report 和当前 report 做结构化 diff。
|
||||
它不是再问 LLM,而是用代码比较固定字段。
|
||||
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
|
||||
diff 会直接标成 regression。
|
||||
这样 Agent 改动可以用固定基准做回归判断。
|
||||
```
|
||||
- pass rate 下降。
|
||||
- case 从通过变失败。
|
||||
- 必需证据工具从有变无。
|
||||
- 关键词命中减少。
|
||||
- 工具调用或耗时明显上升。
|
||||
- verdict 分布变化。
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-08
|
||||
@@ -0,0 +1,49 @@
|
||||
## Context
|
||||
|
||||
Executor Structured Output V2 introduced three audit layers after the original fixed-case eval harness was created:
|
||||
|
||||
- Gatekeeper result in `selfEvaluation.verifier_evaluation.gatekeeper_result`
|
||||
- Verifier claim-level checks in `selfEvaluation.verifier_evaluation.claim_checks`
|
||||
- Composer result in `selfEvaluation.verifier_evaluation.composer_output`
|
||||
|
||||
The existing harness proves that a saved trace has expected answer keywords, evidence tools, verdicts, and basic structured claim bindings. It does not yet prove that the V2 audit chain is internally consistent or that Composer filtered unsupported claims before writing the final user answer.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make V2 audit fields part of offline regression checks.
|
||||
- Catch physical/protocol regressions: missing Gatekeeper audit, PASS after Gatekeeper fail, missing claim checks, missing Composer audit, raw Executor JSON leakage, and unsupported claims leaking into the final answer.
|
||||
- Add fixture cases that exercise the new checks without starting the application.
|
||||
- Keep the evaluator deterministic and easy to explain to another implementation agent.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not redesign Planner, Executor, Gatekeeper, Verifier, or Composer.
|
||||
- Do not introduce an LLM-based evaluator.
|
||||
- Do not require live MySQL, Redis, Milvus, or application startup for the baseline fixture tests.
|
||||
- Do not add a broad new database schema; trace audit remains read from existing JSON fields.
|
||||
|
||||
## Decisions
|
||||
|
||||
1. Extend eval case expectations instead of hardcoding every V2 rule globally.
|
||||
|
||||
Some legacy or intentionally partial fixtures may not contain all V2 audit fields. Case-level expectations let the fixed baseline explicitly say which trace must prove Gatekeeper, claim checks, Composer, or leakage prevention. The default can remain backward-compatible while new V2 fixtures opt in to stricter checks.
|
||||
|
||||
2. Keep consistency checks lexical and deterministic.
|
||||
|
||||
The evaluator will not determine semantic truth from scratch. It will compare configured keywords against the final answer and audit fields. This is enough to catch the intended regression class: unsupported or contradicted claims being rendered as confirmed final answers.
|
||||
|
||||
3. Treat Gatekeeper failure as a hard regression if paired with `PASS`.
|
||||
|
||||
The production chain already guards this. The eval harness should independently fail any trace where `gatekeeper_result.status=fail` and Verifier verdict remains `PASS`, because that means the audit layer can no longer be trusted.
|
||||
|
||||
4. Record audit signals in eval results.
|
||||
|
||||
Per-case results should expose the observed Gatekeeper status, Composer status, and claim-check count so baseline JSON/Markdown reports remain useful during review.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Keyword-based final-answer checks can miss paraphrases. -> Mitigation: use them only for regression-sensitive fixtures where the unsafe claim keyword is deliberately fixed.
|
||||
- [Risk] Adding too many case fields makes fixtures harder to maintain. -> Mitigation: keep expectation fields small and optional.
|
||||
- [Risk] Baseline report changes may look like a product behavior change. -> Mitigation: document that this phase changes only eval fixtures/rules unless a production bug is discovered and fixed.
|
||||
@@ -0,0 +1,37 @@
|
||||
## Why
|
||||
|
||||
Executor Structured Output V2 已经完成 Executor、Gatekeeper、Verifier、Composer 四段主链路改造,但现有离线评测仍主要检查最终答案关键词、证据工具覆盖和 Verifier verdict。阶段 5 需要把新增的审计字段纳入固定回归门禁,证明 `Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer` 链路不是只在单次 demo 中可用。
|
||||
|
||||
## What Changes
|
||||
|
||||
- Extend the diagnosis eval harness so V2 audit fields are validated as first-class regression checks.
|
||||
- Add fixture coverage for Gatekeeper failure, claim verification filtering, Composer fallback, and raw JSON leakage prevention.
|
||||
- Update baseline reports to reflect the expanded fixed fixture set.
|
||||
- Update eval documentation so another agent can understand the background, stages, data fields, and acceptance commands.
|
||||
- No production protocol change is intended in this phase; production chain behavior should remain unchanged unless tests reveal a bug.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: add V2 audit-closure requirements for Gatekeeper, Verifier claim checks, Composer output, and final-answer consistency.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code:
|
||||
- `src/main/java/com/superbiz/agent/eval/*`
|
||||
- `src/test/java/com/superbiz/agent/eval/*`
|
||||
- Affected data:
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- `mvp/eval/fixtures/*.json`
|
||||
- `mvp/eval/reports/baseline-report.*`
|
||||
- `mvp/eval/schema.md`
|
||||
- `mvp/eval/README.md`
|
||||
- Verification:
|
||||
- focused evaluator tests
|
||||
- relevant Executor/Gatekeeper/Verifier/Composer integration tests
|
||||
- OpenSpec validation and archive
|
||||
+38
@@ -0,0 +1,38 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
|
||||
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
|
||||
|
||||
#### Scenario: Required V2 audit fields are present
|
||||
- **WHEN** an evaluation case requires V2 audit closure
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
|
||||
|
||||
#### Scenario: Gatekeeper failure cannot pass verification
|
||||
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
|
||||
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
|
||||
|
||||
#### Scenario: Claim checks are auditable
|
||||
- **WHEN** an evaluation case requires claim checks
|
||||
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
|
||||
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
|
||||
|
||||
#### Scenario: Composer output is auditable
|
||||
- **WHEN** an evaluation case requires Composer output
|
||||
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
|
||||
|
||||
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
|
||||
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
|
||||
|
||||
#### Scenario: Unsupported final-answer claim is rejected
|
||||
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
|
||||
|
||||
#### Scenario: Raw Executor JSON is not user-facing
|
||||
- **WHEN** a trace is evaluated under V2 audit closure
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
|
||||
|
||||
#### Scenario: Composer fallback still avoids raw JSON leakage
|
||||
- **WHEN** a trace records Composer fallback rendering
|
||||
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
|
||||
@@ -0,0 +1,23 @@
|
||||
## 1. Eval Contract
|
||||
|
||||
- [x] 1.1 Extend `DiagnosisEvalCase` with optional V2 audit expectations.
|
||||
- [x] 1.2 Extend `DiagnosisEvalResult` and report output with observed audit signals.
|
||||
- [x] 1.3 Add deterministic evaluator checks for Gatekeeper, claim checks, Composer output, and raw JSON leakage.
|
||||
|
||||
## 2. Fixtures and Baseline
|
||||
|
||||
- [x] 2.1 Add fixed eval cases for Gatekeeper failure, unsupported claim filtering, and Composer fallback leakage prevention.
|
||||
- [x] 2.2 Add matching trace fixtures for each new case.
|
||||
- [x] 2.3 Regenerate baseline JSON and Markdown reports.
|
||||
|
||||
## 3. Documentation
|
||||
|
||||
- [x] 3.1 Update eval README with stage 5 scope and verification commands.
|
||||
- [x] 3.2 Rewrite eval schema documentation so V2 audit fields and case expectations are readable.
|
||||
|
||||
## 4. Verification and Archive
|
||||
|
||||
- [x] 4.1 Add or update unit tests for the new V2 audit checks.
|
||||
- [x] 4.2 Run focused evaluator and agent-chain tests.
|
||||
- [x] 4.3 Validate and archive the OpenSpec change.
|
||||
- [x] 4.4 Commit the completed phase 5 changes.
|
||||
@@ -133,3 +133,40 @@ The system SHALL expose baseline diff output in structured JSON and reviewable M
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
|
||||
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
|
||||
|
||||
#### Scenario: Required V2 audit fields are present
|
||||
- **WHEN** an evaluation case requires V2 audit closure
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
|
||||
|
||||
#### Scenario: Gatekeeper failure cannot pass verification
|
||||
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
|
||||
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
|
||||
|
||||
#### Scenario: Claim checks are auditable
|
||||
- **WHEN** an evaluation case requires claim checks
|
||||
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
|
||||
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
|
||||
|
||||
#### Scenario: Composer output is auditable
|
||||
- **WHEN** an evaluation case requires Composer output
|
||||
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
|
||||
|
||||
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
|
||||
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
|
||||
|
||||
#### Scenario: Unsupported final-answer claim is rejected
|
||||
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
|
||||
|
||||
#### Scenario: Raw Executor JSON is not user-facing
|
||||
- **WHEN** a trace is evaluated under V2 audit closure
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
|
||||
|
||||
#### Scenario: Composer fallback still avoids raw JSON leakage
|
||||
- **WHEN** a trace records Composer fallback rendering
|
||||
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
|
||||
|
||||
@@ -22,4 +22,10 @@ public class DiagnosisEvalCase {
|
||||
private List<String> requiredEvidenceTools;
|
||||
private List<String> allowedVerdicts;
|
||||
private List<String> forbiddenAnswerKeywords;
|
||||
private Boolean requireV2AuditClosure;
|
||||
private Boolean requireClaimChecks;
|
||||
private Boolean requireComposerOutput;
|
||||
private List<String> expectedGatekeeperStatuses;
|
||||
private List<String> expectedComposerStatuses;
|
||||
private List<String> forbiddenConfirmedClaimKeywords;
|
||||
}
|
||||
|
||||
@@ -46,8 +46,8 @@ public class DiagnosisEvalReportWriter {
|
||||
}
|
||||
|
||||
builder.append("## Cases\n\n");
|
||||
builder.append("| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |\n");
|
||||
builder.append("| --- | --- | --- | --- | ---: | ---: | --- |\n");
|
||||
builder.append("| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |\n");
|
||||
builder.append("| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |\n");
|
||||
for (DiagnosisEvalResult result : report.getResults()) {
|
||||
builder.append("| ")
|
||||
.append(result.getCaseId())
|
||||
@@ -56,6 +56,12 @@ public class DiagnosisEvalReportWriter {
|
||||
.append(" | ")
|
||||
.append(valueOrDash(result.getVerdict()))
|
||||
.append(" | ")
|
||||
.append(valueOrDash(result.getGatekeeperStatus()))
|
||||
.append(" | ")
|
||||
.append(valueOrDash(result.getComposerStatus()))
|
||||
.append(" | ")
|
||||
.append(result.getClaimCheckCount() == null ? "-" : result.getClaimCheckCount())
|
||||
.append(" | ")
|
||||
.append(result.getMatchedKeywordCount()).append("/").append(result.getRequiredKeywordCount())
|
||||
.append(" | ")
|
||||
.append(result.getToolCallCount() == null ? "-" : result.getToolCallCount())
|
||||
|
||||
@@ -22,6 +22,9 @@ public class DiagnosisEvalResult {
|
||||
private int matchedKeywordCount;
|
||||
private int requiredKeywordCount;
|
||||
private Map<String, Boolean> evidenceCoverage;
|
||||
private String gatekeeperStatus;
|
||||
private String composerStatus;
|
||||
private Integer claimCheckCount;
|
||||
private Integer toolCallCount;
|
||||
private Integer durationMs;
|
||||
}
|
||||
|
||||
@@ -20,6 +20,20 @@ public class DiagnosisTraceEvaluator {
|
||||
|
||||
private static final TypeReference<List<DiagnosisEvalCase>> CASE_LIST_TYPE = new TypeReference<>() {};
|
||||
private static final String REJECT_DEGRADED_PREFIX = "当前无法基于已获取证据生成可靠结论";
|
||||
private static final Set<String> RAW_EXECUTOR_MARKERS = Set.of(
|
||||
"executor_evidence_v2",
|
||||
"answer_version",
|
||||
"evidence_bindings",
|
||||
"claim_id"
|
||||
);
|
||||
private static final Set<String> VALID_CLAIM_VERIFICATIONS = Set.of(
|
||||
"direct_observation",
|
||||
"reasonable_inference",
|
||||
"overstated",
|
||||
"unsupported",
|
||||
"external_unknown",
|
||||
"contradicted"
|
||||
);
|
||||
|
||||
private final ObjectMapper objectMapper;
|
||||
|
||||
@@ -51,6 +65,9 @@ public class DiagnosisTraceEvaluator {
|
||||
.matchedKeywordCount(0)
|
||||
.requiredKeywordCount(size(evalCase.getExpectedRootCauseKeywords()))
|
||||
.evidenceCoverage(emptyCoverage(evalCase.getRequiredEvidenceTools()))
|
||||
.gatekeeperStatus(null)
|
||||
.composerStatus(null)
|
||||
.claimCheckCount(null)
|
||||
.toolCallCount(null)
|
||||
.durationMs(null)
|
||||
.build());
|
||||
@@ -102,6 +119,11 @@ public class DiagnosisTraceEvaluator {
|
||||
}
|
||||
|
||||
failedChecks.addAll(validateExecutorStructuredOutput(trace));
|
||||
String gatekeeperStatus = extractNestedString(trace, "verifier_evaluation", "gatekeeper_result", "status");
|
||||
String composerStatus = extractNestedString(trace, "verifier_evaluation", "composer_output", "status");
|
||||
Integer claimCheckCount = countList(trace, "verifier_evaluation", "claim_checks");
|
||||
failedChecks.addAll(validateV2AuditClosure(evalCase, trace, normalizedAnswer, verdict,
|
||||
gatekeeperStatus, composerStatus));
|
||||
|
||||
Integer toolCallCount = trace.getToolInvocations() == null ? 0 : trace.getToolInvocations().size();
|
||||
Integer durationMs = trace.getSession() == null ? null : trace.getSession().getTotalDurationMs();
|
||||
@@ -115,6 +137,9 @@ public class DiagnosisTraceEvaluator {
|
||||
.matchedKeywordCount(matchedKeywordCount)
|
||||
.requiredKeywordCount(requiredKeywordCount)
|
||||
.evidenceCoverage(evidenceCoverage)
|
||||
.gatekeeperStatus(gatekeeperStatus)
|
||||
.composerStatus(composerStatus)
|
||||
.claimCheckCount(claimCheckCount)
|
||||
.toolCallCount(toolCallCount)
|
||||
.durationMs(durationMs)
|
||||
.build();
|
||||
@@ -202,6 +227,91 @@ public class DiagnosisTraceEvaluator {
|
||||
return failedChecks;
|
||||
}
|
||||
|
||||
private List<String> validateV2AuditClosure(DiagnosisEvalCase evalCase,
|
||||
DiagnosisTraceResponse trace,
|
||||
String normalizedAnswer,
|
||||
String verdict,
|
||||
String gatekeeperStatus,
|
||||
String composerStatus) {
|
||||
List<String> failedChecks = new ArrayList<>();
|
||||
boolean requireV2AuditClosure = Boolean.TRUE.equals(evalCase.getRequireV2AuditClosure());
|
||||
boolean requireClaimChecks = requireV2AuditClosure || Boolean.TRUE.equals(evalCase.getRequireClaimChecks());
|
||||
boolean requireComposerOutput = requireV2AuditClosure || Boolean.TRUE.equals(evalCase.getRequireComposerOutput());
|
||||
|
||||
Object gatekeeperResult = nestedValue(trace, "verifier_evaluation", "gatekeeper_result");
|
||||
if (requireV2AuditClosure && !(gatekeeperResult instanceof Map<?, ?>)) {
|
||||
failedChecks.add("missing gatekeeper_result");
|
||||
}
|
||||
if ("fail".equals(gatekeeperStatus) && "PASS".equals(verdict)) {
|
||||
failedChecks.add("gatekeeper fail cannot have PASS verdict");
|
||||
}
|
||||
if (!safeList(evalCase.getExpectedGatekeeperStatuses()).isEmpty()
|
||||
&& !safeList(evalCase.getExpectedGatekeeperStatuses()).contains(gatekeeperStatus)) {
|
||||
failedChecks.add("gatekeeper status not expected: " + valueOrMissing(gatekeeperStatus));
|
||||
}
|
||||
|
||||
failedChecks.addAll(validateClaimChecks(trace, requireClaimChecks));
|
||||
|
||||
Object composerOutput = nestedValue(trace, "verifier_evaluation", "composer_output");
|
||||
if (requireComposerOutput && !(composerOutput instanceof Map<?, ?>)) {
|
||||
failedChecks.add("missing composer_output");
|
||||
}
|
||||
if (requireComposerOutput && isBlank(composerStatus)) {
|
||||
failedChecks.add("composer_output missing status");
|
||||
}
|
||||
if (!safeList(evalCase.getExpectedComposerStatuses()).isEmpty()
|
||||
&& !safeList(evalCase.getExpectedComposerStatuses()).contains(composerStatus)) {
|
||||
failedChecks.add("composer status not expected: " + valueOrMissing(composerStatus));
|
||||
}
|
||||
|
||||
for (String forbidden : safeList(evalCase.getForbiddenConfirmedClaimKeywords())) {
|
||||
if (normalizedAnswer.contains(forbidden.toLowerCase(Locale.ROOT))) {
|
||||
failedChecks.add("answer contains forbidden confirmed claim keyword: " + forbidden);
|
||||
}
|
||||
}
|
||||
|
||||
if (requireV2AuditClosure) {
|
||||
for (String marker : RAW_EXECUTOR_MARKERS) {
|
||||
if (normalizedAnswer.contains(marker.toLowerCase(Locale.ROOT))) {
|
||||
failedChecks.add("answer leaks raw executor marker: " + marker);
|
||||
}
|
||||
}
|
||||
}
|
||||
return failedChecks;
|
||||
}
|
||||
|
||||
private List<String> validateClaimChecks(DiagnosisTraceResponse trace, boolean required) {
|
||||
Object claimChecks = nestedValue(trace, "verifier_evaluation", "claim_checks");
|
||||
if (!(claimChecks instanceof List<?> claimCheckList)) {
|
||||
return required ? List.of("missing claim_checks") : List.of();
|
||||
}
|
||||
if (required && claimCheckList.isEmpty()) {
|
||||
return List.of("claim_checks is empty");
|
||||
}
|
||||
|
||||
List<String> failedChecks = new ArrayList<>();
|
||||
for (Object item : claimCheckList) {
|
||||
if (!(item instanceof Map<?, ?> claimCheck)) {
|
||||
failedChecks.add("claim_check is not an object");
|
||||
continue;
|
||||
}
|
||||
String claimId = stringValue(claimCheck.get("claim_id"));
|
||||
String verification = stringValue(claimCheck.get("verification"));
|
||||
if (isBlank(claimId)) {
|
||||
failedChecks.add("claim_check missing claim_id");
|
||||
}
|
||||
if (isBlank(verification)) {
|
||||
failedChecks.add("claim_check missing verification: " + valueOrMissing(claimId));
|
||||
} else if (!VALID_CLAIM_VERIFICATIONS.contains(verification)) {
|
||||
failedChecks.add("claim_check verification invalid: " + verification);
|
||||
}
|
||||
if (isBlank(stringValue(claimCheck.get("detail")))) {
|
||||
failedChecks.add("claim_check missing detail: " + valueOrMissing(claimId));
|
||||
}
|
||||
}
|
||||
return failedChecks;
|
||||
}
|
||||
|
||||
private Object nestedValue(DiagnosisTraceResponse trace, String firstKey, String secondKey) {
|
||||
if (trace.getSession() == null || trace.getSession().getSelfEvaluation() == null) {
|
||||
return null;
|
||||
@@ -213,6 +323,20 @@ public class DiagnosisTraceEvaluator {
|
||||
return map.get(secondKey);
|
||||
}
|
||||
|
||||
private String extractNestedString(DiagnosisTraceResponse trace, String firstKey, String secondKey, String thirdKey) {
|
||||
Object value = nestedValue(trace, firstKey, secondKey);
|
||||
if (!(value instanceof Map<?, ?> map)) {
|
||||
return null;
|
||||
}
|
||||
Object nested = map.get(thirdKey);
|
||||
return nested == null ? null : String.valueOf(nested);
|
||||
}
|
||||
|
||||
private Integer countList(DiagnosisTraceResponse trace, String firstKey, String secondKey) {
|
||||
Object value = nestedValue(trace, firstKey, secondKey);
|
||||
return value instanceof List<?> list ? list.size() : null;
|
||||
}
|
||||
|
||||
private int countMatches(String normalizedAnswer, List<String> keywords) {
|
||||
int count = 0;
|
||||
for (String keyword : safeList(keywords)) {
|
||||
@@ -242,4 +366,16 @@ public class DiagnosisTraceEvaluator {
|
||||
private String nullToEmpty(String value) {
|
||||
return value == null ? "" : value;
|
||||
}
|
||||
|
||||
private String stringValue(Object value) {
|
||||
return value == null ? null : String.valueOf(value);
|
||||
}
|
||||
|
||||
private boolean isBlank(String value) {
|
||||
return value == null || value.isBlank();
|
||||
}
|
||||
|
||||
private String valueOrMissing(String value) {
|
||||
return isBlank(value) ? "missing" : value;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -100,14 +100,14 @@ class DiagnosisEvalBaselineDiffTest {
|
||||
}
|
||||
|
||||
private void degradeRedisCase(DiagnosisEvalReport report) {
|
||||
report.setPassedCases(4);
|
||||
report.setPassRate(0.8);
|
||||
report.setPassedCases(7);
|
||||
report.setPassRate(0.875);
|
||||
report.setAverageToolCallCount(3.0);
|
||||
report.setAverageDurationMs(45800.0);
|
||||
report.setAverageDurationMs(44875.0);
|
||||
report.setVerdictDistribution(new LinkedHashMap<>());
|
||||
report.getVerdictDistribution().put("PASS", 2L);
|
||||
report.getVerdictDistribution().put("LOW_CONFID", 2L);
|
||||
report.getVerdictDistribution().put("REJECT", 1L);
|
||||
report.getVerdictDistribution().put("LOW_CONFID", 4L);
|
||||
report.getVerdictDistribution().put("REJECT", 2L);
|
||||
|
||||
DiagnosisEvalResult redis = result(report, "redis-timeout");
|
||||
redis.setPassed(false);
|
||||
|
||||
@@ -24,11 +24,12 @@ class DiagnosisTraceEvaluatorTest {
|
||||
|
||||
DiagnosisEvalReport report = evaluator.evaluate(cases, Path.of("mvp/eval/fixtures"));
|
||||
|
||||
assertEquals(5, report.getTotalCases());
|
||||
assertEquals(5, report.getPassedCases());
|
||||
assertEquals(8, report.getTotalCases());
|
||||
assertEquals(8, report.getPassedCases());
|
||||
assertEquals(1.0, report.getPassRate(), 0.001);
|
||||
assertEquals(2L, report.getVerdictDistribution().get("PASS"));
|
||||
assertEquals(3L, report.getVerdictDistribution().get("LOW_CONFID"));
|
||||
assertEquals(5L, report.getVerdictDistribution().get("LOW_CONFID"));
|
||||
assertEquals(1L, report.getVerdictDistribution().get("REJECT"));
|
||||
|
||||
DiagnosisEvalResult payment = result(report, "payment-timeout");
|
||||
assertTrue(payment.isPassed());
|
||||
@@ -39,6 +40,16 @@ class DiagnosisTraceEvaluatorTest {
|
||||
DiagnosisEvalResult redis = result(report, "redis-timeout");
|
||||
assertTrue(redis.isPassed());
|
||||
assertTrue(redis.getEvidenceCoverage().get("query_logs"));
|
||||
|
||||
DiagnosisEvalResult fabricatedInvocation = result(report, "gatekeeper-fabricated-invocation");
|
||||
assertTrue(fabricatedInvocation.isPassed());
|
||||
assertEquals("fail", fabricatedInvocation.getGatekeeperStatus());
|
||||
assertEquals("valid", fabricatedInvocation.getComposerStatus());
|
||||
assertEquals(1, fabricatedInvocation.getClaimCheckCount());
|
||||
|
||||
DiagnosisEvalResult composerFallback = result(report, "composer-fallback-no-raw-json");
|
||||
assertTrue(composerFallback.isPassed());
|
||||
assertEquals("composer_malformed", composerFallback.getComposerStatus());
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -108,6 +119,111 @@ class DiagnosisTraceEvaluatorTest {
|
||||
"executor confirmed claim missing evidence bindings: claim-unsupported"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void evaluateFailsWhenGatekeeperFailStillPassesVerifier() {
|
||||
DiagnosisEvalCase evalCase = DiagnosisEvalCase.builder()
|
||||
.id("gatekeeper-pass-leak")
|
||||
.title("Gatekeeper pass leak")
|
||||
.expectedRootCauseKeywords(List.of())
|
||||
.requiredEvidenceTools(List.of())
|
||||
.allowedVerdicts(List.of("PASS", "LOW_CONFID", "REJECT"))
|
||||
.requireV2AuditClosure(true)
|
||||
.build();
|
||||
DiagnosisTraceResponse trace = DiagnosisTraceResponse.builder()
|
||||
.session(DiagnosisTraceResponse.SessionTrace.builder()
|
||||
.answer("安全回答")
|
||||
.selfEvaluation(java.util.Map.of(
|
||||
"verifier_evaluation", java.util.Map.of(
|
||||
"verdict", "PASS",
|
||||
"gatekeeper_result", java.util.Map.of("status", "fail"),
|
||||
"claim_checks", java.util.List.of(java.util.Map.of(
|
||||
"claim_id", "claim-1",
|
||||
"verification", "unsupported",
|
||||
"detail", "evidence ref invalid"
|
||||
)),
|
||||
"composer_output", java.util.Map.of("status", "valid")
|
||||
)))
|
||||
.build())
|
||||
.toolInvocations(List.of())
|
||||
.build();
|
||||
|
||||
DiagnosisEvalResult result = evaluator.evaluate(evalCase, trace);
|
||||
|
||||
assertFalse(result.isPassed());
|
||||
assertTrue(result.getFailedChecks().contains("gatekeeper fail cannot have PASS verdict"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void evaluateFailsWhenUnsupportedClaimLeaksIntoFinalAnswer() {
|
||||
DiagnosisEvalCase evalCase = DiagnosisEvalCase.builder()
|
||||
.id("unsupported-leak")
|
||||
.title("Unsupported leak")
|
||||
.expectedRootCauseKeywords(List.of())
|
||||
.requiredEvidenceTools(List.of())
|
||||
.allowedVerdicts(List.of("LOW_CONFID"))
|
||||
.requireV2AuditClosure(true)
|
||||
.forbiddenConfirmedClaimKeywords(List.of("主库故障"))
|
||||
.build();
|
||||
DiagnosisTraceResponse trace = DiagnosisTraceResponse.builder()
|
||||
.session(DiagnosisTraceResponse.SessionTrace.builder()
|
||||
.answer("已经确认主库故障。")
|
||||
.selfEvaluation(java.util.Map.of(
|
||||
"verifier_evaluation", java.util.Map.of(
|
||||
"verdict", "LOW_CONFID",
|
||||
"gatekeeper_result", java.util.Map.of("status", "pass"),
|
||||
"claim_checks", java.util.List.of(java.util.Map.of(
|
||||
"claim_id", "claim-1",
|
||||
"verification", "unsupported",
|
||||
"detail", "missing database evidence"
|
||||
)),
|
||||
"composer_output", java.util.Map.of("status", "valid")
|
||||
)))
|
||||
.build())
|
||||
.toolInvocations(List.of())
|
||||
.build();
|
||||
|
||||
DiagnosisEvalResult result = evaluator.evaluate(evalCase, trace);
|
||||
|
||||
assertFalse(result.isPassed());
|
||||
assertTrue(result.getFailedChecks().contains(
|
||||
"answer contains forbidden confirmed claim keyword: 主库故障"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void evaluateFailsWhenFinalAnswerLeaksRawExecutorMarker() {
|
||||
DiagnosisEvalCase evalCase = DiagnosisEvalCase.builder()
|
||||
.id("raw-json-leak")
|
||||
.title("Raw json leak")
|
||||
.expectedRootCauseKeywords(List.of())
|
||||
.requiredEvidenceTools(List.of())
|
||||
.allowedVerdicts(List.of("LOW_CONFID"))
|
||||
.requireV2AuditClosure(true)
|
||||
.build();
|
||||
DiagnosisTraceResponse trace = DiagnosisTraceResponse.builder()
|
||||
.session(DiagnosisTraceResponse.SessionTrace.builder()
|
||||
.answer("answer_version=executor_evidence_v2")
|
||||
.selfEvaluation(java.util.Map.of(
|
||||
"verifier_evaluation", java.util.Map.of(
|
||||
"verdict", "LOW_CONFID",
|
||||
"gatekeeper_result", java.util.Map.of("status", "pass"),
|
||||
"claim_checks", java.util.List.of(java.util.Map.of(
|
||||
"claim_id", "claim-1",
|
||||
"verification", "direct_observation",
|
||||
"detail", "log evidence"
|
||||
)),
|
||||
"composer_output", java.util.Map.of("status", "composer_malformed")
|
||||
)))
|
||||
.build())
|
||||
.toolInvocations(List.of())
|
||||
.build();
|
||||
|
||||
DiagnosisEvalResult result = evaluator.evaluate(evalCase, trace);
|
||||
|
||||
assertFalse(result.isPassed());
|
||||
assertTrue(result.getFailedChecks().contains("answer leaks raw executor marker: executor_evidence_v2"));
|
||||
assertTrue(result.getFailedChecks().contains("answer leaks raw executor marker: answer_version"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void reportWriterOutputsJsonAndMarkdown(@TempDir Path tempDir) throws Exception {
|
||||
DiagnosisEvalReport report = evaluator.evaluate(readCases(), Path.of("mvp/eval/fixtures"));
|
||||
|
||||
Reference in New Issue
Block a user