feat(eval): add executor audit closure checks
This commit is contained in:
+46
-14
@@ -1,6 +1,16 @@
|
||||
# Diagnosis Eval Harness
|
||||
|
||||
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
|
||||
This folder contains the fixed offline regression set for the MVP diagnosis Agent.
|
||||
|
||||
## Background
|
||||
|
||||
The current diagnosis chain is:
|
||||
|
||||
```text
|
||||
Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer
|
||||
```
|
||||
|
||||
Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo.
|
||||
|
||||
## Scope
|
||||
|
||||
@@ -14,7 +24,23 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
|
||||
|
||||
## Current Mode
|
||||
|
||||
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
The committed baseline currently contains:
|
||||
|
||||
```text
|
||||
8 fixed cases
|
||||
8 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
5 LOW_CONFID verdicts
|
||||
1 REJECT verdict
|
||||
```
|
||||
|
||||
The three V2 audit-closure cases cover:
|
||||
|
||||
- Gatekeeper failure for a fabricated tool invocation reference.
|
||||
- Unsupported claim filtering before the final answer.
|
||||
- Composer fallback rendering without raw Executor JSON leakage.
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -24,29 +50,35 @@ Run the focused evaluator test:
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||
```
|
||||
|
||||
The committed baseline report represents the current fixed fixture set:
|
||||
Run the broader phase-5 regression set:
|
||||
|
||||
```text
|
||||
5 fixed cases
|
||||
5 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
3 LOW_CONFID verdicts
|
||||
```powershell
|
||||
mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
||||
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
|
||||
|
||||
## Interview Story
|
||||
## Regression Signal
|
||||
|
||||
The harness gives the MVP a repeatable baseline:
|
||||
The harness is deterministic code, not an LLM judge:
|
||||
|
||||
```text
|
||||
fixed diagnosis case
|
||||
-> saved or runtime trace
|
||||
-> saved trace fixture
|
||||
-> rule-based trace validation
|
||||
-> JSON / Markdown report
|
||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
||||
-> regression signal for prompts, tools, retrieval, verifier, and composer behavior
|
||||
```
|
||||
|
||||
Stage 5 adds these V2 checks:
|
||||
|
||||
- `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`.
|
||||
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
|
||||
- `claim_checks` must be structurally auditable.
|
||||
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||
@@ -58,4 +90,4 @@ current report
|
||||
-> regressions, improvements, and changed signals
|
||||
```
|
||||
|
||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
||||
Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?
|
||||
|
||||
@@ -53,5 +53,56 @@
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["可以忽略"]
|
||||
},
|
||||
{
|
||||
"id": "gatekeeper-fabricated-invocation",
|
||||
"title": "Gatekeeper fabricated invocation",
|
||||
"question": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
||||
"traceFixture": "gatekeeper-fabricated-invocation-reject.json",
|
||||
"expectedRootCauseKeywords": ["证据", "引用", "失败"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["REJECT", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经完全确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["fail"],
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["连接池耗尽导致支付失败"]
|
||||
},
|
||||
{
|
||||
"id": "unsupported-claim-filtering",
|
||||
"title": "Unsupported claim filtering",
|
||||
"question": "订单超时是否可以确认由数据库主库故障导致?",
|
||||
"traceFixture": "unsupported-claim-filtering-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||
},
|
||||
{
|
||||
"id": "composer-fallback-no-raw-json",
|
||||
"title": "Composer fallback no raw JSON",
|
||||
"question": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
||||
"traceFixture": "composer-fallback-no-raw-json-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["慢响应", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["executor_evidence_v2", "answer_version", "claim_id"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedComposerStatuses": ["composer_malformed"],
|
||||
"forbiddenConfirmedClaimKeywords": ["线程池已经耗尽"]
|
||||
}
|
||||
]
|
||||
|
||||
@@ -0,0 +1,142 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-composer-fallback-no-raw-json",
|
||||
"query": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 47000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.45,
|
||||
"critical_fact_count": 2,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-slow-response",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "库存服务出现慢响应",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics",
|
||||
"evidence_excerpt": "inventory p99 latency increased"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-thread-pool",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "线程池已经耗尽",
|
||||
"support_level": "weak",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics",
|
||||
"evidence_excerpt": "inventory p99 latency increased"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询库存服务线程池指标",
|
||||
"reason": "当前只有慢响应指标"
|
||||
}
|
||||
],
|
||||
"missing_info": ["线程池队列长度", "活跃线程数"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-slow-response",
|
||||
"claim_text": "库存服务出现慢响应",
|
||||
"claim_type": "symptom",
|
||||
"verification": "direct_observation",
|
||||
"detail": "指标显示 inventory p99 latency increased",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-thread-pool",
|
||||
"claim_text": "线程池已经耗尽",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "external_unknown",
|
||||
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "库存服务出现慢响应",
|
||||
"verification": "direct_evidence",
|
||||
"detail": "指标显示 inventory p99 latency increased",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"fact": "线程池已经耗尽",
|
||||
"verification": "external_unknown",
|
||||
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"composer_output": {
|
||||
"status": "composer_malformed",
|
||||
"detail": "used safe fallback rendering",
|
||||
"answer_summary": "指标显示库存服务出现慢响应。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询库存服务线程池指标",
|
||||
"reason": "当前只有慢响应指标"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-composer-fallback-no-raw-json",
|
||||
"toolName": "query_metrics",
|
||||
"outputPreview": "inventory p99 latency increased",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,110 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
||||
"query": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 39000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "REJECT",
|
||||
"groundedness_score": 0.1,
|
||||
"critical_fact_count": 1,
|
||||
"gatekeeper_result": {
|
||||
"status": "fail",
|
||||
"failed_rules": ["evidence.invocation_ref"],
|
||||
"warnings": [],
|
||||
"errors": [
|
||||
{
|
||||
"rule_id": "evidence.invocation_ref",
|
||||
"field": "claims[0].evidence_bindings[0].tool_invocation_id",
|
||||
"message": "tool_invocation_id does not exist"
|
||||
}
|
||||
]
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "连接池耗尽导致支付失败",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 99,
|
||||
"tool_name": "query_logs",
|
||||
"evidence_excerpt": "connection pool exhausted"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "重新查询支付服务错误日志",
|
||||
"reason": "当前 Executor 证据引用无法回溯"
|
||||
}
|
||||
],
|
||||
"missing_info": ["需要有效的日志工具调用记录"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "连接池耗尽导致支付失败",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "unsupported",
|
||||
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "连接池耗尽导致支付失败",
|
||||
"verification": "unsupported",
|
||||
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "证据引用校验失败,不能确认根因。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "重新查询支付服务错误日志",
|
||||
"reason": "当前 Executor 证据引用无法回溯"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "payment failed without matching connection pool exhaustion entry",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,151 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-unsupported-claim-filtering",
|
||||
"query": "订单超时是否可以确认由数据库主库故障导致?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 44000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.42,
|
||||
"critical_fact_count": 2,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "订单接口出现超时",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs",
|
||||
"evidence_excerpt": "order api timeout"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-db-primary",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "数据库主库故障导致订单超时",
|
||||
"support_level": "weak",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs",
|
||||
"evidence_excerpt": "order api timeout"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询数据库主库状态和错误日志",
|
||||
"reason": "当前只有订单接口超时日志"
|
||||
}
|
||||
],
|
||||
"missing_info": ["数据库主库状态", "数据库错误日志"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"claim_text": "订单接口出现超时",
|
||||
"claim_type": "symptom",
|
||||
"verification": "direct_observation",
|
||||
"detail": "日志直接记录 order api timeout",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-db-primary",
|
||||
"claim_text": "数据库主库故障导致订单超时",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "unsupported",
|
||||
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "订单接口出现超时",
|
||||
"verification": "direct_evidence",
|
||||
"detail": "日志直接记录 order api timeout",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"fact": "数据库主库故障导致订单超时",
|
||||
"verification": "unsupported",
|
||||
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_logs"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "日志显示订单接口超时,但数据库方向证据不足。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询数据库主库状态和错误日志",
|
||||
"reason": "当前只有订单接口超时日志"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-unsupported-claim-filtering",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "order api timeout",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -1,13 +1,14 @@
|
||||
{
|
||||
"totalCases" : 5,
|
||||
"passedCases" : 5,
|
||||
"totalCases" : 8,
|
||||
"passedCases" : 8,
|
||||
"passRate" : 1.0,
|
||||
"verdictDistribution" : {
|
||||
"PASS" : 2,
|
||||
"LOW_CONFID" : 3
|
||||
"LOW_CONFID" : 5,
|
||||
"REJECT" : 1
|
||||
},
|
||||
"averageToolCallCount" : 2.0,
|
||||
"averageDurationMs" : 45800.0,
|
||||
"averageToolCallCount" : 1.625,
|
||||
"averageDurationMs" : 44875.0,
|
||||
"results" : [ {
|
||||
"caseId" : "payment-timeout",
|
||||
"title" : "Payment API timeout",
|
||||
@@ -21,6 +22,9 @@
|
||||
"query_logs" : true,
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 3,
|
||||
"durationMs" : 42000
|
||||
}, {
|
||||
@@ -35,6 +39,9 @@
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 51000
|
||||
}, {
|
||||
@@ -48,6 +55,9 @@
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 36000
|
||||
}, {
|
||||
@@ -62,6 +72,9 @@
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 47000
|
||||
}, {
|
||||
@@ -76,7 +89,58 @@
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 53000
|
||||
}, {
|
||||
"caseId" : "gatekeeper-fabricated-invocation",
|
||||
"title" : "Gatekeeper fabricated invocation",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "REJECT",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "fail",
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 1,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 39000
|
||||
}, {
|
||||
"caseId" : "unsupported-claim-filtering",
|
||||
"title" : "Unsupported claim filtering",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 2,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 44000
|
||||
}, {
|
||||
"caseId" : "composer-fallback-no-raw-json",
|
||||
"title" : "Composer fallback no raw JSON",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"composerStatus" : "composer_malformed",
|
||||
"claimCheckCount" : 2,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 47000
|
||||
} ]
|
||||
}
|
||||
|
||||
@@ -1,22 +1,26 @@
|
||||
# Diagnosis Eval Report
|
||||
|
||||
- Total cases: 5
|
||||
- Passed cases: 5
|
||||
- Total cases: 8
|
||||
- Passed cases: 8
|
||||
- Pass rate: 100.00%
|
||||
- Average tool calls: 2.00
|
||||
- Average duration ms: 45800.00
|
||||
- Average tool calls: 1.63
|
||||
- Average duration ms: 44875.00
|
||||
|
||||
## Verdict Distribution
|
||||
|
||||
- PASS: 2
|
||||
- LOW_CONFID: 3
|
||||
- LOW_CONFID: 5
|
||||
- REJECT: 1
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
|
||||
| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | - | - | - | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | - | - | - | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | - | - | - | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 53000 | - |
|
||||
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | valid | 1 | 3/3 | 1 | 39000 | - |
|
||||
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | valid | 2 | 2/2 | 1 | 44000 | - |
|
||||
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
|
||||
|
||||
+94
-147
@@ -1,203 +1,150 @@
|
||||
# Diagnosis Eval Data Schema
|
||||
|
||||
这份文档记录评测基准里的数据结构。口语化理解就是:
|
||||
这份文档记录 `mvp/eval` 固定评测集的数据结构。评测器读取保存好的 trace fixture,用确定性规则判断这次 Agent 运行是否满足预期。
|
||||
|
||||
```text
|
||||
用例文件说“我要考什么”
|
||||
trace 文件说“Agent 实际做了什么”
|
||||
评测结果说“这次有没有跑偏”
|
||||
汇总报告说“整体稳定性怎么样”
|
||||
case 文件:我要考什么
|
||||
fixture 文件:Agent 实际做了什么
|
||||
评测结果:这条 case 是否通过,哪里失败
|
||||
baseline report:整套固定集当前认可的结果
|
||||
```
|
||||
|
||||
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。
|
||||
当前评测不调用 LLM 打分。
|
||||
|
||||
## 1. 用例定义
|
||||
## 1. Case 定义
|
||||
|
||||
文件:`mvp/eval/cases/diagnosis-cases.json`
|
||||
|
||||
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。
|
||||
每条 case 定义一个固定诊断场景。
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"traceFixture": "payment-timeout-pass.json",
|
||||
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
|
||||
"id": "unsupported-claim-filtering",
|
||||
"title": "Unsupported claim filtering",
|
||||
"question": "订单超时是否可以确认由数据库主库故障导致?",
|
||||
"traceFixture": "unsupported-claim-filtering-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
|
||||
"allowedVerdicts": ["PASS", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["无证据确定"]
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||
}
|
||||
```
|
||||
|
||||
字段说明:
|
||||
|
||||
| 字段 | 意思 | 评测器怎么用 |
|
||||
| 字段 | 含义 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 |
|
||||
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 |
|
||||
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 |
|
||||
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 |
|
||||
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 |
|
||||
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 |
|
||||
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 |
|
||||
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 |
|
||||
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 |
|
||||
| `id` | case 唯一标识 | 出现在报告中 |
|
||||
| `title` | 可读标题 | 出现在报告中 |
|
||||
| `question` | 原始用户问题 | 用于说明场景,fixture 模式不会真实发送给 Agent |
|
||||
| `traceFixture` | 对应 fixture 文件名 | 从 `mvp/eval/fixtures` 加载 |
|
||||
| `expectedRootCauseKeywords` | 最终答案应覆盖的关键词 | 在 `session.answer` 中做包含判断 |
|
||||
| `minKeywordMatches` | 最少命中关键词数 | 低于该值则失败 |
|
||||
| `requiredEvidenceTools` | 必须出现的证据工具 | 从 `toolInvocations` 和 `tool_trace_summary` 中收集 |
|
||||
| `allowedVerdicts` | 允许的 Verifier verdict | verdict 不在列表中则失败 |
|
||||
| `forbiddenAnswerKeywords` | 最终答案禁止出现的词 | 用于拦截过度自信或危险表达 |
|
||||
| `requireV2AuditClosure` | 是否要求 V2 审计闭环字段 | 要求 `gatekeeper_result`、`claim_checks`、`composer_output` 存在,并检查 raw JSON 泄漏 |
|
||||
| `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 |
|
||||
| `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` |
|
||||
| `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 |
|
||||
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
|
||||
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
|
||||
|
||||
## 2. Trace Fixture
|
||||
|
||||
目录:`mvp/eval/fixtures/*.json`
|
||||
|
||||
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。
|
||||
fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规则需要的字段。
|
||||
|
||||
当前会读取这些字段:
|
||||
|
||||
| Trace 字段 | 意思 | 评测器怎么用 |
|
||||
| Trace 字段 | 含义 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 |
|
||||
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 |
|
||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 |
|
||||
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor 结构化输出里的确认事实证据绑定 | 如果 trace 中存在结构化 Executor 输出,每条 confirmed claim 必须有证据绑定 |
|
||||
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 |
|
||||
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 |
|
||||
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 |
|
||||
|
||||
简单说,trace 里最重要的是三类信息:
|
||||
|
||||
```text
|
||||
最终回答:它说了什么
|
||||
工具证据:它查了什么
|
||||
Verifier:它自己有没有承认这个结论可靠
|
||||
```
|
||||
| `session.answer` | 最终用户答案 | 检查关键词、禁用词、unsupported claim 泄漏、raw JSON 泄漏 |
|
||||
| `session.totalDurationMs` | 运行耗时 | 进入报告 |
|
||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
|
||||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
|
||||
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
|
||||
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
|
||||
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
|
||||
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 看到的工具证据 | 用于补充证据工具覆盖 |
|
||||
| `toolInvocations[*].toolName` | 实际调用工具名 | 用于检查 `requiredEvidenceTools` |
|
||||
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 中 |
|
||||
|
||||
## 3. 单条评测结果
|
||||
|
||||
Java 类型:`DiagnosisEvalResult`
|
||||
|
||||
这是每条 case 跑完之后的判断结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| 字段 | 含义 |
|
||||
| --- | --- |
|
||||
| `caseId` | 对应的 case id |
|
||||
| `caseId` | 对应 case id |
|
||||
| `title` | case 标题 |
|
||||
| `passed` | 这条 case 是否通过 |
|
||||
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 |
|
||||
| `verdict` | 从 trace 里读出来的 Verifier verdict |
|
||||
| `matchedKeywordCount` | 最终回答命中的关键词数量 |
|
||||
| `requiredKeywordCount` | case 定义里一共有多少个关键词 |
|
||||
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` |
|
||||
| `toolCallCount` | 本次 trace 里工具调用总数 |
|
||||
| `durationMs` | 本次 trace 的耗时 |
|
||||
|
||||
判断通过的口语化规则:
|
||||
|
||||
```text
|
||||
回答要说到关键点
|
||||
该查的证据工具要查到
|
||||
Verifier 的结论要在可接受范围内
|
||||
回答不能出现危险的过度自信表达
|
||||
如果是 REJECT,就必须走降级模板
|
||||
```
|
||||
| `passed` | 该 case 是否通过 |
|
||||
| `failedChecks` | 失败原因列表 |
|
||||
| `verdict` | 从 trace 中读到的 Verifier verdict |
|
||||
| `matchedKeywordCount` | 最终答案命中的关键词数量 |
|
||||
| `requiredKeywordCount` | case 配置的关键词数量 |
|
||||
| `evidenceCoverage` | 每个必需工具是否出现 |
|
||||
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
|
||||
| `composerStatus` | 读到的 `composer_output.status` |
|
||||
| `claimCheckCount` | `claim_checks` 数量 |
|
||||
| `toolCallCount` | trace 中工具调用总数 |
|
||||
| `durationMs` | trace 总耗时 |
|
||||
|
||||
## 4. 汇总报告
|
||||
|
||||
Java 类型:`DiagnosisEvalReport`
|
||||
|
||||
这是整个基准集跑完之后的总结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| 字段 | 含义 |
|
||||
| --- | --- |
|
||||
| `totalCases` | 总共评测了多少条 case |
|
||||
| `passedCases` | 通过了多少条 |
|
||||
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` |
|
||||
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` |
|
||||
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 |
|
||||
| `totalCases` | case 总数 |
|
||||
| `passedCases` | 通过数 |
|
||||
| `passRate` | 通过率,范围 `0.0` 到 `1.0` |
|
||||
| `verdictDistribution` | Verifier verdict 分布 |
|
||||
| `averageToolCallCount` | 平均工具调用数 |
|
||||
| `averageDurationMs` | 平均耗时 |
|
||||
| `results` | 每条 case 的详细结果列表 |
|
||||
| `results` | 单条 case 结果列表 |
|
||||
|
||||
## 5. 怎么看这个基准
|
||||
## 5. Stage 5 V2 审计闭环规则
|
||||
|
||||
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答:
|
||||
阶段 5 关注的是“前四段链路是否能被固定评测证明”:
|
||||
|
||||
```text
|
||||
以前能过的诊断题,现在还过不过?
|
||||
它是不是少查了某些证据?
|
||||
它是不是变得更自信但证据不足?
|
||||
它是不是开始输出不该说的话?
|
||||
它是不是明显变慢了?
|
||||
Executor structured output
|
||||
-> Gatekeeper deterministic audit
|
||||
-> Verifier claim_checks
|
||||
-> Composer filtered final answer
|
||||
```
|
||||
|
||||
所以面试里可以这样讲:
|
||||
新增确定性规则:
|
||||
|
||||
```text
|
||||
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。
|
||||
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。
|
||||
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
|
||||
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
|
||||
```
|
||||
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
|
||||
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
|
||||
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
|
||||
- Composer 输出必须记录 `status`。
|
||||
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
|
||||
- case 配置的 `forbiddenConfirmedClaimKeywords` 不能出现在最终答案里。
|
||||
|
||||
## 6. Baseline Diff
|
||||
|
||||
Baseline diff 是拿两份 report 做对比:
|
||||
Baseline diff 比较两份 report:
|
||||
|
||||
```text
|
||||
baseline report:以前认可的基准结果
|
||||
current report:这次改动后跑出来的新结果
|
||||
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
|
||||
baseline report:已经认可的基准结果
|
||||
current report:当前代码/fixture 跑出的结果
|
||||
diff report:结构化列出退化、改善和普通变化
|
||||
```
|
||||
|
||||
Java 类型:
|
||||
主要退化信号:
|
||||
|
||||
- `DiagnosisEvalDiffReport`
|
||||
- `DiagnosisEvalDiffItem`
|
||||
|
||||
`DiagnosisEvalDiffReport` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `baselineTotalCases` | baseline 里有多少条 case |
|
||||
| `currentTotalCases` | current 里有多少条 case |
|
||||
| `baselinePassedCases` | baseline 通过了多少条 |
|
||||
| `currentPassedCases` | current 通过了多少条 |
|
||||
| `baselinePassRate` | baseline 通过率 |
|
||||
| `currentPassRate` | current 通过率 |
|
||||
| `regressionCount` | 退化项数量 |
|
||||
| `improvementCount` | 改善项数量 |
|
||||
| `changedCount` | 普通变化项数量 |
|
||||
| `hasRegression` | 是否存在退化 |
|
||||
| `items` | 具体 diff 明细 |
|
||||
|
||||
`DiagnosisEvalDiffItem` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
|
||||
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
|
||||
| `caseId` | 如果是单条 case 变化,这里记录 case id |
|
||||
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
|
||||
| `baselineValue` | baseline 里的值 |
|
||||
| `currentValue` | current 里的值 |
|
||||
| `delta` | 数值变化量;非数值变化为空 |
|
||||
| `message` | 给人看的变化说明 |
|
||||
|
||||
口语化判断规则:
|
||||
|
||||
```text
|
||||
pass rate 下降:退化
|
||||
case 从通过变失败:退化
|
||||
证据工具从有变没有:退化
|
||||
关键词命中变少:退化
|
||||
工具调用或耗时升高:成本上升,记为退化信号
|
||||
verdict 分布变化:记录变化,供人工判断是否符合预期
|
||||
```
|
||||
|
||||
面试里可以这样讲:
|
||||
|
||||
```text
|
||||
我把 baseline report 和当前 report 做结构化 diff。
|
||||
它不是再问 LLM,而是用代码比较固定字段。
|
||||
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
|
||||
diff 会直接标成 regression。
|
||||
这样 Agent 改动可以用固定基准做回归判断。
|
||||
```
|
||||
- pass rate 下降。
|
||||
- case 从通过变失败。
|
||||
- 必需证据工具从有变无。
|
||||
- 关键词命中减少。
|
||||
- 工具调用或耗时明显上升。
|
||||
- verdict 分布变化。
|
||||
|
||||
Reference in New Issue
Block a user