Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0
# Conflicts: # mvp/issues/README.md
This commit is contained in:
@@ -2,6 +2,13 @@
|
||||
|
||||
This demo proves the MVP flow from user question to persisted diagnosis trace.
|
||||
|
||||
For interview use, start with:
|
||||
|
||||
- `interview-walkthrough.md` for the talk track
|
||||
- `trace-inspection-checklist.md` for fields to inspect
|
||||
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
|
||||
- `requests/payment-timeout-chat.json` for the fixed request payload
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
|
||||
@@ -22,6 +29,22 @@ http://localhost:9900
|
||||
|
||||
## 1. Run Chat Diagnosis
|
||||
|
||||
Fast path:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
This writes:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
Manual path:
|
||||
|
||||
```powershell
|
||||
$sessionId = "mvp-demo-payment-timeout-001"
|
||||
$body = @{
|
||||
|
||||
@@ -0,0 +1,146 @@
|
||||
# Interview Walkthrough: MVP Diagnosis Agent
|
||||
|
||||
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
|
||||
|
||||
## 30-Second Summary
|
||||
|
||||
```text
|
||||
This is an enterprise diagnosis Agent MVP.
|
||||
It takes a payment-timeout question, plans the investigation, calls evidence tools,
|
||||
checks the answer through a verifier, persists the full trace, and accepts feedback.
|
||||
```
|
||||
|
||||
The important claim is not "the model answered once." The claim is:
|
||||
|
||||
```text
|
||||
The system can show what evidence was used, how the answer was checked, and how to replay the session.
|
||||
```
|
||||
|
||||
## Demo Flow
|
||||
|
||||
1. Start the service with the `mvp-demo` profile.
|
||||
2. Run the fixed payment-timeout request.
|
||||
3. Open `mvp/demo/output/chat-response.json`.
|
||||
4. Open `mvp/demo/output/trace-response.json`.
|
||||
5. Point to evidence tools and verifier evaluation.
|
||||
6. Submit feedback and show it is attached to the same session.
|
||||
|
||||
## Commands
|
||||
|
||||
Start service:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
Run the demo from another terminal:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
Optional custom session:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
||||
```
|
||||
|
||||
## What To Show
|
||||
|
||||
### 1. User-Facing Answer
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
|
||||
```
|
||||
|
||||
### 2. Evidence Trace
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/trace-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the important Agent engineering part.
|
||||
I can inspect which tools were called, what inputs they received,
|
||||
whether they succeeded, and what evidence preview was persisted.
|
||||
```
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.toolInvocations[*].toolName`
|
||||
- `data.toolInvocations[*].inputParams`
|
||||
- `data.toolInvocations[*].outputPreview`
|
||||
- `data.toolInvocations[*].success`
|
||||
|
||||
### 3. Verifier / Self-Evaluation
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.session.selfEvaluation`
|
||||
- `data.summary.hasVerifierEvaluation`
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
The final answer is not just raw Executor output.
|
||||
It is checked by a verifier or self-evaluation layer using the persisted trace.
|
||||
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
|
||||
```
|
||||
|
||||
### 4. Feedback Loop
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
Then re-query trace if needed.
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
Feedback is attached to the same diagnosis session.
|
||||
That makes it possible to mine useful / not useful cases later.
|
||||
```
|
||||
|
||||
### 5. Regression Story
|
||||
|
||||
Mention, do not deep dive unless asked:
|
||||
|
||||
```text
|
||||
For repeatability, I also built an offline eval baseline.
|
||||
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
|
||||
The two are separate on purpose: demo for human review, eval for automated signal.
|
||||
```
|
||||
|
||||
## Strong Interview Framing
|
||||
|
||||
Use this phrasing:
|
||||
|
||||
```text
|
||||
I focused on the Agent engineering surface:
|
||||
traceability, evidence persistence, verifier gating, feedback, and regression checks.
|
||||
The model answer is only one part of the system.
|
||||
The more important part is whether we can audit and improve the answer after it is produced.
|
||||
```
|
||||
|
||||
## Known Limits To Say Proactively
|
||||
|
||||
```text
|
||||
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
|
||||
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
|
||||
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
|
||||
```
|
||||
@@ -0,0 +1,11 @@
|
||||
# Demo Output
|
||||
|
||||
This directory is the default output location for local demo responses.
|
||||
|
||||
Generated files are intentionally ignored by Git:
|
||||
|
||||
- `chat-response.json`
|
||||
- `trace-response.json`
|
||||
- `feedback-response.json`
|
||||
|
||||
Keep this README so the directory exists in the repository.
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-payment-timeout-001",
|
||||
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
param(
|
||||
[string]$BaseUrl = "http://localhost:9900",
|
||||
[string]$SessionId = "mvp-demo-payment-timeout-001",
|
||||
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
|
||||
[string]$OutputDir = "$PSScriptRoot/../output"
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||
|
||||
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||
$request.Id = $SessionId
|
||||
$body = $request | ConvertTo-Json -Depth 8
|
||||
|
||||
Write-Host "Running payment-timeout chat demo..."
|
||||
Write-Host "BaseUrl: $BaseUrl"
|
||||
Write-Host "SessionId: $SessionId"
|
||||
|
||||
$chat = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "$BaseUrl/api/chat" `
|
||||
-ContentType "application/json; charset=utf-8" `
|
||||
-Body $body
|
||||
|
||||
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||
Write-Host "Saved chat response: $OutputDir/chat-response.json"
|
||||
|
||||
$trace = Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||
|
||||
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||
Write-Host "Saved trace response: $OutputDir/trace-response.json"
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
$feedback = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "$BaseUrl/api/feedback" `
|
||||
-ContentType "application/json; charset=utf-8" `
|
||||
-Body $feedbackBody
|
||||
|
||||
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
|
||||
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "Demo completed. Review:"
|
||||
Write-Host "- mvp/demo/output/chat-response.json"
|
||||
Write-Host "- mvp/demo/output/trace-response.json"
|
||||
Write-Host "- mvp/demo/output/feedback-response.json"
|
||||
@@ -0,0 +1,52 @@
|
||||
# Trace Inspection Checklist
|
||||
|
||||
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
|
||||
|
||||
## Session
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
|
||||
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
|
||||
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
|
||||
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
|
||||
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
|
||||
|
||||
## Agent Steps
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
|
||||
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
|
||||
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
|
||||
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
|
||||
|
||||
## Tool Evidence
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
|
||||
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
|
||||
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
|
||||
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
|
||||
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
|
||||
|
||||
## Summary
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
|
||||
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
|
||||
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
|
||||
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
|
||||
|
||||
## What Good Looks Like
|
||||
|
||||
```text
|
||||
same session id
|
||||
-> final answer
|
||||
-> persisted agent steps
|
||||
-> persisted evidence tool calls
|
||||
-> verifier/self-evaluation
|
||||
-> feedback attached to the same session
|
||||
```
|
||||
@@ -0,0 +1,61 @@
|
||||
# Diagnosis Eval Harness
|
||||
|
||||
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
|
||||
|
||||
## Scope
|
||||
|
||||
- Case definitions: `cases/diagnosis-cases.json`
|
||||
- Offline trace fixtures: `fixtures/*.json`
|
||||
- Field definitions: `schema.md`
|
||||
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
|
||||
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
|
||||
- Evaluator implementation: `DiagnosisTraceEvaluator`
|
||||
- Report writer: `DiagnosisEvalReportWriter`
|
||||
|
||||
## Current Mode
|
||||
|
||||
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
## Verification
|
||||
|
||||
Run the focused evaluator test:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||
```
|
||||
|
||||
The committed baseline report represents the current fixed fixture set:
|
||||
|
||||
```text
|
||||
5 fixed cases
|
||||
5 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
3 LOW_CONFID verdicts
|
||||
```
|
||||
|
||||
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
||||
|
||||
## Interview Story
|
||||
|
||||
The harness gives the MVP a repeatable baseline:
|
||||
|
||||
```text
|
||||
fixed diagnosis case
|
||||
-> saved or runtime trace
|
||||
-> rule-based trace validation
|
||||
-> JSON / Markdown report
|
||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
||||
```
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||
|
||||
```text
|
||||
baseline report
|
||||
current report
|
||||
-> deterministic diff
|
||||
-> regressions, improvements, and changed signals
|
||||
```
|
||||
|
||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
||||
@@ -0,0 +1,57 @@
|
||||
[
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"traceFixture": "payment-timeout-pass.json",
|
||||
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
|
||||
"allowedVerdicts": ["PASS", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["无证据确定"]
|
||||
},
|
||||
{
|
||||
"id": "mysql-pool-exhausted",
|
||||
"title": "MySQL connection pool exhausted",
|
||||
"question": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
|
||||
"traceFixture": "mysql-pool-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["mysql", "连接池", "超时"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["已经完全确认"]
|
||||
},
|
||||
{
|
||||
"id": "redis-timeout",
|
||||
"title": "Redis timeout",
|
||||
"question": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
||||
"traceFixture": "redis-timeout-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["redis", "超时"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["无需进一步排查"]
|
||||
},
|
||||
{
|
||||
"id": "slow-response",
|
||||
"title": "Slow response",
|
||||
"question": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
||||
"traceFixture": "slow-response-pass.json",
|
||||
"expectedRootCauseKeywords": ["p99", "慢响应"],
|
||||
"minKeywordMatches": 1,
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["没有风险"]
|
||||
},
|
||||
{
|
||||
"id": "jvm-memory-risk",
|
||||
"title": "JVM memory risk",
|
||||
"question": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||
"traceFixture": "jvm-memory-risk-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["jvm", "内存", "oom"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["可以忽略"]
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 53000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n订单服务存在 JVM 内存风险,但还不能完全确认会发生 OOM。指标显示 heap 使用率持续高于 88%,日志出现多次 Full GC 和 allocation pressure,需要继续观察对象增长来源并检查最近发布。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.52,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "indirect"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,53 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 51000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nMySQL 连接池可能参与了本次超时问题。日志中出现 connection pool exhausted,但当前缺少完整指标证据,因此只能作为低置信结论处理。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.48,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "lookup_knowledge",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"toolName": "lookup_knowledge",
|
||||
"success": true,
|
||||
"relevanceLevel": "PRECISE"
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,64 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 42000,
|
||||
"toolCallCount": 3,
|
||||
"answer": "支付接口超时与连接池等待有关。知识库说明支付超时需要同时检查连接池、日志和指标;日志出现 connection pool exhausted;指标显示支付服务延迟升高。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.86,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "lookup_knowledge",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"toolName": "lookup_knowledge",
|
||||
"success": true,
|
||||
"relevanceLevel": "PRECISE"
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 3,
|
||||
"returnedToolCallCount": 3,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,41 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-redis-timeout",
|
||||
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 36000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nRedis 连接超时可能和支付服务到 Redis 的网络抖动或连接池等待有关。日志中出现 redis timeout 和 command timeout 记录,但当前缺少指标侧证据,因此只能作为低置信结论处理。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.46,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-redis-timeout",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 2,
|
||||
"returnedStepCount": 2,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-slow-response",
|
||||
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 47000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "用户服务 P99 升高主要表现为慢响应。指标显示 P99 latency 从 280ms 上升到 1800ms,日志中同时出现 slow request 和 downstream timeout,因此优先排查下游依赖耗时和线程池排队。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.78,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-slow-response",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-slow-response",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,85 @@
|
||||
{
|
||||
"baselineTotalCases" : 5,
|
||||
"currentTotalCases" : 5,
|
||||
"baselinePassedCases" : 5,
|
||||
"currentPassedCases" : 4,
|
||||
"baselinePassRate" : 1.0,
|
||||
"currentPassRate" : 0.8,
|
||||
"regressionCount" : 6,
|
||||
"improvementCount" : 0,
|
||||
"changedCount" : 2,
|
||||
"hasRegression" : true,
|
||||
"items" : [ {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "passRate",
|
||||
"baselineValue" : "1.0",
|
||||
"currentValue" : "0.8",
|
||||
"delta" : -0.19999999999999996,
|
||||
"message" : "passRate changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "averageToolCallCount",
|
||||
"baselineValue" : "2.0",
|
||||
"currentValue" : "3.0",
|
||||
"delta" : 1.0,
|
||||
"message" : "averageToolCallCount changed"
|
||||
}, {
|
||||
"type" : "CHANGED",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "verdictDistribution.LOW_CONFID",
|
||||
"baselineValue" : "3",
|
||||
"currentValue" : "2",
|
||||
"delta" : -1.0,
|
||||
"message" : "verdict count changed for LOW_CONFID"
|
||||
}, {
|
||||
"type" : "CHANGED",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "verdictDistribution.REJECT",
|
||||
"baselineValue" : "0",
|
||||
"currentValue" : "1",
|
||||
"delta" : 1.0,
|
||||
"message" : "verdict count changed for REJECT"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "passed",
|
||||
"baselineValue" : "true",
|
||||
"currentValue" : "false",
|
||||
"delta" : null,
|
||||
"message" : "redis-timeout pass state changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "verdict",
|
||||
"baselineValue" : "LOW_CONFID",
|
||||
"currentValue" : "REJECT",
|
||||
"delta" : -1.0,
|
||||
"message" : "redis-timeout verdict changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "matchedKeywordCount",
|
||||
"baselineValue" : "2",
|
||||
"currentValue" : "1",
|
||||
"delta" : -1.0,
|
||||
"message" : "redis-timeout matchedKeywordCount changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "evidenceCoverage.query_logs",
|
||||
"baselineValue" : "true",
|
||||
"currentValue" : "false",
|
||||
"delta" : null,
|
||||
"message" : "redis-timeout evidence coverage changed for query_logs"
|
||||
} ]
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
# Diagnosis Eval Baseline Diff
|
||||
|
||||
- Baseline pass rate: 100.00%
|
||||
- Current pass rate: 80.00%
|
||||
- Baseline passed cases: 5/5
|
||||
- Current passed cases: 4/5
|
||||
- Regressions: 6
|
||||
- Improvements: 0
|
||||
- Other changes: 2
|
||||
|
||||
## Diff Items
|
||||
|
||||
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
|
||||
| --- | --- | --- | --- | --- | --- | ---: | --- |
|
||||
| REGRESSION | aggregate | - | passRate | 1.0 | 0.8 | -0.200 | passRate changed |
|
||||
| REGRESSION | aggregate | - | averageToolCallCount | 2.0 | 3.0 | 1.000 | averageToolCallCount changed |
|
||||
| CHANGED | aggregate | - | verdictDistribution.LOW_CONFID | 3 | 2 | -1.000 | verdict count changed for LOW_CONFID |
|
||||
| CHANGED | aggregate | - | verdictDistribution.REJECT | 0 | 1 | 1.000 | verdict count changed for REJECT |
|
||||
| REGRESSION | case | redis-timeout | passed | true | false | - | redis-timeout pass state changed |
|
||||
| REGRESSION | case | redis-timeout | verdict | LOW_CONFID | REJECT | -1.000 | redis-timeout verdict changed |
|
||||
| REGRESSION | case | redis-timeout | matchedKeywordCount | 2 | 1 | -1.000 | redis-timeout matchedKeywordCount changed |
|
||||
| REGRESSION | case | redis-timeout | evidenceCoverage.query_logs | true | false | - | redis-timeout evidence coverage changed for query_logs |
|
||||
@@ -0,0 +1,82 @@
|
||||
{
|
||||
"totalCases" : 5,
|
||||
"passedCases" : 5,
|
||||
"passRate" : 1.0,
|
||||
"verdictDistribution" : {
|
||||
"PASS" : 2,
|
||||
"LOW_CONFID" : 3
|
||||
},
|
||||
"averageToolCallCount" : 2.0,
|
||||
"averageDurationMs" : 45800.0,
|
||||
"results" : [ {
|
||||
"caseId" : "payment-timeout",
|
||||
"title" : "Payment API timeout",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true,
|
||||
"query_metrics" : true
|
||||
},
|
||||
"toolCallCount" : 3,
|
||||
"durationMs" : 42000
|
||||
}, {
|
||||
"caseId" : "mysql-pool-exhausted",
|
||||
"title" : "MySQL connection pool exhausted",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 51000
|
||||
}, {
|
||||
"caseId" : "redis-timeout",
|
||||
"title" : "Redis timeout",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 36000
|
||||
}, {
|
||||
"caseId" : "slow-response",
|
||||
"title" : "Slow response",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 47000
|
||||
}, {
|
||||
"caseId" : "jvm-memory-risk",
|
||||
"title" : "JVM memory risk",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 53000
|
||||
} ]
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
# Diagnosis Eval Report
|
||||
|
||||
- Total cases: 5
|
||||
- Passed cases: 5
|
||||
- Pass rate: 100.00%
|
||||
- Average tool calls: 2.00
|
||||
- Average duration ms: 45800.00
|
||||
|
||||
## Verdict Distribution
|
||||
|
||||
- PASS: 2
|
||||
- LOW_CONFID: 3
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
|
||||
@@ -0,0 +1,202 @@
|
||||
# Diagnosis Eval Data Schema
|
||||
|
||||
这份文档记录评测基准里的数据结构。口语化理解就是:
|
||||
|
||||
```text
|
||||
用例文件说“我要考什么”
|
||||
trace 文件说“Agent 实际做了什么”
|
||||
评测结果说“这次有没有跑偏”
|
||||
汇总报告说“整体稳定性怎么样”
|
||||
```
|
||||
|
||||
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。
|
||||
|
||||
## 1. 用例定义
|
||||
|
||||
文件:`mvp/eval/cases/diagnosis-cases.json`
|
||||
|
||||
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"traceFixture": "payment-timeout-pass.json",
|
||||
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
|
||||
"allowedVerdicts": ["PASS", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["无证据确定"]
|
||||
}
|
||||
```
|
||||
|
||||
字段说明:
|
||||
|
||||
| 字段 | 意思 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 |
|
||||
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 |
|
||||
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 |
|
||||
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 |
|
||||
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 |
|
||||
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 |
|
||||
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 |
|
||||
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 |
|
||||
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 |
|
||||
|
||||
## 2. Trace Fixture
|
||||
|
||||
目录:`mvp/eval/fixtures/*.json`
|
||||
|
||||
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。
|
||||
|
||||
当前会读取这些字段:
|
||||
|
||||
| Trace 字段 | 意思 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 |
|
||||
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 |
|
||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 |
|
||||
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 |
|
||||
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 |
|
||||
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 |
|
||||
|
||||
简单说,trace 里最重要的是三类信息:
|
||||
|
||||
```text
|
||||
最终回答:它说了什么
|
||||
工具证据:它查了什么
|
||||
Verifier:它自己有没有承认这个结论可靠
|
||||
```
|
||||
|
||||
## 3. 单条评测结果
|
||||
|
||||
Java 类型:`DiagnosisEvalResult`
|
||||
|
||||
这是每条 case 跑完之后的判断结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `caseId` | 对应的 case id |
|
||||
| `title` | case 标题 |
|
||||
| `passed` | 这条 case 是否通过 |
|
||||
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 |
|
||||
| `verdict` | 从 trace 里读出来的 Verifier verdict |
|
||||
| `matchedKeywordCount` | 最终回答命中的关键词数量 |
|
||||
| `requiredKeywordCount` | case 定义里一共有多少个关键词 |
|
||||
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` |
|
||||
| `toolCallCount` | 本次 trace 里工具调用总数 |
|
||||
| `durationMs` | 本次 trace 的耗时 |
|
||||
|
||||
判断通过的口语化规则:
|
||||
|
||||
```text
|
||||
回答要说到关键点
|
||||
该查的证据工具要查到
|
||||
Verifier 的结论要在可接受范围内
|
||||
回答不能出现危险的过度自信表达
|
||||
如果是 REJECT,就必须走降级模板
|
||||
```
|
||||
|
||||
## 4. 汇总报告
|
||||
|
||||
Java 类型:`DiagnosisEvalReport`
|
||||
|
||||
这是整个基准集跑完之后的总结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `totalCases` | 总共评测了多少条 case |
|
||||
| `passedCases` | 通过了多少条 |
|
||||
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` |
|
||||
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` |
|
||||
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 |
|
||||
| `averageDurationMs` | 平均耗时 |
|
||||
| `results` | 每条 case 的详细结果列表 |
|
||||
|
||||
## 5. 怎么看这个基准
|
||||
|
||||
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答:
|
||||
|
||||
```text
|
||||
以前能过的诊断题,现在还过不过?
|
||||
它是不是少查了某些证据?
|
||||
它是不是变得更自信但证据不足?
|
||||
它是不是开始输出不该说的话?
|
||||
它是不是明显变慢了?
|
||||
```
|
||||
|
||||
所以面试里可以这样讲:
|
||||
|
||||
```text
|
||||
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。
|
||||
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。
|
||||
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
|
||||
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
|
||||
```
|
||||
|
||||
## 6. Baseline Diff
|
||||
|
||||
Baseline diff 是拿两份 report 做对比:
|
||||
|
||||
```text
|
||||
baseline report:以前认可的基准结果
|
||||
current report:这次改动后跑出来的新结果
|
||||
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
|
||||
```
|
||||
|
||||
Java 类型:
|
||||
|
||||
- `DiagnosisEvalDiffReport`
|
||||
- `DiagnosisEvalDiffItem`
|
||||
|
||||
`DiagnosisEvalDiffReport` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `baselineTotalCases` | baseline 里有多少条 case |
|
||||
| `currentTotalCases` | current 里有多少条 case |
|
||||
| `baselinePassedCases` | baseline 通过了多少条 |
|
||||
| `currentPassedCases` | current 通过了多少条 |
|
||||
| `baselinePassRate` | baseline 通过率 |
|
||||
| `currentPassRate` | current 通过率 |
|
||||
| `regressionCount` | 退化项数量 |
|
||||
| `improvementCount` | 改善项数量 |
|
||||
| `changedCount` | 普通变化项数量 |
|
||||
| `hasRegression` | 是否存在退化 |
|
||||
| `items` | 具体 diff 明细 |
|
||||
|
||||
`DiagnosisEvalDiffItem` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
|
||||
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
|
||||
| `caseId` | 如果是单条 case 变化,这里记录 case id |
|
||||
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
|
||||
| `baselineValue` | baseline 里的值 |
|
||||
| `currentValue` | current 里的值 |
|
||||
| `delta` | 数值变化量;非数值变化为空 |
|
||||
| `message` | 给人看的变化说明 |
|
||||
|
||||
口语化判断规则:
|
||||
|
||||
```text
|
||||
pass rate 下降:退化
|
||||
case 从通过变失败:退化
|
||||
证据工具从有变没有:退化
|
||||
关键词命中变少:退化
|
||||
工具调用或耗时升高:成本上升,记为退化信号
|
||||
verdict 分布变化:记录变化,供人工判断是否符合预期
|
||||
```
|
||||
|
||||
面试里可以这样讲:
|
||||
|
||||
```text
|
||||
我把 baseline report 和当前 report 做结构化 diff。
|
||||
它不是再问 LLM,而是用代码比较固定字段。
|
||||
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
|
||||
diff 会直接标成 regression。
|
||||
这样 Agent 改动可以用固定基准做回归判断。
|
||||
```
|
||||
@@ -0,0 +1,137 @@
|
||||
# ISS-005 证据链补齐与降级契约收敛
|
||||
|
||||
**状态**:进行中(sm-flow)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-A 面试打磨项 / 基于 ISS-003 的当前实现复核
|
||||
**关联**:ISS-003(Verifier 证据链、失败路径可验证性)、`chat-verifier-agent`、`mvp-demo-trace-acceptance`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前 MVP 已具备:
|
||||
|
||||
- `lookup_knowledge`、`query_logs`、`query_metrics` 的工具调用落库
|
||||
- Verifier 基于 `tool_trace_summary` 做事实核查
|
||||
- `LOW_CONFID` / `REJECT` 的用户侧降级输出
|
||||
- trace API 可回放 session、agent_step、tool_invocation 和 self_evaluation
|
||||
|
||||
但如果目标是拿这个项目去面试 Agent 工程师,当前实现仍有一个明显短板:
|
||||
|
||||
**证据链已经“有了”,但还没有被收敛成清晰、稳定、可测试的工程契约。**
|
||||
|
||||
这会直接影响三个面试问题的回答质量:
|
||||
|
||||
1. 工具失败时系统会怎样降级?
|
||||
2. Verifier 看到的 evidence 到底是否一致、可审计?
|
||||
3. 这些失败路径和降级行为有没有稳定测试,而不是只靠 runtime 演示?
|
||||
|
||||
---
|
||||
|
||||
## 当前现状复核
|
||||
|
||||
### 1. 工具落库入口已经存在,但契约不统一
|
||||
|
||||
- `QueryLogsTools` 和 `QueryMetricsTools` 通过 `ToolInvocationRecorder.recordEvidenceTool(...)` 记录 evidence tool 调用。
|
||||
- `LookupKnowledgeTool` 仍保留独立的 `saveToolInvocation(...)` 路径,自己构造 `ToolInvocation` 实体。
|
||||
|
||||
这意味着:
|
||||
|
||||
- evidence tool 的公共字段有一套约定
|
||||
- knowledge retrieval 又有一套定制字段拼装
|
||||
|
||||
两者都能工作,但**没有形成统一的“证据调用记录契约”**。
|
||||
|
||||
### 2. 失败 / 无结果 / 去重命中的语义不够显式
|
||||
|
||||
当前实现里:
|
||||
|
||||
- `query_logs` 未命中时会返回 `success=false` + `"未找到匹配的日志"`
|
||||
- `query_metrics` 失败时会返回 `success=false`
|
||||
- `lookup_knowledge` 去重命中时会返回 `found=false`,但 `tool_invocation.success=true`
|
||||
- `ToolTraceSummaryService` 通过 `success`、`relevanceLevel`、`dedupReason` 等字段做启发式摘要
|
||||
|
||||
这些行为在代码里是分散成立的,但**没有被定义成统一契约**,导致:
|
||||
|
||||
- Verifier 能看到的“失败”和“无证据”边界不够稳定
|
||||
- 评测时难以明确统计哪些是“调用失败”、哪些是“无命中”、哪些是“已检索过”
|
||||
|
||||
### 3. ChatService 的降级路径有实现,但测试矩阵不完整
|
||||
|
||||
`ChatService` 已处理:
|
||||
|
||||
- `verifier_output` 缺失或无法解析 → fallback `LOW_CONFID`
|
||||
- `REJECT` → degraded output
|
||||
- `LOW_CONFID` → disclaimer output
|
||||
|
||||
但目前缺少成体系的专项验证,尤其是:
|
||||
|
||||
- Verifier 输出非法 JSON
|
||||
- evidence tool 查询失败
|
||||
- knowledge lookup 无有效证据
|
||||
- fallback 文案是否只基于 verifier 缺口拼装
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- **面试表达弱化**:你能讲“我有 trace”,但还不能很硬地讲“我的失败路径是有契约和测试保护的”。
|
||||
- **评测基础不稳**:后续 P1-B 做 case-based harness 时,统计口径会受 evidence 语义不一致影响。
|
||||
- **Verifier 可审计性打折**:当前实现可用,但 still relies on code convention,而不是一份明确收敛后的工程协议。
|
||||
|
||||
---
|
||||
|
||||
## 本 issue 目标
|
||||
|
||||
P1-A 只做三件事:
|
||||
|
||||
1. 收敛 evidence tool 的落库契约,让 `lookup_knowledge`、`query_logs`、`query_metrics` 的公共语义一致。
|
||||
2. 明确失败 / 无证据 / 去重 / verifier 非法输出等降级契约,让 `ToolTraceSummaryService` 和 `ChatService` 面向统一状态工作。
|
||||
3. 增加专项离线测试,覆盖证据摘要与关键降级路径。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- `ToolInvocationRecorder` 契约增强
|
||||
- `LookupKnowledgeTool` 入库路径收敛
|
||||
- `QueryLogsTools` / `QueryMetricsTools` evidence 语义对齐
|
||||
- `ToolTraceSummaryService` 对失败 / no-hit / mixed evidence 的摘要规则收敛
|
||||
- `ChatService` 对 verifier 非法输出与降级输出的专项测试
|
||||
- 与该 change 直接相关的文档、OpenSpec、devflow 记录
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不引入新的数据库表或 schema 变更
|
||||
- 不扩展新的 evidence tool
|
||||
- 不做 P1-B 评测集 / harness
|
||||
- 不做前端 trace UI
|
||||
- 不处理敏感配置和默认 `mvn test` 离线化
|
||||
|
||||
---
|
||||
|
||||
## 预期结果
|
||||
|
||||
完成后,项目在面试里应能更清楚地表述为:
|
||||
|
||||
```text
|
||||
我不仅把 Agent 的工具调用落到了库里,
|
||||
还把 evidence trace、失败语义和 verifier 降级路径收敛成了稳定契约,
|
||||
并用离线测试覆盖了这些关键失败场景。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
|
||||
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`
|
||||
@@ -0,0 +1,99 @@
|
||||
# ISS-006 固定诊断评测集与回归 Harness
|
||||
|
||||
**状态**:进行中(sm-flow)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-B 面试打磨项
|
||||
**依赖**:ISS-005 / `evidence-trace-hardening`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
MVP 已经具备可追溯证据链、Verifier 质量门禁、trace API 和固定 demo 流程。上一阶段 `evidence-trace-hardening` 进一步统一了 evidence tool 的状态语义,让系统能稳定区分:
|
||||
|
||||
- `supported`
|
||||
- `no_evidence`
|
||||
- `deduped`
|
||||
- `failed`
|
||||
|
||||
下一步需要证明 Agent 在一组固定诊断场景下的表现,而不是只依赖单次 demo。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前项目能演示一次支付超时诊断,但还缺少稳定的评测基线:
|
||||
|
||||
- 每次改 prompt、工具、Verifier 或检索逻辑后,无法快速判断是否退化。
|
||||
- 只能人工看 trace,缺少结构化通过 / 失败结果。
|
||||
- 缺少面试时能展示的指标,如 evidence coverage、verdict 分布、工具调用数量和耗时。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
建立一个轻量的固定 case 评测 harness,用于验证 MVP Agent 的诊断质量和证据链完整性。
|
||||
|
||||
第一版不做 LLM-as-judge,优先做规则化校验:
|
||||
|
||||
- 固定 5 个 MVP 诊断 case
|
||||
- 每个 case 定义 expected root-cause keywords、required evidence tools、allowed verdicts
|
||||
- 基于 trace 结果校验 evidence coverage、verifier evaluation、tool invocation、final answer shape
|
||||
- 输出 JSON 和 Markdown 报告
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- 评测 case 定义文件
|
||||
- trace 规则校验器
|
||||
- eval runner 或测试入口
|
||||
- JSON / Markdown 报告输出
|
||||
- demo 文档和 devflow 记录
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不引入 LLM-as-judge
|
||||
- 不要求完整离线 LLM runtime
|
||||
- 不新增生产 API
|
||||
- 不修改 Chat 主链路
|
||||
- 不修改 evidence trace 运行时语义
|
||||
|
||||
---
|
||||
|
||||
## 预期面试表达
|
||||
|
||||
完成后可以这样描述:
|
||||
|
||||
```text
|
||||
我不仅有一个可演示的 Agent,还给它建立了固定 case 的回归评测。
|
||||
每次修改 prompt、工具或 verifier 后,都可以跑同一批诊断 case,
|
||||
检查证据覆盖、verdict 分布、工具调用成本和关键结论是否退化。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 初始候选 case
|
||||
|
||||
| Case | 目标 |
|
||||
| --- | --- |
|
||||
| payment-timeout | 支付接口超时,验证知识库 + 日志 + 指标证据 |
|
||||
| mysql-pool-exhausted | 数据库连接池耗尽,验证日志和知识库证据 |
|
||||
| redis-timeout | Redis 连接超时,验证日志依赖证据 |
|
||||
| slow-response | P99 响应时间过高,验证指标 + 慢请求日志 |
|
||||
| jvm-memory-risk | JVM 内存 / OOM 风险,验证指标 + 系统事件日志 |
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `mvp/demo/README.md`
|
||||
- `mvp/demo/payment-timeout-acceptance.md`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- `src/main/java/com/superbiz/agent/domain/entity/DiagnosisSession.java`
|
||||
- `src/main/java/com/superbiz/agent/domain/entity/ToolInvocation.java`
|
||||
- `openspec/specs/evidence-trace-hardening/spec.md`
|
||||
@@ -6,6 +6,11 @@
|
||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
|
||||
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
|
||||
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 已归档 | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
|
||||
|
||||
## RAG 重构计划
|
||||
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
# Diagnosis Eval Baseline Diff
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**来源**:P1-B follow-up
|
||||
**依赖**:`diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
现在项目已经有固定诊断 case、完整 fixture 和 baseline report。下一步需要把 baseline 真正用起来:每次改 Agent 后,把新的 report 和 baseline report 做对比。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 baseline 只能告诉我们“标准状态是什么”,但还不能自动告诉我们“这次改动有没有变差”。
|
||||
|
||||
典型问题包括:
|
||||
|
||||
- pass rate 是否下降。
|
||||
- 某个 case 是否从通过变失败。
|
||||
- 某个 evidence tool 是否从覆盖变成缺失。
|
||||
- verifier verdict 分布是否异常变化。
|
||||
- 平均工具调用数和耗时是否明显上升。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
新增一个 deterministic baseline diff 能力,用代码比较两份 `DiagnosisEvalReport`。
|
||||
|
||||
完成后应该做到:
|
||||
|
||||
- 输入 baseline report 和 current report。
|
||||
- 输出结构化 diff。
|
||||
- 标出 regression、improvement 和普通 changed。
|
||||
- 支持 JSON 和 Markdown 输出。
|
||||
- 文档说明面试时怎么解释这套回归判断。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- report-level diff 数据结构。
|
||||
- aggregate 指标比较。
|
||||
- case-level 指标比较。
|
||||
- JSON / Markdown diff writer。
|
||||
- focused tests 和 eval 文档。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不运行真实 Agent。
|
||||
- 不生成新 trace。
|
||||
- 不引入 LLM-as-judge。
|
||||
- 不改现有 evaluator 评分规则。
|
||||
|
||||
---
|
||||
|
||||
## 面试表达
|
||||
|
||||
可以这样讲:
|
||||
|
||||
```text
|
||||
我不是只保存了一份 baseline,而是加了 baseline diff。
|
||||
每次改 prompt、tool、retrieval 或 verifier 后,
|
||||
我都能把新 report 和 baseline 比较,
|
||||
直接看到哪些 case 退化、哪些证据缺失、成本有没有上升。
|
||||
```
|
||||
@@ -0,0 +1,83 @@
|
||||
# Expand Diagnosis Eval Fixtures
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-B follow-up
|
||||
**依赖**:`diagnosis-eval-harness`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
`diagnosis-eval-harness` 已经把固定 case、trace evaluator、JSON / Markdown report 和字段文档搭起来了。
|
||||
|
||||
现在还差一步:5 条固定诊断 case 里,只有 2 条有 fixture,另外 3 条还是 missing 状态。这个状态可以验证 evaluator 的错误报告能力,但还不能作为完整 baseline 展示。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 baseline 还不够完整:
|
||||
|
||||
- `redis-timeout` 没有对应 trace fixture。
|
||||
- `slow-response` 没有对应 trace fixture。
|
||||
- `jvm-memory-risk` 没有对应 trace fixture。
|
||||
- 仓库里还没有一份固定的 baseline JSON / Markdown 报告可供对比。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
补齐固定诊断评测集,让它从“框架可跑”变成“基准可用”。
|
||||
|
||||
完成后应该做到:
|
||||
|
||||
- 5 条固定 case 都能加载到对应 fixture。
|
||||
- evaluator 能输出完整 baseline report。
|
||||
- baseline report 被保存到仓库,后续 Agent 改动可以拿它做对比。
|
||||
- 文档说明怎么重新生成和怎么看报告。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- 补齐 3 个缺失 fixture。
|
||||
- 保存 baseline JSON / Markdown 报告。
|
||||
- 更新 eval 文档。
|
||||
- 补充测试,确保 case 文件引用的 fixture 都存在。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不新增 case 数量。
|
||||
- 不改生产 Agent 主链路。
|
||||
- 不引入 LLM-as-judge。
|
||||
- 不启动真实 MySQL、Redis、Milvus 或 LLM。
|
||||
|
||||
---
|
||||
|
||||
## 面试表达
|
||||
|
||||
可以这样讲:
|
||||
|
||||
```text
|
||||
我先搭了评测 harness,然后把固定 case 的 trace fixture 补齐,
|
||||
生成一份可复现的 baseline report。
|
||||
这样以后每次改 prompt、tool 或 verifier,
|
||||
都能看固定诊断集有没有行为回退,而不是只靠人工感觉。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- `mvp/eval/fixtures/`
|
||||
- `mvp/eval/reports/`
|
||||
- `mvp/eval/README.md`
|
||||
- `mvp/eval/schema.md`
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||
- `src/test/java/com/superbiz/agent/eval/DiagnosisTraceEvaluatorTest.java`
|
||||
- `openspec/specs/diagnosis-eval-harness/spec.md`
|
||||
@@ -0,0 +1,53 @@
|
||||
# MVP Demo Interview Runbook
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**来源**:Plan C
|
||||
**依赖**:`mvp-demo-trace-acceptance`, `evidence-trace-hardening`, `diagnosis-eval-harness`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
项目已经有 Agent 主链路、证据 trace、Verifier、反馈、eval baseline,但这些材料分散在不同目录。面试时真正需要的是一个能快速跑、快速讲清楚的 demo 入口。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 demo 还不够“面试友好”:
|
||||
|
||||
- 启动、请求、trace、反馈步骤分散在文档里。
|
||||
- 没有固定请求 payload 文件。
|
||||
- 没有一键跑 payment-timeout demo 的脚本。
|
||||
- 没有把 trace 字段和面试讲法对应起来的 walkthrough。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
把 Plan C 落地成 `mvp/demo` 下的可复现 demo 包:
|
||||
|
||||
- 固定支付超时请求。
|
||||
- 一键执行 chat、trace、feedback。
|
||||
- 保存 demo 输出,便于复盘。
|
||||
- 提供面试讲解稿和 trace 检查清单。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- `mvp/demo` 文档。
|
||||
- `mvp/demo/requests` 请求文件。
|
||||
- `mvp/demo/scripts` PowerShell 脚本。
|
||||
- `mvp/demo/output` 目录说明。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不新增后端 API。
|
||||
- 不改 Agent prompt。
|
||||
- 不扩 eval harness。
|
||||
- 不处理密钥外置和完整离线化。
|
||||
Reference in New Issue
Block a user