Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0

# Conflicts:
#	mvp/issues/README.md
This commit is contained in:
aruo
2026-07-05 01:42:42 +08:00
96 changed files with 4602 additions and 141 deletions
+23
View File
@@ -2,6 +2,13 @@
This demo proves the MVP flow from user question to persisted diagnosis trace.
For interview use, start with:
- `interview-walkthrough.md` for the talk track
- `trace-inspection-checklist.md` for fields to inspect
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
- `requests/payment-timeout-chat.json` for the fixed request payload
## Prerequisites
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
@@ -22,6 +29,22 @@ http://localhost:9900
## 1. Run Chat Diagnosis
Fast path:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
This writes:
```text
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
Manual path:
```powershell
$sessionId = "mvp-demo-payment-timeout-001"
$body = @{
+146
View File
@@ -0,0 +1,146 @@
# Interview Walkthrough: MVP Diagnosis Agent
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
## 30-Second Summary
```text
This is an enterprise diagnosis Agent MVP.
It takes a payment-timeout question, plans the investigation, calls evidence tools,
checks the answer through a verifier, persists the full trace, and accepts feedback.
```
The important claim is not "the model answered once." The claim is:
```text
The system can show what evidence was used, how the answer was checked, and how to replay the session.
```
## Demo Flow
1. Start the service with the `mvp-demo` profile.
2. Run the fixed payment-timeout request.
3. Open `mvp/demo/output/chat-response.json`.
4. Open `mvp/demo/output/trace-response.json`.
5. Point to evidence tools and verifier evaluation.
6. Submit feedback and show it is attached to the same session.
## Commands
Start service:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
Run the demo from another terminal:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
Optional custom session:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
```
## What To Show
### 1. User-Facing Answer
File:
```text
mvp/demo/output/chat-response.json
```
Say:
```text
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
```
### 2. Evidence Trace
File:
```text
mvp/demo/output/trace-response.json
```
Say:
```text
This is the important Agent engineering part.
I can inspect which tools were called, what inputs they received,
whether they succeeded, and what evidence preview was persisted.
```
Point to:
- `data.toolInvocations[*].toolName`
- `data.toolInvocations[*].inputParams`
- `data.toolInvocations[*].outputPreview`
- `data.toolInvocations[*].success`
### 3. Verifier / Self-Evaluation
Point to:
- `data.session.selfEvaluation`
- `data.summary.hasVerifierEvaluation`
Say:
```text
The final answer is not just raw Executor output.
It is checked by a verifier or self-evaluation layer using the persisted trace.
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
```
### 4. Feedback Loop
File:
```text
mvp/demo/output/feedback-response.json
```
Then re-query trace if needed.
Say:
```text
Feedback is attached to the same diagnosis session.
That makes it possible to mine useful / not useful cases later.
```
### 5. Regression Story
Mention, do not deep dive unless asked:
```text
For repeatability, I also built an offline eval baseline.
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
The two are separate on purpose: demo for human review, eval for automated signal.
```
## Strong Interview Framing
Use this phrasing:
```text
I focused on the Agent engineering surface:
traceability, evidence persistence, verifier gating, feedback, and regression checks.
The model answer is only one part of the system.
The more important part is whether we can audit and improve the answer after it is produced.
```
## Known Limits To Say Proactively
```text
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
```
+11
View File
@@ -0,0 +1,11 @@
# Demo Output
This directory is the default output location for local demo responses.
Generated files are intentionally ignored by Git:
- `chat-response.json`
- `trace-response.json`
- `feedback-response.json`
Keep this README so the directory exists in the repository.
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-payment-timeout-001",
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
}
@@ -0,0 +1,54 @@
param(
[string]$BaseUrl = "http://localhost:9900",
[string]$SessionId = "mvp-demo-payment-timeout-001",
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
[string]$OutputDir = "$PSScriptRoot/../output"
)
$ErrorActionPreference = "Stop"
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
$request.Id = $SessionId
$body = $request | ConvertTo-Json -Depth 8
Write-Host "Running payment-timeout chat demo..."
Write-Host "BaseUrl: $BaseUrl"
Write-Host "SessionId: $SessionId"
$chat = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/chat" `
-ContentType "application/json; charset=utf-8" `
-Body $body
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
Write-Host "Saved chat response: $OutputDir/chat-response.json"
$trace = Invoke-RestMethod `
-Method Get `
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
Write-Host "Saved trace response: $OutputDir/trace-response.json"
$feedbackBody = @{
sessionId = $SessionId
feedback = "useful"
} | ConvertTo-Json
$feedback = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/feedback" `
-ContentType "application/json; charset=utf-8" `
-Body $feedbackBody
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
Write-Host ""
Write-Host "Demo completed. Review:"
Write-Host "- mvp/demo/output/chat-response.json"
Write-Host "- mvp/demo/output/trace-response.json"
Write-Host "- mvp/demo/output/feedback-response.json"
+52
View File
@@ -0,0 +1,52 @@
# Trace Inspection Checklist
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
## Session
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
## Agent Steps
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
## Tool Evidence
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
## Summary
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
## What Good Looks Like
```text
same session id
-> final answer
-> persisted agent steps
-> persisted evidence tool calls
-> verifier/self-evaluation
-> feedback attached to the same session
```
+61
View File
@@ -0,0 +1,61 @@
# Diagnosis Eval Harness
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
## Scope
- Case definitions: `cases/diagnosis-cases.json`
- Offline trace fixtures: `fixtures/*.json`
- Field definitions: `schema.md`
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
- Evaluator implementation: `DiagnosisTraceEvaluator`
- Report writer: `DiagnosisEvalReportWriter`
## Current Mode
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
## Verification
Run the focused evaluator test:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
```
The committed baseline report represents the current fixed fixture set:
```text
5 fixed cases
5 passing fixture evaluations
2 PASS verdicts
3 LOW_CONFID verdicts
```
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
## Interview Story
The harness gives the MVP a repeatable baseline:
```text
fixed diagnosis case
-> saved or runtime trace
-> rule-based trace validation
-> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, and verifier behavior
```
## Baseline Diff
Baseline diff compares a current report against `reports/baseline-report.json`.
```text
baseline report
current report
-> deterministic diff
-> regressions, improvements, and changed signals
```
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
+57
View File
@@ -0,0 +1,57 @@
[
{
"id": "payment-timeout",
"title": "Payment API timeout",
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"traceFixture": "payment-timeout-pass.json",
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
"allowedVerdicts": ["PASS", "LOW_CONFID"],
"forbiddenAnswerKeywords": ["无证据确定"]
},
{
"id": "mysql-pool-exhausted",
"title": "MySQL connection pool exhausted",
"question": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
"traceFixture": "mysql-pool-low-confid.json",
"expectedRootCauseKeywords": ["mysql", "连接池", "超时"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["已经完全确认"]
},
{
"id": "redis-timeout",
"title": "Redis timeout",
"question": "支付服务出现 Redis 连接超时,请定位可能原因。",
"traceFixture": "redis-timeout-low-confid.json",
"expectedRootCauseKeywords": ["redis", "超时"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["无需进一步排查"]
},
{
"id": "slow-response",
"title": "Slow response",
"question": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
"traceFixture": "slow-response-pass.json",
"expectedRootCauseKeywords": ["p99", "慢响应"],
"minKeywordMatches": 1,
"requiredEvidenceTools": ["query_metrics", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["没有风险"]
},
{
"id": "jvm-memory-risk",
"title": "JVM memory risk",
"question": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
"traceFixture": "jvm-memory-risk-low-confid.json",
"expectedRootCauseKeywords": ["jvm", "内存", "oom"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["可以忽略"]
}
]
@@ -0,0 +1,52 @@
{
"session": {
"sessionId": "eval-jvm-memory-risk",
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 53000,
"toolCallCount": 2,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n订单服务存在 JVM 内存风险,但还不能完全确认会发生 OOM。指标显示 heap 使用率持续高于 88%,日志出现多次 Full GC 和 allocation pressure,需要继续观察对象增长来源并检查最近发布。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.52,
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "indirect"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-jvm-memory-risk",
"toolName": "query_metrics",
"success": true
},
{
"id": 2,
"sessionId": "eval-jvm-memory-risk",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,53 @@
{
"session": {
"sessionId": "eval-mysql-pool",
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 51000,
"toolCallCount": 2,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nMySQL 连接池可能参与了本次超时问题。日志中出现 connection pool exhausted,但当前缺少完整指标证据,因此只能作为低置信结论处理。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.48,
"tool_trace_summary": [
{
"tool_name": "lookup_knowledge",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-mysql-pool",
"toolName": "lookup_knowledge",
"success": true,
"relevanceLevel": "PRECISE"
},
{
"id": 2,
"sessionId": "eval-mysql-pool",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,64 @@
{
"session": {
"sessionId": "eval-payment-timeout",
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 42000,
"toolCallCount": 3,
"answer": "支付接口超时与连接池等待有关。知识库说明支付超时需要同时检查连接池、日志和指标;日志出现 connection pool exhausted;指标显示支付服务延迟升高。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.86,
"tool_trace_summary": [
{
"tool_name": "lookup_knowledge",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-payment-timeout",
"toolName": "lookup_knowledge",
"success": true,
"relevanceLevel": "PRECISE"
},
{
"id": 2,
"sessionId": "eval-payment-timeout",
"toolName": "query_logs",
"success": true
},
{
"id": 3,
"sessionId": "eval-payment-timeout",
"toolName": "query_metrics",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 3,
"returnedToolCallCount": 3,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,41 @@
{
"session": {
"sessionId": "eval-redis-timeout",
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 36000,
"toolCallCount": 1,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nRedis 连接超时可能和支付服务到 Redis 的网络抖动或连接池等待有关。日志中出现 redis timeout 和 command timeout 记录,但当前缺少指标侧证据,因此只能作为低置信结论处理。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.46,
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-redis-timeout",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 2,
"returnedStepCount": 2,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
+52
View File
@@ -0,0 +1,52 @@
{
"session": {
"sessionId": "eval-slow-response",
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 47000,
"toolCallCount": 2,
"answer": "用户服务 P99 升高主要表现为慢响应。指标显示 P99 latency 从 280ms 上升到 1800ms,日志中同时出现 slow request 和 downstream timeout,因此优先排查下游依赖耗时和线程池排队。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.78,
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-slow-response",
"toolName": "query_metrics",
"success": true
},
{
"id": 2,
"sessionId": "eval-slow-response",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,85 @@
{
"baselineTotalCases" : 5,
"currentTotalCases" : 5,
"baselinePassedCases" : 5,
"currentPassedCases" : 4,
"baselinePassRate" : 1.0,
"currentPassRate" : 0.8,
"regressionCount" : 6,
"improvementCount" : 0,
"changedCount" : 2,
"hasRegression" : true,
"items" : [ {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "passRate",
"baselineValue" : "1.0",
"currentValue" : "0.8",
"delta" : -0.19999999999999996,
"message" : "passRate changed"
}, {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "averageToolCallCount",
"baselineValue" : "2.0",
"currentValue" : "3.0",
"delta" : 1.0,
"message" : "averageToolCallCount changed"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.LOW_CONFID",
"baselineValue" : "3",
"currentValue" : "2",
"delta" : -1.0,
"message" : "verdict count changed for LOW_CONFID"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.REJECT",
"baselineValue" : "0",
"currentValue" : "1",
"delta" : 1.0,
"message" : "verdict count changed for REJECT"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "passed",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout pass state changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "verdict",
"baselineValue" : "LOW_CONFID",
"currentValue" : "REJECT",
"delta" : -1.0,
"message" : "redis-timeout verdict changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "matchedKeywordCount",
"baselineValue" : "2",
"currentValue" : "1",
"delta" : -1.0,
"message" : "redis-timeout matchedKeywordCount changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "evidenceCoverage.query_logs",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout evidence coverage changed for query_logs"
} ]
}
+22
View File
@@ -0,0 +1,22 @@
# Diagnosis Eval Baseline Diff
- Baseline pass rate: 100.00%
- Current pass rate: 80.00%
- Baseline passed cases: 5/5
- Current passed cases: 4/5
- Regressions: 6
- Improvements: 0
- Other changes: 2
## Diff Items
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
| --- | --- | --- | --- | --- | --- | ---: | --- |
| REGRESSION | aggregate | - | passRate | 1.0 | 0.8 | -0.200 | passRate changed |
| REGRESSION | aggregate | - | averageToolCallCount | 2.0 | 3.0 | 1.000 | averageToolCallCount changed |
| CHANGED | aggregate | - | verdictDistribution.LOW_CONFID | 3 | 2 | -1.000 | verdict count changed for LOW_CONFID |
| CHANGED | aggregate | - | verdictDistribution.REJECT | 0 | 1 | 1.000 | verdict count changed for REJECT |
| REGRESSION | case | redis-timeout | passed | true | false | - | redis-timeout pass state changed |
| REGRESSION | case | redis-timeout | verdict | LOW_CONFID | REJECT | -1.000 | redis-timeout verdict changed |
| REGRESSION | case | redis-timeout | matchedKeywordCount | 2 | 1 | -1.000 | redis-timeout matchedKeywordCount changed |
| REGRESSION | case | redis-timeout | evidenceCoverage.query_logs | true | false | - | redis-timeout evidence coverage changed for query_logs |
+82
View File
@@ -0,0 +1,82 @@
{
"totalCases" : 5,
"passedCases" : 5,
"passRate" : 1.0,
"verdictDistribution" : {
"PASS" : 2,
"LOW_CONFID" : 3
},
"averageToolCallCount" : 2.0,
"averageDurationMs" : 45800.0,
"results" : [ {
"caseId" : "payment-timeout",
"title" : "Payment API timeout",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"lookup_knowledge" : true,
"query_logs" : true,
"query_metrics" : true
},
"toolCallCount" : 3,
"durationMs" : 42000
}, {
"caseId" : "mysql-pool-exhausted",
"title" : "MySQL connection pool exhausted",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"lookup_knowledge" : true,
"query_logs" : true
},
"toolCallCount" : 2,
"durationMs" : 51000
}, {
"caseId" : "redis-timeout",
"title" : "Redis timeout",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_logs" : true
},
"toolCallCount" : 1,
"durationMs" : 36000
}, {
"caseId" : "slow-response",
"title" : "Slow response",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_metrics" : true,
"query_logs" : true
},
"toolCallCount" : 2,
"durationMs" : 47000
}, {
"caseId" : "jvm-memory-risk",
"title" : "JVM memory risk",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_metrics" : true,
"query_logs" : true
},
"toolCallCount" : 2,
"durationMs" : 53000
} ]
}
+22
View File
@@ -0,0 +1,22 @@
# Diagnosis Eval Report
- Total cases: 5
- Passed cases: 5
- Pass rate: 100.00%
- Average tool calls: 2.00
- Average duration ms: 45800.00
## Verdict Distribution
- PASS: 2
- LOW_CONFID: 3
## Cases
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
| --- | --- | --- | --- | ---: | ---: | --- |
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
+202
View File
@@ -0,0 +1,202 @@
# Diagnosis Eval Data Schema
这份文档记录评测基准里的数据结构。口语化理解就是:
```text
用例文件说“我要考什么”
trace 文件说“Agent 实际做了什么”
评测结果说“这次有没有跑偏”
汇总报告说“整体稳定性怎么样”
```
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。
## 1. 用例定义
文件:`mvp/eval/cases/diagnosis-cases.json`
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。
```json
{
"id": "payment-timeout",
"title": "Payment API timeout",
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"traceFixture": "payment-timeout-pass.json",
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
"allowedVerdicts": ["PASS", "LOW_CONFID"],
"forbiddenAnswerKeywords": ["无证据确定"]
}
```
字段说明:
| 字段 | 意思 | 评测器怎么用 |
| --- | --- | --- |
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 |
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 |
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 |
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 |
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 |
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 |
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 |
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 |
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 |
## 2. Trace Fixture
目录:`mvp/eval/fixtures/*.json`
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。
当前会读取这些字段:
| Trace 字段 | 意思 | 评测器怎么用 |
| --- | --- | --- |
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 |
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 |
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 |
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 |
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 |
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 |
简单说,trace 里最重要的是三类信息:
```text
最终回答:它说了什么
工具证据:它查了什么
Verifier:它自己有没有承认这个结论可靠
```
## 3. 单条评测结果
Java 类型:`DiagnosisEvalResult`
这是每条 case 跑完之后的判断结果。
| 字段 | 意思 |
| --- | --- |
| `caseId` | 对应的 case id |
| `title` | case 标题 |
| `passed` | 这条 case 是否通过 |
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 |
| `verdict` | 从 trace 里读出来的 Verifier verdict |
| `matchedKeywordCount` | 最终回答命中的关键词数量 |
| `requiredKeywordCount` | case 定义里一共有多少个关键词 |
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` |
| `toolCallCount` | 本次 trace 里工具调用总数 |
| `durationMs` | 本次 trace 的耗时 |
判断通过的口语化规则:
```text
回答要说到关键点
该查的证据工具要查到
Verifier 的结论要在可接受范围内
回答不能出现危险的过度自信表达
如果是 REJECT,就必须走降级模板
```
## 4. 汇总报告
Java 类型:`DiagnosisEvalReport`
这是整个基准集跑完之后的总结果。
| 字段 | 意思 |
| --- | --- |
| `totalCases` | 总共评测了多少条 case |
| `passedCases` | 通过了多少条 |
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` |
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` |
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 |
| `averageDurationMs` | 平均耗时 |
| `results` | 每条 case 的详细结果列表 |
## 5. 怎么看这个基准
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答:
```text
以前能过的诊断题,现在还过不过?
它是不是少查了某些证据?
它是不是变得更自信但证据不足?
它是不是开始输出不该说的话?
它是不是明显变慢了?
```
所以面试里可以这样讲:
```text
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
```
## 6. Baseline Diff
Baseline diff 是拿两份 report 做对比:
```text
baseline report:以前认可的基准结果
current report:这次改动后跑出来的新结果
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
```
Java 类型:
- `DiagnosisEvalDiffReport`
- `DiagnosisEvalDiffItem`
`DiagnosisEvalDiffReport` 字段:
| 字段 | 意思 |
| --- | --- |
| `baselineTotalCases` | baseline 里有多少条 case |
| `currentTotalCases` | current 里有多少条 case |
| `baselinePassedCases` | baseline 通过了多少条 |
| `currentPassedCases` | current 通过了多少条 |
| `baselinePassRate` | baseline 通过率 |
| `currentPassRate` | current 通过率 |
| `regressionCount` | 退化项数量 |
| `improvementCount` | 改善项数量 |
| `changedCount` | 普通变化项数量 |
| `hasRegression` | 是否存在退化 |
| `items` | 具体 diff 明细 |
`DiagnosisEvalDiffItem` 字段:
| 字段 | 意思 |
| --- | --- |
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
| `caseId` | 如果是单条 case 变化,这里记录 case id |
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
| `baselineValue` | baseline 里的值 |
| `currentValue` | current 里的值 |
| `delta` | 数值变化量;非数值变化为空 |
| `message` | 给人看的变化说明 |
口语化判断规则:
```text
pass rate 下降:退化
case 从通过变失败:退化
证据工具从有变没有:退化
关键词命中变少:退化
工具调用或耗时升高:成本上升,记为退化信号
verdict 分布变化:记录变化,供人工判断是否符合预期
```
面试里可以这样讲:
```text
我把 baseline report 和当前 report 做结构化 diff。
它不是再问 LLM,而是用代码比较固定字段。
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
diff 会直接标成 regression。
这样 Agent 改动可以用固定基准做回归判断。
```
@@ -0,0 +1,137 @@
# ISS-005 证据链补齐与降级契约收敛
**状态**:进行中(sm-flow)
**严重程度**:高
**发现时间**:2026-07-04
**来源**:P1-A 面试打磨项 / 基于 ISS-003 的当前实现复核
**关联**:ISS-003(Verifier 证据链、失败路径可验证性)、`chat-verifier-agent`、`mvp-demo-trace-acceptance`
---
## 背景
当前 MVP 已具备:
- `lookup_knowledge`、`query_logs`、`query_metrics` 的工具调用落库
- Verifier 基于 `tool_trace_summary` 做事实核查
- `LOW_CONFID` / `REJECT` 的用户侧降级输出
- trace API 可回放 session、agent_step、tool_invocation 和 self_evaluation
但如果目标是拿这个项目去面试 Agent 工程师,当前实现仍有一个明显短板:
**证据链已经“有了”,但还没有被收敛成清晰、稳定、可测试的工程契约。**
这会直接影响三个面试问题的回答质量:
1. 工具失败时系统会怎样降级?
2. Verifier 看到的 evidence 到底是否一致、可审计?
3. 这些失败路径和降级行为有没有稳定测试,而不是只靠 runtime 演示?
---
## 当前现状复核
### 1. 工具落库入口已经存在,但契约不统一
- `QueryLogsTools` 和 `QueryMetricsTools` 通过 `ToolInvocationRecorder.recordEvidenceTool(...)` 记录 evidence tool 调用。
- `LookupKnowledgeTool` 仍保留独立的 `saveToolInvocation(...)` 路径,自己构造 `ToolInvocation` 实体。
这意味着:
- evidence tool 的公共字段有一套约定
- knowledge retrieval 又有一套定制字段拼装
两者都能工作,但**没有形成统一的“证据调用记录契约”**。
### 2. 失败 / 无结果 / 去重命中的语义不够显式
当前实现里:
- `query_logs` 未命中时会返回 `success=false` + `"未找到匹配的日志"`
- `query_metrics` 失败时会返回 `success=false`
- `lookup_knowledge` 去重命中时会返回 `found=false`,但 `tool_invocation.success=true`
- `ToolTraceSummaryService` 通过 `success`、`relevanceLevel`、`dedupReason` 等字段做启发式摘要
这些行为在代码里是分散成立的,但**没有被定义成统一契约**,导致:
- Verifier 能看到的“失败”和“无证据”边界不够稳定
- 评测时难以明确统计哪些是“调用失败”、哪些是“无命中”、哪些是“已检索过”
### 3. ChatService 的降级路径有实现,但测试矩阵不完整
`ChatService` 已处理:
- `verifier_output` 缺失或无法解析 → fallback `LOW_CONFID`
- `REJECT` → degraded output
- `LOW_CONFID` → disclaimer output
但目前缺少成体系的专项验证,尤其是:
- Verifier 输出非法 JSON
- evidence tool 查询失败
- knowledge lookup 无有效证据
- fallback 文案是否只基于 verifier 缺口拼装
---
## 影响
- **面试表达弱化**:你能讲“我有 trace”,但还不能很硬地讲“我的失败路径是有契约和测试保护的”。
- **评测基础不稳**:后续 P1-B 做 case-based harness 时,统计口径会受 evidence 语义不一致影响。
- **Verifier 可审计性打折**:当前实现可用,但 still relies on code convention,而不是一份明确收敛后的工程协议。
---
## 本 issue 目标
P1-A 只做三件事:
1. 收敛 evidence tool 的落库契约,让 `lookup_knowledge`、`query_logs`、`query_metrics` 的公共语义一致。
2. 明确失败 / 无证据 / 去重 / verifier 非法输出等降级契约,让 `ToolTraceSummaryService` 和 `ChatService` 面向统一状态工作。
3. 增加专项离线测试,覆盖证据摘要与关键降级路径。
---
## 范围
### In scope
- `ToolInvocationRecorder` 契约增强
- `LookupKnowledgeTool` 入库路径收敛
- `QueryLogsTools` / `QueryMetricsTools` evidence 语义对齐
- `ToolTraceSummaryService` 对失败 / no-hit / mixed evidence 的摘要规则收敛
- `ChatService` 对 verifier 非法输出与降级输出的专项测试
- 与该 change 直接相关的文档、OpenSpec、devflow 记录
### Out of scope
- 不引入新的数据库表或 schema 变更
- 不扩展新的 evidence tool
- 不做 P1-B 评测集 / harness
- 不做前端 trace UI
- 不处理敏感配置和默认 `mvn test` 离线化
---
## 预期结果
完成后,项目在面试里应能更清楚地表述为:
```text
我不仅把 Agent 的工具调用落到了库里,
还把 evidence trace、失败语义和 verifier 降级路径收敛成了稳定契约,
并用离线测试覆盖了这些关键失败场景。
```
---
## 相关文件
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`
@@ -0,0 +1,99 @@
# ISS-006 固定诊断评测集与回归 Harness
**状态**:进行中(sm-flow)
**严重程度**:高
**发现时间**:2026-07-04
**来源**:P1-B 面试打磨项
**依赖**:ISS-005 / `evidence-trace-hardening`
---
## 背景
MVP 已经具备可追溯证据链、Verifier 质量门禁、trace API 和固定 demo 流程。上一阶段 `evidence-trace-hardening` 进一步统一了 evidence tool 的状态语义,让系统能稳定区分:
- `supported`
- `no_evidence`
- `deduped`
- `failed`
下一步需要证明 Agent 在一组固定诊断场景下的表现,而不是只依赖单次 demo。
---
## 问题
当前项目能演示一次支付超时诊断,但还缺少稳定的评测基线:
- 每次改 prompt、工具、Verifier 或检索逻辑后,无法快速判断是否退化。
- 只能人工看 trace,缺少结构化通过 / 失败结果。
- 缺少面试时能展示的指标,如 evidence coverage、verdict 分布、工具调用数量和耗时。
---
## 目标
建立一个轻量的固定 case 评测 harness,用于验证 MVP Agent 的诊断质量和证据链完整性。
第一版不做 LLM-as-judge,优先做规则化校验:
- 固定 5 个 MVP 诊断 case
- 每个 case 定义 expected root-cause keywords、required evidence tools、allowed verdicts
- 基于 trace 结果校验 evidence coverage、verifier evaluation、tool invocation、final answer shape
- 输出 JSON 和 Markdown 报告
---
## 范围
### In scope
- 评测 case 定义文件
- trace 规则校验器
- eval runner 或测试入口
- JSON / Markdown 报告输出
- demo 文档和 devflow 记录
### Out of scope
- 不引入 LLM-as-judge
- 不要求完整离线 LLM runtime
- 不新增生产 API
- 不修改 Chat 主链路
- 不修改 evidence trace 运行时语义
---
## 预期面试表达
完成后可以这样描述:
```text
我不仅有一个可演示的 Agent,还给它建立了固定 case 的回归评测。
每次修改 prompt、工具或 verifier 后,都可以跑同一批诊断 case,
检查证据覆盖、verdict 分布、工具调用成本和关键结论是否退化。
```
---
## 初始候选 case
| Case | 目标 |
| --- | --- |
| payment-timeout | 支付接口超时,验证知识库 + 日志 + 指标证据 |
| mysql-pool-exhausted | 数据库连接池耗尽,验证日志和知识库证据 |
| redis-timeout | Redis 连接超时,验证日志依赖证据 |
| slow-response | P99 响应时间过高,验证指标 + 慢请求日志 |
| jvm-memory-risk | JVM 内存 / OOM 风险,验证指标 + 系统事件日志 |
---
## 相关文件
- `mvp/demo/README.md`
- `mvp/demo/payment-timeout-acceptance.md`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
- `src/main/java/com/superbiz/agent/domain/entity/DiagnosisSession.java`
- `src/main/java/com/superbiz/agent/domain/entity/ToolInvocation.java`
- `openspec/specs/evidence-trace-hardening/spec.md`
+5
View File
@@ -6,6 +6,11 @@
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 已归档 | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
## RAG 重构计划
@@ -0,0 +1,73 @@
# Diagnosis Eval Baseline Diff
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-05
**来源**:P1-B follow-up
**依赖**:`diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`
---
## 背景
现在项目已经有固定诊断 case、完整 fixture 和 baseline report。下一步需要把 baseline 真正用起来:每次改 Agent 后,把新的 report 和 baseline report 做对比。
---
## 问题
当前 baseline 只能告诉我们“标准状态是什么”,但还不能自动告诉我们“这次改动有没有变差”。
典型问题包括:
- pass rate 是否下降。
- 某个 case 是否从通过变失败。
- 某个 evidence tool 是否从覆盖变成缺失。
- verifier verdict 分布是否异常变化。
- 平均工具调用数和耗时是否明显上升。
---
## 目标
新增一个 deterministic baseline diff 能力,用代码比较两份 `DiagnosisEvalReport`。
完成后应该做到:
- 输入 baseline report 和 current report。
- 输出结构化 diff。
- 标出 regression、improvement 和普通 changed。
- 支持 JSON 和 Markdown 输出。
- 文档说明面试时怎么解释这套回归判断。
---
## 范围
### In scope
- report-level diff 数据结构。
- aggregate 指标比较。
- case-level 指标比较。
- JSON / Markdown diff writer。
- focused tests 和 eval 文档。
### Out of scope
- 不运行真实 Agent。
- 不生成新 trace。
- 不引入 LLM-as-judge。
- 不改现有 evaluator 评分规则。
---
## 面试表达
可以这样讲:
```text
我不是只保存了一份 baseline,而是加了 baseline diff。
每次改 prompt、tool、retrieval 或 verifier 后,
我都能把新 report 和 baseline 比较,
直接看到哪些 case 退化、哪些证据缺失、成本有没有上升。
```
@@ -0,0 +1,83 @@
# Expand Diagnosis Eval Fixtures
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-04
**来源**:P1-B follow-up
**依赖**:`diagnosis-eval-harness`
---
## 背景
`diagnosis-eval-harness` 已经把固定 case、trace evaluator、JSON / Markdown report 和字段文档搭起来了。
现在还差一步:5 条固定诊断 case 里,只有 2 条有 fixture,另外 3 条还是 missing 状态。这个状态可以验证 evaluator 的错误报告能力,但还不能作为完整 baseline 展示。
---
## 问题
当前 baseline 还不够完整:
- `redis-timeout` 没有对应 trace fixture。
- `slow-response` 没有对应 trace fixture。
- `jvm-memory-risk` 没有对应 trace fixture。
- 仓库里还没有一份固定的 baseline JSON / Markdown 报告可供对比。
---
## 目标
补齐固定诊断评测集,让它从“框架可跑”变成“基准可用”。
完成后应该做到:
- 5 条固定 case 都能加载到对应 fixture。
- evaluator 能输出完整 baseline report。
- baseline report 被保存到仓库,后续 Agent 改动可以拿它做对比。
- 文档说明怎么重新生成和怎么看报告。
---
## 范围
### In scope
- 补齐 3 个缺失 fixture。
- 保存 baseline JSON / Markdown 报告。
- 更新 eval 文档。
- 补充测试,确保 case 文件引用的 fixture 都存在。
### Out of scope
- 不新增 case 数量。
- 不改生产 Agent 主链路。
- 不引入 LLM-as-judge。
- 不启动真实 MySQL、Redis、Milvus 或 LLM。
---
## 面试表达
可以这样讲:
```text
我先搭了评测 harness,然后把固定 case 的 trace fixture 补齐,
生成一份可复现的 baseline report。
这样以后每次改 prompt、tool 或 verifier,
都能看固定诊断集有没有行为回退,而不是只靠人工感觉。
```
---
## 相关文件
- `mvp/eval/cases/diagnosis-cases.json`
- `mvp/eval/fixtures/`
- `mvp/eval/reports/`
- `mvp/eval/README.md`
- `mvp/eval/schema.md`
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- `src/test/java/com/superbiz/agent/eval/DiagnosisTraceEvaluatorTest.java`
- `openspec/specs/diagnosis-eval-harness/spec.md`
+53
View File
@@ -0,0 +1,53 @@
# MVP Demo Interview Runbook
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-05
**来源**:Plan C
**依赖**:`mvp-demo-trace-acceptance`, `evidence-trace-hardening`, `diagnosis-eval-harness`
---
## 背景
项目已经有 Agent 主链路、证据 trace、Verifier、反馈、eval baseline,但这些材料分散在不同目录。面试时真正需要的是一个能快速跑、快速讲清楚的 demo 入口。
---
## 问题
当前 demo 还不够“面试友好”:
- 启动、请求、trace、反馈步骤分散在文档里。
- 没有固定请求 payload 文件。
- 没有一键跑 payment-timeout demo 的脚本。
- 没有把 trace 字段和面试讲法对应起来的 walkthrough。
---
## 目标
把 Plan C 落地成 `mvp/demo` 下的可复现 demo 包:
- 固定支付超时请求。
- 一键执行 chat、trace、feedback。
- 保存 demo 输出,便于复盘。
- 提供面试讲解稿和 trace 检查清单。
---
## 范围
### In scope
- `mvp/demo` 文档。
- `mvp/demo/requests` 请求文件。
- `mvp/demo/scripts` PowerShell 脚本。
- `mvp/demo/output` 目录说明。
### Out of scope
- 不新增后端 API。
- 不改 Agent prompt。
- 不扩 eval harness。
- 不处理密钥外置和完整离线化。