Compare commits

..
9 Commits
Author SHA1 Message Date
aruo 79feed3314 Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0
# Conflicts:
#	mvp/issues/README.md
2026-07-05 01:42:42 +08:00
aruo 98155ae1d8 Merge branch 'aiops-trace-scope' into refactor/mvp1.0 2026-07-05 01:42:23 +08:00
aruo 26e12a8d6b Archive MVP demo interview runbook 2026-07-05 01:34:04 +08:00
aruo cbef3ddd3c Add MVP demo interview runbook 2026-07-05 01:25:20 +08:00
aruo 69deb15330 Add diagnosis eval baseline diff 2026-07-05 00:59:53 +08:00
aruo 4c7c53b024 Expand diagnosis eval fixtures 2026-07-05 00:27:57 +08:00
aruo ca5c61fabf Add diagnosis eval harness 2026-07-04 23:51:43 +08:00
aruo 23ee05c7c3 feat: add traceable scoped AIOps diagnosis 2026-07-04 22:57:28 +08:00
aruo dc6cd32a67 Harden evidence trace semantics 2026-07-04 22:36:30 +08:00
125 changed files with 5781 additions and 166 deletions
+4
View File
@@ -60,3 +60,7 @@ uploads/
### Windows / Runtime Artifacts
*.stackdump
NUL
### MVP Demo Generated Outputs
mvp/demo/output/*.json
!mvp/demo/output/README.md
+7
View File
@@ -4,7 +4,14 @@
| 日期 | slug | 领域 | 关键词 | 状态 |
|---|---|---|---|---|
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
@@ -0,0 +1,14 @@
# Acceptance: aiops-alert-scope-control
## Verification
- [x] Payload-mode prompt focuses the final report on the supplied alert.
- [x] No-payload prompt requires active-alert discovery first.
- [x] Targeted tests pass.
- [x] Compile passes.
- [x] OpenSpec validates.
## Known Limits
- Prompt-only scope control may still require runtime observation.
- AIOps Verifier remains deferred.
@@ -0,0 +1,28 @@
# Brief: aiops-alert-scope-control
## Background
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
## Goal
Make AIOps scope explicit:
- Payload present -> targeted diagnosis for the supplied alert.
- Payload absent -> automatic active-alert discovery and diagnosis.
## Scope
- In scope:
- `AiOpsService.buildTaskPrompt(...)` scope rules.
- Focused tests.
- Demo acceptance wording.
- Out of scope:
- Verifier integration.
- Java-side filtering of tool results.
- API shape changes.
- Database changes.
## Related OpenSpec
`openspec/changes/aiops-alert-scope-control/`
@@ -0,0 +1,42 @@
# Decisions: aiops-alert-scope-control
## Clarify
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
- Slug: `aiops-alert-scope-control`
- Scale: standard-light.
## Context
- AIOps traceability is implemented and verified.
- Mock Prometheus returns multiple active alerts.
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
## GitNexus
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
## Key Decisions
- Payload mode is detected when any alert field is present.
- Payload mode final report must focus on the supplied alert.
- No-payload mode must first call `queryPrometheusAlerts`.
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
@@ -0,0 +1,44 @@
# Evidence: aiops-alert-scope-control
## Local Impact Analysis
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
- No controller, DTO, repository, or database changes are required.
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
## Verification Results
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
- `mvn -q -DskipTests compile` passed.
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
## Runtime Verification
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
- `diagnosis_session` persisted:
- `agent_flow = AI_OPS`
- `status = SUCCESS`
- `total_duration_ms = 69875`
- `step_count = 5`
- `tool_call_count = 8`
- Tool invocation counts:
- `query_metrics = 1`
- `lookup_knowledge = 1`
- `query_logs = 6`
- Report scope check:
- `告警根因分析 - HighCPUUsage` exists.
- `告警根因分析 - HighMemoryUsage` does not exist.
- `告警根因分析 - SlowResponse` does not exist.
- `相关风险告警` exists.
## Runtime Fix
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
- `maximum-pool-size: 5`
- `minimum-idle: 1`
- `connection-timeout: 10000`
- `validation-timeout: 5000`
- `idle-timeout: 60000`
- `max-lifetime: 120000`
- `keepalive-time: 30000`
@@ -0,0 +1,18 @@
# Acceptance: aiops-traceable-diagnosis-entry
## Verification
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
- [x] Targeted AIOps service tests pass.
- [x] Compile verification passes.
- [x] Demo docs describe AIOps request -> session id -> trace query.
## Result
Accepted for implementation scope.
## Known Limits
- AIOps Verifier integration is deferred.
- Runtime still depends on configured model and infrastructure.
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
@@ -0,0 +1,28 @@
# Brief: aiops-traceable-diagnosis-entry
## Background
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
## Goal
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
## Scope
- In scope:
- Optional AIOps alert request body.
- Stable request/session id propagation.
- Persisted AIOps query summary and final answer.
- SSE session id event.
- Demo documentation and focused tests.
- Out of scope:
- Full AIOps and ChatService unification.
- AIOps Verifier integration.
- Database schema changes.
- Sensitive configuration cleanup.
- Fully offline runtime.
## Related OpenSpec
`openspec/changes/aiops-traceable-diagnosis-entry/`
@@ -0,0 +1,68 @@
# Decisions: aiops-traceable-diagnosis-entry
## Clarify
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
- Slug: `aiops-traceable-diagnosis-entry`
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
## Context
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
## User-Interview Confirmations
| Topic | User Words | Decision |
|---|---|---|
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
## Key Decisions
- Keep `/api/ai_ops` and make its body optional.
- Emit `SseMessage.type=session` before long-running analysis starts.
- Store AIOps request summary in `diagnosis_session.query`.
- Store final report in `diagnosis_session.answer`.
- Defer AIOps Verifier to a later change so this slice stays focused.
## Architecture Audit
```text
POST /api/ai_ops
-> optional AIOpsRequest
-> resolve sessionId
-> create diagnosis_session(agentFlow=AI_OPS)
-> run ai_ops_supervisor(planner, executor)
-> AgentLoggingHook persists steps
-> tools persist invocations under SessionContextHolder
-> extract final report
-> persist answer
-> GET /api/diagnosis/{sessionId}/trace replays the run
```
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
@@ -0,0 +1,37 @@
# Evidence: aiops-traceable-diagnosis-entry
## Local Impact Analysis
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
## GitNexus
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
## Expected Verification
- Focused unit tests for AIOps request/session/report helper behavior.
- Compile verification.
- OpenSpec validation if CLI is available.
## Verification Results
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
- `mvn -q -DskipTests compile`: passed.
## Demo Alignment
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
## Metric Alignment Follow-up
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
- Targeted verification:
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
- `mvn -q -DskipTests compile`: passed.
@@ -0,0 +1,36 @@
# Acceptance: diagnosis-eval-harness
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | `ISS-006` and initial OpenSpec artifacts were created. |
| Implementation | Done | Added fixed cases, fixture-mode trace evaluation, aggregate metrics, and JSON / Markdown report writer. |
| Verification | Done | Targeted evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- First implementation uses fixture-mode evaluation.
- Live trace API polling remains a follow-up option.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed
- Notes: Covers fixed case loading, fixture evaluation, missing fixture reporting, reject degraded-output validation, and report writing.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate diagnosis-eval-harness --strict`
- Result: passed
@@ -0,0 +1,31 @@
# Brief: diagnosis-eval-harness
## Background
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
## Goals
1. Define fixed diagnosis cases for the MVP demo domain.
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
3. Produce JSON and Markdown reports for interview and regression use.
4. Keep the first version offline by supporting trace fixtures.
## Scope
- Evaluation case definitions
- Trace fixture shape
- Rule-based evaluator
- JSON / Markdown report output
- Focused offline tests and docs
## Non-Goals
- No LLM-as-judge
- No live end-to-end runtime requirement
- No production API
- No chat or verifier runtime change
## Related OpenSpec
`openspec/changes/diagnosis-eval-harness/`
@@ -0,0 +1,28 @@
# Diagnosis Eval Harness Decisions
## Clarify
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
- Slug: `diagnosis-eval-harness`
- Devflow scale: standard-light
## Context
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
## Key Decisions
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
- Reason: The first regression signal should be deterministic and tied to trace contracts.
- Decision: Support offline fixture traces first.
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
- Decision: Output both JSON and Markdown.
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
## Open Questions
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
@@ -0,0 +1,10 @@
# Diagnosis Eval Harness Evidence
## Evidence
| Source | Evidence | Conclusion | Reported |
|---|---|---|---|
| `openspec/specs/evidence-trace-hardening/spec.md` | Defines stable evidence states and summary behavior | Evaluation can rely on trace semantics rather than ad hoc log parsing | Yes |
| `mvp/demo/README.md` | Documents an end-to-end demo flow with chat, trace, and feedback | Existing demo flow provides the runtime story, but not a reusable evaluation baseline | Yes |
| `DiagnosisTraceService` | Aggregates session, steps, tools, and self-evaluation | Trace response shape can be reused as evaluation input | Yes |
| `ToolTraceSummaryService` | Builds verifier-facing evidence summaries from persisted tool rows | Evaluator can check evidence coverage through persisted trace artifacts | Yes |
@@ -0,0 +1,32 @@
# Acceptance: evidence-trace-hardening
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | `ISS-005` and the initial OpenSpec artifacts were created. |
| Implementation | Done | Recorder contract, lookup persistence path, evidence summary semantics, and degraded-path tests were implemented. |
| Verification | Done | Targeted offline tests and compile verification passed. |
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,LookupKnowledgeToolTest" test`
- Result: passed
- Notes: Covers recorder contract, summary semantics for success/failure/no-evidence, and `ChatService` fallback / degraded paths.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
## Open Questions
| Question | Current position |
| --- | --- |
| Should deduped retrievals be counted separately from generic no-hit events in future evaluation metrics? | Deferred to P1-B; this change preserves enough structure to decide later. |
@@ -0,0 +1,32 @@
# Brief: evidence-trace-hardening
## Background
The MVP already has persisted tool traces and a verifier, but the evidence contract is still only partially standardized. For interview-focused hardening, the project now needs a tighter contract for evidence persistence, no-evidence / failure semantics, and degraded-output behavior.
## Goals
1. Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
2. Make verifier-facing summaries distinguish failed calls, no-hit calls, deduped retrievals, and actual supporting evidence.
3. Add offline tests for verifier fallback and degraded-output paths.
## Scope
- `ToolInvocationRecorder`
- `LookupKnowledgeTool`
- `QueryLogsTools`
- `QueryMetricsTools`
- `ToolTraceSummaryService`
- `ChatService`
- Focused offline tests
## Non-Goals
- No new API or schema
- No evaluation harness yet
- No trace UI
- No security/config cleanup
## Related OpenSpec
`openspec/changes/evidence-trace-hardening/`
@@ -0,0 +1,24 @@
# Evidence Trace Hardening Decisions
## Clarify
- Entry summary: harden the MVP evidence contract before building the P1-B evaluation harness.
- Slug: `evidence-trace-hardening`
- Devflow scale: standard-light
## Context
- `ISS-003` raised verifier traceability and failure-path concerns.
- Current code inspection shows `QueryLogsTools` and `QueryMetricsTools` already use `ToolInvocationRecorder`, while `LookupKnowledgeTool` still persists rows through a local helper.
- `ChatService` already contains fallback behavior for missing/invalid `verifier_output`, but coverage is narrow.
## Key Decisions
- Decision: Treat this as a contract-hardening change, not a new feature change.
- Reason: The project already has the necessary runtime pieces; the gap is semantic consistency and testability.
- Decision: Keep the scope before P1-B.
- Reason: The evaluation harness will rely on stable evidence semantics, so this contract slice should land first.
- Decision: Preserve schema and API stability.
- Reason: The interview value here is engineering rigor, not more surface area.
@@ -0,0 +1,11 @@
# Evidence Trace Hardening Evidence
## Evidence
| Source | Evidence | Conclusion | Reported |
|---|---|---|---|
| `ToolInvocationRecorder` | Provides a common persistence seam for evidence tools | Contract hardening should build on the existing recorder instead of introducing a new store path | Yes |
| `LookupKnowledgeTool` | Still constructs `ToolInvocation` rows through a local helper | Retrieval-aware evidence persistence is not yet unified with the recorder contract | Yes |
| `QueryLogsTools` / `QueryMetricsTools` | Already record evidence invocations through `recordEvidenceTool(...)` | Current gap is semantic alignment, not missing persistence | Yes |
| `ToolTraceSummaryService` | Merges rows by tool and topic domain and infers evidence level heuristically | Summary rules need explicit handling for failure, no-hit, and dedup cases | Yes |
| `ChatService` | Falls back to `LOW_CONFID` when verifier output is missing or invalid | These degraded paths exist and should now be covered by focused offline tests | Yes |
@@ -0,0 +1,37 @@
# Acceptance: expand-diagnosis-eval-fixtures
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
| Implementation | Done | Added remaining fixtures, full baseline reports, and documentation updates. |
| Verification | Done | Evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- Fixture coverage is complete for the five fixed diagnosis cases.
- Baseline reports are saved under `mvp/eval/reports`.
- No production runtime behavior has been changed.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed
- Notes: Covers full fixture coverage, baseline report matching, reject degraded-output validation, and report writing.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate expand-diagnosis-eval-fixtures --strict`
- Result: passed
@@ -0,0 +1,31 @@
# Brief: expand-diagnosis-eval-fixtures
## Background
The diagnosis eval harness is implemented and archived, but the fixed baseline is incomplete because three of the five diagnosis cases still reference missing fixtures.
## Goals
1. Add representative trace fixtures for all remaining fixed diagnosis cases.
2. Save a reproducible baseline report in JSON and Markdown.
3. Document how to regenerate and interpret the baseline.
4. Keep evaluation offline and deterministic.
## Scope
- Redis timeout fixture
- Slow response fixture
- JVM memory risk fixture
- Baseline reports under `mvp/eval/reports`
- Focused tests for full fixture coverage and report generation
## Non-Goals
- No new diagnosis cases
- No production Agent runtime changes
- No LLM-as-judge
- No live infrastructure requirement
## Related OpenSpec
`openspec/changes/expand-diagnosis-eval-fixtures/`
@@ -0,0 +1,28 @@
# Expand Diagnosis Eval Fixtures Decisions
## Clarify
- Entry summary: complete the fixed diagnosis eval baseline after the harness is in place.
- Slug: `expand-diagnosis-eval-fixtures`
- Devflow scale: standard-light
## Context
- `diagnosis-eval-harness` created the evaluator, case file, fixture mode, and report writer.
- The first baseline still has missing fixtures by design.
- This follow-up turns that partial baseline into a full fixed-case baseline.
## Key Decisions
- Decision: Keep this change data-focused.
- Reason: the evaluator rules already landed; this change should not blur fixture expansion with harness behavior changes.
- Decision: Save baseline reports in the repository.
- Reason: interview review and future diffs are easier when the expected baseline is visible.
- Decision: Use deterministic fixture traces instead of live trace generation.
- Reason: this baseline should run without infrastructure or external model calls.
## Open Questions
- Whether a future change should add a CLI or Maven goal for report regeneration.
@@ -0,0 +1,11 @@
# Evidence: expand-diagnosis-eval-fixtures
## Evidence Log
- 2026-07-04: Created slug-based issue `expand-diagnosis-eval-fixtures.md`.
- 2026-07-04: Created OpenSpec change `expand-diagnosis-eval-fixtures`.
- 2026-07-04: Added Redis timeout, slow response, and JVM memory risk fixtures.
- 2026-07-04: Added baseline JSON and Markdown reports under `mvp/eval/reports`.
- 2026-07-04: Verification passed with `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`.
- 2026-07-04: Verification passed with `mvn -q -DskipTests compile`.
- 2026-07-04: Verification passed with `openspec validate expand-diagnosis-eval-fixtures --strict`.
@@ -0,0 +1,37 @@
# Acceptance: diagnosis-eval-baseline-diff
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
- JSON and Markdown diff output are available.
- No production runtime behavior has been changed.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
- Result: passed
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
- Result: passed
@@ -0,0 +1,30 @@
# Brief: diagnosis-eval-baseline-diff
## Background
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
## Goals
1. Compare baseline and current `DiagnosisEvalReport` objects.
2. Detect aggregate and per-case regressions.
3. Output JSON and Markdown diff reports.
4. Document how to read the diff in interview and engineering terms.
## Scope
- Diff data structures
- Deterministic report comparison
- JSON / Markdown diff output
- Focused tests and eval docs
## Non-Goals
- No live Agent execution
- No LLM-as-judge
- No evaluator scoring rule changes
- No production API changes
## Related OpenSpec
`openspec/changes/diagnosis-eval-baseline-diff/`
@@ -0,0 +1,28 @@
# Diagnosis Eval Baseline Diff Decisions
## Clarify
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
- Slug: `diagnosis-eval-baseline-diff`
- Devflow scale: standard-light
## Context
- `diagnosis-eval-harness` created deterministic fixture evaluation.
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
- This change compares new reports against that baseline.
## Key Decisions
- Decision: Diff report DTOs instead of raw traces.
- Reason: the report is the stable contract for regression review.
- Decision: Use deterministic code rules instead of LLM-as-judge.
- Reason: baseline regression checks should be repeatable and explainable.
- Decision: Output both JSON and Markdown.
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
## Open Questions
- Whether a future change should expose this through a CLI or Maven goal.
@@ -0,0 +1,11 @@
# Evidence: diagnosis-eval-baseline-diff
## Evidence Log
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
@@ -0,0 +1,25 @@
# Acceptance: mvp-demo-interview-runbook
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
| Verification | Done | OpenSpec validation passed. |
## Current State
- No backend runtime behavior has been changed.
- Demo is packaged under `mvp/demo` for interview use.
## Verification
### OpenSpec Verification
- Command: `openspec validate mvp-demo-interview-runbook --strict`
- Result: passed
@@ -0,0 +1,29 @@
# Brief: mvp-demo-interview-runbook
## Background
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
## Goals
1. Provide a fixed payment-timeout request payload.
2. Provide a PowerShell script that runs chat, trace, and feedback.
3. Save demo responses under `mvp/demo/output`.
4. Add interview walkthrough and trace checklist.
## Scope
- Demo docs and scripts only
- Existing local APIs only
- Existing `mvp-demo` profile only
## Non-Goals
- No backend code changes
- No eval extension
- No secret cleanup
- No full offline runtime
## Related OpenSpec
`openspec/changes/mvp-demo-interview-runbook/`
@@ -0,0 +1,27 @@
# MVP Demo Interview Runbook Decisions
## Clarify
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
- Slug: `mvp-demo-interview-runbook`
- Devflow scale: standard-light
## Context
- Evidence trace and eval baseline work are already done.
- The next useful step is not more eval tooling, but a runnable demo path.
## Key Decisions
- Decision: Keep this change documentation/script-only.
- Reason: Plan C is about demo packaging, not new runtime capability.
- Decision: Use a stable session id.
- Reason: it makes trace lookup and saved output predictable.
- Decision: Save outputs to `mvp/demo/output`.
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
## Open Questions
- Whether a later change should add a truly offline stubbed demo mode.
@@ -0,0 +1,9 @@
# Evidence: mvp-demo-interview-runbook
## Evidence Log
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
- 2026-07-05: Added fixed payment-timeout request payload.
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
@@ -0,0 +1,81 @@
---
title: AIOps 告警排障 Runbook
keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs]
summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。
category: troubleshooting
---
# AIOps 告警排障 Runbook
## 1. 告警输入处理原则
AIOps 诊断入口有两种触发方式:
- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。
- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。
最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。
## 2. Mock 告警与日志主题映射
| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 |
|---|---|---|---|
| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` |
| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` |
| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` |
| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` |
## 3. HighCPUUsage / payment-service 排障步骤
### 3.1 现象确认
先确认 Prometheus 活动告警中是否存在:
- `alert_name = HighCPUUsage`
- `service = payment-service`
- CPU 使用率超过 80%
- 状态为 firing
如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。
### 3.2 指标与日志取证
推荐工具调用顺序:
1. `queryPrometheusAlerts`:确认当前活动告警。
2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。
3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。
### 3.3 根因判断
可接受的根因结论必须至少满足一项:
- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。
- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。
- 告警持续时间与日志时间线一致。
如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。
## 4. 处理建议
### 临时止血
- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。
- 对高耗时接口开启限流或降级非核心功能。
- 如果近期有发布,检查变更窗口并准备回滚。
### 根因修复
- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。
- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。
- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。
## 5. 报告要求
告警分析报告至少包含:
- 活跃告警清单。
- 告警根因分析。
- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。
- 已执行或建议执行的处理方案。
- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。
+71
View File
@@ -2,6 +2,13 @@
This demo proves the MVP flow from user question to persisted diagnosis trace.
For interview use, start with:
- `interview-walkthrough.md` for the talk track
- `trace-inspection-checklist.md` for fields to inspect
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
- `requests/payment-timeout-chat.json` for the fixed request payload
## Prerequisites
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
@@ -22,6 +29,22 @@ http://localhost:9900
## 1. Run Chat Diagnosis
Fast path:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
This writes:
```text
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
Manual path:
```powershell
$sessionId = "mvp-demo-payment-timeout-001"
$body = @{
@@ -78,6 +101,43 @@ Expected result:
- `success` is `true`.
- A later trace query shows `data.session.feedback` as `useful`.
## 4. Run AIOps Alert Diagnosis
```powershell
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
Expected result:
- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`.
- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload.
- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`.
- `data.session.answer` contains the final alert analysis report when a report is generated.
- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them.
Query the AIOps trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
```
## Demo Story
The important interview story is:
@@ -92,3 +152,14 @@ one session id
-> feedback
-> trace API for replay and audit
```
The AIOps story uses the same audit spine:
```text
one session id
-> alert payload
-> AIOps planner/executor execution
-> evidence tools
-> alert analysis report
-> trace API for replay and audit
```
+38
View File
@@ -0,0 +1,38 @@
# AIOps Alert Acceptance Case
## Goal
Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry.
## Input
- Session id: `mvp-demo-aiops-payment-cpu-001`
- Endpoint: `POST /api/ai_ops`
- Profile: `mvp-demo`
- Alert:
```json
{
"sessionId": "mvp-demo-aiops-payment-cpu-001",
"alertName": "HighCPUUsage",
"service": "payment-service",
"severity": "P1",
"description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。",
"timeRange": "last_15m",
"userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
}
```
## Acceptance Criteria
1. The SSE stream emits a `session` message containing the requested session id.
2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`.
3. The persisted session query contains the alert name, service, severity, time range, and description.
4. If a final report is generated, `diagnosis_session.answer` contains that report.
5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations.
6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections.
## Known Limits
- This slice does not add a Verifier Agent to AIOps.
- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration.
+146
View File
@@ -0,0 +1,146 @@
# Interview Walkthrough: MVP Diagnosis Agent
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
## 30-Second Summary
```text
This is an enterprise diagnosis Agent MVP.
It takes a payment-timeout question, plans the investigation, calls evidence tools,
checks the answer through a verifier, persists the full trace, and accepts feedback.
```
The important claim is not "the model answered once." The claim is:
```text
The system can show what evidence was used, how the answer was checked, and how to replay the session.
```
## Demo Flow
1. Start the service with the `mvp-demo` profile.
2. Run the fixed payment-timeout request.
3. Open `mvp/demo/output/chat-response.json`.
4. Open `mvp/demo/output/trace-response.json`.
5. Point to evidence tools and verifier evaluation.
6. Submit feedback and show it is attached to the same session.
## Commands
Start service:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
Run the demo from another terminal:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
Optional custom session:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
```
## What To Show
### 1. User-Facing Answer
File:
```text
mvp/demo/output/chat-response.json
```
Say:
```text
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
```
### 2. Evidence Trace
File:
```text
mvp/demo/output/trace-response.json
```
Say:
```text
This is the important Agent engineering part.
I can inspect which tools were called, what inputs they received,
whether they succeeded, and what evidence preview was persisted.
```
Point to:
- `data.toolInvocations[*].toolName`
- `data.toolInvocations[*].inputParams`
- `data.toolInvocations[*].outputPreview`
- `data.toolInvocations[*].success`
### 3. Verifier / Self-Evaluation
Point to:
- `data.session.selfEvaluation`
- `data.summary.hasVerifierEvaluation`
Say:
```text
The final answer is not just raw Executor output.
It is checked by a verifier or self-evaluation layer using the persisted trace.
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
```
### 4. Feedback Loop
File:
```text
mvp/demo/output/feedback-response.json
```
Then re-query trace if needed.
Say:
```text
Feedback is attached to the same diagnosis session.
That makes it possible to mine useful / not useful cases later.
```
### 5. Regression Story
Mention, do not deep dive unless asked:
```text
For repeatability, I also built an offline eval baseline.
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
The two are separate on purpose: demo for human review, eval for automated signal.
```
## Strong Interview Framing
Use this phrasing:
```text
I focused on the Agent engineering surface:
traceability, evidence persistence, verifier gating, feedback, and regression checks.
The model answer is only one part of the system.
The more important part is whether we can audit and improve the answer after it is produced.
```
## Known Limits To Say Proactively
```text
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
```
+11
View File
@@ -0,0 +1,11 @@
# Demo Output
This directory is the default output location for local demo responses.
Generated files are intentionally ignored by Git:
- `chat-response.json`
- `trace-response.json`
- `feedback-response.json`
Keep this README so the directory exists in the repository.
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-payment-timeout-001",
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
}
@@ -0,0 +1,54 @@
param(
[string]$BaseUrl = "http://localhost:9900",
[string]$SessionId = "mvp-demo-payment-timeout-001",
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
[string]$OutputDir = "$PSScriptRoot/../output"
)
$ErrorActionPreference = "Stop"
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
$request.Id = $SessionId
$body = $request | ConvertTo-Json -Depth 8
Write-Host "Running payment-timeout chat demo..."
Write-Host "BaseUrl: $BaseUrl"
Write-Host "SessionId: $SessionId"
$chat = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/chat" `
-ContentType "application/json; charset=utf-8" `
-Body $body
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
Write-Host "Saved chat response: $OutputDir/chat-response.json"
$trace = Invoke-RestMethod `
-Method Get `
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
Write-Host "Saved trace response: $OutputDir/trace-response.json"
$feedbackBody = @{
sessionId = $SessionId
feedback = "useful"
} | ConvertTo-Json
$feedback = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/feedback" `
-ContentType "application/json; charset=utf-8" `
-Body $feedbackBody
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
Write-Host ""
Write-Host "Demo completed. Review:"
Write-Host "- mvp/demo/output/chat-response.json"
Write-Host "- mvp/demo/output/trace-response.json"
Write-Host "- mvp/demo/output/feedback-response.json"
+52
View File
@@ -0,0 +1,52 @@
# Trace Inspection Checklist
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
## Session
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
## Agent Steps
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
## Tool Evidence
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
## Summary
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
## What Good Looks Like
```text
same session id
-> final answer
-> persisted agent steps
-> persisted evidence tool calls
-> verifier/self-evaluation
-> feedback attached to the same session
```
+61
View File
@@ -0,0 +1,61 @@
# Diagnosis Eval Harness
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
## Scope
- Case definitions: `cases/diagnosis-cases.json`
- Offline trace fixtures: `fixtures/*.json`
- Field definitions: `schema.md`
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
- Evaluator implementation: `DiagnosisTraceEvaluator`
- Report writer: `DiagnosisEvalReportWriter`
## Current Mode
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
## Verification
Run the focused evaluator test:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
```
The committed baseline report represents the current fixed fixture set:
```text
5 fixed cases
5 passing fixture evaluations
2 PASS verdicts
3 LOW_CONFID verdicts
```
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
## Interview Story
The harness gives the MVP a repeatable baseline:
```text
fixed diagnosis case
-> saved or runtime trace
-> rule-based trace validation
-> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, and verifier behavior
```
## Baseline Diff
Baseline diff compares a current report against `reports/baseline-report.json`.
```text
baseline report
current report
-> deterministic diff
-> regressions, improvements, and changed signals
```
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
+57
View File
@@ -0,0 +1,57 @@
[
{
"id": "payment-timeout",
"title": "Payment API timeout",
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"traceFixture": "payment-timeout-pass.json",
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
"allowedVerdicts": ["PASS", "LOW_CONFID"],
"forbiddenAnswerKeywords": ["无证据确定"]
},
{
"id": "mysql-pool-exhausted",
"title": "MySQL connection pool exhausted",
"question": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
"traceFixture": "mysql-pool-low-confid.json",
"expectedRootCauseKeywords": ["mysql", "连接池", "超时"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["已经完全确认"]
},
{
"id": "redis-timeout",
"title": "Redis timeout",
"question": "支付服务出现 Redis 连接超时,请定位可能原因。",
"traceFixture": "redis-timeout-low-confid.json",
"expectedRootCauseKeywords": ["redis", "超时"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["无需进一步排查"]
},
{
"id": "slow-response",
"title": "Slow response",
"question": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
"traceFixture": "slow-response-pass.json",
"expectedRootCauseKeywords": ["p99", "慢响应"],
"minKeywordMatches": 1,
"requiredEvidenceTools": ["query_metrics", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["没有风险"]
},
{
"id": "jvm-memory-risk",
"title": "JVM memory risk",
"question": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
"traceFixture": "jvm-memory-risk-low-confid.json",
"expectedRootCauseKeywords": ["jvm", "内存", "oom"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["可以忽略"]
}
]
@@ -0,0 +1,52 @@
{
"session": {
"sessionId": "eval-jvm-memory-risk",
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 53000,
"toolCallCount": 2,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n订单服务存在 JVM 内存风险,但还不能完全确认会发生 OOM。指标显示 heap 使用率持续高于 88%,日志出现多次 Full GC 和 allocation pressure,需要继续观察对象增长来源并检查最近发布。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.52,
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "indirect"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-jvm-memory-risk",
"toolName": "query_metrics",
"success": true
},
{
"id": 2,
"sessionId": "eval-jvm-memory-risk",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,53 @@
{
"session": {
"sessionId": "eval-mysql-pool",
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 51000,
"toolCallCount": 2,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nMySQL 连接池可能参与了本次超时问题。日志中出现 connection pool exhausted,但当前缺少完整指标证据,因此只能作为低置信结论处理。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.48,
"tool_trace_summary": [
{
"tool_name": "lookup_knowledge",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-mysql-pool",
"toolName": "lookup_knowledge",
"success": true,
"relevanceLevel": "PRECISE"
},
{
"id": 2,
"sessionId": "eval-mysql-pool",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,64 @@
{
"session": {
"sessionId": "eval-payment-timeout",
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 42000,
"toolCallCount": 3,
"answer": "支付接口超时与连接池等待有关。知识库说明支付超时需要同时检查连接池、日志和指标;日志出现 connection pool exhausted;指标显示支付服务延迟升高。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.86,
"tool_trace_summary": [
{
"tool_name": "lookup_knowledge",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-payment-timeout",
"toolName": "lookup_knowledge",
"success": true,
"relevanceLevel": "PRECISE"
},
{
"id": 2,
"sessionId": "eval-payment-timeout",
"toolName": "query_logs",
"success": true
},
{
"id": 3,
"sessionId": "eval-payment-timeout",
"toolName": "query_metrics",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 3,
"returnedToolCallCount": 3,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,41 @@
{
"session": {
"sessionId": "eval-redis-timeout",
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 36000,
"toolCallCount": 1,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nRedis 连接超时可能和支付服务到 Redis 的网络抖动或连接池等待有关。日志中出现 redis timeout 和 command timeout 记录,但当前缺少指标侧证据,因此只能作为低置信结论处理。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.46,
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-redis-timeout",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 2,
"returnedStepCount": 2,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
+52
View File
@@ -0,0 +1,52 @@
{
"session": {
"sessionId": "eval-slow-response",
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 47000,
"toolCallCount": 2,
"answer": "用户服务 P99 升高主要表现为慢响应。指标显示 P99 latency 从 280ms 上升到 1800ms,日志中同时出现 slow request 和 downstream timeout,因此优先排查下游依赖耗时和线程池排队。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.78,
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-slow-response",
"toolName": "query_metrics",
"success": true
},
{
"id": 2,
"sessionId": "eval-slow-response",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,85 @@
{
"baselineTotalCases" : 5,
"currentTotalCases" : 5,
"baselinePassedCases" : 5,
"currentPassedCases" : 4,
"baselinePassRate" : 1.0,
"currentPassRate" : 0.8,
"regressionCount" : 6,
"improvementCount" : 0,
"changedCount" : 2,
"hasRegression" : true,
"items" : [ {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "passRate",
"baselineValue" : "1.0",
"currentValue" : "0.8",
"delta" : -0.19999999999999996,
"message" : "passRate changed"
}, {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "averageToolCallCount",
"baselineValue" : "2.0",
"currentValue" : "3.0",
"delta" : 1.0,
"message" : "averageToolCallCount changed"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.LOW_CONFID",
"baselineValue" : "3",
"currentValue" : "2",
"delta" : -1.0,
"message" : "verdict count changed for LOW_CONFID"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.REJECT",
"baselineValue" : "0",
"currentValue" : "1",
"delta" : 1.0,
"message" : "verdict count changed for REJECT"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "passed",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout pass state changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "verdict",
"baselineValue" : "LOW_CONFID",
"currentValue" : "REJECT",
"delta" : -1.0,
"message" : "redis-timeout verdict changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "matchedKeywordCount",
"baselineValue" : "2",
"currentValue" : "1",
"delta" : -1.0,
"message" : "redis-timeout matchedKeywordCount changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "evidenceCoverage.query_logs",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout evidence coverage changed for query_logs"
} ]
}
+22
View File
@@ -0,0 +1,22 @@
# Diagnosis Eval Baseline Diff
- Baseline pass rate: 100.00%
- Current pass rate: 80.00%
- Baseline passed cases: 5/5
- Current passed cases: 4/5
- Regressions: 6
- Improvements: 0
- Other changes: 2
## Diff Items
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
| --- | --- | --- | --- | --- | --- | ---: | --- |
| REGRESSION | aggregate | - | passRate | 1.0 | 0.8 | -0.200 | passRate changed |
| REGRESSION | aggregate | - | averageToolCallCount | 2.0 | 3.0 | 1.000 | averageToolCallCount changed |
| CHANGED | aggregate | - | verdictDistribution.LOW_CONFID | 3 | 2 | -1.000 | verdict count changed for LOW_CONFID |
| CHANGED | aggregate | - | verdictDistribution.REJECT | 0 | 1 | 1.000 | verdict count changed for REJECT |
| REGRESSION | case | redis-timeout | passed | true | false | - | redis-timeout pass state changed |
| REGRESSION | case | redis-timeout | verdict | LOW_CONFID | REJECT | -1.000 | redis-timeout verdict changed |
| REGRESSION | case | redis-timeout | matchedKeywordCount | 2 | 1 | -1.000 | redis-timeout matchedKeywordCount changed |
| REGRESSION | case | redis-timeout | evidenceCoverage.query_logs | true | false | - | redis-timeout evidence coverage changed for query_logs |
+82
View File
@@ -0,0 +1,82 @@
{
"totalCases" : 5,
"passedCases" : 5,
"passRate" : 1.0,
"verdictDistribution" : {
"PASS" : 2,
"LOW_CONFID" : 3
},
"averageToolCallCount" : 2.0,
"averageDurationMs" : 45800.0,
"results" : [ {
"caseId" : "payment-timeout",
"title" : "Payment API timeout",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"lookup_knowledge" : true,
"query_logs" : true,
"query_metrics" : true
},
"toolCallCount" : 3,
"durationMs" : 42000
}, {
"caseId" : "mysql-pool-exhausted",
"title" : "MySQL connection pool exhausted",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"lookup_knowledge" : true,
"query_logs" : true
},
"toolCallCount" : 2,
"durationMs" : 51000
}, {
"caseId" : "redis-timeout",
"title" : "Redis timeout",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_logs" : true
},
"toolCallCount" : 1,
"durationMs" : 36000
}, {
"caseId" : "slow-response",
"title" : "Slow response",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_metrics" : true,
"query_logs" : true
},
"toolCallCount" : 2,
"durationMs" : 47000
}, {
"caseId" : "jvm-memory-risk",
"title" : "JVM memory risk",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_metrics" : true,
"query_logs" : true
},
"toolCallCount" : 2,
"durationMs" : 53000
} ]
}
+22
View File
@@ -0,0 +1,22 @@
# Diagnosis Eval Report
- Total cases: 5
- Passed cases: 5
- Pass rate: 100.00%
- Average tool calls: 2.00
- Average duration ms: 45800.00
## Verdict Distribution
- PASS: 2
- LOW_CONFID: 3
## Cases
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
| --- | --- | --- | --- | ---: | ---: | --- |
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
+202
View File
@@ -0,0 +1,202 @@
# Diagnosis Eval Data Schema
这份文档记录评测基准里的数据结构。口语化理解就是:
```text
用例文件说“我要考什么”
trace 文件说“Agent 实际做了什么”
评测结果说“这次有没有跑偏”
汇总报告说“整体稳定性怎么样”
```
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。
## 1. 用例定义
文件:`mvp/eval/cases/diagnosis-cases.json`
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。
```json
{
"id": "payment-timeout",
"title": "Payment API timeout",
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"traceFixture": "payment-timeout-pass.json",
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
"allowedVerdicts": ["PASS", "LOW_CONFID"],
"forbiddenAnswerKeywords": ["无证据确定"]
}
```
字段说明:
| 字段 | 意思 | 评测器怎么用 |
| --- | --- | --- |
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 |
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 |
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 |
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 |
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 |
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 |
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 |
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 |
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 |
## 2. Trace Fixture
目录:`mvp/eval/fixtures/*.json`
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。
当前会读取这些字段:
| Trace 字段 | 意思 | 评测器怎么用 |
| --- | --- | --- |
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 |
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 |
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 |
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 |
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 |
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 |
简单说,trace 里最重要的是三类信息:
```text
最终回答:它说了什么
工具证据:它查了什么
Verifier:它自己有没有承认这个结论可靠
```
## 3. 单条评测结果
Java 类型:`DiagnosisEvalResult`
这是每条 case 跑完之后的判断结果。
| 字段 | 意思 |
| --- | --- |
| `caseId` | 对应的 case id |
| `title` | case 标题 |
| `passed` | 这条 case 是否通过 |
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 |
| `verdict` | 从 trace 里读出来的 Verifier verdict |
| `matchedKeywordCount` | 最终回答命中的关键词数量 |
| `requiredKeywordCount` | case 定义里一共有多少个关键词 |
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` |
| `toolCallCount` | 本次 trace 里工具调用总数 |
| `durationMs` | 本次 trace 的耗时 |
判断通过的口语化规则:
```text
回答要说到关键点
该查的证据工具要查到
Verifier 的结论要在可接受范围内
回答不能出现危险的过度自信表达
如果是 REJECT,就必须走降级模板
```
## 4. 汇总报告
Java 类型:`DiagnosisEvalReport`
这是整个基准集跑完之后的总结果。
| 字段 | 意思 |
| --- | --- |
| `totalCases` | 总共评测了多少条 case |
| `passedCases` | 通过了多少条 |
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` |
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` |
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 |
| `averageDurationMs` | 平均耗时 |
| `results` | 每条 case 的详细结果列表 |
## 5. 怎么看这个基准
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答:
```text
以前能过的诊断题,现在还过不过?
它是不是少查了某些证据?
它是不是变得更自信但证据不足?
它是不是开始输出不该说的话?
它是不是明显变慢了?
```
所以面试里可以这样讲:
```text
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
```
## 6. Baseline Diff
Baseline diff 是拿两份 report 做对比:
```text
baseline report:以前认可的基准结果
current report:这次改动后跑出来的新结果
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
```
Java 类型:
- `DiagnosisEvalDiffReport`
- `DiagnosisEvalDiffItem`
`DiagnosisEvalDiffReport` 字段:
| 字段 | 意思 |
| --- | --- |
| `baselineTotalCases` | baseline 里有多少条 case |
| `currentTotalCases` | current 里有多少条 case |
| `baselinePassedCases` | baseline 通过了多少条 |
| `currentPassedCases` | current 通过了多少条 |
| `baselinePassRate` | baseline 通过率 |
| `currentPassRate` | current 通过率 |
| `regressionCount` | 退化项数量 |
| `improvementCount` | 改善项数量 |
| `changedCount` | 普通变化项数量 |
| `hasRegression` | 是否存在退化 |
| `items` | 具体 diff 明细 |
`DiagnosisEvalDiffItem` 字段:
| 字段 | 意思 |
| --- | --- |
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
| `caseId` | 如果是单条 case 变化,这里记录 case id |
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
| `baselineValue` | baseline 里的值 |
| `currentValue` | current 里的值 |
| `delta` | 数值变化量;非数值变化为空 |
| `message` | 给人看的变化说明 |
口语化判断规则:
```text
pass rate 下降:退化
case 从通过变失败:退化
证据工具从有变没有:退化
关键词命中变少:退化
工具调用或耗时升高:成本上升,记为退化信号
verdict 分布变化:记录变化,供人工判断是否符合预期
```
面试里可以这样讲:
```text
我把 baseline report 和当前 report 做结构化 diff。
它不是再问 LLM,而是用代码比较固定字段。
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
diff 会直接标成 regression。
这样 Agent 改动可以用固定基准做回归判断。
```
@@ -0,0 +1,137 @@
# ISS-005 证据链补齐与降级契约收敛
**状态**:进行中(sm-flow)
**严重程度**:高
**发现时间**:2026-07-04
**来源**:P1-A 面试打磨项 / 基于 ISS-003 的当前实现复核
**关联**:ISS-003(Verifier 证据链、失败路径可验证性)、`chat-verifier-agent`、`mvp-demo-trace-acceptance`
---
## 背景
当前 MVP 已具备:
- `lookup_knowledge`、`query_logs`、`query_metrics` 的工具调用落库
- Verifier 基于 `tool_trace_summary` 做事实核查
- `LOW_CONFID` / `REJECT` 的用户侧降级输出
- trace API 可回放 session、agent_step、tool_invocation 和 self_evaluation
但如果目标是拿这个项目去面试 Agent 工程师,当前实现仍有一个明显短板:
**证据链已经“有了”,但还没有被收敛成清晰、稳定、可测试的工程契约。**
这会直接影响三个面试问题的回答质量:
1. 工具失败时系统会怎样降级?
2. Verifier 看到的 evidence 到底是否一致、可审计?
3. 这些失败路径和降级行为有没有稳定测试,而不是只靠 runtime 演示?
---
## 当前现状复核
### 1. 工具落库入口已经存在,但契约不统一
- `QueryLogsTools` 和 `QueryMetricsTools` 通过 `ToolInvocationRecorder.recordEvidenceTool(...)` 记录 evidence tool 调用。
- `LookupKnowledgeTool` 仍保留独立的 `saveToolInvocation(...)` 路径,自己构造 `ToolInvocation` 实体。
这意味着:
- evidence tool 的公共字段有一套约定
- knowledge retrieval 又有一套定制字段拼装
两者都能工作,但**没有形成统一的“证据调用记录契约”**。
### 2. 失败 / 无结果 / 去重命中的语义不够显式
当前实现里:
- `query_logs` 未命中时会返回 `success=false` + `"未找到匹配的日志"`
- `query_metrics` 失败时会返回 `success=false`
- `lookup_knowledge` 去重命中时会返回 `found=false`,但 `tool_invocation.success=true`
- `ToolTraceSummaryService` 通过 `success`、`relevanceLevel`、`dedupReason` 等字段做启发式摘要
这些行为在代码里是分散成立的,但**没有被定义成统一契约**,导致:
- Verifier 能看到的“失败”和“无证据”边界不够稳定
- 评测时难以明确统计哪些是“调用失败”、哪些是“无命中”、哪些是“已检索过”
### 3. ChatService 的降级路径有实现,但测试矩阵不完整
`ChatService` 已处理:
- `verifier_output` 缺失或无法解析 → fallback `LOW_CONFID`
- `REJECT` → degraded output
- `LOW_CONFID` → disclaimer output
但目前缺少成体系的专项验证,尤其是:
- Verifier 输出非法 JSON
- evidence tool 查询失败
- knowledge lookup 无有效证据
- fallback 文案是否只基于 verifier 缺口拼装
---
## 影响
- **面试表达弱化**:你能讲“我有 trace”,但还不能很硬地讲“我的失败路径是有契约和测试保护的”。
- **评测基础不稳**:后续 P1-B 做 case-based harness 时,统计口径会受 evidence 语义不一致影响。
- **Verifier 可审计性打折**:当前实现可用,但 still relies on code convention,而不是一份明确收敛后的工程协议。
---
## 本 issue 目标
P1-A 只做三件事:
1. 收敛 evidence tool 的落库契约,让 `lookup_knowledge`、`query_logs`、`query_metrics` 的公共语义一致。
2. 明确失败 / 无证据 / 去重 / verifier 非法输出等降级契约,让 `ToolTraceSummaryService` 和 `ChatService` 面向统一状态工作。
3. 增加专项离线测试,覆盖证据摘要与关键降级路径。
---
## 范围
### In scope
- `ToolInvocationRecorder` 契约增强
- `LookupKnowledgeTool` 入库路径收敛
- `QueryLogsTools` / `QueryMetricsTools` evidence 语义对齐
- `ToolTraceSummaryService` 对失败 / no-hit / mixed evidence 的摘要规则收敛
- `ChatService` 对 verifier 非法输出与降级输出的专项测试
- 与该 change 直接相关的文档、OpenSpec、devflow 记录
### Out of scope
- 不引入新的数据库表或 schema 变更
- 不扩展新的 evidence tool
- 不做 P1-B 评测集 / harness
- 不做前端 trace UI
- 不处理敏感配置和默认 `mvn test` 离线化
---
## 预期结果
完成后,项目在面试里应能更清楚地表述为:
```text
我不仅把 Agent 的工具调用落到了库里,
还把 evidence trace、失败语义和 verifier 降级路径收敛成了稳定契约,
并用离线测试覆盖了这些关键失败场景。
```
---
## 相关文件
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`
@@ -0,0 +1,99 @@
# ISS-006 固定诊断评测集与回归 Harness
**状态**:进行中(sm-flow)
**严重程度**:高
**发现时间**:2026-07-04
**来源**:P1-B 面试打磨项
**依赖**:ISS-005 / `evidence-trace-hardening`
---
## 背景
MVP 已经具备可追溯证据链、Verifier 质量门禁、trace API 和固定 demo 流程。上一阶段 `evidence-trace-hardening` 进一步统一了 evidence tool 的状态语义,让系统能稳定区分:
- `supported`
- `no_evidence`
- `deduped`
- `failed`
下一步需要证明 Agent 在一组固定诊断场景下的表现,而不是只依赖单次 demo。
---
## 问题
当前项目能演示一次支付超时诊断,但还缺少稳定的评测基线:
- 每次改 prompt、工具、Verifier 或检索逻辑后,无法快速判断是否退化。
- 只能人工看 trace,缺少结构化通过 / 失败结果。
- 缺少面试时能展示的指标,如 evidence coverage、verdict 分布、工具调用数量和耗时。
---
## 目标
建立一个轻量的固定 case 评测 harness,用于验证 MVP Agent 的诊断质量和证据链完整性。
第一版不做 LLM-as-judge,优先做规则化校验:
- 固定 5 个 MVP 诊断 case
- 每个 case 定义 expected root-cause keywords、required evidence tools、allowed verdicts
- 基于 trace 结果校验 evidence coverage、verifier evaluation、tool invocation、final answer shape
- 输出 JSON 和 Markdown 报告
---
## 范围
### In scope
- 评测 case 定义文件
- trace 规则校验器
- eval runner 或测试入口
- JSON / Markdown 报告输出
- demo 文档和 devflow 记录
### Out of scope
- 不引入 LLM-as-judge
- 不要求完整离线 LLM runtime
- 不新增生产 API
- 不修改 Chat 主链路
- 不修改 evidence trace 运行时语义
---
## 预期面试表达
完成后可以这样描述:
```text
我不仅有一个可演示的 Agent,还给它建立了固定 case 的回归评测。
每次修改 prompt、工具或 verifier 后,都可以跑同一批诊断 case,
检查证据覆盖、verdict 分布、工具调用成本和关键结论是否退化。
```
---
## 初始候选 case
| Case | 目标 |
| --- | --- |
| payment-timeout | 支付接口超时,验证知识库 + 日志 + 指标证据 |
| mysql-pool-exhausted | 数据库连接池耗尽,验证日志和知识库证据 |
| redis-timeout | Redis 连接超时,验证日志依赖证据 |
| slow-response | P99 响应时间过高,验证指标 + 慢请求日志 |
| jvm-memory-risk | JVM 内存 / OOM 风险,验证指标 + 系统事件日志 |
---
## 相关文件
- `mvp/demo/README.md`
- `mvp/demo/payment-timeout-acceptance.md`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
- `src/main/java/com/superbiz/agent/domain/entity/DiagnosisSession.java`
- `src/main/java/com/superbiz/agent/domain/entity/ToolInvocation.java`
- `openspec/specs/evidence-trace-hardening/spec.md`
+5
View File
@@ -6,6 +6,11 @@
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 已归档 | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
## RAG 重构计划
@@ -0,0 +1,73 @@
# Diagnosis Eval Baseline Diff
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-05
**来源**:P1-B follow-up
**依赖**:`diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`
---
## 背景
现在项目已经有固定诊断 case、完整 fixture 和 baseline report。下一步需要把 baseline 真正用起来:每次改 Agent 后,把新的 report 和 baseline report 做对比。
---
## 问题
当前 baseline 只能告诉我们“标准状态是什么”,但还不能自动告诉我们“这次改动有没有变差”。
典型问题包括:
- pass rate 是否下降。
- 某个 case 是否从通过变失败。
- 某个 evidence tool 是否从覆盖变成缺失。
- verifier verdict 分布是否异常变化。
- 平均工具调用数和耗时是否明显上升。
---
## 目标
新增一个 deterministic baseline diff 能力,用代码比较两份 `DiagnosisEvalReport`。
完成后应该做到:
- 输入 baseline report 和 current report。
- 输出结构化 diff。
- 标出 regression、improvement 和普通 changed。
- 支持 JSON 和 Markdown 输出。
- 文档说明面试时怎么解释这套回归判断。
---
## 范围
### In scope
- report-level diff 数据结构。
- aggregate 指标比较。
- case-level 指标比较。
- JSON / Markdown diff writer。
- focused tests 和 eval 文档。
### Out of scope
- 不运行真实 Agent。
- 不生成新 trace。
- 不引入 LLM-as-judge。
- 不改现有 evaluator 评分规则。
---
## 面试表达
可以这样讲:
```text
我不是只保存了一份 baseline,而是加了 baseline diff。
每次改 prompt、tool、retrieval 或 verifier 后,
我都能把新 report 和 baseline 比较,
直接看到哪些 case 退化、哪些证据缺失、成本有没有上升。
```
@@ -0,0 +1,83 @@
# Expand Diagnosis Eval Fixtures
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-04
**来源**:P1-B follow-up
**依赖**:`diagnosis-eval-harness`
---
## 背景
`diagnosis-eval-harness` 已经把固定 case、trace evaluator、JSON / Markdown report 和字段文档搭起来了。
现在还差一步:5 条固定诊断 case 里,只有 2 条有 fixture,另外 3 条还是 missing 状态。这个状态可以验证 evaluator 的错误报告能力,但还不能作为完整 baseline 展示。
---
## 问题
当前 baseline 还不够完整:
- `redis-timeout` 没有对应 trace fixture。
- `slow-response` 没有对应 trace fixture。
- `jvm-memory-risk` 没有对应 trace fixture。
- 仓库里还没有一份固定的 baseline JSON / Markdown 报告可供对比。
---
## 目标
补齐固定诊断评测集,让它从“框架可跑”变成“基准可用”。
完成后应该做到:
- 5 条固定 case 都能加载到对应 fixture。
- evaluator 能输出完整 baseline report。
- baseline report 被保存到仓库,后续 Agent 改动可以拿它做对比。
- 文档说明怎么重新生成和怎么看报告。
---
## 范围
### In scope
- 补齐 3 个缺失 fixture。
- 保存 baseline JSON / Markdown 报告。
- 更新 eval 文档。
- 补充测试,确保 case 文件引用的 fixture 都存在。
### Out of scope
- 不新增 case 数量。
- 不改生产 Agent 主链路。
- 不引入 LLM-as-judge。
- 不启动真实 MySQL、Redis、Milvus 或 LLM。
---
## 面试表达
可以这样讲:
```text
我先搭了评测 harness,然后把固定 case 的 trace fixture 补齐,
生成一份可复现的 baseline report。
这样以后每次改 prompt、tool 或 verifier,
都能看固定诊断集有没有行为回退,而不是只靠人工感觉。
```
---
## 相关文件
- `mvp/eval/cases/diagnosis-cases.json`
- `mvp/eval/fixtures/`
- `mvp/eval/reports/`
- `mvp/eval/README.md`
- `mvp/eval/schema.md`
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- `src/test/java/com/superbiz/agent/eval/DiagnosisTraceEvaluatorTest.java`
- `openspec/specs/diagnosis-eval-harness/spec.md`
+53
View File
@@ -0,0 +1,53 @@
# MVP Demo Interview Runbook
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-05
**来源**:Plan C
**依赖**:`mvp-demo-trace-acceptance`, `evidence-trace-hardening`, `diagnosis-eval-harness`
---
## 背景
项目已经有 Agent 主链路、证据 trace、Verifier、反馈、eval baseline,但这些材料分散在不同目录。面试时真正需要的是一个能快速跑、快速讲清楚的 demo 入口。
---
## 问题
当前 demo 还不够“面试友好”:
- 启动、请求、trace、反馈步骤分散在文档里。
- 没有固定请求 payload 文件。
- 没有一键跑 payment-timeout demo 的脚本。
- 没有把 trace 字段和面试讲法对应起来的 walkthrough。
---
## 目标
把 Plan C 落地成 `mvp/demo` 下的可复现 demo 包:
- 固定支付超时请求。
- 一键执行 chat、trace、feedback。
- 保存 demo 输出,便于复盘。
- 提供面试讲解稿和 trace 检查清单。
---
## 范围
### In scope
- `mvp/demo` 文档。
- `mvp/demo/requests` 请求文件。
- `mvp/demo/scripts` PowerShell 脚本。
- `mvp/demo/output` 目录说明。
### Out of scope
- 不新增后端 API。
- 不改 Agent prompt。
- 不扩 eval harness。
- 不处理密钥外置和完整离线化。
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,54 @@
## Context
The AIOps endpoint has two natural modes:
- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert.
- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them.
The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied.
## Goals / Non-Goals
**Goals:**
- Make AIOps payload mode single-alert focused.
- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior.
- Keep the change prompt-only and low risk.
- Add tests for prompt scope rules.
**Non-Goals:**
- Do not add a Verifier Agent.
- Do not force tool calls in Java code.
- Do not change `/api/ai_ops` request/response contracts.
- Do not modify mock alert data.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. |
| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. |
| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. |
| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. |
## Prompt Rules
Payload mode MUST instruct the agent:
- Treat supplied payload as the primary and only report target.
- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk.
- Do not create root-cause sections for unrelated active alerts.
- Report unrelated alerts only in a brief "关联风险" note if they appear relevant.
Auto-discovery mode MUST instruct the agent:
- First call `queryPrometheusAlerts`.
- Select P0/P1 or longest-running firing alerts.
- Analyze one or more active alerts based on severity and evidence.
## Risks / Trade-offs
- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace.
- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections.
- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier.
@@ -0,0 +1,27 @@
## Why
Runtime verification showed that AIOps now correctly accepts an alert payload and persists a trace, but the generated report still expands to every active mock Prometheus alert. That weakens the product boundary between `/api/chat` and `/api/ai_ops`: an alert payload should mean targeted alert diagnosis, while an empty payload should mean automatic active-alert discovery.
## What Changes
- Tighten the AIOps task prompt so payload mode focuses the final report on the supplied alert.
- Preserve full active-alert discovery when no payload is supplied.
- Allow Prometheus active-alert lookup in payload mode only as supporting evidence, not as permission to expand the report to unrelated alerts.
- Update tests and demo acceptance wording to lock the new behavior.
## Capabilities
### New Capabilities
- `aiops-alert-scope-control`: Defines AIOps diagnosis scope rules for payload mode versus auto-discovery mode.
### Modified Capabilities
- `aiops-traceable-diagnosis-entry`: Keeps the same API and trace behavior but clarifies how AIOps should scope its diagnosis.
## Impact
- Affected code: `AiOpsService.buildTaskPrompt(...)`, focused tests, demo documentation, devflow records.
- Affected API: no endpoint or request/response shape change.
- Affected persistence: no schema change.
- Non-goals: no Verifier integration, no tool implementation change, no prompt rewrite for Chat.
@@ -0,0 +1,25 @@
## ADDED Requirements
### Requirement: AIOps payload mode focuses on supplied alert
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
#### Scenario: Request includes alertName and service
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
- **THEN** the AIOps task prompt identifies payload mode
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
### Requirement: AIOps auto-discovery mode queries active alerts first
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
#### Scenario: Request body is omitted
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
- **THEN** the AIOps task prompt identifies auto-discovery mode
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
### Requirement: Payload mode may use active alerts as supporting context
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
#### Scenario: Prometheus returns multiple active alerts
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
- **THEN** the prompt permits mentioning those alerts only as related risk or context
- **AND** the final report target remains the supplied alert
@@ -0,0 +1,21 @@
## 1. Flow Records
- [x] 1.1 Add devflow brief, decisions with lightweight Grill, evidence, and acceptance records.
- [x] 1.2 Record local impact analysis and GitNexus skip context.
## 2. Prompt Scope Control
- [x] 2.1 Add payload detection helper in `AiOpsService`.
- [x] 2.2 Update `buildTaskPrompt(...)` with payload-mode and auto-discovery-mode rules.
## 3. Tests And Docs
- [x] 3.1 Add tests for payload-mode prompt rules.
- [x] 3.2 Add tests for no-payload auto-discovery prompt rules.
- [x] 3.3 Update AIOps demo acceptance wording for single-alert payload mode.
## 4. Verification
- [x] 4.1 Run targeted tests.
- [x] 4.2 Run compile verification.
- [x] 4.3 Run OpenSpec validation.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,90 @@
## Context
The AIOps endpoint is currently useful as a standalone alert-analysis demo, but it is not aligned with the MVP trace story:
- `ChatController.aiOps()` accepts no request body.
- `AiOpsService.executeAiOpsAnalysis(...)` creates a random 8-character session id internally.
- The caller cannot reliably discover that id and query `GET /api/diagnosis/{sessionId}/trace`.
- `extractFinalReport(...)` returns the report text but does not persist it to `diagnosis_session.answer`.
The existing trace API already aggregates `diagnosis_session`, `agent_step`, and `tool_invocation`, so this change should reuse that storage rather than introduce new persistence.
## Goals / Non-Goals
**Goals:**
- Make `/api/ai_ops` usable as an alert-triggered diagnosis entry point.
- Preserve backward compatibility for callers that post with no request body.
- Return the resolved `sessionId` through SSE.
- Persist the final report into the existing diagnosis session.
- Keep the AIOps path observable through the existing trace API.
**Non-Goals:**
- Do not merge AIOps into `ChatService`.
- Do not add a new Verifier Agent to AIOps in this slice.
- Do not change `GET /api/diagnosis/{sessionId}/trace`.
- Do not add database migrations.
- Do not clean up sensitive configuration.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| API compatibility | Keep `POST /api/ai_ops` SSE and make body optional | Add a new `/api/ai_ops/v2` endpoint | Optional body keeps existing demo callers working while enabling traceable alert input. |
| Session identity | Accept request `sessionId`, otherwise generate UUID | Continue internal-only 8-char id | Reviewers need the id to query trace and submit feedback. |
| Query persistence | Build a concise alert diagnosis query from request fields | Store only "AI Ops 告警分析" | Trace should show what alert was diagnosed. |
| Final answer persistence | Save extracted final report to `diagnosis_session.answer` | Only stream the report | Trace replay must include the final answer without relying on SSE logs. |
| Verifier scope | Defer AIOps Verifier integration | Add Chat Verifier now | The minimum interview value is traceability; Verifier unification can be a follow-up after this entry point is stable. |
| GitNexus | Skip by user decision | Block until MCP available | GitNexus tools are not exposed in this session, and the user explicitly requested skipping GitNexus. Local impact analysis and tests cover this slice. |
## Interface Impact
- Level: L3 API behavior extension.
- Endpoint: `POST /api/ai_ops`
- Compatibility: callers may still omit a body. New callers may send:
```json
{
"sessionId": "mvp-demo-aiops-payment-latency-001",
"alertName": "payment-service-latency-high",
"service": "payment-service",
"severity": "P1",
"description": "支付服务 P95 延迟升高并伴随超时错误",
"timeRange": "last_15m"
}
```
The SSE stream emits a first content message containing the resolved session id:
```text
sessionId: mvp-demo-aiops-payment-latency-001
```
## Data Flow
```text
POST /api/ai_ops
-> ChatController resolves request body and tools
-> AiOpsService.executeAiOpsAnalysis(chatModel, tools, request)
-> create diagnosis_session(agentFlow=AI_OPS, query=<alert summary>)
-> set SessionContextHolder(sessionId)
-> ai_ops_supervisor -> planner_agent -> executor_agent
-> persist agent_step and tool_invocation through existing hooks/tools
-> extract final report
-> persist diagnosis_session.answer/status/counts
-> caller queries GET /api/diagnosis/{sessionId}/trace
```
## Risks / Trade-offs
- [Risk] AIOps still lacks the Chat Verifier quality gate. -> Mitigation: document as follow-up and keep this slice focused on traceability.
- [Risk] SSE clients may not parse the new first message. -> Mitigation: message is additive content; existing clients still receive the final report.
- [Risk] Optional request body in Spring MVC can be easy to mishandle. -> Mitigation: use `@RequestBody(required = false)` and default request values in service code.
- [Risk] AIOps generated reports may still depend on real infrastructure. -> Mitigation: demo profile already enables mock logs/metrics where available; full offline mode remains out of scope.
## Migration Plan
- No database migration.
- Deploy with application restart.
- Rollback by reverting controller/service/DTO changes; existing persisted sessions remain valid.
@@ -0,0 +1,29 @@
## Why
The MVP already has a strong traceable chat diagnosis path, but the legacy `/api/ai_ops` endpoint still behaves like an early standalone demo: it accepts no alert payload, generates an internal session id that callers cannot reuse, and streams a report without reliably persisting the final answer for trace replay. For an Agent Engineer interview project, AIOps should become a second entry point into the same observable diagnosis story rather than a disconnected legacy path.
## What Changes
- Allow `/api/ai_ops` to accept an optional alert diagnosis request body.
- Resolve a stable session id from the request or generate one when omitted.
- Persist the AIOps alert query and final report into `diagnosis_session`.
- Emit the resolved session id in the SSE stream so reviewers can call `GET /api/diagnosis/{sessionId}/trace`.
- Keep the existing AIOps planner/executor flow and evidence tools; do not replace it with the chat flow in this slice.
- Document the AIOps demo path beside the existing MVP demo trace flow.
## Capabilities
### New Capabilities
- `aiops-traceable-diagnosis-entry`: Makes the AIOps alert endpoint traceable by session id and replayable through the existing diagnosis trace API.
### Modified Capabilities
- Existing `/api/ai_ops` behavior is extended from a no-input SSE trigger into an optional request-body alert diagnosis endpoint.
## Impact
- Affected code: `ChatController`, `AiOpsService`, `AIOpsRequest`, focused tests, MVP demo documentation, devflow records.
- Affected API: `POST /api/ai_ops` remains SSE, but now accepts an optional JSON body and streams a first message containing `sessionId`.
- Affected persistence: no schema migration; writes existing `diagnosis_session.query`, `answer`, `status`, timing, and aggregate counts.
- Non-goals: no full AIOps/Chat service unification, no new database table, no production security cleanup, no full offline fake runtime, no mandatory Verifier integration for AIOps in this slice.
@@ -0,0 +1,41 @@
## ADDED Requirements
### Requirement: AIOps endpoint accepts optional alert input
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
#### Scenario: Caller supplies alert input
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
#### Scenario: Caller omits alert input
- **WHEN** a caller posts to `/api/ai_ops` without a body
- **THEN** the system still starts the default AIOps alert-analysis flow
### Requirement: AIOps session id is traceable
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
#### Scenario: Request includes session id
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
- **THEN** the created `diagnosis_session.session_id` equals that value
- **AND** the SSE stream includes the same session id
#### Scenario: Request omits session id
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
- **THEN** the system generates a session id
- **AND** the SSE stream includes the generated session id
### Requirement: AIOps report is persisted for trace replay
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
#### Scenario: AIOps report is generated
- **WHEN** the AIOps planner/executor flow returns a final report
- **THEN** the corresponding diagnosis session is marked successful
- **AND** `diagnosis_session.answer` stores the final report
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
### Requirement: AIOps trace uses existing evidence tables
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
#### Scenario: AIOps uses evidence tools
- **WHEN** the AIOps flow calls available evidence tools
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
@@ -0,0 +1,22 @@
## 1. Flow Records
- [x] 1.1 Add devflow brief, decisions, evidence, and acceptance records for `aiops-traceable-diagnosis-entry`.
- [x] 1.2 Record user-approved GitNexus skip and local impact analysis.
## 2. AIOps API And Service
- [x] 2.1 Extend `AIOpsRequest` with optional session id and alert fields.
- [x] 2.2 Change `/api/ai_ops` to accept an optional request body and emit the resolved session id in SSE.
- [x] 2.3 Change `AiOpsService` to accept the request, resolve session id, build a request-specific prompt, and persist the request summary.
- [x] 2.4 Persist the final AIOps report to `diagnosis_session.answer`.
## 3. Demo Documentation
- [x] 3.1 Add an AIOps alert demo section to `mvp/demo/README.md`.
- [x] 3.2 Add a concrete AIOps acceptance case under `mvp/demo`.
## 4. Verification
- [x] 4.1 Add focused unit coverage for AIOps request/session/report persistence behavior where practical.
- [x] 4.2 Run targeted tests.
- [x] 4.3 Run compile verification.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,54 @@
## Context
The project now has the pieces needed for trace-based evaluation:
- `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation`
- `agent_step` stores ordered agent execution records
- `tool_invocation` stores evidence tool calls with normalized evidence semantics
- `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review
- `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
## Goals / Non-Goals
**Goals:**
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
- Produce JSON and Markdown reports with pass/fail status and key metrics.
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
- Leave room for a later runtime mode that queries the trace API after a demo run.
**Non-Goals:**
- No LLM-as-judge in this slice.
- No automatic prompt optimization.
- No new production API.
- No change to chat, verifier, retrieval, upload, or feedback behavior.
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
## Risks / Trade-offs
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
## Migration Plan
- No deployment migration is required.
- The harness is additive and can be run locally as a test or script.
- Rollback is deleting the eval case files, runner, and report docs.
## Open Questions
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?
@@ -0,0 +1,27 @@
## Why
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
## What Changes
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
- Add JSON and Markdown report output for quick review after a run.
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
## Capabilities
### New Capabilities
- `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
### Modified Capabilities
- None.
## Impact
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
- Affected APIs: none.
- Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions.
- Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.
@@ -0,0 +1,59 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
#### Scenario: Case definition includes expected evidence
- **WHEN** an evaluation case is defined
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
#### Scenario: Case definition can express forbidden behavior
- **WHEN** a case has known unsafe behavior
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
### Requirement: Evaluation harness SHALL validate diagnosis traces
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
#### Scenario: Evidence coverage validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
#### Scenario: Verifier evaluation validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
- **AND** the verdict SHALL be one of the case's allowed verdicts
#### Scenario: Answer keyword validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
#### Scenario: Degraded output validation
- **WHEN** a trace verdict is `REJECT`
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
### Requirement: Evaluation harness SHALL report quality and cost signals
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
#### Scenario: JSON report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
#### Scenario: Markdown report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
#### Scenario: Aggregate metrics
- **WHEN** multiple cases are evaluated
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
### Requirement: Evaluation harness SHALL support offline fixture mode
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
#### Scenario: Fixture trace evaluation
- **WHEN** the evaluator is run against a directory of trace fixture files
- **THEN** it SHALL evaluate each trace file against its matching case definition
- **AND** it SHALL not require a running application service
#### Scenario: Missing fixture is reported clearly
- **WHEN** a case has no matching trace fixture
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
@@ -0,0 +1,25 @@
## 1. Case Definitions
- [x] 1.1 Add evaluation case definition format for fixed MVP diagnosis scenarios.
- [x] 1.2 Add the first 5 case definitions: payment timeout, MySQL pool exhausted, Redis timeout, slow response, and JVM memory risk.
- [x] 1.3 Document the meaning of expected keywords, required evidence tools, allowed verdicts, and forbidden behavior.
## 2. Trace Fixtures
- [x] 2.1 Add fixture trace schema or DTOs that match `DiagnosisTraceResponse` enough for offline evaluation.
- [x] 2.2 Add at least one representative trace fixture for a passing case.
- [x] 2.3 Add at least one fixture covering low-confidence or degraded behavior.
## 3. Evaluator
- [x] 3.1 Implement trace validation rules for evidence coverage, verifier verdict, answer keyword coverage, and degraded-output contract.
- [x] 3.2 Implement aggregate metrics: pass rate, verdict distribution, average tool-call count, and average duration.
- [x] 3.3 Implement JSON report output.
- [x] 3.4 Implement Markdown report output.
## 4. Tests And Documentation
- [x] 4.1 Add focused offline tests for the evaluator.
- [x] 4.2 Add run instructions under `mvp/demo` or `mvp/notes`.
- [x] 4.3 Run targeted tests for the evaluator.
- [x] 4.4 Run compile verification.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,64 @@
## Context
The current MVP already has the core pieces required for traceable agent execution:
- `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation`
- `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder`
- `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries
- `ChatService` already contains fallback behavior for missing or invalid `verifier_output`
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
- `LookupKnowledgeTool` builds `ToolInvocation` rows itself
- the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)`
- “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
- degraded output behavior exists in code but is only lightly covered by tests
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
## Goals / Non-Goals
**Goals:**
- Centralize the common persistence contract for evidence-bearing tools.
- Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
- Define stable summarization semantics for:
- successful evidence
- no-hit / no-usable-evidence
- deduped retrievals
- failed evidence queries
- Make `ChatService` fallback and degraded-output paths testable as explicit product behavior.
- Keep the scope small enough to unblock the next P1-B evaluation harness.
**Non-Goals:**
- No new table, column, or Flyway migration.
- No new public API.
- No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`.
- No attempt to redesign planner/executor routing.
- No full offline runtime or end-to-end benchmark harness in this slice.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. |
| `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
## Risks / Trade-offs
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
- [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic.
## Migration Plan
- No deployment migration is required beyond shipping the code changes.
- Existing `tool_invocation` rows remain valid because this change reuses the same schema.
- Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
## Open Questions
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.
@@ -0,0 +1,26 @@
## Why
The MVP already persists evidence tool invocations and uses a Verifier to judge answer quality, but the current evidence trace semantics are still only partially standardized. For interview-grade agent engineering, the system needs a tighter contract for evidence persistence, no-evidence/failure states, and degraded output behavior, plus focused tests that prove those paths work offline.
## What Changes
- Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
- Align how evidence tools represent success, no-hit, deduped, and failed calls so `ToolTraceSummaryService` can summarize them consistently.
- Harden `ChatService` fallback behavior for invalid or missing verifier output and make the degraded-output paths explicitly testable.
- Add focused offline tests for evidence recording, trace summarization, and verifier fallback / degraded output behavior.
- Record this slice as a dedicated P1-A change tied to the interview-focused MVP hardening track.
## Capabilities
### New Capabilities
- `evidence-trace-hardening`: Covers standardized evidence invocation persistence, verifier-facing evidence summary semantics, and explicit degraded-output contracts for evidence gaps and verifier failures.
### Modified Capabilities
- `chat-verifier-agent`: Tightens verifier input evidence semantics and fallback guarantees without changing the high-level planner/executor/verifier workflow.
## Impact
- Affected code: `ToolInvocationRecorder`, `LookupKnowledgeTool`, `QueryLogsTools`, `QueryMetricsTools`, `ToolTraceSummaryService`, `ChatService`, and focused test classes.
- Affected runtime behavior: evidence-bearing tools will persist more consistent invocation semantics; verifier fallback and degraded outputs remain additive hardening, not a product-flow rewrite.
- Affected APIs: none. No new endpoint or schema is introduced.
- Non-goals: no new evidence tools, no database migration, no evaluation harness, no trace UI, no security/config cleanup in this slice.
@@ -0,0 +1,14 @@
## ADDED Requirements
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
#### Scenario: Failed evidence remains a verifier-visible gap
- **WHEN** an evidence-bearing tool invocation fails
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
- **AND** the verifier flow SHALL continue without crashing
#### Scenario: Deduped retrievals do not count as fresh support
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
- **THEN** those entries SHALL be treated as no-new-evidence
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
@@ -0,0 +1,82 @@
## ADDED Requirements
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
#### Scenario: Common evidence fields are always persisted
- **WHEN** an evidence-bearing tool finishes a call
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
- **WHEN** `lookup_knowledge` persists a tool invocation
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
#### Scenario: Tool failure is preserved as failure
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
- **THEN** the persisted row SHALL set `success=false`
- **AND** it SHALL preserve an `error_message` explaining the failure
#### Scenario: No usable evidence is preserved without pretending success
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
- **THEN** the persisted contract SHALL preserve that the call completed
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
#### Scenario: Deduped retrieval remains auditable
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
- **THEN** the persisted row SHALL preserve the dedup reason
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
#### Scenario: Failed evidence calls remain visible in the summary
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
- **THEN** the summary SHALL retain them
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
#### Scenario: No-hit and deduped calls do not upgrade evidence level
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
- **AND** their counts SHALL still be reflected in the merged summary entry
#### Scenario: Successful evidence keeps the strongest available support
- **WHEN** multiple rows for the same tool and topic domain are merged
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
### Requirement: ChatService SHALL degrade predictably on verifier output failures
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
#### Scenario: Missing verifier output falls back to LOW_CONFID
- **WHEN** the verifier step completes without a usable `verifier_output`
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
- **WHEN** the verifier returns malformed or non-parseable JSON
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the fallback SHALL still persist a verifier-evaluation record
#### Scenario: REJECT output hides unverified raw answer text
- **WHEN** the final verifier decision is `REJECT`
- **THEN** the user-facing output SHALL use the degraded template
- **AND** it SHALL NOT pass through the raw executor answer
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
#### Scenario: Evidence recorder contract is tested offline
- **WHEN** the test suite runs the focused recorder tests
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
#### Scenario: Trace summary hardening is tested offline
- **WHEN** the test suite runs the focused trace-summary tests
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
#### Scenario: Verifier fallback behavior is tested offline
- **WHEN** the test suite runs the focused `ChatService` fallback tests
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
@@ -0,0 +1,22 @@
## 1. Evidence Persistence Contract
- [x] 1.1 Extend `ToolInvocationRecorder` with a richer evidence-recording path that can preserve retrieval-aware fields as well as common evidence fields.
- [x] 1.2 Refactor `LookupKnowledgeTool` to persist `tool_invocation` rows through `ToolInvocationRecorder` instead of its own ad hoc row-construction path.
- [x] 1.3 Align `QueryLogsTools` and `QueryMetricsTools` no-hit / failure payloads with the hardened evidence contract.
## 2. Verifier-Facing Summary Semantics
- [x] 2.1 Harden `ToolTraceSummaryService` so failed, no-hit, and deduped evidence rows are summarized with stable no-evidence semantics.
- [x] 2.2 Preserve merged-call counts for repeated hits, failures, and no-new-evidence rows without overstating evidence strength.
## 3. Chat Degraded Paths
- [x] 3.1 Add focused `ChatService` tests for missing verifier output fallback to `LOW_CONFID`.
- [x] 3.2 Add focused `ChatService` tests for invalid verifier JSON fallback to `LOW_CONFID`.
- [x] 3.3 Add focused `ChatService` tests that `REJECT` output uses the degraded template and does not leak raw executor answer content.
## 4. Verification
- [x] 4.1 Add focused offline tests for the recorder contract and `ToolTraceSummaryService`.
- [x] 4.2 Run targeted test commands for the new/updated offline tests.
- [x] 4.3 Run compile verification.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,39 @@
## Context
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
## Goals / Non-Goals
**Goals:**
- Compare two `DiagnosisEvalReport` objects without requiring external services.
- Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
- Write JSON and Markdown diff outputs for review.
**Non-Goals:**
- Do not run the Agent or regenerate traces.
- Do not introduce LLM-as-judge.
- Do not change evaluator scoring rules.
- Do not block on performance thresholds beyond simple numeric diff signals.
## Decisions
- Decision: Compare report DTOs instead of raw traces.
- Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
- Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
- Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`.
- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
- Decision: Keep thresholds explicit and conservative.
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
## Risks / Trade-offs
- Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed.
- Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
- Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.
@@ -0,0 +1,28 @@
## Why
The evaluation baseline is now complete, but developers still need a repeatable way to decide whether a new Agent run regressed against that baseline. A deterministic baseline diff turns saved reports into an actionable regression signal instead of a static artifact.
## What Changes
- Add a baseline diff model that compares two `DiagnosisEvalReport` objects.
- Detect aggregate changes such as pass-rate drops, verdict distribution shifts, tool-call cost changes, and duration changes.
- Detect per-case changes such as pass/fail regression, verdict changes, keyword coverage changes, evidence coverage loss, and missing/new cases.
- Add JSON and Markdown diff output suitable for review.
- Document how to interpret the diff in the eval docs.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `diagnosis-eval-harness`: Extend the existing evaluation harness so a current report can be compared against the saved baseline report.
## Impact
- Affects eval-only Java code under `src/main/java/com/superbiz/agent/eval`.
- Adds focused tests under `src/test/java/com/superbiz/agent/eval`.
- Updates `mvp/eval` documentation and may add sample diff output.
- No production Agent runtime, API, database schema, or external dependency changes are expected.
@@ -0,0 +1,46 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
#### Scenario: Aggregate regression detection
- **WHEN** the current report has a lower pass rate than the baseline report
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
#### Scenario: Cost signal detection
- **WHEN** average tool-call count or average duration changes between reports
- **THEN** the diff SHALL record the baseline value, current value, and delta
#### Scenario: Verdict distribution comparison
- **WHEN** verdict counts differ between reports
- **THEN** the diff SHALL record the verdict distribution changes
### Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
#### Scenario: Case pass/fail regression
- **WHEN** a case changes from passing in the baseline to failing in the current report
- **THEN** the diff SHALL record a regression for that case
#### Scenario: Evidence coverage regression
- **WHEN** a required evidence tool changes from covered to uncovered for a case
- **THEN** the diff SHALL record a regression naming the case and tool
#### Scenario: Missing case detection
- **WHEN** a baseline case is absent from the current report
- **THEN** the diff SHALL record a regression for the missing case
#### Scenario: New case detection
- **WHEN** a current report contains a case absent from the baseline
- **THEN** the diff SHALL record the case as a non-regression change
### Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
#### Scenario: JSON diff output
- **WHEN** a baseline diff is written as JSON
- **THEN** it SHALL include aggregate summary fields and detailed diff items
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items
@@ -0,0 +1,23 @@
## 1. OpenSpec And Issue Setup
- [x] 1.1 Create slug-based issue and devflow tracking files.
- [x] 1.2 Create OpenSpec proposal, design, delta spec, and tasks.
## 2. Baseline Diff Implementation
- [x] 2.1 Add diff result data structures for summary and per-item changes.
- [x] 2.2 Implement deterministic report comparison rules.
- [x] 2.3 Implement JSON and Markdown diff report writing.
## 3. Documentation
- [x] 3.1 Document baseline diff inputs, outputs, and interpretation in eval docs.
- [x] 3.2 Add sample diff output for a representative regression.
## 4. Tests And Validation
- [x] 4.1 Add focused tests for aggregate and case-level diff behavior.
- [x] 4.2 Add focused tests for JSON and Markdown diff output.
- [x] 4.3 Run evaluator/diff tests.
- [x] 4.4 Run compile verification.
- [x] 4.5 Run OpenSpec validation.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,39 @@
## Context
`diagnosis-eval-harness` already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.
## Goals / Non-Goals
**Goals:**
- Add representative trace fixtures for every fixed diagnosis case.
- Save a baseline report that can be reviewed and compared after future Agent changes.
- Keep the baseline reproducible in offline mode.
- Document how to regenerate the baseline.
**Non-Goals:**
- Do not change production Agent runtime behavior.
- Do not require live infrastructure or a real LLM.
- Do not introduce a new LLM-based grader.
- Do not expand the case set beyond the existing five fixed MVP diagnosis cases.
## Decisions
- Use checked-in fixture traces instead of live service calls.
- Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
- Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
- Save baseline reports under `mvp/eval/reports`.
- Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
- Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
- Keep fixture outcomes representative rather than forcing every case to pass.
- Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
- Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.
## Risks / Trade-offs
- Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like `DiagnosisTraceResponse` and add tests that load every referenced fixture.
- A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
- Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.
@@ -0,0 +1,27 @@
## Why
The diagnosis evaluation harness is implemented, but the baseline is still incomplete because only two of the five fixed cases have trace fixtures. Completing the fixture set and saving a baseline report makes the harness useful as a practical regression signal for interview demos and future Agent changes.
## What Changes
- Add trace fixtures for the remaining fixed diagnosis cases: Redis timeout, slow response, and JVM memory risk.
- Add a reproducible baseline report generated from the full fixture set.
- Document how to regenerate and interpret the baseline.
- Keep the evaluator deterministic and offline; no live MySQL, Redis, Milvus, or LLM service is required.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `diagnosis-eval-harness`: Extend the existing evaluation harness requirement so the fixed MVP case set has complete fixture coverage and a saved baseline report.
## Impact
- Affects `mvp/eval/cases`, `mvp/eval/fixtures`, and eval documentation.
- May add baseline output files under `mvp/eval/reports`.
- May add or update focused evaluator tests to assert full fixture coverage and report generation.
- No production runtime API or database schema changes are expected.
@@ -0,0 +1,27 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
#### Scenario: Every case resolves to a fixture file
- **WHEN** the evaluator loads the fixed case definition file
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
#### Scenario: Fixture files are loadable as diagnosis traces
- **WHEN** each referenced fixture is loaded
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
The system SHALL preserve a generated baseline report for the full fixed fixture set.
#### Scenario: Baseline report includes all fixed cases
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
- **THEN** the report SHALL include one result for every fixed case
#### Scenario: Baseline report is reviewable
- **WHEN** the baseline report is written
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
#### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes fixtures or evaluator rules
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
@@ -0,0 +1,19 @@
## 1. Fixture Coverage
- [x] 1.1 Add Redis timeout trace fixture referenced by the fixed case file.
- [x] 1.2 Add slow response trace fixture referenced by the fixed case file.
- [x] 1.3 Add JVM memory risk trace fixture referenced by the fixed case file.
- [x] 1.4 Verify every `traceFixture` in `diagnosis-cases.json` resolves to an existing fixture file.
## 2. Baseline Reports
- [x] 2.1 Generate a full baseline JSON report for all fixed cases.
- [x] 2.2 Generate a full baseline Markdown report for review.
- [x] 2.3 Document how to regenerate and interpret the baseline reports.
## 3. Tests And Validation
- [x] 3.1 Add or update focused tests for full fixture coverage and baseline report generation.
- [x] 3.2 Run evaluator tests.
- [x] 3.3 Run compile verification.
- [x] 3.4 Run OpenSpec validation.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,35 @@
## Context
The current `mvp/demo` folder documents the core flow, but the steps are embedded in prose. For an interview, the demo needs a sharper entry point: what to start, what to run, what files get produced, and what to point at when explaining Agent engineering quality.
## Goals / Non-Goals
**Goals:**
- Make the payment-timeout demo runnable through a small script.
- Save chat, trace, and feedback responses for review.
- Provide a short interview walkthrough that connects runtime evidence to the engineering story.
- Keep the demo focused on existing APIs and existing `mvp-demo` profile behavior.
**Non-Goals:**
- Do not add new backend endpoints.
- Do not modify Agent prompts or runtime orchestration.
- Do not solve secret cleanup or full offline test isolation in this change.
- Do not expand the eval harness.
## Decisions
- Decision: Use PowerShell scripts.
- Reason: the current runbook already uses PowerShell and the user environment is Windows.
- Decision: Save outputs under `mvp/demo/output`.
- Reason: interview review is easier when chat, trace, and feedback responses are persisted as files.
- Decision: Keep the walkthrough separate from the low-level runbook.
- Reason: `README.md` should tell how to run; `interview-walkthrough.md` should tell how to explain.
## Risks / Trade-offs
- The demo still depends on configured MySQL, Redis, Milvus, and model keys. Mitigation: document this explicitly and keep mock log/metric providers enabled through `mvp-demo`.
- Script assertions are intentionally lightweight. Mitigation: use the trace checklist for human review and keep automated regression in `mvp/eval`.
@@ -0,0 +1,26 @@
## Why
The MVP already has trace, evidence hardening, and evaluation artifacts, but the interview demo path is still too scattered. This change packages the existing capabilities into a repeatable demo runbook that can be executed and explained in a short interview window.
## What Changes
- Add a focused interview walkthrough for the payment-timeout MVP demo.
- Add reusable request payloads and PowerShell scripts under `mvp/demo`.
- Add a trace inspection checklist that maps runtime output to the engineering story.
- Keep the change documentation-only and script-only; no backend runtime behavior changes.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `mvp-demo-trace-acceptance`: Extend the demo acceptance surface with a repeatable interview runbook and executable local demo scripts.
## Impact
- Affects `mvp/demo` documentation and scripts.
- Adds issue and devflow tracking files.
- No Java production code, API contract, database schema, or dependency changes are expected.
@@ -0,0 +1,30 @@
## ADDED Requirements
### Requirement: MVP demo SHALL provide an interview runbook
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
#### Scenario: Walkthrough explains the demo story
- **WHEN** a developer opens the interview walkthrough
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
#### Scenario: Walkthrough stays scoped to existing capabilities
- **WHEN** the walkthrough describes the demo
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
### Requirement: MVP demo SHALL provide executable local demo scripts
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
#### Scenario: Demo script sends the fixed diagnosis request
- **WHEN** the demo script is executed against a running local service
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
#### Scenario: Demo script captures review artifacts
- **WHEN** the demo script finishes successfully
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
### Requirement: MVP demo SHALL provide a trace inspection checklist
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
#### Scenario: Checklist maps fields to interview claims
- **WHEN** a developer reviews a trace response
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
@@ -0,0 +1,16 @@
## 1. Demo Artifacts
- [x] 1.1 Add fixed payment-timeout request payload.
- [x] 1.2 Add PowerShell script to run chat, trace, and feedback steps.
- [x] 1.3 Add output directory documentation without committing generated outputs.
## 2. Interview Documentation
- [x] 2.1 Add interview walkthrough for the demo story.
- [x] 2.2 Add trace inspection checklist.
- [x] 2.3 Update `mvp/demo/README.md` to link the runnable demo package.
## 3. Tracking And Validation
- [x] 3.1 Add slug-based issue and devflow tracking files.
- [x] 3.2 Run OpenSpec validation.
@@ -0,0 +1,28 @@
# aiops-alert-scope-control Specification
## Purpose
TBD - created by archiving change aiops-alert-scope-control. Update Purpose after archive.
## Requirements
### Requirement: AIOps payload mode focuses on supplied alert
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
#### Scenario: Request includes alertName and service
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
- **THEN** the AIOps task prompt identifies payload mode
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
### Requirement: AIOps auto-discovery mode queries active alerts first
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
#### Scenario: Request body is omitted
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
- **THEN** the AIOps task prompt identifies auto-discovery mode
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
### Requirement: Payload mode may use active alerts as supporting context
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
#### Scenario: Prometheus returns multiple active alerts
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
- **THEN** the prompt permits mentioning those alerts only as related risk or context
- **AND** the final report target remains the supplied alert
@@ -0,0 +1,44 @@
# aiops-traceable-diagnosis-entry Specification
## Purpose
TBD - created by archiving change aiops-traceable-diagnosis-entry. Update Purpose after archive.
## Requirements
### Requirement: AIOps endpoint accepts optional alert input
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
#### Scenario: Caller supplies alert input
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
#### Scenario: Caller omits alert input
- **WHEN** a caller posts to `/api/ai_ops` without a body
- **THEN** the system still starts the default AIOps alert-analysis flow
### Requirement: AIOps session id is traceable
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
#### Scenario: Request includes session id
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
- **THEN** the created `diagnosis_session.session_id` equals that value
- **AND** the SSE stream includes the same session id
#### Scenario: Request omits session id
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
- **THEN** the system generates a session id
- **AND** the SSE stream includes the generated session id
### Requirement: AIOps report is persisted for trace replay
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
#### Scenario: AIOps report is generated
- **WHEN** the AIOps planner/executor flow returns a final report
- **THEN** the corresponding diagnosis session is marked successful
- **AND** `diagnosis_session.answer` stores the final report
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
### Requirement: AIOps trace uses existing evidence tables
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
#### Scenario: AIOps uses evidence tools
- **WHEN** the AIOps flow calls available evidence tools
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
@@ -190,3 +190,17 @@ Verifier facts SHALL be linkable to the evidence summaries used during verificat
- **WHEN** the ChatService persists `verifier_evaluation`
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
#### Scenario: Failed evidence remains a verifier-visible gap
- **WHEN** an evidence-bearing tool invocation fails
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
- **AND** the verifier flow SHALL continue without crashing
#### Scenario: Deduped retrievals do not count as fresh support
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
- **THEN** those entries SHALL be treated as no-new-evidence
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
@@ -0,0 +1,135 @@
# diagnosis-eval-harness Specification
## Purpose
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
## Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
#### Scenario: Case definition includes expected evidence
- **WHEN** an evaluation case is defined
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
#### Scenario: Case definition can express forbidden behavior
- **WHEN** a case has known unsafe behavior
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
### Requirement: Evaluation harness SHALL validate diagnosis traces
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
#### Scenario: Evidence coverage validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
#### Scenario: Verifier evaluation validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
- **AND** the verdict SHALL be one of the case's allowed verdicts
#### Scenario: Answer keyword validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
#### Scenario: Degraded output validation
- **WHEN** a trace verdict is `REJECT`
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
### Requirement: Evaluation harness SHALL report quality and cost signals
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
#### Scenario: JSON report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
#### Scenario: Markdown report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
#### Scenario: Aggregate metrics
- **WHEN** multiple cases are evaluated
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
### Requirement: Evaluation harness SHALL support offline fixture mode
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
#### Scenario: Fixture trace evaluation
- **WHEN** the evaluator is run against a directory of trace fixture files
- **THEN** it SHALL evaluate each trace file against its matching case definition
- **AND** it SHALL not require a running application service
#### Scenario: Missing fixture is reported clearly
- **WHEN** a case has no matching trace fixture
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
#### Scenario: Every case resolves to a fixture file
- **WHEN** the evaluator loads the fixed case definition file
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
#### Scenario: Fixture files are loadable as diagnosis traces
- **WHEN** each referenced fixture is loaded
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
The system SHALL preserve a generated baseline report for the full fixed fixture set.
#### Scenario: Baseline report includes all fixed cases
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
- **THEN** the report SHALL include one result for every fixed case
#### Scenario: Baseline report is reviewable
- **WHEN** the baseline report is written
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
#### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes fixtures or evaluator rules
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
### Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
#### Scenario: Aggregate regression detection
- **WHEN** the current report has a lower pass rate than the baseline report
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
#### Scenario: Cost signal detection
- **WHEN** average tool-call count or average duration changes between reports
- **THEN** the diff SHALL record the baseline value, current value, and delta
#### Scenario: Verdict distribution comparison
- **WHEN** verdict counts differ between reports
- **THEN** the diff SHALL record the verdict distribution changes
### Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
#### Scenario: Case pass/fail regression
- **WHEN** a case changes from passing in the baseline to failing in the current report
- **THEN** the diff SHALL record a regression for that case
#### Scenario: Evidence coverage regression
- **WHEN** a required evidence tool changes from covered to uncovered for a case
- **THEN** the diff SHALL record a regression naming the case and tool
#### Scenario: Missing case detection
- **WHEN** a baseline case is absent from the current report
- **THEN** the diff SHALL record a regression for the missing case
#### Scenario: New case detection
- **WHEN** a current report contains a case absent from the baseline
- **THEN** the diff SHALL record the case as a non-regression change
### Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
#### Scenario: JSON diff output
- **WHEN** a baseline diff is written as JSON
- **THEN** it SHALL include aggregate summary fields and detailed diff items
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items
@@ -0,0 +1,86 @@
# evidence-trace-hardening Specification
## Purpose
TBD - created by archiving change evidence-trace-hardening. Update Purpose after archive.
## Requirements
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
#### Scenario: Common evidence fields are always persisted
- **WHEN** an evidence-bearing tool finishes a call
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
- **WHEN** `lookup_knowledge` persists a tool invocation
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
#### Scenario: Tool failure is preserved as failure
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
- **THEN** the persisted row SHALL set `success=false`
- **AND** it SHALL preserve an `error_message` explaining the failure
#### Scenario: No usable evidence is preserved without pretending success
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
- **THEN** the persisted contract SHALL preserve that the call completed
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
#### Scenario: Deduped retrieval remains auditable
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
- **THEN** the persisted row SHALL preserve the dedup reason
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
#### Scenario: Failed evidence calls remain visible in the summary
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
- **THEN** the summary SHALL retain them
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
#### Scenario: No-hit and deduped calls do not upgrade evidence level
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
- **AND** their counts SHALL still be reflected in the merged summary entry
#### Scenario: Successful evidence keeps the strongest available support
- **WHEN** multiple rows for the same tool and topic domain are merged
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
### Requirement: ChatService SHALL degrade predictably on verifier output failures
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
#### Scenario: Missing verifier output falls back to LOW_CONFID
- **WHEN** the verifier step completes without a usable `verifier_output`
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
- **WHEN** the verifier returns malformed or non-parseable JSON
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the fallback SHALL still persist a verifier-evaluation record
#### Scenario: REJECT output hides unverified raw answer text
- **WHEN** the final verifier decision is `REJECT`
- **THEN** the user-facing output SHALL use the degraded template
- **AND** it SHALL NOT pass through the raw executor answer
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
#### Scenario: Evidence recorder contract is tested offline
- **WHEN** the test suite runs the focused recorder tests
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
#### Scenario: Trace summary hardening is tested offline
- **WHEN** the test suite runs the focused trace-summary tests
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
#### Scenario: Verifier fallback behavior is tested offline
- **WHEN** the test suite runs the focused `ChatService` fallback tests
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
@@ -35,3 +35,32 @@ The project SHALL include an end-to-end acceptance case that demonstrates start-
#### Scenario: Reviewer follows the acceptance case
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
### Requirement: MVP demo SHALL provide an interview runbook
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
#### Scenario: Walkthrough explains the demo story
- **WHEN** a developer opens the interview walkthrough
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
#### Scenario: Walkthrough stays scoped to existing capabilities
- **WHEN** the walkthrough describes the demo
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
### Requirement: MVP demo SHALL provide executable local demo scripts
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
#### Scenario: Demo script sends the fixed diagnosis request
- **WHEN** the demo script is executed against a running local service
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
#### Scenario: Demo script captures review artifacts
- **WHEN** the demo script finishes successfully
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
### Requirement: MVP demo SHALL provide a trace inspection checklist
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
#### Scenario: Checklist maps fields to interview claims
- **WHEN** a developer reviews a trace response
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
@@ -131,13 +131,15 @@ public class QueryLogsTools {
output.setMessage(String.format("共有 %d 个可用的日志主题。建议使用默认地域 'ap-guangzhou' 或省略 region 参数", topics.size()));
String response = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
recordInvocation(startTime, "get_available_log_topics", null, null, null, response, true, null, "logs");
recordInvocation("get_available_log_topics", startTime, "get_available_log_topics", null, null, null,
response, true, null, "logs", ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED);
return response;
} catch (Exception e) {
logger.error("获取日志主题列表失败", e);
String response = "{\"success\":false,\"message\":\"获取日志主题列表失败: " + e.getMessage() + "\"}";
recordInvocation(startTime, "get_available_log_topics", null, null, null, response, false, e.getMessage(), "logs");
recordInvocation("get_available_log_topics", startTime, "get_available_log_topics", null, null, null,
response, false, e.getMessage(), "logs", ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
return response;
}
}
@@ -191,8 +193,9 @@ public class QueryLogsTools {
} else {
// 真实模式:调用 CLS API(这里预留接口,后续实现)
String response = buildErrorResponse("CLS 真实查询尚未实现,请启用 mock 模式进行测试");
recordInvocation(startTime, safeQuery, region, logTopic, actualLimit, response, false,
"CLS 真实查询尚未实现,请启用 mock 模式进行测试", normalizeTopicDomain(logTopic));
recordInvocation("query_logs", startTime, safeQuery, region, logTopic, actualLimit, response, false,
"CLS 真实查询尚未实现,请启用 mock 模式进行测试", normalizeTopicDomain(logTopic),
ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
return response;
}
@@ -208,23 +211,26 @@ public class QueryLogsTools {
String jsonResult = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
logger.info("日志查询完成: 找到 {} 条日志", logEntries.size());
recordInvocation(startTime, safeQuery, region, logTopic, actualLimit, jsonResult,
!logEntries.isEmpty(), logEntries.isEmpty() ? "未找到匹配的日志" : null,
normalizeTopicDomain(logTopic));
recordInvocation("query_logs", startTime, safeQuery, region, logTopic, actualLimit, jsonResult,
true, null, normalizeTopicDomain(logTopic),
logEntries.isEmpty()
? ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE
: ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED);
return jsonResult;
} catch (Exception e) {
logger.error("查询日志失败", e);
String response = buildErrorResponse("查询失败: " + e.getMessage());
recordInvocation(startTime, safeQuery, region, logTopic, actualLimit, response, false,
e.getMessage(), normalizeTopicDomain(logTopic));
recordInvocation("query_logs", startTime, safeQuery, region, logTopic, actualLimit, response, false,
e.getMessage(), normalizeTopicDomain(logTopic), ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
return response;
}
}
private void recordInvocation(long startTime, String query, String region, String logTopic, Integer limit,
String output, boolean success, String errorMessage, String topicDomain) {
private void recordInvocation(String toolName, long startTime, String query, String region, String logTopic, Integer limit,
String output, boolean success, String errorMessage, String topicDomain,
String evidenceStatus) {
Map<String, Object> input = new HashMap<>();
input.put("query", query == null || query.isBlank() ? "DEFAULT_QUERY" : query);
if (region != null) {
@@ -239,13 +245,15 @@ public class QueryLogsTools {
input.put("mock_enabled", mockEnabled);
toolInvocationRecorder.recordEvidenceTool(
"query_logs",
toolName,
input,
output,
success,
startTime,
errorMessage,
topicDomain
topicDomain,
evidenceStatus,
Map.of("log_topic", logTopic == null ? "" : logTopic)
);
}
@@ -81,7 +81,7 @@ public class QueryMetricsTools {
if (!"success".equals(result.getStatus())) {
String response = buildErrorResponse("Prometheus API 返回非成功状态: " + result.getStatus(), result.getError());
recordInvocation(startTime, response, false, result.getError());
recordInvocation(startTime, response, false, result.getError(), ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
return response;
}
@@ -119,19 +119,22 @@ public class QueryMetricsTools {
String jsonResult = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
logger.info("Prometheus 告警查询完成: 找到 {} 个告警", simplifiedAlerts.size());
recordInvocation(startTime, jsonResult, true, null);
recordInvocation(startTime, jsonResult, true, null,
simplifiedAlerts.isEmpty()
? ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE
: ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED);
return jsonResult;
} catch (Exception e) {
logger.error("查询 Prometheus 告警失败", e);
String response = buildErrorResponse("查询失败", e.getMessage());
recordInvocation(startTime, response, false, e.getMessage());
recordInvocation(startTime, response, false, e.getMessage(), ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
return response;
}
}
private void recordInvocation(long startTime, String output, boolean success, String errorMessage) {
private void recordInvocation(long startTime, String output, boolean success, String errorMessage, String evidenceStatus) {
toolInvocationRecorder.recordEvidenceTool(
"query_metrics",
Map.of("query", "active_prometheus_alerts", "mock_enabled", mockEnabled),
@@ -139,7 +142,9 @@ public class QueryMetricsTools {
success,
startTime,
errorMessage,
"prometheus_alerts"
"prometheus_alerts",
evidenceStatus,
Map.of("metric_family", "prometheus_alerts")
);
}

Some files were not shown because too many files have changed in this diff Show More