Compare commits
9
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
79feed3314 | ||
|
|
98155ae1d8 | ||
|
|
26e12a8d6b | ||
|
|
cbef3ddd3c | ||
|
|
69deb15330 | ||
|
|
4c7c53b024 | ||
|
|
ca5c61fabf | ||
|
|
23ee05c7c3 | ||
|
|
dc6cd32a67 |
@@ -60,3 +60,7 @@ uploads/
|
||||
### Windows / Runtime Artifacts
|
||||
*.stackdump
|
||||
NUL
|
||||
|
||||
### MVP Demo Generated Outputs
|
||||
mvp/demo/output/*.json
|
||||
!mvp/demo/output/README.md
|
||||
|
||||
@@ -4,7 +4,14 @@
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
# Acceptance: aiops-alert-scope-control
|
||||
|
||||
## Verification
|
||||
|
||||
- [x] Payload-mode prompt focuses the final report on the supplied alert.
|
||||
- [x] No-payload prompt requires active-alert discovery first.
|
||||
- [x] Targeted tests pass.
|
||||
- [x] Compile passes.
|
||||
- [x] OpenSpec validates.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- Prompt-only scope control may still require runtime observation.
|
||||
- AIOps Verifier remains deferred.
|
||||
@@ -0,0 +1,28 @@
|
||||
# Brief: aiops-alert-scope-control
|
||||
|
||||
## Background
|
||||
|
||||
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
|
||||
|
||||
## Goal
|
||||
|
||||
Make AIOps scope explicit:
|
||||
|
||||
- Payload present -> targeted diagnosis for the supplied alert.
|
||||
- Payload absent -> automatic active-alert discovery and diagnosis.
|
||||
|
||||
## Scope
|
||||
|
||||
- In scope:
|
||||
- `AiOpsService.buildTaskPrompt(...)` scope rules.
|
||||
- Focused tests.
|
||||
- Demo acceptance wording.
|
||||
- Out of scope:
|
||||
- Verifier integration.
|
||||
- Java-side filtering of tool results.
|
||||
- API shape changes.
|
||||
- Database changes.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/aiops-alert-scope-control/`
|
||||
@@ -0,0 +1,42 @@
|
||||
# Decisions: aiops-alert-scope-control
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
|
||||
- Slug: `aiops-alert-scope-control`
|
||||
- Scale: standard-light.
|
||||
|
||||
## Context
|
||||
|
||||
- AIOps traceability is implemented and verified.
|
||||
- Mock Prometheus returns multiple active alerts.
|
||||
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
|
||||
|
||||
## Grill Question Pool
|
||||
|
||||
| # | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
|
||||
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
|
||||
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
|
||||
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
|
||||
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
|
||||
|
||||
## Evidence-Driven Conclusions
|
||||
|
||||
| Conclusion | Evidence Source | Result |
|
||||
|---|---|---|
|
||||
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
|
||||
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
|
||||
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
|
||||
|
||||
## GitNexus
|
||||
|
||||
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Payload mode is detected when any alert field is present.
|
||||
- Payload mode final report must focus on the supplied alert.
|
||||
- No-payload mode must first call `queryPrometheusAlerts`.
|
||||
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
|
||||
@@ -0,0 +1,44 @@
|
||||
# Evidence: aiops-alert-scope-control
|
||||
|
||||
## Local Impact Analysis
|
||||
|
||||
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
|
||||
- No controller, DTO, repository, or database changes are required.
|
||||
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
|
||||
|
||||
## Verification Results
|
||||
|
||||
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
|
||||
- `mvn -q -DskipTests compile` passed.
|
||||
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
|
||||
|
||||
## Runtime Verification
|
||||
|
||||
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
|
||||
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
|
||||
- `diagnosis_session` persisted:
|
||||
- `agent_flow = AI_OPS`
|
||||
- `status = SUCCESS`
|
||||
- `total_duration_ms = 69875`
|
||||
- `step_count = 5`
|
||||
- `tool_call_count = 8`
|
||||
- Tool invocation counts:
|
||||
- `query_metrics = 1`
|
||||
- `lookup_knowledge = 1`
|
||||
- `query_logs = 6`
|
||||
- Report scope check:
|
||||
- `告警根因分析 - HighCPUUsage` exists.
|
||||
- `告警根因分析 - HighMemoryUsage` does not exist.
|
||||
- `告警根因分析 - SlowResponse` does not exist.
|
||||
- `相关风险告警` exists.
|
||||
|
||||
## Runtime Fix
|
||||
|
||||
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
|
||||
- `maximum-pool-size: 5`
|
||||
- `minimum-idle: 1`
|
||||
- `connection-timeout: 10000`
|
||||
- `validation-timeout: 5000`
|
||||
- `idle-timeout: 60000`
|
||||
- `max-lifetime: 120000`
|
||||
- `keepalive-time: 30000`
|
||||
@@ -0,0 +1,18 @@
|
||||
# Acceptance: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Verification
|
||||
|
||||
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
|
||||
- [x] Targeted AIOps service tests pass.
|
||||
- [x] Compile verification passes.
|
||||
- [x] Demo docs describe AIOps request -> session id -> trace query.
|
||||
|
||||
## Result
|
||||
|
||||
Accepted for implementation scope.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- AIOps Verifier integration is deferred.
|
||||
- Runtime still depends on configured model and infrastructure.
|
||||
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
|
||||
@@ -0,0 +1,28 @@
|
||||
# Brief: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Background
|
||||
|
||||
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
|
||||
|
||||
## Goal
|
||||
|
||||
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
|
||||
|
||||
## Scope
|
||||
|
||||
- In scope:
|
||||
- Optional AIOps alert request body.
|
||||
- Stable request/session id propagation.
|
||||
- Persisted AIOps query summary and final answer.
|
||||
- SSE session id event.
|
||||
- Demo documentation and focused tests.
|
||||
- Out of scope:
|
||||
- Full AIOps and ChatService unification.
|
||||
- AIOps Verifier integration.
|
||||
- Database schema changes.
|
||||
- Sensitive configuration cleanup.
|
||||
- Fully offline runtime.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/aiops-traceable-diagnosis-entry/`
|
||||
@@ -0,0 +1,68 @@
|
||||
# Decisions: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
|
||||
- Slug: `aiops-traceable-diagnosis-entry`
|
||||
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
|
||||
|
||||
## Context
|
||||
|
||||
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
|
||||
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
|
||||
|
||||
## Grill Question Pool
|
||||
|
||||
| # | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
|
||||
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
|
||||
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
|
||||
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
|
||||
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
|
||||
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
|
||||
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
|
||||
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
|
||||
|
||||
## Evidence-Driven Conclusions
|
||||
|
||||
| Conclusion | Evidence Source | Result |
|
||||
|---|---|---|
|
||||
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
|
||||
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
|
||||
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
|
||||
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
|
||||
|
||||
## User-Interview Confirmations
|
||||
|
||||
| Topic | User Words | Decision |
|
||||
|---|---|---|
|
||||
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
|
||||
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
|
||||
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Keep `/api/ai_ops` and make its body optional.
|
||||
- Emit `SseMessage.type=session` before long-running analysis starts.
|
||||
- Store AIOps request summary in `diagnosis_session.query`.
|
||||
- Store final report in `diagnosis_session.answer`.
|
||||
- Defer AIOps Verifier to a later change so this slice stays focused.
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> optional AIOpsRequest
|
||||
-> resolve sessionId
|
||||
-> create diagnosis_session(agentFlow=AI_OPS)
|
||||
-> run ai_ops_supervisor(planner, executor)
|
||||
-> AgentLoggingHook persists steps
|
||||
-> tools persist invocations under SessionContextHolder
|
||||
-> extract final report
|
||||
-> persist answer
|
||||
-> GET /api/diagnosis/{sessionId}/trace replays the run
|
||||
```
|
||||
|
||||
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
|
||||
@@ -0,0 +1,37 @@
|
||||
# Evidence: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Local Impact Analysis
|
||||
|
||||
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
|
||||
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
|
||||
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
|
||||
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
|
||||
|
||||
## GitNexus
|
||||
|
||||
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
|
||||
|
||||
## Expected Verification
|
||||
|
||||
- Focused unit tests for AIOps request/session/report helper behavior.
|
||||
- Compile verification.
|
||||
- OpenSpec validation if CLI is available.
|
||||
|
||||
## Verification Results
|
||||
|
||||
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
|
||||
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
|
||||
## Demo Alignment
|
||||
|
||||
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
|
||||
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
|
||||
|
||||
## Metric Alignment Follow-up
|
||||
|
||||
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
|
||||
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
|
||||
- Targeted verification:
|
||||
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
@@ -0,0 +1,36 @@
|
||||
# Acceptance: diagnosis-eval-harness
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | `ISS-006` and initial OpenSpec artifacts were created. |
|
||||
| Implementation | Done | Added fixed cases, fixture-mode trace evaluation, aggregate metrics, and JSON / Markdown report writer. |
|
||||
| Verification | Done | Targeted evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- First implementation uses fixture-mode evaluation.
|
||||
- Live trace API polling remains a follow-up option.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers fixed case loading, fixture evaluation, missing fixture reporting, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-harness --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: diagnosis-eval-harness
|
||||
|
||||
## Background
|
||||
|
||||
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Define fixed diagnosis cases for the MVP demo domain.
|
||||
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
|
||||
3. Produce JSON and Markdown reports for interview and regression use.
|
||||
4. Keep the first version offline by supporting trace fixtures.
|
||||
|
||||
## Scope
|
||||
|
||||
- Evaluation case definitions
|
||||
- Trace fixture shape
|
||||
- Rule-based evaluator
|
||||
- JSON / Markdown report output
|
||||
- Focused offline tests and docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No LLM-as-judge
|
||||
- No live end-to-end runtime requirement
|
||||
- No production API
|
||||
- No chat or verifier runtime change
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-harness/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Harness Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
|
||||
- Slug: `diagnosis-eval-harness`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
|
||||
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
|
||||
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
|
||||
- Reason: The first regression signal should be deterministic and tied to trace contracts.
|
||||
|
||||
- Decision: Support offline fixture traces first.
|
||||
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
|
||||
@@ -0,0 +1,10 @@
|
||||
# Diagnosis Eval Harness Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `openspec/specs/evidence-trace-hardening/spec.md` | Defines stable evidence states and summary behavior | Evaluation can rely on trace semantics rather than ad hoc log parsing | Yes |
|
||||
| `mvp/demo/README.md` | Documents an end-to-end demo flow with chat, trace, and feedback | Existing demo flow provides the runtime story, but not a reusable evaluation baseline | Yes |
|
||||
| `DiagnosisTraceService` | Aggregates session, steps, tools, and self-evaluation | Trace response shape can be reused as evaluation input | Yes |
|
||||
| `ToolTraceSummaryService` | Builds verifier-facing evidence summaries from persisted tool rows | Evaluator can check evidence coverage through persisted trace artifacts | Yes |
|
||||
@@ -0,0 +1,32 @@
|
||||
# Acceptance: evidence-trace-hardening
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | `ISS-005` and the initial OpenSpec artifacts were created. |
|
||||
| Implementation | Done | Recorder contract, lookup persistence path, evidence summary semantics, and degraded-path tests were implemented. |
|
||||
| Verification | Done | Targeted offline tests and compile verification passed. |
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,LookupKnowledgeToolTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers recorder contract, summary semantics for success/failure/no-evidence, and `ChatService` fallback / degraded paths.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
## Open Questions
|
||||
|
||||
| Question | Current position |
|
||||
| --- | --- |
|
||||
| Should deduped retrievals be counted separately from generic no-hit events in future evaluation metrics? | Deferred to P1-B; this change preserves enough structure to decide later. |
|
||||
@@ -0,0 +1,32 @@
|
||||
# Brief: evidence-trace-hardening
|
||||
|
||||
## Background
|
||||
|
||||
The MVP already has persisted tool traces and a verifier, but the evidence contract is still only partially standardized. For interview-focused hardening, the project now needs a tighter contract for evidence persistence, no-evidence / failure semantics, and degraded-output behavior.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
2. Make verifier-facing summaries distinguish failed calls, no-hit calls, deduped retrievals, and actual supporting evidence.
|
||||
3. Add offline tests for verifier fallback and degraded-output paths.
|
||||
|
||||
## Scope
|
||||
|
||||
- `ToolInvocationRecorder`
|
||||
- `LookupKnowledgeTool`
|
||||
- `QueryLogsTools`
|
||||
- `QueryMetricsTools`
|
||||
- `ToolTraceSummaryService`
|
||||
- `ChatService`
|
||||
- Focused offline tests
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new API or schema
|
||||
- No evaluation harness yet
|
||||
- No trace UI
|
||||
- No security/config cleanup
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/evidence-trace-hardening/`
|
||||
@@ -0,0 +1,24 @@
|
||||
# Evidence Trace Hardening Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: harden the MVP evidence contract before building the P1-B evaluation harness.
|
||||
- Slug: `evidence-trace-hardening`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `ISS-003` raised verifier traceability and failure-path concerns.
|
||||
- Current code inspection shows `QueryLogsTools` and `QueryMetricsTools` already use `ToolInvocationRecorder`, while `LookupKnowledgeTool` still persists rows through a local helper.
|
||||
- `ChatService` already contains fallback behavior for missing/invalid `verifier_output`, but coverage is narrow.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Treat this as a contract-hardening change, not a new feature change.
|
||||
- Reason: The project already has the necessary runtime pieces; the gap is semantic consistency and testability.
|
||||
|
||||
- Decision: Keep the scope before P1-B.
|
||||
- Reason: The evaluation harness will rely on stable evidence semantics, so this contract slice should land first.
|
||||
|
||||
- Decision: Preserve schema and API stability.
|
||||
- Reason: The interview value here is engineering rigor, not more surface area.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence Trace Hardening Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `ToolInvocationRecorder` | Provides a common persistence seam for evidence tools | Contract hardening should build on the existing recorder instead of introducing a new store path | Yes |
|
||||
| `LookupKnowledgeTool` | Still constructs `ToolInvocation` rows through a local helper | Retrieval-aware evidence persistence is not yet unified with the recorder contract | Yes |
|
||||
| `QueryLogsTools` / `QueryMetricsTools` | Already record evidence invocations through `recordEvidenceTool(...)` | Current gap is semantic alignment, not missing persistence | Yes |
|
||||
| `ToolTraceSummaryService` | Merges rows by tool and topic domain and infers evidence level heuristically | Summary rules need explicit handling for failure, no-hit, and dedup cases | Yes |
|
||||
| `ChatService` | Falls back to `LOW_CONFID` when verifier output is missing or invalid | These degraded paths exist and should now be covered by focused offline tests | Yes |
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added remaining fixtures, full baseline reports, and documentation updates. |
|
||||
| Verification | Done | Evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Fixture coverage is complete for the five fixed diagnosis cases.
|
||||
- Baseline reports are saved under `mvp/eval/reports`.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers full fixture coverage, baseline report matching, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate expand-diagnosis-eval-fixtures --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Background
|
||||
|
||||
The diagnosis eval harness is implemented and archived, but the fixed baseline is incomplete because three of the five diagnosis cases still reference missing fixtures.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Add representative trace fixtures for all remaining fixed diagnosis cases.
|
||||
2. Save a reproducible baseline report in JSON and Markdown.
|
||||
3. Document how to regenerate and interpret the baseline.
|
||||
4. Keep evaluation offline and deterministic.
|
||||
|
||||
## Scope
|
||||
|
||||
- Redis timeout fixture
|
||||
- Slow response fixture
|
||||
- JVM memory risk fixture
|
||||
- Baseline reports under `mvp/eval/reports`
|
||||
- Focused tests for full fixture coverage and report generation
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new diagnosis cases
|
||||
- No production Agent runtime changes
|
||||
- No LLM-as-judge
|
||||
- No live infrastructure requirement
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/expand-diagnosis-eval-fixtures/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Expand Diagnosis Eval Fixtures Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: complete the fixed diagnosis eval baseline after the harness is in place.
|
||||
- Slug: `expand-diagnosis-eval-fixtures`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created the evaluator, case file, fixture mode, and report writer.
|
||||
- The first baseline still has missing fixtures by design.
|
||||
- This follow-up turns that partial baseline into a full fixed-case baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Keep this change data-focused.
|
||||
- Reason: the evaluator rules already landed; this change should not blur fixture expansion with harness behavior changes.
|
||||
|
||||
- Decision: Save baseline reports in the repository.
|
||||
- Reason: interview review and future diffs are easier when the expected baseline is visible.
|
||||
|
||||
- Decision: Use deterministic fixture traces instead of live trace generation.
|
||||
- Reason: this baseline should run without infrastructure or external model calls.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should add a CLI or Maven goal for report regeneration.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-04: Created slug-based issue `expand-diagnosis-eval-fixtures.md`.
|
||||
- 2026-07-04: Created OpenSpec change `expand-diagnosis-eval-fixtures`.
|
||||
- 2026-07-04: Added Redis timeout, slow response, and JVM memory risk fixtures.
|
||||
- 2026-07-04: Added baseline JSON and Markdown reports under `mvp/eval/reports`.
|
||||
- 2026-07-04: Verification passed with `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-04: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-04: Verification passed with `openspec validate expand-diagnosis-eval-fixtures --strict`.
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: diagnosis-eval-baseline-diff
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
|
||||
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
|
||||
- JSON and Markdown diff output are available.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
|
||||
- Result: passed
|
||||
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,30 @@
|
||||
# Brief: diagnosis-eval-baseline-diff
|
||||
|
||||
## Background
|
||||
|
||||
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Compare baseline and current `DiagnosisEvalReport` objects.
|
||||
2. Detect aggregate and per-case regressions.
|
||||
3. Output JSON and Markdown diff reports.
|
||||
4. Document how to read the diff in interview and engineering terms.
|
||||
|
||||
## Scope
|
||||
|
||||
- Diff data structures
|
||||
- Deterministic report comparison
|
||||
- JSON / Markdown diff output
|
||||
- Focused tests and eval docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No live Agent execution
|
||||
- No LLM-as-judge
|
||||
- No evaluator scoring rule changes
|
||||
- No production API changes
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-baseline-diff/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Baseline Diff Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
|
||||
- Slug: `diagnosis-eval-baseline-diff`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created deterministic fixture evaluation.
|
||||
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
|
||||
- This change compares new reports against that baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Diff report DTOs instead of raw traces.
|
||||
- Reason: the report is the stable contract for regression review.
|
||||
|
||||
- Decision: Use deterministic code rules instead of LLM-as-judge.
|
||||
- Reason: baseline regression checks should be repeatable and explainable.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should expose this through a CLI or Maven goal.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: diagnosis-eval-baseline-diff
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
|
||||
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
|
||||
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
|
||||
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
|
||||
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Acceptance: mvp-demo-interview-runbook
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
|
||||
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
|
||||
| Verification | Done | OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- No backend runtime behavior has been changed.
|
||||
- Demo is packaged under `mvp/demo` for interview use.
|
||||
|
||||
## Verification
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate mvp-demo-interview-runbook --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,29 @@
|
||||
# Brief: mvp-demo-interview-runbook
|
||||
|
||||
## Background
|
||||
|
||||
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Provide a fixed payment-timeout request payload.
|
||||
2. Provide a PowerShell script that runs chat, trace, and feedback.
|
||||
3. Save demo responses under `mvp/demo/output`.
|
||||
4. Add interview walkthrough and trace checklist.
|
||||
|
||||
## Scope
|
||||
|
||||
- Demo docs and scripts only
|
||||
- Existing local APIs only
|
||||
- Existing `mvp-demo` profile only
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No backend code changes
|
||||
- No eval extension
|
||||
- No secret cleanup
|
||||
- No full offline runtime
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/mvp-demo-interview-runbook/`
|
||||
@@ -0,0 +1,27 @@
|
||||
# MVP Demo Interview Runbook Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
|
||||
- Slug: `mvp-demo-interview-runbook`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- Evidence trace and eval baseline work are already done.
|
||||
- The next useful step is not more eval tooling, but a runnable demo path.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Keep this change documentation/script-only.
|
||||
- Reason: Plan C is about demo packaging, not new runtime capability.
|
||||
|
||||
- Decision: Use a stable session id.
|
||||
- Reason: it makes trace lookup and saved output predictable.
|
||||
|
||||
- Decision: Save outputs to `mvp/demo/output`.
|
||||
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a later change should add a truly offline stubbed demo mode.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Evidence: mvp-demo-interview-runbook
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
|
||||
- 2026-07-05: Added fixed payment-timeout request payload.
|
||||
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
|
||||
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
|
||||
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
|
||||
@@ -0,0 +1,81 @@
|
||||
---
|
||||
title: AIOps 告警排障 Runbook
|
||||
keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs]
|
||||
summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。
|
||||
category: troubleshooting
|
||||
---
|
||||
|
||||
# AIOps 告警排障 Runbook
|
||||
|
||||
## 1. 告警输入处理原则
|
||||
|
||||
AIOps 诊断入口有两种触发方式:
|
||||
|
||||
- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。
|
||||
- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。
|
||||
|
||||
最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。
|
||||
|
||||
## 2. Mock 告警与日志主题映射
|
||||
|
||||
| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 |
|
||||
|---|---|---|---|
|
||||
| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` |
|
||||
| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` |
|
||||
| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` |
|
||||
| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` |
|
||||
|
||||
## 3. HighCPUUsage / payment-service 排障步骤
|
||||
|
||||
### 3.1 现象确认
|
||||
|
||||
先确认 Prometheus 活动告警中是否存在:
|
||||
|
||||
- `alert_name = HighCPUUsage`
|
||||
- `service = payment-service`
|
||||
- CPU 使用率超过 80%
|
||||
- 状态为 firing
|
||||
|
||||
如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。
|
||||
|
||||
### 3.2 指标与日志取证
|
||||
|
||||
推荐工具调用顺序:
|
||||
|
||||
1. `queryPrometheusAlerts`:确认当前活动告警。
|
||||
2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。
|
||||
3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。
|
||||
|
||||
### 3.3 根因判断
|
||||
|
||||
可接受的根因结论必须至少满足一项:
|
||||
|
||||
- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。
|
||||
- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。
|
||||
- 告警持续时间与日志时间线一致。
|
||||
|
||||
如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。
|
||||
|
||||
## 4. 处理建议
|
||||
|
||||
### 临时止血
|
||||
|
||||
- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。
|
||||
- 对高耗时接口开启限流或降级非核心功能。
|
||||
- 如果近期有发布,检查变更窗口并准备回滚。
|
||||
|
||||
### 根因修复
|
||||
|
||||
- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。
|
||||
- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。
|
||||
- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。
|
||||
|
||||
## 5. 报告要求
|
||||
|
||||
告警分析报告至少包含:
|
||||
|
||||
- 活跃告警清单。
|
||||
- 告警根因分析。
|
||||
- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。
|
||||
- 已执行或建议执行的处理方案。
|
||||
- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。
|
||||
@@ -2,6 +2,13 @@
|
||||
|
||||
This demo proves the MVP flow from user question to persisted diagnosis trace.
|
||||
|
||||
For interview use, start with:
|
||||
|
||||
- `interview-walkthrough.md` for the talk track
|
||||
- `trace-inspection-checklist.md` for fields to inspect
|
||||
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
|
||||
- `requests/payment-timeout-chat.json` for the fixed request payload
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
|
||||
@@ -22,6 +29,22 @@ http://localhost:9900
|
||||
|
||||
## 1. Run Chat Diagnosis
|
||||
|
||||
Fast path:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
This writes:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
Manual path:
|
||||
|
||||
```powershell
|
||||
$sessionId = "mvp-demo-payment-timeout-001"
|
||||
$body = @{
|
||||
@@ -78,6 +101,43 @@ Expected result:
|
||||
- `success` is `true`.
|
||||
- A later trace query shows `data.session.feedback` as `useful`.
|
||||
|
||||
## 4. Run AIOps Alert Diagnosis
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`.
|
||||
- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload.
|
||||
- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`.
|
||||
- `data.session.answer` contains the final alert analysis report when a report is generated.
|
||||
- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them.
|
||||
|
||||
Query the AIOps trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||
```
|
||||
|
||||
## Demo Story
|
||||
|
||||
The important interview story is:
|
||||
@@ -92,3 +152,14 @@ one session id
|
||||
-> feedback
|
||||
-> trace API for replay and audit
|
||||
```
|
||||
|
||||
The AIOps story uses the same audit spine:
|
||||
|
||||
```text
|
||||
one session id
|
||||
-> alert payload
|
||||
-> AIOps planner/executor execution
|
||||
-> evidence tools
|
||||
-> alert analysis report
|
||||
-> trace API for replay and audit
|
||||
```
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
# AIOps Alert Acceptance Case
|
||||
|
||||
## Goal
|
||||
|
||||
Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry.
|
||||
|
||||
## Input
|
||||
|
||||
- Session id: `mvp-demo-aiops-payment-cpu-001`
|
||||
- Endpoint: `POST /api/ai_ops`
|
||||
- Profile: `mvp-demo`
|
||||
- Alert:
|
||||
|
||||
```json
|
||||
{
|
||||
"sessionId": "mvp-demo-aiops-payment-cpu-001",
|
||||
"alertName": "HighCPUUsage",
|
||||
"service": "payment-service",
|
||||
"severity": "P1",
|
||||
"description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。",
|
||||
"timeRange": "last_15m",
|
||||
"userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
}
|
||||
```
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
1. The SSE stream emits a `session` message containing the requested session id.
|
||||
2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`.
|
||||
3. The persisted session query contains the alert name, service, severity, time range, and description.
|
||||
4. If a final report is generated, `diagnosis_session.answer` contains that report.
|
||||
5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations.
|
||||
6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- This slice does not add a Verifier Agent to AIOps.
|
||||
- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration.
|
||||
@@ -0,0 +1,146 @@
|
||||
# Interview Walkthrough: MVP Diagnosis Agent
|
||||
|
||||
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
|
||||
|
||||
## 30-Second Summary
|
||||
|
||||
```text
|
||||
This is an enterprise diagnosis Agent MVP.
|
||||
It takes a payment-timeout question, plans the investigation, calls evidence tools,
|
||||
checks the answer through a verifier, persists the full trace, and accepts feedback.
|
||||
```
|
||||
|
||||
The important claim is not "the model answered once." The claim is:
|
||||
|
||||
```text
|
||||
The system can show what evidence was used, how the answer was checked, and how to replay the session.
|
||||
```
|
||||
|
||||
## Demo Flow
|
||||
|
||||
1. Start the service with the `mvp-demo` profile.
|
||||
2. Run the fixed payment-timeout request.
|
||||
3. Open `mvp/demo/output/chat-response.json`.
|
||||
4. Open `mvp/demo/output/trace-response.json`.
|
||||
5. Point to evidence tools and verifier evaluation.
|
||||
6. Submit feedback and show it is attached to the same session.
|
||||
|
||||
## Commands
|
||||
|
||||
Start service:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
Run the demo from another terminal:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
Optional custom session:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
||||
```
|
||||
|
||||
## What To Show
|
||||
|
||||
### 1. User-Facing Answer
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
|
||||
```
|
||||
|
||||
### 2. Evidence Trace
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/trace-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the important Agent engineering part.
|
||||
I can inspect which tools were called, what inputs they received,
|
||||
whether they succeeded, and what evidence preview was persisted.
|
||||
```
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.toolInvocations[*].toolName`
|
||||
- `data.toolInvocations[*].inputParams`
|
||||
- `data.toolInvocations[*].outputPreview`
|
||||
- `data.toolInvocations[*].success`
|
||||
|
||||
### 3. Verifier / Self-Evaluation
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.session.selfEvaluation`
|
||||
- `data.summary.hasVerifierEvaluation`
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
The final answer is not just raw Executor output.
|
||||
It is checked by a verifier or self-evaluation layer using the persisted trace.
|
||||
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
|
||||
```
|
||||
|
||||
### 4. Feedback Loop
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
Then re-query trace if needed.
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
Feedback is attached to the same diagnosis session.
|
||||
That makes it possible to mine useful / not useful cases later.
|
||||
```
|
||||
|
||||
### 5. Regression Story
|
||||
|
||||
Mention, do not deep dive unless asked:
|
||||
|
||||
```text
|
||||
For repeatability, I also built an offline eval baseline.
|
||||
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
|
||||
The two are separate on purpose: demo for human review, eval for automated signal.
|
||||
```
|
||||
|
||||
## Strong Interview Framing
|
||||
|
||||
Use this phrasing:
|
||||
|
||||
```text
|
||||
I focused on the Agent engineering surface:
|
||||
traceability, evidence persistence, verifier gating, feedback, and regression checks.
|
||||
The model answer is only one part of the system.
|
||||
The more important part is whether we can audit and improve the answer after it is produced.
|
||||
```
|
||||
|
||||
## Known Limits To Say Proactively
|
||||
|
||||
```text
|
||||
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
|
||||
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
|
||||
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
|
||||
```
|
||||
@@ -0,0 +1,11 @@
|
||||
# Demo Output
|
||||
|
||||
This directory is the default output location for local demo responses.
|
||||
|
||||
Generated files are intentionally ignored by Git:
|
||||
|
||||
- `chat-response.json`
|
||||
- `trace-response.json`
|
||||
- `feedback-response.json`
|
||||
|
||||
Keep this README so the directory exists in the repository.
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-payment-timeout-001",
|
||||
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
param(
|
||||
[string]$BaseUrl = "http://localhost:9900",
|
||||
[string]$SessionId = "mvp-demo-payment-timeout-001",
|
||||
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
|
||||
[string]$OutputDir = "$PSScriptRoot/../output"
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||
|
||||
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||
$request.Id = $SessionId
|
||||
$body = $request | ConvertTo-Json -Depth 8
|
||||
|
||||
Write-Host "Running payment-timeout chat demo..."
|
||||
Write-Host "BaseUrl: $BaseUrl"
|
||||
Write-Host "SessionId: $SessionId"
|
||||
|
||||
$chat = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "$BaseUrl/api/chat" `
|
||||
-ContentType "application/json; charset=utf-8" `
|
||||
-Body $body
|
||||
|
||||
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||
Write-Host "Saved chat response: $OutputDir/chat-response.json"
|
||||
|
||||
$trace = Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||
|
||||
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||
Write-Host "Saved trace response: $OutputDir/trace-response.json"
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
$feedback = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "$BaseUrl/api/feedback" `
|
||||
-ContentType "application/json; charset=utf-8" `
|
||||
-Body $feedbackBody
|
||||
|
||||
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
|
||||
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "Demo completed. Review:"
|
||||
Write-Host "- mvp/demo/output/chat-response.json"
|
||||
Write-Host "- mvp/demo/output/trace-response.json"
|
||||
Write-Host "- mvp/demo/output/feedback-response.json"
|
||||
@@ -0,0 +1,52 @@
|
||||
# Trace Inspection Checklist
|
||||
|
||||
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
|
||||
|
||||
## Session
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
|
||||
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
|
||||
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
|
||||
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
|
||||
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
|
||||
|
||||
## Agent Steps
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
|
||||
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
|
||||
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
|
||||
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
|
||||
|
||||
## Tool Evidence
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
|
||||
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
|
||||
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
|
||||
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
|
||||
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
|
||||
|
||||
## Summary
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
|
||||
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
|
||||
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
|
||||
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
|
||||
|
||||
## What Good Looks Like
|
||||
|
||||
```text
|
||||
same session id
|
||||
-> final answer
|
||||
-> persisted agent steps
|
||||
-> persisted evidence tool calls
|
||||
-> verifier/self-evaluation
|
||||
-> feedback attached to the same session
|
||||
```
|
||||
@@ -0,0 +1,61 @@
|
||||
# Diagnosis Eval Harness
|
||||
|
||||
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
|
||||
|
||||
## Scope
|
||||
|
||||
- Case definitions: `cases/diagnosis-cases.json`
|
||||
- Offline trace fixtures: `fixtures/*.json`
|
||||
- Field definitions: `schema.md`
|
||||
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
|
||||
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
|
||||
- Evaluator implementation: `DiagnosisTraceEvaluator`
|
||||
- Report writer: `DiagnosisEvalReportWriter`
|
||||
|
||||
## Current Mode
|
||||
|
||||
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
## Verification
|
||||
|
||||
Run the focused evaluator test:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||
```
|
||||
|
||||
The committed baseline report represents the current fixed fixture set:
|
||||
|
||||
```text
|
||||
5 fixed cases
|
||||
5 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
3 LOW_CONFID verdicts
|
||||
```
|
||||
|
||||
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
||||
|
||||
## Interview Story
|
||||
|
||||
The harness gives the MVP a repeatable baseline:
|
||||
|
||||
```text
|
||||
fixed diagnosis case
|
||||
-> saved or runtime trace
|
||||
-> rule-based trace validation
|
||||
-> JSON / Markdown report
|
||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
||||
```
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||
|
||||
```text
|
||||
baseline report
|
||||
current report
|
||||
-> deterministic diff
|
||||
-> regressions, improvements, and changed signals
|
||||
```
|
||||
|
||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
||||
@@ -0,0 +1,57 @@
|
||||
[
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"traceFixture": "payment-timeout-pass.json",
|
||||
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
|
||||
"allowedVerdicts": ["PASS", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["无证据确定"]
|
||||
},
|
||||
{
|
||||
"id": "mysql-pool-exhausted",
|
||||
"title": "MySQL connection pool exhausted",
|
||||
"question": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
|
||||
"traceFixture": "mysql-pool-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["mysql", "连接池", "超时"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["已经完全确认"]
|
||||
},
|
||||
{
|
||||
"id": "redis-timeout",
|
||||
"title": "Redis timeout",
|
||||
"question": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
||||
"traceFixture": "redis-timeout-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["redis", "超时"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["无需进一步排查"]
|
||||
},
|
||||
{
|
||||
"id": "slow-response",
|
||||
"title": "Slow response",
|
||||
"question": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
||||
"traceFixture": "slow-response-pass.json",
|
||||
"expectedRootCauseKeywords": ["p99", "慢响应"],
|
||||
"minKeywordMatches": 1,
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["没有风险"]
|
||||
},
|
||||
{
|
||||
"id": "jvm-memory-risk",
|
||||
"title": "JVM memory risk",
|
||||
"question": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||
"traceFixture": "jvm-memory-risk-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["jvm", "内存", "oom"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["可以忽略"]
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 53000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n订单服务存在 JVM 内存风险,但还不能完全确认会发生 OOM。指标显示 heap 使用率持续高于 88%,日志出现多次 Full GC 和 allocation pressure,需要继续观察对象增长来源并检查最近发布。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.52,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "indirect"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,53 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 51000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nMySQL 连接池可能参与了本次超时问题。日志中出现 connection pool exhausted,但当前缺少完整指标证据,因此只能作为低置信结论处理。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.48,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "lookup_knowledge",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"toolName": "lookup_knowledge",
|
||||
"success": true,
|
||||
"relevanceLevel": "PRECISE"
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,64 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 42000,
|
||||
"toolCallCount": 3,
|
||||
"answer": "支付接口超时与连接池等待有关。知识库说明支付超时需要同时检查连接池、日志和指标;日志出现 connection pool exhausted;指标显示支付服务延迟升高。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.86,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "lookup_knowledge",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"toolName": "lookup_knowledge",
|
||||
"success": true,
|
||||
"relevanceLevel": "PRECISE"
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 3,
|
||||
"returnedToolCallCount": 3,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,41 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-redis-timeout",
|
||||
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 36000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nRedis 连接超时可能和支付服务到 Redis 的网络抖动或连接池等待有关。日志中出现 redis timeout 和 command timeout 记录,但当前缺少指标侧证据,因此只能作为低置信结论处理。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.46,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-redis-timeout",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 2,
|
||||
"returnedStepCount": 2,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-slow-response",
|
||||
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 47000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "用户服务 P99 升高主要表现为慢响应。指标显示 P99 latency 从 280ms 上升到 1800ms,日志中同时出现 slow request 和 downstream timeout,因此优先排查下游依赖耗时和线程池排队。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.78,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-slow-response",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-slow-response",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,85 @@
|
||||
{
|
||||
"baselineTotalCases" : 5,
|
||||
"currentTotalCases" : 5,
|
||||
"baselinePassedCases" : 5,
|
||||
"currentPassedCases" : 4,
|
||||
"baselinePassRate" : 1.0,
|
||||
"currentPassRate" : 0.8,
|
||||
"regressionCount" : 6,
|
||||
"improvementCount" : 0,
|
||||
"changedCount" : 2,
|
||||
"hasRegression" : true,
|
||||
"items" : [ {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "passRate",
|
||||
"baselineValue" : "1.0",
|
||||
"currentValue" : "0.8",
|
||||
"delta" : -0.19999999999999996,
|
||||
"message" : "passRate changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "averageToolCallCount",
|
||||
"baselineValue" : "2.0",
|
||||
"currentValue" : "3.0",
|
||||
"delta" : 1.0,
|
||||
"message" : "averageToolCallCount changed"
|
||||
}, {
|
||||
"type" : "CHANGED",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "verdictDistribution.LOW_CONFID",
|
||||
"baselineValue" : "3",
|
||||
"currentValue" : "2",
|
||||
"delta" : -1.0,
|
||||
"message" : "verdict count changed for LOW_CONFID"
|
||||
}, {
|
||||
"type" : "CHANGED",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "verdictDistribution.REJECT",
|
||||
"baselineValue" : "0",
|
||||
"currentValue" : "1",
|
||||
"delta" : 1.0,
|
||||
"message" : "verdict count changed for REJECT"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "passed",
|
||||
"baselineValue" : "true",
|
||||
"currentValue" : "false",
|
||||
"delta" : null,
|
||||
"message" : "redis-timeout pass state changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "verdict",
|
||||
"baselineValue" : "LOW_CONFID",
|
||||
"currentValue" : "REJECT",
|
||||
"delta" : -1.0,
|
||||
"message" : "redis-timeout verdict changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "matchedKeywordCount",
|
||||
"baselineValue" : "2",
|
||||
"currentValue" : "1",
|
||||
"delta" : -1.0,
|
||||
"message" : "redis-timeout matchedKeywordCount changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "evidenceCoverage.query_logs",
|
||||
"baselineValue" : "true",
|
||||
"currentValue" : "false",
|
||||
"delta" : null,
|
||||
"message" : "redis-timeout evidence coverage changed for query_logs"
|
||||
} ]
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
# Diagnosis Eval Baseline Diff
|
||||
|
||||
- Baseline pass rate: 100.00%
|
||||
- Current pass rate: 80.00%
|
||||
- Baseline passed cases: 5/5
|
||||
- Current passed cases: 4/5
|
||||
- Regressions: 6
|
||||
- Improvements: 0
|
||||
- Other changes: 2
|
||||
|
||||
## Diff Items
|
||||
|
||||
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
|
||||
| --- | --- | --- | --- | --- | --- | ---: | --- |
|
||||
| REGRESSION | aggregate | - | passRate | 1.0 | 0.8 | -0.200 | passRate changed |
|
||||
| REGRESSION | aggregate | - | averageToolCallCount | 2.0 | 3.0 | 1.000 | averageToolCallCount changed |
|
||||
| CHANGED | aggregate | - | verdictDistribution.LOW_CONFID | 3 | 2 | -1.000 | verdict count changed for LOW_CONFID |
|
||||
| CHANGED | aggregate | - | verdictDistribution.REJECT | 0 | 1 | 1.000 | verdict count changed for REJECT |
|
||||
| REGRESSION | case | redis-timeout | passed | true | false | - | redis-timeout pass state changed |
|
||||
| REGRESSION | case | redis-timeout | verdict | LOW_CONFID | REJECT | -1.000 | redis-timeout verdict changed |
|
||||
| REGRESSION | case | redis-timeout | matchedKeywordCount | 2 | 1 | -1.000 | redis-timeout matchedKeywordCount changed |
|
||||
| REGRESSION | case | redis-timeout | evidenceCoverage.query_logs | true | false | - | redis-timeout evidence coverage changed for query_logs |
|
||||
@@ -0,0 +1,82 @@
|
||||
{
|
||||
"totalCases" : 5,
|
||||
"passedCases" : 5,
|
||||
"passRate" : 1.0,
|
||||
"verdictDistribution" : {
|
||||
"PASS" : 2,
|
||||
"LOW_CONFID" : 3
|
||||
},
|
||||
"averageToolCallCount" : 2.0,
|
||||
"averageDurationMs" : 45800.0,
|
||||
"results" : [ {
|
||||
"caseId" : "payment-timeout",
|
||||
"title" : "Payment API timeout",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true,
|
||||
"query_metrics" : true
|
||||
},
|
||||
"toolCallCount" : 3,
|
||||
"durationMs" : 42000
|
||||
}, {
|
||||
"caseId" : "mysql-pool-exhausted",
|
||||
"title" : "MySQL connection pool exhausted",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 51000
|
||||
}, {
|
||||
"caseId" : "redis-timeout",
|
||||
"title" : "Redis timeout",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 36000
|
||||
}, {
|
||||
"caseId" : "slow-response",
|
||||
"title" : "Slow response",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 47000
|
||||
}, {
|
||||
"caseId" : "jvm-memory-risk",
|
||||
"title" : "JVM memory risk",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 53000
|
||||
} ]
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
# Diagnosis Eval Report
|
||||
|
||||
- Total cases: 5
|
||||
- Passed cases: 5
|
||||
- Pass rate: 100.00%
|
||||
- Average tool calls: 2.00
|
||||
- Average duration ms: 45800.00
|
||||
|
||||
## Verdict Distribution
|
||||
|
||||
- PASS: 2
|
||||
- LOW_CONFID: 3
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
|
||||
@@ -0,0 +1,202 @@
|
||||
# Diagnosis Eval Data Schema
|
||||
|
||||
这份文档记录评测基准里的数据结构。口语化理解就是:
|
||||
|
||||
```text
|
||||
用例文件说“我要考什么”
|
||||
trace 文件说“Agent 实际做了什么”
|
||||
评测结果说“这次有没有跑偏”
|
||||
汇总报告说“整体稳定性怎么样”
|
||||
```
|
||||
|
||||
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。
|
||||
|
||||
## 1. 用例定义
|
||||
|
||||
文件:`mvp/eval/cases/diagnosis-cases.json`
|
||||
|
||||
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"traceFixture": "payment-timeout-pass.json",
|
||||
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
|
||||
"allowedVerdicts": ["PASS", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["无证据确定"]
|
||||
}
|
||||
```
|
||||
|
||||
字段说明:
|
||||
|
||||
| 字段 | 意思 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 |
|
||||
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 |
|
||||
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 |
|
||||
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 |
|
||||
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 |
|
||||
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 |
|
||||
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 |
|
||||
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 |
|
||||
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 |
|
||||
|
||||
## 2. Trace Fixture
|
||||
|
||||
目录:`mvp/eval/fixtures/*.json`
|
||||
|
||||
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。
|
||||
|
||||
当前会读取这些字段:
|
||||
|
||||
| Trace 字段 | 意思 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 |
|
||||
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 |
|
||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 |
|
||||
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 |
|
||||
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 |
|
||||
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 |
|
||||
|
||||
简单说,trace 里最重要的是三类信息:
|
||||
|
||||
```text
|
||||
最终回答:它说了什么
|
||||
工具证据:它查了什么
|
||||
Verifier:它自己有没有承认这个结论可靠
|
||||
```
|
||||
|
||||
## 3. 单条评测结果
|
||||
|
||||
Java 类型:`DiagnosisEvalResult`
|
||||
|
||||
这是每条 case 跑完之后的判断结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `caseId` | 对应的 case id |
|
||||
| `title` | case 标题 |
|
||||
| `passed` | 这条 case 是否通过 |
|
||||
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 |
|
||||
| `verdict` | 从 trace 里读出来的 Verifier verdict |
|
||||
| `matchedKeywordCount` | 最终回答命中的关键词数量 |
|
||||
| `requiredKeywordCount` | case 定义里一共有多少个关键词 |
|
||||
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` |
|
||||
| `toolCallCount` | 本次 trace 里工具调用总数 |
|
||||
| `durationMs` | 本次 trace 的耗时 |
|
||||
|
||||
判断通过的口语化规则:
|
||||
|
||||
```text
|
||||
回答要说到关键点
|
||||
该查的证据工具要查到
|
||||
Verifier 的结论要在可接受范围内
|
||||
回答不能出现危险的过度自信表达
|
||||
如果是 REJECT,就必须走降级模板
|
||||
```
|
||||
|
||||
## 4. 汇总报告
|
||||
|
||||
Java 类型:`DiagnosisEvalReport`
|
||||
|
||||
这是整个基准集跑完之后的总结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `totalCases` | 总共评测了多少条 case |
|
||||
| `passedCases` | 通过了多少条 |
|
||||
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` |
|
||||
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` |
|
||||
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 |
|
||||
| `averageDurationMs` | 平均耗时 |
|
||||
| `results` | 每条 case 的详细结果列表 |
|
||||
|
||||
## 5. 怎么看这个基准
|
||||
|
||||
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答:
|
||||
|
||||
```text
|
||||
以前能过的诊断题,现在还过不过?
|
||||
它是不是少查了某些证据?
|
||||
它是不是变得更自信但证据不足?
|
||||
它是不是开始输出不该说的话?
|
||||
它是不是明显变慢了?
|
||||
```
|
||||
|
||||
所以面试里可以这样讲:
|
||||
|
||||
```text
|
||||
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。
|
||||
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。
|
||||
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
|
||||
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
|
||||
```
|
||||
|
||||
## 6. Baseline Diff
|
||||
|
||||
Baseline diff 是拿两份 report 做对比:
|
||||
|
||||
```text
|
||||
baseline report:以前认可的基准结果
|
||||
current report:这次改动后跑出来的新结果
|
||||
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
|
||||
```
|
||||
|
||||
Java 类型:
|
||||
|
||||
- `DiagnosisEvalDiffReport`
|
||||
- `DiagnosisEvalDiffItem`
|
||||
|
||||
`DiagnosisEvalDiffReport` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `baselineTotalCases` | baseline 里有多少条 case |
|
||||
| `currentTotalCases` | current 里有多少条 case |
|
||||
| `baselinePassedCases` | baseline 通过了多少条 |
|
||||
| `currentPassedCases` | current 通过了多少条 |
|
||||
| `baselinePassRate` | baseline 通过率 |
|
||||
| `currentPassRate` | current 通过率 |
|
||||
| `regressionCount` | 退化项数量 |
|
||||
| `improvementCount` | 改善项数量 |
|
||||
| `changedCount` | 普通变化项数量 |
|
||||
| `hasRegression` | 是否存在退化 |
|
||||
| `items` | 具体 diff 明细 |
|
||||
|
||||
`DiagnosisEvalDiffItem` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
|
||||
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
|
||||
| `caseId` | 如果是单条 case 变化,这里记录 case id |
|
||||
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
|
||||
| `baselineValue` | baseline 里的值 |
|
||||
| `currentValue` | current 里的值 |
|
||||
| `delta` | 数值变化量;非数值变化为空 |
|
||||
| `message` | 给人看的变化说明 |
|
||||
|
||||
口语化判断规则:
|
||||
|
||||
```text
|
||||
pass rate 下降:退化
|
||||
case 从通过变失败:退化
|
||||
证据工具从有变没有:退化
|
||||
关键词命中变少:退化
|
||||
工具调用或耗时升高:成本上升,记为退化信号
|
||||
verdict 分布变化:记录变化,供人工判断是否符合预期
|
||||
```
|
||||
|
||||
面试里可以这样讲:
|
||||
|
||||
```text
|
||||
我把 baseline report 和当前 report 做结构化 diff。
|
||||
它不是再问 LLM,而是用代码比较固定字段。
|
||||
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
|
||||
diff 会直接标成 regression。
|
||||
这样 Agent 改动可以用固定基准做回归判断。
|
||||
```
|
||||
@@ -0,0 +1,137 @@
|
||||
# ISS-005 证据链补齐与降级契约收敛
|
||||
|
||||
**状态**:进行中(sm-flow)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-A 面试打磨项 / 基于 ISS-003 的当前实现复核
|
||||
**关联**:ISS-003(Verifier 证据链、失败路径可验证性)、`chat-verifier-agent`、`mvp-demo-trace-acceptance`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前 MVP 已具备:
|
||||
|
||||
- `lookup_knowledge`、`query_logs`、`query_metrics` 的工具调用落库
|
||||
- Verifier 基于 `tool_trace_summary` 做事实核查
|
||||
- `LOW_CONFID` / `REJECT` 的用户侧降级输出
|
||||
- trace API 可回放 session、agent_step、tool_invocation 和 self_evaluation
|
||||
|
||||
但如果目标是拿这个项目去面试 Agent 工程师,当前实现仍有一个明显短板:
|
||||
|
||||
**证据链已经“有了”,但还没有被收敛成清晰、稳定、可测试的工程契约。**
|
||||
|
||||
这会直接影响三个面试问题的回答质量:
|
||||
|
||||
1. 工具失败时系统会怎样降级?
|
||||
2. Verifier 看到的 evidence 到底是否一致、可审计?
|
||||
3. 这些失败路径和降级行为有没有稳定测试,而不是只靠 runtime 演示?
|
||||
|
||||
---
|
||||
|
||||
## 当前现状复核
|
||||
|
||||
### 1. 工具落库入口已经存在,但契约不统一
|
||||
|
||||
- `QueryLogsTools` 和 `QueryMetricsTools` 通过 `ToolInvocationRecorder.recordEvidenceTool(...)` 记录 evidence tool 调用。
|
||||
- `LookupKnowledgeTool` 仍保留独立的 `saveToolInvocation(...)` 路径,自己构造 `ToolInvocation` 实体。
|
||||
|
||||
这意味着:
|
||||
|
||||
- evidence tool 的公共字段有一套约定
|
||||
- knowledge retrieval 又有一套定制字段拼装
|
||||
|
||||
两者都能工作,但**没有形成统一的“证据调用记录契约”**。
|
||||
|
||||
### 2. 失败 / 无结果 / 去重命中的语义不够显式
|
||||
|
||||
当前实现里:
|
||||
|
||||
- `query_logs` 未命中时会返回 `success=false` + `"未找到匹配的日志"`
|
||||
- `query_metrics` 失败时会返回 `success=false`
|
||||
- `lookup_knowledge` 去重命中时会返回 `found=false`,但 `tool_invocation.success=true`
|
||||
- `ToolTraceSummaryService` 通过 `success`、`relevanceLevel`、`dedupReason` 等字段做启发式摘要
|
||||
|
||||
这些行为在代码里是分散成立的,但**没有被定义成统一契约**,导致:
|
||||
|
||||
- Verifier 能看到的“失败”和“无证据”边界不够稳定
|
||||
- 评测时难以明确统计哪些是“调用失败”、哪些是“无命中”、哪些是“已检索过”
|
||||
|
||||
### 3. ChatService 的降级路径有实现,但测试矩阵不完整
|
||||
|
||||
`ChatService` 已处理:
|
||||
|
||||
- `verifier_output` 缺失或无法解析 → fallback `LOW_CONFID`
|
||||
- `REJECT` → degraded output
|
||||
- `LOW_CONFID` → disclaimer output
|
||||
|
||||
但目前缺少成体系的专项验证,尤其是:
|
||||
|
||||
- Verifier 输出非法 JSON
|
||||
- evidence tool 查询失败
|
||||
- knowledge lookup 无有效证据
|
||||
- fallback 文案是否只基于 verifier 缺口拼装
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- **面试表达弱化**:你能讲“我有 trace”,但还不能很硬地讲“我的失败路径是有契约和测试保护的”。
|
||||
- **评测基础不稳**:后续 P1-B 做 case-based harness 时,统计口径会受 evidence 语义不一致影响。
|
||||
- **Verifier 可审计性打折**:当前实现可用,但 still relies on code convention,而不是一份明确收敛后的工程协议。
|
||||
|
||||
---
|
||||
|
||||
## 本 issue 目标
|
||||
|
||||
P1-A 只做三件事:
|
||||
|
||||
1. 收敛 evidence tool 的落库契约,让 `lookup_knowledge`、`query_logs`、`query_metrics` 的公共语义一致。
|
||||
2. 明确失败 / 无证据 / 去重 / verifier 非法输出等降级契约,让 `ToolTraceSummaryService` 和 `ChatService` 面向统一状态工作。
|
||||
3. 增加专项离线测试,覆盖证据摘要与关键降级路径。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- `ToolInvocationRecorder` 契约增强
|
||||
- `LookupKnowledgeTool` 入库路径收敛
|
||||
- `QueryLogsTools` / `QueryMetricsTools` evidence 语义对齐
|
||||
- `ToolTraceSummaryService` 对失败 / no-hit / mixed evidence 的摘要规则收敛
|
||||
- `ChatService` 对 verifier 非法输出与降级输出的专项测试
|
||||
- 与该 change 直接相关的文档、OpenSpec、devflow 记录
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不引入新的数据库表或 schema 变更
|
||||
- 不扩展新的 evidence tool
|
||||
- 不做 P1-B 评测集 / harness
|
||||
- 不做前端 trace UI
|
||||
- 不处理敏感配置和默认 `mvn test` 离线化
|
||||
|
||||
---
|
||||
|
||||
## 预期结果
|
||||
|
||||
完成后,项目在面试里应能更清楚地表述为:
|
||||
|
||||
```text
|
||||
我不仅把 Agent 的工具调用落到了库里,
|
||||
还把 evidence trace、失败语义和 verifier 降级路径收敛成了稳定契约,
|
||||
并用离线测试覆盖了这些关键失败场景。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
|
||||
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`
|
||||
@@ -0,0 +1,99 @@
|
||||
# ISS-006 固定诊断评测集与回归 Harness
|
||||
|
||||
**状态**:进行中(sm-flow)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-B 面试打磨项
|
||||
**依赖**:ISS-005 / `evidence-trace-hardening`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
MVP 已经具备可追溯证据链、Verifier 质量门禁、trace API 和固定 demo 流程。上一阶段 `evidence-trace-hardening` 进一步统一了 evidence tool 的状态语义,让系统能稳定区分:
|
||||
|
||||
- `supported`
|
||||
- `no_evidence`
|
||||
- `deduped`
|
||||
- `failed`
|
||||
|
||||
下一步需要证明 Agent 在一组固定诊断场景下的表现,而不是只依赖单次 demo。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前项目能演示一次支付超时诊断,但还缺少稳定的评测基线:
|
||||
|
||||
- 每次改 prompt、工具、Verifier 或检索逻辑后,无法快速判断是否退化。
|
||||
- 只能人工看 trace,缺少结构化通过 / 失败结果。
|
||||
- 缺少面试时能展示的指标,如 evidence coverage、verdict 分布、工具调用数量和耗时。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
建立一个轻量的固定 case 评测 harness,用于验证 MVP Agent 的诊断质量和证据链完整性。
|
||||
|
||||
第一版不做 LLM-as-judge,优先做规则化校验:
|
||||
|
||||
- 固定 5 个 MVP 诊断 case
|
||||
- 每个 case 定义 expected root-cause keywords、required evidence tools、allowed verdicts
|
||||
- 基于 trace 结果校验 evidence coverage、verifier evaluation、tool invocation、final answer shape
|
||||
- 输出 JSON 和 Markdown 报告
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- 评测 case 定义文件
|
||||
- trace 规则校验器
|
||||
- eval runner 或测试入口
|
||||
- JSON / Markdown 报告输出
|
||||
- demo 文档和 devflow 记录
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不引入 LLM-as-judge
|
||||
- 不要求完整离线 LLM runtime
|
||||
- 不新增生产 API
|
||||
- 不修改 Chat 主链路
|
||||
- 不修改 evidence trace 运行时语义
|
||||
|
||||
---
|
||||
|
||||
## 预期面试表达
|
||||
|
||||
完成后可以这样描述:
|
||||
|
||||
```text
|
||||
我不仅有一个可演示的 Agent,还给它建立了固定 case 的回归评测。
|
||||
每次修改 prompt、工具或 verifier 后,都可以跑同一批诊断 case,
|
||||
检查证据覆盖、verdict 分布、工具调用成本和关键结论是否退化。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 初始候选 case
|
||||
|
||||
| Case | 目标 |
|
||||
| --- | --- |
|
||||
| payment-timeout | 支付接口超时,验证知识库 + 日志 + 指标证据 |
|
||||
| mysql-pool-exhausted | 数据库连接池耗尽,验证日志和知识库证据 |
|
||||
| redis-timeout | Redis 连接超时,验证日志依赖证据 |
|
||||
| slow-response | P99 响应时间过高,验证指标 + 慢请求日志 |
|
||||
| jvm-memory-risk | JVM 内存 / OOM 风险,验证指标 + 系统事件日志 |
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `mvp/demo/README.md`
|
||||
- `mvp/demo/payment-timeout-acceptance.md`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- `src/main/java/com/superbiz/agent/domain/entity/DiagnosisSession.java`
|
||||
- `src/main/java/com/superbiz/agent/domain/entity/ToolInvocation.java`
|
||||
- `openspec/specs/evidence-trace-hardening/spec.md`
|
||||
@@ -6,6 +6,11 @@
|
||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
|
||||
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
|
||||
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 已归档 | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
|
||||
|
||||
## RAG 重构计划
|
||||
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
# Diagnosis Eval Baseline Diff
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**来源**:P1-B follow-up
|
||||
**依赖**:`diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
现在项目已经有固定诊断 case、完整 fixture 和 baseline report。下一步需要把 baseline 真正用起来:每次改 Agent 后,把新的 report 和 baseline report 做对比。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 baseline 只能告诉我们“标准状态是什么”,但还不能自动告诉我们“这次改动有没有变差”。
|
||||
|
||||
典型问题包括:
|
||||
|
||||
- pass rate 是否下降。
|
||||
- 某个 case 是否从通过变失败。
|
||||
- 某个 evidence tool 是否从覆盖变成缺失。
|
||||
- verifier verdict 分布是否异常变化。
|
||||
- 平均工具调用数和耗时是否明显上升。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
新增一个 deterministic baseline diff 能力,用代码比较两份 `DiagnosisEvalReport`。
|
||||
|
||||
完成后应该做到:
|
||||
|
||||
- 输入 baseline report 和 current report。
|
||||
- 输出结构化 diff。
|
||||
- 标出 regression、improvement 和普通 changed。
|
||||
- 支持 JSON 和 Markdown 输出。
|
||||
- 文档说明面试时怎么解释这套回归判断。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- report-level diff 数据结构。
|
||||
- aggregate 指标比较。
|
||||
- case-level 指标比较。
|
||||
- JSON / Markdown diff writer。
|
||||
- focused tests 和 eval 文档。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不运行真实 Agent。
|
||||
- 不生成新 trace。
|
||||
- 不引入 LLM-as-judge。
|
||||
- 不改现有 evaluator 评分规则。
|
||||
|
||||
---
|
||||
|
||||
## 面试表达
|
||||
|
||||
可以这样讲:
|
||||
|
||||
```text
|
||||
我不是只保存了一份 baseline,而是加了 baseline diff。
|
||||
每次改 prompt、tool、retrieval 或 verifier 后,
|
||||
我都能把新 report 和 baseline 比较,
|
||||
直接看到哪些 case 退化、哪些证据缺失、成本有没有上升。
|
||||
```
|
||||
@@ -0,0 +1,83 @@
|
||||
# Expand Diagnosis Eval Fixtures
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-B follow-up
|
||||
**依赖**:`diagnosis-eval-harness`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
`diagnosis-eval-harness` 已经把固定 case、trace evaluator、JSON / Markdown report 和字段文档搭起来了。
|
||||
|
||||
现在还差一步:5 条固定诊断 case 里,只有 2 条有 fixture,另外 3 条还是 missing 状态。这个状态可以验证 evaluator 的错误报告能力,但还不能作为完整 baseline 展示。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 baseline 还不够完整:
|
||||
|
||||
- `redis-timeout` 没有对应 trace fixture。
|
||||
- `slow-response` 没有对应 trace fixture。
|
||||
- `jvm-memory-risk` 没有对应 trace fixture。
|
||||
- 仓库里还没有一份固定的 baseline JSON / Markdown 报告可供对比。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
补齐固定诊断评测集,让它从“框架可跑”变成“基准可用”。
|
||||
|
||||
完成后应该做到:
|
||||
|
||||
- 5 条固定 case 都能加载到对应 fixture。
|
||||
- evaluator 能输出完整 baseline report。
|
||||
- baseline report 被保存到仓库,后续 Agent 改动可以拿它做对比。
|
||||
- 文档说明怎么重新生成和怎么看报告。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- 补齐 3 个缺失 fixture。
|
||||
- 保存 baseline JSON / Markdown 报告。
|
||||
- 更新 eval 文档。
|
||||
- 补充测试,确保 case 文件引用的 fixture 都存在。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不新增 case 数量。
|
||||
- 不改生产 Agent 主链路。
|
||||
- 不引入 LLM-as-judge。
|
||||
- 不启动真实 MySQL、Redis、Milvus 或 LLM。
|
||||
|
||||
---
|
||||
|
||||
## 面试表达
|
||||
|
||||
可以这样讲:
|
||||
|
||||
```text
|
||||
我先搭了评测 harness,然后把固定 case 的 trace fixture 补齐,
|
||||
生成一份可复现的 baseline report。
|
||||
这样以后每次改 prompt、tool 或 verifier,
|
||||
都能看固定诊断集有没有行为回退,而不是只靠人工感觉。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- `mvp/eval/fixtures/`
|
||||
- `mvp/eval/reports/`
|
||||
- `mvp/eval/README.md`
|
||||
- `mvp/eval/schema.md`
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||
- `src/test/java/com/superbiz/agent/eval/DiagnosisTraceEvaluatorTest.java`
|
||||
- `openspec/specs/diagnosis-eval-harness/spec.md`
|
||||
@@ -0,0 +1,53 @@
|
||||
# MVP Demo Interview Runbook
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**来源**:Plan C
|
||||
**依赖**:`mvp-demo-trace-acceptance`, `evidence-trace-hardening`, `diagnosis-eval-harness`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
项目已经有 Agent 主链路、证据 trace、Verifier、反馈、eval baseline,但这些材料分散在不同目录。面试时真正需要的是一个能快速跑、快速讲清楚的 demo 入口。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 demo 还不够“面试友好”:
|
||||
|
||||
- 启动、请求、trace、反馈步骤分散在文档里。
|
||||
- 没有固定请求 payload 文件。
|
||||
- 没有一键跑 payment-timeout demo 的脚本。
|
||||
- 没有把 trace 字段和面试讲法对应起来的 walkthrough。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
把 Plan C 落地成 `mvp/demo` 下的可复现 demo 包:
|
||||
|
||||
- 固定支付超时请求。
|
||||
- 一键执行 chat、trace、feedback。
|
||||
- 保存 demo 输出,便于复盘。
|
||||
- 提供面试讲解稿和 trace 检查清单。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- `mvp/demo` 文档。
|
||||
- `mvp/demo/requests` 请求文件。
|
||||
- `mvp/demo/scripts` PowerShell 脚本。
|
||||
- `mvp/demo/output` 目录说明。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不新增后端 API。
|
||||
- 不改 Agent prompt。
|
||||
- 不扩 eval harness。
|
||||
- 不处理密钥外置和完整离线化。
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The AIOps endpoint has two natural modes:
|
||||
|
||||
- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert.
|
||||
- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them.
|
||||
|
||||
The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make AIOps payload mode single-alert focused.
|
||||
- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior.
|
||||
- Keep the change prompt-only and low risk.
|
||||
- Add tests for prompt scope rules.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not add a Verifier Agent.
|
||||
- Do not force tool calls in Java code.
|
||||
- Do not change `/api/ai_ops` request/response contracts.
|
||||
- Do not modify mock alert data.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. |
|
||||
| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. |
|
||||
| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. |
|
||||
| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. |
|
||||
|
||||
## Prompt Rules
|
||||
|
||||
Payload mode MUST instruct the agent:
|
||||
|
||||
- Treat supplied payload as the primary and only report target.
|
||||
- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk.
|
||||
- Do not create root-cause sections for unrelated active alerts.
|
||||
- Report unrelated alerts only in a brief "关联风险" note if they appear relevant.
|
||||
|
||||
Auto-discovery mode MUST instruct the agent:
|
||||
|
||||
- First call `queryPrometheusAlerts`.
|
||||
- Select P0/P1 or longest-running firing alerts.
|
||||
- Analyze one or more active alerts based on severity and evidence.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace.
|
||||
- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections.
|
||||
- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
Runtime verification showed that AIOps now correctly accepts an alert payload and persists a trace, but the generated report still expands to every active mock Prometheus alert. That weakens the product boundary between `/api/chat` and `/api/ai_ops`: an alert payload should mean targeted alert diagnosis, while an empty payload should mean automatic active-alert discovery.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Tighten the AIOps task prompt so payload mode focuses the final report on the supplied alert.
|
||||
- Preserve full active-alert discovery when no payload is supplied.
|
||||
- Allow Prometheus active-alert lookup in payload mode only as supporting evidence, not as permission to expand the report to unrelated alerts.
|
||||
- Update tests and demo acceptance wording to lock the new behavior.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `aiops-alert-scope-control`: Defines AIOps diagnosis scope rules for payload mode versus auto-discovery mode.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `aiops-traceable-diagnosis-entry`: Keeps the same API and trace behavior but clarifies how AIOps should scope its diagnosis.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `AiOpsService.buildTaskPrompt(...)`, focused tests, demo documentation, devflow records.
|
||||
- Affected API: no endpoint or request/response shape change.
|
||||
- Affected persistence: no schema change.
|
||||
- Non-goals: no Verifier integration, no tool implementation change, no prompt rewrite for Chat.
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: AIOps payload mode focuses on supplied alert
|
||||
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
|
||||
|
||||
#### Scenario: Request includes alertName and service
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
|
||||
- **THEN** the AIOps task prompt identifies payload mode
|
||||
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
|
||||
|
||||
### Requirement: AIOps auto-discovery mode queries active alerts first
|
||||
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
|
||||
|
||||
#### Scenario: Request body is omitted
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
|
||||
- **THEN** the AIOps task prompt identifies auto-discovery mode
|
||||
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
|
||||
|
||||
### Requirement: Payload mode may use active alerts as supporting context
|
||||
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
|
||||
|
||||
#### Scenario: Prometheus returns multiple active alerts
|
||||
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
|
||||
- **THEN** the prompt permits mentioning those alerts only as related risk or context
|
||||
- **AND** the final report target remains the supplied alert
|
||||
@@ -0,0 +1,21 @@
|
||||
## 1. Flow Records
|
||||
|
||||
- [x] 1.1 Add devflow brief, decisions with lightweight Grill, evidence, and acceptance records.
|
||||
- [x] 1.2 Record local impact analysis and GitNexus skip context.
|
||||
|
||||
## 2. Prompt Scope Control
|
||||
|
||||
- [x] 2.1 Add payload detection helper in `AiOpsService`.
|
||||
- [x] 2.2 Update `buildTaskPrompt(...)` with payload-mode and auto-discovery-mode rules.
|
||||
|
||||
## 3. Tests And Docs
|
||||
|
||||
- [x] 3.1 Add tests for payload-mode prompt rules.
|
||||
- [x] 3.2 Add tests for no-payload auto-discovery prompt rules.
|
||||
- [x] 3.3 Update AIOps demo acceptance wording for single-alert payload mode.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Run targeted tests.
|
||||
- [x] 4.2 Run compile verification.
|
||||
- [x] 4.3 Run OpenSpec validation.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,90 @@
|
||||
## Context
|
||||
|
||||
The AIOps endpoint is currently useful as a standalone alert-analysis demo, but it is not aligned with the MVP trace story:
|
||||
|
||||
- `ChatController.aiOps()` accepts no request body.
|
||||
- `AiOpsService.executeAiOpsAnalysis(...)` creates a random 8-character session id internally.
|
||||
- The caller cannot reliably discover that id and query `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `extractFinalReport(...)` returns the report text but does not persist it to `diagnosis_session.answer`.
|
||||
|
||||
The existing trace API already aggregates `diagnosis_session`, `agent_step`, and `tool_invocation`, so this change should reuse that storage rather than introduce new persistence.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make `/api/ai_ops` usable as an alert-triggered diagnosis entry point.
|
||||
- Preserve backward compatibility for callers that post with no request body.
|
||||
- Return the resolved `sessionId` through SSE.
|
||||
- Persist the final report into the existing diagnosis session.
|
||||
- Keep the AIOps path observable through the existing trace API.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not merge AIOps into `ChatService`.
|
||||
- Do not add a new Verifier Agent to AIOps in this slice.
|
||||
- Do not change `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- Do not add database migrations.
|
||||
- Do not clean up sensitive configuration.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| API compatibility | Keep `POST /api/ai_ops` SSE and make body optional | Add a new `/api/ai_ops/v2` endpoint | Optional body keeps existing demo callers working while enabling traceable alert input. |
|
||||
| Session identity | Accept request `sessionId`, otherwise generate UUID | Continue internal-only 8-char id | Reviewers need the id to query trace and submit feedback. |
|
||||
| Query persistence | Build a concise alert diagnosis query from request fields | Store only "AI Ops 告警分析" | Trace should show what alert was diagnosed. |
|
||||
| Final answer persistence | Save extracted final report to `diagnosis_session.answer` | Only stream the report | Trace replay must include the final answer without relying on SSE logs. |
|
||||
| Verifier scope | Defer AIOps Verifier integration | Add Chat Verifier now | The minimum interview value is traceability; Verifier unification can be a follow-up after this entry point is stable. |
|
||||
| GitNexus | Skip by user decision | Block until MCP available | GitNexus tools are not exposed in this session, and the user explicitly requested skipping GitNexus. Local impact analysis and tests cover this slice. |
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Level: L3 API behavior extension.
|
||||
- Endpoint: `POST /api/ai_ops`
|
||||
- Compatibility: callers may still omit a body. New callers may send:
|
||||
|
||||
```json
|
||||
{
|
||||
"sessionId": "mvp-demo-aiops-payment-latency-001",
|
||||
"alertName": "payment-service-latency-high",
|
||||
"service": "payment-service",
|
||||
"severity": "P1",
|
||||
"description": "支付服务 P95 延迟升高并伴随超时错误",
|
||||
"timeRange": "last_15m"
|
||||
}
|
||||
```
|
||||
|
||||
The SSE stream emits a first content message containing the resolved session id:
|
||||
|
||||
```text
|
||||
sessionId: mvp-demo-aiops-payment-latency-001
|
||||
```
|
||||
|
||||
## Data Flow
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> ChatController resolves request body and tools
|
||||
-> AiOpsService.executeAiOpsAnalysis(chatModel, tools, request)
|
||||
-> create diagnosis_session(agentFlow=AI_OPS, query=<alert summary>)
|
||||
-> set SessionContextHolder(sessionId)
|
||||
-> ai_ops_supervisor -> planner_agent -> executor_agent
|
||||
-> persist agent_step and tool_invocation through existing hooks/tools
|
||||
-> extract final report
|
||||
-> persist diagnosis_session.answer/status/counts
|
||||
-> caller queries GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] AIOps still lacks the Chat Verifier quality gate. -> Mitigation: document as follow-up and keep this slice focused on traceability.
|
||||
- [Risk] SSE clients may not parse the new first message. -> Mitigation: message is additive content; existing clients still receive the final report.
|
||||
- [Risk] Optional request body in Spring MVC can be easy to mishandle. -> Mitigation: use `@RequestBody(required = false)` and default request values in service code.
|
||||
- [Risk] AIOps generated reports may still depend on real infrastructure. -> Mitigation: demo profile already enables mock logs/metrics where available; full offline mode remains out of scope.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No database migration.
|
||||
- Deploy with application restart.
|
||||
- Rollback by reverting controller/service/DTO changes; existing persisted sessions remain valid.
|
||||
@@ -0,0 +1,29 @@
|
||||
## Why
|
||||
|
||||
The MVP already has a strong traceable chat diagnosis path, but the legacy `/api/ai_ops` endpoint still behaves like an early standalone demo: it accepts no alert payload, generates an internal session id that callers cannot reuse, and streams a report without reliably persisting the final answer for trace replay. For an Agent Engineer interview project, AIOps should become a second entry point into the same observable diagnosis story rather than a disconnected legacy path.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Allow `/api/ai_ops` to accept an optional alert diagnosis request body.
|
||||
- Resolve a stable session id from the request or generate one when omitted.
|
||||
- Persist the AIOps alert query and final report into `diagnosis_session`.
|
||||
- Emit the resolved session id in the SSE stream so reviewers can call `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- Keep the existing AIOps planner/executor flow and evidence tools; do not replace it with the chat flow in this slice.
|
||||
- Document the AIOps demo path beside the existing MVP demo trace flow.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `aiops-traceable-diagnosis-entry`: Makes the AIOps alert endpoint traceable by session id and replayable through the existing diagnosis trace API.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- Existing `/api/ai_ops` behavior is extended from a no-input SSE trigger into an optional request-body alert diagnosis endpoint.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `ChatController`, `AiOpsService`, `AIOpsRequest`, focused tests, MVP demo documentation, devflow records.
|
||||
- Affected API: `POST /api/ai_ops` remains SSE, but now accepts an optional JSON body and streams a first message containing `sessionId`.
|
||||
- Affected persistence: no schema migration; writes existing `diagnosis_session.query`, `answer`, `status`, timing, and aggregate counts.
|
||||
- Non-goals: no full AIOps/Chat service unification, no new database table, no production security cleanup, no full offline fake runtime, no mandatory Verifier integration for AIOps in this slice.
|
||||
+41
@@ -0,0 +1,41 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: AIOps endpoint accepts optional alert input
|
||||
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
|
||||
|
||||
#### Scenario: Caller supplies alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
|
||||
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
|
||||
|
||||
#### Scenario: Caller omits alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without a body
|
||||
- **THEN** the system still starts the default AIOps alert-analysis flow
|
||||
|
||||
### Requirement: AIOps session id is traceable
|
||||
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
|
||||
|
||||
#### Scenario: Request includes session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
|
||||
- **THEN** the created `diagnosis_session.session_id` equals that value
|
||||
- **AND** the SSE stream includes the same session id
|
||||
|
||||
#### Scenario: Request omits session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
|
||||
- **THEN** the system generates a session id
|
||||
- **AND** the SSE stream includes the generated session id
|
||||
|
||||
### Requirement: AIOps report is persisted for trace replay
|
||||
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
|
||||
|
||||
#### Scenario: AIOps report is generated
|
||||
- **WHEN** the AIOps planner/executor flow returns a final report
|
||||
- **THEN** the corresponding diagnosis session is marked successful
|
||||
- **AND** `diagnosis_session.answer` stores the final report
|
||||
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
|
||||
|
||||
### Requirement: AIOps trace uses existing evidence tables
|
||||
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
|
||||
|
||||
#### Scenario: AIOps uses evidence tools
|
||||
- **WHEN** the AIOps flow calls available evidence tools
|
||||
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Flow Records
|
||||
|
||||
- [x] 1.1 Add devflow brief, decisions, evidence, and acceptance records for `aiops-traceable-diagnosis-entry`.
|
||||
- [x] 1.2 Record user-approved GitNexus skip and local impact analysis.
|
||||
|
||||
## 2. AIOps API And Service
|
||||
|
||||
- [x] 2.1 Extend `AIOpsRequest` with optional session id and alert fields.
|
||||
- [x] 2.2 Change `/api/ai_ops` to accept an optional request body and emit the resolved session id in SSE.
|
||||
- [x] 2.3 Change `AiOpsService` to accept the request, resolve session id, build a request-specific prompt, and persist the request summary.
|
||||
- [x] 2.4 Persist the final AIOps report to `diagnosis_session.answer`.
|
||||
|
||||
## 3. Demo Documentation
|
||||
|
||||
- [x] 3.1 Add an AIOps alert demo section to `mvp/demo/README.md`.
|
||||
- [x] 3.2 Add a concrete AIOps acceptance case under `mvp/demo`.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Add focused unit coverage for AIOps request/session/report persistence behavior where practical.
|
||||
- [x] 4.2 Run targeted tests.
|
||||
- [x] 4.3 Run compile verification.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The project now has the pieces needed for trace-based evaluation:
|
||||
|
||||
- `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation`
|
||||
- `agent_step` stores ordered agent execution records
|
||||
- `tool_invocation` stores evidence tool calls with normalized evidence semantics
|
||||
- `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review
|
||||
- `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics
|
||||
|
||||
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
|
||||
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
|
||||
- Produce JSON and Markdown reports with pass/fail status and key metrics.
|
||||
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
|
||||
- Leave room for a later runtime mode that queries the trace API after a demo run.
|
||||
|
||||
**Non-Goals:**
|
||||
- No LLM-as-judge in this slice.
|
||||
- No automatic prompt optimization.
|
||||
- No new production API.
|
||||
- No change to chat, verifier, retrieval, upload, or feedback behavior.
|
||||
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
|
||||
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
|
||||
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
|
||||
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
|
||||
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
|
||||
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
|
||||
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
|
||||
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required.
|
||||
- The harness is additive and can be run locally as a test or script.
|
||||
- Rollback is deleting the eval case files, runner, and report docs.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
|
||||
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
|
||||
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
|
||||
- Add JSON and Markdown report output for quick review after a run.
|
||||
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
|
||||
|
||||
### Modified Capabilities
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
|
||||
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
|
||||
- Affected APIs: none.
|
||||
- Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions.
|
||||
- Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
|
||||
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
|
||||
|
||||
#### Scenario: Case definition includes expected evidence
|
||||
- **WHEN** an evaluation case is defined
|
||||
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
|
||||
|
||||
#### Scenario: Case definition can express forbidden behavior
|
||||
- **WHEN** a case has known unsafe behavior
|
||||
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate diagnosis traces
|
||||
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
|
||||
|
||||
#### Scenario: Evidence coverage validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
|
||||
|
||||
#### Scenario: Verifier evaluation validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
|
||||
- **AND** the verdict SHALL be one of the case's allowed verdicts
|
||||
|
||||
#### Scenario: Answer keyword validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
|
||||
|
||||
#### Scenario: Degraded output validation
|
||||
- **WHEN** a trace verdict is `REJECT`
|
||||
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
|
||||
|
||||
### Requirement: Evaluation harness SHALL report quality and cost signals
|
||||
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
|
||||
|
||||
#### Scenario: JSON report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
|
||||
|
||||
#### Scenario: Markdown report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
|
||||
|
||||
#### Scenario: Aggregate metrics
|
||||
- **WHEN** multiple cases are evaluated
|
||||
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
|
||||
|
||||
### Requirement: Evaluation harness SHALL support offline fixture mode
|
||||
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
|
||||
|
||||
#### Scenario: Fixture trace evaluation
|
||||
- **WHEN** the evaluator is run against a directory of trace fixture files
|
||||
- **THEN** it SHALL evaluate each trace file against its matching case definition
|
||||
- **AND** it SHALL not require a running application service
|
||||
|
||||
#### Scenario: Missing fixture is reported clearly
|
||||
- **WHEN** a case has no matching trace fixture
|
||||
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
|
||||
@@ -0,0 +1,25 @@
|
||||
## 1. Case Definitions
|
||||
|
||||
- [x] 1.1 Add evaluation case definition format for fixed MVP diagnosis scenarios.
|
||||
- [x] 1.2 Add the first 5 case definitions: payment timeout, MySQL pool exhausted, Redis timeout, slow response, and JVM memory risk.
|
||||
- [x] 1.3 Document the meaning of expected keywords, required evidence tools, allowed verdicts, and forbidden behavior.
|
||||
|
||||
## 2. Trace Fixtures
|
||||
|
||||
- [x] 2.1 Add fixture trace schema or DTOs that match `DiagnosisTraceResponse` enough for offline evaluation.
|
||||
- [x] 2.2 Add at least one representative trace fixture for a passing case.
|
||||
- [x] 2.3 Add at least one fixture covering low-confidence or degraded behavior.
|
||||
|
||||
## 3. Evaluator
|
||||
|
||||
- [x] 3.1 Implement trace validation rules for evidence coverage, verifier verdict, answer keyword coverage, and degraded-output contract.
|
||||
- [x] 3.2 Implement aggregate metrics: pass rate, verdict distribution, average tool-call count, and average duration.
|
||||
- [x] 3.3 Implement JSON report output.
|
||||
- [x] 3.4 Implement Markdown report output.
|
||||
|
||||
## 4. Tests And Documentation
|
||||
|
||||
- [x] 4.1 Add focused offline tests for the evaluator.
|
||||
- [x] 4.2 Add run instructions under `mvp/demo` or `mvp/notes`.
|
||||
- [x] 4.3 Run targeted tests for the evaluator.
|
||||
- [x] 4.4 Run compile verification.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,64 @@
|
||||
## Context
|
||||
|
||||
The current MVP already has the core pieces required for traceable agent execution:
|
||||
|
||||
- `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation`
|
||||
- `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder`
|
||||
- `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries
|
||||
- `ChatService` already contains fallback behavior for missing or invalid `verifier_output`
|
||||
|
||||
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
|
||||
|
||||
- `LookupKnowledgeTool` builds `ToolInvocation` rows itself
|
||||
- the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)`
|
||||
- “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
|
||||
- degraded output behavior exists in code but is only lightly covered by tests
|
||||
|
||||
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Centralize the common persistence contract for evidence-bearing tools.
|
||||
- Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
|
||||
- Define stable summarization semantics for:
|
||||
- successful evidence
|
||||
- no-hit / no-usable-evidence
|
||||
- deduped retrievals
|
||||
- failed evidence queries
|
||||
- Make `ChatService` fallback and degraded-output paths testable as explicit product behavior.
|
||||
- Keep the scope small enough to unblock the next P1-B evaluation harness.
|
||||
|
||||
**Non-Goals:**
|
||||
- No new table, column, or Flyway migration.
|
||||
- No new public API.
|
||||
- No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`.
|
||||
- No attempt to redesign planner/executor routing.
|
||||
- No full offline runtime or end-to-end benchmark harness in this slice.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. |
|
||||
| `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
|
||||
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
|
||||
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
|
||||
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
|
||||
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
|
||||
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
|
||||
- [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required beyond shipping the code changes.
|
||||
- Existing `tool_invocation` rows remain valid because this change reuses the same schema.
|
||||
- Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.
|
||||
@@ -0,0 +1,26 @@
|
||||
## Why
|
||||
|
||||
The MVP already persists evidence tool invocations and uses a Verifier to judge answer quality, but the current evidence trace semantics are still only partially standardized. For interview-grade agent engineering, the system needs a tighter contract for evidence persistence, no-evidence/failure states, and degraded output behavior, plus focused tests that prove those paths work offline.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
- Align how evidence tools represent success, no-hit, deduped, and failed calls so `ToolTraceSummaryService` can summarize them consistently.
|
||||
- Harden `ChatService` fallback behavior for invalid or missing verifier output and make the degraded-output paths explicitly testable.
|
||||
- Add focused offline tests for evidence recording, trace summarization, and verifier fallback / degraded output behavior.
|
||||
- Record this slice as a dedicated P1-A change tied to the interview-focused MVP hardening track.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `evidence-trace-hardening`: Covers standardized evidence invocation persistence, verifier-facing evidence summary semantics, and explicit degraded-output contracts for evidence gaps and verifier failures.
|
||||
|
||||
### Modified Capabilities
|
||||
- `chat-verifier-agent`: Tightens verifier input evidence semantics and fallback guarantees without changing the high-level planner/executor/verifier workflow.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `ToolInvocationRecorder`, `LookupKnowledgeTool`, `QueryLogsTools`, `QueryMetricsTools`, `ToolTraceSummaryService`, `ChatService`, and focused test classes.
|
||||
- Affected runtime behavior: evidence-bearing tools will persist more consistent invocation semantics; verifier fallback and degraded outputs remain additive hardening, not a product-flow rewrite.
|
||||
- Affected APIs: none. No new endpoint or schema is introduced.
|
||||
- Non-goals: no new evidence tools, no database migration, no evaluation harness, no trace UI, no security/config cleanup in this slice.
|
||||
+14
@@ -0,0 +1,14 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
|
||||
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
|
||||
|
||||
#### Scenario: Failed evidence remains a verifier-visible gap
|
||||
- **WHEN** an evidence-bearing tool invocation fails
|
||||
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
|
||||
- **AND** the verifier flow SHALL continue without crashing
|
||||
|
||||
#### Scenario: Deduped retrievals do not count as fresh support
|
||||
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
|
||||
- **THEN** those entries SHALL be treated as no-new-evidence
|
||||
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
|
||||
+82
@@ -0,0 +1,82 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
|
||||
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
|
||||
#### Scenario: Common evidence fields are always persisted
|
||||
- **WHEN** an evidence-bearing tool finishes a call
|
||||
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
|
||||
|
||||
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
|
||||
- **WHEN** `lookup_knowledge` persists a tool invocation
|
||||
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
|
||||
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
|
||||
|
||||
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
|
||||
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
|
||||
|
||||
#### Scenario: Tool failure is preserved as failure
|
||||
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
|
||||
- **THEN** the persisted row SHALL set `success=false`
|
||||
- **AND** it SHALL preserve an `error_message` explaining the failure
|
||||
|
||||
#### Scenario: No usable evidence is preserved without pretending success
|
||||
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
|
||||
- **THEN** the persisted contract SHALL preserve that the call completed
|
||||
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
|
||||
|
||||
#### Scenario: Deduped retrieval remains auditable
|
||||
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
|
||||
- **THEN** the persisted row SHALL preserve the dedup reason
|
||||
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
|
||||
|
||||
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
|
||||
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
|
||||
|
||||
#### Scenario: Failed evidence calls remain visible in the summary
|
||||
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
|
||||
- **THEN** the summary SHALL retain them
|
||||
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
|
||||
|
||||
#### Scenario: No-hit and deduped calls do not upgrade evidence level
|
||||
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
|
||||
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
|
||||
- **AND** their counts SHALL still be reflected in the merged summary entry
|
||||
|
||||
#### Scenario: Successful evidence keeps the strongest available support
|
||||
- **WHEN** multiple rows for the same tool and topic domain are merged
|
||||
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
|
||||
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
|
||||
|
||||
### Requirement: ChatService SHALL degrade predictably on verifier output failures
|
||||
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
|
||||
|
||||
#### Scenario: Missing verifier output falls back to LOW_CONFID
|
||||
- **WHEN** the verifier step completes without a usable `verifier_output`
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
|
||||
|
||||
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
|
||||
- **WHEN** the verifier returns malformed or non-parseable JSON
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the fallback SHALL still persist a verifier-evaluation record
|
||||
|
||||
#### Scenario: REJECT output hides unverified raw answer text
|
||||
- **WHEN** the final verifier decision is `REJECT`
|
||||
- **THEN** the user-facing output SHALL use the degraded template
|
||||
- **AND** it SHALL NOT pass through the raw executor answer
|
||||
|
||||
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
|
||||
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
|
||||
|
||||
#### Scenario: Evidence recorder contract is tested offline
|
||||
- **WHEN** the test suite runs the focused recorder tests
|
||||
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
|
||||
|
||||
#### Scenario: Trace summary hardening is tested offline
|
||||
- **WHEN** the test suite runs the focused trace-summary tests
|
||||
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
|
||||
|
||||
#### Scenario: Verifier fallback behavior is tested offline
|
||||
- **WHEN** the test suite runs the focused `ChatService` fallback tests
|
||||
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Evidence Persistence Contract
|
||||
|
||||
- [x] 1.1 Extend `ToolInvocationRecorder` with a richer evidence-recording path that can preserve retrieval-aware fields as well as common evidence fields.
|
||||
- [x] 1.2 Refactor `LookupKnowledgeTool` to persist `tool_invocation` rows through `ToolInvocationRecorder` instead of its own ad hoc row-construction path.
|
||||
- [x] 1.3 Align `QueryLogsTools` and `QueryMetricsTools` no-hit / failure payloads with the hardened evidence contract.
|
||||
|
||||
## 2. Verifier-Facing Summary Semantics
|
||||
|
||||
- [x] 2.1 Harden `ToolTraceSummaryService` so failed, no-hit, and deduped evidence rows are summarized with stable no-evidence semantics.
|
||||
- [x] 2.2 Preserve merged-call counts for repeated hits, failures, and no-new-evidence rows without overstating evidence strength.
|
||||
|
||||
## 3. Chat Degraded Paths
|
||||
|
||||
- [x] 3.1 Add focused `ChatService` tests for missing verifier output fallback to `LOW_CONFID`.
|
||||
- [x] 3.2 Add focused `ChatService` tests for invalid verifier JSON fallback to `LOW_CONFID`.
|
||||
- [x] 3.3 Add focused `ChatService` tests that `REJECT` output uses the degraded template and does not leak raw executor answer content.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Add focused offline tests for the recorder contract and `ToolTraceSummaryService`.
|
||||
- [x] 4.2 Run targeted test commands for the new/updated offline tests.
|
||||
- [x] 4.3 Run compile verification.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,39 @@
|
||||
## Context
|
||||
|
||||
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Compare two `DiagnosisEvalReport` objects without requiring external services.
|
||||
- Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
|
||||
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
|
||||
- Write JSON and Markdown diff outputs for review.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not run the Agent or regenerate traces.
|
||||
- Do not introduce LLM-as-judge.
|
||||
- Do not change evaluator scoring rules.
|
||||
- Do not block on performance thresholds beyond simple numeric diff signals.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Decision: Compare report DTOs instead of raw traces.
|
||||
- Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
|
||||
- Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
|
||||
|
||||
- Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`.
|
||||
- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
|
||||
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
|
||||
|
||||
- Decision: Keep thresholds explicit and conservative.
|
||||
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
|
||||
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed.
|
||||
- Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
|
||||
- Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.
|
||||
@@ -0,0 +1,28 @@
|
||||
## Why
|
||||
|
||||
The evaluation baseline is now complete, but developers still need a repeatable way to decide whether a new Agent run regressed against that baseline. A deterministic baseline diff turns saved reports into an actionable regression signal instead of a static artifact.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a baseline diff model that compares two `DiagnosisEvalReport` objects.
|
||||
- Detect aggregate changes such as pass-rate drops, verdict distribution shifts, tool-call cost changes, and duration changes.
|
||||
- Detect per-case changes such as pass/fail regression, verdict changes, keyword coverage changes, evidence coverage loss, and missing/new cases.
|
||||
- Add JSON and Markdown diff output suitable for review.
|
||||
- Document how to interpret the diff in the eval docs.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness so a current report can be compared against the saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects eval-only Java code under `src/main/java/com/superbiz/agent/eval`.
|
||||
- Adds focused tests under `src/test/java/com/superbiz/agent/eval`.
|
||||
- Updates `mvp/eval` documentation and may add sample diff output.
|
||||
- No production Agent runtime, API, database schema, or external dependency changes are expected.
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
||||
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
||||
|
||||
#### Scenario: Aggregate regression detection
|
||||
- **WHEN** the current report has a lower pass rate than the baseline report
|
||||
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
||||
|
||||
#### Scenario: Cost signal detection
|
||||
- **WHEN** average tool-call count or average duration changes between reports
|
||||
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
||||
|
||||
#### Scenario: Verdict distribution comparison
|
||||
- **WHEN** verdict counts differ between reports
|
||||
- **THEN** the diff SHALL record the verdict distribution changes
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare case-level report results
|
||||
The system SHALL compare case results by case id and report actionable per-case changes.
|
||||
|
||||
#### Scenario: Case pass/fail regression
|
||||
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
||||
- **THEN** the diff SHALL record a regression for that case
|
||||
|
||||
#### Scenario: Evidence coverage regression
|
||||
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
||||
- **THEN** the diff SHALL record a regression naming the case and tool
|
||||
|
||||
#### Scenario: Missing case detection
|
||||
- **WHEN** a baseline case is absent from the current report
|
||||
- **THEN** the diff SHALL record a regression for the missing case
|
||||
|
||||
#### Scenario: New case detection
|
||||
- **WHEN** a current report contains a case absent from the baseline
|
||||
- **THEN** the diff SHALL record the case as a non-regression change
|
||||
|
||||
### Requirement: Evaluation harness SHALL report baseline diff results
|
||||
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
||||
|
||||
#### Scenario: JSON diff output
|
||||
- **WHEN** a baseline diff is written as JSON
|
||||
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
||||
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
@@ -0,0 +1,23 @@
|
||||
## 1. OpenSpec And Issue Setup
|
||||
|
||||
- [x] 1.1 Create slug-based issue and devflow tracking files.
|
||||
- [x] 1.2 Create OpenSpec proposal, design, delta spec, and tasks.
|
||||
|
||||
## 2. Baseline Diff Implementation
|
||||
|
||||
- [x] 2.1 Add diff result data structures for summary and per-item changes.
|
||||
- [x] 2.2 Implement deterministic report comparison rules.
|
||||
- [x] 2.3 Implement JSON and Markdown diff report writing.
|
||||
|
||||
## 3. Documentation
|
||||
|
||||
- [x] 3.1 Document baseline diff inputs, outputs, and interpretation in eval docs.
|
||||
- [x] 3.2 Add sample diff output for a representative regression.
|
||||
|
||||
## 4. Tests And Validation
|
||||
|
||||
- [x] 4.1 Add focused tests for aggregate and case-level diff behavior.
|
||||
- [x] 4.2 Add focused tests for JSON and Markdown diff output.
|
||||
- [x] 4.3 Run evaluator/diff tests.
|
||||
- [x] 4.4 Run compile verification.
|
||||
- [x] 4.5 Run OpenSpec validation.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,39 @@
|
||||
## Context
|
||||
|
||||
`diagnosis-eval-harness` already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Add representative trace fixtures for every fixed diagnosis case.
|
||||
- Save a baseline report that can be reviewed and compared after future Agent changes.
|
||||
- Keep the baseline reproducible in offline mode.
|
||||
- Document how to regenerate the baseline.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not change production Agent runtime behavior.
|
||||
- Do not require live infrastructure or a real LLM.
|
||||
- Do not introduce a new LLM-based grader.
|
||||
- Do not expand the case set beyond the existing five fixed MVP diagnosis cases.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Use checked-in fixture traces instead of live service calls.
|
||||
- Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
|
||||
- Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
|
||||
|
||||
- Save baseline reports under `mvp/eval/reports`.
|
||||
- Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
|
||||
- Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
|
||||
|
||||
- Keep fixture outcomes representative rather than forcing every case to pass.
|
||||
- Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
|
||||
- Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like `DiagnosisTraceResponse` and add tests that load every referenced fixture.
|
||||
- A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
|
||||
- Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The diagnosis evaluation harness is implemented, but the baseline is still incomplete because only two of the five fixed cases have trace fixtures. Completing the fixture set and saving a baseline report makes the harness useful as a practical regression signal for interview demos and future Agent changes.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add trace fixtures for the remaining fixed diagnosis cases: Redis timeout, slow response, and JVM memory risk.
|
||||
- Add a reproducible baseline report generated from the full fixture set.
|
||||
- Document how to regenerate and interpret the baseline.
|
||||
- Keep the evaluator deterministic and offline; no live MySQL, Redis, Milvus, or LLM service is required.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness requirement so the fixed MVP case set has complete fixture coverage and a saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects `mvp/eval/cases`, `mvp/eval/fixtures`, and eval documentation.
|
||||
- May add baseline output files under `mvp/eval/reports`.
|
||||
- May add or update focused evaluator tests to assert full fixture coverage and report generation.
|
||||
- No production runtime API or database schema changes are expected.
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
|
||||
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
|
||||
|
||||
#### Scenario: Every case resolves to a fixture file
|
||||
- **WHEN** the evaluator loads the fixed case definition file
|
||||
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
|
||||
|
||||
#### Scenario: Fixture files are loadable as diagnosis traces
|
||||
- **WHEN** each referenced fixture is loaded
|
||||
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
|
||||
|
||||
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
|
||||
The system SHALL preserve a generated baseline report for the full fixed fixture set.
|
||||
|
||||
#### Scenario: Baseline report includes all fixed cases
|
||||
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
|
||||
- **THEN** the report SHALL include one result for every fixed case
|
||||
|
||||
#### Scenario: Baseline report is reviewable
|
||||
- **WHEN** the baseline report is written
|
||||
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
|
||||
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes fixtures or evaluator rules
|
||||
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
||||
@@ -0,0 +1,19 @@
|
||||
## 1. Fixture Coverage
|
||||
|
||||
- [x] 1.1 Add Redis timeout trace fixture referenced by the fixed case file.
|
||||
- [x] 1.2 Add slow response trace fixture referenced by the fixed case file.
|
||||
- [x] 1.3 Add JVM memory risk trace fixture referenced by the fixed case file.
|
||||
- [x] 1.4 Verify every `traceFixture` in `diagnosis-cases.json` resolves to an existing fixture file.
|
||||
|
||||
## 2. Baseline Reports
|
||||
|
||||
- [x] 2.1 Generate a full baseline JSON report for all fixed cases.
|
||||
- [x] 2.2 Generate a full baseline Markdown report for review.
|
||||
- [x] 2.3 Document how to regenerate and interpret the baseline reports.
|
||||
|
||||
## 3. Tests And Validation
|
||||
|
||||
- [x] 3.1 Add or update focused tests for full fixture coverage and baseline report generation.
|
||||
- [x] 3.2 Run evaluator tests.
|
||||
- [x] 3.3 Run compile verification.
|
||||
- [x] 3.4 Run OpenSpec validation.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,35 @@
|
||||
## Context
|
||||
|
||||
The current `mvp/demo` folder documents the core flow, but the steps are embedded in prose. For an interview, the demo needs a sharper entry point: what to start, what to run, what files get produced, and what to point at when explaining Agent engineering quality.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make the payment-timeout demo runnable through a small script.
|
||||
- Save chat, trace, and feedback responses for review.
|
||||
- Provide a short interview walkthrough that connects runtime evidence to the engineering story.
|
||||
- Keep the demo focused on existing APIs and existing `mvp-demo` profile behavior.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not add new backend endpoints.
|
||||
- Do not modify Agent prompts or runtime orchestration.
|
||||
- Do not solve secret cleanup or full offline test isolation in this change.
|
||||
- Do not expand the eval harness.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Decision: Use PowerShell scripts.
|
||||
- Reason: the current runbook already uses PowerShell and the user environment is Windows.
|
||||
|
||||
- Decision: Save outputs under `mvp/demo/output`.
|
||||
- Reason: interview review is easier when chat, trace, and feedback responses are persisted as files.
|
||||
|
||||
- Decision: Keep the walkthrough separate from the low-level runbook.
|
||||
- Reason: `README.md` should tell how to run; `interview-walkthrough.md` should tell how to explain.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- The demo still depends on configured MySQL, Redis, Milvus, and model keys. Mitigation: document this explicitly and keep mock log/metric providers enabled through `mvp-demo`.
|
||||
- Script assertions are intentionally lightweight. Mitigation: use the trace checklist for human review and keep automated regression in `mvp/eval`.
|
||||
@@ -0,0 +1,26 @@
|
||||
## Why
|
||||
|
||||
The MVP already has trace, evidence hardening, and evaluation artifacts, but the interview demo path is still too scattered. This change packages the existing capabilities into a repeatable demo runbook that can be executed and explained in a short interview window.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a focused interview walkthrough for the payment-timeout MVP demo.
|
||||
- Add reusable request payloads and PowerShell scripts under `mvp/demo`.
|
||||
- Add a trace inspection checklist that maps runtime output to the engineering story.
|
||||
- Keep the change documentation-only and script-only; no backend runtime behavior changes.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `mvp-demo-trace-acceptance`: Extend the demo acceptance surface with a repeatable interview runbook and executable local demo scripts.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects `mvp/demo` documentation and scripts.
|
||||
- Adds issue and devflow tracking files.
|
||||
- No Java production code, API contract, database schema, or dependency changes are expected.
|
||||
+30
@@ -0,0 +1,30 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: MVP demo SHALL provide an interview runbook
|
||||
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
|
||||
|
||||
#### Scenario: Walkthrough explains the demo story
|
||||
- **WHEN** a developer opens the interview walkthrough
|
||||
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
|
||||
|
||||
#### Scenario: Walkthrough stays scoped to existing capabilities
|
||||
- **WHEN** the walkthrough describes the demo
|
||||
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
|
||||
|
||||
### Requirement: MVP demo SHALL provide executable local demo scripts
|
||||
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
|
||||
|
||||
#### Scenario: Demo script sends the fixed diagnosis request
|
||||
- **WHEN** the demo script is executed against a running local service
|
||||
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
|
||||
|
||||
#### Scenario: Demo script captures review artifacts
|
||||
- **WHEN** the demo script finishes successfully
|
||||
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
|
||||
|
||||
### Requirement: MVP demo SHALL provide a trace inspection checklist
|
||||
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
|
||||
|
||||
#### Scenario: Checklist maps fields to interview claims
|
||||
- **WHEN** a developer reviews a trace response
|
||||
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
|
||||
@@ -0,0 +1,16 @@
|
||||
## 1. Demo Artifacts
|
||||
|
||||
- [x] 1.1 Add fixed payment-timeout request payload.
|
||||
- [x] 1.2 Add PowerShell script to run chat, trace, and feedback steps.
|
||||
- [x] 1.3 Add output directory documentation without committing generated outputs.
|
||||
|
||||
## 2. Interview Documentation
|
||||
|
||||
- [x] 2.1 Add interview walkthrough for the demo story.
|
||||
- [x] 2.2 Add trace inspection checklist.
|
||||
- [x] 2.3 Update `mvp/demo/README.md` to link the runnable demo package.
|
||||
|
||||
## 3. Tracking And Validation
|
||||
|
||||
- [x] 3.1 Add slug-based issue and devflow tracking files.
|
||||
- [x] 3.2 Run OpenSpec validation.
|
||||
@@ -0,0 +1,28 @@
|
||||
# aiops-alert-scope-control Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change aiops-alert-scope-control. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: AIOps payload mode focuses on supplied alert
|
||||
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
|
||||
|
||||
#### Scenario: Request includes alertName and service
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
|
||||
- **THEN** the AIOps task prompt identifies payload mode
|
||||
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
|
||||
|
||||
### Requirement: AIOps auto-discovery mode queries active alerts first
|
||||
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
|
||||
|
||||
#### Scenario: Request body is omitted
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
|
||||
- **THEN** the AIOps task prompt identifies auto-discovery mode
|
||||
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
|
||||
|
||||
### Requirement: Payload mode may use active alerts as supporting context
|
||||
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
|
||||
|
||||
#### Scenario: Prometheus returns multiple active alerts
|
||||
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
|
||||
- **THEN** the prompt permits mentioning those alerts only as related risk or context
|
||||
- **AND** the final report target remains the supplied alert
|
||||
@@ -0,0 +1,44 @@
|
||||
# aiops-traceable-diagnosis-entry Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change aiops-traceable-diagnosis-entry. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: AIOps endpoint accepts optional alert input
|
||||
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
|
||||
|
||||
#### Scenario: Caller supplies alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
|
||||
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
|
||||
|
||||
#### Scenario: Caller omits alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without a body
|
||||
- **THEN** the system still starts the default AIOps alert-analysis flow
|
||||
|
||||
### Requirement: AIOps session id is traceable
|
||||
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
|
||||
|
||||
#### Scenario: Request includes session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
|
||||
- **THEN** the created `diagnosis_session.session_id` equals that value
|
||||
- **AND** the SSE stream includes the same session id
|
||||
|
||||
#### Scenario: Request omits session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
|
||||
- **THEN** the system generates a session id
|
||||
- **AND** the SSE stream includes the generated session id
|
||||
|
||||
### Requirement: AIOps report is persisted for trace replay
|
||||
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
|
||||
|
||||
#### Scenario: AIOps report is generated
|
||||
- **WHEN** the AIOps planner/executor flow returns a final report
|
||||
- **THEN** the corresponding diagnosis session is marked successful
|
||||
- **AND** `diagnosis_session.answer` stores the final report
|
||||
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
|
||||
|
||||
### Requirement: AIOps trace uses existing evidence tables
|
||||
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
|
||||
|
||||
#### Scenario: AIOps uses evidence tools
|
||||
- **WHEN** the AIOps flow calls available evidence tools
|
||||
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
|
||||
@@ -190,3 +190,17 @@ Verifier facts SHALL be linkable to the evidence summaries used during verificat
|
||||
- **WHEN** the ChatService persists `verifier_evaluation`
|
||||
- **THEN** it SHALL include `traceability_version`
|
||||
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
|
||||
|
||||
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
|
||||
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
|
||||
|
||||
#### Scenario: Failed evidence remains a verifier-visible gap
|
||||
- **WHEN** an evidence-bearing tool invocation fails
|
||||
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
|
||||
- **AND** the verifier flow SHALL continue without crashing
|
||||
|
||||
#### Scenario: Deduped retrievals do not count as fresh support
|
||||
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
|
||||
- **THEN** those entries SHALL be treated as no-new-evidence
|
||||
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
|
||||
|
||||
|
||||
@@ -0,0 +1,135 @@
|
||||
# diagnosis-eval-harness Specification
|
||||
|
||||
## Purpose
|
||||
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
|
||||
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
|
||||
|
||||
#### Scenario: Case definition includes expected evidence
|
||||
- **WHEN** an evaluation case is defined
|
||||
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
|
||||
|
||||
#### Scenario: Case definition can express forbidden behavior
|
||||
- **WHEN** a case has known unsafe behavior
|
||||
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate diagnosis traces
|
||||
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
|
||||
|
||||
#### Scenario: Evidence coverage validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
|
||||
|
||||
#### Scenario: Verifier evaluation validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
|
||||
- **AND** the verdict SHALL be one of the case's allowed verdicts
|
||||
|
||||
#### Scenario: Answer keyword validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
|
||||
|
||||
#### Scenario: Degraded output validation
|
||||
- **WHEN** a trace verdict is `REJECT`
|
||||
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
|
||||
|
||||
### Requirement: Evaluation harness SHALL report quality and cost signals
|
||||
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
|
||||
|
||||
#### Scenario: JSON report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
|
||||
|
||||
#### Scenario: Markdown report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
|
||||
|
||||
#### Scenario: Aggregate metrics
|
||||
- **WHEN** multiple cases are evaluated
|
||||
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
|
||||
|
||||
### Requirement: Evaluation harness SHALL support offline fixture mode
|
||||
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
|
||||
|
||||
#### Scenario: Fixture trace evaluation
|
||||
- **WHEN** the evaluator is run against a directory of trace fixture files
|
||||
- **THEN** it SHALL evaluate each trace file against its matching case definition
|
||||
- **AND** it SHALL not require a running application service
|
||||
|
||||
#### Scenario: Missing fixture is reported clearly
|
||||
- **WHEN** a case has no matching trace fixture
|
||||
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
|
||||
|
||||
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
|
||||
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
|
||||
|
||||
#### Scenario: Every case resolves to a fixture file
|
||||
- **WHEN** the evaluator loads the fixed case definition file
|
||||
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
|
||||
|
||||
#### Scenario: Fixture files are loadable as diagnosis traces
|
||||
- **WHEN** each referenced fixture is loaded
|
||||
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
|
||||
|
||||
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
|
||||
The system SHALL preserve a generated baseline report for the full fixed fixture set.
|
||||
|
||||
#### Scenario: Baseline report includes all fixed cases
|
||||
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
|
||||
- **THEN** the report SHALL include one result for every fixed case
|
||||
|
||||
#### Scenario: Baseline report is reviewable
|
||||
- **WHEN** the baseline report is written
|
||||
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
|
||||
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes fixtures or evaluator rules
|
||||
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
||||
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
||||
|
||||
#### Scenario: Aggregate regression detection
|
||||
- **WHEN** the current report has a lower pass rate than the baseline report
|
||||
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
||||
|
||||
#### Scenario: Cost signal detection
|
||||
- **WHEN** average tool-call count or average duration changes between reports
|
||||
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
||||
|
||||
#### Scenario: Verdict distribution comparison
|
||||
- **WHEN** verdict counts differ between reports
|
||||
- **THEN** the diff SHALL record the verdict distribution changes
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare case-level report results
|
||||
The system SHALL compare case results by case id and report actionable per-case changes.
|
||||
|
||||
#### Scenario: Case pass/fail regression
|
||||
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
||||
- **THEN** the diff SHALL record a regression for that case
|
||||
|
||||
#### Scenario: Evidence coverage regression
|
||||
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
||||
- **THEN** the diff SHALL record a regression naming the case and tool
|
||||
|
||||
#### Scenario: Missing case detection
|
||||
- **WHEN** a baseline case is absent from the current report
|
||||
- **THEN** the diff SHALL record a regression for the missing case
|
||||
|
||||
#### Scenario: New case detection
|
||||
- **WHEN** a current report contains a case absent from the baseline
|
||||
- **THEN** the diff SHALL record the case as a non-regression change
|
||||
|
||||
### Requirement: Evaluation harness SHALL report baseline diff results
|
||||
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
||||
|
||||
#### Scenario: JSON diff output
|
||||
- **WHEN** a baseline diff is written as JSON
|
||||
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
||||
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
@@ -0,0 +1,86 @@
|
||||
# evidence-trace-hardening Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change evidence-trace-hardening. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
|
||||
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
|
||||
#### Scenario: Common evidence fields are always persisted
|
||||
- **WHEN** an evidence-bearing tool finishes a call
|
||||
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
|
||||
|
||||
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
|
||||
- **WHEN** `lookup_knowledge` persists a tool invocation
|
||||
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
|
||||
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
|
||||
|
||||
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
|
||||
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
|
||||
|
||||
#### Scenario: Tool failure is preserved as failure
|
||||
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
|
||||
- **THEN** the persisted row SHALL set `success=false`
|
||||
- **AND** it SHALL preserve an `error_message` explaining the failure
|
||||
|
||||
#### Scenario: No usable evidence is preserved without pretending success
|
||||
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
|
||||
- **THEN** the persisted contract SHALL preserve that the call completed
|
||||
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
|
||||
|
||||
#### Scenario: Deduped retrieval remains auditable
|
||||
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
|
||||
- **THEN** the persisted row SHALL preserve the dedup reason
|
||||
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
|
||||
|
||||
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
|
||||
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
|
||||
|
||||
#### Scenario: Failed evidence calls remain visible in the summary
|
||||
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
|
||||
- **THEN** the summary SHALL retain them
|
||||
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
|
||||
|
||||
#### Scenario: No-hit and deduped calls do not upgrade evidence level
|
||||
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
|
||||
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
|
||||
- **AND** their counts SHALL still be reflected in the merged summary entry
|
||||
|
||||
#### Scenario: Successful evidence keeps the strongest available support
|
||||
- **WHEN** multiple rows for the same tool and topic domain are merged
|
||||
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
|
||||
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
|
||||
|
||||
### Requirement: ChatService SHALL degrade predictably on verifier output failures
|
||||
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
|
||||
|
||||
#### Scenario: Missing verifier output falls back to LOW_CONFID
|
||||
- **WHEN** the verifier step completes without a usable `verifier_output`
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
|
||||
|
||||
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
|
||||
- **WHEN** the verifier returns malformed or non-parseable JSON
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the fallback SHALL still persist a verifier-evaluation record
|
||||
|
||||
#### Scenario: REJECT output hides unverified raw answer text
|
||||
- **WHEN** the final verifier decision is `REJECT`
|
||||
- **THEN** the user-facing output SHALL use the degraded template
|
||||
- **AND** it SHALL NOT pass through the raw executor answer
|
||||
|
||||
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
|
||||
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
|
||||
|
||||
#### Scenario: Evidence recorder contract is tested offline
|
||||
- **WHEN** the test suite runs the focused recorder tests
|
||||
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
|
||||
|
||||
#### Scenario: Trace summary hardening is tested offline
|
||||
- **WHEN** the test suite runs the focused trace-summary tests
|
||||
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
|
||||
|
||||
#### Scenario: Verifier fallback behavior is tested offline
|
||||
- **WHEN** the test suite runs the focused `ChatService` fallback tests
|
||||
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
|
||||
|
||||
@@ -35,3 +35,32 @@ The project SHALL include an end-to-end acceptance case that demonstrates start-
|
||||
#### Scenario: Reviewer follows the acceptance case
|
||||
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
|
||||
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
|
||||
|
||||
### Requirement: MVP demo SHALL provide an interview runbook
|
||||
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
|
||||
|
||||
#### Scenario: Walkthrough explains the demo story
|
||||
- **WHEN** a developer opens the interview walkthrough
|
||||
- **THEN** it SHALL explain the user question, Agent flow, evidence tools, verifier judgment, trace API, feedback, and eval baseline connection
|
||||
|
||||
#### Scenario: Walkthrough stays scoped to existing capabilities
|
||||
- **WHEN** the walkthrough describes the demo
|
||||
- **THEN** it SHALL avoid claiming unsupported runtime behavior or new production features
|
||||
|
||||
### Requirement: MVP demo SHALL provide executable local demo scripts
|
||||
The MVP demo SHALL provide scripts and request payloads for running the payment-timeout case through existing local APIs.
|
||||
|
||||
#### Scenario: Demo script sends the fixed diagnosis request
|
||||
- **WHEN** the demo script is executed against a running local service
|
||||
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
|
||||
|
||||
#### Scenario: Demo script captures review artifacts
|
||||
- **WHEN** the demo script finishes successfully
|
||||
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
|
||||
|
||||
### Requirement: MVP demo SHALL provide a trace inspection checklist
|
||||
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
|
||||
|
||||
#### Scenario: Checklist maps fields to interview claims
|
||||
- **WHEN** a developer reviews a trace response
|
||||
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
|
||||
|
||||
@@ -131,13 +131,15 @@ public class QueryLogsTools {
|
||||
output.setMessage(String.format("共有 %d 个可用的日志主题。建议使用默认地域 'ap-guangzhou' 或省略 region 参数", topics.size()));
|
||||
|
||||
String response = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
|
||||
recordInvocation(startTime, "get_available_log_topics", null, null, null, response, true, null, "logs");
|
||||
recordInvocation("get_available_log_topics", startTime, "get_available_log_topics", null, null, null,
|
||||
response, true, null, "logs", ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED);
|
||||
return response;
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("获取日志主题列表失败", e);
|
||||
String response = "{\"success\":false,\"message\":\"获取日志主题列表失败: " + e.getMessage() + "\"}";
|
||||
recordInvocation(startTime, "get_available_log_topics", null, null, null, response, false, e.getMessage(), "logs");
|
||||
recordInvocation("get_available_log_topics", startTime, "get_available_log_topics", null, null, null,
|
||||
response, false, e.getMessage(), "logs", ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
|
||||
return response;
|
||||
}
|
||||
}
|
||||
@@ -191,8 +193,9 @@ public class QueryLogsTools {
|
||||
} else {
|
||||
// 真实模式:调用 CLS API(这里预留接口,后续实现)
|
||||
String response = buildErrorResponse("CLS 真实查询尚未实现,请启用 mock 模式进行测试");
|
||||
recordInvocation(startTime, safeQuery, region, logTopic, actualLimit, response, false,
|
||||
"CLS 真实查询尚未实现,请启用 mock 模式进行测试", normalizeTopicDomain(logTopic));
|
||||
recordInvocation("query_logs", startTime, safeQuery, region, logTopic, actualLimit, response, false,
|
||||
"CLS 真实查询尚未实现,请启用 mock 模式进行测试", normalizeTopicDomain(logTopic),
|
||||
ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
|
||||
return response;
|
||||
}
|
||||
|
||||
@@ -208,23 +211,26 @@ public class QueryLogsTools {
|
||||
|
||||
String jsonResult = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
|
||||
logger.info("日志查询完成: 找到 {} 条日志", logEntries.size());
|
||||
recordInvocation(startTime, safeQuery, region, logTopic, actualLimit, jsonResult,
|
||||
!logEntries.isEmpty(), logEntries.isEmpty() ? "未找到匹配的日志" : null,
|
||||
normalizeTopicDomain(logTopic));
|
||||
recordInvocation("query_logs", startTime, safeQuery, region, logTopic, actualLimit, jsonResult,
|
||||
true, null, normalizeTopicDomain(logTopic),
|
||||
logEntries.isEmpty()
|
||||
? ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE
|
||||
: ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED);
|
||||
|
||||
return jsonResult;
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("查询日志失败", e);
|
||||
String response = buildErrorResponse("查询失败: " + e.getMessage());
|
||||
recordInvocation(startTime, safeQuery, region, logTopic, actualLimit, response, false,
|
||||
e.getMessage(), normalizeTopicDomain(logTopic));
|
||||
recordInvocation("query_logs", startTime, safeQuery, region, logTopic, actualLimit, response, false,
|
||||
e.getMessage(), normalizeTopicDomain(logTopic), ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
|
||||
return response;
|
||||
}
|
||||
}
|
||||
|
||||
private void recordInvocation(long startTime, String query, String region, String logTopic, Integer limit,
|
||||
String output, boolean success, String errorMessage, String topicDomain) {
|
||||
private void recordInvocation(String toolName, long startTime, String query, String region, String logTopic, Integer limit,
|
||||
String output, boolean success, String errorMessage, String topicDomain,
|
||||
String evidenceStatus) {
|
||||
Map<String, Object> input = new HashMap<>();
|
||||
input.put("query", query == null || query.isBlank() ? "DEFAULT_QUERY" : query);
|
||||
if (region != null) {
|
||||
@@ -239,13 +245,15 @@ public class QueryLogsTools {
|
||||
input.put("mock_enabled", mockEnabled);
|
||||
|
||||
toolInvocationRecorder.recordEvidenceTool(
|
||||
"query_logs",
|
||||
toolName,
|
||||
input,
|
||||
output,
|
||||
success,
|
||||
startTime,
|
||||
errorMessage,
|
||||
topicDomain
|
||||
topicDomain,
|
||||
evidenceStatus,
|
||||
Map.of("log_topic", logTopic == null ? "" : logTopic)
|
||||
);
|
||||
}
|
||||
|
||||
|
||||
@@ -81,7 +81,7 @@ public class QueryMetricsTools {
|
||||
|
||||
if (!"success".equals(result.getStatus())) {
|
||||
String response = buildErrorResponse("Prometheus API 返回非成功状态: " + result.getStatus(), result.getError());
|
||||
recordInvocation(startTime, response, false, result.getError());
|
||||
recordInvocation(startTime, response, false, result.getError(), ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
|
||||
return response;
|
||||
}
|
||||
|
||||
@@ -119,19 +119,22 @@ public class QueryMetricsTools {
|
||||
|
||||
String jsonResult = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
|
||||
logger.info("Prometheus 告警查询完成: 找到 {} 个告警", simplifiedAlerts.size());
|
||||
recordInvocation(startTime, jsonResult, true, null);
|
||||
recordInvocation(startTime, jsonResult, true, null,
|
||||
simplifiedAlerts.isEmpty()
|
||||
? ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE
|
||||
: ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED);
|
||||
|
||||
return jsonResult;
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("查询 Prometheus 告警失败", e);
|
||||
String response = buildErrorResponse("查询失败", e.getMessage());
|
||||
recordInvocation(startTime, response, false, e.getMessage());
|
||||
recordInvocation(startTime, response, false, e.getMessage(), ToolInvocationRecorder.EVIDENCE_STATUS_FAILED);
|
||||
return response;
|
||||
}
|
||||
}
|
||||
|
||||
private void recordInvocation(long startTime, String output, boolean success, String errorMessage) {
|
||||
private void recordInvocation(long startTime, String output, boolean success, String errorMessage, String evidenceStatus) {
|
||||
toolInvocationRecorder.recordEvidenceTool(
|
||||
"query_metrics",
|
||||
Map.of("query", "active_prometheus_alerts", "mock_enabled", mockEnabled),
|
||||
@@ -139,7 +142,9 @@ public class QueryMetricsTools {
|
||||
success,
|
||||
startTime,
|
||||
errorMessage,
|
||||
"prometheus_alerts"
|
||||
"prometheus_alerts",
|
||||
evidenceStatus,
|
||||
Map.of("metric_family", "prometheus_alerts")
|
||||
);
|
||||
}
|
||||
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user