feat: add aiops lightweight verifier

This commit is contained in:
aruo
2026-07-05 13:44:30 +08:00
parent 2658742119
commit ed267d753d
15 changed files with 417 additions and 5 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-05
@@ -0,0 +1,59 @@
## Context
Chat diagnosis has a Verifier Agent that writes structured evaluation into `diagnosis_session.self_evaluation`. AIOps currently focuses on payload scoping, evidence tools, and trace persistence, but it has no quality gate that checks whether the final report stayed on target or used evidence.
The next stage should add a low-risk quality gate before considering a full AIOps LLM verifier.
## Goals / Non-Goals
**Goals:**
- Evaluate AIOps final reports with deterministic rules.
- Persist the evaluation under a dedicated `aiops_rule_evaluation` self-evaluation key.
- Keep trace replay able to show whether AIOps output passed, warned, or failed basic quality checks.
- Add focused unit tests without requiring live LLMs or external tools.
**Non-Goals:**
- Do not add an AIOps Verifier Agent yet.
- Do not route/retry AIOps execution based on the evaluation result.
- Do not change `tool_invocation` schema.
- Do not require new database migrations.
## Decisions
### Decision 1: Rule-Based Before LLM Verifier
The first AIOps verifier is a deterministic evaluator, not an LLM agent.
Rationale:
- AIOps quality risks are concrete at this stage: payload focus, evidence coverage, and report presence.
- Rule evaluation is cheap, stable, and easy to explain in an interview.
- A full verifier agent can be added later once AIOps trace expectations are stable.
### Decision 2: Dedicated Self-Evaluation Channel
Persist under `aiops_rule_evaluation` instead of reusing `rule_evaluation` or `verifier_evaluation`.
Rationale:
- `verifier_evaluation` is already associated with Chat's LLM verifier.
- `rule_evaluation` may be used by generic diagnosis evaluation.
- A dedicated key avoids conflating AIOps-specific checks with other evaluation channels.
### Decision 3: Evaluate After Final Report Persistence
Run the evaluator when `persistFinalReport(...)` is called.
Rationale:
- It has access to the final report and session id.
- It can read persisted tool invocations for the same session.
- It does not disturb the Agent execution path.
## Risks / Trade-offs
- [Risk] Rule evaluation can miss semantic hallucinations. -> Mitigation: position it as lightweight AIOps quality gate, not full groundedness verification.
- [Risk] Strict keyword checks may warn on valid reports with different wording. -> Mitigation: use WARN for missing soft signals and FAIL only for critical absence.
- [Risk] Evaluation after report persistence does not trigger retries. -> Mitigation: keep routing unchanged in this phase; later changes can consume the verdict.
@@ -0,0 +1,26 @@
## Why
AIOps now has traceable payload scope control and improved RAG retrieval, but it still lacks a quality gate comparable to Chat's verifier. A lightweight rule-based verifier can check the most important AIOps risks without introducing another LLM agent.
## What Changes
- Add a rule-based AIOps evaluation service that checks final report quality after the AIOps flow completes.
- Persist the evaluation under `diagnosis_session.self_evaluation.aiops_rule_evaluation`.
- Evaluate payload focus, evidence-tool coverage, and basic report completeness.
- Expose the evaluation through the existing trace API self-evaluation payload.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `aiops-traceable-diagnosis-entry`: AIOps sessions include a lightweight rule evaluation for trace replay.
## Impact
- Affects AIOps session finalization and trace self-evaluation.
- Does not change AIOps API input, Agent flow topology, tool signatures, or database schema.
- Does not add an LLM verifier agent.
@@ -0,0 +1,22 @@
## ADDED Requirements
### Requirement: AIOps sessions SHALL persist lightweight rule evaluation
When an AIOps final report is persisted, the system SHALL evaluate it with deterministic AIOps-specific quality rules and store the result in session self-evaluation.
#### Scenario: Payload-focused report is evaluated
- **WHEN** an AIOps session has alert payload fields and a final report is persisted
- **THEN** the system SHALL evaluate whether the report mentions the supplied alert and service
- **AND** it SHALL store the result under `self_evaluation.aiops_rule_evaluation`
#### Scenario: Evidence coverage is evaluated
- **WHEN** an AIOps final report is evaluated
- **THEN** the system SHALL check whether evidence tool invocations such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted for the session
#### Scenario: Evaluation is traceable
- **WHEN** the diagnosis trace API returns an AIOps session
- **THEN** the session self-evaluation payload SHALL include `aiops_rule_evaluation` when it has been generated
#### Scenario: Evaluation uses stable verdicts
- **WHEN** AIOps rule evaluation completes
- **THEN** it SHALL produce a verdict from `PASS`, `WARN`, or `FAIL`
- **AND** it SHALL include check details and a human-readable rationale
@@ -0,0 +1,19 @@
## 1. Rule Evaluation
- [x] 1.1 Add an AIOps rule evaluation service with PASS/WARN/FAIL verdicts.
- [x] 1.2 Check payload focus, evidence-tool coverage, and report completeness.
## 2. AIOps Integration
- [x] 2.1 Persist AIOps rule evaluation when the final AIOps report is saved.
- [x] 2.2 Make trace summary indicate that AIOps rule evaluation exists.
## 3. Tests And Docs
- [x] 3.1 Add focused unit tests for the evaluator and AIOps integration.
- [x] 3.2 Add interview notes for the lightweight AIOps verifier.
## 4. Verification
- [x] 4.1 Run focused service tests.
- [x] 4.2 Validate the OpenSpec change and review git scope.