Files
SuperBizAgent-java/interview/aiops-lightweight-verifier.md
T

54 lines
2.2 KiB
Markdown

# AIOps Lightweight Verifier
## What Changed
AIOps now has a deterministic post-run quality gate.
After the final AIOps report is persisted, the service evaluates:
- whether the final report exists and is not trivially short
- whether a payload-targeted report mentions the supplied alert and service
- whether evidence tools such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted
The result is stored under:
```text
diagnosis_session.self_evaluation.aiops_rule_evaluation
```
The trace API returns this payload through the existing session self-evaluation field.
## Why Rule-Based First
This is not a full LLM verifier yet.
The first AIOps quality risks are concrete and easy to check with rules:
- Did the report stay focused on the payload?
- Did the run use evidence tools?
- Did the system produce a usable final report?
Rule evaluation is stable, cheap, and easy to explain. It also avoids adding another hidden model call to the AIOps flow before the current trace contract is mature.
## Verdicts
The evaluator emits:
```text
PASS
WARN
FAIL
```
`FAIL` is reserved for critical issues such as a missing or too-short report. Missing payload focus terms or missing evidence tools currently produce `WARN`, because valid reports may use slightly different wording or evidence may be unavailable in a mock/demo environment.
## Interview Answer
If asked why AIOps has a verifier now:
> Chat already has an LLM verifier because the user questions are open-ended. For AIOps, I started with a lighter rule-based verifier because the first quality checks are very concrete: payload focus, evidence coverage, and report completeness. The evaluation is persisted into `self_evaluation`, so the trace can show not only what the Agent did, but also whether the output passed basic quality gates.
If asked why not use the Chat verifier directly:
> AIOps verification is different from Chat verification. It needs to check alert scope, evidence tool coverage, and whether unrelated active alerts were over-expanded. Reusing the Chat verifier directly would blur those semantics. The rule-based evaluator gives us a stable first quality gate; a later AIOps LLM verifier can build on the same trace contract.