54 lines
2.2 KiB
Markdown
54 lines
2.2 KiB
Markdown
# AIOps Lightweight Verifier
|
|
|
|
## What Changed
|
|
|
|
AIOps now has a deterministic post-run quality gate.
|
|
|
|
After the final AIOps report is persisted, the service evaluates:
|
|
|
|
- whether the final report exists and is not trivially short
|
|
- whether a payload-targeted report mentions the supplied alert and service
|
|
- whether evidence tools such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted
|
|
|
|
The result is stored under:
|
|
|
|
```text
|
|
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
|
```
|
|
|
|
The trace API returns this payload through the existing session self-evaluation field.
|
|
|
|
## Why Rule-Based First
|
|
|
|
This is not a full LLM verifier yet.
|
|
|
|
The first AIOps quality risks are concrete and easy to check with rules:
|
|
|
|
- Did the report stay focused on the payload?
|
|
- Did the run use evidence tools?
|
|
- Did the system produce a usable final report?
|
|
|
|
Rule evaluation is stable, cheap, and easy to explain. It also avoids adding another hidden model call to the AIOps flow before the current trace contract is mature.
|
|
|
|
## Verdicts
|
|
|
|
The evaluator emits:
|
|
|
|
```text
|
|
PASS
|
|
WARN
|
|
FAIL
|
|
```
|
|
|
|
`FAIL` is reserved for critical issues such as a missing or too-short report. Missing payload focus terms or missing evidence tools currently produce `WARN`, because valid reports may use slightly different wording or evidence may be unavailable in a mock/demo environment.
|
|
|
|
## Interview Answer
|
|
|
|
If asked why AIOps has a verifier now:
|
|
|
|
> Chat already has an LLM verifier because the user questions are open-ended. For AIOps, I started with a lighter rule-based verifier because the first quality checks are very concrete: payload focus, evidence coverage, and report completeness. The evaluation is persisted into `self_evaluation`, so the trace can show not only what the Agent did, but also whether the output passed basic quality gates.
|
|
|
|
If asked why not use the Chat verifier directly:
|
|
|
|
> AIOps verification is different from Chat verification. It needs to check alert scope, evidence tool coverage, and whether unrelated active alerts were over-expanded. Reusing the Chat verifier directly would blur those semantics. The rule-based evaluator gives us a stable first quality gate; a later AIOps LLM verifier can build on the same trace contract.
|