feat: add aiops lightweight verifier
This commit is contained in:
@@ -0,0 +1,53 @@
|
||||
# AIOps Lightweight Verifier
|
||||
|
||||
## What Changed
|
||||
|
||||
AIOps now has a deterministic post-run quality gate.
|
||||
|
||||
After the final AIOps report is persisted, the service evaluates:
|
||||
|
||||
- whether the final report exists and is not trivially short
|
||||
- whether a payload-targeted report mentions the supplied alert and service
|
||||
- whether evidence tools such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted
|
||||
|
||||
The result is stored under:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
The trace API returns this payload through the existing session self-evaluation field.
|
||||
|
||||
## Why Rule-Based First
|
||||
|
||||
This is not a full LLM verifier yet.
|
||||
|
||||
The first AIOps quality risks are concrete and easy to check with rules:
|
||||
|
||||
- Did the report stay focused on the payload?
|
||||
- Did the run use evidence tools?
|
||||
- Did the system produce a usable final report?
|
||||
|
||||
Rule evaluation is stable, cheap, and easy to explain. It also avoids adding another hidden model call to the AIOps flow before the current trace contract is mature.
|
||||
|
||||
## Verdicts
|
||||
|
||||
The evaluator emits:
|
||||
|
||||
```text
|
||||
PASS
|
||||
WARN
|
||||
FAIL
|
||||
```
|
||||
|
||||
`FAIL` is reserved for critical issues such as a missing or too-short report. Missing payload focus terms or missing evidence tools currently produce `WARN`, because valid reports may use slightly different wording or evidence may be unavailable in a mock/demo environment.
|
||||
|
||||
## Interview Answer
|
||||
|
||||
If asked why AIOps has a verifier now:
|
||||
|
||||
> Chat already has an LLM verifier because the user questions are open-ended. For AIOps, I started with a lighter rule-based verifier because the first quality checks are very concrete: payload focus, evidence coverage, and report completeness. The evaluation is persisted into `self_evaluation`, so the trace can show not only what the Agent did, but also whether the output passed basic quality gates.
|
||||
|
||||
If asked why not use the Chat verifier directly:
|
||||
|
||||
> AIOps verification is different from Chat verification. It needs to check alert scope, evidence tool coverage, and whether unrelated active alerts were over-expanded. Reusing the Chat verifier directly would blur those semantics. The rule-based evaluator gives us a stable first quality gate; a later AIOps LLM verifier can build on the same trace contract.
|
||||
Reference in New Issue
Block a user