Files
SuperBizAgent-java/interview/aiops-lightweight-verifier.md
T

2.2 KiB

AIOps Lightweight Verifier

What Changed

AIOps now has a deterministic post-run quality gate.

After the final AIOps report is persisted, the service evaluates:

  • whether the final report exists and is not trivially short
  • whether a payload-targeted report mentions the supplied alert and service
  • whether evidence tools such as lookup_knowledge, query_metrics, or query_logs were persisted

The result is stored under:

diagnosis_session.self_evaluation.aiops_rule_evaluation

The trace API returns this payload through the existing session self-evaluation field.

Why Rule-Based First

This is not a full LLM verifier yet.

The first AIOps quality risks are concrete and easy to check with rules:

  • Did the report stay focused on the payload?
  • Did the run use evidence tools?
  • Did the system produce a usable final report?

Rule evaluation is stable, cheap, and easy to explain. It also avoids adding another hidden model call to the AIOps flow before the current trace contract is mature.

Verdicts

The evaluator emits:

PASS
WARN
FAIL

FAIL is reserved for critical issues such as a missing or too-short report. Missing payload focus terms or missing evidence tools currently produce WARN, because valid reports may use slightly different wording or evidence may be unavailable in a mock/demo environment.

Interview Answer

If asked why AIOps has a verifier now:

Chat already has an LLM verifier because the user questions are open-ended. For AIOps, I started with a lighter rule-based verifier because the first quality checks are very concrete: payload focus, evidence coverage, and report completeness. The evaluation is persisted into self_evaluation, so the trace can show not only what the Agent did, but also whether the output passed basic quality gates.

If asked why not use the Chat verifier directly:

AIOps verification is different from Chat verification. It needs to check alert scope, evidence tool coverage, and whether unrelated active alerts were over-expanded. Reusing the Chat verifier directly would blur those semantics. The rule-based evaluator gives us a stable first quality gate; a later AIOps LLM verifier can build on the same trace contract.