Files
SuperBizAgent-java/openspec/changes/aiops-lightweight-verifier/design.md
T

2.5 KiB

Context

Chat diagnosis has a Verifier Agent that writes structured evaluation into diagnosis_session.self_evaluation. AIOps currently focuses on payload scoping, evidence tools, and trace persistence, but it has no quality gate that checks whether the final report stayed on target or used evidence.

The next stage should add a low-risk quality gate before considering a full AIOps LLM verifier.

Goals / Non-Goals

Goals:

  • Evaluate AIOps final reports with deterministic rules.
  • Persist the evaluation under a dedicated aiops_rule_evaluation self-evaluation key.
  • Keep trace replay able to show whether AIOps output passed, warned, or failed basic quality checks.
  • Add focused unit tests without requiring live LLMs or external tools.

Non-Goals:

  • Do not add an AIOps Verifier Agent yet.
  • Do not route/retry AIOps execution based on the evaluation result.
  • Do not change tool_invocation schema.
  • Do not require new database migrations.

Decisions

Decision 1: Rule-Based Before LLM Verifier

The first AIOps verifier is a deterministic evaluator, not an LLM agent.

Rationale:

  • AIOps quality risks are concrete at this stage: payload focus, evidence coverage, and report presence.
  • Rule evaluation is cheap, stable, and easy to explain in an interview.
  • A full verifier agent can be added later once AIOps trace expectations are stable.

Decision 2: Dedicated Self-Evaluation Channel

Persist under aiops_rule_evaluation instead of reusing rule_evaluation or verifier_evaluation.

Rationale:

  • verifier_evaluation is already associated with Chat's LLM verifier.
  • rule_evaluation may be used by generic diagnosis evaluation.
  • A dedicated key avoids conflating AIOps-specific checks with other evaluation channels.

Decision 3: Evaluate After Final Report Persistence

Run the evaluator when persistFinalReport(...) is called.

Rationale:

  • It has access to the final report and session id.
  • It can read persisted tool invocations for the same session.
  • It does not disturb the Agent execution path.

Risks / Trade-offs

  • [Risk] Rule evaluation can miss semantic hallucinations. -> Mitigation: position it as lightweight AIOps quality gate, not full groundedness verification.
  • [Risk] Strict keyword checks may warn on valid reports with different wording. -> Mitigation: use WARN for missing soft signals and FAIL only for critical absence.
  • [Risk] Evaluation after report persistence does not trigger retries. -> Mitigation: keep routing unchanged in this phase; later changes can consume the verdict.