## Context Chat diagnosis has a Verifier Agent that writes structured evaluation into `diagnosis_session.self_evaluation`. AIOps currently focuses on payload scoping, evidence tools, and trace persistence, but it has no quality gate that checks whether the final report stayed on target or used evidence. The next stage should add a low-risk quality gate before considering a full AIOps LLM verifier. ## Goals / Non-Goals **Goals:** - Evaluate AIOps final reports with deterministic rules. - Persist the evaluation under a dedicated `aiops_rule_evaluation` self-evaluation key. - Keep trace replay able to show whether AIOps output passed, warned, or failed basic quality checks. - Add focused unit tests without requiring live LLMs or external tools. **Non-Goals:** - Do not add an AIOps Verifier Agent yet. - Do not route/retry AIOps execution based on the evaluation result. - Do not change `tool_invocation` schema. - Do not require new database migrations. ## Decisions ### Decision 1: Rule-Based Before LLM Verifier The first AIOps verifier is a deterministic evaluator, not an LLM agent. Rationale: - AIOps quality risks are concrete at this stage: payload focus, evidence coverage, and report presence. - Rule evaluation is cheap, stable, and easy to explain in an interview. - A full verifier agent can be added later once AIOps trace expectations are stable. ### Decision 2: Dedicated Self-Evaluation Channel Persist under `aiops_rule_evaluation` instead of reusing `rule_evaluation` or `verifier_evaluation`. Rationale: - `verifier_evaluation` is already associated with Chat's LLM verifier. - `rule_evaluation` may be used by generic diagnosis evaluation. - A dedicated key avoids conflating AIOps-specific checks with other evaluation channels. ### Decision 3: Evaluate After Final Report Persistence Run the evaluator when `persistFinalReport(...)` is called. Rationale: - It has access to the final report and session id. - It can read persisted tool invocations for the same session. - It does not disturb the Agent execution path. ## Risks / Trade-offs - [Risk] Rule evaluation can miss semantic hallucinations. -> Mitigation: position it as lightweight AIOps quality gate, not full groundedness verification. - [Risk] Strict keyword checks may warn on valid reports with different wording. -> Mitigation: use WARN for missing soft signals and FAIL only for critical absence. - [Risk] Evaluation after report persistence does not trigger retries. -> Mitigation: keep routing unchanged in this phase; later changes can consume the verdict.