feat(harness): add rag and log projections

This commit is contained in:
zhuyongxin
2026-07-21 20:12:18 +08:00
parent 0dbdd7d8d3
commit 3e602781d6
20 changed files with 1140 additions and 0 deletions
@@ -0,0 +1,69 @@
# Design: single-react-rag-log-projections
## Architecture
```text
typed ACI request + framework envelope
|
v
RagToolAdapter / QueryLogsToolAdapter
| (legacy executor is internal and request controls are injected here)
v
ToolBoundary.execute(..., raw -> requestAwareProjector.project(...))
|
+--> CanonicalInvocationStore (complete raw + bounded agent result)
+--> ToolBoundaryResult (bounded result only)
```
The adapters are the only place that knows how the current legacy tools are called. They parse the typed ACI request before invoking the boundary, serialize the legacy result as raw JSON, and bind the typed request and framework `tool_call_id` to the projector closure. `ToolBoundary` remains responsible for all lifecycle, ownership, budget, size and error rules.
## Request-aware projection decision
The generic `ToolResultProjector.project(String rawResponse)` remains unchanged for stage 3A compatibility. Stage 3B introduces specialized projector methods:
```java
ProjectedToolResult project(RagToolRequest request, String toolCallId, String rawResponse);
ProjectedToolResult project(QueryLogsRequest request, String toolCallId,
LogQueryScope scope, String rawResponse);
```
Adapters pass these methods through a lambda to the generic boundary. This preserves request scope without adding request fields to `ProjectedToolResult`, changing the generic boundary API, or relying on raw JSON conventions.
## RAG projection
- Parse only `evidenceBlocks` from the legacy response.
- Accept an evidence block only when its content/excerpt is non-blank.
- Use `source` as the stable `document_id`; fall back to title, then a deterministic ordinal only for malformed legacy records.
- Deduplicate by document ID while retaining the first exact excerpt.
- Project only `document_id`, `source`, `title`, `breadcrumb`, and bounded `excerpt`.
- Ignore `contextPack`, `retrievalTrace`, `rerankTrace`, scores, hit reasons, domains, and messages.
- Set `NO_EVIDENCE` when no usable excerpt remains.
## Query-log projection
- Parse only the legacy `logs` array and success/total fields needed for semantics.
- Map `LogTopic` to the existing Mock topic names inside the adapter (`APPLICATION -> application-logs`, `DATABASE_SLOW_QUERY -> database-slow-query`, `SYSTEM_EVENTS -> system-events`).
- Inject the legacy default region and an internal source limit; neither is part of the ACI request or result.
- Build `LogQueryScope` from the logical request and `lookback_minutes`, using the adapter clock for `end_time`.
- Always record `source_kind=MOCK` for this stage.
- Exclude `instance` and `metrics`; redact credentials, host/pod identifiers, PIDs, IPs, SQL literals and stack-like suffixes from messages.
- Aggregate patterns by sanitized level, service and normalized message, with count and first/last timestamps.
- Sample timeline events deterministically when the source exceeds the event bound.
- Set `NO_EVIDENCE` for a successful empty `logs` array; legacy error responses become projection errors.
## Bounds
`ToolProjectionLimits` defines maximum evidence/items, excerpt/message characters, patterns, events and total Agent projection UTF-8 bytes. Collection bounds set `truncated=true`. If the serialized projection still exceeds the total budget, the projector removes the last timeline/pattern/evidence items until it fits; if no valid bounded result can be produced, it throws and the boundary records `PROJECTION_ERROR`.
## Compatibility and ownership
- No changes to `LookupKnowledgeTool`, `QueryLogsTools`, `ToolInvocationRecorder`, JPA entities, ChatService, AiOpsService, Controller or public HTTP/SSE payloads.
- Redis access remains in `RedisCanonicalInvocationStore`.
- Legacy tools remain the source of raw data until a later Diagnosis Agent cutover.
## Risks and mitigations
- Legacy output shape drift: strict JSON parsing and focused malformed-response tests fail closed.
- Redaction can remove useful details: only sensitive tokens/identifiers are replaced; service, level and sanitized message context remain.
- Scope clock skew: adapter uses an injected `Clock` and tests use a fixed clock.
- Projection size pressure: deterministic collection bounds and explicit `truncated` prevent silent overflow.
@@ -0,0 +1,51 @@
# Proposal: single-react-rag-log-projections
## Problem
阶段 3A 已提供统一 `ToolBoundary` 和 canonical invocation store,但 RAG 与 `query_logs` 仍只有旧工具输出。旧输出包含检索轨迹、上下文打包、基础设施字段、实例信息和未受约束的日志内容,不能直接作为单体 Diagnosis Agent 的 ACI evidence projection。若分别实现状态机或存储,会重新引入重复的生命周期、Run ownership 和错误处理逻辑。
## Proposed change
- 新增 RAG result projector,将旧 `LookupKnowledgeTool` JSON 转换为冻结的 `RagToolResult`。
- 新增 query-log result projector,将现有 Mock `QueryLogsTools` JSON 转换为冻结的 `QueryLogsToolResult`。
- projector 负责精确摘录、文档去重、证据/事件/模式数量上限、整体字符预算、时间线抽样和脱敏。
- projector 由请求感知适配层调用,使 `topic`、`query` 和 `lookback_minutes` 保留在完整 `scope` 中;不扩展 Agent request,不让 raw response 直接进入 Agent。
- 两个适配器均通过阶段 3A `ToolBoundary` 执行,沿用框架 `tool_call_id`、RunContext、canonical store、`PROJECTING/READY/ERROR` 与 `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 语义。
- 保留旧工具、旧 recorder、Chat/AIOps、Controller 和公开协议不变;本阶段只新增投影/适配层与 focused tests。
## Scope
### In scope
- RAG projector and adapter.
- Query-log projector and adapter for existing Mock source.
- Request-aware projection entry point required to preserve log scope.
- Contract, sanitization, deduplication, truncation, no-evidence and boundary integration tests.
### Out of scope
- Real CLS/MCP log integration.
- MySQL projector (stage 3C).
- Diagnosis Agent cutover, Guards, Chat use-case cutover, SSE changes, or legacy recorder cleanup.
- Changes to the frozen ACI DTO field names.
## Constraints and risks
- Raw payload remains Harness-only canonical data and is never returned to the Agent.
- `NO_EVIDENCE` is scoped to the recorded query and time window; it is not a health claim.
- Log messages may contain credentials, hostnames, pod IDs, SQL literals, or stack traces; projection must redact these before Agent exposure.
- A projector must not silently truncate raw data. It may bound Agent-facing collections and excerpts while setting `truncated=true`.
- The existing generic `ToolResultProjector` API has no request argument. The adapter will carry the typed request alongside the projector invocation, keeping the generic boundary reusable and avoiding a raw JSON scope convention.
## Acceptance direction
- Serialized RAG output contains only the frozen fields and bounded exact excerpts.
- Serialized log output contains complete logical scope, `source_kind=MOCK`, bounded patterns/events, distinct match and returned counts, and no infrastructure controls.
- Both projectors produce `NO_EVIDENCE` for successful empty results and `ERROR` for malformed/unsafe input.
- Boundary tests prove framework ID preservation, canonical lifecycle reuse, raw isolation, and safe errors.
## Context sources
- ISS-014 stage 3B requirements.
- `aci-evidence-tool-contracts` and `canonical-tool-invocation-store` specifications.
- Existing `LookupKnowledgeTool`, `QueryLogsTools`, and frozen contract tests.
@@ -0,0 +1,73 @@
# rag-log-projections Specification
## Purpose
Define bounded RAG and Mock query-log projections that execute through the stage 3A ToolBoundary and expose only the frozen ACI contracts to the Agent.
## ADDED Requirements
### Requirement: RAG projection SHALL expose bounded document evidence only
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result.
#### Scenario: RAG evidence is projected
- WHEN the legacy response contains usable evidence blocks
- THEN the result contains the original query, framework `tool_call_id`, deduplicated evidence with exact bounded excerpts, a returned count, and an explicit truncation flag
#### Scenario: RAG result has no usable excerpt
- WHEN the legacy response has no usable evidence block
- THEN the boundary returns `READY` with `evidence_status=NO_EVIDENCE`, the original query and an empty evidence list
#### Scenario: RAG duplicate documents
- WHEN multiple blocks identify the same source/document
- THEN only the first document is returned and the result remains deterministic
### Requirement: Query-log projection SHALL preserve logical scope and Mock provenance
The query-log adapter SHALL accept only logical topic, query and optional lookback minutes, execute the existing Mock source through ToolBoundary, and project `source_kind=MOCK`, complete scope, match count, returned count, bounded patterns, bounded timeline events and truncation.
#### Scenario: Mock logs are projected
- WHEN the legacy Mock response contains log entries
- THEN the result includes the logical topic/query/time window, sanitized pattern aggregates, sanitized timeline events, distinct match and returned counts, and `source_kind=MOCK`
#### Scenario: Empty Mock result
- WHEN the legacy Mock response is successful with an empty log array
- THEN the boundary returns `READY` with `evidence_status=NO_EVIDENCE` and preserves the full query scope
#### Scenario: Legacy log error
- WHEN the legacy response indicates failure or is malformed
- THEN the boundary returns `ERROR` with a bounded projection error and no Agent result
### Requirement: Projection SHALL redact and bound sensitive log content
The log projector SHALL exclude instance and metrics fields and redact credentials, token-like values, host/pod identifiers, PIDs, IP addresses, SQL literals and stack-like suffixes from Agent-facing messages. It SHALL enforce per-item, collection and total UTF-8 bounds and set `truncated=true` when any bound removes data.
#### Scenario: Sensitive fields are present
- WHEN a log entry includes instance, metrics, a secret assignment, a pod identifier or a SQL literal
- THEN none of those raw values are present in the projected Agent result
#### Scenario: Projection exceeds collection budget
- WHEN evidence, patterns, events or excerpts exceed configured bounds
- THEN the result is valid JSON, contains only bounded collections, and sets `truncated=true`
### Requirement: Adapters SHALL reuse canonical boundary ownership
Both adapters SHALL pass the framework `tool_call_id` and RunContext to the existing ToolBoundary and SHALL NOT create a second ID, write a parallel store, return raw responses, or modify legacy audit paths.
#### Scenario: Framework ID and Run are valid
- WHEN an adapter executes a valid request
- THEN the canonical record and bounded result use the exact framework ID and the current Run ID
#### Scenario: Boundary rejects the request
- WHEN Run ownership, authorization, read-only, duplicate, budget or size preflight fails
- THEN the adapter returns the boundary's safe error without invoking the legacy tool or projector
@@ -0,0 +1,24 @@
# Tasks: single-react-rag-log-projections
## 1. Projection limits and request-aware interfaces
- [x] 1.1 Add `ToolProjectionLimits` with explicit item, excerpt/message, collection and total UTF-8 limits.
- [x] 1.2 Add request-aware projector methods and adapter-facing typed execution helpers without changing the generic ToolBoundary API.
## 2. RAG projection and adapter
- [x] 2.1 Implement JSON parsing, usable excerpt filtering, source/document ID fallback, deduplication, exact excerpt bounding and `RagToolResult` serialization.
- [x] 2.2 Implement the RAG adapter that maps typed request JSON to the legacy knowledge tool and invokes ToolBoundary with the framework ID.
- [x] 2.3 Add tests for field exclusion, no evidence, duplicate documents, excerpt/total bounds, malformed raw response and canonical ID preservation.
## 3. Query-log projection and adapter
- [x] 3.1 Implement logical topic mapping, injected legacy Mock invocation, lookback scope construction and `source_kind=MOCK`.
- [x] 3.2 Implement redaction, pattern aggregation, deterministic timeline sampling, counts and bounded serialization.
- [x] 3.3 Add tests for scope, provenance, pattern/timeline bounds, redaction, empty result, malformed/error response and no topic discovery call.
## 4. Boundary integration and verification
- [x] 4.1 Route both adapters through the existing ToolBoundary and verify raw data remains canonical-only.
- [x] 4.2 Run focused projector/adapter/boundary tests and the prior stage regression suite.
- [x] 4.3 Validate OpenSpec, update devflow acceptance/evidence, archive the change and commit before stage 3C.
@@ -0,0 +1,23 @@
# rag-log-projections Specification
## Purpose
Define bounded RAG and Mock query-log projections that execute through the stage 3A ToolBoundary and expose only the frozen ACI contracts to the Agent.
## Requirements
### Requirement: RAG projection SHALL expose bounded document evidence only
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result.
### Requirement: Query-log projection SHALL preserve logical scope and Mock provenance
The query-log adapter SHALL accept only logical topic, query and optional lookback minutes, execute the existing Mock source through ToolBoundary, and project `source_kind=MOCK`, complete scope, match count, returned count, bounded patterns, bounded timeline events and truncation.
### Requirement: Projection SHALL redact and bound sensitive log content
The log projector SHALL exclude instance and metrics fields and redact credentials, token-like values, host/pod identifiers, PIDs, IP addresses, SQL literals and stack-like suffixes from Agent-facing messages. It SHALL enforce per-item, collection and total UTF-8 bounds and set `truncated=true` when any bound removes data.
### Requirement: Adapters SHALL reuse canonical boundary ownership
Both adapters SHALL pass the framework `tool_call_id` and RunContext to the existing ToolBoundary and SHALL NOT create a second ID, write a parallel store, return raw responses, or modify legacy audit paths.