feat(harness): complete protocol repair stop and archive ISS-016

Add repairable INVALID_PROGRESS_PROTOCOL observations, independent
PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths.
Archive the OpenSpec change after syncing main specs and devflow.
This commit is contained in:
zhuyongxin
2026-07-27 19:10:07 +08:00
parent 5c369f3b6c
commit 38f781b157
44 changed files with 977 additions and 80 deletions
@@ -14,7 +14,6 @@ The system SHALL expose `EVIDENCE_FOUND`, `NO_EVIDENCE`, and `ERROR` as Agent-fa
#### Scenario: Tool execution fails
- **WHEN** schema, authorization, execution, or projection fails
- **THEN** the Agent-facing result uses `evidence_status=ERROR` and SHALL NOT report `NO_EVIDENCE`
### Requirement: Referencable Tool results SHALL use the framework Tool Call ID
Each `EVIDENCE_FOUND` or `NO_EVIDENCE` result SHALL contain the non-blank `tool_call_id` supplied by the framework Tool Call request. Agent inputs SHALL NOT contain `tool_call_id`, and contract code SHALL NOT generate, replace, or derive a second call ID.
@@ -25,18 +24,16 @@ Each `EVIDENCE_FOUND` or `NO_EVIDENCE` result SHALL contain the non-blank `tool_
#### Scenario: Framework ID is invalid
- **WHEN** the Tool request has a missing or invalid framework Tool Call ID
- **THEN** the call returns an `ERROR` result without inventing a referencable ID
### Requirement: RAG Tool contract SHALL expose only bounded document evidence
The RAG Request SHALL contain only `query`. The RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`.
The RAG business Request SHALL contain only `query`. The canonical RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, optional normalized `relevance_level`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`. The RAG projector SHALL accept upstream `relevanceLevel` or `relevance_level` and normalize recognized values without exposing raw relevance scores or retrieval traces.
#### Scenario: RAG evidence is serialized
- **WHEN** a RAG result contains a matching document excerpt
- **THEN** its JSON matches the frozen snake_case fields and excludes ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
- **WHEN** a RAG result contains a matching document excerpt and an upstream relevance level
- **THEN** its canonical JSON preserves the bounded evidence and normalized `relevance_level` while excluding ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
#### Scenario: RAG query has no evidence
- **WHEN** RAG executes successfully without a usable document excerpt
- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, and returns an empty evidence list
- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, returns an empty evidence list, and does not upgrade relevance into evidence
### Requirement: Log Tool contract SHALL use logical scope and retain Mock provenance
The log Request SHALL contain logical `topic`, `query`, and optional `lookback_minutes` only. The log Result SHALL contain `evidence_status`, `tool_call_id`, `source_kind`, complete query `scope`, `match_count`, `returned_count`, bounded `patterns`, bounded timeline `events`, and `truncated`. The initial logical topics SHALL be `APPLICATION`, `DATABASE_SLOW_QUERY`, and `SYSTEM_EVENTS`, and the initial source kind SHALL be `MOCK`.
@@ -47,7 +44,6 @@ The log Request SHALL contain logical `topic`, `query`, and optional `lookback_m
#### Scenario: Agent creates a log request
- **WHEN** the Agent requests log evidence
- **THEN** it selects a logical topic and lookback window without supplying region, physical TopicId, credentials, or result limit and without calling a Topic discovery Tool first
### Requirement: MySQL Tool contract SHALL expose a logical read-only query interface
The MySQL Request SHALL contain only logical `data_source`, parameterized `sql`, and `params`. The MySQL Result SHALL contain `evidence_status`, `tool_call_id`, `columns`, bounded structured `rows`, `returned_count`, and `truncated`; it SHALL NOT expose connection details, credentials, internal stack traces, or resource-limit controls.
@@ -58,17 +54,25 @@ The MySQL Request SHALL contain only logical `data_source`, parameterized `sql`,
#### Scenario: Agent creates a MySQL request
- **WHEN** the Agent requests business database evidence
- **THEN** it supplies a logical data source, SQL placeholders, and parameter values without supplying JDBC connection information or security policy
### Requirement: Tool descriptions SHALL be concise and implementation-neutral
The system SHALL define stable snake_case names and concise descriptions for `lookup_knowledge`, `query_logs`, and `query_mysql`. Each description SHALL state when to call the Tool, its minimal input, and what it cannot query, and SHALL NOT describe retrieval internals, infrastructure coordinates, credentials, audit storage, retries, result limits, or ranking implementation.
#### Scenario: Agent receives Tool definitions
- **WHEN** a future Agent adapter registers the frozen Tool definitions
- **THEN** the definitions describe available actions and boundaries without exposing Milvus, L0/L1, rerank, CLS region/TopicId, Redis, JDBC credentials, topK, limit, or Trace internals
### Requirement: Contract freeze SHALL NOT cut over the current runtime
This change SHALL add contract types, descriptions, and tests without changing the current `LookupKnowledgeTool`, `QueryLogsTools`, ChatService, AiOpsService, Controller, Tool registration, or Agent-visible runtime results.
#### Scenario: Stage 1 tests pass
- **WHEN** all ACI contract tests pass
- **THEN** the current public Chat and AIOps paths still execute the old Tool implementations until their later projector and cutover changes
### Requirement: Tool results SHALL have separate Harness and model views
Each successful evidence Tool result SHALL provide a Harness Control View and a bounded Model Observation derived from the same canonical result. The control view MAY contain returned counts, normalized scope, relevance, truncation and duplicate identity. The Model Observation SHALL contain only fields needed to understand and cite the result and SHALL NOT contain raw responses, internal scores, retrieval traces, duplicate fingerprints, counters, thresholds, budgets or store identities.
#### Scenario: Model receives RAG observation
- **WHEN** a canonical RAG result is READY
- **THEN** the model receives Tool Call ID, actual query scope, bounded evidence, evidence status, optional coarse relevance and truncation, but not raw scores or Harness counters
#### Scenario: Harness evaluates duplicate scope
- **WHEN** the same normalized Tool scope is requested again
- **THEN** the Harness can compare its control view identity without exposing that fingerprint to the model