refactor(harness): freeze single-agent contracts

This commit is contained in:
zhuyongxin
2026-07-21 17:33:25 +08:00
parent 30d3296043
commit 58c39107c5
47 changed files with 2607 additions and 24 deletions
@@ -0,0 +1 @@
Devflow archive prepared and stage verification passed on 2026-07-21.
@@ -0,0 +1 @@
Committed after strict validation on 2026-07-21.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-21
@@ -0,0 +1,29 @@
# Stage 0 Acceptance Evidence
## Static Verification
- Contract package is dependency-free from Redis, JPA, Controller, Agent state and existing Hook classes.
- Repository secret scan covers tracked worktree files and reports no known plaintext credential matches.
- `scripts/query_mysql.py` requires `SUPERBIZ_MYSQL_PASSWORD` and exits before connecting when it is absent.
- Spring AI 1.1.7 `SpringAiRetryProperties` bytecode shows a default `maxAttempts` value of 10; stage 2 must set underlying retries to one attempt and keep retry ownership in Harness.
## Script Verification
- `mvn -q '-Dtest=HarnessContractTest' test` - passed.
- `mvn -q '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
- `mvn -q '-Dtest=HarnessContractTest,ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
- `openspec validate single-react-design-freeze --strict` - passed.
- `git diff --check` - passed; only existing Windows line-ending warnings were reported.
- Secret scan for known committed key/password patterns - zero matches.
- `python scripts/query_mysql.py "SELECT 1"` without `SUPERBIZ_MYSQL_PASSWORD` - exited before connecting with the expected missing-variable error.
## Runtime Behavior
- Public Chat runtime was not switched in stage zero.
- No live model, Redis, MySQL or Milvus E2E was run; full live E2E remains stage 7 scope.
## External Security Prerequisite
- Plaintext credentials previously present in the repository must be rotated in their respective MySQL, Redis, DeepSeek, SiliconFlow and Milvus systems by the credential owner.
- Repository changes can prove removal but cannot prove provider-side rotation.
- Stage 3C and stage 7 must not claim live security/E2E acceptance until required environment variables contain rotated credentials.
@@ -0,0 +1,25 @@
# Brief: single-react-design-freeze
## Background
ISS-014 将当前 Chat 多 Agent/Hook/ThreadLocal 主链路重构为一个 Diagnosis ReAct Agent、一个确定性 Harness 和一个隔离 SemanticGuard。阶段 0 先冻结后续 10 个实施 change 共同依赖的契约和安全边界。
## Goal
产出可执行、可测试、可归档的 contract types、失败语义、安全前置和阶段门禁,同时保持现有公开 Chat 运行行为不变。
## Scope
- 类型化 Draft、Knowledge Answer、Fallback、previous turn 和状态枚举。
- Tool ID、双状态、取消、重试、Redis canonical record 和 MySQL 安全设计冻结。
- 明文脚本凭据清理和 focused baseline。
- ISS-014、OpenSpec、devflow 对齐。
## Non-Goals
- 不实现或接入新 Harness/Agent/Guard。
- 不切换 `/api/chat`、不删除旧链路、不运行 live E2E。
## Source PRD
`mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` 是本 change 的完整 PRD 和总设计来源,不复制为第二份 PRD。
@@ -0,0 +1,80 @@
# Decisions
## Entry
- Parent issue: `ISS-014`
- Change: `single-react-design-freeze`
- Scale: `complex`
- Interface impact: future L4; this change freezes contracts without switching runtime behavior.
- Capability sources: sm-flow built-in clarify/context/propose, `grill-with-docs`, `openspec-propose`, `zoom-out`, `openspec-apply-change`, `openspec-archive-change`.
## Context Evidence
- `session-run-trace-isolation` established Chat Session and Diagnosis Run as separate lifecycles and made `runId` the Trace ownership key.
- `verifier-evidence-reference-fidelity` established that no-evidence is a scoped negative observation, not proof that a problem does not exist.
- `executor-composer-final-answer` established deterministic safe fallback boundaries and prohibited unfiltered raw output from reaching users.
- `modular-rag-pipeline` established that retrieval trace and context packing are audit details rather than direct facts.
- Current Spring AI Alibaba `ToolCallRequest` already provides `tool_call_id`; Harness must validate and persist it rather than create a second identity.
- Current Spring AI `ChatModel.call(Prompt)` has no cancellation token, so cancellation must be expressed as layered, observable semantics rather than an unsupported absolute guarantee.
## Question Pool
| ID | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Terminology | Does `tool_call_id` use the framework ID or a Harness-generated ID? | user-interview | confirmed: framework ID |
| Q2 | Terminology | Are invocation lifecycle and evidence outcome separate fields? | user-interview | confirmed: `status` + `evidence_status` |
| Q3 | Boundary | Is ISS-014 one umbrella Issue with independent OpenSpec changes? | user-interview | confirmed: one Issue, 11 changes |
| Q4 | Boundary | May stage 4 publish before Guards exist? | user-interview | confirmed: no; public cutover only in 6B |
| Q5 | Lifecycle | What is the durable source of truth for `previous_turn` and `last_intent`? | user-interview | confirmed: `diagnosis_run` safe published result |
| Q6 | Contract | How is KNOWLEDGE_QUERY citation validation represented? | evidence-driven | resolved: structured answer items with exact RAG bindings |
| Q7 | Cancellation | What cancellation guarantees are technically enforceable? | evidence-driven | resolved: layered cancellation, no false hard-cancel claim |
| Q8 | Acceptance | Are Apply, Archive and phase Git commit pre-authorized? | user-interview | confirmed: yes, for all phases |
## Confirmed Decisions
- `tool_call_id` is the framework Tool Call protocol ID. Harness validates non-empty, bounded, safe characters and Run-local uniqueness; duplicate/invalid/missing IDs fail closed.
- Redis `status=PROJECTING/READY/ERROR` represents invocation/projector lifecycle.
- `evidence_status=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` represents result semantics.
- `NO_EVIDENCE` may only support `NEGATIVE_OBSERVATION` within the exact query scope; it cannot prove absence, exclusion or health.
- Stage 4 and 6A remain internal. Stage 6B performs the only public Chat cutover after stage 5 release gates pass.
- ISS-014 remains the umbrella Issue. Eleven independent changes run serially; each must complete sm-flow, OpenSpec archive and Git commit before the next starts.
- Full live E2E is deferred to stage 7; earlier stages run focused verification proportional to their change.
- KNOWLEDGE_QUERY uses a dedicated structured draft: each answer item binds the single lookup `tool_call_id` and one or more returned `document_id` values; Harness validates exact membership before rendering `answer + references + limitations`. It does not reuse the Diagnosis Analysis schema and does not enter SemanticGuard in the first version.
- Cancellation is layered: mark cancellation requested, prevent new model/Tool rounds and any final Draft release, invoke framework interruption, cancel owned Tool/JDBC work where supported, and rely on configured HTTP timeouts for an already-blocking synchronous model call. Run finalization is atomic and late results are discarded.
- Public SSE `done` is not emitted after the client has disconnected; internal Run state still reaches `CANCELLED`.
- `diagnosis_run` is the durable source for `intent`, `release_outcome` and `published_result`. Only the latest same-session `DIAGNOSIS + SUCCESS` record with a non-null safe published result may become `previous_turn`; `FALLBACK/FAILED/CANCELLED` remain auditable but are excluded.
- `published_result` stores only `user_query/published_conclusion/scope/limitations/source_documents`; it excludes Tool Call IDs, raw evidence, full Draft and SemanticGuard audit reasons.
## Evidence-Driven Findings To Report
- The framework already exposes `ToolInterceptor`, structured output types, tool execution timeout, model/tool call limit hooks and `ReactAgent.interrupt`; later Harness stages should reuse these extension points.
- Spring AI model dependencies include retry support, while current application configuration does not explicitly freeze all retry layers; stage 0 must define a retry inventory and stage 2 must enforce it.
- Current `DiagnosisRun` stores a text answer and generic status but has no explicit `intent`, `release_outcome` or structured published result; Q5 must be resolved before the previous-turn contract is executable.
- Current KNOWLEDGE_QUERY target behavior promises citation validation, but the issue only defines the Diagnosis Draft binding schema; the committed spec must add the dedicated answer-item contract described above.
- Spring AI retry auto-configuration defaults `maxAttempts` to 10. The target Harness retry matrix requires underlying model/HTTP retries to be set to one attempt, with Router and SemanticGuard retries performed only by Harness.
## OpenSpec Backfill
- All confirmed decisions above must appear in design/specs/tasks before `.committed` is created.
- All user-interview questions are confirmed; no pending decision blocks Commit.
## Architecture Audit
Current input flows from `ChatController` into `ChatService`, which owns routing, ReactAgent construction, multi-Agent orchestration and final rendering; tools persist evidence through `ToolInvocationRecorder`, while Hooks and ThreadLocal bridge Run and verifier state. Stage zero introduces only dependency-free contract types under `harness.contract`; those types must not depend on Controller, Redis, JPA, Spring Agent state or current Hook classes. Later stages move ownership in order: RunContext, invocation store, Tool-specific projection, Diagnosis Agent, Guards, application use case and finally the public SSE adapter. `diagnosis_run` remains durable Run ownership, Redis canonical invocation remains short-lived Harness ownership, and `agent_step/tool_invocation` remain durable audit detail. The principal risk is spec/runtime drift, mitigated by archiving only the stage-zero contract capability now and delaying modifications to existing runtime capabilities until their implementation changes.
## Cross-Artifact Alignment
| Chain | Status | Evidence |
|---|---|---|
| ISS-014/brief goals, scope and non-goals → proposal | aligned | Proposal limits stage zero to contracts, security and baseline with no public cutover. |
| proposal commitments → design | aligned | Design records every ID, status, Draft, fallback, previous-turn, retry, cancellation and phase-gate commitment. |
| design decisions → specs | aligned | The single stage-zero capability has testable requirements for every stable contract boundary. |
| specs observable behavior → tasks | aligned | Tasks create reusable types/tests, remove the secret, align artifacts and verify without switching runtime behavior. |
## Commit Gate Result
- Question pool covers terminology, boundary, lifecycle, contract, cancellation and acceptance.
- All user-interview items are explicitly confirmed.
- Evidence-driven conclusions were reported and written into design/spec/tasks.
- Interface impact is recorded as future L4; this change itself does not switch the public API.
- No devflow/OpenSpec conflict remains.
@@ -0,0 +1,77 @@
## Context
ISS-014 是一次 L4 Chat 重构的总设计来源,但实施被拆成 11 个必须串行归档的 OpenSpec changes。阶段 0 不切换公开协议或 Agent 运行链,只创建后续阶段复用的类型化契约、安全前置、失败语义和 focused baseline。
现有代码已经具备 `runId` Trace、工具调用审计、no-evidence 精确引用、Verifier/Composer fallback 和模块化 RAG,但这些能力分散在 `ChatService`、Hook、ThreadLocal、Tool 和 JSON 字符串中。Spring AI Alibaba 已提供 `ToolInterceptor`、结构化输出类型、工具执行超时、调用限额 Hook 和 `ReactAgent.interrupt`;Spring AI 底层 retry 默认最多 10 次,不能直接满足 ISS-014 的显式 Harness 重试矩阵。
## Goals / Non-Goals
**Goals:**
- 生成后续阶段可直接复用的 Java contract types 和枚举,不实现新 Agent 执行链。
- 冻结 Tool ID、双状态、Draft、Knowledge Answer、Fallback、previous turn、SSE、重试和取消语义。
- 冻结 MySQL fail-closed 允许子集和安全前置。
- 移除仓库脚本和主配置中的明文凭据并建立改造前 focused baseline。
- 保证 ISS-014、OpenSpec、devflow 术语和 11 个阶段门禁一致。
**Non-Goals:**
- 不接入 Harness、Diagnosis Agent、EvidenceGuard 或 SemanticGuard 运行时。
- 不修改 Controller 协议、旧 ChatService 行为、Redis invocation store 或数据库表。
- 不实现 RAG/日志/MySQL Tool 投影。
- 不运行完整 live E2E。
## Decisions
### Contract types are reusable runtime inputs
阶段 0 创建位于 `com.superbiz.agent.harness.contract` 的轻量 record/enum,而不是只写文档或引入 JSON Schema 引擎。后续 ReactAgent `outputType`、Harness validator、持久化和 SSE DTO 可以直接复用这些类型,减少同一字段在多个阶段重复定义。
### Framework Tool Call ID is canonical
`tool_call_id` 使用框架协议 ID。Harness 后续只校验非空、长度/字符安全和 Run 内唯一性,不生成第二套 ID。Redis Key 仍按 `runId + toolCallId` 隔离,真实性来自当前 Run 的 canonical record,而不是 ID 本身。
### Invocation lifecycle and evidence outcome are orthogonal
`InvocationStatus=PROJECTING/READY/ERROR` 只描述调用与投影生命周期;`EvidenceStatus=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 描述结果语义。`NO_EVIDENCE` 只允许绑定 `AnalysisKind=NEGATIVE_OBSERVATION`,并必须保留查询范围和零匹配信息。
### Diagnosis and knowledge answer contracts stay separate
Diagnosis Draft 使用 Analysis ID 与 Tool Call IDs;KNOWLEDGE_QUERY 使用 answer items,每项绑定一次 lookup 的 Tool Call ID 和返回的 document IDs。Harness 后续验证精确成员关系并确定性展开引用。首版 KNOWLEDGE_QUERY 不进入 SemanticGuard。
### Safe published context belongs to Diagnosis Run
`diagnosis_run` 是 `intent/release_outcome/published_result` 的持久化真理源。只有同 Session 最近一个 `DIAGNOSIS + SUCCESS` 且 published result 非空的 Run 可形成 previous turn。Published result 只含用户查询、已发布结论、范围、限制和 RAG 文档元数据。
### Cancellation is observable and layered
取消请求立即阻止新模型/Tool 轮次和最终 Draft 释放;框架中断、可控 Future/JDBC 取消尽力执行;已进入同步 `ChatModel.call` 的请求依靠底层 HTTP timeout。Run 终态通过原子状态转换保证唯一,晚到结果被丢弃,不宣称无法证明的底层硬取消。
### Harness owns all retries
底层 SDK/HTTP/数据库 retry 必须关闭或压为一次 attempt。Intent Router 和 SemanticGuard 的第二次 attempt 由 Harness 显式执行并审计。阶段 0 只冻结矩阵;阶段 2 实现执行器和配置。
### Stage boundaries are release boundaries
阶段 4 和 6A 仅内部运行;阶段 6B 在 Guards 已完成后原子切换公开入口。每个 change 必须 Archive 并 Git commit 后才能进入下一项,阶段 7 才执行统一 live E2E。
## Risks / Trade-offs
- [Risk] Contract records过早绑定实现细节 → Mitigation:只包含跨阶段稳定字段,不包含 Redis、JPA 或框架对象。
- [Risk] Spring AI provider 对 Tool Call ID 行为不同 → Mitigation:阶段 3A 使用 Fake Model 和当前 DeepSeek 路径验证,缺失或重复时 fail closed。
- [Risk] 同步模型调用不能立即取消 → Mitigation:明确 layered semantics、HTTP timeout 和晚到结果丢弃,不把状态更新等同底层资源已终止。
- [Risk] 设计冻结测试增加维护成本 → Mitigation:只保留 focused serialization/validation tests,不复制完整 E2E。
- [Risk] 明文凭据可能已泄露 → Mitigation:仓库中移除并要求外部轮换;轮换证据记录在 acceptance。
## Migration Plan
1. 创建 contract types、契约测试和阶段台账,不接运行链。
2. 清理查询脚本凭据,改为环境变量注入。
3. 记录 focused baseline;Archive 本 change。
4. 后续 10 个 change 逐步实现,并在各自 Archive 时同步对应运行 capability specs。
Rollback:阶段 0 没有公开行为变化;可删除新增 contract package/tests 并恢复文档。已经轮换的凭据不得回滚为旧值。
## Open Questions
无。所有影响实现的用户决策已经在 `decisions.md` 中确认。
@@ -0,0 +1,32 @@
## Why
当前 Chat 诊断把一个 ReAct 生命周期拆成多个 Agent、Hook、ThreadLocal 和重试分支,导致证据契约、失败语义、上下文预算和公开释放边界分散。进入分阶段重构前,需要先把 ISS-014 的跨阶段契约、安全前置和验收基线冻结为唯一可执行规格,避免后续 change 各自解释同一概念。
## What Changes
- 冻结单体 Diagnosis ReAct Agent、确定性 Harness、EvidenceGuard 和隔离 SemanticGuard 的职责边界。
- 冻结 Diagnosis Draft、Analysis、Conclusion、Fallback、Intent Router 和 SSE 事件契约。
- 冻结 Tool Call ID、调用生命周期 `status`、结果语义 `evidence_status`、Redis canonical invocation 和 ToolResultProjector 命名。
- 冻结预算、取消、重试、隐藏重试禁用和 Run 终态语义,但不在本阶段实现新运行链路。
- 冻结只读 MySQL Tool 的 JSqlParser 允许子集、静态 allowlist 和安全前置。
- 移除仓库脚本和主配置中的明文数据库、Redis、模型及向量服务凭据,并记录必须完成外部轮换。
- 建立改造前 focused baseline 和后续 11 个串行 sm-flow change 的阶段台账。
- **BREAKING(后续阶段实施)**:最终仅保留 `POST /api/chat` SSE、移除旧多 Agent/Graph Chat 主链路和旧 Tool Contract;本 change 不执行公开协议切换。
## Capabilities
### New Capabilities
- `single-react-diagnosis-harness`: 冻结单体诊断 Agent、Harness、证据状态、Guard、Fallback、预算、重试和阶段门禁的跨阶段基础契约。
### Modified Capabilities
- None. 本阶段不声明旧运行能力已经迁移;后续 change 在实现对应行为时再修改现有 capability specs。
## Impact
- 设计与规格:ISS-014、OpenSpec 主规格、devflow 词汇表和阶段归档台账。
- 契约测试:Draft/Fallback、状态语义、SSE、Router、预算、重试和 MySQL 安全基线。
- 安全:`scripts/query_mysql.py` 和 `application.yml` 中的明文凭据必须移除并在外部轮换。
- 后续代码范围:ChatController、ChatService、Agent/Hook、Tool、Redis、JPA/Flyway、静态前端、Trace/Eval fixtures。
- 本 change 不切换 Controller 协议、不接入新 Agent、不修改公开运行行为。
@@ -0,0 +1,75 @@
## ADDED Requirements
### Requirement: Design contracts SHALL separate deterministic control from diagnosis reasoning
The frozen contract set SHALL define Diagnosis Agent as the only diagnosis report author, Harness as deterministic execution control, EvidenceGuard as deterministic evidence validation, and SemanticGuard as an isolated single-turn semantic reviewer.
#### Scenario: Contract ownership is inspected
- **WHEN** a later phase reads the stage-zero contracts
- **THEN** no Harness contract assigns Planner, Executor, Composer, workflow routing, or diagnosis reasoning responsibilities to Harness
### Requirement: Tool invocation identity SHALL use the framework Tool Call ID
The contract SHALL use the framework-provided `tool_call_id` as the sole Tool invocation reference and SHALL require later Harness implementations to reject missing, invalid, or duplicate IDs within a Run.
#### Scenario: Duplicate Tool Call ID is proposed
- **WHEN** two Tool actions in one Run present the same framework Tool Call ID
- **THEN** the contract classifies the second action as an error and prohibits overwriting the first canonical invocation
### Requirement: Invocation status and evidence status SHALL be independent
The contract SHALL define `PROJECTING/READY/ERROR` as invocation lifecycle states and `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` as evidence result states.
#### Scenario: Successful query returns no evidence
- **WHEN** a Tool executes successfully and its bounded projection contains zero matching evidence
- **THEN** invocation status is `READY` and evidence status is `NO_EVIDENCE`
#### Scenario: No-evidence result is cited
- **WHEN** a Diagnosis Draft cites a `NO_EVIDENCE` Tool result
- **THEN** the Analysis kind MUST be `NEGATIVE_OBSERVATION` and MUST remain bounded to the Tool query scope
### Requirement: Diagnosis Draft SHALL expose typed report structure
The contract SHALL define conclusion, analysis items, action plan, recommendations, limitations and Tool Call bindings without exposing chain-of-thought or raw Tool payloads.
#### Scenario: Draft contains an analysis item
- **WHEN** the Diagnosis Agent emits a structured Draft
- **THEN** every Analysis has a unique analysis ID, a fixed Analysis kind and at least one Tool Call ID
### Requirement: Knowledge answers SHALL use exact RAG bindings
The KNOWLEDGE_QUERY contract SHALL represent the answer as bounded answer items whose references identify the single lookup Tool Call and returned document IDs.
#### Scenario: Knowledge answer cites an unknown document
- **WHEN** an answer item references a document ID absent from the bounded lookup result
- **THEN** later Harness validation rejects the answer instead of publishing the fabricated citation
### Requirement: Safe fallback SHALL use a fixed schema
The contract SHALL define stable fallback types for evidence validation failure, semantic unsupported and semantic unavailable outcomes, and SHALL exclude unvalidated Draft content and internal errors.
#### Scenario: Evidence validation fails twice
- **WHEN** initial validation and the single no-Tool structural repair both fail
- **THEN** the fallback type is `EVIDENCE_VALIDATION_FAILED` and verified sources are empty
### Requirement: Previous turn SHALL come from a safe durable Run result
The contract SHALL define `diagnosis_run` as the durable source of intent, release outcome and safe published result, and SHALL exclude fallback, failed and cancelled Runs from Diagnosis previous-turn selection.
#### Scenario: Latest Run is a fallback
- **WHEN** the latest same-session Run ended with `FALLBACK`
- **THEN** it is not used as Diagnosis previous turn and selection continues to the latest eligible `DIAGNOSIS + SUCCESS` Run
### Requirement: Cancellation SHALL be layered and observable
The contract SHALL distinguish cancellation request, prevention of new work, framework interruption, cancellable Tool work and HTTP timeout for already-blocking synchronous model calls.
#### Scenario: Client disconnects during a model call
- **WHEN** an SSE client disconnects while a synchronous model call is in flight
- **THEN** the system prevents later Draft release, requests interruption, records an internal cancelled terminal state and discards any late model result
### Requirement: Retry attempts SHALL be owned by Harness
The contract SHALL require underlying SDK, HTTP and database retry layers to execute one attempt, while Harness explicitly owns any allowed Router or SemanticGuard retry.
#### Scenario: SemanticGuard returns an invalid schema
- **WHEN** the first SemanticGuard attempt returns an invalid structured result
- **THEN** Harness may execute one second attempt with the same verified snapshot and records both attempts
### Requirement: Phase gates SHALL remain serial
The implementation plan SHALL contain eleven independent OpenSpec changes and SHALL prohibit starting a change before its predecessor is archived and committed.
#### Scenario: Stage 4 completes internal Agent tests
- **WHEN** stage 4 passes its focused tests but stage 5 Guards are not implemented
- **THEN** the public Chat entry remains on the old path and stage 6B cutover is prohibited
@@ -0,0 +1,21 @@
## 1. Contract Model
- [x] 1.1 Add typed enums for intent, release outcome, invocation status, evidence status, analysis kind, semantic verdict, fallback type and SSE outcome.
- [x] 1.2 Add reusable records for Diagnosis Draft, Knowledge Answer Draft, safe fallback, published result and previous turn.
- [x] 1.3 Add focused serialization and contract-shape tests covering positive evidence, no-evidence negative observation and forbidden raw/internal fields.
## 2. Security And Configuration Baseline
- [x] 2.1 Remove tracked plaintext credentials from `scripts/query_mysql.py` and `application.yml`, require environment-based secrets, and ignore local secret files.
- [x] 2.2 Record the required external credential rotation and the Spring AI hidden-retry baseline without changing the public Chat runtime in this stage.
## 3. Design Freeze Alignment
- [x] 3.1 Keep ISS-014, OpenSpec artifacts and the devflow glossary aligned on the 11 serial changes, framework Tool Call ID, dual status fields and safe previous-turn source.
- [x] 3.2 Add an architecture audit and cross-artifact alignment result to `decisions.md`.
- [x] 3.3 Create the sm-flow `.committed` marker after proposal/design/specs/tasks and all decision gates pass.
## 4. Verification
- [x] 4.1 Run focused contract tests and the smallest existing Chat/evidence baseline needed to prove stage zero did not switch runtime behavior.
- [x] 4.2 Run strict OpenSpec validation and record commands, results and unverified external rotation in acceptance evidence.
@@ -0,0 +1,78 @@
# single-react-diagnosis-harness Specification
## Purpose
TBD - created by archiving change single-react-design-freeze. Update Purpose after archive.
## Requirements
### Requirement: Design contracts SHALL separate deterministic control from diagnosis reasoning
The frozen contract set SHALL define Diagnosis Agent as the only diagnosis report author, Harness as deterministic execution control, EvidenceGuard as deterministic evidence validation, and SemanticGuard as an isolated single-turn semantic reviewer.
#### Scenario: Contract ownership is inspected
- **WHEN** a later phase reads the stage-zero contracts
- **THEN** no Harness contract assigns Planner, Executor, Composer, workflow routing, or diagnosis reasoning responsibilities to Harness
### Requirement: Tool invocation identity SHALL use the framework Tool Call ID
The contract SHALL use the framework-provided `tool_call_id` as the sole Tool invocation reference and SHALL require later Harness implementations to reject missing, invalid, or duplicate IDs within a Run.
#### Scenario: Duplicate Tool Call ID is proposed
- **WHEN** two Tool actions in one Run present the same framework Tool Call ID
- **THEN** the contract classifies the second action as an error and prohibits overwriting the first canonical invocation
### Requirement: Invocation status and evidence status SHALL be independent
The contract SHALL define `PROJECTING/READY/ERROR` as invocation lifecycle states and `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` as evidence result states.
#### Scenario: Successful query returns no evidence
- **WHEN** a Tool executes successfully and its bounded projection contains zero matching evidence
- **THEN** invocation status is `READY` and evidence status is `NO_EVIDENCE`
#### Scenario: No-evidence result is cited
- **WHEN** a Diagnosis Draft cites a `NO_EVIDENCE` Tool result
- **THEN** the Analysis kind MUST be `NEGATIVE_OBSERVATION` and MUST remain bounded to the Tool query scope
### Requirement: Diagnosis Draft SHALL expose typed report structure
The contract SHALL define conclusion, analysis items, action plan, recommendations, limitations and Tool Call bindings without exposing chain-of-thought or raw Tool payloads.
#### Scenario: Draft contains an analysis item
- **WHEN** the Diagnosis Agent emits a structured Draft
- **THEN** every Analysis has a unique analysis ID, a fixed Analysis kind and at least one Tool Call ID
### Requirement: Knowledge answers SHALL use exact RAG bindings
The KNOWLEDGE_QUERY contract SHALL represent the answer as bounded answer items whose references identify the single lookup Tool Call and returned document IDs.
#### Scenario: Knowledge answer cites an unknown document
- **WHEN** an answer item references a document ID absent from the bounded lookup result
- **THEN** later Harness validation rejects the answer instead of publishing the fabricated citation
### Requirement: Safe fallback SHALL use a fixed schema
The contract SHALL define stable fallback types for evidence validation failure, semantic unsupported and semantic unavailable outcomes, and SHALL exclude unvalidated Draft content and internal errors.
#### Scenario: Evidence validation fails twice
- **WHEN** initial validation and the single no-Tool structural repair both fail
- **THEN** the fallback type is `EVIDENCE_VALIDATION_FAILED` and verified sources are empty
### Requirement: Previous turn SHALL come from a safe durable Run result
The contract SHALL define `diagnosis_run` as the durable source of intent, release outcome and safe published result, and SHALL exclude fallback, failed and cancelled Runs from Diagnosis previous-turn selection.
#### Scenario: Latest Run is a fallback
- **WHEN** the latest same-session Run ended with `FALLBACK`
- **THEN** it is not used as Diagnosis previous turn and selection continues to the latest eligible `DIAGNOSIS + SUCCESS` Run
### Requirement: Cancellation SHALL be layered and observable
The contract SHALL distinguish cancellation request, prevention of new work, framework interruption, cancellable Tool work and HTTP timeout for already-blocking synchronous model calls.
#### Scenario: Client disconnects during a model call
- **WHEN** an SSE client disconnects while a synchronous model call is in flight
- **THEN** the system prevents later Draft release, requests interruption, records an internal cancelled terminal state and discards any late model result
### Requirement: Retry attempts SHALL be owned by Harness
The contract SHALL require underlying SDK, HTTP and database retry layers to execute one attempt, while Harness explicitly owns any allowed Router or SemanticGuard retry.
#### Scenario: SemanticGuard returns an invalid schema
- **WHEN** the first SemanticGuard attempt returns an invalid structured result
- **THEN** Harness may execute one second attempt with the same verified snapshot and records both attempts
### Requirement: Phase gates SHALL remain serial
The implementation plan SHALL contain eleven independent OpenSpec changes and SHALL prohibit starting a change before its predecessor is archived and committed.
#### Scenario: Stage 4 completes internal Agent tests
- **WHEN** stage 4 passes its focused tests but stage 5 Guards are not implemented
- **THEN** the public Chat entry remains on the old path and stage 6B cutover is prohibited