refactor(harness): freeze single-agent contracts
This commit is contained in:
@@ -0,0 +1 @@
|
||||
Devflow archive prepared and stage verification passed on 2026-07-21.
|
||||
@@ -0,0 +1 @@
|
||||
Committed after strict validation on 2026-07-21.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,29 @@
|
||||
# Stage 0 Acceptance Evidence
|
||||
|
||||
## Static Verification
|
||||
|
||||
- Contract package is dependency-free from Redis, JPA, Controller, Agent state and existing Hook classes.
|
||||
- Repository secret scan covers tracked worktree files and reports no known plaintext credential matches.
|
||||
- `scripts/query_mysql.py` requires `SUPERBIZ_MYSQL_PASSWORD` and exits before connecting when it is absent.
|
||||
- Spring AI 1.1.7 `SpringAiRetryProperties` bytecode shows a default `maxAttempts` value of 10; stage 2 must set underlying retries to one attempt and keep retry ownership in Harness.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn -q '-Dtest=HarnessContractTest' test` - passed.
|
||||
- `mvn -q '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
|
||||
- `mvn -q '-Dtest=HarnessContractTest,ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
|
||||
- `openspec validate single-react-design-freeze --strict` - passed.
|
||||
- `git diff --check` - passed; only existing Windows line-ending warnings were reported.
|
||||
- Secret scan for known committed key/password patterns - zero matches.
|
||||
- `python scripts/query_mysql.py "SELECT 1"` without `SUPERBIZ_MYSQL_PASSWORD` - exited before connecting with the expected missing-variable error.
|
||||
|
||||
## Runtime Behavior
|
||||
|
||||
- Public Chat runtime was not switched in stage zero.
|
||||
- No live model, Redis, MySQL or Milvus E2E was run; full live E2E remains stage 7 scope.
|
||||
|
||||
## External Security Prerequisite
|
||||
|
||||
- Plaintext credentials previously present in the repository must be rotated in their respective MySQL, Redis, DeepSeek, SiliconFlow and Milvus systems by the credential owner.
|
||||
- Repository changes can prove removal but cannot prove provider-side rotation.
|
||||
- Stage 3C and stage 7 must not claim live security/E2E acceptance until required environment variables contain rotated credentials.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Brief: single-react-design-freeze
|
||||
|
||||
## Background
|
||||
|
||||
ISS-014 将当前 Chat 多 Agent/Hook/ThreadLocal 主链路重构为一个 Diagnosis ReAct Agent、一个确定性 Harness 和一个隔离 SemanticGuard。阶段 0 先冻结后续 10 个实施 change 共同依赖的契约和安全边界。
|
||||
|
||||
## Goal
|
||||
|
||||
产出可执行、可测试、可归档的 contract types、失败语义、安全前置和阶段门禁,同时保持现有公开 Chat 运行行为不变。
|
||||
|
||||
## Scope
|
||||
|
||||
- 类型化 Draft、Knowledge Answer、Fallback、previous turn 和状态枚举。
|
||||
- Tool ID、双状态、取消、重试、Redis canonical record 和 MySQL 安全设计冻结。
|
||||
- 明文脚本凭据清理和 focused baseline。
|
||||
- ISS-014、OpenSpec、devflow 对齐。
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- 不实现或接入新 Harness/Agent/Guard。
|
||||
- 不切换 `/api/chat`、不删除旧链路、不运行 live E2E。
|
||||
|
||||
## Source PRD
|
||||
|
||||
`mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` 是本 change 的完整 PRD 和总设计来源,不复制为第二份 PRD。
|
||||
@@ -0,0 +1,80 @@
|
||||
# Decisions
|
||||
|
||||
## Entry
|
||||
|
||||
- Parent issue: `ISS-014`
|
||||
- Change: `single-react-design-freeze`
|
||||
- Scale: `complex`
|
||||
- Interface impact: future L4; this change freezes contracts without switching runtime behavior.
|
||||
- Capability sources: sm-flow built-in clarify/context/propose, `grill-with-docs`, `openspec-propose`, `zoom-out`, `openspec-apply-change`, `openspec-archive-change`.
|
||||
|
||||
## Context Evidence
|
||||
|
||||
- `session-run-trace-isolation` established Chat Session and Diagnosis Run as separate lifecycles and made `runId` the Trace ownership key.
|
||||
- `verifier-evidence-reference-fidelity` established that no-evidence is a scoped negative observation, not proof that a problem does not exist.
|
||||
- `executor-composer-final-answer` established deterministic safe fallback boundaries and prohibited unfiltered raw output from reaching users.
|
||||
- `modular-rag-pipeline` established that retrieval trace and context packing are audit details rather than direct facts.
|
||||
- Current Spring AI Alibaba `ToolCallRequest` already provides `tool_call_id`; Harness must validate and persist it rather than create a second identity.
|
||||
- Current Spring AI `ChatModel.call(Prompt)` has no cancellation token, so cancellation must be expressed as layered, observable semantics rather than an unsupported absolute guarantee.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| ID | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Terminology | Does `tool_call_id` use the framework ID or a Harness-generated ID? | user-interview | confirmed: framework ID |
|
||||
| Q2 | Terminology | Are invocation lifecycle and evidence outcome separate fields? | user-interview | confirmed: `status` + `evidence_status` |
|
||||
| Q3 | Boundary | Is ISS-014 one umbrella Issue with independent OpenSpec changes? | user-interview | confirmed: one Issue, 11 changes |
|
||||
| Q4 | Boundary | May stage 4 publish before Guards exist? | user-interview | confirmed: no; public cutover only in 6B |
|
||||
| Q5 | Lifecycle | What is the durable source of truth for `previous_turn` and `last_intent`? | user-interview | confirmed: `diagnosis_run` safe published result |
|
||||
| Q6 | Contract | How is KNOWLEDGE_QUERY citation validation represented? | evidence-driven | resolved: structured answer items with exact RAG bindings |
|
||||
| Q7 | Cancellation | What cancellation guarantees are technically enforceable? | evidence-driven | resolved: layered cancellation, no false hard-cancel claim |
|
||||
| Q8 | Acceptance | Are Apply, Archive and phase Git commit pre-authorized? | user-interview | confirmed: yes, for all phases |
|
||||
|
||||
## Confirmed Decisions
|
||||
|
||||
- `tool_call_id` is the framework Tool Call protocol ID. Harness validates non-empty, bounded, safe characters and Run-local uniqueness; duplicate/invalid/missing IDs fail closed.
|
||||
- Redis `status=PROJECTING/READY/ERROR` represents invocation/projector lifecycle.
|
||||
- `evidence_status=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` represents result semantics.
|
||||
- `NO_EVIDENCE` may only support `NEGATIVE_OBSERVATION` within the exact query scope; it cannot prove absence, exclusion or health.
|
||||
- Stage 4 and 6A remain internal. Stage 6B performs the only public Chat cutover after stage 5 release gates pass.
|
||||
- ISS-014 remains the umbrella Issue. Eleven independent changes run serially; each must complete sm-flow, OpenSpec archive and Git commit before the next starts.
|
||||
- Full live E2E is deferred to stage 7; earlier stages run focused verification proportional to their change.
|
||||
- KNOWLEDGE_QUERY uses a dedicated structured draft: each answer item binds the single lookup `tool_call_id` and one or more returned `document_id` values; Harness validates exact membership before rendering `answer + references + limitations`. It does not reuse the Diagnosis Analysis schema and does not enter SemanticGuard in the first version.
|
||||
- Cancellation is layered: mark cancellation requested, prevent new model/Tool rounds and any final Draft release, invoke framework interruption, cancel owned Tool/JDBC work where supported, and rely on configured HTTP timeouts for an already-blocking synchronous model call. Run finalization is atomic and late results are discarded.
|
||||
- Public SSE `done` is not emitted after the client has disconnected; internal Run state still reaches `CANCELLED`.
|
||||
- `diagnosis_run` is the durable source for `intent`, `release_outcome` and `published_result`. Only the latest same-session `DIAGNOSIS + SUCCESS` record with a non-null safe published result may become `previous_turn`; `FALLBACK/FAILED/CANCELLED` remain auditable but are excluded.
|
||||
- `published_result` stores only `user_query/published_conclusion/scope/limitations/source_documents`; it excludes Tool Call IDs, raw evidence, full Draft and SemanticGuard audit reasons.
|
||||
|
||||
## Evidence-Driven Findings To Report
|
||||
|
||||
- The framework already exposes `ToolInterceptor`, structured output types, tool execution timeout, model/tool call limit hooks and `ReactAgent.interrupt`; later Harness stages should reuse these extension points.
|
||||
- Spring AI model dependencies include retry support, while current application configuration does not explicitly freeze all retry layers; stage 0 must define a retry inventory and stage 2 must enforce it.
|
||||
- Current `DiagnosisRun` stores a text answer and generic status but has no explicit `intent`, `release_outcome` or structured published result; Q5 must be resolved before the previous-turn contract is executable.
|
||||
- Current KNOWLEDGE_QUERY target behavior promises citation validation, but the issue only defines the Diagnosis Draft binding schema; the committed spec must add the dedicated answer-item contract described above.
|
||||
- Spring AI retry auto-configuration defaults `maxAttempts` to 10. The target Harness retry matrix requires underlying model/HTTP retries to be set to one attempt, with Router and SemanticGuard retries performed only by Harness.
|
||||
|
||||
## OpenSpec Backfill
|
||||
|
||||
- All confirmed decisions above must appear in design/specs/tasks before `.committed` is created.
|
||||
- All user-interview questions are confirmed; no pending decision blocks Commit.
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
Current input flows from `ChatController` into `ChatService`, which owns routing, ReactAgent construction, multi-Agent orchestration and final rendering; tools persist evidence through `ToolInvocationRecorder`, while Hooks and ThreadLocal bridge Run and verifier state. Stage zero introduces only dependency-free contract types under `harness.contract`; those types must not depend on Controller, Redis, JPA, Spring Agent state or current Hook classes. Later stages move ownership in order: RunContext, invocation store, Tool-specific projection, Diagnosis Agent, Guards, application use case and finally the public SSE adapter. `diagnosis_run` remains durable Run ownership, Redis canonical invocation remains short-lived Harness ownership, and `agent_step/tool_invocation` remain durable audit detail. The principal risk is spec/runtime drift, mitigated by archiving only the stage-zero contract capability now and delaying modifications to existing runtime capabilities until their implementation changes.
|
||||
|
||||
## Cross-Artifact Alignment
|
||||
|
||||
| Chain | Status | Evidence |
|
||||
|---|---|---|
|
||||
| ISS-014/brief goals, scope and non-goals → proposal | aligned | Proposal limits stage zero to contracts, security and baseline with no public cutover. |
|
||||
| proposal commitments → design | aligned | Design records every ID, status, Draft, fallback, previous-turn, retry, cancellation and phase-gate commitment. |
|
||||
| design decisions → specs | aligned | The single stage-zero capability has testable requirements for every stable contract boundary. |
|
||||
| specs observable behavior → tasks | aligned | Tasks create reusable types/tests, remove the secret, align artifacts and verify without switching runtime behavior. |
|
||||
|
||||
## Commit Gate Result
|
||||
|
||||
- Question pool covers terminology, boundary, lifecycle, contract, cancellation and acceptance.
|
||||
- All user-interview items are explicitly confirmed.
|
||||
- Evidence-driven conclusions were reported and written into design/spec/tasks.
|
||||
- Interface impact is recorded as future L4; this change itself does not switch the public API.
|
||||
- No devflow/OpenSpec conflict remains.
|
||||
@@ -0,0 +1,77 @@
|
||||
## Context
|
||||
|
||||
ISS-014 是一次 L4 Chat 重构的总设计来源,但实施被拆成 11 个必须串行归档的 OpenSpec changes。阶段 0 不切换公开协议或 Agent 运行链,只创建后续阶段复用的类型化契约、安全前置、失败语义和 focused baseline。
|
||||
|
||||
现有代码已经具备 `runId` Trace、工具调用审计、no-evidence 精确引用、Verifier/Composer fallback 和模块化 RAG,但这些能力分散在 `ChatService`、Hook、ThreadLocal、Tool 和 JSON 字符串中。Spring AI Alibaba 已提供 `ToolInterceptor`、结构化输出类型、工具执行超时、调用限额 Hook 和 `ReactAgent.interrupt`;Spring AI 底层 retry 默认最多 10 次,不能直接满足 ISS-014 的显式 Harness 重试矩阵。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 生成后续阶段可直接复用的 Java contract types 和枚举,不实现新 Agent 执行链。
|
||||
- 冻结 Tool ID、双状态、Draft、Knowledge Answer、Fallback、previous turn、SSE、重试和取消语义。
|
||||
- 冻结 MySQL fail-closed 允许子集和安全前置。
|
||||
- 移除仓库脚本和主配置中的明文凭据并建立改造前 focused baseline。
|
||||
- 保证 ISS-014、OpenSpec、devflow 术语和 11 个阶段门禁一致。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不接入 Harness、Diagnosis Agent、EvidenceGuard 或 SemanticGuard 运行时。
|
||||
- 不修改 Controller 协议、旧 ChatService 行为、Redis invocation store 或数据库表。
|
||||
- 不实现 RAG/日志/MySQL Tool 投影。
|
||||
- 不运行完整 live E2E。
|
||||
|
||||
## Decisions
|
||||
|
||||
### Contract types are reusable runtime inputs
|
||||
|
||||
阶段 0 创建位于 `com.superbiz.agent.harness.contract` 的轻量 record/enum,而不是只写文档或引入 JSON Schema 引擎。后续 ReactAgent `outputType`、Harness validator、持久化和 SSE DTO 可以直接复用这些类型,减少同一字段在多个阶段重复定义。
|
||||
|
||||
### Framework Tool Call ID is canonical
|
||||
|
||||
`tool_call_id` 使用框架协议 ID。Harness 后续只校验非空、长度/字符安全和 Run 内唯一性,不生成第二套 ID。Redis Key 仍按 `runId + toolCallId` 隔离,真实性来自当前 Run 的 canonical record,而不是 ID 本身。
|
||||
|
||||
### Invocation lifecycle and evidence outcome are orthogonal
|
||||
|
||||
`InvocationStatus=PROJECTING/READY/ERROR` 只描述调用与投影生命周期;`EvidenceStatus=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 描述结果语义。`NO_EVIDENCE` 只允许绑定 `AnalysisKind=NEGATIVE_OBSERVATION`,并必须保留查询范围和零匹配信息。
|
||||
|
||||
### Diagnosis and knowledge answer contracts stay separate
|
||||
|
||||
Diagnosis Draft 使用 Analysis ID 与 Tool Call IDs;KNOWLEDGE_QUERY 使用 answer items,每项绑定一次 lookup 的 Tool Call ID 和返回的 document IDs。Harness 后续验证精确成员关系并确定性展开引用。首版 KNOWLEDGE_QUERY 不进入 SemanticGuard。
|
||||
|
||||
### Safe published context belongs to Diagnosis Run
|
||||
|
||||
`diagnosis_run` 是 `intent/release_outcome/published_result` 的持久化真理源。只有同 Session 最近一个 `DIAGNOSIS + SUCCESS` 且 published result 非空的 Run 可形成 previous turn。Published result 只含用户查询、已发布结论、范围、限制和 RAG 文档元数据。
|
||||
|
||||
### Cancellation is observable and layered
|
||||
|
||||
取消请求立即阻止新模型/Tool 轮次和最终 Draft 释放;框架中断、可控 Future/JDBC 取消尽力执行;已进入同步 `ChatModel.call` 的请求依靠底层 HTTP timeout。Run 终态通过原子状态转换保证唯一,晚到结果被丢弃,不宣称无法证明的底层硬取消。
|
||||
|
||||
### Harness owns all retries
|
||||
|
||||
底层 SDK/HTTP/数据库 retry 必须关闭或压为一次 attempt。Intent Router 和 SemanticGuard 的第二次 attempt 由 Harness 显式执行并审计。阶段 0 只冻结矩阵;阶段 2 实现执行器和配置。
|
||||
|
||||
### Stage boundaries are release boundaries
|
||||
|
||||
阶段 4 和 6A 仅内部运行;阶段 6B 在 Guards 已完成后原子切换公开入口。每个 change 必须 Archive 并 Git commit 后才能进入下一项,阶段 7 才执行统一 live E2E。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Contract records过早绑定实现细节 → Mitigation:只包含跨阶段稳定字段,不包含 Redis、JPA 或框架对象。
|
||||
- [Risk] Spring AI provider 对 Tool Call ID 行为不同 → Mitigation:阶段 3A 使用 Fake Model 和当前 DeepSeek 路径验证,缺失或重复时 fail closed。
|
||||
- [Risk] 同步模型调用不能立即取消 → Mitigation:明确 layered semantics、HTTP timeout 和晚到结果丢弃,不把状态更新等同底层资源已终止。
|
||||
- [Risk] 设计冻结测试增加维护成本 → Mitigation:只保留 focused serialization/validation tests,不复制完整 E2E。
|
||||
- [Risk] 明文凭据可能已泄露 → Mitigation:仓库中移除并要求外部轮换;轮换证据记录在 acceptance。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 创建 contract types、契约测试和阶段台账,不接运行链。
|
||||
2. 清理查询脚本凭据,改为环境变量注入。
|
||||
3. 记录 focused baseline;Archive 本 change。
|
||||
4. 后续 10 个 change 逐步实现,并在各自 Archive 时同步对应运行 capability specs。
|
||||
|
||||
Rollback:阶段 0 没有公开行为变化;可删除新增 contract package/tests 并恢复文档。已经轮换的凭据不得回滚为旧值。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。所有影响实现的用户决策已经在 `decisions.md` 中确认。
|
||||
@@ -0,0 +1,32 @@
|
||||
## Why
|
||||
|
||||
当前 Chat 诊断把一个 ReAct 生命周期拆成多个 Agent、Hook、ThreadLocal 和重试分支,导致证据契约、失败语义、上下文预算和公开释放边界分散。进入分阶段重构前,需要先把 ISS-014 的跨阶段契约、安全前置和验收基线冻结为唯一可执行规格,避免后续 change 各自解释同一概念。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 冻结单体 Diagnosis ReAct Agent、确定性 Harness、EvidenceGuard 和隔离 SemanticGuard 的职责边界。
|
||||
- 冻结 Diagnosis Draft、Analysis、Conclusion、Fallback、Intent Router 和 SSE 事件契约。
|
||||
- 冻结 Tool Call ID、调用生命周期 `status`、结果语义 `evidence_status`、Redis canonical invocation 和 ToolResultProjector 命名。
|
||||
- 冻结预算、取消、重试、隐藏重试禁用和 Run 终态语义,但不在本阶段实现新运行链路。
|
||||
- 冻结只读 MySQL Tool 的 JSqlParser 允许子集、静态 allowlist 和安全前置。
|
||||
- 移除仓库脚本和主配置中的明文数据库、Redis、模型及向量服务凭据,并记录必须完成外部轮换。
|
||||
- 建立改造前 focused baseline 和后续 11 个串行 sm-flow change 的阶段台账。
|
||||
- **BREAKING(后续阶段实施)**:最终仅保留 `POST /api/chat` SSE、移除旧多 Agent/Graph Chat 主链路和旧 Tool Contract;本 change 不执行公开协议切换。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `single-react-diagnosis-harness`: 冻结单体诊断 Agent、Harness、证据状态、Guard、Fallback、预算、重试和阶段门禁的跨阶段基础契约。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 本阶段不声明旧运行能力已经迁移;后续 change 在实现对应行为时再修改现有 capability specs。
|
||||
|
||||
## Impact
|
||||
|
||||
- 设计与规格:ISS-014、OpenSpec 主规格、devflow 词汇表和阶段归档台账。
|
||||
- 契约测试:Draft/Fallback、状态语义、SSE、Router、预算、重试和 MySQL 安全基线。
|
||||
- 安全:`scripts/query_mysql.py` 和 `application.yml` 中的明文凭据必须移除并在外部轮换。
|
||||
- 后续代码范围:ChatController、ChatService、Agent/Hook、Tool、Redis、JPA/Flyway、静态前端、Trace/Eval fixtures。
|
||||
- 本 change 不切换 Controller 协议、不接入新 Agent、不修改公开运行行为。
|
||||
+75
@@ -0,0 +1,75 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Design contracts SHALL separate deterministic control from diagnosis reasoning
|
||||
The frozen contract set SHALL define Diagnosis Agent as the only diagnosis report author, Harness as deterministic execution control, EvidenceGuard as deterministic evidence validation, and SemanticGuard as an isolated single-turn semantic reviewer.
|
||||
|
||||
#### Scenario: Contract ownership is inspected
|
||||
- **WHEN** a later phase reads the stage-zero contracts
|
||||
- **THEN** no Harness contract assigns Planner, Executor, Composer, workflow routing, or diagnosis reasoning responsibilities to Harness
|
||||
|
||||
### Requirement: Tool invocation identity SHALL use the framework Tool Call ID
|
||||
The contract SHALL use the framework-provided `tool_call_id` as the sole Tool invocation reference and SHALL require later Harness implementations to reject missing, invalid, or duplicate IDs within a Run.
|
||||
|
||||
#### Scenario: Duplicate Tool Call ID is proposed
|
||||
- **WHEN** two Tool actions in one Run present the same framework Tool Call ID
|
||||
- **THEN** the contract classifies the second action as an error and prohibits overwriting the first canonical invocation
|
||||
|
||||
### Requirement: Invocation status and evidence status SHALL be independent
|
||||
The contract SHALL define `PROJECTING/READY/ERROR` as invocation lifecycle states and `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` as evidence result states.
|
||||
|
||||
#### Scenario: Successful query returns no evidence
|
||||
- **WHEN** a Tool executes successfully and its bounded projection contains zero matching evidence
|
||||
- **THEN** invocation status is `READY` and evidence status is `NO_EVIDENCE`
|
||||
|
||||
#### Scenario: No-evidence result is cited
|
||||
- **WHEN** a Diagnosis Draft cites a `NO_EVIDENCE` Tool result
|
||||
- **THEN** the Analysis kind MUST be `NEGATIVE_OBSERVATION` and MUST remain bounded to the Tool query scope
|
||||
|
||||
### Requirement: Diagnosis Draft SHALL expose typed report structure
|
||||
The contract SHALL define conclusion, analysis items, action plan, recommendations, limitations and Tool Call bindings without exposing chain-of-thought or raw Tool payloads.
|
||||
|
||||
#### Scenario: Draft contains an analysis item
|
||||
- **WHEN** the Diagnosis Agent emits a structured Draft
|
||||
- **THEN** every Analysis has a unique analysis ID, a fixed Analysis kind and at least one Tool Call ID
|
||||
|
||||
### Requirement: Knowledge answers SHALL use exact RAG bindings
|
||||
The KNOWLEDGE_QUERY contract SHALL represent the answer as bounded answer items whose references identify the single lookup Tool Call and returned document IDs.
|
||||
|
||||
#### Scenario: Knowledge answer cites an unknown document
|
||||
- **WHEN** an answer item references a document ID absent from the bounded lookup result
|
||||
- **THEN** later Harness validation rejects the answer instead of publishing the fabricated citation
|
||||
|
||||
### Requirement: Safe fallback SHALL use a fixed schema
|
||||
The contract SHALL define stable fallback types for evidence validation failure, semantic unsupported and semantic unavailable outcomes, and SHALL exclude unvalidated Draft content and internal errors.
|
||||
|
||||
#### Scenario: Evidence validation fails twice
|
||||
- **WHEN** initial validation and the single no-Tool structural repair both fail
|
||||
- **THEN** the fallback type is `EVIDENCE_VALIDATION_FAILED` and verified sources are empty
|
||||
|
||||
### Requirement: Previous turn SHALL come from a safe durable Run result
|
||||
The contract SHALL define `diagnosis_run` as the durable source of intent, release outcome and safe published result, and SHALL exclude fallback, failed and cancelled Runs from Diagnosis previous-turn selection.
|
||||
|
||||
#### Scenario: Latest Run is a fallback
|
||||
- **WHEN** the latest same-session Run ended with `FALLBACK`
|
||||
- **THEN** it is not used as Diagnosis previous turn and selection continues to the latest eligible `DIAGNOSIS + SUCCESS` Run
|
||||
|
||||
### Requirement: Cancellation SHALL be layered and observable
|
||||
The contract SHALL distinguish cancellation request, prevention of new work, framework interruption, cancellable Tool work and HTTP timeout for already-blocking synchronous model calls.
|
||||
|
||||
#### Scenario: Client disconnects during a model call
|
||||
- **WHEN** an SSE client disconnects while a synchronous model call is in flight
|
||||
- **THEN** the system prevents later Draft release, requests interruption, records an internal cancelled terminal state and discards any late model result
|
||||
|
||||
### Requirement: Retry attempts SHALL be owned by Harness
|
||||
The contract SHALL require underlying SDK, HTTP and database retry layers to execute one attempt, while Harness explicitly owns any allowed Router or SemanticGuard retry.
|
||||
|
||||
#### Scenario: SemanticGuard returns an invalid schema
|
||||
- **WHEN** the first SemanticGuard attempt returns an invalid structured result
|
||||
- **THEN** Harness may execute one second attempt with the same verified snapshot and records both attempts
|
||||
|
||||
### Requirement: Phase gates SHALL remain serial
|
||||
The implementation plan SHALL contain eleven independent OpenSpec changes and SHALL prohibit starting a change before its predecessor is archived and committed.
|
||||
|
||||
#### Scenario: Stage 4 completes internal Agent tests
|
||||
- **WHEN** stage 4 passes its focused tests but stage 5 Guards are not implemented
|
||||
- **THEN** the public Chat entry remains on the old path and stage 6B cutover is prohibited
|
||||
@@ -0,0 +1,21 @@
|
||||
## 1. Contract Model
|
||||
|
||||
- [x] 1.1 Add typed enums for intent, release outcome, invocation status, evidence status, analysis kind, semantic verdict, fallback type and SSE outcome.
|
||||
- [x] 1.2 Add reusable records for Diagnosis Draft, Knowledge Answer Draft, safe fallback, published result and previous turn.
|
||||
- [x] 1.3 Add focused serialization and contract-shape tests covering positive evidence, no-evidence negative observation and forbidden raw/internal fields.
|
||||
|
||||
## 2. Security And Configuration Baseline
|
||||
|
||||
- [x] 2.1 Remove tracked plaintext credentials from `scripts/query_mysql.py` and `application.yml`, require environment-based secrets, and ignore local secret files.
|
||||
- [x] 2.2 Record the required external credential rotation and the Spring AI hidden-retry baseline without changing the public Chat runtime in this stage.
|
||||
|
||||
## 3. Design Freeze Alignment
|
||||
|
||||
- [x] 3.1 Keep ISS-014, OpenSpec artifacts and the devflow glossary aligned on the 11 serial changes, framework Tool Call ID, dual status fields and safe previous-turn source.
|
||||
- [x] 3.2 Add an architecture audit and cross-artifact alignment result to `decisions.md`.
|
||||
- [x] 3.3 Create the sm-flow `.committed` marker after proposal/design/specs/tasks and all decision gates pass.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Run focused contract tests and the smallest existing Chat/evidence baseline needed to prove stage zero did not switch runtime behavior.
|
||||
- [x] 4.2 Run strict OpenSpec validation and record commands, results and unverified external rotation in acceptance evidence.
|
||||
Reference in New Issue
Block a user