Files
SuperBizAgent-java/mvp/architecture/harness-quality-gates.md
T

65 lines
3.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Harness 与质量门禁
**更新日期**:2026-07-23
**状态**:当前可运行架构
## 1. Harness 定位
Harness 是确定性执行边界,不承担业务推理。它统一管理:
- RunContext、deadline、first-terminal-wins lifecycle 与客户端取消。
- 模型调用、Tool 调用、Token、字节数和单 Tool 次数预算。
- 类型化 retry policy;Diagnosis Agent 和 Tool 调用不自动重试。
- ToolBoundary、canonical invocation 与 Agent projection。
- EvidenceGuard、Evidence repair、SemanticGuard 与 Release Policy。
- metadata-only durable audit、统一 Trace Timeline 与独立 reasoning 审计。
## 2. ToolBoundary
```text
framework tool_call_id
-> exact Run / schema / authorization / read-only / budget
-> backend execution
-> raw response -> Redis canonical invocation
-> projector -> bounded agent_result
-> ToolInvocation durable metadata audit
-> Agent observation
```
Redis canonical invocation 可在 TTL 内保存完整 request/raw_response/agent_result,受独立前缀、容量和 Harness-only 访问保护。Durable audit 只保存 identity、Tool 名、状态、耗时和字节数;audit 写入失败可观测但不改变 canonical Tool 结果。
## 3. EvidenceGuard
EvidenceGuard 不调用模型。它校验 Draft schema、analysis ID、当前 Run Tool ownership、READY 状态、evidence status 和每条结论的引用闭包,并生成只包含 Agent projection 的 verified snapshot。
## 4. SemanticGuard
SemanticGuard 使用隔离的单轮模型调用,只接收原始 query、完整 Draft 和 verified snapshot。它无 Tool、无记忆、不访问 Redis、不改写报告;技术失败最多按相同输入重试一次,仍失败则安全降级。
## 5. Release Policy
- `SUPPORTED`:发布 Diagnosis Agent 原始安全 Draft 的 typed report。
- `UNSUPPORTED` 或 evidence failure:发布有界 SAFE_FALLBACK,说明 `failure_stage`、已验证 `observed_facts`、`validation_issues`、限制和 `next_steps`;不得泄露 Prompt、原始 Draft、原始 Tool 载荷或内部异常。
- technical failure:发布 stable failure,不泄漏内部异常。
- cancel/timeout:结束 exact Run,禁止 late content。
## 6. Trace Recorder
`diagnosis_trace_event` 是追加式统一 Timeline。Recorder 按 exact `runId` 分配递增 `sequence_no`,覆盖:
- `RUN`:Run 开始和终态。
- `ROUTING`:路由尝试与最终 intent。
- `AGENT`:模型步骤及有界 metadata。
- `TOOL`:Tool 调用状态和耗时。
- `EVIDENCE`:首次校验、Repair 尝试和复检。
- `SEMANTIC`:语义校验尝试与判定。
- `RELEASE`:最终发布决策。
Trace 写入失败只记录警告,不应改变业务执行结果;`details` 禁止包含 Prompt、Thought、Draft 正文或 raw Tool payload。普通 Trace API 按 `sequence_no, id` 返回 Timeline。
## 7. Audit 安全
AgentStep 不保存 Prompt、消息正文、模型正文、Tool arguments 或 Thought。ToolInvocation 不保存完整 request、SQL/日志 query、raw response 或 Agent projection。Provider reasoning 仅写入独立 `agent_reasoning_audit`,不进入普通 Trace 或发布结果;无 Provider 内容时必须记录 unavailable,不能伪造。应用日志不得打印这些字段。
Reasoning endpoint 当前已与普通 Trace 分离并执行 `sessionId + runId` 归属校验,但访问控制、保留期限和加密要求尚未完成,继续由 ISS-015 跟踪。