Files
SuperBizAgent-java/mvp/architecture/archive/2026-07-05-legacy/confidence-feedback.md

158 lines
4.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 证据评分与用户反馈架构
## 一、整体架构
```
用户对话
↓
ChatService.executeChat / executeChatComplex
↓ SUCCESS 后写入 answer,异步触发
EvaluationService.evaluate(sessionId, answer)
└─ 读取 tool_invocation 事实 → 规则引擎 → 写 selfEvaluation
用户提交反馈
↓
POST /api/feedback { sessionId, feedback: "useful" | "not_useful" }
↓
FeedbackService.submitFeedback
├─ 写 DiagnosisSession.feedback
├─ useful → CaseLibraryService.createFromSession → 写 case_library
└─ not_useful → 仅写 feedback,status 不变
```
---
## 二、评分规则(evidence_score)
### 定位
`evidence_score` 衡量的是**证据收集充分度**,不是答案准确性。
- 能证明的:Agent 是否有尝试收集证据、检索是否命中
- 不能证明的:答案是否有幻觉、推理是否正确
### 数据来源
规则引擎只消费 `tool_invocation` 表的事实记录,不依赖 LLM 判断。
### 规则定义
| 规则名 | 条件 | delta |
|---|---|---|
| `no_tool_call` | 无任何工具调用 | 直接 0 分,不参与加权 |
| `execution_failed` | status = FAILED | 直接 0 分,不参与加权 |
| `has_successful_tool_call` | 至少 1 次成功调用 | +30 |
| `l0_exact_match` | 任意调用有 L0 精确匹配命中 | +35 |
| `l1_semantic_match` | 无 L0 命中但有 L1 语义匹配 | +20 |
| `retrieval_no_hit` | 有检索调用但无任何命中 | -10 |
| `all_tool_calls_failed` | 全部调用失败 | -20 |
> L0 和 L1 互斥取高优先级(L0 命中时跳过 L1 分支)。
### selfEvaluation 字段格式
```json
{
"evidence_score": 65,
"source": "rule",
"factors": [
{"name": "has_successful_tool_call", "delta": 30, "description": "有成功的工具调用(20次)"},
{"name": "l0_exact_match", "delta": 35, "description": "L0 精确匹配命中"}
]
}
```
| 字段 | 说明 |
|---|---|
| `evidence_score` | 0-100 整数 |
| `source` | 当前固定为 `"rule"`;预留 `"llm"` 供后续扩展 |
| `factors` | 命中的规则列表,含 name / delta / description |
| `llm_opinion` | 预留字段(未实现),LLM 观点叠加时在此扩展 |
### 已知边界
- 非检索工具(DateTimeTools、QueryMetricsTools 等)不写 `tool_invocation`,这类 session 的 evidence_score = 0,属于设计边界
- 评分为异步写入(`@Async`),失败时 `selfEvaluation` 保持 null,前端需处理 null
---
## 三、反馈机制
### API
```
POST /api/feedback
Content-Type: application/json
{
"sessionId": "xxx",
"feedback": "useful" | "not_useful"
}
```
**响应**
```json
{
"success": true,
"message": "反馈已记录",
"caseId": "uuid 或 null"
}
```
### 后端行为
| feedback 值 | 操作 |
|---|---|
| `useful` | 写 `DiagnosisSession.feedback = "useful"`,生成 `CaseLibrary` 记录,返回 caseId |
| `not_useful` | 写 `DiagnosisSession.feedback = "not_useful"`,status 不变 |
| 其他值 | 返回 HTTP 400 |
### 重要设计决策
**BAD_CASE 不改 status 字段**
`status` 表示执行状态(RUNNING/SUCCESS/FAILED),是独立维度,不能被质量标签覆盖。
查询 BadCase 使用:`WHERE feedback = 'not_useful'`
**useful 触发案例沉淀规则**
| CaseLibrary 字段 | 来源 |
|---|---|
| caseId | UUID |
| diagnosisId | DiagnosisSession.sessionId |
| sourceType | AUTO |
| faultCategory | GENERAL(暂时,后续人工补充) |
| title | query 前 100 字符 |
| rootCause / solution | DiagnosisSession.answer(完整答案) |
| createdBy | "system" |
**幂等性**:同一 sessionId 重复提交 useful,返回已有 caseId,不重复插入 case_library。
---
## 四、数据库变更
### V008(新增)
```sql
ALTER TABLE diagnosis_session ADD COLUMN answer LONGTEXT COMMENT 'Agent 返回给用户的完整答案';
```
### diagnosis_session 关键字段
| 字段 | 类型 | 说明 |
|---|---|---|
| `answer` | LONGTEXT | Agent 完整回答,useful 案例沉淀的内容来源 |
| `self_evaluation` | JSON | 证据评分结果,格式见上 |
| `feedback` | VARCHAR(16) | useful / not_useful / null |
| `status` | VARCHAR(16) | 执行状态,不受 feedback 影响 |
---
## 五、扩展方向(Phase 2)
- **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑
- **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../../../issues/archived/ISS-001-duplicate-retrieval.md))