feat: add freshrss openclaw pipeline and clean repo
This commit is contained in:
@@ -0,0 +1,328 @@
|
||||
# 规则引擎使用说明
|
||||
|
||||
## 1. 它解决什么问题
|
||||
|
||||
规则引擎负责把上游产出的结构化信号收敛成最终筛选决策:
|
||||
|
||||
`item -> content extraction -> llm summary -> rule engine -> candidate/openclaw`
|
||||
|
||||
这里有一个明确边界:
|
||||
|
||||
- LLM 负责理解正文、生成结构化摘要信号
|
||||
- 规则引擎负责输出稳定、可复现、可审计的 `keep / drop / review`
|
||||
|
||||
也就是说,规则引擎不是“再让 LLM 判断一遍”,而是用确定性规则做最后裁决。
|
||||
|
||||
## 2. 相关文件
|
||||
|
||||
- 规则模型: `src/summary_mcp/models/filtering.py`
|
||||
- 规则执行器: `src/summary_mcp/filters/engine.py`
|
||||
- 默认规则: `configs/filter_rules.json`
|
||||
- 本地脚本: `scripts/run_filter_rules.py`
|
||||
- MCP tool: `filter_summary_result`
|
||||
|
||||
## 3. 输入与输出
|
||||
|
||||
规则引擎统一读取一个 `FilterInput`,包含四部分:
|
||||
|
||||
- `item`
|
||||
- 标准化后的条目对象
|
||||
- `article`
|
||||
- 正文提取结果
|
||||
- `summary`
|
||||
- LLM 结构化摘要结果
|
||||
- `context`
|
||||
- 运行时注入的偏好信息
|
||||
|
||||
当前最常用的判断信号主要来自两类字段:
|
||||
|
||||
- `article.quality_flags.*`
|
||||
- 如 `is_low_content`、`is_truncated`、`is_paywalled`
|
||||
- `summary.*`
|
||||
- 如 `category`、`worth_keeping`、`topics`
|
||||
|
||||
输出是 `FilterDecisionResult`:
|
||||
|
||||
- `decision`
|
||||
- 最终决策,`keep / drop / review`
|
||||
- `matched_rules`
|
||||
- 命中的规则 ID 列表
|
||||
- `reasons`
|
||||
- 命中规则的原因说明
|
||||
- `labels`
|
||||
- 聚合后的标签
|
||||
- `priority`
|
||||
- 命中规则中的最高优先级
|
||||
- `matches`
|
||||
- 每条命中规则的明细
|
||||
|
||||
## 4. 规则文件怎么写
|
||||
|
||||
规则文件位置是 `configs/filter_rules.json`,顶层必须是一个 JSON 数组。
|
||||
|
||||
单条规则结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"rule_id": "keep-worth-keeping-method",
|
||||
"enabled": true,
|
||||
"stop_on_match": false,
|
||||
"conditions_all": [
|
||||
{ "field": "summary.worth_keeping", "op": "eq", "value": true },
|
||||
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
|
||||
],
|
||||
"action": {
|
||||
"decision": "keep",
|
||||
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||
"labels": ["summary", "durable"],
|
||||
"priority": 80
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
字段说明:
|
||||
|
||||
- `rule_id`
|
||||
- 规则唯一标识,建议稳定命名
|
||||
- `enabled`
|
||||
- 是否启用
|
||||
- `stop_on_match`
|
||||
- 命中后是否立即停止继续匹配后续规则
|
||||
- `conditions_all`
|
||||
- 全部命中才算命中
|
||||
- `conditions_any`
|
||||
- 任意命中即可
|
||||
- `action.decision`
|
||||
- 该规则命中时产出的规则级决策
|
||||
- `action.reason`
|
||||
- 命中原因
|
||||
- `action.labels`
|
||||
- 打到结果里的标签
|
||||
- `action.priority`
|
||||
- 执行时排序优先级,越大越先执行
|
||||
|
||||
## 5. 当前支持的操作符
|
||||
|
||||
- `eq`
|
||||
- `ne`
|
||||
- `in`
|
||||
- `not_in`
|
||||
- `contains`
|
||||
- `overlap`
|
||||
- `gte`
|
||||
- `lte`
|
||||
- `exists`
|
||||
|
||||
常见例子:
|
||||
|
||||
```json
|
||||
{ "field": "summary.worth_keeping", "op": "eq", "value": true }
|
||||
```
|
||||
|
||||
```json
|
||||
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
|
||||
```
|
||||
|
||||
```json
|
||||
{ "field": "summary.topics", "op": "overlap", "value": ["知识管理", "阅读工作流"] }
|
||||
```
|
||||
|
||||
## 6. 执行顺序和收敛逻辑
|
||||
|
||||
执行顺序不是按文件书写顺序,而是按 `action.priority` 从高到低排序。
|
||||
|
||||
命中后会先收集所有规则,再做最终收敛:
|
||||
|
||||
- 只要命中过任意 `drop`,最终就是 `drop`
|
||||
- 否则只要命中过任意 `keep`,最终就是 `keep`
|
||||
- 否则只要命中过任意 `review`,最终就是 `review`
|
||||
- 如果完全没有命中,默认 `review`
|
||||
|
||||
这意味着:
|
||||
|
||||
- `drop` 是硬拦截
|
||||
- `keep` 只能在没有更高约束的 `drop` 时生效
|
||||
- `review` 是默认灰区兜底
|
||||
|
||||
如果你希望某条高优规则一旦命中就不再继续匹配,把 `stop_on_match` 设为 `true`。
|
||||
|
||||
## 7. 动态上下文怎么用
|
||||
|
||||
规则支持从 `context` 动态取值,不需要把用户偏好写死进规则文件。
|
||||
|
||||
示例:
|
||||
|
||||
```json
|
||||
{
|
||||
"field": "summary.topics",
|
||||
"op": "overlap",
|
||||
"value": { "from_field": "context.interest_topics" }
|
||||
}
|
||||
```
|
||||
|
||||
对应的 `context` 可以是:
|
||||
|
||||
```json
|
||||
{
|
||||
"interest_topics": ["个人知识管理", "阅读工作流"]
|
||||
}
|
||||
```
|
||||
|
||||
这样同一套规则就可以被不同用户、不同运行场景复用。
|
||||
|
||||
## 8. 本地怎么跑
|
||||
|
||||
最小调用:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--output outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
带上下文:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--context outputs/reference/filter/filter-context.json ^
|
||||
--output outputs/reference/filter/filter-decision.with-context.json
|
||||
```
|
||||
|
||||
也可以显式指定另一份规则文件:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--rules configs/filter_rules.json ^
|
||||
--output outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
## 9. MCP 怎么调用
|
||||
|
||||
MCP tool 名称是 `filter_summary_result`。
|
||||
|
||||
输入参数:
|
||||
|
||||
- `summary_result`
|
||||
- `extracted_article`
|
||||
- `item`
|
||||
- `context`
|
||||
|
||||
其中只有 `summary_result` 是必填,其余都是可选补充信号。
|
||||
|
||||
示例:
|
||||
|
||||
```json
|
||||
{
|
||||
"summary_result": {
|
||||
"title": "如何构建个人阅读工作流",
|
||||
"url": "https://example.com/read-flow",
|
||||
"summary": "文章介绍了从采集、提炼到沉淀的个人阅读工作流设计。",
|
||||
"highlights": ["先采集再提炼", "用规则做稳定筛选", "日报再进入知识库"],
|
||||
"keywords": ["阅读工作流", "知识管理", "RSS", "规则引擎", "日报"],
|
||||
"topics": ["阅读工作流", "知识管理", "信息筛选"],
|
||||
"category": "方法论",
|
||||
"worth_keeping": true,
|
||||
"reason": "提供了可复用的方法框架。"
|
||||
},
|
||||
"context": {
|
||||
"interest_topics": ["阅读工作流", "知识管理"]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 10. 结果怎么看
|
||||
|
||||
一个典型结果会像这样:
|
||||
|
||||
```json
|
||||
{
|
||||
"decision": "keep",
|
||||
"matched_rules": [
|
||||
"keep-worth-keeping-method",
|
||||
"keep-interest-topic"
|
||||
],
|
||||
"reasons": [
|
||||
"Structured summary marked the content as worth keeping in a durable category.",
|
||||
"Topics overlap with current interest profile."
|
||||
],
|
||||
"labels": ["durable", "interest", "summary", "topic-match"],
|
||||
"priority": 80,
|
||||
"matches": [
|
||||
{
|
||||
"rule_id": "keep-worth-keeping-method",
|
||||
"decision": "keep",
|
||||
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||
"labels": ["summary", "durable"],
|
||||
"priority": 80
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
解读方式:
|
||||
|
||||
- 看 `decision`
|
||||
- 最终裁决
|
||||
- 看 `matched_rules`
|
||||
- 哪些规则生效了
|
||||
- 看 `reasons`
|
||||
- 为什么做出这个判断
|
||||
- 看 `matches`
|
||||
- 需要排查时看完整命中明细
|
||||
|
||||
## 11. 在当前整条链路里的位置
|
||||
|
||||
当前生产链路里,规则引擎已经集成在 `run_freshrss_openclaw_pipeline` 中。
|
||||
|
||||
顺序是:
|
||||
|
||||
1. 从 FreshRSS 拉未读
|
||||
2. 从 RSS 项目里读取正文
|
||||
3. 调用 LLM 生成结构化摘要
|
||||
4. 规则引擎输出 `keep / drop / review`
|
||||
5. 构建 `ArticleCandidateRecord`
|
||||
6. 压缩成 `OpenClawCandidateInput`
|
||||
7. 生成 `openclaw-delivery-payload.json`
|
||||
8. 如果开启 `mark_read`,最后再标记已读
|
||||
|
||||
所以在生产模式下,一般不需要单独跑 `run_filter_rules.py`,只有在调规则或排查命中逻辑时才单独跑。
|
||||
|
||||
## 12. 调规则时的建议
|
||||
|
||||
- 把“硬性淘汰”规则放高优先级
|
||||
- 比如低质量正文、明显噪音内容
|
||||
- 把“强 keep”规则放在中高优先级
|
||||
- 比如 `worth_keeping=true` 且类别是方法论
|
||||
- 把“兜底 review”规则放低一些
|
||||
- 避免过早收敛
|
||||
- 用户偏好尽量走 `context`
|
||||
- 不要把临时兴趣直接硬编码到规则里
|
||||
- `rule_id` 保持稳定
|
||||
- 方便后续审计、统计和排障
|
||||
|
||||
## 13. 一个最常见的改法
|
||||
|
||||
如果你想增加一条“命中关注主题就 keep”的规则,可以直接在 `configs/filter_rules.json` 里加:
|
||||
|
||||
```json
|
||||
{
|
||||
"rule_id": "keep-interest-topic",
|
||||
"enabled": true,
|
||||
"stop_on_match": false,
|
||||
"conditions_all": [
|
||||
{ "field": "summary.topics", "op": "overlap", "value": { "from_field": "context.interest_topics" } }
|
||||
],
|
||||
"action": {
|
||||
"decision": "keep",
|
||||
"reason": "Topics overlap with current interest profile.",
|
||||
"labels": ["interest", "topic-match"],
|
||||
"priority": 75
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
如果只是临时停用某条规则,最简单的是把它的 `enabled` 改成 `false`,不要先删规则。
|
||||
Reference in New Issue
Block a user