feat: add freshrss openclaw pipeline and clean repo

This commit is contained in:
zhuyongxin
2026-03-26 16:48:19 +08:00
parent 27fe1e8882
commit 100044e1f7
143 changed files with 1776 additions and 7293 deletions
+328
View File
@@ -0,0 +1,328 @@
# 规则引擎使用说明
## 1. 它解决什么问题
规则引擎负责把上游产出的结构化信号收敛成最终筛选决策:
`item -> content extraction -> llm summary -> rule engine -> candidate/openclaw`
这里有一个明确边界:
- LLM 负责理解正文、生成结构化摘要信号
- 规则引擎负责输出稳定、可复现、可审计的 `keep / drop / review`
也就是说,规则引擎不是“再让 LLM 判断一遍”,而是用确定性规则做最后裁决。
## 2. 相关文件
- 规则模型: `src/summary_mcp/models/filtering.py`
- 规则执行器: `src/summary_mcp/filters/engine.py`
- 默认规则: `configs/filter_rules.json`
- 本地脚本: `scripts/run_filter_rules.py`
- MCP tool: `filter_summary_result`
## 3. 输入与输出
规则引擎统一读取一个 `FilterInput`,包含四部分:
- `item`
- 标准化后的条目对象
- `article`
- 正文提取结果
- `summary`
- LLM 结构化摘要结果
- `context`
- 运行时注入的偏好信息
当前最常用的判断信号主要来自两类字段:
- `article.quality_flags.*`
- 如 `is_low_content`、`is_truncated`、`is_paywalled`
- `summary.*`
- 如 `category`、`worth_keeping`、`topics`
输出是 `FilterDecisionResult`:
- `decision`
- 最终决策,`keep / drop / review`
- `matched_rules`
- 命中的规则 ID 列表
- `reasons`
- 命中规则的原因说明
- `labels`
- 聚合后的标签
- `priority`
- 命中规则中的最高优先级
- `matches`
- 每条命中规则的明细
## 4. 规则文件怎么写
规则文件位置是 `configs/filter_rules.json`,顶层必须是一个 JSON 数组。
单条规则结构:
```json
{
"rule_id": "keep-worth-keeping-method",
"enabled": true,
"stop_on_match": false,
"conditions_all": [
{ "field": "summary.worth_keeping", "op": "eq", "value": true },
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
],
"action": {
"decision": "keep",
"reason": "Structured summary marked the content as worth keeping in a durable category.",
"labels": ["summary", "durable"],
"priority": 80
}
}
```
字段说明:
- `rule_id`
- 规则唯一标识,建议稳定命名
- `enabled`
- 是否启用
- `stop_on_match`
- 命中后是否立即停止继续匹配后续规则
- `conditions_all`
- 全部命中才算命中
- `conditions_any`
- 任意命中即可
- `action.decision`
- 该规则命中时产出的规则级决策
- `action.reason`
- 命中原因
- `action.labels`
- 打到结果里的标签
- `action.priority`
- 执行时排序优先级,越大越先执行
## 5. 当前支持的操作符
- `eq`
- `ne`
- `in`
- `not_in`
- `contains`
- `overlap`
- `gte`
- `lte`
- `exists`
常见例子:
```json
{ "field": "summary.worth_keeping", "op": "eq", "value": true }
```
```json
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
```
```json
{ "field": "summary.topics", "op": "overlap", "value": ["知识管理", "阅读工作流"] }
```
## 6. 执行顺序和收敛逻辑
执行顺序不是按文件书写顺序,而是按 `action.priority` 从高到低排序。
命中后会先收集所有规则,再做最终收敛:
- 只要命中过任意 `drop`,最终就是 `drop`
- 否则只要命中过任意 `keep`,最终就是 `keep`
- 否则只要命中过任意 `review`,最终就是 `review`
- 如果完全没有命中,默认 `review`
这意味着:
- `drop` 是硬拦截
- `keep` 只能在没有更高约束的 `drop` 时生效
- `review` 是默认灰区兜底
如果你希望某条高优规则一旦命中就不再继续匹配,把 `stop_on_match` 设为 `true`。
## 7. 动态上下文怎么用
规则支持从 `context` 动态取值,不需要把用户偏好写死进规则文件。
示例:
```json
{
"field": "summary.topics",
"op": "overlap",
"value": { "from_field": "context.interest_topics" }
}
```
对应的 `context` 可以是:
```json
{
"interest_topics": ["个人知识管理", "阅读工作流"]
}
```
这样同一套规则就可以被不同用户、不同运行场景复用。
## 8. 本地怎么跑
最小调用:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--output outputs/reference/filter/filter-decision.json
```
带上下文:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--context outputs/reference/filter/filter-context.json ^
--output outputs/reference/filter/filter-decision.with-context.json
```
也可以显式指定另一份规则文件:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--rules configs/filter_rules.json ^
--output outputs/reference/filter/filter-decision.json
```
## 9. MCP 怎么调用
MCP tool 名称是 `filter_summary_result`。
输入参数:
- `summary_result`
- `extracted_article`
- `item`
- `context`
其中只有 `summary_result` 是必填,其余都是可选补充信号。
示例:
```json
{
"summary_result": {
"title": "如何构建个人阅读工作流",
"url": "https://example.com/read-flow",
"summary": "文章介绍了从采集、提炼到沉淀的个人阅读工作流设计。",
"highlights": ["先采集再提炼", "用规则做稳定筛选", "日报再进入知识库"],
"keywords": ["阅读工作流", "知识管理", "RSS", "规则引擎", "日报"],
"topics": ["阅读工作流", "知识管理", "信息筛选"],
"category": "方法论",
"worth_keeping": true,
"reason": "提供了可复用的方法框架。"
},
"context": {
"interest_topics": ["阅读工作流", "知识管理"]
}
}
```
## 10. 结果怎么看
一个典型结果会像这样:
```json
{
"decision": "keep",
"matched_rules": [
"keep-worth-keeping-method",
"keep-interest-topic"
],
"reasons": [
"Structured summary marked the content as worth keeping in a durable category.",
"Topics overlap with current interest profile."
],
"labels": ["durable", "interest", "summary", "topic-match"],
"priority": 80,
"matches": [
{
"rule_id": "keep-worth-keeping-method",
"decision": "keep",
"reason": "Structured summary marked the content as worth keeping in a durable category.",
"labels": ["summary", "durable"],
"priority": 80
}
]
}
```
解读方式:
- 看 `decision`
- 最终裁决
- 看 `matched_rules`
- 哪些规则生效了
- 看 `reasons`
- 为什么做出这个判断
- 看 `matches`
- 需要排查时看完整命中明细
## 11. 在当前整条链路里的位置
当前生产链路里,规则引擎已经集成在 `run_freshrss_openclaw_pipeline` 中。
顺序是:
1. 从 FreshRSS 拉未读
2. 从 RSS 项目里读取正文
3. 调用 LLM 生成结构化摘要
4. 规则引擎输出 `keep / drop / review`
5. 构建 `ArticleCandidateRecord`
6. 压缩成 `OpenClawCandidateInput`
7. 生成 `openclaw-delivery-payload.json`
8. 如果开启 `mark_read`,最后再标记已读
所以在生产模式下,一般不需要单独跑 `run_filter_rules.py`,只有在调规则或排查命中逻辑时才单独跑。
## 12. 调规则时的建议
- 把“硬性淘汰”规则放高优先级
- 比如低质量正文、明显噪音内容
- 把“强 keep”规则放在中高优先级
- 比如 `worth_keeping=true` 且类别是方法论
- 把“兜底 review”规则放低一些
- 避免过早收敛
- 用户偏好尽量走 `context`
- 不要把临时兴趣直接硬编码到规则里
- `rule_id` 保持稳定
- 方便后续审计、统计和排障
## 13. 一个最常见的改法
如果你想增加一条“命中关注主题就 keep”的规则,可以直接在 `configs/filter_rules.json` 里加:
```json
{
"rule_id": "keep-interest-topic",
"enabled": true,
"stop_on_match": false,
"conditions_all": [
{ "field": "summary.topics", "op": "overlap", "value": { "from_field": "context.interest_topics" } }
],
"action": {
"decision": "keep",
"reason": "Topics overlap with current interest profile.",
"labels": ["interest", "topic-match"],
"priority": 75
}
}
```
如果只是临时停用某条规则,最简单的是把它的 `enabled` 改成 `false`,不要先删规则。