Files
reader/docs/design/filter-rule-engine-usage.md

8.3 KiB
Raw Permalink Blame History

规则引擎使用说明

1. 它解决什么问题

规则引擎负责把上游产出的结构化信号收敛成最终筛选决策:

item -> content extraction -> llm summary -> rule engine -> candidate/openclaw

这里有一个明确边界:

  • LLM 负责理解正文、生成结构化摘要信号
  • 规则引擎负责输出稳定、可复现、可审计的 keep / drop / review

也就是说,规则引擎不是“再让 LLM 判断一遍”,而是用确定性规则做最后裁决。

2. 相关文件

  • 规则模型: src/summary_mcp/models/filtering.py
  • 规则执行器: src/summary_mcp/filters/engine.py
  • 默认规则: configs/filter_rules.json
  • 本地脚本: scripts/run_filter_rules.py
  • MCP tool: filter_summary_result

3. 输入与输出

规则引擎统一读取一个 FilterInput,包含四部分:

  • item
    • 标准化后的条目对象
  • article
    • 正文提取结果
  • summary
    • LLM 结构化摘要结果
  • context
    • 运行时注入的偏好信息

当前最常用的判断信号主要来自两类字段:

  • article.quality_flags.*
    • 如 is_low_content、is_truncated、is_paywalled
  • summary.*
    • 如 category、worth_keeping、topics

输出是 FilterDecisionResult:

  • decision
    • 最终决策,keep / drop / review
  • matched_rules
    • 命中的规则 ID 列表
  • reasons
    • 命中规则的原因说明
  • labels
    • 聚合后的标签
  • priority
    • 命中规则中的最高优先级
  • matches
    • 每条命中规则的明细

4. 规则文件怎么写

规则文件位置是 configs/filter_rules.json,顶层必须是一个 JSON 数组。

单条规则结构:

{
  "rule_id": "keep-worth-keeping-method",
  "enabled": true,
  "stop_on_match": false,
  "conditions_all": [
    { "field": "summary.worth_keeping", "op": "eq", "value": true },
    { "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
  ],
  "action": {
    "decision": "keep",
    "reason": "Structured summary marked the content as worth keeping in a durable category.",
    "labels": ["summary", "durable"],
    "priority": 80
  }
}

字段说明:

  • rule_id
    • 规则唯一标识,建议稳定命名
  • enabled
    • 是否启用
  • stop_on_match
    • 命中后是否立即停止继续匹配后续规则
  • conditions_all
    • 全部命中才算命中
  • conditions_any
    • 任意命中即可
  • action.decision
    • 该规则命中时产出的规则级决策
  • action.reason
    • 命中原因
  • action.labels
    • 打到结果里的标签
  • action.priority
    • 执行时排序优先级,越大越先执行

5. 当前支持的操作符

  • eq
  • ne
  • in
  • not_in
  • contains
  • overlap
  • gte
  • lte
  • exists

常见例子:

{ "field": "summary.worth_keeping", "op": "eq", "value": true }
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
{ "field": "summary.topics", "op": "overlap", "value": ["知识管理", "阅读工作流"] }

6. 执行顺序和收敛逻辑

执行顺序不是按文件书写顺序,而是按 action.priority 从高到低排序。

命中后会先收集所有规则,再做最终收敛:

  • 只要命中过任意 drop,最终就是 drop
  • 否则只要命中过任意 keep,最终就是 keep
  • 否则只要命中过任意 review,最终就是 review
  • 如果完全没有命中,默认 review

这意味着:

  • drop 是硬拦截
  • keep 只能在没有更高约束的 drop 时生效
  • review 是默认灰区兜底

如果你希望某条高优规则一旦命中就不再继续匹配,把 stop_on_match 设为 true。

7. 动态上下文怎么用

规则支持从 context 动态取值,不需要把用户偏好写死进规则文件。

示例:

{
  "field": "summary.topics",
  "op": "overlap",
  "value": { "from_field": "context.interest_topics" }
}

对应的 context 可以是:

{
  "interest_topics": ["个人知识管理", "阅读工作流"]
}

这样同一套规则就可以被不同用户、不同运行场景复用。

8. 本地怎么跑

最小调用:

python scripts/run_filter_rules.py ^
  --summary outputs/reference/summary/result.loop.json ^
  --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
  --output outputs/reference/filter/filter-decision.json

带上下文:

python scripts/run_filter_rules.py ^
  --summary outputs/reference/summary/result.loop.json ^
  --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
  --context outputs/reference/filter/filter-context.json ^
  --output outputs/reference/filter/filter-decision.with-context.json

也可以显式指定另一份规则文件:

python scripts/run_filter_rules.py ^
  --summary outputs/reference/summary/result.loop.json ^
  --rules configs/filter_rules.json ^
  --output outputs/reference/filter/filter-decision.json

9. MCP 怎么调用

MCP tool 名称是 filter_summary_result。

输入参数:

  • summary_result
  • extracted_article
  • item
  • context

其中只有 summary_result 是必填,其余都是可选补充信号。

示例:

{
  "summary_result": {
    "title": "如何构建个人阅读工作流",
    "url": "https://example.com/read-flow",
    "summary": "文章介绍了从采集、提炼到沉淀的个人阅读工作流设计。",
    "highlights": ["先采集再提炼", "用规则做稳定筛选", "日报再进入知识库"],
    "keywords": ["阅读工作流", "知识管理", "RSS", "规则引擎", "日报"],
    "topics": ["阅读工作流", "知识管理", "信息筛选"],
    "category": "方法论",
    "worth_keeping": true,
    "reason": "提供了可复用的方法框架。"
  },
  "context": {
    "interest_topics": ["阅读工作流", "知识管理"]
  }
}

10. 结果怎么看

一个典型结果会像这样:

{
  "decision": "keep",
  "matched_rules": [
    "keep-worth-keeping-method",
    "keep-interest-topic"
  ],
  "reasons": [
    "Structured summary marked the content as worth keeping in a durable category.",
    "Topics overlap with current interest profile."
  ],
  "labels": ["durable", "interest", "summary", "topic-match"],
  "priority": 80,
  "matches": [
    {
      "rule_id": "keep-worth-keeping-method",
      "decision": "keep",
      "reason": "Structured summary marked the content as worth keeping in a durable category.",
      "labels": ["summary", "durable"],
      "priority": 80
    }
  ]
}

解读方式:

  • 看 decision
    • 最终裁决
  • 看 matched_rules
    • 哪些规则生效了
  • 看 reasons
    • 为什么做出这个判断
  • 看 matches
    • 需要排查时看完整命中明细

11. 在当前整条链路里的位置

当前生产链路里,规则引擎已经集成在 run_freshrss_openclaw_pipeline 中。

顺序是:

  1. 从 FreshRSS 拉未读
  2. 从 RSS 项目里读取正文
  3. 调用 LLM 生成结构化摘要
  4. 规则引擎输出 keep / drop / review
  5. 构建 ArticleCandidateRecord
  6. 压缩成 OpenClawCandidateInput
  7. 生成 openclaw-delivery-payload.json
  8. 如果开启 mark_read,最后再标记已读

所以在生产模式下,一般不需要单独跑 run_filter_rules.py,只有在调规则或排查命中逻辑时才单独跑。

12. 调规则时的建议

  • 把“硬性淘汰”规则放高优先级
    • 比如低质量正文、明显噪音内容
  • 把“强 keep”规则放在中高优先级
    • 比如 worth_keeping=true 且类别是方法论
  • 把“兜底 review”规则放低一些
    • 避免过早收敛
  • 用户偏好尽量走 context
    • 不要把临时兴趣直接硬编码到规则里
  • rule_id 保持稳定
    • 方便后续审计、统计和排障

13. 一个最常见的改法

如果你想增加一条“命中关注主题就 keep”的规则,可以直接在 configs/filter_rules.json 里加:

{
  "rule_id": "keep-interest-topic",
  "enabled": true,
  "stop_on_match": false,
  "conditions_all": [
    { "field": "summary.topics", "op": "overlap", "value": { "from_field": "context.interest_topics" } }
  ],
  "action": {
    "decision": "keep",
    "reason": "Topics overlap with current interest profile.",
    "labels": ["interest", "topic-match"],
    "priority": 75
  }
}

如果只是临时停用某条规则,最简单的是把它的 enabled 改成 false,不要先删规则。