feat: add freshrss openclaw pipeline and clean repo
This commit is contained in:
+20
-12
@@ -1,4 +1,4 @@
|
||||
# 文档索引
|
||||
# 文档索引
|
||||
|
||||
## 当前目录结构
|
||||
|
||||
@@ -19,25 +19,29 @@
|
||||
|
||||
1. `docs/current/context-reset-brief.md`
|
||||
- 当前真实进度与下一步入口
|
||||
2. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
- 给 OpenClaw 的接手说明、环境变量、MCP 调用方式与已知限制
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
||||
3. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
- 提供给 OpenClaw 的批量投递 envelope 说明
|
||||
4. `docs/design/summary-mcp-service-design.md`
|
||||
5. `docs/design/summary-mcp-service-design.md`
|
||||
- 当前 MCP 服务的职责、接口和边界
|
||||
5. `docs/design/filter-rule-engine-design.md`
|
||||
6. `docs/design/filter-rule-engine-design.md`
|
||||
- 过滤层的输入输出、规则结构与当前实现
|
||||
6. `docs/design/markdown-sink-design.md`
|
||||
7. `docs/design/filter-rule-engine-usage.md`
|
||||
- 规则怎么写、怎么跑、结果怎么解读的使用说明
|
||||
8. `docs/design/markdown-sink-design.md`
|
||||
- 第一版 Markdown sink 的输入输出、目录结构与落地方式
|
||||
7. `docs/openclaw/openclaw-daily-digest-refactor.md`
|
||||
9. `docs/openclaw/openclaw-daily-digest-refactor.md`
|
||||
- 为什么要从单篇入库改成 OpenClaw 日报聚合链路
|
||||
8. `docs/openclaw/article-candidate-daily-digest-schema.md`
|
||||
10. `docs/openclaw/article-candidate-daily-digest-schema.md`
|
||||
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的正式设计
|
||||
9. `docs/design/source-schema-design.md`
|
||||
11. `docs/design/source-schema-design.md`
|
||||
- `source -> item -> document` 的对象设计
|
||||
10. `docs/notes/reading-pipeline-design-notes.md`
|
||||
12. `docs/notes/reading-pipeline-design-notes.md`
|
||||
- 更上层的阅读流方案与阶段划分
|
||||
11. `docs/design/summary-loop-explained.md`
|
||||
13. `docs/design/summary-loop-explained.md`
|
||||
- 当前 LLM 摘要校验闭环的解释
|
||||
|
||||
## 当前文档分层
|
||||
@@ -65,11 +69,15 @@
|
||||
- 提取 JSON -> LLM 摘要 JSON -> 校验 的闭环说明
|
||||
- `docs/design/filter-rule-engine-design.md`
|
||||
- 第一版规则过滤引擎设计与落地位置
|
||||
- `docs/design/filter-rule-engine-usage.md`
|
||||
- 规则配置、调用方式与结果解读
|
||||
- `docs/design/markdown-sink-design.md`
|
||||
- 第一版 Markdown sink 设计与落地位置
|
||||
|
||||
### 3. OpenClaw 与下游设计
|
||||
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
- OpenClaw 接手所需的运行说明、工具入口与已知限制
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
@@ -95,4 +103,4 @@
|
||||
- `TODO.md` 记录任务优先级与下一步
|
||||
- `outputs/README.md` 记录当前输出目录约定
|
||||
- `docs/archive/content-extract-mcp-mvp-archive.md` 只当历史快照,不再作为最新事实来源
|
||||
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
|
||||
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
|
||||
|
||||
@@ -1,102 +1,112 @@
|
||||
# 项目当前状态简报
|
||||
|
||||
## 当前结论
|
||||
|
||||
当前仓库已经具备交付给 OpenClaw 的基础条件。
|
||||
|
||||
当前主链路是:
|
||||
|
||||
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
|
||||
|
||||
OpenClaw 应通过 MCP 工具 `run_freshrss_openclaw_pipeline` 调用这条链路,而不是自行拼接脚本。
|
||||
|
||||
## 当前已完成
|
||||
|
||||
- 已明确整体链路:`来源 -> 聚合池 -> 内容提取 MCP -> LLM 摘要 -> 校验 -> 过滤 -> 入库/推送`
|
||||
- 已确定当前 MCP 的职责边界:
|
||||
- 只负责内容提取
|
||||
- 不负责摘要、分类、价值判断
|
||||
- 已完成 Python MCP 骨架
|
||||
- 已实现三个 tool:
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
- `filter_summary_result`
|
||||
- 已完成真实 URL 提取验证
|
||||
- 已完成 LLM 摘要 prompt
|
||||
- 已完成 LLM 摘要结果 schema 校验器
|
||||
- 已完成“提取 JSON -> LLM 摘要 JSON -> 校验”的最小闭环脚本
|
||||
- 已完成 LLM 摘要校验 skill 封装
|
||||
- 已清理旧的启发式 `summarizer.py`
|
||||
- 已将旧的 MCP 设计文档更新为当前“Content Extract MCP”语义
|
||||
- 已完成 FreshRSS `greader` API 接入
|
||||
- 已完成 FreshRSS entry -> `item` 映射
|
||||
- 已产出真实 `item` 样例文件
|
||||
- 已跑通 `FreshRSS -> item -> content extraction` 单条链路
|
||||
- 已定义过滤层输入输出 schema
|
||||
- 已完成第一版规则引擎、本地脚本和 MCP tool
|
||||
- 已定义 sink 输入输出 schema
|
||||
- 已完成第一版 Markdown sink 与本地写入脚本
|
||||
- 已完成 FreshRSS `greader` API 接入与未读拉取
|
||||
- 已完成 FreshRSS 条目到标准化 `item` 的映射
|
||||
- 已完成 RSS-first 提取策略
|
||||
- 已完成 LLM 总结与校验闭环
|
||||
- 已完成规则引擎过滤
|
||||
- 已完成 `ArticleCandidateRecord` 与 `OpenClawCandidateInput` 分层
|
||||
- 已完成 `OpenClawDeliveryPayload` 批量投递结构
|
||||
- 已完成 FreshRSS 已读状态回写
|
||||
- 已完成“仅在最终 payload 成功写盘后再标记已读”的语义
|
||||
- 已完成 MCP 工具 `run_freshrss_openclaw_pipeline`
|
||||
- 已完成默认精简输出模式,减少中间文件
|
||||
|
||||
## 当前 MCP 工具
|
||||
|
||||
当前服务入口:
|
||||
|
||||
- `src/summary_mcp/server.py`
|
||||
|
||||
当前暴露的 MCP 工具:
|
||||
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
- `filter_summary_result`
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
|
||||
其中生产主入口是:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
|
||||
## 当前关键文件
|
||||
|
||||
- MCP 入口:
|
||||
- MCP 服务入口
|
||||
- `src/summary_mcp/server.py`
|
||||
- 提取主流程:
|
||||
- FreshRSS 统一工作流
|
||||
- `src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
- 摘要循环
|
||||
- `src/summary_mcp/core/summary_loop.py`
|
||||
- 提取主流程
|
||||
- `src/summary_mcp/core/pipeline.py`
|
||||
- LLM 结果模型:
|
||||
- `src/summary_mcp/models/llm_result.py`
|
||||
- LLM 校验器:
|
||||
- `src/summary_mcp/validators/llm_result.py`
|
||||
- 过滤模型:
|
||||
- `src/summary_mcp/models/filtering.py`
|
||||
- 过滤引擎:
|
||||
- FreshRSS 集成
|
||||
- `src/summary_mcp/integrations/freshrss.py`
|
||||
- 规则引擎
|
||||
- `src/summary_mcp/filters/engine.py`
|
||||
- sink 模型:
|
||||
- `src/summary_mcp/models/sink.py`
|
||||
- Markdown sink:
|
||||
- `src/summary_mcp/sinks/markdown.py`
|
||||
- 默认过滤规则:
|
||||
- `configs/filter_rules.json`
|
||||
- 校验 CLI:
|
||||
- `src/summary_mcp/validate_llm_result.py`
|
||||
- 最小闭环脚本:
|
||||
- `scripts/run_summary_loop.py`
|
||||
- FreshRSS 拉取脚本:
|
||||
- `scripts/pull_freshrss_items.py`
|
||||
- FreshRSS 提取脚本:
|
||||
- `scripts/run_freshrss_extract.py`
|
||||
- 过滤脚本:
|
||||
- `scripts/run_filter_rules.py`
|
||||
- Markdown sink 脚本:
|
||||
- `scripts/run_markdown_sink.py`
|
||||
- Markdown sink 文档:
|
||||
- `docs/design/markdown-sink-design.md`
|
||||
- 当前提示词:
|
||||
- `outputs/prompts/llm-summary-prompt.txt`
|
||||
- 过滤结果样例:
|
||||
- `outputs/reference/filter/filter-decision.json`
|
||||
- 文档索引:
|
||||
- `docs/README.md`
|
||||
- 当前 TODO:
|
||||
- `TODO.md`
|
||||
- LLM 结果校验
|
||||
- `src/summary_mcp/validators/llm_result.py`
|
||||
- OpenClaw candidate 模型
|
||||
- `src/summary_mcp/models/article_candidate.py`
|
||||
- OpenClaw delivery 模型
|
||||
- `src/summary_mcp/models/openclaw_delivery.py`
|
||||
- 生产脚本入口
|
||||
- `scripts/run_freshrss_pipeline.py`
|
||||
- OpenClaw 交接说明
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
|
||||
## 当前已经验证通过
|
||||
## 当前输出规则
|
||||
|
||||
- 参考文章 URL 可提取为结构化 JSON
|
||||
- LLM 可根据提取结果生成摘要 JSON
|
||||
- validator 可校验摘要 JSON
|
||||
- 最小闭环脚本可直接调用 LLM 接口并产出通过校验的结果
|
||||
- FreshRSS API 可拉取真实 entry
|
||||
- 真实 entry 可映射为标准化 `item`
|
||||
- 标准化 `item` 可继续进入内容提取流程
|
||||
- 规则引擎可对结构化摘要结果输出 `keep / drop / review` 决策
|
||||
- MCP tool 可承接“由上层 LLM/Agent 调用过滤”的模式
|
||||
- Markdown sink 可将过滤后的结果写入本地知识库目录
|
||||
默认生产模式只输出:
|
||||
|
||||
## 当前未开始的下一阶段
|
||||
- `raw/freshrss.raw.json`
|
||||
- `candidates/openclaw-delivery-payload.json`
|
||||
- `run-report.json`
|
||||
|
||||
- 设计 webhook / 推送格式
|
||||
- 把 FreshRSS 拉取与提取流程进一步批量化/调度化
|
||||
- 迭代更细的过滤规则与个性化上下文
|
||||
- 决定是否扩展 Notion / Webhook / 其他 sink
|
||||
如果需要排障,可开启:
|
||||
|
||||
## 收束后建议从这里继续
|
||||
- `debug_artifacts=true`
|
||||
- 或脚本参数 `--debug-artifacts`
|
||||
|
||||
优先从这两个问题继续:
|
||||
这样才会额外输出逐条中间文件。
|
||||
|
||||
1. 先设计 webhook / 推送格式
|
||||
2. 再决定是否扩展其他 sink,而不是继续深化 Markdown sink 本身
|
||||
## 当前验证状态
|
||||
|
||||
已经验证通过:
|
||||
|
||||
- FreshRSS 未读拉取成功
|
||||
- 已读回写成功
|
||||
- MCP 工具入口可直接触发完整链路
|
||||
- 微信公众号样本可直接使用 RSS 提供的 `summary` 内容提取,不再回源抓网页
|
||||
- 精简输出模式已实际跑通
|
||||
|
||||
## 当前已知限制
|
||||
|
||||
- 当前对 FreshRSS 条目采用 RSS-first 策略,不再回源抓原网页
|
||||
- 如果 RSS 中没有足够正文内容,该条会直接跳过,不会进入后续总结
|
||||
- 某些规则仍偏保守,部分内容可能落到 `review`
|
||||
- `paywall` 相关启发式仍可能误判中文文本
|
||||
- Webhook / 主动投递到 OpenClaw 外部接口尚未实现,当前是由 OpenClaw 通过 MCP 主动调用
|
||||
|
||||
## 当前最建议的交接阅读顺序
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
5. `TODO.md`
|
||||
|
||||
## 一句话结论
|
||||
|
||||
当前 MVP 已完成,FreshRSS 上游、第一版规则过滤层和第一版 Markdown sink 都已接通,下一阶段应转向推送层与更完整的下游集成。
|
||||
当前仓库已经从“提取 MCP 原型”演进到“可供 OpenClaw 调用的 FreshRSS -> OpenClaw payload 上游处理器”,可以开始交接,但后续仍建议继续补 webhook / delivery 接线与规则收敛。
|
||||
|
||||
@@ -0,0 +1,328 @@
|
||||
# 规则引擎使用说明
|
||||
|
||||
## 1. 它解决什么问题
|
||||
|
||||
规则引擎负责把上游产出的结构化信号收敛成最终筛选决策:
|
||||
|
||||
`item -> content extraction -> llm summary -> rule engine -> candidate/openclaw`
|
||||
|
||||
这里有一个明确边界:
|
||||
|
||||
- LLM 负责理解正文、生成结构化摘要信号
|
||||
- 规则引擎负责输出稳定、可复现、可审计的 `keep / drop / review`
|
||||
|
||||
也就是说,规则引擎不是“再让 LLM 判断一遍”,而是用确定性规则做最后裁决。
|
||||
|
||||
## 2. 相关文件
|
||||
|
||||
- 规则模型: `src/summary_mcp/models/filtering.py`
|
||||
- 规则执行器: `src/summary_mcp/filters/engine.py`
|
||||
- 默认规则: `configs/filter_rules.json`
|
||||
- 本地脚本: `scripts/run_filter_rules.py`
|
||||
- MCP tool: `filter_summary_result`
|
||||
|
||||
## 3. 输入与输出
|
||||
|
||||
规则引擎统一读取一个 `FilterInput`,包含四部分:
|
||||
|
||||
- `item`
|
||||
- 标准化后的条目对象
|
||||
- `article`
|
||||
- 正文提取结果
|
||||
- `summary`
|
||||
- LLM 结构化摘要结果
|
||||
- `context`
|
||||
- 运行时注入的偏好信息
|
||||
|
||||
当前最常用的判断信号主要来自两类字段:
|
||||
|
||||
- `article.quality_flags.*`
|
||||
- 如 `is_low_content`、`is_truncated`、`is_paywalled`
|
||||
- `summary.*`
|
||||
- 如 `category`、`worth_keeping`、`topics`
|
||||
|
||||
输出是 `FilterDecisionResult`:
|
||||
|
||||
- `decision`
|
||||
- 最终决策,`keep / drop / review`
|
||||
- `matched_rules`
|
||||
- 命中的规则 ID 列表
|
||||
- `reasons`
|
||||
- 命中规则的原因说明
|
||||
- `labels`
|
||||
- 聚合后的标签
|
||||
- `priority`
|
||||
- 命中规则中的最高优先级
|
||||
- `matches`
|
||||
- 每条命中规则的明细
|
||||
|
||||
## 4. 规则文件怎么写
|
||||
|
||||
规则文件位置是 `configs/filter_rules.json`,顶层必须是一个 JSON 数组。
|
||||
|
||||
单条规则结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"rule_id": "keep-worth-keeping-method",
|
||||
"enabled": true,
|
||||
"stop_on_match": false,
|
||||
"conditions_all": [
|
||||
{ "field": "summary.worth_keeping", "op": "eq", "value": true },
|
||||
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
|
||||
],
|
||||
"action": {
|
||||
"decision": "keep",
|
||||
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||
"labels": ["summary", "durable"],
|
||||
"priority": 80
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
字段说明:
|
||||
|
||||
- `rule_id`
|
||||
- 规则唯一标识,建议稳定命名
|
||||
- `enabled`
|
||||
- 是否启用
|
||||
- `stop_on_match`
|
||||
- 命中后是否立即停止继续匹配后续规则
|
||||
- `conditions_all`
|
||||
- 全部命中才算命中
|
||||
- `conditions_any`
|
||||
- 任意命中即可
|
||||
- `action.decision`
|
||||
- 该规则命中时产出的规则级决策
|
||||
- `action.reason`
|
||||
- 命中原因
|
||||
- `action.labels`
|
||||
- 打到结果里的标签
|
||||
- `action.priority`
|
||||
- 执行时排序优先级,越大越先执行
|
||||
|
||||
## 5. 当前支持的操作符
|
||||
|
||||
- `eq`
|
||||
- `ne`
|
||||
- `in`
|
||||
- `not_in`
|
||||
- `contains`
|
||||
- `overlap`
|
||||
- `gte`
|
||||
- `lte`
|
||||
- `exists`
|
||||
|
||||
常见例子:
|
||||
|
||||
```json
|
||||
{ "field": "summary.worth_keeping", "op": "eq", "value": true }
|
||||
```
|
||||
|
||||
```json
|
||||
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
|
||||
```
|
||||
|
||||
```json
|
||||
{ "field": "summary.topics", "op": "overlap", "value": ["知识管理", "阅读工作流"] }
|
||||
```
|
||||
|
||||
## 6. 执行顺序和收敛逻辑
|
||||
|
||||
执行顺序不是按文件书写顺序,而是按 `action.priority` 从高到低排序。
|
||||
|
||||
命中后会先收集所有规则,再做最终收敛:
|
||||
|
||||
- 只要命中过任意 `drop`,最终就是 `drop`
|
||||
- 否则只要命中过任意 `keep`,最终就是 `keep`
|
||||
- 否则只要命中过任意 `review`,最终就是 `review`
|
||||
- 如果完全没有命中,默认 `review`
|
||||
|
||||
这意味着:
|
||||
|
||||
- `drop` 是硬拦截
|
||||
- `keep` 只能在没有更高约束的 `drop` 时生效
|
||||
- `review` 是默认灰区兜底
|
||||
|
||||
如果你希望某条高优规则一旦命中就不再继续匹配,把 `stop_on_match` 设为 `true`。
|
||||
|
||||
## 7. 动态上下文怎么用
|
||||
|
||||
规则支持从 `context` 动态取值,不需要把用户偏好写死进规则文件。
|
||||
|
||||
示例:
|
||||
|
||||
```json
|
||||
{
|
||||
"field": "summary.topics",
|
||||
"op": "overlap",
|
||||
"value": { "from_field": "context.interest_topics" }
|
||||
}
|
||||
```
|
||||
|
||||
对应的 `context` 可以是:
|
||||
|
||||
```json
|
||||
{
|
||||
"interest_topics": ["个人知识管理", "阅读工作流"]
|
||||
}
|
||||
```
|
||||
|
||||
这样同一套规则就可以被不同用户、不同运行场景复用。
|
||||
|
||||
## 8. 本地怎么跑
|
||||
|
||||
最小调用:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--output outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
带上下文:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--context outputs/reference/filter/filter-context.json ^
|
||||
--output outputs/reference/filter/filter-decision.with-context.json
|
||||
```
|
||||
|
||||
也可以显式指定另一份规则文件:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--rules configs/filter_rules.json ^
|
||||
--output outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
## 9. MCP 怎么调用
|
||||
|
||||
MCP tool 名称是 `filter_summary_result`。
|
||||
|
||||
输入参数:
|
||||
|
||||
- `summary_result`
|
||||
- `extracted_article`
|
||||
- `item`
|
||||
- `context`
|
||||
|
||||
其中只有 `summary_result` 是必填,其余都是可选补充信号。
|
||||
|
||||
示例:
|
||||
|
||||
```json
|
||||
{
|
||||
"summary_result": {
|
||||
"title": "如何构建个人阅读工作流",
|
||||
"url": "https://example.com/read-flow",
|
||||
"summary": "文章介绍了从采集、提炼到沉淀的个人阅读工作流设计。",
|
||||
"highlights": ["先采集再提炼", "用规则做稳定筛选", "日报再进入知识库"],
|
||||
"keywords": ["阅读工作流", "知识管理", "RSS", "规则引擎", "日报"],
|
||||
"topics": ["阅读工作流", "知识管理", "信息筛选"],
|
||||
"category": "方法论",
|
||||
"worth_keeping": true,
|
||||
"reason": "提供了可复用的方法框架。"
|
||||
},
|
||||
"context": {
|
||||
"interest_topics": ["阅读工作流", "知识管理"]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 10. 结果怎么看
|
||||
|
||||
一个典型结果会像这样:
|
||||
|
||||
```json
|
||||
{
|
||||
"decision": "keep",
|
||||
"matched_rules": [
|
||||
"keep-worth-keeping-method",
|
||||
"keep-interest-topic"
|
||||
],
|
||||
"reasons": [
|
||||
"Structured summary marked the content as worth keeping in a durable category.",
|
||||
"Topics overlap with current interest profile."
|
||||
],
|
||||
"labels": ["durable", "interest", "summary", "topic-match"],
|
||||
"priority": 80,
|
||||
"matches": [
|
||||
{
|
||||
"rule_id": "keep-worth-keeping-method",
|
||||
"decision": "keep",
|
||||
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||
"labels": ["summary", "durable"],
|
||||
"priority": 80
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
解读方式:
|
||||
|
||||
- 看 `decision`
|
||||
- 最终裁决
|
||||
- 看 `matched_rules`
|
||||
- 哪些规则生效了
|
||||
- 看 `reasons`
|
||||
- 为什么做出这个判断
|
||||
- 看 `matches`
|
||||
- 需要排查时看完整命中明细
|
||||
|
||||
## 11. 在当前整条链路里的位置
|
||||
|
||||
当前生产链路里,规则引擎已经集成在 `run_freshrss_openclaw_pipeline` 中。
|
||||
|
||||
顺序是:
|
||||
|
||||
1. 从 FreshRSS 拉未读
|
||||
2. 从 RSS 项目里读取正文
|
||||
3. 调用 LLM 生成结构化摘要
|
||||
4. 规则引擎输出 `keep / drop / review`
|
||||
5. 构建 `ArticleCandidateRecord`
|
||||
6. 压缩成 `OpenClawCandidateInput`
|
||||
7. 生成 `openclaw-delivery-payload.json`
|
||||
8. 如果开启 `mark_read`,最后再标记已读
|
||||
|
||||
所以在生产模式下,一般不需要单独跑 `run_filter_rules.py`,只有在调规则或排查命中逻辑时才单独跑。
|
||||
|
||||
## 12. 调规则时的建议
|
||||
|
||||
- 把“硬性淘汰”规则放高优先级
|
||||
- 比如低质量正文、明显噪音内容
|
||||
- 把“强 keep”规则放在中高优先级
|
||||
- 比如 `worth_keeping=true` 且类别是方法论
|
||||
- 把“兜底 review”规则放低一些
|
||||
- 避免过早收敛
|
||||
- 用户偏好尽量走 `context`
|
||||
- 不要把临时兴趣直接硬编码到规则里
|
||||
- `rule_id` 保持稳定
|
||||
- 方便后续审计、统计和排障
|
||||
|
||||
## 13. 一个最常见的改法
|
||||
|
||||
如果你想增加一条“命中关注主题就 keep”的规则,可以直接在 `configs/filter_rules.json` 里加:
|
||||
|
||||
```json
|
||||
{
|
||||
"rule_id": "keep-interest-topic",
|
||||
"enabled": true,
|
||||
"stop_on_match": false,
|
||||
"conditions_all": [
|
||||
{ "field": "summary.topics", "op": "overlap", "value": { "from_field": "context.interest_topics" } }
|
||||
],
|
||||
"action": {
|
||||
"decision": "keep",
|
||||
"reason": "Topics overlap with current interest profile.",
|
||||
"labels": ["interest", "topic-match"],
|
||||
"priority": 75
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
如果只是临时停用某条规则,最简单的是把它的 `enabled` 改成 `false`,不要先删规则。
|
||||
@@ -0,0 +1,163 @@
|
||||
# OpenClaw Handoff
|
||||
|
||||
## Purpose
|
||||
|
||||
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
|
||||
|
||||
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
|
||||
|
||||
OpenClaw should treat this repository as an MCP-backed upstream content processor.
|
||||
|
||||
## Production Entrypoint
|
||||
|
||||
OpenClaw should call the MCP tool:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
|
||||
This is the canonical entrypoint for production use.
|
||||
|
||||
## Required Environment Variables
|
||||
|
||||
The MCP server process must have these variables available:
|
||||
|
||||
- `FRESHRSS_API_BASE_URL`
|
||||
- `FRESHRSS_USERNAME`
|
||||
- `FRESHRSS_API_PASSWORD`
|
||||
- `LLM_API_URL`
|
||||
- `LLM_API_KEY`
|
||||
- `LLM_MODEL`
|
||||
|
||||
Example:
|
||||
|
||||
```powershell
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
set LLM_API_URL=https://api.deepseek.com
|
||||
set LLM_API_KEY=your-llm-api-key
|
||||
set LLM_MODEL=deepseek-chat
|
||||
```
|
||||
|
||||
## Server Startup
|
||||
|
||||
Install dependencies:
|
||||
|
||||
```bash
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
Start the MCP server:
|
||||
|
||||
```bash
|
||||
summary-mcp
|
||||
```
|
||||
|
||||
## Recommended MCP Call
|
||||
|
||||
Recommended default call:
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 5,
|
||||
"mark_read": false,
|
||||
"include_read": false,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
Recommended semantics:
|
||||
|
||||
- Use `mark_read=false` while validating integration.
|
||||
- Use `mark_read=true` only after confirming OpenClaw will consume the returned payload successfully.
|
||||
- Keep `debug_artifacts=false` for routine production runs.
|
||||
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
|
||||
|
||||
## What The Tool Returns
|
||||
|
||||
Primary return fields:
|
||||
|
||||
- `run_id`
|
||||
- `output_dir`
|
||||
- `raw_output`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
- `pulled_count`
|
||||
- `delivered_count`
|
||||
- `marked_read_count`
|
||||
- `status_counts`
|
||||
- `delivery_payload`
|
||||
|
||||
Optional:
|
||||
|
||||
- `items`
|
||||
- Returned only when `include_item_reports=true`
|
||||
|
||||
## Minimal Output Files
|
||||
|
||||
By default the pipeline writes only:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_id>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_id>/run-report.json`
|
||||
|
||||
If `debug_artifacts=true`, the pipeline additionally writes per-item intermediate files.
|
||||
|
||||
## Payload Specs
|
||||
|
||||
OpenClaw payload field specs live here:
|
||||
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
|
||||
## Read-State Semantics
|
||||
|
||||
The pipeline reads from FreshRSS unread items by default.
|
||||
|
||||
If `mark_read=true`:
|
||||
|
||||
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
|
||||
- only successfully delivered items are marked as read
|
||||
- failed or skipped items remain unread
|
||||
|
||||
## FreshRSS Content Policy
|
||||
|
||||
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
|
||||
|
||||
Behavior:
|
||||
|
||||
- use `item.raw_content` first
|
||||
- if missing, use `item.raw_summary`
|
||||
- if neither contains usable content, skip the item
|
||||
- do not fetch the original webpage again for FreshRSS items
|
||||
|
||||
This is intentional.
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
|
||||
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
|
||||
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
|
||||
- Paywall heuristics may produce false positives for some Chinese text patterns.
|
||||
|
||||
## Files OpenClaw Should Read First
|
||||
|
||||
Recommended reading order for a new maintainer:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
5. `docs/current/context-reset-brief.md`
|
||||
|
||||
## Current Recommendation
|
||||
|
||||
For integration handoff, the repository is usable now.
|
||||
|
||||
The minimum you need to give OpenClaw is:
|
||||
|
||||
- the repository code
|
||||
- the MCP server startup command
|
||||
- the required environment variables in the target environment
|
||||
- the instruction to call `run_freshrss_openclaw_pipeline`
|
||||
Reference in New Issue
Block a user