feat: add freshrss openclaw pipeline and clean repo

This commit is contained in:
zhuyongxin
2026-03-26 16:48:19 +08:00
parent 27fe1e8882
commit 100044e1f7
143 changed files with 1776 additions and 7293 deletions
+20 -12
View File
@@ -1,4 +1,4 @@
# 文档索引
# 文档索引
## 当前目录结构
@@ -19,25 +19,29 @@
1. `docs/current/context-reset-brief.md`
- 当前真实进度与下一步入口
2. `docs/openclaw/openclaw-candidate-input-field-spec.md`
2. `docs/openclaw/openclaw-handoff.md`
- 给 OpenClaw 的接手说明、环境变量、MCP 调用方式与已知限制
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
- 提供给 OpenClaw 的单篇结构化输入字段说明
3. `docs/openclaw/openclaw-delivery-payload-spec.md`
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
- 提供给 OpenClaw 的批量投递 envelope 说明
4. `docs/design/summary-mcp-service-design.md`
5. `docs/design/summary-mcp-service-design.md`
- 当前 MCP 服务的职责、接口和边界
5. `docs/design/filter-rule-engine-design.md`
6. `docs/design/filter-rule-engine-design.md`
- 过滤层的输入输出、规则结构与当前实现
6. `docs/design/markdown-sink-design.md`
7. `docs/design/filter-rule-engine-usage.md`
- 规则怎么写、怎么跑、结果怎么解读的使用说明
8. `docs/design/markdown-sink-design.md`
- 第一版 Markdown sink 的输入输出、目录结构与落地方式
7. `docs/openclaw/openclaw-daily-digest-refactor.md`
9. `docs/openclaw/openclaw-daily-digest-refactor.md`
- 为什么要从单篇入库改成 OpenClaw 日报聚合链路
8. `docs/openclaw/article-candidate-daily-digest-schema.md`
10. `docs/openclaw/article-candidate-daily-digest-schema.md`
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的正式设计
9. `docs/design/source-schema-design.md`
11. `docs/design/source-schema-design.md`
- `source -> item -> document` 的对象设计
10. `docs/notes/reading-pipeline-design-notes.md`
12. `docs/notes/reading-pipeline-design-notes.md`
- 更上层的阅读流方案与阶段划分
11. `docs/design/summary-loop-explained.md`
13. `docs/design/summary-loop-explained.md`
- 当前 LLM 摘要校验闭环的解释
## 当前文档分层
@@ -65,11 +69,15 @@
- 提取 JSON -> LLM 摘要 JSON -> 校验 的闭环说明
- `docs/design/filter-rule-engine-design.md`
- 第一版规则过滤引擎设计与落地位置
- `docs/design/filter-rule-engine-usage.md`
- 规则配置、调用方式与结果解读
- `docs/design/markdown-sink-design.md`
- 第一版 Markdown sink 设计与落地位置
### 3. OpenClaw 与下游设计
- `docs/openclaw/openclaw-handoff.md`
- OpenClaw 接手所需的运行说明、工具入口与已知限制
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- 提供给 OpenClaw 的单篇结构化输入字段说明
- `docs/openclaw/openclaw-delivery-payload-spec.md`
@@ -95,4 +103,4 @@
- `TODO.md` 记录任务优先级与下一步
- `outputs/README.md` 记录当前输出目录约定
- `docs/archive/content-extract-mcp-mvp-archive.md` 只当历史快照,不再作为最新事实来源
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
+92 -82
View File
@@ -1,102 +1,112 @@
# 项目当前状态简报
## 当前结论
当前仓库已经具备交付给 OpenClaw 的基础条件。
当前主链路是:
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
OpenClaw 应通过 MCP 工具 `run_freshrss_openclaw_pipeline` 调用这条链路,而不是自行拼接脚本。
## 当前已完成
- 已明确整体链路:`来源 -> 聚合池 -> 内容提取 MCP -> LLM 摘要 -> 校验 -> 过滤 -> 入库/推送`
- 已确定当前 MCP 的职责边界:
- 只负责内容提取
- 不负责摘要、分类、价值判断
- 已完成 Python MCP 骨架
- 已实现三个 tool:
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- 已完成真实 URL 提取验证
- 已完成 LLM 摘要 prompt
- 已完成 LLM 摘要结果 schema 校验器
- 已完成“提取 JSON -> LLM 摘要 JSON -> 校验”的最小闭环脚本
- 已完成 LLM 摘要校验 skill 封装
- 已清理旧的启发式 `summarizer.py`
- 已将旧的 MCP 设计文档更新为当前“Content Extract MCP”语义
- 已完成 FreshRSS `greader` API 接入
- 已完成 FreshRSS entry -> `item` 映射
- 已产出真实 `item` 样例文件
- 已跑通 `FreshRSS -> item -> content extraction` 单条链路
- 已定义过滤层输入输出 schema
- 已完成第一版规则引擎、本地脚本和 MCP tool
- 已定义 sink 输入输出 schema
- 已完成第一版 Markdown sink 与本地写入脚本
- 已完成 FreshRSS `greader` API 接入与未读拉取
- 已完成 FreshRSS 条目到标准化 `item` 的映射
- 已完成 RSS-first 提取策略
- 已完成 LLM 总结与校验闭环
- 已完成规则引擎过滤
- 已完成 `ArticleCandidateRecord` 与 `OpenClawCandidateInput` 分层
- 已完成 `OpenClawDeliveryPayload` 批量投递结构
- 已完成 FreshRSS 已读状态回写
- 已完成“仅在最终 payload 成功写盘后再标记已读”的语义
- 已完成 MCP 工具 `run_freshrss_openclaw_pipeline`
- 已完成默认精简输出模式,减少中间文件
## 当前 MCP 工具
当前服务入口:
- `src/summary_mcp/server.py`
当前暴露的 MCP 工具:
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `run_freshrss_openclaw_pipeline`
其中生产主入口是:
- `run_freshrss_openclaw_pipeline`
## 当前关键文件
- MCP 入口:
- MCP 服务入口
- `src/summary_mcp/server.py`
- 提取主流程:
- FreshRSS 统一工作流
- `src/summary_mcp/workflows/freshrss_pipeline.py`
- 摘要循环
- `src/summary_mcp/core/summary_loop.py`
- 提取主流程
- `src/summary_mcp/core/pipeline.py`
- LLM 结果模型:
- `src/summary_mcp/models/llm_result.py`
- LLM 校验器:
- `src/summary_mcp/validators/llm_result.py`
- 过滤模型:
- `src/summary_mcp/models/filtering.py`
- 过滤引擎:
- FreshRSS 集成
- `src/summary_mcp/integrations/freshrss.py`
- 规则引擎
- `src/summary_mcp/filters/engine.py`
- sink 模型:
- `src/summary_mcp/models/sink.py`
- Markdown sink:
- `src/summary_mcp/sinks/markdown.py`
- 默认过滤规则:
- `configs/filter_rules.json`
- 校验 CLI:
- `src/summary_mcp/validate_llm_result.py`
- 最小闭环脚本:
- `scripts/run_summary_loop.py`
- FreshRSS 拉取脚本:
- `scripts/pull_freshrss_items.py`
- FreshRSS 提取脚本:
- `scripts/run_freshrss_extract.py`
- 过滤脚本:
- `scripts/run_filter_rules.py`
- Markdown sink 脚本:
- `scripts/run_markdown_sink.py`
- Markdown sink 文档:
- `docs/design/markdown-sink-design.md`
- 当前提示词:
- `outputs/prompts/llm-summary-prompt.txt`
- 过滤结果样例:
- `outputs/reference/filter/filter-decision.json`
- 文档索引:
- `docs/README.md`
- 当前 TODO:
- `TODO.md`
- LLM 结果校验
- `src/summary_mcp/validators/llm_result.py`
- OpenClaw candidate 模型
- `src/summary_mcp/models/article_candidate.py`
- OpenClaw delivery 模型
- `src/summary_mcp/models/openclaw_delivery.py`
- 生产脚本入口
- `scripts/run_freshrss_pipeline.py`
- OpenClaw 交接说明
- `docs/openclaw/openclaw-handoff.md`
## 当前已经验证通过
## 当前输出规则
- 参考文章 URL 可提取为结构化 JSON
- LLM 可根据提取结果生成摘要 JSON
- validator 可校验摘要 JSON
- 最小闭环脚本可直接调用 LLM 接口并产出通过校验的结果
- FreshRSS API 可拉取真实 entry
- 真实 entry 可映射为标准化 `item`
- 标准化 `item` 可继续进入内容提取流程
- 规则引擎可对结构化摘要结果输出 `keep / drop / review` 决策
- MCP tool 可承接“由上层 LLM/Agent 调用过滤”的模式
- Markdown sink 可将过滤后的结果写入本地知识库目录
默认生产模式只输出:
## 当前未开始的下一阶段
- `raw/freshrss.raw.json`
- `candidates/openclaw-delivery-payload.json`
- `run-report.json`
- 设计 webhook / 推送格式
- 把 FreshRSS 拉取与提取流程进一步批量化/调度化
- 迭代更细的过滤规则与个性化上下文
- 决定是否扩展 Notion / Webhook / 其他 sink
如果需要排障,可开启:
## 收束后建议从这里继续
- `debug_artifacts=true`
- 或脚本参数 `--debug-artifacts`
优先从这两个问题继续:
这样才会额外输出逐条中间文件。
1. 先设计 webhook / 推送格式
2. 再决定是否扩展其他 sink,而不是继续深化 Markdown sink 本身
## 当前验证状态
已经验证通过:
- FreshRSS 未读拉取成功
- 已读回写成功
- MCP 工具入口可直接触发完整链路
- 微信公众号样本可直接使用 RSS 提供的 `summary` 内容提取,不再回源抓网页
- 精简输出模式已实际跑通
## 当前已知限制
- 当前对 FreshRSS 条目采用 RSS-first 策略,不再回源抓原网页
- 如果 RSS 中没有足够正文内容,该条会直接跳过,不会进入后续总结
- 某些规则仍偏保守,部分内容可能落到 `review`
- `paywall` 相关启发式仍可能误判中文文本
- Webhook / 主动投递到 OpenClaw 外部接口尚未实现,当前是由 OpenClaw 通过 MCP 主动调用
## 当前最建议的交接阅读顺序
1. `README.md`
2. `docs/openclaw/openclaw-handoff.md`
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
5. `TODO.md`
## 一句话结论
当前 MVP 已完成,FreshRSS 上游、第一版规则过滤层和第一版 Markdown sink 都已接通,下一阶段应转向推送层与更完整的下游集成。
当前仓库已经从“提取 MCP 原型”演进到“可供 OpenClaw 调用的 FreshRSS -> OpenClaw payload 上游处理器”,可以开始交接,但后续仍建议继续补 webhook / delivery 接线与规则收敛。
+328
View File
@@ -0,0 +1,328 @@
# 规则引擎使用说明
## 1. 它解决什么问题
规则引擎负责把上游产出的结构化信号收敛成最终筛选决策:
`item -> content extraction -> llm summary -> rule engine -> candidate/openclaw`
这里有一个明确边界:
- LLM 负责理解正文、生成结构化摘要信号
- 规则引擎负责输出稳定、可复现、可审计的 `keep / drop / review`
也就是说,规则引擎不是“再让 LLM 判断一遍”,而是用确定性规则做最后裁决。
## 2. 相关文件
- 规则模型: `src/summary_mcp/models/filtering.py`
- 规则执行器: `src/summary_mcp/filters/engine.py`
- 默认规则: `configs/filter_rules.json`
- 本地脚本: `scripts/run_filter_rules.py`
- MCP tool: `filter_summary_result`
## 3. 输入与输出
规则引擎统一读取一个 `FilterInput`,包含四部分:
- `item`
- 标准化后的条目对象
- `article`
- 正文提取结果
- `summary`
- LLM 结构化摘要结果
- `context`
- 运行时注入的偏好信息
当前最常用的判断信号主要来自两类字段:
- `article.quality_flags.*`
- 如 `is_low_content`、`is_truncated`、`is_paywalled`
- `summary.*`
- 如 `category`、`worth_keeping`、`topics`
输出是 `FilterDecisionResult`:
- `decision`
- 最终决策,`keep / drop / review`
- `matched_rules`
- 命中的规则 ID 列表
- `reasons`
- 命中规则的原因说明
- `labels`
- 聚合后的标签
- `priority`
- 命中规则中的最高优先级
- `matches`
- 每条命中规则的明细
## 4. 规则文件怎么写
规则文件位置是 `configs/filter_rules.json`,顶层必须是一个 JSON 数组。
单条规则结构:
```json
{
"rule_id": "keep-worth-keeping-method",
"enabled": true,
"stop_on_match": false,
"conditions_all": [
{ "field": "summary.worth_keeping", "op": "eq", "value": true },
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
],
"action": {
"decision": "keep",
"reason": "Structured summary marked the content as worth keeping in a durable category.",
"labels": ["summary", "durable"],
"priority": 80
}
}
```
字段说明:
- `rule_id`
- 规则唯一标识,建议稳定命名
- `enabled`
- 是否启用
- `stop_on_match`
- 命中后是否立即停止继续匹配后续规则
- `conditions_all`
- 全部命中才算命中
- `conditions_any`
- 任意命中即可
- `action.decision`
- 该规则命中时产出的规则级决策
- `action.reason`
- 命中原因
- `action.labels`
- 打到结果里的标签
- `action.priority`
- 执行时排序优先级,越大越先执行
## 5. 当前支持的操作符
- `eq`
- `ne`
- `in`
- `not_in`
- `contains`
- `overlap`
- `gte`
- `lte`
- `exists`
常见例子:
```json
{ "field": "summary.worth_keeping", "op": "eq", "value": true }
```
```json
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
```
```json
{ "field": "summary.topics", "op": "overlap", "value": ["知识管理", "阅读工作流"] }
```
## 6. 执行顺序和收敛逻辑
执行顺序不是按文件书写顺序,而是按 `action.priority` 从高到低排序。
命中后会先收集所有规则,再做最终收敛:
- 只要命中过任意 `drop`,最终就是 `drop`
- 否则只要命中过任意 `keep`,最终就是 `keep`
- 否则只要命中过任意 `review`,最终就是 `review`
- 如果完全没有命中,默认 `review`
这意味着:
- `drop` 是硬拦截
- `keep` 只能在没有更高约束的 `drop` 时生效
- `review` 是默认灰区兜底
如果你希望某条高优规则一旦命中就不再继续匹配,把 `stop_on_match` 设为 `true`。
## 7. 动态上下文怎么用
规则支持从 `context` 动态取值,不需要把用户偏好写死进规则文件。
示例:
```json
{
"field": "summary.topics",
"op": "overlap",
"value": { "from_field": "context.interest_topics" }
}
```
对应的 `context` 可以是:
```json
{
"interest_topics": ["个人知识管理", "阅读工作流"]
}
```
这样同一套规则就可以被不同用户、不同运行场景复用。
## 8. 本地怎么跑
最小调用:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--output outputs/reference/filter/filter-decision.json
```
带上下文:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--context outputs/reference/filter/filter-context.json ^
--output outputs/reference/filter/filter-decision.with-context.json
```
也可以显式指定另一份规则文件:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--rules configs/filter_rules.json ^
--output outputs/reference/filter/filter-decision.json
```
## 9. MCP 怎么调用
MCP tool 名称是 `filter_summary_result`。
输入参数:
- `summary_result`
- `extracted_article`
- `item`
- `context`
其中只有 `summary_result` 是必填,其余都是可选补充信号。
示例:
```json
{
"summary_result": {
"title": "如何构建个人阅读工作流",
"url": "https://example.com/read-flow",
"summary": "文章介绍了从采集、提炼到沉淀的个人阅读工作流设计。",
"highlights": ["先采集再提炼", "用规则做稳定筛选", "日报再进入知识库"],
"keywords": ["阅读工作流", "知识管理", "RSS", "规则引擎", "日报"],
"topics": ["阅读工作流", "知识管理", "信息筛选"],
"category": "方法论",
"worth_keeping": true,
"reason": "提供了可复用的方法框架。"
},
"context": {
"interest_topics": ["阅读工作流", "知识管理"]
}
}
```
## 10. 结果怎么看
一个典型结果会像这样:
```json
{
"decision": "keep",
"matched_rules": [
"keep-worth-keeping-method",
"keep-interest-topic"
],
"reasons": [
"Structured summary marked the content as worth keeping in a durable category.",
"Topics overlap with current interest profile."
],
"labels": ["durable", "interest", "summary", "topic-match"],
"priority": 80,
"matches": [
{
"rule_id": "keep-worth-keeping-method",
"decision": "keep",
"reason": "Structured summary marked the content as worth keeping in a durable category.",
"labels": ["summary", "durable"],
"priority": 80
}
]
}
```
解读方式:
- 看 `decision`
- 最终裁决
- 看 `matched_rules`
- 哪些规则生效了
- 看 `reasons`
- 为什么做出这个判断
- 看 `matches`
- 需要排查时看完整命中明细
## 11. 在当前整条链路里的位置
当前生产链路里,规则引擎已经集成在 `run_freshrss_openclaw_pipeline` 中。
顺序是:
1. 从 FreshRSS 拉未读
2. 从 RSS 项目里读取正文
3. 调用 LLM 生成结构化摘要
4. 规则引擎输出 `keep / drop / review`
5. 构建 `ArticleCandidateRecord`
6. 压缩成 `OpenClawCandidateInput`
7. 生成 `openclaw-delivery-payload.json`
8. 如果开启 `mark_read`,最后再标记已读
所以在生产模式下,一般不需要单独跑 `run_filter_rules.py`,只有在调规则或排查命中逻辑时才单独跑。
## 12. 调规则时的建议
- 把“硬性淘汰”规则放高优先级
- 比如低质量正文、明显噪音内容
- 把“强 keep”规则放在中高优先级
- 比如 `worth_keeping=true` 且类别是方法论
- 把“兜底 review”规则放低一些
- 避免过早收敛
- 用户偏好尽量走 `context`
- 不要把临时兴趣直接硬编码到规则里
- `rule_id` 保持稳定
- 方便后续审计、统计和排障
## 13. 一个最常见的改法
如果你想增加一条“命中关注主题就 keep”的规则,可以直接在 `configs/filter_rules.json` 里加:
```json
{
"rule_id": "keep-interest-topic",
"enabled": true,
"stop_on_match": false,
"conditions_all": [
{ "field": "summary.topics", "op": "overlap", "value": { "from_field": "context.interest_topics" } }
],
"action": {
"decision": "keep",
"reason": "Topics overlap with current interest profile.",
"labels": ["interest", "topic-match"],
"priority": 75
}
}
```
如果只是临时停用某条规则,最简单的是把它的 `enabled` 改成 `false`,不要先删规则。
+163
View File
@@ -0,0 +1,163 @@
# OpenClaw Handoff
## Purpose
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
OpenClaw should treat this repository as an MCP-backed upstream content processor.
## Production Entrypoint
OpenClaw should call the MCP tool:
- `run_freshrss_openclaw_pipeline`
This is the canonical entrypoint for production use.
## Required Environment Variables
The MCP server process must have these variables available:
- `FRESHRSS_API_BASE_URL`
- `FRESHRSS_USERNAME`
- `FRESHRSS_API_PASSWORD`
- `LLM_API_URL`
- `LLM_API_KEY`
- `LLM_MODEL`
Example:
```powershell
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
```
## Server Startup
Install dependencies:
```bash
pip install -e .
```
Start the MCP server:
```bash
summary-mcp
```
## Recommended MCP Call
Recommended default call:
```json
{
"limit": 5,
"mark_read": false,
"include_read": false,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
Recommended semantics:
- Use `mark_read=false` while validating integration.
- Use `mark_read=true` only after confirming OpenClaw will consume the returned payload successfully.
- Keep `debug_artifacts=false` for routine production runs.
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
## What The Tool Returns
Primary return fields:
- `run_id`
- `output_dir`
- `raw_output`
- `delivery_output`
- `report_output`
- `pulled_count`
- `delivered_count`
- `marked_read_count`
- `status_counts`
- `delivery_payload`
Optional:
- `items`
- Returned only when `include_item_reports=true`
## Minimal Output Files
By default the pipeline writes only:
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_id>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_id>/run-report.json`
If `debug_artifacts=true`, the pipeline additionally writes per-item intermediate files.
## Payload Specs
OpenClaw payload field specs live here:
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- `docs/openclaw/openclaw-delivery-payload-spec.md`
## Read-State Semantics
The pipeline reads from FreshRSS unread items by default.
If `mark_read=true`:
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
- only successfully delivered items are marked as read
- failed or skipped items remain unread
## FreshRSS Content Policy
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
Behavior:
- use `item.raw_content` first
- if missing, use `item.raw_summary`
- if neither contains usable content, skip the item
- do not fetch the original webpage again for FreshRSS items
This is intentional.
## Known Limitations
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
- Paywall heuristics may produce false positives for some Chinese text patterns.
## Files OpenClaw Should Read First
Recommended reading order for a new maintainer:
1. `README.md`
2. `docs/openclaw/openclaw-handoff.md`
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
5. `docs/current/context-reset-brief.md`
## Current Recommendation
For integration handoff, the repository is usable now.
The minimum you need to give OpenClaw is:
- the repository code
- the MCP server startup command
- the required environment variables in the target environment
- the instruction to call `run_freshrss_openclaw_pipeline`