Files
reader/README.md
T
root 7b791ac947 docs: 重写 README 为项目介绍
从 MCP 集成技术文档改为完整项目 README,包含:
- 项目定位与核心能力
- 架构概览(流程图 + MCP 工具层 + CLI)
- 快速开始与环境变量
- 项目结构一览
- 关键技术决策说明
- 关键词治理流程
2026-07-28 16:26:34 +08:00

198 lines
7.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Reader · AI 日报引擎
> 从 FreshRSS 到 AI 日报的自动化流水线,为个人知识管理生成每日 AI 工程化简报。
Reader 是一个端到端的 AI 日报生产系统,定时从自建 FreshRSS 的 RSS 订阅源拉取文章,经过内容提取、LLM 筛选与摘要、关键词索引构建,最终产出两个输出:
1. **公开日报** — 推送到 [Hugo 站点](https://osiman.site/daily/) 的精选技术简报
2. **知识沉淀** — 单篇结构化摘要上传到 IMA 知识库(`daily` KB)
整个流程由 OpenClaw 编排,作为 MCP Workflow Service 对外暴露。
---
## ✨ 核心能力
| 能力 | 说明 |
|:----|:------|
| **RSS 拉取** | 从 FreshRSS API 拉取订阅文章,支持增量读取与已读标记 |
| **内容提取** | 自动提取文章正文、标题、来源等结构化字段 |
| **LLM 筛选** | 基于个人兴趣画像(`filter_context.personal.json`)自动评估文章质量,分为 keep / review / drop 三档 |
| **LLM 摘要** | 并行生成每篇文章的结构化摘要(4 路并发,约 24 秒完成 7 篇) |
| **关键词索引** | 自动构建每日关键词索引,支持别名映射与停用词过滤 |
| **候选简报** | 生成 `digest-brief.json` 供编排层(OpenClaw)决策 |
| **单篇沉淀** | 对选中的文章生成结构化知识笔记,上传到 IMA 知识库 |
| **异步 Job** | 全部生产流程走异步 job,支持恢复与状态查询 |
---
## 🏗 架构概览
```
FreshRSS ──→ 拉取 ──→ 内容提取 ──→ LLM 筛选 ──→ 关键词索引
│
digest-brief.json
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Hugo 日报 IMA 知识库 term_index
(公开简报) (单篇沉淀) (关键词数据)
```
### MCP 工具层
Reader 通过 Hermes MCP 暴露 20+ 个工具,分为三类:
**日报流水线:**
- `start_freshrss_pipeline_job` → 启动异步日报 Job
- `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result` → 轮询结果
**状态查询:**
- `get_run_status` / `get_delivery_payload` / `get_run_report` → 读取运行结果
- `list_runs` / `list_run_artifacts` → 浏览运行历史
**恢复与单篇总结:**
- `inspect_resume_plan` / `start_resume_job` → 恢复失败 Job
- `start_article_summary_job` / `generate_article_summaries` → 单篇文章摘要
### CLI 入口
同步入口,适合本地 debug / fallback:
```bash
# 完整日报流水线
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
# 单篇文章摘要
python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
# 关键词维护
python scripts/build_keyword_index.py
python scripts/generate_term_cleanup_suggestions.py
```
---
## 🚀 快速开始
### 环境变量
```
# FreshRSS
FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
FRESHRSS_USERNAME=bot
FRESHRSS_API_PASSWORD=xxx
# LLM(主流水线)
LLM_API_URL=https://api.deepseek.com
LLM_API_KEY=xxx
LLM_MODEL=deepseek-chat
# LLM(可选,单篇摘要独立模型)
ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
ARTICLE_SUMMARY_LLM_API_KEY=xxx
ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
# IMA 知识库(可选,仅沉淀时需要)
IMA_DAILY_KNOWLEDGE_BASE_ID=xxx
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
```
### 运行
```bash
# 安装
pip install -e .
# 跑日报流水线(CLI 模式)
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
# 启动 MCP 服务(OpenClaw 集成用)
summary-mcp
```
---
## 📁 项目结构
```
reader/
├── configs/ # 配置
│ ├── filter_context.personal.json # 个人兴趣画像
│ ├── term_aliases.json # 关键词别名映射(149 条)
│ ├── term_stopwords.json # 关键词停用词(132 条)
│ └── term_cleanup_policy.json # 关键词清理策略
├── src/
│ └── summary_mcp/ # MCP 服务核心
│ ├── server.py # MCP 服务入口
│ ├── runtime/ # 运行时(Job 管理、状态持久化)
│ └── workflows/ # 工作流(日报流水线逻辑)
├── scripts/ # CLI 入口
├── outputs/ # 运行时产出
│ └── freshrss/
│ ├── rerun/<run_id>/ # 每次运行的全量产物
│ │ ├── candidates/ # digest-brief.json, delivery payload
│ │ ├── extracted/ # item-XX.extracted.json
│ │ └── run-state.json # 运行状态
│ └── single_summaries/ # 单篇摘要输出
├── data/
│ └── term_index/ # 关键词索引数据
│ ├── daily/YYYY-MM-DD.json
│ └── term_stats.json
├── docs/ # 设计文档
└── prompts/ # LLM Prompt 模板
```
---
## ⚙️ 关键技术决策
| 决策 | 选择 | 原因 |
|:----|:----|:------|
| 运行模式 | **异步 Job** 为主,CLI fallback | 避免 MCP 传输层 120s 超时限制 |
| 摘要并发 | **ThreadPoolExecutor(max_workers=4)** | LLM 调用是 I/O 密集型,4 路并行将 7 篇摘要从 2-3 分钟压到 ~24 秒 |
| 环境变量 | **子进程显式注入 .env** | 解决 MCP 服务器环境隔离导致子进程读取不到 LLM_API_KEY 的问题 |
| 关键词过滤 | **别名映射 + 停用词 + 语义清洗** | 先用 `term_aliases.json` 归一化,再用 `term_stopwords.json` 过滤噪声,最后通过 LLM 做语义级清洗 |
| Tag 选择 | **复用已有通用 Tag**,不从 term_index 翻生僻词 | 保持 Hugo 站点 /tags/ 页面整洁,避免大量一次性专有名词 |
---
## 🔧 关键词治理
配置治理走四步流程(`scripts/` 下脚本):
```bash
# 1. 构建评审数据包
python skills/keyword-cleanup-review/scripts/build_review_bundle.py --days 7 --top 50
# 2. 统计规则级建议(大小写、单复数、频次阈值)
python scripts/generate_term_cleanup_suggestions.py
# 3. LLM 语义级建议(中英映射、简称-全称、近义词)
python scripts/generate_term_cleanup_semantic_suggestions.py
# 4. 确认后写入配置
python scripts/apply_term_suggestions.py --accept-watch ... --dry-run
```
详见 `docs/design/keyword-engine-maintenance.md`。
---
## 📄 文档
- `docs/openclaw/README.md` — OpenClaw 集成指南
- `docs/openclaw/openclaw-orchestration-flow.md` — 编排流程
- `docs/openclaw/openclaw-delivery-payload-spec.md` — 字段契约
- `docs/design/README.md` — 设计文档总索引
- `docs/design/filter-rule-engine-design.md` — 过滤规则引擎设计
- `docs/design/filter-rule-engine-usage.md` — 过滤规则使用说明
---
## 📝 License
MIT