Compare commits

..
24 Commits
Author SHA1 Message Date
root 8e27cca166 docs: sync reader-digest-flow skill with Hermes version (absolute path, extract_failed handling, rerun include_read) 2026-08-03 09:55:19 +08:00
zhuyongxin 807027976e docs: update digest template and localize references 2026-07-29 10:20:20 +08:00
zhuyongxin 6dd8cef347 refactor: simplify reader digest skill 2026-07-28 19:14:09 +08:00
root 5eb390e3ed docs: 添加 Agent Skill 到项目仓库
- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/)
- README 新增 Agent Skill 章节说明供 Agent 使用的工作流
2026-07-28 18:38:54 +08:00
root 7b791ac947 docs: 重写 README 为项目介绍
从 MCP 集成技术文档改为完整项目 README,包含:
- 项目定位与核心能力
- 架构概览(流程图 + MCP 工具层 + CLI)
- 快速开始与环境变量
- 项目结构一览
- 关键技术决策说明
- 关键词治理流程
2026-07-28 16:26:34 +08:00
root cdbcdcd485 feat: pipeline 并行摘要 + 子进程 env 注入 + 循环导入修复
源码:
- runtime/__init__.py: resume_jobs/resume_service 改为懒加载,打破循环导入
- freshrss_pipeline_jobs.py / resume_jobs.py: 子进程注入 .env 环境变量
- freshrss_pipeline.py: LLM 摘要串行改并行 (ThreadPoolExecutor, max_workers=4)

配置:
- term_aliases: 19→149 条,大幅扩充别名映射
- term_stopwords: 19→132 条,增加过滤规则
- filter_context.personal.json: +7 个兴趣关键词
2026-07-28 16:24:31 +08:00
root 590d050218 keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions
- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
2026-05-14 17:17:49 +08:00
root 4399c9ca90 feat: digest-brief now includes review candidates (not just keep) 2026-04-28 10:39:54 +08:00
zhuyongxin 8136301ad4 fix(runtime): async article summary job startup to prevent MCP timeout
Root cause:
- RunStore.save() called 4 times synchronously before returning MCP response
- MCP stdio transport blocked by disk I/O and file locks
- Client timed out (-32001) even though job actually started

Fix:
- Use ThreadPoolExecutor to launch job in background thread
- Main thread returns immediately after creating job directory and input file
- Background thread handles subprocess.Popen and finish_stage persistence

Reference: issues/mcp-timeout-job-running.md
2026-04-16 18:37:13 +08:00
root b29cd8f934 docs: add issue - MCP start_article_summary_job timeout but job actually runs 2026-04-16 11:22:04 +08:00
root 4c4d1a45e6 docs: narrow keyword cleanup skill boundaries 2026-04-14 17:09:23 +08:00
root b8727f1885 feat: add async resume jobs and doc navigation 2026-04-14 15:55:02 +08:00
root 6705613aa4 fix(reader): include all keep candidates in digest brief 2026-04-14 10:10:25 +08:00
root 416414ae1d chore: normalize term stopwords ordering 2026-04-14 09:51:57 +08:00
root f3e7488fc8 docs: add reference templates for IMA notes and public digest 2026-04-14 09:51:57 +08:00
root e3663f681d feat: add keyword cleanup docs, skill updates, and delivery compatibility fix 2026-04-14 09:51:57 +08:00
root a06f2a1d08 chore: update term configs, skill docs and Python 3.10 compat
- term_change_log: record watch term add history
- term_watchlist: add initial watch terms from cleanup review
- term_aliases: minor update
- filter_context.personal: reorganize interest keywords
- keyword-cleanup-review skill: clarify JSON as formal artifact, markdown as temp review copy
- openclaw_delivery: add Python <3.11 UTC import compat
- daily-keyword-index-design: update suggestion artifact semantics
2026-04-14 09:51:57 +08:00
zhuyongxin c0194647d8 feat: add sort params 2026-04-13 17:42:39 +08:00
root c528e0abc7 fix(article-summary): 修复 article_summary pipeline 的 3 个 bug
1. Bug 3 (文件名含斜杠导致路径错误):
   - safe_title 生成时用 .replace('/', '-') 处理 '/' 字符
   - 同时清理连续 dash (--+) 和首尾 dash
   - 原本只处理空格,导致含 '/' 的标题写出路径错误

2. Bug 2 (delivery payload 格式不匹配):
   - 新增 '_iter_selected_items' 对 'candidates' 数组格式的支持
   - 自动 normalize selected_ids 的 'cand:' 前缀
   - candidate 条目同时支持 item_id 字段(新增)和 candidate_id(兼容)
   - 传入 delivery payload + candidate_id 时可正常匹配

3. Pipeline 补充 item_id 字段:
   - OpenClawCandidateInput 加 item_id 字段
   - build_openclaw_candidate_input 填充 item_id
   - 使得后续 article-summary 可通过 item_id 关联 extracted 文件
2026-04-13 17:38:40 +08:00
root 52ce6bfdf5 fix: validate article summary extracted path input 2026-04-11 23:04:13 +08:00
root c58d8114cd feat: add async job entrypoint for freshrss pipeline 2026-04-11 22:12:11 +08:00
root 10f9cd088f fix: fallback to repo .env for reader pipeline settings 2026-04-11 16:29:19 +08:00
root 2053cccfef feat(reader): add async article summary jobs 2026-04-10 16:37:43 +08:00
root 87d18e4263 feat: harden article summary timeout and fallback 2026-04-07 21:08:06 +08:00
66 changed files with 9935 additions and 1018 deletions
+187 -313
View File
@@ -1,340 +1,214 @@
# Reader MCP Workflow Service
# Reader · AI 日报引擎
reader 当前已经收口为面向 OpenClaw 的 MCP workflow service。正式能力边界以 FreshRSS 日报工作流为准:启动 run、写入 `run-state.json`、查询运行状态、读取结构化结果,以及最小可用的 `resume_run`。CLI 仍保留,但定位为 debug / fallback,而不是正式集成入口。
> 从 FreshRSS 到 AI 日报的自动化流水线,为个人知识管理生成每日 AI 工程化简报。
## 运行
Reader 是一个端到端的 AI 日报生产系统,定时从自建 FreshRSS 的 RSS 订阅源拉取文章,经过内容提取、LLM 筛选与摘要、关键词索引构建,最终产出两个输出:
1. **公开日报** — 推送到 [Hugo 站点](https://osiman.site/daily/) 的精选技术简报
2. **知识沉淀** — 单篇结构化摘要上传到 IMA 知识库(`daily` KB)
整个流程由 OpenClaw 编排,作为 MCP Workflow Service 对外暴露。
---
## ✨ 核心能力
| 能力 | 说明 |
|:----|:------|
| **RSS 拉取** | 从 FreshRSS API 拉取订阅文章,支持增量读取与已读标记 |
| **内容提取** | 自动提取文章正文、标题、来源等结构化字段 |
| **LLM 筛选** | 基于个人兴趣画像(`filter_context.personal.json`)自动评估文章质量,分为 keep / review / drop 三档 |
| **LLM 摘要** | 并行生成每篇文章的结构化摘要(4 路并发,约 24 秒完成 7 篇) |
| **关键词索引** | 自动构建每日关键词索引,支持别名映射与停用词过滤 |
| **候选简报** | 生成 `digest-brief.json` 供编排层(OpenClaw)决策 |
| **单篇沉淀** | 对选中的文章生成结构化知识笔记,上传到 IMA 知识库 |
| **异步 Job** | 全部生产流程走异步 job,支持恢复与状态查询 |
---
## 🏗 架构概览
```
FreshRSS ──→ 拉取 ──→ 内容提取 ──→ LLM 筛选 ──→ 关键词索引
│
digest-brief.json
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Hugo 日报 IMA 知识库 term_index
(公开简报) (单篇沉淀) (关键词数据)
```
### MCP 工具层
Reader 通过 Hermes MCP 暴露 20+ 个工具,分为三类:
**日报流水线:**
- `start_freshrss_pipeline_job` → 启动异步日报 Job
- `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result` → 轮询结果
**状态查询:**
- `get_run_status` / `get_delivery_payload` / `get_run_report` → 读取运行结果
- `list_runs` / `list_run_artifacts` → 浏览运行历史
**恢复与单篇总结:**
- `inspect_resume_plan` / `start_resume_job` → 恢复失败 Job
- `start_article_summary_job` / `generate_article_summaries` → 单篇文章摘要
### CLI 入口
同步入口,适合本地 debug / fallback:
```bash
# 完整日报流水线
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
# 单篇文章摘要
python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
# 关键词维护
python scripts/build_keyword_index.py
python scripts/generate_term_cleanup_suggestions.py
```
---
## 🚀 快速开始
### 环境变量
```
# FreshRSS
FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
FRESHRSS_USERNAME=bot
FRESHRSS_API_PASSWORD=xxx
# LLM(主流水线)
LLM_API_URL=https://api.deepseek.com
LLM_API_KEY=xxx
LLM_MODEL=deepseek-chat
# LLM(可选,单篇摘要独立模型)
ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
ARTICLE_SUMMARY_LLM_API_KEY=xxx
ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
# IMA 知识库(可选,仅沉淀时需要)
IMA_DAILY_KNOWLEDGE_BASE_ID=xxx
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
```
### 运行
```bash
# 安装
pip install -e .
# 跑日报流水线(CLI 模式)
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
# 启动 MCP 服务(OpenClaw 集成用)
summary-mcp
```
服务当前暴露 11 个工具。
---
正式 workflow service 相关工具:
## 📁 项目结构
- `run_freshrss_openclaw_pipeline`
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
- `resume_run`
单步处理 / 调试相关工具:
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `generate_article_summaries`
## 正式能力边界
- 当前正式 workflow 只有 `freshrss_daily_digest`
- 每次 FreshRSS 主流水线 run 都会在 `outputs/freshrss/rerun/<run_dir>/run-state.json` 落地运行真相
- OpenClaw 正式读取结果应优先使用 `get_delivery_payload` 与 `get_run_report`,而不是自己拼输出目录路径
- `digest-brief.json` 当前会随主流水线产出,但还没有独立的 MCP 读取工具;如需定位它,应通过 `list_run_artifacts` 或 `get_run_report` 返回的信息发现
- `run_freshrss_openclaw_pipeline` 与 `resume_run` 当前都是同步 MCP 调用;仓库里还没有后台队列 / worker / 异步任务管理
## OpenClaw 推荐调用路径
1. 调用 `run_freshrss_openclaw_pipeline` 启动正式日报 run,并保存返回的 `run_id`
2. 后续所有状态判断都基于 `get_run_status(run_id)` 或 `list_runs(...)`
3. 需要看产物列表时用 `list_run_artifacts(run_id)`,不要在 OpenClaw 里硬编码 `outputs/freshrss/rerun/...`
4. 需要消费正式结果时优先用 `get_delivery_payload(run_id)` 与 `get_run_report(run_id)`
5. 仅当 `resume_run` 的最小恢复范围满足时,才对失败 run 调用 `resume_run(run_id)`;否则应重启一个新 run
## `resume_run` 当前最小范围
- 只支持带有效 `run-state.json` 的 run
- 只支持 workflow `freshrss_daily_digest`
- 恢复时继续沿用原 `run_id`,不会新建 retry run
- 当前支持的恢复起点只有:`generate_summaries`、`apply_filters`、`build_delivery_payload`、`write_run_report`
- 当前明确不支持从 `fetch_feed`、`extract_articles` 恢复;这类失败应新开 run
- 恢复前会校验关键中间产物是否齐备,缺失时直接返回不可恢复,而不会自动回退到更早 stage
## 单篇文章总结后处理(可选使用独立 LLM)
### 生产环境推荐输入
- 单篇总结的正式生产输入,优先使用 FreshRSS 主流水线输出的**单篇 extracted 文件**:
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`
- 这些**逐条 extracted 文件**是下游单篇总结的**正式默认产物**。
- 像 `outputs/freshrss/extracted/freshrss.extracted.json` 这样的**批量 extracted 文件**,只作为临时场景、兼容旧流程的输入形态保留,**不是首选生产默认**。
### daily 知识库默认配置
- `IMA_DAILY_KNOWLEDGE_BASE_ID` —— 单篇日报总结默认上传的 IMA 知识库 ID
- `IMA_DAILY_KNOWLEDGE_BASE_NAME` —— 默认知识库名称(预期值:`daily`)
- 上传逻辑在运行时应先校验目标知识库;若配置的目标不存在,应先按名称查找,仍不存在则创建 `daily`
相关能力:
- `generate_article_summaries` MCP 工具(基于已有 extracted payload 做单篇总结)
- `scripts/run_article_summaries.py` CLI 辅助脚本
## 校验 LLM 摘要结果
```bash
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
```
reader/
├── configs/ # 配置
│ ├── filter_context.personal.json # 个人兴趣画像
│ ├── term_aliases.json # 关键词别名映射(149 条)
│ ├── term_stopwords.json # 关键词停用词(132 条)
│ └── term_cleanup_policy.json # 关键词清理策略
├── src/
│ └── summary_mcp/ # MCP 服务核心
│ ├── server.py # MCP 服务入口
│ ├── runtime/ # 运行时(Job 管理、状态持久化)
│ └── workflows/ # 工作流(日报流水线逻辑)
├── scripts/ # CLI 入口
├── outputs/ # 运行时产出
│ └── freshrss/
│ ├── rerun/<run_id>/ # 每次运行的全量产物
│ │ ├── candidates/ # digest-brief.json, delivery payload
│ │ ├── extracted/ # item-XX.extracted.json
│ │ └── run-state.json # 运行状态
│ └── single_summaries/ # 单篇摘要输出
├── data/
│ └── term_index/ # 关键词索引数据
│ ├── daily/YYYY-MM-DD.json
│ └── term_stats.json
├── docs/ # 设计文档
└── prompts/ # LLM Prompt 模板
```
## 跑最小 extraction → summary 循环
---
## ⚙️ 关键技术决策
| 决策 | 选择 | 原因 |
|:----|:----|:------|
| 运行模式 | **异步 Job** 为主,CLI fallback | 避免 MCP 传输层 120s 超时限制 |
| 摘要并发 | **ThreadPoolExecutor(max_workers=4)** | LLM 调用是 I/O 密集型,4 路并行将 7 篇摘要从 2-3 分钟压到 ~24 秒 |
| 环境变量 | **子进程显式注入 .env** | 解决 MCP 服务器环境隔离导致子进程读取不到 LLM_API_KEY 的问题 |
| 关键词过滤 | **别名映射 + 停用词 + 语义清洗** | 先用 `term_aliases.json` 归一化,再用 `term_stopwords.json` 过滤噪声,最后通过 LLM 做语义级清洗 |
| Tag 选择 | **复用已有通用 Tag**,不从 term_index 翻生僻词 | 保持 Hugo 站点 /tags/ 页面整洁,避免大量一次性专有名词 |
---
## 🔧 关键词治理
配置治理走四步流程(`scripts/` 下脚本):
```bash
python scripts/run_summary_loop.py ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--prompt outputs/prompts/llm-summary-prompt.txt ^
--output outputs/reference/summary/result.loop.json
# 1. 构建评审数据包
python skills/keyword-cleanup-review/scripts/build_review_bundle.py --days 7 --top 50
# 2. 统计规则级建议(大小写、单复数、频次阈值)
python scripts/generate_term_cleanup_suggestions.py
# 3. LLM 语义级建议(中英映射、简称-全称、近义词)
python scripts/generate_term_cleanup_semantic_suggestions.py
# 4. 确认后写入配置
python scripts/apply_term_suggestions.py --accept-watch ... --dry-run
```
## 拉取 FreshRSS 条目并映射为标准化 `item`
详见 `docs/design/keyword-engine-maintenance.md`。
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=bot
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/pull_freshrss_items.py --limit 5 --mark-read
```
---
默认会排除已经带 `read` 标签的条目。
如果你想拿到完整阅读列表,可以加 `--include-read`。
启用 `--mark-read` 后,脚本会在执行成功后把本次抓到的条目标记为已读。
## 🤖 Agent Skill
脚本会写出:
Reader 附带一个完整的 OpenClaw Agent Skill,位于 `skills/reader-digest-flow/`,供 AI Agent(Hermes / Claude Code 等)编排每日日报流程使用。
- `outputs/freshrss/raw/freshrss.raw.json`
- `outputs/freshrss/items/freshrss.items.json`
Skill 包含完整的 7 阶段工作流定义:
1. **Phase 1** — 跑 Pipeline(FreshRSS → 提取 → LLM 筛选)
2. **Phase 2** — 汇报候选(展示候选文章给用户决策)
3. **Phase 3** — 用户选文(选择 Hugo 发布文章)
4. **Phase 4** — 生成并发布 Hugo 日报
5. **Phase 5** — 用户选 IMA 沉淀文章
6. **Phase 6** — LLM 摘要生成
7. **Phase 7** — IMA 知识库上传
## 拉取 FreshRSS 条目并逐条做内容提取
以及海量铁律(不重跑 pipeline、编号规则、Tag 选择规范、IMA 上传流程等)和参考文件(`references/` 目录)。
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/run_freshrss_extract.py --limit 1 --mark-read
```
---
默认会排除已经带 `read` 标签的条目。
启用 `--mark-read` 后,只有提取成功的条目才会被标记为已读。
## 📄 文档
脚本会写出:
- `docs/openclaw/README.md` — OpenClaw 集成指南
- `docs/openclaw/openclaw-orchestration-flow.md` — 编排流程
- `docs/openclaw/openclaw-delivery-payload-spec.md` — 字段契约
- `docs/design/README.md` — 设计文档总索引
- `docs/design/filter-rule-engine-design.md` — 过滤规则引擎设计
- `docs/design/filter-rule-engine-usage.md` — 过滤规则使用说明
- `outputs/freshrss/raw/freshrss.raw.json`
- `outputs/freshrss/items/freshrss.items.json`
- `outputs/freshrss/extracted/freshrss.extracted.json`
---
注意:这个**批量 extracted 文件**主要用于独立提取场景和旧流程兼容。下游单篇总结的正式生产默认输入,仍然是 `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` 这类**逐条 extracted 文件**。
## 📝 License
## 跑完整 FreshRSS 流水线,并在最终 delivery payload 写盘成功后再标记已读
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
```
如果你希望过滤时引入个人工程兴趣 / AI Agent 兴趣画像,可以传入 context 文件:
```bash
python scripts/run_freshrss_pipeline.py ^
--limit 5 ^
--context configs/filter_context.personal.json ^
--mark-read
```
这条 CLI 与 MCP `run_freshrss_openclaw_pipeline` 共用同一条主流水线逻辑,但正式生产集成应优先走 MCP;CLI 仅用于本地 debug / fallback。默认会写出这些产物:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`(给 OpenClaw 生成 public digest 用的轻量输入,仅包含 `keep` 候选)
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`(每篇一份)
同时还会更新每日关键词索引运行数据:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
主流水线默认**不会**产出批量级的 `freshrss.extracted.json`。
如果你需要更多逐条中间产物,例如标准化 items、摘要结果、过滤决策、candidate record、candidate input,可以加 `--debug-artifacts`。
当 OpenClaw 接入这个 MCP 服务后,应以 `run_id` 作为稳定句柄:先调用 `run_freshrss_openclaw_pipeline`,再通过 `get_run_status` / `get_delivery_payload` / `get_run_report` 读取状态与结果,而不是直接拼接目录路径。
在排查复杂问题时,也可以把 `debug_artifacts=true` 打开,并结合 `list_run_artifacts` 查看该 run 下实际产物。
## 对结构化摘要结果执行确定性过滤规则
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--output outputs/reference/filter/filter-decision.json
```
如果你希望注入兴趣主题或来源标签,也可以额外传入 context 文件:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--context outputs/reference/filter/filter-context.json ^
--output outputs/reference/filter/filter-decision.with-context.json
```
规则引擎设计和规则编写说明见:
- `docs/design/filter-rule-engine-design.md`
- `docs/design/filter-rule-engine-usage.md`
## 将过滤结果写入 Markdown sink
```bash
python scripts/run_markdown_sink.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json
```
脚本会把 Markdown 笔记写到 `knowledge-base/` 下。
## 构建内部 `ArticleCandidateRecord` 与精简版 `OpenClawCandidateInput`
```bash
python scripts/run_article_candidate.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json ^
--section-hint tools_and_workflows
```
默认会写出:
- `outputs/reference/candidates/article-candidate-record.json`
- `outputs/reference/candidates/openclaw-candidate-input.json`
## 构建批量 OpenClaw delivery payload
```bash
python scripts/build_openclaw_delivery.py ^
--input-dir outputs/freshrss/candidates/batch ^
--sort-by-rank ^
--date 2026-03-25
```
默认会写出:
- `outputs/reference/candidates/openclaw-delivery-payload.json`
输出目录布局说明见 `outputs/README.md`。
## 关键词索引默认配置
相关配置文件位于:
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
你也可以基于已有 delivery payload 重新构建关键词索引:
```bash
python scripts/build_keyword_index.py ^
--input outputs/reference/candidates/openclaw-delivery-payload.json
```
运行期关键词数据存放在:
- `data/term_index/`
关键词清理评审 skill 位于:
- `skills/keyword-cleanup-review/`
构建给 LLM skill 使用的评审数据包(review bundle):
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
--days 7 ^
--top 50 ^
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
现在这个评审数据包(review bundle)还会额外携带治理上下文:
- 来自 `configs/term_cleanup_policy.json` 的清理阈值
- 当前 watch list(`configs/term_watchlist.json`)
- 最近已应用的变更(`configs/term_change_log.json`)
这个 skill 只负责生成 review 输入与建议,不会自动修改:
- `term_aliases`
- `term_stopwords`
- `filter_context.personal.json`
如果你想先预览已接受建议,再决定是否写配置文件:
```bash
python scripts/apply_term_suggestions.py ^
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
--accept-watch Cron Heartbeat Memory ^
--dry-run
```
去掉 `--dry-run` 后才会真正写文件。
这个脚本也支持通过 `--accept-alias`、`--accept-stopword`、`--accept-interest` 应用 alias / stopword / interest keyword 变更。
已接受的 watch 词会写入 `configs/term_watchlist.json`,每次应用动作也会被追加到 `configs/term_change_log.json`。
## 单篇总结 LLM 配置
如果你希望单篇总结后处理使用独立模型,而不影响主流水线,可以设置:
- `ARTICLE_SUMMARY_LLM_API_URL`
- `ARTICLE_SUMMARY_LLM_MODEL`
- `ARTICLE_SUMMARY_LLM_API_KEY`
如果这些变量未设置,单篇总结会回退使用主流程中的 `LLM_*` / `OPENAI_*` 配置。
示例(PowerShell 风格):
```bash
set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
```
然后可以这样调用 CLI。
正式生产环境建议优先使用 `outputs/freshrss/rerun/<run_id>/extracted/` 下的**逐条 extracted 文件**;下面这个**批量 extracted** 示例仅保留为兼容旧流程 / 临时场景输入:
```bash
python scripts/run_article_summaries.py ^
--extracted outputs/freshrss/extracted/freshrss.extracted.json ^
--ids 12345 67890 ^
--output-dir outputs/freshrss/single_summaries
```
也可以通过 `summary_mcp.server` 暴露的 MCP 工具 `generate_article_summaries` 调用:
- `extracted_path`(string):单篇 extracted JSON 路径(例如 `outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json`),或者包含 `results` 数组的 batch extracted JSON
- `selected_ids`(array of strings):要总结的一个或多个 `item_id`。如果传空数组,则对文件中的全部条目做总结
- `output_dir`(optional string):Markdown 输出目录;若不传,则默认写到 extracted 文件旁边的 `single_summaries/` 目录
- `llm_api_key` / `llm_model` / `llm_api_url`(optional strings):单篇总结 LLM 的覆盖配置;不传时会按前文规则回退到 `ARTICLE_SUMMARY_*` 或主 `LLM_*`
该工具返回一个 JSON 数组,内容为生成好的 Markdown 文件路径。
单篇总结使用独立 prompt:`outputs/prompts/article-summary-prompt.txt`。
它与日报 prompt 完全独立,输出的是中文结构化知识笔记,包含这些部分:
- 核心结论
- 主要论点
- 关键方法 / 机制
- 重要细节
- 可复用启发
- 关键词
- 主题
MIT
+141 -5
View File
@@ -12,11 +12,16 @@
开始编码前必须阅读:
1. `plans/reader-mcp-architecture-design.md`
2. `plans/reader-mcp-implementation-plan.md`
3. 本文件
4. `plans/issues/2026-04-06-reader-digest-sigterm.md`
5. `docs/openclaw/openclaw-handoff.md`
1. `README.md`
2. `docs/README.md`
3. `docs/openclaw/README.md`
4. `docs/openclaw/openclaw-handoff.md`
5. `docs/openclaw/openclaw-orchestration-flow.md`
6. `plans/README.md`
7. `plans/reader-mcp-architecture-design.md`
8. `plans/reader-mcp-implementation-plan.md`
9. 本文件
10. `plans/issues/2026-04-06-reader-digest-sigterm.md`
---
@@ -184,6 +189,119 @@
- 2026-04-07:已完成 `resume_run` minimal design 与现有 runtime/workflow/server 代码对齐分析,开始实现最小恢复链路。
- 2026-04-07:已完成 `resume_run` 最小实现编码,新增 runtime 恢复服务并接入 MCP server;当前进入设计对齐与本地自检。
- 2026-04-07:已完成 `resume_run` 架构对齐与本地自检;已验证 `write_run_report` 可恢复,且 `extract_articles` 会被明确拒绝恢复。
- 2026-04-14:已补 `inspect_resume_plan`、artifact-first 恢复判定,以及生产模式下稳定 `summary-batch` / `candidate-batch` artifacts;当前 `resume` 的剩余主问题不再是恢复点判断,而是同步执行模型仍可能让 OpenClaw 恢复阶段超时。
---
### [DONE][P1] 把 `resume_run` 升级为最小真异步 job
目标:
- 解决 `resume_run` 在 OpenClaw → MCP 同步链路里仍可能超时的问题
- 让恢复也具备“启动 / 轮询 / 读取结果”的正式控制面
要求:
- 新增最小异步接口:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
- 状态目录固定落到:
- `outputs/freshrss/resume_jobs/<job_id>/`
- 至少包含:
- `run-state.json`
- `input.json`
- `result.json`(成功时)
- `job-report.json`
- 启动前必须先走 `inspect_resume_plan`
- 业务执行继续复用现有 `resume_service`,不要重写恢复主逻辑
- `resume_run` 保留为同步 debug / fallback 路径,但不再作为 OpenClaw 的默认恢复入口
完成标准:
- 可恢复 run 上,`start_resume_job` 能成功返回 `job_id`
- `get_resume_job_status` 能稳定反映恢复 job 生命周期
- `get_resume_job_result` 能稳定返回 `run_id`、`resume_from_stage`、最终状态与关键产物路径
- 恢复耗时超过单次 MCP 同步窗口时,OpenClaw 仍不会因为同步调用挂住
进展备注:
- 2026-04-14:已落地 `src/summary_mcp/runtime/resume_jobs.py` 与 `scripts/run_resume_job.py`,新增 `start_resume_job` / `get_resume_job_status` / `get_resume_job_result`
- 2026-04-14:启动前会先走 `inspect_resume_plan`;不可恢复 run 会在 job 输入校验阶段直接失败,不进入后台恢复执行
- 2026-04-14:后台执行复用现有 `_resume_freshrss_run(...)`,没有重写恢复主逻辑
- 2026-04-14:已完成本地 synthetic 验证:`write_run_report` 恢复可通过 `start -> poll -> result` 闭环成功收敛
---
### [DONE][P1] 单篇总结改为最小真异步 job
目标:
- 解决 `generate_article_summaries` 在 OpenClaw → MCP 同步链路里易 timeout 的问题
- 将单篇总结正式升级为可启动、可轮询、可读取结果的异步 job
要求:
- 新增最小异步接口:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
- 状态目录固定落到:
- `outputs/freshrss/article_summary_jobs/<job_id>/`
- 至少包含:
- `run-state.json`
- `input.json`
- `result.json`(成功时)
- `job-report.json`
- 执行模型优先使用后台子进程,不使用线程
- 业务逻辑继续复用 `summarize_selected_articles(...)`,不要重写正文总结核心逻辑
- 对 OpenClaw / reader-digest-flow 而言,异步 job 成功后应可继续接 IMA 沉淀闭环
当前进展:
- 2026-04-10:已完成方案文档 `plans/article-summary-async-job-plan.md`
- 2026-04-10:已落地最小代码骨架:
- `src/summary_mcp/runtime/article_summary_jobs.py`
- `scripts/run_article_summary_job.py`
- `src/summary_mcp/server.py` 已新增 3 个 async job tools
- 2026-04-10:已用真实 extracted 文件验证最小异步链路可跑通,job 能成功进入 `running → success`,并可读回结果
- 2026-04-10:已补 README / OpenClaw handoff 文档,并对外统一为 `start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
- 2026-04-10:已完成聚焦自检:
- MCP tool 注册名校验通过
- stubbed async job 成功路径通过
- stubbed async job 失败路径通过,`error_summary` 与 `job-report.json` 可回读
下一步:
- 将 `reader-digest-flow` 正式默认路径切到 async job
- 用真实 LLM 配置再做一次非 stub 的服务端冒烟验证
---
### [DOING][P0] FreshRSS 主日报 run 改为最小真异步 job
目标:
- 解决 `run_freshrss_openclaw_pipeline` 在正式生产链路里仍为同步 MCP 调用、易超时的问题
- 将 FreshRSS 主日报启动路径升级为可启动、可轮询、可读取结果的异步 job
要求:
- 新增最小异步接口:
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
- `get_freshrss_pipeline_job_result`
- 状态目录固定落到:
- `outputs/freshrss/pipeline_jobs/<job_id>/`
- 至少包含:
- `run-state.json`
- `input.json`
- `result.json`(成功时)
- `job-report.json`
- 执行模型优先使用后台子进程,不使用线程
- 主业务逻辑继续复用 `run_freshrss_pipeline(...)`,不要重写日报核心逻辑
- job 成功后结果中必须带回 `run_id` 与关键产物路径
- README / handoff / OpenClaw 生产建议路径需要同步改成 async start path
当前进展:
- 2026-04-11:问题定位完成,确认之前异步化的是 article-summary,不是主日报 run
- 2026-04-11:已新增方案文档 `plans/freshrss-pipeline-async-job-plan.md`
下一步:
- 复用 article-summary job runtime 骨架实现主日报 async job
- 新增后台 runner 脚本
- 暴露 3 个 MCP tools
- 用真实 MCP 冒烟验证 `start -> status -> result`
---
@@ -202,6 +320,23 @@
---
### [DONE][P1] interest/watch 候选引擎从固定阈值改为百分位排名 + 增速因子
目标:
- 解决固定阈值(total_count>=3)不随数据量自适应的问题
- 引入趋势信号(growth 因子),识别近期集中爆发的词
- 支持 7 天、41 天、200 天数据量下取同样的 top 5%/5%-20% 而不需调阈值
要求:
- `build_review_bundle.py`:新增 percentile 和 growth 计算函数;候选池从固定阈值改为百分位 + 增速
- `configs/term_cleanup_policy.json`:升级为 v2 schema,percentile/growth 替代绝对阈值
- 不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
- 全量跑一次对比新旧产出,确认差异合理
方案文档:`plans/keyword-cleanup-interest-watch-engine-improvement.md`
---
### [DONE][P3] 更新 README / handoff / docs,明确 MCP 为正式入口
目标:
@@ -228,6 +363,7 @@
- 2026-04-07:新增架构设计文档 `plans/reader-mcp-architecture-design.md`
- 2026-04-07:新增实施计划文档 `plans/reader-mcp-implementation-plan.md`
- 2026-04-07:完成 `freshrss` pipeline 的 run-state 基础设施,新增 `runtime` 包并覆盖关键 stages 状态持久化。
- 2026-04-14:已补 OpenClaw 文档导航、历史归档、design/notes/plans 导航,并统一当前正式口径为 async job 编排入口。
### 风险提醒
+52 -20
View File
@@ -23,34 +23,66 @@
"前沿科技"
],
"interest_keywords": [
"Java",
"Go",
"Python",
"Spring",
"Agent",
"Agent Skills",
"AgentScope",
"AI Agent",
"AI Coding Agent",
"AliSQL",
"Anthropic",
"Claude",
"Claude Code",
"CLAUDE.md",
"CLI",
"Context Engineering",
"Cursor",
"ChatGPT",
"DeepSeek",
"FastAPI",
"Gin",
"Go",
"gRPC",
"MySQL",
"PostgreSQL",
"Redis",
"Harness Engineering",
"Hermes Agent",
"Java",
"Kafka",
"微服务",
"可观测性",
"Kubernetes",
"云原生",
"AI Agent",
"Agent",
"LLM",
"RAG",
"Loop Engineering",
"MCP",
"Prompt Engineering",
"Workflow",
"向量数据库",
"知识库",
"MoE",
"MySQL",
"MySQL复制延迟",
"OpenAI",
"DeepSeek",
"OpenClaw",
"AliSQL",
"MySQL复制延迟"
"PostgreSQL",
"Prompt Engineering",
"Python",
"RAG",
"ReAct",
"ReActAgent",
"Redis",
"Skill",
"SKILL.md",
"Skills",
"Spring",
"SubAgent",
"TypeScript",
"Vibe Coding",
"Workflow",
"上下文压缩",
"上下文工程",
"上下文管理",
"云原生",
"代码审查",
"可观测性",
"向量数据库",
"多Agent协作",
"大模型",
"子Agent",
"强化学习",
"微服务",
"渐进式披露",
"知识库"
]
}
+146 -2
View File
@@ -1,5 +1,149 @@
{
"AI助手": "AI Agent",
"Agent": "Agent",
"Agent框架": "Agent",
"Agent能力": "Agent Skills",
"智能体": "AI Agent",
"Agentic架构": "Agentic架构",
"多Agent协作": "多Agent协作",
"多智能体架构": "多Agent协作",
"Multi-Agent": "多Agent",
"Subagent": "子Agent",
"Sub Agents验证": "子Agent",
"子Agent": "子Agent",
"子智能体": "子Agent",
"Coding Agent": "AI Coding Agent",
"AI编程": "AI Coding Agent",
"AI辅助编程": "AI Coding Agent",
"代码生成": "AI代码生成",
"代码审查": "Code Review",
"Prompt": "Prompt Engineering",
"Prompt Caching": "提示缓存",
"RAG": "RAG",
"图文RAG": "RAG",
"Prompt架构": "Prompt Engineering"
"向量检索": "向量检索",
"向量嵌入": "向量嵌入",
"Multi-Token Prediction": "多Token预测",
"Pair-In Pair-Out": "PIPO架构",
"PIPO": "PIPO架构",
"上下文管理": "上下文管理",
"上下文卸载": "上下文卸载",
"Self-GC": "上下文压缩",
"记忆管理": "上下文管理",
"会话管理": "上下文管理",
"Harness Engineering": "Harness工程化",
"Harness架构": "Harness工程化",
"Harness": "Harness工程化",
"Loop Engineering": "Loop Engineering",
"推理加速": "推理加速",
"推理深度": "推理深度",
"长链路推理": "长链路推理",
"RLVR": "RLVR",
"GRPO": "GRPO",
"强化学习": "强化学习",
"Multi-Agent RL": "多Agent强化学习",
"Viking AI搜索": "AI搜索",
"Viking AI Search": "AI搜索",
"智能搜索": "AI搜索",
"SearchCLI": "CLI搜索",
"视频生成": "AI视频生成",
"视频生成模型": "AI视频生成",
"LingBot-Video": "AI视频生成",
"视觉自回归模型": "AI视频生成",
"火山云数据库PostgreSQL Serverless版": "Serverless数据库",
"PostgreSQL": "PostgreSQL",
"MySQL": "MySQL",
"OceanBase": "OceanBase",
"StarRocks": "StarRocks",
"Milvus": "Milvus",
"Seal AI Zone": "AI安全",
"NEX沙箱": "沙箱隔离",
"MicroVM": "沙箱隔离",
"安全左移": "安全左移",
"安全中台": "AI安全",
"成本降低": "成本优化",
"成本杠杆": "成本优化",
"Scale-to-Zero": "弹性伸缩",
"Data as Git": "数据分支管理",
"Schema Diff": "Schema对比",
"Time Travel": "数据回溯",
"多端架构": "多端架构",
"契约化": "契约化架构",
"大仓": "大仓工程化",
"Vibe Coding": "Vibe Coding",
"LLM Judge": "LLM评估",
"SWE-Bench": "SWE-Bench",
"SWE Bench Pro": "SWE-Bench",
"SWE-Bench Pro": "SWE-Bench",
"Verification Agent": "验证Agent",
"CLI工具": "CLI",
"CLI": "CLI",
"漏桶算法": "限流架构",
"固定窗口限流": "限流架构",
"Suspend消费控制": "限流架构",
"RocketMQ LiteTopic": "消息队列",
"LLM Wiki": "LLM知识库",
"知识工程": "知识工程",
"语义资产": "语义资产管理",
"知识图谱": "知识图谱",
"知识库沉淀": "知识管理",
"Skill": "Skill",
"Skill Hub": "技能生态",
"具身智能": "具身智能",
"Open X-Embodiment": "具身智能",
"YOLO Classifier": "目标检测",
"MCP": "MCP",
"MCP连接器": "MCP",
"缓存击穿": "缓存优化",
"GPU算力调度": "算力调度",
"异构资源": "异构计算",
"XPU": "异构计算",
"弹性RDMA": "RDMA网络",
"国内主流GPU": "国产芯片",
"国产AI芯片": "国产芯片",
"Paxos协议": "分布式一致性",
"Token": "Token管理",
"Token效率": "Token管理",
"百万token上下文": "长上下文",
"MoE": "MoE架构",
"MoE架构": "MoE架构",
"思维链": "思维链",
"CoT Distillation": "思维链蒸馏",
"自然语言驱动": "自然语言交互",
"NL2SQL": "NL2SQL",
"AI对齐": "AI对齐",
"注意力机制": "注意力机制",
"多模态": "多模态",
"音视频工作台": "音视频处理",
"AI助手": "AI Agent",
"Agent架构": "AI Agent",
"Agent专业化": "AI Agent",
"Agent Teams": "多Agent协作",
"Agentic Engineering": "AI Agent",
"AI智能体": "AI Agent",
"LLM Agent": "AI Agent",
"AI Harness": "Harness Engineering",
"AI代码生成": "AI Coding Agent",
"Memory管理": "上下文管理",
"Agent Skill": "Agent Skills",
"Binlog": "binlog",
"vibe coding": "Vibe Coding",
"Agent组织化协作平台": "Agent协作平台",
"Anthropic": "Anthropic",
"OpenClaw": "OpenClaw",
"WorkBuddy": "WorkBuddy",
"Claude": "Claude",
"ChatGPT": "ChatGPT",
"GPT": "GPT",
"Opus": "Opus",
"Sonnet": "Sonnet",
"Grok": "Grok",
"Qwen": "Qwen",
"GLM": "GLM",
"Claude Code": "Claude Code",
"Cursor": "Cursor",
"Codex": "Codex",
"Pi": "Pi",
"CoT": "CoT",
"SVG": "SVG",
"TTS": "TTS"
}
+592 -1
View File
@@ -1,4 +1,595 @@
{
"schema_version": "v1",
"entries": []
"entries": [
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_watch_term",
"term": "A2A",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_watch_term",
"term": "Agentic Loop",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_watch_term",
"term": "AI Gateway",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_watch_term",
"term": "Claude Skills",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_watch_term",
"term": "Cron",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_watch_term",
"term": "CoPaw",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_interest_keyword",
"term": "Claude Code",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=7, days_seen=4, recent_count=4.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_interest_keyword",
"term": "Agent Skills",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=3, recent_count=0.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_interest_keyword",
"term": "SubAgent",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=2.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_interest_keyword",
"term": "AgentScope",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=1.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-04-08T02:34:14.194320Z",
"action": "add_interest_keyword",
"term": "ReActAgent",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=1.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Anthropic",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=13, days_seen=10, recent_count=13.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Harness Engineering",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=12, days_seen=11, recent_count=12.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Skill",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=11, days_seen=9, recent_count=11.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "上下文工程",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "多Agent协作",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Claude",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "上下文管理",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "渐进式披露",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=5, recent_count=5.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Skills",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "SKILL.md",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "CLAUDE.md",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "上下文压缩",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "AI Coding Agent",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Hermes Agent",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Vibe Coding",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Context Engineering",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Cursor",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "大模型",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=3, recent_count=4.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_interest_keyword",
"term": "TypeScript",
"reason": "Core language for AI agent development (e.g., Claude Code, Cursor) and backend engineering, complements existing Python/Java/Go keywords.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_interest_keyword",
"term": "代码审查",
"reason": "Chinese term for 'code review', a key practice in backend engineering and AI agent development workflows.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Channels",
"reason": "Too generic; could refer to communication channels, YouTube channels, or software channels, not specific to user's focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Memory",
"reason": "Extremely broad term; could refer to computer memory, human memory, or memory in various contexts, not discriminative enough.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Prompt",
"reason": "Already covered by 'Prompt Engineering' as a more specific term; 'Prompt' alone is too broad and matches many unrelated articles.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "AGI",
"reason": "Too broad and speculative; not directly actionable for the user's practical engineering focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "AI日报",
"reason": "Generic news term; not a technical concept or tool, would add noise to the keyword index.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "AIHOT",
"reason": "Unclear meaning, likely a brand or aggregator, not a specific technical term.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "All In Code",
"reason": "Too vague; could refer to a podcast, a philosophy, or a project, not a specific technical concept.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "auto-twitter-campaign",
"reason": "Too specific to a single project/tool, not a general interest keyword for the user's focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "ChangeSet",
"reason": "Generic term used in version control and databases; too broad to be a useful filter.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Lumina",
"reason": "Unclear reference; could be a product, framework, or brand, not clearly aligned with user's focus.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "OpenViking",
"reason": "Unclear reference; not a known tool or concept in the user's stated focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Seedance 2.0",
"reason": "Unclear reference; likely a product or version, not a general technical term.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "质量门禁",
"reason": "Chinese term for 'quality gate', too generic in software engineering; not specific to user's focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agent Skill",
"reason": "Singular variant",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "Agent Skills",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Binlog",
"reason": "Case variant (auto-ranked)",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "binlog",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Coding Agent",
"reason": "Abbreviated form of 'AI Coding Agent', referring to the same concept.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Coding Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Subagent",
"reason": "Case variant",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "SubAgent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Subagents",
"reason": "Plural variant",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "SubAgent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "vibe coding",
"reason": "Case variant",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "Vibe Coding",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agent架构",
"reason": "Chinese translation of 'Agent architecture', a core concept in AI Agent engineering.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agent专业化",
"reason": "Chinese term for 'Agent specialization', directly related to Agent engineering.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agent Teams",
"reason": "English equivalent of 'Multi-Agent collaboration', same concept.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "多Agent协作",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agentic Engineering",
"reason": "Broader term for engineering with AI agents, closely related to Agent engineering focus.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "CLI工具",
"reason": "Chinese translation of 'CLI tool', same concept.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "CLI",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "AI编程",
"reason": "Chinese term for 'AI programming', closely related to AI Coding Agent.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Coding Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "记忆管理",
"reason": "Chinese term for 'memory management', closely related to context management in LLM applications.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "上下文管理",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "会话管理",
"reason": "Chinese term for 'session management', related to context management in LLM applications.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "上下文管理",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "强化学习",
"reason": "top 0.8% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=10, days_seen=10, recent_count=10.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "ReAct",
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=8, recent_count=8.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "CLI",
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=7, recent_count=8.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "子Agent",
"reason": "top 1.3% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=7, days_seen=6, recent_count=7.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "Loop Engineering",
"reason": "top 1.5% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "MoE",
"reason": "top 2.0% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
}
]
}
+10 -9
View File
@@ -1,14 +1,12 @@
{
"schema_version": "v1",
"schema_version": "v2",
"interest_keyword_review": {
"min_total_count": 3,
"min_days_seen": 2
"percentile_max": 0.05,
"growth_promotion": 0.5
},
"watch_term_review": {
"min_total_count": 1,
"min_days_seen": 1,
"max_total_count": 2,
"max_days_seen": 2
"percentile_min": 0.05,
"percentile_max": 0.20
},
"alias_review": {
"min_total_count": 2,
@@ -19,7 +17,10 @@
"max_days_seen": 2
},
"notes": [
"当前阶段采用保守阈值,避免在低样本条件下直接扩充 interest_keywords。",
"watch_terms 先用于观察,后续再决定是否升格为 interest_keywords 或进入 alias/stopword 配置。"
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
"growth = recent_count / total_count,衡量近期活跃度",
"watch_term_review 的 percentile_min 可理解为兴趣边界下限,低于此值的词归入 interest 候选",
"growth_promotion(默认 0.5)用于识别近期集中爆发词,即使排位不高也主动推荐确认"
]
}
+121 -1
View File
@@ -1,6 +1,126 @@
[
"1688",
"AGI",
"AIHOT",
"AI日报",
"All In Code",
"Andrej Karpathy",
"Anthropic",
"auto-twitter-campaign",
"Boundaries",
"ChangeSet",
"Channels",
"Claude Fable 5",
"Claude Mythos",
"Cohere",
"Confidence Head",
"Cosmos 3",
"Databricks",
"DINOv2",
"domain-mapping",
"Dropbox",
"EchoGen",
"FLUX.1-dev VAE",
"GB300 GPU",
"GLM 5.2",
"GLM5.0",
"GPT-5.5",
"GPT-Live",
"GPT5.5",
"Grok 4.5",
"GrowBrain",
"iMedImage",
"iMedLoop",
"iMedMaaS",
"iMedStudio",
"J-space",
"JLens",
"John Jumper",
"J空间",
"KAIROS",
"KubeRay",
"LibTV Agent",
"LingBot-Video",
"Lumina",
"Markdown",
"Marvis",
"MDASH",
"Meta Superintelligence Labs",
"MTS",
"Muse Image",
"Muse Video",
"N-gram Embedding",
"OCP China",
"OCP China 2026",
"On-Policy Distillation",
"OPC训练营",
"OpenAI",
"OpenBMC",
"OpenClaw",
"OpenViking",
"Opus 4.8",
"Qwen3",
"Qwen3-30B-A3B",
"RAS API",
"Redfish",
"ScMoE",
"Seal AI Zone",
"SealRouter",
"Seedance 2.0",
"Sonnet 5",
"Spec模式",
"STE固件团队",
"Three.js",
"Unity AI Gateway",
"Vant Weapp",
"WeTV",
"WorkBuddy",
"wpc",
"YOLO Classifier",
"一人公司",
"中国科学技术大学",
"五大扶持体系",
"出门问问",
"分镜",
"剧本",
"奋斗文化",
"字节跳动",
"小银",
"得力",
"德适科技",
"成都天府长岛",
"扣子",
"星云平台",
"火山引擎",
"百度百舸",
"百炼网关",
"科大讯飞",
"腾讯云开发者社区",
"腾讯混元Hy3",
"蚂蚁灵波",
"贝尔实验室",
"质量门禁",
"配乐",
"配音",
"银行客户经理",
"飞书妙搭",
"飞盘物理",
"奋斗文化"
"自动化",
"定时任务",
"开源模型",
"陌生化",
"AlphaFold",
"Brand Kit",
"DataWorks",
"Enhance-Nanocodec",
"IRIS Codec",
"Lovart",
"MiniMax M3",
"Gemini 3.5 Flash",
"Codex",
"CodeBuddy",
"Claude Cowork",
"AGENTS.md",
"Claude",
"RLVR"
]
+45 -2
View File
@@ -1,5 +1,48 @@
{
"schema_version": "v1",
"updated_at": "2026-03-27T00:00:00Z",
"terms": []
"updated_at": "2026-07-15T02:17:50.155586Z",
"terms": [
{
"term": "A2A",
"added_at": "2026-04-08T02:34:14.194320Z",
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
"status": "watching"
},
{
"term": "Agentic Loop",
"added_at": "2026-04-08T02:34:14.194320Z",
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
"status": "watching"
},
{
"term": "AI Gateway",
"added_at": "2026-04-08T02:34:14.194320Z",
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
"status": "watching"
},
{
"term": "Claude Skills",
"added_at": "2026-04-08T02:34:14.194320Z",
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
"status": "watching"
},
{
"term": "CoPaw",
"added_at": "2026-04-08T02:34:14.194320Z",
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
"status": "watching"
},
{
"term": "Cron",
"added_at": "2026-04-08T02:34:14.194320Z",
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
"status": "watching"
}
]
}
+50 -38
View File
@@ -1,6 +1,25 @@
# 文档索引
## 当前目录结构
## 当前最短阅读路径
1. `README.md`
- 仓库入口与常用脚本
2. `docs/current/context-reset-brief.md`
- 当前状态的最短摘要
3. `docs/openclaw/README.md`
- OpenClaw 集成文档导航
4. `docs/openclaw/openclaw-handoff.md`
- OpenClaw 接手总览
5. `docs/openclaw/openclaw-orchestration-flow.md`
- OpenClaw 正式编排手册
6. `docs/design/README.md`
- 设计文档导航,区分当前有效设计与背景草案
7. `plans/README.md`
- 规划文档导航
8. `TODO.md`
- 当前任务状态
## 目录结构
- `docs/README.md`
- 文档总索引
@@ -8,45 +27,18 @@
- 当前状态、收束入口、阶段导航
- `docs/design/`
- 当前实现的设计文档
- `docs/design/README.md`
- 设计文档导航
- `docs/openclaw/`
- OpenClaw 日报聚合与下游对象设计
- OpenClaw 集成文档、对象规范与历史归档
- `docs/notes/`
- 较上层的方案笔记与非最终设计
- `docs/notes/README.md`
- notes 导航
- `docs/archive/`
- 历史归档,不作为最新事实来源
## 当前推荐阅读顺序
1. `docs/current/context-reset-brief.md`
- 当前真实进度与下一步入口
2. `docs/openclaw/openclaw-handoff.md`
- 给 OpenClaw 的接手说明、环境变量、MCP 调用方式与已知限制
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
- 提供给 OpenClaw 的单篇结构化输入字段说明
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
- 提供给 OpenClaw 的批量投递 envelope 说明
5. `docs/design/summary-mcp-service-design.md`
- 当前 MCP 服务的职责、接口和边界
6. `docs/design/filter-rule-engine-design.md`
- 过滤层的输入输出、规则结构与当前实现
7. `docs/design/filter-rule-engine-usage.md`
- 规则怎么写、怎么跑、结果怎么解读的使用说明
8. `docs/design/daily-keyword-index-design.md`
- 日报级词元库与周期性词元清洗 skill 设计
9. `docs/design/markdown-sink-design.md`
- 第一版 Markdown sink 的输入输出、目录结构与落地方式
10. `docs/openclaw/openclaw-daily-digest-refactor.md`
- 为什么要从单篇入库改成 OpenClaw 日报聚合链路
11. `docs/openclaw/article-candidate-daily-digest-schema.md`
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的正式设计
12. `docs/design/source-schema-design.md`
- `source -> item -> document` 的对象设计
13. `docs/notes/reading-pipeline-design-notes.md`
- 更上层的阅读流方案与阶段划分
14. `docs/design/summary-loop-explained.md`
- 当前 LLM 摘要校验闭环的解释
## 当前文档分层
## 按主题阅读
### 1. 当前状态与导航
@@ -58,11 +50,15 @@
- 当前阶段状态的最短摘要
- `docs/README.md`
- 文档索引与阅读顺序
- `plans/README.md`
- 规划文档导航
### 2. 当前实现设计
- `docs/design/README.md`
- 设计文档导航与状态说明
- `docs/design/summary-mcp-service-design.md`
- 当前内容提取 MCP 的真实设计
- 早期 content-extract MCP 设计草案,现主要保留背景参考价值
- `docs/design/summary-core-interface-design.md`
- 摘要/提取内核的接口抽象
- `docs/design/source-schema-design.md`
@@ -82,19 +78,23 @@
### 3. OpenClaw 与下游设计
- `docs/openclaw/README.md`
- OpenClaw 相关文档导航与归档边界
- `docs/openclaw/openclaw-handoff.md`
- OpenClaw 接手所需的运行说明、工具入口与已知限制
- `docs/openclaw/openclaw-orchestration-flow.md`
- OpenClaw 编排层的正式运行手册
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- 提供给 OpenClaw 的单篇结构化输入字段说明
- `docs/openclaw/openclaw-delivery-payload-spec.md`
- 提供给 OpenClaw 的批量投递 envelope 说明
- `docs/openclaw/openclaw-daily-digest-refactor.md`
- 改造为 OpenClaw 日报聚合链路的原因与目标结构
- `docs/openclaw/article-candidate-daily-digest-schema.md`
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的字段设计与对象关系
### 4. 方案笔记
- `docs/notes/README.md`
- notes 导航与使用边界
- `docs/notes/reading-pipeline-design-notes.md`
- 整体阅读流、规则、sink、push 的方案笔记
@@ -102,11 +102,23 @@
- `docs/archive/content-extract-mcp-mvp-archive.md`
- MVP 阶段归档,部分状态已被后续进展覆盖
- `docs/openclaw/archive/README.md`
- OpenClaw 历史文档归档说明
- `docs/openclaw/archive/formalization-summary-2026-04-07.md`
- 第一阶段正式化总结
- `docs/openclaw/archive/openclaw-daily-digest-refactor.md`
- 早期日报聚合改造背景
- `docs/openclaw/archive/digest-optimization-summary.md`
- 早期 digest 优化总结
- `docs/openclaw/archive/p1-status-reconciliation-plan-2026-04-14.md`
- `resume` / 状态收敛问题的阶段修复计划与回填
## 当前文档维护原则
## 维护原则
- `docs/current/context-reset-brief.md` 记录当前最新状态
- `TODO.md` 记录任务优先级与下一步
- `plans/README.md` 负责规划文档分层与导航
- `docs/design/README.md` 负责设计文档分层与导航
- `outputs/README.md` 记录当前输出目录约定
- `docs/archive/content-extract-mcp-mvp-archive.md` 只当历史快照,不再作为最新事实来源
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
+73 -121
View File
@@ -2,152 +2,104 @@
## 当前结论
当前仓库已经具备交付给 OpenClaw 的基础条件。
当前仓库已经具备作为 OpenClaw 上游服务的正式基础能力。
当前主链路是:
当前正式主链路是:
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
OpenClaw 应通过 MCP 工具 `run_freshrss_openclaw_pipeline` 调用这条链路,而不是自行拼接脚本。
当前正式控制面已经收口为异步 job:
## 当前已完成
- 主日报:`start_freshrss_pipeline_job -> poll -> get result`
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
- 单篇总结:`start_article_summary_job -> poll -> get result`
- 已完成 FreshRSS `greader` API 接入与未读拉取
- 已完成 FreshRSS 条目到标准化 `item` 的映射
- 已完成 RSS-first 提取策略
- 已完成 LLM 总结与校验闭环
- 已完成规则引擎过滤
- 已完成 `ArticleCandidateRecord` 与 `OpenClawCandidateInput` 分层
- 已完成 `OpenClawDeliveryPayload` 批量投递结构
- 已完成 FreshRSS 已读状态回写
- 已完成“仅在最终 payload 成功写盘后再标记已读”的语义
- 已完成 MCP 工具 `run_freshrss_openclaw_pipeline`
- 已完成默认精简输出模式,减少中间文件
- 已完成日报级 `keywords` 词元库与全局词频统计
- 已完成 `keyword-cleanup-review` skill 骨架与 review bundle 脚本
- 已完成低复杂治理层:`term_cleanup_policy` / `term_watchlist` / `term_change_log`
- 已完成采纳建议写回脚本 `scripts/apply_term_suggestions.py`
同步 `run_freshrss_openclaw_pipeline`、`resume_run`、`generate_article_summaries` 仍保留,但只用于 debug / fallback。
## 当前 MCP 工具
## 当前权威入口
当前服务入口:
先看这些文档:
- `src/summary_mcp/server.py`
1. `README.md`
2. `docs/README.md`
3. `docs/openclaw/README.md`
4. `docs/openclaw/openclaw-handoff.md`
5. `docs/openclaw/openclaw-orchestration-flow.md`
6. `plans/README.md`
7. `TODO.md`
当前暴露的 MCP 工具:
如果问题是 OpenClaw 集成、状态分支或恢复策略,优先看 `docs/openclaw/`,不要先翻历史计划。
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `run_freshrss_openclaw_pipeline`
## 当前正式能力
其中生产主入口是:
- FreshRSS 主日报 run 会落地 `run-state.json`
- `get_run_status` / `list_runs` / `list_run_artifacts` 提供 run 级观测
- `get_delivery_payload` / `get_run_report` 提供正式结果读取
- 查询层已经支持 stale state 与终态 artifacts 的状态收敛
- 主日报正式启动已切到 async job
- `resume` 已切到 async job,并在执行前先做 `inspect_resume_plan`
- 生产恢复依赖 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`
- 单篇总结也已补齐 async job 形态
- `run_freshrss_openclaw_pipeline`
## 当前关键代码入口
## 当前关键文件
- MCP 服务入口:`src/summary_mcp/server.py`
- 主日报 workflow:`src/summary_mcp/workflows/freshrss_pipeline.py`
- run / artifact 查询:`src/summary_mcp/runtime/query_service.py`
- 主日报 async job:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
- resume 预检与恢复:`src/summary_mcp/runtime/resume_service.py`
- resume async job:`src/summary_mcp/runtime/resume_jobs.py`
- 单篇总结 async job:`src/summary_mcp/runtime/article_summary_jobs.py`
- 关键词治理:`src/summary_mcp/core/keyword_index.py`
- MCP 服务入口
- `src/summary_mcp/server.py`
- FreshRSS 统一工作流
- `src/summary_mcp/workflows/freshrss_pipeline.py`
- 词元统计核心
- `src/summary_mcp/core/keyword_index.py`
- 词元统计模型
- `src/summary_mcp/models/keyword_index.py`
- 摘要循环
- `src/summary_mcp/core/summary_loop.py`
- 提取主流程
- `src/summary_mcp/core/pipeline.py`
- FreshRSS 集成
- `src/summary_mcp/integrations/freshrss.py`
- 规则引擎
- `src/summary_mcp/filters/engine.py`
- LLM 结果校验
- `src/summary_mcp/validators/llm_result.py`
- OpenClaw candidate 模型
- `src/summary_mcp/models/article_candidate.py`
- OpenClaw delivery 模型
- `src/summary_mcp/models/openclaw_delivery.py`
- 生产脚本入口
- `scripts/run_freshrss_pipeline.py`
- 词元统计重建脚本
- `scripts/build_keyword_index.py`
- 词元清洗 skill
- `skills/keyword-cleanup-review/SKILL.md`
- skill review bundle 脚本
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- 采纳建议写回脚本
- `scripts/apply_term_suggestions.py`
- 清洗治理配置
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
- OpenClaw 交接说明
- `docs/openclaw/openclaw-handoff.md`
## 当前核心产物
## 当前输出规则
主日报稳定产物:
默认生产模式只输出:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
- `raw/freshrss.raw.json`
- `candidates/openclaw-delivery-payload.json`
- `run-report.json`
job 状态目录:
同时会更新本地运行数据:
- `outputs/freshrss/pipeline_jobs/<job_id>/`
- `outputs/freshrss/resume_jobs/<job_id>/`
- `outputs/freshrss/article_summary_jobs/<job_id>/`
关键词运行数据:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
如果需要词元清洗审阅输入,可额外生成:
## 当前已验证
- `outputs/term_index/review/keyword-cleanup-bundle.json`
- FreshRSS 未读拉取与已读回写可用
- RSS-first 提取策略可用
- 主日报 MCP 主链路可触发并写出正式产物
- 状态查询与结果读取接口可用
- stale state / artifacts 收敛逻辑已落地
- `resume` 的 artifact-first 判定已落地
- `start_resume_job -> poll -> result` 已做本地 synthetic 验证
- 单篇总结 async job 可跑通
- 关键词 review bundle 与建议写回脚本可用
如果需要在人工确认后把建议正式写入 watchlist / change log,可使用:
## 当前主要限制
- `scripts/apply_term_suggestions.py`
- 某些源 RSS 正文不足时会被直接跳过
- 规则仍然偏保守,部分内容会落到 `review`
- `paywall` 启发式对中文仍可能误判
- 关键词治理还没有接入周期性调度
- `digest-brief.json` 仍没有独立 MCP 读取工具
- `resume` 目前的剩余主风险不再是恢复点判定,而是缺少真实生产环境的完整恢复验证
如果需要排障,可开启:
## 当前建议
- `debug_artifacts=true`
- 或脚本参数 `--debug-artifacts`
这样才会额外输出逐条中间文件。
## 当前验证状态
已经验证通过:
- FreshRSS 未读拉取成功
- 已读回写成功
- MCP 工具入口可直接触发完整链路
- 微信公众号样本可直接使用 RSS 提供的 `summary` 内容提取,不再回源抓网页
- 精简输出模式已实际跑通
- 日报级词元统计已通过离线样例验证,确认别名、停用词、非 `drop` 过滤和 rerun 覆盖逻辑正常
- `keyword-cleanup-review` skill 已通过 `quick_validate.py` 结构校验
- review bundle 脚本已实际跑通
- `apply_term_suggestions.py` 已通过 dry-run 与临时副本写回验证
## 当前已知限制
- 当前对 FreshRSS 条目采用 RSS-first 策略,不再回源抓原网页
- 如果 RSS 中没有足够正文内容,该条会直接跳过,不会进入后续总结
- 某些规则仍偏保守,部分内容可能落到 `review`
- `paywall` 相关启发式仍可能误判中文文本
- Webhook / 主动投递到 OpenClaw 外部接口尚未实现,当前是由 OpenClaw 通过 MCP 主动调用
- 词元清洗 skill 当前已支持“bundle 构建 -> 建议审阅 -> 人工确认写回 watchlist/change_log”,但尚未接入周期性调度
- 当前词元统计仍以前置 `OpenClawDeliveryPayload` 作为日报前代理输入,真实 `DailyDigest` 接入后还需切换上游
## 当前最建议的交接阅读顺序
1. `README.md`
2. `docs/openclaw/openclaw-handoff.md`
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
5. `docs/design/daily-keyword-index-design.md`
6. `skills/keyword-cleanup-review/SKILL.md`
7. `TODO.md`
## 一句话结论
当前仓库已经从“提取 MCP 原型”演进到“可供 OpenClaw 调用的 FreshRSS -> OpenClaw payload 上游处理器”,并已补上第一阶段的日报级词元统计能力和词元清洗 skill 骨架;后续重点转向 skill 周期调度、知识库状态流转和 webhook 接线。
- 把 `docs/openclaw/openclaw-orchestration-flow.md` 当成正式编排手册
- 把 `docs/openclaw/openclaw-handoff.md` 当成接手总览
- 把 `plans/README.md` 当成规划文档导航
- 把 `docs/openclaw/archive/` 和 `docs/archive/` 当成历史资料,不要当当前事实源
+56
View File
@@ -0,0 +1,56 @@
# 设计文档导航
## 使用原则
`docs/design/` 目录同时包含两类文档:
- 当前实现仍然有效的设计说明
- 早期架构草案和背景设计
不要默认把这里所有文档都当成当前生产事实。
当前生产事实仍以这些入口为准:
1. `README.md`
2. `docs/current/context-reset-brief.md`
3. `docs/openclaw/README.md`
4. `docs/openclaw/openclaw-handoff.md`
5. `docs/openclaw/openclaw-orchestration-flow.md`
6. `TODO.md`
## 当前实现仍然有效
- `filter-rule-engine-design.md`
- 规则过滤层的设计与职责边界
- `filter-rule-engine-usage.md`
- 规则引擎的使用说明
- `daily-keyword-index-design.md`
- 关键词索引与清洗治理设计
- `summary-loop-explained.md`
- LLM 摘要校验闭环说明
- `markdown-sink-design.md`
- Markdown sink 设计
## 当前仍有参考价值,但不是生产真相入口
- `summary-mcp-service-design.md`
- 早期 MCP 服务设计草案,部分定位已被后续 workflow service 演进覆盖
- `source-schema-design.md`
- 更偏对象建模和来源抽象的背景设计
- `summary-core-interface-design.md`
- 更偏早期摘要内核接口抽象
## 建议阅读顺序
如果你是在理解当前实现:
1. `filter-rule-engine-design.md`
2. `filter-rule-engine-usage.md`
3. `daily-keyword-index-design.md`
4. `summary-loop-explained.md`
5. `markdown-sink-design.md`
如果你是在回看背景设计:
1. `summary-mcp-service-design.md`
2. `source-schema-design.md`
3. `summary-core-interface-design.md`
+17 -4
View File
@@ -326,8 +326,13 @@ LLM 可以帮助做清洗建议,但不适合直接维护主词元库。
skill 不直接修改配置文件,而是生成建议文件,例如:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
其中建议语义为:
- JSON 是 review / apply 之间的唯一正式建议产物
- Markdown 是人工临时审阅展示稿,不是长期真相来源
低复杂治理层建议补充三类输入:
@@ -403,8 +408,9 @@ skill 不直接修改配置文件,而是生成建议文件,例如:
2. 程序更新 `daily/YYYY-MM-DD.json`
3. 程序更新 `term_stats.json`
4. 每周或人工触发一次词元清洗 skill
5. skill 输出建议
6. 人工确认后再更新配置文件
5. skill 生成 suggestions JSON(正式建议产物)
6. 如需要人工阅读,再临时生成 Markdown 展示稿
7. 人工确认后再更新配置文件
## 14. 与规则引擎的关系
@@ -445,7 +451,8 @@ skill 不直接修改配置文件,而是生成建议文件,例如:
再补治理层:
- 增加词元清洗 skill
- 输出建议文件
- 输出建议文件(以 JSON 为正式产物)
- Markdown 仅作为按需生成的人工展示层
- 人工确认后更新配置
### Phase 3
@@ -459,3 +466,9 @@ skill 不直接修改配置文件,而是生成建议文件,例如:
## 16. 一句话结论
这套设计选择“只统计日报中的 `keywords`,由程序维护轻量词元库,再由独立 skill 周期性做清洗建议”,目的是在控制数据规模的前提下,为规则配置和长期兴趣演化提供稳定、可审计、可扩展的基础设施。
补充的产物策略是:
- facts/state 长期保留
- suggestions JSON 作为正式建议产物短期保留
- review bundle 与 Markdown 展示稿降级为临时工作文件 / 展示层
@@ -0,0 +1,151 @@
# 关键词清洗流程概述
> 2026-05-14 初版
> 从"数据记录"到"人工确认落盘"的完整链路
---
## 整体数据流
```
每日日报 pipeline
│
▼
term_index/daily/YYYY-MM-DD.json ← 每天一篇候选文章的热词统计
│
▼
term_index/term_stats.json ← 所有 daily 的汇总(1070 个词)
│
├──── build_review_bundle.py ← 打包为审查数据包
│ │
│ ▼
│ review/keyword-cleanup-bundle.json
│ │
│ ▼
│ generate_term_cleanup_suggestions.py
│ │
│ ▼
│ review/term-cleanup-suggestions-YYYY-MM-DD.json ← 正式建议产物
│ │
│ ▼
│ (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md ← 展示稿
│
├──── 人工确认哪些建议 accept
│
▼
apply_term_suggestions.py ← 写入配置
│
├── configs/filter_context.personal.json ← interest_keywords
├── configs/term_aliases.json ← alias
├── configs/term_stopwords.json ← stopword
├── configs/term_watchlist.json ← watch
└── configs/term_change_log.json ← 变更日志
```
---
## 各环节说明
### 阶段 1:数据记录(每日自动)
```bash
# FreshRSS pipeline 跑完后自动产出
data/term_index/daily/2026-05-14.json
```
- 每天一篇,记录当天候选文章中出现的热词
- 包含 term、total_count、days_seen 等信息
- 目前累计 **41 天**,共 **1070 个独立词**
### 阶段 2:全量汇总(每日自动)
```bash
data/term_index/term_stats.json
```
- 从所有 daily 文件重建,会覆盖重跑
- 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13)
### 阶段 3:构建审查数据包(手动触发)
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 \
--top 100 \
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
- 把 term_stats + 当前配置打成一包,方便后续处理
- 输出:`review/keyword-cleanup-bundle.json`
### 阶段 4:生成建议(手动触发)
```bash
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
#### 当前产出能力
| 建议类型 | 状态 | 当前阈值 | 说明 |
|---------|------|----------|------|
| interest_keyword_suggestions | ✅ **已实现** | total≥3, days≥2 | 产出 20 条 |
| watch_terms | ✅ **已实现** | total≤2, days≤2 | 本次 0 条 |
| alias_suggestions | ❌ **硬编码为空** | policy 有阈值(total≥2, days≥2)但脚本未实现 | |
| stopword_suggestions | ❌ **硬编码为空** | policy 有阈值(total≤2, days≤2)但脚本未实现 | |
**关键发现:** alias 和 stopword 不是"阈值太保守",是 **generate 脚本里压根没写对应的生成函数**。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 `[]` 和 `0`。
### 阶段 5:人工确认(手动)
```
OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply
```
本次模式:
- 高频(≥5次/5天以上)→ 强烈推荐 ✅
- 中频(3-4次)→ 附带建议 ✅
- 泛词 → 建议跳过 ❌
### 阶段 6:落盘配置(手动)
```bash
python scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 词1 词2 ...
```
- dry-run 预览 → 确认后正式 apply
- 写入 `configs/filter_context.personal.json`
- 同步记录到 `term_change_log.json`
- **不备份原始配置**(待优化)
- **apply 后不自动清理 review 目录**(待优化)
### 阶段 7:维护清理(按需)
由 OpenClaw 侧 `reader-keyword-maintenance` skill 处理:
- 删除旧 markdown 展示稿
- 保留最近一份 bundle
- 保守保留 suggestions JSON
---
## 当前配置资产
| 文件 | 内容 | 数据量 |
|------|------|--------|
| `filter_context.personal.json` | interest_keywords | 52 个 |
| `term_aliases.json` | 别名映射 | 0 组(未启用) |
| `term_stopwords.json` | 停用词 | 0 个(未启用) |
| `term_watchlist.json` | 观察词 | 6 个 |
| `term_change_log.json` | 所有变更记录 | 已记录 |
---
## 待优化项
1. **alias/stopword 建议生成为空** — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数
2. **apply 前无配置备份** — 建议 apply 前自动 cp 备份
3. **apply 后无自动收尾** — 建议 apply 后自动删旧 markdown 和 bundle
4. **alias 识别依赖规则而非 LLM** — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply
+11
View File
@@ -1,5 +1,16 @@
# Source Schema 设计草案
## 状态说明
本文件偏对象建模和来源抽象,主要用于解释早期 schema 设计思路。
它不是当前生产运行手册,也不是当前 workflow service 的唯一真相来源。
如果你关注当前 OpenClaw 集成或运行状态,应优先看:
- `docs/current/context-reset-brief.md`
- `docs/openclaw/README.md`
- `docs/openclaw/openclaw-orchestration-flow.md`
## 1. 文档目的
本文档用于定义阅读流系统中的来源与内容对象模型,目标是把“来源分类”的讨论收敛成一套可执行的数据结构,供后续的抓取、摘要、过滤、入库和推送流程统一使用。
@@ -1,5 +1,17 @@
# Summary Core Interface 设计草案
## 状态说明
本文件记录的是较早期的摘要内核接口抽象。
它更适合用于理解背景设计,不应直接当成当前生产接口契约。
当前接口与编排真相请优先看:
- `README.md`
- `docs/current/context-reset-brief.md`
- `docs/openclaw/README.md`
- `docs/openclaw/openclaw-handoff.md`
## 1. 文档目的
本文档用于定义 `summary-core` 的输入输出接口,目标是把“页面摘要能力”从概念讨论收敛成一套稳定、可复用、可封装的数据接口。
+14 -1
View File
@@ -1,5 +1,18 @@
# Content Extract MCP Service 设计草案
## 状态说明
本文件主要记录早期 “content extract MCP” 的设计抽象。
它仍有背景参考价值,但不是当前生产事实入口。
当前生产能力已经演进为更完整的 workflow service,正式口径请优先看:
- `README.md`
- `docs/current/context-reset-brief.md`
- `docs/openclaw/README.md`
- `docs/openclaw/openclaw-handoff.md`
- `docs/openclaw/openclaw-orchestration-flow.md`
## 1. 文档目的
本文档用于定义当前仓库中已经落地的 MCP 服务设计,即“内容提取 MCP”。
@@ -12,7 +25,7 @@
- validator 与 LLM 摘要如何接在 MCP 之后
- 当前 MVP 已完成到哪一层
这份文档描述的是当前真实实现,而不是早期“摘要 MCP”设想。
这份文档主要记录当时实现阶段的设计取向,而不是当前生产阶段的唯一事实来源。
---
+22
View File
@@ -0,0 +1,22 @@
# Notes 导航
`docs/notes/` 保存的是更早期、讨论型、背景型方案笔记。
这些文档的用途是:
- 理解项目最初的问题空间
- 回看为什么会形成现在的对象分层和流程划分
这些文档不是当前生产事实来源。
当前如需判断“现在到底怎么跑”,优先看:
- `README.md`
- `docs/current/context-reset-brief.md`
- `docs/openclaw/README.md`
- `docs/openclaw/openclaw-orchestration-flow.md`
当前 notes:
- `reading-pipeline-design-notes.md`
- 早期阅读流方案讨论纪要
+36
View File
@@ -0,0 +1,36 @@
# OpenClaw 文档导航
## 当前有效文档
- `openclaw-handoff.md`
- 面向接手者的总览文档
- 说明 reader 的职责边界、MCP 工具面、环境变量和正式集成约束
- `openclaw-orchestration-flow.md`
- 面向 OpenClaw 编排层的正式运行手册
- 说明启动、轮询、读结果、恢复和人工介入的标准动作
- `openclaw-candidate-input-field-spec.md`
- 单篇 `OpenClawCandidateInput` 字段规范
- `openclaw-delivery-payload-spec.md`
- 批量 `OpenClawDeliveryPayload` 字段规范
- `article-candidate-daily-digest-schema.md`
- 对象分层设计说明
- 用于理解 `ArticleCandidateRecord` / `OpenClawCandidateInput` / `DailyDigest` 的关系
## 当前推荐阅读顺序
1. `openclaw-handoff.md`
2. `openclaw-orchestration-flow.md`
3. `openclaw-candidate-input-field-spec.md`
4. `openclaw-delivery-payload-spec.md`
5. `article-candidate-daily-digest-schema.md`
## 归档说明
`archive/` 下的文档保留历史决策、阶段总结和排障规划,但不再作为当前事实来源。
当前已归档:
- `archive/formalization-summary-2026-04-07.md`
- `archive/openclaw-daily-digest-refactor.md`
- `archive/digest-optimization-summary.md`
- `archive/p1-status-reconciliation-plan-2026-04-14.md`
+9
View File
@@ -0,0 +1,9 @@
# OpenClaw 历史归档
本目录只保留阶段性总结、设计演进记录和排障计划。
使用原则:
- 需要了解“为什么会这样设计”时再看
- 不要把这里的描述当成当前生产事实
- 当前正式口径以 `docs/openclaw/README.md`、`docs/openclaw/openclaw-handoff.md`、`docs/openclaw/openclaw-orchestration-flow.md` 为准
@@ -0,0 +1,163 @@
# reader MCP workflow service formalization summary (2026-04-07)
## Overview
On 2026-04-07, the reader project was formally advanced from a script-first integration model into a reader-centric MCP workflow service model.
The key shift is:
- before: OpenClaw primarily relied on long CLI / exec flows and direct output-path stitching
- now: reader exposes a formal workflow-oriented MCP surface with run-state, status queries, result reads, and minimal recovery
This document records the main outcomes and commits for the first formalization phase.
---
## Completed capability set
### 1. Run-state persistence
Commit:
- `72a6853` — `Add run-state persistence for FreshRSS pipeline`
Delivered:
- `run-state.json`
- `RunState / StageState / ArtifactRecord`
- stage-level state persistence for the FreshRSS pipeline
### 2. Architecture / implementation docs
Commit:
- `7563aa8` — `docs: add reader MCP architecture and implementation plan`
Delivered:
- architecture design
- implementation plan
- TODO-driven collaboration model
### 3. MCP run-status query tools
Commit:
- `d9173fb` — `feat: add MCP run status query tools`
Delivered:
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
### 4. MCP result-read tools
Commit:
- `4a02894` — `Add MCP delivery payload and run report queries`
Delivered:
- `get_delivery_payload`
- `get_run_report`
### 5. Minimal resume design
Commit:
- `d91cdbc` — `docs: narrow resume_run minimal recovery design`
Delivered:
- narrowed design for `resume_run`
- explicit supported / unsupported recovery points
### 6. Minimal `resume_run`
Commit:
- `c622bc6` — `Implement minimal resume_run for freshrss runs`
Delivered:
- minimal `resume_run`
- supports only freshrss runs with `run-state.json`
- supports only recent resumable points
- explicitly rejects `fetch_feed` and `extract_articles`
### 7. Formal handoff / workflow docs
Commits:
- `4f219ef` — `docs: formalize reader MCP workflow service handoff`
- `2df0af5` — `docs: add openclaw orchestration flow for reader MCP`
Delivered:
- formal handoff aligned to actual implementation
- OpenClaw orchestration runbook
- explicit rule that OpenClaw should stop hand-stitching reader paths in the normal production flow
---
## Current formal MCP workflow surface
The current reader MCP workflow surface now includes:
- `run_freshrss_openclaw_pipeline`
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
- `resume_run` (minimal version)
---
## Current boundary
reader is now the upstream workflow engine for:
- FreshRSS pull
- extraction
- summary
- filter
- payload generation
- run-state persistence
- result read
- minimal recovery
OpenClaw / skill remains responsible for:
- digest markdown generation
- Hugo publishing
- chat reporting
- user confirmation
- IMA orchestration
---
## Current limitations
The first formalization phase is complete, but some constraints remain:
- `resume_run` is still minimal and does not support arbitrary stage re-entry
- historical runs without `run-state.json` are not formally recoverable
- some very old runs may still require conservative artifact/path discovery
- `rerun_stage` is not implemented
- deeper runtime consolidation of `run_freshrss_openclaw_pipeline` can still be improved later
---
## Practical conclusion
The reader project should now be treated as a formal MCP workflow service rather than as a long-running CLI-first integration point.
For normal production orchestration:
- start via MCP
- observe via MCP status tools
- read results via MCP result tools
- use `resume_run` only within the documented minimal recovery range
- keep CLI for debug / fallback only
@@ -0,0 +1,367 @@
# Reader 日报链路 P1 状态收敛问题:规划与修复清单(2026-04-14)
## 背景
在 2026-04-14 的 reader 日报正式运行中,出现了以下现象:
- `openclaw-delivery-payload.json`、`digest-brief.json`、`run-report.json` 已真实落盘
- 但 `get_freshrss_pipeline_job_status` / `get_run_status` 仍可能显示:
- `running`
- `failed`
- 或 `current_stage=generate_summaries`
- `resume_run` 在这种状态下可能直接超时
这说明当前 reader 的**状态层(job/run-state)**与**产物层(artifacts/report)**之间没有稳定收敛。
---
## 本次确认的核心结论
### 1. job status 与 run status 是两套独立状态系统
- **job 层状态**:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
- `start_freshrss_pipeline_job()`
- `run_freshrss_pipeline_job()`
- `get_freshrss_pipeline_job_status()`
- 状态文件位于:`outputs/freshrss/pipeline_jobs/<job_id>/run-state.json`
- 只有 4 个粗粒度 stage:
- `prepare_job`
- `load_input`
- `run_pipeline`
- `write_result`
- **run 层状态**:`src/summary_mcp/workflows/freshrss_pipeline.py`
- `run_freshrss_pipeline()`
- 由 `src/summary_mcp/runtime/query_service.py:get_run_status()` 查询
- 状态文件位于:`outputs/freshrss/rerun/<run_dir>/run-state.json`
- 包含 6 个细粒度 stage:
- `fetch_feed`
- `extract_articles`
- `generate_summaries`
- `apply_filters`
- `build_delivery_payload`
- `write_run_report`
**问题:** 两套状态没有统一收敛规则,用户可以同时看到两套不同口径的“当前进度”。
---
### 2. 查询层目前优先信 run-state,不会用 artifacts / run-report 纠偏
代码位置:`src/summary_mcp/runtime/query_service.py`
关键行为:
- `_resolve_run_record()` 只要发现 `run-state.json` 存在,就优先使用 `RunStore.load(...)`
- 即使 `run-report.json`、`delivery_payload`、`digest_brief` 已存在,也不会自动纠偏状态
**结果:**
- 一旦 `run-state.json` 因中断、超时、外层 SIGTERM 或写回未完成而停留在旧值
- `get_run_status()` 就会持续返回过期状态
- 造成“产物已完成,但状态仍显示 running/failed/卡在 summary”的错觉
---
### 3. `generate_summaries` 假卡住,本质上更像 stale state,不像真实业务卡住
代码位置:`src/summary_mcp/workflows/freshrss_pipeline.py`
从执行顺序看:
1. `start_stage(generate_summaries)`
2. summary 循环
3. `finish_stage(generate_summaries)`
4. `start_stage(apply_filters)`
5. `finish_stage(apply_filters)`
6. `start_stage(build_delivery_payload)`
7. 写 payload / digest brief
8. `finish_stage(build_delivery_payload)`
9. `start_stage(write_run_report)`
10. 写 run-report
11. `finish_stage(write_run_report)`
12. `finish_run(...)`
**判断:**
如果 payload / digest brief / run-report 都已经存在,那么“仍显示卡在 `generate_summaries`”更可能是:
- `run-state.json` 没来得及写回最终状态
- 或查询时读到了旧状态
而不是 summary 阶段真实没有跑过去。
---
### 4. `resume_run` 不是轻量恢复,而是同步继续跑工作流
代码位置:`src/summary_mcp/runtime/resume_service.py`
关键行为:
- `resume_run()` 会根据 `resume_from_stage` 直接继续执行:
- `_run_summary_stage(...)`
- `_run_filter_stage(...)`
- `_run_delivery_stage(...)`
- `_run_report_stage(...)`
这意味着它不是“修状态”的工具,而是“同步继续跑剩余工作流”的工具。
**问题:**
- 如果 stale state 把 `resume_from_stage` 定在 `generate_summaries`
- 那么 `resume_run` 会从一个过早阶段重新跑
- 在 MCP 包装层下非常容易超时
---
## 问题分类
### A. 真实 bug
1. **查询层过度信任 stale `run-state.json`**
- 文件:`src/summary_mcp/runtime/query_service.py`
- 影响:产物已完成但状态仍错误
2. **`resume_run` 过度依赖 stale `current_stage` / recovery 信息**
- 文件:`src/summary_mcp/runtime/resume_service.py`
- 影响:从过早阶段重跑,放大 timeout 风险
### B. 状态设计缺陷
3. **job 层与 run 层两套状态源没有统一收敛规则**
- 文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
- 文件:`src/summary_mcp/runtime/query_service.py`
- 影响:用户看到两个互相打架的状态解释
4. **状态系统完全依赖显式写回,不会按产物反推修正**
- 文件:`src/summary_mcp/runtime/run_store.py`
- 影响:一旦中断,状态比产物更容易脏
### C. 调用层误判
5. **把 `resume_run` 当成轻量恢复接口使用**
- 实际上它更接近“同步恢复执行器”
- 影响:在长链路场景下超时是高概率事件
---
## 修复目标
## 当前落地状态(回填)
- [x] Phase 1 已落地:`get_run_status()` 会基于 `run-report.json` 与关键产物做终态收敛,并暴露 `status_source` / `state_conflict`
- [x] Phase 2 已落地第一阶段:`resume_run()` 会拒绝对已有终态 `run-report.json` 的 run 继续恢复
- [x] Phase 2 已继续增强:恢复起点现在会优先根据 artifacts 重算,而不是直接盲信 `run-state.recovery.resume_from_stage`
- [x] 新增 `inspect_resume_plan(run_id)` 作为恢复前置判定接口,避免调用方用 `resume_run` 探路
- [x] Phase 2 已补齐生产恢复 artifacts:正式 run 会稳定写出 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`,`resume_run` / `inspect_resume_plan` 会优先使用它们,而不是依赖 debug per-item 文件
- [x] Phase 3 已落地:job 状态与结果读取会基于 linked run 做收敛,避免 outer job stale state 卡住编排
### 一级目标(必须达成)
1. 当 `run-report.json` / `delivery_payload` / `digest_brief` 已存在时,`get_run_status()` 不应继续盲目展示明显过期的 stage 状态;对调用方暴露的 `status` 必须直接收敛为可用终态,而不是只附加 hint
2. 当状态层与产物层冲突时,查询结果必须显式标注“状态冲突 / stale state”
3. `resume_run()` 在恢复前应优先基于现有 artifacts 判断真实可恢复起点,避免从过早阶段重跑
### 二级目标(建议达成)
4. job 层状态结果中增加对 linked run 的补充解释,避免“job running 但 run 产物已齐”这种情况毫无说明
5. 为后续编排层提供明确可消费的“状态可信度/冲突提示”字段
---
## 最小修复方案
### Phase 1|先修 run 查询层(优先级最高)
#### 目标
让 `get_run_status()` 至少能正确识别:
- run-state 是旧的
- 但关键产物已经齐了
#### 建议改动点
文件:`src/summary_mcp/runtime/query_service.py`
#### 建议动作
- [x] 在 `_resolve_run_record()` 或 `_build_status_response()` 中增加“关键产物存在性检查”
- `run-report.json`
- `candidates/openclaw-delivery-payload.json`
- `candidates/digest-brief.json`
- [x] 如果 `run-state.current_stage` 仍停留在早期阶段,但关键产物已齐:
- 不要继续原样输出为可信最终态
- 应直接把对外 `status` / `current_stage` / `recovery` 收敛成终态语义
- 同时新增解释字段,例如:
- `state_conflict: true`
- `state_conflict_reason: "run_state indicates generate_summaries but run-report.json already proves the workflow reached a terminal state"`
- `status_source: "run_report_reconciliation"`
- [x] 保留 `state_source=run_state`,但增加 `status_source` / `state_quality` / `state_conflict` 之类解释字段
#### 预期收益
- OpenClaw 继续按 `status` 分支时也不会卡住
- 第一时间减少“明明产物齐了却还像没跑完”的误判
- 不需要立刻动 workflow 主链路
---
### Phase 2|修 `resume_run` 的恢复起点判断
#### 目标
避免 stale state 让恢复逻辑从 `generate_summaries` 这类过早阶段重跑。
#### 建议改动点
文件:`src/summary_mcp/runtime/resume_service.py`
#### 建议动作
- [x] 在 `_resolve_resume_from_stage()` 之前/之后加入真实 artifacts 检查
- [x] 如果以下文件已存在:
- `openclaw-delivery-payload.json`
- `digest-brief.json`
- `run-report.json`
则不要再从 `generate_summaries` 或 `apply_filters` 起跑
- [x] 为 `resume_run()` 增加“恢复起点是基于 artifacts 重算还是基于 state 推断”的返回说明
- [x] 必要时增加更保守逻辑:
- `run-report.json` 已存在时,默认拒绝继续 resume,并提示“产物已完成,请先检查状态一致性”
- 补充:默认生产模式下,主链路会稳定写出 `summary-batch` / `candidate-batch`,恢复逻辑优先消费这两个 batch artifacts;若它们缺失或不稳定,才回退到更早的安全 stage 或直接拒绝恢复
- 补充:调用方可先走 `inspect_resume_plan`,只有 `recommended_action=resume` 时再调用 `resume_run`
#### 预期收益
- 降低无意义重跑和 timeout 风险
- 让 `resume_run` 更接近真正的恢复工具,而不是误重跑工具
---
### Phase 3|补 job/run 双状态解释层
#### 目标
让 `get_freshrss_pipeline_job_status()` 和 `get_run_status()` 的关系对调用方更可理解。
#### 建议改动点
文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
#### 建议动作
- [x] 在 `get_freshrss_pipeline_job_status()` 中,读取 linked run 的关键产物存在性(轻量即可)
- [x] 若 job 仍显示 `run_pipeline`,但 linked run 已有 report/payload/digest 产物:
- 不仅增加解释字段,还应直接把 job 对外 `status` 收敛为终态,避免外层永远轮询
- 例如:
- `status_source: "linked_run_reconciliation"`
- `status_note: "linked run artifacts are complete; the job can be treated as completed"`
- [x] 若 `result.json` 缺失,但 linked run 已有 `run-report.json` 与 delivery 产物:
- `get_freshrss_pipeline_job_result()` 应能基于 linked run 产物合成最小结果,至少稳定返回 `run_id`
- [x] 明确文档:job status 是外层异步任务态,不等于内部 workflow 细粒度状态
#### 预期收益
- 减少“job running / run finished”口径冲突带来的误解
- 避免 OpenClaw 因 outer job stale state 卡死在轮询和 result 读取前
---
### Phase 4|把 `resume_run` 改成异步恢复 job
#### 目标
解决当前剩余的核心问题:`resume_run` 虽然恢复判定已经安全,但执行模型仍是同步 MCP 调用,长链路恢复时依然可能超时,导致 OpenClaw 编排层“看起来像又卡住了”。
#### 建议改动点
文件:
- `src/summary_mcp/runtime/resume_jobs.py`(新)
- `scripts/run_resume_job.py`(新)
- `src/summary_mcp/server.py`
- `src/summary_mcp/runtime/__init__.py`
- `src/summary_mcp/runtime/resume_service.py`
#### 建议动作
- [x] 新增最小异步恢复接口:
- `start_resume_job(run_id)`
- `get_resume_job_status(job_id)`
- `get_resume_job_result(job_id)`
- [x] job 目录固定落到:
- `outputs/freshrss/resume_jobs/<job_id>/`
- [x] 最少产物约定:
- `run-state.json`
- `input.json`
- `result.json`(成功时)
- `job-report.json`
- [x] `start_resume_job` 内部先调用 `inspect_resume_plan`
- 只有 `recommended_action=resume` 才允许真正启动
- `read_terminal_result` / `start_new_run` 要直接在 job 输入校验阶段返回,不进入执行器
- [x] 后台执行时复用现有 `_resume_freshrss_run(...)`
- 不重写恢复业务逻辑
- 只把同步入口拆成异步 job 外壳
- [x] `resume_run(run_id)` 保留,但降级为 debug / fallback
- 文档中明确:OpenClaw 编排默认应走 resume async job,而不是同步 `resume_run`
- [x] job result 里至少稳定返回:
- `run_id`
- `resume_from_stage`
- `status`
- `result_source`
- `delivery_output` / `report_output`(若存在)
#### 预期收益
- 彻底切掉恢复阶段的 MCP 同步超时风险
- 让 OpenClaw 对“启动恢复 / 轮询恢复 / 读取恢复结果”的控制面与主 pipeline async job 保持一致
- 把“恢复判定”与“恢复执行”分层,减少误调用和卡住错觉
---
## 不建议现在就做的事
- [ ] **不要先做自动 fallback 修状态**
- 例如:看到 artifacts 齐了就直接把 run-state 强行改成 success
- 原因:这会掩盖真正的状态写回问题
- [ ] **不要先大改 workflow 主链路**
- 当前更像查询层与恢复层的状态解释缺陷
- 先修读取与恢复判断,收益更大、风险更低
---
## 建议执行顺序
1. **先改 `query_service.py`**
- 让 `get_run_status()` 能暴露 stale state / artifact conflict
2. **再改 `resume_service.py`**
- 避免从错误阶段重跑
3. **最后看 `freshrss_pipeline_jobs.py`**
- 给 job status 加 linked run 补充说明
4. **收尾改 `resume async job`**
- 让恢复执行也走正式异步控制面,避免同步恢复再把编排卡住
---
## 验收标准
### 验收 1:状态冲突识别
构造一个场景:
- `run-state.json` 留在 `generate_summaries`
- 但 payload / digest brief / run-report 已存在
期望:
- `get_run_status()` 不再只回“卡在 generate_summaries”
- 会显式返回冲突提示字段
### 验收 2:恢复起点修正
构造一个场景:
- `run-state` 指向 `generate_summaries`
- 但 `delivery_payload` / `run-report` 已存在
期望:
- `resume_run()` 不应再从 summary 阶段重跑
- 至少应拒绝恢复并提示“产物已完成,优先检查状态一致性”
### 验收 3:job/run 双层说明
构造一个场景:
- job status 仍在 `run_pipeline`
- linked run 已有关键产物
期望:
- `get_freshrss_pipeline_job_status()` 能返回补充说明,不再只有生硬 running
### 验收 4:恢复执行不再阻塞编排
构造一个场景:
- run 可恢复
- 恢复点为 `generate_summaries` 或 `apply_filters`
- 恢复执行耗时超过单次 MCP 同步窗口
期望:
- OpenClaw 调用的是 `start_resume_job(...)`,而不是同步 `resume_run(...)`
- `get_resume_job_status(job_id)` 可稳定轮询到终态
- `get_resume_job_result(job_id)` 至少稳定返回 `run_id`、`resume_from_stage` 与最终产物引用
- 即使恢复失败,也能在 job-report / result 中看清失败点,而不是只表现为调用超时
---
## 备注
截至 2026-04-14,本文件中的 Phase 1 / 2 / 3 / 4 已完成主要落地;当前 `resume` 链路已经从“状态收敛 + 安全恢复点判定”进一步补齐到“正式异步恢复执行”。
+189 -243
View File
@@ -1,64 +1,196 @@
# OpenClaw Handoff
## Role
This file is the integration overview for OpenClaw maintainers.
Use it for:
- reader capability boundary
- production MCP entrypoints
- environment requirements
- integration rules and limitations
Do not use it as the step-by-step runbook.
For formal orchestration, read `docs/openclaw/openclaw-orchestration-flow.md`.
For field contracts, read:
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- `docs/openclaw/openclaw-delivery-payload-spec.md`
Historical plans and incident documents live under `docs/openclaw/archive/`.
## Purpose
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
reader is the upstream FreshRSS processing service for OpenClaw:
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
OpenClaw should treat this repository as an MCP-backed upstream content processor.
This repository is responsible only for upstream reading-pipeline work:
reader is responsible for:
- FreshRSS pull
- content extraction
- LLM summary generation/validation
- LLM summary generation and validation
- rule-based filtering
- OpenClaw delivery payload generation
- selected-article summary capability based on existing extracted text
- run-state persistence, status lookup, result lookup, and minimal resume for the FreshRSS workflow
- run-state persistence and run/result lookup
- async resume control for the FreshRSS workflow
- async selected-article summary generation from existing extracted files
This repository should **not** take over downstream orchestration responsibilities that belong to OpenClaw / skills, such as:
reader is not responsible for:
- Hugo publishing
- chat reporting
- user confirmation handling
- IMA upload orchestration
## Production Entrypoint
## Production Surface
reader 当前正式工作流服务入口是 MCP tool:
Current MCP tool count: 21.
- `run_freshrss_openclaw_pipeline`
Main daily workflow:
It starts the only formally supported workflow today: `freshrss_daily_digest`.
OpenClaw should treat the returned `run_id` as the only stable handle for follow-up reads. Do not hand-build `outputs/freshrss/rerun/...` paths in OpenClaw.
## Supported MCP Tools
Current MCP tools: 11 total.
Workflow service tools:
- `run_freshrss_openclaw_pipeline`
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
- `get_freshrss_pipeline_job_result`
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
Resume workflow:
- `inspect_resume_plan`
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
- `resume_run`
Single-step / debug tools:
Selected-article summary workflow:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
- `generate_article_summaries`
Debug / single-step tools:
- `run_freshrss_openclaw_pipeline`
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `generate_article_summaries`
## Required Environment Variables
Production rules:
The MCP server process must have these variables available:
- main production start path is `start_freshrss_pipeline_job`
- production resume path is `inspect_resume_plan -> start_resume_job -> get_resume_job_status -> get_resume_job_result`
- `run_freshrss_openclaw_pipeline` is sync debug / fallback only
- `resume_run` is sync debug / fallback only
- `generate_article_summaries` is sync debug / fallback only
## Production Contract
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads.
OpenClaw should not hand-build these paths:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
If filesystem access is needed for debugging, only consume paths returned by MCP:
- `output_dir`
- `artifact.path`
- `delivery_output`
- `report_output`
Top-level `status` is the only status field callers should branch on.
`status_source` and `state_conflict` are explanatory fields for reconciled status.
## Minimal Production Sequence
Daily workflow:
1. Call `start_freshrss_pipeline_job`
2. Poll `get_freshrss_pipeline_job_status`
3. On success, read `get_freshrss_pipeline_job_result`
4. Persist the returned `run_id`
5. Use `get_run_status`, `get_delivery_payload`, and `get_run_report` for follow-up reads
Resume workflow:
1. Call `inspect_resume_plan(run_id)`
2. Only if `can_resume=true` and `recommended_action=resume`, call `start_resume_job`
3. Poll `get_resume_job_status`
4. Read `get_resume_job_result`
Selected-article summary workflow:
1. Call `start_article_summary_job` with a real extracted file path and non-empty `selected_ids`
2. Poll `get_article_summary_job_status`
3. Read `get_article_summary_job_result`
## Capability Boundary
Formal workflow boundary:
- only workflow `freshrss_daily_digest`
- every current FreshRSS run writes `run-state.json`
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older runs without `run-state.json`
- resume requires a valid `run-state.json`; inferred historical runs are not resumable
Resume boundary:
- resume in place on the original `run_id`
- supported resume points:
- `generate_summaries`
- `apply_filters`
- `build_delivery_payload`
- `write_run_report`
- unsupported resume points:
- `fetch_feed`
- `extract_articles`
- production resume prefers:
- `summary/summary-batch.json`
- `candidates/candidate-batch.json`
- if required artifacts are missing, recovery should return non-resumable instead of silently falling back
Selected-article summary boundary:
- uses existing extracted files as input
- should not re-fetch original URLs
## Output Expectations
Main daily pipeline core artifacts:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
Async job state directories:
- main pipeline job: `outputs/freshrss/pipeline_jobs/<job_id>/`
- resume job: `outputs/freshrss/resume_jobs/<job_id>/`
- article-summary job: `outputs/freshrss/article_summary_jobs/<job_id>/`
Each job directory minimally contains:
- `run-state.json`
- `input.json`
- `result.json` on success
- `job-report.json`
## Environment And Startup
Required environment variables:
- `FRESHRSS_API_BASE_URL`
- `FRESHRSS_USERNAME`
@@ -67,52 +199,13 @@ The MCP server process must have these variables available:
- `LLM_API_KEY`
- `LLM_MODEL`
Example:
```powershell
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
```
## Server Startup
Install dependencies:
Startup:
```bash
pip install -e .
```
Start the MCP server:
```bash
summary-mcp
```
## Recommended MCP Workflow
Recommended production path:
1. Call `run_freshrss_openclaw_pipeline` and persist the returned `run_id`
2. Use `get_run_status(run_id)` as the authoritative run-state read for status, stage, artifacts, and recovery
3. Use `list_runs(...)` when OpenClaw needs recent-run discovery or high-level inspection
4. Use `list_run_artifacts(run_id)` when OpenClaw needs to inspect what this run actually produced
5. Use `get_delivery_payload(run_id)` and `get_run_report(run_id)` as the formal result-reading APIs
6. Use `resume_run(run_id)` only when the run falls inside the minimal supported resume scope
OpenClaw should not directly derive or hardcode:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
If filesystem access is needed for debugging, consume only paths returned by MCP such as `output_dir`, `artifact.path`, `delivery_output`, or `report_output`.
## Recommended MCP Call
Recommended production start call:
```json
@@ -126,198 +219,51 @@ Recommended production start call:
}
```
Recommended semantics:
## Data And Content Policy
- Use `mark_read=true` for normal production runs.
- Use `mark_read=false` only for debug, test, or validation runs.
- Keep `debug_artifacts=false` for routine production runs.
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
- Treat the returned `run_id` as the stable identifier for all follow-up MCP reads.
- If no real `openclaw-delivery-payload.json` was produced, OpenClaw should stop instead of generating a digest from placeholders or examples.
## Formal Capability Boundary
reader 当前正式 MCP workflow service 的边界如下:
- formal workflow: only `freshrss_daily_digest`
- run truth: every FreshRSS run writes `run-state.json`
- state query tools: `get_run_status`, `list_runs`, `list_run_artifacts`
- result read tools: `get_delivery_payload`, `get_run_report`
- `digest-brief.json` is generated and registered as an artifact, but there is no standalone `get_digest_brief` tool yet
- `run_freshrss_openclaw_pipeline` and `resume_run` are synchronous MCP calls today; there is no background queue / worker model yet
- `generate_article_summaries` is supported, but it is outside the formal `resume_run` scope and not part of the FreshRSS workflow-state model
Historical compatibility note:
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older run directories without `run-state.json`
- `resume_run` does **not** support those inferred historical runs; it requires a valid `run-state.json`
## What The Tool Returns
Primary return fields from `run_freshrss_openclaw_pipeline`:
- `run_id`
- `output_dir`
- `raw_output`
- `delivery_output`
- `report_output`
- `digest_brief_output`
- `pulled_count`
- `delivered_count`
- `marked_read_count`
- `status_counts`
- `delivery_payload`
- `keyword_index`
Optional:
- `items`
- returned only when `include_item_reports=true`
Follow-up structured reads should use MCP tools rather than re-reading these files directly.
## Minimal Output Files
By default the pipeline writes these core artifacts:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json` (one per item)
It also updates local runtime keyword data:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
Per-item extracted files live under `extracted/` and are always written.
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
## `resume_run` Minimal Scope
`resume_run` currently supports only the minimum resume contract:
- only runs with a valid `run-state.json`
- only workflow `freshrss_daily_digest`
- resume in place on the original `run_id`
- supported resume points: `generate_summaries`, `apply_filters`, `build_delivery_payload`, `write_run_report`
- unsupported resume points: `fetch_feed`, `extract_articles`
- if required artifacts are missing, the tool returns a non-resumable response instead of silently falling back to an earlier stage
Artifact expectations by resume point:
- `generate_summaries`: requires `raw/freshrss.raw.json` and `extracted/`
- `apply_filters`: requires the above plus per-item summary outputs
- `build_delivery_payload`: requires per-item candidate inputs consistent with filter-stage output
- `write_run_report`: requires `candidates/openclaw-delivery-payload.json`; if `mark_read=true`, raw input must still be present
## Payload Specs
OpenClaw payload field specs live here:
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- `docs/openclaw/openclaw-delivery-payload-spec.md`
## Read-State Semantics
The pipeline reads from FreshRSS unread items by default.
If `mark_read=true`:
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
- only successfully delivered items are marked as read
- failed or skipped items remain unread
## FreshRSS Content Policy
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
Behavior:
FreshRSS processing is RSS-first:
- use `item.raw_content` first
- if missing, use `item.raw_summary`
- if neither contains usable content, skip the item
- do not fetch the original webpage again for FreshRSS items
This is intentional.
Read-state policy:
## Keyword Cleanup Governance
- items are marked read only after successful delivery payload write
- only successfully delivered items are marked read
This repository also includes a lightweight keyword-governance flow for downstream review.
Downstream boundary:
Current pieces:
- runtime keyword stats
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
- governance config
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
- review bundle builder
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- accepted-suggestion writer
- `scripts/apply_term_suggestions.py`
Current status:
- OpenClaw can read the keyword review bundle as maintenance input
- accepted suggestions still require explicit human confirmation
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
- this governance flow is not yet wired into a periodic scheduler inside the repository
Boundary:
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
- the repository does not auto-apply cleanup suggestions without confirmation
- current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest`
## Known Limitations
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
- Paywall heuristics may produce false positives for some Chinese text patterns.
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
## Files OpenClaw Should Read First
Recommended reading order for a new maintainer:
1. `README.md`
2. `docs/openclaw/openclaw-handoff.md`
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
5. `docs/design/daily-keyword-index-design.md`
6. `skills/keyword-cleanup-review/SKILL.md`
7. `docs/current/context-reset-brief.md`
## Downstream Boundary Rules
For the daily-digest workflow:
- the digest should go to Hugo and chat reporting, not directly into IMA
- the full daily digest should **not** be uploaded to IMA
- the daily digest goes to Hugo and chat reporting
- the full daily digest should not be uploaded to IMA
- only explicitly user-selected article summaries should be uploaded to IMA
- selected-article summaries should be generated from existing extracted text, not by re-fetching original URLs
## Current Recommendation
## Related Maintenance Flow
For integration handoff, the repository is usable now.
Keyword cleanup exists as a separate maintenance flow, not the main RSS ingestion path.
The minimum you need to give OpenClaw is:
- the repository code
- the MCP server startup command
- the required environment variables in the target environment
- the instruction to call `run_freshrss_openclaw_pipeline`
- the rule that follow-up state/result reads must go through MCP tools first, not handwritten filesystem paths
If OpenClaw will also participate in keyword-governance review, additionally point it to:
Relevant files:
- `docs/design/daily-keyword-index-design.md`
- `skills/keyword-cleanup-review/SKILL.md`
- `scripts/apply_term_suggestions.py`
## Known Limitations
- some sources expose only partial RSS content; those items may be skipped
- rule behavior is still conservative; many items may land in `review`
- paywall heuristics may still produce false positives on some Chinese text
- keyword cleanup governance is usable but not yet wired to periodic scheduling
## Read First
Recommended reading order for a new maintainer:
1. `README.md`
2. `docs/openclaw/README.md`
3. `docs/openclaw/openclaw-handoff.md`
4. `docs/openclaw/openclaw-orchestration-flow.md`
5. `docs/openclaw/openclaw-candidate-input-field-spec.md`
6. `docs/openclaw/openclaw-delivery-payload-spec.md`
7. `docs/current/context-reset-brief.md`
+167 -134
View File
@@ -2,79 +2,86 @@
## 1. 文档目的
本文档定义 OpenClaw 在正式环境中如何调用 reader 作为上游 MCP workflow service。
本文档只回答一个问题:OpenClaw 在正式环境里应该如何编排 reader。
目标不是描述 reader 内部实现,而是明确 OpenClaw 的编排动作:
这里不重复介绍 reader 内部实现,只定义正式控制面:
- 什么时候启动新 run
- 什么时候查询状态
- 什么时候读取结果
- 什么时候尝试恢复
- 什么时候直接新开 run
- 什么时候需要人工介入
- 如何启动日报
- 如何轮询 job
- 如何读取 run 结果
- 如何判断是否恢复
- 如何走异步恢复
- 什么时候直接新开 run 或人工介入
本文档基于 reader 当前**已真实落地**的能力编写,不描述尚未实现的未来接口。
## 2. 当前正式入口
---
### 2.1 新 run
## 2. 当前 reader 已正式支持的 MCP 能力
正式生产入口:
当前可用能力:
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
- `get_freshrss_pipeline_job_result`
同步入口:
- `run_freshrss_openclaw_pipeline`
同步入口只保留给 debug / fallback,不再是正式编排默认路径。
### 2.2 run 级读取
正式 run 级读取接口:
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
### 2.3 恢复
正式恢复入口:
- `inspect_resume_plan`
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
同步恢复入口:
- `resume_run`
其中:
- `run_freshrss_openclaw_pipeline` 是当前正式启动入口
- `get_run_status` / `list_runs` / `list_run_artifacts` 用于观测
- `get_delivery_payload` / `get_run_report` 用于读取正式结果
- `resume_run` 用于最小恢复能力
---
`resume_run` 只保留给 debug / fallback。
## 3. 编排基本原则
### 3.1 OpenClaw 不再手拼路径
### 3.1 OpenClaw 不手拼路径
OpenClaw 不应再自己拼 reader 输出路径来判断运行状态或读取核心结果。
OpenClaw 不应自己推导这些路径:
优先使用 MCP:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- 查状态 → `get_run_status`
- 读 payload → `get_delivery_payload`
- 读 report → `get_run_report`
- 做恢复 → `resume_run`
需要路径时,只消费 MCP 返回值:
只有在排障/人工核查时,才回退到直接看 reader run 目录。
- `output_dir`
- `artifact.path`
- `delivery_output`
- `report_output`
### 3.2 reader 是上游 workflow engine
### 3.2 顶层 `status` 才是分支依据
reader 负责:
`get_run_status` 和 job status 接口都可能做状态收敛。
- FreshRSS 拉取
- 内容提取
- 摘要
- 过滤
- payload 生成
- run 状态记录
- 最小恢复
因此:
OpenClaw 负责:
- 优先使用顶层 `status`
- `status_source` 用来解释状态来自原始 state 还是收敛结果
- `state_conflict=true` 说明底层状态文件已经落后于真实产物
- 触发执行
- 轮询状态
- 读取结果
- 生成 digest markdown
- Hugo 发布
- 聊天汇报
- 用户确认精选
- IMA 编排
不要再拿旧的 `raw_status`、`raw_current_stage` 或早期阶段名重新做分支。
### 3.3 默认生产语义
@@ -82,19 +89,18 @@ OpenClaw 负责:
- `mark_read=true`
- `debug_artifacts=false`
- 只在 debug/test/validation 时显式放宽
---
只有 debug / test / validation 时才放宽。
## 4. 标准 Happy Path
### Step 1: 启动新 run
### Step 1: 启动新 job
调用:
- `run_freshrss_openclaw_pipeline`
- `start_freshrss_pipeline_job`
推荐参数示例:
推荐参数:
```json
{
@@ -107,131 +113,156 @@ OpenClaw 负责:
}
```
期望:
预期:
- 获得 `run_id`
- 获得 `output_dir`
- 获得初始结果摘要
- 立即返回 `job_id`
- 后续由 OpenClaw 轮询 job,而不是同步等待整条流水线
如果启动阶段直接抛错:
### Step 2: 轮询 job
- 直接判为启动失败
- 不进入后续查询
调用:
### Step 2: 查询运行状态
- `get_freshrss_pipeline_job_status(job_id=...)`
根据返回:
- `status=running`:继续轮询
- `status=success`:读取 job result
- `status=failed`:进入失败处理
额外规则:
- 如果 `status_source=linked_run_reconciliation`,说明 outer job state 已落后,但 linked run 已经给出可用终态
- 如果 `status_source=stale_job_state_timeout`,把它当成终态失败,不要继续无限轮询
### Step 3: 读取 job result
调用:
- `get_freshrss_pipeline_job_result(job_id=...)`
预期读取:
- `run_id`
- `output_dir`
- `delivery_output`
- `report_output`
从这一刻开始,`run_id` 是正式的稳定句柄。
### Step 4: 读取 run 级状态与结果
调用:
- `get_run_status(run_id=...)`
根据返回:
- `status=running` → 继续轮询
- `status=success` → 进入结果读取
- `status=failed` → 进入失败处理
- `status=partial` → 视为未完成,优先看 `recovery` 和当前阶段
### Step 3: 读取正式结果
成功后读取:
- `get_delivery_payload(run_id=...)`
- `get_run_report(run_id=...)`
后续 OpenClaw 编排应以这两个接口为正式结果源,而不是自己拼路径读取 JSON。
根据 `get_run_status`:
### Step 4: 进入下游编排
- `status=running`:继续观察
- `status=success`:继续下游 digest / 发布 / 汇报
- `status=failed`:进入恢复或重跑决策
- `status=partial`:优先检查 report、artifacts 和 recovery
OpenClaw 在拿到正式 payload / report 后,继续执行:
如果 `status_source=run_report_reconciliation`,说明 `run-state.json` 已经过期,但 reader 已经根据终态产物收敛出有效状态。
- public/internal digest 生成
- Hugo 发布
- 聊天汇报
- 用户确认精选
- IMA 沉淀
如果 `status_source=stale_run_state_timeout`,说明 reader 认为该 run 长时间未收敛且没有终态产物,应按失败处理。
---
## 5. 恢复决策
## 5. 状态 → 动作映射
### 5.1 先看预检,不要直接恢复
| reader 状态 | OpenClaw 动作 |
|---|---|
| `running` | 继续轮询 `get_run_status` |
| `success` | 读取 `get_delivery_payload` 和 `get_run_report` |
| `failed` 且 `recovery.resumable=true` | 评估是否调用 `resume_run` |
| `failed` 且 `recovery.resumable=false` | 直接判失败,通常新开 run 或人工介入 |
| `partial` | 先读状态详情和 recovery,再决定继续等 / 恢复 / 人工介入 |
恢复前固定动作:
---
- 先调用 `inspect_resume_plan(run_id)`
## 6. 失败处理与恢复决策
只在以下条件同时成立时才启动恢复:
### 6.1 什么时候优先尝试 `resume_run`
- `can_resume=true`
- `recommended_action=resume`
满足以下条件时,优先考虑恢复而不是新开 run:
重点字段:
- `get_run_status` 返回 `failed`
- `recovery.resumable=true`
- 当前 run 对应的是 freshrss workflow
- 当前失败点在 reader 第一版支持的恢复范围内
- `requested_resume_from_stage`
- `resume_from_stage`
- `resume_decision_source`
- `artifact_resume_from_stage`
- `artifact_snapshot`
### 6.2 `resume_run` 当前支持范围
### 5.2 正式恢复路径
当前最小实现仅支持:
正式恢复控制面:
- 仅对带 `run-state.json` 的 freshrss run
- 仅从最近可恢复点继续
- 支持的恢复点:
1. `start_resume_job(run_id)`
2. `get_resume_job_status(job_id)`
3. `get_resume_job_result(job_id)`
不要再把同步 `resume_run(run_id)` 当成正式恢复入口。
### 5.3 当前支持范围
当前只支持:
- 带有效 `run-state.json` 的 `freshrss_daily_digest` run
- 从以下阶段恢复:
- `generate_summaries`
- `apply_filters`
- `build_delivery_payload`
- `write_run_report`
明确不支持:
当前不支持:
- `fetch_feed`
- `extract_articles`
### 6.3 什么时候不要恢复,直接新开 run
正式生产恢复优先依赖:
以下情况不建议 `resume_run`:
- `summary/summary-batch.json`
- `candidates/candidate-batch.json`
- `recovery.resumable=false`
- run 没有 `run-state.json`
- 失败点是 `fetch_feed` 或 `extract_articles`
- 恢复所需关键产物缺失
- 恢复点语义不明确或结果存在明显漂移风险
### 5.4 什么时候不要恢复
这时更合理的动作通常是:
以下情况直接新开 run 更合理:
- 直接新开 run
- 或人工介入排查
- `recommended_action=start_new_run`
- `recommended_action=read_terminal_result`
- 没有有效 `run-state.json`
- 恢复所需关键 artifacts 缺失
- 连续恢复失败
### 6.4 什么时候需要人工介入
## 6. 状态到动作映射
出现以下任一情况时,建议人工介入:
| 接口 | 状态 | OpenClaw 动作 |
| --- | --- | --- |
| `get_freshrss_pipeline_job_status` | `running` | 继续轮询 job |
| `get_freshrss_pipeline_job_status` | `success` | 读取 `get_freshrss_pipeline_job_result` |
| `get_freshrss_pipeline_job_status` | `failed` | 结束本次 job,必要时读 linked run |
| `get_run_status` | `running` | 继续观察 run |
| `get_run_status` | `success` | 读取 `get_delivery_payload` / `get_run_report` |
| `get_run_status` | `failed` | 先看 `inspect_resume_plan` |
| `inspect_resume_plan` | `recommended_action=resume` | 启动 `start_resume_job` |
| `inspect_resume_plan` | `recommended_action=read_terminal_result` | 直接读 run 结果,不恢复 |
| `inspect_resume_plan` | `recommended_action=start_new_run` | 新开 run 或人工介入 |
## 7. 人工介入条件
出现以下任一情况时,建议不要自动编排:
- 连续恢复失败
- `get_run_status` 与实际产物明显不一致
- payload/report 结构不符合预期
- 恢复依赖的关键文件缺失且原因不明
- FreshRSS / LLM / 外部环境异常
- payload / report 结构不符合预期
- `get_run_status` 与实际产物长期明显冲突
- FreshRSS、LLM 或外部依赖异常
- 恢复判定结果和编排预期不一致
---
## 8. 结论
## 7. 读取结果的标准动作
当前 OpenClaw 的正式调用方式已经收口为两条异步控制面:
### 7.1 `get_delivery_payload`
- 主日报:`start_freshrss_pipeline_job -> poll -> get result -> run reads`
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
用途:
- 获取正式交付给 OpenClaw 的 payload
- 后续 digest 生成应以该返回为准
OpenClaw 应做:
- 读取后直接进入 digest 生成
- 不再自己拼 `candidates/openclaw-delivery-payload.json`
同步 `run_freshrss_openclaw_pipeline` 和 `resume_run` 仅用于 debug / fallback,不应再作为默认正式编排路径。
### 7.2 `get_run_report`
@@ -263,7 +294,9 @@ OpenClaw 应做:
1. 调 `get_run_status`
2. 若 `failed && recovery.resumable=true`:
- 调 `resume_run`
- 调 `start_resume_job`
- 轮询 `get_resume_job_status`
- 读取 `get_resume_job_result`
3. 恢复后再次:
- 调 `get_run_status`
- 若成功,再读 payload / report
@@ -279,7 +312,7 @@ OpenClaw 应做:
- 让 OpenClaw 直接长时间 `exec` reader CLI 作为主要生产入口
- 让 OpenClaw 自己拼 reader 输出路径来判断成功/失败
- 让 OpenClaw 自己读取 `outputs/.../*.json` 作为正式结果源
- 在未确认 `resume_run` 支持范围外的失败点上强行恢复
- 在未确认恢复支持范围外的失败点上强行恢复
CLI 现在的定位是:
@@ -293,7 +326,7 @@ CLI 现在的定位是:
## 10. 当前已知局限
- `resume_run` 仍是最小实现,不支持任意 stage 任意重入
- 恢复能力仍是最小实现,不支持任意 stage 任意重入
- 历史无 `run-state.json` 的 run 不支持正式恢复
- 极旧 run 的结果读取仍可能依赖保守目录扫描
- `write_run_report` 若涉及重新 `mark_read`,仍依赖 FreshRSS 环境和可用凭据
@@ -304,4 +337,4 @@ CLI 现在的定位是:
OpenClaw 当前应把 reader 当作正式 MCP workflow service 使用:
**启动用 `run_freshrss_openclaw_pipeline`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复仅在 `resume_run` 最小支持范围内启用;不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**
**启动用 `start_freshrss_pipeline_job`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复默认用 `inspect_resume_plan` + `start_resume_job`,不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**
+86
View File
@@ -0,0 +1,86 @@
# [bug] MCP `start_article_summary_job` 超时但 job 实际执行了
## 问题描述
连续多个 `start_article_summary_job` 调用报 MCP 超时(-32001 Request timed out),但 job 实际执行了:
- `run_state.json` 显示 `status: failed`,`current_stage: generate_markdown`
- job 目录正常生成,`run-state.json` 存在于 `outputs/freshrss/article_summary_jobs/<job-id>/`
- `generate_markdown` 阶段实际执行过(有结果),但 MCP 响应没能发回来
## 根因分析(已定位)
**根本原因:双阻塞点导致 MCP stdio 响应超时**
### 阻塞点 1:`RunStore.save()` 高频同步写盘
`start_article_summary_job` 执行流程中的写盘次数:
```python
store.save() # ← 第 134 行:初始化后写盘
store.start_stage("prepare_job") # ← 第 68 行:内部 save()
store.register_artifact(...) # ← 第 143 行:内部 save()
store.finish_stage("prepare_job", ...) # ← 第 90 行:内部 save()
return {...} # ← 第 158 行:返回 MCP 响应
```
**4 次同步写盘** 在 `Popen` 之后、`return` 之前完成。磁盘 I/O 慢 + Windows 文件锁 = **响应超时**。
### 阻塞点 2:MCP FastMCP stdio 传输机制
`mcp.run()` 默认使用 **stdio 传输**(进程间管道)。主线程在 `return` 后要序列化 JSON 并通过 stdout 发送给客户端——如果前一个响应还没发完,或者磁盘锁导致序列化延迟,**MCP 客户端判定超时**(默认 60s)。
### 原代码问题
```python
# ← 原代码:同步阻塞
proc = subprocess.Popen(...) # Popen 返回
store.finish_stage("prepare_job", outputs={...}) # ← 同步写盘
return {"job_id": job_id, ...} # ← 这里已经超时了
```
## 修复方案(已实施)
**方案 A:异步解耦** —— 用 `ThreadPoolExecutor` 后台启动 job,主线程立即返回响应。
```python
# ← 修复后:异步非阻塞
_executor = ThreadPoolExecutor(max_workers=4, thread_name_prefix="article_summary_job")
def _launch_job_background(*, job_id, input_payload, store):
proc = subprocess.Popen(...)
store.finish_stage("prepare_job", outputs={...}) # ← 后台线程写盘
def start_article_summary_job(...):
# ... 创建 store 和 input 文件
store.start_stage("prepare_job") # ← 不调用 save()
store.register_artifact(...) # ← 不调用 save()
# 【关键】:后台线程执行 Popen + finish_stage,主线程立即返回
_executor.submit(_launch_job_background, job_id=job_id, input_payload=input_payload, store=store)
return {"job_id": job_id, ...} # ← 立即返回,不等待写盘
```
**修复效果**:
- 主线程:创建 job 目录 → 写 input.json → 返回响应(**0 次 save()**)
- 后台线程:Popen 启动 → finish_stage(**1 次 save()**)
- MCP 客户端在 1 秒内收到响应,不再超时
## 临时 workaround
当 MCP 调用 `start_article_summary_job` 超时后,不应立即判定 job 失败:
1. 调用 `get_article_summary_job_status(job_id)` 查询真实状态
2. 若返回 `status=running` 或 `run-state.json` 存在且 `status=running` → job 在跑,继续等待
3. 若返回 `status=failed` → 查 `run-state.json` 的 `failed_stage` 和 `error_summary`
## 影响范围
- OpenClaw MCP 客户端调用 `start_article_summary_job`
- 任何通过 stdio MCP 通道使用 article summary job 的场景
## 修复时间
- **根因定位**:2026-04-16
- **修复实施**:2026-04-16(方案 A:异步解耦)
+73
View File
@@ -0,0 +1,73 @@
# 规划文档导航
## 使用原则
`plans/` 目录保存架构设计、实施计划、专题方案和历史问题分析。
不要把所有 plan 都当成当前权威事实。
当前事实优先级应是:
1. `README.md`
2. `docs/README.md`
3. `docs/openclaw/README.md`
4. `docs/openclaw/openclaw-handoff.md`
5. `docs/openclaw/openclaw-orchestration-flow.md`
6. `TODO.md`
`plans/` 更适合回答:
- 为什么这样设计
- 某个能力是怎么分阶段落地的
- 某次事故当时是怎么分析的
## 当前权威规划
这些文档仍然是当前协作时应优先阅读的规划基线:
- `reader-mcp-architecture-design.md`
- MCP workflow service 的总体架构方向
- `reader-mcp-implementation-plan.md`
- 实施分阶段计划
- `../TODO.md`
- 当前任务状态与落地进展
## 已完成能力的专题方案
这些方案主要用于回看设计取舍,相关能力已经基本落地:
- `freshrss-pipeline-async-job-plan.md`
- 主日报 async job 方案
- `article-summary-async-job-plan.md`
- 单篇总结 async job 方案
- `resume-run-minimal-design.md`
- `resume_run` 最小恢复语义设计
## 仍有参考价值的专题设计
- `keyword-cleanup-artifact-slimming-v1.md`
- `keyword-cleanup-review-suggestions-layer-design.md`
- `article-summary-prompt-independent.md`
- `article-deep-summary-skill.md`
- `docker-deployment-plan.md`
## 历史问题分析
- `issues/2026-04-06-reader-digest-sigterm.md`
- 一次真实运行事故的分析
## 建议阅读顺序
如果是新接手维护:
1. `reader-mcp-architecture-design.md`
2. `reader-mcp-implementation-plan.md`
3. `../TODO.md`
4. `../docs/openclaw/README.md`
5. `../docs/openclaw/openclaw-handoff.md`
如果是在排查某一类能力:
- 主日报启动/轮询:看 `freshrss-pipeline-async-job-plan.md`
- 恢复:看 `resume-run-minimal-design.md`
- 单篇总结:看 `article-summary-async-job-plan.md`
- 历史故障:看 `issues/2026-04-06-reader-digest-sigterm.md`
+757
View File
@@ -0,0 +1,757 @@
# 单篇总结异步 job 最小版落地方案
## 1. 背景与问题定义
当前 reader 已经把日更 FreshRSS 主流程做成了带 `run-state.json` 的 runtime 模型:
- `src/summary_mcp/runtime/state_models.py`
- `src/summary_mcp/runtime/run_store.py`
- `src/summary_mcp/runtime/query_service.py`
- `src/summary_mcp/workflows/freshrss_pipeline.py`
这条主链路已经具备:
- run / stage / artifact 的结构化状态
- MCP 查询接口:`get_run_status` / `list_runs` / `list_run_artifacts`
- 结果读取接口:`get_delivery_payload` / `get_run_report`
但“单篇总结”这条线目前还是同步调用:
- 核心逻辑:`src/summary_mcp/workflows/article_summary.py`
- MCP 暴露:`src/summary_mcp/server.py` 中的 `generate_article_summaries`
- CLI 辅助:`scripts/run_article_summaries.py`
现状问题已经很明确:
- 在 OpenClaw → MCP tool 这条链路里,`generate_article_summaries` 可能因为 tool 调用时长而 timeout
- 但 reader 项目本体在 `.venv` 下直接跑 article summary,大约 37.5 秒即可成功
- 这说明问题不一定在 summary 业务本身,而更可能在“同步工具调用 + 上层等待模型”这个包装层
所以目标不是先继续调 timeout,而是把单篇总结也纳入 **真正异步、可轮询、可落盘、可恢复基本状态** 的最小 job 模型里。
---
## 2. 为什么同步 MCP 不适合这一步
`generate_article_summaries` 当前在 `server.py` 里直接同步执行 `summarize_selected_articles(...)`,调用方必须一直阻塞等待,直到:
1. 读取 extracted payload
2. 调用 LLM 生成总结
3. 可选 repair retry
4. 渲染 Markdown
5. 写文件完成
6. MCP tool 返回生成路径
这个模式对“几十秒级、依赖外部 LLM、可能重试”的任务不稳,核心问题有三层:
### 2.1 tool 调用时长不可控
`run_loop_payload()` 内部会发起外部 HTTP 请求,还可能做 repair retry。即便单次平均 37.5 秒,也已经接近很多上层编排系统的心理和技术超时边界。
### 2.2 调用方看不到中间状态
现在如果卡住,调用方只能等:
- 不知道是在读输入
- 不知道是在调 LLM
- 不知道是在重试
- 不知道是否已经写出部分结果
这也是同步接口最烦的点:失败时只能看到“tool timeout / tool failed”,而不是“业务跑到哪一步了”。
### 2.3 与 reader 已有 runtime 风格不一致
FreshRSS 主流程已经是“run truth + status query + artifact read”的思路,而单篇总结仍然是黑箱同步函数。继续维持两套风格,只会让 SOP 更复杂:
- 日报主链路用 `run_id`
- 单篇总结却要么同步等,要么退回 CLI fallback
这不利于后续把 `reader-digest-flow` 稳定成正式 SOP。
---
## 3. 本轮目标:最小真异步,不做大而全
这次只做 **单篇总结异步 job 最小版**,目标是:
> 让 OpenClaw 或其他调用方能先“启动单篇总结 job”,立即拿到 `job_id`,再通过状态接口轮询,最后读取输出文件/结果。
### 3.1 本轮必须做到的范围
1. 新增单篇总结 job 的 start/status/result 最小接口
2. job 真正在后台执行,而不是 MCP 请求线程里阻塞到完成
3. 状态落盘到文件,遵循 reader 当前 runtime 风格
4. 复用现有 `summarize_selected_articles` 逻辑,不重写业务
5. 输出仍然是现有 Markdown 文件,不改知识内容 schema
### 3.2 本轮明确不做
1. **不做通用队列系统**
2. **不做数据库**
3. **不做多 worker / 分布式调度**
4. **不做取消 job / kill job**
5. **不做并发配额控制**
6. **不做 resume/retry from stage**
7. **不把 article summary 一次性并入 freshrss `resume_run` 体系**
8. **不改 summary prompt / validator / 输出格式**
9. **不处理批量高吞吐场景优化**
一句话:这轮只解“同步 tool 容易 timeout,但业务本身能跑完”这个问题,不顺手扩成任务调度平台。
---
## 4. 接入当前 reader 结构的建议
### 4.1 复用现有 runtime 设计,但单独建 article summary job 命名空间
不建议把 article summary job 粗暴塞进现有 `freshrss_daily_digest` run 查询里混用一个 schema;更合适的是:
- 复用 `RunState / StageState / ArtifactRecord / RunStore` 这套思维
- 但给单篇总结定义独立 workflow 名称与存储目录
建议:
- workflow: `article_summary_job`
- run_type: `article_summary`
- output root: `outputs/freshrss/article_summary_jobs/<job_id>/`
这样有几个好处:
- 不污染 `outputs/freshrss/rerun/`
- 语义清楚:这是独立 job,不是假装自己是日报 rerun
- 查询和排查时更直观
### 4.2 job 与结果 Markdown 解耦
job 目录只负责:
- 状态文件
- 输入快照
- artifact 索引
- 执行报告
真正生成的总结 Markdown,仍然可以写到用户指定的 `output_dir`(或默认 `single_summaries/`)。
这样不破坏当前下游 SOP:
- `reader-digest-flow` 依然从原来的单篇总结输出目录拿 `.md`
- job 目录只提供状态与索引,不强迫下游改结果路径约定
---
## 5. 新增工具 / API 设计
本轮建议新增 3 个 MCP tool,名字尽量和当前 runtime 风格一致。
## 5.1 `start_article_summary_job`
### 作用
启动一个后台 job,立即返回 `job_id`,不等待总结完成。
### 输入建议
```json
{
"extracted_path": "outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json",
"selected_ids": ["12345"],
"output_dir": "outputs/freshrss/single_summaries/2026-04-10",
"max_retries": 2,
"timeout_seconds": 120,
"llm_api_key": null,
"llm_model": null,
"llm_api_url": null
}
```
### 返回建议
```json
{
"job_id": "article-summary-20260410-144500-ab12cd34",
"workflow": "article_summary_job",
"run_type": "article_summary",
"status": "running",
"output_dir": "outputs/freshrss/article_summary_jobs/article-summary-20260410-144500-ab12cd34",
"message": "Article summary job started successfully. Use get_article_summary_job_status to poll progress."
}
```
### 约束建议
- `selected_ids` 第一版允许多个,但建议由 OpenClaw 每次只传一篇或少量篇,避免一个 job 干太多事
- `extracted_path` 必须存在,否则直接拒绝启动
- `output_dir` 不传则按当前默认逻辑推导
---
## 5.2 `get_article_summary_job_status`
### 作用
查询 job 当前状态、阶段、进度、错误摘要、已注册 artifact。
### 输入
```json
{
"job_id": "article-summary-20260410-144500-ab12cd34"
}
```
### 返回建议
```json
{
"job_id": "article-summary-20260410-144500-ab12cd34",
"workflow": "article_summary_job",
"run_type": "article_summary",
"status": "running",
"current_stage": "generate_markdown",
"started_at": "2026-04-10T14:45:00+08:00",
"updated_at": "2026-04-10T14:45:23+08:00",
"finished_at": null,
"output_dir": "outputs/freshrss/article_summary_jobs/article-summary-20260410-144500-ab12cd34",
"progress": {
"completed_stage_count": 2,
"running_stage_count": 1,
"failed_stage_count": 0,
"pending_stage_count": 1,
"total_stage_count": 4
},
"artifacts": [...],
"error_summary": null
}
```
---
## 5.3 `get_article_summary_job_result`
### 作用
当 job 成功后,返回结构化结果,供 OpenClaw 继续下游 IMA 沉淀。
### 输入
```json
{
"job_id": "article-summary-20260410-144500-ab12cd34"
}
```
### 返回建议
```json
{
"job_id": "article-summary-20260410-144500-ab12cd34",
"status": "success",
"written_paths": [
"outputs/freshrss/single_summaries/2026-04-10/某篇文章标题.md"
],
"artifact": {
"name": "job_result",
"path": "outputs/freshrss/article_summary_jobs/article-summary-20260410-144500-ab12cd34/result.json",
"kind": "json",
"stage": "write_result"
},
"result": {
"selected_ids": ["12345"],
"written_paths": [
"outputs/freshrss/single_summaries/2026-04-10/某篇文章标题.md"
]
}
}
```
### 行为建议
- 若 job 还没完成,返回当前状态 + 提示“not ready”
- 若 job 失败,返回失败摘要,不硬抛文件不存在异常
---
## 6. job 状态文件设计
建议直接复用现有 `RunState` 模型,不另造一套 schema。
job 目录示例:
```text
outputs/freshrss/article_summary_jobs/
article-summary-20260410-144500-ab12cd34/
run-state.json
input.json
result.json
job-report.json
```
## 6.1 `run-state.json`
建议直接沿用当前字段:
```json
{
"run_id": "article-summary-20260410-144500-ab12cd34",
"workflow": "article_summary_job",
"run_type": "article_summary",
"status": "running",
"current_stage": "generate_markdown",
"started_at": "2026-04-10T14:45:00+08:00",
"updated_at": "2026-04-10T14:45:23+08:00",
"finished_at": null,
"input": {
"extracted_path": "outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json",
"selected_ids": ["12345"],
"output_dir": "outputs/freshrss/single_summaries/2026-04-10",
"max_retries": 2,
"timeout_seconds": 120
},
"stages": [...],
"artifacts": [...],
"error": null,
"recovery": {
"resumable": false,
"resume_from_stage": null,
"last_success_stage": "load_input"
}
}
```
### 第一版 stage 建议
建议只切 4 个 stage,够看即可:
1. `prepare_job`
- 校验输入
- 解析路径
- 写 `input.json`
2. `load_input`
- 读取 extracted payload
- 确认 `selected_ids` 可匹配条目
3. `generate_markdown`
- 调 `summarize_selected_articles(...)`
- 这是主要耗时阶段
4. `write_result`
- 写 `result.json`
- 注册输出 artifact
这里不要把 LLM 调用再拆更多细 stage,否则最小版反而过度设计。
---
## 6.2 `input.json`
作用:保留启动请求快照,便于排查。
建议内容与 `run-state.input` 基本一致。
---
## 6.3 `result.json`
成功时写:
```json
{
"job_id": "article-summary-20260410-144500-ab12cd34",
"selected_ids": ["12345"],
"written_paths": [
"outputs/freshrss/single_summaries/2026-04-10/某篇文章标题.md"
],
"completed_at": "2026-04-10T14:45:41+08:00"
}
```
失败时可以不写,或只写失败快照都行。最小版建议:
- 成功写 `result.json`
- 失败只依赖 `run-state.json`
避免双份失败状态不一致。
---
## 7. 执行模型建议:优先子进程,不建议线程
### 7.1 推荐:子进程后台执行
最小真异步推荐模型:
- `start_article_summary_job` 负责:
- 创建 job 目录
- 初始化 `run-state.json`
- 通过 `subprocess.Popen(...)` 启动一个独立 Python 进程执行 job runner
- 立即返回 `job_id`
后台 runner 再去:
- 读取 `input.json`
- 用 `RunStore` 更新状态
- 调用 `summarize_selected_articles(...)`
- 写 `result.json`
- finish/fail run
### 7.2 为什么不推荐线程
虽然线程实现看起来更省事,但不适合作为 reader 的正式最小异步落地:
1. **MCP server 进程重启后线程直接丢失**
2. 线程状态不天然可恢复,容易出现“状态文件还在 running,但线程没了”
3. 未来要做健康检查/孤儿 job 检测时,线程模型更难收口
### 7.3 为什么子进程更贴当前项目风格
reader 现在本来就偏“文件产物 + runtime 状态真相”风格。子进程模式有天然优势:
- 和 CLI/fallback 思维一致
- 进程边界清楚
- `run-state.json` 由实际执行者写,职责清晰
- 将来如果要做 orphan detection / stale running job 修复,也容易补
### 7.4 本轮不做进程管理增强
最小版里,不要求:
- 记录 PID 后做 kill/cancel
- 自动清理僵尸进程
- 守护进程/worker 池
但建议在 `input.json` 或 `run-state.input` 里附带:
- `launcher_pid`
- `runner_command`
方便排障。
---
## 8. 与现有 `summarize_selected_articles` 的复用关系
核心原则:**不重写总结业务,只包一层 job runner。**
### 8.1 直接复用的部分
`src/summary_mcp/workflows/article_summary.py` 已经做了:
- 读取 extracted payload
- 根据 `selected_ids` 找条目
- 调用 `run_loop_payload(...)`
- 渲染 Markdown
- 写到 `output_dir`
- 返回 `list[Path]`
这些都继续用。
### 8.2 最小新增建议
建议只新增一层 runtime/service,例如:
- `src/summary_mcp/runtime/article_summary_jobs.py`
职责:
- 生成 `job_id`
- 创建 job 目录
- 初始化 `RunStore`
- 启动 runner 子进程
- 提供 status/result 查询
- 在 runner 里调用 `summarize_selected_articles`
### 8.3 是否需要改 `summarize_selected_articles`
最小版尽量少改,只建议加两类低风险增强:
1. **可选输入校验增强**
- 如果 `selected_ids` 一个都匹配不到,显式报错
- 避免“成功返回空列表”却让 job 看起来像成功
2. **可选 hook / telemetry(非必须)**
- 如果后面需要更细粒度写 stage 输出,可再加
- 但第一版没必要为了观测性重构函数
结论:
- 第一版优先保持 `summarize_selected_articles` 基本不动
- job 层只把它作为黑盒业务函数调用
---
## 9. 建议新增代码组织
建议新增文件:
```text
src/summary_mcp/runtime/article_summary_jobs.py
scripts/run_article_summary_job.py
```
### 9.1 `article_summary_jobs.py`
建议包含:
- `start_article_summary_job(...)`
- `run_article_summary_job(...)`
- `get_article_summary_job_status(...)`
- `get_article_summary_job_result(...)`
- 若干私有 helper:
- job id 生成
- job dir 解析
- result artifact 读取
### 9.2 `scripts/run_article_summary_job.py`
作用:作为子进程 runner 入口。
例如:
```bash
python scripts/run_article_summary_job.py --job-id article-summary-...
```
runner 只做一件事:
- 根据 `job_id` 找到 job 目录和 `input.json`
- 真正执行 job
这样避免在 `Popen("python -c ...")` 里塞长字符串,也方便本地调试。
---
## 10. 对 `reader-digest-flow` SOP 的影响
这块是重点,因为老大的实际痛点就在这。
### 10.1 现状 SOP
当前 skill 已经有硬规则:
- 优先走 MCP `generate_article_summaries`
- MCP timeout 时,立刻 fallback 到 reader 本地 `.venv`
这个 fallback 现在是必要的,但它本质是在补“单篇总结没有正式异步接口”。
### 10.2 引入异步 job 后的推荐 SOP
建议调整为:
1. OpenClaw 在用户确认保留文章后
2. 调 `start_article_summary_job`
3. 拿到 `job_id`
4. 轮询 `get_article_summary_job_status`
5. 成功后调 `get_article_summary_job_result`
6. 再继续 IMA 格式化与上传
### 10.3 对 skill 文档的影响
`reader-digest-flow` 需要后续补一条新规则:
- 单篇总结的正式生产路径从“同步 MCP + 本地 fallback”升级为“异步 job + 状态轮询”
- 本地 `.venv` CLI fallback 仍保留,但降级为:
- job 启动失败
- job runner 异常
- reader 服务端出现系统性问题时的应急路径
### 10.4 用户体验改善
异步 job 后,OpenClaw 可给出更像正式系统的反馈:
- “已启动单篇总结任务,正在生成”
- “当前状态:generate_markdown”
- “已完成,准备继续沉淀到 IMA”
而不是现在的:
- 直接卡住几十秒
- 然后 tool timeout
- 再走一套 fallback
---
## 11. 验证方案
本轮验证不要追求大而全,按 4 层做就够。
### 11.1 单元级验证
目标:确保 job 状态文件和结果文件行为正确。
建议覆盖:
1. `start_article_summary_job` 能创建 job 目录与 `run-state.json`
2. 输入路径不存在时,启动直接失败
3. `get_article_summary_job_status` 能正确读取状态
4. job 成功后 `get_article_summary_job_result` 返回 `written_paths`
5. job 失败后状态为 `failed`,并带错误摘要
### 11.2 本地集成验证
用一个真实 extracted 文件跑:
1. 启动 job
2. 轮询 status
3. 成功后检查:
- `result.json` 存在
- Markdown 文件存在
- 路径正确
### 11.3 OpenClaw 链路验证
在实际 `reader-digest-flow` 环节,用一篇已确认保留的文章做:
1. 启动 async job
2. 等待成功
3. 继续做 IMA markdown 整理与上传
4. 确认整个链路不再因为同步 tool timeout 中断
### 11.4 异常验证
至少测 3 类异常:
1. `selected_ids` 不存在
2. LLM 接口失败 / 超时
3. runner 进程异常退出
预期:
- `run-state.json` 最终为 `failed`
- `error_summary` 可读
- 调用方能明确知道失败,而不是只看到 transport timeout
---
## 12. 风险与回滚方案
## 12.1 风险
### 风险 1:后台子进程成功启动,但状态长期卡在 running
常见原因:
- runner 进程崩了
- server 重启时某些路径没写全
- 子进程命令不对
缓解:
- runner 启动前就先写 `run-state.json`
- runner 一进来先更新 `prepare_job` / `load_input`
- 后续可加“stale running 超时判定”,但第一版先不做自动修复
### 风险 2:复用旧函数导致“空输出也算成功”
当前 `summarize_selected_articles()` 如果没匹配到文章或全部失败,存在返回空列表的可能。
缓解:
- job runner 里把“`written_paths` 为空”视为失败
- 或者顺手在 `summarize_selected_articles` 里补显式校验
### 风险 3:同一时刻大量 job 并发,LLM 调用被打爆
第一版不解决系统级并发控制。
缓解:
- SOP 层先按单篇/少量串行用
- skill 层避免一口气启动很多 job
### 风险 4:状态目录与结果目录分离,排查时容易迷路
缓解:
- 在 `result.json` 和 artifact 元数据里明确记录 `written_paths`
- 在 `run-state.input.output_dir` 中保留结果目录
---
## 12.2 回滚方案
这个方案很好回滚,因为是“新增,不替换”。
### 回滚原则
- 保留现有 `generate_article_summaries`
- 保留现有 `scripts/run_article_summaries.py`
- 新增 async job 接口如果不稳定,直接停止在 OpenClaw 层使用即可
### 回滚路径
1. 停止调用 `start_article_summary_job`
2. 恢复回原 SOP:
- 先尝试同步 MCP `generate_article_summaries`
- 失败则本地 `.venv` fallback
3. async job 相关代码保留但不作为正式入口
也就是说,这轮改动不会堵死当前生产路径,风险可控。
---
## 13. 实施顺序(建议按这个顺序落地)
### Phase A:先把方案落地成最小代码骨架
1. 新增 `src/summary_mcp/runtime/article_summary_jobs.py`
2. 新增 job 目录常量与 helper
3. 新增 `scripts/run_article_summary_job.py`
4. 实现 runner 内部对 `summarize_selected_articles` 的调用
5. 先本地命令验证 job 可跑通
### Phase B:再把 MCP 接口接上
6. 在 `server.py` 新增:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
7. 本地 MCP 调用验证
### Phase C:最后接 OpenClaw SOP
8. 更新 `reader-digest-flow` skill,把正式路径切到 async job
9. 保留本地 `.venv` fallback 作为应急方案
10. 跑一次真实日报保留文章沉淀闭环
---
## 14. 推荐的最小返回/状态语义
为了跟现有 runtime 风格一致,建议继续用:
- `status`: `running` / `success` / `failed`
- `current_stage`: 当前 stage 名
- `artifacts`: 注册产物列表
- `error_summary`: 结构化错误
第一版不必引入:
- `queued`
- `cancelled`
- `retrying`
- `paused`
避免状态机一开始就复杂化。
---
## 15. 结论 / 拍板建议
拍板建议很直接:
1. **这件事值得做,而且优先级高**,因为它正好卡在当前日报 SOP 的真实痛点上
2. **不建议继续先修同步 timeout**,因为同步模型本身就不适合几十秒级、外部 LLM 驱动的任务
3. **最小真异步应采用“文件状态 + 后台子进程 + start/status/result 三接口”**
4. **业务层严格复用 `summarize_selected_articles`**,不要为异步化重写 summary 核心逻辑
5. **先把单篇总结 job 做成独立小 runtime 命名空间**,不要急着并入 freshrss 主 run 的 resume 体系
如果只允许做一轮最小落地,我建议就做到:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
- `outputs/freshrss/article_summary_jobs/<job_id>/run-state.json`
- 子进程 runner
这套已经足够把当前 OpenClaw timeout 问题从“同步等待”改成“正式异步轮询”,并且几乎不碰无关模块。
+277
View File
@@ -0,0 +1,277 @@
# reader MCP Docker 部署计划
## 1. 目标
将 reader 作为正式 MCP workflow service 以 Docker 方式部署,满足以下原则:
1. 服务运行在容器内
2. 运行态与产物必须外置挂载,不闷在容器内
3. 配置统一记录在 `.env`
4. 读写行为与当前仓库约定保持一致
5. OpenClaw 后续可将该服务作为正式上游 MCP 使用
---
## 2. 部署原则
### 2.1 容器职责
容器只负责:
- 提供 reader MCP 服务运行环境
- 加载 reader 代码与依赖
- 读取挂载进来的配置与状态目录
- 对外暴露 MCP 服务入口
### 2.2 宿主机职责
宿主机负责持久化:
- 配置文件
- 运行态
- outputs 产物
- 数据目录
- configs
### 2.3 配置收口原则
所有环境配置统一放在 `.env`,避免:
- 零散写在 compose 内
- 零散写在 shell 命令里
- 零散写在 OpenClaw skill 里
---
## 3. 建议部署目录
建议在 reader 仓库内准备标准部署结构:
```text
/home/ubuntu/zhu/github/reader/
Dockerfile
docker-compose.yml
.env
outputs/
data/
configs/
knowledge-base/
```
说明:
- `Dockerfile`:构建 reader MCP 服务镜像
- `docker-compose.yml`:单服务部署编排
- `.env`:统一环境变量
- `outputs/`:产物、run-state、digest、payload 等外置持久化
- `data/`:term index 等数据外置持久化
- `configs/`:reader 运行配置外置持久化
- `knowledge-base/`:如当前 reader/skill 仍会依赖本地知识目录,可继续挂载
---
## 4. 必须挂载的目录 / 文件
### 必须挂载
- `.env`
- `outputs/`
- `data/`
- `configs/`
### 建议挂载
- `knowledge-base/`
### 通常不必挂载
- `docs/`
- `plans/`
- `.git/`
---
## 5. `.env` 统一配置建议
至少应包含以下配置:
### FreshRSS
- `FRESHRSS_API_BASE_URL`
- `FRESHRSS_USERNAME`
- `FRESHRSS_API_PASSWORD`
### 主 LLM
- `LLM_API_URL`
- `LLM_API_KEY`
- `LLM_MODEL`
### 单篇总结专用 LLM(如已使用)
- `ARTICLE_SUMMARY_API_URL`
- `ARTICLE_SUMMARY_API_KEY`
- `ARTICLE_SUMMARY_MODEL`
### IMA(如 reader / skill 仍依赖这些配置约定)
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
### 运行控制
- `PYTHONUNBUFFERED=1`
- 视需要增加日志级别等配置
原则:
- 所有会影响服务行为的环境项,都优先进入 `.env`
- compose 文件只引用 `.env`,不在 compose 里硬编码业务参数
---
## 6. Dockerfile 设计建议
### 目标
- 使用 Python 3.11
- 安装 reader 依赖
- 默认启动 MCP 服务入口
### 建议思路
1. 基于 `python:3.11-slim`
2. 设置工作目录到 `/app`
3. 复制仓库代码
4. 安装依赖(如 `pip install -e .`)
5. 默认启动 reader MCP 服务
### 启动入口
优先使用当前正式服务入口,例如:
- `summary-mcp`
如果后续 reader 明确切换到别的稳定入口,再同步更新。
---
## 7. docker-compose 设计建议
建议先保持单服务简单结构,例如:
- service 名称:`reader-mcp`
- `env_file: .env`
- 挂载:
- `./outputs:/app/outputs`
- `./data:/app/data`
- `./configs:/app/configs`
- `./knowledge-base:/app/knowledge-base`(如需要)
- `./.env:/app/.env:ro`(可选,若程序直接读取文件)
- `restart: unless-stopped`
如果当前 MCP 服务是 stdio 型而不是 HTTP 型,需要进一步明确:
- 它是由 OpenClaw 以本地进程方式拉起
- 还是以常驻 sidecar / gateway adapter 方式挂接
因此 compose 的最终 command 需要结合实际接入方式确认。
---
## 8. 部署前确认项
在正式执行前,需要先确认以下问题:
### 8.1 MCP 连接方式
必须确认 reader MCP 服务的正式接入方式是:
1. **stdio 型**:OpenClaw/调用方本地拉起进程
2. **HTTP/SSE 型**:服务常驻监听端口,OpenClaw 远程连接
这会直接影响:
- Docker command
- 是否需要端口映射
- OpenClaw 接入配置
### 8.2 当前 `summary-mcp` 的服务形态
需要确认:
- 现有 `summary-mcp` 是 FastMCP stdio 默认模式
- 还是已有可直接 HTTP 化的运行方式
在这点没确认前,不要盲目写死端口暴露方案。
### 8.3 OpenClaw 侧接入点
部署完成后,还需要明确 OpenClaw 将如何引用该 MCP 服务:
- 本机命令型 MCP
- Docker 内服务桥接
- 或其它现有 OpenClaw MCP 配置方式
---
## 9. 执行顺序(建议)
### Phase A:部署方案落地
1. 确认 MCP 服务连接方式(stdio / HTTP)
2. 确认最终 Dockerfile 启动命令
3. 确认 compose 结构与挂载目录
4. 整理 `.env` 字段
### Phase B:容器化实现
1. 新建/更新 `Dockerfile`
2. 新建/更新 `docker-compose.yml`
3. 检查 `.dockerignore`
4. 核对路径是否与仓库内当前代码一致
### Phase C:本地部署验证
1. `docker compose build`
2. `docker compose up -d`
3. 验证服务启动
4. 验证容器外 `outputs/`、`data/` 等是否正常落盘
### Phase D:OpenClaw 接入验证
1. 让 OpenClaw 通过正式 MCP 路径连接 reader
2. 真跑一轮:
- run
- status
- payload
- report
3. 如有需要,验证一次最小 `resume_run`
---
## 10. 当前不在本轮范围内的事
本轮部署计划不直接处理:
- `rerun_stage`
- 更复杂的后台任务系统
- 多实例部署
- 横向扩展
- 生产告警体系
本轮只做:
- 单实例
- Docker 化
- 配置收口
- 挂载持久化
- OpenClaw 可正式接入
---
## 11. 一句话结论
reader 的下一步不是继续堆内部接口,而是:
**以 Docker 正式部署成 MCP workflow service,配置进 `.env`,状态和产物目录挂载到宿主机,然后由 OpenClaw 按正式 MCP 编排路径真实接入和验证。**
+229
View File
@@ -0,0 +1,229 @@
# FreshRSS 主日报异步 job 方案
## 背景
当前 `run_freshrss_openclaw_pipeline` 虽然已经作为正式 MCP workflow 入口存在,但执行模型仍是**同步 MCP 调用**。这会带来几个现实问题:
1. OpenClaw / MCP wrapper 存在超时风险,尤其是 5-10 篇的正式日报批次。
2. wrapper timeout 与真实 run 是否已落地,容易出现语义分离。
3. 当前已有 `run-state.json`、`get_run_status`、`get_run_report`、`get_delivery_payload`,但**启动层**仍然是同步调用,不利于正式生产链路稳定运行。
4. 单篇总结已经验证了“最小 async job + 轮询状态 + 读取结果”模型可行,主日报 run 应收敛到同一套运行模式。
## 目标
将 FreshRSS 主日报 run 改造成与 article-summary 类似的**最小真异步 job**:
- 启动即返回 `job_id`
- 真正执行由后台子进程完成
- 状态可轮询
- 成功后可读取结构化结果
- 业务逻辑继续复用既有 `run_freshrss_pipeline(...)`
- 不推翻现有 run-state / result query 能力
## 非目标
本阶段不做:
- 分布式任务队列
- 多 worker 调度
- 任意 stage 的后台恢复编排
- 并发控制中心
- 主流程与 article-summary job 的通用抽象框架一次性大重构
先做最小可用。
## 设计原则
1. **启动层异步化,执行核心不重写**
- `run_freshrss_pipeline(...)` 继续是主业务逻辑真相。
- async job 只负责启动、状态持久化、结果回读。
2. **run truth 与 job truth 分层**
- job truth:这次异步任务有没有启动、运行到哪一步、是否成功。
- run truth:真正的 freshrss workflow 输出与 `run-state.json`。
3. **OpenClaw 正式生产默认改为 async start path**
- 启动走 async job
- 状态和结果优先先看 job
- 真正业务产物仍由现有 run 查询工具承接
4. **与 article-summary job 尽量同构**
- 目录结构
- `run-state.json` / `input.json` / `result.json` / `job-report.json`
- 后台 runner 脚本
## 拟新增能力
### MCP tools
新增 3 个工具:
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
- `get_freshrss_pipeline_job_result`
### job 目录
固定目录:
`outputs/freshrss/pipeline_jobs/<job_id>/`
至少包含:
- `run-state.json`
- `input.json`
- `result.json`(成功时)
- `job-report.json`
### 执行模型
- `start_freshrss_pipeline_job` 写入 input + 初始化 job state
- 后台 `subprocess.Popen(...)` 启动 runner
- runner 内部调用 `run_freshrss_pipeline(...)`
- 成功后把 `run_id`、核心产物路径、关键计数写入 `result.json`
## job 输入参数
与现有 `run_freshrss_openclaw_pipeline` 尽量对齐:
- `limit`
- `mark_read`
- `include_read`
- `debug_artifacts`
- `continuation`
- `timeout_seconds`
- `max_retries`
- `stream_id`
- `api_base_url`
- `username`
- `api_password`
- `llm_api_key`
- `llm_model`
- `llm_api_url`
- `context`
- `run_id`
- `date_value`
- `output_dir`
- `include_item_reports`
## 返回语义
### start
返回:
- `job_id`
- `workflow`
- `run_type`
- `status=running`
- `output_dir`
- `message`
### status
返回:
- `job_id`
- `status`
- `current_stage`
- `started_at` / `updated_at` / `finished_at`
- `progress`
- `artifacts`
- `error_summary`
- 若主 run 已创建,可附带 `linked_run_id`
### result
成功时返回:
- `job_id`
- `status=success`
- `run_id`
- `delivery_output`
- `report_output`
- `digest_brief_output`
- `pulled_count`
- `delivered_count`
- `marked_read_count`
- `artifact`
- `result`
## stages 建议
最小 job stages:
1. `prepare_job`
2. `load_input`
3. `run_pipeline`
4. `write_result`
其中 `run_pipeline` 内部仍由现有 freshrss workflow 自己写它的 run-state。
## 与现有同步入口的关系
### 保留
`run_freshrss_openclaw_pipeline` 暂时保留,作为:
- debug / light path
- 本地调试工具
- 向后兼容路径
### 正式语义调整
文档与 OpenClaw handoff 中,主日报正式生产默认启动入口改为:
- `start_freshrss_pipeline_job`
同步入口降级为:
- debug / fallback
- 小批量验证
## OpenClaw 编排建议
新的推荐路径:
1. `start_freshrss_pipeline_job`
2. `get_freshrss_pipeline_job_status`
3. 成功后 `get_freshrss_pipeline_job_result`
4. 后续仍用:
- `get_run_status`
- `get_delivery_payload`
- `get_run_report`
- `list_run_artifacts`
## 风险点
1. **job 成功但 run 部分失败**
- 允许,job 结果应以真实 `run_freshrss_pipeline(...)` 返回为准。
- `run_id` + `report_output` 仍是最终真相。
2. **runner 崩溃但来不及写 result**
- 需保证 `job-report.json` 至少能写下失败摘要。
3. **重复状态源导致混淆**
- 文档必须明确:
- job state 管“启动任务”
- run state 管“业务工作流真相”
4. **同步 / 异步双入口长期漂移**
- 必须要求 async job 内部直接复用 `run_freshrss_pipeline(...)`
- 禁止再实现一套平行主流程
## 验收标准
1. 能通过 MCP 启动一个主日报 async job 并立即返回 `job_id`
2. 能轮询到 `running -> success/failed`
3. 成功后 `result.json` 含 `run_id` 与核心产物路径
4. 对应 run 仍能通过既有 `get_run_status` / `get_run_report` / `get_delivery_payload` 正常读取
5. README / handoff / TODO / plans 同步更新
## 建议实施顺序
1. 复制 article-summary job 骨架到 freshrss pipeline job
2. 新增 runner 脚本
3. server.py 暴露 3 个新工具
4. 补 query/result 读法
5. 更新 README / handoff
6. 将 TODO 主任务切到“主日报 async job”
@@ -0,0 +1,257 @@
# keyword-cleanup 产物精简方案 v1
## 1. 背景
当前 keyword-cleanup 治理链已经从“只有 review bundle”演进到:
- term stats / daily term index
- review bundle
- suggestions json
- suggestions markdown
- apply -> config / watchlist / change_log
这说明链路已经打通,但也带来一个新问题:
> 中间产物偏多,容易让治理系统本身比被治理对象更重。
本方案的目标不是回退功能,而是重新划分:
- 哪些产物是长期资产
- 哪些产物只是决策输入
- 哪些产物只是运行时工作文件
从而把 keyword-cleanup 收敛成一个更轻的治理辅助层,而不是继续长成一个复杂子系统。
---
## 2. 设计目标
本轮精简目标:
1. 保留真正有长期价值的事实层与状态层数据
2. 保留唯一正式建议产物,用于 review / apply
3. 将 review bundle 和 markdown 展示稿降级为临时产物
4. 让主链路收敛到:
`stats -> suggestions json -> apply -> config`
而不是长期依赖:
`stats -> bundle -> suggestions json + md -> review -> apply`
---
## 3. 产物分层建议
### 3.1 长期保留:事实层
这些文件是系统长期事实基础,应继续长期保留:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
原因:
- daily 文件代表每日聚合观察结果
- global stats 是治理决策的核心事实来源
- 二者共同构成 term governance 的历史依据
### 3.2 长期保留:状态层
这些文件代表治理系统当前状态,应继续长期保留:
- `configs/filter_context.personal.json`
- `configs/term_watchlist.json`
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/term_change_log.json`
原因:
- 它们是已确认生效的治理结果
- 后续 reader 行为依赖这些配置
- `change_log` 负责回溯治理动作
### 3.3 短期保留:正式建议层
建议保留:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
定位:
- 这是 review / apply 之间的唯一正式建议产物
- 机器可消费
- 可作为某次治理决策的外部依据
建议策略:
- 默认仅保留最近少量几份
- 或仅保留已经 apply 过的 suggestions JSON
- 避免无限累积所有历史 suggestions 文件
### 3.4 降级为临时产物:review bundle
建议降级:
- `outputs/term_index/review/keyword-cleanup-bundle.json`
定位:
- review 输入打包文件
- 只服务于 suggestions 生成过程
- 不属于长期治理资产
建议策略:
- 默认只保留当前最新一份
- 或迁移到更明确的 working/tmp 目录语义
- 不按日期长期积累
### 3.5 降级为临时产物:Markdown 展示稿
建议降级:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
定位:
- 纯人工审阅展示层
- 不是唯一真相
- 不参与 apply 逻辑
建议策略:
- 默认不长期持久化
- 需要人工审阅时临时生成
- 优先在聊天/界面中直接展示,而不是默认写成长期文件
---
## 4. 精简后的主链路
建议主链路口径收敛为:
1. 更新 daily term index
2. 更新 global term stats
3. 生成 suggestions JSON
4. 人工确认
5. apply 到 config
6. 记录 change log
即:
`stats -> suggestions json -> apply -> config`
其中:
- bundle = 内部工作层
- markdown = 展示层
- suggestions JSON = 唯一正式建议输入
---
## 5. 为什么这样收敛
### 5.1 避免中间层过多
如果 bundle / md / suggestions 都被长期持久化,就容易出现:
- 多份文件语义重叠
- 不知道谁是“准的”
- 哪些只是试跑产物,哪些是正式治理决策不清晰
### 5.2 保持系统重心正确
keyword-cleanup 的最终目的不是维护一个漂亮的 review 文件集合,而是:
- 持续积累稳定的关键词事实数据
- 让 interest/watch/alias/stopword 演化有据可依
- 让 reader 的长期偏好配置从真实日报里长出来
### 5.3 降低治理系统自身复杂度
治理系统应该比主系统更轻,而不是更重。
如果 review 产物越积越多,最终会反过来增加维护和理解成本。
---
## 6. 对现有实现的影响
本轮不要求删除已有能力,而是重新定义口径。
### 6.1 保留
- `build_review_bundle.py`
- `generate_term_cleanup_suggestions.py`
- `apply_term_suggestions.py`
### 6.2 调整口径
- `keyword-cleanup-bundle.json` 从“默认产物”降级为“临时工作文件”
- `term-cleanup-suggestions-YYYY-MM-DD.md` 从“正式产物”降级为“临时展示稿”
- `term-cleanup-suggestions-YYYY-MM-DD.json` 作为唯一正式建议产物保留
### 6.3 后续可选实现动作
- 覆盖式写入 bundle,而不是长期累积
- Markdown 按需生成,而不是默认总是落盘
- 增加清理策略,只保留最近 N 个 suggestions JSON
---
## 7. SOP 调整建议
### 7.1 review 阶段
默认步骤:
1. 生成或更新 term stats
2. 生成最新 bundle(临时)
3. 生成 suggestions JSON(正式)
4. 如需要人工阅读,再临时生成 Markdown 或直接在聊天展示
### 7.2 apply 阶段
apply 后以以下内容作为最终真相:
- config 文件当前值
- `term_change_log.json`
- 如需要,保留对应 suggestions JSON 作为决策依据
### 7.3 清理策略
建议:
- bundle:默认仅保留最新
- markdown:默认不归档
- suggestions JSON:保留少量最近记录或已应用记录
---
## 8. 非目标
本轮不做:
- 删除现有脚本
- 重写治理链路
- 一次性重构所有 review 文档
- 自动 apply 所有 suggestions
- 引入更复杂的存储系统
本轮只做一件事:
> 把 keyword-cleanup 的产物语义分清,长期保留该留的,临时化该临时的。
---
## 9. 一句话结论
keyword-cleanup 应该收敛为:
- **事实层长期保留**:daily / term_stats
- **状态层长期保留**:interest / watch / alias / stopword / change_log
- **正式建议层轻量保留**:suggestions JSON
- **中间输入层与展示层临时化**:bundle / markdown
最终目标是让 reader 的关键词治理成为一个轻量、可持续、可回溯的偏好演化机制,而不是一个不断膨胀的中间文件系统。
@@ -0,0 +1,290 @@
# interest/watch 候选引擎改进方案
> 从固定阈值到自适应排位 + 趋势因子的演进
## 1. 背景
### 1.1 当前实现
`build_review_bundle.py` 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:
| 候选池 | 判断条件 | 依据 |
|--------|---------|------|
| `interest_review_candidates` | `total_count >= 3 AND days_seen >= 2` | `policy.interest_keyword_review` |
| `watch_review_candidates` | `total_count <= 2 AND days_seen <= 2` | `policy.watch_term_review` |
`generate_term_cleanup_suggestions.py` 则直接从这两个候选池过滤、去重、排序后输出。
### 1.2 当前方案的问题
**问题一:固定阈值不随数据量自适应**
```
场景 total_count=3 意味着什么
─────────────────────────────────────────────
7 天数据(~200 词) top 15%,有一定区分度 ✅
41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变
未来 200 天 仍然用 3 次,区分度稀释 ❌
```
同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。
**问题二:固定阈值忽略趋势信号**
- "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
- "Channels":total=3, recent=0 — 早期出现但近期消失
- 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。
**问题三:interest 和 watch 的分界线是硬的**
total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。
### 1.3 讨论结论
与老大讨论后确认:
1. interest/watch 是**统计判断**,不需要大模型介入,纯算法可以解决
2. 当前引擎缺的不是大模型,而是**算法本身没写完**——自适应维度(排位、趋势)还没实现
3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
4. 修改量小,可以在 1 小时内落地
---
## 2. 设计方案
### 2.1 核心思路
引入两个互补维度替代固定阈值:
```
判定维度 含义 数据来源
────────────────────────────────────────────────────────────
percentile(百分位排名) 该词 total_count 在所有词 term_stats
中的排位占比
growth(增速因子) 近期集中度 = recent_count daily 近 N 天
/ total_count
```
两个维度配合:
- **percentile** 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
- **growth** 衡量"这个词是持续出现还是近期爆发"——识别趋势信号
### 2.2 候选池划分逻辑
```
percentile
│
┌─────────────────────┐
│ top 5% │
│ → 建议 interest │ ← 高频稳定词
├─────────────────────┤
│ top 5%-20% │
│ → 建议 watch │ ← 有信号但未达 threshold
├─────────────────────┤
│ bottom 80% │
│ → 暂不处理 │ ← 噪声/低频
└─────────────────────┘
额外规则:
如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
→ 主动提升到 watch / 主动推 confirm
```
这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。
### 2.3 接口变化
**`configs/term_cleanup_policy.json`**:
```json
{
"schema_version": "v2",
"interest_keyword_review": {
"percentile_max": 0.05,
"growth_promotion": 0.5
},
"watch_term_review": {
"percentile_min": 0.05,
"percentile_max": 0.20
}
}
```
`v1` 的 `min_total_count`/`min_days_seen` 等绝对阈值字段不再使用。
**`build_review_bundle.py` 输出的候选项**:
```json
{
"term": "Anthropic",
"total_count": 13,
"days_seen": 10,
"percentile": 0.012,
"growth": 0.54,
"reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
}
```
### 2.4 不需要改动的部分
- `generate_term_cleanup_suggestions.py` — 它只消费候选池,不用改
- `apply_term_suggestions.py` — 消费 suggestions JSON,不用改
- `keyword-cleanup-bundle.json` 结构 — 向后兼容,新增 percentile/growth 字段
---
## 3. 实施计划
### 3.1 改动范围
| 文件 | 改动量 | 内容 |
|------|--------|------|
| `skills/keyword-cleanup-review/scripts/build_review_bundle.py` | ~40 行 | 新增 `_compute_percentile()` 和 `_compute_growth()` 函数;修改候选池生成逻辑;候选项中增加 percentile/growth |
| `configs/term_cleanup_policy.json` | ~10 行 | schema v2:percentile/growth 替代绝对阈值 |
### 3.2 实施步骤
1. **build_review_bundle.py**:在 `top_global_terms` 生成后,增加 percentile 计算函数和 growth 计算函数
2. **build_review_bundle.py**:修改 `interest_review_candidates` 和 `watch_review_candidates` 的生成逻辑,从固定阈值改为 percentile + growth
3. **build_review_bundle.py**:候选项增加 `percentile` 和 `growth` 字段,更新 `reason` 文案
4. **term_cleanup_policy.json**:更新为 v2 schema
5. **验证**:全量跑一次(`--days 365 --top 100`),对比新旧两份输出的差异
### 3.3 验证方法
```bash
# 1. 用旧版生成 baseline
cd /home/ubuntu/zhu/github/reader
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-baseline.json
# 2. 改代码后用新版生成
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-new.json
# 3. 对比 governance_hints
python3 -c "
import json
a = json.load(open('/tmp/bundle-baseline.json'))
b = json.load(open('/tmp/bundle-new.json'))
for key in ['interest_review_candidates', 'watch_review_candidates']:
old = set(i['term'] for i in a['governance_hints'][key])
new = set(i['term'] for i in b['governance_hints'][key])
print(f'{key}: 新增={new-old}, 减少={old-new}')
"
```
### 3.4 风险
| 风险 | 概率 | 应对 |
|------|------|------|
| 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 |
| growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 |
| 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 |
---
## 4. alias/stopword 设计方案
### 4.1 核心判断
alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。
| 类型 | 需要什么 | 判断方式 |
|------|---------|----------|
| 大小写变体 | 表层 | 规则:casefold 去重 |
| 单复数 | 表层 | 规则:去/加 s 后缀匹配 |
| 分词变体(空格/连字符) | 表层 | 规则:去空格归一 |
| 简写全称(MCP→Model Context Protocol) | **语义** | LLM |
| 中英文(上下文工程→Context Engineering) | **语义** | LLM |
| 同义不同名(Rush→猿辅导 Rush 平台) | **语义** | LLM |
| stopword(大模型、AI 太泛) | **语义** | LLM |
### 4.2 分层方案
```
输入:高频未覆盖词 + 已有 interest 词表
│
├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
│ 输出候选 alias 对
│
└── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
做语义聚类
输出 alias 组 + stopword 标记
```
### 4.3 规则层设计
在 `generate_term_cleanup_suggestions.py` 中新增 `_prepare_alias_suggestions()` 函数:
```python
def _prepare_alias_suggestions(top_terms, interest_keywords):
"""
基于表层规则生成 alias 建议。
规则1:casefold 匹配——同一个 casefold 下有多个原文变体
规则2:单复数——去掉/加上末尾 s 后匹配
规则3:分词变体——去空格/连字符后匹配
"""
```
优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。
### 4.4 LLM 层设计
单独脚本,非 generate 主链路的一部分。
```bash
python scripts/generate_term_cleanup_semantic_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
LLM prompt 设计:
```
你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
请做三件事:
1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest
用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地
```
LLM 层输出格式:
```json
{
"alias_suggestions": [
{"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
],
"stopword_suggestions": [
{"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
]
}
```
### 4.5 预期效果
| 覆盖类型 | 规则层 | LLM 层 |
|---------|--------|--------|
| 大小写变体 | ✅ | — |
| 单复数 | ✅ | — |
| 分词变体 | ✅ | — |
| 简写全称 | — | ✅ |
| 中英文映射 | — | ✅ |
| 同义不同名 | — | ✅ |
| stopword 判断 | — | ✅ |
---
## 5. 讨论记录
- 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
- 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
- 2026-05-14:确认本方案不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
- 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先
@@ -0,0 +1,534 @@
# keyword-cleanup-review 建议产物补齐设计
## 1. 背景与目标
当前仓库已经具备 keyword cleanup review 的大部分基础设施:
- 已有 review bundle 构建脚本 `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- 已有 term stats 与 daily term index 数据源
- 已有治理输入:`configs/term_cleanup_policy.json`、`configs/term_watchlist.json`、`configs/term_change_log.json`
- 已有建议落地脚本 `scripts/apply_term_suggestions.py`
当前缺口是:
> 缺少一层“把 `keyword-cleanup-bundle.json` 转成正式建议产物”的实现层。
也就是说,仓库现在能生成 review bundle,也能消费 suggestions JSON,但中间缺少稳定、可复用、可落盘的 suggestions 生成器。
本轮目标是补齐最小闭环,让仓库能够从 review bundle 稳定生成两份正式建议产物:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
并保证 JSON 与 `skills/keyword-cleanup-review/references/suggestion-schema.md` 对齐,且能直接衔接 `scripts/apply_term_suggestions.py`。
补充口径:本方案中的 JSON 是 review / apply 之间的唯一正式建议产物;bundle 与 Markdown 主要作为运行时工作文件和临时展示层,不建议与 facts/configs 一样长期沉淀。
---
## 2. 当前现状
### 2.1 已有输入层
`build_review_bundle.py` 已经把以下输入聚合成单个 bundle:
- `data/term_index/term_stats.json`
- `data/term_index/daily/*.json`
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/filter_context.personal.json`
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
bundle 中已经包含:
- 当前配置快照
- 最近 N 天热点词
- uncovered terms
- interest review candidates
- watch review candidates
这些信息已经足够支撑“保守的、可审查的” suggestions 生成。
### 2.2 已有输出消费层
`scripts/apply_term_suggestions.py` 已经能消费 suggestions JSON,并将接受的建议写回:
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/filter_context.personal.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
这说明落地层已存在,缺的是中间的正式建议产物生成层。
---
## 3. 当前缺口
当前流程停在:
`build_review_bundle.py` -> `keyword-cleanup-bundle.json`
但缺少:
`keyword-cleanup-bundle.json` -> `term-cleanup-suggestions-YYYY-MM-DD.json/.md`
因此出现几个问题:
- README / 设计文档里已经引用 suggestions 产物,但仓库内没有稳定生成脚本
- 人工审阅与后续 apply 之间没有统一的正式交付格式
- 同一份 bundle 无法稳定、幂等地重放为同名 suggestions 产物
- alias / stopword / interest / watch 四类建议缺少统一出入口
---
## 4. 推荐最小闭环架构
推荐新增一层独立脚本:
- `scripts/generate_term_cleanup_suggestions.py`
职责:
- 输入:`outputs/term_index/review/keyword-cleanup-bundle.json`
- 输出:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
推荐最小数据流:
1. `build_review_bundle.py` 生成 bundle
2. `generate_term_cleanup_suggestions.py` 读取 bundle
3. 脚本基于治理 hints 生成 suggestions JSON
4. 同时渲染人类可审阅的 Markdown
5. 审阅后可用 `apply_term_suggestions.py` 选择性写回配置
本轮不引入默认 LLM 路径:
- 默认实现采用确定性规则生成
- 如果未来需要 LLM 参与,应作为显式可选增强,而不是默认路径
---
## 5. 产物设计
### 5.1 JSON 产物
JSON 必须与 `suggestion-schema.md` 对齐,至少包含:
```json
{
"date": "2026-04-08",
"based_on_days": 7,
"alias_suggestions": [],
"stopword_suggestions": [],
"interest_keyword_suggestions": [
{
"term": "Claude Code",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered."
}
],
"watch_terms": [
{
"term": "A2A",
"reason": "Falls into the configured watch-term review range and should be observed first."
}
]
}
```
在不破坏兼容性的前提下,可以补充少量元数据字段,建议仅限:
- `source_bundle`
- `policy_schema_version`
- `summary`
建议项字段口径:
- `alias_suggestions[]`
- `from`
- `to`
- `reason`
- `stopword_suggestions[]`
- `term`
- `reason`
- `interest_keyword_suggestions[]`
- `term`
- `reason`
- 可选:`total_count`、`days_seen`、`recent_count`
- `watch_terms[]`
- `term`
- `reason`
- 可选:`total_count`、`days_seen`、`recent_count`
兼容性要求:
- `apply_term_suggestions.py` 只依赖分类 bucket 与关键字段名
- 因此额外证据字段只能追加,不能替换现有字段名
### 5.2 Markdown 产物
Markdown 推荐结构:
1. 标题与日期
2. 输入 bundle 与策略摘要
3. 当前现状摘要
- top/global 观察
- uncovered terms 概览
- 当前 watchlist / interest 覆盖情况
4. 建议摘要
- interest 建议数量
- watch 建议数量
- alias 建议数量
- stopword 建议数量
5. `interest_keyword_suggestions`
6. `watch_terms`
7. `alias_suggestions`
8. `stopword_suggestions`
9. 应用方式
- 指向生成的 JSON
- 给出 `apply_term_suggestions.py` 的调用示例
这样可以保证:
- 人可以直接审阅
- 机器可以直接消费同名 JSON
- Markdown 与 JSON 始终一一对应
---
## 6. 建议生成策略
### 6.1 本轮主链路:interest / watch
本轮先实现最小可用主链路:
- `interest_keyword_suggestions`
- `watch_terms`
直接复用 bundle 中已有的:
- `governance_hints.interest_review_candidates`
- `governance_hints.watch_review_candidates`
原因:
- 这些候选已经与 policy 对齐
- 这些候选已经排除了大部分已覆盖项
- 能直接与现有 apply 脚本形成闭环
### 6.2 alias / stopword 保守处理
本轮边界明确如下:
- `alias_suggestions` 先保持保守,默认可为空
- `stopword_suggestions` 先保持保守,默认可为空
- 后续如果补充更强证据或人工审查规则,再逐步增强
这样可以避免在证据不足时误伤配置。
### 6.3 Phase 2 新方向:alias review 交给 LLM 整理
对于 alias,不再优先走程序规则匹配。
Phase 2 建议改为:
- 程序继续负责准备 review 输入(term stats / daily / current aliases / stopwords / interest / watchlist)
- LLM 负责整理 alias 候选
- 默认先输出人工审阅汇报,而不是直接 apply
- 人工确认后,再决定是否写入 `term_aliases.json`
这样做的原因:
- alias 更偏语义整理,而不是简单趋势筛选
- 与 watch / interest 相比,alias 一旦错误归并,代价更高
- 对当前低频治理场景来说,LLM + 人工确认更轻,也比在程序里持续堆复杂规则更合适
---
## 7. 错误处理
脚本应做显式校验,并在失败时给出明确错误:
### 7.1 输入错误
- bundle 文件不存在 -> 直接失败
- bundle 不是 JSON object -> 直接失败
- 缺少关键字段(如 `days`、`governance_hints`)-> 直接失败
- 候选 bucket 结构错误 -> 直接失败
### 7.2 输出错误
- 输出目录不存在时自动创建
- JSON / Markdown 写入失败时直接退出非 0
### 7.3 数据去重与冲突
- 同一 term 不能同时出现在 interest 与 watch 中
- 优先级:`interest_keyword_suggestions` > `watch_terms`
- 已在 bundle 当前配置中覆盖的 term 不重复输出
---
## 8. 幂等性
本轮要求具备基础幂等性:
- 同一份 bundle 多次运行,默认生成同名产物
- 同一份 bundle 多次运行,JSON 内容顺序稳定
- Markdown 内容顺序稳定
建议做法:
- 优先使用 bundle 的 `generated_at` 日期作为 suggestions 文件日期
- term 排序按证据强度与 term 名稳定排序
- 不在默认输出中写入“每次运行变化”的当前时间戳
这样可以让生成器作为可重放步骤存在于 review 流程中。
---
## 9. 与 apply_term_suggestions.py 的衔接
正式链路应变成:
1. `build_review_bundle.py`
2. `generate_term_cleanup_suggestions.py`
3. 人工审阅 Markdown
4. `apply_term_suggestions.py --suggestions ...`
衔接要求:
- JSON bucket 名必须与 `apply_term_suggestions.py` 读取逻辑一致
- `date` 与 `based_on_days` 字段保留,用于 change log 回写
- 建议项中的 `reason` 直接沿用到 apply 后的 change log
这保证建议生成层不会成为孤立产物,而是正式进入 repo 治理闭环。
---
## 10. Python 3.11 依赖处理
### 10.1 当前现状
`build_review_bundle.py` 当前使用 `from datetime import UTC`,这要求 Python 3.11。
仓库整体 `pyproject.toml` 当前也声明 `requires-python = ">=3.11"`,因此短期内使用 `/usr/bin/python3.11` 运行是符合仓库现状的。
### 10.2 短期建议
短期先在文档与验证命令中明确:
- bundle 构建使用 `/usr/bin/python3.11`
- suggestions 生成脚本也按仓库当前 3.11 基线运行
### 10.3 中期建议
如果后续希望把 keyword cleanup 工具链下探到 Python 3.10,可做兼容改造:
- 把 `datetime.UTC` 替换为 `datetime.timezone.utc`
- 重新检查相关脚本是否还有其他 3.11-only 语法或库依赖
本轮不做超范围兼容重构,只在文档中把此约束说清楚。
---
## 11. 边界与非目标
本轮明确边界:
- 先实现 interest/watch 主链路
- alias/stopword 先保持保守或留待后续增强
- 不做超范围重构
- 不把 LLM 作为默认生成路径
- 不直接改 `filter_rules.json`
- 不自动 apply 建议到配置
因此,本轮交付定义为:
- 补齐 bundle -> suggestions 的正式实现层
- 让 review 流程可运行、可落盘、可审阅、可应用
而不是一次性做完所有高级治理逻辑。
---
## 12. 实施建议
建议按以下小步落地:
### Phase 1(已完成)
1. 新增 `scripts/generate_term_cleanup_suggestions.py`
2. 读取 bundle 并做结构校验
3. 生成稳定排序的 interest/watch suggestions JSON
4. Markdown 改成按需生成
5. README / skill 文档补一条生成命令
6. 用 `/usr/bin/python3.11` 完整跑通 bundle -> suggestions
完成后,keyword cleanup review 的最小正式链路变为:
`review bundle` -> `suggestions json` -> `apply accepted suggestions`
其中 Markdown 只是按需生成的展示层。
### Phase 2(下一步)
1. 保持程序继续准备 review 输入
2. 引入 LLM 做 alias 候选整理
3. 默认先生成 alias review 汇报,而不是直接 apply
4. 由人工确认后再决定是否写入 `term_aliases.json`
这样 alias review 会成为一个低频治理动作,而不是主链路里的自动归一步骤。
### Phase 3(下一步)
1. 保持程序继续准备 review 输入
2. 引入 LLM 做 stopword 候选整理
3. 默认先生成 stopword review 汇报,而不是直接 apply
4. 由人工确认后再决定是否写入 `term_stopwords.json`
这样 stopword review 会成为一个低频减噪动作,而不是主链路里的自动过滤步骤。
#### Phase 3 输入建议
建议给 LLM 的输入包括:
- `data/term_index/term_stats.json` 中的高频词与 recent evidence
- 最近 N 天 `data/term_index/daily/*.json` 的热点词上下文
- 当前 `configs/term_stopwords.json`
- 当前 `configs/term_aliases.json`
- 当前 `configs/filter_context.personal.json` 中的 `interest_keywords`
- 当前 `configs/term_watchlist.json`
程序层只负责把这些输入整理成紧凑 review context,不负责直接做 stopword 决策。
#### Phase 3 输出建议
建议 LLM 默认输出一份人工审阅汇报,而不是直接写配置:
- `建议加入 stopword`
- 词
- 简短理由
- 证据(如 total_count / days_seen / recent_count)
- `暂不建议加入 stopword`
- 词
- 为什么虽然偏泛,但当前还不能杀
- `需要人工判断`
- 词
- 风险点:可能是噪声,也可能仍保留有价值信号
如需结构化输出,可额外补一份 `stopword_review_candidates.json`,但该文件只作为 review 输入,不直接作为自动 apply 指令。
#### Phase 3 审阅原则
- 宁可少删,不乱杀
- 优先处理过泛、低辨识度、持续污染统计的词
- 对可能仍承载有效技术语义的词保持保守
- 默认先汇报,确认后再执行
#### Phase 3 汇报模板建议
建议 stopword review 默认按以下结构汇报给用户:
1. `建议加入 stopword`
- 词
- 理由:为什么这个词对治理帮助低、噪声高
- 证据:`total_count` / `days_seen` / `recent_count` 或最近出现上下文
2. `暂不建议加入 stopword`
- 词
- 理由:为什么当前不建议删掉
3. `需要人工判断`
- 词
- 风险点:泛词与有效主题词之间边界不清等
推荐汇报风格:
- 简短、保守、可审阅
- 先给判断,再给证据
- 不输出机器式原始 dump
- 不默认承诺“已应用”,只汇报“建议”
#### Phase 2 输入建议
建议给 LLM 的输入包括:
- `data/term_index/term_stats.json` 中的高频词与 recent evidence
- 最近 N 天 `data/term_index/daily/*.json` 的热点词上下文
- 当前 `configs/term_aliases.json`
- 当前 `configs/term_stopwords.json`
- 当前 `configs/filter_context.personal.json` 中的 `interest_keywords`
- 当前 `configs/term_watchlist.json`
程序层只负责把这些输入整理成紧凑 review context,不负责直接做 alias 决策。
#### Phase 2 输出建议
建议 LLM 默认输出一份人工审阅汇报,而不是直接写配置:
- `建议合并`
- `from -> to`
- 简短理由
- 证据(如 total_count / days_seen / recent_count)
- `暂不建议合并`
- 为什么不建议并掉
- `需要人工判断`
- 语义相近但风险较高的项
如需结构化输出,可额外补一份 `alias_review_candidates.json`,但该文件只作为 review 输入,不直接作为自动 apply 指令。
#### Phase 2 审阅原则
- 宁可少提,不乱提
- 优先整理明显同义 / 同概念 / 词形差异
- 不把公司名、产品名、泛概念词强行混并
- 默认先汇报,确认后再执行
#### Phase 2 汇报模板建议
建议 alias review 默认按以下结构汇报给用户:
1. `建议合并`
- `from -> to`
- 理由:为什么判断为同一概念或更合适的标准词
- 证据:`total_count` / `days_seen` / `recent_count` 或最近出现上下文
2. `暂不建议合并`
- 候选对
- 理由:为什么虽然相近,但当前不建议并
3. `需要人工判断`
- 候选对
- 风险点:歧义、范围差异、产品名/公司名混淆等
推荐汇报风格:
- 简短、保守、可审阅
- 先给判断,再给证据
- 不输出机器式原始 dump
- 不默认承诺“已应用”,只汇报“建议”
#### Phase 2 示例输出
建议合并:
- `Claude code -> Claude Code`
- 理由:明显属于同一产品名,仅是大小写写法不一致。
- 证据:`Claude Code` 在最近多日持续出现,而小写写法只是在少量上下文中作为变体出现。
- `Sub-Agent -> SubAgent`
- 理由:更像词形差异,不构成新的独立概念。
- 证据:两者都围绕同一 agent 架构语境出现,且没有稳定区分语义。
暂不建议合并:
- `Skills ↔ Agent Skills`
- 理由:前者过泛,后者更具体,当前强行归并会损失粒度。
- `Anthropic ↔ Claude`
- 理由:公司名与产品名并不等价,不应直接视为一个关键词。
需要人工判断:
- `AI助手 ↔ AI Agent`
- 风险点:语义可能接近,但中文表述范围更宽,是否并入需要结合你的使用语境判断。
+34
View File
@@ -0,0 +1,34 @@
# 示例文章标题
原文链接:
https://example.com/article
## 核心结论
这里先用一段短句概括最重要的判断。
如果结论较长,继续拆成第二个短段,而不是塞成一个大长段。
## 主要论点
先交代文章的核心主张。
再单独起一段解释支撑这个主张的关键论据。
如果还有补充判断,继续拆段,保证在 IMA 中阅读时不会挤成一整坨。
## 关键方法 / 机制
- 要点一
- 要点二
- 要点三
## 重要细节
- 细节一
- 细节二
## 可复用启发
- 启发一
- 启发二
+40
View File
@@ -0,0 +1,40 @@
+++
title = "AI 日报 · 示例"
date = 2026-04-01T16:55:00+08:00
summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实践的当日观察。"
+++
> ⚠️ 格式规范(生成 Hugo 时必须遵守):
> - 文章编号:`1.` `2.` `3.`(阿拉伯数字 + 点),禁止 `① ② ③` / `一、二、三` 等变体
> - 四个 section 缺一不可:`今日概览` → `今日重点` → `趋势观察` → `延伸阅读`
> - 每篇文章结构:标题来源 → 摘要段 → "值得关注:"三点 → "这篇更值得关注的理由"段
# 今日概览
今天的公开候选主要集中在 AI Agent 的架构演进、工具化落地与工程化实践三条线索上。相比早期偏概念展示的讨论,这一批内容更强调模块化能力栈、真实部署路径与系统可维护性,说明行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
## 今日重点
### 1. 学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变
来源:阿里云开发者
文章分析了 AI 智能体架构从单体 Agent 向模块化 Skills 的范式转变。Anthropic 先后推出 MCP 和 Agent Skills 开放标准,构建了知识、工具、协作和运行分层架构。文章通过一个自动化美化相册的真实项目,对比了 Claude Code 与 OpenClaw 两种实现方案,验证了新架构的可复用性与灵活性。
值得关注:
- Anthropic 在 14 个月内先后推出 MCP 和 Agent Skills 两个开放标准,推动 AI 智能体架构分层化。
- 新范式核心是构建薄 Agent 引擎与可组合的 Skills 库,取代为每个用例定制单体 Agent。
- 文章通过自动化美化相册项目,实操演示了 Skills、MCP、OpenClaw 和 A2A 协议如何协同工作。
这篇内容更值得关注的原因在于,它不只是提出了“Agent 要模块化”这个判断,而是把开放标准、分层架构和真实项目案例串成了一条完整论证链,能直接支撑今天日报的主线。
## 趋势观察
1. Agent 正在从单体能力转向可组合的模块化体系。无论是 Skills、MCP、记忆还是运行时编排,这批内容都在强调解耦与复用,而不是把智能体继续当成一个不可拆分的黑箱。
2. 工程化正在变成 AI 应用竞争的主战场。桌面 Agent、企业级架构和部署实践类内容增多,说明真正的差异化开始落在接入现有流程、控制风险和提升可维护性上。
3. AI 能力的竞争点正在上移。模型本身仍重要,但真正可持续的优势越来越来自系统设计、工作流整合和对业务场景的理解。
## 延伸阅读
- [学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变](https://example.com/a)|阿里云开发者
- [Agent Skills:打通可复用专业领域知识的最后一公里](https://example.com/b)|阿里云开发者
- [CoPaw深度解析:源码架构和功能实践](https://example.com/c)|阿里云开发者
+314
View File
@@ -0,0 +1,314 @@
#!/usr/bin/env python3
"""
Generate semantic keyword suggestions using LLM.
Covers what surface-form rules cannot:
- semantic alias (abbreviation ↔ full name, Chinese ↔ English, synonym)
- stopword (overly broad / low-discrimination terms)
- promote (new term that aligns with user's focus areas)
Usage:
python scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-2026-05-14.json
"""
from __future__ import annotations
import argparse
import json
import os
import re
import sys
import time
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
from urllib.request import Request, urlopen
REPO_ROOT = Path(__file__).resolve().parents[1]
DEFAULT_BUNDLE_PATH = REPO_ROOT / "outputs" / "term_index" / "review" / "keyword-cleanup-bundle.json"
DEFAULT_OUTPUT_DIR = REPO_ROOT / "outputs" / "term_index" / "review"
def _load_json(path: Path) -> Any:
return json.loads(path.read_text(encoding="utf-8-sig"))
def _save_json(path: Path, payload: dict[str, Any]) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def _load_env(path: Path) -> dict[str, str]:
"""Load key=value pairs from .env file."""
env: dict[str, str] = {}
if not path.exists():
return env
for line in path.read_text(encoding="utf-8").splitlines():
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
env[key.strip()] = value.strip().strip("\"'")
return env
def _build_prompt(
interest_keywords: list[str],
rule_alias_suggestions: list[dict[str, str]],
candidate_terms: list[dict[str, Any]],
relevant_watch_terms: list[dict[str, Any]],
) -> str:
"""Build the LLM prompt for semantic suggestions."""
interest_bullets = "\n".join(f" - {t}" for t in sorted(interest_keywords))
candidate_bullets = "\n".join(
f" - {t['term']} (count={t['total_count']}, days={t['days_seen']})"
for t in candidate_terms[:40]
)
# Alias from rule layer (for LLM to build on, not duplicate)
rule_alias_text = ""
if rule_alias_suggestions:
rule_alias_text = "\nSurface-form alias (already identified, skip these):\n" + "\n".join(
f" {a['from']} → {a['to']} ({a['reason']})"
for a in rule_alias_suggestions
)
watch_text = ""
if relevant_watch_terms:
watch_text = "\nWatch terms (low-frequency but potentially relevant):\n" + "\n".join(
f" {t['term']} (count={t['total_count']}, days={t['days_seen']})"
for t in relevant_watch_terms[:20]
)
return f"""You are a keyword governance assistant for an AI engineer. Your job is to analyze keyword data and produce structured suggestions.
## User's focus areas
- AI Agent engineering (Skills, Harness, MCP, Agent architecture)
- Backend engineering (Java, Go, Kubernetes, MySQL, distributed systems)
- Open source AI tools and practices (Claude Code, Cursor, DeepSeek, OpenClaw)
- LLM application engineering (context engineering, RAG, prompt engineering)
## Interest keywords (52 already configured)
{interest_bullets}
## Uncovered candidate terms (sorted by frequency)
{candidate_bullets}
{watch_text}{rule_alias_text}
## Task
Analyze the candidate terms and output a JSON object with exactly three keys:
1. "semantic_alias": array of alias suggestions that SURFACE RULES CAN'T CATCH (e.g. abbreviation↔full name, Chinese↔English, different naming for the same concept).
Format: [{{"from": "<variant>", "to": "<canonical interest keyword>", "reason": "<why>"}}]
2. "stopword": array of terms that are too broad/generic to be useful as filters. A stopword is a term that appears frequently but has LOW DISCRIMINATION — it matches too many unrelated articles and clutters the keyword index.
Format: [{{"term": "<term>", "reason": "<why it should be a stopword>"}}]
3. "promote_to_interest": array of uncovered terms that align well with the user's focus areas and should be added as interest keywords.
Format: [{{"term": "<term>", "reason": "<why it fits>"}}]
## Rules
- Be conservative. When in doubt, leave it out.
- Only suggest alias for terms that clearly refer to the SAME concept as an existing interest keyword.
- Only suggest stopword for terms that are genuinely too broad (appear in many unrelated contexts).
- Only suggest promote for terms that clearly match the user's stated focus areas.
- Output valid JSON only, no markdown, no explanation outside the JSON."""
def _call_llm(prompt: str, api_url: str, model: str, api_key: str) -> str:
"""Call LLM API and return the response text."""
payload = json.dumps({
"model": model,
"messages": [{"role": "user", "content": prompt}],
"temperature": 0.1,
"max_tokens": 2048,
}).encode("utf-8")
req = Request(
api_url.rstrip("/") + "/chat/completions",
data=payload,
headers={
"Content-Type": "application/json",
"Authorization": f"Bearer {api_key}",
},
)
max_retries = 3
for attempt in range(max_retries):
try:
with urlopen(req, timeout=120) as resp:
result = json.loads(resp.read().decode("utf-8"))
return result["choices"][0]["message"]["content"]
except Exception as e:
if attempt < max_retries - 1:
wait = 2 ** attempt
print(f" LLM call failed (attempt {attempt+1}/{max_retries}): {e}", file=sys.stderr)
print(f" Retrying in {wait}s...", file=sys.stderr)
time.sleep(wait)
else:
raise
def _parse_llm_response(text: str) -> dict[str, list[dict[str, str]]]:
"""Extract JSON from LLM response (may contain markdown fences)."""
# Try to find JSON block
json_match = re.search(r"```(?:json)?\s*\n?(\{.*?\})\s*\n?```", text, re.DOTALL)
if json_match:
text = json_match.group(1)
# Clean up: remove any text before { or after }
start = text.find("{")
end = text.rfind("}")
if start >= 0 and end > start:
text = text[start : end + 1]
try:
result = json.loads(text)
except json.JSONDecodeError:
# Try partial recovery
print(f" Warning: LLM response not clean JSON, attempting recovery", file=sys.stderr)
print(f" Raw: {text[:500]}", file=sys.stderr)
return {"semantic_alias": [], "stopword": [], "promote_to_interest": []}
# Normalize keys
normalized = {
"semantic_alias": result.get("semantic_alias", result.get("alias", [])),
"stopword": result.get("stopword", result.get("stopword_suggestions", [])),
"promote_to_interest": result.get("promote_to_interest", result.get("promote", [])),
}
# Ensure each is a list
for key in normalized:
if not isinstance(normalized[key], list):
normalized[key] = []
return normalized
def main() -> None:
parser = argparse.ArgumentParser(description="Generate semantic keyword suggestions via LLM.")
parser.add_argument("--bundle", type=Path, default=DEFAULT_BUNDLE_PATH, help="Review bundle JSON path")
parser.add_argument("--suggestions", type=Path, default=None, help="Existing suggestions JSON (for rule alias context)")
parser.add_argument("--output", type=Path, default=None, help="Output JSON path (auto-generated if omitted)")
parser.add_argument("--llm-api-url", type=str, default=None, help="LLM API base URL")
parser.add_argument("--llm-model", type=str, default=None, help="LLM model name")
parser.add_argument("--llm-api-key", type=str, default=None, help="LLM API key")
parser.add_argument("--dry-run", action="store_true", help="Print prompt and exit without calling LLM")
args = parser.parse_args()
# Load config
env_path = REPO_ROOT / ".env"
env = _load_env(env_path) if env_path.exists() else {}
api_url = args.llm_api_url or os.environ.get("LLM_API_URL") or env.get("LLM_API_URL", "https://api.deepseek.com")
# Map OpenClaw model aliases to actual API model names
model_raw = args.llm_model or os.environ.get("LLM_MODEL") or env.get("LLM_MODEL", "deepseek-chat")
MODEL_ALIAS_MAP = {
"deepseek/deepseek-v4-flash": "deepseek-chat",
"deepseek/deepseek-chat": "deepseek-chat",
"deepseek-v4-flash": "deepseek-chat",
"deepseek-chat": "deepseek-chat",
}
model = MODEL_ALIAS_MAP.get(model_raw, model_raw)
api_key = args.llm_api_key or os.environ.get("LLM_API_KEY") or env.get("LLM_API_KEY", "")
if not api_key:
print("Error: No LLM API key found. Set LLM_API_KEY in .env or pass --llm-api-key.", file=sys.stderr)
sys.exit(1)
# Load bundle
if not args.bundle.exists():
print(f"Error: Bundle not found: {args.bundle}", file=sys.stderr)
sys.exit(1)
bundle = _load_json(args.bundle)
current_config = bundle.get("current_config", {})
interest_keywords = current_config.get("interest_keywords", [])
top_global_terms = bundle.get("top_global_terms", [])
governance_hints = bundle.get("governance_hints", {})
# Build candidate list (uncovered terms from interest + watch candidates)
candidate_terms = []
for item in governance_hints.get("interest_review_candidates", []):
if isinstance(item, dict):
candidate_terms.append({
"term": item.get("term", ""),
"total_count": item.get("total_count", 0),
"days_seen": item.get("days_seen", 0),
"percentile": item.get("percentile", 0),
"growth": item.get("growth", 0),
})
for item in governance_hints.get("watch_review_candidates", []):
if isinstance(item, dict):
# Avoid duplicates
if not any(c["term"] == item.get("term") for c in candidate_terms):
candidate_terms.append({
"term": item.get("term", ""),
"total_count": item.get("total_count", 0),
"days_seen": item.get("days_seen", 0),
"percentile": item.get("percentile", 0),
"growth": item.get("growth", 0),
})
# Sort by total_count descending
candidate_terms.sort(key=lambda x: -x["total_count"])
relevant_watch_terms = governance_hints.get("watch_review_candidates", [])[:20]
# Load rule-layer alias suggestions if available
rule_alias = []
if args.suggestions and args.suggestions.exists():
s = _load_json(args.suggestions)
rule_alias = s.get("alias_suggestions", [])
# Build prompt
prompt = _build_prompt(
interest_keywords=interest_keywords,
rule_alias_suggestions=rule_alias,
candidate_terms=candidate_terms,
relevant_watch_terms=relevant_watch_terms,
)
# Determine output path
suggestion_date = datetime.now(timezone.utc).date().isoformat()
output_path = args.output or (DEFAULT_OUTPUT_DIR / f"term-cleanup-semantic-suggestions-{suggestion_date}.json")
if args.dry_run:
print("=== DRY RUN: Prompt ===")
print(prompt)
print("\n=== END ===")
print(f"\nWould write to: {output_path}")
return
# Call LLM
print(f"Calling LLM ({model})...", file=sys.stderr)
response = _call_llm(prompt, api_url, model, api_key)
print(f"LLM response received ({len(response)} chars)", file=sys.stderr)
# Parse
parsed = _parse_llm_response(response)
# Build output
output = {
"date": suggestion_date,
"source_bundle": str(args.bundle),
"model": model,
"interest_keyword_count": len(interest_keywords),
"candidate_count": len(candidate_terms),
**parsed,
}
_save_json(output_path, output)
summary = {
"output": str(output_path),
"semantic_alias": len(output.get("semantic_alias", [])),
"stopword": len(output.get("stopword", [])),
"promote_to_interest": len(output.get("promote_to_interest", [])),
}
print(json.dumps(summary, ensure_ascii=False, indent=2))
if __name__ == "__main__":
main()
@@ -0,0 +1,536 @@
from __future__ import annotations
import argparse
import json
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
REPO_ROOT = Path(__file__).resolve().parents[1]
DEFAULT_BUNDLE_PATH = REPO_ROOT / "outputs" / "term_index" / "review" / "keyword-cleanup-bundle.json"
DEFAULT_OUTPUT_DIR = REPO_ROOT / "outputs" / "term_index" / "review"
def _load_json(path: Path) -> Any:
return json.loads(path.read_text(encoding="utf-8-sig"))
def _save_json(path: Path, payload: dict[str, Any]) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def _save_text(path: Path, content: str) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(content, encoding="utf-8")
def _term_key(value: str) -> str:
return value.strip().casefold()
def _utc_today() -> str:
return datetime.now(timezone.utc).date().isoformat()
def _require_dict(payload: Any, name: str) -> dict[str, Any]:
if not isinstance(payload, dict):
raise RuntimeError(f"{name} must be a JSON object.")
return payload
def _require_list(payload: Any, name: str) -> list[Any]:
if not isinstance(payload, list):
raise RuntimeError(f"{name} must be a JSON array.")
return payload
def _bundle_date(bundle: dict[str, Any]) -> str:
generated_at = bundle.get("generated_at")
if isinstance(generated_at, str) and generated_at.strip():
normalized = generated_at.replace("Z", "+00:00")
try:
return datetime.fromisoformat(normalized).date().isoformat()
except ValueError:
pass
return _utc_today()
def _recent_count_map(top_global_terms: list[dict[str, Any]]) -> dict[str, int]:
counts: dict[str, int] = {}
for item in top_global_terms:
term = item.get("term")
recent_count = item.get("recent_count")
if isinstance(term, str) and isinstance(recent_count, int):
counts[term] = recent_count
return counts
def _covered_term_sets(bundle: dict[str, Any]) -> tuple[set[str], set[str], set[str]]:
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
interest_keywords = _require_list(current_config.get("interest_keywords"), "bundle.current_config.interest_keywords")
stopwords = _require_list(current_config.get("stopwords"), "bundle.current_config.stopwords")
watchlist = _require_list(current_config.get("watchlist"), "bundle.current_config.watchlist")
interest_set = {_term_key(item) for item in interest_keywords if isinstance(item, str) and item.strip()}
stopword_set = {_term_key(item) for item in stopwords if isinstance(item, str) and item.strip()}
watch_set = {
_term_key(str(item.get("term", "")))
for item in watchlist
if isinstance(item, dict) and isinstance(item.get("term"), str) and str(item.get("term", "")).strip()
}
return interest_set, stopword_set, watch_set
def _sort_key(item: dict[str, Any]) -> tuple[int, int, int, str, str]:
total_count = int(item.get("total_count") or 0)
days_seen = int(item.get("days_seen") or 0)
recent_count = int(item.get("recent_count") or 0)
term = str(item.get("term") or "")
return (-total_count, -days_seen, -recent_count, term.casefold(), term)
def _prepare_interest_suggestions(bundle: dict[str, Any]) -> list[dict[str, Any]]:
governance_hints = _require_dict(bundle.get("governance_hints"), "bundle.governance_hints")
candidates = _require_list(
governance_hints.get("interest_review_candidates"),
"bundle.governance_hints.interest_review_candidates",
)
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
recent_counts = _recent_count_map([item for item in top_global_terms if isinstance(item, dict)])
interest_set, stopword_set, watch_set = _covered_term_sets(bundle)
suggestions: list[dict[str, Any]] = []
seen: set[str] = set()
for item in candidates:
if not isinstance(item, dict):
continue
term = item.get("term")
if not isinstance(term, str) or not term.strip():
continue
term_key = _term_key(term)
if term_key in seen or term_key in interest_set or term_key in stopword_set:
continue
total_count = int(item.get("total_count") or 0)
days_seen = int(item.get("days_seen") or 0)
recent_count = recent_counts.get(term, 0)
base_reason = str(item.get("reason") or "Meets the configured interest-keyword review threshold.")
if term_key in watch_set:
base_reason += " It is currently in watchlist and is ready for promotion."
reason = f"{base_reason} Evidence: total_count={total_count}, days_seen={days_seen}, recent_count={recent_count}."
suggestions.append(
{
"term": term,
"reason": reason,
"total_count": total_count,
"days_seen": days_seen,
"recent_count": recent_count,
}
)
seen.add(term_key)
suggestions.sort(key=_sort_key)
return suggestions
def _prepare_watch_suggestions(bundle: dict[str, Any], reserved_terms: set[str]) -> list[dict[str, Any]]:
governance_hints = _require_dict(bundle.get("governance_hints"), "bundle.governance_hints")
candidates = _require_list(
governance_hints.get("watch_review_candidates"),
"bundle.governance_hints.watch_review_candidates",
)
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
recent_counts = _recent_count_map([item for item in top_global_terms if isinstance(item, dict)])
interest_set, stopword_set, watch_set = _covered_term_sets(bundle)
suggestions: list[dict[str, Any]] = []
seen: set[str] = set(reserved_terms)
for item in candidates:
if not isinstance(item, dict):
continue
term = item.get("term")
if not isinstance(term, str) or not term.strip():
continue
term_key = _term_key(term)
if term_key in seen or term_key in interest_set or term_key in stopword_set or term_key in watch_set:
continue
total_count = int(item.get("total_count") or 0)
days_seen = int(item.get("days_seen") or 0)
recent_count = recent_counts.get(term, 0)
base_reason = str(item.get("reason") or "Falls into the configured watch-term review range.")
reason = f"{base_reason} Evidence: total_count={total_count}, days_seen={days_seen}, recent_count={recent_count}."
suggestions.append(
{
"term": term,
"reason": reason,
"total_count": total_count,
"days_seen": days_seen,
"recent_count": recent_count,
}
)
seen.add(term_key)
suggestions.sort(key=_sort_key)
return suggestions
def _prepare_alias_suggestions(
bundle: dict[str, Any],
all_terms: list[dict[str, Any]] | None = None,
) -> list[dict[str, Any]]:
"""
Generate alias suggestions using surface-form rules (no LLM).
Rules:
1. casefold match — same normalized form, different original casing
2. trailing-s singularization — singular/plural variants
3. whitespace/hyphen normalization — word boundary variants
Scans all_terms (full term_stats) if provided; otherwise falls back
to top_global_terms from the bundle.
"""
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
interest_keywords = _require_list(
current_config.get("interest_keywords"), "bundle.current_config.interest_keywords"
)
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
source_terms = all_terms if all_terms is not None else top_global_terms
interest_set = {_term_key(t) for t in interest_keywords if isinstance(t, str)}
interest_originals: set[str] = {t for t in interest_keywords if isinstance(t, str)}
# Build full casefold → [original forms] map
cf_map: dict[str, list[str]] = {}
for item in source_terms:
term = None
if isinstance(item, dict):
term = item.get("term")
elif isinstance(item, str):
term = item
if not isinstance(term, str) or not term.strip():
continue
key = _term_key(term)
if key not in cf_map:
cf_map[key] = []
if term not in cf_map[key]:
cf_map[key].append(term)
suggestions: list[dict[str, Any]] = []
seen_pairs: set[tuple[str, str]] = set()
def _add(from_term: str, to_term: str, reason: str) -> None:
pair = (_term_key(from_term), _term_key(to_term))
if pair in seen_pairs:
return
seen_pairs.add(pair)
suggestions.append({"from": from_term, "to": to_term, "reason": reason})
# Build a set of all term keys from source for quick lookup
source_keys = set(cf_map.keys())
# Rule 1: casefold match — same normalized form, different casing
for key, variants in cf_map.items():
if len(variants) < 2:
continue
canonical = None
alt_forms = []
for v in variants:
if v in interest_originals:
canonical = v
else:
alt_forms.append(v)
if canonical and alt_forms:
for alt in alt_forms:
_add(alt, canonical, "Case variant")
elif len(variants) >= 2 and not canonical:
# None is canonical — suggest the highest-frequency form
ranked = sorted(variants, key=lambda t: -(
next(
(it.get("total_count", 0) for it in top_global_terms if it.get("term") == t),
0,
)
))
for alt in ranked[1:]:
_add(alt, ranked[0], "Case variant (auto-ranked)")
# Rule 2: singular/plural — trailing-s normalization
# Check all source terms (not just interest keys) for bidirectional matching
for key in source_keys:
if key in interest_set:
continue
if key.endswith("s") and len(key) > 2:
singular_key = key.rstrip("s")
if singular_key in interest_set and singular_key != key:
# Find canonical interest keyword
canon = next((t for t in interest_keywords if _term_key(t) == singular_key), None)
from_form = cf_map[key][0]
if canon:
_add(from_form, canon, "Plural variant")
# singular form → interest has plural
plural_key = key + "s"
if plural_key in interest_set and plural_key != key:
canon = next((t for t in interest_keywords if _term_key(t) == plural_key), None)
from_form = cf_map[key][0]
if canon:
_add(from_form, canon, "Singular variant")
# Rule 3: whitespace/hyphen normalization
for key in source_keys:
if key in interest_set:
continue
normalized = key.replace("-", "").replace("_", "").replace(" ", "")
if normalized in interest_set and normalized != key:
canon = next((t for t in interest_keywords if _term_key(t) == normalized), None)
from_form = cf_map[key][0]
if canon:
_add(from_form, canon, "Whitespace/punctuation variant")
suggestions.sort(key=lambda x: (x["from"].casefold(), x["to"].casefold()))
return suggestions
def _render_table(items: list[dict[str, Any]]) -> str:
if not items:
return "_None in this pass._\n"
lines = [
"| Term | Total | Days | Recent | Reason |",
"| --- | ---: | ---: | ---: | --- |",
]
for item in items:
term = str(item.get("term") or "")
total_count = int(item.get("total_count") or 0)
days_seen = int(item.get("days_seen") or 0)
recent_count = int(item.get("recent_count") or 0)
reason = str(item.get("reason") or "").replace("|", "\\|")
lines.append(f"| {term} | {total_count} | {days_seen} | {recent_count} | {reason} |")
return "\n".join(lines) + "\n"
def _render_simple_table(items: list[dict[str, Any]], first_column: str) -> str:
if not items:
return "_None in this pass._\n"
lines = [
f"| {first_column} | Reason |",
"| --- | --- |",
]
for item in items:
value = str(item.get(first_column.casefold()) or item.get(first_column) or "")
reason = str(item.get("reason") or "").replace("|", "\\|")
lines.append(f"| {value} | {reason} |")
return "\n".join(lines) + "\n"
def _render_markdown(
*,
suggestion_date: str,
bundle_path: Path,
json_output_path: Path,
bundle: dict[str, Any],
suggestions: dict[str, Any],
) -> str:
policy = _require_dict(bundle.get("policy"), "bundle.policy")
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
uncovered_terms = _require_list(bundle.get("uncovered_terms"), "bundle.uncovered_terms")
top_preview = [item for item in top_global_terms if isinstance(item, dict)][:5]
uncovered_preview = [item for item in uncovered_terms if isinstance(item, dict)][:5]
interest_items = suggestions["interest_keyword_suggestions"]
watch_items = suggestions["watch_terms"]
alias_items = suggestions["alias_suggestions"]
stopword_items = suggestions["stopword_suggestions"]
lines = [
f"# Term Cleanup Suggestions - {suggestion_date}",
"",
"## Review Context",
"",
f"- Source bundle: `{bundle_path}`",
f"- Suggestions JSON: `{json_output_path}`",
f"- Bundle generated_at: `{bundle.get('generated_at', 'unknown')}`",
f"- Based on days: `{suggestions['based_on_days']}`",
f"- Policy schema version: `{policy.get('schema_version', 'unknown')}`",
"- Scope: implement `interest_keyword_suggestions` and `watch_terms` main path first; keep alias/stopword conservative in this pass.",
"",
"## Current State",
"",
f"- Interest keywords: `{current_config.get('interest_keyword_count', 0)}`",
f"- Watch terms: `{current_config.get('watch_term_count', 0)}`",
f"- Stopwords: `{current_config.get('stopword_count', 0)}`",
f"- Aliases: `{current_config.get('alias_count', 0)}`",
f"- Top global terms considered: `{len(top_global_terms)}`",
f"- Uncovered terms considered: `{len(uncovered_terms)}`",
"",
"### Top Terms Snapshot",
"",
]
if top_preview:
for item in top_preview:
lines.append(
f"- `{item.get('term', '')}`: total_count={item.get('total_count', 0)}, days_seen={item.get('days_seen', 0)}, recent_count={item.get('recent_count', 0)}"
)
else:
lines.append("- No top terms available.")
lines.extend([
"",
"### Uncovered Terms Snapshot",
"",
])
if uncovered_preview:
for item in uncovered_preview:
lines.append(
f"- `{item.get('term', '')}`: total_count={item.get('total_count', 0)}, days_seen={item.get('days_seen', 0)}, recent_count={item.get('recent_count', 0)}"
)
else:
lines.append("- No uncovered terms available.")
lines.extend([
"",
"## Suggestion Summary",
"",
f"- `interest_keyword_suggestions`: `{len(interest_items)}`",
f"- `watch_terms`: `{len(watch_items)}`",
f"- `alias_suggestions`: `{len(alias_items)}`",
f"- `stopword_suggestions`: `{len(stopword_items)}`",
"",
"## Interest Keyword Suggestions",
"",
_render_table(interest_items).rstrip(),
"",
"## Watch Terms",
"",
_render_table(watch_items).rstrip(),
"",
"## Alias Suggestions",
"",
"_Conservative by design in this minimal version; no automatic alias suggestions are emitted yet._" if not alias_items else _render_simple_table(alias_items, "from").rstrip(),
"",
"## Stopword Suggestions",
"",
"_Conservative by design in this minimal version; no automatic stopword suggestions are emitted yet._" if not stopword_items else _render_simple_table(stopword_items, "term").rstrip(),
"",
"## Apply",
"",
"Review the Markdown first, then selectively apply accepted suggestions with the JSON file.",
"",
"```bash",
f"python scripts/apply_term_suggestions.py \\",
f" --suggestions {json_output_path} \\",
" --accept-interest \"Claude Code\" \\",
" --accept-watch \"A2A\" \\",
" --dry-run",
"```",
"",
])
return "\n".join(lines)
def _build_output_paths(
*,
output_dir: Path,
suggestion_date: str,
json_output: Path | None,
markdown_output: Path | None,
) -> tuple[Path, Path]:
stem = f"term-cleanup-suggestions-{suggestion_date}"
resolved_json = json_output or (output_dir / f"{stem}.json")
resolved_markdown = markdown_output or (output_dir / f"{stem}.md")
return resolved_json, resolved_markdown
def main() -> None:
parser = argparse.ArgumentParser(description="Generate term cleanup suggestions JSON and Markdown from review bundle.")
parser.add_argument("--bundle", type=Path, default=DEFAULT_BUNDLE_PATH, help="Review bundle JSON file")
parser.add_argument(
"--output-dir",
type=Path,
default=DEFAULT_OUTPUT_DIR,
help="Directory for generated suggestions outputs when explicit output paths are not provided",
)
parser.add_argument("--date", type=str, default=None, help="Override suggestions date (YYYY-MM-DD)")
parser.add_argument("--json-output", type=Path, default=None, help="Explicit suggestions JSON output path")
parser.add_argument("--markdown-output", type=Path, default=None, help="Explicit suggestions Markdown output path")
parser.add_argument(
"--emit-markdown",
action="store_true",
help="Also write the human-readable Markdown review draft. JSON suggestions are always written.",
)
args = parser.parse_args()
if not args.bundle.exists():
raise RuntimeError(f"Bundle file not found: {args.bundle}")
bundle = _require_dict(_load_json(args.bundle), "bundle")
days = bundle.get("days")
if not isinstance(days, int):
raise RuntimeError("bundle.days must be an integer.")
suggestion_date = args.date or _bundle_date(bundle)
json_output_path, markdown_output_path = _build_output_paths(
output_dir=args.output_dir,
suggestion_date=suggestion_date,
json_output=args.json_output,
markdown_output=args.markdown_output,
)
# Load full term_stats for alias scanning (bundle only has top N)
stats_path = REPO_ROOT / "data" / "term_index" / "term_stats.json"
all_stats_terms: list[str] = []
if stats_path.exists():
stats_payload = _load_json(stats_path)
raw_terms = stats_payload.get("terms") if isinstance(stats_payload, dict) else []
if isinstance(raw_terms, list):
all_stats_terms = [str(t["term"]) for t in raw_terms if isinstance(t, dict) and isinstance(t.get("term"), str)]
interest_items = _prepare_interest_suggestions(bundle)
reserved_terms = {_term_key(str(item.get("term") or "")) for item in interest_items}
watch_items = _prepare_watch_suggestions(bundle, reserved_terms=reserved_terms)
alias_items = _prepare_alias_suggestions(bundle, all_terms=all_stats_terms)
suggestions = {
"date": suggestion_date,
"based_on_days": days,
"source_bundle": str(args.bundle),
"policy_schema_version": _require_dict(bundle.get("policy"), "bundle.policy").get("schema_version", "unknown"),
"summary": {
"interest_keyword_suggestions": len(interest_items),
"watch_terms": len(watch_items),
"alias_suggestions": len(alias_items),
"stopword_suggestions": 0,
},
"alias_suggestions": alias_items,
"stopword_suggestions": [],
"interest_keyword_suggestions": interest_items,
"watch_terms": watch_items,
}
markdown = _render_markdown(
suggestion_date=suggestion_date,
bundle_path=args.bundle,
json_output_path=json_output_path,
bundle=bundle,
suggestions=suggestions,
)
_save_json(json_output_path, suggestions)
if args.emit_markdown:
_save_text(markdown_output_path, markdown)
summary = {
"bundle": str(args.bundle),
"date": suggestion_date,
"json_output": str(json_output_path),
"markdown_output": str(markdown_output_path) if args.emit_markdown else None,
"interest_keyword_suggestions": len(interest_items),
"watch_terms": len(watch_items),
"alias_suggestions": len(alias_items),
"stopword_suggestions": 0,
"emit_markdown": args.emit_markdown,
}
print(json.dumps(summary, ensure_ascii=False, indent=2))
if __name__ == "__main__":
main()
+1 -1
View File
@@ -37,7 +37,7 @@ def main() -> None:
help="Directory to write per-article Markdown summaries",
)
parser.add_argument("--max-retries", type=int, default=2, help="Maximum LLM retry attempts per article")
parser.add_argument("--timeout", type=float, default=60.0, help="LLM request timeout in seconds")
parser.add_argument("--timeout", type=float, default=120.0, help="LLM request timeout in seconds")
parser.add_argument("--api-key", type=str, default=None, help="Override article-summary LLM API key")
parser.add_argument("--model", type=str, default=None, help="Override article-summary LLM model")
parser.add_argument("--api-url", type=str, default=None, help="Override article-summary LLM API URL/base URL")
+24
View File
@@ -0,0 +1,24 @@
from __future__ import annotations
import argparse
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
SRC_ROOT = REPO_ROOT / "src"
if str(SRC_ROOT) not in sys.path:
sys.path.insert(0, str(SRC_ROOT))
from summary_mcp.runtime.article_summary_jobs import run_article_summary_job
def main() -> None:
parser = argparse.ArgumentParser(description="Run a background article-summary job by job_id.")
parser.add_argument("--job-id", required=True, help="Article summary job id")
args = parser.parse_args()
run_article_summary_job(job_id=args.job_id)
if __name__ == "__main__":
main()
+24
View File
@@ -0,0 +1,24 @@
from __future__ import annotations
import argparse
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
SRC_ROOT = REPO_ROOT / "src"
if str(SRC_ROOT) not in sys.path:
sys.path.insert(0, str(SRC_ROOT))
from summary_mcp.runtime.freshrss_pipeline_jobs import run_freshrss_pipeline_job
def main() -> None:
parser = argparse.ArgumentParser(description="Run a background FreshRSS pipeline job by job_id.")
parser.add_argument("--job-id", required=True, help="FreshRSS pipeline job id")
args = parser.parse_args()
run_freshrss_pipeline_job(job_id=args.job_id)
if __name__ == "__main__":
main()
+24
View File
@@ -0,0 +1,24 @@
from __future__ import annotations
import argparse
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
SRC_ROOT = REPO_ROOT / "src"
if str(SRC_ROOT) not in sys.path:
sys.path.insert(0, str(SRC_ROOT))
from summary_mcp.runtime.resume_jobs import run_resume_job
def main() -> None:
parser = argparse.ArgumentParser(description="Run a background resume job by job_id.")
parser.add_argument("--job-id", required=True, help="Resume job id")
args = parser.parse_args()
run_resume_job(job_id=args.job_id)
if __name__ == "__main__":
main()
+134 -13
View File
@@ -1,15 +1,32 @@
---
name: keyword-cleanup-review
description: 审查和整理本仓库的每日关键词索引和频率统计。当用户需要检查 `data/term_index/term_stats.json`、最近的 `data/term_index/daily/*.json`、`configs/term_aliases.json`、`configs/term_stopwords.json` 或 `configs/filter_context.personal.json`,以提议别名合并、停用词、关注词或 `interest_keywords` 更新(但不直接修改配置)时使用。
description: 生成 reader 项目的正式关键词 review 输入。当用户需要基于 `data/term_index/term_stats.json`、最近的 `data/term_index/daily/*.json` 和当前 configs 产出 bundle、正式 suggestions JSON,或按需生成审阅 Markdown 供 OpenClaw 汇报和等待确认时使用。不要用于低频清理 review 目录、删除旧产物或直接 apply 配置。
---
# 关键词清理审查
使用此技能将仓库的关键词统计转化为可审查的清理建议。
使用此技能将仓库的关键词统计转化为**正式 review 输入**,供 OpenClaw 后续做汇报、确认和 apply 编排。
## 角色边界
这个 skill 负责:
- 构建 review bundle
- 生成正式 suggestions JSON
- 按需生成人工审阅 Markdown
- 给 OpenClaw 提供稳定的关键词 review 输入
这个 skill 不负责:
- 清理 `outputs/term_index/review/` 下的旧文件
- 决定删除哪些历史 bundle / suggestions / markdown
- 直接 apply `configs/term_aliases.json` / `configs/term_stopwords.json` / `configs/filter_context.personal.json`
低频维护、清理和 dry-run 校验应由 OpenClaw 侧 maintenance SOP 处理,而不是由本 skill 承担。
## 工作流程
1. 构建精简的审查数据包:
### Phase 1:构建审查数据包
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py
@@ -17,25 +34,89 @@ python skills/keyword-cleanup-review/scripts/build_review_bundle.py
可选参数:
- `--days 7`
- `--top 50`
- `--days 7`(默认 7,建议传 365 覆盖全量)
- `--top 100`(考虑的词数)
- `--output outputs/term_index/review/keyword-cleanup-bundle.json`
2. 阅读生成的数据包和建议模式:
#### 候选引擎策略
- `outputs/term_index/review/keyword-cleanup-bundle.json`
- `skills/keyword-cleanup-review/references/suggestion-schema.md`
根据 `configs/term_cleanup_policy.json` 的 `schema_version` 自动切换:
3. 生成两份输出:
| 版本 | 策略 | 说明 |
|------|------|------|
| v1(旧) | 固定阈值(total≥3/days≥2 → interest) | 小数据集兼容 |
| v2(当前默认) | 百分位排名 + 增速因子 | 自适应数据量,不需要手工调阈值 |
- 一份简短的供人工审阅的 Markdown 报告
- 一份符合模式的 JSON 建议文件
v2 策略说明:
- **percentile**:total_count 在所有词里的排位占比。top 5% → interest 候选,5%-20% → watch 候选
- **growth**:recent_count / total_count,衡量近期活跃度。growth≥0.5 的排位外词也会主动推荐
4. 严格保持边界:
### Phase 2:生成建议(规则层)
```bash
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
```
如需人工审阅展示稿:
```bash
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
#### 产出能力
| 建议类型 | 状态 | 方法 |
|---------|------|------|
| interest 建议 | ✅ 已实现 | 百分位 top 5% + 增速促活 |
| watch 建议 | ✅ 已实现 | 百分位 5%-20% |
| alias 建议 | ✅ 已实现 | 规则层:大小写归一、单复数、去空格/连字符 |
| stopword 建议 | ❌ 规则层空缺 | 见 Phase 3(LLM 层) |
默认生成:
- `term-cleanup-suggestions-YYYY-MM-DD.json`(正式建议产物)
显式加 `--emit-markdown` 额外生成:
- `term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿)
### Phase 3:生成建议(LLM 层,可选)
规则层覆盖不了 alias(中英文对应、缩写展开、同义不同名)和 stopword 判断,需要 LLM 辅助:
```bash
python scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
从 `.env` 读取 LLM 配置(`LLM_API_URL` / `LLM_MODEL` / `LLM_API_KEY`),使用 DeepSeek API。
输出三部分:
| 输出 | 说明 |
|------|------|
| `semantic_alias` | 语义级别名(中英文、缩写、同义不同名) |
| `stopword` | 泛词过滤建议(规则层做不了的需要语义判断的) |
| `promote_to_interest` | 与用户关注方向一致的新词,建议加入 interest |
**注:LLM 层产物是候选,不应自动 apply,需要人工确认后由 OpenClaw 编排 apply。**
### Phase 4:输出给 OpenClaw 编排
- `suggestions JSON` = review / apply 之间唯一正式建议输入
- `semantic-suggestions JSON` = LLM 补充建议,需要人工筛选后合并到 suggestions JSON 再 apply
- Markdown = 临时展示层
- 后续汇报、确认、dry-run、apply、收尾清理由 OpenClaw 编排层执行
### Phase 5:严格保持边界
- 建议 `configs/term_aliases.json` 的修改
- 建议 `configs/term_stopwords.json` 的修改
- 建议 `configs/filter_context.personal.json` 的新增
- **LLM 层产出(semantic-suggestions)不自动 apply**,需人工确认后由 OpenClaw 编排层执行
- 除非用户明确要求,否则不要直接编辑这些文件
- 除非用户要求修改规则逻辑,否则不要建议直接编辑 `configs/filter_rules.json`
@@ -78,10 +159,47 @@ Markdown 输出应:
- 分类别名、停用词、兴趣关键词和关注词建议
- 用简短、具体的句子解释理由
说明:Markdown 主要用于人工临时审阅,不必默认当作长期资产保留。
JSON 输出应遵循:
- `references/suggestion-schema.md`
说明:JSON 是 review / apply 之间的唯一正式建议产物,应优先保留。
## 产物口径
长期保留:
- `data/term_index/daily/*.json`
- `data/term_index/term_stats.json`
- `configs/filter_context.personal.json`
- `configs/term_watchlist.json`
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/term_change_log.json`
短期保留:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
- `outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json`
临时产物:
- `outputs/term_index/review/keyword-cleanup-bundle.json`
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
默认执行口径:
- bundle 只作为运行时工作文件,默认只保留当前最新一份
- Markdown 只作为人工展示层,优先按需生成,不默认长期归档
- JSON suggestions 是 review / apply 之间唯一正式建议输入
说明:
- “是否删除旧 bundle / 旧 markdown / 旧 suggestions” 不属于本 skill 的正式职责
- 这类维护动作应由 OpenClaw 侧的 maintenance skill 处理
## 仓库说明
当前仓库行为:
@@ -97,6 +215,9 @@ JSON 输出应遵循:
## 资源
- 脚本:
- `scripts/build_review_bundle.py`
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- `scripts/generate_term_cleanup_suggestions.py`
- `scripts/generate_term_cleanup_semantic_suggestions.py`(LLM 层)
- 参考文档:
- `references/suggestion-schema.md`
- `plans/keyword-cleanup-interest-watch-engine-improvement.md`(v2 引擎设计)
@@ -9,16 +9,15 @@ from typing import Any
DEFAULT_POLICY: dict[str, Any] = {
"schema_version": "v1",
"schema_version": "v2",
"interest_keyword_review": {
"min_total_count": 3,
"min_days_seen": 2,
"percentile_min": 0.0,
"percentile_max": 0.05,
"growth_promotion": 0.5,
},
"watch_term_review": {
"min_total_count": 1,
"min_days_seen": 1,
"max_total_count": 2,
"max_days_seen": 2,
"percentile_min": 0.05,
"percentile_max": 0.20,
},
"alias_review": {
"min_total_count": 2,
@@ -28,6 +27,11 @@ DEFAULT_POLICY: dict[str, Any] = {
"max_total_count": 2,
"max_days_seen": 2,
},
"notes": [
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
"growth = recent_count / total_count,衡量近期活跃度",
],
}
@@ -126,6 +130,37 @@ def _within_watch_thresholds(item: dict[str, Any], thresholds: dict[str, Any]) -
)
def _compute_percentile(value: int, sorted_values: list[int]) -> float:
"""
Return the percentile rank of `value` in `sorted_values` (ascending).
0.0 = highest frequency (top rank), 1.0 = lowest frequency (bottom rank).
"""
if not sorted_values:
return 1.0
# bisect_left — count of values strictly less than `value`
lo, hi = 0, len(sorted_values)
while lo < hi:
mid = (lo + hi) // 2
if sorted_values[mid] < value:
lo = mid + 1
else:
hi = mid
rank = lo
# invert: smallest value → rank=0 → 1.0 (bottom)
# largest value → rank=len → 0.0 (top)
return 1.0 - (rank / len(sorted_values))
def _compute_growth(recent_count: int, total_count: int) -> float:
"""
Return growth factor: recent_count / total_count.
Only meaningful when total_count >= 3; returns 0.0 for small counts.
"""
if total_count < 3:
return 0.0
return recent_count / total_count
def main() -> None:
parser = argparse.ArgumentParser(
description="Build a compact review bundle for the keyword-cleanup-review skill."
@@ -235,6 +270,13 @@ def main() -> None:
alias_values = _casefold_set(list(aliases.values()))
watch_set = _casefold_set([str(item.get("term", "")) for item in watchlist])
# Build a sorted list of all total_counts for percentile computation
all_total_counts = sorted(
int(item.get("total_count") or 0)
for item in stats_terms
if isinstance(item, dict) and isinstance(item.get("term"), str)
)
top_global_terms = []
for item in stats_terms[: args.top]:
if not isinstance(item, dict):
@@ -256,42 +298,108 @@ def main() -> None:
"is_alias_target": folded in alias_values,
"in_watchlist": folded in watch_set,
"recent_count": recent_counter.get(term, 0),
"percentile": _compute_percentile(
int(item.get("total_count") or 0), all_total_counts
),
"growth": _compute_growth(
recent_counter.get(term, 0),
int(item.get("total_count") or 0),
),
}
)
# Keep more uncovered terms for percentile-based selection
uncovered_terms = [
item for item in top_global_terms if not item["in_interest_keywords"] and not item["is_stopword"]
][:20]
][:100]
policy_version = (policy.get("schema_version") if isinstance(policy, dict) else None) or "v1"
interest_thresholds = policy.get("interest_keyword_review") if isinstance(policy, dict) else {}
watch_thresholds = policy.get("watch_term_review") if isinstance(policy, dict) else {}
if policy_version == "v2" or "percentile_max" in interest_thresholds:
# v2: percentile + growth based selection
pct_min_interest = float(interest_thresholds.get("percentile_min", 0.0))
pct_max_interest = float(interest_thresholds.get("percentile_max", 0.05))
growth_promo = float(interest_thresholds.get("growth_promotion", 0.5))
pct_min_watch = float(watch_thresholds.get("percentile_min", 0.05))
pct_max_watch = float(watch_thresholds.get("percentile_max", 0.20))
interest_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"]
and pct_min_interest <= item["percentile"] <= pct_max_interest
]
watch_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"]
and pct_min_watch < item["percentile"] <= pct_max_watch
]
# Growth boost: terms outside watch range but with strong growth signal
growth_boost_candidates = [
item for item in uncovered_terms
if not item["in_watchlist"]
and item["percentile"] > pct_max_watch
and item["growth"] >= growth_promo
]
else:
# v1 fallback: fixed thresholds
interest_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
]
watch_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"]
and not _meets_min_thresholds(item, interest_thresholds)
and _within_watch_thresholds(item, watch_thresholds)
]
growth_boost_candidates = []
interest_review_candidates = [
{
"term": item["term"],
"total_count": item["total_count"],
"days_seen": item["days_seen"],
"percentile": item["percentile"],
"growth": item["growth"],
"reason": (
"Meets the configured interest-keyword review threshold and is not yet covered "
"by interest keywords or stopwords."
f"top {item['percentile']:.1%} by frequency,"
f"growth={item['growth']:.0%},"
"not yet covered by interest keywords or stopwords."
),
}
for item in uncovered_terms
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
for item in interest_candidates_raw
][:20]
watch_review_candidates = [
{
"term": item["term"],
"total_count": item["total_count"],
"days_seen": item["days_seen"],
"percentile": item["percentile"],
"growth": item["growth"],
"reason": (
"Falls into the configured watch-term review range and should be observed "
"before promotion into interest keywords."
f"top {item['percentile']:.1%} by frequency,"
f"growth={item['growth']:.0%},"
"fell into watch-review range."
),
}
for item in uncovered_terms
if not item["in_watchlist"]
and not _meets_min_thresholds(item, interest_thresholds)
and _within_watch_thresholds(item, watch_thresholds)
for item in watch_candidates_raw
][:20]
growth_boost_review_items = [
{
"term": item["term"],
"total_count": item["total_count"],
"days_seen": item["days_seen"],
"percentile": item["percentile"],
"growth": item["growth"],
"reason": (
f"growth spike: {item['growth']:.0%} of occurrences in recent window "
f"(total={item['total_count']}, days={item['days_seen']})."
),
}
for item in growth_boost_candidates
][:5]
recent_hot_terms = sorted(
({"term": term, "recent_count": count} for term, count in recent_counter.items()),
key=lambda item: (-item["recent_count"], item["term"].casefold(), item["term"]),
@@ -330,6 +438,7 @@ def main() -> None:
"governance_hints": {
"interest_review_candidates": interest_review_candidates,
"watch_review_candidates": watch_review_candidates,
"growth_boost_review_items": growth_boost_review_items,
},
}
_save_json(args.output, bundle)
+186
View File
@@ -0,0 +1,186 @@
---
name: reader-digest-flow
description: 编排 reader 项目的端到端 AI 日报流程。仅在用户要求运行/重跑日报、汇报候选、发布 Hugo 日报、沉淀选中文章或继续已有日报任务时使用;覆盖异步 MCP 任务、候选确认、发布、单篇摘要和 IMA 知识库上传。不要因验证、排障冲动或候选质量不佳自行重跑。
---
# Reader Digest Flow
## 职责边界
本 Skill 负责:
- 通过 reader MCP 启动、观察和恢复日报任务;
- 向用户展示候选并保持稳定编号;
- 根据用户选择生成并发布 Hugo 日报;
- 对用户选中的文章生成知识笔记并编排 IMA 上传;
- 在每个副作用边界执行确认和结果验证。
本 Skill 不负责:
- 实现 reader 内部抓取、摘要、过滤或恢复逻辑;
- 通过手拼目录推导 Run 状态或 Artifact;
- 未经用户要求自行重跑 Pipeline;
- 未经用户确认发布日报或写入知识库;
- 直接维护关键词配置;关键词治理委托给 `keyword-cleanup-review`。
## 核心规则
1. **只按用户指令运行。** 只有用户明确要求“跑日报”“重新跑”“再跑一次”时才启动新 Pipeline。验证、解释排序和排障默认读取已有 Run。
2. **一次对话绑定一个当前 Run。** 以异步 Job 结果返回的 `run_id` 为稳定句柄;新 Run 产生新的候选编号体系,不混用历史编号。
3. **状态以 MCP 返回为准。** Agent 只根据顶层 `status` 和 `recommended_action` 分支;`status_source`、`state_conflict` 仅用于解释。
4. **路径以返回值为准。** 使用 `output_dir`、`artifact.path`、`delivery_output`、`report_output` 和 `written_paths`;不要根据 `run_id` 手拼 `outputs/...`。
5. **候选编号保持稳定。** 用户编号永远对应当前候选列表的原始顺序(1-based);跨产物读取详情时按 URL 或完整 `item_id` 关联,不按数组位置关联。
6. **副作用必须授权。** 用户确认 Hugo 文章后才能发布;用户确认 IMA 文章后才能生成并上传知识笔记。
7. **内容必须有来源。** 日报和知识笔记只能基于当前 Run 的 `article.plain_text`、摘要、highlights 等 Artifact;不得使用通用知识补写原文没有的信息,也不为满足长度而扩写。
## 默认生产参数
用户未显式覆盖时使用:
```json
{
"limit": 7,
"include_read": false,
"mark_read": true,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
- 默认只处理未读文章。
- `include_read=true` 仅在用户明确要求扩大到已读内容时使用。
- debug/test/validation 才允许 `mark_read=false` 或 `debug_artifacts=true`。
- 不随机生成文章数量;用户指定 `limit` 时按用户值执行。
## 正式流程
### Phase 1:启动并观察日报 Job
正式生产入口统一为异步 MCP:
1. 调用 `start_freshrss_pipeline_job`;
2. 轮询 `get_freshrss_pipeline_job_status`;
3. `status=success` 后调用 `get_freshrss_pipeline_job_result`;
4. 保存返回的 `run_id` 和 Artifact 路径;
5. 使用 `get_run_status`、`get_delivery_payload`、`get_run_report` 读取业务状态和结果。
失败时:
1. 使用 `get_run_status(run_id)` 读取关联 Run;
2. 调用 `inspect_resume_plan(run_id)`;
3. `recommended_action=resume` 时启动并轮询异步 Resume Job;
4. `recommended_action=read_terminal_result` 时直接读取已有终态结果;
5. `recommended_action=start_new_run` 时停止并向用户报告,不自行新建 Run。
CLI 仅用于 MCP 不可用时的 fallback、debug 或人工排障,不是默认生产入口。具体调用序列见 `references/flow.md`。
### Phase 2:汇报候选
- 使用当前 Run 返回的 Delivery Payload 或 digest brief Artifact;
- 按候选原始顺序从 1 编号,状态可显示为“已入选/待确认”,但不得重新分组编号;
- 每篇提供标题、来源、2-3 句摘要和筛选理由,避免原始 JSON dump;
- 用户质疑编号或排序时读取当前 Run 产物核对,不重新运行 Pipeline;
- 需要跨 Artifact 取详情时按 URL 或完整 `item_id` 交叉验证。
Feishu 输出不要使用 Markdown 表格,见 `references/feishu-format-notes.md`。
### Phase 3:等待 Hugo 选择
- 等待用户明确选择要发布的文章;
- 用户编号映射到当前候选列表,不映射到 extracted 文件序号;
- 用户拒绝发布时立即停止当日日报后续流程,不劝说、不自动换一批;
- 用户明确要求重跑时才创建新 Run,并重新建立编号体系。
### Phase 4:生成并发布 Hugo 日报
发布前读取 `references/public-digest-example.md`,按其最终页面结构生成:
- `今日概览`
- `今日重点`
- `趋势观察`
每篇 `今日重点` 文章末尾必须添加 `来源:[来源名](原文 URL)`,来源链接跟随对应文章,不再生成独立的 `延伸阅读` 章节或重复链接。
仅发布用户在 Phase 3 选中的文章。公开页面不得出现 `keep/review/drop`、候选、待确认等内部状态。
写入 Hugo 后执行部署,并验证首页、日报列表页和当日详情页均可访问。命令和检查项见 `references/flow.md`。
### Phase 5:等待 IMA 选择
Hugo 发布选择与知识沉淀选择相互独立。询问用户哪些文章值得长期保存:
- 仅处理用户明确选择的文章;
- 不把整份日报上传到 IMA;
- 本次选择本身即授权后续单篇摘要和 IMA 上传,不重复确认。
### Phase 6:生成单篇知识笔记
对每篇选中文章:
1. 通过 URL/完整 `item_id` 找到对应 extracted Artifact;
2. 使用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
3. 轮询 `get_article_summary_job_status`,成功后读取 `get_article_summary_job_result`;
4. 只使用现有 `article.plain_text`,不重新抓取原 URL;
5. 使用返回的 `written_paths` 定位结果并检查 Markdown 内容。
6. **`extracted_path` 必须传绝对路径**(前缀 `/home/ubuntu/zhu/github/reader/`):summary-mcp 工作目录是 `/root/.hermes`,相对路径会报 `extracted_path does not exist`。同一工具连续 3 次失败会触发 MCP 冷却(约 45-60s,报 `MCP server 'reader' is unreachable`),等待冷却后再重试,不要循环重试同一调用。
异步 MCP 不可用时才使用项目 CLI fallback。不要因为内容较短而引入原文之外的知识。
### Phase 7:上传到 IMA
上传前按需读取:
- 格式规则:`references/ima-format-quickref.md`
- API 和上传步骤:`references/ima-upload-api.md`(含 `-200` 版本拦截修复)
- 凭证定位:`references/ima-credential-chain.md`
- 批量上传脚本:`scripts/ima_upload_one.py`(Python 编排,规避中文文件名 bash 引号问题)
硬规则:
- 使用完整文章标题作为文件名;
- 以 Markdown 文件 `media_type=7` 上传到 `daily` knowledge base;
- 保留原文 URL 和 Category,保证来源可追踪;
- 不使用 URL 导入或 Notes 类型替代知识库文件;
- 上传后验证目标知识库中存在对应条目;
- 失败时报告具体阶段,不无限重试。
## 关键词治理路由
只有用户明确要求“清理关键词”“词库治理”等操作时才触发关键词治理。Review bundle 和 Suggestions 生成委托给 `keyword-cleanup-review`,确认与 Apply 仍由当前编排层负责:
1. 生成 review bundle;
2. 生成规则层与可选语义层 Suggestions JSON;
3. 等待人工确认;
4. dry-run 后按精确 accept 参数 Apply。
`reader-digest-flow` 不直接编辑 `term_aliases.json`、`term_stopwords.json` 或兴趣配置。简要路由见 `references/keyword-engine-maintenance.md`。
## 环境坑位(本机部署)
- **MCP 相对路径陷阱**:`summary-mcp` 服务工作目录是 `/root/.hermes`,不是 reader 项目根。MCP 返回的 `output_dir`/`artifact.path` 是相对路径,直接传给 `start_article_summary_job(extracted_path=...)` 会报 `extracted_path does not exist`。传入前必须拼绝对路径前缀 `/home/ubuntu/zhu/github/reader/`。
- **提取失败不等于运行失败**:`status_counts.extract_failed` 的条目(`CONTENT_EXTRACTION_FAILED`,`retryable=false`)跳过即可并如实汇报;失败文章常是推广/活动等低价值内容,不因此自行重跑。`linked_run_status=partial` 时先读 run-report 的 item 级 `error` 确认原因。
- **用户要求"重新跑一批"**:候选质量低(用户主动提出)时重跑,应 `include_read=true` 并调高 `limit`(如 10),否则默认 `include_read=false` 会拉回同一批未读文章。重跑是新 Run,候选编号体系重新建立,汇报时提醒用户按新列表选择。
## 停止与人工介入
出现以下任一情况时停止自动流程并报告:
- 用户没有授权运行、发布或知识库写入;
- Job/Run 返回不可恢复,或连续恢复失败;
- Payload、候选 ID 或 Artifact 之间无法可靠关联;
- 生成内容缺少可追踪来源;
- Hugo 部署验证失败;
- IMA 凭证、目标知识库或上传结果无法验证。
## Reference 路由
- `references/flow.md`:具体 MCP 调用序列、候选映射(含 extracted_path 绝对路径、候选≠文件名顺序)、Hugo 发布和 IMA 主步骤。
- `references/content-extraction.md`:FreshRSS 内容来源与 `plain_text` 质量判断。
- `references/public-digest-example.md`:可直接参考的 Hugo 最终页面结构。
- `references/feishu-format-notes.md`:Feishu 输出格式限制。
- `references/ima-format-quickref.md`:IMA Markdown 格式规则。
- `references/ima-upload-api.md`:IMA Markdown 文件上传 API(含 `-200` 版本拦截修复)。
- `references/ima-credential-chain.md`:IMA 凭证与知识库配置定位。
- `references/keyword-engine-maintenance.md`:关键词治理 Skill 路由。
- `scripts/ima_upload_one.py`:单篇 Markdown 上传 daily 知识库的完整 Python 脚本(preflight→重名→create_media→COS→add_knowledge)。
@@ -0,0 +1,45 @@
# 内容提取流程
本文说明处理流水线如何把 FreshRSS 条目转换为可供摘要使用的文章文本。
## 核心规则:FreshRSS 条目不重新抓取原文 URL
**FreshRSS 是仅提供 RSS 内容的上游。** 对于 FreshRSS 条目,流水线不会向文章原始 URL 发起 HTTP 请求。该行为由 `pipeline.py` 中的 `RSS_ONLY_UPSTREAMS = {"freshrss"}` 强制保证。
唯一例外是非 FreshRSS 上游。未来未设置 `upstream: freshrss` 的其他来源,可以在必要时使用 `fetch_html()` 作为回退。
## 内容来源优先级
`content_loader.py` 按以下顺序检查内容,并使用第一个包含 **至少 500 个可读字符** 的来源:
| 优先级 | 来源 | 含义 |
|--------|------|------|
| 1 | `raw_html` | 通过 `ExtractionInput.raw_html` 预先注入的 HTML;常规 FreshRSS 运行中很少使用。 |
| 2 | `item.raw_content` | RSS `<content:encoded>` 中的文章正文;部分订阅源提供,部分不提供。 |
| 3 | `item.raw_summary` | RSS `<description>` 中的摘要或片段;这是当前运行中最常见的来源。 |
| 4 | `rss_content` | 来自非条目字段的独立 RSS 内容。 |
| — | `none` | 没有可用内容;FreshRSS 不允许回源抓取,因此抛出 `RSS_CONTENT_MISSING`。 |
## `content_source` 与文本质量的关系
每个 `item-XX.extracted.json` 中的 `content_source` 字段表示流水线实际使用的内容来源:
- **`item.raw_content`**:RSS `<content:encoded>` 提供的文章正文,通常质量最好,接近直接阅读原文。
- **`item.raw_summary`**:只有 RSS 摘要或描述,并非完整正文。不同来源长度差异较大,通常为 300-2000 个字符;AI 摘要基于该片段,而不是完整文章。
- **`rss_content`**:来自独立 RSS 内容,质量取决于订阅源。
- **`fetched_html`**:从原始 URL 抓取的 HTML。FreshRSS 条目不会出现该来源,只适用于非 FreshRSS 上游。
## 对日报质量的影响
如果提取结果文件中出现 `content_source: item.raw_summary`,说明 AI 使用的是订阅源摘要或片段,而不是完整正文。日报内容显得较浅时,原因可能只是 RSS 描述过短。
提高质量可以选择提供完整 `<content:encoded>` 的订阅源,或者把内容来源切换到支持全文 RSS 的系统,例如具备全文提取能力的 RSS 代理或 FiveFilters 等服务。
## 快速检查
先调用 `list_run_artifacts(run_id)`,再读取返回的提取结果产物路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。
## 相关代码路径
- `src/summary_mcp/core/pipeline.py`:`RSS_ONLY_UPSTREAMS`、`_should_skip_fetch()`、`extract_content()`。
- `src/summary_mcp/core/content_loader.py`:`choose_inline_content()` 的优先级链和 `fetch_html()`;FreshRSS 条目不会调用后者。
@@ -0,0 +1,17 @@
# 飞书 Markdown 格式说明
## 背景
Hermes 的飞书网关(`gateway/platforms/feishu.py`)通过 `_build_outbound_payload` 发送消息。该方法会检查内容中的 Markdown 特征,并据此决定消息类型:
- 内容匹配 `_MARKDOWN_HINT_RE`(加粗、列表、代码、链接等)时,使用包含 `md` 元素的飞书 `post` 类型发送,可以正常渲染。
- 内容匹配 `_MARKDOWN_TABLE_RE`(Markdown 表头和分隔行)时,整条消息会被强制转换为 `text` 类型,即纯文本,不再渲染 Markdown。
原因是 `_build_markdown_post_payload` 会把内容包装为 `{"tag": "md", "text": "..."}` 元素,而飞书的 `md` 元素不支持表格,也没有把 Markdown 表格转换为飞书原生表格的逻辑。
## 飞书输出规则
- 通过飞书发送的消息不得使用 Markdown 表格;消息中只要出现一个表格,整条消息就会退化为纯文本。
- 需要表达结构化信息时,优先使用分点列表、带标题的分节或行内格式。
- 加粗(`**加粗**`)、行内代码(`` `代码` ``)、无序列表(`- 项目`)、有序列表(`1. 项目`)和链接均可正常使用。
- 围栏式代码块可以使用,但代码块后的尾随内容可能存在渲染边界问题。
@@ -0,0 +1,177 @@
# Reader Digest Flow 操作参考
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
## 1. 日报 Pipeline
### 默认参数
```json
{
"limit": 7,
"include_read": false,
"mark_read": true,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
### 正式调用序列
```text
start_freshrss_pipeline_job
→ get_freshrss_pipeline_job_status
→ get_freshrss_pipeline_job_result
→ get_run_status
→ get_delivery_payload / get_run_report
```
状态动作:
- `running`:按合理间隔继续轮询;
- `success`:读取结果,保存 `run_id`;
- `failed`:读取关联 Run 并执行 Resume Plan;
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
恢复序列:
```text
inspect_resume_plan
→ recommended_action=resume
→ start_resume_job
→ get_resume_job_status
→ get_resume_job_result
```
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
## 2. 候选汇报与选择
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
展示规则:
1. 使用候选数组原始顺序并从 1 编号;
2. 不因 `keep/review` 分组而重新编号;
3. 每篇展示标题、来源、摘要和判断理由;
4. 用户编号只映射当前候选数组;
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
### extracted 文件与候选编号错位(实测 2026-07-31)
- `digest-brief.json` 的 `top_candidates` **没有 `item_id` 字段**,只有 `url`;
- `extracted/item-XX.extracted.json` 的文件名序号与候选编号**可能不一致**(实例:候选2 = item-05、候选3 = item-02);
- 正确做法:用 **URL 交叉匹配**(归一化 `%3D`→`=` 后逐条比对),或用完整 `item_id`(从 candidate-batch.json 的 `items[i].item_id` 按候选数组顺序取)在 extracted 文件里反查;两者都能验证时优先 item_id。
## 3. Hugo 日报
用户确认发布文章后:
1. 读取 `public-digest-example.md`;
2. 仅使用用户选中的文章生成公开内容;
3. 写入 Hugo 当日页面;
4. 前台执行部署,避免把构建日志作为聊天通知;
5. 验证首页、列表页和详情页。
当前部署位置:
```text
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
```
验证至少覆盖:
```text
http://127.0.0.1:14322/
http://127.0.0.1:14322/daily/
http://127.0.0.1:14322/daily/YYYY-MM-DD/
```
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
## 4. 单篇知识笔记
用户确认 IMA 文章后:
1. 从候选中取得 URL 和完整 `item_id`;
2. 从 Run Artifact 中找到匹配的 extracted 文件;
3. 交叉验证 `article.item_id` 或 URL;
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
6. 从 Job 结果的 `written_paths` 读取 Markdown。
### ⚠️ extracted_path 必须用绝对路径
`summary-mcp` 进程的工作目录是 `/root/.hermes`(不是 reader 项目根)。传相对路径(如 `outputs/freshrss/...`)会直接报 `extracted_path does not exist`。必须传绝对路径:
```text
/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json
```
### ⚠️ 候选编号 ≠ extracted 文件名顺序
候选数组顺序与 extracted 文件名(`item-01`…`item-07`)**不一定对齐**(实测候选2 落在 item-05)。`digest-brief.json` 的候选**没有 `item_id` 字段**,只有 URL。可靠匹配方法:
1. 从 `candidate-batch.json` 取每项完整 `item_id`(在 `candidate` 嵌套对象里,顶层 `item_key` 只是 `item-XX` 文件名序号);
2. 或按 URL 匹配:归一化(`%3D`→`=`)后与每个 extracted 文件的 `article.url` / `article.canonical_url` 比对;
3. 绝不要按候选位置对应 extracted 文件序号。
```python
def norm(u): return u.replace('%3D','=').replace('%3d','=').strip()
# 对每个 extracted 文件取 norm(article.url),与候选 norm(url) 精确比对
```
正式序列:
```text
start_article_summary_job
→ get_article_summary_job_status
→ get_article_summary_job_result
```
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
CLI 仅在异步 MCP 不可用或人工排障时使用:
```bash
python scripts/run_article_summaries.py \
--extracted <returned-extracted-path> \
--ids <full-item-id> \
--output-dir <explicit-output-dir>
```
## 5. IMA 上传
用户在知识沉淀阶段的文章选择即为上传授权。
执行顺序:
1. 检查生成的 Markdown 与来源;
2. 文件名规范化为 `<完整文章标题>.md`;
3. 确认目标为 `daily` knowledge base;
4. 执行 preflight、create_media、COS upload、add_knowledge;
5. 验证知识库条目存在。
上传格式与 API 参数分别见:
- `ima-format-quickref.md`
- `ima-upload-api.md`
- `ima-credential-chain.md`
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
## 6. CLI fallback 原则
CLI 仅在以下场景使用:
- MCP 服务不可用;
- Tool transport/launch 失败且无法取得有效 Job;
- 用户明确要求本地调试;
- 人工排障需要直接检查脚本输出。
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。
@@ -0,0 +1,27 @@
# IMA 凭证与安全边界
## 必需配置
- `IMA_OPENAPI_CLIENTID`
- `IMA_OPENAPI_APIKEY`
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
- `IMA_DAILY_KNOWLEDGE_BASE_NAME=daily`
优先使用当前进程环境和 IMA Skill 已支持的凭证加载机制。不要在本 Skill 中复制、迁移或重写密钥文件。
## 缺失处理
Preflight 返回凭证缺失或目标知识库无法解析时:
1. 停止上传;
2. 只报告缺失的变量名或配置项;
3. 等待用户或运行环境补齐配置;
4. 配置恢复后重新执行 preflight,不重复生成知识笔记。
## 安全边界
- 不在聊天、日志或命令输出中打印完整 API key、KB ID 或 COS 临时凭证;
- `create_media` 返回的 COS 凭证仅在同一受控进程内传给上传工具,不写入磁盘;
- 不通过拼接 Shell 字符串传递凭证,使用参数数组或 IMA Skill 的封装;
- 不绕过 Hermes 的脱敏机制;若现有工具链无法安全传递凭证,停止并报告;
- 上传结束后不持久化 COS 临时凭证。
@@ -0,0 +1,47 @@
# IMA Markdown 格式速查
## 文件与标题
- 文件名:`<完整文章标题>.md`
- `add_knowledge.title`:完整文章标题,不包含 `.md`
- 上传类型:Markdown 文件,`media_type=7`
- 目标:`daily` knowledge base
## 内容来源
只能使用当前 Run 的可追踪内容:
1. extracted Artifact 的 `article.plain_text`;
2. 对应文章的结构化摘要;
3. digest brief 的 summary 与 highlights。
不得使用通用知识补写原文没有的信息,不设置固定字数或字节数门槛。内容较短时保持简洁并忠于来源。
## 标准结构
```markdown
# 完整文章标题
Source: https://原文链接
Category: 分类
## 核心结论
## 主要论点
## 关键方法 / 机制
## 重要细节
## 可复用启发
## 关键词
## 主题
```
- 核心结论和主要论点使用连贯段落;
- 方法、细节和启发按完整知识点分项;
- 没有来源支持的 Section 可以简写,不得编造内容填充。
上传 API 见 `ima-upload-api.md`。
@@ -0,0 +1,126 @@
# IMA Markdown 上传 API
用于将用户选中的单篇 Markdown 知识笔记上传到 `daily` knowledge base。
## 凭证
- `IMA_OPENAPI_CLIENTID`
- `IMA_OPENAPI_APIKEY`
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
凭证定位和恢复见 `ima-credential-chain.md`。不要在终端输出完整密钥。
## 上传前检查
- 文件名为 `<完整文章标题>.md`;
- `title` 为完整文章标题,不带 `.md`;
- Markdown 符合 `ima-format-quickref.md`;
- 内容可追溯到当前 Run Artifact;
- 用户已经明确选择该文章;
- 目标知识库已经解析并验证。
## 1. Preflight
调用 IMA Skill 的 `preflight-check.cjs` 检查文件类型、扩展名、大小和 MIME。
预期:
```text
file_ext=md
content_type=text/markdown
media_type=7
```
### ⚠️ IMA skill 版本拦截(-200)
`ima_api.cjs` 每天首次调用会检查更新,若检测到新版(如 1.1.8 > 当前 1.1.7)会以 `code=-200` 拦截原请求。注意:**官方 zip 包内的 `meta.json` 可能没同步版本号**(下载 1.1.8 zip 后 meta 仍写 1.1.7),所以光替换文件无法跳过拦截。
快速修复(脚本本身已是新版,只差版本号):
```bash
cd /root/.hermes/skills/openclaw-imports/ima-skill && python3 -c "
import json
m = json.load(open('meta.json')); m['version'] = '1.1.8'
json.dump(m, open('meta.json','w'), ensure_ascii=False, indent=2)
"
```
先用 `diff -rq` 对比 zip 与安装目录:若只有 `.DS_Store`/meta 差异,说明代码已是最新,直接改 meta.json 版本号即可;若脚本有实质差异才需要整体替换。
## 2. Create Media
```text
POST /openapi/wiki/v1/create_media
```
请求核心字段:
```json
{
"file_name": "<完整文章标题>.md",
"file_size": 0,
"content_type": "text/markdown",
"knowledge_base_id": "<daily-kb-id>",
"file_ext": "md"
}
```
保存返回的 `media_id` 和 `cos_credential`。COS 临时凭证只在进程内传递,不打印到聊天或日志。
## 3. COS Upload
使用 IMA Skill 提供的 `cos-upload.cjs`,通过参数数组调用并检查:
- 进程 `returncode`;
- `stderr`;
- HTTP 上传结果。
不要拼接包含凭证的 Shell 字符串,也不要把多条 JSON 响应重定向到同一个文件。
## 4. Add Knowledge
```text
POST /openapi/wiki/v1/add_knowledge
```
核心字段:
```json
{
"media_type": 7,
"media_id": "<media-id>",
"title": "<完整文章标题>",
"knowledge_base_id": "<daily-kb-id>",
"file_info": {
"cos_key": "<cos-key>",
"file_size": 0,
"file_name": "<完整文章标题>.md"
}
}
```
## 5. 验证
上传成功不能只依据本地命令退出码。验证目标 knowledge base 中存在对应标题或返回对象,并向用户报告最终结果。
失败时明确报告发生在 Preflight、Create Media、COS Upload 或 Add Knowledge 的哪一步,不改用 URL 导入或 Notes 类型绕过错误。
## 6. 已知坑:IMA skill 版本拦截(-200)
`ima_api.cjs` 每天首次调用会检查远端版本,若发现新版本(如 1.1.8 > 1.1.7)会以 exit=1 + stderr `{"code":-200}` 拦截**所有** API 调用,原请求不发送。此前遇到过。
处理方式(不必整包替换):
1. 按 stderr 提示下载新版 zip(如 `https://app-dl.ima.qq.com/skills/ima-skills-1.1.8.zip`)并解压;
2. 对比新旧 `ima_api.cjs` 的 md5——zip 内核心脚本常与本地一致,只是 `meta.json` 的 `version` 未同步(zip 内仍写 1.1.7);
3. 若 `ima_api.cjs` 一致,只需把本地 `meta.json` 的 `version` 改为远端版本号即可跳过拦截,无需替换文件。
调用成功后再执行本文件前面的上传流程。
## 6. 版本拦截与批量上传实测(2026-07-31)
- **`-200` skill 更新拦截**:`ima_api.cjs` 每天首次调用检查版本,发现新版时以 code -200 退出并提示更新。下载 zip 后**先对比 `ima_api.cjs` 的 md5**——实测 zip 内脚本与已装版本完全一致,只是 `meta.json` 版本号未同步。此时只需把 `~/.hermes/skills/openclaw-imports/ima-skill/meta.json` 的 `version` 改为最新版即可跳过拦截,无需替换任何脚本。
- **Python 脚本编排上传**比 bash 可靠:bash 拼接含中文文件名/凭证的 curl 易出错。用 `subprocess` 参数数组依次调 `preflight-check.cjs` → `ima_api.cjs check_repeated_names` → `create_media` → `cos-upload.cjs`(`--secret-id/--secret-key/--token` 走参数数组,不打印)→ `add_knowledge`,每步解析返回 JSON,失败即停。
- **批量上传**:4 篇逐个跑同一脚本即可;同名文件先 `check_repeated_names` 确认无重复。
- 凭证从 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` 环境变量读取(`ima_api.cjs` 自动加载),KB ID 用 `IMA_DAILY_KNOWLEDGE_BASE_ID`。
@@ -0,0 +1,29 @@
# 关键词治理路由
关键词治理不属于 `reader-digest-flow` 的日常执行阶段。
仅当用户明确要求“清理关键词”“词库治理”“生成关键词建议”时,委托:
```text
skills/keyword-cleanup-review/SKILL.md
```
Review 输入生成由该 Skill 定义,后续确认与 Apply 由当前编排层负责:
```text
build review bundle
→ generate rule suggestions
→ optional semantic suggestions
→ human review
→ dry-run
→ apply accepted suggestions
```
约束:
- Suggestions JSON 是 Review 与 Apply 之间的正式契约;
- LLM 语义建议不能自动 Apply;
- 不直接编辑 aliases、stopwords、watchlist 或 interest 配置;
- 不在日报主流程中因 tag 质量不佳自动触发治理。
Review bundle、Suggestions、Schema 和产物保留策略以 `keyword-cleanup-review` 为唯一事实来源;该 Skill 不直接 Apply 配置。
@@ -0,0 +1,31 @@
+++
title = "AI 日报 · 示例"
date = 2026-04-01T09:00:00+08:00
summary = "围绕 Agent 架构分层、Skills 标准化与工程实践的当日观察。"
+++
# 今日概览
今天的公开内容主要集中在 AI Agent 架构演进、工具化落地与工程实践三条线索。行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
## 今日重点
### 1. 从 Agent 到 Skills:AI 智能体架构的范式转变
文章分析了 AI 智能体从单体 Agent 向模块化 Skills 的演进,并结合 MCP、Skills 和真实项目说明能力分层与复用方式。
值得关注:
- Skills 将领域流程从 Agent 主体中拆出,便于复用和维护。
- MCP 为 Agent 与外部工具提供标准化连接方式。
- 工程竞争点逐渐从模型调用转向状态、工具和工作流设计。
这篇内容值得关注的原因在于,它把开放协议、分层架构和真实落地案例连接成了完整论证链。
来源:[示例来源](https://example.com/a)
## 趋势观察
1. Agent 正在从单体能力转向可组合的模块化体系。
2. 工具契约、状态管理和验证机制正在成为 AI 应用的核心工程能力。
3. Human-in-the-loop 仍是控制高风险副作用的重要边界。
@@ -0,0 +1,96 @@
#!/usr/bin/env python3
"""IMA 上传单篇 Markdown 知识笔记到 daily 知识库。
用法: python3 ima_upload_one.py "<绝对路径/summary.md>" "<完整文章标题>"
依赖环境变量: IMA_OPENAPI_CLIENTID / IMA_OPENAPI_APIKEY / IMA_DAILY_KNOWLEDGE_BASE_ID
流程: preflight -> check_repeated_names -> create_media -> cos-upload -> add_knowledge
说明: 用 Python 而非 bash 编排,避免中文文件名/引号转义问题。
退出码 2 = 文件名重复(需与用户确认保留双方或取消),非 0 均为失败。
"""
import json, os, subprocess, sys
SKILL_DIR = "/root/.hermes/skills/openclaw-imports/ima-skill"
IMA_API = os.path.join(SKILL_DIR, "ima_api.cjs")
COS_UPLOAD = os.path.join(SKILL_DIR, "knowledge-base/scripts/cos-upload.cjs")
PREFLIGHT = os.path.join(SKILL_DIR, "knowledge-base/scripts/preflight-check.cjs")
def run_node(script, args):
r = subprocess.run(["node", script] + args, capture_output=True, text=True, timeout=120)
if r.returncode != 0:
raise RuntimeError(f"{script} exit={r.returncode} stderr={r.stderr[:500]}")
return json.loads(r.stdout)
def ima_api(api_path, body):
r = subprocess.run(["node", IMA_API, api_path, json.dumps(body, ensure_ascii=False)],
capture_output=True, text=True, timeout=120)
if r.returncode != 0:
raise RuntimeError(f"ima_api {api_path} exit={r.returncode} stderr={r.stderr[:500]}")
resp = json.loads(r.stdout)
if resp.get("code") != 0:
raise RuntimeError(f"ima_api {api_path} code={resp.get('code')} msg={resp.get('msg')}")
return resp.get("data", {})
def main():
kb_id = os.environ["IMA_DAILY_KNOWLEDGE_BASE_ID"]
file_path = sys.argv[1]
title = sys.argv[2] # 完整文章标题(不含 .md)
pf = run_node(PREFLIGHT, ["--file", file_path])
if not pf.get("pass"):
raise RuntimeError(f"preflight failed: {pf}")
file_name = pf["file_name"]; media_type = pf["media_type"]
content_type = pf["content_type"]; file_size = pf["file_size"]; file_ext = pf["file_ext"]
print(f"[preflight] ok file={file_name} ext={file_ext} size={file_size} media_type={media_type}")
dup = ima_api("openapi/wiki/v1/check_repeated_names", {
"params": [{"name": file_name, "media_type": media_type}],
"knowledge_base_id": kb_id
})
is_rep = dup.get("results", [{}])[0].get("is_repeated", False) if dup.get("results") else False
if is_rep:
print(f"[check_repeated_names] REPEATED: {file_name} — 需要处理")
sys.exit(2)
print("[check_repeated_names] no duplicate")
cm = ima_api("openapi/wiki/v1/create_media", {
"file_name": file_name,
"file_size": file_size,
"content_type": content_type,
"knowledge_base_id": kb_id,
"file_ext": file_ext
})
media_id = cm["media_id"]
cos = cm["cos_credential"]
print(f"[create_media] media_id={media_id} cos_bucket={cos.get('bucket_name')} cos_key={cos.get('cos_key','')[:40]}")
r = subprocess.run(["node", COS_UPLOAD,
"--file", file_path,
"--secret-id", cos["secret_id"],
"--secret-key", cos["secret_key"],
"--token", cos["token"],
"--bucket", cos["bucket_name"],
"--region", cos["region"],
"--cos-key", cos["cos_key"],
"--content-type", content_type,
"--start-time", str(cos.get("start_time", "")),
"--expired-time", str(cos.get("expired_time", "")),
"--timeout", "300000"
], capture_output=True, text=True, timeout=360)
if r.returncode != 0:
raise RuntimeError(f"cos-upload exit={r.returncode} stderr={r.stderr[:800]}")
print(f"[cos-upload] ok rc=0 stdout={r.stdout.strip()[:200]}")
ak = ima_api("openapi/wiki/v1/add_knowledge", {
"media_type": media_type,
"media_id": media_id,
"title": title,
"knowledge_base_id": kb_id,
"file_info": {
"cos_key": cos["cos_key"],
"file_size": file_size,
"file_name": file_name
}
})
print(f"[add_knowledge] ok media_id={ak.get('media_id') or media_id}")
if __name__ == "__main__":
main()
+44 -3
View File
@@ -10,6 +10,7 @@ from pathlib import Path
from typing import Any, Callable
import httpx
from dotenv import dotenv_values
from summary_mcp.validators.llm_result import ValidationReport
from summary_mcp.validators.llm_result import validate_llm_result as validate_llm_result_from_path
@@ -18,6 +19,8 @@ from summary_mcp.validators.llm_result import validate_llm_result_payload
JSON_BLOCK_RE = re.compile(r"```(?:json)?\s*(\{.*\})\s*```", re.DOTALL)
DEFAULT_CHAT_COMPLETIONS_URL = "https://api.openai.com/v1/chat/completions"
REPO_ROOT = Path(__file__).resolve().parents[3]
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
def load_text(path: Path) -> str:
@@ -99,22 +102,60 @@ def normalize_chat_completions_url(api_url: str | None) -> str | None:
return f"{normalized}/chat/completions"
def _load_repo_dotenv() -> dict[str, str]:
if not DEFAULT_DOTENV_PATH.exists():
return {}
return {
key: value
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items()
if isinstance(key, str) and isinstance(value, str) and value
}
def _pick_value(*values: str | None) -> str | None:
for value in values:
if value:
return value
return None
def resolve_llm_settings(
*,
api_key: str | None = None,
model: str | None = None,
api_url: str | None = None,
) -> tuple[str, str, str]:
resolved_api_key = api_key or os.environ.get("LLM_API_KEY") or os.environ.get("OPENAI_API_KEY")
dotenv_map = _load_repo_dotenv()
resolved_api_key = _pick_value(
api_key,
os.environ.get("LLM_API_KEY"),
os.environ.get("OPENAI_API_KEY"),
dotenv_map.get("LLM_API_KEY"),
dotenv_map.get("OPENAI_API_KEY"),
)
if not resolved_api_key:
raise RuntimeError("Missing LLM_API_KEY or OPENAI_API_KEY, or pass an API key.")
resolved_model = model or os.environ.get("LLM_MODEL") or os.environ.get("OPENAI_MODEL")
resolved_model = _pick_value(
model,
os.environ.get("LLM_MODEL"),
os.environ.get("OPENAI_MODEL"),
dotenv_map.get("LLM_MODEL"),
dotenv_map.get("OPENAI_MODEL"),
)
if not resolved_model:
raise RuntimeError("Missing LLM_MODEL or OPENAI_MODEL, or pass a model.")
resolved_api_url = normalize_chat_completions_url(
api_url or os.environ.get("LLM_API_URL") or os.environ.get("OPENAI_API_URL") or DEFAULT_CHAT_COMPLETIONS_URL
_pick_value(
api_url,
os.environ.get("LLM_API_URL"),
os.environ.get("OPENAI_API_URL"),
dotenv_map.get("LLM_API_URL"),
dotenv_map.get("OPENAI_API_URL"),
DEFAULT_CHAT_COMPLETIONS_URL,
)
)
if not resolved_api_url:
raise RuntimeError("Missing LLM_API_URL, OPENAI_API_URL, or pass an API URL.")
+1
View File
@@ -228,6 +228,7 @@ class FreshRSSClient:
params: list[tuple[str, str | int]] = [
("output", "json"),
("n", limit),
("r", "n")
]
if continuation:
params.append(("c", continuation))
@@ -66,6 +66,7 @@ class ArticleCandidateRecord(BaseModel):
class OpenClawCandidateInput(BaseModel):
candidate_id: str
item_id: str | None = None
title: str
url: HttpUrl
canonical_url: HttpUrl | None = None
@@ -169,6 +170,7 @@ def build_openclaw_candidate_input(record: ArticleCandidateRecord) -> OpenClawCa
return OpenClawCandidateInput(
candidate_id=record.candidate_id,
item_id=item.item_id if item is not None else None,
title=title,
url=raw_url,
canonical_url=normalize_candidate_url(raw_url),
+11 -5
View File
@@ -1,12 +1,17 @@
from __future__ import annotations
from datetime import UTC, date, datetime
try:
from datetime import UTC, date, datetime
except ImportError: # Python < 3.11 compatibility
from datetime import timezone, date, datetime
UTC = timezone.utc
from pydantic import BaseModel, Field
from .article_candidate import OpenClawCandidateInput
DIGEST_BRIEF_LIMIT = 5
DIGEST_BRIEF_LIMIT: int | None = None
DIGEST_BRIEF_HIGHLIGHT_LIMIT = 3
@@ -76,11 +81,12 @@ def build_openclaw_delivery_payload(
def build_openclaw_digest_brief(
payload: OpenClawDeliveryPayload,
*,
limit: int = DIGEST_BRIEF_LIMIT,
limit: int | None = DIGEST_BRIEF_LIMIT,
highlight_limit: int = DIGEST_BRIEF_HIGHLIGHT_LIMIT,
) -> OpenClawDigestBrief:
keep_candidates = [candidate for candidate in payload.candidates if candidate.selection_decision == "keep"]
keep_candidates = [candidate for candidate in payload.candidates if candidate.selection_decision in ("keep", "review")]
sorted_candidates = sorted(keep_candidates, key=lambda candidate: candidate.digest_rank, reverse=True)
selected_candidates = sorted_candidates if limit is None else sorted_candidates[:limit]
top_candidates = [
DigestBriefCandidate(
title=candidate.title,
@@ -92,7 +98,7 @@ def build_openclaw_digest_brief(
selection_decision=candidate.selection_decision,
url=str(candidate.url),
)
for candidate in sorted_candidates[:limit]
for candidate in selected_candidates
]
return OpenClawDigestBrief(
schema_version="digest-brief.v1",
+38 -1
View File
@@ -1,10 +1,43 @@
from .query_service import get_delivery_payload, get_run_report, get_run_status, list_run_artifacts, list_runs
from .run_store import RunStore
from .state_models import ArtifactRecord, RecoveryState, RunError, RunState, StageState
from .freshrss_pipeline_jobs import (
get_freshrss_pipeline_job_result,
get_freshrss_pipeline_job_status,
start_freshrss_pipeline_job,
)
from .query_service import get_delivery_payload, get_run_report, get_run_status, list_run_artifacts, list_runs
# NOTE: resume_jobs and resume_service are NOT eagerly imported here to avoid
# a circular import chain:
# workflows/freshrss_pipeline.py -> runtime -> resume_jobs -> resume_service
# -> workflows/freshrss_pipeline.py (circular!)
# They are lazy-loaded via __getattr__ when accessed as summary_mcp.runtime.*
def __getattr__(name):
import importlib
_LAZY = {
"get_resume_job_result": ("resume_jobs", "get_resume_job_result"),
"get_resume_job_status": ("resume_jobs", "get_resume_job_status"),
"start_resume_job": ("resume_jobs", "start_resume_job"),
"inspect_resume_plan": ("resume_service", "inspect_resume_plan"),
"resume_run": ("resume_service", "resume_run"),
}
if name in _LAZY:
mod_name, attr_name = _LAZY[name]
mod = importlib.import_module(f".{mod_name}", __package__)
return getattr(mod, attr_name)
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
__all__ = [
"ArtifactRecord",
"get_delivery_payload",
"get_freshrss_pipeline_job_result",
"get_freshrss_pipeline_job_status",
"get_resume_job_result",
"get_resume_job_status",
"get_run_report",
"RecoveryState",
"RunError",
@@ -12,6 +45,10 @@ __all__ = [
"RunStore",
"StageState",
"get_run_status",
"inspect_resume_plan",
"list_run_artifacts",
"list_runs",
"resume_run",
"start_resume_job",
"start_freshrss_pipeline_job",
]
@@ -0,0 +1,346 @@
from __future__ import annotations
import json
import os
import subprocess
import sys
from concurrent.futures import ThreadPoolExecutor
from datetime import datetime
from pathlib import Path
from typing import Any
from uuid import uuid4
from .run_store import RunStore
from .state_models import RunState
# 后台线程池:用于异步启动 job,避免阻塞 MCP stdio 响应
# max_workers=4 足够应对并发请求,线程复用降低启动开销
_executor = ThreadPoolExecutor(max_workers=4, thread_name_prefix="article_summary_job")
REPO_ROOT = Path(__file__).resolve().parents[3]
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
ARTICLE_SUMMARY_JOBS_ROOT = OUTPUT_ROOT / "article_summary_jobs"
WORKFLOW_NAME = "article_summary_job"
RUN_TYPE = "article_summary"
RUN_STATE_FILENAME = "run-state.json"
INPUT_FILENAME = "input.json"
RESULT_FILENAME = "result.json"
JOB_REPORT_FILENAME = "job-report.json"
DEFAULT_STAGES = [
"prepare_job",
"load_input",
"generate_markdown",
"write_result",
]
def _now() -> datetime:
return datetime.now().astimezone()
def _new_job_id() -> str:
ts = _now().strftime("%Y%m%d-%H%M%S")
return f"article-summary-{ts}-{uuid4().hex[:8]}"
def _job_dir(job_id: str) -> Path:
return ARTICLE_SUMMARY_JOBS_ROOT / job_id
def _run_state_path(job_id: str) -> Path:
return _job_dir(job_id) / RUN_STATE_FILENAME
def _input_path(job_id: str) -> Path:
return _job_dir(job_id) / INPUT_FILENAME
def _result_path(job_id: str) -> Path:
return _job_dir(job_id) / RESULT_FILENAME
def _job_report_path(job_id: str) -> Path:
return _job_dir(job_id) / JOB_REPORT_FILENAME
def _normalize_repo_path(path: Path) -> str:
try:
return str(path.resolve().relative_to(REPO_ROOT.resolve()))
except ValueError:
return str(path)
def _load_run_store(job_id: str) -> RunStore:
return RunStore.load(path=_run_state_path(job_id), repo_root=REPO_ROOT)
def _load_result(job_id: str) -> dict[str, Any] | None:
path = _result_path(job_id)
if not path.exists():
return None
return json.loads(path.read_text(encoding="utf-8-sig"))
def _validate_article_summary_extracted_path(extracted_path: Path) -> None:
if not extracted_path.exists():
raise FileNotFoundError(f"extracted_path does not exist: {extracted_path}")
if extracted_path.is_dir():
raise ValueError(
"extracted_path must be a JSON file, not a directory. "
"Pass a single extracted file like outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json, "
"or a batch extracted JSON file."
)
def _launch_job_background(*, job_id: str, input_payload: dict[str, Any], store: RunStore) -> None:
"""后台线程执行:启动 subprocess 并完成 run_state 写盘。
主线程已经返回 MCP 响应,后台负责:
1. 启动 runner 子进程(Popen)
2. 记录 prepare_job 阶段完成
3. 持久化 run-state.json
"""
runner_script = REPO_ROOT / "scripts" / "run_article_summary_job.py"
cmd = [sys.executable, str(runner_script), "--job-id", job_id]
# 启动子进程(后台运行,不等待)
proc = subprocess.Popen(
cmd,
cwd=str(REPO_ROOT),
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
start_new_session=True,
)
# 记录 prepare_job 完成
store.finish_stage("prepare_job", outputs={"runner_pid": proc.pid, "runner_command": cmd})
def start_article_summary_job(
*,
extracted_path: Path,
selected_ids: list[str],
output_dir: Path | None = None,
max_retries: int = 2,
timeout_seconds: float = 120.0,
llm_api_key: str | None = None,
llm_model: str | None = None,
llm_api_url: str | None = None,
) -> dict[str, Any]:
"""Start an asynchronous article-summary job and return a job_id immediately.
关键:使用后台线程异步启动 job,主线程立即返回 MCP 响应,避免 stdio 阻塞。
"""
_validate_article_summary_extracted_path(extracted_path)
if not selected_ids:
raise ValueError("selected_ids must not be empty")
job_id = _new_job_id()
job_dir = _job_dir(job_id)
job_dir.mkdir(parents=True, exist_ok=True)
started_at = _now()
resolved_output_dir = output_dir or (extracted_path.parent / "single_summaries")
input_payload = {
"extracted_path": str(extracted_path),
"selected_ids": selected_ids,
"output_dir": str(resolved_output_dir),
"max_retries": max_retries,
"timeout_seconds": timeout_seconds,
"llm_api_key": llm_api_key,
"llm_model": llm_model,
"llm_api_url": llm_api_url,
"launcher_pid": os.getpid(),
}
# 创建 run-state(只写初始状态,不启动 subprocess)
store = RunStore.create(
path=_run_state_path(job_id),
run_id=job_id,
workflow=WORKFLOW_NAME,
run_type=RUN_TYPE,
started_at=started_at,
input_payload=input_payload,
repo_root=REPO_ROOT,
)
for stage_name in DEFAULT_STAGES:
store._get_or_create_stage(stage_name)
# 注册输入文件
input_file = _input_path(job_id)
input_file.write_text(json.dumps(input_payload, ensure_ascii=False, indent=2), encoding="utf-8")
# 启动 prepare_job 阶段(只写状态,不调用 save() —— 交给后台线程)
store.start_stage("prepare_job")
store.register_artifact(name="job_input", path=input_file, kind="json", stage="prepare_job")
# 【关键改造】:后台线程执行 Popen + finish_stage,主线程立即返回
_executor.submit(_launch_job_background, job_id=job_id, input_payload=input_payload, store=store)
return {
"job_id": job_id,
"workflow": WORKFLOW_NAME,
"run_type": RUN_TYPE,
"status": "running",
"output_dir": _normalize_repo_path(job_dir),
"message": "Article summary job started successfully. Use get_article_summary_job_status to poll progress.",
}
def run_article_summary_job(*, job_id: str) -> dict[str, Any]:
from summary_mcp.workflows.article_summary import ArticleSummaryConfig, summarize_selected_articles
store = _load_run_store(job_id)
input_payload = json.loads(_input_path(job_id).read_text(encoding="utf-8-sig"))
try:
store.start_stage("load_input")
extracted_path = Path(input_payload["extracted_path"])
output_dir = Path(input_payload["output_dir"])
selected_ids = list(input_payload["selected_ids"])
_validate_article_summary_extracted_path(extracted_path)
if not selected_ids:
raise ValueError("selected_ids must not be empty")
store.finish_stage(
"load_input",
outputs={
"extracted_path": _normalize_repo_path(extracted_path),
"selected_id_count": len(selected_ids),
"output_dir": _normalize_repo_path(output_dir),
},
)
store.start_stage("generate_markdown")
config = ArticleSummaryConfig(
max_retries=int(input_payload.get("max_retries") or 2),
timeout_seconds=float(input_payload.get("timeout_seconds") or 120.0),
)
written_paths = summarize_selected_articles(
extracted_path=extracted_path,
selected_ids=selected_ids,
output_dir=output_dir,
config=config,
api_key=input_payload.get("llm_api_key"),
model=input_payload.get("llm_model"),
api_url=input_payload.get("llm_api_url"),
)
if not written_paths:
raise RuntimeError("Article summary job produced no Markdown outputs.")
normalized_paths = [_normalize_repo_path(Path(p)) for p in written_paths]
store.finish_stage(
"generate_markdown",
outputs={
"written_count": len(written_paths),
"written_paths": normalized_paths,
},
)
for idx, p in enumerate(written_paths, start=1):
store.register_artifact(
name=f"summary_markdown_{idx}",
path=Path(p),
kind="markdown",
stage="generate_markdown",
metadata={"output_type": "single_summary"},
)
store.start_stage("write_result")
result = {
"job_id": job_id,
"selected_ids": selected_ids,
"written_paths": normalized_paths,
"completed_at": _now().isoformat(),
}
result_file = _result_path(job_id)
result_file.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
store.register_artifact(name="job_result", path=result_file, kind="json", stage="write_result")
report_file = _job_report_path(job_id)
report_file.write_text(
json.dumps(
{
"job_id": job_id,
"status": "success",
"written_count": len(normalized_paths),
"written_paths": normalized_paths,
},
ensure_ascii=False,
indent=2,
),
encoding="utf-8",
)
store.register_artifact(name="job_report", path=report_file, kind="json", stage="write_result")
store.finish_stage("write_result", outputs={"result_path": _normalize_repo_path(result_file)})
store.finish_run(status="success")
return result
except Exception as exc:
current_stage = store.state.current_stage or "generate_markdown"
store.fail_stage(current_stage, error=exc)
report_file = _job_report_path(job_id)
report_file.write_text(
json.dumps(
{
"job_id": job_id,
"status": "failed",
"error_type": type(exc).__name__,
"error_message": str(exc),
"failed_stage": current_stage,
},
ensure_ascii=False,
indent=2,
),
encoding="utf-8",
)
try:
store.register_artifact(name="job_report", path=report_file, kind="json", stage=current_stage)
except Exception:
pass
raise
def get_article_summary_job_status(*, job_id: str) -> dict[str, Any]:
store = _load_run_store(job_id)
state = store.state
completed_stage_count = sum(1 for s in state.stages if s.status == "success")
running_stage_count = sum(1 for s in state.stages if s.status == "running")
failed_stage_count = sum(1 for s in state.stages if s.status == "failed")
pending_stage_count = sum(1 for s in state.stages if s.status == "pending")
return {
"job_id": state.run_id,
"workflow": state.workflow,
"run_type": state.run_type,
"status": state.status,
"current_stage": state.current_stage,
"started_at": state.started_at.isoformat(),
"updated_at": state.updated_at.isoformat(),
"finished_at": state.finished_at.isoformat() if state.finished_at else None,
"output_dir": _normalize_repo_path(_job_dir(job_id)),
"progress": {
"completed_stage_count": completed_stage_count,
"running_stage_count": running_stage_count,
"failed_stage_count": failed_stage_count,
"pending_stage_count": pending_stage_count,
"total_stage_count": len(state.stages),
},
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
"error_summary": state.error.model_dump(mode="json") if state.error else None,
}
def get_article_summary_job_result(*, job_id: str) -> dict[str, Any]:
store = _load_run_store(job_id)
state = store.state
result = _load_result(job_id)
if state.status != "success" or result is None:
return {
"job_id": state.run_id,
"status": state.status,
"message": "Article summary job result is not ready.",
"error_summary": state.error.model_dump(mode="json") if state.error else None,
}
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
return {
"job_id": state.run_id,
"status": state.status,
"written_paths": result.get("written_paths", []),
"artifact": artifact,
"result": result,
}
@@ -0,0 +1,672 @@
from __future__ import annotations
import json
import os
import subprocess
import sys
from datetime import UTC, datetime
from pathlib import Path
from typing import Any
from uuid import uuid4
from dotenv import dotenv_values
from .run_store import RunStore
from .query_service import _resolve_run_record
REPO_ROOT = Path(__file__).resolve().parents[3]
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
PIPELINE_JOBS_ROOT = OUTPUT_ROOT / "pipeline_jobs"
WORKFLOW_NAME = "freshrss_pipeline_job"
RUN_TYPE = "daily_digest"
RUN_STATE_FILENAME = "run-state.json"
INPUT_FILENAME = "input.json"
RESULT_FILENAME = "result.json"
JOB_REPORT_FILENAME = "job-report.json"
DEFAULT_STAGES = [
"prepare_job",
"load_input",
"run_pipeline",
"write_result",
]
MIN_JOB_STALE_SECONDS = 30 * 60
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
def _build_subprocess_env() -> dict[str, str]:
"""Build an env dict for subprocess, merging parent env with .env values.
The subprocess inherits the Hermes MCP server's environment, but .env values
may not be in os.environ at the time the subprocess is spawned. This function
loads them from .env and merges them in so the child process sees all needed
variables (LLM_API_KEY, LLM_MODEL, LLM_API_URL, FRESHRSS_*, etc.) directly
in os.environ, avoiding any dotenv-loading timing issues inside the subprocess.
"""
env = os.environ.copy()
if DEFAULT_DOTENV_PATH.exists():
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items():
if isinstance(key, str) and isinstance(value, str) and value:
# Only set if not already present in parent env
env.setdefault(key, value)
return env
def _now() -> datetime:
return datetime.now().astimezone()
def _new_job_id() -> str:
ts = _now().strftime("%Y%m%d-%H%M%S")
return f"freshrss-pipeline-job-{ts}-{uuid4().hex[:8]}"
def _job_dir(job_id: str) -> Path:
return PIPELINE_JOBS_ROOT / job_id
def _run_state_path(job_id: str) -> Path:
return _job_dir(job_id) / RUN_STATE_FILENAME
def _input_path(job_id: str) -> Path:
return _job_dir(job_id) / INPUT_FILENAME
def _result_path(job_id: str) -> Path:
return _job_dir(job_id) / RESULT_FILENAME
def _job_report_path(job_id: str) -> Path:
return _job_dir(job_id) / JOB_REPORT_FILENAME
def _normalize_repo_path(path: Path) -> str:
try:
return str(path.resolve().relative_to(REPO_ROOT.resolve()))
except ValueError:
return str(path)
def _normalize_repo_path_value(path_value: str | None) -> str | None:
if not path_value:
return None
return _normalize_repo_path(Path(path_value))
def _normalize_keyword_index(keyword_index: dict[str, Any] | None) -> dict[str, Any]:
if not isinstance(keyword_index, dict):
return {}
normalized = dict(keyword_index)
for field in ("daily_output", "stats_output"):
normalized[field] = _normalize_repo_path_value(normalized.get(field))
return normalized
def _load_run_store(job_id: str) -> RunStore:
return RunStore.load(path=_run_state_path(job_id), repo_root=REPO_ROOT)
def _load_result(job_id: str) -> dict[str, Any] | None:
path = _result_path(job_id)
if not path.exists():
return None
return json.loads(path.read_text(encoding="utf-8-sig"))
def _build_run_defaults(*, run_id: str | None, output_dir: Path | None) -> tuple[str, Path]:
resolved_output_dir = output_dir
resolved_run_id = run_id
if resolved_run_id is not None and resolved_output_dir is not None:
return resolved_run_id, resolved_output_dir
run_stamp = datetime.now(tz=UTC).strftime("%Y%m%d-%H%M%S")
if resolved_run_id is None:
resolved_run_id = f"freshrss-pipeline-{run_stamp}"
if resolved_output_dir is None:
resolved_output_dir = OUTPUT_ROOT / "rerun" / run_stamp
return resolved_run_id, resolved_output_dir
def _build_result_payload(*, job_id: str, pipeline_result: dict[str, Any], include_item_reports: bool) -> dict[str, Any]:
result = {
"job_id": job_id,
"run_id": pipeline_result["run_id"],
"output_dir": _normalize_repo_path_value(pipeline_result.get("output_dir")),
"raw_output": _normalize_repo_path_value(pipeline_result.get("raw_output")),
"delivery_output": _normalize_repo_path_value(pipeline_result.get("delivery_output")),
"digest_brief_output": _normalize_repo_path_value(pipeline_result.get("digest_brief_output")),
"report_output": _normalize_repo_path_value(pipeline_result.get("report_output")),
"keyword_index": _normalize_keyword_index(pipeline_result.get("keyword_index")),
"pulled_count": pipeline_result.get("pulled_count"),
"delivered_count": pipeline_result.get("delivered_count"),
"marked_read_count": pipeline_result.get("marked_read_count"),
"status_counts": pipeline_result.get("status_counts"),
"debug_artifacts": pipeline_result.get("debug_artifacts"),
"completed_at": _now().isoformat(),
}
if include_item_reports:
result["items"] = pipeline_result.get("items", [])
return result
def _write_job_report(
*,
job_id: str,
payload: dict[str, Any],
) -> Path:
report_file = _job_report_path(job_id)
report_file.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
return report_file
def start_freshrss_pipeline_job(
*,
limit: int = 5,
mark_read: bool = False,
include_read: bool = False,
debug_artifacts: bool = False,
continuation: str | None = None,
timeout_seconds: float = 60.0,
max_retries: int = 2,
stream_id: str = "user/-/state/com.google/reading-list",
api_base_url: str | None = None,
username: str | None = None,
api_password: str | None = None,
llm_api_key: str | None = None,
llm_model: str | None = None,
llm_api_url: str | None = None,
context: dict[str, Any] | None = None,
run_id: str | None = None,
date_value: str | None = None,
output_dir: Path | None = None,
include_item_reports: bool = False,
) -> dict[str, Any]:
job_id = _new_job_id()
job_dir = _job_dir(job_id)
job_dir.mkdir(parents=True, exist_ok=True)
started_at = _now()
resolved_run_id, resolved_output_dir = _build_run_defaults(run_id=run_id, output_dir=output_dir)
input_payload = {
"limit": limit,
"mark_read": mark_read,
"include_read": include_read,
"debug_artifacts": debug_artifacts,
"continuation": continuation,
"timeout_seconds": timeout_seconds,
"max_retries": max_retries,
"stream_id": stream_id,
"api_base_url": api_base_url,
"username": username,
"api_password": api_password,
"llm_api_key": llm_api_key,
"llm_model": llm_model,
"llm_api_url": llm_api_url,
"context": context,
"run_id": resolved_run_id,
"date_value": date_value,
"output_dir": str(resolved_output_dir),
"include_item_reports": include_item_reports,
"launcher_pid": os.getpid(),
}
store = RunStore.create(
path=_run_state_path(job_id),
run_id=job_id,
workflow=WORKFLOW_NAME,
run_type=RUN_TYPE,
started_at=started_at,
input_payload=input_payload,
repo_root=REPO_ROOT,
)
for stage_name in DEFAULT_STAGES:
store._get_or_create_stage(stage_name)
store.save()
store.start_stage("prepare_job")
input_file = _input_path(job_id)
input_file.write_text(json.dumps(input_payload, ensure_ascii=False, indent=2), encoding="utf-8")
store.register_artifact(name="job_input", path=input_file, kind="json", stage="prepare_job")
runner_script = REPO_ROOT / "scripts" / "run_freshrss_pipeline_job.py"
cmd = [sys.executable, str(runner_script), "--job-id", job_id]
try:
proc = subprocess.Popen(
cmd,
cwd=str(REPO_ROOT),
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
start_new_session=True,
env=_build_subprocess_env(),
)
except Exception as exc:
report_file = _write_job_report(
job_id=job_id,
payload={
"job_id": job_id,
"status": "failed",
"error_type": type(exc).__name__,
"error_message": str(exc),
"failed_stage": "prepare_job",
"linked_run_id": resolved_run_id,
"linked_output_dir": _normalize_repo_path(resolved_output_dir),
},
)
store.register_artifact(name="job_report", path=report_file, kind="json", stage="prepare_job")
store.fail_stage("prepare_job", error=exc)
raise
store.finish_stage(
"prepare_job",
outputs={
"runner_pid": proc.pid,
"runner_command": cmd,
"linked_run_id": resolved_run_id,
"linked_output_dir": _normalize_repo_path(resolved_output_dir),
},
)
return {
"job_id": job_id,
"workflow": WORKFLOW_NAME,
"run_type": RUN_TYPE,
"status": "running",
"output_dir": _normalize_repo_path(job_dir),
"linked_run_id": resolved_run_id,
"linked_output_dir": _normalize_repo_path(resolved_output_dir),
"message": "FreshRSS pipeline job started successfully. Use get_freshrss_pipeline_job_status to poll progress.",
}
def run_freshrss_pipeline_job(*, job_id: str) -> dict[str, Any]:
from datetime import date
from summary_mcp.workflows import run_freshrss_pipeline
store = _load_run_store(job_id)
input_payload = json.loads(_input_path(job_id).read_text(encoding="utf-8-sig"))
current_stage = "run_pipeline"
try:
store.start_stage("load_input")
resolved_output_dir = Path(input_payload["output_dir"])
resolved_run_id = str(input_payload["run_id"])
store.finish_stage(
"load_input",
outputs={
"linked_run_id": resolved_run_id,
"linked_output_dir": _normalize_repo_path(resolved_output_dir),
"limit": int(input_payload.get("limit") or 5),
},
)
store.start_stage("run_pipeline")
pipeline_result = run_freshrss_pipeline(
api_base_url=input_payload.get("api_base_url"),
username=input_payload.get("username"),
api_password=input_payload.get("api_password"),
stream_id=input_payload.get("stream_id") or "user/-/state/com.google/reading-list",
limit=int(input_payload.get("limit") or 5),
continuation=input_payload.get("continuation"),
include_read=bool(input_payload.get("include_read")),
mark_read=bool(input_payload.get("mark_read")),
debug_artifacts=bool(input_payload.get("debug_artifacts")),
context=input_payload.get("context"),
max_retries=int(input_payload.get("max_retries") or 2),
timeout_seconds=float(input_payload.get("timeout_seconds") or 60.0),
llm_api_key=input_payload.get("llm_api_key"),
llm_model=input_payload.get("llm_model"),
llm_api_url=input_payload.get("llm_api_url"),
run_id=resolved_run_id,
delivery_date=date.fromisoformat(input_payload["date_value"]) if input_payload.get("date_value") else None,
output_dir=resolved_output_dir,
)
store.finish_stage(
"run_pipeline",
outputs={
"linked_run_id": pipeline_result.get("run_id"),
"linked_output_dir": _normalize_repo_path_value(pipeline_result.get("output_dir")),
"delivery_output": _normalize_repo_path_value(pipeline_result.get("delivery_output")),
"report_output": _normalize_repo_path_value(pipeline_result.get("report_output")),
"delivered_count": pipeline_result.get("delivered_count"),
"marked_read_count": pipeline_result.get("marked_read_count"),
},
)
store.start_stage("write_result")
result = _build_result_payload(
job_id=job_id,
pipeline_result=pipeline_result,
include_item_reports=bool(input_payload.get("include_item_reports")),
)
result_file = _result_path(job_id)
result_file.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
store.register_artifact(name="job_result", path=result_file, kind="json", stage="write_result")
report_file = _write_job_report(
job_id=job_id,
payload={
"job_id": job_id,
"status": "success",
"run_id": result["run_id"],
"delivery_output": result["delivery_output"],
"report_output": result["report_output"],
"digest_brief_output": result["digest_brief_output"],
"pulled_count": result["pulled_count"],
"delivered_count": result["delivered_count"],
"marked_read_count": result["marked_read_count"],
},
)
store.register_artifact(name="job_report", path=report_file, kind="json", stage="write_result")
store.finish_stage(
"write_result",
outputs={
"result_path": _normalize_repo_path(result_file),
"linked_run_id": result["run_id"],
},
)
store.finish_run(status="success")
return result
except Exception as exc:
current_stage = store.state.current_stage or current_stage
report_file = _write_job_report(
job_id=job_id,
payload={
"job_id": job_id,
"status": "failed",
"error_type": type(exc).__name__,
"error_message": str(exc),
"failed_stage": current_stage,
"linked_run_id": input_payload.get("run_id"),
"linked_output_dir": _normalize_repo_path_value(input_payload.get("output_dir")),
},
)
try:
store.register_artifact(name="job_report", path=report_file, kind="json", stage=current_stage)
except Exception:
pass
store.fail_stage(current_stage, error=exc)
raise
def get_freshrss_pipeline_job_status(*, job_id: str) -> dict[str, Any]:
store = _load_run_store(job_id)
state = store.state
effective = _resolve_effective_job_view(job_id=job_id, state=state)
progress = _build_effective_job_progress(state=state, effective_status=effective["status"])
linked_run_id = state.input.get("run_id") if isinstance(state.input, dict) else None
linked_output_dir = _normalize_repo_path_value(state.input.get("output_dir")) if isinstance(state.input, dict) else None
return {
"job_id": state.run_id,
"workflow": state.workflow,
"run_type": state.run_type,
"status": effective["status"],
"current_stage": effective["current_stage"],
"started_at": state.started_at.isoformat(),
"updated_at": effective["updated_at"],
"finished_at": effective["finished_at"],
"output_dir": _normalize_repo_path(_job_dir(job_id)),
"linked_run_id": linked_run_id,
"linked_output_dir": linked_output_dir,
"progress": progress,
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
"error_summary": state.error.model_dump(mode="json") if state.error else None,
"status_source": effective["status_source"],
"state_quality": effective["state_quality"],
"state_conflict": effective["state_conflict"],
"state_conflict_reason": effective["state_conflict_reason"],
"status_note": effective["status_note"],
"raw_status": state.status,
"raw_current_stage": state.current_stage,
"linked_run_status": effective["linked_run_status"],
"linked_run_state_conflict": effective["linked_run_state_conflict"],
}
def get_freshrss_pipeline_job_result(*, job_id: str) -> dict[str, Any]:
store = _load_run_store(job_id)
state = store.state
result = _load_result(job_id)
effective = _resolve_effective_job_view(job_id=job_id, state=state, result=result)
synthesized_result = result or effective["result"]
if effective["status"] != "success" or synthesized_result is None:
message = "FreshRSS pipeline job result is not ready."
if effective["status"] == "failed":
message = "FreshRSS pipeline job did not complete successfully, so no terminal job result is available."
return {
"job_id": state.run_id,
"status": effective["status"],
"message": message,
"linked_run_id": state.input.get("run_id") if isinstance(state.input, dict) else None,
"error_summary": state.error.model_dump(mode="json") if state.error else None,
"status_source": effective["status_source"],
"status_note": effective["status_note"],
}
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
return {
"job_id": state.run_id,
"status": effective["status"],
"run_id": synthesized_result.get("run_id"),
"output_dir": synthesized_result.get("output_dir"),
"raw_output": synthesized_result.get("raw_output"),
"delivery_output": synthesized_result.get("delivery_output"),
"digest_brief_output": synthesized_result.get("digest_brief_output"),
"report_output": synthesized_result.get("report_output"),
"pulled_count": synthesized_result.get("pulled_count"),
"delivered_count": synthesized_result.get("delivered_count"),
"marked_read_count": synthesized_result.get("marked_read_count"),
"status_counts": synthesized_result.get("status_counts"),
"keyword_index": synthesized_result.get("keyword_index"),
"artifact": artifact,
"result": synthesized_result,
"status_source": effective["status_source"],
"status_note": effective["status_note"],
"result_source": synthesized_result.get("result_source", "job_result"),
"linked_run_status": effective["linked_run_status"],
}
def _resolve_effective_job_view(
*,
job_id: str,
state: Any,
result: dict[str, Any] | None = None,
) -> dict[str, Any]:
resolved_result = result or _load_result(job_id)
linked_run_record = _load_linked_run_record(state)
raw_status = state.status
raw_current_stage = state.current_stage
effective_status = raw_status
effective_current_stage = raw_current_stage
effective_updated_at = state.updated_at.isoformat()
effective_finished_at = state.finished_at.isoformat() if state.finished_at else None
status_source = "job_run_state"
state_quality = "trusted"
state_conflict = False
state_conflict_reason = None
status_note = None
if resolved_result is not None:
completed_at = resolved_result.get("completed_at")
effective_status = "success"
effective_current_stage = None
effective_updated_at = completed_at or effective_updated_at
effective_finished_at = completed_at or effective_finished_at
if raw_status != "success" or raw_current_stage is not None:
status_source = "job_result_reconciliation"
state_quality = "reconciled"
state_conflict = True
state_conflict_reason = "result.json already exists, but the job run-state did not converge to success."
status_note = "job result exists; the job can be treated as completed."
elif linked_run_record is not None:
if _linked_run_result_available(linked_run_record):
effective_status = "success"
effective_current_stage = None
effective_updated_at = linked_run_record.get("updated_at") or effective_updated_at
effective_finished_at = linked_run_record.get("finished_at") or effective_finished_at
status_source = "linked_run_reconciliation"
state_quality = "reconciled"
state_conflict = raw_status != "success" or raw_current_stage is not None
state_conflict_reason = (
"The linked run already has terminal artifacts, but the outer job run-state did not converge."
)
status_note = "linked run artifacts are complete; the job can be treated as completed."
resolved_result = _build_result_from_linked_run(job_id=job_id, state=state, linked_run_record=linked_run_record)
elif linked_run_record["status"] == "failed" and raw_status == "running":
effective_status = "failed"
effective_current_stage = None
effective_updated_at = linked_run_record.get("updated_at") or effective_updated_at
effective_finished_at = linked_run_record.get("finished_at") or effective_finished_at
status_source = "linked_run_reconciliation"
state_quality = "reconciled"
state_conflict = True
state_conflict_reason = "The linked run is already failed, but the outer job still reports running."
status_note = "linked run failed; the outer job state appears stale."
elif raw_status == "running":
stale_reason = _build_stale_job_reason(state)
if stale_reason is not None:
effective_status = "failed"
effective_current_stage = raw_current_stage
effective_finished_at = effective_updated_at
status_source = "stale_job_state_timeout"
state_quality = "reconciled"
state_conflict = True
state_conflict_reason = stale_reason
status_note = stale_reason
return {
"status": effective_status,
"current_stage": effective_current_stage,
"updated_at": effective_updated_at,
"finished_at": effective_finished_at,
"status_source": status_source,
"state_quality": state_quality,
"state_conflict": state_conflict,
"state_conflict_reason": state_conflict_reason,
"status_note": status_note,
"linked_run_status": linked_run_record["status"] if linked_run_record is not None else None,
"linked_run_state_conflict": linked_run_record.get("state_conflict") if linked_run_record is not None else None,
"result": resolved_result,
}
def _build_effective_job_progress(*, state: Any, effective_status: str) -> dict[str, int]:
total_stage_count = len(state.stages)
if effective_status == "success":
return {
"completed_stage_count": total_stage_count,
"running_stage_count": 0,
"failed_stage_count": 0,
"pending_stage_count": 0,
"total_stage_count": total_stage_count,
}
if effective_status == "failed" and state.status == "running":
completed_stage_count = sum(1 for s in state.stages if s.status == "success")
pending_stage_count = max(total_stage_count - completed_stage_count - 1, 0)
return {
"completed_stage_count": completed_stage_count,
"running_stage_count": 0,
"failed_stage_count": 1,
"pending_stage_count": pending_stage_count,
"total_stage_count": total_stage_count,
}
return {
"completed_stage_count": sum(1 for s in state.stages if s.status == "success"),
"running_stage_count": sum(1 for s in state.stages if s.status == "running"),
"failed_stage_count": sum(1 for s in state.stages if s.status == "failed"),
"pending_stage_count": sum(1 for s in state.stages if s.status == "pending"),
"total_stage_count": total_stage_count,
}
def _load_linked_run_record(state: Any) -> dict[str, Any] | None:
if not isinstance(state.input, dict):
return None
linked_run_id = state.input.get("run_id")
if not isinstance(linked_run_id, str) or not linked_run_id.strip():
return None
try:
return _resolve_run_record(linked_run_id)
except FileNotFoundError:
return None
def _linked_run_result_available(record: dict[str, Any]) -> bool:
artifact_presence = record.get("artifact_presence")
if not isinstance(artifact_presence, dict):
return False
return bool(artifact_presence.get("run_report")) and bool(artifact_presence.get("delivery_payload"))
def _build_result_from_linked_run(
*,
job_id: str,
state: Any,
linked_run_record: dict[str, Any],
) -> dict[str, Any]:
report = linked_run_record.get("report")
if not isinstance(report, dict):
raise RuntimeError("Cannot synthesize job result because linked run-report.json is missing.")
result = {
"job_id": job_id,
"run_id": linked_run_record["run_id"],
"output_dir": _normalize_repo_path_value(str(linked_run_record["run_dir"])),
"raw_output": _normalize_repo_path_value(report.get("raw_output")),
"delivery_output": _normalize_repo_path_value(report.get("delivery_output")),
"digest_brief_output": _normalize_repo_path_value(report.get("digest_brief_output")),
"report_output": _normalize_repo_path(linked_run_record["run_dir"] / "run-report.json"),
"keyword_index": _normalize_keyword_index(report.get("keyword_index")),
"pulled_count": report.get("pulled_count"),
"delivered_count": report.get("delivered_count"),
"marked_read_count": report.get("marked_read_count"),
"status_counts": report.get("status_counts"),
"debug_artifacts": report.get("debug_artifacts"),
"completed_at": report.get("completed_at") or linked_run_record.get("finished_at"),
"result_source": "linked_run_report",
}
if bool(state.input.get("include_item_reports")):
result["items"] = report.get("items", [])
return result
def _build_stale_job_reason(state: Any) -> str | None:
updated_at = state.updated_at
stale_after_seconds = _estimate_job_stale_seconds(state)
age_seconds = (datetime.now(tz=updated_at.tzinfo) - updated_at).total_seconds()
if age_seconds < stale_after_seconds:
return None
current_stage = state.current_stage or "unknown_stage"
age_minutes = int(age_seconds // 60)
stale_after_minutes = int(stale_after_seconds // 60)
return (
f"job run-state has remained in running state at {current_stage} for about {age_minutes} minutes "
f"without result.json; it exceeded the stale threshold of {stale_after_minutes} minutes."
)
def _estimate_job_stale_seconds(state: Any) -> int:
input_payload = state.input if isinstance(state.input, dict) else {}
limit = _safe_int(input_payload.get("limit"), default=5)
timeout_seconds = _safe_float(input_payload.get("timeout_seconds"), default=60.0)
max_retries = _safe_int(input_payload.get("max_retries"), default=2)
estimated = int((limit * max(timeout_seconds, 1.0) * max(max_retries, 1)) + 20 * 60)
return max(MIN_JOB_STALE_SECONDS, min(MAX_JOB_STALE_SECONDS, estimated))
def _safe_int(value: Any, *, default: int) -> int:
try:
return int(value)
except (TypeError, ValueError):
return default
def _safe_float(value: Any, *, default: float) -> float:
try:
return float(value)
except (TypeError, ValueError):
return default
+221 -2
View File
@@ -45,6 +45,8 @@ DISCOVERED_ARTIFACTS = [
"relative_path": Path("candidates/digest-brief.json"),
},
]
MIN_RUNNING_STALE_SECONDS = 30 * 60
MAX_RUNNING_STALE_SECONDS = 6 * 60 * 60
def get_run_status(*, run_id: str) -> dict[str, Any]:
@@ -186,7 +188,7 @@ def _build_run_record(run_dir: Path) -> dict[str, Any]:
run_store = RunStore.load(path=state_path, repo_root=REPO_ROOT)
state = run_store.state
run_id = state.run_id
return {
record = {
"run_id": run_id,
"workflow": state.workflow,
"run_type": state.run_type,
@@ -204,12 +206,13 @@ def _build_run_record(run_dir: Path) -> dict[str, Any]:
"state": state,
"report": _load_json(report_path) if report_path.exists() else None,
}
return _reconcile_run_record(record)
report = _load_json(report_path) if report_path.exists() else None
run_id = str(report.get("run_id")) if isinstance(report, dict) and report.get("run_id") else run_dir.name
inferred_record = _infer_run_record_from_directory(run_dir=run_dir, report=report, run_id=run_id)
inferred_record["aliases"] = {run_id, run_dir.name}
return inferred_record
return _reconcile_run_record(inferred_record)
def _infer_run_record_from_directory(*, run_dir: Path, report: dict[str, Any] | None, run_id: str) -> dict[str, Any]:
@@ -298,6 +301,13 @@ def _build_status_response(record: dict[str, Any]) -> dict[str, Any]:
"artifacts": _collect_artifacts(record),
"recovery": record["recovery"],
"state_source": record["state_source"],
"status_source": record["status_source"],
"state_quality": record["state_quality"],
"state_conflict": record["state_conflict"],
"state_conflict_reason": record["state_conflict_reason"],
"raw_status": record["raw_status"],
"raw_current_stage": record["raw_current_stage"],
"artifact_presence": record["artifact_presence"],
}
@@ -310,6 +320,10 @@ def _build_run_lookup_response(record: dict[str, Any]) -> dict[str, Any]:
"status": record["status"],
"output_dir": _normalize_repo_path(record["run_dir"]),
"state_source": record["state_source"],
"status_source": record["status_source"],
"state_quality": record["state_quality"],
"state_conflict": record["state_conflict"],
"state_conflict_reason": record["state_conflict_reason"],
}
@@ -330,9 +344,191 @@ def _build_list_response(record: dict[str, Any]) -> dict[str, Any]:
"recovery": record["recovery"],
"artifact_count": len(_collect_artifacts(record)),
"state_source": record["state_source"],
"status_source": record["status_source"],
"state_quality": record["state_quality"],
"state_conflict": record["state_conflict"],
"state_conflict_reason": record["state_conflict_reason"],
"raw_status": record["raw_status"],
"raw_current_stage": record["raw_current_stage"],
}
def _reconcile_run_record(record: dict[str, Any]) -> dict[str, Any]:
raw_status = record["status"]
raw_current_stage = record["current_stage"]
raw_stages = record["stages"]
raw_recovery = record["recovery"]
artifact_presence = _build_artifact_presence(record["run_dir"], report=record.get("report"))
reconciled = dict(record)
reconciled["raw_status"] = raw_status
reconciled["raw_current_stage"] = raw_current_stage
reconciled["status_source"] = record["state_source"]
reconciled["state_quality"] = "trusted"
reconciled["state_conflict"] = False
reconciled["state_conflict_reason"] = None
reconciled["artifact_presence"] = artifact_presence
report = record.get("report")
if not isinstance(report, dict):
stale_reason = _build_stale_running_reason(record)
if stale_reason is not None:
reconciled["status"] = "failed"
reconciled["finished_at"] = record["updated_at"]
reconciled["stages"] = _build_stale_failed_stages(raw_stages, raw_current_stage)
reconciled["error"] = {
"type": "StaleRunState",
"message": stale_reason,
"stage": raw_current_stage,
"details": {
"raw_status": raw_status,
"raw_current_stage": raw_current_stage,
},
}
reconciled["status_source"] = "stale_run_state_timeout"
reconciled["state_quality"] = "reconciled"
reconciled["state_conflict"] = True
reconciled["state_conflict_reason"] = stale_reason
return reconciled
effective_status = _infer_status_from_report(report)
report_completed_at = _maybe_iso(report.get("completed_at"))
conflict = (
raw_status != effective_status
or raw_current_stage is not None
or any(stage["status"] in {"running", "failed"} for stage in raw_stages)
or bool(raw_recovery.get("resumable"))
)
if not conflict:
reconciled["artifact_presence"] = artifact_presence
return reconciled
reconciled["status"] = effective_status
reconciled["current_stage"] = None
reconciled["updated_at"] = report_completed_at or record["updated_at"]
reconciled["finished_at"] = report_completed_at or record["finished_at"]
reconciled["stages"] = _build_terminal_success_stages(raw_stages)
reconciled["error"] = None
reconciled["recovery"] = {
"resumable": False,
"resume_from_stage": None,
"last_success_stage": DEFAULT_STAGES[-1],
}
reconciled["status_source"] = "run_report_reconciliation"
reconciled["state_quality"] = "reconciled"
reconciled["state_conflict"] = True
reconciled["state_conflict_reason"] = (
"run-state.json did not converge, but run-report.json already proves the workflow reached a terminal state."
)
return reconciled
def _build_artifact_presence(run_dir: Path, *, report: dict[str, Any] | None) -> dict[str, bool]:
return {
"run_report": isinstance(report, dict) or (run_dir / "run-report.json").exists(),
"delivery_payload": (run_dir / "candidates" / "openclaw-delivery-payload.json").exists(),
"digest_brief": (run_dir / "candidates" / "digest-brief.json").exists(),
}
def _build_terminal_success_stages(stages: list[dict[str, Any]]) -> list[dict[str, Any]]:
stage_by_name = {stage["name"]: stage for stage in stages}
reconciled_stages: list[dict[str, Any]] = []
for stage_name in DEFAULT_STAGES:
existing = stage_by_name.get(stage_name)
if existing is None:
reconciled_stages.append(_stage_dict(name=stage_name, status="success"))
continue
reconciled_stages.append(
{
"name": stage_name,
"status": "success",
"started_at": existing.get("started_at"),
"finished_at": existing.get("finished_at"),
"outputs": existing.get("outputs", {}),
"error": None,
}
)
return reconciled_stages
def _build_stale_failed_stages(stages: list[dict[str, Any]], current_stage: str | None) -> list[dict[str, Any]]:
stage_by_name = {stage["name"]: stage for stage in stages}
reconciled_stages: list[dict[str, Any]] = []
for stage_name in DEFAULT_STAGES:
existing = stage_by_name.get(stage_name)
if existing is None:
reconciled_stages.append(_stage_dict(name=stage_name, status="pending"))
continue
status = existing.get("status")
if status == "running" or (current_stage is not None and stage_name == current_stage):
status = "failed"
reconciled_stages.append(
{
"name": stage_name,
"status": status,
"started_at": existing.get("started_at"),
"finished_at": existing.get("finished_at") or existing.get("started_at"),
"outputs": existing.get("outputs", {}),
"error": existing.get("error"),
}
)
return reconciled_stages
def _build_stale_running_reason(record: dict[str, Any]) -> str | None:
if record["status"] != "running":
return None
updated_at = _parse_iso_datetime(record.get("updated_at"))
if updated_at is None:
return None
stale_after_seconds = _estimate_running_stale_seconds(record)
age_seconds = (datetime.now(tz=updated_at.tzinfo) - updated_at).total_seconds()
if age_seconds < stale_after_seconds:
return None
current_stage = record.get("current_stage") or "unknown_stage"
age_minutes = int(age_seconds // 60)
stale_after_minutes = int(stale_after_seconds // 60)
return (
f"run-state.json has remained in running state at {current_stage} for about {age_minutes} minutes "
f"without terminal artifacts; it exceeded the stale threshold of {stale_after_minutes} minutes."
)
def _estimate_running_stale_seconds(record: dict[str, Any]) -> int:
state = record.get("state")
input_payload = state.input if isinstance(state, RunState) and isinstance(state.input, dict) else {}
limit = _safe_int(input_payload.get("limit"), default=5)
timeout_seconds = _safe_float(input_payload.get("timeout_seconds"), default=60.0)
max_retries = _safe_int(input_payload.get("max_retries"), default=2)
expected_items = limit
current_stage = record.get("current_stage")
if current_stage == "generate_summaries":
expected_items = _stage_output_from_record(record, "generate_summaries", "expected_items") or limit
elif current_stage == "extract_articles":
expected_items = _stage_output_from_record(record, "extract_articles", "expected_items") or limit
estimated = int((expected_items * max(timeout_seconds, 1.0) * max(max_retries, 1)) + 15 * 60)
return max(MIN_RUNNING_STALE_SECONDS, min(MAX_RUNNING_STALE_SECONDS, estimated))
def _stage_output_from_record(record: dict[str, Any], stage_name: str, key: str) -> Any:
for stage in record.get("stages", []):
if stage.get("name") != stage_name:
continue
outputs = stage.get("outputs")
if isinstance(outputs, dict):
return outputs.get(key)
return None
def _collect_artifacts(record: dict[str, Any]) -> list[dict[str, Any]]:
artifacts: list[dict[str, Any]] = []
seen_names: set[str] = set()
@@ -587,6 +783,29 @@ def _maybe_iso(value: Any) -> str | None:
return None
def _parse_iso_datetime(value: Any) -> datetime | None:
if not isinstance(value, str) or not value.strip():
return None
try:
return datetime.fromisoformat(value)
except ValueError:
return None
def _safe_int(value: Any, *, default: int) -> int:
try:
return int(value)
except (TypeError, ValueError):
return default
def _safe_float(value: Any, *, default: float) -> float:
try:
return float(value)
except (TypeError, ValueError):
return default
def _record_sort_key(record: dict[str, Any]) -> tuple[str, str]:
return (record.get("updated_at") or "", record["run_dir"].name)
+614
View File
@@ -0,0 +1,614 @@
from __future__ import annotations
import json
import os
import subprocess
import sys
from datetime import datetime
from pathlib import Path
from typing import Any
from uuid import uuid4
from dotenv import dotenv_values
from .query_service import _resolve_run_record
from .resume_service import (
SUPPORTED_RESUME_STAGES,
UNSUPPORTED_RESUME_STAGES,
_build_resume_plan,
_resume_freshrss_run,
_validate_resume_artifacts,
inspect_resume_plan,
)
from .run_store import RunStore
from .state_models import RunState
REPO_ROOT = Path(__file__).resolve().parents[3]
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
RESUME_JOBS_ROOT = OUTPUT_ROOT / "resume_jobs"
WORKFLOW_NAME = "freshrss_resume_job"
RUN_TYPE = "resume_job"
RUN_STATE_FILENAME = "run-state.json"
INPUT_FILENAME = "input.json"
RESULT_FILENAME = "result.json"
JOB_REPORT_FILENAME = "job-report.json"
DEFAULT_STAGES = [
"prepare_job",
"validate_resume_plan",
"resume_run",
"write_result",
]
MIN_JOB_STALE_SECONDS = 30 * 60
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
def _build_subprocess_env() -> dict[str, str]:
"""Build an env dict for subprocess, merging parent env with .env values.
Ensures the subprocess sees all needed variables (LLM_API_KEY, LLM_MODEL,
LLM_API_URL, FRESHRSS_*, etc.) directly in os.environ, avoiding dotenv
timing issues in the child process.
"""
env = os.environ.copy()
if DEFAULT_DOTENV_PATH.exists():
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items():
if isinstance(key, str) and isinstance(value, str) and value:
env.setdefault(key, value)
return env
def _now() -> datetime:
return datetime.now().astimezone()
def _new_job_id() -> str:
ts = _now().strftime("%Y%m%d-%H%M%S")
return f"freshrss-resume-job-{ts}-{uuid4().hex[:8]}"
def _job_dir(job_id: str) -> Path:
return RESUME_JOBS_ROOT / job_id
def _run_state_path(job_id: str) -> Path:
return _job_dir(job_id) / RUN_STATE_FILENAME
def _input_path(job_id: str) -> Path:
return _job_dir(job_id) / INPUT_FILENAME
def _result_path(job_id: str) -> Path:
return _job_dir(job_id) / RESULT_FILENAME
def _job_report_path(job_id: str) -> Path:
return _job_dir(job_id) / JOB_REPORT_FILENAME
def _normalize_repo_path(path: Path) -> str:
try:
return str(path.resolve().relative_to(REPO_ROOT.resolve()))
except ValueError:
return str(path)
def _normalize_repo_path_value(path_value: str | None) -> str | None:
if not path_value:
return None
return _normalize_repo_path(Path(path_value))
def _load_run_store(job_id: str) -> RunStore:
return RunStore.load(path=_run_state_path(job_id), repo_root=REPO_ROOT)
def _load_result(job_id: str) -> dict[str, Any] | None:
path = _result_path(job_id)
if not path.exists():
return None
return json.loads(path.read_text(encoding="utf-8-sig"))
def _load_job_report(job_id: str) -> dict[str, Any] | None:
path = _job_report_path(job_id)
if not path.exists():
return None
return json.loads(path.read_text(encoding="utf-8-sig"))
def _write_job_report(*, job_id: str, payload: dict[str, Any]) -> Path:
report_file = _job_report_path(job_id)
report_file.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
return report_file
def _build_result_payload(
*,
job_id: str,
run_id: str,
run_dir: Path,
resume_plan: dict[str, Any],
resume_result: dict[str, Any],
) -> dict[str, Any]:
return {
"job_id": job_id,
"run_id": run_id,
"resume_from_stage": resume_plan["effective_resume_from_stage"],
"requested_resume_from_stage": resume_plan["requested_resume_from_stage"],
"resume_decision_source": resume_plan["decision_source"],
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
"artifact_snapshot": resume_plan["artifact_snapshot"],
"status": resume_result["status"],
"output_dir": _normalize_repo_path(run_dir),
"delivery_output": _normalize_repo_path_value(resume_result.get("delivery_output")),
"report_output": _normalize_repo_path_value(resume_result.get("report_output")),
"digest_brief_output": _normalize_repo_path_value(resume_result.get("digest_brief_output")),
"pulled_count": resume_result.get("pulled_count"),
"delivered_count": resume_result.get("delivered_count"),
"marked_read_count": resume_result.get("marked_read_count"),
"status_counts": resume_result.get("status_counts"),
"keyword_index": resume_result.get("keyword_index"),
"completed_at": _now().isoformat(),
"result_source": "resume_job_result",
}
def _build_result_from_linked_run(*, job_id: str, linked_run_record: dict[str, Any]) -> dict[str, Any] | None:
report = linked_run_record.get("report")
if not isinstance(report, dict):
return None
recovery = linked_run_record.get("recovery")
requested_resume_from_stage = None
if isinstance(recovery, dict):
requested_resume_from_stage = recovery.get("resume_from_stage")
return {
"job_id": job_id,
"run_id": linked_run_record["run_id"],
"resume_from_stage": requested_resume_from_stage,
"requested_resume_from_stage": requested_resume_from_stage,
"resume_decision_source": "linked_run_report",
"artifact_resume_from_stage": None,
"artifact_snapshot": None,
"status": linked_run_record["status"],
"output_dir": _normalize_repo_path(linked_run_record["run_dir"]),
"delivery_output": _normalize_repo_path_value(report.get("delivery_output")),
"report_output": _normalize_repo_path(linked_run_record["run_dir"] / "run-report.json"),
"digest_brief_output": _normalize_repo_path_value(report.get("digest_brief_output")),
"pulled_count": report.get("pulled_count"),
"delivered_count": report.get("delivered_count"),
"marked_read_count": report.get("marked_read_count"),
"status_counts": report.get("status_counts"),
"keyword_index": report.get("keyword_index"),
"completed_at": report.get("completed_at"),
"result_source": "linked_run_report",
}
def _linked_run_result_available(linked_run_record: dict[str, Any]) -> bool:
return linked_run_record["status"] in {"success", "partial"} and isinstance(linked_run_record.get("report"), dict)
def _load_linked_run_record(state: Any) -> dict[str, Any] | None:
if not isinstance(state.input, dict):
return None
run_id = state.input.get("run_id")
if not isinstance(run_id, str) or not run_id.strip():
return None
try:
return _resolve_run_record(run_id)
except FileNotFoundError:
return None
def _build_stale_job_reason(state: Any) -> str | None:
if state.status != "running" or state.current_stage is None or state.finished_at is not None:
return None
now = _now()
age_seconds = max(0.0, (now - state.updated_at).total_seconds())
if age_seconds < MIN_JOB_STALE_SECONDS:
return None
if age_seconds >= MAX_JOB_STALE_SECONDS:
return (
f"Resume job has remained in stage '{state.current_stage}' for more than {int(MAX_JOB_STALE_SECONDS)} seconds "
"without producing a terminal result; treating the job state as stale."
)
return None
def start_resume_job(*, run_id: str) -> dict[str, Any]:
resume_view = inspect_resume_plan(run_id=run_id)
job_id = _new_job_id()
job_dir = _job_dir(job_id)
job_dir.mkdir(parents=True, exist_ok=True)
started_at = _now()
input_payload = {
"run_id": run_id,
"requested_resume_from_stage": resume_view.get("requested_resume_from_stage"),
"resume_from_stage": resume_view.get("resume_from_stage"),
"resume_decision_source": resume_view.get("resume_decision_source"),
"recommended_action": resume_view.get("recommended_action"),
"artifact_resume_from_stage": resume_view.get("artifact_resume_from_stage"),
"artifact_snapshot": resume_view.get("artifact_snapshot"),
"launcher_pid": os.getpid(),
}
store = RunStore.create(
path=_run_state_path(job_id),
run_id=job_id,
workflow=WORKFLOW_NAME,
run_type=RUN_TYPE,
started_at=started_at,
input_payload=input_payload,
repo_root=REPO_ROOT,
)
for stage_name in DEFAULT_STAGES:
store._get_or_create_stage(stage_name)
store.save()
store.start_stage("prepare_job")
input_file = _input_path(job_id)
input_file.write_text(json.dumps(input_payload, ensure_ascii=False, indent=2), encoding="utf-8")
store.register_artifact(name="job_input", path=input_file, kind="json", stage="prepare_job")
if not resume_view.get("can_resume") or resume_view.get("recommended_action") != "resume":
report_file = _write_job_report(
job_id=job_id,
payload={
"job_id": job_id,
"status": "failed",
"error_type": "ResumePreflightRejected",
"error_message": resume_view.get("message"),
"failed_stage": "validate_resume_plan",
"linked_run_id": run_id,
"resume_from_stage": resume_view.get("resume_from_stage"),
"requested_resume_from_stage": resume_view.get("requested_resume_from_stage"),
"recommended_action": resume_view.get("recommended_action"),
},
)
store.register_artifact(name="job_report", path=report_file, kind="json", stage="prepare_job")
store.finish_stage("prepare_job", outputs={"linked_run_id": run_id})
store.fail_stage("validate_resume_plan", error=ValueError(str(resume_view.get("message") or "Resume preflight rejected.")))
return {
"job_id": job_id,
"workflow": WORKFLOW_NAME,
"run_type": RUN_TYPE,
"status": "failed",
"output_dir": _normalize_repo_path(job_dir),
"linked_run_id": run_id,
"resume_from_stage": resume_view.get("resume_from_stage"),
"recommended_action": resume_view.get("recommended_action"),
"message": resume_view.get("message"),
}
runner_script = REPO_ROOT / "scripts" / "run_resume_job.py"
cmd = [sys.executable, str(runner_script), "--job-id", job_id]
try:
proc = subprocess.Popen(
cmd,
cwd=str(REPO_ROOT),
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
start_new_session=True,
env=_build_subprocess_env(),
)
except Exception as exc:
report_file = _write_job_report(
job_id=job_id,
payload={
"job_id": job_id,
"status": "failed",
"error_type": type(exc).__name__,
"error_message": str(exc),
"failed_stage": "prepare_job",
"linked_run_id": run_id,
"resume_from_stage": resume_view.get("resume_from_stage"),
},
)
store.register_artifact(name="job_report", path=report_file, kind="json", stage="prepare_job")
store.fail_stage("prepare_job", error=exc)
raise
store.finish_stage(
"prepare_job",
outputs={
"runner_pid": proc.pid,
"runner_command": cmd,
"linked_run_id": run_id,
"resume_from_stage": resume_view.get("resume_from_stage"),
},
)
return {
"job_id": job_id,
"workflow": WORKFLOW_NAME,
"run_type": RUN_TYPE,
"status": "running",
"output_dir": _normalize_repo_path(job_dir),
"linked_run_id": run_id,
"resume_from_stage": resume_view.get("resume_from_stage"),
"message": "Resume job started successfully. Use get_resume_job_status to poll progress.",
}
def run_resume_job(*, job_id: str) -> dict[str, Any]:
store = _load_run_store(job_id)
input_payload = json.loads(_input_path(job_id).read_text(encoding="utf-8-sig"))
current_stage = "resume_run"
run_id = str(input_payload["run_id"])
try:
store.start_stage("validate_resume_plan")
record = _resolve_run_record(run_id)
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
raise RuntimeError("This run cannot be resumed because run-state.json is missing or could not be loaded.")
run_store = RunStore.load(path=record["run_dir"] / "run-state.json", repo_root=REPO_ROOT)
resume_plan = _build_resume_plan(record=record, state=run_store.state)
if resume_plan["decision"] != "resume":
raise RuntimeError(str(resume_plan["message"]))
resume_from_stage = resume_plan["effective_resume_from_stage"]
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
raise RuntimeError(f"This run cannot be resumed from {resume_from_stage} in the current implementation.")
if resume_from_stage not in SUPPORTED_RESUME_STAGES:
raise RuntimeError(f"This run cannot be resumed because stage '{resume_from_stage}' is not supported.")
missing_artifacts = _validate_resume_artifacts(
record=record,
state=run_store.state,
resume_from_stage=resume_from_stage,
)
if missing_artifacts:
raise RuntimeError(
f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing: {missing_artifacts}"
)
store.finish_stage(
"validate_resume_plan",
outputs={
"linked_run_id": run_id,
"resume_from_stage": resume_from_stage,
"requested_resume_from_stage": resume_plan["requested_resume_from_stage"],
"resume_decision_source": resume_plan["decision_source"],
},
)
current_stage = "resume_run"
store.start_stage("resume_run")
resume_result = _resume_freshrss_run(record=record, run_store=run_store, resume_from_stage=resume_from_stage)
store.finish_stage(
"resume_run",
outputs={
"linked_run_id": run_id,
"resume_from_stage": resume_from_stage,
"linked_run_status": resume_result["status"],
"delivery_output": _normalize_repo_path_value(resume_result.get("delivery_output")),
"report_output": _normalize_repo_path_value(resume_result.get("report_output")),
},
)
store.start_stage("write_result")
result = _build_result_payload(
job_id=job_id,
run_id=run_id,
run_dir=record["run_dir"],
resume_plan=resume_plan,
resume_result=resume_result,
)
result_file = _result_path(job_id)
result_file.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
store.register_artifact(name="job_result", path=result_file, kind="json", stage="write_result")
report_file = _write_job_report(
job_id=job_id,
payload={
"job_id": job_id,
"status": "success",
"run_id": run_id,
"resume_from_stage": result["resume_from_stage"],
"delivery_output": result["delivery_output"],
"report_output": result["report_output"],
"digest_brief_output": result["digest_brief_output"],
},
)
store.register_artifact(name="job_report", path=report_file, kind="json", stage="write_result")
store.finish_stage(
"write_result",
outputs={
"result_path": _normalize_repo_path(result_file),
"linked_run_id": run_id,
},
)
store.finish_run(status="success")
return result
except Exception as exc:
current_stage = store.state.current_stage or current_stage
report_file = _write_job_report(
job_id=job_id,
payload={
"job_id": job_id,
"status": "failed",
"error_type": type(exc).__name__,
"error_message": str(exc),
"failed_stage": current_stage,
"linked_run_id": run_id,
},
)
try:
store.register_artifact(name="job_report", path=report_file, kind="json", stage=current_stage)
except Exception:
pass
store.fail_stage(current_stage, error=exc)
raise
def _resolve_effective_job_view(
*,
job_id: str,
state: Any,
result: dict[str, Any] | None = None,
) -> dict[str, Any]:
resolved_result = result or _load_result(job_id)
linked_run_record = _load_linked_run_record(state)
raw_status = state.status
raw_current_stage = state.current_stage
effective_status = raw_status
effective_current_stage = raw_current_stage
effective_updated_at = state.updated_at.isoformat()
effective_finished_at = state.finished_at.isoformat() if state.finished_at else None
status_source = "job_run_state"
state_quality = "trusted"
state_conflict = False
state_conflict_reason = None
status_note = None
if resolved_result is not None:
completed_at = resolved_result.get("completed_at")
effective_status = "success"
effective_current_stage = None
effective_updated_at = completed_at or effective_updated_at
effective_finished_at = completed_at or effective_finished_at
if raw_status != "success" or raw_current_stage is not None:
status_source = "job_result_reconciliation"
state_quality = "reconciled"
state_conflict = True
state_conflict_reason = "result.json already exists, but the resume job run-state did not converge to success."
status_note = "resume job result exists; the job can be treated as completed."
elif linked_run_record is not None:
if _linked_run_result_available(linked_run_record) and raw_status == "running":
effective_status = "success"
effective_current_stage = None
effective_updated_at = linked_run_record.get("updated_at") or effective_updated_at
effective_finished_at = linked_run_record.get("finished_at") or effective_finished_at
status_source = "linked_run_reconciliation"
state_quality = "reconciled"
state_conflict = raw_status != "success" or raw_current_stage is not None
state_conflict_reason = "The linked run already has terminal artifacts, but the resume job run-state did not converge."
status_note = "linked run artifacts are complete; the resume job can be treated as completed."
resolved_result = _build_result_from_linked_run(job_id=job_id, linked_run_record=linked_run_record)
elif raw_status == "running":
stale_reason = _build_stale_job_reason(state)
if stale_reason is not None:
effective_status = "failed"
effective_current_stage = raw_current_stage
effective_finished_at = effective_updated_at
status_source = "stale_job_state_timeout"
state_quality = "reconciled"
state_conflict = True
state_conflict_reason = stale_reason
status_note = stale_reason
return {
"status": effective_status,
"current_stage": effective_current_stage,
"updated_at": effective_updated_at,
"finished_at": effective_finished_at,
"status_source": status_source,
"state_quality": state_quality,
"state_conflict": state_conflict,
"state_conflict_reason": state_conflict_reason,
"status_note": status_note,
"result": resolved_result,
"linked_run_status": linked_run_record["status"] if linked_run_record is not None else None,
"linked_run_state_conflict": linked_run_record.get("state_conflict") if linked_run_record is not None else None,
}
def _build_effective_job_progress(*, state: Any, effective_status: str) -> dict[str, int]:
completed_stage_count = sum(1 for stage in state.stages if stage.status == "success")
running_stage_count = sum(1 for stage in state.stages if stage.status == "running")
failed_stage_count = sum(1 for stage in state.stages if stage.status == "failed")
pending_stage_count = sum(1 for stage in state.stages if stage.status == "pending")
if effective_status == "success" and running_stage_count > 0:
pending_stage_count += running_stage_count
running_stage_count = 0
return {
"completed_stage_count": completed_stage_count,
"running_stage_count": running_stage_count,
"failed_stage_count": failed_stage_count,
"pending_stage_count": pending_stage_count,
"total_stage_count": len(state.stages),
}
def get_resume_job_status(*, job_id: str) -> dict[str, Any]:
store = _load_run_store(job_id)
state = store.state
effective = _resolve_effective_job_view(job_id=job_id, state=state)
linked_run_id = state.input.get("run_id") if isinstance(state.input, dict) else None
return {
"job_id": state.run_id,
"workflow": state.workflow,
"run_type": state.run_type,
"status": effective["status"],
"current_stage": effective["current_stage"],
"started_at": state.started_at.isoformat(),
"updated_at": effective["updated_at"],
"finished_at": effective["finished_at"],
"output_dir": _normalize_repo_path(_job_dir(job_id)),
"linked_run_id": linked_run_id,
"progress": _build_effective_job_progress(state=state, effective_status=effective["status"]),
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
"error_summary": state.error.model_dump(mode="json") if state.error else None,
"status_source": effective["status_source"],
"state_quality": effective["state_quality"],
"state_conflict": effective["state_conflict"],
"state_conflict_reason": effective["state_conflict_reason"],
"status_note": effective["status_note"],
"raw_status": state.status,
"raw_current_stage": state.current_stage,
"linked_run_status": effective["linked_run_status"],
"linked_run_state_conflict": effective["linked_run_state_conflict"],
}
def get_resume_job_result(*, job_id: str) -> dict[str, Any]:
store = _load_run_store(job_id)
state = store.state
result = _load_result(job_id)
effective = _resolve_effective_job_view(job_id=job_id, state=state, result=result)
synthesized_result = result or effective["result"]
if effective["status"] != "success" or synthesized_result is None:
message = "Resume job result is not ready."
if effective["status"] == "failed":
message = "Resume job did not complete successfully, so no terminal job result is available."
return {
"job_id": state.run_id,
"status": effective["status"],
"message": message,
"linked_run_id": state.input.get("run_id") if isinstance(state.input, dict) else None,
"error_summary": state.error.model_dump(mode="json") if state.error else None,
"status_source": effective["status_source"],
"status_note": effective["status_note"],
}
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
return {
"job_id": state.run_id,
"status": effective["status"],
"run_id": synthesized_result.get("run_id"),
"resume_from_stage": synthesized_result.get("resume_from_stage"),
"requested_resume_from_stage": synthesized_result.get("requested_resume_from_stage"),
"resume_decision_source": synthesized_result.get("resume_decision_source"),
"artifact_resume_from_stage": synthesized_result.get("artifact_resume_from_stage"),
"artifact_snapshot": synthesized_result.get("artifact_snapshot"),
"output_dir": synthesized_result.get("output_dir"),
"delivery_output": synthesized_result.get("delivery_output"),
"report_output": synthesized_result.get("report_output"),
"digest_brief_output": synthesized_result.get("digest_brief_output"),
"pulled_count": synthesized_result.get("pulled_count"),
"delivered_count": synthesized_result.get("delivered_count"),
"marked_read_count": synthesized_result.get("marked_read_count"),
"status_counts": synthesized_result.get("status_counts"),
"keyword_index": synthesized_result.get("keyword_index"),
"artifact": artifact,
"result": synthesized_result,
"status_source": effective["status_source"],
"status_note": effective["status_note"],
"result_source": synthesized_result.get("result_source", "job_result"),
"linked_run_status": effective["linked_run_status"],
}
+426 -18
View File
@@ -25,6 +25,7 @@ from summary_mcp.models.openclaw_delivery import (
)
from summary_mcp.models.summary_io import ExtractionOutput
from summary_mcp.workflows.freshrss_pipeline import (
CANDIDATE_BATCH_ARTIFACT,
DEFAULT_PROMPT_PATH,
DEFAULT_RULES_PATH,
DEFAULT_TERM_ALIASES_PATH,
@@ -37,13 +38,18 @@ from summary_mcp.workflows.freshrss_pipeline import (
FILTER_STAGE,
REPORT_STAGE,
REPO_ROOT,
SUMMARY_BATCH_ARTIFACT,
SUMMARY_STAGE,
WORKFLOW_NAME,
_build_item_context,
_candidate_batch_output,
_final_run_status,
_load_json,
_load_required_env,
_persist_candidate_batch_artifact,
_persist_summary_batch_artifact,
_save_json,
_summary_batch_output,
)
from .query_service import _normalize_repo_path, _resolve_repo_path, _resolve_run_record
@@ -62,6 +68,12 @@ UNSUPPORTED_RESUME_STAGES = {
}
UTC = timezone.utc
DEFAULT_STREAM_ID = "user/-/state/com.google/reading-list"
RESUME_STAGE_ORDER = {
SUMMARY_STAGE: 1,
FILTER_STAGE: 2,
DELIVERY_STAGE: 3,
REPORT_STAGE: 4,
}
def resume_run(*, run_id: str) -> dict[str, Any]:
@@ -71,37 +83,68 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
"workflow": record["workflow"],
"status": record["status"],
"output_dir": _normalize_repo_path(record["run_dir"]),
"state_source": record["state_source"],
"status_source": record.get("status_source"),
}
if record["status"] in {"success", "partial"} and isinstance(record.get("report"), dict):
return {
**base_response,
"resumed": False,
"resume_from_stage": None,
"requested_resume_from_stage": None,
"resume_decision_source": "artifacts",
"recommended_action": "read_terminal_result",
"message": "This run already has a terminal run-report.json; prefer reading get_run_status/get_run_report instead of resuming.",
"missing_artifacts": [],
"state_conflict": record.get("state_conflict", False),
"state_conflict_reason": record.get("state_conflict_reason"),
}
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
return {
**base_response,
"resumed": False,
"resume_from_stage": None,
"requested_resume_from_stage": None,
"resume_decision_source": "unavailable",
"recommended_action": "start_new_run",
"message": "This run cannot be resumed because run-state.json is missing or could not be loaded.",
"missing_artifacts": [],
}
run_store = RunStore.load(path=record["run_dir"] / "run-state.json", repo_root=REPO_ROOT)
state = run_store.state
resume_from_stage = _resolve_resume_from_stage(state)
resume_plan = _build_resume_plan(record=record, state=state)
requested_resume_from_stage = resume_plan["requested_resume_from_stage"]
resume_from_stage = resume_plan["effective_resume_from_stage"]
if state.workflow != WORKFLOW_NAME:
return {
**base_response,
"resumed": False,
"resume_from_stage": resume_from_stage,
"requested_resume_from_stage": requested_resume_from_stage,
"resume_decision_source": resume_plan["decision_source"],
"recommended_action": "start_new_run",
"message": f"This run cannot be resumed because workflow '{state.workflow}' is not supported by the minimal resume_run implementation.",
"missing_artifacts": [],
}
if resume_from_stage is None:
if resume_plan["decision"] != "resume":
return {
**base_response,
"resumed": False,
"resume_from_stage": None,
"message": "This run does not expose a recoverable stage in run-state.json.",
"missing_artifacts": [],
"resume_from_stage": resume_from_stage,
"requested_resume_from_stage": requested_resume_from_stage,
"resume_decision_source": resume_plan["decision_source"],
"recommended_action": resume_plan["recommended_action"],
"message": resume_plan["message"],
"missing_artifacts": resume_plan["missing_artifacts"],
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
"artifact_snapshot": resume_plan["artifact_snapshot"],
"state_conflict": record.get("state_conflict", False),
"state_conflict_reason": record.get("state_conflict_reason"),
}
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
@@ -109,6 +152,9 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
**base_response,
"resumed": False,
"resume_from_stage": resume_from_stage,
"requested_resume_from_stage": requested_resume_from_stage,
"resume_decision_source": resume_plan["decision_source"],
"recommended_action": "start_new_run",
"message": f"This run cannot be resumed from {resume_from_stage} in the current minimal implementation.",
"missing_artifacts": [],
}
@@ -118,6 +164,9 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
**base_response,
"resumed": False,
"resume_from_stage": resume_from_stage,
"requested_resume_from_stage": requested_resume_from_stage,
"resume_decision_source": resume_plan["decision_source"],
"recommended_action": "start_new_run",
"message": f"This run cannot be resumed because stage '{resume_from_stage}' is not supported.",
"missing_artifacts": [],
}
@@ -128,8 +177,13 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
**base_response,
"resumed": False,
"resume_from_stage": resume_from_stage,
"requested_resume_from_stage": requested_resume_from_stage,
"resume_decision_source": resume_plan["decision_source"],
"recommended_action": "start_new_run",
"message": f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing.",
"missing_artifacts": missing_artifacts,
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
"artifact_snapshot": resume_plan["artifact_snapshot"],
}
try:
@@ -141,10 +195,15 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
**base_response,
"resumed": True,
"resume_from_stage": resume_from_stage,
"requested_resume_from_stage": requested_resume_from_stage,
"resume_decision_source": resume_plan["decision_source"],
"recommended_action": "inspect_error",
"status": run_store.state.status,
"message": f"Run resumed from {resume_from_stage} but failed again at {failed_stage}: {error}",
"missing_artifacts": [],
"error_summary": run_store.state.error.model_dump(mode="json") if run_store.state.error else None,
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
"artifact_snapshot": resume_plan["artifact_snapshot"],
}
return {
@@ -152,6 +211,9 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
"workflow": run_store.state.workflow,
"resumed": True,
"resume_from_stage": resume_from_stage,
"requested_resume_from_stage": requested_resume_from_stage,
"resume_decision_source": resume_plan["decision_source"],
"recommended_action": "none",
"status": result["status"],
"output_dir": _normalize_repo_path(record["run_dir"]),
"message": f"Run resumed from {resume_from_stage} and completed with status {result['status']}.",
@@ -166,9 +228,92 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
"marked_read_count": result.get("marked_read_count"),
"status_counts": result.get("status_counts"),
"missing_artifacts": [],
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
"artifact_snapshot": resume_plan["artifact_snapshot"],
}
def inspect_resume_plan(*, run_id: str) -> dict[str, Any]:
record = _resolve_run_record(run_id)
base_response = {
"run_id": record["run_id"],
"workflow": record["workflow"],
"status": record["status"],
"output_dir": _normalize_repo_path(record["run_dir"]),
"state_source": record["state_source"],
"status_source": record.get("status_source"),
"state_conflict": record.get("state_conflict", False),
"state_conflict_reason": record.get("state_conflict_reason"),
}
if record["status"] in {"success", "partial"} and isinstance(record.get("report"), dict):
return {
**base_response,
"can_resume": False,
"resume_from_stage": None,
"requested_resume_from_stage": None,
"resume_decision_source": "artifacts",
"recommended_action": "read_terminal_result",
"message": "This run already has a terminal run-report.json; prefer reading get_run_status/get_run_report instead of resuming.",
"missing_artifacts": [],
}
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
return {
**base_response,
"can_resume": False,
"resume_from_stage": None,
"requested_resume_from_stage": None,
"resume_decision_source": "unavailable",
"recommended_action": "start_new_run",
"message": "This run cannot be resumed because run-state.json is missing or could not be loaded.",
"missing_artifacts": [],
}
state = record["state"]
resume_plan = _build_resume_plan(record=record, state=state)
response = {
**base_response,
"can_resume": False,
"resume_from_stage": resume_plan["effective_resume_from_stage"],
"requested_resume_from_stage": resume_plan["requested_resume_from_stage"],
"resume_decision_source": resume_plan["decision_source"],
"recommended_action": resume_plan["recommended_action"],
"message": resume_plan["message"],
"missing_artifacts": resume_plan["missing_artifacts"],
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
"artifact_snapshot": resume_plan["artifact_snapshot"],
}
if state.workflow != WORKFLOW_NAME:
response["message"] = (
f"This run cannot be resumed because workflow '{state.workflow}' is not supported by the minimal resume_run implementation."
)
return response
resume_from_stage = resume_plan["effective_resume_from_stage"]
if resume_plan["decision"] != "resume":
return response
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
response["message"] = f"This run cannot be resumed from {resume_from_stage} in the current minimal implementation."
response["recommended_action"] = "start_new_run"
return response
if resume_from_stage not in SUPPORTED_RESUME_STAGES:
response["message"] = f"This run cannot be resumed because stage '{resume_from_stage}' is not supported."
response["recommended_action"] = "start_new_run"
return response
missing_artifacts = _validate_resume_artifacts(record=record, state=state, resume_from_stage=resume_from_stage)
if missing_artifacts:
response["message"] = f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing."
response["missing_artifacts"] = missing_artifacts
response["recommended_action"] = "start_new_run"
return response
response["can_resume"] = True
return response
def _resume_freshrss_run(*, record: dict[str, Any], run_store: RunStore, resume_from_stage: str) -> dict[str, Any]:
state = run_store.state
run_dir = record["run_dir"]
@@ -304,6 +449,75 @@ def _build_resume_config(*, state: RunState, run_dir: Path) -> dict[str, Any]:
}
def _find_summary_batch_path(*, run_dir: Path, state: RunState) -> Path | None:
return _find_artifact_path(
run_dir=run_dir,
state=state,
artifact_name=SUMMARY_BATCH_ARTIFACT,
relative_path=Path("summary/summary-batch.json"),
)
def _find_candidate_batch_path(*, run_dir: Path, state: RunState) -> Path | None:
return _find_artifact_path(
run_dir=run_dir,
state=state,
artifact_name=CANDIDATE_BATCH_ARTIFACT,
relative_path=Path("candidates/candidate-batch.json"),
)
def _load_summary_batch_lookup(path: Path | None) -> tuple[bool, dict[str, dict[str, Any]]]:
if path is None or not path.exists():
return False, {}
try:
payload = _load_json(path)
except Exception:
return False, {}
items = payload.get("items")
if not isinstance(items, list):
return False, {}
summaries_by_item_key: dict[str, dict[str, Any]] = {}
for entry in items:
if not isinstance(entry, dict):
return False, {}
item_key = entry.get("item_key")
summary = entry.get("summary")
if not isinstance(item_key, str) or not item_key.strip() or not isinstance(summary, dict):
return False, {}
summaries_by_item_key[item_key] = summary
return True, summaries_by_item_key
def _load_candidate_batch_lookup(path: Path | None) -> tuple[bool, dict[str, OpenClawCandidateInput]]:
if path is None or not path.exists():
return False, {}
try:
payload = _load_json(path)
except Exception:
return False, {}
items = payload.get("items")
if not isinstance(items, list):
return False, {}
candidates_by_item_key: dict[str, OpenClawCandidateInput] = {}
try:
for entry in items:
if not isinstance(entry, dict):
return False, {}
item_key = entry.get("item_key")
candidate_payload = entry.get("candidate")
if not isinstance(item_key, str) or not item_key.strip() or not isinstance(candidate_payload, dict):
return False, {}
candidates_by_item_key[item_key] = OpenClawCandidateInput.model_validate(candidate_payload)
except Exception:
return False, {}
return True, candidates_by_item_key
def _load_item_contexts(
*,
items: list[Any],
@@ -313,6 +527,8 @@ def _load_item_contexts(
include_candidates: bool = False,
delivery_candidates_by_id: dict[str, OpenClawCandidateInput] | None = None,
) -> list[dict[str, Any]]:
summary_batch_valid, summary_batch_by_item_key = _load_summary_batch_lookup(_summary_batch_output(run_dir))
candidate_batch_valid, candidate_batch_by_item_key = _load_candidate_batch_lookup(_candidate_batch_output(run_dir))
contexts: list[dict[str, Any]] = []
for index, item in enumerate(items, start=1):
context = _build_item_context(
@@ -334,16 +550,24 @@ def _load_item_contexts(
else:
context["item_report"]["status"] = "extracted"
if include_summaries and context["summary_output"] is not None and context["summary_output"].exists():
summary_payload = _load_json(context["summary_output"])
context["summary_payload"] = summary_payload
context["item_report"]["status"] = "summarized"
elif include_summaries and context.get("extraction") is not None and context["extraction"].success:
context["item_report"]["status"] = "summary_failed"
if include_summaries:
summary_payload: dict[str, Any] | None = None
if context["summary_output"] is not None and context["summary_output"].exists():
summary_payload = _load_json(context["summary_output"])
elif summary_batch_valid:
summary_payload = summary_batch_by_item_key.get(context["item_key"])
if summary_payload is not None:
context["summary_payload"] = summary_payload
context["item_report"]["status"] = "summarized"
elif context.get("extraction") is not None and context["extraction"].success:
context["item_report"]["status"] = "summary_failed"
candidate: OpenClawCandidateInput | None = None
if include_candidates and context["openclaw_path"] is not None and context["openclaw_path"].exists():
candidate = OpenClawCandidateInput.model_validate(_load_json(context["openclaw_path"]))
elif include_candidates and candidate_batch_valid:
candidate = candidate_batch_by_item_key.get(context["item_key"])
elif delivery_candidates_by_id is not None and context.get("extraction") is not None and context["extraction"].success:
candidate_id = candidate_id_for(item, context["extraction"].article)
candidate = delivery_candidates_by_id.get(candidate_id)
@@ -406,6 +630,11 @@ def _run_summary_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any
},
)
summary_batch_output = _persist_summary_batch_artifact(
run_store=run_store,
run_dir=config["run_dir"],
item_contexts=item_contexts,
)
if config["debug_artifacts"] and (config["run_dir"] / "summary").exists():
run_store.register_artifact(name="summary_dir", path=config["run_dir"] / "summary", kind="directory", stage=SUMMARY_STAGE)
run_store.finish_stage(
@@ -415,6 +644,7 @@ def _run_summary_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any
"completed_items": summary_success_count + summary_failed_count,
"success_count": summary_success_count,
"failed_count": summary_failed_count,
"summary_batch_output": str(summary_batch_output),
},
)
@@ -500,6 +730,11 @@ def _run_filter_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any]
},
)
candidate_batch_output = _persist_candidate_batch_artifact(
run_store=run_store,
run_dir=config["run_dir"],
item_contexts=item_contexts,
)
if config["debug_artifacts"] and (config["run_dir"] / "candidates").exists():
run_store.register_artifact(name="candidate_dir", path=config["run_dir"] / "candidates", kind="directory", stage=FILTER_STAGE)
run_store.finish_stage(
@@ -511,6 +746,7 @@ def _run_filter_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any]
"keep_count": keep_count,
"review_count": review_count,
"drop_count": drop_count,
"candidate_batch_output": str(candidate_batch_output),
},
)
@@ -682,6 +918,165 @@ def _build_keyword_index_result(
return result
def _build_resume_plan(*, record: dict[str, Any], state: RunState) -> dict[str, Any]:
requested_resume_from_stage = _resolve_resume_from_stage(state)
artifact_snapshot = _collect_resume_artifact_snapshot(record=record, state=state)
artifact_resume_from_stage = _resolve_resume_stage_from_artifacts(artifact_snapshot)
if artifact_snapshot["has_run_report"]:
return {
"decision": "reject_terminal",
"decision_source": "artifacts",
"requested_resume_from_stage": requested_resume_from_stage,
"effective_resume_from_stage": None,
"artifact_resume_from_stage": artifact_resume_from_stage,
"recommended_action": "read_terminal_result",
"message": "This run already has a terminal run-report.json; prefer reading get_run_status/get_run_report instead of resuming.",
"missing_artifacts": [],
"artifact_snapshot": artifact_snapshot,
}
if artifact_resume_from_stage is None:
return {
"decision": "reject_unrecoverable",
"decision_source": "artifacts",
"requested_resume_from_stage": requested_resume_from_stage,
"effective_resume_from_stage": None,
"artifact_resume_from_stage": None,
"recommended_action": "start_new_run",
"message": "This run does not expose a safe artifact-backed resume point; start a new run instead.",
"missing_artifacts": artifact_snapshot["missing_for_next_resume"],
"artifact_snapshot": artifact_snapshot,
}
decision_source = "artifacts"
message = f"Resume will continue from {artifact_resume_from_stage} based on available artifacts."
if requested_resume_from_stage == artifact_resume_from_stage:
decision_source = "state_and_artifacts"
message = f"Resume stage {artifact_resume_from_stage} was confirmed by both run-state.json and artifacts."
elif requested_resume_from_stage is not None:
requested_rank = RESUME_STAGE_ORDER.get(requested_resume_from_stage, -1)
artifact_rank = RESUME_STAGE_ORDER.get(artifact_resume_from_stage, -1)
if artifact_rank > requested_rank:
message = (
f"Resume stage was advanced from {requested_resume_from_stage} to {artifact_resume_from_stage} "
f"because artifacts prove the run already progressed further."
)
else:
message = (
f"Resume stage was moved back from {requested_resume_from_stage} to {artifact_resume_from_stage} "
f"because later-stage artifacts are not stable enough for a safe resume."
)
return {
"decision": "resume",
"decision_source": decision_source,
"requested_resume_from_stage": requested_resume_from_stage,
"effective_resume_from_stage": artifact_resume_from_stage,
"artifact_resume_from_stage": artifact_resume_from_stage,
"recommended_action": "resume",
"message": message,
"missing_artifacts": [],
"artifact_snapshot": artifact_snapshot,
}
def _collect_resume_artifact_snapshot(*, record: dict[str, Any], state: RunState) -> dict[str, Any]:
run_dir = record["run_dir"]
raw_output = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="raw_output", relative_path=Path("raw/freshrss.raw.json"))
extracted_dir = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="extracted_dir", relative_path=Path("extracted"))
summary_batch_output = _find_summary_batch_path(run_dir=run_dir, state=state)
candidate_batch_output = _find_candidate_batch_path(run_dir=run_dir, state=state)
delivery_output = _find_artifact_path(
run_dir=run_dir,
state=state,
artifact_name="delivery_payload",
relative_path=Path("candidates/openclaw-delivery-payload.json"),
)
digest_brief_output = _find_artifact_path(
run_dir=run_dir,
state=state,
artifact_name="digest_brief",
relative_path=Path("candidates/digest-brief.json"),
)
report_output = run_dir / "run-report.json"
items: list[Any] = []
raw_output_valid = False
if raw_output is not None:
try:
items = _load_items(raw_output)
raw_output_valid = True
except Exception:
items = []
raw_output_valid = False
summary_batch_valid, _ = _load_summary_batch_lookup(summary_batch_output)
candidate_batch_valid, _ = _load_candidate_batch_lookup(candidate_batch_output)
item_contexts = _load_item_contexts(
items=items,
run_dir=run_dir,
debug_artifacts=bool(state.input.get("debug_artifacts", False)),
include_summaries=True,
include_candidates=True,
)
extraction_success_count = sum(
1 for context in item_contexts if context.get("extraction") is not None and context["extraction"].success
)
summary_count = sum(1 for context in item_contexts if context.get("summary_payload") is not None)
candidate_count = sum(1 for context in item_contexts if context.get("candidate") is not None)
extracted_complete = raw_output_valid and bool(items) and all(context["extracted_path"].exists() for context in item_contexts)
stable_summary_outputs = summary_batch_valid or (extraction_success_count > 0 and extraction_success_count == summary_count)
stable_candidate_outputs = candidate_batch_valid or (summary_count > 0 and summary_count == candidate_count)
missing_for_next_resume: list[str] = []
if raw_output is None or not raw_output_valid:
missing_for_next_resume.append("raw/freshrss.raw.json")
if extracted_dir is None or not extracted_complete:
missing_for_next_resume.append("extracted/")
if extraction_success_count > 0 and not stable_summary_outputs and not stable_candidate_outputs:
missing_for_next_resume.append("summary/summary-batch.json")
if extraction_success_count > 0 and stable_summary_outputs and not stable_candidate_outputs:
missing_for_next_resume.append("candidates/candidate-batch.json")
return {
"has_raw_output": raw_output is not None,
"raw_output_valid": raw_output_valid,
"has_extracted_dir": extracted_dir is not None,
"has_summary_batch": summary_batch_output is not None,
"summary_batch_valid": summary_batch_valid,
"has_candidate_batch": candidate_batch_output is not None,
"candidate_batch_valid": candidate_batch_valid,
"has_delivery_payload": delivery_output is not None,
"has_digest_brief": digest_brief_output is not None,
"has_run_report": report_output.exists(),
"item_count": len(items),
"extraction_success_count": extraction_success_count,
"summary_count": summary_count,
"candidate_count": candidate_count,
"extracted_complete": extracted_complete,
"stable_summary_outputs": stable_summary_outputs,
"stable_candidate_outputs": stable_candidate_outputs,
"missing_for_next_resume": missing_for_next_resume,
}
def _resolve_resume_stage_from_artifacts(snapshot: dict[str, Any]) -> str | None:
if snapshot["has_run_report"]:
return None
if snapshot["has_delivery_payload"]:
return REPORT_STAGE
if snapshot["stable_candidate_outputs"]:
return DELIVERY_STAGE
if snapshot["has_raw_output"] and snapshot["raw_output_valid"] and snapshot["has_extracted_dir"] and snapshot["extracted_complete"]:
if snapshot["extraction_success_count"] <= 0:
return SUMMARY_STAGE
if snapshot["stable_summary_outputs"]:
return FILTER_STAGE
return SUMMARY_STAGE
return None
def _load_filter_context(config: dict[str, Any]) -> FilterContext:
if config["context"] is not None:
return FilterContext.model_validate(config["context"])
@@ -695,6 +1090,8 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
missing_artifacts: list[str] = []
raw_output = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="raw_output", relative_path=Path("raw/freshrss.raw.json"))
extracted_dir = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="extracted_dir", relative_path=Path("extracted"))
summary_batch_output = _find_summary_batch_path(run_dir=run_dir, state=state)
candidate_batch_output = _find_candidate_batch_path(run_dir=run_dir, state=state)
if resume_from_stage in {SUMMARY_STAGE, FILTER_STAGE, DELIVERY_STAGE} and raw_output is None:
missing_artifacts.append("raw/freshrss.raw.json")
if resume_from_stage in {SUMMARY_STAGE, FILTER_STAGE} and extracted_dir is None:
@@ -711,6 +1108,8 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
include_summaries=resume_from_stage in {FILTER_STAGE, DELIVERY_STAGE},
include_candidates=resume_from_stage == DELIVERY_STAGE,
)
summary_batch_valid, _ = _load_summary_batch_lookup(summary_batch_output)
candidate_batch_valid, _ = _load_candidate_batch_lookup(candidate_batch_output)
if resume_from_stage == SUMMARY_STAGE:
for context in item_contexts:
@@ -718,17 +1117,26 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
missing_artifacts.append(_normalize_repo_path(context["extracted_path"]))
if resume_from_stage == FILTER_STAGE:
for context in item_contexts:
if context.get("extraction") is not None and context["extraction"].success:
if context["summary_output"] is None or not context["summary_output"].exists():
missing_artifacts.append(
_normalize_repo_path(context["summary_output"] or (run_dir / "summary" / context["item_key"] / "result.loop.json"))
)
expected_summary_count = int(_stage_output(state, SUMMARY_STAGE, "success_count") or 0)
actual_summary_count = sum(1 for context in item_contexts if context.get("summary_payload") is not None)
if summary_batch_valid:
if expected_summary_count != actual_summary_count:
missing_artifacts.append(_normalize_repo_path(summary_batch_output or _summary_batch_output(run_dir)))
else:
for context in item_contexts:
if context.get("extraction") is not None and context["extraction"].success:
if context["summary_output"] is None or not context["summary_output"].exists():
missing_artifacts.append(
_normalize_repo_path(context["summary_output"] or (run_dir / "summary" / context["item_key"] / "result.loop.json"))
)
if resume_from_stage == DELIVERY_STAGE:
expected_candidate_count = int(_stage_output(state, FILTER_STAGE, "candidate_count") or 0)
actual_candidate_count = sum(1 for context in item_contexts if context.get("candidate") is not None)
if expected_candidate_count != actual_candidate_count:
if candidate_batch_valid:
if expected_candidate_count != actual_candidate_count:
missing_artifacts.append(_normalize_repo_path(candidate_batch_output or _candidate_batch_output(run_dir)))
elif expected_candidate_count != actual_candidate_count:
for context in item_contexts:
if context.get("summary_payload") is not None and (context["openclaw_path"] is None or not context["openclaw_path"].exists()):
missing_artifacts.append(
+135 -4
View File
@@ -1,7 +1,7 @@
from __future__ import annotations
# MCP 服务入口:将内容提取、过滤、FreshRSS 全链路管道暴露为 MCP 工具。
# 生产主入口是 run_freshrss_openclaw_pipeline,其余工具供单步调试使用。
# 主日报正式启动入口是 start_freshrss_pipeline_job;同步入口仅保留给 debug / fallback。
from datetime import date
from pathlib import Path
@@ -16,10 +16,23 @@ from summary_mcp.models.item import Item
from summary_mcp.models.llm_result import LlmSummaryResult
from summary_mcp.models.summary_io import ExtractionInput
from summary_mcp.runtime import get_delivery_payload as load_delivery_payload
from summary_mcp.runtime import get_freshrss_pipeline_job_result as load_freshrss_pipeline_job_result
from summary_mcp.runtime import get_freshrss_pipeline_job_status as load_freshrss_pipeline_job_status
from summary_mcp.runtime import get_resume_job_result as load_resume_job_result
from summary_mcp.runtime import get_resume_job_status as load_resume_job_status
from summary_mcp.runtime import get_run_status as load_run_status
from summary_mcp.runtime import get_run_report as load_run_report
from summary_mcp.runtime import inspect_resume_plan as load_resume_plan
from summary_mcp.runtime import list_run_artifacts as load_run_artifacts
from summary_mcp.runtime import list_runs as load_runs
from summary_mcp.runtime import start_resume_job as launch_resume_job
from summary_mcp.runtime import start_freshrss_pipeline_job as launch_freshrss_pipeline_job
from summary_mcp.runtime.article_summary_jobs import (
get_article_summary_job_result as load_article_summary_job_result,
get_article_summary_job_status as load_article_summary_job_status,
start_article_summary_job as launch_article_summary_job,
_validate_article_summary_extracted_path,
)
from summary_mcp.runtime.resume_service import resume_run as resume_existing_run
from summary_mcp.workflows import run_freshrss_pipeline
from summary_mcp.workflows.article_summary import ArticleSummaryConfig, summarize_selected_articles
@@ -125,6 +138,64 @@ def run_freshrss_openclaw_pipeline(
return result
@mcp.tool()
def start_freshrss_pipeline_job(
limit: int = 5,
mark_read: bool = False,
include_read: bool = False,
debug_artifacts: bool = False,
continuation: str | None = None,
timeout_seconds: float = 60.0,
max_retries: int = 2,
stream_id: str = "user/-/state/com.google/reading-list",
api_base_url: str | None = None,
username: str | None = None,
api_password: str | None = None,
llm_api_key: str | None = None,
llm_model: str | None = None,
llm_api_url: str | None = None,
context: dict | None = None,
run_id: str | None = None,
date_value: str | None = None,
output_dir: str | None = None,
include_item_reports: bool = False,
) -> dict:
"""Start an asynchronous FreshRSS pipeline job and return a job_id immediately."""
return launch_freshrss_pipeline_job(
limit=limit,
mark_read=mark_read,
include_read=include_read,
debug_artifacts=debug_artifacts,
continuation=continuation,
timeout_seconds=timeout_seconds,
max_retries=max_retries,
stream_id=stream_id,
api_base_url=api_base_url,
username=username,
api_password=api_password,
llm_api_key=llm_api_key,
llm_model=llm_model,
llm_api_url=llm_api_url,
context=context,
run_id=run_id,
date_value=date_value,
output_dir=Path(output_dir) if output_dir else None,
include_item_reports=include_item_reports,
)
@mcp.tool()
def get_freshrss_pipeline_job_status(job_id: str) -> dict:
"""Get the current status of an asynchronous FreshRSS pipeline job."""
return load_freshrss_pipeline_job_status(job_id=job_id)
@mcp.tool()
def get_freshrss_pipeline_job_result(job_id: str) -> dict:
"""Get the final result of an asynchronous FreshRSS pipeline job."""
return load_freshrss_pipeline_job_result(job_id=job_id)
@mcp.tool()
def get_run_status(run_id: str) -> dict:
"""Get the current status of a workflow run by run_id."""
@@ -165,6 +236,67 @@ def resume_run(run_id: str) -> dict:
return resume_existing_run(run_id=run_id)
@mcp.tool()
def inspect_resume_plan(run_id: str) -> dict:
"""Inspect the effective resume plan for a FreshRSS workflow run without executing it."""
return load_resume_plan(run_id=run_id)
@mcp.tool()
def start_resume_job(run_id: str) -> dict:
"""Start an asynchronous resume job for a resumable FreshRSS workflow run."""
return launch_resume_job(run_id=run_id)
@mcp.tool()
def get_resume_job_status(job_id: str) -> dict:
"""Get the current status of an asynchronous resume job."""
return load_resume_job_status(job_id=job_id)
@mcp.tool()
def get_resume_job_result(job_id: str) -> dict:
"""Get the final result of an asynchronous resume job."""
return load_resume_job_result(job_id=job_id)
@mcp.tool()
def start_article_summary_job(
*,
extracted_path: str,
selected_ids: list[str],
output_dir: str | None = None,
max_retries: int = 2,
timeout_seconds: float = 120.0,
llm_api_key: str | None = None,
llm_model: str | None = None,
llm_api_url: str | None = None,
) -> dict:
"""Start an asynchronous article-summary job and return a job_id immediately."""
return launch_article_summary_job(
extracted_path=Path(extracted_path),
selected_ids=selected_ids,
output_dir=Path(output_dir) if output_dir else None,
max_retries=max_retries,
timeout_seconds=timeout_seconds,
llm_api_key=llm_api_key,
llm_model=llm_model,
llm_api_url=llm_api_url,
)
@mcp.tool()
def get_article_summary_job_status(job_id: str) -> dict:
"""Get the current status of an asynchronous article-summary job."""
return load_article_summary_job_status(job_id=job_id)
@mcp.tool()
def get_article_summary_job_result(job_id: str) -> dict:
"""Get the final result of an asynchronous article-summary job."""
return load_article_summary_job_result(job_id=job_id)
@mcp.tool()
def generate_article_summaries(
*,
@@ -172,7 +304,7 @@ def generate_article_summaries(
selected_ids: list[str],
output_dir: str | None = None,
max_retries: int = 2,
timeout_seconds: float = 60.0,
timeout_seconds: float = 120.0,
llm_api_key: str | None = None,
llm_model: str | None = None,
llm_api_url: str | None = None,
@@ -197,8 +329,7 @@ def generate_article_summaries(
"""
extracted_path_obj = Path(extracted_path)
if not extracted_path_obj.exists():
raise FileNotFoundError(f"extracted_path does not exist: {extracted_path}")
_validate_article_summary_extracted_path(extracted_path_obj)
if output_dir is None:
default_dir = extracted_path_obj.parent / "single_summaries"
+92 -17
View File
@@ -4,6 +4,9 @@ from dataclasses import dataclass
from pathlib import Path
from typing import Iterable, Mapping, Sequence
import httpx
from dotenv import dotenv_values
from summary_mcp.core.summary_loop import run_loop_payload
from summary_mcp.validators.article_summary import validate_article_summary_payload
@@ -12,13 +15,14 @@ REPO_ROOT = Path(__file__).resolve().parents[3]
OUTPUT_ROOT = REPO_ROOT / "outputs"
FRESHRSS_OUTPUT_ROOT = OUTPUT_ROOT / "freshrss"
DEFAULT_PROMPT_PATH = OUTPUT_ROOT / "prompts" / "article-summary-prompt.txt"
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
@dataclass
class ArticleSummaryConfig:
prompt_path: Path = DEFAULT_PROMPT_PATH
max_retries: int = 2
timeout_seconds: float = 60.0
timeout_seconds: float = 120.0
llm_api_key: str | None = None
llm_model: str | None = None
llm_api_url: str | None = None
@@ -35,6 +39,7 @@ def _resolve_article_llm_settings(
Resolution order for each field:
- explicit function argument
- ARTICLE_SUMMARY_* environment variable
- ARTICLE_SUMMARY_* in repo .env
- main LLM_* / OPENAI_* environment variables (handled by summary_loop.resolve_llm_settings)
This helper intentionally does not validate presence; the summary loop
@@ -43,9 +48,17 @@ def _resolve_article_llm_settings(
import os
resolved_api_key = api_key or os.environ.get("ARTICLE_SUMMARY_LLM_API_KEY")
resolved_model = model or os.environ.get("ARTICLE_SUMMARY_LLM_MODEL")
resolved_api_url = api_url or os.environ.get("ARTICLE_SUMMARY_LLM_API_URL")
dotenv_map = {}
if DEFAULT_DOTENV_PATH.exists():
dotenv_map = {
key: value
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items()
if isinstance(key, str) and isinstance(value, str) and value
}
resolved_api_key = api_key or os.environ.get("ARTICLE_SUMMARY_LLM_API_KEY") or dotenv_map.get("ARTICLE_SUMMARY_LLM_API_KEY")
resolved_model = model or os.environ.get("ARTICLE_SUMMARY_LLM_MODEL") or dotenv_map.get("ARTICLE_SUMMARY_LLM_MODEL")
resolved_api_url = api_url or os.environ.get("ARTICLE_SUMMARY_LLM_API_URL") or dotenv_map.get("ARTICLE_SUMMARY_LLM_API_URL")
return resolved_api_key, resolved_model, resolved_api_url
@@ -97,6 +110,32 @@ def _iter_selected_items(
}
return
# Fallback: pipeline delivery payload with top-level "candidates" array.
# Each entry has item_id (the raw FreshRSS item_id) and candidate_id (with cand: prefix).
# Normalize selected_ids by stripping "cand:" prefix so they match item_id.
candidates = extracted_payload.get("candidates")
if isinstance(candidates, list):
norm_selected = {
sid.removeprefix("cand:") if sid.startswith("cand:") else sid
for sid in selected_ids
}
for entry in candidates:
if not isinstance(entry, Mapping):
continue
# Support both item_id field (new) and candidate_id (legacy fallback)
raw_item_id = entry.get("item_id") or ""
if not raw_item_id and entry.get("candidate_id"):
raw_item_id = entry["candidate_id"].removeprefix("cand:")
item_id = str(raw_item_id) if raw_item_id else None
if not item_id or item_id not in norm_selected:
continue
# Candidates store article fields directly, not nested under "article"
yield item_id, {
"item": {},
"extraction": {"article": entry, "warnings": []},
}
return
# Fallback: legacy payload with top-level "items" array.
items = extracted_payload.get("items")
if isinstance(items, list):
@@ -107,11 +146,10 @@ def _iter_selected_items(
item_id = str(raw_item_id) if raw_item_id is not None else None
if not item_id or item_id not in selected_set:
continue
yield item_id, item
return
# Format 3: single-item extracted file produced by run_freshrss_pipeline debug mode.
# Format 3: single-item extracted file produced by run_freshrss_pipeline.
# Shape: {"success": bool, "article": {"item_id": "...", ...}, "warnings": [...]}
article = extracted_payload.get("article")
if isinstance(article, Mapping):
@@ -165,10 +203,12 @@ def summarize_selected_articles(
output_dir.mkdir(parents=True, exist_ok=True)
written_paths: list[Path] = []
matched_ids: set[str] = set()
from summary_mcp.core.summary_loop import build_summary_input
for item_id, entry in _iter_selected_items(payload, selected_ids):
matched_ids.add(item_id)
# Prefer the real project format where each entry has ``item`` and
# ``extraction.article``; fall back to legacy layout where the
# article fields live directly on the element.
@@ -187,17 +227,45 @@ def summarize_selected_articles(
"warnings": warnings,
}
summary_exit_code, summary_payload, _ = run_loop_payload(
extracted_payload=extracted_payload,
prompt_path=cfg.prompt_path,
max_retries=cfg.max_retries,
timeout_seconds=cfg.timeout_seconds,
api_key=resolved_api_key,
model=resolved_model,
api_url=resolved_api_url,
output_path=None,
validator=validate_article_summary_payload,
summary_exit_code = 1
summary_payload = None
try:
summary_exit_code, summary_payload, _ = run_loop_payload(
extracted_payload=extracted_payload,
prompt_path=cfg.prompt_path,
max_retries=cfg.max_retries,
timeout_seconds=cfg.timeout_seconds,
api_key=resolved_api_key,
model=resolved_model,
api_url=resolved_api_url,
output_path=None,
validator=validate_article_summary_payload,
)
except httpx.TimeoutException:
summary_exit_code = 1
summary_payload = None
should_fallback = (
(summary_exit_code != 0 or summary_payload is None)
and resolved_model is not None
and resolved_model == (cfg.llm_model or model or resolved_model)
and resolved_model.startswith("deepseek")
)
if should_fallback:
summary_exit_code, summary_payload, _ = run_loop_payload(
extracted_payload=extracted_payload,
prompt_path=cfg.prompt_path,
max_retries=cfg.max_retries,
timeout_seconds=cfg.timeout_seconds,
api_key=None,
model=None,
api_url=None,
output_path=None,
validator=validate_article_summary_payload,
)
if summary_exit_code != 0 or summary_payload is None:
continue
@@ -250,12 +318,19 @@ def summarize_selected_articles(
lines.append("、".join(topics))
lines.append("")
# Normalize: collapse spaces/slashes/underscores to single dash, strip punctuation, collapse multi-dashes
normalized = str(title).lower()
for sep in (" ", "/", "_", "——", "―", "‐"):
normalized = normalized.replace(sep, "-")
safe_title = "-".join(
str(title).lower().strip().replace(" ", "-").split()
part for part in normalized.split("-") if part
)[:80]
filename = f"{safe_title or item_id}.md"
output_path = output_dir / filename
output_path.write_text("\n".join(lines), encoding="utf-8")
written_paths.append(output_path)
if not matched_ids:
raise ValueError(f"No extracted entries matched selected_ids: {list(selected_ids)}")
return written_paths
+153 -35
View File
@@ -2,10 +2,13 @@ from __future__ import annotations
import json
import os
from concurrent.futures import ThreadPoolExecutor, as_completed
from datetime import date, datetime, timezone
from pathlib import Path
from typing import Any
from dotenv import dotenv_values
from summary_mcp.core.keyword_index import persist_keyword_indexes
from summary_mcp.core.pipeline import extract_content
from summary_mcp.core.summary_loop import resolve_llm_settings, run_loop_payload
@@ -40,6 +43,7 @@ DEFAULT_TERM_ALIASES_PATH = REPO_ROOT / "configs" / "term_aliases.json"
DEFAULT_TERM_STOPWORDS_PATH = REPO_ROOT / "configs" / "term_stopwords.json"
DEFAULT_TERM_DAILY_DIR = DATA_ROOT / "daily"
DEFAULT_TERM_STATS_PATH = DATA_ROOT / "term_stats.json"
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
WORKFLOW_NAME = "freshrss_daily_digest"
RUN_TYPE = "daily_digest"
FETCH_STAGE = "fetch_feed"
@@ -48,6 +52,10 @@ SUMMARY_STAGE = "generate_summaries"
FILTER_STAGE = "apply_filters"
DELIVERY_STAGE = "build_delivery_payload"
REPORT_STAGE = "write_run_report"
SUMMARY_BATCH_ARTIFACT = "summary_batch"
CANDIDATE_BATCH_ARTIFACT = "candidate_batch"
SUMMARY_BATCH_FILENAME = "summary-batch.json"
CANDIDATE_BATCH_FILENAME = "candidate-batch.json"
def _save_json(path: Path, payload: dict[str, Any] | list[Any]) -> None:
@@ -59,14 +67,27 @@ def _load_json(path: Path) -> dict[str, Any]:
return json.loads(path.read_text(encoding="utf-8-sig"))
def _load_required_env(name: str, value: str | None) -> str:
def _load_repo_dotenv() -> dict[str, str]:
if not DEFAULT_DOTENV_PATH.exists():
return {}
return {
key: value
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items()
if isinstance(key, str) and isinstance(value, str) and value
}
def _load_required_env(name: str, value: str | None, dotenv_map: dict[str, str] | None = None) -> str:
if value:
return value
env_value = os.environ.get(name)
if env_value:
return env_value
dotenv_value = (dotenv_map or {}).get(name)
if dotenv_value:
return dotenv_value
raise RuntimeError(
f"Missing required value '{name}': not passed as argument and not set as environment variable."
f"Missing required value '{name}': not passed as argument, not set as environment variable, and not found in {DEFAULT_DOTENV_PATH}."
)
@@ -121,6 +142,77 @@ def _build_item_context(*, index: int, item: Any, resolved_output_dir: Path, deb
}
def _summary_batch_output(run_dir: Path) -> Path:
return run_dir / "summary" / SUMMARY_BATCH_FILENAME
def _candidate_batch_output(run_dir: Path) -> Path:
return run_dir / "candidates" / CANDIDATE_BATCH_FILENAME
def _build_summary_batch_payload(*, run_id: str, item_contexts: list[dict[str, Any]]) -> dict[str, Any]:
items: list[dict[str, Any]] = []
for item_context in item_contexts:
summary_payload = item_context.get("summary_payload")
if summary_payload is None:
continue
item = item_context["item"]
items.append(
{
"item_key": item_context["item_key"],
"item_id": item.item_id,
"summary": summary_payload,
}
)
return {
"run_id": run_id,
"summary_count": len(items),
"items": items,
}
def _build_candidate_batch_payload(*, run_id: str, item_contexts: list[dict[str, Any]]) -> dict[str, Any]:
items: list[dict[str, Any]] = []
for item_context in item_contexts:
candidate = item_context.get("candidate")
if candidate is None:
continue
item = item_context["item"]
items.append(
{
"item_key": item_context["item_key"],
"item_id": item.item_id,
"candidate_id": candidate.candidate_id,
"candidate": candidate.model_dump(mode="json"),
}
)
return {
"run_id": run_id,
"candidate_count": len(items),
"items": items,
}
def _persist_summary_batch_artifact(*, run_store, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
output_path = _summary_batch_output(run_dir)
_save_json(
output_path,
_build_summary_batch_payload(run_id=run_store.state.run_id, item_contexts=item_contexts),
)
run_store.register_artifact(name=SUMMARY_BATCH_ARTIFACT, path=output_path, kind="json", stage=SUMMARY_STAGE)
return output_path
def _persist_candidate_batch_artifact(*, run_store, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
output_path = _candidate_batch_output(run_dir)
_save_json(
output_path,
_build_candidate_batch_payload(run_id=run_store.state.run_id, item_contexts=item_contexts),
)
run_store.register_artifact(name=CANDIDATE_BATCH_ARTIFACT, path=output_path, kind="json", stage=FILTER_STAGE)
return output_path
def _build_run_report(
*,
resolved_run_id: str,
@@ -243,9 +335,10 @@ def run_freshrss_pipeline(
try:
run_store.start_stage(FETCH_STAGE, outputs={"output_dir": str(resolved_output_dir)})
resolved_api_base_url = _load_required_env("FRESHRSS_API_BASE_URL", api_base_url)
resolved_username = _load_required_env("FRESHRSS_USERNAME", username)
resolved_api_password = _load_required_env("FRESHRSS_API_PASSWORD", api_password)
dotenv_map = _load_repo_dotenv()
resolved_api_base_url = _load_required_env("FRESHRSS_API_BASE_URL", api_base_url, dotenv_map)
resolved_username = _load_required_env("FRESHRSS_USERNAME", username, dotenv_map)
resolved_api_password = _load_required_env("FRESHRSS_API_PASSWORD", api_password, dotenv_map)
resolved_llm_api_key, resolved_llm_model, resolved_llm_api_url = resolve_llm_settings(
api_key=llm_api_key,
model=llm_model,
@@ -369,38 +462,56 @@ def run_freshrss_pipeline(
summary_success_count = 0
summary_failed_count = 0
summary_candidates = [ctx for ctx in item_contexts if ctx["extraction"] is not None and ctx["extraction"].success]
for item_context in summary_candidates:
item_report = item_context["item_report"]
summary_exit_code, summary_payload, summary_report = run_loop_payload(
extracted_payload=item_context["extracted_payload"],
prompt_path=resolved_prompt_path,
output_path=item_context["summary_output"],
max_retries=max_retries,
timeout_seconds=timeout_seconds,
api_key=resolved_llm_api_key,
model=resolved_llm_model,
api_url=resolved_llm_api_url,
)
if summary_exit_code != 0 or summary_payload is None:
item_report["status"] = "summary_failed"
if summary_report is not None:
item_report["summary_errors"] = summary_report.errors
summary_failed_count += 1
else:
item_context["summary_payload"] = summary_payload
item_report["status"] = "summarized"
summary_success_count += 1
# Parallelize LLM summaries — I/O bound calls, independent per article
with ThreadPoolExecutor(max_workers=min(len(summary_candidates) or 1, 4)) as pool:
fut_map = {}
for item_context in summary_candidates:
fut = pool.submit(
run_loop_payload,
extracted_payload=item_context["extracted_payload"],
prompt_path=resolved_prompt_path,
output_path=item_context["summary_output"],
max_retries=max_retries,
timeout_seconds=timeout_seconds,
api_key=resolved_llm_api_key,
model=resolved_llm_model,
api_url=resolved_llm_api_url,
)
fut_map[fut] = item_context
run_store.update_stage(
SUMMARY_STAGE,
outputs={
"expected_items": extracted_success_count,
"completed_items": summary_success_count + summary_failed_count,
"success_count": summary_success_count,
"failed_count": summary_failed_count,
},
)
for fut in as_completed(fut_map):
item_context = fut_map[fut]
item_report = item_context["item_report"]
try:
summary_exit_code, summary_payload, summary_report = fut.result()
except Exception as exc:
summary_exit_code, summary_payload, summary_report = 1, None, None
if summary_exit_code != 0 or summary_payload is None:
item_report["status"] = "summary_failed"
if summary_report is not None:
item_report["summary_errors"] = summary_report.errors
summary_failed_count += 1
else:
item_context["summary_payload"] = summary_payload
item_report["status"] = "summarized"
summary_success_count += 1
run_store.update_stage(
SUMMARY_STAGE,
outputs={
"expected_items": extracted_success_count,
"completed_items": summary_success_count + summary_failed_count,
"success_count": summary_success_count,
"failed_count": summary_failed_count,
},
)
summary_batch_output = _persist_summary_batch_artifact(
run_store=run_store,
run_dir=resolved_output_dir,
item_contexts=item_contexts,
)
if debug_artifacts and (resolved_output_dir / "summary").exists():
run_store.register_artifact(name="summary_dir", path=resolved_output_dir / "summary", kind="directory", stage=SUMMARY_STAGE)
run_store.finish_stage(
@@ -410,6 +521,7 @@ def run_freshrss_pipeline(
"completed_items": summary_success_count + summary_failed_count,
"success_count": summary_success_count,
"failed_count": summary_failed_count,
"summary_batch_output": str(summary_batch_output),
},
)
@@ -493,6 +605,11 @@ def run_freshrss_pipeline(
},
)
candidate_batch_output = _persist_candidate_batch_artifact(
run_store=run_store,
run_dir=resolved_output_dir,
item_contexts=item_contexts,
)
if debug_artifacts and (resolved_output_dir / "candidates").exists():
run_store.register_artifact(name="candidate_dir", path=resolved_output_dir / "candidates", kind="directory", stage=FILTER_STAGE)
run_store.finish_stage(
@@ -504,6 +621,7 @@ def run_freshrss_pipeline(
"keep_count": keep_count,
"review_count": review_count,
"drop_count": drop_count,
"candidate_batch_output": str(candidate_batch_output),
},
)