Compare commits

...
7 Commits
Author SHA1 Message Date
root 8e27cca166 docs: sync reader-digest-flow skill with Hermes version (absolute path, extract_failed handling, rerun include_read) 2026-08-03 09:55:19 +08:00
zhuyongxin 807027976e docs: update digest template and localize references 2026-07-29 10:20:20 +08:00
zhuyongxin 6dd8cef347 refactor: simplify reader digest skill 2026-07-28 19:14:09 +08:00
root 5eb390e3ed docs: 添加 Agent Skill 到项目仓库
- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/)
- README 新增 Agent Skill 章节说明供 Agent 使用的工作流
2026-07-28 18:38:54 +08:00
root 7b791ac947 docs: 重写 README 为项目介绍
从 MCP 集成技术文档改为完整项目 README,包含:
- 项目定位与核心能力
- 架构概览(流程图 + MCP 工具层 + CLI)
- 快速开始与环境变量
- 项目结构一览
- 关键技术决策说明
- 关键词治理流程
2026-07-28 16:26:34 +08:00
root cdbcdcd485 feat: pipeline 并行摘要 + 子进程 env 注入 + 循环导入修复
源码:
- runtime/__init__.py: resume_jobs/resume_service 改为懒加载,打破循环导入
- freshrss_pipeline_jobs.py / resume_jobs.py: 子进程注入 .env 环境变量
- freshrss_pipeline.py: LLM 摘要串行改并行 (ThreadPoolExecutor, max_workers=4)

配置:
- term_aliases: 19→149 条,大幅扩充别名映射
- term_stopwords: 19→132 条,增加过滤规则
- filter_context.personal.json: +7 个兴趣关键词
2026-07-28 16:24:31 +08:00
root 590d050218 keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions
- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
2026-05-14 17:17:49 +08:00
28 changed files with 2971 additions and 456 deletions
+187 -363
View File
@@ -1,390 +1,214 @@
# Reader MCP Workflow Service
# Reader · AI 日报引擎
reader 当前已经收口为面向 OpenClaw 的 MCP workflow service。正式能力边界以 FreshRSS 日报工作流为准:启动 run、写入 `run-state.json`、查询运行状态、读取结构化结果,以及异步恢复 job。CLI 与同步入口仍保留,但定位为 debug / fallback,而不是正式集成入口。
> 从 FreshRSS 到 AI 日报的自动化流水线,为个人知识管理生成每日 AI 工程化简报。
## 运行
Reader 是一个端到端的 AI 日报生产系统,定时从自建 FreshRSS 的 RSS 订阅源拉取文章,经过内容提取、LLM 筛选与摘要、关键词索引构建,最终产出两个输出:
1. **公开日报** — 推送到 [Hugo 站点](https://osiman.site/daily/) 的精选技术简报
2. **知识沉淀** — 单篇结构化摘要上传到 IMA 知识库(`daily` KB)
整个流程由 OpenClaw 编排,作为 MCP Workflow Service 对外暴露。
---
## ✨ 核心能力
| 能力 | 说明 |
|:----|:------|
| **RSS 拉取** | 从 FreshRSS API 拉取订阅文章,支持增量读取与已读标记 |
| **内容提取** | 自动提取文章正文、标题、来源等结构化字段 |
| **LLM 筛选** | 基于个人兴趣画像(`filter_context.personal.json`)自动评估文章质量,分为 keep / review / drop 三档 |
| **LLM 摘要** | 并行生成每篇文章的结构化摘要(4 路并发,约 24 秒完成 7 篇) |
| **关键词索引** | 自动构建每日关键词索引,支持别名映射与停用词过滤 |
| **候选简报** | 生成 `digest-brief.json` 供编排层(OpenClaw)决策 |
| **单篇沉淀** | 对选中的文章生成结构化知识笔记,上传到 IMA 知识库 |
| **异步 Job** | 全部生产流程走异步 job,支持恢复与状态查询 |
---
## 🏗 架构概览
```
FreshRSS ──→ 拉取 ──→ 内容提取 ──→ LLM 筛选 ──→ 关键词索引
│
digest-brief.json
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Hugo 日报 IMA 知识库 term_index
(公开简报) (单篇沉淀) (关键词数据)
```
### MCP 工具层
Reader 通过 Hermes MCP 暴露 20+ 个工具,分为三类:
**日报流水线:**
- `start_freshrss_pipeline_job` → 启动异步日报 Job
- `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result` → 轮询结果
**状态查询:**
- `get_run_status` / `get_delivery_payload` / `get_run_report` → 读取运行结果
- `list_runs` / `list_run_artifacts` → 浏览运行历史
**恢复与单篇总结:**
- `inspect_resume_plan` / `start_resume_job` → 恢复失败 Job
- `start_article_summary_job` / `generate_article_summaries` → 单篇文章摘要
### CLI 入口
同步入口,适合本地 debug / fallback:
```bash
# 完整日报流水线
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
# 单篇文章摘要
python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
# 关键词维护
python scripts/build_keyword_index.py
python scripts/generate_term_cleanup_suggestions.py
```
---
## 🚀 快速开始
### 环境变量
```
# FreshRSS
FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
FRESHRSS_USERNAME=bot
FRESHRSS_API_PASSWORD=xxx
# LLM(主流水线)
LLM_API_URL=https://api.deepseek.com
LLM_API_KEY=xxx
LLM_MODEL=deepseek-chat
# LLM(可选,单篇摘要独立模型)
ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
ARTICLE_SUMMARY_LLM_API_KEY=xxx
ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
# IMA 知识库(可选,仅沉淀时需要)
IMA_DAILY_KNOWLEDGE_BASE_ID=xxx
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
```
### 运行
```bash
# 安装
pip install -e .
# 跑日报流水线(CLI 模式)
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
# 启动 MCP 服务(OpenClaw 集成用)
summary-mcp
```
服务当前暴露 21 个工具。
---
正式集成摘要:
## 📁 项目结构
- 主日报正式入口:`start_freshrss_pipeline_job`
- 主日报正式读取:`get_run_status`、`get_delivery_payload`、`get_run_report`
- 恢复正式入口:`inspect_resume_plan`、`start_resume_job`、`get_resume_job_status`、`get_resume_job_result`
- 单篇总结正式入口:`start_article_summary_job`、`get_article_summary_job_status`、`get_article_summary_job_result`
- `run_freshrss_openclaw_pipeline`、`resume_run`、`generate_article_summaries` 仅用于同步 debug / fallback
## 文档入口
如果你在做 OpenClaw 集成,不要只看这个 README,优先看:
- `docs/openclaw/README.md`
- `docs/openclaw/openclaw-handoff.md`
- `docs/openclaw/openclaw-orchestration-flow.md`
字段契约见:
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- `docs/openclaw/openclaw-delivery-payload-spec.md`
文档总索引见:
- `docs/README.md`
- `docs/current/context-reset-brief.md`
- `docs/design/README.md`
- `plans/README.md`
## 正式能力边界
- 当前正式 workflow 只有 `freshrss_daily_digest`
- 当前生产编排默认走异步 job,而不是同步 MCP / CLI
- 每次 FreshRSS 主流水线 run 都会在 `outputs/freshrss/rerun/<run_dir>/run-state.json` 落地运行真相
- OpenClaw 正式读取结果应优先使用 MCP 返回的 `run_id`、`output_dir`、`delivery_output`、`report_output`
- 正式恢复只支持带有效 `run-state.json` 的当前 run,不处理历史推断 run
- 正式生产恢复依赖 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`
## OpenClaw 最短调用路径
1. 调 `start_freshrss_pipeline_job`
2. 轮询 `get_freshrss_pipeline_job_status`
3. 成功后读 `get_freshrss_pipeline_job_result`,拿 `run_id`
4. 用 `get_run_status`、`get_delivery_payload`、`get_run_report` 做后续读取
5. 如需恢复,先调 `inspect_resume_plan`,只有 `recommended_action=resume` 才走 `start_resume_job`
更完整的状态分支、恢复策略和人工介入条件见 `docs/openclaw/openclaw-orchestration-flow.md`。
## 单篇文章总结后处理(可选使用独立 LLM)
### 生产环境推荐输入
- 单篇总结的正式生产输入,优先使用 FreshRSS 主流水线输出的**单篇 extracted 文件**:
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`
- 这些**逐条 extracted 文件**是下游单篇总结的**正式默认产物**。
- 像 `outputs/freshrss/extracted/freshrss.extracted.json` 这样的**批量 extracted 文件**,只作为临时场景、兼容旧流程的输入形态保留,**不是首选生产默认**。
### daily 知识库默认配置
- `IMA_DAILY_KNOWLEDGE_BASE_ID` —— 单篇日报总结默认上传的 IMA 知识库 ID
- `IMA_DAILY_KNOWLEDGE_BASE_NAME` —— 默认知识库名称(预期值:`daily`)
- 上传逻辑在运行时应先校验目标知识库;若配置的目标不存在,应先按名称查找,仍不存在则创建 `daily`
相关能力:
- 正式路径:`start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
- 同步 debug:`generate_article_summaries`
- CLI:`scripts/run_article_summaries.py`
- 后台 runner:`scripts/run_article_summary_job.py`
## 校验 LLM 摘要结果
```bash
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
```
reader/
├── configs/ # 配置
│ ├── filter_context.personal.json # 个人兴趣画像
│ ├── term_aliases.json # 关键词别名映射(149 条)
│ ├── term_stopwords.json # 关键词停用词(132 条)
│ └── term_cleanup_policy.json # 关键词清理策略
├── src/
│ └── summary_mcp/ # MCP 服务核心
│ ├── server.py # MCP 服务入口
│ ├── runtime/ # 运行时(Job 管理、状态持久化)
│ └── workflows/ # 工作流(日报流水线逻辑)
├── scripts/ # CLI 入口
├── outputs/ # 运行时产出
│ └── freshrss/
│ ├── rerun/<run_id>/ # 每次运行的全量产物
│ │ ├── candidates/ # digest-brief.json, delivery payload
│ │ ├── extracted/ # item-XX.extracted.json
│ │ └── run-state.json # 运行状态
│ └── single_summaries/ # 单篇摘要输出
├── data/
│ └── term_index/ # 关键词索引数据
│ ├── daily/YYYY-MM-DD.json
│ └── term_stats.json
├── docs/ # 设计文档
└── prompts/ # LLM Prompt 模板
```
## 跑最小 extraction → summary 循环
---
## ⚙️ 关键技术决策
| 决策 | 选择 | 原因 |
|:----|:----|:------|
| 运行模式 | **异步 Job** 为主,CLI fallback | 避免 MCP 传输层 120s 超时限制 |
| 摘要并发 | **ThreadPoolExecutor(max_workers=4)** | LLM 调用是 I/O 密集型,4 路并行将 7 篇摘要从 2-3 分钟压到 ~24 秒 |
| 环境变量 | **子进程显式注入 .env** | 解决 MCP 服务器环境隔离导致子进程读取不到 LLM_API_KEY 的问题 |
| 关键词过滤 | **别名映射 + 停用词 + 语义清洗** | 先用 `term_aliases.json` 归一化,再用 `term_stopwords.json` 过滤噪声,最后通过 LLM 做语义级清洗 |
| Tag 选择 | **复用已有通用 Tag**,不从 term_index 翻生僻词 | 保持 Hugo 站点 /tags/ 页面整洁,避免大量一次性专有名词 |
---
## 🔧 关键词治理
配置治理走四步流程(`scripts/` 下脚本):
```bash
python scripts/run_summary_loop.py ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--prompt outputs/prompts/llm-summary-prompt.txt ^
--output outputs/reference/summary/result.loop.json
# 1. 构建评审数据包
python skills/keyword-cleanup-review/scripts/build_review_bundle.py --days 7 --top 50
# 2. 统计规则级建议(大小写、单复数、频次阈值)
python scripts/generate_term_cleanup_suggestions.py
# 3. LLM 语义级建议(中英映射、简称-全称、近义词)
python scripts/generate_term_cleanup_semantic_suggestions.py
# 4. 确认后写入配置
python scripts/apply_term_suggestions.py --accept-watch ... --dry-run
```
## 拉取 FreshRSS 条目并映射为标准化 `item`
详见 `docs/design/keyword-engine-maintenance.md`。
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=bot
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/pull_freshrss_items.py --limit 5 --mark-read
```
---
默认会排除已经带 `read` 标签的条目。
如果你想拿到完整阅读列表,可以加 `--include-read`。
启用 `--mark-read` 后,脚本会在执行成功后把本次抓到的条目标记为已读。
## 🤖 Agent Skill
脚本会写出:
Reader 附带一个完整的 OpenClaw Agent Skill,位于 `skills/reader-digest-flow/`,供 AI Agent(Hermes / Claude Code 等)编排每日日报流程使用。
- `outputs/freshrss/raw/freshrss.raw.json`
- `outputs/freshrss/items/freshrss.items.json`
Skill 包含完整的 7 阶段工作流定义:
1. **Phase 1** — 跑 Pipeline(FreshRSS → 提取 → LLM 筛选)
2. **Phase 2** — 汇报候选(展示候选文章给用户决策)
3. **Phase 3** — 用户选文(选择 Hugo 发布文章)
4. **Phase 4** — 生成并发布 Hugo 日报
5. **Phase 5** — 用户选 IMA 沉淀文章
6. **Phase 6** — LLM 摘要生成
7. **Phase 7** — IMA 知识库上传
## 拉取 FreshRSS 条目并逐条做内容提取
以及海量铁律(不重跑 pipeline、编号规则、Tag 选择规范、IMA 上传流程等)和参考文件(`references/` 目录)。
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/run_freshrss_extract.py --limit 1 --mark-read
```
---
默认会排除已经带 `read` 标签的条目。
启用 `--mark-read` 后,只有提取成功的条目才会被标记为已读。
## 📄 文档
脚本会写出:
- `docs/openclaw/README.md` — OpenClaw 集成指南
- `docs/openclaw/openclaw-orchestration-flow.md` — 编排流程
- `docs/openclaw/openclaw-delivery-payload-spec.md` — 字段契约
- `docs/design/README.md` — 设计文档总索引
- `docs/design/filter-rule-engine-design.md` — 过滤规则引擎设计
- `docs/design/filter-rule-engine-usage.md` — 过滤规则使用说明
- `outputs/freshrss/raw/freshrss.raw.json`
- `outputs/freshrss/items/freshrss.items.json`
- `outputs/freshrss/extracted/freshrss.extracted.json`
---
注意:这个**批量 extracted 文件**主要用于独立提取场景和旧流程兼容。下游单篇总结的正式生产默认输入,仍然是 `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` 这类**逐条 extracted 文件**。
## 📝 License
## 跑完整 FreshRSS 流水线,并在最终 delivery payload 写盘成功后再标记已读
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
```
如果你希望过滤时引入个人工程兴趣 / AI Agent 兴趣画像,可以传入 context 文件:
```bash
python scripts/run_freshrss_pipeline.py ^
--limit 5 ^
--context configs/filter_context.personal.json ^
--mark-read
```
这条 CLI 与 MCP `run_freshrss_openclaw_pipeline` / `start_freshrss_pipeline_job` 共用同一条主流水线逻辑,但正式生产集成应优先走 async MCP job;CLI 与同步 MCP 入口仅用于本地 debug / fallback。默认会写出这些产物:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`(给 OpenClaw 生成 public digest 用的轻量输入,仅包含 `keep` 候选)
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`(每篇一份)
同时还会更新每日关键词索引运行数据:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
主流水线默认**不会**产出批量级的 `freshrss.extracted.json`。
如果你需要更多逐条中间产物,例如标准化 items、摘要结果、过滤决策、candidate record、candidate input,可以加 `--debug-artifacts`。
当 OpenClaw 接入这个 MCP 服务后,应先通过 `start_freshrss_pipeline_job` 启动任务,轮询 `get_freshrss_pipeline_job_status`,再从 `get_freshrss_pipeline_job_result` 读取稳定的 `run_id`。拿到 `run_id` 之后,再通过 `get_run_status` / `get_delivery_payload` / `get_run_report` 读取状态与结果,而不是直接拼接目录路径。`run_freshrss_openclaw_pipeline` 仅保留为同步 debug / fallback 路径。
在排查复杂问题时,也可以把 `debug_artifacts=true` 打开,并结合 `list_run_artifacts` 查看该 run 下实际产物。
## 对结构化摘要结果执行确定性过滤规则
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--output outputs/reference/filter/filter-decision.json
```
如果你希望注入兴趣主题或来源标签,也可以额外传入 context 文件:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--context outputs/reference/filter/filter-context.json ^
--output outputs/reference/filter/filter-decision.with-context.json
```
规则引擎设计和规则编写说明见:
- `docs/design/filter-rule-engine-design.md`
- `docs/design/filter-rule-engine-usage.md`
## 将过滤结果写入 Markdown sink
```bash
python scripts/run_markdown_sink.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json
```
脚本会把 Markdown 笔记写到 `knowledge-base/` 下。
## 构建内部 `ArticleCandidateRecord` 与精简版 `OpenClawCandidateInput`
```bash
python scripts/run_article_candidate.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json ^
--section-hint tools_and_workflows
```
默认会写出:
- `outputs/reference/candidates/article-candidate-record.json`
- `outputs/reference/candidates/openclaw-candidate-input.json`
## 构建批量 OpenClaw delivery payload
```bash
python scripts/build_openclaw_delivery.py ^
--input-dir outputs/freshrss/candidates/batch ^
--sort-by-rank ^
--date 2026-03-25
```
默认会写出:
- `outputs/reference/candidates/openclaw-delivery-payload.json`
输出目录布局说明见 `outputs/README.md`。
## 关键词索引默认配置
相关配置文件位于:
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
你也可以基于已有 delivery payload 重新构建关键词索引:
```bash
python scripts/build_keyword_index.py ^
--input outputs/reference/candidates/openclaw-delivery-payload.json
```
运行期关键词数据存放在:
- `data/term_index/`
关键词清理评审 skill 位于:
- `skills/keyword-cleanup-review/`
构建给关键词治理流程使用的评审数据包(review bundle,临时工作文件):
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
--days 7 ^
--top 50 ^
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
这个评审数据包(review bundle)会额外携带治理上下文:
- 来自 `configs/term_cleanup_policy.json` 的清理阈值
- 当前 watch list(`configs/term_watchlist.json`)
- 最近已应用的变更(`configs/term_change_log.json`)
接下来可以把 bundle 渲染成正式建议产物(默认走确定性规则,不把 LLM 作为默认路径):
```bash
python scripts/generate_term_cleanup_suggestions.py ^
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
```
默认只生成:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`(唯一正式建议产物,建议短期保留)
如需人工审阅展示稿,再显式加:
```bash
python scripts/generate_term_cleanup_suggestions.py ^
--bundle outputs/term_index/review/keyword-cleanup-bundle.json ^
--emit-markdown
```
这时才会额外生成:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿,可按需生成,不必作为长期资产保留)
其中本轮最小版本优先覆盖 `interest_keyword_suggestions` 和 `watch_terms` 主链路;`alias_suggestions` / `stopword_suggestions` 先保持保守。
产物保留策略建议:
- `data/term_index/daily/YYYY-MM-DD.json`、`data/term_index/term_stats.json` 作为事实层长期保留
- `configs/filter_context.personal.json`、`configs/term_watchlist.json`、`configs/term_aliases.json`、`configs/term_stopwords.json`、`configs/term_change_log.json` 作为状态层长期保留
- `term-cleanup-suggestions-YYYY-MM-DD.json` 作为正式建议产物短期保留(最近少量几份或仅保留已应用过的)
- `keyword-cleanup-bundle.json` 仅作为临时工作文件,默认只保留当前最新一份
- `term-cleanup-suggestions-YYYY-MM-DD.md` 仅作为临时展示稿,优先按需生成,不建议默认长期归档
如果你想先预览已接受建议,再决定是否写配置文件:
```bash
python scripts/apply_term_suggestions.py ^
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
--accept-watch Cron Heartbeat Memory ^
--dry-run
```
去掉 `--dry-run` 后才会真正写文件。
这个脚本也支持通过 `--accept-alias`、`--accept-stopword`、`--accept-interest` 应用 alias / stopword / interest keyword 变更。
已接受的 watch 词会写入 `configs/term_watchlist.json`,每次应用动作也会被追加到 `configs/term_change_log.json`。
## 单篇总结 LLM 配置
如果你希望单篇总结后处理使用独立模型,而不影响主流水线,可以设置:
- `ARTICLE_SUMMARY_LLM_API_URL`
- `ARTICLE_SUMMARY_LLM_MODEL`
- `ARTICLE_SUMMARY_LLM_API_KEY`
如果这些变量未设置,单篇总结会回退使用主流程中的 `LLM_*` / `OPENAI_*` 配置。
如果显式指定 DeepSeek 作为单篇总结模型且请求超时或失败,当前实现会自动再用主流程默认模型配置重试一次。
示例(PowerShell 风格):
```bash
set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
```
然后可以这样调用 CLI。
正式生产环境建议优先使用 `outputs/freshrss/rerun/<run_id>/extracted/` 下的**逐条 extracted 文件**;下面这个**批量 extracted** 示例仅保留为兼容旧流程 / 临时场景输入:
```bash
python scripts/run_article_summaries.py ^
--extracted outputs/freshrss/extracted/freshrss.extracted.json ^
--ids 12345 67890 ^
--output-dir outputs/freshrss/single_summaries ^
--timeout 120
```
也可以通过 `summary_mcp.server` 暴露的 MCP 工具 `generate_article_summaries` 调用:
- `extracted_path`(string):单篇 extracted JSON 路径(例如 `outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json`),或者包含 `results` 数组的 batch extracted JSON
- `selected_ids`(array of strings):必填,至少传一个 `item_id`;如果传入的 ID 在 extracted payload 中一个都匹配不到,会直接报错
- `output_dir`(optional string):Markdown 输出目录;若不传,则默认写到 extracted 文件旁边的 `single_summaries/` 目录
- `llm_api_key` / `llm_model` / `llm_api_url`(optional strings):单篇总结 LLM 的覆盖配置;不传时会按前文规则回退到 `ARTICLE_SUMMARY_*` 或主 `LLM_*`
该工具返回一个 JSON 数组,内容为生成好的 Markdown 文件路径。
OpenClaw / 正式集成建议优先走异步 job:
- 调 `start_article_summary_job` 启动任务,立即拿到 `job_id`
- 轮询 `get_article_summary_job_status(job_id)`,直到 `status` 进入 `success` 或 `failed`
- 成功后调用 `get_article_summary_job_result(job_id)` 读取 `written_paths` 与结构化结果
- 失败时优先查看 `error_summary` 与 job 目录中的 `job-report.json`
异步 job 状态目录固定落在 `outputs/freshrss/article_summary_jobs/<job_id>/`,最小会包含:
- `run-state.json`
- `input.json`
- `result.json`(成功时)
- `job-report.json`
单篇总结使用独立 prompt:`outputs/prompts/article-summary-prompt.txt`。
它与日报 prompt 完全独立,输出的是中文结构化知识笔记,包含这些部分:
- 核心结论
- 主要论点
- 关键方法 / 机制
- 重要细节
- 可复用启发
- 关键词
- 主题
MIT
+17
View File
@@ -320,6 +320,23 @@
---
### [DONE][P1] interest/watch 候选引擎从固定阈值改为百分位排名 + 增速因子
目标:
- 解决固定阈值(total_count>=3)不随数据量自适应的问题
- 引入趋势信号(growth 因子),识别近期集中爆发的词
- 支持 7 天、41 天、200 天数据量下取同样的 top 5%/5%-20% 而不需调阈值
要求:
- `build_review_bundle.py`:新增 percentile 和 growth 计算函数;候选池从固定阈值改为百分位 + 增速
- `configs/term_cleanup_policy.json`:升级为 v2 schema,percentile/growth 替代绝对阈值
- 不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
- 全量跑一次对比新旧产出,确认差异合理
方案文档:`plans/keyword-cleanup-interest-watch-engine-improvement.md`
---
### [DONE][P3] 更新 README / handoff / docs,明确 MCP 为正式入口
目标:
+27
View File
@@ -27,18 +27,30 @@
"Agent Skills",
"AgentScope",
"AI Agent",
"AI Coding Agent",
"AliSQL",
"Anthropic",
"Claude",
"Claude Code",
"CLAUDE.md",
"CLI",
"Context Engineering",
"Cursor",
"ChatGPT",
"DeepSeek",
"FastAPI",
"Gin",
"Go",
"gRPC",
"Harness Engineering",
"Hermes Agent",
"Java",
"Kafka",
"Kubernetes",
"LLM",
"Loop Engineering",
"MCP",
"MoE",
"MySQL",
"MySQL复制延迟",
"OpenAI",
@@ -47,15 +59,30 @@
"Prompt Engineering",
"Python",
"RAG",
"ReAct",
"ReActAgent",
"Redis",
"Skill",
"SKILL.md",
"Skills",
"Spring",
"SubAgent",
"TypeScript",
"Vibe Coding",
"Workflow",
"上下文压缩",
"上下文工程",
"上下文管理",
"云原生",
"代码审查",
"可观测性",
"向量数据库",
"多Agent协作",
"大模型",
"子Agent",
"强化学习",
"微服务",
"渐进式披露",
"知识库"
]
}
+146 -2
View File
@@ -1,5 +1,149 @@
{
"AI助手": "AI Agent",
"Agent": "Agent",
"Agent框架": "Agent",
"Agent能力": "Agent Skills",
"智能体": "AI Agent",
"Agentic架构": "Agentic架构",
"多Agent协作": "多Agent协作",
"多智能体架构": "多Agent协作",
"Multi-Agent": "多Agent",
"Subagent": "子Agent",
"Sub Agents验证": "子Agent",
"子Agent": "子Agent",
"子智能体": "子Agent",
"Coding Agent": "AI Coding Agent",
"AI编程": "AI Coding Agent",
"AI辅助编程": "AI Coding Agent",
"代码生成": "AI代码生成",
"代码审查": "Code Review",
"Prompt": "Prompt Engineering",
"Prompt Caching": "提示缓存",
"RAG": "RAG",
"图文RAG": "RAG",
"Prompt架构": "Prompt Engineering"
"向量检索": "向量检索",
"向量嵌入": "向量嵌入",
"Multi-Token Prediction": "多Token预测",
"Pair-In Pair-Out": "PIPO架构",
"PIPO": "PIPO架构",
"上下文管理": "上下文管理",
"上下文卸载": "上下文卸载",
"Self-GC": "上下文压缩",
"记忆管理": "上下文管理",
"会话管理": "上下文管理",
"Harness Engineering": "Harness工程化",
"Harness架构": "Harness工程化",
"Harness": "Harness工程化",
"Loop Engineering": "Loop Engineering",
"推理加速": "推理加速",
"推理深度": "推理深度",
"长链路推理": "长链路推理",
"RLVR": "RLVR",
"GRPO": "GRPO",
"强化学习": "强化学习",
"Multi-Agent RL": "多Agent强化学习",
"Viking AI搜索": "AI搜索",
"Viking AI Search": "AI搜索",
"智能搜索": "AI搜索",
"SearchCLI": "CLI搜索",
"视频生成": "AI视频生成",
"视频生成模型": "AI视频生成",
"LingBot-Video": "AI视频生成",
"视觉自回归模型": "AI视频生成",
"火山云数据库PostgreSQL Serverless版": "Serverless数据库",
"PostgreSQL": "PostgreSQL",
"MySQL": "MySQL",
"OceanBase": "OceanBase",
"StarRocks": "StarRocks",
"Milvus": "Milvus",
"Seal AI Zone": "AI安全",
"NEX沙箱": "沙箱隔离",
"MicroVM": "沙箱隔离",
"安全左移": "安全左移",
"安全中台": "AI安全",
"成本降低": "成本优化",
"成本杠杆": "成本优化",
"Scale-to-Zero": "弹性伸缩",
"Data as Git": "数据分支管理",
"Schema Diff": "Schema对比",
"Time Travel": "数据回溯",
"多端架构": "多端架构",
"契约化": "契约化架构",
"大仓": "大仓工程化",
"Vibe Coding": "Vibe Coding",
"LLM Judge": "LLM评估",
"SWE-Bench": "SWE-Bench",
"SWE Bench Pro": "SWE-Bench",
"SWE-Bench Pro": "SWE-Bench",
"Verification Agent": "验证Agent",
"CLI工具": "CLI",
"CLI": "CLI",
"漏桶算法": "限流架构",
"固定窗口限流": "限流架构",
"Suspend消费控制": "限流架构",
"RocketMQ LiteTopic": "消息队列",
"LLM Wiki": "LLM知识库",
"知识工程": "知识工程",
"语义资产": "语义资产管理",
"知识图谱": "知识图谱",
"知识库沉淀": "知识管理",
"Skill": "Skill",
"Skill Hub": "技能生态",
"具身智能": "具身智能",
"Open X-Embodiment": "具身智能",
"YOLO Classifier": "目标检测",
"MCP": "MCP",
"MCP连接器": "MCP",
"缓存击穿": "缓存优化",
"GPU算力调度": "算力调度",
"异构资源": "异构计算",
"XPU": "异构计算",
"弹性RDMA": "RDMA网络",
"国内主流GPU": "国产芯片",
"国产AI芯片": "国产芯片",
"Paxos协议": "分布式一致性",
"Token": "Token管理",
"Token效率": "Token管理",
"百万token上下文": "长上下文",
"MoE": "MoE架构",
"MoE架构": "MoE架构",
"思维链": "思维链",
"CoT Distillation": "思维链蒸馏",
"自然语言驱动": "自然语言交互",
"NL2SQL": "NL2SQL",
"AI对齐": "AI对齐",
"注意力机制": "注意力机制",
"多模态": "多模态",
"音视频工作台": "音视频处理",
"AI助手": "AI Agent",
"Agent架构": "AI Agent",
"Agent专业化": "AI Agent",
"Agent Teams": "多Agent协作",
"Agentic Engineering": "AI Agent",
"AI智能体": "AI Agent",
"LLM Agent": "AI Agent",
"AI Harness": "Harness Engineering",
"AI代码生成": "AI Coding Agent",
"Memory管理": "上下文管理",
"Agent Skill": "Agent Skills",
"Binlog": "binlog",
"vibe coding": "Vibe Coding",
"Agent组织化协作平台": "Agent协作平台",
"Anthropic": "Anthropic",
"OpenClaw": "OpenClaw",
"WorkBuddy": "WorkBuddy",
"Claude": "Claude",
"ChatGPT": "ChatGPT",
"GPT": "GPT",
"Opus": "Opus",
"Sonnet": "Sonnet",
"Grok": "Grok",
"Qwen": "Qwen",
"GLM": "GLM",
"Claude Code": "Claude Code",
"Cursor": "Cursor",
"Codex": "Codex",
"Pi": "Pi",
"CoT": "CoT",
"SVG": "SVG",
"TTS": "TTS"
}
+491
View File
@@ -99,6 +99,497 @@
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
"suggestion_date": "2026-04-08",
"based_on_days": 7
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Anthropic",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=13, days_seen=10, recent_count=13.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Harness Engineering",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=12, days_seen=11, recent_count=12.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Skill",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=11, days_seen=9, recent_count=11.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "上下文工程",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "多Agent协作",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Claude",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "上下文管理",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "渐进式披露",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=5, recent_count=5.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Skills",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "SKILL.md",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "CLAUDE.md",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "上下文压缩",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "AI Coding Agent",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Hermes Agent",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Vibe Coding",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Context Engineering",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "Cursor",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T08:20:35.134979Z",
"action": "add_interest_keyword",
"term": "大模型",
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=3, recent_count=4.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_interest_keyword",
"term": "TypeScript",
"reason": "Core language for AI agent development (e.g., Claude Code, Cursor) and backend engineering, complements existing Python/Java/Go keywords.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_interest_keyword",
"term": "代码审查",
"reason": "Chinese term for 'code review', a key practice in backend engineering and AI agent development workflows.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Channels",
"reason": "Too generic; could refer to communication channels, YouTube channels, or software channels, not specific to user's focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Memory",
"reason": "Extremely broad term; could refer to computer memory, human memory, or memory in various contexts, not discriminative enough.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Prompt",
"reason": "Already covered by 'Prompt Engineering' as a more specific term; 'Prompt' alone is too broad and matches many unrelated articles.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "AGI",
"reason": "Too broad and speculative; not directly actionable for the user's practical engineering focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "AI日报",
"reason": "Generic news term; not a technical concept or tool, would add noise to the keyword index.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "AIHOT",
"reason": "Unclear meaning, likely a brand or aggregator, not a specific technical term.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "All In Code",
"reason": "Too vague; could refer to a podcast, a philosophy, or a project, not a specific technical concept.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "auto-twitter-campaign",
"reason": "Too specific to a single project/tool, not a general interest keyword for the user's focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "ChangeSet",
"reason": "Generic term used in version control and databases; too broad to be a useful filter.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Lumina",
"reason": "Unclear reference; could be a product, framework, or brand, not clearly aligned with user's focus.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "OpenViking",
"reason": "Unclear reference; not a known tool or concept in the user's stated focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "Seedance 2.0",
"reason": "Unclear reference; likely a product or version, not a general technical term.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_stopword",
"term": "质量门禁",
"reason": "Chinese term for 'quality gate', too generic in software engineering; not specific to user's focus areas.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agent Skill",
"reason": "Singular variant",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "Agent Skills",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Binlog",
"reason": "Case variant (auto-ranked)",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "binlog",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Coding Agent",
"reason": "Abbreviated form of 'AI Coding Agent', referring to the same concept.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Coding Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Subagent",
"reason": "Case variant",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "SubAgent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Subagents",
"reason": "Plural variant",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "SubAgent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "vibe coding",
"reason": "Case variant",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "Vibe Coding",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agent架构",
"reason": "Chinese translation of 'Agent architecture', a core concept in AI Agent engineering.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agent专业化",
"reason": "Chinese term for 'Agent specialization', directly related to Agent engineering.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agent Teams",
"reason": "English equivalent of 'Multi-Agent collaboration', same concept.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "多Agent协作",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "Agentic Engineering",
"reason": "Broader term for engineering with AI agents, closely related to Agent engineering focus.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "CLI工具",
"reason": "Chinese translation of 'CLI tool', same concept.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "CLI",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "AI编程",
"reason": "Chinese term for 'AI programming', closely related to AI Coding Agent.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "AI Coding Agent",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "记忆管理",
"reason": "Chinese term for 'memory management', closely related to context management in LLM applications.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "上下文管理",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-05-14T09:12:50.749877Z",
"action": "add_alias",
"term": "会话管理",
"reason": "Chinese term for 'session management', related to context management in LLM applications.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
"value": "上下文管理",
"suggestion_date": "2026-05-14",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "强化学习",
"reason": "top 0.8% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=10, days_seen=10, recent_count=10.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "ReAct",
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=8, recent_count=8.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "CLI",
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=7, recent_count=8.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "子Agent",
"reason": "top 1.3% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=7, days_seen=6, recent_count=7.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "Loop Engineering",
"reason": "top 1.5% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
},
{
"applied_at": "2026-07-15T02:17:50.155586Z",
"action": "add_interest_keyword",
"term": "MoE",
"reason": "top 2.0% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
"suggestion_date": "2026-07-15",
"based_on_days": 365
}
]
}
+10 -9
View File
@@ -1,14 +1,12 @@
{
"schema_version": "v1",
"schema_version": "v2",
"interest_keyword_review": {
"min_total_count": 3,
"min_days_seen": 2
"percentile_max": 0.05,
"growth_promotion": 0.5
},
"watch_term_review": {
"min_total_count": 1,
"min_days_seen": 1,
"max_total_count": 2,
"max_days_seen": 2
"percentile_min": 0.05,
"percentile_max": 0.20
},
"alias_review": {
"min_total_count": 2,
@@ -19,7 +17,10 @@
"max_days_seen": 2
},
"notes": [
"当前阶段采用保守阈值,避免在低样本条件下直接扩充 interest_keywords。",
"watch_terms 先用于观察,后续再决定是否升格为 interest_keywords 或进入 alias/stopword 配置。"
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
"growth = recent_count / total_count,衡量近期活跃度",
"watch_term_review 的 percentile_min 可理解为兴趣边界下限,低于此值的词归入 interest 候选",
"growth_promotion(默认 0.5)用于识别近期集中爆发词,即使排位不高也主动推荐确认"
]
}
+121 -1
View File
@@ -1,6 +1,126 @@
[
"1688",
"AGI",
"AIHOT",
"AI日报",
"All In Code",
"Andrej Karpathy",
"Anthropic",
"auto-twitter-campaign",
"Boundaries",
"ChangeSet",
"Channels",
"Claude Fable 5",
"Claude Mythos",
"Cohere",
"Confidence Head",
"Cosmos 3",
"Databricks",
"DINOv2",
"domain-mapping",
"Dropbox",
"EchoGen",
"FLUX.1-dev VAE",
"GB300 GPU",
"GLM 5.2",
"GLM5.0",
"GPT-5.5",
"GPT-Live",
"GPT5.5",
"Grok 4.5",
"GrowBrain",
"iMedImage",
"iMedLoop",
"iMedMaaS",
"iMedStudio",
"J-space",
"JLens",
"John Jumper",
"J空间",
"KAIROS",
"KubeRay",
"LibTV Agent",
"LingBot-Video",
"Lumina",
"Markdown",
"Marvis",
"MDASH",
"Meta Superintelligence Labs",
"MTS",
"Muse Image",
"Muse Video",
"N-gram Embedding",
"OCP China",
"OCP China 2026",
"On-Policy Distillation",
"OPC训练营",
"OpenAI",
"OpenBMC",
"OpenClaw",
"OpenViking",
"Opus 4.8",
"Qwen3",
"Qwen3-30B-A3B",
"RAS API",
"Redfish",
"ScMoE",
"Seal AI Zone",
"SealRouter",
"Seedance 2.0",
"Sonnet 5",
"Spec模式",
"STE固件团队",
"Three.js",
"Unity AI Gateway",
"Vant Weapp",
"WeTV",
"WorkBuddy",
"wpc",
"YOLO Classifier",
"一人公司",
"中国科学技术大学",
"五大扶持体系",
"出门问问",
"分镜",
"剧本",
"奋斗文化",
"字节跳动",
"小银",
"得力",
"德适科技",
"成都天府长岛",
"扣子",
"星云平台",
"火山引擎",
"百度百舸",
"百炼网关",
"科大讯飞",
"腾讯云开发者社区",
"腾讯混元Hy3",
"蚂蚁灵波",
"贝尔实验室",
"质量门禁",
"配乐",
"配音",
"银行客户经理",
"飞盘物理"
"飞书妙搭",
"飞盘物理",
"自动化",
"定时任务",
"开源模型",
"陌生化",
"AlphaFold",
"Brand Kit",
"DataWorks",
"Enhance-Nanocodec",
"IRIS Codec",
"Lovart",
"MiniMax M3",
"Gemini 3.5 Flash",
"Codex",
"CodeBuddy",
"Claude Cowork",
"AGENTS.md",
"Claude",
"RLVR"
]
+1 -1
View File
@@ -1,6 +1,6 @@
{
"schema_version": "v1",
"updated_at": "2026-04-08T02:34:14.194320Z",
"updated_at": "2026-07-15T02:17:50.155586Z",
"terms": [
{
"term": "A2A",
@@ -0,0 +1,151 @@
# 关键词清洗流程概述
> 2026-05-14 初版
> 从"数据记录"到"人工确认落盘"的完整链路
---
## 整体数据流
```
每日日报 pipeline
│
▼
term_index/daily/YYYY-MM-DD.json ← 每天一篇候选文章的热词统计
│
▼
term_index/term_stats.json ← 所有 daily 的汇总(1070 个词)
│
├──── build_review_bundle.py ← 打包为审查数据包
│ │
│ ▼
│ review/keyword-cleanup-bundle.json
│ │
│ ▼
│ generate_term_cleanup_suggestions.py
│ │
│ ▼
│ review/term-cleanup-suggestions-YYYY-MM-DD.json ← 正式建议产物
│ │
│ ▼
│ (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md ← 展示稿
│
├──── 人工确认哪些建议 accept
│
▼
apply_term_suggestions.py ← 写入配置
│
├── configs/filter_context.personal.json ← interest_keywords
├── configs/term_aliases.json ← alias
├── configs/term_stopwords.json ← stopword
├── configs/term_watchlist.json ← watch
└── configs/term_change_log.json ← 变更日志
```
---
## 各环节说明
### 阶段 1:数据记录(每日自动)
```bash
# FreshRSS pipeline 跑完后自动产出
data/term_index/daily/2026-05-14.json
```
- 每天一篇,记录当天候选文章中出现的热词
- 包含 term、total_count、days_seen 等信息
- 目前累计 **41 天**,共 **1070 个独立词**
### 阶段 2:全量汇总(每日自动)
```bash
data/term_index/term_stats.json
```
- 从所有 daily 文件重建,会覆盖重跑
- 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13)
### 阶段 3:构建审查数据包(手动触发)
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 \
--top 100 \
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
- 把 term_stats + 当前配置打成一包,方便后续处理
- 输出:`review/keyword-cleanup-bundle.json`
### 阶段 4:生成建议(手动触发)
```bash
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
#### 当前产出能力
| 建议类型 | 状态 | 当前阈值 | 说明 |
|---------|------|----------|------|
| interest_keyword_suggestions | ✅ **已实现** | total≥3, days≥2 | 产出 20 条 |
| watch_terms | ✅ **已实现** | total≤2, days≤2 | 本次 0 条 |
| alias_suggestions | ❌ **硬编码为空** | policy 有阈值(total≥2, days≥2)但脚本未实现 | |
| stopword_suggestions | ❌ **硬编码为空** | policy 有阈值(total≤2, days≤2)但脚本未实现 | |
**关键发现:** alias 和 stopword 不是"阈值太保守",是 **generate 脚本里压根没写对应的生成函数**。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 `[]` 和 `0`。
### 阶段 5:人工确认(手动)
```
OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply
```
本次模式:
- 高频(≥5次/5天以上)→ 强烈推荐 ✅
- 中频(3-4次)→ 附带建议 ✅
- 泛词 → 建议跳过 ❌
### 阶段 6:落盘配置(手动)
```bash
python scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 词1 词2 ...
```
- dry-run 预览 → 确认后正式 apply
- 写入 `configs/filter_context.personal.json`
- 同步记录到 `term_change_log.json`
- **不备份原始配置**(待优化)
- **apply 后不自动清理 review 目录**(待优化)
### 阶段 7:维护清理(按需)
由 OpenClaw 侧 `reader-keyword-maintenance` skill 处理:
- 删除旧 markdown 展示稿
- 保留最近一份 bundle
- 保守保留 suggestions JSON
---
## 当前配置资产
| 文件 | 内容 | 数据量 |
|------|------|--------|
| `filter_context.personal.json` | interest_keywords | 52 个 |
| `term_aliases.json` | 别名映射 | 0 组(未启用) |
| `term_stopwords.json` | 停用词 | 0 个(未启用) |
| `term_watchlist.json` | 观察词 | 6 个 |
| `term_change_log.json` | 所有变更记录 | 已记录 |
---
## 待优化项
1. **alias/stopword 建议生成为空** — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数
2. **apply 前无配置备份** — 建议 apply 前自动 cp 备份
3. **apply 后无自动收尾** — 建议 apply 后自动删旧 markdown 和 bundle
4. **alias 识别依赖规则而非 LLM** — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply
@@ -0,0 +1,290 @@
# interest/watch 候选引擎改进方案
> 从固定阈值到自适应排位 + 趋势因子的演进
## 1. 背景
### 1.1 当前实现
`build_review_bundle.py` 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:
| 候选池 | 判断条件 | 依据 |
|--------|---------|------|
| `interest_review_candidates` | `total_count >= 3 AND days_seen >= 2` | `policy.interest_keyword_review` |
| `watch_review_candidates` | `total_count <= 2 AND days_seen <= 2` | `policy.watch_term_review` |
`generate_term_cleanup_suggestions.py` 则直接从这两个候选池过滤、去重、排序后输出。
### 1.2 当前方案的问题
**问题一:固定阈值不随数据量自适应**
```
场景 total_count=3 意味着什么
─────────────────────────────────────────────
7 天数据(~200 词) top 15%,有一定区分度 ✅
41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变
未来 200 天 仍然用 3 次,区分度稀释 ❌
```
同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。
**问题二:固定阈值忽略趋势信号**
- "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
- "Channels":total=3, recent=0 — 早期出现但近期消失
- 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。
**问题三:interest 和 watch 的分界线是硬的**
total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。
### 1.3 讨论结论
与老大讨论后确认:
1. interest/watch 是**统计判断**,不需要大模型介入,纯算法可以解决
2. 当前引擎缺的不是大模型,而是**算法本身没写完**——自适应维度(排位、趋势)还没实现
3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
4. 修改量小,可以在 1 小时内落地
---
## 2. 设计方案
### 2.1 核心思路
引入两个互补维度替代固定阈值:
```
判定维度 含义 数据来源
────────────────────────────────────────────────────────────
percentile(百分位排名) 该词 total_count 在所有词 term_stats
中的排位占比
growth(增速因子) 近期集中度 = recent_count daily 近 N 天
/ total_count
```
两个维度配合:
- **percentile** 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
- **growth** 衡量"这个词是持续出现还是近期爆发"——识别趋势信号
### 2.2 候选池划分逻辑
```
percentile
│
┌─────────────────────┐
│ top 5% │
│ → 建议 interest │ ← 高频稳定词
├─────────────────────┤
│ top 5%-20% │
│ → 建议 watch │ ← 有信号但未达 threshold
├─────────────────────┤
│ bottom 80% │
│ → 暂不处理 │ ← 噪声/低频
└─────────────────────┘
额外规则:
如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
→ 主动提升到 watch / 主动推 confirm
```
这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。
### 2.3 接口变化
**`configs/term_cleanup_policy.json`**:
```json
{
"schema_version": "v2",
"interest_keyword_review": {
"percentile_max": 0.05,
"growth_promotion": 0.5
},
"watch_term_review": {
"percentile_min": 0.05,
"percentile_max": 0.20
}
}
```
`v1` 的 `min_total_count`/`min_days_seen` 等绝对阈值字段不再使用。
**`build_review_bundle.py` 输出的候选项**:
```json
{
"term": "Anthropic",
"total_count": 13,
"days_seen": 10,
"percentile": 0.012,
"growth": 0.54,
"reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
}
```
### 2.4 不需要改动的部分
- `generate_term_cleanup_suggestions.py` — 它只消费候选池,不用改
- `apply_term_suggestions.py` — 消费 suggestions JSON,不用改
- `keyword-cleanup-bundle.json` 结构 — 向后兼容,新增 percentile/growth 字段
---
## 3. 实施计划
### 3.1 改动范围
| 文件 | 改动量 | 内容 |
|------|--------|------|
| `skills/keyword-cleanup-review/scripts/build_review_bundle.py` | ~40 行 | 新增 `_compute_percentile()` 和 `_compute_growth()` 函数;修改候选池生成逻辑;候选项中增加 percentile/growth |
| `configs/term_cleanup_policy.json` | ~10 行 | schema v2:percentile/growth 替代绝对阈值 |
### 3.2 实施步骤
1. **build_review_bundle.py**:在 `top_global_terms` 生成后,增加 percentile 计算函数和 growth 计算函数
2. **build_review_bundle.py**:修改 `interest_review_candidates` 和 `watch_review_candidates` 的生成逻辑,从固定阈值改为 percentile + growth
3. **build_review_bundle.py**:候选项增加 `percentile` 和 `growth` 字段,更新 `reason` 文案
4. **term_cleanup_policy.json**:更新为 v2 schema
5. **验证**:全量跑一次(`--days 365 --top 100`),对比新旧两份输出的差异
### 3.3 验证方法
```bash
# 1. 用旧版生成 baseline
cd /home/ubuntu/zhu/github/reader
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-baseline.json
# 2. 改代码后用新版生成
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-new.json
# 3. 对比 governance_hints
python3 -c "
import json
a = json.load(open('/tmp/bundle-baseline.json'))
b = json.load(open('/tmp/bundle-new.json'))
for key in ['interest_review_candidates', 'watch_review_candidates']:
old = set(i['term'] for i in a['governance_hints'][key])
new = set(i['term'] for i in b['governance_hints'][key])
print(f'{key}: 新增={new-old}, 减少={old-new}')
"
```
### 3.4 风险
| 风险 | 概率 | 应对 |
|------|------|------|
| 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 |
| growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 |
| 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 |
---
## 4. alias/stopword 设计方案
### 4.1 核心判断
alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。
| 类型 | 需要什么 | 判断方式 |
|------|---------|----------|
| 大小写变体 | 表层 | 规则:casefold 去重 |
| 单复数 | 表层 | 规则:去/加 s 后缀匹配 |
| 分词变体(空格/连字符) | 表层 | 规则:去空格归一 |
| 简写全称(MCP→Model Context Protocol) | **语义** | LLM |
| 中英文(上下文工程→Context Engineering) | **语义** | LLM |
| 同义不同名(Rush→猿辅导 Rush 平台) | **语义** | LLM |
| stopword(大模型、AI 太泛) | **语义** | LLM |
### 4.2 分层方案
```
输入:高频未覆盖词 + 已有 interest 词表
│
├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
│ 输出候选 alias 对
│
└── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
做语义聚类
输出 alias 组 + stopword 标记
```
### 4.3 规则层设计
在 `generate_term_cleanup_suggestions.py` 中新增 `_prepare_alias_suggestions()` 函数:
```python
def _prepare_alias_suggestions(top_terms, interest_keywords):
"""
基于表层规则生成 alias 建议。
规则1:casefold 匹配——同一个 casefold 下有多个原文变体
规则2:单复数——去掉/加上末尾 s 后匹配
规则3:分词变体——去空格/连字符后匹配
"""
```
优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。
### 4.4 LLM 层设计
单独脚本,非 generate 主链路的一部分。
```bash
python scripts/generate_term_cleanup_semantic_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
LLM prompt 设计:
```
你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
请做三件事:
1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest
用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地
```
LLM 层输出格式:
```json
{
"alias_suggestions": [
{"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
],
"stopword_suggestions": [
{"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
]
}
```
### 4.5 预期效果
| 覆盖类型 | 规则层 | LLM 层 |
|---------|--------|--------|
| 大小写变体 | ✅ | — |
| 单复数 | ✅ | — |
| 分词变体 | ✅ | — |
| 简写全称 | — | ✅ |
| 中英文映射 | — | ✅ |
| 同义不同名 | — | ✅ |
| stopword 判断 | — | ✅ |
---
## 5. 讨论记录
- 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
- 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
- 2026-05-14:确认本方案不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
- 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先
+314
View File
@@ -0,0 +1,314 @@
#!/usr/bin/env python3
"""
Generate semantic keyword suggestions using LLM.
Covers what surface-form rules cannot:
- semantic alias (abbreviation ↔ full name, Chinese ↔ English, synonym)
- stopword (overly broad / low-discrimination terms)
- promote (new term that aligns with user's focus areas)
Usage:
python scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-2026-05-14.json
"""
from __future__ import annotations
import argparse
import json
import os
import re
import sys
import time
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
from urllib.request import Request, urlopen
REPO_ROOT = Path(__file__).resolve().parents[1]
DEFAULT_BUNDLE_PATH = REPO_ROOT / "outputs" / "term_index" / "review" / "keyword-cleanup-bundle.json"
DEFAULT_OUTPUT_DIR = REPO_ROOT / "outputs" / "term_index" / "review"
def _load_json(path: Path) -> Any:
return json.loads(path.read_text(encoding="utf-8-sig"))
def _save_json(path: Path, payload: dict[str, Any]) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def _load_env(path: Path) -> dict[str, str]:
"""Load key=value pairs from .env file."""
env: dict[str, str] = {}
if not path.exists():
return env
for line in path.read_text(encoding="utf-8").splitlines():
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
env[key.strip()] = value.strip().strip("\"'")
return env
def _build_prompt(
interest_keywords: list[str],
rule_alias_suggestions: list[dict[str, str]],
candidate_terms: list[dict[str, Any]],
relevant_watch_terms: list[dict[str, Any]],
) -> str:
"""Build the LLM prompt for semantic suggestions."""
interest_bullets = "\n".join(f" - {t}" for t in sorted(interest_keywords))
candidate_bullets = "\n".join(
f" - {t['term']} (count={t['total_count']}, days={t['days_seen']})"
for t in candidate_terms[:40]
)
# Alias from rule layer (for LLM to build on, not duplicate)
rule_alias_text = ""
if rule_alias_suggestions:
rule_alias_text = "\nSurface-form alias (already identified, skip these):\n" + "\n".join(
f" {a['from']} → {a['to']} ({a['reason']})"
for a in rule_alias_suggestions
)
watch_text = ""
if relevant_watch_terms:
watch_text = "\nWatch terms (low-frequency but potentially relevant):\n" + "\n".join(
f" {t['term']} (count={t['total_count']}, days={t['days_seen']})"
for t in relevant_watch_terms[:20]
)
return f"""You are a keyword governance assistant for an AI engineer. Your job is to analyze keyword data and produce structured suggestions.
## User's focus areas
- AI Agent engineering (Skills, Harness, MCP, Agent architecture)
- Backend engineering (Java, Go, Kubernetes, MySQL, distributed systems)
- Open source AI tools and practices (Claude Code, Cursor, DeepSeek, OpenClaw)
- LLM application engineering (context engineering, RAG, prompt engineering)
## Interest keywords (52 already configured)
{interest_bullets}
## Uncovered candidate terms (sorted by frequency)
{candidate_bullets}
{watch_text}{rule_alias_text}
## Task
Analyze the candidate terms and output a JSON object with exactly three keys:
1. "semantic_alias": array of alias suggestions that SURFACE RULES CAN'T CATCH (e.g. abbreviation↔full name, Chinese↔English, different naming for the same concept).
Format: [{{"from": "<variant>", "to": "<canonical interest keyword>", "reason": "<why>"}}]
2. "stopword": array of terms that are too broad/generic to be useful as filters. A stopword is a term that appears frequently but has LOW DISCRIMINATION — it matches too many unrelated articles and clutters the keyword index.
Format: [{{"term": "<term>", "reason": "<why it should be a stopword>"}}]
3. "promote_to_interest": array of uncovered terms that align well with the user's focus areas and should be added as interest keywords.
Format: [{{"term": "<term>", "reason": "<why it fits>"}}]
## Rules
- Be conservative. When in doubt, leave it out.
- Only suggest alias for terms that clearly refer to the SAME concept as an existing interest keyword.
- Only suggest stopword for terms that are genuinely too broad (appear in many unrelated contexts).
- Only suggest promote for terms that clearly match the user's stated focus areas.
- Output valid JSON only, no markdown, no explanation outside the JSON."""
def _call_llm(prompt: str, api_url: str, model: str, api_key: str) -> str:
"""Call LLM API and return the response text."""
payload = json.dumps({
"model": model,
"messages": [{"role": "user", "content": prompt}],
"temperature": 0.1,
"max_tokens": 2048,
}).encode("utf-8")
req = Request(
api_url.rstrip("/") + "/chat/completions",
data=payload,
headers={
"Content-Type": "application/json",
"Authorization": f"Bearer {api_key}",
},
)
max_retries = 3
for attempt in range(max_retries):
try:
with urlopen(req, timeout=120) as resp:
result = json.loads(resp.read().decode("utf-8"))
return result["choices"][0]["message"]["content"]
except Exception as e:
if attempt < max_retries - 1:
wait = 2 ** attempt
print(f" LLM call failed (attempt {attempt+1}/{max_retries}): {e}", file=sys.stderr)
print(f" Retrying in {wait}s...", file=sys.stderr)
time.sleep(wait)
else:
raise
def _parse_llm_response(text: str) -> dict[str, list[dict[str, str]]]:
"""Extract JSON from LLM response (may contain markdown fences)."""
# Try to find JSON block
json_match = re.search(r"```(?:json)?\s*\n?(\{.*?\})\s*\n?```", text, re.DOTALL)
if json_match:
text = json_match.group(1)
# Clean up: remove any text before { or after }
start = text.find("{")
end = text.rfind("}")
if start >= 0 and end > start:
text = text[start : end + 1]
try:
result = json.loads(text)
except json.JSONDecodeError:
# Try partial recovery
print(f" Warning: LLM response not clean JSON, attempting recovery", file=sys.stderr)
print(f" Raw: {text[:500]}", file=sys.stderr)
return {"semantic_alias": [], "stopword": [], "promote_to_interest": []}
# Normalize keys
normalized = {
"semantic_alias": result.get("semantic_alias", result.get("alias", [])),
"stopword": result.get("stopword", result.get("stopword_suggestions", [])),
"promote_to_interest": result.get("promote_to_interest", result.get("promote", [])),
}
# Ensure each is a list
for key in normalized:
if not isinstance(normalized[key], list):
normalized[key] = []
return normalized
def main() -> None:
parser = argparse.ArgumentParser(description="Generate semantic keyword suggestions via LLM.")
parser.add_argument("--bundle", type=Path, default=DEFAULT_BUNDLE_PATH, help="Review bundle JSON path")
parser.add_argument("--suggestions", type=Path, default=None, help="Existing suggestions JSON (for rule alias context)")
parser.add_argument("--output", type=Path, default=None, help="Output JSON path (auto-generated if omitted)")
parser.add_argument("--llm-api-url", type=str, default=None, help="LLM API base URL")
parser.add_argument("--llm-model", type=str, default=None, help="LLM model name")
parser.add_argument("--llm-api-key", type=str, default=None, help="LLM API key")
parser.add_argument("--dry-run", action="store_true", help="Print prompt and exit without calling LLM")
args = parser.parse_args()
# Load config
env_path = REPO_ROOT / ".env"
env = _load_env(env_path) if env_path.exists() else {}
api_url = args.llm_api_url or os.environ.get("LLM_API_URL") or env.get("LLM_API_URL", "https://api.deepseek.com")
# Map OpenClaw model aliases to actual API model names
model_raw = args.llm_model or os.environ.get("LLM_MODEL") or env.get("LLM_MODEL", "deepseek-chat")
MODEL_ALIAS_MAP = {
"deepseek/deepseek-v4-flash": "deepseek-chat",
"deepseek/deepseek-chat": "deepseek-chat",
"deepseek-v4-flash": "deepseek-chat",
"deepseek-chat": "deepseek-chat",
}
model = MODEL_ALIAS_MAP.get(model_raw, model_raw)
api_key = args.llm_api_key or os.environ.get("LLM_API_KEY") or env.get("LLM_API_KEY", "")
if not api_key:
print("Error: No LLM API key found. Set LLM_API_KEY in .env or pass --llm-api-key.", file=sys.stderr)
sys.exit(1)
# Load bundle
if not args.bundle.exists():
print(f"Error: Bundle not found: {args.bundle}", file=sys.stderr)
sys.exit(1)
bundle = _load_json(args.bundle)
current_config = bundle.get("current_config", {})
interest_keywords = current_config.get("interest_keywords", [])
top_global_terms = bundle.get("top_global_terms", [])
governance_hints = bundle.get("governance_hints", {})
# Build candidate list (uncovered terms from interest + watch candidates)
candidate_terms = []
for item in governance_hints.get("interest_review_candidates", []):
if isinstance(item, dict):
candidate_terms.append({
"term": item.get("term", ""),
"total_count": item.get("total_count", 0),
"days_seen": item.get("days_seen", 0),
"percentile": item.get("percentile", 0),
"growth": item.get("growth", 0),
})
for item in governance_hints.get("watch_review_candidates", []):
if isinstance(item, dict):
# Avoid duplicates
if not any(c["term"] == item.get("term") for c in candidate_terms):
candidate_terms.append({
"term": item.get("term", ""),
"total_count": item.get("total_count", 0),
"days_seen": item.get("days_seen", 0),
"percentile": item.get("percentile", 0),
"growth": item.get("growth", 0),
})
# Sort by total_count descending
candidate_terms.sort(key=lambda x: -x["total_count"])
relevant_watch_terms = governance_hints.get("watch_review_candidates", [])[:20]
# Load rule-layer alias suggestions if available
rule_alias = []
if args.suggestions and args.suggestions.exists():
s = _load_json(args.suggestions)
rule_alias = s.get("alias_suggestions", [])
# Build prompt
prompt = _build_prompt(
interest_keywords=interest_keywords,
rule_alias_suggestions=rule_alias,
candidate_terms=candidate_terms,
relevant_watch_terms=relevant_watch_terms,
)
# Determine output path
suggestion_date = datetime.now(timezone.utc).date().isoformat()
output_path = args.output or (DEFAULT_OUTPUT_DIR / f"term-cleanup-semantic-suggestions-{suggestion_date}.json")
if args.dry_run:
print("=== DRY RUN: Prompt ===")
print(prompt)
print("\n=== END ===")
print(f"\nWould write to: {output_path}")
return
# Call LLM
print(f"Calling LLM ({model})...", file=sys.stderr)
response = _call_llm(prompt, api_url, model, api_key)
print(f"LLM response received ({len(response)} chars)", file=sys.stderr)
# Parse
parsed = _parse_llm_response(response)
# Build output
output = {
"date": suggestion_date,
"source_bundle": str(args.bundle),
"model": model,
"interest_keyword_count": len(interest_keywords),
"candidate_count": len(candidate_terms),
**parsed,
}
_save_json(output_path, output)
summary = {
"output": str(output_path),
"semantic_alias": len(output.get("semantic_alias", [])),
"stopword": len(output.get("stopword", [])),
"promote_to_interest": len(output.get("promote_to_interest", [])),
}
print(json.dumps(summary, ensure_ascii=False, indent=2))
if __name__ == "__main__":
main()
+128 -3
View File
@@ -175,6 +175,121 @@ def _prepare_watch_suggestions(bundle: dict[str, Any], reserved_terms: set[str])
return suggestions
def _prepare_alias_suggestions(
bundle: dict[str, Any],
all_terms: list[dict[str, Any]] | None = None,
) -> list[dict[str, Any]]:
"""
Generate alias suggestions using surface-form rules (no LLM).
Rules:
1. casefold match — same normalized form, different original casing
2. trailing-s singularization — singular/plural variants
3. whitespace/hyphen normalization — word boundary variants
Scans all_terms (full term_stats) if provided; otherwise falls back
to top_global_terms from the bundle.
"""
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
interest_keywords = _require_list(
current_config.get("interest_keywords"), "bundle.current_config.interest_keywords"
)
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
source_terms = all_terms if all_terms is not None else top_global_terms
interest_set = {_term_key(t) for t in interest_keywords if isinstance(t, str)}
interest_originals: set[str] = {t for t in interest_keywords if isinstance(t, str)}
# Build full casefold → [original forms] map
cf_map: dict[str, list[str]] = {}
for item in source_terms:
term = None
if isinstance(item, dict):
term = item.get("term")
elif isinstance(item, str):
term = item
if not isinstance(term, str) or not term.strip():
continue
key = _term_key(term)
if key not in cf_map:
cf_map[key] = []
if term not in cf_map[key]:
cf_map[key].append(term)
suggestions: list[dict[str, Any]] = []
seen_pairs: set[tuple[str, str]] = set()
def _add(from_term: str, to_term: str, reason: str) -> None:
pair = (_term_key(from_term), _term_key(to_term))
if pair in seen_pairs:
return
seen_pairs.add(pair)
suggestions.append({"from": from_term, "to": to_term, "reason": reason})
# Build a set of all term keys from source for quick lookup
source_keys = set(cf_map.keys())
# Rule 1: casefold match — same normalized form, different casing
for key, variants in cf_map.items():
if len(variants) < 2:
continue
canonical = None
alt_forms = []
for v in variants:
if v in interest_originals:
canonical = v
else:
alt_forms.append(v)
if canonical and alt_forms:
for alt in alt_forms:
_add(alt, canonical, "Case variant")
elif len(variants) >= 2 and not canonical:
# None is canonical — suggest the highest-frequency form
ranked = sorted(variants, key=lambda t: -(
next(
(it.get("total_count", 0) for it in top_global_terms if it.get("term") == t),
0,
)
))
for alt in ranked[1:]:
_add(alt, ranked[0], "Case variant (auto-ranked)")
# Rule 2: singular/plural — trailing-s normalization
# Check all source terms (not just interest keys) for bidirectional matching
for key in source_keys:
if key in interest_set:
continue
if key.endswith("s") and len(key) > 2:
singular_key = key.rstrip("s")
if singular_key in interest_set and singular_key != key:
# Find canonical interest keyword
canon = next((t for t in interest_keywords if _term_key(t) == singular_key), None)
from_form = cf_map[key][0]
if canon:
_add(from_form, canon, "Plural variant")
# singular form → interest has plural
plural_key = key + "s"
if plural_key in interest_set and plural_key != key:
canon = next((t for t in interest_keywords if _term_key(t) == plural_key), None)
from_form = cf_map[key][0]
if canon:
_add(from_form, canon, "Singular variant")
# Rule 3: whitespace/hyphen normalization
for key in source_keys:
if key in interest_set:
continue
normalized = key.replace("-", "").replace("_", "").replace(" ", "")
if normalized in interest_set and normalized != key:
canon = next((t for t in interest_keywords if _term_key(t) == normalized), None)
from_form = cf_map[key][0]
if canon:
_add(from_form, canon, "Whitespace/punctuation variant")
suggestions.sort(key=lambda x: (x["from"].casefold(), x["to"].casefold()))
return suggestions
def _render_table(items: list[dict[str, Any]]) -> str:
if not items:
return "_None in this pass._\n"
@@ -361,9 +476,19 @@ def main() -> None:
markdown_output=args.markdown_output,
)
# Load full term_stats for alias scanning (bundle only has top N)
stats_path = REPO_ROOT / "data" / "term_index" / "term_stats.json"
all_stats_terms: list[str] = []
if stats_path.exists():
stats_payload = _load_json(stats_path)
raw_terms = stats_payload.get("terms") if isinstance(stats_payload, dict) else []
if isinstance(raw_terms, list):
all_stats_terms = [str(t["term"]) for t in raw_terms if isinstance(t, dict) and isinstance(t.get("term"), str)]
interest_items = _prepare_interest_suggestions(bundle)
reserved_terms = {_term_key(str(item.get("term") or "")) for item in interest_items}
watch_items = _prepare_watch_suggestions(bundle, reserved_terms=reserved_terms)
alias_items = _prepare_alias_suggestions(bundle, all_terms=all_stats_terms)
suggestions = {
"date": suggestion_date,
@@ -373,10 +498,10 @@ def main() -> None:
"summary": {
"interest_keyword_suggestions": len(interest_items),
"watch_terms": len(watch_items),
"alias_suggestions": 0,
"alias_suggestions": len(alias_items),
"stopword_suggestions": 0,
},
"alias_suggestions": [],
"alias_suggestions": alias_items,
"stopword_suggestions": [],
"interest_keyword_suggestions": interest_items,
"watch_terms": watch_items,
@@ -400,7 +525,7 @@ def main() -> None:
"markdown_output": str(markdown_output_path) if args.emit_markdown else None,
"interest_keyword_suggestions": len(interest_items),
"watch_terms": len(watch_items),
"alias_suggestions": 0,
"alias_suggestions": len(alias_items),
"stopword_suggestions": 0,
"emit_markdown": args.emit_markdown,
}
+68 -24
View File
@@ -26,7 +26,7 @@ description: 生成 reader 项目的正式关键词 review 输入。当用户需
## 工作流程
1. 构建精简的审查数据包(临时工作文件):
### Phase 1:构建审查数据包
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py
@@ -34,51 +34,92 @@ python skills/keyword-cleanup-review/scripts/build_review_bundle.py
可选参数:
- `--days 7`
- `--top 50`
- `--days 7`(默认 7,建议传 365 覆盖全量)
- `--top 100`(考虑的词数)
- `--output outputs/term_index/review/keyword-cleanup-bundle.json`
2. 阅读建议模式:
#### 候选引擎策略
- `skills/keyword-cleanup-review/references/suggestion-schema.md`
根据 `configs/term_cleanup_policy.json` 的 `schema_version` 自动切换:
3. 运行 suggestions 生成脚本:
| 版本 | 策略 | 说明 |
|------|------|------|
| v1(旧) | 固定阈值(total≥3/days≥2 → interest) | 小数据集兼容 |
| v2(当前默认) | 百分位排名 + 增速因子 | 自适应数据量,不需要手工调阈值 |
v2 策略说明:
- **percentile**:total_count 在所有词里的排位占比。top 5% → interest 候选,5%-20% → watch 候选
- **growth**:recent_count / total_count,衡量近期活跃度。growth≥0.5 的排位外词也会主动推荐
### Phase 2:生成建议(规则层)
```bash
python scripts/generate_term_cleanup_suggestions.py ^
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
```
默认生成:
- 一份符合模式的 JSON 建议文件(正式建议产物,也是 review / apply 之间唯一正式输入)
如需人工审阅展示稿,再显式加:
如需人工审阅展示稿:
```bash
python scripts/generate_term_cleanup_suggestions.py ^
--bundle outputs/term_index/review/keyword-cleanup-bundle.json ^
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
这时才会额外生成:
#### 产出能力
- 一份简短的供人工审阅的 Markdown 报告(临时展示稿)
| 建议类型 | 状态 | 方法 |
|---------|------|------|
| interest 建议 | ✅ 已实现 | 百分位 top 5% + 增速促活 |
| watch 建议 | ✅ 已实现 | 百分位 5%-20% |
| alias 建议 | ✅ 已实现 | 规则层:大小写归一、单复数、去空格/连字符 |
| stopword 建议 | ❌ 规则层空缺 | 见 Phase 3(LLM 层) |
4. 严格保持边界:
默认生成:
- `term-cleanup-suggestions-YYYY-MM-DD.json`(正式建议产物)
显式加 `--emit-markdown` 额外生成:
- `term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿)
### Phase 3:生成建议(LLM 层,可选)
规则层覆盖不了 alias(中英文对应、缩写展开、同义不同名)和 stopword 判断,需要 LLM 辅助:
```bash
python scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
从 `.env` 读取 LLM 配置(`LLM_API_URL` / `LLM_MODEL` / `LLM_API_KEY`),使用 DeepSeek API。
输出三部分:
| 输出 | 说明 |
|------|------|
| `semantic_alias` | 语义级别名(中英文、缩写、同义不同名) |
| `stopword` | 泛词过滤建议(规则层做不了的需要语义判断的) |
| `promote_to_interest` | 与用户关注方向一致的新词,建议加入 interest |
**注:LLM 层产物是候选,不应自动 apply,需要人工确认后由 OpenClaw 编排 apply。**
### Phase 4:输出给 OpenClaw 编排
- `suggestions JSON` = review / apply 之间唯一正式建议输入
- `semantic-suggestions JSON` = LLM 补充建议,需要人工筛选后合并到 suggestions JSON 再 apply
- Markdown = 临时展示层
- 后续汇报、确认、dry-run、apply、收尾清理由 OpenClaw 编排层执行
### Phase 5:严格保持边界
- 建议 `configs/term_aliases.json` 的修改
- 建议 `configs/term_stopwords.json` 的修改
- 建议 `configs/filter_context.personal.json` 的新增
- **LLM 层产出(semantic-suggestions)不自动 apply**,需人工确认后由 OpenClaw 编排层执行
- 除非用户明确要求,否则不要直接编辑这些文件
- 除非用户要求修改规则逻辑,否则不要建议直接编辑 `configs/filter_rules.json`
5. 输出交接口径:
- 将 JSON suggestions 视为正式 review 输入
- 将 Markdown 视为可选展示层
- 后续汇报、确认、dry-run apply、正式 apply、收尾清理应由 OpenClaw 编排层继续执行
## 审查启发式规则
优先考虑以下决策:
@@ -141,6 +182,7 @@ JSON 输出应遵循:
短期保留:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
- `outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json`
临时产物:
@@ -173,7 +215,9 @@ JSON 输出应遵循:
## 资源
- 脚本:
- `scripts/build_review_bundle.py`
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- `scripts/generate_term_cleanup_suggestions.py`
- `scripts/generate_term_cleanup_semantic_suggestions.py`(LLM 层)
- 参考文档:
- `references/suggestion-schema.md`
- `plans/keyword-cleanup-interest-watch-engine-improvement.md`(v2 引擎设计)
@@ -9,16 +9,15 @@ from typing import Any
DEFAULT_POLICY: dict[str, Any] = {
"schema_version": "v1",
"schema_version": "v2",
"interest_keyword_review": {
"min_total_count": 3,
"min_days_seen": 2,
"percentile_min": 0.0,
"percentile_max": 0.05,
"growth_promotion": 0.5,
},
"watch_term_review": {
"min_total_count": 1,
"min_days_seen": 1,
"max_total_count": 2,
"max_days_seen": 2,
"percentile_min": 0.05,
"percentile_max": 0.20,
},
"alias_review": {
"min_total_count": 2,
@@ -28,6 +27,11 @@ DEFAULT_POLICY: dict[str, Any] = {
"max_total_count": 2,
"max_days_seen": 2,
},
"notes": [
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
"growth = recent_count / total_count,衡量近期活跃度",
],
}
@@ -126,6 +130,37 @@ def _within_watch_thresholds(item: dict[str, Any], thresholds: dict[str, Any]) -
)
def _compute_percentile(value: int, sorted_values: list[int]) -> float:
"""
Return the percentile rank of `value` in `sorted_values` (ascending).
0.0 = highest frequency (top rank), 1.0 = lowest frequency (bottom rank).
"""
if not sorted_values:
return 1.0
# bisect_left — count of values strictly less than `value`
lo, hi = 0, len(sorted_values)
while lo < hi:
mid = (lo + hi) // 2
if sorted_values[mid] < value:
lo = mid + 1
else:
hi = mid
rank = lo
# invert: smallest value → rank=0 → 1.0 (bottom)
# largest value → rank=len → 0.0 (top)
return 1.0 - (rank / len(sorted_values))
def _compute_growth(recent_count: int, total_count: int) -> float:
"""
Return growth factor: recent_count / total_count.
Only meaningful when total_count >= 3; returns 0.0 for small counts.
"""
if total_count < 3:
return 0.0
return recent_count / total_count
def main() -> None:
parser = argparse.ArgumentParser(
description="Build a compact review bundle for the keyword-cleanup-review skill."
@@ -235,6 +270,13 @@ def main() -> None:
alias_values = _casefold_set(list(aliases.values()))
watch_set = _casefold_set([str(item.get("term", "")) for item in watchlist])
# Build a sorted list of all total_counts for percentile computation
all_total_counts = sorted(
int(item.get("total_count") or 0)
for item in stats_terms
if isinstance(item, dict) and isinstance(item.get("term"), str)
)
top_global_terms = []
for item in stats_terms[: args.top]:
if not isinstance(item, dict):
@@ -256,42 +298,108 @@ def main() -> None:
"is_alias_target": folded in alias_values,
"in_watchlist": folded in watch_set,
"recent_count": recent_counter.get(term, 0),
"percentile": _compute_percentile(
int(item.get("total_count") or 0), all_total_counts
),
"growth": _compute_growth(
recent_counter.get(term, 0),
int(item.get("total_count") or 0),
),
}
)
# Keep more uncovered terms for percentile-based selection
uncovered_terms = [
item for item in top_global_terms if not item["in_interest_keywords"] and not item["is_stopword"]
][:20]
][:100]
policy_version = (policy.get("schema_version") if isinstance(policy, dict) else None) or "v1"
interest_thresholds = policy.get("interest_keyword_review") if isinstance(policy, dict) else {}
watch_thresholds = policy.get("watch_term_review") if isinstance(policy, dict) else {}
if policy_version == "v2" or "percentile_max" in interest_thresholds:
# v2: percentile + growth based selection
pct_min_interest = float(interest_thresholds.get("percentile_min", 0.0))
pct_max_interest = float(interest_thresholds.get("percentile_max", 0.05))
growth_promo = float(interest_thresholds.get("growth_promotion", 0.5))
pct_min_watch = float(watch_thresholds.get("percentile_min", 0.05))
pct_max_watch = float(watch_thresholds.get("percentile_max", 0.20))
interest_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"]
and pct_min_interest <= item["percentile"] <= pct_max_interest
]
watch_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"]
and pct_min_watch < item["percentile"] <= pct_max_watch
]
# Growth boost: terms outside watch range but with strong growth signal
growth_boost_candidates = [
item for item in uncovered_terms
if not item["in_watchlist"]
and item["percentile"] > pct_max_watch
and item["growth"] >= growth_promo
]
else:
# v1 fallback: fixed thresholds
interest_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
]
watch_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"]
and not _meets_min_thresholds(item, interest_thresholds)
and _within_watch_thresholds(item, watch_thresholds)
]
growth_boost_candidates = []
interest_review_candidates = [
{
"term": item["term"],
"total_count": item["total_count"],
"days_seen": item["days_seen"],
"percentile": item["percentile"],
"growth": item["growth"],
"reason": (
"Meets the configured interest-keyword review threshold and is not yet covered "
"by interest keywords or stopwords."
f"top {item['percentile']:.1%} by frequency,"
f"growth={item['growth']:.0%},"
"not yet covered by interest keywords or stopwords."
),
}
for item in uncovered_terms
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
for item in interest_candidates_raw
][:20]
watch_review_candidates = [
{
"term": item["term"],
"total_count": item["total_count"],
"days_seen": item["days_seen"],
"percentile": item["percentile"],
"growth": item["growth"],
"reason": (
"Falls into the configured watch-term review range and should be observed "
"before promotion into interest keywords."
f"top {item['percentile']:.1%} by frequency,"
f"growth={item['growth']:.0%},"
"fell into watch-review range."
),
}
for item in uncovered_terms
if not item["in_watchlist"]
and not _meets_min_thresholds(item, interest_thresholds)
and _within_watch_thresholds(item, watch_thresholds)
for item in watch_candidates_raw
][:20]
growth_boost_review_items = [
{
"term": item["term"],
"total_count": item["total_count"],
"days_seen": item["days_seen"],
"percentile": item["percentile"],
"growth": item["growth"],
"reason": (
f"growth spike: {item['growth']:.0%} of occurrences in recent window "
f"(total={item['total_count']}, days={item['days_seen']})."
),
}
for item in growth_boost_candidates
][:5]
recent_hot_terms = sorted(
({"term": term, "recent_count": count} for term, count in recent_counter.items()),
key=lambda item: (-item["recent_count"], item["term"].casefold(), item["term"]),
@@ -330,6 +438,7 @@ def main() -> None:
"governance_hints": {
"interest_review_candidates": interest_review_candidates,
"watch_review_candidates": watch_review_candidates,
"growth_boost_review_items": growth_boost_review_items,
},
}
_save_json(args.output, bundle)
+186
View File
@@ -0,0 +1,186 @@
---
name: reader-digest-flow
description: 编排 reader 项目的端到端 AI 日报流程。仅在用户要求运行/重跑日报、汇报候选、发布 Hugo 日报、沉淀选中文章或继续已有日报任务时使用;覆盖异步 MCP 任务、候选确认、发布、单篇摘要和 IMA 知识库上传。不要因验证、排障冲动或候选质量不佳自行重跑。
---
# Reader Digest Flow
## 职责边界
本 Skill 负责:
- 通过 reader MCP 启动、观察和恢复日报任务;
- 向用户展示候选并保持稳定编号;
- 根据用户选择生成并发布 Hugo 日报;
- 对用户选中的文章生成知识笔记并编排 IMA 上传;
- 在每个副作用边界执行确认和结果验证。
本 Skill 不负责:
- 实现 reader 内部抓取、摘要、过滤或恢复逻辑;
- 通过手拼目录推导 Run 状态或 Artifact;
- 未经用户要求自行重跑 Pipeline;
- 未经用户确认发布日报或写入知识库;
- 直接维护关键词配置;关键词治理委托给 `keyword-cleanup-review`。
## 核心规则
1. **只按用户指令运行。** 只有用户明确要求“跑日报”“重新跑”“再跑一次”时才启动新 Pipeline。验证、解释排序和排障默认读取已有 Run。
2. **一次对话绑定一个当前 Run。** 以异步 Job 结果返回的 `run_id` 为稳定句柄;新 Run 产生新的候选编号体系,不混用历史编号。
3. **状态以 MCP 返回为准。** Agent 只根据顶层 `status` 和 `recommended_action` 分支;`status_source`、`state_conflict` 仅用于解释。
4. **路径以返回值为准。** 使用 `output_dir`、`artifact.path`、`delivery_output`、`report_output` 和 `written_paths`;不要根据 `run_id` 手拼 `outputs/...`。
5. **候选编号保持稳定。** 用户编号永远对应当前候选列表的原始顺序(1-based);跨产物读取详情时按 URL 或完整 `item_id` 关联,不按数组位置关联。
6. **副作用必须授权。** 用户确认 Hugo 文章后才能发布;用户确认 IMA 文章后才能生成并上传知识笔记。
7. **内容必须有来源。** 日报和知识笔记只能基于当前 Run 的 `article.plain_text`、摘要、highlights 等 Artifact;不得使用通用知识补写原文没有的信息,也不为满足长度而扩写。
## 默认生产参数
用户未显式覆盖时使用:
```json
{
"limit": 7,
"include_read": false,
"mark_read": true,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
- 默认只处理未读文章。
- `include_read=true` 仅在用户明确要求扩大到已读内容时使用。
- debug/test/validation 才允许 `mark_read=false` 或 `debug_artifacts=true`。
- 不随机生成文章数量;用户指定 `limit` 时按用户值执行。
## 正式流程
### Phase 1:启动并观察日报 Job
正式生产入口统一为异步 MCP:
1. 调用 `start_freshrss_pipeline_job`;
2. 轮询 `get_freshrss_pipeline_job_status`;
3. `status=success` 后调用 `get_freshrss_pipeline_job_result`;
4. 保存返回的 `run_id` 和 Artifact 路径;
5. 使用 `get_run_status`、`get_delivery_payload`、`get_run_report` 读取业务状态和结果。
失败时:
1. 使用 `get_run_status(run_id)` 读取关联 Run;
2. 调用 `inspect_resume_plan(run_id)`;
3. `recommended_action=resume` 时启动并轮询异步 Resume Job;
4. `recommended_action=read_terminal_result` 时直接读取已有终态结果;
5. `recommended_action=start_new_run` 时停止并向用户报告,不自行新建 Run。
CLI 仅用于 MCP 不可用时的 fallback、debug 或人工排障,不是默认生产入口。具体调用序列见 `references/flow.md`。
### Phase 2:汇报候选
- 使用当前 Run 返回的 Delivery Payload 或 digest brief Artifact;
- 按候选原始顺序从 1 编号,状态可显示为“已入选/待确认”,但不得重新分组编号;
- 每篇提供标题、来源、2-3 句摘要和筛选理由,避免原始 JSON dump;
- 用户质疑编号或排序时读取当前 Run 产物核对,不重新运行 Pipeline;
- 需要跨 Artifact 取详情时按 URL 或完整 `item_id` 交叉验证。
Feishu 输出不要使用 Markdown 表格,见 `references/feishu-format-notes.md`。
### Phase 3:等待 Hugo 选择
- 等待用户明确选择要发布的文章;
- 用户编号映射到当前候选列表,不映射到 extracted 文件序号;
- 用户拒绝发布时立即停止当日日报后续流程,不劝说、不自动换一批;
- 用户明确要求重跑时才创建新 Run,并重新建立编号体系。
### Phase 4:生成并发布 Hugo 日报
发布前读取 `references/public-digest-example.md`,按其最终页面结构生成:
- `今日概览`
- `今日重点`
- `趋势观察`
每篇 `今日重点` 文章末尾必须添加 `来源:[来源名](原文 URL)`,来源链接跟随对应文章,不再生成独立的 `延伸阅读` 章节或重复链接。
仅发布用户在 Phase 3 选中的文章。公开页面不得出现 `keep/review/drop`、候选、待确认等内部状态。
写入 Hugo 后执行部署,并验证首页、日报列表页和当日详情页均可访问。命令和检查项见 `references/flow.md`。
### Phase 5:等待 IMA 选择
Hugo 发布选择与知识沉淀选择相互独立。询问用户哪些文章值得长期保存:
- 仅处理用户明确选择的文章;
- 不把整份日报上传到 IMA;
- 本次选择本身即授权后续单篇摘要和 IMA 上传,不重复确认。
### Phase 6:生成单篇知识笔记
对每篇选中文章:
1. 通过 URL/完整 `item_id` 找到对应 extracted Artifact;
2. 使用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
3. 轮询 `get_article_summary_job_status`,成功后读取 `get_article_summary_job_result`;
4. 只使用现有 `article.plain_text`,不重新抓取原 URL;
5. 使用返回的 `written_paths` 定位结果并检查 Markdown 内容。
6. **`extracted_path` 必须传绝对路径**(前缀 `/home/ubuntu/zhu/github/reader/`):summary-mcp 工作目录是 `/root/.hermes`,相对路径会报 `extracted_path does not exist`。同一工具连续 3 次失败会触发 MCP 冷却(约 45-60s,报 `MCP server 'reader' is unreachable`),等待冷却后再重试,不要循环重试同一调用。
异步 MCP 不可用时才使用项目 CLI fallback。不要因为内容较短而引入原文之外的知识。
### Phase 7:上传到 IMA
上传前按需读取:
- 格式规则:`references/ima-format-quickref.md`
- API 和上传步骤:`references/ima-upload-api.md`(含 `-200` 版本拦截修复)
- 凭证定位:`references/ima-credential-chain.md`
- 批量上传脚本:`scripts/ima_upload_one.py`(Python 编排,规避中文文件名 bash 引号问题)
硬规则:
- 使用完整文章标题作为文件名;
- 以 Markdown 文件 `media_type=7` 上传到 `daily` knowledge base;
- 保留原文 URL 和 Category,保证来源可追踪;
- 不使用 URL 导入或 Notes 类型替代知识库文件;
- 上传后验证目标知识库中存在对应条目;
- 失败时报告具体阶段,不无限重试。
## 关键词治理路由
只有用户明确要求“清理关键词”“词库治理”等操作时才触发关键词治理。Review bundle 和 Suggestions 生成委托给 `keyword-cleanup-review`,确认与 Apply 仍由当前编排层负责:
1. 生成 review bundle;
2. 生成规则层与可选语义层 Suggestions JSON;
3. 等待人工确认;
4. dry-run 后按精确 accept 参数 Apply。
`reader-digest-flow` 不直接编辑 `term_aliases.json`、`term_stopwords.json` 或兴趣配置。简要路由见 `references/keyword-engine-maintenance.md`。
## 环境坑位(本机部署)
- **MCP 相对路径陷阱**:`summary-mcp` 服务工作目录是 `/root/.hermes`,不是 reader 项目根。MCP 返回的 `output_dir`/`artifact.path` 是相对路径,直接传给 `start_article_summary_job(extracted_path=...)` 会报 `extracted_path does not exist`。传入前必须拼绝对路径前缀 `/home/ubuntu/zhu/github/reader/`。
- **提取失败不等于运行失败**:`status_counts.extract_failed` 的条目(`CONTENT_EXTRACTION_FAILED`,`retryable=false`)跳过即可并如实汇报;失败文章常是推广/活动等低价值内容,不因此自行重跑。`linked_run_status=partial` 时先读 run-report 的 item 级 `error` 确认原因。
- **用户要求"重新跑一批"**:候选质量低(用户主动提出)时重跑,应 `include_read=true` 并调高 `limit`(如 10),否则默认 `include_read=false` 会拉回同一批未读文章。重跑是新 Run,候选编号体系重新建立,汇报时提醒用户按新列表选择。
## 停止与人工介入
出现以下任一情况时停止自动流程并报告:
- 用户没有授权运行、发布或知识库写入;
- Job/Run 返回不可恢复,或连续恢复失败;
- Payload、候选 ID 或 Artifact 之间无法可靠关联;
- 生成内容缺少可追踪来源;
- Hugo 部署验证失败;
- IMA 凭证、目标知识库或上传结果无法验证。
## Reference 路由
- `references/flow.md`:具体 MCP 调用序列、候选映射(含 extracted_path 绝对路径、候选≠文件名顺序)、Hugo 发布和 IMA 主步骤。
- `references/content-extraction.md`:FreshRSS 内容来源与 `plain_text` 质量判断。
- `references/public-digest-example.md`:可直接参考的 Hugo 最终页面结构。
- `references/feishu-format-notes.md`:Feishu 输出格式限制。
- `references/ima-format-quickref.md`:IMA Markdown 格式规则。
- `references/ima-upload-api.md`:IMA Markdown 文件上传 API(含 `-200` 版本拦截修复)。
- `references/ima-credential-chain.md`:IMA 凭证与知识库配置定位。
- `references/keyword-engine-maintenance.md`:关键词治理 Skill 路由。
- `scripts/ima_upload_one.py`:单篇 Markdown 上传 daily 知识库的完整 Python 脚本(preflight→重名→create_media→COS→add_knowledge)。
@@ -0,0 +1,45 @@
# 内容提取流程
本文说明处理流水线如何把 FreshRSS 条目转换为可供摘要使用的文章文本。
## 核心规则:FreshRSS 条目不重新抓取原文 URL
**FreshRSS 是仅提供 RSS 内容的上游。** 对于 FreshRSS 条目,流水线不会向文章原始 URL 发起 HTTP 请求。该行为由 `pipeline.py` 中的 `RSS_ONLY_UPSTREAMS = {"freshrss"}` 强制保证。
唯一例外是非 FreshRSS 上游。未来未设置 `upstream: freshrss` 的其他来源,可以在必要时使用 `fetch_html()` 作为回退。
## 内容来源优先级
`content_loader.py` 按以下顺序检查内容,并使用第一个包含 **至少 500 个可读字符** 的来源:
| 优先级 | 来源 | 含义 |
|--------|------|------|
| 1 | `raw_html` | 通过 `ExtractionInput.raw_html` 预先注入的 HTML;常规 FreshRSS 运行中很少使用。 |
| 2 | `item.raw_content` | RSS `<content:encoded>` 中的文章正文;部分订阅源提供,部分不提供。 |
| 3 | `item.raw_summary` | RSS `<description>` 中的摘要或片段;这是当前运行中最常见的来源。 |
| 4 | `rss_content` | 来自非条目字段的独立 RSS 内容。 |
| — | `none` | 没有可用内容;FreshRSS 不允许回源抓取,因此抛出 `RSS_CONTENT_MISSING`。 |
## `content_source` 与文本质量的关系
每个 `item-XX.extracted.json` 中的 `content_source` 字段表示流水线实际使用的内容来源:
- **`item.raw_content`**:RSS `<content:encoded>` 提供的文章正文,通常质量最好,接近直接阅读原文。
- **`item.raw_summary`**:只有 RSS 摘要或描述,并非完整正文。不同来源长度差异较大,通常为 300-2000 个字符;AI 摘要基于该片段,而不是完整文章。
- **`rss_content`**:来自独立 RSS 内容,质量取决于订阅源。
- **`fetched_html`**:从原始 URL 抓取的 HTML。FreshRSS 条目不会出现该来源,只适用于非 FreshRSS 上游。
## 对日报质量的影响
如果提取结果文件中出现 `content_source: item.raw_summary`,说明 AI 使用的是订阅源摘要或片段,而不是完整正文。日报内容显得较浅时,原因可能只是 RSS 描述过短。
提高质量可以选择提供完整 `<content:encoded>` 的订阅源,或者把内容来源切换到支持全文 RSS 的系统,例如具备全文提取能力的 RSS 代理或 FiveFilters 等服务。
## 快速检查
先调用 `list_run_artifacts(run_id)`,再读取返回的提取结果产物路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。
## 相关代码路径
- `src/summary_mcp/core/pipeline.py`:`RSS_ONLY_UPSTREAMS`、`_should_skip_fetch()`、`extract_content()`。
- `src/summary_mcp/core/content_loader.py`:`choose_inline_content()` 的优先级链和 `fetch_html()`;FreshRSS 条目不会调用后者。
@@ -0,0 +1,17 @@
# 飞书 Markdown 格式说明
## 背景
Hermes 的飞书网关(`gateway/platforms/feishu.py`)通过 `_build_outbound_payload` 发送消息。该方法会检查内容中的 Markdown 特征,并据此决定消息类型:
- 内容匹配 `_MARKDOWN_HINT_RE`(加粗、列表、代码、链接等)时,使用包含 `md` 元素的飞书 `post` 类型发送,可以正常渲染。
- 内容匹配 `_MARKDOWN_TABLE_RE`(Markdown 表头和分隔行)时,整条消息会被强制转换为 `text` 类型,即纯文本,不再渲染 Markdown。
原因是 `_build_markdown_post_payload` 会把内容包装为 `{"tag": "md", "text": "..."}` 元素,而飞书的 `md` 元素不支持表格,也没有把 Markdown 表格转换为飞书原生表格的逻辑。
## 飞书输出规则
- 通过飞书发送的消息不得使用 Markdown 表格;消息中只要出现一个表格,整条消息就会退化为纯文本。
- 需要表达结构化信息时,优先使用分点列表、带标题的分节或行内格式。
- 加粗(`**加粗**`)、行内代码(`` `代码` ``)、无序列表(`- 项目`)、有序列表(`1. 项目`)和链接均可正常使用。
- 围栏式代码块可以使用,但代码块后的尾随内容可能存在渲染边界问题。
@@ -0,0 +1,177 @@
# Reader Digest Flow 操作参考
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
## 1. 日报 Pipeline
### 默认参数
```json
{
"limit": 7,
"include_read": false,
"mark_read": true,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
### 正式调用序列
```text
start_freshrss_pipeline_job
→ get_freshrss_pipeline_job_status
→ get_freshrss_pipeline_job_result
→ get_run_status
→ get_delivery_payload / get_run_report
```
状态动作:
- `running`:按合理间隔继续轮询;
- `success`:读取结果,保存 `run_id`;
- `failed`:读取关联 Run 并执行 Resume Plan;
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
恢复序列:
```text
inspect_resume_plan
→ recommended_action=resume
→ start_resume_job
→ get_resume_job_status
→ get_resume_job_result
```
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
## 2. 候选汇报与选择
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
展示规则:
1. 使用候选数组原始顺序并从 1 编号;
2. 不因 `keep/review` 分组而重新编号;
3. 每篇展示标题、来源、摘要和判断理由;
4. 用户编号只映射当前候选数组;
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
### extracted 文件与候选编号错位(实测 2026-07-31)
- `digest-brief.json` 的 `top_candidates` **没有 `item_id` 字段**,只有 `url`;
- `extracted/item-XX.extracted.json` 的文件名序号与候选编号**可能不一致**(实例:候选2 = item-05、候选3 = item-02);
- 正确做法:用 **URL 交叉匹配**(归一化 `%3D`→`=` 后逐条比对),或用完整 `item_id`(从 candidate-batch.json 的 `items[i].item_id` 按候选数组顺序取)在 extracted 文件里反查;两者都能验证时优先 item_id。
## 3. Hugo 日报
用户确认发布文章后:
1. 读取 `public-digest-example.md`;
2. 仅使用用户选中的文章生成公开内容;
3. 写入 Hugo 当日页面;
4. 前台执行部署,避免把构建日志作为聊天通知;
5. 验证首页、列表页和详情页。
当前部署位置:
```text
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
```
验证至少覆盖:
```text
http://127.0.0.1:14322/
http://127.0.0.1:14322/daily/
http://127.0.0.1:14322/daily/YYYY-MM-DD/
```
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
## 4. 单篇知识笔记
用户确认 IMA 文章后:
1. 从候选中取得 URL 和完整 `item_id`;
2. 从 Run Artifact 中找到匹配的 extracted 文件;
3. 交叉验证 `article.item_id` 或 URL;
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
6. 从 Job 结果的 `written_paths` 读取 Markdown。
### ⚠️ extracted_path 必须用绝对路径
`summary-mcp` 进程的工作目录是 `/root/.hermes`(不是 reader 项目根)。传相对路径(如 `outputs/freshrss/...`)会直接报 `extracted_path does not exist`。必须传绝对路径:
```text
/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json
```
### ⚠️ 候选编号 ≠ extracted 文件名顺序
候选数组顺序与 extracted 文件名(`item-01`…`item-07`)**不一定对齐**(实测候选2 落在 item-05)。`digest-brief.json` 的候选**没有 `item_id` 字段**,只有 URL。可靠匹配方法:
1. 从 `candidate-batch.json` 取每项完整 `item_id`(在 `candidate` 嵌套对象里,顶层 `item_key` 只是 `item-XX` 文件名序号);
2. 或按 URL 匹配:归一化(`%3D`→`=`)后与每个 extracted 文件的 `article.url` / `article.canonical_url` 比对;
3. 绝不要按候选位置对应 extracted 文件序号。
```python
def norm(u): return u.replace('%3D','=').replace('%3d','=').strip()
# 对每个 extracted 文件取 norm(article.url),与候选 norm(url) 精确比对
```
正式序列:
```text
start_article_summary_job
→ get_article_summary_job_status
→ get_article_summary_job_result
```
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
CLI 仅在异步 MCP 不可用或人工排障时使用:
```bash
python scripts/run_article_summaries.py \
--extracted <returned-extracted-path> \
--ids <full-item-id> \
--output-dir <explicit-output-dir>
```
## 5. IMA 上传
用户在知识沉淀阶段的文章选择即为上传授权。
执行顺序:
1. 检查生成的 Markdown 与来源;
2. 文件名规范化为 `<完整文章标题>.md`;
3. 确认目标为 `daily` knowledge base;
4. 执行 preflight、create_media、COS upload、add_knowledge;
5. 验证知识库条目存在。
上传格式与 API 参数分别见:
- `ima-format-quickref.md`
- `ima-upload-api.md`
- `ima-credential-chain.md`
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
## 6. CLI fallback 原则
CLI 仅在以下场景使用:
- MCP 服务不可用;
- Tool transport/launch 失败且无法取得有效 Job;
- 用户明确要求本地调试;
- 人工排障需要直接检查脚本输出。
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。
@@ -0,0 +1,27 @@
# IMA 凭证与安全边界
## 必需配置
- `IMA_OPENAPI_CLIENTID`
- `IMA_OPENAPI_APIKEY`
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
- `IMA_DAILY_KNOWLEDGE_BASE_NAME=daily`
优先使用当前进程环境和 IMA Skill 已支持的凭证加载机制。不要在本 Skill 中复制、迁移或重写密钥文件。
## 缺失处理
Preflight 返回凭证缺失或目标知识库无法解析时:
1. 停止上传;
2. 只报告缺失的变量名或配置项;
3. 等待用户或运行环境补齐配置;
4. 配置恢复后重新执行 preflight,不重复生成知识笔记。
## 安全边界
- 不在聊天、日志或命令输出中打印完整 API key、KB ID 或 COS 临时凭证;
- `create_media` 返回的 COS 凭证仅在同一受控进程内传给上传工具,不写入磁盘;
- 不通过拼接 Shell 字符串传递凭证,使用参数数组或 IMA Skill 的封装;
- 不绕过 Hermes 的脱敏机制;若现有工具链无法安全传递凭证,停止并报告;
- 上传结束后不持久化 COS 临时凭证。
@@ -0,0 +1,47 @@
# IMA Markdown 格式速查
## 文件与标题
- 文件名:`<完整文章标题>.md`
- `add_knowledge.title`:完整文章标题,不包含 `.md`
- 上传类型:Markdown 文件,`media_type=7`
- 目标:`daily` knowledge base
## 内容来源
只能使用当前 Run 的可追踪内容:
1. extracted Artifact 的 `article.plain_text`;
2. 对应文章的结构化摘要;
3. digest brief 的 summary 与 highlights。
不得使用通用知识补写原文没有的信息,不设置固定字数或字节数门槛。内容较短时保持简洁并忠于来源。
## 标准结构
```markdown
# 完整文章标题
Source: https://原文链接
Category: 分类
## 核心结论
## 主要论点
## 关键方法 / 机制
## 重要细节
## 可复用启发
## 关键词
## 主题
```
- 核心结论和主要论点使用连贯段落;
- 方法、细节和启发按完整知识点分项;
- 没有来源支持的 Section 可以简写,不得编造内容填充。
上传 API 见 `ima-upload-api.md`。
@@ -0,0 +1,126 @@
# IMA Markdown 上传 API
用于将用户选中的单篇 Markdown 知识笔记上传到 `daily` knowledge base。
## 凭证
- `IMA_OPENAPI_CLIENTID`
- `IMA_OPENAPI_APIKEY`
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
凭证定位和恢复见 `ima-credential-chain.md`。不要在终端输出完整密钥。
## 上传前检查
- 文件名为 `<完整文章标题>.md`;
- `title` 为完整文章标题,不带 `.md`;
- Markdown 符合 `ima-format-quickref.md`;
- 内容可追溯到当前 Run Artifact;
- 用户已经明确选择该文章;
- 目标知识库已经解析并验证。
## 1. Preflight
调用 IMA Skill 的 `preflight-check.cjs` 检查文件类型、扩展名、大小和 MIME。
预期:
```text
file_ext=md
content_type=text/markdown
media_type=7
```
### ⚠️ IMA skill 版本拦截(-200)
`ima_api.cjs` 每天首次调用会检查更新,若检测到新版(如 1.1.8 > 当前 1.1.7)会以 `code=-200` 拦截原请求。注意:**官方 zip 包内的 `meta.json` 可能没同步版本号**(下载 1.1.8 zip 后 meta 仍写 1.1.7),所以光替换文件无法跳过拦截。
快速修复(脚本本身已是新版,只差版本号):
```bash
cd /root/.hermes/skills/openclaw-imports/ima-skill && python3 -c "
import json
m = json.load(open('meta.json')); m['version'] = '1.1.8'
json.dump(m, open('meta.json','w'), ensure_ascii=False, indent=2)
"
```
先用 `diff -rq` 对比 zip 与安装目录:若只有 `.DS_Store`/meta 差异,说明代码已是最新,直接改 meta.json 版本号即可;若脚本有实质差异才需要整体替换。
## 2. Create Media
```text
POST /openapi/wiki/v1/create_media
```
请求核心字段:
```json
{
"file_name": "<完整文章标题>.md",
"file_size": 0,
"content_type": "text/markdown",
"knowledge_base_id": "<daily-kb-id>",
"file_ext": "md"
}
```
保存返回的 `media_id` 和 `cos_credential`。COS 临时凭证只在进程内传递,不打印到聊天或日志。
## 3. COS Upload
使用 IMA Skill 提供的 `cos-upload.cjs`,通过参数数组调用并检查:
- 进程 `returncode`;
- `stderr`;
- HTTP 上传结果。
不要拼接包含凭证的 Shell 字符串,也不要把多条 JSON 响应重定向到同一个文件。
## 4. Add Knowledge
```text
POST /openapi/wiki/v1/add_knowledge
```
核心字段:
```json
{
"media_type": 7,
"media_id": "<media-id>",
"title": "<完整文章标题>",
"knowledge_base_id": "<daily-kb-id>",
"file_info": {
"cos_key": "<cos-key>",
"file_size": 0,
"file_name": "<完整文章标题>.md"
}
}
```
## 5. 验证
上传成功不能只依据本地命令退出码。验证目标 knowledge base 中存在对应标题或返回对象,并向用户报告最终结果。
失败时明确报告发生在 Preflight、Create Media、COS Upload 或 Add Knowledge 的哪一步,不改用 URL 导入或 Notes 类型绕过错误。
## 6. 已知坑:IMA skill 版本拦截(-200)
`ima_api.cjs` 每天首次调用会检查远端版本,若发现新版本(如 1.1.8 > 1.1.7)会以 exit=1 + stderr `{"code":-200}` 拦截**所有** API 调用,原请求不发送。此前遇到过。
处理方式(不必整包替换):
1. 按 stderr 提示下载新版 zip(如 `https://app-dl.ima.qq.com/skills/ima-skills-1.1.8.zip`)并解压;
2. 对比新旧 `ima_api.cjs` 的 md5——zip 内核心脚本常与本地一致,只是 `meta.json` 的 `version` 未同步(zip 内仍写 1.1.7);
3. 若 `ima_api.cjs` 一致,只需把本地 `meta.json` 的 `version` 改为远端版本号即可跳过拦截,无需替换文件。
调用成功后再执行本文件前面的上传流程。
## 6. 版本拦截与批量上传实测(2026-07-31)
- **`-200` skill 更新拦截**:`ima_api.cjs` 每天首次调用检查版本,发现新版时以 code -200 退出并提示更新。下载 zip 后**先对比 `ima_api.cjs` 的 md5**——实测 zip 内脚本与已装版本完全一致,只是 `meta.json` 版本号未同步。此时只需把 `~/.hermes/skills/openclaw-imports/ima-skill/meta.json` 的 `version` 改为最新版即可跳过拦截,无需替换任何脚本。
- **Python 脚本编排上传**比 bash 可靠:bash 拼接含中文文件名/凭证的 curl 易出错。用 `subprocess` 参数数组依次调 `preflight-check.cjs` → `ima_api.cjs check_repeated_names` → `create_media` → `cos-upload.cjs`(`--secret-id/--secret-key/--token` 走参数数组,不打印)→ `add_knowledge`,每步解析返回 JSON,失败即停。
- **批量上传**:4 篇逐个跑同一脚本即可;同名文件先 `check_repeated_names` 确认无重复。
- 凭证从 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` 环境变量读取(`ima_api.cjs` 自动加载),KB ID 用 `IMA_DAILY_KNOWLEDGE_BASE_ID`。
@@ -0,0 +1,29 @@
# 关键词治理路由
关键词治理不属于 `reader-digest-flow` 的日常执行阶段。
仅当用户明确要求“清理关键词”“词库治理”“生成关键词建议”时,委托:
```text
skills/keyword-cleanup-review/SKILL.md
```
Review 输入生成由该 Skill 定义,后续确认与 Apply 由当前编排层负责:
```text
build review bundle
→ generate rule suggestions
→ optional semantic suggestions
→ human review
→ dry-run
→ apply accepted suggestions
```
约束:
- Suggestions JSON 是 Review 与 Apply 之间的正式契约;
- LLM 语义建议不能自动 Apply;
- 不直接编辑 aliases、stopwords、watchlist 或 interest 配置;
- 不在日报主流程中因 tag 质量不佳自动触发治理。
Review bundle、Suggestions、Schema 和产物保留策略以 `keyword-cleanup-review` 为唯一事实来源;该 Skill 不直接 Apply 配置。
@@ -0,0 +1,31 @@
+++
title = "AI 日报 · 示例"
date = 2026-04-01T09:00:00+08:00
summary = "围绕 Agent 架构分层、Skills 标准化与工程实践的当日观察。"
+++
# 今日概览
今天的公开内容主要集中在 AI Agent 架构演进、工具化落地与工程实践三条线索。行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
## 今日重点
### 1. 从 Agent 到 Skills:AI 智能体架构的范式转变
文章分析了 AI 智能体从单体 Agent 向模块化 Skills 的演进,并结合 MCP、Skills 和真实项目说明能力分层与复用方式。
值得关注:
- Skills 将领域流程从 Agent 主体中拆出,便于复用和维护。
- MCP 为 Agent 与外部工具提供标准化连接方式。
- 工程竞争点逐渐从模型调用转向状态、工具和工作流设计。
这篇内容值得关注的原因在于,它把开放协议、分层架构和真实落地案例连接成了完整论证链。
来源:[示例来源](https://example.com/a)
## 趋势观察
1. Agent 正在从单体能力转向可组合的模块化体系。
2. 工具契约、状态管理和验证机制正在成为 AI 应用的核心工程能力。
3. Human-in-the-loop 仍是控制高风险副作用的重要边界。
@@ -0,0 +1,96 @@
#!/usr/bin/env python3
"""IMA 上传单篇 Markdown 知识笔记到 daily 知识库。
用法: python3 ima_upload_one.py "<绝对路径/summary.md>" "<完整文章标题>"
依赖环境变量: IMA_OPENAPI_CLIENTID / IMA_OPENAPI_APIKEY / IMA_DAILY_KNOWLEDGE_BASE_ID
流程: preflight -> check_repeated_names -> create_media -> cos-upload -> add_knowledge
说明: 用 Python 而非 bash 编排,避免中文文件名/引号转义问题。
退出码 2 = 文件名重复(需与用户确认保留双方或取消),非 0 均为失败。
"""
import json, os, subprocess, sys
SKILL_DIR = "/root/.hermes/skills/openclaw-imports/ima-skill"
IMA_API = os.path.join(SKILL_DIR, "ima_api.cjs")
COS_UPLOAD = os.path.join(SKILL_DIR, "knowledge-base/scripts/cos-upload.cjs")
PREFLIGHT = os.path.join(SKILL_DIR, "knowledge-base/scripts/preflight-check.cjs")
def run_node(script, args):
r = subprocess.run(["node", script] + args, capture_output=True, text=True, timeout=120)
if r.returncode != 0:
raise RuntimeError(f"{script} exit={r.returncode} stderr={r.stderr[:500]}")
return json.loads(r.stdout)
def ima_api(api_path, body):
r = subprocess.run(["node", IMA_API, api_path, json.dumps(body, ensure_ascii=False)],
capture_output=True, text=True, timeout=120)
if r.returncode != 0:
raise RuntimeError(f"ima_api {api_path} exit={r.returncode} stderr={r.stderr[:500]}")
resp = json.loads(r.stdout)
if resp.get("code") != 0:
raise RuntimeError(f"ima_api {api_path} code={resp.get('code')} msg={resp.get('msg')}")
return resp.get("data", {})
def main():
kb_id = os.environ["IMA_DAILY_KNOWLEDGE_BASE_ID"]
file_path = sys.argv[1]
title = sys.argv[2] # 完整文章标题(不含 .md)
pf = run_node(PREFLIGHT, ["--file", file_path])
if not pf.get("pass"):
raise RuntimeError(f"preflight failed: {pf}")
file_name = pf["file_name"]; media_type = pf["media_type"]
content_type = pf["content_type"]; file_size = pf["file_size"]; file_ext = pf["file_ext"]
print(f"[preflight] ok file={file_name} ext={file_ext} size={file_size} media_type={media_type}")
dup = ima_api("openapi/wiki/v1/check_repeated_names", {
"params": [{"name": file_name, "media_type": media_type}],
"knowledge_base_id": kb_id
})
is_rep = dup.get("results", [{}])[0].get("is_repeated", False) if dup.get("results") else False
if is_rep:
print(f"[check_repeated_names] REPEATED: {file_name} — 需要处理")
sys.exit(2)
print("[check_repeated_names] no duplicate")
cm = ima_api("openapi/wiki/v1/create_media", {
"file_name": file_name,
"file_size": file_size,
"content_type": content_type,
"knowledge_base_id": kb_id,
"file_ext": file_ext
})
media_id = cm["media_id"]
cos = cm["cos_credential"]
print(f"[create_media] media_id={media_id} cos_bucket={cos.get('bucket_name')} cos_key={cos.get('cos_key','')[:40]}")
r = subprocess.run(["node", COS_UPLOAD,
"--file", file_path,
"--secret-id", cos["secret_id"],
"--secret-key", cos["secret_key"],
"--token", cos["token"],
"--bucket", cos["bucket_name"],
"--region", cos["region"],
"--cos-key", cos["cos_key"],
"--content-type", content_type,
"--start-time", str(cos.get("start_time", "")),
"--expired-time", str(cos.get("expired_time", "")),
"--timeout", "300000"
], capture_output=True, text=True, timeout=360)
if r.returncode != 0:
raise RuntimeError(f"cos-upload exit={r.returncode} stderr={r.stderr[:800]}")
print(f"[cos-upload] ok rc=0 stdout={r.stdout.strip()[:200]}")
ak = ima_api("openapi/wiki/v1/add_knowledge", {
"media_type": media_type,
"media_id": media_id,
"title": title,
"knowledge_base_id": kb_id,
"file_info": {
"cos_key": cos["cos_key"],
"file_size": file_size,
"file_name": file_name
}
})
print(f"[add_knowledge] ok media_id={ak.get('media_id') or media_id}")
if __name__ == "__main__":
main()
+24 -2
View File
@@ -6,8 +6,30 @@ from .freshrss_pipeline_jobs import (
start_freshrss_pipeline_job,
)
from .query_service import get_delivery_payload, get_run_report, get_run_status, list_run_artifacts, list_runs
from .resume_jobs import get_resume_job_result, get_resume_job_status, start_resume_job
from .resume_service import inspect_resume_plan, resume_run
# NOTE: resume_jobs and resume_service are NOT eagerly imported here to avoid
# a circular import chain:
# workflows/freshrss_pipeline.py -> runtime -> resume_jobs -> resume_service
# -> workflows/freshrss_pipeline.py (circular!)
# They are lazy-loaded via __getattr__ when accessed as summary_mcp.runtime.*
def __getattr__(name):
import importlib
_LAZY = {
"get_resume_job_result": ("resume_jobs", "get_resume_job_result"),
"get_resume_job_status": ("resume_jobs", "get_resume_job_status"),
"start_resume_job": ("resume_jobs", "start_resume_job"),
"inspect_resume_plan": ("resume_service", "inspect_resume_plan"),
"resume_run": ("resume_service", "resume_run"),
}
if name in _LAZY:
mod_name, attr_name = _LAZY[name]
mod = importlib.import_module(f".{mod_name}", __package__)
return getattr(mod, attr_name)
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
__all__ = [
"ArtifactRecord",
@@ -9,6 +9,8 @@ from pathlib import Path
from typing import Any
from uuid import uuid4
from dotenv import dotenv_values
from .run_store import RunStore
from .query_service import _resolve_run_record
@@ -29,6 +31,25 @@ DEFAULT_STAGES = [
]
MIN_JOB_STALE_SECONDS = 30 * 60
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
def _build_subprocess_env() -> dict[str, str]:
"""Build an env dict for subprocess, merging parent env with .env values.
The subprocess inherits the Hermes MCP server's environment, but .env values
may not be in os.environ at the time the subprocess is spawned. This function
loads them from .env and merges them in so the child process sees all needed
variables (LLM_API_KEY, LLM_MODEL, LLM_API_URL, FRESHRSS_*, etc.) directly
in os.environ, avoiding any dotenv-loading timing issues inside the subprocess.
"""
env = os.environ.copy()
if DEFAULT_DOTENV_PATH.exists():
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items():
if isinstance(key, str) and isinstance(value, str) and value:
# Only set if not already present in parent env
env.setdefault(key, value)
return env
def _now() -> datetime:
@@ -218,6 +239,7 @@ def start_freshrss_pipeline_job(
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
start_new_session=True,
env=_build_subprocess_env(),
)
except Exception as exc:
report_file = _write_job_report(
+19
View File
@@ -9,6 +9,8 @@ from pathlib import Path
from typing import Any
from uuid import uuid4
from dotenv import dotenv_values
from .query_service import _resolve_run_record
from .resume_service import (
SUPPORTED_RESUME_STAGES,
@@ -38,6 +40,22 @@ DEFAULT_STAGES = [
]
MIN_JOB_STALE_SECONDS = 30 * 60
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
def _build_subprocess_env() -> dict[str, str]:
"""Build an env dict for subprocess, merging parent env with .env values.
Ensures the subprocess sees all needed variables (LLM_API_KEY, LLM_MODEL,
LLM_API_URL, FRESHRSS_*, etc.) directly in os.environ, avoiding dotenv
timing issues in the child process.
"""
env = os.environ.copy()
if DEFAULT_DOTENV_PATH.exists():
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items():
if isinstance(key, str) and isinstance(value, str) and value:
env.setdefault(key, value)
return env
def _now() -> datetime:
@@ -276,6 +294,7 @@ def start_resume_job(*, run_id: str) -> dict[str, Any]:
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
start_new_session=True,
env=_build_subprocess_env(),
)
except Exception as exc:
report_file = _write_job_report(
+46 -32
View File
@@ -2,6 +2,7 @@ from __future__ import annotations
import json
import os
from concurrent.futures import ThreadPoolExecutor, as_completed
from datetime import date, datetime, timezone
from pathlib import Path
from typing import Any
@@ -192,7 +193,7 @@ def _build_candidate_batch_payload(*, run_id: str, item_contexts: list[dict[str,
}
def _persist_summary_batch_artifact(*, run_store: RunStore, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
def _persist_summary_batch_artifact(*, run_store, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
output_path = _summary_batch_output(run_dir)
_save_json(
output_path,
@@ -202,7 +203,7 @@ def _persist_summary_batch_artifact(*, run_store: RunStore, run_dir: Path, item_
return output_path
def _persist_candidate_batch_artifact(*, run_store: RunStore, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
def _persist_candidate_batch_artifact(*, run_store, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
output_path = _candidate_batch_output(run_dir)
_save_json(
output_path,
@@ -461,37 +462,50 @@ def run_freshrss_pipeline(
summary_success_count = 0
summary_failed_count = 0
summary_candidates = [ctx for ctx in item_contexts if ctx["extraction"] is not None and ctx["extraction"].success]
for item_context in summary_candidates:
item_report = item_context["item_report"]
summary_exit_code, summary_payload, summary_report = run_loop_payload(
extracted_payload=item_context["extracted_payload"],
prompt_path=resolved_prompt_path,
output_path=item_context["summary_output"],
max_retries=max_retries,
timeout_seconds=timeout_seconds,
api_key=resolved_llm_api_key,
model=resolved_llm_model,
api_url=resolved_llm_api_url,
)
if summary_exit_code != 0 or summary_payload is None:
item_report["status"] = "summary_failed"
if summary_report is not None:
item_report["summary_errors"] = summary_report.errors
summary_failed_count += 1
else:
item_context["summary_payload"] = summary_payload
item_report["status"] = "summarized"
summary_success_count += 1
# Parallelize LLM summaries — I/O bound calls, independent per article
with ThreadPoolExecutor(max_workers=min(len(summary_candidates) or 1, 4)) as pool:
fut_map = {}
for item_context in summary_candidates:
fut = pool.submit(
run_loop_payload,
extracted_payload=item_context["extracted_payload"],
prompt_path=resolved_prompt_path,
output_path=item_context["summary_output"],
max_retries=max_retries,
timeout_seconds=timeout_seconds,
api_key=resolved_llm_api_key,
model=resolved_llm_model,
api_url=resolved_llm_api_url,
)
fut_map[fut] = item_context
run_store.update_stage(
SUMMARY_STAGE,
outputs={
"expected_items": extracted_success_count,
"completed_items": summary_success_count + summary_failed_count,
"success_count": summary_success_count,
"failed_count": summary_failed_count,
},
)
for fut in as_completed(fut_map):
item_context = fut_map[fut]
item_report = item_context["item_report"]
try:
summary_exit_code, summary_payload, summary_report = fut.result()
except Exception as exc:
summary_exit_code, summary_payload, summary_report = 1, None, None
if summary_exit_code != 0 or summary_payload is None:
item_report["status"] = "summary_failed"
if summary_report is not None:
item_report["summary_errors"] = summary_report.errors
summary_failed_count += 1
else:
item_context["summary_payload"] = summary_payload
item_report["status"] = "summarized"
summary_success_count += 1
run_store.update_stage(
SUMMARY_STAGE,
outputs={
"expected_items": extracted_success_count,
"completed_items": summary_success_count + summary_failed_count,
"success_count": summary_success_count,
"failed_count": summary_failed_count,
},
)
summary_batch_output = _persist_summary_batch_artifact(
run_store=run_store,