- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/) - README 新增 Agent Skill 章节说明供 Agent 使用的工作流
7.7 KiB
7.7 KiB
关键词引擎维护流程
概述
reader 项目内置了完整的关键词清洗链路,用于管理 term_aliases.json、term_stopwords.json、filter_context.personal.json 等配置。本文件说明链路各环节的职责与使用方式。
完整数据流
每日日报 pipeline
│
▼
term_index/daily/YYYY-MM-DD.json ← build_keyword_index.py (pipeline 末环节自动)
│
▼
term_index/term_stats.json ← 全量汇总
│
▼
build_review_bundle.py ← 打包审查数据包 (手动触发)
│
▼
generate_term_cleanup_suggestions.py ← 生成建议 (手动触发)
│
▼
apply_term_suggestions.py ← 人工确认后落盘 (手动触发)
各环节命令
1. 重建每日 term_index(pipeline 末环节自动执行,也可手动)
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/build_keyword_index.py \
--input outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
2. 重建 review bundle(打包当前配置+统计供审查)
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 \
--top 200 \
--output outputs/term_index/review/keyword-cleanup-bundle.json
参数:
--days 365— 考虑最近多少天的统计数据--top 200— 取前 N 个高频词纳入 bundle
3. 生成清洗建议
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
输出:
outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json— 正式建议产物outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md— 人工阅读展示稿
当前脚本能生成的建议类型:
| 类型 | 生成规则 | 当前产出 |
|---|---|---|
| interest_keyword_suggestions | 高频未覆盖词(total≥3, days≥2) | ✅ 按阈值产出 |
| watch_terms | 中频观察词(total≤2, days≤2) | ✅ 按阈值产出 |
| alias_suggestions | 大小写变体/单复数/空格连词符差异 | ✅ 统计规则产出 |
| stopword_suggestions | — | ❌ 脚本硬编码为空 |
4. 应用建议(dry-run → review → apply)
# 先 dry-run 预览
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
--dry-run
# 确认后正式 apply(去掉 --dry-run)
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
支持的 accept 参数:
--accept-interest 词1 词2 ...— 添加到 interest_keywords--accept-watch 词1 词2 ...— 添加到 watchlist--accept-alias 源词1 源词2 ...— 添加 alias 映射--accept-stopword 词1 词2 ...— 添加停用词
脚本能做什么 vs 不能做什么
✅ 脚本能做的(统计级清洗)
- 发现大小写变体(
vibe coding→Vibe Coding) - 发现单复数差异(
Agent Skill→Agent Skills) - 发现空格/连词符差异
- 按频次推荐 should-be-interest / should-be-watch 的词
- 批量 apply 到配置文件,自动记录变更日志
❌ 脚本不能做的(语义级清洗,需要人工判断)
- 识别产品名(
WorkBuddy、LibTV Agent、飞书妙搭)→ 应加 stopword - 识别模型名(
Qwen3-30B-A3B、GLM 5.2、Opus 4.8)→ 应加 stopword - 识别人名/地名(
Andrej Karpathy、成都天府长岛)→ 应加 stopword - 英文专有名词归一化到中文主题词(
Multi-Agent→多Agent) - 长产品名归一化(
火山云数据库PostgreSQL Serverless版→Serverless数据库) - 判断某个词对 Hugo tag 是否有用(
1688、OPC训练营→ 无效)
语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py)
除了统计规则的清洗,项目还有一个 LLM 驱动的语义级清洗脚本,能发现统计规则做不到的事情:
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
参数:
--bundle— review bundle(先跑 build_review_bundle.py)--suggestions— 规则层的建议 JSON(可选,供 LLM 避免重复)--output— 输出路径(自动生成)--dry-run— 只打印 prompt 不调 LLM
语义脚本能做的(而统计规则不能做的)
| 类型 | LLM 能发现什么 | 示例 |
|---|---|---|
| semantic_alias | 中英文映射、简称↔全称、近义词 | AI智能体→AI Agent, AI Harness→Harness Engineering |
| stopword | 语义判断哪些词太泛/不相干 | 自动化(太泛), AlphaFold(生物领域), 陌生化(穿透词) |
| promote_to_interest | 与用户关注领域对齐的新词 | 强化学习, Loop Engineering, MoE, ChatGPT |
两阶段清洗 SOP
当需要清洗关键词时,按此顺序操作:
阶段一:规则统计清洗(脚本发现 + 人工确认)
build_review_bundle.py→ 重建 bundlegenerate_term_cleanup_suggestions.py→ 产出统计级建议- 检查建议,决定哪些 accept
apply_term_suggestions.py --dry-run→ 预览apply_term_suggestions.py→ 正式落地
阶段二:语义级清洗(LLM 发现 + 人工确认)
generate_term_cleanup_semantic_suggestions.py→ 调 LLM 产生语义建议- 检查 LLM 的 semantic_alias / stopword / promote_to_interest
- 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化)
🔴 不要绕过确认流程直接改配置。 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。
完整的语义级清洗操作流程
当需要大量添加 aliases/stopwords 时,推荐流程:
- 跑一次
build_review_bundle.py+generate_term_cleanup_suggestions.py看看统计建议 - 走一遍实际 pipeline 产出,收集所有 unique keywords:
for rid in $(ls outputs/freshrss/rerun/); do python3 -c " import json d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json')) for item in d['items']: for kw in item['summary'].get('keywords', []): print(kw) " done | sort -u - 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的
- 分别更新
term_aliases.json和term_stopwords.json - 重建 term_index:对每个受影响日期的 run 跑
build_keyword_index.py - 更新 Hugo tags:从重建后的
data/term_index/daily/YYYY-MM-DD.json取
当前配置 (2026-07-16)
| 文件 | 条目数 |
|---|---|
configs/term_aliases.json |
142 |
configs/term_stopwords.json |
106 |
configs/filter_context.personal.json |
54 (interest_keywords) |
configs/term_watchlist.json |
6 |
关键词/别名/停用词配置的更新规范
term_aliases.json— 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式term_stopwords.json— 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词- 随时可以加,加了后重建 term_index 即可生效
- 不涉及 pipeline 重新跑——只影响下游展示