Files
reader/skills/reader-digest-flow/references/keyword-engine-maintenance.md
T
root 5eb390e3ed docs: 添加 Agent Skill 到项目仓库
- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/)
- README 新增 Agent Skill 章节说明供 Agent 使用的工作流
2026-07-28 18:38:54 +08:00

7.7 KiB
Raw Blame History

关键词引擎维护流程

概述

reader 项目内置了完整的关键词清洗链路,用于管理 term_aliases.json、term_stopwords.json、filter_context.personal.json 等配置。本文件说明链路各环节的职责与使用方式。

完整数据流

每日日报 pipeline
     │
     ▼
term_index/daily/YYYY-MM-DD.json    ← build_keyword_index.py (pipeline 末环节自动)
     │
     ▼
term_index/term_stats.json          ← 全量汇总
     │
     ▼
build_review_bundle.py              ← 打包审查数据包 (手动触发)
     │
     ▼
generate_term_cleanup_suggestions.py ← 生成建议 (手动触发)
     │
     ▼
apply_term_suggestions.py           ← 人工确认后落盘 (手动触发)

各环节命令

1. 重建每日 term_index(pipeline 末环节自动执行,也可手动)

cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/build_keyword_index.py \
  --input outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json

2. 重建 review bundle(打包当前配置+统计供审查)

cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
  --days 365 \
  --top 200 \
  --output outputs/term_index/review/keyword-cleanup-bundle.json

参数:

  • --days 365 — 考虑最近多少天的统计数据
  • --top 200 — 取前 N 个高频词纳入 bundle

3. 生成清洗建议

cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \
  --bundle outputs/term_index/review/keyword-cleanup-bundle.json \
  --emit-markdown

输出:

  • outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json — 正式建议产物
  • outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md — 人工阅读展示稿

当前脚本能生成的建议类型:

类型 生成规则 当前产出
interest_keyword_suggestions 高频未覆盖词(total≥3, days≥2) ✅ 按阈值产出
watch_terms 中频观察词(total≤2, days≤2) ✅ 按阈值产出
alias_suggestions 大小写变体/单复数/空格连词符差异 ✅ 统计规则产出
stopword_suggestions — ❌ 脚本硬编码为空

4. 应用建议(dry-run → review → apply)

# 先 dry-run 预览
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/apply_term_suggestions.py \
  --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
  --accept-interest 强化学习 ReAct CLI \
  --dry-run

# 确认后正式 apply(去掉 --dry-run)
.venv/bin/python3 scripts/apply_term_suggestions.py \
  --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
  --accept-interest 强化学习 ReAct CLI \

支持的 accept 参数:

  • --accept-interest 词1 词2 ... — 添加到 interest_keywords
  • --accept-watch 词1 词2 ... — 添加到 watchlist
  • --accept-alias 源词1 源词2 ... — 添加 alias 映射
  • --accept-stopword 词1 词2 ... — 添加停用词

脚本能做什么 vs 不能做什么

✅ 脚本能做的(统计级清洗)

  • 发现大小写变体(vibe coding → Vibe Coding)
  • 发现单复数差异(Agent Skill → Agent Skills)
  • 发现空格/连词符差异
  • 按频次推荐 should-be-interest / should-be-watch 的词
  • 批量 apply 到配置文件,自动记录变更日志

❌ 脚本不能做的(语义级清洗,需要人工判断)

  • 识别产品名(WorkBuddy、LibTV Agent、飞书妙搭)→ 应加 stopword
  • 识别模型名(Qwen3-30B-A3B、GLM 5.2、Opus 4.8)→ 应加 stopword
  • 识别人名/地名(Andrej Karpathy、成都天府长岛)→ 应加 stopword
  • 英文专有名词归一化到中文主题词(Multi-Agent → 多Agent)
  • 长产品名归一化(火山云数据库PostgreSQL Serverless版 → Serverless数据库)
  • 判断某个词对 Hugo tag 是否有用(1688、OPC训练营 → 无效)

语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py)

除了统计规则的清洗,项目还有一个 LLM 驱动的语义级清洗脚本,能发现统计规则做不到的事情:

cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \
  --bundle outputs/term_index/review/keyword-cleanup-bundle.json \
  --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
  --output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json

参数:

  • --bundle — review bundle(先跑 build_review_bundle.py)
  • --suggestions — 规则层的建议 JSON(可选,供 LLM 避免重复)
  • --output — 输出路径(自动生成)
  • --dry-run — 只打印 prompt 不调 LLM

语义脚本能做的(而统计规则不能做的)

类型 LLM 能发现什么 示例
semantic_alias 中英文映射、简称↔全称、近义词 AI智能体→AI Agent, AI Harness→Harness Engineering
stopword 语义判断哪些词太泛/不相干 自动化(太泛), AlphaFold(生物领域), 陌生化(穿透词)
promote_to_interest 与用户关注领域对齐的新词 强化学习, Loop Engineering, MoE, ChatGPT

两阶段清洗 SOP

当需要清洗关键词时,按此顺序操作:

阶段一:规则统计清洗(脚本发现 + 人工确认)

  1. build_review_bundle.py → 重建 bundle
  2. generate_term_cleanup_suggestions.py → 产出统计级建议
  3. 检查建议,决定哪些 accept
  4. apply_term_suggestions.py --dry-run → 预览
  5. apply_term_suggestions.py → 正式落地

阶段二:语义级清洗(LLM 发现 + 人工确认)

  1. generate_term_cleanup_semantic_suggestions.py → 调 LLM 产生语义建议
  2. 检查 LLM 的 semantic_alias / stopword / promote_to_interest
  3. 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化)

🔴 不要绕过确认流程直接改配置。 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。

完整的语义级清洗操作流程

当需要大量添加 aliases/stopwords 时,推荐流程:

  1. 跑一次 build_review_bundle.py + generate_term_cleanup_suggestions.py 看看统计建议
  2. 走一遍实际 pipeline 产出,收集所有 unique keywords:
    for rid in $(ls outputs/freshrss/rerun/); do
      python3 -c "
    import json
    d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json'))
    for item in d['items']:
        for kw in item['summary'].get('keywords', []):
            print(kw)
    "
    done | sort -u
    
  3. 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的
  4. 分别更新 term_aliases.json 和 term_stopwords.json
  5. 重建 term_index:对每个受影响日期的 run 跑 build_keyword_index.py
  6. 更新 Hugo tags:从重建后的 data/term_index/daily/YYYY-MM-DD.json 取

当前配置 (2026-07-16)

文件 条目数
configs/term_aliases.json 142
configs/term_stopwords.json 106
configs/filter_context.personal.json 54 (interest_keywords)
configs/term_watchlist.json 6

关键词/别名/停用词配置的更新规范

  • term_aliases.json — 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式
  • term_stopwords.json — 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词
  • 随时可以加,加了后重建 term_index 即可生效
  • 不涉及 pipeline 重新跑——只影响下游展示