Files
reader/docs/design/keyword-cleanup-flow-overview.md
T
root 590d050218 keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions
- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
2026-05-14 17:17:49 +08:00

5.0 KiB
Raw Blame History

关键词清洗流程概述

2026-05-14 初版 从"数据记录"到"人工确认落盘"的完整链路


整体数据流

每日日报 pipeline
     │
     ▼
term_index/daily/YYYY-MM-DD.json    ← 每天一篇候选文章的热词统计
     │
     ▼
term_index/term_stats.json          ← 所有 daily 的汇总(1070 个词)
     │
     ├──── build_review_bundle.py    ← 打包为审查数据包
     │         │
     │         ▼
     │    review/keyword-cleanup-bundle.json
     │         │
     │         ▼
     │    generate_term_cleanup_suggestions.py
     │         │
     │         ▼
     │    review/term-cleanup-suggestions-YYYY-MM-DD.json  ← 正式建议产物
     │         │
     │         ▼
     │    (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md  ← 展示稿
     │
     ├──── 人工确认哪些建议 accept
     │
     ▼
apply_term_suggestions.py           ← 写入配置
     │
     ├── configs/filter_context.personal.json  ← interest_keywords
     ├── configs/term_aliases.json             ← alias
     ├── configs/term_stopwords.json           ← stopword
     ├── configs/term_watchlist.json           ← watch
     └── configs/term_change_log.json          ← 变更日志

各环节说明

阶段 1:数据记录(每日自动)

# FreshRSS pipeline 跑完后自动产出
data/term_index/daily/2026-05-14.json
  • 每天一篇,记录当天候选文章中出现的热词
  • 包含 term、total_count、days_seen 等信息
  • 目前累计 41 天,共 1070 个独立词

阶段 2:全量汇总(每日自动)

data/term_index/term_stats.json
  • 从所有 daily 文件重建,会覆盖重跑
  • 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13)

阶段 3:构建审查数据包(手动触发)

python skills/keyword-cleanup-review/scripts/build_review_bundle.py \
  --days 365 \
  --top 100 \
  --output outputs/term_index/review/keyword-cleanup-bundle.json
  • 把 term_stats + 当前配置打成一包,方便后续处理
  • 输出:review/keyword-cleanup-bundle.json

阶段 4:生成建议(手动触发)

python scripts/generate_term_cleanup_suggestions.py \
  --bundle outputs/term_index/review/keyword-cleanup-bundle.json \
  --emit-markdown

当前产出能力

建议类型 状态 当前阈值 说明
interest_keyword_suggestions ✅ 已实现 total≥3, days≥2 产出 20 条
watch_terms ✅ 已实现 total≤2, days≤2 本次 0 条
alias_suggestions ❌ 硬编码为空 policy 有阈值(total≥2, days≥2)但脚本未实现
stopword_suggestions ❌ 硬编码为空 policy 有阈值(total≤2, days≤2)但脚本未实现

关键发现: alias 和 stopword 不是"阈值太保守",是 generate 脚本里压根没写对应的生成函数。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 [] 和 0。

阶段 5:人工确认(手动)

OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply

本次模式:

  • 高频(≥5次/5天以上)→ 强烈推荐 ✅
  • 中频(3-4次)→ 附带建议 ✅
  • 泛词 → 建议跳过 ❌

阶段 6:落盘配置(手动)

python scripts/apply_term_suggestions.py \
  --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
  --accept-interest 词1 词2 ...
  • dry-run 预览 → 确认后正式 apply
  • 写入 configs/filter_context.personal.json
  • 同步记录到 term_change_log.json
  • 不备份原始配置(待优化)
  • apply 后不自动清理 review 目录(待优化)

阶段 7:维护清理(按需)

由 OpenClaw 侧 reader-keyword-maintenance skill 处理:

  • 删除旧 markdown 展示稿
  • 保留最近一份 bundle
  • 保守保留 suggestions JSON

当前配置资产

文件 内容 数据量
filter_context.personal.json interest_keywords 52 个
term_aliases.json 别名映射 0 组(未启用)
term_stopwords.json 停用词 0 个(未启用)
term_watchlist.json 观察词 6 个
term_change_log.json 所有变更记录 已记录

待优化项

  1. alias/stopword 建议生成为空 — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数
  2. apply 前无配置备份 — 建议 apply 前自动 cp 备份
  3. apply 后无自动收尾 — 建议 apply 后自动删旧 markdown 和 bundle
  4. alias 识别依赖规则而非 LLM — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply