# 关键词清洗流程概述 > 2026-05-14 初版 > 从"数据记录"到"人工确认落盘"的完整链路 --- ## 整体数据流 ``` 每日日报 pipeline │ ▼ term_index/daily/YYYY-MM-DD.json ← 每天一篇候选文章的热词统计 │ ▼ term_index/term_stats.json ← 所有 daily 的汇总(1070 个词) │ ├──── build_review_bundle.py ← 打包为审查数据包 │ │ │ ▼ │ review/keyword-cleanup-bundle.json │ │ │ ▼ │ generate_term_cleanup_suggestions.py │ │ │ ▼ │ review/term-cleanup-suggestions-YYYY-MM-DD.json ← 正式建议产物 │ │ │ ▼ │ (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md ← 展示稿 │ ├──── 人工确认哪些建议 accept │ ▼ apply_term_suggestions.py ← 写入配置 │ ├── configs/filter_context.personal.json ← interest_keywords ├── configs/term_aliases.json ← alias ├── configs/term_stopwords.json ← stopword ├── configs/term_watchlist.json ← watch └── configs/term_change_log.json ← 变更日志 ``` --- ## 各环节说明 ### 阶段 1:数据记录(每日自动) ```bash # FreshRSS pipeline 跑完后自动产出 data/term_index/daily/2026-05-14.json ``` - 每天一篇,记录当天候选文章中出现的热词 - 包含 term、total_count、days_seen 等信息 - 目前累计 **41 天**,共 **1070 个独立词** ### 阶段 2:全量汇总(每日自动) ```bash data/term_index/term_stats.json ``` - 从所有 daily 文件重建,会覆盖重跑 - 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13) ### 阶段 3:构建审查数据包(手动触发) ```bash python skills/keyword-cleanup-review/scripts/build_review_bundle.py \ --days 365 \ --top 100 \ --output outputs/term_index/review/keyword-cleanup-bundle.json ``` - 把 term_stats + 当前配置打成一包,方便后续处理 - 输出:`review/keyword-cleanup-bundle.json` ### 阶段 4:生成建议(手动触发) ```bash python scripts/generate_term_cleanup_suggestions.py \ --bundle outputs/term_index/review/keyword-cleanup-bundle.json \ --emit-markdown ``` #### 当前产出能力 | 建议类型 | 状态 | 当前阈值 | 说明 | |---------|------|----------|------| | interest_keyword_suggestions | ✅ **已实现** | total≥3, days≥2 | 产出 20 条 | | watch_terms | ✅ **已实现** | total≤2, days≤2 | 本次 0 条 | | alias_suggestions | ❌ **硬编码为空** | policy 有阈值(total≥2, days≥2)但脚本未实现 | | | stopword_suggestions | ❌ **硬编码为空** | policy 有阈值(total≤2, days≤2)但脚本未实现 | | **关键发现:** alias 和 stopword 不是"阈值太保守",是 **generate 脚本里压根没写对应的生成函数**。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 `[]` 和 `0`。 ### 阶段 5:人工确认(手动) ``` OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply ``` 本次模式: - 高频(≥5次/5天以上)→ 强烈推荐 ✅ - 中频(3-4次)→ 附带建议 ✅ - 泛词 → 建议跳过 ❌ ### 阶段 6:落盘配置(手动) ```bash python scripts/apply_term_suggestions.py \ --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \ --accept-interest 词1 词2 ... ``` - dry-run 预览 → 确认后正式 apply - 写入 `configs/filter_context.personal.json` - 同步记录到 `term_change_log.json` - **不备份原始配置**(待优化) - **apply 后不自动清理 review 目录**(待优化) ### 阶段 7:维护清理(按需) 由 OpenClaw 侧 `reader-keyword-maintenance` skill 处理: - 删除旧 markdown 展示稿 - 保留最近一份 bundle - 保守保留 suggestions JSON --- ## 当前配置资产 | 文件 | 内容 | 数据量 | |------|------|--------| | `filter_context.personal.json` | interest_keywords | 52 个 | | `term_aliases.json` | 别名映射 | 0 组(未启用) | | `term_stopwords.json` | 停用词 | 0 个(未启用) | | `term_watchlist.json` | 观察词 | 6 个 | | `term_change_log.json` | 所有变更记录 | 已记录 | --- ## 待优化项 1. **alias/stopword 建议生成为空** — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数 2. **apply 前无配置备份** — 建议 apply 前自动 cp 备份 3. **apply 后无自动收尾** — 建议 apply 后自动删旧 markdown 和 bundle 4. **alias 识别依赖规则而非 LLM** — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply