# 关键词引擎维护流程 ## 概述 reader 项目内置了完整的关键词清洗链路,用于管理 `term_aliases.json`、`term_stopwords.json`、`filter_context.personal.json` 等配置。本文件说明链路各环节的职责与使用方式。 ## 完整数据流 ``` 每日日报 pipeline │ ▼ term_index/daily/YYYY-MM-DD.json ← build_keyword_index.py (pipeline 末环节自动) │ ▼ term_index/term_stats.json ← 全量汇总 │ ▼ build_review_bundle.py ← 打包审查数据包 (手动触发) │ ▼ generate_term_cleanup_suggestions.py ← 生成建议 (手动触发) │ ▼ apply_term_suggestions.py ← 人工确认后落盘 (手动触发) ``` ## 各环节命令 ### 1. 重建每日 term_index(pipeline 末环节自动执行,也可手动) ```bash cd /home/ubuntu/zhu/github/reader .venv/bin/python3 scripts/build_keyword_index.py \ --input outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json ``` ### 2. 重建 review bundle(打包当前配置+统计供审查) ```bash cd /home/ubuntu/zhu/github/reader .venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \ --days 365 \ --top 200 \ --output outputs/term_index/review/keyword-cleanup-bundle.json ``` 参数: - `--days 365` — 考虑最近多少天的统计数据 - `--top 200` — 取前 N 个高频词纳入 bundle ### 3. 生成清洗建议 ```bash cd /home/ubuntu/zhu/github/reader .venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \ --bundle outputs/term_index/review/keyword-cleanup-bundle.json \ --emit-markdown ``` 输出: - `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json` — 正式建议产物 - `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md` — 人工阅读展示稿 当前脚本能生成的建议类型: | 类型 | 生成规则 | 当前产出 | |:----|:---------|:---------| | interest_keyword_suggestions | 高频未覆盖词(total≥3, days≥2) | ✅ 按阈值产出 | | watch_terms | 中频观察词(total≤2, days≤2) | ✅ 按阈值产出 | | alias_suggestions | 大小写变体/单复数/空格连词符差异 | ✅ 统计规则产出 | | stopword_suggestions | — | ❌ 脚本硬编码为空 | ### 4. 应用建议(dry-run → review → apply) ```bash # 先 dry-run 预览 cd /home/ubuntu/zhu/github/reader .venv/bin/python3 scripts/apply_term_suggestions.py \ --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \ --accept-interest 强化学习 ReAct CLI \ --dry-run # 确认后正式 apply(去掉 --dry-run) .venv/bin/python3 scripts/apply_term_suggestions.py \ --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \ --accept-interest 强化学习 ReAct CLI \ ``` 支持的 accept 参数: - `--accept-interest 词1 词2 ...` — 添加到 interest_keywords - `--accept-watch 词1 词2 ...` — 添加到 watchlist - `--accept-alias 源词1 源词2 ...` — 添加 alias 映射 - `--accept-stopword 词1 词2 ...` — 添加停用词 ## 脚本能做什么 vs 不能做什么 ### ✅ 脚本能做的(统计级清洗) - 发现大小写变体(`vibe coding` → `Vibe Coding`) - 发现单复数差异(`Agent Skill` → `Agent Skills`) - 发现空格/连词符差异 - 按频次推荐 should-be-interest / should-be-watch 的词 - 批量 apply 到配置文件,自动记录变更日志 ### ❌ 脚本不能做的(语义级清洗,需要人工判断) - 识别产品名(`WorkBuddy`、`LibTV Agent`、`飞书妙搭`)→ 应加 stopword - 识别模型名(`Qwen3-30B-A3B`、`GLM 5.2`、`Opus 4.8`)→ 应加 stopword - 识别人名/地名(`Andrej Karpathy`、`成都天府长岛`)→ 应加 stopword - 英文专有名词归一化到中文主题词(`Multi-Agent` → `多Agent`) - 长产品名归一化(`火山云数据库PostgreSQL Serverless版` → `Serverless数据库`) - 判断某个词对 Hugo tag 是否有用(`1688`、`OPC训练营` → 无效) ## 语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py) 除了统计规则的清洗,项目还有一个 **LLM 驱动的语义级清洗脚本**,能发现统计规则做不到的事情: ```bash cd /home/ubuntu/zhu/github/reader .venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \ --bundle outputs/term_index/review/keyword-cleanup-bundle.json \ --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \ --output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json ``` 参数: - `--bundle` — review bundle(先跑 build_review_bundle.py) - `--suggestions` — 规则层的建议 JSON(可选,供 LLM 避免重复) - `--output` — 输出路径(自动生成) - `--dry-run` — 只打印 prompt 不调 LLM ### 语义脚本能做的(而统计规则不能做的) | 类型 | LLM 能发现什么 | 示例 | |:----|:--------------|:-----| | **semantic_alias** | 中英文映射、简称↔全称、近义词 | `AI智能体`→`AI Agent`, `AI Harness`→`Harness Engineering` | | **stopword** | 语义判断哪些词太泛/不相干 | `自动化`(太泛), `AlphaFold`(生物领域), `陌生化`(穿透词) | | **promote_to_interest** | 与用户关注领域对齐的新词 | `强化学习`, `Loop Engineering`, `MoE`, `ChatGPT` | ### 两阶段清洗 SOP 当需要清洗关键词时,按此顺序操作: **阶段一:规则统计清洗(脚本发现 + 人工确认)** 1. `build_review_bundle.py` → 重建 bundle 2. `generate_term_cleanup_suggestions.py` → 产出统计级建议 3. 检查建议,决定哪些 accept 4. `apply_term_suggestions.py --dry-run` → 预览 5. `apply_term_suggestions.py` → 正式落地 **阶段二:语义级清洗(LLM 发现 + 人工确认)** 1. `generate_term_cleanup_semantic_suggestions.py` → 调 LLM 产生语义建议 2. 检查 LLM 的 semantic_alias / stopword / promote_to_interest 3. 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化) **🔴 不要绕过确认流程直接改配置。** 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。 ## 完整的语义级清洗操作流程 当需要大量添加 aliases/stopwords 时,推荐流程: 1. **跑一次 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py`** 看看统计建议 2. **走一遍实际 pipeline 产出**,收集所有 unique keywords: ```bash for rid in $(ls outputs/freshrss/rerun/); do python3 -c " import json d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json')) for item in d['items']: for kw in item['summary'].get('keywords', []): print(kw) " done | sort -u ``` 3. 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的 4. 分别更新 `term_aliases.json` 和 `term_stopwords.json` 5. **重建 term_index**:对每个受影响日期的 run 跑 `build_keyword_index.py` 6. **更新 Hugo tags**:从重建后的 `data/term_index/daily/YYYY-MM-DD.json` 取 ## 当前配置 (2026-07-16) | 文件 | 条目数 | |:----|:------| | `configs/term_aliases.json` | 142 | | `configs/term_stopwords.json` | 106 | | `configs/filter_context.personal.json` | 54 (interest_keywords) | | `configs/term_watchlist.json` | 6 | ## 关键词/别名/停用词配置的更新规范 - `term_aliases.json` — 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式 - `term_stopwords.json` — 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词 - **随时可以加**,加了后重建 term_index 即可生效 - 不涉及 pipeline 重新跑——只影响下游展示