- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/) - README 新增 Agent Skill 章节说明供 Agent 使用的工作流
196 lines
7.7 KiB
Markdown
196 lines
7.7 KiB
Markdown
# 关键词引擎维护流程
|
||
|
||
## 概述
|
||
|
||
reader 项目内置了完整的关键词清洗链路,用于管理 `term_aliases.json`、`term_stopwords.json`、`filter_context.personal.json` 等配置。本文件说明链路各环节的职责与使用方式。
|
||
|
||
## 完整数据流
|
||
|
||
```
|
||
每日日报 pipeline
|
||
│
|
||
▼
|
||
term_index/daily/YYYY-MM-DD.json ← build_keyword_index.py (pipeline 末环节自动)
|
||
│
|
||
▼
|
||
term_index/term_stats.json ← 全量汇总
|
||
│
|
||
▼
|
||
build_review_bundle.py ← 打包审查数据包 (手动触发)
|
||
│
|
||
▼
|
||
generate_term_cleanup_suggestions.py ← 生成建议 (手动触发)
|
||
│
|
||
▼
|
||
apply_term_suggestions.py ← 人工确认后落盘 (手动触发)
|
||
```
|
||
|
||
## 各环节命令
|
||
|
||
### 1. 重建每日 term_index(pipeline 末环节自动执行,也可手动)
|
||
|
||
```bash
|
||
cd /home/ubuntu/zhu/github/reader
|
||
.venv/bin/python3 scripts/build_keyword_index.py \
|
||
--input outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
||
```
|
||
|
||
### 2. 重建 review bundle(打包当前配置+统计供审查)
|
||
|
||
```bash
|
||
cd /home/ubuntu/zhu/github/reader
|
||
.venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||
--days 365 \
|
||
--top 200 \
|
||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||
```
|
||
|
||
参数:
|
||
- `--days 365` — 考虑最近多少天的统计数据
|
||
- `--top 200` — 取前 N 个高频词纳入 bundle
|
||
|
||
### 3. 生成清洗建议
|
||
|
||
```bash
|
||
cd /home/ubuntu/zhu/github/reader
|
||
.venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \
|
||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||
--emit-markdown
|
||
```
|
||
|
||
输出:
|
||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json` — 正式建议产物
|
||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md` — 人工阅读展示稿
|
||
|
||
当前脚本能生成的建议类型:
|
||
|
||
| 类型 | 生成规则 | 当前产出 |
|
||
|:----|:---------|:---------|
|
||
| interest_keyword_suggestions | 高频未覆盖词(total≥3, days≥2) | ✅ 按阈值产出 |
|
||
| watch_terms | 中频观察词(total≤2, days≤2) | ✅ 按阈值产出 |
|
||
| alias_suggestions | 大小写变体/单复数/空格连词符差异 | ✅ 统计规则产出 |
|
||
| stopword_suggestions | — | ❌ 脚本硬编码为空 |
|
||
|
||
### 4. 应用建议(dry-run → review → apply)
|
||
|
||
```bash
|
||
# 先 dry-run 预览
|
||
cd /home/ubuntu/zhu/github/reader
|
||
.venv/bin/python3 scripts/apply_term_suggestions.py \
|
||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||
--accept-interest 强化学习 ReAct CLI \
|
||
--dry-run
|
||
|
||
# 确认后正式 apply(去掉 --dry-run)
|
||
.venv/bin/python3 scripts/apply_term_suggestions.py \
|
||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||
--accept-interest 强化学习 ReAct CLI \
|
||
```
|
||
|
||
支持的 accept 参数:
|
||
- `--accept-interest 词1 词2 ...` — 添加到 interest_keywords
|
||
- `--accept-watch 词1 词2 ...` — 添加到 watchlist
|
||
- `--accept-alias 源词1 源词2 ...` — 添加 alias 映射
|
||
- `--accept-stopword 词1 词2 ...` — 添加停用词
|
||
|
||
## 脚本能做什么 vs 不能做什么
|
||
|
||
### ✅ 脚本能做的(统计级清洗)
|
||
|
||
- 发现大小写变体(`vibe coding` → `Vibe Coding`)
|
||
- 发现单复数差异(`Agent Skill` → `Agent Skills`)
|
||
- 发现空格/连词符差异
|
||
- 按频次推荐 should-be-interest / should-be-watch 的词
|
||
- 批量 apply 到配置文件,自动记录变更日志
|
||
|
||
### ❌ 脚本不能做的(语义级清洗,需要人工判断)
|
||
|
||
- 识别产品名(`WorkBuddy`、`LibTV Agent`、`飞书妙搭`)→ 应加 stopword
|
||
- 识别模型名(`Qwen3-30B-A3B`、`GLM 5.2`、`Opus 4.8`)→ 应加 stopword
|
||
- 识别人名/地名(`Andrej Karpathy`、`成都天府长岛`)→ 应加 stopword
|
||
- 英文专有名词归一化到中文主题词(`Multi-Agent` → `多Agent`)
|
||
- 长产品名归一化(`火山云数据库PostgreSQL Serverless版` → `Serverless数据库`)
|
||
- 判断某个词对 Hugo tag 是否有用(`1688`、`OPC训练营` → 无效)
|
||
|
||
## 语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py)
|
||
|
||
除了统计规则的清洗,项目还有一个 **LLM 驱动的语义级清洗脚本**,能发现统计规则做不到的事情:
|
||
|
||
```bash
|
||
cd /home/ubuntu/zhu/github/reader
|
||
.venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \
|
||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||
```
|
||
|
||
参数:
|
||
- `--bundle` — review bundle(先跑 build_review_bundle.py)
|
||
- `--suggestions` — 规则层的建议 JSON(可选,供 LLM 避免重复)
|
||
- `--output` — 输出路径(自动生成)
|
||
- `--dry-run` — 只打印 prompt 不调 LLM
|
||
|
||
### 语义脚本能做的(而统计规则不能做的)
|
||
|
||
| 类型 | LLM 能发现什么 | 示例 |
|
||
|:----|:--------------|:-----|
|
||
| **semantic_alias** | 中英文映射、简称↔全称、近义词 | `AI智能体`→`AI Agent`, `AI Harness`→`Harness Engineering` |
|
||
| **stopword** | 语义判断哪些词太泛/不相干 | `自动化`(太泛), `AlphaFold`(生物领域), `陌生化`(穿透词) |
|
||
| **promote_to_interest** | 与用户关注领域对齐的新词 | `强化学习`, `Loop Engineering`, `MoE`, `ChatGPT` |
|
||
|
||
### 两阶段清洗 SOP
|
||
|
||
当需要清洗关键词时,按此顺序操作:
|
||
|
||
**阶段一:规则统计清洗(脚本发现 + 人工确认)**
|
||
1. `build_review_bundle.py` → 重建 bundle
|
||
2. `generate_term_cleanup_suggestions.py` → 产出统计级建议
|
||
3. 检查建议,决定哪些 accept
|
||
4. `apply_term_suggestions.py --dry-run` → 预览
|
||
5. `apply_term_suggestions.py` → 正式落地
|
||
|
||
**阶段二:语义级清洗(LLM 发现 + 人工确认)**
|
||
1. `generate_term_cleanup_semantic_suggestions.py` → 调 LLM 产生语义建议
|
||
2. 检查 LLM 的 semantic_alias / stopword / promote_to_interest
|
||
3. 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化)
|
||
|
||
**🔴 不要绕过确认流程直接改配置。** 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。
|
||
|
||
## 完整的语义级清洗操作流程
|
||
|
||
当需要大量添加 aliases/stopwords 时,推荐流程:
|
||
|
||
1. **跑一次 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py`** 看看统计建议
|
||
2. **走一遍实际 pipeline 产出**,收集所有 unique keywords:
|
||
```bash
|
||
for rid in $(ls outputs/freshrss/rerun/); do
|
||
python3 -c "
|
||
import json
|
||
d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json'))
|
||
for item in d['items']:
|
||
for kw in item['summary'].get('keywords', []):
|
||
print(kw)
|
||
"
|
||
done | sort -u
|
||
```
|
||
3. 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的
|
||
4. 分别更新 `term_aliases.json` 和 `term_stopwords.json`
|
||
5. **重建 term_index**:对每个受影响日期的 run 跑 `build_keyword_index.py`
|
||
6. **更新 Hugo tags**:从重建后的 `data/term_index/daily/YYYY-MM-DD.json` 取
|
||
|
||
## 当前配置 (2026-07-16)
|
||
|
||
| 文件 | 条目数 |
|
||
|:----|:------|
|
||
| `configs/term_aliases.json` | 142 |
|
||
| `configs/term_stopwords.json` | 106 |
|
||
| `configs/filter_context.personal.json` | 54 (interest_keywords) |
|
||
| `configs/term_watchlist.json` | 6 |
|
||
|
||
## 关键词/别名/停用词配置的更新规范
|
||
|
||
- `term_aliases.json` — 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式
|
||
- `term_stopwords.json` — 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词
|
||
- **随时可以加**,加了后重建 term_index 即可生效
|
||
- 不涉及 pipeline 重新跑——只影响下游展示
|