Files
reader/skills/reader-digest-flow/references/keyword-engine-maintenance.md
T
root 5eb390e3ed docs: 添加 Agent Skill 到项目仓库
- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/)
- README 新增 Agent Skill 章节说明供 Agent 使用的工作流
2026-07-28 18:38:54 +08:00

196 lines
7.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 关键词引擎维护流程
## 概述
reader 项目内置了完整的关键词清洗链路,用于管理 `term_aliases.json`、`term_stopwords.json`、`filter_context.personal.json` 等配置。本文件说明链路各环节的职责与使用方式。
## 完整数据流
```
每日日报 pipeline
│
▼
term_index/daily/YYYY-MM-DD.json ← build_keyword_index.py (pipeline 末环节自动)
│
▼
term_index/term_stats.json ← 全量汇总
│
▼
build_review_bundle.py ← 打包审查数据包 (手动触发)
│
▼
generate_term_cleanup_suggestions.py ← 生成建议 (手动触发)
│
▼
apply_term_suggestions.py ← 人工确认后落盘 (手动触发)
```
## 各环节命令
### 1. 重建每日 term_index(pipeline 末环节自动执行,也可手动)
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/build_keyword_index.py \
--input outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
```
### 2. 重建 review bundle(打包当前配置+统计供审查)
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 \
--top 200 \
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
参数:
- `--days 365` — 考虑最近多少天的统计数据
- `--top 200` — 取前 N 个高频词纳入 bundle
### 3. 生成清洗建议
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
输出:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json` — 正式建议产物
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md` — 人工阅读展示稿
当前脚本能生成的建议类型:
| 类型 | 生成规则 | 当前产出 |
|:----|:---------|:---------|
| interest_keyword_suggestions | 高频未覆盖词(total≥3, days≥2) | ✅ 按阈值产出 |
| watch_terms | 中频观察词(total≤2, days≤2) | ✅ 按阈值产出 |
| alias_suggestions | 大小写变体/单复数/空格连词符差异 | ✅ 统计规则产出 |
| stopword_suggestions | — | ❌ 脚本硬编码为空 |
### 4. 应用建议(dry-run → review → apply)
```bash
# 先 dry-run 预览
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
--dry-run
# 确认后正式 apply(去掉 --dry-run)
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
```
支持的 accept 参数:
- `--accept-interest 词1 词2 ...` — 添加到 interest_keywords
- `--accept-watch 词1 词2 ...` — 添加到 watchlist
- `--accept-alias 源词1 源词2 ...` — 添加 alias 映射
- `--accept-stopword 词1 词2 ...` — 添加停用词
## 脚本能做什么 vs 不能做什么
### ✅ 脚本能做的(统计级清洗)
- 发现大小写变体(`vibe coding` → `Vibe Coding`)
- 发现单复数差异(`Agent Skill` → `Agent Skills`)
- 发现空格/连词符差异
- 按频次推荐 should-be-interest / should-be-watch 的词
- 批量 apply 到配置文件,自动记录变更日志
### ❌ 脚本不能做的(语义级清洗,需要人工判断)
- 识别产品名(`WorkBuddy`、`LibTV Agent`、`飞书妙搭`)→ 应加 stopword
- 识别模型名(`Qwen3-30B-A3B`、`GLM 5.2`、`Opus 4.8`)→ 应加 stopword
- 识别人名/地名(`Andrej Karpathy`、`成都天府长岛`)→ 应加 stopword
- 英文专有名词归一化到中文主题词(`Multi-Agent` → `多Agent`)
- 长产品名归一化(`火山云数据库PostgreSQL Serverless版` → `Serverless数据库`)
- 判断某个词对 Hugo tag 是否有用(`1688`、`OPC训练营` → 无效)
## 语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py)
除了统计规则的清洗,项目还有一个 **LLM 驱动的语义级清洗脚本**,能发现统计规则做不到的事情:
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
参数:
- `--bundle` — review bundle(先跑 build_review_bundle.py)
- `--suggestions` — 规则层的建议 JSON(可选,供 LLM 避免重复)
- `--output` — 输出路径(自动生成)
- `--dry-run` — 只打印 prompt 不调 LLM
### 语义脚本能做的(而统计规则不能做的)
| 类型 | LLM 能发现什么 | 示例 |
|:----|:--------------|:-----|
| **semantic_alias** | 中英文映射、简称↔全称、近义词 | `AI智能体`→`AI Agent`, `AI Harness`→`Harness Engineering` |
| **stopword** | 语义判断哪些词太泛/不相干 | `自动化`(太泛), `AlphaFold`(生物领域), `陌生化`(穿透词) |
| **promote_to_interest** | 与用户关注领域对齐的新词 | `强化学习`, `Loop Engineering`, `MoE`, `ChatGPT` |
### 两阶段清洗 SOP
当需要清洗关键词时,按此顺序操作:
**阶段一:规则统计清洗(脚本发现 + 人工确认)**
1. `build_review_bundle.py` → 重建 bundle
2. `generate_term_cleanup_suggestions.py` → 产出统计级建议
3. 检查建议,决定哪些 accept
4. `apply_term_suggestions.py --dry-run` → 预览
5. `apply_term_suggestions.py` → 正式落地
**阶段二:语义级清洗(LLM 发现 + 人工确认)**
1. `generate_term_cleanup_semantic_suggestions.py` → 调 LLM 产生语义建议
2. 检查 LLM 的 semantic_alias / stopword / promote_to_interest
3. 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化)
**🔴 不要绕过确认流程直接改配置。** 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。
## 完整的语义级清洗操作流程
当需要大量添加 aliases/stopwords 时,推荐流程:
1. **跑一次 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py`** 看看统计建议
2. **走一遍实际 pipeline 产出**,收集所有 unique keywords:
```bash
for rid in $(ls outputs/freshrss/rerun/); do
python3 -c "
import json
d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json'))
for item in d['items']:
for kw in item['summary'].get('keywords', []):
print(kw)
"
done | sort -u
```
3. 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的
4. 分别更新 `term_aliases.json` 和 `term_stopwords.json`
5. **重建 term_index**:对每个受影响日期的 run 跑 `build_keyword_index.py`
6. **更新 Hugo tags**:从重建后的 `data/term_index/daily/YYYY-MM-DD.json` 取
## 当前配置 (2026-07-16)
| 文件 | 条目数 |
|:----|:------|
| `configs/term_aliases.json` | 142 |
| `configs/term_stopwords.json` | 106 |
| `configs/filter_context.personal.json` | 54 (interest_keywords) |
| `configs/term_watchlist.json` | 6 |
## 关键词/别名/停用词配置的更新规范
- `term_aliases.json` — 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式
- `term_stopwords.json` — 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词
- **随时可以加**,加了后重建 term_index 即可生效
- 不涉及 pipeline 重新跑——只影响下游展示