docs: 添加 Agent Skill 到项目仓库

- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/)
- README 新增 Agent Skill 章节说明供 Agent 使用的工作流
This commit is contained in:
root
2026-07-28 18:38:54 +08:00
parent 7b791ac947
commit 5eb390e3ed
16 changed files with 2112 additions and 0 deletions
@@ -0,0 +1,195 @@
# 关键词引擎维护流程
## 概述
reader 项目内置了完整的关键词清洗链路,用于管理 `term_aliases.json`、`term_stopwords.json`、`filter_context.personal.json` 等配置。本文件说明链路各环节的职责与使用方式。
## 完整数据流
```
每日日报 pipeline
│
▼
term_index/daily/YYYY-MM-DD.json ← build_keyword_index.py (pipeline 末环节自动)
│
▼
term_index/term_stats.json ← 全量汇总
│
▼
build_review_bundle.py ← 打包审查数据包 (手动触发)
│
▼
generate_term_cleanup_suggestions.py ← 生成建议 (手动触发)
│
▼
apply_term_suggestions.py ← 人工确认后落盘 (手动触发)
```
## 各环节命令
### 1. 重建每日 term_index(pipeline 末环节自动执行,也可手动)
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/build_keyword_index.py \
--input outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
```
### 2. 重建 review bundle(打包当前配置+统计供审查)
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 \
--top 200 \
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
参数:
- `--days 365` — 考虑最近多少天的统计数据
- `--top 200` — 取前 N 个高频词纳入 bundle
### 3. 生成清洗建议
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
输出:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json` — 正式建议产物
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md` — 人工阅读展示稿
当前脚本能生成的建议类型:
| 类型 | 生成规则 | 当前产出 |
|:----|:---------|:---------|
| interest_keyword_suggestions | 高频未覆盖词(total≥3, days≥2) | ✅ 按阈值产出 |
| watch_terms | 中频观察词(total≤2, days≤2) | ✅ 按阈值产出 |
| alias_suggestions | 大小写变体/单复数/空格连词符差异 | ✅ 统计规则产出 |
| stopword_suggestions | — | ❌ 脚本硬编码为空 |
### 4. 应用建议(dry-run → review → apply)
```bash
# 先 dry-run 预览
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
--dry-run
# 确认后正式 apply(去掉 --dry-run)
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
```
支持的 accept 参数:
- `--accept-interest 词1 词2 ...` — 添加到 interest_keywords
- `--accept-watch 词1 词2 ...` — 添加到 watchlist
- `--accept-alias 源词1 源词2 ...` — 添加 alias 映射
- `--accept-stopword 词1 词2 ...` — 添加停用词
## 脚本能做什么 vs 不能做什么
### ✅ 脚本能做的(统计级清洗)
- 发现大小写变体(`vibe coding` → `Vibe Coding`)
- 发现单复数差异(`Agent Skill` → `Agent Skills`)
- 发现空格/连词符差异
- 按频次推荐 should-be-interest / should-be-watch 的词
- 批量 apply 到配置文件,自动记录变更日志
### ❌ 脚本不能做的(语义级清洗,需要人工判断)
- 识别产品名(`WorkBuddy`、`LibTV Agent`、`飞书妙搭`)→ 应加 stopword
- 识别模型名(`Qwen3-30B-A3B`、`GLM 5.2`、`Opus 4.8`)→ 应加 stopword
- 识别人名/地名(`Andrej Karpathy`、`成都天府长岛`)→ 应加 stopword
- 英文专有名词归一化到中文主题词(`Multi-Agent` → `多Agent`)
- 长产品名归一化(`火山云数据库PostgreSQL Serverless版` → `Serverless数据库`)
- 判断某个词对 Hugo tag 是否有用(`1688`、`OPC训练营` → 无效)
## 语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py)
除了统计规则的清洗,项目还有一个 **LLM 驱动的语义级清洗脚本**,能发现统计规则做不到的事情:
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
参数:
- `--bundle` — review bundle(先跑 build_review_bundle.py)
- `--suggestions` — 规则层的建议 JSON(可选,供 LLM 避免重复)
- `--output` — 输出路径(自动生成)
- `--dry-run` — 只打印 prompt 不调 LLM
### 语义脚本能做的(而统计规则不能做的)
| 类型 | LLM 能发现什么 | 示例 |
|:----|:--------------|:-----|
| **semantic_alias** | 中英文映射、简称↔全称、近义词 | `AI智能体`→`AI Agent`, `AI Harness`→`Harness Engineering` |
| **stopword** | 语义判断哪些词太泛/不相干 | `自动化`(太泛), `AlphaFold`(生物领域), `陌生化`(穿透词) |
| **promote_to_interest** | 与用户关注领域对齐的新词 | `强化学习`, `Loop Engineering`, `MoE`, `ChatGPT` |
### 两阶段清洗 SOP
当需要清洗关键词时,按此顺序操作:
**阶段一:规则统计清洗(脚本发现 + 人工确认)**
1. `build_review_bundle.py` → 重建 bundle
2. `generate_term_cleanup_suggestions.py` → 产出统计级建议
3. 检查建议,决定哪些 accept
4. `apply_term_suggestions.py --dry-run` → 预览
5. `apply_term_suggestions.py` → 正式落地
**阶段二:语义级清洗(LLM 发现 + 人工确认)**
1. `generate_term_cleanup_semantic_suggestions.py` → 调 LLM 产生语义建议
2. 检查 LLM 的 semantic_alias / stopword / promote_to_interest
3. 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化)
**🔴 不要绕过确认流程直接改配置。** 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。
## 完整的语义级清洗操作流程
当需要大量添加 aliases/stopwords 时,推荐流程:
1. **跑一次 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py`** 看看统计建议
2. **走一遍实际 pipeline 产出**,收集所有 unique keywords:
```bash
for rid in $(ls outputs/freshrss/rerun/); do
python3 -c "
import json
d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json'))
for item in d['items']:
for kw in item['summary'].get('keywords', []):
print(kw)
"
done | sort -u
```
3. 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的
4. 分别更新 `term_aliases.json` 和 `term_stopwords.json`
5. **重建 term_index**:对每个受影响日期的 run 跑 `build_keyword_index.py`
6. **更新 Hugo tags**:从重建后的 `data/term_index/daily/YYYY-MM-DD.json` 取
## 当前配置 (2026-07-16)
| 文件 | 条目数 |
|:----|:------|
| `configs/term_aliases.json` | 142 |
| `configs/term_stopwords.json` | 106 |
| `configs/filter_context.personal.json` | 54 (interest_keywords) |
| `configs/term_watchlist.json` | 6 |
## 关键词/别名/停用词配置的更新规范
- `term_aliases.json` — 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式
- `term_stopwords.json` — 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词
- **随时可以加**,加了后重建 term_index 即可生效
- 不涉及 pipeline 重新跑——只影响下游展示