keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions

- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
This commit is contained in:
root
2026-05-14 17:17:49 +08:00
parent 4399c9ca90
commit 590d050218
13 changed files with 1592 additions and 57 deletions
@@ -0,0 +1,151 @@
# 关键词清洗流程概述
> 2026-05-14 初版
> 从"数据记录"到"人工确认落盘"的完整链路
---
## 整体数据流
```
每日日报 pipeline
│
▼
term_index/daily/YYYY-MM-DD.json ← 每天一篇候选文章的热词统计
│
▼
term_index/term_stats.json ← 所有 daily 的汇总(1070 个词)
│
├──── build_review_bundle.py ← 打包为审查数据包
│ │
│ ▼
│ review/keyword-cleanup-bundle.json
│ │
│ ▼
│ generate_term_cleanup_suggestions.py
│ │
│ ▼
│ review/term-cleanup-suggestions-YYYY-MM-DD.json ← 正式建议产物
│ │
│ ▼
│ (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md ← 展示稿
│
├──── 人工确认哪些建议 accept
│
▼
apply_term_suggestions.py ← 写入配置
│
├── configs/filter_context.personal.json ← interest_keywords
├── configs/term_aliases.json ← alias
├── configs/term_stopwords.json ← stopword
├── configs/term_watchlist.json ← watch
└── configs/term_change_log.json ← 变更日志
```
---
## 各环节说明
### 阶段 1:数据记录(每日自动)
```bash
# FreshRSS pipeline 跑完后自动产出
data/term_index/daily/2026-05-14.json
```
- 每天一篇,记录当天候选文章中出现的热词
- 包含 term、total_count、days_seen 等信息
- 目前累计 **41 天**,共 **1070 个独立词**
### 阶段 2:全量汇总(每日自动)
```bash
data/term_index/term_stats.json
```
- 从所有 daily 文件重建,会覆盖重跑
- 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13)
### 阶段 3:构建审查数据包(手动触发)
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 \
--top 100 \
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
- 把 term_stats + 当前配置打成一包,方便后续处理
- 输出:`review/keyword-cleanup-bundle.json`
### 阶段 4:生成建议(手动触发)
```bash
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
#### 当前产出能力
| 建议类型 | 状态 | 当前阈值 | 说明 |
|---------|------|----------|------|
| interest_keyword_suggestions | ✅ **已实现** | total≥3, days≥2 | 产出 20 条 |
| watch_terms | ✅ **已实现** | total≤2, days≤2 | 本次 0 条 |
| alias_suggestions | ❌ **硬编码为空** | policy 有阈值(total≥2, days≥2)但脚本未实现 | |
| stopword_suggestions | ❌ **硬编码为空** | policy 有阈值(total≤2, days≤2)但脚本未实现 | |
**关键发现:** alias 和 stopword 不是"阈值太保守",是 **generate 脚本里压根没写对应的生成函数**。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 `[]` 和 `0`。
### 阶段 5:人工确认(手动)
```
OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply
```
本次模式:
- 高频(≥5次/5天以上)→ 强烈推荐 ✅
- 中频(3-4次)→ 附带建议 ✅
- 泛词 → 建议跳过 ❌
### 阶段 6:落盘配置(手动)
```bash
python scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 词1 词2 ...
```
- dry-run 预览 → 确认后正式 apply
- 写入 `configs/filter_context.personal.json`
- 同步记录到 `term_change_log.json`
- **不备份原始配置**(待优化)
- **apply 后不自动清理 review 目录**(待优化)
### 阶段 7:维护清理(按需)
由 OpenClaw 侧 `reader-keyword-maintenance` skill 处理:
- 删除旧 markdown 展示稿
- 保留最近一份 bundle
- 保守保留 suggestions JSON
---
## 当前配置资产
| 文件 | 内容 | 数据量 |
|------|------|--------|
| `filter_context.personal.json` | interest_keywords | 52 个 |
| `term_aliases.json` | 别名映射 | 0 组(未启用) |
| `term_stopwords.json` | 停用词 | 0 个(未启用) |
| `term_watchlist.json` | 观察词 | 6 个 |
| `term_change_log.json` | 所有变更记录 | 已记录 |
---
## 待优化项
1. **alias/stopword 建议生成为空** — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数
2. **apply 前无配置备份** — 建议 apply 前自动 cp 备份
3. **apply 后无自动收尾** — 建议 apply 后自动删旧 markdown 和 bundle
4. **alias 识别依赖规则而非 LLM** — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply