keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions

- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
This commit is contained in:
root
2026-05-14 17:17:49 +08:00
parent 4399c9ca90
commit 590d050218
13 changed files with 1592 additions and 57 deletions
@@ -0,0 +1,290 @@
# interest/watch 候选引擎改进方案
> 从固定阈值到自适应排位 + 趋势因子的演进
## 1. 背景
### 1.1 当前实现
`build_review_bundle.py` 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:
| 候选池 | 判断条件 | 依据 |
|--------|---------|------|
| `interest_review_candidates` | `total_count >= 3 AND days_seen >= 2` | `policy.interest_keyword_review` |
| `watch_review_candidates` | `total_count <= 2 AND days_seen <= 2` | `policy.watch_term_review` |
`generate_term_cleanup_suggestions.py` 则直接从这两个候选池过滤、去重、排序后输出。
### 1.2 当前方案的问题
**问题一:固定阈值不随数据量自适应**
```
场景 total_count=3 意味着什么
─────────────────────────────────────────────
7 天数据(~200 词) top 15%,有一定区分度 ✅
41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变
未来 200 天 仍然用 3 次,区分度稀释 ❌
```
同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。
**问题二:固定阈值忽略趋势信号**
- "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
- "Channels":total=3, recent=0 — 早期出现但近期消失
- 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。
**问题三:interest 和 watch 的分界线是硬的**
total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。
### 1.3 讨论结论
与老大讨论后确认:
1. interest/watch 是**统计判断**,不需要大模型介入,纯算法可以解决
2. 当前引擎缺的不是大模型,而是**算法本身没写完**——自适应维度(排位、趋势)还没实现
3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
4. 修改量小,可以在 1 小时内落地
---
## 2. 设计方案
### 2.1 核心思路
引入两个互补维度替代固定阈值:
```
判定维度 含义 数据来源
────────────────────────────────────────────────────────────
percentile(百分位排名) 该词 total_count 在所有词 term_stats
中的排位占比
growth(增速因子) 近期集中度 = recent_count daily 近 N 天
/ total_count
```
两个维度配合:
- **percentile** 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
- **growth** 衡量"这个词是持续出现还是近期爆发"——识别趋势信号
### 2.2 候选池划分逻辑
```
percentile
│
┌─────────────────────┐
│ top 5% │
│ → 建议 interest │ ← 高频稳定词
├─────────────────────┤
│ top 5%-20% │
│ → 建议 watch │ ← 有信号但未达 threshold
├─────────────────────┤
│ bottom 80% │
│ → 暂不处理 │ ← 噪声/低频
└─────────────────────┘
额外规则:
如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
→ 主动提升到 watch / 主动推 confirm
```
这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。
### 2.3 接口变化
**`configs/term_cleanup_policy.json`**:
```json
{
"schema_version": "v2",
"interest_keyword_review": {
"percentile_max": 0.05,
"growth_promotion": 0.5
},
"watch_term_review": {
"percentile_min": 0.05,
"percentile_max": 0.20
}
}
```
`v1` 的 `min_total_count`/`min_days_seen` 等绝对阈值字段不再使用。
**`build_review_bundle.py` 输出的候选项**:
```json
{
"term": "Anthropic",
"total_count": 13,
"days_seen": 10,
"percentile": 0.012,
"growth": 0.54,
"reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
}
```
### 2.4 不需要改动的部分
- `generate_term_cleanup_suggestions.py` — 它只消费候选池,不用改
- `apply_term_suggestions.py` — 消费 suggestions JSON,不用改
- `keyword-cleanup-bundle.json` 结构 — 向后兼容,新增 percentile/growth 字段
---
## 3. 实施计划
### 3.1 改动范围
| 文件 | 改动量 | 内容 |
|------|--------|------|
| `skills/keyword-cleanup-review/scripts/build_review_bundle.py` | ~40 行 | 新增 `_compute_percentile()` 和 `_compute_growth()` 函数;修改候选池生成逻辑;候选项中增加 percentile/growth |
| `configs/term_cleanup_policy.json` | ~10 行 | schema v2:percentile/growth 替代绝对阈值 |
### 3.2 实施步骤
1. **build_review_bundle.py**:在 `top_global_terms` 生成后,增加 percentile 计算函数和 growth 计算函数
2. **build_review_bundle.py**:修改 `interest_review_candidates` 和 `watch_review_candidates` 的生成逻辑,从固定阈值改为 percentile + growth
3. **build_review_bundle.py**:候选项增加 `percentile` 和 `growth` 字段,更新 `reason` 文案
4. **term_cleanup_policy.json**:更新为 v2 schema
5. **验证**:全量跑一次(`--days 365 --top 100`),对比新旧两份输出的差异
### 3.3 验证方法
```bash
# 1. 用旧版生成 baseline
cd /home/ubuntu/zhu/github/reader
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-baseline.json
# 2. 改代码后用新版生成
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-new.json
# 3. 对比 governance_hints
python3 -c "
import json
a = json.load(open('/tmp/bundle-baseline.json'))
b = json.load(open('/tmp/bundle-new.json'))
for key in ['interest_review_candidates', 'watch_review_candidates']:
old = set(i['term'] for i in a['governance_hints'][key])
new = set(i['term'] for i in b['governance_hints'][key])
print(f'{key}: 新增={new-old}, 减少={old-new}')
"
```
### 3.4 风险
| 风险 | 概率 | 应对 |
|------|------|------|
| 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 |
| growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 |
| 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 |
---
## 4. alias/stopword 设计方案
### 4.1 核心判断
alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。
| 类型 | 需要什么 | 判断方式 |
|------|---------|----------|
| 大小写变体 | 表层 | 规则:casefold 去重 |
| 单复数 | 表层 | 规则:去/加 s 后缀匹配 |
| 分词变体(空格/连字符) | 表层 | 规则:去空格归一 |
| 简写全称(MCP→Model Context Protocol) | **语义** | LLM |
| 中英文(上下文工程→Context Engineering) | **语义** | LLM |
| 同义不同名(Rush→猿辅导 Rush 平台) | **语义** | LLM |
| stopword(大模型、AI 太泛) | **语义** | LLM |
### 4.2 分层方案
```
输入:高频未覆盖词 + 已有 interest 词表
│
├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
│ 输出候选 alias 对
│
└── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
做语义聚类
输出 alias 组 + stopword 标记
```
### 4.3 规则层设计
在 `generate_term_cleanup_suggestions.py` 中新增 `_prepare_alias_suggestions()` 函数:
```python
def _prepare_alias_suggestions(top_terms, interest_keywords):
"""
基于表层规则生成 alias 建议。
规则1:casefold 匹配——同一个 casefold 下有多个原文变体
规则2:单复数——去掉/加上末尾 s 后匹配
规则3:分词变体——去空格/连字符后匹配
"""
```
优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。
### 4.4 LLM 层设计
单独脚本,非 generate 主链路的一部分。
```bash
python scripts/generate_term_cleanup_semantic_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
LLM prompt 设计:
```
你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
请做三件事:
1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest
用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地
```
LLM 层输出格式:
```json
{
"alias_suggestions": [
{"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
],
"stopword_suggestions": [
{"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
]
}
```
### 4.5 预期效果
| 覆盖类型 | 规则层 | LLM 层 |
|---------|--------|--------|
| 大小写变体 | ✅ | — |
| 单复数 | ✅ | — |
| 分词变体 | ✅ | — |
| 简写全称 | — | ✅ |
| 中英文映射 | — | ✅ |
| 同义不同名 | — | ✅ |
| stopword 判断 | — | ✅ |
---
## 5. 讨论记录
- 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
- 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
- 2026-05-14:确认本方案不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
- 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先