Files
reader/plans/keyword-cleanup-interest-watch-engine-improvement.md
root 590d050218 keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions
- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
2026-05-14 17:17:49 +08:00

291 lines
10 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# interest/watch 候选引擎改进方案
> 从固定阈值到自适应排位 + 趋势因子的演进
## 1. 背景
### 1.1 当前实现
`build_review_bundle.py` 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:
| 候选池 | 判断条件 | 依据 |
|--------|---------|------|
| `interest_review_candidates` | `total_count >= 3 AND days_seen >= 2` | `policy.interest_keyword_review` |
| `watch_review_candidates` | `total_count <= 2 AND days_seen <= 2` | `policy.watch_term_review` |
`generate_term_cleanup_suggestions.py` 则直接从这两个候选池过滤、去重、排序后输出。
### 1.2 当前方案的问题
**问题一:固定阈值不随数据量自适应**
```
场景 total_count=3 意味着什么
─────────────────────────────────────────────
7 天数据(~200 词) top 15%,有一定区分度 ✅
41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变
未来 200 天 仍然用 3 次,区分度稀释 ❌
```
同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。
**问题二:固定阈值忽略趋势信号**
- "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
- "Channels":total=3, recent=0 — 早期出现但近期消失
- 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。
**问题三:interest 和 watch 的分界线是硬的**
total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。
### 1.3 讨论结论
与老大讨论后确认:
1. interest/watch 是**统计判断**,不需要大模型介入,纯算法可以解决
2. 当前引擎缺的不是大模型,而是**算法本身没写完**——自适应维度(排位、趋势)还没实现
3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
4. 修改量小,可以在 1 小时内落地
---
## 2. 设计方案
### 2.1 核心思路
引入两个互补维度替代固定阈值:
```
判定维度 含义 数据来源
────────────────────────────────────────────────────────────
percentile(百分位排名) 该词 total_count 在所有词 term_stats
中的排位占比
growth(增速因子) 近期集中度 = recent_count daily 近 N 天
/ total_count
```
两个维度配合:
- **percentile** 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
- **growth** 衡量"这个词是持续出现还是近期爆发"——识别趋势信号
### 2.2 候选池划分逻辑
```
percentile
│
┌─────────────────────┐
│ top 5% │
│ → 建议 interest │ ← 高频稳定词
├─────────────────────┤
│ top 5%-20% │
│ → 建议 watch │ ← 有信号但未达 threshold
├─────────────────────┤
│ bottom 80% │
│ → 暂不处理 │ ← 噪声/低频
└─────────────────────┘
额外规则:
如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
→ 主动提升到 watch / 主动推 confirm
```
这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。
### 2.3 接口变化
**`configs/term_cleanup_policy.json`**:
```json
{
"schema_version": "v2",
"interest_keyword_review": {
"percentile_max": 0.05,
"growth_promotion": 0.5
},
"watch_term_review": {
"percentile_min": 0.05,
"percentile_max": 0.20
}
}
```
`v1` 的 `min_total_count`/`min_days_seen` 等绝对阈值字段不再使用。
**`build_review_bundle.py` 输出的候选项**:
```json
{
"term": "Anthropic",
"total_count": 13,
"days_seen": 10,
"percentile": 0.012,
"growth": 0.54,
"reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
}
```
### 2.4 不需要改动的部分
- `generate_term_cleanup_suggestions.py` — 它只消费候选池,不用改
- `apply_term_suggestions.py` — 消费 suggestions JSON,不用改
- `keyword-cleanup-bundle.json` 结构 — 向后兼容,新增 percentile/growth 字段
---
## 3. 实施计划
### 3.1 改动范围
| 文件 | 改动量 | 内容 |
|------|--------|------|
| `skills/keyword-cleanup-review/scripts/build_review_bundle.py` | ~40 行 | 新增 `_compute_percentile()` 和 `_compute_growth()` 函数;修改候选池生成逻辑;候选项中增加 percentile/growth |
| `configs/term_cleanup_policy.json` | ~10 行 | schema v2:percentile/growth 替代绝对阈值 |
### 3.2 实施步骤
1. **build_review_bundle.py**:在 `top_global_terms` 生成后,增加 percentile 计算函数和 growth 计算函数
2. **build_review_bundle.py**:修改 `interest_review_candidates` 和 `watch_review_candidates` 的生成逻辑,从固定阈值改为 percentile + growth
3. **build_review_bundle.py**:候选项增加 `percentile` 和 `growth` 字段,更新 `reason` 文案
4. **term_cleanup_policy.json**:更新为 v2 schema
5. **验证**:全量跑一次(`--days 365 --top 100`),对比新旧两份输出的差异
### 3.3 验证方法
```bash
# 1. 用旧版生成 baseline
cd /home/ubuntu/zhu/github/reader
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-baseline.json
# 2. 改代码后用新版生成
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-new.json
# 3. 对比 governance_hints
python3 -c "
import json
a = json.load(open('/tmp/bundle-baseline.json'))
b = json.load(open('/tmp/bundle-new.json'))
for key in ['interest_review_candidates', 'watch_review_candidates']:
old = set(i['term'] for i in a['governance_hints'][key])
new = set(i['term'] for i in b['governance_hints'][key])
print(f'{key}: 新增={new-old}, 减少={old-new}')
"
```
### 3.4 风险
| 风险 | 概率 | 应对 |
|------|------|------|
| 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 |
| growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 |
| 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 |
---
## 4. alias/stopword 设计方案
### 4.1 核心判断
alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。
| 类型 | 需要什么 | 判断方式 |
|------|---------|----------|
| 大小写变体 | 表层 | 规则:casefold 去重 |
| 单复数 | 表层 | 规则:去/加 s 后缀匹配 |
| 分词变体(空格/连字符) | 表层 | 规则:去空格归一 |
| 简写全称(MCP→Model Context Protocol) | **语义** | LLM |
| 中英文(上下文工程→Context Engineering) | **语义** | LLM |
| 同义不同名(Rush→猿辅导 Rush 平台) | **语义** | LLM |
| stopword(大模型、AI 太泛) | **语义** | LLM |
### 4.2 分层方案
```
输入:高频未覆盖词 + 已有 interest 词表
│
├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
│ 输出候选 alias 对
│
└── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
做语义聚类
输出 alias 组 + stopword 标记
```
### 4.3 规则层设计
在 `generate_term_cleanup_suggestions.py` 中新增 `_prepare_alias_suggestions()` 函数:
```python
def _prepare_alias_suggestions(top_terms, interest_keywords):
"""
基于表层规则生成 alias 建议。
规则1:casefold 匹配——同一个 casefold 下有多个原文变体
规则2:单复数——去掉/加上末尾 s 后匹配
规则3:分词变体——去空格/连字符后匹配
"""
```
优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。
### 4.4 LLM 层设计
单独脚本,非 generate 主链路的一部分。
```bash
python scripts/generate_term_cleanup_semantic_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
LLM prompt 设计:
```
你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
请做三件事:
1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest
用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地
```
LLM 层输出格式:
```json
{
"alias_suggestions": [
{"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
],
"stopword_suggestions": [
{"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
]
}
```
### 4.5 预期效果
| 覆盖类型 | 规则层 | LLM 层 |
|---------|--------|--------|
| 大小写变体 | ✅ | — |
| 单复数 | ✅ | — |
| 分词变体 | ✅ | — |
| 简写全称 | — | ✅ |
| 中英文映射 | — | ✅ |
| 同义不同名 | — | ✅ |
| stopword 判断 | — | ✅ |
---
## 5. 讨论记录
- 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
- 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
- 2026-05-14:确认本方案不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
- 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先