keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions
- build_review_bundle.py: 新增 _compute_percentile/_compute_growth, 候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy) - term_cleanup_policy.json: 升级 v2 schema - generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions, 规则层输出 alias (大小写/单复数/分词变体) - generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本 (DeepSeek API, 产出 semantic alias/stopword/promote) - SKILL.md: 更新为 5 Phase 工作流程 - 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个 - docs/design/keyword-cleanup-flow-overview.md: 流程文档 - plans/: 引擎设计方案
This commit is contained in:
@@ -0,0 +1,290 @@
|
||||
# interest/watch 候选引擎改进方案
|
||||
|
||||
> 从固定阈值到自适应排位 + 趋势因子的演进
|
||||
|
||||
## 1. 背景
|
||||
|
||||
### 1.1 当前实现
|
||||
|
||||
`build_review_bundle.py` 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:
|
||||
|
||||
| 候选池 | 判断条件 | 依据 |
|
||||
|--------|---------|------|
|
||||
| `interest_review_candidates` | `total_count >= 3 AND days_seen >= 2` | `policy.interest_keyword_review` |
|
||||
| `watch_review_candidates` | `total_count <= 2 AND days_seen <= 2` | `policy.watch_term_review` |
|
||||
|
||||
`generate_term_cleanup_suggestions.py` 则直接从这两个候选池过滤、去重、排序后输出。
|
||||
|
||||
### 1.2 当前方案的问题
|
||||
|
||||
**问题一:固定阈值不随数据量自适应**
|
||||
|
||||
```
|
||||
场景 total_count=3 意味着什么
|
||||
─────────────────────────────────────────────
|
||||
7 天数据(~200 词) top 15%,有一定区分度 ✅
|
||||
41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变
|
||||
未来 200 天 仍然用 3 次,区分度稀释 ❌
|
||||
```
|
||||
|
||||
同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。
|
||||
|
||||
**问题二:固定阈值忽略趋势信号**
|
||||
|
||||
- "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
|
||||
- "Channels":total=3, recent=0 — 早期出现但近期消失
|
||||
- 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。
|
||||
|
||||
**问题三:interest 和 watch 的分界线是硬的**
|
||||
|
||||
total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。
|
||||
|
||||
### 1.3 讨论结论
|
||||
|
||||
与老大讨论后确认:
|
||||
|
||||
1. interest/watch 是**统计判断**,不需要大模型介入,纯算法可以解决
|
||||
2. 当前引擎缺的不是大模型,而是**算法本身没写完**——自适应维度(排位、趋势)还没实现
|
||||
3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
|
||||
4. 修改量小,可以在 1 小时内落地
|
||||
|
||||
---
|
||||
|
||||
## 2. 设计方案
|
||||
|
||||
### 2.1 核心思路
|
||||
|
||||
引入两个互补维度替代固定阈值:
|
||||
|
||||
```
|
||||
判定维度 含义 数据来源
|
||||
────────────────────────────────────────────────────────────
|
||||
percentile(百分位排名) 该词 total_count 在所有词 term_stats
|
||||
中的排位占比
|
||||
growth(增速因子) 近期集中度 = recent_count daily 近 N 天
|
||||
/ total_count
|
||||
```
|
||||
|
||||
两个维度配合:
|
||||
|
||||
- **percentile** 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
|
||||
- **growth** 衡量"这个词是持续出现还是近期爆发"——识别趋势信号
|
||||
|
||||
### 2.2 候选池划分逻辑
|
||||
|
||||
```
|
||||
percentile
|
||||
│
|
||||
┌─────────────────────┐
|
||||
│ top 5% │
|
||||
│ → 建议 interest │ ← 高频稳定词
|
||||
├─────────────────────┤
|
||||
│ top 5%-20% │
|
||||
│ → 建议 watch │ ← 有信号但未达 threshold
|
||||
├─────────────────────┤
|
||||
│ bottom 80% │
|
||||
│ → 暂不处理 │ ← 噪声/低频
|
||||
└─────────────────────┘
|
||||
|
||||
额外规则:
|
||||
如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
|
||||
→ 主动提升到 watch / 主动推 confirm
|
||||
```
|
||||
|
||||
这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。
|
||||
|
||||
### 2.3 接口变化
|
||||
|
||||
**`configs/term_cleanup_policy.json`**:
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5
|
||||
},
|
||||
"watch_term_review": {
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`v1` 的 `min_total_count`/`min_days_seen` 等绝对阈值字段不再使用。
|
||||
|
||||
**`build_review_bundle.py` 输出的候选项**:
|
||||
|
||||
```json
|
||||
{
|
||||
"term": "Anthropic",
|
||||
"total_count": 13,
|
||||
"days_seen": 10,
|
||||
"percentile": 0.012,
|
||||
"growth": 0.54,
|
||||
"reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
|
||||
}
|
||||
```
|
||||
|
||||
### 2.4 不需要改动的部分
|
||||
|
||||
- `generate_term_cleanup_suggestions.py` — 它只消费候选池,不用改
|
||||
- `apply_term_suggestions.py` — 消费 suggestions JSON,不用改
|
||||
- `keyword-cleanup-bundle.json` 结构 — 向后兼容,新增 percentile/growth 字段
|
||||
|
||||
---
|
||||
|
||||
## 3. 实施计划
|
||||
|
||||
### 3.1 改动范围
|
||||
|
||||
| 文件 | 改动量 | 内容 |
|
||||
|------|--------|------|
|
||||
| `skills/keyword-cleanup-review/scripts/build_review_bundle.py` | ~40 行 | 新增 `_compute_percentile()` 和 `_compute_growth()` 函数;修改候选池生成逻辑;候选项中增加 percentile/growth |
|
||||
| `configs/term_cleanup_policy.json` | ~10 行 | schema v2:percentile/growth 替代绝对阈值 |
|
||||
|
||||
### 3.2 实施步骤
|
||||
|
||||
1. **build_review_bundle.py**:在 `top_global_terms` 生成后,增加 percentile 计算函数和 growth 计算函数
|
||||
2. **build_review_bundle.py**:修改 `interest_review_candidates` 和 `watch_review_candidates` 的生成逻辑,从固定阈值改为 percentile + growth
|
||||
3. **build_review_bundle.py**:候选项增加 `percentile` 和 `growth` 字段,更新 `reason` 文案
|
||||
4. **term_cleanup_policy.json**:更新为 v2 schema
|
||||
5. **验证**:全量跑一次(`--days 365 --top 100`),对比新旧两份输出的差异
|
||||
|
||||
### 3.3 验证方法
|
||||
|
||||
```bash
|
||||
# 1. 用旧版生成 baseline
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 --top 100 \
|
||||
--output /tmp/bundle-baseline.json
|
||||
|
||||
# 2. 改代码后用新版生成
|
||||
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 --top 100 \
|
||||
--output /tmp/bundle-new.json
|
||||
|
||||
# 3. 对比 governance_hints
|
||||
python3 -c "
|
||||
import json
|
||||
a = json.load(open('/tmp/bundle-baseline.json'))
|
||||
b = json.load(open('/tmp/bundle-new.json'))
|
||||
for key in ['interest_review_candidates', 'watch_review_candidates']:
|
||||
old = set(i['term'] for i in a['governance_hints'][key])
|
||||
new = set(i['term'] for i in b['governance_hints'][key])
|
||||
print(f'{key}: 新增={new-old}, 减少={old-new}')
|
||||
"
|
||||
```
|
||||
|
||||
### 3.4 风险
|
||||
|
||||
| 风险 | 概率 | 应对 |
|
||||
|------|------|------|
|
||||
| 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 |
|
||||
| growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 |
|
||||
| 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 |
|
||||
|
||||
---
|
||||
|
||||
## 4. alias/stopword 设计方案
|
||||
|
||||
### 4.1 核心判断
|
||||
|
||||
alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。
|
||||
|
||||
| 类型 | 需要什么 | 判断方式 |
|
||||
|------|---------|----------|
|
||||
| 大小写变体 | 表层 | 规则:casefold 去重 |
|
||||
| 单复数 | 表层 | 规则:去/加 s 后缀匹配 |
|
||||
| 分词变体(空格/连字符) | 表层 | 规则:去空格归一 |
|
||||
| 简写全称(MCP→Model Context Protocol) | **语义** | LLM |
|
||||
| 中英文(上下文工程→Context Engineering) | **语义** | LLM |
|
||||
| 同义不同名(Rush→猿辅导 Rush 平台) | **语义** | LLM |
|
||||
| stopword(大模型、AI 太泛) | **语义** | LLM |
|
||||
|
||||
### 4.2 分层方案
|
||||
|
||||
```
|
||||
输入:高频未覆盖词 + 已有 interest 词表
|
||||
│
|
||||
├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
|
||||
│ 输出候选 alias 对
|
||||
│
|
||||
└── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
|
||||
做语义聚类
|
||||
输出 alias 组 + stopword 标记
|
||||
```
|
||||
|
||||
### 4.3 规则层设计
|
||||
|
||||
在 `generate_term_cleanup_suggestions.py` 中新增 `_prepare_alias_suggestions()` 函数:
|
||||
|
||||
```python
|
||||
def _prepare_alias_suggestions(top_terms, interest_keywords):
|
||||
"""
|
||||
基于表层规则生成 alias 建议。
|
||||
规则1:casefold 匹配——同一个 casefold 下有多个原文变体
|
||||
规则2:单复数——去掉/加上末尾 s 后匹配
|
||||
规则3:分词变体——去空格/连字符后匹配
|
||||
"""
|
||||
```
|
||||
|
||||
优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。
|
||||
|
||||
### 4.4 LLM 层设计
|
||||
|
||||
单独脚本,非 generate 主链路的一部分。
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||||
```
|
||||
|
||||
LLM prompt 设计:
|
||||
|
||||
```
|
||||
你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
|
||||
请做三件事:
|
||||
|
||||
1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
|
||||
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
|
||||
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest
|
||||
|
||||
用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地
|
||||
```
|
||||
|
||||
LLM 层输出格式:
|
||||
|
||||
```json
|
||||
{
|
||||
"alias_suggestions": [
|
||||
{"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
|
||||
],
|
||||
"stopword_suggestions": [
|
||||
{"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 4.5 预期效果
|
||||
|
||||
| 覆盖类型 | 规则层 | LLM 层 |
|
||||
|---------|--------|--------|
|
||||
| 大小写变体 | ✅ | — |
|
||||
| 单复数 | ✅ | — |
|
||||
| 分词变体 | ✅ | — |
|
||||
| 简写全称 | — | ✅ |
|
||||
| 中英文映射 | — | ✅ |
|
||||
| 同义不同名 | — | ✅ |
|
||||
| stopword 判断 | — | ✅ |
|
||||
|
||||
---
|
||||
|
||||
## 5. 讨论记录
|
||||
|
||||
- 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
|
||||
- 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
|
||||
- 2026-05-14:确认本方案不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
|
||||
- 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先
|
||||
Reference in New Issue
Block a user