- build_review_bundle.py: 新增 _compute_percentile/_compute_growth, 候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy) - term_cleanup_policy.json: 升级 v2 schema - generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions, 规则层输出 alias (大小写/单复数/分词变体) - generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本 (DeepSeek API, 产出 semantic alias/stopword/promote) - SKILL.md: 更新为 5 Phase 工作流程 - 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个 - docs/design/keyword-cleanup-flow-overview.md: 流程文档 - plans/: 引擎设计方案
10 KiB
10 KiB
interest/watch 候选引擎改进方案
从固定阈值到自适应排位 + 趋势因子的演进
1. 背景
1.1 当前实现
build_review_bundle.py 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:
| 候选池 | 判断条件 | 依据 |
|---|---|---|
interest_review_candidates |
total_count >= 3 AND days_seen >= 2 |
policy.interest_keyword_review |
watch_review_candidates |
total_count <= 2 AND days_seen <= 2 |
policy.watch_term_review |
generate_term_cleanup_suggestions.py 则直接从这两个候选池过滤、去重、排序后输出。
1.2 当前方案的问题
问题一:固定阈值不随数据量自适应
场景 total_count=3 意味着什么
─────────────────────────────────────────────
7 天数据(~200 词) top 15%,有一定区分度 ✅
41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变
未来 200 天 仍然用 3 次,区分度稀释 ❌
同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。
问题二:固定阈值忽略趋势信号
- "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
- "Channels":total=3, recent=0 — 早期出现但近期消失
- 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。
问题三:interest 和 watch 的分界线是硬的
total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。
1.3 讨论结论
与老大讨论后确认:
- interest/watch 是统计判断,不需要大模型介入,纯算法可以解决
- 当前引擎缺的不是大模型,而是算法本身没写完——自适应维度(排位、趋势)还没实现
- alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
- 修改量小,可以在 1 小时内落地
2. 设计方案
2.1 核心思路
引入两个互补维度替代固定阈值:
判定维度 含义 数据来源
────────────────────────────────────────────────────────────
percentile(百分位排名) 该词 total_count 在所有词 term_stats
中的排位占比
growth(增速因子) 近期集中度 = recent_count daily 近 N 天
/ total_count
两个维度配合:
- percentile 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
- growth 衡量"这个词是持续出现还是近期爆发"——识别趋势信号
2.2 候选池划分逻辑
percentile
│
┌─────────────────────┐
│ top 5% │
│ → 建议 interest │ ← 高频稳定词
├─────────────────────┤
│ top 5%-20% │
│ → 建议 watch │ ← 有信号但未达 threshold
├─────────────────────┤
│ bottom 80% │
│ → 暂不处理 │ ← 噪声/低频
└─────────────────────┘
额外规则:
如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
→ 主动提升到 watch / 主动推 confirm
这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。
2.3 接口变化
configs/term_cleanup_policy.json:
{
"schema_version": "v2",
"interest_keyword_review": {
"percentile_max": 0.05,
"growth_promotion": 0.5
},
"watch_term_review": {
"percentile_min": 0.05,
"percentile_max": 0.20
}
}
v1 的 min_total_count/min_days_seen 等绝对阈值字段不再使用。
build_review_bundle.py 输出的候选项:
{
"term": "Anthropic",
"total_count": 13,
"days_seen": 10,
"percentile": 0.012,
"growth": 0.54,
"reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
}
2.4 不需要改动的部分
generate_term_cleanup_suggestions.py— 它只消费候选池,不用改apply_term_suggestions.py— 消费 suggestions JSON,不用改keyword-cleanup-bundle.json结构 — 向后兼容,新增 percentile/growth 字段
3. 实施计划
3.1 改动范围
| 文件 | 改动量 | 内容 |
|---|---|---|
skills/keyword-cleanup-review/scripts/build_review_bundle.py |
~40 行 | 新增 _compute_percentile() 和 _compute_growth() 函数;修改候选池生成逻辑;候选项中增加 percentile/growth |
configs/term_cleanup_policy.json |
~10 行 | schema v2:percentile/growth 替代绝对阈值 |
3.2 实施步骤
- build_review_bundle.py:在
top_global_terms生成后,增加 percentile 计算函数和 growth 计算函数 - build_review_bundle.py:修改
interest_review_candidates和watch_review_candidates的生成逻辑,从固定阈值改为 percentile + growth - build_review_bundle.py:候选项增加
percentile和growth字段,更新reason文案 - term_cleanup_policy.json:更新为 v2 schema
- 验证:全量跑一次(
--days 365 --top 100),对比新旧两份输出的差异
3.3 验证方法
# 1. 用旧版生成 baseline
cd /home/ubuntu/zhu/github/reader
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-baseline.json
# 2. 改代码后用新版生成
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 --top 100 \
--output /tmp/bundle-new.json
# 3. 对比 governance_hints
python3 -c "
import json
a = json.load(open('/tmp/bundle-baseline.json'))
b = json.load(open('/tmp/bundle-new.json'))
for key in ['interest_review_candidates', 'watch_review_candidates']:
old = set(i['term'] for i in a['governance_hints'][key])
new = set(i['term'] for i in b['governance_hints'][key])
print(f'{key}: 新增={new-old}, 减少={old-new}')
"
3.4 风险
| 风险 | 概率 | 应对 |
|---|---|---|
| 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 |
| growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 |
| 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 |
4. alias/stopword 设计方案
4.1 核心判断
alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。
| 类型 | 需要什么 | 判断方式 |
|---|---|---|
| 大小写变体 | 表层 | 规则:casefold 去重 |
| 单复数 | 表层 | 规则:去/加 s 后缀匹配 |
| 分词变体(空格/连字符) | 表层 | 规则:去空格归一 |
| 简写全称(MCP→Model Context Protocol) | 语义 | LLM |
| 中英文(上下文工程→Context Engineering) | 语义 | LLM |
| 同义不同名(Rush→猿辅导 Rush 平台) | 语义 | LLM |
| stopword(大模型、AI 太泛) | 语义 | LLM |
4.2 分层方案
输入:高频未覆盖词 + 已有 interest 词表
│
├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
│ 输出候选 alias 对
│
└── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
做语义聚类
输出 alias 组 + stopword 标记
4.3 规则层设计
在 generate_term_cleanup_suggestions.py 中新增 _prepare_alias_suggestions() 函数:
def _prepare_alias_suggestions(top_terms, interest_keywords):
"""
基于表层规则生成 alias 建议。
规则1:casefold 匹配——同一个 casefold 下有多个原文变体
规则2:单复数——去掉/加上末尾 s 后匹配
规则3:分词变体——去空格/连字符后匹配
"""
优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。
4.4 LLM 层设计
单独脚本,非 generate 主链路的一部分。
python scripts/generate_term_cleanup_semantic_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
LLM prompt 设计:
你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
请做三件事:
1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest
用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地
LLM 层输出格式:
{
"alias_suggestions": [
{"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
],
"stopword_suggestions": [
{"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
]
}
4.5 预期效果
| 覆盖类型 | 规则层 | LLM 层 |
|---|---|---|
| 大小写变体 | ✅ | — |
| 单复数 | ✅ | — |
| 分词变体 | ✅ | — |
| 简写全称 | — | ✅ |
| 中英文映射 | — | ✅ |
| 同义不同名 | — | ✅ |
| stopword 判断 | — | ✅ |
5. 讨论记录
- 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
- 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
- 2026-05-14:确认本方案不改
generate_term_cleanup_suggestions.py和apply_term_suggestions.py - 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先