Files
reader/plans/keyword-cleanup-interest-watch-engine-improvement.md
root 590d050218 keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions
- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
2026-05-14 17:17:49 +08:00

10 KiB
Raw Permalink Blame History

interest/watch 候选引擎改进方案

从固定阈值到自适应排位 + 趋势因子的演进

1. 背景

1.1 当前实现

build_review_bundle.py 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:

候选池 判断条件 依据
interest_review_candidates total_count >= 3 AND days_seen >= 2 policy.interest_keyword_review
watch_review_candidates total_count <= 2 AND days_seen <= 2 policy.watch_term_review

generate_term_cleanup_suggestions.py 则直接从这两个候选池过滤、去重、排序后输出。

1.2 当前方案的问题

问题一:固定阈值不随数据量自适应

场景                   total_count=3 意味着什么
─────────────────────────────────────────────
7 天数据(~200 词)     top 15%,有一定区分度  ✅
41 天数据(1070 词)    top 5%,区分度更高      ✅ 但阈值没变
未来 200 天             仍然用 3 次,区分度稀释  ❌

同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。

问题二:固定阈值忽略趋势信号

  • "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
  • "Channels":total=3, recent=0 — 早期出现但近期消失
  • 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。

问题三:interest 和 watch 的分界线是硬的

total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。

1.3 讨论结论

与老大讨论后确认:

  1. interest/watch 是统计判断,不需要大模型介入,纯算法可以解决
  2. 当前引擎缺的不是大模型,而是算法本身没写完——自适应维度(排位、趋势)还没实现
  3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
  4. 修改量小,可以在 1 小时内落地

2. 设计方案

2.1 核心思路

引入两个互补维度替代固定阈值:

判定维度                   含义                     数据来源
────────────────────────────────────────────────────────────
percentile(百分位排名)    该词 total_count 在所有词     term_stats
                            中的排位占比
growth(增速因子)          近期集中度 = recent_count      daily 近 N 天
                                                    / total_count

两个维度配合:

  • percentile 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
  • growth 衡量"这个词是持续出现还是近期爆发"——识别趋势信号

2.2 候选池划分逻辑

           percentile
              │
    ┌─────────────────────┐
    │  top 5%             │
    │  → 建议 interest    │ ← 高频稳定词
    ├─────────────────────┤
    │  top 5%-20%         │
    │  → 建议 watch       │ ← 有信号但未达 threshold
    ├─────────────────────┤
    │  bottom 80%          │
    │  → 暂不处理          │ ← 噪声/低频
    └─────────────────────┘

额外规则:
  如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
  → 主动提升到 watch / 主动推 confirm

这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。

2.3 接口变化

configs/term_cleanup_policy.json:

{
  "schema_version": "v2",
  "interest_keyword_review": {
    "percentile_max": 0.05,
    "growth_promotion": 0.5
  },
  "watch_term_review": {
    "percentile_min": 0.05,
    "percentile_max": 0.20
  }
}

v1 的 min_total_count/min_days_seen 等绝对阈值字段不再使用。

build_review_bundle.py 输出的候选项:

{
  "term": "Anthropic",
  "total_count": 13,
  "days_seen": 10,
  "percentile": 0.012,
  "growth": 0.54,
  "reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
}

2.4 不需要改动的部分

  • generate_term_cleanup_suggestions.py — 它只消费候选池,不用改
  • apply_term_suggestions.py — 消费 suggestions JSON,不用改
  • keyword-cleanup-bundle.json 结构 — 向后兼容,新增 percentile/growth 字段

3. 实施计划

3.1 改动范围

文件 改动量 内容
skills/keyword-cleanup-review/scripts/build_review_bundle.py ~40 行 新增 _compute_percentile() 和 _compute_growth() 函数;修改候选池生成逻辑;候选项中增加 percentile/growth
configs/term_cleanup_policy.json ~10 行 schema v2:percentile/growth 替代绝对阈值

3.2 实施步骤

  1. build_review_bundle.py:在 top_global_terms 生成后,增加 percentile 计算函数和 growth 计算函数
  2. build_review_bundle.py:修改 interest_review_candidates 和 watch_review_candidates 的生成逻辑,从固定阈值改为 percentile + growth
  3. build_review_bundle.py:候选项增加 percentile 和 growth 字段,更新 reason 文案
  4. term_cleanup_policy.json:更新为 v2 schema
  5. 验证:全量跑一次(--days 365 --top 100),对比新旧两份输出的差异

3.3 验证方法

# 1. 用旧版生成 baseline
cd /home/ubuntu/zhu/github/reader
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
  --days 365 --top 100 \
  --output /tmp/bundle-baseline.json

# 2. 改代码后用新版生成
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
  --days 365 --top 100 \
  --output /tmp/bundle-new.json

# 3. 对比 governance_hints
python3 -c "
import json
a = json.load(open('/tmp/bundle-baseline.json'))
b = json.load(open('/tmp/bundle-new.json'))
for key in ['interest_review_candidates', 'watch_review_candidates']:
    old = set(i['term'] for i in a['governance_hints'][key])
    new = set(i['term'] for i in b['governance_hints'][key])
    print(f'{key}: 新增={new-old}, 减少={old-new}')
"

3.4 风险

风险 概率 应对
百分位阈值对特小数据集(如只有 1 天数据)不适用 低 不足 7 天时降级回绝对阈值
growth 因子对低频词的偏差(total=1, recent=1 → growth=1) 低 growth 只对 total>=3 的词计算
排位突变导致推荐漂移 低 percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位

4. alias/stopword 设计方案

4.1 核心判断

alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。

类型 需要什么 判断方式
大小写变体 表层 规则:casefold 去重
单复数 表层 规则:去/加 s 后缀匹配
分词变体(空格/连字符) 表层 规则:去空格归一
简写全称(MCP→Model Context Protocol) 语义 LLM
中英文(上下文工程→Context Engineering) 语义 LLM
同义不同名(Rush→猿辅导 Rush 平台) 语义 LLM
stopword(大模型、AI 太泛) 语义 LLM

4.2 分层方案

输入:高频未覆盖词 + 已有 interest 词表
     │
     ├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
     │                     输出候选 alias 对
     │
     └── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
                                     做语义聚类
                                     输出 alias 组 + stopword 标记

4.3 规则层设计

在 generate_term_cleanup_suggestions.py 中新增 _prepare_alias_suggestions() 函数:

def _prepare_alias_suggestions(top_terms, interest_keywords):
    """
    基于表层规则生成 alias 建议。
    规则1:casefold 匹配——同一个 casefold 下有多个原文变体
    规则2:单复数——去掉/加上末尾 s 后匹配
    规则3:分词变体——去空格/连字符后匹配
    """

优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。

4.4 LLM 层设计

单独脚本,非 generate 主链路的一部分。

python scripts/generate_term_cleanup_semantic_suggestions.py \
  --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
  --output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json

LLM prompt 设计:

你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
请做三件事:

1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest

用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地

LLM 层输出格式:

{
  "alias_suggestions": [
    {"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
  ],
  "stopword_suggestions": [
    {"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
  ]
}

4.5 预期效果

覆盖类型 规则层 LLM 层
大小写变体 ✅ —
单复数 ✅ —
分词变体 ✅ —
简写全称 — ✅
中英文映射 — ✅
同义不同名 — ✅
stopword 判断 — ✅

5. 讨论记录

  • 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
  • 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
  • 2026-05-14:确认本方案不改 generate_term_cleanup_suggestions.py 和 apply_term_suggestions.py
  • 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先