# interest/watch 候选引擎改进方案 > 从固定阈值到自适应排位 + 趋势因子的演进 ## 1. 背景 ### 1.1 当前实现 `build_review_bundle.py` 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池: | 候选池 | 判断条件 | 依据 | |--------|---------|------| | `interest_review_candidates` | `total_count >= 3 AND days_seen >= 2` | `policy.interest_keyword_review` | | `watch_review_candidates` | `total_count <= 2 AND days_seen <= 2` | `policy.watch_term_review` | `generate_term_cleanup_suggestions.py` 则直接从这两个候选池过滤、去重、排序后输出。 ### 1.2 当前方案的问题 **问题一:固定阈值不随数据量自适应** ``` 场景 total_count=3 意味着什么 ───────────────────────────────────────────── 7 天数据(~200 词) top 15%,有一定区分度 ✅ 41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变 未来 200 天 仍然用 3 次,区分度稀释 ❌ ``` 同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。 **问题二:固定阈值忽略趋势信号** - "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势 - "Channels":total=3, recent=0 — 早期出现但近期消失 - 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。 **问题三:interest 和 watch 的分界线是硬的** total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。 ### 1.3 讨论结论 与老大讨论后确认: 1. interest/watch 是**统计判断**,不需要大模型介入,纯算法可以解决 2. 当前引擎缺的不是大模型,而是**算法本身没写完**——自适应维度(排位、趋势)还没实现 3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内 4. 修改量小,可以在 1 小时内落地 --- ## 2. 设计方案 ### 2.1 核心思路 引入两个互补维度替代固定阈值: ``` 判定维度 含义 数据来源 ──────────────────────────────────────────────────────────── percentile(百分位排名) 该词 total_count 在所有词 term_stats 中的排位占比 growth(增速因子) 近期集中度 = recent_count daily 近 N 天 / total_count ``` 两个维度配合: - **percentile** 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响 - **growth** 衡量"这个词是持续出现还是近期爆发"——识别趋势信号 ### 2.2 候选池划分逻辑 ``` percentile │ ┌─────────────────────┐ │ top 5% │ │ → 建议 interest │ ← 高频稳定词 ├─────────────────────┤ │ top 5%-20% │ │ → 建议 watch │ ← 有信号但未达 threshold ├─────────────────────┤ │ bottom 80% │ │ → 暂不处理 │ ← 噪声/低频 └─────────────────────┘ 额外规则: 如果词在 top 20% 之外,但 growth > 0.5(近期集中度高) → 主动提升到 watch / 主动推 confirm ``` 这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。 ### 2.3 接口变化 **`configs/term_cleanup_policy.json`**: ```json { "schema_version": "v2", "interest_keyword_review": { "percentile_max": 0.05, "growth_promotion": 0.5 }, "watch_term_review": { "percentile_min": 0.05, "percentile_max": 0.20 } } ``` `v1` 的 `min_total_count`/`min_days_seen` 等绝对阈值字段不再使用。 **`build_review_bundle.py` 输出的候选项**: ```json { "term": "Anthropic", "total_count": 13, "days_seen": 10, "percentile": 0.012, "growth": 0.54, "reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal." } ``` ### 2.4 不需要改动的部分 - `generate_term_cleanup_suggestions.py` — 它只消费候选池,不用改 - `apply_term_suggestions.py` — 消费 suggestions JSON,不用改 - `keyword-cleanup-bundle.json` 结构 — 向后兼容,新增 percentile/growth 字段 --- ## 3. 实施计划 ### 3.1 改动范围 | 文件 | 改动量 | 内容 | |------|--------|------| | `skills/keyword-cleanup-review/scripts/build_review_bundle.py` | ~40 行 | 新增 `_compute_percentile()` 和 `_compute_growth()` 函数;修改候选池生成逻辑;候选项中增加 percentile/growth | | `configs/term_cleanup_policy.json` | ~10 行 | schema v2:percentile/growth 替代绝对阈值 | ### 3.2 实施步骤 1. **build_review_bundle.py**:在 `top_global_terms` 生成后,增加 percentile 计算函数和 growth 计算函数 2. **build_review_bundle.py**:修改 `interest_review_candidates` 和 `watch_review_candidates` 的生成逻辑,从固定阈值改为 percentile + growth 3. **build_review_bundle.py**:候选项增加 `percentile` 和 `growth` 字段,更新 `reason` 文案 4. **term_cleanup_policy.json**:更新为 v2 schema 5. **验证**:全量跑一次(`--days 365 --top 100`),对比新旧两份输出的差异 ### 3.3 验证方法 ```bash # 1. 用旧版生成 baseline cd /home/ubuntu/zhu/github/reader python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \ --days 365 --top 100 \ --output /tmp/bundle-baseline.json # 2. 改代码后用新版生成 python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \ --days 365 --top 100 \ --output /tmp/bundle-new.json # 3. 对比 governance_hints python3 -c " import json a = json.load(open('/tmp/bundle-baseline.json')) b = json.load(open('/tmp/bundle-new.json')) for key in ['interest_review_candidates', 'watch_review_candidates']: old = set(i['term'] for i in a['governance_hints'][key]) new = set(i['term'] for i in b['governance_hints'][key]) print(f'{key}: 新增={new-old}, 减少={old-new}') " ``` ### 3.4 风险 | 风险 | 概率 | 应对 | |------|------|------| | 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 | | growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 | | 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 | --- ## 4. alias/stopword 设计方案 ### 4.1 核心判断 alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。 | 类型 | 需要什么 | 判断方式 | |------|---------|----------| | 大小写变体 | 表层 | 规则:casefold 去重 | | 单复数 | 表层 | 规则:去/加 s 后缀匹配 | | 分词变体(空格/连字符) | 表层 | 规则:去空格归一 | | 简写全称(MCP→Model Context Protocol) | **语义** | LLM | | 中英文(上下文工程→Context Engineering) | **语义** | LLM | | 同义不同名(Rush→猿辅导 Rush 平台) | **语义** | LLM | | stopword(大模型、AI 太泛) | **语义** | LLM | ### 4.2 分层方案 ``` 输入:高频未覆盖词 + 已有 interest 词表 │ ├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符 │ 输出候选 alias 对 │ └── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM 做语义聚类 输出 alias 组 + stopword 标记 ``` ### 4.3 规则层设计 在 `generate_term_cleanup_suggestions.py` 中新增 `_prepare_alias_suggestions()` 函数: ```python def _prepare_alias_suggestions(top_terms, interest_keywords): """ 基于表层规则生成 alias 建议。 规则1:casefold 匹配——同一个 casefold 下有多个原文变体 规则2:单复数——去掉/加上末尾 s 后匹配 规则3:分词变体——去空格/连字符后匹配 """ ``` 优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。 ### 4.4 LLM 层设计 单独脚本,非 generate 主链路的一部分。 ```bash python scripts/generate_term_cleanup_semantic_suggestions.py \ --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \ --output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json ``` LLM prompt 设计: ``` 你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。 请做三件事: 1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体 2. STOPWORD:标记哪些词太宽泛/通用,建议排除 3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest 用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地 ``` LLM 层输出格式: ```json { "alias_suggestions": [ {"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"} ], "stopword_suggestions": [ {"term": "大模型", "reason": "过于宽泛,高频率但低区分度"} ] } ``` ### 4.5 预期效果 | 覆盖类型 | 规则层 | LLM 层 | |---------|--------|--------| | 大小写变体 | ✅ | — | | 单复数 | ✅ | — | | 分词变体 | ✅ | — | | 简写全称 | — | ✅ | | 中英文映射 | — | ✅ | | 同义不同名 | — | ✅ | | stopword 判断 | — | ✅ | --- ## 5. 讨论记录 - 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决 - 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地 - 2026-05-14:确认本方案不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py` - 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先