keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions

- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
This commit is contained in:
root
2026-05-14 17:17:49 +08:00
parent 4399c9ca90
commit 590d050218
13 changed files with 1592 additions and 57 deletions
+17
View File
@@ -320,6 +320,23 @@
---
### [DONE][P1] interest/watch 候选引擎从固定阈值改为百分位排名 + 增速因子
目标:
- 解决固定阈值(total_count>=3)不随数据量自适应的问题
- 引入趋势信号(growth 因子),识别近期集中爆发的词
- 支持 7 天、41 天、200 天数据量下取同样的 top 5%/5%-20% 而不需调阈值
要求:
- `build_review_bundle.py`:新增 percentile 和 growth 计算函数;候选池从固定阈值改为百分位 + 增速
- `configs/term_cleanup_policy.json`:升级为 v2 schema,percentile/growth 替代绝对阈值
- 不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
- 全量跑一次对比新旧产出,确认差异合理
方案文档:`plans/keyword-cleanup-interest-watch-engine-improvement.md`
---
### [DONE][P3] 更新 README / handoff / docs,明确 MCP 为正式入口
目标: