Files
reader/docs/design/keyword-cleanup-flow-overview.md
root 590d050218 keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions
- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
2026-05-14 17:17:49 +08:00

152 lines
5.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 关键词清洗流程概述
> 2026-05-14 初版
> 从"数据记录"到"人工确认落盘"的完整链路
---
## 整体数据流
```
每日日报 pipeline
│
▼
term_index/daily/YYYY-MM-DD.json ← 每天一篇候选文章的热词统计
│
▼
term_index/term_stats.json ← 所有 daily 的汇总(1070 个词)
│
├──── build_review_bundle.py ← 打包为审查数据包
│ │
│ ▼
│ review/keyword-cleanup-bundle.json
│ │
│ ▼
│ generate_term_cleanup_suggestions.py
│ │
│ ▼
│ review/term-cleanup-suggestions-YYYY-MM-DD.json ← 正式建议产物
│ │
│ ▼
│ (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md ← 展示稿
│
├──── 人工确认哪些建议 accept
│
▼
apply_term_suggestions.py ← 写入配置
│
├── configs/filter_context.personal.json ← interest_keywords
├── configs/term_aliases.json ← alias
├── configs/term_stopwords.json ← stopword
├── configs/term_watchlist.json ← watch
└── configs/term_change_log.json ← 变更日志
```
---
## 各环节说明
### 阶段 1:数据记录(每日自动)
```bash
# FreshRSS pipeline 跑完后自动产出
data/term_index/daily/2026-05-14.json
```
- 每天一篇,记录当天候选文章中出现的热词
- 包含 term、total_count、days_seen 等信息
- 目前累计 **41 天**,共 **1070 个独立词**
### 阶段 2:全量汇总(每日自动)
```bash
data/term_index/term_stats.json
```
- 从所有 daily 文件重建,会覆盖重跑
- 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13)
### 阶段 3:构建审查数据包(手动触发)
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 \
--top 100 \
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
- 把 term_stats + 当前配置打成一包,方便后续处理
- 输出:`review/keyword-cleanup-bundle.json`
### 阶段 4:生成建议(手动触发)
```bash
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
#### 当前产出能力
| 建议类型 | 状态 | 当前阈值 | 说明 |
|---------|------|----------|------|
| interest_keyword_suggestions | ✅ **已实现** | total≥3, days≥2 | 产出 20 条 |
| watch_terms | ✅ **已实现** | total≤2, days≤2 | 本次 0 条 |
| alias_suggestions | ❌ **硬编码为空** | policy 有阈值(total≥2, days≥2)但脚本未实现 | |
| stopword_suggestions | ❌ **硬编码为空** | policy 有阈值(total≤2, days≤2)但脚本未实现 | |
**关键发现:** alias 和 stopword 不是"阈值太保守",是 **generate 脚本里压根没写对应的生成函数**。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 `[]` 和 `0`。
### 阶段 5:人工确认(手动)
```
OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply
```
本次模式:
- 高频(≥5次/5天以上)→ 强烈推荐 ✅
- 中频(3-4次)→ 附带建议 ✅
- 泛词 → 建议跳过 ❌
### 阶段 6:落盘配置(手动)
```bash
python scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 词1 词2 ...
```
- dry-run 预览 → 确认后正式 apply
- 写入 `configs/filter_context.personal.json`
- 同步记录到 `term_change_log.json`
- **不备份原始配置**(待优化)
- **apply 后不自动清理 review 目录**(待优化)
### 阶段 7:维护清理(按需)
由 OpenClaw 侧 `reader-keyword-maintenance` skill 处理:
- 删除旧 markdown 展示稿
- 保留最近一份 bundle
- 保守保留 suggestions JSON
---
## 当前配置资产
| 文件 | 内容 | 数据量 |
|------|------|--------|
| `filter_context.personal.json` | interest_keywords | 52 个 |
| `term_aliases.json` | 别名映射 | 0 组(未启用) |
| `term_stopwords.json` | 停用词 | 0 个(未启用) |
| `term_watchlist.json` | 观察词 | 6 个 |
| `term_change_log.json` | 所有变更记录 | 已记录 |
---
## 待优化项
1. **alias/stopword 建议生成为空** — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数
2. **apply 前无配置备份** — 建议 apply 前自动 cp 备份
3. **apply 后无自动收尾** — 建议 apply 后自动删旧 markdown 和 bundle
4. **alias 识别依赖规则而非 LLM** — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply