- build_review_bundle.py: 新增 _compute_percentile/_compute_growth, 候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy) - term_cleanup_policy.json: 升级 v2 schema - generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions, 规则层输出 alias (大小写/单复数/分词变体) - generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本 (DeepSeek API, 产出 semantic alias/stopword/promote) - SKILL.md: 更新为 5 Phase 工作流程 - 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个 - docs/design/keyword-cleanup-flow-overview.md: 流程文档 - plans/: 引擎设计方案
152 lines
5.0 KiB
Markdown
152 lines
5.0 KiB
Markdown
# 关键词清洗流程概述
|
||
|
||
> 2026-05-14 初版
|
||
> 从"数据记录"到"人工确认落盘"的完整链路
|
||
|
||
---
|
||
|
||
## 整体数据流
|
||
|
||
```
|
||
每日日报 pipeline
|
||
│
|
||
▼
|
||
term_index/daily/YYYY-MM-DD.json ← 每天一篇候选文章的热词统计
|
||
│
|
||
▼
|
||
term_index/term_stats.json ← 所有 daily 的汇总(1070 个词)
|
||
│
|
||
├──── build_review_bundle.py ← 打包为审查数据包
|
||
│ │
|
||
│ ▼
|
||
│ review/keyword-cleanup-bundle.json
|
||
│ │
|
||
│ ▼
|
||
│ generate_term_cleanup_suggestions.py
|
||
│ │
|
||
│ ▼
|
||
│ review/term-cleanup-suggestions-YYYY-MM-DD.json ← 正式建议产物
|
||
│ │
|
||
│ ▼
|
||
│ (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md ← 展示稿
|
||
│
|
||
├──── 人工确认哪些建议 accept
|
||
│
|
||
▼
|
||
apply_term_suggestions.py ← 写入配置
|
||
│
|
||
├── configs/filter_context.personal.json ← interest_keywords
|
||
├── configs/term_aliases.json ← alias
|
||
├── configs/term_stopwords.json ← stopword
|
||
├── configs/term_watchlist.json ← watch
|
||
└── configs/term_change_log.json ← 变更日志
|
||
```
|
||
|
||
---
|
||
|
||
## 各环节说明
|
||
|
||
### 阶段 1:数据记录(每日自动)
|
||
|
||
```bash
|
||
# FreshRSS pipeline 跑完后自动产出
|
||
data/term_index/daily/2026-05-14.json
|
||
```
|
||
|
||
- 每天一篇,记录当天候选文章中出现的热词
|
||
- 包含 term、total_count、days_seen 等信息
|
||
- 目前累计 **41 天**,共 **1070 个独立词**
|
||
|
||
### 阶段 2:全量汇总(每日自动)
|
||
|
||
```bash
|
||
data/term_index/term_stats.json
|
||
```
|
||
|
||
- 从所有 daily 文件重建,会覆盖重跑
|
||
- 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13)
|
||
|
||
### 阶段 3:构建审查数据包(手动触发)
|
||
|
||
```bash
|
||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||
--days 365 \
|
||
--top 100 \
|
||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||
```
|
||
|
||
- 把 term_stats + 当前配置打成一包,方便后续处理
|
||
- 输出:`review/keyword-cleanup-bundle.json`
|
||
|
||
### 阶段 4:生成建议(手动触发)
|
||
|
||
```bash
|
||
python scripts/generate_term_cleanup_suggestions.py \
|
||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||
--emit-markdown
|
||
```
|
||
|
||
#### 当前产出能力
|
||
|
||
| 建议类型 | 状态 | 当前阈值 | 说明 |
|
||
|---------|------|----------|------|
|
||
| interest_keyword_suggestions | ✅ **已实现** | total≥3, days≥2 | 产出 20 条 |
|
||
| watch_terms | ✅ **已实现** | total≤2, days≤2 | 本次 0 条 |
|
||
| alias_suggestions | ❌ **硬编码为空** | policy 有阈值(total≥2, days≥2)但脚本未实现 | |
|
||
| stopword_suggestions | ❌ **硬编码为空** | policy 有阈值(total≤2, days≤2)但脚本未实现 | |
|
||
|
||
**关键发现:** alias 和 stopword 不是"阈值太保守",是 **generate 脚本里压根没写对应的生成函数**。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 `[]` 和 `0`。
|
||
|
||
### 阶段 5:人工确认(手动)
|
||
|
||
```
|
||
OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply
|
||
```
|
||
|
||
本次模式:
|
||
- 高频(≥5次/5天以上)→ 强烈推荐 ✅
|
||
- 中频(3-4次)→ 附带建议 ✅
|
||
- 泛词 → 建议跳过 ❌
|
||
|
||
### 阶段 6:落盘配置(手动)
|
||
|
||
```bash
|
||
python scripts/apply_term_suggestions.py \
|
||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||
--accept-interest 词1 词2 ...
|
||
```
|
||
|
||
- dry-run 预览 → 确认后正式 apply
|
||
- 写入 `configs/filter_context.personal.json`
|
||
- 同步记录到 `term_change_log.json`
|
||
- **不备份原始配置**(待优化)
|
||
- **apply 后不自动清理 review 目录**(待优化)
|
||
|
||
### 阶段 7:维护清理(按需)
|
||
|
||
由 OpenClaw 侧 `reader-keyword-maintenance` skill 处理:
|
||
- 删除旧 markdown 展示稿
|
||
- 保留最近一份 bundle
|
||
- 保守保留 suggestions JSON
|
||
|
||
---
|
||
|
||
## 当前配置资产
|
||
|
||
| 文件 | 内容 | 数据量 |
|
||
|------|------|--------|
|
||
| `filter_context.personal.json` | interest_keywords | 52 个 |
|
||
| `term_aliases.json` | 别名映射 | 0 组(未启用) |
|
||
| `term_stopwords.json` | 停用词 | 0 个(未启用) |
|
||
| `term_watchlist.json` | 观察词 | 6 个 |
|
||
| `term_change_log.json` | 所有变更记录 | 已记录 |
|
||
|
||
---
|
||
|
||
## 待优化项
|
||
|
||
1. **alias/stopword 建议生成为空** — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数
|
||
2. **apply 前无配置备份** — 建议 apply 前自动 cp 备份
|
||
3. **apply 后无自动收尾** — 建议 apply 后自动删旧 markdown 和 bundle
|
||
4. **alias 识别依赖规则而非 LLM** — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply
|