keyword cleanup: v2 engine, alias rule layer, LLM semantic suggestions

- build_review_bundle.py: 新增 _compute_percentile/_compute_growth,
  候选池从固定阈值改为百分位排名 + 增速因子 (v2 policy)
- term_cleanup_policy.json: 升级 v2 schema
- generate_term_cleanup_suggestions.py: 新增 _prepare_alias_suggestions,
  规则层输出 alias (大小写/单复数/分词变体)
- generate_term_cleanup_semantic_suggestions.py: 新增 LLM 语义建议脚本
  (DeepSeek API, 产出 semantic alias/stopword/promote)
- SKILL.md: 更新为 5 Phase 工作流程
- 首轮清洗 apply: interest 54, aliases 17组, stopwords 17个
- docs/design/keyword-cleanup-flow-overview.md: 流程文档
- plans/: 引擎设计方案
This commit is contained in:
root
2026-05-14 17:17:49 +08:00
parent 4399c9ca90
commit 590d050218
13 changed files with 1592 additions and 57 deletions
+68 -24
View File
@@ -26,7 +26,7 @@ description: 生成 reader 项目的正式关键词 review 输入。当用户需
## 工作流程
1. 构建精简的审查数据包(临时工作文件):
### Phase 1:构建审查数据包
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py
@@ -34,51 +34,92 @@ python skills/keyword-cleanup-review/scripts/build_review_bundle.py
可选参数:
- `--days 7`
- `--top 50`
- `--days 7`(默认 7,建议传 365 覆盖全量)
- `--top 100`(考虑的词数)
- `--output outputs/term_index/review/keyword-cleanup-bundle.json`
2. 阅读建议模式:
#### 候选引擎策略
- `skills/keyword-cleanup-review/references/suggestion-schema.md`
根据 `configs/term_cleanup_policy.json` 的 `schema_version` 自动切换:
3. 运行 suggestions 生成脚本:
| 版本 | 策略 | 说明 |
|------|------|------|
| v1(旧) | 固定阈值(total≥3/days≥2 → interest) | 小数据集兼容 |
| v2(当前默认) | 百分位排名 + 增速因子 | 自适应数据量,不需要手工调阈值 |
v2 策略说明:
- **percentile**:total_count 在所有词里的排位占比。top 5% → interest 候选,5%-20% → watch 候选
- **growth**:recent_count / total_count,衡量近期活跃度。growth≥0.5 的排位外词也会主动推荐
### Phase 2:生成建议(规则层)
```bash
python scripts/generate_term_cleanup_suggestions.py ^
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
```
默认生成:
- 一份符合模式的 JSON 建议文件(正式建议产物,也是 review / apply 之间唯一正式输入)
如需人工审阅展示稿,再显式加:
如需人工审阅展示稿:
```bash
python scripts/generate_term_cleanup_suggestions.py ^
--bundle outputs/term_index/review/keyword-cleanup-bundle.json ^
python scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
这时才会额外生成:
#### 产出能力
- 一份简短的供人工审阅的 Markdown 报告(临时展示稿)
| 建议类型 | 状态 | 方法 |
|---------|------|------|
| interest 建议 | ✅ 已实现 | 百分位 top 5% + 增速促活 |
| watch 建议 | ✅ 已实现 | 百分位 5%-20% |
| alias 建议 | ✅ 已实现 | 规则层:大小写归一、单复数、去空格/连字符 |
| stopword 建议 | ❌ 规则层空缺 | 见 Phase 3(LLM 层) |
4. 严格保持边界:
默认生成:
- `term-cleanup-suggestions-YYYY-MM-DD.json`(正式建议产物)
显式加 `--emit-markdown` 额外生成:
- `term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿)
### Phase 3:生成建议(LLM 层,可选)
规则层覆盖不了 alias(中英文对应、缩写展开、同义不同名)和 stopword 判断,需要 LLM 辅助:
```bash
python scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
从 `.env` 读取 LLM 配置(`LLM_API_URL` / `LLM_MODEL` / `LLM_API_KEY`),使用 DeepSeek API。
输出三部分:
| 输出 | 说明 |
|------|------|
| `semantic_alias` | 语义级别名(中英文、缩写、同义不同名) |
| `stopword` | 泛词过滤建议(规则层做不了的需要语义判断的) |
| `promote_to_interest` | 与用户关注方向一致的新词,建议加入 interest |
**注:LLM 层产物是候选,不应自动 apply,需要人工确认后由 OpenClaw 编排 apply。**
### Phase 4:输出给 OpenClaw 编排
- `suggestions JSON` = review / apply 之间唯一正式建议输入
- `semantic-suggestions JSON` = LLM 补充建议,需要人工筛选后合并到 suggestions JSON 再 apply
- Markdown = 临时展示层
- 后续汇报、确认、dry-run、apply、收尾清理由 OpenClaw 编排层执行
### Phase 5:严格保持边界
- 建议 `configs/term_aliases.json` 的修改
- 建议 `configs/term_stopwords.json` 的修改
- 建议 `configs/filter_context.personal.json` 的新增
- **LLM 层产出(semantic-suggestions)不自动 apply**,需人工确认后由 OpenClaw 编排层执行
- 除非用户明确要求,否则不要直接编辑这些文件
- 除非用户要求修改规则逻辑,否则不要建议直接编辑 `configs/filter_rules.json`
5. 输出交接口径:
- 将 JSON suggestions 视为正式 review 输入
- 将 Markdown 视为可选展示层
- 后续汇报、确认、dry-run apply、正式 apply、收尾清理应由 OpenClaw 编排层继续执行
## 审查启发式规则
优先考虑以下决策:
@@ -141,6 +182,7 @@ JSON 输出应遵循:
短期保留:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
- `outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json`
临时产物:
@@ -173,7 +215,9 @@ JSON 输出应遵循:
## 资源
- 脚本:
- `scripts/build_review_bundle.py`
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- `scripts/generate_term_cleanup_suggestions.py`
- `scripts/generate_term_cleanup_semantic_suggestions.py`(LLM 层)
- 参考文档:
- `references/suggestion-schema.md`
- `plans/keyword-cleanup-interest-watch-engine-improvement.md`(v2 引擎设计)
@@ -9,16 +9,15 @@ from typing import Any
DEFAULT_POLICY: dict[str, Any] = {
"schema_version": "v1",
"schema_version": "v2",
"interest_keyword_review": {
"min_total_count": 3,
"min_days_seen": 2,
"percentile_min": 0.0,
"percentile_max": 0.05,
"growth_promotion": 0.5,
},
"watch_term_review": {
"min_total_count": 1,
"min_days_seen": 1,
"max_total_count": 2,
"max_days_seen": 2,
"percentile_min": 0.05,
"percentile_max": 0.20,
},
"alias_review": {
"min_total_count": 2,
@@ -28,6 +27,11 @@ DEFAULT_POLICY: dict[str, Any] = {
"max_total_count": 2,
"max_days_seen": 2,
},
"notes": [
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
"growth = recent_count / total_count,衡量近期活跃度",
],
}
@@ -126,6 +130,37 @@ def _within_watch_thresholds(item: dict[str, Any], thresholds: dict[str, Any]) -
)
def _compute_percentile(value: int, sorted_values: list[int]) -> float:
"""
Return the percentile rank of `value` in `sorted_values` (ascending).
0.0 = highest frequency (top rank), 1.0 = lowest frequency (bottom rank).
"""
if not sorted_values:
return 1.0
# bisect_left — count of values strictly less than `value`
lo, hi = 0, len(sorted_values)
while lo < hi:
mid = (lo + hi) // 2
if sorted_values[mid] < value:
lo = mid + 1
else:
hi = mid
rank = lo
# invert: smallest value → rank=0 → 1.0 (bottom)
# largest value → rank=len → 0.0 (top)
return 1.0 - (rank / len(sorted_values))
def _compute_growth(recent_count: int, total_count: int) -> float:
"""
Return growth factor: recent_count / total_count.
Only meaningful when total_count >= 3; returns 0.0 for small counts.
"""
if total_count < 3:
return 0.0
return recent_count / total_count
def main() -> None:
parser = argparse.ArgumentParser(
description="Build a compact review bundle for the keyword-cleanup-review skill."
@@ -235,6 +270,13 @@ def main() -> None:
alias_values = _casefold_set(list(aliases.values()))
watch_set = _casefold_set([str(item.get("term", "")) for item in watchlist])
# Build a sorted list of all total_counts for percentile computation
all_total_counts = sorted(
int(item.get("total_count") or 0)
for item in stats_terms
if isinstance(item, dict) and isinstance(item.get("term"), str)
)
top_global_terms = []
for item in stats_terms[: args.top]:
if not isinstance(item, dict):
@@ -256,42 +298,108 @@ def main() -> None:
"is_alias_target": folded in alias_values,
"in_watchlist": folded in watch_set,
"recent_count": recent_counter.get(term, 0),
"percentile": _compute_percentile(
int(item.get("total_count") or 0), all_total_counts
),
"growth": _compute_growth(
recent_counter.get(term, 0),
int(item.get("total_count") or 0),
),
}
)
# Keep more uncovered terms for percentile-based selection
uncovered_terms = [
item for item in top_global_terms if not item["in_interest_keywords"] and not item["is_stopword"]
][:20]
][:100]
policy_version = (policy.get("schema_version") if isinstance(policy, dict) else None) or "v1"
interest_thresholds = policy.get("interest_keyword_review") if isinstance(policy, dict) else {}
watch_thresholds = policy.get("watch_term_review") if isinstance(policy, dict) else {}
if policy_version == "v2" or "percentile_max" in interest_thresholds:
# v2: percentile + growth based selection
pct_min_interest = float(interest_thresholds.get("percentile_min", 0.0))
pct_max_interest = float(interest_thresholds.get("percentile_max", 0.05))
growth_promo = float(interest_thresholds.get("growth_promotion", 0.5))
pct_min_watch = float(watch_thresholds.get("percentile_min", 0.05))
pct_max_watch = float(watch_thresholds.get("percentile_max", 0.20))
interest_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"]
and pct_min_interest <= item["percentile"] <= pct_max_interest
]
watch_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"]
and pct_min_watch < item["percentile"] <= pct_max_watch
]
# Growth boost: terms outside watch range but with strong growth signal
growth_boost_candidates = [
item for item in uncovered_terms
if not item["in_watchlist"]
and item["percentile"] > pct_max_watch
and item["growth"] >= growth_promo
]
else:
# v1 fallback: fixed thresholds
interest_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
]
watch_candidates_raw = [
item for item in uncovered_terms
if not item["in_watchlist"]
and not _meets_min_thresholds(item, interest_thresholds)
and _within_watch_thresholds(item, watch_thresholds)
]
growth_boost_candidates = []
interest_review_candidates = [
{
"term": item["term"],
"total_count": item["total_count"],
"days_seen": item["days_seen"],
"percentile": item["percentile"],
"growth": item["growth"],
"reason": (
"Meets the configured interest-keyword review threshold and is not yet covered "
"by interest keywords or stopwords."
f"top {item['percentile']:.1%} by frequency,"
f"growth={item['growth']:.0%},"
"not yet covered by interest keywords or stopwords."
),
}
for item in uncovered_terms
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
for item in interest_candidates_raw
][:20]
watch_review_candidates = [
{
"term": item["term"],
"total_count": item["total_count"],
"days_seen": item["days_seen"],
"percentile": item["percentile"],
"growth": item["growth"],
"reason": (
"Falls into the configured watch-term review range and should be observed "
"before promotion into interest keywords."
f"top {item['percentile']:.1%} by frequency,"
f"growth={item['growth']:.0%},"
"fell into watch-review range."
),
}
for item in uncovered_terms
if not item["in_watchlist"]
and not _meets_min_thresholds(item, interest_thresholds)
and _within_watch_thresholds(item, watch_thresholds)
for item in watch_candidates_raw
][:20]
growth_boost_review_items = [
{
"term": item["term"],
"total_count": item["total_count"],
"days_seen": item["days_seen"],
"percentile": item["percentile"],
"growth": item["growth"],
"reason": (
f"growth spike: {item['growth']:.0%} of occurrences in recent window "
f"(total={item['total_count']}, days={item['days_seen']})."
),
}
for item in growth_boost_candidates
][:5]
recent_hot_terms = sorted(
({"term": term, "recent_count": count} for term, count in recent_counter.items()),
key=lambda item: (-item["recent_count"], item["term"].casefold(), item["term"]),
@@ -330,6 +438,7 @@ def main() -> None:
"governance_hints": {
"interest_review_candidates": interest_review_candidates,
"watch_review_candidates": watch_review_candidates,
"growth_boost_review_items": growth_boost_review_items,
},
}
_save_json(args.output, bundle)