103 lines
3.2 KiB
Markdown
103 lines
3.2 KiB
Markdown
---
|
|
name: keyword-cleanup-review
|
|
description: Review and curate this repository's daily keyword index and frequency stats. Use when the user wants to inspect `data/term_index/term_stats.json`, recent `data/term_index/daily/*.json`, `configs/term_aliases.json`, `configs/term_stopwords.json`, or `configs/filter_context.personal.json` to propose alias merges, stopwords, watch terms, or `interest_keywords` updates without directly modifying configs.
|
|
---
|
|
|
|
# Keyword Cleanup Review
|
|
|
|
Use this skill to turn the repository's keyword statistics into reviewable cleanup suggestions.
|
|
|
|
## Workflow
|
|
|
|
1. Build a compact review bundle:
|
|
|
|
```bash
|
|
python skills/keyword-cleanup-review/scripts/build_review_bundle.py
|
|
```
|
|
|
|
Optional knobs:
|
|
|
|
- `--days 7`
|
|
- `--top 50`
|
|
- `--output outputs/term_index/review/keyword-cleanup-bundle.json`
|
|
|
|
2. Read the generated bundle and the suggestion schema:
|
|
|
|
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
|
- `skills/keyword-cleanup-review/references/suggestion-schema.md`
|
|
|
|
3. Produce two outputs:
|
|
|
|
- A short Markdown review for humans
|
|
- A JSON suggestion file matching the schema
|
|
|
|
4. Keep the boundary strict:
|
|
|
|
- Suggest changes to `configs/term_aliases.json`
|
|
- Suggest changes to `configs/term_stopwords.json`
|
|
- Suggest additions to `configs/filter_context.personal.json`
|
|
- Do not directly edit these files unless the user explicitly asks
|
|
- Do not suggest direct edits to `configs/filter_rules.json` unless the user asks for rule logic changes
|
|
|
|
## Review Heuristics
|
|
|
|
Prioritize these decisions:
|
|
|
|
- Alias suggestion
|
|
- Same concept with different naming, casing, abbreviation, or Chinese/English variants
|
|
- Stopword suggestion
|
|
- Too generic, too broad, or too noisy to help filtering
|
|
- Interest keyword suggestion
|
|
- High-frequency and aligned with the user's backend engineering, AI-agent, and frontier-tech focus
|
|
- Watch term
|
|
- Recent and potentially important, but evidence is still weak
|
|
|
|
Prefer conservative suggestions. If confidence is low, put the term into `watch_terms`.
|
|
|
|
## Inputs
|
|
|
|
Primary inputs:
|
|
|
|
- `data/term_index/term_stats.json`
|
|
- `data/term_index/daily/*.json`
|
|
- `configs/term_aliases.json`
|
|
- `configs/term_stopwords.json`
|
|
- `configs/filter_context.personal.json`
|
|
- `configs/term_cleanup_policy.json`
|
|
- `configs/term_watchlist.json`
|
|
- `configs/term_change_log.json`
|
|
|
|
The bundled script already compacts these into a single review bundle.
|
|
|
|
## Output Expectations
|
|
|
|
The Markdown output should:
|
|
|
|
- Summarize the current state briefly
|
|
- List the top terms worth acting on
|
|
- Separate alias, stopword, interest-keyword, and watch-term recommendations
|
|
- Explain reasoning in short, concrete sentences
|
|
|
|
The JSON output should follow:
|
|
|
|
- `references/suggestion-schema.md`
|
|
|
|
## Repository Notes
|
|
|
|
Current repository behavior:
|
|
|
|
- Keyword stats are program-maintained, not LLM-maintained
|
|
- Stats are built from `keywords`, not `topics`
|
|
- Stats only include non-`drop` candidates
|
|
- `data/term_index/term_stats.json` is rebuilt from daily files, so reruns overwrite the same day instead of double-counting
|
|
- cleanup policy, watchlist, and change log are repository-managed governance inputs and should be respected during review
|
|
|
|
Keep suggestions aligned with that design.
|
|
|
|
## Resources
|
|
|
|
- Script:
|
|
- `scripts/build_review_bundle.py`
|
|
- Reference:
|
|
- `references/suggestion-schema.md`
|