--- name: keyword-cleanup-review description: Review and curate this repository's daily keyword index and frequency stats. Use when the user wants to inspect `data/term_index/term_stats.json`, recent `data/term_index/daily/*.json`, `configs/term_aliases.json`, `configs/term_stopwords.json`, or `configs/filter_context.personal.json` to propose alias merges, stopwords, watch terms, or `interest_keywords` updates without directly modifying configs. --- # Keyword Cleanup Review Use this skill to turn the repository's keyword statistics into reviewable cleanup suggestions. ## Workflow 1. Build a compact review bundle: ```bash python skills/keyword-cleanup-review/scripts/build_review_bundle.py ``` Optional knobs: - `--days 7` - `--top 50` - `--output outputs/term_index/review/keyword-cleanup-bundle.json` 2. Read the generated bundle and the suggestion schema: - `outputs/term_index/review/keyword-cleanup-bundle.json` - `skills/keyword-cleanup-review/references/suggestion-schema.md` 3. Produce two outputs: - A short Markdown review for humans - A JSON suggestion file matching the schema 4. Keep the boundary strict: - Suggest changes to `configs/term_aliases.json` - Suggest changes to `configs/term_stopwords.json` - Suggest additions to `configs/filter_context.personal.json` - Do not directly edit these files unless the user explicitly asks - Do not suggest direct edits to `configs/filter_rules.json` unless the user asks for rule logic changes ## Review Heuristics Prioritize these decisions: - Alias suggestion - Same concept with different naming, casing, abbreviation, or Chinese/English variants - Stopword suggestion - Too generic, too broad, or too noisy to help filtering - Interest keyword suggestion - High-frequency and aligned with the user's backend engineering, AI-agent, and frontier-tech focus - Watch term - Recent and potentially important, but evidence is still weak Prefer conservative suggestions. If confidence is low, put the term into `watch_terms`. ## Inputs Primary inputs: - `data/term_index/term_stats.json` - `data/term_index/daily/*.json` - `configs/term_aliases.json` - `configs/term_stopwords.json` - `configs/filter_context.personal.json` - `configs/term_cleanup_policy.json` - `configs/term_watchlist.json` - `configs/term_change_log.json` The bundled script already compacts these into a single review bundle. ## Output Expectations The Markdown output should: - Summarize the current state briefly - List the top terms worth acting on - Separate alias, stopword, interest-keyword, and watch-term recommendations - Explain reasoning in short, concrete sentences The JSON output should follow: - `references/suggestion-schema.md` ## Repository Notes Current repository behavior: - Keyword stats are program-maintained, not LLM-maintained - Stats are built from `keywords`, not `topics` - Stats only include non-`drop` candidates - `data/term_index/term_stats.json` is rebuilt from daily files, so reruns overwrite the same day instead of double-counting - cleanup policy, watchlist, and change log are repository-managed governance inputs and should be respected during review Keep suggestions aligned with that design. ## Resources - Script: - `scripts/build_review_bundle.py` - Reference: - `references/suggestion-schema.md`