Files
reader/skills/keyword-cleanup-review/SKILL.md
T

3.2 KiB

name, description
name description
keyword-cleanup-review Review and curate this repository's daily keyword index and frequency stats. Use when the user wants to inspect `data/term_index/term_stats.json`, recent `data/term_index/daily/*.json`, `configs/term_aliases.json`, `configs/term_stopwords.json`, or `configs/filter_context.personal.json` to propose alias merges, stopwords, watch terms, or `interest_keywords` updates without directly modifying configs.

Keyword Cleanup Review

Use this skill to turn the repository's keyword statistics into reviewable cleanup suggestions.

Workflow

  1. Build a compact review bundle:
python skills/keyword-cleanup-review/scripts/build_review_bundle.py

Optional knobs:

  • --days 7
  • --top 50
  • --output outputs/term_index/review/keyword-cleanup-bundle.json
  1. Read the generated bundle and the suggestion schema:
  • outputs/term_index/review/keyword-cleanup-bundle.json
  • skills/keyword-cleanup-review/references/suggestion-schema.md
  1. Produce two outputs:
  • A short Markdown review for humans
  • A JSON suggestion file matching the schema
  1. Keep the boundary strict:
  • Suggest changes to configs/term_aliases.json
  • Suggest changes to configs/term_stopwords.json
  • Suggest additions to configs/filter_context.personal.json
  • Do not directly edit these files unless the user explicitly asks
  • Do not suggest direct edits to configs/filter_rules.json unless the user asks for rule logic changes

Review Heuristics

Prioritize these decisions:

  • Alias suggestion
    • Same concept with different naming, casing, abbreviation, or Chinese/English variants
  • Stopword suggestion
    • Too generic, too broad, or too noisy to help filtering
  • Interest keyword suggestion
    • High-frequency and aligned with the user's backend engineering, AI-agent, and frontier-tech focus
  • Watch term
    • Recent and potentially important, but evidence is still weak

Prefer conservative suggestions. If confidence is low, put the term into watch_terms.

Inputs

Primary inputs:

  • data/term_index/term_stats.json
  • data/term_index/daily/*.json
  • configs/term_aliases.json
  • configs/term_stopwords.json
  • configs/filter_context.personal.json
  • configs/term_cleanup_policy.json
  • configs/term_watchlist.json
  • configs/term_change_log.json

The bundled script already compacts these into a single review bundle.

Output Expectations

The Markdown output should:

  • Summarize the current state briefly
  • List the top terms worth acting on
  • Separate alias, stopword, interest-keyword, and watch-term recommendations
  • Explain reasoning in short, concrete sentences

The JSON output should follow:

  • references/suggestion-schema.md

Repository Notes

Current repository behavior:

  • Keyword stats are program-maintained, not LLM-maintained
  • Stats are built from keywords, not topics
  • Stats only include non-drop candidates
  • data/term_index/term_stats.json is rebuilt from daily files, so reruns overwrite the same day instead of double-counting
  • cleanup policy, watchlist, and change log are repository-managed governance inputs and should be respected during review

Keep suggestions aligned with that design.

Resources

  • Script:
    • scripts/build_review_bundle.py
  • Reference:
    • references/suggestion-schema.md