Add keyword cleanup governance workflow

This commit is contained in:
zhuyongxin
2026-03-27 16:59:52 +08:00
parent 100044e1f7
commit 32ad584348
25 changed files with 1915 additions and 18 deletions
+54
View File
@@ -92,6 +92,11 @@ This is the recommended production entrypoint. By default it writes only:
- `outputs/freshrss/rerun/<timestamp>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<timestamp>/run-report.json`
It also updates the daily keyword index runtime data:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
If you need per-item intermediates, add `--debug-artifacts`.
When OpenClaw is connected to the MCP server, it should call `run_freshrss_openclaw_pipeline` for the same behavior directly through MCP. The tool also supports `debug_artifacts=true` when deeper inspection is needed.
@@ -160,3 +165,52 @@ The script writes by default:
- `outputs/reference/candidates/openclaw-delivery-payload.json`
Output layout details live in `outputs/README.md`.
Keyword index defaults live in:
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
You can also rebuild the keyword index from an existing delivery payload:
```bash
python scripts/build_keyword_index.py ^
--input outputs/reference/candidates/openclaw-delivery-payload.json
```
Runtime keyword data is stored under `data/term_index/`.
The keyword cleanup review skill lives in:
- `skills/keyword-cleanup-review/`
To build a review bundle for the LLM skill:
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
--days 7 ^
--top 50 ^
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
The review bundle now also carries cleanup governance context:
- cleanup thresholds from `configs/term_cleanup_policy.json`
- the current watch list from `configs/term_watchlist.json`
- recent applied changes from `configs/term_change_log.json`
The skill only produces review inputs and suggestions. It does not modify `term_aliases`, `term_stopwords`, or `filter_context.personal.json` automatically.
To preview accepted suggestions before writing any config files:
```bash
python scripts/apply_term_suggestions.py ^
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
--accept-watch Cron Heartbeat Memory ^
--dry-run
```
Remove `--dry-run` to write the accepted changes. The script can also apply accepted `alias`, `stopword`, and `interest keyword` suggestions through `--accept-alias`, `--accept-stopword`, and `--accept-interest`. Accepted watch terms are written into `configs/term_watchlist.json`, and every applied action is appended into `configs/term_change_log.json`.