Extract three internal helpers to reduce the main function from ~330 lines to ~80 lines: - _process_item(): handles per-item extract/summarize/filter/build - _build_and_persist_delivery(): builds payload and persists keyword index - _build_run_report(): assembles status counts and run report dict No behavior changes. External signature and return shape unchanged.
Content Extract MCP
Python MCP scaffold for article content extraction, structured summary validation, deterministic filtering, and Markdown sink output.
Run
pip install -e .
summary-mcp
The server exposes four tools:
extract_url_contentextract_item_contentfilter_summary_resultrun_freshrss_openclaw_pipeline
Article-summary post-processing (separate LLM optional):
Daily knowledge-base defaults:
-
IMA_DAILY_KNOWLEDGE_BASE_ID— default IMA knowledge base ID for daily single-article summaries -
IMA_DAILY_KNOWLEDGE_BASE_NAME— default IMA knowledge base name (expected:daily) -
Runtime upload logic should verify the configured target before upload; if unavailable, try resolving by name and create
dailyif needed -
article-summaryMCP tool (operates on existing extracted payloads) -
scripts/run_article_summaries.pyCLI helper
Validate an LLM summary result:
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
Run the minimal extraction-to-summary loop:
python scripts/run_summary_loop.py ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--prompt outputs/prompts/llm-summary-prompt.txt ^
--output outputs/reference/summary/result.loop.json
Pull FreshRSS entries and map them into normalized item objects:
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=bot
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/pull_freshrss_items.py --limit 5 --mark-read
By default the script excludes entries already tagged as read. Add --include-read if you want the full reading list.
When --mark-read is enabled, fetched entries are marked as read after the script finishes successfully.
The script writes:
outputs/freshrss/raw/freshrss.raw.jsonoutputs/freshrss/items/freshrss.items.json
Pull FreshRSS entries and run content extraction for each mapped item:
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/run_freshrss_extract.py --limit 1 --mark-read
By default the script excludes entries already tagged as read. When --mark-read is enabled, only entries with successful extraction are marked as read.
The script writes:
outputs/freshrss/raw/freshrss.raw.jsonoutputs/freshrss/items/freshrss.items.jsonoutputs/freshrss/extracted/freshrss.extracted.json
Run the full FreshRSS pipeline and mark items as read only after the final OpenClaw delivery payload is written:
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
If you want to apply your personal engineering and AI-agent interest profile during filtering, pass a context file:
python scripts/run_freshrss_pipeline.py ^
--limit 5 ^
--context configs/filter_context.personal.json ^
--mark-read
This is the recommended production entrypoint. By default it writes only:
outputs/freshrss/rerun/<timestamp>/raw/freshrss.raw.jsonoutputs/freshrss/rerun/<timestamp>/candidates/openclaw-delivery-payload.jsonoutputs/freshrss/rerun/<timestamp>/run-report.jsonoutputs/freshrss/rerun/<timestamp>/extracted/item-XX.extracted.json(one per item)
It also updates the daily keyword index runtime data:
data/term_index/daily/YYYY-MM-DD.jsondata/term_index/term_stats.json
The main pipeline does not emit a batch-level freshrss.extracted.json file by default.
If you need additional per-item intermediates such as normalized items, summaries, filter decisions, candidate records, or candidate inputs, add --debug-artifacts.
When OpenClaw is connected to the MCP server, it should call run_freshrss_openclaw_pipeline for the same behavior directly through MCP. The tool also supports debug_artifacts=true when deeper inspection is needed.
Run deterministic filter rules against a structured summary result:
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--output outputs/reference/filter/filter-decision.json
You can optionally pass a context file to inject interest topics or source tags:
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--context outputs/reference/filter/filter-context.json ^
--output outputs/reference/filter/filter-decision.with-context.json
Rule engine details and rule authoring guidance live in:
docs/design/filter-rule-engine-design.mddocs/design/filter-rule-engine-usage.md
Write a filtered result into the Markdown sink:
python scripts/run_markdown_sink.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json
The script writes markdown notes under knowledge-base/.
Build an internal ArticleCandidateRecord and a slim OpenClawCandidateInput:
python scripts/run_article_candidate.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json ^
--section-hint tools_and_workflows
The script writes by default:
outputs/reference/candidates/article-candidate-record.jsonoutputs/reference/candidates/openclaw-candidate-input.json
Build a batch OpenClaw delivery payload:
python scripts/build_openclaw_delivery.py ^
--input-dir outputs/freshrss/candidates/batch ^
--sort-by-rank ^
--date 2026-03-25
The script writes by default:
outputs/reference/candidates/openclaw-delivery-payload.json
Output layout details live in outputs/README.md.
Keyword index defaults live in:
configs/term_aliases.jsonconfigs/term_stopwords.jsonconfigs/term_cleanup_policy.jsonconfigs/term_watchlist.jsonconfigs/term_change_log.json
You can also rebuild the keyword index from an existing delivery payload:
python scripts/build_keyword_index.py ^
--input outputs/reference/candidates/openclaw-delivery-payload.json
Runtime keyword data is stored under data/term_index/.
The keyword cleanup review skill lives in:
skills/keyword-cleanup-review/
To build a review bundle for the LLM skill:
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
--days 7 ^
--top 50 ^
--output outputs/term_index/review/keyword-cleanup-bundle.json
The review bundle now also carries cleanup governance context:
- cleanup thresholds from
configs/term_cleanup_policy.json - the current watch list from
configs/term_watchlist.json - recent applied changes from
configs/term_change_log.json
The skill only produces review inputs and suggestions. It does not modify term_aliases, term_stopwords, or filter_context.personal.json automatically.
To preview accepted suggestions before writing any config files:
python scripts/apply_term_suggestions.py ^
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
--accept-watch Cron Heartbeat Memory ^
--dry-run
Remove --dry-run to write the accepted changes. The script can also apply accepted alias, stopword, and interest keyword suggestions through --accept-alias, --accept-stopword, and --accept-interest. Accepted watch terms are written into configs/term_watchlist.json, and every applied action is appended into configs/term_change_log.json.
Article-summary LLM configuration
Set a dedicated model for post-processing summaries without affecting the main pipeline:
ARTICLE_SUMMARY_LLM_API_URLARTICLE_SUMMARY_LLM_MODELARTICLE_SUMMARY_LLM_API_KEY
If these are not set, the summarizer falls back to the main LLM_* / OPENAI_* settings used elsewhere.
Example (PowerShell style):
set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
Then run, for example, either via the CLI script:
python scripts/run_article_summaries.py ^
--extracted outputs/freshrss/extracted/freshrss.extracted.json ^
--ids 12345 67890 ^
--output-dir outputs/freshrss/single_summaries
…or through the MCP server tool generate_article_summaries exposed by summary_mcp.server:
extracted_path(string): path to the extracted JSON, for exampleoutputs/freshrss/extracted/freshrss.extracted.json.selected_ids(array of strings): one or moreitem_idvalues from the extracted payload to summarize.output_dir(optional string): directory to write Markdown summaries. If omitted, summaries are written undersingle_summaries/next to the extracted file.llm_api_key/llm_model/llm_api_url(optional strings): overrides for article-summary LLM settings. If omitted, the tool falls back toARTICLE_SUMMARY_*or mainLLM_*env vars as described above.
The tool returns a JSON array of file paths for the generated Markdown summaries.