ff3d15b8c4b71a59453bac8d00e37b94b826925b
Content Extract MCP
Python MCP scaffold for article content extraction, structured summary validation, and deterministic filtering.
Run
pip install -e .
summary-mcp
The server exposes three tools:
extract_url_contentextract_item_contentfilter_summary_result
Validate an LLM summary result:
validate-llm-result outputs/result.json --extracted outputs/read-flow-2026.extracted.json
Run the minimal extraction-to-summary loop:
python scripts/run_summary_loop.py ^
--extracted outputs/read-flow-2026.extracted.json ^
--prompt outputs/llm-summary-prompt.txt ^
--output outputs/result.json
Pull FreshRSS entries and map them into normalized item objects:
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=bot
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/pull_freshrss_items.py --limit 5
The script writes:
outputs/freshrss.raw.jsonoutputs/freshrss.items.json
Pull FreshRSS entries and run content extraction for each mapped item:
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/run_freshrss_extract.py --limit 1
The script writes:
outputs/freshrss.raw.jsonoutputs/freshrss.items.jsonoutputs/freshrss.extracted.json
Run deterministic filter rules against a structured summary result:
python scripts/run_filter_rules.py ^
--summary outputs/result.json ^
--extracted outputs/read-flow-2026.extracted.json ^
--output outputs/filter-decision.json
You can optionally pass a context file to inject interest topics or source tags:
python scripts/run_filter_rules.py ^
--summary outputs/result.json ^
--extracted outputs/read-flow-2026.extracted.json ^
--context outputs/filter-context.json ^
--output outputs/filter-decision.with-context.json
Languages
Python
100%