163 lines
5.2 KiB
Markdown
163 lines
5.2 KiB
Markdown
# Content Extract MCP
|
|
|
|
Python MCP scaffold for article content extraction, structured summary validation, deterministic filtering, and Markdown sink output.
|
|
|
|
## Run
|
|
|
|
```bash
|
|
pip install -e .
|
|
summary-mcp
|
|
```
|
|
|
|
The server exposes four tools:
|
|
|
|
- `extract_url_content`
|
|
- `extract_item_content`
|
|
- `filter_summary_result`
|
|
- `run_freshrss_openclaw_pipeline`
|
|
|
|
Validate an LLM summary result:
|
|
|
|
```bash
|
|
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
|
|
```
|
|
|
|
Run the minimal extraction-to-summary loop:
|
|
|
|
```bash
|
|
python scripts/run_summary_loop.py ^
|
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
|
--prompt outputs/prompts/llm-summary-prompt.txt ^
|
|
--output outputs/reference/summary/result.loop.json
|
|
```
|
|
|
|
Pull FreshRSS entries and map them into normalized `item` objects:
|
|
|
|
```bash
|
|
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
|
set FRESHRSS_USERNAME=bot
|
|
set FRESHRSS_API_PASSWORD=your-api-password
|
|
python scripts/pull_freshrss_items.py --limit 5 --mark-read
|
|
```
|
|
|
|
By default the script excludes entries already tagged as `read`. Add `--include-read` if you want the full reading list.
|
|
When `--mark-read` is enabled, fetched entries are marked as read after the script finishes successfully.
|
|
|
|
The script writes:
|
|
|
|
- `outputs/freshrss/raw/freshrss.raw.json`
|
|
- `outputs/freshrss/items/freshrss.items.json`
|
|
|
|
Pull FreshRSS entries and run content extraction for each mapped item:
|
|
|
|
```bash
|
|
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
|
set FRESHRSS_USERNAME=osiman
|
|
set FRESHRSS_API_PASSWORD=your-api-password
|
|
python scripts/run_freshrss_extract.py --limit 1 --mark-read
|
|
```
|
|
|
|
By default the script excludes entries already tagged as `read`. When `--mark-read` is enabled, only entries with successful extraction are marked as read.
|
|
|
|
The script writes:
|
|
|
|
- `outputs/freshrss/raw/freshrss.raw.json`
|
|
- `outputs/freshrss/items/freshrss.items.json`
|
|
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
|
|
|
Run the full FreshRSS pipeline and mark items as read only after the final OpenClaw delivery payload is written:
|
|
|
|
```bash
|
|
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
|
set FRESHRSS_USERNAME=osiman
|
|
set FRESHRSS_API_PASSWORD=your-api-password
|
|
set LLM_API_URL=https://api.deepseek.com
|
|
set LLM_API_KEY=your-llm-api-key
|
|
set LLM_MODEL=deepseek-chat
|
|
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
|
|
```
|
|
|
|
If you want to apply your personal engineering and AI-agent interest profile during filtering, pass a context file:
|
|
|
|
```bash
|
|
python scripts/run_freshrss_pipeline.py ^
|
|
--limit 5 ^
|
|
--context configs/filter_context.personal.json ^
|
|
--mark-read
|
|
```
|
|
|
|
This is the recommended production entrypoint. By default it writes only:
|
|
|
|
- `outputs/freshrss/rerun/<timestamp>/raw/freshrss.raw.json`
|
|
- `outputs/freshrss/rerun/<timestamp>/candidates/openclaw-delivery-payload.json`
|
|
- `outputs/freshrss/rerun/<timestamp>/run-report.json`
|
|
|
|
If you need per-item intermediates, add `--debug-artifacts`.
|
|
|
|
When OpenClaw is connected to the MCP server, it should call `run_freshrss_openclaw_pipeline` for the same behavior directly through MCP. The tool also supports `debug_artifacts=true` when deeper inspection is needed.
|
|
|
|
Run deterministic filter rules against a structured summary result:
|
|
|
|
```bash
|
|
python scripts/run_filter_rules.py ^
|
|
--summary outputs/reference/summary/result.loop.json ^
|
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
|
--output outputs/reference/filter/filter-decision.json
|
|
```
|
|
|
|
You can optionally pass a context file to inject interest topics or source tags:
|
|
|
|
```bash
|
|
python scripts/run_filter_rules.py ^
|
|
--summary outputs/reference/summary/result.loop.json ^
|
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
|
--context outputs/reference/filter/filter-context.json ^
|
|
--output outputs/reference/filter/filter-decision.with-context.json
|
|
```
|
|
|
|
Rule engine details and rule authoring guidance live in:
|
|
|
|
- `docs/design/filter-rule-engine-design.md`
|
|
- `docs/design/filter-rule-engine-usage.md`
|
|
|
|
Write a filtered result into the Markdown sink:
|
|
|
|
```bash
|
|
python scripts/run_markdown_sink.py ^
|
|
--summary outputs/reference/summary/result.loop.json ^
|
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
|
--filter outputs/reference/filter/filter-decision.json
|
|
```
|
|
|
|
The script writes markdown notes under `knowledge-base/`.
|
|
|
|
Build an internal `ArticleCandidateRecord` and a slim `OpenClawCandidateInput`:
|
|
|
|
```bash
|
|
python scripts/run_article_candidate.py ^
|
|
--summary outputs/reference/summary/result.loop.json ^
|
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
|
--filter outputs/reference/filter/filter-decision.json ^
|
|
--section-hint tools_and_workflows
|
|
```
|
|
|
|
The script writes by default:
|
|
|
|
- `outputs/reference/candidates/article-candidate-record.json`
|
|
- `outputs/reference/candidates/openclaw-candidate-input.json`
|
|
|
|
Build a batch OpenClaw delivery payload:
|
|
|
|
```bash
|
|
python scripts/build_openclaw_delivery.py ^
|
|
--input-dir outputs/freshrss/candidates/batch ^
|
|
--sort-by-rank ^
|
|
--date 2026-03-25
|
|
```
|
|
|
|
The script writes by default:
|
|
|
|
- `outputs/reference/candidates/openclaw-delivery-payload.json`
|
|
|
|
Output layout details live in `outputs/README.md`.
|