Files
reader/README.md
T

120 lines
3.4 KiB
Markdown

# Content Extract MCP
Python MCP scaffold for article content extraction, structured summary validation, deterministic filtering, and Markdown sink output.
## Run
```bash
pip install -e .
summary-mcp
```
The server exposes three tools:
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
Validate an LLM summary result:
```bash
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
```
Run the minimal extraction-to-summary loop:
```bash
python scripts/run_summary_loop.py ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--prompt outputs/prompts/llm-summary-prompt.txt ^
--output outputs/reference/summary/result.loop.json
```
Pull FreshRSS entries and map them into normalized `item` objects:
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=bot
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/pull_freshrss_items.py --limit 5
```
The script writes:
- `outputs/freshrss/raw/freshrss.raw.json`
- `outputs/freshrss/items/freshrss.items.json`
Pull FreshRSS entries and run content extraction for each mapped item:
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/run_freshrss_extract.py --limit 1
```
The script writes:
- `outputs/freshrss/raw/freshrss.raw.json`
- `outputs/freshrss/items/freshrss.items.json`
- `outputs/freshrss/extracted/freshrss.extracted.json`
Run deterministic filter rules against a structured summary result:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--output outputs/reference/filter/filter-decision.json
```
You can optionally pass a context file to inject interest topics or source tags:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--context outputs/reference/filter/filter-context.json ^
--output outputs/reference/filter/filter-decision.with-context.json
```
Write a filtered result into the Markdown sink:
```bash
python scripts/run_markdown_sink.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json
```
The script writes markdown notes under `knowledge-base/`.
Build an internal `ArticleCandidateRecord` and a slim `OpenClawCandidateInput`:
```bash
python scripts/run_article_candidate.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json ^
--section-hint tools_and_workflows
```
The script writes by default:
- `outputs/reference/candidates/article-candidate-record.json`
- `outputs/reference/candidates/openclaw-candidate-input.json`
Build a batch OpenClaw delivery payload:
```bash
python scripts/build_openclaw_delivery.py ^
--input-dir outputs/freshrss/candidates/batch ^
--sort-by-rank ^
--date 2026-03-25
```
The script writes by default:
- `outputs/reference/candidates/openclaw-delivery-payload.json`
Output layout details live in `outputs/README.md`.