Files
reader/README.md
T

259 lines
8.7 KiB
Markdown

# Content Extract MCP
Python MCP scaffold for article content extraction, structured summary validation, deterministic filtering, and Markdown sink output.
## Run
```bash
pip install -e .
summary-mcp
```
The server exposes four tools:
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `run_freshrss_openclaw_pipeline`
Article-summary post-processing (separate LLM optional):
- `article-summary` MCP tool (operates on existing extracted payloads)
- `scripts/run_article_summaries.py` CLI helper
Validate an LLM summary result:
```bash
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
```
Run the minimal extraction-to-summary loop:
```bash
python scripts/run_summary_loop.py ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--prompt outputs/prompts/llm-summary-prompt.txt ^
--output outputs/reference/summary/result.loop.json
```
Pull FreshRSS entries and map them into normalized `item` objects:
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=bot
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/pull_freshrss_items.py --limit 5 --mark-read
```
By default the script excludes entries already tagged as `read`. Add `--include-read` if you want the full reading list.
When `--mark-read` is enabled, fetched entries are marked as read after the script finishes successfully.
The script writes:
- `outputs/freshrss/raw/freshrss.raw.json`
- `outputs/freshrss/items/freshrss.items.json`
Pull FreshRSS entries and run content extraction for each mapped item:
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/run_freshrss_extract.py --limit 1 --mark-read
```
By default the script excludes entries already tagged as `read`. When `--mark-read` is enabled, only entries with successful extraction are marked as read.
The script writes:
- `outputs/freshrss/raw/freshrss.raw.json`
- `outputs/freshrss/items/freshrss.items.json`
- `outputs/freshrss/extracted/freshrss.extracted.json`
Run the full FreshRSS pipeline and mark items as read only after the final OpenClaw delivery payload is written:
```bash
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
```
If you want to apply your personal engineering and AI-agent interest profile during filtering, pass a context file:
```bash
python scripts/run_freshrss_pipeline.py ^
--limit 5 ^
--context configs/filter_context.personal.json ^
--mark-read
```
This is the recommended production entrypoint. By default it writes only:
- `outputs/freshrss/rerun/<timestamp>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<timestamp>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<timestamp>/run-report.json`
It also updates the daily keyword index runtime data:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
If you need per-item intermediates, add `--debug-artifacts`.
When OpenClaw is connected to the MCP server, it should call `run_freshrss_openclaw_pipeline` for the same behavior directly through MCP. The tool also supports `debug_artifacts=true` when deeper inspection is needed.
Run deterministic filter rules against a structured summary result:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--output outputs/reference/filter/filter-decision.json
```
You can optionally pass a context file to inject interest topics or source tags:
```bash
python scripts/run_filter_rules.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--context outputs/reference/filter/filter-context.json ^
--output outputs/reference/filter/filter-decision.with-context.json
```
Rule engine details and rule authoring guidance live in:
- `docs/design/filter-rule-engine-design.md`
- `docs/design/filter-rule-engine-usage.md`
Write a filtered result into the Markdown sink:
```bash
python scripts/run_markdown_sink.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json
```
The script writes markdown notes under `knowledge-base/`.
Build an internal `ArticleCandidateRecord` and a slim `OpenClawCandidateInput`:
```bash
python scripts/run_article_candidate.py ^
--summary outputs/reference/summary/result.loop.json ^
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
--filter outputs/reference/filter/filter-decision.json ^
--section-hint tools_and_workflows
```
The script writes by default:
- `outputs/reference/candidates/article-candidate-record.json`
- `outputs/reference/candidates/openclaw-candidate-input.json`
Build a batch OpenClaw delivery payload:
```bash
python scripts/build_openclaw_delivery.py ^
--input-dir outputs/freshrss/candidates/batch ^
--sort-by-rank ^
--date 2026-03-25
```
The script writes by default:
- `outputs/reference/candidates/openclaw-delivery-payload.json`
Output layout details live in `outputs/README.md`.
Keyword index defaults live in:
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
You can also rebuild the keyword index from an existing delivery payload:
```bash
python scripts/build_keyword_index.py ^
--input outputs/reference/candidates/openclaw-delivery-payload.json
```
Runtime keyword data is stored under `data/term_index/`.
The keyword cleanup review skill lives in:
- `skills/keyword-cleanup-review/`
To build a review bundle for the LLM skill:
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
--days 7 ^
--top 50 ^
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
The review bundle now also carries cleanup governance context:
- cleanup thresholds from `configs/term_cleanup_policy.json`
- the current watch list from `configs/term_watchlist.json`
- recent applied changes from `configs/term_change_log.json`
The skill only produces review inputs and suggestions. It does not modify `term_aliases`, `term_stopwords`, or `filter_context.personal.json` automatically.
To preview accepted suggestions before writing any config files:
```bash
python scripts/apply_term_suggestions.py ^
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
--accept-watch Cron Heartbeat Memory ^
--dry-run
```
Remove `--dry-run` to write the accepted changes. The script can also apply accepted `alias`, `stopword`, and `interest keyword` suggestions through `--accept-alias`, `--accept-stopword`, and `--accept-interest`. Accepted watch terms are written into `configs/term_watchlist.json`, and every applied action is appended into `configs/term_change_log.json`.
## Article-summary LLM configuration
Set a dedicated model for post-processing summaries without affecting the main pipeline:
- `ARTICLE_SUMMARY_LLM_API_URL`
- `ARTICLE_SUMMARY_LLM_MODEL`
- `ARTICLE_SUMMARY_LLM_API_KEY`
If these are not set, the summarizer falls back to the main `LLM_*` / `OPENAI_*` settings used elsewhere.
Example (PowerShell style):
```bash
set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
```
Then run, for example, either via the CLI script:
```bash
python scripts/run_article_summaries.py ^
--extracted outputs/freshrss/extracted/freshrss.extracted.json ^
--ids 12345 67890 ^
--output-dir outputs/freshrss/single_summaries
```
…or through the MCP server tool `generate_article_summaries` exposed by `summary_mcp.server`:
- `extracted_path` (string): path to the extracted JSON, for example `outputs/freshrss/extracted/freshrss.extracted.json`.
- `selected_ids` (array of strings): one or more `item_id` values from the extracted payload to summarize.
- `output_dir` (optional string): directory to write Markdown summaries. If omitted, summaries are written under `single_summaries/` next to the extracted file.
- `llm_api_key` / `llm_model` / `llm_api_url` (optional strings): overrides for article-summary LLM settings. If omitted, the tool falls back to `ARTICLE_SUMMARY_*` or main `LLM_*` env vars as described above.
The tool returns a JSON array of file paths for the generated Markdown summaries.