Files
reader/README.md
T

11 KiB

Content Extract MCP

Python MCP scaffold for article content extraction, structured summary validation, deterministic filtering, and Markdown sink output.

Run

pip install -e .
summary-mcp

The server exposes five tools:

  • extract_url_content
  • extract_item_content
  • filter_summary_result
  • run_freshrss_openclaw_pipeline
  • generate_article_summaries

Article-summary post-processing (separate LLM optional):

Recommended production input for selected-article summaries

  • Use the per-item extracted files written by the main FreshRSS pipeline:
    • outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json
  • Treat these per-item extracted files as the formal default artifacts for downstream selected-article summarization.
  • A batch extracted file such as outputs/freshrss/extracted/freshrss.extracted.json is only a compatible input shape for ad hoc or legacy workflows, not the preferred production default.

Daily knowledge-base defaults:

  • IMA_DAILY_KNOWLEDGE_BASE_ID — default IMA knowledge base ID for daily single-article summaries

  • IMA_DAILY_KNOWLEDGE_BASE_NAME — default IMA knowledge base name (expected: daily)

  • Runtime upload logic should verify the configured target before upload; if unavailable, try resolving by name and create daily if needed

  • article-summary MCP tool (operates on existing extracted payloads)

  • scripts/run_article_summaries.py CLI helper

Validate an LLM summary result:

validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json

Run the minimal extraction-to-summary loop:

python scripts/run_summary_loop.py ^
  --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
  --prompt outputs/prompts/llm-summary-prompt.txt ^
  --output outputs/reference/summary/result.loop.json

Pull FreshRSS entries and map them into normalized item objects:

set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=bot
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/pull_freshrss_items.py --limit 5 --mark-read

By default the script excludes entries already tagged as read. Add --include-read if you want the full reading list. When --mark-read is enabled, fetched entries are marked as read after the script finishes successfully.

The script writes:

  • outputs/freshrss/raw/freshrss.raw.json
  • outputs/freshrss/items/freshrss.items.json

Pull FreshRSS entries and run content extraction for each mapped item:

set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
python scripts/run_freshrss_extract.py --limit 1 --mark-read

By default the script excludes entries already tagged as read. When --mark-read is enabled, only entries with successful extraction are marked as read.

The script writes:

  • outputs/freshrss/raw/freshrss.raw.json
  • outputs/freshrss/items/freshrss.items.json
  • outputs/freshrss/extracted/freshrss.extracted.json

Note: this batch extracted file is mainly a compatible artifact for standalone extraction runs and older workflows. The formal production default for downstream selected-article summarization is still the per-item extracted output under outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json.

Run the full FreshRSS pipeline and mark items as read only after the final OpenClaw delivery payload is written:

set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read

If you want to apply your personal engineering and AI-agent interest profile during filtering, pass a context file:

python scripts/run_freshrss_pipeline.py ^
  --limit 5 ^
  --context configs/filter_context.personal.json ^
  --mark-read

This is the recommended production entrypoint. By default it writes only:

  • outputs/freshrss/rerun/<timestamp>/raw/freshrss.raw.json
  • outputs/freshrss/rerun/<timestamp>/candidates/openclaw-delivery-payload.json
  • outputs/freshrss/rerun/<timestamp>/run-report.json
  • outputs/freshrss/rerun/<timestamp>/extracted/item-XX.extracted.json (one per item)

It also updates the daily keyword index runtime data:

  • data/term_index/daily/YYYY-MM-DD.json
  • data/term_index/term_stats.json

The main pipeline does not emit a batch-level freshrss.extracted.json file by default. If you need additional per-item intermediates such as normalized items, summaries, filter decisions, candidate records, or candidate inputs, add --debug-artifacts.

When OpenClaw is connected to the MCP server, it should call run_freshrss_openclaw_pipeline for the same behavior directly through MCP. The tool also supports debug_artifacts=true when deeper inspection is needed.

Run deterministic filter rules against a structured summary result:

python scripts/run_filter_rules.py ^
  --summary outputs/reference/summary/result.loop.json ^
  --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
  --output outputs/reference/filter/filter-decision.json

You can optionally pass a context file to inject interest topics or source tags:

python scripts/run_filter_rules.py ^
  --summary outputs/reference/summary/result.loop.json ^
  --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
  --context outputs/reference/filter/filter-context.json ^
  --output outputs/reference/filter/filter-decision.with-context.json

Rule engine details and rule authoring guidance live in:

  • docs/design/filter-rule-engine-design.md
  • docs/design/filter-rule-engine-usage.md

Write a filtered result into the Markdown sink:

python scripts/run_markdown_sink.py ^
  --summary outputs/reference/summary/result.loop.json ^
  --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
  --filter outputs/reference/filter/filter-decision.json

The script writes markdown notes under knowledge-base/.

Build an internal ArticleCandidateRecord and a slim OpenClawCandidateInput:

python scripts/run_article_candidate.py ^
  --summary outputs/reference/summary/result.loop.json ^
  --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
  --filter outputs/reference/filter/filter-decision.json ^
  --section-hint tools_and_workflows

The script writes by default:

  • outputs/reference/candidates/article-candidate-record.json
  • outputs/reference/candidates/openclaw-candidate-input.json

Build a batch OpenClaw delivery payload:

python scripts/build_openclaw_delivery.py ^
  --input-dir outputs/freshrss/candidates/batch ^
  --sort-by-rank ^
  --date 2026-03-25

The script writes by default:

  • outputs/reference/candidates/openclaw-delivery-payload.json

Output layout details live in outputs/README.md.

Keyword index defaults live in:

  • configs/term_aliases.json
  • configs/term_stopwords.json
  • configs/term_cleanup_policy.json
  • configs/term_watchlist.json
  • configs/term_change_log.json

You can also rebuild the keyword index from an existing delivery payload:

python scripts/build_keyword_index.py ^
  --input outputs/reference/candidates/openclaw-delivery-payload.json

Runtime keyword data is stored under data/term_index/.

The keyword cleanup review skill lives in:

  • skills/keyword-cleanup-review/

To build a review bundle for the LLM skill:

python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
  --days 7 ^
  --top 50 ^
  --output outputs/term_index/review/keyword-cleanup-bundle.json

The review bundle now also carries cleanup governance context:

  • cleanup thresholds from configs/term_cleanup_policy.json
  • the current watch list from configs/term_watchlist.json
  • recent applied changes from configs/term_change_log.json

The skill only produces review inputs and suggestions. It does not modify term_aliases, term_stopwords, or filter_context.personal.json automatically.

To preview accepted suggestions before writing any config files:

python scripts/apply_term_suggestions.py ^
  --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
  --accept-watch Cron Heartbeat Memory ^
  --dry-run

Remove --dry-run to write the accepted changes. The script can also apply accepted alias, stopword, and interest keyword suggestions through --accept-alias, --accept-stopword, and --accept-interest. Accepted watch terms are written into configs/term_watchlist.json, and every applied action is appended into configs/term_change_log.json.

Article-summary LLM configuration

Set a dedicated model for post-processing summaries without affecting the main pipeline:

  • ARTICLE_SUMMARY_LLM_API_URL
  • ARTICLE_SUMMARY_LLM_MODEL
  • ARTICLE_SUMMARY_LLM_API_KEY

If these are not set, the summarizer falls back to the main LLM_* / OPENAI_* settings used elsewhere.

Example (PowerShell style):

set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat

Then run, for example, either via the CLI script. For normal production use, prefer a per-item extracted file from outputs/freshrss/rerun/<run_id>/extracted/. The batch extracted example below is kept only as a compatible legacy/ad hoc input shape:

python scripts/run_article_summaries.py ^
  --extracted outputs/freshrss/extracted/freshrss.extracted.json ^
  --ids 12345 67890 ^
  --output-dir outputs/freshrss/single_summaries

…or through the MCP server tool generate_article_summaries exposed by summary_mcp.server:

  • extracted_path (string): path to a single-item extracted JSON (e.g. outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json) or a batch extracted JSON containing a results array.
  • selected_ids (array of strings): one or more item_id values to summarize. Pass an empty array to summarize all items in the file.
  • output_dir (optional string): directory to write Markdown summaries. If omitted, summaries are written under single_summaries/ next to the extracted file.
  • llm_api_key / llm_model / llm_api_url (optional strings): overrides for article-summary LLM settings. If omitted, the tool falls back to ARTICLE_SUMMARY_* or main LLM_* env vars as described above.

The tool returns a JSON array of file paths for the generated Markdown summaries.

The article summary uses a dedicated prompt (outputs/prompts/article-summary-prompt.txt) that is completely independent from the daily digest prompt. It outputs a structured knowledge note in Chinese with sections: 核心结论、主要论点、关键方法 / 机制、重要细节、可复用启发、关键词、主题.