# Content Extract MCP Python MCP scaffold for article content extraction, structured summary validation, deterministic filtering, and Markdown sink output. ## Run ```bash pip install -e . summary-mcp ``` The server exposes five tools: - `extract_url_content` - `extract_item_content` - `filter_summary_result` - `run_freshrss_openclaw_pipeline` - `generate_article_summaries` Article-summary post-processing (separate LLM optional): Daily knowledge-base defaults: - `IMA_DAILY_KNOWLEDGE_BASE_ID` — default IMA knowledge base ID for daily single-article summaries - `IMA_DAILY_KNOWLEDGE_BASE_NAME` — default IMA knowledge base name (expected: `daily`) - Runtime upload logic should verify the configured target before upload; if unavailable, try resolving by name and create `daily` if needed - `article-summary` MCP tool (operates on existing extracted payloads) - `scripts/run_article_summaries.py` CLI helper Validate an LLM summary result: ```bash validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json ``` Run the minimal extraction-to-summary loop: ```bash python scripts/run_summary_loop.py ^ --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^ --prompt outputs/prompts/llm-summary-prompt.txt ^ --output outputs/reference/summary/result.loop.json ``` Pull FreshRSS entries and map them into normalized `item` objects: ```bash set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php set FRESHRSS_USERNAME=bot set FRESHRSS_API_PASSWORD=your-api-password python scripts/pull_freshrss_items.py --limit 5 --mark-read ``` By default the script excludes entries already tagged as `read`. Add `--include-read` if you want the full reading list. When `--mark-read` is enabled, fetched entries are marked as read after the script finishes successfully. The script writes: - `outputs/freshrss/raw/freshrss.raw.json` - `outputs/freshrss/items/freshrss.items.json` Pull FreshRSS entries and run content extraction for each mapped item: ```bash set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php set FRESHRSS_USERNAME=osiman set FRESHRSS_API_PASSWORD=your-api-password python scripts/run_freshrss_extract.py --limit 1 --mark-read ``` By default the script excludes entries already tagged as `read`. When `--mark-read` is enabled, only entries with successful extraction are marked as read. The script writes: - `outputs/freshrss/raw/freshrss.raw.json` - `outputs/freshrss/items/freshrss.items.json` - `outputs/freshrss/extracted/freshrss.extracted.json` Run the full FreshRSS pipeline and mark items as read only after the final OpenClaw delivery payload is written: ```bash set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php set FRESHRSS_USERNAME=osiman set FRESHRSS_API_PASSWORD=your-api-password set LLM_API_URL=https://api.deepseek.com set LLM_API_KEY=your-llm-api-key set LLM_MODEL=deepseek-chat python scripts/run_freshrss_pipeline.py --limit 5 --mark-read ``` If you want to apply your personal engineering and AI-agent interest profile during filtering, pass a context file: ```bash python scripts/run_freshrss_pipeline.py ^ --limit 5 ^ --context configs/filter_context.personal.json ^ --mark-read ``` This is the recommended production entrypoint. By default it writes only: - `outputs/freshrss/rerun//raw/freshrss.raw.json` - `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` - `outputs/freshrss/rerun//run-report.json` - `outputs/freshrss/rerun//extracted/item-XX.extracted.json` (one per item) It also updates the daily keyword index runtime data: - `data/term_index/daily/YYYY-MM-DD.json` - `data/term_index/term_stats.json` The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default. If you need additional per-item intermediates such as normalized items, summaries, filter decisions, candidate records, or candidate inputs, add `--debug-artifacts`. When OpenClaw is connected to the MCP server, it should call `run_freshrss_openclaw_pipeline` for the same behavior directly through MCP. The tool also supports `debug_artifacts=true` when deeper inspection is needed. Run deterministic filter rules against a structured summary result: ```bash python scripts/run_filter_rules.py ^ --summary outputs/reference/summary/result.loop.json ^ --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^ --output outputs/reference/filter/filter-decision.json ``` You can optionally pass a context file to inject interest topics or source tags: ```bash python scripts/run_filter_rules.py ^ --summary outputs/reference/summary/result.loop.json ^ --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^ --context outputs/reference/filter/filter-context.json ^ --output outputs/reference/filter/filter-decision.with-context.json ``` Rule engine details and rule authoring guidance live in: - `docs/design/filter-rule-engine-design.md` - `docs/design/filter-rule-engine-usage.md` Write a filtered result into the Markdown sink: ```bash python scripts/run_markdown_sink.py ^ --summary outputs/reference/summary/result.loop.json ^ --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^ --filter outputs/reference/filter/filter-decision.json ``` The script writes markdown notes under `knowledge-base/`. Build an internal `ArticleCandidateRecord` and a slim `OpenClawCandidateInput`: ```bash python scripts/run_article_candidate.py ^ --summary outputs/reference/summary/result.loop.json ^ --extracted outputs/reference/extracted/read-flow-2026.extracted.json ^ --filter outputs/reference/filter/filter-decision.json ^ --section-hint tools_and_workflows ``` The script writes by default: - `outputs/reference/candidates/article-candidate-record.json` - `outputs/reference/candidates/openclaw-candidate-input.json` Build a batch OpenClaw delivery payload: ```bash python scripts/build_openclaw_delivery.py ^ --input-dir outputs/freshrss/candidates/batch ^ --sort-by-rank ^ --date 2026-03-25 ``` The script writes by default: - `outputs/reference/candidates/openclaw-delivery-payload.json` Output layout details live in `outputs/README.md`. Keyword index defaults live in: - `configs/term_aliases.json` - `configs/term_stopwords.json` - `configs/term_cleanup_policy.json` - `configs/term_watchlist.json` - `configs/term_change_log.json` You can also rebuild the keyword index from an existing delivery payload: ```bash python scripts/build_keyword_index.py ^ --input outputs/reference/candidates/openclaw-delivery-payload.json ``` Runtime keyword data is stored under `data/term_index/`. The keyword cleanup review skill lives in: - `skills/keyword-cleanup-review/` To build a review bundle for the LLM skill: ```bash python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^ --days 7 ^ --top 50 ^ --output outputs/term_index/review/keyword-cleanup-bundle.json ``` The review bundle now also carries cleanup governance context: - cleanup thresholds from `configs/term_cleanup_policy.json` - the current watch list from `configs/term_watchlist.json` - recent applied changes from `configs/term_change_log.json` The skill only produces review inputs and suggestions. It does not modify `term_aliases`, `term_stopwords`, or `filter_context.personal.json` automatically. To preview accepted suggestions before writing any config files: ```bash python scripts/apply_term_suggestions.py ^ --suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^ --accept-watch Cron Heartbeat Memory ^ --dry-run ``` Remove `--dry-run` to write the accepted changes. The script can also apply accepted `alias`, `stopword`, and `interest keyword` suggestions through `--accept-alias`, `--accept-stopword`, and `--accept-interest`. Accepted watch terms are written into `configs/term_watchlist.json`, and every applied action is appended into `configs/term_change_log.json`. ## Article-summary LLM configuration Set a dedicated model for post-processing summaries without affecting the main pipeline: - `ARTICLE_SUMMARY_LLM_API_URL` - `ARTICLE_SUMMARY_LLM_MODEL` - `ARTICLE_SUMMARY_LLM_API_KEY` If these are not set, the summarizer falls back to the main `LLM_*` / `OPENAI_*` settings used elsewhere. Example (PowerShell style): ```bash set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat ``` Then run, for example, either via the CLI script: ```bash python scripts/run_article_summaries.py ^ --extracted outputs/freshrss/extracted/freshrss.extracted.json ^ --ids 12345 67890 ^ --output-dir outputs/freshrss/single_summaries ``` …or through the MCP server tool `generate_article_summaries` exposed by `summary_mcp.server`: - `extracted_path` (string): path to a single-item extracted JSON (e.g. `outputs/freshrss/rerun//extracted/item-01.extracted.json`) or a batch extracted JSON containing a `results` array. - `selected_ids` (array of strings): one or more `item_id` values to summarize. Pass an empty array to summarize all items in the file. - `output_dir` (optional string): directory to write Markdown summaries. If omitted, summaries are written under `single_summaries/` next to the extracted file. - `llm_api_key` / `llm_model` / `llm_api_url` (optional strings): overrides for article-summary LLM settings. If omitted, the tool falls back to `ARTICLE_SUMMARY_*` or main `LLM_*` env vars as described above. The tool returns a JSON array of file paths for the generated Markdown summaries. The article summary uses a dedicated prompt (`outputs/prompts/article-summary-prompt.txt`) that is completely independent from the daily digest prompt. It outputs a structured knowledge note in Chinese with sections: 核心结论、主要论点、关键方法 / 机制、重要细节、可复用启发、关键词、主题.