Files
reader/skills/reader-digest-flow/references/content-extraction.md
T

2.9 KiB

Content Extraction Pipeline

How the pipeline turns FreshRSS items into extractable article text.

Core Rule: FreshRSS items never re-fetch the original URL

FreshRSS is an RSS-only upstream. The pipeline never makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by RSS_ONLY_UPSTREAMS = {"freshrss"} in pipeline.py.

The only exception: non-FreshRSS upstreams (future sources that don't set upstream: freshrss) may trigger fetch_html() as a fallback.

Content source priority chain

content_loader.py tries sources in this order, using the first one with ≥500 readable characters:

Priority Source Meaning
1 raw_html HTML pre-injected via ExtractionInput.raw_html. Rarely used in normal FreshRSS runs.
2 item.raw_content RSS <content:encoded> — the full article body. Some feeds provide this; many don't.
3 item.raw_summary RSS <description> — the summary/snippet field. This is the most common source in current runs.
4 rss_content RSS content from non-item sources.
— none Nothing usable → raises RSS_CONTENT_MISSING for FreshRSS items (because fetch is skipped).

How content_source maps to actual text quality

The content_source field in every item-XX.extracted.json tells you what the pipeline actually used:

  • item.raw_content → Full article text from RSS <content:encoded>. Best quality, same as reading the original page.
  • item.raw_summary → RSS summary/description only. Not the full article. Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text.
  • rss_content → From standalone RSS content. Quality depends on the feed.
  • fetched_html → HTML fetched from the original URL (never happens for FreshRSS; only for non-FreshRSS upstreams).

What this means for digest quality

If you see content_source: item.raw_summary in the extracted files (current norm), the AI is summarizing from a feed summary/snippet, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions.

To improve quality: either find feeds that provide full <content:encoded>, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters).

Quick check

先调用 list_run_artifacts(run_id),再读取返回的 extracted Artifact 路径,检查 content_source 和 article.plain_text。不要按 run_id 或“最新目录”手拼路径。

Relevant code paths

  • src/summary_mcp/core/pipeline.py — RSS_ONLY_UPSTREAMS, _should_skip_fetch(), extract_content()
  • src/summary_mcp/core/content_loader.py — choose_inline_content() priority chain, fetch_html() (never called for FreshRSS)