2.9 KiB
Content Extraction Pipeline
How the pipeline turns FreshRSS items into extractable article text.
Core Rule: FreshRSS items never re-fetch the original URL
FreshRSS is an RSS-only upstream. The pipeline never makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by RSS_ONLY_UPSTREAMS = {"freshrss"} in pipeline.py.
The only exception: non-FreshRSS upstreams (future sources that don't set upstream: freshrss) may trigger fetch_html() as a fallback.
Content source priority chain
content_loader.py tries sources in this order, using the first one with ≥500 readable characters:
| Priority | Source | Meaning |
|---|---|---|
| 1 | raw_html |
HTML pre-injected via ExtractionInput.raw_html. Rarely used in normal FreshRSS runs. |
| 2 | item.raw_content |
RSS <content:encoded> — the full article body. Some feeds provide this; many don't. |
| 3 | item.raw_summary |
RSS <description> — the summary/snippet field. This is the most common source in current runs. |
| 4 | rss_content |
RSS content from non-item sources. |
| — | none |
Nothing usable → raises RSS_CONTENT_MISSING for FreshRSS items (because fetch is skipped). |
How content_source maps to actual text quality
The content_source field in every item-XX.extracted.json tells you what the pipeline actually used:
item.raw_content→ Full article text from RSS<content:encoded>. Best quality, same as reading the original page.item.raw_summary→ RSS summary/description only. Not the full article. Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text.rss_content→ From standalone RSS content. Quality depends on the feed.fetched_html→ HTML fetched from the original URL (never happens for FreshRSS; only for non-FreshRSS upstreams).
What this means for digest quality
If you see content_source: item.raw_summary in the extracted files (current norm), the AI is summarizing from a feed summary/snippet, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions.
To improve quality: either find feeds that provide full <content:encoded>, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters).
Quick check
先调用 list_run_artifacts(run_id),再读取返回的 extracted Artifact 路径,检查 content_source 和 article.plain_text。不要按 run_id 或“最新目录”手拼路径。
Relevant code paths
src/summary_mcp/core/pipeline.py—RSS_ONLY_UPSTREAMS,_should_skip_fetch(),extract_content()src/summary_mcp/core/content_loader.py—choose_inline_content()priority chain,fetch_html()(never called for FreshRSS)