- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/) - README 新增 Agent Skill 章节说明供 Agent 使用的工作流
2.8 KiB
Content Extraction Pipeline
How the pipeline turns FreshRSS items into extractable article text.
Core Rule: FreshRSS items never re-fetch the original URL
FreshRSS is an RSS-only upstream. The pipeline never makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by RSS_ONLY_UPSTREAMS = {"freshrss"} in pipeline.py.
The only exception: non-FreshRSS upstreams (future sources that don't set upstream: freshrss) may trigger fetch_html() as a fallback.
Content source priority chain
content_loader.py tries sources in this order, using the first one with ≥500 readable characters:
| Priority | Source | Meaning |
|---|---|---|
| 1 | raw_html |
HTML pre-injected via ExtractionInput.raw_html. Rarely used in normal FreshRSS runs. |
| 2 | item.raw_content |
RSS <content:encoded> — the full article body. Some feeds provide this; many don't. |
| 3 | item.raw_summary |
RSS <description> — the summary/snippet field. This is the most common source in current runs. |
| 4 | rss_content |
RSS content from non-item sources. |
| — | none |
Nothing usable → raises RSS_CONTENT_MISSING for FreshRSS items (because fetch is skipped). |
How content_source maps to actual text quality
The content_source field in every item-XX.extracted.json tells you what the pipeline actually used:
item.raw_content→ Full article text from RSS<content:encoded>. Best quality, same as reading the original page.item.raw_summary→ RSS summary/description only. Not the full article. Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text.rss_content→ From standalone RSS content. Quality depends on the feed.fetched_html→ HTML fetched from the original URL (never happens for FreshRSS; only for non-FreshRSS upstreams).
What this means for digest quality
If you see content_source: item.raw_summary in the extracted files (current norm), the AI is summarizing from a feed summary/snippet, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions.
To improve quality: either find feeds that provide full <content:encoded>, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters).
Quick check
# Check content_source for latest run
grep -h "content_source" outputs/freshrss/rerun/*/extracted/item-*.json | sort | uniq -c
Relevant code paths
src/summary_mcp/core/pipeline.py—RSS_ONLY_UPSTREAMS,_should_skip_fetch(),extract_content()src/summary_mcp/core/content_loader.py—choose_inline_content()priority chain,fetch_html()(never called for FreshRSS)