# Content Extraction Pipeline How the pipeline turns FreshRSS items into extractable article text. ## Core Rule: FreshRSS items never re-fetch the original URL **FreshRSS is an RSS-only upstream.** The pipeline *never* makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by `RSS_ONLY_UPSTREAMS = {"freshrss"}` in `pipeline.py`. The only exception: non-FreshRSS upstreams (future sources that don't set `upstream: freshrss`) may trigger `fetch_html()` as a fallback. ## Content source priority chain `content_loader.py` tries sources in this order, using the **first one with ≥500 readable characters**: | Priority | Source | Meaning | |----------|--------|---------| | 1 | `raw_html` | HTML pre-injected via `ExtractionInput.raw_html`. Rarely used in normal FreshRSS runs. | | 2 | `item.raw_content` | RSS `` — the full article body. Some feeds provide this; many don't. | | 3 | `item.raw_summary` | RSS `` — the summary/snippet field. **This is the most common source in current runs.** | | 4 | `rss_content` | RSS content from non-item sources. | | — | `none` | Nothing usable → raises `RSS_CONTENT_MISSING` for FreshRSS items (because fetch is skipped). | ## How `content_source` maps to actual text quality The `content_source` field in every `item-XX.extracted.json` tells you what the pipeline actually used: - **`item.raw_content`** → Full article text from RSS ``. Best quality, same as reading the original page. - **`item.raw_summary`** → RSS summary/description only. **Not the full article.** Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text. - **`rss_content`** → From standalone RSS content. Quality depends on the feed. - **`fetched_html`** → HTML fetched from the original URL (**never happens for FreshRSS**; only for non-FreshRSS upstreams). ## What this means for digest quality If you see `content_source: item.raw_summary` in the extracted files (current norm), the AI is summarizing from a **feed summary/snippet**, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions. To improve quality: either find feeds that provide full ``, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters). ## Quick check ```bash # Check content_source for latest run grep -h "content_source" outputs/freshrss/rerun/*/extracted/item-*.json | sort | uniq -c ``` ## Relevant code paths - `src/summary_mcp/core/pipeline.py` — `RSS_ONLY_UPSTREAMS`, `_should_skip_fetch()`, `extract_content()` - `src/summary_mcp/core/content_loader.py` — `choose_inline_content()` priority chain, `fetch_html()` (never called for FreshRSS)