diff --git a/skills/reader-digest-flow/SKILL.md b/skills/reader-digest-flow/SKILL.md index edbd2dc..66342fb 100644 --- a/skills/reader-digest-flow/SKILL.md +++ b/skills/reader-digest-flow/SKILL.md @@ -99,7 +99,8 @@ Feishu 输出不要使用 Markdown 表格,见 `references/feishu-format-notes. - `今日概览` - `今日重点` - `趋势观察` -- `延伸阅读` + +每篇 `今日重点` 文章末尾必须添加 `来源:[来源名](原文 URL)`,来源链接跟随对应文章,不再生成独立的 `延伸阅读` 章节或重复链接。 仅发布用户在 Phase 3 选中的文章。公开页面不得出现 `keep/review/drop`、候选、待确认等内部状态。 diff --git a/skills/reader-digest-flow/references/content-extraction.md b/skills/reader-digest-flow/references/content-extraction.md index 96e2b2f..deca298 100644 --- a/skills/reader-digest-flow/references/content-extraction.md +++ b/skills/reader-digest-flow/references/content-extraction.md @@ -1,45 +1,45 @@ -# Content Extraction Pipeline +# 内容提取流程 -How the pipeline turns FreshRSS items into extractable article text. +本文说明处理流水线如何把 FreshRSS 条目转换为可供摘要使用的文章文本。 -## Core Rule: FreshRSS items never re-fetch the original URL +## 核心规则:FreshRSS 条目不重新抓取原文 URL -**FreshRSS is an RSS-only upstream.** The pipeline *never* makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by `RSS_ONLY_UPSTREAMS = {"freshrss"}` in `pipeline.py`. +**FreshRSS 是仅提供 RSS 内容的上游。** 对于 FreshRSS 条目,流水线不会向文章原始 URL 发起 HTTP 请求。该行为由 `pipeline.py` 中的 `RSS_ONLY_UPSTREAMS = {"freshrss"}` 强制保证。 -The only exception: non-FreshRSS upstreams (future sources that don't set `upstream: freshrss`) may trigger `fetch_html()` as a fallback. +唯一例外是非 FreshRSS 上游。未来未设置 `upstream: freshrss` 的其他来源,可以在必要时使用 `fetch_html()` 作为回退。 -## Content source priority chain +## 内容来源优先级 -`content_loader.py` tries sources in this order, using the **first one with ≥500 readable characters**: +`content_loader.py` 按以下顺序检查内容,并使用第一个包含 **至少 500 个可读字符** 的来源: -| Priority | Source | Meaning | -|----------|--------|---------| -| 1 | `raw_html` | HTML pre-injected via `ExtractionInput.raw_html`. Rarely used in normal FreshRSS runs. | -| 2 | `item.raw_content` | RSS `` — the full article body. Some feeds provide this; many don't. | -| 3 | `item.raw_summary` | RSS `` — the summary/snippet field. **This is the most common source in current runs.** | -| 4 | `rss_content` | RSS content from non-item sources. | -| — | `none` | Nothing usable → raises `RSS_CONTENT_MISSING` for FreshRSS items (because fetch is skipped). | +| 优先级 | 来源 | 含义 | +|--------|------|------| +| 1 | `raw_html` | 通过 `ExtractionInput.raw_html` 预先注入的 HTML;常规 FreshRSS 运行中很少使用。 | +| 2 | `item.raw_content` | RSS `` 中的文章正文;部分订阅源提供,部分不提供。 | +| 3 | `item.raw_summary` | RSS `` 中的摘要或片段;这是当前运行中最常见的来源。 | +| 4 | `rss_content` | 来自非条目字段的独立 RSS 内容。 | +| — | `none` | 没有可用内容;FreshRSS 不允许回源抓取,因此抛出 `RSS_CONTENT_MISSING`。 | -## How `content_source` maps to actual text quality +## `content_source` 与文本质量的关系 -The `content_source` field in every `item-XX.extracted.json` tells you what the pipeline actually used: +每个 `item-XX.extracted.json` 中的 `content_source` 字段表示流水线实际使用的内容来源: -- **`item.raw_content`** → Full article text from RSS ``. Best quality, same as reading the original page. -- **`item.raw_summary`** → RSS summary/description only. **Not the full article.** Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text. -- **`rss_content`** → From standalone RSS content. Quality depends on the feed. -- **`fetched_html`** → HTML fetched from the original URL (**never happens for FreshRSS**; only for non-FreshRSS upstreams). +- **`item.raw_content`**:RSS `` 提供的文章正文,通常质量最好,接近直接阅读原文。 +- **`item.raw_summary`**:只有 RSS 摘要或描述,并非完整正文。不同来源长度差异较大,通常为 300-2000 个字符;AI 摘要基于该片段,而不是完整文章。 +- **`rss_content`**:来自独立 RSS 内容,质量取决于订阅源。 +- **`fetched_html`**:从原始 URL 抓取的 HTML。FreshRSS 条目不会出现该来源,只适用于非 FreshRSS 上游。 -## What this means for digest quality +## 对日报质量的影响 -If you see `content_source: item.raw_summary` in the extracted files (current norm), the AI is summarizing from a **feed summary/snippet**, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions. +如果提取结果文件中出现 `content_source: item.raw_summary`,说明 AI 使用的是订阅源摘要或片段,而不是完整正文。日报内容显得较浅时,原因可能只是 RSS 描述过短。 -To improve quality: either find feeds that provide full ``, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters). +提高质量可以选择提供完整 `` 的订阅源,或者把内容来源切换到支持全文 RSS 的系统,例如具备全文提取能力的 RSS 代理或 FiveFilters 等服务。 -## Quick check +## 快速检查 -先调用 `list_run_artifacts(run_id)`,再读取返回的 extracted Artifact 路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。 +先调用 `list_run_artifacts(run_id)`,再读取返回的提取结果产物路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。 -## Relevant code paths +## 相关代码路径 -- `src/summary_mcp/core/pipeline.py` — `RSS_ONLY_UPSTREAMS`, `_should_skip_fetch()`, `extract_content()` -- `src/summary_mcp/core/content_loader.py` — `choose_inline_content()` priority chain, `fetch_html()` (never called for FreshRSS) +- `src/summary_mcp/core/pipeline.py`:`RSS_ONLY_UPSTREAMS`、`_should_skip_fetch()`、`extract_content()`。 +- `src/summary_mcp/core/content_loader.py`:`choose_inline_content()` 的优先级链和 `fetch_html()`;FreshRSS 条目不会调用后者。 diff --git a/skills/reader-digest-flow/references/feishu-format-notes.md b/skills/reader-digest-flow/references/feishu-format-notes.md index 7ec8795..25ab37b 100644 --- a/skills/reader-digest-flow/references/feishu-format-notes.md +++ b/skills/reader-digest-flow/references/feishu-format-notes.md @@ -1,26 +1,17 @@ -# Feishu Markdown Format Notes +# 飞书 Markdown 格式说明 -## Background +## 背景 -Hermes' Feishu gateway (`gateway/platforms/feishu.py`) sends outbound messages -through `_build_outbound_payload`, which checks content for markdown patterns -to decide how to send: +Hermes 的飞书网关(`gateway/platforms/feishu.py`)通过 `_build_outbound_payload` 发送消息。该方法会检查内容中的 Markdown 特征,并据此决定消息类型: -- Content matches `_MARKDOWN_HINT_RE` (bold, lists, code, links, etc.) - → sent as Feishu `post` type using `md` elements → renders correctly. +- 内容匹配 `_MARKDOWN_HINT_RE`(加粗、列表、代码、链接等)时,使用包含 `md` 元素的飞书 `post` 类型发送,可以正常渲染。 +- 内容匹配 `_MARKDOWN_TABLE_RE`(Markdown 表头和分隔行)时,整条消息会被强制转换为 `text` 类型,即纯文本,不再渲染 Markdown。 -- Content matches `_MARKDOWN_TABLE_RE` (a markdown table header + separator) - → **entire message** forced to `text` type (plain text) → no rendering. +原因是 `_build_markdown_post_payload` 会把内容包装为 `{"tag": "md", "text": "..."}` 元素,而飞书的 `md` 元素不支持表格,也没有把 Markdown 表格转换为飞书原生表格的逻辑。 -The root cause is that the `_build_markdown_post_payload` helper wraps content -in `{"tag": "md", "text": "..."}` elements, and Feishu's `md` element does not -support table rendering. There is no table-to-native-Feishu-table conversion. +## 飞书输出规则 -## Rules for Feishu output - -- **Never use markdown tables** in any message delivered via Feishu. - A single table anywhere in the message forces the whole message to plain text. -- Prefer bullet lists, sections with headings, or inline formatting instead. -- Bold (`**bold**`), inline code (`` `code` ``), unordered lists (`- item`), - ordered lists (`1. item`), and links all work correctly. -- Code fences (``` ``` ```) work but may have edge cases with trailing content. +- 通过飞书发送的消息不得使用 Markdown 表格;消息中只要出现一个表格,整条消息就会退化为纯文本。 +- 需要表达结构化信息时,优先使用分点列表、带标题的分节或行内格式。 +- 加粗(`**加粗**`)、行内代码(`` `代码` ``)、无序列表(`- 项目`)、有序列表(`1. 项目`)和链接均可正常使用。 +- 围栏式代码块可以使用,但代码块后的尾随内容可能存在渲染边界问题。 diff --git a/skills/reader-digest-flow/references/public-digest-example.md b/skills/reader-digest-flow/references/public-digest-example.md index 22c0c58..088a951 100644 --- a/skills/reader-digest-flow/references/public-digest-example.md +++ b/skills/reader-digest-flow/references/public-digest-example.md @@ -22,12 +22,10 @@ summary = "围绕 Agent 架构分层、Skills 标准化与工程实践的当日 这篇内容值得关注的原因在于,它把开放协议、分层架构和真实落地案例连接成了完整论证链。 +来源:[示例来源](https://example.com/a) + ## 趋势观察 1. Agent 正在从单体能力转向可组合的模块化体系。 2. 工具契约、状态管理和验证机制正在成为 AI 应用的核心工程能力。 3. Human-in-the-loop 仍是控制高风险副作用的重要边界。 - -## 延伸阅读 - -- [从 Agent 到 Skills:AI 智能体架构的范式转变](https://example.com/a)|示例来源