refactor: decouple Hugo publication from pipeline - report first, user selects, then publish
- New flow: pipeline → report candidates → user selects Hugo articles → generate public digest → publish Hugo → user selects IMA articles → IMA deposition - digest-brief.json is now the sole data source (concise, less token consumption) - Removed 来源 line from public-digest-example.md to match SKILL hard rule - Public digest no longer auto-generated with all keep items
This commit is contained in:
+104
-105
@@ -1,6 +1,6 @@
|
||||
---
|
||||
name: reader-digest-flow
|
||||
description: Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, generate a daily report from reader payloads, publish the digest to Hugo, report the digest back in chat, select valuable articles, and then summarize only the selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → digest → selection → knowledge-base flow.
|
||||
description: Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, report candidates for review, let the user decide which articles go into the Hugo daily digest, then generate and publish the public digest, and optionally summarize selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → report → user-selection → Hugo → knowledge-base flow.
|
||||
---
|
||||
|
||||
# Reader Digest Flow
|
||||
@@ -23,8 +23,9 @@ Run the reader-based daily digest as a fixed SOP. Treat this skill as the orches
|
||||
- Generate the daily digest markdown from the payload in OpenClaw.
|
||||
- Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making.
|
||||
- **Public digest and internal review digest are orchestration-layer outputs, and by default are generated by the current OpenClaw session model rather than inheriting reader `LLM_*` settings.**
|
||||
- Publish only the public digest to Hugo.
|
||||
- Once a normal daily digest run succeeds and a public digest is generated, proceed directly with Hugo publication as the default action; do not ask the user for a separate confirmation about generating or publishing the Hugo page unless the user explicitly says to skip Hugo.
|
||||
- **Publish only the public digest to Hugo, and only after the user has explicitly selected which articles go into the daily digest.**
|
||||
- The default flow is: pipeline → report candidates → **user selects Hugo articles** → generate public digest → publish Hugo → (optional) user selects IMA articles → IMA deposition.
|
||||
- Do NOT generate the public digest or publish Hugo before the user has confirmed the Hugo article selection.
|
||||
- Report the internal review digest back to the user in chat.
|
||||
- **Do not upload the full daily digest to IMA.**
|
||||
- **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection. Once the user has confirmed which articles to keep, proceed directly with IMA deposition and do not ask for a second confirmation about uploading to the knowledge base.**
|
||||
@@ -57,10 +58,10 @@ After the async job succeeds, treat the returned `run_id` as the stable handle f
|
||||
|
||||
Minimum expected artifacts:
|
||||
|
||||
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`(public digest 优先读取的轻量输入,仅包含 `keep` 候选)
|
||||
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json` — **精简候选清单**(唯一要读取的文件,含 keep + review)
|
||||
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json` — 完整 payload(pipeline 内部产物,digest-brief 由此推导,流程中无需读取)
|
||||
- `outputs/freshrss/rerun/<run-id>/run-report.json`
|
||||
- extracted article data: `outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json` 等(每个候选一篇,格式为 `{"success": true, "article": {...}}`)
|
||||
- extracted article data: `outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json` 等(每个候选一篇,格式为 `{"success": true, "article": {...}}`,Phase 6 使用)
|
||||
|
||||
If the main pipeline job fails:
|
||||
|
||||
@@ -73,10 +74,65 @@ If the main pipeline job fails:
|
||||
- if `recommended_action=read_terminal_result`, continue from the terminal run result instead of retrying
|
||||
- if `recommended_action=start_new_run`, stop and report the exact failure point
|
||||
|
||||
### Phase 2: Generate daily digest markdown
|
||||
### Phase 2: Report candidates to the user (internal review digest only)
|
||||
|
||||
Read the real run outputs and generate digest markdown for the day.
|
||||
If there is no real payload, stop instead of writing a fake or example digest.
|
||||
After the pipeline run succeeds, do NOT generate the public digest or publish Hugo yet.
|
||||
Read candidates and produce **only the internal review digest** in a concise format (not the full payload dump).
|
||||
|
||||
Recommended **internal review digest** structure:
|
||||
|
||||
- `今日候选概况`
|
||||
- `已入选重点`
|
||||
- `待你确认`
|
||||
- `原始候选清单`
|
||||
|
||||
Data source: use `candidates/digest-brief.json` (`top_candidates` array) — it now includes both keep and review articles in a concise, simplified format (title, source, summary, highlights, category). No need to read the full payload.
|
||||
|
||||
Internal review digest writing rules:
|
||||
|
||||
- Keep this as a human-readable concise review draft, not a raw machine dump.
|
||||
- Do **not** display `rank` or `digest_rank` values.
|
||||
- Convert machine states to Chinese operator-facing labels:
|
||||
- `keep` → `已入选`
|
||||
- `review` → `待确认`
|
||||
- `drop` → `暂不纳入`
|
||||
- For every item under `已入选重点`, include:
|
||||
- title + source
|
||||
- status (`已入选`)
|
||||
- summary paragraph (use the concise version from digest-brief)
|
||||
- highlights (from digest-brief)
|
||||
- a short judgment paragraph explaining why it matters
|
||||
- For every item under `待你确认`, include:
|
||||
- title + source
|
||||
- status (`待确认`)
|
||||
- summary paragraph (concise)
|
||||
- highlights (if available)
|
||||
- reason(为什么要你确认)
|
||||
- recommendation(建议采纳/不采纳)
|
||||
- In `原始候选清单`, present all articles as a brief list (title + status + one-liner), using Chinese status labels.
|
||||
|
||||
**Hard reporting requirement:** the chat report must not be only a title list. For every reported article, include at least:
|
||||
- title
|
||||
- one-sentence summary
|
||||
- a short reason explaining why it matters / why it is recommended or pending confirmation
|
||||
|
||||
At this step:
|
||||
|
||||
- the public digest is **not** generated yet
|
||||
- Hugo is **not** published yet
|
||||
- the internal review digest is sent to the user in chat
|
||||
- the user decides: **which articles go into the Hugo daily digest** AND separately which articles go into IMA deposition
|
||||
|
||||
### Phase 3: User selects Hugo articles
|
||||
|
||||
The user reviews the internal digest and specifies which articles should appear in the Hugo daily digest.
|
||||
- Confirm the selection explicitly before proceeding.
|
||||
- If the user wants to include some `review` candidate articles, respect that choice.
|
||||
- Only the user-selected articles will appear in the public digest.
|
||||
|
||||
### Phase 4: Generate and publish Hugo public digest
|
||||
|
||||
Generate the public digest markdown **only for the user-selected articles**, then publish to Hugo.
|
||||
|
||||
**⚠️ 格式基准 — 开始生成任何 digest 内容之前,必须先完整阅读 `references/public-digest-example.md` 并以此为格式基准,不得凭记忆或直觉写作。**
|
||||
|
||||
@@ -92,87 +148,13 @@ If there is no real payload, stop instead of writing a fake or example digest.
|
||||
- [ ] "这篇更值得关注的原因在于:"段落
|
||||
5. [ ] `延伸阅读` 每条含来源标注:`- [标题](url)|来源`
|
||||
|
||||
For **public digest**, prefer `candidates/digest-brief.json` when present; fall back to `candidates/openclaw-delivery-payload.json` only if the brief file is missing.
|
||||
For **internal review digest**, continue reading the full `candidates/openclaw-delivery-payload.json`.
|
||||
Public digest source material: use `candidates/digest-brief.json` — it has all the info needed (title, summary, highlights, source, url). No need to read the full payload.
|
||||
|
||||
Produce two output views from the same run, preferably in **one generation step**:
|
||||
|
||||
1. **Public digest** — for Hugo / public browsing
|
||||
2. **Internal review digest** — for chat reporting and operator decisions
|
||||
|
||||
Before publishing, persist digest outputs back into the same reader run directory under:
|
||||
Persist the generated digest under:
|
||||
|
||||
- `outputs/freshrss/rerun/<run-id>/digest/public_digest.md`
|
||||
- `outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md`
|
||||
- `outputs/freshrss/rerun/<run-id>/digest/combined.json`
|
||||
|
||||
Then write only the public digest into Hugo using this structure:
|
||||
|
||||
- `content/daily/YYYY-MM-DD/index.md`
|
||||
|
||||
The public digest is the browsing layer, not the long-term knowledge layer.
|
||||
It must not expose internal workflow states or operator-facing review labels.
|
||||
|
||||
Recommended **public digest** structure:
|
||||
|
||||
- `今日概览`
|
||||
- `今日重点`
|
||||
- `趋势观察`
|
||||
- `延伸阅读`
|
||||
|
||||
Public digest writing rules:
|
||||
|
||||
- Use `digest-brief.json` as the default source when available.
|
||||
- Treat it as a public-only view that already excludes non-keep items.
|
||||
- **Hugo public digest must include ALL keep articles from `digest-brief.json`. Do NOT apply an additional manual filter or subset selection. The IMA deposition step (Phase 5–6) is separate and operates on a user-selected subset; Hugo always shows the full keep set.**
|
||||
- Keep the tone suitable for public browsing and Hugo publishing.
|
||||
- Style should follow `references/public-digest-example.md` as the default public-writing example.
|
||||
- Public digest is a **public reading draft / editor-style public note**, not a workflow report.
|
||||
- In `今日概览`, focus on the day’s topic lines, shared signals, and broader industry movement; do **not** describe filtering mechanics or internal selection process.
|
||||
- Do **not** expose internal workflow labels or operator language such as `待确认`, `建议沉淀到 IMA`, `keep/review/drop`, or `selection_decision`.
|
||||
- Explicitly avoid wording such as `共筛出`, `候选`, `保留`, `入选`, `待确认`, `建议沉淀` in public digest.
|
||||
- Prefer concise but information-dense writing.
|
||||
- For each item under `今日重点`, include not only summary and highlights, but also one short editor-style value sentence, for example: `这篇内容更值得关注的原因在于……`.
|
||||
- When rendering highlights in public digest, prefer a short label such as `值得关注:` followed by one item per line, instead of packing multiple points into a single long sentence.
|
||||
- In `延伸阅读`, every item must include source attribution in the form: `- [标题](url)|来源`.
|
||||
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other numbering variant.**
|
||||
- **All four sections are required: `今日概览`, `今日重点`, `趋势观察`, `延伸阅读`. Missing any section is a format violation.**
|
||||
|
||||
Recommended **internal review digest** structure:
|
||||
|
||||
- `今日候选概况`
|
||||
- `已入选重点`
|
||||
- `待你确认`
|
||||
- `建议沉淀到 IMA`
|
||||
- `原始候选清单`
|
||||
|
||||
Internal review digest writing rules:
|
||||
|
||||
- Use the full `openclaw-delivery-payload.json`.
|
||||
- Keep this as a human-readable review draft rather than a raw machine dump.
|
||||
- Do **not** display `rank` values.
|
||||
- Convert machine states to Chinese operator-facing labels:
|
||||
- `keep` → `已入选`
|
||||
- `review` → `待确认`
|
||||
- `drop` → `暂不纳入`
|
||||
- For every item under `已入选重点`, include:
|
||||
- title + source
|
||||
- status
|
||||
- a fuller summary paragraph
|
||||
- a short judgment paragraph explaining why it matters in today's digest
|
||||
- For every item under `待你确认`, include:
|
||||
- title + source
|
||||
- status
|
||||
- a fuller summary paragraph
|
||||
- reason
|
||||
- recommendation
|
||||
- In `原始候选清单`, also use Chinese status labels instead of raw machine values.
|
||||
|
||||
### Phase 3: Publish to Hugo
|
||||
|
||||
Publish only the public digest to Hugo.
|
||||
|
||||
Required execution steps:
|
||||
Then publish to Hugo:
|
||||
|
||||
1. write the public digest markdown to:
|
||||
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
|
||||
@@ -184,49 +166,66 @@ Required execution steps:
|
||||
- detail page works
|
||||
- latest digest is visible
|
||||
|
||||
Do not treat “markdown file written” as equivalent to publish success. Hugo publish for this SOP is complete only after redeploy and page verification pass.
|
||||
Do not treat "markdown file written" as equivalent to publish success. Hugo publish for this SOP is complete only after redeploy and page verification pass.
|
||||
|
||||
Do not block on style polish unless the user explicitly asks.
|
||||
|
||||
### Phase 4: Report digest back to the user
|
||||
The public digest is the browsing layer, not the long-term knowledge layer.
|
||||
It must not expose internal workflow states or operator-facing review labels.
|
||||
|
||||
Send the internal review digest in chat and ask the user which articles should be retained for long-term knowledge.
|
||||
Public digest writing rules:
|
||||
|
||||
**Hard reporting requirement:** the chat report must not be only a title list. For every reported article, include at least:
|
||||
- title
|
||||
- one-sentence summary
|
||||
- a short reason explaining why it matters / why it is recommended or pending confirmation
|
||||
- Use article content from `candidates/digest-brief.json` for the user-selected subset.
|
||||
- **Hugo public digest includes ONLY the articles the user selected in Phase 3.** Not all `keep` items — only the user's explicit selection.
|
||||
- Keep the tone suitable for public browsing and Hugo publishing.
|
||||
- Style should follow `references/public-digest-example.md` as the default public-writing example.
|
||||
- Public digest is a **public reading draft / editor-style public note**, not a workflow report.
|
||||
- In `今日概览`, focus on the day's topic lines, shared signals, and broader industry movement; do **not** describe filtering mechanics or internal selection process.
|
||||
- Do **not** expose internal workflow labels or operator language such as `待确认`, `建议沉淀到 IMA`, `keep/review/drop`, or `selection_decision`.
|
||||
- Explicitly avoid wording such as `共筛出`, `候选`, `保留`, `入选`, `待确认`, `建议沉淀` in public digest.
|
||||
- Prefer concise but information-dense writing.
|
||||
- For each item under `今日重点`, include not only summary and highlights, but also one short editor-style value sentence, for example: `这篇内容更值得关注的原因在于……`.
|
||||
- When rendering highlights in public digest, prefer a short label such as `值得关注:` followed by one item per line, instead of packing multiple points into a single long sentence.
|
||||
- In `延伸阅读`, every item must include source attribution in the form: `- [标题](url)|来源`.
|
||||
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other numbering variant.**
|
||||
- **All four sections are required: `今日概览`, `今日重点`, `趋势观察`, `延伸阅读`. Missing any section is a format violation.**
|
||||
|
||||
### Phase 5: User selects articles for IMA deposition
|
||||
|
||||
After Hugo publication, ask the user which articles should be retained for long-term knowledge (separate from the Hugo selection).
|
||||
- The user can select from all `keep` and `review` candidates, including or excluding articles that were already put in Hugo.
|
||||
- Confirm the selection explicitly before proceeding.
|
||||
- Only the user-selected articles will be summarized and uploaded to IMA.
|
||||
|
||||
At this step:
|
||||
|
||||
- the public digest is already in Hugo
|
||||
- the internal review digest stays in chat / operator workflow
|
||||
- the digest is **not** uploaded to IMA
|
||||
- the user decides which articles are worth preserving
|
||||
- the user decides which articles are worth preserving as knowledge notes
|
||||
- the digest is **not** uploaded to IMA as a whole
|
||||
|
||||
### Phase 5: Summarize selected articles
|
||||
### Phase 6: Summarize selected articles
|
||||
|
||||
For every article explicitly selected by the user:
|
||||
|
||||
1. use the `reader` article-summary capability
|
||||
2. **⚠️ 关键:extracted_path 必须传入 pipeline 产生的 individual item 文件,路径为:**
|
||||
2. **⚠️ 关键:extracted_path 必须传入 pipeline 产生的 individual item 文件,路径为:**
|
||||
```
|
||||
outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json
|
||||
outputs/freshrss/rerun/<run-id>/extracted/item-02.extracted.json
|
||||
...
|
||||
```
|
||||
**不要**传 `candidates/openclaw-delivery-payload.json` 或 `outputs/freshrss/extracted/`(这两个路径的 payload 格式与 article-summary workflow 不兼容,会报错**
|
||||
**不要**传 `candidates/openclaw-delivery-payload.json` 或 `outputs/freshrss/extracted/`(这两个路径的 payload 格式与 article-summary workflow 不兼容,会报错**
|
||||
3. for each selected article, run **one summary job per extracted item file**
|
||||
4. pass the corresponding `item_id` from the `candidates` array **去掉 `cand:` 前缀**
|
||||
5. generate one markdown summary per article
|
||||
|
||||
**⚠️ item_id 前缀注意:** `candidates` 里的 ID 格式是 `cand:sha256:xxx`,但 item 文件里的 `item_id` 是 `sha256:xxx`(无 `cand:` 前缀)。传 selected_ids 时**必须去掉前缀**,否则匹配失败。
|
||||
**⚠️ item_id 前缀注意:** `candidates` 里的 ID 格式是 `cand:sha256:xxx`,但 item 文件里的 `item_id` 是 `sha256:xxx`(无 `cand:` 前缀)。传 selected_ids 时**必须去掉前缀**,否则匹配失败。
|
||||
|
||||
**⚠️ item_id 与 extracted 文件编号的对应关系:candidates 数组顺序 ≠ extracted 文件编号顺序。pipeline 提取阶段和 LLM 筛选阶段是两套独立顺序,不能按"第几篇"的位置来映射。**
|
||||
**⚠️ item_id 与 extracted 文件编号的对应关系:candidates 数组顺序 ≠ extracted 文件编号顺序。pipeline 提取阶段和 LLM 筛选阶段是两套独立顺序,不能按"第几篇"的位置来映射。**
|
||||
|
||||
**正确做法:** 传 selected_ids 前,必须先从 `candidates` 数组找到目标文章的 `item_id`(去前缀),再到对应的 `item-XX.extracted.json` 文件里读取其内部的 `article.item_id` 做交叉验证,确认匹配后再传。禁止仅凭"文章在 candidates 里排第几"来推断应该用哪个 item-XX 文件。
|
||||
**正确做法:** 传 selected_ids 前,必须先从 `candidates` 数组找到目标文章的 `item_id`(去前缀),再到对应的 `item-XX.extracted.json` 文件里读取其内部的 `article.item_id` 做交叉验证,确认匹配后再传。禁止仅凭"文章在 candidates 里排第几"来推断应该用哪个 item-XX 文件。
|
||||
|
||||
**⚠️ single-item 输入约束:** 当 `extracted_path` 指向 `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json` 这类单篇文件时,`selected_ids` 只应包含这一个文件对应的单个 `item_id`。不要对单个 extracted 文件传多个 IDs。
|
||||
**⚠️ single-item 输入约束:** 当 `extracted_path` 指向 `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json` 这类单篇文件时,`selected_ids` 只应包含这一个文件对应的单个 `item_id`。不要对单个 extracted 文件传多个 IDs。
|
||||
|
||||
Preferred routes:
|
||||
|
||||
@@ -249,7 +248,7 @@ Hard fallback rule:
|
||||
|
||||
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
|
||||
|
||||
### Phase 6: Upload selected article notes to IMA
|
||||
### Phase 7: Upload selected article notes to IMA
|
||||
|
||||
Upload only the generated single-article markdown summaries to IMA.
|
||||
|
||||
@@ -291,7 +290,7 @@ IMA Markdown layout guidance for selected article deposition:
|
||||
- Before every upload, open and check the actual markdown file that will be uploaded; do not assume the generator already matched IMA style.
|
||||
- The upload target must be an IMA-facing formatted markdown file, not the raw default output from `reader` if the styles differ.
|
||||
- Remove stiff metadata headers such as `Source:` / `Category:` when preparing the IMA-facing Markdown.
|
||||
- Keep source traceability by placing `原文链接:` near the top, followed by the original URL on the next line.
|
||||
- Keep source traceability by placing `原文链接:` near the top, followed by the original URL on the next line.
|
||||
- Break long prose under `核心结论` and `主要论点` into short paragraphs for IMA readability instead of relying on platform auto-formatting.
|
||||
- The final upload filename should normally be `<文章标题>.md`; only when a same-name file already exists should you append a timestamp suffix before `.md`.
|
||||
- If local working files use internal slugs or prefixes for convenience, create or rename a final upload copy before calling IMA upload APIs.
|
||||
|
||||
Reference in New Issue
Block a user