Compare commits
9
Commits
382afd186e
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
022a23b75e | ||
|
|
ade5c58cbc | ||
|
|
1c0bd65e85 | ||
|
|
1be0c4a8b1 | ||
|
|
22c7f1638e | ||
|
|
8dd76a1f86 | ||
|
|
e97dc5b765 | ||
|
|
732d5a23cf | ||
|
|
9d809117d1 |
+176
-84
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: reader-digest-flow
|
name: reader-digest-flow
|
||||||
description: Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, generate a daily report from reader payloads, publish the digest to Hugo, report the digest back in chat, select valuable articles, and then summarize only the selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → digest → selection → knowledge-base flow.
|
description: Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, report candidates for review, let the user decide which articles go into the Hugo daily digest, then generate and publish the public digest, and optionally summarize selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → report → user-selection → Hugo → knowledge-base flow.
|
||||||
---
|
---
|
||||||
|
|
||||||
# Reader Digest Flow
|
# Reader Digest Flow
|
||||||
@@ -9,19 +9,31 @@ Run the reader-based daily digest as a fixed SOP. Treat this skill as the orches
|
|||||||
|
|
||||||
## Core Rules
|
## Core Rules
|
||||||
|
|
||||||
- Run `reader` / MCP for upstream fetching, extraction, filtering, payload generation, and selected-article post-processing.
|
- Run `reader` as the formal upstream MCP workflow service for fetching, extraction, filtering, payload generation, result reading, and selected-article post-processing.
|
||||||
|
- **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.**
|
||||||
|
- Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path.
|
||||||
|
- **Main FreshRSS production runs should default to the async MCP job path**: `start_freshrss_pipeline_job` → `get_freshrss_pipeline_job_status` → `get_freshrss_pipeline_job_result`. Use the synchronous `run_freshrss_openclaw_pipeline` only for debug / light validation / fallback.
|
||||||
|
- **If the main pipeline job fails, do not immediately abandon the run.** First inspect the linked run through `get_run_status` and `inspect_resume_plan`; when the plan returns `recommended_action=resume`, continue through `start_resume_job` → `get_resume_job_status` → `get_resume_job_result`.
|
||||||
|
- **Operational note from validation:** before the main pipeline async job existed, a debug/test run could still complete successfully inside reader even when the synchronous MCP wrapper returned timeout. For historical/debug cases, continue from real run truth using MCP status/result query tools rather than treating the whole flow as failed.
|
||||||
|
- For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths.
|
||||||
|
- **When reader status APIs return reconciled status, always branch on the top-level `status`.** Treat `status_source` and `state_conflict` as explanatory metadata; do not re-derive flow control from stale stage names.
|
||||||
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
|
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
|
||||||
|
- **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.**
|
||||||
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
|
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
|
||||||
- Generate the daily digest markdown from the payload in OpenClaw.
|
- Generate the daily digest markdown from the payload in OpenClaw.
|
||||||
- Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making.
|
- Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making.
|
||||||
- **Public digest and internal review digest are orchestration-layer outputs, and by default are generated by the current OpenClaw session model rather than inheriting reader `LLM_*` settings.**
|
- **Public digest and internal review digest are orchestration-layer outputs, and by default are generated by the current OpenClaw session model rather than inheriting reader `LLM_*` settings.**
|
||||||
- Publish only the public digest to Hugo.
|
- **Publish only the public digest to Hugo, and only after the user has explicitly selected which articles go into the daily digest.**
|
||||||
|
- The default flow is: pipeline → report candidates → **user selects Hugo articles** → generate public digest → publish Hugo → (optional) user selects IMA articles → IMA deposition.
|
||||||
|
- Do NOT generate the public digest or publish Hugo before the user has confirmed the Hugo article selection.
|
||||||
- Report the internal review digest back to the user in chat.
|
- Report the internal review digest back to the user in chat.
|
||||||
- **Do not upload the full daily digest to IMA.**
|
- **Do not upload the full daily digest to IMA.**
|
||||||
- **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection.**
|
- **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection. Once the user has confirmed which articles to keep, proceed directly with IMA deposition and do not ask for a second confirmation about uploading to the knowledge base.**
|
||||||
- Upload **only user-selected articles** to IMA as individual knowledge notes.
|
- Upload **only user-selected articles** to IMA as individual knowledge notes.
|
||||||
- **The formal IMA upload object must be the generated single-article summary content, uploaded directly into the target knowledge base as a Markdown knowledge item (`media_type=7`), not the original article webpage URL.**
|
- **The formal IMA upload object must be the generated single-article summary content, uploaded directly into the target IMA knowledge base as a Markdown knowledge item (`media_type=7`), not the original article webpage URL.**
|
||||||
|
- **For daily deposition, the target must explicitly be the IMA "knowledge-base" type path, not the IMA "notes" type path.**
|
||||||
- **Do not use webpage URL import or note-then-link (`media_type=11`) as the default daily deposition path. Importing original article URLs into IMA can be used only as temporary source collection, and does not count as completing the daily summary deposition flow.**
|
- **Do not use webpage URL import or note-then-link (`media_type=11`) as the default daily deposition path. Importing original article URLs into IMA can be used only as temporary source collection, and does not count as completing the daily summary deposition flow.**
|
||||||
|
- **Adding content as an IMA note, or uploading to any non-knowledge-base container first and then linking it indirectly, does not count as completing the SOP. Completion requires direct knowledge-base ingestion.**
|
||||||
- Default IMA target for daily single-article summaries is the `daily` knowledge base.
|
- Default IMA target for daily single-article summaries is the `daily` knowledge base.
|
||||||
- Resolve that default target from reader `.env` (`IMA_DAILY_KNOWLEDGE_BASE_ID`, `IMA_DAILY_KNOWLEDGE_BASE_NAME`) and verify it at runtime before upload.
|
- Resolve that default target from reader `.env` (`IMA_DAILY_KNOWLEDGE_BASE_ID`, `IMA_DAILY_KNOWLEDGE_BASE_NAME`) and verify it at runtime before upload.
|
||||||
- If the configured daily knowledge base is missing, try to locate it by name; if still missing, create `daily` and continue.
|
- If the configured daily knowledge base is missing, try to locate it by name; if still missing, create `daily` and continue.
|
||||||
@@ -36,137 +48,214 @@ Run the reader-based daily digest as a fixed SOP. Treat this skill as the orches
|
|||||||
|
|
||||||
In the `reader` project, run the FreshRSS pipeline and obtain a real payload.
|
In the `reader` project, run the FreshRSS pipeline and obtain a real payload.
|
||||||
|
|
||||||
|
Formal production start path:
|
||||||
|
|
||||||
|
- `start_freshrss_pipeline_job`
|
||||||
|
- `get_freshrss_pipeline_job_status`
|
||||||
|
- `get_freshrss_pipeline_job_result`
|
||||||
|
|
||||||
|
After the async job succeeds, treat the returned `run_id` as the stable handle for downstream `get_run_status` / `get_delivery_payload` / `get_run_report` reads.
|
||||||
|
|
||||||
Minimum expected artifacts:
|
Minimum expected artifacts:
|
||||||
|
|
||||||
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json`
|
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json` — **精简候选清单**(唯一要读取的文件,含 keep + review)
|
||||||
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`(public digest 优先读取的轻量输入,仅包含 `keep` 候选)
|
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json` — 完整 payload(pipeline 内部产物,digest-brief 由此推导,流程中无需读取)
|
||||||
- `outputs/freshrss/rerun/<run-id>/run-report.json`
|
- `outputs/freshrss/rerun/<run-id>/run-report.json`
|
||||||
- extracted article data, such as:
|
- extracted article data: `outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json` 等(每个候选一篇,格式为 `{"success": true, "article": {...}}`,Phase 6 使用)
|
||||||
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
|
||||||
|
|
||||||
If the pipeline fails, stop and report the exact failure point.
|
If the main pipeline job fails:
|
||||||
|
|
||||||
### Phase 2: Generate daily digest markdown
|
- read the linked run through `get_run_status`
|
||||||
|
- call `inspect_resume_plan(run_id)`
|
||||||
|
- if `recommended_action=resume`, continue with:
|
||||||
|
- `start_resume_job`
|
||||||
|
- `get_resume_job_status`
|
||||||
|
- `get_resume_job_result`
|
||||||
|
- if `recommended_action=read_terminal_result`, continue from the terminal run result instead of retrying
|
||||||
|
- if `recommended_action=start_new_run`, stop and report the exact failure point
|
||||||
|
|
||||||
Read the real run outputs and generate digest markdown for the day.
|
### Phase 2: Report candidates to the user (internal review digest only)
|
||||||
If there is no real payload, stop instead of writing a fake or example digest.
|
|
||||||
|
|
||||||
For **public digest**, prefer `candidates/digest-brief.json` when present; fall back to `candidates/openclaw-delivery-payload.json` only if the brief file is missing.
|
After the pipeline run succeeds, do NOT generate the public digest or publish Hugo yet.
|
||||||
For **internal review digest**, continue reading the full `candidates/openclaw-delivery-payload.json`.
|
Read candidates and produce **only the internal review digest** in a concise format (not the full payload dump).
|
||||||
|
|
||||||
Produce two output views from the same run, preferably in **one generation step**:
|
|
||||||
|
|
||||||
1. **Public digest** — for Hugo / public browsing
|
|
||||||
2. **Internal review digest** — for chat reporting and operator decisions
|
|
||||||
|
|
||||||
Before publishing, persist digest outputs back into the same reader run directory under:
|
|
||||||
|
|
||||||
- `outputs/freshrss/rerun/<run-id>/digest/public_digest.md`
|
|
||||||
- `outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md`
|
|
||||||
- `outputs/freshrss/rerun/<run-id>/digest/combined.json`
|
|
||||||
|
|
||||||
Then write only the public digest into Hugo using this structure:
|
|
||||||
|
|
||||||
- `content/daily/YYYY-MM-DD/index.md`
|
|
||||||
|
|
||||||
The public digest is the browsing layer, not the long-term knowledge layer.
|
|
||||||
It must not expose internal workflow states or operator-facing review labels.
|
|
||||||
|
|
||||||
Recommended **public digest** structure:
|
|
||||||
|
|
||||||
- `今日概览`
|
|
||||||
- `今日重点`
|
|
||||||
- `趋势观察`
|
|
||||||
- `延伸阅读`
|
|
||||||
|
|
||||||
Public digest writing rules:
|
|
||||||
|
|
||||||
- Use `digest-brief.json` as the default source when available.
|
|
||||||
- Treat it as a public-only view that already excludes non-keep items.
|
|
||||||
- Keep the tone suitable for public browsing and Hugo publishing.
|
|
||||||
- Style should follow `references/public-digest-example.md` as the default public-writing example.
|
|
||||||
- Public digest is a **public reading draft / editor-style public note**, not a workflow report.
|
|
||||||
- In `今日概览`, focus on the day’s topic lines, shared signals, and broader industry movement; do **not** describe filtering mechanics or internal selection process.
|
|
||||||
- Do **not** expose internal workflow labels or operator language such as `待确认`, `建议沉淀到 IMA`, `keep/review/drop`, or `selection_decision`.
|
|
||||||
- Explicitly avoid wording such as `共筛出`, `候选`, `保留`, `入选`, `待确认`, `建议沉淀` in public digest.
|
|
||||||
- Prefer concise but information-dense writing.
|
|
||||||
- For each item under `今日重点`, include not only summary and highlights, but also one short editor-style value sentence, for example: `这篇内容更值得关注的原因在于……`.
|
|
||||||
- When rendering highlights in public digest, prefer a short label such as `值得关注:` followed by one item per line, instead of packing multiple points into a single long sentence.
|
|
||||||
- In `延伸阅读`, every item must include source attribution in the form: `- [标题](url)|来源`.
|
|
||||||
|
|
||||||
Recommended **internal review digest** structure:
|
Recommended **internal review digest** structure:
|
||||||
|
|
||||||
- `今日候选概况`
|
- `今日候选概况`
|
||||||
- `已入选重点`
|
- `已入选重点`
|
||||||
- `待你确认`
|
- `待你确认`
|
||||||
- `建议沉淀到 IMA`
|
|
||||||
- `原始候选清单`
|
- `原始候选清单`
|
||||||
|
|
||||||
|
Data source: use `candidates/digest-brief.json` (`top_candidates` array) — it now includes both keep and review articles in a concise, simplified format (title, source, summary, highlights, category). No need to read the full payload.
|
||||||
|
|
||||||
Internal review digest writing rules:
|
Internal review digest writing rules:
|
||||||
|
|
||||||
- Use the full `openclaw-delivery-payload.json`.
|
- Keep this as a human-readable concise review draft, not a raw machine dump.
|
||||||
- Keep this as a human-readable review draft rather than a raw machine dump.
|
- Do **not** display `rank` or `digest_rank` values.
|
||||||
- Do **not** display `rank` values.
|
|
||||||
- Convert machine states to Chinese operator-facing labels:
|
- Convert machine states to Chinese operator-facing labels:
|
||||||
- `keep` → `已入选`
|
- `keep` → `已入选`
|
||||||
- `review` → `待确认`
|
- `review` → `待确认`
|
||||||
- `drop` → `暂不纳入`
|
- `drop` → `暂不纳入`
|
||||||
- For every item under `已入选重点`, include:
|
- For every item under `已入选重点`, include:
|
||||||
- title + source
|
- title + source
|
||||||
- status
|
- status (`已入选`)
|
||||||
- a fuller summary paragraph
|
- summary paragraph (use the concise version from digest-brief)
|
||||||
- a short judgment paragraph explaining why it matters in today's digest
|
- highlights (from digest-brief)
|
||||||
|
- a short judgment paragraph explaining why it matters
|
||||||
- For every item under `待你确认`, include:
|
- For every item under `待你确认`, include:
|
||||||
- title + source
|
- title + source
|
||||||
- status
|
- status (`待确认`)
|
||||||
- a fuller summary paragraph
|
- summary paragraph (concise)
|
||||||
- reason
|
- highlights (if available)
|
||||||
- recommendation
|
- reason(为什么要你确认)
|
||||||
- In `原始候选清单`, also use Chinese status labels instead of raw machine values.
|
- recommendation(建议采纳/不采纳)
|
||||||
|
- In `原始候选清单`, present all articles as a brief list (title + status + one-liner), using Chinese status labels.
|
||||||
|
|
||||||
### Phase 3: Publish to Hugo
|
**Hard reporting requirement:** the chat report must not be only a title list. For every reported article, include at least:
|
||||||
|
- title
|
||||||
|
- one-sentence summary
|
||||||
|
- a short reason explaining why it matters / why it is recommended or pending confirmation
|
||||||
|
|
||||||
Publish only the public digest to Hugo and verify:
|
At this step:
|
||||||
|
|
||||||
- list page works
|
- the public digest is **not** generated yet
|
||||||
- detail page works
|
- Hugo is **not** published yet
|
||||||
- latest digest is visible
|
- the internal review digest is sent to the user in chat
|
||||||
|
- the user decides: **which articles go into the Hugo daily digest** AND separately which articles go into IMA deposition
|
||||||
|
|
||||||
|
### Phase 3: User selects Hugo articles
|
||||||
|
|
||||||
|
The user reviews the internal digest and specifies which articles should appear in the Hugo daily digest.
|
||||||
|
- Confirm the selection explicitly before proceeding.
|
||||||
|
- If the user wants to include some `review` candidate articles, respect that choice.
|
||||||
|
- Only the user-selected articles will appear in the public digest.
|
||||||
|
|
||||||
|
### Phase 4: Generate and publish Hugo public digest
|
||||||
|
|
||||||
|
Generate the public digest markdown **only for the user-selected articles**, then publish to Hugo.
|
||||||
|
|
||||||
|
**⚠️ 格式基准 — 开始生成任何 digest 内容之前,必须先完整阅读 `references/public-digest-example.md` 并以此为格式基准,不得凭记忆或直觉写作。**
|
||||||
|
|
||||||
|
**⚠️ HARD PRE-WRITE CHECKLIST — 对照 `references/public-digest-example.md` 逐项确认后再开始写:**
|
||||||
|
|
||||||
|
1. [ ] frontmatter `summary` 字段已填写(**完整句子**,不是关键词罗列,参考格式:"围绕 XXX 的当日深度观察。")
|
||||||
|
2. [ ] 编号只用 `1.` `2.` `3.`,不用中文数字或罗马数字
|
||||||
|
3. [ ] 四个 section 全部存在:`今日概览` `今日重点` `趋势观察` `延伸阅读`
|
||||||
|
4. [ ] 每篇 `今日重点` 下有:
|
||||||
|
- [ ] 标题(无"来源:"字样,来源仅在延伸阅读标注)
|
||||||
|
- [ ] 摘要段落
|
||||||
|
- [ ] "值得关注:"要点列表(每篇 3 条)
|
||||||
|
- [ ] "这篇更值得关注的原因在于:"段落
|
||||||
|
5. [ ] `延伸阅读` 每条含来源标注:`- [标题](url)|来源`
|
||||||
|
|
||||||
|
Public digest source material: use `candidates/digest-brief.json` — it has all the info needed (title, summary, highlights, source, url). No need to read the full payload.
|
||||||
|
|
||||||
|
Persist the generated digest under:
|
||||||
|
|
||||||
|
- `outputs/freshrss/rerun/<run-id>/digest/public_digest.md`
|
||||||
|
|
||||||
|
Then publish to Hugo:
|
||||||
|
|
||||||
|
1. write the public digest markdown to:
|
||||||
|
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
|
||||||
|
2. immediately redeploy Hugo:
|
||||||
|
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
|
||||||
|
3. verify before moving on:
|
||||||
|
- homepage works
|
||||||
|
- list page works
|
||||||
|
- detail page works
|
||||||
|
- latest digest is visible
|
||||||
|
|
||||||
|
Do not treat "markdown file written" as equivalent to publish success. Hugo publish for this SOP is complete only after redeploy and page verification pass.
|
||||||
|
|
||||||
Do not block on style polish unless the user explicitly asks.
|
Do not block on style polish unless the user explicitly asks.
|
||||||
|
|
||||||
### Phase 4: Report digest back to the user
|
The public digest is the browsing layer, not the long-term knowledge layer.
|
||||||
|
It must not expose internal workflow states or operator-facing review labels.
|
||||||
|
|
||||||
Send the internal review digest in chat and ask the user which articles should be retained for long-term knowledge.
|
Public digest writing rules:
|
||||||
|
|
||||||
|
- Use article content from `candidates/digest-brief.json` for the user-selected subset.
|
||||||
|
- **Hugo public digest includes ONLY the articles the user selected in Phase 3.** Not all `keep` items — only the user's explicit selection.
|
||||||
|
- Keep the tone suitable for public browsing and Hugo publishing.
|
||||||
|
- Style should follow `references/public-digest-example.md` as the default public-writing example.
|
||||||
|
- Public digest is a **public reading draft / editor-style public note**, not a workflow report.
|
||||||
|
- In `今日概览`, focus on the day's topic lines, shared signals, and broader industry movement; do **not** describe filtering mechanics or internal selection process.
|
||||||
|
- Do **not** expose internal workflow labels or operator language such as `待确认`, `建议沉淀到 IMA`, `keep/review/drop`, or `selection_decision`.
|
||||||
|
- Explicitly avoid wording such as `共筛出`, `候选`, `保留`, `入选`, `待确认`, `建议沉淀` in public digest.
|
||||||
|
- Prefer concise but information-dense writing.
|
||||||
|
- For each item under `今日重点`, include not only summary and highlights, but also one short editor-style value sentence, for example: `这篇内容更值得关注的原因在于……`.
|
||||||
|
- When rendering highlights in public digest, prefer a short label such as `值得关注:` followed by one item per line, instead of packing multiple points into a single long sentence.
|
||||||
|
- In `延伸阅读`, every item must include source attribution in the form: `- [标题](url)|来源`.
|
||||||
|
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other numbering variant.**
|
||||||
|
- **All four sections are required: `今日概览`, `今日重点`, `趋势观察`, `延伸阅读`. Missing any section is a format violation.**
|
||||||
|
|
||||||
|
### Phase 5: User selects articles for IMA deposition
|
||||||
|
|
||||||
|
After Hugo publication, ask the user which articles should be retained for long-term knowledge (separate from the Hugo selection).
|
||||||
|
- The user can select from all `keep` and `review` candidates, including or excluding articles that were already put in Hugo.
|
||||||
|
- Confirm the selection explicitly before proceeding.
|
||||||
|
- Only the user-selected articles will be summarized and uploaded to IMA.
|
||||||
|
|
||||||
At this step:
|
At this step:
|
||||||
|
|
||||||
- the public digest is already in Hugo
|
- the public digest is already in Hugo
|
||||||
- the internal review digest stays in chat / operator workflow
|
- the user decides which articles are worth preserving as knowledge notes
|
||||||
- the digest is **not** uploaded to IMA
|
- the digest is **not** uploaded to IMA as a whole
|
||||||
- the user decides which articles are worth preserving
|
|
||||||
|
|
||||||
### Phase 5: Summarize selected articles
|
### Phase 6: Summarize selected articles
|
||||||
|
|
||||||
For every article explicitly selected by the user:
|
For every article explicitly selected by the user:
|
||||||
|
|
||||||
1. use the `reader` article-summary capability
|
1. use the `reader` article-summary capability
|
||||||
2. point it at the existing extracted payload
|
2. **⚠️ 关键:extracted_path 必须传入 pipeline 产生的 individual item 文件,路径为:**
|
||||||
3. pass the selected `item_id` values
|
```
|
||||||
4. generate one markdown summary per article
|
outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json
|
||||||
|
outputs/freshrss/rerun/<run-id>/extracted/item-02.extracted.json
|
||||||
|
...
|
||||||
|
```
|
||||||
|
**不要**传 `candidates/openclaw-delivery-payload.json` 或 `outputs/freshrss/extracted/`(这两个路径的 payload 格式与 article-summary workflow 不兼容,会报错**
|
||||||
|
3. for each selected article, run **one summary job per extracted item file**
|
||||||
|
4. pass the corresponding `item_id` from the `candidates` array **去掉 `cand:` 前缀**
|
||||||
|
5. generate one markdown summary per article
|
||||||
|
|
||||||
|
**⚠️ item_id 前缀注意:** `candidates` 里的 ID 格式是 `cand:sha256:xxx`,但 item 文件里的 `item_id` 是 `sha256:xxx`(无 `cand:` 前缀)。传 selected_ids 时**必须去掉前缀**,否则匹配失败。
|
||||||
|
|
||||||
|
**⚠️ item_id 与 extracted 文件编号的对应关系:candidates 数组顺序 ≠ extracted 文件编号顺序。pipeline 提取阶段和 LLM 筛选阶段是两套独立顺序,不能按"第几篇"的位置来映射。**
|
||||||
|
|
||||||
|
**正确做法:** 传 selected_ids 前,必须先从 `candidates` 数组找到目标文章的 `item_id`(去前缀),再到对应的 `item-XX.extracted.json` 文件里读取其内部的 `article.item_id` 做交叉验证,确认匹配后再传。禁止仅凭"文章在 candidates 里排第几"来推断应该用哪个 item-XX 文件。
|
||||||
|
|
||||||
|
**⚠️ single-item 输入约束:** 当 `extracted_path` 指向 `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json` 这类单篇文件时,`selected_ids` 只应包含这一个文件对应的单个 `item_id`。不要对单个 extracted 文件传多个 IDs。
|
||||||
|
|
||||||
Preferred routes:
|
Preferred routes:
|
||||||
|
|
||||||
- MCP tool: `generate_article_summaries`
|
- async MCP job path:
|
||||||
|
- `start_article_summary_job`
|
||||||
|
- `get_article_summary_job_status`
|
||||||
|
- `get_article_summary_job_result`
|
||||||
|
- local fallback: run the reader article-summary workflow directly inside the project `.venv`
|
||||||
- CLI fallback: `scripts/run_article_summaries.py`
|
- CLI fallback: `scripts/run_article_summaries.py`
|
||||||
|
|
||||||
|
Formal production rule:
|
||||||
|
- For normal production deposition, selected-article summary generation should default to the async MCP job path instead of synchronous `generate_article_summaries`.
|
||||||
|
- Start the job, poll status until `success` / `failed`, then read `written_paths` from the job result.
|
||||||
|
- Treat synchronous `generate_article_summaries` as a debug / light-weight helper, not the default production entry.
|
||||||
|
|
||||||
|
Hard fallback rule:
|
||||||
|
- If the async article-summary job path fails because of MCP/tool-layer timeout, transport failure, or job-launch failure, do **not** stop the daily deposition flow.
|
||||||
|
- In that case, immediately fall back to running the reader article-summary path locally inside `/home/ubuntu/zhu/github/reader` with the project `.venv`.
|
||||||
|
- Treat a successful local article-summary run as equivalent completion for the summary-generation phase; async MCP job is the preferred entry, not a single point of failure.
|
||||||
|
|
||||||
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
|
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
|
||||||
|
|
||||||
### Phase 6: Upload selected article notes to IMA
|
### Phase 7: Upload selected article notes to IMA
|
||||||
|
|
||||||
Upload only the generated single-article markdown summaries to IMA.
|
Upload only the generated single-article markdown summaries to IMA.
|
||||||
|
|
||||||
**Hard gate before upload:** even if the markdown file was generated by `reader`, do **not** upload it to IMA as-is. You must first reformat/check it against the IMA-facing Markdown rules in this skill, then upload the formatted version. Treat `reader` output as article-summary source material, not automatically as final IMA-ready Markdown.
|
**Hard gate before upload:** even if the markdown file was generated by `reader`, do **not** upload it to IMA as-is. You must first reformat/check it against the IMA-facing Markdown rules in this skill, then upload the formatted version. Treat `reader` output as article-summary source material, not automatically as final IMA-ready Markdown.
|
||||||
|
|
||||||
|
**Hard naming rule before upload:** the final uploaded Markdown filename must use the article's user-facing natural title (normally the original Chinese article title) plus `.md`. Do **not** use internal workflow names, slugs, prefixes, or temp filenames such as `ima-*`, `item-*`, `summary-*`, or English-only shorthand as the final IMA object name.
|
||||||
|
|
||||||
Default target knowledge base for this phase:
|
Default target knowledge base for this phase:
|
||||||
|
|
||||||
- `daily`
|
- `daily`
|
||||||
@@ -178,6 +267,7 @@ Completion criteria for this phase:
|
|||||||
|
|
||||||
- a selected article has been summarized from extracted content
|
- a selected article has been summarized from extracted content
|
||||||
- a single-article Markdown summary file has been generated
|
- a single-article Markdown summary file has been generated
|
||||||
|
- the final upload filename has been normalized to a user-facing article title (not an internal slug / temp name)
|
||||||
- the Markdown file is uploaded directly into the target knowledge base as a Markdown knowledge item (`media_type=7`)
|
- the Markdown file is uploaded directly into the target knowledge base as a Markdown knowledge item (`media_type=7`)
|
||||||
- the uploaded object preserves source link context and uses IMA-friendly layout for readability
|
- the uploaded object preserves source link context and uses IMA-friendly layout for readability
|
||||||
- the upload target is the `daily` knowledge base unless the user explicitly requests otherwise
|
- the upload target is the `daily` knowledge base unless the user explicitly requests otherwise
|
||||||
@@ -200,8 +290,10 @@ IMA Markdown layout guidance for selected article deposition:
|
|||||||
- Before every upload, open and check the actual markdown file that will be uploaded; do not assume the generator already matched IMA style.
|
- Before every upload, open and check the actual markdown file that will be uploaded; do not assume the generator already matched IMA style.
|
||||||
- The upload target must be an IMA-facing formatted markdown file, not the raw default output from `reader` if the styles differ.
|
- The upload target must be an IMA-facing formatted markdown file, not the raw default output from `reader` if the styles differ.
|
||||||
- Remove stiff metadata headers such as `Source:` / `Category:` when preparing the IMA-facing Markdown.
|
- Remove stiff metadata headers such as `Source:` / `Category:` when preparing the IMA-facing Markdown.
|
||||||
- Keep source traceability by placing `原文链接:` near the top, followed by the original URL on the next line.
|
- Keep source traceability by placing `原文链接:` near the top, followed by the original URL on the next line.
|
||||||
- Break long prose under `核心结论` and `主要论点` into short paragraphs for IMA readability instead of relying on platform auto-formatting.
|
- Break long prose under `核心结论` and `主要论点` into short paragraphs for IMA readability instead of relying on platform auto-formatting.
|
||||||
|
- The final upload filename should normally be `<文章标题>.md`; only when a same-name file already exists should you append a timestamp suffix before `.md`.
|
||||||
|
- If local working files use internal slugs or prefixes for convenience, create or rename a final upload copy before calling IMA upload APIs.
|
||||||
- Use `references/ima-markdown-example.md` as the formatting example whenever preparing the final upload file.
|
- Use `references/ima-markdown-example.md` as the formatting example whenever preparing the final upload file.
|
||||||
- Completion for IMA upload requires both: (1) upload API success, and (2) the uploaded markdown having passed the above format checks.
|
- Completion for IMA upload requires both: (1) upload API success, and (2) the uploaded markdown having passed the above format checks.
|
||||||
|
|
||||||
|
|||||||
@@ -34,6 +34,8 @@ Concrete operational checklist for the `reader-digest-flow` skill.
|
|||||||
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
|
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
|
||||||
- Do not expose internal review states or operator-facing labels in the public digest.
|
- Do not expose internal review states or operator-facing labels in the public digest.
|
||||||
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
|
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
|
||||||
|
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
|
||||||
|
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
|
||||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
|
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
|
||||||
- Actively report progress after each completed phase.
|
- Actively report progress after each completed phase.
|
||||||
|
|
||||||
@@ -41,6 +43,29 @@ Concrete operational checklist for the `reader-digest-flow` skill.
|
|||||||
|
|
||||||
### 1. Run reader pipeline
|
### 1. Run reader pipeline
|
||||||
|
|
||||||
|
Use the formal MCP workflow path as the default production route.
|
||||||
|
Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting.
|
||||||
|
|
||||||
|
Formal production startup sequence:
|
||||||
|
|
||||||
|
1. `start_freshrss_pipeline_job`
|
||||||
|
2. `get_freshrss_pipeline_job_status`
|
||||||
|
3. `get_freshrss_pipeline_job_result`
|
||||||
|
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
|
||||||
|
|
||||||
|
If the main pipeline job ends in `failed`:
|
||||||
|
|
||||||
|
1. inspect the linked run with `get_run_status`
|
||||||
|
2. call `inspect_resume_plan(run_id)`
|
||||||
|
3. if `recommended_action=resume`, continue with:
|
||||||
|
- `start_resume_job`
|
||||||
|
- `get_resume_job_status`
|
||||||
|
- `get_resume_job_result`
|
||||||
|
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
|
||||||
|
5. if `recommended_action=start_new_run`, stop and report the failure
|
||||||
|
|
||||||
|
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
|
||||||
|
|
||||||
Project root:
|
Project root:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
@@ -49,8 +74,11 @@ Project root:
|
|||||||
|
|
||||||
Default behavior for a normal production run:
|
Default behavior for a normal production run:
|
||||||
|
|
||||||
|
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
|
||||||
- run with mark-read enabled
|
- run with mark-read enabled
|
||||||
|
- do not enable `debug_artifacts`
|
||||||
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
||||||
|
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
|
||||||
|
|
||||||
Typical artifacts to inspect after a successful run:
|
Typical artifacts to inspect after a successful run:
|
||||||
|
|
||||||
@@ -66,6 +94,7 @@ Notes:
|
|||||||
- `digest-brief.json` is the preferred input for **public digest** generation.
|
- `digest-brief.json` is the preferred input for **public digest** generation.
|
||||||
- It is a lighter public-only view and currently includes only `keep` candidates.
|
- It is a lighter public-only view and currently includes only `keep` candidates.
|
||||||
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
|
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
|
||||||
|
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
|
||||||
|
|
||||||
### 2. Generate digest markdown
|
### 2. Generate digest markdown
|
||||||
|
|
||||||
@@ -123,6 +152,9 @@ Public digest constraints:
|
|||||||
- Optimize for concise public readability with solid information density.
|
- Optimize for concise public readability with solid information density.
|
||||||
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
|
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
|
||||||
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
|
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
|
||||||
|
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other variant.**
|
||||||
|
- **All four sections are required. Missing any one is a format violation.**
|
||||||
|
- **Hugo digest must include ALL keep articles from `digest-brief.json`. The IMA deposition subset is a separate downstream step.**
|
||||||
|
|
||||||
Recommended **internal review digest** structure:
|
Recommended **internal review digest** structure:
|
||||||
|
|
||||||
@@ -143,6 +175,18 @@ Internal review digest constraints:
|
|||||||
### 3. Publish Hugo
|
### 3. Publish Hugo
|
||||||
|
|
||||||
Publish only the public digest to Hugo.
|
Publish only the public digest to Hugo.
|
||||||
|
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
|
||||||
|
|
||||||
|
Hard execution steps:
|
||||||
|
|
||||||
|
1. write the public digest markdown to:
|
||||||
|
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
|
||||||
|
2. redeploy Hugo immediately after writing:
|
||||||
|
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
|
||||||
|
3. verify all three URLs before continuing:
|
||||||
|
- `http://127.0.0.1:14322/`
|
||||||
|
- `http://127.0.0.1:14322/daily/`
|
||||||
|
- `http://127.0.0.1:14322/daily/YYYY-MM-DD/`
|
||||||
|
|
||||||
Expected verification targets:
|
Expected verification targets:
|
||||||
|
|
||||||
@@ -154,13 +198,19 @@ Expected verification targets:
|
|||||||
|
|
||||||
Provide the internal review digest in chat and ask which articles should be retained.
|
Provide the internal review digest in chat and ask which articles should be retained.
|
||||||
|
|
||||||
|
Hard reporting rule:
|
||||||
|
- do not send only article titles
|
||||||
|
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
|
||||||
|
|
||||||
### 5. Generate selected article summaries
|
### 5. Generate selected article summaries
|
||||||
|
|
||||||
Only do this after the user explicitly confirms which articles to retain.
|
Only do this after the user explicitly confirms which articles to retain.
|
||||||
|
|
||||||
Preferred MCP tool:
|
Preferred MCP path:
|
||||||
|
|
||||||
- `generate_article_summaries`
|
- `start_article_summary_job`
|
||||||
|
- `get_article_summary_job_status`
|
||||||
|
- `get_article_summary_job_result`
|
||||||
|
|
||||||
Expected inputs:
|
Expected inputs:
|
||||||
|
|
||||||
@@ -169,22 +219,54 @@ Expected inputs:
|
|||||||
- optional `output_dir`
|
- optional `output_dir`
|
||||||
- optional article-summary LLM overrides
|
- optional article-summary LLM overrides
|
||||||
|
|
||||||
|
Single-item rule:
|
||||||
|
|
||||||
|
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
|
||||||
|
- in that case, `selected_ids` should contain only the matching single `item_id`
|
||||||
|
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
|
||||||
|
|
||||||
Recommended output layout:
|
Recommended output layout:
|
||||||
|
|
||||||
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
||||||
|
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||||
|
|
||||||
|
Recommended production sequence:
|
||||||
|
|
||||||
|
1. call `start_article_summary_job`
|
||||||
|
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
|
||||||
|
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
|
||||||
|
|
||||||
|
Hard fallback rule:
|
||||||
|
|
||||||
|
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
|
||||||
|
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
|
||||||
|
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
|
||||||
|
|
||||||
|
Synchronous helper:
|
||||||
|
|
||||||
|
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
|
||||||
|
|
||||||
CLI fallback:
|
CLI fallback:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scripts/run_article_summaries.py \
|
python scripts/run_article_summaries.py \
|
||||||
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
|
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
|
||||||
--ids <item_id_1> <item_id_2> \
|
--ids <item_id_without_cand_prefix> \
|
||||||
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
||||||
```
|
```
|
||||||
|
|
||||||
### 6. Upload selected summaries to IMA
|
### 6. Upload selected summaries to IMA
|
||||||
|
|
||||||
Upload only the generated markdown files for the selected articles.
|
Upload only the generated markdown files for the selected articles.
|
||||||
|
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
|
||||||
|
|
||||||
|
Hard execution rules before upload:
|
||||||
|
|
||||||
|
1. reformat/check the generated markdown into IMA-facing final content
|
||||||
|
2. normalize the final upload filename to `<文章标题>.md`
|
||||||
|
3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name
|
||||||
|
4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md`
|
||||||
|
5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs
|
||||||
|
|
||||||
Default target knowledge base:
|
Default target knowledge base:
|
||||||
|
|
||||||
@@ -199,5 +281,8 @@ Default target knowledge base:
|
|||||||
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
||||||
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
||||||
- Only explicitly user-selected articles go to IMA.
|
- Only explicitly user-selected articles go to IMA.
|
||||||
|
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
|
||||||
|
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
|
||||||
- Selected article summaries use extracted text, not live refetch.
|
- Selected article summaries use extracted text, not live refetch.
|
||||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|
||||||
|
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
|
||||||
|
|||||||
@@ -4,6 +4,11 @@ date = 2026-04-01T16:55:00+08:00
|
|||||||
summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实践的当日观察。"
|
summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实践的当日观察。"
|
||||||
+++
|
+++
|
||||||
|
|
||||||
|
> ⚠️ 格式规范(生成 Hugo 时必须遵守):
|
||||||
|
> - 文章编号:`1.` `2.` `3.`(阿拉伯数字 + 点),禁止 `① ② ③` / `一、二、三` 等变体
|
||||||
|
> - 四个 section 缺一不可:`今日概览` → `今日重点` → `趋势观察` → `延伸阅读`
|
||||||
|
> - 每篇文章结构:标题 → 摘要段 → "值得关注:"三点 → "这篇更值得关注的理由"段
|
||||||
|
|
||||||
# 今日概览
|
# 今日概览
|
||||||
|
|
||||||
今天的公开候选主要集中在 AI Agent 的架构演进、工具化落地与工程化实践三条线索上。相比早期偏概念展示的讨论,这一批内容更强调模块化能力栈、真实部署路径与系统可维护性,说明行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
今天的公开候选主要集中在 AI Agent 的架构演进、工具化落地与工程化实践三条线索上。相比早期偏概念展示的讨论,这一批内容更强调模块化能力栈、真实部署路径与系统可维护性,说明行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
||||||
@@ -11,7 +16,6 @@ summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实
|
|||||||
## 今日重点
|
## 今日重点
|
||||||
|
|
||||||
### 1. 学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变
|
### 1. 学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变
|
||||||
来源:阿里云开发者
|
|
||||||
|
|
||||||
文章分析了 AI 智能体架构从单体 Agent 向模块化 Skills 的范式转变。Anthropic 先后推出 MCP 和 Agent Skills 开放标准,构建了知识、工具、协作和运行分层架构。文章通过一个自动化美化相册的真实项目,对比了 Claude Code 与 OpenClaw 两种实现方案,验证了新架构的可复用性与灵活性。
|
文章分析了 AI 智能体架构从单体 Agent 向模块化 Skills 的范式转变。Anthropic 先后推出 MCP 和 Agent Skills 开放标准,构建了知识、工具、协作和运行分层架构。文章通过一个自动化美化相册的真实项目,对比了 Claude Code 与 OpenClaw 两种实现方案,验证了新架构的可复用性与灵活性。
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,128 @@
|
|||||||
|
---
|
||||||
|
name: reader-keyword-maintenance
|
||||||
|
description: 维护 reader 项目的关键词 review 目录与低频清理 SOP。当用户要求清理或整理 reader 的关键词 review 目录、控制哪些产物长期保留、检查 suggestions JSON 是否可 apply、或执行“只保留当前 bundle / 仅保留最近少量 suggestions JSON / 删除旧 markdown”等低频维护动作时使用。不要用于正式关键词 review 生成,不要用于 reader 日报主链路。
|
||||||
|
---
|
||||||
|
|
||||||
|
# Reader Keyword Maintenance
|
||||||
|
|
||||||
|
这个 skill 负责 reader 关键词治理里的**低频维护动作**,目标是把清理策略、展示稿生成、dry-run 校验和 review 目录维护从 reader 主链路里拆出来,由 OpenClaw 单独承接。
|
||||||
|
|
||||||
|
## 角色边界
|
||||||
|
|
||||||
|
这个 skill 负责:
|
||||||
|
|
||||||
|
- 检查 `outputs/term_index/review/` 当前有哪些产物
|
||||||
|
- 判断哪些 review 产物应长期保留、短期保留或可清理
|
||||||
|
- 用 `apply_term_suggestions.py --dry-run` 做 apply 前校验
|
||||||
|
- 执行低频清理和收尾动作
|
||||||
|
|
||||||
|
这个 skill 不负责:
|
||||||
|
|
||||||
|
- 正式生成关键词 review 建议
|
||||||
|
- 充当 alias / stopword 正式 review 决策器
|
||||||
|
- 替代 reader 侧的 `keyword-cleanup-review`
|
||||||
|
|
||||||
|
如果用户要的是“做一轮正式关键词 review / 生成正式 suggestions JSON”,应由 OpenClaw 去调用 reader 侧 `keyword-cleanup-review`,而不是由本 skill 直接承担。
|
||||||
|
|
||||||
|
## 适用场景
|
||||||
|
|
||||||
|
当用户要求以下事情时使用:
|
||||||
|
|
||||||
|
- 查看或解释 `reader/outputs/term_index/review/` 下有哪些文件
|
||||||
|
- 清理旧的 review bundle、old suggestions、old markdown
|
||||||
|
- 只保留当前 bundle 或最近少量 suggestions JSON
|
||||||
|
- 检查 suggestions JSON 是否可被 `apply_term_suggestions.py` 消费
|
||||||
|
- 执行一次人工 review 准备动作,但不改变 reader 主链路设计
|
||||||
|
|
||||||
|
不要用于:
|
||||||
|
|
||||||
|
- 日报正式生产运行
|
||||||
|
- 生成 daily term index / term_stats 主链路
|
||||||
|
- 自动写配置作为默认行为
|
||||||
|
- 修改 reader 仓库里的主 SOP 作为低频维护动作的默认入口
|
||||||
|
- 正式生成 alias / stopword / interest / watch 建议
|
||||||
|
|
||||||
|
## 默认口径
|
||||||
|
|
||||||
|
长期保留:
|
||||||
|
|
||||||
|
- `data/term_index/daily/*.json`
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
- `configs/filter_context.personal.json`
|
||||||
|
- `configs/term_watchlist.json`
|
||||||
|
- `configs/term_aliases.json`
|
||||||
|
- `configs/term_stopwords.json`
|
||||||
|
- `configs/term_change_log.json`
|
||||||
|
|
||||||
|
正式建议产物(短期保留):
|
||||||
|
|
||||||
|
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||||
|
|
||||||
|
临时工作文件 / 展示层:
|
||||||
|
|
||||||
|
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||||
|
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||||
|
|
||||||
|
更细的保留/清理判断,读取:
|
||||||
|
|
||||||
|
- `references/cleanup-policy.md`
|
||||||
|
|
||||||
|
## 固定 SOP
|
||||||
|
|
||||||
|
### 1. 检查 review 目录
|
||||||
|
|
||||||
|
- 查看 `outputs/term_index/review/` 当前有哪些文件
|
||||||
|
- 分类成:
|
||||||
|
- 长期保留
|
||||||
|
- 短期保留
|
||||||
|
- 临时工作文件
|
||||||
|
|
||||||
|
### 2. 准备人工审阅
|
||||||
|
|
||||||
|
- 确认最新 bundle 存在
|
||||||
|
- 确认最新 suggestions JSON 存在
|
||||||
|
- 如需正式生成 suggestions 或 Markdown 展示稿,应转到 reader 侧 `keyword-cleanup-review`
|
||||||
|
|
||||||
|
### 3. dry-run 校验 apply
|
||||||
|
|
||||||
|
- 仅通过 `scripts/apply_term_suggestions.py --dry-run` 校验 suggestions JSON 是否可消费
|
||||||
|
- 不直接落配置
|
||||||
|
|
||||||
|
### 4. 清理 review 产物
|
||||||
|
|
||||||
|
- 删除旧 markdown 展示稿
|
||||||
|
- 默认只保留当前/latest bundle
|
||||||
|
- 保守清理旧 suggestions JSON
|
||||||
|
- 不碰 facts/state/config files
|
||||||
|
|
||||||
|
## 常用动作
|
||||||
|
|
||||||
|
### 1. 验证 suggestions JSON 能否被 apply 消费
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd /home/ubuntu/zhu/github/reader && \
|
||||||
|
/usr/bin/python3.11 scripts/apply_term_suggestions.py \
|
||||||
|
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||||
|
--accept-interest "Claude Code" \
|
||||||
|
--dry-run
|
||||||
|
```
|
||||||
|
|
||||||
|
## 清理策略
|
||||||
|
|
||||||
|
保守策略:
|
||||||
|
|
||||||
|
- bundle:默认只保留当前最新一份
|
||||||
|
- markdown:默认不归档,需要时现生成
|
||||||
|
- suggestions JSON:只保留最近少量几份或已应用过的记录
|
||||||
|
|
||||||
|
## 维护原则
|
||||||
|
|
||||||
|
- 优先减少中间产物,不要让 review 工作文件变成长期资产
|
||||||
|
- JSON suggestions 是 review / apply 之间唯一正式输入
|
||||||
|
- Markdown 只是展示层,不是系统真相
|
||||||
|
- 任何清理动作前,先确认不会影响当前待审阅或待 apply 的 suggestions JSON
|
||||||
|
- 正式关键词 review 的生成入口不在这里,而在 reader 侧 `keyword-cleanup-review`
|
||||||
|
|
||||||
|
执行清理或 review-prep 前,先读取:
|
||||||
|
|
||||||
|
- `references/maintenance-checklist.md`
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
# Alias Review Example
|
||||||
|
|
||||||
|
Use this as the default style reference when preparing a low-frequency alias review report.
|
||||||
|
|
||||||
|
## Suggested merges
|
||||||
|
|
||||||
|
- `Claude code -> Claude Code`
|
||||||
|
- 理由:明显属于同一产品名,仅是大小写写法不一致。
|
||||||
|
- 证据:`Claude Code` 在最近多日持续出现,而小写写法只是在少量上下文中作为变体出现。
|
||||||
|
|
||||||
|
- `Sub-Agent -> SubAgent`
|
||||||
|
- 理由:更像词形差异,不构成新的独立概念。
|
||||||
|
- 证据:两者都围绕同一 agent 架构语境出现,且没有稳定区分语义。
|
||||||
|
|
||||||
|
## Not recommended to merge
|
||||||
|
|
||||||
|
- `Skills ↔ Agent Skills`
|
||||||
|
- 理由:前者过泛,后者更具体,当前强行归并会损失粒度。
|
||||||
|
|
||||||
|
- `Anthropic ↔ Claude`
|
||||||
|
- 理由:公司名与产品名并不等价,不应直接视为一个关键词。
|
||||||
|
|
||||||
|
## Needs human judgment
|
||||||
|
|
||||||
|
- `AI助手 ↔ AI Agent`
|
||||||
|
- 风险点:语义可能接近,但中文表述范围更宽,是否并入需要结合实际使用语境判断。
|
||||||
|
|
||||||
|
## Style notes
|
||||||
|
|
||||||
|
- Keep the report concise.
|
||||||
|
- Give judgment first, then evidence.
|
||||||
|
- Do not claim anything is already applied.
|
||||||
|
- Prefer conservative proposals over broad semantic merging.
|
||||||
@@ -0,0 +1,106 @@
|
|||||||
|
# Cleanup Policy
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
This reference defines how `reader-keyword-maintenance` should treat keyword governance artifacts.
|
||||||
|
|
||||||
|
The goal is to keep long-term assets small and stable while allowing review-time working files to exist when needed.
|
||||||
|
|
||||||
|
This policy is for low-frequency maintenance only.
|
||||||
|
It does not replace the formal keyword review generation flow in reader.
|
||||||
|
|
||||||
|
## Artifact classes
|
||||||
|
|
||||||
|
### Long-term assets
|
||||||
|
|
||||||
|
Keep these by default:
|
||||||
|
|
||||||
|
- `data/term_index/daily/*.json`
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
- `configs/filter_context.personal.json`
|
||||||
|
- `configs/term_watchlist.json`
|
||||||
|
- `configs/term_aliases.json`
|
||||||
|
- `configs/term_stopwords.json`
|
||||||
|
- `configs/term_change_log.json`
|
||||||
|
|
||||||
|
These are facts or active state.
|
||||||
|
|
||||||
|
### Short-term decision artifacts
|
||||||
|
|
||||||
|
Keep selectively:
|
||||||
|
|
||||||
|
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||||
|
|
||||||
|
Recommended policy:
|
||||||
|
|
||||||
|
- keep only recent few files, or
|
||||||
|
- keep only files that were actually used for apply decisions
|
||||||
|
|
||||||
|
### Temporary working/display files
|
||||||
|
|
||||||
|
Treat as disposable unless the user explicitly asks to archive them:
|
||||||
|
|
||||||
|
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||||
|
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||||
|
|
||||||
|
Recommended policy:
|
||||||
|
|
||||||
|
- bundle: keep only latest current file
|
||||||
|
- markdown: generate on demand, do not archive by default
|
||||||
|
|
||||||
|
## Default maintenance actions
|
||||||
|
|
||||||
|
### Safe checks
|
||||||
|
|
||||||
|
Before deleting anything:
|
||||||
|
|
||||||
|
1. confirm the target is not the current bundle under active review
|
||||||
|
2. confirm the target suggestions JSON is not the one about to be applied
|
||||||
|
3. never delete configs or facts during review cleanup
|
||||||
|
|
||||||
|
### Safe cleanup order
|
||||||
|
|
||||||
|
1. remove or overwrite old markdown display drafts
|
||||||
|
2. keep only latest bundle file
|
||||||
|
3. prune old suggestions JSON files conservatively
|
||||||
|
|
||||||
|
## Decision rules
|
||||||
|
|
||||||
|
### When user says "clean review artifacts"
|
||||||
|
|
||||||
|
Default action:
|
||||||
|
|
||||||
|
- keep facts/state untouched
|
||||||
|
- keep current suggestions JSON
|
||||||
|
- remove markdown drafts if they are old and clearly derived display files
|
||||||
|
- keep bundle only as current working file
|
||||||
|
|
||||||
|
### When user says "prepare manual review"
|
||||||
|
|
||||||
|
Default action:
|
||||||
|
|
||||||
|
- ensure latest bundle exists
|
||||||
|
- ensure latest suggestions JSON exists
|
||||||
|
- do not generate formal suggestions or markdown here by default
|
||||||
|
- if the user wants formal review material, route to reader-side `keyword-cleanup-review`
|
||||||
|
|
||||||
|
### When user says "review alias candidates" or "review stopword candidates"
|
||||||
|
|
||||||
|
Default action:
|
||||||
|
|
||||||
|
- explain that formal keyword review generation belongs to reader-side `keyword-cleanup-review`
|
||||||
|
- keep this skill focused on maintenance, cleanup, display generation, and dry-run validation
|
||||||
|
- only proceed here if the user explicitly wants a low-frequency maintenance view rather than the formal review flow
|
||||||
|
|
||||||
|
### When user says "what can be deleted"
|
||||||
|
|
||||||
|
Explain in three buckets:
|
||||||
|
|
||||||
|
- must keep
|
||||||
|
- can keep temporarily
|
||||||
|
- safe to regenerate/delete
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
This policy does not change reader production logic.
|
||||||
|
It only governs low-frequency maintenance and cleanup decisions in OpenClaw.
|
||||||
@@ -0,0 +1,55 @@
|
|||||||
|
# Maintenance Checklist
|
||||||
|
|
||||||
|
Use this checklist before doing any cleanup or review-maintenance action for reader keyword artifacts.
|
||||||
|
|
||||||
|
## Pre-check
|
||||||
|
|
||||||
|
1. Confirm the current repo root is `/home/ubuntu/zhu/github/reader`
|
||||||
|
2. Confirm the user asked for a maintenance / cleanup / review-prep action
|
||||||
|
3. If the user actually wants formal keyword review generation, route to reader-side `keyword-cleanup-review` instead of using this maintenance skill
|
||||||
|
4. Identify whether the action targets:
|
||||||
|
- current bundle
|
||||||
|
- suggestions JSON
|
||||||
|
- markdown display draft
|
||||||
|
- old review outputs
|
||||||
|
|
||||||
|
## Safety check
|
||||||
|
|
||||||
|
Before deleting or overwriting anything:
|
||||||
|
|
||||||
|
1. Do not touch:
|
||||||
|
- `data/term_index/daily/*.json`
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
- `configs/filter_context.personal.json`
|
||||||
|
- `configs/term_watchlist.json`
|
||||||
|
- `configs/term_aliases.json`
|
||||||
|
- `configs/term_stopwords.json`
|
||||||
|
- `configs/term_change_log.json`
|
||||||
|
2. Confirm the suggestions JSON to keep is not the one about to be applied
|
||||||
|
3. Treat markdown drafts as disposable only after confirming they are display-only artifacts
|
||||||
|
|
||||||
|
## Review-prep flow
|
||||||
|
|
||||||
|
When preparing manual review:
|
||||||
|
|
||||||
|
1. Ensure the latest bundle exists
|
||||||
|
2. Ensure the latest suggestions JSON exists
|
||||||
|
3. Generate markdown only if the user explicitly wants human-readable review material
|
||||||
|
4. Prefer showing conclusions in chat before creating more files
|
||||||
|
|
||||||
|
## Cleanup flow
|
||||||
|
|
||||||
|
Recommended order:
|
||||||
|
|
||||||
|
1. remove old markdown display drafts
|
||||||
|
2. keep only the current/latest bundle
|
||||||
|
3. prune old suggestions JSON conservatively
|
||||||
|
4. leave facts/state/config untouched
|
||||||
|
|
||||||
|
## Post-check
|
||||||
|
|
||||||
|
After the action:
|
||||||
|
|
||||||
|
1. verify the expected kept files still exist
|
||||||
|
2. verify no config file was accidentally changed
|
||||||
|
3. summarize what was kept vs removed
|
||||||
Reference in New Issue
Block a user