feat(reader-digest-flow): switch selected summaries to async jobs
This commit is contained in:
@@ -10,20 +10,26 @@ Run the reader-based daily digest as a fixed SOP. Treat this skill as the orches
|
|||||||
## Core Rules
|
## Core Rules
|
||||||
|
|
||||||
- Run `reader` as the formal upstream MCP workflow service for fetching, extraction, filtering, payload generation, result reading, and selected-article post-processing.
|
- Run `reader` as the formal upstream MCP workflow service for fetching, extraction, filtering, payload generation, result reading, and selected-article post-processing.
|
||||||
|
- **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.**
|
||||||
- Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path.
|
- Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path.
|
||||||
|
- **Operational note from validation:** a debug/test run may still complete successfully inside reader even when the synchronous MCP wrapper returns timeout. In that case, use MCP status/result query tools (`list_runs`, `get_run_report`, `get_delivery_payload`, article-summary async job status/result tools) to continue from the real run truth, rather than treating the whole flow as failed.
|
||||||
- For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths.
|
- For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths.
|
||||||
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
|
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
|
||||||
|
- **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.**
|
||||||
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
|
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
|
||||||
- Generate the daily digest markdown from the payload in OpenClaw.
|
- Generate the daily digest markdown from the payload in OpenClaw.
|
||||||
- Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making.
|
- Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making.
|
||||||
- **Public digest and internal review digest are orchestration-layer outputs, and by default are generated by the current OpenClaw session model rather than inheriting reader `LLM_*` settings.**
|
- **Public digest and internal review digest are orchestration-layer outputs, and by default are generated by the current OpenClaw session model rather than inheriting reader `LLM_*` settings.**
|
||||||
- Publish only the public digest to Hugo.
|
- Publish only the public digest to Hugo.
|
||||||
|
- Once a normal daily digest run succeeds and a public digest is generated, proceed directly with Hugo publication as the default action; do not ask the user for a separate confirmation about generating or publishing the Hugo page unless the user explicitly says to skip Hugo.
|
||||||
- Report the internal review digest back to the user in chat.
|
- Report the internal review digest back to the user in chat.
|
||||||
- **Do not upload the full daily digest to IMA.**
|
- **Do not upload the full daily digest to IMA.**
|
||||||
- **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection.**
|
- **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection. Once the user has confirmed which articles to keep, proceed directly with IMA deposition and do not ask for a second confirmation about uploading to the knowledge base.**
|
||||||
- Upload **only user-selected articles** to IMA as individual knowledge notes.
|
- Upload **only user-selected articles** to IMA as individual knowledge notes.
|
||||||
- **The formal IMA upload object must be the generated single-article summary content, uploaded directly into the target knowledge base as a Markdown knowledge item (`media_type=7`), not the original article webpage URL.**
|
- **The formal IMA upload object must be the generated single-article summary content, uploaded directly into the target IMA knowledge base as a Markdown knowledge item (`media_type=7`), not the original article webpage URL.**
|
||||||
|
- **For daily deposition, the target must explicitly be the IMA "knowledge-base" type path, not the IMA "notes" type path.**
|
||||||
- **Do not use webpage URL import or note-then-link (`media_type=11`) as the default daily deposition path. Importing original article URLs into IMA can be used only as temporary source collection, and does not count as completing the daily summary deposition flow.**
|
- **Do not use webpage URL import or note-then-link (`media_type=11`) as the default daily deposition path. Importing original article URLs into IMA can be used only as temporary source collection, and does not count as completing the daily summary deposition flow.**
|
||||||
|
- **Adding content as an IMA note, or uploading to any non-knowledge-base container first and then linking it indirectly, does not count as completing the SOP. Completion requires direct knowledge-base ingestion.**
|
||||||
- Default IMA target for daily single-article summaries is the `daily` knowledge base.
|
- Default IMA target for daily single-article summaries is the `daily` knowledge base.
|
||||||
- Resolve that default target from reader `.env` (`IMA_DAILY_KNOWLEDGE_BASE_ID`, `IMA_DAILY_KNOWLEDGE_BASE_NAME`) and verify it at runtime before upload.
|
- Resolve that default target from reader `.env` (`IMA_DAILY_KNOWLEDGE_BASE_ID`, `IMA_DAILY_KNOWLEDGE_BASE_NAME`) and verify it at runtime before upload.
|
||||||
- If the configured daily knowledge base is missing, try to locate it by name; if still missing, create `daily` and continue.
|
- If the configured daily knowledge base is missing, try to locate it by name; if still missing, create `daily` and continue.
|
||||||
@@ -150,6 +156,11 @@ Do not block on style polish unless the user explicitly asks.
|
|||||||
|
|
||||||
Send the internal review digest in chat and ask the user which articles should be retained for long-term knowledge.
|
Send the internal review digest in chat and ask the user which articles should be retained for long-term knowledge.
|
||||||
|
|
||||||
|
**Hard reporting requirement:** the chat report must not be only a title list. For every reported article, include at least:
|
||||||
|
- title
|
||||||
|
- one-sentence summary
|
||||||
|
- a short reason explaining why it matters / why it is recommended or pending confirmation
|
||||||
|
|
||||||
At this step:
|
At this step:
|
||||||
|
|
||||||
- the public digest is already in Hugo
|
- the public digest is already in Hugo
|
||||||
@@ -168,9 +179,23 @@ For every article explicitly selected by the user:
|
|||||||
|
|
||||||
Preferred routes:
|
Preferred routes:
|
||||||
|
|
||||||
- MCP tool: `generate_article_summaries`
|
- async MCP job path:
|
||||||
|
- `start_article_summary_job`
|
||||||
|
- `get_article_summary_job_status`
|
||||||
|
- `get_article_summary_job_result`
|
||||||
|
- local fallback: run the reader article-summary workflow directly inside the project `.venv`
|
||||||
- CLI fallback: `scripts/run_article_summaries.py`
|
- CLI fallback: `scripts/run_article_summaries.py`
|
||||||
|
|
||||||
|
Formal production rule:
|
||||||
|
- For normal production deposition, selected-article summary generation should default to the async MCP job path instead of synchronous `generate_article_summaries`.
|
||||||
|
- Start the job, poll status until `success` / `failed`, then read `written_paths` from the job result.
|
||||||
|
- Treat synchronous `generate_article_summaries` as a debug / light-weight helper, not the default production entry.
|
||||||
|
|
||||||
|
Hard fallback rule:
|
||||||
|
- If the async article-summary job path fails because of MCP/tool-layer timeout, transport failure, or job-launch failure, do **not** stop the daily deposition flow.
|
||||||
|
- In that case, immediately fall back to running the reader article-summary path locally inside `/home/ubuntu/zhu/github/reader` with the project `.venv`.
|
||||||
|
- Treat a successful local article-summary run as equivalent completion for the summary-generation phase; async MCP job is the preferred entry, not a single point of failure.
|
||||||
|
|
||||||
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
|
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
|
||||||
|
|
||||||
### Phase 6: Upload selected article notes to IMA
|
### Phase 6: Upload selected article notes to IMA
|
||||||
|
|||||||
@@ -52,8 +52,11 @@ Project root:
|
|||||||
|
|
||||||
Default behavior for a normal production run:
|
Default behavior for a normal production run:
|
||||||
|
|
||||||
|
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
|
||||||
- run with mark-read enabled
|
- run with mark-read enabled
|
||||||
|
- do not enable `debug_artifacts`
|
||||||
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
||||||
|
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
|
||||||
|
|
||||||
Typical artifacts to inspect after a successful run:
|
Typical artifacts to inspect after a successful run:
|
||||||
|
|
||||||
@@ -69,6 +72,7 @@ Notes:
|
|||||||
- `digest-brief.json` is the preferred input for **public digest** generation.
|
- `digest-brief.json` is the preferred input for **public digest** generation.
|
||||||
- It is a lighter public-only view and currently includes only `keep` candidates.
|
- It is a lighter public-only view and currently includes only `keep` candidates.
|
||||||
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
|
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
|
||||||
|
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
|
||||||
|
|
||||||
### 2. Generate digest markdown
|
### 2. Generate digest markdown
|
||||||
|
|
||||||
@@ -146,6 +150,7 @@ Internal review digest constraints:
|
|||||||
### 3. Publish Hugo
|
### 3. Publish Hugo
|
||||||
|
|
||||||
Publish only the public digest to Hugo.
|
Publish only the public digest to Hugo.
|
||||||
|
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
|
||||||
|
|
||||||
Hard execution steps:
|
Hard execution steps:
|
||||||
|
|
||||||
@@ -168,13 +173,19 @@ Expected verification targets:
|
|||||||
|
|
||||||
Provide the internal review digest in chat and ask which articles should be retained.
|
Provide the internal review digest in chat and ask which articles should be retained.
|
||||||
|
|
||||||
|
Hard reporting rule:
|
||||||
|
- do not send only article titles
|
||||||
|
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
|
||||||
|
|
||||||
### 5. Generate selected article summaries
|
### 5. Generate selected article summaries
|
||||||
|
|
||||||
Only do this after the user explicitly confirms which articles to retain.
|
Only do this after the user explicitly confirms which articles to retain.
|
||||||
|
|
||||||
Preferred MCP tool:
|
Preferred MCP path:
|
||||||
|
|
||||||
- `generate_article_summaries`
|
- `start_article_summary_job`
|
||||||
|
- `get_article_summary_job_status`
|
||||||
|
- `get_article_summary_job_result`
|
||||||
|
|
||||||
Expected inputs:
|
Expected inputs:
|
||||||
|
|
||||||
@@ -186,6 +197,23 @@ Expected inputs:
|
|||||||
Recommended output layout:
|
Recommended output layout:
|
||||||
|
|
||||||
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
||||||
|
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||||
|
|
||||||
|
Recommended production sequence:
|
||||||
|
|
||||||
|
1. call `start_article_summary_job`
|
||||||
|
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
|
||||||
|
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
|
||||||
|
|
||||||
|
Hard fallback rule:
|
||||||
|
|
||||||
|
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
|
||||||
|
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
|
||||||
|
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
|
||||||
|
|
||||||
|
Synchronous helper:
|
||||||
|
|
||||||
|
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
|
||||||
|
|
||||||
CLI fallback:
|
CLI fallback:
|
||||||
|
|
||||||
@@ -199,6 +227,7 @@ python scripts/run_article_summaries.py \
|
|||||||
### 6. Upload selected summaries to IMA
|
### 6. Upload selected summaries to IMA
|
||||||
|
|
||||||
Upload only the generated markdown files for the selected articles.
|
Upload only the generated markdown files for the selected articles.
|
||||||
|
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
|
||||||
|
|
||||||
Hard execution rules before upload:
|
Hard execution rules before upload:
|
||||||
|
|
||||||
@@ -221,6 +250,8 @@ Default target knowledge base:
|
|||||||
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
||||||
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
||||||
- Only explicitly user-selected articles go to IMA.
|
- Only explicitly user-selected articles go to IMA.
|
||||||
|
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
|
||||||
|
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
|
||||||
- Selected article summaries use extracted text, not live refetch.
|
- Selected article summaries use extracted text, not live refetch.
|
||||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|
||||||
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
|
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
|
||||||
|
|||||||
Reference in New Issue
Block a user