feat(reader-digest-flow): switch selected summaries to async jobs

This commit is contained in:
root
2026-04-10 16:59:55 +08:00
parent 732d5a23cf
commit e97dc5b765
2 changed files with 61 additions and 5 deletions
+33 -2
View File
@@ -52,8 +52,11 @@ Project root:
Default behavior for a normal production run:
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
- run with mark-read enabled
- do not enable `debug_artifacts`
- only skip mark-read if the user explicitly says the run is debug, test, or validation
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
Typical artifacts to inspect after a successful run:
@@ -69,6 +72,7 @@ Notes:
- `digest-brief.json` is the preferred input for **public digest** generation.
- It is a lighter public-only view and currently includes only `keep` candidates.
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
### 2. Generate digest markdown
@@ -146,6 +150,7 @@ Internal review digest constraints:
### 3. Publish Hugo
Publish only the public digest to Hugo.
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
Hard execution steps:
@@ -168,13 +173,19 @@ Expected verification targets:
Provide the internal review digest in chat and ask which articles should be retained.
Hard reporting rule:
- do not send only article titles
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
### 5. Generate selected article summaries
Only do this after the user explicitly confirms which articles to retain.
Preferred MCP tool:
Preferred MCP path:
- `generate_article_summaries`
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
Expected inputs:
@@ -186,6 +197,23 @@ Expected inputs:
Recommended output layout:
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
Recommended production sequence:
1. call `start_article_summary_job`
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
Hard fallback rule:
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
Synchronous helper:
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
CLI fallback:
@@ -199,6 +227,7 @@ python scripts/run_article_summaries.py \
### 6. Upload selected summaries to IMA
Upload only the generated markdown files for the selected articles.
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
Hard execution rules before upload:
@@ -221,6 +250,8 @@ Default target knowledge base:
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
- Only explicitly user-selected articles go to IMA.
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
- Selected article summaries use extracted text, not live refetch.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.