docs: 添加 Agent Skill 到项目仓库
- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/) - README 新增 Agent Skill 章节说明供 Agent 使用的工作流
This commit is contained in:
@@ -0,0 +1,48 @@
|
||||
# Content Extraction Pipeline
|
||||
|
||||
How the pipeline turns FreshRSS items into extractable article text.
|
||||
|
||||
## Core Rule: FreshRSS items never re-fetch the original URL
|
||||
|
||||
**FreshRSS is an RSS-only upstream.** The pipeline *never* makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by `RSS_ONLY_UPSTREAMS = {"freshrss"}` in `pipeline.py`.
|
||||
|
||||
The only exception: non-FreshRSS upstreams (future sources that don't set `upstream: freshrss`) may trigger `fetch_html()` as a fallback.
|
||||
|
||||
## Content source priority chain
|
||||
|
||||
`content_loader.py` tries sources in this order, using the **first one with ≥500 readable characters**:
|
||||
|
||||
| Priority | Source | Meaning |
|
||||
|----------|--------|---------|
|
||||
| 1 | `raw_html` | HTML pre-injected via `ExtractionInput.raw_html`. Rarely used in normal FreshRSS runs. |
|
||||
| 2 | `item.raw_content` | RSS `<content:encoded>` — the full article body. Some feeds provide this; many don't. |
|
||||
| 3 | `item.raw_summary` | RSS `<description>` — the summary/snippet field. **This is the most common source in current runs.** |
|
||||
| 4 | `rss_content` | RSS content from non-item sources. |
|
||||
| — | `none` | Nothing usable → raises `RSS_CONTENT_MISSING` for FreshRSS items (because fetch is skipped). |
|
||||
|
||||
## How `content_source` maps to actual text quality
|
||||
|
||||
The `content_source` field in every `item-XX.extracted.json` tells you what the pipeline actually used:
|
||||
|
||||
- **`item.raw_content`** → Full article text from RSS `<content:encoded>`. Best quality, same as reading the original page.
|
||||
- **`item.raw_summary`** → RSS summary/description only. **Not the full article.** Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text.
|
||||
- **`rss_content`** → From standalone RSS content. Quality depends on the feed.
|
||||
- **`fetched_html`** → HTML fetched from the original URL (**never happens for FreshRSS**; only for non-FreshRSS upstreams).
|
||||
|
||||
## What this means for digest quality
|
||||
|
||||
If you see `content_source: item.raw_summary` in the extracted files (current norm), the AI is summarizing from a **feed summary/snippet**, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions.
|
||||
|
||||
To improve quality: either find feeds that provide full `<content:encoded>`, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters).
|
||||
|
||||
## Quick check
|
||||
|
||||
```bash
|
||||
# Check content_source for latest run
|
||||
grep -h "content_source" outputs/freshrss/rerun/*/extracted/item-*.json | sort | uniq -c
|
||||
```
|
||||
|
||||
## Relevant code paths
|
||||
|
||||
- `src/summary_mcp/core/pipeline.py` — `RSS_ONLY_UPSTREAMS`, `_should_skip_fetch()`, `extract_content()`
|
||||
- `src/summary_mcp/core/content_loader.py` — `choose_inline_content()` priority chain, `fetch_html()` (never called for FreshRSS)
|
||||
@@ -0,0 +1,26 @@
|
||||
# Feishu Markdown Format Notes
|
||||
|
||||
## Background
|
||||
|
||||
Hermes' Feishu gateway (`gateway/platforms/feishu.py`) sends outbound messages
|
||||
through `_build_outbound_payload`, which checks content for markdown patterns
|
||||
to decide how to send:
|
||||
|
||||
- Content matches `_MARKDOWN_HINT_RE` (bold, lists, code, links, etc.)
|
||||
→ sent as Feishu `post` type using `md` elements → renders correctly.
|
||||
|
||||
- Content matches `_MARKDOWN_TABLE_RE` (a markdown table header + separator)
|
||||
→ **entire message** forced to `text` type (plain text) → no rendering.
|
||||
|
||||
The root cause is that the `_build_markdown_post_payload` helper wraps content
|
||||
in `{"tag": "md", "text": "..."}` elements, and Feishu's `md` element does not
|
||||
support table rendering. There is no table-to-native-Feishu-table conversion.
|
||||
|
||||
## Rules for Feishu output
|
||||
|
||||
- **Never use markdown tables** in any message delivered via Feishu.
|
||||
A single table anywhere in the message forces the whole message to plain text.
|
||||
- Prefer bullet lists, sections with headings, or inline formatting instead.
|
||||
- Bold (`**bold**`), inline code (`` `code` ``), unordered lists (`- item`),
|
||||
ordered lists (`1. item`), and links all work correctly.
|
||||
- Code fences (``` ``` ```) work but may have edge cases with trailing content.
|
||||
@@ -0,0 +1,328 @@
|
||||
# Reader Digest Flow Reference
|
||||
|
||||
## Purpose
|
||||
|
||||
Concrete operational checklist for the `reader-digest-flow` skill.
|
||||
|
||||
## Default Operating Model
|
||||
|
||||
### Layering
|
||||
|
||||
- `reader` layer:
|
||||
- FreshRSS pull
|
||||
- extraction
|
||||
- summary/filter/payload generation
|
||||
- selected-article summary capability
|
||||
- OpenClaw / skill layer:
|
||||
- public digest generation
|
||||
- internal review digest generation
|
||||
- Hugo publishing
|
||||
- chat reporting
|
||||
- user confirmation handling
|
||||
- calling selected-article summaries
|
||||
- IMA upload orchestration
|
||||
- Hugo layer:
|
||||
- public digest browsing and archive only
|
||||
- IMA layer:
|
||||
- long-term storage for selected article notes only
|
||||
|
||||
### Hard rules
|
||||
|
||||
- Do not upload the full digest to IMA.
|
||||
- Upload only explicitly user-selected articles to IMA.
|
||||
- Do not generate a digest without a real payload.
|
||||
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
|
||||
- Do not expose internal review states or operator-facing labels in the public digest.
|
||||
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
|
||||
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
|
||||
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
|
||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
|
||||
- Actively report progress after each completed phase.
|
||||
|
||||
## Step-by-step checklist
|
||||
|
||||
### 1. Run reader pipeline
|
||||
|
||||
Use the formal MCP workflow path as the default production route.
|
||||
Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting.
|
||||
|
||||
Formal production startup sequence:
|
||||
|
||||
1. `start_freshrss_pipeline_job`
|
||||
2. `get_freshrss_pipeline_job_status`
|
||||
3. `get_freshrss_pipeline_job_result`
|
||||
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
|
||||
|
||||
If the main pipeline job ends in `failed`:
|
||||
|
||||
1. inspect the linked run with `get_run_status`
|
||||
2. call `inspect_resume_plan(run_id)`
|
||||
3. if `recommended_action=resume`, continue with:
|
||||
- `start_resume_job`
|
||||
- `get_resume_job_status`
|
||||
- `get_resume_job_result`
|
||||
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
|
||||
5. if `recommended_action=start_new_run`, stop and report the failure
|
||||
|
||||
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
|
||||
|
||||
Project root:
|
||||
|
||||
```bash
|
||||
/home/ubuntu/zhu/github/reader
|
||||
```
|
||||
|
||||
Default behavior for a normal production run:
|
||||
|
||||
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
|
||||
- run with mark-read enabled
|
||||
- do not enable `debug_artifacts`
|
||||
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
||||
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
|
||||
|
||||
Typical artifacts to inspect after a successful run:
|
||||
|
||||
```text
|
||||
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
||||
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json
|
||||
outputs/freshrss/rerun/<run-id>/run-report.json
|
||||
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
|
||||
```
|
||||
|
||||
Notes:
|
||||
|
||||
- `digest-brief.json` is the preferred input for **public digest** generation.
|
||||
- It is a lighter public-only view and currently includes only `keep` candidates.
|
||||
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
|
||||
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
|
||||
|
||||
**⚠️ Output path divergence when run from Feishu context:**
|
||||
When triggered via MCP from inside a Feishu session, the extracted files may be written to an MCP-managed temp path instead of `outputs/freshrss/rerun/<run-id>/extracted/`. Verify the actual output path from the pipeline result before continuing to Phase 5/6. If the project-relative path doesn't exist, use the MCP result's `extracted_path` directly or copy the files to the expected project location.
|
||||
|
||||
### 2. Generate digest markdown
|
||||
|
||||
Generate two output views from the same run, preferably in one model call:
|
||||
|
||||
1. a **public digest** for Hugo / public readers
|
||||
2. an **internal review digest** for chat / operator workflow
|
||||
|
||||
Input preference:
|
||||
|
||||
- **public digest**: prefer `candidates/digest-brief.json`
|
||||
- **internal review digest**: use `candidates/openclaw-delivery-payload.json`
|
||||
|
||||
Recommended generation pattern:
|
||||
|
||||
- Pass the public brief and the full payload as two explicitly labeled input blocks.
|
||||
- Ask the model to return both outputs in one response.
|
||||
- Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed.
|
||||
|
||||
Before publishing, persist the generated digest artifacts back into the same reader run directory:
|
||||
|
||||
```text
|
||||
outputs/freshrss/rerun/<run-id>/digest/public_digest.md
|
||||
outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md
|
||||
outputs/freshrss/rerun/<run-id>/digest/combined.json
|
||||
```
|
||||
|
||||
Then write only the public digest into Hugo here:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
||||
```
|
||||
|
||||
Recommended public-digest front matter:
|
||||
|
||||
```toml
|
||||
+++
|
||||
title = "AI 日报 · YYYY-MM-DD"
|
||||
date = YYYY-MM-DDTHH:MM:SS+08:00
|
||||
summary = "当日日报摘要"
|
||||
+++
|
||||
```
|
||||
|
||||
Recommended **public digest** structure:
|
||||
|
||||
- `今日概览`
|
||||
- `今日重点`
|
||||
- `趋势观察`
|
||||
- `延伸阅读`
|
||||
|
||||
Public digest constraints:
|
||||
|
||||
- Use only the public brief view when present.
|
||||
- Keep public wording free of internal workflow language.
|
||||
- Optimize for concise public readability with solid information density.
|
||||
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
|
||||
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
|
||||
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other variant.**
|
||||
- **All four sections are required. Missing any one is a format violation.**
|
||||
- **Hugo digest must include ALL keep articles from `digest-brief.json`. The IMA deposition subset is a separate downstream step.**
|
||||
|
||||
Recommended **internal review digest** structure:
|
||||
|
||||
- `今日候选概况`
|
||||
- `已入选重点`
|
||||
- `待你确认`
|
||||
- `建议沉淀到 IMA`
|
||||
- `原始候选清单`
|
||||
|
||||
Internal review digest constraints:
|
||||
|
||||
- Do not show `rank` values.
|
||||
- Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`.
|
||||
- For `已入选重点`, include a fuller summary plus a short judgment paragraph.
|
||||
- For `待你确认`, include a fuller summary, reason, and recommendation.
|
||||
- Keep it readable as an operator review draft, not a raw payload dump.
|
||||
- **Feishu compatibility: never use markdown tables in the digest report.**
|
||||
When delivered via Feishu, a single table forces the whole message to plain
|
||||
text. Use lists and sections instead. See `references/feishu-format-notes.md`.
|
||||
|
||||
### 3. Publish Hugo
|
||||
|
||||
Publish only the public digest to Hugo.
|
||||
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
|
||||
|
||||
**⚠️ Hugo redeploy: always use foreground terminal, never background+notify_on_complete.**
|
||||
|
||||
The Hugo redeploy runs via `./redeploy.sh` and completes in ~10 seconds. **Do NOT use `terminal(background=true, notify_on_complete=true)`** for this step — the Gateway will push the raw Docker build output (compiler logs, layered build output, nginx config) to Feishu as a notification. This output is machine-readable, not human-readable, and the user has explicitly said this is noise.
|
||||
|
||||
**Correct approach:**
|
||||
|
||||
```python
|
||||
# ✅ Foreground terminal with adequate timeout
|
||||
cmd = "cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh"
|
||||
result = terminal(cmd, timeout=120)
|
||||
# Then verify 200 OK on the detail page
|
||||
```
|
||||
|
||||
**Wrong approach (DO NOT use):**
|
||||
|
||||
```python
|
||||
# ❌ Background + notify sends raw build logs to Feishu
|
||||
terminal("cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh", background=True, notify_on_complete=True)
|
||||
```
|
||||
|
||||
This rule only applies to Hugo redeploy (fast, deterministic output). For long-running Reader pipeline jobs (>60s), background + notify_on_complete is still appropriate since the output is meaningful content (article summaries, pipeline stats).
|
||||
|
||||
1. **Pre-check: verify Hugo content directory exists**
|
||||
```bash
|
||||
ls -d /home/ubuntu/zhu/apps/hugo-site/content/daily/ 2>/dev/null || mkdir -p /home/ubuntu/zhu/apps/hugo-site/content/daily/
|
||||
```
|
||||
If the entire `content/` directory is missing, create it before proceeding.
|
||||
**Do not assume the Hugo site has a standard structure.**
|
||||
|
||||
2. write the public digest markdown to:
|
||||
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
|
||||
3. redeploy Hugo immediately after writing:
|
||||
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
|
||||
4. verify all three URLs before continuing:
|
||||
- `http://127.0.0.1:14322/`
|
||||
- `http://127.0.0.1:14322/daily/`
|
||||
- `http://127.0.0.1:14322/daily/YYYY-MM-DD/`
|
||||
|
||||
Expected verification targets:
|
||||
|
||||
- homepage works
|
||||
- `/daily/` works
|
||||
- `/daily/YYYY-MM-DD/` works
|
||||
|
||||
### 4. Report digest in chat
|
||||
|
||||
Provide the internal review digest in chat and ask which articles should be retained.
|
||||
|
||||
Hard reporting rule:
|
||||
- do not send only article titles
|
||||
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
|
||||
- **Feishu**: avoid markdown tables entirely. Use lists with headings.
|
||||
|
||||
### 5. Generate selected article summaries
|
||||
|
||||
Only do this after the user explicitly confirms which articles to retain.
|
||||
|
||||
Preferred MCP path:
|
||||
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
|
||||
Expected inputs:
|
||||
|
||||
- `extracted_path`
|
||||
- `selected_ids`
|
||||
**⚠️ item_id 前缀注意:**
|
||||
The `candidates` array IDs use `cand:sha256:xxx` format, but individual extracted item files use `sha256:xxx` (no `cand:` prefix). When passing `selected_ids` to article-summary tools, strip the `cand:` prefix. If the ID doesn't match, the summary tool won't find the article.
|
||||
|
||||
**Path resolution for article-summary:**
|
||||
After a pipeline run via MCP (Feishu context), the extracted files may live at an MCP-managed temp path, not the expected `outputs/freshrss/rerun/<run-id>/extracted/`. Before calling `generate_article_summaries` or `start_article_summary_job`, verify the extracted path exists. If not, use the `run_id` from the pipeline result to locate the actual output directory through `get_run_report`, or copy the files from the MCP-managed path.
|
||||
|
||||
Single-item rule:
|
||||
|
||||
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
|
||||
- in that case, `selected_ids` should contain only the matching single `item_id`
|
||||
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
|
||||
|
||||
Recommended output layout:
|
||||
|
||||
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
||||
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
|
||||
Recommended production sequence:
|
||||
|
||||
1. call `start_article_summary_job`
|
||||
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
|
||||
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
|
||||
|
||||
Hard fallback rule:
|
||||
|
||||
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
|
||||
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
|
||||
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
|
||||
|
||||
Synchronous helper:
|
||||
|
||||
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
|
||||
|
||||
CLI fallback:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
|
||||
--ids <item_id_without_cand_prefix> \
|
||||
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
||||
```
|
||||
|
||||
### 6. Upload selected summaries to IMA
|
||||
|
||||
Upload only the generated markdown files for the selected articles.
|
||||
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
|
||||
|
||||
**⚠️ IMA upload 500 retry: Always wrap IMA uploads in a retry loop.**
|
||||
IMA's OpenAPI may return HTTP 500 on the first attempt. If the first upload fails, wait ~3 seconds and retry. Normally the second attempt succeeds. If `cos-upload.cjs` is used (for CDN-backed uploads), check subprocess stderr even when exit code is 0 — an HTTP error in the upload service may still produce exit code 0.
|
||||
|
||||
Hard execution rules before upload:
|
||||
|
||||
1. reformat/check the generated markdown into IMA-facing final content
|
||||
2. normalize the final upload filename to `<文章标题>.md`
|
||||
3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name
|
||||
4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md`
|
||||
5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs
|
||||
|
||||
Default target knowledge base:
|
||||
|
||||
- `daily`
|
||||
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
|
||||
- Verify the configured target at runtime before upload
|
||||
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
|
||||
|
||||
## Hard rules recap
|
||||
|
||||
- Never generate a digest from placeholder or example data when a real run is expected.
|
||||
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
||||
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
||||
- Only explicitly user-selected articles go to IMA.
|
||||
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
|
||||
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
|
||||
- Selected article summaries use extracted text, not live refetch.
|
||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|
||||
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
|
||||
@@ -0,0 +1,55 @@
|
||||
# IMA COS 上传凭证处理(Hermes redact_secrets 兼容模式)
|
||||
|
||||
## 问题
|
||||
|
||||
Hermes 配置 `security.redact_secrets: true` 时,`create_media` API 返回的 `cos_credential.token` 字段在 `terminal()` 输出中被替换为 `***`。
|
||||
|
||||
直接通过 `terminal()` 调用 `cos-upload.cjs` 会因 token 截断而失败(HTTP 403 InvalidAccessKeyId)。
|
||||
|
||||
## 安全的工作流(execute_code + subprocess.run)
|
||||
|
||||
不要用 `terminal()` 传递 COS 凭证。改用 `execute_code()` + `subprocess.run()` 模式:
|
||||
|
||||
```python
|
||||
# Phase A: terminal() 中保存原始响应到文件
|
||||
result = terminal("""
|
||||
bash -c '
|
||||
set -a
|
||||
source /home/ubuntu/zhu/github/reader/.env
|
||||
set +a
|
||||
OPTS=$(printf "%s" "{\\"clientId\\":\\""$IMA_OPENAPI_CLIENTID"\\",\\\"apiKey\\":\\\""$IMA_OPENAPI_APIKEY"\\\"}")
|
||||
RESP=$(node /root/.hermes/skills/openclaw-imports/ima-skill/ima_api.cjs "openapi/wiki/v1/create_media" "{...}" "$OPTS" 2>/dev/null)
|
||||
echo "$RESP" > /tmp/create_media_raw.json
|
||||
echo "saved"
|
||||
'
|
||||
""")
|
||||
|
||||
# Phase B: execute_code 中从文件读取凭证
|
||||
import json, subprocess
|
||||
with open("/tmp/create_media_raw.json") as f:
|
||||
data = json.load(f)
|
||||
cred = data["data"]["cos_credential"]
|
||||
|
||||
# Phase C: subprocess.run 直接调用,不经过 terminal()
|
||||
args = ["node", f"{SKILL_DIR}/knowledge-base/scripts/cos-upload.cjs",
|
||||
"--file", FILE_PATH, "--secret-id", cred["secret_id"], "--secret-key", cred["secret_key"],
|
||||
"--token", cred["token"], "--bucket", cred["bucket_name"], "--region", cred["region"],
|
||||
"--cos-key", cred["cos_key"], "--content-type", "text/markdown",
|
||||
"--start-time", cred["start_time"], "--expired-time", cred["expired_time"], "--timeout", "300000"]
|
||||
r = subprocess.run(args, capture_output=True, text=True, timeout=310)
|
||||
```
|
||||
|
||||
## 对比:错误的做法(terminal 直接传递)
|
||||
|
||||
```bash
|
||||
# ⛔ 这样不行!token 会被 redact_secrets 替换为 ***
|
||||
TOKEN=$(echo "$RESP" | jq -r '.data.cos_credential.token')
|
||||
node ... --token "$TOKEN" ... # 会收到 HTTP 403
|
||||
```
|
||||
|
||||
## 关键原则
|
||||
|
||||
- `terminal()` 输出中的敏感字段会被自动脱敏,但不影响底层 JSON 文件写入
|
||||
- `execute_code` 中 `terminal()` 返回的 `output` 已经是脱敏后的文本
|
||||
- **唯一可靠的凭证源**是直接写入磁盘的原始 JSON 文件
|
||||
- `subprocess.run` 在 `execute_code` 中绕过脱敏,因为凭证在 Python 内存中直接被传递给子进程,不经过 Hermes 的 stdout 脱敏管道
|
||||
@@ -0,0 +1,46 @@
|
||||
# IMA 凭证链:从 reader `.env` 到 IMA 上传
|
||||
|
||||
## 凭证来源
|
||||
|
||||
IMA 上传所需的凭证存储在多个位置,优先级如下:
|
||||
|
||||
| 优先级 | 位置 | 说明 |
|
||||
|--------|------|------|
|
||||
| 1 | 环境变量 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` | Hermes session 级 |
|
||||
| 2 | `~/.config/ima/client_id` / `api_key` | ima-skill 的默认检查路径 |
|
||||
| 3 | `/home/ubuntu/zhu/github/reader/.env` | reader 项目配置,含完整的 IMA 凭证和 KB ID |
|
||||
|
||||
## 凭证内容(reader .env 中)
|
||||
|
||||
```
|
||||
IMA_OPENAPI_CLIENTID=<32位hex>
|
||||
IMA_OPENAPI_APIKEY=<base64编码的API密钥>
|
||||
IMA_DAILY_KNOWLEDGE_BASE_ID=<base64编码的KB ID>
|
||||
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
|
||||
```
|
||||
|
||||
## 缺失时的处理流程
|
||||
|
||||
当 IMA 上传失败(`-100` 凭证缺失错误)时:
|
||||
|
||||
1. 从 reader `.env` 读取凭证:
|
||||
```
|
||||
grep -E '^(IMA_OPENAPI_CLIENTID|IMA_OPENAPI_APIKEY)=' /home/ubuntu/zhu/github/reader/.env
|
||||
```
|
||||
2. 同步到 ima-skill 默认检查路径:
|
||||
```
|
||||
echo "<client_id>" > ~/.config/ima/client_id
|
||||
echo "<api_key>" > ~/.config/ima/api_key
|
||||
```
|
||||
3. (可选)追加到 Hermes `.env` 以全局生效:
|
||||
```
|
||||
echo "IMA_OPENAPI_CLIENTID=<client_id>" >> /root/.hermes/.env
|
||||
echo "IMA_OPENAPI_APIKEY=<api_key>" >> /root/.hermes/.env
|
||||
echo "IMA_DAILY_KNOWLEDGE_BASE_ID=<kb_id>" >> /root/.hermes/.env
|
||||
echo "IMA_DAILY_KNOWLEDGE_BASE_NAME=daily" >> /root/.hermes/.env
|
||||
```
|
||||
|
||||
## 执行注意事项
|
||||
|
||||
- **COS 凭证红线**:`create_media` 返回的 `cos_credential` 中的 `token`/`secret_id`/`secret_key` 在 `terminal()` 输出中会被 Hermes 替换为 `***`。必须用 `subprocess.run()` 捕获原始输出,或用 `execute_code` 内联操作。
|
||||
- **凭证格式**:`api_key` 是 base64 字符串(76 字符),`kb_id` 也是 base64 字符串。不要截断或转码。
|
||||
@@ -0,0 +1,97 @@
|
||||
# IMA Markdown 文件上传流程(media_type=7)
|
||||
|
||||
## 什么时候用此流程
|
||||
|
||||
当用户说"沉淀到 IMA"时,必须用此文件上传流程,而不是 URL 导入或笔记导入。
|
||||
|
||||
## 完整流程
|
||||
|
||||
### 1. 写 .md 文件
|
||||
|
||||
```bash
|
||||
mkdir -p /tmp/ima_upload
|
||||
cat > /tmp/ima_upload/文章标题.md << 'EOF'
|
||||
# 文章标题
|
||||
|
||||
> 来源:XXX
|
||||
|
||||
## 摘要
|
||||
|
||||
...
|
||||
|
||||
## 核心亮点
|
||||
|
||||
- ...
|
||||
|
||||
[原文链接](url)
|
||||
EOF
|
||||
```
|
||||
|
||||
### 2. preflight 检查
|
||||
|
||||
```python
|
||||
pf = subprocess.run(
|
||||
["node", f"{SKILL_DIR}/knowledge-base/scripts/preflight-check.cjs",
|
||||
"--file", filepath],
|
||||
capture_output=True, text=True
|
||||
)
|
||||
meta = json.loads(pf.stdout)
|
||||
# meta = {pass, file_name, file_ext, file_size, media_type, content_type}
|
||||
```
|
||||
|
||||
### 3. create_media(获取 COS 凭证)
|
||||
|
||||
```python
|
||||
r = ima_api("openapi/wiki/v1/create_media", {
|
||||
"file_name": fname,
|
||||
"file_size": meta["file_size"],
|
||||
"content_type": meta["content_type"],
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_ext": meta["file_ext"]
|
||||
})
|
||||
cos = r["data"]["cos_credential"]
|
||||
media_id = r["data"]["media_id"]
|
||||
```
|
||||
|
||||
### 4. COS 上传
|
||||
|
||||
```python
|
||||
cu = subprocess.run([
|
||||
"node", f"{SKILL_DIR}/knowledge-base/scripts/cos-upload.cjs",
|
||||
"--file", filepath,
|
||||
"--secret-id", cos["secret_id"],
|
||||
"--secret-key", cos["secret_key"],
|
||||
"--token", cos["token"],
|
||||
"--bucket", cos["bucket_name"],
|
||||
"--region", cos["region"],
|
||||
"--cos-key", cos["cos_key"],
|
||||
"--content-type", meta["content_type"],
|
||||
"--start-time", str(cos["start_time"]),
|
||||
"--expired-time", str(cos["expired_time"]),
|
||||
"--timeout", "300000"
|
||||
], capture_output=True, text=True, timeout=30)
|
||||
# 非0退出 = 上传失败
|
||||
```
|
||||
|
||||
### 5. add_knowledge
|
||||
|
||||
```python
|
||||
r = ima_api("openapi/wiki/v1/add_knowledge", {
|
||||
"media_type": 7,
|
||||
"media_id": media_id,
|
||||
"title": title, # 文章中文标题
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_info": {
|
||||
"cos_key": cos["cos_key"],
|
||||
"file_size": meta["file_size"],
|
||||
"file_name": fname
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
## 注意事项
|
||||
|
||||
- **COS 凭证不能通过 terminal() 读取**(Hermes redact_secrets 会把 token 替换成 `***`)。必须用 Python `subprocess.run()` 捕获原始 stdout。
|
||||
- **create_media 的 content_type 不能带 charset 参数**(如 `text/markdown; charset=utf-8` 会被拒)。用纯 MIME 类型 `text/markdown`。
|
||||
- **文件命名 = `<文章中文标题>.md`**。不要用英文 slug 或 temp name。
|
||||
- **media_type=7** 是 Markdown 文件。**绝对不要用 media_type=11**(笔记)或 `import_urls`。
|
||||
@@ -0,0 +1,42 @@
|
||||
# IMA 上传格式速查表(日报沉淀专用)
|
||||
|
||||
## 文件名 vs 标题 vs 内容对照表
|
||||
|
||||
| 项目 | ✅ 正确 | ❌ 错误 |
|
||||
|------|---------|---------|
|
||||
| **文件名** | `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` | `大仓AI工程化实战.md` |
|
||||
| **add_knowledge title** | `从Vibe Coding到Harness—— 一套大仓AI工程化实战` | `大仓AI工程化实战` / `大仓AI工程化实战.md` |
|
||||
| **上传方式** | `preflight` → `create_media` → `cos-upload` → `add_knowledge(media_type=7)` | `import_doc` + `add_knowledge(media_type=11)` |
|
||||
| **正文内容** | pipeline `extracted/` 的 `article.content` 或 `summary-batch.json` 的 `summary` | 自己写的两三句话 |
|
||||
| **正文长度** | ≥ 500 字 | < 500 字 |
|
||||
|
||||
## 文件名常见错误模式
|
||||
|
||||
| 原始标题 | ❌ 错误文件名 | ✅ 正确文件名 |
|
||||
|----------|-------------|-------------|
|
||||
| `从Vibe Coding到Harness—— 一套大仓AI工程化实战` | `大仓AI工程化实战.md` | `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` |
|
||||
| `契约化多端架构:基于领域模型的Harness实践` | `契约化多端架构Harness实践.md` | `契约化多端架构:基于领域模型的Harness实践.md` |
|
||||
| `Loop Engineering 实战:从日志扫描到预发部署的全自主闭环` | `LoopEngineering实战.md` | `Loop Engineering 实战:从日志扫描到预发部署的全自主闭环.md` |
|
||||
|
||||
## 5 Section 格式要求
|
||||
|
||||
| Section | 格式 | 说明 |
|
||||
|---------|------|------|
|
||||
| **核心结论** | 一段连贯段落(非分点) | 总结文章核心发现或主张 |
|
||||
| **主要论点** | 一段连贯段落(非分点) | 综合多个论点形成连贯叙述 |
|
||||
| **关键方法 / 机制** | `- **方法名**:详细说明` 分点 | 每条展开到能理解原理的程度 |
|
||||
| **重要细节** | `- 每条一个带解释的完整知识点` 分点 | 每条是一个完整知识点,不是关键词 |
|
||||
| **可复用启发** | `- 每条 actionable 的实践启示` 分点 | 附应用场景说明 |
|
||||
|
||||
## 底部额外 section
|
||||
|
||||
- `## 关键词` — 标签式关键词列表
|
||||
- `## 主题` — 主题分类
|
||||
|
||||
## 速查口诀
|
||||
|
||||
> 文件名 = 完整标题.md
|
||||
> 内容取 pipeline,不自己写
|
||||
> 5 个 section,核心结论和主要论点是段落
|
||||
> media_type = 7,不是 11
|
||||
> title 不加 .md
|
||||
@@ -0,0 +1,38 @@
|
||||
# AI Agent 的 Skill 系统设计
|
||||
|
||||
Source: https://mp.weixin.qq.com/s?__biz=MzAxNDEwNjk5OQ==&mid=2650544717&idx=1&sn=b578abf5a81034670900a3b8eb874296
|
||||
Category: 方法论
|
||||
|
||||
## 核心结论
|
||||
好的 Skill 应是一个小而准的行为系统,通过触发、加载、执行、约束、验证和迭代的组织,将通用 Agent 转化为在特定任务上稳定可靠的专用 Agent。核心原则是上下文窗口是公共资源,必须采用渐进披露、按任务风险设置自由度,并通过真实任务前向测试来证明行为改变。
|
||||
|
||||
## 主要论点
|
||||
Skill 设计的本质是行为编程而非文档编写,需要将期望行为转化为 Agent 能稳定执行的结构化工作流。为此,必须同时解决发现(正确场景触发)、加载(最小上下文)、执行(合适自由度)和验证(真实任务测试)四件事,并通过门控、脚本外化、测试防合理化等机制确保 Agent 在复杂压力下不走捷径。
|
||||
|
||||
## 关键方法 / 机制
|
||||
- 渐进披露的三层内容加载:元数据(frontmatter 的 name/description)用于发现,正文(SKILL.md)用于执行,资源(scripts/references/assets)按需读取。description 只做路由器,不包含完整流程,防止 Agent 凭印象执行。
|
||||
- 门控机制(HARD-GATE):在低自由度任务中,使用明确的 <HARD-GATE> 标签禁止 Agent 在条件满足前执行后续动作,减少解释空间,让关键路径更像程序而非建议。
|
||||
- 脚本外化降低上下文消耗和行为漂移:将需要确定性的操作(如 PDF 旋转)封装为 scripts/ 中的可执行脚本,避免 Agent 每次临时生成代码,提升可靠性和节省 token。
|
||||
- 基于 TDD 的前向测试方法:用子代理模拟真实用户任务,只给原始任务和最少上下文,不泄露预期结论;观察行为轨迹、输出文件等原始证据,发现并封堵 Agent 的违规行为。
|
||||
- 反模式自查表与检查表:交付前检查触发条件、自由度设置、资源引用、验证流程等关键项,确保 Skill 不是一份草稿而是一个可用的能力包。
|
||||
- 跨平台适配与优雅降级:Skill 应写行为规则(如 TodoWrite),再通过平台层映射到具体工具名(如 todowrite);平台能力不足时优雅降级,保持 Skill 的可迁移性。
|
||||
|
||||
## 重要细节
|
||||
- SKILL.md 的 frontmatter 和正文职责分离:name/description 用于发现(Agent 触发前可见),正文用于执行(触发后加载)。如果触发条件写在正文里,Agent 在决定是否触发时根本读不到。
|
||||
- 命名规范:短、可触发、动词优先,例如 create-skill 比 skill-creation 更好,这本质上是路由质量——Agent 在技能库里找能力时,name/description 是第一层索引。
|
||||
- 资源组织原则——“信息只放一个地方”:不要在 SKILL.md 和 references/ 中重复同一段规则,重复会带来漂移,导致 Agent 在两个版本间自行解释,增加维护成本。
|
||||
- 门控类型示例:先决条件门控(先理解例子再编辑)、并发冲突门控、未保存内容门控、敏感操作门控(已创建 Skill 需处理影响再修改)。
|
||||
- 流程图用 GraphViz DOT 嵌入 Markdown:对于包含非线性判断、循环、回退的步骤,流程图比纯文本更稳定,能防止 Agent 遗漏关键分支。
|
||||
- 验证时防“合理化”问题:AI Agent 在压力下会为跳过规则编造理由,Skill 需要提前写出这些借口并给出反驳;审查循环应围绕真实失败风险而非措辞偏好。
|
||||
|
||||
## 可复用启发
|
||||
- “上下文窗口是公共资源”原则:设计任何 Agent 指令时,每段内容都要质疑“Agent 真的需要这段解释吗?”和“值得占用的 token 成本吗?”,这适用于提示词、系统消息等所有 Agent 输入设计。
|
||||
- 先收集具体例子再抽象 Skill:不要从抽象能力开始写,而是先收集用户会怎么触发、哪些请求应该触发/不应该触发、成功输出是什么等具体场景,避免写出宽泛不可执行的指令。
|
||||
- 用脚本固化确定性任务、用门控防止关键路径走捷径:对于高脆弱、低变化空间的任务(如文件格式转换),应使用脚本而非描述性建议;对于必须按顺序执行的步骤,用门控打断 Agent 的“合理化”冲动。
|
||||
- 设计防合理化的测试流程:用子代理模拟真实用户,只给原始任务,不泄露预期答案;观察是否存在只有看到结论才能成功的情况——如果这样,说明 Skill 不够清楚或测试设置泄露答案。
|
||||
|
||||
## 关键词
|
||||
SKILL.md、YAML、Markdown、DOT、GraphViz、TDD、HARD-GATE、quick_validate.py
|
||||
|
||||
## 主题
|
||||
AI Agent、行为编程、Token 经济、系统设计、约束机制
|
||||
@@ -0,0 +1,33 @@
|
||||
# IMA 笔记格式参考
|
||||
|
||||
> ⚠️ 格式基准文件为 `references/ima-format-reference.md`(`ai-agent-的-skill-系统设计.md`),每次生成 IMA 沉淀前必须先读该文件。
|
||||
|
||||
## 标准结构
|
||||
|
||||
### 顶部元数据
|
||||
|
||||
```markdown
|
||||
# 文章完整标题
|
||||
|
||||
Source: https://原文链接(纯 URL,不加 "原文链接:" 标签)
|
||||
Category: 分类
|
||||
```
|
||||
|
||||
### 5 个必含 Section
|
||||
|
||||
1. **核心结论** — 一段总结文章核心发现或主张的段落(不是分点),像基准文件一样是一段连贯文字
|
||||
2. **主要论点** — 一段概括文章核心论述的段落(不是分点列表),综合多个论点形成连贯叙述
|
||||
3. **关键方法 / 机制** — `**方法名**:详细说明` 的格式,每条展开到能理解其原理的程度
|
||||
4. **重要细节** — 每条一个带解释的完整知识点,不是关键词
|
||||
5. **可复用启发** — 每条 actionable 的实践启示,附应用场景说明
|
||||
|
||||
### 底部额外 Section
|
||||
|
||||
- **## 关键词** — 标签式关键词列表
|
||||
- **## 主题** — 主题分类
|
||||
|
||||
### 质量要求
|
||||
|
||||
- 每篇文章总内容量应达到 **4,000+ bytes**(基准文件 4,782B)
|
||||
- 每个 section 展开到完整的知识点级别,不能只列关键词
|
||||
- 内容来源:优先 `item-XX.extracted.json` 的 `article.content`,其次 `summary-batch.json` 的 `summary` 字段
|
||||
@@ -0,0 +1,80 @@
|
||||
# IMA Upload API Reference
|
||||
|
||||
Full API flow for uploading markdown articles to the IMA `daily` knowledge base. Used in Phase 7 of the reader-digest-flow.
|
||||
|
||||
## Credentials
|
||||
|
||||
```
|
||||
IMA_OPENAPI_CLIENTID - from reader .env or user-provided
|
||||
IMA_OPENAPI_APIKEY - from reader .env or user-provided
|
||||
IMA_DAILY_KNOWLEDGE_BASE_ID - daily KB UUID
|
||||
```
|
||||
|
||||
The `ima-skill` v1.1.7+ ships with `ima_api.cjs` for credential loading.
|
||||
Legacy auth header: `ima-openapi-ctx: skill_version=1.1.7`.
|
||||
|
||||
Do NOT export the full API key in shell commands — use `execute_code` with `subprocess.run` and Python string variables.
|
||||
|
||||
## Flow (3 steps)
|
||||
|
||||
### 1. create_media
|
||||
|
||||
```
|
||||
POST https://ima.qq.com/openapi/wiki/v1/create_media
|
||||
Headers: ima-openapi-clientid, ima-openapi-apikey, Content-Type: application/json
|
||||
Body: { file_name, file_size, content_type, knowledge_base_id, file_ext }
|
||||
Returns: { code: 0, data: { media_id, cos_credential: { secret_id, secret_key, token, bucket_name, region, cos_key, start_time, expired_time } } }
|
||||
```
|
||||
|
||||
`file_ext` is without the dot (e.g. `md` not `.md`).
|
||||
`file_name` must be the user-facing article title + `.md`.
|
||||
`content_type` for markdown is `text/markdown`; media_type=7.
|
||||
|
||||
### 2. COS upload
|
||||
|
||||
Use `cos-upload.cjs` from `ima-skill/knowledge-base/scripts/`:
|
||||
|
||||
```
|
||||
node <skill_dir>/knowledge-base/scripts/cos-upload.cjs \
|
||||
--file <local_md_file> \
|
||||
--secret-id <from create_media> \
|
||||
--secret-key <from create_media> \
|
||||
--token <from create_media> \
|
||||
--bucket <bucket_name> \
|
||||
--region <region> \
|
||||
--cos-key <cos_key> \
|
||||
--content-type text/markdown \
|
||||
--start-time <start_time> \
|
||||
--expired-time <expired_time>
|
||||
```
|
||||
|
||||
⚠️ Must use Python `subprocess.run(args=[...])` to avoid shell parameter mangling.
|
||||
⚠️ Always capture `returncode` and `stderr` — COS may return exit 0 on HTTP 500.
|
||||
|
||||
### 3. add_knowledge
|
||||
|
||||
```
|
||||
POST https://ima.qq.com/openapi/wiki/v1/add_knowledge
|
||||
Headers: same as create_media
|
||||
Body: { media_type: 7, media_id, title: "<file_name>", knowledge_base_id, file_info: { cos_key, file_size, file_name } }
|
||||
```
|
||||
|
||||
`media_type=7` for markdown. `title` MUST equal `file_name`.
|
||||
|
||||
## Article Markdown reformatting (before upload)
|
||||
|
||||
Generated summaries from `reader` have `Source:` and `Category:` header lines.
|
||||
Before uploading, reformat to IMA style:
|
||||
|
||||
```
|
||||
原文链接:<original article URL>
|
||||
|
||||
## 核心结论
|
||||
...
|
||||
|
||||
## 主要论点
|
||||
...
|
||||
```
|
||||
|
||||
Remove `Source:`, `Category:` lines. Keep `原文链接:` at top with the URL on the next line.
|
||||
Break long prose (>200 chars per paragraph) into shorter paragraphs for IMA readability.
|
||||
@@ -0,0 +1,195 @@
|
||||
# 关键词引擎维护流程
|
||||
|
||||
## 概述
|
||||
|
||||
reader 项目内置了完整的关键词清洗链路,用于管理 `term_aliases.json`、`term_stopwords.json`、`filter_context.personal.json` 等配置。本文件说明链路各环节的职责与使用方式。
|
||||
|
||||
## 完整数据流
|
||||
|
||||
```
|
||||
每日日报 pipeline
|
||||
│
|
||||
▼
|
||||
term_index/daily/YYYY-MM-DD.json ← build_keyword_index.py (pipeline 末环节自动)
|
||||
│
|
||||
▼
|
||||
term_index/term_stats.json ← 全量汇总
|
||||
│
|
||||
▼
|
||||
build_review_bundle.py ← 打包审查数据包 (手动触发)
|
||||
│
|
||||
▼
|
||||
generate_term_cleanup_suggestions.py ← 生成建议 (手动触发)
|
||||
│
|
||||
▼
|
||||
apply_term_suggestions.py ← 人工确认后落盘 (手动触发)
|
||||
```
|
||||
|
||||
## 各环节命令
|
||||
|
||||
### 1. 重建每日 term_index(pipeline 末环节自动执行,也可手动)
|
||||
|
||||
```bash
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 scripts/build_keyword_index.py \
|
||||
--input outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
||||
```
|
||||
|
||||
### 2. 重建 review bundle(打包当前配置+统计供审查)
|
||||
|
||||
```bash
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 \
|
||||
--top 200 \
|
||||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
参数:
|
||||
- `--days 365` — 考虑最近多少天的统计数据
|
||||
- `--top 200` — 取前 N 个高频词纳入 bundle
|
||||
|
||||
### 3. 生成清洗建议
|
||||
|
||||
```bash
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--emit-markdown
|
||||
```
|
||||
|
||||
输出:
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json` — 正式建议产物
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md` — 人工阅读展示稿
|
||||
|
||||
当前脚本能生成的建议类型:
|
||||
|
||||
| 类型 | 生成规则 | 当前产出 |
|
||||
|:----|:---------|:---------|
|
||||
| interest_keyword_suggestions | 高频未覆盖词(total≥3, days≥2) | ✅ 按阈值产出 |
|
||||
| watch_terms | 中频观察词(total≤2, days≤2) | ✅ 按阈值产出 |
|
||||
| alias_suggestions | 大小写变体/单复数/空格连词符差异 | ✅ 统计规则产出 |
|
||||
| stopword_suggestions | — | ❌ 脚本硬编码为空 |
|
||||
|
||||
### 4. 应用建议(dry-run → review → apply)
|
||||
|
||||
```bash
|
||||
# 先 dry-run 预览
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 scripts/apply_term_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--accept-interest 强化学习 ReAct CLI \
|
||||
--dry-run
|
||||
|
||||
# 确认后正式 apply(去掉 --dry-run)
|
||||
.venv/bin/python3 scripts/apply_term_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--accept-interest 强化学习 ReAct CLI \
|
||||
```
|
||||
|
||||
支持的 accept 参数:
|
||||
- `--accept-interest 词1 词2 ...` — 添加到 interest_keywords
|
||||
- `--accept-watch 词1 词2 ...` — 添加到 watchlist
|
||||
- `--accept-alias 源词1 源词2 ...` — 添加 alias 映射
|
||||
- `--accept-stopword 词1 词2 ...` — 添加停用词
|
||||
|
||||
## 脚本能做什么 vs 不能做什么
|
||||
|
||||
### ✅ 脚本能做的(统计级清洗)
|
||||
|
||||
- 发现大小写变体(`vibe coding` → `Vibe Coding`)
|
||||
- 发现单复数差异(`Agent Skill` → `Agent Skills`)
|
||||
- 发现空格/连词符差异
|
||||
- 按频次推荐 should-be-interest / should-be-watch 的词
|
||||
- 批量 apply 到配置文件,自动记录变更日志
|
||||
|
||||
### ❌ 脚本不能做的(语义级清洗,需要人工判断)
|
||||
|
||||
- 识别产品名(`WorkBuddy`、`LibTV Agent`、`飞书妙搭`)→ 应加 stopword
|
||||
- 识别模型名(`Qwen3-30B-A3B`、`GLM 5.2`、`Opus 4.8`)→ 应加 stopword
|
||||
- 识别人名/地名(`Andrej Karpathy`、`成都天府长岛`)→ 应加 stopword
|
||||
- 英文专有名词归一化到中文主题词(`Multi-Agent` → `多Agent`)
|
||||
- 长产品名归一化(`火山云数据库PostgreSQL Serverless版` → `Serverless数据库`)
|
||||
- 判断某个词对 Hugo tag 是否有用(`1688`、`OPC训练营` → 无效)
|
||||
|
||||
## 语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py)
|
||||
|
||||
除了统计规则的清洗,项目还有一个 **LLM 驱动的语义级清洗脚本**,能发现统计规则做不到的事情:
|
||||
|
||||
```bash
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||||
```
|
||||
|
||||
参数:
|
||||
- `--bundle` — review bundle(先跑 build_review_bundle.py)
|
||||
- `--suggestions` — 规则层的建议 JSON(可选,供 LLM 避免重复)
|
||||
- `--output` — 输出路径(自动生成)
|
||||
- `--dry-run` — 只打印 prompt 不调 LLM
|
||||
|
||||
### 语义脚本能做的(而统计规则不能做的)
|
||||
|
||||
| 类型 | LLM 能发现什么 | 示例 |
|
||||
|:----|:--------------|:-----|
|
||||
| **semantic_alias** | 中英文映射、简称↔全称、近义词 | `AI智能体`→`AI Agent`, `AI Harness`→`Harness Engineering` |
|
||||
| **stopword** | 语义判断哪些词太泛/不相干 | `自动化`(太泛), `AlphaFold`(生物领域), `陌生化`(穿透词) |
|
||||
| **promote_to_interest** | 与用户关注领域对齐的新词 | `强化学习`, `Loop Engineering`, `MoE`, `ChatGPT` |
|
||||
|
||||
### 两阶段清洗 SOP
|
||||
|
||||
当需要清洗关键词时,按此顺序操作:
|
||||
|
||||
**阶段一:规则统计清洗(脚本发现 + 人工确认)**
|
||||
1. `build_review_bundle.py` → 重建 bundle
|
||||
2. `generate_term_cleanup_suggestions.py` → 产出统计级建议
|
||||
3. 检查建议,决定哪些 accept
|
||||
4. `apply_term_suggestions.py --dry-run` → 预览
|
||||
5. `apply_term_suggestions.py` → 正式落地
|
||||
|
||||
**阶段二:语义级清洗(LLM 发现 + 人工确认)**
|
||||
1. `generate_term_cleanup_semantic_suggestions.py` → 调 LLM 产生语义建议
|
||||
2. 检查 LLM 的 semantic_alias / stopword / promote_to_interest
|
||||
3. 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化)
|
||||
|
||||
**🔴 不要绕过确认流程直接改配置。** 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。
|
||||
|
||||
## 完整的语义级清洗操作流程
|
||||
|
||||
当需要大量添加 aliases/stopwords 时,推荐流程:
|
||||
|
||||
1. **跑一次 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py`** 看看统计建议
|
||||
2. **走一遍实际 pipeline 产出**,收集所有 unique keywords:
|
||||
```bash
|
||||
for rid in $(ls outputs/freshrss/rerun/); do
|
||||
python3 -c "
|
||||
import json
|
||||
d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json'))
|
||||
for item in d['items']:
|
||||
for kw in item['summary'].get('keywords', []):
|
||||
print(kw)
|
||||
"
|
||||
done | sort -u
|
||||
```
|
||||
3. 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的
|
||||
4. 分别更新 `term_aliases.json` 和 `term_stopwords.json`
|
||||
5. **重建 term_index**:对每个受影响日期的 run 跑 `build_keyword_index.py`
|
||||
6. **更新 Hugo tags**:从重建后的 `data/term_index/daily/YYYY-MM-DD.json` 取
|
||||
|
||||
## 当前配置 (2026-07-16)
|
||||
|
||||
| 文件 | 条目数 |
|
||||
|:----|:------|
|
||||
| `configs/term_aliases.json` | 142 |
|
||||
| `configs/term_stopwords.json` | 106 |
|
||||
| `configs/filter_context.personal.json` | 54 (interest_keywords) |
|
||||
| `configs/term_watchlist.json` | 6 |
|
||||
|
||||
## 关键词/别名/停用词配置的更新规范
|
||||
|
||||
- `term_aliases.json` — 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式
|
||||
- `term_stopwords.json` — 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词
|
||||
- **随时可以加**,加了后重建 term_index 即可生效
|
||||
- 不涉及 pipeline 重新跑——只影响下游展示
|
||||
@@ -0,0 +1,60 @@
|
||||
# Memory Drift Recovery
|
||||
|
||||
## Symptom
|
||||
|
||||
`store_memory(action="add", ...)` fails with:
|
||||
> Refusing to write MEMORY.md: file on disk has content that wouldn't round-trip through the memory tool...
|
||||
|
||||
A `.bak` snapshot is created: `/root/.hermes/memories/MEMORY.md.bak.<timestamp>`
|
||||
|
||||
## Root Cause
|
||||
|
||||
The MEMORY.md file format doesn't match what the memory tool expects — likely because the file was modified externally (by `patch`, `write_file`, shell `>>` append, or a concurrent session). The tool uses a `§` (section sign) delimited format internally and its serialization/deserialization doesn't match the on-disk content.
|
||||
|
||||
## Recovery Procedure
|
||||
|
||||
### Step 1: Read the backup and the current file
|
||||
|
||||
```bash
|
||||
diff /root/.hermes/memories/MEMORY.md.bak.<timestamp> /root/.hermes/memories/MEMORY.md
|
||||
```
|
||||
|
||||
### Step 2: Extract missing entries (if any)
|
||||
|
||||
```bash
|
||||
# List entries from the backup
|
||||
grep '^§' /root/.hermes/memories/MEMORY.md.bak.<timestamp>
|
||||
```
|
||||
|
||||
### Step 3: Re-add each missing entry via store_memory
|
||||
|
||||
For each entry that was in the backup but is now gone from the current file:
|
||||
```bash
|
||||
store_memory(action="add", content="<entry text>", target="memory")
|
||||
```
|
||||
|
||||
### Step 4: Reset to a clean state
|
||||
|
||||
If the file is completely corrupted, the cleanest path is:
|
||||
1. Save any new entries from the backup you want to keep
|
||||
2. Rewrite the file as a clean `§`-delimited list (one entry per `§` line)
|
||||
3. The format is: `§<content>\n` per entry, with `---` or blank line separators
|
||||
|
||||
```bash
|
||||
# Example clean format:
|
||||
echo '§当前重要条目一
|
||||
§当前重要条目二
|
||||
§当前重要条目三' > /root/.hermes/memories/MEMORY.md
|
||||
```
|
||||
|
||||
### Prevention
|
||||
|
||||
- Do NOT use `write_file` or `patch` to modify MEMORY.md directly — always use `store_memory()`
|
||||
- Do NOT use shell `>>` to append to MEMORY.md
|
||||
- If you must bulk-import, use `store_memory` per-entry, not file-level operations
|
||||
|
||||
## Environment
|
||||
|
||||
- Host: Linux (5.15)
|
||||
- Hermes home: `/root/.hermes`
|
||||
- Memory files: `~/.hermes/memories/MEMORY.md`, `~/.hermes/memories/USER.md`
|
||||
@@ -0,0 +1,41 @@
|
||||
+++
|
||||
title = "AI 日报 · 示例"
|
||||
date = 2026-04-01T16:55:00+08:00
|
||||
summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实践的当日观察。"
|
||||
+++
|
||||
|
||||
> ⚠️ 格式规范(生成 Hugo 时必须遵守):
|
||||
> - 文章编号:`1.` `2.` `3.`(阿拉伯数字 + 点),禁止 `① ② ③` / `一、二、三` 等变体
|
||||
> - 四个 section 缺一不可:`今日概览` → `今日重点` → `趋势观察` → `延伸阅读`
|
||||
> - 每篇文章结构:标题 → 摘要段 → "值得关注:"三点 → "这篇更值得关注的理由"段
|
||||
|
||||
# 今日概览
|
||||
|
||||
今天的公开候选主要集中在 AI Agent 的架构演进、工具化落地与工程化实践三条线索上。相比早期偏概念展示的讨论,这一批内容更强调模块化能力栈、真实部署路径与系统可维护性,说明行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
||||
|
||||
## 今日重点
|
||||
|
||||
### 1. 学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变
|
||||
|
||||
文章分析了 AI 智能体架构从单体 Agent 向模块化 Skills 的范式转变。Anthropic 先后推出 MCP 和 Agent Skills 开放标准,构建了知识、工具、协作和运行分层架构。文章通过一个自动化美化相册的真实项目,对比了 Claude Code 与 OpenClaw 两种实现方案,验证了新架构的可复用性与灵活性。
|
||||
|
||||
值得关注:
|
||||
- Anthropic 在 14 个月内先后推出 MCP 和 Agent Skills 两个开放标准,推动 AI 智能体架构分层化。
|
||||
- 新范式核心是构建薄 Agent 引擎与可组合的 Skills 库,取代为每个用例定制单体 Agent。
|
||||
- 文章通过自动化美化相册项目,实操演示了 Skills、MCP、OpenClaw 和 A2A 协议如何协同工作。
|
||||
|
||||
这篇内容更值得关注的原因在于,它不只是提出了“Agent 要模块化”这个判断,而是把开放标准、分层架构和真实项目案例串成了一条完整论证链,能直接支撑今天日报的主线。
|
||||
|
||||
## 趋势观察
|
||||
|
||||
1. Agent 正在从单体能力转向可组合的模块化体系。无论是 Skills、MCP、记忆还是运行时编排,这批内容都在强调解耦与复用,而不是把智能体继续当成一个不可拆分的黑箱。
|
||||
2. 工程化正在变成 AI 应用竞争的主战场。桌面 Agent、企业级架构和部署实践类内容增多,说明真正的差异化开始落在接入现有流程、控制风险和提升可维护性上。
|
||||
3. AI 能力的竞争点正在上移。模型本身仍重要,但真正可持续的优势越来越来自系统设计、工作流整合和对业务场景的理解。
|
||||
|
||||
## 延伸阅读
|
||||
|
||||
- [学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变](https://example.com/a)|阿里云开发者
|
||||
- [Agent Skills:打通可复用专业领域知识的最后一公里](https://example.com/b)|阿里云开发者
|
||||
- [CoPaw深度解析:源码架构和功能实践](https://example.com/c)|阿里云开发者
|
||||
|
||||
> ⚠️ 延伸阅读必须包含当天所有入选文章的原文链接(`- [标题](url)|来源`),一条对应一篇今日重点文章。不要放未入选或 pipeline drop 的文章。这是日报读者获取原文的入口,不是"其他相关阅读"区。
|
||||
@@ -0,0 +1,64 @@
|
||||
# Tag 去重合并计划 (2026-07-22)
|
||||
|
||||
## 问题
|
||||
https://osiman.site/tags/ 页面存在大量重叠 tag,如 `Harness工程化` / `Harness工程` / `Harness Engineering` 三个 tag 指向同一概念。
|
||||
|
||||
## 方案:数据层合并(Option A)
|
||||
|
||||
不修改展示层,直接合并所有日报 md 文件中的 tags。Hugo 自动重建 tag 页面,旧 tag 自动废弃。
|
||||
|
||||
## 规范映射表
|
||||
|
||||
| 废弃 tag | → | 规范名 | 说明 |
|
||||
|:---------|:---|:-------|:----|
|
||||
| `Harness工程化`, `Harness Engineering` | → | `Harness工程` | 三合一 |
|
||||
| `Skills` | → | `Skill` | 单复数 |
|
||||
| `AI Coding`, `AI代码生成`, `AI工程化` | → | `AI Coding Agent` | 统一为 Agent 维度 |
|
||||
| `Agent 框架`, `Agentic架构`, `Agent工程` | → | `Agent工程` | 三合一 |
|
||||
| `Prompt` | → | `Prompt Engineering` | 从简写改全称 |
|
||||
| `多Agent`, `Multi-Agent架构`, `多Agent协作` | → | `多Agent` | 三合一 |
|
||||
| `LLM`, `LLM评估`, `LLM训练`, `大模型应用开发` | → | `LLM` | 四合一 |
|
||||
| `循环工程`, `Loop Engineering`, `Agent Loop` | → | `循环工程` | 三合一 |
|
||||
| `上下文管理`, `Context工程` | → | `上下文管理` | 统一中文 |
|
||||
| `Code Review`, `代码质量` | → | `Code Review` | 统一英文 |
|
||||
| `推理加速`, `长文本推理`, `多步推理` | → | `推理加速` | 三合一 |
|
||||
| `安全`, `安全防御` | → | `安全` | 二合一 |
|
||||
| `技能系统`, `知识工程`, `知识管理` | → | `知识管理` | 三合一 |
|
||||
| `Agent`, `AI` | → | (删除) | 太泛,无信息量 |
|
||||
| `工程化` | → | (删除) | 冗余 |
|
||||
| `架构` | → | (删除) | 冗余 |
|
||||
|
||||
## 保留的独立 tag
|
||||
|
||||
`MCP`, `RAG`, `ACP`, `KV Cache`, `MoE架构`, `State Lake`, `RL决策训练`, `Token管理`, `Vibe Coding`, `Spec工程`, `Skill流水线`, `Hook 链`, `全双工语音交互`, `多端架构`, `契约化架构`, `数据`, `工具链`, `工作流`, `评测`, `部署`, `搜索`, `AI搜索`, `AI原生研发`, `AI Coding Agent`, `企业落地`, `记忆`, `开源`
|
||||
|
||||
## 执行方式
|
||||
|
||||
用 Python 脚本扫描 `content/daily/` 下所有 md 文件,对每个文件的 `tags = [...]` 替换为规范名版本。脚本参考:
|
||||
|
||||
```python
|
||||
import re, os, json
|
||||
|
||||
MAPPING = {
|
||||
"Harness工程化": "Harness工程",
|
||||
"Harness Engineering": "Harness工程",
|
||||
"Skills": "Skill",
|
||||
# ... 完整映射
|
||||
}
|
||||
|
||||
REMOVE = {"Agent", "AI", "工程化", "架构", ...}
|
||||
|
||||
hugo_dir = "/home/ubuntu/zhu/apps/hugo-site/content/daily"
|
||||
for root, _, files in os.walk(hugo_dir):
|
||||
for fname in files:
|
||||
if not fname.endswith(".md"):
|
||||
continue
|
||||
path = os.path.join(root, fname)
|
||||
with open(path, 'r') as f:
|
||||
content = f.read()
|
||||
# parse tags from frontmatter
|
||||
# replace deprecated → canonical
|
||||
# remove items in REMOVE
|
||||
# deduplicate
|
||||
# write back
|
||||
```
|
||||
Reference in New Issue
Block a user