Merge pull request 'feat/assistant-body/start-stop-status' (#4) from feat/assistant-body/start-stop-status into feat/assistant-body/main

Reviewed-on: #4
This commit was merged in pull request #4.
This commit is contained in:
2026-09-01 07:24:29 +00:00
11 changed files with 1513 additions and 76 deletions
+2 -2
View File
@@ -1,6 +1,6 @@
{
"name": "jsc-assist",
"version": "0.0.2",
"version": "0.1.0",
"description": "助理:事件收攏、健康巡檢與待辦簿(MONITOR_{HASH} wiki 頁)",
"skills": "./skills",
"author": {
@@ -18,7 +18,7 @@
"requires": {
"jsc-cli": ">=0.2.7",
"jsc-gitea": ">=0.2.0",
"jsc-hooks": ">=0.3.4",
"jsc-hooks": ">=0.3.7",
"jsc-log": ">=0.1.4"
}
}
+2 -2
View File
@@ -1,13 +1,13 @@
{
"name": "jsc-assist",
"version": "0.0.2",
"version": "0.1.0",
"description": "助理:事件收攏、健康巡檢與待辦簿(MONITOR_{HASH} wiki 頁)",
"skills": "./skills",
"jsc": {
"requires": {
"jsc-cli": ">=0.2.7",
"jsc-gitea": ">=0.2.0",
"jsc-hooks": ">=0.3.4",
"jsc-hooks": ">=0.3.7",
"jsc-log": ">=0.1.4"
}
}
+10 -4
View File
@@ -24,9 +24,9 @@ Marketplace 統一為 `jsc`(https://gitea.jsc.idv.tw/plugins/meta.git),安
<!-- JSC-SKILLS:START -->
### `status`
### `assistant`
查助理現在的狀況,全程唯讀。讀心跳檔判斷助理是不是還在跑——檔案在、而且 `ts` 距現在不到 300 秒才算新鮮,過期就是沒在跑;接著列出待辦簿裡的每一筆,印成一張現況表。心跳檔不存在時印「助理未運行」,不當成錯誤。這支不寫檔、不寫 wiki、不碰閘門。
助理主體,四個操作:`start` 啟動、`status` 查現況、`patrol` 跑一輪巡檢、`stop` 停止。心跳的寫入、判定與清除一律交給 `jsc-hooks` 的 `hooks/heartbeat.sh`,判定只有那一份;系統排程一律交給 `tools/schedule.sh`;一輪巡檢的流程交給 `tools/patrol.sh`。**心跳由巡檢寫,而且只由巡檢寫**:一輪跑完、結果寫上監控頁了,才寫那一次心跳,所以心跳新鮮等於「上一輪巡檢真的做完了」。`start` 先跑一輪巡檢,再裝上巡檢那一筆排程;巡檢週期由心跳的過期門檻算出來,兩個數字綁在一起。`patrol` 讀四項來源(使用統計、版本與重啟閘門、SDLC 階段鎖與工作包鎖、心跳自述),四項各自獨立,一項掛掉其餘三項照跑、照記,結果一律附加到 `MONITOR_{HASH}`、不覆寫。`status` 全程唯讀,讀心跳、排程與待辦簿,印成三塊;助理沒在跑就印「助理未運行」,不當成錯誤。`stop` 先移除排程再清掉心跳,順序不能反。這支不參與閘門判定、不做決策、巡檢那一路全程不問人。
<!-- JSC-SKILLS:END -->
@@ -36,13 +36,15 @@ Marketplace 統一為 `jsc`(https://gitea.jsc.idv.tw/plugins/meta.git),安
| --- | --- | --- |
| `jsc-cli` | `>=0.2.7` | CLI 偵測與委派 |
| `jsc-gitea` | `>=0.2.0` | 監控頁的所有 wiki 讀寫,一律經 `tools/gitea.sh` |
| `jsc-hooks` | `>=0.3.4` | 心跳、閘門與事件來源(`$JSC_HOME` 底下的狀態檔) |
| `jsc-hooks` | `>=0.3.7` | 心跳、閘門與事件來源(`$JSC_HOME` 底下的狀態檔)。心跳的寫入、判定與清除一律走 `hooks/heartbeat.sh`,那支腳本是 `0.3.7` 才有的 |
| `jsc-log` | `>=0.1.4` | 使用統計與工作日誌的資料來源 |
## 參考與工具
| 檔案 | 用途 |
| --- | --- |
| `tools/schedule.sh` | 助理系統排程的安裝、移除與查現況。三個子命令 `install`、`remove`、`status`,只裝 `patrol` 這一筆——心跳由巡檢自己寫,`install heartbeat` 一律回 6,舊版遺留的心跳條目由 `install patrol` 順手清掉。巡檢週期由心跳的過期門檻算出來(`2 × 週期 × 60 < 門檻`,再取能整除一小時的分鐘數):門檻 300 秒是每 2 分鐘一輪,門檻 1800 秒是每 12 分鐘一輪。Linux、WSL 與 macOS 走 crontab,Windows 走 schtasks。條目行尾帶固定標記 `# jsc-assist:assistant {工作}`,只動自己那一筆,別人的排程一行都不碰。裝完會檢查排程服務在不在跑,沒跑就回 1——WSL 預設不啟動 cron。`--dry-run` 只印組出來的條目與寫回後的內容,什麼都不動 |
| `tools/patrol.sh` | 一輪巡檢的收攏與收口。三個子命令:`collect` 取鎖、讀四項來源、組出監控頁要附加的那一節與目錄頁那一列;`finish` 在監控頁寫成之後才寫心跳、換上用量快照、放掉鎖;`abort` 只放掉鎖,不寫心跳。四項來源各自獨立,一項失敗其餘三項照跑,失敗那一項在頁上寫明是「這一項失敗」而不是沒資料。整輪拿一把目錄鎖,上一輪還在跑就回 4 讓開;鎖逾時(門檻取心跳門檻)會被下一輪搶回來,並在頁上記一筆。`version-guard.sh report` 回「查詢失敗」時照原字抄,不補查、不美化 |
| `references/behaviors.md` | 本 domain 的技能行為清單:一支技能一節,五列記下觸發時機、關鍵步驟、外部呼叫、完成條件、可驗證跡象,供稽核與驗證比對。格式合約見 `plugins/meta` 的 `references/guidelines.md`「技能行為清單」 |
| `templates/monitor-contents.md` | 目錄頁 `MONITOR_CONTENTS` 的範本。一列代表一台機器,雜湊來源是 `{主機名}/{登入帳號}`。寫入語意是**只更新自己那一列**:比對主機與帳號兩欄,別台機器的列原樣保留,禁止整頁覆蓋 |
| `templates/monitor-page.md` | 內容頁 `MONITOR_{HASH}` 的範本。記的是這台機器的巡檢軌跡。寫入語意與目錄頁相反,是**一律附加一節、不覆寫**:一次巡檢一節,節標題帶時間戳,既有的節一個字都不動 |
@@ -53,8 +55,12 @@ Marketplace 統一為 `jsc`(https://gitea.jsc.idv.tw/plugins/meta.git),安
| 路徑 | 內容 |
| --- | --- |
| `heartbeat` | 心跳檔,欄位 `ts`、`pid`、`cli`、`session`。判準只看 `ts`,不看 pid 存活——五支 CLI 與容器裡的行程互相看不到彼此的 pid |
| `heartbeat` | 心跳檔,欄位 `ts`、`pid`、`cli`、`session`。只由 `tools/patrol.sh finish` 寫,也就是一輪巡檢跑完、結果記下來之後才寫。判準只看 `ts`,不看 pid 存活——五支 CLI 與容器裡的行程互相看不到彼此的 pid |
| `tasks/{id}` | 待辦簿,一筆一檔。一筆一檔是為了讓並行寫入不互相覆寫 |
| `schedule.log` | 排程條目的輸出。刻意放在存取庫外面:寫進專案會多出未追蹤檔,污染別人的變更盤點 |
| `patrol.lock/` | 一輪巡檢的鎖,是目錄——`mkdir` 是原子操作,搶不到就是別人在跑。裡面的 `info` 記 `round`、`pid`、`started` |
| `patrol/` | 本輪巡檢的暫存檔:`section.md` 是要附加的那一節,`newpage.md` 是頁不存在時要建的整頁,`contents.tsv` 是目錄頁那一列 |
| `usage-prev.tsv` | 上一輪記下來的累計用量。有了它,下一輪的「本輪次數」才算得出來;沒有它的第一輪一律寫「-」,不拿累計冒充本輪 |
## 相關 domain
+2 -2
View File
@@ -1,13 +1,13 @@
{
"name": "jsc-assist",
"version": "0.0.2",
"version": "0.1.0",
"description": "助理:事件收攏、健康巡檢與待辦簿(MONITOR_{HASH} wiki 頁)",
"skills": "./skills/",
"jsc": {
"requires": {
"jsc-cli": ">=0.2.7",
"jsc-gitea": ">=0.2.0",
"jsc-hooks": ">=0.3.4",
"jsc-hooks": ">=0.3.7",
"jsc-log": ">=0.1.4"
}
}
+6 -6
View File
@@ -2,12 +2,12 @@
本頁記錄 jsc-assist 每支技能的行為基準,供技能驗證比對。技能異動時,在同一個 PR 內一起更新這一頁。
## status
## assistant
| 項目 | 內容 |
| --- | --- |
| 觸發時機 | 有人問助理現在還在不在跑,或問待辦簿裡剩下哪幾筆時用。啟動與停止助理不走這支。執行環境健檢不走這支,走 `jsc-cli:doctor`。技能使用次數不走這支,走 `jsc-log:stats` |
| 關鍵步驟 | 解出 `$JSC_HOME`(未設定就退回 `~/.jsc`)並組出 `assistant/` 目錄、讀 `heartbeat` 的 `ts`、`pid`、`cli`、`session`、以 `ts` 距現在是否不到 300 秒判成新鮮或過期、不看 pid 存活、列出 `tasks/` 底下每一個檔案並解析 `state`、`title`、`next_run`、`fail_count`、把心跳區塊與逐筆待辦印成一張表、`fail_count` 大於 0 的列標上「已連續失敗 N 次」 |
| 外部呼叫 | 無。只讀 `$JSC_HOME/assistant/heartbeat` 與 `$JSC_HOME/assistant/tasks/` 底下的檔案。不呼叫腳本、不呼叫其他技能、不碰 wiki、不啟動也不停止助理 |
| 完成條件 | 印出現況表,或印出「助理未運行」並說明是哪個路徑讀不到。心跳檔不存在、待辦簿目錄不存在、待辦簿零筆,三種都算正常結束,不得以非 0 結束 |
| 可驗證跡象 | 無寫入跡象,只有回報內容 |
| 觸發時機 | 要啟動助理、要停止助理、要跑一輪巡檢,或要問助理現在還在不在跑、待辦簿剩下哪幾筆時用。四個操作 `start`、`status`、`patrol`、`stop` 都走這一支。排程每一輪叫起來的也是這一支的 `patrol`。執行環境健檢不走這支,走 `jsc-cli:doctor`。技能使用次數不走這支,走 `jsc-log:stats` |
| 關鍵步驟 | 先認出使用者要的是哪一個操作,`patrol` 那一路全程不問人。`start`:先照 `patrol` 的每一步跑完一輪巡檢,第一次心跳由那一輪寫、不另外寫、跑不完就不算啟動、跑 `heartbeat.sh report` 確認 `state=fresh`、跑 `tools/schedule.sh install patrol` 裝巡檢那一筆排程、依結束碼選一段收尾訊息印出——排程接上、排程寫進去了但 cron 沒在跑、排程沒接上三種各一段。心跳那一筆不裝了,`install heartbeat` 一律回 6。`patrol`:跑 `tools/patrol.sh collect` 取鎖並讀四項來源、結束碼 4 就讓開不寫任何東西、結束碼 1 與 3 照樣把那一節寫上監控頁、`hash` 是空的就 `abort`、把 `section_file` 交給 `jsc-gitea:wiki` 附加到 `MONITOR_{HASH}`、頁不存在(唯有結束碼 4)才用 `newpage_file` 建頁、把 `contents_file` 的 `row` 更新到 `MONITOR_CONTENTS` 自己那一列、兩次寫入任一失敗就 `abort` 且不寫心跳、全部寫成才跑 `tools/patrol.sh finish` 寫心跳、最後印出四項結果與待人處理列。`status`:跑 `heartbeat.sh report` 取心跳現況、把 `state` 對映成新鮮、過期、心跳檔損壞、不存在、不自己解析心跳檔也不自己判定、從 `file=` 解出助理目錄後列出 `tasks/` 底下每一個檔案並解析 `state`、`title`、`next_run`、`fail_count`、跑 `tools/schedule.sh status` 取排程現況與週期、印成心跳、排程、待辦三塊、`fail_count` 大於 0 的列標上「已連續失敗 N 次」、心跳與排程兜起來會誤讀的四種組合各補一句話。`stop`:先跑 `heartbeat.sh report` 留下原本的狀態、再跑 `tools/schedule.sh remove all` 移除排程與舊版遺留的心跳條目、最後才跑 `heartbeat.sh clear` 清掉心跳、印出停止訊息並說明心跳清掉之後閘門會擋人、同時說明閘門還沒接線所以現在擋不到人 |
| 外部呼叫 | `jsc-hooks/hooks/heartbeat.sh` 的 `write`、`report`、`clear` 三個子命令,六個結束碼各有處置:0 往下走、1 與 3 印「助理未運行」、2 回報判不出狀態並停下、4 當成不新鮮並回報心跳檔損壞、5 是嚴重狀況要吵出來且不得回報成功、6 是呼叫寫錯要更正後重跑。`write` 只由 `tools/patrol.sh finish` 呼叫,技能自己不呼叫。本 domain 的 `tools/schedule.sh` 的 `install`、`remove`、`status` 三個子命令,七個結束碼各有處置:0 往下走、1 是條目裝了但 cron 沒在跑要照實講不會執行、2 是缺 jsc-hooks 導致門檻讀不到、3 是這台機器沒有排程機制、4 是排程操作失敗要原樣引用 stderr、5 是回讀驗證失敗要叫人自己去看 `crontab -l`、6 是呼叫寫錯,含 `install heartbeat` 與週期塞不進門檻。本 domain 的 `tools/patrol.sh` 的 `collect`、`finish`、`abort` 三個子命令,七個結束碼各有處置:0 往下走、1 部分失敗照樣寫頁、2 是 finish 找不到 heartbeat.sh 要回報「記下來了但沒有心跳」、3 是四項全失敗照樣寫頁且判定異常、4 是讓開或鎖被搶走一律不寫心跳、5 是檔案系統失敗要吵出來、6 是呼叫寫錯。巡檢那四項讀 `jsc-log/tools/usage-stats.sh`、`jsc-hooks/hooks/version-guard.sh report`、`jsc-hooks/hooks/restart-gate.sh report`、`$JSC_HOME/sessions/*.stage`、`$JSC_HOME/wp/*.pr`、`heartbeat.sh report`,全部只讀,任一項失敗不影響其餘三項。wiki 讀寫一律經 `jsc-gitea:wiki`,技能自己不拼 API 呼叫。crontab 與 schtasks 一律經 `tools/schedule.sh`。另外唯讀 `$JSC_HOME/assistant/tasks/` 底下的檔案。呼叫端沒講清楚要哪一個操作時走 `jsc-ask:ask` 的決策樹問,但 `patrol` 那一路一律不問。不參與閘門判定 |
| 完成條件 | `start` 要那一輪巡檢的 `finish` 回 0 且 `report` 回 `state=fresh`,才算啟動成功;巡檢沒寫成心跳一律回報失敗並停下,不得宣稱啟動;`schedule.sh install patrol` 回 1 要講明條目不會被執行與 `sudo service cron start`,不得宣稱排程會定時執行。`patrol` 要四項各自有 `status`、監控頁附加成功、目錄頁那一列更新成功、`finish` 回 0,才算一輪跑完;`collect` 回 4 是讓開,不算失敗也不寫任何東西;監控頁或目錄頁任一沒寫成就 `abort`,回報「這一輪沒有結果」,心跳一定不寫。`status` 要印出現況表,或印出「助理未運行」並說明原因;心跳不存在、待辦簿目錄不存在、待辦簿零筆、排程沒裝,四種都算正常結束。`stop` 要 `schedule.sh remove all` 先回 0、`clear` 再回 0,並印出帶三段話的停止訊息;`remove` 非 0 就回報排程還在、助理停不掉,不清心跳也不印停止訊息;`clear` 回 5 就回報心跳檔還在、助理沒有確實停掉,不印停止訊息 |
| 可驗證跡象 | `start` 之後 `$JSC_HOME/assistant/heartbeat` 存在,`ts` 是剛才那一輪的時間,`crontab -l` 找得到一筆帶 `# jsc-assist:assistant patrol` 的條目,而且只有一筆,帶 `# jsc-assist:assistant heartbeat` 的舊條目一筆都不剩。`patrol` 跑完之後 wiki 的 `MONITOR_{HASH}` 多一節、節標題帶時間戳、舊的節一字不改,`MONITOR_CONTENTS` 只有自己那一列變動,`$JSC_HOME/assistant/patrol/` 底下有本輪的 `section.md`、`newpage.md`、`contents.tsv`,`$JSC_HOME/assistant/usage-prev.tsv` 換成本輪的累計數,`$JSC_HOME/assistant/patrol.lock` 已經放掉。讓開的那一輪沒有任何寫入跡象。`stop` 之後心跳路徑不存在,`crontab -l` 找不到任何 `# jsc-assist:assistant` 條目。以上都不動別人的排程條目,條目數量前後相同。`status` 無寫入跡象,只有回報內容。四個操作都不動 `tasks/` 底下的檔案,也不動 worktree 與程式碼存取庫。排程的 log 一律在 `$JSC_HOME/assistant/schedule.log`,不落在任何存取庫 |
+218
View File
@@ -0,0 +1,218 @@
---
name: assistant
description: Start, inspect, patrol, or stop the background assistant, with jsc-hooks/hooks/heartbeat.sh owning the single freshness verdict, tools/schedule.sh owning the system scheduler, and tools/patrol.sh owning one patrol round. The heartbeat is written by a completed patrol round and by nothing else, so the schedule carries the patrol entry only and its period is derived from the heartbeat TTL; start runs one round and then installs that entry, status turns heartbeat.sh report, schedule.sh status and the task book into one read-only table, stop removes the entry first and then clears the heartbeat. One round reads four independent sources - skill and chain usage, version gaps and the restart gate, SDLC stage and work-package locks, and the heartbeat's own report - and appends the result to wiki MONITOR_{HASH} through jsc-gitea:wiki before tools/patrol.sh finish writes the heartbeat. A round that cannot record its result writes no heartbeat, and a round that starts while the previous one still holds the lock stands down. Use when someone starts, patrols or stops the assistant, or asks whether it is running and what is queued; not for environment health checks (jsc-cli:doctor), not for skill usage counts (jsc-log:stats).
---
# assistant — start, status, patrol, stop
The background assistant runs where nobody is watching it. Its heartbeat is the only evidence that it is alive, so this skill is the single entry point for the four operations that touch that evidence: `patrol` writes it, `status` reads it, `stop` clears it, and `start` bootstraps the whole loop.
`jsc-hooks/hooks/heartbeat.sh` owns every heartbeat operation, including the freshness verdict. Never read, parse, write or delete `$JSC_HOME/assistant/heartbeat` directly — one verdict, one source.
`tools/schedule.sh` owns every system-scheduler operation: installing an entry, removing it, and reading which entries exist. Never call `crontab` or `schtasks` from this skill, and never edit a crontab by hand.
`tools/patrol.sh` owns one patrol round: taking the round lock, reading the four sources, composing the monitor-page section, and — after that section is on the page — writing the heartbeat. Never re-read a source this skill already handed to that script, and never compose the section by hand; the script prints the file paths.
All three flows have fixed inputs and outputs, so all three live in scripts. The task book is the only thing this skill reads for itself, and that is one directory listing.
## Pick the operation
Run exactly one operation per invocation. Take it from the request: starting, launching or waking the assistant is `start`; asking whether it runs, what it is doing, or what is queued is `status`; running one round, patrolling, or a scheduled wake-up is `patrol`; stopping, halting or shutting it down is `stop`. When the request names none of the four, or names more than one, ask through the `jsc-ask:ask` decision tree with those four as the options, each stating its effect — `start` runs one round and installs the scheduled entry that keeps running rounds, `status` changes nothing, `patrol` runs one round and writes one heartbeat, `stop` removes that entry and deletes the heartbeat. **The one exception: a `patrol` invocation never asks anything at all** (see 界線 1 below). Never guess, and never run a second operation the caller did not ask for. Completion condition: exactly one of `start`, `status`, `patrol`, `stop` is chosen and named in the report.
## Data sources
| Path | Read by | Format |
| --- | --- | --- |
| `$JSC_HOME/assistant/heartbeat` | `heartbeat.sh` only, never this skill | `key=value` lines: `ts`, `pid`, `cli`, `session` |
| `$JSC_HOME/assistant/schedule.log` | nobody here — the scheduled entry appends to it | free text; point the operator at it when a scheduled round misbehaves |
| `$JSC_HOME/assistant/tasks/{id}` | this skill, read-only | `key=value` lines, one task per file: `id`, `kind` (`check` / `todo`), `title`, `action`, `trigger`, `recur`, `repo`, `due`, `state` (`pending` / `done` / `paused`), `last_run`, `next_run`, `fail_count`, `origin` (`user` / `assistant`) |
| `$JSC_HOME/assistant/patrol.lock/` | `patrol.sh` only | the round lock, a directory. `info` holds `round`, `pid`, `started` |
| `$JSC_HOME/assistant/patrol/` | `patrol.sh` only | one round's scratch files, including `section.md`, `newpage.md` and `contents.tsv` |
| `$JSC_HOME/assistant/usage-prev.tsv` | `patrol.sh` only | last recorded round's cumulative usage counts, so the next round can print a real per-round delta |
`$JSC_HOME` defaults to `~/.jsc`. `heartbeat.sh report` prints the resolved heartbeat path in its `file=` field, so take the assistant directory from there rather than rebuilding it.
**The verdict is time-based only.** A heartbeat counts as fresh when the file exists and its `ts` is less than the TTL behind now (300 seconds by default, `JSC_ASSISTANT_HEARTBEAT_TTL` overrides it). `pid` liveness is never tested: five CLIs and container processes cannot see each other's pids, so a live-looking pid proves nothing and a missing one proves nothing either. Report `pid` as a hint for whoever has to find a blocking process, and give it no weight in the verdict.
## heartbeat.sh exit codes
Every call in every operation below is judged by this table. Report the code you got, then take the row's action — never retry a code silently, and never downgrade a failure into a success.
| Code | Meaning | What to do |
| --- | --- | --- |
| 0 | `write` wrote the heartbeat, `clear` finished and the file is gone, `report` printed its line, `check` says fresh | Carry on with the operation's next step. For `report`, the state still has to be read out of the printed `state=` field |
| 1 | `check`: the heartbeat exists but is at or past the TTL — the last patrol round finished more than one TTL ago | Report `助理未運行`, name the age in seconds, and say the assistant has to be started again. `report` returns this state as `state=stale` with exit 0 |
| 2 | The script did not run at all — it failed to load its `lib.sh` | Report that the heartbeat state is unknown, name the script path and the code, and stop the operation. Never claim the assistant is running, and never claim it is stopped |
| 3 | `check`: no heartbeat file — no patrol round has ever finished | Report `助理未運行` and say to run `start`. `report` returns this state as `state=absent` with exit 0. In `stop` this state cannot appear, because `clear` treats a missing file as success |
| 4 | `check`: the heartbeat exists but its `ts` is missing, empty or not a number — the file is damaged, the assistant is not merely stopped | Treat it as not fresh; falling back to fresh is forbidden. Report the file as damaged, say the state cannot be read from it, and tell the operator to run `stop` and then `start` to rebuild it. `report` returns this state as `state=invalid` with exit 0 |
| 5 | Filesystem failure — `write` could not write the file, or `clear` could not delete it and the file is still there | Serious. Report it loudly with the stderr text and the path, and follow the operation's own step for this code. Never report the operation as done |
| 6 | Usage error — an unknown subcommand, or none at all | This is a defect in the call, not a state of the assistant. Report the exact command line that was run, correct it to one of `write`, `check`, `report`, `clear`, and run it once more. Report a second exit 6 as a defect in this skill and stop |
## The scheduler
Nothing in a background assistant runs on its own. The system scheduler is what makes it periodic, and `tools/schedule.sh` is the only thing here that touches it. One job exists, written as exactly one entry carrying the fixed marker `# jsc-assist:assistant patrol`:
| Job | Period | Runs | Installed by `start` |
| --- | --- | --- | --- |
| `patrol` | derived from the heartbeat TTL (`*/2 * * * *` at the default TTL of 300 seconds) | one patrol round through the caller's CLI | yes, always |
| `heartbeat` | — | nothing. This job existed in the previous version and is no longer installable | no — `install heartbeat` exits 6 |
**The heartbeat job is gone on purpose.** It used to call `heartbeat.sh write` every minute, which made a fresh heartbeat prove only that cron was alive. Anything that writes a heartbeat outside a finished patrol round brings that back, so `install heartbeat` is refused, and `install patrol` deletes any leftover `heartbeat` entry from an older install and reports `legacy_removed=1`. Say that number in the report — a surviving legacy entry silently undoes this whole design.
**The period is derived, never guessed.** The heartbeat now moves once per patrol round, so the round period has to fit inside the freshness threshold. `schedule.sh` reads the machine's effective threshold from `heartbeat.sh report`'s `ttl=` field and picks the largest whole-hour-dividing minute count `P` with `2 × P × 60 < ttl`: one missed round still reads fresh, two missed rounds read stale. At the default 300 seconds that is every 2 minutes; raise `JSC_ASSISTANT_HEARTBEAT_TTL` to 1800 and it becomes every 12 minutes. Report both numbers (`ttl=`, `period=`) so the operator can see the trade-off and change it in one place. `--period` overrides the calculation and is checked against the same inequality; a period that does not fit exits 6 rather than installing a schedule that keeps the heartbeat permanently stale.
The mechanism follows the platform: `crontab` on Linux, WSL and macOS, `schtasks` on Windows. macOS keeps `crontab` — a `launchd` user who wants a plist writes it themselves; this skill does not generate one.
Four properties of that script matter enough to state here, because a report that ignores any of them is wrong:
- **It only ever touches its own entries.** Install filters out its own old entries by marker and appends the new one; it never rewrites a crontab it failed to read, and it counts everybody else's lines before and after to prove none went missing. Remove takes out its own markers only. Say this in the report — the operator is entitled to know their own cron entries survived.
- **A written entry is not a running entry.** WSL does not start cron by default, and this is the machine's most likely state. Exit 1 from `install` means the entry is on disk and will never fire. Report that as a failure of the start, name `sudo service cron start`, and say it has to be run again after every WSL restart. Never soften exit 1 into "scheduling is set up".
- **The log lives at `$JSC_HOME/assistant/schedule.log`**, deliberately outside every repository. Do not offer to move it into a project.
- **The entry runs with no human present.** The command is installed with `</dev/null`, so nothing it runs can block on input. A patrol round that stops to ask for a tool permission hangs that round, and the lock it holds stands the next round down until the lock ages out — which is why `patrol` asks nothing, of anybody, ever.
### What a fresh heartbeat actually proves
The heartbeat is written in exactly one place: `tools/patrol.sh finish`, and `finish` is called only after that round's result is on the monitor page. So the verdict 新鮮 now proves one thing that is worth proving — **the last patrol round ran to the end and its result was recorded** — and it still does not prove three others:
- **Not that the round was clean.** Four sources are read independently and a round with three failures still records and still beats. The health of a round is `本輪判定` on the monitor page, never the heartbeat.
- **Not that any task in the book moved.** The task rows — `last_run`, `next_run`, `fail_count` — are the only evidence about work.
- **Not that the round did anything about what it found.** The patrol reports; a human acts. 界線 6.
The failure this design buys is the one worth having: a round that cannot read its sources, cannot reach the wiki, or dies half way writes no heartbeat, so the heartbeat ages past the TTL and every reader sees 過期. **A silent patrol is now indistinguishable from a stopped assistant, which is exactly right.** `status` still prints the heartbeat, the schedule and the task book as three separate facts, and the same limit binds whatever gate reads this heartbeat later: a fresh heartbeat is grounds for not blocking, never grounds for saying the assistant is doing its job.
## schedule.sh exit codes
| Code | Meaning | What to do |
| --- | --- | --- |
| 0 | `install` wrote the entry and read it back, the scheduler service is running; `remove` finished, or there was nothing to remove; `status` printed its lines | Carry on. For `status`, the state still has to be read out of the `installed=` fields |
| 1 | `install` wrote the entry, but the cron service is not running — the entry will never fire | The start did not succeed. Report the entry as installed and inert, quote the fix (`sudo service cron start`, and again after each WSL restart), and never claim the assistant will keep itself alive |
| 2 | `jsc-hooks/hooks/heartbeat.sh` was not found, so the TTL cannot be read and the period cannot be derived | Report that `jsc-hooks` is missing or too old (0.3.7 or newer is required) and stop the operation |
| 3 | No usable scheduler on this machine | Report the platform and that neither `crontab` nor `schtasks` was found, and stop. Never fall back to some other mechanism |
| 4 | The scheduler operation failed — the existing schedule could not be read for a reason other than "no crontab", or the write or delete returned non-zero | Report the stderr text verbatim. A read failure means nothing was written, so the user's other entries are untouched; say so |
| 5 | Read-back verification failed — the entry is missing after a successful write, is present twice, is still there after a delete, or somebody else's line count changed | Serious. Report it loudly with the printed numbers, and tell the operator to inspect `crontab -l` by hand before anything else is run |
| 6 | Usage error — an unknown subcommand or job name, a missing option value, `install heartbeat`, a `--period` that does not fit the TTL, or the patrol CLI could not be determined | A defect in the call, not a state of the machine. Correct the command line and run it once more; report a second exit 6 as a defect in this skill and stop |
## patrol.sh exit codes
One table for all three subcommands. Read `collect`'s codes carefully: **1 and 3 are results, not aborts.** A round with failed items still has a section to write, and refusing to write it would hide the failure instead of recording it.
| Code | Meaning | What to do |
| --- | --- | --- |
| 0 | `collect`: all four items read to the end, empty sources included. `finish`: heartbeat written, snapshot promoted, lock released. `abort`: lock released | Carry on with the operation's next step |
| 1 | `collect`: partial success — at least one item failed and at least one produced a result | **Write the page anyway.** The section already marks the failed items and the round verdict is 警示. Name the failed items and their `note=` text in the report |
| 2 | `finish`: `jsc-hooks/hooks/heartbeat.sh` was not found | The round completed and is recorded, but no heartbeat exists to prove it. Report the round as recorded and the heartbeat as not written, say `jsc-hooks` 0.3.7 or newer has to be installed, and run `tools/patrol.sh abort --round {id}` to release the lock |
| 3 | `collect`: all four items failed | **Write the page anyway**, with verdict 異常. A page listing four failures is the signal; a missing page is not. Then carry on to `finish` as usual — the round did complete |
| 4 | Another round holds the lock (`collect`), or the lock is no longer this round's (`finish`, `abort`) | Not a failure. On `collect`: report 本輪讓開 and name the holder and its age from the printed `lock=busy` line, then write nothing and stop. On `finish`: the previous round overran and was taken over, so this round's result does not count — report it, write no heartbeat, and stop |
| 5 | Filesystem failure — the lock could not be created or released, a scratch file could not be written, the snapshot could not be promoted, or `heartbeat.sh write` returned non-zero | Serious. Report it loudly with the stderr text and the path. On a `finish` failure the round is recorded but unproven: say so plainly and never claim the round beat |
| 6 | Usage error — an unknown subcommand, a missing `--round`, or an option with no value | A defect in the call. Correct it and run it once more; report a second exit 6 as a defect in this skill and stop |
## Boundaries
The six limits in `AGENTS.md`「助理的界線」 hold for all four operations. Four of them need saying out loud here:
- **This skill never judges a gate.** It maintains the heartbeat and prints what the heartbeat says. Whether a stale heartbeat blocks a skill call is decided by a hook, synchronously and offline; nothing in this skill blocks or waves through anything. 界線 2.
- **A patrol round asks nothing.** It runs from cron with nobody present, so there is no one to answer and a question hangs the round. Every branch in the patrol steps below resolves without a question: a missing source is recorded as missing, an ambiguous result is recorded verbatim, and a round that cannot proceed aborts and reports. Never call `jsc-ask:ask` from `patrol`. 界線 1.
- **A patrol round only ever appends to the monitor page.** Read the old page back first, append one section, put the whole page. The contents page gets its own row updated and nobody else's. A page that could not be read is a page that does not get written. 界線 4.
- **A patrol round reports; it never acts on what it found.** The 待人處理 rows name an entry point for a human. The patrol does not run that entry point, does not fix a hook, does not update a plugin and does not touch a repository. 界線 3 and 界線 6.
- **`stop` clearing the heartbeat and removing the schedule is not a breach of 界線 5「不刪除狀態檔」.** That limit protects state that records work — the task book, worktrees, wiki pages — from a background process nobody is watching. The heartbeat records one fact only, "the last patrol round finished", and the schedule entry is what keeps rounds running, so a `stop` that leaves either behind leaves a lie behind. Clearing both is the whole job of `stop`, and they are the only deletions any operation here performs, both of them entries this skill installed itself. `stop` touches nothing under `tasks/`, nobody else's cron entry, no worktree and no wiki page. Do not "restore" this limit later by taking either removal out of `stop`.
## Crash exit needs no cleanup
An assistant that is killed, crashes, or dies with the machine writes no farewell. It does not need to. The heartbeat is a timestamp, not a lock: the last one written stays on disk, ages past the TTL on its own, and every reader from then on sees 過期. No shutdown handler, no cleanup hook and no pid check is involved, so there is nothing left that can fail to run.
The round lock is the one thing a crash does leave behind, and it ages out the same way: `patrol.sh collect` breaks a lock older than the heartbeat TTL, takes it, and prints `lock_broken=1` so the takeover lands on the monitor page instead of happening quietly. The overrun round that lost its lock then gets exit 4 from `finish` and writes no heartbeat, which is correct — it never reached the end.
That property holds only while nothing fakes a heartbeat. **`write` is called by `tools/patrol.sh finish` and nowhere else.** `start` does not call it, `status` does not call it, `stop` does not call it, no scheduled entry calls it, and no other skill calls it. A heartbeat written by anything that is not a finished round says a round finished when none did, and the reader has no way to tell the difference. This is also why `stop` removes the scheduled entry before clearing the heartbeat, and never in the other order.
## start
`start` proves the loop works before it schedules it: one patrol round first, then the scheduled entry. It installs no daemon and writes no bare heartbeat.
1. **Run one patrol round.** Follow every step of the `patrol` operation below, start to finish. This is what writes the first heartbeat — there is no shortcut past it, because a heartbeat that no round produced is exactly the lie this design removes. When that round ends without a heartbeat for any reason (`collect` exit 4, 5 or 6, an empty `hash=`, a failed wiki write, or `finish` exit 2, 4 or 5), the start has failed: report the round's outcome and the code, do not run step 2, and do not claim a started assistant. A round that completed with failed items (`collect` exit 1 or 3) is still a completed round — carry on to step 2 and name the failures in the closing report. Completion condition: `patrol.sh finish` exited 0, or the failure report naming the step and the code has been printed and no start was claimed.
2. **Confirm the heartbeat.** Run `jsc-hooks/hooks/heartbeat.sh report` and read its `state=`, `ts=`, `ttl=`, `pid=`, `cli=`, `session=` and `file=` fields. `state=fresh` is the expected result. Any other state right after a successful round means something rewrote or removed the file in between: report the state, the path and that the heartbeat did not survive its own write, and do not claim a started assistant. Completion condition: the report line was read and either `state=fresh` was recorded with its seven fields, or the mismatch was reported.
3. **Install the patrol entry.** Run `tools/schedule.sh install patrol`. Judge the result by the schedule.sh exit-code table, and keep the printed `entry=`, `ttl=`, `period=`, `legacy_removed=`, `others_kept=` and `service=` fields for the report. Exit 1 is the case to get right: the entry is installed and inert, so step 4 reports a started assistant whose heartbeat will expire, not a scheduled one. On 2, 3, 4, 5 or 6 nothing is scheduled — report the code, say the round ran but no further round will, and do not claim the assistant will stay alive. Completion condition: the exit code is recorded, and on exit 0 the printed entry line, the TTL, the period, the legacy count and the surviving-entry count are recorded with it.
4. **Report the start.** Print the round's verdict and its four item results, the monitor page that was written, the heartbeat path, the local time of `ts`, the TTL in seconds, `pid`, `cli` and `session` as hints, then the scheduler mechanism, the derived period, the installed entry line, how many legacy heartbeat entries were removed, and how many other entries were left untouched. Close with the notice that matches step 3's outcome, printed literally with `{ttl}` replaced by the TTL just read and `{period}` by the derived period:
| Step 3 | Notice |
| --- | --- |
| exit 0 | 助理已啟動,第一輪巡檢跑完了,結果寫上監控頁了,心跳也寫了。排程接上了,之後每 {period} 分鐘跑一輪,每一輪跑完才寫一次心跳。心跳新鮮代表上一輪巡檢真的做完了;那一輪四項有沒有全過,看監控頁的本輪判定。 |
| exit 1 | 助理已啟動,第一輪巡檢跑完了,排程條目也寫進去了,但 cron 服務沒在跑,那一筆一次都不會被執行。心跳過了 {ttl} 秒就會過期。請先跑 `sudo service cron start`,重開 WSL 之後要再跑一次。 |
| 其他結束碼 | 助理已啟動,第一輪巡檢跑完了,但排程沒接上(結束碼 {code})。不會再有下一輪,心跳過了 {ttl} 秒就會過期,屆時請再跑一次 start。 |
Completion condition: the report carries the round verdict, the monitor page name, the path, the local heartbeat time, the TTL, the period, the three hint fields and the scheduler outcome, and exactly one notice above appears with the real numbers.
## patrol
One round: read four sources, record the result, then beat. Everything before the heartbeat is read-only except the round's own scratch files. Ask nobody anything.
1. **Collect.** Run `tools/patrol.sh collect --trigger 排程` (use `--trigger 手動` when a person asked for this round). Judge the exit code by the patrol.sh table. Exit 4 stands the round down — report the holder and its age from the printed `lock=busy` line, and stop; write no page and no heartbeat. Exit 5 and 6 stop the round the same way, with the code and the stderr text. Exit 0, 1 and 3 all carry on to step 2. Record `round=`, `lock_broken=`, `hash=`, `page=`, `verdict=`, `failed_sources=`, every `item=` line, and the three file paths `section_file=`, `newpage_file=` and `contents_file=`. Completion condition: the round id, the page name and the three file paths are recorded, or the stand-down or the failure was reported and the round stopped.
2. **Check the page name.** An empty `hash=` means `jsc-gitea/tools/hash-id` could not be found or could not run, so there is no page to write to and nothing can be recorded. Run `tools/patrol.sh abort --round {round}`, report that the round found its results but has nowhere to put them, name `jsc-gitea` as missing, and stop. Never invent a page name — a hand-made name lands the content on a page nobody reads. Completion condition: `page=` holds a `MONITOR_{HASH}` name, or the abort ran and the round was reported as unrecorded.
3. **Append the section to `MONITOR_{HASH}`.** Hand it to `jsc-gitea:wiki` with page type `MONITOR`: read the page back first, then append the whole content of `section_file` as a new last section and put the whole page. Only exit 4 from the read permits creating the page instead, and then the page body is the whole content of `newpage_file`, which already carries the basic-data section plus this round's section. Exit 7 and exit 8 mean the old content is unknown: create nothing, write nothing. On any write failure — including exit 3 with no wiki repo configured for `MONITOR`, which the patrol cannot ask about — run `tools/patrol.sh abort --round {round}`, report the code, and stop. **No record, no heartbeat.** Completion condition: the append or the create returned success, or the abort ran and the round was reported as unrecorded with its exit code.
4. **Update this machine's row in `MONITOR_CONTENTS`.** Take the `row=` line from `contents_file` — it is already the finished table row. Hand it to `jsc-gitea:wiki`: read the whole page, match the row whose 主機 and 帳號 columns both equal this round's `host=` and `user=`, overwrite that row's remaining columns, and put the whole page back. No matching row means append one. Never overwrite the whole page, and never touch another machine's row — the write semantics here are the opposite of the content page's, and mixing them up deletes other machines' records. On failure, run `tools/patrol.sh abort --round {round}`, report the code, and stop. Completion condition: exactly one row carries this machine's 主機 and 帳號 values, every other row is byte-identical to what was read, and the put returned success.
5. **Write the heartbeat.** Run `tools/patrol.sh finish --round {round}`. This is the last step for a reason: it is the only thing that turns a fresh heartbeat into a true statement. Judge the exit code by the patrol.sh table — 2, 4 and 5 all mean the round is recorded but unproven, and each has its own report line there. Completion condition: `finish` exited 0, or the failure was reported as "recorded but no heartbeat" with its code.
6. **Report the round.** Print the round verdict, one line per item with its `status=` and, for a failure, its `note=`; the monitor page name and the contents row that was written; whether the heartbeat was written; and, when `lock_broken=1`, that the previous round's lock was taken over because it had aged past the TTL. Close with the 待人處理 rows from the section, verbatim, and nothing else — the patrol names an entry point and stops there. Completion condition: all four items appear in the report, the heartbeat outcome is stated as written or not written, and no suggestion in 待人處理 was acted on.
## status
Read-only throughout. This operation creates, modifies and deletes nothing under `$JSC_HOME`, and it never calls `write` or `clear`.
1. **Read the heartbeat through the script.** Run `jsc-hooks/hooks/heartbeat.sh report` and split the line on spaces, taking `file=` last so a path containing spaces stays intact. Map `state=` to the verdict: `fresh` → `新鮮`, `stale` → `過期`, `invalid` → `心跳檔損壞`, `absent` → `不存在`. Print `助理未運行` for `stale`, `invalid` and `absent`. Never re-derive the verdict from `ts` yourself, and never treat `invalid` as fresh. On exit 2 or 6, follow that code's row, record the heartbeat state as unknown, and carry on to step 2 — the task book is still worth printing. Completion condition: the heartbeat state holds one of `新鮮`, `過期`, `心跳檔損壞`, `不存在` or unknown, and `ts`, `age`, `ttl`, `pid`, `cli`, `session` and `file` are recorded as read or as empty.
2. **Read the task book.** Take the assistant directory from the `file=` path of step 1, list the regular files directly under its `tasks/` subdirectory, and parse each one as `key=value` lines. Branch on the outcome.
| Outcome | Do |
| --- | --- |
| Directory absent | Report zero entries. This is a normal result, not an error |
| Directory present, no files | Report zero entries |
| A file cannot be read, or holds no recognisable key | Keep it as one row, put the file name in the title column, name the read or parse error in that row, and carry on with the remaining files |
| A key is missing from a readable file | Print `-` in that column |
Completion condition: every file under `tasks/` produced exactly one row, or zero entries was reported.
3. **Read the schedule.** Run `tools/schedule.sh status`. It writes nothing. Record `mechanism=`, `service=`, `ttl=`, `period=` and the `installed=` value of both jobs. A `heartbeat` job reported as installed is a leftover from an older version: say so, and say `start` or `schedule.sh install patrol` removes it. On exit 2, 3 or 6 nothing was read: record the schedule state as unknown with its code and carry on — the heartbeat and the task book still print. Completion condition: both jobs have an installed state, or the schedule state is recorded as unknown with its code.
4. **Print the status table.** Lead with the heartbeat block — verdict, last heartbeat time rendered from `ts` in local time, age in seconds, TTL, `cli`, `session`, `pid`, and the task count. Follow it with the schedule block — mechanism, service state, derived period, and one line per job saying installed or not. Then one row per task carrying `state`, `title`, `next_run` and `fail_count`, in the order the files were listed. Completion condition: the heartbeat block holds all eight values, the schedule block holds both jobs and the period, and the row count equals the task count from step 2.
5. **Say what the two blocks together mean.** Four combinations get an explicit sentence, because each one reads as something it is not:
| Heartbeat | Schedule | Say |
| --- | --- | --- |
| 新鮮 | patrol installed, service running | 上一輪巡檢跑完了,結果也記上監控頁了,排程還在跑。那一輪四項有沒有全過,要看監控頁的本輪判定 |
| 新鮮 | not installed, or service stopped | 上一輪巡檢跑完了,但沒有排程在叫下一輪,過了 TTL 心跳就會過期 |
| 過期 or 不存在 | patrol installed, service running | 排程裝著卻沒有新的心跳,巡檢自己跑失敗了,去看 `$JSC_HOME/assistant/schedule.log` 與監控頁最新一節 |
| any | `heartbeat` job installed | 舊版的心跳排程還留著,它會蓋掉「心跳等於巡檢跑完」這件事。請跑一次 `start`,或 `schedule.sh install patrol` 把它清掉 |
Completion condition: every matching sentence is printed, or none of the four combinations applied.
6. **Flag the repeatedly failing tasks.** Append 已連續失敗 N 次 to every row whose `fail_count` is above 0, with `N` taken verbatim from the file. A broken entry that retries every round with nobody noticing is the reason this field exists, so let no such row leave the table unmarked. Completion condition: every row with `fail_count` above 0 carries the marker and its number matches the file.
7. **Finish successfully.** `助理未運行`, an absent `tasks/` directory, an empty `tasks/` directory and an uninstalled schedule are normal results — never exit non-zero for any of them. Reserve a failure report for a condition none of the tables above covers, and state which path and which error produced it. Completion condition: the report is printed and nothing under `$JSC_HOME` has been created, modified or deleted.
## stop
1. **Record what is being stopped.** Run `jsc-hooks/hooks/heartbeat.sh report` first and keep its `state=`, `ts=`, `pid=`, `cli=` and `file=` fields for the closing report — after the clear they are gone for good. `state=absent` means no round has finished; say so and still run steps 2 and 3, because a scheduled entry can outlive its heartbeat and `clear` on a missing file is a success, so running both leaves the outcome unambiguous. On exit 2 or 6, follow that code's row, record the previous state as unknown, and carry on to step 2. Completion condition: the previous state and its fields are recorded, or the previous state is recorded as unknown with its code.
2. **Remove the schedule first.** Run `tools/schedule.sh remove all` — both job names, so the patrol entry and any leftover heartbeat entry from an older install both go. This comes before the clear and never after: clear first and the next scheduled round writes a fresh heartbeat over the stopped assistant, and every reader from then on is told a dead assistant is alive. Judge the result by the schedule.sh exit-code table, and keep `removed=` and `others_kept=` for the report. On any non-zero code the schedule is still installed: report the code, say plainly that rounds will keep running and the assistant therefore cannot be stopped, name the manual fix (`crontab -l` to look, then remove the line carrying `# jsc-assist:assistant` by hand), and skip steps 3 and 4 — clearing a heartbeat that the next round rewrites only hides the problem. Completion condition: `remove` exited 0 with its counts recorded, or the failure report has been printed and no stop was claimed.
3. **Clear the heartbeat.** Run `jsc-hooks/hooks/heartbeat.sh clear`. On exit 5 the file is still there: report the failure with the script's stderr line and the path, say plainly that every reader still sees a heartbeat claiming a round just finished and that the assistant is therefore not reliably stopped, name the manual fix (delete that path by hand, then run `status` to confirm `助理未運行`), and skip step 4 — the closing notice must not be printed after a failed clear. On exit 2 or 6, follow that code's row and stop the same way. Completion condition: `clear` exited 0, or the failure report naming the code, the path and the manual fix has been printed and no stop was claimed.
4. **Report the stop and what it means for the gate.** Print the previous state and heartbeat time from step 1 and the entries removed in step 2, then this literally:
> 助理已停止,排程移除了,心跳也清掉了,其他人的排程一筆都沒動。靠心跳判定的 jsc 技能閘門一讀到沒有心跳就會擋下技能呼叫;閘門目前還沒接線,所以這一刻誰都擋不到。要再工作就先跑一次 start。
Say it exactly this way. The blocking is the designed consequence of a cleared heartbeat, and whoever stops the assistant has to know it is coming; the clause about the gate being unwired is the part that keeps the notice honest while that is still true. When the gate is wired, that clause is what gets rewritten — not the rest. Completion condition: the notice appears with all three clauses, and the previous state, the heartbeat time and the removal counts are printed above it.
## Round lock and a round that will not stop leaving one behind
A patrol round holds `$JSC_HOME/assistant/patrol.lock` from `collect` to `finish` or `abort`, which spans the wiki writes — the slow part. Two consequences bind every branch above:
- **Every path out of a started round ends in `finish` or `abort`.** Steps 2, 3 and 4 of `patrol` each name their abort. A round that stops without either leaves the lock standing until it ages out, which stands the next rounds down for up to one TTL. There is no third option.
- **`stop` does not remove the lock.** It is not a state file that records work, but it is also not this skill's to delete while a round may still be using it; it ages out on its own within one TTL. If an operator reports that every round stands down, tell them the holder and age from the `lock=busy` line and let them decide — 界線 5 keeps destructive cleanup with the human.
-55
View File
@@ -1,55 +0,0 @@
---
name: status
description: Report the background assistant's current state read-only, from the heartbeat file and the task book under $JSC_HOME/assistant/. Heartbeat counts as fresh only when the file exists and its ts is less than 300 seconds old, and pid liveness is never checked. Print one table covering heartbeat freshness, last heartbeat time, cli, session, task count, and every task's state, title, next_run and fail_count, flagging each task whose fail_count is above zero. A missing heartbeat file prints 助理未運行 and still counts as a normal result rather than an error. Use when someone asks whether the assistant is running or what is queued; not for starting or stopping it, not for environment health checks (jsc-cli:doctor), and not for skill usage counts (jsc-log:stats).
---
# status — assistant heartbeat and task book snapshot
Read-only snapshot of the background assistant. This skill reads two paths and prints one table. It writes no file, writes no wiki page, calls no gate, and never starts or stops the assistant.
No `tools/` script backs this skill. Two paths and one table stay below the extraction bar; re-evaluate when the assistant body itself lands.
## Data sources
| Path | Format | Keys |
| --- | --- | --- |
| `$JSC_HOME/assistant/heartbeat` | plain text, one `key=value` per line | `ts` (epoch seconds), `pid`, `cli`, `session` |
| `$JSC_HOME/assistant/tasks/{id}` | plain text, one `key=value` per line, one entry per file | `id`, `kind` (`check` or `todo`), `title`, `action`, `trigger`, `recur`, `repo`, `due`, `state` (`pending` / `done` / `paused`), `last_run`, `next_run`, `fail_count`, `origin` (`user` or `assistant`) |
`$JSC_HOME` defaults to `~/.jsc`.
**Freshness is time-based only.** The heartbeat is fresh when the file exists and `ts` is less than 300 seconds behind the current time. An older `ts` means the assistant is not running. Never test whether `pid` is alive: the five CLIs and the processes inside containers cannot see each other's pids, so a live-looking pid proves nothing and a missing one proves nothing either. Report `pid` as a hint for whoever has to find a blocking process, and give it no weight in the verdict.
## Steps
1. **Resolve the assistant directory.** Take `$JSC_HOME` from the environment; when it is unset or empty, use `~/.jsc`. Append `assistant/` to get the directory this skill reads. When that directory is absent or cannot be listed, print `助理未運行`, name the resolved path and the reason (the variable was unset and the default path does not exist, or the listing was denied), skip steps 2 to 6, and finish per step 7. Done when one absolute assistant directory path is recorded, or the not-running report naming that path is printed.
2. **Read the heartbeat.** Read `{assistant}/heartbeat` and split each line on its first `=`. Branch on the outcome.
| Outcome | Do |
| --- | --- |
| File absent | Set heartbeat state to `不存在`, print `助理未運行`, continue at step 4 — the task book is still worth printing |
| File unreadable (permission denied, I/O error) | Set heartbeat state to `不存在`, print `助理未運行`, name the error text as the reason, continue at step 4 |
| File present, `ts` absent or not an integer | Set heartbeat state to `過期`, name the malformed value, continue at step 4 |
| File present with an integer `ts` | Continue at step 3 |
Done when the heartbeat state holds one of `不存在`, `過期`, or a pending verdict handed to step 3, and the values of `pid`, `cli` and `session` are recorded as read or as absent.
3. **Judge freshness.** Subtract `ts` from the current epoch seconds. A difference below 300 sets the state to `新鮮`; 300 or above sets it to `過期`. Done when the state is `新鮮` or `過期` and the age in seconds is recorded.
4. **Read the task book.** List the regular files directly under `{assistant}/tasks/` and parse each one as `key=value` lines. Branch on the outcome.
| Outcome | Do |
| --- | --- |
| Directory absent | Report zero entries. This is a normal result, not an error |
| Directory present, no files | Report zero entries |
| A file cannot be read or holds no recognisable key | Keep it as one row, put the file name in the title column, name the read or parse error in that row, and carry on with the remaining files |
| A key is missing from a readable file | Print `-` in that column |
Done when every file under `tasks/` has produced exactly one row, or zero entries has been reported.
5. **Print the status table.** Lead with the heartbeat block — state (`新鮮` / `過期` / `不存在`), last heartbeat time rendered from `ts` in local time, `cli`, `session`, and the task count. Follow it with one row per task carrying `state`, `title`, `next_run` and `fail_count`, in the order the files were listed. Done when the heartbeat block holds all five values and the row count equals the task count reported in step 4.
6. **Flag the repeatedly failing tasks.** Append `已連續失敗 N 次` to every row whose `fail_count` is above 0, with `N` taken verbatim from the file. A broken entry that retries every round with nobody noticing is the reason this field exists, so let no such row leave the table unmarked. Done when every row with `fail_count` above 0 carries the marker and its number matches the file.
7. **Finish successfully.** `助理未運行`, an absent tasks directory and an empty tasks directory are normal results — never exit non-zero for any of them. Reserve a failure report for a condition none of the tables above covers, and state which path and which error produced it. Done when the report is printed and nothing under `$JSC_HOME` has been created, modified or deleted.
+2 -2
View File
@@ -15,8 +15,8 @@
| 監控頁 | 指向 `MONITOR_{HASH}` 的同 wiki 連結 | 少了連結就要人自己算雜湊才翻得到內容頁 |
| 主機 | 這台機器的主機名,與雜湊第一段相同 | 比對用的兩欄之一,決定要更新哪一列 |
| 帳號 | 助理執行時的登入帳號,與雜湊第二段相同 | 比對用的兩欄之一。同一台機器換帳號就是另一個巡檢對象 |
| 心跳 | 巡檢當下的心跳判定,判準只看 `ts` 距現在是否不到 300 秒 | 一眼看出這台機器的助理還在不在跑,不必逐頁翻 |
| 最後巡檢 | 該頁最新一節的時間戳 | 心跳新鮮而這一欄很舊,代表助理活著卻沒在巡 |
| 心跳 | 巡檢當下(本輪寫入前)的心跳判定,判準只看 `ts` 距現在有沒有超過門檻,預設 300 秒 | 一眼看出這台機器上一輪巡檢有沒有跑完,不必逐頁翻 |
| 最後巡檢 | 該頁最新一節的時間戳 | 心跳由巡檢寫,兩欄理當一致;差很多就代表有一輪寫了心跳卻沒寫頁,那是缺陷 |
| 待辦筆數 | 待辦簿現有筆數 | 心跳新鮮而筆數為 0,代表助理空轉,沒有東西可跑 |
| 連續失敗項 | 該頁最新一節裡 `fail_count` 大於 0 的筆數 | 待辦簿的項目失敗不會自動暫停,每輪都重試。這一欄讓壞掉的項目在目錄頁就現形 |
+10 -3
View File
@@ -7,12 +7,15 @@
```mermaid
flowchart LR
A[巡檢一輪] --> B[收攏六類結果]
A[巡檢一輪] --> B[收攏各類結果]
B --> C[附加一節,節標題帶時間戳]
C --> D[既有的節原樣保留]
D --> E[回頭更新 MONITOR_CONTENTS 自己那一列]
E --> F[最後才寫心跳]
```
心跳排在最後一步,不能提前。心跳新鮮的意思就是「上一輪跑到這一步了」:這一節沒寫上來,心跳就不寫,讓它自己過期。那是巡檢在空轉的唯一訊號。
## 本頁基本資料
建頁時寫一次,之後不再更動。
@@ -26,13 +29,14 @@ flowchart LR
## 巡檢 {yyyy-MM-dd HH:mm}
一輪巡檢就是這樣一節,最新的一節放在最下面。六個子節固定都寫;某個來源讀不到,就在那個子節寫明是哪個路徑讀不到,不要整節略過。
一輪巡檢就是這樣一節,最新的一節放在最下面。六個子節固定都寫;某個來源讀不到,就在那個子節寫明是哪個路徑讀不到,不要整節略過。還沒實作的子節也照寫,寫明「這一輪不做這一項」——空表格會被讀成「查過了,沒問題」。
| 項目 | 內容 |
| --- | --- |
| 巡檢時間 | {yyyy-MM-dd HH:mm} |
| 觸發方式 | {排程、事件、手動 三選一} |
| 本輪判定 | {正常、警示、異常 三選一} |
| 本輪項目 | {這一輪跑了哪幾項,成功幾項、失敗幾項} |
| 讀不到的來源 | {路徑清單,全部讀得到就寫「無」} |
### 心跳與閘門狀態
@@ -45,7 +49,9 @@ flowchart LR
| session | {工作階段代號} |
| pid | {數字}。只給要找行程的人參考,不參與判定 |
心跳的判準只看 `ts` 距現在是否不到 300 秒。不看 pid 存活:五支 CLI 與容器裡的行程互相看不到彼此的 pid。閘門的判定留在 hook,助理只維持心跳。
這一欄讀到的是**上一輪**巡檢寫的心跳:心跳由巡檢寫,本輪那一次要等這一節寫上來之後才寫。
心跳的判準只看 `ts` 距現在有沒有超過門檻,預設 300 秒。不看 pid 存活:五支 CLI 與容器裡的行程互相看不到彼此的 pid。閘門的判定留在 hook,助理只維持心跳。心跳新鮮代表上一輪巡檢跑完了,不代表那一輪四項都成功——那要看本節上面的「本輪判定」。
### 技能與呼叫鏈使用統計
@@ -112,3 +118,4 @@ flowchart LR
- 禁止整頁覆寫。覆寫等於把這台機器的巡檢軌跡刪掉。
- 讀不到舊內容就中止,不附加,也不寫入。
- 附加成功之後,才回頭更新 `MONITOR_CONTENTS` 自己那一列。
- 兩頁都寫成之後,才寫這一輪的心跳。任一頁沒寫成就不寫心跳,讓它過期。
+739
View File
@@ -0,0 +1,739 @@
#!/usr/bin/env sh
# patrol.sh — 助理巡檢一輪的收攏與收口(供 jsc-assist:assistant 的 patrol 操作呼叫)。
#
# 用法:
# patrol.sh collect [--out {目錄}] [--trigger {排程|事件|手動}]
# patrol.sh finish --round {輪次代號} [--out {目錄}] [--dry-run]
# patrol.sh abort --round {輪次代號} [--out {目錄}]
#
# collect 帶了 --out,finish 與 abort 就要帶同一個目錄,不然換不到本輪的用量快照。
#
# collect 讀四項來源、組出監控頁要附加的那一節、把鎖拿在手上。
# finish 在監控頁寫成功之後才呼叫:寫心跳、換上用量快照、放掉鎖。
# abort 在監控頁沒寫成時呼叫:只放掉鎖,不寫心跳。
#
# 結束碼(三個子命令共用一張表,同一碼在不同子命令的成因寫在同一列):
# 0 collect:四項全部讀到底(含「來源在、沒有資料」);finish:心跳寫好、快照換上、
# 鎖放掉;abort:鎖放掉,本來就沒鎖也算
# 1 collect:部分成功——至少一項失敗,也至少一項有結果。**結果照樣印得出來,呼叫端
# 照樣要把這一節寫上監控頁**,只是本輪判定要標成警示
# 2 finish:找不到 jsc-hooks 的 hooks/heartbeat.sh,心跳沒有東西可寫。collect 不會回這
# 一碼——心跳讀不到只是 D-09 這一項失敗,另外三項照跑
# 3 collect:四項全部失敗,一項資料都沒有。這一節還是要寫上監控頁,本輪判定標成異常
# 4 上一輪還在跑,本輪讓開(collect),或鎖已經不在自己手上(finish、abort)。這不是
# 失敗,是刻意讓開:不寫心跳、不寫監控頁,下一輪再來
# 5 檔案系統失敗:鎖建不起來或放不掉、暫存檔寫不進去、快照換不上,或 heartbeat.sh write
# 回非 0。心跳沒寫成就是沒寫成,一律吵出來
# 6 用法錯誤:不認得的子命令、缺 --round、參數缺值
#
# --- 心跳為什麼由這裡寫,不由排程直接寫 ---
#
# 排程每分鐘直接呼叫 heartbeat.sh write 的話,心跳新鮮只證明 cron 活著。巡檢整個壞掉、
# 一項資料都讀不到、監控頁一頁都沒寫成,心跳照樣新鮮,靠心跳判定的閘門照樣放行,沒有
# 任何訊號。所以心跳改由巡檢寫:跑完一輪、而且結果真的記下來了,才寫那一次心跳。
#
# --- 心跳寫不寫,只看結果有沒有記下來 ---
#
# 四項的成敗不決定心跳。四項全失敗但監控頁寫成了,那一輪還是跑完了,證據也留下來了,
# 心跳照寫,頁上判定是異常,看頁的人自己判斷。反過來,監控頁沒寫成就是這一輪沒有結果,
# 心跳一定不寫:讓它自己過期,就是「巡檢在空轉」的唯一訊號。
# 所以寫心跳一定是獨立的 finish,時序上排在監控頁寫成之後,不與 collect 綁在一起。
#
# --- 上一輪還沒跑完,下一輪被叫起來 ---
#
# 一輪巡檢包含一次 CLI 呼叫與兩次 wiki 寫入,跑過一個排程週期是有可能的。所以整輪拿一把
# 鎖:$JSC_HOME/assistant/patrol.lock 是目錄,mkdir 是原子操作,搶不到就是別人在跑。
# 搶不到的那一輪回 4 直接讓開,不排隊、不並行——並行的兩輪會在同一頁附加兩節,還會互相
# 蓋掉用量快照。
# 鎖會逾時自動搶回來,門檻取心跳門檻(heartbeat.sh report 的 ttl 欄):上一輪跑得比門檻
# 還久,它本來就已經維持不住心跳新鮮了,讓新的一輪接手才對。被搶回來的那一輪,finish 會
# 拿 --round 比對出鎖不是自己的,回 4 且不寫心跳。
# 搶回來這件事會記在監控頁上(lock_broken=1),不會安靜發生。
#
# --- 這四項都是純讀取 ---
#
# D-01 技能與呼叫鏈使用統計 jsc-log 的 tools/usage-stats.sh
# D-04 版本落差與重啟閘門 jsc-hooks 的 version-guard.sh report、restart-gate.sh report
# D-07 SDLC 階段鎖與工作包鎖 $JSC_HOME/sessions/*.stage、$JSC_HOME/wp/*.pr
# D-09 心跳與閘門狀態自述 jsc-hooks 的 heartbeat.sh report
# 四項各自獨立:一項的來源不見了、或回非 0,只讓那一項標成失敗,其餘三項照跑、照記。
# 四項都不呼叫別的技能、不寫程式碼存取庫、不做決策。
#
# --- version-guard.sh report 的既有缺陷照實記 ---
#
# 這支腳本在部分機器上對每一個 domain 都回「查詢失敗」,recommend 跟著回 unverifiable。
# 巡檢照抄第四欄原字,不自己補查遠端版本、不把「查詢失敗」寫成「相符」或「最新」。
# 查不到就是沒有證據,寫成別的字等於幫一個既有缺陷蓋章。
#
# --- collect 的輸出 ---
#
# stdout 是 key=value,一行一個鍵,供呼叫端逐行取值。監控頁要用的 markdown 不印在
# stdout,改寫成檔案再把路徑印出來:那一段有好幾百字,經過對話重打一次只會多錯字。
# round= 本輪代號,finish 與 abort 要原樣帶回
# lock= acquired
# lock_broken= 0 或 1。1 代表上一輪的鎖逾時被搶回來
# hash= 監控頁雜湊,來源是 {主機名}/{登入帳號};算不出來時為空
# page= MONITOR_{HASH};hash 為空時為空
# host= user= at= 主機名、登入帳號、本輪時間
# item= 一項一行,欄位 status(ok、empty、fail)、rc、note
# verdict= 正常、警示、異常
# failed_sources= 讀不到的來源路徑,以「、」分隔;全部讀得到就是「無」
# tasks_total= tasks_failing= 待辦簿筆數與連續失敗筆數,只供目錄頁那一列用
# section_file= 要附加到 MONITOR_{HASH} 的那一節
# newpage_file= MONITOR_{HASH} 不存在時要建的整頁內容
# contents_file= MONITOR_CONTENTS 那一列的欄位值
#
# 環境變數:
# JSC_HOME 助理狀態檔的根目錄,預設 ~/.jsc
# JSC_ASSIST_HEARTBEAT_SH heartbeat.sh 路徑覆寫;找不到並排存取庫時才設
# JSC_ASSIST_VERSION_GUARD_SH version-guard.sh 路徑覆寫
# JSC_ASSIST_RESTART_GATE_SH restart-gate.sh 路徑覆寫
# JSC_ASSIST_USAGE_STATS_SH usage-stats.sh 路徑覆寫
# JSC_ASSIST_HASH_ID hash-id 路徑覆寫
# JSC_ASSIST_PATROL_LOCK_TTL 鎖的逾時秒數;未設定時取心跳門檻,取不到就用 300
set -u
JSC_HOME="${JSC_HOME:-$HOME/.jsc}"
STATE_DIR="$JSC_HOME/assistant"
LOCK="$STATE_DIR/patrol.lock"
PREV_SNAP="$STATE_DIR/usage-prev.tsv"
RD="$STATE_DIR/patrol"
SCRIPT_DIR=$(CDPATH= cd -- "$(dirname -- "$0")" 2>/dev/null && pwd)
SCRIPT_DIR="${SCRIPT_DIR:-.}"
TRIGGER=''
ROUND=''
DRYRUN=0
LOCK_BROKEN=0
FAILED_SOURCES=''
OK_COUNT=0
FAIL_COUNT=0
WARN=0
usage() {
cat >&2 <<'EOF'
usage: patrol.sh collect [--out 目錄] [--trigger 排程|事件|手動]
patrol.sh finish --round 輪次代號 [--out 目錄] [--dry-run]
patrol.sh abort --round 輪次代號 [--out 目錄]
EOF
exit 6
}
die() { # $1=結束碼 $2=訊息
printf '[jsc][助理巡檢][ERR]:%s\n' "$2" >&2
exit "$1"
}
# 找一支別的 domain 的腳本。搜尋順序比照 schedule.sh 的 heartbeat_sh():先環境變數覆寫,
# 再開發用的並排存取庫版面,最後已安裝的 plugin 快取版面。
find_tool() { # $1=domain 短名 $2=domain 內相對路徑 $3=環境變數覆寫值(可為空)
if [ -n "$3" ]; then
[ -f "$3" ] && { printf '%s\n' "$3"; return 0; }
return 1
fi
_root="${CLAUDE_PLUGIN_ROOT:-$SCRIPT_DIR/..}"
for _c in "$_root/../$1/$2" "$_root/../jsc-$1/$2"; do
[ -f "$_c" ] && { printf '%s\n' "$_c"; return 0; }
done
_c=$(ls -d "$_root"/../../jsc-"$1"/*/"$2" \
"$_root"/../../"$1"/*/"$2" \
"$HOME"/.claude/plugins/cache/*/jsc-"$1"/*/"$2" 2>/dev/null | sort | tail -n1)
[ -n "$_c" ] && [ -f "$_c" ] && { printf '%s\n' "$_c"; return 0; }
return 1
}
fmt_ts() { # $1=epoch 秒數;轉成當地時間字串,轉不動就原樣印
date -d "@$1" '+%Y-%m-%d %H:%M' 2>/dev/null && return 0
date -r "$1" '+%Y-%m-%d %H:%M' 2>/dev/null && return 0
printf '%s\n' "$1"
}
mtime_of() { # $1=檔案;印出修改時間,取不到印「-」
_t=$(stat -c %Y "$1" 2>/dev/null) || _t=''
[ -n "$_t" ] || _t=$(stat -f %m "$1" 2>/dev/null) || _t=''
[ -n "$_t" ] || { printf '-'; return 0; }
fmt_ts "$_t" | tr -d '\n'
}
# markdown 表格欄位裡的 `|` 會把欄切開,一律跳脫;換行壓成空白。
cell() { printf '%s' "$1" | tr '\n' ' ' | sed 's/|/\\|/g'; }
# 記一個讀不到的來源。同一輪多項失敗就串起來,供監控頁「讀不到的來源」那一列用。
add_failed_source() { # $1=路徑或來源名稱
if [ -z "$FAILED_SOURCES" ]; then FAILED_SOURCES="$1"; else FAILED_SOURCES="$FAILED_SOURCES、$1"; fi
}
# 記一列待人處理。助理只提醒,不代為執行。
add_pending() { # $1=要處理什麼 $2=來源子節 $3=建議入口
printf '| %s | %s | %s |\n' "$(cell "$1")" "$(cell "$2")" "$(cell "$3")" >>"$RD/pend.md"
}
# --- 鎖 ---
lock_ttl() { # 鎖的逾時秒數
_t="${JSC_ASSIST_PATROL_LOCK_TTL:-}"
case "$_t" in ''|*[!0-9]*) _t='' ;; esac
[ -n "$_t" ] && { printf '%s' "$_t"; return 0; }
[ -n "${HEARTBEAT_TTL:-}" ] && { printf '%s' "$HEARTBEAT_TTL"; return 0; }
printf '300'
}
lock_round() { sed -n 's/^round=//p' "$LOCK/info" 2>/dev/null | head -n1; }
lock_age() {
_s=$(sed -n 's/^started=//p' "$LOCK/info" 2>/dev/null | head -n1)
case "$_s" in ''|*[!0-9]*) printf '999999'; return 0 ;; esac
printf '%s' "$(( $(date +%s) - _s ))"
}
# 搶鎖。搶到回 0,別人在跑回 4,建不起來回 5。
lock_acquire() {
mkdir -p "$STATE_DIR" 2>/dev/null || die 5 "建不出助理狀態目錄 $STATE_DIR。"
if mkdir "$LOCK" 2>/dev/null; then
printf 'round=%s\npid=%s\nstarted=%s\n' "$ROUND" "$$" "$(date +%s)" >"$LOCK/info" 2>/dev/null \
|| die 5 "鎖建起來了,卻寫不進 $LOCK/info。"
return 0
fi
_age=$(lock_age); _ttl=$(lock_ttl)
if [ "$_age" -lt "$_ttl" ]; then
printf 'lock=busy holder=%s age=%s ttl=%s\n' "$(lock_round)" "$_age" "$_ttl"
die 4 "上一輪巡檢還在跑(輪次 $(lock_round),已經跑了 $_age 秒,未達 $_ttl 秒門檻),本輪讓開。"
fi
# 逾時搶回來。上一輪跑得比心跳門檻還久,它已經維持不住心跳新鮮了,讓新的一輪接手。
rm -rf "$LOCK" 2>/dev/null
mkdir "$LOCK" 2>/dev/null || die 5 "上一輪的鎖逾時,卻搶不回來:$LOCK。"
printf 'round=%s\npid=%s\nstarted=%s\n' "$ROUND" "$$" "$(date +%s)" >"$LOCK/info" 2>/dev/null \
|| die 5 "鎖搶回來了,卻寫不進 $LOCK/info。"
LOCK_BROKEN=1
WARN=1
return 0
}
# 放鎖。鎖不在自己手上就回 4,絕不硬放——那會把正在跑的那一輪的鎖拆掉。
lock_release() {
[ -d "$LOCK" ] || { printf 'lock=absent\n'; return 0; }
_h=$(lock_round)
[ "$_h" = "$ROUND" ] || die 4 "鎖不在本輪手上(鎖的輪次是 ${_h:-空值},本輪是 $ROUND),不動它,也不寫心跳。"
rm -rf "$LOCK" 2>/dev/null || die 5 "鎖放不掉:$LOCK。"
printf 'lock=released\n'
return 0
}
# --- D-01 技能與呼叫鏈使用統計 ---
# 把 usage-stats.sh 的「次數<TAB>名稱」轉成監控頁的列,順便算出本輪增量。
# 增量要有上一輪的累計快照才算得出來;沒有快照的第一輪一律寫「-」,不拿累計冒充本輪。
usage_rows() { # $1=類別(技能、呼叫鏈) $2=快照鍵(skill、chain) $3=資料檔
while IFS=' ' read -r _cnt _name; do
[ -n "$_name" ] || continue
printf '%s\t%s\t%s\n' "$2" "$_name" "$_cnt" >>"$RD/usage-next.tsv"
_prev=''
[ -f "$PREV_SNAP" ] && _prev=$(awk -F'\t' -v k="$2" -v n="$_name" \
'$1 == k && $2 == n { print $3; exit }' "$PREV_SNAP" 2>/dev/null)
if [ -n "$_prev" ]; then _delta=$(( _cnt - _prev )); else _delta='-'; fi
printf '| %s | %s | %s | %s |\n' "$(cell "$_name")" "$1" "$_delta" "$_cnt"
done <"$3"
}
d01() {
D01_STATUS=fail; D01_RC=0; D01_NOTE=''
: >"$RD/usage-next.tsv"
{
printf '### 技能與呼叫鏈使用統計\n\n'
printf '資料出自 `%s` 與 `%s`,由 `jsc-log` 的 `tools/usage-stats.sh` 聚合。\n\n' \
"\$JSC_HOME/usage/skills.jsonl" "\$JSC_HOME/usage/chains.jsonl"
} >"$RD/d01.md"
if ! _us=$(find_tool log tools/usage-stats.sh "${JSC_ASSIST_USAGE_STATS_SH:-}"); then
D01_NOTE='找不到 jsc-log 的 tools/usage-stats.sh'
D01_RC=127
add_failed_source 'jsc-log/tools/usage-stats.sh'
printf '**這一項失敗**:%s。這一輪沒有使用統計,不是「零次」。\n' "$D01_NOTE" >>"$RD/d01.md"
add_pending '技能用量讀不到,jsc-log 沒裝或版本太舊' '技能與呼叫鏈使用統計' '/jsc-cli:doctor'
return 0
fi
_rc1=0; _rc2=0
"$_us" skills >"$RD/d01.skills" 2>"$RD/d01.err" || _rc1=$?
"$_us" chains >"$RD/d01.chains" 2>>"$RD/d01.err" || _rc2=$?
if [ "$_rc1" -ne 0 ] || [ "$_rc2" -ne 0 ]; then
D01_RC=$(( _rc1 > _rc2 ? _rc1 : _rc2 ))
D01_NOTE="usage-stats.sh 回非 0(skills=$_rc1、chains=$_rc2)"
add_failed_source "$_us"
printf '**這一項失敗**:%s。訊息:%s\n' "$D01_NOTE" "$(cell "$(cat "$RD/d01.err" 2>/dev/null)")" >>"$RD/d01.md"
add_pending '使用統計腳本回非 0' '技能與呼叫鏈使用統計' '/jsc-log:stats'
return 0
fi
{
printf '| 對象 | 類別 | 本輪次數 | 累計次數 |\n'
printf '| --- | --- | ---: | ---: |\n'
usage_rows '技能' skill "$RD/d01.skills"
usage_rows '呼叫鏈' chain "$RD/d01.chains"
} >"$RD/d01.rows"
_n=$(grep -c '^| ' "$RD/d01.rows" 2>/dev/null); [ -n "$_n" ] || _n=0
if [ "$_n" -le 2 ]; then
D01_STATUS=empty
printf '來源檔讀得到,本輪一筆用量都沒有。hook 還沒記過任何一次呼叫就是這個狀態,那是「零次」,不是故障。\n' >>"$RD/d01.md"
return 0
fi
D01_STATUS=ok
cat "$RD/d01.rows" >>"$RD/d01.md"
if [ ! -f "$PREV_SNAP" ]; then
printf '\n第一輪沒有上一輪的累計快照,「本輪次數」一律寫「-」,不拿累計冒充本輪。\n' >>"$RD/d01.md"
fi
return 0
}
# --- D-04 版本落差與重啟閘門 ---
d04() {
D04_STATUS=fail; D04_RC=0; D04_NOTE=''
{
printf '### 版本落差與重啟閘門\n\n'
printf '資料出自 `version-guard.sh report` 與 `restart-gate.sh report`。\n\n'
} >"$RD/d04.md"
_vg=''; _rg=''
find_tool hooks hooks/version-guard.sh "${JSC_ASSIST_VERSION_GUARD_SH:-}" >"$RD/vg" 2>/dev/null \
&& _vg=$(cat "$RD/vg")
find_tool hooks hooks/restart-gate.sh "${JSC_ASSIST_RESTART_GATE_SH:-}" >"$RD/rg" 2>/dev/null \
&& _rg=$(cat "$RD/rg")
if [ -z "$_vg" ] && [ -z "$_rg" ]; then
D04_NOTE='找不到 jsc-hooks 的 version-guard.sh 與 restart-gate.sh'
D04_RC=127
add_failed_source 'jsc-hooks/hooks/version-guard.sh、jsc-hooks/hooks/restart-gate.sh'
printf '**這一項失敗**:%s。這一輪沒有版本證據,不是「版本都相符」。\n' "$D04_NOTE" >>"$RD/d04.md"
add_pending '版本與重啟閘門讀不到,jsc-hooks 沒裝或版本太舊' '版本落差與重啟閘門' '/jsc-cli:doctor'
return 0
fi
_part=0
# 版本比對表。第四欄照 version-guard.sh 原字抄,不改寫、不補查。
if [ -n "$_vg" ]; then
_rc=0
"$_vg" report >"$RD/d04.vg" 2>"$RD/d04.err" </dev/null || _rc=$?
if [ "$_rc" -ne 0 ]; then
_part=1; D04_RC="$_rc"
D04_NOTE="version-guard.sh report 回 $_rc"
add_failed_source "$_vg"
printf '**版本比對失敗**:`version-guard.sh report` 回 %s。訊息:%s\n\n' \
"$_rc" "$(cell "$(cat "$RD/d04.err" 2>/dev/null)")" >>"$RD/d04.md"
else
{
printf '| domain | 本機版本 | 應有版本 | 判定 |\n'
printf '| --- | --- | --- | --- |\n'
} >>"$RD/d04.md"
_unver=0; _rows=0
while IFS=' ' read -r _c1 _c2 _c3 _c4; do
[ -n "$_c1" ] || continue
case "$_c1" in
behind) printf '| (合計) | - | - | 落後 %s 個 |\n' "$(cell "$_c2")" >>"$RD/d04.md"; continue ;;
noregistry)
printf '| (無註冊檔) | - | - | 查不到本機已裝的 plugin:%s |\n' "$(cell "$_c2")" >>"$RD/d04.md"
_unver=1; continue ;;
esac
_rows=$(( _rows + 1 ))
printf '| %s | %s | %s | %s |\n' "$(cell "$_c1")" "$(cell "$_c2")" "$(cell "$_c3")" "$(cell "$_c4")" >>"$RD/d04.md"
[ "$_c4" = '查詢失敗' ] && _unver=$(( _unver + 1 ))
done <"$RD/d04.vg"
[ "$_rows" -eq 0 ] && printf '| (無 domain) | - | - | 這台機器一個 jsc plugin 都沒查到 |\n' >>"$RD/d04.md"
printf '\n判定欄照 `version-guard.sh report` 第四欄原字抄。抄到「查詢失敗」就寫「查詢失敗」,不改寫成「相符」或「最新」,也不自己補查遠端版本——查不到是沒有證據,不是版本沒問題。\n\n' >>"$RD/d04.md"
if [ "$_unver" -gt 0 ]; then
WARN=1
printf '**本輪有 %s 列查不到遠端版本。** 同一支腳本在擋人那條路徑查得到遠端版本,report 這條查不到,這是既有缺陷,不是這台機器的網路問題。\n\n' "$_unver" >>"$RD/d04.md"
add_pending 'version-guard.sh report 查不到遠端版本,版本落差本輪無證據' '版本落差與重啟閘門' '/jsc-cli:doctor'
fi
fi
else
_part=1
add_failed_source 'jsc-hooks/hooks/version-guard.sh'
printf '**版本比對失敗**:找不到 `version-guard.sh`。\n\n' >>"$RD/d04.md"
fi
# 重啟閘門。一行一支還沒重啟的 CLI;一行都沒有就是都放下了。
if [ -n "$_rg" ]; then
_rc=0
"$_rg" report >"$RD/d04.rg" 2>>"$RD/d04.err" </dev/null || _rc=$?
if [ "$_rc" -ne 0 ]; then
_part=1; D04_RC="$_rc"
add_failed_source "$_rg"
printf '**重啟閘門讀不到**:`restart-gate.sh report` 回 %s。\n' "$_rc" >>"$RD/d04.md"
else
{
printf '| CLI | 重啟閘門 | 升起時間 |\n'
printf '| --- | --- | --- |\n'
} >>"$RD/d04.md"
_up=0
while read -r _cli _rest; do
[ -n "$_cli" ] || continue
_up=$(( _up + 1 ))
_at=$(printf '%s' "$_rest" | sed -n 's/.*at=\([^ ]*\).*/\1/p')
printf '| %s | 已升起 | %s |\n' "$(cell "$_cli")" "$(cell "${_at:--}")" >>"$RD/d04.md"
add_pending "$_cli 的重啟閘門還升著,那一支要重新啟動" '版本落差與重啟閘門' '重開該支 CLI'
done <"$RD/d04.rg"
if [ "$_up" -eq 0 ]; then
printf '| (無) | 未升起 | - |\n' >>"$RD/d04.md"
else
WARN=1
fi
fi
else
_part=1
add_failed_source 'jsc-hooks/hooks/restart-gate.sh'
printf '**重啟閘門讀不到**:找不到 `restart-gate.sh`。\n' >>"$RD/d04.md"
fi
if [ "$_part" -eq 1 ]; then
D04_STATUS=fail
[ -n "$D04_NOTE" ] || D04_NOTE='兩份來源有一份讀不到'
else
D04_STATUS=ok
fi
return 0
}
# --- D-07 SDLC 階段鎖與工作包鎖 ---
d07() {
D07_STATUS=fail; D07_RC=0; D07_NOTE=''
_sess="$JSC_HOME/sessions"; _wp="$JSC_HOME/wp"
{
printf '### SDLC 階段鎖與工作包鎖現況\n\n'
printf '資料出自 `%s` 與 `%s`。只讀狀態,不做判定。\n\n' \
"\$JSC_HOME/sessions/{sid}.stage" "\$JSC_HOME/wp/*.pr"
} >"$RD/d07.md"
_bad=0
if [ -d "$_sess" ] && [ ! -r "$_sess" ]; then _bad=1; add_failed_source "$_sess"; fi
if [ -d "$_wp" ] && [ ! -r "$_wp" ]; then _bad=1; add_failed_source "$_wp"; fi
if [ "$_bad" -eq 1 ]; then
D07_NOTE='狀態目錄讀不到(權限)'
D07_RC=13
printf '**這一項失敗**:%s。這一輪沒有階段鎖與工作包鎖的證據,不是「沒有鎖」。\n' "$D07_NOTE" >>"$RD/d07.md"
add_pending 'SDLC 狀態目錄讀不到,權限要修' 'SDLC 階段鎖與工作包鎖現況' '/jsc-cli:setup'
return 0
fi
{
printf '| 工作階段 | 階段 | 必要標籤 | 上鎖時的模型 | 登記時間 |\n'
printf '| --- | --- | --- | --- | --- |\n'
} >>"$RD/d07.md"
_sn=0
for _f in "$_sess"/*.stage; do
[ -f "$_f" ] && [ -r "$_f" ] || continue
_sn=$(( _sn + 1 ))
_sid=$(basename "$_f" .stage)
_stage=$(cut -f1 "$_f" 2>/dev/null | head -n1)
_req=$(cut -f2 "$_f" 2>/dev/null | head -n1)
_mdl=$(cut -f3 "$_f" 2>/dev/null | head -n1)
printf '| %s | %s | %s | %s | %s |\n' \
"$(cell "$_sid")" "$(cell "${_stage:--}")" "$(cell "${_req:--}")" \
"$(cell "${_mdl:--}")" "$(mtime_of "$_f")" >>"$RD/d07.md"
done
[ "$_sn" -eq 0 ] && printf '| (無) | - | - | - | - |\n' >>"$RD/d07.md"
{
printf '\n| 存取庫 | 工作包 | PR | 上鎖時間 |\n'
printf '| --- | --- | --- | --- |\n'
} >>"$RD/d07.md"
_wn=0
for _f in "$_wp"/*.pr; do
[ -f "$_f" ] && [ -r "$_f" ] || continue
_wn=$(( _wn + 1 ))
_repo=$(sed -n 's/^repo=//p' "$_f" 2>/dev/null | head -n1)
_idx=$(sed -n 's/^index=//p' "$_f" 2>/dev/null | head -n1)
_wpn=$(sed -n 's/^wp=//p' "$_f" 2>/dev/null | head -n1)
_lk=$(sed -n 's/^locked=//p' "$_f" 2>/dev/null | head -n1)
printf '| %s | %s | %s | %s |\n' \
"$(cell "${_repo:--}")" "$(cell "${_wpn:--}")" \
"$(cell "${_repo:-?}#${_idx:-?}")" "$(cell "${_lk:--}")" >>"$RD/d07.md"
done
[ "$_wn" -eq 0 ] && printf '| (無) | - | - | - |\n' >>"$RD/d07.md"
printf '\n`.stage` 沒有存取庫欄位,`.pr` 沒有歸屬工作階段欄位,兩欄一律留「-」,不從分支名或目錄名回推——那兩個都會被改。登記時間取檔案的修改時間。\n' >>"$RD/d07.md"
if [ "$_sn" -eq 0 ] && [ "$_wn" -eq 0 ]; then D07_STATUS=empty; else D07_STATUS=ok; fi
return 0
}
# --- D-09 心跳與閘門狀態自述 ---
d09() {
D09_STATUS=fail; D09_RC=0; D09_NOTE=''
HEARTBEAT_STATE='讀不到'
{
printf '### 心跳與閘門狀態\n\n'
printf '資料出自 `heartbeat.sh report`。這一欄讀到的是**上一輪巡檢**寫的心跳:心跳改由巡檢寫,本輪那一次要等監控頁寫成之後才寫。\n\n'
} >"$RD/d09.md"
if ! _hb=$(find_tool hooks hooks/heartbeat.sh "${JSC_ASSIST_HEARTBEAT_SH:-}"); then
D09_NOTE='找不到 jsc-hooks 的 hooks/heartbeat.sh'
D09_RC=127
add_failed_source 'jsc-hooks/hooks/heartbeat.sh'
printf '**這一項失敗**:%s。心跳判不出來,不代表助理沒在跑,也不代表在跑。\n' "$D09_NOTE" >>"$RD/d09.md"
add_pending '心跳腳本找不到,jsc-hooks 沒裝或版本低於 0.3.7' '心跳與閘門狀態' '/jsc-cli:deploy'
return 0
fi
_rc=0
"$_hb" report >"$RD/d09.line" 2>"$RD/d09.err" </dev/null || _rc=$?
if [ "$_rc" -ne 0 ]; then
D09_RC="$_rc"
D09_NOTE="heartbeat.sh report 回 $_rc"
add_failed_source "$_hb"
printf '**這一項失敗**:%s。訊息:%s\n' "$D09_NOTE" "$(cell "$(cat "$RD/d09.err" 2>/dev/null)")" >>"$RD/d09.md"
add_pending '心跳腳本回非 0,心跳判不出來' '心跳與閘門狀態' '/jsc-assist:assistant status'
return 0
fi
_line=$(cat "$RD/d09.line" 2>/dev/null)
# 欄位以空白分隔,路徑擺最後。拆成一行一欄再取,不用貪婪比對——貪婪會抓到後面同名的鍵。
_get() { printf '%s' "$_line" | tr ' ' '\n' | sed -n "s/^$1=//p" | head -n1; }
_st=$(printf '%s' "$_line" | sed -n 's/^state=\([^ ]*\).*/\1/p')
_ts=$(_get ts); _age=$(_get age); _ttl=$(_get ttl)
_pid=$(_get pid); _cli=$(_get cli); _sid=$(_get session)
_file=$(printf '%s' "$_line" | sed -n 's/.*file=//p')
HEARTBEAT_TTL="$_ttl"
case "$_st" in
fresh) HEARTBEAT_STATE='新鮮' ;;
stale) HEARTBEAT_STATE='過期'; WARN=1 ;;
invalid) HEARTBEAT_STATE='心跳檔損壞'; WARN=1 ;;
absent) HEARTBEAT_STATE='不存在'; WARN=1 ;;
*) HEARTBEAT_STATE="判不出(state=${_st:-空值})"; WARN=1 ;;
esac
{
printf '| 項目 | 內容 |\n'
printf '| --- | --- |\n'
printf '| 心跳 | %s |\n' "$(cell "$HEARTBEAT_STATE")"
if [ -n "$_ts" ]; then
printf '| 上次心跳 | %s,距這次巡檢 %s 秒 |\n' "$(fmt_ts "$_ts" | tr -d '\n')" "$(cell "${_age:--}")"
else
printf '| 上次心跳 | - |\n'
fi
printf '| 過期門檻 | %s 秒 |\n' "$(cell "${_ttl:--}")"
printf '| cli | %s |\n' "$(cell "${_cli:--}")"
printf '| session | %s |\n' "$(cell "${_sid:--}")"
printf '| pid | %s。只給要找行程的人參考,不參與判定 |\n' "$(cell "${_pid:--}")"
printf '| 心跳檔 | `%s` |\n' "$(cell "${_file:--}")"
printf '\n心跳的判準只看 `ts` 距現在有沒有超過門檻,不看 pid 存活:五支 CLI 與容器裡的行程互相看不到彼此的 pid。閘門的判定留在 hook,助理只維持心跳,不參與判定。\n'
printf '\n心跳新鮮代表上一輪巡檢跑完了,而且結果記上監控頁了。它不代表那一輪四項都成功——四項的成敗看本節上面的「本輪判定」。\n'
} >>"$RD/d09.md"
case "$_st" in
fresh|stale|invalid|absent) D09_STATUS=ok ;;
*) D09_STATUS=fail; D09_NOTE='report 印不出認得的 state' ;;
esac
[ "$_st" = absent ] && D09_STATUS=empty
return 0
}
# --- 待辦簿筆數(只供目錄頁那一列用)---
count_tasks() {
TASKS_TOTAL=0; TASKS_FAILING=0
_d="$STATE_DIR/tasks"
[ -d "$_d" ] && [ -r "$_d" ] || return 0
for _f in "$_d"/*; do
[ -f "$_f" ] || continue
TASKS_TOTAL=$(( TASKS_TOTAL + 1 ))
_fc=$(sed -n 's/^fail_count=//p' "$_f" 2>/dev/null | head -n1)
case "$_fc" in ''|*[!0-9]*) _fc=0 ;; esac
[ "$_fc" -gt 0 ] && TASKS_FAILING=$(( TASKS_FAILING + 1 ))
done
return 0
}
# --- 組出監控頁那一節 ---
tally() { # $1=項目狀態
case "$1" in
fail) FAIL_COUNT=$(( FAIL_COUNT + 1 )) ;;
*) OK_COUNT=$(( OK_COUNT + 1 )) ;;
esac
}
compose() {
_v='正常'
[ "$WARN" -eq 1 ] && _v='警示'
[ "$FAIL_COUNT" -gt 0 ] && _v='警示'
[ "$OK_COUNT" -eq 0 ] && _v='異常'
VERDICT="$_v"
{
printf '## 巡檢 %s\n\n' "$AT"
printf '| 項目 | 內容 |\n'
printf '| --- | --- |\n'
printf '| 巡檢時間 | %s |\n' "$AT"
printf '| 觸發方式 | %s |\n' "$TRIGGER"
printf '| 本輪判定 | %s |\n' "$VERDICT"
printf '| 本輪項目 | 四項:D-01 使用統計、D-04 版本與重啟閘門、D-07 階段鎖與工作包鎖、D-09 心跳自述。成功 %s 項、失敗 %s 項 |\n' "$OK_COUNT" "$FAIL_COUNT"
printf '| 讀不到的來源 | %s |\n' "$(cell "${FAILED_SOURCES:-無}")"
if [ "$LOCK_BROKEN" -eq 1 ]; then
printf '| 鎖 | 上一輪的鎖逾時,本輪搶回來了。上一輪沒跑完,那一輪不會寫心跳 |\n'
fi
printf '\n'
cat "$RD/d09.md"; printf '\n'
cat "$RD/d01.md"; printf '\n'
printf '### hook 執行期錯誤\n\n'
printf '**這一輪不做這一項。** D-02 hook 錯誤巡檢還沒實作,這一節沒有資料不代表沒有 hook 錯誤。要現在查就跑 `/jsc-hooks:hooks-install` 的錯誤掃描,或直接跑 `jsc-hooks` 的 `tools/scan-hook-errors.sh`。\n\n'
cat "$RD/d04.md"; printf '\n'
cat "$RD/d07.md"; printf '\n'
printf '### 待辦簿到期與逾期\n\n'
printf '**這一輪只數筆數,不逐筆判到期。** 本輪待辦簿共 %s 筆,其中 %s 筆 `fail_count` 大於 0。逐筆的到期與逾期判定還沒實作,要看逐筆內容就跑 `/jsc-assist:assistant status`。\n\n' \
"$TASKS_TOTAL" "$TASKS_FAILING"
printf '### 待人處理\n\n'
printf '助理只提醒,不代為執行。這一節列的是本輪要人接手的項目。\n\n'
printf '| 項目 | 來源子節 | 建議入口 |\n'
printf '| --- | --- | --- |\n'
if [ -s "$RD/pend.md" ]; then cat "$RD/pend.md"; else printf '| (無) | - | - |\n'; fi
} >"$RD/section.md"
# 監控頁不存在時要建的整頁內容。基本資料建頁時寫一次,之後不再更動。
{
printf '# 助理巡檢 — %s/%s\n\n' "$HOST" "$USER_NAME"
printf '> 由 `jsc-assist` 維護。這是監控頁 `%s`。\n' "$PAGE"
printf '> 這頁是這台機器的巡檢軌跡:一次巡檢附加一節,節標題帶時間戳,舊的節一個字都不動。\n'
printf '> 附加是刻意的。助理的寫入是背景行為,覆寫錯了沒人在現場,軌跡被抹掉也看不出斷在哪一輪。\n'
printf '> 目錄頁 `MONITOR_CONTENTS` 只更新自己那一列,寫入語意與這頁不同,不要混用。\n\n'
printf '```mermaid\nflowchart LR\n'
printf ' A[巡檢一輪] --> B[收攏四項結果]\n'
printf ' B --> C[附加一節,節標題帶時間戳]\n'
printf ' C --> D[既有的節原樣保留]\n'
printf ' D --> E[回頭更新 MONITOR_CONTENTS 自己那一列]\n'
printf ' E --> F[最後才寫心跳]\n'
printf '```\n\n'
printf '## 本頁基本資料\n\n'
printf '建頁時寫一次,之後不再更動。\n\n'
printf '| 項目 | 內容 |\n'
printf '| --- | --- |\n'
printf '| 主機 | %s |\n' "$(cell "$HOST")"
printf '| 帳號 | %s |\n' "$(cell "$USER_NAME")"
printf '| 雜湊來源 | `%s/%s` |\n' "$(cell "$HOST")" "$(cell "$USER_NAME")"
printf '| 狀態檔根目錄 | `$JSC_HOME/assistant/`(`$JSC_HOME` 未設定就退回 `~/.jsc`) |\n\n'
cat "$RD/section.md"
} >"$RD/newpage.md"
{
printf 'page=%s\n' "$PAGE"
printf 'host=%s\n' "$HOST"
printf 'user=%s\n' "$USER_NAME"
printf 'heartbeat=%s\n' "$HEARTBEAT_STATE"
printf 'last_patrol=%s\n' "$AT"
printf 'tasks_total=%s\n' "$TASKS_TOTAL"
printf 'tasks_failing=%s\n' "$TASKS_FAILING"
printf 'row=| [[%s]] | %s | %s | %s | %s | %s | %s |\n' \
"$PAGE" "$HOST" "$USER_NAME" "$HEARTBEAT_STATE" "$AT" "$TASKS_TOTAL" "$TASKS_FAILING"
} >"$RD/contents.tsv"
return 0
}
# --- 參數解析 ---
CMD="${1:-}"
[ -n "$CMD" ] || usage
shift
case "$CMD" in collect|finish|abort) ;; *) usage ;; esac
OUT=''
while [ "$#" -gt 0 ]; do
case "$1" in
--out) [ "$#" -ge 2 ] || usage; OUT="$2"; shift 2 ;;
--trigger) [ "$#" -ge 2 ] || usage; TRIGGER="$2"; shift 2 ;;
--round) [ "$#" -ge 2 ] || usage; ROUND="$2"; shift 2 ;;
--dry-run) DRYRUN=1; shift ;;
*) usage ;;
esac
done
[ -n "$OUT" ] && RD="$OUT"
case "$CMD" in
collect)
[ -n "$TRIGGER" ] || { if [ "${JSC_CLI:-}" = cron ]; then TRIGGER='排程'; else TRIGGER='手動'; fi; }
case "$TRIGGER" in 排程|事件|手動) ;; *) usage ;; esac
ROUND="$(date +%s)-$$"
HOST=$(hostname 2>/dev/null || uname -n 2>/dev/null || printf 'unknown')
USER_NAME="${USER:-$(id -un 2>/dev/null || printf 'unknown')}"
AT=$(date '+%Y-%m-%d %H:%M')
# 先讀心跳。門檻要先拿到,鎖的逾時才有依據。
mkdir -p "$RD" 2>/dev/null || die 5 "建不出巡檢暫存目錄 $RD。"
: >"$RD/pend.md"
HEARTBEAT_TTL=''
d09
lock_acquire
d01
d04
d07
count_tasks
tally "$D01_STATUS"; tally "$D04_STATUS"; tally "$D07_STATUS"; tally "$D09_STATUS"
HASH=''
if _hi=$(find_tool gitea tools/hash-id "${JSC_ASSIST_HASH_ID:-}"); then
HASH=$("$_hi" "$HOST/$USER_NAME" 2>/dev/null) || HASH=''
fi
if [ -n "$HASH" ]; then PAGE="MONITOR_$HASH"; else PAGE=''; fi
compose
printf 'round=%s\n' "$ROUND"
printf 'lock=acquired\n'
printf 'lock_broken=%s\n' "$LOCK_BROKEN"
printf 'hash=%s\n' "$HASH"
printf 'page=%s\n' "$PAGE"
printf 'host=%s\n' "$HOST"
printf 'user=%s\n' "$USER_NAME"
printf 'at=%s\n' "$AT"
printf 'item=D-01 status=%s rc=%s note=%s\n' "$D01_STATUS" "$D01_RC" "$D01_NOTE"
printf 'item=D-04 status=%s rc=%s note=%s\n' "$D04_STATUS" "$D04_RC" "$D04_NOTE"
printf 'item=D-07 status=%s rc=%s note=%s\n' "$D07_STATUS" "$D07_RC" "$D07_NOTE"
printf 'item=D-09 status=%s rc=%s note=%s\n' "$D09_STATUS" "$D09_RC" "$D09_NOTE"
printf 'verdict=%s\n' "$VERDICT"
printf 'failed_sources=%s\n' "${FAILED_SOURCES:-無}"
printf 'tasks_total=%s\n' "$TASKS_TOTAL"
printf 'tasks_failing=%s\n' "$TASKS_FAILING"
printf 'section_file=%s\n' "$RD/section.md"
printf 'newpage_file=%s\n' "$RD/newpage.md"
printf 'contents_file=%s\n' "$RD/contents.tsv"
[ "$OK_COUNT" -eq 0 ] && exit 3
[ "$FAIL_COUNT" -gt 0 ] && exit 1
exit 0 ;;
finish)
[ -n "$ROUND" ] || usage
[ -d "$LOCK" ] || die 4 "本輪的鎖已經不在($LOCK),不寫心跳。鎖多半是逾時被下一輪搶走了。"
_h=$(lock_round)
[ "$_h" = "$ROUND" ] \
|| die 4 "鎖不在本輪手上(鎖的輪次是 ${_h:-空值},本輪是 $ROUND),不寫心跳。上一輪跑太久被搶走了,這一輪的結果不算數。"
_hb=$(find_tool hooks hooks/heartbeat.sh "${JSC_ASSIST_HEARTBEAT_SH:-}") \
|| die 2 '找不到 jsc-hooks 的 hooks/heartbeat.sh,心跳沒有東西可寫。請先安裝 jsc-hooks 0.3.7 以上。'
if [ "$DRYRUN" -eq 1 ]; then
printf 'dryrun=1 heartbeat_cmd=%s write snapshot=%s -> %s lock=%s\n' \
"$_hb" "$RD/usage-next.tsv" "$PREV_SNAP" "$LOCK"
exit 0
fi
_rc=0
"$_hb" write >/dev/null 2>"$RD/finish.err" </dev/null || _rc=$?
[ "$_rc" -eq 0 ] \
|| die 5 "heartbeat.sh write 回 $_rc,心跳沒寫成:$(tr '\n' ' ' <"$RD/finish.err" 2>/dev/null)"
printf 'heartbeat=written\n'
# 快照要等這一輪真的收口才換上。半途失敗就換掉的話,下一輪的「本輪次數」會少算。
if [ -f "$RD/usage-next.tsv" ]; then
cp "$RD/usage-next.tsv" "$PREV_SNAP" 2>/dev/null || die 5 "用量快照換不上:$PREV_SNAP。"
printf 'snapshot=promoted\n'
else
printf 'snapshot=skipped\n'
fi
lock_release
exit 0 ;;
abort)
[ -n "$ROUND" ] || usage
printf 'heartbeat=not-written\n'
lock_release
exit 0 ;;
esac
+522
View File
@@ -0,0 +1,522 @@
#!/usr/bin/env sh
# schedule.sh — 助理系統排程的安裝、移除與查現況(供 jsc-assist:assistant 呼叫)。
#
# 用法:
# schedule.sh install [patrol|all] [--dry-run] [--cli {代號}] [--patrol-cmd {指令}] [--period {分鐘}]
# schedule.sh remove [heartbeat|patrol|all] [--dry-run]
# schedule.sh status [--dry-run]
#
# 工作代號省略時一律是 patrol。這一版只裝巡檢那一筆,心跳由巡檢跑完那一輪自己寫。
#
# 結束碼:
# 0 成功:install 條目寫進去也回讀得到、排程服務在跑;remove 移除完成,或本來就沒裝;
# status 印完現況(有沒有裝都算成功,看 installed 欄)
# 1 install 寫進去了,但排程服務沒在跑——條目不會被執行。WSL 預設不啟動 cron,這一碼
# 多半就是它。呼叫端一律照實講「排程裝了但不會執行」,不可以宣稱會定時執行
# 2 找不到 jsc-hooks 的 hooks/heartbeat.sh。心跳門檻查不到,週期算不出來,整支停下
# 3 這台機器沒有可用的排程機制:認不得作業系統,或 crontab 與 schtasks 都找不到
# 4 排程操作失敗:讀不到現有排程(且失敗原因不是「沒有排程」)、寫入或刪除回非 0
# 5 回讀驗證失敗:寫入回 0 但條目不在,或移除回 0 但條目還在,又或其他人的條目數量對不上
# 6 用法錯誤:不認得的子命令、不認得的工作代號、缺參數、判不出要用哪一支 CLI 跑巡檢,
# 或 --period 給的週期塞不進心跳的過期門檻
#
# --- 排程只叫巡檢,心跳由巡檢寫 ---
#
# 舊版排程每分鐘直接呼叫 heartbeat.sh write。那樣心跳新鮮只證明 cron 活著:巡檢整個壞掉、
# 一輪都沒跑成,心跳照樣新鮮,靠心跳判定的閘門照樣放行,沒有任何訊號。
# 現在排程只叫巡檢,巡檢把結果寫上監控頁之後才寫那一次心跳。心跳新鮮於是等於「上一輪巡檢
# 真的做完了,而且結果記下來了」。
# 所以 `install heartbeat` 直接回 6,不給裝:裝了就是有第二個寫心跳的人,心跳的意思立刻回到
# 舊版。舊機器上留著的那一筆 heartbeat 條目,`install patrol` 會順手清掉,並在輸出印
# legacy_removed=1——不清的話它每分鐘照樣寫,這次改動等於白做。
# `remove` 與 `status` 仍然認得 heartbeat 這個代號,就是為了清掉與看得到那一筆舊條目。
#
# --- 巡檢週期怎麼定 ---
#
# 心跳的更新頻率現在等於巡檢週期,所以週期一定要塞得進心跳的過期門檻,不然心跳永遠是過期。
# 門檻不寫死,改讀 `heartbeat.sh report` 的 `ttl` 欄——那是這台機器實際生效的值。
# 週期取「漏掉一輪還算新鮮、漏掉兩輪就過期」:2 × 週期 × 60 < 門檻,再往下取一個能整除一小時
# 的分鐘數,排程間隔才規律。門檻預設 300 秒時算出來是每 2 分鐘一輪。
# 要拉長巡檢週期就先把門檻調大(`JSC_ASSISTANT_HEARTBEAT_TTL`,單位秒),週期會跟著變長:
# 門檻 1800 秒算出每 12 分鐘一輪。週期與門檻兩個數字綁在一起算,不會再各走各的。
#
# --- 只動自己那一筆 ---
#
# 每一筆條目行尾都帶固定標記「# jsc-assist:assistant {工作代號}」,安裝與移除都靠它比對。
# 安裝先用 `grep -vF` 濾掉自己這幾個工作的舊條目,再把新條目追加上去,整份寫回;**絕不**
# 把 crontab 當空的重寫。`crontab -l` 在沒有任何排程時會回非 0,錯誤訊息才分得出是「沒有
# 排程」還是「權限不足」——後者當成空的寫回去,會把使用者整份排程刪光,所以這裡讀不懂
# 錯誤訊息就直接回 4,不猜。
# 寫回之後還會數行數:其他人的條目一行都不能少,少了就回 5。
#
# --- 排程服務沒在跑,等於沒裝 ---
#
# 條目寫進去不代表會被執行。WSL 預設不啟動 cron,要 `sudo service cron start`,而且重開
# WSL 之後要再啟動一次。安裝完一律檢查服務在不在跑,沒跑就回 1 並講清楚,不可以只回報
# 「排程已建立」。
#
# --- 排程是非互動環境 ---
#
# 條目一律接 `</dev/null`,跑的東西讀不到標準輸入,卡不住。heartbeat.sh 本來就不讀 stdin,
# 這一條是給巡檢那一筆用的:助理跳出權限詢問就會整輪卡死,等於排程沒跑。
#
# --- log 放哪裡 ---
#
# $JSC_HOME/assistant/schedule.log(JSC_HOME 未設定時為 ~/.jsc)。刻意放在專案外面:寫進
# 任何存取庫都會多出未追蹤檔,污染別人的變更盤點。
#
# 環境變數:
# JSC_HOME 助理狀態檔的根目錄,預設 ~/.jsc
# JSC_ASSIST_CRONTAB_CMD crontab 執行檔,預設 crontab。crontab 不在標準路徑,或要用
# 假的 crontab 驗濾除邏輯時才設
# JSC_ASSIST_PATROL_CMD 巡檢要跑的指令,優先於 --patrol-cmd 以外的所有推斷
# JSC_CLI 目前是哪一支 CLI,決定巡檢預設指令
# JSC_ASSISTANT_HEARTBEAT_TTL 心跳過期門檻,單位秒。巡檢週期由它算出來
set -u
MARK_PREFIX='# jsc-assist:assistant'
JSC_HOME="${JSC_HOME:-$HOME/.jsc}"
STATE_DIR="$JSC_HOME/assistant"
LOG="$STATE_DIR/schedule.log"
CRONTAB_CMD="${JSC_ASSIST_CRONTAB_CMD:-crontab}"
TASK_PREFIX='jsc-assist-assistant'
DRYRUN=0
CLI=''
PATROL_CMD="${JSC_ASSIST_PATROL_CMD:-}"
PERIOD=''
SCRIPT_DIR=$(CDPATH= cd -- "$(dirname -- "$0")" 2>/dev/null && pwd)
SCRIPT_DIR="${SCRIPT_DIR:-.}"
usage() {
cat >&2 <<'EOF'
usage: schedule.sh install [patrol|all] [--dry-run] [--cli 代號] [--patrol-cmd 指令] [--period 分鐘]
schedule.sh remove [heartbeat|patrol|all] [--dry-run]
schedule.sh status [--dry-run]
EOF
exit 6
}
die() { # $1=結束碼 $2=訊息
printf '[jsc][助理排程][ERR]:%s\n' "$2" >&2
exit "$1"
}
note() { printf '[jsc][助理排程]:%s\n' "$1" >&2; }
# 找出 jsc-hooks 的 hooks/heartbeat.sh 絕對路徑。搜尋順序比照 jsc-hooks lib.sh 的
# jsc_gitea_sh():先環境變數,再開發用的並排存取庫版面,最後已安裝的 plugin 快取版面。
heartbeat_sh() {
if [ -n "${JSC_HOOKS_HOOKS:-}" ] && [ -f "$JSC_HOOKS_HOOKS/heartbeat.sh" ]; then
printf '%s\n' "$JSC_HOOKS_HOOKS/heartbeat.sh"; return 0
fi
_root="${CLAUDE_PLUGIN_ROOT:-$SCRIPT_DIR/..}"
for _c in "$_root/../hooks/hooks/heartbeat.sh" "$_root/../jsc-hooks/hooks/heartbeat.sh"; do
[ -f "$_c" ] && { (CDPATH= cd -- "$(dirname -- "$_c")" && printf '%s/heartbeat.sh\n' "$(pwd)"); return 0; }
done
_c=$(ls -d "$_root"/../../jsc-hooks/*/hooks/heartbeat.sh \
"$_root"/../../hooks/*/hooks/heartbeat.sh \
"$HOME"/.claude/plugins/cache/*/jsc-hooks/*/hooks/heartbeat.sh 2>/dev/null \
| sort | tail -n1)
[ -n "$_c" ] && [ -f "$_c" ] && { printf '%s\n' "$_c"; return 0; }
_c=$(command -v heartbeat.sh 2>/dev/null || true)
[ -n "$_c" ] && { printf '%s\n' "$_c"; return 0; }
return 1
}
os_kind() {
case "$(uname -s 2>/dev/null)" in
Linux) printf 'linux' ;;
Darwin) printf 'darwin' ;;
MINGW*|MSYS*|CYGWIN*|Windows_NT) printf 'windows' ;;
*) printf 'unknown' ;;
esac
}
# 這台機器要用哪一種排程機制。Linux、WSL 與 macOS 都用 crontab;macOS 若要改用 launchd
# 請自行改寫,本腳本不代為產生 plist。
mechanism() {
case "$(os_kind)" in
windows) command -v schtasks >/dev/null 2>&1 && { printf 'schtasks'; return 0; } ;;
linux|darwin) command -v "$CRONTAB_CMD" >/dev/null 2>&1 && { printf 'crontab'; return 0; } ;;
esac
return 1
}
# 排程服務在不在跑。回 running、stopped 或 unknown。
service_state() {
case "$(os_kind)" in
darwin)
# macOS 的 cron 由 launchd 隨用隨起,查不到行程不代表沒在跑,一律回 unknown。
printf 'unknown'; return 0 ;;
windows)
printf 'unknown'; return 0 ;;
esac
if command -v pgrep >/dev/null 2>&1; then
if pgrep -x cron >/dev/null 2>&1 || pgrep -x crond >/dev/null 2>&1; then
printf 'running'
else
printf 'stopped'
fi
return 0
fi
if command -v ps >/dev/null 2>&1; then
if ps -e 2>/dev/null | grep -qE '[ /](cron|crond)$'; then printf 'running'; else printf 'stopped'; fi
return 0
fi
printf 'unknown'
}
marker_of() { printf '%s %s' "$MARK_PREFIX" "$1"; }
# 數行數。用 grep -c '' 不用 wc -l:最後一行沒有換行時 wc -l 會少數一行。
# grep 數到 0 會回非 0,數字照樣印得出來,所以只在完全沒有輸出時才補 0。
count_lines() { _n=$(grep -c '' "$1" 2>/dev/null); [ -n "$_n" ] || _n=0; printf '%s' "$_n"; }
# 數管線進來的行數,語意同 count_lines。
count_stdin() { _n=$(grep -c '' 2>/dev/null); [ -n "$_n" ] || _n=0; printf '%s' "$_n"; }
# 這台機器實際生效的心跳過期門檻。取自 heartbeat.sh report 的 ttl 欄,不自己重算:
# 門檻的唯一來源是那支腳本,兩邊各算一次就會漂移。讀不到就退回 300。
heartbeat_ttl() {
_t=$("$HEARTBEAT" report </dev/null 2>/dev/null | tr ' ' '\n' | sed -n 's/^ttl=//p' | head -n1)
case "$_t" in ''|*[!0-9]*) _t=300 ;; esac
[ "$_t" -gt 0 ] || _t=300
printf '%s' "$_t"
}
# 由門檻算出巡檢週期,單位分鐘。條件是 2 × 週期 × 60 < 門檻:漏掉一輪還算新鮮,漏掉兩輪
# 才過期。再往下取一個能整除一小時的分鐘數,`*/N` 的間隔才規律。
period_for_ttl() { # $1=門檻秒數
_max=$(( ($1 - 1) / 120 ))
[ "$_max" -lt 1 ] && _max=1
_p=1
for _d in 1 2 3 4 5 6 10 12 15 20 30 60; do
[ "$_d" -le "$_max" ] && _p="$_d"
done
printf '%s' "$_p"
}
spec_of() {
case "$1" in
heartbeat) printf '* * * * *' ;;
patrol) printf '*/%s * * * *' "$PERIOD" ;;
esac
}
# 巡檢要跑的指令。優先序:--patrol-cmd 或 JSC_ASSIST_PATROL_CMD > 依 CLI 代號推斷。
# 判不出 CLI 就回非 0,由主流程回 6,不猜——猜錯會每 15 分鐘跑一支不存在的執行檔。
patrol_command() {
[ -n "$PATROL_CMD" ] && { printf '%s' "$PATROL_CMD"; return 0; }
_cli="$CLI"
[ -n "$_cli" ] || _cli="${JSC_CLI:-}"
[ -n "$_cli" ] || { [ -n "${CLAUDE_PLUGIN_ROOT:-}" ] && _cli=claude; }
case "$_cli" in
claude) printf 'claude -p "/jsc-assist:assistant 跑一輪巡檢"' ;;
codex) printf "codex exec '\$assistant 跑一輪巡檢'" ;;
copilot) printf 'copilot -p "跑一輪助理巡檢"' ;;
antigravity) printf 'agy -p "/jsc-assist:assistant 跑一輪巡檢"' ;;
kiro) printf 'kiro-cli -p "跑一輪助理巡檢"' ;;
*) return 1 ;;
esac
}
# 組出一筆 crontab 條目。`%` 在 crontab 是換行符號,一律跳脫。
cron_entry() { # $1=工作代號
_spec=$(spec_of "$1")
case "$1" in
heartbeat) _cmd="JSC_CLI=cron JSC_SESSION_ID=schedule '$HEARTBEAT' write" ;;
patrol) _cmd="JSC_CLI=cron JSC_SESSION_ID=schedule $PATROL_RESOLVED" ;;
esac
printf '%s %s </dev/null >>%s 2>&1 %s' \
"$_spec" "$_cmd" "'$LOG'" "$(marker_of "$1")" | sed 's/%/\\%/g'
}
task_name() { printf '%s-%s' "$TASK_PREFIX" "$1"; }
# 讀現有的 crontab 到 $1。沒有任何排程時 crontab -l 會回非 0,那算正常;讀不懂的錯誤
# 一律回 4,不當成空的——當成空的寫回去會把使用者整份排程刪光。
cron_read() { # $1=輸出檔
_err="$TMPD/err"
if "$CRONTAB_CMD" -l >"$1" 2>"$_err"; then return 0; fi
if [ ! -s "$_err" ] || grep -qiE 'no crontab|沒有 crontab' "$_err"; then
: >"$1"; return 0
fi
die 4 "讀不到現有排程,原因不是「沒有排程」:$(tr '\n' ' ' <"$_err")"
}
cron_lines_for() { # $1=crontab 檔 $2=工作代號;印出該工作的條目
grep -F "$(marker_of "$2")" "$1" 2>/dev/null || true
}
# --- 參數解析 ---
CMD="${1:-}"
[ -n "$CMD" ] || usage
shift
case "$CMD" in install|remove|status) ;; *) usage ;; esac
JOBS='patrol'
if [ "$#" -gt 0 ]; then
case "$1" in
heartbeat) JOBS='heartbeat'; shift ;;
patrol) JOBS='patrol'; shift ;;
all) JOBS='heartbeat patrol'; shift ;;
esac
fi
while [ "$#" -gt 0 ]; do
case "$1" in
--dry-run) DRYRUN=1; shift ;;
--cli) [ "$#" -ge 2 ] || usage; CLI="$2"; shift 2 ;;
--patrol-cmd) [ "$#" -ge 2 ] || usage; PATROL_CMD="$2"; shift 2 ;;
--period) [ "$#" -ge 2 ] || usage; PERIOD="$2"; shift 2 ;;
*) usage ;;
esac
done
# 心跳那一筆不給裝。理由見檔頭「排程只叫巡檢,心跳由巡檢寫」:多一個寫心跳的人,心跳的
# 意思立刻退回舊版。install all 只裝巡檢那一筆。
case "$CMD:$JOBS" in
install:heartbeat)
die 6 '心跳那一筆不裝了。心跳改由巡檢跑完那一輪自己寫,排程只叫巡檢:請跑 `schedule.sh install patrol`。舊機器上留著的 heartbeat 條目,install patrol 會順手清掉。' ;;
install:*heartbeat*) JOBS='patrol' ;;
esac
HEARTBEAT=$(heartbeat_sh) \
|| die 2 '找不到 jsc-hooks 的 hooks/heartbeat.sh,心跳的過期門檻查不到,巡檢週期算不出來。請先安裝 jsc-hooks 0.3.7 以上。'
MECH=$(mechanism) \
|| die 3 "這台機器沒有可用的排程機制(作業系統:$(os_kind),找不到 $CRONTAB_CMD 或 schtasks)。"
# 巡檢週期與心跳門檻綁在一起算。--period 給的值一樣要通過同一條式子,否則裝出來的排程
# 會讓心跳永遠過期,而且沒有人看得出來是週期設錯。
TTL=$(heartbeat_ttl)
if [ -n "$PERIOD" ]; then
case "$PERIOD" in ''|*[!0-9]*) usage ;; esac
[ "$PERIOD" -ge 1 ] || usage
[ $(( 2 * PERIOD * 60 )) -lt "$TTL" ] \
|| die 6 "巡檢週期 $PERIOD 分鐘塞不進心跳的過期門檻 $TTL 秒(要 2 × 週期 × 60 < 門檻)。心跳現在由巡檢寫,週期比門檻長的話心跳永遠是過期。請改短週期,或先把 JSC_ASSISTANT_HEARTBEAT_TTL 調大。"
else
PERIOD=$(period_for_ttl "$TTL")
fi
# 巡檢指令在這裡就解出來。放進 cron_entry 再解的話,那支是在命令替換的子行程裡跑,
# 判不出 CLI 時 die 只結束子行程,主流程會帶著空指令繼續往下裝。
PATROL_RESOLVED=''
case "$CMD:$JOBS" in
install:*patrol*)
PATROL_RESOLVED=$(patrol_command) \
|| die 6 '判不出要用哪一支 CLI 跑巡檢,請帶 --cli {claude|codex|copilot|antigravity|kiro} 或 --patrol-cmd「指令」。' ;;
esac
TMPD=$(mktemp -d) || die 4 '建不出暫存目錄。'
trap 'rm -rf "$TMPD"' EXIT
# --- schtasks(Windows)---
schtasks_install() {
_rc=0
# 舊版的 heartbeat 任務一併刪掉,理由同 crontab 那一邊。刪不掉不算失敗,往下照裝。
_legacy=0
if schtasks /Query /TN "$(task_name heartbeat)" >/dev/null 2>&1; then
schtasks /Delete /TN "$(task_name heartbeat)" /F >/dev/null 2>&1 && _legacy=1
fi
for _job in $JOBS; do
_tn=$(task_name "$_job")
case "$_job" in
patrol) _mo="$PERIOD"; _run="$PATROL_RESOLVED" ;;
*) continue ;;
esac
_tr="cmd /c $_run <NUL >> \"$LOG\" 2>&1"
if [ "$DRYRUN" -eq 1 ]; then
printf 'dryrun=schtasks job=%s task=%s interval=%s cmd=%s\n' "$_job" "$_tn" "$_mo" "$_tr"
continue
fi
# /F 只覆蓋同名任務,也就是本腳本自己那一筆,不影響別人的排程。
schtasks /Create /TN "$_tn" /SC MINUTE /MO "$_mo" /F /TR "$_tr" >/dev/null 2>&1 \
|| die 4 "schtasks 建立 $_tn 失敗。"
schtasks /Query /TN "$_tn" >/dev/null 2>&1 || die 5 "schtasks 建立 $_tn 回 0,卻查不到這個任務。"
printf 'installed=%s task=%s\n' "$_job" "$_tn"
_rc=0
done
printf 'ttl=%s period=%s legacy_removed=%s log=%s\n' "$TTL" "$PERIOD" "$_legacy" "$LOG"
return "$_rc"
}
schtasks_remove() {
for _job in $JOBS; do
_tn=$(task_name "$_job")
if [ "$DRYRUN" -eq 1 ]; then
printf 'dryrun=schtasks job=%s task=%s action=delete\n' "$_job" "$_tn"
continue
fi
if ! schtasks /Query /TN "$_tn" >/dev/null 2>&1; then
note "$_job 本來就沒有排程,不用移除。"
continue
fi
schtasks /Delete /TN "$_tn" /F >/dev/null 2>&1 || die 4 "schtasks 刪除 $_tn 失敗。"
schtasks /Query /TN "$_tn" >/dev/null 2>&1 && die 5 "schtasks 刪除 $_tn 回 0,任務卻還在。"
done
return 0
}
schtasks_status() {
printf 'mechanism=schtasks service=%s ttl=%s period=%s log=%s\n' \
"$(service_state)" "$TTL" "$PERIOD" "$LOG"
for _job in heartbeat patrol; do
_tn=$(task_name "$_job")
if schtasks /Query /TN "$_tn" >/dev/null 2>&1; then
printf 'job=%s installed=yes task=%s\n' "$_job" "$_tn"
else
printf 'job=%s installed=no task=%s\n' "$_job" "$_tn"
fi
done
return 0
}
# --- crontab(Linux、WSL、macOS)---
crontab_install() {
_cur="$TMPD/cur"; _new="$TMPD/new"
cron_read "$_cur"
cp "$_cur" "$_new"
# 先濾掉自己這幾個工作的舊條目,再追加新的。重跑不會疊成兩筆,別人的條目原樣留著。
# 舊版的 heartbeat 條目一併清掉:留著它每分鐘照樣寫心跳,心跳就退回「只證明 cron 活著」。
_legacy=$(cron_lines_for "$_new" heartbeat | count_stdin)
for _job in $JOBS heartbeat; do
grep -vF "$(marker_of "$_job")" "$_new" >"$TMPD/f" 2>/dev/null || true
mv "$TMPD/f" "$_new"
done
_others=$(count_lines "$_new")
for _job in $JOBS; do
cron_entry "$_job" >>"$_new"
printf '\n' >>"$_new"
done
if [ "$DRYRUN" -eq 1 ]; then
for _job in $JOBS; do
printf 'dryrun=crontab job=%s entry=%s\n' "$_job" "$(cron_entry "$_job")"
done
printf 'dryrun=crontab action=write ttl=%s period=%s legacy_removed=%s others_kept=%s total_lines=%s\n' \
"$TTL" "$PERIOD" "$_legacy" "$_others" "$(count_lines "$_new")"
printf -- '--- 寫回後的 crontab ---\n'
cat "$_new"
return 0
fi
mkdir -p "$STATE_DIR" 2>/dev/null || true
"$CRONTAB_CMD" "$_new" >/dev/null 2>"$TMPD/err" \
|| die 4 "crontab 寫入失敗:$(tr '\n' ' ' <"$TMPD/err")"
_chk="$TMPD/chk"; cron_read "$_chk"
for _job in $JOBS; do
[ -n "$(cron_lines_for "$_chk" "$_job")" ] \
|| die 5 "crontab 寫入回 0,卻讀不到 $_job 的條目。"
[ "$(cron_lines_for "$_chk" "$_job" | count_stdin)" -eq 1 ] \
|| die 5 "$_job 的條目不只一筆,排程會重複執行。"
done
_kept=$(grep -vF "$MARK_PREFIX" "$_chk" 2>/dev/null | count_stdin)
[ "$_kept" -eq "$_others" ] \
|| die 5 "別人的排程條目從 $_others 筆變成 $_kept 筆,寫回不完整。"
for _job in $JOBS; do
printf 'installed=%s entry=%s\n' "$_job" "$(cron_lines_for "$_chk" "$_job")"
done
printf 'ttl=%s period=%s legacy_removed=%s others_kept=%s log=%s\n' \
"$TTL" "$PERIOD" "$_legacy" "$_kept" "$LOG"
return 0
}
crontab_remove() {
_cur="$TMPD/cur"; _new="$TMPD/new"
cron_read "$_cur"
_hit=0
cp "$_cur" "$_new"
for _job in $JOBS; do
if [ -n "$(cron_lines_for "$_new" "$_job")" ]; then
_hit=$((_hit + 1))
else
note "$_job 本來就沒有排程,不用移除。"
continue
fi
grep -vF "$(marker_of "$_job")" "$_new" >"$TMPD/f" 2>/dev/null || true
mv "$TMPD/f" "$_new"
done
if [ "$_hit" -eq 0 ]; then
printf 'removed=0 others_kept=%s\n' "$(count_lines "$_cur")"
return 0
fi
if [ "$DRYRUN" -eq 1 ]; then
printf 'dryrun=crontab action=remove jobs=%s removed=%s\n' "$JOBS" "$_hit"
printf -- '--- 寫回後的 crontab ---\n'
cat "$_new"
return 0
fi
"$CRONTAB_CMD" "$_new" >/dev/null 2>"$TMPD/err" \
|| die 4 "crontab 寫入失敗:$(tr '\n' ' ' <"$TMPD/err")"
_chk="$TMPD/chk"; cron_read "$_chk"
for _job in $JOBS; do
[ -z "$(cron_lines_for "$_chk" "$_job")" ] \
|| die 5 "crontab 刪除回 0,$_job 的條目卻還在。"
done
_before=$(grep -vF "$MARK_PREFIX" "$_cur" 2>/dev/null | count_stdin)
_after=$(grep -vF "$MARK_PREFIX" "$_chk" 2>/dev/null | count_stdin)
[ "$_before" -eq "$_after" ] \
|| die 5 "別人的排程條目從 $_before 筆變成 $_after 筆,移除動到了不該動的東西。"
printf 'removed=%s others_kept=%s\n' "$_hit" "$_after"
return 0
}
crontab_status() {
_cur="$TMPD/cur"; cron_read "$_cur"
printf 'mechanism=crontab service=%s ttl=%s period=%s log=%s\n' \
"$(service_state)" "$TTL" "$PERIOD" "$LOG"
for _job in heartbeat patrol; do
_line=$(cron_lines_for "$_cur" "$_job")
if [ -n "$_line" ]; then
printf 'job=%s installed=yes entry=%s\n' "$_job" "$_line"
else
printf 'job=%s installed=no entry=\n' "$_job"
fi
done
return 0
}
# --- 主流程 ---
case "$CMD" in
install)
case "$MECH" in
crontab) crontab_install ;;
schtasks) schtasks_install ;;
esac
[ "$DRYRUN" -eq 1 ] && exit 0
_svc=$(service_state)
if [ "$_svc" = stopped ]; then
printf 'service=stopped\n'
die 1 '排程條目寫進去了,但 cron 服務沒在跑,條目一次都不會被執行。WSL 預設不啟動 cron:請跑 `sudo service cron start`,而且重開 WSL 之後要再啟動一次。'
fi
printf 'service=%s\n' "$_svc"
[ "$_svc" = unknown ] && note '判不出排程服務在不在跑,請自行確認條目真的會被執行。'
exit 0 ;;
remove)
case "$MECH" in
crontab) crontab_remove ;;
schtasks) schtasks_remove ;;
esac
exit 0 ;;
status)
case "$MECH" in
crontab) crontab_status ;;
schtasks) schtasks_status ;;
esac
exit 0 ;;
esac