feat/assistant-body/start-stop-status #4

Merged
admin merged 3 commits from feat/assistant-body/start-stop-status into feat/assistant-body/main 2026-09-01 07:24:30 +00:00
7 changed files with 957 additions and 73 deletions
Showing only changes of commit 2b7176a644 - Show all commits
+7 -3
View File
@@ -26,7 +26,7 @@ Marketplace 統一為 `jsc`(https://gitea.jsc.idv.tw/plugins/meta.git),安
### `assistant`
助理主體,三個操作:`start` 啟動、`status` 查現況、`stop` 停止。心跳的寫入、判定與清除一律交給 `jsc-hooks` 的 `hooks/heartbeat.sh`,判定只有那一份;系統排程一律交給 `tools/schedule.sh`,技能自己不碰 crontab 與 schtasks。`start` 寫下第一次心跳,再裝上每分鐘寫一次心跳的排程;cron 服務沒在跑就照實講條目不會被執行。`status` 全程唯讀,讀心跳、排程與待辦簿,印成三塊;助理沒在跑就印「助理未運行」,不當成錯誤。`stop` 先移除排程再清掉心跳,順序不能反——反了下一分鐘 cron 會再寫一次心跳。這支不碰 wiki、不參與閘門判定。心跳新鮮只證明排程活著,不證明助理做了事。
助理主體,四個操作:`start` 啟動、`status` 查現況、`patrol` 跑一輪巡檢、`stop` 停止。心跳的寫入、判定與清除一律交給 `jsc-hooks` 的 `hooks/heartbeat.sh`,判定只有那一份;系統排程一律交給 `tools/schedule.sh`;一輪巡檢的流程交給 `tools/patrol.sh`。**心跳由巡檢寫,而且只由巡檢寫**:一輪跑完、結果寫上監控頁了,才寫那一次心跳,所以心跳新鮮等於「上一輪巡檢真的做完了」。`start` 先跑一輪巡檢,再裝上巡檢那一筆排程;巡檢週期由心跳的過期門檻算出來,兩個數字綁在一起。`patrol` 讀四項來源(使用統計、版本與重啟閘門、SDLC 階段鎖與工作包鎖、心跳自述),四項各自獨立,一項掛掉其餘三項照跑、照記,結果一律附加到 `MONITOR_{HASH}`、不覆寫。`status` 全程唯讀,讀心跳、排程與待辦簿,印成三塊;助理沒在跑就印「助理未運行」,不當成錯誤。`stop` 先移除排程再清掉心跳,順序不能反。這支不參與閘門判定、不做決策、巡檢那一路全程不問人。
<!-- JSC-SKILLS:END -->
@@ -43,7 +43,8 @@ Marketplace 統一為 `jsc`(https://gitea.jsc.idv.tw/plugins/meta.git),安
| 檔案 | 用途 |
| --- | --- |
| `tools/schedule.sh` | 助理系統排程的安裝、移除與查現況。三個子命令 `install`、`remove`、`status`,兩筆工作 `heartbeat`(每 60 秒)與 `patrol`(每 15 分鐘)。Linux、WSL 與 macOS 走 crontab,Windows 走 schtasks。條目行尾帶固定標記 `# jsc-assist:assistant {工作}`,只動自己那一筆,別人的排程一行都不碰。裝完會檢查排程服務在不在跑,沒跑就回 1——WSL 預設不啟動 cron。巡檢那一筆預設不裝,巡檢本體還沒實作。`--dry-run` 只印組出來的條目與寫回後的內容,什麼都不動 |
| `tools/schedule.sh` | 助理系統排程的安裝、移除與查現況。三個子命令 `install`、`remove`、`status`,只裝 `patrol` 這一筆——心跳由巡檢自己寫,`install heartbeat` 一律回 6,舊版遺留的心跳條目由 `install patrol` 順手清掉。巡檢週期由心跳的過期門檻算出來(`2 × 週期 × 60 < 門檻`,再取能整除一小時的分鐘數):門檻 300 秒是每 2 分鐘一輪,門檻 1800 秒是每 12 分鐘一輪。Linux、WSL 與 macOS 走 crontab,Windows 走 schtasks。條目行尾帶固定標記 `# jsc-assist:assistant {工作}`,只動自己那一筆,別人的排程一行都不碰。裝完會檢查排程服務在不在跑,沒跑就回 1——WSL 預設不啟動 cron。`--dry-run` 只印組出來的條目與寫回後的內容,什麼都不動 |
| `tools/patrol.sh` | 一輪巡檢的收攏與收口。三個子命令:`collect` 取鎖、讀四項來源、組出監控頁要附加的那一節與目錄頁那一列;`finish` 在監控頁寫成之後才寫心跳、換上用量快照、放掉鎖;`abort` 只放掉鎖,不寫心跳。四項來源各自獨立,一項失敗其餘三項照跑,失敗那一項在頁上寫明是「這一項失敗」而不是沒資料。整輪拿一把目錄鎖,上一輪還在跑就回 4 讓開;鎖逾時(門檻取心跳門檻)會被下一輪搶回來,並在頁上記一筆。`version-guard.sh report` 回「查詢失敗」時照原字抄,不補查、不美化 |
| `references/behaviors.md` | 本 domain 的技能行為清單:一支技能一節,五列記下觸發時機、關鍵步驟、外部呼叫、完成條件、可驗證跡象,供稽核與驗證比對。格式合約見 `plugins/meta` 的 `references/guidelines.md`「技能行為清單」 |
| `templates/monitor-contents.md` | 目錄頁 `MONITOR_CONTENTS` 的範本。一列代表一台機器,雜湊來源是 `{主機名}/{登入帳號}`。寫入語意是**只更新自己那一列**:比對主機與帳號兩欄,別台機器的列原樣保留,禁止整頁覆蓋 |
| `templates/monitor-page.md` | 內容頁 `MONITOR_{HASH}` 的範本。記的是這台機器的巡檢軌跡。寫入語意與目錄頁相反,是**一律附加一節、不覆寫**:一次巡檢一節,節標題帶時間戳,既有的節一個字都不動 |
@@ -54,9 +55,12 @@ Marketplace 統一為 `jsc`(https://gitea.jsc.idv.tw/plugins/meta.git),安
| 路徑 | 內容 |
| --- | --- |
| `heartbeat` | 心跳檔,欄位 `ts`、`pid`、`cli`、`session`。判準只看 `ts`,不看 pid 存活——五支 CLI 與容器裡的行程互相看不到彼此的 pid |
| `heartbeat` | 心跳檔,欄位 `ts`、`pid`、`cli`、`session`。只由 `tools/patrol.sh finish` 寫,也就是一輪巡檢跑完、結果記下來之後才寫。判準只看 `ts`,不看 pid 存活——五支 CLI 與容器裡的行程互相看不到彼此的 pid |
| `tasks/{id}` | 待辦簿,一筆一檔。一筆一檔是為了讓並行寫入不互相覆寫 |
| `schedule.log` | 排程條目的輸出。刻意放在存取庫外面:寫進專案會多出未追蹤檔,污染別人的變更盤點 |
| `patrol.lock/` | 一輪巡檢的鎖,是目錄——`mkdir` 是原子操作,搶不到就是別人在跑。裡面的 `info` 記 `round`、`pid`、`started` |
| `patrol/` | 本輪巡檢的暫存檔:`section.md` 是要附加的那一節,`newpage.md` 是頁不存在時要建的整頁,`contents.tsv` 是目錄頁那一列 |
| `usage-prev.tsv` | 上一輪記下來的累計用量。有了它,下一輪的「本輪次數」才算得出來;沒有它的第一輪一律寫「-」,不拿累計冒充本輪 |
## 相關 domain
+5 -5
View File
@@ -6,8 +6,8 @@
| 項目 | 內容 |
| --- | --- |
| 觸發時機 | 要啟動助理、要停止助理,或要問助理現在還在不在跑、待辦簿剩下哪幾筆時用。三個操作 `start`、`status`、`stop` 都走這一支。執行環境健檢不走這支,走 `jsc-cli:doctor`。技能使用次數不走這支,走 `jsc-log:stats` |
| 關鍵步驟 | 先認出使用者要的是哪一個操作。`start`:跑 `heartbeat.sh write` 寫第一次心跳、跑 `heartbeat.sh report` 確認寫進去了、跑 `tools/schedule.sh install heartbeat` 裝每分鐘那一筆排程、依結束碼選一段收尾訊息印出——排程接上、排程寫進去了但 cron 沒在跑、排程沒接上三種各一段。巡檢那一筆預設不裝,要呼叫端指名 `patrol` 或 `all` 才裝,而且要先講明巡檢本體還沒實作、裝了每一輪都會失敗。`status`:跑 `heartbeat.sh report` 取心跳現況、把 `state` 對映成新鮮、過期、心跳檔損壞、不存在、不自己解析心跳檔也不自己判定、從 `file=` 解出助理目錄後列出 `tasks/` 底下每一個檔案並解析 `state`、`title`、`next_run`、`fail_count`、跑 `tools/schedule.sh status` 取排程現況、印成心跳、排程、待辦三塊、`fail_count` 大於 0 的列標上「已連續失敗 N 次」、心跳與排程兜起來會誤讀的三種組合各補一句話。`stop`:先跑 `heartbeat.sh report` 留下原本的狀態、再跑 `tools/schedule.sh remove all` 移除排程、最後才跑 `heartbeat.sh clear` 清掉心跳、印出停止訊息並說明心跳清掉之後閘門會擋人、同時說明閘門還沒接線所以現在擋不到人 |
| 外部呼叫 | `jsc-hooks/hooks/heartbeat.sh` 的 `write`、`report`、`clear` 三個子命令,六個結束碼各有處置:0 往下走、1 與 3 印「助理未運行」、2 回報判不出狀態並停下、4 當成不新鮮並回報心跳檔損壞、5 是嚴重狀況要吵出來且不得回報成功、6 是呼叫寫錯要更正後重跑。本 domain 的 `tools/schedule.sh` 的 `install`、`remove`、`status` 三個子命令,七個結束碼各有處置:0 往下走、1 是條目裝了但 cron 沒在跑要照實講不會執行、2 是缺 jsc-hooks、3 是這台機器沒有排程機制、4 是排程操作失敗要原樣引用 stderr、5 是回讀驗證失敗要叫人自己去看 `crontab -l`、6 是呼叫寫錯。crontab 與 schtasks 一律經那支腳本,技能自己不碰。另外唯讀 `$JSC_HOME/assistant/tasks/` 底下的檔案。呼叫端沒講清楚要哪一個操作時,走 `jsc-ask:ask` 的決策樹問。不碰 wiki、不參與閘門判定 |
| 完成條件 | `start` 要 `write` 回 0 且 `report` 回 `state=fresh`,才算啟動成功;`write` 回 5 一律回報失敗並停下,不得宣稱啟動;`schedule.sh install` 回 1 要講明條目不會被執行與 `sudo service cron start`,不得宣稱排程會定時執行。`status` 要印出現況表,或印出「助理未運行」並說明原因;心跳不存在、待辦簿目錄不存在、待辦簿零筆、排程沒裝,四種都算正常結束。`stop` 要 `schedule.sh remove all` 先回 0、`clear` 再回 0,並印出帶三段話的停止訊息;`remove` 非 0 就回報排程還在、助理停不掉,不清心跳也不印停止訊息;`clear` 回 5 就回報心跳檔還在、助理沒有確實停掉,不印停止訊息 |
| 可驗證跡象 | `start` 之後 `$JSC_HOME/assistant/heartbeat` 存在,`ts` 是剛才的時間,`crontab -l` 找得到一筆帶 `# jsc-assist:assistant heartbeat` 的條目,而且只有一筆。`stop` 之後心跳路徑不存在,`crontab -l` 找不到任何 `# jsc-assist:assistant` 條目。兩者都不動別人的排程條目,條目數量前後相同。`status` 無寫入跡象,只有回報內容。三個操作都不動 `tasks/` 底下的檔案,也不動 worktree 與 wiki 頁。排程的 log 一律在 `$JSC_HOME/assistant/schedule.log`,不落在任何存取庫 |
| 觸發時機 | 要啟動助理、要停止助理、要跑一輪巡檢,或要問助理現在還在不在跑、待辦簿剩下哪幾筆時用。四個操作 `start`、`status`、`patrol`、`stop` 都走這一支。排程每一輪叫起來的也是這一支的 `patrol`。執行環境健檢不走這支,走 `jsc-cli:doctor`。技能使用次數不走這支,走 `jsc-log:stats` |
| 關鍵步驟 | 先認出使用者要的是哪一個操作,`patrol` 那一路全程不問人。`start`:先照 `patrol` 的每一步跑完一輪巡檢,第一次心跳由那一輪寫、不另外寫、跑不完就不算啟動、跑 `heartbeat.sh report` 確認 `state=fresh`、跑 `tools/schedule.sh install patrol` 裝巡檢那一筆排程、依結束碼選一段收尾訊息印出——排程接上、排程寫進去了但 cron 沒在跑、排程沒接上三種各一段。心跳那一筆不裝了,`install heartbeat` 一律回 6。`patrol`:跑 `tools/patrol.sh collect` 取鎖並讀四項來源、結束碼 4 就讓開不寫任何東西、結束碼 1 與 3 照樣把那一節寫上監控頁、`hash` 是空的就 `abort`、把 `section_file` 交給 `jsc-gitea:wiki` 附加到 `MONITOR_{HASH}`、頁不存在(唯有結束碼 4)才用 `newpage_file` 建頁、把 `contents_file` 的 `row` 更新到 `MONITOR_CONTENTS` 自己那一列、兩次寫入任一失敗就 `abort` 且不寫心跳、全部寫成才跑 `tools/patrol.sh finish` 寫心跳、最後印出四項結果與待人處理列。`status`:跑 `heartbeat.sh report` 取心跳現況、把 `state` 對映成新鮮、過期、心跳檔損壞、不存在、不自己解析心跳檔也不自己判定、從 `file=` 解出助理目錄後列出 `tasks/` 底下每一個檔案並解析 `state`、`title`、`next_run`、`fail_count`、跑 `tools/schedule.sh status` 取排程現況與週期、印成心跳、排程、待辦三塊、`fail_count` 大於 0 的列標上「已連續失敗 N 次」、心跳與排程兜起來會誤讀的四種組合各補一句話。`stop`:先跑 `heartbeat.sh report` 留下原本的狀態、再跑 `tools/schedule.sh remove all` 移除排程與舊版遺留的心跳條目、最後才跑 `heartbeat.sh clear` 清掉心跳、印出停止訊息並說明心跳清掉之後閘門會擋人、同時說明閘門還沒接線所以現在擋不到人 |
| 外部呼叫 | `jsc-hooks/hooks/heartbeat.sh` 的 `write`、`report`、`clear` 三個子命令,六個結束碼各有處置:0 往下走、1 與 3 印「助理未運行」、2 回報判不出狀態並停下、4 當成不新鮮並回報心跳檔損壞、5 是嚴重狀況要吵出來且不得回報成功、6 是呼叫寫錯要更正後重跑。`write` 只由 `tools/patrol.sh finish` 呼叫,技能自己不呼叫。本 domain 的 `tools/schedule.sh` 的 `install`、`remove`、`status` 三個子命令,七個結束碼各有處置:0 往下走、1 是條目裝了但 cron 沒在跑要照實講不會執行、2 是缺 jsc-hooks 導致門檻讀不到、3 是這台機器沒有排程機制、4 是排程操作失敗要原樣引用 stderr、5 是回讀驗證失敗要叫人自己去看 `crontab -l`、6 是呼叫寫錯,含 `install heartbeat` 與週期塞不進門檻。本 domain 的 `tools/patrol.sh` 的 `collect`、`finish`、`abort` 三個子命令,七個結束碼各有處置:0 往下走、1 部分失敗照樣寫頁、2 是 finish 找不到 heartbeat.sh 要回報「記下來了但沒有心跳」、3 是四項全失敗照樣寫頁且判定異常、4 是讓開或鎖被搶走一律不寫心跳、5 是檔案系統失敗要吵出來、6 是呼叫寫錯。巡檢那四項讀 `jsc-log/tools/usage-stats.sh`、`jsc-hooks/hooks/version-guard.sh report`、`jsc-hooks/hooks/restart-gate.sh report`、`$JSC_HOME/sessions/*.stage`、`$JSC_HOME/wp/*.pr`、`heartbeat.sh report`,全部只讀,任一項失敗不影響其餘三項。wiki 讀寫一律經 `jsc-gitea:wiki`,技能自己不拼 API 呼叫。crontab 與 schtasks 一律經 `tools/schedule.sh`。另外唯讀 `$JSC_HOME/assistant/tasks/` 底下的檔案。呼叫端沒講清楚要哪一個操作時走 `jsc-ask:ask` 的決策樹問,但 `patrol` 那一路一律不問。不參與閘門判定 |
| 完成條件 | `start` 要那一輪巡檢的 `finish` 回 0 且 `report` 回 `state=fresh`,才算啟動成功;巡檢沒寫成心跳一律回報失敗並停下,不得宣稱啟動;`schedule.sh install patrol` 回 1 要講明條目不會被執行與 `sudo service cron start`,不得宣稱排程會定時執行。`patrol` 要四項各自有 `status`、監控頁附加成功、目錄頁那一列更新成功、`finish` 回 0,才算一輪跑完;`collect` 回 4 是讓開,不算失敗也不寫任何東西;監控頁或目錄頁任一沒寫成就 `abort`,回報「這一輪沒有結果」,心跳一定不寫。`status` 要印出現況表,或印出「助理未運行」並說明原因;心跳不存在、待辦簿目錄不存在、待辦簿零筆、排程沒裝,四種都算正常結束。`stop` 要 `schedule.sh remove all` 先回 0、`clear` 再回 0,並印出帶三段話的停止訊息;`remove` 非 0 就回報排程還在、助理停不掉,不清心跳也不印停止訊息;`clear` 回 5 就回報心跳檔還在、助理沒有確實停掉,不印停止訊息 |
| 可驗證跡象 | `start` 之後 `$JSC_HOME/assistant/heartbeat` 存在,`ts` 是剛才那一輪的時間,`crontab -l` 找得到一筆帶 `# jsc-assist:assistant patrol` 的條目,而且只有一筆,帶 `# jsc-assist:assistant heartbeat` 的舊條目一筆都不剩。`patrol` 跑完之後 wiki 的 `MONITOR_{HASH}` 多一節、節標題帶時間戳、舊的節一字不改,`MONITOR_CONTENTS` 只有自己那一列變動,`$JSC_HOME/assistant/patrol/` 底下有本輪的 `section.md`、`newpage.md`、`contents.tsv`,`$JSC_HOME/assistant/usage-prev.tsv` 換成本輪的累計數,`$JSC_HOME/assistant/patrol.lock` 已經放掉。讓開的那一輪沒有任何寫入跡象。`stop` 之後心跳路徑不存在,`crontab -l` 找不到任何 `# jsc-assist:assistant` 條目。以上都不動別人的排程條目,條目數量前後相同。`status` 無寫入跡象,只有回報內容。四個操作都不動 `tasks/` 底下的檔案,也不動 worktree 與程式碼存取庫。排程的 log 一律在 `$JSC_HOME/assistant/schedule.log`,不落在任何存取庫 |
+98 -40
View File
@@ -1,27 +1,34 @@
---
name: assistant
description: Start, inspect, or stop the background assistant, with jsc-hooks/hooks/heartbeat.sh owning the single freshness verdict and tools/schedule.sh owning the system scheduler. start writes the first heartbeat and installs the one-minute heartbeat entry through crontab or schtasks; status turns heartbeat.sh report, schedule.sh status and the task book under $JSC_HOME/assistant/tasks/ into one read-only table; stop removes the schedule first, then clears the heartbeat, and states what a cleared heartbeat means for the jsc skill gate. Both scripts only ever touch their own marked entry, never the rest of the user's crontab. A heartbeat that cannot be written or cleared, and a schedule installed on a machine whose cron service is not running, are reported as failures, never as success. Use when someone starts the assistant, stops it, or asks whether it is running and what is queued; not for environment health checks (jsc-cli:doctor), not for skill usage counts (jsc-log:stats).
description: Start, inspect, patrol, or stop the background assistant, with jsc-hooks/hooks/heartbeat.sh owning the single freshness verdict, tools/schedule.sh owning the system scheduler, and tools/patrol.sh owning one patrol round. The heartbeat is written by a completed patrol round and by nothing else, so the schedule carries the patrol entry only and its period is derived from the heartbeat TTL; start runs one round and then installs that entry, status turns heartbeat.sh report, schedule.sh status and the task book into one read-only table, stop removes the entry first and then clears the heartbeat. One round reads four independent sources - skill and chain usage, version gaps and the restart gate, SDLC stage and work-package locks, and the heartbeat's own report - and appends the result to wiki MONITOR_{HASH} through jsc-gitea:wiki before tools/patrol.sh finish writes the heartbeat. A round that cannot record its result writes no heartbeat, and a round that starts while the previous one still holds the lock stands down. Use when someone starts, patrols or stops the assistant, or asks whether it is running and what is queued; not for environment health checks (jsc-cli:doctor), not for skill usage counts (jsc-log:stats).
---
# assistant — start, status, stop
# assistant — start, status, patrol, stop
The background assistant runs where nobody is watching it. Its heartbeat is the only evidence that it is alive, so this skill is the single entry point for the three operations that touch that evidence: `start` writes it, `status` reads it, `stop` clears it.
The background assistant runs where nobody is watching it. Its heartbeat is the only evidence that it is alive, so this skill is the single entry point for the four operations that touch that evidence: `patrol` writes it, `status` reads it, `stop` clears it, and `start` bootstraps the whole loop.
`jsc-hooks/hooks/heartbeat.sh` owns every heartbeat operation, including the freshness verdict. Never read, parse, write or delete `$JSC_HOME/assistant/heartbeat` directly — one verdict, one source.
`tools/schedule.sh` owns every system-scheduler operation: installing an entry, removing it, and reading which entries exist. Never call `crontab` or `schtasks` from this skill, and never edit a crontab by hand. Both flows have fixed inputs and outputs, so both live in scripts; the task book is the only thing this skill reads for itself, and that is one directory listing.
`tools/schedule.sh` owns every system-scheduler operation: installing an entry, removing it, and reading which entries exist. Never call `crontab` or `schtasks` from this skill, and never edit a crontab by hand.
`tools/patrol.sh` owns one patrol round: taking the round lock, reading the four sources, composing the monitor-page section, and — after that section is on the page — writing the heartbeat. Never re-read a source this skill already handed to that script, and never compose the section by hand; the script prints the file paths.
All three flows have fixed inputs and outputs, so all three live in scripts. The task book is the only thing this skill reads for itself, and that is one directory listing.
## Pick the operation
Run exactly one operation per invocation. Take it from the request: starting, launching or waking the assistant is `start`; asking whether it runs, what it is doing, or what is queued is `status`; stopping, halting or shutting it down is `stop`. When the request names none of the three, or names more than one, ask through the `jsc-ask:ask` decision tree with those three as the options, each stating its effect — `start` writes a heartbeat and installs the scheduled entry that keeps writing it, `status` changes nothing, `stop` removes that entry and deletes the heartbeat. Never guess, and never run a second operation the caller did not ask for. Completion condition: exactly one of `start`, `status`, `stop` is chosen and named in the report.
Run exactly one operation per invocation. Take it from the request: starting, launching or waking the assistant is `start`; asking whether it runs, what it is doing, or what is queued is `status`; running one round, patrolling, or a scheduled wake-up is `patrol`; stopping, halting or shutting it down is `stop`. When the request names none of the four, or names more than one, ask through the `jsc-ask:ask` decision tree with those four as the options, each stating its effect — `start` runs one round and installs the scheduled entry that keeps running rounds, `status` changes nothing, `patrol` runs one round and writes one heartbeat, `stop` removes that entry and deletes the heartbeat. **The one exception: a `patrol` invocation never asks anything at all** (see 界線 1 below). Never guess, and never run a second operation the caller did not ask for. Completion condition: exactly one of `start`, `status`, `patrol`, `stop` is chosen and named in the report.
## Data sources
| Path | Read by | Format |
| --- | --- | --- |
| `$JSC_HOME/assistant/heartbeat` | `heartbeat.sh` only, never this skill | `key=value` lines: `ts`, `pid`, `cli`, `session` |
| `$JSC_HOME/assistant/schedule.log` | nobody here — the scheduled entries append to it | free text; point the operator at it when a scheduled round misbehaves |
| `$JSC_HOME/assistant/schedule.log` | nobody here — the scheduled entry appends to it | free text; point the operator at it when a scheduled round misbehaves |
| `$JSC_HOME/assistant/tasks/{id}` | this skill, read-only | `key=value` lines, one task per file: `id`, `kind` (`check` / `todo`), `title`, `action`, `trigger`, `recur`, `repo`, `due`, `state` (`pending` / `done` / `paused`), `last_run`, `next_run`, `fail_count`, `origin` (`user` / `assistant`) |
| `$JSC_HOME/assistant/patrol.lock/` | `patrol.sh` only | the round lock, a directory. `info` holds `round`, `pid`, `started` |
| `$JSC_HOME/assistant/patrol/` | `patrol.sh` only | one round's scratch files, including `section.md`, `newpage.md` and `contents.tsv` |
| `$JSC_HOME/assistant/usage-prev.tsv` | `patrol.sh` only | last recorded round's cumulative usage counts, so the next round can print a real per-round delta |
`$JSC_HOME` defaults to `~/.jsc`. `heartbeat.sh report` prints the resolved heartbeat path in its `file=` field, so take the assistant directory from there rather than rebuilding it.
@@ -34,36 +41,44 @@ Every call in every operation below is judged by this table. Report the code you
| Code | Meaning | What to do |
| --- | --- | --- |
| 0 | `write` wrote the heartbeat, `clear` finished and the file is gone, `report` printed its line, `check` says fresh | Carry on with the operation's next step. For `report`, the state still has to be read out of the printed `state=` field |
| 1 | `check`: the heartbeat exists but is at or past the TTL — the assistant ran and has stopped | Report `助理未運行`, name the age in seconds, and say the assistant has to be started again. `report` returns this state as `state=stale` with exit 0 |
| 1 | `check`: the heartbeat exists but is at or past the TTL — the last patrol round finished more than one TTL ago | Report `助理未運行`, name the age in seconds, and say the assistant has to be started again. `report` returns this state as `state=stale` with exit 0 |
| 2 | The script did not run at all — it failed to load its `lib.sh` | Report that the heartbeat state is unknown, name the script path and the code, and stop the operation. Never claim the assistant is running, and never claim it is stopped |
| 3 | `check`: no heartbeat file — the assistant has never been started | Report `助理未運行` and say to run `start`. `report` returns this state as `state=absent` with exit 0. In `stop` this state cannot appear, because `clear` treats a missing file as success |
| 3 | `check`: no heartbeat file — no patrol round has ever finished | Report `助理未運行` and say to run `start`. `report` returns this state as `state=absent` with exit 0. In `stop` this state cannot appear, because `clear` treats a missing file as success |
| 4 | `check`: the heartbeat exists but its `ts` is missing, empty or not a number — the file is damaged, the assistant is not merely stopped | Treat it as not fresh; falling back to fresh is forbidden. Report the file as damaged, say the state cannot be read from it, and tell the operator to run `stop` and then `start` to rebuild it. `report` returns this state as `state=invalid` with exit 0 |
| 5 | Filesystem failure — `write` could not write the file, or `clear` could not delete it and the file is still there | Serious. Report it loudly with the stderr text and the path, and follow the operation's own step for this code. Never report the operation as done |
| 6 | Usage error — an unknown subcommand, or none at all | This is a defect in the call, not a state of the assistant. Report the exact command line that was run, correct it to one of `write`, `check`, `report`, `clear`, and run it once more. Report a second exit 6 as a defect in this skill and stop |
## The scheduler
Nothing in a background assistant runs on its own. The system scheduler is what makes it periodic, and `tools/schedule.sh` is the only thing here that touches it. Two jobs exist, each written as exactly one entry carrying the fixed marker `# jsc-assist:assistant {job}`:
Nothing in a background assistant runs on its own. The system scheduler is what makes it periodic, and `tools/schedule.sh` is the only thing here that touches it. One job exists, written as exactly one entry carrying the fixed marker `# jsc-assist:assistant patrol`:
| Job | Period | Runs | Installed by `start` |
| --- | --- | --- | --- |
| `heartbeat` | every 60 seconds (`* * * * *`) | `jsc-hooks/hooks/heartbeat.sh write` | yes, always |
| `patrol` | every 15 minutes (`*/15 * * * *`) | one round of the assistant's patrol | no — only on an explicit request |
| `patrol` | derived from the heartbeat TTL (`*/2 * * * *` at the default TTL of 300 seconds) | one patrol round through the caller's CLI | yes, always |
| `heartbeat` | — | nothing. This job existed in the previous version and is no longer installable | no — `install heartbeat` exits 6 |
**The patrol body is not implemented yet (M-04).** Installing that entry today buys a failure every fifteen minutes and a log full of them, so `start` never installs it. It goes in only when the caller names it — `schedule.sh install patrol` or `schedule.sh install all` — and the caller has to be told what they are asking for before it is run. Its command line also cannot be guessed: pass `--cli` or `--patrol-cmd`, or the script exits 6 rather than scheduling a binary that may not exist.
**The heartbeat job is gone on purpose.** It used to call `heartbeat.sh write` every minute, which made a fresh heartbeat prove only that cron was alive. Anything that writes a heartbeat outside a finished patrol round brings that back, so `install heartbeat` is refused, and `install patrol` deletes any leftover `heartbeat` entry from an older install and reports `legacy_removed=1`. Say that number in the report — a surviving legacy entry silently undoes this whole design.
**The period is derived, never guessed.** The heartbeat now moves once per patrol round, so the round period has to fit inside the freshness threshold. `schedule.sh` reads the machine's effective threshold from `heartbeat.sh report`'s `ttl=` field and picks the largest whole-hour-dividing minute count `P` with `2 × P × 60 < ttl`: one missed round still reads fresh, two missed rounds read stale. At the default 300 seconds that is every 2 minutes; raise `JSC_ASSISTANT_HEARTBEAT_TTL` to 1800 and it becomes every 12 minutes. Report both numbers (`ttl=`, `period=`) so the operator can see the trade-off and change it in one place. `--period` overrides the calculation and is checked against the same inequality; a period that does not fit exits 6 rather than installing a schedule that keeps the heartbeat permanently stale.
The mechanism follows the platform: `crontab` on Linux, WSL and macOS, `schtasks` on Windows. macOS keeps `crontab` — a `launchd` user who wants a plist writes it themselves; this skill does not generate one.
Four properties of that script matter enough to state here, because a report that ignores any of them is wrong:
- **It only ever touches its own entry.** Install filters out its own old entries by marker and appends the new one; it never rewrites a crontab it failed to read, and it counts everybody else's lines before and after to prove none went missing. Remove takes out its own marker only. Say this in the report — the operator is entitled to know their own cron entries survived.
- **It only ever touches its own entries.** Install filters out its own old entries by marker and appends the new one; it never rewrites a crontab it failed to read, and it counts everybody else's lines before and after to prove none went missing. Remove takes out its own markers only. Say this in the report — the operator is entitled to know their own cron entries survived.
- **A written entry is not a running entry.** WSL does not start cron by default, and this is the machine's most likely state. Exit 1 from `install` means the entry is on disk and will never fire. Report that as a failure of the start, name `sudo service cron start`, and say it has to be run again after every WSL restart. Never soften exit 1 into "scheduling is set up".
- **The log lives at `$JSC_HOME/assistant/schedule.log`**, deliberately outside every repository. Do not offer to move it into a project.
- **The entry runs with no human present.** Every command is installed with `</dev/null`, so nothing it runs can block on input. A patrol command that stops to ask for a tool permission still hangs the round, which is one more reason the patrol entry waits for M-04.
- **The entry runs with no human present.** The command is installed with `</dev/null`, so nothing it runs can block on input. A patrol round that stops to ask for a tool permission hangs that round, and the lock it holds stands the next round down until the lock ages out — which is why `patrol` asks nothing, of anybody, ever.
### What a fresh heartbeat actually proves
Once the heartbeat entry is installed, cron writes a heartbeat every minute for as long as the machine is up. So the verdict 新鮮 proves the scheduler is alive — and nothing more. It does not prove a patrol ran, that any task in the book moved, or that the assistant did a single useful thing. **Never turn a fresh heartbeat into a claim about work done.** `status` prints the heartbeat, the schedule and the task book as three separate facts for exactly this reason, and the task rows — `last_run`, `next_run`, `fail_count` — are the only evidence about work. The same limit binds whatever gate reads this heartbeat later: a fresh heartbeat is grounds for not blocking, never grounds for saying the assistant is doing its job.
The heartbeat is written in exactly one place: `tools/patrol.sh finish`, and `finish` is called only after that round's result is on the monitor page. So the verdict 新鮮 now proves one thing that is worth proving — **the last patrol round ran to the end and its result was recorded** — and it still does not prove three others:
- **Not that the round was clean.** Four sources are read independently and a round with three failures still records and still beats. The health of a round is `本輪判定` on the monitor page, never the heartbeat.
- **Not that any task in the book moved.** The task rows — `last_run`, `next_run`, `fail_count` — are the only evidence about work.
- **Not that the round did anything about what it found.** The patrol reports; a human acts. 界線 6.
The failure this design buys is the one worth having: a round that cannot read its sources, cannot reach the wiki, or dies half way writes no heartbeat, so the heartbeat ages past the TTL and every reader sees 過期. **A silent patrol is now indistinguishable from a stopped assistant, which is exactly right.** `status` still prints the heartbeat, the schedule and the task book as three separate facts, and the same limit binds whatever gate reads this heartbeat later: a fresh heartbeat is grounds for not blocking, never grounds for saying the assistant is doing its job.
## schedule.sh exit codes
@@ -71,44 +86,79 @@ Once the heartbeat entry is installed, cron writes a heartbeat every minute for
| --- | --- | --- |
| 0 | `install` wrote the entry and read it back, the scheduler service is running; `remove` finished, or there was nothing to remove; `status` printed its lines | Carry on. For `status`, the state still has to be read out of the `installed=` fields |
| 1 | `install` wrote the entry, but the cron service is not running — the entry will never fire | The start did not succeed. Report the entry as installed and inert, quote the fix (`sudo service cron start`, and again after each WSL restart), and never claim the assistant will keep itself alive |
| 2 | `jsc-hooks/hooks/heartbeat.sh` was not found | Report that `jsc-hooks` is missing or too old (0.3.7 or newer is required) and stop the operation |
| 2 | `jsc-hooks/hooks/heartbeat.sh` was not found, so the TTL cannot be read and the period cannot be derived | Report that `jsc-hooks` is missing or too old (0.3.7 or newer is required) and stop the operation |
| 3 | No usable scheduler on this machine | Report the platform and that neither `crontab` nor `schtasks` was found, and stop. Never fall back to some other mechanism |
| 4 | The scheduler operation failed — the existing schedule could not be read for a reason other than "no crontab", or the write or delete returned non-zero | Report the stderr text verbatim. A read failure means nothing was written, so the user's other entries are untouched; say so |
| 5 | Read-back verification failed — the entry is missing after a successful write, is present twice, is still there after a delete, or somebody else's line count changed | Serious. Report it loudly with the printed numbers, and tell the operator to inspect `crontab -l` by hand before anything else is run |
| 6 | Usage error — an unknown subcommand or job name, a missing option value, or the patrol CLI could not be determined | A defect in the call, not a state of the machine. Correct the command line and run it once more; report a second exit 6 as a defect in this skill and stop |
| 6 | Usage error — an unknown subcommand or job name, a missing option value, `install heartbeat`, a `--period` that does not fit the TTL, or the patrol CLI could not be determined | A defect in the call, not a state of the machine. Correct the command line and run it once more; report a second exit 6 as a defect in this skill and stop |
## patrol.sh exit codes
One table for all three subcommands. Read `collect`'s codes carefully: **1 and 3 are results, not aborts.** A round with failed items still has a section to write, and refusing to write it would hide the failure instead of recording it.
| Code | Meaning | What to do |
| --- | --- | --- |
| 0 | `collect`: all four items read to the end, empty sources included. `finish`: heartbeat written, snapshot promoted, lock released. `abort`: lock released | Carry on with the operation's next step |
| 1 | `collect`: partial success — at least one item failed and at least one produced a result | **Write the page anyway.** The section already marks the failed items and the round verdict is 警示. Name the failed items and their `note=` text in the report |
| 2 | `finish`: `jsc-hooks/hooks/heartbeat.sh` was not found | The round completed and is recorded, but no heartbeat exists to prove it. Report the round as recorded and the heartbeat as not written, say `jsc-hooks` 0.3.7 or newer has to be installed, and run `tools/patrol.sh abort --round {id}` to release the lock |
| 3 | `collect`: all four items failed | **Write the page anyway**, with verdict 異常. A page listing four failures is the signal; a missing page is not. Then carry on to `finish` as usual — the round did complete |
| 4 | Another round holds the lock (`collect`), or the lock is no longer this round's (`finish`, `abort`) | Not a failure. On `collect`: report 本輪讓開 and name the holder and its age from the printed `lock=busy` line, then write nothing and stop. On `finish`: the previous round overran and was taken over, so this round's result does not count — report it, write no heartbeat, and stop |
| 5 | Filesystem failure — the lock could not be created or released, a scratch file could not be written, the snapshot could not be promoted, or `heartbeat.sh write` returned non-zero | Serious. Report it loudly with the stderr text and the path. On a `finish` failure the round is recorded but unproven: say so plainly and never claim the round beat |
| 6 | Usage error — an unknown subcommand, a missing `--round`, or an option with no value | A defect in the call. Correct it and run it once more; report a second exit 6 as a defect in this skill and stop |
## Boundaries
The six limits in `AGENTS.md`「助理的界線」 hold for all three operations. Two of them need saying out loud here:
The six limits in `AGENTS.md`「助理的界線」 hold for all four operations. Four of them need saying out loud here:
- **This skill never judges a gate.** It maintains the heartbeat and prints what the heartbeat says. Whether a stale heartbeat blocks a skill call is decided by a hook, synchronously and offline; nothing in this skill blocks or waves through anything.
- **`stop` clearing the heartbeat and removing the schedule is not a breach of 界線 5「不刪除狀態檔」.** That limit protects state that records work — the task book, worktrees, wiki pages — from a background process nobody is watching. The heartbeat records one fact only, "the assistant is alive", and the schedule entry is what keeps writing it, so a `stop` that leaves either behind leaves a lie behind: the next minute cron writes a fresh heartbeat over a stopped assistant. Clearing both is the whole job of `stop`, and they are the only deletions any operation here performs, both of them entries this skill installed itself. `stop` touches nothing under `tasks/`, nobody else's cron entry, no worktree and no wiki page. Do not "restore" this limit later by taking either removal out of `stop`.
- **This skill never judges a gate.** It maintains the heartbeat and prints what the heartbeat says. Whether a stale heartbeat blocks a skill call is decided by a hook, synchronously and offline; nothing in this skill blocks or waves through anything. 界線 2.
- **A patrol round asks nothing.** It runs from cron with nobody present, so there is no one to answer and a question hangs the round. Every branch in the patrol steps below resolves without a question: a missing source is recorded as missing, an ambiguous result is recorded verbatim, and a round that cannot proceed aborts and reports. Never call `jsc-ask:ask` from `patrol`. 界線 1.
- **A patrol round only ever appends to the monitor page.** Read the old page back first, append one section, put the whole page. The contents page gets its own row updated and nobody else's. A page that could not be read is a page that does not get written. 界線 4.
- **A patrol round reports; it never acts on what it found.** The 待人處理 rows name an entry point for a human. The patrol does not run that entry point, does not fix a hook, does not update a plugin and does not touch a repository. 界線 3 and 界線 6.
- **`stop` clearing the heartbeat and removing the schedule is not a breach of 界線 5「不刪除狀態檔」.** That limit protects state that records work — the task book, worktrees, wiki pages — from a background process nobody is watching. The heartbeat records one fact only, "the last patrol round finished", and the schedule entry is what keeps rounds running, so a `stop` that leaves either behind leaves a lie behind. Clearing both is the whole job of `stop`, and they are the only deletions any operation here performs, both of them entries this skill installed itself. `stop` touches nothing under `tasks/`, nobody else's cron entry, no worktree and no wiki page. Do not "restore" this limit later by taking either removal out of `stop`.
## Crash exit needs no cleanup
An assistant that is killed, crashes, or dies with the machine writes no farewell. It does not need to. The heartbeat is a timestamp, not a lock: the last one written stays on disk, ages past the TTL on its own, and every reader from then on sees 過期. No shutdown handler, no cleanup hook and no pid check is involved, so there is nothing left that can fail to run.
That property holds only while nothing fakes a heartbeat. **`write` is called by `start` and by the scheduled entry that keeps a running assistant alive — nowhere else.** `status` never writes one, `stop` never writes one, and no other skill writes one. A heartbeat written by anything that is not a live assistant says a dead assistant is alive, and the reader has no way to tell the difference. This is also why `stop` removes the scheduled entry before clearing the heartbeat, and never in the other order.
The round lock is the one thing a crash does leave behind, and it ages out the same way: `patrol.sh collect` breaks a lock older than the heartbeat TTL, takes it, and prints `lock_broken=1` so the takeover lands on the monitor page instead of happening quietly. The overrun round that lost its lock then gets exit 4 from `finish` and writes no heartbeat, which is correct — it never reached the end.
That property holds only while nothing fakes a heartbeat. **`write` is called by `tools/patrol.sh finish` and nowhere else.** `start` does not call it, `status` does not call it, `stop` does not call it, no scheduled entry calls it, and no other skill calls it. A heartbeat written by anything that is not a finished round says a round finished when none did, and the reader has no way to tell the difference. This is also why `stop` removes the scheduled entry before clearing the heartbeat, and never in the other order.
## start
`start` writes the first heartbeat and installs the heartbeat entry in the system scheduler, so the heartbeat keeps being refreshed once this session is gone. It installs no patrol entry and no daemon.
`start` proves the loop works before it schedules it: one patrol round first, then the scheduled entry. It installs no daemon and writes no bare heartbeat.
1. **Write the first heartbeat.** Run `jsc-hooks/hooks/heartbeat.sh write`. On exit 5 the assistant cannot start: without a heartbeat its own gate reads it as not running, so report the failure, quote the script's stderr line and the heartbeat path, name the likely causes (a full disk, a permission problem on `$JSC_HOME/assistant/`, or something other than a regular file sitting at the heartbeat path), and stop — do not run step 2, and do not report a started assistant. On exit 2 or 6, follow that code's row in the exit-code table and stop. Completion condition: `write` exited 0, or the failure report naming the code and the path has been printed and no start was claimed.
1. **Run one patrol round.** Follow every step of the `patrol` operation below, start to finish. This is what writes the first heartbeat — there is no shortcut past it, because a heartbeat that no round produced is exactly the lie this design removes. When that round ends without a heartbeat for any reason (`collect` exit 4, 5 or 6, an empty `hash=`, a failed wiki write, or `finish` exit 2, 4 or 5), the start has failed: report the round's outcome and the code, do not run step 2, and do not claim a started assistant. A round that completed with failed items (`collect` exit 1 or 3) is still a completed round — carry on to step 2 and name the failures in the closing report. Completion condition: `patrol.sh finish` exited 0, or the failure report naming the step and the code has been printed and no start was claimed.
2. **Confirm what was written.** Run `jsc-hooks/hooks/heartbeat.sh report` and read its `state=`, `ts=`, `ttl=`, `pid=`, `cli=`, `session=` and `file=` fields. `state=fresh` is the expected result. Any other state right after a successful `write` means something rewrote or removed the file in between: report the state, the path and that the heartbeat did not survive its own write, and do not claim a started assistant. Completion condition: the report line was read and either `state=fresh` was recorded with its seven fields, or the mismatch was reported.
2. **Confirm the heartbeat.** Run `jsc-hooks/hooks/heartbeat.sh report` and read its `state=`, `ts=`, `ttl=`, `pid=`, `cli=`, `session=` and `file=` fields. `state=fresh` is the expected result. Any other state right after a successful round means something rewrote or removed the file in between: report the state, the path and that the heartbeat did not survive its own write, and do not claim a started assistant. Completion condition: the report line was read and either `state=fresh` was recorded with its seven fields, or the mismatch was reported.
3. **Install the heartbeat entry.** Run `tools/schedule.sh install heartbeat` — that job name only, never `patrol` and never `all` unless the caller asked for the patrol entry in this same request and was told it fails every round until M-04 lands. Judge the result by the schedule.sh exit-code table, and keep the printed `entry=`, `others_kept=` and `service=` fields for the report. Exit 1 is the case to get right: the entry is installed and inert, so step 4 reports a started assistant whose heartbeat will expire, not a scheduled one. On 2, 3, 4, 5 or 6 nothing is scheduled — report the code, say the heartbeat was written but will expire in one TTL, and do not claim the assistant will stay alive. Completion condition: the exit code is recorded, and on exit 0 the printed entry line and the surviving-entry count are recorded with it.
3. **Install the patrol entry.** Run `tools/schedule.sh install patrol`. Judge the result by the schedule.sh exit-code table, and keep the printed `entry=`, `ttl=`, `period=`, `legacy_removed=`, `others_kept=` and `service=` fields for the report. Exit 1 is the case to get right: the entry is installed and inert, so step 4 reports a started assistant whose heartbeat will expire, not a scheduled one. On 2, 3, 4, 5 or 6 nothing is scheduled — report the code, say the round ran but no further round will, and do not claim the assistant will stay alive. Completion condition: the exit code is recorded, and on exit 0 the printed entry line, the TTL, the period, the legacy count and the surviving-entry count are recorded with it.
4. **Report the start.** Print the heartbeat path, the local time of `ts`, the TTL in seconds, `pid`, `cli` and `session` as hints, then the scheduler mechanism, the installed entry line, and how many other entries were left untouched. Close with the notice that matches step 3's outcome, printed literally with `{ttl}` replaced by the TTL just read:
4. **Report the start.** Print the round's verdict and its four item results, the monitor page that was written, the heartbeat path, the local time of `ts`, the TTL in seconds, `pid`, `cli` and `session` as hints, then the scheduler mechanism, the derived period, the installed entry line, how many legacy heartbeat entries were removed, and how many other entries were left untouched. Close with the notice that matches step 3's outcome, printed literally with `{ttl}` replaced by the TTL just read and `{period}` by the derived period:
| Step 3 | Notice |
| --- | --- |
| exit 0 | 助理已啟動,第一次心跳寫好了,排程也接上了,之後每分鐘寫一次心跳。心跳新鮮只證明排程活著,不證明助理做了事——巡檢本體還沒實作,這一輪沒有裝巡檢那一筆。 |
| exit 1 | 助理已啟動,第一次心跳寫好了,排程條目也寫進去了,但 cron 服務沒在跑,那一筆一次都不會被執行。心跳過了 {ttl} 秒就會過期。請先跑 `sudo service cron start`,重開 WSL 之後要再跑一次。 |
| 其他結束碼 | 助理已啟動,第一次心跳寫好了,但排程沒接上(結束碼 {code})。心跳不會自動更新,過了 {ttl} 秒就會過期,屆時請再跑一次 start。 |
| exit 0 | 助理已啟動,第一輪巡檢跑完了,結果寫上監控頁了,心跳也寫了。排程接上了,之後每 {period} 分鐘跑一輪,每一輪跑完才寫一次心跳。心跳新鮮代表上一輪巡檢真的做完了;那一輪四項有沒有全過,看監控頁的本輪判定。 |
| exit 1 | 助理已啟動,第一輪巡檢跑完了,排程條目也寫進去了,但 cron 服務沒在跑,那一筆一次都不會被執行。心跳過了 {ttl} 秒就會過期。請先跑 `sudo service cron start`,重開 WSL 之後要再跑一次。 |
| 其他結束碼 | 助理已啟動,第一輪巡檢跑完了,但排程沒接上(結束碼 {code})。不會再有下一輪,心跳過了 {ttl} 秒就會過期,屆時請再跑一次 start。 |
Completion condition: the report carries the path, the local heartbeat time, the TTL, the three hint fields and the scheduler outcome, and exactly one notice above appears with the real TTL, and the real code where the row calls for it.
Completion condition: the report carries the round verdict, the monitor page name, the path, the local heartbeat time, the TTL, the period, the three hint fields and the scheduler outcome, and exactly one notice above appears with the real numbers.
## patrol
One round: read four sources, record the result, then beat. Everything before the heartbeat is read-only except the round's own scratch files. Ask nobody anything.
1. **Collect.** Run `tools/patrol.sh collect --trigger 排程` (use `--trigger 手動` when a person asked for this round). Judge the exit code by the patrol.sh table. Exit 4 stands the round down — report the holder and its age from the printed `lock=busy` line, and stop; write no page and no heartbeat. Exit 5 and 6 stop the round the same way, with the code and the stderr text. Exit 0, 1 and 3 all carry on to step 2. Record `round=`, `lock_broken=`, `hash=`, `page=`, `verdict=`, `failed_sources=`, every `item=` line, and the three file paths `section_file=`, `newpage_file=` and `contents_file=`. Completion condition: the round id, the page name and the three file paths are recorded, or the stand-down or the failure was reported and the round stopped.
2. **Check the page name.** An empty `hash=` means `jsc-gitea/tools/hash-id` could not be found or could not run, so there is no page to write to and nothing can be recorded. Run `tools/patrol.sh abort --round {round}`, report that the round found its results but has nowhere to put them, name `jsc-gitea` as missing, and stop. Never invent a page name — a hand-made name lands the content on a page nobody reads. Completion condition: `page=` holds a `MONITOR_{HASH}` name, or the abort ran and the round was reported as unrecorded.
3. **Append the section to `MONITOR_{HASH}`.** Hand it to `jsc-gitea:wiki` with page type `MONITOR`: read the page back first, then append the whole content of `section_file` as a new last section and put the whole page. Only exit 4 from the read permits creating the page instead, and then the page body is the whole content of `newpage_file`, which already carries the basic-data section plus this round's section. Exit 7 and exit 8 mean the old content is unknown: create nothing, write nothing. On any write failure — including exit 3 with no wiki repo configured for `MONITOR`, which the patrol cannot ask about — run `tools/patrol.sh abort --round {round}`, report the code, and stop. **No record, no heartbeat.** Completion condition: the append or the create returned success, or the abort ran and the round was reported as unrecorded with its exit code.
4. **Update this machine's row in `MONITOR_CONTENTS`.** Take the `row=` line from `contents_file` — it is already the finished table row. Hand it to `jsc-gitea:wiki`: read the whole page, match the row whose 主機 and 帳號 columns both equal this round's `host=` and `user=`, overwrite that row's remaining columns, and put the whole page back. No matching row means append one. Never overwrite the whole page, and never touch another machine's row — the write semantics here are the opposite of the content page's, and mixing them up deletes other machines' records. On failure, run `tools/patrol.sh abort --round {round}`, report the code, and stop. Completion condition: exactly one row carries this machine's 主機 and 帳號 values, every other row is byte-identical to what was read, and the put returned success.
5. **Write the heartbeat.** Run `tools/patrol.sh finish --round {round}`. This is the last step for a reason: it is the only thing that turns a fresh heartbeat into a true statement. Judge the exit code by the patrol.sh table — 2, 4 and 5 all mean the round is recorded but unproven, and each has its own report line there. Completion condition: `finish` exited 0, or the failure was reported as "recorded but no heartbeat" with its code.
6. **Report the round.** Print the round verdict, one line per item with its `status=` and, for a failure, its `note=`; the monitor page name and the contents row that was written; whether the heartbeat was written; and, when `lock_broken=1`, that the previous round's lock was taken over because it had aged past the TTL. Close with the 待人處理 rows from the section, verbatim, and nothing else — the patrol names an entry point and stops there. Completion condition: all four items appear in the report, the heartbeat outcome is stated as written or not written, and no suggestion in 待人處理 was acted on.
## status
@@ -127,19 +177,20 @@ Read-only throughout. This operation creates, modifies and deletes nothing under
Completion condition: every file under `tasks/` produced exactly one row, or zero entries was reported.
3. **Read the schedule.** Run `tools/schedule.sh status`. It writes nothing. Record `mechanism=`, `service=` and the `installed=` value of both jobs. On exit 2, 3 or 6 nothing was read: record the schedule state as unknown with its code and carry on — the heartbeat and the task book still print. Completion condition: both jobs have an installed state, or the schedule state is recorded as unknown with its code.
3. **Read the schedule.** Run `tools/schedule.sh status`. It writes nothing. Record `mechanism=`, `service=`, `ttl=`, `period=` and the `installed=` value of both jobs. A `heartbeat` job reported as installed is a leftover from an older version: say so, and say `start` or `schedule.sh install patrol` removes it. On exit 2, 3 or 6 nothing was read: record the schedule state as unknown with its code and carry on — the heartbeat and the task book still print. Completion condition: both jobs have an installed state, or the schedule state is recorded as unknown with its code.
4. **Print the status table.** Lead with the heartbeat block — verdict, last heartbeat time rendered from `ts` in local time, age in seconds, TTL, `cli`, `session`, `pid`, and the task count. Follow it with the schedule block — mechanism, service state, and one line per job saying installed or not. Then one row per task carrying `state`, `title`, `next_run` and `fail_count`, in the order the files were listed. Completion condition: the heartbeat block holds all eight values, the schedule block holds both jobs, and the row count equals the task count from step 2.
4. **Print the status table.** Lead with the heartbeat block — verdict, last heartbeat time rendered from `ts` in local time, age in seconds, TTL, `cli`, `session`, `pid`, and the task count. Follow it with the schedule block — mechanism, service state, derived period, and one line per job saying installed or not. Then one row per task carrying `state`, `title`, `next_run` and `fail_count`, in the order the files were listed. Completion condition: the heartbeat block holds all eight values, the schedule block holds both jobs and the period, and the row count equals the task count from step 2.
5. **Say what the two blocks together mean.** Three combinations get an explicit sentence, because each one reads as something it is not:
5. **Say what the two blocks together mean.** Four combinations get an explicit sentence, because each one reads as something it is not:
| Heartbeat | Schedule | Say |
| --- | --- | --- |
| 新鮮 | heartbeat installed, service running | 排程活著,心跳是它寫的。這不代表助理做了事,做了什麼看下面的待辦表 |
| 新鮮 | not installed, or service stopped | 心跳還新鮮,但沒有排程在維持它,過了 TTL 就會過期 |
| 過期 or 不存在 | heartbeat installed, service running | 排程裝著卻沒有心跳,排程那一筆自己失敗了,去看 `$JSC_HOME/assistant/schedule.log` |
| 新鮮 | patrol installed, service running | 上一輪巡檢跑完了,結果也記上監控頁了,排程還在跑。那一輪四項有沒有全過,要看監控頁的本輪判定 |
| 新鮮 | not installed, or service stopped | 上一輪巡檢跑完了,但沒有排程在叫下一輪,過了 TTL 心跳就會過期 |
| 過期 or 不存在 | patrol installed, service running | 排程裝著卻沒有新的心跳,巡檢自己跑失敗了,去看 `$JSC_HOME/assistant/schedule.log` 與監控頁最新一節 |
| any | `heartbeat` job installed | 舊版的心跳排程還留著,它會蓋掉「心跳等於巡檢跑完」這件事。請跑一次 `start`,或 `schedule.sh install patrol` 把它清掉 |
Completion condition: the matching sentence is printed, or none of the three combinations applied.
Completion condition: every matching sentence is printed, or none of the four combinations applied.
6. **Flag the repeatedly failing tasks.** Append 已連續失敗 N 次 to every row whose `fail_count` is above 0, with `N` taken verbatim from the file. A broken entry that retries every round with nobody noticing is the reason this field exists, so let no such row leave the table unmarked. Completion condition: every row with `fail_count` above 0 carries the marker and its number matches the file.
@@ -147,14 +198,21 @@ Read-only throughout. This operation creates, modifies and deletes nothing under
## stop
1. **Record what is being stopped.** Run `jsc-hooks/hooks/heartbeat.sh report` first and keep its `state=`, `ts=`, `pid=`, `cli=` and `file=` fields for the closing report — after the clear they are gone for good. `state=absent` means the assistant was already stopped; say so and still run steps 2 and 3, because a scheduled entry can outlive its heartbeat and `clear` on a missing file is a success, so running both leaves the outcome unambiguous. On exit 2 or 6, follow that code's row, record the previous state as unknown, and carry on to step 2. Completion condition: the previous state and its fields are recorded, or the previous state is recorded as unknown with its code.
1. **Record what is being stopped.** Run `jsc-hooks/hooks/heartbeat.sh report` first and keep its `state=`, `ts=`, `pid=`, `cli=` and `file=` fields for the closing report — after the clear they are gone for good. `state=absent` means no round has finished; say so and still run steps 2 and 3, because a scheduled entry can outlive its heartbeat and `clear` on a missing file is a success, so running both leaves the outcome unambiguous. On exit 2 or 6, follow that code's row, record the previous state as unknown, and carry on to step 2. Completion condition: the previous state and its fields are recorded, or the previous state is recorded as unknown with its code.
2. **Remove the schedule first.** Run `tools/schedule.sh remove all` — both jobs, so a patrol entry somebody installed by hand goes too. This comes before the clear and never after: clear first and the next cron minute writes a fresh heartbeat over the stopped assistant, and every reader from then on is told a dead assistant is alive. Judge the result by the schedule.sh exit-code table, and keep `removed=` and `others_kept=` for the report. On any non-zero code the schedule is still installed: report the code, say plainly that the entry will keep writing heartbeats and the assistant therefore cannot be stopped, name the manual fix (`crontab -l` to look, then remove the line carrying `# jsc-assist:assistant` by hand), and skip steps 3 and 4 — clearing a heartbeat that cron rewrites a minute later only hides the problem. Completion condition: `remove` exited 0 with its counts recorded, or the failure report has been printed and no stop was claimed.
2. **Remove the schedule first.** Run `tools/schedule.sh remove all` — both job names, so the patrol entry and any leftover heartbeat entry from an older install both go. This comes before the clear and never after: clear first and the next scheduled round writes a fresh heartbeat over the stopped assistant, and every reader from then on is told a dead assistant is alive. Judge the result by the schedule.sh exit-code table, and keep `removed=` and `others_kept=` for the report. On any non-zero code the schedule is still installed: report the code, say plainly that rounds will keep running and the assistant therefore cannot be stopped, name the manual fix (`crontab -l` to look, then remove the line carrying `# jsc-assist:assistant` by hand), and skip steps 3 and 4 — clearing a heartbeat that the next round rewrites only hides the problem. Completion condition: `remove` exited 0 with its counts recorded, or the failure report has been printed and no stop was claimed.
3. **Clear the heartbeat.** Run `jsc-hooks/hooks/heartbeat.sh clear`. On exit 5 the file is still there: report the failure with the script's stderr line and the path, say plainly that every reader still sees a heartbeat claiming the assistant is running and that the assistant is therefore not reliably stopped, name the manual fix (delete that path by hand, then run `status` to confirm `助理未運行`), and skip step 4 — the closing notice must not be printed after a failed clear. On exit 2 or 6, follow that code's row and stop the same way. Completion condition: `clear` exited 0, or the failure report naming the code, the path and the manual fix has been printed and no stop was claimed.
3. **Clear the heartbeat.** Run `jsc-hooks/hooks/heartbeat.sh clear`. On exit 5 the file is still there: report the failure with the script's stderr line and the path, say plainly that every reader still sees a heartbeat claiming a round just finished and that the assistant is therefore not reliably stopped, name the manual fix (delete that path by hand, then run `status` to confirm `助理未運行`), and skip step 4 — the closing notice must not be printed after a failed clear. On exit 2 or 6, follow that code's row and stop the same way. Completion condition: `clear` exited 0, or the failure report naming the code, the path and the manual fix has been printed and no stop was claimed.
4. **Report the stop and what it means for the gate.** Print the previous state and heartbeat time from step 1 and the entries removed in step 2, then this literally:
> 助理已停止,排程移除了,心跳也清掉了,其他人的排程一筆都沒動。靠心跳判定的 jsc 技能閘門一讀到沒有心跳就會擋下技能呼叫;閘門目前還沒接線,所以這一刻誰都擋不到。要再工作就先跑一次 start。
Say it exactly this way. The blocking is the designed consequence of a cleared heartbeat, and whoever stops the assistant has to know it is coming; the clause about the gate being unwired is the part that keeps the notice honest while that is still true. When the gate is wired, that clause is what gets rewritten — not the rest. Completion condition: the notice appears with all three clauses, and the previous state, the heartbeat time and the removal counts are printed above it.
## Round lock and a round that will not stop leaving one behind
A patrol round holds `$JSC_HOME/assistant/patrol.lock` from `collect` to `finish` or `abort`, which spans the wiki writes — the slow part. Two consequences bind every branch above:
- **Every path out of a started round ends in `finish` or `abort`.** Steps 2, 3 and 4 of `patrol` each name their abort. A round that stops without either leaves the lock standing until it ages out, which stands the next rounds down for up to one TTL. There is no third option.
- **`stop` does not remove the lock.** It is not a state file that records work, but it is also not this skill's to delete while a round may still be using it; it ages out on its own within one TTL. If an operator reports that every round stands down, tell them the holder and age from the `lock=busy` line and let them decide — 界線 5 keeps destructive cleanup with the human.
+2 -2
View File
@@ -15,8 +15,8 @@
| 監控頁 | 指向 `MONITOR_{HASH}` 的同 wiki 連結 | 少了連結就要人自己算雜湊才翻得到內容頁 |
| 主機 | 這台機器的主機名,與雜湊第一段相同 | 比對用的兩欄之一,決定要更新哪一列 |
| 帳號 | 助理執行時的登入帳號,與雜湊第二段相同 | 比對用的兩欄之一。同一台機器換帳號就是另一個巡檢對象 |
| 心跳 | 巡檢當下的心跳判定,判準只看 `ts` 距現在是否不到 300 秒 | 一眼看出這台機器的助理還在不在跑,不必逐頁翻 |
| 最後巡檢 | 該頁最新一節的時間戳 | 心跳新鮮而這一欄很舊,代表助理活著卻沒在巡 |
| 心跳 | 巡檢當下(本輪寫入前)的心跳判定,判準只看 `ts` 距現在有沒有超過門檻,預設 300 秒 | 一眼看出這台機器上一輪巡檢有沒有跑完,不必逐頁翻 |
| 最後巡檢 | 該頁最新一節的時間戳 | 心跳由巡檢寫,兩欄理當一致;差很多就代表有一輪寫了心跳卻沒寫頁,那是缺陷 |
| 待辦筆數 | 待辦簿現有筆數 | 心跳新鮮而筆數為 0,代表助理空轉,沒有東西可跑 |
| 連續失敗項 | 該頁最新一節裡 `fail_count` 大於 0 的筆數 | 待辦簿的項目失敗不會自動暫停,每輪都重試。這一欄讓壞掉的項目在目錄頁就現形 |
+10 -3
View File
@@ -7,12 +7,15 @@
```mermaid
flowchart LR
A[巡檢一輪] --> B[收攏六類結果]
A[巡檢一輪] --> B[收攏各類結果]
B --> C[附加一節,節標題帶時間戳]
C --> D[既有的節原樣保留]
D --> E[回頭更新 MONITOR_CONTENTS 自己那一列]
E --> F[最後才寫心跳]
```
心跳排在最後一步,不能提前。心跳新鮮的意思就是「上一輪跑到這一步了」:這一節沒寫上來,心跳就不寫,讓它自己過期。那是巡檢在空轉的唯一訊號。
## 本頁基本資料
建頁時寫一次,之後不再更動。
@@ -26,13 +29,14 @@ flowchart LR
## 巡檢 {yyyy-MM-dd HH:mm}
一輪巡檢就是這樣一節,最新的一節放在最下面。六個子節固定都寫;某個來源讀不到,就在那個子節寫明是哪個路徑讀不到,不要整節略過。
一輪巡檢就是這樣一節,最新的一節放在最下面。六個子節固定都寫;某個來源讀不到,就在那個子節寫明是哪個路徑讀不到,不要整節略過。還沒實作的子節也照寫,寫明「這一輪不做這一項」——空表格會被讀成「查過了,沒問題」。
| 項目 | 內容 |
| --- | --- |
| 巡檢時間 | {yyyy-MM-dd HH:mm} |
| 觸發方式 | {排程、事件、手動 三選一} |
| 本輪判定 | {正常、警示、異常 三選一} |
| 本輪項目 | {這一輪跑了哪幾項,成功幾項、失敗幾項} |
| 讀不到的來源 | {路徑清單,全部讀得到就寫「無」} |
### 心跳與閘門狀態
@@ -45,7 +49,9 @@ flowchart LR
| session | {工作階段代號} |
| pid | {數字}。只給要找行程的人參考,不參與判定 |
心跳的判準只看 `ts` 距現在是否不到 300 秒。不看 pid 存活:五支 CLI 與容器裡的行程互相看不到彼此的 pid。閘門的判定留在 hook,助理只維持心跳。
這一欄讀到的是**上一輪**巡檢寫的心跳:心跳由巡檢寫,本輪那一次要等這一節寫上來之後才寫。
心跳的判準只看 `ts` 距現在有沒有超過門檻,預設 300 秒。不看 pid 存活:五支 CLI 與容器裡的行程互相看不到彼此的 pid。閘門的判定留在 hook,助理只維持心跳。心跳新鮮代表上一輪巡檢跑完了,不代表那一輪四項都成功——那要看本節上面的「本輪判定」。
### 技能與呼叫鏈使用統計
@@ -112,3 +118,4 @@ flowchart LR
- 禁止整頁覆寫。覆寫等於把這台機器的巡檢軌跡刪掉。
- 讀不到舊內容就中止,不附加,也不寫入。
- 附加成功之後,才回頭更新 `MONITOR_CONTENTS` 自己那一列。
- 兩頁都寫成之後,才寫這一輪的心跳。任一頁沒寫成就不寫心跳,讓它過期。
+739
View File
@@ -0,0 +1,739 @@
#!/usr/bin/env sh
# patrol.sh — 助理巡檢一輪的收攏與收口(供 jsc-assist:assistant 的 patrol 操作呼叫)。
#
# 用法:
# patrol.sh collect [--out {目錄}] [--trigger {排程|事件|手動}]
# patrol.sh finish --round {輪次代號} [--out {目錄}] [--dry-run]
# patrol.sh abort --round {輪次代號} [--out {目錄}]
#
# collect 帶了 --out,finish 與 abort 就要帶同一個目錄,不然換不到本輪的用量快照。
#
# collect 讀四項來源、組出監控頁要附加的那一節、把鎖拿在手上。
# finish 在監控頁寫成功之後才呼叫:寫心跳、換上用量快照、放掉鎖。
# abort 在監控頁沒寫成時呼叫:只放掉鎖,不寫心跳。
#
# 結束碼(三個子命令共用一張表,同一碼在不同子命令的成因寫在同一列):
# 0 collect:四項全部讀到底(含「來源在、沒有資料」);finish:心跳寫好、快照換上、
# 鎖放掉;abort:鎖放掉,本來就沒鎖也算
# 1 collect:部分成功——至少一項失敗,也至少一項有結果。**結果照樣印得出來,呼叫端
# 照樣要把這一節寫上監控頁**,只是本輪判定要標成警示
# 2 finish:找不到 jsc-hooks 的 hooks/heartbeat.sh,心跳沒有東西可寫。collect 不會回這
# 一碼——心跳讀不到只是 D-09 這一項失敗,另外三項照跑
# 3 collect:四項全部失敗,一項資料都沒有。這一節還是要寫上監控頁,本輪判定標成異常
# 4 上一輪還在跑,本輪讓開(collect),或鎖已經不在自己手上(finish、abort)。這不是
# 失敗,是刻意讓開:不寫心跳、不寫監控頁,下一輪再來
# 5 檔案系統失敗:鎖建不起來或放不掉、暫存檔寫不進去、快照換不上,或 heartbeat.sh write
# 回非 0。心跳沒寫成就是沒寫成,一律吵出來
# 6 用法錯誤:不認得的子命令、缺 --round、參數缺值
#
# --- 心跳為什麼由這裡寫,不由排程直接寫 ---
#
# 排程每分鐘直接呼叫 heartbeat.sh write 的話,心跳新鮮只證明 cron 活著。巡檢整個壞掉、
# 一項資料都讀不到、監控頁一頁都沒寫成,心跳照樣新鮮,靠心跳判定的閘門照樣放行,沒有
# 任何訊號。所以心跳改由巡檢寫:跑完一輪、而且結果真的記下來了,才寫那一次心跳。
#
# --- 心跳寫不寫,只看結果有沒有記下來 ---
#
# 四項的成敗不決定心跳。四項全失敗但監控頁寫成了,那一輪還是跑完了,證據也留下來了,
# 心跳照寫,頁上判定是異常,看頁的人自己判斷。反過來,監控頁沒寫成就是這一輪沒有結果,
# 心跳一定不寫:讓它自己過期,就是「巡檢在空轉」的唯一訊號。
# 所以寫心跳一定是獨立的 finish,時序上排在監控頁寫成之後,不與 collect 綁在一起。
#
# --- 上一輪還沒跑完,下一輪被叫起來 ---
#
# 一輪巡檢包含一次 CLI 呼叫與兩次 wiki 寫入,跑過一個排程週期是有可能的。所以整輪拿一把
# 鎖:$JSC_HOME/assistant/patrol.lock 是目錄,mkdir 是原子操作,搶不到就是別人在跑。
# 搶不到的那一輪回 4 直接讓開,不排隊、不並行——並行的兩輪會在同一頁附加兩節,還會互相
# 蓋掉用量快照。
# 鎖會逾時自動搶回來,門檻取心跳門檻(heartbeat.sh report 的 ttl 欄):上一輪跑得比門檻
# 還久,它本來就已經維持不住心跳新鮮了,讓新的一輪接手才對。被搶回來的那一輪,finish 會
# 拿 --round 比對出鎖不是自己的,回 4 且不寫心跳。
# 搶回來這件事會記在監控頁上(lock_broken=1),不會安靜發生。
#
# --- 這四項都是純讀取 ---
#
# D-01 技能與呼叫鏈使用統計 jsc-log 的 tools/usage-stats.sh
# D-04 版本落差與重啟閘門 jsc-hooks 的 version-guard.sh report、restart-gate.sh report
# D-07 SDLC 階段鎖與工作包鎖 $JSC_HOME/sessions/*.stage、$JSC_HOME/wp/*.pr
# D-09 心跳與閘門狀態自述 jsc-hooks 的 heartbeat.sh report
# 四項各自獨立:一項的來源不見了、或回非 0,只讓那一項標成失敗,其餘三項照跑、照記。
# 四項都不呼叫別的技能、不寫程式碼存取庫、不做決策。
#
# --- version-guard.sh report 的既有缺陷照實記 ---
#
# 這支腳本在部分機器上對每一個 domain 都回「查詢失敗」,recommend 跟著回 unverifiable。
# 巡檢照抄第四欄原字,不自己補查遠端版本、不把「查詢失敗」寫成「相符」或「最新」。
# 查不到就是沒有證據,寫成別的字等於幫一個既有缺陷蓋章。
#
# --- collect 的輸出 ---
#
# stdout 是 key=value,一行一個鍵,供呼叫端逐行取值。監控頁要用的 markdown 不印在
# stdout,改寫成檔案再把路徑印出來:那一段有好幾百字,經過對話重打一次只會多錯字。
# round= 本輪代號,finish 與 abort 要原樣帶回
# lock= acquired
# lock_broken= 0 或 1。1 代表上一輪的鎖逾時被搶回來
# hash= 監控頁雜湊,來源是 {主機名}/{登入帳號};算不出來時為空
# page= MONITOR_{HASH};hash 為空時為空
# host= user= at= 主機名、登入帳號、本輪時間
# item= 一項一行,欄位 status(ok、empty、fail)、rc、note
# verdict= 正常、警示、異常
# failed_sources= 讀不到的來源路徑,以「、」分隔;全部讀得到就是「無」
# tasks_total= tasks_failing= 待辦簿筆數與連續失敗筆數,只供目錄頁那一列用
# section_file= 要附加到 MONITOR_{HASH} 的那一節
# newpage_file= MONITOR_{HASH} 不存在時要建的整頁內容
# contents_file= MONITOR_CONTENTS 那一列的欄位值
#
# 環境變數:
# JSC_HOME 助理狀態檔的根目錄,預設 ~/.jsc
# JSC_ASSIST_HEARTBEAT_SH heartbeat.sh 路徑覆寫;找不到並排存取庫時才設
# JSC_ASSIST_VERSION_GUARD_SH version-guard.sh 路徑覆寫
# JSC_ASSIST_RESTART_GATE_SH restart-gate.sh 路徑覆寫
# JSC_ASSIST_USAGE_STATS_SH usage-stats.sh 路徑覆寫
# JSC_ASSIST_HASH_ID hash-id 路徑覆寫
# JSC_ASSIST_PATROL_LOCK_TTL 鎖的逾時秒數;未設定時取心跳門檻,取不到就用 300
set -u
JSC_HOME="${JSC_HOME:-$HOME/.jsc}"
STATE_DIR="$JSC_HOME/assistant"
LOCK="$STATE_DIR/patrol.lock"
PREV_SNAP="$STATE_DIR/usage-prev.tsv"
RD="$STATE_DIR/patrol"
SCRIPT_DIR=$(CDPATH= cd -- "$(dirname -- "$0")" 2>/dev/null && pwd)
SCRIPT_DIR="${SCRIPT_DIR:-.}"
TRIGGER=''
ROUND=''
DRYRUN=0
LOCK_BROKEN=0
FAILED_SOURCES=''
OK_COUNT=0
FAIL_COUNT=0
WARN=0
usage() {
cat >&2 <<'EOF'
usage: patrol.sh collect [--out 目錄] [--trigger 排程|事件|手動]
patrol.sh finish --round 輪次代號 [--out 目錄] [--dry-run]
patrol.sh abort --round 輪次代號 [--out 目錄]
EOF
exit 6
}
die() { # $1=結束碼 $2=訊息
printf '[jsc][助理巡檢][ERR]:%s\n' "$2" >&2
exit "$1"
}
# 找一支別的 domain 的腳本。搜尋順序比照 schedule.sh 的 heartbeat_sh():先環境變數覆寫,
# 再開發用的並排存取庫版面,最後已安裝的 plugin 快取版面。
find_tool() { # $1=domain 短名 $2=domain 內相對路徑 $3=環境變數覆寫值(可為空)
if [ -n "$3" ]; then
[ -f "$3" ] && { printf '%s\n' "$3"; return 0; }
return 1
fi
_root="${CLAUDE_PLUGIN_ROOT:-$SCRIPT_DIR/..}"
for _c in "$_root/../$1/$2" "$_root/../jsc-$1/$2"; do
[ -f "$_c" ] && { printf '%s\n' "$_c"; return 0; }
done
_c=$(ls -d "$_root"/../../jsc-"$1"/*/"$2" \
"$_root"/../../"$1"/*/"$2" \
"$HOME"/.claude/plugins/cache/*/jsc-"$1"/*/"$2" 2>/dev/null | sort | tail -n1)
[ -n "$_c" ] && [ -f "$_c" ] && { printf '%s\n' "$_c"; return 0; }
return 1
}
fmt_ts() { # $1=epoch 秒數;轉成當地時間字串,轉不動就原樣印
date -d "@$1" '+%Y-%m-%d %H:%M' 2>/dev/null && return 0
date -r "$1" '+%Y-%m-%d %H:%M' 2>/dev/null && return 0
printf '%s\n' "$1"
}
mtime_of() { # $1=檔案;印出修改時間,取不到印「-」
_t=$(stat -c %Y "$1" 2>/dev/null) || _t=''
[ -n "$_t" ] || _t=$(stat -f %m "$1" 2>/dev/null) || _t=''
[ -n "$_t" ] || { printf '-'; return 0; }
fmt_ts "$_t" | tr -d '\n'
}
# markdown 表格欄位裡的 `|` 會把欄切開,一律跳脫;換行壓成空白。
cell() { printf '%s' "$1" | tr '\n' ' ' | sed 's/|/\\|/g'; }
# 記一個讀不到的來源。同一輪多項失敗就串起來,供監控頁「讀不到的來源」那一列用。
add_failed_source() { # $1=路徑或來源名稱
if [ -z "$FAILED_SOURCES" ]; then FAILED_SOURCES="$1"; else FAILED_SOURCES="$FAILED_SOURCES、$1"; fi
}
# 記一列待人處理。助理只提醒,不代為執行。
add_pending() { # $1=要處理什麼 $2=來源子節 $3=建議入口
printf '| %s | %s | %s |\n' "$(cell "$1")" "$(cell "$2")" "$(cell "$3")" >>"$RD/pend.md"
}
# --- 鎖 ---
lock_ttl() { # 鎖的逾時秒數
_t="${JSC_ASSIST_PATROL_LOCK_TTL:-}"
case "$_t" in ''|*[!0-9]*) _t='' ;; esac
[ -n "$_t" ] && { printf '%s' "$_t"; return 0; }
[ -n "${HEARTBEAT_TTL:-}" ] && { printf '%s' "$HEARTBEAT_TTL"; return 0; }
printf '300'
}
lock_round() { sed -n 's/^round=//p' "$LOCK/info" 2>/dev/null | head -n1; }
lock_age() {
_s=$(sed -n 's/^started=//p' "$LOCK/info" 2>/dev/null | head -n1)
case "$_s" in ''|*[!0-9]*) printf '999999'; return 0 ;; esac
printf '%s' "$(( $(date +%s) - _s ))"
}
# 搶鎖。搶到回 0,別人在跑回 4,建不起來回 5。
lock_acquire() {
mkdir -p "$STATE_DIR" 2>/dev/null || die 5 "建不出助理狀態目錄 $STATE_DIR。"
if mkdir "$LOCK" 2>/dev/null; then
printf 'round=%s\npid=%s\nstarted=%s\n' "$ROUND" "$$" "$(date +%s)" >"$LOCK/info" 2>/dev/null \
|| die 5 "鎖建起來了,卻寫不進 $LOCK/info。"
return 0
fi
_age=$(lock_age); _ttl=$(lock_ttl)
if [ "$_age" -lt "$_ttl" ]; then
printf 'lock=busy holder=%s age=%s ttl=%s\n' "$(lock_round)" "$_age" "$_ttl"
die 4 "上一輪巡檢還在跑(輪次 $(lock_round),已經跑了 $_age 秒,未達 $_ttl 秒門檻),本輪讓開。"
fi
# 逾時搶回來。上一輪跑得比心跳門檻還久,它已經維持不住心跳新鮮了,讓新的一輪接手。
rm -rf "$LOCK" 2>/dev/null
mkdir "$LOCK" 2>/dev/null || die 5 "上一輪的鎖逾時,卻搶不回來:$LOCK。"
printf 'round=%s\npid=%s\nstarted=%s\n' "$ROUND" "$$" "$(date +%s)" >"$LOCK/info" 2>/dev/null \
|| die 5 "鎖搶回來了,卻寫不進 $LOCK/info。"
LOCK_BROKEN=1
WARN=1
return 0
}
# 放鎖。鎖不在自己手上就回 4,絕不硬放——那會把正在跑的那一輪的鎖拆掉。
lock_release() {
[ -d "$LOCK" ] || { printf 'lock=absent\n'; return 0; }
_h=$(lock_round)
[ "$_h" = "$ROUND" ] || die 4 "鎖不在本輪手上(鎖的輪次是 ${_h:-空值},本輪是 $ROUND),不動它,也不寫心跳。"
rm -rf "$LOCK" 2>/dev/null || die 5 "鎖放不掉:$LOCK。"
printf 'lock=released\n'
return 0
}
# --- D-01 技能與呼叫鏈使用統計 ---
# 把 usage-stats.sh 的「次數<TAB>名稱」轉成監控頁的列,順便算出本輪增量。
# 增量要有上一輪的累計快照才算得出來;沒有快照的第一輪一律寫「-」,不拿累計冒充本輪。
usage_rows() { # $1=類別(技能、呼叫鏈) $2=快照鍵(skill、chain) $3=資料檔
while IFS=' ' read -r _cnt _name; do
[ -n "$_name" ] || continue
printf '%s\t%s\t%s\n' "$2" "$_name" "$_cnt" >>"$RD/usage-next.tsv"
_prev=''
[ -f "$PREV_SNAP" ] && _prev=$(awk -F'\t' -v k="$2" -v n="$_name" \
'$1 == k && $2 == n { print $3; exit }' "$PREV_SNAP" 2>/dev/null)
if [ -n "$_prev" ]; then _delta=$(( _cnt - _prev )); else _delta='-'; fi
printf '| %s | %s | %s | %s |\n' "$(cell "$_name")" "$1" "$_delta" "$_cnt"
done <"$3"
}
d01() {
D01_STATUS=fail; D01_RC=0; D01_NOTE=''
: >"$RD/usage-next.tsv"
{
printf '### 技能與呼叫鏈使用統計\n\n'
printf '資料出自 `%s` 與 `%s`,由 `jsc-log` 的 `tools/usage-stats.sh` 聚合。\n\n' \
"\$JSC_HOME/usage/skills.jsonl" "\$JSC_HOME/usage/chains.jsonl"
} >"$RD/d01.md"
if ! _us=$(find_tool log tools/usage-stats.sh "${JSC_ASSIST_USAGE_STATS_SH:-}"); then
D01_NOTE='找不到 jsc-log 的 tools/usage-stats.sh'
D01_RC=127
add_failed_source 'jsc-log/tools/usage-stats.sh'
printf '**這一項失敗**:%s。這一輪沒有使用統計,不是「零次」。\n' "$D01_NOTE" >>"$RD/d01.md"
add_pending '技能用量讀不到,jsc-log 沒裝或版本太舊' '技能與呼叫鏈使用統計' '/jsc-cli:doctor'
return 0
fi
_rc1=0; _rc2=0
"$_us" skills >"$RD/d01.skills" 2>"$RD/d01.err" || _rc1=$?
"$_us" chains >"$RD/d01.chains" 2>>"$RD/d01.err" || _rc2=$?
if [ "$_rc1" -ne 0 ] || [ "$_rc2" -ne 0 ]; then
D01_RC=$(( _rc1 > _rc2 ? _rc1 : _rc2 ))
D01_NOTE="usage-stats.sh 回非 0(skills=$_rc1、chains=$_rc2)"
add_failed_source "$_us"
printf '**這一項失敗**:%s。訊息:%s\n' "$D01_NOTE" "$(cell "$(cat "$RD/d01.err" 2>/dev/null)")" >>"$RD/d01.md"
add_pending '使用統計腳本回非 0' '技能與呼叫鏈使用統計' '/jsc-log:stats'
return 0
fi
{
printf '| 對象 | 類別 | 本輪次數 | 累計次數 |\n'
printf '| --- | --- | ---: | ---: |\n'
usage_rows '技能' skill "$RD/d01.skills"
usage_rows '呼叫鏈' chain "$RD/d01.chains"
} >"$RD/d01.rows"
_n=$(grep -c '^| ' "$RD/d01.rows" 2>/dev/null); [ -n "$_n" ] || _n=0
if [ "$_n" -le 2 ]; then
D01_STATUS=empty
printf '來源檔讀得到,本輪一筆用量都沒有。hook 還沒記過任何一次呼叫就是這個狀態,那是「零次」,不是故障。\n' >>"$RD/d01.md"
return 0
fi
D01_STATUS=ok
cat "$RD/d01.rows" >>"$RD/d01.md"
if [ ! -f "$PREV_SNAP" ]; then
printf '\n第一輪沒有上一輪的累計快照,「本輪次數」一律寫「-」,不拿累計冒充本輪。\n' >>"$RD/d01.md"
fi
return 0
}
# --- D-04 版本落差與重啟閘門 ---
d04() {
D04_STATUS=fail; D04_RC=0; D04_NOTE=''
{
printf '### 版本落差與重啟閘門\n\n'
printf '資料出自 `version-guard.sh report` 與 `restart-gate.sh report`。\n\n'
} >"$RD/d04.md"
_vg=''; _rg=''
find_tool hooks hooks/version-guard.sh "${JSC_ASSIST_VERSION_GUARD_SH:-}" >"$RD/vg" 2>/dev/null \
&& _vg=$(cat "$RD/vg")
find_tool hooks hooks/restart-gate.sh "${JSC_ASSIST_RESTART_GATE_SH:-}" >"$RD/rg" 2>/dev/null \
&& _rg=$(cat "$RD/rg")
if [ -z "$_vg" ] && [ -z "$_rg" ]; then
D04_NOTE='找不到 jsc-hooks 的 version-guard.sh 與 restart-gate.sh'
D04_RC=127
add_failed_source 'jsc-hooks/hooks/version-guard.sh、jsc-hooks/hooks/restart-gate.sh'
printf '**這一項失敗**:%s。這一輪沒有版本證據,不是「版本都相符」。\n' "$D04_NOTE" >>"$RD/d04.md"
add_pending '版本與重啟閘門讀不到,jsc-hooks 沒裝或版本太舊' '版本落差與重啟閘門' '/jsc-cli:doctor'
return 0
fi
_part=0
# 版本比對表。第四欄照 version-guard.sh 原字抄,不改寫、不補查。
if [ -n "$_vg" ]; then
_rc=0
"$_vg" report >"$RD/d04.vg" 2>"$RD/d04.err" </dev/null || _rc=$?
if [ "$_rc" -ne 0 ]; then
_part=1; D04_RC="$_rc"
D04_NOTE="version-guard.sh report 回 $_rc"
add_failed_source "$_vg"
printf '**版本比對失敗**:`version-guard.sh report` 回 %s。訊息:%s\n\n' \
"$_rc" "$(cell "$(cat "$RD/d04.err" 2>/dev/null)")" >>"$RD/d04.md"
else
{
printf '| domain | 本機版本 | 應有版本 | 判定 |\n'
printf '| --- | --- | --- | --- |\n'
} >>"$RD/d04.md"
_unver=0; _rows=0
while IFS=' ' read -r _c1 _c2 _c3 _c4; do
[ -n "$_c1" ] || continue
case "$_c1" in
behind) printf '| (合計) | - | - | 落後 %s 個 |\n' "$(cell "$_c2")" >>"$RD/d04.md"; continue ;;
noregistry)
printf '| (無註冊檔) | - | - | 查不到本機已裝的 plugin:%s |\n' "$(cell "$_c2")" >>"$RD/d04.md"
_unver=1; continue ;;
esac
_rows=$(( _rows + 1 ))
printf '| %s | %s | %s | %s |\n' "$(cell "$_c1")" "$(cell "$_c2")" "$(cell "$_c3")" "$(cell "$_c4")" >>"$RD/d04.md"
[ "$_c4" = '查詢失敗' ] && _unver=$(( _unver + 1 ))
done <"$RD/d04.vg"
[ "$_rows" -eq 0 ] && printf '| (無 domain) | - | - | 這台機器一個 jsc plugin 都沒查到 |\n' >>"$RD/d04.md"
printf '\n判定欄照 `version-guard.sh report` 第四欄原字抄。抄到「查詢失敗」就寫「查詢失敗」,不改寫成「相符」或「最新」,也不自己補查遠端版本——查不到是沒有證據,不是版本沒問題。\n\n' >>"$RD/d04.md"
if [ "$_unver" -gt 0 ]; then
WARN=1
printf '**本輪有 %s 列查不到遠端版本。** 同一支腳本在擋人那條路徑查得到遠端版本,report 這條查不到,這是既有缺陷,不是這台機器的網路問題。\n\n' "$_unver" >>"$RD/d04.md"
add_pending 'version-guard.sh report 查不到遠端版本,版本落差本輪無證據' '版本落差與重啟閘門' '/jsc-cli:doctor'
fi
fi
else
_part=1
add_failed_source 'jsc-hooks/hooks/version-guard.sh'
printf '**版本比對失敗**:找不到 `version-guard.sh`。\n\n' >>"$RD/d04.md"
fi
# 重啟閘門。一行一支還沒重啟的 CLI;一行都沒有就是都放下了。
if [ -n "$_rg" ]; then
_rc=0
"$_rg" report >"$RD/d04.rg" 2>>"$RD/d04.err" </dev/null || _rc=$?
if [ "$_rc" -ne 0 ]; then
_part=1; D04_RC="$_rc"
add_failed_source "$_rg"
printf '**重啟閘門讀不到**:`restart-gate.sh report` 回 %s。\n' "$_rc" >>"$RD/d04.md"
else
{
printf '| CLI | 重啟閘門 | 升起時間 |\n'
printf '| --- | --- | --- |\n'
} >>"$RD/d04.md"
_up=0
while read -r _cli _rest; do
[ -n "$_cli" ] || continue
_up=$(( _up + 1 ))
_at=$(printf '%s' "$_rest" | sed -n 's/.*at=\([^ ]*\).*/\1/p')
printf '| %s | 已升起 | %s |\n' "$(cell "$_cli")" "$(cell "${_at:--}")" >>"$RD/d04.md"
add_pending "$_cli 的重啟閘門還升著,那一支要重新啟動" '版本落差與重啟閘門' '重開該支 CLI'
done <"$RD/d04.rg"
if [ "$_up" -eq 0 ]; then
printf '| (無) | 未升起 | - |\n' >>"$RD/d04.md"
else
WARN=1
fi
fi
else
_part=1
add_failed_source 'jsc-hooks/hooks/restart-gate.sh'
printf '**重啟閘門讀不到**:找不到 `restart-gate.sh`。\n' >>"$RD/d04.md"
fi
if [ "$_part" -eq 1 ]; then
D04_STATUS=fail
[ -n "$D04_NOTE" ] || D04_NOTE='兩份來源有一份讀不到'
else
D04_STATUS=ok
fi
return 0
}
# --- D-07 SDLC 階段鎖與工作包鎖 ---
d07() {
D07_STATUS=fail; D07_RC=0; D07_NOTE=''
_sess="$JSC_HOME/sessions"; _wp="$JSC_HOME/wp"
{
printf '### SDLC 階段鎖與工作包鎖現況\n\n'
printf '資料出自 `%s` 與 `%s`。只讀狀態,不做判定。\n\n' \
"\$JSC_HOME/sessions/{sid}.stage" "\$JSC_HOME/wp/*.pr"
} >"$RD/d07.md"
_bad=0
if [ -d "$_sess" ] && [ ! -r "$_sess" ]; then _bad=1; add_failed_source "$_sess"; fi
if [ -d "$_wp" ] && [ ! -r "$_wp" ]; then _bad=1; add_failed_source "$_wp"; fi
if [ "$_bad" -eq 1 ]; then
D07_NOTE='狀態目錄讀不到(權限)'
D07_RC=13
printf '**這一項失敗**:%s。這一輪沒有階段鎖與工作包鎖的證據,不是「沒有鎖」。\n' "$D07_NOTE" >>"$RD/d07.md"
add_pending 'SDLC 狀態目錄讀不到,權限要修' 'SDLC 階段鎖與工作包鎖現況' '/jsc-cli:setup'
return 0
fi
{
printf '| 工作階段 | 階段 | 必要標籤 | 上鎖時的模型 | 登記時間 |\n'
printf '| --- | --- | --- | --- | --- |\n'
} >>"$RD/d07.md"
_sn=0
for _f in "$_sess"/*.stage; do
[ -f "$_f" ] && [ -r "$_f" ] || continue
_sn=$(( _sn + 1 ))
_sid=$(basename "$_f" .stage)
_stage=$(cut -f1 "$_f" 2>/dev/null | head -n1)
_req=$(cut -f2 "$_f" 2>/dev/null | head -n1)
_mdl=$(cut -f3 "$_f" 2>/dev/null | head -n1)
printf '| %s | %s | %s | %s | %s |\n' \
"$(cell "$_sid")" "$(cell "${_stage:--}")" "$(cell "${_req:--}")" \
"$(cell "${_mdl:--}")" "$(mtime_of "$_f")" >>"$RD/d07.md"
done
[ "$_sn" -eq 0 ] && printf '| (無) | - | - | - | - |\n' >>"$RD/d07.md"
{
printf '\n| 存取庫 | 工作包 | PR | 上鎖時間 |\n'
printf '| --- | --- | --- | --- |\n'
} >>"$RD/d07.md"
_wn=0
for _f in "$_wp"/*.pr; do
[ -f "$_f" ] && [ -r "$_f" ] || continue
_wn=$(( _wn + 1 ))
_repo=$(sed -n 's/^repo=//p' "$_f" 2>/dev/null | head -n1)
_idx=$(sed -n 's/^index=//p' "$_f" 2>/dev/null | head -n1)
_wpn=$(sed -n 's/^wp=//p' "$_f" 2>/dev/null | head -n1)
_lk=$(sed -n 's/^locked=//p' "$_f" 2>/dev/null | head -n1)
printf '| %s | %s | %s | %s |\n' \
"$(cell "${_repo:--}")" "$(cell "${_wpn:--}")" \
"$(cell "${_repo:-?}#${_idx:-?}")" "$(cell "${_lk:--}")" >>"$RD/d07.md"
done
[ "$_wn" -eq 0 ] && printf '| (無) | - | - | - |\n' >>"$RD/d07.md"
printf '\n`.stage` 沒有存取庫欄位,`.pr` 沒有歸屬工作階段欄位,兩欄一律留「-」,不從分支名或目錄名回推——那兩個都會被改。登記時間取檔案的修改時間。\n' >>"$RD/d07.md"
if [ "$_sn" -eq 0 ] && [ "$_wn" -eq 0 ]; then D07_STATUS=empty; else D07_STATUS=ok; fi
return 0
}
# --- D-09 心跳與閘門狀態自述 ---
d09() {
D09_STATUS=fail; D09_RC=0; D09_NOTE=''
HEARTBEAT_STATE='讀不到'
{
printf '### 心跳與閘門狀態\n\n'
printf '資料出自 `heartbeat.sh report`。這一欄讀到的是**上一輪巡檢**寫的心跳:心跳改由巡檢寫,本輪那一次要等監控頁寫成之後才寫。\n\n'
} >"$RD/d09.md"
if ! _hb=$(find_tool hooks hooks/heartbeat.sh "${JSC_ASSIST_HEARTBEAT_SH:-}"); then
D09_NOTE='找不到 jsc-hooks 的 hooks/heartbeat.sh'
D09_RC=127
add_failed_source 'jsc-hooks/hooks/heartbeat.sh'
printf '**這一項失敗**:%s。心跳判不出來,不代表助理沒在跑,也不代表在跑。\n' "$D09_NOTE" >>"$RD/d09.md"
add_pending '心跳腳本找不到,jsc-hooks 沒裝或版本低於 0.3.7' '心跳與閘門狀態' '/jsc-cli:deploy'
return 0
fi
_rc=0
"$_hb" report >"$RD/d09.line" 2>"$RD/d09.err" </dev/null || _rc=$?
if [ "$_rc" -ne 0 ]; then
D09_RC="$_rc"
D09_NOTE="heartbeat.sh report 回 $_rc"
add_failed_source "$_hb"
printf '**這一項失敗**:%s。訊息:%s\n' "$D09_NOTE" "$(cell "$(cat "$RD/d09.err" 2>/dev/null)")" >>"$RD/d09.md"
add_pending '心跳腳本回非 0,心跳判不出來' '心跳與閘門狀態' '/jsc-assist:assistant status'
return 0
fi
_line=$(cat "$RD/d09.line" 2>/dev/null)
# 欄位以空白分隔,路徑擺最後。拆成一行一欄再取,不用貪婪比對——貪婪會抓到後面同名的鍵。
_get() { printf '%s' "$_line" | tr ' ' '\n' | sed -n "s/^$1=//p" | head -n1; }
_st=$(printf '%s' "$_line" | sed -n 's/^state=\([^ ]*\).*/\1/p')
_ts=$(_get ts); _age=$(_get age); _ttl=$(_get ttl)
_pid=$(_get pid); _cli=$(_get cli); _sid=$(_get session)
_file=$(printf '%s' "$_line" | sed -n 's/.*file=//p')
HEARTBEAT_TTL="$_ttl"
case "$_st" in
fresh) HEARTBEAT_STATE='新鮮' ;;
stale) HEARTBEAT_STATE='過期'; WARN=1 ;;
invalid) HEARTBEAT_STATE='心跳檔損壞'; WARN=1 ;;
absent) HEARTBEAT_STATE='不存在'; WARN=1 ;;
*) HEARTBEAT_STATE="判不出(state=${_st:-空值})"; WARN=1 ;;
esac
{
printf '| 項目 | 內容 |\n'
printf '| --- | --- |\n'
printf '| 心跳 | %s |\n' "$(cell "$HEARTBEAT_STATE")"
if [ -n "$_ts" ]; then
printf '| 上次心跳 | %s,距這次巡檢 %s 秒 |\n' "$(fmt_ts "$_ts" | tr -d '\n')" "$(cell "${_age:--}")"
else
printf '| 上次心跳 | - |\n'
fi
printf '| 過期門檻 | %s 秒 |\n' "$(cell "${_ttl:--}")"
printf '| cli | %s |\n' "$(cell "${_cli:--}")"
printf '| session | %s |\n' "$(cell "${_sid:--}")"
printf '| pid | %s。只給要找行程的人參考,不參與判定 |\n' "$(cell "${_pid:--}")"
printf '| 心跳檔 | `%s` |\n' "$(cell "${_file:--}")"
printf '\n心跳的判準只看 `ts` 距現在有沒有超過門檻,不看 pid 存活:五支 CLI 與容器裡的行程互相看不到彼此的 pid。閘門的判定留在 hook,助理只維持心跳,不參與判定。\n'
printf '\n心跳新鮮代表上一輪巡檢跑完了,而且結果記上監控頁了。它不代表那一輪四項都成功——四項的成敗看本節上面的「本輪判定」。\n'
} >>"$RD/d09.md"
case "$_st" in
fresh|stale|invalid|absent) D09_STATUS=ok ;;
*) D09_STATUS=fail; D09_NOTE='report 印不出認得的 state' ;;
esac
[ "$_st" = absent ] && D09_STATUS=empty
return 0
}
# --- 待辦簿筆數(只供目錄頁那一列用)---
count_tasks() {
TASKS_TOTAL=0; TASKS_FAILING=0
_d="$STATE_DIR/tasks"
[ -d "$_d" ] && [ -r "$_d" ] || return 0
for _f in "$_d"/*; do
[ -f "$_f" ] || continue
TASKS_TOTAL=$(( TASKS_TOTAL + 1 ))
_fc=$(sed -n 's/^fail_count=//p' "$_f" 2>/dev/null | head -n1)
case "$_fc" in ''|*[!0-9]*) _fc=0 ;; esac
[ "$_fc" -gt 0 ] && TASKS_FAILING=$(( TASKS_FAILING + 1 ))
done
return 0
}
# --- 組出監控頁那一節 ---
tally() { # $1=項目狀態
case "$1" in
fail) FAIL_COUNT=$(( FAIL_COUNT + 1 )) ;;
*) OK_COUNT=$(( OK_COUNT + 1 )) ;;
esac
}
compose() {
_v='正常'
[ "$WARN" -eq 1 ] && _v='警示'
[ "$FAIL_COUNT" -gt 0 ] && _v='警示'
[ "$OK_COUNT" -eq 0 ] && _v='異常'
VERDICT="$_v"
{
printf '## 巡檢 %s\n\n' "$AT"
printf '| 項目 | 內容 |\n'
printf '| --- | --- |\n'
printf '| 巡檢時間 | %s |\n' "$AT"
printf '| 觸發方式 | %s |\n' "$TRIGGER"
printf '| 本輪判定 | %s |\n' "$VERDICT"
printf '| 本輪項目 | 四項:D-01 使用統計、D-04 版本與重啟閘門、D-07 階段鎖與工作包鎖、D-09 心跳自述。成功 %s 項、失敗 %s 項 |\n' "$OK_COUNT" "$FAIL_COUNT"
printf '| 讀不到的來源 | %s |\n' "$(cell "${FAILED_SOURCES:-無}")"
if [ "$LOCK_BROKEN" -eq 1 ]; then
printf '| 鎖 | 上一輪的鎖逾時,本輪搶回來了。上一輪沒跑完,那一輪不會寫心跳 |\n'
fi
printf '\n'
cat "$RD/d09.md"; printf '\n'
cat "$RD/d01.md"; printf '\n'
printf '### hook 執行期錯誤\n\n'
printf '**這一輪不做這一項。** D-02 hook 錯誤巡檢還沒實作,這一節沒有資料不代表沒有 hook 錯誤。要現在查就跑 `/jsc-hooks:hooks-install` 的錯誤掃描,或直接跑 `jsc-hooks` 的 `tools/scan-hook-errors.sh`。\n\n'
cat "$RD/d04.md"; printf '\n'
cat "$RD/d07.md"; printf '\n'
printf '### 待辦簿到期與逾期\n\n'
printf '**這一輪只數筆數,不逐筆判到期。** 本輪待辦簿共 %s 筆,其中 %s 筆 `fail_count` 大於 0。逐筆的到期與逾期判定還沒實作,要看逐筆內容就跑 `/jsc-assist:assistant status`。\n\n' \
"$TASKS_TOTAL" "$TASKS_FAILING"
printf '### 待人處理\n\n'
printf '助理只提醒,不代為執行。這一節列的是本輪要人接手的項目。\n\n'
printf '| 項目 | 來源子節 | 建議入口 |\n'
printf '| --- | --- | --- |\n'
if [ -s "$RD/pend.md" ]; then cat "$RD/pend.md"; else printf '| (無) | - | - |\n'; fi
} >"$RD/section.md"
# 監控頁不存在時要建的整頁內容。基本資料建頁時寫一次,之後不再更動。
{
printf '# 助理巡檢 — %s/%s\n\n' "$HOST" "$USER_NAME"
printf '> 由 `jsc-assist` 維護。這是監控頁 `%s`。\n' "$PAGE"
printf '> 這頁是這台機器的巡檢軌跡:一次巡檢附加一節,節標題帶時間戳,舊的節一個字都不動。\n'
printf '> 附加是刻意的。助理的寫入是背景行為,覆寫錯了沒人在現場,軌跡被抹掉也看不出斷在哪一輪。\n'
printf '> 目錄頁 `MONITOR_CONTENTS` 只更新自己那一列,寫入語意與這頁不同,不要混用。\n\n'
printf '```mermaid\nflowchart LR\n'
printf ' A[巡檢一輪] --> B[收攏四項結果]\n'
printf ' B --> C[附加一節,節標題帶時間戳]\n'
printf ' C --> D[既有的節原樣保留]\n'
printf ' D --> E[回頭更新 MONITOR_CONTENTS 自己那一列]\n'
printf ' E --> F[最後才寫心跳]\n'
printf '```\n\n'
printf '## 本頁基本資料\n\n'
printf '建頁時寫一次,之後不再更動。\n\n'
printf '| 項目 | 內容 |\n'
printf '| --- | --- |\n'
printf '| 主機 | %s |\n' "$(cell "$HOST")"
printf '| 帳號 | %s |\n' "$(cell "$USER_NAME")"
printf '| 雜湊來源 | `%s/%s` |\n' "$(cell "$HOST")" "$(cell "$USER_NAME")"
printf '| 狀態檔根目錄 | `$JSC_HOME/assistant/`(`$JSC_HOME` 未設定就退回 `~/.jsc`) |\n\n'
cat "$RD/section.md"
} >"$RD/newpage.md"
{
printf 'page=%s\n' "$PAGE"
printf 'host=%s\n' "$HOST"
printf 'user=%s\n' "$USER_NAME"
printf 'heartbeat=%s\n' "$HEARTBEAT_STATE"
printf 'last_patrol=%s\n' "$AT"
printf 'tasks_total=%s\n' "$TASKS_TOTAL"
printf 'tasks_failing=%s\n' "$TASKS_FAILING"
printf 'row=| [[%s]] | %s | %s | %s | %s | %s | %s |\n' \
"$PAGE" "$HOST" "$USER_NAME" "$HEARTBEAT_STATE" "$AT" "$TASKS_TOTAL" "$TASKS_FAILING"
} >"$RD/contents.tsv"
return 0
}
# --- 參數解析 ---
CMD="${1:-}"
[ -n "$CMD" ] || usage
shift
case "$CMD" in collect|finish|abort) ;; *) usage ;; esac
OUT=''
while [ "$#" -gt 0 ]; do
case "$1" in
--out) [ "$#" -ge 2 ] || usage; OUT="$2"; shift 2 ;;
--trigger) [ "$#" -ge 2 ] || usage; TRIGGER="$2"; shift 2 ;;
--round) [ "$#" -ge 2 ] || usage; ROUND="$2"; shift 2 ;;
--dry-run) DRYRUN=1; shift ;;
*) usage ;;
esac
done
[ -n "$OUT" ] && RD="$OUT"
case "$CMD" in
collect)
[ -n "$TRIGGER" ] || { if [ "${JSC_CLI:-}" = cron ]; then TRIGGER='排程'; else TRIGGER='手動'; fi; }
case "$TRIGGER" in 排程|事件|手動) ;; *) usage ;; esac
ROUND="$(date +%s)-$$"
HOST=$(hostname 2>/dev/null || uname -n 2>/dev/null || printf 'unknown')
USER_NAME="${USER:-$(id -un 2>/dev/null || printf 'unknown')}"
AT=$(date '+%Y-%m-%d %H:%M')
# 先讀心跳。門檻要先拿到,鎖的逾時才有依據。
mkdir -p "$RD" 2>/dev/null || die 5 "建不出巡檢暫存目錄 $RD。"
: >"$RD/pend.md"
HEARTBEAT_TTL=''
d09
lock_acquire
d01
d04
d07
count_tasks
tally "$D01_STATUS"; tally "$D04_STATUS"; tally "$D07_STATUS"; tally "$D09_STATUS"
HASH=''
if _hi=$(find_tool gitea tools/hash-id "${JSC_ASSIST_HASH_ID:-}"); then
HASH=$("$_hi" "$HOST/$USER_NAME" 2>/dev/null) || HASH=''
fi
if [ -n "$HASH" ]; then PAGE="MONITOR_$HASH"; else PAGE=''; fi
compose
printf 'round=%s\n' "$ROUND"
printf 'lock=acquired\n'
printf 'lock_broken=%s\n' "$LOCK_BROKEN"
printf 'hash=%s\n' "$HASH"
printf 'page=%s\n' "$PAGE"
printf 'host=%s\n' "$HOST"
printf 'user=%s\n' "$USER_NAME"
printf 'at=%s\n' "$AT"
printf 'item=D-01 status=%s rc=%s note=%s\n' "$D01_STATUS" "$D01_RC" "$D01_NOTE"
printf 'item=D-04 status=%s rc=%s note=%s\n' "$D04_STATUS" "$D04_RC" "$D04_NOTE"
printf 'item=D-07 status=%s rc=%s note=%s\n' "$D07_STATUS" "$D07_RC" "$D07_NOTE"
printf 'item=D-09 status=%s rc=%s note=%s\n' "$D09_STATUS" "$D09_RC" "$D09_NOTE"
printf 'verdict=%s\n' "$VERDICT"
printf 'failed_sources=%s\n' "${FAILED_SOURCES:-無}"
printf 'tasks_total=%s\n' "$TASKS_TOTAL"
printf 'tasks_failing=%s\n' "$TASKS_FAILING"
printf 'section_file=%s\n' "$RD/section.md"
printf 'newpage_file=%s\n' "$RD/newpage.md"
printf 'contents_file=%s\n' "$RD/contents.tsv"
[ "$OK_COUNT" -eq 0 ] && exit 3
[ "$FAIL_COUNT" -gt 0 ] && exit 1
exit 0 ;;
finish)
[ -n "$ROUND" ] || usage
[ -d "$LOCK" ] || die 4 "本輪的鎖已經不在($LOCK),不寫心跳。鎖多半是逾時被下一輪搶走了。"
_h=$(lock_round)
[ "$_h" = "$ROUND" ] \
|| die 4 "鎖不在本輪手上(鎖的輪次是 ${_h:-空值},本輪是 $ROUND),不寫心跳。上一輪跑太久被搶走了,這一輪的結果不算數。"
_hb=$(find_tool hooks hooks/heartbeat.sh "${JSC_ASSIST_HEARTBEAT_SH:-}") \
|| die 2 '找不到 jsc-hooks 的 hooks/heartbeat.sh,心跳沒有東西可寫。請先安裝 jsc-hooks 0.3.7 以上。'
if [ "$DRYRUN" -eq 1 ]; then
printf 'dryrun=1 heartbeat_cmd=%s write snapshot=%s -> %s lock=%s\n' \
"$_hb" "$RD/usage-next.tsv" "$PREV_SNAP" "$LOCK"
exit 0
fi
_rc=0
"$_hb" write >/dev/null 2>"$RD/finish.err" </dev/null || _rc=$?
[ "$_rc" -eq 0 ] \
|| die 5 "heartbeat.sh write 回 $_rc,心跳沒寫成:$(tr '\n' ' ' <"$RD/finish.err" 2>/dev/null)"
printf 'heartbeat=written\n'
# 快照要等這一輪真的收口才換上。半途失敗就換掉的話,下一輪的「本輪次數」會少算。
if [ -f "$RD/usage-next.tsv" ]; then
cp "$RD/usage-next.tsv" "$PREV_SNAP" 2>/dev/null || die 5 "用量快照換不上:$PREV_SNAP。"
printf 'snapshot=promoted\n'
else
printf 'snapshot=skipped\n'
fi
lock_release
exit 0 ;;
abort)
[ -n "$ROUND" ] || usage
printf 'heartbeat=not-written\n'
lock_release
exit 0 ;;
esac
+96 -20
View File
@@ -2,23 +2,43 @@
# schedule.sh — 助理系統排程的安裝、移除與查現況(供 jsc-assist:assistant 呼叫)。
#
# 用法:
# schedule.sh install [heartbeat|patrol|all] [--dry-run] [--cli {代號}] [--patrol-cmd {指令}]
# schedule.sh install [patrol|all] [--dry-run] [--cli {代號}] [--patrol-cmd {指令}] [--period {分鐘}]
# schedule.sh remove [heartbeat|patrol|all] [--dry-run]
# schedule.sh status [--dry-run]
#
# 工作代號省略時一律是 heartbeat。心跳每 60 秒寫一次,巡檢每 15 分鐘跑一輪。
# 巡檢那一筆要自己指名(`patrol` 或 `all`)才會裝——巡檢本體還沒實作,裝了每一輪都會失敗。
# 工作代號省略時一律是 patrol。這一版只裝巡檢那一筆,心跳由巡檢跑完那一輪自己寫。
#
# 結束碼:
# 0 成功:install 條目寫進去也回讀得到、排程服務在跑;remove 移除完成,或本來就沒裝;
# status 印完現況(有沒有裝都算成功,看 installed 欄)
# 1 install 寫進去了,但排程服務沒在跑——條目不會被執行。WSL 預設不啟動 cron,這一碼
# 多半就是它。呼叫端一律照實講「排程裝了但不會執行」,不可以宣稱會定時執行
# 2 找不到 jsc-hooks 的 hooks/heartbeat.sh。心跳沒有東西可跑,整支停下
# 2 找不到 jsc-hooks 的 hooks/heartbeat.sh。心跳門檻查不到,週期算不出來,整支停下
# 3 這台機器沒有可用的排程機制:認不得作業系統,或 crontab 與 schtasks 都找不到
# 4 排程操作失敗:讀不到現有排程(且失敗原因不是「沒有排程」)、寫入或刪除回非 0
# 5 回讀驗證失敗:寫入回 0 但條目不在,或移除回 0 但條目還在,又或其他人的條目數量對不上
# 6 用法錯誤:不認得的子命令、不認得的工作代號、缺參數,或判不出要用哪一支 CLI 跑巡檢
# 6 用法錯誤:不認得的子命令、不認得的工作代號、缺參數、判不出要用哪一支 CLI 跑巡檢,
# 或 --period 給的週期塞不進心跳的過期門檻
#
# --- 排程只叫巡檢,心跳由巡檢寫 ---
#
# 舊版排程每分鐘直接呼叫 heartbeat.sh write。那樣心跳新鮮只證明 cron 活著:巡檢整個壞掉、
# 一輪都沒跑成,心跳照樣新鮮,靠心跳判定的閘門照樣放行,沒有任何訊號。
# 現在排程只叫巡檢,巡檢把結果寫上監控頁之後才寫那一次心跳。心跳新鮮於是等於「上一輪巡檢
# 真的做完了,而且結果記下來了」。
# 所以 `install heartbeat` 直接回 6,不給裝:裝了就是有第二個寫心跳的人,心跳的意思立刻回到
# 舊版。舊機器上留著的那一筆 heartbeat 條目,`install patrol` 會順手清掉,並在輸出印
# legacy_removed=1——不清的話它每分鐘照樣寫,這次改動等於白做。
# `remove` 與 `status` 仍然認得 heartbeat 這個代號,就是為了清掉與看得到那一筆舊條目。
#
# --- 巡檢週期怎麼定 ---
#
# 心跳的更新頻率現在等於巡檢週期,所以週期一定要塞得進心跳的過期門檻,不然心跳永遠是過期。
# 門檻不寫死,改讀 `heartbeat.sh report` 的 `ttl` 欄——那是這台機器實際生效的值。
# 週期取「漏掉一輪還算新鮮、漏掉兩輪就過期」:2 × 週期 × 60 < 門檻,再往下取一個能整除一小時
# 的分鐘數,排程間隔才規律。門檻預設 300 秒時算出來是每 2 分鐘一輪。
# 要拉長巡檢週期就先把門檻調大(`JSC_ASSISTANT_HEARTBEAT_TTL`,單位秒),週期會跟著變長:
# 門檻 1800 秒算出每 12 分鐘一輪。週期與門檻兩個數字綁在一起算,不會再各走各的。
#
# --- 只動自己那一筆 ---
#
@@ -51,6 +71,7 @@
# 假的 crontab 驗濾除邏輯時才設
# JSC_ASSIST_PATROL_CMD 巡檢要跑的指令,優先於 --patrol-cmd 以外的所有推斷
# JSC_CLI 目前是哪一支 CLI,決定巡檢預設指令
# JSC_ASSISTANT_HEARTBEAT_TTL 心跳過期門檻,單位秒。巡檢週期由它算出來
set -u
MARK_PREFIX='# jsc-assist:assistant'
@@ -63,13 +84,14 @@ TASK_PREFIX='jsc-assist-assistant'
DRYRUN=0
CLI=''
PATROL_CMD="${JSC_ASSIST_PATROL_CMD:-}"
PERIOD=''
SCRIPT_DIR=$(CDPATH= cd -- "$(dirname -- "$0")" 2>/dev/null && pwd)
SCRIPT_DIR="${SCRIPT_DIR:-.}"
usage() {
cat >&2 <<'EOF'
usage: schedule.sh install [heartbeat|patrol|all] [--dry-run] [--cli 代號] [--patrol-cmd 指令]
usage: schedule.sh install [patrol|all] [--dry-run] [--cli 代號] [--patrol-cmd 指令] [--period 分鐘]
schedule.sh remove [heartbeat|patrol|all] [--dry-run]
schedule.sh status [--dry-run]
EOF
@@ -155,10 +177,31 @@ count_lines() { _n=$(grep -c '' "$1" 2>/dev/null); [ -n "$_n" ] || _n=0; printf
# 數管線進來的行數,語意同 count_lines。
count_stdin() { _n=$(grep -c '' 2>/dev/null); [ -n "$_n" ] || _n=0; printf '%s' "$_n"; }
# 這台機器實際生效的心跳過期門檻。取自 heartbeat.sh report 的 ttl 欄,不自己重算:
# 門檻的唯一來源是那支腳本,兩邊各算一次就會漂移。讀不到就退回 300。
heartbeat_ttl() {
_t=$("$HEARTBEAT" report </dev/null 2>/dev/null | tr ' ' '\n' | sed -n 's/^ttl=//p' | head -n1)
case "$_t" in ''|*[!0-9]*) _t=300 ;; esac
[ "$_t" -gt 0 ] || _t=300
printf '%s' "$_t"
}
# 由門檻算出巡檢週期,單位分鐘。條件是 2 × 週期 × 60 < 門檻:漏掉一輪還算新鮮,漏掉兩輪
# 才過期。再往下取一個能整除一小時的分鐘數,`*/N` 的間隔才規律。
period_for_ttl() { # $1=門檻秒數
_max=$(( ($1 - 1) / 120 ))
[ "$_max" -lt 1 ] && _max=1
_p=1
for _d in 1 2 3 4 5 6 10 12 15 20 30 60; do
[ "$_d" -le "$_max" ] && _p="$_d"
done
printf '%s' "$_p"
}
spec_of() {
case "$1" in
heartbeat) printf '* * * * *' ;;
patrol) printf '*/15 * * * *' ;;
patrol) printf '*/%s * * * *' "$PERIOD" ;;
esac
}
@@ -170,10 +213,10 @@ patrol_command() {
[ -n "$_cli" ] || _cli="${JSC_CLI:-}"
[ -n "$_cli" ] || { [ -n "${CLAUDE_PLUGIN_ROOT:-}" ] && _cli=claude; }
case "$_cli" in
claude) printf 'claude -p "/jsc-assist:patrol"' ;;
codex) printf "codex exec '\$patrol'" ;;
claude) printf 'claude -p "/jsc-assist:assistant 跑一輪巡檢"' ;;
codex) printf "codex exec '\$assistant 跑一輪巡檢'" ;;
copilot) printf 'copilot -p "跑一輪助理巡檢"' ;;
antigravity) printf 'agy -p "/jsc-assist:patrol"' ;;
antigravity) printf 'agy -p "/jsc-assist:assistant 跑一輪巡檢"' ;;
kiro) printf 'kiro-cli -p "跑一輪助理巡檢"' ;;
*) return 1 ;;
esac
@@ -214,7 +257,7 @@ CMD="${1:-}"
shift
case "$CMD" in install|remove|status) ;; *) usage ;; esac
JOBS='heartbeat'
JOBS='patrol'
if [ "$#" -gt 0 ]; then
case "$1" in
heartbeat) JOBS='heartbeat'; shift ;;
@@ -228,15 +271,36 @@ while [ "$#" -gt 0 ]; do
--dry-run) DRYRUN=1; shift ;;
--cli) [ "$#" -ge 2 ] || usage; CLI="$2"; shift 2 ;;
--patrol-cmd) [ "$#" -ge 2 ] || usage; PATROL_CMD="$2"; shift 2 ;;
--period) [ "$#" -ge 2 ] || usage; PERIOD="$2"; shift 2 ;;
*) usage ;;
esac
done
# 心跳那一筆不給裝。理由見檔頭「排程只叫巡檢,心跳由巡檢寫」:多一個寫心跳的人,心跳的
# 意思立刻退回舊版。install all 只裝巡檢那一筆。
case "$CMD:$JOBS" in
install:heartbeat)
die 6 '心跳那一筆不裝了。心跳改由巡檢跑完那一輪自己寫,排程只叫巡檢:請跑 `schedule.sh install patrol`。舊機器上留著的 heartbeat 條目,install patrol 會順手清掉。' ;;
install:*heartbeat*) JOBS='patrol' ;;
esac
HEARTBEAT=$(heartbeat_sh) \
|| die 2 '找不到 jsc-hooks 的 hooks/heartbeat.sh,心跳沒有東西可跑。請先安裝 jsc-hooks 0.3.7 以上。'
|| die 2 '找不到 jsc-hooks 的 hooks/heartbeat.sh,心跳的過期門檻查不到,巡檢週期算不出來。請先安裝 jsc-hooks 0.3.7 以上。'
MECH=$(mechanism) \
|| die 3 "這台機器沒有可用的排程機制(作業系統:$(os_kind),找不到 $CRONTAB_CMD 或 schtasks)。"
# 巡檢週期與心跳門檻綁在一起算。--period 給的值一樣要通過同一條式子,否則裝出來的排程
# 會讓心跳永遠過期,而且沒有人看得出來是週期設錯。
TTL=$(heartbeat_ttl)
if [ -n "$PERIOD" ]; then
case "$PERIOD" in ''|*[!0-9]*) usage ;; esac
[ "$PERIOD" -ge 1 ] || usage
[ $(( 2 * PERIOD * 60 )) -lt "$TTL" ] \
|| die 6 "巡檢週期 $PERIOD 分鐘塞不進心跳的過期門檻 $TTL 秒(要 2 × 週期 × 60 < 門檻)。心跳現在由巡檢寫,週期比門檻長的話心跳永遠是過期。請改短週期,或先把 JSC_ASSISTANT_HEARTBEAT_TTL 調大。"
else
PERIOD=$(period_for_ttl "$TTL")
fi
# 巡檢指令在這裡就解出來。放進 cron_entry 再解的話,那支是在命令替換的子行程裡跑,
# 判不出 CLI 時 die 只結束子行程,主流程會帶著空指令繼續往下裝。
PATROL_RESOLVED=''
@@ -253,11 +317,16 @@ trap 'rm -rf "$TMPD"' EXIT
schtasks_install() {
_rc=0
# 舊版的 heartbeat 任務一併刪掉,理由同 crontab 那一邊。刪不掉不算失敗,往下照裝。
_legacy=0
if schtasks /Query /TN "$(task_name heartbeat)" >/dev/null 2>&1; then
schtasks /Delete /TN "$(task_name heartbeat)" /F >/dev/null 2>&1 && _legacy=1
fi
for _job in $JOBS; do
_tn=$(task_name "$_job")
case "$_job" in
heartbeat) _mo=1; _run="sh \"$HEARTBEAT\" write" ;;
patrol) _mo=15; _run=$(patrol_command) || die 6 '判不出要用哪一支 CLI 跑巡檢,請帶 --cli 或 --patrol-cmd。' ;;
patrol) _mo="$PERIOD"; _run="$PATROL_RESOLVED" ;;
*) continue ;;
esac
_tr="cmd /c $_run <NUL >> \"$LOG\" 2>&1"
if [ "$DRYRUN" -eq 1 ]; then
@@ -268,8 +337,10 @@ schtasks_install() {
schtasks /Create /TN "$_tn" /SC MINUTE /MO "$_mo" /F /TR "$_tr" >/dev/null 2>&1 \
|| die 4 "schtasks 建立 $_tn 失敗。"
schtasks /Query /TN "$_tn" >/dev/null 2>&1 || die 5 "schtasks 建立 $_tn 回 0,卻查不到這個任務。"
printf 'installed=%s task=%s\n' "$_job" "$_tn"
_rc=0
done
printf 'ttl=%s period=%s legacy_removed=%s log=%s\n' "$TTL" "$PERIOD" "$_legacy" "$LOG"
return "$_rc"
}
@@ -291,7 +362,8 @@ schtasks_remove() {
}
schtasks_status() {
printf 'mechanism=schtasks service=%s log=%s\n' "$(service_state)" "$LOG"
printf 'mechanism=schtasks service=%s ttl=%s period=%s log=%s\n' \
"$(service_state)" "$TTL" "$PERIOD" "$LOG"
for _job in heartbeat patrol; do
_tn=$(task_name "$_job")
if schtasks /Query /TN "$_tn" >/dev/null 2>&1; then
@@ -311,7 +383,9 @@ crontab_install() {
cp "$_cur" "$_new"
# 先濾掉自己這幾個工作的舊條目,再追加新的。重跑不會疊成兩筆,別人的條目原樣留著。
for _job in $JOBS; do
# 舊版的 heartbeat 條目一併清掉:留著它每分鐘照樣寫心跳,心跳就退回「只證明 cron 活著」。
_legacy=$(cron_lines_for "$_new" heartbeat | count_stdin)
for _job in $JOBS heartbeat; do
grep -vF "$(marker_of "$_job")" "$_new" >"$TMPD/f" 2>/dev/null || true
mv "$TMPD/f" "$_new"
done
@@ -325,8 +399,8 @@ crontab_install() {
for _job in $JOBS; do
printf 'dryrun=crontab job=%s entry=%s\n' "$_job" "$(cron_entry "$_job")"
done
printf 'dryrun=crontab action=write others_kept=%s total_lines=%s\n' \
"$_others" "$(count_lines "$_new")"
printf 'dryrun=crontab action=write ttl=%s period=%s legacy_removed=%s others_kept=%s total_lines=%s\n' \
"$TTL" "$PERIOD" "$_legacy" "$_others" "$(count_lines "$_new")"
printf -- '--- 寫回後的 crontab ---\n'
cat "$_new"
return 0
@@ -350,7 +424,8 @@ crontab_install() {
for _job in $JOBS; do
printf 'installed=%s entry=%s\n' "$_job" "$(cron_lines_for "$_chk" "$_job")"
done
printf 'others_kept=%s log=%s\n' "$_kept" "$LOG"
printf 'ttl=%s period=%s legacy_removed=%s others_kept=%s log=%s\n' \
"$TTL" "$PERIOD" "$_legacy" "$_kept" "$LOG"
return 0
}
@@ -402,7 +477,8 @@ crontab_remove() {
crontab_status() {
_cur="$TMPD/cur"; cron_read "$_cur"
printf 'mechanism=crontab service=%s log=%s\n' "$(service_state)" "$LOG"
printf 'mechanism=crontab service=%s ttl=%s period=%s log=%s\n' \
"$(service_state)" "$TTL" "$PERIOD" "$LOG"
for _job in heartbeat patrol; do
_line=$(cron_lines_for "$_cur" "$_job")
if [ -n "$_line" ]; then