Files
assist/skills/assistant/SKILL.md
T
jiantw83 2b7176a644 feat(patrol): 一輪巡檢落地,心跳改由巡檢跑完才寫
What:
- 新增 tools/patrol.sh,三個子命令:collect 收集、finish 收尾、abort 中止。這一輪的巡檢項目是四項純讀取:使用統計、版本落差與重啟閘門、階段鎖與工作包鎖、心跳自述。
- 心跳的寫入從排程移到巡檢收尾。排程只呼叫巡檢,不再直接寫心跳。
- 排程週期由心跳門檻推導,不再寫死。install heartbeat 這個工作代號改為拒絕。

Why:
- 排程直接寫心跳的話,心跳新鮮只證明排程活著。巡檢整個壞掉、每輪都失敗,心跳照樣新鮮,閘門照樣放行,而且沒有任何錯誤訊息——這是無聲失效,是最難發現的一種。
- 改成巡檢寫,心跳新鮮才等於上一輪真的跑完了。閘門判的才是工作訊號,不是行程存活訊號。
- 門檻是讀取端的設定,心跳檔裡不存它。所以週期與門檻各寫死一個數字一定會撞:門檻五分鐘、巡檢十五分鐘,心跳永遠是過期的。

How:
- 週期取「滿足漏掉一輪還算新鮮、漏掉兩輪才過期」的最大值,並且要能整除一小時。要拉長巡檢週期就調大門檻,週期自動跟著長,兩個數字不會各走各的。
- 心跳只看「這一輪有沒有把結果記下來」,不看四項的成敗。四項有失敗但監控頁寫成了就寫心跳,頁上判定標警示;頁寫不成就中止,一定不寫,讓心跳自己過期。頁每輪都寫失敗卻照樣寫心跳,等於把這次要修掉的缺陷原封不動搬過去。
- 整輪拿一把目錄鎖,搶不到就讓開,不排隊也不並行。並行兩輪會在同一頁附兩節,還會互相蓋掉用量快照。鎖逾時可被下一輪搶走,被搶走的那一輪收尾時對不上就不寫心跳——它沒跑到底,不該蓋章。
- 任一項失敗不影響其餘項目。失敗要在監控頁上看得出來是失敗,不是沒資料;來源是空的則明寫「那是零次,不是故障」。
- 版本盤點照抄腳本原字。抄到查詢失敗就寫查詢失敗,不改寫成相符、也不自己補查遠端版本——查不到是沒有證據,不是版本沒問題。

Who:
助理落地的第四塊。排程那一輪留下的問題在這裡解掉了。
2026-09-01 15:06:10 +08:00

36 KiB
Raw Blame History

name, description
name description
assistant Start, inspect, patrol, or stop the background assistant, with jsc-hooks/hooks/heartbeat.sh owning the single freshness verdict, tools/schedule.sh owning the system scheduler, and tools/patrol.sh owning one patrol round. The heartbeat is written by a completed patrol round and by nothing else, so the schedule carries the patrol entry only and its period is derived from the heartbeat TTL; start runs one round and then installs that entry, status turns heartbeat.sh report, schedule.sh status and the task book into one read-only table, stop removes the entry first and then clears the heartbeat. One round reads four independent sources - skill and chain usage, version gaps and the restart gate, SDLC stage and work-package locks, and the heartbeat's own report - and appends the result to wiki MONITOR_{HASH} through jsc-gitea:wiki before tools/patrol.sh finish writes the heartbeat. A round that cannot record its result writes no heartbeat, and a round that starts while the previous one still holds the lock stands down. Use when someone starts, patrols or stops the assistant, or asks whether it is running and what is queued; not for environment health checks (jsc-cli:doctor), not for skill usage counts (jsc-log:stats).

assistant — start, status, patrol, stop

The background assistant runs where nobody is watching it. Its heartbeat is the only evidence that it is alive, so this skill is the single entry point for the four operations that touch that evidence: patrol writes it, status reads it, stop clears it, and start bootstraps the whole loop.

jsc-hooks/hooks/heartbeat.sh owns every heartbeat operation, including the freshness verdict. Never read, parse, write or delete $JSC_HOME/assistant/heartbeat directly — one verdict, one source.

tools/schedule.sh owns every system-scheduler operation: installing an entry, removing it, and reading which entries exist. Never call crontab or schtasks from this skill, and never edit a crontab by hand.

tools/patrol.sh owns one patrol round: taking the round lock, reading the four sources, composing the monitor-page section, and — after that section is on the page — writing the heartbeat. Never re-read a source this skill already handed to that script, and never compose the section by hand; the script prints the file paths.

All three flows have fixed inputs and outputs, so all three live in scripts. The task book is the only thing this skill reads for itself, and that is one directory listing.

Pick the operation

Run exactly one operation per invocation. Take it from the request: starting, launching or waking the assistant is start; asking whether it runs, what it is doing, or what is queued is status; running one round, patrolling, or a scheduled wake-up is patrol; stopping, halting or shutting it down is stop. When the request names none of the four, or names more than one, ask through the jsc-ask:ask decision tree with those four as the options, each stating its effect — start runs one round and installs the scheduled entry that keeps running rounds, status changes nothing, patrol runs one round and writes one heartbeat, stop removes that entry and deletes the heartbeat. The one exception: a patrol invocation never asks anything at all (see 界線 1 below). Never guess, and never run a second operation the caller did not ask for. Completion condition: exactly one of start, status, patrol, stop is chosen and named in the report.

Data sources

Path Read by Format
$JSC_HOME/assistant/heartbeat heartbeat.sh only, never this skill key=value lines: ts, pid, cli, session
$JSC_HOME/assistant/schedule.log nobody here — the scheduled entry appends to it free text; point the operator at it when a scheduled round misbehaves
$JSC_HOME/assistant/tasks/{id} this skill, read-only key=value lines, one task per file: id, kind (check / todo), title, action, trigger, recur, repo, due, state (pending / done / paused), last_run, next_run, fail_count, origin (user / assistant)
$JSC_HOME/assistant/patrol.lock/ patrol.sh only the round lock, a directory. info holds round, pid, started
$JSC_HOME/assistant/patrol/ patrol.sh only one round's scratch files, including section.md, newpage.md and contents.tsv
$JSC_HOME/assistant/usage-prev.tsv patrol.sh only last recorded round's cumulative usage counts, so the next round can print a real per-round delta

$JSC_HOME defaults to ~/.jsc. heartbeat.sh report prints the resolved heartbeat path in its file= field, so take the assistant directory from there rather than rebuilding it.

The verdict is time-based only. A heartbeat counts as fresh when the file exists and its ts is less than the TTL behind now (300 seconds by default, JSC_ASSISTANT_HEARTBEAT_TTL overrides it). pid liveness is never tested: five CLIs and container processes cannot see each other's pids, so a live-looking pid proves nothing and a missing one proves nothing either. Report pid as a hint for whoever has to find a blocking process, and give it no weight in the verdict.

heartbeat.sh exit codes

Every call in every operation below is judged by this table. Report the code you got, then take the row's action — never retry a code silently, and never downgrade a failure into a success.

Code Meaning What to do
0 write wrote the heartbeat, clear finished and the file is gone, report printed its line, check says fresh Carry on with the operation's next step. For report, the state still has to be read out of the printed state= field
1 check: the heartbeat exists but is at or past the TTL — the last patrol round finished more than one TTL ago Report 助理未運行, name the age in seconds, and say the assistant has to be started again. report returns this state as state=stale with exit 0
2 The script did not run at all — it failed to load its lib.sh Report that the heartbeat state is unknown, name the script path and the code, and stop the operation. Never claim the assistant is running, and never claim it is stopped
3 check: no heartbeat file — no patrol round has ever finished Report 助理未運行 and say to run start. report returns this state as state=absent with exit 0. In stop this state cannot appear, because clear treats a missing file as success
4 check: the heartbeat exists but its ts is missing, empty or not a number — the file is damaged, the assistant is not merely stopped Treat it as not fresh; falling back to fresh is forbidden. Report the file as damaged, say the state cannot be read from it, and tell the operator to run stop and then start to rebuild it. report returns this state as state=invalid with exit 0
5 Filesystem failure — write could not write the file, or clear could not delete it and the file is still there Serious. Report it loudly with the stderr text and the path, and follow the operation's own step for this code. Never report the operation as done
6 Usage error — an unknown subcommand, or none at all This is a defect in the call, not a state of the assistant. Report the exact command line that was run, correct it to one of write, check, report, clear, and run it once more. Report a second exit 6 as a defect in this skill and stop

The scheduler

Nothing in a background assistant runs on its own. The system scheduler is what makes it periodic, and tools/schedule.sh is the only thing here that touches it. One job exists, written as exactly one entry carrying the fixed marker # jsc-assist:assistant patrol:

Job Period Runs Installed by start
patrol derived from the heartbeat TTL (*/2 * * * * at the default TTL of 300 seconds) one patrol round through the caller's CLI yes, always
heartbeat — nothing. This job existed in the previous version and is no longer installable no — install heartbeat exits 6

The heartbeat job is gone on purpose. It used to call heartbeat.sh write every minute, which made a fresh heartbeat prove only that cron was alive. Anything that writes a heartbeat outside a finished patrol round brings that back, so install heartbeat is refused, and install patrol deletes any leftover heartbeat entry from an older install and reports legacy_removed=1. Say that number in the report — a surviving legacy entry silently undoes this whole design.

The period is derived, never guessed. The heartbeat now moves once per patrol round, so the round period has to fit inside the freshness threshold. schedule.sh reads the machine's effective threshold from heartbeat.sh report's ttl= field and picks the largest whole-hour-dividing minute count P with 2 × P × 60 < ttl: one missed round still reads fresh, two missed rounds read stale. At the default 300 seconds that is every 2 minutes; raise JSC_ASSISTANT_HEARTBEAT_TTL to 1800 and it becomes every 12 minutes. Report both numbers (ttl=, period=) so the operator can see the trade-off and change it in one place. --period overrides the calculation and is checked against the same inequality; a period that does not fit exits 6 rather than installing a schedule that keeps the heartbeat permanently stale.

The mechanism follows the platform: crontab on Linux, WSL and macOS, schtasks on Windows. macOS keeps crontab — a launchd user who wants a plist writes it themselves; this skill does not generate one.

Four properties of that script matter enough to state here, because a report that ignores any of them is wrong:

  • It only ever touches its own entries. Install filters out its own old entries by marker and appends the new one; it never rewrites a crontab it failed to read, and it counts everybody else's lines before and after to prove none went missing. Remove takes out its own markers only. Say this in the report — the operator is entitled to know their own cron entries survived.
  • A written entry is not a running entry. WSL does not start cron by default, and this is the machine's most likely state. Exit 1 from install means the entry is on disk and will never fire. Report that as a failure of the start, name sudo service cron start, and say it has to be run again after every WSL restart. Never soften exit 1 into "scheduling is set up".
  • The log lives at $JSC_HOME/assistant/schedule.log, deliberately outside every repository. Do not offer to move it into a project.
  • The entry runs with no human present. The command is installed with </dev/null, so nothing it runs can block on input. A patrol round that stops to ask for a tool permission hangs that round, and the lock it holds stands the next round down until the lock ages out — which is why patrol asks nothing, of anybody, ever.

What a fresh heartbeat actually proves

The heartbeat is written in exactly one place: tools/patrol.sh finish, and finish is called only after that round's result is on the monitor page. So the verdict 新鮮 now proves one thing that is worth proving — the last patrol round ran to the end and its result was recorded — and it still does not prove three others:

  • Not that the round was clean. Four sources are read independently and a round with three failures still records and still beats. The health of a round is 本輪判定 on the monitor page, never the heartbeat.
  • Not that any task in the book moved. The task rows — last_run, next_run, fail_count — are the only evidence about work.
  • Not that the round did anything about what it found. The patrol reports; a human acts. 界線 6.

The failure this design buys is the one worth having: a round that cannot read its sources, cannot reach the wiki, or dies half way writes no heartbeat, so the heartbeat ages past the TTL and every reader sees 過期. A silent patrol is now indistinguishable from a stopped assistant, which is exactly right. status still prints the heartbeat, the schedule and the task book as three separate facts, and the same limit binds whatever gate reads this heartbeat later: a fresh heartbeat is grounds for not blocking, never grounds for saying the assistant is doing its job.

schedule.sh exit codes

Code Meaning What to do
0 install wrote the entry and read it back, the scheduler service is running; remove finished, or there was nothing to remove; status printed its lines Carry on. For status, the state still has to be read out of the installed= fields
1 install wrote the entry, but the cron service is not running — the entry will never fire The start did not succeed. Report the entry as installed and inert, quote the fix (sudo service cron start, and again after each WSL restart), and never claim the assistant will keep itself alive
2 jsc-hooks/hooks/heartbeat.sh was not found, so the TTL cannot be read and the period cannot be derived Report that jsc-hooks is missing or too old (0.3.7 or newer is required) and stop the operation
3 No usable scheduler on this machine Report the platform and that neither crontab nor schtasks was found, and stop. Never fall back to some other mechanism
4 The scheduler operation failed — the existing schedule could not be read for a reason other than "no crontab", or the write or delete returned non-zero Report the stderr text verbatim. A read failure means nothing was written, so the user's other entries are untouched; say so
5 Read-back verification failed — the entry is missing after a successful write, is present twice, is still there after a delete, or somebody else's line count changed Serious. Report it loudly with the printed numbers, and tell the operator to inspect crontab -l by hand before anything else is run
6 Usage error — an unknown subcommand or job name, a missing option value, install heartbeat, a --period that does not fit the TTL, or the patrol CLI could not be determined A defect in the call, not a state of the machine. Correct the command line and run it once more; report a second exit 6 as a defect in this skill and stop

patrol.sh exit codes

One table for all three subcommands. Read collect's codes carefully: 1 and 3 are results, not aborts. A round with failed items still has a section to write, and refusing to write it would hide the failure instead of recording it.

Code Meaning What to do
0 collect: all four items read to the end, empty sources included. finish: heartbeat written, snapshot promoted, lock released. abort: lock released Carry on with the operation's next step
1 collect: partial success — at least one item failed and at least one produced a result Write the page anyway. The section already marks the failed items and the round verdict is 警示. Name the failed items and their note= text in the report
2 finish: jsc-hooks/hooks/heartbeat.sh was not found The round completed and is recorded, but no heartbeat exists to prove it. Report the round as recorded and the heartbeat as not written, say jsc-hooks 0.3.7 or newer has to be installed, and run tools/patrol.sh abort --round {id} to release the lock
3 collect: all four items failed Write the page anyway, with verdict 異常. A page listing four failures is the signal; a missing page is not. Then carry on to finish as usual — the round did complete
4 Another round holds the lock (collect), or the lock is no longer this round's (finish, abort) Not a failure. On collect: report 本輪讓開 and name the holder and its age from the printed lock=busy line, then write nothing and stop. On finish: the previous round overran and was taken over, so this round's result does not count — report it, write no heartbeat, and stop
5 Filesystem failure — the lock could not be created or released, a scratch file could not be written, the snapshot could not be promoted, or heartbeat.sh write returned non-zero Serious. Report it loudly with the stderr text and the path. On a finish failure the round is recorded but unproven: say so plainly and never claim the round beat
6 Usage error — an unknown subcommand, a missing --round, or an option with no value A defect in the call. Correct it and run it once more; report a second exit 6 as a defect in this skill and stop

Boundaries

The six limits in AGENTS.md「助理的界線」 hold for all four operations. Four of them need saying out loud here:

  • This skill never judges a gate. It maintains the heartbeat and prints what the heartbeat says. Whether a stale heartbeat blocks a skill call is decided by a hook, synchronously and offline; nothing in this skill blocks or waves through anything. 界線 2.
  • A patrol round asks nothing. It runs from cron with nobody present, so there is no one to answer and a question hangs the round. Every branch in the patrol steps below resolves without a question: a missing source is recorded as missing, an ambiguous result is recorded verbatim, and a round that cannot proceed aborts and reports. Never call jsc-ask:ask from patrol. 界線 1.
  • A patrol round only ever appends to the monitor page. Read the old page back first, append one section, put the whole page. The contents page gets its own row updated and nobody else's. A page that could not be read is a page that does not get written. 界線 4.
  • A patrol round reports; it never acts on what it found. The 待人處理 rows name an entry point for a human. The patrol does not run that entry point, does not fix a hook, does not update a plugin and does not touch a repository. 界線 3 and 界線 6.
  • stop clearing the heartbeat and removing the schedule is not a breach of 界線 5「不刪除狀態檔」. That limit protects state that records work — the task book, worktrees, wiki pages — from a background process nobody is watching. The heartbeat records one fact only, "the last patrol round finished", and the schedule entry is what keeps rounds running, so a stop that leaves either behind leaves a lie behind. Clearing both is the whole job of stop, and they are the only deletions any operation here performs, both of them entries this skill installed itself. stop touches nothing under tasks/, nobody else's cron entry, no worktree and no wiki page. Do not "restore" this limit later by taking either removal out of stop.

Crash exit needs no cleanup

An assistant that is killed, crashes, or dies with the machine writes no farewell. It does not need to. The heartbeat is a timestamp, not a lock: the last one written stays on disk, ages past the TTL on its own, and every reader from then on sees 過期. No shutdown handler, no cleanup hook and no pid check is involved, so there is nothing left that can fail to run.

The round lock is the one thing a crash does leave behind, and it ages out the same way: patrol.sh collect breaks a lock older than the heartbeat TTL, takes it, and prints lock_broken=1 so the takeover lands on the monitor page instead of happening quietly. The overrun round that lost its lock then gets exit 4 from finish and writes no heartbeat, which is correct — it never reached the end.

That property holds only while nothing fakes a heartbeat. write is called by tools/patrol.sh finish and nowhere else. start does not call it, status does not call it, stop does not call it, no scheduled entry calls it, and no other skill calls it. A heartbeat written by anything that is not a finished round says a round finished when none did, and the reader has no way to tell the difference. This is also why stop removes the scheduled entry before clearing the heartbeat, and never in the other order.

start

start proves the loop works before it schedules it: one patrol round first, then the scheduled entry. It installs no daemon and writes no bare heartbeat.

  1. Run one patrol round. Follow every step of the patrol operation below, start to finish. This is what writes the first heartbeat — there is no shortcut past it, because a heartbeat that no round produced is exactly the lie this design removes. When that round ends without a heartbeat for any reason (collect exit 4, 5 or 6, an empty hash=, a failed wiki write, or finish exit 2, 4 or 5), the start has failed: report the round's outcome and the code, do not run step 2, and do not claim a started assistant. A round that completed with failed items (collect exit 1 or 3) is still a completed round — carry on to step 2 and name the failures in the closing report. Completion condition: patrol.sh finish exited 0, or the failure report naming the step and the code has been printed and no start was claimed.

  2. Confirm the heartbeat. Run jsc-hooks/hooks/heartbeat.sh report and read its state=, ts=, ttl=, pid=, cli=, session= and file= fields. state=fresh is the expected result. Any other state right after a successful round means something rewrote or removed the file in between: report the state, the path and that the heartbeat did not survive its own write, and do not claim a started assistant. Completion condition: the report line was read and either state=fresh was recorded with its seven fields, or the mismatch was reported.

  3. Install the patrol entry. Run tools/schedule.sh install patrol. Judge the result by the schedule.sh exit-code table, and keep the printed entry=, ttl=, period=, legacy_removed=, others_kept= and service= fields for the report. Exit 1 is the case to get right: the entry is installed and inert, so step 4 reports a started assistant whose heartbeat will expire, not a scheduled one. On 2, 3, 4, 5 or 6 nothing is scheduled — report the code, say the round ran but no further round will, and do not claim the assistant will stay alive. Completion condition: the exit code is recorded, and on exit 0 the printed entry line, the TTL, the period, the legacy count and the surviving-entry count are recorded with it.

  4. Report the start. Print the round's verdict and its four item results, the monitor page that was written, the heartbeat path, the local time of ts, the TTL in seconds, pid, cli and session as hints, then the scheduler mechanism, the derived period, the installed entry line, how many legacy heartbeat entries were removed, and how many other entries were left untouched. Close with the notice that matches step 3's outcome, printed literally with {ttl} replaced by the TTL just read and {period} by the derived period:

    Step 3 Notice
    exit 0 助理已啟動,第一輪巡檢跑完了,結果寫上監控頁了,心跳也寫了。排程接上了,之後每 {period} 分鐘跑一輪,每一輪跑完才寫一次心跳。心跳新鮮代表上一輪巡檢真的做完了;那一輪四項有沒有全過,看監控頁的本輪判定。
    exit 1 助理已啟動,第一輪巡檢跑完了,排程條目也寫進去了,但 cron 服務沒在跑,那一筆一次都不會被執行。心跳過了 {ttl} 秒就會過期。請先跑 sudo service cron start,重開 WSL 之後要再跑一次。
    其他結束碼 助理已啟動,第一輪巡檢跑完了,但排程沒接上(結束碼 {code})。不會再有下一輪,心跳過了 {ttl} 秒就會過期,屆時請再跑一次 start。

    Completion condition: the report carries the round verdict, the monitor page name, the path, the local heartbeat time, the TTL, the period, the three hint fields and the scheduler outcome, and exactly one notice above appears with the real numbers.

patrol

One round: read four sources, record the result, then beat. Everything before the heartbeat is read-only except the round's own scratch files. Ask nobody anything.

  1. Collect. Run tools/patrol.sh collect --trigger 排程 (use --trigger 手動 when a person asked for this round). Judge the exit code by the patrol.sh table. Exit 4 stands the round down — report the holder and its age from the printed lock=busy line, and stop; write no page and no heartbeat. Exit 5 and 6 stop the round the same way, with the code and the stderr text. Exit 0, 1 and 3 all carry on to step 2. Record round=, lock_broken=, hash=, page=, verdict=, failed_sources=, every item= line, and the three file paths section_file=, newpage_file= and contents_file=. Completion condition: the round id, the page name and the three file paths are recorded, or the stand-down or the failure was reported and the round stopped.

  2. Check the page name. An empty hash= means jsc-gitea/tools/hash-id could not be found or could not run, so there is no page to write to and nothing can be recorded. Run tools/patrol.sh abort --round {round}, report that the round found its results but has nowhere to put them, name jsc-gitea as missing, and stop. Never invent a page name — a hand-made name lands the content on a page nobody reads. Completion condition: page= holds a MONITOR_{HASH} name, or the abort ran and the round was reported as unrecorded.

  3. Append the section to MONITOR_{HASH}. Hand it to jsc-gitea:wiki with page type MONITOR: read the page back first, then append the whole content of section_file as a new last section and put the whole page. Only exit 4 from the read permits creating the page instead, and then the page body is the whole content of newpage_file, which already carries the basic-data section plus this round's section. Exit 7 and exit 8 mean the old content is unknown: create nothing, write nothing. On any write failure — including exit 3 with no wiki repo configured for MONITOR, which the patrol cannot ask about — run tools/patrol.sh abort --round {round}, report the code, and stop. No record, no heartbeat. Completion condition: the append or the create returned success, or the abort ran and the round was reported as unrecorded with its exit code.

  4. Update this machine's row in MONITOR_CONTENTS. Take the row= line from contents_file — it is already the finished table row. Hand it to jsc-gitea:wiki: read the whole page, match the row whose 主機 and 帳號 columns both equal this round's host= and user=, overwrite that row's remaining columns, and put the whole page back. No matching row means append one. Never overwrite the whole page, and never touch another machine's row — the write semantics here are the opposite of the content page's, and mixing them up deletes other machines' records. On failure, run tools/patrol.sh abort --round {round}, report the code, and stop. Completion condition: exactly one row carries this machine's 主機 and 帳號 values, every other row is byte-identical to what was read, and the put returned success.

  5. Write the heartbeat. Run tools/patrol.sh finish --round {round}. This is the last step for a reason: it is the only thing that turns a fresh heartbeat into a true statement. Judge the exit code by the patrol.sh table — 2, 4 and 5 all mean the round is recorded but unproven, and each has its own report line there. Completion condition: finish exited 0, or the failure was reported as "recorded but no heartbeat" with its code.

  6. Report the round. Print the round verdict, one line per item with its status= and, for a failure, its note=; the monitor page name and the contents row that was written; whether the heartbeat was written; and, when lock_broken=1, that the previous round's lock was taken over because it had aged past the TTL. Close with the 待人處理 rows from the section, verbatim, and nothing else — the patrol names an entry point and stops there. Completion condition: all four items appear in the report, the heartbeat outcome is stated as written or not written, and no suggestion in 待人處理 was acted on.

status

Read-only throughout. This operation creates, modifies and deletes nothing under $JSC_HOME, and it never calls write or clear.

  1. Read the heartbeat through the script. Run jsc-hooks/hooks/heartbeat.sh report and split the line on spaces, taking file= last so a path containing spaces stays intact. Map state= to the verdict: fresh → 新鮮, stale → 過期, invalid → 心跳檔損壞, absent → 不存在. Print 助理未運行 for stale, invalid and absent. Never re-derive the verdict from ts yourself, and never treat invalid as fresh. On exit 2 or 6, follow that code's row, record the heartbeat state as unknown, and carry on to step 2 — the task book is still worth printing. Completion condition: the heartbeat state holds one of 新鮮, 過期, 心跳檔損壞, 不存在 or unknown, and ts, age, ttl, pid, cli, session and file are recorded as read or as empty.

  2. Read the task book. Take the assistant directory from the file= path of step 1, list the regular files directly under its tasks/ subdirectory, and parse each one as key=value lines. Branch on the outcome.

    Outcome Do
    Directory absent Report zero entries. This is a normal result, not an error
    Directory present, no files Report zero entries
    A file cannot be read, or holds no recognisable key Keep it as one row, put the file name in the title column, name the read or parse error in that row, and carry on with the remaining files
    A key is missing from a readable file Print - in that column

    Completion condition: every file under tasks/ produced exactly one row, or zero entries was reported.

  3. Read the schedule. Run tools/schedule.sh status. It writes nothing. Record mechanism=, service=, ttl=, period= and the installed= value of both jobs. A heartbeat job reported as installed is a leftover from an older version: say so, and say start or schedule.sh install patrol removes it. On exit 2, 3 or 6 nothing was read: record the schedule state as unknown with its code and carry on — the heartbeat and the task book still print. Completion condition: both jobs have an installed state, or the schedule state is recorded as unknown with its code.

  4. Print the status table. Lead with the heartbeat block — verdict, last heartbeat time rendered from ts in local time, age in seconds, TTL, cli, session, pid, and the task count. Follow it with the schedule block — mechanism, service state, derived period, and one line per job saying installed or not. Then one row per task carrying state, title, next_run and fail_count, in the order the files were listed. Completion condition: the heartbeat block holds all eight values, the schedule block holds both jobs and the period, and the row count equals the task count from step 2.

  5. Say what the two blocks together mean. Four combinations get an explicit sentence, because each one reads as something it is not:

    Heartbeat Schedule Say
    新鮮 patrol installed, service running 上一輪巡檢跑完了,結果也記上監控頁了,排程還在跑。那一輪四項有沒有全過,要看監控頁的本輪判定
    新鮮 not installed, or service stopped 上一輪巡檢跑完了,但沒有排程在叫下一輪,過了 TTL 心跳就會過期
    過期 or 不存在 patrol installed, service running 排程裝著卻沒有新的心跳,巡檢自己跑失敗了,去看 $JSC_HOME/assistant/schedule.log 與監控頁最新一節
    any heartbeat job installed 舊版的心跳排程還留著,它會蓋掉「心跳等於巡檢跑完」這件事。請跑一次 start,或 schedule.sh install patrol 把它清掉

    Completion condition: every matching sentence is printed, or none of the four combinations applied.

  6. Flag the repeatedly failing tasks. Append 已連續失敗 N 次 to every row whose fail_count is above 0, with N taken verbatim from the file. A broken entry that retries every round with nobody noticing is the reason this field exists, so let no such row leave the table unmarked. Completion condition: every row with fail_count above 0 carries the marker and its number matches the file.

  7. Finish successfully. 助理未運行, an absent tasks/ directory, an empty tasks/ directory and an uninstalled schedule are normal results — never exit non-zero for any of them. Reserve a failure report for a condition none of the tables above covers, and state which path and which error produced it. Completion condition: the report is printed and nothing under $JSC_HOME has been created, modified or deleted.

stop

  1. Record what is being stopped. Run jsc-hooks/hooks/heartbeat.sh report first and keep its state=, ts=, pid=, cli= and file= fields for the closing report — after the clear they are gone for good. state=absent means no round has finished; say so and still run steps 2 and 3, because a scheduled entry can outlive its heartbeat and clear on a missing file is a success, so running both leaves the outcome unambiguous. On exit 2 or 6, follow that code's row, record the previous state as unknown, and carry on to step 2. Completion condition: the previous state and its fields are recorded, or the previous state is recorded as unknown with its code.

  2. Remove the schedule first. Run tools/schedule.sh remove all — both job names, so the patrol entry and any leftover heartbeat entry from an older install both go. This comes before the clear and never after: clear first and the next scheduled round writes a fresh heartbeat over the stopped assistant, and every reader from then on is told a dead assistant is alive. Judge the result by the schedule.sh exit-code table, and keep removed= and others_kept= for the report. On any non-zero code the schedule is still installed: report the code, say plainly that rounds will keep running and the assistant therefore cannot be stopped, name the manual fix (crontab -l to look, then remove the line carrying # jsc-assist:assistant by hand), and skip steps 3 and 4 — clearing a heartbeat that the next round rewrites only hides the problem. Completion condition: remove exited 0 with its counts recorded, or the failure report has been printed and no stop was claimed.

  3. Clear the heartbeat. Run jsc-hooks/hooks/heartbeat.sh clear. On exit 5 the file is still there: report the failure with the script's stderr line and the path, say plainly that every reader still sees a heartbeat claiming a round just finished and that the assistant is therefore not reliably stopped, name the manual fix (delete that path by hand, then run status to confirm 助理未運行), and skip step 4 — the closing notice must not be printed after a failed clear. On exit 2 or 6, follow that code's row and stop the same way. Completion condition: clear exited 0, or the failure report naming the code, the path and the manual fix has been printed and no stop was claimed.

  4. Report the stop and what it means for the gate. Print the previous state and heartbeat time from step 1 and the entries removed in step 2, then this literally:

    助理已停止,排程移除了,心跳也清掉了,其他人的排程一筆都沒動。靠心跳判定的 jsc 技能閘門一讀到沒有心跳就會擋下技能呼叫;閘門目前還沒接線,所以這一刻誰都擋不到。要再工作就先跑一次 start。

    Say it exactly this way. The blocking is the designed consequence of a cleared heartbeat, and whoever stops the assistant has to know it is coming; the clause about the gate being unwired is the part that keeps the notice honest while that is still true. When the gate is wired, that clause is what gets rewritten — not the rest. Completion condition: the notice appears with all three clauses, and the previous state, the heartbeat time and the removal counts are printed above it.

Round lock and a round that will not stop leaving one behind

A patrol round holds $JSC_HOME/assistant/patrol.lock from collect to finish or abort, which spans the wiki writes — the slow part. Two consequences bind every branch above:

  • Every path out of a started round ends in finish or abort. Steps 2, 3 and 4 of patrol each name their abort. A round that stops without either leaves the lock standing until it ages out, which stands the next rounds down for up to one TTL. There is no third option.
  • stop does not remove the lock. It is not a state file that records work, but it is also not this skill's to delete while a round may still be using it; it ages out on its own within one TTL. If an operator reports that every round stands down, tell them the holder and age from the lock=busy line and let them decide — 界線 5 keeps destructive cleanup with the human.