feat(assistant): 存取庫掃描接上 {repo} 代入點
新增 tools/scan-repos.sh:掃出工作目錄底下的存取庫,交出 {repo} 要代的
那份清單,順便對出每一個存取庫的 REPO_{HASH} 盤點頁頁名(雜湊規則與
盤點頁那一支相同,實測值一致)。
掃描起點刻意沒有預設值。給一個預設就等於猜,而猜錯的後果不是掃不到,
是在猜錯的那些目錄底下跑指令——往上一層是家目錄、再往上是整台機器,
而那一輪沒有人看得到它跑到哪裡去了。沒設就回 3 並說要設哪一個變數。
深度預設一層、上限三層;點開頭的目錄、符號連結、路徑帶空白或殼層特殊
字元的三種一律不交出去,但逐個印出來——那是真的存在卻沒被盤點到的
存取庫,只印一個總數會讀成整批都掃過了。
run-due.sh 改成兩個代入點都代:{cli} 取偵測到的 CLI 代號、{repo} 取
掃到的存取庫,兩個都帶的那一筆目標數是乘積。待辦簿那一筆的 repo 欄
有值時只代那一個,值可以是路徑也可以是 {owner}/{repo};對不出來就
印 held=,不退回全部存取庫——那一筆指名了一個目標。代入之後再驗一次
禁止字元,因為代進去的值是這一輪現場算出來的。
schedule.sh 把 JSC_ASSIST_SCAN_ROOT 與 JSC_ASSIST_SCAN_EXCLUDE 快照進
排程條目。少了掃描起點,帶 {repo} 的內建項每一輪都代不出目標,而那一輪
只印一行 held=,看起來像這一批還沒接上、不像一個變數沒進條目。
同時補兩個只有真的執行才發現得了的缺陷:
一、cron 一行的長度上限只在寫入那一刻由 crontab 自己擋,--dry-run 那一路
根本不碰它,所以預演每一次都過。實測踩過一次:整條 PATH 快照進條目那一批
預演全綠、安裝回結束碼 4。改成建好條目就量,兩個模式都印 entry_len= 與
上限,達到上限當場拒絕並指出長度是從哪幾個快照變數來的。
二、run-due.sh 截錯誤訊息用的 cut -c 數的是位元組不是字元,中文字剛好被
截在中間就在報告上留下一個替代字元。改成截完再過一次 iconv -c 丟掉那個
不完整的序列。這一種不會有人來報:亂碼不影響結束碼。
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -76,6 +76,7 @@ Every tool below is addressed through `{CURRENT}/{plugin}`, with `{CURRENT}` sta
|
||||
| the task book, the only writer there is | `{CURRENT}/jsc-assist/tools/tasks.sh` |
|
||||
| the built-in check items, reconciled against the delegation list | `{CURRENT}/jsc-assist/tools/seed-tasks.sh` |
|
||||
| the due built-in check items, actually run | `{CURRENT}/jsc-assist/tools/run-due.sh` |
|
||||
| the repositories under the working directory, scanned for `{repo}` | `{CURRENT}/jsc-assist/tools/scan-repos.sh` |
|
||||
| the heartbeat | `{CURRENT}/jsc-hooks/hooks/heartbeat.sh` |
|
||||
| the status event stream | `{CURRENT}/jsc-hooks/tools/report-status.sh` |
|
||||
| the wiki, through `jsc-gitea:wiki` | `{CURRENT}/jsc-gitea/tools/gitea.sh` |
|
||||
@@ -185,7 +186,7 @@ Four properties of that script matter enough to state here, because a report tha
|
||||
- **A written entry is not a running entry.** WSL does not start cron by default, and this is the machine's most likely state. Exit 1 from `install` means the entry is on disk and will never fire. Report that as a failure of the start, name `sudo service cron start`, and say it has to be run again after every WSL restart. Never soften exit 1 into "scheduling is set up".
|
||||
- **The log lives at `$JSC_HOME/assistant/schedule.log`**, deliberately outside every repository. Do not offer to move it into a project.
|
||||
- **The entry runs with no human present.** The command is installed with `</dev/null`, so nothing it runs can block on input. A patrol round that stops to ask for a tool permission hangs that round, and the lock it holds stands the next round down until the lock ages out — which is why `patrol` asks nothing, of anybody, ever.
|
||||
- **The entry carries its own environment.** cron gives it a short `PATH`, no settings file and no tty, so `schedule.sh` writes three things into the entry: the CLI resolved to an absolute path with `command -v`, a snapshot of the wiki variables taken at install time (`GITEA_HOST`, `GITEA_TOKEN`, `JSC_HOME`, `JSC_ASSISTANT_HEARTBEAT_TTL` and every set `JSC_WIKI_REPO*` — `JSC_WIKI_REPO`, `JSC_WIKI_REPO_MONITOR` for the monitor page and `JSC_WIKI_REPO_CONTENTS` for the directory page, which the script picks up from the environment rather than from a hardcoded list), and `JSC_GITEA_CONFIRM=yes`, because the write confirmation only recognises a tty and an unattended round has nobody to confirm. Two consequences belong in every report: the entry holds a copy of the token, so the crontab file has to stay readable by its owner alone, and a changed variable only reaches the entry after another `install`. `install` prints the snapshotted names in `env_snapshot=` and masks the token in every entry it prints — never print an entry read from `crontab -l` yourself.
|
||||
- **The entry carries its own environment.** cron gives it a short `PATH`, no settings file and no tty, so `schedule.sh` writes three things into the entry: the CLI resolved to an absolute path with `command -v`, a snapshot of the wiki variables taken at install time (`GITEA_HOST`, `GITEA_TOKEN`, `JSC_HOME`, `JSC_ASSISTANT_HEARTBEAT_TTL`, `JSC_ASSIST_SCAN_ROOT` and `JSC_ASSIST_SCAN_EXCLUDE` for the repository scan, and every set `JSC_WIKI_REPO*` — `JSC_WIKI_REPO`, `JSC_WIKI_REPO_MONITOR` for the monitor page and `JSC_WIKI_REPO_CONTENTS` for the directory page, which the script picks up from the environment rather than from a hardcoded list), and `JSC_GITEA_CONFIRM=yes`, because the write confirmation only recognises a tty and an unattended round has nobody to confirm. Two consequences belong in every report: the entry holds a copy of the token, so the crontab file has to stay readable by its owner alone, and a changed variable only reaches the entry after another `install`. `install` prints the snapshotted names in `env_snapshot=` and masks the token in every entry it prints — never print an entry read from `crontab -l` yourself.
|
||||
- **`install` prints the permission rules that round needs.** One `allow_rule=` line each, every path a full literal under `current` — no variable, no tilde, and no wildcard inside the path, because a rule holding one matches nothing. Hand them to the operator verbatim: an unattended round that hits a permission prompt hangs until the lock ages out, and nobody is there to approve it. `Write(...)` rules do nothing for file writes — only `Edit(...)` is recognised — so never turn a printed `Edit` rule into a `Write` one.
|
||||
- **`install` writes the tool root into the entry.** The round it schedules cannot resolve the root for itself, so `schedule.sh` puts the literal path into the entry's prompt as `工具根目錄={path}` and prints the same value as `patrol_root=`. Report that value, and treat any hand-edit of the entry that drops it as breaking every future round: from then on each one stops at step 0 with nothing recorded.
|
||||
|
||||
@@ -263,7 +264,7 @@ The delegation list holds one row per jsc skill and records whether that skill c
|
||||
|
||||
**A `probe` that cannot be substituted falls back to `remind` and is reported, never dropped.** The path is not `{root}/jsc-{domain}/...`, the command carries a `$` or a `~`, a brace other than the three known holes survived, the root could not be resolved, or the script is not on this machine: each prints `probe_bad=` and the entry is still seeded, as a reminder. Not seeding it would mean removing it on `apply`, so one mistyped cell upstream would delete an entry carrying its own `last_run` and `fail_count` history.
|
||||
|
||||
**`{root}` is substituted here; `{cli}` and `{repo}` are deliberately left in the value.** The CLI list has to be detected and the repositories have to be scanned, and neither is known at seeding time. Expanding into several entries instead would give one `spec_key` several files — the reconcile treats that as `dup=` and refuses to touch any of them — and would go stale the moment a CLI is installed or a repository cloned, with re-expansion costing the history it just protected. So the hole stays, and one rule pays for it: **an `action` containing a brace is not yet a runnable command and must never be executed as written.** `run-due.sh` is what executes one, and it is the only thing that does: `tasks.sh` stores the value, `due.sh` prints it, `status` and the monitor page print it. That tool substitutes `{cli}` from `jsc-cli/tools/detect-clis.sh` and runs once per target, with the targets' results together counting as the entry's one success or failure. `{repo}` has no source yet — the repository scan is not built — so an entry carrying that hole is held, reported, and left with its `last_run` and `fail_count` untouched. **Holding is not failing**: an entry that cannot be given a target did nothing wrong, and recording it as a failure grows a counter nobody can bring down by fixing that entry, which is exactly what that counter exists to make visible.
|
||||
**`{root}` is substituted here; `{cli}` and `{repo}` are deliberately left in the value.** The CLI list has to be detected and the repositories have to be scanned, and neither is known at seeding time. Expanding into several entries instead would give one `spec_key` several files — the reconcile treats that as `dup=` and refuses to touch any of them — and would go stale the moment a CLI is installed or a repository cloned, with re-expansion costing the history it just protected. So the hole stays, and one rule pays for it: **an `action` containing a brace is not yet a runnable command and must never be executed as written.** `run-due.sh` is what executes one, and it is the only thing that does: `tasks.sh` stores the value, `due.sh` prints it, `status` and the monitor page print it. That tool substitutes `{cli}` from `jsc-cli/tools/detect-clis.sh` and `{repo}` from `jsc-assist/tools/scan-repos.sh`, then runs once per target, with the targets' results together counting as the entry's one success or failure. An entry carrying both holes runs once per pair — three repositories and five CLIs is fifteen targets. An entry whose `repo=` field names a repository is given that one repository only, matched against the scan by working directory or by `{owner}/{repo}`; a name the scan did not find is held rather than widened back to every repository, because that entry asked for one target. **The scan refuses to guess its starting point**: without `JSC_ASSIST_SCAN_ROOT` it reports that and nothing else, because guessing one layer too high turns the working directory into the home directory and then into the whole machine, and the commands would run there with nobody watching. So an entry carrying `{repo}` on a machine where that variable never reached the scheduled entry is held with the scan's own words as the reason — that is what tells a person which variable to set and that a fresh `schedule.sh install patrol` is what carries it into the entry. **Holding is not failing**: an entry that cannot be given a target did nothing wrong, and recording it as a failure grows a counter nobody can bring down by fixing that entry, which is exactly what that counter exists to make visible.
|
||||
|
||||
**A `pending` row stays a reminder but is never reported as an ordinary one.** The reason lives in the report, as a `pending=` line, and nothing is written into the task file. Putting the marker in the title would change the `id` and cut that entry's history, and the drift check compares `kind`, `action`, `trigger` and `recur` but not the title, so a stale marker would never be caught; adding a sixteenth key would change the task book's fixed storage format for a piece of upstream prose the task book has no way to edit later. `status` runs `plan`, so `pending=`, `held=` and `drift=` all reach a human in the same place.
|
||||
|
||||
@@ -329,7 +330,7 @@ That property holds only while nothing fakes a heartbeat. **`write` is called by
|
||||
|
||||
3. **Confirm the heartbeat.** Run `{CURRENT}/jsc-hooks/hooks/heartbeat.sh report` and read its `state=`, `ts=`, `ttl=`, `pid=`, `cli=`, `session=` and `file=` fields. `state=fresh` is the expected result. Any other state right after a successful round means something rewrote or removed the file in between: report the state, the path and that the heartbeat did not survive its own write, and do not claim a started assistant. Completion condition: the report line was read and either `state=fresh` was recorded with its seven fields, or the mismatch was reported.
|
||||
|
||||
4. **Install the patrol entry.** Run `{CURRENT}/jsc-assist/tools/schedule.sh install patrol`. Judge the result by the schedule.sh exit-code table, and keep the printed `entry=`, `ttl=`, `period=`, `legacy_removed=`, `others_kept=`, `env_snapshot=`, `patrol_root=`, every `allow_rule=` line and `service=` for the report. Exit 1 is the case to get right: the entry is installed and inert, so step 5 reports a started assistant whose heartbeat will expire, not a scheduled one. Exit 6 with a CLI executable that is not on `PATH` is the second one: nothing was installed, and the fix is to install that CLI or to pass `--patrol-cmd`, not to write a bare command name into the entry. On 2, 3, 4, 5 or 6 nothing is scheduled — report the code, say the round ran but no further round will, and do not claim the assistant will stay alive. Completion condition: the exit code is recorded, and on exit 0 the entry line, the TTL, the period, the legacy count, the surviving-entry count, the snapshotted variable names, the tool root the entry carries and the allow rules are recorded with it.
|
||||
4. **Install the patrol entry.** Run `{CURRENT}/jsc-assist/tools/schedule.sh install patrol`. Judge the result by the schedule.sh exit-code table, and keep the printed `entry=`, `entry_len=`, `ttl=`, `period=`, `legacy_removed=`, `others_kept=`, `env_snapshot=`, `patrol_root=`, every `allow_rule=` line and `service=` for the report. `entry_len=` is measured against cron's per-line limit in both the real install and `--dry-run`, because that limit is enforced by `crontab` itself: a snapshot that pushes the entry over it fails the install with a bare "command too long" while every dry run stays green. Exit 1 is the case to get right: the entry is installed and inert, so step 5 reports a started assistant whose heartbeat will expire, not a scheduled one. Exit 6 with a CLI executable that is not on `PATH` is the second one: nothing was installed, and the fix is to install that CLI or to pass `--patrol-cmd`, not to write a bare command name into the entry. On 2, 3, 4, 5 or 6 nothing is scheduled — report the code, say the round ran but no further round will, and do not claim the assistant will stay alive. Completion condition: the exit code is recorded, and on exit 0 the entry line, the TTL, the period, the legacy count, the surviving-entry count, the snapshotted variable names, the tool root the entry carries and the allow rules are recorded with it.
|
||||
|
||||
5. **Report the start.** Print the round's verdict and its four item results, the monitor page that was written, the heartbeat path, the local time of `ts`, the TTL in seconds, `pid`, `cli` and `session` as hints, then the scheduler mechanism, the derived period, the installed entry line as the script printed it with the token already masked, the `patrol_root=` the entry carries — that is what every later round reads its tool root from — how many legacy heartbeat entries were removed, and how many other entries were left untouched. Then hand over the two operator items the install printed: the `allow_rule=` lines verbatim, so the unattended round never meets a permission prompt, and the reminder that the entry holds a snapshot of the listed variables including the token — keep the crontab file readable by its owner alone, and run `install` again after any of those variables changes.
|
||||
|
||||
@@ -399,7 +400,7 @@ One round: read five sources, record the result, then beat. Everything before th
|
||||
|
||||
Completion condition: `link-check.sh` exited 0 over the block's URL and the script exited 0 with exactly one `## MONITOR_{HASH}` block on the page carrying this round's values, or exit 3 from the upsert or a non-zero `link-check.sh` was reported as an unwritten directory entry and the round carried on, or one of the other non-zero codes — `wiki-url`'s included — was reported after the abort ran.
|
||||
|
||||
5. **Run the built-in check items that are due.** Run `{CURRENT}/jsc-assist/tools/run-due.sh run --root {CURRENT} --rows {the `due_rows_file=` step 1 printed}`. **That path is not optional and there is no default.** The judging step writes its output into the directory whoever called it chose, so a default would point somewhere else — and it did: the tool shipped with one, and every round read a file a person had left behind by running the judge by hand. That file sat on this machine for 67 hours while each round acted on it, one round even running an entry that had already been removed. Nothing looked wrong, because the stale list had been correct when it was written and its contents happened not to change. Passing the path makes each round say which round's data it is acting on; the tool also refuses a list older than the heartbeat TTL, because a list older than that cannot describe this round. This is the one step of the round that changes something outside the round's own files, and it is deliberately narrow: it runs only the entries whose `action` is a command and whose `spec_key` is set, so a reminder, a skill name and anything a person entered by hand are all left alone. Judge the exit code by the run-due.sh table, and keep every `done=`, `failed=`, `held=`, `skip=`, `write_failed=`, `target_ok=` and `target_fail=` line plus the summary counts for the report. **`cli_missing=` is the one to read carefully.** It names the CLIs the scheduled entry recorded at install time that this round could not detect, and it is the only place that gap shows: every target that was detected still ran, still passed, and the round still reports "all of them ran", because without that comparison nothing knows how many there should have been. One round reported four CLIs all passing on a machine with five installed. The usual cause is a CLI whose executable sits in a directory that expires — one here lives under a process-numbered multishell path — while the entry's `PATH` was snapshotted at install time. Report the named CLIs, say that a clean pass is not the same as a full pass, and name re-running the schedule install as what refreshes the snapshot. **No exit code from this step stops the round.** Exit 1 means an entry's command failed and that entry now carries one more failure — that is a finding, not a broken round; exit 4 means a write-back failed, so the same entry will run again next round, which is worth saying out loud; exit 2, 5 and 6 mean nothing ran, and the round still has a result to record. Completion condition: the exit code and the summary counts are recorded, and step 6 was reached whatever that code was.
|
||||
5. **Run the built-in check items that are due.** Run `{CURRENT}/jsc-assist/tools/run-due.sh run --root {CURRENT} --rows {the `due_rows_file=` step 1 printed}`. **That path is not optional and there is no default.** The judging step writes its output into the directory whoever called it chose, so a default would point somewhere else — and it did: the tool shipped with one, and every round read a file a person had left behind by running the judge by hand. That file sat on this machine for 67 hours while each round acted on it, one round even running an entry that had already been removed. Nothing looked wrong, because the stale list had been correct when it was written and its contents happened not to change. Passing the path makes each round say which round's data it is acting on; the tool also refuses a list older than the heartbeat TTL, because a list older than that cannot describe this round. This is the one step of the round that changes something outside the round's own files, and it is deliberately narrow: it runs only the entries whose `action` is a command and whose `spec_key` is set, so a reminder, a skill name and anything a person entered by hand are all left alone. Judge the exit code by the run-due.sh table, and keep every `done=`, `failed=`, `held=`, `skip=`, `write_failed=`, `target_ok=` and `target_fail=` line plus the summary counts for the report. **`cli_missing=` is the one to read carefully.** It names the CLIs the scheduled entry recorded at install time that this round could not detect, and it is the only place that gap shows: every target that was detected still ran, still passed, and the round still reports "all of them ran", because without that comparison nothing knows how many there should have been. One round reported four CLIs all passing on a machine with five installed. The usual cause is a CLI whose executable sits in a directory that expires — one here lives under a process-numbered multishell path — while the entry's `PATH` was snapshotted at install time. Report the named CLIs, say that a clean pass is not the same as a full pass, and name re-running the schedule install as what refreshes the snapshot. **`repos=` and `repo_scan=` are the same kind of reading for `{repo}`.** `repos=0` with `repo_scan=rc3` means the scan had no starting point, so every entry carrying `{repo}` was held this round — report the variable to set, not "there is nothing to inventory". `repo_scan=rc1` means the scan handed over what it could and named the rest: each `repo_scan_note=` line is a repository that exists on this machine and was **not** inventoried this round, either because its path carries a space or a shell metacharacter or because it has no `origin` to compute a `REPO_{HASH}` page name from. Those lines go into the report as findings; a round that prints only the count reads as a full sweep. **No exit code from this step stops the round.** Exit 1 means an entry's command failed and that entry now carries one more failure — that is a finding, not a broken round; exit 4 means a write-back failed, so the same entry will run again next round, which is worth saying out loud; exit 2, 5 and 6 mean nothing ran, and the round still has a result to record. Completion condition: the exit code and the summary counts are recorded, and step 6 was reached whatever that code was.
|
||||
|
||||
6. **Write the heartbeat.** Run `{CURRENT}/jsc-assist/tools/patrol.sh finish --round {round}`. This is the last step for a reason: it is the only thing that turns a fresh heartbeat into a true statement. Judge the exit code by the patrol.sh table — 2, 4 and 5 all mean the round is recorded but unproven, and each has its own report line there. Completion condition: `finish` exited 0, or the failure was reported as "recorded but no heartbeat" with its code.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user