Files
meta/skills/skill-check/SKILL.md
jiantw83 2e237b7674 feat(behaviors): 新增技能行為清單與檢查腳本
What:新增 references/behaviors.md,一支技能一節,共七支技能。每節五列,記下觸發時機、關鍵步驟、外部呼叫、完成條件、可驗證跡象。新增 tools/check-behaviors.sh,比對 skills/ 與這份清單。skill-new、skill-update、skill-delete、skillset-update 加上同步更新清單的步驟。skill-check 把這支腳本併進第一組稽核。

Why:技能驗證原本沒有基準,稽核只能靠眼睛比對 SKILL.md。十個 domain 每輪都要重做一遍,還會漏掉。行為清單當基準,技能改了、清單沒跟著改,就是漂移。漂移交給程式判定才穩。

How:腳本檢查節數、節名、節序、每節一張表、五個欄位齊全、內容欄非空。退出碼 0 代表相符,1 代表不符,2 代表用法錯誤,3 代表找不到清單或找不到 skills 目錄。domain 名以 plugin.json 的 name 為準,checkout 目錄名只是退路。四支異動技能在同一個 PR 內改清單,並照退出碼分流。

Who:屬於「技能行為清單」這件需求,提供技能驗證的參考基準。
2026-08-31 13:37:04 +08:00

14 KiB

name, description
name description
skill-check Routine compliance, script, hook, flow-efficiency, and cost-efficiency audit of the whole jsc skill set with no change request in hand. Sync every domain repo from the Gitea canonical marketplace, then run three parallel groups - lint-scripts.sh plus check-behaviors.sh plus hook smoke, the guidelines.md checklist audit, and a review of parallelism, tool extraction, repeated interaction, redundant checks, misplaced gates, and avoidable token, sub-agent, API, scan, or interaction cost. Confirm compliance fixes and optimization suggestions before applying them, re-check until accepted fixes pass, then open a PR per affected repo via jsc-git pr. Use for periodic or on-demand skill-set checks; not for applying a change request (use skillset-update) or editing one skill (use skill-update).

skill-check — audit compliance, flow efficiency, and cost efficiency

Single source of guidelines: ../../references/guidelines.md.

Flow

  1. Run tools/sync-domains.sh to sync every domain repo of the Gitea canonical marketplace. Completion condition: the script exits 0 and prints one domain<TAB>path line per marketplace domain — exit 0 is the only code that means every repo is present and current. Exit 3 means some repos were not updated: reconcile every path named on stderr (commit or stash the dirty tree, or fix the failing pull) and rerun; when the user confirms a dirty tree is intentional local work, record that decision and continue on the local version — never read exit 3 as current. Exit 2 means a domain could not be cloned. Exit 1 means the root could not be derived, gitea.sh was not found, or the canonical marketplace was unreadable; when stderr says the root could not be derived, set JSC_PLUGINS_ROOT to the directory that holds the domain repos and rerun, because under a plugin install the script sits in the CLI's plugin cache and its built-in guess lands there instead of the domain workspace. Resolve 2 and 1 before continuing.

    The three review groups of step 2 all read this synced tree, so the sync finishes first.

  2. Run the three review groups over the synced repos. They are independent — every one only reads, none writes a file — so launch all three in parallel and merge their results in step 3.

    Group 1 — validate scripts, behavior lists, and hooks.

    1. For every synced domain repo, run tools/lint-scripts.sh {domain-path}. One run per domain, and the runs go in parallel — no domain's verdict depends on another's. The tool covers three checks in one pass: sh -n syntax, executable bit, and an exit-code declaration in the file header. Route each exit code: 0 — the domain's scripts pass all three; 1 — the failing items are printed as {file}:{check}:{detail}, so report each one; 2 — usage error, the tool takes exactly one argument; 3 — nothing was scanned, because the path is missing or the domain has neither tools/ nor hooks/. Record exit 3 as 「無腳本可掃」; a domain with no script directory is not a failure, but exit 3 is never a pass.
    2. For every synced domain repo, run tools/check-behaviors.sh {domain-path}. One run per domain, and the runs go in parallel alongside the lint-scripts.sh runs — no domain's verdict depends on another's. It compares references/behaviors.md against skills/: section per skill, dictionary order, one table per section, five rows, no empty content cell. Route each exit code: 0 — that domain's behavior list matches; 1 — the mismatches are printed on stderr as {檔案}:{技能名}:{說明}, so report every one as a compliance failure with the skill it belongs to; 2 — usage error, the tool takes exactly one argument; 3 — nothing was checked, because references/behaviors.md is missing, skills/ is missing, or no SKILL.md was found. Record exit 3 as 「無清單可查」with the cause from stderr and carry it into the step 3 merge; a domain with no behavior list is a compliance failure, and exit 3 is never a pass.
    3. For every shell script directly named by a SKILL.md, confirm the skill routes every exit code the script's header declares. lint-scripts.sh proves the script exists and declares its codes; this check is the other half — that the caller branches on each of them. Report evidence as skill file:line -> script path.
    4. When the jsc-hooks domain is present, run jsc-hooks/tools/wire-cli.sh smoke {cli} for every CLI reported by jsc-cli/tools/detect-clis.sh; the per-CLI smokes run in parallel. When no CLI is detected, run jsc-hooks/tools/wire-cli.sh smoke codex as the minimum hook behavior check and label it 「預設 hook smoke」 in the report. Use smoke, not purge or rewiring actions, and set JSC_READONLY=1 for the whole audit so a mistyped sub-command is refused in code (exit 6) instead of rewiring the machine; status and smoke are unaffected by that variable. Route each smoke exit code: 0 — the run passed its own assertions; 2 — usage error, so fix the CLI code and rerun; 4 — the smoke failed, which includes the script's own result-line count not matching what it expected. Read the count from the script's lines<TAB>{數量} output line; never write the number into this skill. The script counts its own result lines and asserts them, so a hardcoded number here goes stale the moment a hook or a decision path is added — an out-of-date count in a SKILL.md is exactly what misled the previous audit.
    5. When a hook or script smoke fails, route it as a compliance failure with script name, exit code, output summary, and proposed fix. Do not continue to report the affected hook as compliant.

    Group 2 — audit every skill of every domain against the guidelines.md audit checklist. This group MUST run as a sub agent, one sub agent per domain repo, and those sub agents run in parallel. Each sub agent reports its findings: skill, failed checklist item, evidence (file:line), proposed fix. Cover the checklist's four flow checks by name, not only the naming and language items:

    • Every step number, file path and section title the skill references — inside itself and in other files — really exists (the pointer points at something).
    • Every step ends in a checkable completion condition, with no vague wording.
    • Every external call (script, API, other skill) states what to do on failure and routes every exit code.
    • No gate the skill installs blocks the only path that lifts that gate.

    Four checklist items are already decided by group 1 and must not be re-run here: sh -n on every tools/ and hooks/ script, script existence with the executable bit, the hook smoke, and the references/behaviors.md match. Tell each sub agent to skip those four and leave them blank; the main agent fills them in from the group 1 verdicts when merging in step 3. Re-scanning the same files in every domain sub agent buys nothing — group 1 already scanned them all, with the same tools, on the same synced tree.

    Group 3 — a flow and cost optimization review, kept separate from the compliance audit. Each aspect MUST run as a sub agent, and the six aspects run in parallel with each other and with groups 1 and 2:

    Aspect Scope
    1 Parallelism Steps that run in series today but have no data dependency and can run in parallel
    2 Tool extraction SKILL.md text flows with clear inputs and outputs that should move to tools/, including hook-enforceable rules that still rely on prompts
    3 Repeated interaction The same user question, repository fact, wiki page, API result, or file content being collected more than once
    4 Redundant checks Completion conditions or verification steps that overlap, or a later step that necessarily covers an earlier check
    5 Gate timing Gates that run too early or too late, causing wasted work before a block or blocking the only path that clears the gate
    6 Cost efficiency Avoidable token, sub-agent, API, file-scan, full-repo audit, or user-interaction cost that can be reduced without weakening correctness

    Each optimization finding reports skill, aspect, evidence (file:line), current flow step count, proposed flow step count, what time or interaction it saves, what cost it saves, current cost driver, proposed cost driver, whether correctness decreases, and which protection would be weakened if any. Cost savings may be token volume, sub-agent count, API calls, file scans, full-repo audits, or user prompts. Keep optimization findings separate from compliance failures.

    Completion condition for all three groups: every domain has a lint-scripts.sh verdict and a check-behaviors.sh verdict, every script named by a SKILL.md has an exit-code-routing verdict, and every smoked CLI has a smoke exit code plus the lines value the script printed for it; every domain has a group 2 audit result that names a verdict for all checklist items — the four flow checks included, and the four group 1 items left blank for the step 3 merge rather than re-scanned; and every one of the six aspects has returned a verdict for every domain, 「無發現」 where an aspect found nothing.

  3. Merge the three groups, then present compliance failures and optimization findings separately via the jsc-ask:ask decision tree. Merging means one thing in code: fill the four skipped checklist items of every group 2 sub agent report from the matching group 1 verdicts, so each domain ends with one complete checklist and no item counted twice.

    • Compliance failure options: apply the proposed fix / skip / custom fix. Every option states its impact scope, for example skipping leaves the skill non-compliant until the next audit.
    • Optimization options: apply / defer / custom. Any suggestion that weakens a protection must name the protection it removes and must not be applied unless the user explicitly accepts that tradeoff. Cost optimization may move, merge, cache, or narrow checks; it must not delete a compliance check only because it is expensive.

    Completion condition: every domain's checklist is complete after the merge, and every compliance failure and every optimization finding has a recorded decision.

  4. Apply the confirmed fixes and accepted optimizations — the file-change part MUST run as a sub agent, one sub agent per affected domain repo, and those sub agents run in parallel: each repo's files are independent. A fix that changes a skill's behavior also updates that skill's ## {name} section in the same repo's references/behaviors.md, in the same pass, so the fix and the behavior list land in one PR. Then run tools/sync-skill-manifest.sh {domain-path} directly (no sub agent needed) for each affected domain repo to refresh that domain README's 「Skills 目錄」 section and bump the version in all three manifests. Route each exit code: 0 — the README block and all three manifests are synced; 1 — the domain path, skills/, README.md, the JSC-SKILLS markers, a SKILL.md, a manifest, or a manifest version field is missing, so fix the named cause on stderr and rerun; 2 — usage error, the script takes exactly one argument; any other code — the script runs under set -e, so treat it as an environment fault and stop, never as a successful sync. Completion condition: every affected repo carries the changes, the matching references/behaviors.md update for every fix that changed a skill's behavior, and the manifest bump.

  5. Sync the canonical marketplace — a required step, never optional. The canonical pair lives in plugins/meta and every domain repo carries a byte-identical copy, so a fix that leaves the copies apart makes some repos register a stale plugin set. Run tools/sync-marketplace.sh {domain} {repo-url} {description} once with an existing entry's own current values (rewriting the same entry is idempotent); the script rewrites both canonical files and copies them into every domain repo. Route each exit code:

    • Exit 3 — written, but some domain repo is not present locally. Run tools/sync-domains.sh, then rerun this step.
    • Exit 2 — usage error: the script takes exactly three arguments. Fix them and rerun.
    • Exit 1 — the root could not be derived, python3 is missing, a canonical file was unreadable, or copies differ byte for byte. Read stderr and fix the named cause: install python3 for the second; for the root case set JSC_PLUGINS_ROOT to the directory that holds the domain repos, because under a plugin install the script sits in the CLI's plugin cache and its built-in guess lands there. Then rerun.
    • Exit 0 — every copy holds identical bytes; the script verifies that itself.

    Completion condition: the script exits 0 and prints the touched paths.

  6. Re-run the group 1 script, behavior-list, and hook validation, re-check the guidelines.md audit checklist for every touched skill, then re-run the optimization aspect that produced each accepted optimization. These three re-runs are as independent as the first pass, so run them in parallel and merge them the same way step 3 did. On any compliance failure, return to step 3: confirm and fix again, until all accepted compliance fixes pass. On an accepted optimization that does not produce the promised step reduction or cost reduction, or still weakens correctness beyond the recorded decision, return to step 3 for a new decision. Completion condition: tools/lint-scripts.sh exits 0 or 3 for every domain, tools/check-behaviors.sh exits 0 for every domain, every hook smoke exits 0 with the lines count the script itself asserted, all checklist items pass, and every accepted optimization has a matching verification result.

  7. Call jsc-git:pr once per affected domain repo to open a Push Request. Completion condition: every affected repo has a PR URL, and all URLs are reported in one table with the format in ../../references/pr-report.md.