Files
gitea/tools/normalize-name.sh
T
jiantw83andClaude Sonnet 5 96a0772d65 feat(gitea): 為 hash-id 新增 key 子命令並拆出獨立的名稱正規化模組
為什麼:多個技能各自手動拼接計畫名稱與 owner/repo 字串再算雜湊,重複邏輯容易長出分歧的正規化寫法,也讓 hash-id 混雜了業務規則,職責變得不單純。

做了什麼:
- hash-id 新增 key 子命令(hash-id key <計畫名稱> <owner/repo>),將計畫名稱正規化後以 | 分隔接上 owner/repo 再計算雜湊,提供統一的組法入口,讓其他技能不必各自接字串。
- 將正規化規則(Unicode NFC、去除頭尾空白、內部連續空白壓成一個、全形英數轉半形,大小寫維持不變)獨立成 tools/normalize-name.sh,hash-id 只呼叫它,不再內嵌業務規則,保持雜湊職責單純,也方便日後其他工具重用同一套正規化邏輯。

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 10:32:24 +08:00

47 lines
1.6 KiB
Bash
Executable File
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env sh
# normalize-name — 正規化一個名稱字串,讓同一個名稱的不同打法算出同一個結果。
# 用法:
# normalize-name.sh <text>
# 規則(依序套用):
# 1. Unicode NFC 正規化
# 2. 去除頭尾空白
# 3. 內部連續空白(含全形空白)壓成一個半形空白
# 4. 英數全形字元轉半形(只轉英數,其他全形字元如標點不動)
# 大小寫不在正規化範圍內,維持原樣、視為不同名稱。
# 例子:
# normalize-name.sh "我的 計畫" -> 我的 計畫
# normalize-name.sh " Abc1 " -> Abc1
# 結束碼: 0=成功 1=這台機器沒有 python3,算不出正規化結果 2=用法錯誤:沒有給文字
set -eu
if [ "$#" -lt 1 ]; then
echo 'usage: normalize-name.sh <text>' >&2
exit 2
fi
if ! command -v python3 >/dev/null 2>&1; then
echo 'no python3 found' >&2
exit 1
fi
printf '%s' "$1" | python3 -c '
import sys, re, unicodedata
FULLWIDTH_OFFSET = 0xFEE0 # Unicode 全形英數區塊相對半形區塊的固定位移量
def to_halfwidth_alnum(ch):
if ("0" <= ch <= "9") or ("A" <= ch <= "Z") or ("a" <= ch <= "z"):
return chr(ord(ch) - FULLWIDTH_OFFSET)
return ch
name = unicodedata.normalize("NFC", sys.stdin.read())
name = name.strip()
name = re.sub(r"[ \t ]+", " ", name).strip()
# 只轉英數的全形字元,其他全形字元(例如全形標點)不動——只轉英數是規則本身
# 的範圍,不是整段跑 NFKC;NFKC 連標點都會改寫,會讓名稱的視覺呈現跟著變。
name = "".join(to_halfwidth_alnum(c) for c in name)
sys.stdout.write(name)
'