fix(議題解析): 圍欄與表格的邊界,五個會污染下游的解析缺陷
code review 逐項驗出來的,全部可重現:
- `~~~` 圍欄完全沒被認出來,裡面的井字號會被當成段落標題。
- 段落內的圍欄不影響清單解析,於是程式碼範例裡的減號變成假的驗收標準。
這是最嚴重的一個——輸出多出一條沒有人寫過的標準。
- 表格欄位裡逸脫的直線 `\|` 會把欄位切斷,內容整段消失。
- 同一段落裡若出現第二條分隔列,它會變成一筆 {term:'---'} 的假名詞。
- 圍欄開了沒關時,其後內容的歸屬沒有明確定義。
改法是把「圍欄裡的東西不是內容」這件事收斂到 eachLine 處理一次,段落切分、
清單、表格三者都靠它,而不是各自寫一份半套的判斷。圍欄需同種標記才算關閉;
沒關就到結尾時,其後內容一律算在圍欄內,這與 markdown 的實際渲染一致。
巢狀清單改為明確攤平並寫進註解:需求議題的模板沒有巢狀,真的出現時寧可多帶
一項,也不要無聲吃掉內容。
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+68
-23
@@ -3,11 +3,42 @@
|
||||
*
|
||||
* 純函式,不碰網路也不碰檔案系統:抽取類腳本的解析全部走這裡,
|
||||
* 解析規則只寫一次,需求議題與工作包議題共用同一套。
|
||||
*
|
||||
* 全檔的共同前提是「圍欄裡的東西不是內容」:議題裡放 mermaid 或程式碼是常態,
|
||||
* 那裡面的井字號不是標題、減號不是清單項、直線不是表格。這件事只在 eachLine
|
||||
* 處理一次,其餘函式都靠它。
|
||||
*/
|
||||
|
||||
/**
|
||||
* 依 `## 標題` 切出各段落。
|
||||
* 圍欄區塊(```)內的內容一律不當成標題,否則 mermaid 或程式碼裡的井字號會把段落切碎。
|
||||
* 逐行走過內容並標註圍欄狀態。
|
||||
* 圍欄以 ``` 或 ~~~ 開啟,且要同一種標記才算關閉——混用時後者只是普通文字。
|
||||
* 圍欄沒關就到結尾時,其後的內容一律算在圍欄內,這與 markdown 的實際渲染一致。
|
||||
* @param {string} text
|
||||
* @returns {Generator<{line: string, inFence: boolean, isFence: boolean}>}
|
||||
*/
|
||||
function* eachLine(text) {
|
||||
let fence = null;
|
||||
|
||||
for (const line of (text ?? '').split('\n')) {
|
||||
const marker = line.match(/^\s*(`{3,}|~{3,})/)?.[1]?.[0];
|
||||
|
||||
if (marker && fence === null) {
|
||||
fence = marker;
|
||||
yield { line, inFence: true, isFence: true };
|
||||
continue;
|
||||
}
|
||||
if (marker && marker === fence) {
|
||||
fence = null;
|
||||
yield { line, inFence: true, isFence: true };
|
||||
continue;
|
||||
}
|
||||
yield { line, inFence: fence !== null, isFence: false };
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 依 `## 標題` 切出各段落。圍欄內的行原樣保留在段落內容裡
|
||||
* ——流程圖那一段要的就是整個 mermaid 區塊。
|
||||
* @param {string} body 議題 body
|
||||
* @returns {Map<string, string>} 段落名稱 → 該段內容(前後空白已修掉)
|
||||
*/
|
||||
@@ -15,16 +46,13 @@ export function parseSections(body) {
|
||||
const sections = new Map();
|
||||
const buffer = [];
|
||||
let current = null;
|
||||
let inFence = false;
|
||||
|
||||
const flush = () => {
|
||||
if (current !== null) sections.set(current, buffer.join('\n').trim());
|
||||
buffer.length = 0;
|
||||
};
|
||||
|
||||
for (const line of (body ?? '').split('\n')) {
|
||||
if (/^\s*```/.test(line)) inFence = !inFence;
|
||||
|
||||
for (const { line, inFence } of eachLine(body)) {
|
||||
const heading = inFence ? null : line.match(/^##\s+(.+?)\s*$/);
|
||||
if (heading) {
|
||||
flush();
|
||||
@@ -50,22 +78,25 @@ export function textSection(sections, name) {
|
||||
}
|
||||
|
||||
/**
|
||||
* 取出列表型段落的每一項。符號清單與編號清單一視同仁,
|
||||
* checkbox 只留文字不留標記 —— 需求議題的驗收標準不追蹤勾選狀態。
|
||||
* 取出列表型段落的每一項。符號清單與編號清單一視同仁,checkbox 只留文字不留標記
|
||||
* ——需求議題的驗收標準不追蹤勾選狀態。巢狀項目一律攤平:需求議題的模板沒有巢狀,
|
||||
* 真的出現時寧可多帶一項,也不要無聲吃掉內容。
|
||||
* @param {Map<string, string>} sections
|
||||
* @param {string} name
|
||||
* @returns {string[]}
|
||||
*/
|
||||
export function listSection(sections, name) {
|
||||
const content = sections.get(name);
|
||||
if (!content) return [];
|
||||
const items = [];
|
||||
|
||||
return content
|
||||
.split('\n')
|
||||
.map((line) => line.match(/^\s*(?:[-*+]|\d+\.)\s+(.*)$/))
|
||||
.filter((match) => match !== null)
|
||||
.map((match) => match[1].replace(/^\[[ xX]\]\s*/, '').trim())
|
||||
.filter((text) => text !== '');
|
||||
for (const { line, inFence } of eachLine(sections.get(name))) {
|
||||
if (inFence) continue;
|
||||
const item = line.match(/^\s*(?:[-*+]|\d+\.)\s+(.*)$/);
|
||||
if (!item) continue;
|
||||
|
||||
const text = item[1].replace(/^\[[ xX]\]\s*/, '').trim();
|
||||
if (text !== '') items.push(text);
|
||||
}
|
||||
return items;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -76,19 +107,33 @@ export function listSection(sections, name) {
|
||||
* @returns {{term: string, def: string}[]}
|
||||
*/
|
||||
export function tableSection(sections, name) {
|
||||
const content = sections.get(name);
|
||||
if (!content) return [];
|
||||
const rows = [];
|
||||
|
||||
const rows = content
|
||||
.split('\n')
|
||||
.filter((line) => line.trim().startsWith('|'))
|
||||
.map((line) => line.trim().replace(/^\||\|$/g, '').split('|').map((cell) => cell.trim()));
|
||||
for (const { line, inFence } of eachLine(sections.get(name))) {
|
||||
if (inFence || !line.trim().startsWith('|')) continue;
|
||||
rows.push(splitRow(line));
|
||||
}
|
||||
|
||||
const separator = rows.findIndex((cells) => cells.every((cell) => /^:?-+:?$/.test(cell)));
|
||||
const separator = rows.findIndex(isSeparator);
|
||||
if (separator === -1) return [];
|
||||
|
||||
return rows
|
||||
.slice(separator + 1)
|
||||
// 同一段落裡若不慎貼了第二張表,它的分隔列不該變成一筆 {term:'---'}
|
||||
.filter((cells) => !isSeparator(cells))
|
||||
.filter((cells) => cells.length >= 2 && cells.some((cell) => cell !== ''))
|
||||
.map(([term, def]) => ({ term, def }));
|
||||
}
|
||||
|
||||
/** 以未被逸脫的直線切欄,再把 `\|` 還原成內容裡的直線 */
|
||||
function splitRow(line) {
|
||||
return line
|
||||
.trim()
|
||||
.replace(/^\||\|$/g, '')
|
||||
.split(/(?<!\\)\|/)
|
||||
.map((cell) => cell.replace(/\\\|/g, '|').trim());
|
||||
}
|
||||
|
||||
function isSeparator(cells) {
|
||||
return cells.length > 0 && cells.every((cell) => /^:?-+:?$/.test(cell));
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user