Files
superpaper/DESIGN.md
T

1825 lines
93 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Superpaper Family:论文笔记编排层与推导图工具包
| 字段 | 值 |
|---|---|
| **Title** | Superpaper Family:论文脉络 / 推导 / 出图分层设计 |
| **Author** | TBD |
| **Date** | 2026-08-17 |
| **Status** | Final |
| **Audience** | 将在 `/home/carry/myprj/tools/skills/` 落地 `superpaper/`(及后续 `superderive/`)的工程师与 agent |
| **Companion** | `superfig/`、`supertensor/`、`wdkns-skills/skills/youtube-render-pdf/` |
---
## Overview
用户要的三件事——**文章脉络梳理**、**公式推导**、**抽象公式图像化**——不是「再画一种图」,而是一层**读论文 / 写讲义**的编排问题,外加一类尚未存在的**逐步代数图**。现有 `superfig` 的契约是「一条 paper claim → 一张 dense、slide-ready 的节点/边图」(见 `superfig/SKILL.md`);`supertensor` 的契约是「一条张量计算路径 → 一张形状对齐的 face 图」(见 `supertensor/SKILL.md`)。把笔记编排、符号表、引理依赖、逐步改写塞进 `superfig.sty`,会同时破坏两边已经用 lint-on-warning 钉死的图尺度不变量(一图一个 `\sfcallout`、`\sftopformula` 必须在最后 `\sfrowend` 之后、每图 ≤4 个活跃色相)。
本设计按 **superfig / supertensor 已经验证过的 sibling 分层** 落一套 family,而不是升级某一个 `.sty`:
```
superpaper NEW(v1):读论文 / 笔记编排 skill(文档层)
├─ 脉络: claims, definitions, assumptions, lemma dependencies, symbol table
├─ 推导链: 哪些代数步骤必须在笔记里展开
└─ 委派出图
├─ superfig EXISTING:architecture / pipeline / state / dependency
├─ supertensor EXISTING:shape-aligned tensor formula figures
└─ superderive NEW(phase 2):stepwise algebraic derivation figures
```
`superpaper` 是论文侧的 `youtube-render-pdf`:输入一篇论文,输出结构化中文 `.tex` 笔记 + 编译 PDF。它**不**生长 `superfig.sty`。所谓「抽象公式图像化」在 v1 是一个 **figure router**(决策表 + 委派契约),不是第三套绘图引擎。逐步改写图在 phase 2 由 `superderive` 承接;v1 笔记用 `align` / `aligned` 顶上。
---
## Background & Motivation
### 当前状态
仓库里已经有两套 **figure toolkit** sibling,形状几乎同构:
| | `superfig/` | `supertensor/` |
|---|---|---|
| 入口 | `SKILL.md`(瘦;细节在 `references/`) | 同 |
| 宏包 | `assets/superfig.sty`,`\sf*` | `assets/supertensor.sty`,`\st*` |
| 原语 | `\sfnode` `\sfconn` `\sfarrow` `\sfgroup` `\sfcallout` `\sfmeaningbox` | `\stface` `\stdim` `\ststack` `\stindexface` `\stlink` `\stcomm` |
| 工作流 | preflight → 收成一条 claim → **semantics ledger** → 选最小语法 → `build.sh` | 同,但 ledger 是 **shape & semantics + geometry** |
| 构建纪律 | `Missing character` / overfull / `Package superfig Warning` = 失败 | 同,包名换成 `supertensor` |
| 交付物 | PNG 预览 + 一段 mechanism 说明 + `.tex` + PDF/SVG | 同 |
| 明确不做 | 张量形状、数值绘图 | 无形状的架构框图、数值绘图 |
House style 是**复制**而非抽取的:muted palette `#4F8FA5 / #EE995B / #C95B5B / #8A74B5 / #85898F`,一 role 一色,≤4 活跃色相 + gray。`supertensor/README.md` Provenance 写明它抽自 `wdkns-skills/skills/tensor-formula-viz/`,后者保持不动。
文档/编排层的现成类比是 `wdkns-skills/skills/youtube-render-pdf/`(及 `bilibili-render-pdf/`):
- 教学序列:动机 → 想法 → 机制 → 证据 → takeaway
- 公式规则:先中文讲清楚 → display math → 扁平符号表
- 盒子:`importantbox` / `knowledgebox` / `warningbox`(定义在 `assets/notes-template.tex`)
- 可视化委派:TikZ 或 matplotlib,**skill 自己不拥有图原语**
- 长视频切段 + outline / writers / figure / consistency 多 agent(见 `wdkns-skills/README.md`「subagents 的触发」)
- 召回不足的缓解:独立 reviewer 对照源材料查漏
`wdkns-skills/templates/writing/readme.md` 把写作当成 structure(骨架)+ contents(血肉),并指向 paper-slides 前作。这和「先 ledger 再写笔记」是同一条纪律。
### 痛点
1. **图层错位。** 脉络是跨章节的 claim / 假设 / 符号生命周期问题;`superfig` 的 lint(`scripts/lint.py`:一图一个 callout、formula 不得出现在最后一行之前、禁止 raw `\draw`)是为 **standalone 一页图** 写的。把 15 页笔记塞进这套规则,只可能有两种结果:要么放松 lint(明确非目标),要么笔记根本编不过。
2. **原语错位。** 「逐步消去 / 代入」不是 node/edge,也不是 tensor face。硬用 `\sfnode` 表示一行等式,会画出一张**编译干净、看起来整齐、教错东西**的图——这正是 `superfig/references/antipatterns.md` 和 `supertensor` 存在的理由。
3. **出图引擎错位。** 「把抽象公式变成图」在本仓库里已经分裂成两种合法引擎。缺的是**路由**,不是第三种「什么都能画」的宏。
4. **没有论文侧的 ledger。** `superfig` 要求画之前先有 nodes/edges/groups ledger;`supertensor` 要求 shape & geometry ledger。论文笔记今天没有对应物,于是符号漂移、漏 claim、图和正文各说各话——这是 `youtube-render-pdf` 用 consistency agent + 事后 reviewer 在视频域打过的同一类仗。
### 为何现在做
产品决策已定:**不要塞进 superfig;按层拆,对齐现有 sibling 模式。** 仓库里 `superfig/` 与 `supertensor/` 的目录形状、`SKILL.md` 瘦身方式、`scripts/{preflight.sh,lint.py,build.sh,test.sh}` 四件套已经足够当脚手架模板。`youtube-render-pdf` 提供文档层教学契约。缺的是一份足够具体、第一张 PR 不用再发明产品决策的设计。
---
## Goals & Non-Goals
### Goals
- **v1 落地 `superpaper/`**:agent skill + 笔记模板 + ledger schema + 可执行 figure router + 笔记构建/一致性 lint。输入一篇论文(PDF / arXiv / 本地 tex / 摘录),输出结构化中文笔记 `.tex` + PDF。
- **Ledger-first**:先写 claims / definitions / assumptions / symbols / derivation chains / figure plan,再写正文、再委派出图。Ledger 是 machine-checkable 的单一事实源。
- **路由而非绘制**:按决策表把每张计划中的图交给 `superfig`、`supertensor`、(v2)`superderive`、screenshot / matplotlib,或 `none`(纯公式)。
- **委派契约可执行**:figure agent 收到的 request 文件能直接当 sibling skill 的输入;交付物与 `superfig/SKILL.md`「Output」一节同构。
- **零改动 sibling 契约(v1)**:不改 `superfig.sty` / `supertensor.sty` / 它们的 lint 规则 / `wdkns-skills/`。
- **v2 规格写清**:`superderive` 的视觉语法、宏草图、与 `align` 的分界,足以单独开 PR,不必再做产品讨论。
### Non-Goals
- 不把 `superfig` 变成笔记器、证明助手或厨房水槽宏包。
- 不合并 `superfig` 与 `supertensor`。
- v1 **不**抽取共享 `superstyle` 包。现有两套已经复制 palette;现在抽取比再复制一次更贵。
- 不实现数值绘图引擎(loss / bar / scatter 出家族,见路由表)。
- 不原地改 `wdkns-skills`;新工作只出现在 `/home/carry/myprj/tools/skills/` 的 sibling 目录。
- 不把 `superfig` 的图尺度 lint(一 callout、formula 顺序、4 色相)「升级」成文档尺度规则。
- 不做多篇论文对比、文献综述自动生成、论文 PDF 重排版、或论文原文的 TeX 再发布。
- 不做交互式证明助手(Lean/Isabelle);推导链是教学展开,不是形式化核验。
- 不把 arXiv 源码拿来 **编译**(可读、可抄符号,不执行)。
---
## Proposed Design
### 家族分层与数据流
```mermaid
flowchart TB
subgraph sources [Input]
PDF[Local PDF]
ARX[arXiv id]
TEX[Local tex]
EX[Excerpt / markdown]
end
subgraph sp [superpaper — document layer]
ING[ingest.sh]
LED[ledger.yaml]
OUTL[outline.md]
NOTES[notes.tex]
REQ[figures/Fid.request.md]
LINT[lint.py]
NBUILD[scripts/build.sh]
end
subgraph fig [figure siblings — standalone layer]
SF[superfig]
ST[supertensor]
SD[superderive v2]
SC[screenshot / matplotlib]
end
PDF --> ING
ARX --> ING
TEX --> ING
EX --> ING
ING --> LED
LED --> OUTL
LED --> NOTES
LED --> REQ
REQ --> SF
REQ --> ST
REQ --> SD
REQ --> SC
SF -->|PDF vector| NOTES
ST -->|PDF vector| NOTES
SD -->|PDF vector| NOTES
SC -->|PDF/PNG + 出处脚注| NOTES
NOTES --> LINT
LED --> LINT
LINT --> NBUILD
NBUILD --> PDFN[notes.pdf]
```
两条硬边界:
1. **笔记文档不 `\usepackage{superfig}` / `supertensor`。** 它们的 `\documentclass{standalone}` + 图尺度 lint 与 `article` 笔记互斥。委派产物以 **PDF 向量** `\includegraphics` 嵌入(与 `youtube-render-pdf` 对外生图「export pdf, include」的约定一致,见该 SKILL「Visualization」)。
2. **图尺度不变量仍由 sibling `build.sh` 执行。** `superpaper` 的 `lint.py` 管 ledger 完整性、claim 覆盖、符号表、toolkit 枚举、禁止把 `.sty` 拉进笔记;不管 callout 预算和 4 色相。
---
### Layer 1 — `superpaper`(先做)
#### 1. 输入契约
**v1 范围内的源(按信息质量降序):**
| `source.kind` | 用户怎么给 | ingest 做什么 | 质量含义 |
|---|---|---|---|
| `tex` | 本地 `.tex` / 工程目录 | 复制到 `source/tex/`;不编译 | 符号与公式最佳 |
| `arxiv` | `1706.03762` 或 `arxiv.org/abs/...` | 校验 id;拉 PDF;**尽力**拉 e-print(tar),只解包阅读 | PDF 保底 + 可能有源码 |
| `pdf` | 本地 PDF 路径 | 复制到 `source/paper.pdf` | 正文靠 `pdftotext` + 页渲染 |
| `excerpt` | 粘贴 / `.md` / 选段 | 写入 `source/excerpt.md` | `coverage.mode = excerpt`,不假装读完全文 |
| `markdown` | 已有笔记 / HTML dump | 复制为 `source/excerpt.md` | 同 excerpt |
arXiv id 校验(防 shell 注入):
```text
^(ar[Xx]iv:)?(\d{4}\.\d{4,5}(v\d+)?|[a-z-]+/\d{7})$
```
下载:
- PDF:`https://arxiv.org/pdf/<id>.pdf`(失败则 `export.arxiv.org`)
- e-print(可选、best-effort):`https://arxiv.org/e-print/<id>` → `source/eprint/`。**禁止**对解出的 TeX 跑 `latex`。只当符号/公式的高信息文本。
**明确不在 v1:** 付费出版社 HTML 抓取、扫描件 OCR 当主路径、视频论文(走 `youtube-render-pdf`)、ar5iv HTML。扫描件若 `pdftotext` 几乎为空:标 `degraded: scanned`,进入页渲染视觉模式(对标 `bilibili-render-pdf` 的 visual-only),并在交付里说出口。
**论文里已有图/表怎么处理(相对重绘):**
ingest 后 outline agent 必须给**每一张源图 / 主表**打标签,写入 `ledger.yaml` 的 `source_assets[]` 与/或 `figures[]`:
| `handling` | 何时 | 笔记里长什么样 |
|---|---|---|
| `redraw-superfig` | 架构、流水线、数据流、状态、依赖、「谁吃谁」 | 委派 `superfig`;caption 写「重绘自论文 Figure N」 |
| `redraw-supertensor` | 轴、形状、转置、broadcast、gather、shard、收缩 | 委派 `supertensor` |
| `redraw-superderive` | 必须被**看见**的逐步改写(v2;v1 降级为 `align`) | v1:`align`;v2:委派 `superderive` |
| `screenshot` | 数值曲线、照片、复杂实验装置、UI | `pdftoppm -r 200` 该页(+ 可选 `crop_bbox`);出处脚注 |
| `matplotlib` | 用户明确要求重画数值图 | 族外脚本;导出 PDF 再嵌入 |
| `omit` | 装饰、重复、无教学价值 | ledger 记原因,不进笔记 |
规则:
- **不要**把论文每一张图都截进笔记。按教学需要选,对标 `youtube-render-pdf`「Select figures by necessity」。
- 架构类图**优先重绘**,不要截原图——原图常常违反 house style,且无法进入 sibling 的语义审计。
- 数值图**不要**用 `superfig` 假装;路由到 `screenshot` 或 `matplotlib`,并在笔记里说「本家族不出数值图」。
- 出处脚注是视频时间戳的论文对应物:`论文 Figure 2,§4.1,p.7`。重绘图同样要写源 Figure 编号。
- 理解源图必须**看页渲染**(`pdftoppm`),不要只靠 caption 文本猜内容。对标 youtube skill 的 `view image` 纪律,禁止用 tesseract 代替看图。
**页预算与切段(对标长视频策略;与 §7 单/多 agent 表同一套谓词):**
| 条件 | 策略 |
|---|---|
| **必须拆** writer/figure agent:正文 **> 12 页**,或论文顶层节 **> 4**,或计划重绘图 **≥ 2**,或用户显式 spawn | 按 `outline.md` 的**讲义节**并行(不是按 PDF 页序) |
| **否则单 agent**(含 excerpt、以及 9–12 页 ∧ 顶层节 ≤ 4 ∧ 重绘图 ≤ 1) | 一遍写完 |
| 正文 > 40 页(综述 / 教材) | 默认 **只做用户点名的节 + 引言/结论**;否则先问范围。禁止单次塞完整 40 页 |
附录默认跳过,写入 `coverage.sections_skipped`。用户点名再纳入。
**工作目录(唯一根;所有脚本吃 `--work <work>`):**
```text
<work>/ # 默认 ./work/<paper-id>/
source/
meta.yaml # kind, arxiv, pages, sha256
paper.pdf
paper.txt # pdftotext -layout
pages/pg-001.png # 一律三位:ingest 把 pdftoppm 输出重命名为 pg-%03d.png
eprint/ # 可选
tex/ # 可选
excerpt.md # excerpt / markdown
ledger.yaml # SSOT(始终在 <work>/ 根,不在 notes/)
outline.md
notes/
notes.tex # 主文件;cwd 编译;\input{sections/sec-01.tex}
sections/sec-01.tex
sections/symbols.tex # render_ledger.py --work 写入此处
figures/
F1/
F1.request.md
F1.tex
F1-mechanism.md
build/F1.{pdf,svg,png} # sibling 向量产物
orig.png # 仅 screenshot
plot.py # 仅 matplotlib
run.log
out/notes.pdf
```
**路径约定(三处必须同形,相对 `notes/`):**
| 位置 | 向量图 | 截图 |
|---|---|---|
| `figures[].include` | `figures/F1/build/F1.pdf` | `figures/F2/orig.png` |
| `\spfig` / `\spscreenshot` | `figures/#2/build/#2.pdf` | `figures/#2/orig.png` |
| 磁盘 | `<work>/notes/figures/F1/build/F1.pdf` | `<work>/notes/figures/F2/orig.png` |
| `figures[].request` | `figures/F1/F1.request.md` | `null`(截图不写 request) |
禁止再写 `notes/figures/...` 进 ledger(那是相对 `<work>/` 的另一套根)。API 层只承诺「相对 `notes/` 的 `figures/F<id>/...`」。
`paper-id`:arXiv id **原样保留点号**(`1706.03762` → 目录 `1706.03762`)。非 arXiv:NFKC → 小写 → 非 `[a-z0-9.]` 换成 `-` → 压缩连字符 → 截断 40 字符 → 必须以字母或数字开头结尾;空则用 `paper`。仓库内 `examples/` 只放 **自造 fixture**,不提交受版权保护的 PDF。
存储量级(单篇 12 页 ML 论文):PDF 1–15 MB + 页渲染约 8–20 MB + 笔记/图 2–8 MB ≈ **20–50 MB**。运行时 `<work>/` 默认在调用方 cwd 下的 `work/`,gitignore,不进 skill 仓库。
---
#### 2. Ledger-first 工作流
对标 `superfig/SKILL.md` 步骤 3「Build a semantics ledger before drawing」和 `supertensor/SKILL.md` 步骤 3 的双 ledger。`superpaper` 的 ledger 在**文档层**:
1. `preflight.sh`(笔记工具链,不是图工具链)。
2. `ingest.sh` 落 `source/`。
3. **先写 `ledger.yaml`**,再写 `outline.md`,再写任何 `sections/*.tex`。
4. figure plan 填完才能开 figure agent。
5. 改图或改术语:先改 ledger,再改那一处正文或那一张 request——对标 sibling 的「Iterating:不要重画整张图」。
**格式选 YAML,不选 Markdown。** 理由:claim 覆盖、toolkit 枚举、`retired_ids` 交叠检查必须机器可解析;figure plan 是带 enum 的分发表。Markdown 标题约定会逼 lint 写第二套解析器。数学用字面块 `|` 或单引号标量。`assets/ledger.example.yaml` 既是 schema 样例,也是 outline agent 的**空模板**(复制到 `<work>/ledger.yaml` 再填)。不要用 `agents/openai.yaml` 当「agent 会写 YAML」的证据——那只是 4 行 `interface:` 块。
Markdown 只作为**投影**:`render_ledger.py --work <work>` 写入 `notes/sections/symbols.tex`(附录 `\input`)和可选的 `<work>/claims.md`。投影只读。
**id 规则:**
- 前缀 + **正整数**:`C1`、`Q1`、`D1`、`A1`、`L1`、`E1`、`DER1`、`F1`、`SA1`。正则:`^(C|Q|D|A|L|E|DER|F|SA)[1-9][0-9]*$`。
- **禁止** `F1a` / `F1b`。跨类拆分时删掉(或 retire)混信号行,用**下一个整数**建新行(见 §4)。
- 同一 ledger 内 id 唯一(含已 `status: dropped` 的行)。引用只用 id。
- 删除或永久废弃一个 id 时:把它写入顶层 `retired_ids`,并从对应数组移除**或**保留为 `status: dropped` 且不再改语义。新对象只许拿下一个未用整数。
- 笔记用 `\splabel{C1}` 把段落钉到 id 上(PR 4 起强制 core claim)。
---
#### 3. Ledger schema(`superpaper.ledger/v1`)
权威副本:`superpaper/assets/ledger.schema.yaml`(JSON Schema draft-07)。`lint.py` 顺序固定,禁止「先 validate 再认别名」:
1. **Normalize(写回内存副本,不改磁盘):** 把每个 `symbols[].kind` 按下面别名表收到 canonical。`kind: activation` → `value`,`kind: shape-parameter` / `shape_parameter` → `shape parameter`。不在别名表且不在 canonical 的值原样留下,下一步会 `SP001`。
2. **硬失败 `SP001`:** 对**已经 normalize** 的对象跑 `jsonschema.validate`(依赖 **`jsonschema`**)。schema 的 `kind` / `steps[].rule` enum **只含 canonical**。未知 `rule`(如 `foo`)是 `SP001`,不是警告。`additionalProperties: true`,未知键不因此失败。
3. **警告(不失败):** validate 通过后,扫顶层与一层嵌套的未知键、`coverage.mode=full` 且 `sections_in` 空、符号从未在笔记出现。**不要**再对 `steps[].rule` 发警告——那条路径已被 schema enum 关掉。
excerpt-toy 必须带齐 **required 数组**(可空):`claims: []` 合法。缺 `paper.id` 不合法。
**Required:**
| 对象 | required |
|---|---|
| 根 | `schema`, `paper`, `coverage` |
| `paper` | `id`, `title`, `source` |
| `paper.source` | `kind` |
| `coverage` | `mode` |
| `questions[]` | `id`, `text` |
| `claims[]` | `id`, `text`, `kind`, `status` |
| `definitions[]` | `id`, `name`, `text` |
| `assumptions[]` | `id`, `text` |
| `lemmas[]` | `id`, `text` |
| `symbols[]` | `name`, `latex`, `meaning`, `kind` |
| `derivations[]` | `id`, `claim`, `title`, `expand` |
| `derivations[].steps[]` | `id`, `from`, `to`, `rule` |
| `figures[]` | `id`, `claim`, `title`, `grammar`, `toolkit`, `signals`, `status` |
| `evidence[]` | `id`, `kind`, `source`, `supports`, `handling` |
| `terms[]` | `canonical` |
| `source_assets[]` | `id`, `kind`, `handling` |
| `skip_reasons[]` | `section`, `reason` |
未列的数组键缺省视为 `[]`;schema 仍把它们标为 optional。`retired_ids` 缺省 `[]`。
**类型与可空:**
- `paper.year`:integer 或省略。
- `paper.source.pages`:integer ≥ 1 或省略。
- `paper.degraded`:string 数组,enum `scanned | no-eprint | no-pdftotext`。
- `figures[].source_pages`:integer 数组(页码,从 1)。
- `figures[].include` / `request` / `drop_reason` / `derivations[].figure`:`string` **或** `null`。路径/空值与 `status`/`toolkit` 的组合**不**写进 draft-07 `if`/`then`(避免和 `additionalProperties: true` 扭在一起);由 `check_row` 的 `SP013–SP016` 执行,见 §4 / §10。
- `figures[].crop_bbox`:`null` 或恰好 4 个 number `[x0,y0,x1,y1]`(200 dpi 页像素,左上原点)。仅 screenshot。
- `figures[].signals`:非空 string 数组(`toolkit: none` 的 bookkeeping 行允许 `["notation"]` 或空数组)。
**Enums(闭集;写进 schema):**
| 字段 | enum |
|---|---|
| `schema` | `superpaper.ledger/v1` |
| `source.kind` | `arxiv \| pdf \| tex \| excerpt \| markdown` |
| `coverage.mode` | `full \| excerpt \| body-only` |
| `claims[].kind` | `contribution \| theoretical \| empirical \| methodological` |
| `claims[].status` | `core \| supporting \| dropped` |
| `figures[].toolkit` | `superfig \| supertensor \| superderive \| screenshot \| matplotlib \| none \| align` |
| `figures[].status` | `planned \| delegated \| built \| included \| dropped` |
| `figures[].grammar` | `architecture \| pipeline \| data-flow \| state \| time \| dependency \| argument-map \| tensor-face \| derivation \| screenshot \| plot \| notation \| none` |
| `evidence[].kind` | `table \| plot \| ablation \| theorem \| example` |
| `evidence[].handling` / `source_assets[].handling` | `redraw-superfig \| redraw-supertensor \| redraw-superderive \| screenshot \| matplotlib \| omit` |
| `source_assets[].kind` | `figure \| table` |
| `steps[].rule` | `definition \| substitute \| cancel \| factor \| scale \| take-limit \| approx \| cite \| rearrange \| introduce` |
| `symbols[].kind` | 见下表 |
**`symbols[].kind`:** 权威左列抄 `supertensor/references/semantics.md`「Kinds worth separating」。canonical 取斜杠左侧(`shape parameter` 保留空格,与原文一致)。schema enum **只列 canonical**。别名合法当且仅当走第 1 步 normalize;`SP001` 在 normalize **之后**跑。这是有文档的扩展,不是「零新增类型学」——`scalar` / `set` 只用于文档层、不进张量图。
| canonical(schema enum) | lint 接受的别名 | 例 |
|---|---|---|
| `value` | `activation`, `value / activation` | `X`, `H` |
| `score` | `logit`, `score / logit` | `S = QK^T/√d_h` |
| `probability` | | `A = softmax(S)` |
| `index` | `coordinate`, `index / coordinate` | `I` |
| `rank` | `order`, `rank / order` | 选中槽轴 `r` |
| `count` | | `n_e` |
| `id` | | token / device id |
| `mask` | `support`, `mask / support` | 因果 `M` |
| `permutation` | | gather 序 |
| `shape parameter` | `shape-parameter`, `shape_parameter` | `d_k`, `h`, `p`(**不是** `scalar`) |
| `scalar`(文档层扩展) | | 损失 `L`、温度 `τ` 这类 0 维量,且不是轴长 |
| `set`(文档层扩展) | | 词表 `V` |
draft-07 正文(检入 `assets/ledger.schema.yaml`;实现不得再猜 required/enum):
```yaml
$schema: "http://json-schema.org/draft-07/schema#"
$id: "https://local/superpaper.ledger/v1"
type: object
required: [schema, paper, coverage]
additionalProperties: true
properties:
schema: {const: superpaper.ledger/v1}
retired_ids:
type: array
items: {type: string, pattern: "^(C|Q|D|A|L|E|DER|F|SA)[1-9][0-9]*$"}
paper:
type: object
required: [id, title, source]
additionalProperties: true
properties:
id: {type: string, minLength: 1}
title: {type: string}
authors: {type: array, items: {type: string}}
year: {type: integer}
venue: {type: string}
notes_language: {type: string, enum: [zh, en]}
degraded:
type: array
items: {type: string, enum: [scanned, no-eprint, no-pdftotext]}
source:
type: object
required: [kind]
additionalProperties: true
properties:
kind: {enum: [arxiv, pdf, tex, excerpt, markdown]}
arxiv: {type: string}
local_pdf: {type: string}
pages: {type: integer, minimum: 1}
language: {type: string}
coverage:
type: object
required: [mode]
additionalProperties: true
properties:
mode: {enum: [full, excerpt, body-only]}
sections_in: {type: array, items: {type: string}}
sections_skipped: {type: array, items: {type: string}}
skip_reasons:
type: array
items:
type: object
required: [section, reason]
properties:
section: {type: string}
reason: {type: string}
questions:
type: array
items:
type: object
required: [id, text]
properties:
id: {type: string, pattern: "^Q[1-9][0-9]*$"}
text: {type: string}
source: {type: string}
claims:
type: array
items:
type: object
required: [id, text, kind, status]
properties:
id: {type: string, pattern: "^C[1-9][0-9]*$"}
text: {type: string}
kind: {enum: [contribution, theoretical, empirical, methodological]}
status: {enum: [core, supporting, dropped]}
supports: {type: array, items: {type: string}}
depends_on: {type: array, items: {type: string}}
evidence: {type: array, items: {type: string}}
source: {type: string}
definitions:
type: array
items:
type: object
required: [id, name, text]
properties:
id: {type: string, pattern: "^D[1-9][0-9]*$"}
name: {type: string}
text: {type: string}
source: {type: string}
assumptions:
type: array
items:
type: object
required: [id, text]
properties:
id: {type: string, pattern: "^A[1-9][0-9]*$"}
text: {type: string}
source: {type: string}
used_by: {type: array, items: {type: string}}
lemmas:
type: array
items:
type: object
required: [id, text]
properties:
id: {type: string, pattern: "^L[1-9][0-9]*$"}
text: {type: string}
depends_on: {type: array, items: {type: string}}
used_by: {type: array, items: {type: string}}
source: {type: string}
symbols:
type: array
items:
type: object
required: [name, latex, meaning, kind]
properties:
name: {type: string}
latex: {type: string}
meaning: {type: string}
domain: {type: string}
kind:
enum: [value, score, probability, index, rank, count, id, mask,
permutation, "shape parameter", scalar, set]
introduced: {type: string}
used: {type: array, items: {type: string}}
aliases: {type: array, items: {type: string}}
derivations:
type: array
items:
type: object
required: [id, claim, title, expand]
properties:
id: {type: string, pattern: "^DER[1-9][0-9]*$"}
claim: {type: string}
title: {type: string}
source: {type: string}
expand: {type: boolean}
figure: {type: ["string", "null"]}
steps:
type: array
items:
type: object
required: [id, from, to, rule]
properties:
id: {type: string}
from: {type: string}
to: {type: string}
rule:
enum: [definition, substitute, cancel, factor, scale,
take-limit, approx, cite, rearrange, introduce]
cite: {type: string}
justify: {type: string}
figures:
type: array
items:
type: object
required: [id, claim, title, grammar, toolkit, signals, status]
properties:
id: {type: string, pattern: "^F[1-9][0-9]*$"}
claim: {type: string}
title: {type: string}
grammar:
enum: [architecture, pipeline, data-flow, state, time, dependency,
argument-map, tensor-face, derivation, screenshot, plot,
notation, none]
toolkit:
enum: [superfig, supertensor, superderive, screenshot,
matplotlib, none, align]
signals: {type: array, items: {type: string}}
source_fig: {type: string}
source_pages: {type: array, items: {type: integer, minimum: 1}}
crop_bbox:
type: ["array", "null"]
minItems: 4
maxItems: 4
items: {type: number}
request: {type: ["string", "null"]}
include: {type: ["string", "null"]}
status: {enum: [planned, delegated, built, included, dropped]}
drop_reason: {type: ["string", "null"]}
evidence:
type: array
items:
type: object
required: [id, kind, source, supports, handling]
properties:
id: {type: string, pattern: "^E[1-9][0-9]*$"}
kind: {enum: [table, plot, ablation, theorem, example]}
source: {type: string}
supports: {type: array, items: {type: string}}
handling:
enum: [redraw-superfig, redraw-supertensor, redraw-superderive,
screenshot, matplotlib, omit]
terms:
type: array
items:
type: object
required: [canonical]
properties:
canonical: {type: string}
aliases: {type: array, items: {type: string}}
first_defined: {type: string}
source_assets:
type: array
items:
type: object
required: [id, kind, handling]
properties:
id: {type: string, pattern: "^SA[1-9][0-9]*$"}
kind: {enum: [figure, table]}
paper_ref: {type: string}
handling:
enum: [redraw-superfig, redraw-supertensor, redraw-superderive,
screenshot, matplotlib, omit]
pages: {type: array, items: {type: integer}}
note: {type: string}
```
**样例(字段合法;`include`/`request` 已改为相对 `notes/`):**
```yaml
schema: superpaper.ledger/v1
retired_ids: []
paper:
id: "1706.03762"
title: "Attention Is All You Need"
authors: ["Vaswani et al."]
year: 2017
venue: "NeurIPS"
source: {kind: arxiv, arxiv: "1706.03762", local_pdf: "source/paper.pdf", pages: 15, language: en}
notes_language: zh
degraded: []
coverage:
mode: full
sections_in: ["1", "2", "3", "4", "5"]
sections_skipped: ["6", "appendix"]
skip_reasons:
- {section: "6", reason: "default skip appendix-like extras"}
questions:
- {id: Q1, text: "能否去掉 recurrence / convolution,只用注意力做序列转导?", source: "§1"}
claims:
- id: C1
text: "多头自注意力可以替代循环与卷积作为序列转导的主模块"
kind: contribution
status: core
supports: [Q1]
depends_on: [A1, D1]
evidence: [E1]
source: "§1, §3.2.2, §5"
definitions:
- {id: D1, name: "scaled dot-product attention",
text: "Attention(Q,K,V) = softmax(QK^T / sqrt(d_k)) V", source: "§3.2.1, Eq.(1)"}
assumptions:
- {id: A1, text: "位置信息由显式 positional encoding 注入", source: "§3.5", used_by: [C1]}
lemmas: []
symbols:
- {name: Q, latex: "Q", meaning: "query 矩阵", domain: "R^{n x d_k}",
kind: value, introduced: "§3.2.1, Eq.(1)", used: ["§3.2.1", "§3.2.2"]}
- {name: d_k, latex: "d_k", meaning: "key 维", kind: "shape parameter"}
derivations:
- id: DER1
claim: C1
title: "scaled dot-product 的缩放从何而来"
source: "§3.2.1"
expand: true
figure: null # v1:笔记 align;见 §4
steps:
- {id: S1, from: "QK^T", to: "QK^T / sqrt(d_k)", rule: scale,
cite: "§3.2.1", justify: "点积方差随 d_k 增长"}
figures:
- id: F1
claim: C1
title: "Encoder–decoder 块"
grammar: architecture
toolkit: superfig
signals: [architecture, what-eats-what]
source_fig: "Figure 1"
source_pages: [3]
crop_bbox: null
request: "figures/F1/F1.request.md"
include: "figures/F1/build/F1.pdf"
status: planned
drop_reason: null
evidence:
- {id: E1, kind: table, source: "Table 2", supports: [C1], handling: screenshot}
terms:
- {canonical: "multi-head attention", aliases: ["MHA", "多头注意力"], first_defined: "§3.2.2"}
source_assets:
- {id: SA1, kind: table, paper_ref: "Table 2", handling: screenshot, pages: [6],
note: "BLEU 表;数值结果走 screenshot,signals 用 table 不是 loss-curve"}
```
**PR 1 起就必须存在的 schema fixture 与期望码**(只测本 PR 已实现的检查;完整码表见 §10):
| 文件 | 期望 |
|---|---|
| `tests/ledger-valid.yaml` | 退出 0 |
| `tests/ledger-invalid/missing-paper-id.yaml` | `SP001` |
| `tests/ledger-invalid/bad-toolkit.yaml` | `SP001`(enum) |
| `tests/ledger-invalid/duplicate-id.yaml` | `SP002` |
| `tests/notes-loads-sty.tex` + 最小 ledger | `SP003` |
---
#### 4. Figure router(「抽象公式图像化」= 路由)
权威规则在 `references/router.md`;可执行副本在 `scripts/router.py`(**PR 1 即完整实现**,含单测;lint 接线分 PR,见 §10 / PR Plan)。
`suggest()` 返回的是 **class**,不是 toolkit。lint(PR 2+)检查 `figures[].toolkit ∈ ACCEPTED[class]`。
**信号 → class:**
| class | 若 `signals` 与下列集合相交 |
|---|---|
| `numeric` | `{loss-curve, bar, scatter, histogram, numeric-plot}` |
| `raster` | `{table, photo, apparatus, ui}` |
| `tensor` | `{axis, shape, transpose, broadcast, gather, shard, contraction, face}` |
| `derive` | `{rewrite-figure, cancel-visual, subst-visual}` |
| `fig` | `{architecture, pipeline, data-flow, state, time, dependency, what-eats-what, argument-map}` |
| `none` | 不相交(含仅 `notation` 或 `[]`) |
**class → 合法 toolkit 集合:**
| class | `ACCEPTED[class]` | 笔记呈现 |
|---|---|---|
| `numeric` | `{screenshot, matplotlib}` | agent 选:原图截图或族外 PDF;正文声明「本家族不画数值图」 |
| `raster` | `{screenshot}` | 表 / 照片 / 装置 / UI;**不能**选 matplotlib |
| `tensor` | `{supertensor}` | 委派 sibling |
| `fig` | `{superfig}` | 委派 sibling |
| `derive` | phase 1:`{align}`;phase 2:`{align, superderive}` | 见下 |
| `none` | `{none}` | 散文 + display math + 符号表 |
`toolkit: matplotlib` + `signals: [loss-curve]` ⇒ class `numeric` ⇒ **合法**。`toolkit: screenshot` + `signals: [table]` ⇒ class `raster` ⇒ **合法**。原先草图把 `suggest` 直接返回 `"screenshot"`,会误杀这两条。
**`align` 与 `derivations[].figure`:**
| 情况 | `derivations[].figure` | 是否需要 `figures[]` 行 |
|---|---|---|
| 默认(v1 几乎全部):笔记里 `align` | **必须 `null`** | **不需要**。推导不是图 |
| 可选 bookkeeping:想在 figure plan 里留下「已路由到 align」 | 可填该 `F*` | 一行 `toolkit: align`,`include: null`,`request: null`,`signals` ⊆ derive;`status: included` 表示 writer 已把 `align` 写进讲义节 |
| v2 且必须看见取消/代入,或 >4 步 | 填 `F*` | `toolkit: superderive`,走 sibling 委派 |
v1:任何 derive-class 行的 toolkit 必须是 `align`;`toolkit: superderive` → `SP022`(仅在 `SUPERPAPER_PHASE=1` 时触发;缺省已是 2)。phase 2 仍遵守 ≤4 步且无 `cancel-visual`/`subst-visual` ⇒ **必须** `align`(即使包已存在)。
`toolkit: none` 只用于「考虑过、决定不出图」:`status: dropped`,`include: null`。这些组合**不是** `SP010`(`SP010` 只表示 toolkit ∉ `ACCEPTED[class]`)。由下面 `check_row(..., status, include, fig_id)` 给出独立码。
**冲突拆分(硬规则,禁 `F1a`):**
- **任何一行** `classify(signals)` 不得返回多个 class。混类 → `SP011`,消息列出 class 集合并要求 outline 用**下一个整数 id** 拆成 N 行。
- 拆分算法:retire(写入 `retired_ids`)或删除混类行;新建 `F{max+1}`、`F{max+2}`、…;每子行 `signals` 是**单类子集**;每子行 `toolkit ∈ ACCEPTED[class]`。三路(`numeric+tensor+fig`)就是三个新整数。不保留 parent 行。
- 子行 `signals` 必须 ⊆ 该类的信号集合。
```python
# scripts/router.py — PR 1 检入的完整实现(不是草图)
from __future__ import annotations
NUMERIC = frozenset({"loss-curve", "bar", "scatter", "histogram", "numeric-plot"})
RASTER = frozenset({"table", "photo", "apparatus", "ui"})
TENSOR = frozenset({"axis", "shape", "transpose", "broadcast", "gather",
"shard", "contraction", "face"})
DERIVE = frozenset({"rewrite-figure", "cancel-visual", "subst-visual"})
FIG = frozenset({"architecture", "pipeline", "data-flow", "state", "time",
"dependency", "what-eats-what", "argument-map"})
CLASS_SIGNALS = {
"numeric": NUMERIC, "raster": RASTER, "tensor": TENSOR,
"derive": DERIVE, "fig": FIG,
}
def classify(signals: set[str]) -> list[str]:
s = set(signals)
hit = [c for c, vocab in CLASS_SIGNALS.items() if s & vocab]
return hit or ["none"]
def suggest(signals: set[str], *, phase: int = 1) -> str:
"""Return a class name, or 'SPLIT:a+b+...' in CLASS_SIGNALS order."""
classes = classify(signals)
if len(classes) > 1:
return "SPLIT:" + "+".join(classes)
return classes[0]
def accepted(cls: str, *, phase: int = 1) -> frozenset[str]:
if cls == "numeric":
return frozenset({"screenshot", "matplotlib"})
if cls == "raster":
return frozenset({"screenshot"})
if cls == "tensor":
return frozenset({"supertensor"})
if cls == "fig":
return frozenset({"superfig"})
if cls == "derive":
return frozenset({"align", "superderive"} if phase >= 2 else {"align"})
if cls == "none":
return frozenset({"none"})
raise KeyError(cls)
VECTOR = frozenset({"superfig", "supertensor", "superderive", "matplotlib"})
def include_pdf(fig_id: str) -> str:
return f"figures/{fig_id}/build/{fig_id}.pdf"
def include_png(fig_id: str) -> str:
return f"figures/{fig_id}/orig.png"
def check_row(
signals: set[str],
toolkit: str,
*,
status: str = "planned",
include: str | None = None,
fig_id: str = "F1",
phase: int = 1,
) -> str | None:
"""First matching code, or None. SP010 ≠ include/status rules."""
classes = classify(signals)
if len(classes) > 1:
return "SP011"
if toolkit not in accepted(classes[0], phase=phase):
return "SP010"
if toolkit == "none" and status != "dropped":
return "SP013"
if toolkit in {"align", "none"} and include is not None:
return "SP014"
if status == "included" and toolkit in VECTOR and include != include_pdf(fig_id):
return "SP015"
if status == "included" and toolkit == "screenshot" and include != include_png(fig_id):
return "SP016"
return None
```
谓词(lint-only;不写进 draft-07 `if`/`then`):
| 码 | 条件 |
|---|---|
| `SP010` | 单类,但 `toolkit ∉ ACCEPTED[class]` |
| `SP011` | `classify` 返回多个 class |
| `SP013` | `toolkit=none` ∧ `status ≠ dropped` |
| `SP014` | `toolkit ∈ {align,none}` ∧ `include is not None` |
| `SP015` | `status=included` ∧ 向量 toolkit ∧ `include ≠ figures/{id}/build/{id}.pdf`(含 `include` 为 null) |
| `SP016` | `status=included` ∧ `toolkit=screenshot` ∧ `include ≠ figures/{id}/orig.png` |
`status=included` ∧ `toolkit=align` ∧ `include=null` 合法(bookkeeping)。`SP021`(PR 4)只检查「`include` 已是非空字符串但文件不存在」,不再兼职这些结构规则。
`router_test.py`(PR 1 必须绿)至少覆盖:
- `[loss-curve] → numeric`;`accepted` 含 `matplotlib` 与 `screenshot`
- `[table] → raster`;`matplotlib` 不在 `accepted`
- `[axis, architecture] → SPLIT:tensor+fig` 且 `check_row` = `SP011`
- `[rewrite-figure]` phase 1:`align` ok,`superderive` → `SP010`
- `[]` 或 `[notation] → none`
- `none` + `status=included` → `SP013`;`align` + `include="figures/F1/build/F1.pdf"` → `SP014`
- `superfig` + `included` + `include=null` → `SP015`;`screenshot` + `included` + 错路径 → `SP016`
**默认偏向 `none`。** 公式不是图。只有「机制靠空间关系才能一次看清」才出图。
---
#### 5. 笔记教学法
**直接复用** `youtube-render-pdf/SKILL.md`「Pedagogical Standard」+「Writing Rules」1–2、5–11,论文侧改下列几处(写在 `references/pedagogy.md`,不要在 SKILL.md 复述 youtube 全文):
| youtube | superpaper |
|---|---|
| 时间戳 / 章节 | **§ / Eq.(n) / Figure n / Table n / p.** |
| 片头封面图 | **书目卡片**(题名、作者、年份、venue、arXiv、本地 PDF)。不做论文首页截图当封面 |
| `$$...$$` | **`\[` / `align` / `aligned`**(多步推导是一等公民;不沿用 `$$`) |
| `dialoguebox` | **`quotebox`**:短原文 + `§/Eq` 出处。禁止大段粘贴 PDF |
| 关键帧截图 | 源图 screenshot **或** sibling 重绘;重绘优先于架构类原图 |
| `\subsection{本章小结}` | 保留 |
| `\section{总结与延伸}` | 保留:论文 limitations / future work + 笔记作者的压缩与追问 |
| 默认中文 | 保留;用户明确要求才英文 |
| 禁止 `[cite]` 占位 | 保留;出处写进 `\spsource{...}` 或正文 |
公式三拍(不可拆):
1. 先用中文讲这式子在主张什么、为什么在这里出现;
2. display math(`\[` 或 `align`);
3. 立刻跟扁平符号表(`itemize`,每个符号一行:符号 — 含义 — 定义处)。
盒子语义与 youtube 相同,定义抄进 `assets/notes-template.tex`(颜色也抄:蓝/黄/红。**不要**把笔记盒子改成 figure 的 muted palette——笔记是文档层,图是图层次;v1 不统一)。图必须留在盒子外面(youtube Writing Rule 9 原句)。
每节内部顺序:动机 → 想法 → 机制 → 证据 → takeaway(本章小结)。不要按 PDF 页序复述。
建议顶层骨架(模板注释里写死;outline agent 可增删中间节,不可删首尾):
```tex
\section{这篇论文在问什么} % questions[] + 动机
\section{主张与贡献} % claims[status=core],\splabel{C*}
\section{预备:定义、假设、符号} % D* / A* / 符号表摘要
% --- 中间节由 outline 按论文机制切 ---
\section{实验与证据} % 若有 E*;可 omit
\section{总结与延伸}
\appendix
\section{符号表} % render_ledger.py 投影
\section{推导链一览}
\section{图表清单}
```
主张依赖图(argument map)**默认不出**。仅当 `claims + assumptions + lemmas + evidence` 的节点 ≥ 6 且用户没反对时,加一条 `figures[]`:`grammar: argument-map`,`toolkit: superfig`,用现有 `\sfnode`/`\sfarrow`(见 §「Argument-map figures」)。
---
#### 6. 委派时的 figure request 与嵌入约定
**谁画什么:**
| toolkit | 谁做 | 输入 | 输出 |
|---|---|---|---|
| `superfig` / `supertensor` / `superderive` | **figure agent**(每图一个) | 仅 `figures/F*/F*.request.md` + sibling `SKILL.md` | standalone `.tex` + sibling `build.sh` 产物 |
| `screenshot` | **主 agent**(不是 sibling figure agent) | `scripts/screenshot.sh --work <work> --id F*` | `figures/F*/orig.png` |
| `matplotlib` | **主 agent** 跑 `figures/F*/plot.py` | cwd = `<work>/notes/figures/F*` | `build/F*.pdf` |
| `align` / `none` | writer / 不画 | 无 request | 无 PDF |
lint `SP012`(PR 2):`toolkit ∈ {superfig,supertensor,superderive}` ∧ `status != dropped` ⇒ `<work>/notes/` + `figures[].request` 指向的文件必须存在。
Sibling skill **不**读 `ledger.yaml`。request 字段必须能 1:1 译成真实宏(`\sfnode` keys 只有 `role, level, gap, bracket, at=`;`\sfconn{name}{label}` 是 cursor 对象,**没有** from/to)。
**`allow-raw-tikz`:** sibling lint 只认源码注释 `% superfig-lint: allow-raw-tikz`(`superfig/scripts/lint.py` `directives()`)。request 里写 true 而不把该注释抄进 `.tex`,`build.sh` 仍失败。figure agent 必须两者一起写。
**superfig request(PR 2 example;与 `superfig/examples/pipeline.tex` 逐宏对应):**
```markdown
# Figure request F1
toolkit: superfig
language: cjk
claim: 一次前向是 x → f_θ → ŷ;损失不在主路上。
grammar: pipeline
# 以下路径一律相对 notes/
work_rel_dir: figures/F1
# 主 agent 必须这样调 sibling(第二个参数是 outdir,不要丢 PDF 在 work_rel_dir 根上):
# superfig/scripts/build.sh <work>/notes/figures/F1/F1.tex \
# <work>/notes/figures/F1/build
## Roles # \sfsetrole;宏吃 ROLE 不吃颜色
- {role: input, color: sfTeal}
- {role: model, color: sfOrange}
- {role: loss, color: sfCoral}
- {role: output, color: sfViolet}
## Flow (cursor). \sfconn 没有 endpoints。
stage: {name: SA, text: "推理流程:一次前向"}
row: {name: R1, height: 16mm}
in_row:
- {macro: sfnode, name: x, role: input, label: "输入 $x$", w: 16mm, h: 12mm}
- {macro: sfconn, name: e1, label: 预处理}
- {macro: sfnode, name: f, role: model, label: "模型 $f_\\theta$", w: 18mm, h: 12mm}
- {macro: sfconn, name: e2, label: logits}
- {macro: sfnode, name: y, role: output, label: "预测 $\\hat y$", w: 16mm, h: 12mm}
## Fixed topology(非线性才用;at= 必须是带 ($…$) 的 calc 坐标)
# 译成 \sfnode[..., at={($(y.south)+(0,-22mm)$)}]{s}{...}
- {macro: sfnode, name: s, role: loss, label: "损失 $L$", w: 14mm, h: 12mm,
at: "($(y.south)+(0,-22mm)$)"}
- {macro: sfarrowlabel, from: "y.south", to: "s.north", label: "$L(\\hat y, y)$"}
## After row
lane: R1
captions:
- {on: x, symbol: "$x$", detail: 原始输入}
- {on: f, symbol: "$f_\\theta$", detail: 可学习参数}
nolane: true
extra_captions:
- {on: s, symbol: "$L$", detail: 与真值比较}
topformula: "$x \\xrightarrow{f_\\theta} \\hat y,\\quad \\min_\\theta L(f_\\theta(x), y)$"
meaning:
idea: 一次从输入到预测的前向;损失在预测之后单独计算
objects: 数据、模型参数、预测、损失
mechanism: 预处理后送入模型;模型产生 logits 得到预测;损失比较预测与真值
## Lint
allow-raw-tikz: false
# 若为 true:F1.tex 顶部必须有 `% superfig-lint: allow-raw-tikz`
```
译出的 `.tex` 必须能被现有 `superfig/scripts/build.sh` 编过。`at` 字段是 **TikZ `calc` 表达式**,必须含外层 `($…$)`(与 `superfig/examples/pipeline.tex` 的 `at={($(y.south)+(0,-22mm)$)}` 一致)。禁止 `at: "below y"`,也禁止漏掉括号写成 `$(y.south)+(0,-22mm)$`。
**supertensor request(PR 3;字段对齐 `semantics.md` + `geometry.md` + `mha-causal.tex` 的 ledger,图可裁到 stage A 以免 example 过大):**
```markdown
# Figure request F2
toolkit: supertensor
language: cjk
claim: 每头打分沿 d_h 收缩;K^T 必须物理换面。
grammar: tensor-face
work_rel_dir: figures/F2
# build.sh <work>/notes/figures/F2/F2.tex <work>/notes/figures/F2/build
## Roles # \stsetrole
- {role: q, color: stTeal}
- {role: k, color: stOrange}
- {role: s, color: stCoral}
## Geometry # \stdim{axis}{cells}
- {axis: T, cells: 6}
- {axis: dh, cells: 3}
## Symbols # semantics.md 每条
- {name: Q, kind: value, dtype: R, shape: "h x T x d_h",
axes: {h: heads, T: time, dh: head dim}, producer: "linear_q"}
- {name: KT, kind: value, dtype: R, shape: "h x d_h x T",
note: "physical transpose of K; contracted axis dh matches Q width"}
- {name: S, kind: score, dtype: R, shape: "h x T x T"}
## Flow
stage: {name: SA, text: "每头打分:沿 $d_h$ 收缩"}
row: {name: rowA, height: T}
# coord: "" 必填。译码器对 ststack/stface/stglyph/stindexface
# 永远输出空坐标参数:\ststack[role=q, bracket=true]{Q}{}{T}{dh}{3}
# 丢掉 {} 会把 T 当成 coord,编出来的图与 stage A 对不上。
in_row:
- {macro: ststack, name: Q, role: q, coord: "", rows: T, cols: dh, sheets: 3, bracket: true}
- {macro: stglyph, name: mA, coord: "", glyph: "$\\times$"}
- {macro: ststack, name: KT, role: k, coord: "", rows: dh, cols: T, sheets: 3, bracket: true}
- {macro: stglyph, name: eA, coord: "", glyph: "$=$"}
- {macro: ststack, name: S, role: s, coord: "", rows: T, cols: T, sheets: 3}
captions:
- {on: Q, symbol: "$\\mathbf Q^{(i)}$", shape: "$h\\times T\\times d_h$"}
- {on: KT, symbol: "$\\mathbf K^{(i)\\top}$", shape: "$h\\times d_h\\times T$"}
- {on: S, symbol: "$\\mathbf S^{(i)}$", shape: "$h\\times T\\times T$"}
topformula: "$S^{(i)} = Q^{(i)} K^{(i)\\top}$"
meaning:
axes: "T 时间;d_h 头维;h 头数画成 stack 深度"
objects: "Q/K 是 value;S 是 score,不是 mask"
mechanism: "沿 d_h 收缩;K^T 换面,收缩边等长"
## Lint
# 源码若需要豁免,写 `% supertensor-lint: ...`,不要只写在本文件
```
不要在 supertensor request 里发明 `\sf*`。
**Figure agent 交付(与 `superfig/SKILL.md` Output 对齐):** PNG 预览、`F<id>-mechanism.md` 一段、`.tex` + `build/` 下 PDF/SVG/PNG。
```bash
/home/carry/myprj/tools/skills/superfig/scripts/build.sh \
"<work>/notes/figures/F1/F1.tex" \
"<work>/notes/figures/F1/build"
```
`F1.tex` 头与 golden 相同:`\documentclass[border=10pt]{standalone}` + `\usepackage[cjk]{superfig}`。
**Screenshot CLI(主 agent):**
```bash
./scripts/screenshot.sh --work <work> --id F2
```
算法(**页渲染是唯一主路径**;`crop_bbox` 的坐标系只对它有定义):
1. 读 `figures[F2]`。`source_pages` 必填,取**第一个**页码 `N`(多页截图拆成多行 `F*`,不在一页脚本里拼)。
2. 主路径:`pdftoppm -png -r 200 -f N -l N <work>/source/paper.pdf <tmp>/pg` → 得到一张 200 dpi 整页 PNG。这与 `crop_bbox`「200 dpi 页像素、左上原点」同一空间。
3. 若 `crop_bbox` 非 null:必须有 `magick`。没有 → `screenshot.sh` 非零退出(不要 silently 交整页)。有则 `magick <page>.png -crop WxH+X+Y +repage` 写出 `orig.png`(`W=x1-x0` 等)。
4. 若 `crop_bbox` 为 null:把整页复制为 `orig.png`。
5. **不要**用 `pdfimages -all` 当默认。它无页码、分辨率不是 200 dpi 页像素,无法实现「Table 2 on p.6 + crop_bbox」。v1 不提供 `raster_backend: embedded`。
`source_pages` 缺失或 `N` 越界 → 非零退出。
**Matplotlib:** `<work>/notes/figures/F3/plot.py` 必须在 cwd=`.../F3` 下接受 `--out build/F3.pdf` 并写出向量 PDF。主 agent:`python3 plot.py --out build/F3.pdf`。`include` 仍是 `figures/F3/build/F3.pdf`,用 `\spfig` 不是 `\spscreenshot`。
**笔记嵌入:**
```tex
\spfig{F1}{一次前向与侧路损失。}{重绘自论文 Figure~1(§3.1);toolkit: \texttt{superfig};ledger id: F1。}
\spscreenshot{F2}{WMT 2014 BLEU。}{论文 Table~2,p.6;toolkit: screenshot。}
```
宏签名见 API 节。禁止:`\input`/`\includestandalone` sibling 源;向量图用 PNG 当主嵌入;图进盒子;笔记 `\usepackage{superfig|supertensor|superderive}`(`SP003`)。
---
#### 7. Agent 拆分
对标 `wdkns-skills/README.md` 里 youtube 的四类角色,论文侧改成 **ledger 先行**。`references/agents.md` 是权威;SKILL.md 只留触发阈值和一份可粘贴 spawn 配方。
```mermaid
sequenceDiagram
participant U as User
participant M as Main / integrator
participant O as Outline agent
participant W as Writer agents
participant F as Figure agents
participant C as Consistency agent
participant R as Reviewer (optional follow-up)
U->>M: paper source + 范围
M->>M: preflight + ingest
M->>O: source/ + assets/ledger.example.yaml
O-->>M: ledger.yaml + outline.md + notes 骨架
alt 用户要求确认 ledger
M->>U: 展示 claims / figure plan
U-->>M: 修订
end
par 按节
M->>W: ledger + 本节边界 + 邻节 overlap
W-->>M: sections/sec-XX.tex
end
par 每张计划图一张 agent
M->>F: F*.request.md + 对应 sibling SKILL
F-->>M: tex + build/ + mechanism.md
end
M->>M: 核对 notes.tex 的 \\input 与 outline map;更新 figures[].status
M->>C: ledger + notes + figure 产物
C-->>M: consistency.yaml(id 列表)
M->>M: lint.py --work + build.sh --work → notes.pdf
opt 用户显式要求查漏
M->>R: 对照 source/paper.txt 与页渲染
R-->>M: 仅反馈,不改(与 youtube reviewer 相同)
end
M->>U: notes.pdf + ledger + 图产物
```
**单遍 vs 多 agent(闭合 9–12 页空洞;与 §1 表同一谓词):**
| 必须拆 | 单 agent(其余一切,含 9–12 页 ∧ 顶层节 ≤ 4 ∧ 重绘 ≤ 1) |
|---|---|
| 正文 **> 12** 页 | excerpt,或用户只要 ledger |
| 论文顶层节 **> 4** | 上列三个「必须拆」都不成立 |
| 计划重绘图 **≥ 2**(每图一个 figure agent) | |
| 用户 query **显式**要求 spawn | |
**Codex / 部分 runtime 的硬约束**(抄 `wdkns-skills/README.md`):只有用户 query **显式**要求 sub-agent / 并行时才会 spawn。因此 `SKILL.md` 必须附一份可粘贴配方,而不是假设运行时总会拆:
```text
$superpaper <源> 请 spawn 多 sub agents,隔离上下文:
- 1 个 outline agent:ingest 之后写 ledger.yaml + outline.md + 节边界
- N 个 writer agents:各写 sections/sec-XX.tex,引用 ledger id
- 每个计划重绘图 1 个 figure agent:只读 F*.request.md,调用对应 sibling skill
- 1 个 consistency agent:符号、术语、claim 覆盖、路由一致性
完成后用 superpaper/scripts/lint.py 与 build.sh 收口。
```
**Writer 工作单元 = `outline.md` 里的讲义节,不是 `coverage.sections_in`。**
`coverage.sections_in` 只限制**可以引用的论文节**(防 writer 去写被 skip 的附录)。切段、文件名、并行度全部看讲义 map。
`outline.md` **必须**含下列标题(否则 consistency 视为未完成):
```markdown
# Outline: <paper.title>
## Lecture map
| file | lecture_title | paper_sections | ledger_ids |
|---|---|---|---|
| sec-01.tex | 这篇论文在问什么 | 1 | Q1 |
| sec-02.tex | 主张与贡献 | 1, 3 | C1 |
| sec-03.tex | 预备:定义、假设、符号 | 2, 3.1 | D1, A1 |
| sec-04.tex | <outline 自定机制节> | 3.2 | C1, DER1, F1 |
| sec-05.tex | 实验与证据 | 5 | E1, F2 |
| sec-06.tex | 总结与延伸 | 6, 7 | C1 |
| sec-app-a.tex | 符号表 | — | (generated) |
| sec-app-b.tex | 推导链一览 | — | DER* |
| sec-app-c.tex | 图表清单 | — | F* |
## Locked
- 首节标题必须是「这篇论文在问什么」
- 末节(appendix 前)必须是「总结与延伸」
- 三个 appendix 行必须存在;`sec-app-a.tex` 只 `\input{sections/symbols.tex}`
```
`sec-XX` 的 XX 是**讲义序号**(01, 02, …),不是论文节号。不要按 PDF 页序复述;map 的 `paper_sections` 只是引用许可。
**谁写哪份文件:**
| 文件 | 作者 | 之后谁可以改 |
|---|---|---|
| `ledger.yaml` | outline(从 `assets/ledger.example.yaml` 复制) | 仅 outline / 主 agent(writer **禁止**) |
| `outline.md` | outline | 仅 outline |
| `notes/notes.tex` | outline:填元数据 + 按 map `\input{sections/sec-XX.tex}` | 仅当 map 增删行时由 outline 改。主 agent「缝合」= **核对 `\input` 列表与 map 一致**,不是重写正文 |
| `notes/sections/sec-XX.tex` | 每个文件 **一个 writer**(含首节、总结) | 该 writer |
| `notes/sections/symbols.tex` | `render_ledger.py` | 无人手改 |
| `notes/figures/F*/F*.request.md` | outline | figure agent 只读 |
首节/末节不由 outline 代写正文:outline 只落 `\section{...}` 空壳,writer 填。
**handoff:**
- Outline 输入:`source/` + `assets/ledger.example.yaml`。输出:磁盘上的 `ledger.yaml` + `outline.md` + `notes.tex` 骨架 + 空 `sec-XX.tex` + 各 `F*.request.md`。
- Writer 输入:`ledger.yaml` + 本行 map(file / lecture_title / paper_sections / ledger_ids)+ 邻接讲义节最后一段 overlap + `references/pedagogy.md`。不得新增 `status: core` claim。
- Figure 输入:仅 request + sibling SKILL。不读 `source/`。
- Consistency 输入:ledger + 全部 `sec-*.tex` + 各 `F*-mechanism.md`。输出 `<work>/consistency.yaml`:
```yaml
schema: superpaper.consistency/v1
unlabeled_core_claims: [] # C*
symbol_drift: [] # {name, ledger_kind, notes_hint}
term_mismatch: [] # {used, canonical}
router_mismatches: [] # F*
missing_artifacts: [] # F*
missing_requests: [] # F*
orphan_labels: [] # \splabel 指向不存在的 id
```
主 agent 按该文件改磁盘,再跑 `lint.py`。
**并发度:** 典型 12 页 = 1 outline + 3–5 writers + 0–3 figures + 1 consistency。Figure 目录互不共享。Writer 不改 ledger。
---
#### 8. 仓库布局(v1 第一张 PR 就必须按此建目录)
**Git 拓扑(已落地):** SuperPaper 是母仓库。`superfig/` 与 `supertensor/` 是
git submodule,远程分别为 `carrydela/SuperFig` 与 `carrydela/SuperTensor`。
`superderive/` 以后同样以子仓库加入。本地 `skills/superfig` 与
`skills/supertensor` 是指进本仓库的符号链接。子仓库的 `.sty` / lint 契约不变。
v1 笔记脚手架仍落在母仓库根下(与子仓库并列),不是塞进某一个 child:
```text
/home/carry/myprj/tools/skills/superpaper/
SKILL.md
README.md
agents/openai.yaml
assets/
notes-template.tex
ledger.schema.yaml
ledger.example.yaml
references/
input.md # 源类型、ingest、原图 vs 重绘
ledger.md # schema、id、retired_ids、kind 别名表
router.md # class / ACCEPTED / SPLIT 整数 id
figure-request.md # 两套 request 全文、screenshot.sh
pedagogy.md # 教学序列、公式三拍、盒子、骨架
agents.md # 讲义 map、阈值、spawn、consistency.yaml
api.md # 全部 \sp* / quotebox 签名
checklist.md # 笔记交付清单(不是图清单)
antipatterns.md # 编译干净但教错的笔记
fallback.md # 无 LaTeX / 扫描件 / sibling 缺失
scripts/
preflight.sh
ingest.sh
screenshot.sh # PR 2
lint.py
router.py # PR 1 完整实现
render_ledger.py
build.sh
test.sh
examples/
excerpt-toy/ # 就是一个 --work 树(见下)
pipeline-delegate/
tensor-delegate/
tests/
ledger-valid.yaml # --ledger only
ledger-invalid/ # --ledger only;按 PR 切片
loads-sty/ # 迷你 --work 树,专测 SP003
notes-smoke/ # 迷你 --work 树,只编过
router_test.py
```
对标 `superfig/README.md` 的「目录即契约」:`SKILL.md` 瘦、`references/` 按需加载、`scripts/{preflight.sh,lint.py,build.sh,test.sh}` + `ingest` / `router` / `render_ledger` / `screenshot`。
**不出现:** `assets/superpaper.sty`(v1 盒子和宏都放进 `notes-template.tex`,避免再长一个包)、`work/`、真实论文 PDF。
`agents/openai.yaml`:
```yaml
interface:
display_name: "Superpaper"
short_description: "把论文编成带 ledger 与路由出图的中文讲义"
default_prompt: >-
Use $superpaper to reconstruct this paper's argument thread,
expand the key derivations, and produce structured Chinese notes.
Build ledger.yaml first. Route figures to superfig / supertensor / align;
do not draw inside superpaper and do not \usepackage{superfig} in the notes.
```
---
#### 9. `SKILL.md` 形态(给第一张 PR 的正文大纲)
遵守 `~/.grok/bundled/skills/create-skill/SKILL.md` 的 frontmatter,以及 `skill-design-principles`:一事一处、SKILL 是 agent 提示词不是文档、细节指针到 `references/`。对标 `superfig/SKILL.md` 的瘦身程度(约 80 行),**不要**写成 youtube 那样 260 行的全抄。
建议 frontmatter(触发词必须能自动召回):
```yaml
---
name: superpaper
description: >-
Reconstruct a paper's argument thread and write structured Chinese
lecture notes with a claims/symbols/derivation ledger, then route
figures to superfig, supertensor, or (later) superderive. Use when
the user asks for 读论文, 论文笔记, 脉络梳理, 公式推导, 抽象公式图像化,
arXiv notes, paper notes, or $superpaper. Not for drawing a single
architecture figure (use superfig) or a tensor-shape figure (use
supertensor) or for turning a lecture video into notes
(use youtube-render-pdf / bilibili-render-pdf).
---
```
正文只保留:
1. 一句话定位:论文侧的 `youtube-render-pdf`,**不是** figure 包。
2. When to use / When not(指向 sibling 与视频 skill)。
3. 编号工作流:preflight → ingest → **ledger** → outline → writers → router/delegate → consistency → build。
4. Non-negotiables:ledger SSOT;笔记不加载 figure `.sty`;公式三拍;路由表;不发明 superfig 原语。
5. Output:`ledger.yaml` + `notes.tex` + `notes.pdf` + 各图产物。
6. Iterating:改 ledger 再改一处。
7. 指针:`references/*.md`。
8. 可粘贴 spawn 配方。
---
#### 10. 脚本契约
所有面向工作副本的脚本统一:
```bash
./scripts/ingest.sh --work <work> --arxiv 1706.03762
./scripts/ingest.sh --work <work> --pdf /path/p.pdf
./scripts/ingest.sh --work <work> --tex /path/main.tex
./scripts/ingest.sh --work <work> --excerpt /path/clip.md
./scripts/lint.py --work <work>
./scripts/lint.py --ledger tests/ledger-valid.yaml
# YAML-only:SP001/SP002,以及已启用的
# SP010/SP011/SP013–SP016/SP022/SP023/SP024
# 跳过必须看见笔记或磁盘文件的码:SP003、SP012、SP020、SP021
./scripts/lint.py --ledger X.yaml --notes path/to/notes.tex
# 另加 SP003(及已启用的 SP020);仍无 --work 则跳过
# SP012 / SP021(它们解析 <work>/notes/ 下的文件)
./scripts/render_ledger.py --work <work> # 写 notes/sections/symbols.tex
./scripts/build.sh --work <work>
./scripts/screenshot.sh --work <work> --id F2
```
兼容别名:`--out` = `--work`(ingest 文档里两名同义,实现只保留 `--work`)。
**每个 `examples/*` 目录就是一棵 `--work` 树**,不是扁的 `notes.tex`:
```text
examples/excerpt-toy/
ledger.yaml
outline.md
source/excerpt.md
notes/notes.tex
notes/sections/sec-01.tex
```
`test.sh` 对 example / `tests/notes-smoke` / `tests/loads-sty` 一律 `build.sh --work <that-dir>` 或 `lint.py --work <that-dir>`。`tests/ledger-invalid/*.yaml` 继续 `--ledger`(无笔记)。
**`preflight.sh`** — 0 完整 / 1 降级 / 2 无引擎。对标 sibling:缺 CJK 是 1 不是 2。
| 档 | 检查 | 缺了 |
|---|---|---|
| 引擎 | `xelatex`、`article.cls` | **exit 2**(「无 LaTeX」只用于这一档) |
| 模板/CJK(exit 1) | `ctex.sty`、`FandolSong-Regular.otf`、`amsmath.sty`、`amssymb.sty`、`tcolorbox.sty`、`graphicx.sty`、`hyperref.sty`、`geometry.sty`、`listings.sty`、`booktabs.sty`、`subcaption.sty`、`float.sty`(`figure[H]`)、`tikz.sty`、`etoolbox.sty` | 降级:无 Fandol 则英文笔记或稍后 Missing character |
| Python(exit 1) | `python3`、`import yaml`、`import jsonschema` | 不能 lint |
| ingest(exit 1) | `pdftotext`、`pdftoppm` | 只吃 tex/excerpt |
| 可选 | `magick`、`pdftocairo`、`curl`/`wget`、sibling 目录 | 无 `magick` 时 `crop_bbox` 必须失败(见 `screenshot.sh`),不是静默整页;不升到 2 |
不要求 `standalone.cls`。
**`ingest.sh`:** 写 `source/meta.yaml`。`pdftoppm -png -r 120` 之后**重命名**为 `pages/pg-%03d.png`(永远三位数,避免 8 页 `pg-1` vs 15 页 `pg-01`)。网络失败可重入。
**`lint.py` 码表(按 PR 启用;未启用的码不跑、对应 fixture 不进 `test.sh`):**
| 码 | PR | 失败条件 |
|---|---|---|
| `SP001` | 1 | `jsonschema` 失败(缺 required、坏 enum、坏类型) |
| `SP002` | 1 | 任意数组内 id 重复 |
| `SP003` | 1 | 笔记源 `\usepackage` 匹配 `superfig\|supertensor\|superderive` |
| `SP010` | 2 | 单类但 toolkit ∉ `ACCEPTED[class]`(**不含** include/status) |
| `SP011` | 2 | 单行 `classify` 返回多个 class |
| `SP012` | 2 | sibling toolkit ∧ `status != dropped` ∧ request 文件不存在 |
| `SP013` | 2 | `toolkit=none` ∧ `status ≠ dropped` |
| `SP014` | 2 | `toolkit ∈ {align,none}` ∧ `include is not None` |
| `SP015` | 2 | `included` ∧ 向量 toolkit ∧ `include ≠ figures/{id}/build/{id}.pdf` |
| `SP016` | 2 | `included` ∧ `screenshot` ∧ `include ≠ figures/{id}/orig.png` |
| `SP020` | 4 | `status: core` 的 claim 在 `notes/**/*.tex` 无 `\splabel{C*}`(`coverage.mode=excerpt` 同样查 ledger 里列出的 core) |
| `SP021` | 4 | `include` 已是非空路径但 `<work>/notes/<include>` 不存在 |
| `SP022` | 4 | v1(`SUPERPAPER_PHASE` 缺省 2,PR 7 已落地;仅当显式设为 1 时触发)出现 `toolkit: superderive` |
| `SP023` | 4 | 非 dropped 的新 id ∈ `retired_ids` |
| `SP024` | 4 | `figures[].claim` / `derivations[].claim` 不是已有 `C*` |
stderr 格式**抄** `superfig/scripts/lint.py`:先 `!! {path}`,随后缩进行 ` line N: SP00X message`(无行号则 ` SP00X message`)。
警告(不失败):未知顶层键、`coverage.mode=full` 且 `sections_in` 空、符号从未在笔记出现。未知 `steps[].rule` **不是**警告(`SP001`)。
**禁止**复制 superfig 的 raw-tikz / callout / 4-hue 规则。
**`build.sh --work <work>`:**
1. `lint.py --work <work>`
2. `render_ledger.py --work <work>`
3. `cd <work>/notes && xelatex -halt-on-error notes.tex` 两遍(`\includegraphics` 相对 `notes/`)
4. `Missing character` / undefined references / multiply-defined = 失败
5. **Overfull/Underfull 不失败**
6. 复制 PDF 到 `<work>/out/notes.pdf`
**`test.sh`:** 只跑**当前已启用码**对应的 fixture。example 与 `tests/{notes-smoke,loads-sty}` 按 `--work` 调。PR 2/3 example 的图 PDF 预构建并提交;`test.sh` 编笔记,不重跑 sibling `build.sh`。
**`render_ledger.py --work`:** 写 `notes/sections/symbols.tex`;可选 `<work>/claims.md`。
**`screenshot.sh`:** 见 §6。
---
#### 11. Argument-map figures(本设计允许的**唯一** superfig「扩展」,且是后置 PR)
v1 **零改** `superfig/`。主张图先用现有原语 + 文档化 role 约定:
```tex
\sfsetrole{claim}{sfTeal}
\sfsetrole{assumption}{sfOrange}
\sfsetrole{lemma}{sfViolet}
\sfsetrole{evidence}{sfCoral}
% 然后 \sfnode / \sfarrow;边语义只能是 dependency 或 causality
% (superfig/references/grammar.md:data | control | dependency | causality)
```
四点约束:
- **不新增原语。** 不需要 `\sfclaim`。role 名只是 `\sfsetrole` 字符串。
- **不改 lint。** 4 色相预算刚好够这四个 role + gray。第五种节点用 `level=` 或拆图。
- **后置、可加性。** PR 5 最多:`superfig/examples/argument-map.tex` 一张 golden + `references/grammar.md` 表格加一行「argument map = 现有 node/edge + 上述 role」。`SKILL.md`「When to use」可加半句。不改 `.sty`。
- 推荐 **不要** 为 argument map 改 `superfig` 的 description 触发词;路由由 `superpaper` 拥有。
---
### Layer 2 — `superderive`(现在写清,以后再实现)
新的 sibling **figure** toolkit:逐步代数推导图。视觉语法既不是 node/edge,也不是 tensor face。
**不要**往 `superfig.sty` 加 `align` 风格宏。
#### 视觉语法
三区仍对齐家族(顶 claim / 中结构 / 底 meaning),但中区是 **改写行** 不是流水线:
| 要暴露的事实 | 语法 | 宏(草图) |
|---|---|---|
| 一行合法改写 | step | `\sdstep{name}{lhs}{rhs}` |
| 为何合法 | 右轨 justification | `\sdreason{name}{text}` |
| 消去 | strike + cancel 色 | `\sdcancel{name}{subterm}` |
| 代入 | 源/目标高亮 | `\sdsubst{name}{from}{to}` |
| 焦点子式 | 不改变项的框 | `\sdbox{name}{subterm}` |
| 非法 vs 合法对照 | 两列,非法列 role=warn | `\sdcol` … `\sdcolend` |
| 底注 | 三行 meaning | `\sdmeaningbox`(idea / rewrite / caveat) |
House style **复制** `superfig` 骨架(palette 名改为 `sdTeal` 等,HEX 相同;一 role 一色;≤4 色相;lint-on-warning;`scripts/{preflight.sh,lint.py,build.sh,test.sh}`)。v1/v2 都不抽 `superstyle`。
Role 建议:`keep`(保留项)、`rewrite`(本步触碰)、`cancel`、`intro`(新引入)。不要把每一步换成新色相。
#### 何时由 superpaper 调用
见上表。额外:若 `derivations[].expand: true` 且步数 ≤ 4 且没有任何 `cancel-visual` / `subst-visual`,即使以后有 `superderive` 也 **仍用 `align`**。图是为「必须看见」服务的,不是为「有包就用」服务的。
#### API 草图(实现时可以微调名字,但前缀锁定 `\sd`)
```tex
\documentclass[border=10pt]{standalone}
\usepackage[cjk]{superderive}
\sdsetrole{keep}{sdTeal}
\sdsetrole{rewrite}{sdOrange}
\sdsetrole{cancel}{sdCoral}
\begin{document}\begin{tikzpicture}
\sdstage{D1}{缩放来自方差,不是装饰}
\sdrow{R1}
\sdstep{s1}{QK^{\top}}{QK^{\top}/\sqrt{d_k}}
\sdreason{s1}{§3.2.1,避免 softmax 饱和}
\sdrowend
\sdbbox{all}
\sdtopformula{F}{\mathrm{Attention}(Q,K,V)=\mathrm{softmax}(QK^{\top}/\sqrt{d_k})V}
\sdmeaningbox{mb}{120mm}{all}
{点积方差随 $d_k$ 线性涨}
{缩放是改写,不是新算子}
{没有缩放,softmax 进饱和区,梯度消失}
\end{tikzpicture}\end{document}
```
Lint(v2,对标 `superfig/scripts/lint.py` 的职责切分):未声明 role、>4 色相、raw `\draw`、`\sdtopformula` 出现在最后 `\sdrowend` 之前、一图多个「非法对照」列(预算:一列对照)。**不要**把「公式必须在最后一行之后」理解成禁止中间出现数学——数学就是 step 的内容;禁的是**顶栏 claim** 提前居中。
#### `superderive/` 目录(phase 2 才建,形状抄 `superfig/`)
```text
/home/carry/myprj/tools/skills/superderive/
SKILL.md
README.md
agents/openai.yaml
assets/superderive.sty
references/{grammar,layout,style,api,checklist,antipatterns,fallback}.md
scripts/{preflight.sh,lint.py,build.sh,test.sh}
examples/{rewrite-cancel.tex,subst-chain.tex,antipatterns.tex}
tests/{smoke.tex,lint-invalid/,invalid/}
```
Provenance 写:「第三 sibling;不是从 superfig 长出来的。superpaper router 在 phase≥2 才发出 `toolkit: superderive`。」`tensor-formula-viz` 与 `superfig` 都不动。
`agents/openai.yaml`:`Use $superderive to turn a rewrite sequence into a step/justification figure. Not for architecture (superfig) or tensor faces (supertensor).`
---
## API / Interface Changes
### 新建
| 接口 | 消费者 | 稳定承诺 |
|---|---|---|
| `superpaper` skill | 用户 / 主 agent | 见 SKILL description |
| `ledger.yaml` `superpaper.ledger/v1` | outline / writers / lint | §3 schema;未知键警告 |
| `router.classify` / `suggest` / `accepted` / `check_row` | lint + agent | class 或 `SPLIT:a+b`;不是 toolkit |
| `--work <dir>` CLI | 全部 scripts | ledger 在 `<work>/ledger.yaml` |
| `\sp*` / `\quotebox` | 笔记源 | 下方签名 |
| `F*.request.md` | figure agent | §6 两套模板 |
| `superderive` `\sd*`(v2) | figure agent | v1 不存在 |
### 笔记宏(检入 `assets/notes-template.tex` 与 `references/api.md`)
```tex
% 锚:hypertarget + label,供 PR 4 用正则 \\splabel\{C1\} 扫描
\newcommand{\splabel}[1]{\hypertarget{sp:#1}{}\label{sp:#1}}
% 可点击引用;排版为等宽 id
\newcommand{\spref}[1]{\hyperlink{sp:#1}{\texttt{#1}}}
% 出处脚注
\newcommand{\spsource}[1]{\footnote{来源:#1}}
% 向量图:读 figures/#2/build/#2.pdf
\newcommand{\spfig}[4][0.92\textwidth]{%
\begin{figure}[H]\centering
\includegraphics[width=#1]{figures/#2/build/#2.pdf}%
\caption{#3\protect\footnotemark}\end{figure}
\footnotetext{#4}}
% 截图:读 figures/#2/orig.png
\newcommand{\spscreenshot}[4][0.92\textwidth]{%
\begin{figure}[H]\centering
\includegraphics[width=#1]{figures/#2/orig.png}%
\caption{#3\protect\footnotemark}\end{figure}
\footnotetext{#4}}
% 短原文;#1 = 标题(含 §/Eq)
\newtcolorbox{quotebox}[1]{
enhanced, breakable,
colback=black!3!white, colframe=black!55, colbacktitle=black!55,
coltitle=white, fonttitle=\bfseries, title=#1, sharp corners}
```
用法:`\splabel{C1}` 放在陈述该 claim 的段首;`\spref{C1}` 在后文回指;`\spsource{§3.2, Eq.(4)}` 作句末脚注。`quotebox` 正文禁止超过约 8 行。图不得放入任何盒子。需要 `float`(`[H]`)和 `hyperref`(`\hypertarget`)。
### 对现有包
| 包 | v1 | 后置 |
|---|---|---|
| `superfig` | **零改动** | PR 5:additive example + grammar 一行。不改 `.sty`、不改 lint |
| `supertensor` | **零改动** | 无 |
| `tensor-formula-viz` | **不动**(与 `supertensor/README.md` Provenance 一致) | 无 |
| `youtube-render-pdf` / `bilibili-render-pdf` | **不动**;盒子定义复制进 superpaper 模板 | 不抽共享模板 |
无 before/after 宏变更。笔记与图的唯一运行时耦合是 **相对 `notes/` 的路径** `figures/F<id>/build/F<id>.pdf`(截图为 `figures/F<id>/orig.png`)。
---
## Data Model Changes
无数据库。磁盘上的 SSOT 是 `ledger.yaml`。
**迁移:** `schema: superpaper.ledger/v1`。破坏性变更加 `v2`。`render_ledger.py` 投影随时可删可再生。
**id 稳定性:** 顶层 `retired_ids: [C1, F2, …]`。废弃时把 id 写入该数组,并从工作数组删除或保留为不可变的 `status: dropped` 行。新对象只用下一个未占用整数。lint `SP023`:非 dropped id ∩ `retired_ids` ⇒ 失败。不存在「两条同 id(一条 dropped 一条新)」的检查——那已经是 `SP002`。
---
## Alternatives Considered
### A. 做大 `superfig`
把笔记宏、`align` 推导、主张 role、甚至「文档模式」option 塞进 `superfig.sty`。
| 利 | 弊 |
|---|---|
| 一个 skill 名 | 图尺度 lint 与文档尺度互相污染(一 callout、4 色相、formula 顺序、raw tikz) |
| | `standalone` 与 `article` 的构建失败语义不同(`build.sh` 对 overfull 零容忍,中文 article 做不到) |
| | 逐步改写会逼出与 `\sfnode` 无关的宏,包变成厨房水槽 |
| | 与「为何从 tensor-formula-viz 拆出 supertensor」的理由直接相反 |
**否决。** 产品决策已定;仓库证据也支持:`superfig.sty` 文件头自称 “A smaller sibling of `supertensor.sty`… primitives are generic nodes and edges, not tensor faces.” 它的价值就是小。
### B. 一个 mega-skill(读论文 + 两种图 + 推导 + 视频?)
| 利 | 弊 |
|---|---|
| 用户只记一个名字 | SKILL.md 无法保持「瘦 + 按需 references」;与 skill-design-principles 冲突 |
| | youtube 的帧召回纪律和 superfig 的 4 色相纪律写在同一提示词里,agent 会串规则 |
| | 无法对「只画一张架构图」保持 `superfig` 的短 description 触发 |
**否决。** 触发词污染:用户说「画一张 pipeline」不应拉起读论文流水线。
### C. Sibling family(采纳)
与 `superfig` / `supertensor` 的拆分理由相同,再加一层文档:
- **失败模式不同。** 图的失败是「编译干净但几何/语义在撒谎」(shards 没铺满、未转置却标 \(K^\top\))。笔记的失败是「编译干净但漏 claim、符号漂移、路由错引擎」。推导图的失败是「把改写画成假拓扑」。三种失败要三套不变量、三套 lint。
- **原语不同。** nodes/edges ≠ faces/axes ≠ steps/justification。共用宏包只会让最小语法选择失效。
- **构建目标不同。** standalone 一页图 vs article 讲义 vs(v2)又一页 standalone 图。
- **风格复制可接受。** 两套包已经复制 palette;第三套 figure 包(`superderive`)继续复制。文档层连 palette 都不共享(youtube 盒子颜色)。抽取 `superstyle` 的时机是第四个 **figure** 包出现之后,不是现在。
`superpaper` 这个名字放文档层,避免 `paper-render-pdf` 暗示它是 youtube 的克隆(源、出处、公式规则都不同)。
### 曾短暂考虑、已丢弃的小方案
- **Markdown ledger:** 对人友好,对 PR 4 的 lint 不友好。改为 YAML + 只读 Markdown 投影。
- **`\input` standalone TikZ 进笔记:** 省一次 PDF。会把 `Package superfig Warning` 打进笔记 log,并迫使笔记 `build.sh` 解释图尺度警告。否决。
- **v1 就抽 notes 模板共享包:** 要改 `wdkns-skills`,违反 non-goal。复制盒子。
---
## Security & Privacy Considerations
本 skill 是**本地、单用户、不可信 PDF** 的批处理,不是服务。威胁模型小但具体:
| 威胁 | 严重度 | 缓解 |
|---|---|---|
| arXiv id / 路径注入 `ingest.sh` | 高 | id 正则;路径不当作 shell 拼接;用数组传参 |
| 解包 e-print 后编译/执行 | 高 | **禁止**对 e-print 跑 latex/make;只当文本读 |
| 恶意 PDF 的工具链漏洞(poppler) | 中 | 本地可信用户假设;不把 ingest 暴露成网络服务 |
| 把受版权论文提交进 git | 中 | `work/` gitignore;`examples/` 只用自造 fixture |
| ledger / 笔记里粘贴隐私批注 | 低 | 不上传;无遥测 |
| figure agent 读整篇论文导致提示词膨胀 / 数据扩散 | 低 | request 文件隔离;figure agent 禁止读 `source/` |
无认证、无多租户、无密钥。`ingest.sh` 只访问用户给出的本地路径与 arXiv。
---
## Observability
不是线上服务。可观察性 = **本地构建纪律**(抄 sibling,但指标不同):
| 信号 | 来源 | 处理 |
|---|---|---|
| preflight 0/1/2 | `preflight.sh` | 1/2 必须在交付中说出口(对标 `superfig/references/fallback.md`) |
| ingest 页数、`sha256`、是否拿到 e-print | `source/meta.yaml` + `run.log` | 主 agent 写进笔记书目卡片 |
| lint 失败列表 | `!! {path}` 后跟缩进 ` line N: SP00X …`(与 `superfig/scripts/lint.py` 相同) | `build.sh` 非零退出 |
| 路由冲突 | `SPLIT:` / toolkit mismatch | 失败,不静默改 toolkit |
| 笔记 TeX | `Missing character`、undefined ref | 失败 |
| 笔记 TeX | Overfull | **记录在 `run.log`,不失败** |
| 图 TeX | sibling `build.sh` | 图目录内失败;主笔记 build 若 `status: included` 缺 PDF 再失败 |
| claim 覆盖 | core 缺 `\splabel` | 失败 |
| 耗时 | `run.log` 时间戳 | 人工;目标见下 |
**量级目标(单用户,12 页、2 张重绘图):**
- ingest(已有 PDF):< 30 s(`pdftotext` + 120 dpi 页渲染)
- ingest(arXiv 冷下载):受网络限制,60–120 s 可接受
- 单张 sibling 图:`build.sh` 5–20 s(与现仓库 `superfig/examples/pipeline.tex` 同量级)
- 笔记两遍 xelatex:10–40 s
- Agent 墙钟(多 agent):15–40 min,主要在模型,不在脚本
无 metrics backend、无 alerting。CI 就是 `scripts/test.sh`。
---
## Rollout Plan
本地 skill。开关是 **lint 码启用表 + `SUPERPAPER_PHASE`(缺省 2,PR 7 已合入)**。依赖链 1→2→3→4、PR 5/6 并行,均已完成;**只有 PR 7** 负责 phase 缺省翻转,2026-08-17 已随 `e6918ff` 落地。
1. PR 1:skill 可跑无委派笔记。`router.py` **完整**(class / `ACCEPTED` / `SPLIT` / `check_row`),但 lint **不**跑 `SP010–SP016`。SKILL 写「委派是后续 PR」。
2. PR 2:启用 `SP010–SP016`;superfig 委派 example(出生即带 `\splabel`)。
3. PR 3:supertensor 委派 example(同样自带 `\splabel`)。
4. PR 4:启用 `SP020–SP024`。example 不得返工。
5. PR 5:optional,只动 `superfig/examples` + grammar 一行。
6. PR 6:落地 `superderive/` 宏包与测试。**不改** router 缺省 phase,不改 `SP022`。
7. PR 7:`SUPERPAPER_PHASE` 缺省 2;关掉 `SP022`;加 derive-delegate example。
**回滚:** 删/回退对应目录即可。v1 对 sibling 零改动,回滚 superpaper **不会**留下 sty 垃圾。PR 5 回滚只删 example 与 grammar 那一行。PR 7 回滚即回到 phase 1,宏包可留。
**降级路径**(`references/fallback.md`):
- 无 XeLaTeX:仍交 `ledger.yaml` + Markdown 笔记;声明 PDF 未产出。
- 无 `pdftotext`:只接受 `tex` / `excerpt`。
- sibling 缺失或 figure `build.sh` 失败:该图降级为 `align` 或 screenshot,ledger 写 `status: dropped` + 原因;笔记仍须能编过。
---
## Risks
| 风险 | 严重度 | 机制 | 缓解 |
|---|---|---|---|
| 选错 toolkit(框图当 tensor 画,或反过来) | 高 | 图「看起来对」但撒谎;正是 sibling 存在的理由 | 可执行 `router.py`;跨类强制拆图;request 里写死 grammar |
| 符号漂移(正文、符号表、图 caption 三套名字) | 高 | 长论文 + 多 writer | YAML SSOT;`\splabel`;consistency agent;lint;`terms[].aliases` |
| 长论文召回失败 | 高 | youtube README 已承认的「AI extraction 漏召回」 | 切段 + overlap;页渲染召回;可选 reviewer **只反馈不改**;excerpt/body-only 默认 |
| superfig 契约被慢慢污染 | 高 | 「顺便加个宏吧」 | v1 零改动;本设计只允许 PR 5 加 example;评审检查表写明 |
| 笔记 build 误用图尺度失败条件 | 中 | 中文 article 全是 overfull | `build.sh` 明确不把 overfull 当失败;lint 禁止笔记加载 figure sty |
| standalone 嵌入方式选错 | 中 | class 冲突 / 警告泄漏 | 只 `\includegraphics` PDF |
| YAML 数学引号把公式写坏 | 中 | outline agent 产出非法 YAML | example + schema;lint 先解析;字面块 `|` |
| Codex 不 spawn,单上下文丢细节 | 中 | 平台约束 | SKILL 内可粘贴配方;单 agent 阈值写清 |
| 把论文原图当「已经可视化」 | 中 | 架构图截图无法审计语义 | 路由表:架构优先重绘 |
| `superderive` 延期导致推导体验差 | 低 | v1 用 `align` 是刻意的 | 公式三拍 + 展开规则已够教学;图是增强 |
| e-print 解包炸弹 | 低 | 恶意 tar | 限制解压大小/文件数(ingest.sh);永不编译 |
| 工作目录塞进 git | 低 | 大二进制 + 版权 | gitignore;CI 不跑 ingest 网络 |
---
## What “done” looks like
### v1(本设计的实现完成线)
- 目录 `superpaper/` 按上文树存在,`./scripts/preflight.sh` 与 `./scripts/test.sh` 在干净 TeX 环境下绿。
- Agent 仅读 `SKILL.md` + 按需 `references/` 能走通:摘录 → `ledger.yaml` → 中文 `notes.tex` → `notes.pdf`。
- 公式三拍、`\splabel`、书目卡片、本章小结、总结与延伸,都在模板注释和 pedagogy 里写死。
- Router 对决策表有单测;跨类 signals 必须拆图。
- 至少一个 **superfig** 委派 example、一个 **supertensor** 委派 example(图 PDF 预构建)。
- 笔记源加载 `superfig.sty` 会被 lint 拒绝。
- `toolkit: superderive` 会被 lint 拒绝,并提示用 `align`。
- `superfig/`、`supertensor/`、`wdkns-skills/` 的 diff 为空。
### v2
- `superderive/` 以 superfig 同构形状存在,一张 golden rewrite 图 + 负例 lint。
- PR 7 把 `accepted("derive")` 扩为 `{align, superderive}` 并停掉 `SP022`;example 增补一个委派。`suggest()` 仍然返回 class `derive`,不是 toolkit 名。
- PR 5 的 argument-map golden 可先于或后于 superderive 独立合并。
---
## Key Decisions
1. **按层拆 sibling,不升级 `superfig`。** 文档失败模式、图失败模式、推导图失败模式需要三套不变量。这与 `superfig` 从「通用图」里把张量面拆给 `supertensor` 的理由相同。
2. **`superpaper` 是编排 skill,不是 `.sty`。** 对标 `youtube-render-pdf`,不对标 `superfig.sty`。v1 不引入 `superpaper.sty`。
3. **「抽象公式图像化」是 router,不是第三套画笔。** `suggest()` 返回 **class**(`numeric|raster|tensor|derive|fig|none|SPLIT:…`);`ACCEPTED[class]` 才是 toolkit 集合(`numeric → {screenshot, matplotlib}`,`raster → {screenshot}`)。
4. **Ledger 用 YAML + draft-07 + `jsonschema`。** 未知键 lint 警告、schema 不 `additionalProperties: false`。不靠 `agents/openai.yaml` 当理由。
5. **v1 对 `superfig` / `supertensor` / `wdkns-skills` 零改动。** 唯一允许的后续触碰是 PR 5 的可加性 example。
6. **唯一工作根 `--work`;ledger 在 `<work>/ledger.yaml`;笔记路径一律相对 `notes/`。** `\spfig` 与 `figures[].include` 同形。
7. **House style:figure 包继续复制 muted palette;笔记盒子继续复制 youtube 的蓝/黄/红。** v1 不抽 `superstyle`。
8. **v1 逐步推导用 `align`(`derivations[].figure = null`),不用假框图。** 可选 `toolkit: align` bookkeeping 行没有 PDF。
9. **主张图用现有 `\sfnode`/`\sfarrow` + 四个 role 名。** 不改 `.sty`。
10. **`symbols[].kind` 以 `semantics.md` 左列为 canonical;斜杠右侧是别名;`scalar`/`set` 是文档层扩展(`d_k` 是 `shape parameter`)。**
11. **必须拆 agent 当且仅当:页 > 12 ∨ 顶层节 > 4 ∨ 重绘 ≥ 2 ∨ 用户 spawn。** 9–12 页落在单 agent。Writer 按讲义节切,不按 `coverage.sections_in`。
12. **examples 用自造 fixture,不提交论文 PDF。出生即带 `\splabel`。**
13. **笔记 `build.sh` 不因 overfull 失败;图的 `build.sh` 保持零容忍。**
14. **Figure agent 只看见 `F*.request.md`;截图走 `screenshot.sh`(主 agent)。**
15. **e-print 只读不编译。id 退役进 `retired_ids`。拆分只用下一个整数,不用 `F1a`。**
16. **PR 链 1→2→3→4 可审但后条依赖前条;phase 默认只在 PR 7 翻转。**
17. **`lint.py` 先 normalize `symbols[].kind` 再 `SP001`。** `check_row` 的 class/toolkit 失配是 `SP010`/`SP011`;include/status 是独立码 `SP013–SP016`。截图主路径是 `pdftoppm -r 200`,不是 `pdfimages`。
---
## Open Questions
在已给产品决策和仓库约束下,实施所需的产品选择已经闭合。下列两项**不阻塞 v1**,有偏好再改:
1. **PR 2/3 的委派 example 用哪条自造 claim?** 默认改编 `superfig/examples/pipeline.tex`(前向 + 侧路损失)和 `supertensor/examples/mha-causal.tex` 的教学点,写成「论文摘录 fixture」,避免引入新的科学内容。若希望 example 更「像一篇真论文」,再换。
2. **`SUPERPAPER_PHASE` 升 2 的时机。** 已定为 PR 7(`superderive/scripts/test.sh` 全绿之后),不绑在 PR 6 的 sty。
无需用户在实现前回答即可开工 PR 1。
---
## References
- `/home/carry/myprj/tools/skills/superfig/SKILL.md` — 文档层对标的「瘦 skill + ledger-before-draw + Output 三件套」
- `/home/carry/myprj/tools/skills/superfig/README.md` — sibling 目录契约、lint-on-warning 哲学
- `/home/carry/myprj/tools/skills/superfig/references/{grammar,api,style,checklist,layout,antipatterns,fallback}.md`
- `/home/carry/myprj/tools/skills/superfig/assets/superfig.sty` — `\sf*`、palette、role registry
- `/home/carry/myprj/tools/skills/superfig/scripts/{preflight.sh,lint.py,build.sh,test.sh}`
- `/home/carry/myprj/tools/skills/superfig/examples/{pipeline,branch-architecture,state-flow,dependency-graph}.tex`
- `/home/carry/myprj/tools/skills/supertensor/SKILL.md`、`README.md`(含 Provenance)
- `/home/carry/myprj/tools/skills/supertensor/references/semantics.md` — 符号 kind 的单一来源
- `/home/carry/myprj/tools/skills/supertensor/references/api.md` — `\st*` / `\stdim` / `\stsetrole`
- `/home/carry/myprj/tools/skills/wdkns-skills/skills/youtube-render-pdf/SKILL.md`
- `/home/carry/myprj/tools/skills/wdkns-skills/skills/youtube-render-pdf/assets/notes-template.tex` — 盒子与书目页
- `/home/carry/myprj/tools/skills/wdkns-skills/skills/bilibili-render-pdf/SKILL.md` — 降级阶梯
- `/home/carry/myprj/tools/skills/wdkns-skills/README.md` — 多 agent 配方与 reviewer 查漏
- `/home/carry/myprj/tools/skills/wdkns-skills/skills/tensor-formula-viz/` — 保持原地不动
- `/home/carry/myprj/tools/skills/wdkns-skills/templates/writing/readme.md`
- `/home/carry/.grok/bundled/skills/skill-design-principles/SKILL.md`
- `/home/carry/.grok/bundled/skills/create-skill/SKILL.md`
---
## PR Plan
**1→2→3→4 是依赖链**:每条可单独审查,但 2/3/4/7 **不能**在没有前驱的情况下合并。5 与 6 可并行。不要再说「每条独立可合并」。
### PR 1 — `superpaper` 脚手架 + SKILL + ledger + 笔记模板(无出图委派)
- **Title:** `superpaper: scaffold skill, ledger schema, and notes template`
- **Files/components:**
- 新建 `superpaper/SKILL.md`、`README.md`、`agents/openai.yaml`
- `assets/{notes-template.tex,ledger.schema.yaml,ledger.example.yaml}`(schema = §3 全文;模板含全部 `\sp*` / `\quotebox`)
- `references/{input,ledger,router,pedagogy,api,checklist,antipatterns,fallback}.md`;`figure-request.md` / `agents.md` 可先骨架
- `scripts/{preflight.sh,ingest.sh,lint.py,router.py,render_ledger.py,build.sh,test.sh}`
- `examples/excerpt-toy/`(一棵 `--work` 树:`ledger.yaml` + `notes/notes.tex` + `notes/sections/`;**每个 core claim 已有 `\splabel`**;`align` 推导;无 sibling)
- `tests/ledger-valid.yaml`
- `tests/ledger-invalid/{missing-paper-id.yaml,bad-toolkit.yaml,duplicate-id.yaml}`
- `tests/loads-sty/`、`tests/notes-smoke/`(迷你 `--work` 树,不是扁的 `.tex`)、`tests/router_test.py`
- **Lint 码启用:** 仅 `SP001 SP002 SP003`
- **`test.sh` 跑:** `--ledger` 测 schema/id;`lint.py --work tests/loads-sty` 测 `SP003`;`build.sh --work` 编 `notes-smoke` 与 excerpt-toy;再跑 `router_test.py`(含 `check_row` 的 `SP010–SP016` 单测,但 lint 进程不启用这些码)。**不**放跨类未拆分、缺 `\splabel`、缺 request 的 fixture。
- **Dependencies:** 无
- **Description:** `--work` 约定落地。`router.py` **完整**(`classify`/`suggest`/`accepted`/`check_row`),单测钉 class、`SPLIT` 与 include/status 谓词,但 lint 不调用 `check_row`。ingest 支持 `--arxiv/--pdf/--tex/--excerpt`,页渲染改名为 `pg-%03d.png`。preflight 按 §10 分档。`jsonschema` + PyYAML 是依赖。
### PR 2 — 路由 lint 接线 + 嵌入约定 + superfig 委派 example
- **Title:** `superpaper: wire router lint and superfig include convention`
- **Files/components:**
- `references/{figure-request,agents}.md` 写满(§6 / §7)
- `scripts/lint.py`:启用 `SP010 SP011 SP012 SP013 SP014 SP015 SP016`(即开始调用 `check_row`)
- `scripts/screenshot.sh`(`pdftoppm -r 200` 主路径;`crop_bbox` 缺 `magick` 则失败)
- `examples/pipeline-delegate/`:一棵 `--work` 树;§6 的 F1.request(1:1 `pipeline.tex`,`at` 含 `($…$)`)+ 预构建 `notes/figures/F1/build/F1.pdf` + `\spfig` + **`\splabel` 已在**
- `tests/ledger-invalid/{mixed-signals.yaml,toolkit-mismatch.yaml,missing-request.yaml,none-not-dropped.yaml,align-has-include.yaml,included-null-include.yaml,screenshot-bad-include.yaml}`
- `SKILL.md` 补委派步骤
- **不改:** `router.py` 决策表(已在 PR 1)。本 PR 只接线。
- **Lint 码启用:** `SP001–SP003` + `SP010–SP016`
- **Dependencies:** PR 1
- **Description:** 笔记相对路径 + sibling `build.sh` 命令行写进 example README。CI 编笔记,不重跑 superfig。零改 `superfig/`。
### PR 3 — `supertensor` 委派 example
- **Title:** `superpaper: supertensor delegation example`
- **Files/components:**
- `examples/tensor-delegate/`(§6 的 F2 request + 预构建 PDF + `\splabel`)
- `references/figure-request.md` 挂上 supertensor 剖面(若 PR 2 已写入则本 PR 只加 example)
- 可选:再加一个 mixed-signals fixture 若 PR 2 已覆盖则不必
- **新 lint 码:** 无
- **Dependencies:** PR 2
- **Description:** 证明 `axis` 只能配 `supertensor`。零改 `supertensor/`。
### PR 4 — 笔记一致性 lint(符号表、claim 覆盖)
- **Title:** `superpaper: consistency lint for claims and symbols`
- **Files/components:**
- `lint.py`:启用 `SP020 SP021 SP022 SP023 SP024`
- `tests/ledger-invalid/{unlabeled-core.yaml, missing-include.yaml, superderive-v1.yaml, retired-reuse.yaml, dangling-claim.yaml}`
- `references/{checklist,antipatterns}.md` 补漏 claim / 符号漂移
- **禁止:** 回头改 PR 1–3 example 补 `\splabel`(它们出生就必须带)
- **Lint 码启用:** 全表 `SP001–SP024`(v1)
- **Dependencies:** PR 3(链上;技术上只依赖 PR 1 的 example 已有 label,但按链合并以免测试矩阵分叉)
- **Description:** consistency 的机械部分变成 build 失败。符号未使用仍只警告。`render_ledger.py` 若 PR 1 已有则本 PR 不改。
### PR 5 —(可选)superfig argument-map golden(可加性)
- **Title:** `superfig: additive argument-map example using existing primitives`
- **Files/components:**
- **仅** `superfig/examples/argument-map.tex`(+ 其 `build/` 产物若仓库习惯提交)
- `superfig/references/grammar.md` 表格加一行
- 可选:`superfig/SKILL.md`「When to use」加半句
- `superfig/scripts/test.sh` 会自动捡起 `examples/*.tex`(现有通配符)——必须让这张图 lint/build 干净
- **Dependencies:** 无(可与 superpaper PR 并行)。若要在笔记 example 里引用,则依赖 PR 2
- **Description:** 四个 role(`claim`/`assumption`/`lemma`/`evidence`)+ `\sfnode`/`\sfarrow`。不改 `.sty`,不改 lint 规则,不加原语。4 色相预算内。这是本设计允许的唯一 superfig 触碰。
### PR 6 — phase 2:`superderive` 脚手架
- **Title:** `superderive: scaffold stepwise derivation figure toolkit`
- **Files/components:**
- 新建 `superderive/` 整树(见 Layer 2 目录)
- `assets/superderive.sty` 最小:`\sdsetrole`、`\sdstep`、`\sdreason`、`\sdmeaningbox`、cursor 行、palette
- `examples/rewrite-cancel.tex` + `tests/smoke.tex` + lint 负例
- `scripts/{preflight.sh,lint.py,build.sh,test.sh}`(复制 superfig 骨架,不抽库)
- `superpaper/scripts/router.py` **不**改缺省 phase;**不**动 `SP022`
- **Dependencies:** 无硬依赖。建议 PR 1–4 已合并
- **Description:** 证明「取消一项」能被看见。不往 `superfig.sty` 加宏。phase 翻转留给 PR 7。
### PR 7 — router 打开 `superderive`(唯一改缺省 phase 的提交)
- **Title:** `superpaper: route rewrite-figure to superderive`
- **Files/components:**
- `router.py`:`phase` 缺省 2;仍可读 `SUPERPAPER_PHASE`
- `lint.py`:停用 `SP022`
- `examples/derive-delegate/` + `references/router.md`
- **Dependencies:** PR 6、PR 4
- **Description:** 单独一记,回滚路由不必回滚宏包。`suggest()` 仍返回 class `derive`。
---
*本文是实施规格。第一张 PR 按 PR 1 的文件列表脚手架即可开工,不必再做一次产品选型。*