# Pedagogy Job: teach the argument thread. The paper is the source of claims, not the outline of the notes. Section order is still motive → idea → mechanism → evidence → takeaway. That order is not a license to recap the PDF in a new sequence. ## Recite vs teach The default failure is a **closed-book failure**: a colleague who has not read this paper still cannot use the core claim, because the notes assume everything the paper assumes. | | Recite | Teach | |---|---|---| | Starts from | the paper's next subsection | what the reader still lacks | | Formula | "本节给出 Eq.(n)" then symbols | a world without the formula, then the formula | | Gap | whatever the paper already wrote | the jump the paper did not write | | Stops when | every cited paragraph is covered | the takeaway is usable | Reader model: a competent colleague in the ambient field, textbook level. They have **not** read this paper, do not know its notation, and will not reconstruct skipped algebra. Do not reteach the ambient textbook (softmax, SGD, …) unless a core claim hangs on a twist. If deleting the PDF would make the section unreadable as a lecture, it was a recap. ## Before each section At the top of every writer `sec-XX.tex` (not `symbols.tex` / generated appendix), answer these four lines **before** the `\section`: ```tex % teach: % gap: <读者此刻还缺什么,不是论文小节标题> % takeaway: <离开本节必须带走的一句> % jump: <论文跳过、读者会卡住的那一步;没有则 none> % omit: <论文这里有、本讲义故意不写的> ``` - `gap` — entering deficit, in the reader's words. - `takeaway` — one sentence. This is also the `本章小结`. - `jump` — the missing why / scale / naive alternative / extreme case. - `omit` — ceremony you will not import "for completeness". Do not start the body until all four are filled. `lint.py` enforces the four keys on every `notes/sections/sec-*.tex` as **`SP025`** (presence only — it cannot judge whether the answers are honest). ## Coverage density `coverage.sections_in` is **cite permission**, not a to-do list. `ledger_ids` on the outline row are the teaching budget. - Outline: fewer lecture rows than paper subsections. Merge until each row has **one** takeaway. A row per paper `\subsection` is too dense. - Writer: teach this row's ids. A paper paragraph that is not required for a `status: core` claim stays out. - Supporting claims: in only if they unblock a core claim. - Dropped claims: never. Default skip (paper ceremony): - "the rest of this paper is organized as follows" - related-work tour (keep the 2–3 contrasts that **define the gap**) - contribution bullets that restate the abstract - dataset / hyperparameter / hardware laundry lists, unless a core empirical claim hangs on a specific number - every ablation, every appendix proof, every lemma not on an `expand: true` derivation - notation paragraphs that only duplicate the symbol appendix ## Intuition before formulas Display math is a three-beat that must not be split: 1. **Intuition in Chinese, with the formula still off-stage.** Use at least one of: analogy, contrast with the naive move, extreme case (0 / 1 / ∞), or tiny numbers. 「下面给出公式」/「本节引入 Eq.(n)」 is not a motive. 2. `\[` or `align` / `aligned`. Never `$$`. 3. Flat symbol list: one item per symbol — 符号 — 含义 — 定义处. The section that owns the core mechanism then puts **one** `importantbox{如果你只记一件事}` with the takeaway sentence. At most one such box per major section. `本章小结` restates that sentence; it is not a new inventory. Pick one intuition tool, not all four. A running numeric example is allowed when scale or cancellation is the point; it is not mandatory. ## Paper jumps Papers skip the why, the scale, and the obvious alternative. Before mechanism, name the stall and fill it: - Why this form, not the naive one? - What happens at 0 / 1 / ∞, or if the term is dropped? - Which quantity is a count vs a rate vs a scale? - Where does the gradient / information / mass actually flow? Fill in prose or `warningbox`. If the paper is silent and you would have to invent a result, say it is silent — do not fabricate. ## Boxes and citations | box | payload | |---|---| | `importantbox` | walk-away: core claim, mechanism, 「如果你只记一件事」 | | `knowledgebox` | not the main thread: prerequisite, analogy, running example | | `warningbox` | naive-vs-correct, hidden assumption, paper-is-silent | | `quotebox` | short quotation + `§` / `Eq.(n)` in the title; ≲ 8 lines; no PDF paste | Routine exposition stays in prose. Figures stay outside every box. Cite `§` / `Eq.(n)` / Figure / Table / page, not timestamps. No `[cite]` placeholders. `\spsource{...}` at the point of use. Front page is a bibliography card (title, authors, year, venue, arXiv), not a PDF-cover screenshot. Default language is Chinese. End every major section with `\subsection{本章小结}`. End the notes with `\section{总结与延伸}` (limitations + compressed takeaways + open questions). Locked first/last titles: 「这篇论文在问什么」…「总结与延伸」, then appendix 符号表 / 推导链一览 / 图表清单. Outline may add or drop middle sections; it may not rename or drop the lock. Suggested skeleton (same lock as `assets/notes-template.tex`): ```tex \section{这篇论文在问什么} % questions[] + motive \section{主张与贡献} % claims[status=core],\splabel{C*} \section{预备:定义、假设、符号} % D* / A* / 符号摘要;不是论文 §2 巡游 % --- outline-chosen mechanism sections --- \section{实验与证据} % only E* that isolate a core claim \section{总结与延伸} ``` ## Antipatterns Each pair is the same lecture beat. Left recites; right teaches. ### 1. Paper order, new wording ```tex % recite 本文第 3 节提出 scaled dot-product attention。 作者将 $Q$、$K$ 做点积,除以 $\sqrt{d_k}$,再 softmax 乘 $V$。 ``` ```tex % teach % gap: 还以为「对齐再加权」就完了,不知道维数一涨会出什么事 读者已经会「用相似度加权」。先看极端情况:$d_k$ 很大时, 点积方差跟着涨,softmax 进饱和,梯度没了。 除以 $\sqrt{d_k}$ 不是新算子,是把方差按回 1。 ``` ### 2. Formula with no off-stage intuition ```tex % recite 本节引入如下公式: \[ \mathrm{Attention}(Q,K,V)=\mathrm{softmax}(QK^{\top}/\sqrt{d_k})V. \] \begin{itemize} \item $Q$ — query \item $K$ — key \item $d_k$ — key 维 \end{itemize} ``` ```tex % teach 两个随机向量的点积有多尖,随维数走。不先压方差,后面的 softmax 是死的。 \[ \mathrm{Attention}(Q,K,V)=\mathrm{softmax}(QK^{\top}/\sqrt{d_k})V. \] \begin{itemize} \item $Q,K,V$ — query / key / value \item $d_k$ — key 维(shape parameter;缩放跟它走) \end{itemize} \begin{importantbox}{如果你只记一件事} $\sqrt{d_k}$ 是改写,不是新注意力。 \end{importantbox} ``` ### 3. Paper jump left as 「如式所示」 ```tex % recite 如 Eq.(3) 所示,取 $A=d/H_{\mathrm{act}}$。 ``` ```tex % teach % jump: 论文直接写下 A,没说为什么不是 1/N 走得越宽,越要把输出缩小,否则残差被宽隐层撑爆。 所以是 $A=d/H_{\mathrm{act}}$:宽 $4096$ 时 $A=1/4$,宽 $1024$ 时 $A=1$。 不是 $1/N$——专家总数还没进这把尺子。 ``` ### 4. Ceremony imported for completeness ```tex % recite 文献三条线:…(半页 related work) 数据集为 ImageNet / CIFAR / …,硬件为 8$\times$A100,超参见表 7。 ``` ```tex % teach % omit: related-work 巡游、数据集表、硬件 现成方法把问题当静态参数;缺的是带延迟和互馈的那一截。 经验数字只在它能隔离这个机制时才进正文。 ``` ### 5. One lecture row per paper paragraph ```tex % recite \subsection{Lemma 2} \subsection{Lemma 3} \subsection{Ablation on dropout} \subsection{Ablation on warmup} ``` ```tex % teach % omit: 不在 DER1 上的 lemma;不能隔离 C1 的 ablation 只展开 DER1 用到的那一步。Ablation 只留「去掉这项,C1 是否还成立」。 ```