Rewrite references/pedagogy.md around recite-vs-teach: reader model,
coverage density (sections_in is cite permission, not a to-do list),
intuition-before-formula three-beat, and paper-jump filling.
Every notes/sections/sec-*.tex must now answer gap / takeaway / jump /
omit above the first \section. lint.py checks presence only (it cannot
judge honesty); files opening with "% generated by" are exempt.
- scripts/lint.py: SP025 + is_writer_section / teach_gaps helpers
- tests/no-teach-block/: fixture with takeaway only, wired into test.sh
- examples/*: all 14 writer sections get real teach blocks
- SKILL.md, references/{agents,antipatterns,checklist}.md, DESIGN.md,
assets/notes-template.tex: route writers and consistency agent
through pedagogy.md
24 lines
868 B
TeX
24 lines
868 B
TeX
% teach:
|
||
% gap: 会算内积,但不知道维数一涨为什么会把后面的非线性弄坏
|
||
% takeaway: 除 $\sqrt{d}$ 是把方差按回 1 的改写,不是新算子
|
||
% jump: 摘录直接写下 $1/\sqrt{d}$,没说维数涨会让点积方差跟着涨
|
||
% omit: 数据集、超参、硬件
|
||
\section{一次前向与损失}
|
||
内积的方差会随维数 $d$ 涨。为了不让后续非线性饱和,要把点积除掉 $\sqrt{d}$。
|
||
|
||
\begin{align}
|
||
u^\top v &\longrightarrow \frac{u^\top v}{\sqrt{d}}.
|
||
\end{align}
|
||
|
||
\begin{itemize}
|
||
\item $u,v$ — 两个 $d$ 维向量
|
||
\item $d$ — 特征维(shape parameter)
|
||
\end{itemize}
|
||
|
||
\begin{importantbox}{主路与损失}
|
||
损失比较 $\hat y$ 与 $y$,梯度再回到 $\theta$。不要把 $L$ 画成前向的一站。
|
||
\end{importantbox}
|
||
|
||
\subsection{本章小结}
|
||
前向只负责预测;缩放是改写,不是新算子。
|