feat: land superpaper v1 notes scaffold
Add the ledger schema, class router, lint codes, ingest/build pipeline, and three work-tree examples: align derivation, superfig delegation, and supertensor delegation.
This commit is contained in:
@@ -0,0 +1,3 @@
|
||||
\section{这篇论文在问什么}
|
||||
要把输入变成预测,最简单的机制是什么?损失要不要走在前向主路上?
|
||||
这是摘录 fixture,不假装读完全文。
|
||||
@@ -0,0 +1,3 @@
|
||||
\section{主张与贡献}
|
||||
\splabel{C1}
|
||||
一次前向是 $x \to f_\theta \to \hat y$。损失 $L(\hat y,y)$ 在预测之后单独计算,不是主路上的一站。
|
||||
@@ -0,0 +1,2 @@
|
||||
\section{预备:定义、假设、符号}
|
||||
预测器定义为 $\hat y = f_\theta(x)$。符号见附录。
|
||||
@@ -0,0 +1,18 @@
|
||||
\section{一次前向与损失}
|
||||
内积的方差会随维数 $d$ 涨。为了不让后续非线性饱和,要把点积除掉 $\sqrt{d}$。
|
||||
|
||||
\begin{align}
|
||||
u^\top v &\longrightarrow \frac{u^\top v}{\sqrt{d}}.
|
||||
\end{align}
|
||||
|
||||
\begin{itemize}
|
||||
\item $u,v$ — 两个 $d$ 维向量
|
||||
\item $d$ — 特征维(shape parameter)
|
||||
\end{itemize}
|
||||
|
||||
\begin{importantbox}{主路与损失}
|
||||
损失比较 $\hat y$ 与 $y$,梯度再回到 $\theta$。不要把 $L$ 画成前向的一站。
|
||||
\end{importantbox}
|
||||
|
||||
\subsection{本章小结}
|
||||
前向只负责预测;缩放是改写,不是新算子。
|
||||
@@ -0,0 +1,2 @@
|
||||
\section{总结与延伸}
|
||||
摘录只保留一条机制:一次前向加侧路损失。更长的论文用 ledger 把 claim 钉住,再按路由出图。
|
||||
Reference in New Issue
Block a user