Document LatentMoE sigmoid routing, sparse dispatch, and K3 block figures
Ledger C10–C12 match the permute-pad-bmm path and Switch aux/z-loss. Section 8 adds overview and component TikZ; MoE capacity is C_moe so it does not collide with KDA chunk size.
This commit is contained in:
@@ -30,7 +30,7 @@ $n_r$ & routed 专家数 & 16 \\
|
||||
$k$ & Top-$k$ & 2 \\
|
||||
$n_s$ & shared 专家数 & 2 \\
|
||||
$d_{\mathrm{ff}}$ & 专家中间维度 & 96 \\
|
||||
$C$ & MoE 专家容量(pad 宽度) & 动态 \\
|
||||
$C_{\mathrm{moe}}$ & MoE 专家容量(pad 宽度) & 动态 \\
|
||||
$\alpha_{\mathrm{aux}}$ & Switch/GShard aux 系数 & $10^{-2}$ \\
|
||||
$\alpha_z$ & router z-loss 系数 & $10^{-3}$ \\
|
||||
$N$ & AttnRes 原子层数 ($= 2L$) & 8 \\
|
||||
@@ -113,8 +113,8 @@ $s$ & \shape{B, T, n_r} & sigmoid 分数 $\sigma(\mathrm{logits})$ \\
|
||||
$b$ & \shape{n_r} & expert bias(非持久,只进 TopK) \\
|
||||
ids & \shape{B, T, k} & Top-$k$ 专家索引 \\
|
||||
$p_i$ & \shape{B, T, k} & sigmoid-L1 权重 $s_i/\sum_{j\in T}s_j$ \\
|
||||
padded & \shape{n_r, C, \ell} & dispatch 后 pad 到容量 $C$ \\
|
||||
$C$ & 标量 & 最大专家负载(pad 宽度) \\
|
||||
padded & \shape{n_r, C_{\mathrm{moe}}, \ell} & dispatch 后 pad 到容量 $C_{\mathrm{moe}}$ \\
|
||||
$C_{\mathrm{moe}}$ & 标量 & 最大专家负载(pad 宽度) \\
|
||||
$u$ & \shape{B, T, \ell} & routed 加权输出 \\
|
||||
$s_{\mathrm{sh}}$ & \shape{B, T, D} & shared 专家求和 \\
|
||||
$y$ & \shape{B, T, D} & $s_{\mathrm{sh}} + W_\uparrow \mathrm{RMSNorm}(u)$ \\
|
||||
@@ -164,6 +164,11 @@ MLA 解压 & \texttt{'bhtj,hvj->bhtv'} & $\tilde{o}$ \shape{B,H,T,d_v} \\
|
||||
AttnRes 深度打分 & \texttt{'d,nbtd->nbt'} & $s_{l,i}$ \shape{n,B,T} \\
|
||||
AttnRes 深度加权和 & \texttt{'nbt,nbtd->btd'} & $h_l$ \shape{B,T,D} \\
|
||||
AttnRes 批量打分(inter) & \texttt{'qd,nbtd->qnbt'} & logits \shape{S,n,B,T} \\
|
||||
MoE dispatch pad & \texttt{index\_put} & padded \shape{R,C_{\mathrm{moe}},\ell} \\
|
||||
MoE gate 投影(grouped) & \texttt{bmm(padded, w\_g.T)} & $wg$ \shape{R,C_{\mathrm{moe}},ff} \\
|
||||
MoE up 投影(grouped) & \texttt{bmm(padded, w\_u.T)} & $wu$ \shape{R,C_{\mathrm{moe}},ff} \\
|
||||
MoE 输出投影(grouped) & \texttt{bmm(g$\odot$h, w\_o.T)} & out \shape{R,C_{\mathrm{moe}},\ell} \\
|
||||
MoE scatter-add & \texttt{index\_add(0, tok, ...)} & $u$ \shape{N,\ell} \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{center}
|
||||
@@ -177,7 +182,9 @@ AttnRes 批量打分(inter) & \texttt{'qd,nbtd->qnbt'} & logits \shape{S,n,B
|
||||
\item \textbf{分块} = chunk 内下三角解 + chunk 间状态递推,等价于 naive recurrent
|
||||
\item \textbf{GVA} = $H_V = G \cdot H$,forward repeat\_interleave / backward view+sum
|
||||
\item \textbf{MLA} = 低秩 latent + 矩阵吸收,KV cache 从 $2Hd$ 降到 $r$
|
||||
\item \textbf{LatentMoE} = shared 全宽 + routed 半宽 latent + SiTU-GLU 防溢出
|
||||
\item \textbf{LatentMoE} = shared 全宽 + routed 半宽 latent + SiTU-GLU 防溢出;
|
||||
K3 sigmoid-TopK 路由 + 稀疏 permute-dispatch(每 token 只算 $k$ 个专家)+
|
||||
Switch/GShard aux \& z-loss 防塌缩
|
||||
\item \textbf{K3 Hybrid} = 3 KDA + 1 MLA,KDA 提供位置感知
|
||||
\item \textbf{AttnRes} = 深度维 softmax 残差,Block 版把源数压到 $O(N/S)$,
|
||||
两阶段 = inter 批量 + intra online-softmax 合并
|
||||
|
||||
Reference in New Issue
Block a user