Document LatentMoE sigmoid routing, sparse dispatch, and K3 block figures

Ledger C10–C12 match the permute-pad-bmm path and Switch aux/z-loss.
Section 8 adds overview and component TikZ; MoE capacity is C_moe so it
does not collide with KDA chunk size.
This commit is contained in:
dela
2026-08-26 14:43:58 +08:00
parent ea7167b3f7
commit a2c4217dae
6 changed files with 239 additions and 18 deletions
+11 -4
View File
@@ -30,7 +30,7 @@ $n_r$ & routed 专家数 & 16 \\
$k$ & Top-$k$ & 2 \\
$n_s$ & shared 专家数 & 2 \\
$d_{\mathrm{ff}}$ & 专家中间维度 & 96 \\
$C$ & MoE 专家容量(pad 宽度) & 动态 \\
$C_{\mathrm{moe}}$ & MoE 专家容量(pad 宽度) & 动态 \\
$\alpha_{\mathrm{aux}}$ & Switch/GShard aux 系数 & $10^{-2}$ \\
$\alpha_z$ & router z-loss 系数 & $10^{-3}$ \\
$N$ & AttnRes 原子层数 ($= 2L$) & 8 \\
@@ -113,8 +113,8 @@ $s$ & \shape{B, T, n_r} & sigmoid 分数 $\sigma(\mathrm{logits})$ \\
$b$ & \shape{n_r} & expert bias(非持久,只进 TopK) \\
ids & \shape{B, T, k} & Top-$k$ 专家索引 \\
$p_i$ & \shape{B, T, k} & sigmoid-L1 权重 $s_i/\sum_{j\in T}s_j$ \\
padded & \shape{n_r, C, \ell} & dispatch 后 pad 到容量 $C$ \\
$C$ & 标量 & 最大专家负载(pad 宽度) \\
padded & \shape{n_r, C_{\mathrm{moe}}, \ell} & dispatch 后 pad 到容量 $C_{\mathrm{moe}}$ \\
$C_{\mathrm{moe}}$ & 标量 & 最大专家负载(pad 宽度) \\
$u$ & \shape{B, T, \ell} & routed 加权输出 \\
$s_{\mathrm{sh}}$ & \shape{B, T, D} & shared 专家求和 \\
$y$ & \shape{B, T, D} & $s_{\mathrm{sh}} + W_\uparrow \mathrm{RMSNorm}(u)$ \\
@@ -164,6 +164,11 @@ MLA 解压 & \texttt{'bhtj,hvj->bhtv'} & $\tilde{o}$ \shape{B,H,T,d_v} \\
AttnRes 深度打分 & \texttt{'d,nbtd->nbt'} & $s_{l,i}$ \shape{n,B,T} \\
AttnRes 深度加权和 & \texttt{'nbt,nbtd->btd'} & $h_l$ \shape{B,T,D} \\
AttnRes 批量打分(inter) & \texttt{'qd,nbtd->qnbt'} & logits \shape{S,n,B,T} \\
MoE dispatch pad & \texttt{index\_put} & padded \shape{R,C_{\mathrm{moe}},\ell} \\
MoE gate 投影(grouped) & \texttt{bmm(padded, w\_g.T)} & $wg$ \shape{R,C_{\mathrm{moe}},ff} \\
MoE up 投影(grouped) & \texttt{bmm(padded, w\_u.T)} & $wu$ \shape{R,C_{\mathrm{moe}},ff} \\
MoE 输出投影(grouped) & \texttt{bmm(g$\odot$h, w\_o.T)} & out \shape{R,C_{\mathrm{moe}},\ell} \\
MoE scatter-add & \texttt{index\_add(0, tok, ...)} & $u$ \shape{N,\ell} \\
\bottomrule
\end{tabular}
\end{center}
@@ -177,7 +182,9 @@ AttnRes 批量打分(inter) & \texttt{'qd,nbtd->qnbt'} & logits \shape{S,n,B
\item \textbf{分块} = chunk 内下三角解 + chunk 间状态递推,等价于 naive recurrent
\item \textbf{GVA} = $H_V = G \cdot H$,forward repeat\_interleave / backward view+sum
\item \textbf{MLA} = 低秩 latent + 矩阵吸收,KV cache 从 $2Hd$ 降到 $r$
\item \textbf{LatentMoE} = shared 全宽 + routed 半宽 latent + SiTU-GLU 防溢出
\item \textbf{LatentMoE} = shared 全宽 + routed 半宽 latent + SiTU-GLU 防溢出;
K3 sigmoid-TopK 路由 + 稀疏 permute-dispatch(每 token 只算 $k$ 个专家)+
Switch/GShard aux \& z-loss 防塌缩
\item \textbf{K3 Hybrid} = 3 KDA + 1 MLA,KDA 提供位置感知
\item \textbf{AttnRes} = 深度维 softmax 残差,Block 版把源数压到 $O(N/S)$,
两阶段 = inter 批量 + intra online-softmax 合并