Step 0 generate was writing a 5GB success=0 snapshot and leaving the
32GB card fragmented, so the next Adam step OOM'd after batch-32 eval.
Only promote _best after step 0 and empty_cache when eval returns.
train_sft used to torch.save only after the full epoch budget, so Ctrl+C
dropped all translation weights. Write _last every --ckpt-every steps
and on KeyboardInterrupt; write _best when frozen eval (success, chrF)
improves; --resume continues from _last.
Route with σ(W_r x), Top-k(s+b), then L1-normalize over the selected set.
Add Switch/GShard aux and router z-loss into train_k3 and train_sft.
Wiki parquet URLs honor HF_ENDPOINT for mirrored downloads.