OpenBayes sets SWANLAB_PROJECT as a string; swanlab>=0.9 parses that as
ProjectSettings and raises QuoteAwareEnvSettingsSource. Drop it, keep
SWANLAB_PROJ_NAME, and share run-id extraction with train_k3.
swanlab.init always opened a new experiment on --resume. Save the run
id in the checkpoint and pass resume=True, id=... on the next start.
--swanlab-id overrides; --swanlab-new forces a fresh experiment.
Every micro-step was hitting SwanLab, and every 100 steps wrote a 5GB
ckpt plus greedy decode. 0.5b now logs every 20, eval/held-out every 500,
saves _last every 1000, samples every 2000.
Whole-mixer checkpoint plus T×T MLA scores OOM'd a 31GB GPU on backward.
Checkpoint each AttnRes block, run absorbed MLA through SDPA, and compute
CE in vocab chunks so [B,T,V] logits are never materialized.
--max-tokens is now the training budget; default --steps 2000 no longer
caps a 1B-token run at 250 optimizer steps.
Route with σ(W_r x), Top-k(s+b), then L1-normalize over the selected set.
Add Switch/GShard aux and router z-loss into train_k3 and train_sft.
Wiki parquet URLs honor HF_ENDPOINT for mirrored downloads.