Commit Graph

  • a2c4217dae Document LatentMoE sigmoid routing, sparse dispatch, and K3 block figures main dela 2026-08-26 14:43:58 +08:00
  • ea7167b3f7 Size vocab by max token id so duplicate-piece vocabs (Yi-6B) don't overflow embedding dela 2026-08-26 14:42:12 +08:00
  • 94d0f2ff6a Do not let HF tokenizers truncate wiki articles at 4096 dela 2026-08-26 14:31:07 +08:00
  • 24c9d56b72 Keep attnres on resume, fix final chunk_index, default 0.5b to Yi-6B dela 2026-08-26 14:23:29 +08:00
  • 5cc0555563 Fix chrF fallback so English wiki garbage fails translation_success dela 2026-08-26 14:23:22 +08:00
  • 5a7d949b01 Skip the cold-start SFT best ckpt and free CUDA cache after eval dela 2026-08-26 10:08:31 +08:00
  • 9652a9a7eb Save SFT last/best checkpoints during training and on interrupt dela 2026-08-26 10:00:11 +08:00
  • 071dfaf42c Sanitize SwanLab env before login so 0.9 nested project does not crash dela 2026-08-26 10:00:06 +08:00
  • 9a4862a866 Resume the same SwanLab run from the id stored in the ckpt dela 2026-08-25 22:08:15 +08:00
  • 47c72e5bb8 Warn and reset chunk_index when resume changes batch or seq_len dela 2026-08-25 22:00:28 +08:00
  • e7185cbf49 Pull OPUS-100 en-zh for SFT instead of a checked-in jsonl dela 2026-08-25 21:39:58 +08:00
  • 53d0f4b17a Cut 1B-run I/O: rarer SwanLab, ckpt, and generate dela 2026-08-25 20:59:29 +08:00
  • 8442f92c58 Keep LatentMoE routed bmm in activation dtype under bf16 autocast dela 2026-08-25 20:22:51 +08:00
  • 49aede9cb2 Fit 0.5b training on 32GB: SDPA MLA, block checkpoint, chunked CE dela 2026-08-25 20:09:27 +08:00
  • 7a12f61de1 LatentMoE: K3 sigmoid routing and Switch aux/z-loss dela 2026-08-25 19:50:07 +08:00
  • d1da0816f2 LatentMoE: sparse permute-dispatch + padded bmm dela 2026-08-25 17:47:44 +08:00
  • 584f7e9e73 Initial K3 snapshot: 0.5B KDA/MLA/MoE train path dela 2026-08-25 14:43:17 +08:00