-
a2c4217dae
Document LatentMoE sigmoid routing, sparse dispatch, and K3 block figures
main
dela
2026-08-26 14:43:58 +08:00
-
ea7167b3f7
Size vocab by max token id so duplicate-piece vocabs (Yi-6B) don't overflow embedding
dela
2026-08-26 14:42:12 +08:00
-
94d0f2ff6a
Do not let HF tokenizers truncate wiki articles at 4096
dela
2026-08-26 14:31:07 +08:00
-
24c9d56b72
Keep attnres on resume, fix final chunk_index, default 0.5b to Yi-6B
dela
2026-08-26 14:23:29 +08:00
-
5cc0555563
Fix chrF fallback so English wiki garbage fails translation_success
dela
2026-08-26 14:23:22 +08:00
-
5a7d949b01
Skip the cold-start SFT best ckpt and free CUDA cache after eval
dela
2026-08-26 10:08:31 +08:00
-
9652a9a7eb
Save SFT last/best checkpoints during training and on interrupt
dela
2026-08-26 10:00:11 +08:00
-
071dfaf42c
Sanitize SwanLab env before login so 0.9 nested project does not crash
dela
2026-08-26 10:00:06 +08:00
-
9a4862a866
Resume the same SwanLab run from the id stored in the ckpt
dela
2026-08-25 22:08:15 +08:00
-
47c72e5bb8
Warn and reset chunk_index when resume changes batch or seq_len
dela
2026-08-25 22:00:28 +08:00
-
e7185cbf49
Pull OPUS-100 en-zh for SFT instead of a checked-in jsonl
dela
2026-08-25 21:39:58 +08:00
-
53d0f4b17a
Cut 1B-run I/O: rarer SwanLab, ckpt, and generate
dela
2026-08-25 20:59:29 +08:00
-
8442f92c58
Keep LatentMoE routed bmm in activation dtype under bf16 autocast
dela
2026-08-25 20:22:51 +08:00
-
49aede9cb2
Fit 0.5b training on 32GB: SDPA MLA, block checkpoint, chunked CE
dela
2026-08-25 20:09:27 +08:00
-
7a12f61de1
LatentMoE: K3 sigmoid routing and Switch aux/z-loss
dela
2026-08-25 19:50:07 +08:00
-
d1da0816f2
LatentMoE: sparse permute-dispatch + padded bmm
dela
2026-08-25 17:47:44 +08:00
-
584f7e9e73
Initial K3 snapshot: 0.5B KDA/MLA/MoE train path
dela
2026-08-25 14:43:17 +08:00