5 Commits
Author SHA1 Message Date
dela ea7167b3f7 Size vocab by max token id so duplicate-piece vocabs (Yi-6B) don't overflow embedding 2026-08-26 14:42:12 +08:00
dela 94d0f2ff6a Do not let HF tokenizers truncate wiki articles at 4096
Yi-6B sets model_max_length=4096. encode() would clip long Wikipedia
pages before we pack seq_len chunks. Raise the cap so only our
chunker limits context.
2026-08-26 14:31:07 +08:00
dela e7185cbf49 Pull OPUS-100 en-zh for SFT instead of a checked-in jsonl
train_sft --data opus-100 streams Helsinki-NLP/opus-100, writes both
directions, and skips frozen eval sentences. Runtime cache stays under
data/sft/ (gitignored).
2026-08-25 21:39:58 +08:00
dela 7a12f61de1 LatentMoE: K3 sigmoid routing and Switch aux/z-loss
Route with σ(W_r x), Top-k(s+b), then L1-normalize over the selected set.
Add Switch/GShard aux and router z-loss into train_k3 and train_sft.
Wiki parquet URLs honor HF_ENDPOINT for mirrored downloads.
2026-08-25 19:50:07 +08:00
dela 584f7e9e73 Initial K3 snapshot: 0.5B KDA/MLA/MoE train path
Standalone tree split from LLMRL/projects/kda. Includes Triton dt_bias
backward fix, train_k3 --preset 0.5b, SFT, Docker runtime, and tests.
2026-08-25 14:43:17 +08:00