Sacrebleu errors used to fall back to set-overlap unigrams, so any English hyp scored ~70–80 against English refs and SFT reported success 1.0. Use count-based char n-grams, keep BLEU failures from clobbering chrF, and print a few hyps during eval.
Standalone tree split from LLMRL/projects/kda. Includes Triton dt_bias backward fix, train_k3 --preset 0.5b, SFT, Docker runtime, and tests.