Sacrebleu errors used to fall back to set-overlap unigrams, so any English hyp scored ~70–80 against English refs and SFT reported success 1.0. Use count-based char n-grams, keep BLEU failures from clobbering chrF, and print a few hyps during eval.
Sacrebleu errors used to fall back to set-overlap unigrams, so any English hyp scored ~70–80 against English refs and SFT reported success 1.0. Use count-based char n-grams, keep BLEU failures from clobbering chrF, and print a few hyps during eval.