K2-Horizon separates token preference from stopping behavior.
Midtraining and SFT checkpoints already support the teacher-preferred token and remain relatively stable. A pretrained student can learn that surface token, yet later lose its tendency to stop.
With semantic EOS alignment, the pretrained student still shows initial length growth, an elevated plateau, and late-stage re-inflation. The plateau is not recovery to the teacher’s response length, and surface-token identity alone cannot explain the renewed collapse.
Pretrain → Final: 400 steps. Midtrain/SFT → Final: 200 steps. Labels 1–3 mark initial growth, an intermediate plateau, and late-stage re-inflation; gray dashed lines indicate approximate phase boundaries.





