Before training, about half of Flash’s answers ran out of room. By step 20 it was about one in sixteen, and the average answer had gone from 5,217 tokens to 2,424.
AIME 2024 accuracy over 20 training steps, evaluated every five steps. Each dot represents one Flash response; the number reaching the 8,192-token cap fell from 166 to 22 out of 360.
Evaluation. 30 problems, with 32 responses per problem for GLM-5.3 and 12 for Flash: 960 and 360 responses per checkpoint, respectively. All responses count, including capped responses, which score zero. Each model had its own run and recipe.