We live in a time of incredible model progress, with new frontier open source models being released every few months. At Trajectory, we’re building a platform that lets anyone own their intelligence, so we bring every new model that pushes the Pareto frontier on board as quickly as possible.
We recently did this for GLM-5.3 and GLM-5.3 Flash with the SkyRL team. GLM-5.3 is a leading open source model on par with the closed source frontier; Flash adds the benefit of being incredibly fast.
One analogy we like to use is that each new model is a new expedition. Every expedition moves through the same series of camps, and every model does too: model support, scaling, the training loop, validation, recipes and serving. We've run enough expeditions that the camps have become familiar, even if the terrain never is. Only once every camp is set up can anyone train and use the model.
In this post, we dive deep into one of those camps: the training loop, with a first look at validation. The challenge isn’t just getting things to run, it’s getting them to run fast enough to iterate. At GLM-5.3’s scale, each iteration cycle can take hours, so we explore how we can make that cycle dramatically faster.
Here's the route we'll take:
- What stands in the way.
- Obstacle one: silent mismatches. How we check that the handoff (between Trainer and Sampler) works without waiting for a training run.
- Obstacle two: slow handoffs. How we cut the wait inside every training step from minutes to seconds.
- The first trek. How do we know we’ve built the camp?
What stands in the way
RL training has two jobs. A sampler generates responses, and a trainer learns from them by computing log probabilities and gradients and updating the policy. Each job wants a different tool. The sampler needs fast inference, so we use vLLM. The trainer needs precise, heavily parallel computation, so we use Megatron. That means the loop runs two copies of the same model, one in each system. We train with LoRA, so the base weights stay frozen and only a small adapter changes. After each update, that adapter has to move from the trainer to the sampler before the next round of generation.
Two questions about that handoff kept costing us time.
The first is whether the handoff works at all. Both copies have to behave like the same policy, and each update has to reach the sampler. The end-to-end way to find out is to train and watch the reward. On GLM-5.3, one training cycle took ~150 minutes.
The second is how long the handoff takes. Even when everything is correct, generation can't start until the sampler has the new adapter. On full GLM-5.3, that took more than six minutes.
We built a tool for each. Our five minute test answers the first question, and an in-memory sync shortens the second. We'll start with the test.
Obstacle one: silent mismatches
Megatron and vLLM implement the same architecture, load the same weights and produce reasonable answers on their own. But despite this, they can still run different policies. Why? The two systems split the model across GPUs differently, use different kernels and do arithmetic in a different order. Some of the resulting difference is floating-point noise. But sometimes, it’s a bug.
What makes these bugs insidious is they tend not to crash anything—at least not right away. A flat reward curve could come from the task, the recipe or the adapter implementation. Waiting hours for a curve is slow, and the answer is ambiguous when it arrives. We needed a check that tested the handoff directly. So we built one.
It shows the same response tokens to both copies and compares their log probabilities token by token. Agreement within tolerance gives us a direct check of policy consistency on those tokens; it does not prove equivalence on every possible input.
But that alone isn't enough. With a zero-initialized adapter, both copies can agree even if the synchronization path is broken. A check of that initial state alone would pass on exactly the bug it should catch. So our numerical regression test breaks agreement on purpose. It runs five steps on a fixed set of tokens:
- Check that the trainer and sampler agree.
- Perturb only the trainer's adapter, leaving the sampler unchanged.
- Check that the sampler now disagrees, since it still has the old adapter.
- Send the new adapter to the sampler.
- Check that the two agree again.
In pseudocode:
assert_logprobs_match(trainer, sampler, tokens, mean_tol=0.05) # 1. same starting policy
perturb_trainer_lora_weights(trainer) # 2. change only the trainer
assert_stale_sampler_differs(trainer, sampler, tokens, mean_error_above=0.05) # 3. sampler should now be stale
publish_updated_lora_to_sampler(trainer, sampler) # 4. send the update
assert_logprobs_match(trainer, sampler, tokens, mean_tol=0.05) # 5. agreement returns
Step 3, the deliberate disagreement, is what makes the test especially meaningful. If the sampler disagrees after we change only the trainer, the change was large enough to detect. So when the two agree again at step 5, we know that the update actually arrived.
To speed up even more, the check skips the optimizer and uses short responses. Our full GLM-5.3 check finished in 5.24 minutes after initialization, compared with the ~150 minutes for a training cycle, about 28× faster feedback.
A question that used to cost a training cycle now takes about five minutes. That's the first tool for setting up this camp. The second deals with the wait inside every training step.
Obstacle two: slow handoffs
With the test in place, we could trust the handoff (i.e, the numerics check out!), and that training worked well. The next problem at this camp? The wait for weight synchronization between the trainer and sampler.
Much of our training is done via LoRAs. LoRA changes a small part of the model, so theoretically, sending the adapter should be fast. On full GLM-5.3, the file-based path took more than six minutes.
The problem was duplication. GLM-5.3 is a mixture-of-experts model with a shared expert, and the export repeated the shared adapter under many expert names. The naive implementation wrote all of those tensors to disk, and the inference workers read them back, compounding the time quite significantly.
We implemented an in-memory path that finds duplicate tensors and sends each unique tensor only once. Inference workers rebuild the adapter names on their side. Separate trainer and inference workers use NCCL; workers that share GPUs use CUDA IPC. That allows us to avoid writing to the disk entirely.
With this implementation, mean synchronization time fell from about 390 seconds to 13 seconds, a 29.4× speedup.
Each training step now waits seconds for the update instead of minutes. With both tools in place, our camp is almost ready. Now we just need to figure out if it can stand on its own.
The first trek
A camp is only useful if the expedition can actually rely on it. To test it, we trained both models on DAPO-math with a deliberately tight 8,192-token response limit, then evaluated on AIME 2024.
This token limit makes this a real task even for strong models. At baseline, 46.11% of Flash's AIME responses hit the cap. Many of those answers weren't wrong so much as unfinished, and an unfinished solution gets no credit. To improve, the models had to solve more problems within the budget.
Over 20 training steps, AIME accuracy rose from 57% to 81% for GLM-5.3. For Flash, from 54% to 83%.
The models also learned to use less of their token budget. For Flash, AIME response length fell from 5,217 to 2,424 tokens, and the share of responses hitting the cap fell from 46% to just 6%.
These results show better accuracy within 8K tokens. Shorter responses also mean less decoding, and a promising thing to measure would be the impact on end-to-end latency.
Both models improved on this task within the fixed response budget.
Try it
The model integrations, regression tests and synchronization work have been upstreamed to SkyRL. Start with the full-model recipe or the two-node Flash recipe. Data and methods contains the run settings, plotted values and measurement details.
If you're training frontier scale, foundation models for your own workloads, feel free to reach out at hello@trajectory.ai. To work on the infrastructure behind these experiments, write to hiring@trajectory.ai.
Here’s to many more expeditions 🧭
Acknowledgements
We'd like to thank Avi Basnet for his help and collaboration on the weight-synchronization work; Eric Tang and Sumanth R Hegde for their work enabling GLM-5.3 and GLM-5.3 Flash; and Philipp Moritz for his support of the Trajectory x SkyRL collaboration.
