Across a 20-step run, that’s about 2 hours of waiting cut to about 4 minutes, assuming one sync per step.
Mean synchronization time fell from 389.73 seconds to 13.26 seconds. Both lanes share a clock accelerated 40×; the repeated transfers illustrate the ratio, rather than a measured sequence of training updates.
Measurement. Three trials per route on two B300 nodes, with full GLM-5.3-BF16, separate trainer and inference workers, trainer TP8/EP8, vLLM TP8 and rank-32 LoRA with shared-expert adapters. Tensor copies are schematic.