AWS has published measurements and configuration guidance for reinforcement learning on Mixture-of-Experts models running across its own network fabric, in a post dated 25 September. Its headline result: across 48 P5en instances, 16 dedicated to training and 32 to inference, running what it calls a super-sparse MoE model, enabling DeepEP over Elastic Fabric Adapter increased aggregate rollout throughput by 40 percent.
The more durable part is the engineering underneath. DeepEP, the expert-parallel communication library released by DeepSeek, was written against a CUDA-specific RDMA backend. Amazon says its upstream contributions move those primitives to libfabric, making the transport portable across EFA-supported configurations and reducing per-message overhead for the sparse, fine-grained traffic that expert parallelism generates.
Why it matters: sparse MoE models route tokens to experts sitting on different machines, so post-training them at scale is a networking problem as much as a compute one. That has tended to favour whichever interconnect handles small, irregular messages best. A hyperscaler contributing open-source work to a Chinese lab’s library in order to make its own fabric competitive is a fair indication of where the bottleneck now sits.
AWS files the post as a technical how-to rather than an announcement, and the figure arrives with named software versions on both sides of the comparison, which is more disclosure than most vendor benchmarks carry.
