InfrastructureCompute, chips & cloud
NVIDIA reports optimized JAX stack for dropless MoE training with Transformer Engine
NVIDIA says Transformer Engine and MaxText now provide a faster dropless MoE training path in JAX, combining grouped GEMM, expert-parallel communication, MXFP8 quantization, host offloading and XLA scheduling. The headline speedups are vendor-reported and not independently reproduced.
NVIDIA reports a DeepSeek-V3-scale dropless MoE training improvement from 103 to 1,068 TFLOPS per GPU using its JAX and Transformer Engine stack; the reviewed research treats this as vendor-reported only. [1]
The implementation combines grouped GEMM for ragged expert batches, expert-parallel dispatch and combine, MXFP8 quantization, host offloading and XLA collective scheduling rather than a single isolated optimization. [1] [6] [7]
NVIDIA has published a configuration-oriented recipe for dropless MoE training in JAX, and prior research independently supports the view that routing, all-to-all communication and ragged expert computation are central MoE bottlenecks.
Read the full assessment
Implications: for AI teams training large sparse models on NVIDIA systems, software stack choices may materially affect accelerator utilization and cost. For business leaders, the takeaway is narrower than a general 10× claim: integrated training infrastructure can be as strategic as hardware procurement.
Executive brief
On September 14, 2026, NVIDIA published a primary technical blog describing an optimized dropless Mixture-of-Experts training path in JAX using NVIDIA Transformer Engine, MaxText, grouped GEMM, MXFP8 quantization, expert-parallel dispatch/combine kernels, JAX host offloading, and XLA collective scheduling. NVIDIA reports a large performance improvement for DeepSeek-V3-scale MoE training: an “unoptimized baseline” at 103 TFLOPS/GPU rising to 1,068 TFLOPS/GPU, or 10.4×, with inter-GPU communication originally consuming 84% of accumulated kernel time. The most important caveat is that the public evidence is largely NVIDIA-controlled: the blog, Transformer Engine release notes, MaxText configuration documentation, and NVIDIA/JAX Toolbox guidance.
Read the full section
On September 14, 2026, NVIDIA published a primary technical blog describing an optimized dropless Mixture-of-Experts training path in JAX using NVIDIA Transformer Engine, MaxText, grouped GEMM, MXFP8 quantization, expert-parallel dispatch/combine kernels, JAX host offloading, and XLA collective scheduling. NVIDIA reports a large performance improvement for DeepSeek-V3-scale MoE training: an “unoptimized baseline” at 103 TFLOPS/GPU rising to 1,068 TFLOPS/GPU, or 10.4×, with inter-GPU communication originally consuming 84% of accumulated kernel time. Treat these numbers as vendor-reported, not independently validated. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
The engineering claim is plausible in direction: dropless MoE creates irregular expert token counts, which makes naive all-to-all and per-expert GEMMs inefficient; prior systems work such as MegaBlocks also identified the token-dropping/padding tradeoff and proposed sparse kernels to avoid it. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts NVIDIA’s specific JAX stack, however, has not yet surfaced in independent reporting or third-party reproduction in the searches conducted for this dossier. The most important caveat is that the public evidence is largely NVIDIA-controlled: the blog, Transformer Engine release notes, MaxText configuration documentation, and NVIDIA/JAX Toolbox guidance.
For practitioners, the near-term implication is not “JAX MoE is now automatically 10× faster.” It is narrower: if you are training very large MoE models on recent NVIDIA systems, and you can adopt the NVIDIA MaxText container and Transformer Engine TE MoEBlock path, NVIDIA now documents a reproducible configuration and a performance-oriented recipe. The business implication is that software stack maturity—not only GPU count—remains a central determinant of MoE training economics.
What changed and event timeline
- Publication event NVIDIA’s blog, titled “Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine,” is dated September 14, 2026 and authored by Seonghee Lee, Jeremy Berchtold, Phuong Nguyen, Teddy Do, and Tejash Shah.
- What NVIDIA says changed The article says NVIDIA combined multiple optimizations for dropless MoE in JAX: cuBLAS GroupedGEMM, XLA multistream collectives, MXFP8 GroupQuant, host activation offloading, and optimized expert-parallel dispatch/combine via Transformer Engine/NCCL EP.
- Availability timeline NVIDIA says the optimizations ship in an NGC MaxText container with Transformer Engine built in, and instructs users to use
ghcr.io/nvidia/jax:maxtext-2026-09-09or newer. - Important evidence gap Searches found NVIDIA’s article and related official documentation, plus prior MoE systems papers, but did not find independent third-party reproduction or independent reporting of the specific September 14 NVIDIA JAX/Transformer Engine result.
Capabilities and access
The documented access path is through MaxText plus NVIDIA Transformer Engine, not a new foundation model release. The MaxText documentation says this requires sparse_matmul=True, prefuse_moe_weights=True, and Transformer Engine JAX with expert-parallel MoE support, listed as TE 2.19+.
Read the full section
The documented access path is through MaxText plus NVIDIA Transformer Engine, not a new foundation model release. The relevant user-facing switch is te_moe_block: true, which MaxText documentation describes as using Transformer Engine’s fused expert-parallel MoEBlock for routing, dispatch, grouped GEMMs, and combining expert outputs. The MaxText documentation says this requires sparse_matmul=True, prefuse_moe_weights=True, and Transformer Engine JAX with expert-parallel MoE support, listed as TE 2.19+. maxtext/docs/reference/core_concepts/moe_configuration.md at main · AI-Hypercomputer/maxtext · GitHub
The NVIDIA blog’s minimal MaxText flags are:
te_moe_block: true
te_gmm_quantization: "te_mxfp8"
ragged_buffer_factor: 2.0
te_ep_overflow_check_every_n_steps: 20
sparse_matmul: true
prefuse_moe_weights: true
NVIDIA says the exact DeepSeek-V3 reproduction path additionally uses model_name: "deepseek3-671b", max_target_length: 4096, per_device_batch_size: 6, steps: 15, attention: "cudnn_flash_te", quantization: "te_fp8_currentscaling", te_gmm_quantization: "te_mxfp8", a custom rematerialization/offload policy, and a 128-GPU parallelism layout with FSDP and expert parallelism. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
A notable documentation wrinkle: the Transformer Engine 2.19 API docs expose JAX mesh concepts including an ep_resource axis for expert parallelism, but the specific “JAX: Expert Parallelism with TransformerEngine” tutorial page currently says “TODO — Coming soon.” Jax — Transformer Engine 2.19.0 JAX: Expert Parallelism with TransformerEngine — Transformer Engine 2.19.0 For developers, that means the capability appears to exist in the stack, but the public tutorial documentation is incomplete as of retrieval.
Technical analysis for researchers and developers
In a standard dense Transformer feed-forward block, every token uses the same MLP weights. In MoE, a router sends each token to one or more expert MLPs. Switch Transformer described MoE as selecting different parameters per input, yielding sparse activation but also noted that MoE adoption is hindered by complexity, communication cost, and training instability.
Read the full section
Architecture: why dropless MoE is hard
In a standard dense Transformer feed-forward block, every token uses the same MLP weights. In MoE, a router sends each token to one or more expert MLPs. That gives conditional computation: more total parameters can be available without activating all parameters for every token. Switch Transformer described MoE as selecting different parameters per input, yielding sparse activation but also noted that MoE adoption is hindered by complexity, communication cost, and training instability. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
The hard systems problem is that expert assignment is data-dependent. Each expert may receive a different number of tokens per step, creating ragged expert batches rather than uniform rectangular matrices. NVIDIA frames the critical bottlenecks as token routing, expert dispatch/gather, all-to-all communication, and ragged expert GEMMs. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
Capacity-based MoE avoids some of this irregularity by imposing a fixed per-expert capacity. Overflow tokens may be dropped or padding may be used to preserve regular shapes. MegaBlocks identified the same core tradeoff: existing frameworks restricted dynamic routing, forcing users to choose between dropping tokens or wasting compute/memory on padding. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Optimization 1: grouped GEMM / ragged dot
NVIDIA’s first core optimization is grouped GEMM: rather than running one GEMM per expert or padding every expert to a worst-case capacity, grouped GEMM handles all expert matmuls in one call using each expert’s actual token count. NVIDIA says Transformer Engine’s grouped_gemm / ragged_dot path uses cuBLAS and cuBLASLt and enables MXFP8 block scaling on Blackwell GPUs. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
This is conceptually aligned with prior dropless MoE work, but the implementation path differs. MegaBlocks reformulated MoE computation as block-sparse matrix multiplication and built block-sparse GPU kernels; the NVIDIA JAX path emphasizes grouped GEMM, ragged buffers, MXFP8 quantization, and Tensor Core utilization through NVIDIA libraries. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Optimization 2: expert-parallel dispatch and combine
MoE expert parallelism requires moving tokens to the devices hosting their selected experts, running the expert MLPs, and combining results back into the original token order. NVIDIA says Transformer Engine integrates dispatch and combine into a fused kernel path powered by NCCL EP, including token deduplication so a token routed to multiple experts on the same rank or remote InfiniBand node traverses the network once and is replicated on receipt. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
This targets the communication-heavy part of MoE. Prior work such as Tutel also focused on MoE all-to-all and dynamic workload handling, including adaptive pipelining and a two-dimensional hierarchical all-to-all algorithm. [](https://proceedings.mlsys.org/paper_files/paper/2023/file/5616d34cf8ff73942cfd5aa922842556-Paper-mlsys2023.pdf) The repeated pattern across systems papers is that MoE performance is rarely about matmul alone; routing, packing, communication, and load imbalance are first-order constraints.
Optimization 3: host offloading and XLA collectives
NVIDIA says its DeepSeek-V3 configuration uses host offloading for selected activations and XLA multistream collectives. In the blog, NVIDIA specifically says intermediate activations need not remain on device for the entire forward pass, and mentions offloading query and value projection results for DSv3 training. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog NVIDIA’s separate host-offloading blog reports that QKV activation offloading plus the Latency Hiding Scheduler improved throughput for Llama 3.1 405B, but that is a different model and should be treated as supporting context, not proof of the DeepSeek-V3 MoE result. Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading | NVIDIA Technical Blog
OpenXLA’s own flags guidance says xla_gpu_enable_latency_hiding_scheduler enables schedulers to overlap asynchronous communication with computation, and that optimization level -O1 enables advanced GPU passes including collective pipelining and latency hiding scheduling. XLA Flags Guidance | OpenXLA Project NVIDIA/JAX Toolbox documentation similarly says LHS is now enabled by O1 and documents XLA flags for asynchronous collectives. JAX-Toolbox/docs/GPU_performance.md at main · NVIDIA/JAX-Toolbox · GitHub
Evaluation methodology and reproducibility
The public method is configuration-centric, not a full benchmark artifact. NVIDIA provides a DeepSeek-V3-specific MaxText configuration, XLA flags, and environment variables, and says different models require different tuning. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog The disclosed run uses steps: 15, which is enough to reproduce a short performance experiment but not, by itself, evidence of long-horizon training stability, convergence, or final model quality. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
There is also a hardware-label ambiguity in the public text: the article body reports the 103-to-1,068 TFLOPS/GPU result for NVIDIA GB200, while later figure captions describe end-to-end DeepSeek-V3 671B performance and scaling on GB300 NVL72 hardware. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog A careful reproduction report should specify exact GPU type, rack topology, CUDA/cuDNN/NCCL versions, Transformer Engine commit/version, MaxText commit, XLA build, fabric configuration, and measurement window.
Claims and evidence
- NVIDIA published the article on September 14, 2026. — Supported by NVIDIA primary source.
- Transformer Engine with JAX improved DeepSeek-V3 MoE training from 103 to 1,068 TFLOPS/GPU.
- The stack sustains 97% scaling efficiency at 1,024 GPUs.
Read the full section
| Material claim | Evidence status |
| NVIDIA published the article on September 14, 2026. | Supported by NVIDIA primary source. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog |
| Transformer Engine with JAX improved DeepSeek-V3 MoE training from 103 to 1,068 TFLOPS/GPU. | Vendor-reported only; no independent reproduction found. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog |
| The stack sustains 97% scaling efficiency at 1,024 GPUs. | Vendor-reported only; source says GB300 NVL72 and scaling relative to 128-GPU baseline. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog |
| Dropless MoE avoids token dropping/padding tradeoffs. | Supported by NVIDIA description and independently consistent with MegaBlocks paper. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog MegaBlocks: Efficient Sparse Training with Mixture-of-Experts |
| MaxText exposes a TE MoEBlock requiring TE 2.19+ expert-parallel support. | Supported by MaxText documentation. maxtext/docs/reference/core_concepts/moe_configuration.md at main · AI-Hypercomputer/maxtext · GitHub |
| Transformer Engine v2.19 includes JAX MoE-related changes. | Supported by NVIDIA release notes. Transformer Engine v2.19 Release Notes — Transformer Engine 2.19.0 |
| Public JAX expert-parallel tutorial documentation is complete. | Not supported; the page currently says “TODO — Coming soon.” JAX: Expert Parallelism with TransformerEngine — Transformer Engine 2.19.0 |
Context and prior work
MoE is an old idea, but its modern LLM role is to increase parameter capacity while activating only a subset of experts per token. Switch Transformer simplified routing and showed sparse models could train faster than dense baselines under matched resources, while also highlighting complexity, communication costs, and instability. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity [](https://proceedings.mlsys.org/paper_files/paper/2023/file/5616d34cf8ff73942cfd5aa922842556-Paper-mlsys2023.pdf)
Read the full section
MoE is an old idea, but its modern LLM role is to increase parameter capacity while activating only a subset of experts per token. Switch Transformer simplified routing and showed sparse models could train faster than dense baselines under matched resources, while also highlighting complexity, communication costs, and instability. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity GLaM and later MoE models further established the scaling appeal of sparsely activated experts. GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
DeepSeek-V3 is a relevant stress test because it is itself a large MoE model. Its technical report describes 671B total parameters with 37B activated per token, using MLA and DeepSeekMoE architectures, an auxiliary-loss-free load-balancing strategy, and a multi-token prediction objective. DeepSeek-V3 Technical Report
MegaBlocks is the closest prior dropless MoE systems reference surfaced in search. It directly addressed the problem of dynamic expert token counts without token dropping, using block-sparse operations and new GPU kernels. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts Tutel is another important systems predecessor, focusing on adaptive MoE parallelism and optimized all-to-all communication. [](https://proceedings.mlsys.org/paper_files/paper/2023/file/5616d34cf8ff73942cfd5aa922842556-Paper-mlsys2023.pdf)
Limitations, safety, and contested findings
The main limitation is evidentiary: the performance numbers are NVIDIA’s own measurements. The stack is tightly coupled to NVIDIA GPUs, Transformer Engine, cuBLAS/cuBLASLt, NCCL EP, MaxText, and XLA GPU behavior. Teams on AMD, TPU, Trainium, or generic PyTorch stacks should not assume comparable gains.
Read the full section
The main limitation is evidentiary: the performance numbers are NVIDIA’s own measurements. No independent reproduction, profiler traces, raw logs, or third-party benchmark report were found. The configuration is useful, but a 15-step reproduction setting is not equivalent to full pretraining evidence. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
The second limitation is portability. The stack is tightly coupled to NVIDIA GPUs, Transformer Engine, cuBLAS/cuBLASLt, NCCL EP, MaxText, and XLA GPU behavior. Teams on AMD, TPU, Trainium, or generic PyTorch stacks should not assume comparable gains.
The third limitation is model-quality evidence. Dropless routing avoids dropping routed tokens, but the NVIDIA post does not present downstream quality or convergence comparisons for the optimized path. Prior work motivates avoiding token dropping, but it does not independently validate NVIDIA’s DeepSeek-V3 JAX result.
Safety implications are indirect. This is training infrastructure, not a model capability release. The main safety-relevant effect is economic: faster large-MoE training can lower the cost of frontier-scale or near-frontier-scale model development, increasing the number of actors able to train large sparse models if they also have hardware access.
Business and practitioner implications
For AI infrastructure leaders, this is a reminder that MoE economics depend heavily on the software stack. For developers, the practical adoption path is to start with the NVIDIA MaxText container, verify correctness on a smaller MoE, instrument step time, grouped GEMM latency, dispatch/combine latency, MFU, and exposed collective time, then scale.
Read the full section
For AI infrastructure leaders, this is a reminder that MoE economics depend heavily on the software stack. Buying more accelerators is insufficient if routing, all-to-all, grouped matmuls, memory pressure, and compiler scheduling leave GPUs idle.
For JAX users, the announcement is strategically important: it suggests NVIDIA is investing in JAX as a first-class large-scale training stack alongside PyTorch. Transformer Engine officially supports both PyTorch and JAX. Getting Started — Transformer Engine 2.19.0
For developers, the practical adoption path is to start with the NVIDIA MaxText container, verify correctness on a smaller MoE, instrument step time, grouped GEMM latency, dispatch/combine latency, MFU, and exposed collective time, then scale. NVIDIA itself recommends starting small before scaling and tracking these metrics. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
For procurement and strategy teams, the main takeaway is that Blackwell-era MoE training performance may increasingly depend on vendor-integrated software features such as MXFP8/NVFP4, NCCL EP, and compiler-level communication overlap. That can improve time-to-train, but it may also deepen platform lock-in.
Sources
- NVIDIA Technical Blog, “Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine,” September 14, 2026. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
- MaxText MoE configuration documentation. maxtext/docs/reference/core_concepts/moe_configuration.md at main · AI-Hypercomputer/maxtext · GitHub
- NVIDIA Transformer Engine v2.19 release notes. Transformer Engine v2.19 Release Notes — Transformer Engine 2.19.0
Read the full section
- NVIDIA Technical Blog, “Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine,” September 14, 2026. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine | NVIDIA Technical Blog
- MaxText MoE configuration documentation. maxtext/docs/reference/core_concepts/moe_configuration.md at main · AI-Hypercomputer/maxtext · GitHub
- NVIDIA Transformer Engine v2.19 release notes. Transformer Engine v2.19 Release Notes — Transformer Engine 2.19.0
- NVIDIA Transformer Engine JAX API documentation. Jax — Transformer Engine 2.19.0
- OpenXLA GPU flags guidance. XLA Flags Guidance | OpenXLA Project
- NVIDIA/JAX Toolbox GPU performance guidance. JAX-Toolbox/docs/GPU_performance.md at main · NVIDIA/JAX-Toolbox · GitHub
- MegaBlocks: Efficient Sparse Training with Mixture-of-Experts. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
- DeepSeek-V3 Technical Report. DeepSeek-V3 Technical Report
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Tutel: Adaptive Mixture-of-Experts at Scale. [](https://proceedings.mlsys.org/paper_files/paper/2023/file/5616d34cf8ff73942cfd5aa922842556-Paper-mlsys2023.pdf)