Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Autoencoders vs. PCA: I Rigged the Test and PCA Still Won

    RFK Jr. thinks AI will free us from the “tyranny” of medical facts, expertise

    Open, scalable training infrastructure for large MoEs

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI Tutorials»Open, scalable training infrastructure for large MoEs
    AI Tutorials

    Open, scalable training infrastructure for large MoEs

    By No Comments7 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Open, scalable training infrastructure for large MoEs
    Share
    Facebook Twitter LinkedIn Pinterest Email


    đź“„ Tech Report | đź’» Code | đź§© Interactive demo

    Today we’re releasing Olmo-core 3, a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system.

    Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. It’s one of the core systems behind the next generation of Olmo, and part of our ongoing commitment to open up the tools and training infrastructure behind each new model.

    Training large AI models takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. MoE models offer a more efficient approach—they can contain many more learned components, or parameters, without requiring every input to use all of them. But the full model still has to be stored across GPU memory and updated during training, and directing inputs to the right experts – the specialized components within an MoE – across a cluster creates its own communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input.

    Olmo-core 3 is built to close that gap. In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.

    The same infrastructure has been benchmarked at over one trillion total parameters.

    Olmo-core 3: expert-pool scaling and training throughput



    Building a training stack around how MoEs actually work

    Olmo-core has evolved with each generation of Olmo.

    Our work on sparse models goes back to OlmoE, which used an MoE architecture with 64 routed experts. Olmo 3, by contrast, used a dense architecture, meaning nearly all of the model was active for every token and its training stack was built around that design. Olmo-core 3 extends the framework with a training system designed for much larger MoE models.

    Our earlier MoE implementation in Olmo-core used fully sharded data parallelism (FSDP), configured to gather and reshard model weights for each small batch of training data. Olmo-core 3 switches to a system based on distributed data parallelism (DDP). It keeps experts resident on GPUs and routes the relevant data to them, avoiding that repeated weight gathering.

    NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.

    Olmo-core 3 training throughput compared with the earlier implementation



    Scaling and optimizing MoE training

    Olmo-core 3 combines several techniques for distributing large MoEs across GPU clusters with optimizations that make routing and computation more efficient.

    Three techniques determine how the model and its training state are split across hardware:

    • Expert parallelism spreads the experts across GPUs, so each GPU stores only part of the full expert pool.
    • Pipeline parallelism splits the model’s layers – the successive stages that transform an input – across groups of GPUs, reducing how much of the model each GPU needs to keep in memory.
    • A distributed optimizer spreads the optimizer state – the additional data used to calculate and apply updates during training – across GPUs instead of storing a full copy on every GPU.

    Together, these techniques allow an MoE to scale without requiring every GPU to keep the entire model and its training state in memory.

    Olmo-core 3 also reduces the cost of routing data to the right experts and running their computations. Rowwise expert parallelism places routed data directly into expert input buffers, minimizing the extra work needed to rearrange it. GPU-resident routing keeps routing metadata on the GPUs, so the CPU can queue work without waiting for that information to be copied back. And grouped GEMM combines many small expert computations so GPUs can execute them more efficiently.

    Finally, Olmo-core 3 supports MXFP8, a lower-precision number format that represents some values with fewer bits. This can reduce computation and the amount of data moved between GPUs, as long as those savings outweigh the cost of converting between number formats.

    We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.

    These techniques and optimizations have to work together. Speeding up one part of training can create costs elsewhere; faster computation may require more data movement, while moving fewer bits may not help if converting the data takes too long. Olmo-core 3 is built around those trade-offs across the full training process, giving us – and researchers using the open stack – control over how the pieces fit together.

    Explore our interactive walkthrough to see how data, expert, and pipeline parallelism work together to scale MoE training—from a single GPU to many.

    Blog Banner



    Scaling into the trillion-parameter range

    We’ve benchmarked Olmo-core 3 across a range of configurations on NVIDIA B300 GPUs, including a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs. Its highest observed throughput was 858 TFLOP/s/GPU—a measure of useful model computation per second on each GPU. These tests used random routing to measure system performance, rather than the quality of a trained model.

    We’ve also experimented with DeepEP v2, an alternative way of handling communication between experts across GPUs, reaching a configuration with 2.38 trillion total parameters. This was a short-capacity test rather than a full training run, so it demonstrates the scale Olmo-core 3 can reach rather than sustained training performance.

    At these scales, systems performance is only part of the picture. Our technical report also documents experiments that informed how we train MoEs and measure their performance. For example:

    • A score intended to encourage balanced routing could improve even as the actual workload became less balanced. We call this failure token gerrymandering.
    • Lowering experts’ learning rates – the size of their training updates – because they process fewer tokens did not improve results in the model family we tested.
    • GPU calculations took different amounts of time when the values being processed changed, even with the same matrix dimensions. Performance comparisons therefore need matching input values as well as matching shapes.
    • Overlapping communication and computation on separate GPU streams did not always make training faster. In some tests, it slowed end-to-end execution—a reminder that more overlap does not necessarily mean higher throughput.

    The report explains these findings alongside the approaches we tested and chose not to adopt.



    Built for the next generation of Olmo, open for everyone

    Olmo-core 3 is the foundation for what we’re building next. Our next-generation Olmo will use an MoE architecture, and we’re aiming for it to be our most capable Olmo yet, trained on our largest dataset and with our longest context window.

    The new stack lets us scale beyond our previous MoE work while giving us more flexibility to adapt training as models and hardware evolve. And it’s fully open—researchers and developers can use Olmo-core 3 to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism, and other parts of the system.

    That’s part of how we think about open model development—model weights are more useful when the infrastructure and training decisions behind them are open too.

    For a deeper look at the systems design, experiments, ablations, and approaches we tested along the way, read our technical report and explore Olmo-core 3 on GitHub.

    infrastructure large MoEs open scalable Training
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticlePhoton held a funeral for mobile apps. Now it has $4.5M to help replace them with agents
    Next Article RFK Jr. thinks AI will free us from the “tyranny” of medical facts, expertise
    • Website

    Related Posts

    AI Tutorials

    Databricks Academy: What the Courses, Certs and Costs Really Get You

    AI Tutorials

    Snowflake University Explained: Free Training That Actually Makes You Data-Fluent

    AI Tutorials

    He Built This City

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Autoencoders vs. PCA: I Rigged the Test and PCA Still Won

    0 Views

    RFK Jr. thinks AI will free us from the “tyranny” of medical facts, expertise

    0 Views

    Open, scalable training infrastructure for large MoEs

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Autoencoders vs. PCA: I Rigged the Test and PCA Still Won

    0 Views

    RFK Jr. thinks AI will free us from the “tyranny” of medical facts, expertise

    0 Views

    Open, scalable training infrastructure for large MoEs

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.