MLPerf v6.0 Results

Train bigger and serve faster on the only AI cloud leading MLPerf in both training and inference

More useful work per GPU, on infrastructure purpose-built to keep up with the models you're shipping

What is MLPerf?

MLPerf is an industry-standard benchmark suite that measures machine learning performance across realistic deployment scenarios. The speed at which systems process inputs and generate outputs from a trained model directly influences performance and user experience, making the MLPerf Inference benchmark a critical performance metric for both CoreWeave and our customers. 

CoreWeave delivers leading MLPerf v6.0 results

CoreWeave’s latest MLPerf submission results reflect AI cloud leadership across both inference and training. CoreWeave is proven to give you faster iteration, stronger throughput, and more useful work per GPU.

Higher inference throughput

Leading tokens/sec per GPU for DeepSeek-R1 in server and offline mode on NVIDIA GB300 NVL72, among all v6.0 submitters

Record-setting performance

Largest-ever submission with 8,192 NVIDIA GB300 NVL72 GPUs, reaching the MLPerf quality target for the DeepSeek-V3 671B benchmark in 2.02 minutes

2.80x faster training

On the same Llama 3.1 405B benchmark, CoreWeave improved training time from 27.33 minutes in v5.0 to 9.77 minutes in v6.0, with 1.70x more useful work per GPU-minute

Left
Right

The pattern across MLPerf v6.0 results is consistent: each new GPU platform on CoreWeave converts directly into shorter training runs and higher inference throughput.

Copied
MLPERF V6.0 RESULTS

‍CoreWeave's MLPerf v6.0 results for training and inference

MLPERF TRAINING V6.0

CoreWeave helps AI pioneers train faster at scale

CoreWeave trained DeepSeek-V3 671B benchmark to the MLPerf quality target in just 2.02 minutes on 8,192 NVIDIA GB300 NVL72 GPUs. CoreWeave also trained Llama 3.1 405B in 9.77 minutes on 4,096 NVIDIA GB300 GPUs, delivering a 2.80x wall-clock speedup year over year on an identical benchmark. Scaling from 2,048 to 4,096 GPUs reached 89.6% per-doubling efficiency, and scaling from 4,096 to 8,192 GPUs reached 76.5% per-doubling efficiency.

MLPerf Inference v6.0

CoreWeave delivers unparalleled throughput leadership

For all of CoreWeave’s submissions in this round, NVIDIA GB300 NVL72 delivered the strongest DeepSeek-R1 performance, leading in offline throughput and per-GPU efficiency. DeepSeek-R1 server throughput increased by more than 14% and offline throughput increased by more than 34% on NVIDIA GB300 NVL72 versus NVIDIA GB200 NVL72. CoreWeave is turning newer GPU platforms into more inference throughput.

MLPERF V5.0 RESULTS

‍CoreWeave’s MLPerf v5.0 results for training and inference

CoreWeave’s MLPerf v5.0 results established clear leadership across both training and inference, with early NVIDIA GB200 NVL72 inference performance and record-setting training scale. Download the whitepaper for the full story.

MLPerf Training v5.0

More than 2 x faster training performance

CoreWeave, NVIDIA, and IBM partnered on MLPerf Training v5.0 results, showcasing an NVIDIA GB200 cluster 34x larger than the next largest submission. That scale cut training times to up to twice as fast and gave teams more room to iterate.

MLPerf Inference v5.0

Inference performance built for production speed

CoreWeave is the first and only cloud provider to submit MLPerf Inference v5.0 results on NVIDIA GB200 Grace Blackwell instances, clearing 800 tokens per second on the Llama 3.1 405B model – a 2.86x per-chip gain over NVIDIA H200 GPUs. Our NVIDIA H200 GPU instances reached 33,000 tokens per second on the Llama 2 70B model, 40% more throughput than NVIDIA H100 GPUs. That speed keeps GPUs working and shortens the loop between idea and result.

Trusted by leading AI labs, enterprises, and startups
Rev.comRev.com
DecartDecart
CloudflareCloudflare
AbridgeAbridge
OpenAIOpenAI
Jane StreetJane Street
CohereCohere
GoogleGoogle
WaveForms AIWaveForms AI
InflectionInflection
Fireworks AIFireworks AI
AugmentAugment
AltumAltum
ConjectureConjecture
ChaiChai
MistralAIMistralAI
NovelAINovelAI

Frequently asked questions

 What is MLPerf?

MLPerf is an industry-standard benchmark suite developed by MLCommons to measure and compare machine learning training and inference performance across hardware and platforms.

Why are MLPerf benchmarks important?

MLPerf benchmarks provide a fair, transparent method for evaluating AI hardware and cloud platforms, helping businesses choose solutions that offer optimal performance, scalability, and cost efficiency.

What does MLPerf Training v6.0 measure?

MLPerf Training v6.0 measures how quickly a computing system can train complex machine learning models—such as Meta's Llama 3.1 405B and the new DeepSeek-V3 671B mixture-of-experts benchmark—from initialization to a specified quality target, enabling fair and transparent performance comparisons across hardware platforms and cloud providers.

 What does MLPerf Inference v5.0 measure?

MLPerf Inference v5.0 measures how quickly computing systems process inputs and generate outputs using fully trained machine learning models, focusing specifically on throughput (tokens per second) and latency across realistic deployment scenarios to evaluate and compare the inference performance of hardware and cloud infrastructure providers.

How does CoreWeave's performance compare in MLPerf benchmarks?

In MLPerf Training v6.0, CoreWeave set new records—training DeepSeek-V3 671B in just 2.02 minutes on 8,192 NVIDIA GB300 GPUs, the fastest time in the round, and improving its Llama 3.1 405B time-to-train 2.8x year over year. These results were achieved on the same production GB300 NVL72 infrastructure customers run on every day.

 What makes CoreWeave GPUs unique for MLPerf results?

CoreWeave's results come from the full platform, not hardware alone. By combining the latest NVIDIA GB300 NVL72 systems with NVIDIA Spectrum-X networking, the topology-aware SUNK scheduler, and CoreWeave Mission Control, CoreWeave sustains industry-leading performance and scaling efficiency from 64 to 8,192 GPUs—on the same infrastructure available to customers.

Left
Right

Ready to accelerate your roadmap?

Train bigger, serve faster, and get more performance out of every GPU. Let’s talk about what that means for your AI roadmap.

Disclaimer

Result verified by MLCommons Association. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See www.mlcommons.org for more information.