Building with AI

Roofline analysis reveals why ZeRO-3 fails for DeepSeek-V3 on Hopper

Speed of light modeling shows that FSDP creates a communication bottleneck on InfiniBand for DeepSeek-V3, making pipeline parallelism the necessary choice for efficient training.

Infrastructure engineers analyzing the training configuration for DeepSeek-V3 on H800 clusters have used roofline analysis to determine optimal parallelism strategies. Published in October 2026, this technical deep dive demonstrates that Fully Sharded Data Parallelism (FSDP), also known as ZeRO-3, creates a severe communication bottleneck when scaling this specific model architecture.

The analysis concludes that pipeline parallelism is required to keep the system compute-bound rather than communication-bound. By comparing theoretical hardware limits against actual measured performance, the authors provide a method for predicting training efficiency without building and benchmarking multiple full-scale systems.

What happened

In the second part of a series on training DeepSeek-V3, the authors examined how parallelism choices, activation checkpointing, and low-precision arithmetic affect memory requirements. They identified a configuration that allows the model to fit on a 2048-node H800 cluster. The authors note that this outcome was not accidental; fixing the model architecture and targeting the specific hardware DeepSeek used naturally leads to the configuration the original team selected.

The core challenge addressed is selecting a parallelism strategy that maximizes useful floating-point operations per second (FLOPs/sec) without the prohibitive cost of implementing and benchmarking multiple systems on a full-size cluster. Implementing pipeline parallelism is significantly more complex than FSDP, so engineers need a reliable way to decide between them before writing code. The authors argue that even if one were to build both systems, benchmarking errors could invalidate the comparison.

To solve this, they applied roofline analysis, a classic performance modeling technique. This method treats the total required FLOPs as fixed and identifies whether the system is limited by computation speed or distributed communication bandwidth. The analysis reveals that FSDP is a poor choice for DeepSeek-V3 because it becomes bound by InfiniBand bandwidth, whereas other strategies can keep the GPUs busy with computation.

How it works

Roofline analysis relies on the concept of "speed of light" (SOL) performance, which represents the strict upper bound of what hardware can achieve. In physics, nothing exceeds the speed of light; in computing, no software can exceed the theoretical peak of the underlying silicon. However, using marketing-spec peaks is often misleading. NVIDIA’s advertised TFLOP/s figures assume ideal conditions, such as zeroed tensors and perfect instruction scheduling, which rarely occur in real matrix multiplications.

Real-world performance diverges from spec sheets due to memory loading, cache hierarchy latency, and power limits. For instance, running non-zero data can trigger power caps that reduce clock speeds, meaning performance depends on input data values. To address this, the authors use achievable FLOPs measured via microbenchmarks from HuggingFace’s Smol Training Playbook. For BF16 precision, the H800 achieves 758 TFLOP/s, which is 76.6% of its 989 TFLOP/s theoretical peak. For FP8, it achieves 1.46 PFLOP/s, or 73.6% of the 1.98 PFLOP/s peak.

Network bandwidth is treated similarly. While InfiniBand specs claim 50 GB/s, which is largely achievable, NVLink performance on H800s is lower than on H100s. DeepSeek reported achieving only 160 GB/s unidirectional bandwidth on H800 NVLink, compared to the 200 GB/s spec. The analysis uses these measured values to calculate the time spent on communication versus computation.

In distributed training, the bottleneck shifts from high-bandwidth memory (HBM) to cross-node bandwidth. The total FLOPs required for a training step remain constant regardless of parallelism, much like slicing a pizza does not change its total size. However, the communication overhead varies drastically. If communication time exceeds computation time, the system is communication-bound, and GPUs sit idle waiting for data. The goal is to ensure computation takes longer than communication, allowing the trainer to overlap these operations and hide latency.

Key details

  • Hardware target: The analysis focuses on a 2048-node cluster of NVIDIA H800 GPUs.
  • Achievable compute: Measured BF16 performance is 758 TFLOP/s (76.6% of peak), and FP8 is 1.46 PFLOP/s (73.6% of peak).
  • Network constraints: InfiniBand bandwidth is assumed at 50 GB/s per GPU, while NVLink is capped at 160 GB/s based on DeepSeek’s reports.
  • Parallelism verdict: FSDP (ZeRO-3) is rejected because it becomes bound by InfiniBand bandwidth for this model size.
  • Methodology: The approach compares computed "speed of light" times for FLOPs and byte transfers to determine if a step is compute-bound or communication-bound.
  • Simplification: Multi-Token Prediction (MTP) is omitted from this specific analysis to maintain clarity.

Why it matters

For machine learning infrastructure teams, this analysis provides a rigorous framework for making high-stakes architectural decisions without massive upfront investment. Building and benchmarking multiple parallelism strategies on thousands of GPUs is prohibitively expensive and time-consuming. Roofline analysis offers a spreadsheet-based alternative that can predict performance bottlenecks with reasonable accuracy. It prevents teams from pursuing implementation paths, such as complex pipeline parallelism setups, only to discover later that a simpler approach like FSDP would have failed due to network limits.

Furthermore, the distinction between marketing specs and achievable performance is critical for accurate capacity planning. Relying on theoretical peaks can lead to significant overestimation of training throughput. By using measured microbenchmarks, engineers can create realistic timelines and budget expectations. This is especially important as models grow larger and precision drops, which accelerates computation but does not reduce communication volume, potentially exacerbating bandwidth bottlenecks.

What you can do

  • Benchmark your own hardware’s GEMM performance using libraries like cuBLAS to determine achievable FLOP/s rather than relying on spec sheets.
  • Measure actual NVLink and InfiniBand bandwidth in your cluster environment, as real-world topology and power caps may reduce throughput.
  • Use roofline analysis to compare parallelism strategies before implementation, focusing on whether communication time exceeds computation time.
  • Account for power limits in your models, recognizing that non-zero data inputs may throttle GPU clock speeds during training.
  • Prioritize compute-bound configurations where possible, ensuring that communication overhead can be overlapped with calculation.
  • Re-evaluate assumptions when changing precision levels, as lower precision increases compute speed but may shift the bottleneck to network bandwidth.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories