Building with AI

Burn 0.22 removes backend generics to speed up Rust ML builds

Burn 0.22 eliminates backend type parameters from user APIs, cutting rebuild times by up to 15x and adding LoRA support.

Rust gear mechanism interlocking with glowing circuit boards
Illustration generated for this article

The Burn machine learning framework for Rust has released version 0.22, a significant update that simplifies application code and drastically reduces compilation times. Released in October 2026, this version removes backend generic types from the user-facing API, allowing developers to select execution devices at runtime rather than compile time. The update also introduces native support for LoRA and QLoRA fine-tuning, improved memory management, and faster build processes for complex models.

What happened

Previous versions of Burn required developers to propagate backend type parameters, such as B: Backend, throughout their entire application stack. This approach offered flexibility but created a heavy dependency chain that slowed down compilation whenever model structures changed. With version 0.22, these backend generics are removed from user code. Instead, execution context is selected via the device initialization, such as Device::cuda(0) or Device::wgpu(). This shift decouples high-level tensor operations from specific backend implementations, streamlining the development workflow.

The impact on build times is substantial. In benchmarks provided by the development team, removing a hidden layer from a small convolutional neural network reduced median release rebuild times from 28.42 seconds to 4.57 seconds. For a transformer model with a custom training loop, alternating equivalent feedforward expressions saw rebuild times drop from 14.73 seconds to just 1.00 second. These improvements stem from breaking the dependency chain that previously forced recompilation of large portions of the codebase upon minor model edits.

Beyond API changes, the release focuses on runtime performance and developer experience. It introduces adaptive memory pools that adjust allocation sizes based on workload statistics, reducing peak VRAM usage by nearly half in some CNN benchmarks. The update also adds support for fine-tuning existing models using LoRA and QLoRA, enabling efficient adaptation of large models without modifying their base layers. Additionally, the framework now supports exporting models to ONNX and integrates remote compute capabilities via the Iroh transport protocol.

How it works

The core architectural change involves a new execution path: Tensor → Bridge → Dispatch → Backend. The bridge layer hides concrete backend representations from the high-level tensor API, effectively erasing types that previously tied application code to specific backends. This type erasure allows the dispatch system to route operations to the appropriate backend at runtime. While the Backend trait remains central for implementing custom operations, application code no longer needs to carry generic constraints. Autodiff contexts are now configured on the device and inherited by tensors, moving precondition checks to runtime.

Memory management has been overhauled through CubeCL’s new adaptive memory pools. Instead of static pool sizes, the system monitors allocation statistics during a dry run and adjusts page sizes dynamically. It releases outdated pages when they become empty and moves live allocations to free memory sooner. Small, frequently changing allocations are kept in a separate pool to minimize fragmentation. This approach ensures that reserved memory closely matches actual usage, significantly lowering peak VRAM requirements without manual configuration.

Custom operations are now integrated via the #[backend_extension] macro, which connects user-defined kernels to the dispatch system. This allows developers to expose custom functions, such as fused matrix multiplication with bias and ReLU, through standard tensor interfaces without reintroducing backend generics. The macro generates registration for lazy execution and handles fusion graph boundaries, ensuring that custom kernels can be optimized alongside built-in operations. This mechanism underpins new libraries like burn-linalg and burn-signal.

Key details

  • Build Speed: Median rebuild times decreased by up to 15x, with transformer models dropping from 14.73s to 1.00s for equivalent code changes.
  • API Simplification: Backend type parameters (B: Backend) are removed from user-facing structs and functions, replaced by runtime device selection.
  • Memory Efficiency: Adaptive memory pools reduced peak VRAM usage by 49% for CNNs (956 MiB to 486 MiB) and 17.7% for transformers.
  • Fine-Tuning Support: Native integration for LoRA and QLoRA allows freezing base weights while training low-rank adapters with independent optimizer configurations.
  • Remote Compute: Burn Remote now uses Iroh for authenticated, encrypted peer-to-peer connections, supporting graph replay to reduce communication overhead.
  • Compiler Updates: CubeCL migrated to Pliron for kernel representation and added LLVM targets for AMD and NVIDIA GPUs, improving portability and optimization.

Why it matters

For software engineers building ML products in Rust, compilation speed is a major productivity bottleneck. The removal of backend generics means that iterative development—tweaking model architectures or debugging training loops—becomes significantly faster. Developers no longer need to wait tens of seconds for every minor change to recompile, enabling a more responsive feedback loop. This change lowers the barrier to entry for using Rust in ML, where C++ and Python have traditionally dominated due to easier iteration cycles.

The addition of LoRA and QLoRA support addresses a critical need for deploying large language models and other foundation models in resource-constrained environments. By allowing developers to fine-tune models efficiently without storing full gradient states for all parameters, Burn 0.22 makes it feasible to adapt large models on consumer hardware. The adaptive memory management further enhances this capability, ensuring that available VRAM is used effectively, which is crucial for running larger batch sizes or more complex models on limited GPU memory.

What you can do

  • Update your Burn dependencies to version 0.22 and remove backend generic parameters from your model structs and function signatures.
  • Replace compile-time backend selection with runtime device initialization using Device::cuda(), Device::wgpu(), or Device::flex().
  • Experiment with the new Lora module to fine-tune existing models, using ParamGroup to control which layers receive adapters.
  • Monitor memory usage with the updated device API to observe how adaptive pools reduce VRAM consumption in your specific workloads.
  • If you use custom kernels, refactor them to use the #[backend_extension] macro to integrate with the new dispatch system and enable fusion.
  • Explore remote execution options by setting up a Burn Remote server with Iroh transport for distributed training or inference tasks.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories