Open & local AI

GGUF replaces bitsandbytes for low-VRAM LoRA training on local hardware

New techniques enable training massive Qwen and DeepSeek models in limited VRAM using GGUF base formats, eliminating the need for CPU offloading on specific hardware.

A small glowing chip representing efficient local training next to fading server racks.
Illustration generated for this article

A new open-source recipe demonstrates how to perform Low-Rank Adaptation (LoRA) training on large language models using significantly less video memory than previously thought possible. By leveraging the GGUF file format instead of traditional quantization libraries, developers can now fine-tune models like Qwen3.6-35B on just 16 GiB of VRAM without relying on slow CPU offloading. This shift suggests that GGUF is poised to replace bitsandbytes as the standard base model format for efficient local training.

What happened

The developer behind the repository woct0rdho/transformers5-qwen3.5-recipe has published a technical deep dive into low-VRAM training methods tailored for the Strix Halo hardware architecture. The work challenges the conventional wisdom that training large models requires massive GPU clusters or extensive memory swapping. Instead, it shows that with the right combination of quantized base models and optimized kernels, even consumer-grade or integrated high-performance hardware can handle substantial fine-tuning tasks.

The core achievement involves training several state-of-the-art open-weight models entirely within VRAM limits that were previously considered insufficient. For instance, the Qwen3.6-35B-A3B model was trained using only 16 GiB of VRAM. The author notes that this efficiency implies larger variants, such as the Qwen3.5-122B-A10B, could be trained in 64 GiB, and the massive Qwen3.5-397B-A17B in 192 GiB. Similarly, the DeepSeek-V4-Flash model, which contains 284 billion parameters, was trained in 90 GiB of VRAM, while the Qwen3.8-Flash-Next model required 40 GiB.

This development is significant because it treats open-weight AI similarly to open-source software, where users not only run the weights but also modify them. The ability to modify weights locally without prohibitive hardware costs lowers the barrier to entry for customizing foundational models. However, the current implementation is specifically tuned for Strix Halo, meaning additional engineering effort is required to port these optimizations to other GPU architectures.

How it works

The method relies on replacing the bitsandbytes library with GGUF as the base model format for loading quantized weights. GGUF, originally popularized by llama.cpp for inference, is now being adapted for training workflows through a custom fork of the Transformers library. This approach uses a GGUF quantizer that integrates directly with the training loop, allowing the model to remain in a compressed state during computation rather than decompressing fully into high-precision formats that consume excessive memory.

Several specialized kernels and optimization techniques make this possible. The system employs tuned General Matrix Multiply (GEMM) operations and Mixed Precision Quantization (MMQ) similar to those found in llama.cpp. For Mixture of Experts (MoE) layers, the recipe uses AITER Triton kernels with specific configurations for non-quantized LoRA adapters. It also implements fast backward formulas for LoRA updates, akin to those used in Unsloth, to accelerate gradient calculations for both linear and MoE layers.

Memory savings are further achieved by disabling certain features that are not strictly necessary for training stability, such as the autoregressive decoding cache and load balancing loss. The system uses non-reentrant gradient checkpointing and an 8-bit AdamW optimizer from bitsandbytes to minimize memory footprint. Additionally, torch.compile is applied to the GGUF dequantization function to reduce VRAM usage during the loading phase, ensuring that the overhead of handling compressed weights remains manageable.

Key details

  • Qwen3.6-35B-A3B trains in 16 GiB VRAM using APEX-I-Mini quantization, which occupies only 13.3 GiB.
  • DeepSeek-V4-Flash (284B parameters) trains in 90 GiB VRAM using IQ2_XXS quantization.
  • Qwen3.8-Flash-Next trains in 40 GiB VRAM plus 27 GiB for engrams, using GSQ-RCO Q2_0 quantization.
  • The solution uses a custom Transformers fork with GGUF support, tracked in Hugging Face issue #40070.
  • Optimizations include tuned GEMM, MMQ, AITER Triton kernels, and RMSNorm from Liger Kernel.
  • Current kernels and parameters are specifically tuned for Strix Halo hardware, requiring adaptation for other GPUs.

Why it matters

For engineers building products with AI, this shift reduces the cost and complexity of fine-tuning large models. Traditionally, training required either expensive cloud instances with hundreds of gigabytes of VRAM or complex setups involving CPU offloading, which drastically slows down training times. By keeping the entire training process in VRAM using efficient quantization, developers can iterate faster and experiment with larger models on more accessible hardware. This democratizes access to model customization, allowing smaller teams to tailor foundational models to their specific domains without massive infrastructure investments.

The move toward GGUF for training also signals a convergence between inference and training toolchains. Previously, developers had to maintain separate pipelines for converting models for inference (often using GGUF or similar formats) and for training (using full precision or bitsandbytes). Unifying these formats simplifies the workflow, reducing the risk of errors during conversion and ensuring that the model behavior during training closely matches its behavior during deployment. This consistency is crucial for maintaining model performance and reliability in production environments.

What you can do

  • Experiment with the provided recipe on Strix Halo hardware to benchmark training speeds and memory usage for your specific use cases.
  • Monitor the Hugging Face Transformers issue #40070 to track the potential merger of GGUF quantization support into the main library.
  • Evaluate if your current fine-tuning workflows can benefit from switching from bitsandbytes to GGUF-based quantization for base models.
  • Investigate the custom Triton kernels and MMQ implementations to understand how they might be adapted for your specific GPU architecture if not using Strix Halo.
  • Consider the memory savings from disabling autoregressive decoding cache and load balancing loss when designing your training loops for similar large models.
  • Explore the use of torch.compile on dequantization functions to further optimize VRAM usage during model loading and training.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories