Fine-tuning NVIDIA Nemotron ASR for Saudi Arabic dialects
A technical workflow using NVIDIA NeMo reduces word error rates for Najdi and Hijazi speech by combining weighted replay mixing and partial encoder unfreezing.
NVIDIA engineers have detailed a specific fine-tuning workflow for the Nemotron 3.5 Automatic Speech Recognition (ASR) model, targeting underrepresented Saudi Arabic dialects. Published in October 2026, the guide demonstrates how to adapt a multilingual streaming model to handle Najdi and Hijazi speech without degrading performance in English or other Arabic variants. The approach relies on the NVIDIA NeMo framework and emphasizes careful data curation over brute-force retraining.
What happened
Multilingual ASR models often struggle with regional dialects and local recording conditions that differ from their pretraining data. While Nemotron 3.5 supports transcription across 40 language-locales, it showed significant gaps when processing Saudi dialects like Najdi and Hijazi. To address this, NVIDIA developed a targeted adaptation pipeline using 133.7 hours of speech data from the SADA 2022 dataset. The goal was to lower the word error rate (WER) for these specific dialects while maintaining the model’s existing capabilities in Modern Standard Arabic and English.
The team started with a baseline experiment that fine-tuned the model solely on the target dialect data. This initial attempt yielded modest improvements, with validation WER plateauing after 45 epochs. The results indicated that training exclusively on the new dialect caused the model to forget previously learned patterns, a phenomenon known as catastrophic forgetting. The baseline WER for the target Najdi and Hijazi test split remained high at 55.05%, suggesting that a simple full fine-tune was insufficient for low-resource dialect adaptation.
To overcome these limitations, the engineers introduced three key changes: minimal but strict data curation, weighted replay mixing, and duration-based bucketing. They also experimented with partial encoder unfreezing to balance computational cost and accuracy. The final workflow reduced the WER for Najdi and Hijazi speech to 29.96%, a substantial improvement of over 25 percentage points. Surprisingly, the model’s performance on English also improved slightly, dropping from 11.04% to 10.42% WER, proving that targeted specialization can enhance overall robustness if managed correctly.
How it works
The core of the solution is a weighted replay mixing strategy that prevents the model from overwriting its general knowledge. Instead of training only on Saudi dialect data, the pipeline mixes in 10% of data from the FLEURS dataset, split into 7% English and 3% Arabic. This ensures the model continues to practice its original languages while learning the new dialect. The mixing is declared via configuration weights rather than simple file concatenation, ensuring the replay data is evenly distributed throughout training batches via sliding-window shuffling.
Data curation plays a critical role in removing noise without discarding difficult but valid speech. The team filtered out clips with inaudible markers, such as the Arabic annotation for "unclear," and removed utterances with unrealistic character-to-duration ratios. They used NVIDIA NeMo Curator stages to score audio quality, setting custom thresholds for UTMOS and SIGMOS metrics based on the specific distribution of the SADA dataset. This preserved 82.5% of the original utterances, keeping challenging accents that generic filters might have rejected.
Training efficiency is achieved through duration-based bucketing, which groups utterances of similar lengths into the same batch. This reduces padding and memory usage, allowing for larger effective batch sizes. The team also adjusted the attention context to 13 lookahead frames and used beam-8 MALSD decoding for offline evaluation, which lowered WER by an additional 2.71 absolute points without requiring retraining. For speaker-attributed transcription, the workflow integrates NVIDIA Nemotron 3 Diarization to align speaker boundaries with ASR timestamps for up to eight speakers.
Key details
- Dataset: 133.7 hours of Najdi and Hijazi speech from SADA 2022, mixed with 10% FLEURS data for replay.
- Performance gain: WER for target dialects dropped from 55.05% to 29.96%; English WER improved from 11.04% to 10.42%.
- Hardware: Experiments ran on two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs for approximately 4.5 hours.
- Model architecture: Cache-Aware FastConformer-RNNT with prompted multilingual streaming capabilities.
- Parameter efficiency: Updating all 24 encoder layers involved 230.4 million trainable parameters; partial unfreezing offers a cheaper alternative.
- Decoding optimization: Switching to beam-8 MALSD decoding and a larger attention context reduced WER by 2.71 points without retraining.
Why it matters
For software engineers building voice interfaces in diverse linguistic regions, this workflow provides a reproducible recipe for handling low-resource dialects. It demonstrates that you do not need massive datasets to specialize a large multilingual model. By using replay mixing and careful curation, teams can deploy accurate ASR for specific regional variants without sacrificing the broad language support that makes pre-trained models attractive. This is particularly relevant for applications in customer service, healthcare, and education where local dialect accuracy is critical for user trust.
The technical insights also highlight the importance of data engineering over model architecture changes. The significant performance gains came from how the data was mixed, filtered, and batched, rather than from altering the underlying neural network structure. This shifts the focus for AI practitioners toward building robust data pipelines that can handle noisy, real-world audio. The ability to improve English performance while specializing in Arabic suggests that well-managed fine-tuning can act as a form of regularization, strengthening the model’s general features.
What you can do
- Reproduce the workflow: Use the provided NVIDIA NeMo ASR fine-tuning notebook to run the Saudi Arabic adaptation recipe on your own hardware.
- Curate carefully: Inspect the distribution of audio quality scores in your dataset before applying generic filters to avoid discarding valid dialect speech.
- Implement replay mixing: When fine-tuning on a narrow domain, mix in a small percentage of general-purpose data to prevent catastrophic forgetting.
- Enable bucketing: Ensure
use_bucketingis set to true in your NeMo configuration to reduce padding and improve training efficiency. - Test partial unfreezing: If compute resources are limited, try unfreezing only the top encoder layers and the decoder instead of the entire model.
- Evaluate independently: Always measure performance on a held-out test set that includes both target dialects and general languages to monitor for degradation.

