Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages

NVIDIA published a workflow for adapting its Nemotron 3.5 ASR model to Saudi Arabic dialects using the NVIDIA NeMo framework. The model supports multilingual streaming transcription across 40 language-locales, including Arabic, but NVIDIA said deployment-specific dialects such as Najdi and Hijazi still benefit from targeted fine-tuning.
The workflow combines minimal corpus curation, weighted replay mixing with FLEURS data, duration-based bucketing, and optional partial encoder unfreezing. Fine-tuning on 133.7 hours of Najdi and Hijazi speech reduced word error rate from 55.05% to 29.96% on the target test split, and English performance improved from 11.04% to 10.42% WER without degrading other Arabic dialects, according to the reported results.
Curation retained 103,559 of 125,490 utterances, or 82.5%, drawn from SADA 2022. The replay mix used 90% Saudi speech, 7% FLEURS English and 3% FLEURS Arabic. Training ran for 12,000 steps over roughly 4.5 hours on two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs.
Updating all 24 encoder layers produced the lowest error rates at the cost of 230.4 million trainable parameters, with 407.6 million frozen, while partial freezing costs 2.4 points against a full fine-tune. NVIDIA said the full-encoder choice reflects this data volume rather than a general rule, and that freezing suits thinner corpora or memory constraints.
At inference, switching to the widest supported attention context of 13 lookahead frames lowered WER by 1.31 absolute points without retraining, adding approximately 800 ms of latency. A later sweep found NeMo's batched malsd_batch decoding with beam 4 reduced SADA WER from 29.96% to 28.81% at 0.59 times greedy runtime, while beam 8 reached 28.62% at 0.64 times. NVIDIA said it did not evaluate MAES or NGPU-LM fusion.
NVIDIA also released Nemotron 3 Diarization, an open model for speaker-attributed transcription of up to eight speakers that aligns speaker boundaries with ASR timestamps. NVIDIA noted the workflow is not evidence for every Arabic dialect or deployment environment, and that guidance for other languages should start with an audited manifest, tokenizer testing, and appropriate normalization.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Automatic speech recognition must handle how people actually speak, not only the languages and styles that dominate pretraining data. Regional dialects and...