Hugging Face published a per-language fine-tuning guide for the 'Nemotron 3.5 ASR' speech model. The composition is strongly academic—one official HF source plus four arXiv papers—research accumulation bundled into implementation steps rather than an announcement, taking the form of a how-to anyone can follow. The focus is fine-tuning a general ASR model to each language's characteristics to raise accuracy; the through-line is a move from 'one big model' toward lightly adapting to each language in practice. The notable part is organizing multilingual ASR so practitioners can optimize it themselves. But it's right after release—the degree of per-language accuracy gains and adoption in non-English regions are to confirm.
HF guides Nemotron ASR fine-tuning
HF guides Nemotron ASR fine-tuning
How to Fine-Tune Nemotron 3.5 ASR for Your Language, Domain, or Accent
NVIDIA guide: fine-tuning Nemotron 3.5 ASR for language/domain/accent
Academic (arxiv etc.) 6 ▾
USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding
USAD 2.0 scales representation distillation for universal audio understanding
From Self to Other: Evaluating Demographic Perspective-Taking in LLM Hate Speech Annotation
Evaluating demographic perspective-taking in LLM hate-speech annotation
FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition
FiLM-based speaker conditioning of a SpeechLLM for pathological speech
Revisiting Lexicon Evaluation in Unsupervised Word Discovery
Revisiting lexicon evaluation in unsupervised word discovery
Ouvia: a user-centered framework for speech-translation usability
ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity
ProSarc: audio-only sarcasm detection via prosodic incongruity