v0.6

modelscope/FunASRv0.6Aug 18, 2025by liPatrick

AI Summary

Ultravox release improving language understanding, adding noise tokenization for non-human audio, and introducing new model variants like gemma3 and qwen3.

Key Highlights

  • Improved Hindi language understanding
  • Produces <noise> token on noisy or non-human audio
  • Improved background noise and audio quality robustness
  • New gemma3 and qwen3 variants

New Features

  • Language improvements
  • New model variants
  • Noise robustness

Full Release Notes

We're releasing **Ultravox v0.6** today. The [weights](https://huggingface.co/fixie-ai) have been pushed to Hugging Face. If you're using the [Ultravox Realtime APIs](https://docs.ultravox.ai), v0.6 is the new default.

## What's New
* v0.6 improves upon 0.5 in the following ways:
* Improvements on Hindi language understanding.
* Produces <noise> token on noisy or non-human audio. 
* Improved background noise and audio quality robustness in all languages. 
* New gemma3 and qwen3 variants in addition to our base llama 3.3 models.

## Evals
New eval support for [VoiceBench](https://github.com/MatthewCYM/VoiceBench). The benchmark tests speech-language models on 9 different tasks ranging from open-form text generation to question-answering and instruction following. 

## Training
This version of Ultravox continues to use a frozen Llama pre-trained core (3.1 for 8B and 3.3 for 70B), along with new gemma3 27b and qwen3 32b variants.

## What's Changed
Training Stability: Patch HF Hub and Datasets methods and update [datasets.py](http://datasets.py/) by @farzadab #280
General improvement. by @zqhuang211 #281
Ultravox v0.6 + general improvements by @liPatrick #309
Allow response generation with no user message by @matthewclso #310
Add voicebench evaluation suite by @zqhuang211 #312

**Full Changelog**: https://github.com/fixie-ai/ultravox/compare/v0.5...v0.6