v0.6

unclecode/crawl4aiv0.6Aug 18, 2025by liPatrick

AI Summary

This release introduces new model variants (gemma3 and qwen3), enhances audio quality robustness across languages, and adds support for the VoiceBench evaluation suite.

Key Highlights

  • New gemma3 and qwen3 model variants
  • Improved background noise and audio quality robustness
  • Added VoiceBench evaluation suite support
  • Enhanced Hindi language understanding

New Features

  • gemma3 and qwen3 variants
  • VoiceBench evaluation suite support
  • Noise token generation for non-human audio

Full Release Notes

We're releasing **Ultravox v0.6** today. The [weights](https://huggingface.co/fixie-ai) have been pushed to Hugging Face. If you're using the [Ultravox Realtime APIs](https://docs.ultravox.ai), v0.6 is the new default.

## What's New
* v0.6 improves upon 0.5 in the following ways:
* Improvements on Hindi language understanding.
* Produces <noise> token on noisy or non-human audio. 
* Improved background noise and audio quality robustness in all languages. 
* New gemma3 and qwen3 variants in addition to our base llama 3.3 models.

## Evals
New eval support for [VoiceBench](https://github.com/MatthewCYM/VoiceBench). The benchmark tests speech-language models on 9 different tasks ranging from open-form text generation to question-answering and instruction following. 

## Training
This version of Ultravox continues to use a frozen Llama pre-trained core (3.1 for 8B and 3.3 for 70B), along with new gemma3 27b and qwen3 32b variants.

## What's Changed
Training Stability: Patch HF Hub and Datasets methods and update [datasets.py](http://datasets.py/) by @farzadab #280
General improvement. by @zqhuang211 #281
Ultravox v0.6 + general improvements by @liPatrick #309
Allow response generation with no user message by @matthewclso #310
Add voicebench evaluation suite by @zqhuang211 #312

**Full Changelog**: https://github.com/fixie-ai/ultravox/compare/v0.5...v0.6