v0.2.24
fixie-ai/ultravoxv0.2.24Jul 11, 2026by QuentinFuxa
AI Summary
This release introduces a new streaming AlignAtt translation backend and refreshes benchmarks to H100 standards. It also fixes issues with `mlx-whisper` producing empty captions and improves error handling by making server failures more explicit.
Key Highlights
- New AlignAtt translation backend for streaming LLM translation
- Benchmarks re-measured on H100 to correct previous data
- Fixed empty captions issue with `mlx-whisper`
- Server now fails loudly on ASR backend failures
New Features
- AlignAtt translation backend (`--translation-backend alignatt`)
Full Release Notes
Streaming LLM translation, refreshed benchmarks and loud failures. ## Added - AlignAtt translation backend (`--translation-backend alignatt`): streaming LLM translation through an [Alignatt4LLM](https://github.com/QuentinFuxa/Alignatt4LLM) sidecar. The model drafts ahead over the unstable ASR tail and commits only target words whose attention lands on committed source words, so translations are append-only and released the instant the ASR commits. See [docs/translation-alignatt.md](https://github.com/QuentinFuxa/WhisperLiveKit/blob/main/docs/translation-alignatt.md). - Benchmark figures re-measured on an H100 with the current backends (the README scatter plots dated from March and misrepresented the qwen3 backends). Raw results in `benchmarks/h100_scatter/`. ## Fixed - mlx-whisper produced empty captions with torch >= 2.13 (#383, thanks @RobertBartelds-FleetEnergies). - The server now fails loudly when the ASR backend produces nothing: warmup errors abort startup with the cause, and a watchdog logs an explicit error if audio flows but no text is ever produced.