whisperX Releases
103 releases of m-bain/whisperX
- v0.0.7
Minimal release notes: "### Added - pipecat smart turn training ([#16](https://github.com/wavekat/wavekat-turn/pull/16))"
May 10, 2026
- v3.8.5
This release addresses compatibility issues with Torch 2.8 by pinning specific versions of torchvision and torchcodec.
Apr 1, 2026
- v1.4.0v1.4.0 — 上游大迁移:更稳定、更简单、更可靠
Performs a major migration of upstream tools to improve stability and reduce deployment complexity across 6 channels.
Mar 31, 2026
- v0.0.6
Minimal release notes: "### Added - expose audio_duration_ms in TurnPrediction ([#15](https://github.com/wavekat/wavekat-turn/pull/15)) - add T"
Mar 31, 2026
- v1.0.0v1.0.0 — Multi-Provider Council
The first stable release introducing a multi-provider system where 18 historical thinkers deliberate on complex problems. It features automatic provider detection, streamlined protocols, and specific deliberation modes.
Mar 30, 2026
- v3.7.9
This release backports a bug fix to restore word-level timestamps for words containing digits, symbols, or foreign scripts.
Mar 25, 2026
- v3.6.2
A backport of word-level timestamp fixes from v3.8.4, specifically targeting the restoration of timestamps for unalignable characters and improving the testing infrastructure.
Mar 25, 2026
- v3.5.2
This release backports a bug fix to restore word-level timestamps for words containing digits, symbols, or foreign scripts.
Mar 25, 2026
- v3.4.5
Backport of word-level timestamp fixes from v3.8.4, focusing on restoring functionality for special characters and improving test coverage.
Mar 25, 2026
- v3.3.6
Backport of word-level timestamp fixes from v3.8.4, aiming to correct alignment issues for words containing special characters.
Mar 25, 2026
- v3.8.4
Adds progress callbacks for transcription, alignment, and diarization functions, along with fixes for file handle leaks and word-level timestamps. Requires faster-whisper version 1.2.0 or higher.
Mar 25, 2026
- v3.3.5
Backport of alignment fixes from v3.8.2, restoring the original CTC forced-alignment logic and fixing blank_id handling.
Mar 10, 2026
- v3.4.4
Backport of alignment fixes from v3.8.2, restoring the original CTC forced-alignment logic and fixing blank_id handling.
Mar 10, 2026
- v3.5.1
Backport of alignment fixes from v3.8.2, restoring the original CTC forced-alignment logic and fixing blank_id handling.
Mar 10, 2026
- v3.6.1
Backport of alignment fixes from v3.8.2, correcting the CTC forced-alignment behavior and blank_id handling.
Mar 10, 2026
- v3.7.8
Backport of alignment fixes from v3.8.2, restoring the original CTC forced-alignment logic and fixing blank_id handling for specific model types.
Mar 10, 2026
- v3.8.2
Exposes average log probability per segment from ctranslate2 beam search and reverts a previous wildcard alignment change that negatively impacted word-level timestamps.
Mar 10, 2026
- v1.3.0v1.3.0 — WeChat 公众号 + Windows 修复
Adds WeChat Official Account support and resolves Windows-specific compatibility issues.
Mar 4, 2026
- v3.8.1
This release addresses alignment model directory handling and adds support for accessing gated Hugging Face models via a token argument.
Feb 14, 2026
- v3.8.0
The project migrates its speaker diarization pipeline to pyannote-audio version 4 to improve speaker identification capabilities.
Feb 13, 2026