v6.2

Vanessa219/vditorv6.2Nov 6, 2025by snakers4

AI Summary

This release focuses on enhancing the Voice Activity Detection (VAD) model's robustness and accuracy, specifically targeting difficult audio scenarios like muted speech, unusual voices, and lower-quality audio.

Key Highlights

  • Reworked training paradigm for improved model quality and stability.
  • Higher stability on out-of-distribution and rare data.
  • Significant quality improvements on muted voices and speech.
  • Enhanced handling of child voices, cartoon voices, and phone calls.

New Features

  • Improved handling of muted voices and speech
  • Better performance on lower quality audio inputs
  • Enhanced stability on unusual and unique data sets

Full Release Notes

- VAD training paradigm reworked;
- Overall slight quality improvement (no metrics update);
- Higher stability on OOD / rare / strange / unique data;
- Significant quality improvements on various known edge cases:
  - Unusual voices
  - Child voices
  - Cartoon voices
  - Muted voices
  - Muted speech
  - Lower quality phone calls

## What's Changed

* Fix type hint for min_silence_at_max_speech (float -> int) by @Purfview in https://github.com/snakers4/silero-vad/pull/714
* Adamnsandle by @adamnsandle in https://github.com/snakers4/silero-vad/pull/717
* Adamnsandle by @adamnsandle in https://github.com/snakers4/silero-vad/pull/719

**Full Changelog**: https://github.com/snakers4/silero-vad/compare/v6.1...v6.2