v0.3.103

shivangi-aneja/GaussianSpeechv0.3.103Apr 19, 2025by KoljaB

AI Summary

Introduces robust improvements to the RealtimeSTT library, focusing on thread safety, audio normalization, and asynchronous event handling to ensure consistent transcription quality.

Key Highlights

  • Thread-safe IPC using `SafePipe`
  • Audio normalization to -0.95 dBFS
  • Asynchronous event callbacks for all event types
  • Nanosecond-precision timestamps
  • CLI flags for VAD filtering and debugging

New Features

  • SafePipe replacement for `mp.Pipe`
  • Audio normalization option
  • Async event callback overhaul
  • Rich metadata with nanosecond precision
  • Faster Whisper VAD configuration
  • Debug WebSocket flags

Full Release Notes

# RealtimeSTT 0.3.103

### Features & Improvements
- **Thread‑safe IPC**: Introduce `SafePipe` to replace `mp.Pipe`, hopefully ensuring robust inter-process communication (needs more tests).
- **Audio normalization**: New `normalize_audio` option scales input to –0.95 dBFS for consistent transcription quality.
- **Callback overhaul**: All event callbacks (VAD, wake‑word, turn detection, recording, realtime updates) now run asynchronously via helper threads.
- **Wake word & VAD**: Add `wakeword_backend` config and `faster_whisper_vad_filter` flag; improved error messages when misconfigured.
- **Rich metadata**: Embed nanosecond‑precision timestamps in both client and server, serialized as formatted strings.
- **CLI enhancements**: `--faster_whisper_vad_filter` and `--debug_websockets` flags give finer control over server behavior.
- **Testing updates**: Adjusted parameters in `realtimestt_test` and added a new `type_into_textbox.py` example.