v0.3.103
shivangi-aneja/GaussianSpeechv0.3.103Apr 19, 2025by KoljaB
AI Summary
Introduces robust improvements to the RealtimeSTT library, focusing on thread safety, audio normalization, and asynchronous event handling to ensure consistent transcription quality.
Key Highlights
- Thread-safe IPC using `SafePipe`
- Audio normalization to -0.95 dBFS
- Asynchronous event callbacks for all event types
- Nanosecond-precision timestamps
- CLI flags for VAD filtering and debugging
New Features
- SafePipe replacement for `mp.Pipe`
- Audio normalization option
- Async event callback overhaul
- Rich metadata with nanosecond precision
- Faster Whisper VAD configuration
- Debug WebSocket flags
Full Release Notes
# RealtimeSTT 0.3.103 ### Features & Improvements - **Thread‑safe IPC**: Introduce `SafePipe` to replace `mp.Pipe`, hopefully ensuring robust inter-process communication (needs more tests). - **Audio normalization**: New `normalize_audio` option scales input to –0.95 dBFS for consistent transcription quality. - **Callback overhaul**: All event callbacks (VAD, wake‑word, turn detection, recording, realtime updates) now run asynchronously via helper threads. - **Wake word & VAD**: Add `wakeword_backend` config and `faster_whisper_vad_filter` flag; improved error messages when misconfigured. - **Rich metadata**: Embed nanosecond‑precision timestamps in both client and server, serialized as formatted strings. - **CLI enhancements**: `--faster_whisper_vad_filter` and `--debug_websockets` flags give finer control over server behavior. - **Testing updates**: Adjusted parameters in `realtimestt_test` and added a new `type_into_textbox.py` example.