v1.0.0
tt-a1i/archifyv1.0.0Sep 6, 2026by andimarafioti
AI Summary
This is the first major release introducing a unified Realtime engine, explicit `serve` and `local` commands, expanded speech backend support, and improved conversation lifecycle handling.
Key Highlights
- Unified command structure with three distinct commands: `serve`, `talk`, and `local`.
- Expanded backend support including Qwen3-ASR, OpenAI-compatible endpoints, vLLM, OmniVoice, and Supertonic.
- Improved reliability for real-time conversations with ordered output, cancellation support, and session cleanup.
- Updated Apple Silicon setup with refreshed MLX dependencies and optimized default models.
- Enhanced diagnostics and logging options, including opt-in full transcript logging.
Breaking Changes
- Command invocation changes: `speech-to-speech` is now `speech-to-speech serve`; `--mode realtime` is deprecated; `--mode socket` and `raw-websocket` are removed.
- Removed redundant extras: `facebook-mms`, `language-detection`, and `websocket` are no longer installable as extras.
- Default hosted LLM changed from previous versions to `gpt-5.6-terra`.
- `--local_mac_optimal_settings` no longer selects a command; users must explicitly choose between `serve` or `local`.
New Features
- New commands: `serve` (server), `talk` (client), and `local` (local tool execution with `--tool-module`).
- New speech backends: Qwen3-ASR, OpenAI-compatible STT/TTS, stateful OpenAI Realtime transcription, vLLM Realtime, OmniVoice, and Supertonic.
- New logging feature: `--log_transcripts` flag for full transcript logging.
- New audio feature: `--local_audio_block_mic_during_playback` to pause microphone capture during playback.
- Improved reliability features: ordered text and audio output, remote-request cancellation, and better session cleanup.
Full Release Notes
# speech-to-speech v1.0.0 The first major release brings a single Realtime engine, explicit server and microphone-client commands, more speech backends, and improved conversation lifecycle handling. These notes cover changes since v0.2.12. ## Highlights - **Three commands: `serve`, `talk`, and `local`.** Run the Realtime server, connect the packaged microphone/speaker client, or start both in one terminal. The client also supports local tool execution with `--tool-module`. - **More speech backends.** Add Qwen3-ASR speech recognition, OpenAI-compatible STT and TTS endpoints, stateful OpenAI Realtime transcription, experimental vLLM Realtime transcription, and optional OmniVoice and Supertonic TTS. - **More reliable realtime conversations.** Improve ordered text, audio, and tool-call output; transcript and content-part events; interrupted responses; remote-request cancellation; and session cleanup. The browser demo uses the OpenAI Agents SDK, with integration tests for its WebSocket and WebRTC transports against the implemented core Realtime event set. - **Updated Apple Silicon setup.** Refresh the MLX dependencies and default to the 4-bit `mlx-community/Qwen3-4B-Instruct-2507-4bit` LLM with 6-bit Qwen3-TTS. The Mac preset supplies defaults while respecting explicit overrides. - **Clearer setup and diagnostics.** The README provides local Mac, local NVIDIA, and hosted-LLM starting configurations, a shared llama.cpp recipe, offline setup guidance, and a speaker-feedback workaround. Full transcript logging is now opt-in with `--log_transcripts`; operational logs omit conversation text by default. ## Upgrading from v0.2.12 In your Python environment: ```bash pip install --upgrade "speech-to-speech==1.0.0" ``` Update scripts and service definitions to use the explicit commands: | Previous invocation | v1.0.0 invocation | |---|---| | `speech-to-speech` | `speech-to-speech serve` | | `speech-to-speech --mode realtime` | `speech-to-speech serve` | | `speech-to-speech --mode local` | `speech-to-speech local` | | `speech-to-speech --local_mac_optimal_settings` | `speech-to-speech local --mac-optimal-settings` | | `python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765` | `speech-to-speech talk --url ws://127.0.0.1:8765/v1/realtime` | The `--mode realtime` and `--mode local` forms still work temporarily with a deprecation warning. Other mode values, including `socket` and `raw-websocket`, have been removed. Migrate raw PCM clients to the Realtime WebSocket or WebRTC API; the old raw transport flags are no longer accepted. The `--mac-optimal-settings` preset no longer selects a command: use `local` for microphone/speaker interaction or `serve` for a server. `serve` binds to `127.0.0.1` by default; set `--host` explicitly for network access. `local` uses loopback only. The default hosted LLM is now `gpt-5.6-terra`, with reasoning effort `none`. Pin `--model_name` if your deployment depends on a particular model. Local Mac users can likewise override the preset's model with `--model_name`. The redundant `facebook-mms`, `language-detection`, and `websocket` installation extras were removed; their dependencies are included in the standard package. Install `speech-to-speech[webrtc]` for WebRTC, `speech-to-speech[omnivoice]` for OmniVoice, or `speech-to-speech[supertonic]` for Supertonic. Known options for inactive backends are accepted but ignored with a warning. ## Compatibility and setup - Python 3.10+ is supported; the installation smoke tests use Python 3.11 on Linux and Apple Silicon. Check the README's CUDA wheel guidance for Linux. - Realtime compatibility covers the documented core event set, not every OpenAI Realtime API feature. See the [protocol reference](https://github.com/huggingface/speech-to-speech/blob/v1.0.0/src/speech_to_speech/api/openai_realtime/README.md). - `--local_audio_block_mic_during_playback` pauses microphone capture during assistant playback to reduce speaker feedback. It disables spoken interruptions during playback; it does not perform acoustic echo cancellation. See the [README](https://github.com/huggingface/speech-to-speech/blob/v1.0.0/README.md) for starting configurations and the [full changelog](https://github.com/huggingface/speech-to-speech/compare/v0.2.12...v1.0.0) for all changes and contributors.