v0.1.0
NVIDIA/NeMo-Speech.cppv0.1.0Aug 19, 2026by pskrunner14
AI Summary
NeMo-Speech.cpp v0.1.0 is the first versioned release, introducing a comprehensive toolkit for speech processing including ASR, TTS, and translation. Key features include a unified CLI, OpenAI-compatible HTTP API, native C API integration, and support for multiple hardware backends.
Key Highlights
- Support for diverse speech models (ASR, TTS, Translation, Diarization)
- OpenAI-compatible HTTP API and WebSocket transcription capabilities
- Native C API integration for application development
- Multi-backend support (CPU, CUDA, Vulkan, Metal)
New Features
- Speech recognition with Nemotron 3.5, Nemotron Speech Streaming, Parakeet TDT, and Parakeet CTC
- Speaker diarization with Streaming Sortformer v2
- Speech synthesis with MagpieTTS and NanoCodec
- Text and speech translation with Riva Translate 4B Instruct v2
- Unified nemo-speech CLI with managed model discovery, downloads, and caching
- Local HTTP API with OpenAI-compatible transcription and playground
- Native C API, shared libraries, and exported CMake package
- Support for CPU, CUDA, Vulkan, and Metal backends
Full Release Notes
`0.1.0` is the first versioned release of NeMo-Speech.cpp. ## Highlights - Speech recognition with Nemotron 3.5 ASR, Nemotron Speech Streaming, Parakeet TDT, and Parakeet CTC - Speaker diarization with Streaming Sortformer v2, standalone or integrated with ASR - Speech synthesis with MagpieTTS and NanoCodec - Text and speech translation with Riva Translate 4B Instruct v2 - Subtitles, VAD, endpointing, punctuation, capitalization, and structured output - Unified nemo-speech CLI with managed model discovery, downloads, and caching - Local HTTP API with OpenAI-compatible transcription and speech subsets, realtime WebSocket transcription, and playground - Native C API, shared libraries, and exported CMake package for application integration - CPU, CUDA, Vulkan, and Metal backend support Prebuilt archives cover Linux and Windows on x86-64 and ARM64, plus Apple Silicon macOS, using the platform-appropriate CPU, CUDA, Vulkan, or Metal backend. Note: this is an early release and API interfaces may evolve.