runtime-llamacpp-v0.1.2
QwenAudio/SenseVoiceruntime-llamacpp-v0.1.2Jun 21, 2026by LauraGPT
AI Summary
First prebuilt, self-contained binaries for running SenseVoice locally without Python or build steps, featuring built-in FSMN-VAD. This release introduces support for q8 GGUF models to reduce size without sacrificing accuracy.
Key Highlights
- New support for q8 GGUF models (~half the size of f16 with similar accuracy).
- No Python, no build, no torch required.
- Strong performance on Chinese and Cantonese.
- Supports Linux (x64/arm64), macOS (arm64), and Windows (x64).
New Features
- q8 GGUF model support
- Zero-dependency standalone binaries
- FSMN-VAD integration
Full Release Notes
Prebuilt self-contained binaries for running **SenseVoice** (and Paraformer / Fun-ASR-Nano) locally with the FunASR llama.cpp / GGUF runtime — built-in FSMN-VAD, whisper.cpp-style on-device ASR, strong on Chinese & Cantonese. **New:** q8 GGUF models are ~half the size of f16 with the same accuracy. ```bash bash download-funasr-model.sh sensevoice ./gguf llama-funasr-sensevoice -m ./gguf/sensevoice-small-q8.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav ``` No Python, no build. Linux (x64/arm64), macOS (arm64), Windows (x64). Docs: `runtime/llama.cpp/README.md`.