runtime-llamacpp-v0.2.1
modelscope/FunASRruntime-llamacpp-v0.2.1Aug 26, 2026by github-actions[bot]
AI Summary
Introduces prebuilt self-contained binaries for the FunASR llama.cpp runtime, supporting SenseVoice, Paraformer, and Fun-ASR-Nano models with built-in FSMN-VAD. It offers various hardware optimizations including AVX2, Vulkan, and CUDA backends to enhance performance across different platforms.
Key Highlights
- Support for SenseVoice, Paraformer, and Fun-ASR-Nano with built-in FSMN-VAD
- Multiple backend options: AVX2 for CPU throughput, Vulkan for GPU, and CUDA for NVIDIA
- Self-contained binaries requiring no Python runtime or local build
- Download helper script for default quantized models
New Features
- Vulkan backend support for SenseVoiceSmall graph execution
- Windows x64 CUDA backend support targeting CUDA architecture 86
- Built-in FSMN-VAD for voice activity detection
- Prebuilt binaries for Linux, macOS, and Windows platforms
Full Release Notes
Prebuilt self-contained binaries for the FunASR llama.cpp / GGUF runtime: SenseVoice, Paraformer and Fun-ASR-Nano with built-in FSMN-VAD. Download the default quantized model with `bash download-funasr-model.sh <sensevoice|paraformer|nano>` (the helper requires the Hugging Face CLI: `pip install -U huggingface_hub`), then run `llama-funasr-cli` / `llama-funasr-sensevoice` / `llama-funasr-paraformer`. Use the default x64 asset for maximum CPU compatibility; use the x64-avx2 asset on CPUs with AVX2/FMA/F16C/BMI2 for higher throughput. The Vulkan assets are `linux-x64-vulkan` and `windows-x64-vulkan`; they require a working Vulkan driver/ICD and enable SenseVoiceSmall graph execution with `llama-funasr-sensevoice ... --backend vulkan`. Build from source with `-DGGML_VULKAN=ON` to validate platform-specific GPU stacks. The Windows CUDA asset is `windows-x64-cuda`; it requires an NVIDIA driver compatible with the CUDA Toolkit version configured by the release workflow, targets CUDA architecture 86, and enables SenseVoiceSmall graph execution with `llama-funasr-sensevoice ... --backend cuda`. Build from source for other GPU architectures. No Python ASR runtime or local build is required. Docs: https://github.com/modelscope/FunASR/blob/runtime-llamacpp-v0.2.1/runtime/llama.cpp/README.md