runtime-llamacpp-v0.1.4
VoltAgent/voltagentruntime-llamacpp-v0.1.4Jun 29, 2026by github-actions[bot]
AI Summary
This release provides prebuilt self-contained binaries for the FunASR llama.cpp runtime, enabling on-device ASR for SenseVoice, Paraformer, and Fun-ASR-Nano without requiring Python or build steps.
Key Highlights
- Prebuilt self-contained binaries for FunASR llama.cpp/GGUF runtime
- Supports SenseVoice, Paraformer, and Fun-ASR-Nano with built-in FSMN-VAD
- Includes CLI tools for SenseVoice and Paraformer
- Offers x64 and x64-avx2 CPU asset options for compatibility
New Features
- Self-contained binaries for SenseVoice
- Self-contained binaries for Paraformer
- Self-contained binaries for Fun-ASR-Nano
- Built-in FSMN-VAD support
- Command-line interface tools (llama-funasr-cli, etc.)
Full Release Notes
Prebuilt self-contained binaries for the FunASR llama.cpp / GGUF runtime — SenseVoice, Paraformer and Fun-ASR-Nano with built-in FSMN-VAD (a whisper.cpp-style on-device ASR, strong on Chinese). Get a model with `bash download-funasr-model.sh <sensevoice|paraformer|nano>`, then run `llama-funasr-cli` / `llama-funasr-sensevoice` / `llama-funasr-paraformer`. Use the default x64 asset for maximum CPU compatibility; use the x64-avx2 asset on CPUs with AVX2/FMA/F16C/BMI2 for higher throughput. No Python, no build. Docs: runtime/llama.cpp/README.md