runtime-llamacpp-v0.2.5
modelscope/FunASRruntime-llamacpp-v0.2.5Aug 29, 2026by github-actions[bot]
AI Summary
This runtime release fixes a critical Vulkan weight upload bug and provides updated prebuilt binaries for various platforms including Vulkan and CUDA.
Key Highlights
- Fix for Vulkan host-to-device weight upload.
- Support for Vulkan backend on Linux and Windows.
- CUDA runtime support targeting architecture 86.
- Prebuilt binaries for SenseVoice, Paraformer, and Fun-ASR-Nano.
New Features
- Vulkan host-to-device weight upload fix.
- Vulkan backend support.
- CUDA architecture 86 support.
Full Release Notes
Prebuilt self-contained binaries for the FunASR llama.cpp / GGUF runtime: SenseVoice, Paraformer and Fun-ASR-Nano with built-in FSMN-VAD. Download the default quantized model with `bash download-funasr-model.sh <sensevoice|paraformer|nano>` (the helper requires the Hugging Face CLI: `pip install -U huggingface_hub`), then run `llama-funasr-cli` / `llama-funasr-sensevoice` / `llama-funasr-paraformer`. ## v0.2.5 boundary This release fixes Vulkan host-to-device weight upload so model tensors are copied into the backend buffer before graph execution. The exact release commit is `f371370d4c5e4c61d13d4eb9c55cda2f4dd95e4f`. Local Linux validation passed for Vulkan llvmpipe with Q8 and F16 weights. This does **not** establish that the reported AMD Windows hardware crash is fixed. Windows AMD users should use `funasr-llamacpp-windows-x64-vulkan.zip`, capture the new initialization-boundary logs, and report the GPU, driver version, model/quantization, command, and complete output on [FunASR #3479](https://github.com/modelscope/FunASR/issues/3479). ## Platform guidance - Use the default x64 asset for maximum CPU compatibility. - Use x64-avx2 on CPUs with AVX2/FMA/F16C/BMI2 for higher throughput. - Vulkan assets are `linux-x64-vulkan` and `windows-x64-vulkan`; they require a working Vulkan driver/ICD and enable SenseVoiceSmall graph execution with `llama-funasr-sensevoice ... --backend vulkan`. - The Windows CUDA asset targets CUDA architecture 86 and requires a compatible NVIDIA driver. Build from source for other GPU architectures. - Build from source with `-DGGML_VULKAN=ON` to validate platform-specific GPU stacks. No Python ASR runtime or local build is required. [Runtime documentation](https://github.com/modelscope/FunASR/blob/runtime-llamacpp-v0.2.5/runtime/llama.cpp/README.md) ## SHA256 ```text 4350ed49b1220374231ddafc84fed2dbe6a1d5e19d830edcb6a252d19d78b5db funasr-llamacpp-linux-arm64.tar.gz f9bb3fe78776ef86feb334b02488d3b1a5c16926a9a1cc79b3d5b1219ef1fb30 funasr-llamacpp-linux-x64-avx2.tar.gz 88cd52ba1ce92af9bf344944521a9d47ed9ceb424b312a5993369228c1e8c4fb funasr-llamacpp-linux-x64-vulkan.tar.gz 69d47bf724216093bedc89a50efe0a97c80a2f4b89367378a72d2d3d94e4f241 funasr-llamacpp-linux-x64.tar.gz c6b3cea2c5238ae48175c931365231b1c47351f95ca289455536de0a959648af funasr-llamacpp-macos-arm64.tar.gz 6f19b8af0ce623122c767a0160bb8faabbc9666ae4f0cef1675cdf6250a9ea26 funasr-llamacpp-windows-x64-avx2.zip 3bb704e6b86f2e2e613aba84c361a51273e729c107183c7b69a1f2781c95cce8 funasr-llamacpp-windows-x64-cuda.zip d51e34afed66c0898dba3d944cb17d2abc92c5fb7758fa4ee376561d5ce35a4c funasr-llamacpp-windows-x64-vulkan.zip dc90e85bc9a22477c5a8b556427fa5a6e0caf8951b5b6ec0237317406ce11bf5 funasr-llamacpp-windows-x64.zip ```