runtime-llamacpp-v0.1.9

luminal-ai/luminalruntime-llamacpp-v0.1.9Jul 24, 2026by LauraGPT

AI Summary

This release provides prebuilt self-contained binaries for the FunASR llama.cpp runtime, supporting SenseVoice, Paraformer, and Fun-ASR-Nano with built-in FSMN-VAD. A critical security update moves the SenseVoice source to the current main branch to fix a YAML loader vulnerability.

Key Highlights

  • Prebuilt self-contained binaries for llama.cpp runtime
  • Support for SenseVoice, Paraformer, and Fun-ASR-Nano models
  • Security fix for YAML loader vulnerability in SenseVoice
  • Multiple backend options including AVX2, Vulkan, and CUDA

New Features

  • Built-in FSMN-VAD (Voice Activity Detection)
  • Vulkan backend support
  • CUDA backend support targeting architecture 86

Full Release Notes

Prebuilt self-contained binaries for the FunASR llama.cpp / GGUF runtime, with SenseVoice, Paraformer, and Fun-ASR-Nano support plus built-in FSMN-VAD.

This release moves the SenseVoice **Latest Release** source snapshot onto current `main`, including the `yaml.SafeLoader` security fix from [#312](https://github.com/QwenAudio/SenseVoice/pull/312). Users downloading GitHub's generated source archives no longer receive the older unsafe YAML loader from `runtime-llamacpp-v0.1.4`.

## Choose an asset

- Default `linux-x64` / `windows-x64`: maximum CPU compatibility.
- `linux-x64-avx2` / `windows-x64-avx2`: higher throughput on CPUs with AVX2, FMA, F16C, and BMI2.
- `linux-x64-vulkan` / `windows-x64-vulkan`: Vulkan backend; requires a working Vulkan driver and ICD.
- `windows-x64-cuda`: CUDA backend targeting architecture 86; build from source for other GPU architectures.
- Native `linux-arm64` and `macos-arm64` packages are included.

Download a default quantized model with:

```bash
pip install -U huggingface_hub
bash download-funasr-model.sh sensevoice
```

Then run `llama-funasr-sensevoice`; no Python ASR runtime or local compilation is required.

All nine files are byte-for-byte identical to the digest-verified assets from [`modelscope/FunASR` runtime v0.1.9](https://github.com/modelscope/FunASR/releases/tag/runtime-llamacpp-v0.1.9).

Documentation: https://github.com/QwenAudio/SenseVoice/blob/runtime-llamacpp-v0.1.9/runtime/llama.cpp/README.md