runtime-llamacpp-v0.1.1

docling-project/doclingruntime-llamacpp-v0.1.1Jun 21, 2026by LauraGPT

AI Summary

A prebuilt, self-contained binary release for the FunASR llama.cpp runtime, enabling local execution of SenseVoice and Paraformer models with built-in VAD and optimized performance for Chinese and Cantonese.

Key Highlights

  • Prebuilt, self-contained binaries for Linux, macOS, and Windows
  • Strong performance on Chinese and Cantonese (8.01% CER)
  • Built-in FSMN-VAD for speech activity detection
  • whisper.cpp-style on-device ASR

New Features

  • Self-contained binaries requiring no Python, build, or torch
  • SenseVoice and Paraformer/Fun-ASR-Nano support
  • FSMN-VAD integration
  • whisper.cpp-style on-device ASR
  • Optimized performance for Chinese and Cantonese languages

Full Release Notes

Prebuilt, self-contained binaries to run **SenseVoice** (and Paraformer / Fun-ASR-Nano) locally with the FunASR llama.cpp / GGUF runtime — built-in FSMN-VAD, whisper.cpp-style on-device ASR, strong on Chinese & Cantonese.

```bash
bash download-funasr-model.sh sensevoice ./gguf
llama-funasr-sensevoice -m ./gguf/SenseVoiceSmall-f16.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav
```

No Python, no build, no torch. Binaries for Linux (x64/arm64), macOS (arm64), Windows (x64). Docs: `runtime/llama.cpp/README.md`. CER (micro-avg): SenseVoice 8.01% vs whisper.cpp small 22%.