MNN Releases
86 releases of alibaba/MNN
- v1.121.0Hyperswitch v1.121.0
This release focuses on infrastructure updates for Hyperswitch, including significant enhancements to the UCS tunnel, connector capabilities, and authentication services. It adds support for various payment methods like BNPL and wallets while improving analytics and webhook integrations.
Feb 24, 2026
- 3.4.0
MNN 3.4.0 focuses on deepening GPU/QNN backend capabilities, optimizing Attention computation and long-context memory usage, and improving GPU runtime stability. Key improvements include Vulkan LLM support with CoopMat acceleration, CPU/Metal Flash Attention, KV Cache quantization, and new unified attention_mode configuration.
Feb 7, 2026
- v2.36.0
This release introduces CSV upload support for Bulk Send and significantly expands the Get Template API response with detailed metadata. It also adds Korean language support and resolves several rendering and interaction issues in the document editor.
Feb 6, 2026
- v2.35.0
This release focuses on enhancing document management workflows, introducing new preferences for Date Widgets, and improving prefill capabilities for both API and UI.
Jan 7, 2026
- v2.34.0
This update improves the signing experience with UI enhancements like stamps and text inputs, alongside backend improvements including webhook authentication and SSO support.
Dec 22, 2025
- 3.3.03.3.0 NPU 支持 / SME2 指令加速 / EAGLE 投机解码加速
MNN 3.3.0 introduces NPU support for Qualcomm QNN and MTK, SME2 instruction acceleration for Armv9 devices, and EAGLE-3 speculative decoding achieving 2.24x speedup. The release also adds HQQ quantization, new model support, and CUDA backend LLM capabilities.
Oct 31, 2025
- 3.2.03.2.0 MoE 架构和Omni支持 / 投机解码实现 / 新增模板引擎 / KlediAI 功能更新 / 启动性能优化
MNN 3.2 introduces MoE architecture and Qwen Omni support, speculative decoding for 2-3x decoding efficiency improvement, new jinja template engine, KlediAI integration updates, and significant startup performance optimizations.
Jun 6, 2025
- 3.1.03.1.0 新增LLM移动端应用
MNN 3.1.0 introduces new mobile LLM applications for Android and iOS, significant performance improvements in CPU prefill and GPU inference, and expanded model support including Qwen2-VL/Audio and DeepSeek-R1.
Feb 24, 2025
- 3.0.03.0发布【多模态大模型支持、动态量化功能完善、Web SIMD支持及其他Bug修复】
MNN 3.0 is a major release introducing comprehensive LLM inference support, Stable Diffusion support, dynamic quantization for arbitrary convolution types, and WebAssembly SIMD support.
Nov 20, 2024
- 2.9.0【LLM相关性能优化,形状缓存机制】2.9.0
MNN 2.9.0 officially integrates the MNN-LLM module with transformer-related operators and graph optimization, introduces shape cache mechanism for speech/text models, and adds HarmonyOS support.
May 15, 2024
- 2.8.1
Dec 29, 2023
- 2.8.0
MNN 2.8.0 adds JSON format for quantization parameters, new operators, significant LLM CPU performance improvements through dynamic quantization, and various bug fixes.
Dec 5, 2023
- 2.7.2
Dec 4, 2023
- 2.6.3
Sep 28, 2023
- 2.7.1
Sep 28, 2023
- 2.7.0【内存分配优化、性能优化、Bugfix】2.7.0
MNN 2.7.0 focuses on memory allocation optimization with DeferAllocator (19.13% average reduction), new LayerNorm int8 quantization, CUDA TopKV2 support, and various performance improvements.
Sep 4, 2023
- 2.6.0【新增Int8 量化算子,OpenCL后端适配 recordable queue】2.6.0
MNN 2.6.0 adds comprehensive int8 quantization operator support, OpenCL recordable queue adaptation for Qualcomm GPUs, and low memory inference mode supporting ChatGLM-6B in 3GB.
Jul 5, 2023
- 2.5.0
MNN 2.5.0 adds extensive new OpenCV operators, Tflite int8 quantization support, CUDA bf16 support, and significant performance optimizations using Intel subgroup extensions.
Apr 27, 2023
- 2.4.02.4.0 NNAPI后端/CUDA后端支持量化模型
MNN 2.4.0 is a significant release introducing NNAPI int8 quantization support, online operator Fuse for OpenCL/Metal backends, and automated CI/CD via GitHub Actions. The release also includes major CUDA optimizations for Softmax/DepthwiseConv, OpenCL 3x3 convolution improvements, and removes deprecated LLVMJit/Codegen backends.
Mar 1, 2023
- 2.3.02.3.0 基于几何计算实现求导/支持模型权重分离模式
MNN 2.3.0 introduces significant backend improvements including CUDA high-precision mode (FP32) and SM60 support, OpenCL low power mode, and experimental Vulkan Buffer memory layout. The release also adds MNN-Train derivative optimizations with CONTENT mode for geometric computation, experimental support for separating model structure from weights, and includes various performance optimizations and bug fixes.
Dec 30, 2022