3.6.1
guilinhu/proactive_hearing_assistant3.6.1Jul 23, 2026by github-actions[bot]
AI Summary
MNN 3.6.1 focuses on the new Qualcomm Hexagon NPU backend, Transformer C4 Fuse optimization across all backends, and various stability fixes. It introduces direct programming support for Hexagon NPU and significantly enhances Metal and OpenCL performance for LLM inference.
Key Highlights
- Hexagon NPU Backend: New direct-programming backend offering 7.9x faster prefill speeds than CPU.
- Transformer C4 Fuse: Fused LLM operators (RoPE/Attention/LayerNorm) enabled across CPU, Metal, OpenCL, and CUDA.
- Metal Optimizations: New Fused Q4/Q8 dequant+GEMM kernels and Op Profile tool.
- MNN-aware QLoRA: New capability for QLoRA finetuning specifically for MNN deployment.
New Features
- Hexagon NPU Backend (DSP kernel support)
- Fused LLM Operators (RoPE, Attention, LayerNorm)
- MNN-aware QLoRA fine-tuning
- Metal Op Profile tool
- llm_bench -fa option for Flash Attention
Full Release Notes
# MNN 3.6.1 版本发布说明 **发布日期**: 2026年7月 📦 **安装**: `pip install mnn==3.6.1` | [PyPI](https://pypi.org/project/mnn/3.6.1/) ## 📌 版本概述 MNN 3.6.1 版本聚焦于 **Hexagon NPU 后端**、**Transformer C4 Fuse 全后端性能优化** 与 **稳定性修复** 三大核心主题: * **Hexagon NPU 后端**:新增高通 Hexagon NPU 直接编程后端,Qwen3-0.6B(W4 对称量化,block64)在骁龙 8 Elite(Hexagon v79)上 prefill 达 2667 tok/s,为 CPU 的 7.9×。 * **Transformer C4 Fuse 全后端性能优化**:新增 fused LLM 算子(RoPE / Attention / LayerNorm 等 C4 路径)并在 CPU / Metal / OpenCL / CUDA 全后端启用 Transformer C4 图;优化 C4 导出与 packed RoPE 路径;Metal 新增 LLM prefill/decode 优化 kernel 与 Fused Q4/Q8 dequant+GEMM(M_TILE=64);OpenCL 修复 NC4HW4 raster L2 cache-set 抖动并优化 int4 GEMM。 * **稳定性修复与工具增强**:新增 MNN-aware QLoRA 微调;修复 mmap 分配失败崩溃、推测解码 KVCache 索引、Arm82 LinearAttention fp16 精度等多项问题。 --- ## 🚀 版本亮点 * **Hexagon 后端(高通 NPU 直接编程)**:Qwen3-0.6B 在骁龙 8 Elite 上 pp512 达 2667 tok/s(CPU 的 7.9×,OpenCL GPU 的 1.8×) * **Transformer C4 图全后端启用**:fused LLM 算子 + C4 后端路径覆盖 CPU / Metal / OpenCL / CUDA,包含 packed RoPE fuse 优化 * **Metal LLM Prefill 大幅优化**:Fused Q4/Q8 dequant+GEMM kernel、M_TILE=64 变体(面向 M5)、prefill/decode kernel 与 shape/RoPE 基础设施重构,并新增 Metal Op Profile * **MNN-aware QLoRA 微调**:新增面向 MNN 部署的 QLoRA 微调能力 --- ## ✨ 新功能 ### 后端:Hexagon(新增) * **Hexagon NPU 后端**:新增高通 Hexagon NPU 直接编程后端(HexagonBackend / HexagonRuntime / DSP kernel),支持 Attention 等 LLM 关键算子;附带 `skills/hexagon` 开发文档 ### LLM / 推理加速 * **Fused LLM 算子与 C4 后端路径**:新增 CPURoPE、SharedGather 等 fused 算子及 CPU/Metal C4 计算路径 * **C4 导出与 packed RoPE 优化**:优化 C4 图导出;CPU / Metal / OpenCL 的 RoPE fuse(packed RoPE)路径优化 * **Transformer C4 图全后端启用**:CPU / Metal / OpenCL / CUDA(新增 RoPEExecution.cu、LayerNorm/LinearAttention C4 适配)统一启用 C4 图 * **MNN-aware QLoRA 微调**:新增 QLoRA 微调支持 ### GPU 后端 * **Metal LLM prefill/decode kernel 优化**:prefill/decode kernel、shape/RoPE 基础设施与 MetalConvolution1x1 SIMD-group 路径重构,新增 Metal Op Profile 工具 * **Metal Fused Q4/Q8 dequant+GEMM**:新增融合反量化+GEMM kernel 及 M_TILE=64 变体,优化 M5 上的 LLM prefill * **OpenCL raster 与 int4 GEMM 优化**:转置遍历顺序修复 NC4HW4 L2 cache-set 抖动;优化 int4 GEMM 性能;拆分 binary_layernorm_c4 kernel * **Vulkan Range 算子**:新增 Range 算子 image 实现 * **QNN Interp/Resize 算子**:QNN 后端新增 Interp/Resize 支持 ### CPU 后端 * **RVV CommonOptFunction 适配**:为 RISC-V RVV 架构适配 CommonOptFunction ### 工具 * **llm_bench `-fa` 选项**:新增 flash attention 开关(默认 1,启用 flash attention) --- ## 🐛 缺陷修复 ### LLM / VLM * **Qwen3.5 fused 推理正确性**:修复 Qwen3.5 fused 推理结果错误 * **C4 fuse 保护与 Metal raster**:为 C4 fuse 增加 guard 条件并修复 Metal raster 问题 * **RoPE schema 编译错误**:修复 rope schema 引入的编译错误 * **MPS 大 tensor Q4 量化静默全零**:修复在 MPS 上量化大 tensor 时 Q4 权重静默全零的问题 * **通用 embedding scale 导出**:修复导出时 embedding scale 缺失 * **WikiText 校准数据集加载**:修复量化校准数据集加载失败 * **QNN 导出 chunk_size**:修复 QNN prefill IO 生成未遵循 `--chunk_size` 的问题 * **decoder embedding 导出 causal mask**:decoder 类 embedding 模型导出改用 causal mask ### GPU 后端 * **Metal ATTENTION_C4 输出分支缺失**:修复 prefill_qkv_tensor 缺少 ATTENTION_C4 输出分支 * **Metal attention kv_cache flag**:修复 Metal attention 未遵循 kv_cache flag * **OpenCL attention c4 kernel 编译**:修复 attention c4 的 OpenCL kernel 编译错误 * **OpenCL c4 长 prefill attention**:修复 GEMM tune 时 V-buffer 越界与 NC4HW4 输出布局缺失 * **OpenCL maskless attention setArg 溢出**:修复导致 Adreno 上 SD UNet 输出错误的 setArg 溢出(issue #4653),同时修复 LinearAttention 越界写、共享 runtime gpuMode 被覆盖及 shared_gather 构建错误 * **CUDA pooling pads 与 TopKV2 largest flag**:修复显式 pooling pads 与 TopKV2 largest flag 处理,并重组相关测试 * **QNN mmpose 模型**:修复运行 mmpose 模型时的若干问题 * **QNN context 生成栈溢出**:提高 context 生成的栈上限 ### CPU 后端 * **Arm82 LinearAttention fp16 精度**:改用 fp32 累加器,修复 fp16 精度损失 * **推测解码 KVCache 索引**:修复 CPU KVCache value 索引错误导致 Eagle3 输出质量下降(issue #4591) * **ConvInt8 SIGSEGV**:修复 Conv 缺少量化权重时 ConvInt8 崩溃 * **RVV 源文件缺失 Macro.h**:补全 RISC-V RVV 源文件的头文件包含 * **RVV Int8 函数修复**:修复 MNNLineDepthWiseInt8AddBiasScaleUnit 等 RVV int8 函数问题(issue #4454) ### 框架 / 其它 * **mmap 分配失败崩溃**:修复各后端 mmap 分配失败时的崩溃 * **RuntimeManager::setCache 空指针**:增加 null runtime 保护 * **Android 构建缺失源文件**:CMakeLists 补充 diffusion_session 与 sana_session 源文件 --- ## 📚 其他改进 * **开发流程文档**:新增 development workflow guidelines * **CI 脚本整合**:整合 CI 测试脚本与文档 * **Shader 构建脚本现代化**:支持复杂路径 * **NeuroPilot 代码移出公开仓库**:相关实现改由内部维护 * **iOS 发布包更名**:移除包名中的 CoreML 后缀 * **Copybara 同步脚本可配置化** * **PyMNN 打包更新**:更新 3.6.1 发布的 PyMNN 打包配置 --- ## 📊 性能基准 <img width="3900" height="1650" alt="combined_bench" src="https://github.com/user-attachments/assets/26920ba4-d0df-4df8-af46-8085964a4001" /> 相比 3.6.0(同设备、同量化配置,模型用 3.6.1 重新导出): * **Android 性能大幅回升**:Qwen3-4B OpenCL prefill (pp512) **+101.7%**,Qwen3-4B CPU decode (tg128) **+72.2%**,Qwen3-0.6B OpenCL decode **+76.1%** * **macOS Metal 全线提升**:Qwen3-0.6B prefill (pp1024) **+50.1%**、decode **+36.1%**;Qwen3-4B prefill (pp512) **+18.8%** > 测试环境: Mac mini M4 (16GB) / OnePlus PJZ110 (Snapdragon 8 Elite) | 模型: Qwen3-0.6B / Qwen3-4B | 量化: W4 (block=64) | 工具: llm_bench (rep=5, median) ### Hexagon NPU 后端性能 Qwen3-0.6B(4-bit 对称量化,block64,C4)在 SM8750(骁龙 8 Elite,Hexagon v79)上三后端对比: | 后端 | pp512 (tok/s) | tg128 (tok/s) | vs CPU prefill | |------|--------------|--------------|----------------| | Hexagon NPU | 2667 | 68.5 | **7.9×** | | OpenCL GPU | 1491 | 71.4 | 4.4× | | CPU | 336 | 94.3 | 1.0× | Hexagon NPU 在 prefill 上优势显著(7.9×),decode 目前仍以 CPU 最优。Hexagon SDK 5.x 与 6.x 编译的 DSP skeleton 性能一致(差异 <2%,误差范围内)。 --- ## 🙏 致谢 衷心感谢所有贡献者对本版本的宝贵贡献(按 GitHub login 列出): * **@jxt1234** — Hexagon 后端 (#4633)、fused LLM 算子与 C4 后端路径 (#4554) * **@wangzhaode** — Transformer C4 图全后端启用 (#4629)、C4 导出与 packed RoPE 优化 (#4598)、QLoRA 微调 (#4560)、Qwen3.5 fused 推理修复 (#4631)、C4 fuse guard 与 Metal raster 修复 (#4607)、Vulkan Range 算子 (#4563)、Metal kv_cache flag 修复 (#4556)、embedding scale 导出修复 (#4615) * **@v0jiuqi** — Metal LLM prefill/decode kernel 与 RoPE 基础设施 (#4650)、Fused Q4/Q8 dequant+GEMM (#4651)、llm_bench `-fa` 选项 (#4658)、Metal ATTENTION_C4 分支修复 (#4654)、MPS Q4 全零权重修复 (#4649)、Arm82 LinearAttention fp16 修复 (#4634)、mmap 失败崩溃修复 (#4630) * **@Qxinyu** — OpenCL NC4HW4 raster / int4 GEMM 优化 (#4606)、c4 长 prefill attention 修复 (#4609)、maskless attention setArg 溢出等修复 (#4659)、attention c4 kernel 编译修复 (#4644)、QNN mmpose 修复 (#4626)、setCache 空指针保护 (#4617) * **@Fatemanx (xly)** — CUDA pooling pads 与 TopKV2 largest flag 修复 (#4627) * **@LudovicoYIN** — QNN 导出 chunk_size 修复 (#4611)、WikiText 校准数据集修复 (#4566)、QNN context 生成栈上限 * **@huluhuluu** — 推测解码 KVCache 索引修复 (#4600) * **@lanling-47** — QNN Interp/Resize 算子 (#4582) * **@jxgxxx** — RVV CommonOptFunction 适配 (#4426) * **@Sherlockzhangjinge** — RVV Int8 函数修复 (#4552) * **@xuhaonan999** — RVV Macro.h 包含修复 (#4588) * **@pingji2016** — Android 构建源文件修复 (#4569) * **@RinZ27** — Shader 构建脚本现代化 (#4550) * **@devin-lai** — decoder embedding causal mask 修复 (#4548) --- # MNN 3.6.1 Release Notes **Release Date**: July 2026 📦 **Install**: `pip install mnn==3.6.1` | [PyPI](https://pypi.org/project/mnn/3.6.1/) ## 📌 Overview MNN 3.6.1 focuses on three core themes: **Hexagon NPU Backend**, **Transformer C4 Fuse Optimization Across Backends**, and **Stability Fixes**: * **Hexagon NPU Backend**: New Qualcomm Hexagon NPU direct-programming backend. Qwen3-0.6B (W4 symmetric, block64) reaches 2667 tok/s prefill on Snapdragon 8 Elite (Hexagon v79), 7.9× faster than CPU. * **Transformer C4 Fuse Optimization Across Backends**: New fused LLM ops (RoPE / Attention / LayerNorm C4 paths) with Transformer C4 graphs enabled across CPU / Metal / OpenCL / CUDA; optimized C4 export and packed RoPE (RoPE fuse) paths; Metal adds fused Q4/Q8 dequant+GEMM (M_TILE=64) and reworked LLM prefill/decode kernels; OpenCL fixes NC4HW4 raster L2 cache-set thrashing and optimizes int4 GEMM. * **Stability Fixes & Tooling**: New MNN-aware QLoRA finetuning; fixes for mmap allocation failure crashes, speculative-decoding KVCache indexing, Arm82 LinearAttention fp16 accuracy, and more. --- ## 🚀 Highlights * **Hexagon Backend (Qualcomm NPU Direct Programming)**: Qwen3-0.6B pp512 reaches 2667 tok/s on Snapdragon 8 Elite (7.9× vs CPU, 1.8× vs OpenCL GPU) * **Transformer C4 Graphs Across Backends**: Fused LLM ops + C4 backend paths across CPU / Metal / OpenCL / CUDA, including packed RoPE fuse optimization * **Metal LLM Prefill Optimization**: Fused Q4/Q8 dequant+GEMM kernel, M_TILE=64 variant (for M5), reworked prefill/decode kernels with shape/RoPE infrastructure, plus a new Metal Op Profile tool * **MNN-aware QLoRA Finetuning**: New QLoRA finetuning targeting MNN deployment --- ## ✨ New Features ### Backend: Hexagon (New) * **Hexagon NPU Backend**: New Qualcomm Hexagon NPU direct-programming backend (HexagonBackend / HexagonRuntime / DSP kernels) with LLM key ops such as Attention; ships with `skills/hexagon` development docs ### LLM / Inference Acceleration * **Fused LLM Ops & C4 Backend Paths**: New fused ops (CPURoPE, SharedGather, etc.) and CPU/Metal C4 compute paths * **C4 Export & Packed RoPE Optimization**: Optimized C4 graph export; packed RoPE (RoPE fuse) paths across CPU / Metal / OpenCL * **Transformer C4 Graphs Enabled Across Backends**: CPU / Metal / OpenCL / CUDA (new RoPEExecution.cu, LayerNorm/LinearAttention C4 adaptation) * **MNN-aware QLoRA Finetuning** ### GPU Backends * **Metal LLM prefill/decode Kernel Optimization**: Reworked prefill/decode kernels, shape/RoPE infrastructure, MetalConvolution1x1 SIMD-group paths, and a new Metal Op Profile tool * **Metal Fused Q4/Q8 dequant+GEMM**: New fused dequant+GEMM kernel with M_TILE=64 variant for LLM prefill on M5 * **OpenCL Raster & int4 GEMM Optimization**: Fixed NC4HW4 L2 cache-set thrashing by transposing traversal order; optimized int4 GEMM; split binary_layernorm_c4 kernel * **Vulkan Range Op**: New image implementation * **QNN Interp/Resize Ops**: Added to QNN backend ### CPU Backend * **RVV CommonOptFunction Adaptation**: Adapted CommonOptFunction for RISC-V RVV ### Tools * **llm_bench `-fa` Option**: New flash attention toggle (default 1, flash attention enabled) --- ## 🐛 Bug Fixes ### LLM / VLM * **Qwen3.5 fused inference correctness**: Fixed * **C4 fuse guard & Metal raster**: Added guard conditions for C4 fuse and fixed Metal raster * **RoPE schema compile error**: Fixed * **Silent all-zero Q4 weights on MPS**: Fixed when quantizing large tensors * **Generic embedding scale export**: Fixed missing embedding scale on export * **WikiText calibration dataset loading**: Fixed * **QNN export chunk-size**: Honor `--chunk_size` when generating QNN prefill IO * **Decoder embedding export causal mask**: Use causal mask for decoder embedding export ### GPU Backends * **Metal missing ATTENTION_C4 output branch**: Fixed in prefill_qkv_tensor * **Metal attention kv_cache flag**: Now respected * **OpenCL attention c4 kernel compile**: Fixed * **OpenCL c4 long-prefill attention**: Fixed V-buffer OOB in GEMM tune and missing NC4HW4 output layout * **OpenCL maskless attention setArg overflow**: Fixed setArg overflow causing wrong SD UNet output on Adreno (issue #4653), plus LinearAttention OOB writes, shared runtime gpuMode clobbering, and shared_gather build error * **CUDA pooling pads & TopKV2 largest flag**: Fixed explicit pooling pads and TopKV2 largest flag handling; reorganized related tests * **QNN mmpose model**: Fixed several runtime bugs * **QNN context generation stack limit**: Raised ### CPU Backend * **Arm82 LinearAttention fp16 accuracy**: Use fp32 accumulators * **Speculative decoding KVCache indexing**: Fixed CPU KVCache value index degrading Eagle3 output quality (issue #4591) * **ConvInt8 SIGSEGV**: Fixed crash when Conv lacks quantized weight * **Missing Macro.h in RVV sources**: Added * **RVV Int8 function fixes**: Fixed MNNLineDepthWiseInt8AddBiasScaleUnit and related RVV int8 functions (issue #4454) ### Core / Misc * **mmap allocation failure crashes**: Fixed across all backends * **RuntimeManager::setCache null runtime**: Guarded * **Android build missing sources**: Added diffusion_session and sana_session to CMakeLists --- ## 📚 Other Improvements * **Development workflow guidelines**: Added docs * **CI test scripts consolidation** * **Shader build scripts modernization**: Handle complex paths * **NeuroPilot implementation removed from public repository** * **iOS release package renamed**: Removed CoreML suffix * **Configurable Copybara sync script** * **PyMNN packaging update**: Updated PyMNN packaging for the 3.6.1 release --- ## 📊 Performance Benchmarks <img width="3900" height="1650" alt="combined_bench" src="https://github.com/user-attachments/assets/26920ba4-d0df-4df8-af46-8085964a4001" /> Compared with 3.6.0 (same devices and quantization config, models re-exported with 3.6.1): * **Major Android recovery**: Qwen3-4B OpenCL prefill (pp512) **+101.7%**, Qwen3-4B CPU decode (tg128) **+72.2%**, Qwen3-0.6B OpenCL decode **+76.1%** * **macOS Metal improvements across the board**: Qwen3-0.6B prefill (pp1024) **+50.1%**, decode **+36.1%**; Qwen3-4B prefill (pp512) **+18.8%** > Environment: Mac mini M4 (16GB) / OnePlus PJZ110 (Snapdragon 8 Elite) | Models: Qwen3-0.6B / Qwen3-4B | Quantization: W4 (block=64) | Tool: llm_bench (rep=5, median) ### Hexagon NPU Backend Performance Qwen3-0.6B (4-bit symmetric, block64, C4) on SM8750 (Snapdragon 8 Elite, Hexagon v79), three-backend comparison: | Backend | pp512 (tok/s) | tg128 (tok/s) | vs CPU prefill | |------|--------------|--------------|----------------| | Hexagon NPU | 2667 | 68.5 | **7.9×** | | OpenCL GPU | 1491 | 71.4 | 4.4× | | CPU | 336 | 94.3 | 1.0× | Hexagon NPU excels at prefill (7.9×); CPU remains fastest for decode. DSP skeletons built with Hexagon SDK 5.x and 6.x perform identically (<2% difference, within noise). --- ## 🙏 Acknowledgements We sincerely thank all contributors for their valuable contributions to this release (listed by GitHub login): * **@jxt1234** — Hexagon backend (#4633), fused LLM ops & C4 backend paths (#4554) * **@wangzhaode** — Transformer C4 graphs across backends (#4629), C4 export & packed RoPE optimization (#4598), QLoRA finetuning (#4560), Qwen3.5 fused inference fix (#4631), C4 fuse guard & Metal raster fix (#4607), Vulkan Range op (#4563), Metal kv_cache flag fix (#4556), embedding scale export fix (#4615) * **@v0jiuqi** — Metal LLM prefill/decode kernels & RoPE infrastructure (#4650), fused Q4/Q8 dequant+GEMM (#4651), llm_bench `-fa` option (#4658), Metal ATTENTION_C4 branch fix (#4654), MPS all-zero Q4 weights fix (#4649), Arm82 LinearAttention fp16 fix (#4634), mmap failure crash fix (#4630) * **@Qxinyu** — OpenCL NC4HW4 raster / int4 GEMM optimization (#4606), c4 long-prefill attention fix (#4609), maskless attention setArg overflow & related fixes (#4659), attention c4 kernel compile fix (#4644), QNN mmpose fixes (#4626), setCache null guard (#4617) * **@Fatemanx (xly)** — CUDA pooling pads & TopKV2 largest flag fix (#4627) * **@LudovicoYIN** — QNN export chunk-size fix (#4611), WikiText calibration fix (#4566), QNN context generation stack limit * **@huluhuluu** — Speculative decoding KVCache index fix (#4600) * **@lanling-47** — QNN Interp/Resize ops (#4582) * **@jxgxxx** — RVV CommonOptFunction adaptation (#4426) * **@Sherlockzhangjinge** — RVV Int8 function fixes (#4552) * **@xuhaonan999** — RVV Macro.h include fix (#4588) * **@pingji2016** — Android build sources fix (#4569) * **@RinZ27** — Shader build scripts modernization (#4550) * **@devin-lai** — Decoder embedding causal mask fix (#4548) --- **Full Changelog**: https://github.com/alibaba/MNN/compare/3.6.0...3.6.1