2.1.0

alibaba/MNN2.1.0Aug 31, 2022by jxt1234

AI Summary

MNN 2.1.0 is a significant release focusing on framework generality, performance optimization, and model compression. Key improvements include CUDA performance enhancements via cutlass-based matrix multiplication and Winograd algorithms, Eager mode support in MNN-Express, and Winograd Int8 optimization for quantized convolutions. The release also includes Metal shader compilation improvements and various bug fixes.

Key Highlights

  • CUDA performance improvement through cutlass-based matrix multiplication and Winograd algorithm optimization for convolutions
  • MNN-Express now supports Eager mode (default in Python, Lazy in C++) for direct result computation without saving computation graphs
  • MNN-CoreML backend now supports zero-copy operations
  • Winograd Int8 optimization support for quantized convolutions with kernel_size > 1
  • MNN-CV added solvepnp/svd function implementations

Breaking Changes

  • Default removal of TFlite Uint8 operator support (can be re-enabled with MNN_SUPPORT_DEPRECATED_OP macro)
  • Removed libTorch library from Linux Torchscript compilation - now downloads from network during compilation

New Features

  • MNN-CV: Added solvepnp and svd function implementations
  • MNN-Train: Added derivative implementations for Unary/Binary/Reduction operations
  • MNN-Express: Eager mode support for direct computation without saving computation graphs
  • New Markdown+Sphinx based documentation system
  • Winograd Int8 optimization for quantized convolutions with configurable offline quantization tool
  • MNN Metal: Online shader compilation to avoid MNN.metallib compatibility issues
  • ScatterND operator now implemented based on Loop operator

Full Release Notes

# 一、框架通用性
- MNN-CV 增加 solvepnp / svd 等函数实现
- MNN-Train 补充 Unary / Binary / Reduction 的求导实现
- MNN-Express 支持 Eager 模式,该模式下不保存计算图,直接计算结果,可通过 Executor 的 lazyEval 配置
   - 在C++中默认使用Lazy模式
   - 在Python中默认使用Eager模式
- 新增基于Markdown+Sphinx的文档

# 二、性能优化

- 服务端推理 CUDA 性能提升
   - 基于 cutlass 重新实现了矩阵乘,对卷积应用 Winograd算法优化;
![image](https://user-images.githubusercontent.com/5484403/187678826-683fcf10-0eda-466d-8a2e-ab2da06437aa.png)


- MNN-CoreML 后端支持免拷贝
![image](https://user-images.githubusercontent.com/5484403/187679144-54f1b8af-c74c-4a2f-a5d6-aa888ba0eaf0.png)


# 三、模型压缩

- 支持 Winograd Int8对`kernel_size > 1`的量化卷积进行优化 ,离线量化工具(C++: `quantized.out`,python: `mnnquant`)json配置文件中增加`"winogradOpt": true`,并将特征量化方法设置为`"feature_quantize_method":"EMA"`即可使用

![image](https://user-images.githubusercontent.com/5484403/187679192-e1aff6e4-cc1e-43bc-aba3-a32c87cfb841.png)


# 四、其他
- 进行了部分代码重构(包括但不限于)
   - MNN Metal 改为在线生成 Shader 编译,避免集成 MNN.metallib 的兼容性问题
   - 移除 CPU / Geometry  / Arm82 部分冗余代码
   - 默认移除原先 TFlite - Uint8 的算子支持,但可以通过 MNN_SUPPORT_DEPRECATED_OP 宏打开
   - 移除 linux 系统下编译 Torchscript 所需要的 libTorch 库,改为编译时从网络下载
   - ScatterND 改为基于 Loop 算子实现
- 修复了如下 Bug(包括但不限于)
   - CPU - AVX512 int8 在转换 NC4HW4 格式时内存访问越界
   - GeometryBinary 处理 NC4HW4 输入,两边Channel上对齐大小相同,但Channel不同时计算出错
   - Arm82 Interp 算子多 Batch 情况下计算出错问题
   - Windows 上 json2MNN 工具写入结果有误
   - Onnx GatherElement 算子在输入大小不确定时转换失败
   - 修复Python中Module,RuntimeManager内存泄露问题
   - 修复控制流模型转换时输出算子Name不匹配的问题
   - 修正 ROIPooling / ROIAlign 低精度计算 crash 的问题