v0.1.6
PaddlePaddle/PaddleOCRv0.1.6May 26, 2026by afourney
AI Summary
This release introduces new OCR capabilities for embedded images and PDFs, adds an Azure Content Understanding converter, and fixes several bugs including memory leaks and HTML parsing errors.
Key Highlights
- Added OCR layer service for embedded images and PDF scans
- Added Azure Content Understanding converter
- Fixed O(n) memory growth in PDF conversion
- Fixed RecursionError for deeply nested HTML
New Features
- OCR layer service for embedded images and PDF scans
- Azure Content Understanding converter
Full Release Notes
## What's Changed * [MS] Add OCR layer service for embedded images and PDF scans by @lesyk in https://github.com/microsoft/markitdown/pull/1541 * Fix O(n) memory growth in PDF conversion by calling page.close() afte… by @lesyk in https://github.com/microsoft/markitdown/pull/1612 * Updated warning about binding to non-local interfaces. by @afourney in https://github.com/microsoft/markitdown/pull/1653 * fix: handle deeply nested HTML that triggers RecursionError by @jigangz in https://github.com/microsoft/markitdown/pull/1644 * Clarify security posture in READMEs by @afourney in https://github.com/microsoft/markitdown/pull/1807 * feat: Add Azure Content Understanding converter by @chienyuanchang in https://github.com/microsoft/markitdown/pull/1865 * Bump version to 0.1.6 by @afourney in https://github.com/microsoft/markitdown/pull/1914 ## New Contributors * @jigangz made their first contribution in https://github.com/microsoft/markitdown/pull/1644 * @chienyuanchang made their first contribution in https://github.com/microsoft/markitdown/pull/1865 **Full Changelog**: https://github.com/microsoft/markitdown/compare/v0.1.5...v0.1.6