v0.1.6
Gen-Verse/MMaDAv0.1.6May 26, 2026by afourney
AI Summary
Introduced OCR capabilities and an Azure Content Understanding converter while resolving memory leaks and HTML parsing errors.
Key Highlights
- Added OCR layer service for extracting text from images and PDF scans
- Implemented Azure Content Understanding converter
- Fixed memory growth in PDF conversion by properly closing pages
- Resolved RecursionError caused by deeply nested HTML
New Features
- Azure Content Understanding converter
- OCR support for images and PDFs
Full Release Notes
## What's Changed * [MS] Add OCR layer service for embedded images and PDF scans by @lesyk in https://github.com/microsoft/markitdown/pull/1541 * Fix O(n) memory growth in PDF conversion by calling page.close() afte… by @lesyk in https://github.com/microsoft/markitdown/pull/1612 * Updated warning about binding to non-local interfaces. by @afourney in https://github.com/microsoft/markitdown/pull/1653 * fix: handle deeply nested HTML that triggers RecursionError by @jigangz in https://github.com/microsoft/markitdown/pull/1644 * Clarify security posture in READMEs by @afourney in https://github.com/microsoft/markitdown/pull/1807 * feat: Add Azure Content Understanding converter by @chienyuanchang in https://github.com/microsoft/markitdown/pull/1865 * Bump version to 0.1.6 by @afourney in https://github.com/microsoft/markitdown/pull/1914 ## New Contributors * @jigangz made their first contribution in https://github.com/microsoft/markitdown/pull/1644 * @chienyuanchang made their first contribution in https://github.com/microsoft/markitdown/pull/1865 **Full Changelog**: https://github.com/microsoft/markitdown/compare/v0.1.5...v0.1.6