v0.1.6
MinishLab/model2vecv0.1.6May 26, 2026by afourney
AI Summary
This update brings significant new capabilities including OCR support for embedded images and PDFs, alongside an Azure Content Understanding converter. It also resolves critical memory leaks and recursion errors in PDF and HTML processing.
Key Highlights
- Added OCR layer service for images and PDF scans
- Added Azure Content Understanding converter
- Fixed O(n) memory growth in PDF conversion
- Fixed RecursionError in deeply nested HTML handling
New Features
- OCR layer service
- Azure Content Understanding converter
Full Release Notes
## What's Changed * [MS] Add OCR layer service for embedded images and PDF scans by @lesyk in https://github.com/microsoft/markitdown/pull/1541 * Fix O(n) memory growth in PDF conversion by calling page.close() afte… by @lesyk in https://github.com/microsoft/markitdown/pull/1612 * Updated warning about binding to non-local interfaces. by @afourney in https://github.com/microsoft/markitdown/pull/1653 * fix: handle deeply nested HTML that triggers RecursionError by @jigangz in https://github.com/microsoft/markitdown/pull/1644 * Clarify security posture in READMEs by @afourney in https://github.com/microsoft/markitdown/pull/1807 * feat: Add Azure Content Understanding converter by @chienyuanchang in https://github.com/microsoft/markitdown/pull/1865 * Bump version to 0.1.6 by @afourney in https://github.com/microsoft/markitdown/pull/1914 ## New Contributors * @jigangz made their first contribution in https://github.com/microsoft/markitdown/pull/1644 * @chienyuanchang made their first contribution in https://github.com/microsoft/markitdown/pull/1865 **Full Changelog**: https://github.com/microsoft/markitdown/compare/v0.1.5...v0.1.6