v0.1.6

mishushakov/llm-scraperv0.1.6May 26, 2026by afourney

AI Summary

This release adds an OCR layer for processing embedded images and PDF scans, introduces an Azure Content Understanding converter, and fixes memory growth and HTML parsing issues.

Key Highlights

  • Added OCR layer service for embedded images and PDF scans
  • Added Azure Content Understanding converter
  • Fixed O(n) memory growth in PDF conversion
  • Fixed deeply nested HTML handling causing RecursionError
  • Updated security posture and warnings

New Features

  • OCR layer service
  • Azure Content Understanding converter
  • HTML parsing fix for deep nesting

Full Release Notes

## What's Changed
* [MS] Add OCR layer service for embedded images and PDF scans by @lesyk in https://github.com/microsoft/markitdown/pull/1541
* Fix O(n) memory growth in PDF conversion by calling page.close() afte… by @lesyk in https://github.com/microsoft/markitdown/pull/1612
* Updated warning about binding to non-local interfaces. by @afourney in https://github.com/microsoft/markitdown/pull/1653
* fix: handle deeply nested HTML that triggers RecursionError by @jigangz in https://github.com/microsoft/markitdown/pull/1644
* Clarify security posture in READMEs by @afourney in https://github.com/microsoft/markitdown/pull/1807
* feat: Add Azure Content Understanding converter by @chienyuanchang in https://github.com/microsoft/markitdown/pull/1865
* Bump version to 0.1.6 by @afourney in https://github.com/microsoft/markitdown/pull/1914

## New Contributors
* @jigangz made their first contribution in https://github.com/microsoft/markitdown/pull/1644
* @chienyuanchang made their first contribution in https://github.com/microsoft/markitdown/pull/1865

**Full Changelog**: https://github.com/microsoft/markitdown/compare/v0.1.5...v0.1.6