v17.10.0
ocrmypdf/OCRmyPDFv17.10.0Aug 5, 2026by github-actions[bot]
AI Summary
The watcher helper has been modernized with a new library and strict security checks to prevent running from the same directory as the interpreter. It also fixes handling of password-protected files and removes unnecessary JPEG re-encoding for newer Ghostscript versions.
Key Highlights
- Switched from `watchdog` to `watchfiles` for better filesystem notifications.
- Harvard architecture security enforcement to prevent data/code overlap with interpreter paths.
- Password-protected PDFs no longer crash the watcher; failures are logged and watching continues.
- Ghostscript 10.7.0+ no longer triggers JPEG re-encoding workaround.
Breaking Changes
- Existing deployments co-locating data with the interpreter must relocate data or the application.
New Features
- watchfiles integration
- Harvard architecture enforcement
Full Release Notes
- The `watcher.py` watched-folder helper (the `watcher` extra) has been
modernized and security-hardened:
- It now uses `watchfiles` instead of `watchdog`. Installing
`ocrmypdf[watcher]` now pulls in `watchfiles`; native OS filesystem
notifications are used by default, with `OCR_USE_POLLING=1` to force
polling.
- It enforces a "Harvard architecture" separation between data and code:
at startup it refuses to run (exit code 9) if the input, output or
archive directory overlaps any Python interpreter path (`sys.path`,
the virtual environment, site-packages, or `$PATH`), if
`OCR_JSON_SETTINGS` points at a file inside a data directory or one that
is group/world-writable, or if it specifies a plugin located inside a
data directory. It also refuses to run when the output or archive
directory is the input directory or a subdirectory of it, which would
otherwise cause OCRmyPDF output to be reprocessed in an endless loop.
- At runtime it no longer follows symlinks or processes non-regular files
(fifos, devices, etc.) in the watched directory, and refuses to write
output onto a destination occupied by a non-regular file.
- A password-protected PDF dropped into the watched folder no longer stops
the watcher ({issue}`1715`). `pikepdf.PasswordError` does not derive from
`pikepdf.PdfError`, so it escaped the handler that waits for a file to be
fully written and tore down the watch loop, leaving files that arrived
afterwards unprocessed. Encrypted files are now logged and skipped
immediately — no amount of retrying will supply the password. More
generally, no per-file error can stop the watcher now: failures are
logged and watching continues. Thanks @christophdb for the report and a
fix ({issue}`1716`).
See the "Watcher security model" section of the batch processing
documentation for details. Existing
deployments where the data directories are kept separate from the application
are unaffected; deployments that co-located data with the interpreter or its
environment will need to relocate one or the other.
- Ghostscript 10.7.0 and later are no longer treated as affected by the JPEG
passthrough truncation bug, which Ghostscript fixed in 10.07.0
({issue}`1726`). The version check had no upper bound, so users on a fixed
Ghostscript still saw the "JPEG encoding errors" warning and, worse, silently
had every JPEG lossily re-encoded at `--optimize 1` (the default) to work
around a bug their Ghostscript did not have. Thanks @zuentec-droid for the
detailed measurements and upstream analysis.
- The same JPEG re-encoding workaround no longer applies when Ghostscript did
not produce the file at all. It was previously triggered by the mere presence
of an affected Ghostscript, so `--output-type pdf` and files converted by the
speculative PDF/A path — neither of which runs Ghostscript — paid the quality
loss for nothing.