v17.10.0
macro-inc/macrov17.10.0Aug 5, 2026by github-actions[bot]
AI Summary
The `watcher.py` helper is modernized with `watchfiles` and hardened against security risks, while Ghostscript version handling is refined to prevent unnecessary JPEG re-encoding.
Key Highlights
- Switched from `watchdog` to `watchfiles` for filesystem notifications.
- Implemented a 'Harvard architecture' security model to prevent data and code overlap.
- The watcher no longer crashes on password errors or generic per-file errors.
- Ghostscript 10.7.0+ no longer triggers the JPEG re-encoding workaround.
Breaking Changes
- The watcher now exits with code 9 (instead of 0) if security checks fail.
- Deployments where data is co-located with the interpreter are no longer supported.
New Features
- Native OS filesystem notifications via `watchfiles`.
- Security hardening checks for input, output, and archive directories.
- Improved error resilience allowing the watcher to continue processing after file errors.
Full Release Notes
- The `watcher.py` watched-folder helper (the `watcher` extra) has been
modernized and security-hardened:
- It now uses `watchfiles` instead of `watchdog`. Installing
`ocrmypdf[watcher]` now pulls in `watchfiles`; native OS filesystem
notifications are used by default, with `OCR_USE_POLLING=1` to force
polling.
- It enforces a "Harvard architecture" separation between data and code:
at startup it refuses to run (exit code 9) if the input, output or
archive directory overlaps any Python interpreter path (`sys.path`,
the virtual environment, site-packages, or `$PATH`), if
`OCR_JSON_SETTINGS` points at a file inside a data directory or one that
is group/world-writable, or if it specifies a plugin located inside a
data directory. It also refuses to run when the output or archive
directory is the input directory or a subdirectory of it, which would
otherwise cause OCRmyPDF output to be reprocessed in an endless loop.
- At runtime it no longer follows symlinks or processes non-regular files
(fifos, devices, etc.) in the watched directory, and refuses to write
output onto a destination occupied by a non-regular file.
- A password-protected PDF dropped into the watched folder no longer stops
the watcher ({issue}`1715`). `pikepdf.PasswordError` does not derive from
`pikepdf.PdfError`, so it escaped the handler that waits for a file to be
fully written and tore down the watch loop, leaving files that arrived
afterwards unprocessed. Encrypted files are now logged and skipped
immediately — no amount of retrying will supply the password. More
generally, no per-file error can stop the watcher now: failures are
logged and watching continues. Thanks @christophdb for the report and a
fix ({issue}`1716`).
See the "Watcher security model" section of the batch processing
documentation for details. Existing
deployments where the data directories are kept separate from the application
are unaffected; deployments that co-located data with the interpreter or its
environment will need to relocate one or the other.
- Ghostscript 10.7.0 and later are no longer treated as affected by the JPEG
passthrough truncation bug, which Ghostscript fixed in 10.07.0
({issue}`1726`). The version check had no upper bound, so users on a fixed
Ghostscript still saw the "JPEG encoding errors" warning and, worse, silently
had every JPEG lossily re-encoded at `--optimize 1` (the default) to work
around a bug their Ghostscript did not have. Thanks @zuentec-droid for the
detailed measurements and upstream analysis.
- The same JPEG re-encoding workaround no longer applies when Ghostscript did
not produce the file at all. It was previously triggered by the mere presence
of an affected Ghostscript, so `--output-type pdf` and files converted by the
speculative PDF/A path — neither of which runs Ghostscript — paid the quality
loss for nothing.