v17.10.0

macro-inc/macrov17.10.0Aug 5, 2026by github-actions[bot]

AI Summary

The `watcher.py` helper is modernized with `watchfiles` and hardened against security risks, while Ghostscript version handling is refined to prevent unnecessary JPEG re-encoding.

Key Highlights

  • Switched from `watchdog` to `watchfiles` for filesystem notifications.
  • Implemented a 'Harvard architecture' security model to prevent data and code overlap.
  • The watcher no longer crashes on password errors or generic per-file errors.
  • Ghostscript 10.7.0+ no longer triggers the JPEG re-encoding workaround.

Breaking Changes

  • The watcher now exits with code 9 (instead of 0) if security checks fail.
  • Deployments where data is co-located with the interpreter are no longer supported.

New Features

  • Native OS filesystem notifications via `watchfiles`.
  • Security hardening checks for input, output, and archive directories.
  • Improved error resilience allowing the watcher to continue processing after file errors.

Full Release Notes

- The `watcher.py` watched-folder helper (the `watcher` extra) has been
  modernized and security-hardened:
    - It now uses `watchfiles` instead of `watchdog`. Installing
      `ocrmypdf[watcher]` now pulls in `watchfiles`; native OS filesystem
      notifications are used by default, with `OCR_USE_POLLING=1` to force
      polling.
    - It enforces a "Harvard architecture" separation between data and code:
      at startup it refuses to run (exit code 9) if the input, output or
      archive directory overlaps any Python interpreter path (`sys.path`,
      the virtual environment, site-packages, or `$PATH`), if
      `OCR_JSON_SETTINGS` points at a file inside a data directory or one that
      is group/world-writable, or if it specifies a plugin located inside a
      data directory. It also refuses to run when the output or archive
      directory is the input directory or a subdirectory of it, which would
      otherwise cause OCRmyPDF output to be reprocessed in an endless loop.
    - At runtime it no longer follows symlinks or processes non-regular files
      (fifos, devices, etc.) in the watched directory, and refuses to write
      output onto a destination occupied by a non-regular file.
    - A password-protected PDF dropped into the watched folder no longer stops
      the watcher ({issue}`1715`). `pikepdf.PasswordError` does not derive from
      `pikepdf.PdfError`, so it escaped the handler that waits for a file to be
      fully written and tore down the watch loop, leaving files that arrived
      afterwards unprocessed. Encrypted files are now logged and skipped
      immediately — no amount of retrying will supply the password. More
      generally, no per-file error can stop the watcher now: failures are
      logged and watching continues. Thanks @christophdb for the report and a
      fix ({issue}`1716`).

  See the "Watcher security model" section of the batch processing
  documentation for details. Existing
  deployments where the data directories are kept separate from the application
  are unaffected; deployments that co-located data with the interpreter or its
  environment will need to relocate one or the other.
- Ghostscript 10.7.0 and later are no longer treated as affected by the JPEG
  passthrough truncation bug, which Ghostscript fixed in 10.07.0
  ({issue}`1726`). The version check had no upper bound, so users on a fixed
  Ghostscript still saw the "JPEG encoding errors" warning and, worse, silently
  had every JPEG lossily re-encoded at `--optimize 1` (the default) to work
  around a bug their Ghostscript did not have. Thanks @zuentec-droid for the
  detailed measurements and upstream analysis.
- The same JPEG re-encoding workaround no longer applies when Ghostscript did
  not produce the file at all. It was previously triggered by the mere presence
  of an affected Ghostscript, so `--output-type pdf` and files converted by the
  speculative PDF/A path — neither of which runs Ghostscript — paid the quality
  loss for nothing.