v0.0.9

srbhr/Resume-Matcherv0.0.9Jun 26, 2026by paulpierre

AI Summary

This release resolves six open issues, including critical bugs related to encoding and link handling, and introduces community-requested features like path exclusion and configurable markdown heading styles. It also upgrades the project structure to support modern Python packaging standards and adds a comprehensive test suite.

Key Highlights

  • Fixed `UnicodeEncodeError` by enforcing UTF-8 encoding on all file writes.
  • Added `exclude_paths` parameter to filter URLs during crawling and CLI usage.
  • Configured markdownify to support various heading styles (`ATX`, `UNDERLINE`, etc.).
  • Implemented a comprehensive test suite with 60 tests and 95% code coverage.
  • Updated minimum Python version to 3.8 and migrated `pyproject.toml` to PEP 621 standards.

Breaking Changes

  • Minimum Python version bumped to 3.8.
  • Configuration structure updated in `pyproject.toml` from `[tool.poetry.scripts]` to `[project.scripts]`.

New Features

  • `exclude_paths` parameter for URL filtering in `crawl()`, `get_target_links()`, `worker()`, `md_crawl()`, and CLI.
  • Configurable `heading_style` for markdownify (`ATX`, `ATX_CLOSED`, `UNDERLINE`, `SETEXT`).
  • Support for `uvx` via proper `dependencies` in `pyproject.toml`.
  • Developer dependencies (`pytest`, `pytest-cov`) added for contributors.

Full Release Notes

## What's New

This release resolves **6 open issues** and introduces several community-requested features.

### Fixed
- **#7** — `UnicodeEncodeError` on non-ASCII content. All file writes now use `encoding='utf-8'`.
- **#8** — Bypass JavaScript checks. A browser-like `User-Agent` header is now sent on all `requests.get` calls.
- **#10** — `target_content` CSS selector handling. Null `href` links are now skipped in `get_target_links`.
- **#11** — `UnboundLocalError` in `get_target_content` when no tags are found. `main_content` is now initialized to `None` and checked before `str()` conversion.

### Added
- **#9** — `exclude_paths` parameter for URL filtering. Available in `crawl()`, `get_target_links()`, `worker()`, `md_crawl()`, and the CLI (`-x` / `--exclude-paths`).
- **#20** — Configurable `heading_style` for markdownify (`ATX`, `ATX_CLOSED`, `UNDERLINE`, `SETEXT`). CLI flag: `-s` / `--heading-style`.
- **#17** — Proper `dependencies` in `pyproject.toml` (`beautifulsoup4`, `markdownify`, `requests`) enabling `uvx` support out of the box.
- Comprehensive test suite — 60 tests with 95% coverage across `__init__.py` (96%) and `cli.py` (88%).
- `dev` optional dependencies (`pytest`, `pytest-cov`) for contributors.

### Changed
- Minimum Python version bumped to 3.8.
- `pyproject.toml` now uses `[project.scripts]` instead of `[tool.poetry.scripts]`.

**Full changelog:** https://github.com/paulpierre/markdown-crawler/blob/main/CHANGELOG.md