v2.0.0
adbar/trafilaturav2.0.0Dec 3, 2024by adbar
AI Summary
Major version release with significant breaking changes including Python 3.6/3.7 deprecation, refactored bare_extraction() to return Document class instances, removal of deprecated GUI, and extensive type hinting additions.
Key Highlights
- Python 3.6 and 3.7 deprecated
- bare_extraction() now returns Document class by default instead of dict
- Deprecated graphical user interface removed
- Added comprehensive type hinting throughout codebase
- CLI improvements: early URL printing for feeds/sitemaps and exit code 126 for high error ratio
Breaking Changes
- Python 3.6 and 3.7 deprecated
- bare_extraction() returns Document class instead of dict
- no_fallback parameter deprecated → use 'fast' instead
- decode argument removed in fetch_url() → use fetch_response
- Deprecated GUI removed
- max_tree_size parameter moved to settings.cfg
New Features
- Document class for extraction results
- fetch_response替代decode argument
- Type hinting throughout
- CLI --list prints URLs early for feeds/sitemaps
- 126 exit code for high error ratio in CLI
Full Release Notes
Breaking changes: - Python 3.6 and 3.7 deprecated (#709) - `bare_extraction()`: - now returns an instance of the `Document` class by default - `as_dict` deprecation warning → use `.as_dict()` method on return value (#730) - `bare_extraction()` and `extract()`: `no_fallback` deprecation warning → use `fast` instead (#730) - downloads: remove `decode` argument in `fetch_url()` → use `fetch_response` instead (#724) - deprecated graphical user interface now removed (#713) - extraction: move `max_tree_size` parameter to `settings.cfg` (#742) - use type hinting (#721, #723, #748) - see [Python](https://trafilatura.readthedocs.io/en/latest/usage-python.html#deprecations) and [CLI](https://trafilatura.readthedocs.io/en/latest/usage-cli.html#deprecations) deprecations in the docs Fixes: - set `options.source` before raising error on empty doc tree by @dmoklaf (#707) - robust encoding in `options.source` (#717) - more robust mapping for conversion to HTML (#721) - CLI downloads: use all information in settings file (#734) - downloads: cleaner urllib3 code (#736) - refine table markdown output by @unsleepy22 (#752) - extraction fix: images in text nodes by @unsleepy22 (#757) Metadata: - more robust URL extraction (#710) Command-line interface: - CLI: print URLs early for feeds and sitemaps with `--list` with @gremid (#744) - CLI: add 126 exit code for high error ratio (#747) Maintenance: - remove already deprecated functions and args (#716) - add type hints (#723, #728) - setup: use `pyproject.toml` file (#715) - simplify code (#708, #709, #727) - better debug messages in `main_extractor` (#714) - evaluation: review data, update packages, add magic_html (#731) - setup: explicit exports through `__all__` (#740) - tests: extend coverage (#753) Documentation: - fix link in `docs/index.html` by @nzw0301 (#711) - remove docs from published packages (#743) - update docs (#745)