v2.0.0
ByteDance-Seed/SeedVRv2.0.0Dec 3, 2024by adbar
AI Summary
This major release introduces significant breaking changes to the web content extraction library, including deprecated Python versions, API changes, and removal of deprecated GUI functionality.
Key Highlights
- Python 3.6 and 3.7 are now deprecated
- Bare extraction now returns Document class by default
- Removed deprecated graphical user interface
- Added comprehensive type hinting throughout the codebase
Breaking Changes
- Python 3.6 and 3.7 deprecated
- Bare_extraction() now returns Document class by default
- no_fallback parameter deprecated, use 'fast' instead
- decode argument removed from fetch_url(), use fetch_response instead
- Deprecated GUI removed
- max_tree_size parameter moved to settings.cfg
New Features
- Added CLI exit code for high error ratio
- Improved table markdown output
- Better debug messages in main_extractor
Full Release Notes
Breaking changes: - Python 3.6 and 3.7 deprecated (#709) - `bare_extraction()`: - now returns an instance of the `Document` class by default - `as_dict` deprecation warning → use `.as_dict()` method on return value (#730) - `bare_extraction()` and `extract()`: `no_fallback` deprecation warning → use `fast` instead (#730) - downloads: remove `decode` argument in `fetch_url()` → use `fetch_response` instead (#724) - deprecated graphical user interface now removed (#713) - extraction: move `max_tree_size` parameter to `settings.cfg` (#742) - use type hinting (#721, #723, #748) - see [Python](https://trafilatura.readthedocs.io/en/latest/usage-python.html#deprecations) and [CLI](https://trafilatura.readthedocs.io/en/latest/usage-cli.html#deprecations) deprecations in the docs Fixes: - set `options.source` before raising error on empty doc tree by @dmoklaf (#707) - robust encoding in `options.source` (#717) - more robust mapping for conversion to HTML (#721) - CLI downloads: use all information in settings file (#734) - downloads: cleaner urllib3 code (#736) - refine table markdown output by @unsleepy22 (#752) - extraction fix: images in text nodes by @unsleepy22 (#757) Metadata: - more robust URL extraction (#710) Command-line interface: - CLI: print URLs early for feeds and sitemaps with `--list` with @gremid (#744) - CLI: add 126 exit code for high error ratio (#747) Maintenance: - remove already deprecated functions and args (#716) - add type hints (#723, #728) - setup: use `pyproject.toml` file (#715) - simplify code (#708, #709, #727) - better debug messages in `main_extractor` (#714) - evaluation: review data, update packages, add magic_html (#731) - setup: explicit exports through `__all__` (#740) - tests: extend coverage (#753) Documentation: - fix link in `docs/index.html` by @nzw0301 (#711) - remove docs from published packages (#743) - update docs (#745)