v2.0.0
anthropics/anthropic-tokenizer-typescriptv2.0.0Dec 3, 2024by adbar
AI Summary
Major version 2.0.0 of trafilatura with comprehensive updates including Python version deprecations, API changes to bare_extraction(), removal of deprecated GUI, and full type hinting implementation.
Key Highlights
- Python 3.6 and 3.7 deprecated
- bare_extraction() now returns Document class by default
- Type hinting added throughout codebase
- Deprecated GUI removed
- max_tree_size parameter moved to settings.cfg
Breaking Changes
- Python 3.6 and 3.7 deprecated
- bare_extraction() return type changed to Document class
- no_fallback parameter renamed to fast
- decode argument removed from fetch_url()
- Deprecated GUI removed
- max_tree_size parameter moved to settings.cfg
New Features
- Full type hinting implementation
- Enhanced robustness in extraction and downloads
- Improved table markdown output
- CLI improvements with exit codes
Full Release Notes
Breaking changes: - Python 3.6 and 3.7 deprecated (#709) - `bare_extraction()`: - now returns an instance of the `Document` class by default - `as_dict` deprecation warning → use `.as_dict()` method on return value (#730) - `bare_extraction()` and `extract()`: `no_fallback` deprecation warning → use `fast` instead (#730) - downloads: remove `decode` argument in `fetch_url()` → use `fetch_response` instead (#724) - deprecated graphical user interface now removed (#713) - extraction: move `max_tree_size` parameter to `settings.cfg` (#742) - use type hinting (#721, #723, #748) - see [Python](https://trafilatura.readthedocs.io/en/latest/usage-python.html#deprecations) and [CLI](https://trafilatura.readthedocs.io/en/latest/usage-cli.html#deprecations) deprecations in the docs Fixes: - set `options.source` before raising error on empty doc tree by @dmoklaf (#707) - robust encoding in `options.source` (#717) - more robust mapping for conversion to HTML (#721) - CLI downloads: use all information in settings file (#734) - downloads: cleaner urllib3 code (#736) - refine table markdown output by @unsleepy22 (#752) - extraction fix: images in text nodes by @unsleepy22 (#757) Metadata: - more robust URL extraction (#710) Command-line interface: - CLI: print URLs early for feeds and sitemaps with `--list` with @gremid (#744) - CLI: add 126 exit code for high error ratio (#747) Maintenance: - remove already deprecated functions and args (#716) - add type hints (#723, #728) - setup: use `pyproject.toml` file (#715) - simplify code (#708, #709, #727) - better debug messages in `main_extractor` (#714) - evaluation: review data, update packages, add magic_html (#731) - setup: explicit exports through `__all__` (#740) - tests: extend coverage (#753) Documentation: - fix link in `docs/index.html` by @nzw0301 (#711) - remove docs from published packages (#743) - update docs (#745)