py-1.43.0

mantinedev/mantinepy-1.43.0Jul 21, 2026by github-actions[bot]

AI Summary

A major release introducing numerous deprecations, performance enhancements like global DNS caching, and new expressions such as `ewm_sum` and trigonometric functions, alongside improved Iceberg and Arrow support.

Key Highlights

  • Multiple deprecations for casting to categoricals, bitwise ops, and `LazyFrame.profile()`.
  • Performance improvements including split multiplexers and global DNS caching.
  • New `ewm_sum` and `ewm_sum_by` expressions.
  • Support for Iceberg V3 tables and `scan_arrow_c_stream`.
  • New `Series.degrees/radians` and `list` expression.

New Features

  • `ewm_sum` and `ewm_sum_by` expressions.
  • `scan_arrow_c_stream` function.
  • `Series.degrees` and `Series.radians`.
  • `list` expression for packing elements into a new List type.
  • `read_database` Arrow fast-path for 'python-oracledb'.
  • `Expr.cat.to` and `Expr.cat.physical` methods.
  • `POLARS_OOC_DISK_BUDGET_MB` environment variable.

Full Release Notes

## ⚠️ Deprecations

- Deprecate casting numeric types to categoricals (#28349)
- Deprecate `cat.get_categories()` and `cat.to_local()` (#28299)
- Deprecate `LazyFrame.profile()` (#28275)
- Deprecate casts from non-nested dtypes into Lists (#28193)
- Deprecate bitwise ops between integers and booleans (#28186)
- Deprecate `list/arr.to_struct()` calls that don't pass field names (#28243)
- Rename `missing_utf8_is_empty_string` to `empty_string_is_null` (#28173)

## 🚀 Performance improvements

- Split multiplexers that directly scan from in-memory DataFrame (#28376)
- Make DNS cache global (#28352)
- Do not remove cache if predicates not pushed to all inputs (#28341)
- Environment variable for logging slow DNS lookup (#28211)
- Pre-partition on left, right and semi joins on hive partitioned data (#28374)
- Rewrite inner joins on hive partitioned data to a union of partitions filtered joins (#28327)
- Tighten redundant comparisons and propagate filter constraints (#27999)
- Canonicalize predicates (#28212)
- Remove excess memory copy in streaming IPC sink (#28141)
- Parallelize scan\_iceberg resolve (#28168)
- Chunk-aware morsel splitting on single-file IO sink (#27950)
- Don't allocate for `full_null()` with small lengths (#28181)
- Use monotonic deque for rolling min\_by/max\_by (O(n)) (#27436)

## ✨ Enhancements

- Pre-partition group-by on hive keys  (#28444)
- Add ewm\_sum and ewm\_sum\_by #28151 (#28215)
- Serve stale records from object\_store DNS cache (#28256)
- Make partition\_hive an opt flag (#28345)
- Add `scan_arrow_c_stream` (#28340)
- When displaying IRExprs, only add parentheses when needed for disambiguation (#28319)
- Support Iceberg field `initial-default` in native `scan_iceberg` (#28324)
- Show only differing dtypes in `assert_frame_equal` error (#27816)
- Expose ewm and rolling \*\_by expressions in Python visitor (#28286)
- More accurate rolling moments (#28309)
- Remove hard block on native scan for Iceberg V3 tables (#28277)
- Add Series.degrees/radians and clarify trig functions use radians (#22961)
- Merge sorted with multikey (#28067)
- Add `list` expression which consistently packs elements together into new List type (#27990)
- Add `read_database` Arrow fast-path for "python-oracledb" (#28230)
- Optimize joins with redundant equi-join keys (#28017)
- Implement missing datetime options for `CastColumnsPolicy` (#28084)
- Add Expr.cat.to and Expr.cat.physical (#28096)
- Add POLARS\_OOC\_DISK\_BUDGET\_MB (#28045)
- Add build side preference parameter for join (#28154)

## 🐞 Bug fixes

- Avoid panic when union slice skips all rows (#28420)
- Solve panic in `dt.replace` when there were multiple chunks (#28437)
- Propagate `is_scalar` from the input to the output of `.sort()` and `.sort_by()` (#28438)
- Resolve CSV column names overwrite in DSL->IR conversion (#28383)
- Avoid IEJoin rewrite for Categorical comparisons (#28427)
- Do not rewrite `sort().reverse()` to `sort(descending=True)` when `maintain_order=True` (#28403)
- Ensure BinaryView offset+len does not exceed i32::MAX where possible (#28048)
- Invalid offset in strptime (#28388)
- Incorrect schema type for decimal \<-> primitive division (#28373)
- Panic in in-memory CSEE handling (#28371)
- Fix offset in arrow ffi export of sliced struct arrays (#28369)
- Fix cross filter not applied with sink and CSE (#28297)
- Fix panic on projection pushdown with caches (#28280)
- Fix `write_json()` null values in `Array` columns being written incorrectly as `null` (#28330)
- Fix EntityTooSmall on sink\_ipc to S3 (#28255)
- Raise error instead of silent wrapping for `select(len())` (#28355)
- Float16 groupby aggregates (#28361)
- Respect lexical ordering of Categorical in `top_k`/`bottom_k` (#28359)
- Support dtype-list schema overrides in scan\_csv (#28226)
- Resolve unknown type in fused multiply add (#28321)
- Honor `aws_checksum_algorithm` when provided (#28328)
- Remove usage of deprecated `pyarrow.feather.read_table` (#28323)
- Fix regression in sourcing HF\_TOKEN (#28317)
- Raise on addition/subtraction between temporal and non-temporal Series (#28266)
- Avoid panic constructing Series from unaligned numpy arrays (#28120)
- Rewrite with\_columns on empty df to select during DSL -> IR (#28298)
- Fix sink deadlock regression (#28296)
- Avoid panic when consumer of vertical concat is done (#28272)
- Harden IPC mmap API (#28258)
- Classify S3 region request failures as IO errors (#28225)
- Incorrect inner dtype in ArrayChunked::from\_aligned\_values (#28249)
- Raise on implode of `Object` dtype instead of creating invalid `List(Object)` (#28246)
- Return null not 0 for the first unbiased `ewm_var`/`ewm_std` value (#28235)
- Database appends with `ADBC` shouldn't require "CREATE" privileges unless the table is confirmed not to exist (#28228)
- Avoid slice pushdown panic on shared cache inputs (#28157)
- Incorrect dtype for Unknown right literal in arithmetic (#28184)
- Remove unnecessary unsafe Python object downcast (#28189)
- Normalize non-native byte order in numpy array construction (#28197)
- Use-after-free in numpy ufunc FFI (#28195)
- Unpickling for pl.when().then() (#28165)
- Reduce unknown casts (#28145)
- Consider a `rolling_rank` position invalid if `ost.len() < min_samples` (#28142)
- Incorrect broadcast with empty len in pl.concat\_list (#28134)

## 📖 Documentation

- Update comparison page (#28418)
- Update GPU support documentation with the cudf-polars 26.06 release (#27830)
- Fix dev docs being canonical for SEO (#28306)
- Document experimental HDFS support on-premises (#28115)
- Add Series.degrees/radians and clarify trig functions use radians (#22961)
- Add polars-bio to community plugins (#28103)
- Add AWS Marketplace walkthrough to Cloud User Guide (#28214)
- Correct `merge_sorted()` docs also for `DataFrame` (#28183)
- Clarify DataFrame.remove combines predicates with AND (#28171)
- Update config reference (#28150)
- Correct that `merge_sorted()` null inputs should be nulls first (#28177)
- Clarify parts of the Polars Cloud and On-prem user guide (#28170)
- Add On-Premise release 0.6.3 to releases page (#28175)
- Update PolarsCloud snippets for 0.9.0 release  (#28166)
- Fix to\_titlecase incorrectly claiming alphanumeric separation (#28136)

## 📦 Build system

- Set the default `dev` profile to `line-tables-only` (#28358)
- Reduce binary size by optimizing for size in the SQL transpilation (#28336)

## 🛠️ Other improvements

- Update AI policy for comments (#28436)
- Fix flaky ordering expectation in top\_k\_by (#28435)
- Ignore all deprecation warnings in Python doctest (#28398)
- Re-work `shuffle` parameter for `sample()` (#27460)
- Ignore deprecation warnings on `to_struct()` in doctest (#28370)
- Update `spin` dependency to `v0.10.1` (#28360)
- Bloat dispatch reduction (#28337)
- Show only differing dtypes in `assert_frame_equal` error (#27816)
- Accept IR in `optimize` (#28318)
- Annotate untyped module constants (#27907)
- Remove unencountered deny entries (#28260)
- Bump crossbeam-epoch dependency (#28259)
- Change `object_store` custom DNS resolver patch to rev-based (#28241)
- Mark shape tests as passing (#28133)
- Attribute annotations for `*GroupBy` (#27903)
- Clarify AI policy, explicitly forbid agents from interacting with our repository (#28206)
- Tighten first-time-contributor screenshot requirements (#28207)
- Covariant `classinstmethod` and attribute annotation (#27904)
- Ignore quickxml DoS rustsec (#28185)
- Update `typos` to `v1.48.0` (#28162)

Thank you to all our contributors for making this release possible!
@0guban0v, @AnirudhRahul, @Bharath-970, @CodingSelim, @EndPositive, @Jesse-Bakker, @Kevin-Patyk, @LarryHu0217, @MarcoGorelli, @MatteoPossamai, @TNieuwdorp, @alex-h-sun, @alexander-beedie, @azimafroozeh, @borchero, @dancsi, @dependabot[bot], @dsprenkels, @harrywhalen, @jorenham, @kdn36, @makinzm, @malhotrashivam, @mdavis-xyz, @mikhail5555, @mkzung, @mroeschke, @nameexhaustion, @orlp, @r-brink, @raphaelroshan, @ritchie46, @tylerriccio33, @useredsa, @wence-, @yannbolliger, @zichen0116 and [dependabot[bot]](https://github.com/apps/dependabot)