1.6.5
Mintplex-Labs/anything-llm1.6.5Apr 22, 2026by Borda
AI Summary
This release focuses on stability fixes for segmentation training on specific GPU hardware and resolves issues related to custom resolution handling. It includes a breaking change regarding how source images are stored within detection objects and fixes several bugs affecting model validation and checkpoint management.
Key Highlights
- Fixed segmentation training crashes on T4 and P100 GPUs by implementing a custom autograd function for depthwise convolution.
- Fixed EMA segmentation metrics to ensure they are calculated from the correct accumulator rather than base metrics.
- Added automatic interpolation of pretrained positional embeddings to support custom model resolutions.
- Fixed `BestModelCallback` to persist the best EMA score during training resume.
Breaking Changes
- source_image moved from detections.data to detections.metadata
Full Release Notes
## ⚠️ Breaking Changes
- **`source_image` moved from `data` to `metadata`.** `predict()` now stores the original input image in `detections.metadata["source_image"]` instead of `detections.data["source_image"]`. This fixes `IndexError` when boolean- or integer-indexing `sv.Detections` with `include_source_image=True` (the default) — supervision indexes every value in `data` by the detection mask, but passes `metadata` through unchanged. Update any code that reads the source image: (#972)
```python
# Before (1.6.3–1.6.4)
image = detections.data["source_image"]
# After (1.6.5+)
image = detections.metadata["source_image"]
```
## 🔧 Fixed
- **Fixed segmentation training crash on T4 and P100 GPUs.** cuDNN engine selection fails for depthwise convolution backward on some CUDA stacks (Kaggle, Colab). The previous workaround only disabled cuDNN in the forward pass; backward kernels still ran with cuDNN enabled. A custom `autograd.Function` now disables cuDNN in both forward and backward. (#967)
- **Fixed EMA segmentation mAP logged from base metrics.** `ema_segm_mAP_50_95` and `ema_segm_mAP_50` were computed from the base (non-EMA) metric accumulator instead of the EMA accumulator, producing misleading validation scores for segmentation models. Both metrics now use the correct EMA source. (#980)
- **Fixed `BestModelCallback` losing best EMA score on resume.** The `_best_ema` value was not persisted in `state_dict()`, so resuming training reset the EMA best-model tracker and could overwrite a better earlier checkpoint. (#973)
- **Fixed `positional_encoding_size` not updating with custom resolution.** Setting `resolution` at construction time (e.g. `RFDETRLarge(resolution=640)`) did not update the positional encoding grid size, causing shape mismatches during forward. A model validator now auto-syncs `positional_encoding_size` when a custom resolution is provided. (#956)
- **Fixed pretrained weight loading crash with custom resolution.** Loading COCO pretrained weights into a model with a non-default resolution failed because the DINOv2 positional embeddings had an incompatible shape. Pretrained PE tensors are now bicubic-interpolated to match the target grid before loading. (#964)
```python
# This now works — PE is automatically interpolated from 560px grid to 640px
model = RFDETRLarge(resolution=640)
```
- **Fixed cryptic error on `patch_size` mismatch.** When loading checkpoints without explicit `args.patch_size`, the first error was an opaque `RuntimeError` from `load_state_dict`. The compatibility check now infers `patch_size` from the DINOv2 projection weight shape and raises a descriptive `ValueError` explaining the mismatch. (#971)
- **Fixed `source_shape` causing `TypeError` on `sv.Detections` iteration.** `predict()` stored `source_shape` as a Python `tuple`, which `sv.Detections` cannot index per-detection. It is now an `np.ndarray` of shape `(N, 2)` with dtype `int64`, where each row is `[height, width]`. (#966)
- **Fixed spurious "class_id out of range" warning for background class.** RF-DETR uses `num_classes + 1` logits internally; class index `num_classes` is the background/no-object class and is expected. Background-class detections now map `class_name` to `"__background__"` without a warning. (#970)
---
## 🏆 Contributors
Welcome to our new contributors, and thank you to everyone who helped with this release:
- **Md Faruk Alam** (@farukalamai) ([LinkedIn](https://linkedin.com/in/farukalamai))— *cuDNN depthwise conv backward fix for T4/P100*
- **M. Fazri Nizar** (@mfazrinizar) ([LinkedIn](https://linkedin.com/in/mfazrinizar)) — *EMA segmentation metrics fix*
- **Jirka Borovec** (@Borda) ([LinkedIn](https://www.linkedin.com/in/jirka-borovec)) — *release coordination, reviews*
*Automated contributions: @copilot-swe-agent[bot], @pre-commit-ci[bot]*
---
**Full changelog**: https://github.com/roboflow/rf-detr/compare/1.6.4...1.6.5