inference Releases
132 releases of roboflow/inference
- v1.5.2
v1.5.2 introduces Action Recognition as a new video task, adds support for Roboflow Cosmos 3 Edge fine-tunes via LoRA adapters, and enables RF-DETR keypoint detection on TensorRT. The release also includes performance improvements for Camera Focus v2 on GPU, a new health probe for secure gateway monitoring, and fixes for WebSocket connections and VLM parsing.
Sep 4, 2026
- v1.5.1
This release introduces an optional S3-compatible shared cache for model files to reduce cold-start latency, adds token-usage outputs for remote VLM Workflow blocks, and ships an experimental CUDA 13 server build. It also includes telemetry for the GStreamer CUDA video producer and fixes several critical issues regarding Modal boundary serialization, workflow worker context, MQTT sink lifecycle, and TRT CUDA-graph capture.
Aug 28, 2026
- v1.5.0
This release adds header-based API-key authentication, rebuilds `OFFLINE_MODE` in `inference-models` with an explicit registry, and introduces the Qwen VLM v2 workflow block. It also adds visual prompts for SAM3 Video Tracker, a String Template block, and a SequenceJoin UQL operation, alongside fixes for legacy RF-DETR preprocessing and Triton kernel errors.
Aug 21, 2026
- v1.4.1
This release fixes a critical Google Colab import crash caused by PyTorch and TorchAudio CUDA version mismatches and adds new AI model integrations like Grok and Gemini 3.7 Flash.
Aug 14, 2026
- v1.4.0
Introduces experimental tensor-native Workflow execution for GPU acceleration and hardware-accelerated video decoding, alongside JetPack 7.2 support.
Aug 12, 2026
- v1.3.10
Fixes a critical bug affecting all ONNX models on MacOS and adds a new contributor.
Aug 10, 2026
- v1.3.9
Adds new fusion and sink blocks for tamper detection and event bundling, expands VLM support with Qwen 3.8 Max, and improves RTSPS streaming security and performance.
Aug 7, 2026
- v1.3.8
Adds opt-in model pre-loading for Workflows, introduces Rich Label visualization and Label v2, upgrades Gemini to v4 with native detection coordinates, and unifies region/environment selection.
Jul 31, 2026
- v1.3.7
This release introduces air-gapped deployment support, adds the YOLO26 depth estimation model, and introduces new workflow blocks like Detections Nearest Neighbor alongside performance optimizations for memory usage and TensorRT engine builds.
Jul 27, 2026
- v1.3.6
This release previews NVIDIA Cosmos 3 Edge, accelerates RF-DETR detection with TensorRT, and adds new workflow blocks for image rotation and frame delays, along with enhanced observability for CUDA memory.
Jul 22, 2026
- v1.3.5
This release simplifies WebRTC streaming, adds a CUDA memory reclamation watchdog, supports PP-OCRv6, and fixes a critical SSRF vulnerability in URL image loading.
Jul 10, 2026
- v1.3.4
This release focuses on stability fixes, including HTTPS gateway handling for secure edge deployments, and introduces new workflow blocks for visual search classification and geotagging.
Jul 6, 2026
- v25.8.28.1-ltsRelease v25.8.28.1-lts
Jul 5, 2026
- v25.8.27.1-ltsRelease v25.8.27.1-lts
Jul 4, 2026
- v25.8.26.11-ltsRelease v25.8.26.11-lts
Jul 3, 2026
- v26.5.5.8-stableRelease v26.5.5.8-stable
Jul 1, 2026
- v25.8.25.37-ltsRelease v25.8.25.37-lts
Jun 30, 2026
- v1.3.3
This release introduces enterprise-level PLC Reader/Writer blocks and a vLLM proxy backend for serving VLMs, alongside bug fixes for SAM2 cache paths and workflow transformations.
Jun 26, 2026
- v1.10.11.10.1
v1.10.1 focuses on CI/CD stability, security hardening, and bug fixes for agents and model handling. It introduces first-class local-model support and addresses numerous security vulnerabilities related to SSRF and code injection.
Jun 23, 2026
- v1.3.2
The update adds a Switch Case flow-control block and restores SAM3 interactive fast paths, while resolving dependency issues and fixing Jetson runtime introspection.
Jun 19, 2026