inference Releases
132 releases of roboflow/inference
- 3.6.0
MNN 3.6.0 focuses on multimodal LLM support, GPU backend optimizations, and low-bit quantization enhancements. It introduces support for new models like Gemma4 and Wan2.1, alongside significant performance improvements in decoding and inference acceleration.
Jun 16, 2026
- v1.3.1
This release enhances SAM 3 capabilities with point prompting and tracking, introduces Claude Fable 5 as a workflow model, and improves serverless route support.
Jun 12, 2026
- v1.10.01.10.0
v1.10.0 introduces extensive features including knowledge base infrastructure overhaul, memory UI improvements, multi-version flow support, and enterprise RBAC foundations. It also adds support for Astra DB, IBM DB2, and enhanced internationalization.
Jun 9, 2026
- v1.3.0
A major update that adds RF-DETR Keypoint detection, YOLO26 Semantic Segmentation support, and the `current_time` workflow block, accompanied by security hardening guides for self-hosted deployments.
Jun 5, 2026
- v1.9.61.9.6
This release focuses on stability improvements for the Agent LLM and CI/CD workflows. It fixes streaming behavior for incremental tokens, resolves docling remote component processing errors, and synchronizes nightly versions to ensure smooth extension bundle resolution.
Jun 2, 2026
- v1.2.13
A maintenance release that fixes semantic segmentation remote execution, resolves model aliases, and removes deprecated Gemini model versions.
Jun 2, 2026
- v1.2.12
This release adds an EditImageMetadata workflow block, enforces dense representation for instance segmentation masks, and adds YOLO26 semantic segmentation support (ONNX, TorchScript, TRT). It also updates the Execution Engine to v1.10.1 and fixes issues with inner workflow dynamic block setup and RLE mask emission.
May 29, 2026
- v1.2.11
This release focuses on security and stability, hardening auth middleware against CVE-2026-48710 and fixing TorchScript serialization behind a global lock. It also fixes the SAM3 visual_segment adapter and updates requirements on inference_models.
May 27, 2026
- v2.1
This release introduces automated cross-platform installation scripts and enhances the Camera Live feature. It includes a fix for SMS dumping errors on Windows and refactors related documentation.
May 25, 2026
- v1.2.10
This release adds support for Gemini 3.5 Flash and 3.1 Flash-Lite, ports SAM3 to `inference-models`, and introduces a batch processing CLI flag. It also includes optimizations for detections list rollup and fixes RLE mask parsing for remote execution.
May 22, 2026
- v1.2.9
This release introduces the OpenRouter passthrough with a unified Qwen-VL workflow block, adds the BoT-SORT tracking block, and includes Qwen3.5 4b support. It also fixes the usage collector cache miss and updates Jetson dockerfiles.
May 15, 2026
- v1.2.8
This release deprecates the Gaze detection feature (returns HTTP 410 Gone) and introduces new workflow blocks for Image Stack and OpenAI-compatible LLM endpoints. It also adds opt-in HTTPS support for the inference server and includes security hardening updates.
May 13, 2026
- v1.2.7
This release adds support for Google Gemma, Qwen 3.5/3.6, and Kimi workflow blocks via OpenRouter, along with a Per-Class Confidence Filter block. It also introduces GPT-5.5 support and fixes various bugs including confidence types and lambda builds.
May 1, 2026
- v1.2.6
This release fixes a confidence type widening issue in the legacy HTTP API that was introduced in the previous version.
Apr 27, 2026
- v1.2.5
This release fixes compilation and monkey-patching issues with OWLv2 and `torch.compile` in older inference versions.
Apr 24, 2026
- v1.2.4
This release introduces RLE mask representation for memory-efficient instance segmentation and adds support for the Gemma-4 model in the `inference-models` backend.
Apr 24, 2026
- v0.2.83
This release adds support for cross-origin storage and model artifact integrity verification. It also introduces subgroup support and refactors the LLM chat function registration to be state-aware.
Apr 24, 2026
- v1.8.5v1.8.5 - Spanish Locale & Token-Based Chunking
This release introduces Spanish (es-ES) language support with extensive translation coverage and improves embedding chunking to use tokens instead of characters for better handling of CJK and mixed-language content. It also fixes crashes in the credentials endpoint caused by encryption key mismatches.
Apr 19, 2026
- v1.2.3
Introduces nested Workflows (Inner Workflows) for reusability and a new `inference-compiler` extension for ahead-of-time model compilation to TRT engines.
Apr 17, 2026
- v1.2.2
Focuses on cost optimization and stability improvements, including WebRTC local stream processing, Jetson 6.x fixes, and various bug fixes.
Apr 10, 2026