inference Releases
132 releases of roboflow/inference
- v1.8.4v1.8.4 - Security Fix
A critical security update addressing three vulnerabilities: Remote Code Execution via Server-Side Template Injection, Arbitrary file write via path traversal, and Arbitrary file read via Local File Inclusion.
Apr 9, 2026
- v1.2.1
Adds introspection capabilities for locally cached models, Workflow Profiling documentation, and expands SAM3 capabilities.
Apr 7, 2026
- v1.8.3v1.8.3 - Security Fix
Addresses a critical security vulnerability involving SurrealQL injection in the notebooks API and source chat queries.
Apr 7, 2026
- 3.5.0
A major update focusing on expanding LLM inference capabilities across multiple backends (Vulkan, MUSA, QNN), introducing high-performance quantization (TurboQuant), and enhancing voice interaction features.
Apr 7, 2026
- v1.8.2
Introduces support for DashScope (Qwen) and MiniMax AI providers, full Bengali language localization, and fixes various persistence and credential management issues.
Apr 6, 2026
- v2.0
Refactors the architecture from a monolithic script to a modular package structure and introduces a modern Rich-based console interface with advanced ADB utilities.
Apr 5, 2026
- v1.2.0
A major update switching `inference-models` to the default inference engine and introducing new trackers blocks to Workflows.
Mar 27, 2026
- v1.1.2
Adds GLM-OCR support, creates a Semantic Segmentation Model workflow block, and adds an S3 Sink block.
Mar 20, 2026
- v1.1.1
Updates the Execution Engine to v1.8.0, adding control flow lineage and fixing a nested batch filtering issue.
Mar 13, 2026
- v0.2.82
Focuses on improving project stability and developer workflow through CI/CD enhancements and contribution guidelines.
Mar 13, 2026
- v1.1.0
Deprecates Python 3.9, introduces Qwen3.5 multimodal support, GPT-5.4 support, and selectable inference backends for batch processing.
Mar 11, 2026
- v1.8.1
This release introduces Bengali i18n support, podcast functionality, and upgrades the default Azure API version. It includes critical fixes for Tiktoken network errors in offline Docker deployments and resolves a SurrealDB hang issue.
Mar 11, 2026
- v1.0.5
Adds model cold start response headers, supports YOLOLite in `inference_models`, and exposes health endpoints without an API key.
Mar 6, 2026
- 3.4.1
MNN 3.4.1 brings full support for Qwen3.5 models and Linear Attention operators across CPU, Metal, OpenCL, and Vulkan backends. It optimizes LLM resource management by making the Executor built-in to instances and fixes several memory safety vulnerabilities in shape and execution operators.
Mar 5, 2026
- v1.0.4
Updates the SAM3 3D TDFY commit and fixes a class remapping issue for RFDetR-segmentation models.
Mar 4, 2026
- v1.0.3
This release focuses on stability and security by fixing container build issues and updating dependencies. It introduces AV codec support and refines dependency constraints for better compatibility.
Mar 3, 2026
- v1.0.2
This release adds comprehensive JetPack 7.1 support and expands SAM3-3D workflow capabilities with remote execution and custom image support.
Feb 27, 2026
- v1.0.1
A hotfix release addressing a specific issue with RF-Detr model post-processing in TensorRT (TRT). This is a bug fix release following the major v1.0.0 release.
Feb 23, 2026
- v1.0.0
Major 1.0 release introducing a new prediction engine called inference-models (v0.19.0) with faster loading, improved resource utilization, TensorRT support, and better modularity. Includes semantic segmentation and area measurement workflow blocks.
Feb 20, 2026
- v0.2.81
This release overhauls the engine and API by refactoring ChatModule into Engine/MLCEngine. It adds OpenAI API compatibility, expands prebuilt model support (Llama, Mistral, Gemma, Qwen, Phi), and improves WebGPU runtime performance.
Feb 18, 2026