v0.2.81

roboflow/inferencev0.2.81Feb 18, 2026by akaashrp

AI Summary

This release overhauls the engine and API by refactoring ChatModule into Engine/MLCEngine. It adds OpenAI API compatibility, expands prebuilt model support (Llama, Mistral, Gemma, Qwen, Phi), and improves WebGPU runtime performance.

Key Highlights

  • Engine and API overhaul (ChatModule -> Engine/MLCEngine)
  • OpenAI API compatibility
  • Expanded prebuilt model support
  • WebGPU performance improvements and reliability
  • XGrammar integration for constrained generation

Breaking Changes

  • Engine and API refactoring (ChatModule -> Engine/MLCEngine)

New Features

  • Engine/MLCEngine refactoring
  • OpenAI API compatibility (chat/completions, function calling)
  • Expanded model support (Llama 2/3/3.1/3.2, Mistral, Gemma 2, Qwen2/2.5/3, Phi)
  • WebGPU improvements and better OOM/deviceLost handling
  • XGrammar integration for JSON-schema/grammar-constrained generation
  • IndexedDB caching support
  • ServiceWorkerEngine + updated Chrome extension demos

Full Release Notes

* Engine and API overhaul: ChatModule refactored into Engine/MLCEngine, consolidated constructor/reload behavior, multi-model loading, better worker lifecycle, concurrency handling
* OpenAI API: mirror chat/completions APIs, stateful options, function calling and embeddings support
* Conversation templates: unified conversation template schema with custom templates
* Expanded prebuilt model support: added support for more models (Llama 2/3/3.1/3.2, Mistral variants, Gemma 2, Qwen2/2.5/3, Phi family, including vision)
* Runtime and caching: WebGPU performance/reliability improvements (more GPU-side kernels, better OOM/deviceLost handling), wasm/prebuilt versioning updates, support IndexedDB caching
* XGrammar integration: JSON-schema/grammar-constrained generation, XGrammar structural tag
* TVM-FFI integration: refactor for compatibility with more recent TVM commits and TVM FFI 
* Examples: ServiceWorkerEngine + updated Chrome extension demos, new RAG/doc-chat examples, tool calls via structural tag
* CI: GitHub actions for linting and pre-commit hooks 

**Full Changelog**: https://github.com/mlc-ai/web-llm/compare/v0.2.0...v0.2.81