v.2.0.0
microsoft/VibeVoicev.2.0.0Feb 13, 2025by yadong-lu
AI Summary
This release significantly enhances the model with a larger dataset and improves performance, including a 60% reduction in latency and strong accuracy benchmarks. It introduces OmniTool, a unified agent tool designed for seamless control of Windows 11 VMs with various large language models.
Key Highlights
- Larger and cleaner icon caption and grounding dataset
- 60% improvement in latency compared to V1
- Strong performance: 39.6 average accuracy on ScreenSpot Pro
- Unified OmniTool for Windows 11 VM control
- Native support for OpenAI, DeepSeek, Qwen, and Anthropic models
New Features
- Larger and cleaner icon caption + grounding dataset
- OmniTool agent integration
- OpenAI model support (4o, o1, o3-mini)
- DeepSeek R1 model support
- Qwen 2.5VL model support
- Anthropic Computer Use support
Full Release Notes
# What's new in V2.0.0? - Larger and cleaner set of icon caption + grounding dataset - 60% improvement in latency compared to V1 [model checkpoints](https://huggingface.co/microsoft/OmniParser-v2.0) - Strong performance: 39.6 average accuracy on [ScreenSpot Pro](https://github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding) - Your agent only need one tool: OmniTool. Control a Windows 11 VM with OmniParser + your vision model of choice. OmniTool supports out of the box the following large language models - OpenAI (4o/o1/o3-mini), DeepSeek (R1), Qwen (2.5VL) or Anthropic Computer Use.