v.2.0.0

dimartarmizi/map-to-posterv.2.0.0Feb 13, 2025by yadong-lu

AI Summary

This major version introduces a significantly improved dataset and performance enhancements, including 60% lower latency and high accuracy on the ScreenSpot Pro benchmark. It consolidates functionality into a single OmniTool that supports multiple large language models for controlling Windows 11 VMs.

Key Highlights

  • 60% improvement in latency compared to V1
  • 39.6 average accuracy on ScreenSpot Pro benchmark
  • Consolidated into a single OmniTool for VM control
  • Native support for OpenAI (4o/o1/o3-mini), DeepSeek (R1), Qwen (2.5VL), and Anthropic
  • Expanded icon caption and grounding dataset

New Features

  • OmniTool support for OpenAI models
  • OmniTool support for DeepSeek models
  • OmniTool support for Qwen models
  • OmniTool support for Anthropic Computer Use

Full Release Notes

# What's new in V2.0.0?
- Larger and cleaner set of icon caption + grounding dataset
- 60% improvement in latency compared to V1 [model checkpoints](https://huggingface.co/microsoft/OmniParser-v2.0)
- Strong performance: 39.6 average accuracy on [ScreenSpot Pro](https://github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding)
- Your agent only need one tool: OmniTool. Control a Windows 11 VM with OmniParser + your vision model of choice. OmniTool supports out of the box the following large language models - OpenAI (4o/o1/o3-mini), DeepSeek (R1), Qwen (2.5VL) or Anthropic Computer Use.