v.2.0.0

songquanpeng/one-apiv.2.0.0Feb 13, 2025by yadong-lu

AI Summary

A significant update to grounding datasets and model checkpoints, featuring 60% lower latency and improved accuracy on benchmark tests like ScreenSpot Pro.

Key Highlights

  • 60% improvement in latency compared to V1
  • Enhanced icon caption and grounding dataset
  • 39.6 average accuracy on ScreenSpot Pro
  • OmniTool supports multiple LLMs (OpenAI, DeepSeek, Qwen, Anthropic)

New Features

  • Updated model checkpoints
  • Multi-LLM support via OmniTool
  • Reduced latency

Full Release Notes

# What's new in V2.0.0?
- Larger and cleaner set of icon caption + grounding dataset
- 60% improvement in latency compared to V1 [model checkpoints](https://huggingface.co/microsoft/OmniParser-v2.0)
- Strong performance: 39.6 average accuracy on [ScreenSpot Pro](https://github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding)
- Your agent only need one tool: OmniTool. Control a Windows 11 VM with OmniParser + your vision model of choice. OmniTool supports out of the box the following large language models - OpenAI (4o/o1/o3-mini), DeepSeek (R1), Qwen (2.5VL) or Anthropic Computer Use.