H Company's Holo2 nails 78.5% on tough UI benchmark by looking twice
Curated by the Inblix editorial team
H Company just dropped a preview of its largest UI localization model yet, and the numbers are worth paying attention to. Holo2-235B-A22B Preview hits 78.5% on the ScreenSpot-Pro benchmark in agent mode, beating the previous state-of-the-art. On OSWorld G, it scores 79.0%. This isn’t just a marginal bump — it’s a new record on what the company calls the most challenging GUI grounding test around. The model is a research release, available now on Hugging Face, and it’s built to solve a very specific headache: pinpointing tiny UI elements on high-resolution 4K screens.
What makes this work is something H Company calls “agentic localization.” Instead of taking one guess and hoping for the best, the model iterates. It looks at the screen, makes a prediction, then refines that prediction over multiple steps. This approach unlocks a 10-20% relative improvement across all Holo2 sizes. On ScreenSpot-Pro, the single-step accuracy sits at 70.6%. Letting it run for up to three steps in agent mode pushes that number to 78.5%. That gap tells you a lot about how much low-hanging fruit remains in just giving models a second look at a problem.
Under the hood, the training story is also interesting. H Company used SkyPilot to orchestrate workloads across multiple cloud providers and Kubernetes clusters. The pitch is that SkyPilot abstracts away the infrastructure spaghetti so researchers can focus on the model, not on managing k8s manifests. For a small company pushing out a 235-billion-parameter model in preview, that kind of tooling isn’t just a convenience — it’s probably the only reason they could ship it without a dedicated platform engineering army.
Skepticism is healthy with any research preview. Benchmarks are clean; real-world UIs are messy. But 78.5% on ScreenSpot-Pro and the iterative approach suggest a genuine advance in a problem space that has quietly held back a lot of agentic products. If your AI agent can’t reliably find the “submit” button, none of the reasoning intelligence matters. Holo2 is a bet that finding the button is a problem worth solving with scale and a bit of patience.
💡 Key Takeaways
- Agentic localization — letting a model iteratively refine its prediction over multiple steps — delivers a 10-20% relative accuracy boost across all Holo2 model sizes.
- Holо2-235B-A22B Preview sets a new state-of-the-art score of 78.5% on ScreenSpot-Pro by using up to three reasoning steps instead of a single-shot guess.
- H Company used SkyPilot to abstract away multi-cloud Kubernetes complexity during training, allowing a small research team to ship a 235B-parameter model.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.