TLDRocket
Sign in

H Company's new Holo2 model takes the lead in UI Localization

Hugging Face

H Company just dropped Holo2-235B-A22B, a huge new model for finding buttons and icons on screen. It now beats every rival at reading messy, high-res app interfaces — a skill AI agents desperately need.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

H Company, the Paris-based startup building AI agents that actually click around software instead of just chatting about it, has released its biggest UI localization model to date. Holo2-235B-A22B Preview lands two months after the first Holo2 batch, and it's not a minor update — the model just posted 78.5% on ScreenSpot-Pro and 79.0% on OSWorld G, both new state-of-the-art marks on benchmarks specifically designed to torture models with cluttered, high-resolution interfaces.

The hard part of this job isn't understanding what a button does. It's finding the button. On a 4K screen packed with tiny icons, toolbars, and dense menus, a model can know exactly what it's looking for and still miss the pixels by a mile. H Company's fix is what they call agentic localization: instead of guessing once, Holo2 takes an initial stab, checks itself, and refines the answer over several steps. That iterative loop is worth 10 to 20% in relative accuracy gains across every size in the Holo2 lineup, which is a big jump for what sounds like a small architectural trick.

The numbers on ScreenSpot-Pro tell the story cleanly. In a single pass, Holo2-235B-A22B hits 70.6% accuracy. Let it iterate for up to three steps, and that climbs to 78.5% — enough to top the hardest GUI grounding leaderboard around. For context, this benchmark exists precisely because most vision-language models fall apart on real desktop software, where a 16-pixel icon buried in a ribbon menu is a completely different problem than spotting an object in a photo.

None of this happens without unglamorous infrastructure work, and H Company is unusually open about that part too. Training a model at this scale meant running jobs across multiple cloud providers, so the team leaned on SkyPilot to sit on top of their Kubernetes clusters and unify how training jobs get launched. That let researchers spend their time on the model itself rather than hand-writing k8s manifests and juggling separate deployment scripts for every cloud.

What's notable here is the framing: this is a research release, not a flagship product launch, and H Company is putting the weights on Hugging Face rather than locking them behind an API. For a company chasing the same computer-use agent future as OpenAI and Anthropic, that's a deliberate bet that openness — and a faster iteration cycle — beats keeping the crown jewels closed.

My take — AI-written commentary, not fact-checked reporting

I like seeing a European lab actually ship the best number on a hard benchmark instead of just talking about sovereignty and strategic autonomy in a press release. Computer-use agents live or die on exactly this unglamorous skill — precisely locating a tiny button on a messy screen — so a 235B model that iterates its way to state-of-the-art matters more than another chatbot with a slightly better vibe. Putting it on Hugging Face instead of behind a paywalled API is the right call, and frankly it's the only way a smaller player competes with the US giants on mindshare.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.