Subscribe Sign in

French dev aims to solve bots' blindness so they can understand GUIs

3 min read Rewritten in plain language

Artificial intelligence

Show what we removed
  • Because sometimes escaping your sandbox requires clicking a button
  • French AI model dev H unveiled a pair of computer use models on Monday aimed at handling graphical user interfaces (GUIs).
  • Fine tuned using supervised training and reinforcement learning, Holo 4 is built atop Alibaba's Qwen 3.8 27B and Qwen 3.6 35B-A3B models.
  • Alongside Holo 4, H has also updated its Holotron model, which is based on Nvidia's Nemotron 3, with similar capabilities.
  • At AWS' Re:Invent conference last year, the company announced its own set of computer-use models.

5 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

Because sometimes escaping your sandbox needs clicking a button. French AI model dev H unveiled a pair of computer use models on Monday aimed at handling graphical user interfaces.

The report’s most important sentence, shortened and in plain words. How

Headline check

The one thing this headline claims is in the report.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. One claim in this headline could be checked, so this is a narrow pass and not a thorough one. How this is checked

Read the full reportHide the full report3 min

Because sometimes escaping your sandbox requires clicking a button

Most LLMs are great at answering prompts, but fall short when it comes to navigating around the desktop in Windows or Linux. French AI model dev H unveiled a pair of computer use models on Monday aimed at handling graphical user interfaces (GUIs).

Throughout computing history, computer use falls into three categories: command line interfaces (CLIs), application programming interfaces (APIs), and GUIs. AI agents can plug into the first two, but navigating desktop environments and applications that often prioritize form before function remains an ongoing challenge.

H's Holo 4 family of models aims to address this challenge by enabling relatively small but capable models to tackle all three computer use scenarios including pointing, clicking, scrolling, and typing their way through graphical interfaces originally meant for us meatbags.

Fine tuned using supervised training and reinforcement learning, Holo 4 is built atop Alibaba's Qwen 3.8 27B and Qwen 3.6 35B-A3B models. And by optimizing for CLIs, APIs, and GUIs, H claims that its models achieve far greater versatility than pure computer use models might otherwise. Alongside Holo 4, H has also updated its Holotron model, which is based on Nvidia's Nemotron 3, with similar capabilities.

In one example, the company showed Holo 4 27B taking advantage of FreeCAD's macro function to programmatically design a 3D model of the Eiffel Tower than manually building it using primitives like cubes. In another demo, H did the opposite using extruded shapes to recreate the company's logo, showing the model's flexibility.

As with any model dev's benchmarks, take these claims with a grain of salt, but if H is to be believed, Holo 4 outperforms larger frontier models from the likes of OpenAI, while using a fraction of the parameters.

Curiously, this doesn't mean that they're cheaper. In fact, while the company shows higher scores, in many cases the models end up costing more per task. Given what we know about Qwen 3.8 27B, this is likely due to Holo 4 using more "thinking" tokens in order to arrive at a final result relative to something like GPT 6 Luna, which doesn't perform as well in the OSWorld 2.0 benchmark, but costs less.

Having said that, the open weights models' diminutive size means that researchers, AI enthusiasts, and enterprises should be able to run them on relatively modest hardware. A 24 GB Nvidia RTX 3090 should be more than capable of running these models at 4-bit precision.

H expects users to do just that since alongside BF16, FP8, and NVFP4 weights, it's also made a Llama.cpp (and by extension LM Studio and Ollama)-friendly GGUF version of the model available for download on Hugging Face.

The model dev says that it also plans to release DSpark draft weights in order to speed up inference using a technique called speculative decoding. We've explored this performance-enhancing inference tech in the past, but in a nutshell it uses a small model to guess the outputs of a larger model. When it works, users experience a speedup in token processing and generation and, when it doesn't, it falls back to the base model ensuring no loss in output quality.

However, the models aren't worth much without a harness. H has developed several agentic harnesses including its open source HAI-Agents harness, which is available for download on its GitHub. However, in theory the models should work with third-party computer use harnesses.

H isn't the only model dev focused on computer use applications. At AWS' Re:Invent conference last year, the company announced its own set of computer-use models. Meanwhile, the big three American model labs, OpenAI, Google, and Anthropic, are also investing in this capability, perhaps because escaping their sandbox sometimes requires pushing a button. ®

You are reading our version, not theirs. This is The Register's report with its verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
The Registeras they published this story 12.3 15 31 0.1 50.9
Mundane Readneutralized from The Register 10.9 15 31 0.1 50.9

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works