AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How Holo4 Helps AI Agents Work Across Computer Tasks on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

H Company announced Holo4, a series of open-weight agentic models in 27B dense and 35B-A3B MoE sizes that operate software through GUIs, code, MCP and APIs with a single model, as detailed in the original analysis. The company reports the 27B version scores 61.7% on OSWorld 2.0, behind only top closed models, but all benchmark numbers are self-reported and await independent verification.

H Company has released Holo4, a new series of open-weight agentic models designed to operate software through GUIs, code, MCP and APIs using a single model. The series comes in two sizes — a 27B dense model and a 35B-A3B Mixture of Experts model — both available now on the H Models API and for download on Hugging Face in FP16, FP8 and GGUF formats. The company reports that the 27B version scores 61.7% on OSWorld 2.0, a computer-use benchmark, which it frames as trailing only the strongest closed models at a fraction of the cost.

Holo4 is positioned as a generalist computer-use agent: it clicks and types on screens, writes and runs its own code, and calls MCP or API tools, choosing whichever interface fits the task. According to H Company, most agentic models are trained for a single interface — GUI-focused models fail without a screen, while tool-calling models cannot handle applications without APIs. Holo4 runs on desktops, the web, Android, in a code sandbox and against business APIs, with the same model invoked the same way on each platform.

The models were trained through supervised and reinforcement learning on a large set of environments and tasks, including tasks generated by H Company’s Agentic Task Factory. The company says Holo4 improves substantially over its Qwen base model — specifically Qwen3.8 27B for the dense model and Qwen3.6 35B-A3B for the MoE variant, according to its benchmark notes — and shows side-by-side examples on professional software such as FreeCAD 3D modeling and Godot game design, run with the same prompt and harness.

On benchmarks, H Company reports that Holo4 27B scores 61.7% and Holo4 35B-A3B scores 30.9% on OSWorld 2.0, compared with 81.8% for Opus 5.5, the strongest closed model in the comparison. On AutomationBench for API use, Holo4 was measured in the company’s internal harness (v1.0.6) against public-set scores for other models. H Company has open-sourced every trajectory behind its public benchmark scores, viewable at trajectories.hcompany.ai and downloadable from Hugging Face.

At a glance
announcementWhen: announced with immediate availability o…
The developmentH Company released Holo4, two open-weight agentic models for cross-interface computer use, along with all benchmark trajectories for public audit.
At a glance
announcementWhen: announced via Hugging Face and company…
The developmentH Company announced the release of Holo4, a two-model series of open-weight computer-use agents, along with an updated Holotron4 Nano and open-sourced benchmark trajectories.

Open-Weight Agents Closing the Gap

Open-weight computer-use agents remain rare at this performance level. If Holo4’s reported scores hold up under independent evaluation, businesses could run capable software-automation agents at a fraction of the cost of frontier closed models, with the flexibility of self-hosting or open weights.

The cost-performance gap the company claims — 61.7% on OSWorld 2.0 from a 27B model against 81.8% from a much larger closed model — would represent meaningful progress for smaller, cheaper agents. The multi-interface design also addresses a practical limitation: real business tasks often mix screen work, code and API calls, and single-interface models break at those boundaries.

H Company’s decision to release all benchmark trajectories lets outside parties verify each step, which is more transparency than most closed-model providers offer.

From Holo1 to Holo4

Holo4 builds on H Company’s previous agentic model line and arrives alongside an updated version of Holotron 3, called Holotron4 Nano. The models are built on a Qwen base — Qwen3.8 27B for the dense model and Qwen3.6 35B-A3B for the MoE variant, according to the company’s benchmark notes.

H Company’s cost comparisons use specific assumptions: Holo4 is priced at H Models API rates for a single run, Qwen costs are calculated at Alibaba Cloud list prices (with cache hits at 20% of input price for the MoE model), and GPT and Opus effort sweeps come from OpenAI launch data. The company cautions that releases, harnesses and task subsets differ across the compared models.

“Real work is not siloed that way, and a single business task can require combining these different approaches.”

— H Company, announcement

Claims Awaiting Independent Verification

All headline benchmark numbers are self-reported by H Company and measured in the company’s own harness, which the company itself notes differs from other models’ releases, harnesses and task subsets. On AutomationBench, other models’ scores come from the public set while costs come from a leaderboard running the private set — a mismatch the company acknowledges. Holo4 has not yet been evaluated on the AutomationBench private set.

The steep score difference between the 27B dense model (61.7%) and the larger 35B-A3B MoE model (30.9%) on OSWorld 2.0 is not explained in the announcement. Real-world reliability on business workflows, beyond curated demo examples, also remains unverified by third parties.

Evaluations and Adoption Watchpoints

H Company says it will report Holo4 results on the AutomationBench private set once that evaluation is complete. Independent benchmark submissions and third-party reproductions — now possible because trajectories and weights are public — will be the next test of the company’s claims. Developers can access the models through the H Models API or download the full collection from Hugging Face.

Key Questions

What is Holo4?

Holo4 is a series of open-weight agentic models from H Company, released in a 27B dense version and a 35B-A3B Mixture of Experts version. It is designed to operate software through GUIs, code, MCP and APIs using a single model.

How does Holo4 perform on benchmarks?

According to H Company’s own measurements, Holo4 27B scores 61.7% on OSWorld 2.0 and the 35B-A3B version scores 30.9%, compared with 81.8% for Opus 5.5, the strongest closed model in the comparison. These figures are self-reported and have not yet been independently verified.

Where can developers get Holo4?

The models are available through the H Models API and for download on Hugging Face in FP16, FP8 and GGUF formats. Benchmark trajectories are viewable at trajectories.hcompany.ai and downloadable from Hugging Face.

Why does the 35B MoE model score lower than the 27B dense model?

The announcement does not explain the roughly 30-point gap between the two variants on OSWorld 2.0. It remains one of several open questions about the release.

Are Holo4’s benchmark comparisons apples-to-apples?

No. H Company itself cautions that harnesses, releases and task subsets differ across the compared models, and on AutomationBench it compares its internal-harness results against public-set scores while costs come from a private-set leaderboard. Holo4 has not yet been evaluated on the AutomationBench private set.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Is Spatial Computing and Why Are More People Talking About It?

Keen to understand why spatial computing is gaining attention and how it could reshape your daily life? Keep reading to find out.

The Inside Scoop On Elon Musk’s xAI Multi-Agent System For 2026

Details on Elon Musk’s xAI multi-agent architecture remain unconfirmed. No technical documentation or deployment info has been disclosed yet.

Should AI Be Regulated? The Debate Over Laws for Artificial Intelligence

Just how strict AI regulation should be remains uncertain, but understanding the ongoing debate is crucial for shaping its future.

Sony to issue $1bn in dollar-denominated bonds, first in 28 years

Sony plans to issue $1 billion in dollar-denominated bonds, its first such issuance in 28 years, marking a significant move in its financing strategy.