On September 28, 2026, H company published Holo4 on Hugging Face. The series has two sizes: a 27B dense model and a 35B-A3B mixture-of-experts model. Both are on the H Models API. The post also releases Holotron4 Nano, an update of Holotron 3.

Holo4 is described as using whichever interface fits: a GUI, code, MCP, or an API. The same model, called the same way, runs on desktops, the web, Android, a code sandbox, and business APIs. The post says most agent models are trained for one interface. A GUI model is blind without a screen. A tool-calling model gets stuck in front of an application that has no API.

[1]
Stipple: one dotted line reaches both a window and a keyboard.
One line arrives at a window and a keyboard and stays one line. That is the same model using more than one interface. An illustration, not a screenshot., AI-generated illustration, not a news photograph

Training used supervised learning and reinforcement learning. The environments and tasks include ones from the company’s Agentic Task Factory. The factory has produced about 10,000 tasks so far, across web apps, MCP servers, and desktop environments, including hybrid environments that expose the same state through a GUI and MCP.

On OSWorld 2.0, Holo4 27B scores 61.7%, Opus 5.5 scores 81.8%, and Holo4 35B-A3B scores 30.9%. The dense 27B model scores higher than the mixture-of-experts model. The post says this is done with orders of magnitude fewer parameters and at a much lower cost. The page does not state Opus 5.5’s parameter count, so that phrase is not turned into a multiple.

The post says every trajectory behind the public-benchmark scores is open, for replay at trajectories.hcompany.ai or download from Hugging Face.

[1]

The notes under the cost charts have to be read separately. OSWorld 2.0 costs are estimated from the input and output tokens of each agentic run. Holo4 is priced at H Models API rates for a single run. Other closed and open-weight points use the official OSWorld 2.0 leaderboard. The post says releases, harnesses, and task subsets differ. The line connects non-dominated score and cost pairs among the closed models. Holo4 is excluded from that line.

AutomationBench is not the same measurement. Holo4, Qwen3.8 27B, and Qwen3.6 35B-A3B use AutomationBench v1.0.6, with scores and costs from the company’s internal harness. Other models’ public-set scores come from the README, and the cost per task comes from the official leaderboard, which runs on the private set. Holo4 on that private set will be reported once it is evaluated.

The FreeCAD and Godot examples give call counts and token counts. Holo4 27B’s Eiffel Tower run used 84 calls and 1.3 million tokens. The post does not give a success rate for these examples.

Weight formats are written two ways on the same page. The collection is introduced as FP16, FP8, and GGUF. Later the post says the Hugging Face weights are BF16, FP8, NVFP4, and 4-bit GGUF. Those two lists are not collapsed into one precision. The extracted text does not say Apache, so no license is stated. Open trajectories are a separate sentence and are not treated as a weight license.

Holotron4 Nano applies the same post-training stack to Nemotron 3 Nano Omni. The post says the gains are absolute percentage points over that base. The extracted text does not list those points. Optimized DSpark drafter checkpoints are described as coming in the next few days.

[1]

要点

  • Both sizes are on the H Models API: a 27B dense model and a 35B-A3B mixture of experts.
  • On OSWorld 2.0 the 27B scores 61.7%, below Opus 5.5 at 81.8%. The 35B-A3B scores 30.9%.
  • Trajectories behind the public scores are described as open. The page does not state an Apache weight license.
  • AutomationBench mixes an internal harness with private-set costs. Holo4’s private-set result is not in yet.