Skip to content

H Company Ships Holo4, Open-Weight Agents That Work Any Software Interface

Short answer

H company released Holo4, agentic models in 27B dense and 35B-A3B mixture-of-experts sizes, plus an updated Holotron4 Nano, all available on the H Models API and as open weights. Holo4 can drive GUIs, write code, and call MCP or API tools with the same model, scoring 61.7% (27B) and 30.9% (35B-A3B) on OSWorld 2.0 versus 81.8% for Opus 5.5, at much lower cost.

What this means for operators

For a 10-200 person B2B company, the practical draw is not the benchmark score but the interface flexibility: most internal tools, legacy CRMs, and vendor portals have no usable API, forcing manual clicking today. A model that can fall back to GUI control when an API is missing, then switch to API calls where one exists, is directly applicable to automating order entry, ticket triage, or data reconciliation across tools that were never built to be automated. The open weights (BF16, FP8, NVFP4, GGUF) mean a technical team could self-host rather than pay per-call frontier pricing, though the reported OSWorld 2.0 gap versus Opus 5.5 (61.7% vs 81.8%) signals this is a cost-performance tradeoff, not a like-for-like replacement, and should be piloted on a specific workflow before being trusted with production tasks.

H company published details of Holo4, a new series of agentic models built to operate software through whichever interface is available: graphical interfaces, written code, MCP, or direct API calls. The release includes two sizes, a 27B dense model and a 35B-A3B mixture-of-experts model, both available now on the H Models API and as open weights on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF formats. The company also released Holotron4 Nano, an updated version of Holotron 3 built by applying the same training recipe to NVIDIA's Nemotron 3 Nano Omni model, as part of the NVIDIA Nemotron Coalition.

According to the post, Holo4 is trained to use one model across desktops, the web, Android, code sandboxes, and business APIs, rather than requiring a different model per platform. The company frames this against a common limitation: agentic models optimised for GUI clicking are unusable without a screen, while models built for tool-calling fail against applications that expose no API, and real business tasks often need both.

On OSWorld 2.0, a benchmark for long desktop workflows, Holo4 27B scored 61.7% against 81.8% for Opus 5.5, while Holo4 35B-A3B reached 30.9%. H company says this trails the strongest closed models but at a much lower cost per task, and it has published every benchmark trajectory publicly for replay. The models were trained using supervised and reinforcement learning on environments generated by the company's internal Agentic Task Factory, which has so far produced roughly 10,000 tasks spanning web apps, MCP servers and desktop environments, including hybrid environments exposing the same state through both a GUI and MCP.

H company also rebuilt its training harness, the loop that executes an agent's actions and manages its context over long task sequences, adding a persistent memory across hundreds of steps and direct shell access on the desktop machine. Optimized DSpark drafter checkpoints for faster inference are planned for release shortly.

Source: Hugging Face

Next step

Visibility Analyzer

This is what the Visibility Analyzer measures on a real site: which answers cite you, which pages an engine cannot retrieve, and what to fix first. Free to run.

Run a free visibility audit

Free to run. No card.

Fee
Free
Length
One run, minutes

Free tier: two analyses a day, no card required.