AI agents

Holo4: a generalist agent for desktop, web, and mobile workflows

H Company released Holo4, a series of agentic models that interact with software via GUIs, code, and APIs. The 27B model scores 85.2% on OSWorld at $0.08 per task.

H Company launched Holo4 on September 28, 2026, introducing a new series of agentic models designed to operate across desktop, web, and mobile environments. The release includes two primary model sizes, a 27B dense variant and a 35B-A3B Mixture of Experts variant, both available via the H Models API.

What happened

The launch marks a shift toward generalist computer-use agents that are not limited to a single interface. While many existing agentic models are trained exclusively for graphical user interfaces or strictly for tool calling via APIs, Holo4 is built to switch between these modes as needed. It can click and type on a screen, write and execute code, and call Model Context Protocol (MCP) or API tools within the same workflow. This flexibility allows it to handle business tasks that require combining different interaction methods, such as navigating a legacy desktop application while simultaneously querying a modern web API.

Alongside the main Holo4 models, H Company released Holotron4 Nano, an updated version of its post-training stack applied to the Nemotron 3 Nano Omni model. This smaller model demonstrates that the training recipe used for Holo4 can transfer effectively to other base models, significantly improving their performance on GUI workflows and coding sandboxes. The company also open-sourced every trajectory behind its benchmark scores, allowing developers to replay each step of the agent’s decision-making process.

How it works

Holo4 was trained using a combination of supervised fine-tuning and reinforcement learning on a vast dataset generated by H Company’s internal Agentic Task Factory. This factory automatically builds interactive environments and verifiable tasks from documentation, screenshots, and real software. It has produced approximately 10,000 tasks covering web apps, MCP servers, and desktop environments. The training process ensures that tasks are only kept if they pass strict verification gates, including failing on untouched seeds and passing on golden states.

The training pipeline involves three key stages. First, the model undergoes supervised fine-tuning on 127 billion tokens, with about 75% consisting of successful agentic trajectories from the task factory. Second, two specialized LoRA experts are trained using asynchronous online reinforcement learning on long-horizon tasks: one expert focuses on desktop and web interactions, while the other handles terminal, MCP, and API usage. Finally, these experts are merged back into the fine-tuned model with equal weight, creating a single model that combines general skills with specialized capabilities.

Key details

  • Holo4 27B achieves an 85.2% score on OSWorld at a cost of $0.08 per task.
  • The 35B-A3B Mixture of Experts model scores 80.8% on OSWorld at $0.05 per task.
  • On OSWorld 2.0, which tests long computer workflows, Holo4 27B scores 61.7% with a 41.5% success rate.
  • The model supports a context window of 256K input tokens and 1M output tokens.
  • Weights are available on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF formats.
  • Holotron4 Nano improves Nemotron 3 Nano Omni’s OSWorld CUA GUI score from 21.0% to 76.3%.

Why it matters

For engineers building AI-driven automation, the ability to interact with any interface is critical. Real-world business workflows are rarely siloed into just APIs or just GUIs. A task might require extracting data from a PDF viewer, entering it into a web form, and then triggering a backend process via an API. Holo4’s generalist approach means developers do not need to maintain separate models or complex routing logic for different platforms. This simplifies the architecture of agentic systems and reduces the friction of integrating AI into legacy systems that lack robust APIs.

Cost efficiency is another significant factor. Holo4 competes with frontier closed-source models on many benchmarks but at a fraction of the cost. For example, while Opus 5.5 scores higher on OSWorld 2.0, it costs $8.48 per task compared to Holo4 27B’s $1.22. This cost difference makes it feasible to run more extensive evaluations and deploy agents in production environments where margin constraints are tight. The open-weight nature of the models also allows for local deployment and customization, addressing data privacy concerns that often hinder enterprise adoption of cloud-based AI services.

What you can do

  • Access the Holo4 27B and 35B-A3B models via the H Models API to test their performance on your specific workflows.
  • Download the model weights from Hugging Face in your preferred precision format for local inference or fine-tuning.
  • Review the open-sourced trajectories on trajectories.hcompany.ai to understand how the model approaches complex multi-step tasks.
  • Experiment with Holotron4 Nano if you are working with smaller infrastructure constraints or want to apply the agentic training recipe to other base models.
  • Use the provided benchmarks as a reference point for evaluating your own agentic systems, particularly for long-horizon tasks on OSWorld 2.0.
  • Integrate MCP tools into your environment to leverage Holo4’s native support for standardized tool calling alongside GUI interactions.

Tools from the Bytechap store

Keep reading

All stories