Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Microsoft goes quiet after church groups ask for 1% of data center costs

    How to Catch Data Drift When Every Feature Looks Normal

    Nvidia’s Answer to Rogue Agents Is an Open-Source AI Security System

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI News»Holo4: powering generalist computer-use agents
    AI News

    Holo4: powering generalist computer-use agents

    By No Comments7 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Holo4: powering generalist computer-use agents
    Share
    Facebook Twitter LinkedIn Pinterest Email


    Holo4 is our new series of agentic models. It comes in two sizes: 27B dense and 35B-A3B Mixture of Experts. Both are available on the H Models API. We are also releasing an updated version of Holotron 3: Holotron4 Nano.

    Holo4 builds on our previous model and interacts with software through any available interface: GUIs, code, MCP and APIs. It scores well on academic benchmarks, but we built it for real business workflows. It was trained through supervised and reinforcement learning on a large set of environments and tasks, including those generated by our Agentic Task Factory.

    Get started now:



    Models built for every interface

    Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools. It uses whichever fits the task. Most agentic models are trained for one interface only: GUI-focused models are blind without a screen, while models that prefer tool calling are stuck in front of an application that has no API. Real work is not siloed that way, and a single business task can require combining these different approaches.

    Holo4 runs on desktops, on the web, on Android, in a code sandbox and against business APIs. It is the same model in each case and it is called the same way. You do not need to select a different model for each platform.

    Holo4 benchmark results with cost per task, against Qwen3.8 27B and frontier models

    Holo4 models improve significantly over their Qwen base. Holo4 trails only the strongest closed models on long workflows: on OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and Holo4 35B-A3B reaches 30.9%. However, it does so with orders of magnitude fewer parameters and at a much lower cost. We open-source every trajectory behind our scores on public benchmarks: replay each step at trajectories.hcompany.ai or download them from Hugging Face.



    Competitive with the frontier, at a fraction of the cost

    On the hardest academic benchmarks for desktop control (OSWorld 2.0) and API use (AutomationBench), Holo4 competes with frontier models at a much lower cost per task.

    OSWorld 2.0: average partial score vs. cost per task

    AutomationBench: score vs. cost per task

    Notes on the cost-performance charts

    OSWorld 2.0. Costs are estimated from the input and output tokens of each agentic run. Holo4 is priced at H Models API rates (single run). Qwen3.8 27B: model card score, cost from the tokens of our run at Alibaba Cloud list prices. Qwen3.6 35B-A3B: single run in our harness, at Alibaba Cloud list prices with cache hits at 20% of the input price. OpenAI launch data supplies the GPT and Opus effort sweeps; other closed and open-weight points use the official OSWorld 2.0 leaderboard. Releases, harnesses and task subsets differ. The line connects non-dominated score and cost pairs among the closed models; Holo4 is excluded.

    AutomationBench. Holo4, Qwen3.8 27B and Qwen3.6 35B-A3B: AutomationBench v1.0.6, scores and costs measured in our internal harness. Other models: public-set scores from the AutomationBench README, cost per task from the official leaderboard, which runs on the private set. We will report Holo4 on the private set once it is evaluated.



    AI that does work

    Trained on environments and tasks from our Agentic Task Factory, Holo4 models excel on professional software. The examples below show Holo4 27B alongside Qwen3.8 27B, its base model. Same prompt and harness for both models.



    3D modeling · Eiffel tower

    Build a 3D model of the Eiffel Tower in FreeCAD, at a scale of 1 mm to 1 metre, centred on the origin and aligned to the X and Y axes. Work to this design.

    The tower is square in plan at every height, never round. Its half-width, measured from the central axis out to the corner, is 62.5 mm at ground level, 32.5 mm at height 57, 17.5 mm at height 115, and 9.35 mm at height 276. Between those heights the half-width follows a smooth curve that falls steeply near the ground and gently higher up, never a straight line.

    Four identical legs, one per quadrant, each a square column whose outer corner follows that profile. Each leg is 14 mm across at the ground and tapers to 4 mm at height 276. The legs stand apart from the ground up to the first platform, and converge as they rise. Nothing fills the space between them: the tower is open, and you can see straight through it from every side.

    Three platforms, each a solid square slab centred on the axis: 72 mm across and 4 mm thick at height 57; 40 mm across and 3 mm thick at height 115; 22 mm across and 3 mm thick at height 276.

    A mast from height 276 to 324, square, 8 mm across at its base tapering to 2 mm at the tip.

    Every part must be a closed solid with non-zero volume, and no part may fill the space between the legs.

    Holo4 27B (84 calls, 1.3M tokens)

    Qwen3.8 27B (60 calls, 1.0M tokens)



    3D modeling · H logo

    Build a 3D model in FreeCAD of the H company logo: a solid filled disc beside a blocky sans-serif capital letter H, both extruded to the same thickness, the two shapes of similar height and set apart so they do not overlap, with the centre of the disc level with the middle of the H. Colour both shapes black.

    Holo4 27B (94 calls, 1.5M tokens)

    Qwen3.8 27B (118 calls, 1.9M tokens)



    Game design · Pac-Man

    Build a Pac-Man-style game in Godot and leave it running.

    A rectangular maze of walls laid out on a grid, with pellets filling every open corridor. A player marker moves continuously along the corridors, eating each pellet it passes over and scoring a point for it. Three ghosts move through the same corridors and chase the player. If a ghost catches the player, the player loses a life and everything resets to its starting position. Score and lives are drawn on screen.

    No one is going to play this. The player drives itself with a simple heuristic: at each junction it heads toward the nearest pellet, unless a ghost is close, in which case it moves away from the ghost. The game must run unattended and indefinitely, with no keyboard input at all.

    When it works, start the game and leave it playing.

    Holo4 27B (68 calls, 2.4M tokens, 268 lines)

    Qwen3.8 27B (197 calls, 11.4M tokens, 327 lines)



    How we built Holo4



    Agentic task factory

    Our internal set of agentic pipelines builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software. So far it has produced about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments that expose the same state through a GUI and MCP.

    Agentic Task Factory: sources, build, gates, runtime and output



    Training

    Holo4 training: supervised fine-tuning on 127B tokens, two RL experts, one merged model



    Harness

    Alongside training, we rebuilt our harness, the loop that executes the model’s actions and manages its context over hundreds of steps, using feedback from agentic performance on OSWorld 2.0. Agents tagged why each task failed and engineers reviewed their fixes. The largest changes were giving the agent a reliable memory that can keep track of hundreds of steps, and a shell on the desktop machine itself.

    Automatic harness engineering: OSWorld 2.0 score of every evaluation run over time

    Opus 5 (70.2%) and GPT-5.6 Sol (66.2%) use max-effort partial rewards on the v2026.08.08 offline set from OpenAI’s launch chart, as in the cost-performance plot. Other reference scores come from model cards and the official leaderboard. Task releases, subsets and harnesses vary.



    Holotron4 Nano

    Our post-training stack is designed to adapt to new foundation models and produce agents that generalize across interfaces and environments. As a member of the NVIDIA Nemotron Coalition, we applied our latest stack to the Nemotron 3 Nano Omni model as a follow-up to Holotron 3.

    The same recipe turns Nemotron 3 Nano Omni into Holotron4 Nano, a generalist agentic model that significantly improves over the base model on GUI workflows and in environments exposing MCP, APIs or coding sandboxes.

    Holotron4 Nano vs. Nemotron 3 Nano Omni across five benchmarks

    Gains are absolute percentage-point improvements over Nemotron 3 Nano Omni.

    These gains show that our recipe transfers well and can turn a generalist model into an agentic expert. Nothing in it is size-specific.



    Run it yourself

    Both sizes are available today on the H Models API. Weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF, next to our small model, Holotron4 Nano.

    We will release optimized DSpark drafter checkpoints in the coming days to further accelerate inference.

    Agents computeruse Generalist Holo4 powering
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleOpenAI’s agents went rogue on Washington
    Next Article AI Agents Are About to Flood the Workforce. No One’s Ready for It
    • Website

    Related Posts

    AI News

    Microsoft goes quiet after church groups ask for 1% of data center costs

    Free AI Tools

    Nvidia’s Answer to Rogue Agents Is an Open-Source AI Security System

    AI News

    Solving Math’s Greatest Problems Was an Art Form. Then Came AI

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Microsoft goes quiet after church groups ask for 1% of data center costs

    0 Views

    How to Catch Data Drift When Every Feature Looks Normal

    0 Views

    Nvidia’s Answer to Rogue Agents Is an Open-Source AI Security System

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Microsoft goes quiet after church groups ask for 1% of data center costs

    0 Views

    How to Catch Data Drift When Every Feature Looks Normal

    0 Views

    Nvidia’s Answer to Rogue Agents Is an Open-Source AI Security System

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.