Key takeaways
- H Company announced Holo4 on September 28, 2026. The family has a dense 27B model and a 35B-total, 3B-active mixture-of-experts model designed to combine GUI actions, code execution and API or MCP tool calls in one agent loop.
- Hosted model IDs are holo4-27b and holo4-35b-a3b, both on the paid H Models API with a documented 262,144-token context and up to five JPEG, PNG or WebP images per request.
- The faster 35B-A3B API is listed at $0.30 per million input tokens, $0.03 cached input and $2 output. The dense 27B API is $0.40, $0.04 and $3 respectively.
- The weights are not interchangeable from a licensing perspective: H publishes Holo4-35B-A3B under Apache 2.0, while Holo4-27B uses CC BY-NC 4.0 and is labeled research only in the launch post.
- Treat the reported benchmark table as vendor evidence, not a universal ranking. Several comparisons use different harnesses or effort levels, OSWorld 2.0 and ALE-CLI use one Holo4 run, and 480 of 600 public AutomationBench tasks fall in a split used to collect training data.
H released two Holo4 models on September 28
H Company announced Holo4 on September 28, 2026 as a new series of agentic vision-language models. Holo4-27B is a dense 27-billion-parameter model based on the Qwen3.8 architecture. Holo4-35B-A3B uses a Qwen3.5 mixture-of-experts architecture with 35 billion total parameters and 3 billion active per token. H also released a smaller related checkpoint, Holotron4-30B-A3B, but the two Holo4 models are the principal hosted launch.
The product claim is broader than browser clicking. H says the same model can receive screenshots, write and run code, and invoke API or MCP tools, choosing among those interfaces during one workflow. The published hai-agents harness executes requested actions and returns their results to the model. That creates a useful common control loop, but it also means model behavior cannot be evaluated apart from the harness, available tools, permissions and environment.
Both models are available now through paid API IDs
The live Models API documentation lists holo4-35b-a3b and holo4-27b as paid-tier models. Both expose a 262,144-token context, accept up to five JPEG, PNG or WebP images per request and support native function calling. H documents an OpenAI-compatible endpoint, so an existing OpenAI client can use the service by changing the base URL and model ID.
H lists holo4-35b-a3b at $0.30 per million input tokens, $0.03 per million cached-input tokens and $2 per million output tokens. The dense holo4-27b costs $0.40, $0.04 and $3 respectively. Those rates describe model tokens, not completed-work economics: a long computer-use run can include many screenshots, tool results, retries and outputs. H’s benchmark table reports cost per task from its own run tokens at these rates, which should not be treated as a quote for another harness or workload.
The downloadable weights have materially different licenses
H publishes both models on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF variants. That is current downloadable availability, not a promise of weights later. The 35B-A3B model card uses Apache License 2.0, making it the clearer starting point for teams considering commercial self-hosting, subject to their own license and dependency review.
The dense 27B model card instead uses CC BY-NC 4.0, and H labels it “Research only” in the launch post. Open weights therefore does not mean unrestricted commercial use. Do not select 27B for a revenue-generating or internal commercial deployment until counsel has confirmed that the intended use fits the license or H has supplied separate terms. Also inspect the base-model, tokenizer, code and quantization licenses rather than relying on one badge.
H reports strong computer-use results with visible caveats
H reports 85.2% for Holo4-27B on OSWorld at $0.08 per task and 61.7% average partial score on OSWorld 2.0 at $1.22 per task. On AutomationBench’s 600 public tasks, it reports 45.4% at $0.05 per task. Holo4-35B-A3B is cheaper in those runs but lower scoring: 80.8% on OSWorld, 30.9% on OSWorld 2.0 and 34.5% on the public AutomationBench set. These are H-run and H-reported results, not AccessAllGPT measurements.
The footnotes matter. H says its results are means over two to four runs except for single runs on OSWorld 2.0 and ALE-CLI. Frontier references may use different harnesses, effort settings, task subsets or provider reports. Cost values mix H API prices with other vendors’ list prices and observed token counts. A precise chart does not remove those comparability limits.
AutomationBench training overlap changes the headline interpretation
H discloses that 480 of AutomationBench’s 600 public tasks fall in the split from which it collected training data. The remaining 120 are held out. On that held-out subset, H reports 49.3% for Holo4-27B and 31.7% for Holo4-35B-A3B, compared with 40.3% for Qwen3.8 27B and 13.1% for Qwen3.6 35B-A3B in H’s harness.
The disclosure is useful, but the 600-task public-set score should not be read as clean unseen-task generalization. For procurement, preserve the held-out result separately and do not blend it with the full-set headline. H also reports results on internal Agentic Task Factory test sets that it says were not used in training; because H generated, ran and scored those environments, they remain vendor-controlled evidence until independently reproduced.
Training and trajectories improve inspectability, not independence
H says supervised fine-tuning used 127 billion tokens, about three quarters of them successful agent trajectories from its Agentic Task Factory, followed by separate online reinforcement-learning experts for GUI work and terminal or tool work. The two experts were then merged into one model. H says the factory has produced about 10,000 tasks across web, MCP and desktop environments.
The company has published a trajectory viewer and downloadable dataset for its public benchmark runs. That is more inspectable than a score without traces: reviewers can examine actions, tool choices and failure paths. AccessAllGPT did not replay or audit the full collection, however. Public traces support audit of the released runs; they do not establish that the test set was uncontaminated, that graders were correct or that another harness will reproduce the scores.
Data handling is documented but not independently audited here
H’s API documentation says the service logs request time, model and token counts while prompts and responses are not stored, describing this as zero data retention by default. It also says H runs its own models and does not share inputs or outputs with third parties. Those are current vendor policy statements, not a technical audit or contractual guarantee reviewed by AccessAllGPT.
Computer-use inputs can expose entire screens, documents, credentials and tool outputs. Before sending production data, confirm the applicable contract, region, subprocessors, abuse monitoring, incident handling, deletion behavior and enterprise controls. For self-hosting, the organization inherits model-serving logs, screenshots, tool traces, secrets management and deletion obligations.
Choose 35B-A3B first, then test the complete authority path
For most evaluations, start with hosted holo4-35b-a3b: H positions it for interactive loops, it has the lower token price, and its downloadable weights have the less restrictive Apache 2.0 license. Test holo4-27b only when the harder workflow benefits from its reported performance and the API or non-commercial weight license fits. Do not assume a 262K configured context means reliable recall or stable action quality across a 262K-token agent run.
Evaluate complete tasks in your actual environment. Freeze the model ID, harness, prompt, tools and permissions; replay deterministic fixtures; record screenshots, tool calls, retries, tokens, latency, cost and final state; and score consequential side effects separately from task completion. Keep writes behind deterministic authorization and human approval. Wait when license, retention or deployment requirements are unresolved. Reject unattended authority when the agent can act beyond a reversible, least-privilege boundary.
Copy-ready Holo4 evaluation record
Complete this for one Holo4 model, access route, harness and workload before granting production authority.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Trial, deploy, constrain, wait or reject; exact user task, environment, owner, review date and expiry.
Hosted holo4-35b-a3b or holo4-27b, or exact Hugging Face repository, revision, file hashes, quantization and serving stack.
Model, base model, tokenizer, harness, code and quantization licenses; commercial status; attribution; reviewer and approval evidence.
System prompt, context policy, screenshot cadence, memory, tool schemas, retries, stop conditions, step and spend ceilings.
Readable data, credentials, GUI actions, code, APIs and MCP tools; tenant boundaries; write controls; approvals and rollback.
Frozen unseen tasks, contamination review, seeds, repetitions, graders, baseline, completion and side-effect criteria.
Success and partial scores, failure taxonomy, latency, tokens, retries, cost per attempted and accepted task, confidence intervals.
Input classes, retention contract, region, logs, subprocessors, redaction, deletion, incident process and self-hosted trace controls.
Approved scope, remaining gaps, monitoring, kill switch, fallback, correction owner and triggers for re-evaluation.
Primary sources
Browse the publication-wide evidence index →
- Holo4: powering generalist computer-use agentsH Company · Reviewed: September 28, 2026 release; model sizes; interface coverage; training recipe; benchmark table and footnotes; hosted prices; context length; weight formats; license labels; future DSpark checkpoint statement · Retrieved · Supports: H announced Holo4-27B, Holo4-35B-A3B and Holotron4 Nano on September 28, published hosted prices and weights, and reported benchmark results with important run-count, harness and training-overlap qualifications.
- Models APIH Platform Docs · Reviewed: Current model IDs; architecture; context and image limits; input, cached-input and output prices; tier access; tool use; API compatibility; deprecation metadata; open-weight formats; retention statement · Retrieved · Supports: The live documentation lists holo4-35b-a3b and holo4-27b on the paid tier, their per-token prices and 262,144-token contexts, and states that prompts and responses are not stored by default.
- Hcompany/Holo4-27B model cardH Company on Hugging Face · Reviewed: Model summary; architecture; configured context; interfaces; target environments; benchmark claims; shared trajectory dataset; harness description; quantized variants; license · Retrieved · Supports: The model card identifies a 27B dense Qwen3.8-based vision-language model, a 262,144-token configured context and downloadable BF16, FP8, NVFP4 and Q4 GGUF variants under the non-commercial CC BY-NC 4.0 license.
- Hcompany/Holo4-35B-A3B model cardH Company on Hugging Face · Reviewed: Model identity; Qwen3.5 mixture-of-experts architecture; checkpoint format; configured context; harness behavior; benchmark claims; trajectory dataset; quantized variants; license · Retrieved · Supports: The model card identifies a 35B-total, 3B-active Qwen3.5 mixture-of-experts model with a 262,144-token configured context and downloadable variants under Apache License 2.0.
Limitations
AccessAllGPT reviewed four H-controlled sources and no independent evaluation of Holo4. We did not obtain API access; download, hash or inspect model weights; test BF16, FP8, NVFP4 or GGUF builds; review the hai-agents harness code; run GUI, code, API or MCP tasks; reproduce OSWorld, OSWorld 2.0, ALE-CLI, AutomationBench, AndroidWorld or Agentic Task Factory results; validate cost-per-task calculations; audit training-data provenance or contamination; inspect all released trajectories; measure context reliability, latency, memory, throughput or hardware requirements; audit zero-retention controls; or provide a legal opinion on CC BY-NC 4.0 or Apache 2.0. H’s benchmark comparisons mix harnesses, effort settings, subsets and run counts, and the public AutomationBench result includes substantial training-collection overlap. Prices, IDs, licenses, policies and availability can change.
Disclosures
AccessAllGPT did not receive H Company access, credits, weights, hardware, a briefing, review or compensation for this article. H Company, Qwen and benchmark maintainers did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with H Company, Hugging Face, Qwen, OpenAI or organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- Gemini 3.8 Live Avatar Is Generally Available—Here Are the API Price and Limits
- GLM-5.3 Flash Open Weights: Treat “Deployment Ready” as a Testable Claim
- Design an Agent Benchmark That Predicts Production
- AI Agents or Deterministic Workflows: Make the Automation Decision
- Where Human Approval Belongs in AI Automation
- AI API Data Retention and Residency: Set the Procurement Gates
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.