Key takeaways
- Ai2 reported on October 9 that it replaced a priority-based scheduler with GPU-time budgets, hierarchical fair share and minimum-runtime contracts across clusters ranging from 88 to 1,024 H100, B200 and B300 GPUs.
- During a provider-reported 30-day test, teams received 98% of owed GPU hours and cluster occupancy held at 98%. Eighteen percent of delivered GPU time was unallocated, interruptible capacity used to keep hardware occupied.
- On Ai2’s largest H100 cluster, median queue wait reportedly fell from five minutes to 24 seconds and p90 wait from 2.8 hours to 1.8 hours. Debug-job p90 fell from two hours to 30 seconds, but Ai2 notes the baseline debug sample was smaller and more variable.
- The scheduler is not generally available software in the report, and Ai2 publishes no price, repository, API or deployment package. This is a production design report, not a product launch or independently reproduced benchmark.
- Teams should copy the contract, not the headline: measure owed versus delivered GPU time, occupancy, preemption waste, queue latency by job size, checkpoint recovery and interactive-session disruption before changing production policy.
Ai2 reported a production scheduling change on October 9
Ai2 published the results of a new scheduler for its AI research clusters on October 9, 2026. The institute says it manages thousands of NVIDIA H100, B200 and B300 GPUs in clusters ranging from 88 to 1,024 accelerators for roughly 150 internal researchers. Work includes LLM and vision-language-model training, robotics reinforcement learning and post-training for scientific agents.
This is an internal infrastructure report, not a software release. Ai2 does not provide a public repository, package, API, managed service, price or installation instructions. The available artifact is a detailed account of the policy, simulation and reported production outcomes. Builders can assess the design now, but cannot assume they can deploy Ai2’s implementation.
Priority labels failed under two-to-three-times excess demand
Ai2 says submitted workloads requested two to three times the GPUs available at any moment. Its former policy combined priority labels with optional preemptibility and team limits for protected jobs. Over time, every scheduled workload used HIGH priority, lower levels starved, users parked no-op jobs to preserve fast interactive access, and on-call engineers negotiated shutdowns on unhealthy hosts.
The failure was economic as much as technical: declaring high priority or retaining a non-preemptible slot imposed no explicit cost on the requesting team. Ai2’s response was to move strategic allocation out of individual job labels and into a budget process, then make the scheduler account for consumption against those budgets.
Managers allocate GPU time while the scheduler allocates GPUs
Managers assign proportional GPU-time budgets through a hierarchy of programs, projects and researchers. The scheduler tracks occupancy over a sliding lookback window—seven days by default—and places workloads from under-served allocations ahead of workloads from over-served ones. A request must be funded to receive protection from preemption.
This resembles established hierarchical fair-share systems rather than a new scheduling theorem. SchedMD’s current Slurm documentation describes Fair Tree ranking through account associations using normalized shares and usage. Ai2 says its distinctive input is a research-program tree weighted by managerial budgets, not that it invented fair share. The cited Dominant Resource Fairness paper supplies historical context, not validation of Ai2’s production implementation.
A minimum-runtime contract makes long training jobs interruptible
Each workload declares the shortest runtime needed to make useful progress and whether it can resume. The job is protected during that minimum window and charged to its allocation. Afterward, the scheduler can preempt and requeue it to rebalance shares. Ai2 capped protected minimum runtime at eight hours during the rollout. A zero minimum requests unallocated, always-preemptible GPU time that does not consume budget.
That contract matters because large training jobs can otherwise occupy GPUs for days or weeks. It also makes checkpoint quality part of scheduling policy: a resumable label is useful only if state can be saved and restored within acceptable time and without corrupting progress. Ai2 reports the policy let unhealthy hosts drain automatically and reduced repairs requiring a person by 74%; that is an Ai2 measurement, not a portable expectation.
The 30-day results preserved occupancy and improved queue times
Ai2 says teams received 98% of the GPU hours owed to them during a 30-day test, where owed time was capped hour by hour at actual demand. Thirteen of 15 team allocations received at least 95%, and the lowest received 90%. Reported cluster occupancy was 98% both before and after the change. 18% of delivered GPU time was unallocated capacity, allowing idle funded shares to be used by interruptible jobs.
On the largest H100 cluster, median queue wait fell from five minutes to 24 seconds and p90 wait fell from 2.8 hours to 1.8 hours. Debug-workload p90 fell from two hours to 30 seconds. Ai2 had simulated a six-hour-to-five-minute change using constructed scenarios, because historical data lacked enough debug-like workloads. It says the smaller baseline debug sample had higher variance, so the 30-second result should not be treated as a universal scheduler benchmark.
Interactive sessions and very large jobs remain unresolved
The policy made long-lived interactive sessions worse. Researchers had relied on volatile sessions that could remain alive for a week; under the eight-hour protection cap, preemption could force them to recreate state manually. Ai2 plans a nearby CPU-only environment and restorable sessions, but those remedies are roadmap items rather than verified parts of the reported deployment.
Ai2 is also investigating capacity fragmentation for the largest jobs. Protecting many smaller jobs for their minimum runtime can leave too few simultaneous placement opportunities for a large distributed workload. The report does not publish throughput per training run, checkpoint overhead, failed-resume rate, energy use, complete cluster topology, workload counts or statistical intervals. Occupancy measures assignment, not useful accelerator computation.
What AI infrastructure teams should do next
Trial the policy in simulation and on one non-critical partition before changing a shared training fleet. Reconstruct historical arrivals, GPU counts, durations and preemption behavior; add explicit debug, interactive, large-job and maintenance scenarios; and predeclare acceptable allocation error, queue latency, checkpoint loss, failed resumes, occupancy and useful utilization. Keep a baseline policy available for comparison.
Adopt only when funded groups receive their promised time, opportunistic work keeps spare capacity occupied, large jobs remain schedulable and the operational savings exceed checkpoint and user-state costs. Constrain minimum runtimes and interactive access separately. Wait when traces or restart evidence are incomplete. Reject automatic preemption for workloads that cannot restore safely or whose interruption cost is unknown.
Copy-ready AI cluster scheduling trial record
Complete this record before replacing priority rules on a shared accelerator cluster.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Cluster, GPU types and count, scheduler version, workload classes, participating teams, accountable infrastructure owner and trial dates.
Arrival trace, requested versus available GPUs, queue latency by size, protected-job share, occupancy, useful utilization, maintenance interruptions and current gaming behavior.
Programs, projects, allocation percentages, decision owners, review cadence, borrowing rules, idle-share policy and dispute path.
Minimum and maximum protected runtime, resumability declaration, checkpoint interval, restore target, preemption warning and zero-budget opportunistic mode.
Trace window, constructed cases, scheduler configurations, random seeds, repetitions, missing data, expected failure modes and acceptance thresholds.
Owed and delivered GPU hours, allocation error, occupancy, useful utilization, p50/p90 queue latency by job size, preemptions, wasted GPU time and failed restores.
Interactive-session loss, large-job fragmentation, on-call interventions, unhealthy-host drain time, support tickets and communication plan.
Decision, evidence date, unresolved risks, rollback trigger, policy owner, next review and links to raw traces and configuration.
Primary sources
Browse the publication-wide evidence index →
- Impactful scheduling for GPU clustersAi2 on Hugging Face · Reviewed: October 9, 2026 publication date; cluster and workload scope; priority-scheduler failure modes; GPU-time budgets; hierarchical fair share; scheduling contract; simulation method; 30-day rollout results; queue latency; occupancy; maintenance; limitations and roadmap · Retrieved · Supports: Ai2 reports replacing priority-based scheduling across research GPU clusters with budget-weighted hierarchical fair share and minimum-runtime contracts, while preserving 98% occupancy during a 30-day production comparison.
- Fair Tree Fairshare Algorithm (Slurm 26.05)SchedMD · Reviewed: Introduction; end-user overview; association-tree behavior; Level Fairshare calculation; user ranking; configuration and implementation notes for the current Slurm documentation · Retrieved · Supports: SchedMD documents hierarchical Fair Tree behavior in Slurm: sibling associations are ordered by normalized shares and usage, and that ordering determines fair-share priority through the account tree.
- Dominant Resource Fairness: Fair Allocation of Multiple Resource TypesUSENIX NSDI 2011 · Reviewed: Publication record, authorship and open-access paper links for the NSDI 2011 resource-allocation paper cited by Ai2 as background rather than as validation of the new production scheduler · Retrieved · Supports: The USENIX record establishes the cited Dominant Resource Fairness paper and its original research context. It does not independently verify Ai2’s 2026 implementation or operating results.
Limitations
AccessAllGPT reviewed public documents but did not access Ai2 infrastructure; inspect the scheduler, simulator, traces, budgets or accounting system; confirm cluster counts and topology; reproduce allocation or latency metrics; validate the 30-day comparison window; observe preemption; measure checkpoint loss, failed resumes, model throughput, useful GPU utilization, energy or cost; interview researchers; or test interactive and very-large-job regressions. Ai2’s figures are provider-reported, come from one organization and may depend on workload mix, excess demand, scheduler integration, checkpoint behavior and local governance. The report does not release deployable software or pricing, and roadmap mitigations are not available-now capabilities.
Disclosures
AccessAllGPT received no Ai2 cluster access, code, workload traces, briefing, review, payment or compensation for this article. Ai2, Hugging Face, SchedMD, USENIX and NVIDIA did not sponsor, review or endorse it. AccessAllGPT did not benchmark, score or rank the scheduler. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Ai2 or the organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- NVIDIA Publishes the Nemotron Recipe for IMO and IOI
- AWS Adds SageMaker Inference Benchmarking to Coding Agents
- Managed LLM API vs Self-Hosting
- Choose a Model Without Chasing the Leaderboard
- Design an Agent Benchmark That Predicts Production
- AccessAllGPT Research methodology
- Evidence standards
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.