Key takeaways
- Choose the access surface before the model: WeatherNext 3 operational arrays are not the same product as the documented WeatherNext 2 managed custom-inference endpoint.
- The technical change is deeper than finer pixels: WeatherNext 3 ingests low-latency satellite mosaics, predicts multiple observation targets, and initializes hourly while retaining a six-hour autoregressive outer step.
- Do not turn marginal skill into spatial realism. Google discloses mesh-shaped precipitation artifacts, station-head boundary jumps and changing member-level bias.
- Treat Google’s 2024 and six-week 2026 studies as strong release evidence, not a workload acceptance test. Brightband adds independent live evaluation but uses model-specific verification targets.
- Deploy forecast data behind freshness, completeness, calibration, official-warning and rollback gates; wait on custom WeatherNext 3 inference until the exact managed product contract names it.
The decision is consume, pilot, wait or reject
Consume WeatherNext 3 operational forecast data when a global probabilistic feed can improve a bounded downstream decision and you can validate the exact variables, issue times and ensemble statistics. Pilot it against an incumbent on archived decisions. Wait if the requirement is on-demand custom WeatherNext 3 inference: Google’s current managed specification explicitly names WeatherNext 2. Reject any design that treats experimental output as an official severe-weather warning or allows a missing cycle to fail open.
This is AccessAllGPT guidance, not Google’s. “WeatherNext is on Cloud” hides at least two contracts: reading provider-generated WeatherNext 3 arrays and invoking a managed model to generate tailored forecasts. Architecture, pricing, quotas, data lineage, reproducibility and rollback differ between them.
What changed on September 3
Google launched WeatherNext 3 on September 3, 2026 and integrated forecasts into Search, Gemini, Maps, Google Maps Platform and Cloud. The developer overview describes a 15-day global probabilistic forecast, 64 ensemble members and hourly initialization. The paper dates the system after WeatherNext 2 and WeatherNext-Cyclones and says the production model was trained through June 30, 2026.
WeatherNext 2 used a 0.25° grid and six-hour cadence. WeatherNext 3 adds hourly 0.1° single-level output, 0.05° station-head queries in the evaluation, direct satellite input, precipitation heads and tabular cyclone outputs. The chronology matters: this is a new operational model family member, not merely a new visualization over the old feed.
The original finding is an access-contract split
Google’s overview says operational WeatherNext 3 data is available through Cloud Storage in Zarr, Earth Engine and BigQuery. On the same site, the model-specification page states: “The model that is accessible on Gemini Enterprise Agent Platform is WeatherNext 2.” Its table describes 0.25° resolution and six-hour initialization—the older contract.
That does not make WeatherNext 3 unavailable. It makes availability surface-specific. A team that needs provider-issued forecast arrays can proceed to a data-feed pilot. A team that needs custom initial conditions, ensemble sizing or on-demand generation must verify which checkpoint the managed endpoint actually serves and must not infer WeatherNext 3 from the umbrella product name.
How the model combines analysis and observations
WeatherNext 3 receives two analysis frames six hours apart and recent geostationary satellite mosaics. At 14 UTC in the paper’s example, the latest mosaic is 13 UTC while the latest analysis-derived input is valid at 08 UTC. The satellite stream is therefore roughly five hours fresher and lets the system issue a new forecast every hour.
The model is an encode–process–decode Functional Generative Network. Separate encoders map 0.1° and 0.25° modalities to a shared icosahedral processor mesh. WeatherNext 3 raises latent width from 768 to 1,024 and transformer depth from 24 to 32 layers, then generates 64 ensemble members with noise and epistemic dropout.
Hourly does not mean every internal field is hourly
The outer autoregressive step remains six hours. Single-level variables and selected pressure-level fields are predicted at hourly offsets inside that window, while other atmospheric variables retain six-hour steps. The continuous station decoder can be queried in space and time, but its outputs are target-only rather than fed back autoregressively.
This distinction predicts one documented failure mode: station-head jumps at the boundary between six-hour windows. Procurement language should name each required variable’s cadence and head, not advertise the whole product as uniformly hourly.
The evaluation is broad, but it is not one clean benchmark
The authors use full-year 2024 evaluation with a model trained through 2023 for most results, then a July 1–August 11, 2026 quasi-real-time comparison with the production checkpoint. Baselines change by task: WeatherNext 2, ECMWF ENS and AIFS ENS v2 appear in different slices. Ground truth includes HRES analysis, held-out stations, IMERG, MRMS, rain gauges and IBTrACS.
That breadth is useful, but it prevents one score from summarizing the release. IMERG is a target during training and therefore is not independent for the IMERG head; station coverage is geographically uneven; lower-resolution baselines are interpolated; and the six-week 2026 sample is explicitly too small for over-interpreting minor differences.
The strongest claimed gains belong to bounded slices
The paper reports roughly 5% average improvement over WeatherNext 2 on upper-level medium-range variables, interpreted as about six hours of lead-time gain at matched skill. For short-lead station temperature, it reports up to 30% lower CRPS than WeatherNext 2 and 40% than ENS. Its PARDIG precipitation head reports reductions up to 60% against IMERG, 30% against MRMS and 10% against rain gauges at early leads.
These are author-reported, metric- and dataset-specific results. The precipitation evaluation caps accumulation thresholds at 4 mm per six hours and leaves extreme precipitation to future work. “Up to” values cannot be used as an expected production uplift without preserving variable, horizon, region, metric and verification target.
Fresh satellite input buys hours, not certainty
In a latency-adjusted comparison against an otherwise identical six-hour initialization schedule, the paper reports a two-to-three-hour skill lead for early precipitation forecasts. That matches the average freshness advantage of hourly cycles and is operationally plausible for developing storms.
The paper assumes seven hours from nominal analysis initialization to operational availability, including input and dissemination. A buyer should measure actual publication lag, late or revised cycles and end-to-end decision latency. A more recent forecast is not automatically better calibrated for the event, geography or threshold that drives the application.
Marginal skill can hide implausible joint samples
The model trains on marginal continuous ranked probability score. Google explicitly says this can let the model “cheat” on covariance structure. Individual precipitation and station-head members show hexagonal patterns inherited from the mesh; PARDIG artifacts remain in the ensemble mean and are less visible in the median.
The station head also develops global warm-or-cold member bias that changes at each six-hour noise draw. Quantiles are largely cleaner, but a path-dependent simulator, grid optimizer or catastrophe workflow may consume individual trajectories rather than only medians. Those applications need spatial and temporal coherence tests, not only pointwise CRPS.
Cyclone improvement comes with under-spread
The homogeneous 2024 cyclone study reports small, consistent track and intensity improvements and larger extent gains at one-to-three-day leads. It also finds WeatherNext 3 generally more under-spread than WeatherNext 2 for intensity and extent, possibly because the larger model overfits cyclone variables.
Under-spread is a decision risk: an ensemble can improve mean error while assigning too little probability to bad outcomes. Evaluate rank histograms, coverage and threshold reliability for the exact basin and action, and keep official meteorological guidance authoritative.
Independent evaluation corroborates operation, with a caveat
Brightband’s Operational WeatherBench independently lists WeatherNext 3 in a live WeatherBench-X pipeline and exposes deterministic and probabilistic comparisons. At retrieval, the dashboard showed an August 26, 2026 12Z latest ranked cycle and regional links led by WeatherNext 3, while WeatherNext 2 held the displayed three-cycle active streak for the selected temperature metric.
Brightband’s methodology says WeatherNext 3 has no analysis of its own, so it is verified against ECMWF IFS analysis, while other models may use their own lead-zero analysis. This is transparent and useful, but the target is not identical across every row. Live rank is corroborating operational evidence, not universal proof.
The paper is candid about data and geography
Training combines ERA5 from 1959, HRES-era data, an 11-channel satellite mosaic, two precipitation products, station networks and cyclone records. The paper reports major station gaps and initial biases in places including the Andes, Himalayas and high-latitude oceans. It adds 2,000 pseudo-stations per hour filled from analysis data to reduce those biases.
That correction can improve global behavior while partially reintroducing the analysis dependency the station head was meant to escape. Validate local elevation, coast, mountain and sparse-observation slices independently. Global resolution is not evidence of local calibration.
Open source does not mean this production checkpoint is open
Google’s overview points to open-source research models, pretrained weights and notebooks. The WeatherNext 3 launch and paper do not provide a downloadable production-checkpoint hash, training code bundle or complete serving artifact for the operational system reviewed here.
Do not write “WeatherNext 3 is open source” into architecture or procurement records without binding the exact repository, revision, weights, license and parity evidence. A related research implementation and a provider-operated production forecast are different supply-chain objects.
A forecast feed needs an operational contract
Record issue time, nominal initialization, availability time, revision, variable, vertical level, grid, unit, accumulation window, ensemble member semantics and missing-value policy. Reject stale data before it enters downstream optimization. Store the provider cycle identifier with every decision so the result can be reconstructed.
Check the feed at the expected cadence and alert on lateness, schema drift, partial members and impossible values. Never silently substitute yesterday’s cycle or a deterministic median where a probabilistic threshold was expected. Make the fallback explicit: incumbent forecast, no-action state or human meteorological review.
Run a shadow evaluation around decisions
Backtest both forecasts on the historical decisions the system actually makes: renewable dispatch, routing, inventory, staffing, insurance triage or flood-model input. Freeze decision timestamps so no model receives observations unavailable at the time. Score calibration, proper scoring rules and economic loss by region, season, horizon and threshold.
Then shadow live cycles for enough time to include ordinary misses, delayed data and rare conditions. Do not optimize a threshold on the same window used to approve it. Pre-register the incumbent, acceptance margin, uncertainty method, severe-error ceiling and rollback trigger.
Test artifacts against the downstream consumer
Feed individual ensemble members—not only summary statistics—through the downstream pipeline. Probe whether mesh hexagons create false localized peaks, whether six-hour jumps trigger control actions, and whether globally biased members distort portfolio risk. Compare mean, median and exceedance products without selecting the best after seeing outcomes.
If a use case depends on coherent trajectories, require domain constraints and independent observation checks. If it only needs marginal exceedance probabilities, verify reliability at those thresholds and still monitor distribution shift. The paper’s artifact disclosure is a test plan, not a reason to dismiss the model wholesale.
Keep official warnings outside the model authority boundary
The paper labels WeatherNext an automated experimental AI system, says forecasts are provided as-is, and says they are not official severe-weather warnings. Any life-safety interface must preserve official local and national alerts as the authoritative channel.
An AI forecast may support internal analysis without becoming the sole trigger for evacuation, shutdown or public warning. Document who can override it, which source wins on conflict and how users see provenance and uncertainty.
Set rollback and re-evaluation conditions now
Rollback or fail closed when issue latency breaches the freshness budget, required ensemble members are absent, schema or units change, local calibration falls below the incumbent, rare-event misses exceed the severe-error ceiling, downstream artifact tests fail, or the access surface changes model identity.
Re-evaluate after a provider checkpoint, upstream ECMWF cycle, precipitation-head, schema or dissemination change. The paper notes that the production model includes a new ECMWF cycle; upstream distribution changes are part of the effective model version even when the WeatherNext name stays fixed.
Approve one surface and one decision
Consume WeatherNext 3 data where a named decision benefits under frozen, time-correct evaluation. Pilot before replacement. Wait for explicit WeatherNext 3 managed-inference documentation when custom generation is the requirement. Reject life-safety authority and unmonitored stale-data fallback.
The durable finding is not simply that WeatherNext 3 is more accurate. It is that a major AI model can be operational while its build surface, managed-inference surface and evidence surface remain distinct. Deployment starts by naming which one you are actually buying.
Copy-ready WeatherNext deployment record
Complete one record per decision and access surface. A forecast-data feed and a managed custom-inference deployment require separate records.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Business action, geography, users, horizon, thresholds, life-safety relevance, owner and approval expiry.
Operational Zarr, BigQuery or Earth Engine data versus managed inference; exact WeatherNext version; project and region.
Nominal initialization, publication time, freshness SLO, revision policy, member completeness, late-cycle and missing-cycle behavior.
Head, variable, units, grid, level, cadence, accumulation window, interpolation and metadata version.
Frozen historical decisions, time-correct inputs, incumbent, regions, seasons, horizons, proper scores, economic loss and uncertainty.
Threshold reliability, coverage, under-spread tests, extreme-event sample, severe-error ceiling and official-warning comparison.
Individual-member spatial coherence, mesh patterns, six-hour jumps, global member bias and downstream sensitivity.
Provider cycle ID, upstream analysis cycle, checkpoint/training cutoff, schema hash, archive and reproducibility record.
Official alert precedence, human review, incumbent/no-action fallback, stale-data rejection and user-visible provenance.
Pricing, quota, retention, availability, support, deprecation, redistribution terms, monitoring and incident owner.
Freshness, completeness, skill, calibration, artifact, schema, checkpoint and upstream-cycle triggers; next review date.
Primary sources
Browse the publication-wide evidence index →
- WeatherNext 3: Increasing resolution and performance of global weather models with raw observationsGoogle DeepMind and Google Research · Reviewed: Abstract; model architecture; inputs and targets; evaluation methodology; results; joint forecast artifacts; discussion; disclaimer; appendices · Retrieved · Supports: The 43-page paper describes a 15-day, 64-member FGN ensemble with hourly initialization, mixed 0.1° and 0.25° grids, a 0.05° station head, satellite inputs, evaluation windows and baselines. It also discloses precipitation caps, limited real-time sampling, mesh artifacts, six-hour boundary jumps, cyclone under-spread and experimental-use warnings.
- Introducing WeatherNext 3, our most advanced and accurate global weather AI modelGoogle · Reviewed: September 3 launch date; architecture diagram; resolution and cadence; satellite ingestion; product integrations; model access; claimed evaluations · Retrieved · Supports: Google announced WeatherNext 3 on September 3, 2026, describing hourly refreshes, up-to-5-kilometer output, direct geostationary-satellite input and integrations across Search, Gemini, Maps, Google Maps Platform and Cloud. These capability and superiority statements are vendor claims unless separately corroborated.
- WeatherNext developer overviewGoogle for Developers · Reviewed: WeatherNext 3 operational summary; forecast-data access; custom inference; open-source positioning; Weather Lab; product navigation · Retrieved · Supports: The overview says WeatherNext 3 provides 15-day probabilistic forecasts initialized hourly across 64 members and exposes operational forecasts through Cloud Storage Zarr, Earth Engine and BigQuery. It separately advertises custom inference and open-source research models without saying the WeatherNext 3 production checkpoint or training code is downloadable.
- WeatherNext benefits and limitationsGoogle for Developers · Reviewed: Accuracy; ensembles; speed; resolution; precipitation; downstream use; reanalysis bias; smoothing; artifacts; temporal discontinuities; per-member bias · Retrieved · Supports: Google documents WeatherNext 3 resolution and output benefits but also says individual samples can show mesh artifacts, experimental precipitation-head hexagons, six-hour station-head jumps and changing per-sample global bias. It points custom ensemble sizing to WeatherNext 2 managed inference rather than WeatherNext 3.
- Model specifications and data schemaGoogle for Developers · Reviewed: Managed model identity; architecture; spatial and temporal resolution; initialization; horizon; training data; output ensemble; schema; update date · Retrieved · Supports: The managed-inference specification explicitly says the model accessible on Gemini Enterprise Agent Platform is WeatherNext 2. It documents a 0.25° model initialized every six hours, despite the adjacent site now presenting WeatherNext 3 as the flagship operational data product.
- Operational WeatherBench methodologyBrightband · Reviewed: Models; verification targets; deterministic and probabilistic metrics; aggregation; lead times; WeatherNext 3 notes; licensing and references · Retrieved · Supports: Brightband independently includes WeatherNext 3 in a live leaderboard computed with WeatherBench-X. Its methodology says WeatherNext 3 is initialized from ECMWF IFS analysis and therefore uses IFS analysis as the substituted verification target; this is operational corroboration with an important target-comparability caveat.
- Operational WeatherBench live leaderboardBrightband · Reviewed: Latest ranked cycle; active streak; ranking selector; regional leaders; scorecard, skill and time-series navigation; live-status labeling · Retrieved · Supports: At retrieval, Brightband displayed a latest ranked cycle of August 26, 2026 12Z, identified WeatherNext 2 Mean as the three-cycle active streak, and listed WeatherNext 3 as regional leader in several links. The live page corroborates evaluation availability but is mutable and does not establish universal superiority.
Limitations
AccessAllGPT did not have Google Cloud credentials, query Cloud Storage, BigQuery or Earth Engine, inspect account-level availability or pricing, run managed inference, download forecast arrays, reproduce WeatherBench-X scores, evaluate forecast skill, test calibration, inspect individual ensemble members, or validate a downstream weather decision. The paper extraction audited 43 public PDF pages and the checked-in script audited public web contracts; neither is a forecast experiment. Most quantitative results are reported by Google authors. Brightband is independent operational evidence, but its live data is mutable, the retrieved latest cycle lagged the September 3 launch date, and its methodology uses model-specific verification targets. The paper’s principal 2024 model differs from its production checkpoint; its 2026 comparison covers six weeks; precipitation thresholds exclude extremes above 4 mm per six hours in the stated analysis; station and rain-gauge coverage is uneven; and Google discloses spatial artifacts, temporal jumps and cyclone under-spread. Availability, schemas, terms and model identity can change. This is not meteorological, emergency, legal, procurement or safety advice.
Disclosures
AccessAllGPT did not receive, use, benchmark or receive advance access to WeatherNext 3, WeatherNext 2 managed inference, Google Cloud forecast data or Brightband’s underlying database. Google, Google DeepMind, Google Research and Brightband did not sponsor, review or provide private data for this work and receive no endorsement. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Google, Google DeepMind, OpenAI, Brightband or the other organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- Gemini 3.8 Flash: Same Rate, 40% Higher Cost in One Agent Suite
- GLM-5.3-Flash: Cheap, Open and Multimodal—Decide From the Artifact
- Choose a Model Without Chasing the Leaderboard
- Design an Agent Benchmark That Predicts Production
- From Paper to Production Without Losing the Plot
- AI API Data Retention and Residency: Set the Procurement Gates
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.