Regional forecasting

A full year, measured at 3 km.

Hyper-CONUS v2 is our trained regional model. We benchmark it against the CONUS-24 v2 baseline, operational models, persistence, and global machine-learning forecasts using one matched 2024 test set.

2024verification year
702 / 720matched initializations
+6 to +24 hfour lead times
CONUS landarea-weighted scoring

Primary result

Error by lead time.

Root mean square error summarizes field accuracy. Every curve uses the same available 00 and 12 UTC initializations and the same land-area weighting.

Figure 01 / 2024 RMSE

Click to enlarge
Lower is better. Hyper-CONUS v2 is emphasized in blue; the CONUS-24 v2 baseline is dashed. Global forecasts are bilinearly upsampled for this view and should not be read as native 3 km systems.

Four-lead mean

The selected system and the baseline.

Mean RMSE across +6, +12, +18, and +24 hours. The HRRR row provides an operational reference.

SystemTemperature
K
Dewpoint
K
East wind
m/s
North wind
m/s
Wind speed
m/s
Precipitation
mm
Hyper-CONUS v2 · selected1.2901.4651.6081.6501.6440.754
CONUS-24 v2 · baseline2.4852.6661.9152.0072.0420.790
HRRR v41.7132.5641.6541.7051.6410.991

Units follow the evaluated variable. Values are rounded to three decimals. A lower precipitation amount RMSE does not by itself establish better event detection.

Extreme temperature / Full year

Lower error in warm and cold departures.

Observed anomalies are measured against a 2021–2023 monthly and UTC-hour climatology. The comparison uses a common 0.25° grid so regional and global systems are scored on identical pixels.

Figure 02 / Anomaly thresholds

Click to enlarge
At ±5 K, Hyper-CONUS v2 leads every comparator. RMSE is 11.2% lower than HRRR for warm anomalies and 16.3% lower for cold anomalies; reductions versus the three GFS-initialized AI forecasts range from 21.8% to 37.5%. The advantage holds at every evaluated lead. At ≥9 K warm anomalies, Hyper-CONUS v2 and HRRR are statistically tied.

Spatial case / 11 May

Warm-anomaly structure at 24 hours.

The full-year score measures consistency; this case shows spatial placement. At verification time, almost 9% of CONUS land was at least 5 K warmer than its 2021–2023 monthly and UTC-hour mean.

Figure 03 / Temperature anomaly

Click to enlarge
One spatial case from the matched benchmark. Within the observed ≥5 K region, Hyper-CONUS v2 RMSE was 1.36 K, compared with 2.63 K for CONUS-24 v2 and 1.50 K for HRRR.

Interpretation

What the benchmark supports.

The result separates the selected system, the research baseline, and the limits of the precipitation score.

01

Selected system

Hyper-CONUS v2 records the lowest four-lead mean RMSE in this native-grid table for temperature, dewpoint, and hourly precipitation amount. Mean wind-speed RMSE is effectively level with HRRR.

02

Warm and cold anomalies

At the ±5 K threshold, Hyper-CONUS v2 beats HRRR and all three evaluated global AI systems. The paired advantage persists from +6 through +24 hours.

03

Precipitation caveat

Domain-average amount error rewards dry forecasts. Event thresholds and neighborhood scores are needed before making a precipitation-skill claim.

Protocol

Like-for-like where it matters.

The comparison is restricted to initialization times available from every source. Surface analyses and radar estimates provide the verification fields.

  1. Grid
    CONUS 3 km evaluation grid

    Regional systems retain their native-scale structure; global sources are explicitly identified as upsampled.

  2. Truth
    URMA and MRMS

    Temperature, dewpoint, and wind use URMA surface analysis. Hourly precipitation uses MRMS gauge-corrected estimates.

  3. Time
    00 and 12 UTC

    Lead times are +6, +12, +18, and +24 hours across the matched 2024 subset.

Technical discussion

Ask about the benchmark, model design, or evaluation data.