Extending Medium-Range Global Flood Forecasts: The Google Global Flood Forecasting Model Version 2
Code and Analysis for "Extending Medium-Range Global Flood Forecasts: The Google Global Flood Forecasting Model Version 2" Overview This repository contains the comprehensive evaluation codebase and analysis workflow accompanying the manuscript, "Extending Medium-Range Globa
Code and Analysis for “Extending Medium-Range Global Flood Forecasts: The Google Global Flood Forecasting Model Version 2”
Overview
This repository contains the comprehensive evaluation codebase and analysis workflow accompanying the manuscript, “Extending Medium-Range Global Flood Forecasts: The Google Global Flood Forecasting Model Version 2”.
The paper documents the operational updates to the Google Global Flood Forecasting system, transitioning from an Encoder-Decoder LSTM (ED-LSTM) in Version 1 to a continuous Mean Embedding LSTM (ME-LSTM) architecture in Version 2. This upgrade incorporates an expanded training dataset from the Caravan community dataset and integrates GraphCast AI-based meteorological forcings. The primary finding of the paper demonstrates that the v2 system extends the reliable predictive horizon by 6 days in gauged basins and 2 days in ungauged basins relative to the v1 nowcast.
This repository provides the complete Python analysis pipeline (nextgen_river_model_analysis.ipynb) and the required datasets to evaluate these models, compute all metrics, and reproduce the figures presented in the manuscript.
Repository Contents
This repository is organized into two primary components:
1. Analysis Code (nextgen_river_model_analysis.ipynb)
The core analysis script. This single notebook encapsulates the entirety of the data analysis, statistical testing, and visualization code required to reproduce the paper’s findings.
Key Features of the Analysis Code:
Metric Computation: Includes highly optimized, native
xarraylogic to compute multi-dimensional predictive performance metrics without routing through Pandas. Computed metrics include:Nash-Sutcliffe Efficiency (NSE)
Kling-Gupta Efficiency (KGE) and its specific components: Correlation ($r$), Bias Ratio ($$), and Variability Ratio ($$).
Statistical Lead Time Analysis: Contains the logic for the one-sided Wilcoxon signed-rank tests used to measure statistically significant lead time extension between models.
Catchment Attribute Analysis: Translates raw HydroATLAS feature names into human-readable descriptions and computes Pearson/Spearman correlation coefficients to evaluate model performance across diverse hydro-climatic settings (e.g., aridity, elevation, snow cover).
Figure Generation: The script is chronologically organized to perfectly mirror the manuscript.
2. Data Archive (data.tgz)
A compressed tarball containing all the required datasets, cached metrics, and geographic metadata used by the notebook. The archive includes the following files:
model_runs.zarr: The core target Zarr store📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.