Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

LSTM Baseline

# Regional LSTM Baseline for Hydrological Prediction ## Overview This repository contains a reproducible regional LSTM baseline for hydrological modeling. It is designed as the benchmark counterpart to the proposed HydroMoE framework and follows a NeuralHydrology-based workflow.

👤
CreatorYuan, Wenrui
📅
Published2026-04-27
🔗
DOI10.5281/zenodo.19804505
📊
Downloads46
⚖️
Licensecc-by-4.0
File Size323.5 MB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views97
Total Downloads46

# Regional LSTM Baseline for Hydrological Prediction

## Overview

This repository contains a reproducible regional LSTM baseline for hydrological modeling. It is designed as the benchmark counterpart to the proposed HydroMoE framework and follows a NeuralHydrology-based workflow.

The baseline uses a long-table hydrometeorological dataset as input, converts it into the GenericDataset format required by NeuralHydrology, and performs runoff prediction using the following dynamic inputs:

– Precipitation
– Temperature
– Potential evapotranspiration

No static catchment attributes are used in this baseline. The standard temporal split is:

– Training: 1980-01-01 to 1999-12-31
– Validation: 2000-01-01 to 2007-12-31
– Testing: 2008-01-01 to 2014-09-30

The sequence length is fixed to 96 days.

## Purpose

This folder is intended for:

– A transparent and reproducible LSTM benchmark
– Hyperparameter search and ensemble retraining
– Comparison against HydroMoE under the same data constraints
– Archival release on Zenodo for scientific reproducibility

## Workflow

The project follows four main stages:

1. Data adaptation  
   Converts the long-table input into the NeuralHydrology GenericDataset structure.

2. Random search  
   Samples and evaluates multiple LSTM configurations to identify strong candidates.

3. Top-10 ensemble retraining  
   Retrains the best configurations with multiple random seeds and aggregates predictions.

4. Cost reporting  
   Summarizes runtime, model complexity, GPU memory usage, and other experiment costs.

## Repository Layout

### Core project files

– Main pipeline driver
– Dataset adaptation module
– Random search module
– Ensemble retraining module
– Native PyTorch LSTM training and evaluation module
– GPU cost and memory tracking utilities
– NeuralHydrology interface helpers
– Shared configuration and path management

### Output folders

– Prepared dataset outputs
– Basin split lists
– Random search configurations
– Ensemble retraining configurations
– Training run logs
– Metrics tables and prediction files
– Cost summaries and report documents

### Generated artifacts

– Validation metrics from random search
– Top-10 configuration summary
– Ensemble member execution records
– Epoch-level loss curves
– Aggregated prediction outputs
– Runtime and GPU memory summaries
– Publication-ready comparison tables and figures

## Software Requirements

The baseline was developed for a Python environment with the following key dependencies:

– NeuralHydrology
– PyTorch
– NumPy
– Pandas
– Xarray
– Scikit-learn
– DuckDB
– PyYAML
– TQDM

The project assumes a Windows-based scientific computing workfl

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
LSTM Baseline (Full Dataset)323.5 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Yuan, Wenrui (2026). LSTM Baseline. https://doi.org/10.5281/zenodo.19804505