LSTM Baseline
# Regional LSTM Baseline for Hydrological Prediction ## Overview This repository contains a reproducible regional LSTM baseline for hydrological modeling. It is designed as the benchmark counterpart to the proposed HydroMoE framework and follows a NeuralHydrology-based workflow.
# Regional LSTM Baseline for Hydrological Prediction
## Overview
This repository contains a reproducible regional LSTM baseline for hydrological modeling. It is designed as the benchmark counterpart to the proposed HydroMoE framework and follows a NeuralHydrology-based workflow.
The baseline uses a long-table hydrometeorological dataset as input, converts it into the GenericDataset format required by NeuralHydrology, and performs runoff prediction using the following dynamic inputs:
– Precipitation
– Temperature
– Potential evapotranspiration
No static catchment attributes are used in this baseline. The standard temporal split is:
– Training: 1980-01-01 to 1999-12-31
– Validation: 2000-01-01 to 2007-12-31
– Testing: 2008-01-01 to 2014-09-30
The sequence length is fixed to 96 days.
## Purpose
This folder is intended for:
– A transparent and reproducible LSTM benchmark
– Hyperparameter search and ensemble retraining
– Comparison against HydroMoE under the same data constraints
– Archival release on Zenodo for scientific reproducibility
## Workflow
The project follows four main stages:
1. Data adaptation
Converts the long-table input into the NeuralHydrology GenericDataset structure.
2. Random search
Samples and evaluates multiple LSTM configurations to identify strong candidates.
3. Top-10 ensemble retraining
Retrains the best configurations with multiple random seeds and aggregates predictions.
4. Cost reporting
Summarizes runtime, model complexity, GPU memory usage, and other experiment costs.
## Repository Layout
### Core project files
– Main pipeline driver
– Dataset adaptation module
– Random search module
– Ensemble retraining module
– Native PyTorch LSTM training and evaluation module
– GPU cost and memory tracking utilities
– NeuralHydrology interface helpers
– Shared configuration and path management
### Output folders
– Prepared dataset outputs
– Basin split lists
– Random search configurations
– Ensemble retraining configurations
– Training run logs
– Metrics tables and prediction files
– Cost summaries and report documents
### Generated artifacts
– Validation metrics from random search
– Top-10 configuration summary
– Ensemble member execution records
– Epoch-level loss curves
– Aggregated prediction outputs
– Runtime and GPU memory summaries
– Publication-ready comparison tables and figures
## Software Requirements
The baseline was developed for a Python environment with the following key dependencies:
– NeuralHydrology
– PyTorch
– NumPy
– Pandas
– Xarray
– Scikit-learn
– DuckDB
– PyYAML
– TQDM
The project assumes a Windows-based scientific computing workfl
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.