TANVOM v1: Dataset for Orienteering Map Vectorization
This dataset accompanies research presented at ICAI 2026 on topology-aware neural vectorization of printed orienteering maps. The broader research framework targets multiple ISOM symbol classes, but the present release is primarily designed to support reproducible experiments on processed raster
This dataset accompanies research presented at ICAI 2026 on topology-aware neural vectorization of printed orienteering maps. The broader research framework targets multiple ISOM symbol classes, but the present release is primarily designed to support reproducible experiments on processed raster tiles and derived supervision masks for map digitization, contour extraction, thin-line segmentation, and downstream raster-to-vector evaluation.
The dataset is derived from high-resolution raster exports of printed orienteering maps following the International Specification for Orienteering Maps (ISOM). Ground-truth supervision was generated from symbol-specific exports and converted into task-specific raster masks using deterministic preprocessing rules. In the associated research pipeline, neural models predict raster outputs, which are then converted into editable vector objects using deterministic post-processing and vectorization. The main methodological motivation is that practical map reconstruction quality cannot be assessed reliably from raster overlap alone; downstream vector structure, topology, and editability must also be considered.
Important note on data release scope:
For privacy and data-protection reasons, this Zenodo release contains only processed data products. The original full map sheets, original OCAD files, and full-sheet source exports are not redistributed here. Instead, the release provides cropped tile-level RGB inputs and the corresponding derived masks and metadata required for reproducible machine learning experiments. This means that the dataset is suitable for training, validation, benchmarking, and methodological comparison, but it is not intended to recreate or redistribute the original source maps in their full form.
Data organization
The shared dataset is organized at tile level. Full map rasters were partitioned into 512×512 pixel tiles with 128-pixel stride, and each retained base tile was expanded into four augmented variants. The train/validation split is performed at map-sheet level rather than tile level in order to reduce spatial leakage and better reflect generalization to unseen maps. In the current processed snapshot, the dataset contains 21,544 samples in total, with 18,880 training samples and 2,664 validation samples.
Tasks and labels
Depending on the included subset, the release may contain one or more of the following targets:
1. Contours:
A binary contour mask in which ISOM contour symbols 101 and 102 are merged into a single geometry target. This design favors continuity and downstream vectorizability over subtype discrimination.
2. Roads-black:
A binary union mask for black road symbols (ISOM 503–508), intended for geometry-first learning. In the associated pipeline, semantic road typing is performed later, after vectorization.
3. Rare line symbols:
Optional per-code binary masks for rare symbols such as 107 and 108, s
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.