Floating Car Data Collection for Processing and Benchmarking
The dataset is outcome of a paper "Floating Car Data Map-matching Utilizing the Dijkstra Algorithm" accepted for 3rd International Conference on Data Management, Analytics & Innovation held in Kuala Lumpur, Malaysia in 2019. The floating car data (FCD representing movement o
The dataset is outcome of a paper "Floating Car Data Map-matching Utilizing the Dijkstra Algorithm" accepted for 3rd International Conference on Data Management, Analytics & Innovation held in Kuala Lumpur, Malaysia in 2019.
The floating car data (FCD representing movement of cars with their position in time) is produced by the traffic simulator software (further referred to as Simulator) published in [1] and can be used as an input for data processing and benchmarking. The dataset contains FCD of various quality levels based on the routing graph of the Czech Republic derived from Open Street Map openstreetmap.org.
Should the dataset be exploited in scientific or other way, any acknowledgement or references to our paper [1] and dataset are welcomed and highly appreciated.
Archive contents
The archive contains following folders.
city_oneway and city_roadtrip – FCD from the city of Brno, Czech Republic where FCD is based on Origin-Destination in case of oneway and Origin-Destination-Origin in case of a road trip
intercity_oneway and intercity_roadtrip – FCD from cities of Brno, Ostrava, Olomouc and Zlin, all Czech Republic where FCD is based on Origin-Destination in case of oneway and Origin-Destination-Origin in case of a road trip
Content explanation
All four of mentioned folders contain raw FCD as they come from our Simulator, post-processed FCD enriching Simulator FCD, and obfuscated raw FCD (of both low and high obfuscation level). In the both obfuscated data sets, each measured point was moved in a random direction a number of meters given by drawing a number from a Gaussian distribution. We utilized two Gaussian distributions, one for the roads outside the city (N(0,10) for the lower and N(0,20) for the higher obfuscation level) and one for the roads inside the city (N(0,15) and N(0,30) respectively). Then some predefined number of randomly chosen points were removed (3% in our case). This approach should roughly represent real conditions encountered by FCD data as described by El Abbous and Samanta [2].
In case of post-processed road trip data, there is one extra dataset with "cache" suffix representing the very same dataset limited to a 5-minute session memoization. This folder also contains a picture of processed FCD represented on a map.
Data format
Standard UTF-8 encoded CSV files, separated by a semicolon with the following columns:
RAW
Header
session_id;timestamp;lat;lon;speed;bearing;segment_id
Data
session_id: (Type: unsigned INT) – session (car) identifier
timestamp: (Type: datetime) – timestamp in UTC
lat: (Type: unsigned long) – latitude as used in Google maps
lon: (Type: un
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.