Generating Public Transport Data based on Population Distributions for RDF Benchmarking

Taelman, Ruben; Colpaert, Pieter; Mannens, Erik; Verborgh, Ruben

Generating Public Transport Data based on Population Distributions for RDF Benchmarking

Ruben Taelman, Pieter Colpaert, Erik Mannens, and Ruben Verborgh

When benchmarking RDF data management systems such as public transport route planners, system evaluation needs to happen under various realistic circumstances, which requires a wide range of datasets with different properties. Real-world datasets are almost ideal, as they offer these realistic circumstances, but they are often hard to obtain and inflexible for testing. For these reasons, synthetic dataset generators are typically preferred over real-world datasets due to their intrinsic flexibility. Unfortunately, many synthetic dataset that are generated within benchmarks are insufficiently realistic, raising questions about the generalizability of benchmark results to real-world scenarios. In order to benchmark geospatial and temporal RDF data management systems such as route planners with sufficient external validity and depth, we designed PODIGG, a highly configurable generation algorithm for synthetic public transport datasets with realistic geospatial and temporal characteristics comparable to those of their real-world variants. The algorithm is inspired by real-world public transit network design and scheduling methodologies. This article discusses the design and implementation of PODIGG and validates the properties of its generated datasets. Our findings show that the generator achieves a sufficient level of realism, based on the existing coherence metric and new metrics we introduce specifically for the public transport domain. Thereby, PODIGG provides a flexible foundation for benchmarking RDF data management systems with geospatial and temporal data.

full text BibTeX other citation formats

Published in 2019 in Semantic Web Journal.

Keywords:

public transport
dataset generator
benchmarking
RDF
Linked Data

Read this article online

Read the full text online.
Request a digital copy of this article.

Cite this article in your work

Cite this article easily using its BibTeX entry:

@article{taelman_swj_2019,
  author = {Taelman, Ruben and Colpaert, Pieter and Mannens, Erik and Verborgh, Ruben},
  title = {Generating Public Transport Data based on Population Distributions for {RDF} Benchmarking},
  journal = {Semantic Web Journal},
  year = 2019,
  month = jan,
  volume = 10,
  number = 2,
  pages = {305--328},
  publisher = {IOS Press},
  doi = {10.3233/SW-180319},
  url = {http://www.semantic-web-journal.net/system/files/swj1926.pdf},
}

Alternatively, pick a reference of your choice below:

ACM: Ruben Taelman, Pieter Colpaert, Erik Mannens, and Ruben Verborgh. 2019. Generating Public Transport Data based on Population Distributions for RDF Benchmarking. Semantic Web Journal 10, 2 (January 2019), 305–328.
APA: Taelman, R., Colpaert, P., Mannens, E., & Verborgh, R. (2019). Generating Public Transport Data based on Population Distributions for RDF Benchmarking. Semantic Web Journal, 10(2), 305–328.
IEEE: R. Taelman, P. Colpaert, E. Mannens, and R. Verborgh, “Generating Public Transport Data based on Population Distributions for RDF Benchmarking,” Semantic Web Journal, vol. 10, no. 2, pp. 305–328, Jan. 2019.
LNCS: Taelman, R., Colpaert, P., Mannens, E., Verborgh, R.: Generating Public Transport Data based on Population Distributions for RDF Benchmarking. Semantic Web Journal. 10, 305–328 (2019).
MLA: Taelman, Ruben, et al. “Generating Public Transport Data Based on Population Distributions for RDF Benchmarking.” Semantic Web Journal, vol. 10, no. 2, IOS Press, Jan. 2019, pp. 305–28.