Hier finden Sie wissenschaftliche Publikationen aus den Fraunhofer-Instituten.

Constance: An intelligent data lake system

: Hai, Rihan; Geisler, Sandra; Quix, Christoph


Association for Computing Machinery -ACM-; Association for Computing Machinery -ACM-, Special Interest Group on Management of Data -SIGMOD-:
International Conference on Management of Data, SIGMOD 2016. Proceedings : San Francisco, California, USA, June 26 - July 01, 2016
New York: ACM, 2016
ISBN: 978-1-4503-3531-7
International Conference on Management of Data <2016, San Francisco/Calif.>
Conference Paper
Fraunhofer FIT ()

As the challenge of our time, Big Data still has many research hassles, especially the variety of data. The high diversity of data sources often results in information silos, a collection of non-integrated data management systems with heterogeneous schemas, query languages, and APIs. Data Lake systems have been proposed as a solution to this problem, by providing a schema-less repository for raw data with a common access interface. However, just dumping all data into a data lake without any metadata management, would only lead to a 'data swamp'. To avoid this, we propose Constance1, a Data Lake system with sophisticated metadata management over raw data extracted from heterogeneous data sources. Constance discovers, extracts, and summarizes the structural metadata from the data sources, and annotates data and metadata with semantic information to avoid ambiguities.