Holistic primary key and foreign key detection

Lan Jiang,Felix Naumann

doi:10.1007/s10844-019-00562-z

Abstract

Primary keys (PKs) and foreign keys (FKs) are important elements of relational schemata in various applications, such as query optimization and data integration. However, in many cases, these constraints are unknown or not documented. Detecting them manually is time-consuming and even infeasible in large-scale datasets. We study the problem of discovering primary keys and foreign keys automatically and propose an algorithm to detect both, namely Holistic Primary Key and Foreign Key Detection (HoPF). PKs and FKs are subsets of the sets of unique column combinations (UCCs) and inclusion dependencies (INDs), respectively, for which efficient discovery algorithms are known. Using score functions, our approach is able to effectively extract the true PKs and FKs from the vast sets of valid UCCs and INDs. Several pruning rules are employed to speed up the procedure. We evaluate precision and recall on three benchmarks and two real-world datasets. The results show that our method is able to retrieve on average 88% of all primary keys, and 91% of all foreign keys. We compare the performance of HoPF with two baseline approaches that both assume the existence of primary keys.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Holistic primary key and foreign key detection

Abstract

Talk to us

Similar Papers

More From: Journal of Intelligent Information Systems

Lead the way for us

Journal: Journal of Intelligent Information Systems	Publication Date: Jun 10, 2019
Citations: 15

Similar Papers

Dependency Discovery
Ziawasch Abedjan ... Felix Naumann
-
Ziawasch Abedjan, et. al.Ziawasch Abedjan ... Felix Naumann
01 Jan 2019
01 Jan 2019

Discovery of Unique Column Combinations with Hadoop
Shupeng Han ... Haiwei Zhang
-
Shupeng Han, et. al.Shupeng Han ... Haiwei Zhang
01 Jan 2014
01 Jan 2014

Data profiling with metanome
Thorsten Papenbrock ... Jakob Zwiener
Proceedings of the VLDB Endowment | VOL. 8
Thorsten Papenbrock, et. al.Thorsten Papenbrock ... Jakob Zwiener
01 Aug 2015
Proceedings of the VLDB Endowment | VOL. 8

Discovering Foreign Keys on Web Tables with the Crowd
Xiaoyu Wu ... Ning Wang
Computing and Informatics | VOL. 38
Xiaoyu Wu, et. al.Xiaoyu Wu ... Ning Wang
01 Jan 2019
Computing and Informatics | VOL. 38

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Holistic primary key and foreign key detection

Abstract

Talk to us

Similar Papers

More From: Journal of Intelligent Information Systems