Distributed non-negative RESCAL with automatic model selection for exascale data

Manish Bhattarai,Namita Kharat,Ismael Boureima,Erik Skau,Benjamin Nebgen,Hristo Djidjev,Sanjay Rajopadhye,James P Smith,Boian Alexandrov

doi:10.1016/j.jpdc.2023.04.010

Abstract

With the boom in the development of computer hardware and software, social media, IoT platforms, and communications, there has been exponential growth in the volume of data produced worldwide. Among these data, relational datasets are growing in popularity as they provide unique insights regarding the evolution of communities and their interactions. Relational datasets are naturally non-negative, sparse, and extra-large. Relational data usually contain triples (subject, relation, object) and are represented as graphs/multigraphs, called knowledge graphs, which need to be embedded into a low-dimensional dense vector space. Among various embedding models, RESCAL allows the learning of relational data to extract the posterior distributions over the latent variables and to make predictions of missing relations. However, RESCAL is computationally demanding and requires a fast and distributed implementation to analyze extra-large real-world datasets. Here we introduce a distributed non-negative RESCAL algorithm for heterogeneous CPU/GPU architectures with automatic selection of the number of latent communities (model selection), called pyDRESCALk. We demonstrate the correctness of pyDRESCALk with real-world and large synthetic tensors and the efficacy showing near-linear scaling that concurs with the theoretical complexities. Finally, pyDRESCALk determines the number of latent communities in an 11-terabyte dense and 9-exabyte sparse synthetic tensor.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Journal of Parallel and Distributed Computing	Publication Date: May 3, 2023
Citations: 3	License type: cc-by-nc-nd

R Discovery Prime

R Discovery Prime

Distributed non-negative RESCAL with automatic model selection for exascale data

Abstract

Talk to us

Similar Papers

More From: Journal of Parallel and Distributed Computing

Lead the way for us

Similar Papers

Key performance indicators in Australian sub-elite rugby union
Tim J Mosey ... Lachlan J.G Mitchell
Journal of Science and Medicine in Sport | VOL. 23
Tim J Mosey, et. al.Tim J Mosey ... Lachlan J.G Mitchell
22 Aug 2019
Journal of Science and Medicine in Sport | VOL. 23

SIoT Framework to Build Smart Garage Sensors Based Recommendation System
A Soumya Mahalakshmi ... G S Sharvani
-
A Soumya Mahalakshmi, et. al.A Soumya Mahalakshmi ... G S Sharvani
01 Jan 2018
01 Jan 2018

Automatic detection of the support points in relational clustering
Parisa Rastin ... Rosanna Verde
-
Parisa Rastin, et. al.Parisa Rastin ... Rosanna Verde
01 Jul 2019
01 Jul 2019

Oil -Water Relative Permeability Data for Reservoir Simulation Input, Part-I: Systematic Quality Assessment and Consistency Evaluation
Mohammed Idrees Al-Mossawy ... Hossein Ali Algdamsi
-
Mohammed Idrees Al-Mossawy, et. al.Mohammed Idrees Al-Mossawy ... Hossein Ali Algdamsi
10 Dec 2014
10 Dec 2014

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Distributed non-negative RESCAL with automatic model selection for exascale data

Abstract

Talk to us

Similar Papers

More From: Journal of Parallel and Distributed Computing