Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems

Hao-Jun Michael Shi,Dheevatsa Mudigere,Maxim Naumov,Jiyan Yang

doi:10.1145/3394486.3403059

Abstract

Modern deep learning-based recommendation systems exploit hundreds to thousands of different categorical features, each with millions of different categories ranging from clicks to posts. To respect the natural diversity within the categorical data, embeddings map each category to a unique dense representation within an embedded space. Since each categorical feature could take on as many as tens of millions of different possible categories, the embedding tables form the primary memory bottleneck during both training and inference. We propose a novel approach for reducing the embedding size in an end-to-end fashion by exploiting complementary partitions of the category set to produce a unique embedding vector for each category without explicit definition. By storing multiple smaller embedding tables based on each complementary partition and combining embeddings from each table, we define a unique embedding for each category at smaller memory cost. This approach may be interpreted as using a specific fixed codebook to ensure uniqueness of each category's representation. Our experimental results demonstrate the effectiveness of our approach over the hashing trick for reducing the size of the embedding tables in terms of model loss and accuracy, while retaining a similar reduction in the number of parameters.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Impact of categorical and numerical features in ensemble machine learning frameworks for heart disease prediction
Chandan Pan ... Ajoy Kumar Ray
Biomedical Signal Processing and Control | VOL. 76
Chandan Pan, et. al.Chandan Pan ... Ajoy Kumar Ray
05 Apr 2022
Biomedical Signal Processing and Control | VOL. 76

The use of autoencoders for training neural networks with mixed categorical and numerical features
Łukasz Delong ... Anna Kozak
ASTIN Bulletin | VOL. 53
Łukasz Delong, et. al.Łukasz Delong ... Anna Kozak
24 Apr 2023
ASTIN Bulletin | VOL. 53

Adversarial attack on deep learning-based dermatoscopic image recognition systems: Risk of misdiagnosis due to undetectable image perturbations.
Jérôme Allyn ... Amélie Renou
Medicine | VOL. 99
Jérôme Allyn, et. al.Jérôme Allyn ... Amélie Renou
11 Dec 2020
Medicine | VOL. 99

Effectiveness of Data Augmentation in Cellular-based Localization Using Deep Learning
Hamada Rizk ... Ahmed Shokry
-
Hamada Rizk, et. al.Hamada Rizk ... Ahmed Shokry
01 Apr 2019
01 Apr 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems

Abstract

Talk to us

Similar Papers