Ember

Sahaana Suri,Theodoros Rekatsinas,Ihab F Ilyas,Christopher Ré

doi:10.14778/3494124.3494149

Abstract

Structured data, or data that adheres to a pre-defined schema, can suffer from fragmented context: information describing a single entity can be scattered across multiple datasets or tables tailored for specific business needs, with no explicit linking keys. Context enrichment, or rebuilding fragmented context, using keyless joins is an implicit or explicit step in machine learning (ML) pipelines over structured data sources. This process is tedious, domain-specific, and lacks support in now-prevalent no-code ML systems that let users create ML pipelines using just input data and high-level configuration files. In response, we propose Ember, a system that abstracts and automates keyless joins to generalize context enrichment. Our key insight is that Ember can enable a general keyless join operator by constructing an index populated with task-specific embeddings. Ember learns these embeddings by leveraging Transformer-based representation learning techniques. We describe our architectural principles and operators when developing Ember, and empirically demonstrate that Ember allows users to develop no-code context enrichment pipelines for five domains, including search, recommendation and question answering, and can exceed alternatives by up to 39% recall, with as little as a single line configuration change.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Ember

Abstract

Talk to us

Similar Papers

More From: Proceedings of the VLDB Endowment

Lead the way for us

Journal: Proceedings of the VLDB Endowment	Publication Date: Nov 1, 2021
Citations: 7

Similar Papers

Toward Rapid Development and Deployment of Machine Learning Pipelines across Cloud-Edge
Anirban Bhattacharjee ... Thomas Damiano
-
Anirban Bhattacharjee, et. al.Anirban Bhattacharjee ... Thomas Damiano
12 Aug 2021
12 Aug 2021

XAutoML: A Visual Analytics Tool for Understanding and Validating Automated Machine Learning
Marc-André Zöller ... Waldemar Titov
ACM Transactions on Interactive Intelligent Systems | VOL. 13
Marc-André Zöller, et. al.Marc-André Zöller ... Waldemar Titov
08 Dec 2023
ACM Transactions on Interactive Intelligent Systems | VOL. 13

A Comprehensive Machine Learning Benchmark Study for Radiomics-Based Survival Analysis of CT Imaging Data in Patients With Hepatic Metastases of CRC.
Anna Theresa Stüber ... David Rügamer
Investigative Radiology | VOL. 58
Anna Theresa Stüber, et. al.Anna Theresa Stüber ... David Rügamer
28 Jul 2023
Investigative Radiology | VOL. 58

AVATAR - Machine Learning Pipeline Evaluation Using Surrogate Model
Tien-Dung Nguyen ... Tomasz Maszczyk
-
Tien-Dung Nguyen, et. al.Tien-Dung Nguyen ... Tomasz Maszczyk
01 Jan 2020
01 Jan 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Ember

Abstract

Talk to us

Similar Papers

More From: Proceedings of the VLDB Endowment