An empirical study of on-line models for relational data streams

Ashwin Srinivasan,Michael Bain

doi:10.1007/s10994-016-5596-2

Abstract

To date, Inductive Logic Programming (ILP) systems have largely assumed that all data needed for learning have been provided at the onset of model construction. Increasingly, for application areas like telecommunications, astronomy, text processing, financial markets and biology, machine-generated data are being generated continuously and on a vast scale. We see at least four kinds of problems that this presents for ILP: (1) it may not be possible to store all of the data, even in secondary memory; (2) even if it were possible to store the data, it may be impractical to construct an acceptable model using partitioning techniques that repeatedly perform expensive coverage or subsumption-tests on the data; (3) models constructed at some point may become less effective, or even invalid, as more data become available (exemplified by the problem when identifying concepts); and (4) the representation of the data instances may need to change as more data become available (a kind of language drift problem). In this paper, we investigate the adoption of a stream-based on-line learning approach to relational data. Specifically, we examine the representation of relational data in both an infinite-attribute setting, and in the usual fixed-attribute setting, and develop implementations that use ILP engines in combination with on-line model-constructors. The behaviour of each program is investigated using a set of controlled experiments, and performance in practical settings is demonstrated by constructing complete theories for some of the largest biochemical datasets examined by ILP systems to date, including one with a million examples; to the best of our knowledge, the first time this has been empirically demonstrated with ILP on a real-world data set.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

An empirical study of on-line models for relational data streams

Abstract

Talk to us

Similar Papers

More From: Machine Learning

Lead the way for us

Journal: Machine Learning	Publication Date: Dec 3, 2016
Citations: 8

Similar Papers

Learning relational rule from examples that are neither positive nor negative
Ryutaro Ichise ... Masayuki Numao
Systems and Computers in Japan | VOL. 32
Ryutaro Ichise, et. al.Ryutaro Ichise ... Masayuki Numao
26 Nov 2001
Systems and Computers in Japan | VOL. 32

Feature Construction Using Theory-Guided Sampling and Randomised Search
Sachindra Joshi ... Ganesh Ramakrishnan
-
Sachindra Joshi, et. al.Sachindra Joshi ... Ganesh Ramakrishnan
10 Sep 2008
10 Sep 2008

Cautious induction: An alternative to clause-at-a-time hypothesis construction in inductive logic programming
Simon Anthony ... Alan M Frisch
New Generation Computing | VOL. 17
Simon Anthony, et. al.Simon Anthony ... Alan M Frisch
01 Mar 1999
New Generation Computing | VOL. 17

FastLAS: Scalable Inductive Logic Programming Incorporating Domain-Specific Optimisation Criteria
Mark Law ... Alessandra Russo
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34
Mark Law, et. al.Mark Law ... Alessandra Russo
03 Apr 2020
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An empirical study of on-line models for relational data streams

Abstract

Talk to us

Similar Papers

More From: Machine Learning