A Probabilistic Address Parser Using Conditional Random Fields and Stochastic Regular Grammar

Minlue Wang,Amos Yeo,John Howroyd,J Mark Bishop,Andrew Martin,Valeriia Haberland

doi:10.1109/icdmw.2016.0039

Abstract

Automatic semantic annotation of data from databases or the web is an important pre-process for data cleansing and record linkage. It can be used to resolve the problem of imperfect field alignment in a database or identify comparable fields for matching records from multiple sources. The annotation process is not trivial because data values may be noisy, such as abbreviations, variations or misspellings. In particular, overlapping features usually exist in a lexicon-based approach. In this work, we present a probabilistic address parser based on linear-chain conditional random fields (CRFs), which allow more expressive token-level features compared to hidden Markov models (HMMs). In additions, we also proposed two general enhancement techniques to improve the performance. One is taking original semi-structure of the data into account. Another is post-processing of the output sequences of the parser by combining its conditional probability and a score function, which is based on a learned stochastic regular grammar (SRG) that captures segment-level dependencies. Experiments were conducted by comparing the CRF parser to a HMM parser and a semi-Markov CRF parser in two real-world datasets. The CRF parser out-performed the HMM parser and the semi-Markov CRF in both datasets in terms of classification accuracy. Leveraging the structure of the data and combining the linear-chain CRF with the SRG further improved the parser to achieve an accuracy of 97% on a postal dataset and 96% on a company dataset.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A Probabilistic Address Parser Using Conditional Random Fields and Stochastic Regular Grammar

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Semi-Markov Conditional Random Fields のための損失関数スムージング
Kenta Fukuoka ... Yuji Matsumoto
Transactions of the Japanese Society for Artificial Intelligence | VOL. 22
Kenta Fukuoka, et. al.Kenta Fukuoka ... Yuji Matsumoto
01 Jan 2007
Transactions of the Japanese Society for Artificial Intelligence | VOL. 22

Machine Learning Basics
Krasimira Kapitanova ... Sang H Son
-
Krasimira Kapitanova, et. al.Krasimira Kapitanova ... Sang H Son
15 Dec 2012
15 Dec 2012

Equivalence between LC-CRF and HMM, and Discriminative Computing of HMM-Based MPM and MAP
Elie Azeraf ... Wojciech Pieczynski
Algorithms | VOL. 16
Elie Azeraf, et. al.Elie Azeraf ... Wojciech Pieczynski
21 Mar 2023
Algorithms | VOL. 16

On Equivalence between Linear-chain Conditional Random Fields and Hidden Markov Chains
Elie Azeraf ... Emmanuel Monfrini
-
Elie Azeraf, et. al.Elie Azeraf ... Emmanuel Monfrini
01 Jan 2021
01 Jan 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Probabilistic Address Parser Using Conditional Random Fields and Stochastic Regular Grammar

Abstract

Talk to us

Similar Papers