A Transformer-Based Framework for Scene Text Recognition

Prabu Selvam,Abolfazl Mehbodniya,Meshal Alharbi,Sudhakar Sengan,Carlos Andres Tavera Romero,Joseph Abraham Sundar Koilraj,Julian L Webber

doi:10.1109/access.2022.3207469

Abstract

Scene Text Recognition (STR) has become a popular and long-standing research problem in computer vision communities. Almost all the existing approaches mainly adopt the connectionist temporal classification (CTC) technique. However, these existing approaches are not much effective for irregular STR. In this research article, we introduced a new encoder-decoder framework to identify both regular and irregular natural scene text, which is developed based on the transformer framework. The proposed framework is divided into four main modules: Image Transformation, Visual Feature Extraction (VFE), Encoder and Decoder. Firstly, we employ a Thin Plate Spline (TPS) transformation in the image transformation module to normalize the original input image to reduce the burden of subsequent feature extraction. Secondly, in the VFE module, we use ResNet as the Convolutional Neural Network (CNN) backbone to retrieve text image features maps from the rectified word image. However, the VFE module generates one-dimensional feature maps that are not suitable for locating a multi-oriented text on two-dimensional word images. We proposed 2D Positional Encoding (2DPE) to preserve the sequential information. Thirdly, the feature aggregation and feature transformation are carried out simultaneously in the encoder module. We replace the original scaled dot-product attention model as in the standard transformer framework with an Optimal Adaptive Threshold-based Self-Attention (OATSA) model to filter noisy information effectively and focus on the most contributive text regions. Finally, we introduce a new architectural level bi-directional decoding approach in the decoder module to generate a more accurate character sequence. Eventually, We evaluate the effectiveness and robustness of the proposed framework in both horizontal and arbitrary text recognition through extensive experiments on seven public benchmarks including IIIT5K-Words, SVT, ICDAR 2003, ICDAR 2013, ICDAR 2015, SVT-P and CUTE80 datasets. We also demonstrate that our proposed framework outperforms most of the existing approaches by a substantial margin.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: IEEE Access	Publication Date: Jan 1, 2022
Citations: 14	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

A Transformer-Based Framework for Scene Text Recognition

Abstract

Talk to us

Similar Papers

More From: IEEE Access

Lead the way for us

Similar Papers

Occluded Text Detection and Recognition in the Wild
Zobeir Raisi ... John Zelek
-
Zobeir Raisi, et. al.Zobeir Raisi ... John Zelek
01 May 2022
01 May 2022

Improved Telugu Scene Text Recognition with Thin Plate Spline Transform
Nandam Srinivasa Rao ... Atul Negi
-
Nandam Srinivasa Rao, et. al.Nandam Srinivasa Rao ... Atul Negi
01 Jan 2021
01 Jan 2021

Scene text detection and recognition with advances in deep learning: a survey
Xiyan Liu ... Chunhong Pan
International Journal on Document Analysis and Recognition (IJDAR) | VOL. 22
Xiyan Liu, et. al.Xiyan Liu ... Chunhong Pan
27 Mar 2019
International Journal on Document Analysis and Recognition (IJDAR) | VOL. 22

Cursive Text Recognition in Natural Scene Images Using Deep Convolutional Recurrent Neural Network
Asghar Ali Chandio ... Mark R Pickering
IEEE Access | VOL. 10
Asghar Ali Chandio, et. al.Asghar Ali Chandio ... Mark R Pickering
01 Jan 2021
IEEE Access | VOL. 10

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Transformer-Based Framework for Scene Text Recognition

Abstract

Talk to us

Similar Papers

More From: IEEE Access