End-to-End Scene Text Recognition in Videos Based on Multi Frame Tracking

Xiaobing Wang,Yingying Jiang,Zhenbo Luo,Shuli Yang,Xiangyu Zhu,Wei Li,Hua Wang,Pei Fu

doi:10.1109/icdar.2017.207

Xiaobing Wang, Yingying Jiang + Show 6 more

https://doi.org/10.1109/icdar.2017.207

Copy DOI

Export

Save

Cite

Publication Date: Nov 1, 2017

Citations: 22

Affiliation: Samsung (China)

Abstract
Full-Text
Similar Papers

Abstract

Listen

Text detection and recognition in scene images and videos attract much attention in computer vision recently. However, most existing text detection and recognition methods only focus on static images. In this paper an end-to-end scene text recognition method based on multi frame tracking is proposed for text in videos, in which temporal information is employed to improve performance. First, an end-to-end text recognition method based on a unified deep neural network is used to detect and recognize text in each frame of the input video. Then, multi frame text tracking is employed through associations of texts in current frame and several previous frames to obtain final results. Experiments on ICDAR datasets demonstrate that the proposed method outperforms the state-of-the-art methods in end-to-end video text recognition.

Full Text