On the Use of Transformers for End-to-End Optical Music Recognition

Antonio Ríos-Vila,José M Iñesta,Jorge Calvo-Zaragoza

doi:10.1007/978-3-031-04881-4_37

Abstract

State-of-the-art end-to-end Optical Music Recognition (OMR) systems use Recurrent Neural Networks to produce music transcriptions, as these models retrieve a sequence of symbols from an input staff image. However, recent advances in Deep Learning have led other research fields that process sequential data to use a new neural architecture: the Transformer, whose popularity has increased over time. In this paper, we study the application of the Transformer model to the end-to-end OMR systems. We produced several models based on all the existing approaches in this field and tested them on various corpora with different types of encodings for the output. The obtained results allow us to make an in-depth analysis of the advantages and disadvantages of applying this architecture to these systems. This discussion leads us to conclude that Transformers, as they were conceived, do not seem to be appropriate to perform end-to-end OMR, so this paper raises interesting lines of future research to get the full potential of this architecture in this field.KeywordsOptical Music RecognitionTransformersConnectionist Temporal ClassificationImage-to-sequence

Full Text