Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions

Ce Zhang,Guowei Zhang,Weilan Wang

doi:10.3390/electronics11233919

Ce Zhang, Guowei Zhang + Show 1 more

Open Access

https://doi.org/10.3390/electronics11233919

Copy DOI

Abstract

The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Electronics	Publication Date: Nov 27, 2022
Citations: 1	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions

Abstract

Talk to us

Similar Papers

More From: Electronics

Lead the way for us

Similar Papers

A Script Independent Hybrid Feature Extraction Technique for Offline Handwritten Devanagari and Bangla Character Recognition
Raghunath Dey ... Jayashree Piriz
-
Raghunath Dey, et. al.Raghunath Dey ... Jayashree Piriz
19 Dec 2021
19 Dec 2021

Full depth CNN classifier for handwritten and license plate characters recognition
Mohammed Salemdeeb ... Sarp Ertürk
PeerJ Computer Science | VOL. 7
Mohammed Salemdeeb, et. al.Mohammed Salemdeeb ... Sarp Ertürk
18 Jun 2021
PeerJ Computer Science | VOL. 7

Developing children's ASR system under low-resource conditions using end-to-end architecture
Ankita ... S Shahnawazuddin
Digital Signal Processing | VOL. 146
Ankita, et. al. Ankita ... S Shahnawazuddin
08 Jan 2024
Digital Signal Processing | VOL. 146

Morphology based Character Recognition of Overlapped and Touched Objects
Nafis Uddin Khan ... Pallavi Sharma
-
Nafis Uddin Khan, et. al.Nafis Uddin Khan ... Pallavi Sharma
01 Mar 2019
01 Mar 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions

Abstract

Talk to us

Similar Papers

More From: Electronics