Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

A machine learning approach for nominative record linkage in Chinese historical databases

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

We introduce a generic machine learning-based pipeline for nominative linkage of records within and across Chinese historical datasets. The pipeline addresses key challenges, including character variations, incomplete data, and scalability issues specific to historical datasets in which names and other attributes are recorded with Chinese characters, not just for China, but potentially for Korea, Japan and Vietnam. Techniques developed for attributes recorded in phonetic alphabets are of limited use for Chinese characters not only because homonyms are common, but characters that are similar enough in appearance to be mistaken for each other may sound different. Our approach integrates stroke-based character embeddings for efficient blocking, supervised classification with active learning for record matching, and graph-based clustering for final linkage. We demonstrate the effectiveness of this pipeline using the career records of officials in the China Government Employee Database-Qing Jinshenlu (CGED-Q JSL). We achieve improved linkage quality compared to standard probabilistic methods, with longer linked sequences of career records and fewer aberrant transitions. To validate the generalizability, we also successfully apply the pipeline to another database and a cross-database linkage task. By minimizing the need for manual tuning, our pipeline offers a more accessible and effective solution for Chinese historical data linkage.

Similar Papers
  • Research Article
  • 10.54254/2753-7048/2025.km25972
Analyzing the Irreplaceable Characteristic of Chinese Characters: A Case Study of Phonetic Chinese Alphabet
  • Aug 13, 2025
  • Lecture Notes in Education Psychology and Public Media
  • Zhiqing Ling

As the only official writing system that have been popularizing in China, Chinese characters has an impregnable position. Some researchers in movements for re-forming Chinese characters in history, however, had contested Chinese characters and tried to popularize various phonetic Chinese alphabet schemes to illiterate people in the public. This research using the argument between Chinese characters and Phonetic Chinese alphabet and the limitations of phonetic Chinese alphabet during the period of the phonetic alphabet campaign as an example to demonstrate the irreplaceable and valuable of Chinese characters. Ultimately, this research can demonstrate that whenever in cultural transmission aspect or that of practical apply the current form of Chinese characters cannot be re-placed. However, as phonetic Chinese alphabet has limitations so it can only be a supporting tool to learn Chinese characters and Chinese language (including dialects of Chinese). This research can deepen the understanding of Chinese characters and phonetic Chinese alpha-bet, and demonstrate the unique value of Chinese characters and Chinese culture and history.

  • Research Article
  • 10.62517/jnme.202410106
Analysis on the Design Method of Chinese Character Logo Graphics
  • Jan 1, 2024
  • Journal of New Media and Economics
  • Ju Song

Chinese characters, the recorded symbols of the Chinese language, are one of the oldest written characters in the world, with a history of more than 6,000 years, and are considered to be the model of phonetic, morphological and ideographic writing. The amount of information, the richness of meaning and the clarity of a single Chinese character far exceed that of the phonetic Latin alphabet. As the most basic visual elements of information transmission, Chinese characters themselves have strong intuitiveness and cognition, and are often used in modern logo design. Firstly, this paper summarizes the advantages of Chinese characters, Chinese character signs and Chinese character signs. Secondly, the design method of Chinese character logo is analyzed from the aspects of Chinese character font design, Chinese character graphic design and the application of calligraphy form by using excellent Chinese character logo cases. Finally, this paper analyzes the design principles of Chinese character logo graphics, and points out several design points that should be paid attention to in Chinese character logo graphics design. The purpose of this paper is to arouse the deep thinking and research of logo designers on the modern logo graphic design based on Chinese characters, to dig deeply into the traditional cultural connotation and implication of Chinese characters, and to inherit and carry forward the excellent traditional culture of the Chinese nation.

  • Conference Article
  • 10.2991/sschd-16.2016.27
Analyzing of Chinese Character Form Database (CCFD) and the Study of CCFD
  • Jan 1, 2016
  • Jian-Yu Liu

After the combination of the study of Chinese character and the computer information science, Chinese Character Form Database (CCFD) has been a new thing, which can meet the objective requirement of the current collecting and research work for Chinese character. CCFD is an important foundation of modernization and informatization of Chinese character. The Study of CCFD is an interdisciplinary subject, based on theoretical and practical researches, and the study of CCFD can provide a powerful theoretical weapon for the collecting and the research of Chinese character in the information era. At present, the hypostatic construction of CCFD and the systematic research of the study of CCFD are still in the primary stage, but there will be a promising and significant development in the future. The 21st century is the information age and Big Data era. Language and script are the two most important information carriers. Language informatization is the basis of the informatization of the whole society. By the means of information, the systematic, comprehensive and in-depth study of Chinese and Chinese character is necessary for China to be in an invincible position in the 21st century. Innovation is the life of academic scholarship, and academic development is inseparable from the use of new methods and new techniques. In the past 20 years, with the gradual popularization of computers, computer technology has been applied to study Chinese character and gotten a lot of theoretical and practical results; Chinese information processing has become a core area in the current study of Chinese and Chinese character. However, comparing with the mature Chinese corpus and Chinese corpus linguistics, it is still relatively weak to collect and study Chinese character by computer technology and fall far behind the objective requirements of character collection normative work. To change this situation, we must use Chinese Character Form Database (CCFD), and establish a new discipline—the study of CCFD as soon as possible to guide practical work of Chinese character collecting and research.

  • Conference Article
  • Cite Count Icon 4
  • 10.1109/iscid.2016.2037
Chinese Character Recognition in Natural Scenes
  • Dec 1, 2016
  • Shuyou He + 1 more

This paper focuses in particular on the problem of Chinese characters recognition in natural scenes. Due to large variation in fonts, sizes, illumination, cluttered backgrounds, geometric distortions, etc., scene text recognition in the wild is a challenging problem. We proposed a novel method which based on Integral Channel Feature and pooling technology to extract informative features from scenes images. We concatenated many different low-level features and selected typically features to represent the Chinese characters. In this work, we make use of Support Vector Machines as the classifier, and rank the features by the weights of training model in LinearSVM. Thus features representation of characters is compact and it is effective to express distinctive spatial structures of text character. At the same time, for comparative purpose, we evaluated approach extensively on two standard dataset (ICHAR03, Char74K). Because of the absence of Chinese character datasets, we took 311 photos by cellphone and digit camera in Guangzhou and cropped them to pieces manually to form two Chinese datasets. Our experiment results denote that our proposed technology was performed better than the current state-of-the-art methods.

  • Research Article
  • 10.58224/2687-0428-2024-6-1-153-161
К вопросу об использовании пиньиня при формировании лексических навыков на начальной ступени обучения китайскому языку
  • Feb 9, 2024
  • Review of pedagogical research
  • Т.Г Чарчоглян

Аннотация: в настоящей статье рассматривается проблема освоения лексики китайского языка в условиях наличия особого вида письменности – иероглифики, представляющей собой зрительное выражение смысла слова; обосновывается использование в процессе обучения китайскому языку фонетического алфавита китайского языка – пиньиня как вспомогательной письменной системы. Целью работы сталоисследование уровня сформированности лексических навыков у студентов первого курса лингвистического факультета педагогического вуза на предмет его корреляции с параллельным использованием фонетического алфавита китайского языка – пиньинем. В качестве методов исследования были выбраны как теоретические (изучение и анализ научно-методических работ российских и зарубежных авторов, сравнение, обобщение и синтез актуальных для данной работы результатов научных исследований), так и эмпирически (тестирование и опрос обучающихся с последующей декомпозицией и статистической обработкой результатов) методы. Результаты исследования подтвердили выдвинутую гипотезу о негативном влиянии сопровождения фонографическим вариантом письменности приводимых в иероглифической записи текстов УМК на формирование начальных лексических навыков студентов в части оперирования основным (иероглифическим) вариантом на продуктивном и рецептивном уровнях. Автором предлагается в обучении китайскому языку даже на начальном этапе ограничить традиционное дублирование текстов пиньинем, для чего при необходимости внести изменения в материалы УМК. Abstract: this article researches the problem of mastering the vocabulary of the Chinese language in the presence of a special type of writing – Chinese characters, which is a visual expression of the meaning of a word; the use of the Chinese phonetic alphabet, Pinyin, as an auxiliary writing system in the process of teaching Chinese is substantiated. The purpose of the work is to study the level of development of lexical skills among first-year students of the linguistic faculty of a pedagogical university for its correlation with the parallel use of the phonetic alphabet of the Chinese language – pinyin. The research methods chosen are both theoretical (study and analysis of scientific and methodological works of Russian and foreign authors, comparison, generalization and synthesis of scientific research results relevant to this work) and empirical (testing and survey of students with subsequent decomposition and statistical processing of the results) methods. The results of the study confirm the hypothesis put forward about the negative impact of accompanying the phonographic version of writing of teaching texts given in hieroglyphic notation on the formation of students’ initial lexical skills in terms of operating with the main (characters) version at the productive and receptive levels. The author proposes to limit the traditional duplication of texts in pinyin in teaching the Chinese language, even at the initial stage, for which, if necessary, make changes to the teaching materials.

  • Conference Article
  • Cite Count Icon 5
  • 10.1109/cvidl51233.2020.00-28
Chinese Character Captcha Sequential Selection System Based on Convolutional Neural Network
  • Jul 1, 2020
  • Xueting Bi + 1 more

To ensure security, Completely Automated Public Turing test to tell Computers and Humans Apart (CAPTCHA) is widely used in people's online lives. This paper presents a Chinese character captcha sequential selection system based on convolutional neural network (CNN). Captchas composed of English and digits can already be identified with extremely high accuracy, but Chinese character captcha recognition is still challenging. The task we need to complete is to identify Chinese characters with different colors and different fonts that are not on a straight line with rotation and affine transformation on pictures with complex backgrounds, and then perform word order restoration on the identified Chinese characters. We divide the task into several sub-processes: Chinese character detection based on Faster R-CNN, Chinese character recognition and word order recovery based on N-Gram. In the Chinese character recognition sub-process, we have made outstanding contributions. We constructed a single Chinese character data set and built a 10-layer convolutional neural network. Eventually we achieved an accuracy of 98.43%, and completed the task perfectly.

  • Conference Article
  • Cite Count Icon 1
  • 10.1109/ijcnn.2008.4634254
Chinese Character identification by visual features using self-organizing map sets and relevance feedback
  • Jun 1, 2008
  • James S Kirk

Because of its ability to condense a data set in a non-linear, dimension-reducing, topology-preserving way, the self-organizing map (SOM) has proven useful in a wide variety of applications. The Chinese Character Identifier (CCI) uses a set of SOMs along with other natural computation tools to address the problem of identifying an unknown Chinese character by its visual features. By repeatedly presenting small sets of Chinese characters to the user and analyzing which characters are chosen as visually similar to the target character, the system is intended to estimate the visual features upon which the user is presently basing his/her notion of visual similarity. An SOM is then chosen that organizes the universe of characters according to the userpsilas feedback. A simple radial basis function network with basis functions defined in the output space of the selected SOM is used to select a set of characters to present to the user next. The result is a trajectory across the 10-dimensional feature space of the Chinese characters in the direction of the target character. The CCI illustrates the promises and the challenges of using a method of searching high-dimensional data based on relevance feedback that may be termed ldquopiecewise topography preservationrdquo (PTP). This paper discusses the application of PTP to a set of 10-dimensional Chinese character data and explains why certain data sets, exemplified by the Chinese character data, pose a problem for the PTP approach.

  • Conference Article
  • Cite Count Icon 3
  • 10.1109/icdar.2019.00222
HITHCD-2018: Handwritten Chinese Character Database of 21K-Category
  • Sep 1, 2019
  • Tonghua Su + 2 more

Current state of handwritten Chinese character recognition (HCCR) conducted on well-confined character set, far from meeting industrial requirements. The paper describes the creation of a large-scale handwritten Chinese character database. Constructing the database is an effort to scale up Chinese handwritten character classification task to cover the full list of GBK character set specification. It consists of 21-thousand Chinese character categories and 20-million character images, larger than previous databases both in scale and diversity. We present solutions to the challenges of collecting and annotating such large-scale handwritten character samples. We elaborately design the sampling strategy, extract salient signals in a systematic way, annotate the tremendous characters through three distinct stages. Experiments are conducted the generalization to other handwritten character databases and our database demonstrates great values. Surely, its scale opens unprecedented opportunities both in evaluation of character recognition algorithms and in developing new techniques.

  • Research Article
  • Cite Count Icon 7
  • 10.1142/s021800141250005x
RADICAL EXTRACTION USING AFFINE SPARSE MATRIX FACTORIZATION FOR PRINTED CHINESE CHARACTERS RECOGNITION
  • May 1, 2012
  • International Journal of Pattern Recognition and Artificial Intelligence
  • Jun Tan + 3 more

Each Chinese character is comprised of radicals, where a single character (compound character) contains one (or more than one) radicals. For human cognitive perspective, a Chinese character can be recognized by identifying its radicals and their spatial relationship. This human cognitive law may be followed in computer recognition. However, extracting Chinese character radicals automatically by computer is still an unsolved problem. In this paper, we propose using an improved sparse matrix factorization which integrates affine transformation, namely affine sparse matrix factorization (ASMF), for automatically extracting radicals from Chinese characters. Here the affine transformation is vitally important because it can address the poor-alignment problem of characters that may be caused by internal diversity of radicals and image segmentation. Consequently we develop a radical-based Chinese character recognition model. Because the number of radicals is much less than the number of Chinese characters, the radical-based recognition performs a far smaller category classification than the whole character-based recognition, resulting in a more robust recognition system. The experiments on standard Chinese character datasets show that the proposed method gets higher recognition rates than related Chinese character recognition methods.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.3390/su14052920
Combined Searches of Chinese Language and English Language Databases Provide More Comprehensive Data on the Distribution of Five Pest Thrips Species in China for Use in Pest Risk Assessment
  • Mar 2, 2022
  • Sustainability
  • Bingqin Xu + 1 more

Background: Globally, China and the USA are thought to present the greatest biosecurity threat from invasive species given the invasive species they already contain and their trade patterns. A proportion of Chinese scientific publications are published in Chinese language journals in Chinese characters, thus, they are not easily available to the international biosecurity community. Information in these journals may be important for invasive species biosecurity risk assessment. Methods: To assess the need for retrieving information from non-international databases, such as Chinese databases, we compared quantitative and qualitative information on the presence and distribution of five invasive pest thrips species (Frankliniella schultzei, Selenothrips rubrocinctus, Scirtothrips dorsalis, Thrips hawaiiensis, and Thrips palmi) in China, retrieved from an international English language database (Web of Science/WOS) and a Chinese language database (Chinese National Knowledge Infrastructure/CNKI). Such information is necessary for climate matching models which are used regularly for pest risk assessment. Results: Few publications on Frankliniella schultzei were found in either database. For the other species, more publications were sourced from CNKI than WOS. More publications on the provincial distribution of S. rubrocinctus and S. dorsalis in China were found in CNKI than the Crop Protection Compendium (CPC); the two sources had equivalent publications on T. palmi and T. hawaiiensis. The combined provincial distributional data from WOS, CNKI and CPC for the four species provided distribution records at a higher latitude than a recently published checklist—information that is important for optimised climate matching. Additionally, CNKI provided sub-provincial distributional data not available in CPC that will enable a more refined approach for climate matching. Data on the relative proportion of publications found in different databases were constant over time. Conclusions: This study, focusing on pest distributional data, illustrates the importance of searching in Chinese databases in combination with standard searches in international databases, to gain a comprehensive understanding of invasive species for biosecurity risk assessment.

  • Book Chapter
  • Cite Count Icon 1
  • 10.1007/978-3-540-87599-4_51
A Mechanism for Solving the Unencoded Chinese Character Problem on the Web
  • Jan 1, 2008
  • Te-Jun Lin + 5 more

The unencoded Chinese character problem that occurs when digitizing historical Chinese documents makes digital archiving difficult. Expanding the character coding space, such as by using the Unicode Standard, does not solve the problem completely due to the extensibility of Chinese characters. In this paper, we propose a mechanism based on a Chinese glyph structure database, which contains glyph expressions that represent the composition of Chinese characters. Users can search for Chinese characters through our web interface and browse the search results. Each Chinese character can be embedded in a web document using a specific Java Script code. When the web document is opened, the Java Script code will load the image of the Chinese character in an appropriate font size for display. Even if the Chinese characters are not available in the database, their images can be generated through the dynamic character composition function. As the proposed mechanism is cross-platform, users can easily access unencoded Chinese characters without installing any additional font files in their personal computers. A demonstration system is available at http://char.ndap.org.tw.KeywordsChinese glyph structure databasedigital archiveunencoded Chinese characters

  • Research Article
  • Cite Count Icon 1
  • 10.1353/lan.1998.0184
Writing and literacy in Chinese, Korean and Japanese By Insup Taylor and M. Martin Taylor (review)
  • Jun 1, 1998
  • Language
  • Mary S Erbaugh

376LANGUAGE, VOLUME 74, NUMBER 2 (1998) and nonperipheral. These classes define themselves in English dialects by their differential behavior in ongoing processes of sound change. The peripheral vowels in general rise while the nonperipheral ones fall. In addition, for American English, at least, the peripheral vowels have all developed colored off-glides (the source of the infamous IyI and /w/ glides in the TragerSmith transcription), while the nonperipheral ones have in general developed colorless s-offglides : 'bead' [biid] vs. 'bid' [bisd], for example. Similarly, in Quebec French the high vowels [i,y,u] have developed lax counterparts [i, y, u] in closed syllables, and a process of lax harmony spreads the lax feature leftwards across the rest of the word: abusif [abY'zif] 'abusive', inutile [my'tsil] 'useless'. The close relationship between vowels that appear to be adjacent in height 'within a zone' (much as we find between related members of the major articulatory zones) suggests that there is a single feature controlling this relationship, and tense/lax still seems the best available suggestion, despite the unease of the phonetics community. With the exception ofthe issue oftenseness, I am unable to present a single significant criticism ofthis book. On the back ofthe paperback edition John Goldsmith, Michael Kenstowicz, William Hardcastle, and W. Barry are all quoted praising this book as something that needs to be on everyone's bookshelf. I can only endorse the nomination. REFERENCES Halle, Morris. 1983. On distinctive features and their articulatory implementation. Natural Language and Linguistic Theory 1.91-105. Hurch, Bernhard. 1988. Über aspiration: Ein Kapitel aus der natürlichen Phonologie. Tübingen: Gunter Narr Verlag. Labov, William. 1994. Principles of linguistic change. Volume 1: Internal factors. Cambridge, MA: Blackwell . Lindblom, Björn, and Ian Maddieson. 1988. Phonetic universale in consonant systems. Language, speech and mind: Studies in honor of Victoria A. Fromkin, ed. by Larry Hyman and Charles Li, 62-80. London and New York: Routledge. McCarthy, John. 1988. Feature geometry and depency: A review. Phonetica 45:84-108. Nathan, Geoffrey S. 1989. Preliminaries to a theory of phonological substance: The substance of sonority. Linguistic categorization, ed. by Roberta Corrigan, Fred Eckman, and Michael Noonan, 55-67. Amsterdam and Philadelphia: Benjamins. -----. 1995. How the phoneme inventory gets its shape—cognitive grammar' s view of phonological systems. Rivista di Lingüistica 6.275-88. Sagey, Elizabeth. 1990. The representation of features in non-linear phonology: The articulator node hierarchy . New York: Garland. Stevens, Kenneth. 1989. On the quantal nature of speech. Journal of Phonetics 17.3-46. ----- and Samuel Jay Keyser. 1989. Primary features and their enhancement in consonants. Language 65.81-106. Department of Linguistics Southern Illinois University Carbondale, IL 62901-4517 [geoffn@siu.edu] Writing and literacy in Chinese, Korean and Japanese. By Insup Taylor and M. Martin Taylor. Amsterdam & Philadelphia: John Benjamins, 1995. Pp. 412. $68.00. Reviewed by Mary S. Erbaugh, City University of Hong Kong Chinese characters are called hanzi in China, kanji in Japan, and hancha in Korea. About twothirds of characters contain phonetic elements which Chinese speakers find accessible enough to read and write exclusively in characters. Non-Chinese-speakers, Koreans, and Japanese who borrowed the classical Chinese script a millennium ago have developed highly efficient syllabaries to supplement the characters, which are still used for a decreasing number of culturally REVIEWS377 important content words as well as many proper names. Insup Taylor, Korean-born and also fluent in Japanese, enlisted M. Martin Taylor in taking on two huge, rich topics: the psycholinguistics of reading characters in Chinese, Japanese, and Korean; and the historical sociolinguistics ofliteracy in those languages. No other work compares to this in depth, scope, or in the richness of its primary sources in Chinese, Japanese, Korean, and English. The authors argue strongly for the efficiency of characters . Characters, they say, should be strengthened and retained, not only in China but in Korea and Japan as well, especially if they are systematically selected and 'streamlined'. I. Taylor is a member of the Comparative Literacy Project in the McLuhan Program in Culture and Technology , but the book refutes McLuhan's 1962 statement, 'Cultures can rise far above civilization artistically, but without the phonetic alphabet they remain...

  • Research Article
  • Cite Count Icon 2
  • 10.2352/issn.2470-1173.2018.2.vipc-174
Generative Adversarial Networks for Open Set Historical Chinese Character Recognition
  • Jan 28, 2018
  • Electronic Imaging
  • Xiaoyi Yu + 2 more

Historical Chinese character recognition has been suffering from the problem of samples labeling, not only the problem of lacking sufficient labeled training samples, but also of sample classes. So the scenario for Historical Chinese character recognition is "open set" recognition, where incomplete labeling of sample classes is present at training time, and unknown classes can be submitted to the system during testing. This paper proposes a method for open set Historical Chinese Character Recognition. For open set recognition, the features available in the training data cannot effectively characterize different kinds of unknown classes. We assume that the features which characterize unknown classes can be derived or learned from other similar data sets. We utilize an auxiliary data set combined with the open set training data set to learn good features to represent historical Chinese characters. The auxiliary data set is translated using Generative Adversarial Networks (GAN) to make sure that the translated data set is as close to the historical Chinese character dataset as possible. Then we construct a neural network for features extraction. The neural network is trained using an alternative training method with the translated auxiliary dataset and incomplete labeled historical Chinese character data set. Last, features are extracted from certain layer of the trained neural network. Unknown samples are detected using statistical modelling of the Euclidean metric between samples. Experimental results show that the proposed method is effective.

  • Conference Article
  • Cite Count Icon 14
  • 10.1109/icpr48806.2021.9412439
A Transformer-based Radical Analysis Network for Chinese Character Recognition
  • Jan 10, 2021
  • Chen Yang + 5 more

Recently, a novel radical analysis network (RAN) has the capability of effectively recognizing unseen Chinese character classes and largely reducing the requirement of training data by treating a Chinese character as a hierarchical composition of radicals rather than a single character class. However, when dealing with more challenging issues, such as the recognition of complicated characters, low-frequency character categories, and characters in natural scenes, RAN still has a lot of room for improvement. In this paper, we explore options to further improve the structure generalization and robustness capability of RAN with the Transformer architecture, which has achieved start-of-the-art results for many sequence-to-sequence tasks. More specifically, we propose to replace the original attention module in RAN with the transformer decoder, which is named as a transformer-based radical analysis network (RTN). The experimental results show that the proposed approach can significantly outperform the RAN on both printed Chinese character database and natural scene Chinese character database. Meanwhile, further analysis proves that RTN can be better generalized to complex samples and low-frequency characters, and has better robustness in recognizing Chinese characters with different attributes.

  • Book Chapter
  • Cite Count Icon 19
  • 10.1007/978-3-030-69544-6_39
Self-supervised Learning of Orc-Bert Augmentor for Recognizing Few-Shot Oracle Characters
  • Jan 1, 2021
  • Wenhui Han + 4 more

This paper studies the recognition of oracle character, the earliest known hieroglyphs in China. Essentially, oracle character recognition suffers from the problem of data limitation and imbalance. Recognizing the oracle characters of extremely limited samples, naturally, should be taken as the few-shot learning task. Different from the standard few-shot learning setting, our model has only access to large-scale unlabeled source Chinese characters and few labeled oracle characters. In such a setting, meta-based or metric-based few-shot methods are failed to be efficiently trained on source unlabeled data; and thus the only possible methodologies are self-supervised learning and data augmentation. Unfortunately, the conventional geometric augmentation always performs the same global transformations to all samples in pixel format, without considering the diversity of each part within a sample. Moreover, to the best of our knowledge, there is no effective self-supervised learning method for few-shot learning. To this end, this paper integrates the idea of self-supervised learning in data augmentation. And we propose a novel data augmentation approach, named Orc-Bert Augmentor pre-trained by self-supervised learning, for few-shot oracle character recognition. Specifically, Orc-Bert Augmentor leverages a self-supervised BERT model pre-trained on large unlabeled Chinese characters datasets to generate sample-wise augmented samples. Given a masked input in vector format, Orc-Bert Augmentor can recover it and then output a pixel format image as augmented data. Different mask proportion brings diverse reconstructed output. Concatenated with Gaussian noise, the model further performs point-wise displacement to improve diversity. Experimentally, we collect two large-scale datasets of oracle characters and other Chinese ancient characters for few-shot oracle character recognition and Orc-Bert Augmentor pre-training. Extensive experiments on few-shot learning demonstrate the effectiveness of our Orc-Bert Augmentor on improving the performance of various networks in the few-shot oracle character recognition.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant