The Anglicization of LGBTQ+ Slang in Spanish
This study analyzes the typology and pragmatic functions of anglicisms found in Spanish LGBTQ+ slang using language data from a self-created corpus, Queer_ES_Anglicisms, extracted from the Spanish queer magazines Shangay and Togayther. The anglicisms are first classified according to Pulcini et al.’s typology and then arranged and examined in terms of their semantic fields and their morphopragmatic features. Furthermore, the paper includes a case study of the lemmas queer, drag, and gay as well as an analysis of the ideational and expressive pragmatic functions of the anglicisms to unravel the motivations behind their use in the context of Spanish LGBTQ+ slang.
- Research Article
8
- 10.1145/3582496
- Jun 17, 2023
- ACM Transactions on Asian and Low-Resource Language Information Processing
Extracting and analysing meaning-related information from natural language data has attracted the attention of researchers in various fields, such as natural language processing, corpus linguistics, information retrieval, and data science. An important aspect of such automatic information extraction and analysis is the annotation of language data using semantic tagging tools. Different semantic tagging tools have been designed to carry out various levels of semantic analysis, for instance, named entity recognition and disambiguation, sentiment analysis, word sense disambiguation, content analysis, and semantic role labelling. Common to all of these tasks, in the supervised setting, is the requirement for a manually semantically annotated corpus, which acts as a knowledge base from which to train and test potential word and phrase-level sense annotations. Many benchmark corpora have been developed for various semantic tagging tasks, but most are for English and other European languages. There is a dearth of semantically annotated corpora for the Urdu language, which is widely spoken and used around the world. To fill this gap, this study presents a large benchmark corpus and methods for the semantic tagging task for the Urdu language. The proposed corpus contains 8,000 tokens in the following domains or genres: news, social media, Wikipedia, and historical text (each domain having 2K tokens). The corpus has been manually annotated with 21 major semantic fields and 232 sub-fields with the USAS (UCREL Semantic Analysis System) semantic taxonomy which provides a comprehensive set of semantic fields for coarse-grained annotation. Each word in our proposed corpus has been annotated with at least one and up to nine semantic field tags to provide a detailed semantic analysis of the language data, which allowed us to treat the problem of semantic tagging as a supervised multi-target classification task. To demonstrate how our proposed corpus can be used for the development and evaluation of Urdu semantic tagging methods, we extracted local, topical and semantic features from the proposed corpus and applied seven different supervised multi-target classifiers to them. Results show an accuracy of 94% on our proposed corpus which is free and publicly available to download.
- Research Article
- 10.17846/topling-2024-0009
- Dec 23, 2024
- Topics in Linguistics
The paper presents the results of longitudinal research based on authentic data analysis. The research subject was a Czech-speaking monolingual boy. Recordings of the subject’s dialogues in communication with adults from the age of three years to the age of 10 years were analysed; parental diaries were used as supplementary material. The author observed the appearance of words of foreign origin in the child’s speech. The frequency of words, the degree of their adaptation, and the correctness of use in terms of meaning, grammatical form and pronunciation were monitored. The research also focused on which language the words were taken from and in which semantic field they were located. Some tendencies were traced, e.g. the increasing frequency of words of foreign origin in the child’s production over time. An interesting factor was incorrect declension; the child did not respect the difference of the declension paradigm and classified the lexeme with prototypical Czech noun patterns. A shift in semantic fields was also evident. Initially, the child used commonly adapted lexemes. Later, lexemes from English were added, falling into the sociolect of computer-game players and lexemes of the youth sociolect. With the onset of schooling, the number of foreign-language terms from various areas increased, but those lexemes were not often used in the recordings as they were not in the child’s area of interest.
- Research Article
- 10.1111/1467-9582.00040
- Apr 1, 1999
- Studia Linguistica
The purpose of this paper is to demonstrate, following Jackendoff (1992), that parallelism across semantic fields is constrained by properties specific to each field, and that these field‐specific properties play a crucial role in accounting for the non‐parallels which particular lexical items exhibit across semantic fields.Three semantic fields (Temporal, Possessional, and Identificational) are shown to have different properties with respect to three parameters, dimensionality, directedness and continuousness. These differing, field‐specific properties account for where parallelism obtains and where its does not.The examination of field‐specific properties reveals a number of cross‐field behavioral differences that have received little recognition, as shown by case studies of spread, between, over, and extent verbs.
- Research Article
8
- 10.1016/j.dcm.2019.100317
- Aug 6, 2019
- Discourse, Context & Media
The emergence and evolution of master terms in the public debate about livestock farming: Semantic fields, communication strategies and policy practices
- Research Article
6
- 10.2478/stap-2021-0013
- Dec 1, 2021
- Studia Anglica Posnaniensia
Research on anglicisms in Polish has nearly a century-long tradition, yet it was Jacek Fisiak’s 1960s–1980s studies on English loanwords that initiated continuous academic interest in anglicisms, coinciding with more intensive English-Polish language contact in post-war Poland. While English loans have been well-researched in the last four decades, the ongoing intensity of English lexical influence on Polish, yielding not only new loans but also new loan types, calls for further studies, especially in the area of quickly developing professional jargons and sociolects. The influx of English-sourced lexis is reflected in the diversity of semantic fields, whose number has grown from 18 (identified in Słownik warszawski 1900–1927) to 45 (Mańczak-Wohlfeld 1995). A semantic field that has been underresearched in studies on Polish anglicisms is the LGBTQ+-related lexis, which has drawn from American English gayspeak, shaped by the post-Stonewall gay rights movement initiated in the 1970s. The language data analysed in this study have been collected in a two-stage procedure, which included manual extraction of anglicisms sourced in a diversified corpus of LGBTQ+-related written texts, published in Polish between 2004 and 2020. The second stage involved oral interviews which served a verification function. The aim of this study is to contribute to the lexicographic attempts at researching English-sourced LGBTQ+-related vocabulary in Polish through its identification, excerption, and classification. Assuming an onomasiological approach to borrowing, we arrange LGBTQ+-related anglicisms on a decreasing foreignness scale to identify the borrowing techniques adopted by the recipient language speakers in the loan nativization process. We also address issues related to the identification and semantics of loans, and sketch areas of research on loan pragmatic functions that need further studies.
- Research Article
- 10.21248/idsopen.16.2026.85
- Mar 9, 2026
- Online-only Publikationen des Leibniz-Instituts für Deutsche Sprache
The study examines pandemic-related language change using historical lexicons on cholera. It is based on a cholera corpus generated from the historical archives of the DWDS and DeReKo, which was analysed using Sketch Engine and examined for frequencies, compound words and semantic patterns. Historical dictionaries and secondary literature were also consulted to explore concepts such as the feminisation of the disease. The analysis shows that pandemic language change is particularly visible in thematically focused fields such as infrastructure and medical measures. Frequency data show wave-like patterns, with pandemic compound words following productive patterns. Semantic fields are cyclically reactivated across pandemics. Methodologically, the study combines corpus linguistic methods with cultural studies contextualisation and links the results in Linked Open Data structures, enabling comparative analyses between pandemics. The study demonstrates how historical language data can be integrated into modern digital infrastructures, thus opening new perspectives for diachronic analyses of pandemic vocabulary.
- Research Article
- 10.15294/sutasoma.v9i1.46309
- Jul 1, 2021
- Sutasoma : Jurnal Sastra Jawa
This study examines the maintenance of Javanese language by students in the family sphere in Jepangpakis, Kudus Regency and the factors that influence it. The purpose of this research is to describe the form of Javanese language maintenance by students in the family sphere in Kudus and to reveal the factors that influence it. This research uses a qualitative descriptive approach with a case study model. The data source of this study were 10 informants who were selected based on purposive sampling technique. The data of this research is in the form of informants' language data in the form of utterances which is indicated the use of Javanese language either in full or in part. The data is obtained through data collection techniques, namely interviewing, observation, listening, recording, and note-taking techniques. The collected data were analyzed using qualitative descriptive techniques and presented with formal and informal techniques. The results showed that: 1) the form of Javanese language maintenance by students in the family sphere in Kudus can be classified into: a) full Javanese utterances with two types, namely ngoko lugu and ngoko alus varieties; and b) utterances which are the result of the code mixing process of Javanese and Indonesian; 2) factors that influence the maintenance of Javanese language by informants, including: a) Javanese language shows the identity of the speaker; b) positive attitudes towards Javanese by speakers and speech partners; and c) positive attitudes towards Javanese are still found by people around the speakers.
 Keywords: Javanese language, students, language maintenance
- Book Chapter
100
- 10.1163/9789004253216_002
- Jan 1, 2009
The focus of this volume is on semantic and pragmatic change, its causes and mechanisms. The papers gathered here offer both theoretical proposals of more general scope and in-depth studies of language-specific cases of meaning change in particular notional domains. The analyses include data from English, several Romance languages, German, Scandinavian languages, and Oceanic languages. Detailed case-studies covering central semantic domains, such as concession, evidentiality, intensification, modality, negation, scalarity, subjectivity, and temporality, allow the authors to test and refine current models of semantic change, by focusing, for instance, on the respective roles of speakers and hearers in the process and on the relationship between semantic and syntactic reanalysis. Key theoretical notions, such as presuppositions, paradigms, word order, and discourse status are revisited in a diachronic perspective to provide innovative accounts of causes and motivations for linguistic changes. A prominent theme is the evolution of procedural meanings of various kinds. Thus, several papers feature different types of pragmatic markers as their object of study, while others are concerned with items and constructions expressing modality, evidentiality, negation, and relational meanings. Closely related themes are: the interface between semantics and pragmatics/discourse, with figurative uses of language, rhetorical-argumentational strategies, discourse traditions, information structure, and the importance of dialogic contexts in change playing a salient role in several papers; the relationship between meaning change and processes such as grammaticalization, subjectification and pragmaticalization; and, the thorny issue of the categorization of linguistic items such as discourse markers or modal particles, evidentials or epistemic modals, to which the diachronic data are shown to contribute substantially. The volume will be of interest to graduate students and researchers in the fields of semantics, pragmatics, discourse analysis, grammaticalization, and historical linguistics.
- Supplementary Content
4
- 10.17635/lancaster/thesis/831
- Sep 30, 2019
- University of Lancaster
Extracting and analysing meaning-related information from natural language data has attracted the attention of researchers in various fields, such as Natural Language Processing (NLP), corpus linguistics, data sciences, etc. An important aspect of such automatic information extraction and analysis is the semantic annotation of language data using semantic annotation tool (a.k.a semantic tagger). Generally, different semantic annotation tools have been designed to carry out various levels of semantic annotations, for instance, sentiment analysis, word sense disambiguation, content analysis, semantic role labelling, etc. These semantic annotation tools identify or tag partial core semantic information of language data, moreover, they tend to be applicable only for English and other European languages. A semantic annotation tool that can annotate semantic senses of all lexical units (words) is still desirable for the Urdu language based on USAS (the UCREL Semantic Analysis System) semantic taxonomy, in order to provide comprehensive semantic analysis of Urdu language text. This research work report on the development of an Urdu semantic tagging tool and discuss challenging issues which have been faced in this Ph.D. research work. Since standard NLP pipeline tools are not widely available for Urdu, alongside the Urdu semantic tagger a suite of newly developed tools have been created: sentence tokenizer, word tokenizer and part-of-speech tagger. Results for these proposed tools are as follows: word tokenizer reports $F_1$ of 94.01\%, and accuracy of 97.21\%, sentence tokenizer shows F$_1$ of 92.59\%, and accuracy of 93.15\%, whereas, POS tagger shows an accuracy of 95.14\%. The Urdu semantic tagger incorporates semantic resources (lexicon and corpora) as well as semantic field disambiguation methods. In terms of novelty, the NLP pre-processing tools are developed either using rule-based, statistical, or hybrid techniques. Furthermore, all semantic lexicons have been developed using a novel combination of automatic or semi-automatic approaches: mapping, crowdsourcing, statistical machine translation, GIZA++, word embeddings, and named entity. A large multi-target annotated corpus is also constructed using a semi-automatic approach to test accuracy of the Urdu semantic tagger, proposed corpus is also used to train and test supervised multi-target Machine Learning classifiers. The results show that Random k-labEL Disjoint Pruned Sets and Classifier Chain multi-target classifiers outperform all other classifiers on the proposed corpus with a Hamming Loss of 0.06\% and Accuracy of 0.94\%. The best lexical coverage of 88.59\%, 99.63\%, 96.71\% and 89.63\% are obtained on several test corpora. The developed Urdu semantic tagger shows encouraging precision on the proposed test corpus of 79.47\%.
- Research Article
- 10.15388/respectus.2014.25.30.2
- Apr 25, 2014
- Respectus Philologicus
This paper presents a contrastive linguo-cognitive study of the phraseological semantic field, here viewed as a means of manifesting the corresponding fragment of the phraseological picture of the world built in consciousness and defined as a semantic unity of phraseological units which are connected with some phenomenon of thereal or imagined world and reveal the phraseological concept of this phenomenon.The paper focuses upon three main aspects of a contrastive linguo-cognitive study of the same phraseological semantic fields in different languages, highlighting their cognitive value: 1) a semantic inventory of the phraseological semantic field; 2) a phraseological representation of individual phraseological semantic groups as components of the phraseological semantic field and manifestants of separate features of the phraseological concept; and 3) a configuration of the component structure of the phraseological semantic field. The paper presents a method of contrastive linguo-cognitive study which is carried out in the three mentioned directions and allows explication on the basis of language data, as well as comparison and characterization of the contents of the different languages’ concepts and accents available in these contents. The method emphasizes the need to process language data and the importance of employing statistical procedures, both to estimate the relevance of possible interlanguage differences and to assess the possibility of their interpretation in terms of cognitive and cultural specificity. Use of the method is demonstrated through the example of a contrastive linguo-cognitive study of the Russian and English phraseological semantic field of arguing.
- Conference Article
43
- 10.1109/cloudcom.2014.40
- Dec 1, 2014
Social media data consists of feedback, critiques and other comments that are posted online by internet users. Collectively, these comments may reflect sentiments that are sometimes not captured in traditional data collection methods such as administering a survey questionnaire. Thus, social media data offers a rich source of information, which can be adequately analyzed and understood. In this paper, we survey the extant research literature on sentiment analysis and discuss various limitations of the existing analytical methods. A major limitation in the large majority of existing research is the exclusive focus on social media data in the English language. There is a need to plug this research gap by developing effective analytic methods and approaches for sentiment analysis of data in non-English languages. These analyses of non-English language data should be integrated with the analysis of data in English language to better understand sentiments and address people-centric issues, particularly in multilingual societies. In addition, developing a high accuracy method, in which the customization of training datasets is not required, is also a challenge in current sentiment analysis. To address these various limitations and issues in current research, we propose a method that employs a new sentiment analysis scheme. The new scheme enables us to derive dominant valence as well as prominent positive and negative emotions by using an adaptive fuzzy inference method (FIM) with linguistics processors to minimize semantic ambiguity as well as multi-source lexicon integration and development. Our proposed method overcomes the limitations of the existing methods by not only improving the accuracy of the algorithm but also having the capability to perform analysis on non-English languages. Several case studies are included in this paper to illustrate the application and utility of our proposed method.
- Research Article
2
- 10.31940/jasl.v2i1.820
- Jun 11, 2018
- Journal of Applied Studies in Language
The four layers of human language – form, meaning, function, and value – are systematically integrated in order to play the communicative functions of human interaction. It is not an easy job to explore and to explain the nature of human language as the four layers are systematically integrated in complex ways. Thus, the linguistic studies should be held in specific domains and topics by means of appropriate theoretical bases and frameworks. This paper, which is mainly inspired by the grammatical-typological analysis on prefix ba- in Minangkabaunese, particularly discusses how the language features are linguistically analyzed in order to come to logic, valid, reliable findings and conclusion. The discussion presented in this paper aims at proposing logical and reasonable ways of doing linguistic analyses on available data of language. In short, this paper deals with how to begin and to do linguistic analyses toward a group of language data collected. In this paper, the prefix ba- of Minangkabaunese is used as the example of case. The discussion presented in this paper respectively answers two main questions; (i) What should be firstly analyzed dealing with the prefix ba- of Minangkabaunese?; and (ii) How are the linguistic analyses toward the prefix ba- of Minangkabaunese logically continued?
- Research Article
1
- 10.1016/j.sbspro.2015.07.442
- Jul 1, 2015
- Procedia - Social and Behavioral Sciences
A Case Study on Oral Corpus: The Use of Mother Tongue in Class by Brazilian Teachers of Spanish as Foreign Language
- Research Article
204
- 10.1093/applin/amh043
- Jun 1, 2005
- Applied Linguistics
In the past few years researchers have begun to show an interest in humour and language play as it relates to second language learning (SLL). Tarone (2000) has suggested that L2 language play may be facilitative of SLL, in particular by developing sociolinguistic competence, as learners experiment with L2 voices; and by destabilizing the interlanguage (IL) system, thus allowing growth to continue. She recommends research examining the ways in which adult L2 speakers interacting outside the classroom play with language as a way of learning more about this issue. Using case study methodology to document the ways in which L2 verbal humour was negotiated and constructed by three advanced non-native speakers (NNSs) of English as they interacted with native speakers (NSs) of English, this study contributes to this knowledge base by showing patterns of interaction that arise during humorous language play between NSs and NNSs and how these may benefit second language acquisition (SLA). Results suggest that language play can be a marker of proficiency, as more advanced participants used L2 linguistic resources in more creative ways. Language play may also result in deeper processing of lexical items, making them more memorable, thus it may be especially helpful in the acquisition of vocabulary and semantic fields.
- Research Article
19
- 10.3176/tr.2002.4.03
- Jan 1, 2002
- Trames. Journal of the Humanities and Social Sciences
Introduction Emotions can be treated as a natural part of human experience. It is equally natural to constantly experience and to think and talk about this experience. Words and concepts can be treated as the main tools of talking and thinking, respectively. Yet what are the interrelations of ubiquitous experiential units (emotions), units of cognitive processing (concepts) and units of verbal communication (words) is far from obvious. There are figurative and literal expressions in languages for both expressing and describing emotional experience (Kovesces 2000). Though there are differences across languages in the range and scope of specific emotion terms, the very principles of conceptualising have been claimed to be universal (Wierzbicka 1999). Some cognitive linguists have argued that in the vocabulary of a specific domain a folk theory or layperson's model of the domain is built up (Oim 1999). A layperson's model represents the socially relevant common sense of a topic in a given culture, the basic level knowledge that most people share and by which most part of their everyday experience is interpreted. It is not clear, however, whether a layperson's model is mostly influenced by the realm it intermediates (e.g. emotions), the realm it serves (social norms and interactions) or the realm it is carried by (a specific language). The universality vs specificity of emotions, emotion terms and emotion concepts across cultures and languages is a topic of interdisciplinary interest to anthropologists, psychologists and linguists (e.g. Scherer & Wallbott 1994, Russell et al 1995, Hupka et al 1999, Wiercbicka 1999). The field methods originally used in anthropology and psychology have been introduced into linguistics. A tradition of empirical studies based on field methods and reliable data was derived from the cross-cultural study of folk colour terms by Brent Berlin and Paul Kay emphasising the evolutionary universality of vocabularies (Berlin & Kay 1969). Different semantic fields have been studied with similar methodology, e.g. terms of botanical and zoological life-forms (C. H. Brown, 1977, 1979), etc. Also an attempt has been made to demonstrate the universal development of emotion categories in 64 natural languages (Hupka et al 1999). The present study explores the folk model of as it presents itself in the Estonian emotion vocabulary. Two interrelated topics are discussed: the role of emotions, emotion terms and concepts in the layperson's model and the relevant facets of the popular emotion category in Estonian. 2. A case study: emotion vocabulary of Estonians 2.1. Background Estonians are a nation of about 1 million situated on the southern coast of the Gulf of Finland. Although they speak a Finno-Ugric language, relation to Western cultures (especially German) is supposed to be dominant by some researchers (e.g. Ross 2002). As in any other language there are plenty of words in Estonian, referring to and differentiating between the qualitative and quantitative aspects of emotional experience. The boundaries of the natural category emotions itself are yet not clear in Estonian as this category seems to be mixed and blended with another closely related natural category feelings. (2) There is no linguistic nor anthropological analysis of Estonian emotion terms available so far. The earlier attempts to explore the Estonian vocabulary referring to emotional experience (Veski 1996, Allik 1997, Kastik 2000) belong to the field of psychology. The goal of these investigations has been to ascertain not a layperson's emotion vocabulary per se, but the use of the vocabulary for the description of experience. Juri Allik has found out that most of the variation of emotion vocabulary is accounted for by two dimensions: Positive Affect and Negative Affect, which are claimed to be unipolar dimensions, not to be regarded as opposites (Allik 1997, Allik & Realo 1997). …