Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This article focuses on comparing and contrasting languages, using the points of view of a linguistic criterion and seeks to offer a multi-level and systematic comparison of the linguistic structures. The discussion is done at major scales of linguistic description, which would be phonology, morphology, syntax, semantics, and pragmatics, to demonstrate how languages come together based on universal principles and how they go apart based on their language-specific patterns. Some of the fundamental linguistic universals identified by the study are the existence of lexical categories including nouns and verbs, simple clause structure, universal meaning and grammatical relationship expression mechanisms. Meanwhile, it focuses on idiosyncratic peculiarities which differ among languages, such as sound inventories, morphological typologies, word patterns, and pragmatic norms that are informed by cultural and social circumstances. Moreover, the article also talks about typological categories of languages taking into consideration how languages may be categorized based on structural characteristics such as the coping of analytics or synthetic morphology or the vocabulary order as fixed or flexible.
ABSTRACT Translating local research into English as a lingua franca (ELF) connects local scholarship with global readership, but this process remains constrained by language barriers. Large language models (LLMs) offer advanced accessible solutions, but their responsible integration into academic translation requires a deeper understanding of the linguistic profiles they produce. To address this gap, this corpus‐based study quantitatively evaluates three linguistic dimensions of research article abstracts (RAAs), i.e., lexical complexity, syntactic complexity, and cohesive features, that can influence knowledge dissemination of research articles. Using Chinese‐to‐English RAAs from soft and hard disciplines, we compare translations by two LLMs (LLMTs), GPT‐4o and DeepSeek‐V3, against their corresponding de facto languaging practice by human translators (HTs). Findings reveal that both LLMTs consistently produce higher lexical complexity through varied and sophisticated vocabulary beyond academic stylistic norms but feature lower syntactic complexity with less subordination, reduced phrasal complexity, and weaker cohesive strength. HTs, however, balance lower lexical complexity with higher syntactic complexity and stronger cohesive ties. These findings highlight the communicative affordances and constraints of LLM‐mediated academic translation, offering practical and pedagogical insights for English for Academic Purposes practitioners, academic translators, and non‐anglophone scholars in ELF academic contexts.
The study explores the linguistic representation of gender categories in English and Uzbek through a corpus-based comparative typological approach. Gender has become an important topic in contemporary linguistics, as language not only reflects social structures but also contributes to the construction of social identities. The research examines the semantic and functional behavior of gender-related lexical units in two languages that belong to different linguistic families and cultural traditions. Data for the analysis were obtained from large language corpora, including the Corpus of Contemporary American English (COCA) and the British National Corpus for English, as well as available Uzbek language corpora and digital text collections. The analysis focuses on several key gender-related lexical items such as woman, man, gender, masculinity, femininity, and their Uzbek equivalents. Quantitative and qualitative methods were used to investigate frequency patterns, collocational behavior, and contextual usage. The results reveal both similarities and differences in the linguistic representation of gender categories. While English demonstrates a broader range of discourse contexts related to gender identity and social roles, Uzbek usage patterns appear to be more closely linked to cultural and social norms. The findings contribute to comparative linguistics and gender studies by providing empirical evidence about the ways gender categories are encoded in two typologically different languages. The study also highlights the importance of corpus-based methods in identifying linguistic patterns that might not be immediately visible through traditional qualitative analysis.
This article analyzes the phenomenon of homonymy in the Uzbek language from a linguistic perspective. The study elucidates the essence of homonymy and, with illustrative examples, identifies its manifestations in Uzbek, including lexical homonymy, contextual (speech) homonymy, affixal homonymy, phraseological homonymy, homonymy between fixed expressions and word combinations, as well as homonymy between fixed expressions and sentences. Furthermore, the types of homonymy – namely absolute (proper) homonyms and conditional homonyms, including homographs, homoforms, and homophones – are systematically described. Particular attention is paid to a comparative analysis of the phonetic, semantic, and functional differences between homophones and paronyms. The article also discusses the role of homonymic units in discourse, their impact on the communicative process, and issues related to linguistic norms. The findings of the study are of significant scholarly value for further in-depth research on homonymy in Uzbek linguistics and for practical application in lexicography and language education
This article explores the linguistic representation of notional values and national mentality in Uzbek newspaper advertising texts. Focusing on the historical development of advertising discourse in both English and Uzbek contexts, the study highlights their distinctive features shaped by cultural and social factors. Particular attention is paid to the lexical, semantic, and pragmatic means through which value concepts and mentality are expressed in Uzbek print media advertisements. The findings demonstrate that advertising language serves not only as a persuasive tool but also as a reflection of cultural identity and societal norms. Keywords: Advertising Discourse, National Values, Mentality, Uzbek Newspapers, Linguo-Culture, Pragmatics, Cultural Identity.
Language reflects human cognition, particularly in the expression of mental states such as emotions, beliefs, and intentions. This study investigates how English and Uzbek speakers express mental states in everyday speech, analyzing lexical, grammatical, and metaphorical patterns. English favors the direct labeling of mental states through verbs and adjectives, whereas Uzbek relies heavily on descriptive and metaphorical constructions, often involving bodily imagery. Cultural norms and communicative conventions play a significant role in shaping these differences. The findings contribute to cross-linguistic psycholinguistics, intercultural communication, and applied linguistics, providing insights into how cognition and culture influence language.
This article examines the intrinsic relationship between language and culture, emphasizing the crucial role of linguistic systems in representing, transmitting, and shaping cultural knowledge. Language serves not merely as a communicative tool but as a cultural medium that encodes values, beliefs, social norms, and historical experiences of a community. The study highlights how lexical choices, idiomatic expressions, discourse structures, and pragmatic markers reflect cultural priorities and worldview. It further explores the mechanisms through which culture is preserved and communicated through language, including the use of proverbs, metaphors, culturally specific terms, and pragmatic markers in both spoken and written discourse. Drawing on theoretical frameworks from sociolinguistics, linguistic anthropology, and cultural studies, the article demonstrates that understanding the interdependence of language, culture, and pragmatic markers is essential for interpreting texts, cross-cultural communication, and the study of national identity. The findings underscore that language acts as both a repository and a mirror of cultural heritage, mediating between generations and social groups.
Abstract This study examines the pastoral lexicon of northern Albania as a socially embedded system through which cultural identity, collective memory, and environmental ethics are articulated and sustained. Drawing on an ethnolinguistic corpus derived from Gjovalin Shkurtaj’s lexicographic documentation of the Malësia e Madhe region, the research explores how language mediates relationships among landscape, livelihood, and social organization. Rather than treating pastoral vocabulary as a purely technical register, the study approaches it as a living archive in which lexical items encode moral values, customary law, and patterns of coexistence among humans, livestock, and the mountain environment. Using a qualitative, interpretive methodology, the analysis focuses on lexical and phraseological units related to spatial orientation, mobility, herding practices, ritual temporality, and animal symbolism within their cultural contexts. The findings indicate that pastoral language reflects a collective worldview shaped by seasonal migration, communal governance of resources, and reciprocal human–nature relations. Expressions associated with grazing rights, shelters, livestock groupings, ritual departures, and euphemistic naming practices illustrate how social cohesion and ethical norms are linguistically constructed and transmitted across generations. By integrating linguistic evidence with anthropological and ecological perspectives, the study situates Albanian pastoral culture within broader European and Mediterranean traditions of mobile livelihoods. It argues that the preservation of pastoral vocabulary is not merely a matter of linguistic heritage but a crucial component of cultural continuity, identity formation, and sustainable relationships with the environment.
The integration of large language models such as ChatGPT has raised concerns about stylistic homogenization in scholarly writing. While scientific literature shows clear LLM-driven shifts, e.g., increased lexical markers and reduced cohesion (Bao et al., 2025; Kousha & Thelwall, 2024), this study examines whether similar changes appear in humanities thesis and dissertation titles. Drawing on 8,631 unique MA and PhD titles from ProQuest in History, Religion, Literature, Philosophy, and Musicology, linguistic features were compared between 2015 (pre-AI) and 2025 (post-AI stabilization). Five dimensions were analyzed: word length, informativity, lexical diversity, syntactic structure, and semantic content. Results reveal remarkable stability across most metrics (title length ~12–13 words, informativity ~67%, lexical diversity near 100%). Only a modest increase in compound structures (70% to 74%) occurred, reflecting amplification of existing humanities conventions rather than disruption. The brevity of titles and extended human supervision appear to limit deep LLM intervention. These findings contrast with scientific fields and highlight the resilience of disciplinary norms in graduate scholarship.
This article explores the stylistic functions of neologisms in English language chick lit, a genre characterized by its wit, relational focus, and female-centered narratives. While lexical innovation has been widely studied in science fiction and children’s literature, its role in popular women’s fiction remains underexplored. This study examines how neologisms in chick lit are deliberately formed to reflect character identity, enhance humor, dramatize emotional states, and critique consumerism and gender norms. Through qualitative textual analysis, it is shown that these coinages – ranging from playful blends to metalinguistic jokes – function as stylistic tools with strong social and expressive charge. The findings contribute to the broader study of lexical innovation in fiction by situating chick lit as a genre where language is actively shaped to reflect contemporary cultural dynamics.
This article explores the dynamics of linguistic norms in the Russian literary language within the context of social media communication. The relevance of the study lies in the fact that digital platforms create new communicative conditions that reshape traditional language norms. The paper analyzes variability, simplification, and expressive tendencies in syntax, lexicon, and orthography. Furthermore, it demonstrates that linguistic norms in social media are not static but dynamic and context-dependent. The findings suggest that these processes reflect not the degradation of the literary language, but its adaptation to digital discourse.
While language enables meaning, constituting knowledge in courts, schools, or parliaments, who gets to decide what can be known? Is meaning only use or a result of power too? Pitting Wittgenstein's forms of life against Foucault's regimes of discourse makes linguistic norms appear as instruments of exclusion. Marginalised speakers – subaltern, indigenous, and non-normative are often rendered unintelligible. Epistemic justice demands more than inclusion; it demands considering how rules are set, who enforces them, and how meaning is being contextually built. A discourse-sensitive, epistemic theory of justice is proposed, based on Kripke's rule-following paradox and Dijk's discourse analysis, to show that language is not neutral but a battleground of struggle over meaning, recognition, and epistemic authority.
Abstract This paper examines the social contexts in Thomas Mann’s novel Buddenbrooks where the North German dialect Plattdeutsch is spoken. Beyond the technical challenges of translating these passages, the analysis focuses on the literary representation of code-switching that functions primarily as socially and symbolically charged act. Drawing on Pierre Bourdieu’s sociological theory, the study interprets deviations from linguistic norms and dialect use as instances of double negation – a strategy that appears to challenge social conventions but, in reality, affirms the most valuable social capital in classical bourgeois society: the certainty that one’s high status remains unthreatened. The difficulty of translating such passages stems from the specific cultural parameters embedded in the novel. Ultimately, the paper argues that culture is not an immediate given but requires analytical frameworks from the social sciences for proper understanding and interpretation.
The study is devoted to the analysis of the impact of globalization and digital technologies on the transformation of language norms in modern digital discourse. The relevance of the work is due to the growing role of digital communication, in which language becomes a key factor of cultural identity and social integration. The goal is to identify interdependencies between the level of digital maturity of states, the intensity of language hybridization and the dynamics of sociolinguistic variability. The object of the study is the digital language space of the global communication environment. The methodology is based on a combination of systemic, institutional, econometric and corpus-linguistic approaches using official statistical databases and corpora of digital texts for 2015–2024. The results show that during this period ICT Development Index increased from 5.32 to 7.41 points, DESI Index – from 48.7 to 70.4 points, and the share of Internet users – from 58.2% to 84.3%. At the same time, the number of languages with digital representation increased from 312 to 387, and the share of English-language content decreased from 55.1% to 49.3%. The hybridization index increased from 0.42 to 0.61, clearly indicating the establishment of a polycentric mode or multi-node digital discourse. Hybridization is both method and degree: different languages or their structural elements used in the same message (verbal or symbolic), as well as fully/partially mixed code messages; hashtags expressed in different linguistic forms/fonts, etc. (multimedia message). This means how various linguistic components are arranged within one structural unit up to multimedia messages comprising codes written with different fonts‐forms on various levels of hybridity. The econometric model registered a very high positive correlation between digital maturity and hybridization (r=0.82).
ABSTRACT Objectives This study examined therapist–client physiological synchronisation, indexed by heart rate variability (HRV), in meditation versus emotionally supportive sessions (control condition) and its association with clients' affective evaluations. Methods This was a randomised experimental study. Participants were assigned to a meditation session ( n = 20) or an emotional support session ( n = 20). HRV was assessed at baseline and post‐session; an additional mid‐session HRV assessment was obtained in the meditation condition immediately following the meditation segment. Clients reported positive and negative affect and evaluated the session. Results No significant therapist–client HRV correlations were observed at baseline or at the end of the session in either condition. However, in the meditation condition, changes in therapists' HRV from baseline to mid‐session and from baseline to post‐session were associated with clients' affective evaluations, whereas no such associations emerged in the control condition. Finally, a factor analysis revealed two non‐coherent factors of therapists' and clients' HRV before and after the session in controls, and a single factor including all HRV measures of therapists and clients in the meditation condition. Conclusions Therapist autonomic regulation during meditation may relate to clients' immediate emotional experience. Further research using longitudinal and dynamic designs is needed to clarify mechanisms of physiological synchronisation in psychotherapy.
The article explores the value orientations typical of residents in monotowns, i.e., single-industry towns. The linguistic data were obtained from an indirect associative experiment through an analysis of stimulusresponse speech acts. The methods from corpus linguistics and psycholinguistics made it possible to identify the hierarchical framework of key values in the consciousness of respondents from the diamond mining town of Mirny, Republic of Sakha. The core values included health, family, love, friendship, and security, as well as their semantic associations. The methodology for calculating speech act characteristics included the intersection index (O – Number of Overlapping Associates) and intersection strength (OSG – Overlapping Associate Strength), which reflected the psychological proximity of value meanings in communicative practices. The method revealed the integrative and regulatory role of values based on parameters of speech acts as social actions. The results show how language captures and transmits value orientations, reflecting both universal and ethnocultural perceptions. The findings contribute to the development of interdisciplinary linguistics, enhancing the understanding of value system in single-industry settlements. The method can be used for social planning, academic programs, and cross-cultural studies.
This article examines the “Golden Age” of Arabic linguistics under the Abbasid Caliphate (750–1258). It traces how Arabic evolved from a primarily religious and literary medium into a universal language of science. The study highlights the scholarly rivalry between the Basra and Kufa grammatical schools—especially their debates over qiyās (analogy) and samʿ (attested usage/auditory transmission)—and shows how these methodological differences contributed to the codification of linguistic norms. The article also analyzes the foundational role of Sibawayh’s Al-Kitāb and al-Khalīl ibn Aḥmad’s Kitāb al-ʿAyn in systematizing Arabic syntax, phonetics, and lexicography. Finally, it evaluates the Translation Movement at the House of Wisdom (Bayt al-Ḥikma) and explains how the integration of Greek, Persian, and Indian learning enriched Arabic vocabulary and helped establish a broad scientific terminology.
The study aim was to evaluate differences in ratings of valence made for a set of nonspeech sounds varying in spectrotemporal modulation. For 17 auditory-frequency bands (from 125 to 9474 Hz), modulation index values were extracted at five rates (from 2 to 32 Hz). Higher ratings of valence were associated with lower modulation values in the frequency region below 500 Hz and higher modulation values in the 1620-4000 Hz region. Listeners with hearing loss demonstrated significantly shallower relationships between modulation index and ratings of valence compared to adults with normal hearing. This could partially explain previously demonstrated compressed valence associated with hearing loss.
This paper introduces the first steps towards the creation of a novel resource for contemporary Sardinian within the Universal Dependencies framework. Sardinian is a Romance language spoken in Sardinia, an island belonging to the Italian Republic and located in the center of the western Mediterranean. It is a minority and endangered language, traditionally transmitted mainly orally, and characterized by a multiplicity of varieties (usually grouped into two macro-varieties Logudorese and Campidanese), all recognized as part of the Sardinian linguistic continuum. These varieties share basic morphosyntactic features, while presenting differences at the lexical level and in the realization of specific constructions. This internal variation can be particularly challenging with regard to the normalization of lemmas and the linguistic characterization of certain phenomena. The development of the treebank therefore aims to provide an annotated resource for contemporary Sardinian that takes into account the specificities of the different varieties, using Universal Dependencies to represent them within a unified theoretical framework, in order to facilitate both linguistic analysis and automatic processing. The present paper thus describes some linguistic characteristics of Sardinian and the attempts to encode them within the UD framework. Finally, we present the results of our evaluation of an NLP pipeline for Sardinian, trained on our corpus, for the Stanford Stanza parser.
This study has investigated how teachers at a rural school and an urban school perceive the interplay between dialect, dialect levelling and linguistic norms in teaching. The researcher also observed four Swedish language lessons in grade 4 and 6 at each school. In addition, one lesson in each grade was observed in other subjects, such as mathematics, history, geography and home and consumer studies, resulting in a total of eight observed lessons. To explore the teachers’ perceptions, four semi-structured interviews were conducted with teachers who teach Swedish. The semi-structured interviews, in combination with participant observations, complemented each other well and contributed to strengthening, problematizing, and nuancing the findings. The results show that students’ spoken language varies between a Västerbotten dialect, a more standard variety of Swedish, and a form of language influenced by social media, including slang, abbreviations, informal chat language, and English words and expressions. It emerged that the teachers perceived the influence of social media on language as relatively strong, and that words becoming popular on social media spread quickly and become a natural part of students’ spoken language. The observations also revealed variations in how students’ spoken language was expressed depending on the teaching situation and context. Furthermore, the results show that different forms of dialect are still present in students’ speech through dialectal words and expressions, as well as features such as stress, prosody, and pronunciation. Signs of dialect levelling were also identified, as some students used a more standardized form of language in different situations and contexts.
AgriDTB is a domain-specific discourse dependency treebank constructed from English agricultural academic abstracts. The dataset includes elementary discourse unit segmentation and dependency-based discourse relation annotation, designed for research in discourse analysis and NLP.
During the recent years, the use of linguistic data for language processing increased progressively. Such data are now commonly called language resources. Most of the language resources used for this purpose are collections of texts as the Brown Corpus and the Penn Treebank, but electronic lexicons (WordNet, FrameNet, VerbNet, ComLex, Lexicon-Grammar...) and formal grammars (TAG...) developed recently. Most processes of construction of lexicons and grammars are manual, whereas the construction of corpora has always been highly automated. However, more and more specialists of language processing realize that the information content of lexicons and grammars is richer than that of corpora, and hence the former make more elaborate processing possible. The difference in construction time is likely to be connected with the difference in information content: the handcrafting of lexicons and grammars by linguists would make them more informative than automatically generated data. This situation can evolve into two directions: either specialists of language technology get progressively used to handling manually constructed resources, which are more informative and more complex, or the process of construction of lexicons and grammars is automated and industrialized, which is the mainstream perspective. Both evolutions are already in progress, and a tension exists between them. The relation between linguists and computer scientists depends on the future of these evolutions, since the first implies training and hiring numerous linguists, whereas the other depends essentially on solutions elaborated by computer engineers. The aim of this article is to analyse practical examples of the language resources in question, and to discuss about which of the two trends, handcrafting or generating industrially, or a combination of both, can give the best results or is the most realistic.
Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing image captioning benchmarks typically suffer from limited diversity in caption length, the absence of recent advanced MLLMs, and insufficient human annotations, which potentially introduces bias and limits the ability to comprehensively assess the performance of modern MLLMs. To address these limitations, we present a new large-scale image captioning benchmark, termed, ICBench, which covers 12 content categories and consists of both short and long captions generated by 10 advanced MLLMs on 2K images, resulting in 40K captions in total. We conduct extensive human subjective studies to obtain mean opinion scores (MOSs) across fine-grained evaluation dimensions, where short captions are assessed in terms of fluency, relevance, and conciseness, while long captions are evaluated based on fluency, relevance, and completeness. Furthermore, we propose an automated evaluation metric, \textbf{ITIScore}, based on an image-to-text-to-image framework, which measures caption quality through reconstruction consistency. Experimental results demonstrate strong alignment between our automatic metric and human judgments, as well as robust zero-shot generalization ability on other public captioning datasets. Both the dataset and model will be released upon publication.
Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.
Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.
Negation is a central phenomenon in linguistics: every language has some way of expressing the difference between an affirmative sentence and a negative one (Horn and Wansing, 2025).However, the treatment of negation remains uneven in Natural Language Processing (Jimenez-Zafra et al., 2017;Jiménez-Zafra et al., 2020).This paper presents the enrichment of a Brazilian Portuguese corpus with negation-related morphological information within the Universal Dependencies (UD) framework (Nivre et al., 2020;de Marneffe et al., 2021).We enrich the Porttinari-base corpus (Duran et al., 2023) by systematically adding the UD morphological features Polarity=Neg and PronType=Neg for 18 negation-related lexical items.The enrichment only modifies the morphological features, leaving tokenization and dependency structure unchanged.To evaluate the computational results of this enrichment, we present an experiment using the Brazilian Portuguese parser PortParser (Lopes and Pardo, 2024), which we trained both on the original Porttinari-base data (Duran et al., 2023) and on our enriched version.Our results show that after enrichment, the parser's performance remains stable, and the newly introduced features are being learned.
Czech has been part of Universal Dependencies since its first release in 2015. It has also been one of the best represented languages, with the Prague Dependency Treebank being order of magnitude larger than most other UD treebanks. More recently, three other datasets from the Prague family were added and the annotations thoroughly revisited, forming the "Prague Dependency Treebank-Consolidated" (PDT-C). In comparison to the original PDT, PDT-C is more than twice as large, but it is also much more diverse in terms of genres and domains. In this paper, we describe the conversion of the new resource to Universal Dependencies. While the two annotation schemes are relatively similar at the first sight, there are numerous small differences in topology of the dependency structures and in granularity of the POS and relation type inventories. We demonstrate a selection of such differences on examples, discuss the diverging motivations, as well as ways to overcome the differences during conversion. We argue that while PDT is less "universal" and more tightly bound to one language, its multi-layer annotation is rich and provides all information needed for basic UD trees, and much more.
Abstract This paper presents a novel treebank-driven approach to comparing syntactic structures in speech and writing using dependency-parsed corpora. Adopting a fully inductive, bottom-up method, we define syntactic structures as delexicalized dependency (sub)trees and extract them from spoken and written Universal Dependencies (UD) treebanks in two syntactically distinct languages, English and Slovenian. For each corpus, we analyze the size, diversity, and distribution of syntactic inventories, their overlap across modalities, and the structures most characteristic of speech. Results show that, across both languages, spoken corpora contain fewer and less diverse syntactic structures than their written counterparts, with consistent cross-linguistic preferences for certain structural types across modalities. Strikingly, the overlap between spoken and written syntactic inventories is very limited: most structures attested in speech do not occur in writing, pointing to modality-specific preferences in syntactic organization that reflect the distinct demands of real-time interaction and elaborated writing. This contrast is further supported by a keyness analysis of the most frequent speech-specific structures, which highlights patterns associated with interactivity, context-grounding, and economy of expression. We argue that this scalable, language-independent framework offers a useful general method for systematically studying syntactic variation across corpora, laying the groundwork for more comprehensive data-driven theories of grammar in use.
This paper presents the steps taken to integrate data from the UD_Latin-PROIEL treebank into the LiLa Knowledge Base of interoperable linguistic resources for Latin.It describes how the lexical, morphological, syntactic, and citation information from the source was modeled using the Linked Open Data principles as adopted by the LiLa Knowledge Base.The process of linking tokens to the LiLa collection of Latin lemmas is detailed, addressing challenges such as ambiguities, new lemmas, and errors encountered in the source.The outcome is a syntactically annotated textual resource that is interoperable with the (meta)data of other Latin linguistic resources linked within the LiLa Knowledge Base.This integration enables new ways of analyzing linguistic information and using the content as a starting point to explore connections with other interlinked resources.A use case demonstrates this interoperability.
Thermal imaging, which is contact-free, light-independent, and effective in detecting skin temperature changes that reflect autonomic nervous system activity, is expected to be useful for emotion sensing. A recent thermography study demonstrated a linear relationship between ear temperatures and emotional arousal ratings. However, whether and how ear thermal changes may be nonlinearly related to subjective emotions remains untested. To address this issue, we reanalyzed a dataset that included ear thermal images and self-reported arousal ratings obtained while participants watched emotion-eliciting films. We employed linear regression and two nonlinear machine learning models: a random forest model and a ResNet-50 convolutional neural network. Model evaluation using mean squared error and correlation coefficients between actual arousal ratings and model predictions indicated that both machine learning models outperformed linear regression and that the ResNet-50 model outperformed the random forest model. Interpretation of the ResNet-50 model using Gradient-weighted Class Activation Mapping and Shapley additive explanation methods revealed nonlinear associations between temperature changes in specific ear regions and subjective arousal ratings. These findings imply that ear thermal imaging combined with machine learning, particularly deep learning, holds promise for emotion sensing.
We release a sentence-level genre layer for Universal Dependencies as a separate, joinable dataset, computed across UD revisions and linked back to the underlying treebanks via a release-aware composite key comprising treebank, split, sent_id, and UD release metadata.The annotations are derived rather than authoritative and are accompanied by provenance and uncertainty indicators, enabling downstream users to choose appropriate precision-coverage trade-offs and to re-run the pipeline as UD evolves.To support both parity tracking and deployment-oriented interpretation, we report results under two complementary regimes: a fixed-partition setting aligned with earlier protocols, and a language-grouped 10-fold generalisation setting that highlights cross-language heterogeneity and anchor sparsity as operational constraints.The resulting resource is intended to make genre a practical control variable for UD-based experimentation, including genre-stratified evaluation and training data selection for POS tagging and parsing, where performance varies substantially across text types.Finally, we note that reduced genre spaces aligned with recurring robustness profiles (e.g.transcribed speech versus interactional web/social text versus edited prose/news) appear pragmatically useful, but should be treated as a community coordination task implemented through explicit, versioned mapping tables.
This article evaluates the integration of data extracted from a French syntactic lexicon, the Lexicon-Grammar (Gross, 1994), into a probabilistic parser. We show that by applying clustering methods on verbs of the French Treebank (Abeillé et al., 2003), we obtain accurate performances on French with a parser based on a Probabilistic Context-Free Grammar (Petrov et al., 2006).
Monolingualism, native-speakerism and standard language ideology have been identified as dominant ideologies in language teaching with severe effects on second language teacher identities. Such ideologies offer alleged certainties but also detach teachers from the actual uses and value of language in multilingual and multidialectal contexts. As a consequence, educators might feel constrained by rigid linguistic norms, hindering their capacity to re-evaluate their approaches to accommodate the diverse linguistic realities and communicative needs of English learners in an increasingly interconnected and multilingual global landscape. This chapter intends first to offer a broad perspective of how these ideologies have shaped language teaching and how they have clashed with research-based observations of multilingual and multidialectal communicative settings. We will give an overview of the relevant literature, ranging from foundational texts to more recent ones challenging the ‘ideal’ monolingual native speaker, and we will show how the above ideologies are still found in a rather pervasive way in the language teacher profession. This will be followed by an account of recent research conducted in teacher training environments aimed at showing ways to successfully gear future language teachers towards a new vision of language that contemplates diversity and hybridization as fundamental pillars on which teachers’ identities need to be based.
This study examines the impact of social media on the linguistic behavior of Jordanian Gen Z (born 1997–2012) through the lens of their daily use of colloquial speech as a reflection of sociocultural change. It delineates the dominant linguistic features of the language they use and attempts to address how these linguistic practices reflect the construction of identity and socio-cultural shifts among Jordanian Generation Z. Social media platforms such as TikTok, Instagram, and Snapchat heavily influence Generation Z's vernacular. This study employs a qualitative research approach to analyze pertinent data on code-switching, meme-driven expressions, and abbreviation combinations. Two primary methods of data collection were employed: social media data collection for discourse analysis and semi-structured interviews aimed at identifying the most frequently used expressions among Generation Z. Findings show that the vernacular of Jordanian Gen Z is dynamic, hybrid, and highly integrative in terms of global linguistic resources. This new digital Arabic sociolect poses numerous linguistic and cultural challenges for individuals. These include the necessity for extensive code-switching, the establishment of distinct online linguistic norms, the adaptation to cultural hybridity in language use, and the confrontation of linguistic divergence between generations.
Facial expressions are powerful signals of human emotion, shaping both human–human and human–computer interaction. As interactive technologies, from adaptive interfaces to emotion-aware agents, become more pervasive, systems are increasingly expected to recognize and respond to users’ emotions naturally. But what if a system misreads your face? Such misinterpretation is particularly likely when cultural differences in emotion perception are overlooked. This problem may be compounded by the fact that most facial emotion recognition (FER) models are trained on datasets that reflect the norms of a particular cultural group that assume universality, limiting their reliability in multicultural contexts. Surprise, in particular, is an emotion whose valence can be either positive or negative depending on context, making it a critical case for investigating cultural bias in FER. To address this, we examined how cultural background shapes the recognition and valence interpretation of surprise facial expressions among South Korean (N=36) and American (N=34) participants. Participants labeled 200 facial expressions (surprise and fear), rated their perceived valence, and described personal experiences of surprise. Results show that South Korean-labeled surprise expressions exhibited stronger negative Action Unit (AU) activation and lower valence ratings, whereas American-labeled ones showed more balanced or positive facial cues. Qualitative accounts further revealed that South Koreans framed surprise as tense or socially cautious, while Americans viewed it as open and situationally flexible. These findings bridge recognition and interpretation in cross-cultural emotion research and highlight the need for culturally adaptive FER systems that can interpret ambiguous emotions like surprise more inclusively.