Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Ο γλωσσικός πόρος san-Corpus περιλαμβάνει σώμα κειμένων γραπτού λόγου της Νέας Ελληνικής, έκτασης περίπου 9 εκατομμυρίων λέξεων. Ο πόρος συγκροτήθηκε στο πλαίσιο διδακτορικής διατριβής, με στόχο τη μελέτη των συγκρίσεων ομοιότητας στη Νέα Ελληνική. Μέγεθος & Πηγές Το corpus περιλαμβάνει τρία ισομεγέθη υποσώματα, ώστε να επιτρέπονται οι συγκρίσεις μεταξύ τους: (α) Δημοσιογραφικός Λόγος: 7.774 άρθρα από τέσσερις διαδικτυακές εφημερίδες (Η ΑΥΓΗ, Η ΚΑΘΗΜΕΡΙΝΗ, ΕΘΝΟΣ, ΤΟ ΒΗΜΑ - έτος 2015). Συνολική Έκταση: 2,9 εκατ. λέξεις. (β) Εκπαιδευτικός Λόγος: 96 σχολικά εγχειρίδια δημοτικού και γυμνασίου. Συνολική Έκταση: 3,4 εκατ. λέξεις (μελετώνται 2,8 εκατ.). (γ) Λογοτεχνικός Λόγος: 28 μυθιστορήματα (βραβεία αναγνωσιμότητας περιόδου 2010-2015). Συνολική Έκταση: 2,5 εκατ. λέξεις. Κατανομή Σχολικών Εγχειριδίων ανά Γνωστικό Αντικείμενο Πλήθος Εγχειριδίων Ελληνική Λογοτεχνία 13 Ελληνική Γλώσσα 12 Ιστορία 9 Φυσική – Χημεία – Βιολογία 9 Μαθηματικά 9 Γεωγραφία – Γεωλογία – Περιβάλλον 8 Θρησκευτικά 7 Αγωγή Αισθητική (Εικαστικά – Μουσική – Θέατρο) 15 Αγωγή Υγείας (Φυσική Αγωγή – Οικιακή Οικονομία) 5 Πληροφορική – Τεχνολογία 5 Αγωγή Κοινωνική – Πολιτική 3 Αγωγή Σταδιοδρομίας (ΣΕΠ) 1 Κατάλογος μυθιστορημάτων: Συγγραφέας, Τίτλος Έτος 1ης έκδοσης Δούκα, Μάρω - Το δίκιο είναι ζόρικο πολύ 2010 Θέμελης, Νίκος - Η συμφωνία των ονείρων 2010 Καρυστιάνη, Ιωάννα - Τα σακιά 2010 Μιχαλοπούλου, Αμάντα - Πώς να κρυφτείς 2010 Ελευθερίου, Μάνος - Πριν απ' το ηλιοβασίλεμα 2011 Ζουργός, Ισίδωρος - Ανεμώλια 2011 Μακριδάκης, Γιάννης - Η άλωση της Κωνσταντίας 2011 Μπουραζοπούλου, Ιωάννα - Η ενοχή της αθωότητας 2011 Πανσέληνος, Αλέξης - Σκοτεινές επιγραφές 2011 Παπαδημητρίου, Χίλντα - Για μια χούφτα βινύλια 2011 Παπαθεοδώρου, Θοδωρής - Οι καιροί της μνήμης 2011 Τριανταφύλλου, Σώτη - Για την αγάπη της γεωμετρίας 2011 Φακίνος, Μιχάλης - Η έρημος έρχεται 2011 Βαμβουνάκη, Μάρω - Κυριακή απόγευμα στη Βιέννη 2012 Διβάνη, Λένα - Εγώ, ο Ζάχος Ζάχαρης 2012 Στεφανάκης, Δημήτρης - Φιλμ νουάρ 2012 Ακρίβος, Κώστας - Αλλάζει πουκάμισο το φίδι 2013 Ζέη, Άλκη - Με μολύβι φάμπερ νούμερο δύο 2013 Κορτώ, Αύγουστος - Το βιβλίο της Κατερίνας 2013 Κωνσταντούρου, Μαρία - Αγεφύρωτες σιωπές 2013 Μαντά, Λένα - Με λένε Ντάτα 2013 Ξανθούλης, Γιάννης - Κωνσταντινούπολη των ασεβών μου φόβων 2013 Ρώσση–Ζαΐρη, Ρένα - Άρωμα βανίλιας 2013 Ανδρουλάκης, Μίμης - Αλλέγκρα 2014 Δημουλίδου, Χρυσηίδα - Το κελάρι της ντροπής 2014 Παπαδοπούλου, Ελισάβετ - Μέρες και νύχτες που δεν ήταν δικές μας 2014 Χατζή, Αθηνά - Η θάλασσα έφυγε 2014 Χωμενίδης, Χρήστος - Νίκη 2014 Τεχνικές προδιαγραφές & Μορφότυπος Για την αναπαράσταση των δεδομένων και των μεταδεδομένων υιοθετήθηκε η πολυεπίπεδη οπτική των XML σχημάτων και τροποποιήθηκε το διεθνές πρότυπο TEI P5, 4.0.0 (Text Encoding Initiative). Δημιουργήθηκε ειδικός χώρος ονομάτων sanCorpus (sanC) με σχήμα τύπου RELAX-NG. Το σώμα κειμένων διατίθεται σε TXT και σε XML σε τρεις εκδοχές: Βάθος 0: απλό κείμενο (TXT). Περιλαμβάνει το main core (κείμενο βάσει του οποίου εξετάζονται οι συγκρίσεις ομοιότητας) και το out of core (κείμενο εκτός εμβέλειας της διατριβής, στο οποίο περιλαμβάνονται κείμενα που πλαισιώνουν το κυρίως κείμενο, π.χ. κείμενα διδασκαλίας, πίνακες περιεχομένων, εξώφυλλα) Βάθος 1 = κείμενα στην απλούστερη δυνατή XML κωδικοποίηση Βάθος 2 = κείμενα με πιο λεπτομερείς XML κωδικοποιήσεις. Αυτή η έκδοση (san-Corpus v1.0, Depth 0: Plain Text) περιλαμβάνει το σώμα κειμένων σε μορφή απλού κειμένου (Βάθος 0) στην αρχική του διάταξη (βλ. Επεξεργασία). Στόχος είναι ο σταδιακός εμπλουτισμός με επισημειωμένα δεδομένα, καθώς και με τις εκδοχές Βάθους 1 και 2. Επεξεργασία (Processing) Η μεθοδολογία συλλογής των δεδομένων, η θεωρητική τεκμηρίωση και το σχήμα επισημείωσης επεξηγούνται στις μελέτες Αφεντουλίδου (2022, 2021, 2013, 2012) και Afentoulidou (2009). Η πρώτη εκδοχή (Βάθος 0) χρησιμοποιήθηκε αποκλειστικά για τη λημματοποίηση που απαιτούσε η Collostruction Analysis (Gries, 2024). Στάδια επεξεργασίας (για τη λημματοποίηση): Τμηματοποίηση σε προτάσεις (sentence segmentation) με τη χρήση της βιβλιοθήκης Stanza (Stanford NLP Group, Qi et al. 2020), η οποία βασίζεται στο μοντέλο Greek Dependency Treebank (GDT) του Ινστιτούτου Επεξεργασίας του Λόγου / ΕΚ «Αθηνά». Τυχαία αναδιάταξη των προτάσεων για την προστασία της ακεραιότητας των πρωτότυπων έργων. Λημματοποίηση με τον ILSP Lemmatizer μέσω της Υποδομής Clarin-EL. Για την ανάλυση συμφράσεων του δείκτη σαν απομονώθηκαν συγκεκριμένοι λεκτικοί τύποι. Αδειοδότηση & Δικαιώματα Το san-Corpus συγκροτήθηκε για τις ανάγκες της διδακτορικής διατριβής και προστατεύεται από το δικαίωμα ειδικής φύσης σύμφωνα με την Οδηγία 96/9/ΕΟΚ και το άρθρο 45Α του Ν. 2121/1993. Η χρήση του περιεχομένου γίνεται αποκλειστικά για ερευνητικούς σκοπούς βάσει των εξαιρέσεων της Οδηγίας 2001/29 και της Οδηγίας (ΕΕ) 2019/790 (Text and Data Mining exceptions / Fair Use). Η πρόσβαση είναι περιορισμένη (Restricted Access) και παρέχεται αποκλειστικά σε μέλη της ακαδημαϊκής κοινότητας για σκοπούς επαλήθευσης των αποτελεσμάτων της διατριβής και περαιτέρω μη εμπορική έρευνα. Η πηγή προέλευσης δικαιούται να ζητήσει οποιαδήποτε τροποποιητική ενέργεια (π.χ. αφαίρεση) επί του πρωτότυπου περιεχομένου. Τέλος, η άδεια CC BY-NC-ND 4.0 ισχύει για την επιμέλεια (curation), τα μεταδεδομένα και τη γλωσσολογική επισημείωση του σώματος κειμένων. Βιβλιογραφικές αναφορές Αφεντουλίδου, Β. (2022). Σώμα ελληνικών κειμένων για τη μελέτη δομών ομοιότητας της Νέας Ελληνικής: σχεδιασμός και υλοποίηση. Στο Πρακτικά του 10ου Συνεδρίου Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας (σσ. 67-94). ΕΚΠΑ. Αφεντουλίδου, B. (2021). Δομές ομοιότητας στη Νέα Ελληνική. Σωματοκειμενικές παρατηρήσεις για τον πολυλειτουργικό δείκτη σαν. Προφορική ανακοίνωση στην 41η Ετήσια Συνάντηση του Τομέα Γλωσσολογίας, 13–15 Μαΐου 2021. ΑΠΘ. Αφεντουλίδου, Β. (2013). Και σου απάντησα κάτι σαν ‘τέλεια, εντάξει’. Δείκτης σαν + ευθύς λόγος;. Προφορική ανακοίνωση στο 7ο Συνέδριο Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας, 16–18 Μαΐου. ΕΚΠΑ. Αφεντουλίδου, Β. (2012). Συγκρίσεις ομοιότητας στα Νέα Ελληνικά: ο δείκτης σαν. Στο Z. Gavriilidou, A. Efthymiou, E. Thomadaki & P. Kambakis-Vougiouklis (Επιμ.), Selected papers of the 10th International Conference on Greek Linguistics (σσ. 696-707). DUTH. Afentoulidou, V. (2009). Sketching the σαν conditional construction in Modern Greek. Submitted essay, 2009 Linguistic Institute, Linguistic Structure and Language Ecologies, Linguistic Society of America and UC Berkeley. Gries, Stefan Th. 2024. Coll.analysis 4.1. A script for R to compute perform collostructional analyses. https://www.stgries.info/teaching/groningen/index.html Institute for Language and Speech Processing - Athena Research Center (2015). ILSP Lemmatizer. Version 1. [Software (Tool/Service)]. CLARIN:EL. http://hdl.handle.net/11500/ATHENA-0000-0000-23EE-D Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 101–108). Online: Association for Computational Linguistics.
This study explores the intercultural dimensions of academic and media communication through a discourse-analytical framework. In the context of globalization and rapid digital transformation, cross-cultural communication has become increasingly complex, necessitating a deeper understanding of linguistic, pragmatic, and sociocultural factors. The research examines how discourse practices differ between academic and media contexts under the influence of cultural norms, values, and communicative conventions. Employing qualitative discourse analysis, the study identifies key strategies such as rhetorical organization, lexical choices, and narrative structures in selected texts. The findings reveal that intercultural differences significantly shape both the production and interpretation of discourse, affecting argumentation patterns, levels of formality, and communicative intentions. The study concludes that enhancing intercultural competence and critical discourse awareness is essential for effective communication in a globalized world.
This chapter explores how artificial intelligence (AI) tools mediate the identity development, emotional labor, and academic adaptation of international students in U.S. higher education. Framed through the lenses of intersectionality, resilience, and self-authorship, the chapter draws on duoethnography to examine how AI is used not merely as a technical aid but as a scaffold for rewriting the self in unfamiliar academic terrain. While AI offers immediate access to academic conventions, its reliance on dominant linguistic norms often flattens cultural expression and obscures opportunities for deeper growth. Through personal narrative, peer reflection, and theoretical analysis, this chapter interrogates what is gained and what is lost when AI supplements or replaces human-centered support systems. It argues that international student engagement with AI reveals a broader story about survival, belonging, and identity negotiation in an increasingly technologized and globalized university landscape.
This preprint presents a systematic, research-oriented practicum that guides the reader through the entire modern NLP pipeline: from tokenisation and vectorisation to fine-tuning of large language models, retrieval-augmented generation, and reinforcement learning from human feedback. A distinctive feature of the work is its consistent attention to low-resource and morphologically rich languages -- original contributions on Tajik and Tatar, including subword tokenisers, word embeddings, lexical databases, and transliteration benchmarks, are woven throughout the twelve sessions, demonstrating how modern NLP can be adapted to data-scarce environments without sacrificing rigour. Each session combines concise theory with detailed implementation plans, formalised evaluation metrics, and transparent assessment criteria. The work is not a conventional textbook: it is designed as a reproducible research artefact where every session requires publishing code, models, and reports in public repositories. All experiments are conducted on a single evolving corpus, and the work advocates open-weight models over commercial APIs, with special attention to the Hugging Face ecosystem. Designed for senior undergraduates, graduate students, and practising developers seeking to implement, compare, and deploy methods from classical ML to state-of-the-art LLM-based systems.
Reviewer assignment is increasingly critical yet challenging in the LLM era, where rapid topic shifts render many pre-2023 benchmarks outdated and where proxy signals poorly reflect true reviewer familiarity. We address this evaluation bottleneck by introducing LR-bench, a high-fidelity, up-to-date benchmark curated from 2024-2025 AI/NLP manuscripts with five-level self-assessed familiarity ratings collected via a large-scale email survey, yielding 1055 expert-annotated paper-reviewer-score annotations. We further propose RATE, a reviewer-centric ranking framework that distills each reviewer's recent publications into compact keyword-based profiles and fine-tunes an embedding model with weak preference supervision constructed from heuristic retrieval signals, enabling matching each manuscript against a reviewer profile directly. Across LR-bench and the CMU gold-standard dataset, our approach consistently achieves state-of-the-art performance, outperforming strong embedding baselines by a clear margin. We release LR-bench at https://huggingface.co/datasets/Gnociew/LR-bench, and a GitHub repository at https://github.com/Gnociew/RATE-Reviewer-Assign.
This article explores the pragmatic problems in translation, with a focus on achieving pragmatic adequacy between the original and the translated text. The article discusses various challenges, including genre characteristics, background knowledge of the intended reader, communicative purpose, and sociolinguistic factors. Pragmatic adequacy is defined as a translation that fully reflects the original, emphasizing the importance of maintaining linguistic norms and rules. The study delves into the role of dialects, modernization, and other factors in ensuring pragmatic adequacy. The article addresses the challenges posed by regional dialects in translation, emphasizing the loss of pragmatic features when elements specific to dialects are not translated.
Abstract Research on how non-natives process and learn binomials ( black and white ) is limited. The present study addresses this gap using online (eye-tracking) and offline (familiarity rating) tasks. Sixty non-native speakers of English (L1 = Arabic) read six stories seeded with 21 novel binomials in three conditions: one exposure, six exposures, and no exposure (i.e., only in post-test) in a counter-balanced design. Each item was also presented in the reversed order ( white and black ). The non-natives read the stories as their eye movements were monitored and answered comprehension questions. In addition to the novel binomials, 12 existing binomials (congruent with Arabic) were included in the passages as a baseline for comparison. After completing the reading task, the participants completed an offline rating task as a measure of declarative knowledge of the binomial configuration (i.e., word order). All items were rated twice, once in the forward direction and once in the reversed direction. Online results showed that non-natives were not sensitive to the configuration of existing binomials, and there was limited evidence of any sensitivity to novel binomials. Offline, non-natives showed sensitivity to the configuration restrictions of existing binomials but not novel ones.
This article examines the role of dialectal lexis as a vital source for enriching the literary language, with a particular focus on the Uzbek linguistic tradition and comparable processes in other major languages. Dialects preserve lexical units that have been lost, marginalized, or never codified in the standard variety, and therefore function as a living archive of semantic, morphological, and cultural resources. Drawing on descriptive, comparative-historical, and sociolinguistic methods, the study analyzes pathways through which dialectal words migrate into the literary norm: literary creativity, lexicographic fixation, terminological need, and media diffusion. The paper argues that a measured, scientifically grounded integration of dialectal lexis expands the expressive capacity of the standard language, strengthens national identity, and supports terminological development in rapidly modernizing domains. The findings are relevant to linguists, lexicographers, translators, educators, and language-policy specialists working on the dynamics of standardization in multilingual societies.
This study examined how pre-listening information influences music appreciation among 107 Japanese junior and senior high school students. Two songs were used: The Italian Sogno, where musical tone aligns with lyrics, and the German Im wunderschönen Monat Mai, where they do not align. Participants were assigned to three groups differing in the amount of prior information: none (“No Information Group”), brief lyric explanations (“Lyrics Explanation Group”), and detailed explanations including lyrics, background, and acoustic features (“Lyrics and Background Explanation Group”). When lyrics and tone were incongruent, the No Information Group’s emotional valence ratings aligned more with the tone than did those with prior information; this effect was absent in the congruent condition. Open-ended responses showed the No Information Group focused on surface features like the languages of lyrics rather than thematic content. These findings highlight the educational value of emphasizing lyric understanding in Japanese music education.
This article presents a linguocultural analysis of the prose works of O‘tkir Hoshimov, focusing on the interaction between language and culture in literary discourse. The study aims to identify and interpret culturally marked linguistic units that reflect the national worldview and value system of the Uzbek people. The research material consists of selected novels and short stories that depict everyday life, social relations, and moral norms. The methodological framework is based on linguoculturology and integrates descriptive, contextual, conceptual, and interpretative methods. The results of the analysis reveal that culturally specific lexical units, phraseological expressions, proverbs, and metaphorical constructions play a central role in representing key cultural concepts such as family relations, respect for elders, social responsibility, patience, and humanity. These linguistic elements function as carriers of collective experience and cultural memory, ensuring the transmission of national values through literary language. The findings confirm that O‘tkir Hoshimov’s prose constitutes a coherent linguocultural system in which language serves not only as a means of artistic expression but also as a tool for preserving cultural identity. The study contributes to the development of linguocultural research in Uzbek literary studies and highlights the relevance of linguocultural analysis for interpreting national literary heritage.
Pedagogical bilingual dictionaries are expected to do more than provide translation equivalents: they must support comprehension, accurate production, and the gradual formation of lexical and grammatical competence in learners who operate between two linguistic systems. Traditional bilingual lexicography often treats equivalence as a stable word-to-word relation and represents meaning through short glosses, while contrastive linguistics repeatedly demonstrates that cross-language correspondences are frequently partial, context-dependent, and shaped by differences in semantic segmentation, collocational norms, pragmatic conventions, and culture-specific conceptualization. This article argues that the most productive way to modernize bilingual pedagogical dictionaries is to convert contrastive linguistic results into explicit semantic design principles: a contrastive sense inventory, an “equivalence gradient” (full/partial/functional/zero equivalence), frame-informed meaning explanations, corpus-based collocational templates, and learner-oriented usage warnings that directly target typical interference and errors. The results indicate that contrastive-semantic modeling reduces ambiguity in polysemy alignment, improves learners’ productive choices, and increases the dictionary’s diagnostic value as a tool for preventing negative transfer.
This study investigates mythologemes core mythic concepts encoded in language in Kazakh and English cultures and examines their role in shaping cultural identity. Drawing on linguoculturology and cultural semantics, the research analyzes such figures as “Zhalmauyz Kempir”, “Azireyil”, “the Banshee”, and “the Grim Reaper”. A comparative qualitative design was employed, incorporating data from the British National Corpus, the national corpus of the Kazakh language, literary texts, and folkloric sources. Each mythologeme was examined in its linguistic and cultural context, with particular attention to patterns of lexicalization and metaphorization, and subsequently contrasted across the two languages. The findings demonstrate that although these figures share universal archetypal features primarily as representations of death or evil their linguistic realizations and semantic nuances are culturally specific. For instance, in Kazakh tradition, “Zhalmauyz Kempir” is portrayed as an active, child-devouring entity, whereas the Celtic “Banshee” functions as a passive harbinger of death. Similarly, “Azireyil” in Kazakh discourse evokes the notion of divine fate, while the English “Grim Reaper” embodies the personification of fearful inevitability. Despite cultural differences, both sets of mythologemes remain productive in contemporary language through idioms, metaphors, and media discourse, functioning as dynamic carriers of collective memory and value systems. The study demonstrates that mythologemes operate as adaptable linguistic-cultural constructs that preserve core symbolic meanings while continuously adjusting to evolving social contexts. This comparative analysis highlights how shared mythic motifs are reinterpreted within distinct cultural-linguistic frameworks, thereby reinforcing cultural identity, ethical norms, and national narratives.
We revisit punctuation-aware tree binarization for constituency parsing and ask whether dependency-induced headedness improves binary parser supervision. Although learned heads substantially outperform rule-based heads in intrinsic head prediction, they do not yield consistent parsing gains after debinarization. In particular, punctuation-conditioned evaluation shows that learned headedness underperforms rule-based binarization in macro-average punctuation-sensitive $F_1$, despite a small overall gain on CTB. Similar instability appears under cross-treebank transfer. These results suggest that \ycc{linguistically grounded} headedness is not necessarily parser-optimal when used as a binarization control signal. The paper presents a negative result: better head prediction does not imply better punctuation-sensitive constituency parsing.
The practice of web form submission has emerged as a prime conduit for attackers, enabling them to infiltrate modern web applications and illegally harvest sensitive user data. Traditional defense mechanisms, such as static security reviews and server-side validation, are proving insufficiently agile for real-time detection of client-side vulnerabilities. This inadequacy arises directly from the rapid evolution of modern interfaces, which involves spontaneous DOM changes, dynamic element generation, and semantic interpretation that varies based on context and culture. This article presents an innovative browser extension framework that leverages a heuristic-based, multi-dimensional analytical engine combined with deep DOM inspection to identify insecure form submissions the moment they occur. The proposed methodology introduces five fundamental innovations: a contextual risk scoring system that models the complex interdependencies among form fields; an adaptive weighting scheme for risk patterns, accommodating diverse cultural and linguistic norms; a predictive vulnerability estimator that anticipates future threats; intelligent DOM mutation filtering designed to significantly optimize runtime performance; and cross linguistic semantic recognition to determine the true purpose of fields globally. Based on theoretical projections, this combined approach promises to enhance vulnerability detection accuracy while simultaneously reducing computational demands by approximately. Critically, all security analysis is executed exclusively on the user's local machine, guaranteeing privacy by ensuring no sensitive data is transmitted externally. A proof of concept application confirms the framework's practical feasibility and high efficacy for client side security assessment and catalyzing the development of flexible, scalable, and privacy respecting browser-based protections.
Abstract: The integration of artificial intelligence (AI) into English as a Foreign Language (EFL) education has brought about transformative changes in how learners develop intercultural communicative competence (ICC). This systematic literature review examines how AI-mediated language production and adaptive feedback mechanisms reshape ICC among EFL learners. Following PRISMA guidelines, this study analysed 35 peer-reviewed articles published between 2020 and 2025. The review focuses on ELT-relevant dimensions, including automated writing evaluation, generative AI in language learning, and AI-mediated cross-cultural exchange. Findings indicate that AI facilitates ICC by providing real-time adaptive feedback that helps learners negotiate cultural nuances and linguistic norms. The study concludes that AI serves as a "cultural mediator," offering a triadic interaction model that enhances learners' knowledge, skills, and attitudes in intercultural settings.
This dataset contains multimodal neuroimaging and physiological data from a study investigating the effects of Targeted Memory Reactivation (TMR) during REM sleep on emotional reactivity. Participants encoded affective images paired with sounds, received auditory cues during subsequent REM sleep, and were rescanned 48 hours later during arousal rating tasks in an fMRI scanner. The dataset includes structural and functional MRI, polysomnographic recordings with EEG during sleep, heart rate measurements, and behavioral ratings across three sessions spanning two weeks.
The emergence of a distinct Gen-Z sociolect, often termed "Genzie" or "Internet Slang," represents one of the most rapid and transformative linguistic developments of the digital age. This language is not a random collection of slang but a complex, rule-governed system born from the intersection of technology, social change, and identity formation. A comprehensive, data-driven analysis of its historical evolution, structural properties, and socio-pragmatic functions is critical to understanding contemporary communication, as it reflects fundamental shifts in how a generation conceptualizes interaction, community, and self-expression. This study aims to deconstruct the Gen-Z sociolect by tracing its historical development over a key 36-month period (2021-2023), analyzing its core structural components (lexical, semantic, syntactic, multimodal), and explaining its social functions within digital communities. The research seeks to move beyond anecdotal description to provide a rigorous, empirical account of this dynamic linguistic phenomenon, thereby establishing a benchmark for the academic study of internet-native dialects. We position this sociolect not as a degradation of Standard English, but as a legitimate linguistic innovation worthy of serious scholarly attention, with its own internal logic and systemic coherence. A mixed-methods, diachronic approach was employed, integrating the scale of computational linguistics with the nuance of qualitative discourse analysis. A large-scale corpus of approximately 6000 posts was compiled from three core platforms—Twitter/X, Instagram, and TikTok—across the 12-month timeframe, ensuring a representative sample of public-facing Gen-Z communication. Computational linguistics methods were used for quantitative analysis, including time-series modeling for lexical diffusion, diachronic word embeddings for semantic shift, and supervised machine learning for stylometric identification. This was complemented by qualitative discourse and pragmatic analysis of a stratified sample of posts to understand language-in-use, focusing on the interplay between text, image, and platform-specific conventions. The analysis reveals a clear, platform-influenced historical trajectory for Gen-Z language, with terms originating on niche, visually-driven forums like TikTok and Twitch before achieving mass diffusion on the text-centric environment of Twitter and finally being normalized on the broader social canvas of Instagram. We identified and modeled three primary mechanisms of lexical creation: neologism (e.g., "skibidi," "gyatt"), semantic reappropriation (e.g., "cap," "based," "fire"), and phono-semantic matching from online cultures (e.g., "ratio," "L + RIP bozo"). Gen-Z language is a legitimate and sophisticated dialect of the digital era, a natural linguistic adaptation to a hyper-connected, attention-economy-driven world. Its evolution is not chaotic but follows predictable patterns of cultural transmission that are dramatically amplified and accelerated by social media algorithms. Its structure efficiently manages cognitive load in fast-paced digital environments while its primary functions are the performance of a specific digital identity, the creation and policing of digital community boundaries, and a form of resistance to traditional linguistic and social norms.
This study explores the pragmatic typology of the functional-semantic field (FSF) of degree in English and Uzbek, focusing on how gradability, intensity, and comparison are expressed and interpreted across two typologically different languages. The concept of degree is treated as a universal semantic category realized through a range of linguistic means, including morphological forms, lexical items, and syntactic constructions. The research aims to identify both common patterns and language-specific features in the expression of degree, as well as to analyze the role of pragmatic factors in shaping its meaning. The findings demonstrate that English primarily relies on analytic and morphological devices, such as comparative and superlative forms and intensifiers, while Uzbek employs agglutinative mechanisms, lexical markers, and expressive forms such as reduplication. Despite these structural differences, both languages share a common semantic core based on scalarity and gradation. However, the interpretation of degree is highly context-dependent and influenced by speaker intention, discourse context, and cultural norms. The study also shows that degree expressions serve not only as markers of quantitative or qualitative comparison but also as pragmatic tools for expressing evaluation, emphasis, politeness, and implicature. The functional-semantic field of degree is organized into core and peripheral zones, where core elements provide basic gradation and peripheral elements introduce stylistic and contextual variation. Uzbek demonstrates a stronger tendency toward expressive and emphatic usage, while English often relies on more implicit and context-driven strategies. In conclusion, the research highlights the importance of integrating semantic and pragmatic approaches in the study of degree and contributes to a deeper understanding of cross-linguistic variation in functional-semantic categories. The results have practical implications for language teaching, translation, and intercultural communication.
ABSTRACT This study investigates the perceptions of Americanisms among three generations of Nigerians. While prior research has provided quantitative evidence for American influence in contemporary Nigerian English, the role of language beliefs and ideologies in mediating such changes remains underexplored. Developing a sociolinguistic perspective of mobile linguistic resources, this study construes an individual's linguistic repertoire as an identity‐construction resource, agentively mobilised across geographical, social and digital spaces. Interview data indicate that younger speakers orient towards multiple linguistic norms, while older speakers remain critical of Americanisms and favour British norms. Reading task results further indicate that American realisations are most frequent among younger speakers. The study demonstrates that multinormativity extends beyond linguistic production to speakers’ evaluative orientations and perceived repertoires. This finding advances the sociolinguistics of mobility and World Englishes research by showing that shifting language ideologies – rather than usage patterns alone – constitute a key mechanism driving linguistic change in postcolonial varieties.
This study conducts a corpus-based comparative analysis of the translation styles of ChatGPT4o and the official human translation of the 2025 Chinese Government Work Report. Drawing on a multi-level stylistic framework, it integrates quantitative and qualitative analysis to examine salient features at the lexical, syntactic, and textual levels. Findings show systematic stylistic divergences. Lexically, the official translation exhibits a significant overuse of “will” and underuse of “can” relative to ChatGPT, and it consistently employs the explicitation stategy to clarify China-specific terms and abbreviations. ChatGPT, by contrast, tends toward literal translation and occasional transliteration, producing a more compressed but less interpretively informative rendering. Syntactically, the official version strongly favors agentive declaratives, whereas ChatGPT relies heavily on imperative structures; the two also differ significantly in passive usage. Textually, the official translation prefers additive progression aligned with recurring institutional frames, while ChatGPT more frequently uses “while + V-ing” to condense inter-clausal relations, altering tone and perceived authority. The study attributes these differences to institutional skopos and norms, cross-linguistic discourse tendencies, and the absence of political-communicative constraints in LLM output, and it outlines implications for prompt design, materials of varied genres and registers, and different LLMs.
This paper presents a small-scale dependency treebank for Tunisian Arabic (TADT) developed within the Universal Dependencies framework, addressing the scarcity of linguistic resources for the Arabic varieties.The approach employs domain adaptation, leveraging a machine learning model (UDPipe 1.0) trained on Algerian Arabic data to annotate 100 Tunisian Arabic social media comments, followed by manual correction.This pilot study evaluates the feasibility of using machine learning-assisted annotation to scale resource development for spoken Arabic and identifies key challenges in cross-dialectal transfer for improving annotation quality and efficiency.This work contributes to more inclusive and fair representation of Arabic linguistic varieties in academic research and NLP applications.
Background: Body image dissatisfaction, disordered eating, and eating disorders represent significant public health concerns; however, many affected individuals never access evidence-based support. We co-designed and developed a rule-based chatbot, JEM, which conducts conversations addressing evidence-based psychoeducation and psychotherapeutic microinterventions. We previously demonstrated the feasibility, acceptability, and preliminary satisfaction of the JEM chatbot in a research setting. However, broader satisfaction, experiences, and user-reported outcomes in real-world settings have not yet been investigated. Objective: This study aims to conduct a real-world evaluation of the JEM chatbot in Australia and Canada, the two countries that have hosted a deployment of the chatbot to date. Specifically, we aim to explore user satisfaction and experiences with the chatbot and within-session differences in user mood and body image satisfaction when completing the chatbot's microinterventions. Methods: Respondents were users of the JEM chatbot aged 13 to 64 years who self-selected to complete a web-based overall evaluation survey (N=230; n=122 in Australia and n=108 in Canada) over a 6-month period. This evaluation survey included user demographic characteristics, satisfaction measures, and the System Usability Scale. Respondents for the within-session pre-post analyses were JEM chatbot users who chose to complete brief web-based surveys immediately before and after completing one of the chatbot's microinterventions during the same 6-month period. Sample sizes varied across microinterventions, ranging from 75 to 276 respondents overall (Australia: n=34-146; Canada: n=39-130). These surveys included validated visual analog scales (VAS) measuring mood (anxiety, depression, happiness, confidence) and body image satisfaction (body size satisfaction, body shape satisfaction, physical attractiveness). Results: Demographic characteristics showed that survey respondents were commonly young adult cisgender women and nonbinary individuals across Australia and Canada. Respondent satisfaction with the chatbot was high in both countries (Australia: mean 76.1, SD 22.7; Canada: mean 78.8, SD 14.3), and the usability of the chatbot was rated as "excellent" in both countries (Australia: mean 86.5, SD 16.9; Canada: mean 89.5, SD 11.6) according to the System Usability Scale. Across completed microintervention surveys, patterns of within-session pre-post ratings were broadly similar in Australia and Canada, with effect sizes generally ranging from very small to large across VAS-measured mood and body image outcomes. Conclusions: The JEM chatbot achieved high satisfaction and usability ratings. Among respondents who completed pre-post surveys, immediate within-session differences in mood and body image ratings were observed following the completion of chatbot microinterventions. The study findings were broadly similar across Australia and Canada. These results provide evidence of user experience and within-session differences following engagement with JEM and support continued evaluation in future studies.
Linguistic negation has been described as a natural foregrounding device, with some researchers noting that foregrounding with negators is not always marked grammatically, but also semantically or through textual effect. However, most studies focus on the use of sentential negation (e.g., not, no ), whereas other expressions of negation, such as affixal negation, remain understudied. Recent research indicates that affixal negation is used by speakers to convey a variety of opposite meanings, often in creative and subtle ways. This study is a systematic investigation of the use of negative affixes in an oppositional discourse, namely discourse on race in the USA. The dataset is a specialized corpus, The Corpus of the Non-Fictional Writings by Ta-Nehisi Coates (COCO) (468,899 words, 1996–2018). Methodologically, the study combines corpus linguistic techniques with co(n)textual discourse analysis. The results show that affixal negation in COCO, particularly with the prefixes un-, non- and anti-, is used by Coates to disrupt patterns of collocations, fixed expressions or phrases, producing either lexically or semantically deviant instances. Thus, affixal negation has a foregrounding effect and a potential to introduce evaluative clashes in a discourse. The findings, examined through the lens of linguistic creativity, indicate that Coates employs affixal negation to perform various functions: from explicit foregrounding (e.g., as attention-seeking devices) to a subtle critique of societal norms and established institutional order (e.g., renaming concepts to offer an alternative perspective on reality).
This study investigates Chinese-English code-mixing in the film Everything Everywhere All at Once (2022), focusing specifically on insertion as theorized by Muysken (2000). It examines how insertional code-mixing in scripted multilingual dialogue functions as a narrative strategy to convey identity, emotion, and cultural hybridity in diasporic contexts. Using a qualitative descriptive method, data were drawn from manually transcribed utterances inMandarin, Cantonese, and English. Analysis applied Muysken’s typology alongside Halliday’s sociopragmatics framework to interpret the pragmatic and social factors underlying the insertions. A total of twenty-nine insertions were identified and categorized into phrasal or clause insertions (48.28%), discourse features (41.38%), and lexical items (10.34%). Phrasal and clause insertions most frequently occurred in emotionally intense scenes, expressing affect, authority, or familial conflict. Discourse features such as interjections served to convey emotion and establish shared cultural understanding, while lexical functioned as cultural markers rooted in tradition. The findings demonstrate that insertional code-mixing is a deliberate narrative tool that enhances character depth, cultural resonance, and cinematic authenticity. This study contributes to a broader understanding of how multilingual media can represent bilingual subjectivity, challenge monolingual norms, and reflect complex sociocultural identities. By linking linguistic analysis with filmic representation, the research highlights the significance of studying multilingual cinema as a site where language, identity, and emotion intersect in diasporic storytelling.
This article examines the characteristics of developing lexical competence among university students while studying the Ukrainian language for professional purposes, specifically the restrictions on the use of slang terms; it characterizes the concept of slang and its status within the system of non-literary vocabulary; identifies the main sources of student slang; classifies common slang units in contemporary Ukrainian youth speech; and analyzes their functions in communication. The didactic essence of the issue at hand in the context of teaching the Ukrainian language for professional purposes is demonstrated: it is emphasized that efforts to reduce the use of slang terms contribute to the formation of the linguistic culture of future specialists and their professional orientation. It has been established that slang not only reflects the specifics of the younger generation’s communication, its cultural orientations, social practices, and the influence of the globalized information space, but also pollutes professional speech. The modern Ukrainian language is undergoing active development, driven by both internal linguistic processes and the influence of social, cultural, and technological factors. One of the most dynamic layers of the lexicon is slang, which is particularly prevalent in the linguistic environment of young students. Youth slang reflects the specific nature of the younger generation’s communication, its cultural references, social practices, and the influence of the globalized information space. The relevance of this study stems from the fact that in the 21st century, youth and student speech is undergoing significant changes under the influence of digital technologies, social media, and intercultural communication. These processes contribute to the active borrowing of foreign-language vocabulary, primarily from English, as well as the formation of new slang units that spread rapidly in the communicative space. In this regard, the issue of slang as a linguistic phenomenon has attracted significant attention from researchers and has been extensively addressed in both domestic and foreign linguistics. The issue of contemporary youth slang in Ukrainian speech is thoroughly examined in the work by V. Zayets, O. Stepanenko, and Y. Stepchuk, where it is characterized as an open and dynamic system that actively responds to changes in society and the digital environment. The researchers propose a thematic classification of slang units and emphasize their expressiveness. At the same time, K. A. Brovko focuses on the mechanisms of online slang formation, highlighting such productive methods as abbreviation, truncation, word formation, and semantic shifts, which are also characteristic of Ukrainian youth speech. Despite linguists’ significant interest in the issue of non-literary vocabulary, the question of the functioning and classification of contemporary student slang in the Ukrainian language requires further study. Of particular relevance is the analysis of the sources of slang vocabulary formation, its structural features, and its role in shaping the linguistic identity of young people. Youth slang is an integral and one of the most dynamic components of modern Ukrainian speech, reflecting the active sociocultural and technological transformations of society. It is shaped by both intralinguistic processes (word formation, semantic reinterpretation, and linguistic play) and external factors, primarily globalization and the English-language information space. In formal business style, its use is undesirable and limited, but possible under certain conditions. The analysis conducted has shown that slang is not a homogeneous phenomenon and can be classified according to social, communicative, functional, and genetic criteria. This allows us to view it as a complex system encompassing professional, age-based, subcultural, and marginal varieties, as well as one that functions in various communicative environments – from everyday communication to the digital space. It has been demonstrated that youth slang performs a number of important functions: expressive, identificatory, speech-economizing, conspiratorial, and creative. It is precisely through these functions that it ensures not only effective communication but also the formation of group identity, serving as a means of self-expression for young people. A characteristic feature of slang is its instability and rapid changeability: some units disappear or transition into general vocabulary, losing their stylistic distinctiveness. This confirms the openness of the language system and its ability to adapt to new conditions. Thus, youth slang should be viewed not as a deviation from the linguistic norm, but as a natural manifestation of language development that reflects the current communication needs and cultural orientations of modern society.
This paper examines the specific features of the speech of hearing-impaired individuals in the process of orientation to social life. The relevance of the topic is determined by the fact that hearing loss influences not only auditory perception but also speech production, language development, communicative behavior, social adaptation, and participation in education and community life. The purpose of the study is to identify the main linguistic, psychopedagogical, and social characteristics of the speech of hearing-impaired individuals and to explain how these features affect their orientation to social life. The research is based on theoretical analysis, comparison, interpretation, and synthesis of pedagogical, linguistic, medical, and inclusive-education sources. The findings show that the speech of hearing-impaired individuals is characterized by a specific combination of phonetic, lexical, grammatical, prosodic, and pragmatic features, the severity of which depends on the degree and type of hearing loss, the age of identification, access to an accessible language, early intervention, family support, and educational conditions. The paper argues that speech should not be evaluated only from the standpoint of deviation from hearing norms. It must be understood within a broader framework of communication, identity, inclusion, and social participation. It is concluded that successful orientation to social life requires early and accessible language input, individualized educational support, speech and language intervention where appropriate, communicatively rich environments, and inclusive conditions that promote self-expression and social participation.
This chapter explores the uneven geographies of access, participation, and belonging in global higher education, focusing on how language shapes the lived experiences of international students and faculty. Drawing on classroom-based narratives from India and Oman, it examines how linguistic norms, institutional expectations, and cultural assumptions determine inclusion and exclusion in academic spaces. While student mobility is often framed as a success of globalization, the chapter argues that it remains embedded in hierarchies of language, identity, and geography. Multilingual learners from rural or non-elite backgrounds, and faculty from “non-native” English-speaking contexts, often face marginalization due to misalignment with dominant academic norms. Using Bourdieu's linguistic capital, postcolonial critiques, and critical internationalization studies, the chapter calls for moving beyond tokenistic diversity. It advocates for multilingual pedagogies, translanguaging, and ethical student mobility to build a more equitable and culturally responsive global education landscape.
This paper examines how national mentality is reflected through epistemic modality by comparing English and Uzbek. Epistemic markers encode a speaker’s assessment of certainty, doubt, and probability, thereby revealing culturally preferred ways of presenting knowledge. Using descriptive and contrastive analysis, the study outlines key epistemic resources in English (modal verbs and stance adverbs such as must, may/might, probably, perhaps) and in Uzbek (modal words such as ehtimol, balki, chamasi, shekilli, as well as grammatical constructions and suffixes including -sa kerak, -dir, -ekan/-kan, and -ibdi). The comparison suggests that English typically expresses epistemic stance through separate lexical items, whereas Uzbek often integrates evidential and epistemic nuances into verbal morphology. These differences align with discourse norms: English favors explicit speaker positioning, while Uzbek commonly employs mitigated, context-sensitive formulations that support politeness and social harmony. The article argues that epistemic modal analysis provides a productive route for linking linguistic form with culturally shaped worldviews.
The rules that determine which assets count as eligible collateral for central bank operations, and at what haircuts, are not just operational details. They have become a first-order determinant of asset prices, liquidity allocation, and financial stability. This review synthesizes the literature on three transmission channels: convenience yield, collateral scarcity, and market liquidity. I trace the intellectual lineage from Singh and Stella (2012) through Williamson (2016) to the empirical studies of Nyborg and Woschitz (2021), Lengwiler and Orphanides (2024), and Fang, Wang, and Wu (2020). My central argument is that this literature, taken as a whole, reveals a fundamental policy trilemma. Central banks must choose among unconditional acceptance of their own government’s debt (which risks fiscal dominance), rating-based eligibility (which risks self-fulfilling sovereign crises), and discretionary policy-driven eligibility (which risks politicization). No design is safe. I also identify four open questions: the nonlinearity of the collateral channel, its interaction with bank portfolio behavior, the systemic risk of cliff effects, and the external validity of evidence from China.
This article delves into the syntagmatics of occasional phraseological derivatives, exploring their formation, structure, and contextual usage in language. It highlights the interplay between linguistic norms, individual creativity, and communicative intentions in shaping these occasional expressions. Drawing from linguistic theories and examples, the study emphasizes the nuanced meanings and functional characteristics of occasional phraseological units compared to their conventional counterparts. The analysis also touches upon the role of context, creativity, and communicative strategies in the construction and interpretation of these linguistic constructs, shedding light on their dynamic nature within discourse.
Media plays a crucial role in shaping public discourse and reinforcing cultural ideologies, particularly regarding gender representation. In Mandailing society, where patriarchal norms are prevalent, media discourse provides valuable insights into the positioning of women in social and political spheres. This study investigates the linguistic representations of women in local online news media in Mandailing and explores how these discourses either uphold or challenge patriarchal values. Utilizing a Corpus-Assisted Discourse Studies (CADS) approach, the research combines quantitative corpus analysis with qualitative Critical Discourse Analysis (CDA). The data consists of 180 news articles published between 2020 and 2025 from four local digital news platforms. AntConc 4.3.1 was used to perform word list, concordance, and collocation analyses to uncover patterns in word frequency and lexical associations. These patterns are interpreted through Van Dijk’s socio-cognitive model to understand how discursive structures reflect underlying social ideologies. The analysis reveals that while the media often depicts women as active agents and supporters of solidarity, it also contains elements that reinforce traditional roles and patriarchal norms. The findings underscore the complexity of gender representation in this local context, suggesting that media portrayals can significantly influence societal perceptions of gender roles, promoting both empowerment and the reinforcement of patriarchal values.
The paper examines newly discovered lexical substitutions in the Life of Saint Theodore the Studite, edited by Nil Sorsky. This Life originally belongs to the ancient translations from Greek made by South Slavic scribes in the territory of Old Rus’. It is shown that, in the Life edited by Nil, among the majority of lexical substitutions, which are freely differentiated into groups, there are examples of lexical improvements that require particular consideration. Such exclusive substitutions include corrections of translator’s errors based solely on the Church Slavonic context of the extract (изложенїе – низложенїе), the introduction of Slavic explanations to Greek words preserved in the text (ѥпитимиꙗ – запрѣщениѥ) in accordance with the traditions of the Athos-Tyrnovo school of book reading, as well as the replacement of lexical and semantic archaisms with the synonyms commonly used in the 15th – 16th century texts (позоръ – позорище, притъча – ѡбразъ). It has been established that a striking feature of the lexical revisions carried out by Nil is the desire to adapt the original text to the perception of a scribe of the late 15th – early 16th centuries. The discovered lexical substitutions enable the conclusion that the Sorsky ascetic, undoubtedly guided by the lexical norms of the new Svyatogorsk translations and taking into account the development of the Middle Russian variant of Church Slavonic, approached the correction of each lexeme individually.
Affective word values have been widely studied across languages, often focusing on isolated words due to the difficulty of assessing emotionality in texts. This study examines whether written emotional content can be reliably captured using a specific software tool (Watson Natural Language Understanding). Thirty-three Spanish undergraduates wrote 150-word autobiographical texts in their L2 (English) before and after a training with emotional vocabulary. Normative valence ratings of content words obtained in the pre- and post-training phases were compared with sentiment scores generated by Watson NLU. Strong positive correlations were found between sentiment and normative valence scores in both phases, with stronger relations at post-training. Regression analyses confirmed that sentiment scores significantly predicted normative valence. Importantly, while normative valence did not differ between phases, sentiment scores increased after training. These results suggest that Watson NLU is a valid and sensitive tool for assessing emotionality in written language and its modulation through training.
Public trust in journalism is waning, yet research on disinformation has focused predominantly on non-institutional sources such as social media or partisan websites. Far less is known about how deception can emerge within mainstream newswriting that outwardly adheres to professional norms. This study addresses that gap through a corpus-assisted discourse analysis of fabricated and verified reporting by Stephen Glass, a former journalist for The New Republic whose fabrications were exposed in 1998. A purpose-built, matched-author corpus of twenty-three articles (twelve fabricated, eleven verified) was examined using frequency profiling, keyness, collocation, and concordance analysis informed by Appraisal Theory. The analysis identifies systematic linguistic contrasts between fabricated and verified journalism: fabricated texts display lower lexical density, heavier use of verbs and pronouns, and a more personalised narrative style. Evaluative items such as most, just, and not occur more frequently and perform rhetorical functions of emphasis, mitigation, and denial. Collocational and attributional evidence shows that fabricated articles embed stance more often in the journalist’s own voice, projecting confidence and sincerity while limiting alternative readings. By holding author, outlet, and register constant, the study isolates linguistic traces of deception from broader stylistic variation. Methodologically, it demonstrates how corpus tools can be integrated with discourse analysis to reveal how deception is enacted through patterned use of grammatical and evaluative resources. The findings contribute to ongoing work on disinformation by showing that credibility in fabricated journalism is linguistically performed rather than merely asserted, with implications for media literacy and computational detection of deceptive news.
Multilingualism is defined as a mode of communication in contemporary world. The multilingualism teaches us the important values to understand the context. This study analyzes dual point of view about the multilingualism: the foreign languages that appear in it, i.e. explicit multilingualism and the universal aspect or hidden languages that are indirectly described, i.e. implicit multilingualism. Thismay comprise linguistic norms, reader and text interaction, among others. The aim of this study is to highlight the impacts of elements of multilingualism used in Amélie Nothomb’s novels. It focuses essentially on the works of the contemporary Francophone writer, notably, Amélie Nothomb. She articulates the enriching elements of multilingualism in French and Japanese languages through herwritings. Her breakthrough works mainly articulate the diversity of multilingualism and also the essential meaning of understanding the different elements or expressions related to the French and Japanese language through the richness of culture from a geographical point of view and also the other elements. These elements are articulated about expressions which show the impact ofmultilingualism in her writings that refer either to French, Japanese, or other languages.
omnes flores is an NLP framework based on Universal Dependencies (UD) that utilizes multilingual Large Language Models (LLMs), and its default model is trained on data from 40 UD languages comprising 40 treebanks.For the EvaLatin 2026 Dependency Parsing Tasks, we extended the training data of omnes flores by incorporating six public Latin treebanks from UD and trained a dependency parsing model using the extended training data.The dependency parser of omnes flores normally takes a list of word FORM values as input.However, since the EvaLatin 2026 test data includes an UPOS column, we investigated whether incorporating both FORM and UPOS during both training and inference could improve parsing accuracy.Our experiments show that training using both FORM and UPOS improves performance by 0.5-1.0LAS points on Prose compared with training using only FORM, but decreases performance by 5 points on Poetry.
Chinese word segmentation is especially fragile in non-standard text, where language learner errors and other character-level divergences disrupt the word boundaries assumed by downstream annotation and evaluation. This paper formulates Chinese word boundary recovery as an alignment-based projection task. Given a noisy source sentence and a cleaner target counterpart, we first align the two strings at the character level and then project target-side word boundaries back onto the source. Beyond the recovery method itself, we introduce two evaluation resources: a manually checked learner Chinese benchmark based on MuCGEC and a controlled synthetic benchmark derived from the Chinese Penn Treebank. Experiments show that direct segmentation remains vulnerable to compound fragmentation in learner input, whereas the proposed two step projection method corrects many over-segmentation errors by using the corrected target to recover source-side word spans. The results show that word boundary recovery is distinct from ordinary segmentation and that alignment projection provides a principled mechanism for stabilizing Chinese annotation and evaluation under noisy input.
This article provides a comprehensive analysis of the sociolinguistic and pragmatic foundations of the concept of social distance. Social distance is interpreted as a system of relations between communicants determined by social status, age, gender, professional position, and cultural norms. The study identifies the mechanisms of expressing social distance at the lexical, grammatical, and pragmatic levels of the language system, as well as reveals their functional characteristics in the communicative process. Based on Uzbek language material, the linguocultural nature of social distance and its close connection with national mentality and norms of speech etiquette are substantiated.
This article analyzes the systemic crisis of Arabic culture from the 13th to 18th centuries and the pivotal role of Arab Christians in initiating the Arab Renaissance. The author explores the causes of stagnation in Muslim society, primarily the institutional dominance of taqlīd (imitation) and the widening gap between sacralized linguistic norms (fuṣḥa) and living speech. Central to the study is the scholarly contribution of the Maronite scholar Ibn Farhat, whose work bridged Western rationalism and Eastern tradition. Special focus is placed on his treatise «Baḥṯ аl-maṭālib wа ḥаṯṯ аl-ṭālib», which simplified Arabic pedagogy and integrated the language into the daily and liturgical practices of Christian communities. The paper emphasizes that the transition from Karshuni script to classical Arabic, alongside Ibn Farhat’s reforms, provided the ideological foundation for overcoming cultural isolation. It concludes that Lebanon’s Christian intellectuals, educated through European models like the Pontifical Maronite College, were the primary catalysts for modernizing Arabic philology and precursors to the 19th-century Enlightenment.
Relevance of the research. This work is primarily based on the personal scientific project of P. J. Piaseckyj, a Ukrainian diaspora scholar from the USA. It was the project titled "Anglo Surzhyk," which was initiated on February 9, 2018 (and still ongoing), that provided the author with the further impetus to write this paper. The project itself is a study of the prevalence of Anglicisms in everyday, academic, cultural, and professional Ukrainian communication. It is based on articles from Ukrainian media published on the Internet and currently contains over 2,500 borrowings. The project is continuously updated and thus not available in print; an electronic copy may be requested directly from the author via email or by visiting its namesake Facebook page. [2] The author began contemplating the Anglicization of the Ukrainian language as early as 1949, at the age of six, upon arriving in New York and hearing the Anglicized Ukrainian spoken by American Ukrainians (from the first and second waves of emigration). Mr. Piaseckyj himself is fully proficient in English and possesses a "keen sensibility toward our language." On the other hand, Oleh Rudnyk (who is also fully proficient in both mentioned languages) focuses this work on the linguistic purism movement, specifically its historical continuity and its relevance to societal needs within the context of contemporary Ukrainian national realities. This focus is grounded in the ideas and works of American researchers such as Edward Sapir, Benjamin Lee Whorf, and Welsh scholar Rhianwen Daniel [23, 24, 25]. This work also serves as an appeal to the Ukrainian academic community to more urgently address the issue of protecting the Ukrainian language from the phenomenon of creolization in the current conditions of global advancement. The purpose of this work is to analyze the historical preconditions and current manifestations of linguistic distortion—caused by Rossification and the influx of mediated Anglicisms—to fully comprehend this influx. We investigate the consequences of these phenomena for the development of the lexical richness and word-formation capacity of the Ukrainian language. Concurrently, we advocate for the necessity of implementing effective measures for its protection. These conclusions are grounded in the empirical analysis of a corpus of words, gathered within the framework of the “Anglo Surzhyk” project [2], which attests to the Rossian-mediated provenance of a considerable portion of these borrowings. Emphasis is placed on the essential role of governmental involvement in the defense and standardization of the Ukrainian language amid globalization and persistent external influence. Conclusions. The cumulative effect of centuries of Rossification, coupled with the percolation of Anglicisms mediated through the Russian language, poses a considerable threat to the evolution of modern Ukrainian. This pervasive process risks the language's creolization and the potential erosion of its distinct identity. Notwithstanding the substantial lexical richness of the Ukrainian vocabulary, there exists an urgent necessity for proactive measures aimed at linguistic protection and norming. The establishment of a specialized state ministry, such as a Ministry for Language Purity—modeled after the French system—is a pivotal step to ensuring the oversight of linguistic standards, the development of specialized terminology, and the preservation of Ukrainian's uniqueness amid global challenges. Furthermore, it is essential to recover and republish dictionaries dating from the 1920’s from archives, universities, libraries, private collections, and even the Security Service of Ukraine (SBU). Ultimately, the defense of the language is not merely a linguistic pursuit but a national priority that underpins cultural and state identity.
• Combined EEG, voice morphing, ERP, mTRF and MVPA to unravel neural mechanisms of ambiguous attitudinal vocal expression processing. • Pinpointed LSN (700–1600 ms) as the core neural signature distinguishing ambiguous from typical attitudinal voices. • Discovered early–late functional coupling, challenging serial models by linking acoustic encoding to late socio-cognitive inference. • Dissociated acoustic-driven (N1/P2) and valence-specific (LSN) effects via covariate-controlled LMM and mTRF analyses. • Extended multi-stage prosody models to attitudinal processing, integrating ambiguity in real-world social communication. Vocal attitudes (e.g., confidence, desire) convey rich acoustic cues that transmit speaker's intentions and beliefs, playing a pivotal role in natural speech communication. Neurocognitive research has largely centered on inferring attitudes from voices with unambiguous, clearly defined meanings (“typical voices”), while the neural mechanisms underlying ambiguous voices remain underexplored—particularly compared to vocal emotions, leaving a critical gap in understanding paralinguistic socio-cognitive processing. Here, we employed voice morphing to blended two typical attitudinal voices of opposing valences, recording participants’ valence ratings and electroencephalographic (EEG) responses. Data analysis combined conventional ERP analysis using linear mixed-effects modeling on single-trial data with multivariate approaches (multivariate temporal response function [mTRF], multivariate pattern analysis [MVPA]). Behaviorally, ambiguous voices elicited longer reaction times and intermediate valence ratings. Neurally, ambiguous voices showed a P2 (274–324 ms) resembling positive voices, an N400-like negativity (400–450 ms) resembling negative voices, and a robust a Late Sustained Negativity (LSN; 700–1600 ms) distinct from typical voices. Controlling for acoustic parameters eliminated early effects (N1/P2/N4), confirming they reflect acoustic processing, while the LSN persisted—indexing neural responses to attitudinal ambiguity. mTRF validated stronger late-stage neural tracking of ambiguous voices after accounting for acoustics; MVPA revealed cross-temporal early-late functional coupling between acoustic encoding and pragmatic inference. Together, these findings demonstrate the brain treats ambiguous attitudinal prosody as a distinct category, engaging a specialized cascade: enhanced early acoustic discrimination, graded valence evaluation, refined semantic processing, and effortful pragmatic inference. This work extends multi-stage models from emotional to attitudinal prosody, challenging strictly serial accounts by highlighting interactive neural dynamics in ambiguity resolution.
This study examines women’s language features and micro-narcissistic self-presentation in digitally mediated religious discourse, focusing on question-and-answer (Q&A) interactions in the online “Pengajian Sabilu Taubah” led by Gus Iqdam. Adopting a qualitative descriptive design with a sociopragmatic orientation, the study aims to explore how female participants use language to negotiate emotion, politeness, authority, and self-visibility in a publicly streamed religious forum. The data were drawn from five publicly accessible livestream recordings and were selected through purposive sampling based on the presence of direct interaction with female participants, extended utterances, and adequate audio-visual quality. Three women participants were analyzed as primary data sources, while two additional participants were used for confirmatory analysis. The primary research instrument was detailed discourse transcription, including lexical, prosodic, and paralinguistic features. Data analysis followed a theory-driven qualitative content analysis guided by Lakoff’s framework of women’s language and Pearson’s functional classification, with data validation ensured through triangulation and confirmatory analysis. The findings show that women’s language in the Q&A sessions is characterized by expressive-affective features, mitigation strategies, and response-oriented utterances that function to elicit recognition and maintain politeness toward religious authority. Furthermore, micro-narcissistic self-presentation is realized through subtle and socially acceptable linguistic practices, such as admiration-seeking expressions and self-referential narratives, rather than overt self-promotion. This study contributes to sociopragmatic and gender-based discourse research by highlighting women’s linguistic agency in digital religious interaction and by conceptualizing micro-narcissism as an interactional phenomenon shaped by religious norms and public visibility.
Conventional research on phoney speech in legal circumstances typically views deception as a moral or cognitive defect at the individual level that can be recognised by consistent linguistic indicators. By suggesting that deception is an institutionally created discourse practice that results from the interplay of cognitive load, procedural limitations, and power imbalances in courtroom communication, this research proposes a theoretical reorientation. The study uses a mixed-methods strategy that combines qualitative forensic analysis with natural language processing techniques applied to specific Indian criminal court rulings, drawing on forensic linguistics, discourse analysis, and computational language modelling. Patterns of strategic ambiguity, evasive coherence, emotional modulation, and pragmatic indeterminacy are found in the testimonies of witnesses and accused individuals. To track how institutional forces influence communication behaviour, computational methods such as sentiment trajectory mapping, stance identification, and lexical dispersion metrics are combined with careful language reading. The analysis shows that deceitful discourse in legal contexts functions more as an adaptive, situationally sensible tactic conditioned by juridical norms and interpretive authority than as a sign of personal dishonesty. This study offers a reusable analytical framework for forensic linguistics and legal discourse studies by modelling deception as a situated, procedural, and culturally mediated phenomenon. This framework has implications for judicial interpretation, evidentiary evaluation, and the moral use of computational tools in legal contexts.
Dictionaries have historically served as instruments of linguistic standardisation and national consolidation, with the Oxford English Dictionary ( OED ) standing as a paradigmatic example of this function. However, Han Shaogong’s A Dictionary of Maqiao subverts this role through its fictionalised lexicon format. Narrated by an Educated Youth sent to the fictional village of Maqiao, the novel compiles an idiosyncratic dictionary shaped by local dialects and cultural practices, opposing the homogenising aims of official lexicography. Han’s playful allusion to the OED – through the invented place-name “Maqiao” (literally, “horse-bridge”) – invites a satirical comparison with Oxford, foregrounding the novel’s challenge to dominant linguistic and lexicographic models epitomised by the OED. While the dictionary format of DMQ has been explored in previous scholarship, this chapter introduces a new dimension by analysing Julia Lovell’s English translation. It argues that Lovell’s version, which adapts the indexing from stroke-based to alphabetical order, inadvertently marginalises the novel’s local particularities and reflects the broader hegemonies of alphabetic, Western-centric linguistic norms. By tracing DMQ ’s friction with both the OED and its English translation, this chapter situates the novel within a transnational discourse of translation, literature and modernity, demonstrating how Han’s work contests the imagined coherence of language and history by foregrounding the improvised, the unfamiliar, and the minor.
This article investigates national-cultural specificity in language and its impact on translation through a comparative analysis of English and Uzbek. National-cultural specificity is examined as a linguistic and cultural phenomenon that encodes a community’s worldview, social norms, and value systems in lexical choices, phraseology, and pragmatic conventionsThe article further explores the challenges these cultural differences pose for translation, including untranslatability, pragmatic mismatch, and semantic gaps, and discusses strategies such as borrowing, cultural substitution, explicitation, and adaptation to preserve meaning.
Crafting effective academic titles is a challenging task that requires balancing informativeness, conciseness, and reader engagement. This paper investigates titles produced by master’s students of Linguistics at the Faculty of Letters and Humanities of Sfax (FLSHS) and compares them with research article titles written by expert authors and with AI-generated alternatives. The study aims to evaluate students’ titles in relation to expert norms and to explore the potential of Large Language Models (LLMs) in academic title generation. A corpus of 659 titles, including master’s dissertations, research articles, and AI-generated titles, was analysed quantitatively and qualitatively using a synthesised model based on Ken Hyland and Zou (2022) and Swales and Feak (2012). The findings show that students adhere more closely to academic title conventions by producing informative and lexically dense titles, whereas expert authors prioritise reader engagement. AI-generated titles, although capable of producing useful lexical content, rely heavily on formulaic expressions and often fail to recognise the genre-specific conventions of academic discourse. They tend to be longer, more clausal, and more question-based than human-crafted titles. The study suggests that AI tools can serve as valuable brainstorming resources in academic writing pedagogy when their output is critically evaluated and adapted by users.
This record contains the data, code, result tables, figures, and supplementary material for the study “How Much Does Corpus Choice Change Dependency-Distance Estimates?”. The archive includes derived analysis data, sentence-level dependency-distance metrics, treebank inclusion and exclusion records, manual provenance/comparability review tables, analysis scripts, configuration files, reproducibility tests, final result tables, figures, model/audit outputs, and the supplementary material submitted with the article. The raw corpora are public, versioned third-party resources from Universal Dependencies 2.18 and Glottolog CLDF 5.3. They are not redistributed in this archive because they are large and governed by their original license terms. Exact source URLs, versions, checksums, release dates, and retrieval information are documented in the included source manifest files. The main reproducibility files are README_ZENODO.md, data/source_manifest.csv, config/analysis.yaml, scripts/, src/, data/processed/, and results/.
Negation is a central phenomenon in linguistics: every language has some way of expressing the difference between an affirmative sentence and a negative one (Horn and Wansing, 2025).However, the treatment of negation remains uneven in Natural Language Processing (Jimenez-Zafra et al., 2017;Jiménez-Zafra et al., 2020).This paper presents the enrichment of a Brazilian Portuguese corpus with negation-related morphological information within the Universal Dependencies (UD) framework (Nivre et al., 2020;de Marneffe et al., 2021).We enrich the Porttinari-base corpus (Duran et al., 2023) by systematically adding the UD morphological features Polarity=Neg and PronType=Neg for 18 negation-related lexical items.The enrichment only modifies the morphological features, leaving tokenization and dependency structure unchanged.To evaluate the computational results of this enrichment, we present an experiment using the Brazilian Portuguese parser PortParser (Lopes and Pardo, 2024), which we trained both on the original Porttinari-base data (Duran et al., 2023) and on our enriched version.Our results show that after enrichment, the parser's performance remains stable, and the newly introduced features are being learned.