Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Abstract The age-related positivity bias refers to the finding that older adults recount events more positively (or less negatively) as compared to younger adults (i.e., a main effect of age on memory valence). This bias is closely related to the positivity effect, which reflects an interaction between age and valence of information to be remembered. We examined the age-related positivity bias and positivity effect using a one-year longitudinal design with a sample that spanned adulthood (N = 374; age range 19-90; M= 47.41; SD= 16.75). Participants answered questions regarding their memories of learning about the outcome of the 2020 U.S. presidential election. Analyses examined the association between age and valence ratings (positive, negative) and ratings of feelings (happy, elated, upset, and shaken) at Time 1, as well as the association with age between change scores for each of those variables, while controlling for who the participant voted for in the election. Results indicate that increased age was associated with reporting feeling less negative at the time of the event, and also remembering feeling more positive (elated and happy) when reconstructing the event one year later, thereby providing evidence of the positivity bias. There was no evidence of an age by valence interaction in a 2 (Valence) x 3 (Age) mixed ANCOVA on the positive and negative change scores, indicating there was not a positivity effect. Depressive symptoms partially mediated the relationship between age and valence variables, indicating that depressive symptoms may be one mechanism for explaining the age-related positivity bias.
The present study extends recent work on Universal Dependencies annotations for secondlanguage (L2) Korean by introducing a semiautomated framework that identifies morphosyntactic constructions from XPOS sequences and aligns those constructions with corresponding UPOS categories.We also broaden the existing L2-Korean corpus by annotating 2,998 new sentences from argumentative essays.To evaluate the impact of XPOS-UPOS alignments, we fine-tune L2-Korean morphosyntactic analysis models on datasets both with and without these alignments, using two NLP toolkits.Our results indicate that the aligned dataset not only improves consistency across annotation layers but also enhances morphosyntactic tagging and dependency-parsing accuracy, particularly in cases of limited annotated data.
This study explores the syntactic network characteristics of English e-commerce live-streaming discourse by employing a syntactic treebank and syntactic complex network analysis. The main findings are: (1) The syntactic network of English e-commerce live-streaming discourse exhibits small-world and scale-free properties, which are hallmark traits of complex networks. (2) The central nodes of the network are be, I, and the, with be serving as the most central node, while I and the act as local central nodes. (3) The central node be demonstrates both strong centrifugal and centripetal forces. Its centrifugal force is most frequently associated with subject relations and adjective complements, while its centripetal force is characterized by auxiliary and clausal complements. These findings indicate that the syntactic structure of English e-commerce live-streaming discourse is highly robust. This robustness underscores the discourse’s functional purpose: to convey information clearly while engaging users through personalization and specificity. Furthermore, the study highlights the critical role of be in attributive and descriptive constructions. Overall, this research provides insights into the syntactic organization of e-commerce discourse and demonstrates the effectiveness of complex network analysis in linguistic studies.
Treebanks are critical resources in Natural Language Processing (NLP), supporting parser development, linguistic research, and the evaluation of large language models. While Tamil has seen progress in Universal Dependencies (UD) treebanking, existing corpora have been restricted to prose texts, leaving its vast poetic tradition underrepresented. This paper presents the first effort to be made to construct a syntactic treebank for Tamil poetry, specifically focussing on the ThirukkuRaḷ, which is composed in kuRaḷ veṇpā form. A central challenge in this work is posed by the treatment of multiword tokens (MWTs) and elliptical constructions, both of which are observed to occur frequently in Tamil verse due to its agglutinative morphology and metrical constraints. An annotation strategy is proposed within the Enhanced UD (EUD) framework to systematically address five major types of ellipsis—casal, verbal, adjectival, comparative/simile, and cumulative—alongside complex MWT patterns. These annotations not only enhance the representation of Tamil poetic syntax but also broaden the applicability of UD guidelines to underrepresented genres. The contribution is shown to underscore the linguistic and computational importance of capturing the structural specificities of Tamil poetry, while establishing a foundation for future cross-linguistic and literary treebanking efforts.
This study investigated the neurophysiological and affective responses elicited by nature-inspired indoor design elements, including curvilinear forms (CL), nature views (N), and wooden interiors (W), in a virtual environment, and their effects on cognitive performance. Thirty-six participants experienced one control and three experimental conditions in a within-subject design. Electroencephalography (EEG) was used to record neural activity, relaxation and valence ratings assessed affective states, and standardized tasks measured cognitive performance. The W condition elicited EEG patterns indicative of relaxed attentional engagement, including increased alpha-to-theta (ATR) and alpha-to-beta (ABR) ratios, and a decreased theta-to-beta (TBR) ratio. These neural patterns were associated with higher self-reported relaxation and positive affect, and with enhanced cognitive performance relative to the control condition. In contrast, the CL and N conditions did not improve cognitive performance, and the N condition showed elevated physiological arousal, likely due to heightened visual stimulation. Regression analysis identified ATR and relaxation as significant predictors of cognitive performance, emphasizing the role of emotional stability and neural balance in supporting task engagement. Overall, the findings highlight the potential of nature-inspired design to foster a synergy between psychological relaxation and cognitive attention, though further research is needed across diverse spatial typologies to isolate specific design parameters.
The nouns of our language refer to either concrete entities (like a table) or abstract concepts (like justice or love), and cognitive psychology has established that concreteness influences how words are processed. Accordingly, understanding how concreteness is represented in our mind and brain is a central question in psychology, neuroscience, and computational linguistics. While the advent of powerful language models has allowed for quantitative inquiries into the nature of semantic representations, it remains largely underexplored how they represent concreteness. Here, we used behavioral judgments to estimate semantic distances implicitly used by humans, for a set of carefully selected abstract and concrete nouns. Using Representational Similarity Analysis, we find that the implicit representational space of participants and the semantic representations of language models are significantly aligned. We also find that both representational spaces are implicitly aligned to an explicit representation of concreteness, which was obtained from our participants using an additional concreteness rating task. Importantly, using ablation experiments, we demonstrate that the human-to-model alignment is substantially driven by concreteness, but not by other important word characteristics established in psycholinguistics. These results indicate that humans and language models converge on the concreteness dimension, but not on other dimensions.
This essay deals with two colour-related adjectives, badius and baietus, in the Medieval Latin documentary sources of Catalonia studied by the Glossarium Mediae Latinitatis Cataloniae (GMLC). The collection of documental testimonies has been conducted through the lexical database Corpus Documentale Latinum Cataloniae (CODOLCAT), a digital corpus of the Latin texts of this linguistic domain. These sources of notarial and juridical nature contain a considerable amount of colour adjectives, which usually serves to identify and differentiate lexical elements within the same referential class. One of the most attested colour adjectives is badius “bay, brown”, frequently documented with its variant baius and always referring to equines. There are also few occurrences of baietus “brownish”, a lexical innovation derived from badius. To refine their precise definitions, this study explores the forms and uses of both adjectives, taking into consideration the contexts in which they appear. Additionally, this study highlights the importance of the integration of digital tools into lexicographical research and emphasizes the need to incorporate insights from other linguistic domains to achieve a more comprehensive understanding of the words under analysis.
Peer reviewed: True
Cette thèse présente le corpus NaijaSynCor-Prosody, une ressource innovante qui intègre des annotations phonétiques détaillées dans un corpus syntaxique existant. Chaque token du corpus est associée à des annotations décrivant la hauteur, la durée, l’intensité et d’autres attributs prosodiques de chaque syllabe. Ces annotations incluent notamment des contours prosodiques stylisés produits à l’aide du modèle SLAM 3, développé au cours de cette thèse. Cette fusion de données morphosyntaxiques et prosodiques permet des analyses quantitatives de l’interaction entre intonation et syntaxe dans le Naijá, ou pidgin nigérian. À l’aide de ce corpus, nous étudions la différenciation prosodique des unités lexicales qui remplissent également une fonction grammaticale, par exemple des auxiliaires préverbaux marqueurs de TAM. Les contrastes prosodiques entre usages lexicaux et grammaticaux ne sont pas uniformes selon les éléments, bien que notre corpus révèle deux grandes classes prosodiques d’auxiliaires. Un groupe (dey, go, bin) présente une hauteur faible et une courte durée, tandis que les autres se caractérisent globalement par une hauteur plus élevée et une durée plus longue. Deux d’entre eux, make et no, sont particuliers car ils combinent une hauteur élevée et une faible durée. Une analyse plus approfondie de leurs environnements syntaxiques montre qu’ils ont également des distributions atypiques parmi les auxiliaires. Par ailleurs, plusieurs adverbes annotés présentent une distribution syntaxique analogue à celle des auxiliaires et un profil prosodique similaire aux auxiliaires à haute hauteur et longue durée. Ce travail soulève des questions plus générales sur l’utilisation des annotations prosodiques pour remettre en question les étiquettes morphosyntaxiques existantes. La thèse fournit ainsi à la fois une nouvelle ressource linguistique et un cadre méthodologique pour intégrer l’analyse prosodique dans l’étude syntaxe–prosodie à partir de corpus.
This repository contains data accompanying the publication "Auditory localization and subjective assessment of autonomous cleaning robot sounds: A VR experiment on speed, operating mode and alerting signals", submitted for review to the Acta Acustica. The dataset contains: Audio and video material Stimuli consisting of robot recordings under all evaluated conditions (0.3 m/s and 0.8 m/s speed, with and without cleaning, with and without noise AVAS or multi-tone AVAS, both with and without added amplitude modulation). All sounds were exported as 32-bit float wav files; i.e., reading the files into Matlab with audioread results in calibrated Pa values. The files uploaded here were used as source signals in the auralization, assuming a distance of 1 m. The final binaural stimuli were rendered by TASCAR and include an attenuation corresponding to the simulated 7 m distance. 30cms_cleaning_noAVAS.wav 30cms_noCleaning_multiTone.wav 30cms_noCleaning_multiToneAM.wav 30cms_noCleaning_noAVAS.wav 30cms_noCleaning_noise.wav 30cms_noCleaning_noiseAM.wav 80cms_cleaning_noAVAS.wav 80cms_noCleaning_multiTone.wav 80cms_noCleaning_multiToneAM.wav 80cms_noCleaning_noAVAS.wav 80cms_noCleaning_noise.wav 80cms_noCleaning_noiseAM.wav ambienceNoise.wav Excerpt of background noise played back during the experiment. localizationTaskDemo.mp4 Participant POV recording of localization task. This recording was done with a fixed head position, in the actual experiment participants were turning their heads freely. Experiment results and analysis localizationData.csv Table containing the mean and standard deviation of absolute localization error, aggregated for each participant and stimulus. subjectiveData.csv Table containing mean and z-scored annoyance, arousal, trust, and valence ratings for each participant and stimulus. stimuliAnalysis.csv Table containing results of level, loudness, sharpness, roughness, tonality, fluctuation strength, and impulsiveness analysis for all stimuli.
The study examines how children, parents and staff in a kindergarten create a social space in the kindergarten’s cloakroom through (linguistic) actions and language choices. Based on one year of ethnographic fieldwork, which includes participant observation, documentation of the kindergarten’s semiotic landscape, field conversations and interviews with staff, the study shows how the cloakroom becomes a site for multilingual practices, while the kindergarten otherwise is dominated by a Norwegian language norm. The study demonstrates how children and adults take ownership of this space and create a multilingual environment through their actions. The analysis is grounded in Lefebvre’s theory of the production of social space – through spatial practice, representations of space and lived space – Gadamer’s perspectives on play and Pascual-de-Sans’ concept of idiotopy. The findings reveal that language choices and the construction of the cloakroom as a social space are influenced by complex processes related to place, time and agency. When the cloakroom is less in focus for the staff, it can become an important site for children’s play, where they negotiate both linguistic norms and other rules. The study argues for the significance of such social spaces as part of the kindergarten’s linguistic environment, where children can take ownership and make language choices that deviate from established majority language norms. At the same time, the study highlights the importance of time in research on language and place, both theoretically and methodologically, as well as the roles of researchers in gaining access to such spaces through invitations from the children.
This paper examines the phonetic value of the semivowel in Gavrilo Stefanović Venclović’s Služabna knjiga. The work was written between 1711 and 1716 in the Old Church Slavonic language using old Cyrillic script. The script is of the semi-uncial type with elements of cursive writing. The analysis is based on photographs of the manuscript (РГБ, Собрание Н. П. Румянцева Ф. 256 № 401). The text was transcribed using the Transkribus software platform, while most of the excerpted examples were processed using the AntConc software. The main findings of the research can be summarized as follows: (a) semivowel signs and the apostrophe at the end of a word have a purely orthographic function (e.g. даръ, хотѣщимь, готов); (b) as in the vernacular, in the Serbian Slavonic language epoch, the semivowel in a strong position generates the reflex a (e.g. вѣнацъ, диванъ, четвртакъ); (c) the secondary semivowel is vocalized as a (e.g. оганъ, петарь, жизан), or not recorder at all (e.g. жизнъ, огнъ, помыслъ); (d) the semivowel in a weak position, according to the rules of the Serbian Slavonic linguistic norm, is pronounced as a within the prepositions въ (e.g. въ бꙑтїе), съ (e.g. съ нами), and къ (e.g. къ г҃ꙋ), prefixes въ- (e.g. въходиⷮ), въз- (e.g. въⷥдиханїе), and съ- (e.g. съдѣлавъ), and the root въс- (e.g. въсакаа), as evidenced by examples where the letter а appears in place of the former weak position semivowel (e.g. васака, множаствѣ, саблюденїе)
Researchers often assess processes underlying human perception by measuring participants’ judgements of image stimuli. However, traditional methods for quantifying subjective judgements, such as Likert scales, sliding scales, and pairwise comparisons, are vulnerable to biases or demand extensive time and resources from researchers and participants. The present study compared the efficiency, reliability, and validity of these established methods against the Fast Image Rating Experiment (FIRE), our force-choice-based paradigm for assessing perceptions of visual stimuli. When used to rate image preference and naturalness, the FIRE was five times faster than established methods, highly reliable, and valid. FIRE achieved high reliability in less than half the time required to reach equivalent reliability with the Likert or sliding scale, which could save researchers thousands of dollars. The scalability and cost-effectiveness of the FIRE make it a valuable resource for supporting large-scale behavioral science.
Cette thèse vise à présenter une méthode complète permettant d'extraire des règles grammaticales quantitatives et interprétables à partir de treebanks syntaxiques. Cet objectif s'inscrit dans le cadre des grammaires descriptives, qui nécessitent des analyses fines pour saisir des phénomènes linguistiques complexes tout en intégrant les propriétés fondamentales du langage. Pour y parvenir, nous proposons une formalisation générale des règles grammaticales issues de corpus, facile à généraliser et à mettre en œuvre à l'aide de méthodes automatisées. Les règles extraites sont concises, ont différents niveaux de granularité, permettent une sélection flexible et sont ordonnées par importance. Ces descriptions formelles sont extraites de corpus à l'aide de modèles de régression logistique parcimonieux, qui favorisent l'interprétabilité tout en produisant un ensemble concis de règles présentant les propriétés souhaitées. Les résultats de notre méthode sont évalués dans plusieurs langues afin d'examiner sa capacité à répondre aux besoins descriptifs et comparés à d'autres approches existantes. La méthodologie est ensuite étendue à la description contrastive des langues, mettant en évidence les différences et les similitudes entre les langues. Cette approche permet d'obtenir des signatures linguistiques, des ensembles de modèles communs et distinctifs qui profilent chaque langue. Les expériences portent sur plusieurs paires de langues, familles de langues et genres textuels. Une attention particulière est accordée à la nature des règles extraites et au rôle des connaissances linguistiques théoriques dans le processus d'extraction. Cette thèse vise en fin de compte à démontrer la faisabilité de l'extraction d'une grammaire guidée par le corpus pour la description des langues. Ainsi, elle étend les liens entre la linguistique descriptive, formelle et computationnelle dans le contexte plus large de la description des langues.
The article discusses the design and application of dependency treebanks for Biblical Hebrew, focusing on their potential for linguistic research. It highlights the importance of such treebanks in studying historical languages, where native speaker intuition is unavailable. The study compares dependency grammar frameworks, such as Universal Dependencies and Prague Dependencies, examining their suitability for different research goals, including syntax-semantics interface, word order analysis, and phonological-syntactic relationships. Specific criticisms of Universal Dependencies, particularly its hybrid nature prioritising semantic over surface-syntactic relations, are addressed alongside alternatives like the multilayered Prague Dependencies. The article emphasises the need for research-driven design, recommending adaptations based on the intended linguistic applications and underlying theoretical assumptions.
Despite the growing use of NLP in second language (L2) research, model accuracy in L2 settings remains underexplored. This study addresses this gap by evaluating and fine-tuning a Korean language model to extract morphosyntactic features (i.e., morpheme tokenization/tagging and dependency parsing) from L2-Korean texts. We begin by evaluating a domain-general Korean language model on a gold-annotated L2-Korean treebank. We then fine-tune the model on L2-Korean data and quantify the resulting gains across diverse L1- and L2- datasets. Finally, we examine how model reliability varies with learner proficiency scores. Three key findings emerge: while the domain-general model excels at morpheme tokenization, it underperforms on morpheme tagging and dependency parsing; fine-tuning substantially improves adaptability to L2 morphosyntax; and proficiency has minimal effect on morpheme-level tasks but significantly affects dependency-parsing reliability. These results highlight the importance of incorporating L2 training data to improve morphosyntactic analysis in L2 settings and caution against uncritical reliance on automated dependency annotations, especially when performance varies across proficiency levels.
In recent years, syntactic and semantic analysis tools have become increasingly important in various subfields of Natural Language Processing (NLP). These tools enable automatic parsing of large-scale sentences in language corpora, allowing researchers to uncover syntactic structures and statistical regularities of a given language. This study focuses on the development and evaluation of syntactic parsing models for the Uzbek language, employing two widely used approaches: constituency parsing and dependency parsing. For constituency parsing, a rule-based system was developed to identify noun and verb phrases along with their internal constituents. For dependency parsing, a set of hand-crafted linguistic rules was created and applied to syntactically analyze simple Uzbek sentences. As a result of this work, a dependency-based syntactic treebank for Uzbek-Named UzTreebank was constructed. The treebank includes 20,000 automatically parsed simple sentences, of which 10,000 were manually annotated. Additionally, 36 syntactic templates of simple sentences were identified, and 50 linguistic rules were formalized and integrated into the system. The suboptimal performance of the system at its current stage is primarily attributed to the absence of hybrid modeling approaches and the limited size of the training corpus. The paper presents an overview of the rule-based architecture, parsing results, and the current stage of syntactic resource development for the Uzbek language.
The article explores orthographic interference in the process of improving the normative tools of national writing. Orthographic interference is a linguistic phenomenon that arises from the interaction of different language systems, manifesting as deviations or violations of writing norms. The study analyzes changes and adaptations of linguistic norms and their impact on the writing skills of language learners. It focuses on identifying linguistic mechanisms to improve the national writing system by preventing and reducing orthographic interference. The study employs content analysis, comparative methods, and qualitative analysis. Over 60 students’ written works were examined to determine the frequency, typology, and causes of orthographic errors. The analysis identified common types of interlingual and intralingual interference. Interlingual interference results from the influence of Russian and English graphic, phonological, and morphological features, while intralingual interference arises from inadequate understanding of the phonetic and phonological foundations of the language. Content analysis identified the most frequent orthographic errors, with primary causes including insufficient mastery of spelling rules, differences between native and target language writing systems, and teaching method shortcomings. These errors serve as indicators of students’writing experience and language proficiency. The identified types of interference contribute to improving normative tools, assessing the effectiveness of educational programs and methodologies, standardizing orthographic norms, and enhancing writing culture in multilingual societies. The findings support the codification of Kazakh orthographic norms and the development of a scientific and methodological foundation for reducing linguistic interference.
The present paper describes the building of STAF, a Universal Dependencies treebank for Albanian. STAF was bootstrapped using a Stanza model trained on previously unreleased data and then manually corrected by three Albanian speakers supervised by the author, who also revised all sentences. STAF focuses on the fiction genre, featuring 200 sentences selected from nine literary texts written by Albanian contemporary authors.
Tabiiy tilni qayta ishlashda (NLP) parsing, treebank usullari gaplarning sintaktik tahlili, gaplardagi birliklarning sintaktik aloqalari va gap tuzilmasini belgilash bilan shug‘ullanadi. O‘zbek tili morfoanalizatorida 35 000 ta turli uzunlikdagi gaplar POS teglangan bo‘lib, keyingi bosqich esa o‘zbek tilida sintaktik teglash tizimini ishlab chiqish, teglarni taklif etish hisoblanadi. Maqolada o‘zbek tilida sintaktik parsing masalasi, gap bo‘laklarini aniqlash modellari va o‘zbek tili matnlarini dependency parsing qilish muammolari, sintaktik teglar masalasi tadqiq qilindi. O‘zbek tilida sintaktik parsing va treebank masalasi boshqa agglyutinativ tillar bilan solishtirilgan holda tahlil qilindi.
International audience
This paper presents a real-time American Sign Language (ASL) recognition system utilizing a hybrid deep learning architecture combining 3D Convolutional Neural Networks (3D CNN) with Long Short-Term Memory (LSTM) networks. The system processes webcam video streams to recognize word-level ASL signs, addressing communication barriers for over 70 million deaf and hard-of-hearing individuals worldwide. Our architecture leverages 3D convolutions to capture spatial-temporal features from video frames, followed by LSTM layers that model sequential dependencies inherent in sign language gestures. Trained on the WLASL dataset (2,000 common words), ASL-LEX lexical database (~2,700 signs), and a curated set of 100 expert-annotated ASL signs, the system achieves F1-scores ranging from 0.71 to 0.99 across sign classes. The model is deployed on AWS infrastructure with edge deployment capability on OAK-D cameras for real-time inference. We discuss the architecture design, training methodology, evaluation metrics, and deployment considerations for practical accessibility applications.
Michael Riffaterre’s semiotic theory, with its emphasis on ungrammaticality and multi-layered reading, has provided an influential model for analyzing the distinctive structure of poetic language. This study investigates the application of Riffaterre’s semiotic framework to Qur’anic interpretation, with a specific focus on Angelika Neuwirth’s intertextual readings of the Qur’an. Neuwirth views the Qur’an as a poetic and dialogical text that engages with earlier religious and cultural traditions. She employs Riffaterre’s model to reveal the text's semantic depth and internal coherence through the notions of ungrammaticality and dual signification. Using Surah al-Ikhlāṣ as a case study, this paper critically evaluates Neuwirth’s application of Riffaterre’s theory by examining her treatment of the supposed “ungrammaticality” regarding the use of the word aḥad. The paper argues that the notion of ungrammaticality in the Qur’an can be reinterpreted not as a violation of linguistic norms, but rather as a semiotic cue that signals deeper intertextual and theological meanings. Accordingly, evaluating alleged irregularities requires a contextual analysis of lexical patterns across the entire Qur’an, where usage, frequency, and semantic range reveal a consistent theological logic. By integrating insights from classical Arabic grammar, lexicography, and tafsīr with modern semiotic theory, this study reassesses the scope and limits of applying Riffaterre’s model to sacred text analysis. It concludes that while semiotic and intertextual approaches can illuminate the Qur’an’s structural and semantic complexity, they must operate within a balanced hermeneutical framework that respects the text's revelatory nature and linguistic precision. Furthermore, this study demonstrates that the intentional utilization of ungrammaticality in the Qur’an effectively serves specific theological functions.
This paper analyzes the impact of economic globalization on cultural and linguistic norms in contemporary Serbian society. Special attention is devoted to the role of the English language and Anglo-Saxon culture in shaping linguistic practices and identities. The methodological framework is based on a qualitative analysis of relevant literature in the fields of linguistics, cultural studies, and globalization, with particular reference to the works of key authors (Prćić, 2005; Bugarski, 2001; Filipović, 1990). The findings point to the intensive influx of Anglicisms into Serbian and the emergence of a hybrid idiom known as "Anglo-Serbian." The discussion addresses the consequences of global cultural influences for the preservation of local identity, while the conclusion synthesizes the results on the interaction between economic and sociocultural processes.
Abstract Morphological analysis is a foundational task in natural language processing (NLP) and is particularly challenging for low-resourced and morphologically rich languages such as Kangri. Despite substantial numbers of speakers, Kangri lacks annotated corpora, computational tools, and lexicons, making linguistic analysis and downstream processing difficult. This paper presents a hybrid morphological analyzer for the Kangri language that integrates rule-based suffix analysis, lexicon extraction, and efficient machine learning models. A lexicon and suffix transformation rules were automatically induced from the Universal Dependencies (UD) Kangri Treebank. The rule-based morphological analyzer achieved an accuracy of 59\% on the UD test set. A machine learning baseline using TF--IDF character n-grams with Logistic Regression achieved 64.61\% accuracy, while an enhanced model incorporating POS tags improved performance to 67.40%. The results demonstrate that combining linguistic heuristics with statistical learning substantially improves lemma prediction and morphological interpretation for Kangri. This work establishes an initial computational morphology framework for Kangri and provides a foundation for further NLP tool development.
Mood, an individual’s emotional state, fundamentally shapes how the brain interprets sensory input by providing a continuous affective context for prediction and evaluation. In language processing, mood may bias the interpretation of emotionally valenced words, amplifying or dampening their perceived affect. Yet, the temporal dynamics of these mood-valence interactions remain poorly understood. To clarify inconsistent evidence on the timing and nature of mood-valence interactions, we examined how induced mood influences early stages of emotional word processing using EEG. Participants performed a valence-rating task for positive, negative, and neutral words in a baseline condition and following positive or negative mood induction. Event-related potentials were analysed across early processing windows (N1, P2, EPN) using cluster-based permutation statistics. Positive mood selectively attenuated N1 amplitudes for highly valenced words, consistent with reduced prediction error under mood-congruent expectations. Later components (P2, EPN) showed decreased amplitudes for both high and neutral valence, suggesting reduced model updating under mood-congruent expectations. Negative mood, in contrast, produced weaker and temporally delayed modulations. Behaviourally, participants responded more quickly to valenced words under induced mood conditions, supporting the neural findings. Interpreted within a predictive coding framework, these results support the theoretical view that mood functions as a hyperprior, tuning the precision of predictive models during language comprehension. Positive mood appears to enhance predictive flexibility and facilitate the processing of affectively congruent words, whereas induced negative mood reduces positive affect. Taken together, the findings highlight how affective states dynamically modulate early predictive mechanisms in emotional language processing.
This manuscript proposes the S M Nazmuz Sakib Dependency-Focus Principle for Bengali sentence structure and introduces a derived scalar quantity, the Sakib constant, defined over dependency treebanks. Informally, the principle states that in attested Bengali usage, core arguments (subjects and objects) cluster closer to the verbal head than peripheral modifiers (adverbials and clausal adjuncts), and that the ratio between these average distances is numerically stable across corpora. Using real statistics from the UD Bengali-BRU treebank and the Bengali section of the Bengali-Magahi PUD treebank, we define the Sakib constant K Sakib as the ratio between average dependency lengths of core versus peripheral relations, and compute its value for UD Bengali-BRU. Ten figures based on genuine counts and averages illustrate tense and case distributions, relation frequencies, and core versus non-core dependency lengths for Bengali and, for comparison, Magahi. The proposal is presented as a precise hypothesis, mathematically well-defined and empirically grounded in existing treebank data, but still requiring broader testing for confirmation and cross-linguistic generalisation.
Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages; however, they often include significantly less data for low-resource languages such as Urdu compared to high-resource languages like English. To assess the linguistic knowledge of LLMs in Urdu, we present the Urdu Benchmark of Linguistic Minimal Pairs (UrBLiMP) i.e. pairs of minimally different sentences that contrast in grammatical acceptability. UrBLiMP comprises 5,696 minimal pairs targeting ten core syntactic phenomena, carefully curated using the Urdu Treebank and diverse Urdu text corpora. A human evaluation of UrBLiMP annotations yielded a 96.10% inter-annotator agreement, confirming the reliability of the dataset. We evaluate twenty multilingual LLMs on UrBLiMP, revealing significant variation in performance across linguistic phenomena. While LLaMA-3-70B achieves the highest average accuracy (94.73%), its performance is statistically comparable to other top models such as Gemma-3-27B-PT. These findings highlight both the potential and the limitations of current multilingual LLMs in capturing fine-grained syntactic knowledge in low-resource languages.
The article presents the experience of conducting an integrated Ukrainian language class in the format of a language and song talk show titled ‘In the Ukrainian Rhythm’, implemented with first-year students of the Philological Faculty of Educational Technologies at Kyiv National Linguistic University, who are studying within the educational programme ‘The Ukrainian Language and Literature, English, Foreign Literature’. The main goal of the event was to combine the study of the language norms of the modern Ukrainian literary language with elements of the national musical code, specifically song folklore and the modern Ukrainian musical space. The talk show involved an interactive communication format: students not only answered questions but also worked in teams and played games with the audience, showcasing their level of language awareness and their ability to recognise lexical, grammatical, and phonetic phenomena, find language errors, and decode hidden meanings. The programme of the event was structured as a series of competitions (“Name in the Shadow of Periphrases”, “Song, Who is Your Father?”, “LingvoHIT”, “Language Error Catchers”, “Muzrebus”, “Cipherers”, etc.), each of which had a clear linguistic focus, a logical structure, and a focus on interdisciplinary connections encompassing linguistics, literature, media literacy, and cultural studies. Particular attention was paid to fostering participants’ sensitivity to the verbal part of the lyrics. By analysing examples of songs by both classical and contemporary Ukrainian performers, students identified linguistic phenomena (aphorism, paronomasia, homonymy, metaphor, etc.), explored violated norms, expanded their vocabulary, and trained their interpretive skills. Reflection was an important component of the class. In the form of posts on Facebook, students expressed their impressions and emotions, assessed the novelty and pedagogical appropriateness of the lesson, and emphasised its unifying, inspirational, and educational functions. It is concluded that the talk show has a high potential as a pedagogical tool that promotes deeper learning of the norms of the modern Ukrainian language, activates communicative interaction, forms linguistic flair, and, at the same time, develops emotional intelligence, creativity, critical thinking, and national identity in student youth.
We examine how the elements that introduce relative clauses, namely relative complementizers and relative pronouns, evolve over the history of Icelandic using the phrase structure analysis of the IcePaHC treebank.The rate of these elements changes over time and, in the case of relative pronouns, is subject to effects of genre and the type of gap in the relative clause in question.Our paper is a digital humanities study of historical linguistics which would not be possible without a parsed corpus that spans all centuries involved in the change.We relate our findings to studies on the Constant Rate Effect by analyzing these effects in detail.
This article presents the methodology and outcomes of a multi-stage research project aimed at building the first comprehensive lexical database of the Azerbaijani language, developed within the framework of corpus linguistics and statistical lexicography.The database functions as a foundational linguistic module within the national corpus, providing structured lexical inventories for both academic research and technological applications.During the research, a corpus of 520 million word-forms from various functional styles was collected, cleaned, and structured using specialized software.As a result of automatic processing, 2, 918, 910 word-forms were initially identified; through several subsequent stages of technical and lexical filtering, the database was refined.Ultimately, 175, 521 lexical units were compiled into structured word lists, sorted by frequency and in alphabetical order.Additionally, 93, 287 words not found in the orthographic dictionary of Azerbaijani language were identified and presented in a separate list.The article elaborates on the structure of the national corpus, its components (fiction, publisistic, scientific, official, spoken, and educational texts), concordances, and linguistic analyzers.The lexical database is presented as a strategic resource for both theoretical research and practical applications such as natural language processing, automatic translation, text-to-speech systems, and artificial intelligence models.The study demonstrates significant scientific and practical outcomes in the development of digital lexical resources for the Azerbaijani language.
The article examines speech culture as a key component of language competence among higher education students. The author emphasizes that mastering the norms of the literary language, adhering to ethical and stylistic standards in communication, is an indicator not only of a person’s general education but also of their readiness for professional and social interaction. The main components of speech culture are analyzed, including accuracy, clarity, logic, appropriateness, purity, expressiveness, and aesthetic quality of speech. Particular attention is paid to common violations of linguistic norms observed in the student environment: the use of colloquial, slang, and foreign words without necessity, unjustified calques, bureaucratic expressions, as well as syntactic and orthoepic errors. The article outlines the main causes of linguistic carelessness, such as low reading culture, the influence of social media, and the decline in linguistic standards in everyday and educational communication. The author proposes a number of pedagogical and methodological strategies aimed at cultivating a high level of speech culture among students. These include the integration of communicative training into the educational process, regular involvement of students in stylistic text analysis, and the activation of creative language practices. Examples of typical speech situations are provided to demonstrate the contrast between cultured and uncultured language use, highlighting the importance of speaker selfreflection in improving overall language competence. The relevance of the study is due to the growing importance of speech culture in the modern educational environment, where effective communication is a key component of a specialist’s professional training.
This study, based on one of the pioneering studies of József Herman (published in 2000) and using the capabilities and tools of the Computerized Historical Linguistic Database of Latin Inscriptions of the Imperial Age, attempts to describe and analyse the linguistic and dialectological characteristics of Latin in late imperial and early medieval (4-7th century) Italy as evidenced in inscriptions. Limited to the field of phonology, a consonantal and a vocalic phenomenon (the story of the word-final ‑s as well as the e-i and o-u mergers) is the subject of this study which aims to find out whether the late Latin territorial distribution of these changes predicts or anticipates the characteristic territorial differences within the Italian dialects and between the Italian and Sardinian dialects? The study shows that the question posed can be answered in an essentially positive way and that the time of geographically determined occurrence (or non-occurrence) of both discussed phenomena can be given with relative certainty.
In our contemporary world lots of social networks have emerged having an immense influence on the language development. Facebook, Twitter, Tiktok and others not only change the way we communicate, but also fundamentally transformed the nature of the language. Generally, the nature of the language development has always been a slow process. What once took decades or centuries now happens in months or years. New words, phrases, and expressions can spread globally within hours through viral posts, memes, and trending hashtags But social media significantly accelerates this process. The speed of spreading new words and terms is getting higher and higher. The dictionaries, academic institutions, and formal media—no longer control the pace of linguistic innovation. Linguistic innovations like Memes, Hashtags and viral phrases create new forms of language that don’t conform to traditional linguistic norms. Social media has made the trend toward informal language use much quicker and inevitable. Casual spelling, grammar, and vocabulary that were once characteristics of personal correspondence are now common in public discourse. Social media has indeed dismantled traditional boundaries between formal and informal language registers.As language evolvement is a gradual process it requires our great attention to follow the changes that can determine the true nature of the language.
This chapter explores the usage of third-person singular pronouns in English, arguing for the Sex-Tracking View, where pronouns (‘he/him’ and ‘she/her’) traditionally entail information about biological sex. Linguistic evidence supports this view, as pronouns are used consistently to denote sex in contexts involving non-human animals, suggesting that English pronouns have evolved to track biological sex rather than gender or “gender identity,” that is, one’s attitude about the sexes. I respond to objections from Michael Glanzberg and Cameron Domenico Kirk-Giannini, who propose a Gender-First View, by arguing, among other things, that their objections to the Sex-Tracking View may well confuse linguistic norms with social conventions. The chapter then transitions to normative considerations, questioning whether we should maintain the practice of sex-tracking with pronouns. Arguments in favor include the social importance of distinguishing between males and females for safety, the facilitation of sexual reproduction, and the moral implications of pronoun use as signaling acceptance or rejection of specific ideologies. However, I also consider objections to this practice: it may not affirm individuals’ attitudes about the sexes, including his own; it could be seen as exclusive; it may potentially invade privacy; and it might lead to social or professional repercussions for non-compliance. I then critically evaluate these objections.
Street’s (1984) concept of “ideological model” advocates for the plural character of literacy, validating all models of writing. From this perspective, marginalized literacy practices are as legitimate as dominant literacy practices, despite being socially stigmatized. In this article, I aim to investigate in what ways the literacy narratives of two adult students transgress social and linguistic norms. I employ a narrative analysis methodology that incorporates the discursive, situated and performative nature of the stories (MOITA LOPES, 2021) I analyze. At the social level, access to reading and writing is, for the students, a subversion of imposed norms that deprived them of the right to education. As women who migrated from the northeast to the southeast of Brazil in search of better living conditions, returning to school and learning to read and write are acts of resistance that challenge the status quo. At the linguistic level, the students reaffirm their enunciative intentions by transgressing the standard norm of the language, expanding the possibilities of meaning in their texts through the subversive use of punctuation and deixis. The outcomes emphasize the need to recognize the students’ literacy journey as a social criticism against insufficient policies on the provision of quality education for all. Furthermore, the outcomes point to the need for language teaching and learning approaches that recognize non-institutionalized models of literacy as valid knowledge that enhances and enriches language learning, minimizing the abyssal line (SANTOS, 2010) between orality and literacy as well as between school and non-school knowledge.
This study examines the linguistic landscape (LL) of Chikan Old Street, a historic district in Zhanjiang, Guangdong, through the lens of the SPEAKING model. As one of the most well-preserved historical districts in southern China, Chikan Old Street embodies a rich maritime heritage and commercial traditions, making it a compelling site for LL research. The study investigates the interaction between official language policies, regional linguistic identity, and globalization, providing insights into how language hierarchies are constructed in heritage sites. A mixed-methods approach is employed, integrating quantitative corpus analysis, qualitative semiotic interpretation, and public perception surveys. The findings reveal a clear stratification of language use: Chinese dominates official signage, with pinyin and English in subordinate positions, reflecting state-imposed linguistic norms. Private signage, however, demonstrates greater linguistic flexibility, incorporating Cantonese expressions, traditional Chinese characters, and creative bilingual adaptations. By highlighting the negotiation between top-down language standardization and bottom-up linguistic agency, this study contributes to broader discussions on language policy, cultural heritage preservation, and multilingual accessibility in historical districts. The findings underscore the need for improved linguistic planning, standardized translation policies, and greater public engagement in signage design to ensure that linguistic landscapes in heritage sites are both culturally authentic and globally navigable.
Modern Internet-mediated communication has a number of features related to external and internal factors influencing the process of such communication. Within the framework of the “communicative field” of modern social networks, Despite the widespread use of currently available technologies, the main tool of indirect communication is written speech. In addition to the manifestations of simple or intentional illiteracy in online communication, attention should be paid to the euphemization of content, as well as to the direct use of obscene vocabulary, and this does not depend much on the gender and age of the participants in the dialogue. Using the example of user content samples of the social network “In Contact”, the problem of speech behavior of participants in modern Internet communication was considered. It was noted that in conditions of free use of obscene vocabulary, its euphemization can be considered a manifestation of the intellectual superiority of a particular participant in communication, a certain level of mastery of the art of language play.
This chapter critically engages with the author’s book AI for Critical Interculturality (2025), the tenth in a series of retrospective analyses examining his decade-long contributions to critical intercultural studies. The book interrogates the role of AI in shaping and potentially disrupting ICER, challenging binary views of AI as either a threat or a solution. Framing AI as a ‘critical bosom friend’, the author explores its capacity to expose ideological biases in intercultural knowledge production, particularly neoliberal and Us-Euro-centric narratives. Methodologically, the book introduces three innovative frameworks: AI as mirror (revealing users’ epistemic shortcomings), language stratagems (disrupting AI’s linguistic norms through, e.g., paradox and hybridity) and ethical co-creation (rejecting extractive AI partnerships). The chapter highlights the key interventions of the book, decentring human exceptionalism, critiquing AI’s (hidden) curricula and advancing linguistic justice, while acknowledging limitations like technological determinism and Western-centric case studies. The author argues for considering AI as a tool for epistemic humility, urging students, scholars and educators to confront their own complicity in perpetuating uncritical discourses about technology and interculturality.
Despite recent advances in digital resources for other Coptic dialects, especially Sahidic, Bohairic Coptic, the main Coptic dialect for pre-Mamluk, late Byzantine Egypt, and the contemporary language of the Coptic Church, remains critically under-resourced. This paper presents and evaluates the first syntactically annotated corpus of Bohairic Coptic, sampling data from a range of works, including Biblical text, saints' lives and Christian ascetic writing. We also explore some of the main differences we observe compared to the existing UD treebank of Sahidic Coptic, the classical dialect of the language, and conduct joint and cross-dialect parsing experiments, revealing the unique nature of Bohairic as a related, but distinct variety from the more often studied Sahidic.
The design of Korean constituency treebanks raises a central representational question concerning the choice of terminal units. Although Korean words are morphologically complex, treating morphemes as constituency terminals can obscure the distinction between word-internal morphology and phrase-level syntactic structure, and can create mismatches with eojeol-based dependency resources. This paper argues for an eojeol-based constituency representation, with morphological segmentation and fine-grained POS information encoded in a separate, non-constituent layer. A comparative analysis shows that, under explicit normalization assumptions, the Sejong, Penn Korean, and KAIST treebanks can be compared over a shared eojeol-based constituency backbone. Building on this result, we outline an eojeol-based annotation scheme that preserves interpretable constituency, supports cross-treebank comparison and constituency-dependency alignment, and provides a surface-form terminal layer for future end-to-end Korean constituency parsing.
In 2025, we held the fourth iteration of the DIS-RPT Shared Task (Discourse Relation Parsing and Treebanking) dedicated to discourse parsing across formalisms.Following the success of the 2019, 2021, and 2023 tasks on Elementary Discourse Unit Segmentation, Connective Detection, and Relation Classification, this iteration added 13 new datasets, including three new languages (Czech, Polish, Nigerian Pidgin) and two new frameworks: the ISO framework and Enhanced Rhetorical Structure Theory, in addition to the previously included frameworks: RST, SDRT, DEP, and PDTB.In this paper, we review the data included in DISRPT 2025, which covers 39 datasets across 16 languages, survey and compare submitted systems, and report on system performance on each task for both treebanked and plain-tokenized versions of the data.The best systems obtain a mean accuracy of 71.19% for relation classification, a mean F 1 of 91.57(Treebanked Track) and 87.38 (Plain Track) for segmentation, and a mean F 1 of 81.53 (Treebanked Track) and 79.92 (Plain Track) for connective detection.The data and trained models of several participants can be found at https://huggingface. co/multilingual-discourse-hub.
This article examines how Vietnamese audiences interpret and reconstruct meaning from nonstandard lyrics, phonological deviations, and translingual mixing in contemporary vocal performance. While traditional approaches often treat such deviations as errors, this study reframes them as manifestations of post-standardized vocality, where expressive force arises from the voice, affective resonance, and performance context rather than lexical accuracy. Drawing on ethnographic fieldwork and listener interviews across northern, central, and southern Vietnam (2022–2024), the analysis documents phonological irregularities (e.g., chẩy vs. chảy) and emergent multilingual blending (e.g., Vietnamese–English and Vietnamese–Korean). Findings indicate that lyric intelligibility is sustained through a mechanism of cross-affective inference: audiences rely on timbre, embodied cues, and cultural familiarity to make sense of hybrid expressions. To theorize this process, the study introduces two concepts—post-standardized vocality and post-monolingual vocality—to explain how meaning is co-constructed beyond linguistic norms. These frameworks extend debates in lyric pragmatics, ethnomusicology, and postcolonial sound studies, highlighting the agency of listeners and proposing a comparative model for analyzing voice, affect, and intelligibility in global popular music.
This research explores how Micro, Small, and Medium Enterprises (MSMEs) in West Sumatra employ code-switching in their advertising discourse to construct linguistic identity, express cultural belonging, and project entrepreneurial modernity. Using Fairclough’s Critical Discourse Analysis (CDA) as the analytical framework, this study examines linguistic features, forms of code-switching, and the underlying ideological meanings within promotional banners and billboards that combine English and Indonesian. The findings reveal that code-switching serves as more than a marketing strategy it functions as a socio-symbolic practice through which entrepreneurs negotiate between local authenticity and global aspirations. The frequent use of English, despite notable errors in diction, spelling, and syntax, underscores its symbolic power as a marker of prestige and progress in the post-pandemic economic landscape. However, these linguistic inaccuracies also indicate challenges in language proficiency and access to educational resources, exposing power asymmetries between local entrepreneurs and global linguistic norms. From a sociolinguistic standpoint, code-switching embodies both empowerment and vulnerability: it enables small businesses to gain visibility in global markets while simultaneously revealing structural inequalities in linguistic capital. The study concludes that language operates as a key site of negotiation where identity, economy, and ideology intersect. It recommends enhancing critical language awareness and multilingual marketing literacy in MSME training programs. Future research is encouraged to examine digital advertising discourses to understand how linguistic entrepreneurship evolves in online spaces. This study contributes to the growing body of literature on sociolinguistics, linguistic entrepreneurship, and the politics of language in Indonesia’s evolving marketplace.
This study explores the impact of social media on the evolution of discourse structures and the linguistic identity of Generation Z. Using a qualitative, literature review methodology, the research analyzes existing studies to understand how English is adapted and reshaped within digital communication practices. Social media platforms, such as Twitter, Instagram, and TikTok, have led to a shift from traditional linguistic norms towards more informal, brief, and multimodal forms of communication. The review of scholarly works reveals that the structure of language has changed, with an increased use of abbreviations, emojis, and other non-textual elements, reflecting a more fluid and adaptive form of discourse. Additionally, the study highlights how social media serves as a platform for linguistic identity formation, where Generation Z constructs and negotiates their self-representation through language, often incorporating hybrid linguistic forms and code-switching. These changes, while facilitating new ways of expressing identity and fostering digital affiliation, may also pose challenges regarding the preservation of formal language skills in academic and professional contexts. The findings underscore the need for future research to explore the long-term effects of these shifts on language proficiency and their implications for education systems. This paper contributes to the growing understanding of the intersection between digital communication and linguistic evolution in the age of social media.