Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The article systematizes and comprehensively analyzes the phenomena of the phonetic level of the Internet vocabulary of the modern Kazakh language. Instagram, Facebook, social networks (Threads, Instagram, Facebook), and instant messengers (WhatsApp, Telegram), which have been actively used in recent years, were chosen as the object of the study. The research used methods of observation, generalization, comparative and descriptive analysis. As a result, it is revealed that new forms of linguistic usage are being formed in the Internet space, characterized by a mixture of elements of spoken and written speech. At the phonetic level, phenomena such as sound compression of words and, conversely, the repetition of graphemes to convey emotions in writing are widespread. The active use of Latin graphics and the development of foreign-language sounds indicate a new stage of phonetic adaptation in the Kazakh-speaking Internet space. The article provides specific examples of these phenomena, reveals their causes and impact on the modern linguistic norm and writing culture. According to the results of the study, it was found that the phonetic features of the Internet vocabulary reflect the natural development and adaptability of the Kazakh language.
This study investigates differences in artificial intelligence (AI) literacy and adoption between engineering students and faculty in a Middle Eastern higher-education institution. Parallel surveys were administered to undergraduate engineering students (N = 73) and faculty members (N = 20), each rating their familiarity with 20 AI tools covering learning, coding, productivity, and engineering applications. An AI Literacy Index was computed by assigning numerical values to familiarity ratings (A = 2, B = 1, C = 0) and normalizing the total to a 0–1 scale. Results from Welch’s t-test indicated that students demonstrated significantly higher literacy than faculty (0.454 vs. 0.356, p ≈ 0.042). Students also reported strong AI adoption for academic tasks (71.2%) and high perceived learning benefits (83.6%). Conversely, faculty expressed substantial concern about student over-reliance on AI (90%) while indicating readiness for professional development through AI training workshops (75%) and reporting assessment redesign efforts (75%). Overall, the findings highlight a meaningful literacy and perception gap with implications for engineering pedagogy, curriculum development, and assessment practices. Recommendations are provided to support the alignment of student and faculty AI competencies within engineering programs.
Kyrgyz, a Turkic language with over 4.4 million speakers concentrated primarily in Kyrgyzstan and adjacent regions of Central Asia, faces a significant disparity in computational linguistic resources compared to languages with similar or even smaller speaker populations. Despite its status as a government language and cultural cornerstone, Kyrgyz remains underrepresented in the digital linguistic landscape. This investigation examines the application of the Universal Dependencies (UD) framework – an annotation system engineered to facilitate cross-linguistic syntactic comparability – to the structural complexities of Kyrgyz. We endeavor to identify optimal annotation strategies that faithfully represent Kyrgyz-specific syntactic phenomena while adhering to the principled constraints of the UD paradigm. The establishment of standardized syntactic resources for Kyrgyz carries dual significance: it advances linguistic typology by incorporating data from an underrepresented language family, while simultaneously laying groundwork for practical natural language processing applications crucial for Kyrgyz speakers’ participation in the digital sphere. Our methodological approach encompasses rigorous analysis of nascent Kyrgyz treebanks, comparative evaluation of annotation strategies employed for genetically related Turkic languages, and systematic examination of four fundamental annotation challenges: the representation of Kyrgyz’s defective copula system, the classification of multifunctional grammatical particles, the annotation of constructions with implicit heads, and the demarcation between inflectional and derivational morphology in this highly agglutinative language. Our analysis reveals that achieving the dual objectives of linguistic fidelity and cross-linguistic consistency necessitates judicious adaptation of UD guidelines to accommodate Kyrgyz-specific structures. We advance unified annotation solutions that preserve the integrity of Kyrgyz linguistic patterns while facilitating meaningful cross-linguistic comparison. This research not only contributes substantively to computational resources for Kyrgyz but also establishes annotation principles with broader applicability to typologically similar agglutinative languages. The practical implications extend to enhanced guidelines for Kyrgyz treebank development, which will consequently improve parser accuracy and catalyze the development of essential language technology tools for Kyrgyz speakers.
This article explores the lexical-semantic and stylistic characteristics of paremiological units derived from moral and ethical lexis in English and Uzbek languages. Drawing on the principles of cognitive and cultural linguistics, the research examines how universal and culture-specific moral concepts — such as honesty, kindness, respect, patience, and decency — are conceptualized through proverbs and aphorisms. Comparative analysis reveals both shared humanistic values and national distinctions in semantic structures and stylistic expression. The findings demonstrate that paremiological units not only serve as repositories of cultural wisdom but also as linguistic manifestations of moral norms, ethical ideals, and social values in both languages.
Chinese word segmentation is a foundational task in natural language processing (NLP), with far-reaching effects on syntactic analysis. Unlike alphabetic languages like English, Chinese lacks explicit word boundaries, making segmentation both necessary and inherently ambiguous. This study highlights the intricate relationship between word segmentation and syntactic parsing, providing a clearer understanding of how different segmentation strategies shape dependency structures in Chinese. Focusing on the Chinese GSD treebank, we analyze multiple word boundary schemes, each reflecting distinct linguistic and computational assumptions, and examine how they influence the resulting syntactic structures. To support detailed comparison, we introduce an interactive web-based visualization tool that displays parsing outcomes across segmentation methods.
This article investigates the stylistic and linguistic features of Gustave Flaubert’s prose, with a particular focus on lexical contradiction and idiomatic tension. Through a close analysis of key works such as Madame Bovary, L’Éducation sentimentale, and Bouvard et Pécuchet, the study highlights how Flaubert’s writing systematically juxtaposes opposing semantic registers—romantic idealism and mundane realism, poetic elevation and trivial detail. These lexical contradictions not only enrich narrative depth but also underscore the disillusionment and irony characteristic of Flaubert’s modern vision.The article further explores how Flaubert manipulates idiomatic expressions, either by subtly distorting them or by integrating them ironically into character discourse. This tension between conventional language and authorial critique reveals Flaubert’s ambivalent relationship with linguistic norms and his pursuit of lemot juste. Drawing on French and Francophone critical literature, the study situates Flaubert’s stylistic innovation within broader debates about the literary function of cliché, the evolution of free indirect discourse, and the modern fragmentation of narrative voice.By analyzing the paradoxes at the heart of Flaubert’s style, the article demonstrates how lexical contradiction and idiomatic tension function not only as aesthetic devices but also as means of epistemological inquiry—interrogating language,meaning, and the act of writing itself.
Inflection of Belarusian surnames is a topic of nationwide importance in Belarus. To this day, there is a great diversity in reference literature, and the standardization of surname inflection has not been established. The article describes the declension variations of Belarusian surnames in the Belarusian language, influenced by Russian patterns that are incompatible with standard practices of Belarusian speakers. Discussion includes the formation of the Belarusian linguistic norm, which has experienced historical and continuing influence by extra-linguistic factors. Complex instances of declension of Belarusian surnames in both Belarusian and Polish as well as of Polish surnames in both Polish and Belarusian are thoroughly discussed. On the one hand, comparative analysis shows the subtle similarities and differences between the two Slavic languages, from which interference errors can result. On the other hand, there are noticeable tendencies for 1) Belarusian surnames to decline according to the Russian models, as well as 2) Belarusians living in Poland to express preferences to decline Belarusian surnames used in the Polish language, according to the rules of Belarusian language.
The study of linguistic variation within the administrative structures of small-town America reveals a complex intersection between language, social identity, and institutional behavior. When approaching the linguistic environment of these communities from a purely academic perspective, without relying on personal immersion narratives or experiential accounts, one must begin with the foundational premise that English in the United States is profoundly regionalized. This regionalization is not a superficial matter of accent or vocabulary; it is a system of deeply embedded linguistic norms that shape how communication occurs, how authority is interpreted, and how institutional legitimacy is constructed. My interest as a researcher lies not in documenting local flavor or collecting curiosities from rural life but in understanding the mechanisms by which language operates as a structural force within governance. This requires an examination of sociolinguistic corpora, regional dialect research, institutional discourse studies, and the extensive literature on American dialect geography that has accumulated since the mid-twentieth century.
This article analyzes the impact of information and communication technologies (ICT) on the dynamics of contemporary language, focusing on Russian and Bulgarian. In the context of globalization and digitalization, it traces the main trends in lexical, grammatical, and word-formation transformations caused by new technologies and virtual communication. The study is based on a corpus of over one thousand jargon units from the ICT field, through which the process of adaptation and assimilation of foreign borrowings, primarily from English, is examined. Special attention is given to hybridization, digraphy, and visual communication (emoticons, hashtags, memes), which form a new type of linguistic reality. The results show that ICT accelerate linguistic changes, expand the norms of the standard language, and stimulate the interaction between standard and substandard vocabulary.
This research aims to analyze the Politeness and Speech Acts of the Community (Ojol Community). This research uses a qualitative approach with descriptive method. Data were obtained through direct observation and recording of conversations between online ojek drivers and customers in real situations. Recording is done naturally without intervention to reflect authentic speech acts. The audio data is then transcribed and analyzed using Searle's speech act theory. The analysis is done descriptively qualitative by classifying and interpreting the form and function of utterances in the context of the conversation. The results show that nonstandard language is more dominantly used in informal communication, such as conversations between online ojek drivers and passengers, because it is considered more familiar, relaxed, and efficient. However, mastery of standardized language is still important, especially in official contexts, to maintain clarity and politeness. People are expected to be able to adjust the use of language according to the context so that communication remains effective and in accordance with linguistic norms. Keywords:,,,,.
In response to widespread leadership crises that favor popularity over competence, undermining strategic decision‑making and ethical standards, this study articulates the Prophet Muhammad's (PBUH) communication and diplomatic principles in his royal correspondence as a leadership archetype grounded in justice, meritocracy, meticulous composition, and mutual respect. Employing a qualitative literature approach with historical content and comparative analysis, it examines primary manuscripts of prophetic letters and their contexts alongside classical sirah texts and peer-reviewed studies. The analysis uncovers four pivotal elements: concentrated da'wah message summaries; the "Muhammad Rasul Allah" seal for authentication; envoy selection tailored to each court's linguistic norms; and Qur'anic citations for spiritual authority. These elements demonstrate a synergistic blend of prophetic legitimacy and diplomatic courtesy, offering a framework for religious rhetoric and ethical leadership development. The study also recommends rigorous comparative diplomacy across global traditions and innovative, strategic interdisciplinary collaboration for future scholarly inquiry.
In ”Bartleby, the Scrivener,” Herman Melville presents a character whose passive refusal, encapsulated in the repeated phrase “I would prefer not to,” challenges power, agency, and social norms. This essay examines how Bartleby’s refrain acts as both an assertion of autonomy and a critique of the violence inherent in language. By rejecting his employer’s commands, Bartleby disrupts the rational, efficiency-driven logic of the workplace, exposing the violence embedded in linguistic norms. Slavoj Žižek’s concept of language as inherently violent—through its imposition of norms and standards— illuminates how Bartleby’s refusal goes beyond protest, creating a space of resistance that defies interpretation and subverts power dynamics. Bartleby’s language, neither a clear denial nor an expression of desire, becomes a radical negation that questions the very nature of meaning. Ultimately, Bartleby’s refusal does not propose a new order but disrupts the structures of meaning and authority, forcing us to confront the limits of language itself.
As it is well known, sociolinguistics is based on the idea that language is a social institution and an acquired phenomenon, established by members of society through mutual agreement to fulfill their needs and desires. Ibn Jinni defines language as "sounds through which each group expresses its purposes." One of the main reasons for the emergence of Arabic grammar was the social motivation of pride in the Arabic language and the need to preserve it from those who entered Islam, as they required learning Arabic to study and memorize the Qur’an. Since this interaction had an impact on the language, I found it necessary to explore this interdisciplinary sociolinguistic approach in the phenomenon of lahn (linguistic errors). The study is structured into an introduction, two main sections, a conclusion, and a bibliography. The introduction discusses the interdisciplinary study within sociolinguistics. The first section examines the impact of societies during the Islamic conquests on the Arabic language. The second section addresses the dangers of lahn and its effects on Arabic linguistic norms. The conclusion summarizes the key findings of the study.
BACKGROUND AND OBJECTIVES: The Seeking Proxies for Internal States (SPIS) model of OCD posits that reduced access to internal states plays a key role in the development and maintenance of the disorder. The current work sought to provide further support for the model's central claim that obsessive-compulsive tendencies are associated with reduced access to internal states. METHOD: Participants (N = 170) listened to 60 sound stimuli, rated how each one made them feel, and completed a measure of obsessive-compulsive tendencies. Following past procedure, we compared participants' ratings to each sound's normative valence rating, such that higher deviations between the ratings reflect a noisier perception of affective internal states. RESULTS: As hypothesized, higher obsessive-compulsive tendencies predicted greater deviations for both normatively-positive and normatively-negative sounds. CONCLUSIONS: The current work provides additional, novel support for the SPIS model, showing that with increasing obsessive-compulsive tendencies, people exhibited reduced attunement to how auditory stimuli made them feel.
Abstract This paper argues that the study of binary personal pronouns needs to move beyond European languages and the focus on third-person pronouns, and it supports this argument by presenting the problem of the Japanese linguistic norm of binarily gendered first-person pronouns. The Japanese case illustrates five ways in which this field of study should expand. First, binary pronoun studies cannot continue to neglect first-person pronouns. Second, in addition to pronoun systems, the norms of usage must be considered. Third, Japanese requires different forms of resolution for the problem of binary pronouns. Fourth, research should explore linguistic self- and other-referring expressions by examining the distinct functions of first- versus third-person pronouns. Finally, the Japanese case demonstrates that the parties directly affected by binarily gendered pronouns include girls and women as well as gender-nonconforming and trans people. Considering the Japanese case thus potentially contributes to expanding the field in five important ways.
The incorporation of real-world contexts into mathematical problems has become increasingly significant on a global scale, which is also reflected in Vietnamese education reform. This emphasis has resulted in the development of different series of textbooks designed to enhance student engagement and foster active learning through real-world problem contexts. In particular, the geometry and measurement strand presents substantial opportunities for the integration of contexts. In this study, we investigate the theoretical foundations of contexts, context familiarity, and context authenticity in mathematics education, with a focus on the plane geometry content in ninth-grade Vietnamese textbooks. The analysis is supported by four illustrative examples drawn from three reformed textbook series, demonstrating how context familiarity and authenticity are represented. A 5-point Likert-type survey of 109 students (72 males and 37 females) revealed that the soccer context was rated as less familiar than the airplane context overall, with no significant difference in familiarity ratings between boys and girls. Both the bike-riding and tripod contexts were rated as moderately authentic but not highly so. The authors also acknowledge several limitations and offers certain recommendations for future research in the area of contextual mathematics education.
French is often celebrated for its clarity and precision – a legacy shaped by Cartesian rationalism and prescriptive language policies. However, the evolving forms of spoken French challenge this ideal of fixed linguistic norms. This study examines one such feature: the right-peripheral duplication of the subject pronoun je with its tonic counterpart moi, a recurrent but underexplored phenomenon in spoken French. The primary objective is to understand how this syntactic feature functions pragmatically and emotionally in real-life discourse. Using a corpus of movie dialogues, the analysis shows that duplication plays a role in managing conversational flow, expressing personal stance, and enabling self-repair. Through a multidisciplinary lens that draws from sociolinguistics, pragmatics, and applied linguistics, the study argues that such variation enriches the expressive potential of French and complicates the rigid divide between written norms and spoken practice. It also suggests that incorporating these features into language pedagogy can support a more inclusive, realistic understanding of French as a living language.
This research paper examines the sociological significance of dialects and accents, analyzing their role in shaping social identities, reinforcing hierarchies, and influencing systemic biases. Grounded in sociological theories, particularly those of Pierre Bourdieu, Erving Goffman, Max Weber, and Michel Foucault, the study explores how language functions as a form of symbolic capital that dictates access to social mobility and power. Through a critical analysis of language as a site of inclusion and exclusion, the paper highlights how dominant linguistic norms marginalize non-standard dialects, perpetuating social stratification. Additionally, the study investigates the role of media, globalization, and cultural representation in shaping linguistic perceptions and maintaining or challenging linguistic hegemony. While dialects and accents often serve as markers of discrimination, they are also powerful tools for cultural identity and resistance. This paper underscoresthe need for greater linguistic inclusivity in institutional, educational, and social contextsto combat entrenched biases and promote equitable linguistic representation.
The relevance of the article is due to the fact that language plays a decisive role in understanding the world around us and establishing communication between people. Language is a dynamic system, constantly developing and changing. At any given time, language contains both new and obsolete elements. The reasons, mechanisms and algorithms for the emergence of new linguistic units often lead to the variant use of these elements. Variation is thus an integral part of the synchronic and diachronic development of language. It is also a fundamental feature characterizing the functioning of all linguistic units. Therefore, the problem of variability remains relevant in linguistics and requires further study. The presence of linguistic variants can create difficulties for the speaker or writer, causing doubts about choosing one of them. Therefore, the problem of variability is closely related to the norm of the literary language. Although the norm should ensure the stability of the language, it is not static and allows for the emergence of new variants. The presence of different variants of a word or its forms leads to paradigmatic diversity in language. Accordingly, some scientists consider it necessary to eliminate the options in order to preserve the literary norm. However, as most scientists believe, paradigmatic diversity is a necessary stage in the transformation of the linguistic system elements of any literary language, since it has a positive effect on the development of literary norms. Variants are also used as a means of stylistic enrichment of speech. In our research, we attempt to find out the factors of the occurrence of spelling variants of words in the modern Tatar language in terms of the functional and structural classification of the language. The practical part analyzes the collected file (spelling variants of the same word). In the course of the study, the term “a variable nest”, based on the linguistic phenomenon “a lexical nest”, was introduced.
Presidential Speeches are key instruments of political communication. They are powerful tools used to shape public opinion or set nation-level agendas as a presidential leader. Analyzing these speeches helps uncover the underlying ideologies, emotional appeals, and rhetoric strategies used by presidents to influence the perception of the government. This paper analyzes speeches delivered by the presidents of the United States of America from George Washington in the year 1789 to Donald Trump in the year 2020, using the U.S. Presidential Speeches Dataset from Miller Center. It presents a computational linguistic analysis of historical presidential speeches spanning from the late 18th century to the 21st century until January of the year 2020. Utilizing natural language processing techniques, we analyze various linguistic dimensions including lexical diversity, sentiment, rhetorical structure, thematic content, and ideological positioning. Our key findings reveal a gradual decrease of 5-10% in lexical diversity every decade and the strong use of repetition across contexts. In addition, we identified trends correlated with historical events, party affiliation, and changing communication norms. This research provides insights into how presidential communication has evolved and how linguistic patterns reflect broader societal and political changes throughout American history.
<p xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" class="first" dir="auto" id="d5762881e86">Some translators of the New Philosophy viewed linguistic purism as one of the ways to making the Dutch language fit for the purpose of communicating rationalist knowledge. Previous scholars argued that their lexical preferences were determined by the purist norms proposed by Lodewijk Meijer and Adriaan Koerbagh, who used lexicography and etymology as support for their radical criticism on orthodoxy and Calvinist theology. In this chapter, computational methods are applied to test this hypothesis that translators of the New Philosophy were more likely than their contemporaries to follow the purist norms propagated by Meijer and Koerbagh. It describes and evaluates a method designed to automatically detect and quantify loanwords and philosophical terms in early modern Dutch texts.
Rasheed Hasan Khan (1925–2006) occupies a central position in Urdu scholarship for his rigorous approach to research (tahqeeq) and textual editing (tadween). This article examines his scholarly temperament and methodological contribution with particular reference to his influential work Zaban aur Qawaid, a collection of essays addressing core issues of Urdu usage, orthography, pronunciation, lexicography, and grammatical convention. The study highlights how Khan transforms complex linguistic questions into accessible arguments through careful citation, comparative consultation of dictionaries, and close reading of classical poetic and prose sources. Special attention is given to his critical engagement with prescriptive claims about “correct” forms, especially in cases where Arabic–Persian norms intersect with Urdu’s historical and evolutionary language practices. By analyzing selected discussions—such as the standardization of pronunciation and spelling, the evaluation of contested lexical forms, and the treatment of shared-gender (mushtarak) words—this article demonstrates that Khan’s approach is neither merely conservative nor casually permissive; rather, it is a principled, evidence-based model that privileges actual Urdu literary usage while remaining alert to etymology and linguistic structure. The paper concludes that Zaban aur Qawaid represents an enduring scholarly framework for balancing norm, usage, and linguistic change, and it continues to inform contemporary debates on Urdu standardization, dictionary-making, and editorial practice.
The current study aimed to examine the effects of topic familiarity and language proficiency on linguistic complexity, accuracy, and fluency of argumentative essays in an EFL context. This study involved 64 college freshmen, who were divided into two groups according to a TOEIC score of 700, which corresponds to the B1 level on the CEFR scale: a high group (n = 31) and an intermediate group (n = 33). Participants were asked to write two argumentative essays on different topics, controlling for order effects through counterbalancing: one familiar (driving) and the other unfamiliar (smoking). They also completed a questionnaire that included background information and a topic familiarity rating on a 10-point Likert scale. The participants’ writing samples were analyzed in terms of lexical complexity, syntactic complexity, accuracy, and fluency. The results indicated that, in terms of topic familiarity, EFL learners tended to produce texts with lower levels of lexical and syntactic complexity, as well as reduced accuracy, when writing about unfamiliar topics. With regard to language proficiency, advanced learners demonstrated a broader vocabulary range, employed longer and more complex sentence structures, and produced more accurate and extensive texts compared to their intermediate-level peers. In-depth analysis and pedagogical implications are discussed.
The article analyzes the key theoretical and practical aspects of developing lexical competence in the process of learning Ukrainian as a foreign language. Lexical competence is considered an essential component of foreign language communicative competence, ensuring effective communication in accordance with linguistic and cultural norms. The author explores the peculiarities of vocabulary acquisition at the initial (A1), basic (A2), and threshold (B1) levels of Ukrainian language proficiency. The study examines teaching methods that include systemic-linguistic, conditional-communicative, and communicative types of exercises, as well as interactive methods such as language games, working with texts, and visual learning aids. Through the communicative approach, which involves working with realistic texts, role-playing games, and situational dialogues, the process of immersing foreign learners in the language environment is implemented. Based on the topic “Professions,” both traditional systemic-linguistic and communicative tasks adapted to the needs of foreign learners are proposed. The presented set of exercises is aimed at the gradual development of lexical skills, ranging from familiarization with new words to their active use in speech. It facilitates the development of reading, speaking, listening, and writing skills. These tasks help foreign learners work with texts and create their own. The use of such methods contributes to increasing students’ motivation and enhancing the effectiveness of vocabulary acquisition. Reading or listening to texts about famous Ukrainians fosters linguistic and cultural competence. The described approaches can be applied both in classroom settings and for independent student work. The practical significance of this study lies in the introduction of modern approaches to teaching Ukrainian vocabulary as a foreign language. The proposed approaches and types of exercises can be adapted for studying other topics in a foreign language audience at all proficiency levels. Key words: methods of teaching Ukrainian as a foreign language, lexical competence, language proficiency levels, exercises.
While Large Language Models (LLMs) have shown remarkable performance in various Natural Language Processing (NLP) tasks, their effectiveness seems to be heavily biased toward high-resource languages.This proposal aims to address this gap by developing efficient training strategies for low-resource languages.We propose various techniques for efficient learning in simulated low-resource settings for English.We then plan to adapt these methods for lowresource languages.We plan to experiment with both natural language generation and understanding models.We evaluate the models on similar benchmarks as the BabyLM challenge for English.For other languages, we plan to use treebanks and translation techniques to create our own silver test set to evaluate the low-resource LMs.
The post-independence era has marked a new phase in the linguistic relationship between Azerbaijani and Turkish, characterized by an increased lexical exchange. During this period, many loanwords from other languages have been systematically replaced by borrowings from modern Turkish. The trend toward linguistic convergence between Azerbaijani and Turkish has been particularly evident in the language of the press and mass media. Additionally, words that were historically common to both languages but had remained in limited usage within Azerbaijani have been reactivated and reintegrated into everyday discourse. A significant portion of the lexical borrowings from Turkish during this period has served as a substitute for Russian and European-origin words introduced via Russian, as well as for Arabic and Persian loanwords. The presence of Turkish-origin vocabulary in contemporary Azerbaijani continues to expand, with borrowed words exhibiting diverse morphological structures, including simple, derived, and compound forms. However, a notable concern in recent linguistic developments is the unregulated incorporation of Turkish words into Azerbaijani, sometimes without genuine necessity. This phenomenon raises questions about the potential impact on the structural integrity of the Azerbaijani language. To safeguard the purity and coherence of the language, it is imperative to adhere to established literary norms and linguistic standards
We present a family of encodings for sequence labeling dependency parsing, based on the concept of hierarchical bracketing. We prove that the existing 4-bit projective encoding belongs to this family, but it is suboptimal in the number of labels used to encode a tree. We derive an optimal hierarchical bracketing, which minimizes the number of symbols used and encodes projective trees using only 12 distinct labels (vs. 16 for the 4-bit encoding). We also extend optimal hierarchical bracketing to support arbitrary non-projectivity in a more compact way than previous encodings. Our new encodings yield competitive accuracy on a diverse set of treebanks.
We investigate the application of Neural Quantum Embedding and Quantum Neural Networks for sentiment analysis using the Stanford Sentiment Treebank dataset. We adapt the Neural Quantum Embedding framework, originally worked with image classification, to textual data by employing quantum embedding techniques that maximize trace distance for better data separability. Additionally, we incorporate advanced quantum feature maps and preprocessing techniques from recent quantum machine learning studies to enhance classification performance. Our experimental results demonstrate that the quantum models catch up with the classical baselines in sentiment classification accuracy, particularly in noisy intermediate-scale quantum settings. This work highlights the feasibility and potential of quantum-assisted sentiment analysis.
Analysis code and data for "Asymmetric admixture decouples gene–language coevolution in Eastern Eurasia" This repository contains all computational code and the TyDEE (Typological Dataset of Eastern Eurasia) linguistic database used in our study examining gene-language relationships across Eastern Eurasia. The dataset comprises 541 language varieties across 10 major language families, paired with genome-wide genetic data from 135 populations.
The digital era has brought profound changes to language, particularly visible in the rapid emergence of new lexical units. This article explores how internet communication, social media, and technological innovations influence language change, leading to the creation and diffusion of novel words and expressions. By analyzing examples from contemporary digital discourse, the study highlights the dynamic nature of language and the impact of digital culture on vocabulary expansion. The article also discusses the implications of these changes for language norms and linguistic identity.
Kyrgyz remains a low-resource language with limited foundational NLP tools. To address this gap, we introduce KyrgyzBERT, the first publicly available monolingual BERT-based language model for Kyrgyz. The model has 35.9M parameters and uses a custom tokenizer designed for the language's morphological structure. To evaluate performance, we create kyrgyz-sst2, a sentiment analysis benchmark built by translating the Stanford Sentiment Treebank and manually annotating the full test set. KyrgyzBERT fine-tuned on this dataset achieves an F1-score of 0.8280, competitive with a fine-tuned mBERT model five times larger. All models, data, and code are released to support future research in Kyrgyz NLP.
Salient sexual cues (erect penis, attractive individuals) are thought to capture initial attention and automatically trigger genital arousal in women. Conscious appraisal activates subjective sexual arousal and further visual attention. The present study tested whether the attractiveness category (attractive/unattractive) and/or sexual arousal condition (in underwear, naked with flaccid penis, naked with erect penis) of male stimuli predicted female sexual responding and attentional patterns. Genital arousal, visual attention, and subjective ratings (subjective sexual arousal, pleasantness) of 26 predominantly heterosexual women (Mage = 31.2, SDage = 6.8) were measured while exposed to male stimuli across the experimental conditions. Neither attractiveness nor sexual arousal condition of male models significantly predicted genital arousal, subjective sexual arousal ratings or visual attention patterns of women. Results however showed high concordance between genital and subjective sexual arousal measures, and the pleasantness rating of the stimuli positively predicted subjective sexual arousal, suggesting a positive feedback loop of female sexual response.
The literary works of Oscar Wilde are characterized by wit, aestheticism and artistic use of language. This paper examines linguistic anomalies in Wilde's plays, prose, and poetry, particularly the ways he subverts language to generate humour, irony, and social critique. Focusing on deviation at the lexical, syntactic, semantic and phonological levels, the study explores the use of style in The Importance of Being Earnest, The Picture of Dorian Gray, selecting epigrams. The results show that Wilde's studied departures from linguistic norms serve to undermine Victorian convention, provide the maximum satirical impact, and draw attention to aesthetic beauty.
People's language and behavior can be greatly influenced by gender differences.Female communication patterns have historically been characterized by tentative expressions and men use nonstandard language and more slang than women.However, with the development of society and the changes of cultural values, gender language develops and the language differences are not static.Women are asked to behave politely and humbly in the past but now they show more confidence.Gender differences in language manifest in many aspects, and literary works are no exception.In Little Women, Louisa May Alcott challenged traditional gender-based linguistic norms to present an independent and self-reliant girl Jo March who is fond of using slang, acts like a boy and treats herself as a boy.Jo uses masculine language and exhibits more male characteristics in communication.Jo 's unique character also reflects the profound exploration of female independence and freedom, and her thoughts and actions become a model for women to pursue self-worth and independent life.
Kyrgyz remains a low-resource language with limited foundational NLP tools. To address this gap, we introduce KyrgyzBERT, the first publicly available monolingual BERT-based language model for Kyrgyz. The model has 35.9M parameters and uses a custom tokenizer designed for the language's morphological structure. To evaluate performance, we create kyrgyz-sst2, a sentiment analysis benchmark built by translating the Stanford Sentiment Treebank and manually annotating the full test set. KyrgyzBERT fine-tuned on this dataset achieves an F1-score of 0.8280, competitive with a fine-tuned mBERT model five times larger. All models, data, and code are released to support future research in Kyrgyz NLP.
This essay reimagines public speaking education through a culturally sustaining and transgressive lens that challenges dominant norms of language, professionalism, and communication competence. It critiques the ways in which public speaking courses often reinforce linguistic supremacy by privileging standardized English and marginalizing multilingual and culturally grounded speech practices. Drawing on concepts such as translanguaging, Culturally Sustaining Pedagogy (CSP), and transgressive pedagogies, the author calls for a shift in pedagogy that centers students’ lived experiences, community-rooted knowledge, and linguistic norms. Rather than asking students to conform to hegemonic standards, this approach empowers them to speak on their own terms, resist assimilationist pressures, and use language as a tool for identity, resistance, and liberation. By transforming the public speaking classroom into a space for critical reflection and empowerment, educators can cultivate more inclusive and equitable models of communication instruction.
This chapter synthesises the findings and discusses how sociomaterial processes shape languages. Challenging modernist linguistic paradigms, it examines how language categories emerge through diverse cultural, historical, and material practices. The chapter critiques binary linguistic models and universalist, teleological assumptions of standardisation, showing that stable linguistic systems are not ‘natural’, but result from specific sociopolitical and material conditions. In contrast, fluid linguistic practices in postcolonial and globalised contexts exhibit variability, innovation, and complex indexicality. Belize’s multilingual environment exemplifies a setting without a hegemonic linguistic centre, producing liquid linguistic norms. The chapter argues for decolonial approaches to linguistics that embrace heterogeneity and that challenge exclusionary, Eurocentric models. Ultimately, it positions fluid linguistic practices as a cultural avant-garde and understands postcolonial environments as inspiring insights into future global sociolinguistic orders shaped by digitalisation and transnationalism.
Abstract Combining research in developmental sociolinguistics and L1 acquisition, this study explores how caregivers may orient children towards (socio)linguistic norms through parental feedback. Based on self-recorded family interactions in the Belgian-Dutch setting, it applies a top-down quantitative perspective to examine feedback on non-conventional versus non-standard language use, alongside a bottom-up qualitative perspective highlighting factors that influence parental feedback occurrence. Findings reveal limited feedback on children’s non-standard language use, with participation frameworks and multiactivity contexts emerging as possible constraints. The combined approach also foregrounds possible tensions between researcher categorisations and participants’ perspectives. Overall, this study offers a first step in bridging research on parental feedback and sociolinguistic variation, identifying patterns that merit further investigation.
The purpose of this study is to evaluate the English language proficiency of Sayed Jamaluddin Afghani University students enrolled in the English Department. It primarily looks at the grammatical mistakes that impede kids' language development and how they affect their ability to communicate. A structured questionnaire was used in a quantitative survey with a sample of 120 students in order to collect data. The results show that improper application of linguistic norms, infrequent practice, and the effect of students' native language, which frequently interferes with English usage, are the primary causes of grammatical errors. The efficacy and clarity of students' speech are greatly impacted by these errors. According to the study's findings, English professors should implement specialized courses that emphasize communicative grammar instruction. It also recommends giving students regular chances to practice active language usage, which can improve their overall communication skills and grammatical precision, ultimately increasing their English proficiency.
This study aims to analyze the interference of Indonesian in Arabic translation among students of the Arabic Language Education program, focusing on morphological, syntactic, and lexical errors. The research employed quantitative, qualitative, and descriptive approaches, with data collected through a document study of students’ theses translated from Indonesian into Arabic. The analysis was conducted to identify the types of errors, their frequency, and the underlying factors affecting translation quality. The findings indicate that Indonesian interference occurs at multiple linguistic levels, affecting the coherence, cohesion, and stylistic appropriateness (Uslub) of the translated texts. Morphological errors included word-for-word translation, incorrect verb conjugation, and gender disagreement; syntactic errors involved word order, misuse of conjunctions, and improper clause combination; while lexical errors consisted of inappropriate word choice, literal translation of idiomatic expressions, and inaccurate use of technical terms. These results underscore the need for targeted training in linguistic rules, stylistic norms, and discourse practice, alongside the development of cultural and pragmatic awareness. The study concludes that Indonesian interference significantly influences Arabic translation, manifesting at morphological, syntactic, and lexical levels.
This article examines the cross-cultural features of lexical intensification in English and Uzbek political news discourse. The study investigates how journalists in both languages use intensifiers such as scalar adverbs, extreme adjectives, hyperbolic expressions, and culturally embedded evaluative units to shape ideological framing and influence audience perception. By comparing representative political news texts, the research identifies structural, semantic, and pragmatic similarities and differences in the use of intensified vocabulary. The findings show that English political discourse tends to employ graded lexical choices for subtle persuasion, whereas Uzbek discourse relies more heavily on emotionally charged and culturally resonant expressions. Overall, the analysis reveals how linguistic and cultural norms shape the communicative strategies of political media.
The development of social media as a public communication space has had a significant impact on language use, particularly grammar. This article discusses the use of grammar on social media, which is often caused by the freedom of expression not being balanced with an awareness of proper and correct language. Through a qualitative approach with discourse analysis, this article highlights how deviations from grammatical rules in social media posts reflect a shift in societal attitudes towards linguistic norms. On one hand, linguistic freedom on social media is considered a form of creativity and self-expression. However, on the other hand, it has the potential to weaken language skills that adhere to rules, especially among the younger generation. This article will certainly provide a recommendation on the importance of language literacy as an effort to maintain the quality of public communication amidst the wave of digitalization and the expanding freedom of expression.