Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This chapter explores how linguistic, cultural, and ethnic identities are displayed, intertwined, and negotiated by YouTube creators in their reaction videos on BTS content. The reaction video, “a meta-video that documents in real-time individuals watching another video usually for the first time,” has emerged as a popular genre of user-generated content and a venue for online participation on YouTube for K-pop. Employing multimodal analysis, the chapter focuses on one YouTube channel run by a Korean American reactor and the comments and interactions made by the reactors and viewers on various aspects of BTS content. The analysis focuses on how reactors construct and perform their fan identity and position themselves vis-à-vis BTS in terms of language, culture, and ethnicity, including the hybrid language use of Korean and English and cultural aspects observed in BTS content. It is argued that reaction videos function as a new third space, in which the reactors’ different identities are (re)constructed and mediated as developing YouTubers and Korean American fans, and linguistic norms and cultural hierarchies are reordered and renegotiated.
There are limited discussions on how translanguaging practices may be tailored according to needs and contexts by providing examples of the implementation of translanguaging in different countries. Thus, this chapter reports on implementing translanguaging premises and pedagogies in light of critical needs analysis to compare and offer practical recommendations. Türkiye (at a state university where medical English was delivered using CLIL), Brazil (at a state university where linguists and computer scientists develop annotated treebanks of diverse dialects to be used in Natural Language Processing applications, among them a language used by the Warao refugees in Brazil, emerging from the contact between the Warao language, Venezuelan Spanish, and Brazilian Portuguese). Here, the target situation is a site for possible transformation and a translanguaging approach allows. To the knowledge of this chapter's authors, this study in two different contexts is the first of its kind on translanguaging.
This article offers a reflection on the meaning effects produced by the use of foreign languages in contemporary playwriting, taking as a starting point Henri Meschonnic’s notion of rhythm, which displaces language beyond signification and code. In dialogue with Marion Chénetier-Alev, who identifies the use of other languages as a critical and political tool in the European stage context, the text turns to the work of Brazilian playwrights who mobilize foreignness as a poetic operation. The focus then shifts to the contemporary Brazilian stage, observing experiences in which the foreign—through language, through rhythm—crosses the text. These practices challenge the hegemony of linguistic norms and activate Brazilian Portuguese as a sensitive matter, embedded in histories, violence, and translation. Playwriting is thus conceived as an operation that constructs new meanings in the passage between languages, establishing a zone of contagion where language no longer recognizes itself in a stabilized form. The article ultimately proposes the idea of a theatrical language—unstable, rhythmic, continuous, ephemeral—that invents and reconfigures the relations between language, body, and world.
Uncertainty of scientific findings are typically reported through statistical metrics such as $p$-values, confidence intervals, etc. The magnitude of this objective uncertainty is reflected in the language used by the authors to report their findings primarily through expressions carrying uncertainty-inducing terms or phrases. This language uncertainty is a subjective concept and is highly dependent on the writing style of the authors. There is evidence that such subjective uncertainty influences the impact of science on public audience. In this work, we turned our focus to scientists themselves, and measured/analyzed the subjective uncertainty and its impact within scientific communities across different disciplines. We showed that the level of this type of uncertainty varies significantly across different fields, years of publication and geographical locations. We also studied the correlation between subjective uncertainty and several bibliographical metrics, such as number/gender of authors, centrality of the field's community, citation count, etc. The underlying patterns identified in this work are useful in identification and documentation of linguistic norms in scientific communication in different communities/societies.
This study investigates differences in artificial intelligence (AI) literacy and adoption between engineering students and faculty in a Middle Eastern higher-education institution. Parallel surveys were administered to undergraduate engineering students (N = 73) and faculty members (N = 20), each rating their familiarity with 20 AI tools covering learning, coding, productivity, and engineering applications. An AI Literacy Index was computed by assigning numerical values to familiarity ratings (A = 2, B = 1, C = 0) and normalizing the total to a 0–1 scale. Results from Welch’s t-test indicated that students demonstrated significantly higher literacy than faculty (0.454 vs. 0.356, p ≈ 0.042). Students also reported strong AI adoption for academic tasks (71.2%) and high perceived learning benefits (83.6%). Conversely, faculty expressed substantial concern about student over-reliance on AI (90%) while indicating readiness for professional development through AI training workshops (75%) and reporting assessment redesign efforts (75%). Overall, the findings highlight a meaningful literacy and perception gap with implications for engineering pedagogy, curriculum development, and assessment practices. Recommendations are provided to support the alignment of student and faculty AI competencies within engineering programs.
The article systematizes and comprehensively analyzes the phenomena of the phonetic level of the Internet vocabulary of the modern Kazakh language. Instagram, Facebook, social networks (Threads, Instagram, Facebook), and instant messengers (WhatsApp, Telegram), which have been actively used in recent years, were chosen as the object of the study. The research used methods of observation, generalization, comparative and descriptive analysis. As a result, it is revealed that new forms of linguistic usage are being formed in the Internet space, characterized by a mixture of elements of spoken and written speech. At the phonetic level, phenomena such as sound compression of words and, conversely, the repetition of graphemes to convey emotions in writing are widespread. The active use of Latin graphics and the development of foreign-language sounds indicate a new stage of phonetic adaptation in the Kazakh-speaking Internet space. The article provides specific examples of these phenomena, reveals their causes and impact on the modern linguistic norm and writing culture. According to the results of the study, it was found that the phonetic features of the Internet vocabulary reflect the natural development and adaptability of the Kazakh language.
This paper explores the integration of sound files into wordnets, transforming them from static lexical databases into multimodal tools for linguistics, language learning and maintenance.Traditionally, wordnets focus on textual representations.Adding sound improves usability for language learners and linguists, especially in less-documented or endangered languages.We extracted sound data for basic vocabulary in 24 languages from the TUFS Basic Vocabulary Modules, link them to senses and make them available as small wordnets.We also discuss the issues involved with merging the data into an existing wordnet, looking at the Open English Wordnet.In addition, this paper outlines the process of integrating audio, discusses potential use cases, and evaluates the technical challenges involved.Finally we suggest an extension to the wordnet formats to allow sound for examples and definitions as well.
The article examines the word-formation competence of artificial intelligence in translating German feminatives into Ukrainian, with a focus on competing suffix models and compliance with contemporary linguistic norms. The object of the study is feminatives as a component of the Ukrainian word-formation system, while the subject is the translation strategies used by different versions of the ChatGPT language model. The research material includes German official texts containing feminatives denoting professions, positions, and social status. The methodology is based on contrastive and quantitative analysis of German feminatives with the suffix -in/-innen and their Ukrainian equivalents generated by three ChatGPT versions (GPT-4.0, GPT-5.1, and GPT-5.2), correlated with data from the General Ukrainian Language Corpus (HRAK). The results reveal clear differences in word-formation strategies. GPT-5.2 shows the highest sensitivity to modern Ukrainian tendencies, while GPT-5.1 demonstrates a transitional pattern and GPT-4.0 prefers traditional forms. The findings confirm the increasing sensitivity of language models to word-formation variability and corpus-based data.
Kyrgyz, a Turkic language with over 4.4 million speakers concentrated primarily in Kyrgyzstan and adjacent regions of Central Asia, faces a significant disparity in computational linguistic resources compared to languages with similar or even smaller speaker populations. Despite its status as a government language and cultural cornerstone, Kyrgyz remains underrepresented in the digital linguistic landscape. This investigation examines the application of the Universal Dependencies (UD) framework – an annotation system engineered to facilitate cross-linguistic syntactic comparability – to the structural complexities of Kyrgyz. We endeavor to identify optimal annotation strategies that faithfully represent Kyrgyz-specific syntactic phenomena while adhering to the principled constraints of the UD paradigm. The establishment of standardized syntactic resources for Kyrgyz carries dual significance: it advances linguistic typology by incorporating data from an underrepresented language family, while simultaneously laying groundwork for practical natural language processing applications crucial for Kyrgyz speakers’ participation in the digital sphere. Our methodological approach encompasses rigorous analysis of nascent Kyrgyz treebanks, comparative evaluation of annotation strategies employed for genetically related Turkic languages, and systematic examination of four fundamental annotation challenges: the representation of Kyrgyz’s defective copula system, the classification of multifunctional grammatical particles, the annotation of constructions with implicit heads, and the demarcation between inflectional and derivational morphology in this highly agglutinative language. Our analysis reveals that achieving the dual objectives of linguistic fidelity and cross-linguistic consistency necessitates judicious adaptation of UD guidelines to accommodate Kyrgyz-specific structures. We advance unified annotation solutions that preserve the integrity of Kyrgyz linguistic patterns while facilitating meaningful cross-linguistic comparison. This research not only contributes substantively to computational resources for Kyrgyz but also establishes annotation principles with broader applicability to typologically similar agglutinative languages. The practical implications extend to enhanced guidelines for Kyrgyz treebank development, which will consequently improve parser accuracy and catalyze the development of essential language technology tools for Kyrgyz speakers.
Poor relationship quality common among individuals with borderline personality disorder (BPD) may result, in part, from biased interpersonal decision-making. We examined memory biases for hypothetical interpersonal partner choices varying in the degree of familiarity. In Part 1 of our study, participants (n = 192) were asked to choose between novel or familiar partners based on lists of traits across six vignettes, and in Part 2, they completed a trait recognition task 36-60 hours later. Lower perceived social support was associated with a memory bias toward novel (over familiar) partners. BPD features were negatively related to an overall interpersonal memory bias (i.e., remembering both partners more negatively). However, when accounting for idiographic valence ratings, BPD features were positively related to this bias among those also low in social support. Memory biases may be related to partner choices associated with BPD features; however, it is critical to assess the role of perceived social support.
This paper explores the internal logic of ChatGPT, a generative large language model, to understand how it “thinks” before producing a response. By examining the model's architecture, linguistic behavior, and epistemic limitations, the paper reveals how AI simulates thought without engaging in reflection, intention, or creativity. While the system excels at producing fluent and plausible language, it does so through reconsolidation and probabilistic patterning rather than cognitive depth. As AI-generated content becomes ubiquitous in education, media, and research, humans are increasingly tempted to outsource imaginative and interpretive labor to machines. This paper argues that the result is a subtle erosion of originality, creativity, and critical thinking-driven not by overt misuse, but by the normalization of fluent but hollow language. Through an analysis of how ChatGPT generates responses, mimics creativity, reinforces linguistic norms, and shapes user cognition, the paper calls for a renewed commitment to interpretive sovereignty, friction-based learning, and human-centered authorship in the age of AI. Understanding how AI “thinks” is not only a technical question, but an ethical and educational one-crucial for preserving the distinctiveness of human thought in an increasingly synthetic linguistic environment.
This research aims to analyze the Politeness and Speech Acts of the Community (Ojol Community). This research uses a qualitative approach with descriptive method. Data were obtained through direct observation and recording of conversations between online ojek drivers and customers in real situations. Recording is done naturally without intervention to reflect authentic speech acts. The audio data is then transcribed and analyzed using Searle's speech act theory. The analysis is done descriptively qualitative by classifying and interpreting the form and function of utterances in the context of the conversation. The results show that nonstandard language is more dominantly used in informal communication, such as conversations between online ojek drivers and passengers, because it is considered more familiar, relaxed, and efficient. However, mastery of standardized language is still important, especially in official contexts, to maintain clarity and politeness. People are expected to be able to adjust the use of language according to the context so that communication remains effective and in accordance with linguistic norms. Keywords:,,,,.
The study of linguistic variation within the administrative structures of small-town America reveals a complex intersection between language, social identity, and institutional behavior. When approaching the linguistic environment of these communities from a purely academic perspective, without relying on personal immersion narratives or experiential accounts, one must begin with the foundational premise that English in the United States is profoundly regionalized. This regionalization is not a superficial matter of accent or vocabulary; it is a system of deeply embedded linguistic norms that shape how communication occurs, how authority is interpreted, and how institutional legitimacy is constructed. My interest as a researcher lies not in documenting local flavor or collecting curiosities from rural life but in understanding the mechanisms by which language operates as a structural force within governance. This requires an examination of sociolinguistic corpora, regional dialect research, institutional discourse studies, and the extensive literature on American dialect geography that has accumulated since the mid-twentieth century.
Inflection of Belarusian surnames is a topic of nationwide importance in Belarus. To this day, there is a great diversity in reference literature, and the standardization of surname inflection has not been established. The article describes the declension variations of Belarusian surnames in the Belarusian language, influenced by Russian patterns that are incompatible with standard practices of Belarusian speakers. Discussion includes the formation of the Belarusian linguistic norm, which has experienced historical and continuing influence by extra-linguistic factors. Complex instances of declension of Belarusian surnames in both Belarusian and Polish as well as of Polish surnames in both Polish and Belarusian are thoroughly discussed. On the one hand, comparative analysis shows the subtle similarities and differences between the two Slavic languages, from which interference errors can result. On the other hand, there are noticeable tendencies for 1) Belarusian surnames to decline according to the Russian models, as well as 2) Belarusians living in Poland to express preferences to decline Belarusian surnames used in the Polish language, according to the rules of Belarusian language.
In response to widespread leadership crises that favor popularity over competence, undermining strategic decision‑making and ethical standards, this study articulates the Prophet Muhammad's (PBUH) communication and diplomatic principles in his royal correspondence as a leadership archetype grounded in justice, meritocracy, meticulous composition, and mutual respect. Employing a qualitative literature approach with historical content and comparative analysis, it examines primary manuscripts of prophetic letters and their contexts alongside classical sirah texts and peer-reviewed studies. The analysis uncovers four pivotal elements: concentrated da'wah message summaries; the "Muhammad Rasul Allah" seal for authentication; envoy selection tailored to each court's linguistic norms; and Qur'anic citations for spiritual authority. These elements demonstrate a synergistic blend of prophetic legitimacy and diplomatic courtesy, offering a framework for religious rhetoric and ethical leadership development. The study also recommends rigorous comparative diplomacy across global traditions and innovative, strategic interdisciplinary collaboration for future scholarly inquiry.
BACKGROUND AND OBJECTIVES: The Seeking Proxies for Internal States (SPIS) model of OCD posits that reduced access to internal states plays a key role in the development and maintenance of the disorder. The current work sought to provide further support for the model's central claim that obsessive-compulsive tendencies are associated with reduced access to internal states. METHOD: Participants (N = 170) listened to 60 sound stimuli, rated how each one made them feel, and completed a measure of obsessive-compulsive tendencies. Following past procedure, we compared participants' ratings to each sound's normative valence rating, such that higher deviations between the ratings reflect a noisier perception of affective internal states. RESULTS: As hypothesized, higher obsessive-compulsive tendencies predicted greater deviations for both normatively-positive and normatively-negative sounds. CONCLUSIONS: The current work provides additional, novel support for the SPIS model, showing that with increasing obsessive-compulsive tendencies, people exhibited reduced attunement to how auditory stimuli made them feel.
Colour is a fundamental determinant of affective experience in immersive virtual reality (VR), yet the emotional and physiological impact of individual hues remains poorly characterised. This study investigated how fifteen calibrated Munsell hues influence subjective and autonomic responses when presented in immersive VR. Thirty-six adults (18–45 years) viewed each hue in a within-subject design while pupil diameter and skin conductance were recorded continuously, and self-reported emotions were assessed using the Self-Assessment Manikin across pleasure, arousal, and dominance. Repeated-measures ANOVAs revealed robust hue effects on all three self-report dimensions and on pupil dilation, with medium-to-large effect sizes. Reds and red–purple hues elicited the highest arousal and dominance, whereas blue–green hues were rated most pleasurable. Pupil dilation closely tracked arousal ratings, while skin conductance showed no reliable hue differentiation, likely due to the brief (30 s) exposures. Individual differences in cognitive style and personality modulated overall reactivity but did not alter the relative ranking of hues. Taken together, these findings provide the first systematic hue-by-hue mapping of affective and physiological responses in immersive VR. They demonstrate that calibrated colour shapes both experience and ocular physiology, while also offering practical guidance for educational, clinical, and interface design in virtual environments.
As it is well known, sociolinguistics is based on the idea that language is a social institution and an acquired phenomenon, established by members of society through mutual agreement to fulfill their needs and desires. Ibn Jinni defines language as "sounds through which each group expresses its purposes." One of the main reasons for the emergence of Arabic grammar was the social motivation of pride in the Arabic language and the need to preserve it from those who entered Islam, as they required learning Arabic to study and memorize the Qur’an. Since this interaction had an impact on the language, I found it necessary to explore this interdisciplinary sociolinguistic approach in the phenomenon of lahn (linguistic errors). The study is structured into an introduction, two main sections, a conclusion, and a bibliography. The introduction discusses the interdisciplinary study within sociolinguistics. The first section examines the impact of societies during the Islamic conquests on the Arabic language. The second section addresses the dangers of lahn and its effects on Arabic linguistic norms. The conclusion summarizes the key findings of the study.
In ”Bartleby, the Scrivener,” Herman Melville presents a character whose passive refusal, encapsulated in the repeated phrase “I would prefer not to,” challenges power, agency, and social norms. This essay examines how Bartleby’s refrain acts as both an assertion of autonomy and a critique of the violence inherent in language. By rejecting his employer’s commands, Bartleby disrupts the rational, efficiency-driven logic of the workplace, exposing the violence embedded in linguistic norms. Slavoj Žižek’s concept of language as inherently violent—through its imposition of norms and standards— illuminates how Bartleby’s refusal goes beyond protest, creating a space of resistance that defies interpretation and subverts power dynamics. Bartleby’s language, neither a clear denial nor an expression of desire, becomes a radical negation that questions the very nature of meaning. Ultimately, Bartleby’s refusal does not propose a new order but disrupts the structures of meaning and authority, forcing us to confront the limits of language itself.
Abstract This paper argues that the study of binary personal pronouns needs to move beyond European languages and the focus on third-person pronouns, and it supports this argument by presenting the problem of the Japanese linguistic norm of binarily gendered first-person pronouns. The Japanese case illustrates five ways in which this field of study should expand. First, binary pronoun studies cannot continue to neglect first-person pronouns. Second, in addition to pronoun systems, the norms of usage must be considered. Third, Japanese requires different forms of resolution for the problem of binary pronouns. Fourth, research should explore linguistic self- and other-referring expressions by examining the distinct functions of first- versus third-person pronouns. Finally, the Japanese case demonstrates that the parties directly affected by binarily gendered pronouns include girls and women as well as gender-nonconforming and trans people. Considering the Japanese case thus potentially contributes to expanding the field in five important ways.
This article analyzes linguistic issues encountered in the translation of Uzbek auxiliary verbs into English. In the Uzbek language, auxiliary verbs are used to express various grammatical meanings, including the continuity or completion of an action, the manner of execution, and the speaker’s attitude. In English, these meanings are conveyed through different grammatical constructions and independent verbs. Within the scope of this study, approximately 29 auxiliary verbs were identified to express around 70 grammatical meanings. The article examines their translation features and discusses the key methods for translating auxiliary verbs and the challenges in constructing a linguistic database.
Chinese word segmentation is a foundational task in natural language processing (NLP), with far-reaching effects on syntactic analysis. Unlike alphabetic languages like English, Chinese lacks explicit word boundaries, making segmentation both necessary and inherently ambiguous. This study highlights the intricate relationship between word segmentation and syntactic parsing, providing a clearer understanding of how different segmentation strategies shape dependency structures in Chinese. Focusing on the Chinese GSD treebank, we analyze multiple word boundary schemes, each reflecting distinct linguistic and computational assumptions, and examine how they influence the resulting syntactic structures. To support detailed comparison, we introduce an interactive web-based visualization tool that displays parsing outcomes across segmentation methods.
This article discusses the issues of forming a database of phraseological units based on the Uzbek language corpus. A linguistic database is a structured collection of information that stores words, phrases, grammatical forms, idioms and other linguistic elements, and is used in lexicography to create dictionaries and reference books, in natural language processing to support technologies related to automatic translation, speech recognition and text generation, in linguistic research to analyze the structure of the language, the change and use of language units, and in language education to create educational materials and tools for language learning.
While Large Language Models (LLMs) have shown remarkable performance in various Natural Language Processing (NLP) tasks, their effectiveness seems to be heavily biased toward high-resource languages.This proposal aims to address this gap by developing efficient training strategies for low-resource languages.We propose various techniques for efficient learning in simulated low-resource settings for English.We then plan to adapt these methods for lowresource languages.We plan to experiment with both natural language generation and understanding models.We evaluate the models on similar benchmarks as the BabyLM challenge for English.For other languages, we plan to use treebanks and translation techniques to create our own silver test set to evaluate the low-resource LMs.
The current study aimed to examine the effects of topic familiarity and language proficiency on linguistic complexity, accuracy, and fluency of argumentative essays in an EFL context. This study involved 64 college freshmen, who were divided into two groups according to a TOEIC score of 700, which corresponds to the B1 level on the CEFR scale: a high group (n = 31) and an intermediate group (n = 33). Participants were asked to write two argumentative essays on different topics, controlling for order effects through counterbalancing: one familiar (driving) and the other unfamiliar (smoking). They also completed a questionnaire that included background information and a topic familiarity rating on a 10-point Likert scale. The participants’ writing samples were analyzed in terms of lexical complexity, syntactic complexity, accuracy, and fluency. The results indicated that, in terms of topic familiarity, EFL learners tended to produce texts with lower levels of lexical and syntactic complexity, as well as reduced accuracy, when writing about unfamiliar topics. With regard to language proficiency, advanced learners demonstrated a broader vocabulary range, employed longer and more complex sentence structures, and produced more accurate and extensive texts compared to their intermediate-level peers. In-depth analysis and pedagogical implications are discussed.
We investigate the application of Neural Quantum Embedding and Quantum Neural Networks for sentiment analysis using the Stanford Sentiment Treebank dataset. We adapt the Neural Quantum Embedding framework, originally worked with image classification, to textual data by employing quantum embedding techniques that maximize trace distance for better data separability. Additionally, we incorporate advanced quantum feature maps and preprocessing techniques from recent quantum machine learning studies to enhance classification performance. Our experimental results demonstrate that the quantum models catch up with the classical baselines in sentiment classification accuracy, particularly in noisy intermediate-scale quantum settings. This work highlights the feasibility and potential of quantum-assisted sentiment analysis.
We present a family of encodings for sequence labeling dependency parsing, based on the concept of hierarchical bracketing. We prove that the existing 4-bit projective encoding belongs to this family, but it is suboptimal in the number of labels used to encode a tree. We derive an optimal hierarchical bracketing, which minimizes the number of symbols used and encodes projective trees using only 12 distinct labels (vs. 16 for the 4-bit encoding). We also extend optimal hierarchical bracketing to support arbitrary non-projectivity in a more compact way than previous encodings. Our new encodings yield competitive accuracy on a diverse set of treebanks.
This research paper examines the sociological significance of dialects and accents, analyzing their role in shaping social identities, reinforcing hierarchies, and influencing systemic biases. Grounded in sociological theories, particularly those of Pierre Bourdieu, Erving Goffman, Max Weber, and Michel Foucault, the study explores how language functions as a form of symbolic capital that dictates access to social mobility and power. Through a critical analysis of language as a site of inclusion and exclusion, the paper highlights how dominant linguistic norms marginalize non-standard dialects, perpetuating social stratification. Additionally, the study investigates the role of media, globalization, and cultural representation in shaping linguistic perceptions and maintaining or challenging linguistic hegemony. While dialects and accents often serve as markers of discrimination, they are also powerful tools for cultural identity and resistance. This paper underscoresthe need for greater linguistic inclusivity in institutional, educational, and social contextsto combat entrenched biases and promote equitable linguistic representation.
People's language and behavior can be greatly influenced by gender differences.Female communication patterns have historically been characterized by tentative expressions and men use nonstandard language and more slang than women.However, with the development of society and the changes of cultural values, gender language develops and the language differences are not static.Women are asked to behave politely and humbly in the past but now they show more confidence.Gender differences in language manifest in many aspects, and literary works are no exception.In Little Women, Louisa May Alcott challenged traditional gender-based linguistic norms to present an independent and self-reliant girl Jo March who is fond of using slang, acts like a boy and treats herself as a boy.Jo uses masculine language and exhibits more male characteristics in communication.Jo 's unique character also reflects the profound exploration of female independence and freedom, and her thoughts and actions become a model for women to pursue self-worth and independent life.
Salient sexual cues (erect penis, attractive individuals) are thought to capture initial attention and automatically trigger genital arousal in women. Conscious appraisal activates subjective sexual arousal and further visual attention. The present study tested whether the attractiveness category (attractive/unattractive) and/or sexual arousal condition (in underwear, naked with flaccid penis, naked with erect penis) of male stimuli predicted female sexual responding and attentional patterns. Genital arousal, visual attention, and subjective ratings (subjective sexual arousal, pleasantness) of 26 predominantly heterosexual women (Mage = 31.2, SDage = 6.8) were measured while exposed to male stimuli across the experimental conditions. Neither attractiveness nor sexual arousal condition of male models significantly predicted genital arousal, subjective sexual arousal ratings or visual attention patterns of women. Results however showed high concordance between genital and subjective sexual arousal measures, and the pleasantness rating of the stimuli positively predicted subjective sexual arousal, suggesting a positive feedback loop of female sexual response.
Kyrgyz remains a low-resource language with limited foundational NLP tools. To address this gap, we introduce KyrgyzBERT, the first publicly available monolingual BERT-based language model for Kyrgyz. The model has 35.9M parameters and uses a custom tokenizer designed for the language's morphological structure. To evaluate performance, we create kyrgyz-sst2, a sentiment analysis benchmark built by translating the Stanford Sentiment Treebank and manually annotating the full test set. KyrgyzBERT fine-tuned on this dataset achieves an F1-score of 0.8280, competitive with a fine-tuned mBERT model five times larger. All models, data, and code are released to support future research in Kyrgyz NLP.
Kyrgyz remains a low-resource language with limited foundational NLP tools. To address this gap, we introduce KyrgyzBERT, the first publicly available monolingual BERT-based language model for Kyrgyz. The model has 35.9M parameters and uses a custom tokenizer designed for the language's morphological structure. To evaluate performance, we create kyrgyz-sst2, a sentiment analysis benchmark built by translating the Stanford Sentiment Treebank and manually annotating the full test set. KyrgyzBERT fine-tuned on this dataset achieves an F1-score of 0.8280, competitive with a fine-tuned mBERT model five times larger. All models, data, and code are released to support future research in Kyrgyz NLP.
Abstract Combining research in developmental sociolinguistics and L1 acquisition, this study explores how caregivers may orient children towards (socio)linguistic norms through parental feedback. Based on self-recorded family interactions in the Belgian-Dutch setting, it applies a top-down quantitative perspective to examine feedback on non-conventional versus non-standard language use, alongside a bottom-up qualitative perspective highlighting factors that influence parental feedback occurrence. Findings reveal limited feedback on children’s non-standard language use, with participation frameworks and multiactivity contexts emerging as possible constraints. The combined approach also foregrounds possible tensions between researcher categorisations and participants’ perspectives. Overall, this study offers a first step in bridging research on parental feedback and sociolinguistic variation, identifying patterns that merit further investigation.
This chapter synthesises the findings and discusses how sociomaterial processes shape languages. Challenging modernist linguistic paradigms, it examines how language categories emerge through diverse cultural, historical, and material practices. The chapter critiques binary linguistic models and universalist, teleological assumptions of standardisation, showing that stable linguistic systems are not ‘natural’, but result from specific sociopolitical and material conditions. In contrast, fluid linguistic practices in postcolonial and globalised contexts exhibit variability, innovation, and complex indexicality. Belize’s multilingual environment exemplifies a setting without a hegemonic linguistic centre, producing liquid linguistic norms. The chapter argues for decolonial approaches to linguistics that embrace heterogeneity and that challenge exclusionary, Eurocentric models. Ultimately, it positions fluid linguistic practices as a cultural avant-garde and understands postcolonial environments as inspiring insights into future global sociolinguistic orders shaped by digitalisation and transnationalism.
Analysis code and data for "Asymmetric admixture decouples gene–language coevolution in Eastern Eurasia" This repository contains all computational code and the TyDEE (Typological Dataset of Eastern Eurasia) linguistic database used in our study examining gene-language relationships across Eastern Eurasia. The dataset comprises 541 language varieties across 10 major language families, paired with genome-wide genetic data from 135 populations.
This essay reimagines public speaking education through a culturally sustaining and transgressive lens that challenges dominant norms of language, professionalism, and communication competence. It critiques the ways in which public speaking courses often reinforce linguistic supremacy by privileging standardized English and marginalizing multilingual and culturally grounded speech practices. Drawing on concepts such as translanguaging, Culturally Sustaining Pedagogy (CSP), and transgressive pedagogies, the author calls for a shift in pedagogy that centers students’ lived experiences, community-rooted knowledge, and linguistic norms. Rather than asking students to conform to hegemonic standards, this approach empowers them to speak on their own terms, resist assimilationist pressures, and use language as a tool for identity, resistance, and liberation. By transforming the public speaking classroom into a space for critical reflection and empowerment, educators can cultivate more inclusive and equitable models of communication instruction.
French is often celebrated for its clarity and precision – a legacy shaped by Cartesian rationalism and prescriptive language policies. However, the evolving forms of spoken French challenge this ideal of fixed linguistic norms. This study examines one such feature: the right-peripheral duplication of the subject pronoun je with its tonic counterpart moi, a recurrent but underexplored phenomenon in spoken French. The primary objective is to understand how this syntactic feature functions pragmatically and emotionally in real-life discourse. Using a corpus of movie dialogues, the analysis shows that duplication plays a role in managing conversational flow, expressing personal stance, and enabling self-repair. Through a multidisciplinary lens that draws from sociolinguistics, pragmatics, and applied linguistics, the study argues that such variation enriches the expressive potential of French and complicates the rigid divide between written norms and spoken practice. It also suggests that incorporating these features into language pedagogy can support a more inclusive, realistic understanding of French as a living language.
We investigate the performance of state-ofthe-art (SotA) neural grammar induction (GI) models on a morphemically tokenised English dataset based on the CHILDES treebank (Pearl and Sprouse, 2013).Using implementations from Yang et al. (2021b), we train models and evaluate them with the standard F1 score.We introduce novel evaluation metrics-depth-ofmorpheme and sibling-of-morpheme-which measure phenomena around bound morpheme attachment.Our results reveal that models with the highest F1 scores do not necessarily induce linguistically plausible structures for bound morpheme attachment, highlighting a key challenge for cognitively plausible GI.
Human–human interaction studies have shown that live performances of dynamic emotional facial expressions, compared to pre-recorded videos, enhance emotion contagion and spontaneous facial mimicry. While robotic emotional facial expressions can also induce emotion contagion and facial mimicry, the statistical significance of the live presence effect has not been demonstrated. This study utilized a live image relay system to deliver real-time performances of positive (smiling) and negative (frowning) facial expressions by the android Nikola to participants, alongside prerecorded video presentations. Subjective valence and arousal ratings were collected, along with facial electromyography (EMG) from the corrugator supercilii and zygomaticus major muscles. Results indicated that live negative facial expressions elicited lower valence and higher arousal compared to their video counterparts. Facial EMG revealed that live facial expressions induced greater congruent facial muscular activity than pre-recorded videos. These findings suggest that the robotic live presence may enhance affective engagement in socially interactive contexts.
This study examines the use of the Low variety of Arabic, commonly known as colloquial or spoken Arabic, in email communications among Saudi university youth, specifically in their correspondence with academic affairs unit. Drawing on sociolinguistic frameworks of diglossia, this project investigates the extent to which colloquial Arabic is employed and the underlying factors influencing this usage. Through a quantitative analysis of 100 email samples and qualitative analysis of 4 focus group discussions with Arabic language instructors, the findings of this study indicated a noticeable shift toward the incorporation of colloquial Arabic in academic email communication among youth, signaling broader transformations in linguistic norms influenced by technological advancements, generational attitudes and educational factors. This phenomenon underscores the need to adapt language education, institutional guidelines, and cultural expectations to align with these changes, ensuring a balance between linguistic evolution and the preservation of traditional standards.
The development of social media as a public communication space has had a significant impact on language use, particularly grammar. This article discusses the use of grammar on social media, which is often caused by the freedom of expression not being balanced with an awareness of proper and correct language. Through a qualitative approach with discourse analysis, this article highlights how deviations from grammatical rules in social media posts reflect a shift in societal attitudes towards linguistic norms. On one hand, linguistic freedom on social media is considered a form of creativity and self-expression. However, on the other hand, it has the potential to weaken language skills that adhere to rules, especially among the younger generation. This article will certainly provide a recommendation on the importance of language literacy as an effort to maintain the quality of public communication amidst the wave of digitalization and the expanding freedom of expression.