Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The increasing use of English in many professional and academic contexts has played a pivotal role not only in expanding the teaching of English at universities worldwide but also in determining the growth of teaching in English. This has heightened the need to teach thesis writing in many English-medium contexts. Focusing on the tension between individuality – expressing an individual point of view – and commonality – adopting the rhetorical and linguistic norms of the specific discourse community addressed, we discuss our experience of designing an introduction to thesis writing for an EMI (English as a Medium of Instruction) multidisciplinary MA programme in Italy. The chapter begins by exploring the interplay between commonality and individuality, then proceeds to describe the course’s context and its educational underpinnings. The general principles are illustrated through an analysis of activities meant to teach the students how to engage with the discourse community of their choice and to express their stance: from engaging with discourse communities to finding relevant sources, incorporating other voices into a literature review, and working collaboratively on the conclusion section. The chapter closes with a brief summing up of the issues raised.
This paper explores the integration of sound files into wordnets, transforming them from static lexical databases into multimodal tools for linguistics, language learning and maintenance.Traditionally, wordnets focus on textual representations.Adding sound improves usability for language learners and linguists, especially in less-documented or endangered languages.We extracted sound data for basic vocabulary in 24 languages from the TUFS Basic Vocabulary Modules, link them to senses and make them available as small wordnets.We also discuss the issues involved with merging the data into an existing wordnet, looking at the Open English Wordnet.In addition, this paper outlines the process of integrating audio, discusses potential use cases, and evaluates the technical challenges involved.Finally we suggest an extension to the wordnet formats to allow sound for examples and definitions as well.
Uncertainty of scientific findings are typically reported through statistical metrics such as $p$-values, confidence intervals, etc. The magnitude of this objective uncertainty is reflected in the language used by the authors to report their findings primarily through expressions carrying uncertainty-inducing terms or phrases. This language uncertainty is a subjective concept and is highly dependent on the writing style of the authors. There is evidence that such subjective uncertainty influences the impact of science on public audience. In this work, we turned our focus to scientists themselves, and measured/analyzed the subjective uncertainty and its impact within scientific communities across different disciplines. We showed that the level of this type of uncertainty varies significantly across different fields, years of publication and geographical locations. We also studied the correlation between subjective uncertainty and several bibliographical metrics, such as number/gender of authors, centrality of the field's community, citation count, etc. The underlying patterns identified in this work are useful in identification and documentation of linguistic norms in scientific communication in different communities/societies.
The article systematizes and comprehensively analyzes the phenomena of the phonetic level of the Internet vocabulary of the modern Kazakh language. Instagram, Facebook, social networks (Threads, Instagram, Facebook), and instant messengers (WhatsApp, Telegram), which have been actively used in recent years, were chosen as the object of the study. The research used methods of observation, generalization, comparative and descriptive analysis. As a result, it is revealed that new forms of linguistic usage are being formed in the Internet space, characterized by a mixture of elements of spoken and written speech. At the phonetic level, phenomena such as sound compression of words and, conversely, the repetition of graphemes to convey emotions in writing are widespread. The active use of Latin graphics and the development of foreign-language sounds indicate a new stage of phonetic adaptation in the Kazakh-speaking Internet space. The article provides specific examples of these phenomena, reveals their causes and impact on the modern linguistic norm and writing culture. According to the results of the study, it was found that the phonetic features of the Internet vocabulary reflect the natural development and adaptability of the Kazakh language.
This study investigates differences in artificial intelligence (AI) literacy and adoption between engineering students and faculty in a Middle Eastern higher-education institution. Parallel surveys were administered to undergraduate engineering students (N = 73) and faculty members (N = 20), each rating their familiarity with 20 AI tools covering learning, coding, productivity, and engineering applications. An AI Literacy Index was computed by assigning numerical values to familiarity ratings (A = 2, B = 1, C = 0) and normalizing the total to a 0–1 scale. Results from Welch’s t-test indicated that students demonstrated significantly higher literacy than faculty (0.454 vs. 0.356, p ≈ 0.042). Students also reported strong AI adoption for academic tasks (71.2%) and high perceived learning benefits (83.6%). Conversely, faculty expressed substantial concern about student over-reliance on AI (90%) while indicating readiness for professional development through AI training workshops (75%) and reporting assessment redesign efforts (75%). Overall, the findings highlight a meaningful literacy and perception gap with implications for engineering pedagogy, curriculum development, and assessment practices. Recommendations are provided to support the alignment of student and faculty AI competencies within engineering programs.
Kyrgyz, a Turkic language with over 4.4 million speakers concentrated primarily in Kyrgyzstan and adjacent regions of Central Asia, faces a significant disparity in computational linguistic resources compared to languages with similar or even smaller speaker populations. Despite its status as a government language and cultural cornerstone, Kyrgyz remains underrepresented in the digital linguistic landscape. This investigation examines the application of the Universal Dependencies (UD) framework – an annotation system engineered to facilitate cross-linguistic syntactic comparability – to the structural complexities of Kyrgyz. We endeavor to identify optimal annotation strategies that faithfully represent Kyrgyz-specific syntactic phenomena while adhering to the principled constraints of the UD paradigm. The establishment of standardized syntactic resources for Kyrgyz carries dual significance: it advances linguistic typology by incorporating data from an underrepresented language family, while simultaneously laying groundwork for practical natural language processing applications crucial for Kyrgyz speakers’ participation in the digital sphere. Our methodological approach encompasses rigorous analysis of nascent Kyrgyz treebanks, comparative evaluation of annotation strategies employed for genetically related Turkic languages, and systematic examination of four fundamental annotation challenges: the representation of Kyrgyz’s defective copula system, the classification of multifunctional grammatical particles, the annotation of constructions with implicit heads, and the demarcation between inflectional and derivational morphology in this highly agglutinative language. Our analysis reveals that achieving the dual objectives of linguistic fidelity and cross-linguistic consistency necessitates judicious adaptation of UD guidelines to accommodate Kyrgyz-specific structures. We advance unified annotation solutions that preserve the integrity of Kyrgyz linguistic patterns while facilitating meaningful cross-linguistic comparison. This research not only contributes substantively to computational resources for Kyrgyz but also establishes annotation principles with broader applicability to typologically similar agglutinative languages. The practical implications extend to enhanced guidelines for Kyrgyz treebank development, which will consequently improve parser accuracy and catalyze the development of essential language technology tools for Kyrgyz speakers.
Chinese word segmentation is a foundational task in natural language processing (NLP), with far-reaching effects on syntactic analysis. Unlike alphabetic languages like English, Chinese lacks explicit word boundaries, making segmentation both necessary and inherently ambiguous. This study highlights the intricate relationship between word segmentation and syntactic parsing, providing a clearer understanding of how different segmentation strategies shape dependency structures in Chinese. Focusing on the Chinese GSD treebank, we analyze multiple word boundary schemes, each reflecting distinct linguistic and computational assumptions, and examine how they influence the resulting syntactic structures. To support detailed comparison, we introduce an interactive web-based visualization tool that displays parsing outcomes across segmentation methods.
Inflection of Belarusian surnames is a topic of nationwide importance in Belarus. To this day, there is a great diversity in reference literature, and the standardization of surname inflection has not been established. The article describes the declension variations of Belarusian surnames in the Belarusian language, influenced by Russian patterns that are incompatible with standard practices of Belarusian speakers. Discussion includes the formation of the Belarusian linguistic norm, which has experienced historical and continuing influence by extra-linguistic factors. Complex instances of declension of Belarusian surnames in both Belarusian and Polish as well as of Polish surnames in both Polish and Belarusian are thoroughly discussed. On the one hand, comparative analysis shows the subtle similarities and differences between the two Slavic languages, from which interference errors can result. On the other hand, there are noticeable tendencies for 1) Belarusian surnames to decline according to the Russian models, as well as 2) Belarusians living in Poland to express preferences to decline Belarusian surnames used in the Polish language, according to the rules of Belarusian language.
The study of linguistic variation within the administrative structures of small-town America reveals a complex intersection between language, social identity, and institutional behavior. When approaching the linguistic environment of these communities from a purely academic perspective, without relying on personal immersion narratives or experiential accounts, one must begin with the foundational premise that English in the United States is profoundly regionalized. This regionalization is not a superficial matter of accent or vocabulary; it is a system of deeply embedded linguistic norms that shape how communication occurs, how authority is interpreted, and how institutional legitimacy is constructed. My interest as a researcher lies not in documenting local flavor or collecting curiosities from rural life but in understanding the mechanisms by which language operates as a structural force within governance. This requires an examination of sociolinguistic corpora, regional dialect research, institutional discourse studies, and the extensive literature on American dialect geography that has accumulated since the mid-twentieth century.
This research aims to analyze the Politeness and Speech Acts of the Community (Ojol Community). This research uses a qualitative approach with descriptive method. Data were obtained through direct observation and recording of conversations between online ojek drivers and customers in real situations. Recording is done naturally without intervention to reflect authentic speech acts. The audio data is then transcribed and analyzed using Searle's speech act theory. The analysis is done descriptively qualitative by classifying and interpreting the form and function of utterances in the context of the conversation. The results show that nonstandard language is more dominantly used in informal communication, such as conversations between online ojek drivers and passengers, because it is considered more familiar, relaxed, and efficient. However, mastery of standardized language is still important, especially in official contexts, to maintain clarity and politeness. People are expected to be able to adjust the use of language according to the context so that communication remains effective and in accordance with linguistic norms. Keywords:,,,,.
In response to widespread leadership crises that favor popularity over competence, undermining strategic decision‑making and ethical standards, this study articulates the Prophet Muhammad's (PBUH) communication and diplomatic principles in his royal correspondence as a leadership archetype grounded in justice, meritocracy, meticulous composition, and mutual respect. Employing a qualitative literature approach with historical content and comparative analysis, it examines primary manuscripts of prophetic letters and their contexts alongside classical sirah texts and peer-reviewed studies. The analysis uncovers four pivotal elements: concentrated da'wah message summaries; the "Muhammad Rasul Allah" seal for authentication; envoy selection tailored to each court's linguistic norms; and Qur'anic citations for spiritual authority. These elements demonstrate a synergistic blend of prophetic legitimacy and diplomatic courtesy, offering a framework for religious rhetoric and ethical leadership development. The study also recommends rigorous comparative diplomacy across global traditions and innovative, strategic interdisciplinary collaboration for future scholarly inquiry.
In ”Bartleby, the Scrivener,” Herman Melville presents a character whose passive refusal, encapsulated in the repeated phrase “I would prefer not to,” challenges power, agency, and social norms. This essay examines how Bartleby’s refrain acts as both an assertion of autonomy and a critique of the violence inherent in language. By rejecting his employer’s commands, Bartleby disrupts the rational, efficiency-driven logic of the workplace, exposing the violence embedded in linguistic norms. Slavoj Žižek’s concept of language as inherently violent—through its imposition of norms and standards— illuminates how Bartleby’s refusal goes beyond protest, creating a space of resistance that defies interpretation and subverts power dynamics. Bartleby’s language, neither a clear denial nor an expression of desire, becomes a radical negation that questions the very nature of meaning. Ultimately, Bartleby’s refusal does not propose a new order but disrupts the structures of meaning and authority, forcing us to confront the limits of language itself.
As it is well known, sociolinguistics is based on the idea that language is a social institution and an acquired phenomenon, established by members of society through mutual agreement to fulfill their needs and desires. Ibn Jinni defines language as "sounds through which each group expresses its purposes." One of the main reasons for the emergence of Arabic grammar was the social motivation of pride in the Arabic language and the need to preserve it from those who entered Islam, as they required learning Arabic to study and memorize the Qur’an. Since this interaction had an impact on the language, I found it necessary to explore this interdisciplinary sociolinguistic approach in the phenomenon of lahn (linguistic errors). The study is structured into an introduction, two main sections, a conclusion, and a bibliography. The introduction discusses the interdisciplinary study within sociolinguistics. The first section examines the impact of societies during the Islamic conquests on the Arabic language. The second section addresses the dangers of lahn and its effects on Arabic linguistic norms. The conclusion summarizes the key findings of the study.
BACKGROUND AND OBJECTIVES: The Seeking Proxies for Internal States (SPIS) model of OCD posits that reduced access to internal states plays a key role in the development and maintenance of the disorder. The current work sought to provide further support for the model's central claim that obsessive-compulsive tendencies are associated with reduced access to internal states. METHOD: Participants (N = 170) listened to 60 sound stimuli, rated how each one made them feel, and completed a measure of obsessive-compulsive tendencies. Following past procedure, we compared participants' ratings to each sound's normative valence rating, such that higher deviations between the ratings reflect a noisier perception of affective internal states. RESULTS: As hypothesized, higher obsessive-compulsive tendencies predicted greater deviations for both normatively-positive and normatively-negative sounds. CONCLUSIONS: The current work provides additional, novel support for the SPIS model, showing that with increasing obsessive-compulsive tendencies, people exhibited reduced attunement to how auditory stimuli made them feel.
The incorporation of real-world contexts into mathematical problems has become increasingly significant on a global scale, which is also reflected in Vietnamese education reform. This emphasis has resulted in the development of different series of textbooks designed to enhance student engagement and foster active learning through real-world problem contexts. In particular, the geometry and measurement strand presents substantial opportunities for the integration of contexts. In this study, we investigate the theoretical foundations of contexts, context familiarity, and context authenticity in mathematics education, with a focus on the plane geometry content in ninth-grade Vietnamese textbooks. The analysis is supported by four illustrative examples drawn from three reformed textbook series, demonstrating how context familiarity and authenticity are represented. A 5-point Likert-type survey of 109 students (72 males and 37 females) revealed that the soccer context was rated as less familiar than the airplane context overall, with no significant difference in familiarity ratings between boys and girls. Both the bike-riding and tripod contexts were rated as moderately authentic but not highly so. The authors also acknowledge several limitations and offers certain recommendations for future research in the area of contextual mathematics education.
The current study aimed to examine the effects of topic familiarity and language proficiency on linguistic complexity, accuracy, and fluency of argumentative essays in an EFL context. This study involved 64 college freshmen, who were divided into two groups according to a TOEIC score of 700, which corresponds to the B1 level on the CEFR scale: a high group (n = 31) and an intermediate group (n = 33). Participants were asked to write two argumentative essays on different topics, controlling for order effects through counterbalancing: one familiar (driving) and the other unfamiliar (smoking). They also completed a questionnaire that included background information and a topic familiarity rating on a 10-point Likert scale. The participants’ writing samples were analyzed in terms of lexical complexity, syntactic complexity, accuracy, and fluency. The results indicated that, in terms of topic familiarity, EFL learners tended to produce texts with lower levels of lexical and syntactic complexity, as well as reduced accuracy, when writing about unfamiliar topics. With regard to language proficiency, advanced learners demonstrated a broader vocabulary range, employed longer and more complex sentence structures, and produced more accurate and extensive texts compared to their intermediate-level peers. In-depth analysis and pedagogical implications are discussed.
While Large Language Models (LLMs) have shown remarkable performance in various Natural Language Processing (NLP) tasks, their effectiveness seems to be heavily biased toward high-resource languages.This proposal aims to address this gap by developing efficient training strategies for low-resource languages.We propose various techniques for efficient learning in simulated low-resource settings for English.We then plan to adapt these methods for lowresource languages.We plan to experiment with both natural language generation and understanding models.We evaluate the models on similar benchmarks as the BabyLM challenge for English.For other languages, we plan to use treebanks and translation techniques to create our own silver test set to evaluate the low-resource LMs.
We investigate the application of Neural Quantum Embedding and Quantum Neural Networks for sentiment analysis using the Stanford Sentiment Treebank dataset. We adapt the Neural Quantum Embedding framework, originally worked with image classification, to textual data by employing quantum embedding techniques that maximize trace distance for better data separability. Additionally, we incorporate advanced quantum feature maps and preprocessing techniques from recent quantum machine learning studies to enhance classification performance. Our experimental results demonstrate that the quantum models catch up with the classical baselines in sentiment classification accuracy, particularly in noisy intermediate-scale quantum settings. This work highlights the feasibility and potential of quantum-assisted sentiment analysis.
Salient sexual cues (erect penis, attractive individuals) are thought to capture initial attention and automatically trigger genital arousal in women. Conscious appraisal activates subjective sexual arousal and further visual attention. The present study tested whether the attractiveness category (attractive/unattractive) and/or sexual arousal condition (in underwear, naked with flaccid penis, naked with erect penis) of male stimuli predicted female sexual responding and attentional patterns. Genital arousal, visual attention, and subjective ratings (subjective sexual arousal, pleasantness) of 26 predominantly heterosexual women (Mage = 31.2, SDage = 6.8) were measured while exposed to male stimuli across the experimental conditions. Neither attractiveness nor sexual arousal condition of male models significantly predicted genital arousal, subjective sexual arousal ratings or visual attention patterns of women. Results however showed high concordance between genital and subjective sexual arousal measures, and the pleasantness rating of the stimuli positively predicted subjective sexual arousal, suggesting a positive feedback loop of female sexual response.
People's language and behavior can be greatly influenced by gender differences.Female communication patterns have historically been characterized by tentative expressions and men use nonstandard language and more slang than women.However, with the development of society and the changes of cultural values, gender language develops and the language differences are not static.Women are asked to behave politely and humbly in the past but now they show more confidence.Gender differences in language manifest in many aspects, and literary works are no exception.In Little Women, Louisa May Alcott challenged traditional gender-based linguistic norms to present an independent and self-reliant girl Jo March who is fond of using slang, acts like a boy and treats herself as a boy.Jo uses masculine language and exhibits more male characteristics in communication.Jo 's unique character also reflects the profound exploration of female independence and freedom, and her thoughts and actions become a model for women to pursue self-worth and independent life.
Kyrgyz remains a low-resource language with limited foundational NLP tools. To address this gap, we introduce KyrgyzBERT, the first publicly available monolingual BERT-based language model for Kyrgyz. The model has 35.9M parameters and uses a custom tokenizer designed for the language's morphological structure. To evaluate performance, we create kyrgyz-sst2, a sentiment analysis benchmark built by translating the Stanford Sentiment Treebank and manually annotating the full test set. KyrgyzBERT fine-tuned on this dataset achieves an F1-score of 0.8280, competitive with a fine-tuned mBERT model five times larger. All models, data, and code are released to support future research in Kyrgyz NLP.
This essay reimagines public speaking education through a culturally sustaining and transgressive lens that challenges dominant norms of language, professionalism, and communication competence. It critiques the ways in which public speaking courses often reinforce linguistic supremacy by privileging standardized English and marginalizing multilingual and culturally grounded speech practices. Drawing on concepts such as translanguaging, Culturally Sustaining Pedagogy (CSP), and transgressive pedagogies, the author calls for a shift in pedagogy that centers students’ lived experiences, community-rooted knowledge, and linguistic norms. Rather than asking students to conform to hegemonic standards, this approach empowers them to speak on their own terms, resist assimilationist pressures, and use language as a tool for identity, resistance, and liberation. By transforming the public speaking classroom into a space for critical reflection and empowerment, educators can cultivate more inclusive and equitable models of communication instruction.
This chapter synthesises the findings and discusses how sociomaterial processes shape languages. Challenging modernist linguistic paradigms, it examines how language categories emerge through diverse cultural, historical, and material practices. The chapter critiques binary linguistic models and universalist, teleological assumptions of standardisation, showing that stable linguistic systems are not ‘natural’, but result from specific sociopolitical and material conditions. In contrast, fluid linguistic practices in postcolonial and globalised contexts exhibit variability, innovation, and complex indexicality. Belize’s multilingual environment exemplifies a setting without a hegemonic linguistic centre, producing liquid linguistic norms. The chapter argues for decolonial approaches to linguistics that embrace heterogeneity and that challenge exclusionary, Eurocentric models. Ultimately, it positions fluid linguistic practices as a cultural avant-garde and understands postcolonial environments as inspiring insights into future global sociolinguistic orders shaped by digitalisation and transnationalism.
The purpose of this study is to evaluate the English language proficiency of Sayed Jamaluddin Afghani University students enrolled in the English Department. It primarily looks at the grammatical mistakes that impede kids' language development and how they affect their ability to communicate. A structured questionnaire was used in a quantitative survey with a sample of 120 students in order to collect data. The results show that improper application of linguistic norms, infrequent practice, and the effect of students' native language, which frequently interferes with English usage, are the primary causes of grammatical errors. The efficacy and clarity of students' speech are greatly impacted by these errors. According to the study's findings, English professors should implement specialized courses that emphasize communicative grammar instruction. It also recommends giving students regular chances to practice active language usage, which can improve their overall communication skills and grammatical precision, ultimately increasing their English proficiency.
The development of social media as a public communication space has had a significant impact on language use, particularly grammar. This article discusses the use of grammar on social media, which is often caused by the freedom of expression not being balanced with an awareness of proper and correct language. Through a qualitative approach with discourse analysis, this article highlights how deviations from grammatical rules in social media posts reflect a shift in societal attitudes towards linguistic norms. On one hand, linguistic freedom on social media is considered a form of creativity and self-expression. However, on the other hand, it has the potential to weaken language skills that adhere to rules, especially among the younger generation. This article will certainly provide a recommendation on the importance of language literacy as an effort to maintain the quality of public communication amidst the wave of digitalization and the expanding freedom of expression.
This study examines the use of the Low variety of Arabic, commonly known as colloquial or spoken Arabic, in email communications among Saudi university youth, specifically in their correspondence with academic affairs unit. Drawing on sociolinguistic frameworks of diglossia, this project investigates the extent to which colloquial Arabic is employed and the underlying factors influencing this usage. Through a quantitative analysis of 100 email samples and qualitative analysis of 4 focus group discussions with Arabic language instructors, the findings of this study indicated a noticeable shift toward the incorporation of colloquial Arabic in academic email communication among youth, signaling broader transformations in linguistic norms influenced by technological advancements, generational attitudes and educational factors. This phenomenon underscores the need to adapt language education, institutional guidelines, and cultural expectations to align with these changes, ensuring a balance between linguistic evolution and the preservation of traditional standards.
The article examines the evolution of Judaeo-Arabic translations, ranging from early biblical renditions to modern adaptations of European literature. Early pre-Saadian translations, written in a phonetic transcription, preceded Saʿadya Gaon’s Tafsīr (Bible translation), which adhered to ‘classical’ linguistic norms and became the authoritative translation for centuries. However, later translations introduced local dialectal features to meet the needs of diverse Jewish communities. The theoretical framework of ‘centre versus periphery’ is employed to analyse the dynamics of translation traditions, highlighting the interaction between cultural centres like Meknes, Morocco, and Constantine, Algeria, and their peripheries. By the nineteenth and twentieth centuries, Judaeo-Arabic translations extended to Haskala novels and French and English classics such as Robinson Crusoe, demonstrating the influence of global cultural trends. The study emphasises the dual role of these translations in preserving Jewish identity and adapting to contemporary linguistic and cultural shifts.
This article discusses the issues of Named Entity Recognition (NER) and their interpretation in the Uzbek language corpus. Based on the morphological and syntactic rules of Uzbek language grammar, a lexical database (in Uzbek) has been developed, and a rule-based system has been formulated for the automatic identification of Named Entities (NERs) within texts. The article classifies named entities (such as personal names, place names, organization names, etc.) for NER models based on the Uzbek language corpus and describes the methodology used for their interpretation. Additionally, the process of identifying NER entities is presented through specialized diagrams, and the entities themselves are illustrated in tables.
The rapid growth of TikTok as a popular social media platform has significantly influenced written language practices, particularly in caption writing, which tends to be informal and spontaneous. This study aims to identify, classify, and analyze language errors found in TikTok captions created by Indonesian content creators. Employing a descriptive qualitative research design, the data were collected through documentation of publicly accessible TikTok captions and analyzed using qualitative content analysis. The findings reveal that language errors predominantly occur in the forms of redundancy, grammatical errors, spelling errors, and semantic inconsistency. These errors are not categorized as intentional code-switching or stylistic language mixing but rather as violations of standard linguistic norms. Theoretically, this study confirms that the informal nature of social media influences written language use and encourages non-standard forms in digital communication contexts.
Abstract Introduction Dream reports provide unique insights into emotional processing during sleep. While self-reports are the gold standard for assessing dream affect, natural language processing (NLP) tools like ChatGPT may offer scalable alternatives. This study evaluated ChatGPT’s ability to estimate positive and negative affect from dream reports, comparing its performance against self-reported ratings. Methods A total of 136 participants provided one dream report each. Participants rated their positive and negative dream affect on a 0–10 scale. ChatGPT 3.5 Turbo was accessed via the OpenAI API using Python, where each dream report was analyzed and rated on identical scales. The model’s outputs were programmatically appended to a CSV file for further analysis. Agreement between ChatGPT and self-reports was assessed using intra-class correlation coefficients (ICC3k) for consistency, mean absolute error (MAE) for deviation, and Bland-Altman plots for visual inspection of agreement. Results For positive affect ratings, ChatGPT demonstrated excellent agreement with self-reports (ICC3k = 0.844, 95% CI [0.781, 0.889], p <.001), with an MAE of 1.778. Negative affect ratings similarly showed excellent agreement (ICC3k = 0.857, 95% CI [0.799, 0.898], p <.001), with an MAE of 1.681. Bland-Altman plots for both affective dimensions indicated no systematic bias and acceptable limits of agreement upon visual inspection. Conclusion ChatGPT demonstrated strong agreement with self-reported dream affect ratings, supporting its potential as a scalable tool for analyzing emotional content in dream reports. These findings suggest that large language models can provide valid and reliable estimates of dream affect, which may advance sleep and affective science. Support (if any)
We present a family of encodings for sequence labeling dependency parsing, based on the concept of hierarchical bracketing. We prove that the existing 4-bit projective encoding belongs to this family, but it is suboptimal in the number of labels used to encode a tree. We derive an optimal hierarchical bracketing, which minimizes the number of symbols used and encodes projective trees using only 12 distinct labels (vs. 16 for the 4-bit encoding). We also extend optimal hierarchical bracketing to support arbitrary non-projectivity in a more compact way than previous encodings. Our new encodings yield competitive accuracy on a diverse set of treebanks.
Мақалада түркі тілдерінің синтаксистік құрылымын формалды грамматика тұрғысынан және заманауи аннотациялық модельдер негізінде сипаттаудың тәжірибесі қарастырылады. Синтаксистік аннотация тілдің грамматикалық жүйесін формалды түрде сипаттайтын және оны автоматты өңдеуге мүмкіндік беретін маңызды құрал ретінде танылады. Зерттеу барысында «Universal Dependencies» (UD), «MaTT» (Multilingual Aligned Treebank of Turkic) және «Kazakh Dependency Treebank» (KazDT) сияқты жобаларға сүйеніп, түркі тілдеріне тән морфологиялық және синтаксистік ерекшеліктер сипатталды. Синтаксистік белгіленім модельдері: «құрамдық», «аралас», «басыңқы-бағыныңқылық грамматикасы» т.б. тәсілдердің сипаты, ерекшеліктері, түркі тілдері үшін ұтымды тұстары мен кемшіліктері сараланды. Нәтижесінде басыңқы-бағыныңқы қатынастар грамматикасы негізінде жасалған синтаксистік аннотация моделі түркі тілінің құрылымын тиімді сипаттауға мүмкіндік беретіні дәлелденді. Басыңқы-бағыныңқы грамматикасының (басыңқы-бағыныңқы қатынастар) теориялық негіздері, синтаксистік аннотацияның форматы мен стандарттары сараланды. Түркі тілдерінің жалғамалы табиғаты мен еркін сөз тәртібінің «UD» сияқты әмбебап жобаларға бейімделуі талдауға түсті. Сонымен қатар, қазақ тілінің аннотацияланған корпустарын жетілдіру, автоматты парсинг, тілдік білім беру жүйесіне енгізу секілді болашақтағы бағыттары көрсетілді. Мақала түркі тілдерінің синтаксистік белгіленім тәжірибесі негізінде қазақ тілін цифрлық кеңістікке енгізудің маңызды қадамдарының бірі ретінде синтаксистік аннотацияны ғылыми тұрғыда негіздеуді мақсат етті. Түйін сөздер: түркі тілдері, синтаксистік аннотация, басыңқы-бағыныңқы грамматикасы, «UD», KazDT, формалды модельдер, парсинг.
This study describes the morpho-stylistic and the semantic-stylistic features used in the top four songs of Arctic Monkeys’ AM album, namely, “I Wanna Be Yours,” “Do I Wanna Know?”, “Why’d You Only Call Me When You’re High?” and “No. 1 Party Anthem.” This research conducted using a qualitative descriptive method derived from the linguistic deviation theory by Leech analyses of the lyrical texts for their morphological and semantic deviations like informal contractions, neologism, objectification, metaphors and irony. Taking language out of the box, the findings show that Arctic Monkeys have consistently broken linguistic norms in order to produce emotionality, stylistic nuance and lyrical uniqueness. These deviations greatly enhance the band’s lyrical identity of poetic and aesthetic qualities.
This paper addresses the syntactic annotation of the verbs ER and BOL in copular constructions in Turkic-language treebanks in Universal Dependencies. ER is a defective verb, which has semantically empty content linking subjects to non-verbal predicates. BOL is a full verb with copula and semi-copula uses. The paper proposes a refined annotation strategy that treats ER consistently as an auxiliary copula and BOL variably as either an auxiliary copula or a lexical verb, depending on its semantic contribution in context. The aim of this proposal is to improve the consistency and cross-linguistic comparability of Turkic language annotations in UD.
Characteristics of emotion, such as valence and arousal can be evaluated using self-reported affective ratings and electroencephalogram (EEG) to gain a better understanding of individual differences in diversified populations. The International Affective Picture System (IAPS) and AI-generated counterparts were used to elicit emotional responses that were collected from the self-assessment manakin (SAM) rating scale for valence and arousal. EEG data were used to observe biomarkers of emotional processes related to the presentation of these stimuli. These methods were correlated to individual differences such as sex and depression and anxiety related questionaries. The study showed significant sex differences in the self-reported affective ratings, where females showed greater aversive affect compared to males. EEG data showed that there was less alpha reduction relative to baseline in percent for individuals who scored high on the BDI-II, which suggests biomarkers of emotional dysregulation, such as anhedonia. Overall, the study highlights the differences in emotional processes for variable populations, which has implications for intervention and targeted treatment efforts.
The article is theoretical in nature. It focuses on the language awareness among young people. Language awareness enables code-switching and balancing between the sociolect known as the youthlanguage and the official, standard-compliant language. Depending on the context, language is usedby young people to identify themselves with the group and to manifest their independence, distinctiveness and uniqueness. It also reflects their integration into linguistic norms and their creative transformation of these norms. The level of language awareness determines which variety young people choose to use in particular situations. Drawing on the levels of language awareness identified in the literature, the author presents examples of both low and high levels of language awareness in thelinguistic competence of young people. The conclusion is presented within the context of the core curriculum, highlighting the need to correlate literary education with language education.
Bengali (Bangla) is typologically rich in complex predicates, especially verb-verb compounds and verbo-nominal light verb constructions. While these constructions have been studied from theoretical and annotation perspectives, there is no simple corpus-level quantity that summarizes how "saturated" a parsed Bengali corpus is with complex predicates. This paper introduces two related corpus-level constructs attributed to S M Nazmuz Sakib. First, the S M Nazmuz Sakib Constant for Bengali Compound Predicate Saturation (short: Sakib Constant) is defined as the ratio between the number of compound-type dependency relations (compound and compound:lvc) and the number of verbal tokens in a Universal Dependencies (UD) Bengali treebank. Second, the S M Nazmuz Sakib Triangle (short: Sakib Triangle) is a normalized triple giving the relative shares of compound, obj, and advmod relations, interpreted geometrically as barycentric coordinates inside a triangle. Using published statistics for the UD Bengali-BRU and Bengali PUD treebanks, we compute concrete values of the Sakib Constant and Sakib Triangle and visualize them through ten data-based diagrams. We also formulate the S M Nazmuz Sakib Compound Predicate Saturation Hypothesis and the S M Nazmuz Sakib Bipolar Headedness Partition Principle, linking these quantities to word order tendencies in Bengali. The definitions are simple and intended to be testable on larger parsed corpora, such as BDNC, bnTenTen and IndicCorp v2, as more UD-style parsers become available for Bengali.
This paper proposes a conceptual framework for integrating social robots into Islamic Arab communities in a manner that aligns with local cultural, ethical, and linguistic norms. Addressing critical gaps—such as navigating Arabic dialect diversity, adhering to Islamic ethical principles, and maintaining privacy in IoT-enabled environments—the framework comprises five core modules: a Cultural Knowledge Base, Behavior Adaptation Module, Ethical Integration Module, Arabic Language Processing System, and Privacy Management Module. This model fosters trust and acceptance by enabling robots to function effectively in healthcare, education, and public services. Future research will involve iterative testing and real-world evaluations to refine and validate the framework's applicability across diverse global and multi-faith contexts.
Background/Objectives: Obesity and insulin resistance (IR) increase the risk of mood disorders, which often manifest during young adulthood. However, neuroelectrophysiological investigations of whether adiposity and IR modify electrocortical activity and emotional processing outcomes remain underexplored, particularly in young adults. Therefore, this study used electroencephalography (EEG) to investigate whether obesity and/or IR moderate the relationships between brain potentials and affective processing in younger adults. Methods: Thirty younger adults completed a passive picture-viewing task utilizing the International Affective Picture System while real-time electroencephalography was simultaneously recorded. Two event-related potentials—early posterior negativity (EPN) and late positive potential (LPP)—were quantified. Affective processing parameters included mean valence ratings and stimulus-to-response-onset reaction times in response to unpleasant, pleasant, and neutral images. Body fat percentage and Homeostatic Model Assessment for Insulin Resistance values were measured. Hierarchical moderated regression analysis was utilized to test the interrelationships between brain potentials, adiposity, IR, and affective processing. Results: In the Negative−Neutral condition, lean and insulin-sensitive participants gave less negative valence ratings to unpleasant versus neutral images when late-window LPP amplitudes were larger, whereas this relationship was reversed in participants with obesity and absent in those with IR. Contrariwise, neither obesity nor IR moderated LPP responses to affective processing parameters in the Positive−Neutral or Negative−Positive valence conditions. Additionally, obesity and IR did not moderate the links between EPN responses and affective processing parameters in any of the valence conditions. Conclusions: Lean, insulin-sensitive young adults showed attenuated affective processing of unpleasant stimuli through stronger neural responses, whereas neural responses to pleasant stimuli did not vary across levels of body fat or IR. These preliminary findings suggest that both obesity and IR increase the vulnerability to mood disorders in young adulthood.
This paper focuses on data-driven dependency parsing for Vedic Sanskrit.We propose and evaluate a transfer learning approach that benefits from syntactic analysis of typologically related languages, including Ancient Greek and Latin, and a descendant language -Classical Sanskrit.Experiments on the Vedic TreeBank demonstrate the effectiveness of cross-lingual transfer, demonstrating improvements from the biaffine baseline as well as outperforming the current state of the art benchmark, the deep contextualised self-training algorithm, across a wide range of experimental setups.
Abstract Cognitive reappraisal is a fundamental emotion regulation strategy for mental and physical well-being, but how its neural mechanisms relate to individual differences remains poorly understood. In a consortium effort analyzing 40 fMRI datasets ( N =2,175), we examined the relationship between neural activation during reappraisal tasks and three core individual difference indices of reappraisal capabilities: (1) trait questionnaires, (2) task-based affective ratings, and (3) amygdala down-regulation. Strikingly, there was no shared overlap across these three common indices. Only a very weak correlation emerged between amygdala down-regulation and task-based affective ratings. Whole-brain analyses revealed no reliable neural associations with trait questionnaires, and associations with task-based affective ratings fell outside canonical emotion regulation networks (e.g., prefrontal circuitry). Moreover, amygdala down-regulation, often interpreted as a stable individual marker, was confounded by person-specific whole-brain responses — a limitation extending to fMRI research beyond the emotion regulation domain. These findings challenge the assumption that an individual’s prefrontal activity is a valid indicator of their reappraisal capabilities and suggest that common trait, behavioral, and neural measures might capture distinct facets of emotion regulation. More broadly, our results highlight concrete methodological challenges for fMRI research on individual differences, with implications extending beyond emotion regulation to the neuroscience of personality, psychopathology, and general well-being.
Early word learning is a critical milestone for children, yet autistic children often experience delays in language development. Social communication differences are a core feature of autism and may contribute to variability in learning experiences. Prior research has shown that word-level features such as iconicity, concreteness, and input frequency shape the timing of word learning, but less is known about the role of social word features. This study examined whether social word ratings predict when words tend to be acquired by autistic and non-autistic children. Social word ratings were examined as a predictor of word-level autistic and non-autistic acquisition normative data, while accounting for word input frequency. Regression analyses demonstrated that social ratings significantly predicted vocabulary acquisition, even after controlling for word frequency. Additional analyses demonstrated that socialness ratings continued to be a unique predictor of word acquisition when other affective features of words were included in the model (i.e., arousal and valence); this was also the case when iconicity and concreteness were included. Importantly, differences in group and interactions with social ratings and group were not statistically significant in any of the models. Lastly, the pattern of highly social words being acquired later in vocabulary development was strongest for nouns; the association was non-significant when examining verbs separately. Thus, in addition to previously studied word features like concreteness, imageability, and iconicity, social word features are predictive of vocabulary acquisition. These findings highlight an overlap in word features that influence learning in autistic and non-autistic children.
This article investigates the pragmatic functions of first-person pronouns (I/we) in political speeches and formal writing across multiple language systems. By analyzing the discourse of prominent political leaders such as Shavkat Mirziyoyev, Joe Biden, and Emmanuel Macron, the study explores how first-person pronouns function to construct authority, inclusiveness, and responsibility. The research highlights variation across languages in terms of politeness, formality, and rhetorical strategy. Methodologically, the study employs comparative discourse analysis and pragmatic interpretation of political and academic texts. The findings demonstrate that the usage of "I" versus "we" reflects not only linguistic norms but also culturally embedded leadership styles.