Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This research explores the historical emergence of linguistic terminology in three languages—English, Uzbek, and Karakalpak—with special attention to the role of Latin, Greek, and Arabic heritage. It traces how borrowed concepts were nativized and localized in each linguistic setting. By juxtaposing five evolutionary stages in English with analogous processes in Uzbek and Karakalpak, the paper illustrates the interplay between international scholarly traditions and indigenous linguistic norms. The conclusions highlight both universal tendencies and language-specific particularities in the growth of terminological systems.
This study investigates the diachronic drift of Arabic future-marking parti cles, empirically testing the shift from synthetic (سـ, سوف) to analytic (راح, قاعد+ح) forms across registers and regions. Leveraging a multi-register corpus (Penn Arabic Treebank, Corpus of Contemporary Arabic, Arabic Gigaword Fifth Edition) and advanced NLP tools (MADA, CAMeL Tools, AraBERT), we extracted and analyzed future-marker tokens annotated for register (newswire, opinion, social media, religious), region (EG, SA, LB, MA), and time slice (1990-2000, 2000-2010, 2010-2023). Mixed-effects logistic regression revealed significant effects of time, register, and region, confirming a clear diachronic shift towards analytic markers, particularly راح and قاعد+ح. Interaction terms highlighted that this shift is more pronounced in informal registers and certain regions, indicating dialectal pressure and diffusion of innovations (e.g., Gulf Arabic راح) into broader usage. Hierarchical clustering of contextual embeddings would further validate semantic-pragmatic shifts. This research provides robust evidence for ongoing linguistic change in Arabic, contributing to theories of grammaticalization and language contact.
With the increasing prevalence of mental health issues, music therapy has gained attention as a non-pharmacological intervention, and deep learning techniques have shown promise in music emotion recognition and preference prediction. This study constructed a deep neural network model (CNN+RNN/EEGNet) to efficiently identify music type preferences and examine the influence of user familiarity on prediction accuracy. EEG signals were collected using a four-channel Muse S wearable device, and user familiarity scores were used as input features. The study followed a four-stage workflow: preparation, experimental design, model construction, and result analysis. In the experimental design, music was categorized into rock, ballad, and folk, and EEG data and familiarity ratings were collected for each category. Data was trained and tested using CNN+RNN or EEGNet models, and model performance was evaluated via subject-level 10-fold cross-validation. Results indicated that predicting all music types with EEG data alone achieved an accuracy of 82.28 ± 3.42%. For individual music types, accuracies were 91.13 ± 3.60% (rock), 91.83 ± 2.07% (ballad), and 87.87 ± 4.76% (folk). When incorporating user familiarity as a feature and using a multi-level rating output, overall prediction accuracy increased to 94.94 ± 1.61%, while individual music type accuracies reached 99.15 ± 1.56% (rock), 98.51 ± 2.30% (ballad), and 98.21 ± 2.60% (folk). These results demonstrate that combining familiarity features with a multi-level scoring system significantly improves the prediction of music preferences. By using an affordable, wearable Muse S EEG device and leveraging user familiarity, this study successfully developed a highly effective deep neural network model (CNN+RNN/EEGNet) for recognizing music type preferences. The findings indicate that both overall and individual music-type predictions benefit from the inclusion of familiarity information, highlighting the potential of this approach for personalized music recommendations and music therapy applications.
This repository contains GSD-NP and GSD-DiNoS, both derived from Universal Dependencies' (UD) GSD Treebank. GSD-NP (.conllu) is a subset of UD-GSD and comprises its simplex noun phrases (NP): Common nouns (NN/NOUN) and their direct dependents (determiners, adnominal adjectives, nmods, adpositions, adverbs). It consists of 49,425 NPs (119.0k tokens) and has an improved feature annotation coverage (gender, case, number). Breaking with UD annotation, a total of 3,649 APPRART tokens were reconstructed in GSD-NP to restore the original orthographic forms. GSD-DiNoS (.json) is a custom data-driven lexion-like data structure built on GSD-NP, which aggregates NPs with the same head lemma. For each lemma, absolute frequencies of the lemma and its word forms are captured. Moreover, each occurrence feeds into three areas of interest within the word form entry: morphosyntactic features in isolation (gender, case, number), in combination with groups of dependents (collocations), and in combination with the syntactic function (dependency relations). GSD-DiNoS spans 17,433 unique lemmas and 20,190 unique word forms, stemming from 49,416 NPs. Lemmas were relemmatised to assign unique lemmas to nominal compounds, a highly productive and often lexicalised construction in German.
This study presents a lexicon-based semantic tagging information system developed for the Uzbek language corpus. The system employs the six-volume Explanatory Dictionary of the Uzbek Language (OʻzTIL) as its primary lexical resource, which contains over 85,000 entries with full semantic definitions, making it the most authoritative normative lexicographic source for Uzbek. An ontological model organized in three hierarchical levels – top, mid, and low – was designed to categorize lexical units extracted from the dictionary. Five core semantic categories were formed: animal names (approximately 100–200 units), bird names (approximately 100–150 units), personal nouns (approximately 500+ units), place names (approximately 300+ units), and occupation names (approximately 200+ units), totaling approximately 1,200–1,400 lexical units. A rule-based automatic tagging algorithm was developed to annotate corpus tokens against this structured lexical database, assigning standardized semantic tags. The system addresses key challenges inherent to Uzbek, including agglutinative morphology and lexical ambiguity. Compared to international systems such as WordNet and USAS, the proposed dictionary-based approach demonstrates superior normative grounding and cultural adequacy for Uzbek. The system is intended to serve as a foundational open resource for downstream natural language processing tasks, including machine translation, information retrieval, and intelligent educational applications.
We present a large-scale evaluation of the Menzerath-Altmann law (MAL) in the verbal domain across 180 languages, using the Universal Dependencies (UD) treebank collection (v2.17).MAL predicts that as the number of constituents of a linguistic unit increases, their average size decreases.We propose a metric to estimate the MAL effect across corpora of widely varying sizes and define threshold-based categories to classify languages along a MAL preference cline.Crucially, we analyse the preverbal and postverbal domains separately, in addition to the standard bilateral MAL, and control for potential sampling bias by comparing results across language families (Indo-European vs. non-Indo-European) and syntactic types (VO, OV and no dominant order).Our results confirm MAL as a typologically widespread preference but not an absolute universal: several languages display a trivial or even opposite (anti-MAL) tendency.Furthermore, we uncover a significant asymmetry between the two sides of the verb: the MAL effect is stronger in the postverbal domain, while anti-MAL is stronger in the preverbal domain.VO languages tend to show a stronger MAL preference postverbally, whereas OV languages do so preverbally.These findings challenge the widespread assumption that length-based ordering constraints apply symmetrically on both sides of the verb and contribute new cross-linguistic evidence to the debate on the interaction between dependency length minimization and constituent size.
Recent advances in Large Reasoning Models (LRMs), particularly those leveraging Chain-of-Thought reasoning (CoT), have opened brand new possibility for Machine Translation (MT). This position paper argues that LRMs substantially transformed traditional neural MT as well as LLMs-based MT paradigms by reframing translation as a dynamic reasoning task that requires contextual, cultural, and linguistic understanding and reasoning. We identify three foundational shifts: 1) contextual coherence, where LRMs resolve ambiguities and preserve discourse structure through explicit reasoning over cross-sentence and complex context or even lack of context; 2) cultural intentionality, enabling models to adapt outputs by inferring speaker intent, audience expectations, and socio-linguistic norms; 3) self-reflection, LRMs can perform self-reflection during the inference time to correct the potential errors in translation especially extremely noisy cases, showing better robustness compared to simply mapping X->Y translation. We explore various scenarios in translation including stylized translation, document-level translation and multimodal translation by showcasing empirical examples that demonstrate the superiority of LRMs in translation. We also identify several interesting phenomenons for LRMs for MT including auto-pivot translation as well as the critical challenges such as over-localisation in translation and inference efficiency. In conclusion, we think that LRMs redefine translation systems not merely as text converters but as multilingual cognitive agents capable of reasoning about meaning beyond the text. This paradigm shift reminds us to think of problems in translation beyond traditional translation scenarios in a much broader context with LRMs - what we can achieve on top of it.
The paper presents a prototype of a web-app designed to automatically generate verb valency lexica based on the Universal Dependencies (UD) treebanks.It offers an overview of the structure of the app, its core functionality, and functional extensions designed to handle treebank-specific features.Besides, the paper highlights the limitations of the prototype and the potential of its further development.
The limitless semantic potencies of communication is within the framework of the language conventional semantics, which imposes a number of restrictions, including on the explication of emotional experiences by the speaker. The latter either chooses a read y-made preset formula, or directs communicative efforts to search for and objectify emotional and semantic shades of meaning with an uncodified form of verbalization. If the form of expression of an emotional experience is new, atypical, unconventional, we should talk about the representation of diffuse emotive semantics, approaching the actual emotional experience. Diffusivity (fuzziness, vagueness, multiple inconsistencies, ambiguity) is an immanent property of semantics that corresponds to both the natur e of the sign and the environment in which the sign acts. The assumption is that, depending on the characteristics of the discourse and the genre characteristics of the elements included in it, artistic communication was considered. The author of a work of art must go beyond the linguistic prescription, which allows him to have the desired effect on the addressee. The analysis of a dramatic work shows a variety of forms of explication of diffuse emotivity, when emotional experiences become a discursive and genre-forming category: the expression of complex vague emotions allows creating an image of a multifaceted and interesting character. The texts of modern plays allow tracing a similar trend towards diversifying the form of expression of diffuse emotivity, but the emotional tonality is less diverse: negative emotional experiences set the emotional dominant, therefore, the ‘consolidation’ of the emotive occurs rather than through the vector of mixing positive and negative assessments, but the intensification and concretization of the negative evaluative component. The author also postulates that language always approximately describes emotions, but in artistic communication such approximativeness is expressed in conscious and creative imitation, which transforms and develops the linguistic norm.
BACKGROUND: Psychopathic characteristics are associated with an elevated risk for violent behavior and are therefore of interest in research studies. Despite extensive research, the role of emotional and attentional anomalies in subclinical psychopathic traits remains a subject of ongoing debate, possibly attributed to the multifaceted nature of the construct. The study aims to explore how distinct psychopathic traits may differently relate to underlying emotional and attentional mechanisms. METHODS: To further explore the emotional and attentional anomalies underpinning the three Triarchic Psychopathy constructs, boldness, meanness, and disinhibition, this study employed an optimized picture-startle paradigm to address the limitations in commonly used paradigms that capture these dynamics only after 1000 ms post-image onset. This paradigm included negative, positive, and neutral images to elicit varied emotional responses, while auditory startle probes were presented at 50, 700, or 4500 ms post-image onset to measure emotional and attentional fluctuations. In this study, it was proposed that each psychopathic trait correlates with distinct emotional and attentional anomalies. A mixed-gender community sample of 115 participants was included. Eyeblink startle amplitudes (ESAs) were recorded via smartphone technology that utilizes facial landmark data captured via phone cameras, while the P3a and late positive potential (LPP) were measured through electroencephalography (EEG). RESULTS: The results revealed an exaggerated attentional bottleneck associated with boldness in males, indicated by increased P3a amplitudes in response to negative images. Meanness was associated with lower empathy scores and arousal ratings, and reduced ESAs at 4500 ms for negative images, supporting socio-emotional difficulties in meanness. In contrast, disinhibition showed no significant emotional or attentional deviations in this study. CONCLUSIONS: Our findings highlight trait-specific differences in neurocognitive functioning and validate the effectiveness of the optimized picture-startle paradigm for dissociating attentional versus emotional anomalies across triarchic psychopathic traits. The study also demonstrates the feasibility of using BlinkLab’s integrated stimulus presentation and camera-based eyelid tracking with concurrent EEG measures of attentional and affective processing (P3a and LPP), providing a complementary approach that may facilitate scalable data collection.
The article describes the principles of forming linguistic and communicative competence in future doctors during their studies at medical higher education institutions with a view to popularisation medical knowledge, and substantiates the content and structure of a special linguistic discipline in the Ukrainian language for popularisation medical knowledge. The regulatory documents of the Ministry of Health of Ukraine, which describe the qualification characteristics of professionals in the field of medicine and dentistry, emphasise the importance of the multifaceted activities of doctors in disseminating medical knowledge among the population. The current educational standard in the healthcare sector does not include a component that would promote the development of language skills for popularisation medical knowledge. Therefore, it is important to introduce the study of the peculiarities of the linguistic structure and production of relevant genres of popular scientific medical language into the curricula of higher medical education institutions. It is proposed to introduce a special component in one of the senior courses to develop future doctors’ linguistic and communicative competence in popularisation medical knowledge – the academic discipline ‘Ukrainian Language Popularisation of Medical Knowledge’. This discipline is intended for higher education students who have already studied several disciplines in the professional training cycle and have sufficient background knowledge to independently create high-quality popular science texts on the medical disciplines they have studied. As a result of studying the discipline, higher education students should know the terminology of popular science medical text creation, the genre structure of popular medical discourse, be able to create works in relevant genres of popular medical discourse, and master the linguistic norms of popular medical style. The possibility of introducing topics for the formation of linguistic and communicative competence in the popularisation of medical knowledge into the compulsory course of the Ukrainian language (for professional purposes) with the addition of credits in the fourth and fifth years of study is justified. The possibility of introducing a similar discipline for higher education students of various fields and specialities is indicated.
Digital Chinese landscape painting can be enhanced through multisensory fusion, yet culturally adapted cross-modal mappings and real-time adaptation remain under-tested. This laboratory, between-subjects virtual reality experiment (N = 80) compared a visual-only condition with a multisensory condition delivering synchronized auditory, haptic, and olfactory cues, with cue mapping implemented as culturally congruent or deliberately incongruent. Outcomes included standardized immersion/presence measures and affect ratings, behavioral logs of interaction and exploration (including time-on-task and coverage), and physiological indices of arousal derived from electrodermal activity and photoplethysmography-based pulse metrics, assessed immediately after exposure and at one-week and one-month follow-ups. Culturally congruent multisensory cueing produced higher immersion/presence and engagement than visual-only presentation and the incongruent multisensory format, alongside more frequent and diverse interaction behavior, longer time-on-task, and broader exploration of the virtual scenes. Self-reported emotion showed higher positive valence and arousal under congruent cueing with convergent physiological arousal patterns, and differences remained observable at follow-up assessments.
Background: Along with the development of sociolinguistic studies and language education, the orientation of research and learning practices is shifting from a monolingual approach to a more inclusive multilingual framework. In this context, translanguaging is seen as a communicative practice that allows individuals to use the entirety of their linguistic resources dynamically to build meaning, form identity, and establish social relationships. Purpose: This study aims to explore the practice of translanguaging within English Area communities that function as Community of Practice (CoP), focusing on the role of such practices in strengthening social cohesion as well as in the identity negotiation process between core members and outsiders of the community. Method: This research applies ethnographic case study approach. Data collection was conducted through participatory observation, semi-structured interviews, and review of community documents involving various actors, ranging from core administrators, active members, new members, former members, to external observers. Data analysis was conducted using Thematic Analysis to reveal patterns of language use, linguistic norms that develop, and the dynamics of social identity formation in the community. Results and Discussion: Research findings show that the English Area acts as a community of practice that supports the collaborative and sustainable English learning process. The practice of translanguaging not only serves as a pedagogical strategy to facilitate the understanding of the material, but also as an effective strategy that contributes to decreasing language anxiety and increasing member engagement. The use of language in the community is flexible and situational, where English is applied in accordance with the goals of the activity and the level of readiness of the participants. Although the ability to speak English serves as a symbol of membership, the application of language flexibility actually strengthens the sense of community and reduces the potential for social exclusion. Conclusions and Implications: This study confirms that translanguaging plays an important role in language learning, identity formation, and the sustainability of communities of practice, and has implications for the development of more inclusive and contextual English learning practices.
KIParla is a large, modular corpus of spontaneous spoken Italian originally transcribed in ELAN using Jefferson-style conventions. While this representation preserves fine-grained interactional detail and time alignment, it limits interoperability, large-scale querying, and computational reuse. This paper presents the design and implementation of a pseudo-tokenized, verticalized pivot format developed to support validation, maintenance, and infrastructural integration without sacrificing descriptive richness. The proposed format makes explicit the analytical units implicit in Jefferson transcription—transcription units, spans, and tokens—and enforces well-formedness constraints at character, span, and unit levels. Overlap, the most complex relational phenomenon, is resolved through a graph-based algorithm that derives temporal overlap events from alignment data and deterministically matches them to textual spans. Each token is represented as a structured record enriched with lexical, prosodic, interactional, and alignment features, anchored through explicit character offsets. The vertical format functions as a maintained pivot representation from which alternative formats, including ELAN files and UD-compatible treebank representations, can be reproducibly derived. This architecture enables large-scale lemmatization, part-of-speech tagging, syntactic annotation, and cross-layer querying, while supporting version-controlled, DevOps-inspired workflows for sustainable corpus growth. The KIParla pivot format thus reconciles interaction-oriented transcription practices with computational standards and provides a model for reuse-oriented spoken-language data engineering.
Fine-tuning large language models on sensitive data poses significant privacy risks, as membership inference attacks can reveal whether individual records were used during training. While Differential Privacy (DP) provides formal protection, applying DP to conventional Parameter-Efficient Fine-Tuning (PEFT) methods such as Low-Rank Adaptation (LoRA) often incurs substantial utility loss. In this work, we show that a more structurally constrained PEFT architecture, Tensor Train Low-Rank Adaptation (TTLoRA), can improve the privacy-utility tradeoff by shrinking the effective parameter space while preserving expressivity. To this end, we develop TTLoRA-DP, a differentially private training framework for TTLoRA. Specifically, we extend the ghost clipping algorithm to Tensor Train cores via cached contraction states, enabling efficient Differentially Private Stochastic Gradient Descent (DP-SGD) with exact per-example gradient norm computation without materializing full per-example gradients. Experiments on GPT-2 fine-tuning over the Enron and Penn Treebank datasets show that TTLoRA-DP consistently strengthens privacy protection relative to LoRA-DP while maintaining comparable or better downstream utility. Moreover, TTLoRA exhibits lower membership leakage even without DP training, using substantially smaller adapters and requiring on average 7.6X fewer parameters than LoRA. Overall, our results demonstrate that TTLoRA offers a practical path to improving the privacy-utility tradeoff in parameter-efficient language model adaptation.
Abstract Introduction Sleep supports emotion regulation by preferentially consolidating emotional memories while attenuating reactivity. We have shown that dream recall plays an active role by increasing negative over neutral memories and reducing reactivity. In women, fluctuating reproductive hormones across the menstrual cycle influence sleep features implicated in emotional memory, yet whether menstrual phases influence how dreams shape emotional processing remains unknown. This study investigates how dreams shape sleep-dependent emotional processing across the menstrual cycle in naturally cycling women. Methods 128 women (Mage = 32.85 ±11.93 years) completed up to four visits across verified menstrual phases (menses, late-follicular, mid-luteal, late-luteal). At each visit, participants performed the Emotional Picture Task with negative and neutral IAPS images in the evening (Test 1) and the next morning (Test 2). Participants rated old/new, arousal, and valence of images shown at each test. Dream reports were collected upon waking prior to Test 2. Linear mixed-effects models tested main and interaction effects of menstrual phase and dream recall. Results The menstrual cycle altered how dreaming shaped overnight emotional memory. Dream recall typically benefited the emotional trade-off effect —favoring consolidation of negative relative to neutral images (Δd′; t(410)=1.95, p=0.05)—but this pattern reversed during the late-luteal phase (dream × menstrual cycle: t(381)=-2.29, p=0.02). Dreaming showed independent effects on emotional reactivity. Higher valence and arousal ratings for negative images during Test 1 predicted greater dream recall (valence: t(344)=2.05, p=0.04; arousal: t(327)=2.04, p=0.04). Additionally, the more negatively participants rated the images at Test 1, the more negative their dreams tended to be (t(166)=-2.11, p=0.04). Dream recall was linked to reduced next-morning emotional reactivity (valence: t(413)=-2.89, p=0.004; arousal: t(413)=-2.65, p=0.01), with stronger reductions following more negative dreams (β=0.15, t(182)=2.86, p=0.005). Conclusion Menstrual cycle phase influenced how dreams shaped overnight emotional memory. Negative waking experiences increased dream recall and shaped dream content—and recalling dreams, especially negative ones, reduced emotional reactivity and typically strengthened emotional memory—but this benefit disappeared in the late-luteal phase when there are declining reproductive hormones. These findings suggest a novel interaction between the menstrual cycle and dreaming, showing that hormonal fluctuations reshape how sleep and dreams regulate emotional experience and memory. Support (if any) RF1AG061355 (Baker/Mednick)
Contemporary scholarly discourse on gender-inclusive communication remains predominantly descriptive, often avoiding a systematic critical analysis of its internal contradictions and social consequences. However, the growing social tension around new linguistic norms, their ideologization, and direct impact on public institutions demand unbiased examination. The aim of this article is to identify and analyze the key paradoxes generated by gender-inclusive communication, which persist despite its proclaimed goals of equality and respect. The research material comprises English-language texts of various genres and styles, including scholarly articles, media publications, documents from university websites, healthcare institutions, governmental, non-governmental, and commercial organizations, as well as data from blogs and social networks from the period 2017 to 2025. This allowed us to examine gender-inclusive communication both in the sphere of academic reflection and within the context of public practice. The methodological framework is based on critical discourse analysis, which interprets linguistic changes as a struggle for power, and Lotman’s theory of the semiosphere, which views inclusive language as a phenomenon of cultural dynamics. The study establishes that inclusive communication generates a complex of systemic contradictions across different dimensions: linguistic (between the striving to erase and simultaneously multiply gender differences, leading to semantic tautology and a violation of linguistic conventionality); social (where inclusivity in practice becomes a tool for excluding dissenting voices and marginalizing the experiences of traditional groups); ethical (encompassing the conflict between the ideology of self-identification and biological realities, as well as the imposition of Anglo-centric models onto other linguacultures). Interpreting the results through the chosen methodological lens reveals that inclusive language functions not only as a discourse of power but also as a tool of auto-communication, aimed at redefining the core of the cultural semiosphere and consolidating a “progressive” identity. The findings open perspectives for comparative studies on the reception of inclusive practices in different linguacultures and for interdisciplinary research into the long-term social effects of linguistic reform.
Abstract Mood states strongly influence episodic memory processing, yet the neural mechanisms through which mood interacts with stimulus valence during encoding remain unclear. The present study examined how experimentally induced negative mood would modulate neural processing and behavioural outcomes during an associative memory task, and whether mindfulness intervention can alter these effects. Twenty healthy adults completed a memory-encoding (word-stimulus pairs) task in which mood induction (negative vs. neutral) prior to encoding was crossed with stimulus valence (neutral vs. negative), producing four conditions. Continuous EEG was recorded using a 64-channel system, and event-related potentials (ERPs) and oscillatory dynamics were analysed during stimulus-locked encoding epochs. After an initial retrieval session, participants were randomly assigned to either a brief mindfulness meditation intervention or an active podcast-listening control, followed by a second retrieval session. Our behavioral data indicated that, recognition accuracy, arousal, valence ratings, and confidence differed significantly across conditions. Participants formed stronger word–image associations for neutral than negative stimuli across conditions, with associative memory being most impaired when negative stimuli were encoded during a negative mood state. At the neural level, early perceptual–affective processing showed robust Mood × Valence interactions: central P200 and occipito‑parietal EPN amplitudes differentiated negative from neutral images during neutral mood, but this valence effect was markedly reduced under negative mood, indicating an early blunting of neural sensitivity to emotional content. Following the mindfulness intervention, the negative images encoded under negative mood showed the most pronounced reduction in arousal, along with heightened alpha activity during mediation. Together, these findings show that negative mood undermines associative binding and influences early stages of visual-affective processing, while mindfulness may primarily positively influence the associated affective responses.
Background: Auricular asymmetry affects facial aesthetics; yet, the degree of asymmetry discernible by a casual observer remains poorly defined. Objective: To measure the casual observer’s perception of auricular size, angulation, protrusion, and vertical position. Methods: Artificial intelligence-generated facial images were utilized and unilateral variations created in (1) size, (2) angulation, (3) protrusion, and (4) vertical position. A total of 200 responders reviewed 46 images, rating the ear appearance on a Likert scale of 0–10. Results: Causal observers rated a ≥10% change in size significantly different than baseline ( p < 0.001). Retrograde angulation ≥10° and anterograde angulation at 20° were rated significantly different than baseline ( p < 0.001). Positive asymmetrical protrusion was rated more severely than negative asymmetrical protrusion. The vertical position difference was rated least critically on average compared with baseline. Conclusion: Casual observers more critically perceived auricular size asymmetry and retrograde angulation when compared with protrusion and vertical position asymmetry.
Pronouns indicate significant importance in both pedagogical and communicative contexts, as they shape the way individuals are addressed and understood in social interactions and educational settings. Beyond the traditional pronouns “he” and “she,” the American Psychological Association endorses the scholarly use of the singular pronoun “they,” recognizing its relevance in promoting inclusive language practices. In addition, the popularity of neopronouns continues to rise, providing non-binary individuals with a broader range of linguistic options to express their identities. Despite this growing recognition, there remains a dearth of empirical research that systematically investigates the awareness, knowledge, and preferences regarding pronoun use among non-binary populations. Addressing this gap, the present quantitative inquiry examined the level of awareness, knowledge, and preference of pronouns among non-binary college students at a state university in the Philippines. The study involved 80 participants, including 20 lesbians, 20 gays, 20 bisexual males, and 20 bisexual females, selected through criterion sampling, who responded to a four-part researcher-developed survey questionnaire. The results indicate that, overall, non-binary college students are aware of the categories of pronouns (M=2.48, SD=0.39 for traditional pronouns; M=2.79, SD=0.25 for gender-neutral pronouns; and M=2.64, SD=0.32 for neopronouns) and knowledgeable about them (M=2.49, SD=0.31 for traditional pronouns; M=2.68, SD=0.23 for gender-neutral pronouns; and M=2.47, SD=0.32 for neopronouns). However, despite this awareness and knowledge, participants expressed a preference for using traditional pronouns (“he” and “she”) when being referred to. These findings underscore the persistence of traditional linguistic norms in educational settings and highlight the potential influence of formal language instruction on pronoun preference. Empirically, this study contributes to the limited body of research on non-binary pronoun use speficically in the Philippines, providing a foundational dataset that can inform inclusive language policies, pedagogical strategies, and future sociolinguistic investigations. Its significance lies not only in documenting patterns of pronoun awareness and preference but also in offering evidence-based insights for educators, policymakers, and advocates seeking to foster more inclusive and affirming learning environments.
This paper quantifies how parameter-efficient fine-tuning narrows the performance gap between compact transformers and massive Large Language Models (LLMs) on sentiment analysis while reducing memory, energy, and financial budgets. Experiments use the publicly released BERT-base-uncased checkpoint (12 layers, 110 M parameters), the Stanford Sentiment Treebank-2 benchmark (67,349 movie phrases with binary polarity labels), PyTorch and Hugging Face Transformers 4.41, Weights & Biases for experiment tracking, and an NVIDIA A100 GPU configured for mixed-precision computation. Low-Rank Adaptation (LoRA) adapters (rank 16) are injected into the query/key/value projections, so only 0.54 % of weights are updated. The model was trained for five epochs with AdamW, cosine scheduling, batch size 64, and optional 4-bit post-training quantization. Accuracy and macro-averaged F1 are logged every 25 steps. The LoRA-tuned model achieves 91% accuracy and 0.91 F1 on the SST-2 test set, an improvement of 40 percentage points over the off-the-shelf checkpoint and comparable to GPT-4o-mini (93%), while using fewer than 1⁄1500 of its parameters. Validation loss plateaued without overfitting, and 4-bit quantization compressed the model to 27 MB with <0.5 point accuracy loss. Energy profiling shows a 73% reduction in GPU consumption compared with full-parameter fine-tuning. Purpose-built adapters and quantization can unlock high-quality NLP on edge devices. Future work should extend the protocol to multilingual corpora, streaming inference, federated learning, and on-device continual adaptation to preserve accuracy under concept drift and safeguard user privacy.
This paper presents a small-scale dependency treebank for Tunisian Arabic (TADT) developed within the Universal Dependencies framework, addressing the scarcity of linguistic resources for the Arabic varieties.The approach employs domain adaptation, leveraging a machine learning model (UDPipe 1.0) trained on Algerian Arabic data to annotate 100 Tunisian Arabic social media comments, followed by manual correction.This pilot study evaluates the feasibility of using machine learning-assisted annotation to scale resource development for spoken Arabic and identifies key challenges in cross-dialectal transfer for improving annotation quality and efficiency.This work contributes to more inclusive and fair representation of Arabic linguistic varieties in academic research and NLP applications.
Currently, East Asia, especially South Korea, is facing social problems such as population decline and regional inequality. To address these challenges, tourism has been leveraged as a means of economic revitalization, especially in fishing villages that are economically disadvantaged. This study examined the authentic food experiences of tourists who visited fishing villages. Tourists’ food experiences, dimensions of destination food image, and destination loyalty were assessed. In September 2024, 448 responses from Korean tourists were collected and analyzed using confirmatory factor analysis and structural equation modeling to test 15 hypotheses. Local food authenticity showed a significant effect on all dimensions of destination food image. Of the five dimensions of destination food image in this study, food taste, health and hygiene, and unique cultural experiences significantly influenced destination loyalty. In addition, geographic differences moderated the relationship between local marine food authenticity and the perceived food image of the destination. Tourists in the southern coastal regions reported the highest destination food image ratings, driven by authentic local cuisine, while those in the western regions reported the lowest. This study offers both practical and theoretical implications related to sustainable coastal tourism.
Aims and objectives/purpose/research questions: This study examines how digital platforms shape bilingual advertising in post-Soviet Georgia, addressing three questions: (1) What linguistic strategies (loanwords, code-switching, hybrid forms) prevail? (2) How do urban and rural consumers perceive these strategies? (3) How do state regulations and platform constraints shape advertisers’ choices? Design/methodology/approach: We employ a mixed-methods design combining corpus analysis (200 digital advertisements from Facebook, Instagram, TikTok, Google Ads), consumer surveys (225 respondents stratified by age, education, settlement type), and focus groups with 16 advertising professionals. The approach applies offline–online nexus frameworks to examine platform-mediated language practices. Data and analysis: The corpus (January–June 2024) was coded for linguistic strategies using inter-coder reliability (Cohen’s κ = 0.85). Survey data (Cronbach’s α =.85) were analyzed via analysis of variance (ANOVA) to compare perceptions across demographics. Focus group transcripts underwent thematic analysis (κ = 0.78). Findings/conclusions: Partial English use (43%) prevails, driven by code-switching and hybrid forms like “ქეშბექი” (keshbeki, cashback). Younger respondents aged 18–50 (67%) view English ads as innovative, while 35% of all respondents express linguistic marginalization concerns. Advertisers navigate tensions between Georgia’s Advertising Act mandating Georgian-language content and Google Ads’ technical inability to support Georgian script, creating “platform-enforced bilingualism.” Originality: This is the first systematic study of English in Georgian advertising. It introduces “platform-enforced bilingualism” as a concept bridging prestige models and structural constraint perspectives, demonstrating how corporate platform governance supersedes national language policy in digital commercial spaces. Significance/implications: The study extends nexus scholarship to commercial contexts and post-Soviet language policy frameworks. It reveals asymmetric feedback loops where online platform constraints shape offline linguistic norms more powerfully than policies shape online practices, with implications for language policy revision and bilingual education in digitally mediated environments.
Species-level tree identification is a fundamental task in forest monitoring, biodiversity assessment, and climate-smart ecosystem modeling. Close-range laser scanning technologies have become indispensable tools for forest mapping because they provide high-resolution, three-dimensional structural data at the individual-tree level. However, species-level identification remains a major challenge in global environmental monitoring and AI-driven ecological assessment owing to high species diversity, structural plasticity, and variability across sensing platforms. Here, we propose the cognition-inspired multimodal attention fusion network (CI-MAFusion), a dual-branch deep learning framework that integrates point cloud data with multi-view imagery. Guided by expert dendrological reasoning and cognitive neuroscience principles, CI-MAFusion incorporates a structural branch based on an improved graph attention-based point network for encoding 3D morphological patterns and a visual branch that processes standardized multi-view projections to extract textural features. A cross-gate attention mechanism adaptively fuses structural and visual features. Each branch uses an enhanced convolutional block attention module to highlight salient features, analogous to selective attention in the human visual system. We tested CI-MAFusion using Global LiDAR TreeBank, which contains 12,057 trees from 36 species across four continents and six Köppen climate zones. The model achieved 87.50% overall accuracy at the genus level and 86.12% at the species level, outperforming unimodal and existing fusion approaches by up to 8.1%. Additionally, it further achieved > 90% overall accuracy across regions and > 80% across climate zones, with attention visualizations highlighting biologically diagnostic features such as crown contours, bark textures, and branch junctions. This cognitively inspired architecture improves generalization and advances AI-based systems toward robust recognition of biological structures in complex environments.
Building on evidence for experience-specific grounding of word meaning and interindividual differences therein, this study investigated how specific aspects of empathy modulate the processing and representation of abstract emotional words. We investigated single-trial N400 amplitudes as a measure of semantic retrieval in 78 healthy adults during a delayed lexical decision task with emotion-label, emotion-laden, and neutral abstract words. We further measured the participants' levels of empathic concern, fantasy, personal distress, and perspective taking. Additionally, ratings on valence, arousal, and emotional experience quantified the words' emotional representational content. While direct comparison yielded no evidence for N400 differences between word types, N400 amplitudes in response to emotion-label words decreased with increasing fantasy scores, with this modulation being stronger than for emotion-laden and neutral words. Additionally, participants with higher fantasy scores rated emotional words higher in absolute valence. The observed N400 reductions thus seem to reflect fantasy-driven processing facilitation graded by the words' emotionality level. In contrast, we found no evidence for N400 modulations by empathic concern, personal distress, or perspective taking while affective ratings on all scales increased with increasing empathic concern scores. Our findings suggest that fantasy facilitates emotion-label word processing, and empathic concern enriches emotional word meaning representations, demonstrating interindividual differences in the experiential grounding of emotional abstract concepts.
<sec> <title>UNSTRUCTURED</title> Generic sentiment analysis tools are widely deployed in digital mental health and longitudinal text research to infer psychological change. Lexicon-based approaches (eg, NRC), supervised emotion classifiers (eg, GoEmotions), and single-shot large language model (LLM) prompts typically operationalize affect as the frequency or probability of valenced tokens at the message level. While effective for detecting overt affective shifts, the limits of this operationalization remain underexamined. This Viewpoint presents a single-subject longitudinal corpus (n=531 reflective messages across six months) to illustrate a construct–measurement misalignment. The subject reported a substantial psychological transformation characterized by reframing, integration, and narrative restructuring rather than shifts in emotional tone. Phase A (lexicon counts), Phase B (supervised emotion probabilities), and Phase C (LLM affect ratings) were applied under standardized aggregation schemes (per-message scoring, arithmetic mean, binning). Across methods, no consistent longitudinal trend emerged in affective indices. We argue that this null result is not a failure of AI per se, but a measurement blind spot: when transformation occurs at the level of narrative meaning rather than valence frequency, generic sentiment tools may remain insensitive. We propose a construct-specific analytic framework emphasizing alignment between psychological target constructs and computational operationalization. Implications are discussed for responsible deployment of AI in longitudinal digital health and narrative analysis contexts. </sec>
Adequate recovery and habituation to acute stressors in daily life are essential for mental health. One potential moderator might be the use of different emotion regulation (ER) strategies contributing to interindividual differences in vulnerability to chronic stress. Rumination has been linked to impaired endocrine adaptation, whereas reappraisal tendencies have been associated with boosted habituation to psychosocial stress. Yet, no experimental study has directly compared the causal effects of these strategies on psychoneuroendocrine responses to repeated stress. To address this gap, 91 healthy participants (47 women) underwent a short Trier Social Stress Test (TSST) twice on two consecutive days and were randomly assigned to a rumination, reappraisal, or control intervention in between. Cognitive-affective ratings indexed psychological stress, while salivary cortisol, alpha-amylase (sAA), and heart rate (HR) served as biomarkers. Across the entire sample, reduced increases in negative affect and cortisol in response to the second TSST confirmed successful habituation. As expected, rumination immediately increased negative affect, reduced positive affect, and lowered perceived coping abilities indicating successful induction of ruminative thinking. Moreover, it prevented physiological habituation to repeated stress, evidenced by stable cortisol and HR responses. Unexpectedly, reappraisal also lowered perceived coping abilities in men, followed by prolonged sAA reactivity to the first stress exposure, hinting at a sex-specific impairing effect of reappraisal on noradrenergic recovery. However, reappraisal neither affected cortisol recovery nor habituation, which may result from generally poor reappraisal performance. Together, these findings provide initial evidence for differential effects of rumination and reappraisal on psychophysiological adaptations to repeated TSST exposure.
Our brain maps the space immediately surrounding the body, the peripersonal space (PPS), to sharpen sensory-motor coordination whenever an object enters it. Within PPS, past research demonstrated how several factors influence motor readiness: from stimulus characteristics, such as body-object distance and stimulus semantics, to personality traits like anxiety. However, most paradigms infer PPS boundaries using tactile and visual stimuli, with auditory cues often playing an ancillary role, rather than directly measuring response modulation to stimuli entering the PPS. Here, we measured anticipatory postural adjustments, as direct physiological indices of motor planning, to examine whether semantic content and individual suggestibility modulate responses to looming sounds stopping within PPS. Thirty-three adults heard affective semantic sounds (positive - applause, negative - dentist drill) or neutral non-semantic (pink noise) sounds halting at five within-PPS distances while we recorded muscle activation timing, distance estimates, affective ratings, and sensory suggestibility. Motor responses were faster as sounds stopped nearer the body but systematically delayed and less precise for semantic sounds compared to non-semantic sounds, with higher suggestibility predicting longer and more variable latencies, particularly for non-semantic sounds. These findings demonstrate that semantic evaluation imposes measurable processing costs that systematically delay motor preparation and increase perceived distance. Individual suggestibility produces freeze-like response patterns that amplify motor uncertainty, specifically when semantic context is absent, revealing distinct cognitive and trait-based mechanisms that jointly regulate defensive behaviour within PPS.
BACKGROUND AND AIMS: Cannabis cue reactivity paradigms are instrumental in studying the behavioral and neurocognitive mechanisms of cannabis use and cannabis use disorders; however, image sets used for cannabis cue reactivity paradigms vary between studies, and the lack of reliability and validity assessment hinders the quality of evidence they generate. The main aim of this study was to create a novel, open access, standardized and representative database of cannabis use-related images including control images matched by resolution, luminosity and complexity: The Cannabis Research Image Database (CRESIDA). The secondary aim was to examine whether subjective cannabis cue-induced craving was associated with cannabis use severity and whether this relationship was moderated by image type. As an illustrative example of how our open data can be used and how sample characteristics can shape cue reactivity, we also explored the role of cannabis-tobacco mixing by comparing cannabis cue induced cannabis and tobacco craving between individuals who did and did not mix the substances. DESIGN: An online survey was administered to participants recruited via online platforms, community advertisement and snowballing. SETTING: USA, the Netherlands and Australia. PARTICIPANTS/CASES: 689 participants who consumed cannabis monthly to daily (385 men, 298 women, 6 other) were recruited between January 2022 and May 2024. MEASUREMENTS: Out of 93 cannabis images and 93 matched neutral images, participants each rated 31 image pairs for cannabis craving (the primary outcome), arousal, valence and tobacco craving. Participants were characterized for socio-demographic data, level of cannabis use and related problems and mixing cannabis and tobacco. A subset of 78 images was selected for further analysis based on cannabis craving results. Image ratings were evaluated for internal consistency (α). Furthermore, we examined the association between cannabis cravings and cannabis use characteristics, and explored if cannabis craving ratings were affected by image type (i.e. product, paraphernalia and actions) and by using cannabis alone vs. mixing cannabis and tobacco. FINDINGS: The database showed excellent reliability (α = 0.995-0.965). Cannabis craving, valence and arousal discriminated cannabis and control images. More cannabis use days [unstandardized beta (β) = 0.162, P < 0.001] and cannabis use-related problems (β = 0.268, P < 0.001) were statistically significantly associated with higher image-related cannabis craving. Mixing cannabis with tobacco, compared with using cannabis alone, was associated with the presence of tobacco craving in relation to cannabis images, and with greater cannabis craving in relation to cannabis images (β = -0.457, P < 0.001). CONCLUSIONS: Images in the open access Cannabis Research Image Database (CRESIDA, https://osf.io/dc9nz/) appear to be reliable and valid for the scientific study of cue reactivity internationally, providing a broad range of free to use cannabis and control images.
Language, Reason, and the Costs of Abandoning StandardsI have been doing some reading in recent weeks of one of my favorite authors, Stephen Pinker, specifically a work, which has become something of a classic, published some time ago, The Language Instinct.1 In this fascinating work, Pinker includes a discussion of certain strands within contemporary language studies that advocate treating nonstandard dialects as pedagogically equivalent to standard academic language in analytic and institutional contexts. Increasingly, the trend to “deprioritize” the rules of the so-called “hegemonic” idiom has spread not only in the academic literature (especially in discourse within the “critical studies” community), but also into elementary and secondary educational contexts, where traditional instruction gives way to curricula which eschew formal semantics, structural rigor, analytical grammar and attention to conventions of standard speech. These are, in the view of many, tedious and “prescriptivist,” if not altogether oppressive, in effect. The implications of this trend bear attention.The desire for inclusivity, and a respect for the identity, and the legitimacy, of non-standard dialects has served an important corrective function. It has countered naïve assumptions about intelligence, worth, or expressive capacity based on the accent of speakers (particularly, within an educational context, members of traditionally “disadvantaged” or marginalized communities), or their native idiom, and they have rightly underscored that all natural languages, and dialects, as tools for the conveyance of meaning, are systematic, rule-governed, and capable of sustaining coherence within their speech communities. The aforementioned scholarship, among critical theorists, repetitively makes this case, and often advocates for the “disruption” of normative practice in prioritizing standard language in education. While many proponents of critical language awareness advocate additive models (combining standard instruction with critical reflection), there remains a growing pedagogical tendency—particularly in some applied contexts—to de-emphasize explicit instruction in standard forms. Alongside the salutary recognition of the potential of nonstandard idioms to convey meaning, the advocates of disruption and those who decentralize standard language – both in the study of the conventions of semantics, grammar, syntax and literature of the so-called hegemonic culture, and even (increasingly) in second language studies – are encouraged to develop in students “critical language awareness.” This means they must analyze how standard language ideologies function to maintain power, reproduce social inequality, and marginalize non-standard speakers. Standard language in this view should be depreciated, and “traditional” tenets of rigor and formality loosened, if not abandoned.2This, in my view, is a troubling tendency. A growing demand within the academic community that teachers should disparage the role of standard language as a distinct and indispensable instrument of analytic reasoning, education, and public discourse, and should ignore or relax altogether the requirements of standard language within those domains, comes at a genuine intellectual cost. This trend derives from a conflation of two claims that ought to be kept rigorously distinct: 1) that non-standard dialects are linguistically legitimate, and 2) that they are equally suited to the purposes of rigorous analysis, formal exposition, and institutional persuasion. The first claim is well supported; the second is not. Linguistic legitimacy concerns internal rule-governed structure. Analytic suitability concerns the external demands of institutional reasoning. These are different evaluative criteria.The function of standard language is not merely social or aesthetic. It is epistemic. Standard language, and its concomitant conventions, are a deliberately constrained linguistic instrument, refined over time to facilitate explicit reasoning among large, heterogeneous populations lacking shared background assumptions. Their norms - precision of reference, explicit premising, linear argumentation, controlled affect, and stable semantics—are not simply useless, arbitrary conventions imposed by cultural elites.They are, ultimately, technologies of clarity. Like any technology, they are historically contingent and socially distributed — but their function is instrumental rather than symbolic. Their value lies not in who first wielded them, but in what they enable: durable inferenceacross differences.This point becomes clearer when we consider the nature of analytic reasoning itself. Formal and informal logic alike require that premises be identifiable, that terms be defined and used consistently, that inferential steps be traceable, and that conclusions follow recognizably from what precedes them. These requirements impose significant cognitive demands, particularly on novice thinkers – such as students in elementary school and in secondary education. Language varieties or discourse registers that tolerate imprecision, ambiguity, shifting reference, lack of logical parallelism, ellipsis, emotional compression or implicit premises, and which rely heavily on coterie-defined metaphorical, metonymical and other figurative devices may function admirably in high-context, oral, or in-group settings. But when such registers are imported, unaltered, into analytic contexts, they tend to obscure inferential structure and permit, rather than resist, intellectual shortcuts.This observation should not be misunderstood as a claim about the intellectual capacities of speakers. The capacity to reason analytically is not distributed along linguistic lines. Rather, the issue is whether a linguistic system - or, more precisely, a set of linguistic norms - enforces the disciplines that analytic reasoning requires. Standard academic language, as institutionally codified, has evolved to externalize these disciplines; informal and high-context registers are typically not structured with those analytic demands in mind. What is permitted is frequently exploited, particularly by students still learning how to discipline their thought, to the detriment of cogency. The advantage of standard language is not structural superiority, but the institutionalization of constraints that externalize logical discipline.Consider, for example, features often cited in discussions of non-standard speech: double negatives, emotionally charged vocabulary, topic drift, or reliance on shared contextual assumptions. Within the speech communities in which these features are normative, they are not logically defective; they are, indeed, rule-governed and intelligible. But they are poorly aligned with analytic goals. Double negation, while unambiguous within a dialect, can complicate scope in formal reasoning. It takes the rigidity of formal logic, the agreement of the linguistic community as to application of meaning within the norms of “standard” language to resolve this. This is the case for example in Iberian languages, Langue d’oc, Langue d’oil and Italo Romance. In these contexts, the double negative construction and, for example, dative redundancy are standardized as to their interpretation within a closely framed set of linguistic applications. Emotional loading, meanwhile, collapses descriptive rigor into subjectivism. Topic drift, imprecise reference and highly figurative language often obfuscate reasoning and undermine argumentative coherence. Implicit premises and non sequitur render reasoning opaque for readers not already inclined to agree. Along with a failure to rigorously define conventions of semantics, grammar, syntax and punctuation, as well as spelling and pronunciation, these features may not make analytic reasoning entirely impossible. They do, however, make it more difficult, and far more subject to misunderstanding and misinterpretation. They detract from rigorous discourse, they do not assist it. And, in educational settings and communication requiring rigor, assistance matters.Standard language functions as a kind of intellectual scaffolding. Its stylistic constraints externalize cognitive discipline. They force writers to slow down, to specify what they mean, to anticipate objections, and to make inferential commitments explicit. This is extremely important in academic contexts, in the public domain, and in business. In this sense, standard language “carries the water” for analytic thought not because it is inherently more logical, but because it institutionalizes habits of clarity. To abandon those habits in the name of linguistic equivalence is a major category error. It confuses respect of language evolutionary realities and cultural diversity with utility and rigor. Castilian was once the vulgar language of disenfranchised peasants within the complex of the pax romana. Over millennia, it evolved into a highly logical discursive language spectacularly suited for analytical rigor. Standard language, agreed upon by all, fulfills this highly utilitarian function as well as aesthetic objectives. It does so by default; non-standard and informal registers generally do not - much as Basil Bernstein’s distinction between restricted (high-context, assumption-laden) and elaborated (explicit, low-context) codes – which highlights the institutional mismatch when schools deprecate traditional standards or allow expression for heterogeneous consumption exclusively in high context idioms.3The consequences of this confusion in the educational domain and the acceptance, indeed the encouragement, of nonstandard language are increasingly visible. Students are often encouraged to view demands for standard language as matters of taste. Worse, the teaching of standard language is framed – especially within the so-called “critical” language studies community – through an academic perspective that interprets language primarily through the lens of power, domination, and resistance. In this view, standard language is disparaged as the language of the hegemon, the colonialist. It must cede its place to the patois, the languages and dialectical peculiarities of the so-called disenfranchised. Language study and language teaching must no longer prioritize the “standard” language of power elites. Rather, enlightened language theory must be, to use the term of consensus, “disruptive”.In reading such statements, I rarely understand, specifically, what harms are being remediated and what, other than the value of clarity, coherence and logical discourse, is being disrupted. This iconoclasm is seen somehow as a meritorious or virtuous redress of historical injustices which must be corrected. It seems that among such theoreticians, little thought is given to the fact that standard language – because it is the language of the hegemon – has evolved precisely to provide tools for rational thought – as such rational discourse, and its logic, provides the framework for advancement of the social, scientific and artistic achievements of the dominant culture. Institutional centers of power have historically codified linguistic norms that support the forms of reasoning valued within those institutions (reasoning construed as “effective” for achieving the utilitarian objectives of their societies and cultural norms upon which the reigning societal edifices are erected). Instead of disparaging the teaching of the language which fosters access to these norms – including not only the logic, the grammar, the syntax and agreed upon semantics of the hegemonic dialect, but also the indispensable linguistic niceties which constitute the phatic web within which all productive discourse takes place (is experienced as inviting, nonthreatening and conducive to dialogue) –, some pedagogical approaches, in seeking to resist linguistic hierarchy, actually end limiting students’ access to the linguistic tools most rewarded by institutions. In so doing, I suggest they are being cognitively dissonant – they are subverting their own implied desire to provide access to power to those whom they encourage to disparage standard forms. Research in educational sociology and writing studies suggests that explicit instruction in elaborated codes improves students’ ability to produce decontextualized, analytically structured arguments.4,5In today‘s public discourse, including political discourse, regrettably there is a trend increasingly to favor immediacy and affect over argument. This can be seen in both social and legacy media. Commentary frequently substitutes assertion for analysis, confident that expressive force will compensate for logical thinness. In such an environment, the erosion of linguistic standards is not merely a cultural shift; it is an epistemic one.To insist on standard language in analytic domains and in the public discourse, and to require students to learn and to wield it effectively, and with rigor, is not to denigrate non-standard dialects or to deny their expressive richness. It is simply to recognize that different communicative goals require different linguistic disciplines. It is a recognition that, in fact, not all registers are interchangeable. We do not accuse mathematics of elitism because it demands symbolic precision, nor chemistry because it requires a specialized vocabulary. We understand that certain forms of inquiry demand particular tools. Language is no different.This recognition carries clear implications for education, academia, and public life. Students should be taught—explicitly and unapologetically—that mastery of standard language is not a betrayal of identity but an acquisition of intellectual power and access to what Lisa Delpit calls “the language of power”.5 Academic institutions should resist the temptation to relax linguistic standards under the mistaken belief that rigor is exclusionary. Public figures, especially those who aspire to persuade across differences, should model the disciplined use of language appropriate to analytic, business and civic reasoning. It is important to emphasize that nothing in this argument entails that nonstandard dialects are cognitively deficient or aesthetically inferior. Nor does it deny that standard language carries some historical associations with exclusion. The question at issue is not dignity, but function. Educational institutions must decide which linguistic norms best serve the development of transferable analytic competence. While analytic reasoning is in principle independent of linguistic form, the absence of externalized constraints increases cognitive load, particularly for novice thinkers. None of this requires suppressing dialects, policing informal speech, or denying the legitimacy of linguistic variation in its proper contexts. It requires only the courage to say what was once taken for granted: that clarity is not oppressive, and precision is not prejudice. Recognizing the functional value of standard language does not require denying expressive richness or cultural legitimacy in other forms. I think we are better served with a stance not advocating replacement, but rather breadth of repertoire. This allows for expressive richness while recalling that a shared standard language remains one of the most powerful instruments we possess for thinking together in common. NOTES1. Stephen Pinker, The Language Instinct: How the Mind Creates Language, 9th ed, Harper Perennial, 2007.2. Examples in the literature are myriad. A few: The new Journal from the U of Pennsylvania, Racial Justice in Multilingual Education(RJME), the first edition of which was published in August of 2025; The “Standard Language Ideology Statement”, published by the the Department of Linguistics at the University of Michigan, July 2021: https://lsa.umich.edu/linguistics/about-us/values-statement/standard-language-ideology-statement.In this latter, one reads the following: “Linguists do not support the widely held assumption that there is a standard language that should be adopted by all, and our department condemns penalties that come with not using such language. Standard language ideology is a construct that establishes a hierarchy between varieties. It misleads language users into believing that some varieties are better than others and can perpetuate harmful patterns of linguistic discrimination - discrimination that is often a proxy for ethnic, gender, class, and regional discrimination.”3. Basil Bernstein, Class, Codes and Control, Vol. 1 (Routledge, 1971); see also his 1962 article in American Anthropologist. Bernstein argues convincingly that elaborated codes are essential for formal reasoning and institutional success, and failing to teach them disadvantages students structurally, independent of any deficit in the restricted code itself.4. Pierre Bourdieu, Language and Symbolic Power, ed. John B. Thompson (Harvard University Press, 1991), esp. ”The Economics of Linguistic Exchanges” and ”Authorized Language.” Bourdieu argues that refusing to teach the dominant language in the name of equality deprives repressed groups of instruments for social mobility and institutional participation.5. Lisa Delpit, ”Skills and Other Dilemmas of a Progressive Black Educator,” Harvard Educational Review (1986); see also, Lisa Delpit, “The Silenced Dialogue: Power and Pedagogy in Educating Other People’s Children,” Harvard Educational Review 58:280–298 (1988). Delpit argues that explicit instruction in the ”language of power” (standard/edited English) is an ethical obligation, especially for marginalized students, and that withholding it is paternalistic rather than liberatory. This complements pedagogical work on writing as cognitive scaffolding, such as Flower and Hayes, “A Cognitive Process Theory of Writing.” College Composition and Communication, 1981.
In this article, we aim to highlight the correspondences between theAromanian dialect and the dialects of the Romance languages, especially those ofthe Italian language, based on the digital maps extracted from: Nicolae Saramandu,Manuela Nevaci, Atlasul lingvistic al dialectului aromân, vol. II, Editura AcademieiRomâne, Bucharest, 2020; Atlasul lingvistic român pe regiuni. Sinteză – ALRR. Sinteză,vol. III (coordinator: Nicolae Saramandu), Editura Academiei Române, Bucharest,2018 (authors: Mihaela Mariana Morcov, Manuela Nevaci, Irina Floarea, DanielaRăuțu, Carmen-Ioana Radu, Mara Iuliana Manta, Ionuț Geană); Lídia Pons i Griera,Joan Veny, L’Atles Lingüístic del Domini Català, vol. III, Familia Institut d'EstudisCatalans, Barcelona, 2001–2018; Karl Jaberg und Jakob Jud, Sprach- und SachatlasItaliens und der Südschweiz, (vol. I), Zofingen, Ringier, 1928–1940. For the productionof the maps, a digital program developed by CSII Dr. Vasile Apopei and CS II Dr.Silviu-Ioan Bejinariu, Institute of Theoretical Informatics Iași, of the RomanianAcademy, was used. In our analysis we used Map 279. STEPFATHER ‛beau-père’[469] from Atlasul lingvistic al dialectului aromân, vol. II, Editura Academiei Române,Bucharest, 2020 (authors: Nicolae Saramandu, Manuela Nevaci) and Map 329.STEPFATHER [469] from Atlasul lingvistic român pe regiuni. Sinteză – ALRR. Sinteză,vol. III (coordinator: Nicolae Saramandu), Editura Academiei Române, Bucharest.To compare the Aromanian dialect with Daco-Romanian and with the dialects ofRomance languages, such as Italian and Catalan, we used Map 275. Mother [464]from Atlasul lingvistic al dialectului aromân, vol. II, Editura Academiei Române,Bucharest, 2020 (authors: Nicolae Saramandu, Manuela Nevaci), Map 325. Mother[464] from Atlasul lingvistic român pe regiuni. Sinteză – ALRR. Sinteză, vol. III(coordinator: Nicolae Saramandu), Editura Academiei Române, Bucharest, Map 8.Sua madre from Karl Jaberg und Jakob Jud, Sprach- und Sachatlas Italiens und derSüdschweiz, (vol. I), Zofingen, Ringier, 1928–1940, and List 21. La mare (L’AtlesLingüístic del Domini Català, vol. III, Familia Institut d'Estudis Catalans, Barcelona,2001–2018, authors: Lídia Pons i Griera, Joan Veny). The information system usedfor producing the ALAR maps provides essential support in the automaticgeneration, editing, and printing of plates with linguistic maps. “The system usesdata and information grouped into two databases: the geographical database, whichcontains the template of the ALDRO linguistic maps, and the linguistic database,which stores all information related to words (notions), survey points, and thephonetic transcriptions associated with each word in all survey points” (Silviu-IoanBejinariu, UEFISCDI Report, PCE ALDRO, 2021). The AlarMaps program is easy toaccess for a linguist. In this program, we created the map with the localities wherewe conducted fieldwork. In my doctoral thesis, I will make 100 linguistic maps withsix new points (localities in Dobrogea, Tulcea and Constanța counties: Stejaru, Ceamurlia de Sus, Sinoe [Grămostean dialect]; Cogealac and Poiana [Fărșerotdialect]; and Techirghiol [Pindean dialect]). The new linguistic maps, created indigital format, will complete the maps of Aromanian dialects spoken in the Balkans.The recorded responses will be phonetically transcribed and interpreted within themaps. Below, we present the map with the six new points.
Abstract The International Affective Picture System (IAPS) is an extremely valuable resource and among the most widely used stimulus databases. Notably lacking though are separate measures for positive and negative emotion. Thus, researchers cannot distinguish images that evoke a neutral affective state from those that evoke both positive and negative emotion. Further, IAPS ratings are not provided for racial or ethnic groups, leaving it unclear how different groups respond. The current study aimed to aid in image selection by providing information beyond what is currently available. In Study 1, we tested the effects of the Self-Assessment Manikin, a 9-point, pictorial scale used to make the original normative IAPS ratings ( N = 115). When the visual scale was used, ratings were lower and had a restricted range, compared to use of a non-visual, Likert-type, 9-point scale. In Study 2 ( N = 1061), using this non-visual scale, we surveyed a racially and ethnically diverse sample. We provide mean ratings for positive emotion, negative emotion, and arousal, as well as an emotional categorization for each image (positive, negative, or no emotion). Ratings are provided by gender and for Hispanic/Latino/a, East Asian, Southeast Asian, and White, non-Hispanic groups as well as for the combined sample. Preliminary analyses show some group differences in ratings (e.g., East Asian > Hispanic for negative and arousal ratings for one cluster of images), suggesting that demographic variables are related to ratings. The ultimate goal of the current project is to facilitate advances in affective science by providing additional information about IAPS to improve stimuli selection.
Arabic WordNet 4.0 is a comprehensive lexical database for Modern Standard Arabic, derived from the Open English WordNet using the expand approach. **Features:**- 109,823 synsets (100% OEWN coverage)- 124,653 lexical entries- 166,643 senses- 265,676 synset relations- 97.2% ILI coverage for cross-linguistic linking- Full WN-LMF 1.4 XML format compliance **Methodology:**Translations were generated using AI-assisted translation (Google Gemini 3 Pro Preview) following the expand approach for WordNet construction. **Attribution:**This resource is derived from Open English WordNet (CC BY 4.0), which is based on Princeton WordNet 3.0.
Arabic WordNet 4.0 is a comprehensive lexical database for Modern Standard Arabic, derived from the Open English WordNet using the expand approach. Features:109,901 synsets (partial OEWN 2024 coverage — satellite adjectives not yet included)~124,653 lexical entries~166,843 senses~265,676 synset relations97.3% ILI coverage for cross-linguistic linkingFull WN-LMF 1.4 XML format compliance Note: This is an intermediate release (v4.0.1) representing AWN4 after upper-ontology noun additions (+78 synsets) but before satellite adjective coverage was completed. For the complete OEWN 2024 parity release (120,630 synsets), see v4.1.0. Methodology:Translations were generated using AI-assisted translation (Google Gemini 3 Pro Preview) following the expand approach for WordNet construction. Attribution: This resource is derived from Open English WordNet (CC BY 4.0), which is based on Princeton WordNet 3.0.
Arabic WordNet 4.0 is a comprehensive lexical database for Modern Standard Arabic, derived from the Open English WordNet using the expand approach. Features:109,901 synsets (partial OEWN 2024 coverage — satellite adjectives not yet included)~124,653 lexical entries~166,843 senses~265,676 synset relations97.3% ILI coverage for cross-linguistic linkingFull WN-LMF 1.4 XML format compliance Note: This is an intermediate release (v4.0.1) representing AWN4 after upper-ontology noun additions (+78 synsets) but before satellite adjective coverage was completed. For the complete OEWN 2024 parity release (120,630 synsets), see v4.1.0. Methodology:Translations were generated using AI-assisted translation (Google Gemini 3 Pro Preview) following the expand approach for WordNet construction. Attribution: This resource is derived from Open English WordNet (CC BY 4.0), which is based on Princeton WordNet 3.0.
Arabic WordNet 4.0 is a comprehensive lexical database for Modern Standard Arabic, derived from the Open English WordNet 2024 using the expand approach. Features:120,630 synsets (full OEWN 2024 parity)136,041 lexical entries184,238 senses297,150 synset relations (0 skipped — exact parity with OEWN 2024)97.3% ILI coverage for cross-linguistic linkingFull WN-LMF 1.4 XML format complianceAll synsets include Arabic definitions with full tashkeel (diacritical marks) on lemmas What's new in v4.1.0:+10,720 satellite adjectives (pos=s) — completing full OEWN 2024 adjective coverage+9 missing hub verbs (act/move, change, travel, make, communicate, and others)+78 upper-ontology noun synsets completing the noun hierarchyAll 8 validation checks pass against OEWN 2024 Methodology:Initial 109,823 synsets (nouns, verbs, adjectives, adverbs) were generated using AI-assisted translation (Google Gemini 3 Pro Preview). The remaining 10,807 synsets (satellite adjectives, hub verbs, upper-ontology nouns) were translated using Anthropic Claude via an automated Docker pipeline. Attribution:Derived from Open English WordNet 2024 (https://en-word.net/) and Princeton WordNet 3.0 (https://wordnet.princeton.edu/), both licensed under CC BY 4.0.
This dataset contains the coded lexical and contextual data used in a corpus-informed analysis of emotion-related vocabulary in Italian as a foreign language (IFL) textbooks at the beginner level (CEFR A1) used in Polish lower secondary education. The dataset is based on four textbooks from two series: Progetto Italiano Junior (Marin, 2017; 2018) and Va bene! (Kaliska & Kostecka-Szewc, 2021). All materials were analysed in their printed form. The dataset includes all lexical items identified as emotion-related based on their presence in the ANEW-IT database (Montefinese et al., 2014), which provides normative ratings of valence and arousal for Italian words. Each entry in the dataset corresponds to a single token occurrence of an emotion-related lexical item. The dataset includes both surface forms as they appear in the textbooks and their corresponding base forms (lemmas) as listed in ANEW-IT. For each token, the dataset provides contextual, linguistic, and affective information. The variables included are: • textbook and series identification • unit/chapter, page number, and exercise reference • material type (e.g., dialogue, reading text, exercise, review) • pedagogical focus (e.g., grammar, vocabulary, comprehension, mixed) • word form (surface form) and lemma • part of speech (POS) • contextual sentence or description of occurrence • frequency measures (FreqColfis, Ln_Colfis) • affective ratings (valence and arousal, scale 1–9) The dataset enables replication of the quantitative analyses reported in the study, including token counts, type–token ratios, valence and arousal distributions, and comparisons across textbook series and pedagogical contexts.
This article examines the role of advertisements and signboards in shaping and reflecting public attitudes toward language. In modern society, linguistic culture is not only preserved in literature and education, but also manifested in everyday public texts such as commercial advertisements, street signs, shop names, and information boards. The study analyzes the linguistic quality of advertising texts, the influence of globalization on language use, and the social consequences of neglecting linguistic norms. Special attention is given to the relationship between language accuracy and cultural identity. The article also discusses the responsibility of businesses, media representatives, and educational institutions in maintaining linguistic standards in public communication.
The Spoken Corpus of the Southern Dutch Dialects (GCND) is a linguistically annotated, audio-aligned corpus of 653 recordings from 639 locations across Belgium, northern France, and the southern Netherlands. Most recordings (623) date from 1963–1976 and involve speakers born around 1900 (beginning 1871); 30 new recordings (2020–2024) fill geographic gaps. All recordings were transcribed using a two-tier orthographic protocol and annotated with POS tags, lemmas, and syntactic parses via the Alpino parser, supplemented by manual pre- and post-processing. The corpus is searchable through the Dutch Language Institute (INT) using the BlackLab and GrETEL interfaces. GCND forms an unprecedented resource for research on Dutch dialect variation, spoken corpus annotation, but also dialect-sensitive speech technology. Words, lemmata, metadata, and syntactic structures are accessible for non-commercial research use. The search interface of the corpus can be accessed via https://hdl.handle.net/10032/tm-a2-z8 (CLARIN login required). It is hosted by the Institute for the Dutch Language (www.ivdnt.org). Via a web-based frontend (https://blacklab-frontend.ivdnt.org/), users can search transcription tiers, metadata, POS tags, and syntactic structures through a dual platform consisting of BlackLab (for token‑based search)(https://blacklab.ivdnt.org/) and GrETEL (for treebank queries). In this OSF repository, you can find 1. The project documentation, including the syntactic annotation protocol 2. The transcription protocol 3. The metadata of the recordings and the recorded speakers The metadata file contains two tabs: (1) "Opname" ('recording') and (2) "Spreker" ('speaker'). Under (1), each recording is identified by a code for the location, called the Kloeke-code (after the dialectologist G.G. Kloeke who developed the system in the 1920s), which is the standard way in Dutch dialectology. This is further extended by a number with an underscore, as there may be more than one recordings from one location. For instance, I175p_4 is the 4th recording from Sint Niklaas. There is furthermore information on the dialect region, the province, and which collection the recording is from (UGent or Meertens Institute), whether the place is also a sampling place of earlier dialect surveys (RND and SAND), and the date of the recording. Under tab (2), there is information on the speakers, of which there may be more than one per recording. On the recording I175p_4 from Sint Niklaas, for instance, there were two speakers, identified by adding a number with an underscore to the ID of the recording: I175p_4_1 and I175p_4_2. The tab on the speakers furthermore contains information on the education, profession, and mobility of the speaker, as well as information on the mobility of their parents and spouse.
The Darbest dataset is a Universal Dependencies (UD) treebank dataset for Standard Sorani Kurdish written in the Perso-Arabic script. It contains 69,000 annotated sentences and 1,205,855 tokens collected from nine textual domains. The corpus was collected from seven Kurdish online news websites and supplemented with texts from published books. Before preprocessing, the collected corpus contained 1,250,275 words from 5,627 web pages together with book-based texts and was preprocessed using a Python-based pipeline involving text cleaning, Unicode and punctuation normalization, sentence segmentation, and tokenization. The dataset was developed to provide a large-scale syntactically and morphologically annotated resource for Standard Sorani Kurdish. A separate 100-sentence gold-standard set was manually annotated according to the Universal Dependencies v2 guidelines. Sorani Kurdish linguistic experts supported the selection of sentences representing diverse and linguistically complex structures and reviewed LLM-generated annotations for errors. The gold-standard set was used to construct few-shot prompts for annotating the remaining corpus. The resulting annotations were represented in the standard CoNLL-U format and validated using the official Universal Dependencies validation tool, followed by manual correction and quality review. The released treebank is divided into training, development, and test sets and includes lemmas, Universal Part-of-Speech (UPOS) tags, morphological features, syntactic heads, and dependency relations. The dataset can be used to train, evaluate, and benchmark NLP models for part-of-speech tagging, lemmatization, morphological analysis, dependency parsing, and related computational linguistics tasks. The accompanying repository also contains the separate 100-sentence gold-standard set, plain-text corpus splits, README documentation, and a dataset statistics spreadsheet.
Abstract This study examines associations between brain levels of amyloid-β and tau with representational pattern similarity in amygdalar reactivity to negative and neutral images in older adults without dementia. 81 participants viewed affective images during functional magnetic resonance imaging (fMRI). Participants rated the images on valence and arousal outside the scanner. Amyloid-β and tau were measured with 11 C-Pittsburgh compound B and 18 F-MK-6240 positron emission tomography (PET) imaging. Representational pattern similarity analyses compared amygdalar responses to negative and neutral stimuli, while preserving the voxel-wise pattern of fMRI activation. Greater differentiation in the pattern of responding across the amygdala was indicated by lower similarity between negative and neutral stimuli. Greater tau levels in the entorhinal cortex and amygdalae were associated with less pattern similarity in left amygdalar reactivity to negative and neutral images in participants whose tau level was below standard positivity thresholds. Less pattern similarity in right amygdalar reactivity to negative and neutral images was also associated with the participants’ valence and arousal ratings of the stimuli. No associations were found with global amyloid levels and only the association between entorhinal tau and amygdalar similarity remained significant after Bonferroni correction. Greater amygdalar pattern separation in response to negative and neutral stimuli with higher levels of entorhinal tau suggest greater amygdalar sensitivity to negative compared to neutral information as tau accumulates in the brain, suggesting a potential underlying mechanism for the disrupted emotional processes often observed in preclinical Alzheimer’s Disease and other Related Dementias.
This paper presents a direct framework for sequence models with hidden states on closed subgroups of U(d). We use a minimal axiomatic setup and derive recurrent and transformer templates from a shared skeleton in which subgroup choice acts as a drop-in replacement for state space, tangent projection, and update map. We then specialize to O(d) and evaluate orthogonal-state RNN and transformer models on Tiny Shakespeare and Penn Treebank under parameter-matched settings. We also report a general linear-mixing extension in tangent space, which applies across subgroup choices and improves finite-budget performance in the current O(d) experiments.
Negative biases in emotional processing are central to cognitive models of Major Depressive Disorder (MDD), yet it remains unclear whether negatively biased valence attribution covaries with depressive symptoms or represents a more stable vulnerability. In this longitudinal study, we examined emotional valence ratings during a subliminal affective priming task in n = 232 MDD patients and n = 496 healthy controls (HC's) across multiple follow-up assessments, yielding 1395 observations. Multilevel models tested effects of depressive symptom severity, diagnostic group (MDD vs. HC), affective prime (happy, sad, neutral), and their interactions. To separate within-person fluctuations from between-person differences, depressive symptom severity was decomposed (Mundlak within-between specification). Exploratory analyses examined whether long-term symptom trajectory clusters were associated with emotional bias. Greater depressive symptom severity was associated with more negative valence ratings. Within-between analyses indicated that this association was primarily driven by within-person symptom fluctuations, indicating more negative ratings when symptom severity exceeded individual average. In contrast, diagnostic group, prime condition, and symptom severity × prime interactions did not significantly predict valence ratings, and no group differences emerged between HC's and remitted MDD patients. Exploratory trajectory analyses suggested that patients with recurrent symptom patterns may show more negative valence ratings, although cluster stability was limited and findings should be interpreted cautiously. Overall, the results suggest that negative valence attribution in MDD is more closely linked to current depressive symptom severity than to diagnostic status alone. Longitudinal approaches separating within- and between-person processes may help clarify when negative biases emerge and persist during depression progression.
This study has made new contributions to Pauline epistolography by applying the results of sociolinguistic and semantic analyses of the remembrance motif and litotic disclosure formula in epistolary papyri to Paul’s forms of the formulae using an interpretive lens of lexical norms and exploitations. The analyses presented in the study reveal hitherto unnoticed elements in Paul’s communicative strategy and demonstrate how a sociolinguistic approach to documentary papyri opens up new avenues for NT research. Chapter 12 summarises this tripartite study, discusses its significance for Pauline epistolography, and suggests potential avenues of further research.
Understanding how memories of past experiences shape subjective feelings is complicated by the fact that we constantly update our memories. These updates are particularly impactful when individuals are reminded of emotionally positive or negative attributes of the original event. Yet, it remains unclear how such memory updating influences subjective feelings. Here, we investigated how the reactivation of emotional information affects episodic memory, subjective feelings, and their interaction. Across three experiments, participants first learned both positive and negative attributes associated with unfamiliar individuals. Then, they were reminded of a single positive or negative attribute for each individual to reactivate the memory partially. Finally, we reassessed memory for and subjective feelings about each individual’s attributes. In Experiments 1 and 2, these procedures were distributed across three days, while in Experiment 3, they occurred on a single day. Across these three experiments, reminding with negative attributes shifted subjective feelings in a negative direction. Reminded attributes were also better remembered, particularly for negative ones, and changes in subjective feelings were more strongly associated with reminded attributes. However, positive reminders only influenced subjective feelings to change positively when all procedures occurred on the same day. Together, these findings support a model in which memory updating shapes both episodic memory and emotional experience in a valence-dependent manner.