Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Abstract This paper presents a novel treebank-driven approach to comparing syntactic structures in speech and writing using dependency-parsed corpora. Adopting a fully inductive, bottom-up method, we define syntactic structures as delexicalized dependency (sub)trees and extract them from spoken and written Universal Dependencies (UD) treebanks in two syntactically distinct languages, English and Slovenian. For each corpus, we analyze the size, diversity, and distribution of syntactic inventories, their overlap across modalities, and the structures most characteristic of speech. Results show that, across both languages, spoken corpora contain fewer and less diverse syntactic structures than their written counterparts, with consistent cross-linguistic preferences for certain structural types across modalities. Strikingly, the overlap between spoken and written syntactic inventories is very limited: most structures attested in speech do not occur in writing, pointing to modality-specific preferences in syntactic organization that reflect the distinct demands of real-time interaction and elaborated writing. This contrast is further supported by a keyness analysis of the most frequent speech-specific structures, which highlights patterns associated with interactivity, context-grounding, and economy of expression. We argue that this scalable, language-independent framework offers a useful general method for systematically studying syntactic variation across corpora, laying the groundwork for more comprehensive data-driven theories of grammar in use.
We release a sentence-level genre layer for Universal Dependencies as a separate, joinable dataset, computed across UD revisions and linked back to the underlying treebanks via a release-aware composite key comprising treebank, split, sent_id, and UD release metadata.The annotations are derived rather than authoritative and are accompanied by provenance and uncertainty indicators, enabling downstream users to choose appropriate precision-coverage trade-offs and to re-run the pipeline as UD evolves.To support both parity tracking and deployment-oriented interpretation, we report results under two complementary regimes: a fixed-partition setting aligned with earlier protocols, and a language-grouped 10-fold generalisation setting that highlights cross-language heterogeneity and anchor sparsity as operational constraints.The resulting resource is intended to make genre a practical control variable for UD-based experimentation, including genre-stratified evaluation and training data selection for POS tagging and parsing, where performance varies substantially across text types.Finally, we note that reduced genre spaces aligned with recurring robustness profiles (e.g.transcribed speech versus interactional web/social text versus edited prose/news) appear pragmatically useful, but should be treated as a community coordination task implemented through explicit, versioned mapping tables.
Thermal imaging, which is contact-free, light-independent, and effective in detecting skin temperature changes that reflect autonomic nervous system activity, is expected to be useful for emotion sensing. A recent thermography study demonstrated a linear relationship between ear temperatures and emotional arousal ratings. However, whether and how ear thermal changes may be nonlinearly related to subjective emotions remains untested. To address this issue, we reanalyzed a dataset that included ear thermal images and self-reported arousal ratings obtained while participants watched emotion-eliciting films. We employed linear regression and two nonlinear machine learning models: a random forest model and a ResNet-50 convolutional neural network. Model evaluation using mean squared error and correlation coefficients between actual arousal ratings and model predictions indicated that both machine learning models outperformed linear regression and that the ResNet-50 model outperformed the random forest model. Interpretation of the ResNet-50 model using Gradient-weighted Class Activation Mapping and Shapley additive explanation methods revealed nonlinear associations between temperature changes in specific ear regions and subjective arousal ratings. These findings imply that ear thermal imaging combined with machine learning, particularly deep learning, holds promise for emotion sensing.
Using semantic dependency analysis, this study examines narrative productions from Mandarin-speaking preschool children aged three to six to investigate how semantic organization develops with age in early childhood. Four semantic dependency treebanks were constructed from a Chinese narrative corpus available in the CHILDES database. By comparing semantic dependency types and semantic dependency distances across the four age groups, we found that (1) semantic organization shifted from experiencer and classification relations toward agent and patient relations, situational-role relations (particularly those involving measurement, individuation, and direction), and structural relations; (2) mean semantic dependency distance (MSDD) increased with age, as adjacent dependencies decreased and longer dependencies became more frequent. This increase in MSDD indicates growing complexity in semantic organization and is largely driven by significant increases in the MSDD values of specific semantic dependency types. These findings provide new evidence for semantic organization development in preschool children.
Background: Emotion processing is critical in the neuropathology of major depressive disorder (MDD), while its relationship with clinical treatment remains unclear. This study aims to indicate the associations between emotion processing and treatment effects following a sequential dual-site accelerated repetitive transcranial magnetic stimulation (rTMS) protocol. Methods: MDD patients were recruited to receive rTMS treatment with four sessions per day for four consecutive days, with stimulation sequentially delivered to the left dorsolateral prefrontal cortex (dlPFC) and the dorsomedial prefrontal cortex (dmPFC). Symptoms were assessed at baseline, end of treatment, and week 4 using the Montgomery–Åsberg Depression Rating Scale (MADRS), Snaith-Hamilton Pleasure Scale (SHAPS), and Fatigue Severity Scale (FSS). Emotional valence and arousal were evaluated with the Affect Rating Task (ART). Results: A total of 51 participants completed the clinical assessments and ART, with two excluded due to missing baseline data in the SHAPS and FSS. The linear mixed-effects models revealed significant improvement in depressive (p < 0.001, d = −0.343) and fatigue symptoms (p = 0.010, d = −0.572) following rTMS treatment. Neutral valence was correlated with MADRS scores at baseline (R2 = 0.096, p = 0.027). In addition, changes in arousal for positive images (p = 0.047, adjusted R2 = 0.097) and neutral images (p = 0.019, adjusted R2 = 0.160) at treatment end were significantly correlated with MADRS improvement at week 4. Conclusions: Our study highlights the association between changes in emotional arousal and improvement in MDD following accelerated dlPFC-dmPFC dual-site rTMS treatment.
This article evaluates the integration of data extracted from a French syntactic lexicon, the Lexicon-Grammar (Gross, 1994), into a probabilistic parser. We show that by applying clustering methods on verbs of the French Treebank (Abeillé et al., 2003), we obtain accurate performances on French with a parser based on a Probabilistic Context-Free Grammar (Petrov et al., 2006).
ABSTRACT Objectives This study examined therapist–client physiological synchronisation, indexed by heart rate variability (HRV), in meditation versus emotionally supportive sessions (control condition) and its association with clients' affective evaluations. Methods This was a randomised experimental study. Participants were assigned to a meditation session ( n = 20) or an emotional support session ( n = 20). HRV was assessed at baseline and post‐session; an additional mid‐session HRV assessment was obtained in the meditation condition immediately following the meditation segment. Clients reported positive and negative affect and evaluated the session. Results No significant therapist–client HRV correlations were observed at baseline or at the end of the session in either condition. However, in the meditation condition, changes in therapists' HRV from baseline to mid‐session and from baseline to post‐session were associated with clients' affective evaluations, whereas no such associations emerged in the control condition. Finally, a factor analysis revealed two non‐coherent factors of therapists' and clients' HRV before and after the session in controls, and a single factor including all HRV measures of therapists and clients in the meditation condition. Conclusions Therapist autonomic regulation during meditation may relate to clients' immediate emotional experience. Further research using longitudinal and dynamic designs is needed to clarify mechanisms of physiological synchronisation in psychotherapy.
<div> This paper examines register variation in Latin from the third century BCE to the fourteenth century CE using Key Feature Analysis (KFA) (Egbert and Biber, 2023), a quantitative method for identifying statistically over-and underrepresented linguistic features. Registers are defined as text varieties linked to communicative situations and characterized by distributions of lexico-grammatical features (Biber, 1988, 1995). Six dependency-parsed Universal Dependencies (UD) treebanks are classified a priori into ten register categories based on established scholarship. Additionally, Principal Component Analysis (PCA) is used to reduce dimensionality, in order to explore the texts major patterns of variation and clusters of linguistically similar texts. KFA reveals systematic register-specific grammatical profiles consistent with previous research (e.g. Biber (2014b)). Registers with involved language use (e.g. letters and speeches) show higher frequencies of personal reference and engagement, while philosophical texts favor subordination and impersonal constructions. Registers containing narrative elements (e.g. historiography, satire) contain high frequency of verbs in past tense. PCA places the charter register into a distinct cluster, while other registers form more closely grouped patterns. The strongest components reflect contrasts in number, aspect, tense and person, alongside subordination and cordination distributions. The results are largely confirmatory: KFA produces coherent and interpretable groupings of grammatical features consistent with previous findings, providing a proof of concept for quantitative register analysis in historical corpora. Data and code are openly available for future research. </div>
Code-mixed text from social media poses significant challenges for syntactic analysis due to irregular grammar, non-standard usage, and frequent language switching. For Telugu-English code-mixed text, the absence of large-scale syntactic resources and specialized parsing models limits progress in downstream multilingual NLP applications. In this work, we address this gap by introducing the first substantial manually annotated Telugu-English code-mixed dependency treebank of 4,152 sentences, developed using Universal Dependencies (UD) 2.0 guidelines. We further propose enhancements to a biaffine dependency parser by incorporating a language-aware head-dependent bias and relation-specific structural weights to better capture cross-lingual syntactic patterns. Our approach improves parsing performance, achieving 75.53% UAS and 61.86% LAS, with consistent gains over a strong baseline. In addition, we demonstrate that integrating dependency-derived syntactic features into a BiLSTM-CRF model improves part-of-speech tagging, achieving a macro-F1 score of 83.73%, with statistically validated gains. We also re-annotate an existing Telugu-English dataset using UD 2.0 to ensure compatibility with modern syntactic frameworks. Overall, this work provides new annotated resources and modeling strategies that advance syntactic processing for Telugu-English code-mixed text, with broader implications for developing robust NLP systems in low-resource and multilingual settings.
The Dictionary of Jutlandic dialects (Jysk Ordbog, JO) documents the dialects of Jutland and its surrounding islands as they were primarily spoken by the rural population during the period 1870–1930. As a historical documentation dictionary, JO is descriptive in its practical lexicographic foundation. In JO’s meta-text and descriptive language, reference is made to a linguistic norm, often referred to as rigsdansk (Standard Danish), although the term is used ambiguously in JO. In some cases, it refers to Danish orthographic conventions as outlined by The Danish Language Council, and therefore often shares headwords with The Danish Dictionary (Den Danske Ordbog). In other cases, rigsdansk refers to a norm in spoken Danish, but as there is none, this is far more ambiguous; the term as well as the development of the Danish spoken language community is shortly accounted for. The article investigates the use of rigsdansk in JO as well as what it refers to; it proposes a more consistent use of rigsdansk as a linguistic norm against which the traditional Jutlandic dialects are evaluated. More clarity in both meta-text and descriptive language will enable users to get the most out of JO at any given time.
This article examines the role and significance of language corpora and linguistic databases in linguistic expertise. It discusses the possibilities of conducting objective semantic, pragmatic, and stylistic analyses of texts through corpus-based methods. The study also highlights the contribution of linguistic databases and artificial intelligence technologies to improving the accuracy, reliability, and efficiency of expert conclusions. Furthermore, the relevance of developing specialized corpora and databases for forensic linguistics in Uzbekistan is substantiated.
Surzhyk, as a Ukrainian–Russian mixed subcode, has long been a subject of extensive research. Despite this, it remains an ambiguous phenomenon that elicits diverse assessments and attitudes. Within Ukrainian academic circles, a tradition has emerged of evaluating this communicative subcode through the prism of linguistic norms, predominantly characterising it as a negative phenomenon. However, in everyday communication, Surzhyk functions dynamically, reflecting the varied attitudes of its speakers. Significant shifts in its perception became particularly evident following the Russian Federation's full-scale invasion of Ukraine. Consequently, Russian-speaking individuals have increasingly adopted Surzhyk as a means of distancing themselves from the Russian language, which is widely perceived as the language of the aggressor. This article presents the results of a sociolinguistic study conducted in 2020–2021 and 2023–2024 among residents of the Odesa and Mykolaiv regions. The study examines the evolution of attitudes towards Surzhyk in the wake of the invasion, with particular attention given to the role of mass media discourse in shaping these perceptions. The analysis reveals a growing tendency towards a more favourable view of mixed speech; in the context of war, Surzhyk has come to symbolise solidarity and serves as an informal linguistic bridge for transitioning from Russian-dominant to Ukrainian-speaking communication.
In today’s society, social media has become a site for grammatical and orthographic errors in comments under Uzbek Instagram posts, providing a window into how informal digital communication influences written language use. The findings reveal that social media discourse contains a high frequency of deviations from standard Uzbek, with subject–verb agreement errors emerging as the most common category, followed by sentence fragments and spelling inaccuracies. In addition, phonetic spelling, vowel substitution, and consonant confusion occur. The results demonstrate that users often transfer oral speech patterns directly into writing and grammatical accuracy. Such tendencies contribute to the normalization of non-standard forms in online environments. These patterns highlight the strong impact of digital practices on contemporary Uzbek writing and reflect a gradual shift toward informal linguistic norms in public online communication. While platforms like Instagram promote spontaneous expression, they also play a significant role in shaping language habits, particularly among younger users. Therefore, raising awareness about standard language norms and integrating digital literacy into education may help preserve grammatical accuracy and support clearer communication in Uzbek online discourse.
Monolingualism, native-speakerism and standard language ideology have been identified as dominant ideologies in language teaching with severe effects on second language teacher identities. Such ideologies offer alleged certainties but also detach teachers from the actual uses and value of language in multilingual and multidialectal contexts. As a consequence, educators might feel constrained by rigid linguistic norms, hindering their capacity to re-evaluate their approaches to accommodate the diverse linguistic realities and communicative needs of English learners in an increasingly interconnected and multilingual global landscape. This chapter intends first to offer a broad perspective of how these ideologies have shaped language teaching and how they have clashed with research-based observations of multilingual and multidialectal communicative settings. We will give an overview of the relevant literature, ranging from foundational texts to more recent ones challenging the ‘ideal’ monolingual native speaker, and we will show how the above ideologies are still found in a rather pervasive way in the language teacher profession. This will be followed by an account of recent research conducted in teacher training environments aimed at showing ways to successfully gear future language teachers towards a new vision of language that contemplates diversity and hybridization as fundamental pillars on which teachers’ identities need to be based.
This article examines the influence of user-generated content (UGC) on the style and credibility of regional news posted in public messaging channels. The aim of the study is to identify how active audience engagement in content creation transforms professional journalism standards, alters linguistic norms, and creates new challenges for information verification. This research contributes to the field of media studies by offering a systematic analysis of hybrid news models, where professional journalists and users jointly shape the agenda. This is particularly relevant given the growing influence of messaging apps as a key source of regional news. The theoretical basis for the study is based on works on participatory journalism, the digital transformation of media, and the ethical challenges of converging professional and user-generated content. The methodological framework includes content analysis, discourse analysis, and a comparative approach, enabling a comprehensive examination of stylistic changes, verification mechanisms, and editorial strategies for working with UGC. The empirical data consisted of 600 publications from three of Nizhny Novgorod’s largest news channels - “My Nizhny Novgorod,” “Nizhny Novgorod Without Censorship,” and “Ni Mash” - from April to May 2025. The study’s results demonstrate that user-generated content significantly influences news style, favoring informal, emotionally charged language, colloquialisms, and emojis. Furthermore, the reliability of information is often compromised by insufficient verification of user-submitted materials, as well as by a lack of clear source labeling. In some cases, elements of hate speech and aggressive verbal aggression are observed, particularly in materials with a high proportion of user-generated content.
Background and Aims: Most men consume pornography, with a small but significant percentage losing control over their use. Since ICD-11, problematic pornography use can be diagnosed as "compulsive sexual behavior disorder." Debate persists on whether problematic pornography use is an impulse-control disorder or a behavioral addiction. Mechanisms of learning and memory play a central role in addictive disorders but are presumably less relevant for impulse control disorders. Methods: One hundred thirty-nine heterosexual male users of pornography and gaming participated in our study which was part of a multi-center research project on internet use disorders in Germany. We focus on a subsample of fifty-eight non-problematic (n = 35) and problematic pornography users (n = 23, labeled pathological). FMRI data were collected during appetitive conditioning, extinction and recall. Pornographic, game, and money images served as unconditioned stimuli, geometric shapes as conditioned stimuli (CS). Results: During appetitive conditioning pathological pornography users showed a generally stronger response in ventral striatum to all CSs, whereas altered activations in extinction and recall were specific to the porn-associated CS. Greater activations in the dorsal anterior cingulate cortex during extinction and in the medial orbitofrontal cortex during recall suggest persistence of appetitive memory for pornography in pathological users, supported by valence ratings and skin conductance responses (SCR). Sensitization to the monetary cue also emerged in SCR. Discussion and Conclusions: Based on these new neurobiological findings, which are consistent with current addiction theories about stimulus-specific altered reward sensitivity and appetitive memory, we argue that problematic pornography use should be considered a behavioral addiction.
BACKGROUND: L. (caraway) essential oils (EOs) on aging. First, we assessed, in 402 participants, the age-related changes in olfactory functions (odor threshold, discrimination, and identification), gustatory perceptions (sweet, sour, salty, and bitter taste), cognitive functions (focusing on attention, memory, language, and visuospatial/executive functions), and their possible correlations with aging. To achieve this, olfactory function, gustatory perception, and cognitive abilities were evaluated in healthy participants across different age groups. Then, to evaluate the age-related decrease in trigeminal function (59 participants), we used rosemary and caraway EOs that contain carvone, limonene, and 1,8-cineole, all of which are considered typical trigeminal stimuli. METHODS: Olfactory function was assessed with the Sniffin' Sticks test, gustatory function by the Taste Strips test, and rosemary and caraway EOs by the ratings of odor pleasantness, intensity, and familiarity using a labeled hedonic Likert-type scale. RESULTS: Olfactory function could be a potential early indicator of attentional, memory, language, and visuospatial/executive dysfunctions. Our data indicated that rosemary and caraway EOs were perceived without any significant decrease in odor pleasantness, intensity, and familiarity ratings in relation to aging. CONCLUSION: Our results suggest the potential bioactive effects of rosemary and caraway natural EOs as a new strategy to promote healthy aging.
Prediction systems grounded in textual data have become indispensable across high-stakes domains including clinical decision support, financial signal detection, and digital misinformation analysis. Classical statistical approaches and shallow machine learning methods have demonstrated satisfactory performance on narrow, well-curated datasets, but they struggle to generalise once input distributions shift or domain vocabulary diverges from training corpora. Deep learning, and more specifically the pre-trained transformer paradigm, has substantially narrowed this gap; nevertheless, single-architecture solutions routinely leave accuracy on the table when applied to tasks that demand both rich contextual encoding and explicit sequential reasoning. This paper presents a cohesive, end-to-end AI- powered prediction framework that fuses BERT- derived contextual representations with a two-layer bidirectional LSTM (BiLSTM) classification head augmented by an additive attention mechanism. The system is designed as a modular pipeline: text acquisition and normalisation, augmentation-based imbalance handling, deep encoding, sequential modelling, and post-hoc probability calibration are treated as independent, replaceable stages. Experimental evaluation across three publicly available benchmark datasets — the LIAR fake news corpus, Stanford Sentiment Treebank v2, and a health-claim verification collection — confirms that the hybrid BERT-BiLSTM-Attention architecture outperforms five competitive baselines on macro-averaged F1 and area under the ROC curve. Ablation experiments quantify the individual contributions of the attention layer, recurrent head, augmentation strategy, and temperature scaling. A discussion of deployment trade-offs addresses inference latency, continual adaptation, and algorithmic fairness..
Screen use pervades daily life, shaping work, leisure, and social connections while raising concerns for digital wellbeing. Yet, reducing screen time alone risks oversimplifying technology’s role and neglecting its potential for meaningful engagement. We posit self-awareness—reflecting on one’s digital behavior—as a critical pathway to digital wellbeing. We developed WellScreen, a lightweight probe that scaffolds daily reflection by asking people to estimate and report smartphone use. In a two-week deployment with college students (\(\mathtt {N}\)=25) focused on generating formative insights, we examined how discrepancies between estimated and actual usage shaped digital awareness and wellbeing. Participants often underestimated productivity and social media while overestimating entertainment app use. They showed a 10% improvement in positive affect, rating WellScreen as moderately useful. Interviews revealed that structured reflection supported recognition of patterns, adjustment of expectations, and more intentional engagement with technology. Our findings highlight the promise of lightweight reflective interventions for supporting self-awareness and intentional digital engagement, offering implications for designing digital wellbeing tools.
The naming of professions and job titles reflects not only linguistic norms but also social structure, cultural values, historical development, and gender ideology of a society. This article examines the specific features of professional and occupational naming systems in the Russian and Uzbek languages. Special attention is paid to morphological, semantic, grammatical, and sociolinguistic aspects of profession names, including gender marking, borrowing processes, word-formation models, and modernization trends. A comparative analysis reveals both shared characteristics and significant differences conditioned by typological distinctions between the Slavic and Turkic language families. The study demonstrates that professional nomenclature functions as a dynamic linguistic subsystem closely connected with societal changes.
Abstract Arousal and valence are fundamental dimensions of affective experience signifying levels of activation and pleasantness, respectively. These dimensions play a crucial role in shaping emotional responses and behaviors, with significant implications for psychopathology. Previous machine learning studies had some success decoding these states from brain activation patterns observed during task-based functional magnetic resonance imaging (fMRI), but the results have varied across studies. Moreover, prior studies have often been limited by small sample sizes, weak decoding performance, and non-whole-brain analyses, leaving the neural representations of arousal and valence largely unresolved. Here we successfully decoded arousal and valence from whole-brain task-fMRI data collected from 132 participants during exposure to 300 unique emotional stimuli, including 150 movie clips and 150 text scenarios that reliably induced a wide range of arousal and valence states. Mass univariate general linear models identified block-level activation (emotion stimuli > washout) from all gray matter voxels. Multivariate regression analysis predicted arousal and valence ratings based on these gray matter activations. Patterns in the fMRI data underlying arousal and valence were robust, as they were successfully decoded across both induction modalities using five different linear multivariate regression models. Although significant, decoding from scenarios was less successful than from movies, likely due to their more imaginative nature. In particular, decoding arousal from scenarios only showed low predictive utility. Representations of arousal and valence were widespread throughout the brain, and we reveal cerebellar and brainstem contributions that have largely been absent in past fMRI decoding studies. These findings clarify the distributed neural basis of arousal and valence and provide a foundation for future clinical research on the role of these constructs in affective dysregulation.
Part of speech and syntactically annotated dataset for modern Mongolian. The dataset is a fully annotated corpus of modern Mongolian texts written in Mongolian Cyrillic.
While the influence of state-dependent factors on appetitive processing has received considerable attention, the role of stable personality traits remains comparatively unexplored. Extraversion, characterized by heightened positive emotionality, represents a compelling candidate in this regard, as it may shape individual differences in Positive Valence System (PVS) functioning. The present study examined how extraversion modulates neural and subjective responses to pleasant stimuli; the role of neuroticism was additionally explored, given its established association with affective reactivity. Sixty-eight Italian university students (40 females) completed an online version of the Big Five Inventory (BFI-44) before the laboratory session. Then, participants completed a passive viewing task of pleasant and neutral images while undergoing an electroencephalographic (EEG) recording. Appetitive stimulus processing was indexed by the peak amplitude of the P300-LPP complex and subjective SAM ratings. Results revealed that extraversion was positively associated with larger P300-LPP complex amplitudes to pleasant relative to neutral stimuli and with higher arousal ratings across both emotional categories. Additionally, neuroticism was associated with lower valence ratings regardless of stimulus category, with no significant effect on neural responses to emotional stimuli. These findings highlight extraversion as a stable personality trait shaping PVS functioning. Specifically, low extraversion was associated with reduced P300–LPP amplitudes and lower arousal ratings to pleasant stimuli, paralleling neural patterns documented in psychopathological conditions involving blunted PVS activation. These results underscore the utility of ERP-based measures in capturing personality-related differences in appetitive processing relevant to psychopathology risk.
This dataset contains imageability and familiarity ratings for Ukrainian and English work-related proverbs collected from Ukrainian university students. The data were gathered as part of a cross-linguistic study examining how bodily grounding influences the mental imagery associated with proverbial expressions in a first language (L1) and a second language (L2). The participants (N = 49) were students at Vasyl’ Stus Donetsk National University. Ukrainian was their first language (L1), and English was their second language (L2). Participants evaluated Ukrainian and English work-related proverbs using 7-point Likert scales measuring imageability and familiarity. The stimulus set consisted of two proverb categories: body-based (BOD) proverbs containing explicit references to bodily actions, body parts, or sensorimotor experiences, and abstract (ABS) proverbs expressing work-related meanings without direct bodily imagery. Ratings were collected separately for Ukrainian and English proverb sets. The dataset includes raw participant responses, worksheet-level calculations, category means, language-specific means, and derived variables used for hypothesis testing. Statistical calculations included comparisons between BOD and ABS proverb categories as well as between L1 and L2 proverb processing. All participant data are fully anonymized. No personally identifiable information is included. The dataset may be useful for research on embodied cognition, conceptual metaphor theory, psycholinguistics, figurative language processing, proverb comprehension, imageability, familiarity, and cross-linguistic studies of language representation. File contents • Raw imageability ratings for Ukrainian proverbs • Raw imageability ratings for English proverbs • Raw familiarity ratings for Ukrainian proverbs • Raw familiarity ratings for English proverbs • Calculated category means (BOD and ABS) • Derived variables for hypothesis testing (H1–H3) • Statistical summary tables Variables Participant_ID – anonymous participant identifier Proverb_Rating – participant rating assigned to a proverb Imageability – perceived ease of forming a mental image (1–7) Familiarity – perceived familiarity with the proverb (1–7) Language – Ukrainian (L1) or English (L2) Category – Body-Based (BOD) or Abstract (ABS) Mean_Score – average score calculated for a participant, proverb category, or language condition License CC BY 4.0
By studying how individuals in an "at-risk" state of psychosis learn about threat and safety cues - specifically, how they develop and unlearn fear responses to neutral cues - we might better understand the mechanisms leading to heightened arousal and fear that are characteristic of acute psychotic episodes. At-risk individuals (N = 88; of which 28 fulfilled ultra-high-risk criteria on the Comprehensive Assessment of At-Risk Mental States interview and 60 scored above a predefined threshold on the Community Assessment of Psychic Experiences) and healthy controls (N = 44) underwent a standardized and validated differential fear conditioning paradigm including an acquisition, generalization, and extinction phase. The main outcomes of interest were the late positive potential, fear-potentiated startle, and self-reported ratings of valence, arousal, fear, and expectancy elicited by the conditioned stimuli (CS). The at-risk group exhibited diminished fear learning, evident in significantly reduced differentiation between the CS+ vs. CS- in the valence ratings compared to controls. Additionally, they demonstrated impaired fear extinction, evident in valence and arousal ratings, in which their CS differentiation showed a slower reduction than the controls. There were no group differences in late positive potential responses. At risk mental states appear to be associated with problems in distinguishing dangerous from safe stimuli and a diminished ability to adjust affective responses to conditioned stimuli based on new information, while the late-positive potential and fear-potentiated startle are unaltered. Early interventions could focus on recalibrating subjective emotional evaluations of fear-associated events.
This study investigates the influence of three biophilic interior design variables: natural light, interior vegetation (vertical green wall), and biomorphic form (biomorphic wall panel) on affective and physiological responses in a design studio interior utilizing immersive virtual reality (IVR) and wearable biofeedback technology. This study was a within-participant 23 factorial design that included one baseline and eight IVR studio conditions. Participants experienced all conditions while reporting affects using the Self-Assessment Manikin (SAM) valence and arousal scales, electrodermal activity (EDA), and skin temperature (ST). Cybersickness was measured with the Simulator Sickness Questionnaire (SSQ) and presence was assessed using the Igroup Presence Questionnaire and Slater-Usoh-Steed presence measures (IPQ, SUS), while baseline anxiety (STAI) was controlled. The results demonstrated a significant primary influence of natural light on SAM valence ratings: conditions with natural light were evaluated as more pleasant than the non-variable and baseline condition, whereas interior vegetation and biomorphic form had smaller, context-dependent effects that were most evident when layered with natural light. Differences in SAM arousal ratings were modest and non-systematic. EDA did not differentiate, and ST showed only small shifts, indicating that during calm exploratory monitoring, subjective affect was more responsive. The circumplex findings guided to an activity-specific zoned interior rather than a single uniform design studio.
FrameNet is an English-based lexical database that shows how words are used by providing information as to which participants and relations are evoked by a certain concept. Recent efforts toward a multilingual FrameNet have not targeted either ancient languages or different historical stages of the same language. In our paper we propose creating a multilingual FrameNet for Ancient Indo-European languages starting with a set of 80 verb meanings annotated in the Pavia Verb Database. Our pilot study includes four verb meanings: RAIN, THUNDER, SEE, LOOK AT. As the adequacy of the semantic frames developed for English turns out not to be appropriate for the languages in our sample, we propose two new frames that can account for the analyzed data.
The integration of generative artificial intelligence into everyday interpersonal communication - through AI-assisted writing, real-time translation, automated summarization, and conversational suggestion tools - represents a structural transformation in the mediation of human-human communication. Unlike earlier communication technologies that mainly transmitted or stored human-generated content, generative AI actively co-produces communicative output through a system architecture that includes user interfaces, prompt and context management, large language model inference, personalization modules, disclosure controls, and post-generation human editing. This paper examines AI-mediated interpersonal communication (AIMIC) as both a social phenomenon and an implementable communication system. It analyzes three interconnected effects: linguistic convergence, redistribution of communicative labor, and challenges to relational authenticity. The revised AIMIC framework links these effects to concrete implementation layers in email assistants, messaging platforms, translation tools, collaborative writing systems, and organizational communication software. The paper argues that responsible AIMIC design requires preserving human communicative agency, making AI involvement contextually transparent, monitoring linguistic and relational impacts, and embedding governance mechanisms into the communication pipeline. The paper concludes with practical examples, evaluation recommendations, and design principles for platform developers, educators, organizations, and policymakers.
This dataset contains the results of a computational operationalization of Greenberg's Universal 45 applied to the Universal Dependencies (UD) corpus (version 2.14). The data covers 339 treebanks across 186 languages. For each treebank and Universal POS tag (UPOS) category, the dataset records whether gender distinctions are present in singular and/or plural tokens, and whether this combination constitutes a violation of the implication universal. The dataset also includes aggregated counts of gender-marked tokens by treebank, language, and UPOS category.
This paper presents a direct framework for sequence models with hidden states on closed subgroups of U(d). We use a minimal axiomatic setup and derive recurrent and transformer templates from a shared skeleton in which subgroup choice acts as a drop-in replacement for state space, tangent projection, and update map. We then specialize to O(d) and evaluate orthogonal-state RNN and transformer models on Tiny Shakespeare and Penn Treebank under parameter-matched settings. We also report a general linear-mixing extension in tangent space, which applies across subgroup choices and improves finite-budget performance in the current O(d) experiments.
Previous research has produced conflicting findings on how sleep affects emotional memories, suggesting it can either strengthen or weaken their emotional intensity. Rather than having a uniform effect, sleep's influence may depend on factors that determine the most adaptive outcome. To explore this, we examined whether the future importance of an emotional experience shapes how sleep alters the emotional intensity of the associated memory. Emotional memories were induced using a novel evaluative learning paradigm featuring a short film clip depicting an acted version of the Trier Social Stress Test. Some of the actors playing the evaluative panel with a critical or neutral demeanour were then introduced as committee members during the participant's own presentation a week later, thereby assigning future relevance or future irrelevance to the formed memories. We recorded changes in emotional responses to memory cues subjectively (valence and arousal ratings) and objectively (skin conductance) after a 12-h period of either daytime wakefulness (N = 32) or containing nighttime sleep (N = 34) and again one week later. Learning was effective, as indicated by more negative feelings and subjective arousal (but not physiological measures) in response to memory cues of the aversive committee members. Regardless of sleep, arousal decreased over 12 h and one week for the future-irrelevant compared to the future-relevant negative memory cue. Valence ratings remained unchanged. These findings suggest that future relevance influences emotional memories independently of sleep. While the benefits of healthy sleep may be subtle, emotional memory processing might be more vulnerable to disruption from poor sleep quality.
This article investigates whether lexical frequency drives stress assignment in Greek, a morphology-dependent stress system. We examine the stress distributions of trisyllabic nouns from seven inflection classes across three lexical databases (GreekLex 2, the HNC Golden Corpus, and HelexKids 2.0) and compare them with the stress patterns produced by adults and elementary school children (Grades 3 and 6) in a pseudo-noun elicitation task. Using hypothesis testing for proportions, frequentist and Bayesian generalized linear mixed-effects models, Bayes Factor analyses, and Monte Carlo simulations, we show that speakers' stress choices do not replicate lexical frequency distributions. Adults show systematic inflection-class conditioning: their stress patterns broadly follow the lexicon's dominant stress pattern per inflection class but depart from its distributional frequencies. Children exhibit a strong, frequency-insensitive preference for penultimate stress that persists across inflection classes, with only traces of suffix-specific stress preferences emerging in Grade 6. We argue that lexical frequency operates within grammatically imposed boundaries: the phonological default and inflection-class conditioning together determine the extent to which distributional patterns in the lexicon shape stress assignment.
This research explores the historical emergence of linguistic terminology in three languages—English, Uzbek, and Karakalpak—with special attention to the role of Latin, Greek, and Arabic heritage. It traces how borrowed concepts were nativized and localized in each linguistic setting. By juxtaposing five evolutionary stages in English with analogous processes in Uzbek and Karakalpak, the paper illustrates the interplay between international scholarly traditions and indigenous linguistic norms. The conclusions highlight both universal tendencies and language-specific particularities in the growth of terminological systems.
Executive functioning (EF) supports goal-directed behavior and is influenced by affective state. While prior research has largely focused on the effects of emotional valence during task performance, less is known about the influence of pre-task self-reported valence on subsequent EF performance. Our ongoing study will examine whether self-reported ratings of emotional valence assessed prior to the start of a task will predict performance. Participants will report valence ratings during a resting baseline period before performing two EF tasks: the Vertical Stroop and the Go/No-Go. Performance will be assessed using reaction time and accuracy. Regression analyses will be conducted to examine the relationship between pre-task valence ratings and task performance. We hypothesize that resting-state valence will predict EF performance, such that more positive baseline valence will be associated with faster reaction times and/or higher accuracy. If correct, these findings will show that baseline affective states may influence executive control processes before cognitive demands are imposed.
This page contains behavioral data of an encoding and a temporal memory task in two experiments. Experiment 2 also includes an emotional valence rating that was conducted at the end of the experiment. For information about the study, please see the published manuscript in Psychological Research.
Analyzing the structure of complex sentences, comprising multiple clauses within a single sentence, has been recognized as a challenging aspect of dependency parsing. This challenge arises from the syntactic ambiguity inherent in clausal relations, wherein the main verb of each clause may have multiple candidate dependency heads. Minami's Scope Preference Theory (1964, 1974) hypothesized that clausal relations are determined by preferences between pairs of subordinate clauses. Building upon this theoretical foundation, we introduce a neural model fine-tuned to resolve clausal relations, with a focus on the verbs within each clause. To address the challenges in dependency parsing, we propose a forest reranking approach that enables our reranking model to grasp global context. Our approach utilizes our model as a reranker on dependency forests by leveraging cube-pruning for efficient tree enumeration. Through a series of experiments and analyses conducted on the Penn Treebank (PTB) and Penn Chinese Treebank (CTB), we find that our approach is effective in resolving the ambiguity inherent in complex sentence structures.
<p>For Arabic-English bilingual students, writing is a particularly challenging task, as it requires learning how to produce meaning through written scripts, distinctively different from those of the mother tongue. In this study, subjective familiarity ratings and vocabulary knowledge scores of the printed words of the Peabody Picture Vocabulary Test (PPVT) were collected from native Arabic speakers whose English proficiency was at least at the competent-user level (as per International English Language Testing System [IELTS] criteria). This level of competency is generally the precondition for admission of second-language learners to English-medium universities. The aim of the study was two-fold: (a) to determine the information upon which participants’ vocabulary knowledge relies through an examination of the extent to which such knowledge is predicted by key subjective and objective word properties; and (b) to assess the degree to which participants’ vocabulary knowledge, estimated from printed word comprehension (WC), is related to second-language students’ writing performance, as well as writing anxiety. In this study, objective word properties (e.g., frequency counts and the number of semantic neighbors), as well as subjective familiarity ratings of printed words, notably contributed to vocabulary knowledge (as indexed by WC scores). Furthermore, participants’ vocabulary knowledge was related to writing performance, as well as writing anxiety. Thus, the printed words of the PPVT could be used to predict not only the vocabulary knowledge of Arabic-English speakers admitted to university-level courses but also writing difficulties, thereby informing selective preemptive interventions.</p>
Abstract Word frequency databases like SPALEX and SUBTLEX-ESP treat Spanish as a uniform language, but prior studies and an initial survey (Experiment 1) revealed significant lexical differences between Spanish in Spain and Latin American countries, especially Chile. To establish subjective frequencies of Spanish word usage, an extended survey (Experiment 2) was conducted with Chilean participants, categorizing words by usage area: General, Spain, Chile, and Latin America. Consistent with the initial survey, Chilean participants assigned subjective higher ratings to General and Chilean words. In a lexical decision experiment (Experiment 3), participants responded faster and more accurately to words from these categories. Using survey data, simulations with Multilink+ (Experiment 4) revealed that subjective word ratings better predicted Chilean reaction times than frequencies from existing databases. These findings emphasize the need to address Spanish dialectal differences in research, with word ratings offering a more accurate measure of region-specific lexical nuances than current databases.
Abstract In this paper, I aimed to develop a neural parser for Bangla based on simplified Head-driven Phrase Structure Grammar (HPSG) with neural network-based models. The initial stage in natural language processing is to break down the text into separate tokens. When the text corpus is huge, covering all words is inefficient regarding size of vocabulary. The effectiveness of a specific tokenization method varies on various factors, such as size of the dataset, the nature of the task, and the morphological complexity of the dataset. Due to the lack of existing HPSG-compliant treebanks for Bangla, we utilized syntactically annotated resources from existing Bangla corpora and modified them to align with simplified HPSG rule-based restructuring and data permutation. After that we modified a neural parser architecture originally designed for the Penn Treebank, replacing its encoder with multilingual pre-trained models such as XLM-RoBERTa and IndicBERT to better capture the syntactic and lexical entries of Bangla. We conducted experimental evaluations on the modified dataset, and the parser demonstrated promising results in both constituency and dependency parsing tasks. Our extensive experiments showed that the simplified HPSG Neural Parser achieved a new state-of-the-art for constituency parsing when using the same predicted part-of-speech (POS) tags as the self-attentive constituency parser. Additionally, it outperformed previous studies in dependency parsing with a higher Unlabeled Attachment Score (UAS). However, our parser remained lower Labeled Attachment Score (LAS) scores likely due to integrating HPSG with neural approaches for Bangla syntax parsing and underscoring the importance of linguistically informed treebank development in low-resource languages. Lastly, the research findings of this paper suggest that simplified HPSG should be given more attention to linguistic experts when developing treebanks for Bangla Natural Language Processing (BNLP).
We describe THIVLVC, a two-stage system for the EvaLatin 2026 Dependency Parsing task.Given a Latin sentence, we retrieve structurally similar entries from the CIRCSE treebank using sentence length and POS n-gram similarity, then prompt a large language model to refine the baseline parse from UDPipe using the retrieved examples and UD annotation guidelines.We submit two configurations: one without retrieval and one with retrieval (RAG).On poetry (Seneca), THIVLVC improves CLAS by +17 points over the UDPipe baseline; on prose (Thomas Aquinas), the gain is +1.5 CLAS.A double-blind error analysis of 300 divergences between our system and the gold standard reveals that, among unanimous annotator decisions, 53.3% favour THIVLVC, showing annotation inconsistencies both within and across treebanks.
The majority of secondary school pupils in Tanzania are multilingual, speaking at least three languages which include: ethnic community language (there are currently more than 120 of them), (Ki)swahili, the national and first official language and, at varying levels of competency, English which is accorded the status of second official language. A very small number of pupils have access to French, since the language is taught only in a few schools as an optional subject. In public primary schools, pupils are generally bilingual, speaking their ethnic community language and (Ki) swahili. Due to their bilingual/multilingual knowledge, they are expected to activate each of the languages in their repertoire according to the situation of communication and to its level of formality as well as to the purpose of communication, the participants and their various characteristics, identity factors, etc. However, in certain institutional settings, the activation of the language repertoire is determined by the norms established by the schools. This paper is intended to: firstly, describe the formation of the bi/multilingual repertoire of Tanzanian primary and secondary school pupils and the nature of their language practices outside of school settings; secondly, indicate how the language practices are modified by the school and, thirdly, explain the ideological and/or pedagogical origins of the linguistic norms set by schools. The conclusion will attempt to explain the impact of the linguistic norms on the perception that the pupils have about the different languages in contact.
The study explores the reflection of politeness strategies in the translation process, focusing on English source texts and on their Uzbek translations. Using Brown and Levinson`s (1987) framework of positive and negative politeness, the research examines how translators preserve, adapt or omit these strategies in several contexts. The findings indicate that positive politeness strategies are more frequently maintained, whereas negative politeness often undergoes modification to suit Uzbek cultural and linguistic norms. The study highlights the importance of cultural and pragmatic sensitivity in translation.
The majority of secondary school pupils in Tanzania are multilingual, speaking at least three languages which include: ethnic community language (there are currently more than 120 of them), (Ki)swahili, the national and first official language and, at varying levels of competency, English which is accorded the status of second official language. A very small number of pupils have access to French, since the language is taught only in a few schools as an optional subject. In public primary schools, pupils are generally bilingual, speaking their ethnic community language and (Ki) swahili. Due to their bilingual/multilingual knowledge, they are expected to activate each of the languages in their repertoire according to the situation of communication and to its level of formality as well as to the purpose of communication, the participants and their various characteristics, identity factors, etc. However, in certain institutional settings, the activation of the language repertoire is determined by the norms established by the schools. This paper is intended to: firstly, describe the formation of the bi/multilingual repertoire of Tanzanian primary and secondary school pupils and the nature of their language practices outside of school settings; secondly, indicate how the language practices are modified by the school and, thirdly, explain the ideological and/or pedagogical origins of the linguistic norms set by schools. The conclusion will attempt to explain the impact of the linguistic norms on the perception that the pupils have about the different languages in contact.
This article delves into the diverse of translation practices, examining various types of translations across different linguistic contexts especially used by Uzbek translators. It investigates the nuances and challenges inherent in translation processes, ranging from word-for-word translations to adaptations tailored to specific cultural and linguistic norms. By exploring examples and scholarly perspectives, the article sheds light on the complexities involved in conveying meaning accurately across languages. It underscores the importance of understanding the distinct characteristics of each translation type and the significance of employing appropriate strategies to ensure fidelity to the original text. Overall, the article can serve as a comprehensive exploration of the multifaceted nature of translation endeavors.
In filmmaking, a neutral face paired with an emotional context can convey a congruent emotion-a phenomenon known as the Kuleshov effect. However, past research has yielded mixed findings on the existence of the Kuleshov effect, possibly due to methodological variability; we therefore implemented a boundary-test paradigm that combined normed static context images with dynamic facial stimuli separated by an explicit context-rating step to examine how visual context shapes the evaluation and categorization of faces under such conditions. We hypothesized that positive context would elevate facial valence ratings (and vice versa), while evoking different-than-neutral emotions in a neutral face when categorized. Thirty-two participants first rated the valence of a context image on a 7-point Likert scale. They then viewed a video of a neutral face, evaluated its valence on the same scale, and explicitly categorized the expression by selecting one of five predefined emotion labels (neutral, angry, happy, sad, disgusted). Analyses using linear mixed-effects models confirmed that positive contexts led to higher facial valence ratings, while negative contexts led to lower ones; by contrast, multinomial regression revealed no effect of context on emotional categorization. Our findings suggest that while visual context can bias evaluative judgments of facial expressions, it does not necessarily alter their emotional categorization under conditions where cinematic continuity is absent, highlighting the boundaries of the Kuleshov effect in line with previous studies.