Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
As language-based AI systems become more anthropomorphic, the question of whether they can have subjective experience is increasingly pressing. I focus here on the tractability of research questions in the space of AI consciousness. I argue that the fundamental problem of whether AI systems can be conscious is currently intractable in its direct form, given the absence of a universally accepted scientific theory of consciousness, as well as the historical open-endedness of the philosophical mind-body problem. In contrast, questions around the adjacent subject of perceived AI consciousness are tractable, timely, and highly consequential for society. The general public is increasingly open to the possibility of consciousness in AI systems and routinely adopts the vocabulary of human cognition and subjective experience to describe them. This phenomenon is already driving societal shifts across user experience, ethical standards, and linguistic norms. I therefore propose an increased research focus on uncovering the causes and effects of perceived AI consciousness, which ultimately shape how we see our own human subjective experience relative to artificial entities. To support this, I map the current landscape of AI consciousness perception and discuss its key potential drivers and societal consequences. Finally, I urge developers, decision-makers, and the broader scientific community to commit to clear and accurate communication regarding the topic of AI consciousness, explicitly acknowledging its inherent uncertainties.
Contemporary large language models (LLMs) rely on sub-word tokenizers and at attention mechanisms that treat every language as a statistical surface-form distribution. This paper proposes Dhatu-Former, a transformer architecture that internalizes the formal linguistic machinery of Pan.ini’s As..tadhyay the oldest known generative grammar. We hypothesize that (i) morphologically-aware, root-based (dhatu-based) tokenization can reduce vocabulary size and sequence length by 40 60%, (ii) hierarchical attention guided by Pan.inian derivation trees can yield sparse, interpretable attention with O(nlogn) complexity, and (iii) a hybrid symbolic neural reasoning layer that executes sutra-style rewrite rules can substantially reduce hallucination while enabling uni ed language math logic reasoning. We further introduce a modular Retrieval-Augmented Generation (RAG) subsystem grounded in Sanskrit lexical databases (Amarakos.a, Dhatupat.ha) and a continual learning framework inspired by the paribhas.a sutra (meta-rules) of the As..tadhyay. We present order-of-magnitude parameter reduction estimates, architectural blueprints with TikZ diagrams, and a research roadmap for empirical validation. This is a position paper; no experiments have been conducted.
Do vision--language models (VLMs) develop more human-like sensitivity to linguistic concreteness than text-only large language models (LLMs) when both are evaluated with text-only prompts? We study this question with a controlled comparison between matched Llama text backbones and their Llama Vision counterparts across multiple model scales, treating multimodal pretraining as an ablation on perceptual grounding rather than access to images at inference. We measure concreteness effects at three complementary levels: (i) output behavior, by relating question-level concreteness to QA accuracy; (ii) embedding geometry, by testing whether representations organize along a concreteness axis; and (iii) attention dynamics, by quantifying context reliance via attention-entropy measures. In addition, we elicit token-level concreteness ratings from models and evaluate alignment to human norm distributions, testing whether multimodal training yields more human-consistent judgments. Across benchmarks and scales, VLMs show larger gains on more concrete inputs, exhibit clearer concreteness-structured representations, produce ratings that better match human norms, and display systematically different attention patterns consistent with increased grounding.
Do vision--language models (VLMs) develop more human-like sensitivity to linguistic concreteness than text-only large language models (LLMs) when both are evaluated with text-only prompts? We study this question with a controlled comparison between matched Llama text backbones and their Llama Vision counterparts across multiple model scales, treating multimodal pretraining as an ablation on perceptual grounding rather than access to images at inference. We measure concreteness effects at three complementary levels: (i) output behavior, by relating question-level concreteness to QA accuracy; (ii) embedding geometry, by testing whether representations organize along a concreteness axis; and (iii) attention dynamics, by quantifying context reliance via attention-entropy measures. In addition, we elicit token-level concreteness ratings from models and evaluate alignment to human norm distributions, testing whether multimodal training yields more human-consistent judgments. Across benchmarks and scales, VLMs show larger gains on more concrete inputs, exhibit clearer concreteness-structured representations, produce ratings that better match human norms, and display systematically different attention patterns consistent with increased grounding.
While language enables meaning, constituting knowledge in courts, schools, or parliaments, who gets to decide what can be known? Is meaning only use or a result of power too? Pitting Wittgenstein's forms of life against Foucault's regimes of discourse makes linguistic norms appear as instruments of exclusion. Marginalised speakers – subaltern, indigenous, and non-normative are often rendered unintelligible. Epistemic justice demands more than inclusion; it demands considering how rules are set, who enforces them, and how meaning is being contextually built. A discourse-sensitive, epistemic theory of justice is proposed, based on Kripke's rule-following paradox and Dijk's discourse analysis, to show that language is not neutral but a battleground of struggle over meaning, recognition, and epistemic authority.
This study investigates the influence of artificial intelligence on the evolution of language in modern society. It analyzes how AI-driven technologies—including machine translation, natural language processing, chatbots, and speech recognition systems—affect language use, transformation, and development. The research highlights that the growing integration of AI into everyday communication has accelerated multilingual interaction, contributed to the reshaping of linguistic norms, and facilitated the emergence of new vocabulary associated with digital environments. Furthermore, the article examines the dual impact of AI on language practices. On the one hand, it enhances linguistic diversity and expands access to global communication; on the other hand, it raises critical concerns regarding linguistic standardization, the erosion of cultural nuances, and increasing dependence on automated systems. Particular emphasis is placed on the role of AI in language education, where adaptive learning platforms and intelligent tutoring systems enable more personalized instructional pathways. The findings indicate that artificial intelligence should be understood not merely as a technological advancement, but as a significant sociolinguistic force that is reshaping patterns of communication, literacy practices, and cultural identity in the digital era.
Static concreteness ratings are widely used in NLP, yet a word's concreteness can shift with context, especially in figurative language such as metaphor, where common concrete nouns can take abstract interpretations. While such shifts are evident from context, it remains unclear how LLMs understand concreteness internally. We conduct a layer-wise and geometric analysis of LLM hidden representations across four model families, examining how models distinguish literal vs figurative uses of the same noun and how concreteness is organized in representation space. We find that LLMs separate literal and figurative usage in early layers, and that mid-to-late layers compress concreteness into a one-dimensional direction that is consistent across models. Finally, we show that this geometric structure is practically useful: a single concreteness direction supports efficient figurative-language classification and enables training-free steering of generation toward more literal or more figurative rewrites.
Respect plays a crucial role in successful interpersonal and intercultural communication. However, differences in linguistic norms and pragmatic conventions often lead to misunderstandings between speakers of different languages. This article examines pragmatic failures in expressing respect in English-Russian cross-cultural communication. The study aims to identify linguistic and cultural factors that cause misinterpretations of respect and to analyze how respect is pragmatically encoded in both languages. The findings suggest that pragmatic failures frequently arise from divergent politeness strategies, speech act realizations, and sociocultural expectations embedded in English and Russian communicative practices. The study emphasizes the importance of pragmatic awareness in developing intercultural communicative competence.
This article explores the dynamics of linguistic norms in the Russian literary language within the context of social media communication. The relevance of the study lies in the fact that digital platforms create new communicative conditions that reshape traditional language norms. The paper analyzes variability, simplification, and expressive tendencies in syntax, lexicon, and orthography. Furthermore, it demonstrates that linguistic norms in social media are not static but dynamic and context-dependent. The findings suggest that these processes reflect not the degradation of the literary language, but its adaptation to digital discourse.
The study is devoted to the analysis of the impact of globalization and digital technologies on the transformation of language norms in modern digital discourse. The relevance of the work is due to the growing role of digital communication, in which language becomes a key factor of cultural identity and social integration. The goal is to identify interdependencies between the level of digital maturity of states, the intensity of language hybridization and the dynamics of sociolinguistic variability. The object of the study is the digital language space of the global communication environment. The methodology is based on a combination of systemic, institutional, econometric and corpus-linguistic approaches using official statistical databases and corpora of digital texts for 2015–2024. The results show that during this period ICT Development Index increased from 5.32 to 7.41 points, DESI Index – from 48.7 to 70.4 points, and the share of Internet users – from 58.2% to 84.3%. At the same time, the number of languages with digital representation increased from 312 to 387, and the share of English-language content decreased from 55.1% to 49.3%. The hybridization index increased from 0.42 to 0.61, clearly indicating the establishment of a polycentric mode or multi-node digital discourse. Hybridization is both method and degree: different languages or their structural elements used in the same message (verbal or symbolic), as well as fully/partially mixed code messages; hashtags expressed in different linguistic forms/fonts, etc. (multimedia message). This means how various linguistic components are arranged within one structural unit up to multimedia messages comprising codes written with different fonts‐forms on various levels of hybridity. The econometric model registered a very high positive correlation between digital maturity and hybridization (r=0.82).
This study aims to examine how non-dominant languages challenge the current language hierarchy in Greek public administration and encourage reconsideration of the status quo. The data are drawn from a broader research project aimed at examining multilingualism within Greek public administration. A questionnaire, specifically designed for the purposes of the research, was administered across all Greek ministries and interviews were conducted with civil servants from various ministries. A total of 1,156 responses along with qualitative data from interviews reveal a notable discrepancy between the languages employees speak and those they consider most useful for their work. The analysis shows that while English remains the most frequently cited language for professional use, there is a growing recognition of the importance of other languages, including Middle Eastern, Balkan, Afghan – Pakistani and Southern European languages. The findings suggest that the rigid focus on English and other ‘prestigious’ languages fails to acknowledge the practical needs of a multilingual public sector. The study advocates for a more inclusive approach to language policy and planning in Greek public administration, one that promotes linguistic diversity and addresses the diverse needs of both citizens and civil employees.
The article examines the phenomenon of code-switching in Internet discourse as one of the key strategies of modern digital communication. Theoretical foundations from R. Jakobson to P. Blom, J. Gumperz, Sh. Poplack, and C. Myers-Scotton are discussed, along with their application in online interaction. Based on examples from Instagram, Telegram, and YouTube, it is shown that code-switching is not a deviation from linguistic norms but a means of expressing emotions, irony, and identity, adapting messages to multilingual audiences. The study concludes that a hybrid norm of online communication is emerging within English-, Russian-, and Uzbek-speaking digital spaces.
The article explores the value orientations typical of residents in monotowns, i.e., single-industry towns. The linguistic data were obtained from an indirect associative experiment through an analysis of stimulusresponse speech acts. The methods from corpus linguistics and psycholinguistics made it possible to identify the hierarchical framework of key values in the consciousness of respondents from the diamond mining town of Mirny, Republic of Sakha. The core values included health, family, love, friendship, and security, as well as their semantic associations. The methodology for calculating speech act characteristics included the intersection index (O – Number of Overlapping Associates) and intersection strength (OSG – Overlapping Associate Strength), which reflected the psychological proximity of value meanings in communicative practices. The method revealed the integrative and regulatory role of values based on parameters of speech acts as social actions. The results show how language captures and transmits value orientations, reflecting both universal and ethnocultural perceptions. The findings contribute to the development of interdisciplinary linguistics, enhancing the understanding of value system in single-industry settlements. The method can be used for social planning, academic programs, and cross-cultural studies.
FrameNet is an English-based lexical database that shows how words are used by providing information as to which participants and relations are evoked by a certain concept. Recent efforts toward a multilingual FrameNet have not targeted either ancient languages or different historical stages of the same language. In our paper we propose creating a multilingual FrameNet for Ancient Indo-European languages starting with a set of 80 verb meanings annotated in the Pavia Verb Database. Our pilot study includes four verb meanings: RAIN, THUNDER, SEE, LOOK AT. As the adequacy of the semantic frames developed for English turns out not to be appropriate for the languages in our sample, we propose two new frames that can account for the analyzed data.
This article examines the role of advertisements and signboards in shaping and reflecting public attitudes toward language. In modern society, linguistic culture is not only preserved in literature and education, but also manifested in everyday public texts such as commercial advertisements, street signs, shop names, and information boards. The study analyzes the linguistic quality of advertising texts, the influence of globalization on language use, and the social consequences of neglecting linguistic norms. Special attention is given to the relationship between language accuracy and cultural identity. The article also discusses the responsibility of businesses, media representatives, and educational institutions in maintaining linguistic standards in public communication.
Data, code, figures, and the manuscript for The architecture of internet aesthetics (Y. J. Lin, Cornell University).Contents: (1) the internet-aesthetics ecosystem network of 1,115 aesthetics joined by 5,889 community-authored related-links, with community, affect, and era attributes and a self-contained interactive HTML explorer (pan / zoom / search); (2) the are.na human-label and co-curation records; (3) an openly-licensed 681-image stimulus corpus from Wikimedia Commons with per-image attribution (CREDITS.md) and a 681x37 perceptual and mechanism feature table; (4) three-model frontier vision-language affect ratings over the are.na images, from gemini-2.5-flash, gpt-4o, and claude-sonnet-4-6; (5) OASIS benchmarking and mechanistic-intervention (image-edit and prompt-reframe) outputs; and (6) all harvest, analysis, and figure code.The are.na images themselves are copyright-held and are NOT included in this archive; only derived features and model ratings are released. The 681 Wikimedia images ARE included, each retaining its own CC / public-domain licence (see CREDITS.md and LICENSES.md). The dataset compilation (manifests, feature tables, ratings, network, and code) is released under CC-BY-4.0.
We present a manually curated dataset of Spanish second language and heritage learner writing following the Universal Dependencies framework.
In this paper we introduce an example-based method for exploring dependency treebanks that is based on principles of vector symbolic architectures.It leverages key properties of this framework to provide fast and flexible search capabilities, since all combinations of query parameters can be compared with a given parse tree in parallel via a single vector operation.The framework also allows for graded similarity and the natural integration of various kinds of information, such as word embeddings.After some background on the framework and an explanation of our implementation, we provide a few examples of the system's output and draw comparisons to similar applications.
The aim of the paper is to present a first attempt at annotating Information Structure roles in syntactic treebanks of the Universal Dependencies collection, discussing theoretical considerations and practical methodological questions while presenting our core annotation principles.We focus on constructions in which Topic or Focus is overtly marked through (morpho)syntactic means.The proposed annotation is illustrated using examples from five languages: Wolof, Japanese, Tundra Nenets, Hungarian, and Italian.
During the recent years, the use of linguistic data for language processing increased progressively. Such data are now commonly called language resources. Most of the language resources used for this purpose are collections of texts as the Brown Corpus and the Penn Treebank, but electronic lexicons (WordNet, FrameNet, VerbNet, ComLex, Lexicon-Grammar...) and formal grammars (TAG...) developed recently. Most processes of construction of lexicons and grammars are manual, whereas the construction of corpora has always been highly automated. However, more and more specialists of language processing realize that the information content of lexicons and grammars is richer than that of corpora, and hence the former make more elaborate processing possible. The difference in construction time is likely to be connected with the difference in information content: the handcrafting of lexicons and grammars by linguists would make them more informative than automatically generated data. This situation can evolve into two directions: either specialists of language technology get progressively used to handling manually constructed resources, which are more informative and more complex, or the process of construction of lexicons and grammars is automated and industrialized, which is the mainstream perspective. Both evolutions are already in progress, and a tension exists between them. The relation between linguists and computer scientists depends on the future of these evolutions, since the first implies training and hiring numerous linguists, whereas the other depends essentially on solutions elaborated by computer engineers. The aim of this article is to analyse practical examples of the language resources in question, and to discuss about which of the two trends, handcrafting or generating industrially, or a combination of both, can give the best results or is the most realistic.
This study examines multilingual neural machine translation (MNMT) for a diverse group of low-resource Asian languages-Bengali, Filipino, Indonesian, Japanese, Khmer, Malay, and Vietnamese-which differ substantially in linguistic families, writing systems, and typology. This paper evaluates state-of-the-art MNMT systems and introduces a Compact & Language-Sensitive MNMT model designed to improve translation performance while reducing computational cost. The proposed approach shares parameters through a compact multilingual representation, and enhances language discrimination using language-sensitive embeddings, a language-sensitive discriminator, and an adaptive cross-attention mechanism that selects attention parameters based on specific language pairs. Integrated with a multi-stage fine-tuning strategy, this model effectively strengthens cross-lingual transfer while maintaining robust language-specific representations. Experiments on the ALT multi-parallel corpus and the KFTT English-Japanese dataset demonstrate that multilingual models significantly outperform single-language NMT baselines. Despite its smaller size, the proposed Compact & Language-Sensitive MNMT achieves competitive or superior BLEU scores compared to Google’s MNMT, confirming the effectiveness of guided parameter sharing and language-sensitive training. These results highlight the value of compact multilingual architectures and multi-parallel datasets for advancing low-resource Asian machine translation.
Brain Treebank is a large-scale intracranial EEG dataset comprising 43 hours of iEEG recordings from 10 epilepsy patients watching naturalistic Hollywood movies, with 1,688 electrodes sampled at 2048 Hz. The dataset includes time-aligned linguistic annotations with word-level transcripts and Universal Dependencies syntax trees, providing a unique resource for studying neural language processing during naturalistic stimulation.
Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing image captioning benchmarks typically suffer from limited diversity in caption length, the absence of recent advanced MLLMs, and insufficient human annotations, which potentially introduces bias and limits the ability to comprehensively assess the performance of modern MLLMs. To address these limitations, we present a new large-scale image captioning benchmark, termed, ICBench, which covers 12 content categories and consists of both short and long captions generated by 10 advanced MLLMs on 2K images, resulting in 40K captions in total. We conduct extensive human subjective studies to obtain mean opinion scores (MOSs) across fine-grained evaluation dimensions, where short captions are assessed in terms of fluency, relevance, and conciseness, while long captions are evaluated based on fluency, relevance, and completeness. Furthermore, we propose an automated evaluation metric, \textbf{ITIScore}, based on an image-to-text-to-image framework, which measures caption quality through reconstruction consistency. Experimental results demonstrate strong alignment between our automatic metric and human judgments, as well as robust zero-shot generalization ability on other public captioning datasets. Both the dataset and model will be released upon publication.
Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.
While the linguistic shifts between Ottoman and modern Turkish are well-documented qualitatively, quantitative analyses remain scarce.This study addresses this by conducting a comparative computational analysis using two Universal Dependencies treebanks: OTA-DUDU for Ottoman Turkish and TR-BOUN for modern Turkish.By employing descriptive statistics and a log-likelihood ratio test, we demonstrate the change and quantify the magnitude of diachronic variation.The analysis yields three primary statistical findings.First, our data reveals a 77% compliance rate with labial vowel harmony for suffixes, while this value is 98% in modern Turkish.This discrepancy can be explained by the presence of rounding in Ottoman Turkish, which disappears in modern Turkish.On the other hand, the compliance rate of palatal vowel harmony is quite high for both languages, 96% for Ottoman Turkish and 99% for modern Turkish.Second, some suffixes, such as the converb -(y)Ip 1 and the dative infinitive -mAyA, changed by reducing their allomorphs in modern Turkish.Third, we demonstrate that Arabic and Persian pluralization rules, which constituted 28% of plural nouns in Ottoman Turkish, lost their pluralizing function in modern Turkish, although the words remain with singular meaning.
Abstract This paper presents a novel treebank-driven approach to comparing syntactic structures in speech and writing using dependency-parsed corpora. Adopting a fully inductive, bottom-up method, we define syntactic structures as delexicalized dependency (sub)trees and extract them from spoken and written Universal Dependencies (UD) treebanks in two syntactically distinct languages, English and Slovenian. For each corpus, we analyze the size, diversity, and distribution of syntactic inventories, their overlap across modalities, and the structures most characteristic of speech. Results show that, across both languages, spoken corpora contain fewer and less diverse syntactic structures than their written counterparts, with consistent cross-linguistic preferences for certain structural types across modalities. Strikingly, the overlap between spoken and written syntactic inventories is very limited: most structures attested in speech do not occur in writing, pointing to modality-specific preferences in syntactic organization that reflect the distinct demands of real-time interaction and elaborated writing. This contrast is further supported by a keyness analysis of the most frequent speech-specific structures, which highlights patterns associated with interactivity, context-grounding, and economy of expression. We argue that this scalable, language-independent framework offers a useful general method for systematically studying syntactic variation across corpora, laying the groundwork for more comprehensive data-driven theories of grammar in use.
Using semantic dependency analysis, this study examines narrative productions from Mandarin-speaking preschool children aged three to six to investigate how semantic organization develops with age in early childhood. Four semantic dependency treebanks were constructed from a Chinese narrative corpus available in the CHILDES database. By comparing semantic dependency types and semantic dependency distances across the four age groups, we found that (1) semantic organization shifted from experiencer and classification relations toward agent and patient relations, situational-role relations (particularly those involving measurement, individuation, and direction), and structural relations; (2) mean semantic dependency distance (MSDD) increased with age, as adjacent dependencies decreased and longer dependencies became more frequent. This increase in MSDD indicates growing complexity in semantic organization and is largely driven by significant increases in the MSDD values of specific semantic dependency types. These findings provide new evidence for semantic organization development in preschool children.
Thermal imaging, which is contact-free, light-independent, and effective in detecting skin temperature changes that reflect autonomic nervous system activity, is expected to be useful for emotion sensing. A recent thermography study demonstrated a linear relationship between ear temperatures and emotional arousal ratings. However, whether and how ear thermal changes may be nonlinearly related to subjective emotions remains untested. To address this issue, we reanalyzed a dataset that included ear thermal images and self-reported arousal ratings obtained while participants watched emotion-eliciting films. We employed linear regression and two nonlinear machine learning models: a random forest model and a ResNet-50 convolutional neural network. Model evaluation using mean squared error and correlation coefficients between actual arousal ratings and model predictions indicated that both machine learning models outperformed linear regression and that the ResNet-50 model outperformed the random forest model. Interpretation of the ResNet-50 model using Gradient-weighted Class Activation Mapping and Shapley additive explanation methods revealed nonlinear associations between temperature changes in specific ear regions and subjective arousal ratings. These findings imply that ear thermal imaging combined with machine learning, particularly deep learning, holds promise for emotion sensing.
This paper presents a novel treebank-driven approach to comparing syntactic structures in speech and writing using dependency-parsed corpora. Adopting a fully inductive, bottom-up method, we define syntactic structures as delexicalized dependency (sub)trees and extract them from spoken and written Universal Dependencies (UD) treebanks in two syntactically distinct languages, English and Slovenian. For each corpus, we analyze the size, diversity, and distribution of syntactic inventories, their overlap across modalities, and the structures most characteristic of speech. Results show that, across both languages, spoken corpora contain fewer and less diverse syntactic structures than their written counterparts, with consistent cross-linguistic preferences for certain structural types across modalities. Strikingly, the overlap between spoken and written syntactic inventories is very limited: most structures attested in speech do not occur in writing, pointing to modality-specific preferences in syntactic organization that reflect the distinct demands of real-time interaction and elaborated writing. This contrast is further supported by a keyness analysis of the most frequent speech-specific structures, which highlights patterns associated with interactivity, context-grounding, and economy of expression. We argue that this scalable, language-independent framework offers a useful general method for systematically studying syntactic variation across corpora, laying the groundwork for more comprehensive data-driven theories of grammar in use.
Background: Emotion processing is critical in the neuropathology of major depressive disorder (MDD), while its relationship with clinical treatment remains unclear. This study aims to indicate the associations between emotion processing and treatment effects following a sequential dual-site accelerated repetitive transcranial magnetic stimulation (rTMS) protocol. Methods: MDD patients were recruited to receive rTMS treatment with four sessions per day for four consecutive days, with stimulation sequentially delivered to the left dorsolateral prefrontal cortex (dlPFC) and the dorsomedial prefrontal cortex (dmPFC). Symptoms were assessed at baseline, end of treatment, and week 4 using the Montgomery–Åsberg Depression Rating Scale (MADRS), Snaith-Hamilton Pleasure Scale (SHAPS), and Fatigue Severity Scale (FSS). Emotional valence and arousal were evaluated with the Affect Rating Task (ART). Results: A total of 51 participants completed the clinical assessments and ART, with two excluded due to missing baseline data in the SHAPS and FSS. The linear mixed-effects models revealed significant improvement in depressive (p < 0.001, d = −0.343) and fatigue symptoms (p = 0.010, d = −0.572) following rTMS treatment. Neutral valence was correlated with MADRS scores at baseline (R2 = 0.096, p = 0.027). In addition, changes in arousal for positive images (p = 0.047, adjusted R2 = 0.097) and neutral images (p = 0.019, adjusted R2 = 0.160) at treatment end were significantly correlated with MADRS improvement at week 4. Conclusions: Our study highlights the association between changes in emotional arousal and improvement in MDD following accelerated dlPFC-dmPFC dual-site rTMS treatment.
Abstract Quranic Arabic has motivated sustained morphological and syntactic annotation, yet Quranic treebanks remain hard to compare and reuse in modern natural language processing (NLP) because they diverge in clitic segmentation, feature inventories, and syntactic formalisms. We present UD-Quran, a Universal Dependencies (UD) v2 conversion of the Extended Quranic Treebank with hybrid syntactic annotations (EQTB). The conversion treats EQTB morpho-syntactic segments as UD tokens, maps EQTB part-of-speech (POS) categories to 12 UD universal part-of-speech (UPOS) tags, derives UD features from explicit EQTB columns, and collapses EQTB dependency labels into a compact UD relation inventory with deterministic normalization aligned to UD content-head conventions. Two releases are provided: a surface variant aligned to the observable Quranic string by excluding analytically inserted nodes, and an augmented variant that retains inserted material to preserve EQTB’s modeling of ellipsis and implied pronominals. The surface release contains 11,693 sentences and 128,219 UD tokens; the augmented release contains 139,376 tokens. Conversion coverage is quantified by restricting unspecified dependency (dep) to 1,129 tokens (0.881% of surface tokens). UD-Quran includes fixed training/development/test (train/dev/test) splits (seed 42) and lightweight Stanza baselines scored with the CoNLL (Conference on Computational Natural Language Learning) 2018 UD evaluation script. On the test sets, parsing with gold tags reaches labeled attachment score (LAS) 80.66 (surface) and 82.47 (augmented), while the end-to-end pipeline reaches LAS 64.52 and 68.54. UD-Quran is intended as an interoperability layer that supports standard UD tooling while preserving sentence-level traceability to EQTB.
BACKGROUND: L. (caraway) essential oils (EOs) on aging. First, we assessed, in 402 participants, the age-related changes in olfactory functions (odor threshold, discrimination, and identification), gustatory perceptions (sweet, sour, salty, and bitter taste), cognitive functions (focusing on attention, memory, language, and visuospatial/executive functions), and their possible correlations with aging. To achieve this, olfactory function, gustatory perception, and cognitive abilities were evaluated in healthy participants across different age groups. Then, to evaluate the age-related decrease in trigeminal function (59 participants), we used rosemary and caraway EOs that contain carvone, limonene, and 1,8-cineole, all of which are considered typical trigeminal stimuli. METHODS: Olfactory function was assessed with the Sniffin' Sticks test, gustatory function by the Taste Strips test, and rosemary and caraway EOs by the ratings of odor pleasantness, intensity, and familiarity using a labeled hedonic Likert-type scale. RESULTS: Olfactory function could be a potential early indicator of attentional, memory, language, and visuospatial/executive dysfunctions. Our data indicated that rosemary and caraway EOs were perceived without any significant decrease in odor pleasantness, intensity, and familiarity ratings in relation to aging. CONCLUSION: Our results suggest the potential bioactive effects of rosemary and caraway natural EOs as a new strategy to promote healthy aging.
Monolingualism, native-speakerism and standard language ideology have been identified as dominant ideologies in language teaching with severe effects on second language teacher identities. Such ideologies offer alleged certainties but also detach teachers from the actual uses and value of language in multilingual and multidialectal contexts. As a consequence, educators might feel constrained by rigid linguistic norms, hindering their capacity to re-evaluate their approaches to accommodate the diverse linguistic realities and communicative needs of English learners in an increasingly interconnected and multilingual global landscape. This chapter intends first to offer a broad perspective of how these ideologies have shaped language teaching and how they have clashed with research-based observations of multilingual and multidialectal communicative settings. We will give an overview of the relevant literature, ranging from foundational texts to more recent ones challenging the ‘ideal’ monolingual native speaker, and we will show how the above ideologies are still found in a rather pervasive way in the language teacher profession. This will be followed by an account of recent research conducted in teacher training environments aimed at showing ways to successfully gear future language teachers towards a new vision of language that contemplates diversity and hybridization as fundamental pillars on which teachers’ identities need to be based.
This study examines the impact of social media on the linguistic behavior of Jordanian Gen Z (born 1997–2012) through the lens of their daily use of colloquial speech as a reflection of sociocultural change. It delineates the dominant linguistic features of the language they use and attempts to address how these linguistic practices reflect the construction of identity and socio-cultural shifts among Jordanian Generation Z. Social media platforms such as TikTok, Instagram, and Snapchat heavily influence Generation Z's vernacular. This study employs a qualitative research approach to analyze pertinent data on code-switching, meme-driven expressions, and abbreviation combinations. Two primary methods of data collection were employed: social media data collection for discourse analysis and semi-structured interviews aimed at identifying the most frequently used expressions among Generation Z. Findings show that the vernacular of Jordanian Gen Z is dynamic, hybrid, and highly integrative in terms of global linguistic resources. This new digital Arabic sociolect poses numerous linguistic and cultural challenges for individuals. These include the necessity for extensive code-switching, the establishment of distinct online linguistic norms, the adaptation to cultural hybridity in language use, and the confrontation of linguistic divergence between generations.
Screen use pervades daily life, shaping work, leisure, and social connections while raising concerns for digital wellbeing. Yet, reducing screen time alone risks oversimplifying technology’s role and neglecting its potential for meaningful engagement. We posit self-awareness—reflecting on one’s digital behavior—as a critical pathway to digital wellbeing. We developed WellScreen, a lightweight probe that scaffolds daily reflection by asking people to estimate and report smartphone use. In a two-week deployment with college students (\(\mathtt {N}\)=25) focused on generating formative insights, we examined how discrepancies between estimated and actual usage shaped digital awareness and wellbeing. Participants often underestimated productivity and social media while overestimating entertainment app use. They showed a 10% improvement in positive affect, rating WellScreen as moderately useful. Interviews revealed that structured reflection supported recognition of patterns, adjustment of expectations, and more intentional engagement with technology. Our findings highlight the promise of lightweight reflective interventions for supporting self-awareness and intentional digital engagement, offering implications for designing digital wellbeing tools.
This paper examines the ability of language models to capture semantic relations between words in a low-resource language. We describe experiments on automatic prediction of lexical-semantic relations in Belarusian using models of Word2Vec, BERT, and LLM families which differ in neural architecture, feature types and NLP applications. Training and f ine-tuning of the models was carried out on the datasets compiled for our study: Belarusian corpora with UD POS tagging and a database of synonyms and antonyms extracted from Belarusian dictionaries. Model performance was evaluated by pseudo-disambiguation test (Word2Vec CBOW and skip-grams) as well as by expert assessments (roberta-small-Belarusian, Gemini 2.5 Pro). The results proved to be valid and can be applied to create and enrich lexical databases, to analyse word co-occurrence, to improve machine translation, paraphrasing, summarization, and other systems related to automatic processing of the Belarusian language.