Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Abstract Dependency distance (DD) and hierarchical distance (HD) measure syntactic complexity linearly and hierarchically, reflecting comprehension and production preferences, respectively. Using syntactic dependency treebanks of English e-commerce live-streaming by American (CAL) and Chinese (CCL) hosts, this study compares the two hosts’ syntactic complexity via mean dependency distance (MDD) and mean hierarchical distance (MHD). The results show significant MHD differences between CAL and CCL but no MDD differences. Both treebanks follow Zipf’s law in DD/HD distributions, with short sentences dominating. While nine dependency types overlap in the top 10 frequent relations across three sentence lengths, some dependency types, such as discourse and aux, are more frequent in CAL. Notably, CCL exhibits significantly higher MHDs, a finding that is mainly attributed to structures like complex sentences, coordinate clauses, that -clauses, and quantifier phrases. The findings further suggest a universal tendency to minimize syntactic complexity in English live-streaming, with native language transfer primarily influencing hierarchical complexity (MHD).
Analyzing the structure of complex sentences, comprising multiple clauses within a single sentence, has been recognized as a challenging aspect of dependency parsing. This challenge arises from the syntactic ambiguity inherent in clausal relations, wherein the main verb of each clause may have multiple candidate dependency heads. Minami's Scope Preference Theory (1964, 1974) hypothesized that clausal relations are determined by preferences between pairs of subordinate clauses. Building upon this theoretical foundation, we introduce a neural model fine-tuned to resolve clausal relations, with a focus on the verbs within each clause. To address the challenges in dependency parsing, we propose a forest reranking approach that enables our reranking model to grasp global context. Our approach utilizes our model as a reranker on dependency forests by leveraging cube-pruning for efficient tree enumeration. Through a series of experiments and analyses conducted on the Penn Treebank (PTB) and Penn Chinese Treebank (CTB), we find that our approach is effective in resolving the ambiguity inherent in complex sentence structures.
ABSTRACT The widespread use of TikTok among elementary school students has brought noticeable changes to the way children communicate in their daily lives. The platform is no longer used merely as a source of digital entertainment, but has also begun to shape students’ word choices, speaking styles, and language habits. This condition can be observed among students at MIS Al-Khairaat Pombewe, who have become increasingly familiar with viral expressions, popular abbreviations, slang, and the mixing of Indonesian with foreign languages in everyday conversations. Such circumstances have raised concerns regarding the declining use of proper and standard Indonesian within the school environment. This study employed a descriptive qualitative approach involving the principal, teachers, and students selected purposively as research informants. Data were collected through observations, interviews, and documentation, then analyzed through the stages of data reduction, data presentation, and conclusion drawing. The findings reveal that TikTok exerts a dual influence on children’s language development. On the one hand, the platform contributes to vocabulary expansion, enhances students’ creativity in language use, and broadens their digital knowledge. On the other hand, the intensity of TikTok usage encourages the frequent use of informal language in formal situations, leading to a gradual decline in the use of proper Indonesian according to linguistic norms. Therefore, the involvement of teachers and parents is necessary to guide children toward wiser social media use without neglecting the development of their language abilities. ABSTRAK Fenomena penggunaan TikTok di lingkungan sekolah dasar memperlihatkan perubahan yang cukup nyata pada cara siswa berkomunikasi sehari-hari. Platform ini tidak lagi sekadar dimanfaatkan sebagai hiburan digital, tetapi turut membentuk pilihan kata, gaya berbicara, hingga kebiasaan berbahasa anak. Kondisi tersebut terlihat pada siswa MIS Al-Khairaat Pombewe yang semakin akrab dengan istilah viral, singkatan populer, bahasa gaul, serta pencampuran bahasa Indonesia dengan bahasa asing dalam percakapan mereka. Situasi ini memunculkan perhatian terhadap menurunnya penggunaan bahasa Indonesia yang baik dan benar di lingkungan sekolah. Kajian ini memanfaatkan pendekatan deskriptif kualitatif dengan melibatkan kepala sekolah, guru, dan siswa sebagai informan yang dipilih secara purposive. Informasi penelitian diperoleh melalui observasi, wawancara, dan dokumentasi, kemudian dipahami melalui tahapan reduksi data, penyajian data, dan penarikan kesimpulan. Temuan penelitian memperlihatkan bahwa TikTok memberi pengaruh ganda terhadap perkembangan bahasa anak. Di satu sisi, media sosial tersebut membantu siswa memperluas kosakata, meningkatkan kreativitas dalam berbahasa, dan memperkaya wawasan digital mereka. Di sisi lain, intensitas penggunaan TikTok ikut mendorong penggunaan bahasa informal dalam situasi formal sehingga kebiasaan menggunakan bahasa Indonesia sesuai kaidah menjadi semakin berkurang. Karena itu, keterlibatan guru dan orang tua dibutuhkan agar penggunaan media sosial dapat diarahkan secara lebih bijak tanpa mengabaikan perkembangan kemampuan berbahasa siswa.
Language-model representations provide structured, high-dimensional annotations of naturalistic language stimuli and can serve as informative neural predictors during comprehension. We analyzed locked derived data from Brain Treebank, MEG-MASC, and Podcast ECoG with eight frozen language models, blocked encoding models, and matched temporal, nuisance, and representation-capacity controls. Positive held-out prediction and gains over low-level baselines were widespread in source-level summaries. Across Brain Treebank and Podcast ECoG, 67 of 432 evaluable rows met a controlled predictive-only criterion, and model-side feature ablations changed prediction scores in most evaluable source rows. Brain-derived, timing-linked, acoustic, and implanted-signal controls confirmed component-level sensitivity of the analysis pipeline. These findings show that language-model-derived quantities can annotate neural activity during natural speech and text comprehension. Participant-level matched-control advantages were localized rather than uniform, response-profile and feature-specificity contrasts bounded representational or computational interpretations, and complete co-indexed integrated interpretation will require future jointly indexed coverage. Together, the analyses identify language-model features as useful neural predictors and separate predictive usefulness from claims about shared neural organization or language-processing computations.
The purpose of the article is to identify the features of communicative practices implemented in the educational content of Ukrainian language TikTok blogs and to determine their role in the development of users’ language skills in the online communicative environment. Research methodology involved the use of the following methods: content analysis – to identify communicative practices for promoting the Ukrainian language on TikTok; analysis and synthesis, observation, and generalisation – to process educational linguistic content and identify communicative strategies for audience engagement; sociocultural analysis – to outline the role of social networks in shaping language identity and collective awareness of linguistic norms. Scientific novelty of the article lies in the comprehensive analysis of communicative practices of TikTok blogs promoting the Ukrainian language as a factor in the development of users’ language skills, in presenting interactive strategies used by bloggers to engage audiences in mastering language norms, and in forming language identity. Conclusions. Communicative practices of TikTok bloggers develop language skills through interactive communication, word games, and role-playing dialogues, involving the audience in joint analysis of the features of the Ukrainian language. Users become co-creators of content: they comment, suggest alternative forms, create humorous mnemonic constructions, and analyse the presented examples, thereby becoming aware of the cultural and historical context of linguistic norms. Such communicative technologies take the form of collective language activity, shaping language identity and strengthening the sense of belonging to the linguistic community. The format of online platforms, particularly TikTok – short videos, recommendation algorithms, interactivity, and virality – contributes to the promotion of language culture. At the same time, the specifics of clip-based thinking and the tendency toward conciseness may lead to fragmented presentation of material and simplified representation of complex historical-linguistic processes. Digital communicative practices act as a factor in the linguistic mobilisation of Ukrainian society, while simultaneously creating potential challenges in the context of preserving and disseminating language norms.
The authors see the purpose of the study as a comparative analysis of political discourse on the example of D. Trump’s speeches during his election campaigns in 2016 and 2024. The scientific novelty lies in the confirmation of the concept of using language as a weapon, which acts in Trump’s speeches as a tool to manipulate and control people through different discursive means and in different periods of time. Comparing the changes that discourse has undergone over time allowed the authors to analyse the priorities that reflect the general political tendency for positive emotional encouragement rather than threats and aggression. The relevance of this article is determined by the shift that has occurred in political discourse, which requires a rethinking of how the political actors select linguistic norms and how this selection will affect the formation of modern political language.
Cet article présente la dernière version du treebank Rhapsodie, un corpus de français parlé multi-genres annoté en syntaxe et prosodie. Les deux principales innovations sont une annotation morphosyntaxique réellement basée sur la version orale du corpus (ce qui est prononcé) et non sur sa transcription orthographique et une intégration de l’ensemble des niveaux d’annotations, syntaxe, prosodie et métadonnées, dans une même structure, permettant ainsi des requêtes croisées.
What constrains working memory capacity? Classic theories place visual working memory close to perceptual systems, with fixed limits. Yet, emerging evidence shows that visual working memory capacity is increased for real-world objects compared to simple or abstract stimuli. The present study demonstrates that this memory advantage arises from semantic understanding of real-world objects – contrary to classic perceptual accounts of this cognitive system. Using counterfeit objects generated by generative adversarial networks that match real objects in terms of object form and visual similarity, we show that improvements in behavioral performance and increases in neural delay activity emerge solely for semantically meaningful, real objects. Correlation analyses indicate that subjective familiarity ratings predict memory for real objects, whereas stimulus colourfulness predicts memory for artificial objects, suggesting distinct mechanisms support memory for different stimulus types. Thus, conceptual knowledge exerts strong effects on visual working memory, significantly extending current theories that emphasize low-level perceptual features.
This study focused on structural and semantic changes in Kazakh borrowed vocabulary adaption. The study examined how imported concepts were absorbed into Kazakh, altered by the national linguistic system, and contributed to current terminology. Various linguistic methodologies were used, including historical and contemporary text analysis, structural and comparative term analysis, and hybrid word classification. The study covered both present and historical English, allowing vocabulary changes to be tracked. The findings showed that complicated historical and cultural processes in active intercultural exchanges caused Kazakh terminology hybridisation. The Greco-Latin, Persian, and Arabic languages enriched Kazakh lexicon with science, religion, culture, and daily life concepts. Phonetic, morphological, and visual changes were made to borrowed terminology to make them useful and conform to Kazakh linguistic norms. The study showed that hybrid terms are crucial to borrowing integration.
The study explores the attitudes and opinions of Pakistani English teachers on Standard British English (SBE) and Standard American English (SAE). Although most studies have been done on learner attitudes, this paper refracts the same to the teacher, whose role is very important in influencing linguistic norms in EFL. A survey involving 60 English teachers in the government and privately owned institutions in Pakistan was conducted using a mixed-method approach, whereby a questionnaire, which included closed-ended and open-ended questions, was used to gather the data. The results indicate that there is an acute effect of academic preparation of teachers, exposure to media, and the practices of the institution on the preference of variety among the teachers. The paper also examines the relationships between the linguistic backgrounds and pedagogical decision-making amongst teachers. The findings can be added to the current discussion of World Englishes and can be applied to the teaching training and language policy in Pakistan.
Reviewer assignment is increasingly critical yet challenging in the LLM era, where rapid topic shifts render many pre-2023 benchmarks outdated and where proxy signals poorly reflect true reviewer familiarity. We address this evaluation bottleneck by introducing LR-bench, a high-fidelity, up-to-date benchmark curated from 2024-2025 AI/NLP manuscripts with five-level self-assessed familiarity ratings collected via a large-scale email survey, yielding 1055 expert-annotated paper-reviewer-score annotations. We further propose RATE, a reviewer-centric ranking framework that distills each reviewer's recent publications into compact keyword-based profiles and fine-tunes an embedding model with weak preference supervision constructed from heuristic retrieval signals, enabling matching each manuscript against a reviewer profile directly. Across LR-bench and the CMU gold-standard dataset, our approach consistently achieves state-of-the-art performance, outperforming strong embedding baselines by a clear margin. We release LR-bench at https://huggingface.co/datasets/Gnociew/LR-bench, and a GitHub repository at https://github.com/Gnociew/RATE-Reviewer-Assign.
Abstract Research on how non-natives process and learn binomials ( black and white ) is limited. The present study addresses this gap using online (eye-tracking) and offline (familiarity rating) tasks. Sixty non-native speakers of English (L1 = Arabic) read six stories seeded with 21 novel binomials in three conditions: one exposure, six exposures, and no exposure (i.e., only in post-test) in a counter-balanced design. Each item was also presented in the reversed order ( white and black ). The non-natives read the stories as their eye movements were monitored and answered comprehension questions. In addition to the novel binomials, 12 existing binomials (congruent with Arabic) were included in the passages as a baseline for comparison. After completing the reading task, the participants completed an offline rating task as a measure of declarative knowledge of the binomial configuration (i.e., word order). All items were rated twice, once in the forward direction and once in the reversed direction. Online results showed that non-natives were not sensitive to the configuration of existing binomials, and there was limited evidence of any sensitivity to novel binomials. Offline, non-natives showed sensitivity to the configuration restrictions of existing binomials but not novel ones.
In Spanish linguistics today, it is widely recognized that Spanish corresponds to the image of a pluricentric language. Different normative centers coexist, which are perceived as such by the speakers. However, the degree of recognition, status, and prestige of these linguistic norms varies greatly. This article examines the extent to which the pluricentrism of Spanish is represented in textbooks for Spanish as a foreign language. As a case study, three textbooks – Puente Nuevo, ¡Adelante!, and A_tope.com – used in high schools in Basel, Switzerland, are analyzed both quantitatively and qualitatively in terms of their pluricentric character. The study examines the textbooks as a whole (thematic priorities, topics of units, maps) and selected linguistic phenomena, namely: the forms of address (morphosyntax), seseo (pronunciation), and the use of regional vocabulary. At various levels, it can be demonstrated that the textbooks are based on a Eurocentric view of language, which gives Castilian Spanish a superior role to that of other language norms, without this being explicitly stated. The article concludes with some practical recommendations for a more pluricentric approach to teaching Spanish.
This chapter explores the uneven geographies of access, participation, and belonging in global higher education, focusing on how language shapes the lived experiences of international students and faculty. Drawing on classroom-based narratives from India and Oman, it examines how linguistic norms, institutional expectations, and cultural assumptions determine inclusion and exclusion in academic spaces. While student mobility is often framed as a success of globalization, the chapter argues that it remains embedded in hierarchies of language, identity, and geography. Multilingual learners from rural or non-elite backgrounds, and faculty from “non-native” English-speaking contexts, often face marginalization due to misalignment with dominant academic norms. Using Bourdieu's linguistic capital, postcolonial critiques, and critical internationalization studies, the chapter calls for moving beyond tokenistic diversity. It advocates for multilingual pedagogies, translanguaging, and ethical student mobility to build a more equitable and culturally responsive global education landscape.
Affective word values have been widely studied across languages, often focusing on isolated words due to the difficulty of assessing emotionality in texts. This study examines whether written emotional content can be reliably captured using a specific software tool (Watson Natural Language Understanding). Thirty-three Spanish undergraduates wrote 150-word autobiographical texts in their L2 (English) before and after a training with emotional vocabulary. Normative valence ratings of content words obtained in the pre- and post-training phases were compared with sentiment scores generated by Watson NLU. Strong positive correlations were found between sentiment and normative valence scores in both phases, with stronger relations at post-training. Regression analyses confirmed that sentiment scores significantly predicted normative valence. Importantly, while normative valence did not differ between phases, sentiment scores increased after training. These results suggest that Watson NLU is a valid and sensitive tool for assessing emotionality in written language and its modulation through training.
The rules that determine which assets count as eligible collateral for central bank operations, and at what haircuts, are not just operational details. They have become a first-order determinant of asset prices, liquidity allocation, and financial stability. This review synthesizes the literature on three transmission channels: convenience yield, collateral scarcity, and market liquidity. I trace the intellectual lineage from Singh and Stella (2012) through Williamson (2016) to the empirical studies of Nyborg and Woschitz (2021), Lengwiler and Orphanides (2024), and Fang, Wang, and Wu (2020). My central argument is that this literature, taken as a whole, reveals a fundamental policy trilemma. Central banks must choose among unconditional acceptance of their own government’s debt (which risks fiscal dominance), rating-based eligibility (which risks self-fulfilling sovereign crises), and discretionary policy-driven eligibility (which risks politicization). No design is safe. I also identify four open questions: the nonlinearity of the collateral channel, its interaction with bank portfolio behavior, the systemic risk of cliff effects, and the external validity of evidence from China.
ABSTRACT This study investigates the perceptions of Americanisms among three generations of Nigerians. While prior research has provided quantitative evidence for American influence in contemporary Nigerian English, the role of language beliefs and ideologies in mediating such changes remains underexplored. Developing a sociolinguistic perspective of mobile linguistic resources, this study construes an individual's linguistic repertoire as an identity‐construction resource, agentively mobilised across geographical, social and digital spaces. Interview data indicate that younger speakers orient towards multiple linguistic norms, while older speakers remain critical of Americanisms and favour British norms. Reading task results further indicate that American realisations are most frequent among younger speakers. The study demonstrates that multinormativity extends beyond linguistic production to speakers’ evaluative orientations and perceived repertoires. This finding advances the sociolinguistics of mobility and World Englishes research by showing that shifting language ideologies – rather than usage patterns alone – constitute a key mechanism driving linguistic change in postcolonial varieties.
Child-directed fingerspelling is an approach used by Deaf parents for communication, language, and literacy development. This study reports on findings from a qualitative intrinsic case study aimed at understanding how Deaf parents use fingerspelling with their young children. The research questions were: (1) What are the cultural beliefs of Deaf parents regarding fingerspelling with young children? (2) What are their patterns of use of child-directed fingerspelling in natural settings? Twenty-one Deaf families with 27 deaf children ages 5 years and under were interviewed via recorded Zoom meetings conducted in American Sign Language. Data were analyzed using grounded theory to develop a new theoretical contribution with the core category: Deaf families socialize their children into Deaf visual-linguistic norms through fingerspelling. This new theoretical insight aligns with Holcomb's Deaf epistemological framework (2010) and Ochs and Schieffelin's (2008, 2011) language socialization theory. Limitations and recommendations for future research are also included.
AI-mediated communication refers to communicative processes in which artificial intelligence systems actively generate, interpret, modify, or facilitate language. With the rapid advancement of language technologies such as large language models, conversational agents, speech recognition systems and machine translation tools, AI has evolved from a supportive tool into an active mediator of human communication. This paper conceptualizes AI-mediated communication as a form of language technology that reshapes traditional models of interaction, meaning-making and authorship. The study examines the technological foundations of AI-driven language processing and their broader socio-cultural, pedagogical and ethical implications. It explores how AI mediation influences linguistic norms, accessibility and power relations, while also raising critical concerns related to bias, surveillance and linguistic homogenization. By situating AI-mediated communication within contemporary digital culture, the paper argues that understanding AI as a language technology is essential for critically engaging with evolving communicative practices and for developing responsible, inclusive and ethically grounded AI-driven communication systems.
SRC, an acronym for Stimulus-response correlation, refers to determining the relationship between stimulus and corresponding brain responses. The neural aesthetic resonance hypothesis proposes that the level of enjoyment or familiarity can be distinguishable based on the relationship between stimulus and brain responses. To test this hypothesis, we use EEG data of 20 participants listening to 12 songs with their enjoyment and familiarity ratings. We aim to classify the low and high ratings of familiarity and enjoyment based on SRC. Eighteen musical features are extracted and transformed into the first principal component (PC1). In addition, root mean square (RMS) and spectral flux are used for analysis. Canonical Correlation Analysis (CCA), an unsupervised AI optimization method, is employed to compute the SRC between musical features and ten regions of brain responses, followed by considering four principal CCA features for classification using the Random Forest classifier with cross-subject evaluation. Our results demonstrate that the right frontal and right parietal regions provide significant predictive ability. Our empirical finding suggests that RMS features preserve the predictive ability for familiarity, whereas PC1 is for enjoyment prediction. Maximum familiarity and enjoyment accuracy reach nearly 76% and 73% accuracy. This work leverages AI techniques to decode sensor-derived neural signals, advancing real-time applications in affective computing and wearable EEG devices.
This study examines how linguistic adaptation, inclusion, and student diversity are constructed in the Swedish national curriculum from 2025 for upper secondary education (Gy25), with a particular focus on the subject syllabus for Swedish. The aim is to analyse the assumptions about language, learning, and student roles embedded in the curriculum text. The analysis is based on a qualitative text analysis with elements of critical discourse analysis. Selected sections of the curriculum, including the general aims and the subject syllabus for Swedish, constitute the primary material. The analysis focuses on key concepts, modal expressions, and how students and language are represented in the policy text. The results indicate that language is constructed as a central, norm-governed competence that students are expected to develop through education. At the same time, the curriculum emphasises inclusion and the need to adapt teaching to students’ different conditions. The analysis suggests that the curriculum contains a tension between inclusive ideals and established linguistic norms that students are expected to meet. The study highlights how curriculum texts contribute to shaping assumptions about language, participation, and student roles in the Swedish subject.
Pre-trained language models (PLMs) achieve high accuracy on standard benchmarks for sentiment analysis. However, this performance can hide systematic weaknesses in determining the sentiment of negated sentences, for example when the phrase “not good” is still classified as positive. In this study, we use sentiment classification of English movie reviews in the Stanford Sentiment Treebank 2 (SST-2) as a case study to specifically examine and improve how BERT handles negated sentences. We perform a brief additional fine-tuning of the existing BERT model on a small, automatically constructed set of lexicon-based counterfactual examples that target simple lexical negation. Experimental results on carefully paired original-negated sentences show that this procedure substantially reduces prediction errors on negated inputs while leaving overall performance on SST-2 almost unchanged.
Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurally via percolation rules. We treat this notion of constituent headedness as an explicit representational layer and learn it as a supervised prediction task over aligned constituency and dependency annotations, inducing supervision by defining each constituent head as the dependency span head. On aligned English and Chinese data, the resulting models achieve near-ceiling intrinsic accuracy and substantially outperform Collins-style rule-based percolation. Predicted heads yield comparable parsing accuracy under head-driven binarization, consistent with the induced binary training targets being largely equivalent across head choices, while increasing the fidelity of deterministic constituency-to-dependency conversion and transferring across resources and languages under simple label-mapping interfaces.
FrameNet is an English-based lexical database that shows how words are used by providing information as to which participants and relations are evoked by a certain concept. Recent efforts toward a multilingual FrameNet have not targeted either ancient languages or different historical stages of the same language. In our paper we propose creating a multilingual FrameNet for Ancient Indo-European languages starting with a set of 80 verb meanings annotated in the Pavia Verb Database. Our pilot study includes four verb meanings: RAIN, THUNDER, SEE, LOOK AT. As the adequacy of the semantic frames developed for English turns out not to be appropriate for the languages in our sample, we propose two new frames that can account for the analyzed data.
This research explores the historical emergence of linguistic terminology in three languages—English, Uzbek, and Karakalpak—with special attention to the role of Latin, Greek, and Arabic heritage. It traces how borrowed concepts were nativized and localized in each linguistic setting. By juxtaposing five evolutionary stages in English with analogous processes in Uzbek and Karakalpak, the paper illustrates the interplay between international scholarly traditions and indigenous linguistic norms. The conclusions highlight both universal tendencies and language-specific particularities in the growth of terminological systems.
This paper presents a small-scale dependency treebank for Tunisian Arabic (TADT) developed within the Universal Dependencies framework, addressing the scarcity of linguistic resources for the Arabic varieties.The approach employs domain adaptation, leveraging a machine learning model (UDPipe 1.0) trained on Algerian Arabic data to annotate 100 Tunisian Arabic social media comments, followed by manual correction.This pilot study evaluates the feasibility of using machine learning-assisted annotation to scale resource development for spoken Arabic and identifies key challenges in cross-dialectal transfer for improving annotation quality and efficiency.This work contributes to more inclusive and fair representation of Arabic linguistic varieties in academic research and NLP applications.
The paper presents a prototype of a web-app designed to automatically generate verb valency lexica based on the Universal Dependencies (UD) treebanks.It offers an overview of the structure of the app, its core functionality, and functional extensions designed to handle treebank-specific features.Besides, the paper highlights the limitations of the prototype and the potential of its further development.
Abstract Introduction Sleep supports emotion regulation by preferentially consolidating emotional memories while attenuating reactivity. We have shown that dream recall plays an active role by increasing negative over neutral memories and reducing reactivity. In women, fluctuating reproductive hormones across the menstrual cycle influence sleep features implicated in emotional memory, yet whether menstrual phases influence how dreams shape emotional processing remains unknown. This study investigates how dreams shape sleep-dependent emotional processing across the menstrual cycle in naturally cycling women. Methods 128 women (Mage = 32.85 ±11.93 years) completed up to four visits across verified menstrual phases (menses, late-follicular, mid-luteal, late-luteal). At each visit, participants performed the Emotional Picture Task with negative and neutral IAPS images in the evening (Test 1) and the next morning (Test 2). Participants rated old/new, arousal, and valence of images shown at each test. Dream reports were collected upon waking prior to Test 2. Linear mixed-effects models tested main and interaction effects of menstrual phase and dream recall. Results The menstrual cycle altered how dreaming shaped overnight emotional memory. Dream recall typically benefited the emotional trade-off effect —favoring consolidation of negative relative to neutral images (Δd′; t(410)=1.95, p=0.05)—but this pattern reversed during the late-luteal phase (dream × menstrual cycle: t(381)=-2.29, p=0.02). Dreaming showed independent effects on emotional reactivity. Higher valence and arousal ratings for negative images during Test 1 predicted greater dream recall (valence: t(344)=2.05, p=0.04; arousal: t(327)=2.04, p=0.04). Additionally, the more negatively participants rated the images at Test 1, the more negative their dreams tended to be (t(166)=-2.11, p=0.04). Dream recall was linked to reduced next-morning emotional reactivity (valence: t(413)=-2.89, p=0.004; arousal: t(413)=-2.65, p=0.01), with stronger reductions following more negative dreams (β=0.15, t(182)=2.86, p=0.005). Conclusion Menstrual cycle phase influenced how dreams shaped overnight emotional memory. Negative waking experiences increased dream recall and shaped dream content—and recalling dreams, especially negative ones, reduced emotional reactivity and typically strengthened emotional memory—but this benefit disappeared in the late-luteal phase when there are declining reproductive hormones. These findings suggest a novel interaction between the menstrual cycle and dreaming, showing that hormonal fluctuations reshape how sleep and dreams regulate emotional experience and memory. Support (if any) RF1AG061355 (Baker/Mednick)
Le French Treebank: une ressource lexicale et syntaxique richement annotée (et validée manuellement) pour les linguistes, utilisable en TAL, dans sa version 2.0 Projet initié en 1997, avec le soutien de l'IUF, du CNRS et du CNRTL21 550 phrases (environ 664 500 tokens) du journal Le Monde (1990-1993)Métadonnées: auteur, date, domaine (par article)Annotations lexicales (catégories, sous-catégories, flexion, mots composés avec composants) et syntaxiques (constituants majeurs, fonctions grammaticales) validéesPlusieurs formats disponibles: XML, PTB, CoNLL, codage UTF-8 (ligature œ notée oe)Nouveautés de la version 2.0: plusieurs erreurs d'annotation ont été corrigées depuis la publication en 2016 de la version 1.0
The Junction Grammar of English A Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and Consonant–Vowel Information Asymmetry Boicho Dimitrov Temelakiev Saxon Ventura Research Ltd 28th of May,2026 CC BY Abstract This paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling. The morpheme boundary—the junction between a stem and an affix—is treated as the primary object of measurement, and the distribution of the letters that may occupy each side of a junction is read across the corpus. Three results are established. First, the junction carries a stable structure: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a sparsity that scales inversely with a suffix’s productivity. Second, the relationship between any two affixes is quantified by the correlation of their boundary distributions, which measures the degree to which they share a stem population; this correlation ranges from ~0.99 for etymological doublets to 0.57 for productivity-asymmetric near-twins. Third, the written word decomposes into a consonant skeleton carrying lexical identity (53.0% of the vocabulary is uniquely recoverable from consonants alone) and a vowel tissue carrying grammatical form (1.8% recoverable from vowels alone), a ~29-fold information asymmetry. Multiple independent measures—consonant recoverability, derivational classmarking, and free-stem fraction—partition the lexicon at a single boundary separating a transparent Germanic core from a bound Latinate superstructure. The method, its corrections, and its limits are reported in full, and a program of remaining work is stated. 1. Corpus and Method All results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution. The corpus is the dwyl/english-words list (words_alpha.txt), comprising 455,246 alphabetic entries with a total of 4,254,354 letter occurrences and a mean word length of 9.345 letters. Each letter is assigned its ordinal value (A=1 through Z=26). Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter, so that each suffix position returns its exact ordinal value as a sanity check; the first stem letter preceding the suffix—the linker—is then read as a full A–Z frequency distribution rather than as a mean. The governing methodological constraint is that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, and that every claim is backed by a count. Page 1 Method: a derived column applies the ordinal map; end-anchored COUNTIF and MID/CODE formulas extract suffix positions and the linker; per-letter tallies at the linker yield the boundary distribution. Nesting among suffixes (for example ‑MENT within ‑ENT, or ‑ATION within ‑TION within ‑ION) is controlled by excluding longer relatives before counting. No statistical software, machine-learning library, or external lexical database is used at any stage of the core analysis. The constraint that the analysis remain computable by elementary means is not incidental; it ensures every figure in this paper can be independently reconstructed from the public corpus with a spreadsheet alone. 2. The Junction and Its Backbone The investigation began as a survey of unconditioned letter bigrams, which returned no morphological signal; structure appeared only when letter distributions were conditioned on a single morpheme boundary, and that conditioning is the method’s foundation. A preliminary tabulation of adjacent letter pairs across the corpus—which letters follow which, without regard to position within the word—yielded frequency patterns reducible to general orthographic regularities and carrying no isolable morphological content. The signal emerged only when a specific suffix was fixed and the letters preceding it were read as a population: conditioning on the junction, rather than measuring adjacency as such, is what renders the boundary structure visible. The set of letters that may legally occupy the stem side of a morpheme boundary then proves narrow, structured, and stable across suffixes. Reading the linker distribution across the mapped suffix inventory reveals a three-tier structure. A backbone of six letters—T, L, N, R, S, and I—is admitted at high frequency by nearly every suffix; T is the single most frequent linker across the inventory and recurs as the dominant boundary letter in suffix after suffix. An absolute floor of two letters—J and Q—is admitted by no suffix at measurable frequency, a prohibition confirmed corpus-wide and consistent with their status as the two rarest letters overall (J at 0.18%, Q at 0.19% of all letter occurrences). Between backbone and floor lies a selective middle whose composition varies by suffix and in which each suffix’s identity resides. Two descriptive terms are used throughout. Because each letter carries an ordinal value (A=1 through Z=26), a linker distribution may be summarized by where its mass falls on that scale: a distribution concentrated on early-alphabet letters (low ordinal values, A through roughly M) is termed cold, and one concentrated on late-alphabet letters (high ordinal values, roughly N through Z) is termed warm. The terms refer solely to ordinal position on the A–Z scale and carry no semantic content; the backbone letters, for instance, span both ends (cold I and L against warm N, R, S, T). The degree of selectivity is itself a measurement. The count of forbidden letters at a junction—its sparsity—scales inversely with the suffix’s productivity: derivational suffixes that attach choosily to a constrained stem class forbid many letters, whereas inflectional or highly productive suffixes forbid few. Sparsity is therefore not noise but signal: the pattern of exclusion characterizes the suffix as informatively as the pattern of admission. Page 2 Method: forbidden-letter counts are read directly from the linker distribution (frequency below 0.5% taken as floor); cross-checks against COUNTIF totals confirm the populations; J/Q absence is verified against the bare ‑S population of 84,208 words and against corpus-wide letter frequencies. The junction thus behaves as a filter whose admitted and forbidden letters together encode the morphological role of the boundary, with the consonant backbone bearing the structural load and the rare letters marking its limits.
Language is a living organism that evolves alongside technological and social advancements. This paper examines the phenomenon of neologisms—newly coined words or expressions—and their pervasive role in contemporary English mass media. The study categorizes recent neologisms based on their morphological formation processes, such as blending, compounding, and functional shift. Furthermore, it analyzes how mass media acts as a primary catalyst for the popularization of these terms. By investigating digital journals, social media platforms, and news broadcasts, the research highlights the pragmatic functions of neologisms in creating concise, engaging, and culturally relevant communication. The findings provide insights into the current trends of English lexicology and the impact of the digital age on linguistic norms.
This paper presents a direct framework for sequence models with hidden states on closed subgroups of U(d). We use a minimal axiomatic setup and derive recurrent and transformer templates from a shared skeleton in which subgroup choice acts as a drop-in replacement for state space, tangent projection, and update map. We then specialize to O(d) and evaluate orthogonal-state RNN and transformer models on Tiny Shakespeare and Penn Treebank under parameter-matched settings. We also report a general linear-mixing extension in tangent space, which applies across subgroup choices and improves finite-budget performance in the current O(d) experiments.
Abstract This paper examines the social contexts in Thomas Mann’s novel Buddenbrooks where the North German dialect Plattdeutsch is spoken. Beyond the technical challenges of translating these passages, the analysis focuses on the literary representation of code-switching that functions primarily as socially and symbolically charged act. Drawing on Pierre Bourdieu’s sociological theory, the study interprets deviations from linguistic norms and dialect use as instances of double negation – a strategy that appears to challenge social conventions but, in reality, affirms the most valuable social capital in classical bourgeois society: the certainty that one’s high status remains unthreatened. The difficulty of translating such passages stems from the specific cultural parameters embedded in the novel. Ultimately, the paper argues that culture is not an immediate given but requires analytical frameworks from the social sciences for proper understanding and interpretation.
This article examines the role and significance of language corpora and linguistic databases in linguistic expertise. It discusses the possibilities of conducting objective semantic, pragmatic, and stylistic analyses of texts through corpus-based methods. The study also highlights the contribution of linguistic databases and artificial intelligence technologies to improving the accuracy, reliability, and efficiency of expert conclusions. Furthermore, the relevance of developing specialized corpora and databases for forensic linguistics in Uzbekistan is substantiated.
Part of speech and syntactically annotated dataset for modern Mongolian. The dataset is a fully annotated corpus of modern Mongolian texts written in Mongolian Cyrillic.
This paper presents the new Universal Dependencies tree bank for the Macedonian language, marking a significant step towards the comprehensive linguistic analysis of Macedonian within the UD framework.It briefly addresses dependency grammar from a theoretical perspective and moves on to describing the treebank development process, from sentence selection, word segmentation and lemmatization, to POS-tagging, morphological features, and dependency tagging.Given the mostly manual labor invested in developing the treebank, semiautomatic NLP tools specifically designed for this purpose have also been presented and commented.Based on the defined tagset, the paper provides examples and visualizations of annotated sentences.The creation of the first Macedonian UD treebank enhances Macedonian linguistic resources and provides valuable insights into the morphological and syntactic structures of this Balkan Sprachbund language.It also contributes to the broader understanding of language-specific challenges within the UD framework and facilitates cross-linguistic comparisons in the Balkan and broader region.
Introduction: Blood flow restriction (BFR) walking elicits improved fitness, but participants often report higher ratings of perceived exertion (RPE) and pain during BFR walking compared to non-BFR walking. The primary aim was to investigate how multiple BFR walking exposures might affect RPE and pain. Methods: 14 healthy, trained participants completed three BFR walking sessions on separate days. The treadmill speed that elicited 3/10 RPE while BFR was not applied was determined and that same speed was used during all BFR walking sessions. Participants walked for 15 minutes at the predetermined speed while 60% limb occlusion pressure was applied bilaterally to the thighs. RPE (0-10) and pain (0-10) were recorded during each minute of exercise. Two-way repeated measures analysis of variance determined if session (1-3) and/or time (1-15 minutes) affected RPE or pain. Statistical significance was established at p<0.05. Results: RPE was higher during session 1 compared to session 2 (minutes 8-15, 4.4±1.4 versus 3.6±0.8). RPE was higher during session 1 compared to session 3 (minutes 3-5, 3.5±0.7 versus 2.9±0.6; minutes 7-15, 4.3±1.3 versus 3.5±1.1). No significant differences were observed for pain. Conclusion: Participants might tolerate BFR walking better after completing two BFR walking sessions as lowered RPE responses were observed.
We present a manually curated dataset of Spanish second language and heritage learner writing following the Universal Dependencies framework.
Experiment 1 and Baseline Correction Materials The selected pictures (N = 320; negative and neutral) were divided into eight subsets (Image-set: affect_labeling, affect_matching, new_shorter_interval, new_long_interval). Immediate exposure included the affect-labeling and affect-matching sets. The short-interval and long-interval re-exposure phases additionally included their corresponding new sets. •The folder titled “Open_Data_Affect_Labeling_Self-reports_Reexposure” contains an Excel file named “image_selected_exp1_exp2,” which provides comprehensive information about the images used in the experiment 1 (sheet, exp1). The file includes each image’s filename, transformed normative ratings (1–9 Likert scale) for arousal and valence, and Picture Valence (negative, neutral). Additionally, the file indicates how each image was pseudorandomly assigned to a specific Image-set (affect_labeling, affect_matching, new_shorter_interval, new_long_interval). The Image-set corresponds to the Intervention Condition (affect labeling, affect matching, new_shorter_interval, new_long_interval). In the present study, only the affect labeling and affect matching conditions were included. DATA · The folder ‘Open_Data_Affect_Labeling_Self-reports_Reexposure’ contains an Excel file named ‘Exp1_arousal_3phases_ratings_diff’, and ‘Exp1_valence_3phases_ratings_diff’, which includes raw data on arousal and valence ratings collected for each trial per participant, along with their response times during the rating task, respectively. · The same columns in both files are: Subject, Image, Intervention_Condition, Pic_Valence, Emotion, Group, Phase · Within-subject factors o Intervention_Condition (labeled, matched) o Picture_Valence (negative, neutral) o Phase (immediate, short-interval, long-interval) · Between-subject factor: Group (high ADS, low ADS) Arousal Rating · Arousal_Tnorm: normative arousal rating of each image obtained from the Dutch norm data, with a 1-9 scale. · Dependent variables o Arousal_post: participants’ arousal ratings collected during the three reexposure phases following the intervention o ArousalRT_post: Response time (in ms) required to provide the arousal rating during the reexposure phases o diff_arousal: difference score calculated (Arousal_post − Arousal_Tnorm), which reflects the deviation of participants’ post-intervention ratings from the normative image rating (used for baseline correction) Valence Rating · Valence_norm: normative valence rating of each image obtained from the Dutch norm data, with a 1-9 scale. · Dependent variables o Valence_post: participants’ valence ratings collected during the three reexposure phases following the intervention o ValenceRT_post: Response time (in ms) required to provide the valence rating during the reexposure phases o diff_valence: difference score calculated (Valence_post − Valence_Tnorm), which reflects the deviation of participants’ post-intervention ratings from the normative image rating (used for baseline correction) Experiment 2 and Baseline Correction Materials The primary modification in Experiment 2 was the random assignment of images to intervention conditions. In addition, a feature labeling condition was introduced in the label-or-match task prior to the immediate reexposure phase. As a result, a total of 240 images were required: 40 negative and 40 neutral images for each of the three intervention conditions (affect labeling, affect matching, and feature labeling). Of these, 232 images were drawn from the stimulus set used in Experiment 1, and 8 additional images were newly included. •The folder titled “Open_Data_Affect_Labeling_Self-reports_Reexposure” contains an Excel file named “image_selected_exp1_exp2,” which provides comprehensive information about the images used in the experiment 2 (sheet, exp2). The file includes each image’s filename, transformed normative ratings (1–9 Likert scale) for arousal and valence, and Picture Valence (negative, neutral). All images are randomly assigned to the Intervention Condition(affect labeling, affect matching, feature labeling). DATA · The folder ‘Open_Data_Affect_Labeling_Self-reports’ contains an Excel file named ‘Exp2_arousal_3phases_ratings_diff’, and ‘Exp2_valence_3phases_ratings_diff’, which includes raw data on arousal and valence ratings collected for each trial per participant, along with their response times during the rating task, respectively. · The column structure is the same as in the Experiment 1 data files, with two exceptions: o There is no Group factor in Experiment 2 o The Intervention Condition factor includes an additional level: feature labeling (in addition to affect labeling and affect matching).