Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Following the reconceptualization of language learning in Chapter 2 as socially mediated participation rather than individual acquisition, one of the most deeply entrenched categories in language education becomes problematic: the native speaker. The native/non-native binary presupposes that linguistic competence is a stable possession anchored in biography. It assumes that authority derives from origin and that legitimacy correlates with proximity to a homogeneous and internally coherent linguistic norm. Yet a relational framework for understanding language learning destabilizes precisely these assumptions.
Abstract Color-evasive language practices and ideologies are central to the creation of clinical spaces as white public spaces, which are inherently dangerous for non-white people. This chapter explores how race, language, and disability are shaped by the color-evasive logic taught to medical students through linguistic practices within autism diagnostic processes. In these color-evasive practices, culture can index race in a socially acceptable way, as it allows individuals to talk about race without talking about race. In this clinical vignette, a Black child’s use of Black linguistic norms is described by the white clinician as being “inappropriate, disruptive, and echolalic.” This is reinforced when the clinician tells his students that scoring an autism diagnostic test is “not about culture, it’s about the response.” The chapter explores how the scoring of the diagnostic test is grounded in the subjective assessment of the clinician administering the test and how they understand the patient’s behavior, which is shaped by underlying anti-Black logics about race, language, and disability. The clinician’s response shows us that despite his production of a color-evasive framework, he perceives and hears the patient as a racialized other. The chapter argues that color-evasive ideologies create white public space where clinicians police non-white bodies and conflate a Black child’s use of Black language with disability.
Viktor Nozadze’s unpublished manuscript “The Social Norms in The Knight in the Panther’s Skin” consists of three envelopes and is unfinished, lacking both an introduction and a conclusion. The first envelope examines the lexeme “life,” its etymology, semantic nuances, dictionary definitions, and equivalents in foreign languages, though “death,” despite being mentioned, is not analyzed. The second focuses on the lexical units “face” and “faceless,” offering rich examples and insightful observations drawn from the poem. The third envelope briefly addresses various lexical units but is less thoroughly developed. Nevertheless, the manuscript represents significant scholarly effort and is valuable for researchers and readers interested in Georgian literature and culture.
This article examines the structural and semantic organization of advertising texts in Uzbek and English from a linguistic perspective. The study focuses on how advertising language functions not only as a means of providing information but also as a persuasive tool that influences consumer psychology and behavior. Special attention is given to semantic features such as denotative and connotative meanings, as well as Geoffrey Leech’s seven types of meaning: conceptual, connotative, social, affective, reflected, collocative, and thematic meaning. The research analyzes advertising slogans from different fields, including beverages, cosmetics, and clothing products, in order to identify the role of lexical choice, emotional coloring, stylistic devices, and gender-oriented language in advertising discourse. The article also compares the linguistic and cultural characteristics of Uzbek and English advertisements and demonstrates how advertising texts reflect social values, cultural norms, and consumer expectations. The findings show that advertising language is carefully structured to create emotional impact, attract attention, and increase the persuasive power of the message.
данная статья посвящена анализу этрусского текста на каменном памятнике из Сьерра-де-Медина (провинция Тукуман, Аргентина), с привлечением кабардино-черкесского языка (абхазо-адыгская языковая группа) к чтению и осмыслению исторических событий этрусков. Выявлено типологическое сходство структуры фраз с прямым порядком слов СПО (подлежащее-сказуемое-объект), что характерно для этрусского и кабардино-черкесского языков: эргативность (маркирование субъекта переходного глагола); агглютинативность (присоединение аффиксов с отдельными грамматическими значениями). Лексико-синтаксическая структура текстов соответствует эпиграфическим нормам древневосточных царских надписей. this article is devoted to the analysis of the Etruscan text on the stone monument from Sierra de Medina (Tucuman Province, Argentina), with the involvement of the Kabardian-Cherkess language (Abkhaz-Adyghe language group) in reading and understanding the historical events of the Etruscans. Typological similarities in the structure of phrases with the direct word order SVO (subject-verb-object) have been identified, which is typical for Etruscan and Kabardian-Cherkessian languages: ergativity (marking the subject of a transitive verb); agglutination (the addition of affixes with specific grammatical meanings). The lexical and syntactic structure of the texts corresponds to the epigraphic norms of ancient Eastern royal inscriptions.
This OSF repository contains the lexical data, reproducible code, derived phylogenetic character matrices, analysis outputs, and supplementary materials associated with the study “A Phonology-Aware Phylogeny for Asia Minor Greek”. The study investigates the internal phylogenetic structure of Asia Minor Greek varieties by comparing different representations of historical linguistic signal. Three principal phylogenetic datasets are constructed: (i) binary cognate characters, (ii) recurrent sound-correspondence characters inferred from phonetic alignments of etymologically related forms using LingPy and LingRex, and (iii) a combined cognate and sound-correspondence matrix. Phylogenetic inference is conducted using Bayesian analysis in MrBayes. The repository documents the complete workflow from the comparative lexical database to the final phylogenetic analyses. It includes the cleaned input data, data-format specifications, preprocessing scripts, random seeds and software settings required for reproducibility, intermediate alignment and correspondence-pattern outputs, cognate and sound-correspondence character matrices, Bayesian tree files, and diagnostic outputs. Detailed information about the structure of the input data and instructions for reproducing the analyses is provided in `README.md`.
The relevance of the study is determined by the limited number of works devoted to the linguistic problems of scientific dental discourse. In order to improve the professional language competence of specialists, the article presents a comparative analysis of the reference and actual competency models of the language personality of a modern dentist within the paronymic component of scientific professional discourse. Particular attention is paid to the importance of linguistic standardization of various genres of scientific dental texts for the advancement of medical science and professional communication. Materials and Methods. The empirical basis of the study consisted of various genres of scientific dental discourse, including research articles, conference abstracts, monographs, textbooks, teaching manuals, reference books, dissertations, and dissertation abstracts. The study employed general linguistic methods of scientific description, including generalization, systematization, classification, and interpretation of linguistic material. In addition, a descriptive-analytical method was applied to investigate the professional language practices of dentists, while comparative analysis was used to examine differences between the actual and reference competency models within the paronymic component. Results. The actual competency model of the dentist’s language personality was characterized and compared with the reference model in the paronymic domain. Numerous violations of language norms were identified, particularly lexical-grammatical and semantic-stylistic inconsistencies. Conclusions. Significant differences were identified between the actual and reference competency models of the dentist’s language personality, indicating an insufficient level of mastery and practical application of language norms. Difficulties in scientific dental discourse are largely associated with the incorrect use of paronyms and terminological inaccuracies, which may distort semantic interpretation. Bringing the actual language personality of a dentist closer to the reference model requires consideration of the social, cognitive, and verbal-semantic factors involved in its formation, as well as systematic improvement of linguistic competence.
The Self-Assessment Manikin (SAM) is one of the most widely used tools for measuring affect along the dimensions of valence and arousal, yet its abstract humanoid icons have been criticized for ambiguity, especially in representing arousal, and for lacking full gender neutrality. Despite newer alternatives, challenges remain in achieving clarity, inclusivity, and dimensional precision. To address this, we developed the Weather-Based Emotion Reporting (WER), a visual digital scale that represents affective states using familiar weather scenes. Grounded in normative affective data, WER was constructed to systematically map weather imagery onto the valence-arousal circumplex. In a within-subjects online experiment (N = 100), participants rated affective words using either WER or the SAM, allowing comparison of convergent validity, reaction times, and subjective usability.WER demonstrated a strong convergence with SAM, particularly for valence, while arousal showed moderate and more variable agreement, replicating well-documented asymmetry between affective dimensions. Reaction time analyses showed that WER responses were slightly faster than SAM responses, although the effect size was small. Participants consistently preferred WER, reporting greater clarity, comfort, and ease of use. Visual complexity analyses confirmed that WER and SAM differ qualitatively in their visual structure. Complementary image-complexity analyses highlighted the role of visual richness as a contextual factor in affective reporting. Together, these findings support WER as a valid and intuitive alternative to traditional schematic affective rating tools, particularly for communicating valence while maintaining comparable performance for arousal, and suggest that ecologically grounded visual metaphors may address some limitations of existing instruments.
With the increasing popularity and feasibility of implementing Ecological Momentary Assessment (EMA), research on affective dynamics has expanded considerably (1–3). In parallel, advances in artificial intelligence (AI) have enabled scalable extraction of rich multimodal features (e.g., facial, vocal, and linguistic) from video data, which may serve as adjunct complementary behavioural indicators to subjective ratings of affective states typically captured in EMAs. Emerging research has demonstrated the promising utility of these features in predicting the diagnostic status of mental health problems and the presence of transdiagnostic symptomology (4–6). However, how these behavioural indicators relate to momentary affect in naturalistic settings, and whether they explain unique variance in mental health symptoms beyond self-reported mood, remains unexplored. Given the recent rise in perinatal depression and anxiety (7), exploring these questions in the perinatal period will inform the potential clinical value of implementing naturalistic video screeners of postpartum mental health symptoms in perinatal care provision. The present study therefore has two primary aims: 1) to examine the extent to which multimodal behavioural indicators map onto momentary affect ratings across a one-week EMA protocol (3 pings daily) in birthing parents who are 1-6 months postpartum; 2) to test whether behavioural indicators explain variance in mental health symptoms above and beyond self-reported momentary mood, thereby evaluating their incremental validity as markers of emotional functioning.
The Darbest dataset is a Universal Dependencies (UD) treebank dataset for Standard Sorani Kurdish written in the Perso-Arabic script. It contains 69,000 annotated sentences and 1,205,855 tokens collected from nine textual domains. The corpus was collected from seven Kurdish online news websites and supplemented with texts from published books. Before preprocessing, the collected corpus contained 1,250,275 words from 5,627 web pages together with book-based texts and was preprocessed using a Python-based pipeline involving text cleaning, Unicode and punctuation normalization, sentence segmentation, and tokenization. The dataset was developed to provide a large-scale syntactically and morphologically annotated resource for Standard Sorani Kurdish. A separate 100-sentence gold-standard set was manually annotated according to the Universal Dependencies v2 guidelines. Sorani Kurdish linguistic experts supported the selection of sentences representing diverse and linguistically complex structures and reviewed LLM-generated annotations for errors. The gold-standard set was used to construct few-shot prompts for annotating the remaining corpus. The resulting annotations were represented in the standard CoNLL-U format and validated using the official Universal Dependencies validation tool, followed by manual correction and quality review. The released treebank is divided into training, development, and test sets and includes lemmas, Universal Part-of-Speech (UPOS) tags, morphological features, syntactic heads, and dependency relations. The dataset can be used to train, evaluate, and benchmark NLP models for part-of-speech tagging, lemmatization, morphological analysis, dependency parsing, and related computational linguistics tasks. The accompanying repository also contains the separate 100-sentence gold-standard set, plain-text corpus splits, README documentation, and a dataset statistics spreadsheet.
The article is devoted to the study of the functioning of youth slang in modern Polish and to clarifying the relationship between the linguistic norm and variation within this dynamic lexical subsystem. Youth slang is considered an important sociolinguistic phenomenon that reflects the linguistic creativity of the younger generation, its values, communicative needs, and aspirations for linguistic identity. Particular attention is paid to the sources of the formation of slang units, among which foreign borrowings, semantic transformations, word-formation models, as well as the influence of the digital environment and online communication play a significant role. The article analyzes the main tendencies in the development of youth slang in the Polish-speaking environment and outlines its stylistic, functional, and pragmatic characteristics. It has been determined that youth slang is formed under the influence of sociocultural factors, mass culture, media, and interlingual contacts. It has been established that slang performs not only expressive and identificational functions but also contributes to the formation of group solidarity, informal communication, and distancing from the official linguistic norm. At the same time, it is emphasized that the boundary between normative vocabulary and slang units is variable: certain elements of youth speech gradually integrate into broader language use, spread beyond the youth environment, and may acquire the status of generally accepted lexical items. The methodological basis of the study includes descriptive, comparative, and contextual methods of linguistic analysis. As a result of the research, the key mechanisms of variation of slang units, their structural and semantic features, as well as their role in contemporary linguistic processes have been identified. It is concluded that youth slang is an important factor in language dynamics that reflects sociocultural changes and contributes to the continuous renewal of the lexical system of modern Polish.
PURPOSE: This study aimed to provide normative affective ratings for middle-aged to older adults with normal hearing using the Marcell database. The database includes 120 diverse environmental sounds representing various real-world acoustic events, such as those produced by animals, humans, musical instruments, tools, signals, and liquids. Secondary aims were to examine the relationship between valence and arousal, and whether factors such as age, gender, hearing-related lifestyle, hearing demands, and self-reported hearing difficulties influence these affective evaluations. STUDY SAMPLE: Ninety participants (mean age: 54 ± 9 years) with self-reported normal hearing. DATA COLLECTION AND ANALYSIS: The study was conducted online. Participants rated each sound token in the Marcell database in terms of valence (pleasantness) and arousal (intensity). They also completed a hearing-related lifestyle questionnaire assessing the frequency, importance, and difficulty of hearing in everyday situations. Normative data for valence and arousal were calculated as means and standard deviations across participants for each individual sound token. To explore the relationship between valence and arousal, correlation analyses were performed. Furthermore, mixed-effects linear regression was used to examine whether any of the measured variables could predict participants' affective responses to the sounds. RESULTS: Normative data for each individual sound token were obtained and reported. A moderately strong negative correlation was found between valence and arousal ratings, indicating that pleasant sounds were generally perceived as calming, while unpleasant sounds were seen as more arousing. Regression analyses revealed that age, gender, hearing-related lifestyle, hearing demands, and self-reported hearing difficulties did not significantly predict valence or arousal ratings in the current dataset of middle-aged to older adults. CONCLUSIONS: The normative data presented in this study offer a benchmark for the hearing research community when examining emotional responses across middle-aged to older adult populations.
The article is devoted to the theoretical substantiation of the essence of the grammatical aspect of foreign language speech as a key component of foreign language communicative competence. The relevance of the study is determined by the need to understand the structural content of the grammatical aspect of speech in the context of the requirements of modern educational standards. The paper presents a comparative analysis of the approaches of foreign and Russian researchers to understanding the grammatical aspect of speech. D. Larsen-Freeman's three-dimensional model, which includes form, meaning, and use of grammatical phenomena, is examined and illustrated with examples. The position of S. Thornbury, who defines grammar through morphology and syntax and emphasizes its meaning-making potential realized in representational and interpersonal functions, is analyzed. Attention is paid to M. Lewis's lexical approach, which assigns a secondary role to grammar. The article presents the views of Russian methodologists (N.D. Galskova, N.I. Gez, E.N. Solovova), who consider grammar as a fundamental component of speech activity that ensures practical language proficiency for solving communicative tasks. Based on the analysis conducted, the author formulates a definition of the grammatical aspect of foreign language speech as a complex of automated actions for selecting, combining, and using grammatical structures in accordance with communicative intention and language norms. It is concluded that an insufficient level of mastery of the grammatical aspect of speech leads to difficulties in the formation of foreign language communicative competence as a whole.
Abstract Cognitive reappraisal is an emotion regulation strategy that involves changing one’s interpretation of a situation to change its emotional impact. Cognitive reappraisal studies typically assess changes in valence ratings following reappraisal as the primary dependent variable, but scholars have recently documented that valence ratings can take two forms: affective valence (i.e., the positivity/negativity of one’s hedonic internal feelings) and semantic valence (i.e., the positivity/negativity of one’s cognitive evaluation of the stimulus). Because typical rating instructions do not distinguish between these forms of valence, it remains unknown how strongly cognitive reappraisal shifts either one. We address this gap through two preregistered studies. In Study 1, N = 155 participants completed a classic cognitive reappraisal task in which they either responded naturally to or cognitively reappraised negative images. Critically, we manipulated whether participants rated either affective valence, semantic valence, or “default” valence (conventional instructions). Reappraisal significantly improved valence ratings in all conditions, but reappraisal most strongly influenced semantic valence and least strongly affective valence. Interestingly, default valence behaved most similarly to semantic valence, suggesting that existing reappraisal studies largely document changes in semantic evaluations, more than affective experiences. Study 2 ( N = 105) replicated these effects within-subjects. These results extend research on affective and semantic valence by demonstrating that they vary in their responsiveness to regulation. Additionally, they suggest that existing research likely overestimates the impact of reappraisal on felt emotional experiences, prompting new perspectives on past findings and more focused methods for assessing the hedonic impact of emotion regulation.
Abstract This chapter examines Ntozake Shange’s for colored girls as a revolutionary choreopoem that emerges from Black feminist thought, the Black Arts Movement, and United States Black Language to articulate Black women’s interior lives through embodied language and performance. It situates the work within its historical, political, and theatrical contexts and argues that Shange’s written theatricality renders African American Women’s Language visible on the page through phonetic spelling, syntax, punctuation, rhythm, and structure as practices of cultural memory, resistance, and self-definition. The chapter analyzes form, symbolism, and aesthetic strategies—including the rainbow, the slash, eye dialect, musicality, and lowercase typography—to demonstrate how language functions as choreography that directs reading, hearing, and feeling while rejecting standardized English norms and the white gaze. Through close readings of key phases such as “no more love poems #4,” “somebody almost walked off wid alla my stuff,” and “layin on of hands,” it demonstrates how Black lexical items, discourse practices, and sonic rituals enact rhetorical healing, rememory, and collective restoration. Finally, the chapter argues that Shange’s choreopoem functions as a performative Black feminist theory of language that transforms personal and communal trauma into embodied affirmation, spiritual renewal, and an enduring declaration that Black women’s voices, bodies, and lives are already and fully enough.
This study addresses the destabilization of moral frameworks among youth in contemporary Islamic societies, a phenomenon accelerated by rapid digitization. We investigate how the structural and stylistic mechanisms of traditional oral poetry can serve as pedagogical instruments to build ethical resilience among Gen Z and millennial cohorts in modern Aceh. Methodologically, the inquiry adopts a qualitative descriptive design grounded in pedagogical stylistics to conduct a close textual analysis of twenty purposively sampled advice quatrains from the Bunga Rampai Maluku Kie Raha anthology. Our analytical framework evaluates these texts across specific operational metrics: phonology, syntax, and conceptual metaphor. The empirical results reveal that the rigid structural blueprint of the quatrain, particularly the architectural division between the foreshadowing lines (sampiran) and the semantic core (isi), functions as an efficient cognitive delivery system that optimizes mnemonic retention. Additionally, the deliberate deployment of high-impact lexical markers combined with abrupt interrogative rhetorical shifts elevates generic ethical advice into absolute behavioral boundaries. These boundaries directly reinforce Acehnese socio-legal norms. Consequently, curriculum designers and language educators in Aceh can leverage trans-regional oral literatures to diversify character education frameworks, offering a vital alternative as localized oral traditions fade from daily usage. Ultimately, this research demonstrates that UNESCO-recognized traditional poetic structures operate not merely as historical artifacts, but as dynamic socio-pedagogical interventions capable of maintaining ethical stability during acute cultural transitions.
In March 2021, the EU Parliament adopted Resolution 2021/2557, a legally binding measure that mandates all 27 member states to recognize the right to gender self-identification and to implement juridical norms aligned with this principle. Among its most transformative provisions, the Resolution calls for eliminating the male-female binary in favor of a more expansive framework that currently recognizes at least twenty-one gender identities - a number expected to grow. It also urges the revision of national languages to dismantle patriarchal structures and ensure that legal and institutional language reflects principles of gender plurality and inclusivity. Widely seen as a landmark victory for trans-feminist individuals and advocacy groups, this measure has sparked both support and controversy. The research examines whether such linguistic reforms foster inclusion or provoke democratic tensions in Italy, where gendered language is deeply rooted in historical, grammatical, and cultural traditions. It further investigates how trans-feminist advocacy - supported ideologically and financially by EU bodies (Commission, Parliament, and Council) - has gained significant influence, particularly as left-wing progressive political forces currently hold the majority within these institutions. These actors play a central role in shaping the narrative and enforcement of gender policies across EU member states. Employing a qualitative case study methodology, the analysis draws on a diverse range of materials, including press articles, televised debates, public messaging, lexical usage, multimedia content, and ideologically charged propaganda to assess the impact of EU gender policy on Italy’s linguistic landscape. Findings suggest that while these interventions promote visibility and recognition for gender-diverse individuals, they also raise concerns about linguistic autonomy, democratic principles, and the broader cultural consequences of ideologically driven legal mandates.
The Orthographic Junctions of English A Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and ConsonantVowel Information Asymmetry (p. 1) Boicho Dimitrov Temelakiev Saxon Ventura Research Ltd 28th of May, 2026 CC BY Abstract This paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling (p. 1). While the distinct roles of consonants and vowels in language processing are well-recognized psycholinguistically (p. 5), and the dual-stratum organization of English morphophonology is established theoretically (p. 6), this study provides an original, datadriven quantification of these properties directly at the orthographic level. The morpheme boundary—the orthographic junction between a stem and an affix—is treated as the primary object of measurement, reading the distribution of boundary characters across the corpus (p. 1). Three results are established: 1. 2. 3. The junction carries a stable, structured filter: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a distributional sparsity that scales inversely with an affix’s productivity (pp. 1-2). Affix relationships are structural: The relationship between any two affixes is quantified by the correlation of their boundary distributions, measuring their shared stem population (\(r \approx 0.99\) for etymological doublets down to \(r = 0.57\) for productivity-asymmetric near-twins) (pp. 1, 4). A massive information asymmetry partitions the lexicon: The written word decomposes into a invariant consonant skeleton carrying lexical identity (53.0% unique recoverability) and a mobile vowel tissue carrying grammatical form (1.8% unique recoverability) (pp. 1, 5). Multiple independent measures—consonant recoverability, derivational class-marking, and freestem fraction—converge on a single partition separating a transparent Germanic core from a bound Latinate superstructure (pp. 1, 6). The method, its corrections, and its limits are reported in full (p. 1). 1. Introduction and Method Traditional models of English morphophonology have long recognized that the lexicon is organized into distinct, historical strata—principally a native Germanic core and a bound Latinate superstructure (pp. 1, 6). Classic frameworks in generative phonology and lexical morphology, such as those pioneered by Chomsky and Halle (1968) and expanded by Kiparsky (1982), demonstrate that affixes of differing origins impose strict constraints on the phonetic and structural traits of the stems they recruit. Concurrently, cognitive and psycholinguistic research 1 has established a foundational "consonant-vowel functional asymmetry," demonstrating that human language processing systematically relies on consonants to preserve lexical and lexical-root identity, while vowels are dynamically manipulated to signal grammatical operations (Nespor et al., 2003; Bonatti et al., 2005). While these qualitative boundaries and cognitive patterns are deeply documented, this paper presents a mean-free, data-driven methodology to extract, quantify, and map these structural phenomena directly from corporate-scale English orthography without relying on heavy linguistic machinery. We introduce The Orthographic Junctions of English (OJE), an empirical approach that frames the morpheme boundary—the exact character interface between a stem and an affix—as an informational filter whose statistical properties reveal the historical, structural, and cognitive divisions of the vocabulary. All results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution (p. 1). The corpus utilized is the opensource dwyl/english-words repository (words_alpha.txt), comprising 455,246 alphabetic entries with a total of 4,254,354 letter occurrences and a mean word length of 9.345 letters (p. 1). Each letter is assigned its ordinal value (\(A=1\) through \(Z=26\)) (p. 1). Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter, so that each suffix position returns its exact ordinal value as a safety check (p. 1); the first stem letter preceding the suffix—the linker—is then read as a full A–Z frequency distribution rather than as a mean (p. 1). The governing methodological constraint is that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, and that every claim is backed by a precise empirical count (p. 1). By avoiding any dependency on complex machine-learning libraries or external lexical databases, the framework ensures that every architectural pattern discovered can be verified using standard data operations. 2. The Orthographic Junction and Its Backbone 2 The initial phase of this investigation examined unconditioned letter bigrams across the corpus, which yielded no meaningful morphological signal. Structural regularities appeared only when character distributions were explicitly conditioned on a single morpheme boundary—the character interface linking a stem to an affix. This structural conditioning serves as the foundation of the OJE framework. A preliminary tabulation of adjacent letter pairs across the corpus—measuring which letters follow which, without regard to structural position within the word—yielded baseline frequency patterns. These patterns are entirely reducible to general English orthographic constraints and carry no isolable morphological content. The structural signal emerged only when a specific suffix was fixed and the preceding characters were read as a discrete population. Conditioning on the junction, rather than measuring adjacency as a flat sequence, renders the underlying boundary constraints visible. The set of characters that legally occupy the stem side of a morpheme boundary proves narrow, highly structured, and remarkably stable across suffixes. Reading the linker distribution across the mapped suffix inventory reveals a highly stratified, three-tier architectural filter: A structural backbone of six letters—T, L, N, R, S, and I—is admitted at high frequency by nearly every suffix in the English lexicon. Within this backbone, T serves as the single most frequent linker across the inventory and recurs as the dominant boundary letter across independent suffixes. Conversely, an absolute floor of two letters—J and Q—is admitted by no suffix at a measurable frequency. This absolute prohibition is confirmed corpus-wide and is statistically consistent with their status as the two rarest letters in English orthography overall (with J accounting for 0.18% and Q for 0.19% of all letter occurrences). Between the backbone and the floor lies a selective middle whose specific character composition varies dynamically by affix, providing the distinct orthographic footprint wherein an individual suffix’s identity resides...
Emotional responses to sounds are often described along two dimensions: arousal (exciting–calming) and valence (pleasant–unpleasant). Whereas sound level is a robust cue for arousal, the acoustic cues supporting valence are less well defined, particularly across listeners with different hearing abilities. Valence ratings were obtained from 236 adults (with and without hearing loss) for 75 sounds from the International Affective Digitized Sounds corpus (IADS-2). These data were pooled from several studies. For each sound, acoustic characteristics reflecting spectral shape and its time-varying changes were extracted. For each listener, age, gender, pure-tone average hearing threshold (0.5–2 kHz), and anxiety/depression symptom scores were collected. Valence ratings for previously unseen listeners were predicted using an interpretable machine-learning model (gradient-boosted trees with a mixed-effects structure). Out-of-fold performance was moderate (R² = 0.458; Spearman ρ = 0.679). The strongest predictors included variability in spectral “peakiness” (spectral crest), mid-frequency spectral contrast near 1–2 kHz (mean and variability), and loudness variability. Hearing loss also contributed, suggesting audibility and/or altered sound representation affects pleasantness judgments. Overall, listeners appear to rely on time-varying spectral structure, particularly in mid frequencies, when forming pleasant–unpleasant judgments, providing candidate acoustic cues for affective sound processing in listeners with and without hearing loss.
Nick Joaquin (1917-2004), one of the most celebrated Filipino writers in English, masterfully appropriates the language of the colonizer to create a distinctive style of English known as Joaquinesque. His works challenge Western linguistic norms by integrating indigenous and Hispanic influences into English, thereby resisting the erasure of Filipino identity. Through his creative manipulation of English, Joaquin engages in what Joy Harjo describes as "reinventing the enemy’s language" (Harjo et al., 1998, p. 22), transforming it into a medium of cultural assertion and artistic defiance.
The thesis investigates phraseological transformation, contamination and occasional paradigm expansion as sources of semantic ambivalence. The analysis is based on English, Russian and Uzbek examples in which a stable expression is reinterpreted through context. The study shows that phraseological ambivalence appears when an idiom is understood both as a fixed expression and as a free combination of words. Contamination creates a hybrid form in which two phraseological prototypes are simultaneously recognizable. Occasional paradigm expansion occurs when a text temporarily includes a foreign unit into an existing semantic series on the basis of phonetic, graphic or morphological similarity. These mechanisms reveal the semantic productivity of the tension between linguistic norm and contextual innovation.
This research aims to study the stylistic choice of words and trace their semantic and argumentative movement by shedding light on the precision of the stylistic choice of Quranic words, which carry argumentative potential and persuasive semantic dimensions through their context. The recipient has no choice but to submit and acquiesce. The study examines these words in terms of their lexical and semantic content, not as a departure from the linguistic norm. The study examines examples from the Holy Quran, where the word is evaluative, lively, and charged with semantic dimensions and persuasive argumentative potential. These words compel the recipient to submit and acquiesce to what is presented to them. They possess semantic features that connect them to the context in which they appear, directing them in an influential argumentative direction.
The practice of web form submission has emerged as a prime conduit for attackers, enabling them to infiltrate modern web applications and illegally harvest sensitive user data. Traditional defense mechanisms, such as static security reviews and server-side validation, are proving insufficiently agile for real-time detection of client-side vulnerabilities. This inadequacy arises directly from the rapid evolution of modern interfaces, which involves spontaneous DOM changes, dynamic element generation, and semantic interpretation that varies based on context and culture. This article presents an innovative browser extension framework that leverages a heuristic-based, multi-dimensional analytical engine combined with deep DOM inspection to identify insecure form submissions the moment they occur. The proposed methodology introduces five fundamental innovations: a contextual risk scoring system that models the complex interdependencies among form fields; an adaptive weighting scheme for risk patterns, accommodating diverse cultural and linguistic norms; a predictive vulnerability estimator that anticipates future threats; intelligent DOM mutation filtering designed to significantly optimize runtime performance; and cross linguistic semantic recognition to determine the true purpose of fields globally. Based on theoretical projections, this combined approach promises to enhance vulnerability detection accuracy while simultaneously reducing computational demands by approximately. Critically, all security analysis is executed exclusively on the user's local machine, guaranteeing privacy by ensuring no sensitive data is transmitted externally. A proof of concept application confirms the framework's practical feasibility and high efficacy for client side security assessment and catalyzing the development of flexible, scalable, and privacy respecting browser-based protections.
The rapid proliferation of social media platforms over the past two decades has precipitated a fundamental restructuring of human communicative behaviour, with measurable consequences for lexical innovation, syntactic convention, orthographic norms, and the distribution of linguistic authority. This article investigates the transformative impact of hyperconnected digital ecosystems — principally Facebook, X (formerly Twitter), Instagram, and TikTok — on the trajectory of language evolution in the twenty-first century. Drawing on theoretical frameworks from sociolinguistics, cognitive linguistics, and digital communication studies, the study analyses four interconnected dimensions of linguistic change: the acceleration of neologism formation and the entrenchment of digital vernacular; the structural and pragmatic shifts in written syntax facilitated by platform-specific constraints; the democratization of linguistic authority through network-driven, bottom-up innovation; and the long-term evolutionary implications of digital code-switching as an emerging literacy competency. The analysis integrates empirical evidence from corpus-based studies of social media language, sociolinguistic network theory, and recent scholarship on the interaction between algorithmic mediation and human linguistic choice. The findings demonstrate that social media does not merely reflect linguistic change but actively drives it, compressing evolutionary timescales, redistributing prestige, and generating new hybrid registers that challenge the prescriptive boundaries between formal and informal communication. The study concludes with a discussion of future research directions, including the emerging role of generative artificial intelligence in shaping human linguistic production within social media environments.
The Orthographic Junctions of EnglishA Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and Consonant–Vowel Information AsymmetryBoicho Dimitrov Temelakiev · Saxon Ventura Research Ltd · CC BY · Version 3, June 2026Changes in this version (v3)This version corrects and strengthens the deposited record. (i) The corpus provenance is corrected: the analysed file is the dwyl `words.txt` list after cleaning, not `words_alpha.txt` as stated in v1–v2; the exact 455,246-entry corpus is deposited so the source is unambiguous. (ii) The descriptive "warm/cold" parameter (v1 §2/v2 §3) is withdrawn as a measurement-scale error: it performs arithmetic on the alphabet's ordinal position, which is a nominal code, and we show it carries no order-invariant signal. (iii) The associated "conjugation temperature" claim is restated without any appeal to alphabet position. (iv) All doublet correlations were recomputed on the deposited corpus under explicit nesting controls; the values are reported as measured, and one figure (−ANT/−ANCE) is corrected. (v) A permutation-null robustness test is added as a standing methodological control. No surviving result depends on the alphabet's ordering.AbstractThis paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling. The morpheme boundary — the orthographic junction between a stem and an affix — is treated as the primary object of measurement, and the distribution of the letters that may occupy the stem side of a junction is read across the corpus. Three results are established. First, the junction carries a stable, structured filter: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a sparsity that scales inversely with a suffix's productivity. Second, the relationship between any two affixes is quantified by the correlation of their boundary distributions, which measures the degree to which they share a stem population; this correlation ranges from ~0.97 for etymological doublets to 0.58 for productivity-asymmetric near-twins. Third, the written word decomposes into a consonant skeleton carrying lexical identity (53.0% of the vocabulary uniquely recoverable from consonants alone) and a vowel tissue carrying grammatical form (1.8% recoverable from vowels alone), a ~29-fold information asymmetry. Multiple independent measures — consonant recoverability, derivational class-marking, and free-stem fraction — partition the lexicon at a single boundary separating a transparent Germanic core from a bound Latinate superstructure. Every quantitative claim is tested for invariance under permutation of the alphabet, so that no result depends on the arbitrary ordering of the letters; the method, its corrections, and its limits are reported in full.1. Corpus and MethodAll results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution.The corpus derives from the public dwyl/english-words list (`words.txt`), from which non-alphabetic entries were removed and the remainder case-folded and deduplicated, yielding 455,246 unique alphabetic entries with 4,254,354 letter occurrences and a mean word length of 9.345 letters. Because the upstream list drifts over time, the exact 455,246-entry file analysed here is deposited with this record and is the corpus of record; a reader downloading the live upstream list today will not recover the same entry count.Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter; the first stem letter preceding the suffix — the linker — is read as a full per-letter (A–Z) frequency distribution rather than as a mean. Each letter may be referred to by its position in the alphabet purely as a label; no quantity in this paper depends on treating that position as a number (see §9). The governing methodological constraints are that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, that every claim is backed by a count, and — new in this version — that every claim is invariant under relabelling of the alphabet.Nesting among suffixes (for example −MENT within −ENT, or −ATION within −TION within −ION) is controlled by excluding longer relatives before counting. No statistical software, machine-learning library, or external lexical database is used at any stage of the core analysis; every figure can be reconstructed from the deposited corpus with a spreadsheet alone.2. The Junction and Its BackboneThe investigation began as a survey of unconditioned letter bigrams, which returned no morphological signal; structure appeared only when letter distributions were conditioned on a single morpheme boundary, and that conditioning is the method's foundation.Conditioning character statistics on a morpheme boundary places this work within the successor-variety tradition of boundary detection introduced by Harris (1955) and first implemented computationally by Hafer and Weiss (1974). That tradition uses transitional letter predictability to segment words into morphemes; the present method inverts the emphasis, holding a known boundary fixed and characterizing the distribution of stem-side letters it admits — a characterization of the junction rather than a segmentation of the word.A preliminary tabulation of adjacent letter pairs across the corpus — which letters follow which, without regard to position within the word — yielded frequency patterns reducible to general orthographic regularities and carrying no isolable morphological content. The signal emerged only when a specific suffix was fixed and the letters preceding it were read as a population. The set of letters that may legally occupy the stem side of a morpheme boundary then proves narrow, structured, and stable across suffixes.Reading the linker distribution across the mapped suffix inventory reveals a three-tier structure. A backbone of six letters — T, L, N, R, S, and I — is admitted at high frequency by nearly every suffix; T is the single most frequent linker across the inventory and recurs as the dominant boundary letter in suffix after suffix. An absolute floor of two letters — J and Q — is admitted by no suffix at measurable frequency, a prohibition confirmed corpus-wide and consistent with their status as the two rarest letters overall (J at 0.18%, Q at 0.19% of all letter occurrences). Between backbone and floor lies a selective middle whose composition varies by suffix and in which each suffix's identity resides. (The backbone, floor, and selective middle are stated as sets of letters; nothing in the three-tier description depends on the order of the alphabet, and all of it is invariant under the permutation test of §9.)The degree of selectivity is itself a measurement. The count of forbidden letters at a junction — its sparsity — scales inversely with the suffix's productivity: derivational suffixes that attach choosily to a constrained stem class forbid many letters, whereas inflectional or highly productive suffixes forbid few. Sparsity is therefore not noise but signal: the pattern of exclusion characterizes the suffix as informatively as the pattern of admission.3. The Suffix AtlasSuffixes are described by the distribution of letters that survive at their boundary — read in full, never reduced to a mean. The corrective lesson is explicit and was learned in this program: −NESS was first misjudged from a summary statistic and only described correctly once its full distribution was read (E 27%, D 15%, S 15%, I 13%). The general rule that follows is that a junction must be read as a distribution, not a single number, and in particular not as a mean of letter positions — a point developed formally in §9.Two representative distributions illustrate the contrast between a concentrated and a broad boundary: −ABLE (T-led and broad) and −IBLE (a sparse, frozen Latinate boundary). Both are reported as per-letter frequencies; the comparison between them is made by correlation (§4), which is invariant under relabelling of the letters.Table 1. Linker distribution of −ABLE (n = 4,694). Letters ≥3% shown.Linker Count PercentT 846 18.0%R 513 10.9%N 420 9.0%E 392 8.4%S 330 7.0%I 292 6.2%D 286 6.1%L 286 6.1% Table 2. Linker distribution of −IBLE (n = 738). Five letters carry ~90% of the population.Linker Count PercentS 248 33.7%T 208 28.3%C 86 11.7%D 63 8.6%G 59 8.0%N 20 2.7% Read by distribution, the suffix inventory resolves into a small set of boundary shapes, each fixed by the population of stems the suffix recruits rather than by its function, origin, or spelling.4. The Doublet Principle, QuantifiedTwo suffixes that draw on the same stem population share the same boundary distribution, and the correlation between their distributions measures the extent of that shared population directly. Because correlation is computed component-by-component over the same set of letters in both vectors, it is invariant under any relabelling of the alphabet — it is one of the order-independent quantities the program now requires (§9).All correlations below were recomputed for this version on the deposited corpus under explicit nesting controls (Pearson r over the 26-component per-letter frequency vectors). Across three classes of suffix pairs the boundary correlation forms an interpretable gradient.Etymological doublets — the same Latin stem class in two guises — correlate highly: −ENT/−ENCE at r = 0.97 (nesting-controlled, excluding −MENT from the −ENT population) and −ANT/−ANCE at r = 0.94. The −ANT/−ANCE value is corrected here: earlier versions reported r ≈ 0.99, which does not reproduc
This study examines the role of institutional frameworks and public organizational arrangements in supporting social innovation within rural tourism in the Marrakech–Safi region, using the theory of change as the central analytical framework. In a territorial context marked by persistent socio-economic vulnerabilities, strong heterogeneity across rural areas, and increasing dependence on tourism activities, social innovation represents a key lever for promoting inclusive and sustainable territorial development. Public action plays a structuring role by defining strategic orientations, norms, and organizational instruments that shape the emergence and diffusion of innovative initiatives at the local level. The study adopts a qualitative approach based on the lexicometric analysis of institutional discourses related to public action in support of social innovation in rural tourism. Results drawn from lexical frequency distribution, correspondence factor analysis, similarity analysis, and word cloud visualization reveal an institutional discourse strongly structured around normative and instrumental registers. Institutional frameworks emerge as central inputs in the change process, guiding territorial development objectives toward social inclusion, sustainability, and territorial embeddedness. Public organizational arrangements act as intermediate mechanisms, translating these orientations into concrete action capacities through governance, financing, and support for local initiatives. However, the findings also highlight a limited explicit articulation of learning, adjustment, and evaluation mechanisms—elements that are essential in the theory of change—suggesting that the prevailing conception of change remains largely linear and top-down.
The current study examines the Pakistani English and Indian English newspapers’ discursive construction of the 2025 flood crisis, grounding the analysis within the framework of World Englishes and Critical Discourse Analysis (CDA). Drawing on Fairclough’s three-dimensional model (1989), the research investigates textual features, discursive practices, and socio-cultural contexts to reveal how language mediates disaster narratives in two neighbouring South Asian countries. A sample of thirty news reports from six leading English-language newspapers in Pakistan and India were taken, employing a qualitative, comparative analysis. The findings demonstrate clear divergences in disaster representation. Pakistani English newspapers predominantly frame floods as humanitarian emergencies, employing emotive lexicalization, passive constructions, and crisis-oriented narratives that foreground vulnerability, climate risk, and governance limitations. Indian English newspapers, by contrast, adopt a more procedural and bureaucratic discourse, emphasizing administrative control, technical expertise, and institutional accountability through active agency and policy-focused framing. Despite these differences, both varieties rely heavily on elite institutional sources, marginalizing the voices of affected communities. From a World Englishes perspective, the study shows how Pakistani and Indian English function as localized outer-circle varieties that balance global journalistic norms with national socio-political ideologies. The article contributes to disaster discourse scholarship by highlighting how English, as a shared transnational medium, simultaneously enables cross-border circulation of information and reproduces distinct national identities, power relations, and models of governance in climate crisis reporting.
This paper presents the steps taken to integrate data from the UD_Latin-PROIEL treebank into the LiLa Knowledge Base of interoperable linguistic resources for Latin.It describes how the lexical, morphological, syntactic, and citation information from the source was modeled using the Linked Open Data principles as adopted by the LiLa Knowledge Base.The process of linking tokens to the LiLa collection of Latin lemmas is detailed, addressing challenges such as ambiguities, new lemmas, and errors encountered in the source.The outcome is a syntactically annotated textual resource that is interoperable with the (meta)data of other Latin linguistic resources linked within the LiLa Knowledge Base.This integration enables new ways of analyzing linguistic information and using the content as a starting point to explore connections with other interlinked resources.A use case demonstrates this interoperability.
In the Azerbaijani language, there exists a group of words that do not have broad possibilities of usage in the lexicon. Such words either reflect a certain socio-cultural context or are used in the speech of specific specialists, including artists and professionals. These words belong to the general lexical layer of the language and reflect its rich lexical diversity. Words of this type, which do not acquire an active and general-usage character in the Azerbaijani lexicon, can be considered a group of relatively limited-use vocabulary, as they are still part of the overall lexical system. This group includes archaisms, dialectisms, and terminological vocabulary. Archaisms, within the synchronic layer of the language, have limited usage and reflect historical, social, everyday, and other ethnographic concepts, creating a national and ethnic color. This constitutes the general norm of archaisms. Dialectisms express both historical-diachronic and contemporary synchronic concepts. They also reflect local ethnic features of mentality. A certain part of dialectisms may enter the general vocabulary, as a result of which instability of norm is observed in their local and regional forms, whereas in the general vocabulary stable normative functioning is established. Terms related to specific fields of activity function in accordance with the norm of obligatory speech usage. Those that enter the general lexical fund may acquire a stable norm in both spoken and written language.
The Core Idea Classical combinatorics counts the number of ways to partition a set of n elements into non-empty blocks. That count is the Bell number B(n), and it depends only on how many elements there are, never on how they relate to each other. This manuscript introduces a strict refinement. Given a connected graph G on n labeled vertices, define a connected set-partition as a partition of V(G) in which every block induces a connected subgraph. The geometric partition function is p₍𝔾₎(G) = |{ {V₁, …, Vₖ}: each G[Vᵢ] is connected }| This count depends on G, not just n. Two graphs on the same number of vertices can have wildly different geometric partition counts. The entire theory follows from taking that dependence seriously. Main Theorems Theorem 1 (Subgraph Lattice Theorem). If G and H share the same vertex set and E(G) ⊆ E(H), then p₍𝔾₎(G) ≤ p₍𝔾₎(H). Adding edges can only increase the connected partition count, never decrease it. Theorem 2 (Tree Evaluation). For any tree T on n vertices, p₍𝔾₎(T) = 2ⁿ⁻¹. Every tree on n vertices, regardless of shape, has exactly the same geometric partition count. This is the universal floor for connected graphs. Theorem 3 (Complete-Graph Evaluation). For Kₙ, p₍𝔾₎(Kₙ) = B(n), the n-th Bell number. When every possible edge is present, every set-partition is automatically connected, so the geometric count recovers the classical count. This is the universal ceiling. Theorem 4 (Cycle Closed Form). For the cycle graph Cₙ with n ≥ 3, p₍𝔾₎(Cₙ) = 2ⁿ − n. This matches OEIS sequence A000325. Corollary (Sandwich Hierarchy). For every connected graph G on n vertices with n ≥ 4: 2ⁿ⁻¹ = p₍𝔾₎(Tₙ) < p₍𝔾₎(Cₙ) = 2ⁿ − n < p₍𝔾₎(Kₙ) = B(n) Both inequalities are strict. Trees sit at the floor, complete graphs at the ceiling, and the cycle sits strictly between them. What the Numbers Look Like For cycles versus trees versus complete graphs: n = 5: Tree = 16, Cycle = 27, Complete = 52 n = 6: Tree = 32, Cycle = 58, Complete = 203 n = 7: Tree = 64, Cycle = 121, Complete = 877 n = 8: Tree = 128, Cycle = 248, Complete = 4,140 n = 10: Tree = 512, Cycle = 1,014, Complete = 115,975 The gap between floor and ceiling grows super-exponentially. At n = 10, the complete graph has 227× more connected partitions than a tree on the same vertices. At fixed n, different topologies produce sharply different counts. At n = 6 for example: P₆ = 32 < C₆ = 58 < Ladder = 74 < Fan = 89 < K₂,₄ = 96 < Prism = 114 < W₆ = 118 < K₆ = 203 The ordering tracks graph density with Spearman ρ = +0.955 (p < 10⁻⁶) across all tested sizes n = 4 through n = 10. The Connection Curvature Define the partition connection on cycles as Γₙ = p₍𝔾₎(Cₙ₊₁) / p₍𝔾₎(Cₙ). The curvature κₙ = Γₙ − 2 measures deviation from the flat tree-growth rate: κₙ = (n − 1) / (2ⁿ − n) This is strictly positive and monotonically decaying to zero: n = 3: κ = 0.4000 n = 5: κ = 0.1481 n = 7: κ = 0.0496 n = 10: κ = 0.0089 The cycle approaches tree-like growth exponentially fast. The successive ratio κₙ / κₙ₋₁ converges toward ½ from above. Empirical Validation Program The deposit includes a preregistered validation suite with predeclared falsifiers, frozen metrics, and external ground truth. All outcomes, including failures, are retained as first-class results. Test 1 (OEIS Path Recursion). Brute-force enumeration confirms p₍𝔾₎(Pₙ) = 2ⁿ⁻¹ for n = 1 through 20. The sequence matches OEIS A000079 (powers of 2) exactly. Algebraic DP and direct bitmask enumeration agree at every tested value. Test 2b (Same-n Topology Comparison). At each n from 4 to 10, connected graphs on the same vertex count were compared. Spearman rank correlation between vertex connectivity and p₍𝔾₎ is ρ ≥ 0.83 at every tested n. Pooled edge-density correlation: ρ = 0.955 with p < 10⁻⁶. Negative control (max degree): ρ ≈ 0.35, not significant at any n. The lattice theorem's monotonicity prediction holds empirically, and counterexamples to vertex-connectivity ordering exist only between spanning-subgraph-incomparable pairs, exactly as the theory permits. Test 4 (Internet Autonomous Systems, 2010–2024). Using 180 monthly CAIDA BGP snapshots: AS count grew from 33,778 to 77,495 while edges grew from 93,998 to 501,473. The edge-succession to vertex-succession ratio Sₑ/Sᵥ increased with a linear slope of 3.68 per year, R² = 0.88, p = 2.23 × 10⁻⁷, Spearman ρ = 0.968. The pre-registered threshold was a slope above 0.05. The observed slope is 73× that threshold. Edge growth dominates vertex growth in a real infrastructure graph, consistent with the lattice theorem's prediction that denser graphs have exponentially more connected partitions. Test 5 (OpenAlex Citation Subgraphs). Among 10,000 co-citation subgraphs from 2020–2024: only 3 out of 10,000 were path graphs (0.03%). Among connected subgraphs, paths were 0.47%. For n ≥ 7, zero path graphs appeared among 527 connected subgraphs. Real citation networks avoid the tree floor almost entirely. Test 6/6b (Universal Dependencies Treebanks). The primary test fired its falsifier: English had 19.63% path-topology sentences, exceeding the 10% threshold. However, the pre-registered follow-up (Test 6b) conditioning on n ≥ 5 tokens found all six languages below 10%, with a maximum of 1.57% (Russian). At n ≥ 7 tokens, all languages drop below 0.15%. Short sentences are path-like; longer sentences are not. Test 7 (C. elegans Connectome). 279 neurons, 1,961 synapses. The biological network sits near the tree floor: 81% of 4-vertex connected induced subgraphs achieve p₍𝔾₎ = 2ⁿ⁻¹ exactly. KS tests showed no significant difference from Erdős–Rényi at matched density. Pre-registered hypothesis not supported. The sandwich bound is tight in sparse biological networks. Test 10 (Crystallographic Space Groups). Commutation graphs of all 230 space groups tested against a Ramsey proxy bound. The bound holds for 230/230 groups (100%). Saturation target (≥ 3 groups at n = 5 with zero triangles and zero independent triples) was not achieved. Verdict: partial. Tests 2, 8 (Falsifier Fired). Cycloalkane ring strain shows no monotonic observable matching p₍𝔾₎ (Pearson r = −0.11, p = 0.79). KEGG metabolic pathways achieved 14/18 hits for the Sₑ ≥ 2 × Sᵥ criterion versus a target of 15. Both failures are preserved in the record. Validation Ledger Supported: Tests 2b, 4, 5, 6b (4 of 10 completed tests) Falsifier fired: Tests 1 (by design), 2, 6, 8 (4 of 10) Not supported: Test 7 (1 of 10) Partial: Test 10 (1 of 10) Blocked/Deferred: Tests 9, 11 Negative and null results are first-class outcomes, not omissions. The program-level verdict is assessed across the full ledger, not cherry-picked from successes. Limitations Some tests depend on external APIs (OpenAlex, CAIDA, KEGG) whose upstream data evolves. Test 9 is blocked by endpoint observability constraints. This archive captures a dated snapshot; future reruns under changed data conditions should be treated as independent replications, not as invalidations. Citation Cite the Zenodo DOI for this archived package and the manuscript title. If reusing specific test scripts or results, cite the relevant script and JSON artifact in your methods section.
Abstract This study investigates the multifunctionality of the Turkish discourse connective ve ‘and’ in simultaneous interpreting from Turkish into Turkish Sign Language (TİD) in broadcast news. Using a corpus of interpreted news texts, it examines how this polyfunctional connective is transferred across modalities. Drawing on Halliday and Hasan’s cohesion model and the Penn Discourse Treebank typology, the study identifies discourse relations expressed by ve — such as elaboration, temporality, causality, and comparison — and analyzes how they are conveyed in TİD through explicit signs, implicit marking, or omission. Most relations are preserved, though often implicitly, reflecting the economy principle in signed languages. Manual signs used to express ve, including ve -1 and palm-up, vary systematically with relation type. This points to a cross-modal tendency for explicit realization under increased processing demands. As the data derive from hearing interpreters rather than native Deaf signers, the findings reflect interpreter-mediated strategies rather than monolingual Deaf discourse norms.
Anatomical education has historically relied on binary terminology, reinforcing the conflation of sex, gender, and gender identity through long-standing social and linguistic norms. To better include gender and sex diverse students and to respect the delineations between gender and sex, educators are incorporating diversity into their courses through educational primers; however, there is a lack of literature on the subject. Thus, this study explores learners' perceptions and attitudes about adding sex and gender diversity into anatomical education programs at a Canadian university. A total of 73 participants responded to a survey informed by grounded theory, and descriptive statistics are used for analysis. Nearly half of learners felt their overall educational experience was enhanced by the inclusion of sex and gender diversity content in their anatomy course and around half of learners felt more knowledgeable about inclusive language. Some learners were inspired to begin using inclusive language following their anatomy course. Overall, these findings suggest that integration of gender and sex diversity could be well-received and valued by anatomy students for humanistic skill development. Emerging themes requiring qualitative insights, including learners' definitions of "inclusive language" and the viability of peer-based learning, will be explored in future work.
Abstract Recognizing others’ emotions is central to social interaction. Traditional biological psychology infers emotional responding via laboratory measures, whereas contemporary computer vision algorithms claim to identify emotions unobtrusively from facial video. However, the validity of such algorithms for classifying spontaneous emotional responses occurring without explicit communicative intent remains debated. We compared established psychophysiological measures (EEG, facial EMG, EDA activity) with the open-source facial behavior toolkit OpenFace for classifying participants’ spontaneous responses during free viewing of happiness-inducing, disgust-inducing, and neutral pictures. Participants provided valence and arousal ratings and later selected the basic emotion that best matched their reaction which served as the classification criterion. Using within-participants single-trial support vector machine (SVM) classification, EEG achieved the highest accuracy (40%), followed by facial EMG (37%); OpenFace reached 36%. All methods except EDA exceeded chance performance (33.3%) and were lower compared to human raters (48%). Predictions declined slightly for across-participants SVMs, being at chance for OpenFace and EDA. The results indicate that in principle both, psychophysiological measures and video-derived facial action units, can capture diagnostically relevant aspects of emotional responding during picture viewing, but that their performance is limited when expressions are spontaneous and not produced for communicative purposes. Inter-individual variability in expressivity and physiological responding likely contributes to these limitations and should be considered when deploying automatic emotion recognition in research or applied settings.
Abstract Virtual reality (VR) is increasingly adopted across various fields, due to its ability to immerse people in virtual environments (VEs) and induce emotions. A key factor in this experience is the sense of presence, which is the feeling of being in the VE and the perceived realism of the experience. While prior research has demonstrated the importance of presence in driving emotional outcomes, gaps remain in understanding how the type of VEs and individual differences may influence this relationship. The present study addressed these gaps by comparing the strength of the relationship between presence and emotional outcomes across fear-inducing and relaxation-inducing VEs. The study also investigated whether this relationship was moderated by individual differences such as trait absorption and neuroticism. 125 participants were randomly assigned to one of two VEs. Participants completed baseline assessments, experienced the VE, then completed post-test assessments. Emotional outcomes were assessed through subjective emotional valence and arousal ratings. Results showed that the relationship between presence and emotional arousal was significantly stronger in the fear-inducing VE than in the relaxation-inducing VE. Furthermore, trait absorption and neuroticism significantly moderated the relationship between presence and emotional valence in only fear-inducing VEs. No moderation effects were found for the relationship between presence and emotional arousal. The study points that the relationship between presence and emotional outcomes is not consistent and may be influenced by the type of VE, trait absorption, and neuroticism. These findings provide theoretical and practical implications for designing effective VEs for various applications.
This study explores the sociolinguistic evolution of the English language within the contemporary global Muslim community, focusing on the emergence of "Islamic English." As globalization de-centers English from its native Western origins, non-Arab Muslim populations increasingly adopt it as a vital medium for religious expression and identity construction. Utilizing Critical Discourse Analysis (CDA) and a qualitative case-study approach, this research examines how speakers in Indonesia, Pakistan, Turkey, and Western Muslim diasporas navigate the inherent tensions between Anglo-centric linguistic norms and Islamic values. The analysis focuses on lexical borrowing, semantic transformations of Arabic roots, and pragmatic code-mixing within digital and academic discourses. The findings suggest that Islamic English functions not merely as a passive translation tool, but as an active, translingual instrument that validates localized religious identities. Ultimately, this study contributes to the fields of World Englishes and the sociolinguistics of religion by challenging traditional "Standard English" biases and documenting the legitimacy of religious linguistic variations.
Language change is continuous and vital for the sociocultural adaptation and status development of language varieties. Perhaps the most important linguistic feature of Nigerian English language change is semantic extension which refers to the process whereby terms already existing in English (standard English) acquire new meanings. This paper investigated the usage of semantic extension in Nigerian English (NIE), the social motivations responsible for such semantic innovations and the relevance of linguistic innovations for the description and teaching of English in Nigeria. Using descriptive qualitative research design, data for the research were collected using naturally occurring language forms while other information sources included print and social media, educational, political and religious institutions. It was found that semantic extension in Nigerian English is rule-governed, based on culture, multilingualism, technological innovations, socio-economic experiences and indigenous conceptual framework. The data further revealed that the semantic extensions are sociolinguistically recognized by educated users of Nigerian English and linguistic norm and not mere linguistic deviation from Standard English. It was observed that semantic extension promotes lexical reduction, language creativity and the institutionalization of Nigerian English as a variety of the World Englishes paradigm. It was recommended that Nigerian English new lexical items should be incorporated into English textbooks and dictionaries, Nigerian English lexical innovations should form the basis for the teaching of English in Nigeria.
This seminar paper, written in 1984, examines the divergent conceptions of 'Sprachkultur' (language culture) in the German Democratic Republic (GDR) and the Federal Republic of Germany against the background of their differing political systems and world-views. As a contemporary document of German division, the paper offers insights into the ideologically shaped linguistics of the Cold War. The study first analyses the fundamental differences between the Marxist-Leninist and the pluralist-liberal understandings of politics, scholarship, and culture. Taking as its starting point the concept of the Prague School of linguistics, which in the 1930s developed the first scholarly grounded programme of language culture, it traces the later instrumentalisation of this approach in the linguistics of the GDR. Whereas in the GDR 'Sprachkultur' served as a state educational programme for the formation of the 'socialist personality', in the Federal Republic a descriptive, pluralist approach emerged, one that recognises differing linguistic norms as of equal worth. The paper shows how linguistic concepts are shaped by their social framework conditions, and argues that, despite the shared socio-economic structures of modern industrial societies, the differing political systems gave rise to fundamentally distinct conceptions of language cultivation and language policy. As a historical document, the paper vividly attests to the communicative conflicts and ideological barriers that impeded scholarly exchange between East and West. This is a translation of the original paper published under doi 10.5281/zenodo.21925229. It was generated with the aid of Claude.Ai.
Front mid-vowel phonemes /e/ and /ɛ/ are contrastive in Gallego, while Spanish contains the phoneme /e/ and English contains the phoneme /ɛ/. This study investigates how Language Mode, Language Dominance, and L3 proficiency in Spanish-Gallego early bilinguals affect front mid-vowel production in L3 English. Language Mode (Grosjean 1985) represents the real-time activation of a speaker’s languages and the associated processing frameworks, heavily influenced by environmental and psychosocial factors. We manipulate participants' Language Modes using story-retelling and sentence-reading tasks across their languages. The resulting Euclidean distances between vowel tokens will be measured and analyzed against Bilingual Language Profile (Birdsong, Gerken, and Amengual 2012) Language Dominance ratings and self-reported English proficiency. We expect to find that participants with higher proficiency in English will have greater distinction between /e/ and /ɛ/ phonemes across all three languages, that stronger Language Dominance of Gallego over Spanish will correlate with greater distinction between these phonemes, and that activation of Gallego or English Language Mode will also correlate with a greater distinction on average for speakers. We anticipate that participants with low self-reported English proficiency will see greater distinction between these phonemes in the Gallego Language Mode conditions, but not within English Language Mode, following Escudero's (2005) Second Language Linguistic Perception model and it's extension to phonological production as proposed by Casillas and Simonet (2018).
Abstract This study examines high sensitivity (HS) in canine units of the Spanish State Security Forces, assessing its psychological, operational, and decision-making impact. A mixed-methods design combined HS screening in handlers and dogs with a qualitative phase of 13 in-depth interviews, coded into 23 categories and analysed through co-occurrence, Jaccard index, and prosodic analysis on a subsample of 11 audio recordings. Results depict a pattern of multisensory hyper-reactivity, fast learning, high correction sensitivity, and strong emotional attunement to the handler. Performance follows a biphasic pattern, vulnerable to overstimulation but outstanding when management is aligned with the dog’s sensory threshold. The handler’s perception of HS emerges as a critical factor shaping the entire operational cycle: from puppy selection and potential identification to role assignment, performance evaluation, and adult dog dismissal. Institutional norms emphasizing emotional control may contribute to under-recognition of sensitivity traits, potentially affecting the identification and retention of high-potential individuals. We conclude HS is a context-sensitive trait and propose incorporating psychological, sensory, and behavioural indicators into selection and training, complemented by lexical–semantic and prosodic metrics as diagnostic tools when classical tests are unsuitable for specialised populations.
Examining specific linguistic aspects in isolation, traditional language tests fail to capture the dynamic interactions that support discourse production abilities. The accurate assessment of discourse production is therefore crucial for identifying language difficulties and procedures of discourse analysis have emerged as a valid methodological solution. Despite this, heterogeneity in discourse measures limits comparability across studies, and the lack of normative data across the adult lifespan complicates the differentiation between healthy and pathological aging. The present study addresses this issue by providing the first standardization of linguistic measures extracted using a MultiLevel procedure of discourse Analysis (MLA). Narrative samples from 717 healthy Italian-speaking adults (aged 20 - 94) were elicited through a picture description task using two single images and three vignettes. Speech samples were transcribed and analyzed using a semi-automatic pipeline. Normative data, adjusted for age and education, and data-driven age bands were calculated for linguistic measures assessing productivity, lexical difficulties, grammatical construction, macrolinguistic difficulties, and lexical informativeness. Results provide standardized norms across multiple linguistic measures and reveal distinct age-related shifts in performance. Together, these findings offer the most comprehensive adult lifespan framework to date for narrative discourse production and highlight the importance of data-driven age bands for research and clinical assessment in healthy aging. Furthermore, they show that healthy aging disproportionately affects higher-order integrative discourse mechanisms rather than core lexical and morphosyntactic encoding processes.
Introduction. In the present-day scientific discourse, there is a great number of theoretical research focusing on the problem of objectifying the semiotic nature of law and analysing the functional construct of legal semantics. Whereas, many practical legal issues, such as: interpretative ambiguity in the meaning-formation and meaning-application of normative acts, lexical vagueness and contextual dependence of legal notions and the incoherence of legal terminology across different legal systems, remain neglected, which leads to contradictions and inaccuracies in legal practice. The aim of the study is to define the methodological principles fostering establishment of the acceptable scope of semantic interpretation of legal notions in the context of building a legal thinking culture. Materials and Methods. The research methodology was based on the principle of jurisprudential definition of legal norm meaning-formation in socio-legal discourse. Analytical, systematizing and pragmatic methods were used to reveal a complex nature of the semantics of law in the context of legal thinking development. The semiotic analysis of the objectivity and normativity of legal notions taking into account the contextual differences of legal definitions, was used as a specialised research method. Results. It was established that normative notions are the complex semantic constructs encompassing a conceptual sphere (normativity) and social reality. For building sustainable models of legal behaviour and legal culture, it is necessary to overcome external and internal conflicts in interpretation of law. In this regard, a number of advisory measures were proposed aimed at establishing acceptable scope of semantic interpretation: differentiation between the informational nature of prescriptive and descriptive notions, semantic monitoring of legal phenomena, and implementation of the principle of discourse contextualism, which makes it possible to formulate the normativity of law requirements based on the specific contextual interpretations. Discussion and Conclusion. A justified conclusion about possibility of a properly selected semantic toolkit to determine the objectivity of perception of the legal norms and, consequently, to improve the process of building a legal culture was drawn. The main advantage of the principle of discourse contextualism such as conjunction of the semantics and pragmatics of legal notions was identified, which provides a fruitful foundation for further theorizing on the nature and metaphysics of law.
Visual perception of built environments contributes to the affective impressions that people form in everyday life. However, how these impressions are represented within vision foundation models remains largely unexplored. To support the systematic investigation of this subject, we introduce the Emotional Impression of Spaces (EMOIS) dataset, comprising 1,544 real-world built-environment images. Each image is annotated with image-evoked valence and arousal ratings collected from Japanese adults by conducting a large-scale web-based survey, with approximately 120 ratings per image. Using Contrastive Language--Image Pre-training (CLIP) representations, we perform predictive and geometric analyses to systematically investigate how valence and arousal are encoded and organized within the representation space. These analyses reveal that valence exhibited stronger and more coherent organization than arousal. Cross-dataset analyses with the Open Affective Standardized Image Set (OASIS), a benchmark dataset of general affective photographs, reveal differences in affective organization between the two datasets. Regression analyses demonstrate high predictive performance for valence and arousal within EMOIS, with mean coefficients of determination of 0.865 and 0.807, respectively, across repeated internal hold-out evaluations. Finally, we present an example-based interface illustrating how learned representations can support qualitative interpretation of predicted affective values. These findings can help elucidate affective representations of built environments and establish EMOIS as a densely annotated resource for future affective computing research in this domain.
Abstract:In natural language processing, text segmentation is a crucial task that can be approached in various ways, depending on the level of detail required. Examples include segmenting a document into different topical segments or dividing a sentence into smaller units called elementary discourse units (EDUs). Traditional methods for these tasks relied heavily on carefully crafted features, but they have limitations. SEGBOT is our proposed solution to address these limitations. SEGBOT is an all-in-one segmentation model that leverages a bidirectional recurrent neural network to encode an input text sequence. In addition, SEGBOT incorporates another recurrent neural network and a pointer network to identify text boundaries within the input sequence. Our hierarchical model can simultaneously utilize both word-level and EDU-level information for sentence-level sentiment analysis. Through our experiments, we have demonstrated that our model surpasses previous approaches on the benchmarks of Movie Review and Stanford Sentiment Treebank.
This research is situated at the intersection of digital humanities, the history of emotions, and computational linguistics. The article presents the results of the sentiment analysis of the epistolary heritage and diary entries of three key figures of the Russian monarchy: Catherine II, Alexander I, and Nicholas I. The total corpus of analyzed texts amounted to over 2 million word usages. Using models of deep learning BERT (XLMRoBERTaLarge and Conversational RuBERT) adapted for historical texts, the authors reconstruct the emotional dynamics of communication throughout a turbulent century—from Enlightened absolutism to the crisis of the Nicholas system. The study confirms the hypothesis of a stable correlation between the genre of the document (official letter vs. private diary) and the degree of emotional expressiveness, as well as identifies specific lexical markers of anxiety during periods of political instability (on the eve of the Decembrist revolt and during the Crimean War). Methodologically, the research is based on three conceptual foundations. Firstly, it is the theory of emotional communities by B. Rosenwein, according to which emotions are constructed within social groups with shared values and norms of expression. Secondly, it adopts the semiotic approach of Yu. M. Lotman in studying the everyday behavior of the Russian nobility. Thirdly, it employs methodologies of computational text analysis. The hypothesis of this study is as follows: the emotional tone of the personal correspondence and diaries of Russian monarchs is not so much a spontaneous expression of an individual psychological state, but rather a ritualized social action, subject to the cultural codes of the era and genre canon. The aim of the study is to conduct a comprehensive historical-linguistic analysis of the emotional tone of the epistolary and diary heritage of Russian statesmen of the 18th-19th centuries using digital text processing methods, to identify stable emotional patterns and their connection with historical-biographical context. It has been established that the sentimentalist tradition of the late 18th century paradoxically combined with hypertrophied emotional restraint in official communication, creating an effect of "emotional dissonance," which was resolved in the literature and journalism of the 19th century. This work contributes to the methodology of analyzing historical texts, demonstrating the possibilities and limitations of NLP tools when working with archaic vocabulary and bilingual corpora (Russian-French linguistic dualism).
We developed a Chinese version of the "Reading the Mind in the Eyes" Test (RMET-C) and conducted an initial examination of its psychometric properties in a sample of healthy Chinese adults. Candidate images were selected from publicly available face databases through a rating procedure in a small sample. We then administered three tasks—target-word matching, sex recognition, and emotional valence rating in a larger sample, with iterative image screening based on item analysis results. Reliability analyses indicated that the target-word matching task had acceptable internal consistency and good test-retest reliability. Inter-rater reliability and test-retest reliability for emotional valence ratings reached good levels, and sex determination showed exceptionally strong cross-rater consensus. Confirmatory factor analysis supported an acceptable fit for a unidimensional model in the calibration sample. Within-subject comparisons in Phase 3 revealed that Chinese participants performed significantly better on the RMET-C than on a matched set of original Western RMET items. The total scores of the two sets were moderately positively correlated, providing initial evidence for comprehensibility and convergent validity of the RMET-C in a native Chinese sample. This test may provide a measurement tool with clear item properties and preliminary psychometric evidence for future theory-of-mind research in the Chinese cultural context, suitable for group-level assessment of social cognition.
This study investigates the influence of artificial intelligence on the evolution of language in modern society. It analyzes how AI-driven technologies—including machine translation, natural language processing, chatbots, and speech recognition systems—affect language use, transformation, and development. The research highlights that the growing integration of AI into everyday communication has accelerated multilingual interaction, contributed to the reshaping of linguistic norms, and facilitated the emergence of new vocabulary associated with digital environments. Furthermore, the article examines the dual impact of AI on language practices. On the one hand, it enhances linguistic diversity and expands access to global communication; on the other hand, it raises critical concerns regarding linguistic standardization, the erosion of cultural nuances, and increasing dependence on automated systems. Particular emphasis is placed on the role of AI in language education, where adaptive learning platforms and intelligent tutoring systems enable more personalized instructional pathways. The findings indicate that artificial intelligence should be understood not merely as a technological advancement, but as a significant sociolinguistic force that is reshaping patterns of communication, literacy practices, and cultural identity in the digital era.