Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This article examines the “Golden Age” of Arabic linguistics under the Abbasid Caliphate (750–1258). It traces how Arabic evolved from a primarily religious and literary medium into a universal language of science. The study highlights the scholarly rivalry between the Basra and Kufa grammatical schools—especially their debates over qiyās (analogy) and samʿ (attested usage/auditory transmission)—and shows how these methodological differences contributed to the codification of linguistic norms. The article also analyzes the foundational role of Sibawayh’s Al-Kitāb and al-Khalīl ibn Aḥmad’s Kitāb al-ʿAyn in systematizing Arabic syntax, phonetics, and lexicography. Finally, it evaluates the Translation Movement at the House of Wisdom (Bayt al-Ḥikma) and explains how the integration of Greek, Persian, and Indian learning enriched Arabic vocabulary and helped establish a broad scientific terminology.
Abstract Introduction Sleep supports emotion regulation by preferentially consolidating emotional memories while attenuating reactivity. We have shown that dream recall plays an active role by increasing negative over neutral memories and reducing reactivity. In women, fluctuating reproductive hormones across the menstrual cycle influence sleep features implicated in emotional memory, yet whether menstrual phases influence how dreams shape emotional processing remains unknown. This study investigates how dreams shape sleep-dependent emotional processing across the menstrual cycle in naturally cycling women. Methods 128 women (Mage = 32.85 ±11.93 years) completed up to four visits across verified menstrual phases (menses, late-follicular, mid-luteal, late-luteal). At each visit, participants performed the Emotional Picture Task with negative and neutral IAPS images in the evening (Test 1) and the next morning (Test 2). Participants rated old/new, arousal, and valence of images shown at each test. Dream reports were collected upon waking prior to Test 2. Linear mixed-effects models tested main and interaction effects of menstrual phase and dream recall. Results The menstrual cycle altered how dreaming shaped overnight emotional memory. Dream recall typically benefited the emotional trade-off effect —favoring consolidation of negative relative to neutral images (Δd′; t(410)=1.95, p=0.05)—but this pattern reversed during the late-luteal phase (dream × menstrual cycle: t(381)=-2.29, p=0.02). Dreaming showed independent effects on emotional reactivity. Higher valence and arousal ratings for negative images during Test 1 predicted greater dream recall (valence: t(344)=2.05, p=0.04; arousal: t(327)=2.04, p=0.04). Additionally, the more negatively participants rated the images at Test 1, the more negative their dreams tended to be (t(166)=-2.11, p=0.04). Dream recall was linked to reduced next-morning emotional reactivity (valence: t(413)=-2.89, p=0.004; arousal: t(413)=-2.65, p=0.01), with stronger reductions following more negative dreams (β=0.15, t(182)=2.86, p=0.005). Conclusion Menstrual cycle phase influenced how dreams shaped overnight emotional memory. Negative waking experiences increased dream recall and shaped dream content—and recalling dreams, especially negative ones, reduced emotional reactivity and typically strengthened emotional memory—but this benefit disappeared in the late-luteal phase when there are declining reproductive hormones. These findings suggest a novel interaction between the menstrual cycle and dreaming, showing that hormonal fluctuations reshape how sleep and dreams regulate emotional experience and memory. Support (if any) RF1AG061355 (Baker/Mednick)
Prediction systems grounded in textual data have become indispensable across high-stakes domains including clinical decision support, financial signal detection, and digital misinformation analysis. Classical statistical approaches and shallow machine learning methods have demonstrated satisfactory performance on narrow, well-curated datasets, but they struggle to generalise once input distributions shift or domain vocabulary diverges from training corpora. Deep learning, and more specifically the pre-trained transformer paradigm, has substantially narrowed this gap; nevertheless, single-architecture solutions routinely leave accuracy on the table when applied to tasks that demand both rich contextual encoding and explicit sequential reasoning. This paper presents a cohesive, end-to-end AI- powered prediction framework that fuses BERT- derived contextual representations with a two-layer bidirectional LSTM (BiLSTM) classification head augmented by an additive attention mechanism. The system is designed as a modular pipeline: text acquisition and normalisation, augmentation-based imbalance handling, deep encoding, sequential modelling, and post-hoc probability calibration are treated as independent, replaceable stages. Experimental evaluation across three publicly available benchmark datasets — the LIAR fake news corpus, Stanford Sentiment Treebank v2, and a health-claim verification collection — confirms that the hybrid BERT-BiLSTM-Attention architecture outperforms five competitive baselines on macro-averaged F1 and area under the ROC curve. Ablation experiments quantify the individual contributions of the attention layer, recurrent head, augmentation strategy, and temperature scaling. A discussion of deployment trade-offs addresses inference latency, continual adaptation, and algorithmic fairness..
<div> This paper examines register variation in Latin from the third century BCE to the fourteenth century CE using Key Feature Analysis (KFA) (Egbert and Biber, 2023), a quantitative method for identifying statistically over-and underrepresented linguistic features. Registers are defined as text varieties linked to communicative situations and characterized by distributions of lexico-grammatical features (Biber, 1988, 1995). Six dependency-parsed Universal Dependencies (UD) treebanks are classified a priori into ten register categories based on established scholarship. Additionally, Principal Component Analysis (PCA) is used to reduce dimensionality, in order to explore the texts major patterns of variation and clusters of linguistically similar texts. KFA reveals systematic register-specific grammatical profiles consistent with previous research (e.g. Biber (2014b)). Registers with involved language use (e.g. letters and speeches) show higher frequencies of personal reference and engagement, while philosophical texts favor subordination and impersonal constructions. Registers containing narrative elements (e.g. historiography, satire) contain high frequency of verbs in past tense. PCA places the charter register into a distinct cluster, while other registers form more closely grouped patterns. The strongest components reflect contrasts in number, aspect, tense and person, alongside subordination and cordination distributions. The results are largely confirmatory: KFA produces coherent and interpretable groupings of grammatical features consistent with previous findings, providing a proof of concept for quantitative register analysis in historical corpora. Data and code are openly available for future research. </div>
The practice of web form submission has emerged as a prime conduit for attackers, enabling them to infiltrate modern web applications and illegally harvest sensitive user data. Traditional defense mechanisms, such as static security reviews and server-side validation, are proving insufficiently agile for real-time detection of client-side vulnerabilities. This inadequacy arises directly from the rapid evolution of modern interfaces, which involves spontaneous DOM changes, dynamic element generation, and semantic interpretation that varies based on context and culture. This article presents an innovative browser extension framework that leverages a heuristic-based, multi-dimensional analytical engine combined with deep DOM inspection to identify insecure form submissions the moment they occur. The proposed methodology introduces five fundamental innovations: a contextual risk scoring system that models the complex interdependencies among form fields; an adaptive weighting scheme for risk patterns, accommodating diverse cultural and linguistic norms; a predictive vulnerability estimator that anticipates future threats; intelligent DOM mutation filtering designed to significantly optimize runtime performance; and cross linguistic semantic recognition to determine the true purpose of fields globally. Based on theoretical projections, this combined approach promises to enhance vulnerability detection accuracy while simultaneously reducing computational demands by approximately. Critically, all security analysis is executed exclusively on the user's local machine, guaranteeing privacy by ensuring no sensitive data is transmitted externally. A proof of concept application confirms the framework's practical feasibility and high efficacy for client side security assessment and catalyzing the development of flexible, scalable, and privacy respecting browser-based protections.
Data, code, figures, and the manuscript for The architecture of internet aesthetics (Y. J. Lin, Cornell University).Contents: (1) the internet-aesthetics ecosystem network of 1,115 aesthetics joined by 5,889 community-authored related-links, with community, affect, and era attributes and a self-contained interactive HTML explorer (pan / zoom / search); (2) the are.na human-label and co-curation records; (3) an openly-licensed 681-image stimulus corpus from Wikimedia Commons with per-image attribution (CREDITS.md) and a 681x37 perceptual and mechanism feature table; (4) three-model frontier vision-language affect ratings over the are.na images, from gemini-2.5-flash, gpt-4o, and claude-sonnet-4-6; (5) OASIS benchmarking and mechanistic-intervention (image-edit and prompt-reframe) outputs; and (6) all harvest, analysis, and figure code.The are.na images themselves are copyright-held and are NOT included in this archive; only derived features and model ratings are released. The 681 Wikimedia images ARE included, each retaining its own CC / public-domain licence (see CREDITS.md and LICENSES.md). The dataset compilation (manifests, feature tables, ratings, network, and code) is released under CC-BY-4.0.
Abstract: The Emperor Marcus Aurelius and the former slave Epictetus represent the social poles of the Roman Empire, yet both are cornerstones of late Stoic thought. This study employs digital humanities tools to investigate how their disparate life experiences and professional roles produced divergent philosophical "signatures" in their extant literature. By analyzing the Lemmatized Ancient Greek Texts (LAGT) corpus, we identify a distinct linguistic polarity: Marcus Aurelius demonstrates a significant preference for physical and cosmological terminology, reflecting a Stoicism centered on the providential order of the universe. Conversely, Epictetus’s lexicon shifts toward terms of ethical practice, pedagogy, and the transformation of the moral will. While Marcus Aurelius employs a more poetically diverse and intellectually wide-ranging vocabulary, Epictetus utilizes a more repetitive, concentrated technical vocabulary suited for the classroom. Despite these differences, a high degree of overlap reveals a "common core" of Stoic concepts shared by both authors, such as the nature of impressions and the primacy of the divine. These findings quantitatively highlight the adaptability of Stoicism, illustrating how a robust philosophical core was reframed to serve both the private reflections of a struggling ruler and the public exhortations of a committed teacher. Technical Context & Methodology This research integrates philology with a computational pipeline to analyze late Stoic literature. The following technical components are included in this repository: Computational Environment: All analyses were performed using Python 3.11. The pipeline utilizes Pandas and PyArrow for high-speed data processing, and the Classical Language Toolkit (CLTK) for part-of-speech tagging and grammatical filtering. Corpus Data: The primary linguistic data was extracted from the Lemmatized Ancient Greek Texts (LAGT) v4.1 dataset, which provides advanced lemmatization via the GLAUx treebank and GreCy models. Lexicographical Mapping: English definitions were integrated using the LSJ Dictionary (JSON v1.0.0). A custom normalization pipeline was used to standardize lemmata into Normalization Form Canonical Composition (NFC). Lexical Metrics: Vocabulary richness was assessed using Type-Token Ratio (TTR), Guiraud’s Index (R) to compensate for corpus size differences, and the percentage of hapax legomena (terms appearing only once). Generative AI Integration: A Gemma-3-27b-it model was utilized for the thematic classification and translation of 5,371 sentences. Sentences were tagged into the traditional Stoic tripartite division—Logic, Physics, or Ethics—based on the framework established by Pierre Hadot. Visualizations: The included scripts generate Lexical Volcano Plots (mapping total relative frequency against authorial skew) and Weighted Word Clouds that distinguish between author-specific signatures and the "Shared Stoic Core". Files included in this record: Supplementary File S1: Complete Python computational pipeline, README, and requirements.txt. Supplementary File S2: stoic_master_comparison.tsv containing comprehensive lemma frequencies and delta-RF values. Supplementary File S3: Statistical visualizations, including KDE overlap plots and delta-RF histograms. Supplementary File S4: Thematic analysis CSV containing 5,371 sentences with original Greek, English translations, and AI-generated thematic tags.
The study aim was to evaluate differences in ratings of valence made for a set of nonspeech sounds varying in spectrotemporal modulation. For 17 auditory-frequency bands (from 125 to 9474 Hz), modulation index values were extracted at five rates (from 2 to 32 Hz). Higher ratings of valence were associated with lower modulation values in the frequency region below 500 Hz and higher modulation values in the 1620-4000 Hz region. Listeners with hearing loss demonstrated significantly shallower relationships between modulation index and ratings of valence compared to adults with normal hearing. This could partially explain previously demonstrated compressed valence associated with hearing loss.
BACKGROUND: This pilot randomized controlled trial evaluated the effectiveness of an artificial intelligence (AI)–assisted solo workflow for intraoral photography training. The study examined whether real‑time AI feedback could enhance photographic quality, procedural efficiency, learner self‑efficacy, and patient comfort compared with conventional approaches. METHODS: Fifty-four first‑year dental students were randomly assigned to one of three groups: assistant‑supported workflow (four‑handed technique, control), solo workflow without AI support, and solo workflow with AI‑driven real‑time feedback. All participants performed standardized intraoral photography tasks. The primary outcome was a composite photographic quality score derived from expert ratings of three standardized intraoral views (frontal intercuspal, frontal open-bite, and lateral intercuspal), each rated on a 0–10 scale (total range 0–30). Data were analyzed using ANOVA; mean differences (MD) with 95% confidence intervals (CI) were calculated. RESULTS: Inter‑rater reliability for expert image ratings was good (ICC = 0.84, 95% CI: 0.72 to 0.90). The AI-supported solo group achieved the highest composite quality scores (18.2 ± 2.7). This was significantly superior to the unassisted solo group (15.8 ± 3.3), with a mean difference (MD) of 2.4 points (95% CI: 0.45 to 4.35; p = 0.027) and a large effect size (Cohen’s d = 0.80). Compared to the assistant-supported group (17.1 ± 2.1), the difference was not statistically significant (MD = 1.1; 95% CI: -0.65 to 2.85; p = 0.28). Secondary outcomes, including task completion time (F(2,51) = 1.25, p = 0.30), self‑efficacy (all p > 0.40), and patient‑reported comfort (χ²(4, N = 54) = 5.2, p = 0.27), showed no significant between‑group differences. CONCLUSION: In this single‑centre pilot trial, an AI‑assisted solo workflow enabled novice dental students to achieve higher intraoral photographic quality than unguided solo operation, with performance broadly comparable to a conventional four‑handed assistant‑supported workflow and without detectable compromises in efficiency, self‑efficacy, or patient‑reported comfort. These preliminary findings may serve as a valuable adjunct for autonomous skill acquisition, warranting further validation in larger, multi-institutional cohorts. CLINICAL TRIAL NUMBER: Not applicable. This study evaluated an educational training intervention rather than a clinical treatment, and prospective trial registration was not required under institutional policy at the time of initiation. Ethical approval was obtained from Shanghai Ninth People’s Hospital Ethics Committee (SH9H-2022-T30-1).
This study has investigated how teachers at a rural school and an urban school perceive the interplay between dialect, dialect levelling and linguistic norms in teaching. The researcher also observed four Swedish language lessons in grade 4 and 6 at each school. In addition, one lesson in each grade was observed in other subjects, such as mathematics, history, geography and home and consumer studies, resulting in a total of eight observed lessons. To explore the teachers’ perceptions, four semi-structured interviews were conducted with teachers who teach Swedish. The semi-structured interviews, in combination with participant observations, complemented each other well and contributed to strengthening, problematizing, and nuancing the findings. The results show that students’ spoken language varies between a Västerbotten dialect, a more standard variety of Swedish, and a form of language influenced by social media, including slang, abbreviations, informal chat language, and English words and expressions. It emerged that the teachers perceived the influence of social media on language as relatively strong, and that words becoming popular on social media spread quickly and become a natural part of students’ spoken language. The observations also revealed variations in how students’ spoken language was expressed depending on the teaching situation and context. Furthermore, the results show that different forms of dialect are still present in students’ speech through dialectal words and expressions, as well as features such as stress, prosody, and pronunciation. Signs of dialect levelling were also identified, as some students used a more standardized form of language in different situations and contexts.
The intricate relationship between truth and language has long fascinated philosophers, linguists, and scholars across disciplines. In this work, Prof. Dr. Yoesoep Edhie Rachmad, Ph.D., DBA., embarks on a profound exploration of how language shapes, conveys, and sometimes distorts truth. Published in 2000 under The United Nations and The Education Training Centre, this book critically examines the philosophical foundations of linguistic representation and its implications for human understanding. Through an analysis of classical and contemporary theories, the work delves into the roles of meaning, interpretation, and the social constructs that govern communication. By questioning whether absolute truth can ever be expressed without ambiguity, this book invites readers to reconsider their assumptions about knowledge, semantics, and reality itself. Language is the primary medium through which humans express ideas, communicate beliefs, and establish shared realities. However, can language truly capture the essence of truth? This book is born from a deep intellectual curiosity regarding the limitations and potentials of linguistic structures in truth-seeking endeavors. With rapid advancements in technology, media, and cross-cultural exchanges, the question of how truth is conveyed in different languages and frameworks becomes increasingly urgent. Addressing these concerns, this book seeks to bridge the philosophical and practical aspects of truth in language. Understanding truth within the linguistic paradigm requires an examination of key philosophical theories, such as the correspondence, coherence, and pragmatic theories of truth. This book dissects the works of foundational thinkers, including Wittgenstein, Austin, Quine, and Derrida, to present an encompassing view of how meaning is formed and interpreted. Concepts such as semiotics, hermeneutics, and speech act theory play central roles in deciphering the mechanisms through which language represents reality. The complexities of truth in language manifest in various real-world phenomena, including political discourse, media manipulation, translation challenges, and legal interpretation. This book investigates how different languages construct reality in unique ways, leading to variations in perception and understanding. It also considers the implications of artificial intelligence in language processing and whether machines can ever truly comprehend truth. By integrating classical and modern linguistic theories, this book establishes a framework for analyzing truth in communication. It presents a structured approach to evaluating how meaning is derived, how context shapes interpretation, and how linguistic structures influence perception. The core principle is that truth is not merely an objective reality but is also shaped by the languages and systems within which it is communicated. Key indicators of linguistic truth include coherence, consistency, factual alignment, and pragmatic effectiveness. This book identifies operational variables such as cultural context, linguistic ambiguity, speaker intent, and audience interpretation as fundamental factors in understanding how truth is conveyed through language. Numerous elements influence the relationship between language and truth, including cognitive biases, historical contexts, power structures, and evolving linguistic norms. By analyzing these factors, this book provides insights into why different cultures and societies perceive truth differently and how language can both reveal and obscure reality. Applying the theoretical insights from this book, readers are guided through strategies for enhancing clarity, reducing misinterpretation, and fostering more precise communication in various fields, from academia to politics. The book also addresses how language policies, media regulations, and ethical considerations shape the pursuit of truth in public discourse. While language serves as a bridge to truth, it is also a source of distortion, manipulation, and misunderstanding. This book highlights the challenges posed by misinformation, propaganda, and ideological biases, while also recognizing the supporting role of linguistic diversity and education in fostering more nuanced truth-seeking. Truth and language are inseparable in human thought and communication. Through this in-depth exploration, the book underscores the importance of linguistic awareness in navigating the complexities of truth. Whether in philosophy, politics, law, or everyday conversation, understanding how language constructs and conveys truth is vital for clearer and more meaningful human interactions.
Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing image captioning benchmarks typically suffer from limited diversity in caption length, the absence of recent advanced MLLMs, and insufficient human annotations, which potentially introduces bias and limits the ability to comprehensively assess the performance of modern MLLMs. To address these limitations, we present a new large-scale image captioning benchmark, termed, ICBench, which covers 12 content categories and consists of both short and long captions generated by 10 advanced MLLMs on 2K images, resulting in 40K captions in total. We conduct extensive human subjective studies to obtain mean opinion scores (MOSs) across fine-grained evaluation dimensions, where short captions are assessed in terms of fluency, relevance, and conciseness, while long captions are evaluated based on fluency, relevance, and completeness. Furthermore, we propose an automated evaluation metric, \textbf{ITIScore}, based on an image-to-text-to-image framework, which measures caption quality through reconstruction consistency. Experimental results demonstrate strong alignment between our automatic metric and human judgments, as well as robust zero-shot generalization ability on other public captioning datasets. Both the dataset and model will be released upon publication.
Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.
Supplementary materials for manuscript Location-scale models improve within-participant held-out trial prediction in Stroop interference and attractiveness and dominance ratings
Lexical Generativity in English Lexical Generativity in English: An Empirical Study of Verb Classes, Noun Polysemy, and Prepositions Author: Pablo Nogueira Grossi · G6 LLC · Newark NJ ORCID: 0009-0000-6496-2186 Series root: https://doi.org/10.5281/zenodo.19117399 Submitted to: International Journal of Lexicography (Oxford University Press) License: CC BY 4.0 Abstract This article examines lexical generativity in English: the capacity of a finite lexical inventory to support a theoretically unbounded range of context-sensitive meanings in use. Drawing on three converging empirical resources — the verb classification system of Levin (1991, 1993), the frame-semantic architecture of FrameNet (Fillmore, Johnson & Petruck, 2003), and the class-membership and alternation structure of VerbNet (Kipper, Korhonen, Ryant & Palmer, 2008) — the article proposes a layered model of lexical generativity that distinguishes among: (a) stored semantic primitives and qualia structure (Level I) (b) argument-structure templates licensed by class membership (Level II) (c) event-type composition rules governing productive meaning extension (Level III) The empirical core consists of detailed analysis of twelve English verb classes: manner-of-motion, change-of-state, causative-inchoative alternation, communication verbs, psychological verbs, creation-and-transformation verbs, aspectual verbs, verbs of putting, spray-load verbs, contact-by-impact verbs, perception verbs, emission verbs, and verbs of appearance and disappearance. The analysis is extended to systematic noun polysemy — dot objects, type coercion, and metonymic transfer — and to the generative semantics of English spatial prepositions, with a case study of over. Throughout, the article argues that apparent lexicographic irregularities are systematic consequences of a small set of generative principles that can be stated precisely, incorporated into lexicographic description, and exploited in computational lexicography. Implications for dictionary design, large-scale lexical database annotation, and natural language processing are discussed. Keywords: lexical generativity · verb classes · noun polysemy · prepositions · FrameNet · VerbNet · Levin classes · generative lexicon · computational lexicography · argument structure · type coercion · metonymy · qualia structure Deposit Contents File Description lexical_generativity_en.pdf Main article, ~14,000 words Article Structure Section Content 1 Introduction: three forms of lexical generativity 2 Background: Levin verb classes, FrameNet, VerbNet 3 Theoretical framework: the three-level model 4 Verb class analyses (twelve classes) 5 Noun polysemy: dot objects, type coercion, metonymic transfer 6 Prepositions and spatial semantics (over case study) 7 Computational implications: dictionary design, database annotation, NLP 8 Discussion 9 Conclusion Submission Status Submitted to the International Journal of Lexicography (Oxford University Press). This preprint is posted in accordance with OUP's preprint policy. Related Deposits Work DOI Principia Orthogona series root https://doi.org/10.5281/zenodo.19117399 GCM Institutional Edition (context for this article) https://doi.org/10.5281/zenodo.19513913 Coherence Bridge v8.4 coherence_bridge_v8_4.yaml is the machine-readable synchronisation file between the TOGT five-operator grammar, GCM contact geometry, and all three formal pillars. Key changes from v8.3: Anantharaman–Monk source expanded to full arXiv series (arXiv:2304.02678, arXiv:2403.12576, arXiv:2502.12268); Hide–Macera–Thomas polynomial-rate follow-up (arXiv:2508.14874) noted; all five TOGT entries sharpened to precise mathematical statements. Wang–Zahl source expanded to arXiv:2502.17655 + precursor arXiv:2210.09581 + Guth surveys arXiv:2505.07695 / arXiv:2508.05475; conjecture-proved status corrected throughout (was mislabelled open); claim_level_note field added to prevent misreading of analogical tag; all five TOGT entries fully populated. Three formal pillars — sorry inventory Pillar Lean file Proved Sorry Discrete (Collatz) DiscreteDm3.lean v1.6 operatorDecomposition, contactForm meanContraction, lyapunovDescent, hasStructuredCycle Continuous (Navier–Stokes) Dm3Cont.lean v1.0 operatorDecomposition, contactForm meanContraction_cont, lyapunovDescent_cont, hasStructuredAttractor Arithmetic-analytic (BSD) BSD_dm3.lean v1.0 operatorDecomposition, contactForm meanContraction_BSD, lyapunovDescent_BSD, hasStructuredCycle_BSD Closing the three admits on any pillar turns the corresponding conjecture into a categorical corollary of the dm³ framework. Python Simulation — Reproduce Figures pip install numpy matplotlib python3 autophagy_dm3.py --out figures/ Generates all four paper figures. The nbonacci_criticality.py and nbonacci_critical_lambda.py scripts in the AXLE repository reproduce the DNLS / n-bonacci criticality figures from the companion paper (DOI: 10.5281/zenodo.20026942). Related Deposits Paper DOI Principia Orthogona series root 10.5281/zenodo.19117400 This deposit (Autophagy / Triple-alpha) 10.5281/zenodo.20168812 DNLS / n-bonacci companion paper 10.5281/zenodo.20026942 Fruit-fly / MultiOrbitBioSwarm 10.5281/zenodo.19210136 GCM Institutional Edition (manifesto) 10.5281/zenodo.19513913 Keywords dm³ operator · contact geometry · Whitney fold · autophagy · triple-alpha process · Lean 4 · Mathlib4 · formal verification · TOGT · operator grammar · coherence bridge · Collatz · Navier–Stokes · BSD conjecture · stability radius · ε₀ = 1/3 · Principia Orthogona · G6 LLC
Paper 6 (Silva 2026) introduced BPE Mean Vocabulary Morpheme Length (VMML) as a writing system classifier and showed that the Voynich Manuscript occupies a discriminant zone (VMML = 5.918, 95% CI 5.77-6.05) above all 15 tested alphabetic natural languages. This paper (v2.5) expands to 71 corpora across 40+ languages and reports six extended analyses: (1) Alphabetic ceiling confirmed at 5.76; (2) Tagalog (VMML=5.914) is the sole natural-language entry into the Voynich CI, but BC=0.202 distinguishes it from Voynich (BC=0.361); (3) Romanization inflates VMML by 2.4-5.3 units (methodological confound). Extended analyses: (4) Currier A vs B: delta VMML=+1.27, delta CBMI=+0.16 bits -- two quantifiably distinct writing registers; (5) BC coherent across all 7 manuscript sections (CV=6.7%) -- single writing system confirmed; (6) 3D discriminant (VMML x BC x CBMI): Voynich isolated, nearest natural-language neighbor Irish at distance 0.17; (7) Six named hoax mechanisms (monoalphabetic, Vigenere/barbavara, Vigenere/Italian-Knowles 2026, null insertion, syllabic compression, vocabulary shuffle) each fail all three criteria simultaneously; (8) BC orthogonal to all classical textual metrics (|r| < 0.23 vs entropy, TTR, hapax, Zipf) -- genuinely new structural dimension. All code and six extension scripts publicly available in companion repository. v2.3 (2026-06-08): Section 5.9 added - per-folio Currier A/B reanalysis using the Gaskell and Bowern (2022) canonical corpus (36,361 tokens, min_freq=5 BPE). Cross-boundary mutual information (CBMI) identified as primary discriminant: CBMI_A = 1.97 bits vs CBMI_B = 1.51 bits, Cohen d = -1.01, permutation p less than 0.001 (n = 10,000 shuffles, Bonferroni-corrected). CBMI survives within-quire control (pooled nA=46, nB=33; permutation p = 0.0008; Fisher combined within-quire p = 0.001), ruling out manuscript section as a confound. All three metrics (BC, BPE-ratio, CBMI) show A greater than B direction. Fisher combined full-corpus: chi-squared(6) = 40.66, p less than 0.000002. Section 5.1 corrected: direction is A greater than B on BC and CBMI. Finding is orthogonal to Parisel (2026) vowel-selection model. Conclusion 12 added. v2.4 (2026-06-09): §5.10 added — Currier-preserving null model (n = 200 iterations, size-matched) quantifying each metric's section-discrimination sensitivity independently of dialect. Key result: CBMI is the weakest section discriminant (mean |z| = 1.20 across six sections), confirming that the large CBMI A/B gap (§5.9) is not a section-composition artifact. STTR@100 is the strongest section discriminant (mean |z| = 4.75). Herbal section shows anomalously low vocabulary diversity (STTR z = -13.9); Stars shows anomalously high unique vocabulary (Hapax@500 z = +4.4). Demonstrates two independent organizational layers: CBMI tracks dialect, STTR tracks content domain. Conclusion #13 added. v2.5 (2026-06-10): Corpus expanded from 55 to 71 corpora across 40+ languages. §5.11 adds five medieval European corpora in native script via Universal Dependencies treebanks (Gothic transliteration, Old Church Slavonic, Old East Slavic, Ancient Greek PROIEL and Perseus; VMML 3.54-5.18 — all below alphabetic ceiling of 5.748). §5.12 adds 11 Australian Aboriginal language corpora via BibleNLP/eBible (Pama-Nyungan Western Desert, Ngumpin-Yapa, Arandic; Yolngu; Gunwinyguan; Daly; VMML 6.09-8.00 — predominantly above the Voynich zone). Warlpiri (VMML 5.851) is the sole near-entry on VMML but fails BC (0.233) and CBMI (0.244); 3D normalized distance from Voynich = 0.746 (vs. Irish = 0.200, the nearest neighbor from §5.4). Voynich zone is now charted on both sides: fusional alphabetic below (VMML 3.5-5.75), agglutinative-to-polysynthetic above (VMML 6.0-8.0). Voynich occupies a structural configuration not replicated by any of the 71 corpora tested. To our knowledge, this is the first systematic BPE profiling of Pama-Nyungan languages in the computational linguistics literature. Conclusions #14 and #15 added. v2.6 (2026-06-12): Section 5.10.1 adds a prose-only robustness check for the Section 5.10 Currier-preserving null model. Excluding all label, circular and radial loci (8.7% of tokens), every headline deviation survives essentially unchanged: Herbal STTR z = -13.5, Balneological z = -10.7, Stars Hapax z = +4.3; the sensitivity ranking is unchanged with CBMI last in both conditions. A mean-vs-median distributional note (both summaries rank lexical-diversity metrics first, boundary metrics last) and a coverage note (Astro/Zodiac folios carry no Currier tags and are outside any Currier-preserving design) are added. Erratum: Section 5.10 folio count corrected to 226 parsed / 196 Currier-labeled.
Ο γλωσσικός πόρος san-Corpus περιλαμβάνει σώμα κειμένων γραπτού λόγου της Νέας Ελληνικής, έκτασης περίπου 9 εκατομμυρίων λέξεων. Ο πόρος συγκροτήθηκε στο πλαίσιο διδακτορικής διατριβής, με στόχο τη μελέτη των συγκρίσεων ομοιότητας στη Νέα Ελληνική. Μέγεθος & Πηγές Το corpus περιλαμβάνει τρία ισομεγέθη υποσώματα, ώστε να επιτρέπονται οι συγκρίσεις μεταξύ τους: (α) Δημοσιογραφικός Λόγος: 7.774 άρθρα από τέσσερις διαδικτυακές εφημερίδες (Η ΑΥΓΗ, Η ΚΑΘΗΜΕΡΙΝΗ, ΕΘΝΟΣ, ΤΟ ΒΗΜΑ - έτος 2015). Συνολική Έκταση: 2,9 εκατ. λέξεις. (β) Εκπαιδευτικός Λόγος: 96 σχολικά εγχειρίδια δημοτικού και γυμνασίου. Συνολική Έκταση: 3,4 εκατ. λέξεις (μελετώνται 2,8 εκατ.). (γ) Λογοτεχνικός Λόγος: 28 μυθιστορήματα (βραβεία αναγνωσιμότητας περιόδου 2010-2015). Συνολική Έκταση: 2,5 εκατ. λέξεις. Κατανομή Σχολικών Εγχειριδίων ανά Γνωστικό Αντικείμενο Πλήθος Εγχειριδίων Ελληνική Λογοτεχνία 13 Ελληνική Γλώσσα 12 Ιστορία 9 Φυσική – Χημεία – Βιολογία 9 Μαθηματικά 9 Γεωγραφία – Γεωλογία – Περιβάλλον 8 Θρησκευτικά 7 Αγωγή Αισθητική (Εικαστικά – Μουσική – Θέατρο) 15 Αγωγή Υγείας (Φυσική Αγωγή – Οικιακή Οικονομία) 5 Πληροφορική – Τεχνολογία 5 Αγωγή Κοινωνική – Πολιτική 3 Αγωγή Σταδιοδρομίας (ΣΕΠ) 1 Κατάλογος μυθιστορημάτων: Συγγραφέας, Τίτλος Έτος 1ης έκδοσης Δούκα, Μάρω - Το δίκιο είναι ζόρικο πολύ 2010 Θέμελης, Νίκος - Η συμφωνία των ονείρων 2010 Καρυστιάνη, Ιωάννα - Τα σακιά 2010 Μιχαλοπούλου, Αμάντα - Πώς να κρυφτείς 2010 Ελευθερίου, Μάνος - Πριν απ' το ηλιοβασίλεμα 2011 Ζουργός, Ισίδωρος - Ανεμώλια 2011 Μακριδάκης, Γιάννης - Η άλωση της Κωνσταντίας 2011 Μπουραζοπούλου, Ιωάννα - Η ενοχή της αθωότητας 2011 Πανσέληνος, Αλέξης - Σκοτεινές επιγραφές 2011 Παπαδημητρίου, Χίλντα - Για μια χούφτα βινύλια 2011 Παπαθεοδώρου, Θοδωρής - Οι καιροί της μνήμης 2011 Τριανταφύλλου, Σώτη - Για την αγάπη της γεωμετρίας 2011 Φακίνος, Μιχάλης - Η έρημος έρχεται 2011 Βαμβουνάκη, Μάρω - Κυριακή απόγευμα στη Βιέννη 2012 Διβάνη, Λένα - Εγώ, ο Ζάχος Ζάχαρης 2012 Στεφανάκης, Δημήτρης - Φιλμ νουάρ 2012 Ακρίβος, Κώστας - Αλλάζει πουκάμισο το φίδι 2013 Ζέη, Άλκη - Με μολύβι φάμπερ νούμερο δύο 2013 Κορτώ, Αύγουστος - Το βιβλίο της Κατερίνας 2013 Κωνσταντούρου, Μαρία - Αγεφύρωτες σιωπές 2013 Μαντά, Λένα - Με λένε Ντάτα 2013 Ξανθούλης, Γιάννης - Κωνσταντινούπολη των ασεβών μου φόβων 2013 Ρώσση–Ζαΐρη, Ρένα - Άρωμα βανίλιας 2013 Ανδρουλάκης, Μίμης - Αλλέγκρα 2014 Δημουλίδου, Χρυσηίδα - Το κελάρι της ντροπής 2014 Παπαδοπούλου, Ελισάβετ - Μέρες και νύχτες που δεν ήταν δικές μας 2014 Χατζή, Αθηνά - Η θάλασσα έφυγε 2014 Χωμενίδης, Χρήστος - Νίκη 2014 Τεχνικές προδιαγραφές & Μορφότυπος Για την αναπαράσταση των δεδομένων και των μεταδεδομένων υιοθετήθηκε η πολυεπίπεδη οπτική των XML σχημάτων και τροποποιήθηκε το διεθνές πρότυπο TEI P5, 4.0.0 (Text Encoding Initiative). Δημιουργήθηκε ειδικός χώρος ονομάτων sanCorpus (sanC) με σχήμα τύπου RELAX-NG. Το σώμα κειμένων διατίθεται σε TXT και σε XML σε τρεις εκδοχές: Βάθος 0: απλό κείμενο (TXT). Περιλαμβάνει το main core (κείμενο βάσει του οποίου εξετάζονται οι συγκρίσεις ομοιότητας) και το out of core (κείμενο εκτός εμβέλειας της διατριβής, στο οποίο περιλαμβάνονται κείμενα που πλαισιώνουν το κυρίως κείμενο, π.χ. κείμενα διδασκαλίας, πίνακες περιεχομένων, εξώφυλλα) Βάθος 1 = κείμενα στην απλούστερη δυνατή XML κωδικοποίηση Βάθος 2 = κείμενα με πιο λεπτομερείς XML κωδικοποιήσεις. Αυτή η έκδοση (san-Corpus v1.0, Depth 0: Plain Text) περιλαμβάνει το σώμα κειμένων σε μορφή απλού κειμένου (Βάθος 0) στην αρχική του διάταξη (βλ. Επεξεργασία). Στόχος είναι ο σταδιακός εμπλουτισμός με επισημειωμένα δεδομένα, καθώς και με τις εκδοχές Βάθους 1 και 2. Επεξεργασία (Processing) Η μεθοδολογία συλλογής των δεδομένων, η θεωρητική τεκμηρίωση και το σχήμα επισημείωσης επεξηγούνται στις μελέτες Αφεντουλίδου (2022, 2021, 2013, 2012) και Afentoulidou (2009). Η πρώτη εκδοχή (Βάθος 0) χρησιμοποιήθηκε αποκλειστικά για τη λημματοποίηση που απαιτούσε η Collostruction Analysis (Gries, 2024). Στάδια επεξεργασίας (για τη λημματοποίηση): Τμηματοποίηση σε προτάσεις (sentence segmentation) με τη χρήση της βιβλιοθήκης Stanza (Stanford NLP Group, Qi et al. 2020), η οποία βασίζεται στο μοντέλο Greek Dependency Treebank (GDT) του Ινστιτούτου Επεξεργασίας του Λόγου / ΕΚ «Αθηνά». Τυχαία αναδιάταξη των προτάσεων για την προστασία της ακεραιότητας των πρωτότυπων έργων. Λημματοποίηση με τον ILSP Lemmatizer μέσω της Υποδομής Clarin-EL. Για την ανάλυση συμφράσεων του δείκτη σαν απομονώθηκαν συγκεκριμένοι λεκτικοί τύποι. Αδειοδότηση & Δικαιώματα Το san-Corpus συγκροτήθηκε για τις ανάγκες της διδακτορικής διατριβής και προστατεύεται από το δικαίωμα ειδικής φύσης σύμφωνα με την Οδηγία 96/9/ΕΟΚ και το άρθρο 45Α του Ν. 2121/1993. Η χρήση του περιεχομένου γίνεται αποκλειστικά για ερευνητικούς σκοπούς βάσει των εξαιρέσεων της Οδηγίας 2001/29 και της Οδηγίας (ΕΕ) 2019/790 (Text and Data Mining exceptions / Fair Use). Η πρόσβαση είναι περιορισμένη (Restricted Access) και παρέχεται αποκλειστικά σε μέλη της ακαδημαϊκής κοινότητας για σκοπούς επαλήθευσης των αποτελεσμάτων της διατριβής και περαιτέρω μη εμπορική έρευνα. Η πηγή προέλευσης δικαιούται να ζητήσει οποιαδήποτε τροποποιητική ενέργεια (π.χ. αφαίρεση) επί του πρωτότυπου περιεχομένου. Τέλος, η άδεια CC BY-NC-ND 4.0 ισχύει για την επιμέλεια (curation), τα μεταδεδομένα και τη γλωσσολογική επισημείωση του σώματος κειμένων. Βιβλιογραφικές αναφορές Αφεντουλίδου, Β. (2022). Σώμα ελληνικών κειμένων για τη μελέτη δομών ομοιότητας της Νέας Ελληνικής: σχεδιασμός και υλοποίηση. Στο Πρακτικά του 10ου Συνεδρίου Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας (σσ. 67-94). ΕΚΠΑ. Αφεντουλίδου, B. (2021). Δομές ομοιότητας στη Νέα Ελληνική. Σωματοκειμενικές παρατηρήσεις για τον πολυλειτουργικό δείκτη σαν. Προφορική ανακοίνωση στην 41η Ετήσια Συνάντηση του Τομέα Γλωσσολογίας, 13–15 Μαΐου 2021. ΑΠΘ. Αφεντουλίδου, Β. (2013). Και σου απάντησα κάτι σαν ‘τέλεια, εντάξει’. Δείκτης σαν + ευθύς λόγος;. Προφορική ανακοίνωση στο 7ο Συνέδριο Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας, 16–18 Μαΐου. ΕΚΠΑ. Αφεντουλίδου, Β. (2012). Συγκρίσεις ομοιότητας στα Νέα Ελληνικά: ο δείκτης σαν. Στο Z. Gavriilidou, A. Efthymiou, E. Thomadaki & P. Kambakis-Vougiouklis (Επιμ.), Selected papers of the 10th International Conference on Greek Linguistics (σσ. 696-707). DUTH. Afentoulidou, V. (2009). Sketching the σαν conditional construction in Modern Greek. Submitted essay, 2009 Linguistic Institute, Linguistic Structure and Language Ecologies, Linguistic Society of America and UC Berkeley. Gries, Stefan Th. 2024. Coll.analysis 4.1. A script for R to compute perform collostructional analyses. https://www.stgries.info/teaching/groningen/index.html Institute for Language and Speech Processing - Athena Research Center (2015). ILSP Lemmatizer. Version 1. [Software (Tool/Service)]. CLARIN:EL. http://hdl.handle.net/11500/ATHENA-0000-0000-23EE-D Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 101–108). Online: Association for Computational Linguistics.
Ο γλωσσικός πόρος san-Corpus περιλαμβάνει σώμα κειμένων γραπτού λόγου της Νέας Ελληνικής, έκτασης περίπου 9 εκατομμυρίων λέξεων. Ο πόρος συγκροτήθηκε στο πλαίσιο διδακτορικής διατριβής, με στόχο τη μελέτη των συγκρίσεων ομοιότητας στη Νέα Ελληνική. Μέγεθος & Πηγές Το corpus περιλαμβάνει τρία ισομεγέθη υποσώματα, ώστε να επιτρέπονται οι συγκρίσεις μεταξύ τους: (α) Δημοσιογραφικός Λόγος: 7.774 άρθρα από τέσσερις διαδικτυακές εφημερίδες (Η ΑΥΓΗ, Η ΚΑΘΗΜΕΡΙΝΗ, ΕΘΝΟΣ, ΤΟ ΒΗΜΑ - έτος 2015). Συνολική Έκταση: 2,9 εκατ. λέξεις. (β) Εκπαιδευτικός Λόγος: 96 σχολικά εγχειρίδια δημοτικού και γυμνασίου. Συνολική Έκταση: 3,4 εκατ. λέξεις (μελετώνται 2,8 εκατ.). (γ) Λογοτεχνικός Λόγος: 28 μυθιστορήματα (βραβεία αναγνωσιμότητας περιόδου 2010-2015). Συνολική Έκταση: 2,5 εκατ. λέξεις. Κατανομή Σχολικών Εγχειριδίων ανά Γνωστικό Αντικείμενο Πλήθος Εγχειριδίων Ελληνική Λογοτεχνία 13 Ελληνική Γλώσσα 12 Ιστορία 9 Φυσική – Χημεία – Βιολογία 9 Μαθηματικά 9 Γεωγραφία – Γεωλογία – Περιβάλλον 8 Θρησκευτικά 7 Αγωγή Αισθητική (Εικαστικά – Μουσική – Θέατρο) 15 Αγωγή Υγείας (Φυσική Αγωγή – Οικιακή Οικονομία) 5 Πληροφορική – Τεχνολογία 5 Αγωγή Κοινωνική – Πολιτική 3 Αγωγή Σταδιοδρομίας (ΣΕΠ) 1 Κατάλογος μυθιστορημάτων: Συγγραφέας, Τίτλος Έτος 1ης έκδοσης Δούκα, Μάρω - Το δίκιο είναι ζόρικο πολύ 2010 Θέμελης, Νίκος - Η συμφωνία των ονείρων 2010 Καρυστιάνη, Ιωάννα - Τα σακιά 2010 Μιχαλοπούλου, Αμάντα - Πώς να κρυφτείς 2010 Ελευθερίου, Μάνος - Πριν απ' το ηλιοβασίλεμα 2011 Ζουργός, Ισίδωρος - Ανεμώλια 2011 Μακριδάκης, Γιάννης - Η άλωση της Κωνσταντίας 2011 Μπουραζοπούλου, Ιωάννα - Η ενοχή της αθωότητας 2011 Πανσέληνος, Αλέξης - Σκοτεινές επιγραφές 2011 Παπαδημητρίου, Χίλντα - Για μια χούφτα βινύλια 2011 Παπαθεοδώρου, Θοδωρής - Οι καιροί της μνήμης 2011 Τριανταφύλλου, Σώτη - Για την αγάπη της γεωμετρίας 2011 Φακίνος, Μιχάλης - Η έρημος έρχεται 2011 Βαμβουνάκη, Μάρω - Κυριακή απόγευμα στη Βιέννη 2012 Διβάνη, Λένα - Εγώ, ο Ζάχος Ζάχαρης 2012 Στεφανάκης, Δημήτρης - Φιλμ νουάρ 2012 Ακρίβος, Κώστας - Αλλάζει πουκάμισο το φίδι 2013 Ζέη, Άλκη - Με μολύβι φάμπερ νούμερο δύο 2013 Κορτώ, Αύγουστος - Το βιβλίο της Κατερίνας 2013 Κωνσταντούρου, Μαρία - Αγεφύρωτες σιωπές 2013 Μαντά, Λένα - Με λένε Ντάτα 2013 Ξανθούλης, Γιάννης - Κωνσταντινούπολη των ασεβών μου φόβων 2013 Ρώσση–Ζαΐρη, Ρένα - Άρωμα βανίλιας 2013 Ανδρουλάκης, Μίμης - Αλλέγκρα 2014 Δημουλίδου, Χρυσηίδα - Το κελάρι της ντροπής 2014 Παπαδοπούλου, Ελισάβετ - Μέρες και νύχτες που δεν ήταν δικές μας 2014 Χατζή, Αθηνά - Η θάλασσα έφυγε 2014 Χωμενίδης, Χρήστος - Νίκη 2014 Τεχνικές προδιαγραφές & Μορφότυπος Για την αναπαράσταση των δεδομένων και των μεταδεδομένων υιοθετήθηκε η πολυεπίπεδη οπτική των XML σχημάτων και τροποποιήθηκε το διεθνές πρότυπο TEI P5, 4.0.0 (Text Encoding Initiative). Δημιουργήθηκε ειδικός χώρος ονομάτων sanCorpus (sanC) με σχήμα τύπου RELAX-NG. Το σώμα κειμένων διατίθεται σε TXT και σε XML σε τρεις εκδοχές: Βάθος 0: απλό κείμενο (TXT). Περιλαμβάνει το main core (κείμενο βάσει του οποίου εξετάζονται οι συγκρίσεις ομοιότητας) και το out of core (κείμενο εκτός εμβέλειας της διατριβής, στο οποίο περιλαμβάνονται κείμενα που πλαισιώνουν το κυρίως κείμενο, π.χ. κείμενα διδασκαλίας, πίνακες περιεχομένων, εξώφυλλα) Βάθος 1 = κείμενα στην απλούστερη δυνατή XML κωδικοποίηση Βάθος 2 = κείμενα με πιο λεπτομερείς XML κωδικοποιήσεις. Αυτή η έκδοση (san-Corpus v1.0, Depth 0: Plain Text) περιλαμβάνει το σώμα κειμένων σε μορφή απλού κειμένου (Βάθος 0) στην αρχική του διάταξη (βλ. Επεξεργασία). Στόχος είναι ο σταδιακός εμπλουτισμός με επισημειωμένα δεδομένα, καθώς και με τις εκδοχές Βάθους 1 και 2. Επεξεργασία (Processing) Η μεθοδολογία συλλογής των δεδομένων, η θεωρητική τεκμηρίωση και το σχήμα επισημείωσης επεξηγούνται στις μελέτες Αφεντουλίδου (2022, 2021, 2013, 2012) και Afentoulidou (2009). Η πρώτη εκδοχή (Βάθος 0) χρησιμοποιήθηκε αποκλειστικά για τη λημματοποίηση που απαιτούσε η Collostruction Analysis (Gries, 2024). Στάδια επεξεργασίας (για τη λημματοποίηση): Τμηματοποίηση σε προτάσεις (sentence segmentation) με τη χρήση της βιβλιοθήκης Stanza (Stanford NLP Group, Qi et al. 2020), η οποία βασίζεται στο μοντέλο Greek Dependency Treebank (GDT) του Ινστιτούτου Επεξεργασίας του Λόγου / ΕΚ «Αθηνά». Τυχαία αναδιάταξη των προτάσεων για την προστασία της ακεραιότητας των πρωτότυπων έργων. Λημματοποίηση με τον ILSP Lemmatizer μέσω της Υποδομής Clarin-EL. Για την ανάλυση συμφράσεων του δείκτη σαν απομονώθηκαν συγκεκριμένοι λεκτικοί τύποι. Αδειοδότηση & Δικαιώματα Το san-Corpus συγκροτήθηκε για τις ανάγκες της διδακτορικής διατριβής και προστατεύεται από το δικαίωμα ειδικής φύσης σύμφωνα με την Οδηγία 96/9/ΕΟΚ και το άρθρο 45Α του Ν. 2121/1993. Η χρήση του περιεχομένου γίνεται αποκλειστικά για ερευνητικούς σκοπούς βάσει των εξαιρέσεων της Οδηγίας 2001/29 και της Οδηγίας (ΕΕ) 2019/790 (Text and Data Mining exceptions / Fair Use). Η πρόσβαση είναι περιορισμένη (Restricted Access) και παρέχεται αποκλειστικά σε μέλη της ακαδημαϊκής κοινότητας για σκοπούς επαλήθευσης των αποτελεσμάτων της διατριβής και περαιτέρω μη εμπορική έρευνα. Η πηγή προέλευσης δικαιούται να ζητήσει οποιαδήποτε τροποποιητική ενέργεια (π.χ. αφαίρεση) επί του πρωτότυπου περιεχομένου. Τέλος, η άδεια CC BY-NC-ND 4.0 ισχύει για την επιμέλεια (curation), τα μεταδεδομένα και τη γλωσσολογική επισημείωση του σώματος κειμένων. Βιβλιογραφικές αναφορές Αφεντουλίδου, Β. (2022). Σώμα ελληνικών κειμένων για τη μελέτη δομών ομοιότητας της Νέας Ελληνικής: σχεδιασμός και υλοποίηση. Στο Πρακτικά του 10ου Συνεδρίου Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας (σσ. 67-94). ΕΚΠΑ. Αφεντουλίδου, B. (2021). Δομές ομοιότητας στη Νέα Ελληνική. Σωματοκειμενικές παρατηρήσεις για τον πολυλειτουργικό δείκτη σαν. Προφορική ανακοίνωση στην 41η Ετήσια Συνάντηση του Τομέα Γλωσσολογίας, 13–15 Μαΐου 2021. ΑΠΘ. Αφεντουλίδου, Β. (2013). Και σου απάντησα κάτι σαν ‘τέλεια, εντάξει’. Δείκτης σαν + ευθύς λόγος;. Προφορική ανακοίνωση στο 7ο Συνέδριο Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας, 16–18 Μαΐου. ΕΚΠΑ. Αφεντουλίδου, Β. (2012). Συγκρίσεις ομοιότητας στα Νέα Ελληνικά: ο δείκτης σαν. Στο Z. Gavriilidou, A. Efthymiou, E. Thomadaki & P. Kambakis-Vougiouklis (Επιμ.), Selected papers of the 10th International Conference on Greek Linguistics (σσ. 696-707). DUTH. Afentoulidou, V. (2009). Sketching the σαν conditional construction in Modern Greek. Submitted essay, 2009 Linguistic Institute, Linguistic Structure and Language Ecologies, Linguistic Society of America and UC Berkeley. Gries, Stefan Th. 2024. Coll.analysis 4.1. A script for R to compute perform collostructional analyses. https://www.stgries.info/teaching/groningen/index.html Institute for Language and Speech Processing - Athena Research Center (2015). ILSP Lemmatizer. Version 1. [Software (Tool/Service)]. CLARIN:EL. http://hdl.handle.net/11500/ATHENA-0000-0000-23EE-D Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 101–108). Online: Association for Computational Linguistics.
The Orthographic Junctions of English A Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and ConsonantVowel Information Asymmetry (p. 1) Boicho Dimitrov Temelakiev Saxon Ventura Research Ltd 28th of May, 2026 CC BY Abstract This paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling (p. 1). While the distinct roles of consonants and vowels in language processing are well-recognized psycholinguistically (p. 5), and the dual-stratum organization of English morphophonology is established theoretically (p. 6), this study provides an original, datadriven quantification of these properties directly at the orthographic level. The morpheme boundary—the orthographic junction between a stem and an affix—is treated as the primary object of measurement, reading the distribution of boundary characters across the corpus (p. 1). Three results are established: 1. 2. 3. The junction carries a stable, structured filter: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a distributional sparsity that scales inversely with an affix’s productivity (pp. 1-2). Affix relationships are structural: The relationship between any two affixes is quantified by the correlation of their boundary distributions, measuring their shared stem population (\(r \approx 0.99\) for etymological doublets down to \(r = 0.57\) for productivity-asymmetric near-twins) (pp. 1, 4). A massive information asymmetry partitions the lexicon: The written word decomposes into a invariant consonant skeleton carrying lexical identity (53.0% unique recoverability) and a mobile vowel tissue carrying grammatical form (1.8% unique recoverability) (pp. 1, 5). Multiple independent measures—consonant recoverability, derivational class-marking, and freestem fraction—converge on a single partition separating a transparent Germanic core from a bound Latinate superstructure (pp. 1, 6). The method, its corrections, and its limits are reported in full (p. 1). 1. Introduction and Method Traditional models of English morphophonology have long recognized that the lexicon is organized into distinct, historical strata—principally a native Germanic core and a bound Latinate superstructure (pp. 1, 6). Classic frameworks in generative phonology and lexical morphology, such as those pioneered by Chomsky and Halle (1968) and expanded by Kiparsky (1982), demonstrate that affixes of differing origins impose strict constraints on the phonetic and structural traits of the stems they recruit. Concurrently, cognitive and psycholinguistic research 1 has established a foundational "consonant-vowel functional asymmetry," demonstrating that human language processing systematically relies on consonants to preserve lexical and lexical-root identity, while vowels are dynamically manipulated to signal grammatical operations (Nespor et al., 2003; Bonatti et al., 2005). While these qualitative boundaries and cognitive patterns are deeply documented, this paper presents a mean-free, data-driven methodology to extract, quantify, and map these structural phenomena directly from corporate-scale English orthography without relying on heavy linguistic machinery. We introduce The Orthographic Junctions of English (OJE), an empirical approach that frames the morpheme boundary—the exact character interface between a stem and an affix—as an informational filter whose statistical properties reveal the historical, structural, and cognitive divisions of the vocabulary. All results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution (p. 1). The corpus utilized is the opensource dwyl/english-words repository (words_alpha.txt), comprising 455,246 alphabetic entries with a total of 4,254,354 letter occurrences and a mean word length of 9.345 letters (p. 1). Each letter is assigned its ordinal value (\(A=1\) through \(Z=26\)) (p. 1). Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter, so that each suffix position returns its exact ordinal value as a safety check (p. 1); the first stem letter preceding the suffix—the linker—is then read as a full A–Z frequency distribution rather than as a mean (p. 1). The governing methodological constraint is that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, and that every claim is backed by a precise empirical count (p. 1). By avoiding any dependency on complex machine-learning libraries or external lexical databases, the framework ensures that every architectural pattern discovered can be verified using standard data operations. 2. The Orthographic Junction and Its Backbone 2 The initial phase of this investigation examined unconditioned letter bigrams across the corpus, which yielded no meaningful morphological signal. Structural regularities appeared only when character distributions were explicitly conditioned on a single morpheme boundary—the character interface linking a stem to an affix. This structural conditioning serves as the foundation of the OJE framework. A preliminary tabulation of adjacent letter pairs across the corpus—measuring which letters follow which, without regard to structural position within the word—yielded baseline frequency patterns. These patterns are entirely reducible to general English orthographic constraints and carry no isolable morphological content. The structural signal emerged only when a specific suffix was fixed and the preceding characters were read as a discrete population. Conditioning on the junction, rather than measuring adjacency as a flat sequence, renders the underlying boundary constraints visible. The set of characters that legally occupy the stem side of a morpheme boundary proves narrow, highly structured, and remarkably stable across suffixes. Reading the linker distribution across the mapped suffix inventory reveals a highly stratified, three-tier architectural filter: A structural backbone of six letters—T, L, N, R, S, and I—is admitted at high frequency by nearly every suffix in the English lexicon. Within this backbone, T serves as the single most frequent linker across the inventory and recurs as the dominant boundary letter across independent suffixes. Conversely, an absolute floor of two letters—J and Q—is admitted by no suffix at a measurable frequency. This absolute prohibition is confirmed corpus-wide and is statistically consistent with their status as the two rarest letters in English orthography overall (with J accounting for 0.18% and Q for 0.19% of all letter occurrences). Between the backbone and the floor lies a selective middle whose specific character composition varies dynamically by affix, providing the distinct orthographic footprint wherein an individual suffix’s identity resides...
This article examines the key challenges and development directions of modern linguistics. Focus is placed on the processes of digital language transformation, the influence of the internet on linguistic norms, and the cognitive and sociolinguistic aspects of communication. The author emphasizes the need for an interdisciplinary approach to the study of linguistic phenomena in the global information space.
While the linguistic shifts between Ottoman and modern Turkish are well-documented qualitatively, quantitative analyses remain scarce.This study addresses this by conducting a comparative computational analysis using two Universal Dependencies treebanks: OTA-DUDU for Ottoman Turkish and TR-BOUN for modern Turkish.By employing descriptive statistics and a log-likelihood ratio test, we demonstrate the change and quantify the magnitude of diachronic variation.The analysis yields three primary statistical findings.First, our data reveals a 77% compliance rate with labial vowel harmony for suffixes, while this value is 98% in modern Turkish.This discrepancy can be explained by the presence of rounding in Ottoman Turkish, which disappears in modern Turkish.On the other hand, the compliance rate of palatal vowel harmony is quite high for both languages, 96% for Ottoman Turkish and 99% for modern Turkish.Second, some suffixes, such as the converb -(y)Ip 1 and the dative infinitive -mAyA, changed by reducing their allomorphs in modern Turkish.Third, we demonstrate that Arabic and Persian pluralization rules, which constituted 28% of plural nouns in Ottoman Turkish, lost their pluralizing function in modern Turkish, although the words remain with singular meaning.
This paper presents the steps taken to integrate data from the UD_Latin-PROIEL treebank into the LiLa Knowledge Base of interoperable linguistic resources for Latin.It describes how the lexical, morphological, syntactic, and citation information from the source was modeled using the Linked Open Data principles as adopted by the LiLa Knowledge Base.The process of linking tokens to the LiLa collection of Latin lemmas is detailed, addressing challenges such as ambiguities, new lemmas, and errors encountered in the source.The outcome is a syntactically annotated textual resource that is interoperable with the (meta)data of other Latin linguistic resources linked within the LiLa Knowledge Base.This integration enables new ways of analyzing linguistic information and using the content as a starting point to explore connections with other interlinked resources.A use case demonstrates this interoperability.
Using semantic dependency analysis, this study examines narrative productions from Mandarin-speaking preschool children aged three to six to investigate how semantic organization develops with age in early childhood. Four semantic dependency treebanks were constructed from a Chinese narrative corpus available in the CHILDES database. By comparing semantic dependency types and semantic dependency distances across the four age groups, we found that (1) semantic organization shifted from experiencer and classification relations toward agent and patient relations, situational-role relations (particularly those involving measurement, individuation, and direction), and structural relations; (2) mean semantic dependency distance (MSDD) increased with age, as adjacent dependencies decreased and longer dependencies became more frequent. This increase in MSDD indicates growing complexity in semantic organization and is largely driven by significant increases in the MSDD values of specific semantic dependency types. These findings provide new evidence for semantic organization development in preschool children.
INTRODUCTION: The experience of emotions is accompanied by distinct bodily sensations, consistent across cultures. Irritable bowel syndrome is characterized by altered interoceptive and affective processing, suggesting that individuals with IBS may experience emotions differently in their bodies. This study investigated whether so-called bodily maps of emotions differ between individuals with IBS and healthy controls (HC). METHODS: Forty-three individuals with IBS (Rome IV) and 54 HC used the topographical mapping tool EmBODY to color bodily silhouettes marking where sensations were perceived during 13 emotions and a neutral negative affective state. Region-based (abdomen, head, thorax) and whole-body pixel-wise analyses were performed on the resulting body maps to compare emotion-evoked sensations between individuals with IBS and HC using non-parametric one-way ANOVA. RESULTS: IBS participants showed consistently elevated abdominal sensations across emotional states, and emotional state only modulated abdominal sensations in HC (p < 0.001) but not in IBS (p = 0.17). After correcting for neutral activation, positive emotions (love, happiness) elicited smaller increases in abdominal activation in IBS than in HC (p < 0.03). IBS participants did not show greater abdominal activation during negative emotions relative to HC. Several positive emotions were associated with reduced head-region activation in IBS (p < 0.041), while no group differences emerged in the thorax region. CONCLUSION: Individuals with IBS demonstrate altered embodiment of emotional states, characterized by persistent abdominal sensations, limited emotional differentiation, and blunted emotion-specific modulation for positive emotions. Future studies should incorporate concurrent affect ratings and physiological measures to clarify emotion-body interactions in IBS.
This paper presents a novel treebank-driven approach to comparing syntactic structures in speech and writing using dependency-parsed corpora. Adopting a fully inductive, bottom-up method, we define syntactic structures as delexicalized dependency (sub)trees and extract them from spoken and written Universal Dependencies (UD) treebanks in two syntactically distinct languages, English and Slovenian. For each corpus, we analyze the size, diversity, and distribution of syntactic inventories, their overlap across modalities, and the structures most characteristic of speech. Results show that, across both languages, spoken corpora contain fewer and less diverse syntactic structures than their written counterparts, with consistent cross-linguistic preferences for certain structural types across modalities. Strikingly, the overlap between spoken and written syntactic inventories is very limited: most structures attested in speech do not occur in writing, pointing to modality-specific preferences in syntactic organization that reflect the distinct demands of real-time interaction and elaborated writing. This contrast is further supported by a keyness analysis of the most frequent speech-specific structures, which highlights patterns associated with interactivity, context-grounding, and economy of expression. We argue that this scalable, language-independent framework offers a useful general method for systematically studying syntactic variation across corpora, laying the groundwork for more comprehensive data-driven theories of grammar in use.
Abstract Quranic Arabic has motivated sustained morphological and syntactic annotation, yet Quranic treebanks remain hard to compare and reuse in modern natural language processing (NLP) because they diverge in clitic segmentation, feature inventories, and syntactic formalisms. We present UD-Quran, a Universal Dependencies (UD) v2 conversion of the Extended Quranic Treebank with hybrid syntactic annotations (EQTB). The conversion treats EQTB morpho-syntactic segments as UD tokens, maps EQTB part-of-speech (POS) categories to 12 UD universal part-of-speech (UPOS) tags, derives UD features from explicit EQTB columns, and collapses EQTB dependency labels into a compact UD relation inventory with deterministic normalization aligned to UD content-head conventions. Two releases are provided: a surface variant aligned to the observable Quranic string by excluding analytically inserted nodes, and an augmented variant that retains inserted material to preserve EQTB’s modeling of ellipsis and implied pronominals. The surface release contains 11,693 sentences and 128,219 UD tokens; the augmented release contains 139,376 tokens. Conversion coverage is quantified by restricting unspecified dependency (dep) to 1,129 tokens (0.881% of surface tokens). UD-Quran includes fixed training/development/test (train/dev/test) splits (seed 42) and lightweight Stanza baselines scored with the CoNLL (Conference on Computational Natural Language Learning) 2018 UD evaluation script. On the test sets, parsing with gold tags reaches labeled attachment score (LAS) 80.66 (surface) and 82.47 (augmented), while the end-to-end pipeline reaches LAS 64.52 and 68.54. UD-Quran is intended as an interoperability layer that supports standard UD tooling while preserving sentence-level traceability to EQTB.
This article evaluates the integration of data extracted from a French syntactic lexicon, the Lexicon-Grammar (Gross, 1994), into a probabilistic parser. We show that by applying clustering methods on verbs of the French Treebank (Abeillé et al., 2003), we obtain accurate performances on French with a parser based on a Probabilistic Context-Free Grammar (Petrov et al., 2006).
Background and Aims: Most men consume pornography, with a small but significant percentage losing control over their use. Since ICD-11, problematic pornography use can be diagnosed as "compulsive sexual behavior disorder." Debate persists on whether problematic pornography use is an impulse-control disorder or a behavioral addiction. Mechanisms of learning and memory play a central role in addictive disorders but are presumably less relevant for impulse control disorders. Methods: One hundred thirty-nine heterosexual male users of pornography and gaming participated in our study which was part of a multi-center research project on internet use disorders in Germany. We focus on a subsample of fifty-eight non-problematic (n = 35) and problematic pornography users (n = 23, labeled pathological). FMRI data were collected during appetitive conditioning, extinction and recall. Pornographic, game, and money images served as unconditioned stimuli, geometric shapes as conditioned stimuli (CS). Results: During appetitive conditioning pathological pornography users showed a generally stronger response in ventral striatum to all CSs, whereas altered activations in extinction and recall were specific to the porn-associated CS. Greater activations in the dorsal anterior cingulate cortex during extinction and in the medial orbitofrontal cortex during recall suggest persistence of appetitive memory for pornography in pathological users, supported by valence ratings and skin conductance responses (SCR). Sensitization to the monetary cue also emerged in SCR. Discussion and Conclusions: Based on these new neurobiological findings, which are consistent with current addiction theories about stimulus-specific altered reward sensitivity and appetitive memory, we argue that problematic pornography use should be considered a behavioral addiction.
Monolingualism, native-speakerism and standard language ideology have been identified as dominant ideologies in language teaching with severe effects on second language teacher identities. Such ideologies offer alleged certainties but also detach teachers from the actual uses and value of language in multilingual and multidialectal contexts. As a consequence, educators might feel constrained by rigid linguistic norms, hindering their capacity to re-evaluate their approaches to accommodate the diverse linguistic realities and communicative needs of English learners in an increasingly interconnected and multilingual global landscape. This chapter intends first to offer a broad perspective of how these ideologies have shaped language teaching and how they have clashed with research-based observations of multilingual and multidialectal communicative settings. We will give an overview of the relevant literature, ranging from foundational texts to more recent ones challenging the ‘ideal’ monolingual native speaker, and we will show how the above ideologies are still found in a rather pervasive way in the language teacher profession. This will be followed by an account of recent research conducted in teacher training environments aimed at showing ways to successfully gear future language teachers towards a new vision of language that contemplates diversity and hybridization as fundamental pillars on which teachers’ identities need to be based.
Facial expressions are powerful signals of human emotion, shaping both human–human and human–computer interaction. As interactive technologies, from adaptive interfaces to emotion-aware agents, become more pervasive, systems are increasingly expected to recognize and respond to users’ emotions naturally. But what if a system misreads your face? Such misinterpretation is particularly likely when cultural differences in emotion perception are overlooked. This problem may be compounded by the fact that most facial emotion recognition (FER) models are trained on datasets that reflect the norms of a particular cultural group that assume universality, limiting their reliability in multicultural contexts. Surprise, in particular, is an emotion whose valence can be either positive or negative depending on context, making it a critical case for investigating cultural bias in FER. To address this, we examined how cultural background shapes the recognition and valence interpretation of surprise facial expressions among South Korean (N=36) and American (N=34) participants. Participants labeled 200 facial expressions (surprise and fear), rated their perceived valence, and described personal experiences of surprise. Results show that South Korean-labeled surprise expressions exhibited stronger negative Action Unit (AU) activation and lower valence ratings, whereas American-labeled ones showed more balanced or positive facial cues. Qualitative accounts further revealed that South Koreans framed surprise as tense or socially cautious, while Americans viewed it as open and situationally flexible. These findings bridge recognition and interpretation in cross-cultural emotion research and highlight the need for culturally adaptive FER systems that can interpret ambiguous emotions like surprise more inclusively.
Screen use pervades daily life, shaping work, leisure, and social connections while raising concerns for digital wellbeing. Yet, reducing screen time alone risks oversimplifying technology’s role and neglecting its potential for meaningful engagement. We posit self-awareness—reflecting on one’s digital behavior—as a critical pathway to digital wellbeing. We developed WellScreen, a lightweight probe that scaffolds daily reflection by asking people to estimate and report smartphone use. In a two-week deployment with college students (\(\mathtt {N}\)=25) focused on generating formative insights, we examined how discrepancies between estimated and actual usage shaped digital awareness and wellbeing. Participants often underestimated productivity and social media while overestimating entertainment app use. They showed a 10% improvement in positive affect, rating WellScreen as moderately useful. Interviews revealed that structured reflection supported recognition of patterns, adjustment of expectations, and more intentional engagement with technology. Our findings highlight the promise of lightweight reflective interventions for supporting self-awareness and intentional digital engagement, offering implications for designing digital wellbeing tools.
BACKGROUND: Schools have the potential to promote equitable health from early life onwards yet require sufficient organizational capacity to achieve sustained action. Structured improvement approaches, such as PDSA cycles, may help strengthen this capacity by guiding systematic implementation processes. However, their potential in school health promotion remains insufficiently understood, particularly regarding the heterogeneous contextual factors shaping their application. This study examined which contextual determinants shape schools' perceived implementability of the PDSA cycle for health promotion and how these conditions differ across schools. METHODS: Nine German primary schools participating in a holistic health promotion program were purposively sampled to capture heterogeneity across federal states, socioeconomic contexts, and urban-rural settings. Semi-structured qualitative group interviews in a workshop format were conducted with school principals, teachers, and parents and analyzed using the framework method guided by the CFIR. To facilitate cross-case comparison, color-coded valence ratings (facilitator/barrier/mixed) were visualized in a Matrix Heat Map, enabling identification of contextual tendencies. RESULTS: Fifteen contextual factors emerged across the CFIR domains of Outer Setting, Inner Setting, and Individual. Schools with prior experience using structured processes similar to PDSA cycles reported more facilitators, such as established communication structures, while schools without such experience perceived more barriers, notably financial constraints. Common barriers across schools included limited parental engagement and staff shortages, whereas leadership support and compatibility of program components were consistent facilitators. Some factors interacted dynamically, with resource constraints reinforcing other barriers or with strong mission alignment amplifying engagement. CONCLUSION: Schools' prior structured experience seemed to be associated with how they perceived the implementability of PDSA cycles for health promotion implementation, with more experienced schools anticipating more facilitators and fewer barriers. While causality cannot be inferred, these exploratory findings are hypothesis-generating and suggest that prior structured experience may be an important factor to consider for tailoring implementation support and building organizational capacity. Beyond these insights, extending the framework method with a color-coded Matrix Heat Map proved valuable for visualizing contextual heterogeneity and revealing tendencies across cases. This combined approach may inspire further research on how contextual configurations shape the use of structured processes in complex, multi-site implementation settings.
While the influence of state-dependent factors on appetitive processing has received considerable attention, the role of stable personality traits remains comparatively unexplored. Extraversion, characterized by heightened positive emotionality, represents a compelling candidate in this regard, as it may shape individual differences in Positive Valence System (PVS) functioning. The present study examined how extraversion modulates neural and subjective responses to pleasant stimuli; the role of neuroticism was additionally explored, given its established association with affective reactivity. Sixty-eight Italian university students (40 females) completed an online version of the Big Five Inventory (BFI-44) before the laboratory session. Then, participants completed a passive viewing task of pleasant and neutral images while undergoing an electroencephalographic (EEG) recording. Appetitive stimulus processing was indexed by the peak amplitude of the P300-LPP complex and subjective SAM ratings. Results revealed that extraversion was positively associated with larger P300-LPP complex amplitudes to pleasant relative to neutral stimuli and with higher arousal ratings across both emotional categories. Additionally, neuroticism was associated with lower valence ratings regardless of stimulus category, with no significant effect on neural responses to emotional stimuli. These findings highlight extraversion as a stable personality trait shaping PVS functioning. Specifically, low extraversion was associated with reduced P300–LPP amplitudes and lower arousal ratings to pleasant stimuli, paralleling neural patterns documented in psychopathological conditions involving blunted PVS activation. These results underscore the utility of ERP-based measures in capturing personality-related differences in appetitive processing relevant to psychopathology risk.
Individuals with hearing loss, even when using hearing aids, often perceive pleasant environmental sounds as less pleasant than do those with normal hearing. This bias in emotional response may negatively impact well-being, leading to decreased social participation and increased loneliness. The present study examined whether the Positive Focus intervention-encouraging hearing aid users to focus on positive listening experiences-could influence emotional response to environmental sounds. Thirty participants were randomly assigned to either a Positive Focus or a Control group. At the initial laboratory visit, all participants were fitted with study hearing aids and performed affective ratings of 120 environmental sounds. Over 3 weeks, both groups wore the hearing aids; the Positive Focus group additionally reported daily positive listening experiences via a text message. At the end of the three-week period, participants completed questionnaires on hearing aid outcomes and repeated the affective ratings. The Positive Focus intervention did not alter emotional responses to environmental sounds in a laboratory setting. However, regression analyses revealed that valence ratings of typically pleasant sounds moderated the effectiveness of Positive Focus on hearing aid benefit; the intervention was more effective for individuals less naturally inclined to respond positively to such sounds. These findings suggest that valence screening may help identify individuals most likely to benefit from Positive Focus, supporting more personalized hearing care strategies.
By studying how individuals in an "at-risk" state of psychosis learn about threat and safety cues - specifically, how they develop and unlearn fear responses to neutral cues - we might better understand the mechanisms leading to heightened arousal and fear that are characteristic of acute psychotic episodes. At-risk individuals (N = 88; of which 28 fulfilled ultra-high-risk criteria on the Comprehensive Assessment of At-Risk Mental States interview and 60 scored above a predefined threshold on the Community Assessment of Psychic Experiences) and healthy controls (N = 44) underwent a standardized and validated differential fear conditioning paradigm including an acquisition, generalization, and extinction phase. The main outcomes of interest were the late positive potential, fear-potentiated startle, and self-reported ratings of valence, arousal, fear, and expectancy elicited by the conditioned stimuli (CS). The at-risk group exhibited diminished fear learning, evident in significantly reduced differentiation between the CS+ vs. CS- in the valence ratings compared to controls. Additionally, they demonstrated impaired fear extinction, evident in valence and arousal ratings, in which their CS differentiation showed a slower reduction than the controls. There were no group differences in late positive potential responses. At risk mental states appear to be associated with problems in distinguishing dangerous from safe stimuli and a diminished ability to adjust affective responses to conditioned stimuli based on new information, while the late-positive potential and fear-potentiated startle are unaltered. Early interventions could focus on recalibrating subjective emotional evaluations of fear-associated events.
Background: By studying how individuals in an "at-risk" state of psychosis learn about threat and safety cues – specifically, how they develop and unlearn fear responses to neutral cues - we might better understand the mechanisms leading to heightened arousal and fear that are characteristic of acute psychotic episodes. Methods: At-risk individuals (N = 88; of which 28 fulfilled ultra-high-risk criteria on the Comprehensive Assessment of At-Risk Mental States interview and 60 scored above a predefined threshold in the Community assessment of Psychic Experiences questionnaire) and healthy controls (N = 44) underwent a standardized and validated differential fear conditioning paradigm including an acquisition, generalization, and extinction phase. The main outcomes of interest were the late positive potential, fear-potentiated startle, and self-reported ratings of valence, arousal, fear, and expectancy elicited by the conditioned stimuli (CS). Results: The at-risk group exhibited diminished fear learning, evident in significantly reduced differentiation between the CS+ vs. CS- in the valence ratings, compared to controls. Additionally, they demonstrated impaired fear extinction, evident in valence and arousal ratings, in which their CS differentiation showed a slower reduction than the controls. There were no group differences in late positive potential responses. Conclusion: At risk mental states appear to be associated with problems in distinguishing dangerous from safe stimuli and a diminished ability to adjust affective responses to conditioned stimuli based on new information, while the late-positive potential and fear-potentiated startle are unaltered. Early interventions could focus on recalibrating subjective emotional evaluations of fear-associated events.
This study investigates the influence of three biophilic interior design variables: natural light, interior vegetation (vertical green wall), and biomorphic form (biomorphic wall panel) on affective and physiological responses in a design studio interior utilizing immersive virtual reality (IVR) and wearable biofeedback technology. This study was a within-participant 23 factorial design that included one baseline and eight IVR studio conditions. Participants experienced all conditions while reporting affects using the Self-Assessment Manikin (SAM) valence and arousal scales, electrodermal activity (EDA), and skin temperature (ST). Cybersickness was measured with the Simulator Sickness Questionnaire (SSQ) and presence was assessed using the Igroup Presence Questionnaire and Slater-Usoh-Steed presence measures (IPQ, SUS), while baseline anxiety (STAI) was controlled. The results demonstrated a significant primary influence of natural light on SAM valence ratings: conditions with natural light were evaluated as more pleasant than the non-variable and baseline condition, whereas interior vegetation and biomorphic form had smaller, context-dependent effects that were most evident when layered with natural light. Differences in SAM arousal ratings were modest and non-systematic. EDA did not differentiate, and ST showed only small shifts, indicating that during calm exploratory monitoring, subjective affect was more responsive. The circumplex findings guided to an activity-specific zoned interior rather than a single uniform design studio.
This article examines the role and significance of language corpora and linguistic databases in linguistic expertise. It discusses the possibilities of conducting objective semantic, pragmatic, and stylistic analyses of texts through corpus-based methods. The study also highlights the contribution of linguistic databases and artificial intelligence technologies to improving the accuracy, reliability, and efficiency of expert conclusions. Furthermore, the relevance of developing specialized corpora and databases for forensic linguistics in Uzbekistan is substantiated.
We describe THIVLVC, a two-stage system for the EvaLatin 2026 Dependency Parsing task. Given a Latin sentence, we retrieve structurally similar entries from the CIRCSE treebank using sentence length and POS n-gram similarity, then prompt a large language model to refine the baseline parse from UDPipe using the retrieved examples and UD annotation guidelines. We submit two configurations: one without retrieval and one with retrieval (RAG). On poetry (Seneca), THIVLVC improves CLAS by +17 points over the UDPipe baseline; on prose (Thomas Aquinas), the gain is +1.5 CLAS. A double-blind error analysis of 300 divergences between our system and the gold standard reveals that, among unanimous annotator decisions, 53.3% favour THIVLVC, showing annotation inconsistencies both within and across treebanks.
Despite growing understanding of the ways in which sexual disgust operates, significant gaps remain, particularly around the prevalence of participation in sexual behaviours perceived to be disgusting, how this differs by type of behaviour and involvement of fluids (e.g., blood, semen) and barriers (e.g., condoms), as well as factors surrounding participation in these behaviours. Furthermore, no research has examined the direct relationship between disgust and arousal ratings of the same behaviour, how this differs by gender, and how this relationship relates to past participation in and desire to avoid a certain behaviour. The proposed project aims to address these gaps, which will advance knowledge about the prevalence of suboptimal sexual experiences and potential failures of the sexual inhibition and excitation system. The primary research questions are: (1.1 & 1.2) To what extent does past participation converge or diverge from behaviours that individuals indicate they would want to avoid? Does the rate of overlap differ for certain behaviours as compared to others? (2.1 & 2.2) How do the disgust and arousal ratings of a behaviour relate to one another? Does this differ for men and women? (3) How do the disgust and arousal ratings of a behaviour predict past participation or a desire to avoid? And (4.1 & 4.2) What contexts and motivations surround sexual experiences that are perceived to be disgusting? Do people have different experiences for different sexual behaviours?
Existing neurocognitive reading models highlight a left-lateralized brain network supporting word- to discourse-level processing, but they largely overlook emotion. Although emotion-related brain regions are active during discourse processing, the role of arousal (i.e., emotional intensity) remains underexplored. Prior neuroimaging work has shown that isolated words or whole passages varying in arousal evoke activity in brain regions associated with emotion and situation model processing. However, how arousal at the phrase level within passages may modulate neural activity is unclear, nor have studies investigated how individual differences in arousal responsiveness may be linked to reading comprehension ability, particularly in developing readers. Here, we used functional magnetic resonance imaging (fMRI) to examine the neural correlates of lexical arousal in 86 third-graders as they read passages. A parametric modulation analysis, using phrase-level arousal ratings from a validated lexical database, was used to investigate how fluctuations in arousal during passage reading correlated with neural activity. We demonstrated that phrase-level arousal was associated with increased activity in regions implicated in emotional processing and situation model construction: the right amygdala, striatum, and posterior insula, and the left dorsomedial prefrontal cortex (dmPFC). Additionally, dmPFC activity was associated with better reading comprehension ability, aligning with prior literature linking dmPFC to situation model building. This work highlights the importance of integrating lexical emotional dimensions into cognitive models of reading and supports the idea of using emotionally engaging materials to enhance comprehension for developing readers of all abilities.
This page contains behavioral data of an encoding and a temporal memory task in two experiments. Experiment 2 also includes an emotional valence rating that was conducted at the end of the experiment. For information about the study, please see the published manuscript in Psychological Research.