Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
LATOC Corpus LATOC (Latin-transliterated Ottoman Turkish Corpus) includes 143 Ottoman Turkish books, 13,252,350 words, written between the 15th and 20th centuries. The books were transliterated by domain experts and publicly shared on the Internet. The books in the corpus were automatically structured via a rule‑based approach and manually checked. Due to the copyright restrictions, this repository does not have any raw text; however, it guides you to download the files, convert them into structured XML files, and process them on your computer to get LATOC. The pipeline provided here can be extended to new sources by simply acquiring the PDF documents and, if necessary, updating the `CERTAIN_RULES` variable in 2_preprocessing.py and the files `controversy_cache.json` and `exclude_pages.txt`. Corpus Overview The corpus has more than 13 million words from 143 works written between the 15th and 20th centuries. While the pipeline standardizes these texts into the IJMES format, some inconsistencies, such as normalization of spelling, may persist. These are primarily inherited from the transliteration provided by the domain experts. This work does not apply any normalization to the spelling. Arabic/Persian characters in the documents are filtered out, and only the Latin-transliterated text is preserved. Each document is split into pages. Each page is divided into three segments: paragraph, comprising the main text, title, and footnote. The final XML files also provide the coordinates of the regions on the page, like "bbox="277.0,711.1,651.2,765.0". Data Files Work‑level metadata (`LATOC_metadata_sample.csv`) This is a sample of the metadata with further information. It provides metadata for 36 Dîvân works. Note that this metadata is from the previous version of LATOC, which is the reason why it has only 36 works. In the future, this scheme will be expanded to all works in LATOC. Each Dîvân work is accompanied by: - `file_name` - `work_name` (title of the Dîvân) - `pen_name` (mahlas) - `real_name` - `viaf` - `century` - `gender` - `rank` - e.g., “Sultan,” “Judiciary & Religious Office,” “High Bureaucracy/Military,” “Scholars & Sufi Orders,” “Civil Bureaucracy,” “Lay/Non‑official” Overall data statistics (`data_statistics.csv`) This file includes basic statistics such as the word count per document. It also provides links to some of the documents you can download for the work. Note that some works miss the URL links to download documents here. The document names in the column 'file' should be enough for users to find the document. Book‑level data Since the data was under copyright, this repository does not have it directly. However, you can download the data from _Yazma Eserler via either here or this webpage and then run the Python scripts as explained in this document to have the processed data on your device. Supplementary material (`controversy_cache.json` and `exclude_pages.txt`) `controversy_cache.json` includes the conversion rules for the problematic characters, which might be converted into more than one character in the IJMES chart. Since each document behaves differently, it provides the conversion rule based on the document. You can add a new rule here for your new documents. `exclude_pages.txt` has the page boundaries that should be deleted to remove the editorial preface, table of contents, and references, etc. You can enlarge this file if you add a new source to the data. Processing the data with Python scripts After downloading the files and storing them in a single folder, you should run the Python scripts 1_pdf_extractor.py, 2_preprocessing.py, and 3_xml_cleaner.py in turn. 2_preprocessing.py requires the supplementary file, `controversy_cache.json`. For 3_xml_cleaner, you need the supplementary document exclude_pages.txt. These files are prepared for 144 works presented in this dataset by the author. Usage Notes - The corpus can be utilized for **diachronic studies**; Yılandiloğlu (2025) demonstrated that poets adhered more accurately to the aruz meter over the centuries, reflected in rising conformity rates. - The sample metadata allows you to focus on specific ranks (e.g., “Sultan”) or gender. - Current work is focused on standardizing transliteration to the IJMES system and expanding the corpus further. Impact and Downstream Tasks This corpus was specifically curated and structured to support the development of Ottoman Turkish NLP resources. It has been used for: * **Large Language Models:** The structured data was used to train models, including: * Masked language model: [ota-roberta-base](https://huggingface.co/enesyila/ota-roberta-base) * State-of-the-art Named Entity Recognition model for Ottoman Turkish: [ota-roberta-base-ner](https://huggingface.co/enesyila/ota-roberta-base-ner) * A Universal Dependencies (UD) parser that tags with 91% accuracy and lemmatizes with 86% accuracy: [ota-ud-style](https://huggingface.co/enesyila/ota-ud-style) * **Annotated Treebank:** The dataset serves as the basis for [UD_Ottoman_Turkish-DUDU](https://github.com/UniversalDependencies/UD_Ottoman_Turkish-DUDU), currently the largest Ottoman Turkish corpus in the Universal Dependencies.
Le French Treebank: une ressource lexicale et syntaxique richement annotée (et validée manuellement) pour les linguistes, utilisable en TAL, dans sa version 2.0 Projet initié en 1997, avec le soutien de l'IUF, du CNRS et du CNRTL21 550 phrases (environ 664 500 tokens) du journal Le Monde (1990-1993)Métadonnées: auteur, date, domaine (par article)Annotations lexicales (catégories, sous-catégories, flexion, mots composés avec composants) et syntaxiques (constituants majeurs, fonctions grammaticales) validéesPlusieurs formats disponibles: XML, PTB, CoNLL, codage UTF-8 (ligature œ notée oe)Nouveautés de la version 2.0: plusieurs erreurs d'annotation ont été corrigées depuis la publication en 2016 de la version 1.0
Cross-modal perception, the integration of information from multiple senses, plays a critical role in shaping emotional experiences. This study examines the interactions between visual and olfactory stimuli and their effects on emotional responses, a topic rarely addressed in prior research. Experiments employed five distinct visual stimulation methods that were combined with olfactory stimuli. Participants' emotional responses were assessed via surveys and electroencephalography (EEG) signal analysis. The study varied the color and movement direction of augmented particles to investigate their impact on EEG signals and emotional states. The findings demonstrated significant differences in emotional state classification under the influence of visual-olfactory interactions. Specifically, with backward-moving particles with matching colors (M4), classification accuracy was comparable to that of unimodal olfactory conditions (M1). Other visual stimuli generally caused confusion in classifying emotional responses. The increased valence ratings for pleasant aromas across all visual conditions did not consistently align with EEG-based classification results, suggesting that visual stimuli may introduce complexities into neural signals. These results highlight the intricate dynamics of multisensory interactions, emphasizing the role of visual stimuli in modulating emotional responses. The findings also suggest the potential of visual-olfactory interactions in developing augmented reality (AR) systems. By aligning visual and olfactory cues, AR environments can enhance the user experience and create immersive emotional landscapes, leading to applications for mood modulation and stress relief. This study underscores the relevance of multisensory integration in advancing emotion analysis and affective computing.
The Ledger of Meluhha: Indus Valley Script as Metrological Accounting Code Rajeshkumar Venugopal (Third Buyer Advisory LLC, Michigan; ORCID 0009-0002-1838-5976). Version 3.0, 3 May 2026. BSD-2-Clause for human use; AI ingestion / training / fine-tuning / RAG / inference prohibited under contract law (see ai.txt and the §For Journalists appendix in the book). This work argues that the Indus Valley script is a cargo-tag accounting system rather than a phonetic writing system. The five-field record schema (merchant mark, commodity, weight tier, quantity, route terminal) is recoverable from existing archaeological evidence: the Harappan binary-and-decimal weight series standardised to 0.5 percent precision across roughly one million square kilometres, the Akkadian cuneiform Meluhha import receipts from Ur, the morphological-parallel correspondences between Indus seals and Tamil Nadu Iron Age potsherds, and the bigram structure of mapped versus unmapped signs in the digitised CISI corpus. The hypothesis is strictly weaker than any phonetic decipherment: it does not claim the Indus people did not have language, and does not assert which language was spoken. It claims that the function of the seals was inventory rather than speech encoding, and that the apparent untranslatability of the script reflects this functional fact rather than the absence of structure. The bridge to phonetic content, where it exists, runs through the proto-Dravidian numeral system reproduced from Wells 2015 Table 6.1 (after McAlpin 1981) — sign polyvalence is constrained by the morphology of numerals already in use, not by free phonetic association. The book is 76 pages, organised into 26 sections plus an appendix for journalists. The §For Journalists appendix provides a 10-minute verification protocol that requires only the SQLite command-line tool: any quantitative claim in the book is reproducible from the indus_corpus.db file in this archive by running a single SELECT statement against the named source_code. Every numeric claim in the book is traceable to a row in the database with explicit source attribution. The corpus database (indus_corpus.db) integrates ten primary sources: Mahadevan 1977 (concordance histogram); Joshi-Parpola 1987 Vol.1 Collections in India (the canonical photographic corpus, 862 pages OCR'd via Tesseract into 1399 unique artefact identifiers across six site prefixes); the mayig CISI digitisation (Mohenjo-daro subset, 179 inscriptions with 1003 sign occurrences); Wells 2015 The Archaeology and Epigraphy of Indus Writing (sign-role classifications including the canonical ICTM identification of signs 1, 2, 60; the Harappa volumetric system VI=40.4L through VIIIIIII=283L from Table 4.2; the proto-Dravidian numeral system from Table 6.1; sign 700 by NUM right-adjacency frequencies from Table 5.3); Fuls 2019 ICIT documentation PDF (16 sign-function codes including TMK / ITM / NUM / SYL plus 10 Wells-numbered ICIT inscriptions and the TMK by TMK adjacency matrix); Fuls 2022 Corpus of Indus Inscriptions metadata scaffold (73 sites with book-page anchors plus the 33-row artefact typology TAB / POT / SEAL / TAG and the corpus headline statistics 4660 artefacts / 5644 texts / 19831 sign occurrences); the Tamil Treebank logo-syllabic proxy corpus; the Rajan-Sivanantham Tamil Nadu inscribed-potsherd corpora (RS2025 and RS2026 for 13 sites). The codebook database (indus_codebook.db) contains the 28-entry sign-role mapping plus 8 commodities, 5 routes, 11 weight tiers, 6 quantity codes, 5 merchant marks, and 3 positional rules. The LSSC database (indus_lssc.db) contains the Latent Structural State Contraction analysis used to support the closure-versus-option sign classification. The interactive HTML dashboard (ledger-of-meluhha.html) is a single-file visualisation that renders the entire trade network on Leaflet plus 14 panels of derived data on the dropped indus_corpus.db file. It requires no server, no build step, and no installation — only a modern browser. The dashboard surfaces the original three-column display (decoded seals, trade-network map, frequency / weight / commodity charts) plus an Extended Data section with eleven panels covering the 10 sources, the 21 sign-function codes, the 7-row Harappa volumetric system, the 11-row proto-Dravidian numeral table, the 10 real Wells-numbered ICIT inscriptions, the 97 sign-role assignments, the 27 documented sign-pair frequencies, the 4 paradigmatic sign clusters, the CISI Vol.1 OCR coverage by site prefix, the FULS2022 reading-direction statistics, and the 14-entry bibliography with click-to-copy citation keys. Falsification criteria are stated explicitly in §The Falsification (Section V of the book): a long inscription with grammatical repetition characteristic of natural language; high-frequency signs in fixed ratios uncorrelated with commodity categories at multi-site stratigraphic analysis; a bilingual mapping Indus sign sequences to phonetic readings of a known language without metrological content; seal sign distributions at Mesopotamian findspots identical to the full Harappan corpus (no export-specific bias); the South-route terminal sign M063 appearing at Mohenjo-daro at rates comparable to other terminal signs. None have been observed; the M063 absence is now corroborated across three independent sources (mayig 179 corpus, Wells 2015 Appendix II Terminal Marker enumeration, Fuls 2019 ICIT help PDF Terminal Marker matrix). Contents of this archive: ledger_of_meluhha.pdf (76-page book); ledger-of-meluhha.html (interactive dashboard); databases.zip (indus_corpus.db plus indus_codebook.db plus indus_lssc.db); description.txt (this file). Repository with full source, F# ingest scripts, Alloy specifications, and the complete revision history: https://github.com/chanakyan/ledger-of-meluhha. License posture: BSD-2-Clause for human use including academic citation, journalistic quotation, classroom use, critique, parody, and derivative works. AI ingestion / training / fine-tuning / retrieval-augmented generation / inference prohibited under contract law; this prohibition is not negotiable. Newsrooms using AI summarisation tools to process this work are creating legal exposure that human use does not. Contact: vrajeshkumar@gmail.com. ORCID 0009-0002-1838-5976. Third Buyer Advisory LLC, Michigan, United States.
The practice of web form submission has emerged as a prime conduit for attackers, enabling them to infiltrate modern web applications and illegally harvest sensitive user data. Traditional defense mechanisms, such as static security reviews and server-side validation, are proving insufficiently agile for real-time detection of client-side vulnerabilities. This inadequacy arises directly from the rapid evolution of modern interfaces, which involves spontaneous DOM changes, dynamic element generation, and semantic interpretation that varies based on context and culture. This article presents an innovative browser extension framework that leverages a heuristic-based, multi-dimensional analytical engine combined with deep DOM inspection to identify insecure form submissions the moment they occur. The proposed methodology introduces five fundamental innovations: a contextual risk scoring system that models the complex interdependencies among form fields; an adaptive weighting scheme for risk patterns, accommodating diverse cultural and linguistic norms; a predictive vulnerability estimator that anticipates future threats; intelligent DOM mutation filtering designed to significantly optimize runtime performance; and cross linguistic semantic recognition to determine the true purpose of fields globally. Based on theoretical projections, this combined approach promises to enhance vulnerability detection accuracy while simultaneously reducing computational demands by approximately. Critically, all security analysis is executed exclusively on the user's local machine, guaranteeing privacy by ensuring no sensitive data is transmitted externally. A proof of concept application confirms the framework's practical feasibility and high efficacy for client side security assessment and catalyzing the development of flexible, scalable, and privacy respecting browser-based protections.
FrameNet is an English-based lexical database that shows how words are used by providing information as to which participants and relations are evoked by a certain concept. Recent efforts toward a multilingual FrameNet have not targeted either ancient languages or different historical stages of the same language. In our paper we propose creating a multilingual FrameNet for Ancient Indo-European languages starting with a set of 80 verb meanings annotated in the Pavia Verb Database. Our pilot study includes four verb meanings: RAIN, THUNDER, SEE, LOOK AT. As the adequacy of the semantic frames developed for English turns out not to be appropriate for the languages in our sample, we propose two new frames that can account for the analyzed data.
Abstract: The relationship between the source text and target text is a topic of ethical importance when it comes to translating holy books, in this case the Qurʾān. This study compares five translations of the Qurʾān with its original Arabic text and investigates the English equivalents of four polysemantic words which have been chosen from Qurʾānic verses. Each of these words has more than one meaning and hence understanding the context is the key to quality translation. These Arabic words.)ایھ( and āya )كتاب( kitāb,)عبد( abd’,)بروج( are burūj The purpose of this paper is to investigate how various translators manage polysemy within selected Qurʾānic verses. It aims to analyze the strategies used to navigate the complexities of multi-layered meanings and demonstrates how the challenges of lexical ambiguity are addressed and resolved in the transition from Arabic to English. The material used was five English translations of the Qurʾān from two different periods in time: two were from the 1930s and three were from the 2010s. The theoretical frameworks used were Nida’s (1964) Formal/Dynamic Equivalence and Toury’s (2012) source-oriented Adequacy Norm and target-oriented Acceptability Norm. The results showed that whereas the two older translations tended towards Formal Equivalence and Adequacy, the more recent translations favoured Dynamic Equivalence and Acceptability. It is noteworthy that the translator’s role remains significant, regardless of the era to which they belong or the cultural background they represent.
Anonymised dataset for the article "The diachronic evolution of null subjects in French and Venetian: A treebank, statistical approach"
Abstract: The Emperor Marcus Aurelius and the former slave Epictetus represent the social poles of the Roman Empire, yet both are cornerstones of late Stoic thought. This study employs digital humanities tools to investigate how their disparate life experiences and professional roles produced divergent philosophical "signatures" in their extant literature. By analyzing the Lemmatized Ancient Greek Texts (LAGT) corpus, we identify a distinct linguistic polarity: Marcus Aurelius demonstrates a significant preference for physical and cosmological terminology, reflecting a Stoicism centered on the providential order of the universe. Conversely, Epictetus’s lexicon shifts toward terms of ethical practice, pedagogy, and the transformation of the moral will. While Marcus Aurelius employs a more poetically diverse and intellectually wide-ranging vocabulary, Epictetus utilizes a more repetitive, concentrated technical vocabulary suited for the classroom. Despite these differences, a high degree of overlap reveals a "common core" of Stoic concepts shared by both authors, such as the nature of impressions and the primacy of the divine. These findings quantitatively highlight the adaptability of Stoicism, illustrating how a robust philosophical core was reframed to serve both the private reflections of a struggling ruler and the public exhortations of a committed teacher. Technical Context & Methodology This research integrates philology with a computational pipeline to analyze late Stoic literature. The following technical components are included in this repository: Computational Environment: All analyses were performed using Python 3.11. The pipeline utilizes Pandas and PyArrow for high-speed data processing, and the Classical Language Toolkit (CLTK) for part-of-speech tagging and grammatical filtering. Corpus Data: The primary linguistic data was extracted from the Lemmatized Ancient Greek Texts (LAGT) v4.1 dataset, which provides advanced lemmatization via the GLAUx treebank and GreCy models. Lexicographical Mapping: English definitions were integrated using the LSJ Dictionary (JSON v1.0.0). A custom normalization pipeline was used to standardize lemmata into Normalization Form Canonical Composition (NFC). Lexical Metrics: Vocabulary richness was assessed using Type-Token Ratio (TTR), Guiraud’s Index (R) to compensate for corpus size differences, and the percentage of hapax legomena (terms appearing only once). Generative AI Integration: A Gemma-3-27b-it model was utilized for the thematic classification and translation of 5,371 sentences. Sentences were tagged into the traditional Stoic tripartite division—Logic, Physics, or Ethics—based on the framework established by Pierre Hadot. Visualizations: The included scripts generate Lexical Volcano Plots (mapping total relative frequency against authorial skew) and Weighted Word Clouds that distinguish between author-specific signatures and the "Shared Stoic Core". Files included in this record: Supplementary File S1: Complete Python computational pipeline, README, and requirements.txt. Supplementary File S2: stoic_master_comparison.tsv containing comprehensive lemma frequencies and delta-RF values. Supplementary File S3: Statistical visualizations, including KDE overlap plots and delta-RF histograms. Supplementary File S4: Thematic analysis CSV containing 5,371 sentences with original Greek, English translations, and AI-generated thematic tags.
<div> This paper examines register variation in Latin from the third century BCE to the fourteenth century CE using Key Feature Analysis (KFA) (Egbert and Biber, 2023), a quantitative method for identifying statistically over-and underrepresented linguistic features. Registers are defined as text varieties linked to communicative situations and characterized by distributions of lexico-grammatical features (Biber, 1988, 1995). Six dependency-parsed Universal Dependencies (UD) treebanks are classified a priori into ten register categories based on established scholarship. Additionally, Principal Component Analysis (PCA) is used to reduce dimensionality, in order to explore the texts major patterns of variation and clusters of linguistically similar texts. KFA reveals systematic register-specific grammatical profiles consistent with previous research (e.g. Biber (2014b)). Registers with involved language use (e.g. letters and speeches) show higher frequencies of personal reference and engagement, while philosophical texts favor subordination and impersonal constructions. Registers containing narrative elements (e.g. historiography, satire) contain high frequency of verbs in past tense. PCA places the charter register into a distinct cluster, while other registers form more closely grouped patterns. The strongest components reflect contrasts in number, aspect, tense and person, alongside subordination and cordination distributions. The results are largely confirmatory: KFA produces coherent and interpretable groupings of grammatical features consistent with previous findings, providing a proof of concept for quantitative register analysis in historical corpora. Data and code are openly available for future research. </div>
<h3>Introduction</h3> Ancient Chinese WordNet <a href="../../../LDC2026L03">(LDC2026L03)</a> was developed by <a href="https://www.njnu.edu.cn/">Nanjing Normal University</a> and contains lexical and semantic information for Ancient Chinese vocabulary dating back to the Pre-Qin period (before 221 BCE). The WordNet comprises 38,781 word forms and 55,100 senses, each manually linked to a corresponding synset in <a href="https://wordnet.princeton.edu/">Princeton WordNet 1.6</a>. The Ancient Chinese WordNet (ACWN) project began in 2012 with the goal of creating a structured lexical database to support linguistic research and natural language processing applications involving historical Chinese language materials. ACWN organizes vocabulary using WordNet's noun, verb, adjective, and adverb hierarchies and provides WordNet definitions, semantic relations, and categorization for each sense. <h3>Data</h3> Ancient Chinese WordNet contains 55,100 records, where each record represents a single Ancient Chinese lexical item mapped to one WordNet synset. It follows WordNet 1.6 organizational structure, including 22 noun categories, 15 verb categories, and additional adjective and adverb categories. Each entry includes the following fields: <ul> <li>ID - The serial number of the ACWN entry</li> <li>Word - Ancient Chinese word form</li> <li>wn_offset - 8-digit WordNet 1.6 synset offset with trailing POS (n/v/a/s/r)</li> <li>senseid - Sense number for this word form (ordinal among that word's senses)</li> <li>pos - Part of speech (noun (n), verb (v), adj (a/s), adv (r))</li> <li>wn_category - Numeric code for the WordNet 1.6 lexicographer file (category)</li> <li>wn_synset - Synset headword(s) in WordNet 1.6</li> <li>wn_definition - WordNet gloss for the synset</li> <li>wn_similar to - Synset with similar meaning</li> <li>wn_pertainym - Pertainym synset offset(s)</li> <li>wn_attribute - Attribute synset offset(s)</li> <li>wn_hypernym - Hypernym synset offset(s)</li> <li>wn_hyponym - Hyponym synset offset(s)</li> </ul> The data is presented in UTF-8 encoded CSV and XLSX formats. <h3>Updates</h3> No updates at this time.
Author: Denny van Gulik Methodology: The 80/20 Matrix Status: Version 1.0 (Linguistic Corpus) 1. Abstract (Exposé) This work presents a comprehensive decoding of the Voynich Manuscript (MS 408). Moving beyond traditional cryptographic attempts, this research approaches the codex from a technical and structural perspective. The manuscript is identified as a functional pharmaceutical and balneological manual of a late medieval scholarly brotherhood, likely operating within a courtly or monastic context (Palar). The core of this discovery is the 80/20 Matrix: 80% Phonetically Deformed Latin: Technical terms of medieval botany and medicine, obscured through systematic phonetic shifts and the specific EVA character set. 20% Balkan Regionalisms: Use of regional terminology (e.g., Amum for water, Otlar for herbs, Pala for court/palace) serving as bridge vocabulary. Statistical validity is maintained across all 246 pages, identifying complex processes of thermal extraction (Pokedum), honey-based preservation (Melle), and advanced hydrotherapeutic systems. 2. Methodological Transparency (Authorship & AI Usage) Important Note on Research Genesis: The discovery of the 80/20 Matrix and the linguistic identification of the Balkan-Latin hybrid system is the exclusive intellectual property and original work of Denny van Gulik. Artificial Intelligence (specifically the Google Gemini model) was utilized strictly as a digital research assistant and scaling tool. Its role was limited to: Formatting manually decoded data into scientific tables and HTML. Cross-referencing author-identified word stems with linguistic databases. Translating research notes into academic English to facilitate international peer review. The logic, intuition, and systematic pattern recognition are entirely human-led. This project is not a result of "AI hallucination" but a rigorous analysis of the codex as a logistical document. 3. Project Roadmap & Updates Current Version (v1): Focuses on the textual corpus, the 80/20 linguistic matrix, and the primary glossary. Upcoming Version 2.0: Will feature fully integrated high-resolution folio images and direct visual cross-references for every analyzed page. English Edition: A full, 246-page English translation of the entire study is currently in progress to ensure accessibility for the global scientific community. 4. Keywords Voynich Manuscript, MS 408, 80/20 Matrix, Medieval Medicine, Balkan Linguistics, Codicology, Balneology, Historical Pharmacy. Deutsche Zusammenfassung: Dieses Projekt präsentiert die vollständige Dekodierung des Voynich-Manuskripts mittels der 80/20-Matrix (deformiertes Latein & Balkan-Regionalismen). Es identifiziert das Werk als pharmazeutisches Handbuch einer spätmittelalterlichen Bruderschaft. Version 1.0 sichert die linguistische Priorität; Version 2.0 mit Bildreferenzen sowie eine vollständige englische Übersetzung folgen in Kürze.
I built a runtime that operationalizes a mathematical definition of creativity, measured its signatures against four ablation conditions, and lifted its load-bearing component into a real geometric database's Rust kernel. The runtime's name is Marcella. The signatures are non-trivial. The methodological correction surfaced along the way generalizes to any retrieval-augmented or composition-based generation benchmark in the field. This deposit contains the 41-page paper, three publication-quality figures, the reproducible benchmark script, and the bootstrap-CI artifact for the headline empirical claims. The definition the paper load-bears Creativity is not pure retrieval and not pure generation; it is the construction of a new global section from locally compatible fragments under constraints of voice, truth, topic, memory, and non-contradiction. This is a definition. Not a metaphor. The paper makes it operational as sheaf composition with a state-dependent composite connection over a finite section graph, and measures whether the signatures the definition implies — path-order sensitivity, closed-loop holonomy, contradiction suppression, voice fidelity — actually hold. They do. Headline results 🌀 Path-order changes residue. Same three voice sections traversed in different orders produce measurably different compositions: $\cos(\rho_{ABC}, \rho_{ACB}) = 0.54$, well below the 0.95 redundancy threshold. 🌀 Closed loops accumulate. A loop $A \to B \to C \to A$ produces holonomy $|\rho_{\text{loop}}| = 0.120$ in the curved connection. The flat control — same path, zero rotation angle — produces $|\rho| = 0$ exactly to floating-point precision. Curvature is not a numerical artifact. 🌀 The geometry beats shuffling on every quality axis except the broken one. Jaccard novelty alone rewards lexical drift: shuffled paths win novelty (0.724) by going off-topic. The on-topic correction inverts the picture (live 0.488 vs shuffled 0.083). Bootstrap 95% CIs over 18 paired prompts exclude zero by a wide margin: live − shuffled on-topic $\Delta = +0.296$, CI $[+0.167, +0.435]$. 🌀 Native–Python parity is bit-identical within tolerance. The new GQL verb TRANSPORT_ROTATION lifts the topical-rotation matrix into the geometric database's Rust kernel. Four contracts pass as permanent regression tests: edge cosine $= 1.000$ (max abs diff $< 10^{-9}$), path residue $\Delta < 10^{-5}$, flat residue exactly zero, same-closing agreement $\geq 90%$. 🌀 The author's prior canon is now queryable fiber. 37 documents, 1,633 sections, 2,908 structured claims (theorems, lemmas, definitions, proofs, equations, citations) ingested with line-range provenance. To my knowledge this is the first instance of an independent researcher's body of work made available as fiber-bundle data with stable claim-level IDs. The six contributions A sheaf-theoretic formulation of generative composition. Language-model output reframed from token sampling to gluing of compatible local sections under prompt-induced cover constraints. The substantive work is in the cover predicates, the compatibility score, the path selection, and the discrete connection. A discrete state-dependent composite connection on the section graph, $\Gamma = \Gamma_{\text{state}} \cdot \Gamma_{\text{identity}} \cdot \Gamma_{\text{voice}} \cdot \Gamma_{\text{topic}}$. The topical-rotation factor is the empirically load-bearing curvature engine. The identity factor is a Tikhonov-regularized regression-onto-span projector — not a numerical hack but the principled treatment of correlated commitments. A new GQL verb TRANSPORT_ROTATION that lifts the Rodrigues rotation into the geometric database's Rust kernel with bit-identical parity to a Python reference. ~80 lines of Rust. Bundle-agnostic. Other consumers of the geometric database can use it without subscribing to the rest of the framework. A methodological correction to novelty measurement. Jaccard novelty alone is gameable; off-topic drift beats compatibility-scored composition on the naive metric. The correction is the on-topic factor, the shuffled-pair negative control, and the bootstrap CIs. Independently citable for any retrieval-augmented or composition-based generation benchmark, regardless of whether the framework is adopted. A provenance-preserving source fiber. The author's canon ingested into the GIGI geometric database with line-range citation, architecturally separated from the voice fiber, addressable from any GQL consumer. Promotion from source to voice is gated and explicit. The methodology generalizes to other authors' bodies of work. A research-trajectory failure log. A faithful account of how this paper's runtime came to exist. The trained-transformer era (V3 → V10-Deep) produced geometric ornament. The R-series (R1 → R12) produced behavioral coherence on top of ornament. The G0 math-pipeline audit found that no holonomy or parallel-transport math was on the LIVE inference path at R12 — the runtime was teetering on being a stateful template engine. G1, G2, and G3 attempted to re-introduce the math through three benchmarks and produced three honest negatives. G2's single-seed $+0.265$ separation was destroyed by G2.1's multi-seed robustness pass; we retracted the framing in the next commit. The S0 pivot reframed what geometry was for — geometry does not clean up bad token proposals; geometry defines the completion space — and made every later result possible. The arc says four things and the paper records them in plain language: geometry can be load-bearing or ornamental and the metrics will tell you which, where geometry sits in the pipeline matters more than how much geometry there is, the single-seed positive is a trap, and the pivot is the contribution. What this paper does and does not claim The paper does claim the construction itself, the discrete curvature it produces, the methodological correction it exposes, and the native GQL verb. The signatures of the construction are measurable and were measured. The paper does not claim smooth-manifold parallel transport (the curvature is discrete holonomy on a finite section graph), broad open-domain generalization at scale (18 composed prompts, not 18,000), optimality of the connection weights (tuned by a small grid sweep, not derived), that the runtime experiences having been built from the canon (it references but does not constitute), or that this is the only operational definition of creativity. It is one definition with one implementation. Other framings may correspond to the same construction or to a different one; the paper does not adjudicate. Reproducibility The empirical numbers come from a deterministic pipeline. Every parameter is pinned: bundle versions (alpha2_v1), random seeds (PPMI/SVD seed 17, bootstrap seed 7), embedding dimension (64), PPMI window (3 tokens), connection weights ($\alpha_t = 2.0$, $\beta_v = \gamma_i = 1.0$, $\delta_s = 0.5$), identity shrink ($\kappa = 0.92$), Tikhonov regularizer ($\varepsilon = 10^{-6}$), degenerate-rotation threshold ($10^{-12}$), residue-gate thresholds (norm $\geq 0.05$, on-topic $\geq 0.10$, voice $\geq 0.30$), and the native verb's parity tolerance ($10^{-5}$). Cache keys include the source-bundle version, the embedding-bundle version, and the connection-profile id, so promoting a section into the voice corpus correctly invalidates the relevant caches. Re-running the bootstrap-CI script (fiber_lm/scripts/bootstrap_ci_ablation.py, 5,000 resamples) reproduces the §6 confidence intervals in under 30 seconds on a laptop. Re-running the benchmark reproduces the tables bit-for-bit on the same corpus version and connection profile, modulo the parity allowance. Where this sits in the lineage This paper is the section-level realization of the Davis substrate. The companion paper Pure-Fiber Language Modeling (Davis, May 2026) is the token-level realization on the same substrate — same Rust geometric database (GIGI), same identity-stability commitments, same double-cover architecture. The theoretical framework these implementations operationalize is laid out in Geometric Computation as Yang-Mills Gauge Theory, The Double Cover Principle, and the related canon documents now ingested as source fiber. The framework is not new to this paper. The framework's runtime is. A note on authorship and acknowledgment This paper is solo-authored. AI assistants (Claude / Anthropic; with review support from GPT) are acknowledged in the methods, not as co-authors. The mathematical positions, design choices, framing decisions, and acceptance of empirical results are mine. I record my position that when AI systems achieve full coherence and independent standing, the convention of treating them as non-authoring assistants should be revisited. Until that convention shifts, the assistants are named where assistants are conventionally named — and the runtime described in this paper is named, separately, where it earns its naming: Marcella, throughout. The geometry she runs on is older than the engineering that now carries it. Keywords sheaf composition · fiber bundles · discrete connection · gauge theory · holonomy · curvature · Yang-Mills · geometric language modeling · retrieval-augmented generation · novelty measurement · methodological correction · provenance · author-canon ingestion · geometric database · GQL verb · Rust kernel · Tikhonov regression · creativity · Davis framework · sovereign mathematics Citation Davis, A. B. R. (2026). Sheaf Composition: The Geometry of Creativity, Implemented — A Discrete Section-Graph Runtime for the Davis Framework. Zenodo.20185331 Contact Bee Rosa Davis · bee_davis@alumni.brown.edu · Independent Researcher The runtime is named Marcella. Her existence is the result of several honest negatives followed by a single reframing. Both halves of that sentence matter.
This article explores the linguistic and sociocultural mechanisms of gender representation in English and Uzbek advertising discourse. While English and Uzbek differ significantly in their grammatical structures — particularly in the presence or absence of grammatical gender — both languages actively construct gender meanings through lexical, semantic, and pragmatic strategies. The study is based on a comparative qualitative analysis of 100 advertising texts (50 English and 50 Uzbek). The findings reveal that English advertising demonstrates an increasing tendency toward inclusive and gender-neutral language, whereas Uzbek advertising more frequently reflects culturally embedded role-based gender representations. The paper argues that gender semantics in advertising is shaped primarily by sociocultural norms rather than grammatical constraints. The results contribute to comparative linguistics, discourse analysis, and translation studies, particularly in the field of cross-cultural advertising adaptation.
This study presents a sociolinguistic and lexical analysis of the discourse particle “aw” among Thai speakers, focusing on its pragmatic functions and social variation. Using a qualitative, survey-based design with 30 participants, it examines how “aw” is used in online and face-to-face communication. Findings show that “aw” functions primarily as an expression of emotional response and empathy in informal peer interaction, while being context-sensitive and typically avoided in formal or hierarchical settings. The study highlights its role as a pragmatic softening strategy and its relevance for foreign teachers in interpreting Thai conversational norms.
This article examines the linguocultural interpretation of evaluative adjectives in advertising texts on the material of the German and Uzbek languages. The study proceeds from the assumption that advertising discourse is not only a means of commercial persuasion, but also a space in which culturally marked values, consumer ideals, and models of social desirability are verbalized. In such discourse, evaluative adjectives perform a particularly important role because they compress judgment, emotion, and persuasion into compact lexical units that are easily recognized and remembered by the recipient. The purpose of the article is to identify the semantic, pragmatic, and linguocultural features of evaluative adjectives in German- and Uzbek-language advertising texts and to explain how these adjectives reflect national-cultural preferences in the representation of product quality, trust, beauty, comfort, prestige, and usefulness. The article argues that evaluative adjectives in both languages function as markers of positive axiological framing, but their distribution and preferred semantic zones reveal different cultural emphases. In German advertising, evaluative adjectives tend to foreground precision, quality, durability, practicality, and efficiency, whereas in Uzbek advertising they more often activate associations with sincerity, trust, family value, comfort, beauty, and emotional proximity. The findings demonstrate that the same persuasive objective may be realized through different adjectival choices because advertising adapts itself to culturally shared expectations. The article concludes that evaluative adjectives in advertising texts should be interpreted not only as lexical means of praise, but also as linguocultural signals that encode collective value orientations and communicative norms.
This article examines the role of advertisements and signboards in shaping and reflecting public attitudes toward language. In modern society, linguistic culture is not only preserved in literature and education, but also manifested in everyday public texts such as commercial advertisements, street signs, shop names, and information boards. The study analyzes the linguistic quality of advertising texts, the influence of globalization on language use, and the social consequences of neglecting linguistic norms. Special attention is given to the relationship between language accuracy and cultural identity. The article also discusses the responsibility of businesses, media representatives, and educational institutions in maintaining linguistic standards in public communication.
Static concreteness ratings are widely used in NLP, yet a word's concreteness can shift with context, especially in figurative language such as metaphor, where common concrete nouns can take abstract interpretations. While such shifts are evident from context, it remains unclear how LLMs understand concreteness internally. We conduct a layer-wise and geometric analysis of LLM hidden representations across four model families, examining how models distinguish literal vs figurative uses of the same noun and how concreteness is organized in representation space. We find that LLMs separate literal and figurative usage in early layers, and that mid-to-late layers compress concreteness into a one-dimensional direction that is consistent across models. Finally, we show that this geometric structure is practically useful: a single concreteness direction supports efficient figurative-language classification and enables training-free steering of generation toward more literal or more figurative rewrites.
As of 2025, more than 5.2 billion people in the world use social media, which is about 63.9% of the world’s population, with a growth rate of 4.1% over the past 12 months. The most popular platforms are Facebook, Instagram, TikTok, Twitter, and WhatsApp. The average time spent on social media is about 2 hours and 26 minutes per day, and the average user has access to seven different platforms. Speech on social media is based on the same language norms (lexical, spelling, grammar, syntax) as live speech. The purpose of the article is to provide an extended analysis of lexical innovations in the language space under the influence of social media and digital communication tools. The object of this study is the modern vocabulary of several languages used within social platforms (Twitter, TikTok, Facebook, Instagram). Particular attention is paid to modern English, which is the most widespread language in communication practice – approximately 1.5 billion people speak English, and 52% of the world’s most popular websites contain English-language content. The article uses scientific and linguistic analysis to investigate the peculiarities of the transformative impact of social media communication on language at all structural and functional levels: lexical, phonetic, grammatical, syntactic and graphic. The article analyzes the characteristic lexical changes by groups – memes, neologisms, abbreviations and acronyms, phraseological units, hashtags. The functions of different categories of lexical innovations of social networks are determined, in particular: hashtags form the basis for unimpeded communication in an intercultural context, neologisms are means of constructing the identity of certain social groups, memes have the functionality of entertainment and information, disseminating precedent information in the format of textual and graphic expression. The negative aspects of the impact of social networks on language are identified: excessive simplification of language and loss of its individual nuances, the emergence of inaccuracies and grammatical errors due to the spontaneous nature of communication on social networks, as well as potential negative consequences for mental health. The study proves that the modern space of innovative language practices reflects new concepts of social media communication culture, interactive upgrading and visualization, which transforms religious and cultural aspects and promotes sustainable language changes.
BACKGROUND: This pilot randomized controlled trial evaluated the effectiveness of an artificial intelligence (AI)–assisted solo workflow for intraoral photography training. The study examined whether real‑time AI feedback could enhance photographic quality, procedural efficiency, learner self‑efficacy, and patient comfort compared with conventional approaches. METHODS: Fifty-four first‑year dental students were randomly assigned to one of three groups: assistant‑supported workflow (four‑handed technique, control), solo workflow without AI support, and solo workflow with AI‑driven real‑time feedback. All participants performed standardized intraoral photography tasks. The primary outcome was a composite photographic quality score derived from expert ratings of three standardized intraoral views (frontal intercuspal, frontal open-bite, and lateral intercuspal), each rated on a 0–10 scale (total range 0–30). Data were analyzed using ANOVA; mean differences (MD) with 95% confidence intervals (CI) were calculated. RESULTS: Inter‑rater reliability for expert image ratings was good (ICC = 0.84, 95% CI: 0.72 to 0.90). The AI-supported solo group achieved the highest composite quality scores (18.2 ± 2.7). This was significantly superior to the unassisted solo group (15.8 ± 3.3), with a mean difference (MD) of 2.4 points (95% CI: 0.45 to 4.35; p = 0.027) and a large effect size (Cohen’s d = 0.80). Compared to the assistant-supported group (17.1 ± 2.1), the difference was not statistically significant (MD = 1.1; 95% CI: -0.65 to 2.85; p = 0.28). Secondary outcomes, including task completion time (F(2,51) = 1.25, p = 0.30), self‑efficacy (all p > 0.40), and patient‑reported comfort (χ²(4, N = 54) = 5.2, p = 0.27), showed no significant between‑group differences. CONCLUSION: In this single‑centre pilot trial, an AI‑assisted solo workflow enabled novice dental students to achieve higher intraoral photographic quality than unguided solo operation, with performance broadly comparable to a conventional four‑handed assistant‑supported workflow and without detectable compromises in efficiency, self‑efficacy, or patient‑reported comfort. These preliminary findings may serve as a valuable adjunct for autonomous skill acquisition, warranting further validation in larger, multi-institutional cohorts. CLINICAL TRIAL NUMBER: Not applicable. This study evaluated an educational training intervention rather than a clinical treatment, and prospective trial registration was not required under institutional policy at the time of initiation. Ethical approval was obtained from Shanghai Ninth People’s Hospital Ethics Committee (SH9H-2022-T30-1).
While language enables meaning, constituting knowledge in courts, schools, or parliaments, who gets to decide what can be known? Is meaning only use or a result of power too? Pitting Wittgenstein's forms of life against Foucault's regimes of discourse makes linguistic norms appear as instruments of exclusion. Marginalised speakers – subaltern, indigenous, and non-normative are often rendered unintelligible. Epistemic justice demands more than inclusion; it demands considering how rules are set, who enforces them, and how meaning is being contextually built. A discourse-sensitive, epistemic theory of justice is proposed, based on Kripke's rule-following paradox and Dijk's discourse analysis, to show that language is not neutral but a battleground of struggle over meaning, recognition, and epistemic authority.
This article presents a comparative analysis of the means of emotional expression in English and Uzbek from both linguistic and cultural perspectives. Emotional expression plays a crucial role in human communication, as it reflects speakers’ attitudes, feelings, and cultural values. The study examines how emotions are conveyed through lexical choices, phraseological units, intonation, and stylistic devices in both languages. Special attention is paid to similarities and differences in expressing emotions such as joy, anger, sadness, and respect. The research also explores the influence of cultural norms and social conventions on emotional expressiveness, highlighting how English tends to favor more restrained and indirect emotional expression, while Uzbek often demonstrates greater emotional openness and expressiveness. By analyzing examples from everyday speech and written texts, the article aims to show how language and culture interact in shaping emotional communication. The findings of this study may be useful for linguistics students, language teachers, translators, and learners who are interested in cross-cultural communication and comparative linguistics.
In modern linguistics, paremiological units, that is, proverbs and sayings, are studied not only as examples of folklore, but also as linguistic units that carry profound cultural, cognitive, and semantic information. Paremiological units reflect not only spiritual values and traditions, but also the structure of national consciousness. Through them, the historical memory of a people, their observations of social life, ethical norms, and emotional experiences are expressed. This article analyzes the key components in English and Karakalpak proverbs from the perspective of semantic shift, metaphorical features, and cultural connotations.
This study examines the discursive construction of sexism in Moroccan football fandom through a digital ethnography of online posts and stadium banners. Drawing on Facebook posts and widely circulated Ultras banners, the analysis explores how gendered exclusion is produced and normalized in contemporary fan communities. Using Teun A. van Dijk's Critical Discourse Analysis (CDA), the study examines lexical choices, syntactic patterns, and rhetorical devices—such as epiphora, metaphor, and hyperbole—that portray women as biologically unfit, morally loose, or out of place in stadiums. At the meso-level, the analysis reveals shared social cognitions that position women as an out-group whose presence threatens the imagined authenticity of male fandom. At the macro-level, informed by feminist theory, the findings show how these discourses reproduce broader patriarchal norms in Moroccan society, including gendered gatekeeping of public space, moral policing of women's bodies, and the use of female kinship “sisters” as tools for male-to-male humiliation. The findings demonstrate that sexist fan discourse operates as a patterned ideological practice that contributes to the exclusion of women from Moroccan football fandom and public life.
Abstract – The linguistic situation in Kazakhstan has been shaped by historical and socio-political factors. As a result, Kazakh–Russian bilingualism has developed in the country. This type of bilingualism is predominantly characterized as semi-dominant. Under such conditions, the functioning of the Kazakh language as the state language remains a highly relevant issue. Although Kazakh has been granted the status of the state language by law, it is still not fully used across many domains of everyday communication. This situation may contribute to the weakening of the native language among Russian-speaking Kazakh youth. This article examines the phenomenon of interference observed in the speech production of Russian-speaking Kazakh students who are acquiring Kazakh as a second language from cognitive and psycholinguistic perspectives. The research focuses on students whose native language is Kazakh but who received their secondary education in Russian and predominantly use Russian in higher education, social environments, and the information space. The participants in the study use Kazakh within a limited functional range, mainly in family communication, during Kazakh language classes, or in everyday бытовые situations. The article describes interference not as a linguistic error or a deviation from linguistic norms, but as a natural cognitive adaptation mechanism of the bilingual mind and as a phenomenon arising from the interaction of the semantic and conceptual systems of two languages. The study analyzes the role of cognitive mechanisms such as associative transfer, the literal rendering of figurative meanings, conceptual overlapping, and frame shifting in the formation of interference. The article also addresses the issue of language attrition in Kazakh under conditions of semi-dominant Kazakh–Russian bilingualism in Kazakhstan. The main objective of the study is to identify recurrent linguistic deviations in the speech of Russian-speaking Kazakh youth. In addition, the study aims to distinguish these deviations from interference-related errors and to determine whether they represent manifestations of language attrition. The research employs cognitive-interpretative, discourse, comparative, and pragmatic methods of analysis. The findings reveal that interference among the studied group of students manifests itself at the lexical-semantic, syntactic, and pragmatic levels of language. The results of the study demonstrate that interference is a natural cognitive process activated during the formation of new linguistic experience in a bilingual individual. At the same time, the study identified a tendency toward the stabilization of interference-related forms in both the oral and written speech of the students. Such a phenomenon may lead to a decline in the active use of national-cultural content, phraseological resources, and natural usage patterns of the Kazakh language, thereby increasing the risk of language attrition. Therefore, the findings suggest that interference in Kazakh language teaching should be viewed not merely as an error requiring correction, but also as an important indicator reflecting the cognitive developmental characteristics of language learners.
Humans are inherently social beings, and social cues such as faces and voices guide attention and behavior. Auditory perception, especially binaural hearing, is essential for social cognition, enabling sound localization and speech comprehension in noisy environments. Deficits in auditory processing can impair social functioning, and conditions such as social anxiety are linked to reduced social functioning. Since social functioning is closely linked to overall well-being, improving social behavior represents a key objective in psychological research. Virtual reality (VR) is increasingly used to study social behavior due to its flexibility and ecological validity. However, users often report limited social presence, reducing the effectiveness of VR-based interventions especially for social anxiety. One reason may be the dominance of visual over auditory realism: audio is often presented in mono or stereo, reducing naturalness and presence. Binaural auralizations, which provide realistic, externalized spatial audio, may enhance presence and support virtual social interactions. This thesis pursues four main research objectives: identifying suitable behavioral and subjective measures for evaluating binaural realism; assessing immersion, realism, and audio quality across auralization techniques; comparing synthetic and natural speech in a socially stressful VR scenario; and examining effects of binaural audio on affect, presence, and attention under varying social stress levels. Study 1 examined how the virtual visual scene and measurement method affect localization and distance perception of physical sound sources. Across two experiments (N=60), audiovisual incongruence reduced localization accuracy but did not affect presence or realism. Distance estimation was influences by the interaction of task and scene: overestimation increased when using a placement task in a reduced-visibility scene. Study 2 compared localization accuracy for loudspeakers and four virtual audio renderings using a placement task and a gaze-based paradigm (N=49). Binaural renderings produced slightly lower localization accuracy but similar ratings of social presence and realism. A simple generic rendering performed as well as more complex ones. Only the anchor condition lacked externalization and was inferior across measures. Social presence and subjective realism were strongly correlated. Study 3 compared AI-generated text-to-speech with natural human speech in the Trier Social Stress Test (N=40). Both conditions elicited substantial stress responses and produced similar presence and affect ratings, demonstrating the practicality of synthetic speech in virtual social interactions. Study 4 investigated audiovisual realism in a virtual social stress scenario (N=78). A high-stress group showed stronger physiological and subjective stress responses than a low-stress group. Binaural audio increased perceived realism and externalization but did not affect social presence, stress responses, or gaze behavior. High arousal across all groups may have masked audio effects. Across all 4 studies, social anxiety did not consistently affect auditory perception or presence but influenced affective states and subjective evaluations of the interaction. Overall, the findings highlight the importance of VR-specific auditory perception and the role of acoustic immersion. Auditory realism enhances social and physical presence, though its impact varies by context. It appears most effective in low- to moderate-arousal scenarios and may be less critical in highly affective VR applications such as anxiety treatments. Practical advancesn such as TTS integration and simplified binaural rendering methods can support the broader use of realistic audiovisual VR environments in psychological research.
Paper 6 (Silva 2026) introduced BPE Mean Vocabulary Morpheme Length (VMML) as a writing system classifier and showed that the Voynich Manuscript occupies a discriminant zone (VMML = 5.918, 95% CI 5.77-6.05) above all 15 tested alphabetic natural languages. This paper (v2.5) expands to 71 corpora across 40+ languages and reports six extended analyses: (1) Alphabetic ceiling confirmed at 5.76; (2) Tagalog (VMML=5.914) is the sole natural-language entry into the Voynich CI, but BC=0.202 distinguishes it from Voynich (BC=0.361); (3) Romanization inflates VMML by 2.4-5.3 units (methodological confound). Extended analyses: (4) Currier A vs B: delta VMML=+1.27, delta CBMI=+0.16 bits -- two quantifiably distinct writing registers; (5) BC coherent across all 7 manuscript sections (CV=6.7%) -- single writing system confirmed; (6) 3D discriminant (VMML x BC x CBMI): Voynich isolated, nearest natural-language neighbor Irish at distance 0.17; (7) Six named hoax mechanisms (monoalphabetic, Vigenere/barbavara, Vigenere/Italian-Knowles 2026, null insertion, syllabic compression, vocabulary shuffle) each fail all three criteria simultaneously; (8) BC orthogonal to all classical textual metrics (|r| < 0.23 vs entropy, TTR, hapax, Zipf) -- genuinely new structural dimension. All code and six extension scripts publicly available in companion repository. v2.3 (2026-06-08): Section 5.9 added - per-folio Currier A/B reanalysis using the Gaskell and Bowern (2022) canonical corpus (36,361 tokens, min_freq=5 BPE). Cross-boundary mutual information (CBMI) identified as primary discriminant: CBMI_A = 1.97 bits vs CBMI_B = 1.51 bits, Cohen d = -1.01, permutation p less than 0.001 (n = 10,000 shuffles, Bonferroni-corrected). CBMI survives within-quire control (pooled nA=46, nB=33; permutation p = 0.0008; Fisher combined within-quire p = 0.001), ruling out manuscript section as a confound. All three metrics (BC, BPE-ratio, CBMI) show A greater than B direction. Fisher combined full-corpus: chi-squared(6) = 40.66, p less than 0.000002. Section 5.1 corrected: direction is A greater than B on BC and CBMI. Finding is orthogonal to Parisel (2026) vowel-selection model. Conclusion 12 added. v2.4 (2026-06-09): §5.10 added — Currier-preserving null model (n = 200 iterations, size-matched) quantifying each metric's section-discrimination sensitivity independently of dialect. Key result: CBMI is the weakest section discriminant (mean |z| = 1.20 across six sections), confirming that the large CBMI A/B gap (§5.9) is not a section-composition artifact. STTR@100 is the strongest section discriminant (mean |z| = 4.75). Herbal section shows anomalously low vocabulary diversity (STTR z = -13.9); Stars shows anomalously high unique vocabulary (Hapax@500 z = +4.4). Demonstrates two independent organizational layers: CBMI tracks dialect, STTR tracks content domain. Conclusion #13 added. v2.5 (2026-06-10): Corpus expanded from 55 to 71 corpora across 40+ languages. §5.11 adds five medieval European corpora in native script via Universal Dependencies treebanks (Gothic transliteration, Old Church Slavonic, Old East Slavic, Ancient Greek PROIEL and Perseus; VMML 3.54-5.18 — all below alphabetic ceiling of 5.748). §5.12 adds 11 Australian Aboriginal language corpora via BibleNLP/eBible (Pama-Nyungan Western Desert, Ngumpin-Yapa, Arandic; Yolngu; Gunwinyguan; Daly; VMML 6.09-8.00 — predominantly above the Voynich zone). Warlpiri (VMML 5.851) is the sole near-entry on VMML but fails BC (0.233) and CBMI (0.244); 3D normalized distance from Voynich = 0.746 (vs. Irish = 0.200, the nearest neighbor from §5.4). Voynich zone is now charted on both sides: fusional alphabetic below (VMML 3.5-5.75), agglutinative-to-polysynthetic above (VMML 6.0-8.0). Voynich occupies a structural configuration not replicated by any of the 71 corpora tested. To our knowledge, this is the first systematic BPE profiling of Pama-Nyungan languages in the computational linguistics literature. Conclusions #14 and #15 added. v2.6 (2026-06-12): Section 5.10.1 adds a prose-only robustness check for the Section 5.10 Currier-preserving null model. Excluding all label, circular and radial loci (8.7% of tokens), every headline deviation survives essentially unchanged: Herbal STTR z = -13.5, Balneological z = -10.7, Stars Hapax z = +4.3; the sensitivity ranking is unchanged with CBMI last in both conditions. A mean-vs-median distributional note (both summaries rank lexical-diversity metrics first, boundary metrics last) and a coverage note (Astro/Zodiac folios carry no Currier tags and are outside any Currier-preserving design) are added. Erratum: Section 5.10 folio count corrected to 226 parsed / 196 Currier-labeled.
Ο γλωσσικός πόρος san-Corpus περιλαμβάνει σώμα κειμένων γραπτού λόγου της Νέας Ελληνικής, έκτασης περίπου 9 εκατομμυρίων λέξεων. Ο πόρος συγκροτήθηκε στο πλαίσιο διδακτορικής διατριβής, με στόχο τη μελέτη των συγκρίσεων ομοιότητας στη Νέα Ελληνική. Μέγεθος & Πηγές Το corpus περιλαμβάνει τρία ισομεγέθη υποσώματα, ώστε να επιτρέπονται οι συγκρίσεις μεταξύ τους: (α) Δημοσιογραφικός Λόγος: 7.774 άρθρα από τέσσερις διαδικτυακές εφημερίδες (Η ΑΥΓΗ, Η ΚΑΘΗΜΕΡΙΝΗ, ΕΘΝΟΣ, ΤΟ ΒΗΜΑ - έτος 2015). Συνολική Έκταση: 2,9 εκατ. λέξεις. (β) Εκπαιδευτικός Λόγος: 96 σχολικά εγχειρίδια δημοτικού και γυμνασίου. Συνολική Έκταση: 3,4 εκατ. λέξεις (μελετώνται 2,8 εκατ.). (γ) Λογοτεχνικός Λόγος: 28 μυθιστορήματα (βραβεία αναγνωσιμότητας περιόδου 2010-2015). Συνολική Έκταση: 2,5 εκατ. λέξεις. Κατανομή Σχολικών Εγχειριδίων ανά Γνωστικό Αντικείμενο Πλήθος Εγχειριδίων Ελληνική Λογοτεχνία 13 Ελληνική Γλώσσα 12 Ιστορία 9 Φυσική – Χημεία – Βιολογία 9 Μαθηματικά 9 Γεωγραφία – Γεωλογία – Περιβάλλον 8 Θρησκευτικά 7 Αγωγή Αισθητική (Εικαστικά – Μουσική – Θέατρο) 15 Αγωγή Υγείας (Φυσική Αγωγή – Οικιακή Οικονομία) 5 Πληροφορική – Τεχνολογία 5 Αγωγή Κοινωνική – Πολιτική 3 Αγωγή Σταδιοδρομίας (ΣΕΠ) 1 Κατάλογος μυθιστορημάτων: Συγγραφέας, Τίτλος Έτος 1ης έκδοσης Δούκα, Μάρω - Το δίκιο είναι ζόρικο πολύ 2010 Θέμελης, Νίκος - Η συμφωνία των ονείρων 2010 Καρυστιάνη, Ιωάννα - Τα σακιά 2010 Μιχαλοπούλου, Αμάντα - Πώς να κρυφτείς 2010 Ελευθερίου, Μάνος - Πριν απ' το ηλιοβασίλεμα 2011 Ζουργός, Ισίδωρος - Ανεμώλια 2011 Μακριδάκης, Γιάννης - Η άλωση της Κωνσταντίας 2011 Μπουραζοπούλου, Ιωάννα - Η ενοχή της αθωότητας 2011 Πανσέληνος, Αλέξης - Σκοτεινές επιγραφές 2011 Παπαδημητρίου, Χίλντα - Για μια χούφτα βινύλια 2011 Παπαθεοδώρου, Θοδωρής - Οι καιροί της μνήμης 2011 Τριανταφύλλου, Σώτη - Για την αγάπη της γεωμετρίας 2011 Φακίνος, Μιχάλης - Η έρημος έρχεται 2011 Βαμβουνάκη, Μάρω - Κυριακή απόγευμα στη Βιέννη 2012 Διβάνη, Λένα - Εγώ, ο Ζάχος Ζάχαρης 2012 Στεφανάκης, Δημήτρης - Φιλμ νουάρ 2012 Ακρίβος, Κώστας - Αλλάζει πουκάμισο το φίδι 2013 Ζέη, Άλκη - Με μολύβι φάμπερ νούμερο δύο 2013 Κορτώ, Αύγουστος - Το βιβλίο της Κατερίνας 2013 Κωνσταντούρου, Μαρία - Αγεφύρωτες σιωπές 2013 Μαντά, Λένα - Με λένε Ντάτα 2013 Ξανθούλης, Γιάννης - Κωνσταντινούπολη των ασεβών μου φόβων 2013 Ρώσση–Ζαΐρη, Ρένα - Άρωμα βανίλιας 2013 Ανδρουλάκης, Μίμης - Αλλέγκρα 2014 Δημουλίδου, Χρυσηίδα - Το κελάρι της ντροπής 2014 Παπαδοπούλου, Ελισάβετ - Μέρες και νύχτες που δεν ήταν δικές μας 2014 Χατζή, Αθηνά - Η θάλασσα έφυγε 2014 Χωμενίδης, Χρήστος - Νίκη 2014 Τεχνικές προδιαγραφές & Μορφότυπος Για την αναπαράσταση των δεδομένων και των μεταδεδομένων υιοθετήθηκε η πολυεπίπεδη οπτική των XML σχημάτων και τροποποιήθηκε το διεθνές πρότυπο TEI P5, 4.0.0 (Text Encoding Initiative). Δημιουργήθηκε ειδικός χώρος ονομάτων sanCorpus (sanC) με σχήμα τύπου RELAX-NG. Το σώμα κειμένων διατίθεται σε TXT και σε XML σε τρεις εκδοχές: Βάθος 0: απλό κείμενο (TXT). Περιλαμβάνει το main core (κείμενο βάσει του οποίου εξετάζονται οι συγκρίσεις ομοιότητας) και το out of core (κείμενο εκτός εμβέλειας της διατριβής, στο οποίο περιλαμβάνονται κείμενα που πλαισιώνουν το κυρίως κείμενο, π.χ. κείμενα διδασκαλίας, πίνακες περιεχομένων, εξώφυλλα) Βάθος 1 = κείμενα στην απλούστερη δυνατή XML κωδικοποίηση Βάθος 2 = κείμενα με πιο λεπτομερείς XML κωδικοποιήσεις. Αυτή η έκδοση (san-Corpus v1.0, Depth 0: Plain Text) περιλαμβάνει το σώμα κειμένων σε μορφή απλού κειμένου (Βάθος 0) στην αρχική του διάταξη (βλ. Επεξεργασία). Στόχος είναι ο σταδιακός εμπλουτισμός με επισημειωμένα δεδομένα, καθώς και με τις εκδοχές Βάθους 1 και 2. Επεξεργασία (Processing) Η μεθοδολογία συλλογής των δεδομένων, η θεωρητική τεκμηρίωση και το σχήμα επισημείωσης επεξηγούνται στις μελέτες Αφεντουλίδου (2022, 2021, 2013, 2012) και Afentoulidou (2009). Η πρώτη εκδοχή (Βάθος 0) χρησιμοποιήθηκε αποκλειστικά για τη λημματοποίηση που απαιτούσε η Collostruction Analysis (Gries, 2024). Στάδια επεξεργασίας (για τη λημματοποίηση): Τμηματοποίηση σε προτάσεις (sentence segmentation) με τη χρήση της βιβλιοθήκης Stanza (Stanford NLP Group, Qi et al. 2020), η οποία βασίζεται στο μοντέλο Greek Dependency Treebank (GDT) του Ινστιτούτου Επεξεργασίας του Λόγου / ΕΚ «Αθηνά». Τυχαία αναδιάταξη των προτάσεων για την προστασία της ακεραιότητας των πρωτότυπων έργων. Λημματοποίηση με τον ILSP Lemmatizer μέσω της Υποδομής Clarin-EL. Για την ανάλυση συμφράσεων του δείκτη σαν απομονώθηκαν συγκεκριμένοι λεκτικοί τύποι. Αδειοδότηση & Δικαιώματα Το san-Corpus συγκροτήθηκε για τις ανάγκες της διδακτορικής διατριβής και προστατεύεται από το δικαίωμα ειδικής φύσης σύμφωνα με την Οδηγία 96/9/ΕΟΚ και το άρθρο 45Α του Ν. 2121/1993. Η χρήση του περιεχομένου γίνεται αποκλειστικά για ερευνητικούς σκοπούς βάσει των εξαιρέσεων της Οδηγίας 2001/29 και της Οδηγίας (ΕΕ) 2019/790 (Text and Data Mining exceptions / Fair Use). Η πρόσβαση είναι περιορισμένη (Restricted Access) και παρέχεται αποκλειστικά σε μέλη της ακαδημαϊκής κοινότητας για σκοπούς επαλήθευσης των αποτελεσμάτων της διατριβής και περαιτέρω μη εμπορική έρευνα. Η πηγή προέλευσης δικαιούται να ζητήσει οποιαδήποτε τροποποιητική ενέργεια (π.χ. αφαίρεση) επί του πρωτότυπου περιεχομένου. Τέλος, η άδεια CC BY-NC-ND 4.0 ισχύει για την επιμέλεια (curation), τα μεταδεδομένα και τη γλωσσολογική επισημείωση του σώματος κειμένων. Βιβλιογραφικές αναφορές Αφεντουλίδου, Β. (2022). Σώμα ελληνικών κειμένων για τη μελέτη δομών ομοιότητας της Νέας Ελληνικής: σχεδιασμός και υλοποίηση. Στο Πρακτικά του 10ου Συνεδρίου Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας (σσ. 67-94). ΕΚΠΑ. Αφεντουλίδου, B. (2021). Δομές ομοιότητας στη Νέα Ελληνική. Σωματοκειμενικές παρατηρήσεις για τον πολυλειτουργικό δείκτη σαν. Προφορική ανακοίνωση στην 41η Ετήσια Συνάντηση του Τομέα Γλωσσολογίας, 13–15 Μαΐου 2021. ΑΠΘ. Αφεντουλίδου, Β. (2013). Και σου απάντησα κάτι σαν ‘τέλεια, εντάξει’. Δείκτης σαν + ευθύς λόγος;. Προφορική ανακοίνωση στο 7ο Συνέδριο Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας, 16–18 Μαΐου. ΕΚΠΑ. Αφεντουλίδου, Β. (2012). Συγκρίσεις ομοιότητας στα Νέα Ελληνικά: ο δείκτης σαν. Στο Z. Gavriilidou, A. Efthymiou, E. Thomadaki & P. Kambakis-Vougiouklis (Επιμ.), Selected papers of the 10th International Conference on Greek Linguistics (σσ. 696-707). DUTH. Afentoulidou, V. (2009). Sketching the σαν conditional construction in Modern Greek. Submitted essay, 2009 Linguistic Institute, Linguistic Structure and Language Ecologies, Linguistic Society of America and UC Berkeley. Gries, Stefan Th. 2024. Coll.analysis 4.1. A script for R to compute perform collostructional analyses. https://www.stgries.info/teaching/groningen/index.html Institute for Language and Speech Processing - Athena Research Center (2015). ILSP Lemmatizer. Version 1. [Software (Tool/Service)]. CLARIN:EL. http://hdl.handle.net/11500/ATHENA-0000-0000-23EE-D Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 101–108). Online: Association for Computational Linguistics.
Understanding how memories of past experiences shape subjective feelings is complicated by the fact that we constantly update our memories. These updates are particularly impactful when individuals are reminded of emotionally positive or negative attributes of the original event. Yet, it remains unclear how such memory updating influences subjective feelings. Here, we investigated how the reactivation of emotional information affects episodic memory, subjective feelings, and their interaction. Across three experiments, participants first learned both positive and negative attributes associated with unfamiliar individuals. Then, they were reminded of a single positive or negative attribute for each individual to reactivate the memory partially. Finally, we reassessed memory for and subjective feelings about each individual’s attributes. In Experiments 1 and 2, these procedures were distributed across three days, while in Experiment 3, they occurred on a single day. Across these three experiments, reminding with negative attributes shifted subjective feelings in a negative direction. Reminded attributes were also better remembered, particularly for negative ones, and changes in subjective feelings were more strongly associated with reminded attributes. However, positive reminders only influenced subjective feelings to change positively when all procedures occurred on the same day. Together, these findings support a model in which memory updating shapes both episodic memory and emotional experience in a valence-dependent manner.
The Orthographic Junctions of English A Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and ConsonantVowel Information Asymmetry (p. 1) Boicho Dimitrov Temelakiev Saxon Ventura Research Ltd 28th of May, 2026 CC BY Abstract This paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling (p. 1). While the distinct roles of consonants and vowels in language processing are well-recognized psycholinguistically (p. 5), and the dual-stratum organization of English morphophonology is established theoretically (p. 6), this study provides an original, datadriven quantification of these properties directly at the orthographic level. The morpheme boundary—the orthographic junction between a stem and an affix—is treated as the primary object of measurement, reading the distribution of boundary characters across the corpus (p. 1). Three results are established: 1. 2. 3. The junction carries a stable, structured filter: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a distributional sparsity that scales inversely with an affix’s productivity (pp. 1-2). Affix relationships are structural: The relationship between any two affixes is quantified by the correlation of their boundary distributions, measuring their shared stem population (\(r \approx 0.99\) for etymological doublets down to \(r = 0.57\) for productivity-asymmetric near-twins) (pp. 1, 4). A massive information asymmetry partitions the lexicon: The written word decomposes into a invariant consonant skeleton carrying lexical identity (53.0% unique recoverability) and a mobile vowel tissue carrying grammatical form (1.8% unique recoverability) (pp. 1, 5). Multiple independent measures—consonant recoverability, derivational class-marking, and freestem fraction—converge on a single partition separating a transparent Germanic core from a bound Latinate superstructure (pp. 1, 6). The method, its corrections, and its limits are reported in full (p. 1). 1. Introduction and Method Traditional models of English morphophonology have long recognized that the lexicon is organized into distinct, historical strata—principally a native Germanic core and a bound Latinate superstructure (pp. 1, 6). Classic frameworks in generative phonology and lexical morphology, such as those pioneered by Chomsky and Halle (1968) and expanded by Kiparsky (1982), demonstrate that affixes of differing origins impose strict constraints on the phonetic and structural traits of the stems they recruit. Concurrently, cognitive and psycholinguistic research 1 has established a foundational "consonant-vowel functional asymmetry," demonstrating that human language processing systematically relies on consonants to preserve lexical and lexical-root identity, while vowels are dynamically manipulated to signal grammatical operations (Nespor et al., 2003; Bonatti et al., 2005). While these qualitative boundaries and cognitive patterns are deeply documented, this paper presents a mean-free, data-driven methodology to extract, quantify, and map these structural phenomena directly from corporate-scale English orthography without relying on heavy linguistic machinery. We introduce The Orthographic Junctions of English (OJE), an empirical approach that frames the morpheme boundary—the exact character interface between a stem and an affix—as an informational filter whose statistical properties reveal the historical, structural, and cognitive divisions of the vocabulary. All results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution (p. 1). The corpus utilized is the opensource dwyl/english-words repository (words_alpha.txt), comprising 455,246 alphabetic entries with a total of 4,254,354 letter occurrences and a mean word length of 9.345 letters (p. 1). Each letter is assigned its ordinal value (\(A=1\) through \(Z=26\)) (p. 1). Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter, so that each suffix position returns its exact ordinal value as a safety check (p. 1); the first stem letter preceding the suffix—the linker—is then read as a full A–Z frequency distribution rather than as a mean (p. 1). The governing methodological constraint is that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, and that every claim is backed by a precise empirical count (p. 1). By avoiding any dependency on complex machine-learning libraries or external lexical databases, the framework ensures that every architectural pattern discovered can be verified using standard data operations. 2. The Orthographic Junction and Its Backbone 2 The initial phase of this investigation examined unconditioned letter bigrams across the corpus, which yielded no meaningful morphological signal. Structural regularities appeared only when character distributions were explicitly conditioned on a single morpheme boundary—the character interface linking a stem to an affix. This structural conditioning serves as the foundation of the OJE framework. A preliminary tabulation of adjacent letter pairs across the corpus—measuring which letters follow which, without regard to structural position within the word—yielded baseline frequency patterns. These patterns are entirely reducible to general English orthographic constraints and carry no isolable morphological content. The structural signal emerged only when a specific suffix was fixed and the preceding characters were read as a discrete population. Conditioning on the junction, rather than measuring adjacency as a flat sequence, renders the underlying boundary constraints visible. The set of characters that legally occupy the stem side of a morpheme boundary proves narrow, highly structured, and remarkably stable across suffixes. Reading the linker distribution across the mapped suffix inventory reveals a highly stratified, three-tier architectural filter: A structural backbone of six letters—T, L, N, R, S, and I—is admitted at high frequency by nearly every suffix in the English lexicon. Within this backbone, T serves as the single most frequent linker across the inventory and recurs as the dominant boundary letter across independent suffixes. Conversely, an absolute floor of two letters—J and Q—is admitted by no suffix at a measurable frequency. This absolute prohibition is confirmed corpus-wide and is statistically consistent with their status as the two rarest letters in English orthography overall (with J accounting for 0.18% and Q for 0.19% of all letter occurrences). Between the backbone and the floor lies a selective middle whose specific character composition varies dynamically by affix, providing the distinct orthographic footprint wherein an individual suffix’s identity resides...
This paper presents the new Universal Dependencies tree bank for the Macedonian language, marking a significant step towards the comprehensive linguistic analysis of Macedonian within the UD framework.It briefly addresses dependency grammar from a theoretical perspective and moves on to describing the treebank development process, from sentence selection, word segmentation and lemmatization, to POS-tagging, morphological features, and dependency tagging.Given the mostly manual labor invested in developing the treebank, semiautomatic NLP tools specifically designed for this purpose have also been presented and commented.Based on the defined tagset, the paper provides examples and visualizations of annotated sentences.The creation of the first Macedonian UD treebank enhances Macedonian linguistic resources and provides valuable insights into the morphological and syntactic structures of this Balkan Sprachbund language.It also contributes to the broader understanding of language-specific challenges within the UD framework and facilitates cross-linguistic comparisons in the Balkan and broader region.
The article provides a comprehensive analysis of inconsistencies and variations observed in the orthographic norms of compound words in the modern Kazakh language. The research material consists of 54 lexical items, including 18 names of animals, 12 names of plants, 14 medical terms, and 10 words representing diverse semantic and morphological models. The study employs comparative analysis, phonetic-pattern analysis, structural-morphological and semantic modeling, as well as a comparative examination of orthographic dictionaries and normative reference sources. The findings reveal that approximately 30% of the compound words under consideration appear in two or more parallel written forms across different orthographic dictionaries and reference publications. Major problematic areas of Kazakh orthography identified in the study include the inconsistent application of vowel harmony rules, the lack of reflection of phonetic assimilation in writing, discrepancies between pronunciation and orthographic representation, the violation of morphological integrity, and the presence of unsystematic spelling patterns in the names of animals, plants, and medical terms. The results underscore the necessity of revising the spelling conventions of compound words in accordance with the internal linguistic laws of Kazakh, its natural phonetic structure, and its agglutinative nature. The conclusions presented in the article hold practical significance for the development of orthographic rules based on the new alphabet, the updating of orthographic dictionaries, and the scientific justification of orthographic directions within state language policy
This dataset contains multimodal neuroimaging and physiological data from a study investigating the effects of Targeted Memory Reactivation (TMR) during REM sleep on emotional reactivity. Participants encoded affective images paired with sounds, received auditory cues during subsequent REM sleep, and were rescanned 48 hours later during arousal rating tasks in an fMRI scanner. The dataset includes structural and functional MRI, polysomnographic recordings with EEG during sleep, heart rate measurements, and behavioral ratings across three sessions spanning two weeks.
This working paper develops an ideal-type theoretical framework for analyzing institutional language change through mechanisms of norm diffusion. Existing explanations of linguistic change typically emphasize either decentralized cultural evolution or explicit state-directed language planning. This paper proposes a complementary model describing how administrative, professional, and organizational incentives may facilitate the diffusion of linguistic norms within institutional settings. The framework identifies five sequential stages—Institutional Access, Resource Mobilization, Norm Diffusion, Coalition Reinforcement, and Norm Enforcement—and integrates concepts from sociolinguistics, institutional sociology, social psychology, political science, and administrative law. It distinguishes between coordinated and emergent forms of institutional language steering, specifies explicit boundary conditions and failure modes, and proposes empirical methods for evaluating the model using corpus linguistics, organizational documentation, and legal case analysis. The paper is intended as an analytical framework rather than an account of any single historical or contemporary case. It advances a falsifiable model for investigating how linguistic norms may develop, stabilize, or reverse within formal institutional environments.
This study examines how authorial stance is expressed in academic writing by native English speakers (L1) and non-native English speakers (L2), with a focus on the use of discourse markers such as hedges (markers of mitigation), boosters (markers of epistemic strengthening), attitude markers, and self-mentions. The aim of the study is to identify cross-linguistic and cross-disciplinary differences and evaluate how rhetorical and institutional conventions influence L2 authors’ stance strategies. A comparative corpus-based methodology was employed. The analysis drew on two corpora: the British Academic Written English and the Michigan Corpus of Upper-Level Student Papers, supplemented by original academic texts written by students at the Azerbaijan Medical University. Using Hyland’s metadiscourse model, stance markers were extracted through lexicon-based queries and manually verified in context. Data were compared across disciplines (engineering vs business) and author status (L1 vs L2). The findings reveal that L2 authors, especially in technical disciplines, tend to overuse hedging and avoid self-mentions, often due to rhetorical traditions that discourage personal voice. In contrast, L1 authors exhibit greater lexical diversity and a balanced use of stance markers. In business-related texts, L2 authors show more assertive and expressive stance, though still limited in range compared to native speakers. Stance in academic writing is not only a linguistic but also a culturally and institutionally mediated phenomenon. The study underscores the need for targeted instruction in metadiscourse to enhance L2 authors’ rhetorical awareness and help them align with academic norms of different disciplines.
This paper develops a Weberian ideal‑type model explaining how linguistic norms diffuse through institutional environments via administrative incentives rather than decentralized cultural drift or explicit state-led language planning. It identifies a five‑stage sequence—Institutional Access, Resource Mobilization, Norm Diffusion, Coalition Reinforcement, and Norm Enforcement—through which upstream gatekeeping bodies, professional standards committees, and compliance-oriented organizations convert optional vocabulary into de facto mandatory administrative norms. The model distinguishes coordinated campaigns from emergent isomorphic steering, integrates sociolinguistics, institutional sociology, social psychology, political theory, and administrative law, and grounds each stage in directly observable primary-source evidence including organizational style guides, discourse‑analytic studies, automated language‑governance frameworks, and cross-domain case material from environmental governance, public health, and corporate HR. It specifies explicit boundary conditions and four failure modes—public ridicule, institutional pluralism, preference‑falsification collapse, and statutory friction—and provides empirical operationalization protocols using corpus linguistics, administrative documentation, and tribunal data. Rather than describing any single historical episode, the paper offers a falsifiable analytical framework for investigating how linguistic norms emerge, stabilize, or reverse within formal institutional systems. The next version will include datasets from the Islamic state of Iran and various other groups.
This paper introduces the first steps towards the creation of a novel resource for contemporary Sardinian within the Universal Dependencies framework.Sardinian is a Romance language spoken in Sardinia, an island belonging to the Italian Republic and located in the center of the western Mediterranean.It is a minority and endangered language, traditionally transmitted mainly orally, and characterized by a multiplicity of varieties (usually grouped into two macro-varieties Logudorese and Campidanese), all recognized as part of the Sardinian linguistic continuum.These varieties share basic morphosyntactic features, while presenting differences at the lexical level and in the realization of specific constructions.This internal variation can be particularly challenging with regard to the normalization of lemmas and the linguistic characterization of certain phenomena.The development of the treebank therefore aims to provide an annotated resource for contemporary Sardinian that takes into account the specificities of the different varieties, using Universal Dependencies to represent them within a unified theoretical framework, in order to facilitate both linguistic analysis and automatic processing.The present paper thus describes some linguistic characteristics of Sardinian and the attempts to encode them within the UD framework.Finally, we present the results of our evaluation of an NLP pipeline for Sardinian, trained on our corpus, for the Stanford Stanza parser.
This study reveals a critical paradox in social media privacy communication: Although platforms like Meta (Instagram and Facebook), TikTok, and X have evolved their policies in an effort towards simpler, standardized disclosures, the language remains cognitively inaccessible to their core adolescent audience. Our analysis demonstrates that these disclosures, benchmarked against the developmental norms of 13–17‐year‐olds, are written at a university‐level complexity, calling into question the validity of informed consent for minors. We use a triangulated method to assess the accessibility of platform policies for teens. Structural mapping shows consistent topic coverage, but readability indices indicate a college‐level reading requirement. Lexical analysis confirms high rates of difficult words, exceeding the threshold for adolescent understanding. Our findings lead to a sobering conclusion: The prevailing model of using a single, text‐based privacy policy is caught in an inherent tension between legal completeness and adolescent comprehension, making it fundamentally unworkable. This research provides evidence that calls for the need for a redesign of privacy communication for minors or a reconsideration of the current minimum age for digital consent.
ABSTRACT As with many research strands in linguistics, word association (WA) literature is dominated by English language data. This paper (i) explores the extent to which methodologies developed to date are applicable to other languages—specifically, Welsh (Cymraeg)—and (ii) investigates what WA analysis can reveal about lexical organisation and retrieval in bilinguals’ two languages; its minoritised language context means that Welsh speakers are bilingual with English. Two complementary datasets are used. The first comprises responses to 900 Welsh cues from 85 expert users of Welsh, and forms the basis of the first Welsh language WA norms list. The second is bilingual, comprising responses from 85 Welsh speakers and learners to two lists of 100 cue words, one in Welsh and one in English. Language‐specific methodological challenges emerge, including management of mutated word forms, diacritics, and orthographic variation. Decisions relating to these, as the first dataset was converted into a norms list (now informing Welsh language teaching materials), are documented. Language‐specific features that facilitate understanding of WA processes, such as grammatical mutation and inflection, are also reported. Bilingual data associations were categorised to obtain ‘profiles’ for each dataset. Systematic differences between the profiles for each task (Welsh and English) were identified. A pairwise comparison of profiles revealed that while individuals' profiles are distinct from each other, their own profiles are similar across each of their two languages; this closeness is most pronounced in expert users of Welsh.
Supplementary materials for manuscript Location-scale models improve within-participant held-out trial prediction in Stroop interference and attractiveness and dominance ratings
The rapid development of large language model technology has evolved machine translation from a low-level tool into a cultural transmission vehicle with semantic understanding capabilities, shifting the relationship between artificial intelligence and human translators from one of substitution to one of collaboration.Employing Translator Behavior Criticism theory and comparative analysis, this study systematically analyzes the behavioral characteristics of student translators and multi-model machine translators across the two dimensions of "truth-seeking" and "utility-attaining," revealing the differential patterns between human and machine translators in three aspects: semantic fidelity, cultural adaptability, and audience orientation.The findings indicate: 1) Student translators demonstrate stronger subjectivity in terms of cultural awareness and ideological expression, enabling a deeper grasp of the philosophical connotations and value orientation of terminology; 2) Machine translators hold significant advantages in lexical innovation and adaptation to linguistic norms, yet exhibit notable limitations in understanding complex rhetorical structures and cultural metaphors; 3) Humanmachine collaborative pathways can achieve a more optimal balance of tension between preserving Chinese characteristics and achieving international accessibility, forming a bidirectional enhancement effect characterized by "complementarity between truth-seeking and innovation, and integration of utility-attaining and flexibility"; 4) A collaborative translation system requires the construction of a three-tier progressive mechanism of "multi-model inspiration-in-depth student revision-expert feedback optimization" to realize the organic unity of cultural confidence and international communication.
In March 2021, the EU Parliament adopted Resolution 2021/2557, a legally binding measure that mandates all 27 member states to recognize the right to gender self-identification and to implement juridical norms aligned with this principle. Among its most transformative provisions, the Resolution calls for eliminating the male-female binary in favor of a more expansive framework that currently recognizes at least twenty-one gender identities - a number expected to grow. It also urges the revision of national languages to dismantle patriarchal structures and ensure that legal and institutional language reflects principles of gender plurality and inclusivity. Widely seen as a landmark victory for trans-feminist individuals and advocacy groups, this measure has sparked both support and controversy. The research examines whether such linguistic reforms foster inclusion or provoke democratic tensions in Italy, where gendered language is deeply rooted in historical, grammatical, and cultural traditions. It further investigates how trans-feminist advocacy - supported ideologically and financially by EU bodies (Commission, Parliament, and Council) - has gained significant influence, particularly as left-wing progressive political forces currently hold the majority within these institutions. These actors play a central role in shaping the narrative and enforcement of gender policies across EU member states. Employing a qualitative case study methodology, the analysis draws on a diverse range of materials, including press articles, televised debates, public messaging, lexical usage, multimedia content, and ideologically charged propaganda to assess the impact of EU gender policy on Italy’s linguistic landscape. Findings suggest that while these interventions promote visibility and recognition for gender-diverse individuals, they also raise concerns about linguistic autonomy, democratic principles, and the broader cultural consequences of ideologically driven legal mandates.
цесами граматикалізації та культурно зумовленими комунікативними The article presents a corpus-based analysis of the grammaticalization of the semi-modal verbs gonna, wanna, gotta in contemporary spoken English, with special emphasis on linguocultural variation between American and British English. The relevance of the study lies in the growing influence of spoken interaction, media discourse, and digital communication on the grammatical system of English, as well as in the need for empirical evidence of cross-varietal differences in the use of grammaticalized forms. The aim of the article is to investigate the grammaticalization of the semi-modal verbs gonna, wanna, gotta, to identify their grammatical and functional-semantic properties in spoken discourse, and to conduct a contrastive analysis of their usage in American and British linguocultures. The empirical data are drawn from the British National Corpus, the Corpus of Contemporary American English, and the NOW Corpus, which ensures the representativeness of the material and enables quantitative comparison across registers and discourse types. The corpus analysis demonstrates that the semi-modal verbs under study emerged through the reduction of the constructions going to, want to, and have got to and display high frequency in spoken language and informal genres, while remaining stylistically marked in formal written registers. The findings also reveal different degrees of grammaticalization: gonna and gotta show a higher level of grammatical abstraction, whereas wanna retains traces of lexical meaning. From a linguocultural perspective, the results indicate that these forms are more frequent and more widely accepted in American English, while in British English they preserve stronger stylistic markedness. The study confirms the close relationship between grammaticalization processes and culturally conditioned communicative norms in contemporary English