Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This article presents an extended linguocultural analysis of the emotion of joy as expressed through the concepts of “smile” and “laughter” in English, Russian, and Uzbek. The study integrates approaches from cognitive linguistics, cultural linguistics, pragmatics, and discourse analysis to investigate how emotional experience is encoded, structured, and interpreted across languages and cultures. Special attention is given to lexical semantics, phraseology, metaphorical models, and communicative norms. The findings reveal that while smile and laughter are universal physiological and emotional responses, their linguistic representation is shaped by culturally specific values, social expectations, and communicative traditions. English tends toward conventionalized politeness and frequent smiling, Russian demonstrates emotional authenticity and restraint, and Uzbek reflects collectivist warmth and hospitality. The article contributes to cross-cultural linguistics by demonstrating that emotional concepts are culturally mediated and linguistically structured. The research also has implications for intercultural communication, translation studies, and language teaching.
The present study offers an original multi-level linguistic analysis of child speech in the 3-5 age group, combining a structured observational methodology with a qualitative-quantitative framework developed by the author. Unlike existing survey-based accounts, this paper proposes a three-tier Phonetic–Lexical–Grammatical (PLG) Developmental Profile as a practical diagnostic tool for assessing speech maturity in preschool children. The research was conducted on the basis of 120 spontaneous speech samples collected from 40 Kyrgyz-Russian bilingual children aged 3 to 5 years attending preschool institutions in Bishkek. The study identifies age-specific patterns of phonological simplification, lexical growth strategies, and syntactic complexity, and introduces a composite Speech Maturity Index (SMI) that integrates quantitative indicators across all three levels. The findings demonstrate that bilingual children in the studied cohort display a systematically accelerated grammatical development compared to monolingual norms while exhibiting specific phonological transfer patterns not previously described in the literature on Kyrgyz–Russian bilingualism at preschool age. The proposed PLG Profile and SMI are presented as replicable instruments suitable for use by preschool educators and speech therapists.
This paper develops a hierarchical account of concept formation that distinguishes the manifestation of originary qualitative feel from the public, norm-governed concepts through which subjects later articulate experience. Its central claim is not merely that concepts differ in abstractness, but that they differ in constraint profile: some remain tightly corrigible by recurrent felt manifestations delivered through relatively stable experiential channels, whereas others depend far more heavily on social reweighting, comparison, and narrative organization. To capture this difference, I distinguish four strata: (1) the manifestation of originary qualitative feel, (2) channel-bound concepts, (3) integrative judgments, and (4) narrative concepts. The pivotal case is smart. In some contexts, smart functions as a third-level integrative judgment anchored by a relatively stable cue-cluster such as rapid learning, flexible problem solving, and quick comprehension. In other contexts, the same lexical item is drawn into fourth-level use, where application depends on broader evaluations of which forms of knowledge, achievement, or tradition count as genuinely worthwhile. This shift shows that semantic drift does not abolish embodiment; it redistributes which embodied, affective, and inferential resources are recruited as operative sense changes. The paper then reorganizes its dialogue with contemporary theory around this problem of functional migration. Barsalou and Dove help explain embodied anchoring but do not sufficiently account for the reweighting of application rules within public evaluative space. Borghi and collaborators illuminate the role of language, sociality, inner speech, and conversational coordination in abstract concepts, yet their frameworks can be sharpened by distinguishing lexical stabilization from discursive reorganization and by connecting internal cue-stabilization with sentence-level and discourse-level renegotiation. Brandom and Jorem and Lohr clarify the public inferential dimension of concept use, but require a stronger account of differential experiential corrigibility. Searle helps explain how high-level concepts approach recognition and status, but not every narrative concept is thereby reduced to a full institutional fact. The conclusion is that the decisive issue is not whether a concept is social, but how lower-level cues, public norms, and narrative frames are combined in use.
The project Entangled Histories used early modern printed normative texts. The computer used to have significant problems being able to read Dutch Gothic print, which is used in the vast majority of the sources. Using the Handwritten Text Recognition suite Transkribus (v.1.07-v.1.10), we reprocessed the original scans that had poor quality OCR, obtaining a Character Error Rate (CER) much lower than our initial expectations of <5% CER. This result is a significant improvement that enables the searching through 75,000 pages of printed normative texts from the seventeen provinces, also known as the Low Countries. The books of ordinances are compilations; thus, segmentation is essential to retrace the individual norms. We have applied – and compared – four different methods: ABBYY, P2PaLA, NLE Document Recognition and a custom rule-based tool that combines lexical features with font recognition. Each text (norm) in the books concerns one or more topics or categories. A selection of normative texts was manually labelled with internationally used (hierarchical) categories. Using Annif, a tool for automatic subject indexing, the computer was trained to apply the categories by itself. Automatic metadata makes it easier to search relevant texts and allows further analysis. Text recognition, segmentation and categorisation of norms together constitute the datafication of the Early Modern Ordinances. Our experiments for automating these steps have resulted in a provisional process for datafication of this and similar collections.
This paper takes a short sample of student spoken interaction and analyses it in detail to identify some of the features of the talk which may be at variance with what would be recognized as more proficient and fluent interaction.The goal is to identify points of interactional practice that can be judged as areas for consciousness raising and explicit instruction by the teacher.Several of the points raised in the analysis are suggested to stem from a complex set of influences including transfer of lexical, grammatical, and interactional practices from the L1.There also may be a habituation to classroom discourse when speaking in the L2 which is then used unconsciously as a template for non-institutional mundane social interactions.Recognition of the special nature of learner interactions alongside an understanding of the possible causes of this kind of speaking can, it is suggested, inform focused and empirically based teaching that develops learners' interactional competence.The competence versus performance duality posited by Chomsky (1965) is based indirectly on a notion of native speaker intuition.That is, when a native speaker of a language hears an utterance in that language, they can make an immediate judgment of its acceptability in grammatical terms.Deviations from the norm in
This article deals with the study of similarities and differences in expression of the correspondence of the norm to English and Tatar linguistic cultures using the material of phraseological units with a gender component. The paper defines the concept of norm in language and phraseology. One singles out the main groups of phraseological units according to whether they conform to the norm or not. The gender direction in linguistics, the subject of which is the interrelation of language and gender as a social factor, considers the concepts such as “genderâ€, “femalenessâ€, “malenessâ€. Gender is expressed in semantics and in the grammar of the language, forming a linguistic image of the world, which in turn depends on the conceptual image. The gender image of the world is not biologically determined, and the concepts of femaleness and maleness are determined by cultural and historical factors, in particular, by language stereotypes in different cultures and language communities. Gender metaphor also influences the formation of a conceptual and linguistic image of the world. A gender metaphor is understood as “the transfer not only of the physical, but also of the totality of spiritual qualities and properties, united by the nominations of femininity and masculinity to the objects that are not connected with sexâ€. In different language communities, femininity and masculinity referents often do not coincide, which creates difficulties in intercultural communication and translation.
This article explores the relationship between language and gender identity in English and Uzbek social media discourse. It examines how linguistic choices, including lexical items, pronouns, and stylistic markers, reflect and construct gender identities in online communication. By analyzing social media posts, comments, and interactions, the study highlights patterns of gendered language use and the cultural influences shaping these practices. The findings suggest that social media provides a dynamic space where users perform, negotiate, and challenge traditional gender norms through language. Cross-linguistic comparison reveals both similarities and differences in how English and Uzbek speakers encode gender identities in digital communication.
The present article examines the content and linguistic features of the realization of the concept of “happiness” in the English language. Drawing on semantic, pragmatic, and cognitive approaches, the study explores how happiness is encoded through lexical units, phraseological expressions, and contextual usage. Special attention is given to the interaction between language and culture, as well as to the role of discourse in shaping the interpretation of emotional concepts. The findings demonstrate that happiness is a multifaceted concept that reflects not only individual emotional states but also cultural values, social norms, and communicative intentions.
There has been a long felt need to investigate what the study of poetic language contributes to an understanding of ordinary language, and how posing the question in this way may indeed shift some of the assumptions about the way ordinary language works. Metaphor remains an important case in point: How do we get from “He kicked the bucket?” to “He died”? In this metaphorical idiom, the process is unavailable, without going into etymological hypotheses, because the meaning is already given – it’s lexicalized, a “dead metaphor” -- no semantic construal is involved and the literal meaning doesn’t play a role in understanding the meaning of the expression. A poet might break up the idiom and say, “He kicked the bucket and broke it to pieces” – defamiliarizing the automatic reading of the idiom as “he died” and opening up suggestions, for example, of overcoming death; thus lending plasticity to and making present both the figurative and literal meaning. But true access to the process of metaphorical meaning-making is only available in cases of novel poetic metaphor, where the domains brought together are often disparate. I argue that there are good reasons for seeing poetry, fiction and literature in general not as speech acts but as representations, or semblances of speech acts (Searle’s notion of a pretend speech act doesn’t do this justice). Poems stage dramatic situations in which both poetic speaker (that is, the lyrical “I,”) and addressee, if any, are part of a fictional world, rather than the biographical poets themselves addressing a reader. The often ironic gaps and dialogical tensions between the norms and values projected by the poem itself and those expressed by any one of the voices in the poem – including the lyrical ‘I’s perspective – have a crucial expressive force unto themselves. As Searle realized, the indirect speech act model for poetry does not work; whereas “Could you pass the salt?” is a request masquerading as a question, there is no such one-to-one relationship between the poetic utterance and a single determinate speech act hiding behind it. Searle’s and Austin’s approaches couldn’t fully succeed, on my view, because there are key aspects of the language-game of poetry that distinguish poetry from ordinary language speech acts. Let’s say these are three aspects of a Gricean poetic contract between speaker and hearer: 1) the distinction between the lyrical ‘I’ and the biographical poet; 2) what I call the "quasipropositionality" of utterance in poetry (and fiction as well); and 3) poetic form and the ways in which it makes meaning through textual strategies of mediation, like point of view, parody, irony, enjambment, rhyme, meter, rhythm, sound-play, allusion, and metaphor. These forms of aesthetic mediation make it very problematic to isolate the serious or not-serious utterance of a single-line as a primary semantic unit.
This paper presents the new Universal Dependencies tree bank for the Macedonian language, marking a significant step towards the comprehensive linguistic analysis of Macedonian within the UD framework.It briefly addresses dependency grammar from a theoretical perspective and moves on to describing the treebank development process, from sentence selection, word segmentation and lemmatization, to POS-tagging, morphological features, and dependency tagging.Given the mostly manual labor invested in developing the treebank, semiautomatic NLP tools specifically designed for this purpose have also been presented and commented.Based on the defined tagset, the paper provides examples and visualizations of annotated sentences.The creation of the first Macedonian UD treebank enhances Macedonian linguistic resources and provides valuable insights into the morphological and syntactic structures of this Balkan Sprachbund language.It also contributes to the broader understanding of language-specific challenges within the UD framework and facilitates cross-linguistic comparisons in the Balkan and broader region.
These data contain three rows of data: 1. The unique identifier of the video and its order shown; 2. The subject's Likert rating of engagement, where 1-5 from least to most; 3. The subject's Likert rating of familiarity with the video's subject matter, also likert 1-5 from least to most.
The study is devoted to the analysis of the impact of globalization and digital technologies on the transformation of language norms in modern digital discourse. The relevance of the work is due to the growing role of digital communication, in which language becomes a key factor of cultural identity and social integration. The goal is to identify interdependencies between the level of digital maturity of states, the intensity of language hybridization and the dynamics of sociolinguistic variability. The object of the study is the digital language space of the global communication environment. The methodology is based on a combination of systemic, institutional, econometric and corpus-linguistic approaches using official statistical databases and corpora of digital texts for 2015–2024. The results show that during this period ICT Development Index increased from 5.32 to 7.41 points, DESI Index – from 48.7 to 70.4 points, and the share of Internet users – from 58.2% to 84.3%. At the same time, the number of languages with digital representation increased from 312 to 387, and the share of English-language content decreased from 55.1% to 49.3%. The hybridization index increased from 0.42 to 0.61, clearly indicating the establishment of a polycentric mode or multi-node digital discourse. Hybridization is both method and degree: different languages or their structural elements used in the same message (verbal or symbolic), as well as fully/partially mixed code messages; hashtags expressed in different linguistic forms/fonts, etc. (multimedia message). This means how various linguistic components are arranged within one structural unit up to multimedia messages comprising codes written with different fonts‐forms on various levels of hybridity. The econometric model registered a very high positive correlation between digital maturity and hybridization (r=0.82).
The Junction Grammar of English A Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and Consonant–Vowel Information Asymmetry Boicho Dimitrov Temelakiev Saxon Ventura Research Ltd 28th of May,2026 CC BY Abstract This paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling. The morpheme boundary—the junction between a stem and an affix—is treated as the primary object of measurement, and the distribution of the letters that may occupy each side of a junction is read across the corpus. Three results are established. First, the junction carries a stable structure: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a sparsity that scales inversely with a suffix’s productivity. Second, the relationship between any two affixes is quantified by the correlation of their boundary distributions, which measures the degree to which they share a stem population; this correlation ranges from ~0.99 for etymological doublets to 0.57 for productivity-asymmetric near-twins. Third, the written word decomposes into a consonant skeleton carrying lexical identity (53.0% of the vocabulary is uniquely recoverable from consonants alone) and a vowel tissue carrying grammatical form (1.8% recoverable from vowels alone), a ~29-fold information asymmetry. Multiple independent measures—consonant recoverability, derivational classmarking, and free-stem fraction—partition the lexicon at a single boundary separating a transparent Germanic core from a bound Latinate superstructure. The method, its corrections, and its limits are reported in full, and a program of remaining work is stated. 1. Corpus and Method All results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution. The corpus is the dwyl/english-words list (words_alpha.txt), comprising 455,246 alphabetic entries with a total of 4,254,354 letter occurrences and a mean word length of 9.345 letters. Each letter is assigned its ordinal value (A=1 through Z=26). Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter, so that each suffix position returns its exact ordinal value as a sanity check; the first stem letter preceding the suffix—the linker—is then read as a full A–Z frequency distribution rather than as a mean. The governing methodological constraint is that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, and that every claim is backed by a count. Page 1 Method: a derived column applies the ordinal map; end-anchored COUNTIF and MID/CODE formulas extract suffix positions and the linker; per-letter tallies at the linker yield the boundary distribution. Nesting among suffixes (for example ‑MENT within ‑ENT, or ‑ATION within ‑TION within ‑ION) is controlled by excluding longer relatives before counting. No statistical software, machine-learning library, or external lexical database is used at any stage of the core analysis. The constraint that the analysis remain computable by elementary means is not incidental; it ensures every figure in this paper can be independently reconstructed from the public corpus with a spreadsheet alone. 2. The Junction and Its Backbone The investigation began as a survey of unconditioned letter bigrams, which returned no morphological signal; structure appeared only when letter distributions were conditioned on a single morpheme boundary, and that conditioning is the method’s foundation. A preliminary tabulation of adjacent letter pairs across the corpus—which letters follow which, without regard to position within the word—yielded frequency patterns reducible to general orthographic regularities and carrying no isolable morphological content. The signal emerged only when a specific suffix was fixed and the letters preceding it were read as a population: conditioning on the junction, rather than measuring adjacency as such, is what renders the boundary structure visible. The set of letters that may legally occupy the stem side of a morpheme boundary then proves narrow, structured, and stable across suffixes. Reading the linker distribution across the mapped suffix inventory reveals a three-tier structure. A backbone of six letters—T, L, N, R, S, and I—is admitted at high frequency by nearly every suffix; T is the single most frequent linker across the inventory and recurs as the dominant boundary letter in suffix after suffix. An absolute floor of two letters—J and Q—is admitted by no suffix at measurable frequency, a prohibition confirmed corpus-wide and consistent with their status as the two rarest letters overall (J at 0.18%, Q at 0.19% of all letter occurrences). Between backbone and floor lies a selective middle whose composition varies by suffix and in which each suffix’s identity resides. Two descriptive terms are used throughout. Because each letter carries an ordinal value (A=1 through Z=26), a linker distribution may be summarized by where its mass falls on that scale: a distribution concentrated on early-alphabet letters (low ordinal values, A through roughly M) is termed cold, and one concentrated on late-alphabet letters (high ordinal values, roughly N through Z) is termed warm. The terms refer solely to ordinal position on the A–Z scale and carry no semantic content; the backbone letters, for instance, span both ends (cold I and L against warm N, R, S, T). The degree of selectivity is itself a measurement. The count of forbidden letters at a junction—its sparsity—scales inversely with the suffix’s productivity: derivational suffixes that attach choosily to a constrained stem class forbid many letters, whereas inflectional or highly productive suffixes forbid few. Sparsity is therefore not noise but signal: the pattern of exclusion characterizes the suffix as informatively as the pattern of admission. Page 2 Method: forbidden-letter counts are read directly from the linker distribution (frequency below 0.5% taken as floor); cross-checks against COUNTIF totals confirm the populations; J/Q absence is verified against the bare ‑S population of 84,208 words and against corpus-wide letter frequencies. The junction thus behaves as a filter whose admitted and forbidden letters together encode the morphological role of the boundary, with the consonant backbone bearing the structural load and the rare letters marking its limits.
This research explores the historical emergence of linguistic terminology in three languages—English, Uzbek, and Karakalpak—with special attention to the role of Latin, Greek, and Arabic heritage. It traces how borrowed concepts were nativized and localized in each linguistic setting. By juxtaposing five evolutionary stages in English with analogous processes in Uzbek and Karakalpak, the paper illustrates the interplay between international scholarly traditions and indigenous linguistic norms. The conclusions highlight both universal tendencies and language-specific particularities in the growth of terminological systems.
<div> This paper examines register variation in Latin from the third century BCE to the fourteenth century CE using Key Feature Analysis (KFA) (Egbert and Biber, 2023), a quantitative method for identifying statistically over-and underrepresented linguistic features. Registers are defined as text varieties linked to communicative situations and characterized by distributions of lexico-grammatical features (Biber, 1988, 1995). Six dependency-parsed Universal Dependencies (UD) treebanks are classified a priori into ten register categories based on established scholarship. Additionally, Principal Component Analysis (PCA) is used to reduce dimensionality, in order to explore the texts major patterns of variation and clusters of linguistically similar texts. KFA reveals systematic register-specific grammatical profiles consistent with previous research (e.g. Biber (2014b)). Registers with involved language use (e.g. letters and speeches) show higher frequencies of personal reference and engagement, while philosophical texts favor subordination and impersonal constructions. Registers containing narrative elements (e.g. historiography, satire) contain high frequency of verbs in past tense. PCA places the charter register into a distinct cluster, while other registers form more closely grouped patterns. The strongest components reflect contrasts in number, aspect, tense and person, alongside subordination and cordination distributions. The results are largely confirmatory: KFA produces coherent and interpretable groupings of grammatical features consistent with previous findings, providing a proof of concept for quantitative register analysis in historical corpora. Data and code are openly available for future research. </div>
This article scientifically analyzes the development trends of linguistics and contemporary linguistic problems in the context of digital transformation. The study examines the impact of artificial intelligence, corpus linguistics, natural language processing technologies, digital communication, and globalization on language development. Particular attention is paid to the role of the Uzbek language in the digital environment, terminology issues, language policy, linguistic identity, and transformations in digital education. The article argues that modern technologies not only expand the functional capabilities of language but also generate challenges related to linguistic norms, preservation of national languages, and cultural identity. The research is based on international scientific literature, statistical data, and нормативe legal documents.
Static concreteness ratings are widely used in NLP, yet a word's concreteness can shift with context, especially in figurative language such as metaphor, where common concrete nouns can take abstract interpretations. While such shifts are evident from context, it remains unclear how LLMs understand concreteness internally. We conduct a layer-wise and geometric analysis of LLM hidden representations across four model families, examining how models distinguish literal vs figurative uses of the same noun and how concreteness is organized in representation space. We find that LLMs separate literal and figurative usage in early layers, and that mid-to-late layers compress concreteness into a one-dimensional direction that is consistent across models. Finally, we show that this geometric structure is practically useful: a single concreteness direction supports efficient figurative-language classification and enables training-free steering of generation toward more literal or more figurative rewrites.
While language enables meaning, constituting knowledge in courts, schools, or parliaments, who gets to decide what can be known? Is meaning only use or a result of power too? Pitting Wittgenstein's forms of life against Foucault's regimes of discourse makes linguistic norms appear as instruments of exclusion. Marginalised speakers – subaltern, indigenous, and non-normative are often rendered unintelligible. Epistemic justice demands more than inclusion; it demands considering how rules are set, who enforces them, and how meaning is being contextually built. A discourse-sensitive, epistemic theory of justice is proposed, based on Kripke's rule-following paradox and Dijk's discourse analysis, to show that language is not neutral but a battleground of struggle over meaning, recognition, and epistemic authority.
This article explores the dynamics of linguistic norms in the Russian literary language within the context of social media communication. The relevance of the study lies in the fact that digital platforms create new communicative conditions that reshape traditional language norms. The paper analyzes variability, simplification, and expressive tendencies in syntax, lexicon, and orthography. Furthermore, it demonstrates that linguistic norms in social media are not static but dynamic and context-dependent. The findings suggest that these processes reflect not the degradation of the literary language, but its adaptation to digital discourse.
Ο γλωσσικός πόρος san-Corpus περιλαμβάνει σώμα κειμένων γραπτού λόγου της Νέας Ελληνικής, έκτασης περίπου 9 εκατομμυρίων λέξεων. Ο πόρος συγκροτήθηκε στο πλαίσιο διδακτορικής διατριβής, με στόχο τη μελέτη των συγκρίσεων ομοιότητας στη Νέα Ελληνική. Μέγεθος & Πηγές Το corpus περιλαμβάνει τρία ισομεγέθη υποσώματα, ώστε να επιτρέπονται οι συγκρίσεις μεταξύ τους: (α) Δημοσιογραφικός Λόγος: 7.774 άρθρα από τέσσερις διαδικτυακές εφημερίδες (Η ΑΥΓΗ, Η ΚΑΘΗΜΕΡΙΝΗ, ΕΘΝΟΣ, ΤΟ ΒΗΜΑ - έτος 2015). Συνολική Έκταση: 2,9 εκατ. λέξεις. (β) Εκπαιδευτικός Λόγος: 96 σχολικά εγχειρίδια δημοτικού και γυμνασίου. Συνολική Έκταση: 3,4 εκατ. λέξεις (μελετώνται 2,8 εκατ.). (γ) Λογοτεχνικός Λόγος: 28 μυθιστορήματα (βραβεία αναγνωσιμότητας περιόδου 2010-2015). Συνολική Έκταση: 2,5 εκατ. λέξεις. Κατανομή Σχολικών Εγχειριδίων ανά Γνωστικό Αντικείμενο Πλήθος Εγχειριδίων Ελληνική Λογοτεχνία 13 Ελληνική Γλώσσα 12 Ιστορία 9 Φυσική – Χημεία – Βιολογία 9 Μαθηματικά 9 Γεωγραφία – Γεωλογία – Περιβάλλον 8 Θρησκευτικά 7 Αγωγή Αισθητική (Εικαστικά – Μουσική – Θέατρο) 15 Αγωγή Υγείας (Φυσική Αγωγή – Οικιακή Οικονομία) 5 Πληροφορική – Τεχνολογία 5 Αγωγή Κοινωνική – Πολιτική 3 Αγωγή Σταδιοδρομίας (ΣΕΠ) 1 Κατάλογος μυθιστορημάτων: Συγγραφέας, Τίτλος Έτος 1ης έκδοσης Δούκα, Μάρω - Το δίκιο είναι ζόρικο πολύ 2010 Θέμελης, Νίκος - Η συμφωνία των ονείρων 2010 Καρυστιάνη, Ιωάννα - Τα σακιά 2010 Μιχαλοπούλου, Αμάντα - Πώς να κρυφτείς 2010 Ελευθερίου, Μάνος - Πριν απ' το ηλιοβασίλεμα 2011 Ζουργός, Ισίδωρος - Ανεμώλια 2011 Μακριδάκης, Γιάννης - Η άλωση της Κωνσταντίας 2011 Μπουραζοπούλου, Ιωάννα - Η ενοχή της αθωότητας 2011 Πανσέληνος, Αλέξης - Σκοτεινές επιγραφές 2011 Παπαδημητρίου, Χίλντα - Για μια χούφτα βινύλια 2011 Παπαθεοδώρου, Θοδωρής - Οι καιροί της μνήμης 2011 Τριανταφύλλου, Σώτη - Για την αγάπη της γεωμετρίας 2011 Φακίνος, Μιχάλης - Η έρημος έρχεται 2011 Βαμβουνάκη, Μάρω - Κυριακή απόγευμα στη Βιέννη 2012 Διβάνη, Λένα - Εγώ, ο Ζάχος Ζάχαρης 2012 Στεφανάκης, Δημήτρης - Φιλμ νουάρ 2012 Ακρίβος, Κώστας - Αλλάζει πουκάμισο το φίδι 2013 Ζέη, Άλκη - Με μολύβι φάμπερ νούμερο δύο 2013 Κορτώ, Αύγουστος - Το βιβλίο της Κατερίνας 2013 Κωνσταντούρου, Μαρία - Αγεφύρωτες σιωπές 2013 Μαντά, Λένα - Με λένε Ντάτα 2013 Ξανθούλης, Γιάννης - Κωνσταντινούπολη των ασεβών μου φόβων 2013 Ρώσση–Ζαΐρη, Ρένα - Άρωμα βανίλιας 2013 Ανδρουλάκης, Μίμης - Αλλέγκρα 2014 Δημουλίδου, Χρυσηίδα - Το κελάρι της ντροπής 2014 Παπαδοπούλου, Ελισάβετ - Μέρες και νύχτες που δεν ήταν δικές μας 2014 Χατζή, Αθηνά - Η θάλασσα έφυγε 2014 Χωμενίδης, Χρήστος - Νίκη 2014 Τεχνικές προδιαγραφές & Μορφότυπος Για την αναπαράσταση των δεδομένων και των μεταδεδομένων υιοθετήθηκε η πολυεπίπεδη οπτική των XML σχημάτων και τροποποιήθηκε το διεθνές πρότυπο TEI P5, 4.0.0 (Text Encoding Initiative). Δημιουργήθηκε ειδικός χώρος ονομάτων sanCorpus (sanC) με σχήμα τύπου RELAX-NG. Το σώμα κειμένων διατίθεται σε TXT και σε XML σε τρεις εκδοχές: Βάθος 0: απλό κείμενο (TXT). Περιλαμβάνει το main core (κείμενο βάσει του οποίου εξετάζονται οι συγκρίσεις ομοιότητας) και το out of core (κείμενο εκτός εμβέλειας της διατριβής, στο οποίο περιλαμβάνονται κείμενα που πλαισιώνουν το κυρίως κείμενο, π.χ. κείμενα διδασκαλίας, πίνακες περιεχομένων, εξώφυλλα) Βάθος 1 = κείμενα στην απλούστερη δυνατή XML κωδικοποίηση Βάθος 2 = κείμενα με πιο λεπτομερείς XML κωδικοποιήσεις. Αυτή η έκδοση (san-Corpus v1.0, Depth 0: Plain Text) περιλαμβάνει το σώμα κειμένων σε μορφή απλού κειμένου (Βάθος 0) στην αρχική του διάταξη (βλ. Επεξεργασία). Στόχος είναι ο σταδιακός εμπλουτισμός με επισημειωμένα δεδομένα, καθώς και με τις εκδοχές Βάθους 1 και 2. Επεξεργασία (Processing) Η μεθοδολογία συλλογής των δεδομένων, η θεωρητική τεκμηρίωση και το σχήμα επισημείωσης επεξηγούνται στις μελέτες Αφεντουλίδου (2022, 2021, 2013, 2012) και Afentoulidou (2009). Η πρώτη εκδοχή (Βάθος 0) χρησιμοποιήθηκε αποκλειστικά για τη λημματοποίηση που απαιτούσε η Collostruction Analysis (Gries, 2024). Στάδια επεξεργασίας (για τη λημματοποίηση): Τμηματοποίηση σε προτάσεις (sentence segmentation) με τη χρήση της βιβλιοθήκης Stanza (Stanford NLP Group, Qi et al. 2020), η οποία βασίζεται στο μοντέλο Greek Dependency Treebank (GDT) του Ινστιτούτου Επεξεργασίας του Λόγου / ΕΚ «Αθηνά». Τυχαία αναδιάταξη των προτάσεων για την προστασία της ακεραιότητας των πρωτότυπων έργων. Λημματοποίηση με τον ILSP Lemmatizer μέσω της Υποδομής Clarin-EL. Για την ανάλυση συμφράσεων του δείκτη σαν απομονώθηκαν συγκεκριμένοι λεκτικοί τύποι. Αδειοδότηση & Δικαιώματα Το san-Corpus συγκροτήθηκε για τις ανάγκες της διδακτορικής διατριβής και προστατεύεται από το δικαίωμα ειδικής φύσης σύμφωνα με την Οδηγία 96/9/ΕΟΚ και το άρθρο 45Α του Ν. 2121/1993. Η χρήση του περιεχομένου γίνεται αποκλειστικά για ερευνητικούς σκοπούς βάσει των εξαιρέσεων της Οδηγίας 2001/29 και της Οδηγίας (ΕΕ) 2019/790 (Text and Data Mining exceptions / Fair Use). Η πρόσβαση είναι περιορισμένη (Restricted Access) και παρέχεται αποκλειστικά σε μέλη της ακαδημαϊκής κοινότητας για σκοπούς επαλήθευσης των αποτελεσμάτων της διατριβής και περαιτέρω μη εμπορική έρευνα. Η πηγή προέλευσης δικαιούται να ζητήσει οποιαδήποτε τροποποιητική ενέργεια (π.χ. αφαίρεση) επί του πρωτότυπου περιεχομένου. Τέλος, η άδεια CC BY-NC-ND 4.0 ισχύει για την επιμέλεια (curation), τα μεταδεδομένα και τη γλωσσολογική επισημείωση του σώματος κειμένων. Βιβλιογραφικές αναφορές Αφεντουλίδου, Β. (2022). Σώμα ελληνικών κειμένων για τη μελέτη δομών ομοιότητας της Νέας Ελληνικής: σχεδιασμός και υλοποίηση. Στο Πρακτικά του 10ου Συνεδρίου Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας (σσ. 67-94). ΕΚΠΑ. Αφεντουλίδου, B. (2021). Δομές ομοιότητας στη Νέα Ελληνική. Σωματοκειμενικές παρατηρήσεις για τον πολυλειτουργικό δείκτη σαν. Προφορική ανακοίνωση στην 41η Ετήσια Συνάντηση του Τομέα Γλωσσολογίας, 13–15 Μαΐου 2021. ΑΠΘ. Αφεντουλίδου, Β. (2013). Και σου απάντησα κάτι σαν ‘τέλεια, εντάξει’. Δείκτης σαν + ευθύς λόγος;. Προφορική ανακοίνωση στο 7ο Συνέδριο Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας, 16–18 Μαΐου. ΕΚΠΑ. Αφεντουλίδου, Β. (2012). Συγκρίσεις ομοιότητας στα Νέα Ελληνικά: ο δείκτης σαν. Στο Z. Gavriilidou, A. Efthymiou, E. Thomadaki & P. Kambakis-Vougiouklis (Επιμ.), Selected papers of the 10th International Conference on Greek Linguistics (σσ. 696-707). DUTH. Afentoulidou, V. (2009). Sketching the σαν conditional construction in Modern Greek. Submitted essay, 2009 Linguistic Institute, Linguistic Structure and Language Ecologies, Linguistic Society of America and UC Berkeley. Gries, Stefan Th. 2024. Coll.analysis 4.1. A script for R to compute perform collostructional analyses. https://www.stgries.info/teaching/groningen/index.html Institute for Language and Speech Processing - Athena Research Center (2015). ILSP Lemmatizer. Version 1. [Software (Tool/Service)]. CLARIN:EL. http://hdl.handle.net/11500/ATHENA-0000-0000-23EE-D Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 101–108). Online: Association for Computational Linguistics.
This article examines the role and significance of language corpora and linguistic databases in linguistic expertise. It discusses the possibilities of conducting objective semantic, pragmatic, and stylistic analyses of texts through corpus-based methods. The study also highlights the contribution of linguistic databases and artificial intelligence technologies to improving the accuracy, reliability, and efficiency of expert conclusions. Furthermore, the relevance of developing specialized corpora and databases for forensic linguistics in Uzbekistan is substantiated.
The article examines „Georgian“ jokes as a speech genre in which the linguistic norm is violated through connotative reduplication and other forms of language play. Based on Bakhtin’s theory of speech genres and the study of humor, an anecdote is interpreted as a metalanguage practice reflecting social attitudes towards normality, deviation, and linguistic authority. The analysis is carried out in the context of the linguistically polyphonic urban space of Old Tiflis, where Russian historically served as the language of the empire, comparable to the role of French in Southeastern Europe. This sociolinguistic situation contributed to active language contact and the formation of hybrid speech forms, which later became a source of anecdotal texts. Special attention is paid to reduplication as a key mechanism for creating a comic effect in „Georgian“ jokes. Echo constructions such as shashlik-mashlyk, salad-malat, culture-multur do not have an independent denotation. However, they perform pragmatic and semiotic functions: they enhance the expressiveness of the utterance, mark irony, and refer to a conscious deviation from the norm. An analysis of Soviet and post-Soviet anecdotal empirical material, including the „Georgian“ typical characters Gogi and Givi, shows that reduplication serves to carnivalize the norm and symbolically explore the center-periphery relationship. At the end of the article, the author concludes that reduplication in „Georgian“ jokes is a key mechanism of the comic effect. The „echoes“ of main parts of words, such as multur, malat, and mashlik, lack independent denotation and serve to carnivalize linguistic norms and increase the expressiveness of speech. They assume knowledge of the normative code and turn deviation from the norm into a conscious language game and a form of collective language memory.
Understanding how memories of past experiences shape subjective feelings is complicated by the fact that we constantly update our memories. These updates are particularly impactful when individuals are reminded of emotionally positive or negative attributes of the original event. Yet, it remains unclear how such memory updating influences subjective feelings. Here, we investigated how the reactivation of emotional information affects episodic memory, subjective feelings, and their interaction. Across three experiments, participants first learned both positive and negative attributes associated with unfamiliar individuals. Then, they were reminded of a single positive or negative attribute for each individual to reactivate the memory partially. Finally, we reassessed memory for and subjective feelings about each individual’s attributes. In Experiments 1 and 2, these procedures were distributed across three days, while in Experiment 3, they occurred on a single day. Across these three experiments, reminding with negative attributes shifted subjective feelings in a negative direction. Reminded attributes were also better remembered, particularly for negative ones, and changes in subjective feelings were more strongly associated with reminded attributes. However, positive reminders only influenced subjective feelings to change positively when all procedures occurred on the same day. Together, these findings support a model in which memory updating shapes both episodic memory and emotional experience in a valence-dependent manner.
The paper presents a prototype of a web-app designed to automatically generate verb valency lexica based on the Universal Dependencies (UD) treebanks.It offers an overview of the structure of the app, its core functionality, and functional extensions designed to handle treebank-specific features.Besides, the paper highlights the limitations of the prototype and the potential of its further development.
Language is a living organism that evolves alongside technological and social advancements. This paper examines the phenomenon of neologisms—newly coined words or expressions—and their pervasive role in contemporary English mass media. The study categorizes recent neologisms based on their morphological formation processes, such as blending, compounding, and functional shift. Furthermore, it analyzes how mass media acts as a primary catalyst for the popularization of these terms. By investigating digital journals, social media platforms, and news broadcasts, the research highlights the pragmatic functions of neologisms in creating concise, engaging, and culturally relevant communication. The findings provide insights into the current trends of English lexicology and the impact of the digital age on linguistic norms.
ABSTRACT The widespread use of TikTok among elementary school students has brought noticeable changes to the way children communicate in their daily lives. The platform is no longer used merely as a source of digital entertainment, but has also begun to shape students’ word choices, speaking styles, and language habits. This condition can be observed among students at MIS Al-Khairaat Pombewe, who have become increasingly familiar with viral expressions, popular abbreviations, slang, and the mixing of Indonesian with foreign languages in everyday conversations. Such circumstances have raised concerns regarding the declining use of proper and standard Indonesian within the school environment. This study employed a descriptive qualitative approach involving the principal, teachers, and students selected purposively as research informants. Data were collected through observations, interviews, and documentation, then analyzed through the stages of data reduction, data presentation, and conclusion drawing. The findings reveal that TikTok exerts a dual influence on children’s language development. On the one hand, the platform contributes to vocabulary expansion, enhances students’ creativity in language use, and broadens their digital knowledge. On the other hand, the intensity of TikTok usage encourages the frequent use of informal language in formal situations, leading to a gradual decline in the use of proper Indonesian according to linguistic norms. Therefore, the involvement of teachers and parents is necessary to guide children toward wiser social media use without neglecting the development of their language abilities. ABSTRAK Fenomena penggunaan TikTok di lingkungan sekolah dasar memperlihatkan perubahan yang cukup nyata pada cara siswa berkomunikasi sehari-hari. Platform ini tidak lagi sekadar dimanfaatkan sebagai hiburan digital, tetapi turut membentuk pilihan kata, gaya berbicara, hingga kebiasaan berbahasa anak. Kondisi tersebut terlihat pada siswa MIS Al-Khairaat Pombewe yang semakin akrab dengan istilah viral, singkatan populer, bahasa gaul, serta pencampuran bahasa Indonesia dengan bahasa asing dalam percakapan mereka. Situasi ini memunculkan perhatian terhadap menurunnya penggunaan bahasa Indonesia yang baik dan benar di lingkungan sekolah. Kajian ini memanfaatkan pendekatan deskriptif kualitatif dengan melibatkan kepala sekolah, guru, dan siswa sebagai informan yang dipilih secara purposive. Informasi penelitian diperoleh melalui observasi, wawancara, dan dokumentasi, kemudian dipahami melalui tahapan reduksi data, penyajian data, dan penarikan kesimpulan. Temuan penelitian memperlihatkan bahwa TikTok memberi pengaruh ganda terhadap perkembangan bahasa anak. Di satu sisi, media sosial tersebut membantu siswa memperluas kosakata, meningkatkan kreativitas dalam berbahasa, dan memperkaya wawasan digital mereka. Di sisi lain, intensitas penggunaan TikTok ikut mendorong penggunaan bahasa informal dalam situasi formal sehingga kebiasaan menggunakan bahasa Indonesia sesuai kaidah menjadi semakin berkurang. Karena itu, keterlibatan guru dan orang tua dibutuhkan agar penggunaan media sosial dapat diarahkan secara lebih bijak tanpa mengabaikan perkembangan kemampuan berbahasa siswa.
Paper 6 (Silva 2026) introduced BPE Mean Vocabulary Morpheme Length (VMML) as a writing system classifier and showed that the Voynich Manuscript occupies a discriminant zone (VMML = 5.918, 95% CI 5.77-6.05) above all 15 tested alphabetic natural languages. This paper (v2.5) expands to 71 corpora across 40+ languages and reports six extended analyses: (1) Alphabetic ceiling confirmed at 5.76; (2) Tagalog (VMML=5.914) is the sole natural-language entry into the Voynich CI, but BC=0.202 distinguishes it from Voynich (BC=0.361); (3) Romanization inflates VMML by 2.4-5.3 units (methodological confound). Extended analyses: (4) Currier A vs B: delta VMML=+1.27, delta CBMI=+0.16 bits -- two quantifiably distinct writing registers; (5) BC coherent across all 7 manuscript sections (CV=6.7%) -- single writing system confirmed; (6) 3D discriminant (VMML x BC x CBMI): Voynich isolated, nearest natural-language neighbor Irish at distance 0.17; (7) Six named hoax mechanisms (monoalphabetic, Vigenere/barbavara, Vigenere/Italian-Knowles 2026, null insertion, syllabic compression, vocabulary shuffle) each fail all three criteria simultaneously; (8) BC orthogonal to all classical textual metrics (|r| < 0.23 vs entropy, TTR, hapax, Zipf) -- genuinely new structural dimension. All code and six extension scripts publicly available in companion repository. v2.3 (2026-06-08): Section 5.9 added - per-folio Currier A/B reanalysis using the Gaskell and Bowern (2022) canonical corpus (36,361 tokens, min_freq=5 BPE). Cross-boundary mutual information (CBMI) identified as primary discriminant: CBMI_A = 1.97 bits vs CBMI_B = 1.51 bits, Cohen d = -1.01, permutation p less than 0.001 (n = 10,000 shuffles, Bonferroni-corrected). CBMI survives within-quire control (pooled nA=46, nB=33; permutation p = 0.0008; Fisher combined within-quire p = 0.001), ruling out manuscript section as a confound. All three metrics (BC, BPE-ratio, CBMI) show A greater than B direction. Fisher combined full-corpus: chi-squared(6) = 40.66, p less than 0.000002. Section 5.1 corrected: direction is A greater than B on BC and CBMI. Finding is orthogonal to Parisel (2026) vowel-selection model. Conclusion 12 added. v2.4 (2026-06-09): §5.10 added — Currier-preserving null model (n = 200 iterations, size-matched) quantifying each metric's section-discrimination sensitivity independently of dialect. Key result: CBMI is the weakest section discriminant (mean |z| = 1.20 across six sections), confirming that the large CBMI A/B gap (§5.9) is not a section-composition artifact. STTR@100 is the strongest section discriminant (mean |z| = 4.75). Herbal section shows anomalously low vocabulary diversity (STTR z = -13.9); Stars shows anomalously high unique vocabulary (Hapax@500 z = +4.4). Demonstrates two independent organizational layers: CBMI tracks dialect, STTR tracks content domain. Conclusion #13 added. v2.5 (2026-06-10): Corpus expanded from 55 to 71 corpora across 40+ languages. §5.11 adds five medieval European corpora in native script via Universal Dependencies treebanks (Gothic transliteration, Old Church Slavonic, Old East Slavic, Ancient Greek PROIEL and Perseus; VMML 3.54-5.18 — all below alphabetic ceiling of 5.748). §5.12 adds 11 Australian Aboriginal language corpora via BibleNLP/eBible (Pama-Nyungan Western Desert, Ngumpin-Yapa, Arandic; Yolngu; Gunwinyguan; Daly; VMML 6.09-8.00 — predominantly above the Voynich zone). Warlpiri (VMML 5.851) is the sole near-entry on VMML but fails BC (0.233) and CBMI (0.244); 3D normalized distance from Voynich = 0.746 (vs. Irish = 0.200, the nearest neighbor from §5.4). Voynich zone is now charted on both sides: fusional alphabetic below (VMML 3.5-5.75), agglutinative-to-polysynthetic above (VMML 6.0-8.0). Voynich occupies a structural configuration not replicated by any of the 71 corpora tested. To our knowledge, this is the first systematic BPE profiling of Pama-Nyungan languages in the computational linguistics literature. Conclusions #14 and #15 added. v2.6 (2026-06-12): Section 5.10.1 adds a prose-only robustness check for the Section 5.10 Currier-preserving null model. Excluding all label, circular and radial loci (8.7% of tokens), every headline deviation survives essentially unchanged: Herbal STTR z = -13.5, Balneological z = -10.7, Stars Hapax z = +4.3; the sensitivity ranking is unchanged with CBMI last in both conditions. A mean-vs-median distributional note (both summaries rank lexical-diversity metrics first, boundary metrics last) and a coverage note (Astro/Zodiac folios carry no Currier tags and are outside any Currier-preserving design) are added. Erratum: Section 5.10 folio count corrected to 226 parsed / 196 Currier-labeled.
This paper presents a direct framework for sequence models with hidden states on closed subgroups of U(d). We use a minimal axiomatic setup and derive recurrent and transformer templates from a shared skeleton in which subgroup choice acts as a drop-in replacement for state space, tangent projection, and update map. We then specialize to O(d) and evaluate orthogonal-state RNN and transformer models on Tiny Shakespeare and Penn Treebank under parameter-matched settings. We also report a general linear-mixing extension in tangent space, which applies across subgroup choices and improves finite-budget performance in the current O(d) experiments.
Textbooks are fundamental educational tools that not only deliver curricular content but also convey societal and linguistic values. In the context of minority language education, textbooks have particular significance, as students’ language attitudes, identity, and self-perception are closely linked to the status and presentation of their native language. Language ideologies—often implicit beliefs about language and its use—shape how communities perceive linguistic norms, varieties, and speakers.
The Orthographic Junctions of English A Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and ConsonantVowel Information Asymmetry (p. 1) Boicho Dimitrov Temelakiev Saxon Ventura Research Ltd 28th of May, 2026 CC BY Abstract This paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling (p. 1). While the distinct roles of consonants and vowels in language processing are well-recognized psycholinguistically (p. 5), and the dual-stratum organization of English morphophonology is established theoretically (p. 6), this study provides an original, datadriven quantification of these properties directly at the orthographic level. The morpheme boundary—the orthographic junction between a stem and an affix—is treated as the primary object of measurement, reading the distribution of boundary characters across the corpus (p. 1). Three results are established: 1. 2. 3. The junction carries a stable, structured filter: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a distributional sparsity that scales inversely with an affix’s productivity (pp. 1-2). Affix relationships are structural: The relationship between any two affixes is quantified by the correlation of their boundary distributions, measuring their shared stem population (\(r \approx 0.99\) for etymological doublets down to \(r = 0.57\) for productivity-asymmetric near-twins) (pp. 1, 4). A massive information asymmetry partitions the lexicon: The written word decomposes into a invariant consonant skeleton carrying lexical identity (53.0% unique recoverability) and a mobile vowel tissue carrying grammatical form (1.8% unique recoverability) (pp. 1, 5). Multiple independent measures—consonant recoverability, derivational class-marking, and freestem fraction—converge on a single partition separating a transparent Germanic core from a bound Latinate superstructure (pp. 1, 6). The method, its corrections, and its limits are reported in full (p. 1). 1. Introduction and Method Traditional models of English morphophonology have long recognized that the lexicon is organized into distinct, historical strata—principally a native Germanic core and a bound Latinate superstructure (pp. 1, 6). Classic frameworks in generative phonology and lexical morphology, such as those pioneered by Chomsky and Halle (1968) and expanded by Kiparsky (1982), demonstrate that affixes of differing origins impose strict constraints on the phonetic and structural traits of the stems they recruit. Concurrently, cognitive and psycholinguistic research 1 has established a foundational "consonant-vowel functional asymmetry," demonstrating that human language processing systematically relies on consonants to preserve lexical and lexical-root identity, while vowels are dynamically manipulated to signal grammatical operations (Nespor et al., 2003; Bonatti et al., 2005). While these qualitative boundaries and cognitive patterns are deeply documented, this paper presents a mean-free, data-driven methodology to extract, quantify, and map these structural phenomena directly from corporate-scale English orthography without relying on heavy linguistic machinery. We introduce The Orthographic Junctions of English (OJE), an empirical approach that frames the morpheme boundary—the exact character interface between a stem and an affix—as an informational filter whose statistical properties reveal the historical, structural, and cognitive divisions of the vocabulary. All results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution (p. 1). The corpus utilized is the opensource dwyl/english-words repository (words_alpha.txt), comprising 455,246 alphabetic entries with a total of 4,254,354 letter occurrences and a mean word length of 9.345 letters (p. 1). Each letter is assigned its ordinal value (\(A=1\) through \(Z=26\)) (p. 1). Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter, so that each suffix position returns its exact ordinal value as a safety check (p. 1); the first stem letter preceding the suffix—the linker—is then read as a full A–Z frequency distribution rather than as a mean (p. 1). The governing methodological constraint is that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, and that every claim is backed by a precise empirical count (p. 1). By avoiding any dependency on complex machine-learning libraries or external lexical databases, the framework ensures that every architectural pattern discovered can be verified using standard data operations. 2. The Orthographic Junction and Its Backbone 2 The initial phase of this investigation examined unconditioned letter bigrams across the corpus, which yielded no meaningful morphological signal. Structural regularities appeared only when character distributions were explicitly conditioned on a single morpheme boundary—the character interface linking a stem to an affix. This structural conditioning serves as the foundation of the OJE framework. A preliminary tabulation of adjacent letter pairs across the corpus—measuring which letters follow which, without regard to structural position within the word—yielded baseline frequency patterns. These patterns are entirely reducible to general English orthographic constraints and carry no isolable morphological content. The structural signal emerged only when a specific suffix was fixed and the preceding characters were read as a discrete population. Conditioning on the junction, rather than measuring adjacency as a flat sequence, renders the underlying boundary constraints visible. The set of characters that legally occupy the stem side of a morpheme boundary proves narrow, highly structured, and remarkably stable across suffixes. Reading the linker distribution across the mapped suffix inventory reveals a highly stratified, three-tier architectural filter: A structural backbone of six letters—T, L, N, R, S, and I—is admitted at high frequency by nearly every suffix in the English lexicon. Within this backbone, T serves as the single most frequent linker across the inventory and recurs as the dominant boundary letter across independent suffixes. Conversely, an absolute floor of two letters—J and Q—is admitted by no suffix at a measurable frequency. This absolute prohibition is confirmed corpus-wide and is statistically consistent with their status as the two rarest letters in English orthography overall (with J accounting for 0.18% and Q for 0.19% of all letter occurrences). Between the backbone and the floor lies a selective middle whose specific character composition varies dynamically by affix, providing the distinct orthographic footprint wherein an individual suffix’s identity resides...
The practice of web form submission has emerged as a prime conduit for attackers, enabling them to infiltrate modern web applications and illegally harvest sensitive user data. Traditional defense mechanisms, such as static security reviews and server-side validation, are proving insufficiently agile for real-time detection of client-side vulnerabilities. This inadequacy arises directly from the rapid evolution of modern interfaces, which involves spontaneous DOM changes, dynamic element generation, and semantic interpretation that varies based on context and culture. This article presents an innovative browser extension framework that leverages a heuristic-based, multi-dimensional analytical engine combined with deep DOM inspection to identify insecure form submissions the moment they occur. The proposed methodology introduces five fundamental innovations: a contextual risk scoring system that models the complex interdependencies among form fields; an adaptive weighting scheme for risk patterns, accommodating diverse cultural and linguistic norms; a predictive vulnerability estimator that anticipates future threats; intelligent DOM mutation filtering designed to significantly optimize runtime performance; and cross linguistic semantic recognition to determine the true purpose of fields globally. Based on theoretical projections, this combined approach promises to enhance vulnerability detection accuracy while simultaneously reducing computational demands by approximately. Critically, all security analysis is executed exclusively on the user's local machine, guaranteeing privacy by ensuring no sensitive data is transmitted externally. A proof of concept application confirms the framework's practical feasibility and high efficacy for client side security assessment and catalyzing the development of flexible, scalable, and privacy respecting browser-based protections.
Prediction systems grounded in textual data have become indispensable across high-stakes domains including clinical decision support, financial signal detection, and digital misinformation analysis. Classical statistical approaches and shallow machine learning methods have demonstrated satisfactory performance on narrow, well-curated datasets, but they struggle to generalise once input distributions shift or domain vocabulary diverges from training corpora. Deep learning, and more specifically the pre-trained transformer paradigm, has substantially narrowed this gap; nevertheless, single-architecture solutions routinely leave accuracy on the table when applied to tasks that demand both rich contextual encoding and explicit sequential reasoning. This paper presents a cohesive, end-to-end AI- powered prediction framework that fuses BERT- derived contextual representations with a two-layer bidirectional LSTM (BiLSTM) classification head augmented by an additive attention mechanism. The system is designed as a modular pipeline: text acquisition and normalisation, augmentation-based imbalance handling, deep encoding, sequential modelling, and post-hoc probability calibration are treated as independent, replaceable stages. Experimental evaluation across three publicly available benchmark datasets — the LIAR fake news corpus, Stanford Sentiment Treebank v2, and a health-claim verification collection — confirms that the hybrid BERT-BiLSTM-Attention architecture outperforms five competitive baselines on macro-averaged F1 and area under the ROC curve. Ablation experiments quantify the individual contributions of the attention layer, recurrent head, augmentation strategy, and temperature scaling. A discussion of deployment trade-offs addresses inference latency, continual adaptation, and algorithmic fairness..
FrameNet is an English-based lexical database that shows how words are used by providing information as to which participants and relations are evoked by a certain concept. Recent efforts toward a multilingual FrameNet have not targeted either ancient languages or different historical stages of the same language. In our paper we propose creating a multilingual FrameNet for Ancient Indo-European languages starting with a set of 80 verb meanings annotated in the Pavia Verb Database. Our pilot study includes four verb meanings: RAIN, THUNDER, SEE, LOOK AT. As the adequacy of the semantic frames developed for English turns out not to be appropriate for the languages in our sample, we propose two new frames that can account for the analyzed data.
Part of speech and syntactically annotated dataset for modern Mongolian. The dataset is a fully annotated corpus of modern Mongolian texts written in Mongolian Cyrillic.
Abstract Introduction Sleep supports emotion regulation by preferentially consolidating emotional memories while attenuating reactivity. We have shown that dream recall plays an active role by increasing negative over neutral memories and reducing reactivity. In women, fluctuating reproductive hormones across the menstrual cycle influence sleep features implicated in emotional memory, yet whether menstrual phases influence how dreams shape emotional processing remains unknown. This study investigates how dreams shape sleep-dependent emotional processing across the menstrual cycle in naturally cycling women. Methods 128 women (Mage = 32.85 ±11.93 years) completed up to four visits across verified menstrual phases (menses, late-follicular, mid-luteal, late-luteal). At each visit, participants performed the Emotional Picture Task with negative and neutral IAPS images in the evening (Test 1) and the next morning (Test 2). Participants rated old/new, arousal, and valence of images shown at each test. Dream reports were collected upon waking prior to Test 2. Linear mixed-effects models tested main and interaction effects of menstrual phase and dream recall. Results The menstrual cycle altered how dreaming shaped overnight emotional memory. Dream recall typically benefited the emotional trade-off effect —favoring consolidation of negative relative to neutral images (Δd′; t(410)=1.95, p=0.05)—but this pattern reversed during the late-luteal phase (dream × menstrual cycle: t(381)=-2.29, p=0.02). Dreaming showed independent effects on emotional reactivity. Higher valence and arousal ratings for negative images during Test 1 predicted greater dream recall (valence: t(344)=2.05, p=0.04; arousal: t(327)=2.04, p=0.04). Additionally, the more negatively participants rated the images at Test 1, the more negative their dreams tended to be (t(166)=-2.11, p=0.04). Dream recall was linked to reduced next-morning emotional reactivity (valence: t(413)=-2.89, p=0.004; arousal: t(413)=-2.65, p=0.01), with stronger reductions following more negative dreams (β=0.15, t(182)=2.86, p=0.005). Conclusion Menstrual cycle phase influenced how dreams shaped overnight emotional memory. Negative waking experiences increased dream recall and shaped dream content—and recalling dreams, especially negative ones, reduced emotional reactivity and typically strengthened emotional memory—but this benefit disappeared in the late-luteal phase when there are declining reproductive hormones. These findings suggest a novel interaction between the menstrual cycle and dreaming, showing that hormonal fluctuations reshape how sleep and dreams regulate emotional experience and memory. Support (if any) RF1AG061355 (Baker/Mednick)
This paper presents a small-scale dependency treebank for Tunisian Arabic (TADT) developed within the Universal Dependencies framework, addressing the scarcity of linguistic resources for the Arabic varieties.The approach employs domain adaptation, leveraging a machine learning model (UDPipe 1.0) trained on Algerian Arabic data to annotate 100 Tunisian Arabic social media comments, followed by manual correction.This pilot study evaluates the feasibility of using machine learning-assisted annotation to scale resource development for spoken Arabic and identifies key challenges in cross-dialectal transfer for improving annotation quality and efficiency.This work contributes to more inclusive and fair representation of Arabic linguistic varieties in academic research and NLP applications.
This article examines the role of advertisements and signboards in shaping and reflecting public attitudes toward language. In modern society, linguistic culture is not only preserved in literature and education, but also manifested in everyday public texts such as commercial advertisements, street signs, shop names, and information boards. The study analyzes the linguistic quality of advertising texts, the influence of globalization on language use, and the social consequences of neglecting linguistic norms. Special attention is given to the relationship between language accuracy and cultural identity. The article also discusses the responsibility of businesses, media representatives, and educational institutions in maintaining linguistic standards in public communication.
Humans are inherently social beings, and social cues such as faces and voices guide attention and behavior. Auditory perception, especially binaural hearing, is essential for social cognition, enabling sound localization and speech comprehension in noisy environments. Deficits in auditory processing can impair social functioning, and conditions such as social anxiety are linked to reduced social functioning. Since social functioning is closely linked to overall well-being, improving social behavior represents a key objective in psychological research. Virtual reality (VR) is increasingly used to study social behavior due to its flexibility and ecological validity. However, users often report limited social presence, reducing the effectiveness of VR-based interventions especially for social anxiety. One reason may be the dominance of visual over auditory realism: audio is often presented in mono or stereo, reducing naturalness and presence. Binaural auralizations, which provide realistic, externalized spatial audio, may enhance presence and support virtual social interactions. This thesis pursues four main research objectives: identifying suitable behavioral and subjective measures for evaluating binaural realism; assessing immersion, realism, and audio quality across auralization techniques; comparing synthetic and natural speech in a socially stressful VR scenario; and examining effects of binaural audio on affect, presence, and attention under varying social stress levels. Study 1 examined how the virtual visual scene and measurement method affect localization and distance perception of physical sound sources. Across two experiments (N=60), audiovisual incongruence reduced localization accuracy but did not affect presence or realism. Distance estimation was influences by the interaction of task and scene: overestimation increased when using a placement task in a reduced-visibility scene. Study 2 compared localization accuracy for loudspeakers and four virtual audio renderings using a placement task and a gaze-based paradigm (N=49). Binaural renderings produced slightly lower localization accuracy but similar ratings of social presence and realism. A simple generic rendering performed as well as more complex ones. Only the anchor condition lacked externalization and was inferior across measures. Social presence and subjective realism were strongly correlated. Study 3 compared AI-generated text-to-speech with natural human speech in the Trier Social Stress Test (N=40). Both conditions elicited substantial stress responses and produced similar presence and affect ratings, demonstrating the practicality of synthetic speech in virtual social interactions. Study 4 investigated audiovisual realism in a virtual social stress scenario (N=78). A high-stress group showed stronger physiological and subjective stress responses than a low-stress group. Binaural audio increased perceived realism and externalization but did not affect social presence, stress responses, or gaze behavior. High arousal across all groups may have masked audio effects. Across all 4 studies, social anxiety did not consistently affect auditory perception or presence but influenced affective states and subjective evaluations of the interaction. Overall, the findings highlight the importance of VR-specific auditory perception and the role of acoustic immersion. Auditory realism enhances social and physical presence, though its impact varies by context. It appears most effective in low- to moderate-arousal scenarios and may be less critical in highly affective VR applications such as anxiety treatments. Practical advancesn such as TTS integration and simplified binaural rendering methods can support the broader use of realistic audiovisual VR environments in psychological research.
Abstract This paper examines the social contexts in Thomas Mann’s novel Buddenbrooks where the North German dialect Plattdeutsch is spoken. Beyond the technical challenges of translating these passages, the analysis focuses on the literary representation of code-switching that functions primarily as socially and symbolically charged act. Drawing on Pierre Bourdieu’s sociological theory, the study interprets deviations from linguistic norms and dialect use as instances of double negation – a strategy that appears to challenge social conventions but, in reality, affirms the most valuable social capital in classical bourgeois society: the certainty that one’s high status remains unthreatened. The difficulty of translating such passages stems from the specific cultural parameters embedded in the novel. Ultimately, the paper argues that culture is not an immediate given but requires analytical frameworks from the social sciences for proper understanding and interpretation.