Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper introduces an updated, publicly accessible version of the Film, Music and Emotion Dataset (FME-24), designed to examine how perceived emotion in film music evolves over time. It provides a comprehensive introduction to the dataset and explores its potential applications across music information retrieval (MIR), psychology, and AI-training contexts. The FME-24 dataset utilises film's immersive qualities to study emotional perception in a naturalistic yet controlled setting. It contains data from 275 film scores spanning the past two decades, including experimental and mainstream works. The dataset integrates high-quality film compositions with time-stamped valence-arousal (V-A) annotations, emotion sentences, familiarity ratings, and detailed metadata. 98 Participants contributed to annotating these temporal emotion features. For each time-stamped point, a two-second audio segment was analysed, and 78 features were extracted, including low-level timbral descriptors (MFCC statistics, spectral centroid), rhythmic descriptors (onset density, tempo), and higher-level psychoacoustic and tonal features (inharmonicity, roughness, chord transitions, tonal entropy). Although full audio files are unavailable due to licensing, reproducibility is ensured via ISRC codes, precise segment timings, and open access to all metadata and feature files in CSV format. The paper details the dataset's structure, annotation, and feature-extraction procedures, highlighting applications in computational and perceptual research and laying a foundation for future studies on emotion, perception, and narrative in film music.
The article highlights the necessity of adhering to the literary norms of the Ukrainian language in contemporary medical terminology, particularly the importance of understanding its lexical-grammatical and stylistic levels. The study analyzes term-lexemes that form paronymic relations, as well as the causes and consequences of the unmotivated use of paronyms in scientific discourse (semantic similarity, insufficient understanding of lexical meanings, speakers’ lack of competence). It is found that the erroneous use of paronyms leads to distortion of expressed meaning, linguistic paradoxes, and speech errors. Correct (normative) variants of paronymic medical terms appropriate for the professional language of healthcare practitioners are proposed. The research employs the methods of analysis and synthesis, comparative (contrastive) analysis, and the general linguistic method of scientific description. Conclusions. Quantitative and qualitative changes within the national terminological system result from the interaction of linguistic and extralinguistic factors and regularities. A thorough linguistic analysis of core paronymic pairs (series), their inclusion in lexical minima, and active instructional work with them constitute an important aspect of successfully mastering the language of medicine and improving the quality of specialized medical literature.
. The primary task of a dictionary is to collect, describe, and systematize the vocabulary of the Uzbek literary language. This process consists of determining the lexical structure of the language, defining the meaning and scope of use of words, as well as strengthening and stabilizing the norms of the literary language. The dictionary not only explains the meanings of words but also provides their correct pronunciation, grammatical forms, and stylistic features, which helps in the correct use of the language. At the same time, according to the requirements of the time, the dictionary makes a significant contribution to improving the culture of speech, as it teaches students and users the most important rules of the literary language.
This paper explores the systemic structure of the Nepali language through the lens of Bloomfield’s linguistic norms, grounded in the general characteristics of structural linguistics. It aims to explore how far Bloomfieldian norms are applicable to the way the Nepali language functions. This study is based on secondary sources, including books and articles available in the library and in open-access databases. The major findings of the research reveal that these structural principles are not confined to the English language alone; rather, every language possesses its own systematic sentence structure. Accordingly, Nepali also exhibits constituent relationships that conform to its own structural organization. This suggests that Bloomfield's concepts of endocentric and exocentric constructions are also applicable to Nepali sentence structures. In both English and Nepali, meaningful relationships exist between preceding and following constituents. These constituent relationships illustrate the hierarchical organization of sentence structure and the functional interdependence of its elements.
The study analyzes decorative texts from a linguistic-axiological perspective. The relevance of the research is explained by the wide spread of textualized objects of reality in the modern linguistic space. The aim is to analyze the axiological parameters of Russian decorative texts and identify the dominant values in the axiological sphere of the collective consciousness of contemporary Russian society. The study material included Russian decorative texts written on clothing, cars, bento cakes, gifts, disposable coffee cups, jewelry, and interior design. The value parameterization was conducted with the help of methods of linguisticaxiological interpretation, definition analysis, and conceptual analysis of key words. As a result, decorative texts have been proved to be a new form of the language on objects of reality, alongside with study and electronic media; such texts function as catalysts of value meanings in the modern linguistic space. I-mentality as a predominant model of self-identification of the linguistic personality has been identified. This is determined by the inherent self-presentational function of decorative texts and by the expansion of mass culture with self-promotion as its norm. It has been found that deliberate violation of linguistic norms distorts axiological norms and offsets value meanings, which are displaced by simulacra. The following axiological parameters of Russian-language decorative texts have been established: the importance of material values, a hedonistic world perception, and egoistic, assertive, and antisocial behavior. The prospects for the research lie in expanding the corpus of Russian decorative texts to enhance the objectivity of linguistic-axiological analysis and in studying the axiologemes of Russian decorative texts in the form of aphoristic statements.
The Orthographic Junctions of EnglishA Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and Consonant–Vowel Information AsymmetryBoicho Dimitrov Temelakiev · Saxon Ventura Research Ltd · CC BY · Version 3, June 2026Changes in this version (v3)This version corrects and strengthens the deposited record. (i) The corpus provenance is corrected: the analysed file is the dwyl `words.txt` list after cleaning, not `words_alpha.txt` as stated in v1–v2; the exact 455,246-entry corpus is deposited so the source is unambiguous. (ii) The descriptive "warm/cold" parameter (v1 §2/v2 §3) is withdrawn as a measurement-scale error: it performs arithmetic on the alphabet's ordinal position, which is a nominal code, and we show it carries no order-invariant signal. (iii) The associated "conjugation temperature" claim is restated without any appeal to alphabet position. (iv) All doublet correlations were recomputed on the deposited corpus under explicit nesting controls; the values are reported as measured, and one figure (−ANT/−ANCE) is corrected. (v) A permutation-null robustness test is added as a standing methodological control. No surviving result depends on the alphabet's ordering.AbstractThis paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling. The morpheme boundary — the orthographic junction between a stem and an affix — is treated as the primary object of measurement, and the distribution of the letters that may occupy the stem side of a junction is read across the corpus. Three results are established. First, the junction carries a stable, structured filter: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a sparsity that scales inversely with a suffix's productivity. Second, the relationship between any two affixes is quantified by the correlation of their boundary distributions, which measures the degree to which they share a stem population; this correlation ranges from ~0.97 for etymological doublets to 0.58 for productivity-asymmetric near-twins. Third, the written word decomposes into a consonant skeleton carrying lexical identity (53.0% of the vocabulary uniquely recoverable from consonants alone) and a vowel tissue carrying grammatical form (1.8% recoverable from vowels alone), a ~29-fold information asymmetry. Multiple independent measures — consonant recoverability, derivational class-marking, and free-stem fraction — partition the lexicon at a single boundary separating a transparent Germanic core from a bound Latinate superstructure. Every quantitative claim is tested for invariance under permutation of the alphabet, so that no result depends on the arbitrary ordering of the letters; the method, its corrections, and its limits are reported in full.1. Corpus and MethodAll results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution.The corpus derives from the public dwyl/english-words list (`words.txt`), from which non-alphabetic entries were removed and the remainder case-folded and deduplicated, yielding 455,246 unique alphabetic entries with 4,254,354 letter occurrences and a mean word length of 9.345 letters. Because the upstream list drifts over time, the exact 455,246-entry file analysed here is deposited with this record and is the corpus of record; a reader downloading the live upstream list today will not recover the same entry count.Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter; the first stem letter preceding the suffix — the linker — is read as a full per-letter (A–Z) frequency distribution rather than as a mean. Each letter may be referred to by its position in the alphabet purely as a label; no quantity in this paper depends on treating that position as a number (see §9). The governing methodological constraints are that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, that every claim is backed by a count, and — new in this version — that every claim is invariant under relabelling of the alphabet.Nesting among suffixes (for example −MENT within −ENT, or −ATION within −TION within −ION) is controlled by excluding longer relatives before counting. No statistical software, machine-learning library, or external lexical database is used at any stage of the core analysis; every figure can be reconstructed from the deposited corpus with a spreadsheet alone.2. The Junction and Its BackboneThe investigation began as a survey of unconditioned letter bigrams, which returned no morphological signal; structure appeared only when letter distributions were conditioned on a single morpheme boundary, and that conditioning is the method's foundation.Conditioning character statistics on a morpheme boundary places this work within the successor-variety tradition of boundary detection introduced by Harris (1955) and first implemented computationally by Hafer and Weiss (1974). That tradition uses transitional letter predictability to segment words into morphemes; the present method inverts the emphasis, holding a known boundary fixed and characterizing the distribution of stem-side letters it admits — a characterization of the junction rather than a segmentation of the word.A preliminary tabulation of adjacent letter pairs across the corpus — which letters follow which, without regard to position within the word — yielded frequency patterns reducible to general orthographic regularities and carrying no isolable morphological content. The signal emerged only when a specific suffix was fixed and the letters preceding it were read as a population. The set of letters that may legally occupy the stem side of a morpheme boundary then proves narrow, structured, and stable across suffixes.Reading the linker distribution across the mapped suffix inventory reveals a three-tier structure. A backbone of six letters — T, L, N, R, S, and I — is admitted at high frequency by nearly every suffix; T is the single most frequent linker across the inventory and recurs as the dominant boundary letter in suffix after suffix. An absolute floor of two letters — J and Q — is admitted by no suffix at measurable frequency, a prohibition confirmed corpus-wide and consistent with their status as the two rarest letters overall (J at 0.18%, Q at 0.19% of all letter occurrences). Between backbone and floor lies a selective middle whose composition varies by suffix and in which each suffix's identity resides. (The backbone, floor, and selective middle are stated as sets of letters; nothing in the three-tier description depends on the order of the alphabet, and all of it is invariant under the permutation test of §9.)The degree of selectivity is itself a measurement. The count of forbidden letters at a junction — its sparsity — scales inversely with the suffix's productivity: derivational suffixes that attach choosily to a constrained stem class forbid many letters, whereas inflectional or highly productive suffixes forbid few. Sparsity is therefore not noise but signal: the pattern of exclusion characterizes the suffix as informatively as the pattern of admission.3. The Suffix AtlasSuffixes are described by the distribution of letters that survive at their boundary — read in full, never reduced to a mean. The corrective lesson is explicit and was learned in this program: −NESS was first misjudged from a summary statistic and only described correctly once its full distribution was read (E 27%, D 15%, S 15%, I 13%). The general rule that follows is that a junction must be read as a distribution, not a single number, and in particular not as a mean of letter positions — a point developed formally in §9.Two representative distributions illustrate the contrast between a concentrated and a broad boundary: −ABLE (T-led and broad) and −IBLE (a sparse, frozen Latinate boundary). Both are reported as per-letter frequencies; the comparison between them is made by correlation (§4), which is invariant under relabelling of the letters.Table 1. Linker distribution of −ABLE (n = 4,694). Letters ≥3% shown.Linker Count PercentT 846 18.0%R 513 10.9%N 420 9.0%E 392 8.4%S 330 7.0%I 292 6.2%D 286 6.1%L 286 6.1% Table 2. Linker distribution of −IBLE (n = 738). Five letters carry ~90% of the population.Linker Count PercentS 248 33.7%T 208 28.3%C 86 11.7%D 63 8.6%G 59 8.0%N 20 2.7% Read by distribution, the suffix inventory resolves into a small set of boundary shapes, each fixed by the population of stems the suffix recruits rather than by its function, origin, or spelling.4. The Doublet Principle, QuantifiedTwo suffixes that draw on the same stem population share the same boundary distribution, and the correlation between their distributions measures the extent of that shared population directly. Because correlation is computed component-by-component over the same set of letters in both vectors, it is invariant under any relabelling of the alphabet — it is one of the order-independent quantities the program now requires (§9).All correlations below were recomputed for this version on the deposited corpus under explicit nesting controls (Pearson r over the 26-component per-letter frequency vectors). Across three classes of suffix pairs the boundary correlation forms an interpretable gradient.Etymological doublets — the same Latin stem class in two guises — correlate highly: −ENT/−ENCE at r = 0.97 (nesting-controlled, excluding −MENT from the −ENT population) and −ANT/−ANCE at r = 0.94. The −ANT/−ANCE value is corrected here: earlier versions reported r ≈ 0.99, which does not reproduc
The Ledger of Meluhha: Indus Valley Script as Metrological Accounting Code Rajeshkumar Venugopal (Third Buyer Advisory LLC, Michigan; ORCID 0009-0002-1838-5976). Version 3.0, 3 May 2026. BSD-2-Clause for human use; AI ingestion / training / fine-tuning / RAG / inference prohibited under contract law (see ai.txt and the §For Journalists appendix in the book). This work argues that the Indus Valley script is a cargo-tag accounting system rather than a phonetic writing system. The five-field record schema (merchant mark, commodity, weight tier, quantity, route terminal) is recoverable from existing archaeological evidence: the Harappan binary-and-decimal weight series standardised to 0.5 percent precision across roughly one million square kilometres, the Akkadian cuneiform Meluhha import receipts from Ur, the morphological-parallel correspondences between Indus seals and Tamil Nadu Iron Age potsherds, and the bigram structure of mapped versus unmapped signs in the digitised CISI corpus. The hypothesis is strictly weaker than any phonetic decipherment: it does not claim the Indus people did not have language, and does not assert which language was spoken. It claims that the function of the seals was inventory rather than speech encoding, and that the apparent untranslatability of the script reflects this functional fact rather than the absence of structure. The bridge to phonetic content, where it exists, runs through the proto-Dravidian numeral system reproduced from Wells 2015 Table 6.1 (after McAlpin 1981) — sign polyvalence is constrained by the morphology of numerals already in use, not by free phonetic association. The book is 76 pages, organised into 26 sections plus an appendix for journalists. The §For Journalists appendix provides a 10-minute verification protocol that requires only the SQLite command-line tool: any quantitative claim in the book is reproducible from the indus_corpus.db file in this archive by running a single SELECT statement against the named source_code. Every numeric claim in the book is traceable to a row in the database with explicit source attribution. The corpus database (indus_corpus.db) integrates ten primary sources: Mahadevan 1977 (concordance histogram); Joshi-Parpola 1987 Vol.1 Collections in India (the canonical photographic corpus, 862 pages OCR'd via Tesseract into 1399 unique artefact identifiers across six site prefixes); the mayig CISI digitisation (Mohenjo-daro subset, 179 inscriptions with 1003 sign occurrences); Wells 2015 The Archaeology and Epigraphy of Indus Writing (sign-role classifications including the canonical ICTM identification of signs 1, 2, 60; the Harappa volumetric system VI=40.4L through VIIIIIII=283L from Table 4.2; the proto-Dravidian numeral system from Table 6.1; sign 700 by NUM right-adjacency frequencies from Table 5.3); Fuls 2019 ICIT documentation PDF (16 sign-function codes including TMK / ITM / NUM / SYL plus 10 Wells-numbered ICIT inscriptions and the TMK by TMK adjacency matrix); Fuls 2022 Corpus of Indus Inscriptions metadata scaffold (73 sites with book-page anchors plus the 33-row artefact typology TAB / POT / SEAL / TAG and the corpus headline statistics 4660 artefacts / 5644 texts / 19831 sign occurrences); the Tamil Treebank logo-syllabic proxy corpus; the Rajan-Sivanantham Tamil Nadu inscribed-potsherd corpora (RS2025 and RS2026 for 13 sites). The codebook database (indus_codebook.db) contains the 28-entry sign-role mapping plus 8 commodities, 5 routes, 11 weight tiers, 6 quantity codes, 5 merchant marks, and 3 positional rules. The LSSC database (indus_lssc.db) contains the Latent Structural State Contraction analysis used to support the closure-versus-option sign classification. The interactive HTML dashboard (ledger-of-meluhha.html) is a single-file visualisation that renders the entire trade network on Leaflet plus 14 panels of derived data on the dropped indus_corpus.db file. It requires no server, no build step, and no installation — only a modern browser. The dashboard surfaces the original three-column display (decoded seals, trade-network map, frequency / weight / commodity charts) plus an Extended Data section with eleven panels covering the 10 sources, the 21 sign-function codes, the 7-row Harappa volumetric system, the 11-row proto-Dravidian numeral table, the 10 real Wells-numbered ICIT inscriptions, the 97 sign-role assignments, the 27 documented sign-pair frequencies, the 4 paradigmatic sign clusters, the CISI Vol.1 OCR coverage by site prefix, the FULS2022 reading-direction statistics, and the 14-entry bibliography with click-to-copy citation keys. Falsification criteria are stated explicitly in §The Falsification (Section V of the book): a long inscription with grammatical repetition characteristic of natural language; high-frequency signs in fixed ratios uncorrelated with commodity categories at multi-site stratigraphic analysis; a bilingual mapping Indus sign sequences to phonetic readings of a known language without metrological content; seal sign distributions at Mesopotamian findspots identical to the full Harappan corpus (no export-specific bias); the South-route terminal sign M063 appearing at Mohenjo-daro at rates comparable to other terminal signs. None have been observed; the M063 absence is now corroborated across three independent sources (mayig 179 corpus, Wells 2015 Appendix II Terminal Marker enumeration, Fuls 2019 ICIT help PDF Terminal Marker matrix). Contents of this archive: ledger_of_meluhha.pdf (76-page book); ledger-of-meluhha.html (interactive dashboard); databases.zip (indus_corpus.db plus indus_codebook.db plus indus_lssc.db); description.txt (this file). Repository with full source, F# ingest scripts, Alloy specifications, and the complete revision history: https://github.com/chanakyan/ledger-of-meluhha. License posture: BSD-2-Clause for human use including academic citation, journalistic quotation, classroom use, critique, parody, and derivative works. AI ingestion / training / fine-tuning / retrieval-augmented generation / inference prohibited under contract law; this prohibition is not negotiable. Newsrooms using AI summarisation tools to process this work are creating legal exposure that human use does not. Contact: vrajeshkumar@gmail.com. ORCID 0009-0002-1838-5976. Third Buyer Advisory LLC, Michigan, United States.
This article analyzes lexical norms and their violations in media language through a comparison of Azerbaijani and English. Globalization and social media have increased lexical deviations, including misuse of foreign words, semantic inconsistencies, and slang. These issues reduce clarity and accuracy in communication. Strengthening editorial control and journalists’ language competence is essential. Keywords: lexical norm, media, digital, style, text, conceptual, global
This study explores the role of media linguistics in shaping the linguistic norms of contemporary mass media in Kazakhstan, an area that remains insufficiently examined in relation to bilingual media practices, genre variation, and digital communication platforms. It focuses on the influence of global communication trends, digital technologies, and national language policy on linguistic change. The research examines the introduction of English-language borrowings, their adaptation in Kazakh and Russian, and the modification of traditional linguistic structures. These shifts are driven by the widespread use of English-language platforms such as social media (Facebook, Instagram, Telegram), streaming services, and international content formats. The findings show that English loanwords are most common in advertising and news, where there is a need for rapid adaptation to global trends. In contrast, analytical and official publications adhere to traditional linguistic norms, highlighting a balance between formal and informal communication. The study also emphasizes the importance of Kazakhstan’s national language policy in preserving linguistic identity, with measures regulating foreign words in the media to develop sustainable language standards. Digital technologies also shape informal media discourse through memes, hashtags, and hybrid language forms. These changes are particularly evident among younger audiences, who adapt more quickly to language innovations, while older generations tend to be more critical of these shifts. In conclusion, media linguistics proves to be an effective tool for analyzing the intersection of globalization, national identity, and digital transformation, offering valuable insights into the adaptation of linguistic norms amid rapid technological development.
Pure speech language models aim to learn language directly from raw audio without textual resources. A key challenge is that discrete tokens from self-supervised speech encoders result in excessively long sequences, motivating recent work on syllable-like units. However, methods like Sylber and SyllableLM rely on intricate multi-stage training pipelines. We propose ZeroSyl, a simple training-free method to extract syllable boundaries and embeddings directly from a frozen WavLM model. Using L2 norms of features in WavLM's intermediate layers, ZeroSyl achieves competitive syllable segmentation performance. The resulting segments are mean-pooled, discretized using K-means, and used to train a language model. ZeroSyl outperforms prior syllabic tokenizers across lexical, syntactic, and narrative benchmarks. Scaling experiments show that while finer-grained units are beneficial for lexical tasks, our discovered syllabic units exhibit better scaling behavior for syntactic modeling.
Abstract This chapter explores how linguistically annotated corpora for Ancient Greek and Latin occupy a pivotal space between editions, language resources, and language technologies. Drawing on the author’s experience in two landmark projects (the Ancient Greek and Latin Dependency Treebank and LiLa: Linking Latin), it highlights how digital annotated editions and corpora not only encode linguistic information but also participate in textual scholarship. The chapter discusses the history of the projects, the implications of adopting shared standards, and situates these developments in the broader context of the semantic web. Far from being merely technical artifacts, language resources emerge as tools for interpretation and learning. Ultimately, they embody what the critical edition has long represented: a form through which philologists articulate their vision of language and its relationship to the present.
This article traces the historical policing of pronunciation in eighteenth- and nineteenth-century Britain, with particular attention to regional variation and its entanglement with class, authority, and linguistic legitimacy. It begins by contextualizing how regional accents were socially charged, often eliciting responses of ridicule or exclusion. The analysis then turns to early prescriptive works by orthoepists such as Thomas Sheridan (1762 and 1780), William Kenrick (1783), and John Walker (1791), who helped construct a linguistic hierarchy in which ‘proper’ pronunciation was aligned with moral and social superiority. Building on these foundations, nineteenth-century newspapers and periodicals extended prescriptive ideologies to a wider public, naturalising linguistic norms through humour, commentary, and complaint. Drawing on a corpus of editorials, advertisements, and letters to the editor from sources including The Times, The Morning Chronicle, and others, the article highlights how the press functioned as both a conduit and creator of metadiscourses on speech. These texts reveal that standard language ideologies were not solely imposed from above, but were also taken up, reproduced, and contested by the reading public. In examining how pronunciation became a symbolic site for the performance of class identity and cultural legitimacy, the article underscores the long-standing role of accent in structuring social inclusion and exclusion. <br>
Non-binary bottom-up constituency parsing is usually taken to require arity actions: reductions such as \(\textsc{Reduce-}X\#k\) specify both the mother label and the number of children to be composed. We show that this arity parameter is not a necessary transition primitive. Our parser introduces constituent labels separately and recovers reduction spans from delimiter-bounded stack configurations. In a well-formed reduction configuration, arity is uniquely determined by the active delimiter and the label marker, making it a derived property of parser state rather than an action label. This factorization removes label--arity-specific reduce actions while preserving direct construction of original non-binary trees. Experiments on PTB and CTB show that the delimiter-guided parser remains competitive with an arity-specific bottom-up baseline under the same implementation framework, with substantially smaller action inventories. Analyses further show that its predicted arity profile remains close to the gold treebanks and that high-arity constituents do not collapse when arity actions are removed.
Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers -- the Biaffine LSTM, Stack-Pointer Network, AfroXLMR-large, and RemBERT -- across ten typologically diverse languages, with a focus on low-resource African languages. We find that the Biaffine LSTM consistently outperforms transformer models in low-resource regimes, with transformers recovering their advantage as training data increases. The crossover falls within a resource range typical of treebanks for under-resourced languages. Morphological complexity (measured via MATTR) emerges as a significant secondary predictor of transformers' relative disadvantage after controlling for corpus size. These results indicate that the Biaffine LSTM may be better suited for syntactic tool development in low-resource regimes until sufficient annotated data is available to leverage the representational capacity of pre-trained transformers.
This article examines the concept of occasional words (occasionalisms) and their linguistic features within modern discourse. It explores their definitions, structural characteristics, and functional roles from morphological, semantic, pragmatic, and stylistic perspectives. The study highlights the context-dependent nature of occasionalisms, emphasizing their expressive and evaluative functions as products of individual linguistic creativity. Special attention is given to their formation through analogy, their role in literary and journalistic texts, and their position between language norm and speech innovation. The article also differentiates occasionalisms from neologisms and nonce formations, focusing on their degree of lexicalization and communicative purpose. The findings demonstrate that occasional words are not random deviations but systematic, rule-governed innovations that contribute to linguistic development and reflect the dynamic interaction between language structure and usage.
Abstract Efficient emotion regulation is central to everyday functioning, plays a key role in mental health, and is implicated in vulnerability to affective disorders. While many strategies can be used to regulate emotion, direct comparisons of their effects on subjective affect are often limited by differences in study design, stimulus material, and control conditions. To address these limitations, the present study compared four widely studied emotion regulation strategies—distraction, reappraisal, distancing, and suppression—within a single paradigm. The study focused on three central questions: (i) whether antecedent-focused strategies (distraction, reappraisal, distancing) are more effective than suppression in down-regulating negative affect; (ii) whether applying these strategies to neutral stimuli alters affective experience; and (iii) whether baseline emotional reactivity relates to variation in regulation success. Fifty-six healthy young adults viewed negative and neutral images either passively or while regulating their emotional responses and subsequently provided subjective valence and arousal ratings. All participants regulated negative images using each strategy; a subsample additionally applied the strategies to neutral images to quantify strategy effects independently of responses elicited by negative stimuli. All strategies reduced negative affect relative to passive viewing, and antecedent-focused strategies were more effective than suppression. Reappraisal produced the largest improvement in valence, whereas antecedent-focused strategies showed comparable reductions in arousal. Regulation also altered responses to neutral stimuli: reappraisal selectively increased reported valence, whereas distraction and distancing reduced arousal. Individual-difference analyses indicated that regulation success for antecedent-focused strategies did not vary with emotional reactivity, whereas higher reactivity on the valence dimension was associated with lower regulation success in suppression. Together, these findings support a process-model distinction between antecedent- and response-focused regulation and show that regulation strategies can shape affective experience even in the absence of negative emotional content.
This article presents a comparative analysis of the linguistic, stylistic, and structural features of official document texts in the Uzbek and Turkish languages. The study examines state documents, official correspondence, directives, orders, contracts, applications, and decisions as illustrative materials. The terminological systems, syntactic structures, lexical units, and stylistic norms of document texts in both languages are identified, and their similarities and differences are analyzed on a scientific basis. The results reveal the common Turkic roots of Uzbek and Turkish document language, contemporary trends in the development of official style, and factors related to national language policy.
Materials The selected pictures (248 in total: 120 neutral and 128 negative) were drawn from five different standardized affective picture sets (see Table 1 for details on which images were selected from each set). Participants were recruited from Utrecht University via campus advertisements. The Dutch sample consisted of 103 participants (90 females), all native Dutch speakers, with a mean age of 20.93 years (SD = 2.42). Their responses provide the arousal and valence ratings with these 248 images. Task Design This is a emotion rating test conducted online via the Gorilla platform. Each trial began with a 1500-ms fixation, followed by a 2000-ms presentation of the target image. After a 1000-ms delay, the SAM scales appeared and remained on screen until the participant responded, with arousal rated first and valence second (Figure 1). DATA · The folder titled “Open_Data_Dutch_Norm_Data” contains an Excel files named “V1_Dutch_norm_data_arousal_pp103” and “V1_Dutch_norm_data_valence_pp103”, which includes raw data on arousal and valence, along with their response times during the rating task, respectively. · The same columns in both files are: Subject, Gender, Image, Scale, Picture_Valence, Response, Response_Duration. · Within-subject factor: Picture_Valence (negative, neutral)
Abstract – The linguistic situation in Kazakhstan has been shaped by historical and socio-political factors. As a result, Kazakh–Russian bilingualism has developed in the country. This type of bilingualism is predominantly characterized as semi-dominant. Under such conditions, the functioning of the Kazakh language as the state language remains a highly relevant issue. Although Kazakh has been granted the status of the state language by law, it is still not fully used across many domains of everyday communication. This situation may contribute to the weakening of the native language among Russian-speaking Kazakh youth. This article examines the phenomenon of interference observed in the speech production of Russian-speaking Kazakh students who are acquiring Kazakh as a second language from cognitive and psycholinguistic perspectives. The research focuses on students whose native language is Kazakh but who received their secondary education in Russian and predominantly use Russian in higher education, social environments, and the information space. The participants in the study use Kazakh within a limited functional range, mainly in family communication, during Kazakh language classes, or in everyday бытовые situations. The article describes interference not as a linguistic error or a deviation from linguistic norms, but as a natural cognitive adaptation mechanism of the bilingual mind and as a phenomenon arising from the interaction of the semantic and conceptual systems of two languages. The study analyzes the role of cognitive mechanisms such as associative transfer, the literal rendering of figurative meanings, conceptual overlapping, and frame shifting in the formation of interference. The article also addresses the issue of language attrition in Kazakh under conditions of semi-dominant Kazakh–Russian bilingualism in Kazakhstan. The main objective of the study is to identify recurrent linguistic deviations in the speech of Russian-speaking Kazakh youth. In addition, the study aims to distinguish these deviations from interference-related errors and to determine whether they represent manifestations of language attrition. The research employs cognitive-interpretative, discourse, comparative, and pragmatic methods of analysis. The findings reveal that interference among the studied group of students manifests itself at the lexical-semantic, syntactic, and pragmatic levels of language. The results of the study demonstrate that interference is a natural cognitive process activated during the formation of new linguistic experience in a bilingual individual. At the same time, the study identified a tendency toward the stabilization of interference-related forms in both the oral and written speech of the students. Such a phenomenon may lead to a decline in the active use of national-cultural content, phraseological resources, and natural usage patterns of the Kazakh language, thereby increasing the risk of language attrition. Therefore, the findings suggest that interference in Kazakh language teaching should be viewed not merely as an error requiring correction, but also as an important indicator reflecting the cognitive developmental characteristics of language learners.
Chinese Treebank 7.0, Linguistic Data Consortium (LDC) catalog number LDC2010T07 and isbn 1-58563-542-1, consists of over one million words of annotated and parsed text from Chinese newswire, magazine news, various broadcast news and broadcast conversation programs, web newsgroups and weblogs.
A Universal Dependencies treebank for Old Georgian, containing morphosyntactic annotations of the Book of Ezra (Oshki Bible) in CoNLL-U format. We are grateful to Dr. Paul Meurer for his discussion and guidance and to Prof. Dr. Jost Gippert for generously making the Old Georgian corpus available online on TITUS (https://titus.uni-frankfurt.de/texte/etcs/cauc/ageo/at/oskijer/oskij.htm)
ToS-100 contains 100 Terms of Service documents from online platforms, splitinto 20,417 clauses. Each clause is annotated for five categories of potentialunfairness: arbitration (A), unilateral change (CH), content removal (CR),limitation of liability (LTD) and unilateral termination (TER). A clause maycarry none of them (18,843 clauses, 92.3%) or several at once, so the task ismulti-label rather than multi-class. This deposit is the corpus of Ruggeri et al. (2022), itself built on the ToScorpus of Lippi et al. (2019), redistributed as three CSV files for Assignment 1of the Natural Language Processing course (Prof. Paolo Torroni, University ofBologna, a.y. 2026-2027). Files: train.csv 80 documents, 15,837 clauses validation.csv 10 documents, 2,548 clauses test.csv 10 documents, 2,032 clauses Columns: document_ID, document, text, A, CH, CR, LTD, TER. The split is by document, not by clause: all clauses of one contract stay inthe same split. Clauses of a single contract repeat each other almost verbatim,so a clause-level split would leak the test set into training. Positive clauses per category (train / validation / test): A 75 / 20 / 11 CH 268 / 43 / 33 CR 165 / 32 / 19 LTD 504 / 75 / 47 TER 318 / 59 / 43 Text is lowercased and tokenized in Penn Treebank style: brackets appear as-lrb- / -rrb-, quotes as `` and '', and clitics are detached (mozilla 's,do n't). The knowledge-base columns of the original release (*_targets), which point tofree-text legal rationales rather than to spans over the clause, are notincluded. Please cite the original works: Lippi et al., 2019. CLAUDETTE: an Automated Detector of Potentially Unfair Clauses in Online Terms of Service. Artificial Intelligence and Law. Ruggeri et al., 2022. Detecting and Explaining Unfairness in Consumer Contracts through Memory Networks. Artificial Intelligence and Law.
Research on emotional perception often relies on 2D stimuli or highly affective images, limiting ecological validity. We introduce the Emotional Daily Life Library (E-DLL), a database of 132 rotating 3D everyday objects with comprehensive perceptual, cognitive, and emotional normative ratings. In 52 adults, participants provided dimensional (valence–arousal–dominance) and categorical emotion ratings, alongside assessments of recognition, naming, familiarity, contact, usage, and visual complexity. Personality traits (NEO-FFI) and depressive symptoms (BDI-II) were measured to examine individual differences. Cumulative Link Mixed Models (CLMMs) revealed that valence ratings were negatively influenced by the interaction of Neuroticism and subclinical depressive symptoms. For arousal, higher Neuroticism and Conscientiousness demonstrated marginal positive associations, while dominance ratings were unaffected. Generalized Linear Mixed Models (GLMMs) for categorical labels indicated that while emotional attributions were primarily driven by stimulus properties, traits such as Neuroticism and Extraversion significantly predicted the perception of negative emotions (e.g., Sadness). Spearman correlations identified interrelationships among cognitive and perceptual dimensions, and redundant variables (Contact and Usage) were combined into a composite Object Interaction score. By integrating neutral, immersive 3D stimuli with rich multidimensional annotations, E-DLL provides a controlled and ecologically valid tool for experimental and clinical research. Its applicability includes cognitive training, VR-based interventions, and personalized neurorehabilitation platforms such as NeuroAIreh@b, supporting investigations of affective biases and optimizing daily life–oriented therapeutic interventions. • Introduces E-DLL: 132 everyday 3D objects with cognitive & emotional ratings. • Neutral, low-arousal objects provide ecologically valid stimuli for clinical research. • Personality traits & subclinical depression symptoms modulate neutral object ratings. • Composite metrics and cognitive dimensions enhance object-level characterization. • Supports clinical applications like VR training and personalized interventions.
Author: Denny van Gulik Methodology: The 80/20 Matrix Status: Version 1.0 (Linguistic Corpus) 1. Abstract (Exposé) This work presents a comprehensive decoding of the Voynich Manuscript (MS 408). Moving beyond traditional cryptographic attempts, this research approaches the codex from a technical and structural perspective. The manuscript is identified as a functional pharmaceutical and balneological manual of a late medieval scholarly brotherhood, likely operating within a courtly or monastic context (Palar). The core of this discovery is the 80/20 Matrix: 80% Phonetically Deformed Latin: Technical terms of medieval botany and medicine, obscured through systematic phonetic shifts and the specific EVA character set. 20% Balkan Regionalisms: Use of regional terminology (e.g., Amum for water, Otlar for herbs, Pala for court/palace) serving as bridge vocabulary. Statistical validity is maintained across all 246 pages, identifying complex processes of thermal extraction (Pokedum), honey-based preservation (Melle), and advanced hydrotherapeutic systems. 2. Methodological Transparency (Authorship & AI Usage) Important Note on Research Genesis: The discovery of the 80/20 Matrix and the linguistic identification of the Balkan-Latin hybrid system is the exclusive intellectual property and original work of Denny van Gulik. Artificial Intelligence (specifically the Google Gemini model) was utilized strictly as a digital research assistant and scaling tool. Its role was limited to: Formatting manually decoded data into scientific tables and HTML. Cross-referencing author-identified word stems with linguistic databases. Translating research notes into academic English to facilitate international peer review. The logic, intuition, and systematic pattern recognition are entirely human-led. This project is not a result of "AI hallucination" but a rigorous analysis of the codex as a logistical document. 3. Project Roadmap & Updates Current Version (v1): Focuses on the textual corpus, the 80/20 linguistic matrix, and the primary glossary. Upcoming Version 2.0: Will feature fully integrated high-resolution folio images and direct visual cross-references for every analyzed page. English Edition: A full, 246-page English translation of the entire study is currently in progress to ensure accessibility for the global scientific community. 4. Keywords Voynich Manuscript, MS 408, 80/20 Matrix, Medieval Medicine, Balkan Linguistics, Codicology, Balneology, Historical Pharmacy. Deutsche Zusammenfassung: Dieses Projekt präsentiert die vollständige Dekodierung des Voynich-Manuskripts mittels der 80/20-Matrix (deformiertes Latein & Balkan-Regionalismen). Es identifiziert das Werk als pharmazeutisches Handbuch einer spätmittelalterlichen Bruderschaft. Version 1.0 sichert die linguistische Priorität; Version 2.0 mit Bildreferenzen sowie eine vollständige englische Übersetzung folgen in Kürze.
Abstract Impaired musical emotion perception is common in hearing loss, yet how reduced spectral audibility shapes cortical processing during emotion judgments remains unclear. Because alpha-band activity can index top-down control when sensory evidence is degraded, we examined whether spectral degradation increases the cognitive demand required to form stable affective judgments in music. Forty-eight healthy participants were divided into three groups: high-frequency hearing loss simulation (HF sim ), low-frequency hearing loss simulation (LF sim ), and normal hearing (NH). Participants rated the arousal and valence of filtered musical stimuli (happy, sad, neutral) during EEG recording. HF sim showed dimension- and context-dependent alpha modulations. In the happy condition, arousal ratings and alpha power were comparable across groups, whereas valence judgments showed behavioral differences and late-stage alpha increases in HF sim, consistent with reduced certainty when high-frequency cues supporting positive valence are degraded. In the sad condition, behavioral ratings were preserved, yet HF sim showed sustained alpha increases during arousal judgments, suggesting compensatory inhibitory-gating processes that may support stable appraisal under degraded listening. Overall, spectral degradation appears to elicit compensatory cognitive processing: alpha power increases index higher demands for happy-valence with reduced HF cues, and compensatory gating that maintains sad low-arousal appraisal.
The article aimed to describe the psychoanalytic structure of the character as a model of child self-regulation, realised through narrative interaction with norms, organisation of everyday life and a system of representations. The research methodology combined a close reading of the text, a procedural narrative analysis involving the coding of episodes according to the following scheme: source of the norm, form of control, action of the heroine, consequence, a comparative analysis of translations and a visual-narrative comparison. Six codes of self-regulation were identified during the analysis: inversion of authority; play as a way of cognition; comic redefinition; control of the environment through space and objects; everyday choice through food and sweets; and visual consolidation of autonomy. The analysis involved procedural coding of episodes, followed by a comparison of the translated variants and visual markers, and the generalisations are presented in analytical tables. The analysis distinguished six codes of self-regulation: inversion of authority; play as a method of cognition; comic redefinition; control of the environment through space and objects; everyday choice through food and sweets; and visual consolidation of autonomy. It was established that the “institution-child” conflict functions as a transfer of the rule from the sanctioned sphere to the testing sphere, while laughter creates a cognitive distance that reduces dependence on external evaluation. Five spatial scenarios of interaction with norms and four levels of representation (text, image, stage or screen, and critical description) were described. The stability of reference to the character was explained through a three-level model of identification conditions (core, attributes, and context). A comparison of translations showed that although lexical mitigation may alter the intensity with which rebellion is interpreted, it does not destroy the procedural core. The practical significance lies in the potential application of the proposed codes as an analytical framework for interpreting children’s narratives, comparing translations and describing adaptations within educational and editorial practices
We present EmoWork, a multimodal, multi-label dataset designed to support emotion and stress detection in realistic interpersonal work settings. Interpersonal work-common in occupations such as customer service-often requires workers to regulate their emotional expressions in response to strong affective stimuli. These demands, shaped by organizational display rules, present a unique challenge for affective computing systems, particularly in scenarios where internal emotional states diverge from observable behaviors. Despite this, no public datasets exist that capture such dynamics of affect in naturalistic settings. To address this gap, we collected physiological, behavioral, and self-reported data from call center workers who engaged in role-play scenarios simulating customer service interactions with professional actors portraying dissatisfied customers. The dataset includes self-reported affective ratings, which are used as labels for classification, synchronized recordings from three wearable devices (i.e., Polar H10, Empatica E4, and Muse S), and features extracted from video and audio data. The EmoWork dataset advances affective computing by offering context-rich, multimodal data grounded in realistic interpersonal work scenarios.
Ο γλωσσικός πόρος san-Corpus περιλαμβάνει σώμα κειμένων γραπτού λόγου της Νέας Ελληνικής, έκτασης περίπου 9 εκατομμυρίων λέξεων. Ο πόρος συγκροτήθηκε στο πλαίσιο διδακτορικής διατριβής, με στόχο τη μελέτη των συγκρίσεων ομοιότητας στη Νέα Ελληνική. Μέγεθος & Πηγές Το corpus περιλαμβάνει τρία ισομεγέθη υποσώματα, ώστε να επιτρέπονται οι συγκρίσεις μεταξύ τους: (α) Δημοσιογραφικός Λόγος: 7.774 άρθρα από τέσσερις διαδικτυακές εφημερίδες (Η ΑΥΓΗ, Η ΚΑΘΗΜΕΡΙΝΗ, ΕΘΝΟΣ, ΤΟ ΒΗΜΑ - έτος 2015). Συνολική Έκταση: 2,9 εκατ. λέξεις. (β) Εκπαιδευτικός Λόγος: 96 σχολικά εγχειρίδια δημοτικού και γυμνασίου. Συνολική Έκταση: 3,4 εκατ. λέξεις (μελετώνται 2,8 εκατ.). (γ) Λογοτεχνικός Λόγος: 28 μυθιστορήματα (βραβεία αναγνωσιμότητας περιόδου 2010-2015). Συνολική Έκταση: 2,5 εκατ. λέξεις. Κατανομή Σχολικών Εγχειριδίων ανά Γνωστικό Αντικείμενο Πλήθος Εγχειριδίων Ελληνική Λογοτεχνία 13 Ελληνική Γλώσσα 12 Ιστορία 9 Φυσική – Χημεία – Βιολογία 9 Μαθηματικά 9 Γεωγραφία – Γεωλογία – Περιβάλλον 8 Θρησκευτικά 7 Αγωγή Αισθητική (Εικαστικά – Μουσική – Θέατρο) 15 Αγωγή Υγείας (Φυσική Αγωγή – Οικιακή Οικονομία) 5 Πληροφορική – Τεχνολογία 5 Αγωγή Κοινωνική – Πολιτική 3 Αγωγή Σταδιοδρομίας (ΣΕΠ) 1 Κατάλογος μυθιστορημάτων: Συγγραφέας, Τίτλος Έτος 1ης έκδοσης Δούκα, Μάρω - Το δίκιο είναι ζόρικο πολύ 2010 Θέμελης, Νίκος - Η συμφωνία των ονείρων 2010 Καρυστιάνη, Ιωάννα - Τα σακιά 2010 Μιχαλοπούλου, Αμάντα - Πώς να κρυφτείς 2010 Ελευθερίου, Μάνος - Πριν απ' το ηλιοβασίλεμα 2011 Ζουργός, Ισίδωρος - Ανεμώλια 2011 Μακριδάκης, Γιάννης - Η άλωση της Κωνσταντίας 2011 Μπουραζοπούλου, Ιωάννα - Η ενοχή της αθωότητας 2011 Πανσέληνος, Αλέξης - Σκοτεινές επιγραφές 2011 Παπαδημητρίου, Χίλντα - Για μια χούφτα βινύλια 2011 Παπαθεοδώρου, Θοδωρής - Οι καιροί της μνήμης 2011 Τριανταφύλλου, Σώτη - Για την αγάπη της γεωμετρίας 2011 Φακίνος, Μιχάλης - Η έρημος έρχεται 2011 Βαμβουνάκη, Μάρω - Κυριακή απόγευμα στη Βιέννη 2012 Διβάνη, Λένα - Εγώ, ο Ζάχος Ζάχαρης 2012 Στεφανάκης, Δημήτρης - Φιλμ νουάρ 2012 Ακρίβος, Κώστας - Αλλάζει πουκάμισο το φίδι 2013 Ζέη, Άλκη - Με μολύβι φάμπερ νούμερο δύο 2013 Κορτώ, Αύγουστος - Το βιβλίο της Κατερίνας 2013 Κωνσταντούρου, Μαρία - Αγεφύρωτες σιωπές 2013 Μαντά, Λένα - Με λένε Ντάτα 2013 Ξανθούλης, Γιάννης - Κωνσταντινούπολη των ασεβών μου φόβων 2013 Ρώσση–Ζαΐρη, Ρένα - Άρωμα βανίλιας 2013 Ανδρουλάκης, Μίμης - Αλλέγκρα 2014 Δημουλίδου, Χρυσηίδα - Το κελάρι της ντροπής 2014 Παπαδοπούλου, Ελισάβετ - Μέρες και νύχτες που δεν ήταν δικές μας 2014 Χατζή, Αθηνά - Η θάλασσα έφυγε 2014 Χωμενίδης, Χρήστος - Νίκη 2014 Τεχνικές προδιαγραφές & Μορφότυπος Για την αναπαράσταση των δεδομένων και των μεταδεδομένων υιοθετήθηκε η πολυεπίπεδη οπτική των XML σχημάτων και τροποποιήθηκε το διεθνές πρότυπο TEI P5, 4.0.0 (Text Encoding Initiative). Δημιουργήθηκε ειδικός χώρος ονομάτων sanCorpus (sanC) με σχήμα τύπου RELAX-NG. Το σώμα κειμένων διατίθεται σε TXT και σε XML σε τρεις εκδοχές: Βάθος 0: απλό κείμενο (TXT). Περιλαμβάνει το main core (κείμενο βάσει του οποίου εξετάζονται οι συγκρίσεις ομοιότητας) και το out of core (κείμενο εκτός εμβέλειας της διατριβής, στο οποίο περιλαμβάνονται κείμενα που πλαισιώνουν το κυρίως κείμενο, π.χ. κείμενα διδασκαλίας, πίνακες περιεχομένων, εξώφυλλα) Βάθος 1 = κείμενα στην απλούστερη δυνατή XML κωδικοποίηση Βάθος 2 = κείμενα με πιο λεπτομερείς XML κωδικοποιήσεις. Αυτή η έκδοση (san-Corpus v1.0, Depth 0: Plain Text) περιλαμβάνει το σώμα κειμένων σε μορφή απλού κειμένου (Βάθος 0) στην αρχική του διάταξη (βλ. Επεξεργασία). Στόχος είναι ο σταδιακός εμπλουτισμός με επισημειωμένα δεδομένα, καθώς και με τις εκδοχές Βάθους 1 και 2. Επεξεργασία (Processing) Η μεθοδολογία συλλογής των δεδομένων, η θεωρητική τεκμηρίωση και το σχήμα επισημείωσης επεξηγούνται στις μελέτες Αφεντουλίδου (2022, 2021, 2013, 2012) και Afentoulidou (2009). Η πρώτη εκδοχή (Βάθος 0) χρησιμοποιήθηκε αποκλειστικά για τη λημματοποίηση που απαιτούσε η Collostruction Analysis (Gries, 2024). Στάδια επεξεργασίας (για τη λημματοποίηση): Τμηματοποίηση σε προτάσεις (sentence segmentation) με τη χρήση της βιβλιοθήκης Stanza (Stanford NLP Group, Qi et al. 2020), η οποία βασίζεται στο μοντέλο Greek Dependency Treebank (GDT) του Ινστιτούτου Επεξεργασίας του Λόγου / ΕΚ «Αθηνά». Τυχαία αναδιάταξη των προτάσεων για την προστασία της ακεραιότητας των πρωτότυπων έργων. Λημματοποίηση με τον ILSP Lemmatizer μέσω της Υποδομής Clarin-EL. Για την ανάλυση συμφράσεων του δείκτη σαν απομονώθηκαν συγκεκριμένοι λεκτικοί τύποι. Αδειοδότηση & Δικαιώματα Το san-Corpus συγκροτήθηκε για τις ανάγκες της διδακτορικής διατριβής και προστατεύεται από το δικαίωμα ειδικής φύσης σύμφωνα με την Οδηγία 96/9/ΕΟΚ και το άρθρο 45Α του Ν. 2121/1993. Η χρήση του περιεχομένου γίνεται αποκλειστικά για ερευνητικούς σκοπούς βάσει των εξαιρέσεων της Οδηγίας 2001/29 και της Οδηγίας (ΕΕ) 2019/790 (Text and Data Mining exceptions / Fair Use). Η πρόσβαση είναι περιορισμένη (Restricted Access) και παρέχεται αποκλειστικά σε μέλη της ακαδημαϊκής κοινότητας για σκοπούς επαλήθευσης των αποτελεσμάτων της διατριβής και περαιτέρω μη εμπορική έρευνα. Η πηγή προέλευσης δικαιούται να ζητήσει οποιαδήποτε τροποποιητική ενέργεια (π.χ. αφαίρεση) επί του πρωτότυπου περιεχομένου. Τέλος, η άδεια CC BY-NC-ND 4.0 ισχύει για την επιμέλεια (curation), τα μεταδεδομένα και τη γλωσσολογική επισημείωση του σώματος κειμένων. Βιβλιογραφικές αναφορές Αφεντουλίδου, Β. (2022). Σώμα ελληνικών κειμένων για τη μελέτη δομών ομοιότητας της Νέας Ελληνικής: σχεδιασμός και υλοποίηση. Στο Πρακτικά του 10ου Συνεδρίου Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας (σσ. 67-94). ΕΚΠΑ. Αφεντουλίδου, B. (2021). Δομές ομοιότητας στη Νέα Ελληνική. Σωματοκειμενικές παρατηρήσεις για τον πολυλειτουργικό δείκτη σαν. Προφορική ανακοίνωση στην 41η Ετήσια Συνάντηση του Τομέα Γλωσσολογίας, 13–15 Μαΐου 2021. ΑΠΘ. Αφεντουλίδου, Β. (2013). Και σου απάντησα κάτι σαν ‘τέλεια, εντάξει’. Δείκτης σαν + ευθύς λόγος;. Προφορική ανακοίνωση στο 7ο Συνέδριο Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας, 16–18 Μαΐου. ΕΚΠΑ. Αφεντουλίδου, Β. (2012). Συγκρίσεις ομοιότητας στα Νέα Ελληνικά: ο δείκτης σαν. Στο Z. Gavriilidou, A. Efthymiou, E. Thomadaki & P. Kambakis-Vougiouklis (Επιμ.), Selected papers of the 10th International Conference on Greek Linguistics (σσ. 696-707). DUTH. Afentoulidou, V. (2009). Sketching the σαν conditional construction in Modern Greek. Submitted essay, 2009 Linguistic Institute, Linguistic Structure and Language Ecologies, Linguistic Society of America and UC Berkeley. Gries, Stefan Th. 2024. Coll.analysis 4.1. A script for R to compute perform collostructional analyses. https://www.stgries.info/teaching/groningen/index.html Institute for Language and Speech Processing - Athena Research Center (2015). ILSP Lemmatizer. Version 1. [Software (Tool/Service)]. CLARIN:EL. http://hdl.handle.net/11500/ATHENA-0000-0000-23EE-D Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 101–108). Online: Association for Computational Linguistics.
Large language models encode rich semantic information, but how concreteness is represented across layers remains unclear. We examine layer-wise linear separability of concreteness by training linear probes on hidden representations from two opensource model families at multiple scales: Qwen3 and Gemma3-<br/>Instruct. Using human concreteness ratings, we build balanced prompt datasets with four difficulty levels: an extreme abstract–concrete contrast and three finer boundary comparisons at the abstract end, mid-range, and concrete end. Probes achieve high accuracy on the extreme contrast in shallow layers, showing that endpoint differences are strongly linearly separable. For finer distinctions, performance follows a stable hierarchy: midrange concreteness is easiest to separate, abstract-end distinctions are hardest, and concrete-end distinctions are intermediate. Across models, accuracy rises rapidly in early layers, peaks in<br/>middle layers, and declines in later layers. Together, these findings clarify how the linear accessibility of concreteness varies across LLM layers.
We present MorfFlex, a morphological dictionary architecture suitable for languages with extensive regularity in both inflection and derivation. As the primary example of MorfFlex in use we introduce MorfFlex CZ, a morphological dictionary of Czech. It is distributed as a simple, unstructured list of <wordform, lemma, tag> triplets, however, its manually maintained, unpublished source files and conversion scripts encode a sophisticated system of inflectional and derivational patterns. These patterns dramatically reduce the otherwise enormous size of the dictionary, which currently contains over 100 million wordforms and more than 1 million lemmas. The MorfFlex CZ dictionary serves as an essential resource for ensuring the consistency of manual morphological annotation in the Prague Dependency Treebanks and underpins state-of-the-art automatic tools such as MorphoDiTa. In this paper, we focus on: (i) presenting an effective method for managing the rich morphological system within the dictionary, and (ii) demonstrating the utility of such a language resource for maintaining annotation consistency in corpora and supporting the development of advanced NLP applications.
This article investigates national-cultural specificity in language and its impact on translation through a comparative analysis of English and Uzbek. National-cultural specificity is examined as a linguistic and cultural phenomenon that encodes a community’s worldview, social norms, and value systems in lexical choices, phraseology, and pragmatic conventionsThe article further explores the challenges these cultural differences pose for translation, including untranslatability, pragmatic mismatch, and semantic gaps, and discusses strategies such as borrowing, cultural substitution, explicitation, and adaptation to preserve meaning.
The expression of an association between a conditioned stimulus (CS) and an aversive unconditioned stimulus (US) can be weakened by presenting the CS by itself (extinction [Ext]), pairing it with an appetitive US (counterconditioning [CC]), or pairing it with a neutral stimulus (novelty-facilitated extinction [NFE]). The present research tested whether NFE is less susceptible to ABC renewal than Ext and CC. In two experiments, participants viewed streams of rapid trials. After each stream, participants rated how likely it was that the target CS would be followed by the target US (i.e., predictive learning) as well as the valence of the target CS (i.e., evaluative conditioning). A stream was composed of two phases: Phase 1 established an association between the target CS and target US while Phase 2 aimed at disrupting the expression of this association through Ext, CC, or NFE. Phase 1 occurred in Context A while Phase 2 occurred in Context B. Prediction and valence ratings occurred in either Context A, B, or C. Neither Experiment 1 nor Experiment 2 found differences across interference conditions with predictive testing, regardless of test context. In Experiment 2, better controlled for context effect, CC and NFE altered the CS valence (CC more than NFE) when testing occurred in B, but the difference disappeared when testing occurred in either A or C. The present data do not support the hypothesis that NFE is less susceptible to ABC renewal than either Ext or CC. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Buildings shape how people feel, yet the mechanisms through which specific facade properties drive affective states remain empirically underspecified. Here we introduce the Cambridge Facade Affect Dataset (CFAD), 86 orthogonally rectified facade images annotated with continuous arousal and valence ratings from 85 participants, and establish a validated pipeline linking machine-vision-derived surface metrics to human affective responses. Focusing on three quantifiable attributes, complexity, transparency (window-to-wall ratio), and materiality (proportion of natural versus artificial surface composition), we show that perceived complexity is the dominant affective predictor, with significant positive associations for both arousal (beta = 0.507, p < 0.001) and valence (beta = 0.376, p < 0.001) and a curvilinear amplification at higher complexity levels. Transparency exhibits an inverted-U relationship with valence, while increasing surface artificiality suppresses arousal and reduces pleasantness consistent with biophilic response theory. Critically, machine-derived metrics show limited direct predictive power over affective outcomes; mediation analyses reveal that human perceptual evaluation functions as a necessary intermediate layer, with perceived materiality significantly mediating the machine-valence relationship (indirect effect = -0.205, p = 0.003). Cross-context validation demonstrates moderate stability of complexity and materiality ratings across image-based and in-situ conditions, while affective responses, particularly valence, exhibit significant context-dependence (ICC = 0.332). These findings advance facade research from descriptive morphological analysis toward predictive, perception-grounded modelling, and provide an empirically validated basis for affect-informed design of the urban environment.
This study proposes a dual-stream late fusion approach to dimensional emotion recognition by fusing facial landmarks and electrocardiogram data. Unlike category-based methods, the proposed method identifies emotions in the valence-arousal space. By individually optimizing the classifiers for both modalities and then fusing their outputs, the model encourages flexibility, interpretability, and robustness. The ASCERTAIN dataset is used, which consists of facial landmark trajectories, physiological signals like electroencephalograms, electrocardiograms, galvanic skin responses, and electromyograms, and self-reported arousal and valence ratings of fifty-eight subjects who viewed thirty-six videos. Arousal and valence were each modeled using an individual classifier. The random forest-based feature selection achieved eighty-six point ninety-one percent accuracy in valence prediction, while a soft voting ensemble of support vector machines, K-nearest neighbors, and random forest achieved sixty-one point eighteen percent for arousal. These results were merged in order to classify four emotional quadrants: High Arousal High Valence, High Arousal Low Valence, Low Arousal High Valence, and Low Arousal Low Valence, with an aggregate accuracy of sixty-nine point sixteen percent. The findings indicate that the model introduced accurately makes use of complementary modalities and a late fusion strategy, and is suitable for real-world emotion-aware systems.
This data article describes an Uzbek Universal Dependencies (UD) treebank released as a manually curated gold-standard dataset. The resource contains 681 sentences (7542 tokens) drawn from literary and educational Uzbek texts, providing a domain-specific complement to previously available web-based or news-oriented materials. Annotation was carried out in the INCEpTION environment by a five-member team comprising three linguists and two NLP engineers. The workflow followed the UD v2 framework and included calibration-stage agreement assessment, full-corpus double annotation, and adjudication to improve annotation consistency. Agreement measured on the shared calibration material was high across lemmatization, universal part-of-speech annotation, and complete morphological feature-value bundles. The released dataset contains final adjudicated gold-standard annotations, including lemmas, UPOS tags, morphological features, and basic dependency relations in standard CoNLL-U format, and has been validated for compatibility with the Universal Dependencies ecosystem. As an openly reusable Uzbek syntactic resource, it can support the development and evaluation of POS taggers, morphological analyzers, and dependency parsers, while also enabling comparative and cross-lingual studies for low-resource languages.
Crafting effective academic titles is a challenging task that requires balancing informativeness, conciseness, and reader engagement. This paper investigates titles produced by master’s students of Linguistics at the Faculty of Letters and Humanities of Sfax (FLSHS) and compares them with research article titles written by expert authors and with AI-generated alternatives. The study aims to evaluate students’ titles in relation to expert norms and to explore the potential of Large Language Models (LLMs) in academic title generation. A corpus of 659 titles, including master’s dissertations, research articles, and AI-generated titles, was analysed quantitatively and qualitatively using a synthesised model based on Ken Hyland and Zou (2022) and Swales and Feak (2012). The findings show that students adhere more closely to academic title conventions by producing informative and lexically dense titles, whereas expert authors prioritise reader engagement. AI-generated titles, although capable of producing useful lexical content, rely heavily on formulaic expressions and often fail to recognise the genre-specific conventions of academic discourse. They tend to be longer, more clausal, and more question-based than human-crafted titles. The study suggests that AI tools can serve as valuable brainstorming resources in academic writing pedagogy when their output is critically evaluated and adapted by users.
Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training corpora. We present a systematic methodology for transforming structured linguistic resources into specialized AI systems, demonstrating that expert-curated lexical databases can serve as effective foundations for conversational AI development. Our approach converts Hindi WordNet into 1.25 million diverse instruction-response pairs, fine-tunes a 12B-parameter language model using resource-efficient LoRA with 4-bit quantization. Evaluation through a Hindi language learning chatbot demonstrates that structured-knowledge-based systems achieve superior pedagogical effectiveness (91.0 vs. 79.4-83.6 for general-purpose models) while maintaining competitive semantic performance and exceptional consistency. The complete pipeline demonstrates a proof-of-concept methodology using Hindi for developing specialized AI systems for any languages with WordNet resources. This work addresses the critical gap in AI accessibility for low-resource languages, offering a practical alternative to corpus-intensive approaches and potentially enabling specialized AI development for the hundreds of languages with existing WordNet resources.
ObjectiveAchilles tendon protective shoes are medical orthoses that must be worn after Achilles tendon rupture surgery. Currently, these shoes predominantly feature black and gray color schemes and lack conspicuous warning markings, making them visually obscure and prone to accidental impacts, thereby increasing the risk of re-rupture. By conducting perceptual research on the application of warning colors and textures in the shoe upper, this study aims to explore design combinations that can effectively enhance visual warning, thereby improving the safety of the product during use and reducing the risk of secondary injury.MethodsBased on the theory of warning colors and combined with the visual warning arousal mechanism and eye-tracking physiological indicators, three typical biological warning textures (dots, cracks, stripes) and three warning color combinations (orange-black, yellow-black, purple-yellow) were selected to create nine experimental combinations. These were applied to a uniform model of Achilles tendon protective shoes as visual stimulus materials for the experiment. After collecting subjective valence ratings and eye-tracking data (including Time to First Fixation and pupil diameter), the data were analyzed using repeated-measures ANOVA, followed by post-hoc multiple comparisons using the LSD method. Pearson correlation analysis was conducted to examine the consistency between subjective and objective indicators, thereby comprehensively evaluating the perceived warning intensity of each combination.ResultsThe nine combinations of warning textures and colors showed no significant difference in Time to First Fixation, but exhibited significant differences in pupil diameter response and subjective warning valence ratings, with a significant interaction effect between texture and color. A moderate positive correlation was observed between subjective warning valence ratings and the increment of pupil diameter. Among these, subjects exhibited stronger alertness and a greater tendency to make avoidance decisions when viewing striped yellow-black and striped orange-black patterns. In contrast, alertness levels were significantly lower when observing striped purple-yellow or dotted pattern combinations. This indicates that while stripes are superior to dots in eliciting visual alertness, their actual effectiveness is significantly influenced by color interactions. The results of multiple comparisons revealed that the differences identified by subjective warning assessments were not fully captured by the objective physiological indicator of pupil diameter. The striped yellow-black pattern received significantly higher subjective ratings, influenced by cultural learning, symbolic semantics, and individual experience. In contrast, while the cracked-orange-black pattern elicited significant pupil dilation, it was not explicitly perceived as highly alerting. The striped orange-black pattern elicited a pupil diameter response ratio reaching the sympathetic activation threshold, accompanied by high alertness ratings, forming a complete alertness perception cycle that integrates both bottom-up and top-down processing. In contrast, the dotted yellow-black pattern produced a pupil diameter response ratio reaching the parasympathetic activation threshold, which can suppress sympathetic activity during the early stage of alertness stress. The findings suggest that the design of warning colors for Achilles tendon protective footwear should prioritize the selection of colors based on texture, rather than independently selecting texture or color.ConclusionThe stripe pattern in yellow-black serves as the optimal biological warning color paradigm for Achilles tendon protective shoes, significantly enhancing safety. The orange-black stripe pattern can be considered a secondary option, while the dot pattern in yellow-black should be avoided in warning signs.
Textbooks are fundamental educational tools that not only deliver curricular content but also convey societal and linguistic values. In the context of minority language education, textbooks have particular significance, as students’ language attitudes, identity, and self-perception are closely linked to the status and presentation of their native language. Language ideologies—often implicit beliefs about language and its use—shape how communities perceive linguistic norms, varieties, and speakers.
ABSTRACT The widespread use of TikTok among elementary school students has brought noticeable changes to the way children communicate in their daily lives. The platform is no longer used merely as a source of digital entertainment, but has also begun to shape students’ word choices, speaking styles, and language habits. This condition can be observed among students at MIS Al-Khairaat Pombewe, who have become increasingly familiar with viral expressions, popular abbreviations, slang, and the mixing of Indonesian with foreign languages in everyday conversations. Such circumstances have raised concerns regarding the declining use of proper and standard Indonesian within the school environment. This study employed a descriptive qualitative approach involving the principal, teachers, and students selected purposively as research informants. Data were collected through observations, interviews, and documentation, then analyzed through the stages of data reduction, data presentation, and conclusion drawing. The findings reveal that TikTok exerts a dual influence on children’s language development. On the one hand, the platform contributes to vocabulary expansion, enhances students’ creativity in language use, and broadens their digital knowledge. On the other hand, the intensity of TikTok usage encourages the frequent use of informal language in formal situations, leading to a gradual decline in the use of proper Indonesian according to linguistic norms. Therefore, the involvement of teachers and parents is necessary to guide children toward wiser social media use without neglecting the development of their language abilities. ABSTRAK Fenomena penggunaan TikTok di lingkungan sekolah dasar memperlihatkan perubahan yang cukup nyata pada cara siswa berkomunikasi sehari-hari. Platform ini tidak lagi sekadar dimanfaatkan sebagai hiburan digital, tetapi turut membentuk pilihan kata, gaya berbicara, hingga kebiasaan berbahasa anak. Kondisi tersebut terlihat pada siswa MIS Al-Khairaat Pombewe yang semakin akrab dengan istilah viral, singkatan populer, bahasa gaul, serta pencampuran bahasa Indonesia dengan bahasa asing dalam percakapan mereka. Situasi ini memunculkan perhatian terhadap menurunnya penggunaan bahasa Indonesia yang baik dan benar di lingkungan sekolah. Kajian ini memanfaatkan pendekatan deskriptif kualitatif dengan melibatkan kepala sekolah, guru, dan siswa sebagai informan yang dipilih secara purposive. Informasi penelitian diperoleh melalui observasi, wawancara, dan dokumentasi, kemudian dipahami melalui tahapan reduksi data, penyajian data, dan penarikan kesimpulan. Temuan penelitian memperlihatkan bahwa TikTok memberi pengaruh ganda terhadap perkembangan bahasa anak. Di satu sisi, media sosial tersebut membantu siswa memperluas kosakata, meningkatkan kreativitas dalam berbahasa, dan memperkaya wawasan digital mereka. Di sisi lain, intensitas penggunaan TikTok ikut mendorong penggunaan bahasa informal dalam situasi formal sehingga kebiasaan menggunakan bahasa Indonesia sesuai kaidah menjadi semakin berkurang. Karena itu, keterlibatan guru dan orang tua dibutuhkan agar penggunaan media sosial dapat diarahkan secara lebih bijak tanpa mengabaikan perkembangan kemampuan berbahasa siswa.
The practice of web form submission has emerged as a prime conduit for attackers, enabling them to infiltrate modern web applications and illegally harvest sensitive user data. Traditional defense mechanisms, such as static security reviews and server-side validation, are proving insufficiently agile for real-time detection of client-side vulnerabilities. This inadequacy arises directly from the rapid evolution of modern interfaces, which involves spontaneous DOM changes, dynamic element generation, and semantic interpretation that varies based on context and culture. This article presents an innovative browser extension framework that leverages a heuristic-based, multi-dimensional analytical engine combined with deep DOM inspection to identify insecure form submissions the moment they occur. The proposed methodology introduces five fundamental innovations: a contextual risk scoring system that models the complex interdependencies among form fields; an adaptive weighting scheme for risk patterns, accommodating diverse cultural and linguistic norms; a predictive vulnerability estimator that anticipates future threats; intelligent DOM mutation filtering designed to significantly optimize runtime performance; and cross linguistic semantic recognition to determine the true purpose of fields globally. Based on theoretical projections, this combined approach promises to enhance vulnerability detection accuracy while simultaneously reducing computational demands by approximately. Critically, all security analysis is executed exclusively on the user's local machine, guaranteeing privacy by ensuring no sensitive data is transmitted externally. A proof of concept application confirms the framework's practical feasibility and high efficacy for client side security assessment and catalyzing the development of flexible, scalable, and privacy respecting browser-based protections.
<div> This paper examines register variation in Latin from the third century BCE to the fourteenth century CE using Key Feature Analysis (KFA) (Egbert and Biber, 2023), a quantitative method for identifying statistically over-and underrepresented linguistic features. Registers are defined as text varieties linked to communicative situations and characterized by distributions of lexico-grammatical features (Biber, 1988, 1995). Six dependency-parsed Universal Dependencies (UD) treebanks are classified a priori into ten register categories based on established scholarship. Additionally, Principal Component Analysis (PCA) is used to reduce dimensionality, in order to explore the texts major patterns of variation and clusters of linguistically similar texts. KFA reveals systematic register-specific grammatical profiles consistent with previous research (e.g. Biber (2014b)). Registers with involved language use (e.g. letters and speeches) show higher frequencies of personal reference and engagement, while philosophical texts favor subordination and impersonal constructions. Registers containing narrative elements (e.g. historiography, satire) contain high frequency of verbs in past tense. PCA places the charter register into a distinct cluster, while other registers form more closely grouped patterns. The strongest components reflect contrasts in number, aspect, tense and person, alongside subordination and cordination distributions. The results are largely confirmatory: KFA produces coherent and interpretable groupings of grammatical features consistent with previous findings, providing a proof of concept for quantitative register analysis in historical corpora. Data and code are openly available for future research. </div>
Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has focused primarily on recognition tasks such as meal identification, ingredient detection, and nutrition estimation. Image-based prediction of multisensory experience remains largely unexplored. We introduce FoodSense, a human-annotated dataset for cross-sensory inference containing 66,842 participant-image pairs across 2,987 unique food images. Each pair includes numeric ratings (1-5) and free-text descriptors for four sensory dimensions: taste, smell, texture, and sound. To enable models to both predict and explain sensory expectations, we expand short human annotations into image-grounded reasoning traces. A large language model generates visual justifications conditioned on the image, ratings, and descriptors. Using these annotations, we train FoodSense-VL, a vision language benchmark model to produce both multisensory ratings and grounded explanations directly from food images. This work connects cognitive science findings on cross-sensory perception with modern instruction tuning for multimodal models and shows that many popular evaluation metrics are insufficient for visually sensory inference.
Abstract Arousal and valence are fundamental dimensions of affective experience signifying levels of activation and pleasantness, respectively. These dimensions play a crucial role in shaping emotional responses and behaviors, with significant implications for psychopathology. Previous machine learning studies had some success decoding these states from brain activation patterns observed during task-based functional magnetic resonance imaging (fMRI), but the results have varied across studies. Moreover, prior studies have often been limited by small sample sizes, weak decoding performance, and non-whole-brain analyses, leaving the neural representations of arousal and valence largely unresolved. Here we successfully decoded arousal and valence from whole-brain task-fMRI data collected from 132 participants during exposure to 300 unique emotional stimuli, including 150 movie clips and 150 text scenarios that reliably induced a wide range of arousal and valence states. Mass univariate general linear models identified block-level activation (emotion stimuli > washout) from all gray matter voxels. Multivariate regression analysis predicted arousal and valence ratings based on these gray matter activations. Patterns in the fMRI data underlying arousal and valence were robust, as they were successfully decoded across both induction modalities using five different linear multivariate regression models. Although significant, decoding from scenarios was less successful than from movies, likely due to their more imaginative nature. In particular, decoding arousal from scenarios only showed low predictive utility. Representations of arousal and valence were widespread throughout the brain, and we reveal cerebellar and brainstem contributions that have largely been absent in past fMRI decoding studies. These findings clarify the distributed neural basis of arousal and valence and provide a foundation for future clinical research on the role of these constructs in affective dysregulation.