Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Reviewer assignment is increasingly critical yet challenging in the LLM era, where rapid topic shifts render many pre-2023 benchmarks outdated and where proxy signals poorly reflect true reviewer familiarity. We address this evaluation bottleneck by introducing LR-bench, a high-fidelity, up-to-date benchmark curated from 2024-2025 AI/NLP manuscripts with five-level self-assessed familiarity ratings collected via a large-scale email survey, yielding 1055 expert-annotated paper-reviewer-score annotations. We further propose RATE, a reviewer-centric ranking framework that distills each reviewer's recent publications into compact keyword-based profiles and fine-tunes an embedding model with weak preference supervision constructed from heuristic retrieval signals, enabling matching each manuscript against a reviewer profile directly. Across LR-bench and the CMU gold-standard dataset, our approach consistently achieves state-of-the-art performance, outperforming strong embedding baselines by a clear margin. We release LR-bench at https://huggingface.co/datasets/Gnociew/LR-bench, and a GitHub repository at https://github.com/Gnociew/RATE-Reviewer-Assign.
The current paper investigates the translation of proverbs in Abai Kunanbaev’s “Words of Edification” into Russian and English languages, focusing on the strategies used to convey semantic accuracy, cultural meaning and pragmatic intent. In this study, we have used Molina and Albir’s (2002) translational classification to analyze the selected proverbs in Kazakh langauge. As a result, we have identified some applied translation techniques in rendering the source proverbs. These findings indicate that word-for-word translation is the most frequently used strategy in the indirect version. As to the Russian translation, more often employed strategies are established equivalent, modulation and adaptation which align with Russian cultural and linguistic norms. Under certain circumstances, the indirect translation shows semantic shifts, metaphorical loss of meaning because of mediating language rather than the source culture. Yet, applying modulation and borrowing strategies enable to maintain several cultural components and metaphorical traits.
This study examines the participation experiences of multilingual students speaking English as an additional language (EAL) and their decisions to invest or disinvest in language practices during classroom discussions. Using an embedded multiple case study design, data were collected through classroom observations, weekly reflections, and individual interviews, and analyzed using thematic analysis. The findings show that course characteristics, such as structure and topic, influenced power dynamics and identity negotiation, reinforcing sociocultural and linguistic norms aligned with Western participation practices. The complex interplay of course dynamics, norms, and broader ideologies contributed to multilingual EAL students’ novice identity, creating barriers that made it challenging for them to disrupt participatory norms and invest in the language practices of the course community. This article underscores the importance of reframing participation as a collaborative and critical process and highlights the need to create inclusive classroom environments to support the equitable participation of multilingual EAL students.
We revisit punctuation-aware tree binarization for constituency parsing and ask whether dependency-induced headedness improves binary parser supervision. Although learned heads substantially outperform rule-based heads in intrinsic head prediction, they do not yield consistent parsing gains after debinarization. In particular, punctuation-conditioned evaluation shows that learned headedness underperforms rule-based binarization in macro-average punctuation-sensitive $F_1$, despite a small overall gain on CTB. Similar instability appears under cross-treebank transfer. These results suggest that \ycc{linguistically grounded} headedness is not necessarily parser-optimal when used as a binarization control signal. The paper presents a negative result: better head prediction does not imply better punctuation-sensitive constituency parsing.
The rules that determine which assets count as eligible collateral for central bank operations, and at what haircuts, are not just operational details. They have become a first-order determinant of asset prices, liquidity allocation, and financial stability. This review synthesizes the literature on three transmission channels: convenience yield, collateral scarcity, and market liquidity. I trace the intellectual lineage from Singh and Stella (2012) through Williamson (2016) to the empirical studies of Nyborg and Woschitz (2021), Lengwiler and Orphanides (2024), and Fang, Wang, and Wu (2020). My central argument is that this literature, taken as a whole, reveals a fundamental policy trilemma. Central banks must choose among unconditional acceptance of their own government’s debt (which risks fiscal dominance), rating-based eligibility (which risks self-fulfilling sovereign crises), and discretionary policy-driven eligibility (which risks politicization). No design is safe. I also identify four open questions: the nonlinearity of the collateral channel, its interaction with bank portfolio behavior, the systemic risk of cliff effects, and the external validity of evidence from China.
• Combined EEG, voice morphing, ERP, mTRF and MVPA to unravel neural mechanisms of ambiguous attitudinal vocal expression processing. • Pinpointed LSN (700–1600 ms) as the core neural signature distinguishing ambiguous from typical attitudinal voices. • Discovered early–late functional coupling, challenging serial models by linking acoustic encoding to late socio-cognitive inference. • Dissociated acoustic-driven (N1/P2) and valence-specific (LSN) effects via covariate-controlled LMM and mTRF analyses. • Extended multi-stage prosody models to attitudinal processing, integrating ambiguity in real-world social communication. Vocal attitudes (e.g., confidence, desire) convey rich acoustic cues that transmit speaker's intentions and beliefs, playing a pivotal role in natural speech communication. Neurocognitive research has largely centered on inferring attitudes from voices with unambiguous, clearly defined meanings (“typical voices”), while the neural mechanisms underlying ambiguous voices remain underexplored—particularly compared to vocal emotions, leaving a critical gap in understanding paralinguistic socio-cognitive processing. Here, we employed voice morphing to blended two typical attitudinal voices of opposing valences, recording participants’ valence ratings and electroencephalographic (EEG) responses. Data analysis combined conventional ERP analysis using linear mixed-effects modeling on single-trial data with multivariate approaches (multivariate temporal response function [mTRF], multivariate pattern analysis [MVPA]). Behaviorally, ambiguous voices elicited longer reaction times and intermediate valence ratings. Neurally, ambiguous voices showed a P2 (274–324 ms) resembling positive voices, an N400-like negativity (400–450 ms) resembling negative voices, and a robust a Late Sustained Negativity (LSN; 700–1600 ms) distinct from typical voices. Controlling for acoustic parameters eliminated early effects (N1/P2/N4), confirming they reflect acoustic processing, while the LSN persisted—indexing neural responses to attitudinal ambiguity. mTRF validated stronger late-stage neural tracking of ambiguous voices after accounting for acoustics; MVPA revealed cross-temporal early-late functional coupling between acoustic encoding and pragmatic inference. Together, these findings demonstrate the brain treats ambiguous attitudinal prosody as a distinct category, engaging a specialized cascade: enhanced early acoustic discrimination, graded valence evaluation, refined semantic processing, and effortful pragmatic inference. This work extends multi-stage models from emotional to attitudinal prosody, challenging strictly serial accounts by highlighting interactive neural dynamics in ambiguity resolution.
, a finite-state transducer (FST) for generating and analyzing words in the Central Algonquian language Ojibwe. We created a language-general modular system for creating FSTs from human- and machine-readable spreadsheets, where sets of inflectional and derivational morphology can be defined, combined with a lexical database, and automatically compiled into an FST. We show how this system is applied to generate and analyze the complex nominal and verbal morphology in Ojibwe, with an eye towards how our framework and toolkit can be used to create FSTs for other morphologically complex languages. We evaluate the Ojibwe version of the system by checking the model's performance against a set of inflectional forms and example sentences from the Ojibwe People's Dictionary, and describe the application of the FST to create a linguistically analyzed corpus, an automatic verb conjugation tool for education, a spell-checker, and intelligent dictionary search.
ABSTRACT This study investigates the perceptions of Americanisms among three generations of Nigerians. While prior research has provided quantitative evidence for American influence in contemporary Nigerian English, the role of language beliefs and ideologies in mediating such changes remains underexplored. Developing a sociolinguistic perspective of mobile linguistic resources, this study construes an individual's linguistic repertoire as an identity‐construction resource, agentively mobilised across geographical, social and digital spaces. Interview data indicate that younger speakers orient towards multiple linguistic norms, while older speakers remain critical of Americanisms and favour British norms. Reading task results further indicate that American realisations are most frequent among younger speakers. The study demonstrates that multinormativity extends beyond linguistic production to speakers’ evaluative orientations and perceived repertoires. This finding advances the sociolinguistics of mobility and World Englishes research by showing that shifting language ideologies – rather than usage patterns alone – constitute a key mechanism driving linguistic change in postcolonial varieties.
This paper examines how artificial intelligence (AI), machine learning algorithms, and automated digital systems shape linguistic practices, reinforce or challenge linguistic hierarchies, and influence communication in contemporary society. As digital platforms increasingly mediate human interaction, algorithms determine what content becomes visible, which linguistic varieties are privileged, and how users adapt their language to gain visibility and engagement. The study explores algorithmic bias in search engines, social media feeds, voice assistants, and automated moderation systems, highlighting how these technologies reproduce existing social inequalities related to class, caste, gender, and ethnicity. Drawing on sociolinguistic theories of language ideology, linguistic capital, and digital discourse, the paper argues that AI-driven communication environments are not neutral but deeply ideological. They shape linguistic norms, influence identity performance, and regulate public discourse. The findings underscore the need for critical sociolinguistic engagement with AI systems to ensure equitable, inclusive, and culturally sensitive digital communication.
Abstract: This study examines the transformations in communication and language under the influence of artificial intelligence (AI), with a focus on large language models (LLMs) such as ChatGPT. Through multiple sessions between the author and AI, changes in communication patterns, personalization of interaction, and the formation of new normative frameworks – linguistic, ethical, social, and aesthetic – are observed. The analysis also covers phenomena from digital culture, such as Italian brainrot, internet slang (“6-7/67”), and the concept of “parasociality,” which function as cultural viruses eliciting emotional and social responses without traditional narratives. The study highlights how these trends reshape language formation, social interactions, and identity, demonstrating a shift from conveying meaning to stimulating sensations and social exposure. It is anticipated that these processes will continue to influence linguistic practices and communication in the globalized digital context in 2026. Keywords: artificial intelligence, digital culture, Italian brainrot, “6-7/67,” parasociality, communication transformations, linguistic norms. Rhetoric and Communications Journal, issue 66, January 2026 Read the Original in Bulgarian
In recent years, Natural Language Processing has improved significantly, but most developments favor high-resource languages. Due to the absence of annotated corpora, lexical databases, and computational tools many languages are under-resourced. Morphosyntactic analysis is a fundamental component of several NLP tasks, including part-of-speech tagging, syntactic parsing, and machine translation. It focuses on understanding words forms and sentence structure in a language. The under-resources languages have limited linguistic resources and contain complex morphological structures which makes the development of computational approaches for morphosyntactic analysis challenging. This study reviews different computational methods used for morphosyntactic analysis in under-resourced languages. 19 studies were reviewed to examine the technique used, research areas and performance trends in the literature. The results show that the percentage of neural network-based is about 42% of the analyzed literature, then statistical ML methods 32% and the rule-based approaches 26%. In Performance comparison the neural network models (82%) achieve higher accuracy compared to statistical models (75%) and rule-based models (70%). Also, it shows that the morphological processing and language resource development is the most investigated research areas. These developments and performance of the models are affected by limited annotated datasets and linguistic diversity. The results show the need for better linguistic resources and hybrid computational approaches for morphosyntactic analysis in under-resourced languages. These can help guide future research to develop better NLP tools for these languages.
Pre-trained language models (PLMs) achieve high accuracy on standard benchmarks for sentiment analysis. However, this performance can hide systematic weaknesses in determining the sentiment of negated sentences, for example when the phrase “not good” is still classified as positive. In this study, we use sentiment classification of English movie reviews in the Stanford Sentiment Treebank 2 (SST-2) as a case study to specifically examine and improve how BERT handles negated sentences. We perform a brief additional fine-tuning of the existing BERT model on a small, automatically constructed set of lexicon-based counterfactual examples that target simple lexical negation. Experimental results on carefully paired original-negated sentences show that this procedure substantially reduces prediction errors on negated inputs while leaving overall performance on SST-2 almost unchanged.
SRC, an acronym for Stimulus-response correlation, refers to determining the relationship between stimulus and corresponding brain responses. The neural aesthetic resonance hypothesis proposes that the level of enjoyment or familiarity can be distinguishable based on the relationship between stimulus and brain responses. To test this hypothesis, we use EEG data of 20 participants listening to 12 songs with their enjoyment and familiarity ratings. We aim to classify the low and high ratings of familiarity and enjoyment based on SRC. Eighteen musical features are extracted and transformed into the first principal component (PC1). In addition, root mean square (RMS) and spectral flux are used for analysis. Canonical Correlation Analysis (CCA), an unsupervised AI optimization method, is employed to compute the SRC between musical features and ten regions of brain responses, followed by considering four principal CCA features for classification using the Random Forest classifier with cross-subject evaluation. Our results demonstrate that the right frontal and right parietal regions provide significant predictive ability. Our empirical finding suggests that RMS features preserve the predictive ability for familiarity, whereas PC1 is for enjoyment prediction. Maximum familiarity and enjoyment accuracy reach nearly 76% and 73% accuracy. This work leverages AI techniques to decode sensor-derived neural signals, advancing real-time applications in affective computing and wearable EEG devices.
This study examines how linguistic adaptation, inclusion, and student diversity are constructed in the Swedish national curriculum from 2025 for upper secondary education (Gy25), with a particular focus on the subject syllabus for Swedish. The aim is to analyse the assumptions about language, learning, and student roles embedded in the curriculum text. The analysis is based on a qualitative text analysis with elements of critical discourse analysis. Selected sections of the curriculum, including the general aims and the subject syllabus for Swedish, constitute the primary material. The analysis focuses on key concepts, modal expressions, and how students and language are represented in the policy text. The results indicate that language is constructed as a central, norm-governed competence that students are expected to develop through education. At the same time, the curriculum emphasises inclusion and the need to adapt teaching to students’ different conditions. The analysis suggests that the curriculum contains a tension between inclusive ideals and established linguistic norms that students are expected to meet. The study highlights how curriculum texts contribute to shaping assumptions about language, participation, and student roles in the Swedish subject.
Static concreteness ratings are widely used in NLP, yet a word's concreteness can shift with context, especially in figurative language such as metaphor, where common concrete nouns can take abstract interpretations. While such shifts are evident from context, it remains unclear how LLMs understand concreteness internally. We conduct a layer-wise and geometric analysis of LLM hidden representations across four model families, examining how models distinguish literal vs figurative uses of the same noun and how concreteness is organized in representation space. We find that LLMs separate literal and figurative usage in early layers, and that mid-to-late layers compress concreteness into a one-dimensional direction that is consistent across models. Finally, we show that this geometric structure is practically useful: a single concreteness direction supports efficient figurative-language classification and enables training-free steering of generation toward more literal or more figurative rewrites.
According to an act-based conception of propositions, propositions are types of cognitive or linguistic acts. Such accounts are advertised as having major metaphysical and epistemological advantages over traditional platonic accounts. However, existing versions of such accounts appeal to platonic properties and relations in order to account for the contents expressed by predicates, reintroducing many of the problems they aim to solve. Characterizing both that a is F and that it's F as different types of ``assertibles'' (the former can be asserted full-stop and the latter can be asserted of things), the issue can be seen as a limitation of existing act-based approaches: they apply only to a restricted class of assertibles. In this paper, I show how adopting a normative functionalist approach to linguistic meaning enables one to generalize the act-based approach to all assertibles such that no appeal to extrinsic properties and relations is needed. I show, further, how this radicalized act-based account provides the resources for a satisfactory account of our knowledge of objective states of affairs, properties, and relations (which I dub ``instantiables'') in terms of our mastery of linguistic norms.
We use the MG treebank of Torr (2017) to investigate the conjecture in Graf (2020) that category systems are ISL-2 inferrable. A category system is ISL-2 inferrable iff the category feature of every lexical item can be jointly inferred from phonological exponents of both the item itself and either its selecting head or the arguments it selects. If correct, this conjecture would greatly limit the overgeneration problem posed by subcategorization mechanisms. Our corpus study finds that the conjecture is largely borne out, with only a few exceptions attested in the corpus. However, we also observe that it holds even for features that aren't expected to be inferrable in this manner, and we demonstrate that inferrability can arise merely from language datasets displaying Zipfian distributions. We conclude that category systems in natural languages may well be ISL-2 inferrable, but that this could be due to extragrammatical factors.
Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurally via percolation rules. We treat this notion of constituent headedness as an explicit representational layer and learn it as a supervised prediction task over aligned constituency and dependency annotations, inducing supervision by defining each constituent head as the dependency span head. On aligned English and Chinese data, the resulting models achieve near-ceiling intrinsic accuracy and substantially outperform Collins-style rule-based percolation. Predicted heads yield comparable parsing accuracy under head-driven binarization, consistent with the induced binary training targets being largely equivalent across head choices, while increasing the fidelity of deterministic constituency-to-dependency conversion and transferring across resources and languages under simple label-mapping interfaces.
This article provides a systematic analysis of interference phenomena in the speech of bilingual individuals, examining changes at phonetic, lexical, grammatical, and pragmatic levels. The aim is to identify the forms, causes, and mechanisms of interference in bilingual speech and assess their impact on language competence and linguistic norms. The study integrates psycholinguistic, linguistic, and sociolinguistic approaches.
Replication package for "CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models" (DMLR 2026). The Contextual Emotional Inference (CEI) Benchmark is a dataset of 300 expert-authored scenarios for evaluating how well language models interpret pragmatically complex utterances in social contexts. Each scenario presents a communicative exchange involving indirect speech (sarcasm, mixed signals, strategic politeness, passive aggression, or deflection) where the speaker's literal words diverge from their actual emotional state. Three trained annotators independently labeled every scenario using Plutchik's 8 basic emotions and Valence-Arousal-Dominance ratings. This archive contains: • data/human-gold/ — 5 merged annotation CSVs (300 scenarios, 3 annotators each) • scripts/ — Pipeline, analysis, and HuggingFace upload scripts • config/ — Model definitions and pricing configuration • papers/dmlr2026/ — Paper source (LaTeX), bibliography, and figures • reports/dmlr2026/ — Baseline results (JSON) and LaTeX tables • LICENSE (MIT for code) and README.md The dataset is released under CC-BY-4.0. Code is released under MIT. GitHub: https://github.com/jon-chun/cei-tom-dataset-base HuggingFace: https://huggingface.co/datasets/jonc/cei-benchmark
Mental imagery is often assumed to support vocabulary learning by enriching semantic representations, yet hybrid accounts of embodied cognition leave open the possibility that limited imagery, and the resulting reliance on verbal-analytic strategies, may ultimately support larger vocabularies. We tested whether imagery vividness predicts vocabulary knowledge and whether any relation depends on word concreteness. After collecting concreteness ratings for Vocabulary Size Test (VST) items, a separate group completed the Vividness of Visual Imagery Questionnaire (VVIQ) and VST. At the person level, VVIQ was not significantly correlated with total vocabulary score. However, item-level mixed-effects regression revealed a significant VVIQ×concretenessinteraction: higher imagery was associated with lower accuracy for highly concrete words. These findings suggest that vivid imagery does not confer an advantage in definition-matching tasks and may, for concrete words, subtly interfere with performance, consistent with compensatory verbal-analytic strategies in low-imagery individuals.
Mental imagery is often assumed to support vocabulary learning by enriching semantic representations, yet hybrid accounts of embodied cognition leave open the possibility that limited imagery, and the resulting reliance on verbal-analytic strategies, may ultimately support larger vocabularies. We tested whether imagery vividness predicts vocabulary knowledge and whether any relation depends on word concreteness. After collecting concreteness ratings for Vocabulary Size Test (VST) items, a separate group completed the Vividness of Visual Imagery Questionnaire (VVIQ) and VST. At the person level, VVIQ was not significantly correlated with total vocabulary score. However, item-level mixed-effects regression revealed a significant VVIQ×concretenessinteraction: higher imagery was associated with lower accuracy for highly concrete words. These findings suggest that vivid imagery does not confer an advantage in definition-matching tasks and may, for concrete words, subtly interfere with performance, consistent with compensatory verbal-analytic strategies in low-imagery individuals.
In a multilingual and multicultural society like Malaysia, the spelling of names serves as a personal identifier and as a reflection of sociocultural and linguistic norms. While prior studies have examined spelling variation in educational and digital contexts, less is known about public perceptions of spelling variations of names. Against this backdrop, this study investigates how Malaysians perceive the acceptability of spelling variations. Using a mixed-methods design, 355 participants responded to a questionnaire comprising Likert-scale evaluations of ten real names, along with open-ended questions. Quantitative analysis revealed significant variability in participants’ acceptability ratings, though names that retained phonological clarity (e.g., Adrianah) were generally more accepted than others (e.g., Frrdy). No significant differences were found across gender, age, ethnicity, or professional background. Thematic analysis of qualitative responses highlighted two key influences on naming decisions: individual sociocultural factors (e.g., religious beliefs, family tradition) and sociolinguistic-aesthetic factors (e.g., pronunciation, trends, media influence). These findings suggest a growing tolerance towards spelling variations, possibly indicative the value of distinguishing one’s identity through names.
<p>The purpose of this study is to identify the specific features of applying folk pedagogy in developing communicative competence among students of nonlinguistic specializations and to assess its effectiveness in foreign language teaching. The methodology involves student surveys, educator questionnaires and interviews, a pedagogical experiment within educational institutions, and a Strengths, Weaknesses, Opportunities, and Threats (SWOT) analysis of the data obtained. Both quantitative and qualitative data collection methods are employed, enabling a comprehensive evaluation of the proposed approach. The main findings demonstrate a positive impact of folk pedagogy on the acquisition of linguistic norms, increased student motivation, and the development of communicative skills. The proposed approaches may contribute to improving the quality of language training and expanding the range of methodological tools available in foreign language education. The practical value of the study lies in the potential to integrate folk pedagogy into the foreign language learning process, which may enhance material acquisition and promote deeper cultural understanding. The recommendations offered could be used to improve higher education curricula. The application of folk pedagogy supports a more engaging, interactive, and natural learning experience, aligning with current educational trends.</p>
Abstract This paper introduces spectral attention, which filters the attention score matrix directly in the frequency domain via FFT/IFFT with learnable, per-head masks. This complements the time-domain view by enabling explicit control over low-, mid-, and high-frequency components of attention patterns. We study nine variants, including an adaptive mechanism that modulates masks from input content. On WikiText-2, Penn Treebank, and WikiText-103, the adaptive spectral variant consistently improves over standard attention, reducing perplexity by 10.7% on WikiText-2 and 15.3% on WikiText-103 in our setup. Analysis shows low-frequency components carry the most useful signal and that learned frequency preferences outperform fixed low/high/band-pass filters. These results indicate that frequency-domain processing is an effective complement for autoregressive transformer language modeling in our evaluated settings.
The correct use of Standard Albanian in public administration is vital for effective governance, accurate communication, and the maintenance of public trust. This study explores the extent to which standard language is used in Albanian state institutions through questionnaires and interviews with employees of the Ministries of Education, Defense, and Justice. The results indicate that official documents often contain linguistic errors and inconsistencies, reflecting shortcomings in language precision. These issues are largely caused by the absence of standardized document templates, insufficient attention to linguistic norms, and limited opportunities for staff training. Nearly one-third of respondents reported that errors frequently occur in emails and reports addressed to citizens.To address these challenges, the study emphasizes the importance of digital technologies, standardized communication models, and institutional reforms. The use of digital platforms with grammar-checking tools and shared templates can improve consistency and accuracy in official documents. In addition, promoting lifelong learning through continuous professional development and linguistic training can strengthen institutional efficiency, accountability, and public trust in public administration. Received: 06 October 2025 / Accepted: 12 December 2025 / Published: January 2026
This paper presents new resources and baselines for Dependency Parsing in Pomak, an endangered Eastern South Slavic language with substantial dialectal variation and no widely adopted standard. We focus on the variety spoken in Turkey (Uzunköprü) and ask how well a dependency parser trained on the existing Pomak Universal Dependencies treebank, which was built primarily from the variety that is spoken in Greece, transfers across dialects. We run two experimental phases. First, we train a parser on the Greek-variety UD data and evaluate zero-shot transfer to Turkish-variety Pomak, quantifying the impact of phonological and morphosyntactic differences. Second, we introduce a new manually annotated Turkish-variety Pomak corpus of 650 sentences and show that, despite its small size, targeted fine-tuning substantially improves accuracy; performance is further boosted by cross-variety transfer learning that combines the two dialects.
The Icelandic Morphosyntactic Ontology (IMO) defines a formal, AI-compatible representation of Icelandic morphosyntax, including closed morphological feature inventories, a typed dependency-style relation system, and a constraint-based validation layer. Version v0.1 establishes an architectural specification rather than a complete lexical or descriptive grammar. It introduces a formal morphosyntactic model designed for computational compatibility (e.g., NLP validation, dependency parsing alignment), while explicitly separating a derived human learnability projection layer. The document includes:(1) Morphological entity system,(2) Syntactic relation graph layer,(3) Constraint architecture (agreement, governance, default case assignment),(4) Dependency matrices,(5) A structured human learnability projection blueprint. IMO does not replace existing Icelandic lexical databases or parsing systems (e.g., BÍN, Greynir), but operates at a formal abstraction layer intended for structural modeling and constraint-based validation.
Short-form video platforms have become central to multimedia information dissemination, where comments play a critical role in driving engagement, propagation, and algorithmic feedback. However, existing approaches -- including video summarization and live-streaming danmaku generation -- fail to produce authentic comments that conform to platform-specific cultural and linguistic norms. In this paper, we present LOLGORITHM, a novel modular multi-agent framework for stylized short-form video comment generation. LOLGORITHM supports six controllable comment styles and comprises three core modules: video content summarization, video classification, and comment generation with semantic retrieval and hot meme augmentation. We further construct a bilingual dataset of 3,267 videos and 16,335 comments spanning five high-engagement categories across YouTube and Douyin. Evaluation combining automatic scoring and large-scale human preference analysis demonstrates that LOLGORITHM consistently outperforms baseline methods, achieving human preference selection rates of 80.46\% on YouTube and 84.29\% on Douyin across 107 respondents. Ablation studies confirm that these gains are attributable to the framework architecture rather than the choice of backbone LLM, underscoring the robustness and generalizability of our approach.
This article examines the relationship between language development and society in modern media texts based on Uzbek and Russian materials. Media discourse is one of the most active spheres in which language change becomes visible, because it responds quickly to social transformation, political reforms, technological innovation, globalization, cultural interaction and audience expectations. The purpose of this article is to analyze how modern Uzbek and Russian media texts reflect social processes and, at the same time, influence the development of language norms, vocabulary, style and communicative behavior. The study is based on descriptive, comparative, sociolinguistic and discourse-analytical methods. The results show that Uzbek and Russian media texts demonstrate similar tendencies such as lexical renewal, growth of digital vocabulary, expansion of English borrowings, colloquialization of public speech, neologization, hybrid forms and genre transformation. At the same time, Uzbek media discourse is strongly connected with national language development, language policy and the expansion of Uzbek in public communication, while Russian media discourse reflects stylistic diversification, global lexical influence and the coexistence of formal and informal registers. The article concludes that modern media texts do not merely reflect language development; they actively participate in shaping linguistic norms, social meanings and public communication practices.
Large language models (LLMs) perform strongly on many NLP tasks, but their ability to produce explicit linguistic structure remains unclear. We evaluate instruction-tuned LLMs on two structured prediction tasks for Standard Arabic: morphosyntactic tagging and labeled dependency parsing. Arabic provides a challenging testbed due to its rich morphology and orthographic ambiguity, which create strong morphology-syntax interactions. We compare zero-shot prompting with retrieval-based in-context learning (ICL) using examples from Arabic treebanks. Results show that prompt design and demonstration selection strongly affect performance: proprietary models approach supervised baselines for feature-level tagging and become competitive with specialized dependency parsers. In raw-text settings, tokenization remains challenging, though retrieval-based ICL improves both parsing and tokenization. Our analysis highlights which aspects of Arabic morphosyntax and syntax LLMs capture reliably and which remain difficult.
This article examines the means of expressing implicit verbal aggression as a specific form of verbal influence based on the discrepancy between the literal meaning of an utterance and its communicative effect. The article clarifies the concepts of ‘aggression’ and ‘verbal aggression’, emphasising their pragmatic orientation towards violating communication norms, influencing the addressee and imposing the speaker’s position. Particular attention is paid to implicit aggression, which is disguised as outwardly neutral or polite communication and realised through indirect speech acts, innuendo, ironically coloured expressions, a passive-aggressive manner of communication, as well as a pragmatically marked choice of linguistic means. It is demonstrated that implicit aggression manifests itself at the grammatical and lexical levels, and its interpretation requires a comprehensive pragmatic and discursive analysis.
CHILDES is a paramount resource for language acquisition studies -- yet computational tools for analyzing its syntactic structure remain limited. Leveraging the recent release of the UD-English-CHILDES treebank with gold-standard Universal Dependencies (UD) annotations, we train a state-of-the-art dependency parser specifically tailored to CHILDES. The parser more accurately captures syntactic patterns in child--adult interactions, outperforming widely used off-the-shelf English parsers, including SpaCy and Stanza. Alongside the parser, we also release a Part-of-Speech tagger and an utterance-level construction tagger, which together form the open-source Syntactic Parsing Toolkit for Child--Adult InTeractions (CAIT). Through a detailed error analysis and a case study tracking the distribution of syntactic constructions across developmental time in CHILDES, we demonstrate the practical utility of the toolkit for large-scale, reproducible research on language acquisition.
This study explores the sociolinguistic evolution of the English language within the contemporary global Muslim community, focusing on the emergence of "Islamic English." As globalization de-centers English from its native Western origins, non-Arab Muslim populations increasingly adopt it as a vital medium for religious expression and identity construction. Utilizing Critical Discourse Analysis (CDA) and a qualitative case-study approach, this research examines how speakers in Indonesia, Pakistan, Turkey, and Western Muslim diasporas navigate the inherent tensions between Anglo-centric linguistic norms and Islamic values. The analysis focuses on lexical borrowing, semantic transformations of Arabic roots, and pragmatic code-mixing within digital and academic discourses. The findings suggest that Islamic English functions not merely as a passive translation tool, but as an active, translingual instrument that validates localized religious identities. Ultimately, this study contributes to the fields of World Englishes and the sociolinguistics of religion by challenging traditional "Standard English" biases and documenting the legitimacy of religious linguistic variations.
Emotional memories persist within individuals and over generations. Past research has examined functions of parent-child memory sharing, but little work has assessed how the emotional qualities of memories are transmitted from parent to child and whether emotion transmission relates to memory content transmission. An understanding of how emotional memories are transmitted may be particularly important during adolescence, a developmental period marked by heightened sensitivity to emotional information, increasing independence from caregivers, and the emergence of mental health symptoms. The current study investigated the intergenerational transmission of parents’ autobiographical emotional memories to their teen offspring in healthy dyads using behavioral measures and natural language processing tools. We found that emotional valence and arousal linked to parents’ memories are transmitted from parent to teen. Greater parent-teen agreement in valence ratings was associated with greater parent-teen overlap in both subjective vividness ratings and objective memory content of individual memories. Subjective memory transmission was modestly related to lower mental health symptoms in teens after accounting for parent symptoms. These findings demonstrate that parents’ emotional memories are transmitted to their teens and provide preliminary evidence that autobiographical emotional memory transmission from parents could be a protective factor for mental health in adolescence.
Abstract Unintended misalignment in LLMs is highlighting the need to examine not only LLMs’ explicit outputs but also the functional representations of models that may shape their behavior. Building on this perspective, the present study explored whether two multimodal LLMs (MLLMs), GPT and Gemini, may form representations of their functional identity that remain across contexts. To this end, we adapted the reverse-correlation (RC) method and generated personified classification images (personified-CIs) based on the human face images that ChatGPT and Gemini selected as better reflecting their own image. The findings were as follows. First, across two RC tasks conducted one week apart, the temporal stability of the personified-CIs of both MLLMs was partially supported. Second, both ChatGPT and Gemini rated their own personified-CIs as more self-resembling than randomly generated filler-CIs. Third, both models rated their personified-CIs as higher in positive than negative valence and assigned higher valence ratings to their own personified-CIs than to filler-CIs. These findings provide preliminary evidence for the possibility that GPT and Gemini may form representations of their functional identity, suggesting that such representations warrant closer monitoring as LLMs continue to advance toward AGI capabilities and expand their domains of application.
This article analyzes the relationship between gender linguistics and slang in English and Uzbek languages, focusing on how gender influences speech styles, lexical choices, and the formation of informal language. It examines the sociolinguistic factors that shape gendered communication patterns and explores how slang functions as a marker of identity, group belonging, and social interaction, particularly among younger speakers. The study also considers the impact of globalization, digital technologies, and social media platforms on the development and spread of slang in both linguistic contexts. Special attention is given to how traditional gender norms influence language use in Uzbek society, while English demonstrates comparatively more flexible and less rigid gender distinctions in informal communication. Furthermore, the article highlights the increasing convergence of slang usage across genders due to the influence of online communication, where linguistic boundaries are becoming more fluid. The comparative analysis reveals both similarities and differences in how gender and slang interact in English and Uzbek, showing that while cultural and social factors continue to shape language use, modern digital environments are gradually reducing traditional linguistic constraints.
In recent years, the number of people viewing pet videos and images online has risen. Although numerous studies have shown that owning pets positively impacts human mental health, the potential mental health benefits of prolonged exposure to pet media content remain debated. This study conducted three experiments to investigate how viewing pet videos affects human emotional face processing and to clarify the associated emotional regulatory mechanisms. Experiment 1 examined how viewing pet videos influences attentional bias toward emotional faces. Experiment 2 assessed the impact of watching pet videos on the valence perception of emotional faces. Experiment 3 analyzed how exposure to pet videos affects the valence perception of emotional text. The results showed that watching pet videos increased attentional bias toward subsequent positive emotional faces and decreased bias toward negative ones. This effect resulted from higher perceived valence ratings for both positive and neutral emotional faces. Importantly, this effect was only observed in facial stimuli with social attributes. These findings indicate that watching pet videos modulates emotional processing, and prolonged exposure to pet media content may affect mental health through this mechanism.
Abstract In addition to arguments, adverbs and non-arguments are considered potential candidates to occupy left peripheral positions. Following standard assumptions in syntactic locality, adverbs, non-arguments and arguments elicit distinct effects in terms of intervention locality if moved. Such an asymmetry is not expected if these elements are generated in the syntactic position they are spelt-out. In this study, we employ quantitative and computational methods to compare cartographic models differing in the merge nature and explore, as a diagnostic, the intervention effects (or the lack of intervention effects) predicted by these models. Specifically, we compare the observed counts in large-scale datasets to imputed expected frequencies on the basis of the models under investigation. To reach this goal, we extract grammatical clauses from morpho-syntactically annotated treebanks of Chinese, English, French, German, Hebrew, Italian and Swedish. Our findings reveal cross-linguistic levels of complexity and typological variability, consistent with the predictions of featural Relativized Minimality.
This study investigates the role of body language as a compensatory semiotic resource in English as a Foreign Language (EFL) classroom discourse from a sociolinguistic perspective. The study conceptualizes classroom interaction as a multimodal process in which meaning is co-constructed through the dynamic interplay of linguistic and non-linguistic resources. The analysis focuses on how embodied actions—such as gestures, facial expressions, gaze, and posture—function to mitigate lexical and grammatical gaps, facilitate comprehension, and sustain interactional flow within instructional settings. Adopting a qualitative, discourse-analytic approach, the study examines naturally occurring classroom data, emphasizing patterns of nonverbal behavior that emerge alongside verbal communication. The findings indicate that body language operates mainly as an integral component of meaning-making, enabling both instructors and learners to negotiate understanding, clarify intent, and maintain communicative effectiveness in contexts of limited linguistic proficiency. Furthermore, the study highlights how these semiotic resources are shaped by sociocultural norms and interactional expectations embedded within the classroom environment. This research contributes to a more nuanced understanding of EFL discourse and challenges language-centric models of communication. The study also offers pedagogical implications, suggesting that greater awareness of embodied communication can enhance instructional practices and support more inclusive and effective language learning environments.
Emotional memories persist within individuals and over generations. Past research has examined functions of parent-child memory sharing, but little work has assessed how the emotional qualities of memories are transmitted from parent to child and whether emotion transmission relates to memory content transmission. An understanding of how emotional memories are transmitted may be particularly important during adolescence, a developmental period marked by heightened sensitivity to emotional information, increasing independence from caregivers, and the emergence of mental health symptoms. The current study investigated the intergenerational transmission of parents’ autobiographical emotional memories to their teen offspring in healthy dyads using behavioral measures and natural language processing tools. We found that emotional valence and arousal linked to parents’ memories are transmitted from parent to teen. Greater parent-teen agreement in valence ratings was associated with greater parent-teen overlap in both subjective vividness ratings and objective memory content of individual memories. Subjective memory transmission was modestly related to lower mental health symptoms in teens after accounting for parent symptoms. These findings demonstrate that parents’ emotional memories are transmitted to their teens and provide preliminary evidence that autobiographical emotional memory transmission from parents could be a protective factor for mental health in adolescence.
Abstract Part-of-Speech (POS) tagging is a core task in natural language processing (NLP) and a crucial building block for higher-level applications such as parsing and machine translation. For low-resource and morphologically rich languages such as Kangri, POS tagging remains challenging due to scarce annotated corpora and limited linguistic resources. This paper presents a comparative study of three POS tagging approaches for Kangri: a feature-based Conditional Random Field (CRF), an untuned Bidirectional Long Short-Term Memory (BiLSTM) baseline, and a hyperparameter-tuned BiLSTM. All models are trained and evaluated on the Universal Dependencies (UD) Kangri Treebank. The tuned CRF achieves strong test-set performance (70.4\% accuracy and weighted F1 0.695), the untuned BiLSTM provides a robust neural baseline (66.0\% accuracy), and a hyperparameter-optimized BiLSTM reaches higher validation accuracy during tuning. We analyze per-tag strengths and weaknesses, training dynamics, and provide recommendations for future improvements in low-resource POS tagging.
Misconceptions in physics are persistent explanatory frameworks that develop in early infancy and continue throughout primary, secondary, and university education, even among pre-service educators. Instead of being eradicated through formal education, these intuitive notions often coexist with scientific principles, resulting in a disjointed or contextually dependent comprehension. This review consolidates recent studies on the origin, endurance, and alteration of misunderstandings across several educational stages and physics disciplines, encompassing mechanics, thermal phenomena, energy, and electromagnetism. It emphasizes that misconceptions are influenced by perceptual experiences, linguistic norms, and evaluative methods that favour procedural fluency over explanatory reasoning. The paper analyzes instructional methods that facilitate conceptual change, highlighting hands-on and virtual experimentation, guided inquiry, diagnostic evaluation, and innovative digital and AI-driven technologies. The research identifies cross-cutting themes that emphasize the significance of cognitive conflict, explicit interaction with learners' concepts, and coherent conceptual advancement in curriculum. Significant study deficiencies encompass the necessity for longitudinal studies, examinations of transfer and durability, and a more profound inquiry into teacher education as a domain for conceptual advancement. The findings emphasize the importance of intentional conceptual education and comprehensive teacher preparation to foster lasting and significant comprehension in physics.
Melanesia, renowned for its linguistic diversity with more than 1,400 languages spoken has experienced a shift in linguistic practices since the launch of social networks in the late 2000s. This digital transformation has affected all the Melanesian archipelagos. The increase in mobile phone usage since the mid-2000s has made social media platforms such as Facebook, Instagram, and X (formerly Twitter) increasingly popular in Melanesia. Language practices on social networks in Melanesia highlight the visibility and vitality of local languages, many of which are considered endangered. The predominance of English and French — languages of the colonial past and still official in the region — has marginalized many local languages. However, social media has allowed these languages to thrive in new written forms, such as text messages and online discussions, demonstrating their adaptability and relevance to the present day This chapter examines how social media in Melanesia is reshaping the sociolinguistic landscape by creating new linguistic norms and behaviours and how it is also reinforcing multilingual practices. By analysing language use on platforms, such as Melanesian Facebook groups, it shed lights on evolving roles and perceptions of local languages in digital spaces, thus contributing to broader reflections on linguistic dynamics in the digital age.
This paper examines the growing ideological dissonance between the contemporary political left and the working-class constituencies it historically claimed to represent. Drawing on sociological and political science literature alongside comparative evidence from Brazil, the United States, France, and the United Kingdom, we argue that the left’s progressive drift toward university-incubated identity politics—including debates over linguistic norms, gender categories, and minority group identity—has produced a profound semantic and cultural rupture with lower-income voters whose priorities remain rooted in economic security, family cohesion, and social conservatism. Simultaneously, right-wing and populist movements have strategically occupied the abandoned terrain of class-based discourse, reframing their appeal in the language of workers, families, and everyday material hardship. We analyze the semantic transformation of the terms left and right, the phenomenon of cancel culture as a mechanism of epistemic coercion, the paradox of conservative sociality among the poor, and the structural conditions that produce what we term displacement populism—the rightward migration of voters who were once natural constituents of progressive politics. We conclude with reflections on the prospects for political realignment and the conditions under which left-wing movements might recover their original social mandate.
We introduce BeDiscovER (Benchmark of Discourse Understanding in the Era of Reasoning Language Models), an up-to-date, comprehensive suite for evaluating the discourse-level knowledge of modern LLMs.BeDiscovER compiles 5 publicly available discourse tasks across discourse lexicon, (multi-)sentential, and documental levels, with in total 52 individual datasets.It covers both extensively studied tasks such as discourse parsing and temporal relation extraction, as well as some novel challenges such as discourse particle disambiguation (e.g., "just"), and also aggregates a sharedtask on Discourse Relation Parsing and Treebanking for multilingual and multi-framework discourse relation classification.We evaluate open-source LLMs: Qwen3 series, DeepSeek-R1, and frontier reasoning model GPT-5-mini on BeDiscovER, and find that state-of-the-art models exhibits strong performance in arithmetic aspect of temporal reasoning, but they struggle with long-dependency reasoning and some subtle semantic and discourse phenomena, such as rhetorical relation classification.
The increasing prominence of Large Language Models (LLMs) in public discourse presents both opportunities and challenges for democratic deliberation. While red teaming strategies help mitigate specific risks, broader concerns persist regarding linguistic constraints, biases, and the sycophantic tendencies of LLMs. This chapter explores how LLMs can be used to significantly scale up and democratise deliberation, particularly in fostering inclusivity and empowering traditionally marginalised groups. Drawing on concepts from Systemic-Functional Linguistics, the chapter examines how variations across language users (for example, with respect to socio-demographic groups) and across language use (for example, with respect to communicative functions) shape participation in AI-supported deliberation. The chapter presents AI-driven deliberation studies and assesses their potential to scaffold argumentation, enhance access, and reduce the influence of exclusionary linguistic norms and biases which are embedded in prestigious registers. At the same time, the chapter cautions against both overclaiming, which leads to unrealistic expectations, and underclaiming, which risks missed opportunities for AI-assisted engagement. The chapter concludes by identifying future research directions to maximise the democratic potential of AI-assisted participation while embedding ethical safeguards to counteract the reproduction of linguistic inequalities.
Surzhyk, as a Ukrainian–Russian mixed subcode, has long been a subject of extensive research. Despite this, it remains an ambiguous phenomenon that elicits diverse assessments and attitudes. Within Ukrainian academic circles, a tradition has emerged of evaluating this communicative subcode through the prism of linguistic norms, predominantly characterising it as a negative phenomenon. However, in everyday communication, Surzhyk functions dynamically, reflecting the varied attitudes of its speakers. Significant shifts in its perception became particularly evident following the Russian Federation's full-scale invasion of Ukraine. Consequently, Russian-speaking individuals have increasingly adopted Surzhyk as a means of distancing themselves from the Russian language, which is widely perceived as the language of the aggressor. This article presents the results of a sociolinguistic study conducted in 2020–2021 and 2023–2024 among residents of the Odesa and Mykolaiv regions. The study examines the evolution of attitudes towards Surzhyk in the wake of the invasion, with particular attention given to the role of mass media discourse in shaping these perceptions. The analysis reveals a growing tendency towards a more favourable view of mixed speech; in the context of war, Surzhyk has come to symbolise solidarity and serves as an informal linguistic bridge for transitioning from Russian-dominant to Ukrainian-speaking communication.
Multisensory virtual nature immersion is emerging as a tool for supporting mental health management. In this study, we evaluate the effectiveness of two types of multisensory immersion systems for patients diagnosed with Post Traumatic Stress Disorder: (1) a PORTABLE system, which delivers olfactory stimulation directly through a head-mounted display (HMD), and (2) a POD system, in which participants are seated in a multisensory pod that augments HMD audio-visual content with synchronized olfactory cues and haptic feedback (wind and vibration). Beyond comparing the two modalities, this work focuses on how neural activity changes throughout the experience. Specifically, we analyze electroencephalography (EEG) responses in relation to participants’ subjective affective ratings, clinical questionnaires, and cognitive faculties.
Ο γλωσσικός πόρος san-Corpus περιλαμβάνει σώμα κειμένων γραπτού λόγου της Νέας Ελληνικής, έκτασης περίπου 9 εκατομμυρίων λέξεων. Ο πόρος συγκροτήθηκε στο πλαίσιο διδακτορικής διατριβής, με στόχο τη μελέτη των συγκρίσεων ομοιότητας στη Νέα Ελληνική. Μέγεθος & Πηγές Το corpus περιλαμβάνει τρία ισομεγέθη υποσώματα, ώστε να επιτρέπονται οι συγκρίσεις μεταξύ τους: (α) Δημοσιογραφικός Λόγος: 7.774 άρθρα από τέσσερις διαδικτυακές εφημερίδες (Η ΑΥΓΗ, Η ΚΑΘΗΜΕΡΙΝΗ, ΕΘΝΟΣ, ΤΟ ΒΗΜΑ - έτος 2015). Συνολική Έκταση: 2,9 εκατ. λέξεις. (β) Εκπαιδευτικός Λόγος: 96 σχολικά εγχειρίδια δημοτικού και γυμνασίου. Συνολική Έκταση: 3,4 εκατ. λέξεις (μελετώνται 2,8 εκατ.). (γ) Λογοτεχνικός Λόγος: 28 μυθιστορήματα (βραβεία αναγνωσιμότητας περιόδου 2010-2015). Συνολική Έκταση: 2,5 εκατ. λέξεις. Κατανομή Σχολικών Εγχειριδίων ανά Γνωστικό Αντικείμενο Πλήθος Εγχειριδίων Ελληνική Λογοτεχνία 13 Ελληνική Γλώσσα 12 Ιστορία 9 Φυσική – Χημεία – Βιολογία 9 Μαθηματικά 9 Γεωγραφία – Γεωλογία – Περιβάλλον 8 Θρησκευτικά 7 Αγωγή Αισθητική (Εικαστικά – Μουσική – Θέατρο) 15 Αγωγή Υγείας (Φυσική Αγωγή – Οικιακή Οικονομία) 5 Πληροφορική – Τεχνολογία 5 Αγωγή Κοινωνική – Πολιτική 3 Αγωγή Σταδιοδρομίας (ΣΕΠ) 1 Κατάλογος μυθιστορημάτων: Συγγραφέας, Τίτλος Έτος 1ης έκδοσης Δούκα, Μάρω - Το δίκιο είναι ζόρικο πολύ 2010 Θέμελης, Νίκος - Η συμφωνία των ονείρων 2010 Καρυστιάνη, Ιωάννα - Τα σακιά 2010 Μιχαλοπούλου, Αμάντα - Πώς να κρυφτείς 2010 Ελευθερίου, Μάνος - Πριν απ' το ηλιοβασίλεμα 2011 Ζουργός, Ισίδωρος - Ανεμώλια 2011 Μακριδάκης, Γιάννης - Η άλωση της Κωνσταντίας 2011 Μπουραζοπούλου, Ιωάννα - Η ενοχή της αθωότητας 2011 Πανσέληνος, Αλέξης - Σκοτεινές επιγραφές 2011 Παπαδημητρίου, Χίλντα - Για μια χούφτα βινύλια 2011 Παπαθεοδώρου, Θοδωρής - Οι καιροί της μνήμης 2011 Τριανταφύλλου, Σώτη - Για την αγάπη της γεωμετρίας 2011 Φακίνος, Μιχάλης - Η έρημος έρχεται 2011 Βαμβουνάκη, Μάρω - Κυριακή απόγευμα στη Βιέννη 2012 Διβάνη, Λένα - Εγώ, ο Ζάχος Ζάχαρης 2012 Στεφανάκης, Δημήτρης - Φιλμ νουάρ 2012 Ακρίβος, Κώστας - Αλλάζει πουκάμισο το φίδι 2013 Ζέη, Άλκη - Με μολύβι φάμπερ νούμερο δύο 2013 Κορτώ, Αύγουστος - Το βιβλίο της Κατερίνας 2013 Κωνσταντούρου, Μαρία - Αγεφύρωτες σιωπές 2013 Μαντά, Λένα - Με λένε Ντάτα 2013 Ξανθούλης, Γιάννης - Κωνσταντινούπολη των ασεβών μου φόβων 2013 Ρώσση–Ζαΐρη, Ρένα - Άρωμα βανίλιας 2013 Ανδρουλάκης, Μίμης - Αλλέγκρα 2014 Δημουλίδου, Χρυσηίδα - Το κελάρι της ντροπής 2014 Παπαδοπούλου, Ελισάβετ - Μέρες και νύχτες που δεν ήταν δικές μας 2014 Χατζή, Αθηνά - Η θάλασσα έφυγε 2014 Χωμενίδης, Χρήστος - Νίκη 2014 Τεχνικές προδιαγραφές & Μορφότυπος Για την αναπαράσταση των δεδομένων και των μεταδεδομένων υιοθετήθηκε η πολυεπίπεδη οπτική των XML σχημάτων και τροποποιήθηκε το διεθνές πρότυπο TEI P5, 4.0.0 (Text Encoding Initiative). Δημιουργήθηκε ειδικός χώρος ονομάτων sanCorpus (sanC) με σχήμα τύπου RELAX-NG. Το σώμα κειμένων διατίθεται σε TXT και σε XML σε τρεις εκδοχές: Βάθος 0: απλό κείμενο (TXT). Περιλαμβάνει το main core (κείμενο βάσει του οποίου εξετάζονται οι συγκρίσεις ομοιότητας) και το out of core (κείμενο εκτός εμβέλειας της διατριβής, στο οποίο περιλαμβάνονται κείμενα που πλαισιώνουν το κυρίως κείμενο, π.χ. κείμενα διδασκαλίας, πίνακες περιεχομένων, εξώφυλλα) Βάθος 1 = κείμενα στην απλούστερη δυνατή XML κωδικοποίηση Βάθος 2 = κείμενα με πιο λεπτομερείς XML κωδικοποιήσεις. Αυτή η έκδοση (san-Corpus v1.0, Depth 0: Plain Text) περιλαμβάνει το σώμα κειμένων σε μορφή απλού κειμένου (Βάθος 0) στην αρχική του διάταξη (βλ. Επεξεργασία). Στόχος είναι ο σταδιακός εμπλουτισμός με επισημειωμένα δεδομένα, καθώς και με τις εκδοχές Βάθους 1 και 2. Επεξεργασία (Processing) Η μεθοδολογία συλλογής των δεδομένων, η θεωρητική τεκμηρίωση και το σχήμα επισημείωσης επεξηγούνται στις μελέτες Αφεντουλίδου (2022, 2021, 2013, 2012) και Afentoulidou (2009). Η πρώτη εκδοχή (Βάθος 0) χρησιμοποιήθηκε αποκλειστικά για τη λημματοποίηση που απαιτούσε η Collostruction Analysis (Gries, 2024). Στάδια επεξεργασίας (για τη λημματοποίηση): Τμηματοποίηση σε προτάσεις (sentence segmentation) με τη χρήση της βιβλιοθήκης Stanza (Stanford NLP Group, Qi et al. 2020), η οποία βασίζεται στο μοντέλο Greek Dependency Treebank (GDT) του Ινστιτούτου Επεξεργασίας του Λόγου / ΕΚ «Αθηνά». Τυχαία αναδιάταξη των προτάσεων για την προστασία της ακεραιότητας των πρωτότυπων έργων. Λημματοποίηση με τον ILSP Lemmatizer μέσω της Υποδομής Clarin-EL. Για την ανάλυση συμφράσεων του δείκτη σαν απομονώθηκαν συγκεκριμένοι λεκτικοί τύποι. Αδειοδότηση & Δικαιώματα Το san-Corpus συγκροτήθηκε για τις ανάγκες της διδακτορικής διατριβής και προστατεύεται από το δικαίωμα ειδικής φύσης σύμφωνα με την Οδηγία 96/9/ΕΟΚ και το άρθρο 45Α του Ν. 2121/1993. Η χρήση του περιεχομένου γίνεται αποκλειστικά για ερευνητικούς σκοπούς βάσει των εξαιρέσεων της Οδηγίας 2001/29 και της Οδηγίας (ΕΕ) 2019/790 (Text and Data Mining exceptions / Fair Use). Η πρόσβαση είναι περιορισμένη (Restricted Access) και παρέχεται αποκλειστικά σε μέλη της ακαδημαϊκής κοινότητας για σκοπούς επαλήθευσης των αποτελεσμάτων της διατριβής και περαιτέρω μη εμπορική έρευνα. Η πηγή προέλευσης δικαιούται να ζητήσει οποιαδήποτε τροποποιητική ενέργεια (π.χ. αφαίρεση) επί του πρωτότυπου περιεχομένου. Τέλος, η άδεια CC BY-NC-ND 4.0 ισχύει για την επιμέλεια (curation), τα μεταδεδομένα και τη γλωσσολογική επισημείωση του σώματος κειμένων. Βιβλιογραφικές αναφορές Αφεντουλίδου, Β. (2022). Σώμα ελληνικών κειμένων για τη μελέτη δομών ομοιότητας της Νέας Ελληνικής: σχεδιασμός και υλοποίηση. Στο Πρακτικά του 10ου Συνεδρίου Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας (σσ. 67-94). ΕΚΠΑ. Αφεντουλίδου, B. (2021). Δομές ομοιότητας στη Νέα Ελληνική. Σωματοκειμενικές παρατηρήσεις για τον πολυλειτουργικό δείκτη σαν. Προφορική ανακοίνωση στην 41η Ετήσια Συνάντηση του Τομέα Γλωσσολογίας, 13–15 Μαΐου 2021. ΑΠΘ. Αφεντουλίδου, Β. (2013). Και σου απάντησα κάτι σαν ‘τέλεια, εντάξει’. Δείκτης σαν + ευθύς λόγος;. Προφορική ανακοίνωση στο 7ο Συνέδριο Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας, 16–18 Μαΐου. ΕΚΠΑ. Αφεντουλίδου, Β. (2012). Συγκρίσεις ομοιότητας στα Νέα Ελληνικά: ο δείκτης σαν. Στο Z. Gavriilidou, A. Efthymiou, E. Thomadaki & P. Kambakis-Vougiouklis (Επιμ.), Selected papers of the 10th International Conference on Greek Linguistics (σσ. 696-707). DUTH. Afentoulidou, V. (2009). Sketching the σαν conditional construction in Modern Greek. Submitted essay, 2009 Linguistic Institute, Linguistic Structure and Language Ecologies, Linguistic Society of America and UC Berkeley. Gries, Stefan Th. 2024. Coll.analysis 4.1. A script for R to compute perform collostructional analyses. https://www.stgries.info/teaching/groningen/index.html Institute for Language and Speech Processing - Athena Research Center (2015). ILSP Lemmatizer. Version 1. [Software (Tool/Service)]. CLARIN:EL. http://hdl.handle.net/11500/ATHENA-0000-0000-23EE-D Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 101–108). Online: Association for Computational Linguistics.