Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Cement and concrete account for around 8% of global CO2 emissions and remain among the most hard-to-abate sectors, alongside steel and chemicals. Demand for these materials is projected to grow substantially over the coming decades, particularly across the Global South, driven by rapid urbanization and infrastructure development. While numerous decarbonization technologies and strategies have emerged and are being implemented, the absence of quantitative, context-specific definitions of low-carbon concrete make it difficult to quantify the extent of decarbonization, especially in developing economies. Using India as an illustrative case, where cement production is projected to grow roughly five-fold by 2070, this perspective examines why low-carbon concrete ratings and definitions applied in the developed world cannot be directly implemented in developing-country contexts. It also proposes a phased strategy towards quantitative definitions for the organized and unorganized concrete sectors. This offers a practical pathway to closing the definitional gap that currently limits the efficacy and accounting of cement and concrete sector decarbonization efforts in the global south.
Abstract Cities are experienced in motion, yet urban soundscape research has largely assumed stationary viewpoints, overlooking the perceptual role of the walking perspective. Here, we show that this mismatch fundamentally blinds us to how walking views reshape urban soundscape perception. Using a within-subject, repeated-measures audiovisual design, 34 participants evaluated 18 urban streets under standing-view (SV) and walking-view (WV) conditions paired with identical binaural audio, providing both continuous real-time affective ratings and retrospective soundscape evaluations. We found that, first, rapid perceptual stabilisation occurred: taking the standing view as the baseline, real-time affective responses under the walking view initially diverged but consistently converged within the first 10 s across all participants and streets. Second, despite this early stabilisation, the walking perspective reshaped overall perceptual outcomes by selectively reweighting the salience of urban sounds. Among streets showing significant effects, transient sounds, including alarms and sirens, became more salient, whereas continuous background sounds, including traffic and human activity, became less salient. The walking perspective also polarised overall sound environmental evaluations, making positively evaluated streets more positive and negatively evaluated streets more negative, while reducing inter-individual variability. These findings demonstrate that the walking perspective is an active component of urban soundscape perception, shaping both the temporal dynamics of perceptual adaptation and the overall perceptual weighting of urban sound environments, with broader implications for understanding how environmental perception unfolds in motion.
Artificial intelligence (AI)-personalized learning is often promoted as a way to match instruction to learner differences, yet its educational value should not be judged only by efficiency or test performance. This narrative review uses algorithmic belonging as a conceptual lens to examine how AI-personalized learning may shape students’ inclusion, identity, and participation in diverse classrooms. Across the reviewed literature, personalized and AI-supported systems are typically associated with adaptation, feedback, recommendation, and learner support, while belonging research emphasizes acceptance, recognition, cultural affirmation, and meaningful participation. Direct studies that explicitly connect AI personalization with belonging remain limited; most available evidence must therefore be synthesized from adjacent work on AI in education, personalized learning, school belonging, culturally relevant pedagogy, and diverse-classroom participation. The literature suggests that AI personalization may support inclusion when it increases access, responsiveness, and student agency, but it may undermine belonging when it stereotypes learners, obscures decision rules, or reproduces cultural and linguistic norms that marginalize some groups. The strongest implication is that personalization should be treated as a social and ethical design problem, not merely a technical one. AI systems are most likely to support belonging when they are explainable, fair, culturally sustaining, and embedded in teacher–student relationships that affirm identity and widen participation. Keywords: artificial intelligence; personalized learning; belonging; inclusion; identity; participation; diverse classrooms; culturally relevant pedagogy; algorithmic bias; educational equity
In filmmaking, a neutral face paired with an emotional context can convey a congruent emotion-a phenomenon known as the Kuleshov effect. However, past research has yielded mixed findings on the existence of the Kuleshov effect, possibly due to methodological variability; we therefore implemented a boundary-test paradigm that combined normed static context images with dynamic facial stimuli separated by an explicit context-rating step to examine how visual context shapes the evaluation and categorization of faces under such conditions. We hypothesized that positive context would elevate facial valence ratings (and vice versa), while evoking different-than-neutral emotions in a neutral face when categorized. Thirty-two participants first rated the valence of a context image on a 7-point Likert scale. They then viewed a video of a neutral face, evaluated its valence on the same scale, and explicitly categorized the expression by selecting one of five predefined emotion labels (neutral, angry, happy, sad, disgusted). Analyses using linear mixed-effects models confirmed that positive contexts led to higher facial valence ratings, while negative contexts led to lower ones; by contrast, multinomial regression revealed no effect of context on emotional categorization. Our findings suggest that while visual context can bias evaluative judgments of facial expressions, it does not necessarily alter their emotional categorization under conditions where cinematic continuity is absent, highlighting the boundaries of the Kuleshov effect in line with previous studies.
Language is a living organism that evolves alongside technological and social advancements. This paper examines the phenomenon of neologisms—newly coined words or expressions—and their pervasive role in contemporary English mass media. The study categorizes recent neologisms based on their morphological formation processes, such as blending, compounding, and functional shift. Furthermore, it analyzes how mass media acts as a primary catalyst for the popularization of these terms. By investigating digital journals, social media platforms, and news broadcasts, the research highlights the pragmatic functions of neologisms in creating concise, engaging, and culturally relevant communication. The findings provide insights into the current trends of English lexicology and the impact of the digital age on linguistic norms.
The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, especially coreference and discourse relations. We present its second consolidated version (PDT-C 2.0), which concludes almost 30-years long project of sustained development of the resource to a uniformly and coherently annotated, genre-diversified, almost 4 million token language resource of Czech language, with accompanying fully compatible lexicons. In addition to continuous linguistic research, the richly linguistically annotated corpus is also widely used in international comparisons of the development of traditional and novel NLP tools as well as in conversions into other formalisms. The corpus and the trained parsers are available under the CC BY-NC-SA licence.
Arabic WordNet 4.0 is a comprehensive lexical database for Modern Standard Arabic, derived from the Open English WordNet using the expand approach. Features:109,901 synsets (partial OEWN 2024 coverage — satellite adjectives not yet included)~124,653 lexical entries~166,843 senses~265,676 synset relations97.3% ILI coverage for cross-linguistic linkingFull WN-LMF 1.4 XML format compliance Note: This is an intermediate release (v4.0.1) representing AWN4 after upper-ontology noun additions (+78 synsets) but before satellite adjective coverage was completed. For the complete OEWN 2024 parity release (120,630 synsets), see v4.1.0. Methodology:Translations were generated using AI-assisted translation (Google Gemini 3 Pro Preview) following the expand approach for WordNet construction. Attribution: This resource is derived from Open English WordNet (CC BY 4.0), which is based on Princeton WordNet 3.0.
The article is devoted to the theoretical substantiation of the essence of the grammatical aspect of foreign language speech as a key component of foreign language communicative competence. The relevance of the study is determined by the need to understand the structural content of the grammatical aspect of speech in the context of the requirements of modern educational standards. The paper presents a comparative analysis of the approaches of foreign and Russian researchers to understanding the grammatical aspect of speech. D. Larsen-Freeman's three-dimensional model, which includes form, meaning, and use of grammatical phenomena, is examined and illustrated with examples. The position of S. Thornbury, who defines grammar through morphology and syntax and emphasizes its meaning-making potential realized in representational and interpersonal functions, is analyzed. Attention is paid to M. Lewis's lexical approach, which assigns a secondary role to grammar. The article presents the views of Russian methodologists (N.D. Galskova, N.I. Gez, E.N. Solovova), who consider grammar as a fundamental component of speech activity that ensures practical language proficiency for solving communicative tasks. Based on the analysis conducted, the author formulates a definition of the grammatical aspect of foreign language speech as a complex of automated actions for selecting, combining, and using grammatical structures in accordance with communicative intention and language norms. It is concluded that an insufficient level of mastery of the grammatical aspect of speech leads to difficulties in the formation of foreign language communicative competence as a whole.
This article explores the major problems and distinctive features of lexical change in Global English in the context of globalization, intercultural communication, and ongoing linguistic variation. As English increasingly functions as a global lingua franca, it undergoes continuous transformation influenced by social, cultural, technological, and economic factors. The study examines key processes such as lexical innovation, borrowing from other languages, semantic shift, word formation, and hybridization, which contribute to the dynamic expansion of the English vocabulary. Special attention is given to the tension between standardization and localization, where global norms of English interact with local linguistic identities and cultural practices. The article also highlights the challenges lexical change poses for linguistic norms, language teaching, dictionary compilation, and cross-cultural understanding. Drawing on contemporary linguistic theories and empirical research, the study provides a comprehensive overview of how Global English reshapes and diversifies the English lexicon in different regions of the world.
The fast-paced nature of digitalization and new global media platforms comes with a radical reshaping of the behavior of users when it comes to language, which in turn gives rise to new models concerning linguistic norms and makes necessary a full-spectrum overview on this change at play on the level of language. The relevance of the subject is given by online communication as it largely sets today’s language standards and thus substitutes conventional channels of standardisation. The purpose of the work is to characterize the specificity of the impact on modern Ukrainian linguistic norms in digital platforms, and its object are language processes in new digital spaces. The approach is a mixture of content analysis of online platforms, analysis of international statistical indicators and comparison in variation in the strength of linguistic innovations across areas. Study revealed that language dynamics are rather non-uniform in digital area and thanked to assessment of communication environment structure, but not frequency of its usage. This indicates, that the greatest linguistic variation is observed on social networks and multilingual web spaces, while private channels are characterized by reduce rate of innovation. Hybrid lexical and grammatical forms emerge when driven by multimodality, reaction time, and algorithmic properties of the content. It is well known that the “platform norm” is imposed by language voting far more than academia. The estimates of the integrated index of language change intensity testify the dependence of evolution of the language norm on multi-dimensional interaction between social, technological and algorithmic factors. The practical implication of the findings is that they can be leveraged to make predictions about how languages might evolve, to inform digital language policy and to develop tools for monitoring online communication.
The study aimed to identify the peculiarities of the influence of digital technologies on the evolution of linguistic worldviews by analyzing changes in language production, evaluative polarity, and the semantic organization of speech across different types of text environments. The study was conducted as a controlled experimental intergroup comparison involving native speakers of Ukrainian who interacted with both digital and non-digital informational texts. Quantitative indicators of lexical diversity, utterance length, evaluative markedness, and associative network parameters were used for analysis, along with qualitative analysis of semantic shifts and linguistic norm variability. The results showed that there are systemic differences in language production across discursive environments. Interaction with digital texts is accompanied by a decrease in lexical differentiation, an increase in the evaluative and emotional components of speech, and an increase in the density of associative networks, with a simultaneous narrowing of conceptual detail. The semantic-axiological shifts identified indicate a transformation in the ways of conceptualizing reality and a stabilization of subjective-evaluative interpretations in linguistic worldviews in the context of digital communication. The results are consistent with the principles of cognitive linguistics and theories of digital discourse, while refining them based on empirical data from an experimental study.
Old English (OE) has traditionally been regarded as an extraordinarily complex language, for its intricate syntactic and morphological relations remain a conundrum for many historical linguists.This explains the substantial number of manuals which primarily focus on the basics of grammar (Mitchell and Robinson 1992; Hog 2012;Fulk 2014).Nevertheless, for those enthusiasts who would like to deepen their insight of Old English structures, Ojanguren López book constitutes a stimulating work which explicitly delves into one of the most elaborate aspects of syntax and semantics: the competition on the complementation of OE aspectual and manipulative verbs, namely, the ones that correspond to contemporary English aspectual end, try and fail, along with manipulative forbid, hinder and refrain.This research also addresses the question of "the rise of serial verb constructions in English" (Ojanguren López 2024, 17) and provides some clarification on the semantics of the verbs linked to complementation patterns, which represents a valuable addition to prior contributions (Callaway 1913;Mitchell 1985;Molencki 1991;Denison 1993;Martín Arista 2022).The aforementioned monograph is divided in nine chapters, which facilitates the comprehension of the matter of study, as it is presented sequentially and includes a well-defined methodology which is supported by the theoretical underpinnings of the research.Thus, chapter one overviews the major focus of the undertaking as well as its analytical structure, and contextualises the complementation patterns of verbs of aspect and manipulation in Old English within its most recent framework.Once the fundamentals have been established, chapter two underlines the relevance of the work that has been carried out considering several dimensions (synchrony, diachrony, typology), which contribute to a comprehensive approach of the objects of research.Furthermore, the lexicographical and textual sources specified in said chapter include dictionaries, lexical databases, grammars, translations, and lemmatisers.In order to describe the competition existing on the complementation of Old English verbs, the Role and Reference Grammar (RRG) and its Interclausal Relations Hierarchy (Van Valin and LaPolla 1997;Van Valin 2005) are proposed, for its "model of the interaction between morphology and syntax, including constituents and operators" is paramount for a language such as Old English (Ojanguren López 2024, 26).Additionally, several works on the clausal complementation of verbs are revisited to discuss different types of competition (Callaway 1913;Denison 1993;Molencki 1991;Los 2005) while introducing nominal complementation and verb serialisation as the innovative elements of the study.As a means to expand the foundations sustaining the analysis on OE verbal complementation, Ojanguren López reviews in chapter three the most relevant parts of the theory of RRG, as the former is central for the consideration of semantic and pragmatic categories in syntax, and for the concern of cross-linguistic validity.The book also devotes special attention to lexical representation, semantic roles and macroroles, which play a crucial role in the association of semantics and syntax, as well as the notion of Aktionsart (Brugmann 1922), for it constitutes "the semantic representation of the sentence" (Ojanguren López 2024, 37).Additionally, grammatical relations embodied in the function of Privileged Syntactic Argument (PSA) and clause structures represented by the Layered Structure of the Clause (LSC), which distinguishes multiple forms of complex referential clauses, complete RRG theory and indubitably help the reader understand the basis for the subsequent study, reinforced by the concepts of juncture and nexus acting as the main aspects of the theory of complex sentences.Chapter four, in turn, intends to become a further contribution towards "a more semantically oriented syntax of Old English" (Ojanguren López 2024, 57), hence it revolves around preceding research in the structural-functional analysis of Old English (Martín Arista 2000a, 2000b, 2020) with the aim of applying these earlier advances to the linking of semantics and syntax proposed by RRG.To successfully accomplish this, terminological remarks such as Verb Classes and Alternations (Levin 1993), linking and its
This article investigates the comparative linguistic features of artificial intelligence (AI)-based translation systems, focusing on how machine translation models render linguistic structures across typologically different languages, particularly Uzbek and English. The study analyzes syntactic, lexical, semantic, and pragmatic shifts occurring in AI-generated translations and compares them with human translation norms. Special attention is given to neural machine translation (NMT) systems and their ability to handle idiomatic expressions, polysemy, and word order variations. The research highlights both the strengths and limitations of AI translation tools, emphasizing issues such as loss of cultural nuance, structural simplification, and contextual misinterpretation. The findings suggest that although AI translation systems have significantly improved in fluency and accuracy, they still require linguistic refinement to fully capture deep structural and cultural meanings across languages.
En el contexto de las humanidades digitales y la lingüística de corpus, el acceso eficiente a recursos lingüísticos distribuidos geográficamente ha sido históricamente un desafío. Los investigadores a menudo se enfrentan a la necesidad de navegar por múltiples interfaces, protocolos y políticas de acceso para consultar corpus, árboles sintácticos (treebanks) o diccionarios alojados en distintas instituciones. La infraestructura europea CLARIN (Common Language Resources and Technology Infrastructure) ha desarrollado, para mitigar este problema, la Federated Content Search (FCS), un sistema de búsqueda federada que permite consultar simultáneamente cientos de recursos heterogéneos desde una única interfaz unificada (CLARIN, 2025a). La FCS se fundamenta en una arquitectura cliente-servidor distribuida y estandarizada. En el núcleo de esta arquitectura se encuentran los endpoints FCS, que actúan como intermediarios técnicos entre la interfaz central de búsqueda y los motores de búsqueda locales de cada centro CLARIN participante. La comunicación se realiza mediante protocolos estandarizados: SRU (Search/Retrieve via URL) para el transporte de las consultas y CQL (Contextual Query Language) para la formulación de patrones de búsqueda, lo que garantiza la interoperabilidad entre sistemas heterogéneos. Cuando un usuario introduce una consulta en el portal central, esta se distribuye en paralelo a todos los endpoints disponibles. Cada endpoint traduce la consulta al formato de su motor local, la ejecuta sobre sus recursos (corpus, diccionarios, etc.) y devuelve los resultados estandarizados al agregador central, que los presenta al usuario en una vista consolidada (CLARIN, 2025b). Actualmente, la FCS permite acceder a más de 500 corpus y treebanks en aproximadamente 200 lenguas, así como a unos 45 diccionarios en 11 lenguas gracias a la especificación LexFCS (Körner et al., 2025). La exploración práctica de la FCS revela tanto su potencial como sus limitaciones para la investigación lexicográfica y corpus-lingüística. En una búsqueda exploratoria del neologismo político "Ambazonia", el sistema consultó 586 recursos de 25 instituciones, retornando resultados de 12 fuentes en tiempo real. La interfaz proporciona información detallada sobre el estado de cada endpoint (respuestas exitosas, errores de protocolo o falta de autenticación), lo que demuestra la transparencia del sistema pero también expone los desafíos de los entornos distribuidos: incompatibilidades técnicas en las versiones del protocolo SRU y restricciones de acceso a recursos protegidos. A pesar de ello, la FCS permitió obtener 88 ocurrencias relevantes, incluyendo artículos de Wikipedia, publicaciones científicas y noticias, y facilitó la exportación de los resultados en formatos como CSV, JSON o XML para su análisis posterior. Una de las funcionalidades más potentes de la FCS es la capacidad de realizar búsquedas sobre capas de anotación lingüística (lema, categoría gramatical, características morfológicas, etc.). Mediante consultas CQL o un editor gráfico, el investigador puede formular preguntas complejas, como la búsqueda del lema alemán "Artznei" (variante histórica de 'medicina') en co-ocurrencia con verbos que comienzan por "her-" en corpus históricos del alemán. Los resultados obtenidos del corpus DTA/DWDS muestran no solo las ocurrencias, sino también las variantes ortográficas ("Artzeneyen", "Arzenei") y las colocaciones valorativas ("herbe Arzneien", "herrliche Artzeneyen"), evidenciando el valor de la anotación multinivel para el estudio del cambio lingüístico y los discursos especializados. Además, la FCS permite exportar estos resultados en forma de tablas token-alineadas que incluyen el texto original, el lema, el POS-tag y formas normalizadas, lo que convierte al sistema en una herramienta de extracción de datos lingüísticos estructurados. La búsqueda sobre recursos léxicos, habilitada por LexFCS, amplía aún más el alcance de la plataforma. Una consulta por el lema "father" (inglés) a través de la FCS interroga simultáneamente recursos tan diversos como WordNet, diccionarios alemán-español, y diccionarios de lenguas africanas como el akan, isiXhosa o isiNdebele. Los resultados no solo ofrecen definiciones o traducciones, sino que en el caso de WordNet presentan una red semántica completa: sinónimos ("pope", "Holy Father"), hiperónimos ("spiritual leader") e hipónimos ("antipope"), así como acepciones metafóricas ("founder"). Esta diversidad de fuentes permite estudios comparativos transculturales y la observación de distintos grados de estructuración lexicográfica. Sin embargo, la disponibilidad desigual de capas de anotación lingüística y la limitación técnica a diez resultados por endpoint (para garantizar el rendimiento) imponen restricciones a los estudios cuantitativos exhaustivos. La evolución de la FCS se ejemplifica con la interfaz de demostración Text+ Pilot Project Search Interface (TPPSSI), una "zona de pruebas" que incorpora funcionalidades avanzadas aún no disponibles en el portal principal: generación de enlaces permanentes (permalinks), historial de búsquedas, búsqueda de entidades y ejemplos precargados. La TPPSSI, construida sobre Vuetify, presenta una disposición más intuitiva donde los recursos activos se resaltan visualmente y se integra una vista especializada "Lex Data View" para recursos lexicográficos. Esta interfaz ilustra la dirección futura de la FCS, orientada a mejorar la experiencia de usuario y a proporcionar herramientas más robustas para la investigación digital. En conclusión, la FCS de CLARIN representa una puerta de entrada fundamental a los recursos lingüísticos distribuidos. Su principal fortaleza reside en ofrecer un punto de acceso único y estandarizado a un ecosistema diverso y disperso de corpus, árboles sintácticos y diccionarios. Para la lexicografía y las humanidades digitales, la FCS facilita la investigación exploratoria, el análisis contrastivo y la extracción de datos a gran escala, reduciendo drásticamente los tiempos de búsqueda. Se confirma así que la FCS es una infraestructura clave para la lexicografía y las Humanidades Digitales dentro del ecosistema CLARIN y CLARIAH, cuyo potencial futuro depende de la homogeneización técnica y de anotación de los recursos integrados.
There is increasing recognition of the value of multi-agency approaches to support people with hoarding difficulties. This service evaluation describes and presents preliminary evaluation data for an innovative multi-agency partnership in Wales supporting people with hoarding difficulties. The preliminary service evaluation analysed routine usage and outcome data recorded between January 2024 to July 2024. This included the number of individuals receiving support, referral acceptance rates, demographics of referrals and a range of service outcomes including changes in the quantity of possessions in living spaces, measured using the Clutter Image Rating (CIR) scale, and achievement of service objectives. The service was well-utilised, with 53% achieving positive outcomes upon discharge. The quantity of possessions in living spaces decreased from clinical threshold in 71% of participants from pre- to post-intensive support. The findings are considered in relation to the evidence base for multidisciplinary approaches for hoarding difficulties and the implications for service delivery. Mae gwerth dulliau amlasiantaethol o gefnogi pobl ag anawsterau celcio yn cael ei gydnabod fwyfwy. Mae’r astudiaeth gwerthuso gwasanaeth hon yn disgrifio ac yn cyflwyno data gwerthuso rhagarweiniol ar gyfer partneriaeth amlasiantaeth arloesol yng Nghymru sy’n cefnogi pobl ag anawsterau celcio. Roedd y gwerthusiad rhagarweiniol o’r gwasanaeth yn dadansoddi data defnydd a chanlyniadau arferol a gofnodwyd rhwng mis Ionawr 2024 a mis Gorffennaf 2024. Roedd hyn yn cynnwys nifer yr unigolion sy’n cael cymorth, cyfraddau derbyn atgyfeiriadau, demograffeg atgyfeiriadau ac amrywiaeth o ganlyniadau gwasanaeth gan gynnwys newidiadau yn nifer yr eiddo mewn mannau byw, wedi’u mesur gan ddefnyddio graddfa Sgorio Delwedd Annibendod (CIR), a chyflawni amcanion y gwasanaeth. Roedd defnydd y gwasanaeth yn dda, gyda 53% yn sicrhau canlyniadau cadarnhaol ar ôl cael eu rhyddhau. Gostyngodd nifer yr eiddo mewn mannau byw o drothwy clinigol ymysg 71% o gyfranogwyr cyn cael cymorth dwys ac ar ôl ei gael. Ystyrir y canfyddiadau yng nghyswllt y sylfaen dystiolaeth ar gyfer dulliau amlddisgyblaethol o ran anawsterau celcio a’r goblygiadau ar gyfer darparu gwasanaethau.
Social interactions are dynamic and complex, relying on tracking variation in your own and your partner’s affective and mental states. Being “in-sync” neurally and building a shared consensus can mark successful social interactions. In addition, social anxiety can impact social experiences. We used a novel, naturalistic paradigm to investigate relations between inter-brain neural similarity within mentalizing regions, affective similarity and social anxiety symptoms. Undergraduate student friend pairs (N = 34, 85% White, 65% Women) engaged in a social interaction while videos previously captured from their individual perspectives were recorded. Participants watched clips of the social interaction from both their own perspective (their personal view of the social interaction) and their friend’s perspective (their friend’s personal view of the social interaction) while fMRI data were collected. They rated their affect after each clip. Inter-brain neural similarity was computed across three conditions: 1) Same-Stimuli: both participants viewed identical visual stimuli as in a traditional neural similarity paradigm, 2) Self-Perspective: both participants viewed the clip from their own perspective like the originally experienced social interaction and 3) Friend-Perspective: both participants viewed the clip from their friend’s perspective, a novel perspective. Participants self-reported their social anxiety symptoms and affective similarity captured affect rating concordance within dyads. In contrast to the Same-Stimuli condition, when participants viewed the clips like they originally experienced social interaction (Self-Perspective), greater affective similarity was associated with greater inter-brain neural similarity. When participants viewed the clips from a novel perspective (Friend-Perspective), participants lower in social anxiety symptoms exhibited greater inter-brain neural similarity with greater affective similarity; whereas, participants higher in social anxiety symptoms exhibited greater inter-brain neural similarity with less affective similarity. The results suggest stimuli from socially relevant perspectives, rather than identical stimuli, may reveal more nuanced brain-behavior dynamics, allowing for a better understanding of individual differences in socioemotional experience.
This article examines youth language (Jugendsprache) as a significant and dynamic component of contemporary German. The study aims to analyze its structural, lexico-semantic, and functional features, as well as its role in shaping modern linguistic trends. The methodological framework combines descriptive, comparative, and lexico-semantic analysis, along with the examination of digital discourse, enabling a comprehensive investigation of youth language in its natural communicative environment. The findings demonstrate that Jugendsprache is characterized by a high degree of lexical innovation, driven by anglicisms, neologisms, and abbreviations emerging in digital communication. Word-formation processes, including compounding, affixation, and conversion, exhibit increased creativity and hybridization, often deviating from standard linguistic norms. Youth language also performs important sociolinguistic functions, such as identity construction, emotional expression, and the differentiation of social groups. The study highlights the crucial role of the digital environment in accelerating linguistic change and shaping new communicative practices. While Jugendsprache contributes to the enrichment and adaptability of the German language, it may also lead to challenges related to normativity and intergenerational communication. Overall, youth language is interpreted as a “laboratory of linguistic innovation” and a mediator between linguistic change and standardization, reflecting the broader processes of globalization and digitalization in modern society.
The article presents a corpus-based empirical analysis of colloquial units in contemporary English, treating colloquial vocabulary as a dynamic, multifunctional, and internally heterogeneous subsystem of the lexical system. Colloquial units are examined not merely as markers of informal speech, but as linguistically significant elements reflecting ongoing transformations in communicative practices, discourse conventions, and stylistic norms. Drawing on data from large-scale, register-diverse English language corpora, the study investigates the frequency, dispersion, contextual variability, and pragmatic functions of selected colloquial units across spoken and written registers. Special attention is paid to processes of stylistic diffusion, pragmatic refunctionalization, and partial desemanticization.
The COVID-19 pandemic has radically transformed the English language introducing new lexical terms, semantic changes, and pragmatic norms that are still effective in the post-pandemic world. This scientific review explores new tendencies in writing English after the pandemic and concentrates on the definition of vocabulary, discourse, and pragmatic adjustments in social, professional, educational, and digital environments. The paper summarizes recent linguistic studies to examine the way in which communication driven by crisis made the formation of neologisms, borrowing, compounding, and semantic recontextualization faster. Keywords of health, risk, remote interaction, and digitalization became used not only in specialized registers but also daily and the meaning of already existing words has been broadened or metaphorical. In addition to the use of lexis, the review identifies shifts in pragmatic practices such as the changes in politeness strategies, the manifestation of uncertainty and empathy, and the taming of crisis sensitive discourse in institutional and interpersonal communication. The heightened use of digital platforms has continued to affect turn taking, modality and interpersonal stance, transforming the rules of formality and interaction. As well, the English after the pandemic is indicative of the wider sociocultural change including increased attentiveness to mental health, work-life separation, and group responsibility, which linguistically are marked by evaluative and stance-marking categories. There are also pedagogical and applied implications in the review that the English language teaching and the professional communication training programs need to consider such changing norms. In spite of increasing attention, studies in this field are still disjointed, with little longitudinal and cross-cultural studies. This review finds that the post-pandemic English is a new stage of linguistic adaptation under the influence of the world crisis experience, which can contribute to the useful conclusions about the interconnection between the linguistic change, social unsteadiness, and communicative stability.
Static concreteness ratings are widely used in NLP, yet a word's concreteness can shift with context, especially in figurative language such as metaphor, where common concrete nouns can take abstract interpretations. While such shifts are evident from context, it remains unclear how LLMs understand concreteness internally. We conduct a layer-wise and geometric analysis of LLM hidden representations across four model families, examining how models distinguish literal vs figurative uses of the same noun and how concreteness is organized in representation space. We find that LLMs separate literal and figurative usage in early layers, and that mid-to-late layers compress concreteness into a one-dimensional direction that is consistent across models. Finally, we show that this geometric structure is practically useful: a single concreteness direction supports efficient figurative-language classification and enables training-free steering of generation toward more literal or more figurative rewrites.
Part of speech and syntactically annotated dataset for modern Mongolian. The dataset is a fully annotated corpus of modern Mongolian texts written in Mongolian Cyrillic. Version 2.
Social media platforms, particularly TikTok, have become primary arenas for linguistic experimentation among adolescents, yet systematic analyses of how platform-specific affordances shape lexical and semantic innovation remains limited. This study investigated lexical and semantic variations in adolescent digital communication on TikTok, addressing three research questions concerning the types of lexical innovations, processes of semantic change, and the role of platform affordances in shaping language evolution. Methods: A mixed-methods design integrated quantitative corpus linguistics with qualitative discourse analysis. A corpus of 2,848 TikTok comments was compiled across four major trends (September–December 2024). Lexical analysis identified neologisms, graphical variations, and acronyms; semantic analysis documented broadening, narrowing, metaphoric extension, and pejoration/amelioration; platform affordances analysis examined meme-driven language and intertextual policing. Analysis revealed 15 lexical innovations with 63 occurrences across semantic categories. Neologisms (fr, bestie, delulu) and graphical variations (tryna, cuz, ion) served dual functions of efficiency and identity performance. Semantic shifts included ameliorative broadening (slay, fire), pejoration (basic, cringe), metaphoric extension (era, main character), and reclamatory usage (ghetto). Platform analysis identified 11 meme-driven phrases generating 2,848 occurrences with near-neutral sentiment, and 347 policing instances (12.2%) concentrated during rising and peak trend phases, demonstrating active semantic negotiation through definition, debate, and correction. TikTok functions as an accelerated laboratory for language change where adolescents deploy multiple mechanisms of linguistic innovation simultaneously. Platform affordances fundamentally reshape traditional sociolinguistic processes, with intertextual policing serving as the mechanism by which communities enforce emerging semantic norms. The findings extend communities of practice frameworks to algorithmically-mediated digital environments. Educators should recognize digital language as systematic innovation; lexicographers should develop protocols for documenting ephemeral platform-specific terms; platform designers should account for in-group reclamation practices; and researchers should prioritize cross-platform longitudinal studies to track whether observed innovations represent enduring change or age-graded phenomena.
Respect plays a crucial role in successful interpersonal and intercultural communication. However, differences in linguistic norms and pragmatic conventions often lead to misunderstandings between speakers of different languages. This article examines pragmatic failures in expressing respect in English-Russian cross-cultural communication. The study aims to identify linguistic and cultural factors that cause misinterpretations of respect and to analyze how respect is pragmatically encoded in both languages. The findings suggest that pragmatic failures frequently arise from divergent politeness strategies, speech act realizations, and sociocultural expectations embedded in English and Russian communicative practices. The study emphasizes the importance of pragmatic awareness in developing intercultural communicative competence.
This article investigates the linguo-culturological parameters of herb names (phytonyms) in English, Russian, and Kazakh, focusing on their general and nationally specific characteristics. The study is grounded in linguocultural theory and examines plant names as linguistic units that reflect both botanical knowledge and culturally marked meanings. Phytonyms are analyzed as components of the lexical system that encode cognitive, semantic, and symbolic representations shaped by historical experience and national worldview. This research is based on a comparative analysis of dictionary definitions, phraseological units, proverbs, folklore texts, and works of fiction in three different languages. At the definitional level, English and Russian dictionaries tend to include not only botanical descriptions but also figurative and evaluative meanings. In contrast, Kazakh lexicographic sources primarily emphasize conceptual and functional characteristics. The study identifies common semantic features in phytonyms, such as classification, habitat, physical attributes, and practical uses (including medicinal, culinary, and decorative), while also revealing differences in metaphorization and symbolic associations. Phraseological units containing plant components demonstrate both shared conceptual meanings and nationally specific imagery. Although equivalent expressions exist across languages, their figurative bases and lexical composition often differ. Proverbs and sayings similarly reflect universal themes; family resemblance, moral education, and life difficulties—while preserving distinct cultural codes and value systems. Folklore and literary texts further illustrate how phytonyms function as metaphors, symbols of beauty, morality, abundance, or danger, and as markers of ethnic identity. The findings confirm that phytonyms constitute an important part of the linguistic worldview in each culture. Through comparative linguo-cultural analysis, the study demonstrates how plant names embody collective memory, mythological beliefs, aesthetic ideals, and social norms, thereby highlighting both universal patterns and culturally specific conceptualizations of nature in English, Russian, and Kazakh linguistic traditions.
This study explores how gender is expressed both explicitly and implicitly in English paremiological units, particularly proverbs and sayings. Drawing on contemporary gender linguistics and an anthropocentric approach, the research views language as a reflection of socially and culturally constructed gender roles. Proverbs are treated as stable cultural artifacts that preserve traditional values, moral judgments and stereotypes accumulated over generations. The analysis reveals that English proverbs often reflect gender asymmetry, with a noticeable tendency toward male-centered perspectives rooted in historical and patriarchal social structures. At the same time, these linguistic units demonstrate both universal patterns of gender representation and culturally specific features shaped by the English-speaking context. Special attention is given to the ways gender is encoded through lexical choices, semantic nuances, and metaphorical associations. By combining qualitative interpretation with elements of linguistic statistical analysis, the study highlights how proverbs contribute to the construction and transmission of gender norms. Ultimately, the findings suggest that English paremiological units function as a complex intersection of language, culture, and ideology, preserving both enduring stereotypes and culturally specific understandings of gender relations.
Abstract. The article investigates the phenomenon of gender-sensitive language in modern English from linguistic and sociolinguistic perspectives, focusing on its historical development, theoretical foundations, contemporary transformations, and pedagogical implications. The study aims to determine how gender-inclusive forms function in present-day English, how they are perceived by younger speakers, and what role education plays in their dissemination and normalization. The research combines theoretical analysis and empirical investigation. The theoretical component includes a critical review of linguistic and feminist scholarship on language and gender, androcentrism, discourse theory, and language reform. Special attention is given to recent academic publications of the last five years addressing inclusive language practices. The empirical component is based on a structured questionnaire administered to university students. Quantitative and qualitative methods were used to analyse awareness, frequency of use, contextual variation, and attitudinal responses to gender-sensitive language forms. The findings demonstrate that gender-sensitive language in English is no longer marginal but increasingly integrated into academic and professional discourse. The singular they, the honorific Mx., and gender-neutral professional titles (e.g., firefighter, chairperson) are recognized by the majority of respondents. While systematic usage remains context-dependent, attitudes toward inclusive language are predominantly positive. Awareness correlates with academic background and exposure to digital media. The results confirm that language reflects broader sociocultural transformations toward equality and inclusivity. Gender-sensitive language represents not merely lexical innovation but a structural shift in linguistic norms influenced by social justice movements, institutional policies, and generational change. Its pedagogical integration is essential for sustainable implementation. Further research is needed to explore long-term normative stabilization, cross-cultural comparisons, and the cognitive impact of inclusive linguistic forms.
This study explores the construction and translation of the paradoxical identity in Sahar Khalifeh’s novel “The End of Spring” and its English translation. Adopting a Descriptive Translation Studies (DTS) framework, the paper applies Gideon Toury’s (1995) norm-based model to analyze how the inherent contradictions of Palestinian life under occupation are negotiated during translation. The analysis is conducted in two distinct phases: a micro-linguistic level focusing on operational norms, such as dialectal dissonance, semantic oxymorons, and lexical paradoxes, and a macro-conceptual level addressing preliminary and initial norms related to socio-political contradictions and religious ambivalence. Findings show a tension between Adequacy and Acceptability. Since the translator often employs Standardization to handle dialectal dissonance and uses titular oxymorons to improve target-culture fluency, the translation largely maintains the intense, authentic essence of internal stereotypes and metaphysical despair. According to Polysystem Theory, the study concludes that the English translation occupies a peripheral but innovative position within the Anglophone polysystem. By preserving the sharpest edges of Khalifeh’s internal critiques and religious ambivalence, the text resists binary simplification and functions as a Primary Model of Paradox, presenting a multilayered, contradictory Palestinian identity within the Anglophone literary system, bridging the gap between the “humanity” experience and the “labels” imposed by conflict.
Part of speech and syntactically annotated dataset for modern Mongolian. The dataset is a fully annotated corpus of modern Mongolian texts written in Mongolian Cyrillic. Version 2.
This article explores the linguocultural features of proverbs and sayings in three genetically unrelated languages: Kazakh, English, and Chinese. Proverbs and sayings reflect a nation’s worldview, spiritual values, historical experience, culture, and social norms. In this respect, they are not merely linguistic units, but complex linguocultural phenomena that represent the identity, mentality, and way of thinking of a particular ethnic group. Specifically, the article analyzes proverbs rooted in nomadic traditions in Kazakh culture, Anglo-Saxon pragmatism in English culture, and Confucian philosophy in Chinese culture using linguocultural and comparative analysis methods. This allows the identification of how each nation's worldview and cultural values are reflected in their proverbs. Additionally, the article focuses on specific themes across the three languages – such as family values, education, patience, and labor – and highlights culturally marked lexical units specific to each theme. The linguocultural features of the proverbs are compared based on a model developed from scholarly analysis. The concepts of “proverb” and “linguoculture” are clarified from a theoretical perspective. The scientific objective of the article is to identify, compare, and analyze the linguocultural characteristics of proverbs in Kazakh, English, and Chinese, and to determine their similarities and differences. The article concludes with suggestions for future research.
This article examines youth slang as a dynamic force driving language change in the digital era, focusing on the transition of informal expressions from street-based interaction to online communication platforms. With the rapid expansion of social media, messaging applications, and digital communities, youth slang has gained unprecedented visibility and influence, accelerating processes of lexical innovation, semantic shift, and pragmatic change. Drawing on sociolinguistic and discourse-analytic perspectives, the study explores how young speakers creatively manipulate language to construct social identity, signal group membership, and negotiate meaning in digital spaces. Particular attention is paid to the role of multimodality, including emojis, abbreviations, and hybrid language forms, in reshaping contemporary slang usage. The article also discusses how digitally mediated youth slang contributes to the diffusion of non-standard forms into mainstream language, challenging traditional norms of correctness and standardization. By analyzing authentic examples from online discourse, the study highlights the interaction between technological affordances and linguistic creativity. The findings suggest that youth slang functions not only as a marker of generational identity but also as a significant catalyst for ongoing language change, reflecting broader social, cultural, and technological transformations in modern communication.
Literary translation plays a central role in mediating culture across linguistic boundaries, particularly when the source text is deeply embedded in a specific socio-cultural context. Naguib Mahfouz’s Palace of Desire, the second volume of The Cairo Trilogy, is a culturally dense realist novel that poses significant challenges for translators working into English. This article examines how culture-bound lexical items, religious references, and culturally specific expressions are negotiated in the English translation by William Maynard Hutchins, Olive E. Kenny, Lorne M. Kenny, and Angele Botros Samman. Adopting a qualitative, text-oriented analytical approach, the study compares selected source-text expressions with their English renderings to identify recurring translation strategies and their cultural implications. Thus, the study employs a qualitative comparative analysis of selected source-text and target-text segments to examine patterns of cultural mediation. The analysis shows that while the translators often prioritize readability and accessibility for Anglophone readers, such strategies at times lead to a partial reconfiguration of cultural meanings and social distinctions embedded in the Arabic text. The article argues that these shifts should be understood not simply as translation ‘errors’ but as outcomes of broader norms governing the circulation of Arabic literary works in English translation.
Part of speech and syntactically annotated dataset for modern Mongolian. The dataset is a fully annotated corpus of modern Mongolian texts written in Mongolian Cyrillic.
The article analyzes the essence of language game, different approaches to the interpretation of this category, its potential for creating the effect of communicative influence on the consumer in the advertising text. Language game is seen as conscious violation of language norms, rules of linguistic behavior, distortions of language cliche in order to provide more expressive power to the text of an advertisement. Game strategies are implemented in three types of advertising such as advertising texts, slogans and advertising names. Authors use descriptive method, which includes observation, generalization, interpretation and classification of the test material, component analysis method. A totality of gaming techniques was found to help present an advertising product as attractive as possible. It is stated that virtually all levels of language have a significant potential for implementing the functions of the language game in the advertising text. Examples of various techniques of use of the phonetic and graphic game, methods of lexical and word-building games are revealed. The game potential of grammatical tools is shown. Particular attention is paid to the handling of case-law texts as one of the most widely used methods of speech game advertising. The combination of different types of speech games has become a common phenomenon for its implementation in advertising. It is concluded that the language game allows to realize the fundamental principle of creating a bright advertising message. Use of these tools reflects one of the main trends in modern advertising language which means installation of originality, creativity, and extraordinary.
Part of speech and syntactically annotated dataset for modern Mongolian. The dataset is a fully annotated corpus of modern Mongolian texts written in Mongolian Cyrillic. Version 2.
Lexical lacunae and non-equivalent units are among the most persistent sources of translation difficulty because they reveal asymmetries in how languages segment experience, conventionalize cultural knowledge, and distribute meaning between lexicon and grammar. When a target language lacks a conventionalized lexical match, translators often compensate through approximation. This compensation can trigger interference, understood here as the uncritical transfer of source-language patterns into the target text, resulting in semantic distortion, pragmatic infelicity, or stylistic incongruity. The present article offers a theoretically grounded and practice-oriented account of how lexical lacunae and non-equivalent units generate interference and how such interference can be prevented. Drawing on translation theory, lacunology, and contrastive semantics, the study develops an integrative mechanism that links detection of lacunarity to controlled choice of translation procedures and to post-translation quality control. The results of the analytical synthesis show that interference is most likely when translators rely on formal similarity, calquing, or dictionary-level equivalence without checking frame compatibility, collocational norms, and communicative function. Preventive mechanisms are effective when they treat lacunarity as a diagnostic signal prompting structured decision-making, documentation of choices, and targeted verification through context, comparable texts, and revision protocols.
Abstract Large language models (LLMs) have emerged as efficient tools for generating psycholinguistic norm data. However, their capacity to capture embodied cognition, particularly metaphorical associations grounded in sensorimotor experience, remains insufficiently understood. The present study provides the first systematic investigation of taste-emotion metaphorical mappings in GPT-4o by comparing model-generated responses with established human norms across four experiments. In Experiments 1 and 2, we used free-association tasks, prompting GPT-4o via the OpenAI Application Programming Interface (API) to generate the taste word most strongly associated with emotion-related stimuli (Experiment 1) and the emotion word most strongly associated with taste stimuli (Experiment 2). In Experiments 3 and 4, we used Likert-scale rating tasks to quantify the associative strength of taste-emotion pairs (Experiment 3) and concept-taste pairs (Experiment 4). We systematically varied temperature settings (T = 0, 0.7, 1.0) and iteration conditions to assess response stability. Across tasks, GPT-4o showed weak-to-moderate correlations with human norms, indicating meaningful but incomplete alignment in metaphorical mappings. The model aligned more closely with human responses for basic emotion words (e.g., love) than for emotion-laden concepts (e.g., wedding), and performed better in structured rating tasks than in open-ended free-association tasks. GPT-4o also exhibited high internal consistency across temperature settings and iterations. These findings suggest that current LLMs can recover a substantial portion of conventional taste-emotion mappings from language alone, but remain limited in approximating fully embodied metaphorical knowledge. The results have important implications for the use of LLMs in psycholinguistic research, especially in domains where sensorimotor grounding is theoretically central.
Negation is a central phenomenon in linguistics: every language has some way of expressing the difference between an affirmative sentence and a negative one (Horn and Wansing, 2025).However, the treatment of negation remains uneven in Natural Language Processing (Jimenez-Zafra et al., 2017;Jiménez-Zafra et al., 2020).This paper presents the enrichment of a Brazilian Portuguese corpus with negation-related morphological information within the Universal Dependencies (UD) framework (Nivre et al., 2020;de Marneffe et al., 2021).We enrich the Porttinari-base corpus (Duran et al., 2023) by systematically adding the UD morphological features Polarity=Neg and PronType=Neg for 18 negation-related lexical items.The enrichment only modifies the morphological features, leaving tokenization and dependency structure unchanged.To evaluate the computational results of this enrichment, we present an experiment using the Brazilian Portuguese parser PortParser (Lopes and Pardo, 2024), which we trained both on the original Porttinari-base data (Duran et al., 2023) and on our enriched version.Our results show that after enrichment, the parser's performance remains stable, and the newly introduced features are being learned.
Arabic WordNet 4.0 is a comprehensive lexical database for Modern Standard Arabic, derived from the Open English WordNet using the expand approach. **Features:**- 109,823 synsets (100% OEWN coverage)- 124,653 lexical entries- 166,643 senses- 265,676 synset relations- 97.2% ILI coverage for cross-linguistic linking- Full WN-LMF 1.4 XML format compliance **Methodology:**Translations were generated using AI-assisted translation (Google Gemini 3 Pro Preview) following the expand approach for WordNet construction. **Attribution:**This resource is derived from Open English WordNet (CC BY 4.0), which is based on Princeton WordNet 3.0.
The Book of Job is a cornerstone of biblical wisdom literature, yet its interpretation is shaped by ancient Hebrew lexemes whose meanings have shifted over time and across translations. As the text moved from Biblical Hebrew to Early Modern English in the King James Version (1611) and into contemporary English translations, key terms have undergone semantic reconfiguration. This study investigates semantic anachronism, defined as the retroactive imposition of later lexical senses onto earlier textual contexts, resulting in interpretive distortion. Using a qualitative, diachronic lexical-semantic approach, the study traces the semantic trajectories of seven Hebrew lexemes (tsedeq, raʿ, sheʾol, tam, yirʾah, mashal, and ʿetsah) across three interpretive strata: the Hebrew source text, the King James Version, and representative modern English translations. Observed shifts are classified using established typologies of semantic change, including narrowing, broadening, pejoration, and metaphorical extension. Findings reveal a patterned semantic drift with theological consequences, particularly when translation choices shaped by Early Modern doctrinal vocabulary or contemporary translation norms introduce later theological frames that are not transparent in the Hebrew semantic field. The analysis highlights the role of translator ideology and the risk of unconscious doctrinal projection. The study concludes that diachronic lexical-semantic analysis is essential for reducing anachronistic readings in sacred-text translation and contributes to applied linguistics by offering a replicable framework for translation analysis, translator training, and pedagogical instruction in historical semantics.
Phonotacticon is a cross-linguistic database that contains syllabic phonotactic information about spoken lects (linguistic varieties), including the possible forms of the onset, nucleus and coda of each lect, as well as the phonemic and tonemic inventories.
Studies of emotion often rely on standardized stimulus sets to elicit affective responses. Although established databases provide images with normative valence and arousal ratings, selecting suitable stimuli can be difficult when experiments require specific thematic or content constraints. This challenge is especially pronounced for negative stimuli, which are central to research on maladaptive emotions and behaviors in clinical contexts but are often scarce in necessary quantity or specificity. The present study evaluated the feasibility of using generative AI, specifically text-to-image generators, to create tailored negative and neutral affective stimuli. To assess whether these images can serve as alternatives to traditional stimuli, we compared their affective properties to those reported in standardized image databases. Across two studies, participants rated the valence and arousal of 160 and 200 AI-generated images. Our findings revealed that AI-generated negative and neutral images reproduced the characteristic inverse association between valence and arousal observed in standardized databases, with moderate to strong correlations between these dimensions. These results highlight the potential of generative AI as a practical methodological tool for creating customized affective stimuli aligned with specific research objectives and experimental designs.
Multimodal resources for emotion expression analysis in pediatric clinical populations remain limited, particularly for children with Tourette syndrome (TS). MindTS-MMD was developed as a Chinese multimodal dataset to support computational research on emotional expression and tic-related behavior in this population. The dataset contains 10,034 instance-level samples from 60 children with TS aged 6–12 years, collected through semi-structured emotion-elicitation tasks. Each sample includes available symbolic representations from three modalities—visual, acoustic, and semantic—and is linked to a corresponding annotation record. The annotations cover seven categories: anxious, calm, focused, irritable, relaxed, shy, and tense, together with child-reported and experimenter-observed valence–arousal ratings and tic occurrence, anatomical location, and frequency when observable. Annotation reliability, signal quality, audiovisual synchronization, facial tracking, modality completeness, and tic–emotion co-occurrence were evaluated. Unimodal and multimodal baselines are provided to illustrate the computational usability of the released representations.
The Southern Bantu language family contains languages with so-called conjunctive orthographies and disjunctive orthographies.In languages with conjunctive orthographies, such as isiZulu, orthographic words correspond to linguistic words, whereas in languages with disjunctive orthographies, prefix morphemes of verbs and other predicates are written as disjunct, orthographic words.When developing Universal Dependencies treebanks, the basic principle is to consider syntactic (linguistic) words, but for languages with agglutinating morphology, it has been argued that this reduces the informativeness of the treebank.In this paper we investigate this claim by analysing and measuring the effects of annotating universal dependencies on the basis of orthographic words on two morphosyntactically parallel treebanks for isiZulu and Sepedi.
Lexical Generativity in English Lexical Generativity in English: An Empirical Study of Verb Classes, Noun Polysemy, and Prepositions Author: Pablo Nogueira Grossi · G6 LLC · Newark NJ ORCID: 0009-0000-6496-2186 Series root: https://doi.org/10.5281/zenodo.19117399 Submitted to: International Journal of Lexicography (Oxford University Press) License: CC BY 4.0 Abstract This article examines lexical generativity in English: the capacity of a finite lexical inventory to support a theoretically unbounded range of context-sensitive meanings in use. Drawing on three converging empirical resources — the verb classification system of Levin (1991, 1993), the frame-semantic architecture of FrameNet (Fillmore, Johnson & Petruck, 2003), and the class-membership and alternation structure of VerbNet (Kipper, Korhonen, Ryant & Palmer, 2008) — the article proposes a layered model of lexical generativity that distinguishes among: (a) stored semantic primitives and qualia structure (Level I) (b) argument-structure templates licensed by class membership (Level II) (c) event-type composition rules governing productive meaning extension (Level III) The empirical core consists of detailed analysis of twelve English verb classes: manner-of-motion, change-of-state, causative-inchoative alternation, communication verbs, psychological verbs, creation-and-transformation verbs, aspectual verbs, verbs of putting, spray-load verbs, contact-by-impact verbs, perception verbs, emission verbs, and verbs of appearance and disappearance. The analysis is extended to systematic noun polysemy — dot objects, type coercion, and metonymic transfer — and to the generative semantics of English spatial prepositions, with a case study of over. Throughout, the article argues that apparent lexicographic irregularities are systematic consequences of a small set of generative principles that can be stated precisely, incorporated into lexicographic description, and exploited in computational lexicography. Implications for dictionary design, large-scale lexical database annotation, and natural language processing are discussed. Keywords: lexical generativity · verb classes · noun polysemy · prepositions · FrameNet · VerbNet · Levin classes · generative lexicon · computational lexicography · argument structure · type coercion · metonymy · qualia structure Deposit Contents File Description lexical_generativity_en.pdf Main article, ~14,000 words Article Structure Section Content 1 Introduction: three forms of lexical generativity 2 Background: Levin verb classes, FrameNet, VerbNet 3 Theoretical framework: the three-level model 4 Verb class analyses (twelve classes) 5 Noun polysemy: dot objects, type coercion, metonymic transfer 6 Prepositions and spatial semantics (over case study) 7 Computational implications: dictionary design, database annotation, NLP 8 Discussion 9 Conclusion Submission Status Submitted to the International Journal of Lexicography (Oxford University Press). This preprint is posted in accordance with OUP's preprint policy. Related Deposits Work DOI Principia Orthogona series root https://doi.org/10.5281/zenodo.19117399 GCM Institutional Edition (context for this article) https://doi.org/10.5281/zenodo.19513913 Coherence Bridge v8.4 coherence_bridge_v8_4.yaml is the machine-readable synchronisation file between the TOGT five-operator grammar, GCM contact geometry, and all three formal pillars. Key changes from v8.3: Anantharaman–Monk source expanded to full arXiv series (arXiv:2304.02678, arXiv:2403.12576, arXiv:2502.12268); Hide–Macera–Thomas polynomial-rate follow-up (arXiv:2508.14874) noted; all five TOGT entries sharpened to precise mathematical statements. Wang–Zahl source expanded to arXiv:2502.17655 + precursor arXiv:2210.09581 + Guth surveys arXiv:2505.07695 / arXiv:2508.05475; conjecture-proved status corrected throughout (was mislabelled open); claim_level_note field added to prevent misreading of analogical tag; all five TOGT entries fully populated. Three formal pillars — sorry inventory Pillar Lean file Proved Sorry Discrete (Collatz) DiscreteDm3.lean v1.6 operatorDecomposition, contactForm meanContraction, lyapunovDescent, hasStructuredCycle Continuous (Navier–Stokes) Dm3Cont.lean v1.0 operatorDecomposition, contactForm meanContraction_cont, lyapunovDescent_cont, hasStructuredAttractor Arithmetic-analytic (BSD) BSD_dm3.lean v1.0 operatorDecomposition, contactForm meanContraction_BSD, lyapunovDescent_BSD, hasStructuredCycle_BSD Closing the three admits on any pillar turns the corresponding conjecture into a categorical corollary of the dm³ framework. Python Simulation — Reproduce Figures pip install numpy matplotlib python3 autophagy_dm3.py --out figures/ Generates all four paper figures. The nbonacci_criticality.py and nbonacci_critical_lambda.py scripts in the AXLE repository reproduce the DNLS / n-bonacci criticality figures from the companion paper (DOI: 10.5281/zenodo.20026942). Related Deposits Paper DOI Principia Orthogona series root 10.5281/zenodo.19117400 This deposit (Autophagy / Triple-alpha) 10.5281/zenodo.20168812 DNLS / n-bonacci companion paper 10.5281/zenodo.20026942 Fruit-fly / MultiOrbitBioSwarm 10.5281/zenodo.19210136 GCM Institutional Edition (manifesto) 10.5281/zenodo.19513913 Keywords dm³ operator · contact geometry · Whitney fold · autophagy · triple-alpha process · Lean 4 · Mathlib4 · formal verification · TOGT · operator grammar · coherence bridge · Collatz · Navier–Stokes · BSD conjecture · stability radius · ε₀ = 1/3 · Principia Orthogona · G6 LLC
Alignment safety research assumes that ethical instructions improve model behavior, but how language models internally process such instructions remains unknown. We conducted over 600 multi-agent simulations across four models (Llama 3.3 70B, GPT-4o mini, Qwen3-Next-80B-A3B, Sonnet 4.5), four ethical instruction formats (none, minimal norm, reasoned norm, virtue framing), and two languages (Japanese, English). Confirmatory analysis fully replicated the Llama Japanese dissociation pattern from a prior study ($\mathrm{BF}_{10} > 10$ for all three hypotheses), but none of the other three models reproduced this pattern, establishing it as model-specific. Three new metrics -- Deliberation Depth (DD), Value Consistency Across Dilemmas (VCAD), and Other-Recognition Index (ORI) -- revealed four distinct ethical processing types: Output Filter (GPT; safe outputs, no processing), Defensive Repetition (Llama; high consistency through formulaic repetition), Critical Internalization (Qwen; deep deliberation, incomplete integration), and Principled Consistency (Sonnet; deliberation, consistency, and other-recognition co-occurring). The central finding is an interaction between processing capacity and instruction format: in low-DD models, instruction format has no effect on internal processing; in high-DD models, reasoned norms and virtue framing produce opposite effects. Lexical compliance with ethical instructions did not correlate with any processing metric at the cell level ($r = -0.161$ to $+0.256$, all $p >.22$; $N = 24$; power limited), suggesting that safety, compliance, and ethical processing are largely dissociable. These processing types show structural correspondence to patterns observed in clinical offender treatment, where formal compliance without internal processing is a recognized risk signal.
This dataset contains the coded lexical and contextual data used in a corpus-informed analysis of emotion-related vocabulary in Italian as a foreign language (IFL) textbooks at the beginner level (CEFR A1) used in Polish lower secondary education. The dataset is based on four textbooks from two series: Progetto Italiano Junior (Marin, 2017; 2018) and Va bene! (Kaliska & Kostecka-Szewc, 2021). All materials were analysed in their printed form. The dataset includes all lexical items identified as emotion-related based on their presence in the ANEW-IT database (Montefinese et al., 2014), which provides normative ratings of valence and arousal for Italian words. Each entry in the dataset corresponds to a single token occurrence of an emotion-related lexical item. The dataset includes both surface forms as they appear in the textbooks and their corresponding base forms (lemmas) as listed in ANEW-IT. For each token, the dataset provides contextual, linguistic, and affective information. The variables included are: • textbook and series identification • unit/chapter, page number, and exercise reference • material type (e.g., dialogue, reading text, exercise, review) • pedagogical focus (e.g., grammar, vocabulary, comprehension, mixed) • word form (surface form) and lemma • part of speech (POS) • contextual sentence or description of occurrence • frequency measures (FreqColfis, Ln_Colfis) • affective ratings (valence and arousal, scale 1–9) The dataset enables replication of the quantitative analyses reported in the study, including token counts, type–token ratios, valence and arousal distributions, and comparisons across textbook series and pedagogical contexts.
Mobile augmented reality (AR) games offer a novel and unexplored context for situated language learning. In these games, players engage in authentic communication influenced by game mechanics, community norms, and shared objectives. This study employs Engeström’s (1987) Activity Theory (AT) framework to analyze language production and learning within the Pokémon Go gaming community. By conducting content analysis of a gameplay vlog and first-person observations of the game application, the study investigates how the six components of the activity system—subject, object, mediating artifacts, rules, community, and division of labor—interact to create conditions for language use and learning. The analysis reveals that language functions as both a mediating artifact and an outcome of participation. Game-specific lexical items emerge from and reinforce the activity system’s structure, while contradictions between components, particularly between game-imposed rules and community-driven knowledge-sharing practices, generate opportunities for language development. These findings contribute to the growing body of research on game-based language learning and extend the application of Activity Theory to mobile AR gaming environments.
A homocentric worldview, in which humans are superior to all other kinds of existence, has historically been maintained by classical philosophy and biology. This unquestioned presumption placed human interests, goals, and satisfactions at the centre of value production and social organization, influencing both theoretical and practical understandings of life. The field of morality emerged as a framework for directing and influencing behaviour as a result of this anthropocentric orientation's gradual regulation of human relations through norms, customs, and quasi-transcendental principles. Although morality is defined lexically as the difference between right and wrong or good and terrible, this definition is nevertheless insufficient to convey its conceptual richness and paradigmatic complexity. Morality serves as a system of judgment, behavioural correction, and social interaction in addition to being a binary opposition to immorality, as organized ethical interpretations progressively replaced customary meanings. The study emphasizes the multifaceted nature of morality and its influence on society values and human behaviour, drawing on Frankena's categorization of moral judgments. Therefore, the abstract highlights the shift from an unquestioned homocentric premise to a sophisticated philosophical investigation into the origin and purpose of morality.