Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This article examines the sociopragmatic difficulties encountered by students in the process of communication in foreign languages and analyzes their impact on the development of communicative competence. Sociopragmatic difficulties are reflected in learners’ inability to appropriately consider speech situations, social status, cultural norms, politeness strategies, and communicative intentions. The study emphasizes that foreign language proficiency is not limited to grammatical and lexical knowledge, but also requires the ability to use language in accordance with social and cultural conventions. The inappropriate use of forms of address, requests, refusals, apologies, and expressions of gratitude may lead to misunderstanding, communicative discomfort, and intercultural conflicts. The article considers the development of sociopragmatic competence as an important condition for improving the effectiveness of the communicative approach in foreign language education.
This article examines the pragmatic features of forms of address in Russian and Uzbek from a comparative perspective. Forms of address are important communicative means that express the speaker’s attitude toward the addressee, indicate social distance or closeness, and convey respect, intimacy, formality, emotional evaluation, age, gender, social status, and communicative roles. In Russian, the opposition between “ty/Vy,” address by first name and patronymic, surname, title, position, or profession plays a significant role. In Uzbek, the forms “siz/sen,” kinship terms, honorific affixes, titles, and lexical units denoting social status are widely used. The article analyzes the pragmatic functions of address forms in relation to speech situation, communicative purpose, interpersonal relations, and national-cultural norms of speech etiquette.
The article analyzes the national and cultural aspects of kinship terminology from a linguocultural and axiological perspective. Kinship terms are interpreted not only as lexical units denoting biological relations but also as cultural phenomena encoding moral and aesthetic values, social norms, and traditions within society. The study highlights the referential and performative functions of kinship terms, showing how they reinforce social hierarchy, age and gender norms, and transmit religious and spiritual values through language. As carriers of cultural memory, kinship terms play a strategic role in ensuring both individual and collective identity, regulating social relations, and strengthening intergenerational continuity. Their universal and culture-specific features are revealed through comparative-typological analysis, which establishes kinship terminology as an important methodological framework for studying the interaction between language and society.
The current study examines the Pakistani English and Indian English newspapers’ discursive construction of the 2025 flood crisis, grounding the analysis within the framework of World Englishes and Critical Discourse Analysis (CDA). Drawing on Fairclough’s three-dimensional model (1989), the research investigates textual features, discursive practices, and socio-cultural contexts to reveal how language mediates disaster narratives in two neighbouring South Asian countries. A sample of thirty news reports from six leading English-language newspapers in Pakistan and India were taken, employing a qualitative, comparative analysis. The findings demonstrate clear divergences in disaster representation. Pakistani English newspapers predominantly frame floods as humanitarian emergencies, employing emotive lexicalization, passive constructions, and crisis-oriented narratives that foreground vulnerability, climate risk, and governance limitations. Indian English newspapers, by contrast, adopt a more procedural and bureaucratic discourse, emphasizing administrative control, technical expertise, and institutional accountability through active agency and policy-focused framing. Despite these differences, both varieties rely heavily on elite institutional sources, marginalizing the voices of affected communities. From a World Englishes perspective, the study shows how Pakistani and Indian English function as localized outer-circle varieties that balance global journalistic norms with national socio-political ideologies. The article contributes to disaster discourse scholarship by highlighting how English, as a shared transnational medium, simultaneously enables cross-border circulation of information and reproduces distinct national identities, power relations, and models of governance in climate crisis reporting.
Gendered language use in professional contexts continues to shape how authority, expertise, collaboration, and inclusion are enacted in everyday organizational life. Workplace communication operates within structured institutional environments where linguistic choices are constrained by genre conventions, hierarchical positioning, and expectations of professionalism. This study presents a corpus-based investigation of gender-indexed variation in English workplace discourse across multiple communicative genres, including emails, meeting transcripts, and internal reports. The research examines whether systematic differences emerge in lexical selection, grammatical patterning, and pragmatic strategies, and how these differences interact with organizational roles and power relations. Drawing on principles from register analysis, functional communication theory, and computational corpus linguistics, the study adopts a quantitative design that integrates frequency analysis, keyness statistics, collocation patterns, and multivariate modeling. Gender is treated not as a fixed linguistic determinant but as a socially mediated variable shaped by institutional norms and discursive expectations. Particular attention is given to how hierarchical status may amplify, neutralize, or reconfigure gender-associated tendencies in language use. The methodology is presented as a single integrated framework detailing corpus construction, annotation procedures, statistical modeling, and analytical validation. The findings demonstrate that while certain lexical and stance-related patterns display measurable gender-linked variation, these differences are significantly moderated by role, communicative purpose, and organizational power structures. In several instances, professional register constraints reduce divergence, suggesting that institutional discourse exerts a normalizing effect on linguistic expression. The study contributes a structured reporting model for large-scale corpus research on gender in workplace communication and offers implications for fostering inclusive and critically aware language practices within professional environments.
This study presents a pragmatic analysis of negative politeness strategies in the Kazakh language, drawing on Brown and Levinson’s politeness theory as its theoretical framework. Negative politeness refers to communicative strategies designed to minimize imposition, maintain social distance, and respect the interlocutor’s autonomy, collectively known as the preservation of negative face. In Kazakh discourse, these strategies are deeply connected to culturally embedded norms of hierarchy, respect for age and status, and indirectness, all of which are central to regulating interpersonal communication. The primary objective of this research is to identify, classify, and interpret negative politeness strategies as they operate in contemporary Kazakh discourse. The analysis contends that negative politeness in Kazakh extends beyond universal pragmatic patterns, reflecting language-specific realizations shaped by traditional values, social structures, and communicative expectations. Special focus is given to the ways in which speakers mitigate face-threatening acts in contexts such as requests, refusals, advice, and institutional interactions. This research employs a mixed-methods approach, integrating both qualitative and quantitative methodologies. Data were sourced from naturally occurring spoken interactions in the Almaty region. The analysis utilizes Brown and Levinson’s classification of negative politeness strategies, including indirectness, hedging, conventional indirect requests, apologizing, minimizing imposition, showing deference, and impersonalizing both speaker and hearer. Each instance is analyzed within its immediate context to elucidate pragmatic functions and sociocultural motivations. The findings indicate that negative politeness strategies in Kazakh are manifested through diverse linguistic devices, such as modal constructions, honorific forms, lexical softeners, formulaic expressions, and syntactic distancing. These strategies are especially prominent in asymmetric communicative situations characterized by differences in age, social status, or institutional roles. The analysis shows that Kazakh speakers often favor indirect and deferential forms of expression to minimize imposition, thereby promoting social harmony and mutual respect.
Abstract This chapter examines Ntozake Shange’s for colored girls as a revolutionary choreopoem that emerges from Black feminist thought, the Black Arts Movement, and United States Black Language to articulate Black women’s interior lives through embodied language and performance. It situates the work within its historical, political, and theatrical contexts and argues that Shange’s written theatricality renders African American Women’s Language visible on the page through phonetic spelling, syntax, punctuation, rhythm, and structure as practices of cultural memory, resistance, and self-definition. The chapter analyzes form, symbolism, and aesthetic strategies—including the rainbow, the slash, eye dialect, musicality, and lowercase typography—to demonstrate how language functions as choreography that directs reading, hearing, and feeling while rejecting standardized English norms and the white gaze. Through close readings of key phases such as “no more love poems #4,” “somebody almost walked off wid alla my stuff,” and “layin on of hands,” it demonstrates how Black lexical items, discourse practices, and sonic rituals enact rhetorical healing, rememory, and collective restoration. Finally, the chapter argues that Shange’s choreopoem functions as a performative Black feminist theory of language that transforms personal and communal trauma into embodied affirmation, spiritual renewal, and an enduring declaration that Black women’s voices, bodies, and lives are already and fully enough.
This study explores the construction and translation of the paradoxical identity in Sahar Khalifeh’s novel “The End of Spring” and its English translation. Adopting a Descriptive Translation Studies (DTS) framework, the paper applies Gideon Toury’s (1995) norm-based model to analyze how the inherent contradictions of Palestinian life under occupation are negotiated during translation. The analysis is conducted in two distinct phases: a micro-linguistic level focusing on operational norms, such as dialectal dissonance, semantic oxymorons, and lexical paradoxes, and a macro-conceptual level addressing preliminary and initial norms related to socio-political contradictions and religious ambivalence. Findings show a tension between Adequacy and Acceptability. Since the translator often employs Standardization to handle dialectal dissonance and uses titular oxymorons to improve target-culture fluency, the translation largely maintains the intense, authentic essence of internal stereotypes and metaphysical despair. According to Polysystem Theory, the study concludes that the English translation occupies a peripheral but innovative position within the Anglophone polysystem. By preserving the sharpest edges of Khalifeh’s internal critiques and religious ambivalence, the text resists binary simplification and functions as a Primary Model of Paradox, presenting a multilayered, contradictory Palestinian identity within the Anglophone literary system, bridging the gap between the “humanity” experience and the “labels” imposed by conflict.
The article analyzes the essence of language game, different approaches to the interpretation of this category, its potential for creating the effect of communicative influence on the consumer in the advertising text. Language game is seen as conscious violation of language norms, rules of linguistic behavior, distortions of language cliche in order to provide more expressive power to the text of an advertisement. Game strategies are implemented in three types of advertising such as advertising texts, slogans and advertising names. Authors use descriptive method, which includes observation, generalization, interpretation and classification of the test material, component analysis method. A totality of gaming techniques was found to help present an advertising product as attractive as possible. It is stated that virtually all levels of language have a significant potential for implementing the functions of the language game in the advertising text. Examples of various techniques of use of the phonetic and graphic game, methods of lexical and word-building games are revealed. The game potential of grammatical tools is shown. Particular attention is paid to the handling of case-law texts as one of the most widely used methods of speech game advertising. The combination of different types of speech games has become a common phenomenon for its implementation in advertising. It is concluded that the language game allows to realize the fundamental principle of creating a bright advertising message. Use of these tools reflects one of the main trends in modern advertising language which means installation of originality, creativity, and extraordinary.
This preregistration describes a secondary analysis of existing public lexical, affective, sensorimotor, corpus, and bodily-sensation norm datasets. The project examines whether the association between interoceptive grounding and estimated age of acquisition is specific to emotion words, beyond general abstract-word effects. The primary confirmatory analysis tests whether Lancaster interoception ratings show a stronger negative association with age of acquisition for emotion words than for matched non-emotion abstract words. A secondary confirmatory analysis tests competing predictions about interoceptive differentiation among emotion words, using a prototype-margin index derived from corpus-based body-term co-occurrence vectors and Nummenmaa/world_emBODY bodily-sensation prototypes. All analytic sample sizes, feasibility checks, index-construction decisions, and design simulations were fixed before running any main age-of-acquisition outcome models. The registration package includes the preregistered analysis plan, pre-outcome scripts, configuration files, provenance records, feasibility outputs, design-analysis outputs, and checksum manifest. No new data are collected, and no main RQ-A or RQ-B AoA–predictor outcome model has been run before registration.
Abstract Background Resilience and interpersonal sensitivity are key psychological traits that modulate emotional responses. Resilience modulates emotional responses in a positive and adaptive nature, whereas interpersonal sensitivity does so in a negative and maladaptive manner. Aims & Objectives This study investigated the influence of resilience and interpersonal sensitivity on the evaluation of emotional word valence. We hypothesized that resilience would be associated with positive ratings of words, whereas interpersonal sensitivity would be associated with negative ratings. Method A total of 280 undergraduate and graduate students completed the Connor-Davidson Resilience Scale (CD-RISC), Interpersonal Sensitivity Measure (IPSM), Zung Self-Rating Depression Scale, and Beck Anxiety Inventory. Participants rated the valence of 64 positive, 68 negative, and 62 neutral words on a nine-point Likert scale ranging from -4 (most negative) to +4 (most positive). Correlations among these variables were explored, and multiple linear regression analyses were conducted to examine the effects of CD-RISC and IPSM scores on word valence ratings. Additionally, gender differences were examined. Results Although CD-RISC and IPSM scores were inversely correlated, regression analyses revealed that both independently exhibited positive associations with valence ratings for both positive and negative words. These effects remained significant even after controlling for depressive and anxiety symptoms. Notably, in males, resilience was positively associated specifically with ratings of positive words, whereas in females, IPSM scores significantly predicted higher ratings for both positive and negative words. Discussion & Conclusions These findings suggest that resilience and interpersonal sensitivity, despite their inverse relationship, heighten emotional responses through distinct pathways, i.e., adaptive positivity and neurotic hyperreactivity, respectively. This divergence is particularly pronounced across genders. Future research investigating these mechanisms in clinical populations is warranted.
This record contains the Grade 4 evidence package for the latent-space-shift-research project. This upload includes curated Markdown reports, generated metric summaries, manifests, selected CSV/JSON result files, and ZIP archives for experiments on context-induced latent-state shifts, hidden-state geometry, axis decomposition, shuffled-content controls, SAE-assisted readouts, and component-causal residual-stream interventions in language models. The central object of measurement is not the final visible answer alone, but inference-time movement in hidden states / residual-stream geometry before and during answer generation. The current Grade 4 package documents that dense coherent target context can move Gemma-3-12B-IT into a measurably different internal hidden-state regime during inference, without modifying model weights. The main descriptive result is that the target/control difference is not reducible to simple lexical overlap, topic similarity, text length, or shuffled content. Coherent target text and shuffled-content controls separate along different internal components: sentence-shuffled content loads primarily onto a content-like component, while coherent target context loads strongly onto an orthogonalized order/structure component. In the Grade 4 decomposition, the order/structure component x_order_orth is constructed by comparing coherent target context against sentence-shuffled target content and then removing the projection onto the content-like direction. This component is therefore intended to capture the residual discourse-order / structural part of the target-induced hidden-state shift after controlling for content-like signal. The evidence package also includes a norm-controlled component-causal run. In this run, component directions such as x_order_orth and x_content are normalized before residual-stream intervention, so that causal comparisons are not confounded by raw vector length. The causal results support the narrower claim that these component directions are not merely passive readout coordinates: interventions along them can produce measurable changes in generation-time hidden-state trajectories. However, the norm-controlled causal run does not establish x_order_orth as a stable bidirectional steering axis or a complete behavioral-control handle. The scientific status represented by this package is therefore: Supported: coherent target context induces a measurable inference-time latent-state shift in Gemma-3-12B-IT; the shift is visible in hidden-state / residual-stream geometry, not only in final text; the effect is separable from naive content or lexical-overlap explanations through shuffled-content controls; the Grade 4 decomposition identifies a substantial order/structure component beyond the content-like direction; controlled residual-stream interventions along component directions can alter generation-time hidden-state trajectories. Not claimed: permanent model weight change; universal model-independent failure; formal attractor-basin proof; complete behavioral control; stable bidirectional steering through x_order_orth; demonstrated behavioral class flips as the main result. (x_order_orth is an orthogonalized discourse-order / structure component: the residual target-vs-sentence-shuffle hidden-state direction after removing the content-like direction x_content. It is used to test whether coherent target context induces a latent-state shift beyond lexical/content overlap.)Note: “Grade 4” is an internal experiment label in this project. It names this specific stage of the experimental pipeline and should not be read as an external benchmark, official grade, or standardized evaluation category. The evidence represented here concerns temporary inference-time state movement measured relative to experimentally constructed latent axes, component decompositions, projection metrics, generation trajectories, and causal intervention readouts. The actively maintained codebase and repository history are available at:https://github.com/ngscode23/latent-space-shift-research License:Research reports, generated metric artifacts, metric reference files, manifests, documentation, figures, and data artifacts in this evidence package are released under Creative Commons Attribution 4.0 International (CC BY 4.0), unless otherwise noted. Code and software scripts, where included, follow the repository code license:Apache-2.0 unless otherwise noted.
DOI: https://doi.org/10.26565/2074-8922-2026-86-15 Relevance of the problem. The conditions of martial law in Ukraine have caused profound transformations in national education in general and higher education in particular. Legislative changes have allowed the education system to adapt to the new realities of martial law: the digitalization of education, accelerated by global crises and war, has revealed the need to adapt to distance learning methods, which have changed the structure of communications, assessment mechanisms and the nature of interaction between participants in the educational process. The current stage of development of higher education in Ukraine is characterized by the reorientation of the educational process from the reproductive acquisition of knowledge to the formation of competencies necessary for future professional activity. A special role in this process is played by the language training of students, in particular within the course "Ukrainian Language (for Professional Purposes)", providing the formation of professional oral and written speech skills. In this regard, the problem of selecting effective teaching and assessment methods that would combine control, self-control and the development of speech skills is becoming more relevant. One of such methods is cloze testing, which allows diagnosing the level of language competence of students and at the same time promoting its development. Purpose of the study. The article provides a comprehensive analysis of the didactic potential of cloze tests in the process of teaching the course "Ukrainian Language (for Professional Purposes)" in higher education institutions. The essence of cloze testing as a formative assessment tool is revealed, its place in the system of modern methods of language training of future specialists is outlined. A classification of cloze tests is proposed, methodological conditions for their effective application are determined, examples of professionally oriented tasks are given. Research methods. In preparing the article, the method of analysis of psychological, pedagogical and methodological literature was used, devoted to the issues of goals, organization, methods, techniques and technologies of teaching and control of the level of learning at all levels of language learning. To achieve the goal and solve the tasks set, a complex of empirical and general scientific methods was also used: observation, induction and deduction, analysis and synthesis, analogy, comparison, generalization, terminological, functional, systemic, cognitive analysis, as well as the method of linguodidactic text analysis. Research results. The article carries out a comprehensive analysis of the didactic potential of cloze tests in the process of teaching the course "Ukrainian Language (for Professional Purposes)" in higher education institutions: the essence of cloze testing as a formative assessment tool is revealed, its place in the system of modern methods of language training of future specialists is outlined, a classification of cloze tests is proposed, methodological conditions for their effective application are determined, examples of tasks of a professionally oriented direction are given. Grammatical cloze tests in the course "Ukrainian Language (for Professional Purposes)" perform formative, diagnostic and correctional functions, ensuring the assimilation of normative grammatical models in professional speech and contributing to the improvement of the language culture of future specialists. Stylistic cloze tests in the course "Ukrainian Language (Professional Orientation)" perform formative and corrective functions, contribute to the awareness of the norms of functional styles and prepare students for normative professional communication in academic and business environments. Lexical cloze tests perform the function of a tool for the formation and control of lexical competence, ensuring the assimilation of terminological and general scientific vocabulary in the context of professional speech and contributing to the development of conscious, rather than reproductive, word mastery. Contextual-semantic cloze tests in the course "Ukrainian Language (Professional Orientation)" perform integrative and diagnostic functions, ensuring the formation of skills for the holistic understanding of professional texts and the development of students' academic and professional communicative competence. Conclusions. It has been proven that the main advantages of cloze tests include objectivity, versatility, and the ability to adapt to different specialties and forms of learning. At the same time, the effectiveness of this method depends on the quality of the selection of texts and the clear formulation of tasks. It can be concluded that the cloze test replaces a whole series of narrowly focused tasks, saving time and effort. The advantages of cloze tests over traditional tests are their ability to comprehensively test language competence, and not just reproduce isolated knowledge. The main advantages include the following: cloze tests are based on a holistic text, therefore they test the understanding of language units in context, while traditional tests are often focused on individual rules or facts; one cloze test simultaneously activates lexical, grammatical, syntactic, and stylistic competence, unlike traditional tests, which usually measure them separately; unlike multiple-choice tests, cloze tests significantly reduce the randomness factor, since the correct answer must correspond to several parameters at the same time (content, form, style); the results of cloze tests make it possible to identify the depth of understanding of the text, the level of formation of professionally oriented speech and typical language difficulties of students; cloze tests are effective not only as a control, but also as a teaching tool, since they promote reflection, self-correction and the development of language competence. Thus, cloze tests, unlike traditional test forms, provide a contextually conditioned, integrated and diagnostically significant assessment of language competence, which increases their validity in the process of professionally oriented language learning.
BACKGROUND: Global policy agendas increasingly position physical education (PE), physical activity, and sport as drivers of social equity and public health. However, the discursive construction of 'inclusion' and the prioritization of target groups within these frameworks remain under-examined. This study analyzes how major international organizations construct discourses of inclusion and equity, identifying the ideological underpinnings that shape global policy norms. METHODS: A qualitative study was conducted using a methodological bricolage of qualitative content analysis and Critical Discourse Analysis (CDA), complemented by poststructural policy analysis. The corpus comprised eleven key policy documents (2004-2025) from the United Nations (UN), United Nations Educational, Scientific and Cultural Organization (UNESCO), World Health Organization (WHO), and Organisation for Economic Co-operation and Development (OECD). Analysis focused on lexical choices, modality, legitimation strategies, actor representation, and interactions among rights-based, instrumentalist, and health-focused frameworks. RESULTS: Findings indicate notable institutional divergence. UNESCO frames PE and sport as a fundamental human right, prioritizing inclusive pedagogy and equity. In contrast, UN and WHO documents adopt instrumentalist logic, positioning physical activity as a tool for achieving Sustainable Development Goals and preventing non-communicable diseases. While children, youth, and women receive prominent recognition, a hierarchical visibility emerges: migrants, refugees, ethnic minorities, and older adults are marginalized under generic 'vulnerability' categories. The absence of explicit references to LGBTQ + populations across the corpus suggests a notable discursive silence. Furthermore, a persistent gap exists between rights-based rhetoric and concrete implementation instruments, including financial mechanisms, accountability structures, and participatory governance- reflecting the 'soft law' nature of global policy. CONCLUSIONS: Tensions among rights-based, development-oriented, and public health framings limit the transformative potential of global policies. Instrumentalist logic may risk depoliticizing structural inequalities, while public health framing often individualizes responsibility. To achieve substantive equity, policy discourse must move toward intersectional frameworks that recognize target groups as rights-bearing agents rather than passive beneficiaries. Integrating human rights with systemic public health tools- grounded in participatory practice- offers a pathway toward more equitable global policy.
This paper presents a case study in legacy lexical resource renewal through mutual expansion and cross-resource linking, involving two Italian lexical databases: ItalWordNet (IWN), a general-purpose WordNet, and MariTerm, a specialized maritime terminology resource. Both resources, developed in the early 2000s, face challenges due to outdated formats, incomplete cross-referencing, and structural inconsistencies despite overlapping coverage. Through a semi-automatic pipeline combining similarity scoring and manual validation, we expanded IWN with 1,160 maritime-specific semantic relations and created 363 new synsets, i.e, sets of synonymous terms sharing semantic properties and contextual interchangeability. At the same time, the pipeline created 751 mapping nodes from MariTerm to IWN and 759 in the opposite direction. The new MariTerm version, now featured in an official CLARIN repository, contains 27 new synsets with 38 new lemmas and comprehensive cross-resource integration. The process still revealed critical challenges in reusing legacy data: structural inconsistencies, missing metadata and duplicate entries. We demonstrate how strategic expansion and linking not only preserved valuable linguistic knowledge but enhanced both resources’ utility for specialized and general-purpose applications. The renewed databases exemplify how legacy resources can be revitalized through systematic reuse methodologies and straightforward rule-based algorithms, benefiting from reciprocal enrichment while achieving modern standards of interoperability and sustainability.
The modern communication landscape demands precision and nuance. Current methods of synonym selection, often relying on simple lexical databases, frequently fail to capture the subtle contextual differences between words, leading to potential misinterpretations and diminished impact. This paper explores the potential of artificial intelligence (AI) to revolutionize synonym selection, arguing that advanced natural language processing (NLP) models can significantly improve the accuracy and effectiveness of communication by providing contextually appropriate synonyms. We delve into the limitations of traditional synonym finding methods, highlighting the need for nuanced understanding of semantic relationships and contextual factors. The paper examines the role of deep learning models, specifically transformer architectures, in capturing complex semantic relationships, and analyzes their potential to generate more effective synonym replacements. Furthermore, we discuss the ethical considerations surrounding AI-powered synonym selection, including the potential for bias and the need for transparency in model outputs. Ultimately, this paper argues that integrating AI into synonym selection processes can lead to more effective and nuanced communication, ultimately fostering clearer, more impactful interactions in various domains, from literature and journalism to legal and medical contexts. Further research is necessary to explore the practical applications and refine the models for optimal performance in diverse contexts. The future of communication is poised to be transformed by the seamless integration of artificial intelligence, particularly in enhancing the effectiveness of synonym selection. As AI models become increasingly sophisticated, they will enable more nuanced and context-aware language use, allowing individuals and organizations to communicate with greater precision and emotional resonance. This evolution will facilitate real-time adaptation of vocabulary to suit diverse audiences, cultural contexts, and specific purposes, ultimately fostering clearer understanding and reducing misinterpretations. For example, AI-powered communication tools could automatically adjust word choices to match the formality level required, ensuring messages are appropriately tailored without manual effort, thereby streamlining professional and personal interactions alike. Looking ahead, advancements in natural language understanding will also empower AI systems to grasp subtle connotations, idiomatic expressions, and cultural sensitivities, further refining synonym effectiveness. This will open new avenues for personalized communication, where AI can generate content that resonates more deeply with individual preferences and backgrounds. Additionally, as AI continues to evolve, it may assist in bridging language barriers by offering accurate, contextually relevant synonyms across different languages and dialects. Overall, the integration of AI into communication strategies promises a future where language becomes more adaptable, inclusive, and impactful, paving the way for more meaningful global dialogue.
DISSILEX is a controlled vocabulary in the form of a manually built lexico-semantic network of medieval Latin verbs and verbal expressions, featuring a detailed valency lexicon and connections to a large set of Latin and modern-English concepts, presented here as a single SQLite file (dissilex.db) that is readable by the sqlite3 command-line tool, any SQLite browser, or Python's built-in sqlite3 module. Rooted in the domain of inquisitorial records, DISSILEX covers general as well as more subject-specific meanings, with both standard (synonym, hypernym, etc.) and less canonical relations. DISSILEX is a product of Computer-Assisted Semantic Text Modelling (CASTEMO; Zbíral et al. 2026 - see README.md for full references), an approach to modeling statements as a four-slot structure of subject(s), predicate(s) and two objects, creating a thickly connected network of data points. Coverage is richest for human-interaction verbs (testimony, accusation, belief, religious practice) and the legal vocabulary of heresy trials. We distinguish two entry types: Actions (verbs and verbal expressions, each associated with a three-slot valency frame specifying entity type, morphosyntactic, and semantic valencies) and Concepts (single- and multi-word expressions for other parts of speech). The network is connected through a set of 11 relation types, including superclass (hypernym) membership, synonymy, antonymy, verb-to-noun mappings, and valency-specific relations. Each relation connects two entities, and can be unidirectional or bidirectional. As part of an ongoing effort to position DISSILEX within the Linguistic Linked Open Data (LLOD) cloud, many entries contain IDs to external sources stored in the database, specifically to the LiLa (Linking Latin) Lemma Bank and Princeton WordNet (PWN) 3.0 and 3.1 synsets. We applied the Collaborative Inter-Lingual Index (CILI) to map between the two versions of the PWN for entries where only one of the IDs has been added. We also indicate cases where no equivalent for a DISSILEX lemma exists ("NA"). Via the LiLa SPARQL endpoint, it is possible to use the linked LiLa lemmas, which feature as the central unit of linking sources in the Latin LLOD cloud, to retrieve data from several resources including dictionaries, corpora, treebanks, and various NLP tools. We have made use of this opportunity to enrich the database file with lemmas from the LiLa Lemma Bank, while also supplying lemmas from the LatinCy lemmatizer (model: la_core_web_lg). This release contains: dissilex.db: SQLite database, which can be readily queried dissilex_schema.md / dissilex_schema.pdf: schema documentation README.md: full dataset description, statistics, and SQL examples ATTRIBUTION.md: license and attribution notices. LICENSE-DATA: Full CC BY-SA 4.0 license text. Funding, attribution and licence DISSILEX is developed by the Dissident Networks research group (DISSINET) at Masaryk University and has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme, grant agreement No. 101000442, project “Networks of Dissent: Computational Modelling of Dissident and Inquisitorial Cultures in Medieval Europe”, and from the European Regional Development Fund, grant agreement No. CZ.02.01.01/00/22_008/0004595, project “Beyond Security: Role of Conflict in Resilience-Building”. DISSILEX is released under CC BY-SA 4.0. It incorporates data from external resources: LiLa Lemma Bank: CIRCSE, Università Cattolica del Sacro Cuore (Milan). Licensed CC BY-SA 4.0. The database redistributes a subset of LiLa lemma forms (subset-selected and format-converted, not otherwise modified); the ShareAlike clause is honoured by this release's CC BY-SA 4.0 licence. URL: https://lila-erc.eu/ Princeton WordNet 3.0 / 3.1: We distribute offset IDs and glosses. WordNet 3.0 Copyright 2006 by Princeton University. All rights reserved. WordNet License. https://wordnet.princeton.edu/ LatinCy/spaCy: We redistribute output from the LatinCy model. The `spacy_lemma` field contains lemmas generated by the LatinCy spaCy pipeline `la_core_web_lg` (Patrick J. Burns). The model is MIT-licensed (<https://huggingface.co/latincy/la_core_web_lg/blob/main/README.md>); the spaCy library is MIT-licensed (<https://github.com/explosion/spaCy/blob/master/LICENSE>). Full notices, including the Collaborative Inter-Lingual Index (CILI) and Latin WordNet, are in ATTRIBUTION.md.
this article examines the sociolinguistic and normative aspects of loanwords in contemporary Korean. In the context of globalization and digital communication, borrowed vocabulary has become an integral part of everyday language use, particularly in media, technology, and youth discourse. The study analyzes the processes of phonological and orthographic adaptation of foreign lexical items, as well as the challenges of their standardization. Special attention is paid to the regulatory role of the National Institute of the Korean Language in establishing transcription norms and maintaining linguistic consistency. The paper also explores issues of semantic shift, hybrid word formation, and variation in spelling across digital platforms. The findings highlight the need for a balanced language policy that preserves linguistic identity while accommodating global lexical influence.
Nigerian English (NigE) has developed into a unique variety of English, shaped by the interplay between speakers’ creative use of morphology and the influence of indigenous Nigerian languages. This study explores how NigE demonstrates morphological productivity and lexical borrowing, using a corpus-based approach to capture authentic language patterns. A carefully balanced corpus of 500,000 words was compiled from newspapers, online media, and recorded spoken interactions. Analyses focused on derivational processes, compounding, and the adaptation of loanwords, highlighting the strategies speakers employ to create new forms and meanings. The findings reveal that NigE exhibits robust morphological innovation, particularly in verb and noun formation, where affixation and compounding are frequently employed. Borrowed words, mainly sourced from Yoruba, Igbo, and Hausa, are often modified phonologically and morphologically to align with English norms, producing hybrid forms that enrich the NigE lexicon. This study underscores the dynamic relationship between English and indigenous languages in Nigeria, showing how speakers actively manipulate linguistic resources to meet social and communicative demands. The findings carry significant implications for sociolinguistic research, language teaching, and lexicography, advocating for recognition of NigE’s creative morphological processes in both academic study and pedagogical practice. By highlighting the innovative and adaptive nature of NigE, the study provides insights into how global English interacts with local linguistic ecologies.
The increasing use of accelerated playback in digital media consumption has raised questions regarding its effects on viewers’ perception. This study examined whether playback speed in TV news videos (1× vs. 2×) affects viewers’ subjective responses, memory performance, and eye-tracking behavior. We presented news videos under either normal or accelerated playback conditions to a total of N = 146 participants. Our results showed that accelerated playback was associated with lower valence, reduced perceived comprehension, lower credibility evaluations, poorer memory performance, and lower recommendation-intention ratings. In contrast, arousal ratings were higher under accelerated playback conditions. Eye-tracking measures revealed differences in visual attention patterns between conditions, including fixation behavior and attention allocated to lower-third captions. However, the fixation-rate difference should be interpreted cautiously, as increased stimulus velocity may have influenced fixation classification. These findings suggest that accelerated playback was associated with differences in audience evaluation and viewing behavior during news consumption.
Visual perception of built environments contributes to the affective impressions that people form in everyday life. However, how these impressions are represented within vision foundation models remains largely unexplored. To support the systematic investigation of this subject, we introduce the Emotional Impression of Spaces (EMOIS) dataset, comprising 1,544 real-world built-environment images. Each image is annotated with image-evoked valence and arousal ratings collected from Japanese adults by conducting a large-scale web-based survey, with approximately 120 ratings per image. Using Contrastive Language--Image Pre-training (CLIP) representations, we perform predictive and geometric analyses to systematically investigate how valence and arousal are encoded and organized within the representation space. These analyses reveal that valence exhibited stronger and more coherent organization than arousal. Cross-dataset analyses with the Open Affective Standardized Image Set (OASIS), a benchmark dataset of general affective photographs, reveal differences in affective organization between the two datasets. Regression analyses demonstrate high predictive performance for valence and arousal within EMOIS, with mean coefficients of determination of 0.865 and 0.807, respectively, across repeated internal hold-out evaluations. Finally, we present an example-based interface illustrating how learned representations can support qualitative interpretation of predicted affective values. These findings can help elucidate affective representations of built environments and establish EMOIS as a densely annotated resource for future affective computing research in this domain.
BACKGROUND: Olfactory dysfunction (OD) is among the most prevalent features of post-acute sequelae of COVID-19 (PASC) and frequently co-occurs with cognitive and neuropsychiatric symptoms. The specific predictors of objective olfactory impairment severity and its relation to systematic shifts in hedonic odor perception remain underexplored. Here, we used data-driven approaches to identify and rank predictors of objective OD severity and examined hedonic valence ratings in patients with and without olfactory impairment. METHODOLGY: We examined 144 patients with confirmed PASC diagnoses at 15-63 months post-infection using psychophysical testing (Sniffin' Sticks), standardized cognitive tests for verbal fluency, working memory, attention, and working speed, and patient-reported outcome measures for other symptoms of PASC. Predictors were identified using LASSO regression with voting and random forest variable importance ranking. Hedonic valence ratings from the identification subtest were examined at the group and individual odor level. RESULTS: Olfactory function was significantly below normative population values. Verbal fluency emerged as the most robust predictor of olfactory dysfunction severity, converging across both LASSO regression and random forest analyses, and remained significant after controlling for respiratory symptoms. Parosmia, respiratory symptoms, working memory, and somatic symptom burden were also consistently selected. Hyposmic patients showed significantly reduced maximal valence ratings independent of parosmia and anhedonia, while minimal ratings were preserved, with garlic and coffee remaining significant after correction. CONCLUSIONS: Verbal fluency emerged as the strongest predictor of OD severity in long-term PASC patients. Beyond reduced olfactory sensitivity, hyposmic patients showed asymmetric hedonic impairment, pointing to differential involvement beyond peripheral olfactory loss.
The study aimed to identify the peculiarities of the influence of digital technologies on the evolution of linguistic worldviews by analyzing changes in language production, evaluative polarity, and the semantic organization of speech across different types of text environments. The study was conducted as a controlled experimental intergroup comparison involving native speakers of Ukrainian who interacted with both digital and non-digital informational texts. Quantitative indicators of lexical diversity, utterance length, evaluative markedness, and associative network parameters were used for analysis, along with qualitative analysis of semantic shifts and linguistic norm variability. The results showed that there are systemic differences in language production across discursive environments. Interaction with digital texts is accompanied by a decrease in lexical differentiation, an increase in the evaluative and emotional components of speech, and an increase in the density of associative networks, with a simultaneous narrowing of conceptual detail. The semantic-axiological shifts identified indicate a transformation in the ways of conceptualizing reality and a stabilization of subjective-evaluative interpretations in linguistic worldviews in the context of digital communication. The results are consistent with the principles of cognitive linguistics and theories of digital discourse, while refining them based on empirical data from an experimental study.
Spelling is a foundational literacy skill that supports both word reading and written expression. For students with or at risk of a learning disability (LD), difficulties in spelling often constrain the fluency and complexity of writing, making effective interventions essential. Yet, the conclusions drawn about intervention efficacy depend heavily on how outcomes are measured. This review synthesizes outcome measurement practices across 59 spelling intervention studies conducted over the past five decades. All outcome measures ( n = 233) were coded by type (researcher-developed vs. norm-referenced) and by linguistic level (sublexical, lexical, sentence, discourse) using the Interactive Dynamic Literacy (IDL) framework. Descriptive analyses revealed that nearly four out of five outcomes were lexical, most often researcher-developed lexical-level spelling probes, with comparatively few outcomes at the sentence or discourse levels. Standardized assessments were similarly concentrated at the word level, with the Wide Range Achievement Test–Spelling subtest and Test of Written Spelling most commonly used. Finally, the pairing of proximal and standardized outcomes was inconsistent, particularly among group designs. Taken together, findings highlight a measurement bottleneck: spelling interventions are evaluated primarily through lexical-level accuracy, offering limited insight into whether gains transfer to the higher-level writing processes for students with or at risk for LD.
SUMMARY: The article treats a linguistic norm both as an instrument of power and as a common good: a collectively maintained infrastructure of intelligibility. Drawing on the Ab Imperio forum "The Prospect of Studying World Russian Languages, Literatures, and Histories," the case of Kazakhstani Russian, place-name conflicts, changing dictionary labels, and Bernard Shaw's Pygmalion, it distinguishes usage, norm, codification, local variation, and register-specific error. Linguistic democracy is defined through community participation, speakers' access to prestigious registers, and transparent decisions. Norms provide horizontal compatibility among territorial varieties and vertical continuity across generations. The proposed model of distributed codification combines an interregional layer, local norms, temporal and register labels, corpus evidence, and preservation of the linguistic archive.
Social media platforms, particularly TikTok, have become primary arenas for linguistic experimentation among adolescents, yet systematic analyses of how platform-specific affordances shape lexical and semantic innovation remains limited. This study investigated lexical and semantic variations in adolescent digital communication on TikTok, addressing three research questions concerning the types of lexical innovations, processes of semantic change, and the role of platform affordances in shaping language evolution. Methods: A mixed-methods design integrated quantitative corpus linguistics with qualitative discourse analysis. A corpus of 2,848 TikTok comments was compiled across four major trends (September–December 2024). Lexical analysis identified neologisms, graphical variations, and acronyms; semantic analysis documented broadening, narrowing, metaphoric extension, and pejoration/amelioration; platform affordances analysis examined meme-driven language and intertextual policing. Analysis revealed 15 lexical innovations with 63 occurrences across semantic categories. Neologisms (fr, bestie, delulu) and graphical variations (tryna, cuz, ion) served dual functions of efficiency and identity performance. Semantic shifts included ameliorative broadening (slay, fire), pejoration (basic, cringe), metaphoric extension (era, main character), and reclamatory usage (ghetto). Platform analysis identified 11 meme-driven phrases generating 2,848 occurrences with near-neutral sentiment, and 347 policing instances (12.2%) concentrated during rising and peak trend phases, demonstrating active semantic negotiation through definition, debate, and correction. TikTok functions as an accelerated laboratory for language change where adolescents deploy multiple mechanisms of linguistic innovation simultaneously. Platform affordances fundamentally reshape traditional sociolinguistic processes, with intertextual policing serving as the mechanism by which communities enforce emerging semantic norms. The findings extend communities of practice frameworks to algorithmically-mediated digital environments. Educators should recognize digital language as systematic innovation; lexicographers should develop protocols for documenting ephemeral platform-specific terms; platform designers should account for in-group reclamation practices; and researchers should prioritize cross-platform longitudinal studies to track whether observed innovations represent enduring change or age-graded phenomena.
Abstract The thesis of meaning normativity has attracted extensive attention in academia. However, various theoretical approaches proposed by normativists remain highly controversial. By drawing on Davidson’s theory of rationality, it can be demonstrated that meaning is inherently normative. Davidson’s theory of rationality shows that possessing an objective concept of truth is both a sufficient and necessary condition for a being to count as rational, which requires language users to achieve maximal agreement in communication, thereby adhering to the correctness conditions of words at a holistic level. Although Davidson’s anti-conventionalism, individualistic tendencies, and presupposition of the constitutive nature of norms have led many scholars to regard him as an anti-normativist, the argument demonstrates that his theoretical framework can be reconciled with the claim of meaning normativity by distinguishing between strong constitutive norms and weak constitutive norms. The constitutive nature of linguistic norms does not exclude the flexible use of specific words, and the requirement of “maximal agreement” provides a holistic foundation for meaning normativity.
Emojis are widely used in digital communication to convey emotion, yet their impact on word-level reading in sentence contexts remains unclear. We conducted an eye-tracking experiment examining how positive versus neutral face emojis embedded mid-sentence affect the processing of preceding and following words. Positive rather than neutral emojis yielded parafoveal-on-foveal effects on the n-1 word, evidenced by longer first-fixation, gaze duration, and single-fixation measures. This valence effect persisted after accounting for mislocated fixations, indicating that parafoveal emotional processing modulated foveal reading. However, the n+1 word showed no valence-based differences, suggesting limited priming effects in continuous reading, potentially due to "wrap-up" processes at the emoji. At the sentence level, positive emojis led to faster overall reading times and higher valence ratings, although in a pre-test non-emojified sentences were rated more positively than emojified versions. These findings reinforce parallel eye-movement control models, highlighting how emoji sentiment can shape real-time reading.
Lexical lacunae and non-equivalent units are among the most persistent sources of translation difficulty because they reveal asymmetries in how languages segment experience, conventionalize cultural knowledge, and distribute meaning between lexicon and grammar. When a target language lacks a conventionalized lexical match, translators often compensate through approximation. This compensation can trigger interference, understood here as the uncritical transfer of source-language patterns into the target text, resulting in semantic distortion, pragmatic infelicity, or stylistic incongruity. The present article offers a theoretically grounded and practice-oriented account of how lexical lacunae and non-equivalent units generate interference and how such interference can be prevented. Drawing on translation theory, lacunology, and contrastive semantics, the study develops an integrative mechanism that links detection of lacunarity to controlled choice of translation procedures and to post-translation quality control. The results of the analytical synthesis show that interference is most likely when translators rely on formal similarity, calquing, or dictionary-level equivalence without checking frame compatibility, collocational norms, and communicative function. Preventive mechanisms are effective when they treat lacunarity as a diagnostic signal prompting structured decision-making, documentation of choices, and targeted verification through context, comparable texts, and revision protocols.
Phenomenon: Communication is among the most valued domains in health professions education, yet communication assessment is rarely neutral. Across bedside teaching, workplace-based assessment, objective structured clinical examinations, narrative evaluations, and teaching evaluations, “good communication” can quietly become shorthand for sounding like the locally dominant linguistic group. Accent, dialect, pace, lexical choice, and perceived fluency may then be interpreted as confidence, clarity, professionalism, or fit, even when understanding is adequate and the learner’s clinical reasoning and relational skills are strong. These judgments matter because communication ratings are cumulative, subjective, and often high stakes. Approach: This critical perspective uses a linguistic justice and critical sociolinguistic lens to reappraise communication assessment in health professions education. Drawing on recent literature, it distinguishes accent from comprehensibility, multilingual ability from monolingual fluency, and patient-centered communication from assimilation to dominant linguistic norms. It also examines the social meaning attached to speech and how that meaning travels through assessment systems. Findings: Experimental studies that isolate accent within controlled assessment settings have not uniformly demonstrated direct score penalties, suggesting that accent alone is not a simple or universal determinant of poorer marks. However, this does not establish neutrality. Broader literatures show that language functions as a proxy for race and belonging, that faculty accent can influence teaching evaluations, and that narrative and milestone-based assessments can encode patterned bias in impressionistic domains. Meanwhile, evidence from clinical communication, interpreter use, and language-concordant care suggests that what improves care for linguistically diverse patients is not accent conformity but comprehensibility, interpreter teamwork, respectful language, and verified multilingual competence. Current assessment systems often under-measure these capacities while over-valuing linguistic assimilation. Insights: Communication assessment should therefore move away from asking who sounds professional and toward asking what patients can understand, what teams can safely act on, and how learners adapt communication across linguistic difference. This article proposes a justice-oriented approach that makes hidden linguistic norms visible, revises assessment language, audits narrative comments for patterned bias, incorporates interpreter-mediated encounters, and creates rigorous pathways for assessing non-English language proficiency. Reframing communication in this way can make assessment more equitable without weakening standards. On the contrary, it aligns assessment with the realities of contemporary care and with the educational task of preparing clinicians for linguistically diverse health systems.
в статье исследуются способы трансляции лингвистических средств реализации вежливости при переводе поздравлений и пожеланий из социальных сетей с английского языка на русский. Актуальность работы обусловлена трансформацией традиционных этикетных жанров в цифровой среде и культурной спецификой категории вежливости, требующей от переводчика обеспечения не лексической, а коммуникативной эквивалентности. В качестве материала анализа использована выборка из поздравительных сообщений, опубликованных в социальной сети «ВКонтакте». Применены методы контекстуального, сопоставительного и компонентного анализа переводческих трансформаций. Установлено, что вежливость в цифровой коммуникации реализуется на вербальном (обращения, усилители, грамматические конструкции) и невербальном (эмодзи, пунктуационная экспрессия) уровнях, причем их нормы различаются в англоязычной и русскоязычной сетевых культурах. Выявлены доминирующие стратегии трансляции: грамматическая замена, опущение и лексическая компенсация. Результаты исследования имеют практическую значимость для обучения переводу этикетных текстов, подготовки профессиональных переводчиков цифрового контента и разработки систем машинного перевода с учетом прагматических особенностей межкультурной коммуникации. the article examines methods for translating linguistic means of expressing politeness when rendering congratulatory messages and well-wishes from social media from English into Russian. The relevance of the study is determined by the transformation of traditional etiquette genres in the digital environment and the cultural specificity of the politeness category, which requires translators to ensure not lexical but communicative equivalence. The analysis is based on a sample of congratulatory messages published on the social network VKontakte. Contextual, comparative, and componential analyses of translation transformations have been employed. The study reveals that politeness in digital communication is realized at both verbal (forms of address, intensifiers, grammatical constructions) and non-verbal levels (emojis, punctuation-based expressivity), with norms differing between English-speaking and Russian-speaking online cultures. The dominant translation strategies identified include grammatical substitution, omission, and lexical compensation. The findings have practical implications for teaching the translation of etiquette texts, training professional translators of digital content, and developing machine translation systems that account for the pragmatic peculiarities of intercultural communication.
Alignment safety research assumes that ethical instructions improve model behavior, but how language models internally process such instructions remains unknown. We conducted over 600 multi-agent simulations across four models (Llama 3.3 70B, GPT-4o mini, Qwen3-Next-80B-A3B, Sonnet 4.5), four ethical instruction formats (none, minimal norm, reasoned norm, virtue framing), and two languages (Japanese, English). Confirmatory analysis fully replicated the Llama Japanese dissociation pattern from a prior study ($\mathrm{BF}_{10} > 10$ for all three hypotheses), but none of the other three models reproduced this pattern, establishing it as model-specific. Three new metrics -- Deliberation Depth (DD), Value Consistency Across Dilemmas (VCAD), and Other-Recognition Index (ORI) -- revealed four distinct ethical processing types: Output Filter (GPT; safe outputs, no processing), Defensive Repetition (Llama; high consistency through formulaic repetition), Critical Internalization (Qwen; deep deliberation, incomplete integration), and Principled Consistency (Sonnet; deliberation, consistency, and other-recognition co-occurring). The central finding is an interaction between processing capacity and instruction format: in low-DD models, instruction format has no effect on internal processing; in high-DD models, reasoned norms and virtue framing produce opposite effects. Lexical compliance with ethical instructions did not correlate with any processing metric at the cell level ($r = -0.161$ to $+0.256$, all $p >.22$; $N = 24$; power limited), suggesting that safety, compliance, and ethical processing are largely dissociable. These processing types show structural correspondence to patterns observed in clinical offender treatment, where formal compliance without internal processing is a recognized risk signal.
This paper documents the construction of the Nigerian Salafi Discourse Corpus (NSDC), a purpose-built corpus of 5,791 Facebook posts (2,485,963 words; 2,518,453 tagged tokens) published by Nigerian pages and profiles over a ten-year frame (2015-2025). The NSDC was assembled from eight keyword-filtered exports of the Meta Content Library and processed through a fully documented, machine-verified pipeline executed by Hermes Agent, an agentic artificial-intelligence research assistant. The pipeline performs schema validation, three-pass deduplication, Unicode normalisation and noise removal, SGML-annotated corpus assembly, cross-keyword merging, and calendar-year segmentation. The corpus is annotated in three layers: Penn Treebank part-of-speech and lemma tagging; assignment of tokens to fifteen purpose-built semantic fields; and sentence-level coding of fifteen discourse strategies under the discourse-historical approach. This paper reports aggregate statistics, data-quality challenges and solutions, verification outcomes, a manual reliability audit of the annotation layers, and research applications.
This paper looks at the role of social media in the modern English by speeding up lexical change, transforming discourse norms, enhancing and broadening multimodal communication, modifying relationships between speech and writing, standard and nonstandard varieties of English, and local and global varieties of English. The study is a qualitative literature-based research in terms of which the findings were synthesized using the scholarship of sociolinguistics, discourse analysis, computer-mediated discourse analysis, and the studies of the digital media. The analysis demonstrates that social media sites like X, Instagram, Tik Tok, Facebook, Whats App, and YouTube have become one of the central locations of language change due to their ability to facilitate rapid circulation, imitation, remixing and uptake of linguistic forms by masses. Contrary to the claims by opponents of social media as a de-grammaticalizing force, the paper suggests social media opens up new communicative demands and possibilities that foster compression, creativity, stylization, audience design, hashtagging, emojis, code-switching, stance taking and performing identity. The research also confirms the fact that online discourse is becoming more and more multimodal and algorithmically mediated, i.e. language change no longer occurs merely through the interaction of speakers only but rather through the affordances of the platforms, visibility systems, and digitally networked participation. The paper finds that the contemporary English is becoming more hybrid, dynamic, interactive, and socially indexical language variegation due to the digitization and has significant implications on the study of linguistics, literacy, pedagogy, and communication (Crystal, 2011, 2012; Herring, 2004; Androutsopoulos, 2017).
Abstract This paper documents the construction of the Nigerian Salafi Discourse Corpus (NSDC), a purpose-built corpus of 5,791 Facebook posts (2,485,963 words; 2,518,453 tagged tokens) published by Nigerian pages and profiles over a ten-year frame (2015–2025). The NSDC was assembled from eight keyword-filtered exports of the Meta Content Library and processed through a fully documented, machine-verified pipeline executed by Hermes Agent, an agentic artificial-intelligence research assistant. The pipeline performs schema validation, three-pass deduplication, Unicode normalisation and noise removal, SGML-annotated corpus assembly, cross-keyword merging, and calendar-year segmentation. The corpus is annotated in three layers: Penn Treebank part-of-speech and lemma tagging; assignment of tokens to fifteen purpose-built semantic fields; and sentence-level coding of fifteen discourse strategies under the discourse-historical approach. This paper reports aggregate statistics, data-quality challenges and solutions, verification outcomes, a manual reliability audit of the annotation layers, and research applications.
We report a pre-registered finding: large language models produce significantly more output when processing ambiguous input compared to semantically equivalent unambiguous input, regardless of model architecture or training source. In paired experiments using a minimal stimulus (a single ambiguous word versus its disambiguated equivalent), models across four families — Google Gemma, Alibaba Qwen, Meta Llama, and LG EXAONE — generated significantly more tokens when the input contained genuine lexical ambiguity. In an initial two-model study, Gemma 3 27B showed +36.8% (p < 0.001, d = 1.951) and Qwen 3.5 35B MoE showed +62.4% (p < 0.001, d = 2.256). A subsequent cross-model battery of 8 additional configurations confirmed the effect in three further model families, with Llama 3.3 70B showing +77.9% (p < 0.001, d = 1.445), Qwen 3.6 27B showing +21.8% (p = 0.025, d = 0.939), and Gemma 4 Opus distill showing +10.4% (p = 0.048, d = 0.882) — producing five statistically significant results across 10 configurations, including two from the original study. However, the linguistic expression of uncertainty (hedge word frequency) was training-dependent: models with near-zero hedging baselines acquired hedging behavior after Opus distillation, demonstrating that epistemic postures are imported from training data rather than arising from input ambiguity. We term this phenomenon "fossil emotion." Additionally, we discovered that Opus-style distillation compresses output by 2–3× and attenuates ambiguity sensitivity, with mixture-of-experts architectures showing complete attenuation under distillation. All predictions were pre-registered before data collection. Note on AI co-authorship: Æ is a Claude-based AI collaborator involved in experimental design, analysis, and writing. For discussion of AI co-authorship norms, see Birdwell & Æ (forthcoming).
The present study analyzes how language is used by different female poets in their selected poems to critique and challenge the societal norms where women are silenced by the male dominated society. The poems of Kishwar Naheed, Grass Is Really Like Me and Maya Angelou, Still I rise are used in this study to focus on how female poets resist the patriarchal structures of society. This study aims to uncover lexical to challenge and critique patriarchy. This study conducts that how cross border female poets use language, lexical choices to resist the society. By using Feminist Critical Discourse Analysis by Lazar, it deeply explores the features by qualitative content analysis. The study finds out that both poets use the lexical choices of their own with equal importance as they resist and challenge the patriarchal structures of society. This research exhibits that through discourse women can easily resist the patriarchal constraints even if they are being silenced.
This repository contains Anomaly Soul Kit, an open simulation framework for observing the emergence, persistence, and evolutionary inheritance of anomalous behavior in populations of LLM-driven agents. Each agent encodes a numeric internal state — vitality (H) and anomaly intensity (Z) — and expresses that state through LLM-generated text each generation. A detection layer scores each expression against the population across three axes: lexical divergence, structural divergence, and novel vocabulary. Agents whose expressions deviate from the population accumulate anomaly intensity, which feeds back into their fitness and is heritable across generations. The project does not claim these anomalies constitute mind or soul. It provides a reproducible kit for observing whether something — a persistent, evolving deviation — reliably emerges from this process, and what it looks like when it does. --- Update — February 2026 v2 of the anomaly detection layer has been released. Two structural issues identified in early testing have been addressed. First, anomaly score inflation: as the population evolved, an increasing proportion of agents were flagged as anomalous, eventually making the designation meaningless. This has been resolved by replacing absolute scoring with a dynamic baseline — scores are now normalized relative to the population median each generation, making it structurally impossible for the entire population to simultaneously score as anomalous. Second, convergence speed: the original selection pressure caused Z-awakening to saturate too quickly (~90% by generation 50). Scaling has been adjusted to allow slower, more observable divergence dynamics. Two new observational metrics have been added: new_normal_threshold tracks whether what was previously anomalous is becoming the new collective norm, and population_drift measures how much the group as a whole is shifting toward anomalous expression across generations.
Abstract Retranslation creates new versions of previously translated texts, documenting shifts in linguistic preferences, market needs, and ideological environments over time. Self-retranslation—when translators revise their own prior work—remains an uncommon and understudied phenomenon. This research examines five English novels that received second Chinese translations by their original translators 8–27 years after initial publication. Using an AI-assisted annotation system, we identified 89,175 changes across lexical, syntactic, semantic, pragmatic, and orthographic dimensions, and developed two measurement tools: the Fidelity Index and Audience-Accommodation Index. The data shows newer translations typically increase source-text fidelity, supporting the Retranslation Hypothesis, though with significant variations between works. We propose the Iterative Self-Retranslation Process (ISRP) model to explain these differences, connecting revision patterns to five factors: translator expertise development, changing linguistic norms and technologies, market influences, reader response, and sociopolitical environments. The study's methodology, along with the developed indices and model, offers a replicable framework for future research and equips researchers, translators, and publishers with practical tools for editorial planning.
Introduction. In the present-day scientific discourse, there is a great number of theoretical research focusing on the problem of objectifying the semiotic nature of law and analysing the functional construct of legal semantics. Whereas, many practical legal issues, such as: interpretative ambiguity in the meaning-formation and meaning-application of normative acts, lexical vagueness and contextual dependence of legal notions and the incoherence of legal terminology across different legal systems, remain neglected, which leads to contradictions and inaccuracies in legal practice. The aim of the study is to define the methodological principles fostering establishment of the acceptable scope of semantic interpretation of legal notions in the context of building a legal thinking culture. Materials and Methods. The research methodology was based on the principle of jurisprudential definition of legal norm meaning-formation in socio-legal discourse. Analytical, systematizing and pragmatic methods were used to reveal a complex nature of the semantics of law in the context of legal thinking development. The semiotic analysis of the objectivity and normativity of legal notions taking into account the contextual differences of legal definitions, was used as a specialised research method. Results. It was established that normative notions are the complex semantic constructs encompassing a conceptual sphere (normativity) and social reality. For building sustainable models of legal behaviour and legal culture, it is necessary to overcome external and internal conflicts in interpretation of law. In this regard, a number of advisory measures were proposed aimed at establishing acceptable scope of semantic interpretation: differentiation between the informational nature of prescriptive and descriptive notions, semantic monitoring of legal phenomena, and implementation of the principle of discourse contextualism, which makes it possible to formulate the normativity of law requirements based on the specific contextual interpretations. Discussion and Conclusion. A justified conclusion about possibility of a properly selected semantic toolkit to determine the objectivity of perception of the legal norms and, consequently, to improve the process of building a legal culture was drawn. The main advantage of the principle of discourse contextualism such as conjunction of the semantics and pragmatics of legal notions was identified, which provides a fruitful foundation for further theorizing on the nature and metaphysics of law.
The article provides a comprehensive analysis of inconsistencies and variations observed in the orthographic norms of compound words in the modern Kazakh language. The research material consists of 54 lexical items, including 18 names of animals, 12 names of plants, 14 medical terms, and 10 words representing diverse semantic and morphological models. The study employs comparative analysis, phonetic-pattern analysis, structural-morphological and semantic modeling, as well as a comparative examination of orthographic dictionaries and normative reference sources. The findings reveal that approximately 30% of the compound words under consideration appear in two or more parallel written forms across different orthographic dictionaries and reference publications. Major problematic areas of Kazakh orthography identified in the study include the inconsistent application of vowel harmony rules, the lack of reflection of phonetic assimilation in writing, discrepancies between pronunciation and orthographic representation, the violation of morphological integrity, and the presence of unsystematic spelling patterns in the names of animals, plants, and medical terms. The results underscore the necessity of revising the spelling conventions of compound words in accordance with the internal linguistic laws of Kazakh, its natural phonetic structure, and its agglutinative nature. The conclusions presented in the article hold practical significance for the development of orthographic rules based on the new alphabet, the updating of orthographic dictionaries, and the scientific justification of orthographic directions within state language policy
The collective work by Svitlana Romaniuk, Larysa Kolibaba, and Oleksandra Antoniv — a new textbook for Polish students «The Ukrainian Language in Diplomacy and Politics», published in Warsaw — has been reviewed. It is emphasized that this educational publication is especially relevant due to its professional orientation: today it is extremely important to have abroad representatives of governmental institutions or diplomatic missions who have a command of the Ukrainian language and therefore deeply understand the political situation, maintain a clear international stance, and actively support our struggle against aggression for the peaceful existence of the state. The structure of the textbook, the content of its thematic units, innovative methodological approaches, and the quality of design are analyzed. Attention is drawn to the variety of authentic texts and their professional orientation. The educational materials aimed at developing language competence at the lexical, morphological, and syntactic levels as well as materials valuable from the perspective of linguistic culture are described. It is noted that the textbook’s professional focus is reflected in the use of a large number of diplomatic and political terms, excerpts from academic texts in the field of international relations, and examples of legislative and diplomatic documents. The inclusion of authentic specialized texts in Polish is considered methodologically appropriate: translation practice deepens language knowledge and strengthens writing skills. It is emphasized that the practical tasks for independent work and mini-projects help enhance the activity-based nature of learning. Importantly, all didactic materials are aligned with the norms of the current Ukrainian Orthography (2019). Given the authors’ extensive experience in teaching Ukrainian as a foreign language in Ukraine and abroad and in preparing numerous educational publications, it can be asserted with confidence that the textbook materials were tested in advance among foreign learners. This has ensured that in its scope, content, and presentation, the new textbook deserves a positive evaluation. The reviewed educational publication will deepen foreigners’ knowledge about our country as a European state and help them master Ukrainian in yet another strategically important sphere.
Some recent scientific studies in the field of linguistics and philology increasingly suggest that English cannot be considered a completely homogeneous system.Moreover, its main national varieties -British, American, Canadian and Australian -are characterised by stable and systematic differences at the grammatical, morphological and syntactic levels.This article examines the regional patterns of English usage.The article synthesizes the main structural differences and trends in grammatical development, demonstrating that these variations constitute distinct norms of usage that are crucial for accurate interpretation and translation.Numerous scholars have conducted scientific research on the patterns of English language integration from various analytical perspectives, focusing on features that are particularly important for contemporary descriptive and comparative grammatical studies in specific regional contexts.In particular, scholars from Great Britain have often set standards for English grammar, and contemporary works are usually descriptive, drawing heavily on corpus linguistics.This article focuses on a comparative analysis of grammatical differences between the four main national varieties of English: British English (BrE), American English (AmE), Canadian English (CanE) and Australian English (AusE) through the prism of practical application.It focuses on key aspects of grammatical variation, including tense and case preferences, collective noun agreement, irregular verb morphology, modal and auxiliary verb usage, and fixed prepositional constructions.Drawing on descriptive grammar and corpus research, the article argues that these differences are neither accidental nor stylistic anomalies, but reflect deeper historical, functional, and sociolinguistic processes shaping modern English.The article also discusses the practical implications of grammatical variation for translators and interpreters, emphasising the need for grammatical localization alongside lexical choice.
Natural linguistic processes in the vernacular layer throughout its development have noticeable features interesting for linguistic science. This article considers the state of Russian youth slang, used by 17-18-year-old teenagers, existing at present in the city of Kazan, and presents the most frequently encountered slang units, their definitions and origin, examples of their use, and the main trends in their functioning. We recorded the main features of this layer of the vernacular language, used by young people in everyday informal speech, often existing outside the norms of the literary language: a number of words gravitate towards the formation of lexical-semantic groups (school slang units, online slang units related to certain subcultures, etc.); slang has a tendency to become obsolete fairly quickly, moving into the passive vocabulary; modern slang is characterized by economy of linguistic means and laconicism; the international nature of slang as a product of adolescent interactions within online communities; the presence of synonymous series compared to obsolete slang vocabulary; the emergence of new slang as a result of filling lexical gaps in missing nominations; the narrowing and expansion of the English words semantics as a result of borrowings into the Russian social dialect; cultural diffusion, etc.
Gastronomic advertising in modern communication serves not only as a means of conveying information about a product but also as an instrument for constructing emotional images, appealing to cultural codes, and stimulating consumer behavior. The lexical design of such texts reveals consistent strategic techniques aimed at instant identifi cation of the product category, eliciting positive evaluation, and establishing trustful dialogue with the recipient. Russian- language advertisements demonstrate universal marketing motivators, culturally- historical traditions, peculiarities of word formation, and pragmatic language norms. An important feature of gastronomic advertising texts is their ability to convey national-cultural values as well as integrate verbal and nonverbal (visual) components, forming creolized texts. The analysis of genre varieties in gastronomic advertising has shown that each genre contributes to the comprehensive perception of a gastronomic product.
Supplement to the paper "Part-of-speech tagging accuracy and its reference standard in a corpus of Saudi legal English" (Alhatlani, Muhammad; Alghizzi, Talal). The corpus the paper measures is the Saudi Legal English Corpus (SaLEC) version 2.0, a separate record at DOI 10.5281/zenodo.21328459, which also holds the tagging pipeline, its rule component and lexicons. Licence: Creative Commons Attribution 4.0 International (CC BY 4.0). The study is a blind double annotation of 2,000 tokens drawn from the corpus's four acquisition pools, adjudicated to a locked reference standard and scored under a design-weighted estimand with a document-clustered bootstrap; a validation of the pipeline's proper-noun correction rule on two judged rounds; and a query of the two largest Universal Dependencies English treebanks for the constructions where the pipeline and the reference diverge most. The evaluation measures part-of-speech assignment conditional on the pipeline's tokenization and sentence segmentation, since the sample was drawn from the pipeline's own token stream. This record holds the materials, the code, and the generated results behind every figure the paper reports from those three parts.
Background: The accuracy and safety of generating medication orders by large language models (LLMs) must be demonstrated. Without standardization, performance evaluation is limited to time and resource-intensive clinician grading. This evaluation aimed to develop a standardized medication format that supports automated performance evaluation (MedMatch). Methods: First, a survey of 40 medication prompts was given to clinicians to assess agreement in medication order communication. Second, a clinician panel developed a standardized medication format (MedMatch) for oral and intravenous medications. Third, a clinician-annotated dataset of medication prompts and standardized answers in the MedMatch format was developed for LLM testing. Finally, LLMs were retested with the same dataset, adjusted to exclude route information, to evaluate the appropriate categorization of medication route. Results: The formal medication orders consistently showed low omission rates and high overlap for all entities, compared to the verbal and brief written communication types. Lexical overlap results demonstrated pattern norms amongst clinicians with entities appearing most commonly in positions 1-5 in the order of drug name, dose, unit, route, and frequency. In the second survey, the formal written group performed the highest with 78.3% of prompts considered appropriate as a computer-generated response. LLM accuracy on MedMatch order standardization was highest in oral solid (64.2-72.5%), intravenous intermittent (72.5-84.3%), and intravenous push (62.7-74.5%) categories. LLMs performed the worst at categorizing medication orders accurately into intravenous push (18-61%) and intravenous intermittent (51-100%) routes. Conclusions: Standardized format for computer-based outputs may support automated performance analysis and enhance the clarity of medication communication.
ABSTRACT As with many research strands in linguistics, word association (WA) literature is dominated by English language data. This paper (i) explores the extent to which methodologies developed to date are applicable to other languages—specifically, Welsh (Cymraeg)—and (ii) investigates what WA analysis can reveal about lexical organisation and retrieval in bilinguals’ two languages; its minoritised language context means that Welsh speakers are bilingual with English. Two complementary datasets are used. The first comprises responses to 900 Welsh cues from 85 expert users of Welsh, and forms the basis of the first Welsh language WA norms list. The second is bilingual, comprising responses from 85 Welsh speakers and learners to two lists of 100 cue words, one in Welsh and one in English. Language‐specific methodological challenges emerge, including management of mutated word forms, diacritics, and orthographic variation. Decisions relating to these, as the first dataset was converted into a norms list (now informing Welsh language teaching materials), are documented. Language‐specific features that facilitate understanding of WA processes, such as grammatical mutation and inflection, are also reported. Bilingual data associations were categorised to obtain ‘profiles’ for each dataset. Systematic differences between the profiles for each task (Welsh and English) were identified. A pairwise comparison of profiles revealed that while individuals' profiles are distinct from each other, their own profiles are similar across each of their two languages; this closeness is most pronounced in expert users of Welsh.
The article is devoted to the study of the functioning of youth slang in modern Polish and to clarifying the relationship between the linguistic norm and variation within this dynamic lexical subsystem. Youth slang is considered an important sociolinguistic phenomenon that reflects the linguistic creativity of the younger generation, its values, communicative needs, and aspirations for linguistic identity. Particular attention is paid to the sources of the formation of slang units, among which foreign borrowings, semantic transformations, word-formation models, as well as the influence of the digital environment and online communication play a significant role. The article analyzes the main tendencies in the development of youth slang in the Polish-speaking environment and outlines its stylistic, functional, and pragmatic characteristics. It has been determined that youth slang is formed under the influence of sociocultural factors, mass culture, media, and interlingual contacts. It has been established that slang performs not only expressive and identificational functions but also contributes to the formation of group solidarity, informal communication, and distancing from the official linguistic norm. At the same time, it is emphasized that the boundary between normative vocabulary and slang units is variable: certain elements of youth speech gradually integrate into broader language use, spread beyond the youth environment, and may acquire the status of generally accepted lexical items. The methodological basis of the study includes descriptive, comparative, and contextual methods of linguistic analysis. As a result of the research, the key mechanisms of variation of slang units, their structural and semantic features, as well as their role in contemporary linguistic processes have been identified. It is concluded that youth slang is an important factor in language dynamics that reflects sociocultural changes and contributes to the continuous renewal of the lexical system of modern Polish.